Semi-automated extraction of details corresponding to narratological fabula from a corpus of narrative interviews on a single event provides decontextualized building blocks for transversal, or cross-document, narratives. With information extracted from 503 World Trade Center Task Force Interviews comprising 12,000 pages of testimony and novel visualization techniques, this article proposes a computational method for the emergence of narratives that cross beyond the boundaries of one interview. These assembled narratives, in cases like that of Chief Ganci, can document those who did not survive to tell their own story.
Identifying similar narrative sections across longer documents would help identify key events within a corpus, enrich understanding of those events, provide a mechanism for organizing corpora according to their event content, and allow for bottom-up testing of theories of narrative. This paper proposes an automated method for narrative alignment across large textual corpora using techniques from natural language processing and similarity-based image segmentation. This method proceeds by segmenting each document into a series of events, constructs sequences of abstracted representations of those events, compares pairs of sequences to generate image matrices, segments the images, identifies similar segments to discover commonly occurring narrative units, and, finally, returns the source sentences to make the clusters of narrative similarity readable. Preliminary tests of elements of this method were conducted on a small heterogeneous corpus (<; 100 documents) and a moderate heterogeneous corpus (10k documents). Further implementation as described in this position paper is necessary to scale to the full 251k document corpus from which the moderate corpus was drawn.
Automated cross-document comparison of narrative facilitates co-reference and event similarity identification in the retellings of stories from different perspectives. With attention to these outcomes, we introduce a method for the unsupervised generation and comparison of graph representations of narrative texts. Composed of the entity-entity relations that appear in the events of a narrative, these graphs are represented by adjacency matrices populated with text extracted using various natural language processing tools. Graph similarity analysis techniques are then used to measure the similarity of events and the similarity of character function between stories. Designed as an automated process, our first application of this method is against a test corpus of 10 variations of the Aarne-Thompson type 333 story, Little Red Riding Hood. Preliminary experiments correctly co-referenced differently named entities from story variations and indicated the relative similarity of events in different iterations of the tale despite their order differences. Though promising, this work in progress also indicated some incorrect correlations between dissimilar entities.
Analyzing the relationship between location and time in a spatio-temporal data is not trivial. It is even more challenging if the data contains uncertainty. In this paper, we present a new method that visualizes spatio-temporal data with uncertainty. This method is an extension of our 2D visualization technique called Storygraph, and it handles two types of data uncertainty: (1) the spatial and temporal uncertainty about an event; (2) the spatial and temporal uncertainty between two events. We applied this method to ac ase study that involves data extracted from witness testimonies and field reports containing uncertainties inherent to natural language.
Distributed Denial-Of-Service (DDoS) is a common network attack where multiple computers attempt to disable a single system with overwhelming network traffic. Various data visualization methods have been developed to help explain, analyze, and deal with DDoS attacks. However, most of the existing visualization methods do not effectively present the temporal aspect of the DDoS attack data. In this paper, we present a novel DDoS visualization technique, NetTimeView, that applies spatio-temporal data visualization to DDoS data. This technique integrates network traffic data and temporal data in a single view. Its multi-layered visualization technique is able to handle very large data sets with efficient use of visualization space. This tool is particularly useful for system administrators and network security analysts to conduct network forensic analysis. We demonstrate our method with a case study of a large DDoS data set.
Visual analytic techniques are useful for studying patterns and relationships hidden in very large data sets. However, a common problem in visualizing large data sets is visual cluttering. When data get large, visual elements often get crowded, making it difficult for viewers to conduct analysis. In this paper, we present a novel method that reduces visual cluttering in spatial-temporal data visualization. This visualization technique consists of two vertically parallel axes to represent location and a horizontal axis to represent time. Locations are represented as location lines connecting data points on the two vertical axes. As the number of unique locations increase, the number of location lines increase as well, resulting in visual clutter. Our solution is to introduce a new edge bundling technique that reduces the number of crossed lines. We discuss the constraints and simplification of the edge bundling idea and demonstrate this method using two large data sets.
Analysis of spatio-temporal data often involves correlating different events in time and location to uncover relationships between them. It is also desirable to identify different patterns in the data. Visualizing time and space in the same chart is not trivial. Common methods includes plotting the latitude, longitude and time as three dimensions of a 3D chart. Drawbacks of these 3D charts include not being able to scale well due to cluttering, occlusion and difficulty to track time in case of clustered events. In this paper we present a novel 2D visualization technique called Storygraph which provides an integrated view of time and location to address these issues. We also present storylines based on Storygraph which show movement of the actors over time. Lastly, we present case studies to show the applications of Storygraph.
Unidentified victims, perpetrators, and details of human rights violations are camouflaged by the scale of collections of witness reports. This project responds to that problem by developing an approach for identifying the appearance of often unnamed individuals across a corpus by integrating semi-supervised machine-learning based phrase mining, network visualization, and an event model from human rights reporting known as “Who Did What to Whom.” This natural language processing method facilitates the retrieval of cross-document narratives of victims -and perpetrators -of population level events.
Free/Libre and Open source software are generally maintained by a group of developers contributing to the software voluntarily without the presence of any governing institution. Online collaboration platforms and sub-version systems allow developers from all over the world with the necessary skills to contribute to the software. Recently, there has been much interest in analyzing the developers' geographic location since it also serves as a socio-economic marker. In this paper we present our spatio-temporal visualization technique called Storygraph that shows the developers, developer locations and the frequency of commits based on the commit log in one integrated view. We also apply our techniques to the VCS of Rails, Homebrew and D3.js obtained from GitHub.
A major task of spatio-temporal data analysis is to discover relationships and patterns among spatially and temporally scattered events. A most common analytic method is to plot them on a 3D chart with latitude, longitude and time being the three dimensions. The first drawback of this technique is that it fails to scale well when there are thousands of concentrated events since they suffer from cluttering, occlusion and other limitations of 3D plots. Second, it is hard to track the time component if the events are clustered in a region. To overcome these, we present a novel 2D visualization technique called Storygraph that provides an integrated view of location and time. Based on Storygraph, we also present storylines which show the movement of the characters over time. Finally, we present two case studies to demonstrate the effectiveness of the Storygraph.
Archives of human rights violations reports, by virtue of their poor metadata, basis in natural language, and scale, obscure fine grain analyses of violation event patterns. Cross-document coreference of victim or perpetrator occurrences from across a corpus is challenging, particularly when those mentions relate to different events. These challenges are emblematic of the transition from small scale to big data analysis in the humanities. This paper discusses these issues and proposes a framework to address these challenges so as to explore narrative construction and the formation of collective memory. Though our framework is based on processing human rights violation reports, it can be readily extended to support other big data problems in the humanities.
Order copies of this and other ACL proceedings from: Narratives are at the heart of information sharing. Ever since people began to share their experiences, they have connected them to form narratives. The study of storytelling and the field of literary theory called narratology have developed complex frameworks and models related to various aspects of narrative such as plots structures, narrative embeddings, characters' perspectives, reader response, point of view, narrative voice, narrative goals, and many others. These notions from narratology have been applied mainly in Artificial Intelligence and to model formal semantic approaches to narratives (e.g. Plot Units developed by Lehnert (1981)). In recent years, computational narratology has qualified as an autonomous field of study and research. Narrative has been the focus of a number of workshops and conferences (AAAI Symposia, Interactive Storytelling Conference (ICIDS), Computational Models of Narrative). Furthermore, reference annotation schemes for narratives have been proposed (NarrativeML by Mani (2013)). The majority of the previous work on narratives and narrative structures have mainly focused on the analysis of fictitious texts. However, modern day news reports still reflect this narrative structure, but they have proven difficult for automatic tools to summarise, structure, or connect to other reports. This difficulty is partly rooted in the fact that most text processing tools focus on extracting relatively simple structures from the local lexical environment, and concentrate on the document as a unit or on even smaller units such as sentences or phrases, rather than cross-document connections. However, current information needs demand a move towards multidimensional and distributed representations which take into account the connections between all relevant elements involved in a " story ". Additionally, most work on cross-document temporal processing focuses on linear timelines, i.e. representations of chronologically ordered events in time (for instance, the Event Narrative Event Chains by Chambers (2011), or the SemEval 2015 Task 4: Cross Document TimeLines by Minard et al. (2014)). Storylines, though, are more complex, and must take into account temporal, causal and subjective dimensions. How storylines should be represented and annotated, how they can be extracted automatically, and how they can be evaluated are open research questions in the NLP and AI communities. The workshop aimed to bring together researchers from different communities working on representing and extracting narrative structures in news, a text genre which is highly used in NLP but which has received little attention with respect to narrative structure, representation and …