Machine coding of conflict event datasets has recently emerged as a time-effective method which can back up predictive models for conflict escalation at national and sub-national level. However, the event record duplication issue, caused by large news coverage of major conflict events, signifi-cantly degrades the accuracy of these datasets and makes them unreliable for micro-analysis of con-flict processes. In this paper, we assess the effectiveness of two automatic approaches for mitigating the event duplication issue. The first approach (Cluster Linking) consists of linking news article clusters across time, prior to event extraction, while the second one (Event Linking) is based on clas-sification and aggregation of related events. The comparative evaluation is performed by measuring the correlation of the output from an automatic event detection system with human-coded conflict events from the ACLED project, spatially aggregated on administrative units. We find out that, while both methods effectively reduce the automatic system’s large outlier event and victim counts (with a slight prevalence of Event Linking), they can only increase the correlation coefficients with human-coded data significantly if coupled with an accurate and fine-grained geocoding module.
This paper reports on an ongoing development of a tool for extracting structured information on events a given target entity participated in from massive collections of textual documents and anchoring these events on a timescale. An overview of the current version of the tool and the underlying timeline extraction process is given. Some evaluation figures that reflect system output quality are provided too. The paper will be accompanied by a live demo of the timeline extraction tool.
The paper reports on exploring various machine learning techniques and a range of textual and meta-data features to train classifiers for linking related event templates automatically extracted from online news. With the best model using textual features only we achieved 94.7% (92.9%) F1 score on GOLD (SILVER) dataset. These figures were further improved to 98.6% (GOLD) and 97% (SILVER) F1 score by adding meta-data features, mainly thanks to the strong discriminatory power of automatically extracted geographical information related to events.
This paper reports on an effort of creating a corpus of structured information on security-related events automatically extracted from on-line news, part of which has been manually curated. The main motivation behind this effort is to provide material to the NLP community working on event extraction that could be used both for training and evaluation purposes.
Any large organisation, be it public or private, monitors the media for information to keep abreast of developments in their field of interest, and usually also to become aware of positive or negative opinions expressed towards them. At least for the written media, computer programs have become very efficient at helping the human analysts significantly in their monitoring task by gathering media reports, analysing them, detecting trends and - in some cases - even to issue early warnings or to make predictions of likely future developments. We present here trend recognition-related functionality of the Europe Media Monitor (EMM) system, which was developed by the European Commission's Joint Research Centre (JRC) for public administrations in the European Union (EU) and beyond. EMM performs large-scale media analysis in up to seventy languages and recognises various types of trends, some of them combining information from news articles written in different languages and from social media posts. EMM also lets users explore the huge amount of multilingual media data through interactive maps and graphs, allowing them to examine the data from various view points and according to multiple criteria. A lot of EMM's functionality is accessibly freely over the internet or via apps for hand-held devices.
Part 1: The Nature of Language 1. Are Humans Unique? 1.1. On Origins 1.2. Rules and Creativity 1.3. Animal Communication and the 'Design Features' of Language 1.4. Genetic Transmission of Language 1.5. Human-Like Language in Higher Primates? 1.6. The Functional Significance of Hockett's Design Features 1.7. Structure and Function in Language 1.8. Saussure's Structuralist Linguistics. Exercises. Bibliography 2. The Data of Linguistics and the Nature of Learning 2.1. Structuralist Linguistics and Behaviourist Psychology 2.2. Objections to a Corpus-Based Approach 2.3. Rules and Intuitions - Mentalist Linguistics 2.4. Objections to Mentalist Linguistics 2.5. Native Language Learning Empiricism v. Rationalism 2.6. External Evidence for Language Innateness 2.7. An Internal Argument for Innateness. Exercises. Bibliography Part 2: The Structure of Language 3. Phonetics 3.1. Primacy of a Spoken Language 3.2. Sound Waves 3.3. Air Vibration 3.4. Voice production 3.5. Respiration and Speech Airstream Mechanism 3.6. Analysis and Classification of Vowels 3.7. Liquids and Fricatives 3.8. Places of Articulation 3.9. Stops and Affricates 3.10. Voicing and Nasalisation 3.11. Suprasegmentals. Exercises. Bibliography 4. Phonology 4.1. Physical Sound and Linguistic Sound 4.2. Contextual Variation of Sound 4.3. Criteria of Analysis 4.4. Daniel Jones and the Phoneme 4.5. Sapir's Psychological Approach 4.6. Discovery Procedures 4.7. Classification of Phonemes 4.8. Distinctive Features 4.9. Rules and Formalism. Exercises. Bibliography 5. Morphology 5.1. The Morpheme as the Basic Unit 5.2. Phonologically Conditioned Morphological Variation 5.3. Boundaries between Morphemes - the Morph 5.4. the Word - Definitional Criteria 5.5. Morphological Classification of Languages 5.6. The Priority Question - Words v. Morphemes 5.7. Lexical Productivity - the Creation of Words 5.8. Approaches to Morphological Description. Exercises. Bibliography 6. Syntax 6.1. The Domain of Syntax 6.2. Representing Constituency: Phrase Structure Grammar 6.3. Justifying Constituency: Empirical Diagnostics 6.4. Subcategorisation Restrictions 6.5. Transformations. Exercises. Bibliography 7. Semantics 7.1. Word-Meaning and Sentence-Meaning 7.2. Semantic Properties and Relations of Words 7.3. Semantic Properties and Relations of Sentences 7.4. Theories of Word-Meaning 7.5. Theories of Sentence-Meaning 7.6. Semantics and Pragmatics. Exercises. Bibliography 8. Rules and Principles in the Theory of Grammar 8.1. Restricting the Base 8.2. Constraining Transformational Rules 8.3. Abstract Principles in Syntax 8.4. Generative Phonology 8.5. Autosegmental Representations 8.6. Template-Based Morphology 8.7. Metrical Structures 8.8. Stress in Syntax. Exercises. Bibliography Part 3: The Use of Language 9. Psycholinguistics 9.1. Linguistics, Psycholinguistics and Cognitive Psychology 9.2. Psychological Reality of Distinctive Features 9.3. Psychological Reality of Constituent-Structure 9.4. Psychological Reality of Deep Structures 9.5. Psychological Reality of Transformational Rules 9.6. Against the Psychological Reality of Transformational Rules 9.7. An Alternative to the Derivational Theory of Complexity 9.8. Semantics and Sentence Memory 9.9. The Psychological Lexicon 9.10 Universal Categories of Thought. Exercises. Bibliography 10. Language Development in Children 10.1. Description and Explanation in Language Acquisitional Research 10.2. Phonological Development 10.3. Early Syntactic Development 10.4. Transformational Rules in Language Development 10.5. Semantic Development: Relational Meanings 10.6. Semantic Development: Referential Meanings 10.7. The Development of Speech-Acts 10.8. Linguistic Environment and Language Learning. Exercises. Bibliography 11. Comparative Linguistics 11.1. The Problem of 'Language' 11.2. Types of Similarity and their Significance 11.3. Universals and Typology of Language 11.4. The Nature of Language Change 11.5. Change and Reconstruction 11.6. Linguistic geography 11.7. Mechanisms of Linguistic Change. Exercises. Bibliography 12. Sociolinguistics 12.1. Language and Socialisation 12.2. Language Varieties 12.3. Class, Codes and Control 12.4. Variable Rules 12.5. Social Variables 12.6. Register 12.7. Community Grammars 12.8. Pidgins and Creoles 12.9. Conclusion. Exercises. Bibliography. Addenda to Bibliographies. Author Index. Subject Index
This chapter presents a number of techniques for multilingual event extraction, the main task is to accurately and efficiently detect key information about security-related events from electronic news media and summarize it in the form of database-like structures. Gathering such information over time is an important task for developing global news surveillance systems, particularly in the context of security threats and mass emergencies. In particular, this chapter describes novel techniques for dealing with specific extraction tasks, including: an event type classification method based on domain-specific inference rules, an approach to event geo-tagging based on utilisation of lexico-semantic patterns, a simple method for cross-lingual event information fusion, and techniques for scoring the relevance rank of automatically extracted facts.
We propose a real-time machine translation system that allows users to select a news category and to translate the related live news articles from Arabic, Czech, Danish, Farsi, French, German, Italian, Polish, Portuguese, Spanish and Turkish into English. The Moses-based system was optimised for the news domain and differs from other available systems in four ways: (1) News items are automatically categorised on the source side, before translation; (2) Named entity translation is optimised by recognising and extracting them on the source side and by re-inserting their translation in the target language, making use of a separate entity repository; (3) News titles are translated with a separate translation system which is optimised for the specific style of news titles; (4) The system was optimised for speed in order to cope with the large volume of daily news articles.
In May 2011, an outbreak of enterohemorrhagic Escherichia coli (EHEC) occurred in northern Germany. The Shiga toxin-producing strain O104:H4 infected several thousand people, frequently leading to haemolytic uremic syndrome (HUS) and gastroenteritis (GI). First reports about the outbreak appeared in the German media on Saturday 21st of May 2011; the media attention rose to high levels in the following two weeks, with up to 2000 articles categorized per day by the automatic threat detection system MedISys (Medical Information System). In this article, we illustrate how MedISys detected the sudden increase in reporting on E. coli on 21st of May and how automatic analysis of the reporting provided epidemic intelligence information to follow the event. Categorization, filtering and clustering allowed identifying different aspects within the unfolding news event, analyzing general media and official sites in parallel.
This article presents a real-time and multilingual news event extraction system developed at the Joint Research Centre of the European Commission. It is capable of accurately and efficiently extracting violent and natural disaster events from online news. In particular, a linguistically relatively lightweight approach is deployed, in which clustered news are heavily exploited at all stages of processing. Furthermore, the technique applied for event extraction assumes the inverted-pyramid style of writing news articles, i.e., the most important parts of the story are placed in the beginning and the least important facts are left toward the end. The article focuses on the system's architecture, real-time news clustering, geo-locating and geocoding clusters, event extraction grammar development, adapting the system to the processing of new languages, cluster-level information fusion, visual event tracking, event extraction accuracy evaluation, and detecting event reporting boundaries in news article streams. This article is an extended version of [20].
An ever-growing amount of information relevant for early detection of certain threats can be extracted from on-line news. This led to an emergence of news mining tools to help analysts to digest the overflow of information and to extract valuable knowledge from on line news sources. This paper gives an overview of the fully operational Real-time News Event Extraction Framework developed for Frontex, the EU Border Agency, to facilitate the process of extracting structured information on border security-related events from on-line news. In particular, a hybrid event extraction system has been constructed, which is applied to the stream of news articles continuously gathered and pre-processed by the Europe Media Monitor - a large-scale multilingual news aggregation engine. The framework consists also of an earth browser, in which events are visualized and an event moderation tool, which allows to access the database of automatically extracted event descriptions and to clean, validate, group, enhance and export them into other knowledge repositories.
Nowadays, many influential security-related facts are reported multiple times by different sources and in different languages. Therefore, in the recent years, the research on advancing event extraction technology shifted from classical single-document extraction toward cross-document information aggregation and fact validation. However, relatively little work has been reported on cross-lingual information fusion in this area. This paper presents the results of some preliminary experiments on deploying cross-lingual information fusion techniques for refining the results of a large-scale multilingual news event extraction system. The first technique is based on fusing the responses of the mono-lingual event extraction systems, whereas the second one uses state-of-the-art machine translation to convert all news articles reporting on a given event into one common language and subsequently applies the corresponding mono-lingual event extraction system on the translated articles. An evaluation of the aforementioned techniques on a news article corpus, whose articles refer to 523 real-world crisis-related events (violent events, man-made and natural disasters), revealed that the descriptions of circa 10% of the events could be refined through fusing the event descriptions returned by the mono-lingual event extraction systems. The overall gain in recall and precision against the best mono-lingual system was 6,4% and 4,8% respectively. The second approach, based on machine translation, turned to perform significantly worse compared to the former technique and the best mono-lingual system (English).
This chapter gives an overview of tools developed for Frontex, the European Agency for the Management of Operational Cooperation at the External Borders of the Member States of the European Union, to facilitate the process of extracting structured information on events related to border security from on-line news articles, with a particular focus on incidents and developments in the context of illegal migration, cross-border crime, and related crisis situations at the EU external borders and in third countries. A hybrid event extraction system has been constructed, which consists of two core event extraction engines, namely, NEXUS, developed at the Joint Research Centre (JRC) of the European Commission and PULS, developed at the University of Helsinki. These systems are applied to the stream of news articles continuously gathered and pre-processed by the Europe Media Monitor (EMM) - a large-scale multilingual news aggregation engine, developed at the JRC. In order to bridge the automated analysis phase with in-depth human analysis phase an event moderation tool has been developed, which allows the user to access the database of automatically extracted event descriptions and to clean, validate, group, enhance, and export them into other knowledge repositories.
Nowadays, many influential facts are reported multiple times by different sources and in different languages. This paper presents the results of an experiment on deploying cross-lingual information fusion techniques for refining the results of a large-scale multilingual news event extraction system. An evaluation on a test corpus consisting of 618 event descriptions which refer to 523 real-world events revealed that the description of circa 10% of the events extracted by the mono-lingual systems could be refined. In particular, an overall gain of 6,4% and 4,8% in recall and precision against the best monolingual system could be obtained respectively.
This paper gives an overview of an ongoing effort to construct tools for automating the process of extracting structured information about border-security related events from on-line news. The paper describes our overall approach to the problem, the system architecture and event information access and moderation.
We describe a methodology for building event extraction systems. The approach is based on multilingual domain-specific grammars and exploits weakly supervised machine learning algorithms for lexical acquisition. We report on the process of adapting an already existing event extraction system for the domain of conflicts and crises to the Portuguese language.
This paper presents a visual analytics approach to explore large news article collections in the domains of polarity and spatial analysis.The exploration is performed on the data collected with Europe Media Monitor (EMM), a system which monitors over 2500 online sources and processes 90,000 articles per day.By analyzing the news feeds, we want to find out which topics are important in different countries and what is the general polarity of the articles within these topics.To assess the polarity of a news article, automatic techniques for polarity analysis are employed and the results are represented using Literature Fingerprinting for visualization.In the spatial description of the news feeds, every article can be represented by two geographic attributes, the news origin and the location of the event itself.In order to assess these spatial properties of news articles, we conducted our geo-analysis, which is able to cope with the size and spatial distribution of the data.Within this application framework, we show opportunities how real-time news feed data can be analyzed efficiently.