We present a novel approach for applying Large Language Models (LLMs) to threat assessment in the context of foreign peacekeeping missions. Building on the PINPOINT project and its use case, the EU Monitoring Mission in Georgia, we combine an interdisciplinary risk-model with OSINT-based media collection and LLM-supported threat extraction. The proposed workflow maps media contents to mission-relevant threats, extracts structured information and applies several additional LLM-based processing steps to improve relevance and grounding. An evaluation of threats extracted from media documents shows high agreement between automatically generated results and human judgment for core aspects such as threat and mission relevance. These results indicate that LLMs provide a promising approach to support analysts in the context of peacekeeping missions.
Entity matching, also known as user identity linkage, is a critical task in data integration. While established techniques primarily focus on large-scale networks, there are several applications where small networks pose challenges due to limited training data and sparsity. This study addresses entity matching in the field of criminology, where small networks are common and the number of known matching nodes is restricted. To support this research, we exploit a multimodal dataset, collected as part of a security-related project, consisting of an intercepted telephone calls network (i.e., ROXSD data) and a network of social forum interactions (i.e., ROXHOOD data) collected in a simulated environment, although following real investigation scenario. To improve accuracy and efficiency, we propose a novel approach for entity matching across these two small networks using node attributes. Existing techniques often merely focus on topology consistency between two networks and overlook valuable information, such as network node attributes, making them vulnerable to structural changes. Inspired by the remarkable success of deep learning, we present UGC-DeepLink, an end-to-end semi-supervised learning framework that leverages user-generated content. UGC-DeepLink encodes network nodes into vector representations, capturing both local and global network structures to align anchor nodes using deep neural networks. A dual learning paradigm and the policy gradient method transfer knowledge and update the linkage. Additionally, node attributes, such as call contents and forum exchanged texts, enhance the ranking of matching nodes. Experimental results on ROXSD and ROXHOOD demonstrate that UGC-DeepLink surpasses baselines and state-of-the-art methods in terms of identity-match ranking. The code and dataset are available at https://github.com/erichoang/UGC-DeepLink.
AbstractThis chapter provides an in-depth account of current research activities and applications in the field of Speech Technology (ST). It discusses technical, scientific, commercial and societal aspects in various ST sub-fields and relates ST to the wider areas of Natural Language Processing and Artificial Intelligence. Furthermore, it outlines breakthroughs needed, main technology visions and provides an outlook towards 2030 as well as a broad view of how ST may fit into and contribute to a wider vision of Deep Natural Language Understanding and Digital Language Equality in Europe. The chapter integrates the views of several companies and institutions involved in research and commercial application of ST.
AbstractThis chapter provides a comprehensive overview of innovation and the ELG marketplace as core elements for the generation of value and the creation of an active, attractive and vibrant community surrounding the European Language Grid. Innovation is an essential element in making ELG a credible and sustainable undertaking. However, it does not happen by itself nor materialise in a vacuum. Consequently, ELG provides a habitat for various kinds of innovation and a home for the necessary community to put innovation into action. The marketplace is essential for attracting participants supplying and demanding services, resources, components and technologies on a European scale. Innovation and marketplace – as well as the overall business model – are tightly connected and need to be developed and managed in a joint manner. Clearly, this is not a one-off activity, but rather needs to be carried out continuously and extend into the future. ELG is designed and created to promote the excellence and growth of the European LT market, creating new jobs and business opportunities and supporting European digital sovereignty. Encompassing a wide array of technologies and resources for many languages spoken across Europe and in neighbouring regions, it contributes to the Multilingual Digital Single Market as a cross-European driver for innovation.
AbstractWhen preparing the European Language Grid EU project proposal and designing the overall concept of the platform, the need for drawing up a long-term sustainability plan was abundantly evident. Already in the phase of developing the proposal, the centrepiece of the sustainability plan was what we called the “ELG legal entity”, i. e., an independent organisation that would be able to take over operations, maintenace, extension and governance of the European Language Grid platform as well as managing and helping to coordinate its community. This chapter describes our current state of planning with regard to this legal entity. It explains the different options discussed and it presents the different products specified, which can be offered by the legal entity in the medium to long run. We also describe which legal form the organisation will take and how it will ensure the sustainability of ELG.
Criminal investigations contain sensitive and confidential material and are nonpublic by nature. Access to investigation data is very limited and restricted to only selected groups of individuals. Even for research purposes, data typically cannot be accessed freely. Within criminal investigations, data is still processed manually to a large extent. Solutions provided for automation of this processing — or even of individual processing steps — can be assumed to have a significant impact on the work of Law Enforcement Agencies (LEAs). Automation may effectively be key to handle large and complex amounts of data in an efficient manner under the typical operating conditions of LEAs. This paper introduces the ROXANNE Simulated Dataset (ROXSD), a dataset with unique properties prepared by the ROXANNE Project with assistance from several LEAs, to facilitate the development and evaluation of novel tools and technologies for criminal investigations. ROXSD consists of a set of simulated intercepted telephone conversations in a variety of languages. The story follows a realistic setting and includes the conditions and constraints of a real investigation. The network topology corresponding to the conversations was created by partner LEAs to reflect various typical organized crime groups. Conversations have been transcribed carefully and annotated in the original language and in English. The dataset is expected to provide a sound basis for further research and is available to download for researchers under signed agreement.
Fake news and misinformation is a widespread phenomenon these days, affecting social media, alternative and traditional media. In a climate of increasing polarization and perceived societal injustice, the topic of migration is one domain that is frequently the target of fake news, addressing both migrants and citizens in host countries. The problem is inherently a multi-lingual and multi-modal one in that it involves information in an array of languages, material in textual, visual and auditory form and often involves communication in a language which may be unfamiliar to recipients or which these recipients only may have basic knowledge of. We argue that semi-automatic approaches, empowering users to gain a clearer picture and base their decisions on sound information, are needed to counter the problem of misinformation. In order to deal with the scale of the problem, such approaches involve a variety of technologies from the field of Artificial Intelligence (AI). In this paper we identify a number of challenges related to implementing approaches for the detection of fake news in the context of migration. These include collecting multi-lingual and multi-modal datasets related to the migration domain, providing explanations of AI tools used in verification to both media professionals and consumers. Further efforts in truly collaborative AI will be needed.
Europe is a multilingual society, in which dozens of languages are spoken. The only option to enable and to benefit from multilingualism is through Language Technologies (LT), i.e., Natural Language Processing and Speech Technologies. We describe the European Language Grid (ELG), which is targeted to evolve into the primary platform and marketplace for LT in Europe by providing one umbrella platform for the European LT landscape, including research and industry, enabling all stakeholders to upload, share and distribute their services, products and resources. At the end of our EU project, which will establish a legal entity in 2022, the ELG will provide access to approx. 1300 services for all European languages as well as thousands of data sets.
Our perception of the situation in a country or a region is strongly influenced by the reflection of this situation in mass and social media channels. This effect is even more pronounced for geographically and culturally distant regions, for which no firsthand experience is available. To avoid information overload, news outlets typically filter the available news from foreign countries based on the expected interest of the target audiences. Such filtering imposes an inherent bias in the reporting and can create a distorted perception of a region among the consumers of news of other regions. This might lead to misunderstandings between countries and unsubstantiated political and individual decisions (e.g., in the context of migration). In this article, we systematically analyze the bias created in news reports. We consider Europe, or more precisely the European Union (EU) as our zone of concern, and examine its image in the media (news outlets) of other regions, Europe(NON-EU), Africa, Asia, Middle-East, America, and Oceania. An analysis of the year 2018 (January–December 2018) of news published in those regions reveals marked differences in the editorial policies and presented narrative when dealing with EU-related news. We observe a significant variation in the sentiment polarity of the reported EU-related stories between the European and other regional news outlets. We further analyze the polarity variation among different subregions of large geographical areas, such as Africa, Asia, and America. We observe a contrasting difference in their editorial policies. This trend also holds for news related to different topics, such as politics, business, economy, health, and international relation.
With 24 official EU and many additional languages, multilingualism in Europe and an inclusive Digital Single Market can only be enabled through Language Technologies (LTs). European LT business is dominated by hundreds of SMEs and a few large players. Many are world-class, with technologies that outperform the global players. However, European LT business is also fragmented, by nation states, languages, verticals and sectors, significantly holding back its impact. The European Language Grid (ELG) project addresses this fragmentation by establishing the ELG as the primary platform for LT in Europe. The ELG is a scalable cloud platform, providing, in an easy-to-integrate way, access to hundreds of commercial and non-commercial LTs for all European languages, including running tools and services as well as data sets and resources. Once fully operational, it will enable the commercial and non-commercial European LT community to deposit and upload their technologies and data sets into the ELG, to deploy them through the grid, and to connect with other resources. The ELG will boost the Multilingual Digital Single Market towards a thriving European LT community, creating new jobs and opportunities. Furthermore, the ELG project organises two open calls for up to 20 pilot projects. It also sets up 32 National Competence Centres (NCCs) and the European LT Council (LTC) for outreach and coordination purposes.
Verhetzung, Hasskriminalität, bewusste Falschinformation und Verschwörungstheorien auf Social Media stellen Justiz, Kriminalforschung und damit die Gesellschaft vor nahezu unüberblickbare Herausforderungen. Wann sprechen wir noch von freier Meinungsäußerung bzw. ab wann stellt sich die Frage nach der strafrechtlichen Relevanz? Im Rahmen einer semantischen automatisierten Textanalyse und einer konventionellen Inhaltsanalyse wurde dieser Frage nachgegangen und zur Annäherung ein Klassifizierungsschema zur näheren Charakterisierung von strafrechtlich relevanten Aussagen auf Social Media entwickelt.
The current scientific and technological landscape is characterised by the increasing availability of data resources and processing tools and services. In this setting, metadata have emerged as a key factor facilitating management, sharing and usage of such digital assets. In this paper we present ELG-SHARE, a rich metadata schema catering for the description of Language Resources and Technologies (processing and generation services and tools, models, corpora, term lists, etc.), as well as related entities (e.g., organizations, projects, supporting documents, etc.). The schema powers the European Language Grid platform that aims to be the primary hub and marketplace for industry-relevant Language Technology in Europe. ELG-SHARE has been based on various metadata schemas, vocabularies, and ontologies, as well as related recommendations and guidelines.
Die Arbeit widmet sich der Entwicklung eines (konzeptionellen) Frameworks und Modells, das generisch genug ist, um die grose Vielfalt von Datenquellen und Akteuren darzustellen, die typischerweise bei solchen Vorfallen anzutreffen sind, wobei insbesondere die Verbindungen zwischen ihnen berucksichtigt werden. Die Darstellung all dieser Elemente innerhalb eines einzigen Modells bietet den Vorteil, Gemeinsamkeiten zu erkennen und zu nutzen, um schnell auf neu entstehende Plattformen reagieren und Daten effizient und umfassend verarbeiten zu konnen. Die Entwicklung des Modells erfolgt inkrementell; es wird dabei untersucht, welche generischen Elemente und Attribute existieren, welche Verbindungen zwischen ihnen bestehen und wie die einzelnen Medien und Plattformen dargestellt werden konnen. Das Modell ermoglicht es, die verschiedenen Arten von Elementen aus relevanten Plattformen und Medien, welche in fruheren Arbeiten, durch Gesprache mit Hilfsorganisationen und Betrachtung des Stand-der-Technik bestimmt wurden, in einem medienubergreifenden Ansatz zu kombinieren und das daraus resultierende Framework fur die Analyse der Kommunikation im Katastrophenfall anzuwenden. Daruber hinaus bietet es eine Referenz, anhand derer sich neue Technologien und Tools vergleichen lassen und welche zu Planungs- und Managementaktivitaten herangezogen werden kann.
Multilingualism is a cultural cornerstone of Europe and firmly anchored in the European treaties including full language equality. However, language barriers impacting business, cross-lingual and cross-cultural communication are still omnipresent. Language Technologies (LTs) are a powerful means to break down these barriers. While the last decade has seen various initiatives that created a multitude of approaches and technologies tailored to Europe's specific needs, there is still an immense level of fragmentation. At the same time, AI has become an increasingly important concept in the European Information and Communication Technology area. For a few years now, AI, including many opportunities, synergies but also misconceptions, has been overshadowing every other topic. We present an overview of the European LT landscape, describing funding programmes, activities, actions and challenges in the different countries with regard to LT, including the current state of play in industry and the LT market. We present a brief overview of the main LT-related activities on the EU level in the last ten years and develop strategic guidance with regard to four key dimensions.
In today’s attention-driven news economy, rapid changes of topics and events go hand in hand with rapid changes of vocabulary and of language use. ASR systems aimed at transcribing contents pertaining to this fluid media landscape need to keep upto-date in a continuous and dynamic manner. Static models, potentially created a long time ago, are hopelessly outdated within a short period of time. The frequent changes in vocabulary and wording need to be reflected in the models employed for optimal performance of transcription if one does not want to risk falling behind. In this demonstration paper we present the audio processing capabilities of the SAIL LABS Media Mining Indexer, and the CAVA Framework allowing semi-automatic and periodic updates of the ASR vocabulary and language model from relevant and new data.
In 2004, Information Systems for Crisis Response and Management (ISCRAM) was a new area of research. Pioneering researchers from different continents and disciplines found fellowship at the first ISCRAM workshop. Around the same time, the use of social media in crises was first recognized in academia. In 2018, the 15th ISCRAM conference will take place, which gives us the possibility to look back on what has already been achieved with regard to IT support in crises using social media. With this article, we examine trends and developments with a specific focus on social media. We analyzed all papers published at previous ISCRAMs (n=1339). Our analysis shows that various platforms, the use of language and coverage of different types of disasters follow certain trends – most noticeably a dominance of Twitter, English and crises with large impacts such as hurricanes or earthquakes can be seen.
Gerald Quirchmayr合作论文数Institute for Computer Science and Business Informatics;University of Vienna13