The rapid progress of Large Language Models (LLMs) has transformed natural language processing and broadened its impact across research and society. Yet, systematic evaluation of these models, especially for languages beyond English, remains limited. "Challenging the Abilities of LAnguage Models in ITAlian" (CALAMITA) is a large-scale collaborative benchmarking initiative for Italian, coordinated under the Italian Association for Computational Linguistics. Unlike existing efforts that focus on leaderboards, CALAMITA foregrounds methodology: it federates more than 80 contributors from academia, industry, and the public sector to design, document, and evaluate a diverse collection of tasks, covering linguistic competence, commonsense reasoning, factual consistency, fairness, summarization, translation, and code generation. Through this process, we not only assembled a benchmark of over 20 tasks and almost 100 subtasks, but also established a centralized evaluation pipeline that supports heterogeneous datasets and metrics. We report results for four open-weight LLMs, highlighting systematic strengths and weaknesses across abilities, as well as challenges in task-specific evaluation. Beyond quantitative results, CALAMITA exposes methodological lessons: the necessity of fine-grained, task-representative metrics, the importance of harmonized pipelines, and the benefits and limitations of broad community engagement. CALAMITA is conceived as a rolling benchmark, enabling continuous integration of new tasks and models. This makes it both a resource – the most comprehensive and diverse benchmark for Italian to date – and a framework for sustainable, community-driven evaluation. We argue that this combination offers a blueprint for other languages and communities seeking inclusive and rigorous LLM evaluation practices.
In this paper we present the development of an ontology-based bilingual (IT-EN) lexicographic resource of Italian clitic Verbal MultiWord Expressions ( VMWEs) to support machine translation. Starting from an analysis of these units and their linguistic features, we examine how Neural Machine Translation ( NMT) handles complex VMWEs and the related translations issues. Finally, we propose a bilingual resource, formalised by means of the OntoLex-Lemon model, which accounts for morphological, syntactic, and semantic features of Italian clitic verbs, in order to enhance automatic translation of VMWEs.
Abstract In this paper we present the development of an ontology-based bilingual (IT-EN) lexicographic resource of Italian clitic Verbal MultiWord Expressions (VMWEs) to support machine translation. Starting from an analysis of these units and their linguistic features, we examine how Neural Machine Translation (NMT) handles complex VMWEs and the related translations issues. Finally, we propose a bilingual resource, formalised by means of the OntoLex-Lemon model, which accounts for morphological, syntactic, and semantic features of Italian clitic verbs, in order to enhance automatic translation of VMWEs.
Terminology translation plays a significant role in domain-specific machine translation. However, some knowledge domains and languages still suffer from the lack of high-quality machine translation results due to the mistranslation of terminology. This is the case in the legal domain and the Arabic language. Most machine translation systems fail in their results to produce the exact equivalence of most legal terms for Arabic into other languages, mainly English and French. This failure highlights the lack of terminology resources related to the legal domain, the unfamiliarity of the legal systems to render the appropriate equivalences and the terminology linguistic characteristics of this type of discourse. This difficulty recalls the need for more legal terminology resources. In fact, even though there are many Arabic legal dictionaries, most of them are not machine-readable, and cannot be used in machine translation or other Natural Language Processing applications. As a pipeline, we first extract our terms using NooJ grammars, and then proceed with the creation of our dictionary using NooJ morpho-syntactic information (part of speech (POS), gender, number, etc.), syntactic information (transitive, intransitive, Naqis, etc.), and the creation of our semantic tags that describe our domain-knowledge terms including legal, Juri-religion, etc., and geoUsage to indicate where a given term is adapted to express a legal practice. Finally, we propose the translation. In this phase, the process relies on consulting many sources, including EUR-Lex, EuroVoc and IATE, to be then validated by our legal expert. Our electronic dictionary should enable the automatic annotation of the majority of legal documents in Arabic.
This short paper describes a collection of BERT models for the archaeology domain. We took existing language specific BERT models in English, German, and Dutch, and further pre-trained them with archaeology specific training data. We then took each of these three archaeology specific models and fine-tuned them for Named Entity Recognition.
Despite the advantages, Linguistic Linked Data (LLD) best practices and principles seem far from being widely adopted. Such a situation can be related to existing challenges in the creation, reusing, and exposing of LLD resources. In this paper, we present the results of a survey which examined users’ perspective and experience in the use and application of LLD principles, to evaluate the impact, prospects, requirements, or challenges encountered in LLD adoption. The survey was organized in several sections to collect information about participants’ background, LLD knowledge, use, development, publishing, and metadata use. The results show that some bounds have to be over-stepped to ensure the penetration of LLD principles in a wider community and fully exploit their potential.
The lack of annotated datasets affects the development of Natural Language Processing applications and heavily impacts the access to textual data, in particular for specific domains and specific languages. In this paper, we propose a methodology to annotate texts concerning domain-specific knowledge, to provide a reliable source of data for the task of Named Entity Recognition (NER) in the domain of archaeology for the Italian laguage. This method integrates syntactic and semantic information from several structured sources to annotate entities' mentions in unstructured texts. Furthermore, we make use of an ontology to label entities with the specific type they refer to. By using a corpus made up of item descriptions from Europeana's Archaeology Collection, we first test our proposed methodology on a mock dataset composed of 1,000 texts. After several steps of improvements, we use the final process to create a complete dataset composed of 5,000 descriptions. The resulting dataset, Named Entities in Archaeological Texts has a total of 41,002 spans of texts annotated with their domain-specific entity classification according to the CIDOC Conceptual Reference Model.
Extracting events from news stories as the aim of several Natural Language Processing (NLP) applications (e.g., question answering, news recommendation, news summarization) is not a trivial task, due to the complexity of natural language and the fact that news reporting is characterized by journalistic style and norms. Those aspects entail scattering an event description over several sentences within one document (or more documents), applying a mechanism of gradual specification of event-related information. This implies a widespread use of co-reference relations among the textual elements, conveying non-linear temporal information. In addition to this, despite the achievement of state-of-the-art results in several tasks, high-quality training datasets for non-English languages are rarely available. This paper presents our preliminary study to develop an annotated Dataset for Italian Crime Event news (DICE). The contribution of the paper are: (1) the creation of a corpus of 10,395 crime news; (2) the annotation schema; (3) a dataset of 10,395 news with automatic annotations; (4) a preliminary manual annotation using the proposed schema of 1000 documents. The first tests on DICE have compared the performance of a manual annotator with that of single-span and multi-span question answering models and shown there is still a gap in the models, especially when dealing with more complex annotation tasks and limited training data. This underscores the importance of investing in the creation of high-quality annotated datasets like DICE, which can provide a solid foundation for training and testing a wide range of NLP models.
Terminological resources (TRs) are indispensable tools for accessing a specialized domain of knowledge. In this article, we propose a methodology for extracting terms and relevant linguistic information, useful for different users, hinging on the nature of some special linguistic structures: appositive constructions. The case study for our proof of concept in this investigation is the domain of Cultural Heritage. For such purpose, we compile a parallel Italian–English domain corpus in the specialized sub-domain of archaeology, for extracting data to be later stored in TRs tailored to three target users: (1) translators and interpreters, (2) technical communicators, and (3) the general public or laypeople.
The Project Archaeo-Term: Initial Results This article aims at describing the objectives, the theoretical and methodological background, the development, and the first results of the Archaeo-Term project of the University of Naples "L'Orientale", Department of Literary, Linguistic and Comparative Studies. The Archaeo-Term project has been developed within the YourTermCULT project promoted by the Terminology Without Borders Project of the Terminology Coordination Unit (TermCoord) of the European Parliament - Directorate-General for Translation (DGT) specifically for collecting terminology in different aspects related to culture. The aim of the Archaeo-Term project is to enhance the access to the archaeological data in several formats and languages. It represents a common effort to contribute to the creation of linguistic and terminological resources for the domain of Cultural Heritage (CH) and, in particular, for the sub-domain of archaeology, which is notably highly complex and fragmented. One of the first results of the Archaeo-Term project is the creation of a multilingual terminological resource for the domain of archaeology, which can be conveniently employed in different Natural Language Processing (NLP) tasks, including Machine Translation (MT). The first version of the Archaeo-Term multilingual terminological resource is available in 5 languages: Italian, English, Spanish, German, and Dutch and is publicly accessible online. With the objective of promoting a common and shared termbase across different languages, the Archaeo-Term terminological resource is addressed not only to a specialized audience such as experts in the field of archaeology but also as terminological support for translators and interpreters during their professional practice, as well as for a more general audience. The terminological resource is the result of an extraction and aggregation process carried out starting from two already existing thesauri: the Italian "Thesaurus per la definizione dei reperti archeologici" developed by the Italian Central Institute for Catalogue and Documentation (Istituto Centrale per il Catalogo e la Documentazione - ICCD) and the multilingual Art and Architecture Thesaurus (AAT) developed by the Getty Research Institute, which is among the most trustworthy and accurate resources in the domain of Cultural Heritage. Taking advantage of the Semantic Web formalisms applied to these terminological resources, we are able to extract and merge information from the aforementioned thesauri using SPARQL queries. Indeed, we run different queries against the SPARQL endpoint to enrich our multilingual terminological resource by extracting useful information about the different terminological entries. The information extracted and merged from these thesauri by means of several consecutive queries is as follows: the equivalent terms in the foreseen languages, the alternative terms and the plural forms, the domains and sub-domains, the definitions of the terms, and their sources. Furthermore, the extraction phase has been followed by an evaluation step aimed at checking missing information, verifying and adjusting possible misalignments among entries, and setting potential future implementations. As an ongoing project, we are also planning to enlarge the terminological resource with equivalent terms in other languages such as French, Swedish, Polish, Russian, and Chinese, with the aim of extending the language coverage also to non- European languages which are usually under-represented and low-resourced. As a first implementation with regards to the first version of the terminological resource, we have currently collected: 1.059 entries in Italian, 1.055 in Spanish, 1.053 in English, 843 in Russian, 600 in Polish, 460 in German, 193 in French, and 82 in Chinese. To conclude, the Archaeo-Term project aims at promoting the creation of high-quality and trustworthy multilingual terminological resources for the domain of archaeology by also collaborating at the same time with institutions, experts in the field of terminology, linguistics, and cultural heritage.
The need for reusable, interoperable, and interlinked linguistic resources in Natural Language Processing downstream tasks has been proved by the increasing efforts to develop standards and metadata suitable to represent several layers of information. Nevertheless, despite these efforts, the achievement of full compatibility for metadata in linguistic resource production is still far from being reached. Access to resources observing these standards is hindered either by (i) lack of or incomplete information, (ii) inconsistent ways of coding their metadata, and (iii) lack of maintenance. In this paper, we offer a quantitative and qualitative analysis of descriptive metadata and resources availability of two main metadata repositories: LOD Cloud and Annohub. Furthermore, we introduce a metadata enrichment, which aims at improving resource information, and a metadata alignment to META-SHARE ontology, suitable for easing the accessibility and interoperability of such resources.
Terminological resources are indispensable tools for accessing a specialized domain of knowledge. Several actors gravitate around a specialized domain and, therefore, the creation of terminological resources for different kind of users is a challenging task. In this paper, we propose a methodology for extracting terms and linguistic information useful for different receivers, hinging on appositive constructions.