
The landscape of Conversational Artificial Intelligence has become extremely varied, in the last years, as the capability of using language in machines has become the main showcase for generative applications. From a language philosophy point of view, however, the reason why Generative AI produces linguistic content is not aligned with the one motivating humans. In this paper, we provide a multidisciplinary integrative literature review spanning linguistics, cognitive science and computer science to propose a systematic way of organising this knowledge around raison d’exprimer: the reason why a machine should use language. We present both a horizontal model describing how different aspects of communication blend into each other, the hypertriangle of communication, and a vertical model providing a common axis to align multiple disciplines involved in Conversational AI: the illocutionary gradient. By showing how concepts belonging to multiple disciplines align themselves in these models, we provide an extensible theoretical tool to study conceptual alignments in different fields.
The automatic segmentation of raw text into individual sentences, known as sentence splitting or sentence segmentation, is a fundamental task in text processing. Although it is often considered to be solved in standard domains such as news articles and Wikipedia pages, the performance of the system can vary significantly between different textual genres. This study evaluates eight sentence splitting tools employing rule-based, supervised, semi-supervised, and unsupervised approaches, and additionally tests two Large Language Models in a zero-shot setting, on a corpus of 19th-century Italian novels, namely “I Promessi Sposi”, “I Malavoglia”, “Le avventure di Pinocchio”, and “Cuore”. In addition, we train new sentence splitting models using the Stanza pipeline, creating individual models for each novel as well as a combined model trained on all available data. This work aims to highlight that, although literary texts have received relatively little attention in sentence segmentation research, they offer a rich and promising intersection between Natural Language Processing, Italian linguistics, and the Digital Humanities.1
Transformer-based Pre-trained Language Models (PLMs) for Ancient Greek are rapidly advancing in both number and performance. Among the opportunities enabled by these models is semantic statement retrieval, but significant efforts to refine the technique and make it accessible to the scholarly community are still in their early stages. This survey reviews advancements in computational approaches to Ancient Greek semantics, describing the transition from traditional vector-based models to state-of-the-art transformer-based architectures, with a particular focus on their application to statement retrieval. It examines advantages and limitations of applying transformer-based PLMs to tasks such as the study of semantic change, lexical analysis, and statement retrieval. Although these models provide innovative methodologies for addressing semantics in the analysis of Ancient Greek sources, challenges such as limited training data, the absence of shared evaluation standards, and the "black box" nature of deep learning remain significant obstacles. The paper highlights the potential of these technologies to enhance digital scholarship on ancient texts and calls for further refinement and benchmarking to overcome existing limitations.
Voice Activity Detection (VAD) refers to the task of identifying human speech in noisy settings, playing a crucial role in fields like speech recognition and audio surveillance. However, most VAD research has predominantly focused on English, leaving other languages — such as Italian —underexplored. This study aims to evaluate and improve VAD systems for Italian speech, with the ultimate goal of enhancing the speech segmentation component of the Digital Linguistic Biomarkers (DLBs) extraction pipeline for early mental disorder screening. We experimented with multiple VAD systems and proposed a novel ensemble approach that demonstrates improved speech event detection performance. This advancement provides a robust foundation for more accurate early detection of mental health conditions using DLBs in the Italian language.
In this paper, we present a benchmark of texts manually annotated with gustatory information, following a FrameNet-like approach previously applied to olfactory language and here adapted to capture taste-related events. We explore the benchmark to illustrate the possible insights this approach can offer, focusing in particular on the expression of emotional valence across different textual genres. Building on this resource, we train a supervised system for the automatic extraction of gustatory information from both historical and contemporary texts. The system is then applied to a variety of corpora, and we provide a publicly available notebook for exploring the extracted data, along with an analysis of the system’s output in the literary domain.
This paper investigates how human annotators and Large Language Models (LLMs) assign and justify semantic similarity judgments in a Semantic Textual Similarity (STS) task. To this end, we present a new version of SimilEx, the first Italian dataset containing human similarity judgments and natural language explanations for sentence pairs, extended with LLM-generated scores and explanations, enabling a direct comparison between human and LLM behaviour under parallel annotation conditions. Within this framework, we examine the extent to which humans and LLMs align in their perception of sentence similarity. We explore this question from multiple perspectives, including the relationship between sentence-level stylistic features and similarity scores, the consistency of judgments across annotator types, and the alignment of human and LLM explanations. Our findings show that LLMs tend to express more moderate judgments than humans, resulting in higher agreement. At the same time, stylistic features of the evaluated sentences are related to similarity judgments in both groups. As for explanations, humans typically produce shorter, often nominal constructions, reflecting more individually driven strategies for justifying similarity judgments, whereas LLMs generate more canonical sentence structures whose content is also more consistent across models, suggesting that justification is a more subjective process for humans.
The rapid progress of Large Language Models (LLMs) has transformed natural language processing and broadened its impact across research and society. Yet, systematic evaluation of these models, especially for languages beyond English, remains limited. "Challenging the Abilities of LAnguage Models in ITAlian" (CALAMITA) is a large-scale collaborative benchmarking initiative for Italian, coordinated under the Italian Association for Computational Linguistics. Unlike existing efforts that focus on leaderboards, CALAMITA foregrounds methodology: it federates more than 80 contributors from academia, industry, and the public sector to design, document, and evaluate a diverse collection of tasks, covering linguistic competence, commonsense reasoning, factual consistency, fairness, summarization, translation, and code generation. Through this process, we not only assembled a benchmark of over 20 tasks and almost 100 subtasks, but also established a centralized evaluation pipeline that supports heterogeneous datasets and metrics. We report results for four open-weight LLMs, highlighting systematic strengths and weaknesses across abilities, as well as challenges in task-specific evaluation. Beyond quantitative results, CALAMITA exposes methodological lessons: the necessity of fine-grained, task-representative metrics, the importance of harmonized pipelines, and the benefits and limitations of broad community engagement. CALAMITA is conceived as a rolling benchmark, enabling continuous integration of new tasks and models. This makes it both a resource – the most comprehensive and diverse benchmark for Italian to date – and a framework for sustainable, community-driven evaluation. We argue that this combination offers a blueprint for other languages and communities seeking inclusive and rigorous LLM evaluation practices.
Over the past decade, Computational Linguistics (CL) and Natural Language Processing (NLP) have evolved rapidly, especially with the advent of Transformer-based Large Language Models (LLMs). This shift has transformed research goals and priorities, from Lexical and Semantic Resources to Language Modelling and Multimodality. In this study, we track the research trends of the Italian CL and NLP community through an analysis of the contributions to CLiC-it, arguably the leading Italian conference in the field. We compile the proceedings from the first 10 editions of the CLiC-it conference (from 2014 to 2024) into the CLiC-it Corpus, providing a comprehensive analysis of both its metadata, including author provenance, gender, affiliations, and more, as well as the content of the papers themselves, which address various topics. Our goal is to provide the Italian and international research communities with valuable insights into emerging trends and key developments over time, supporting informed decisions and future directions in the field.
The formal study of argumentation-based dialogue lacks a comprehensive reference framework, particularly from a linguistic perspective. This work addresses part of this gap by analysing whether the relationship between the semantic-syntactic structure of questions and their pragmatic features could increase the utility of the corresponding answers. A preliminary experiment was carried out to identify which forms of information-seeking request best prompt useful human responses, thereby efficiently solving decision problems. Four question types were tested alongside alternative formulations to examine the influence of pragmatic features on answer quality and cognitive effort. Results suggest that pragmatic features can affect the informativeness of answers, with polar questions tending to elicit informative or over-informative responses, while content questions displayed greater variability. Additionally, questions containing polarity items appeared to increase cognitive load, as reflected in response patterns. While these findings are necessarily tentative due to the exploratory nature of the study, they offer promising indications for the design of linguistically grounded, argumentation-based dialogue systems and point to several avenues for further research, including broader experimental designs and an expanded set of pragmatic variables.
This paper outlines the evolving interplay between Linguistics and Computational Linguistics, aiming to map the current state of their interactions and to identify areas where deeper integration could drive significant advancements in both areas. Since the early days of Computational Linguistics as an autonomous discipline, the synergy has developed in parallel with progress in both computational methods and linguistic theory. Computational modeling of language offers a powerful framework to investigate core questions of linguistics, from how language works and is acquired, to how it changes across time, space, communicative situations, and domains.Despite this potential, the capabilities of state-of-the-art computational methods remain only partially exploited within linguistic research, leaving a gap between advances in Natural Language Processing and the needs of linguistics. This paper seeks to examine the current landscape of this synergy, its scientific and practical implications, and the challenges that must be addressed to fully harness its potential. A pilot study is presented to illustrate how linguistic resources and computational modeling can provide answers to long-standing research questions and, at the same time, open up new avenues for investigating open issues in language typology.
This paper discusses the results of various experiments assessing the morphosyntactic and semantic competence in Italian of four very large language models (vLLMs): davinci (GPT-3/ChatGPT), davinci-002, davinci-003 (both GPT-3.5 models) and gpt-4-1106-preview (GPT-4). We evaluated these models on (i) acceptability, (ii) complexity, and (iii) coherence judgments using 7-point Likert scales and on (iv) syntactic development through a forced choice task. The test sets were drawn from shared NLP tasks and standard linguistic assessments. The results suggest that, although fine-tuned transformers outperform all GPT models, GPT-4 represents a significant improvement over third-generation GPT models. According to our tests, even if GPT-4 and fine-tuned transformers cannot be considered descriptively or explanatorily adequate, they nonetheless pose a challenge to the poverty of the stimulus hypothesis. The "theory" expressed by GPT models is not linguistically intelligible in any relevant sense, and their training data is orders of magnitude larger than the primary linguistic input available to children. Nevertheless, GPT-4 captures certain generalizations, such as the constraints blocking the insertion of an overt resumptive clitic in specific gap positions, that are arguably unlearnable from just primary positive data.
This article presents a study on annotating explicit discourse relations in Italian student essays, comparing human annotations with outputs from generative large language models and examining their alignment with theoretical models of textuality. We review prior work on automatic discourse relation annotation in Italian, highlighting limitations in language coverage, especially in out-of-domain scenarios and how these have been addressed. Our experiments explore the use of generative models to mitigate the scarcity of domain-specific training data, while assessing their ability to reflect the intended theoretical framework. We evaluate two generative models in detecting connectives and classifying their senses, comparing results to human annotation. For our evaluation sample, we use a string-matching algorithm combined with a rule-based approach to pre-annotate essays with possible connective forms and their senses, based on their presence in the Lexicon of Italian Connectives (LICO). These annotations were manually corrected by two expert annotators, resulting in a publicly available evaluation sample. The study raises significant theoretical questions about the definition of connectives, its relationship to text segmentation and the challenges both human and machines face when annotating discourse relations. Our findings show how computational approaches can shed light on linguistic theories and, vice versa, how linguistic theories can guide the application of computational resources.
Automatic translation to and from Italian Sign Language (LIS) requires the development of computational models, such as avatars, capable of accurately reproducing both manual and non-manual articulators of signed discourse. This, in turn, demands the creation of machine-processable and linguistically robust data collections, built through the segmentation, transcription and systematic categorization of signs, to capture their internal structure and relational dynamics. Such a framework should reflect the multilinear organization of LIS, which poses several challenges. These include the visual-gestural and simultaneous nature of LIS, the absence of a standardized written form and the scarcity of available resources. Key challenges arise at multiple levels, including the very development of LIS resources, the identification of suitable tools for capturing signed data and the lack of a standardized coding system for signed languages. These aspects were addressed in the development of the MultiMedaLIS Dataset (MULTImodal MEDicAl LIS Dataset), a preliminary dataset in the medical domain, collected using multimodal capturing tools. Annotation, performed using ELAN, followed the principles of simplicity and readability by employing multilayered labelling in both Italian and English, along with a dedicated annotation system for signed languages. In this way, the Dataset is accessible to both signers and non-signers, currently serving as a resource for linguistic analyses, as well as for training algorithms for automatic sign recognition.
The linguistic competence of Large Language Models (LLMs) has been the focus of extensive investigation in recent years. Yet, the syntax-semantics interface remains a relatively understudied aspect of LLMs’ linguistic abilities. This study aims to address this gap by focusing on the Instrumental role in Italian. In this language, Instruments can always be syntactically omitted, yet they remain semantically present, as they are recoverable either from the verb meaning alone (when the verb is presented in isolation) or from the interaction between the verb meaning and that of its internal argument (when the verb appears within a syntactic context).To assess the ability of LLMs to semantically determine the most appropriate Instrument(s) from the verb meaning, we conducted two experiments based on psycholinguistically inspired tasks, comparing the performance of GePpeTto and Minerva models (350M, 1B, 3B and 7B) to that of Italian speakers. In the first experiment, verbs were presented in isolation, while in the second, they were presented within a syntactic context. Our findings indicate that the performance of LLMs is influenced by the semantic selectivity of verbs, the presence or absence of a clausal context and model characteristics.
We introduce the ’e-RTE-3-it’ dataset, an enriched version of the Italian RTE-3 dataset, where each text-hypothesis pair, in addition to the ’entailment’, ’contradiction’, or ’neutrality’ label, has been combined with an explanation for the relation. Moreover, the dataset includes the level of confidence with which the annotators wrote the explanation as well as an optional alternative label, along with its explanation, which the annotators could express when they did not agree with the original label. This offers the opportunity to analyse cases of uncertainty in annotation and take into account different perspectives on natural language understanding and generation.
In this paper, we describe the creation of a treebank for Dante’s Comedy in Universal Dependencies, the first syntactically annotated text for Old Italian following a dependency-based paradigm. We detail the phase of treebanking the first part of the Comedy, the Inferno, and we discuss some annotation issues, specifically ellipses and comparative structures. Then, we perform an evaluation of automated dependency parsing with models trained on the currently available annotated portion of the text.1
Educational crosswords offer numerous benefits for students, including increased engagement, improved understanding, critical thinking, and memory retention. Creating high-quality educational crosswords can be challenging, but recent advances in natural language processing and machine learning have made it possible to use language models to generate nice wordplays. The exploitation of cutting-edge language models like GPT3-DaVinci, GPT3-Curie, GPT3-Babbage, GPT3-Ada, and BERT-uncased has led to the development of a comprehensive system for generating and verifying crossword clues. A large dataset of clue-answer pairs was compiled to fine-tune the models in a supervised manner to generate original and challenging clues from a given keyword. On the other hand, for generating crossword clues from a given text, Zero/Few-shot learning techniques were used to extract clues from the input text, adding variety and creativity to the puzzles. We employed the fine-tuned model to generate data and labeled the acceptability of clue-answer parts with human supervision. To ensure quality, we developed a classifier by fine-tuning existing language models on the labeled dataset. Conversely, to assess the quality of clues generated from the given text using zero/few-shot learning, we employed a zero-shot learning approach to check the quality of generated clues. The results of the evaluation have been very promising, demonstrating the effectiveness of the approach in creating high-standard educational crosswords that offer students engaging and rewarding learning experiences. In this new extended version of (Zeinalipour, Iaquinta, Zanollo, et al. 2023) we also propose a linguistic analysis of crossword clues developed from a syntactic perspective. This preliminary analysis provides insights into how clues are derived from complete sentences and what structures, if any, are preferred. The present linguistic study represents a useful support, in future research, for the generation of personalized puzzles in educational and medical environments.
Sentiment analysis is the field of study that analyzes people’s opinions and sentiments towards entities such as products, services and organizations. Brand reputation analysis, competitive intelligence and social network analysis are just a few areas that can benefit from sentiment analysis. Most studies on sentiment analysis have only focused on domains like product reviews and social network content, leaving sentiment inference in the news domain under-investigated. In this work, we use a case study of a company specialized in the analysis of brand reputation to evaluate machine learning models for sentiment analysis on multilingual news articles. Several models were tested, including traditional machine learning models like KNN, and transformer-based models like BERT, Llama and GPT. The implemented models were evaluated on a dataset of Italian, German and Ladin news articles annotated with their sentiment polarity. Overall, our experiments show state-of-the-art results and confirm the outcomes of previous studies, i.e. that sentiment analysis of news articles remains a complex task. Machine learning systems can support manual annotators in accelerating the annotation process. Our findings can provide a benchmark for researchers in natural language processing when performing sentiment analysis of news articles.
In this paper, we address the problem of automatic misogynous meme recognition by dealing with potentially biased elements that could lead to unfair models. In particular, a bias estimation technique is used to identify those textual and visual elements that unintendedly affect the model prediction, and a few bias mitigation methods are proposed, investigating two different types of debiasing strategies, i.e., at training time and at inference time. The proposed approaches achieve remarkable results both in terms of prediction and generalization capabilities.