Conversation topics may vary in abstractness. This might impact the effort required by speakers to reach a common ground and, ultimately, an interactive alignment. In fact, people typically feel less confident with abstract concepts and single-words rating studies suggest abstract concepts are more associated with social interactions than concrete concepts—hence suggesting increasing levels of abstractness enhance inner and mutual monitoring processes. However, experimental studies addressing conversational dynamics afforded by abstract concepts are still sparse. In three preregistered experiments we ask whether abstract sentences are associated with specific constructs in dialogue, i.e., higher uncertainty, more curiosity and willingness to continue a conversation, and more questions related to causal and agency aspects. We do so by asking participants to evaluate the plausibility of linguistic exchanges referring to concrete and abstract concepts. Results support theories proposing that abstract concepts involve more inner monitoring and social dynamics compared to concrete concepts and suggest that reaching alignment in dialogue is more effortful with abstract than with concrete concepts.
The concreteness effect has long been associated with embodied theories of language, which propose that concrete words are easier to process than abstract ones because they more directly engage perceptual and motor simulations. However, empirical findings on this effect remain mixed. This paper argues that such variability stems from overlooking a crucial semantic dimension: word specificity. Drawing on evidence from the ERC-funded ABSTRACTION project, I defend (based on classic and more recent empirical studies) that specificity, defined as a word's position within a conceptual hierarchy and corresponding to the inclusiveness of its category, plays a key role in shaping lexical access and conceptual organization, alone and in interaction with concreteness. The relationship between these two dimensions, and its implications for embodied language processing, has so far remained largely unexplored. Integrating specificity into models of embodied semantic representation offers a more nuanced account of how language supports both abstraction and embodiment in cognition.
The use of metaphors, whether linguistic or visual, has been shown to enhance advertisement effectiveness, and sensory marketing research highlights the positive effects of appealing to consumers’ sensory perception. Synaesthetic metaphors, which involve metaphor and sensory experiences, are ideal for studying the effects of both metaphor and (multi)sensory cues in advertisements. We experimentally tested the hypothesis that the presence of (linguistic and/or visual) metaphor and the evocation of multiple senses will enhance advertisement appreciation and the intention to purchase the advertised product. We manipulated eight print advertisements, each of which was presented in the following conditions: (1) visual and linguistic synaesthetic metaphor; (2) linguistic but no visual synaesthetic metaphor; (3) visual but no linguistic synaesthetic metaphor; and (4) neither visual nor linguistic synaesthetic metaphor. Each advertisement was also rated for its multisensoriality, that is, its association with the five basic senses. Results partly supported the hypothesis, showing that advertisements with both visual and linguistic synaesthetic metaphors and those perceived as more multisensory were most appreciated. However, purchase intentions were not influenced by either metaphor or multisensoriality. This indicates that higher aesthetic appreciation does not necessarily translate into higher purchase intentions, suggesting the need for further research into additional influencing factors.
The rapid progress of Large Language Models (LLMs) has transformed natural language processing and broadened its impact across research and society. Yet, systematic evaluation of these models, especially for languages beyond English, remains limited. "Challenging the Abilities of LAnguage Models in ITAlian" (CALAMITA) is a large-scale collaborative benchmarking initiative for Italian, coordinated under the Italian Association for Computational Linguistics. Unlike existing efforts that focus on leaderboards, CALAMITA foregrounds methodology: it federates more than 80 contributors from academia, industry, and the public sector to design, document, and evaluate a diverse collection of tasks, covering linguistic competence, commonsense reasoning, factual consistency, fairness, summarization, translation, and code generation. Through this process, we not only assembled a benchmark of over 20 tasks and almost 100 subtasks, but also established a centralized evaluation pipeline that supports heterogeneous datasets and metrics. We report results for four open-weight LLMs, highlighting systematic strengths and weaknesses across abilities, as well as challenges in task-specific evaluation. Beyond quantitative results, CALAMITA exposes methodological lessons: the necessity of fine-grained, task-representative metrics, the importance of harmonized pipelines, and the benefits and limitations of broad community engagement. CALAMITA is conceived as a rolling benchmark, enabling continuous integration of new tasks and models. This makes it both a resource – the most comprehensive and diverse benchmark for Italian to date – and a framework for sustainable, community-driven evaluation. We argue that this combination offers a blueprint for other languages and communities seeking inclusive and rigorous LLM evaluation practices.
Iconicity, defined as the potential of linguistic signs to resemble properties or features of their referents, is increasingly recognized as a general property of language. One common approach for quantifying iconicity is to collect iconicity ratings. Although iconicity datasets have been developed for several languages, no comprehensive dataset of iconicity ratings is currently available for Italian. The current study presents IconicITA, the first dataset of Italian iconicity ratings for the 1,121 words of the Italian adaptation of Affective Norms for English Words (ANEW). Ratings were collected from both Italian native speakers (L1) and English native speakers with Italian as a second language (L2). Including L2 participants allowed us to contribute to the debate on whether iconicity ratings genuinely measure form-meaning resemblance, rather than exclusively reflecting semantic properties. We showed that L1 Italian iconicity ratings are positively associated with perceptual strength in the auditory and haptic modalities, and with specificity ratings. Conversely, we found a negative correlation between iconicity and concreteness, age of acquisition, word frequency, and letter frequency. In general, the relationship between Italian iconicity norms and various psycholinguistic variables largely replicated previous findings in the literature on iconicity. Considering L2 data, the ratings provided by L2 speakers correlated more strongly with the Italian L1 data compared to the translation-equivalent English L1 data. This finding suggests that participants' judgments were influenced not only by the semantic information of the words but also by language-specific form-level properties. We take this result as evidence of the validity of iconicity ratings to operationalize the degree of resemblance between words' form and meaning.
Recent advances in artificial intelligence (AI)—including generative approaches—have resulted in technology that can support humans in scientific discovery and forming decisions, but may also disrupt democracies and target individuals. The responsible use of AI and its participation in human–AI teams increasingly shows the need for AI alignment, that is, to make AI systems act according to our preferences. A crucial yet often overlooked aspect of these interactions is the different ways in which humans and machines generalize. In cognitive science, human generalization commonly involves abstraction and concept learning. By contrast, AI generalization encompasses out-of-domain generalization in machine learning, rule-based reasoning in symbolic AI, and abstraction in neurosymbolic AI. Here we combine insights from AI and cognitive science to identify key commonalities and differences across three dimensions: notions of, methods for, and evaluation of generalization. We map the different conceptualizations of generalization in AI and cognitive science along these three dimensions and consider their role for alignment in human–AI teaming. This results in interdisciplinary challenges across AI and cognitive science that must be tackled to support effective and cognitively supported alignment in human–AI teaming scenarios. Ilievski et al. examine differences and similarities in the various ways human and AI systems generalize. The insights are important for effectively supporting alignment in human–AI teams.
People can categorize the same entity at multiple taxonomic levels, such as basic (bear), superordinate (animal), and subordinate (grizzly bear). While prior research has focused on basic-level categories, this study is the first attempt to examine the organization of categories by analyzing exemplars produced at the subordinate level. We present a new Italian psycholinguistic dataset of human-generated exemplars for 187 concrete words. We then use these data to evaluate whether textual and vision LLMs produce meaningful exemplars that align with human category organization across three key tasks: exemplar generation, category induction, and typicality judgment. Our findings show a low alignment between humans and LLMs, consistent with previous studies. However, their performance varies notably across different semantic domains. Ultimately, this study highlights both the promises and the constraints of using AI-generated exemplars to support psychological and linguistic research.
WordNet has long served as a benchmark for approximating the mechanisms of semantic categorization in the human mind, particularly through its hierarchical structure of word synsets, most notably the IS-A relation. However, these semantic relations have traditionally been curated manually by expert lexicographers, relying on external resources like dictionaries and corpora. In this paper, we explore whether large language models (LLMs) can be leveraged to approximate these hierarchical semantic relations, potentially offering a scalable and more dynamic alternative for maintaining and updating the WordNet taxonomy. This investigation addresses the feasibility and implications of automating this process with LLMs by testing a set of prompts encoding different sociodemographic traits and finds that adding age and job information to the prompt affects the model ability to generate text in agreement with hierarchical semantic relations while gender does not have a statistically significant impact.
Political debates are a peculiar type of political discourse, in which candidates directly confront one another, addressing not only the the moderator's questions, but also their opponent's statements, as well as the concerns of voters from both parties and undecided voters. Therefore, language is adjusted to meet specific expectations and achieve persuasion. We analyse how the language of Trump and Harris during the Presidential debate (September 10th, 2024) differs in relation to semantic and pragmatic features, for which we formulated targeted hypotheses: framing values and ideology, appealing to emotion, using words with different degrees of concreteness and specificity, addressing others through singular or plural pronouns. Our findings include: differences in the use of figurative frames (Harris often framing issues around recovery and empowerment, Trump often focused on crisis and decline); similar use of emotional language, with Trump showing a slightly higher tendency toward negativity and toward less subjective language compared to Harris; no significant difference in the specificity of candidates' responses; similar use of abstract language, with Trump showing more variability than Harris, depending on the subject discussed; differences in addressing the opponent, with Trump not mentioning Harris by name, while Harris referring to Trump frequently; different uses of pronouns, with Harris using both singular and plural pronouns equally, while Trump using more singular pronouns. The results are discussed in relation to previous literature on Red and Blue language, which refers to distinct linguistic patterns associated with Republican (Red) and Democratic (Blue) political ideologies.
The processes involve two variables that are often confused with one another: concreteness (banana versus belief) and specificity (chair versus furniture or Buddhism versus religion). Researchers are investigating the relationship between them, but many questions remain open, such as: What type of semantics characterizes words with varying degrees of concreteness and specificity? We tackle this topic through an in-depth semantic analysis of 1049 Italian words for which human-generated concreteness and specificity ratings are available. Our findings show that (as expected) the semantics of concrete and abstract concepts differs, but most interestingly when specificity is considered, the variance in concreteness ratings explained by semantic types increases substantially, suggesting the need to carefully control word specificity in future research. For instance, mathematical concepts (phase) are on average abstract and generic, while behavioral qualities (arrogant) are on average abstract but specific. Moreover, through cluster analyses based on concreteness and specificity ratings, we observe the bottom-up emergence of four subgroups of semantically coherent words. Overall, this study provides empirical evidence and theoretical insight into the interplay of concreteness and specificity in shaping semantic categorization.
This study investigates how Large Language Models (LLMs) interpret generics, drawing upon psycholinguistic experimental methodologies. Understanding how LLMs interpret generic statements serves not only as a measure of their ability to abstract but also arguably plays a role in their encoding of stereotypes. Given that generics interpretation necessitates a comparison with explicitly quantified sentences, we explored i.) whether LLMs can correctly associate a quantifier with the generic structure, and ii.) whether the presence of a generic sentence as context influences the outcomes of quantifiers. We evaluated LLMs using both Surprisal distributions and prompting techniques. The findings indicate that models do not exhibit a strong sensitivity to quantification. Nevertheless, they seem to encode a meaning linked with the generic structure, which leads them to adjust their answers accordingly when a generalization is provided as context.
Noun-noun compounds interpretation is the task where a model is given one of such constructions, and it is asked to provide a paraphrase, making the semantic relation between the nouns explicit, as in carrot cake is "a cake made of carrots." Such a task requires the ability to understand the implicit structured representation of the compound meaning. In this paper, we test to what extent the recent Large Language Models can interpret the semantic relation between the constituents of lexicalized English compounds and whether they can abstract from such semantic knowledge to predict the semantic relation between the constituents of similar but novel compounds by relying on analogical comparisons (e.g., carrot dessert). We test both Surprisal metrics and prompt-based methods to see whether i.) they can correctly predict the relation between constituents, and ii.) the semantic representation of the relation is robust to paraphrasing. Using a dataset of lexicalized and annotated noun-noun compounds, we find that LLMs can infer some semantic relations better than others (with a preference for compounds involving concrete concepts). When challenged to perform abstractions and transfer their interpretations to semantically similar but novel compounds, LLMs show serious limitations.
This paper delves into the integration of gamification techniques within the field of linguistics to enhance data collection for academic research purposes. Through an exploration of the Word Ladders mobile application, designed to elicit hierarchical word associations and therefore linguistic data, the study investigates the potential benefits of gamification in terms of data quality, user experience, and motivation in taking part to the research and to the data collection task. The experimental design examines the advantages of a gamified approach compared to traditional research methods (online surveys), through an experimental session followed by a survey (n=189). Results showed that competition between users is a powerful motivator that can be easily integrated in gamified approaches and less so in classic online surveys, driving engagement and potentially enhancing the scalability of data collection while retaining the quality of data collected in classic lab settings. While challenges persist, our research contributes to the understanding of gamification’s impact on data collection, user experience, and motivation, laying the foundation for transformative advancements in the field of language and communication sciences.
Word Ladders is a free mobile application for Android and iOS, developed for collecting linguistic data, specifically lists of words related to each other through semantic relations of categorical inclusion, within the Abstraction project (ERC-2021-STG-101039777). We hereby provide an overview of Word Ladders, explaining its game logic, motivation and expected results and applications to nlp tasks as well as to the investigation of cognitive scientific open questions
Metaphors can provide a conceptual framework for understanding complex topics and as such, they have frequently been used in COVID-19 discourse. As previous research indicates that conceptual metaphors can influence how people reason about complex topics, the metaphors used to communicate about the pandemic can influence how it is understood and how people respond. This paper investigates the influence of metaphorical framing on emotions and reasoning. An experimental study compares BATTLE and JOURNEY metaphor frames in a hypothetical text (adapted from previous studies) about the pandemic. Our aim is to examine the influence of these frames on readers’ affective state, inferences about the attitudes of others and suggestions to stop the virus from spreading. The results suggest that compared to JOURNEY metaphors, BATTLE metaphors can negatively influence affective state and may influence people to suggest more restrictive solutions to reduce the spread of the virus. Yet, metaphors were not found to influence reasoning about the attitudes of others. We interpret these results in relation to previous studies, suggesting that metaphors may influence emotions and reasoning in some ways only. The results have implications for public communication surrounding the pandemic. Limitations and avenues for future research are discussed.
Tulving (1972) characterized semantic memory as a vast repository of meaning that underlies language and many other cognitive processes. This perspective on lexical and conceptual knowledge galvanized a new era of research undertaken by numerous fields, each with their own idiosyncratic methods and terminology. For example, ‘concept’ has different meanings in philosophy, linguistics, and psychology. As such, many fundamental constructs used to delineate semantic theories remain underspecified and/or opaque. Weak construct specificity is among the leading causes of the replication crisis now facing psychology and related fields. Term ambiguity hinders cross-disciplinary communication, falsifiability, and incremental theory-building. Numerous cognitive subdisciplines (e.g., vision, affective neuroscience) have recently addressed these limitations via the development of consensus-based guidelines and definitions. The project to follow represents our effort to produce a multidisciplinary semantic glossary consisting of succinct definitions, background, principled dissenting views, ratings of agreement, and subjective confidence for 17 target constructs (e.g., abstractness, abstraction, concreteness, concept, embodied cognition, event semantics, lexical-semantic, modality, representation, semantic control, semantic feature, simulation, semantic distance, semantic dimension). We discuss potential benefits and pitfalls (e.g., implicit bias, prescriptiveness) of these efforts to specify a common nomenclature that other researchers might index in specifying their own theoretical perspectives (e.g., They said X, but I mean Y).
The Metaphor Compass: Directions for Metaphor Research in Language, Cognition, Communication, and Creativity provides a roadmap to navigate the recent findings and cutting-edge research conducted around the world on metaphor, focusing on the following four themes: Metaphor and Linguistic Diversity, Metaphor and Cognition, Metaphor and Communication, and Metaphor and Creativity. The research presented in this book employs a variety of empirical methods, ranging from neuroimaging to corpus analyses and from behavioral experimentation to computational modeling. Divided into four parts, it offers an array of pedagogical material including activities at the ends of the chapters to help the reader to consolidate the notions discussed in the chapter. This is a useful resource for students, researchers, and scholars of linguistics, communication, anthropology, psychology, and cognitive science looking to learn about figurative language and creativity.
A dataset of specificity ratings for English words is hereby presented, analyzed and discussed in relation with other collections of speaker-generated ratings, including concreteness. Both, specificity and concreteness are analyzed in their ability to explain decision latencies in lexical and semantic tasks, showing important individual contributions. Specificity ratings are collected through best-worst scaling method on the words included in the ANEW dataset (Bradley and Lang in Affective norms for English words (ANEW): instruction manual and affective ratings (Tech. Rep.). Technical report C-1, the center for research in psychophysiology, 1999), chosen for its compatibility with many other collections of rating resources, and for its comparability with Italian specificity data (Bolognesi and Caselli in Behav Res Methods 55(7):3531-3548, 2023), allowing for cross-linguistic comparisons. Results suggest that specificity plays an important role in word processing and the importance of taking specificity into consideration when investigating concreteness effects.