In collaborative dialogue, shared perception does not guarantee shared interpretation. Mutual understanding must be established through interaction. We investigate whether vision-language models (VLMs) can distinguish what could be shared from what has been shared between dialogue participants through grounding. We formulate this as an interpretation-matching task on 13,077 annotated reference expressions from HCRC MapTask dialogues, and evaluate VLMs under systematically controlled manipulations of dialogue context and map-information access. Our results show that providing authentic map images improves overall performance but shifts models toward over-predicting alignment. Textual descriptions of the same map content reproduce this bias, while non-informative images suppress alignment predictions entirely, indicating that the bias is driven by task-relevant map content, not the visual channel. This improvement comes at the cost of degraded accuracy on non-aligned cases. Calibration analysis and reference-chain tracking further suggest that models rely on static referential cues on the maps rather than tracking how grounding unfolds through dialogue history. We observe these patterns most clearly in Qwen3-VL-8B-Instruct and, to varying degrees, in four additional models from two architecture families. In models that exhibit the bias, map content, whether presented visually or textually, is treated as evidence of mutual understanding, conflating potential with established common ground.
Collaborative dialogue relies on participants incrementally establishing common ground, yet in asymmetric settings they may believe they agree while referring to different entities. We introduce a perspectivist annotation scheme for the HCRC MapTask corpus (Anderson et al., 1991) that separately captures speaker and addressee grounded interpretations for each reference expression, enabling us to trace how understanding emerges, diverges, and repairs over time. Using a scheme-constrained LLM annotation pipeline, we obtain 13k annotated reference expressions with reliability estimates and analyze the resulting understanding states. The results show that full misunderstandings are rare once lexical variants are unified, but multiplicity discrepancies systematically induce divergences, revealing how apparent grounding can mask referential misalignment. Our framework provides both a resource and an analytic lens for studying grounded misunderstanding and for evaluating (V)LLMs' capacity to model perspective-dependent grounding in collaborative dialogue.
Games With A Purpose (GWAPs) for linguistic annotation often struggle to sustain engagement. We investigate whether puzzles and emojis enhance engagement in a word-sense disambiguation (WSD) task. Participants completed three within-subject conditions: (1) puzzles with emojis, (2) puzzles, and (3) annotation alone. Novices found visual cues motivating; experts preferred streamlined tasks. A reflexive thematic analysis revealed distinct preferences regarding puzzle complexity, visual aids, and perceived cognitive load.
Central to human-aligned AI is understanding the benefits of human-elicited labels over synthetic alternatives. While human soft-labels improve calibration by capturing uncertainty, prior studies conflate these benefits with the implicit correction of mislabeled data (mode shifts), obscuring true effects of soft-labels. We present a controlled audit of soft-label learning across MNIST and a synthetic variant, re-annotating subsets to extract human uncertainty. By decoupling soft-label supervision from underlying label mode shifts, we show that while human soft-labels do provide accuracy gains, their larger value lies in acting as a regularizer that improves model calibration on difficult samples and promotes stable convergence across training runs. Dataset cartography reveals models trained on human soft-labels mirror human uncertainty, whereas those trained on synthetic labels fail to align with humans. Broadly, this work provides a diagnostic testbed for human-AI uncertainty alignment.
The current research discusses the thematic alignment of narratives and its impact on user engagement in gamified apps, using NLP Game-with-a-Purpose (GWAP) as the experimental environment. The experimental game is made up of a three-dimensional environment wherein a scene-based narrative is contextually integrated with the game world, and there is no such integration in the case of a non-scene-based narrative. Quantitative data gathered from 80 participants shows that a scene-based narrative contributes much better to user engagement, scoring higher in Focused Attention (FA) and Reward (RW) compared to the non-scene-based narrative. These findings are supported by qualitative feedback and think-aloud protocols, which showed that the participants found the scene-based narrative to be more engaging and involved than the non-scene-based material. Even though there were no statistically significant differences in cognitive load, trends of mental demand and frustration were observed indicating that thematic alignment has the potential to streamline the user experience. These findings have design implications in developing narrative-based gamified systems that can be used to improve interaction in language-related tasks.
Cross-linguistic studies have demonstrated that individuals with schizophrenia-particularly those exhibiting formal thought disorder (FTD)-show distinctive distributions of noun phrases (NPs) in spontaneous speech. NPs (e.g., the picture; a husband) serve to organize the referential structure of meaning. Extracting such referential NP features, however, has traditionally required manual annotations. In this study we applied state-of-the-art large language models (LLMs) to extract these features automatically, using an existing, manually annotated dataset, in which English-speaking participants described a comic strip: 30 individuals with schizophrenia (SZ) (15 with moderate or severe FTD (SZ + FTD), 15 with minimal or no FTD (SZ-FTD), 15 neurotypical controls (NC). We first show that LLM-based analyses replicate the findings based on manual annotation, particularly highlighting that definite NPs tied to prior discourse, markers of grammatical and cognitive complexity and narrative coherence, were underused in the SZ + FTD group. Secondly, we demonstrate that LLMs, especially when used with in-context (few-shot) learning, offer a promising avenue for the automatic extraction of referential features. These results show that a cross-linguistically validated and clinically important linguistic pattern of deviance is accessible to automatized assessment with NLP.
Dialogue-Based Generalized Referring Expressions Comprehension (GREC) requires models to ground the expression and unlimited targets in complex visual scenes while resolving coreference across a long dialogue context. However, existing systems struggle under distribution shift between training and evaluation domains, a gap exacerbated by the scarcity of annotated dialogue grounding data. We address this challenge with a three-tier data-synthesis method that balances realism and controllability to produce scalable supervision for dialogue-conditioned grounding. Fine-tuning on the synthesized data yields consistent, substantial improvements over prior approaches across standard evaluation metrics.
There is growing recognition that many NLP tasks lack a single ground truth, as human judgments reflect diverse perspectives. To capture this variation, models have been developed to predict full annotation distributions rather than majority labels, while perspectivist models aim to reproduce the interpretations of individual annotators. In this work, we compare these approaches on Implicit Discourse Relation Recognition (IDRR), a highly ambiguous task where disagreement often arises from cognitive complexity rather than ideological bias. Our experiments show that existing annotator-specific models perform poorly in IDRR unless ambiguity is reduced, whereas models trained on label distributions yield more stable predictions. Further analysis indicates that frequent cognitively demanding cases drive inconsistency in human interpretation, posing challenges for perspectivist modeling in IDRR.
Coreference resolution is an important task in contextual reasoning. In this paper, we investigate the mechanism for representing and retrieving singular and plural entities for plural reference. We use a combination of mechanistic interpretability and attention pattern analysis to study the process in which LLMs predict a pronoun to refer back to previously mentioned entities. Using a range of causal intervention techniques, we find a set of attention heads that are responsible for (1) representing coreference information in the input, (2) identifying entities that form a plural reference, (3) transferring the information to the component that is responsible for selecting the antecedents and predicting the pronoun. We also find that LLMs align with humans in preference for plural pronoun. Specifically, entities in a plural construction are more likely to be referred to as a plural entity if they are ontologically similar and are linked by the conjunction "and".
This paper examines whether AI-generated images improve user experience and performance in gamified text-labelling tasks. Across two experiments, we investigate how images interact with both task complexity and data ambiguity. In Experiment 1, participants completed either a simple part-of-speech (POS) tagging task or a more complex natural language inference (NLI) task, with or without AI-generated images. Results showed that task complexity strongly influenced accuracy, engagement, enjoyment, and cognitive load, while AI-generated images provided no consistent benefit. Notably, in the POS task, engagement was higher in the absence of images. Experiment 2 focused on data complexity within NLI, comparing AI-generated images, Flickr images, and a no-image condition across low- and high-ambiguity items. Even when images were designed to be semantically aligned with the text, visual context did not improve engagement or accuracy, nor mitigate the effects of ambiguity; accuracy was often highest without images. Together, these findings suggest that in linguistic annotation tasks, performance and experience are driven primarily by textual complexity rather than visual support, challenging assumptions about benefits of adding images to gamified text annotation.
We introduce ClarQ4LLM, an evaluation framework consisting of bilingual English-Chinese conversation tasks, conversational agents and evaluation metrics, designed to serve as a strong benchmark for assessing agents' ability to ask clarification questions in task-oriented dialogues. The benchmark includes 31 different task types, each with 10 unique dialogue scenarios between information seeker and provider agents. The scenarios require the seeker to ask questions to resolve uncertainty and gather necessary information to complete tasks. Unlike traditional benchmarks that evaluate agents based on fixed dialogue content, ClarQ4LLM includes a provider conversational agent to replicate the original human provider in the benchmark. This allows both current and future seeker agents to test their ability to complete information gathering tasks through dialogue by directly interacting with our provider agent. In tests, LLAMA3.1 405B seeker agent managed a maximum success rate of only 60.05%, showing that ClarQ4LLM presents a strong challenge for future research(1).
Proper names are linguistic expressions referring to unique entities, such as individual people or places. This sets them apart from other words like common nouns, which refer to generic concepts. And yet, despite both being individual entities, one's closest friend and one's favorite city are intuitively associated with very different pieces of knowledge—face, voice, social relationship, autobiographical experiences for the former, and mostly visual and spatial information for the latter. Neuroimaging research has revealed the existence of both domain-general and domain-specific brain correlates of semantic processing of individual entities; however, it remains unclear how such commonalities and similarities operate over a fine-grained temporal scale. In this work, we tackle this question using EEG and multivariate (time-resolved and searchlight) decoding analyses. We look at when and where we can accurately decode the semantic category of a proper name and whether we can find person- or place-specific effects of familiarity, which is a modality-independent dimension and therefore avoids sensorimotor differences inherent among the two categories. Semantic category can be decoded in a time window and with spatial localization typically associated with lexical semantic processing. Regarding familiarity, our results reveal that it is easier to distinguish patterns of familiarity-related evoked activity for people, as opposed to places, in both early and late time windows. Second, we discover that within the early responses, both domain-general (left posterior-lateral) and domain-specific (right fronto-temporal, only for people) neural patterns can be individuated, suggesting the existence of person-specific processes.
Coreference Resolution (CR) is crucial for many NLP tasks, but existing LLMs struggle with hallucination and under-performance. In this paper, we investigate the limitations of existing LLM-based approaches to CR-specifically the Question-Answering (QA) Template and Document Template methods-and propose two novel techniques: Reversed Training with Joint Inference and Iterative Document Generation. Our experiments show that Reversed Training improves the QA Template method, while Iterative Document Generation eliminates hallucinations in the generated source text and boosts coreference resolution. Integrating these methods and techniques offers an effective and robust solution to LLM-based coreference resolution (1).
Coreference resolution is the task of resolving mentions that refer to the same entity into clusters. The area and its tasks are crucial in natural language processing (NLP) applications. Extensive surveys of this task have been conducted for English and Chinese; not too much for Arabic. The few Arabic surveys do not cover recent progress and the challenges for Arabic anaphora; and do not cover zero resolution and comprehensive resolution of zeros and full mentions, or anaphora resolution beyond coreference (e.g., bridging). In this paper, we examine the state-of-the-art for Arabic anaphora resolution, highlighting the challenges and advances in this field. We provide a comprehensive survey of the methods employed for Arabic coreference resolution, as well as an overview of the existing datasets and challenges. The goal is to equip researchers with a thorough understanding of Arabic anaphora resolution and to suggest potential future directions in the field.
This study investigates how task complexity influences the impact of AI-generated images on user experience in gamified text-labelling tasks. Participants completed one of two tasks of varying complexity: part-ofspeech (POS) tagging simple and natural language inference (NLI) complex under conditions with and without AIgenerated images. We measured accuracy, engagement, enjoyment, and cognitive load using factorial ANOVAs. Results revealed that task complexity significantly influenced all outcomes: the NLI task yielded lower accuracy, reduced engagement and enjoyment, and increased cognitive load compared with the POS task. AI-generated images had limited impact, negatively affecting perceived usability and overall engagement, with engagement notably higher in the absence of images for the POS task. These findings suggest that while task complexity strongly shapes user experience in gamified labelling tasks, AI-generated images may not enhance performance or engagement.
Understanding the sources of variability in annotations is crucial for developing fair NLP systems, especially for tasks like sexism detection where demographic bias is a concern. This study investigates the extent to which annotator demographic features influence labeling decisions compared to text content. Using a Generalized Linear Mixed Model, we quantify this inf luence, finding that while statistically present, demographic factors account for a minor fraction ( 8%) of the observed variance, with tweet content being the dominant factor. We then assess the reliability of Generative AI (GenAI) models as annotators, specifically evaluating if guiding them with demographic personas improves alignment with human judgments. Our results indicate that simplistic persona prompting often fails to enhance, and sometimes degrades, performance compared to baseline models. Furthermore, explainable AI (XAI) techniques reveal that model predictions rely heavily on content-specific tokens related to sexism, rather than correlates of demographic characteristics. We argue that focusing on content-driven explanations and robust annotation protocols offers a more reliable path towards fairness than potentially persona simulation.
How effective are the communication choices of Multimodal Large Language Models when pursuing a common goal? Can they make use of common human dialogical patterns? We address these questions by engaging two agents based on the Mistral model in a collaborative building task, where one has to instruct the other how to build a specific target structure. The aim of this work is to investigate whether different prompting techniques with varying degrees of multimodality can influence the performance of MLLM-based agents in the proposed task. Code and data available in the project's GitHub repository.
With large language models (LLMs) on the rise, in-game interactions are shifting from rigid commands to natural conversations. However, the impacts of LLMs on player performance and game experience remain underexplored. This work explores LLM's role as a co-builder during gameplay, examining its impact on task performance, usability, and player experience. Using Minecraft as a sandbox, we present an LLM-assisted interface that engages players through natural language, aiming to facilitate creativity and simplify complex gaming commands. We conducted a mixed-methods study with 30 participants, comparing LLM-assisted and command-based interfaces across simple and complex game tasks. Quantitative and qualitative analyses reveal that the LLM-assisted interface significantly improves player performance, engagement, and overall game experience. Additionally, task complexity has a notable effect on player performance and experience across both interfaces. Our findings highlight the potential of LLM-assisted interfaces to revolutionize virtual experiences, emphasizing the importance of balancing intuitiveness with predictability, transparency, and user agency in AI-driven, multimodal gaming environments.