Language is often used strategically, particularly in high-stakes, adversarial settings, yet most work on pragmatics and LLMs centers on cooperativity. This leaves a gap in the systematic understanding of strategic communication in adversarial settings. To address this, we introduce SDA (Strategic Dialogue Assessment), a framework grounded in Gricean and game-theoretic pragmatics to assess strategic use of language. It adapts the ME Game jury function to make it empirically estimable for analyzing dialogue. Our approach incorporates two key adaptations: a commitment-based taxonomy of discourse moves, which provides a finer-grained account of strategic effects, and the use of estimable proxies grounded in Gricean maxims to operationalize abstract constructs such as credibility. Together, these adaptations build on discourse theory by treating discourse as the strategic management of commitments, enabling systematic evaluation of how conversational moves advance or undermine discourse goals. We further derive three interpretable metrics-Benefit at Turn (BAT), Penalty at Turn (PAT), and Normalized Relative Benefit at Turn (NRBAT)-to quantify the perceived strategic effects of discourse moves. We also present CPD (the Crooked Path Dataset), an annotated dataset of real courtroom cross-examinations, to demonstrate the framework's effectiveness. Using these tools, we evaluate a range of LLMs and show that LLMs generally exhibit limited pragmatic understanding of strategic language. While model size shows an increase in performance on our metrics, reasoning ability does not help and largely hurts, introducing overcomplication and internal confusion.
We introduce HERO'S JOURNEY, a benchmark for rule induction in goal-directed episodic tasks, where agents must infer hidden rules from demonstrations and act on them through multi-step execution. HERO'S JOURNEY covers eight tasks across attribute and procedural induction families, each with four structural rule forms, controllable lexical grounding, and identifiability conditions. Evaluating state-of-the-art LLMs, we find that models show evidence of rule induction, but the ability is limited and uneven across tasks. Meanwhile, process execution adds an execution bottleneck for models, whereas surface semantics has minimal effect. Induction-specific steering methods improve performance on attribute tasks but show no reliable gains on procedural tasks, suggesting the gap in procedural induction remains an open challenge.
This study probes how the semantics of ordinals relates to the semantics of comparatives and superlatives. We examine this question with the help of a picture task in which participants are asked to locate objects described by nested descriptions like the candle on the first/closer/closest table, with an ordinal, comparative or superlative modifier in the inner noun phrase. We show that ordinals systematically lack the ‘relative readings’ observed for unmodified nested descriptions like the rabbit in the hat, in which the inner definite is understood with enriched content, as in the rabbit in the hat with a rabbit in it, in contrast to superlatives. Our explanation for this relies on the idea that an ordinal expects an ordering that can be provided by context.
The variations between in-group and out-group speech (intergroup bias) are subtle and could underlie many social phenomena like stereotype perpetuation and implicit bias. In this paper, we model intergroup bias as a tagging task on English sports comments from forums dedicated to fandom for NFL teams. We curate a dataset of over 6 million game-time comments from opposing perspectives (the teams in the game), each comment grounded in a non-linguistic description of the events that precipitated these comments (live win probabilities for each team). Expert and crowd annotations justify modeling the bias through tagging of implicit and explicit referring expressions and reveal the rich, contextual understanding of language and the world required for this task. For large-scale analysis of intergroup variation, we use LLMs for automated tagging, and discover that LLMs occasionally perform better when prompted with linguistic descriptions of the win probability at the time of the comment, rather than numerical probability. Further, large-scale tagging of comments using LLMs uncovers linear variations in the form of referent across win probabilities that distinguish in-group and out-group utterances.
While existing work on studying bias in NLP focues on negative or pejorative language use, Govindarajan et al. (2023) offer a revised framing of bias in terms of intergroup social context, and its effects on language behavior. In this paper, we investigate if two pragmatic features (specificity and affect) systematically vary in different intergroup contexts -- thus connecting this new framing of bias to language output. Preliminary analysis finds modest correlations between specificity and affect of tweets with supervised intergroup relationship (IGR) labels. Counterfactual probing further reveals that while neural models finetuned for predicting IGR labels reliably use affect in classification, the model's usage of specificity is inconclusive. Code and data can be found at: https://github.com/venkatasg/intergroup-probing
Current studies of bias in NLP rely mainly on identifying (unwanted or negative) bias towards a specific demographic group. While this has led to progress recognizing and mitigating negative bias, and having a clear notion of the targeted group is necessary, it is not always practical. In this work we extrapolate to a broader notion of bias, rooted in social science and psychology literature. We move towards predicting interpersonal group relationship (IGR) - modeling the relationship between the speaker and the target in an utterance-using fine-grained interpersonal emotions as an anchor. We build and release a dataset of English tweets by US Congress members annotated for interpersonal emotion - the first of its kind, and 'found supervision' for IGR labels; our analyses show that subtle emotional signals are indicative of different biases. While humans can perform better than chance at identifying IGR given an utterance, we show that neural models perform much better; furthermore, a shared encoding between IGR and interpersonal perceived emotion enabled performance gains in both tasks.
AbstractWe introduce a framework for studying repair initiation in the face of miscommunication. Our aim is to seed development of models that both predict when conversational repair is a likely communicative strategy and explain why interlocutors would not engage in repair in the face of conversational difficulty. We identify three factors as critical to the predictability of repair: (i) the extent to which a misalignment is (un)recognized by participants (ignorance); (ii) the significance of misalignment relative to some cluster of goals (cost of misalignment); and (iii) the significance of engaging in repair relative to some cluster of goals (cost of repair). We offer a simple method for graphically depicting relevant aspects of communicative situations and exemplify the framework with examples of non-repaired miscommunication before discussing its applicability to different empirical domains.
The ability of language to perpetuate inequality is most evident when individuals refer to, or talk about, other individuals in their utterances. While current studies of bias in NLP rely mainly on identifying hate speech or bias towards a specific group, we believe we can reach a more subtle and nuanced understanding of the interaction between bias and language use by modeling the speaker, the text, and the target in the text. In this paper, we introduce a dataset of 3033 English tweets by US Congress members annotated for interpersonal emotion, and ‘found supervision’ for interpersonal group membership labels. We find that negative emotions such as anger and disgust are used predominantly in out-group situations, and directed predominantly at leaders of opposite parties. While humans can perform better than chance at identifying interpersonal group membership given an utterance, neural models perform much better; furthermore, a shared encoding between interpersonal group membership and interpersonal perceived emotion enabled some performance gains in the latter. This work aims to re-align the study of bias in NLP away from specific instances of bias to one which encapsulates the relationship between speaker, text, target and social dynamics.
Over a century of scholarship on presupposition has worked towards reconciling two seemingly contrary properties of these types of inferences: the ability to project through embedding like negation, and the ability to be cancelled explicitly. Describing these properties has been key to not only diagnosing presuppositions, but also differentiating them from other types of inferences like implicatures and entailment and understanding how a theory of presupposition could apply cross-linguistically. This chapter outlines different accounts of presupposition and negation, focusing on six different broad approaches: scope ambiguity, trivalent ambiguity, underspecification, metalinguistic negation, cancellation, and accommodation. These accounts differ with respect to whether they account for default projection, the mechanisms through which projection is derived, and whether entailments and implicatures are targeted by the same negation operators as presuppositions.
While much prior literature on the meaning of clefts-such as the English form "it is X who Z-ed"-concentrates on the nature and status of the exhaustivity inference ("nobody/nothing other than X Z"), we report on experiments examining the role of the doxastic status of alternatives on the naturalness of c'est-clefts in French and it-clefts in English. Specifically, we study the hypothesis that clefts indicate a conflict with a doxastic commitment held by some discourse participant. Results from naturalness tasks suggest that clefts are improved by a property we term "contrariness" (along the lines of Zimmermann, 2008). This property has a gradient effect on felicity judgments: the more strongly interlocutors appear committed to an apparently false notion, the better it is to repudiate them with a cleft.
We discuss presupposition, the phenomenon whereby speakers mark linguistically the information that is presupposed or taken for granted, rather than being part of the main propositional content of a speech act. Expressions and constructions carrying presuppositions are called “presupposition triggers”, which is a large class including definites and factive verbs. The article (an abridged and adapted version of Beaver & Geurts 2010), first introduces the range of triggers, the basic properties of presuppositions such as projection and cancellability, and the diagnostic tests used to identify them. The reader is then introduced to major models of presupposition from the last 50 years, separated into three classes: Frege-Strawson derived semantic models, pragmatic models such as that offered by Stalnaker, and dynamic models. Finally we discuss some of the main current issues in presupposition theory, including accommodation, which occurs when a hearer’s knowledge state is adjusted to meet the speaker’s presuppositions; presupposition failure, and the interaction between presuppositions and attitudes.
Projective content is utterance content that a speaker may be taken to be committed to even when the expression associated with the content occurs embedded under an entailment-canceling operator (e.g., Chierchia & McConnell-Ginet, 1990). It has long been observed that projective content varies in how projective it is (e.g., Karttunen, 1971; Simons, 2001; Abusch, 2010), though preliminary experimental research has been able to confirm only some of the intuitions about projection variability (e.g., Smith & Hall, 2011; Xue & Onea, 2011). Given the sparse empirical evidence for projection variability, the first goal of this paper was to investigate projection variability for projective content associated with 19 expressions of American English. The second goal was to explore the hypothesis, called the Gradient Projection Principle, that content projects to the extent that it is not at-issue. The findings of two pairs of experiments provide robust empirical evidence for projection variability and for the Gradient Projection Principle. We show that many analyses of projection cannot account for the observed projection variability and discuss the implications of our finding that projective content varies in its at-issueness for an empirically adequate analysis of projection.