The paper outlines an account of how the brain might process questions and answers in linguistic interaction, focusing on accessing answers in memory and combining questions and answers into propositions. To enable this, we provide an approximation of the lambda calculus implemented in the Semantic Pointer Architecture (SPA), a neural implementation of a Vector Symbolic Architecture. The account builds a bridge between the type-based accounts of propositions in memory (as in the treatments of belief by Ranta, 1994 and Cooper, 2023) and the suggestion for question answering made by Eliasmith (2013), where question answering is described in terms of transformations of structured representations in memory providing an answer. We will take such representations to correspond to beliefs of the agent. On Cooper's analysis, beliefs are considered to be types which have a record structure closely related to the structure which Eliasmith codes in vector representations (Larsson et al., 2023). Thus the act of answering a question can be seen to have a neural base in a vector transformation translatable in Eliasmith's system to activity of spiking neurons and to correspond to using an item in memory (a belief) to provide an answer to the question.
One account of slurring utterances is that they function as a power grab in conversation (Popa-Wyatt and Wyatt in Philos Stud 175(11):2879-2906, 2018). The speaker assigns a powerless role to the target while taking on a dominant role. This establishes a power imbalance in the conversation. However, this power asymmetry creates further conversational effects that differ across the participants. For example, the target is humiliated, while the speaker is not; the speaker's status is elevated while that of the target is diminished. This is the foundation of an act of humiliation. A robust dialogical theory should be able to capture these divergent effects. In this paper, we propose to model this power asymmetry using an existing formal model of dialogue KoS (in: Ginzburg The interactive stance: meaning for conversation. Oxford University Press, Oxford, 2012). Core to KoS is the idea that conversations and interactions are modelled in terms of individual but coupled cognitive states. This framework enables us to model how slurring utterances have differential effects across the speaker and target by capturing these variations through participant-sensitive update rules. We develop a formal account which relates the perceived emotional signals ('Mood') and power relations evolving or already operative in society. However, slurs do more than merely humiliate; they enact sustained oppression over time. To address this, we introduce the concept of oppressive history which draws on two interrelated processes: a cumulative history of slur usage and its emotionally charged impact.
Exception phrases have recently gained renewed attention in semantics. Our contribution is two-fold. (1) We provide a range of empirical data that show a much broader use of exceptives than acknowledged in previous accounts. (2) Based on advances in linguistics and philosophy—namely type-theoretical semantics and a two-dimensional denotational underpinning of plural NPs—we propose a new analysis of exception phrases. We develop a balance scale model of exceptions in terms of reference and complement sets, a model that captures previous analyses as well as the new data. The theoretical generalization to be drawn is that exceptions are linguistic manifestations of reference repair strategies, strategies that are linguistically constrained by the quantificational properties of the NP constituent from which exceptions are to be drawn. Restrictions of different types of exception phrases are formulated as restrictions on an NP’s balance scale within a compositional grammar fragment.
One of the success stories of formal semantics is explicating responsive moves like answers to questions. There is, however, a significant lacune concerning the characterization of initiating utterances , which are strongly tied to the conversational activity [language game (Wittgenstein), speech genre (Bakhtin)], or—our terminology— conversational type , one is engaged in. To date there has been no systematic proposal trying to account for the range of possible language games / speech genres / conversational types and their global structure. In particular, concerning the range of subject matter that can and needs to be discussed and by whom—ultimately a semantic analogue of Laplace’s demon. We suggest that the subject matter problem for conversational types is a central task for any semantic theory for conversation. This paper develops a theory of conversational types, which embedded in the theory of conversational interaction KoS, enables this problem to be tackled for a wide range of conversational types drawn from the British National Corpus classification of conversational domains. The theory we develop treats conversational types as first class, not metatheoretical entities, in contrast to explications of corresponding notions in game theoretic approaches. We demonstrate that this allows us to explicate the possibilities interlocutors have to refer to and seek clarification about the types of conversations they are engaged in.
Laughter serves a wide variety of functions in adult interaction, some of which are quite sophisticated from a pragmatic perspective. Nevertheless, it is a vocalization that emerges very early in ontogeny. We present a longitudinal semantic and pragmatic study of laughter among four American English mother-child dyads interacting freely at home from 12 to 36 months of age. Our data show differences in child laughter-use compared to mothers and a developmental trajectory in terms of the entities the laughter is related to, of laughter pragmatic functions, and in the amount of shared attention on the object of mothers’ laughter. We observe differences in mother laughter-use by comparison to patterns observed in adult-adult interaction. This suggests that laughter production, conveying meaning in a manner akin to speech, gets modulated in child directed interactions similarly to spoken utterances. Our data show that laughter-use can be informative about the neuro-psychological development of babies very early on: mirroring acquisition of physical world knowledge, development of social cognition, linguistic and pragmatic abilities. Our results constitute the basis for hypotheses about the co-option trajectory of laughter in humans and suggest that laughter production in interaction is a valuable resource for early pragmatic development evaluations.
This paper considers how the kind of formal semantic objects used in TTR (a theory of types with records, Cooper, 2023) might be related to the vector representations used in Eliasmith (2013). An advantage of doing this is that it would immediately give us a neural representation for TTR objects as Eliasmith relates vectors to neural activity in his semantic pointer architecture (SPA). This would be an alternative using convolution to the suggestions made by Cooper (2019a) based on the phasing of neural activity. The project seems potentially hopeful since all complex TTR objects are constructed from labelled sets (essentially sets of ordered pairs consisting of labels and values) which might be seen as corresponding to the representation of structured objects which Eliasmith achieves using superposition and circular convolution.
Indirect answers, crucial in human communication, serve to maintain politeness, avoid conflicts, and align with social customs. Although there has been a substantial number of studies on recognizing and understanding indirect answers to polar questions (often known as yes/no questions), there is a dearth of such work regarding wh-questions. This study takes up the challenge by constructing what is, to our knowledge, the first corpus of indirect answers to wh-questions. We analyze and interpret indirect answers to different wh-questions based on our carefully compiled corpus. In addition, we conducted a pilot study on generating indirect answers to wh-questions by fine-tuning the pre-trained generative language model DialoGPT (Zhang et al., 2020). Our results suggest this is a task that GPT finds difficult.
Exclamative sluices (What a goal!, How odd!) have not been part of the debate on the nature of ellipsis constructions. As we show here their profile is quite different from the usual suspects, VP ellipsis and interrogative sluicing – they occur much more frequently than their clausal counterparts and their resolution is more often than not exophoric (i.e., not based on a linguistic antecedent.). On our analysis exclamative sluices are simply scaled up predications of a contextually given entity. We show how this, taken together with an existing semantically-based account of interrogative sluicing, offers an account of the disparity in exophoric potential between the two types of sluicing, an account which is not available to standard, deletion-based accounts. We argue – against most existing views (though with important exceptions which we note and on whose insights we build) – that exclamatives are propositional in nature and develop an account of exclaiming already applicable to laughter, smiling, and head shakes. This account provides a direct link between scaled up predication and increased arousal, itself taken to involve scalar incrementation.
The paper extends a referentially transparent approach which has been successfully applied to the analysis of declarative quantified NPs to wh-phrases. This uses data from dialogical phenomena such as clarification interaction, anaphora, and incrementality as a guide to the design of wh-phrase meanings.
The neurocognition of multimodal interaction – the embedded, embodied, predictive processing of vocal and non-vocal communicative behaviour – has developed into an important subfield of cognitive science. It leaves a glaring lacuna, however, namely the dearth of a precise investigation of the meanings of the verbal and non-verbal communication signals that constitute multimodal interaction. Cognitively construable dialogue semantics provides a detailed and context-aware notion of meaning, and thereby contributes content-based identity conditions needed for distinguishing syntactically or form-based defined multimodal constituents. We exemplify this by means of two novel empirical examples: dissociated uses of negative polarity utterances and head shaking, and attentional clarification requests addressing speaker/hearer roles. On this view, interlocutors are described as co-active agents, thereby motivating a replacement of sequential turn organisation as a basic organising principle with notions of leading and accompanying voices. The Multimodal Serialisation Hypothesis is formulated: multimodal natural language processing is driven in part by a notion of vertical relevance – relevance of utterances occurring simultaneously – which we suggest supervenes on sequential (‘horizontal’) relevance – relevance of utterances succeeding each other temporally.
The main goal of this work is to conduct a pilot study on the automatic classification of the response space of questions in English. We aim for a relatively fine-grained understanding of the learning problem of this response space; hence, we conducted classical machine learning studies to automatically identify different response classes based on carefully designed features. Moreover, we compared the results from feature-based classical machine learning algorithms to the classification results obtained from a large-scale pre-trained BERT language model. Experimental results show that the feature-based classical machine learning algorithms can achieve performance results which are close to the results obtained by BERT model on this novel task. The overall trend of the classification results for each response class are also similar in both models. Learnability trends similar to corpus-based studies presented in previous literatures emerge.
The main aim of this paper is to provide a characterization of the response space for questions using a taxonomy grounded in a dialogical formal semantics. As a starting point we take the typology for responses in the form of questions provided in \cite{lupginz-jlm}. This work develops a wide coverage taxonomy for question/question sequences observable in corpora including the BNC, CHILDES, and BEE, as well as formal modeling of all the postulated classes. Our aim is to extend this work to cover \emph{all} responses to questions. We present the extended typology of responses to questions based on a corpus studies of BNC, BEE, Maptask and CornellMovie with include 506, 262, 467, and 678 question/response pairs respectively. We compare the data for English with data from Polish using the Spokes corpus (694 question/response pairs). We discuss annotation reliability and disagreement analysis. We sketch how each class can be formalized using a dialogical semantics appropriate for dialogue management.
In this paper, we introduce a carefully designed and collected language resource: UgChDial - a Uyghur dialogue corpus based on a chatroom environment. The Uyghur Chat-based Dialogue Corpus (UgChDial) is divided into two parts: (1). Two-party dialogues and (2). Multi-party dialogues. We ran a series of 25, 120-minutes each, two-party chat sessions, totaling 7323 turns and 1581 question-response pairs. We created 16 different scenarios and topics to gather these two-party conversations. The multi-party conversations were compiled from chitchats in general channels as well as free chats in topic-oriented public channels, yielding 5588 unique turns and 838 question-response pairs. The initial purpose of this corpus is to study query-response pairs in Uyghur, building on an existing fine-grained response space taxonomy for English. We provide here initial annotation results on the Uyghur response space classification task using UgChDial.
Laughter is a valuable means for communicating and engaging in interaction since the earliest months of life. Nevertheless, there is a dearth of work on how its use develops in early interactions—given its putative reflexive nature, it has often been disregarded from studies on pre-linguistic vocalizations. We provide a longitudinal characterization of laughter use analyzing interactions of 4 babies with their mothers at five time-points (12, 18, 24, 30, and 36 months). We show how child laughter is very distinct from mothers’ (and adults’ generally), in terms of frequency, duration, level of arousal displayed, overlap with speech, and responsiveness to others’ laughter. Notably, contrary to what might be expected, we observed that children laugh significantly less than their mothers, especially at the first time-points analyzed. We indeed observe an increasing developmental trajectory in the production of laughter overall and in the contingent multimodal response to mothers’ laughter, showing the child’s increasing attunement to the social environment, interest in others’ appraisals and mental states, and awareness of its communicative value. We also show how mothers’ contingent responses to child laughter change over time, going from high-frequency mimicry, to a lower rate of diversified multimodal responses, in line with the child’s neuro-psychological development. Our data support a dynamic view of dialogue where interactants influence each other bidirectionally and emphasizes the crucial communicative value of laughter. When language is not fully developed, laughter might be an early means, in its already fully available expressiveness, to hold the conversational turn and enable meaningful vocal contribution in interaction at the same level of the interlocutor. Our study aims to provide a benchmark for typical laughter development, since we believe it can be an early means, along with other commonly analyzed behaviors (e.g., smiling, gazing, pointing, etc.), to gain insight into early child neuro-psychological development.
Laughter is a crucial signal for communication and managing interactions. Until now no consensual approach has emerged for classifying laughter. We propose a new framework for laughter analysis and classification, based on the pivotal assumption that laughter has propositional content. We propose an annotation scheme to classify the pragmatic functions of laughter taking into account the form, the laughable, the social, situational, and linguistic context. We apply the framework and taxonomy proposed in a multilingual corpus study (French, Mandarin Chinese, and English), involving a variety of situational contexts. Our results give rise to novel generalizations about the range of meanings laughter exhibits, the placement of the laughable, and how placement and arousal relate to the functions of laughter. We have tested and refuted the validity of the commonly accepted assumption that laughter directly follows its laughable. In the concluding section, we discuss the implications our work has for spoken dialogue systems. We stress that laughter integration in spoken dialogue systems is not only crucial for emotional and affective computing aspects, but also for aspects related to natural language understanding and pragmatic reasoning. We formulate the emergent computational challenges for incorporating laughter in spoken dialogue systems.
In many instances, the head shake can be used instead of or in addition to verbal ˋNo'. Based on previous work on negation in dialogue, we observe head shaking as answer particles and as responding to an implicit or an exophoric (i.e., real world situation) antecedent. Exophoric head shake, however, seems to come in two flavours: with positive and with negative emotional valuation of the antecedent situation. We provide semantic analyses for all three uses (and a head nod) within an HPSG version which is implemented in Type Theory with Records and the dialogue framewok KoS. In particular, we extend on previous work by grounding ˋˋexophoric negation'' in positive or negative appraisal. Finally, we briefly speculate about differences between verbal ˋNo' and head shaking due to (the lack of) simultaneity.
This chapter portrays some phenomena, technical developments and discussions that are pertinent to analysing natural language use in face-to-face interaction from the perspective of HPSG and closely related frameworks. The use of the CONTEXT attribute in order to cover basic pragmatic meaning aspects is sketched. With regard to the notion of common ground, it is argued how to complement CONTEXT by a dynamic update semantics. Furthermore, this chapter discusses challenges posed by dialogue data such as clarification requests to constrained-based, model-theoretic grammars. Responses to these challenges in terms of a type-theoretical underpinning (TTR, a Type Theory with Records) of both the semantic theory and the grammar formalism are reviewed. Finally, the dialogue theory KoS that emerged in this way from work in HPSG is sketched.
In multimodal natural language interaction both speech and non-speech gestures are involved in the basic mechanism of grounding and repair. We discuss a couple of multimodal clarifica- tion requests and argue that gestures, as well as speech expressions, underlie comparable paral- lelism constraints. In order to make this precise, we slightly extend the formal dialogue frame- work KoS to cover also gestural counterparts of verbal locutionary propositions.