We address a longstanding issue in visual world studies of speech and spoken word recognition: What role does retrieval of object names play in mediating eye movements to displayed objects? We propose a novel “linguistic-visual routines” linking hypothesis in which the representations of spoken words are linked to schematic visual representations and visuomotor routines, without involvement of object names. When triggered during reference resolution, these routines direct gaze to objects in a co-present visual world that are consistent with the schematic visual representations. To evaluate possible picture-naming in this task, we created materials that included “synonym” objects (e.g., a picture of an object that could be named as either “couch” or “sofa”, where norms established that each name was an equally good fit for the object, but one name was dominant, in this case, “couch”. Name typicality effects were used to diagnose if and when object names affect eye movements, using targets that were cohorts of either the dominant or subordinate name of a competitor. We reasoned that if names were retrieved, then there would be larger cohort effects and delayed looks to targets when the target was a cohort of the dominant name. We ran conceptual simulations using normalized recurrence localist attractor networks, illustrating predicted data patterns for models where: (1) object names do not affect fixations; (2) retrieval of object names mediates fixations, and (3) sub-threshold priming of object names weakly affects eye-movements. Experiment 1 used a standard Visual World design, finding minimal evidence, if any, for effects of name typicality. In Experiments 2 and 3, participants were cued that a particular object would be occluded before the instruction begins and thus were likely to retrieve the object name to help maintain it in memory. On critical trials, synonym objects were occluded in Experiment 2, and distractor objects were occluded in Experiment 3. Simulations were again used to illustrate predicted data patterns, namely name-retrieval effects only when the synonym object was occluded. In Experiment 2, where the synonym object was occluded, there were strong name typicality effects, with larger cohort effects and delayed target looks for the dominant name, matching the predictions of the model with name retrieval. In Experiment 3, where a distractor object was occluded, cohort effects were not affected by name typicality. We then introduce models simulating conditions from a pilot study in which visually drawing attention to an object, but not occluding it, also does not result in name typicality effects. Taken together, the results are clearly consistent with the linguistic-visual routines hypothesis and inconsistent with name retrieval. We conclude that object name representations play at most a minimal role in mediating fixations in Visual World experiments of spoken word recognition.
Although the term Visual World Paradigm (henceforth VWP) is used to refer to the broad class of studies in which participants eye movements are measured as they listen to language, that is about a circumscribed visual display (henceforth the visual world), there are, in fact, two broadly used variants of the paradigm. The first, introduced by researchers at Rochester in the mid-1990 s, typically used the visual world as a type of workspace that participants interact with, for example following instructions to perform an action or sequence of actions (e.g., "Put the apple on the towel in the box"; "Put the big candle into the trash. Now put the small tie into the blue square."). The second, introduced by Gerry Altmann and colleagues, typically narrates an event or sequence of events, using a display with depicted objects and people (e.g., "The boy will eat the cake."; "Donald is bringing some mail to Mickey while a violent storm is beginning. He's carrying an umbrella…") without asking participants to perform an accompanying action. While the approaches are often used to address similar questions, there are some, often implicit, differences between the assumptions that motivate the different approaches. But what are these assumptions? Are there types of questions for which one of the approaches is better suited than the other? Does the choice of approach affect linking hypotheses? We address these issues in a paper that takes the form of a dialogue, with MKT making the case for including tasks with actions and FH making the case for experiments without an additional action. After responding to each other's arguments, we conclude by: (1) separating principled differences from associations that are tied to the types of questions that were first addressed in some of the foundational studies; (2) making suggestions for factors that should guide researchers' choice of approach; and (3) proposing new avenues of research.
The diversity of contexts in which a word occurs, operationalized as CD, is strongly correlated with response times in visual word recognition, with higher CD words being recognized faster. CD and token word frequency (WF) are highly correlated but in behavioral studies when other variables that affect word visual recognition are controlled for, the WF effect is eliminated when contextual diversity (CD) is controlled. In contrast, the only event-related potential (ERP) study to examine CD and WF Vergara-Martínez et al., Cognitive, Affective, Behavioral Neuroscience, 17, 461–474, (2017) found effects of both WF and CD with different distributions in the 225- to 325-ms time window. We conducted an ERP study with Chinese characters to explore the neurocognitive dynamics of WF and CD. We compared three groups of characters: (1) characters high in frequency and low in CD; (2) characters low in frequency and low in CD; and (3) characters high in frequency and high in CD. Behavioral data showed significant effects of CD but not WF. Character CD, but not character frequency, modulated the late positive component (LPC): high-CD characters elicited a larger LPC, widely distributed, with largest amplitude at the posterior sites compared to low-CD characters in the 400-to 600-ms time window, consistent with earlier ERP studies of WF in Chinese, and with the hypothesis that CD affects semantic and context-based processes. No WF effect on any ERP components was observed when CD was controlled. The results are consistent with behavioral results showing CD but not WF effects, and in particular with a “context constructionist” framework.
Linguistic communication requires interlocutors to consider differences in each other’s knowledge (perspective-taking). However, perspective-taking might either be spontaneous or strategic. We monitored listeners’ eye movements in a referential communication task. A virtual speaker gave temporally ambiguous instructions with scalar adjectives (“big” in “big cubic block”). Scalar adjectives assume a contrasting object (a small cubic block). We manipulated whether the contrasting object (a small triangle) for a competitor object (a big triangle) was in common ground (visible to both speaker and listener) or was occluded so it was in the listener’s privileged ground, in which case perspective-taking would allow earlier reference-resolution. We used a complex visual context with multiple objects, making strategic perspective-taking unlikely when all objects are in the listener’s referential domain. A turn-taking, puzzle-solving task manipulated whether participants could anticipate a more restricted referential domain. Pieces were either confined to a small area (requiring fine-grained coordination) or distributed across spatially distinct regions (requiring only coarse-grained coordination). Results strongly supported spontaneous perspective-taking: Although comprehension was less time-locked in the coarse-grained condition, participants in both conditions used perspective information to identify the target referent earlier when the competitor contrast was in privileged ground, even when participants believed instructions were computer-generated.
Accurate word recognition is facilitated by context. Some relevant context, however, occurs after the word. Rational use of such "right context" would require listeners to have maintained uncertainty or subcategorical information about the word, thus allowing for consideration of possible alternatives when they encounter relevant right context. A classic study continues to be widely cited as evidence that subcategorical information maintenance is limited to highly ambiguous percepts and short time spans (Connine et al., 1991). More recent studies, however, using other phonological contrasts, and sometimes other paradigms, have returned mixed results. We identify procedural and analytical issues that provide an explanation for existing results. We address these issues in two reanalyses of previously published results and two new experiments. In all four cases, we find consistent evidence against both limitations reported in Connine et al.'s seminal work, at least within the classic paradigms. Key to our approach is the introduction of an ideal observer framework to derive normative predictions for human word recognition expected if listeners maintain and integrate subcategorical information about preceding speech input rationally with subsequent context. We test these predictions in Bayesian mixed-effect analyses, including at the level of individual participants. While we find that the ideal observer fits participants' behavior better than models based on previously proposed limitations, we also find one previously unrecognized aspect of listeners' behavior that is unexpected under any existing model, including the ideal observer.
Perspective-taking, which is important for communication and social activities, can be cultivated through joint actions, including musical activities in children. We examined how rhythmic activities requiring coordination affect perspective-taking in a referential communication task with 100 Chinese 4- to 6-year-old children. In Study 1, 5- to 6-year-old children played an instrument with a virtual partner in one of three coordination conditions: synchrony, asynchrony, and antiphase synchrony. Eye movements were then monitored with the partner giving instructions to identify a shape referent which included a pre-nominal scalar adjective (e.g., big cubic block). When the target contrast (a small cubic block) was in the shared ground and a competitor contrast was occluded for the partner, participants who used perspective differences could, in principle, identify the intended referent before the shape was named. We hypothesized that asynchronous and antiphase synchronous musical activities, which require self-other distinction, might have stronger effects on perspective-taking than synchronous activity. Children in the asynchrony and antiphase synchrony conditions, but not the synchrony condition, showed anticipatory looks at the target, demonstrating real-time use of the partner's perspective. Study 2 was conducted to determine if asynchrony and antiphase asynchrony resulted in perspective-taking that otherwise would not have been observed, or if synchronous coordination inhibited perspective-taking that would otherwise have occurred. We found no evidence for online perspective-taking in 4- to 6-year-old children without music manipulation. Therefore, playing instruments asynchronously or in alternation, but not synchronously, increases perspective-taking in children of this age, likely by training self-other distinction and control. A video abstract of this article can be viewed at https://youtu.be/TM9h_GpFlsA. RESEARCH HIGHLIGHTS: This study is the first to show that rhythmic coordination, a form of non-linguistic interaction, can affect children's performance in a subsequent linguistic task. Eye-movement data revealed that children's perspective-taking in language processing was facilitated by prior asynchronous and antiphase synchronous musical interactions, but not by synchronous coordination. The results challenge the common "similar is better" view, suggesting that maintaining self-other distinction may benefit social interactions that involve representing individual differences.
While recent studies find that contextual diversity (CD) is a better determinant of visual word recognition than token frequency, there is a dearth of work comparing contextual diversity and token frequency in developing readers. In two sets of character and lexical decision experiments we examined token frequency and contextual diversity effects for fourth-grade children in Chinese. Experiments 1a and 1b used Chinese characters and words from SUBTLEX-CH. Experiments 2a and 2b used characters and words from a new corpus developed from Chinese primary school textbooks and reading materials from Grade 1 to 4. In both sets of experiments, CD affected character and lexical decision times but token frequency did not. The results are discussed in terms of recent context-based accounts of word learning and lexical processing, and implications are presented for models of skilled and developing reading.
Purpose The purpose of the current study was to examine the lexical and pragmatic factors that may contribute to turn-by-turn failures in communication (i.e., miscommunication) that arise regularly in interactive communication. Method Using a corpus from a collaborative dyadic building task, we investigated what differentiated successful from unsuccessful communication and potential factors associated with the choice to provide greater lexical information to a conversation partner. Results We found that more successful dyads' language tended to be associated with greater lexical density, lower ambiguity, and fewer questions. We also found participants were more lexically dense when accepting and integrating a partner's information (i.e., grounding) but were less lexically dense when responding to a question. Finally, an exploratory analysis suggested that dyads tended to spend more lexical effort when responding to an inquiry and used assent language accurately—that is, only when communication was successful. Conclusion Together, the results suggest that miscommunication both emerges and benefits from ambiguous and lexically dense utterances.
A classic problem in spoken language comprehension is how listeners perceive speech as being composed of discrete words, given the variable time-course of information in continuous signals. We propose a syllable inference account of spoken word recognition and segmentation, according to which alternative hierarchical models of syllables, words, and phonemes are dynamically posited, which are expected to maximally predict incoming sensory input. Generative models are combined with current estimates of context speech rate drawn from neural oscillatory dynamics, which are sensitive to amplitude rises. Over time, models which result in local minima in error between predicted and recently experienced signals give rise to perceptions of hearing words. Three experiments using the visual world eye-tracking paradigm with a picture-selection task tested hypotheses motivated by this framework. Materials were sentences that were acoustically ambiguous in numbers of syllables, words, and phonemes they contained (cf. English plural constructions, such as "saw (a) raccoon(s) swimming," which have two loci of grammatical information). Time-compressing, or expanding, speech materials permitted determination of how temporal information at, or in the context of, each locus affected looks to, and selection of, pictures with a singular or plural referent (e.g., one or more than one raccoon). Supporting our account, listeners probabilistically interpreted identical chunks of speech as consistent with a singular or plural referent to a degree that was based on the chunk's gradient rate in relation to its context. We interpret these results as evidence that arriving temporal information, judged in relation to language model predictions generated from context speech rate evaluated on a continuous scale, informs inferences about syllables, thereby giving rise to perceptual experiences of understanding spoken language as words separated in time.
We investigated whether fine-grained coordination in a screenbased puzzle task with a (virtual) partner would influence online perspective-taking. Participants played a screen-based puzzle game with a computer player. In the high-coordination condition, the player presented participants with puzzle pieces that could be placed near their partner’s last piece. In the lowcoordination condition, pieces could only be placed further away from their partner’s last piece. Participant’s eye movements were then measured in a referential communication task, with the partner giving the instructions, and whether possible competitor referents were in shared or privileged ground. The results demonstrate clear effects of ground and coordination. Participants in both coordination groups were sensitive to the perspective of the interlocutor. In addition, participants in the high-level coordination condition were more sensitive to statistical regularities in the input and their comprehension was more time-locked to the utterance of the speaker.
Incremental Referential Domain Circumscription during Processing of Natural and Synthesized Speech Mary D. Swift (mswift@ling.rochester.edu) Department of Linguistics, University of Rochester Rochester, NY 14627 Ellen Campana (ecampana@bcs.rochester.edu) Department of Brain and Cognitive Sciences, University of Rochester Rochester, NY 14627 James F. Allen (james@cs.rochester.edu) Department of Computer Sciences, University of Rochester Rochester, NY 14627 Michael K. Tanenhaus (mtan@bcs.rochester.edu) Department of Brain and Cognitive Sciences, University of Rochester Rochester, NY 14627 Abstract We present experimental evidence from a study in which we monitor eye movements as people respond to pre-recorded instructions generated by a human speaker and by two text-to- speech synthesizers. We replicate findings demonstrating that people process spoken language incrementally, making partial commitments as the instruction unfolds. Specifically, they establish different referential domains on the fly depending on whether a definite or indefinite article is used. Importantly, incremental understanding is observed for both natural speech instructions and synthesized text-to-speech instructions. These results, including some suggestive differences in responses with the two text-to-speech systems, establish the potential for using eye-tracking as a new method for fine-grained evaluation of dialogue systems and for using dialogue systems as a theoretical and experimental tool for psycholinguistic experimentation. Background Rapid increases in the accuracy and speed of automatic speech recognition and the increased availability of off-the- shelf text-to-speech systems has fueled great interest in spoken dialogue systems (e.g., Allen, Byron, Dzikovska, Ferguson, Galescu & Stent, 2001; Zue, Seneff, Glass, Polifroni, Pao, Hazen & Hetherington, 2000). As the sophistication of such systems increases, we can expect applications to more open-ended domains with larger vocabularies and more varied utterance types. The feasibility of such systems raises both applied and theoretical issues for work on natural language processing that crosses disciplinary boundaries. We focus on two issues here. The first, a computational issue, addresses the need for developing better evaluation tools for dialogue systems, especially tools that can evaluate comprehension on an utterance-by-utterance and within-utterance basis. The second, a psycholinguistic issue, is the possibility that in the near future implemented dialogue systems could serve as a powerful tool for developing and testing psycholinguistic models by allowing stimuli to be generated ‘on the fly,’ conditioned on the current state of the discourse. A necessary prerequisite for enabling both of these goals is that people respond to synthesized speech in much the same way as they do to natural speech. We present experimental evidence from a study in which we monitor eye movements as people respond to pre-recorded instructions generated by a human speaker and by two text- to-speech synthesizers. We replicate findings demonstrating that people process spoken language incrementally, making partial commitments as the instruction unfolds. More specifically, listeners establish referential domains on the fly depending on whether a definite or indefinite article is used. Eye movements as an evaluation tool Spoken utterances unfold over time, resulting in a stream of temporary ambiguities. For example, as the instruction Click on the beaker unfolds, the word beaker is briefly consistent with multiple candidates, including beetle, beeper, and speaker. Numerous psycholinguistic studies demonstrate that people comprehend utterances continuously, entertaining multiple lexical candidates (e.g., Marslen- Wilson, 1987), making provisional commitments at points of syntactic ambiguity, and resolving reference incrementally (e.g., Altmann, 1998; Tanenhaus & Trueswell, 1995). Recent studies using eye movements to a task-relevant object in a visual workspace as people follow spoken instructions provide striking evidence for both incremental understanding and rapid integration of multiple constraints (Tanenhaus, Spivey-Knowlton, Eberhard & Sedivy, 1995; 1996; Tanenhaus, Magnuson & Chambers, forthcoming). For example, if the instruction Click on a beaker is presented in a context in which there are two icons of beakers and two icons of beetles, then reference will be delayed until the word beaker is disambiguated phonetically
Processing language requires integrating information from multiple sources, including context, world knowledge, and the linguistic signal itself. How is this information integrated? A range of positions on the issue is possible, spanned by two extreme positions: extreme informational privilege—certain types of information are processed earlier in online processing and weighted most heavily in the resulting utterance interpretation; and extreme parallelism—all information is processed in parallel and weighted equally in the resulting interpretation. In reviewing the current empirical landscape on scalar implicature processing, the chapter argues for a constraint-based approach to pragmatic processing, which is closer in spirit to the parallelism account than the informational privilege account. The approach is also extended to other pragmatic phenomena.
Immediate Integration of Syntactic and Referential Constraints on Spoken Word Recognition James S. Magnuson (magnuson@psych.columbia.edu) Department of Psychology, Columbia University 1190 Amsterdam Ave., MC 5501 New York, NY 10027 USA Michael K. Tanenhaus (mtan@bcs.rochester.edu) and Richard N. Aslin (aslin@cvs.rochester.edu) Department of Brain & Cognitive Sciences, University of Rochester Rochester, NY 14627 USA Abstract We tested the hypothesis that syntactic constraints on spoken word recognition are integrated immediately when they are highly predictive. We used an artificial lexicon paradigm to create a lexicon of nouns (referring to shapes) and adjectives (referring to textures). Each word had phonological competitors in both form classes. We created strong form class expectations by using visual displays that either required adjective use or made adjectives infelicitous. We found evidence for immediate integration of form class expectations based on the pragmatic visual cues: similar-sounding words competed when they were from the same form class, but not when they were from different form classes. Top-down constraints on word recognition It is clear that we integrate top-down information when we interpret language. If someone tells us they put money in a bank, we understand that their money is in a vault and not buried next to a river. What is less clear is when and how we integrate top-down knowledge with bottom-up linguistic input. One possibility is that language is processed in stages, with top-down information integrated after an encapsulated first-pass on the bottom-up input (e.g., Frazier & Clifton, 1996; Norris, McQueen & Cutler, 2000). The theory behind this genre of model is that optimal efficiency can be achieved by applying automatic processes that will almost always yield a correct result. In the rare event that the automatic result cannot be reconciled with top-down information, reanalysis would be required. A second possibility is that top-down constraints are integrated immediately, with weights proportional to their predictive power (e.g., McClelland & Elman, 1986; MacDonald, Pearlmutter & Seidenberg, 1994; Tanenhaus & Trueswell, 1994). The theory behind constraint-based approaches is that a system can be made more efficient by allowing any sufficiently predictive information source to be integrated with processing as soon as it is relevant. While a variety of results support constraint-based theories of sentence processing (see MacDonald et al., 1994), there is reason to believe that spoken word recognition is initially encapsulated from top-down constraints. Swinney (1979) and Tanenhaus, Leiman & Seidenberg (1979) provided the seminal results on this issue by examining whether all homophones are activated independent of context. Tanenhaus et al. presented participants with spoken sentences that ended with a syntactically ambiguous word (e.g., “they all rose” vs. “they bought a rose”). If participants were asked to name a visual target immediately at the offset of the ambiguous word, priming was found for associates both of the alternative suggested by the context (e.g., “stood” given “they all rose”) and of homophones that would not fit the syntactic frame (e.g., “flower”). Given a 200-ms delay prior to the presentation of the visual stimulus, priming was found only for associates of the syntactically appropriate word. This suggests that lexical activation is initially based only on bottom-up information, and top-down information is a relatively late-acting constraint. Tanenhaus & Lucas (1987) argued that this made sense given the predictive power of a form-class expectation. Knowing that the next word will be one of tens of thousands of nouns would afford virtually no advantage for most nouns (those without homophones in different form classes). Furthermore, expectations for classes like noun or verb might be very weak because modifiers can almost always be inserted before either class (e.g., “they just rose”, “they bought a very pretty red rose”; cf. Shillcock & Bard, 1993). Shillcock & Bard (1993) pointed out that there are form classes that should be more predictive than noun or verb, because they have few members: those made up of closed-class words. They examined whether /wUd/ in a sentence context favoring the closed-class item, “would” (e.g., “John said that he didn’t want to do the job, but his brother would, as we later found out”), would prime associates of its homophone, “wood”, such as “timber” (compared with a context like “John said he didn’t want to do the job with his brother’s
We examined how naïve conversational participants circumscribed referential domains during the production and comprehension of referring expressions by monitoring participants' eye movements during a referential communication task. The results replicated some well-established results, e.g., incremental reference resolution, demonstrating the feasibility of studying real-time language comprehension in interactive conversation. We also observed a high proportion of underspecified referential expressions that were easily understood by addresses because of discourse and pragmatic constraints, including constraints developed as a result of the conversation.
In a block-assembly task with 138, 4-year-old Chinese kindergarten children, tested in pairs, we manipulated whether fine-grained coordination was required for accomplishing a shared goal with the same end product: building two adjoined towers with alternating levels of orange and green colored blocks to match a depicted model. In the coordination condition, each child had blocks of only one color and built the towers together. In the shared-goal-only condition, each child had both color blocks and built one of the towers, which they then adjoined. We predicted that children in the coordination condition would be more prosocial than children in the shared-goal-only condition. Studies with Western children typically find that girls are more generous than boys. However, we predicted the opposite pattern because Chinese culture emphasizes the importance of generosity more for males than females. Children in the coordination condition were more willing to help their partner complete an unrelated task and were more generous in sharing stickers with unknown children in a dictator game. These results demonstrate that level of coordination affects prosociality above and beyond having a shared goal, and are the first demonstration that prosocial effects of a collaborative task with children generalize beyond the participants to anonymous strangers. Boys shared more stickers with unknown children than girls, suggesting that gender differences in generosity are, in part, culturally conditioned.