There is growing recognition that short-term changes in speech perception influence speech production. These effects shed light on phonetic convergence – the subtle alignment of speech patterns that emerges between communication partners – and offer insights into interactions of perception and production. Across three experiments, we investigate the representations underlying perceptual effects on speech production. Building from the strong influence of preceding context on speech perception, we strategically pair contexts and target syllables to enhance or diminish target syllables’ perceptual distinctiveness and examine the impact of these perceptual effects on speech production. Experiment 1 shows that speech contexts rich in articulatory-phonetic information shift speech perception and alter acoustic patterns of speech production. Experiment 2 demonstrates that continuous natural speech filtered to possess subtly different spectral profiles that do not impact articulatory-phonetic information also impact both perception and production. Strikingly, Experiment 3 reveals that even nonspeech tones induce perceptual context effects that influence speech production. The findings point to a much broader scope of perception-production transfer than reported previously, and challenge the necessity of social interaction, covert imitation, and articulatory-phonetic information in sensorimotor speech interactions. This points to the need to extend models of speech motor control to account for perceptual influences of other talkers’ speech on speech production, and to accommodate general auditory processes in perception-production interactions.
Learning a second language (L2) is challenging partly due to perceptual strategies inherited from learners' first language. For example, speakers of tone languages like Mandarin over-use pitch in English prosody perception and production. We developed a novel training paradigm to help Mandarin learners adopt more native-like strategies by enhancing their use of duration relative to pitch cues during prosody categorization. After prosody training, participants used duration more during phrase boundary categorization but showed no clear change for contrastive focus and lexical stress, suggesting that cue weighting training is most effective when targeting a feature's primary cue. The control group, who practiced English vocabulary, relied more on pitch in lexical stress categorization and phrase boundary production after training, suggesting that without targeted instruction, listeners default to existing strategies. Our findings demonstrate that although default strategies in L2 speech perception are difficult to resist, lifelong perceptual habits can be adjusted with training.
With growing interest in ecologically valid stimuli and the proliferation of recognition and classification algorithms, environmental sounds are increasingly a topic of interest and study. However, there is a lack of consensus as to which sound categories should be selected, and which exemplars should represent these categories. Indeed, many existing datasets include only one exemplar per sound class. This is not characteristic of the actual acoustic environment and can lead to brittle and unrepresentative computational solutions and experimental results. Importantly, the existing literature provides relatively little information about how human listeners perceive the recognizability, similarity and representativeness of acoustically and perceptually varying sounds. Here, we introduce the Environmental Sounds (EnviSounds) dataset. It provides data from over 1000 human listeners performing different tasks (recognition, similarity judgement, goodness of fit) on 53 commonly encountered environmental sound categories of human, natural and man-made sounds, each represented by 10 different exemplars. We provide behavioural indices of within-class similarity and goodness of exemplar for all samples. Furthermore, we include normative data for recognition and identification latency, accuracy, and imageability, enabling researchers to select items for experiments based on pre-defined criteria across these dimensions. With this dataset, environmental sounds can be considered at various levels of complexity: sounds produced by the same sources varying in their acoustics (within-class acoustic variability) and acoustically similar sounds produced by different sources (between-category confusion). The dataset can be used for specifying training sets with individual sound sources, defining the categorical boundaries, or providing additional variables to building prediction models.
Auditory selective attention, the ability to focus on specific sounds while ignoring competitors, enables communication in complex soundscapes. Though attention clearly modulates cortical responses to sound, whether and where this modulation occurs in subcortical structures remains disputed. Here, we use electroencephalography to record cortical and subcortical (auditory brainstem responses, ABRs) activity during a selective attention task. Human participants attend to a 3-note melody in one pitch range presented to one ear while ignoring a competing, interleaved melody in a different pitch range played to the other ear (Laffere et al., 2020, 2021). The melodies consist of pitch-evoking pseudo-tones formed by convolving a periodic impulse train with a brief tone pip. These stimuli allow us to measure both ABRs (elicited by each individual tone pip within a pseudo-note) and cortical responses (elicited by the onsets of pseudo-notes) simultaneously. We observed robust ABRs, but no evidence of modulation by attention. Conversely, cortical responses, measured by event related potentials (ERPs), demonstrated attentional modulation of the P1-N1 peak. We conclude that attentional modulation within the brainstem is not measurable in the well-defined peaks of the ABR, which themselves reflect processing up to the input stage to the inferior colliculus.
Computational models of auditory salience predict that acoustic change and divergence from prediction increase the salience of sound streams. Confirming these predictions, prior research has shown that acoustic change and unpredictable sound features are linked to increases in physiological arousal and disruption of concurrent task performance. However, it remains unclear whether linguistic features, such as phonemic and lexical/semantic surprisal, help drive attentional orienting, or whether instead attentional capture takes place prior to linguistic analysis. To address this question, we introduce a new technique for assessing attentional capture by naturalistic task-irrelevant speech. In this paradigm participants tap to a metronome while ignoring a spoken passage from an audiobook. Salient features of the task-irrelevant speech capture attention, increase arousal, and expand subjective time, leading to shifts in tap timing. We show that distortions of subjective time are driven not only by acoustic change but also by phonemic surprisal. Thus, attentional orienting to sound takes place after the initial stages of linguistic analysis.
Recent advances in fast fMRI now enable whole-brain imaging with a TR of ≤1 s, which has helped to rekindle interest in characterizing the blood oxygen level dependent hemodynamic response function (BOLD HRF). Recent studies in the visual system have found intra-areal differences in temporal response characteristics, as well as HRFs that were faster and narrower than predicted by standard models. The auditory system presents a unique challenge, in that neuronal populations must operate across timescales of microseconds to minutes, and the surface of auditory cortex in particular is intricately and heavily vascularized. Here, we used fast fMRI to characterise voxelwise auditory HRFs evoked by short naturalistic sounds, assessing HRF reproducibility and variability across sessions, participants, and independent datasets. Across two studies at 3 T, participants passively listened to short environmental sounds while fMRI and quantitative MRI data were acquired with a 1s TR. Voxelwise HRFs were estimated via novel cross-session alignment routine, and exclusion of large vascular contributions. We identified a diverse set of hemodynamically plausible response shapes, which were not consistently captured by standard HRF approaches. These responses were reproducible within participants across sessions and robust across two independent acquisitions. Using data-driven gamma models, we achieved stable estimates with relatively few runs, particularly in auditory temporal regions. Within auditory cortex, we observed reproducible spatial gradients in response timing and shape, with faster and higher magnitude responses in medial regions, and slower and lower magnitude responses laterally. Together, these findings demonstrate that auditory HRFs are diverse, reliable, and regionally specific, and highlight the value of fast fMRI and data-driven modelling for advancing interpretation of fMRI data.
Grid cells in human entorhinal cortex encode spatial layouts for real-world navigation, yet their role in conceptual navigation remains unclear. Here we show that mentally transforming tones within a purely auditory pitch-duration space engages spatial circuits, and that such resources are causally necessary. In Experiment 1, participants trained for five days to navigate through a purely conceptual pitch-duration auditory space, then underwent fMRI on Day 6. We observed a six-fold modulation of entorhinal BOLD signals aligned to each participant's trajectory angles, similar to grid-cell firing in physical space. Stronger grid-like coding predicted larger training-related gains. In Experiment 2, a new cohort performed the same task under either a spatial or non-spatial interference load. Only the spatial condition selectively disrupted performance on trials requiring mental "movement," indicating a causal reliance on spatial resources. These findings provide evidence that auditory conceptual transformations recruit--and depend on--spatial grid-like computations in the entorhinal-hippocampal system, pointing to a domain-general role for spatial coding in organizing new knowledge along continuous dimensions. ### Competing Interest Statement The authors have declared no competing interest.
Humans and other animals use information about how likely it is for something to happen. The absolute and relative probability of an event influences a remarkable breadth of behaviors, from foraging for food to comprehending linguistic constructions -- even when these probabilities are learned implicitly. It is less clear how, and under what circumstances, statistical learning of simple probabilities might drive changes in perception and cognition. Here, across a series of 29 experiments, we probe listeners' sensitivity to task-irrelevant changes in the probability distribution of tones' acoustic frequency across tone-in-noise detection and tone duration decisions. We observe that the task-irrelevant frequency distribution influences the ability to detect a sound and the speed with which perceptual decisions about its duration are made. The shape of the probability distribution, its range, and a tone's relative position within that range impact observed patterns of suppression and enhancement of tone detection and decision making. Perceptual decisions are also modulated by a newly discovered perceptual bias, with lower frequencies in the distribution more often and more rapidly perceived as longer, and higher frequencies as shorter. Perception is sensitive to rapid distribution changes, but distributional learning from previous probability distributions also carries over. In fact, massed exposure to a single point along the dimension results in seemingly maladaptive loss of sensitivity - occurring entirely in the absence of feedback or reward - along a range of subsequently encountered frequencies. This points to a gain mechanism that suppresses sensitivity to regions along a perceptual dimension that are less likely to be encountered.
Humans implicitly pick up on probabilities of stimuli and events, yet it remains unclear how statistical learning builds expectations that affect perception. Across 29 experiments, we examine the influence of task-irrelevant distributions-defined across acoustic frequency-on both tone detection in noise and tone duration judgments. The shape and range of the frequency distributions impact suppression and enhancement effects, as does a given tone's position within the range. Perception adapts quickly to changing distributions, but past distributions influence future judgments. Massed exposure to a single frequency impacts perception along a range of subsequently encountered frequencies. A novel bias emerges as well: lower frequencies are perceived as longer and higher ones as shorter. Probability-driven learning dynamically shapes perception, driven by interacting influences of sensory processing, distributional learning, and selective attention that sculpt a gain function involving modest enhancement of more-likely stimuli, and robust suppression of less-likely stimuli.
Humans and other animals develop remarkable perceptual and cognitive specializations for identifying, differentiating, and acting on classes of ecologically important signals. This expertise is flexible enough to support diverse perceptual judgments: a voice, for example, simultaneously conveys what a talker says, as well as myriad cues about her identity and state. Expert perception across complex signals thus involves discovering and learning regularities that best inform diverse perceptual judgments, as well as weighting this information flexibly as task demands change. Here, we test whether this flexibility may involve endogenous attentional gain. We use two prospective auditory category learning tasks to relate a complex, entirely novel soundscape to four classes of "alien identity" and two classes of "alien size." Identity, but not size, categorization requires discovery and learning of patterned acoustic input situated in one of two simultaneous, non-overlapping frequency bands. This allows us to capitalize on the coarsely segregated frequency-band-specific channels tiling auditory cortex, using fMRI to ask whether category-relevant perceptual information present in one frequency band is prioritized relative to simultaneous, uninformative information in the other frequency band. Among participants expert at alien identity categorization, we observe prioritization of the identity-diagnostic frequency band that persists even when the diagnostic information becomes irrelevant in the size categorization task. Tellingly, the neural selectivity evoked implicitly in the identity categorization task aligns with that in an independent task, where activation is driven by explicit and sustained selective attention to pure tones in one or the other frequency band. Additionally, the learning trajectories taken to achieve expert-level categorization leave fingerprints on the patterns of neural activity associated with the diagnostic dimension. In all, this indicates that acquiring categories can drive the emergence of acquired attentional gain to category-diagnostic input dimensions.
Understanding how the brain organises natural categories is a central challenge in neuroscience. While prior work has shown that categories can be decoded from distributed activity patterns in auditory cortex, it remains unclear how these categories are globally arranged relative to one another, and how low-level acoustic and higher-level semantic structure jointly shape this organisation. Here, we addressed these questions by deriving low-dimensional functional gradients from high-depth functional magnetic resonance imaging (fMRI) data (three participants, ~4.7 hours each) acquired during a category-specific one-back task. These gradients captured the principal axes of population activity in auditory cortex. Gradient-based models of the auditory cortex explained category structure more accurately than region-of-interest or whole-brain approaches, revealing that category information is distributed across multiple continuous axes rather than aligned with any single organisational dimension. Projecting acoustic (gammatone filter-bank) and behavioural similarity spaces directly into a shared framework with the fMRI functional axes showed that both contribute to the brain's category geometry, with acoustic structure exerting a somewhat stronger influence. However, representational relationships varied across category pairs: some reflected primarily acoustic similarity, others semantic distinctions, and many a combination of both. This pairwise heterogeneity shows how auditory cortex may integrate multiple representational dimensions that define higher-level categories.
The development of non-invasive methods to study brain structure and function has enabled a flowering of cognitive neuroscience in humans and nonhuman species. Herein, we describe the development of protocols for functional magnetic resonance imaging (fMRI) of a bottlenose dolphin (Tursiops truncatus), including protocols to monitor the health and welfare of the subject over the course of our five-year study. A Welfare Control Plan (WCP) was designed to monitor, enhance, and protect our subject’s welfare throughout the course of the study. The WCP was developed so our team of marine mammal veterinarians, trainers, and researchers could (1) identify study procedures that might negatively impact the individual’s welfare and propose measures to mitigate them, (2) define and implement protocols for monitoring the individual’s welfare throughout the study, and (3) determine the study’s temporary or final endpoints. Overall, behavioral, physiological, and health welfare indicators showed that the dolphin’s quality of life was not negatively impacted by participating in our functional neuroimaging study. Our study provides an example of how innovative, ambitious, and logistically complex animal studies can successfully be performed while protecting the welfare of participating animals through adequate planning, enough human and economic resources, and full human/institutional commitment to animal welfare.
Multilingual speakers can find speech recognition in everyday environments like restaurants and open-plan offices particularly challenging. In a world where speaking multiple languages is increasingly common, effective clinical and educational interventions will require a better understanding of how factors like multilingual contexts and listeners’ language proficiency interact with adverse listening environments. For example, word and phrase recognition is facilitated when competing voices speak different languages. Is this due to a “release from masking” from lower-level acoustic differences between languages and talkers, or higher-level cognitive and linguistic factors? To address this question, we created a “one-man bilingual cocktail party” selective attention task using English and Mandarin speech from one bilingual talker to reduce low-level acoustic cues. In Experiment 1, 58 listeners more accurately recognized English targets when distracting speech was Mandarin compared to English. Bilingual Mandarin–English listeners experienced significantly more interference and intrusions from the Mandarin distractor than did English listeners, exacerbated by challenging target-to-masker ratios. In Experiment 2, 29 Mandarin–English bilingual listeners exhibited linguistic release from masking in both languages. Bilinguals experienced greater release from masking when attending to English, confirming an influence of linguistic knowledge on the “cocktail party” paradigm that is separate from primarily energetic masking effects. Effects of higher-order language processing and expertise emerge only in the most demanding target-to-masker contexts. The “one-man bilingual cocktail party” establishes a useful tool for future investigations and characterization of communication challenges in the large and growing worldwide community of Mandarin–English bilinguals.
What factors determine the importance placed on different sources of evidence during speech and music perception? Attention-to-dimension theories suggest that, through prolonged exposure to their first language (L1), listeners become biased to attend to acoustic dimensions especially informative in that language. Given that selective attention can modulate cortical tracking of sound, attention-to-dimension accounts would predict that tone language speakers would show greater cortical tracking of pitch in L2 speech, even when it is not task-relevant, as well as an enhanced ability to attend to pitch in both speech and music. Here we test these hypotheses by examining neural sound encoding, dimension-selective attention and cue-weighting strategies in 54 native English and 60 Mandarin Chinese speakers. Our results show that Mandarin speakers, compared to native English speakers, are better at attending to pitch and worse at attending to duration in verbal and non-verbal stimuli; moreover, they place more importance on pitch and less on duration during speech and music categorization. The effects of language background were moderated by musical experience, however, with Mandarin-speaking musicians better able to attend to duration and using duration more as a cue to phrase boundary perception. There was no effect of L1 on cortical tracking of acoustic dimensions. Nevertheless, the frequency-following response to stimulus pitch was enhanced in Mandarin speakers, suggesting that speaking a tone language can boost processing of early pitch encoding. These findings suggest that tone language experience does not increase the tendency for pitch to capture attention, regardless of task; instead, tone language speakers may benefit from an enhanced ability to direct attention to pitch when it is task-relevant, without affecting pitch salience.
From auditory perception to general cognition, the ability to play a musical instrument has been associated with skills both related and unrelated to music. However, it is unclear if these effects are bound to the specific characteristics of musical instrument training, as little attention has been paid to other populations whose auditory expertise could match or surpass that of musicians in specific auditory tasks or more naturalistic acoustic scenarios. We explored this possibility by comparing conservatory-trained instrumentalists to students of audio engineering (along with naive controls) on measures of auditory discrimination, auditory scene analysis, and speech in noise perception. We found that both musicians and audio engineers had generally lower psychophysical thresholds than controls, with pitch perception showing the largest effect size. Musicians performed best in a sustained selective attention task with two competing streams of tones, while audio engineers could better memorise and recall auditory scenes composed of non-musical sounds, when compared to controls. Additionally, in a diotic speech-in-babble task, musicians showed lower signal-to-noise-ratio thresholds than both controls and engineers. We also observed differences in personality that might account for group-based self-selection biases. Overall, we showed that investigating a wider range of forms of auditory expertise can help us corroborate (or challenge) the specificity of the advantages previously associated with musical instrument training.
How does the brain track and process rapidly changing sensory information? Current computational accounts suggest that our sensations and decisions arise from the intricate interplay between bottom-up sensory signals and constantly changing expectations regarding the statistics of the surrounding world. A significant focus of recent research is determining which statistical properties are tracked by the brain as it monitors the rapid progression of sensory information. Here, by combining EEG (three experiments N ≥ 22 each) and computational modelling, we examined how the brain processes rapid and stochastic sound sequences that simulate key aspects of dynamic sensory environments. Passively listening participants were exposed to structured tone-pip arrangements that contained transitions between a range of stochastic patterns. Predictions were guided by a Bayesian predictive inference model. We demonstrate that listeners automatically track the statistics of unfolding sounds, even when these are irrelevant to behaviour. Transitions between sequence patterns drove a shift in the sustained EEG response. This was observed to a range of distributional statistics, and even in situations where behavioural detection of these transitions was at floor. These observations suggest that the modulation of the EEG sustained response reflects a process of belief updating within the brain. By establishing a connection between the outputs of the computational model and the observed brain responses, we demonstrate that the dynamics of these transition-related responses align with the tracking of "precision" - the confidence or reliability assigned to a predicted sensory signal - shedding light on the intricate interplay between the brain's statistical tracking mechanisms and its response dynamics.
The brain is increasingly viewed as a statistical learning machine, where our sensations and decisions arise from the intricate interplay between bottom-up sensory signals and constantly changing expectations regarding the surrounding world. Which statistics does the brain track while monitoring the rapid progression of sensory information?Here, by combining EEG (three experiments N≥22 each) and computational modelling, we examined how the brain processes rapid and stochastic sound sequences that simulate key aspects of dynamic sensory environments. Passively listening participants were exposed to structured tone-pip arrangements that contained transitions between a range of stochastic patterns. Predictions were guided by a Bayesian predictive inference model. We demonstrate that listeners automatically track the statistics of unfolding sounds, even when these are irrelevant to behaviour. Transitions between sequence patterns drove an increase of the sustained EEG response. This was observed to a range of distributional statistics, and even in situations where behavioural detection of these transitions was at floor. These observations suggest that the modulation of the EEG sustained response reflects a universal process of belief updating within the brain. By establishing a connection between the outputs of the computational model and the observed brain responses, we demonstrate that the dynamics of these transition-related responses align with the tracking of ‘precision’ – the confidence or reliability assigned to a predicted sensory signal - shedding light on the intricate interplay between the brain’s statistical tracking mechanisms and its response dynamics.### Competing Interest StatementThe authors have declared no competing interest.