There is growing recognition that short-term changes in speech perception influence speech production. These effects shed light on phonetic convergence – the subtle alignment of speech patterns that emerges between communication partners – and offer insights into interactions of perception and production. Across three experiments, we investigate the representations underlying perceptual effects on speech production. Building from the strong influence of preceding context on speech perception, we strategically pair contexts and target syllables to enhance or diminish target syllables’ perceptual distinctiveness and examine the impact of these perceptual effects on speech production. Experiment 1 shows that speech contexts rich in articulatory-phonetic information shift speech perception and alter acoustic patterns of speech production. Experiment 2 demonstrates that continuous natural speech filtered to possess subtly different spectral profiles that do not impact articulatory-phonetic information also impact both perception and production. Strikingly, Experiment 3 reveals that even nonspeech tones induce perceptual context effects that influence speech production. The findings point to a much broader scope of perception-production transfer than reported previously, and challenge the necessity of social interaction, covert imitation, and articulatory-phonetic information in sensorimotor speech interactions. This points to the need to extend models of speech motor control to account for perceptual influences of other talkers’ speech on speech production, and to accommodate general auditory processes in perception-production interactions.
Past research has shown that short-term exposure to speech carrying certain acoustic statistics transfers robustly to speech production. However, all studies reporting such transfer have used auditory repetition tasks. Therefore, it is unclear whether perception-production transfer in the acoustic-phonetic domain extends to tasks without an auditory model to probe production. We answer this question in two experiments. Experiment 1 shows that people read aloud the words BEER and PEER differently after exposure to auditory samples of “beer” and “peer” drawn from a distribution of standard American English vs. a distribution of slightly accented speech. Experiments 2A and 2B replicate this finding and show generalization to reading a new word pair (BEACH/PEACH) and a new nonword pair (BEETH/PEETH). Collectively, these results demonstrate that the perception-production transfer in the acoustic-phonetic domain extends beyond auditory repetition tasks to production tasks without an explicit auditory model, and that this transfer generalizes to new syllables.
Computational models of auditory salience predict that acoustic change and divergence from prediction increase the salience of sound streams. Confirming these predictions, prior research has shown that acoustic change and unpredictable sound features are linked to increases in physiological arousal and disruption of concurrent task performance. However, it remains unclear whether linguistic features, such as phonemic and lexical/semantic surprisal, help drive attentional orienting, or whether instead attentional capture takes place prior to linguistic analysis. To address this question, we introduce a new technique for assessing attentional capture by naturalistic task-irrelevant speech. In this paradigm participants tap to a metronome while ignoring a spoken passage from an audiobook. Salient features of the task-irrelevant speech capture attention, increase arousal, and expand subjective time, leading to shifts in tap timing. We show that distortions of subjective time are driven not only by acoustic change but also by phonemic surprisal. Thus, attentional orienting to sound takes place after the initial stages of linguistic analysis.
Statistical learning (SL) is typically assumed to be a core mechanism by which organisms learn covarying structures and recurrent patterns in the environment, with the main purpose of facilitating processing of expected events. Within this theoretical framework, the environment is viewed as relatively stable, and SL "captures" the regularities therein through implicit unsupervised learning by mere exposure. Focusing primarily on language-the domain in which SL theory has been most influential-we review evidence that the environment is far from fixed: It is dynamic, in continual flux, and learners are far from passive absorbers of regularities; they interact with their environments, thereby selecting and even altering the patterns they learn from. We therefore argue for an alternative cognitive architecture, where SL serves as a subcomponent of an information foraging (IF) system. IF aims to detect and assimilate novel recurrent patterns in the input that deviate from randomness, for which SL supplies a baseline. The broad implications of this viewpoint and their relevance to recent debates in cognitive neuroscience are discussed. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
Our ability to predict upcoming events is a fundamental component of human cognition. One way in which we do so is by exploiting temporal regularities in sensory signals: the ticking of a clock, falling of footsteps and the motion of waves each provide a structure that may facilitate anticipation. But how strong is the effect of rhythmic anticipation on perception? And to what degree do people vary in their ability to capitalize on these regularities? In 2015, Hickok et al. introduced a behavioural paradigm to assess how a rhythmic auditory stimulus affects perception of subsequent targets (Hickok G, Farahbod H, Saberi K. 2015 The rhythm of perception: entrainment to acoustic rhythms induces subsequent perceptual oscillation. Psychol. Sci. 26, 1006–1013. (doi:10.1177/0956797615576533)). They tested five listeners and found that perception (target detection accuracy) fluctuated rhythmically just like the sound rhythm. Here, we replicate the original finding, assess how likely the finding is to be observed for any individual, and quantify effect size in a large sample of adult listeners (n = 149). We introduce a model-based analysis approach that allows separate estimates of amplitude and phase information in target detection responses, and quantifies effect size for individual listeners. Together our results strongly support the presence of oscillatory influences on target detection accuracy, as well as substantial variability in the magnitude of this effect across listeners.
Humans and other animals use information about how likely it is for something to happen. The absolute and relative probability of an event influences a remarkable breadth of behaviors, from foraging for food to comprehending linguistic constructions -- even when these probabilities are learned implicitly. It is less clear how, and under what circumstances, statistical learning of simple probabilities might drive changes in perception and cognition. Here, across a series of 29 experiments, we probe listeners' sensitivity to task-irrelevant changes in the probability distribution of tones' acoustic frequency across tone-in-noise detection and tone duration decisions. We observe that the task-irrelevant frequency distribution influences the ability to detect a sound and the speed with which perceptual decisions about its duration are made. The shape of the probability distribution, its range, and a tone's relative position within that range impact observed patterns of suppression and enhancement of tone detection and decision making. Perceptual decisions are also modulated by a newly discovered perceptual bias, with lower frequencies in the distribution more often and more rapidly perceived as longer, and higher frequencies as shorter. Perception is sensitive to rapid distribution changes, but distributional learning from previous probability distributions also carries over. In fact, massed exposure to a single point along the dimension results in seemingly maladaptive loss of sensitivity - occurring entirely in the absence of feedback or reward - along a range of subsequently encountered frequencies. This points to a gain mechanism that suppresses sensitivity to regions along a perceptual dimension that are less likely to be encountered.
There is considerable lab-based evidence for successful incidental learning, in which a learner's attention is directed away from the to-be-learned stimulus and towards another stimulus. In this study, we extend incidental learning research into the language learning classroom. Three groups of adult second language (L2) learners (N = 52) engaged in structured classroom Mandarin learning took part in an 8-week study. One group served as a classroom-only control group. The second group underwent additional intentional auditory training involving Mandarin speech and explicit feedback. The third group underwent additional incidental learning combined with nonspeech "perceptual building block" categories-categories that share critical perceptual dimensions with target L2 speech categories but that are not perceived as speech. We demonstrate that when supplemented with structured classroom learning, incidental learning involving nonspeech analogs promotes phonetic, category, and word learning equivalent to learning from more traditional intentional auditory training.
Speech conveys both linguistic messages and a wealth of social and identity information about a talker. This information arrives as complex variations across many acoustic dimensions. Ultimately, speech communication depends on experience within a language community to develop shared long-term knowledge of the mapping from acoustic patterns to the category distinctions that support word recognition, emotion evaluation, and talker identification. A great deal of research has focused on the learning involved in acquiring long-term knowledge to support speech categorization. Inadvertently, this focus may give the impression of a mature learning endpoint. Instead, there seems to be no firm line between perception and learning in speech. The contributions of acoustic dimensions are malleably reweighted continuously as a function of regularities evolving in short-term input. In this way, continuous learning across speech impacts the very nature of the mapping from sensory input to perceived category. This article presents a case study in understanding how incoming sensory input—and the learning that takes place across it—interacts with existing knowledge to drive predictions that tune the system to support future behavior.
Listening to another speaker’s voice can lead to predictable changes in the listener’s own voice. This means that perception can alter production. A key question is whether overt production and its auditory consequences are critical for observing such changes. We answer this question in two experiments (N = 269) by passively exposing participants to speech that carries different acoustic patterns and investigating changes to production. Experiment 1 shows that decreasing the number of productions by an order of magnitude does not decrease the influence of perception on production. Experiment 2 takes this further by demonstrating that perceptual influence manifests on the very first overt production after exposure to new speech regularities. Collectively, these results show that perception can alter production without relying on feedback from overt production and its auditory consequences. This finding, in turn, points strongly to the need for extending current speech production models to include internal simulations.
Perception changes rapidly and implicitly as a function of passive exposure to speech that samples different acoustic distributions. Past research has shown that this statistical learning generalizes across talkers and, to some extent, new items, but these studies involved listeners’ active engagement in processing statistics-bearing stimuli. In this study, we manipulated the relationship between voice onset time (VOT) and fundamental frequency (F0) to establish distributional regularities either aligned with American English or reversed to create a subtle foreign accent. We then tested whether statistical learning across passive exposure to these distributions generalized to new items never experienced in the accent. Experiment 1 showed statistical learning across passive exposure but no generalization of learning when exposure and test items shared the same initial consonant but differed in vowels (bear/pear → beer/pier) or when they differed in initial consonant but shared distributional regularities across VOT and F0 dimensions (deer/tear → beer/pier). Experiment 2 showed generalization to stimuli that shared the statistics-bearing phoneme (bear/pear → beer/pier), but only when the response set included tokens from both exposure and generalization stimuli. Moreover, statistical learning transferred to influence the subtle acoustics of listeners’ own speech productions but did not generalize to influence productions of stimuli not heard in the accent. In sum, passive exposure is thus sufficient to support statistical learning and its generalization, but task demands modulate this dynamic. Moreover, production does not simply mirror perception: generalization in perception was not accompanied by transfer to production.
Humans implicitly pick up on probabilities of stimuli and events, yet it remains unclear how statistical learning builds expectations that affect perception. Across 29 experiments, we examine the influence of task-irrelevant distributions-defined across acoustic frequency-on both tone detection in noise and tone duration judgments. The shape and range of the frequency distributions impact suppression and enhancement effects, as does a given tone's position within the range. Perception adapts quickly to changing distributions, but past distributions influence future judgments. Massed exposure to a single frequency impacts perception along a range of subsequently encountered frequencies. A novel bias emerges as well: lower frequencies are perceived as longer and higher ones as shorter. Probability-driven learning dynamically shapes perception, driven by interacting influences of sensory processing, distributional learning, and selective attention that sculpt a gain function involving modest enhancement of more-likely stimuli, and robust suppression of less-likely stimuli.
Humans and other animals develop remarkable perceptual and cognitive specializations for identifying, differentiating, and acting on classes of ecologically important signals. This expertise is flexible enough to support diverse perceptual judgments: a voice, for example, simultaneously conveys what a talker says, as well as myriad cues about her identity and state. Expert perception across complex signals thus involves discovering and learning regularities that best inform diverse perceptual judgments, as well as weighting this information flexibly as task demands change. Here, we test whether this flexibility may involve endogenous attentional gain. We use two prospective auditory category learning tasks to relate a complex, entirely novel soundscape to four classes of "alien identity" and two classes of "alien size." Identity, but not size, categorization requires discovery and learning of patterned acoustic input situated in one of two simultaneous, non-overlapping frequency bands. This allows us to capitalize on the coarsely segregated frequency-band-specific channels tiling auditory cortex, using fMRI to ask whether category-relevant perceptual information present in one frequency band is prioritized relative to simultaneous, uninformative information in the other frequency band. Among participants expert at alien identity categorization, we observe prioritization of the identity-diagnostic frequency band that persists even when the diagnostic information becomes irrelevant in the size categorization task. Tellingly, the neural selectivity evoked implicitly in the identity categorization task aligns with that in an independent task, where activation is driven by explicit and sustained selective attention to pure tones in one or the other frequency band. Additionally, the learning trajectories taken to achieve expert-level categorization leave fingerprints on the patterns of neural activity associated with the diagnostic dimension. In all, this indicates that acquiring categories can drive the emergence of acquired attentional gain to category-diagnostic input dimensions.
Efficient behavior is supported by humans' ability to rapidly recognize acoustically distinct sounds as members of a common category. Within the auditory cortex, critical unanswered questions remain regarding the organization and dynamics of sound categorization. We performed intracerebral recordings during epilepsy surgery evaluation as 20 patient-participants listened to natural sounds. We then built encoding models to predict neural responses using sound representations extracted from different layers within a deep neural network (DNN) pretrained to categorize sounds from acoustics. This approach yielded accurate models of neural responses throughout the auditory cortex. The complexity of a cortical site's representation (measured by the depth of the DNN layer that produced the best model) was closely related to its anatomical location, with shallow, middle, and deep layers associated with core (primary auditory cortex), lateral belt, and parabelt regions, respectively. Smoothly varying gradients of representational complexity existed within these regions, with complexity increasing along a posteromedial-to-anterolateral direction in core and lateral belt and along posterior-to-anterior and dorsal-to-ventral dimensions in parabelt. We then characterized the time (relative to sound onset) when feature representations emerged; this measure of temporal dynamics increased across the auditory hierarchy. Finally, we found separable effects of region and temporal dynamics on representational complexity: sites that took longer to begin encoding stimulus features had higher representational complexity independent of region, and downstream regions encoded more complex features independent of temporal dynamics. These findings suggest that hierarchies of timescales and complexity represent a functional organizational principle of the auditory stream underlying our ability to rapidly categorize sounds.
Humans implicitly extract statistical regularities from sensory input across multiple perceptual dimensions. In audition, both transitional probabilities (relationships between successive sounds) and global probabilities (overall frequency of occurrence) shape perception. While each has been extensively studied in isolation, their combined influence remains less understood, especially in active tasks. Here, we investigated how these two forms of statistical structure jointly influence perceptual judgments. Participants heard a 70 ms cue tone followed by a 50 or 90 ms target tone and judged whether the target tone was short or long. Crucially, statistical regularities were embedded along a task-irrelevant dimension: tone frequency. This allowed us to assess how implicit statistical learning unfolds when regularities are orthogonal to task goals. Across three experiments, high transitional probabilities between cue and target frequencies facilitated faster duration judgments, independent of whether frequencies matched. High global probability also sped responses. However, when both regularities co-occurred, the effect of transitional probability is eliminated. Additionally, a direct sensory match between cue and target frequencies provided no perceptual advantage. These findings demonstrate that, even in active tasks demanding attention to a specific dimension, listeners implicitly learn statistical structure along irrelevant dimensions. While both transitional and global probabilities shape perception, global regularities dominate when statistics co-occur. This work advances our understanding of how layered statistical regularities guide behavior in goal-directed listening contexts.
Multilingual speakers can find speech recognition in everyday environments like restaurants and open-plan offices particularly challenging. In a world where speaking multiple languages is increasingly common, effective clinical and educational interventions will require a better understanding of how factors like multilingual contexts and listeners’ language proficiency interact with adverse listening environments. For example, word and phrase recognition is facilitated when competing voices speak different languages. Is this due to a “release from masking” from lower-level acoustic differences between languages and talkers, or higher-level cognitive and linguistic factors? To address this question, we created a “one-man bilingual cocktail party” selective attention task using English and Mandarin speech from one bilingual talker to reduce low-level acoustic cues. In Experiment 1, 58 listeners more accurately recognized English targets when distracting speech was Mandarin compared to English. Bilingual Mandarin–English listeners experienced significantly more interference and intrusions from the Mandarin distractor than did English listeners, exacerbated by challenging target-to-masker ratios. In Experiment 2, 29 Mandarin–English bilingual listeners exhibited linguistic release from masking in both languages. Bilinguals experienced greater release from masking when attending to English, confirming an influence of linguistic knowledge on the “cocktail party” paradigm that is separate from primarily energetic masking effects. Effects of higher-order language processing and expertise emerge only in the most demanding target-to-masker contexts. The “one-man bilingual cocktail party” establishes a useful tool for future investigations and characterization of communication challenges in the large and growing worldwide community of Mandarin–English bilinguals.
Speech provides a rich context for understanding how cortical interactions with the basal ganglia contribute to unique human behaviors, but opportunities for direct human intracranial recordings across cortical-basal ganglia networks are rare. Here we have recorded electrocorticographic signals in the cortex synchronously with single units in the basal ganglia during awake neurosurgeries where participants spoke syllable repetitions. We have discovered that individual subthalamic nucleus (STN) neurons have transient (200 ms) spike-phase coupling (SPC) events with multiple cortical regions. The spike timing of STN neurons is locked to the phase of theta-alpha oscillations in the supramarginal and posterior superior temporal gyrus during speech planning and production. Speech sound errors occur when this STN-cortical interaction is delayed. Our results suggest that timely interactions between the STN and the posterior perisylvian cortex support auditory-motor coordinate transformation or phonological working memory during speech planning. These findings establish a framework for understanding cortical-basal ganglia interaction in other human behaviors, and additionally indicate that firing-rate based models are insufficient for explaining basal ganglia circuit behavior.
The human auditory system consists of both peripheral and central components, both of which play a role but contribute distinctly to overall auditory functioning and can be differentially impacted by pathophysiologic states. Hemispheric surgery (HS), a procedure used for the treatment of drug-resistant epilepsy, involves complete disconnection of the auditory cortex in the operative hemisphere, leaving hearing acuity (peripheral function) intact but having heavy implications for auditory processing (central function). The literature describing pre- and post-operative auditory processing abilities of individuals who have undergone HS is sparse, but the research available provides evidence that several central auditory processes including auditory spatial analysis and temporal processing may be impacted. Deficits noted in standardized testing within the clinical or research environment have concrete functional impacts that may be currently under-appreciated and could lead to under-utilization of appropriate therapeutic strategies and accommodations. This review describes the profile of central auditory processing abilities in patients who have undergone HS by synthesizing available literature and incorporating research in other clinical populations to help fill critical gaps in our understanding of how cerebral disconnection impacts the central auditory system.
Regions in the superior temporal sulcus and gyrus have been heavily implicated in voice-selective responses in human auditory cortex. Despite an apparent specialization for the encoding of human voice, research outside the auditory domain suggests that these areas likely participate in additional neural processes including speech processing and production. The aim of the current study was to combine results of electrophysiological recording and clinical stimulation mapping procedures in patients undergoing stereoelectroencephalography (sEEG) to explore potential functional heterogeneity in voice-encoding cortex. Both channels that demonstrated voice-encoding properties and channels critically implicated in language functioning were heavily concentrated in the left STG/S. Analysis of functional overlap revealed channels in the posterior STG/S that appear to be involved in both voice encoding and language. Strength of voice encoding in these functionally diverse sites was not significantly different from sites that were implicated in voice encoding alone. Our findings add to prior observations of functional heterogeneity in the STG/S and contribute to proposed models of speech perception. We discuss these results in the context of the utility of electrophysiological methods in mapping cortical networks and identifying regions essential for functioning.### Competing Interest StatementThe authors have declared no competing interest.
Prior lesion, noninvasive-imaging, and intracranial-electroencephalography (iEEG) studies have documented hierarchical, parallel, and distributed characteristics of human speech processing. Yet, there have not been direct, intracranial observations of the latency with which regions outside the temporal lobe respond to speech, or how these responses are impacted by task demands. We leveraged human intracranial recordings via stereo-EEG to measure responses from diverse forebrain sites during (i) passive listening to /bi/ and /pi/ syllables, and (ii) active listening requiring /bi/-versus-/pi/ categorization. We find that neural response latency increases from a few tens of ms in Heschl's gyrus (HG) to several tens of ms in superior temporal gyrus (STG), superior temporal sulcus (STS), and early parietal areas, and hundreds of ms in later parietal areas, insula, frontal cortex, hippocampus, and amygdala. These data also suggest parallel flow of speech information dorsally and ventrally, from HG to parietal areas and from HG to STG and STS, respectively. Latency data also reveal areas in parietal cortex, frontal cortex, hippocampus, and amygdala that are not responsive to the stimuli during passive listening but are responsive during categorization. Furthermore, multiple regions-spanning auditory, parietal, frontal, and insular cortices, and hippocampus and amygdala-show greater neural response amplitudes during active versus passive listening (a task-related effect). Overall, these results are consistent with hierarchical processing of speech at a macro level and parallel streams of information flow in temporal and parietal regions. These data also reveal regions where the speech code is stimulus-faithful and those that encode task-relevant representations.
Humans and machines rarely have access to explicit external feedback or supervision, yet manage to learn. Most modern machine learning systems succeed because they benefit from unsupervised data. Humans are also expected to benefit and yet, mysteriously, empirical results are mixed. Does unsupervised learning help humans or not? Here, we argue that the mixed results are not conflicting answers to this question, but reflect that humans self-reinforce their predictions in the absence of supervision, which can help or hurt depending on whether predictions and task align. We use this framework to synthesize empirical results across various domains to clarify when unsupervised learning will help or hurt. This provides new insights into the fundamentals of learning with implications for instruction and lifelong learning.