Developmental dyslexia is a specific, highly prevalent and often debilitating reading and spelling disorder with unknown neurocomputational mechanisms. Here we discovered, in a preregistered functional magnetic resonance imaging study optimized for the subcortical sensory pathway, that dyslexia is characterized by altered predictive coding in left-hemispheric auditory sensory pathway nuclei. The neurocomputational alterations were related to one of the two main dyslexia risk scores, indicating a crucial role for dyslexia pathophysiology.
Responses in the sensory thalamic nuclei are modulated by perceptual tasks. Whether such response modulations rely on feedback from cerebral cortex in humans is unknown. Here, we addressed this question in the context of visual speech recognition: the visual sensory thalamus, i.e. the lateral geniculate nucleus (LGN), has differential BOLD-responses to visual speech than non-speech control tasks. We tested whether such response modulation relies on the function of the visual association cortex, specifically visual-motion area V5/MT. We applied inhibitory transcranial magnetic stimulation (TMS) over bilateral visual-motion sensitive areas V5/MT on 26 healthy adults. Subsequently, participants performed a visual speech and a colour recognition task on identical muted videos of speaking faces during functional magnetic resonance imaging (fMRI). The LGN showed a significant signal change between the visual speech task and the colour task following Vertex stimulation as active control region. This modulation was significantly reduced following inhibitory V5/MT stimulation. V5/MT stimulation also reduced task-dependent functional connectivity between V5/MT and the LGN. These results identify corticothalamic feedback as integral mechanism in visual processing. In particular, the visual association cortex has a causal role in modulating LGN responses during speech recognition.
How does the brain know what is out there and what is not? Living organisms cannot rely solely on sensory signals for perception because they are noisy and ambiguous. To transform sensory signals into stable percepts, the brain uses its prior knowledge or beliefs. Current theories describe perceptual beliefs as probability distributions over the features of the stimuli, summarised by their mean and variance. Beliefs are updated by feature prediction errors: the mismatch between expected and observed feature values. This framework explains how the brain encodes unexpected changes in stimulus features (e.g., higher or lower pitch, stronger or weaker motion). How the brain updates beliefs about a stimulus' presence or absence is, however, unclear. We propose that the detection of absence relies on a distinct form of prediction error dedicated to reducing the beliefs on stimulus occurrence. We call this signal absence prediction error. Using the human auditory system as a model for sensory processing, we developed a paradigm designed to test this hypothesis. fMRI results showed that absence prediction error is encoded in the auditory thalamus and cortex, indicating that absence is explicitly represented in subcortical sensory pathways. Moreover, while feature prediction error is already encoded in the auditory midbrain, absence prediction error was not, implying that absence-related error signals are supported by a different circuit. These results identify a neural mechanism for the detection of sensory absence. Such mechanisms may be disrupted in conditions such as psychosis, where predictions about absence and presence are impaired.
The long-standing hypothesis that autism is linked to changes in the visual magnocellular system of the human brain has never been directly examined due to technological constraints. Here, we used a recently developed 7-Tesla functional MRI (fMRI) approach to investigate this hypothesis within the visual sensory thalamus (lateral geniculate nucleus, LGN). The LGN is a crucial component of the primary visual pathway. It is particularly suited to investigate the magnocellular visual system, because within the LGN, the magnocellular (mLGN) uniquely segregates from the parvocellular (pLGN) system. Our results revealed diminished mLGN blood-oxygenation-level-dependent (BOLD) responses in the autism group compared to controls. pLGN responses were comparable across groups. The mLGN alterations were observed specifically for stimuli optimized for mLGN function, i.e., visual displays with low spatial frequency and high temporal flicker frequency. The results confirm the long-standing hypothesis of magnocellular visual system alterations in autism. They substantiate the emerging perspective that sensory processing variations are part of autism symptomatology.
Expectations aid and bias our perception. In speech, expected words are easier to recognise than unexpected words, particularly in noisy environments, and incorrect expectations can make us misunderstand our conversational partner. Expectations are combined with the output from the sensory pathways to form representations of speech in the cerebral cortex. However, it is unclear whether expectations are propagated further down to subcortical structures to aid the encoding of the basic dynamic constituent of speech: fast frequency-modulation (FM). Fast FM-sweeps are the basic invariant constituent of consonants, and their correct encoding is fundamental for speech recognition. Here we tested the hypothesis that subjective expectations drive the encoding of fast FM-sweeps characteristic of speech in the human subcortical auditory pathway. We used fMRI to measure neural responses in the human auditory midbrain (inferior colliculus) and thalamus (medial geniculate body). Participants listened to sequences of FM-sweeps for which they held different expectations based on the task instructions. We found robust evidence that the responses in auditory midbrain and thalamus encode the difference between the acoustic input and the subjective expectations of the listener. The results indicate that FM-sweeps are already encoded at the level of the human auditory midbrain and that encoding is mainly driven by subjective expectations. We conclude that the subcortical auditory pathway is integrated in the cortical network of predictive speech processing and that expectations are used to optimise the encoding of even the most basic acoustic constituents of speech.
The key assumption of the predictive coding framework is that internal representations are used to generate predictions on how the sensory input will look like in the immediate future. These predictions are tested against the actual input by the so-called prediction error units, which encode the residuals of the predictions. What happens to prediction errors, however, if predictions drawn by different stages of the sensory hierarchy contradict each other? To answer this question, we conducted two fMRI experiments while female and male human participants listened to sequences of sounds: pure tones in the first experiment and frequency-modulated sweeps in the second experiment. In both experiments, we used repetition to induce predictions based on stimulus statistics (stats-informed predictions) and abstract rules disclosed in the task instructions to induce an orthogonal set of (task-informed) predictions. We tested three alternative scenarios: neural responses in the auditory sensory pathway encode prediction error with respect to (1) the stats-informed predictions, (2) the task-informed predictions, or (3) a combination of both. Results showed that neural populations in all recorded regions (bilateral inferior colliculus, medial geniculate body, and primary and secondary auditory cortices) encode prediction error with respect to a combination of the two orthogonal sets of predictions. The findings suggest that predictive coding exploits the non-linear architecture of the auditory pathway for the transmission of predictions. Such non-linear transmission of predictions might be crucial for the predictive coding of complex auditory signals like speech.Significance Statement Sensory systems exploit our subjective expectations to make sense of an overwhelming influx of sensory signals. It is still unclear how expectations at each stage of the processing pipeline are used to predict the representations at the other stages. The current view is that this transmission is hierarchical and linear. Here we measured fMRI responses in auditory cortex, sensory thalamus, and midbrain while we induced two sets of mutually inconsistent expectations on the sensory input, each putatively encoded at a different stage. We show that responses at all stages are concurrently shaped by both sets of expectations. The results challenge the hypothesis that expectations are transmitted linearly and provide for a normative explanation of the non-linear physiology of the corticofugal sensory system.
Expectations can substantially influence perception. Predictive coding is a theory of sensory processing that aims to explain the neural mechanisms underlying the effect of expectations in sensory processing. Its main assumption is that sensory neurons encode prediction error with respect to expected sensory input. Neural populations encoding prediction error have been previously reported in the human auditory cortex (AC); however, most studies focused on the encoding of pure tones and induced expectations by stimulus repetition, potentially confounding prediction error with effects of neural habituation. Here, we systematically studied prediction error to pure tones and fast frequency modulated (FM) sweeps across different auditory cortical fields in humans. We conducted two fMRI experiments, each using one type of stimulus. We measured BOLD responses across the bilateral auditory cortical fields Te1.0, Te1.1, Te1.2, and Te3 while participants listened to sequences of sounds. We induced subjective expectations on the incoming sounds independently of stimulus repetition using abstract rules. Our results indicate that pure tones and FM-sweeps are encoded as prediction error with respect to the participants' expectations across auditory cortical fields. The topographical distribution of neural populations encoding prediction error to pure tones and FM-sweeps was highly correlated in left Te1.1 and Te1.2, and in bilateral Te3, suggesting that predictive coding is the general encoding mechanism in AC.
Autism spectrum disorder (ASD) is characterised by social communication difficulties. These difficulties have been mainly explained by cognitive, motivational, and emotional alterations in ASD. The communication difficulties could, however, also be associated with altered sensory processing of communication signals. Here, we assessed the functional integrity of auditory sensory pathway nuclei in ASD in three independent functional magnetic resonance imaging experiments. We focused on two aspects of auditory communication that are impaired in ASD: voice identity perception, and recognising speech-in-noise. We found reduced processing in adults with ASD as compared to typically developed control groups (pairwise matched on sex, age, and full-scale IQ) in the central midbrain structure of the auditory pathway (inferior colliculus [IC]). The right IC responded less in the ASD as compared to the control group for voice identity, in contrast to speech recognition. The right IC also responded less in the ASD as compared to the control group when passively listening to vocal in contrast to non-vocal sounds. Within the control group, the left and right IC responded more when recognising speech-in-noise as compared to when recognising speech without additional noise. In the ASD group, this was only the case in the left, but not the right IC. The results show that communication signal processing in ASD is associated with reduced subcortical sensory functioning in the midbrain. The results highlight the importance of considering sensory processing alterations in explaining communication difficulties, which are at the core of ASD.
Predictive coding is the leading algorithmic framework to understand how expectations shape our experience of reality. Its main tenet is that sensory neurons encode prediction error: the residuals between a generative model of the sensory world and the actual sensory input. However, it is yet unclear how this scheme generalises to the multi-level hierarchical architecture of sensory processing. Theoretical accounts of predictive coding agree that neurons computing prediction error and the generative model exist at all levels of the processing hierarchy. Some accounts hypothesise that generative models at each level inform prediction error only at the immediately lower level. Other accounts assume that predictions from the highest available generative model propagate downwards in the hierarchy overwriting subsequently lower generative models. Here we tested these two hypotheses in the auditory pathway. We used two paradigms where participants listened to sequences of either pure tones or FM-sweeps while we recorded BOLD responses in inferior colliculus (IC), medial geniculate body (MGB), and auditory cortex (AC). We used the task instructions to manipulate participants expectations on the incoming stimuli independently of the local stimulus statistics. We assumed that the auditory pathway would keep two generative models: one based on local stimulus statistics; and a model based on the subjective expectations induced by the task instruction. We used Bayesian model comparison to test whether neural responses in IC, MGB, and AC encoded prediction error with respect to either of the two generative models, or a combination of both. Results showed that neural populations in bilateral IC, MGB, and AC encode prediction error with respect to a combination of the two generative models, indicating that the predictive architecture of predictive coding is more complex than previously hypothesised.
Predictive processing, a leading theoretical framework for sensory processing, suggests that the brain constantly generates predictions on the sensory world and that perception emerges from the comparison between these predictions and the actual sensory input. This requires two distinct neural elements: generative units, which encode the model of the sensory world; and prediction error units, which compare these predictions against the sensory input. Although predictive processing is generally portrayed as a theory of cerebral cortex function, animal and human studies over the last decade have robustly shown the ubiquitous presence of prediction error responses in several nuclei of the auditory, somatosensory, and visual subcortical pathways. In the auditory modality, prediction error is typically elicited using so-called oddball paradigms, where sequences of repeated pure tones with the same pitch are at unpredictable intervals substituted by a tone of deviant frequency. Repeated sounds become predictable promptly and elicit decreasing prediction error; deviant tones break these predictions and elicit large prediction errors. The simplicity of the rules inducing predictability make oddball paradigms agnostic about the origin of the predictions. Here, we introduce two possible models of the organizational topology of the predictive processing auditory network: (1) the global view, that assumes that predictions on the sensory input are generated at high-order levels of the cerebral cortex and transmitted in a cascade of generative models to the subcortical sensory pathways; and (2) the local view, that assumes that independent local models, computed using local information, are used to perform predictions at each processing stage. In the global view information encoding is optimized globally but biases sensory representations along the entire brain according to the subjective views of the observer. The local view results in a diminished coding efficiency, but guarantees in return a robust encoding of the features of sensory input at each processing stage. Although most experimental results to-date are ambiguous in this respect, recent evidence favors the global model.
Frequency modulation (FM) is a basic constituent of vocalisation in many animals as well as in humans. In human speech, short rising and falling FM-sweeps of around 50 ms duration, called formant transitions, characterise individual speech sounds. There are two representations of FM in the ascending auditory pathway: a spectral representation, holding the instantaneous frequency of the stimuli; and a sweep representation, consisting of neurons that respond selectively to FM direction. To-date computational models use feedforward mechanisms to explain FM encoding. However, from neuroanatomy we know that there are massive feedback projections in the auditory pathway. Here, we found that a classical FM-sweep perceptual effect, the sweep pitch shift, cannot be explained by standard feedforward processing models. We hypothesised that the sweep pitch shift is caused by a predictive feedback mechanism. To test this hypothesis, we developed a novel model of FM encoding incorporating a predictive interaction between the sweep and the spectral representation. The model was designed to encode sweeps of the duration, modulation rate, and modulation shape of formant transitions. It fully accounted for experimental data that we acquired in a perceptual experiment with human participants as well as previously published experimental results. We also designed a new class of stimuli for a second perceptual experiment to further validate the model. Combined, our results indicate that predictive interaction between the frequency encoding and direction encoding neural representations plays an important role in the neural processing of FM. In the brain, this mechanism is likely to occur at early stages of the processing hierarchy.
The subcortical sensory pathways are the fundamental channels for mapping the outside world to our minds. Sensory pathways efficiently transmit information by adapting neural responses to the local statistics of the sensory input. The long-standing mechanistic explanation for this adaptive behaviour is that neural activity decreases with increasing regularities in the local statistics of the stimuli. An alternative account is that neural coding is directly driven by expectations of the sensory input. Here, we used abstract rules to manipulate expectations independently of local stimulus statistics. The ultra-high-field functional-MRI data show that abstract expectations can drive the response amplitude to tones in the human auditory pathway. These results provide first unambiguous evidence of abstract processing in a subcortical sensory pathway. They indicate that the neural representation of the outside world is altered by our prior beliefs even at initial points of the processing hierarchy.
Autism spectrum disorder (ASD) is a clinical condition that is associated with deficient processing of communication signals. To-date, most neuroscience research on ASD focuses on explaining these symptoms at the cerebral cortex level or in limbic structures. This research implicitly or explicitly assumes that the subcortical sensory pathways are intact in ASD. To date, however, the integrity of subcortical sensory pathway nuclei in ASD has never been investigated in humans in vivo. Here, we assessed the functional integrity of auditory sensory pathway nuclei in ASD in three independent functional magnetic resonance imaging (fMRI) experiments. We focused on two aspects of auditory communication that are impaired in ASD: voice identity perception, and recognising speech-in noise. Adults with ASD (n = 16 in experiments 1 and 2 and n = 17 in another sample in experiment 3) and typically developed control groups were pairwise matched on sex, age, handedness, and full-scale IQ. We found reduced processing in the ASD as compared to the control groups in the central midbrain structure of the auditory pathway (inferior colliculus, IC; all results FWE corrected and Bonferroni corrected for the number of regions tested). The right IC responded less in the ASD as compared to the control group for voice identity recognition, in contrast to a speech recognition task. The right IC also responded less in the ASD as compared to the control group when passively listening to vocal sounds in contrast to non-vocal sounds. For speech-in-noise recognition, we found that within the control group, the left and right IC responded more when performing a speech-in-noise recognition task as compared to when performing a speech recognition task without additional noise. In the ASD group, this was only the case in the left, but not the right IC. There was no interaction between noise and group. The results show that communication signal processing in ASD is associated with reduced subcortical sensory functioning in the midbrain. The results highlight the importance of considering sensory processing alterations in explaining social communication difficulties, which are at the core of ASD.
The subcortical sensory pathways are the fundamental channels for mapping the outside world to our minds. Sensory pathways efficiently transmit information by adapting neural responses to the local statistics of the sensory input. The longstanding mechanistic explanation for this adaptive behaviour is that neuronal habituation scales activity to the local statistics of the stimuli. An alternative account is that neural coding is directly driven by expectations of the sensory input. Here we used abstract rules to manipulate expectations independently of local stimulus statistics. The ultra-high-field functional-MRI data show that expectations, and not habituation, are the main driver of the response amplitude to tones in the human auditory pathway. These results provide first unambiguous evidence of predictive coding and abstract processing in a subcortical sensory pathway, indicating that the brain only holds subjective representations of the outside world even at initial points of the processing hierarchy.
Pitch is a fundamental attribute of auditory perception. The interaction of concurrent pitches gives rise to a sensation that can be characterized by its degree of consonance or dissonance. In this work, we propose that human auditory cortex (AC) processes pitch and consonance through a common neural network mechanism operating at early cortical levels. First, we developed a new model of neural ensembles incorporating realistic neuronal and synaptic parameters to assess pitch processing mechanisms at early stages of AC. Next, we designed a magnetoencephalography (MEG) experiment to measure the neuromagnetic activity evoked by dyads with varying degrees of consonance or dissonance. MEG results show that dissonant dyads evoke a pitch onset response (POR) with a latency up to 36 ms longer than consonant dyads. Additionally, we used the model to predict the processing time of concurrent pitches; here, consonant pitch combinations were decoded faster than dissonant combinations, in line with the experimental observations. Specifically, we found a striking match between the predicted and the observed latency of the POR as elicited by the dyads. These novel results suggest that consonance processing starts early in human auditory cortex and may share the network mechanisms that are responsible for (single) pitch processing.
Pitch is the perceptual correlate of sound's periodicity and a fundamental property of the auditory sensation. The interaction of two or more pitches gives rise to a sensation that can be characterized by its degree of consonance or dissonance. In the current study, we investigated the neuromagnetic representations of consonant and dissonant musical dyads using a new model of cortical activity, in an effort to assess the possible involvement of pitch-specific neural mechanisms in consonance processing at early cortical stages. In the first step of the study, we developed a novel model of cortical pitch processing designed to explain the morphology of the pitch onset response (POR), a pitch-specific subcomponent of the auditory evoked N100 component in the human auditory cortex. The model explains the neural mechanisms underlying the generation of the POR and quantitatively accounts for the relation between its peak latency and the perceived pitch. Next, we applied magnetoencephalography (MEG) to record the POR as elicited by six consonant and dissonant dyads. The peak latency of the POR was strongly modulated by the degree of consonance within the stimuli; specifically, the most dissonant dyad exhibited a POR with a latency that was about 30ms longer than that of the most consonant dyad, an effect that greatly exceeds the expected latency difference induced by a single pitch sound. Our model was able to predict the POR latency pattern observed in the neuromagnetic data, and to generalize this prediction to additional dyads. These results indicate that the neural mechanisms responsible for pitch processing exhibit an intrinsic differential response to concurrent consonant and dissonant pitch combinations, suggesting that the perception of consonance and dissonance might be an emergent property of the pitch processing system in human auditory cortex.
Pitch is the perceptual correlate of sound's periodicity and a fundamental property of the auditory sensation. The interaction of two or more pitches gives rise to a sensation that can be characterized by its degree of consonance or dissonance. In the current study, we investigated the neuromagnetic representations of consonant and dissonant musical dyads using a new model of cortical activity, in an effort to assess the possible involvement of pitch-specific neural mechanisms in consonance processing at early cortical stages. In the first step of the study, we developed a novel model of cortical pitch processing designed to explain the morphology of the pitch onset response (POR), a pitch-specific subcomponent of the auditory evoked N100 component in the human auditory cortex. The model explains the neural mechanisms underlying the generation of the POR and quantitatively accounts for the relation between its peak latency and the perceived pitch. Next, we applied magnetoencephalography (MEG) to record the POR as elicited by six consonant and dissonant dyads. The peak latency of the POR was strongly modulated by the degree of consonance within the stimuli; specifically, the most dissonant dyad exhibited a POR with a latency that was about 30ms longer than that of the most consonant dyad, an effect that greatly exceeds the expected latency difference induced by a single pitch sound. Our model was able to predict the POR latency pattern observed in the neuromagnetic data, and to generalize this prediction to additional dyads. These results indicate that the neural mechanisms responsible for pitch processing exhibit an intrinsic differential response to concurrent consonant and dissonant pitch combinations, suggesting that the perception of consonance and dissonance might be an emergent property of the pitch processing system in human auditory cortex.