Statistical learning in deaf and hard of hearing (DHH) children is poorly understood and the impact of age of sign language acquisition has not been systematically investigated. In the present study, we examined visual statistical learning and transfer effects in early (age of sign language acquisition before the age of three years) and late (age of sign language acquisition after the age of three years) DHH child signers (aged 6-8 years) of German Sign Language, compared to agematched hearing children. The children completed a visual artificial grammar learning task across three learning sessions, each consisting of alternating learning and test phases. In a fourth session, a new stimulus set was introduced using the same underlying artificial grammar to assess transfer abilities. Early signers tended to show higher learning performance in the first session, while no difference between early and late signers was observed in the third session. Compared to agematched hearing children, both DHH groups were indistinguishable in the first session but performed lower in the third session. Crucially, only early signers successfully transferred the learned sequence rules to the new stimulus set comparable to hearing children. Correlation analyses indicated that higher visuo-spatial working memory capacity in DHH children was associated with higher performance during transfer and in late signers during learning as well. Overall, the findings suggest that early exposure to natural language input affects visual statistical learning abilities in DHH children.
Across various types of learning and memory, when a new training session follows a previous one after a certain temporal interval, the previously acquired learning can be disrupted—an effect known as retrograde interference (RI) or catastrophic forgetting.1,2,3,4,5,6,7,8,9,10,11,12,13,14,15 This disruption is thought to result from disrupting interactions between the learning of the first-trained task and the learning of the second-trained task while the former has not yet stabilized.6,7,9,10,16 Such destructive interactions have been considered characteristic not only of RI but also of related phenomena.9,10,11,17,18,19,20,21,22,23,24,25,26,27,28,29 However, we found that when the trained feature was subthreshold, the new learning session unexpectedly improved—rather than impaired—performance on the first-trained task, indicating a retrograde facilitation (RF) effect. We demonstrated this in visual perceptual learning (VPL) by conducting two successive training sessions on different coherent motion directions without any temporal gap. Consistent with previous research, when these directions were suprathreshold (10% coherent motion), the second session disrupted improvements from the first, reflecting RI. By contrast, when the trained directions were subthreshold (5% coherent motion), performance improvement on the first-trained direction was greater with a second session than without—indicating RF. Notably, RF was not observed when a 1-h interval separated the two subthreshold training sessions. This finding suggests that facilitative interactions occur only before the learning of the first-trained direction is stabilized. These results provide a new insight: a stimulus detection system related to conscious awareness transforms what would otherwise be facilitative interactions between successive VPL sessions into disruptive ones.
Multisensory processing critically depends on the perceived timing of stimuli in the different sensory modalities. Crossmodal stimuli that fall within rather than outside an individual temporal binding window (TBW) are more likely to be bound into a multisensory percept. A number of studies have shown that a short perceptual training in which participants receive feedback on their responses in an audiovisual simultaneity judgment (SJ) task can substantially decrease the size of the TBW and hence increase crossmodal temporal acuity. Here we tested whether multisensory perceptual learning in the SJ task is specific for the spatial locations at which the audiovisual stimuli are presented during training. Participants received feedback about the correctness of their SJ responses for audiovisual stimuli which were presented in one hemifield only. The TBW was assessed separately for audiovisual stimuli in each hemifield before and one day after the training. In line with previous findings, the size of the TBW was significantly reduced after the training phase. Importantly, an equally strong reduction of TBW size was observed in both the trained and the untrained hemifield. Thus, multisensory temporal learning completely generalized to the untrained hemifield, suggesting that the improvement in crossmodal temporal acuity was mediated by higher, location-invariant processing stages. These findings have implications for the design of multisensory training protocols in applied settings such as clinical interventions by showing that training at multiple spatial locations might not be necessary to achieve robust improvements in crossmodal temporal acuity.
Oculomotor signals influence the neural processing of auditory input. Recent studies have shown that this connection extends to the auditory periphery: The phase and amplitude of eardrum oscillations was systematically influenced by eye movement direction and magnitude, a phenomenon called eye movement-related eardrum oscillations (EMREOs). Previous findings have suggested that EMREOs occur independently from auditory stimulation, but it is unknown whether they depend on the presence of visual sensory input or solely reflect efference copies of the oculomotor system. To distinguish between these two alternatives, we measured eye movements and eardrum oscillations in sighted human participants who performed free saccadic eye movements in darkness. Despite the lack of any sensory stimulation during eye movements, significant EMREOs occurred in all participants. EMREO characteristics were comparable to a separate control experiment in which participants performed guided saccades to visual targets and were robust to different types of eye tracker calibration methods. Thus, our results suggest that EMREOs are not driven by bottom-up sensory signals but rather reflect a pure influence of oculomotor signals on peripheral auditory processing. This indicates that EMREOs might play a crucial role in reference frame transformations which are needed for audio-visual spatial integration.
In a recent study, we reported that multisensory enhancement (ME) of auditory localization after exposure to spatially congruent audiovisual stimuli and crossmodal recalibration in the ventriloquism aftereffect (VAE) are differently affected by the temporal stimulation frequency with which the audiovisual exposure stimuli are presented. Because audiovisual stimulation at 10 Hz rather than at 2 Hz selectively abolished the VAE but did not affect the ME, we concluded that distinct underlying neural mechanisms are involved in the two effects. A commentary on our paper challenged this interpretation and argued that the ME might have been spared simply because participants had acquired higher order knowledge about the loudspeaker locations from the visual stimulus locations in the ME condition, or because the ME was generally more reliable than the VAE. To test this alternative explanation of our results, we conducted an additional control experiment in which participants localized sounds before and after exposure to unimodal visual stimulation at the loudspeaker locations. No significant reduction of auditory localization errors was found after unimodal visual exposure, suggesting that higher order visual location learning cannot sufficiently explain the significant ME that was observed after audiovisual exposure in our previous study. These new results confirm previous findings pointing toward dissociable neural mechanisms underlying ME and VAE.
It has been proposed that dysfunctions in emotional multisensory integration (MSI) could contribute to the development of psychosis. To further substantiate this proposition, we investigated whether impaired MSI of emotional cues can be observed in people with high psychosis proneness without a diagnosis of psychosis and whether it is associated with aberrant perception and psychotic experiences. Adults scoring high vs. low on the positive subscale of the Community Assessment of Psychic Experiences (score ≥9 or <9, respectively; n = 36 each) categorized the perceived emotion and rated the intensity of unimodal, bimodal emotionally congruent and bimodal emotionally incongruent dynamic face-voice stimuli. In different blocks, participants were asked to attend to one modality and to ignore the other modality input. Additionally, participants completed self-report questionnaires on anomalous perceptual experiences, hallucinations and paranoia. Participants with high and low psychosis proneness did not differ in emotion categorization performance as indicated by similar inverse efficiency (IE) scores (i.e., mean reaction time divided by accuracy) in all conditions, nor did they differ in intensity ratings in any condition. Correlation analyses did not reveal significant associations between crossmodal (in)congruency effects and self-reported anomalous perceptual experiences, hallucinations or paranoia. Our findings, thus, do not provide support for the assumption that MSI of emotional cues is linked to altered perception or subclinical psychotic symptoms, nor for the notion that MSI of emotional cues is already altered at a very early stage in the developmental trajectory of psychosis.
Acquiring sequential information is of utmost importance, for example, for language acquisition in children. Yet, the long-term storage of statistical learning in children is poorly understood. To address this question, 27 7-year-olds and 28 young adults completed four sessions of visual sequence learning (Year 1). From this sample, 16 7-year-olds and 20 young adults participated in another four equivalent sessions after a 12-month-delay (Year 2). The first three sessions of each year used Stimulus Set 1, and the last session used Stimulus Set 2 to investigate transfer effects. Each session consisted of alternating learning and test phases in a modified artificial grammar learning task. In Year 1, 7-year-olds and adults learned the regularities and showed transfer to Stimulus Set 2. Both groups retained their final performance level over the 1-year period. In Year 2, children and adults continued to improve with Stimulus Set 1 but did not show additional transfer gains. Adults overall outperformed children, but transfer effects were indistinguishable between both groups. The current results suggest that long-term memory traces are formed from repeated sequence learning that can be used to generalize sequence rules to new visual input. However, the current study did not provide evidence for a childhood advantage in learning and remembering sequence rules.
The ability to detect the absolute location of sensory stimuli can be quantified with either error-based metrics derived from single-trial localization errors or regression-based metrics derived from a linear regression of localization responses on the true stimulus locations. Here we tested the agreement between these two approaches in estimating accuracy and precision in a large sample of 188 subjects who localized auditory stimuli from different azimuthal locations. A subsample of 57 subjects was subsequently exposed to audiovisual stimuli with a consistent spatial disparity before performing the sound localization test again, allowing us to additionally test which of the different metrics best assessed correlations between the amount of crossmodal spatial recalibration and baseline localization performance. First, our findings support a distinction between accuracy and precision. Localization accuracy was mainly reflected in the overall spatial bias and was moderately correlated with precision metrics. However, in our data, the variability of single-trial localization errors (variable error in error-based metrics) and the amount by which the eccentricity of target locations was overestimated (slope in regression-based metrics) were highly correlated, suggesting that intercorrelations between individual metrics need to be carefully considered in spatial perception studies. Secondly, exposure to spatially discrepant audiovisual stimuli resulted in a shift in bias toward the side of the visual stimuli (ventriloquism aftereffect) but did not affect localization precision. The size of the aftereffect shift in bias was at least partly explainable by unspecific test repetition effects, highlighting the need to account for inter-individual baseline differences in studies of spatial learning.
The accurate estimation of time-to-collision (TTC) is essential for the survival of organisms. Previous studies have revealed that the emotional properties of approaching stimuli can influence the estimation of TTC, indicating that approaching threatening stimuli are perceived to collide with the observers earlier than they actually do, and earlier than non-threatening stimuli. However, not only are threatening stimuli more negative in valence, but they also have higher arousal compared to non-threatening stimuli. Up to now, the effect of arousal on TTC estimation remains unclear. In addition, inconsistent findings may result from the different experimental settings employed in previous studies. To investigate whether the underestimation of TTC is attributed to threat or high arousal, three experiments with the same settings were conducted. In Experiment 1, the underestimation of TTC estimation of threatening stimuli was replicated when arousal was not controlled, in comparison to non-threatening stimuli. In Experiments 2 and 3, the underestimation effect of threatening stimuli disappeared when compared to positive stimuli with similar arousal. These findings suggest that being threatening alone is not sufficient to explain the underestimation effect, and arousal also plays a significant role in the TTC estimation of approaching stimuli. Further studies are required to validate the effect of arousal on TTC estimation, as no difference was observed in Experiment 3 between the estimated TTC of high and low arousal stimuli.
Auditory and visual information involve different coordinate systems, with auditory spatial cues anchored to the head and visual spatial cues anchored to the eyes. Information about eye movements is therefore critical for reconciling visual and auditory spatial signals. The recent discovery of eye movement-related eardrum oscillations (EMREOs) suggests that this process could begin as early as the auditory periphery. How this reconciliation might happen remains poorly understood. Because humans and monkeys both have mobile eyes and therefore both must perform this shift of reference frames, comparison of the EMREO across species can provide insights to shared and therefore important parameters of the signal. Here we show that rhesus monkeys, like humans, have a consistent, significant EMREO signal that carries parametric information about eye displacement as well as onset times of eye movements. The dependence of the EMREO on the horizontal displacement of the eye is its most consistent feature, and is shared across behavioural tasks, subjects and species. Differences chiefly involve the waveform frequency (higher in monkeys than in humans) and patterns of individual variation (more prominent in monkeys than in humans), and the waveform of the EMREO when factors due to horizontal and vertical eye displacements were controlled for.This article is part of the theme issue 'Decision and control processes in multisensory perception'.
Acquiring sequential information is of utmost importance, e.g., for language acquisition in children. Yet, the long-term storage of statistical learning in children is poorly understood. To address this question, 27 seven-year-olds and 28 young adults completed four sessions of visual sequence learning (Year 1). From this sample, 16 seven-year-olds and 20 young adults participated in another four equivalent sessions after a 12-month-delay (Year 2). The first three sessions of each year used stimulus set-1, while the last session used stimulus set-2 to investigate transfer effects. Each session consisted of alternating learning and test phases in a modified artificial grammar learning task. In Year 1, seven-year-olds and adults learned the regularities and showed transfer to stimulus set-2. Both groups retained their final performance level over the one-year-period. In Year 2, children and adults continued to improve with stimulus set-1, but did not show additional transfer gains. Adults overall outperformed children, but transfer effects were indistinguishable between both groups. The present results suggest that long-term memory traces are formed from repeated sequence learning which can be used to generalize sequence rules to new visual input. However, the present study did not provide evidence for a childhood advantage in learning and remembering sequence rules.
Human infant learning happens during exploration of the environment, by interaction with objects, and by listening to and repeating utterances casually, which is analogous to unsupervised learning. Only occasionally, a learning infant would receive a matching verbal description of an action it is committing, which is similar to supervised learning. Such a learning mechanism can be mimicked with deep learning. We model this weakly supervised learning paradigm using our Paired Gated Autoencoders (PGAE) model, which combines an action and a language autoencoder. After observing a performance drop when reducing the proportion of supervised training, we introduce the Paired Transformed Autoencoders (PTAE) model, using Transformer-based crossmodal attention. PTAE achieves significantly higher accuracy in language-to-action and action-to-language translations, particularly in realistic but difficult cases when only few supervised training samples are available. We also test whether the trained model behaves realistically with conflicting multimodal input. In accordance with the concept of incongruence in psychology, conflict deteriorates the model output. Conflicting action input has a more severe impact than conflicting language input, and more conflicting features lead to larger interference. PTAE can be trained on mostly unlabelled data where labeled data is scarce, and it behaves plausibly when tested with incongruent input.
Multisensory spatial processes are fundamental for efficient interaction with the world. They include not only the integration of spatial cues across sensory modalities, but also the adjustment or recalibration of spatial representations to changing cue reliabilities, crossmodal correspondences, and causal structures. Yet how multisensory spatial functions emerge during ontogeny is poorly understood. New results suggest that temporal synchrony and enhanced multisensory associative learning capabilities first guide causal inference and initiate early coarse multisensory integration capabilities. These multisensory percepts are crucial for the alignment of spatial maps across sensory systems, and are used to derive more stable biases for adult crossmodal recalibration. The refinement of multisensory spatial integration with increasing age is further promoted by the inclusion of higher-order knowledge.
To clarify the role of sensory experience during early development for adult multisensory learning capabilities, we probed audiovisual spatial processing in human individuals who had been born blind because of dense congenital cataracts (CCs) and who subsequently had received cataract removal surgery, some not before adolescence or adulthood. Their ability to integrate audio-visual input and to recalibrate multisensory spatial representations was compared to normally sighted control participants and individualswith a history of developmental (later onset) cataracts. Results in CC individuals revealed both normal multisensory integration in audiovisual trials (ventriloquism effect) and normal recalibration of unimodal auditory localization following audiovisual discrepant exposure (ventriloquismaftereffect) as observed in the control groups. In addition, only the CC group recalibrated unimodal visual localization after audiovisual exposure. Thus, in parallel to typical multisensory integration and learning, atypical crossmodal mechanisms coexisted in CC individuals, suggesting that multisensory recalibration capabilities are defined during a sensitive period in development.
Reliability-based cue combination is a hallmark of multisensory integration, while the role of cue reliability for crossmodal recalibration is less understood. The present study investigated whether visual cue reliability affects audiovisual recalibration in adults and children. Participants had to localize sounds, which were presented either alone or in combination with a spatially discrepant high- or low-reliability visual stimulus. In a previous study we had shown that the ventriloquist effect (indicating multisensory integration) was overall larger in the children groups and that the shift in sound localization toward the spatially discrepant visual stimulus decreased with visual cue reliability in all groups. The present study replicated the onset of the immediate ventriloquist aftereffect (a shift in unimodal sound localization following a single exposure of a spatially discrepant audiovisual stimulus) at the age of 6-7 years. In adults the immediate ventriloquist aftereffect depended on visual cue reliability, whereas the cumulative ventriloquist aftereffect (reflecting the audiovisual spatial discrepancies over the complete experiment) did not. In 6-7-year-olds the immediate ventriloquist aftereffect was independent of visual cue reliability. The present results are compatible with the idea of immediate and cumulative crossmodal recalibrations being dissociable processes and that the immediate ventriloquist aftereffect is more closely related to genuine multisensory integration.
It has been hypothesized that crossmodal recalibration plays a crucial role for the development of multisensory integration capabilities [1]. To test the developmental trajectory of multisensory integration and crossmodal recalibration, we used a combined ventriloquist/ventriloquist aftereffect paradigm [2] in children aged 5-9 years. The ventriloquist effect (indicating multisensory integration), that is, the shift of auditory localization toward simultaneously presented but spatially discrepant visual stimuli, was larger in children than in adults, which was attributed to a lower auditory localization precision in the children. In fact, the size of the ventriloquist effect depended on the visual stimulus reliability in both children and adults. In all groups, the ventriloquist effect was best explained by a causal inference model. In contrast to their multisensory integration capabilities, 5-year-old children did not recalibrate. The immediate ventriloquist aftereffect (indicating recalibration after a single exposure to a spatially discrepant audio-visual stimulus) emerged in 6- to 7-year-old children, whereas the cumulative ventriloquist aftereffect (reflecting recalibration to the audio-visual spatial discrepancies over the complete experiment) was not observed before the age of 8 years. First, in contrast to common beliefs, the present results provide evidence that multisensory integration precedes rather than follows crossmodal recalibration during development. Second, we report developmental evidence for a dissociation of the processes involved in multisensory integration and immediate as well as cumulative recalibration. We speculate that multisensory integration is a prerequisite for crossmodal recalibration, because the multisensory percept, rather than unimodal cues, might comprise a crucial signal for the calibration of the sensory systems.
According to the Bayesian framework of multisensory integration, audiovisual stimuli associated with a stronger prior belief that they share a common cause (i.e., causal prior) are predicted to result in a greater degree of perceptual binding and therefore greater audiovisual integration. In the present psychophysical study, we systematically manipulated the causal prior while keeping sensory evidence constant. We paired auditory and visual stimuli during an association phase to be spatiotemporally either congruent or incongruent, with the goal of driving the causal prior in opposite directions for different audiovisual pairs. Following this association phase, every pairwise combination of the auditory and visual stimuli was tested in a typical ventriloquism-effect (VE) paradigm. The size of the VE (i.e., the shift of auditory localization towards the spatially discrepant visual stimulus) indicated the degree of multisensory integration. Results showed that exposure to an audiovisual pairing as spatiotemporally congruent compared to incongruent resulted in a larger subsequent VE (Experiment1). This effect was further confirmed in a second VE paradigm, where the congruent and the incongruent visual stimuli flanked the auditory stimulus, and a VE in the direction of the congruent visual stimulus was shown (Experiment2). Since the unisensory reliabilities for the auditory or visual components did not change after the association phase, the observed effects are likely due to changes in multisensory binding by association learning. As suggested by Bayesian theories of multisensory processing, our findings support the existence of crossmodal causal priors that are flexibly shaped by experience in a changing world.
In an ever-changing environment, crossmodal recalibration is crucial to maintain precise and coherent spatial estimates across different sensory modalities. Accordingly, it has been found that perceived auditory space is recalibrated toward vision after consistent exposure to spatially misaligned audio-visual stimuli (VS). While this so-called ventriloquism aftereffect (VAE) yields internal consistency between vision and audition, it does not necessarily lead to consistency between the perceptual representation of space and the actual environment. For this purpose, feedback about the true state of the external world might be necessary. Here, we tested whether the size of the VAE is modulated by external feedback and reward. During adaptation audio-VS with a fixed spatial discrepancy were presented. Participants had to localize the sound and received feedback about the magnitude of their localization error. In half of the sessions the feedback was based on the position of the VS and in the other half it was based on the position of the auditory stimulus. An additional monetary reward was given if the localization error fell below a certain threshold that was based on participants' performance in the pretest. As expected, when error feedback was based on the position of the VS, auditory localization during adaptation trials shifted toward the position of the VS. Conversely, feedback based on the position of the auditory stimuli reduced the visual influence on auditory localization (i.e., the ventriloquism effect) and improved sound localization accuracy. After adaptation with error feedback based on the VS position, a typical auditory VAE (but no visual aftereffect) was observed in subsequent unimodal localization tests. By contrast, when feedback was based on the position of the auditory stimuli during adaptation, no auditory VAE was observed in subsequent unimodal auditory trials. Importantly, in this situation no visual aftereffect was found either. As feedback did not change the physical attributes of the audio-visual stimulation during adaptation, the present findings suggest that crossmodal recalibration is subject to top-down influences. Such top-down influences might help prevent miscalibration of audition toward conflicting visual stimulation in situations in which external feedback indicates that visual information is inaccurate.
Extracting information from noisy signals is of fundamental importance for both biological and artificial perceptual systems. To provide tractable solutions to this challenge, the fields of human perception and machine signal processing (SP) have developed powerful computational models, including Bayesian probabilistic models. However, little true integration between these fields exists in their applications of the probabilistic models for solving analogous problems, such as noise reduction, signal enhancement, and source separation. In this mini review, we briefly introduce and compare selective applications of probabilistic models in machine SP and human psychophysics. We focus on audio and audio-visual processing, using examples of speech enhancement, automatic speech recognition, audio-visual cue integration, source separation, and causal inference to illustrate the basic principles of the probabilistic approach. Our goal is to identify commonalities between probabilistic models addressing brain processes and those aiming at building intelligent machines. These commonalities could constitute the closest points for interdisciplinary convergence.