Fearful faces attract gaze even when they remain 'unseen', suggesting that the oculomotor system prioritizes threat-related facial information in the absence of visual awareness. As congruent voices enhance face perception, we tested whether emotional voices modulate the effects of suppressed fearful and neutral faces on observers' gaze. Using continuous flash suppression, we suppressed fearful and neutral faces from visual awareness and paired them with either a fearful voice, a neutral voice, or no voice. Concurrently, we tracked eye movements. We verified absence of visual awareness of the face using objective measures (chance performance in face localization and categorization tasks) and subjective measures (reports of zero face visibility). Eye tracking data from trials with successful suppression revealed distinct gaze patterns: Fearful faces attracted gaze regardless of vocal context, confirming that visual threats actively guide the eyes in the absence of awareness. By contrast, unseen neutral faces attracted gaze only when paired with a neutral voice, indicating a form of multisensory enhancement whereby emotionally matching sounds enhance weak visual signals to be prioritized by the oculomotor system.
Sensory processing is fundamentally shaped by stimulation history. For example, in visual cortex, neural responses are reduced for repeated or sustained stimuli (adaptation). These phenomena are well characterized and effectively modeled by divisive normalization. We asked whether these same computational principles govern somatosensory processing. We used fMRI (6 participants) and intracranial electroencephalography (iEEG, 2 participants) to measure responses to time-varying vibrotactile stimuli in human somatosensory cortex. Stimuli consisted of single- and paired-pulses with durations and interstimulus intervals ranging from 0.05 to 1.2 s. We extracted BOLD time courses to capture neural response amplitudes, and high-frequency iEEG broadband envelopes to capture fast neural dynamics. In both experiments, we observed pronounced sub-additive temporal summation. Responses to longer or repeated stimuli were consistently lower than predicted by linear integration. Computational modeling revealed that divisive normalization models outperformed linear models in cross-validated accuracy across both datasets. These results demonstrate that somatosensory temporal dynamics closely mirror those in the visual system. Our findings suggest that the nervous system employs similar computational principles across modalities to encode sensory information across time. Significance statement:How the brain integrates sensory information over time is a fundamental question in neuroscience. While nonlinear temporal integration is well documented in visual cortex, it has not been extensively mapped in the human somatosensory system. By combining fMRI with intracranial EEG in humans, we demonstrate that somatosensory responses to tactile stimulation exhibit subadditive temporal summation. This nonlinearity is accurately captured by a divisive normalization model, matching observations in the visual system. Our results suggest that normalization is a canonical computation shared across different modalities to manage temporal dynamics, providing a unified framework for understanding how the brain encodes dynamic sensory stimuli.
Visual stimuli that are not consciously perceived can nevertheless guide the eyes. Yet, which stimulus aspects drive such oculomotor guidance by "invisible" images remains unclear. We tested whether image representations strengthened by congruent actions could guide the eyes even in the absence of visual awareness. Images of snapping fingers, tapping feet, and static noses were continuously suppressed from visual awareness using continuous flash suppression; participants' eye gaze was tracked as they snapped their fingers, tapped their foot, or executed no action. The eyes were attracted towards suppressed foot and finger images, but not suppressed nose images, when participants executed an action, but not when they remained motionless. Thus, suppressed images of action-related body parts selectively attracted the eyes when an action was performed. These results indicate that an effector-unspecific strengthening of image representations through concurrently executed actions drives oculomotor behavior even in the absence of visual awareness.
Cross-modal temporal recalibration guarantees stable temporal perception across ever-changing environments. Yet, the mechanisms of cross-modal temporal recalibration remain unknown. Here, we conducted an experiment to measure how participants’ temporal perception was affected by exposure to audiovisual stimuli with constant temporal delays that we varied across sessions. Consistent with previous findings, recalibration effects plateaued with increasing audiovisual asynchrony (nonlinearity) and varied by which modality led during the exposure phase (asymmetry). We compared six observer models that differed in how they update the audiovisual temporal bias during the exposure phase and in whether they assume a modality-specific or modality-independent precision of arrival latency. The causal-inference observer shifts the audiovisual temporal bias to compensate for perceived asynchrony, which is inferred by considering two causal scenarios: when the audiovisual stimuli have a common cause or separate causes. The asynchrony-contingent observer updates the bias to achieve simultaneity of auditory and visual measurements, modulating the update rate by the likelihood of the audiovisual stimuli originating from a simultaneous event. In the asynchrony-correction model, the observer first assesses whether the sensory measurement is asynchronous; if so, she adjusts the bias proportionally to the magnitude of the measured asynchrony. Each model was paired with either modality-specific or modality-independent precision of arrival latency. A Bayesian model comparison revealed that both the causal-inference process and modality-specific precision in arrival latency are required to capture the nonlinearity and asymmetry observed in audiovisual temporal recalibration. Our findings support the hypothesis that audiovisual temporal recalibration relies on the same causal-inference processes that govern cross-modal perception.
Making decisions based on noisy sensory information is a crucial function of the brain. Various decisions take each sensory signal's uncertainty into account. Here, we investigated whether perceptual inferences rely on accurate estimates of sensory uncertainty. Participants completed a set of auditory, visual, and audiovisual spatial as well as temporal tasks. We fitted Bayesian observer models of each task to every participant's full dataset. Crucially, in some model variants, the uncertainty estimates employed for perceptual inferences were independent of the actual uncertainty associated with the sensory signals. Model comparisons and analysis of the best-fitting parameters revealed that, in unimodal and bimodal contexts, participants' perceptual decisions relied on underestimates of auditory spatial and audiovisual temporal uncertainty. These findings challenge the ubiquitous assumption that human behaviour optimally accounts for sensory uncertainty regardless of sensory domain.
Humans show perceptual biases that suggest distorted internal representations of their own body. New research reveals that these perceptual biases can reflect integration of prior assumptions about body posture rather than a misshaped representation of the body’s geometry.
A central question about the human mind is whether perception is an encapsulated process driven purely by sensory information or whether it is intricately linked with cognitive processes. This debate about the cognitive penetrability of perception is discussed in psychology, cognitive neuroscience and philosophy. Thus far, the debate has centred on vision, without major attempts to examine other senses. In this Review, we provide an overview of the key empirical evidence about cognitive penetrability of perception in vision, audition, somatosensation (including proprioception and pain perception), vestibular perception and chemosensation (gustation, chemesthesis and olfaction). We conclude that many (but not all) of the senses are cognitively penetrable. Specifically, cognitive penetrability seems to vary with the extent to which a sense is intrinsically multimodal, the extent to which it receives indirect cognitive influences, and whether hedonic evaluation is an integral aspect of the perceptual experience. We suggest that the debate about cognitive penetrability needs to be more differentiated with respect to the sensory modality of the perceptual experience and the diversity of cognitive influences on that modality. The debate over cognitive penetrability of perception, which has been largely limited to vision, remains unsolved; in this Review, Vetter and colleagues detail cognitive influences on perception across vision, audition, somatosensation, vestibular perception and chemosensation to advance the debate.
The ability to judge the temporal alignment of visual and auditory information is a prerequisite for multisensory integration and segregation. However, each temporal measurement is subject to error. Thus, when judging whether a visual and auditory stimulus were presented simultaneously, observers must rely on a subjective decision boundary to distinguish between measurement error and truly misaligned audiovisual signals. Here, we tested whether these decision boundaries are relaxed with increasing temporal sensory uncertainty, i.e., whether participants make the same type of adjustment an ideal observer would make. Participants judged the simultaneity of audiovisual stimulus pairs with varying temporal offset, while being immersed in different virtual environments. To obtain estimates of participants’ temporal sensory uncertainty and simultaneity criteria in each environment, an independent-channels model was fitted to their simultaneity judgments. In two experiments, participants’ simultaneity decision boundaries were predicted by their temporal uncertainty, which varied unsystematically with the environment. Hence, observers used a flexibly updated estimate of their own audiovisual temporal uncertainty to establish subjective criteria of simultaneity. This finding implies that, under typical circumstances, audiovisual simultaneity windows reflect an observer’s cross-modal temporal uncertainty.
The human brain is sensitive to threat-related information even when we are not aware of this information. For example, fearful faces attract gaze in the absence of visual awareness. Moreover, information in different sensory modalities interacts in the absence of awareness, for example, the detection of suppressed visual stimuli is facilitated by simultaneously presented congruent sounds or tactile stimuli. Here, we combined these two lines of research and investigated whether threat-related sounds could facilitate visual processing of threat-related images suppressed from awareness such that they attract eye gaze. We suppressed threat-related images of cars and neutral images of human hands from visual awareness using continuous flash suppression and tracked observers’ eye movements while presenting congruent or incongruent sounds (finger snapping and car engine sounds). Indeed, threat-related car sounds guided the eyes toward suppressed car images, participants looked longer at the hidden car images than at any other part of the display. In contrast, neither congruent nor incongruent sounds had a significant effect on eye responses to suppressed finger images. Overall, our results suggest that only in a danger-related context semantically congruent sounds modulate eye movements to images suppressed from awareness, highlighting the prioritisation of eye responses to threat-related stimuli in the absence of visual awareness.
Combining information from multiple senses enhances our perception of the world. Whether we need to be aware of all stimuli to benefit from multisensory integration, however, is still under investigation. Here, we tested whether tactile frequency perception benefits from the presence of congruent visual flicker even if the flicker is so rapid that it is perceptually fused into a steady light and therefore invisible. Our participants completed a tactile frequency discrimination task given either unisensory tactile or congruent tactile-visual stimulation. Tactile and tactile-visual test frequencies ranged from far below to far above participants' flicker fusion threshold (determined separately). For frequencies distinctively below their flicker fusion threshold, participants performed significantly better given tactile-visual stimulation than when presented with only tactile stimuli. Yet, for frequencies above their flicker fusion threshold, participants' tactile frequency perception did not profit from the presence of congruent but likely fused and thus invisible visual flicker. The results matched the predictions of an ideal-observer model in which tactile-visual integration is conditional on awareness of both stimuli. In contrast, it was impossible to reproduce the observed results with a model that assumed tactile-visual integration proceeds irrespective of stimulus awareness. In sum, we revealed that the benefits of congruent visual stimulation for tactile flutter frequency perception depend on the visibility of the visual flicker, suggesting that multisensory integration requires awareness.
Multisensory integration depends on causal inference about the sensory signals. We tested whether implicit causal-inference judgements pertain to entire objects or focus on task-relevant object features. Participants in our study judged virtual visual, haptic and visual-haptic surfaces with respect to two features-slant and roughness-against an internal standard in a two-alternative forced-choice task. Modelling of participants' responses revealed that the degree to which their perceptual judgements were based on integrated visual-haptic information varied unsystematically across features. For example, a perceived mismatch between visual and haptic roughness would not deter the observer from integrating visual and haptic slant. These results indicate that participants based their perceptual judgements on a feature-specific selection of information, suggesting that multisensory causal inference proceeds not at the object level but at the level of single object features. This article is part of the theme issue 'Decision and control processes in multisensory perception'.
PURPOSE:Hepatic steatosis is often diagnosed non-invasively. Various measures and accompanying diagnostic thresholds based on contrast-enhanced CT and virtual non-contrast images have been proposed. We compare these established criteria to novel and fully automated measures.METHOD:CT data sets of 197 patients were analyzed. Regions of interest (ROIs) were manually drawn for the liver, spleen, portal vein, and aorta to calculate four established measures of liver-fat. Two novel measures capturing the deviation between the empirical distributions of HU measurements across all voxels within the liver and spleen were calculated. These measures were calculated with both manual ROIs and using fully automated organ segmentations. Agreement between the different measures was evaluated using correlational analysis, as well as their ability to discriminate between fatty and healthy liver.RESULTS:Established and novel measures of fatty liver were at a high level of agreement. Novel methods were statistically indistinguishable from the established ones when taking established diagnostic thresholds or physicians' diagnoses as ground truth and this high performance level persisted for automatically selected ROIs.CONCLUSION:Automatically generated organ segmentations led to comparable results as manual ROIs, suggesting that the implementation of automated methods can prove to be a valuable tool for incidental diagnosis. Differences in the distribution of HU measurements across voxels between liver and spleen can serve as surrogate markers for the liver-fat-content. Novel measures do not exhibit a measurable disadvantage over established methods based on simpler measures such as across-voxel averages in a population with low incidence of fatty liver.
Our skin is a two-dimensional sheet that can be folded into a multitude of configurations due to the mobility of our body parts. Parts of the human tactile system might account for this flexibility by being tuned to locations in the world rather than on the skin. Using adaptation, we scrutinized the spatial selectivity of two tactile perceptual mechanisms for which the visual equivalents have been reported to be selective in world coordinates: tactile motion and the duration of tactile events. Participants’ hand position—uncrossed or crossed—as well as the stimulated hand varied independently across adaptation and test phases. This design distinguished among somatotopic selectivity for locations on the skin and spatiotopic selectivity for locations in the environment, but also tested spatial selectivity that fits neither of these classical reference frames and is based on the default position of the hands. For both features, adaptation consistently affected subsequent tactile perception at the adapted hand, reflecting skin-bound spatial selectivity. Yet, tactile motion and temporal adaptation also transferred across hands but only if the hands were crossed during the adaptation phase, that is, when one hand was placed at the other hand’s typical location. Thus, selectivity for locations in the world was based on default rather than online sensory information about the location of the hands. These results challenge the prevalent dichotomy of somatotopic and spatiotopic selectivity and suggest that prior information about the hands’ default position —right hand at the right side—is embedded deep in the tactile sensory system.
In daily life, we are bombarded with an abundance of sensory information from several modalities. Due to internal noise in the brain and external noise in the environment, two sensory cues originating from the same object will not perfectly agree with each other in terms of space and time. Therefore, to maintain a coherent percept of the world, the brain integrates cues that are likely to come from the same source and segregates those that are not. The goal of the current study is to understand how temporal discrepancies impact the influence of spatially discrepant visual stimuli on auditory localization as well as perceived unity of the spatially and temporally discrepant cues. To this aim, we presented participants with a visual (a Gaussian blob) and an auditory stimulus (a brief burst of white noise) with various spatial and temporal discrepancies between them. Participants first localized the auditory stimulus and then reported whether they perceived the two stimuli as originating from the same source. We computed the ventriloquism effect, the size of auditory shifts towards the accompanying visual stimulus by comparing the auditory localization responses to those when the auditory stimulus was presented alone in a control experiment. As expected, the bias, the ventriloquism effect relative to the spatial discrepancy, first increased and then decreased as the two stimuli were further apart in space. Importantly, the bias also declined as the two stimuli were more temporally discrepant. Consistent results were observed in the unity judgments: the probability of reporting a common cause fell off as a function of both spatial and temporal discrepancy. Our results provide robust evidence that temporal alignment is taken into account even in spatial tasks and thus spatial and temporal discrepancy are driving factors that determine audiovisual integration and inference of a common cause.
To estimate an environmental property such as object location from multiple sensory signals, the brain must infer their causal relationship. Only information originating from the same source should be integrated. This inference relies on the characteristics of the measurements, the information the sensory modalities provide on a given trial, as well as on a cross-modal common-cause prior: accumulated knowledge about the probability that cross-modal measurements originate from the same source. We examined the plasticity of this cross-modal common-cause prior. In a learning phase, participants were exposed to a series of audiovisual stimuli that were either consistently spatiotemporally congruent or consistently incongruent; participants’ audiovisual spatial integration was measured before and after this exposure. We fitted several Bayesian causal-inference models to the data; the models differed in the plasticity of the common-source prior. Model comparison revealed that, for the majority of the participants, the common-cause prior changed during the learning phase. Our findings reveal that short periods of exposure to audiovisual stimuli with a consistent causal relationship can modify the common-cause prior. In accordance with previous studies, both exposure conditions could either strengthen or weaken the common-cause prior at the participant level. Simulations imply that the direction of the prior-update might be mediated by the degree of sensory noise, the variability of the measurements of the same signal across trials, during the learning phase.