
In this viewpoint article, Zhuanghua Shi and Virginie van Wassenhove discuss the past, present, and future of multisensory time perception. Drawing on decades of research, they argue that the field has undergone an important shift from the centralized internal clock toward timing as an active inference process. This framework, rooted in Bayesian models and predictive coding, suggests that our sense of time is not a passive readout but a construction that weighs evidence against contextual expectations. One thread in the discussion concerns the temporal window of integration: the authors contrast the narrow, potentially misleading windows seen with simple beeps and flashes with those, more stable, found in audiovisual speech. The authors also explore the interplay between automatic integration and top-down processes - such as attention and expectation - noting the brain's remarkable ability to rewrite recent history through postdictive inference to maintain a unified 'now'. Key controversies are raised and include whether timing relies on dedicated circuits or emerges from neural population dynamics, as well as the specific role of neural oscillations in parsing simultaneous inputs. The authors advocate for biological plausibility and the use of time-resolved encoding-decoding modeling approaches to predict brain activity. They emphasize the use of open-science practices, including the sharing of computational models and datasets. Looking ahead, the authors anticipate that AI will revolutionize the field by handling the computational complexity of massive datasets through 'vibe coding', allowing researchers to focus on fundamental theoretical insights. They advise young scientists to develop strong scientific intuitions before expanding into modeling and to remain bold in bridging multisensory perception with cognition (incl. consciousness and linguistics). The authors emphasize that multisensory timing is central for developing sensory prostheses, and a deeper understanding of how the brain achieves a seamless experience of reality.
In this viewpoint article, Monica Gori and Micah Murray share their perspectives on past, present, and future developments in multisensory technologies and their impact on society, as well as career perspectives for trainees. Using seminal works on the neurobiological bases of multisensory processes and their development as springboards, they discuss how they foresee continued advancements in the technologies for studying multisensory processes and rehabilitating/replacing sensory impairments. They see a strong tendency toward understanding inter- and intra-individual variability in multisensory processes and their plasticity as well as a societal impetus to transition from laboratory-based to applied research with bespoke solutions for individuals. These tendencies also encapsulate current controversies regarding the boundaries of neuroplasticity as well as the potential double-edged sword of AI-based technologies that have tremendous positive impact though perhaps at the 'cost' of negative collateral impact on functional capabilities. The domain of multisensory technologies offers trainees a wide breadth of career opportunities not only in academic research, but also in applied/industrial fields. Continued ameliorations in multisensory technologies are an exciting domain at the intersection of applied, basic, and clinical research wherein many inroads remain to be achieved.
Accurately perceiving the duration of stimuli is important for many human behaviours, especially those involving movement such as when catching a ball. However, perceived duration is reportedly different depending on which sense is being monitored, with auditory stimuli often judged as longer than a matched visual stimulus. Recent experiments have implicated gravity in duration perception, but rigorous psychophysical methods have been lacking. Here, we developed a novel psychometric method for measuring perceived duration and manipulated the reliability of gravity by altering posture and by disruptive galvanic vestibular stimulation (dGVS). We assessed people's judgements of 1 s using visual (an LED), auditory (200 Hz at 56 dB played through headphones) or tactile (200 Hz played through tactors mounted on the backs of the hands) stimuli. Participants judged whether each stimulus, the duration of which was controlled by a QUick Estimation by Sequential Testing (QUEST) staircase, was longer or shorter than their internal estimate of 1 s in two postures: supine and standing. We also compared visual duration judgements in the presence of dGVS with those collected during sham stimulation to the neck. The perceived durations of auditory and tactile events were not affected by changing posture, but the perceived duration of a visual second was significantly longer when supine compared to when upright (by 45 ms) or during dGVS compared to during sham stimulation (by 66 ms). The duration of a light or tactile stimulus perceived as having a duration of 1 s was longer than the time needed for an auditory event to be perceived as being on for 1 s. We conclude that disrupting the vestibular system by lying supine or by electrical stimulation influences visual duration perception, but the perceived duration of auditory and tactile stimuli remains unaffected.
In this narrative historical review, we take a critical look at the multisensory experience of atmospheric moisture. While not sensed directly (rather, at its most basic, it is a touch blend), this sensation, which can be either positive or negatively valenced (in part, depending on whether we are in charge of the experience), reflects an inference about the environmental conditions as potentially assessed at the skin surface. The perception of humidity has been largely ignored by the long history of research on tactile psychophysics and instead appears only recently in the writings of environmental psychologists. Nevertheless, the perception of humidity is likely to become a growing environmental issue (e.g., in the era of global climate change). We look at the multisensory influences on the perception of humidity, a crucial component of thermal comfort, and possible individual, group, and cultural differences.
A growing body of literature suggests that there may be deficits in audiovisual multisensory integration in schizophrenia at both the nonsocial and social levels of stimuli. The present study provides a meta-analytic synthesis of the literature on audiovisual multisensory integration in schizophrenia to evaluate if there are deficits in multisensory integration in schizophrenia relative to healthy controls (HCs) and whether these deficits are more prominent with certain types of stimuli. Initial searches identified 1552 articles for review, with 17 studies deemed appropriate for inclusion in the meta-analysis. Findings confirm that individuals with schizophrenia exhibit impaired multisensory integration relative to HCs, d = 0.60, 95% CI [0.46, 0.75]. This impairment was seen regardless of stimulus type, across both social, d = 0.71, 95% CI [0.50, 0.91], and nonsocial, d = 0.49, 95% CI [0.27, 0.71] stimuli and did not significantly differ ( p = 0.150). These findings strengthen the past literature suggesting individuals with schizophrenia do exhibit deficits in multisensory integration across multiple levels of stimuli. Clinical implications are discussed.
Verticality perception is a multisensory process that relies primarily on visual and vestibular input. Previous studies show that sensory conflict induced by optic flow causes increased errors in verticality perception. However, the time course over which sensory conflict alters verticality perception and the processes through which perceptual stability is subsequently restored, remain unknown. Experiment 1 aimed to examine how verticality perception changes over time in response to sensory conflict and Experiment 2 aimed to determine whether verticality perception recovers following the removal of a sensory conflict. Sensory conflict was induced using optic flow, such that visual cues signalled self-motion while vestibular input provided no self-motion signal. Sensitivity to verticality was measured using a Vertical Detection Task (VDT) in which participants judged whether lines were vertically oriented or not. Signal detection theory was used to analyse responses, where d ' was an index of perceptual sensitivity to verticality and criterion was a measure of response bias. In Experiment 1 participants completed the detection task whilst either roll optic flow or randomly moving dots were presented in the background. Results showed that d ' was significantly lower in the roll-optic flow condition, with this effect varying as a function of time. Specifically, d ' was significantly lower in the roll versus random visual stimulation condition for up to six minutes of exposure. Experiment 2 extended these findings by examining recovery from sensory conflict, comparing VDT performance during roll optic flow with performance at successive timepoints following its removal. Consistent with Experiment 1, d ' was significantly reduced during optic flow exposure. However, d ' returned to baseline levels by the first post-flow timepoint, indicating a rapid recovery of verticality sensitivity. Together, these findings demonstrate that sensory conflict reduces verticality sensitivity in a time-dependent manner, followed by rapid recovery within minutes after the conflict is removed.
Musical modes (i.e., different patterns of pitch organisation within a scale) are widely recognised for their emotional character, yet little is known about the broader multisensory associations that listening to them might trigger. In the present exploratory study, we examine whether the full set of Western musical modes elicits systematic patterns of visual and olfactory mental imagery and whether these patterns correspond with how listeners spontaneously group the modes. In all, 249 participants generated open-ended visual and olfactory responses while listening to short modal excerpts and subsequently completed a free-sorting task. The results revealed a number of structured and recurrent imagery themes across participants (e.g., Nature-, Daytime-, and Happiness-related imagery for major modes, versus Dark-, Stress- and Sadness-related imagery for minor modes, alongside Floral-and-Fresh- versus Damp-, Dusty- and Smoke-related olfactory associations). Major and minor modes occupied distinct regions of similarity space across both visual and olfactory modalities; however, meaningful differentiation was also evident at the level of individual modes. This consistent relational structure across visual and olfactory imagery was likewise reflected in the grouping patterns observed in the free-sorting task. Together, these findings indicate that musical modes are associated with multiple sensory representations that extend beyond simple feature-level correspondences. These results therefore provide an exploratory mapping of visual and olfactory mental imagery and establish a foundation for future confirmatory research.
We automatically integrate sensory information from various modalities all the time. Although previous research has shown that our taste evaluation can be influenced by tactile texture of a beverage receptacle, the contextual and temporal tolerance of this kind of sensation transference remains unclear. In Experiment 1, 92 blindfolded Japanese participants in Tokyo evaluated a coffee (Rich Blend capsule, Nestlé) as less acid when they drank it from a smooth-texture cup compared to when they drank from a rough-texture cup, indicating the sensation transference from tactile to taste. This effect was also observed when the tactile input came from an object independent to the beverage in Experiment 2 ( n = 24), but absent when the stimuli were not synchronized in Experiment 3 ( n = 48). The present study tested the limits of sensation transference, and found that the stimuli must happen at the same time but do not have to come from the same object.
Multisensory integration is known to facilitate attentional selection, yet its influence on depth perception remains poorly understood, partly because depth has been difficult to manipulate and measure under controlled conditions. Here, we investigated how attentional benefits elicited by the audiovisual integration affect distance perception by adapting the classic pip-and-pop paradigm to immersive virtual reality. Participants ( N = 30) searched for a visual target whose transient colour change occurred either alone or synchronously with a nonspatial auditory stimulus. After identifying the target, participants indicated where they believed it had appeared in three-dimensional space (in terms of laterality, elevation and depth). Across trials, synchronous audiovisual pairing produced two reliable effects. First, participants responded faster when the target event was paired with a tone, replicating the core facilitation associated with the pip-and-pop effect in an immersive three-dimensional setting. Second, audiovisual pairing systematically biased depth perception: targets accompanied by a tone were judged as closer to the observer than identical targets presented without a tone, indicating a consistent underestimation of target distance. Exploratory analyses suggested that performance varied when the target's vertical position (height above the ground) was manipulated and also differed as a function of lateral target position. Together, these findings show that multisensory integration does not only alter the speed of attentional selection but also the perceived distance of a target. This extends pip-and-pop effects from attentional facilitation to spatial perception biases. In particular, it demonstrates that spatially mismatched but temporally synchronous audiovisual events can bias distance estimates in immersive environments.
In this narrative (historical) review, the emerging literature documenting those crossmodal correspondences that involve material properties such as texture and viscosity, as well as object properties, including shape and size, are both summarized and critically evaluated. Emphasis is placed on those cues to material properties that are available to the nonvisual senses. In this review, we explicitly connect the literature on multisensory shitsukan with the growing literature on crossmodal correspondences, specifically those involving material properties. This narrative historical review also draws attention to the possible material dependence of crossmodal correspondences - meaning, for example, that the existence of specific shape-taste crossmodal correspondences may depend on the material from which an object appears to have been made. This suggestion can be framed in terms of the context dependency of crossmodal correspondences, as has also been highlighted elsewhere in the literature.
Perceiving postural instability accurately is crucial for fall prevention. While sensory integration of visual, vestibular, and somatosensory inputs is known to influence balance, the specific impact of high-consequence visual contexts, such as exposure to height, remains under-investigated due to safety constraints in physical environments. This study serves as a proof-of-concept investigation into the use of virtual reality (VR) for manipulating visual context during a postural temporal-order judgement task. In Experiment 1, participants performed the task in real-world conditions (eyes closed and eyes open). Perceived onset of instability was delayed in both conditions (eyes closed: 25.78 ms; eyes open: 12.33 ms), but these did not differ significantly from true simultaneity. Experiment 2 used VR to safely place participants at the edge of a virtual skyscraper. While perceptual delays remained nonsignificant, precision increased significantly in VR (26.89% increase) compared to the real-world eyes-open condition. These results suggest that while the perceived timing of instability is robust, the presence of a high-arousal visual context in VR enhances the precision of multisensory decision-making. As a foundational step in validating VR-based psychophysical balance assessments, these findings demonstrate the feasibility of using virtual environments to study complex sensory motor integration that is difficult to replicate in naturalistic settings.
The ability to integrate information from different sensory modalities, such as vision and hearing, enables us to perceive and understand complex stimuli. The relevance of this ability can be seen in speech perception, for example, which requires the integration of auditory (e.g., voice) and visual (e.g., lip movements) information. Recent studies have shown that the accuracy of this integration process, and thus speech comprehension, can be improved by brief audiovisual feedback training. However, the long-term stability as well as associated cognitive mechanisms are not yet fully understood. In the present study, we evaluated the effect of a short audiovisual feedback training on multisensory integration and speech perception using three different paradigms. Audiovisual integration was evaluated at three distinct time points: pretraining, immediately post-training and 12 days later. To investigate the relationship between training efficacy and specific cognitive abilities, participants underwent the German version of the Wechsler Adult Intelligence Scale (WAIS-IV), which assesses language comprehension, perceptional reasoning, processing speed, working memory and overall intelligence quotient. Notably, our results indicate a sustained enhancement in audiovisual integration and speech perception that persisted for up to 12 days following the training. No associations with cognitive abilities were found. The findings indicate that a brief audiovisual feedback training enables long-term improvements in multisensory integration, independent of cognitive abilities.
Sensory substitution transforms information normally received by impaired senses into alternate sensory inputs, leveraging brain plasticity to provide better access to visual information for individuals with sensory limitations. Visual to Auditory Sensory Substitution (VASS) converts visual data into auditory cues, aiding individuals with visual impairments. Despite prior research validating sensory substitution, effective VASS training programs remain underdeveloped. This study aims to address this gap by designing and evaluating a VASS training program based on educational theories like conceptual learning and information processing. A training program was developed for learning 23 shapes and figures, incorporating instructional strategies from educational psychology. The learning tasks were organized into modules that included review sessions, concept learning, repetitive practice, and assessments. The program's impact was assessed through a randomized experiment comparing a theory-based program with traditional repetitive training methods. Short-term and sustained outcomes were evaluated using midpoint test, immediate post-test, delayed post-test, and transfer test. The study involved adult learners without visual or auditory impairments to promote efficient program development and effectiveness evaluation. Statistically significant improvements were observed in the experimental group's midpoint and immediate post-test scores compared to the control group. However, no significant differences were found in the delayed post-test and transfer test. The findings suggest that integrating instructional design and learning strategies can lead to more effective training programs to support rehabilitation efforts for individuals with visual impairments. Further implications for rehabilitation are discussed.
Auditory and informational stimuli are part of the broader multisensory experience of eating. Research shows that foreign accents can decrease the credibility and appeal of products in commercial contexts such as ads, customer service, or sales pitches, but few studies have tested regional accents, and none have tested post-consumption effects in the liking and sensory perception of food. Three studies (two tastings, one online) were conducted with audio descriptions in four Swedish accents (Southern, Western, Eastern, Northern), three different food stimuli (vegan cream cheese sandwich, pork and vegan dried sausages, pictures of regional dishes) and a total of 805 participants. The results revealed that while the speaker was perceived differently depending on his accent, the effects were not transferred to consumer liking, willingness to try (WTT), or sensory and emotional associations to the food stimuli. Altogether, they suggest that regional diversity of voices in ads and jobs such as salespeople and waiters are unlikely to influence the preference and perception of food products.
Although taste and touch are processed through separate sensory channels, both elicit strong emotional responses that may share an underlying affective organization. This study investigated whether gustatory and tactile experiences are structured within a common emotional coordinate system defined by valence and arousal. To address this question, we re-analyzed two previously collected datasets, one for taste-based emotion ratings and one for tactile emotion ratings. In two independent datasets, 30 participants evaluated four taste stimuli (sweet, sour, bitter, and salty) on ten affective adjectives, while a separate group of 27 participants rated four tactile textures (rough-hard, rough-soft, smooth-hard, and smooth-soft) on twenty affective adjectives. Principal component analyses of the group-averaged emotion matrices revealed comparable two-dimensional affective spaces in both modalities, corresponding to pleasant-unpleasant and high-low arousal dimensions. Procrustes alignment and consensus mapping demonstrated strong geometric correspondence between taste and touch, with sweet aligning most closely with smooth-soft, sour with smooth-hard, salty with rough-hard, and bitter with rough-soft. Quantitatively, the Procrustes distance was low ( d = 0.094), the mean cosine similarity across pairs was 0.70, and the Mantel test showed a significant correlation ( r = 0.90, p = 0.04), confirming robust cross-modal affective similarity. A combined multidimensional scaling solution further integrated both modalities into a unified affective map with negligible stress. Together, these results provide converging evidence that affective responses to taste and touch follow a shared low-dimensional structure, supporting the idea that affective responses across modalities can be represented within a common evaluative coordinate space. This finding highlights emotion as a unifying representational framework that bridges distinct perceptual domains and informs multisensory theories of affect.
Evidence suggests musicians have enhanced audiovisual and emotion recognition abilities. However, these two lines of research have generally been separated in the literature, despite these processes being similarly altered in certain populations (e.g., autism, schizophrenia). The current systematic review presents a comprehensive picture of the effect of music training on behavioural and neural changes in audiovisual and emotion recognition processes, to better understand where they might overlap or share any similarities. It additionally assessed the impact of different music training factors (i.e., training onset, length, type of musical instrument and the type of research task). Finally, this review aimed to produce a clearer understanding of whether the effects of music training extend beyond the music and sound domain. Following PRISMA guidelines, 64 papers were identified, of which 41 examined audiovisual processing, 20 investigated emotion processing, and three examined both processes. The available evidence revealed a consistent musician's advantage for some audiovisual processes (e.g., audiovisual temporal correspondence), with some evidence that this advantage extended beyond the music domain. Consistent musician's advantages were also found for processing basic emotions from speech prosody, with some evidence that this extended to complex emotions. A shared brain network for these effects was identified comprising the anterior cingulate cortex and superior frontal gyrus. Together, our findings suggest that audiovisual and emotion recognition processes share a number of similarities in how music training can shape them. Further research should directly explore the combined effect of music training on multisensory and emotion recognition to inform effective music interventions aimed at enhancing these processes.
We sometimes feel compelled to touch objects after having looked at them. Is this because touch gives us more accurate information, or is it simply because we tend to trust what we feel more than what we see? To disentangle these possibilities, participants were subjected to a variant of the 'Vertical-Horizontal' illusion where the length of a vertical bar is consistently overestimated relative to the length of a horizontal bar. Our variant focuses on points of subjective equality where the bars almost look the same, and possesses two key characteristics: it creates ambiguity as stimuli can be subjectively similar yet objectively distinct and operates both in vision and touch. In a forced resampling bimodal experiment, participants inspected the stimuli using both vision and touch, judged which bar appeared longer, and rated their confidence. They then re-examined the same objects using either vision or touch, as instructed, and provided a second judgement with another confidence rating. Resampling did not significantly improve the overall accuracy, but participants were more likely to change their mind when using touch. In a second, free resampling bimodal experiment, participants chose their preferred modality for reinspection. Neither modality led to more accurate responses, but participants chose to resample by touch increasingly more frequently when the vertical and horizontal bars' lengths were closer in perceived size, leading to higher ambiguity. Additionally, they were more likely to change their mind under these ambiguous conditions. These findings suggest a selective bias toward touch in case of ambiguity, even when it offers no objective advantage over vision.
Vection (i.e., the experience of self-motion in the absence of actual motion), has traditionally been considered a visual phenomenon. However, recent work on yaw vection (i.e., illusory rotations around the vertical axis), suggests that auditory cues may contribute to vection as well; specifically, if the sounds are similar to those heard in the real world and come from a sound source that typically remains in the same spatial location. In the present study, we sought to determine if roll vection (i.e., illusory rotations around the longitudinal axis) is also enhanced by auditory cues. To test this possibility, we had 44 participants experience three different combinations of sensory cues (audiovisual, visual-only and auditory-only), which were presented via virtual reality a total of three times each. We found that participants experienced vection sooner and more convincingly the first two times they were exposed to the audiovisual condition relative to the visual-only and auditory-only conditions. However, there was no difference between the audiovisual and visual-only condition the third time they experienced them. Regardless of the number of times the participants experienced each condition, the audiovisual and visual-only conditions were always characterized by higher convincingness ratings and lower onset latencies than the auditory-only condition. In tandem, these results suggest that audiovisual stimuli can indeed elicit roll vection that starts earlier and is more convincing than unisensory variants of the stimuli (i.e., visual-only and auditory-only). However, the benefit of including auditory cues diminishes over the course of repetitions suggesting that the brain may down-weight auditory cues. Future studies are needed to better characterize the mechanisms underlying this finding and other factors which may impact this effect (e.g., sound type and field-of-view size).
In this Introduction, we have the pleasure of introducing the twelve articles of this Special Issue of Multisensory Research celebrating the life and works of Vincent Hayward. Vincent was a prolific scientist, collaborator, and colleague. As you will see by the variety of contributed papers, his influence spanned several fields and topics; from engineering to neurophysiology; from skin mechanics to olfactory metacognition. We and many others had the pleasure of knowing and working with Vincent. His boundless curiosity shines through in the papers of this Special Issue, and hopefully in our Introduction as well. Though gone, he is not forgotten; His legacy and influence lives on in the hearts and minds of colleagues studying the (neuro)science of body perception.
Tapping surfaces with a tool tip elicits both sound and haptic information. Sound and haptic information is captured by a microphone and an accelerometer attached to the tip, respectively. In relation to the task of distinguishing objects, this paper investigates the following two questions: (1) how do both the signals (sound and acceleration) help us to assess the surface individually? (2) How does the integration of both modalities affect this assessment? We approach the problem of texture assessment as an unsupervised learning problem. For this purpose, perceptual filter banks are designed based on Weber's law of frequency perception for both modalities to extract the corresponding features. Furthermore, we introduce a symmetric KL divergence-based texture similarity metric, which helped us to compare the sound and haptic (acceleration signals) modalities. Based on our similarity metric comparisons, we argue that proximity between object types is preserved across modalities. Finally, using a permutation test-based approach, we demonstrate that the complementarity of sound domain information to the haptic domain varies depending on the type of object.