Independent traveling remains challenging for blind and visually impaired (BVI) individuals. While the white cane is effective at detecting ground-level obstacles, it provides no information about elevated obstacles or object characteristics. Recent technologies have been designed to support navigation as well as object detection. In our study, we compared the performance of 13 BVI participants who separately used two secondary electronic travel aids (ETAs) versus cane use alone. One ETA was a camera-based mobility vest (NOA), and the other was an ultrasonic sensor-based wearable (BuzzClip). Participants completed an obstacle avoidance task with both ETAs and an object detection task using two versions of NOA’s object-finding functionality. Quantitative performance measures and semi-structured interviews were collected. NOA resulted in enhanced obstacle avoidance. Participants used their canes less and collided less with obstacles when using NOA than the BuzzClip or the white cane alone. NOA resulted in lower frustration and higher perceived performance, as well as greater perceived safety and obstacle detection than the BuzzClip. Object-finding performance outcomes were similar across both versions, suggesting potential benefit from a dynamic combination of approaches tailored for each user. Collectively, these data underscore how ETAs may be integrated into use by the BVI community.
Interception refers to goal-directed motor actions aimed at interacting with moving objects and is essential for both motor co-ordination and social engagement. In childhood, interceptive skills support environmental exploration, peer interaction, and participation in play and sports. For children with visual impairments, the lack of visual cues compromises the development of these skills, potentially limiting motor competence and opportunities for social interaction. Despite its clinical and developmental relevance, research on interception in visually impaired (VI) children is extremely limited. This mini review synthesizes findings from studies on interceptive skills in VI adults, as well as in sighted and VI children. We discuss how vision contributes to interceptive actions, and how alternative sensory pathways can compensate in its absence. We highlight major limitations of current literature, including poor ecological validity, a lack of longitudinal data, and scarce attention to multisensory and social aspects. To address these gaps, we propose future research directions that include cross-sectional and longitudinal studies, multisensory paradigms, and the use of virtual reality technologies to simulate naturalistic environments. These approaches may inform inclusive and rehabilitative interventions that support the motor and social development of VI children through accessible, engaging, and developmentally appropriate interceptive experiences.
Understanding speech in noise can be facilitated by integrating auditory and visual speech cues. Audiovisual temporal acuity, which can be indexed by the temporal binding window (TBW), is critical for this process and can be enhanced through simultaneity judgment training. We hypothesized that multisensory training would narrow the TBW and improve speech understanding in noise. Participants were randomized to receive either training and testing (n = 15) or testing-only (n = 15) over three days. Trained participants demonstrated significant narrowing in their mean TBW size (403ms to 345ms; p = 0.030), whereas control participants did not (409ms to 474ms; p = 0.061). Although there were no group-level changes in word recognition scores, trained participants with larger TBW decreases exhibited larger improvements in auditory word recognition in noise (R2 = 0.291; p = 0.038). Individual differences in responses to training were found to be related to differences in cortical speech processing using functional near-infrared spectroscopy. Low audiovisual-evoked activity in the left middle temporal gyrus (R2 = 0.87; p = 0.006), left angular and superior temporal gyrus (R2 = 0.85; p = 0.006), and visual cortices (R2 = 0.74; p = 0.041) was associated with larger improvements in auditory word recognition after training. Multisensory training transfers benefits to speech comprehension in noise, and this effect may be mediated by upregulating activity in multisensory cortical networks for individuals with low baseline activity.
Autism is a neurodevelopmental condition that presents with significant changes in sensory processing, and which has recently been associated with differences in sensory expectations. One method for measuring sensory expectations (i.e., predictions) is via oddball paradigms, in which a deviant stimulus is presented following a series of repeated stimuli. In EEG signals, this deviance elicits a characteristic mismatch negativity (MMN) response, which acts as a neural signature of deviance detection and perception. Given the growing focus on sensory prediction in autism, a number of studies have now employed the oddball paradigm, with mixed results. We conducted a meta-analysis to better understand the utility of oddball paradigms in evaluating sensory prediction differences in the autism population. A comprehensive literature search queried the PubMed database for empirical auditory and visual oddball studies comparing autistic and non-autistic individuals. Statistical analyses were all conducted in R. We estimated true effect sizes and characterized the effects of various study characteristics on effect size using a multi-level random effects model and robust variance estimation (RVE). Publication bias and study quality were also assessed. Although individual studies have reported differences, the results of this meta-analysis suggest no significant group differences between autistic and non-autistic individuals in auditory or visual oddball perception, recognition, or neural signatures. When used in autism research, auditory and visual oddball MMN responses may not inherently capture changes in sensory prediction, and significant findings may be related more to individual variability than diagnostic group.
In 2024, the United States House of Representatives passed ruling H.R.7213, the Autism CARES Act, which, if passed by the Senate, will reauthorize funding to extant national autism research programs, with an emphasis on including autistic individuals significantly affected by the disorder. This shift toward research inclusion across the autism spectrum clearly highlights the lack of representation in the past. In the field of multisensory integration, it is well documented that there are changes to how autistic individuals integrate stimuli across different sensory modalities, and the relationship between atypical (multi)sensory processing and the core features of autism is well documented. However, much of this research utilizes samples of autistic individuals with high cognitive, verbal, and functional ability. The purpose of this review is to draw attention to disparities in the samples used in multisensory research in autism. We conducted a systematic review of all studies examining multisensory function in autism to date and provide basic descriptive statistics of the studies. We observed that the vast majority of multisensory research is focused on young, low support needs autistic individuals, with very little investigation in autistic individuals with high support needs (HSN). Additionally, we found investigation into the effect of sex or comorbidities to be lacking. We propose methodological improvements addressing gaps in the research in order to make multisensory research in autism more inclusive to HSN autistics.
Audiovisual (AV) (a)synchrony perception, measured by a simultaneity judgment task, provides a proxy measure for the temporal binding window (TBW), the interval of time within which participants perceive individual auditory and visual events as synchronous. The TBW is a sensitive measure showing group level differences across the lifespan. However, the significance of these findings hinges on whether AV synchrony perception is characteristic to an individual and whether it is reliable across sessions. As there is little evidence to this latter assumption, this work aimed to establish test-retest reliability of the TBW. Eighteen participants completed a simultaneity judgment task, twice, on different days, where they judged whether or not the audio and video of word-length speech stimuli were presented at the same or different times. Results showed effects of task familiarity indicated by significantly faster response times observed at the second timepoint. The slope and amplitude asymptote of the TBW also increased, however no change in TBW suggests no change in sensitivity to AV (a)synchrony. The combination of high within-subject correlations (R = 0.71) and substantial intersubject variability provide strong support for the TBW as a robust measure of AV temporal perception, and suggest a conserved mechanistic underpinning within individuals.
Multisensory integration (MSI) is a core neurobehavioral operation that enhances our ability to perceive, decide, and act by combining information from different sensory modalities. This integrative capability is essential for efficiently navigating complex environments and responding to their multisensory nature. One of the powerful behavioral benefits of MSI is in speeding responses. To evaluate this speeding, traditional research in MSI often relies on so-called race models, which predict reaction times (RTs) based on the assumption that information from the different sensory modalities is initially processed independently. When observed RTs are faster than those predicted by these models, it indicates the presence of true convergence and integration of multisensory information prior to the initiation of the motor response. Despite the strong applicability of race models in MSI research, analysis of multisensory RT data often poses challenges for researchers, particularly in managing, interpreting and modeling large datasets or a collection of datasets. To surmount these challenges, we developed a user-friendly graphical user interface (GUI) packaged into a freely available software application that is compatible with both Windows and Mac and that requires no programming expertise. This tool simplifies the processes of data loading, filtering, and statistical analysis. It allows the calculation and visualization of RTs across different sensory modalities, the performance of robust statistical tests, and the testing of race model violations. By integrating these capabilities into a single platform, the CART-GUI facilitates MSI analyses and makes it accessible to a wider range of users, from novice researchers to experts in the field. The GUI’s user-friendly design and advanced analytical features will allow for valuable insights into the mechanisms underlying MSI and contribute to the advancement of research in this domain.
Natural environments are typically multisensory, comprising information from multiple sensory modalities. It is in the integration of these incoming sensory signals that we form our perceptual gestalt that allows us to navigate through the world with relative ease. However, differences in multisensory integration (MSI) ability are found in a number of clinical conditions. Throughout this chapter, we discuss how MSI differences contribute to phenotypic characterization of autism and schizophrenia. Although these clinical populations are often described as opposite each other on a number of spectra, we describe similarities in behavioral performance and neural functions between the two conditions. Understanding the shared features of autism and schizophrenia through the lens of MSI research allows us to better understand the neural and behavioral underpinnings of both disorders. We provide potential avenues for remediation of MSI function in these populations.
The integration of human and artificial intelligence represents a scientific opportunity to advance our understanding of information processing, as each system offers unique computational insights that can enhance and inform the other. The synthesis of human cognitive principles with artificial intelligence has the potential to produce more interpretable and functionally aligned computational models, while simultaneously providing a formal framework for investigating the neural mechanisms underlying perception, learning, and decision-making through systematic model comparisons and representational analyses. In this study, we introduce personalized brain-inspired modeling that integrates human behavioral embeddings and neural data to align with cognitive processes. We took a stepwise approach, fine-tuning the Contrastive Language-Image Pre-training (CLIP) model with large-scale behavioral decisions, group-level neural data, and finally, participant-level neural data within a broader framework that we have named CLIP-Human-Based Analysis (CLIP-HBA). We found that fine-tuning on behavioral data enhances its ability to predict human similarity judgments while indirectly aligning it with dynamic representations captured via MEG. To further gain mechanistic insights into the temporal evolution of cognitive processes, we introduced a model specifically fine-tuned on millisecond-level MEG neural dynamics (CLIP-HBA-MEG). This model resulted in enhanced temporal alignment with human neural processing while still showing improvement on behavioral alignment. Finally, we trained individualized models on participant-specific neural data, effectively capturing individualized neural dynamics and highlighting the potential for personalized AI systems. These personalized systems have far-reaching implications for the fields of medicine, cognitive research, human-computer interfaces, and AI development.
In the real world our body moves within a multisensory environment at different speeds. However, how the speed of our movements influences multisensory perception of time is still unclear. To fill this gap, a new experimental setup that combines the Haply Inverse3 and the MSI Caterpillar devices was developed. The former was used to track participants' arm movements and record their speed; the latter to deliver spatially congruent audio-tactile stimuli during a Temporal Order Judgement (TOJ) task. Participants held the devices in their hands and were asked to move their arms at three different speeds, while performing the TOJ task. Using this multidevice system, the speed of individual participants could be measured. Results showed a positive association between participants' actual speed of movement and precision in TOJ performance, despite no differences between static and active movement conditions. These findings provide valuable insights on the interplay between motor control and perceptual timing, with potential applications for studying development and in rehabilitation protocols for visually impaired participants.
We introduce Perceptual-Initialization (PI), a paradigm shift in visual representation learning that incorporates human perceptual structure during the initialization phase rather than as a downstream fine-tuning step. By integrating human-derived triplet embeddings from the NIGHTS dataset to initialize a CLIP vision encoder, followed by self-supervised learning on YFCC15M, our approach demonstrates significant zero-shot performance improvements, without any task-specific fine-tuning, across 29 zero shot classification and 2 retrieval benchmarks. On ImageNet-1K, zero-shot gains emerge after approximately 15 epochs of pretraining. Benefits are observed across datasets of various scales, with improvements manifesting at different stages of the pretraining process depending on dataset characteristics. Our approach consistently enhances zero-shot top-1 accuracy, top-5 accuracy, and retrieval recall (e.g., R@1, R@5) across these diverse evaluation tasks, without requiring any adaptation to target domains. These findings challenge the conventional wisdom of using human-perceptual data primarily for fine-tuning and demonstrate that embedding human perceptual structure during early representation learning yields more capable and vision-language aligned systems that generalize immediately to unseen tasks. Our work shows that "beginning with you", starting with human perception, provides a stronger foundation for general-purpose vision-language intelligence.
Despite growing evidence of neural and behavioral plasticity following sensory loss, it remains unclear how multisensory processing varies across clinical hearing loss phenotypes. This study investigated visual perception and audiovisual (AV) integration in adults with varying degrees of hearing loss and hearing technology use. Participants included individuals with normal hearing (NH), hearing aid (HA) users, cochlear implant (CI) candidates, and CI users. To assess visual and multisensory processing, we administered a visual temporal order judgment (vTOJ) task, the McGurk illusion, a monosyllabic lipreading task, and an AV word recognition task. Results revealed a trend toward improved visual temporal resolution with increasing hearing loss severity, though this was confounded by age. McGurk illusion responses indicated that the presence of hearing loss decreased auditory weighting, while the severity of hearing loss increased visual weighting. Lipreading performance significantly improved as hearing loss progressed, with CI users outperforming all other groups-possibly due to the use of rehabilitation exercises in the CI clinical protocol. In contrast, AV benefit did not vary systematically with hearing loss, but showed a significant effect of age. Together, these findings suggest that visual performance and visual sensory weighting-more than AV integration per se-are modulated by hearing loss. These differences may reflect underlying plasticity of the cortical regions responsible for processing multisensory input. Furthermore, these findings highlight the potential utility of visual tasks in characterizing sensory phenotypes and informing clinical decision-making for individuals with hearing loss.
Motion perception is a key aspect of sensory processing that enables successful interaction with the environment. While visual motion perception has been extensively studied, little is known about the determinants of auditory motion perception. Our study explores how the perception of auditory motion direction changes with manipulations of low-level stimulus parameters in nonhuman primates (NHPs). Macaque monkeys were trained to perform a 2-AFC task in which they judged the direction of noisy auditory motion stimuli. We systematically manipulated stimulus duration, velocity, and displacement to evaluate their respective influence on motion sensitivity. Displacement had the greatest impact, while the relative influence of duration versus velocity depended upon the duration of the stimulus. These findings suggest that auditory motion direction is most likely processed by a snapshot mechanism, in which stimulus velocity is inferred by sequential snapshots of auditory stimulus location, rather than by velocity-selective motion detectors similar to those found in the visual system. To our knowledge, this study is the first to characterize the influence of low-level stimulus parameters on auditory motion perception in awake, behaving NHPs, and forms the basis for future neurophysiological investigations.
The ability to perceive and interact with the surrounding environment is essential for human adaptation. For visually impaired individuals, touch plays a fundamental role in compensating for the absence of vision, enabling spatial awareness and navigation. Given the importance of touch in this population, various experiments have been carried out to explore their tactile abilities. However, reliable tools for assessing tactile perception of orientation and direction remain limited. To address this gap, we developed TaK (Tactile Knob), a tactile goniometer designed to measure how individuals perceive object orientation and direction. TaK features a manually adjustable spinning arrow, a 3D-printed base, electronic components for precise angle measurement, and a serial port to minimize reporting errors. We conducted two experiments to evaluate its effectiveness. In the first, sighted participants performed an orientation reproduction task using both TaK and a commercial knob. Results showed no significant differences between the two devices, confirming TaK's reliability in measuring tactile orientation perception, comparable to existing commercial tools. The second experiment compared sighted and visually impaired participants using TaK. Again, no significant differences were found, demonstrating its inclusivity and validity for assessing tactile perception in spatial tasks. These findings highlight TaK's potential as a reliable and accessible tool for spatial perception research and assistive technology development, contributing to a better understanding of tactile processing in both sighted and visually impaired individuals.
Although not considered a core feature of autism, autistic children often present with difficulties in reading comprehension, which is a multisensory process involving translation of print to speech sounds (i.e., decoding) and interpreting words in context (i.e., language comprehension). This study tested the hypothesis that audiovisual integration may explain individual differences in reading comprehension, through its relations with decoding and language comprehension, in autistic and non-autistic children. To test our hypothesis, we conducted a concurrent correlational study involving 50 autistic and 50 non-autistic school-aged children (8–17 years of age) matched at the group level on biological sex and chronological age. Participants completed a battery of tests probing their reading comprehension, decoding, and language comprehension, as well as a psychophysical task assessing audiovisual integration as indexed by susceptibility to the McGurk illusion. A series of regression analyses was carried out to test relations of interest. Audiovisual integration was significantly associated with reading comprehension, decoding, and language comprehension, with moderate-to-large effect sizes. Mediation analyses revealed that the relation between audiovisual integration and reading comprehension was completely mediated by decoding and language comprehension, with standardized indirect effects indicating significant mediation through both pathways. These associations did not vary according to diagnostic group. This work highlights the potential role of audiovisual integration in language and literacy development and underscores the potential for multisensory-based interventions to improve reading outcomes in autistic and non-autistic children. Future research should employ longitudinal designs and more diverse samples to replicate and extend these findings.
Introduction:The speech-to-song illusion is a robust effect where repeated speech induces the perception of singing; this effect has been extended to repeated excerpts of environmental sounds (sound-to-music effect). Here we asked whether repetition could elicit musical percepts in cochlear implant (CI) users, who experience challenges with perceiving music due to both physiological and device limitations. Methods:Thirty adult CI users and thirty age-matched controls with normal hearing (NH) completed two repetition experiments for speech and nonspeech sounds (water droplets). We hypothesized that CI users would experience the sound-to-music effect from temporal/rhythmic cues alone, but to a lesser magnitude compared to NH controls, given the limited access to spectral information CI users receive from their implants. Results:We found that CI users did experience the sound-to-music effect but to a lesser degree compared to NH participants. Musicality ratings were not associated with musical training or frequency resolution, and among CI users, clinical variables like duration of hearing loss also did not influence ratings. Discussion:Cochlear implants provide a strong clinical model for disentangling the effects of spectral and temporal information in an acoustic signal; our results suggest that temporal cues are sufficient to perceive the sound-to-music effect when spectral resolution is limited. Additionally, incorporating short repetitions into music specially designed for CI users may provide a promising way for them to experience music.
The progress of developing an effective closed-loop neuromodulation system for many neurological pathologies is hindered by the difficulties in accurately capturing a useful representation of a brain’s instantaneous functional state. Existing approaches rely on expert labeling of electroencephalography data to develop biomarkers of neurophysiological pathology. These techniques do not capture the highly complex functional states of the brain that are presumed to exist between labeled states or allow for the likely possibility of variation among identically labeled states. Thus, we propose BrainState, a self-supervised technique to model an arbitrarily complex instantaneous functional state of a brain using neural multivariate timeseries data. Application of BrainState to intracranial electroencephalography data from patients with epilepsy was able to capture diverse pre-seizure states and quantify nuanced effects of neuromodulation. We anticipate that BrainState will enable the development of sophisticated closed-loop neuromodulation systems for a diverse array of neurological pathologies. ### Competing Interest Statement The authors have declared no competing interest.
Autistic youth demonstrate differences in processing multisensory information, particularly in temporal processing of multisensory speech. Extensive research has identified several key brain regions for multisensory speech processing in non-autistic adults, including the superior temporal sulcus (STS) and insula, but it is unclear to what extent these regions are involved in temporal processing of multisensory speech in autistic youth. As a first step in exploring the neural substrates of multisensory temporal processing in this clinical population, we employed functional magnetic resonance imaging (fMRI) with a simultaneity-judgment audiovisual speech task. Eighteen autistic youth and a comparison group of 20 non-autistic youth matched on chronological age, biological sex, and gender participated. Results extend prior findings from studies of non-autistic adults, with non-autistic youth demonstrating responses in several similar regions as previously implicated in adult temporal processing of multisensory speech. Autistic youth demonstrated responses in fewer of the multisensory regions identified in adult studies; responses were limited to visual and motor cortices. Group responses in the middle temporal gyrus significantly interacted with age; younger autistic individuals showed reduced MTG responses whereas older individuals showed comparable MTG responses relative to non-autistic controls. Across groups, responses in the precuneus covaried with task accuracy, and anterior temporal and insula responses covaried with nonverbal IQ. These preliminary findings suggest possible differences in neural mechanisms of audiovisual processing in autistic youth while highlighting the need to consider participant characteristics in future, larger-scale studies exploring the neural basis of multisensory function in autism.