Reverberation, a ubiquitous feature of real-world acoustic environments, exhibits statistical regularities that human listeners leverage to self-orient, facilitate auditory perception, and understand their environment. Despite the extensive research on sound source representation in the auditory system, it remains unclear how the brain represents real-world reverberant environments. Here, we characterized the neural response to reverberation of varying realism by applying multivariate pattern analysis to electroencephalographic (EEG) brain signals. Human listeners (12 males and 8 females) heard speech samples convolved with real-world and synthetic reverberant impulse responses and judged whether the speech samples were in a "real" or "fake" environment, focusing on the reverberant background rather than the properties of speech itself. Participants distinguished real from synthetic reverberation with ∼75% accuracy; EEG decoding reveals a multistage decoding time course, with dissociable components early in the stimulus presentation and later in the perioffset stage. The early component predominantly occurred in temporal electrode clusters, while the later component was prominent in centroparietal clusters. These findings suggest distinct neural stages in perceiving natural acoustic environments, likely reflecting sensory encoding and higher-level perceptual decision-making processes. Overall, our findings provide evidence that reverberation, rather than being largely suppressed as a noise-like signal, carries relevant environmental information and gains representation along the auditory system. This understanding also offers various applications; it provides insights for including reverberation as a cue to aid navigation for blind and visually impaired people. It also helps to enhance realism perception in immersive virtual reality settings, gaming, music, and film production.
Experience-based plasticity of the human cortex mediates the influence of individual experience on cognition and behavior. The complete loss of a sensory modality is among the most extreme such experiences. Investigating such a selective, yet extreme change in experience allows for the characterization of experience-based plasticity at its boundaries. Here, we investigated information processing in individuals who lost vision at birth or early in life by probing the processing of braille letter information. We characterized the transformation of braille letter information from sensory representations depending on the reading hand to perceptual representations that are independent of the reading hand. Using a multivariate analysis framework in combination with functional magnetic resonance imaging (fMRI), electroencephalography (EEG), and behavioral assessment, we tracked cortical braille representations in space and time, and probed their behavioral relevance. We located sensory representations in tactile processing areas and perceptual representations in sighted reading areas, with the lateral occipital complex as a connecting ‘hinge’ region. This elucidates the plasticity of the visually deprived brain in terms of information processing. Regarding information processing in time, we found that sensory representations emerge before perceptual representations. This indicates that even extreme cases of brain plasticity adhere to a common temporal scheme in the progression from sensory to perceptual transformations. Ascertaining behavioral relevance through perceived similarity ratings, we found that perceptual representations in sighted reading areas, but not sensory representations in tactile processing areas are suitably formatted to guide behavior. Together, our results reveal a nuanced picture of both the potentials and limits of experience-dependent plasticity in the visually deprived brain.
Visual deprivation does not silence the visual cortex, which is responsive to auditory, tactile, and other nonvisual tasks in blind persons. However, the underlying functional dynamics of the neural networks mediating such crossmodal responses remain unclear. Here, using braille reading as a model framework to investigate these networks, we presented sighted (N=13) and blind (N=12) readers with individual visual print and tactile braille alphabetic letters, respectively, during MEG recording. Using time-resolved multivariate pattern analysis and representational similarity analysis, we traced the alphabetic letter processing cascade in both groups of participants. We found that letter representations unfolded more slowly in blind than in sighted brains, with decoding peak latencies ~200 ms later in braille readers. Focusing on the blind group, we found that the format of neural letter representations transformed within the first 500 ms after stimulus onset from a low-level structure consistent with peripheral nerve afferent coding to high-level format reflecting pairwise letter embeddings in a text corpus. The spatiotemporal dynamics of the transformation suggest that the processing cascade proceeds from a starting point in somatosensory cortex to early visual cortex and then to inferotemporal cortex. Together our results give insight into the neural mechanisms underlying braille reading in blind persons and the dynamics of functional reorganization in sensory deprivation.
Active echolocation allows blind individuals to explore their surroundings via self-generated sounds, similarly to dolphins and other echolocating animals. Echolocators emit sounds, such as finger snaps or mouth clicks, and parse the returning echoes for information about their surroundings, including the location, size, and material composition of objects. Because a crucial function of perceiving objects is to enable effective interaction with them, it is important to understand the degree to which three-dimensional shape information extracted from object echoes is useful in the context of other modalities such as haptics or vision. Here, we investigated the resolution of crossmodal transfer of object-level information between acoustic echoes and other senses. First, in a delayed match-to-sample task, blind expert echolocators and sighted control participants inspected common (everyday) and novel target objects using echolocation, then distinguished the target object from a distractor using only haptic information. For blind participants, discrimination accuracy was overall above chance and similar for both common and novel objects, whereas as a group, sighted participants performed above chance for the common, but not novel objects, suggesting that some coarse object information (a) is available to both expert blind and novice sighted echolocators, (b) transfers from auditory to haptic modalities, and (c) may be facilitated by prior object familiarity and/or material differences, particularly for novice echolocators. Next, to estimate an equivalent resolution in visual terms, we briefly presented blurred images of the novel stimuli to sighted participants (N = 22), who then performed the same haptic discrimination task. We found that visuo-haptic discrimination performance approximately matched echo-haptic discrimination for a Gaussian blur kernel σ of ~2.5°. In this way, by matching visual and echo-based contributions to object discrimination, we can estimate the quality of echoacoustic information that transfers to other sensory modalities, predict theoretical bounds on perception, and inform the design of assistive techniques and technology available for blind individuals.
Experience shapes the visual brain, but the degree to which its functional organization is plastic remains unknown. A unique opportunity to detect the boundaries of functional plasticity is by measuring brain plasticity in individuals born with congenital blindness. We used Braille letter reading to probe the brains of blind individuals. Driven by the analogy to visual processing that proceeds from viewing condition-dependent to viewing condition-independent representations in recognition, we investigated how the Braille letter processing proceeds from a hand-dependent to a hand-independent representation. For this we measured fMRI (N=15) and EEG (N=11) while congenitally blind participants read Braille letters with either their left or right index finger. For both imaging modalities, we applied equivalent multivariate classification schemes: a) to assess hand-dependent letter representations we classified between letter pairs trained and tested within the same hand; b) to assess hand-independent letter representations we trained and tested classifiers across different hands. Finally, we integrated spatial and temporal information using EEG-fMRI representation fusion. The fMRI results indicated hand-dependent representations across somatosensory areas, intraparietal sulcus, insula, early visual cortex and ventral areas. Hand-independent information was present across early visual cortex, ventral areas, and insula. This reveals a transformation from hand-dependent to hand-independent representations along known pathways for both tactile object recognition (S1, S2, parietal cortex, insula) and sighted reading (early visual cortex, LFA, VWFA). The EEG results showed that hand-dependent representations emerged earlier than hand-independent representations. This suggests hierarchical processing of Braille letters with somatosensory, hand-dependent information being processed before rather high-level, hand-independent information. The integrated spatiotemporal analysis revealed that hand-dependent representations emerge before hand-independent representations in the same network, encompassing lateral occipital cortex and intraparietal sulcus. Together our findings reveal the spatio-temporal dynamics of Braille letter representations, indicating functional reorganization of the visual brain in mapping tactile sensory signals to meaning.
Echolocation is an active sensing strategy that some blind people use to detect, discriminate and localize objects in their surroundings. Trained echolocators emit tongue clicks and may vary their clicking pattern dynamically to improve perception under challenging circumstances. However, it is unknown how echoacoustic information is integrated across individual samples (clicks) and how individual echoes are represented neurally. To address these questions, here we recorded the brain activity of blind and sighted individuals using EEG while they performed an echoacoustic localization task. On each trial, subjects listened to a train of 2, 5, 8 or 11 synthesized mouth clicks, and spatialized echoes from a reflecting object located at azimuths of ±5° to ±25° relative to the midsagittal plane. The task was to report whether the echo reflector was to the left or right of center. We hypothesized that the number of clicks in each trial, in addition to the echo azimuth, would modulate performance. The blind expert performed at over 93%, with lateralization thresholds decreasing linearly from 2- to 8-click trials; sighted controls performed at chance, with no effect of echo eccentricity or click count, although they easily lateralized the echoes when the emitted click was removed. Left vs. right location was reliably decoded from the EEG response in both groups after only one click. In the sighted group, perceptual reports were decoded more reliably in the last two clicks of a trial relative to the first two, suggesting a cumulative perceptual decision-making process independent of the stimulus representation. In proficient blind observers, successive click-echo samples linearly sharpen echoacoustic representations until saturation; in novice sighted controls, the spatial information in the EEG response was unavailable to conscious access. These results suggest that echolocation expertise relies on extracting echoes from other masking sounds and integrating them across samples.
Braille is a haptic modality based on a system of raised dots to represent text. Analogously to eye movements in visual reading, braille readers move the reading hand(s) over the printed material to acquire text. In contrast to visual reading, bimanual braille reading raises the question of each hand’s contributing role, and the possibility of perceptual mechanisms distinct from visual reading. Here we analyzed the hand movement patterns of blind braille readers in order to examine the relationship between reading style, hand kinematics, and reading speed. Participants read standardized IReST text passages aloud while their hand movements were recorded with a specialized tracking system. Preliminary results suggest that reading styles strongly affected hand kinematics and performance. Participants who used a more independent hand movement style (e.g. scissors) completed trials faster on average than those who used more interdependent style (e.g. parallel or left marks). These styles were also characterized by lower intermanual correlation in kinematic markers such as regressive movements. In addition, we found that scissors-style readers were disproportionately likely to exhibit simultaneous disjoint reading, in which the two hands read different parts of the text in parallel. Notably, this suggests a neural “memory buffer” mechanism distinct from visual print or serial braille reading, as input acquired in parallel is sorted on the fly to reconstruct a serial text stream. Taken together, our results quantitatively support previous work suggesting that the ability to use independent hand movements may be an important factor in the development of efficient braille reading skills. While further research is needed to fully understand the mechanisms underlying these effects, our results provide new insights into the benefits of bimanual braille reading strategies and may have implications for the teaching of braille to blind individuals.
How does the auditory system categorize natural sounds? Here we apply multimodal neuroimaging to illustrate the progression from acoustic to semantically dominated representations. Combining magnetoencephalographic (MEG) and functional magnetic resonance imaging (fMRI) scans of observers listening to naturalistic sounds, we found superior temporal responses beginning ∼55 ms post-stimulus onset, spreading to extratemporal cortices by ∼100 ms. Early regions were distinguished less by onset/peak latency than by functional properties and overall temporal response profiles. Early acoustically-dominated representations trended systematically toward category dominance over time (after ∼200 ms) and space (beyond primary cortex). Semantic category representation was spatially specific: Vocalizations were preferentially distinguished in frontotemporal voice-selective regions and the fusiform; scenes and objects were distinguished in parahippocampal and medial place areas. Our results are consistent with real-world events coded via an extended auditory processing hierarchy, in which acoustic representations rapidly enter multiple streams specialized by category, including areas typically considered visual cortex.
Recent research has explored the use of active echolocation by blind individuals, who, by generating mouth-clicks, elicit echoes and use them to perceive and interact with their surroundings. In prior work we showed that expert practitioners can distinguish the positions of objects separated by as little as ~1.5°, the approximate threshold of visual letter recognition at 35° retinal eccentricity. They can also echolocate household-sized objects, then distinguish them haptically from a distractor with significantly above-chance accuracy (~60%). Here we investigated whether the spatial resolution of crossmodal echo-haptic object discrimination is similar to that measured for localization. We found that blindfolded sighted participants tested on the same crossmodal match-to-sample design performed similarly, but with greater inter-individual variability. Performance was similar for both common household objects and novel (Lego) objects of arbitrary shape. This suggests that some coarse object information a) is available to both expert blind and novice sighted echolocators, b) transfers from auditory to haptic modalities, c) is not dependent on prior object familiarity, and d) may require a larger angular size than was subtended by our test objects. Thus, we repeated the match-to-sample experiments using stimuli enlarged by 50% along each dimension. Preliminary results do not show improved performance with larger object size; feedback after each trial in future sessions may improve accuracy. Next, we aimed to directly estimate the equivalent visual resolution of echoic object perception. In a pilot experiment, sighted participants examined target objects visually at 35° eccentricity and, subsequently, identified the target haptically. Performance was ~85%, suggesting that haptic recognition is better informed by visual object information at 35° than by object echoes at the scales we tested. Manipulating visual blur to equate visual and echoic performance will reveal more precisely the spatial resolution of echo-based object perception.
Scene analysis is fundamental to successfully perceiving, interacting with, and navigating through the environment. In the absence of visual information, scene analysis relies largely on auditory signals such as reverberation, the aggregated acoustic reflections from multiple nearby surfaces. Sound sources and the spaces surrounding them are separably coded (Teng et al., 2017), an operation contingent on spectral and temporal statistics of the reverberant background (Traer & McDermott, 2016). It remains unclear how these perceptual heuristics develop or how they are influenced by experience. Here we investigated whether visual experience modulates reverberant perception. The experience-dependence hypothesis predicts higher fidelity in early- and congenitally blind listeners, who are more heavily dependent on acoustic cues. Alternatively, the calibration hypothesis predicts that, deprived of a visual "scaffold," blind listeners would be impaired in reverberant coding. To test these predictions, we conducted an online experiment in which sighted and blind participants listened to pairs of spoken sentences, each convolved with a reverberant impulse response (IR). The IRs were recorded from a real-world space or synthesized to match or deviate from the temporal and spectral features of that same space. Manipulations included the temporal decay rate and the spectral distribution of the IRs. The task was a 2AFC judgment to indicate which of the spaces was "real." We found that blind as well as sighted participants were highly sensitive to temporal deviations from ecologically valid reverberation, and less sensitive to spectral deviations. Interestingly, while sighted listeners reliably distinguished spectrally altered reverberation at above-chance levels, preliminary results indicate markedly reduced performance in blindness. Some below-chance performance in these conditions suggests a sensitivity to spectral alterations (cf. Voss et al., 2011), but an ambiguity in assigning them to the correct category. Taken together, our results suggest that visual experience modulates representations of auditory environmental statistics.
ABSTRACT As the human brain transforms incoming sounds, it remains unclear whether semantic meaning is assigned via distributed, domain-general architectures or specialized hierarchical streams. Here we show that the spatiotemporal progression from acoustic to semantically dominated representations is consistent with a hierarchical processing scheme. Combining magnetoencephalography (MEG) and functional magnetic resonance imaging (fMRI) patterns, we found superior temporal responses beginning ~80 ms post-stimulus onset, spreading to extratemporal cortices by ~130 ms. Early acoustically-dominated representations trended systematically toward semantic category dominance over time (after ~200 ms) and space (beyond primary cortex). Semantic category representation was spatially specific: vocalizations were preferentially distinguished in temporal and frontal voice-selective regions and the fusiform face area; scene and object sounds were distinguished in parahippocampal and medial place areas. Our results are consistent with an extended auditory processing hierarchy in which acoustic representations give rise to multiple streams specialized by category, including areas typically considered visual cortex.
The fusiform face area responds selectively to faces and is causally involved in face perception. How does face-selectivity in the fusiform arise in development, and why does it develop so systematically in the same location across individuals? Preferential cortical responses to faces develop early in infancy, yet evidence is conflicting on the central question of whether visual experience with faces is necessary. Here, we revisit this question by scanning congenitally blind individuals with fMRI while they haptically explored 3D-printed faces and other stimuli. We found robust face-selective responses in the lateral fusiform gyrus of individual blind participants during haptic exploration of stimuli, indicating that neither visual experience with faces nor fovea-biased inputs is necessary for face-selectivity to arise in the lateral fusiform gyrus. Our results instead suggest a role for long-range connectivity in specifying the location of face-selectivity in the human brain.
How does the FFA arise in development, and why does it develop so systematically in the same location across individuals? Preferential fMRI responses to faces arise early, by around 6 months of age in humans (Deen et al., 2017). Arcaro et al (2017) have further shown that monkeys reared without ever seeing a face show no face-selective patches, and regions that later become face selective are correlated in resting fMRI with foveal retinotopic cortex in newborn monkeys. These findings have been taken to argue that 1) seeing faces is necessary for the development of face-selective patches and 2) face patches arise in previously fovea-biased cortex because early experience with faces is foveally biased. Here we present evidence against both these claims. We scanned congenitally blind subjects (N = 6) with fMRI while they performed a one-back haptic shape discrimination task, sequentially palpating 3D-printed photorealistic models of faces, hands, mazes and chairs in a blocked design. Five out of six participants showed significantly higher responses to faces than other categories in the lateral fusiform gyrus (see Figure 1 A,B). Overall, the face selectivity of fusiform regions for tactile faces in congenitally blind participants was comparable to the face selectivity in sighted subjects (N = 8) for videos of the same stimuli rotating in depth (see Figure1 C,D). Evidently, the development of strongly face-selective responses in the lateral fusiform gyrus does not require a) seeing faces, b) foveating faces, or c) perceptual expertise with faces. We speculate that face selectivity in congenitally blind participants reflects either amodal representations of shape and/or the interpretation of faces as social stimuli (Van den Hurk et al, 2017;Powell et al., 2018).
Echolocating organisms ensonify their surroundings, then extract object and spatial information from the echoes. This behavior has been observed in some blind humans, but the computations underlying this strategy remain extremely poorly understood. Here we tracked the movements and echo emissions of an expert blind echolocator performing a target detection and localization task. We found that the precision of responses as well as target acquisition movements depended significantly on the size of the target and availability of active echo cues. The distribution of click directions suggested that the maximal energy of each click was always directed at the target. Our results pave the way toward characterizing human echolocation in the context of other active sensing behaviors, constraining the types of perceptual mechanisms mediating its behavior, and at a practical level, building a quantitative evidence base for optimizing therapeutic training interventions.
It is well established that areas of high-level visual cortex are selectively driven by visual categories such as places, objects, and faces. These areas include the scene-selective parahippocampal place area (PPA), occipital place area (OPA), and retrosplenial cortex (RSC), the object-selective lateral occipital complex (LOC), and the face-selective fusiform face area (FFA). Here we sought to determine whether neural representations in these regions are evoked without visual input, and if so, how these representations emerge across space and time in the human brain. Using an event-related design, we presented participants (n = 15) with 80 real-world sounds from various sources (animals, human voices, objects, and spaces) and instructed them to form a corresponding mental image with their eyes closed. To trace the emergence of neural representations at both the millisecond and millimeter level, we acquired spatial data from functional magnetic resonance imaging (fMRI) and temporal data from magnetoencephalography (MEG) in independent sessions. Regions of interest (ROIs) were independently localized in auditory and visual cortex. Using similarity-based fusion (Cichy et al., 2014), we correlated MEG and fMRI data to reveal correspondence between temporal and spatial neural dynamics. Our results reveal neural representations evoked from auditory stimuli emerge rapidly in the face-selective FFA, in addition to voice-selective auditory areas (< 100ms). In contrast, representations in scene- and object-selective cortex emerged later (>130ms). We found no evidence for neural representations in early visual cortex, as expected. By tracing the emergence of neural representations in cascade across the human brain, we therefore reveal the differential spatiotemporal neural dynamics of these representations in high-level visual cortex evoked in the absence of visual input. Our findings thus support a multimodal neural framework for sensory representations, and track these emerging neural representations across space and time in the human brain.
When read by blind individuals, tactile braille characters undergo a cascade of transformations over space, time, and representational format. This processing stream is known to recruit typically visual cortical regions as it changes somatosensory dot patterns to semantically meaningful representations. To elucidate the poorly understood spatiotemporal, as well as representational, dynamics of this reorganized functional network, we applied multivariate decoding and representational similarity analysis to magnetoencephalography (MEG) recordings of blind participants' brain responses to braille letters presented to the finger pads. Following previous work suggesting largely idiosyncratic patterns in a sample of braille readers, we presented single alphabetical braille letters in random order during MEG recording of early-blind, braille-proficient individuals in multiple sessions. Subjects performed a 1-back task in which they passively read presented letters and responded via button press to occasional repeated trials, which we excluded from further analysis. To increase the spatial specificity of the signal, we extracted left and right sensorimotor, early "visual" (EVC), and fusiform regions of interest (ROIs) from individual MRI anatomical scans. We then used MVPA to decode letter identity pairwise over time and construct a representational dissimilarity matrix (RDM) of pairwise relationships between letter signals for each time point. Earliest and strongest decoding signals were found specific to the sensorimotor ROIs contralateral to the stimulated finger. EVC and fusiform ROIs showed later onsets and noisier, more sustained representations. Early representational patterns as operationalized by model RDMs are consistent with a sensitivity to low-level letter complexity (e.g., number of dots) in somatosensory cortex, while later representations in downstream ROIs exhibited weaker adherence to this scheme. Our results offer a window into the sensory-to-semantic transformation of braille stimuli as well as a model for investigating the dynamics of crossmodal plasticity in sensory loss generally. Meeting abstract presented at VSS 2018