The cochlea behaves like a bank of band-pass filters, segregating information into different frequency channels. Some aspects of perception reflect processing within individual channels, but others involve the integration of information across them. One instance of this is sound localization, which improves with increasing bandwidth. The processing of binaural cues for sound location has been studied extensively. However, although the advantage conferred by bandwidth is clear, we currently know little about how this additional information is combined to form our percept of space. We investigated the ability of cells in the auditory system of guinea pigs to compare interaural level differences (ILDs), a key localization cue, between tones of disparate frequencies in each ear. Cells in auditory cortex believed to be integral to ILD processing (excitatory from one ear, inhibitory from the other: EI cells) compare ILDs separately over restricted frequency ranges which are not consistent with their monaural tuning. In contrast, cells that are excitatory from both ears (EE cells) show no evidence of frequency-specific processing. Both cell types are explained by a model in which ILDs are computed within separate frequency channels and subsequently combined in a single cortical cell. Interestingly, ILD processing in all inferior colliculus cell types (EE and EI) is largely consistent with processing within single, matched-frequency channels from each ear. Our data suggest a clear constraint on the way that localization cues are integrated: cortical ILD tuning to broadband sounds is a composite of separate, frequency-specific, binaurally sensitive channels. This frequency-specific processing appears after the level of the midbrain. SIGNIFICANCE STATEMENT For some sensory modalities (e.g., somatosensation, vision), the spatial arrangement of the outside world is inherited by the brain from the periphery. The auditory periphery is arranged spatially by frequency, not spatial location. Therefore, our auditory perception of location must be synthesized from physical cues in separate frequency channels. There are multiple cues (e.g., timing, level, spectral cues), but even single cues (e.g., level differences) are frequency dependent. The synthesis of location must account for this frequency dependence, but it is not known how this might occur. Here, we investigated how interaural-level differences are combined across frequency along the ascending auditory system. We found that the integration in auditory cortex preserves the independence of the different-level cues in different frequency regions.
Visual displays in passive sonar based on the Fourier spectrogram are underpinned by detection models that rely on signal and noise power statistics. Time-frequency representations specialised for sparse signals achieve a sharper signal representation, either by reassigning signal energy based on temporal structure or by conveying temporal structure directly. However, temporal representations involve nonlinear transformations that make it difficult to reason about how they respond to additive noise. This article analyses the effect of noise on temporal fine structure measurements such as zero crossings and instantaneous frequency. Detectors that rely on zero crossing intervals, intervals and peak amplitudes, and instantaneous frequency measurements are developed, and evaluated for the detection of a sinusoid in Gaussian noise, using the power detector as a baseline. Detectors that rely on fine structure outperform the power detector under certain circumstances; and detectors that rely on both fine structure and power measurements are superior. Reassigned spectrograms assume that the statistics used to reassign energy are reliable, but the derivation of the fine structure detectors indicates the opposite. The article closes by proposing and demonstrating the concept of a doubly reassigned spectrogram, wherein temporal measurements are reassigned according to a statistical model of the noise background.
Predictive accounts of perception have received increasing attention in the past 20 years. Detecting violations of auditory regularities, as reflected by the Mismatch Negativity (MMN) auditory event-related potential, is amongst the phenomena seamlessly fitting this approach. Largely based on the MMN literature, we propose a psychological conceptual framework called the Auditory Event Representation System (AERS), which is based on the assumption that auditory regularity violation detection and the formation of auditory perceptual objects are based on the same predictive regularity representations. Based on this notion, a computational model of auditory stream segregation, called CHAINS, has been developed. In CHAINS, the auditory sensory event representation of each incoming sound is considered for being the continuation of likely combinations of the preceding sounds in the sequence, thus providing alternative interpretations of the auditory input. Detecting repeating patterns allows predicting upcoming sound events, thus providing a test and potential support for the corresponding interpretation. Alternative interpretations continuously compete for perceptual dominance. In this paper, we briefly describe AERS and deduce some general constraints from this conceptual model. We then go on to illustrate how these constraints are computationally specified in CHAINS.
Classical signal detection theory attributes bias in perceptual decisions to a threshold criterion, against which sensory excitation is compared. The optimal criterion setting depends on the signal level, which may vary over time, and about which the subject is naïve. Consequently, the subject must optimise its threshold by responding appropriately to feedback. Here a series of experiments was conducted, and a computational model applied, to determine how the decision bias of the ferret in an auditory signal detection task tracks changes in the stimulus level. The time scales of criterion dynamics were investigated by means of a yes-no signal-in-noise detection task, in which trials were grouped into blocks that alternately contained easy- and hard-to-detect signals. The responses of the ferrets implied both long- and short-term criterion dynamics. The animals exhibited a bias in favour of responding “yes” during blocks of harder trials, and vice versa. Moreover, the outcome of each single trial had a strong influence on the decision at the next trial. We demonstrate that the single-trial and block-level changes in bias are a manifestation of the same criterion update policy by fitting a model, in which the criterion is shifted by fixed amounts according to the outcome of the previous trial and decays strongly towards a resting value. The apparent block-level stabilisation of bias arises as the probabilities of outcomes and shifts on single trials mutually interact to establish equilibrium. To gain an intuition into how stable criterion distributions arise from specific parameter sets we develop a Markov model which accounts for the dynamic effects of criterion shifts. Our approach provides a framework for investigating the dynamics of decisions at different timescales in other species (e.g., humans) and in other psychological domains (e.g., vision, memory).
The ability of the auditory system to parse complex scenes into component objects in order to extract information from the environment is very robust, yet the processing principles underlying this ability are still not well understood. This study was designed to investigate the proposal that the auditory system constructs multiple interpretations of the acoustic scene in parallel, based on the finding that when listening to a long repetitive sequence listeners report switching between different perceptual organizations. Using the “ABA-” auditory streaming paradigm we trained listeners until they could reliably recognize all possible embedded patterns of length four which could in principle be extracted from the sequence, and in a series of test sessions investigated their spontaneous reports of those patterns. With the training allowing them to identify and mark a wider variety of possible patterns, participants spontaneously reported many more patterns than the ones traditionally assumed (Integrated vs. Segregated). Despite receiving consistent training and despite the apparent randomness of perceptual switching, we found individual switching patterns were idiosyncratic; i.e., the perceptual switching patterns of each participant were more similar to their own switching patterns in different sessions than to those of other participants. These individual differences were found to be preserved even between test sessions held a year after the initial experiment. Our results support the idea that the auditory system attempts to extract an exhaustive set of embedded patterns which can be used to generate expectations of future events and which by competing for dominance give rise to (changing) perceptual awareness, with the characteristics of pattern discovery and perceptual competition having a strong idiosyncratic component. Perceptual multistability thus provides a means for characterizing both general mechanisms and individual differences in human perception.
Many sound sources can only be recognised from the pattern of sounds they emit, and not from the individual sound events that make up their emission sequences. Auditory scene analysis addresses the difficult task of interpreting the sound world in terms of an unknown number of discrete sound sources (causes) with possibly overlapping signals, and therefore of associating each event with the appropriate source. There are potentially many different ways in which incoming events can be assigned to different causes, which means that the auditory system has to choose between them. This problem has been studied for many years using the auditory streaming paradigm, and recently it has become apparent that instead of making one fixed perceptual decision, given sufficient time, auditory perception switches back and forth between the alternatives-a phenomenon known as perceptual bi- or multi-stability. We propose a new model of auditory scene analysis at the core of which is a process that seeks to discover predictable patterns in the ongoing sound sequence. Representations of predictable fragments are created on the fly, and are maintained, strengthened or weakened on the basis of their predictive success, and conflict with other representations. Auditory perceptual organisation emerges spontaneously from the nature of the competition between these representations. We present detailed comparisons between the model simulations and data from an auditory streaming experiment, and show that the model accounts for many important findings, including: the emergence of, and switching between, alternative organisations; the influence of stimulus parameters on perceptual dominance, switching rate and perceptual phase durations; and the build-up of auditory streaming. The principal contribution of the model is to show that a two-stage process of pattern discovery and competition between incompatible patterns can account for both the contents (perceptual organisations) and the dynamics of human perception in auditory streaming.
Sound sources often emit trains of discrete sounds, such as a series of footsteps. Previously, two different principles have been suggested for how the human auditory system binds discrete sounds together into perceptual units. The feature similarity principle is based on linking sounds with similar characteristics over time. The predictability principle is based on linking sounds that follow each other in a predictable manner. The present study compared the effects of these two principles. Participants were presented with tone sequences and instructed to continuously indicate whether they perceived a single coherent sequence or two concurrent streams of sound. We investigated the influence of separate manipulations of similarity and predictability on these perceptual reports. Both grouping principles affected perception of the tone sequences, albeit with different characteristics. In particular, results suggest that whereas predictability is only analyzed for the currently perceived sound organization, feature similarity is also analyzed for alternative groupings of sound. Moreover, changing similarity or predictability within an ongoing sound sequence led to markedly different dynamic effects. Taken together, these results provide evidence for different roles of similarity and predictability in auditory scene analysis, suggesting that forming auditory stream representations and competition between alternatives rely on partly different processes.
Auditory stream segregation involves linking temporally separate acoustic events into one or more coherent sequences. For any non-trivial sequence of sounds, many alternative descriptions can be formed, only one or very few of which emerge in awareness at any time. Evidence from studies showing bi-/multistability in auditory streaming suggest that some, perhaps many of the alternative descriptions are represented in the brain in parallel and that they continuously vie for conscious perception. Here, based on a predictive coding view, we consider the nature of these sound representations and how they compete with each other. Predictive processing helps to maintain perceptual stability by signalling the continuation of previously established patterns as well as the emergence of new sound sources. It also provides a measure of how well each of the competing representations describes the current acoustic scene. This account of auditory stream segregation has been tested on perceptual data obtained in the auditory streaming paradigm.
When people experience an unchanging sensory input for a long period of time, their perception tends to switch stochastically and unavoidably between alternative interpretations of the sensation; a phenomenon known as perceptual bi-stability or multi-stability. The huge variability in the experimental data obtained in such paradigms makes it difficult to distinguish typical patterns of behaviour, or to identify differences between switching patterns. Here we propose a new approach to characterising switching behaviour based upon the extraction of transition matrices from the data, which provide a compact representation that is well-understood mathematically. On the basis of this representation we can characterise patterns of perceptual switching, visualise and simulate typical switching patterns, and calculate the likelihood of observing a particular switching pattern. The proposed method can support comparisons between different observers, experimental conditions and even experiments. We demonstrate the insights offered by this approach using examples from our experiments investigating multi-stability in auditory streaming. However, the methodology is generic and thus widely applicable in studies of multi-stability in any domain.
Visuospatial working memory (vsWM), which is commonly impaired in schizophrenia, involves information processing across the primary visual cortex, association visual cortex, posterior parietal cortex, and dorsolateral prefrontal cortex (DLPFC). Within these regions, vsWM requires inhibition from parvalbumin-expressing basket cells (PVBCs). Here, we analyzed indices of PVBC axon terminals across regions of the vsWM network in schizophrenia.For 20 matched pairs of subjects with schizophrenia and unaffected comparison subjects, tissue sections from the primary visual cortex, association visual cortex, posterior parietal cortex, and DLPFC were immunolabeled for PV, the 65- and 67-kDa isoforms of glutamic acid decarboxylase (GAD65 and GAD67) that synthesize GABA (gamma-aminobutyric acid), and the vesicular GABA transporter. The density of PVBC terminals and of protein levels per terminal was quantified in layer 3 of each cortical region using fluorescence confocal microscopy.In comparison subjects, all measures, except for GAD65 levels, exhibited a caudal-to-rostral decline across the vsWM network. In subjects with schizophrenia, the density of detectable PVBC terminals was significantly lower in all regions except the DLPFC, whereas PVBC terminal levels of PV, GAD67, and GAD65 proteins were lower in all regions. A composite measure of inhibitory strength was lower in subjects with schizophrenia, although the magnitude of the diagnosis effect was greater in the primary visual, association visual, and posterior parietal cortices than in the DLPFC.In schizophrenia, alterations in PVBC terminals across the vsWM network suggest the presence of a shared substrate for cortical dysfunction during vsWM tasks. However, regional differences in the magnitude of the disease effect on an index of PVBC inhibitory strength suggest region-specific alterations in information processing during vsWM tasks.
ABSTRACT We report on the design and the collection of a multi-modal data corpus for cognitive acoustic scene analysis. Sounds are generated by stationary and moving sources (people), that is by omni-directional speakers mounted on people's heads. One or two subjects walk along predetermined systematic and random paths, in synchrony and out of sync. Sound is captured in multiple microphone systems, including a four MEMS microphone directional array, two electret microphones situated in the ears of a stuffed gerbil head, and a Head Acoustics, head-shoulder unit with ICP microphones. Three micro-Doppler units operating at different frequencies were employed to capture gait and the articulatory signatures as well as location of the people in the scene. Three ground vibration sensors were recording the footsteps of the walking people. A 3D MESA camera as well as a web-cam provided 2D and 3D visual data for system calibration and ground truth. Data were collected in three environments ranging from a well controlled environment (anechoic chamber), an indoor environment (large classroom) and the natural environment of an outside courtyard. A software tool has been developed for the browsing and visualization of the data. II.
If, as is widely believed, perception is based upon the responses of neurons that are tuned to stimulus features, then precisely what features are encoded and how do neurons in the system come to be sensitive to those features? Here we show differential responses to ripple stimuli can arise through exposure to formative stimuli in a recurrently connected model of the thalamocortical system which exhibits delays, lateral and recurrent connections, and learning in the form of spike timing dependent plasticity.
We report on the design and the collection of a multi-modal data corpus for cognitive acoustic scene analysis. Sounds are generated by stationary and moving sources (people), that is by omni-directional speakers mounted on people's heads. One or two subjects walk along predetermined systematic and random paths, in synchrony and out of sync. Sound is captured in multiple microphone systems, including a four MEMS microphone directional array, two electret microphones situated in the ears of a stuffed gerbil head, and a Head Acoustics, head-shoulder unit with ICP microphones. Three micro-Doppler units operating at different frequencies were employed to capture gait and the articulatory signatures as well as location of the people in the scene. Three ground vibration sensors were recording the footsteps of the walking people. A 3D MESA camera as well as a web-cam provided 2D and 3D visual data for system calibration and ground truth. Data were collected in three environments ranging from a well controlled environment (anechoic chamber), an indoor environment (large classroom) and the natural environment of an outside courtyard. A software tool has been developed for the browsing and visualization of the data.
The response of an auditory neuron to a tone is often affected by the context in which the tone appears. For example, when measuring the response to a random sequence of tones, frequencies that appear rarely elicit a greater number of spikes than those that appear often. This phenomenon is called stimulus-specific adaptation (SSA). This article presents a neural field model in which SSA arises through selective adaptation to the frequently-occurring inputs. Formulating the network as a field model allows one to obtain an analytical expression for the expected response of a simple two-layer model to tones in a random sequence. The sequences of stimuli used in SSA experiments contain hundreds-and sometimes thousands-of tones, and these experiments routinely measure the response to many such sequences. A conventional neural network model (e.g., integrate-and-fire) would require numerical integration over long time periods to obtain results. Consequently, a field model that offers an immediate, analytical solution for a given input sequence is helpful. Two routes to obtaining this solution are discussed. The first involves the convolution of two closed-form expressions; the second relies on a series of approximations involving Gaussian curves. The purpose of the paper is to describe the model, to develop the approximations that allow an analytical solution, and finally, to comment on the output of the model in light of the SSA results published in the physiology literature. This article is part of a Special Issue entitled "Neural Coding".
Charalambos M. Andreou合作论文数University of Cyprus
Holistic Electronics Research Lab1