Auditory-evoked potentials are classically defined as the summations of synchronous firing along the auditory neuraxis. Converging evidence supports a model whereby timing jitter in neural coding compromises listening and causes variable scalp-recorded potentials. Yet the intrinsic noise of human scalp recordings precludes a full understanding of the biological origins of individual differences in listening skills. To delineate the mechanisms contributing to these phenomena, in vivo extracellular activity was recorded from inferior colliculus in guinea pigs to speech in quiet and noise. Here we show that trial-by-trial timing jitter is a mechanism contributing to auditory response variability. Identical variability patterns were observed in scalp recordings in human children, implicating jittered timing as a factor underlying reduced coding of dynamic speech features and speech in noise. Moreover, intertrial variability in human listeners is tied to language development. Together, these findings suggest that variable timing in inferior colliculus blurs the neural coding of speech in noise, and propose a consequence of this timing jitter for human behavior. These results hint both at the mechanisms underlying speech processing in general, and at what may go awry in individuals with listening difficulties.
The NIH Toolbox project has assembled measurement tools to assess a wide range of human perception and ability across the lifespan. As part of this initiative, a small but comprehensive battery of auditory tests has been assembled. The main tool of this battery, pure-tone thresholds, measures the ability of people to hear at specific frequencies. Pure-tone thresholds have long been considered the "gold standard" of auditory testing, and are normally obtained in a clinical setting by highly trained audiologists. For the purposes of the Toolbox project, an automated procedure (NIH Toolbox Threshold Hearing Test) was developed that allows nonspecialists to administer the test reliably. Three supplemental auditory tests are also included in the Toolbox auditory test battery: assessment of middle-ear function (tympanometry), speech perception in noise (the NIH Toolbox Words-in-Noise Test), and self-assessment of hearing impairment (the NIH Toolbox Hearing Handicap Inventory Ages 18-64 and the NIH Toolbox Hearing Handicap Inventory Ages 64+). Tympanometry can help differentiate conductive from sensorineural pathology. The NIH Toolbox Words-in-Noise Test measures a listener's ability to perceive words in noisy situations. This ability is not necessarily predicted by a person's pure-tone thresholds; some people with normal hearing have difficulty extracting meaning from speech sounds heard in a noisy context. The NIH Toolbox Hearing Handicap Inventory focuses on how a person's perceived hearing status affects daily life. The test was constructed to include emotional and social/situational subscales, with specific questions about how hearing impairment may affect one's emotional state or limit participation in specific activities. The 4 auditory tests included in the Toolbox auditory test battery cover a range of auditory abilities and provide a snapshot of a participant's auditory capacity.
The human auditory brainstem is known to be exquisitely sensitive to fine-grained spectro-temporal differences between speech sound contrasts, and the ability of the brainstem to discriminate between these contrasts is important for speech perception. Recent work has described a novel method for translating brainstem timing differences in response to speech contrasts into frequency-specific phase differentials. Results from this method have shown that the human brainstem response is surprisingly sensitive to phase differences inherent to the stimuli across a wide extent of the spectrum. Here we use an animal model of the auditory brainstem to examine whether the stimulus-specific phase signatures measured in human brainstem responses represent an epiphenomenon associated with far-field (i.e., scalp-recorded) measurement of neural activity, or alternatively whether these specific activity patterns are also evident in auditory nuclei that contribute to the scalp-recorded response, thereby representing a more fundamental temporal processing phenomenon. Responses in anaesthetized guinea pigs to three minimally-contrasting consonant-vowel stimuli were collected simultaneously from the cortical surface vertex and directly from central nucleus of the inferior colliculus (ICc), measuring volume conducted neural activity and multiunit, near-field activity, respectively. Guinea pig surface responses were similar to human scalp-recorded responses to identical stimuli in gross morphology as well as phase characteristics. Moreover, surface-recorded potentials shared many phase characteristics with near-field ICc activity. Response phase differences were prominent during formant transition periods, reflecting spectro-temporal differences between syllables, and showed more subtle differences during the identical steady state periods. ICc encoded stimulus distinctions over a broader frequency range, with differences apparent in the highest frequency ranges analyzed, up to 3000 Hz. Based on the similarity of phase encoding across sites, and the consistency and sensitivity of response phase measured within ICc, results suggest that a general property of the auditory system is a high degree of sensitivity to fine-grained phase information inherent to complex acoustical stimuli. Furthermore, results suggest that temporal encoding in ICc contributes to temporal features measured in speech-evoked scalp-recorded responses.
The way in which normal variations in human neuroanatomy relate to brain function remains largely uninvestigated. This study addresses the question by relating anatomical measurements of Heschl's gyrus (HG), the structure containing human primary auditory cortex, to how this region processes temporal and spectral acoustic information. In this study, subjects' right and left HG were identified and manually indicated on anatomical magnetic resonance imaging scans. Volumes of gray matter, white matter, and total gyrus were recorded, and asymmetry indices were calculated. Additionally, cortical auditory activity in response to noise stimuli varying orthogonally in temporal and spectral dimensions was assessed and related to the volumetric measurements. A high degree of anatomical variability was seen, consistent with other reports in the literature. The auditory cortical responses showed the expected leftward lateralization to varying rates of stimulus change and rightward lateralization of increasing spectral information. An explicit link between auditory structure and function is then established, in which anatomical variability of auditory cortex is shown to relate to individual differences in the way that cortex processes acoustic information. Specifically, larger volumes of left HG were associated with larger extents of rate-related cortex on the left, and larger volumes of right HG related to larger extents of spectral-related cortex on the right. This finding is discussed in relation to known microanatomical asymmetries of HG, including increased myelination of its fibers, and implications for language learning are considered.
STUDYING SIMILARITIES AND DIFFERENCES BETWEEN speech and song provides an opportunity to examine music's role in human culture. Forty participants divided into groups of musicians and nonmusicians spoke and sang lyrics to two familiar songs. The spectral structures of speech and song were analyzed using a statistical analysis of frequency ratios. Results showed that speech and song have similar spectral structures, with song having more energy present at frequency ratios corresponding to those ratios associated with the 12-tone scale. This difference may be attributed to greater fundamental frequency variability in speech, and was not affected by musical experience. Higher levels of musical experience were associated with decreased energy at frequency ratios not corresponding to the 12-tone scale in both speech and song. Thus, musicians may invoke multisensory (auditory/vocal-motor) mechanisms to fine-tune their vocal production to more closely align their speaking and singing voices according to their vast music listening experience.
Research on the contributions of the human nervous system to language processing and learning has generally been focused on the association regions of the brain without considering the possible contribution of primary and adjacent sensory areas. We report a study examining the relationship between the anatomy of Heschl's Gyrus (HG), which includes predominately primary auditory areas and is often found to be associated with nonlinguistic pitch processing and language learning. Unlike English, most languages of the world use pitch patterns to signal word meaning. In the present study, native English-speaking adult subjects learned to incorporate foreign pitch patterns in word identification. Subjects who were less successful in learning showed a smaller HG volume on the left (especially gray matter volume), but not on the right, relative to learners who were successful. These results suggest that HG, typically shown to be associated with the processing of acoustic cues in nonspeech processing, is also involved in speech learning. These results also suggest that primary auditory regions may be important for encoding basic acoustic cues during the course of spoken language learning.
The physiological mechanisms that contribute to abnormal encoding of speech in children with learning problems are yet to be well understood. Furthermore, speech perception problems appear to be particularly exacerbated by background noise in this population. This study compared speech-evoked cortical responses recorded in a noisy background to those recorded in quiet in normal children (NL) and children with learning problems (LP). Timing differences between responses recorded in quiet and in background noise were assessed by cross-correlating the responses with each other. Overall response magnitude was measured with root-mean-square (RMS) amplitude. Cross-correlation scores indicated that 23% of LP children exhibited cortical neural timing abnormalities such that their neurophysiological representation of speech sounds became distorted in the presence of background noise. The latency of the N2 response in noise was isolated as being the root of this distortion. RMS amplitudes in these children did not differ from NL children, indicating that this result was not due to a difference in response magnitude. LP children who participated in a commercial auditory training program and exhibited improved cortical timing also showed improvements in phonological perception. Consequently, auditory pathway timing deficits can be objectively observed in LP children, and auditory training can diminish these deficits.
In this study, spectral timbre’s effect on pitch perception is examined in varying contexts. In two experiments, subjects detected pitch deviations of tones differing in brightness in an isolated context in which they compared two tones, in a tone-series context in which they judged whether the last tone of a simple sequence was in or out of tune, and in a melodic context in which they determined whether the last note of familiar melodies was in or out of tune. Timbre influenced pitch judgments in all the conditions, but increasing tonal context allowed the subjects to extract pitch information more accurately. This appears to be due to two factors: (1) The presence of extra tones creates a stronger reference point from which to judge pitch, and (2) the melodies’ tonal structure gives more cues that facilitate pitch extraction, even in the face of conflicting spectral information.