OBJECTIVES:A major issue in the rehabilitation of children with cochlear implants (CIs) is unexplained variance in their language skills, where many of them lag behind children with normal hearing (NH). Here, we assess links between generative language skills and the perception of prosodic stress, and with musical and parental activities in children with CIs and NH. Understanding these links is expected to guide future research and toward supporting language development in children with a CI.DESIGN:Twenty-one unilaterally and early-implanted children and 31 children with NH, aged 5 to 13, were classified as musically active or nonactive by a questionnaire recording regularity of musical activities, in particular singing, and reading and other activities shared with parents. Perception of word and sentence stress, performance in word finding, verbal intelligence (Wechsler Intelligence Scale for Children (WISC) vocabulary), and phonological awareness (production of rhymes) were measured in all children. Comparisons between children with a CI and NH were made against a subset of 21 of the children with NH who were matched to children with CIs by age, gender, socioeconomic background, and musical activity. Regression analyses, run separately for children with CIs and NH, assessed how much variance in each language task was shared with each of prosodic perception, the child's own music activity, and activities with parents, including singing and reading. All statistical analyses were conducted both with and without control for age and maternal education.RESULTS:Musically active children with CIs performed similarly to NH controls in all language tasks, while those who were not musically active performed more poorly. Only musically nonactive children with CIs made more phonological and semantic errors in word finding than NH controls, and word finding correlated with other language skills. Regression analysis results for word finding and VIQ were similar for children with CIs and NH. These language skills shared considerable variance with the perception of prosodic stress and musical activities. When age and maternal education were controlled for, strong links remained between perception of prosodic stress and VIQ (shared variance: CI, 32%/NH, 16%) and between musical activities and word finding (shared variance: CI, 53%/NH, 20%). Links were always stronger for children with CIs, for whom better phonological awareness was also linked to improved stress perception and more musical activity, and parental activities altogether shared significantly variance with word finding and VIQ.CONCLUSIONS:For children with CIs and NH, better perception of prosodic stress and musical activities with singing are associated with improved generative language skills. In addition, for children with CIs, parental singing has a stronger positive association to word finding and VIQ than parental reading. These results cannot address causality, but they suggest that good perception of prosodic stress, musical activities involving singing, and parental singing and reading may all be beneficial for word finding and other generative language skills in implanted children.
An interactive method for training speech perception in noise was assessed with adult cochlear implant users. The method employed recordings of connected narratives divided into phrases of 4 to 10 words, presented in babble. After each phrase, the listener identified key words from the phrase from among similar sounding foil words. Nine postlingually deafened adult cochlear implant users carried out 12 hr of training over a 4-week period. Training was carried out at home on tablet computers. The primary outcome measure was sentence recognition in babble. Vowel and consonant identification in speech-shaped noise were also assessed, along with digit span in noise, intended as a measure of some important underlying cognitive abilities. Talkers for speech tests were different from those used in training. To control for procedural learning, the test battery was administered repeatedly prior to training. Performance was assessed immediately after training and again after a further 4 weeks during which no training occurred. Sentence recognition in babble improved significantly after training, with an improvement in speech reception threshold of approximately 2 dB, which was maintained at the 4-week follow-up. There was little evidence of improvement in the other measures. It appears that the method has potential as a clinical intervention. However, the underlying sources of improvement and the extent to which benefits generalize to real-world situations remain to be determined.
The perception of speech in noise is challenging for children with cochlear implants (CIs). Singing and musical instrument playing have been associated with improved auditory skills in normal-hearing (NH) children. Therefore, we assessed how children with CIs who sing informally develop in the perception of speech in noise compared to those who do not. We also sought evidence of links of speech perception in noise with MMN and P3a brain responses to musical sounds and studied effects of age and changes over a 14–17 month time period in the speech-in-noise performance of children with CIs. Compared to the NH group, the entire CI group was less tolerant of noise in speech perception, but both groups improved similarly. The CI singing group showed better speech-in-noise perception than the CI non-singing group. The perception of speech in noise in children with CIs was associated with the amplitude of MMN to a change of sound from piano to cymbal, and in the CI singing group only, with earlier P3a for changes in timbre. While our results cannot address causality, they suggest that singing and musical instrument playing may have a potential to enhance the perception of speech in noise in children with CIs.
Sentence recognition in 20-talker babble was measured in eight Nucleus cochlear implant (CI) users with contralateral residual acoustic hearing. Speech reception thresholds (SRTs) were measured both in standard configurations, with some frequency regions presented both acoustically and electrically, and in configurations with no spectral overlap. In both cases a continuous interleaved sampling strategy was used. Mean SRTs were around 3 dB better with bimodal presentation than with CI alone in overlap configurations. A spherical head model was used to simulate azimuthal separation of speech and noise and provided no evidence of a contribution of spatial cues to bimodal benefit. There was no effect on bimodal performance of whether spectral overlap was present or was eliminated by switching off electrodes assigned to frequencies below the upper limit of acoustic hearing. In a subsequent experiment the CI was acutely re-mapped so that all available electrodes were used to cover frequencies not presented acoustically. This gave increased spectral resolution via the CI as assessed by formant frequency discrimination, but no improvement in bimodal performance compared to the configuration with overlap.
Fastl, H. (1987). “Ein Störgeräusch für die Sprachaudiometrie,” Audiol. Akustik, 26, 2-13. Healy, A.F., Havas, D.A., and Parker, J.T. (2000). “Comparing serial position effects in semantic and episodic memory using reconstruction of order tasks,” J. Memo. Lang., 42, 147-167. Jones, T., and Oberauer, K. (2013). “Serial-position effects for items and relations in short-term memory,” Memory, 21, 347-365. Larsby, B., Hällgren, M., Lyxell, B., and Arlinger, S. (2005). “Cognitive performance and perceived effort in speech processing tasks: effects of different noise backgrounds in normal-hearing and hearing-impaired subjects,” Int. J. Audiol., 44, 131-143. Leiss, E. (2002) “Die Wortart Verb”, in Lexikologie. Ein internationales Handbuch zur Natur und Struktur von Wörtern und Wortschätzen. XVII. Die Architektur des Wortschatzes I: Die Wortarten. 2. Halbband. Edited by D.A. Cruse, (de Gruyter, Berlin, New York), pp. 605-616. Ljung, R., Israelsson, K., and Hygge, S. (2013). “Speech intelligibility and recall of spoken material heard at different signal-to-noise ratios and the role played by working memory capacity.” Appl. Cognitive Psych., 27, 198-203. Miller, G.A. (1956). “The magical number seven, plus or minus two. Some limits on our capacity for processing information,” Psychol. Rev., 63, 82-97. Oberauer, K. (2003). “Understanding serial position curves in short-term recognition and recall,” J. Mem. Lang., 49, 469-483. Uslar, V., Ruigendijk, E., Hamann, C., Brand, T., and Kollmeier, B. (2011). “How does linguistic complexity influence intelligibility in a German audiometric sentence intelligibility test?” Int. J. Audiol., 50, 621-631. Vigliocco, G., Vinson, D.P., Druks, J., Barber, H., and Cappa, S.F. (2011). “Nouns and verbs in the brain: A review of behavioural, electrophysiological, neuropsychological and imaging studies,” Neurosci. Biobehav. R., 35, 407-426. Wagener, K., Brand, T., and Kollmeier, B. (1999). “Entwicklung und Evaluation eines Satztests für die deutsche Sprache I: Design des Oldenburger Satztests. Development and evaluation of a German sentence test I: Design of the Oldenburg sentence test,” Z. Audiol., 38, 5-15.
Objective: To study prosodic perception in early-implanted children in relation to auditory discrimination, auditory working memory, and exposure to music. Design: Word and sentence stress perception, discrimination of fundamental frequency (F0), intensity and duration, and forward digit span were measured twice over approximately 16 months. Musical activities were assessed by questionnaire. Study sample: Twenty-one early-implanted and age-matched normal-hearing (NH) children (4-13 years). Results: Children with cochlear implants (CIs) exposed to music performed better than others in stress perception and F0 discrimination. Only this subgroup of implanted children improved with age in word stress perception, intensity discrimination, and improved over time in digit span. Prosodic perception, F0 discrimination and forward digit span in implanted children exposed to music was equivalent to the NH group, but other implanted children performed more poorly. For children with CIs, word stress perception was linked to digit span and intensity discrimination: sentence stress perception was additionally linked to F0 discrimination. Conclusions: Prosodic perception in children with CIs is linked to auditory working memory and aspects of auditory discrimination. Engagement in music was linked to better performance across a range of measures, suggesting that music is a valuable tool in the rehabilitation of implanted children.
Much recent interest surrounds listeners' abilities to adapt to various transformations that distort speech. An extreme example is spectral rotation, in which the spectrum of low-pass filtered speech is inverted around a center frequency (2 kHz here). Spectral shape and its dynamics are completely altered, rendering speech virtually unintelligible initially. However, intonation, rhythm, and contrasts in periodicity and aperiodicity are largely unaffected. Four normal hearing adults underwent 6 h of training with spectrally-rotated speech using Continuous Discourse Tracking. They and an untrained control group completed pre- and post-training speech perception tests, for which talkers differed from the training talker. Significantly improved recognition of spectrally-rotated sentences was observed for trained, but not untrained, participants. However, there were no significant improvements in the identification of medial vowels in /bVd/ syllables or intervocalic consonants. Additional tests were performed with speech materials manipulated so as to isolate the contribution of various speech features. These showed that preserving intonational contrasts did not contribute to the comprehension of spectrally-rotated speech after training, and suggested that improvements involved adaptation to altered spectral shape and dynamics, rather than just learning to focus on speech features relatively unaffected by the transformation.
Background: The accurate perception of prosody assists a listener in deriving meaning from natural speech. Few studies have addressed the ability of cochlear implant (CI) listeners to perceive the brief duration prosodic cues involved in contrastive vowel length, word stress, and compound word and phrase identification. Purpose: To compare performance in the perception of brief duration prosodic contrasts by CI participants and a control group of normal hearing participants. This study investigated the ability to perceive these cues in quiet and noise conditions, and to identify auditory perceptual factors that might predict prosodic perception in the CI group. Prosodic perception was studied both in noise and quiet because noise is a pervasive feature of everyday environments. Research Design: A quasi-experimental correlation design was employed. Study Sample: Twenty-one CI recipients participated along with a control group of 10 normal hearing participants. All CI participants were unilaterally implanted adults who had considerable experience with oral language prior to implantation. Data Collection and Analysis: Speech identification testing measured the participants' ability to identify word stress, vowel length, and compound words or phrases all of which were presented with minimal-pair response choices. Tests were performed in quiet and in speech-spectrum shaped noise at a 10 dB signal-to-noise ratio. Also, discrimination thresholds for four acoustic properties of a synthetic vowel were measured as possible predictors of prosodic perception. Testing was carried out during one session, and participants used their clinically assigned speech processors. Results: The CI group could not identify brief prosodic cues as well as the control group, and their performance decreased significantly in the noise condition. Regression analysis showed that the discrimination of intensity predicted performance on the prosodic tasks. The performance decline measured with the older participants meant that age also emerged as a predictor. Conclusions: This study provides a portrayal of CI recipients' ability to perceive brief prosodic cues. This is of interest in the preparation of rehabilitation materials used in training and in developing realistic expectations for potential CI candidates.
Due to the inherent device limitations of cochlear implants (CI) and of auditory perception via an electrical-neural interface, the ability of CI listeners to perceive prosody is often reported as being worse than that of normal-hearing listeners. We tested the perceptual ability of postlingually-deafened adult CI listeners with stimuli where prosodic features signalled distinctive semantic contrasts. These contrasts were tested with a Danish (n=18) and a Swedish (n=21) cohort in quiet and in noise. We also tested other speech perceptual abilities that could be linked to prosody perception. These included word recognition, sentence perception in noise, and vowel identification. Results of this study show that speech-in-noise ability by CI listeners is related to abilities that underlie vowel identification, while word recognition is related to the identification of compound words and phrases. Comparison of the mean identification rates of the prosodic tasks showed that there was a disparity in the performance of Danish and Swedish CI listeners over tasks that are similar in both languages.
Introduction Speech perception requires a receiver to make decisions both about trends and also about the language patterns of a speaker. Hearing people use the speaker’s speech signal as the direct sensory evidence. In the special case of visual speech perception, otherwise known as speechreading, this sensory evidence is derived from the visible articulation movements of speech. Unfortunately, many important articulation movements are invisible. On the segmental level, some articulation positions are difficult or impossible to distinguish. Even for those sounds that are visible, few have unique visual cues. At the suprasegmental level, the basic speech elements intonation and lexical tone which are mainly conveyed by the vibrating frequencies of the larynx, are totally missing in visible facial gestures. Therefore, a successful speechreader has to compensate for the resulting shortage of sensory data by taking maximal advantage of prior knowledge of language structure and of the linguistic and situational context. This process is demanding, however, and even the most competent speechreaders have to accept a substantial probability of error.
OBJECTIVES:This study investigated whether low frequency information from a hearing aid improved the perception of stress and intonation by English-speaking children with cochlear implants. As pitch information is limited for cochlear implant users, this study also investigated if users rely more on the cues of duration and amplitude to perceive stress and intonation.METHODS:Nine children with bimodal stimulation (cochlear implant and hearing aid) participated in two experiments. The first measured the just audible change in F0 (pitch) and amplitude for a speech-like word 'baba'. The second experiment examined the children's ability to identify focus in natural and manipulated sentences.RESULTS:Overall, group results did not show a bimodal advantage in perceiving stress and intonation. However, the children were significantly better at perceiving focus in sentences with natural speech compared with manipulated speech in both the CI and bimodal conditions. The results suggest that in the absence of pitch cues, amplitude and duration cues are used to perceive stress and intonation. However, the majority of children only perceived amplitude changes greater than the changes typically found in speech, implying duration cues were the most valuable.DISCUSSION:Taken together the findings suggest that for children with cochlear implants, cues to F0 may not be essential for prosody perception and in the absence of cues to F0 and amplitude, duration may offer an alternative cue.CONCLUSION:Although a bimodal advantage was not demonstrated for all participants, it is recommended that if clinically appropriate, a contralateral hearing aid is fitted and trialled to exploit any residual hearing.
Acoustic simulations were used to study the contributions of spatial hearing that may arise from combining a cochlear implant with either a second implant or contralateral residual low-frequency acoustic hearing. Speech reception thresholds (SRTs) were measured in twenty-talker babble. Spatial separation of speech and noise was simulated using a spherical head model. While low-frequency acoustic information contralateral to the implant simulation produced substantially better SRTs there was no effect of spatial cues on SRT, even when interaural differences were artificially enhanced. Simulated bilateral implants showed a significant head shadow effect, but no binaural unmasking based on interaural time differences, and weak, inconsistent overall spatial release from masking. There was also a small but significant non-spatial summation effect. It appears that typical cochlear implant speech processing strategies may substantially reduce the utility of spatial cues, even in the absence of degraded neural processing arising from auditory deprivation.
Two experimental groups were trained for 2 h with live or recorded speech that was noise-vocoded and spectrally shifted and was from the same text and talker. These two groups showed equivalent improvements in performance for vocoded and shifted sentences, and the group trained with recorded speech showed consistently greater improvements than untrained controls. Another group trained with unshifted noise-vocoded speech improved no more than untrained controls. Computer-based training thus appears at least as effective as labor-intensive live-voice training for improving the perception of spectrally shifted noise-vocoded speech, and by implication, for training of users of cochlear implants.
Impaired auditory processing represents an important test of what we understand about speech perception and its auditory basis. The World Health Organization estimates that hearing impairment of a degree that severely affects speech perception occurs in around 2% of the population. The goal of improving speech perception for the hearing impaired also represents a challenge to speech technology-to produce robust speech analysis techniques that can maximize the transmission of useful speech information while minimizing the transmission of noise. The development of cochlear implants and other aids to match very limited auditory abilities has often been informed by an understanding of more central aspects of auditory pattern processing in speech perception. The perceptual role of temporal structure in speech has until recently been rather neglected in comparison to spectral structure. The study of speech perception in hearing-impaired listeners can be seen to be leading to important advances in our understanding of the roles of spectral and temporal processing.
Speech comprehension is a complex human skill, the performance of which requires the perceiver to combine information from several sources - e.g. voice, face, gesture, linguistic context - to achieve an intelligible and interpretable percept. We describe a functional imaging investigation of how auditory, visual and linguistic information interact to facilitate comprehension. Our specific aims were to investigate the neural responses to these different information sources, alone and in interaction, and further to use behavioural speech comprehension scores to address sites of intelligibility-related activation in multifactorial speech comprehension. In fMRI, participants passively watched videos of spoken sentences, in which we varied Auditory Clarity (with noise-vocoding), Visual Clarity (with Gaussian blurring) and Linguistic Predictability. Main effects of enhanced signal with increased auditory and visual clarity were observed in overlapping regions of posterior STS. Two-way interactions of the factors (auditory × visual, auditory × predictability) in the neural data were observed outside temporal cortex, where positive signal change in response to clearer facial information and greater semantic predictability was greatest at intermediate levels of auditory clarity. Overall changes in stimulus intelligibility by condition (as determined using an independent behavioural experiment) were reflected in the neural data by increased activation predominantly in bilateral dorsolateral temporal cortex, as well as inferior frontal cortex and left fusiform gyrus. Specific investigation of intelligibility changes at intermediate auditory clarity revealed a set of regions, including posterior STS and fusiform gyrus, showing enhanced responses to both visual and linguistic information. Finally, an individual differences analysis showed that greater comprehension performance in the scanning participants (measured in a post-scan behavioural test) were associated with increased activation in left inferior frontal gyrus and left posterior STS. The current multimodal speech comprehension paradigm demonstrates recruitment of a wide comprehension network in the brain, in which posterior STS and fusiform gyrus form sites for convergence of auditory, visual and linguistic information, while left-dominant sites in temporal and frontal cortex support successful comprehension.
Objective: To assess the reliability of across-ear, acoustic-electric pitch/timbre comparisons for determining effective characteristic frequencies of cochlear implant electrodes. Study sample: Nine CI users with contralateral residual acoustic hearing. Design: Absolute acoustic thresholds in the unimplanted ear were measured and frequency selectivity was assessed via psychophysical tuning curves. An adjustment method was used to match the percepts elicited by pulse trains on individual electrodes with various acoustic signals (pure tones, narrow-band noises, and bandpass filtered pulse trains). The starting frequency of the acoustic signal was roved and matches were obtained at different loudness levels. Results: Acoustic frequency selectivity varied widely. Two subjects showed clear evidence of frequency selectivity extending above 500 Hz. Only these subjects produced consistent pitch matches over repeated measurements. For other subjects, the acoustic frequency eventually selected tended to correlate with the initially presented frequency. There was limited evidence of level effects and these were inconsistent across subjects and electrodes. Conclusions: Across-modality pitch/timbre matching appears unlikely to provide a generally applicable method for determining the effective characteristic frequencies of cochlear implant electrodes. Frequency selectivity above 500 Hz may be necessary for consistent pitch/timbre matches.
Objective: We studied the neurocognitive mechanisms of musical instrument sound perception in children with Cochlear Implants (CIs) and in children with normal hearing (NH).Methods: ERPs were recorded in a new multi-feature change-detection paradigm. Three magnitudes of change in fundamental frequency, musical instrument, duration, intensity increments and decrements, and presence of a temporal gap were presented amongst repeating 295 Hz piano tones. Independent Component Analysis was utilized to remove artifacts caused by the Cochlear Implants.Results: The ERPs were similar in the two groups across all perceptual dimensions except for intensity increment deviants. CI children had smaller and earlier P1 responses compared to controls, and their MMN responses showed less accurate neural detection of changes of musical instrument, sound duration, and temporal structure. P3a responses suggested that poor neural detection of musical instruments affected their involuntary attention shift.Conclusions: The similarities of neurocognitive processing are surprising in the light of the limited auditory input provided by the CI, suggesting that many types of changes are adequately processed by the CI children.Significance: Our results indicate that CI children's auditory cortical functioning may be enhanced, and difficulties in auditory perception and in attention switching towards sound events alleviated, by multisensory musical activities. (C) 2012 International Federation of Clinical Neurophysiology. Published by Elsevier Ireland Ltd. All rights reserved.
Objectives: A major focus of recent attempts to enhance cochlear implant (CI) systems has been to increase the rate at which pulses are delivered to the electrode array. One basis for these attempts has been the expectation that faster stimulation rates would lead to an enhanced representation of temporal modulation information. However, there is recent physiological and behavioral evidence to suggest that the reverse may be the case. Here, the effects of stimulation rate on the perception of amplitude modulation were assessed using both modulation detection and modulation frequency discrimination tasks for a range of pulse rates extending considerably higher than the highest rate tested in previous studies and for different speech-relevant modulation frequencies. Design: Detection of sinusoidal amplitude modulation was assessed in five CI users using monopolar pulse trains presented to a single electrode at rates of 482, 723, 1447, 2894, and 5787 pulses per second (pps). Adaptive procedures were used to find the minimal detectable modulation depth at modulation frequencies of 10 and 100 Hz and at carrier levels of 25%, 50%, and 75% of the electrode’s dynamic range. Discrimination of modulation frequency was examined for the same range of pulse rates for the highest carrier level. Similar adaptive procedures determined the minimum increase in modulation frequency that could be detected relative to reference modulation frequencies of 10, 100, and 200 Hz. In both tasks, level roving was implemented to minimize possible loudness cues. Results: Consistent with previous evidence, modulation detection thresholds were better for higher carrier levels and lower modulation frequencies. When modulation depth at threshold was expressed in terms of the ratio of the depth of the modulation and the carrier level in dB (i.e., 20 log m), performance was significantly better at lower pulse rates. However, when modulation depth was expressed relative to dynamic range, the effect of pulse rate was no longer significant, reflecting the fact that dynamic range increases with pulse rate. Modulation frequency discrimination clearly worsened with increasing modulation frequency, but there was no significant effect of pulse rate. Conclusions: In contrast to some recent evidence, no clearly harmful effect of higher pulse rates on modulation perception was found. However, even with very fast stimulation rates, tested over a wide range of modulation frequencies and with two different tasks, there is no evidence of benefit from faster stimulation rates in the perception of amplitude modulation.
In the present study, we investigated the processing of word stress related acoustic features in a word context. In a passive oddball multi-feature MMN experiment, we presented a disyllabic pseudo-word with two acoustically similar syllables as standard stimulus, and five contrasting deviants that differed from the standard in that they were either stressed on the first syllable or contained a vowel change. Stress was realized by an increase of f0, intensity, vowel duration or consonant duration. The vowel change was used to investigate if phonemic and prosodic changes elicit different MMN components. As a control condition, we presented non-speech counterparts of the speech stimuli.Results showed all but one feature (non-speech intensity deviant) eliciting the MMN component, which was larger for speech compared to non-speech stimuli. Two other components showed stimulus related effects: the N350 and the LDN (Late Discriminative Negativity). The N350 appeared to the vowel duration and consonant duration deviants, specifically to features related to the temporal characteristics of stimuli, while the LDN was present for all features, and it was larger for speech than for non-speech stimuli. We also found that the f0 and consonant duration features elicited a larger MMN than other features.These results suggest that stress as a phonological feature is processed based on long-term representations, and listeners show a specific sensitivity to segmental and suprasegmental cues signaling the prosodic boundaries of words. These findings support a two-stage model in the perception of stress and phoneme related acoustical information.