OBJECTIVE:Shifting the mean fundamental frequency (F0) of target speech down in frequency may be a way to provide the benefits of electric-acoustic stimulation (EAS) to cochlear implant (CI) users whose limited residual hearing precludes a benefit typically, even with amplification. However, previous study showed a decline in the amount of benefit at the greatest downward frequency shifts, and the authors hypothesized that this might be related to F0 variation. Thus, in the present study, the authors sought to determine the relationship between mean F0, F0 variation, and the benefits of combining electric stimulation from a CI with low-frequency residual acoustic hearing.DESIGN:The authors measured speech intelligibility in normal-hearing listeners using an EAS simulation consisting of a sine vocoder combined either with speech low-pass filtered at 500 Hz, or with a pure tone representing target F0. The authors used extracted target voice pitch information to modulate the tone, and manipulated both the frequency of the carrier (mean F0), as well as the standard deviation of the voice pitch information (F0 variation).RESULTS:A decline in EAS benefit was observed at the lowest mean F0 tested, but this decline disappeared when F0 variation was reduced to be proportional to the amount of the shift in frequency (i.e., when F0 was shifted logarithmically instead of linearly).CONCLUSION:Lowering mean F0 by shifting the frequency of a pure tone carrying target voice pitch information can provide as much EAS benefit as an unshifted tone, at least in the current simulation of EAS. These results may have implications for CI users with extremely limited residual acoustic hearing.
Cochlear implant users report difficulty understanding speech in both noisy and reverberant environments. Electric-acoustic stimulation (EAS) is known to improve speech intelligibility in noise. However, little is known about the potential benefits of EAS in reverberation, or about how such benefits relate to those observed in noise. The present study used EAS simulations to examine these questions. Sentences were convolved with impulse responses from a model of a room whose estimated reverberation times were varied from 0 to 1 sec. These reverberated stimuli were then vocoded to simulate electric stimulation, or presented as a combination of vocoder plus low-pass filtered speech to simulate EAS. Monaural sentence recognition scores were measured in two conditions: reverberated speech and speech in a reverberated noise. The long-term spectrum and amplitude modulations of the noise were equated to the reverberant energy, allowing a comparison of the effects of the interferer (speech vs noise). Results indicate that, at least in simulation, (1) EAS provides significant benefit in reverberation; (2) the benefits of EAS in reverberation may be underestimated by those in a comparable noise; and (3) the EAS benefit in reverberation likely arises from partially preserved cues in this background accessible via the low-frequency acoustic component.
Three experiments were designed to provide psychophysical evidence for the existence of envelope information in the temporal fine structure (TFS) of stimuli that were originally amplitude modulated (AM). The original stimuli typically consisted of the sum of a sinusoidally AM tone and two unmodulated tones so that the envelope and TFS could be determined a priori. Experiment 1 showed that normal-hearing listeners not only perceive AM when presented with the Hilbert fine structure alone but AM detection thresholds are lower than those observed when presenting the original stimuli. Based on our analysis, envelope recovery resulted from the failure of the decomposition process to remove the spectral components related to the original envelope from the TFS and the introduction of spectral components related to the original envelope, suggesting that frequency- to amplitude-modulation conversion is not necessary to recover envelope information from TFS. Experiment 2 suggested that these spectral components interact in such a way that envelope fluctuations are minimized in the broadband TFS. Experiment 3 demonstrated that the modulation depth at the original carrier frequency is only slightly reduced compared to the depth of the original modulator. It also indicated that envelope recovery is not specific to the Hilbert decomposition.
In reverberant environments, reflections occurring within 30–50 ms of the direct speech assist intelligibility by perceptually fusing with the source, effectively increasing its level. Fusion occurs at similar delays in monaurally deaf and normal-hearing listeners, suggesting an independence from binaural processes [Litovsky et al., J. Acoust. Soc. Am. 106, 1633–1654 (1999)]. Its effects can thus be examined in unilateral cochlear implant (CI) users. Simulated CI listening shows little benefit from early reflections, and a detriment when delays exceed 20 ms [Whitmal and Poissant, Conference on Implantable Auditory Prostheses (2009)]. Data from our laboratory show that CI users with residual acoustic hearing benefit from electric-acoustic stimulation (EAS) in reverberation. The role of early reflections in this EAS benefit remains unclear. The present study examined the effect of a single unattenuated reflection at ten delay times on intelligibility in simulated EAS. Target stimuli consisted of sentences combined with unattenuated copies delayed by 0–66 ms. Four-talker babble was added to increase difficulty. A four-channel vocoder simulated electric stimulation; low-pass filtered speech represented residual acoustic hearing. Monaural intelligibility scores from normal-hearing listeners under anechoic and reflected conditions with electric and electric-acoustic processing suggest that EAS may facilitate a benefit from early reflections. [Work supported by the NIDCD.]
The amount of benefit in speech intelligibility in noise can vary widely for cochlear-implant (CI) users when low-frequency acoustic stimulation is combined with electric stimulation. This variability may, in part, be explained by the listener’s ability to follow fundamental frequency (f0) changes with electric stimulation only: those who benefit the least may be most adept at processing f0 in the electric region. The present study measured the ability of CI users to discriminate stochastic FM patterns when presented acoustically or electrically only. The stimuli were 200-Hz carriers randomly modulated in frequency by 5-Hz lowpass noise. A cued 2IFC procedure was used. In one of the two observation intervals, the cue pattern was repeated; in the other, the modulator was inverted so that the pattern of instantaneous frequency fluctuation mirrored that of the cue pattern at the center frequency. Discrimination thresholds were measured in terms of the FM depth needed to discriminate the deviant from the cue pattern. Additional conditions disrupted FM-to-AM conversion by adding extraneous AM to either the carrier or a noise masker. The discrimination results will be compared to the benefit in speech intelligibility obtained by CI users when presented with low-frequency acoustic cues. [Work supported by NIDCD.]
Previous experiments have shown significant improvement in speech intelligibility under both simulated [Brown, C. A., and Bacon, S. P. (2009a). J. Acoust. Soc. Am. 125, 1658-1665; Brown, C. A., and Bacon, S. P. (2010). Hear. Res. 266, 52-59] and real [Brown, C. A., and Bacon, S. P. (2009b). Ear Hear. 30, 489-493] electric-acoustic stimulation when the target speech in the low-frequency region was replaced with a tone modulated in frequency to track the changes in the target talker's fundamental frequency (F0), and in amplitude with the amplitude envelope of the target speech. The present study examined the effects in simulation of applying these cues to a tone lower in frequency than the mean F0 of the target talker. Results showed that shifting the frequency of the tonal carrier downward by as much as 75 Hz had no negative impact on the benefit to intelligibility due to the tone, and that even a shift of 100 Hz resulted in a significant benefit over simulated electric-only stimulation when the sensation level of the tone was comparable to that of the tones shifted by lesser amounts.
We have previously demonstrated that much or all of the benefits of electric-acoustic stimulation (EAS) can be achieved when the low-frequency speech is replaced with a tone that is modulated in both frequency and amplitude with F0 and amplitude envelope cues derived from the target speech. One advantage of this approach is that the frequency of the carrier tone can be lowered with little decline in benefit. Lowering mean F0 has the potential to provide EAS benefit to CI users who have very limited residual hearing. One drawback to this approach is that it relies on the efficacy of the pitch extraction algorithm. This is problematic because pitch extractors have trouble in background noise, an environment in which F0 is particularly useful. Here, an alternative way of lowering mean F0 that is unaffected by the presence of noise is examined under simulated EAS conditions. Speech intelligibility was measured using an algorithm based on resampling and compared to performance with the pitch-based method we have used previously, at frequency shifts of 0, 0.5, and 1 octave. At the 0.5-octave shift, the resampling-based approach provided more benefit than the tone. However, at the 1-octave shift, resampling was less beneficial. [Work supported by NIDCD.]
Speech reception in noise is an especially difficult problem for listeners with hearing impairment as well as for users of cochlear implants (CIs). One likely cause of this is an inability to ‘glimpse’ a target talker in a fluctuating background, which has been linked to deficits in temporal fine-structure processing. A fine-structure cue that has the potential to be beneficial for speech reception in noise is fundamental frequency (F0). A challenging problem, however, is delivering the cue to these individuals. The benefits to speech intelligibility of F0 for both listeners with hearing impairment and users of CIs are reviewed, as well as various methods of delivering F0 to these listeners.
Objective: When either real or simulated electric stimulation from a cochlear implant (Cl) is combined with low-frequency acoustic stimulation (electric-acoustic stimulation [EAS]), speech intelligibility in noise can improve dramatically. We recently showed that a similar benefit to intelligibility can be observed in simulation when the low-frequency acoustic stimulation (low-pass target speech) is replaced with a tone that is modulated both in frequency with the fundamental frequency (F0) of the target talker and in amplitude with the amplitude envelope of the low-pass target speech (Brown & Bacon 2009). The goal of the current experiment was to examine the benefit of the modulated tone to intelligibility in Cl patients.Design: Eight Cl users who had some residual acoustic hearing either in the implanted ear, the unimplanted ear, or both ears participated in this study. Target speech was combined with either multitalker babble or a single competing talker and presented to the implant. Stimulation to the acoustic region consisted of no signal, target speech, or a tone that was modulated in frequency to track the changes in the target talker's F0 and in amplitude to track the amplitude envelope of target speech low-pass filtered at 500 Hz.Results: All patients showed improvements in intelligibility over electric-only stimulation when either the tone or target speech was presented acoustically. The average improvement in intelligibility was 46 percentage points due to the tone and 55 percentage points due to target speech.Conclusions: The results demonstrate that a tone carrying F0 and amplitude envelope cues of target speech can provide significant benefit to CI users and may lead to new technologies that could offer EAS benefit to many patients who would not benefit from current EAS approaches.
The addition of low-frequency acoustic information to real or simulated electric stimulation (so-called electric-acoustic stimulation or EAS) often results in large improvements in intelligibility, particularly in competing backgrounds. This may reflect the availability of fundamental frequency (F0) information in the acoustic region. The contributions of F0 and the amplitude envelope (as well as voicing) of speech to simulated EAS was examined by replacing the low-frequency speech with a tone that was modulated in frequency to track the F0 of the speech, in amplitude with the envelope of the low-frequency speech, or both. A four-channel vocoder simulated electric hearing. Significant benefit over vocoder alone was observed with the addition of a tone carrying F0 or envelope cues, and both cues combined typically provided significantly more benefit than either alone. The intelligibility improvement over vocoder was between 24 and 57 percentage points, and was unaffected by the presence of a tone carrying these cues from a background talker. These results confirm the importance of the F0 of target speech for EAS (in simulation). They indicate that significant benefit can be provided by a tone carrying F0 and amplitude envelope cues. The results support a glimpsing account of EAS and argue against segregation.
Quinine causes a temporary disruption of outer hair cell (OHC) function, and thus can be used to examine the role of OHCs on auditory perception. In the present study, frequency selectivity, temporal resolution, and speech recognition were measured before, during, and after a quinine-induced hearing loss. Normal-hearing listeners ingested 5.76–11.43 mg/kg body weight of quinine, resulting in 5–15 dB of hearing loss. Frequency selectivity was estimated by comparing the level of a noise masker needed to mask a fixed-level, 2-kHz signal when the masker contained a spectral gap at 2 kHz and when it contained no gap. Similarly, temporal resolution was estimated by comparing the masker level needed to mask the 2-kHz signal when the noise masker contained a temporal gap and when it did not. Speech recognition thresholds were measured in quiet and in the presence of a masker (speech-shaped noise or time-reversed speech) fixed at 70-dB SPL. Signal level was varied adaptively to estimate 50% correct recognition. Quinine resulted in reduced frequency selectivity and reduced temporal resolution. Quinine also elevated speech recognition thresholds in quiet (by 7 dB on average), but the thresholds in the presence of the maskers were unaffected. [Work supported by NIDCD and AAA.]
Electric-acoustic stimulation (EAS) occurs when a cochlear implant recipient receives some amount of acoustic stimulation, typically as a consequence of residual low-frequency hearing. This combined stimulation often leads to significant improvements in speech intelligibility, especially in the presence of a competing background. We have demonstrated EAS benefit when target speech in the low-frequency region is replaced with a tone modulated in frequency with the target talker’s F0, and in amplitude with the amplitude envelope of the target speech. We have hypothesized that in some instances, the modulated tone may be more beneficial than speech. For example, a patient possessing extremely limited residual hearing (e.g., only up to 100 Hz) may not show EAS benefit because the F0 of many talkers is too high in frequency to be audible. In support of this hypothesis, we demonstrated that the tone can be shifted down considerably in frequency without affecting its contribution to speech intelligibility. In the current experiment, while generating the tone, we manipulated various signal-processing parameters to examine their effects on the benefits of EAS. One goal of this work is to develop a real-time processor that provides EAS benefit to implant users with limited residual hearing. [Work supported by NIDCD.]
Two experiments investigated the effects of critical bandwidth and frequency region on the use of temporal envelope cues for speech. In both experiments, spectral details were reduced using vocoder processing. In experiment 1, consonant identification scores were measured in a condition for which the cutoff frequency of the envelope extractor was half the critical bandwidth (HCB) of the auditory filters centered on each analysis band. Results showed that performance is similar to those obtained in conditions for which the envelope cutoff was set to 160Hz or above. Experiment 2 evaluated the impact of setting the cutoff frequency of the envelope extractor to values of 4, 8, and 16Hz or to HCB in one or two contiguous bands for an eight-band vocoder. The cutoff was set to 16Hz for all the other bands. Overall, consonant identification was not affected by removing envelope fluctuations above 4Hz in the low- and high-frequency bands. In contrast, speech intelligibility decreased as the cutoff frequency was decreased in the midfrequency region from 16to4Hz. The behavioral results were fairly consistent with a physical analysis of the stimuli, suggesting that clearly measurable envelope fluctuations cannot be attenuated without affecting speech intelligibility.
The present study sought to establish whether speech recognition can be disrupted by the presence of amplitude modulation (AM) at a remote spectral region, and whether that disruption depends upon the rate of AM. The goal was to determine whether this paradigm could be used to examine which modulation frequencies in the speech envelope are most important for speech recognition. Consonant identification for a band of speech located in either the low- or high-frequency region was measured in the presence of a band of noise located in the opposite frequency region. The noise was either unmodulated or amplitude modulated by a sinusoid, a band of noise with a fixed absolute bandwidth, or a band of noise with a fixed relative bandwidth. The frequency of the modulator was 4, 16, 32, or 64 Hz. Small amounts of modulation interference were observed for all modulator types, irrespective of the location of the speech band. More important, the interference depended on modulation frequency, clearly supporting the existence of selectivity of modulation interference with speech stimuli. Overall, the results suggest a primary role of envelope fluctuations around 4 and 16 Hz without excluding the possibility of a contribution by faster rates.
In the newest implementation of cochlear implant surgery, electrode arrays of 10 or 20 mm are inserted into the cochlea with the aim of preserving hearing in the region apical to the tip of the electrode array. In the current study two measures were used to assess hearing preservation: changes in audiometric threshold and changes in psychophysical estimates of nonlinear cochlear processing. Nonlinear cochlear processing was evaluated at signal frequencies of 250 and 500 Hz using Schroeder phase maskers with various indices of masker phase curvature. A total of 15 normal-hearing listeners and 13 cochlear implant patients (7 with a 10 mm insertion and 6 with a 20 mm insertion) were tested. Following surgery the mean low-frequency threshold elevation was 12.7 dB (125-750 Hz). Nine patients had postimplant thresholds within 5-10 dB of preimplant thresholds. Only one patient, however, demonstrated a completely normal nonlinear cochlear function following surgery--although most retained some degree of residual nonlinear processing. This result indicates (i) that Schroeder phase masking functions are a more sensitive index of surgical trauma than audiometric threshold and (ii) that preservation of a normal cochlear function in the apex of the cochlea is relatively uncommon but possible.
The improvement in amplitude modulation (AM) detection thresholds with increasing level of a sinusoidal carrier has been attributed to listening on the high-frequency side of the excitation pattern, where the growth of excitation is more linear, or to an increase in the number of “channels” via spread of excitation. In the present study, AM detection thresholds were measured using a 1000-Hz sinusoidal carrier. Thresholds for modulation frequencies of 4–64Hz improved by about 10–20dB as the carrier level increased from 10dB SL (14.5dB SPL on average) to 80dB SPL. To minimize the use of spread of excitation with an 80-dB carrier, tonal “restrictors” with frequencies of 501, 801, 1210, and 1510Hz were used alone and in combination. High-frequency restrictors elevated AM detection thresholds, whereas low-frequency restrictors did not, indicating that excitation on the high side is more important for detecting AM. Results of modeling suggest that the improvement in AM detection thresholds at high levels is likely due to the use of a relatively linear growth of response on the high-frequency side of the excitation pattern.
When low-frequency acoustic stimulation is combined with either real or simulated electric stimulation from a cochlear implant (electric-acoustic stimulation, or EAS), speech intelligibility in noise can improve dramatically. This improvement has been shown in simulation to be due in part to the presence of fundamental frequency (F0) and amplitude envelope information in the low-frequency region. The current experiment extends those findings to implant patients. Six patients who had residual low-frequency hearing in either their implanted or unimplanted ear participated. A target talker was combined with multitalker babble and presented to the implant. In the low-frequency region, patients heard either no stimulus, target speech, or a tone that was modulated in frequency to track the dynamic changes in F0, and in amplitude with the amplitude envelope of the low-pass target speech. Results showed that the tone provided, on average, about 58 percentage points of improvement over electric-only stimulation. Both the tone and target speech provided a statistically significant benefit over electric stimulation only (p<0.0001), and were statistically equivalent to each other (p>0.05). These results demonstrate that a tone that conveys F0 and amplitude envelope information can provide significant benefit in EAS.
Purpose To compare speech intelligibility in the presence of a 10-Hz square-wave noise masker in younger and older listeners and to relate performance to recovery from forward masking. Method The signal-to-noise ratio required to achieve 50% sentence identification in the presence of a 10-Hz square-wave noise masker was obtained for each of the 8 younger/older listener pairs. Listeners were matched according to their quiet thresholds for frequencies from 600 to 4800 Hz in octave steps. Forward masking was also measured in 2 younger/older threshold-matched groups for signal delays of 2–40 ms. Results Older listeners typically required a significantly higher signal-to-noise ratio than younger listeners to achieve 50% correct sentence recognition. This effect may be understood in terms of increased forward-masked thresholds throughout the range of signal delays corresponding to the silent intervals in the modulated noise (e.g., <50 ms). Conclusions Significant differences were observed between older and younger listeners on measures of both speech intelligibility in a modulated background and forward masking over a range of signal delays (0–40 ms). Age-related susceptibility to forward masking at relatively short delays may reflect a deficit in processing at a fairly central level (e.g., broader temporal windows or less efficient processing).