
A recent study from our lab showed that normal-hearing adults rely on cues in the region of 0.5-1 kHz for target talker localization and speech-in-speech recognition when the target can only be identified based on spatial cues. The present study evaluated whether similar weights are observed for speech-in-speech recognition that relies on access to vowels or consonants, given differences in the distributions of cues required for vowel and consonant recognition. Spectral weights were estimated by filtering stimuli into 1-octave-wide bands, dispersing those bands on the horizontal plane, and assessing the association between each band's position and participants' responses. Data were obtained for two tasks: 1) localization of the target relative to the masker (left vs. right), and 2) closed-set word recognition, where response alternatives differed either with respect to the initial consonant, vowel, or final consonant. As observed previously, weights for both tasks peaked at 0.5-1 kHz for all three stimulus sets. This result is consistent with the idea that the spatial cues supporting spatial release from masking are similar regardless of the frequency regions containing speech cues required for correct recognition of the target.
Unequal weighting of binaural information across frequency can reduce sensitivity in the presence of competing but uninformative cues (“binaural interference”), a potentially serious problem for listeners who use combined electric and acoustic (EAS) hearing. Here, we used virtual-reality techniques to measure spectral weighting functions (SWF) during localization of simulated EAS stimuli [see van Ginkel et al., JASA 145, 2445 (2019)]: low-frequency “acoustic” noise bands and high-frequency “electric” click trains. This study was conducted remotely at each participants’ location, using sanitized and calibrated research equipment (Oculus Quest 2 headsets and Sennheiser HD 280 Pro headphones) delivered by lab personnel. Sounds were presented with random binaural-cue jitter across seven components on each trial. SWFs were computed by statistical regression of localization responses onto binaural cue values. SWFs confirmed the dominance of 400–1000 Hz components reported previously [Stecker, Folkerts, & Stecker 2019(A), JASA 145:1720]. Introducing a gap between noise and click-train components increased weight on neighboring components, consistent with reduced interference in such conditions [van Ginkel et al., 2019]. Finally, presenting noises and click trains from competing locations reduced weight on non-target components for most but not all participants, a potential marker of susceptibility to binaural interference. [Work supported by NIH R01-DC016643, NIH T35-008757.]
The vocal folds experience repeated collision during phonation. The resulting contact pressure is often considered to play an important role in vocal fold injury, and has been the focus of many experimental studies. In this study, vocal fold contact pattern and contact pressure during phonation were numerically investigated. The results show that vocal fold contact in general occurs within a horizontal strip on the medial surface, first appearing at the inferior medial surface and propagating upward. Because of the localized and travelling nature of vocal fold contact, sensors of a finite size may significantly underestimate the peak vocal fold contact pressure, particularly for vocal folds of low transverse stiffness. This underestimation also makes it difficult to identify the contact pressure peak in the intraglottal pressure waveform. These results showed that the vocal fold contact pressure reported in previous experimental studies may have significantly underestimated the actual values. It is recommended that contact pressure sensors with a diameter no greater than 0.4 mm are used in future experiments to ensure adequate accuracy in measuring the peak vocal fold contact pressure during phonation.
Cochlear implant (CI) users experience considerable difficulty in understanding speech in reverberant listening environments. This issue is commonly addressed with time-frequency masking, where a time-frequency decomposed reverberant signal is multiplied by a matrix of gain values to suppress reverberation. However, mask estimation is challenging in reverberant environments due to the large spectro-temporal variations in the speech signal. To overcome this variability, we previously developed a phoneme-based algorithm that selects a different mask estimation model based on the underlying phoneme. In the ideal case where knowledge of the phoneme was assumed, the phoneme-based approach provided larger benefits than a phoneme-independent approach when tested in normal-hearing listeners using an acoustic model of CI processing. The current work investigates the phoneme-based mask estimation algorithm in the real-time feasible case where the prediction from a phoneme classifier is used to select the phoneme-specific mask. To further ensure real-time feasibility, both the phoneme classifier and mask estimation algorithm use causal features extracted from within the CI processing framework. We conducted experiments in normal-hearing listeners using an acoustic model of CI processing, and the results showed that the phoneme-specific algorithm benefitted the majority of subjects.