
Cochlear implants (CIs) are considered effective at restoring some hearing abilities to individuals with severe hearing loss. However, there is considerable variability in performance outcomes among adults with CIs. One likely source of variability is the quality of the electrode-neuron interface (ENI), which describes how well implanted electrodes stimulate nearby spiral-ganglion neurons (SGNs) in the cochlea. Absolute thresholds, measured via focused electrical stimulation, have been proposed as a behavioral estimate of ENI. In this retrospective study, focused-stimulation threshold patterns were compared for 92 ears from 74 adult cochlear implant listeners with various etiologies, durations of deafness, and electrode array types. The results showed that CI listeners implanted with a pre-curved array tended to have lower overall focused thresholds than listeners with straight arrays, as well as a threshold profile with lower (better) thresholds at each end of the array and highest thresholds near the middle of the array. CI listeners with etiologies related to auto-immune dysfunction, such as Meniere’s disease, tended to have the opposite profile, with lowest thresholds near the middle of the array. No significant effect of duration of deafness on focused thresholds was observed. Further investigation into which demographic and device factors affect focused thresholds, and therefore potentially ENI, could shed light on the variability of CI outcomes and guide development of individualized treatment strategies.
Digit-perception tasks are widely used to assess auditory function in clinical and research settings, but performance may also be shaped by non-auditory participant characteristics. This study examined whether educational attainment, forward digit span, socioeconomic deprivation, and musical experience predicted performance on three digit-perception tasks that differed in masker type and spatial demands: digits in speech-shaped noise (SSN), digits in spatially separated babble, and a spatial-selective-attention task in which digit streams were separated by 15°. Participants were 220 young adults (aged 16-17 years at baseline) with bilaterally normal audiometric thresholds, each tested twice. The narrow age range and normal hearing constrained variation in peripheral auditory function, increasing sensitivity to non-auditory influences. Linear mixed-effects models were used to examine single-predictor and multivariate associations between the predictors and each outcome. Digits-in-SSN performance was unrelated to all predictors. Digits-in-babble thresholds showed modest associations with deprivation and educational attainment in single-predictor models, but neither survived multivariate adjustment. By contrast, spatial-selective-attention performance showed robust independent associations with educational attainment, deprivation, and musical experience after mutual adjustment. Forward digit span was not a significant predictor of any task. These findings indicate that spatially demanding digit-perception tasks are vulnerable to non-auditory confounds, and that controlling for short-term memory may be insufficient to address this problem. A secondary analysis examined whether the association between musical experience and spatial selective attention might reflect confounding rather than auditory training. The musician advantage was attenuated by approximately one third after adjustment for sociodemographic factors, consistent with a substantial contribution of confounding.
Hearing protection devices (HPDs) are widely used to mitigate the risks of noise-induced hearing loss in industrial environments. However, conventional HPDs can often hinder communication for workers with hearing loss. While those workers may rely on hearing aids in their daily lives, the potential benefits and associated risks of using such devices in noisy workplaces remain unexplored. This study evaluated the effects of different hearing aid signal processing algorithms on speech intelligibility and estimated noise exposure under industrial noise conditions to characterize the balance between communication and protection needed to develop solutions for workers with hearing loss. A digital hearing protection research platform integrating in-ear receivers and external microphones was configured as a real-time soundcard. Typical hearing aid algorithms, including linear amplification, Wide Dynamic Range Compression (WDRC), digital noise reduction (DNR), and compression limiting, were implemented in different configurations. Two groups, one with normal hearing and one with hearing loss, completed the Hearing in Noise Test (HINT) with a wood-working factory noise at two presentation levels, along with subjective ratings of listening effort and voice clarity and naturalness . Noise exposure for each configuration was also estimated. Results showed significant effects of signal processing strategies and noise presentation levels on both intelligibility and noise exposure. Combining DNR, WDRC and compression limiting showed promising results to balance intelligibility and noise exposure, particularly for participants with hearing loss. The findings emphasize the trade-off between auditory benefit and exposure risk, and call for greater awareness regarding the use of hearing devices in industrial noise.
Previous works that have analyzed electroencephalographic (EEG) speech envelope encoding in hearing-aid users using stimulus reconstruction (SR) approaches have shown that noise reduction (NR) algorithms can improve the neural tracking of speech in noise. Beyond enhancing speech envelope encoding, NR may also restore the perceptibility of acoustic edges that trigger auditory late responses (ALRs), likely promoting more consistent ALR elicitation and thus enhancing brain-to-speech synchronization. To investigate this hypothesis, 26 hearing-impaired listeners were fitted with bilateral hearing aids and participated in an EEG experiment with the instruction to listen to a continuous single-speaker speech stimulus under speech-in-quiet conditions and speech-in-noise conditions with and without NR. Event-related potential analyses of neural responses to acoustic edges in the speech onset envelope revealed that P1 component amplitudes in averaged ALRs decreased significantly for speech-in-noise compared to speech-in-quiet conditions, but partially recovered when NR was activated. These observations correlated significantly with the degree of spectrotemporal consistency across single-trials for the associated ALR time-frequency component. Additionally, a SR analysis for the speech onset envelope produced qualitatively identical results that were highly correlated with those from the spectrotemporal consistency analysis. These findings suggest that hearing aid NR algorithms partly enhance neural speech tracking by restoring neural responses to acoustic edges. In turn, our results may support the future development of hearing aid processing strategies that could facilitate neural tracking of speech based on neural principles.
Pupillometry has been shown to be a correlate for expected mental workload in a variety of cognitively demanding tasks, including those related to listening. Recent work has extended the application of pupillometry to interactive conversation by measuring how the pupil responds around turn-taking. This study furthers the application of pupillometry to conversation by evaluating differences in pupil responses during conversations taking place between individuals with and without hearing loss. Pairs of older Danish talkers (in which one had hearing loss, and one did not) held conversations in a variety of acoustic conditions, and their pupil responses were recorded throughout. Our findings show that both background noise and hearing loss modulate pupil responses during conversation, and further show correspondence between pupil dilation responses and measures of turn-taking behavior which have been suggested to be indicative of increased difficulty.
Reverberation can degrade speech intelligibility and comprehension, increase cognitive load, and impair memory. Many studies examining reverberation effects use audio-only conditions and stationary background noise. Thus, it remains unclear to what extent findings generalise beyond laboratory setups. In the present study, we examine the influence of reverberation on memory and behavioural listening effort in a complex, audiovisual virtual reality (VR) environment. We implemented a virtual reality (VR) open-plan office with realistic room acoustics and time-varying background noise, including animated virtual agents. Three simulated room-acoustic conditions were compared: a free-field ( T 30 = 0 s), an acoustically optimised office ( T 30 = 0.5 s), and an untreated office ( T 30 = 1.0 s). Under these conditions, N = 30 participants performed a dual-task. The primary Heard-Text Recall (HTR) task measured speech comprehension and memory. Participants listened to short stories narrated by two speakers and answered content-related questions. The vibrotactile secondary task measured listening effort. In single-tasking, memory for spoken content was significantly better in free-field than in both offices, with no significant difference between offices. In dual-tasking, vibrotactile response times indicated that listening effort increased in the presence of reverberation. Across single- and dual-task conditions, vibrotactile error rates decreased in the presence of reverberation, possibly due to elevated concentration in more difficult conditions. These findings suggest that room acoustics, compared to free-field environments, influence memory and listening effort in complex audiovisual VR and highlight the importance of incorporating reverberations when designing audiovisual environments for listening research.
Reliable hearing assessment at home can improve accessibility and reduce dependence on in-clinic testing. To be viable, home-based procedures must provide accurate results within short measurement times and remain robust to factors such as ambient noise and variable user attention. In this proof-of-concept study with a limited sample size, we evaluated two such procedures—a Graded Response Bracketing method (GRaBr) for pure-tone threshold estimation and a reinforced adaptive categorical loudness scaling method (rACALOS) for loudness-growth assessment - using remote, smartphone-based testing, and compared their outcomes with those obtained inside the laboratory using established reference procedures. Fifteen young adults with normal hearing completed the tasks at home and in the laboratory. Test–retest reliability was assessed by repeating the home measurements within one week of the initial session. Ambient noise levels in home environments were also recorded. The intraclass correlation coefficients for GRaBr measured at home exceeded 0.75, indicating good test–retest reliability. Similarly, home-based rACALOS showed generally high reliability, with across-run biases below 5 dB at all frequencies; however, mean interquartile ranges reached up to 18.4 dB at some loudness categories. Remote GRaBr measurements showed small, non-significant deviations from laboratory audiometry (0.4 ± 7.1 dB); however, frequency-specific biases of approximately ±5 dB were observed at 250 Hz and 4 kHz, with root-mean-square errors ranging from 4.6 to 8.0 dB across frequencies. Compared with the in-lab ACALOS, remote rACALOS showed an overall bias of 3.38 ± 11.8 dB across loudness categories, with no significant overall effect of test environment. These findings suggest that smartphone-based pure-tone audiometry and loudness-scaling assessments provide accurate and reliable results when using these procedures at home, at least in normal-hearing young adults, who are experienced loudness raters, under suitable acoustic conditions with low ambient noise.
This study evaluated listening-related fatigue among adolescents with prelingual hearing loss (HL) and long-term device use; and examined demographic, audiological, and functional factors associated with increased fatigue. A total of 271 16–19-yr-old adolescents fitted with their first hearing device by 3 years of age, and their parents completed the self- and parent-report versions of the Vanderbilt Fatigue Scale, respectively. Of these, 53.9% (n=136) used hearing aids (HA), 46.0% (n=116) used at least one cochlear implant (CI), <2% (n=4) used bone conduction (BC) devices, and 7% (n=15) were unaided. Group mean fatigue scores of adolescents were within one standard deviation of the normative mean. Yet, scores of 24.5–50.5% of HA and CI users showed clinically significant fatigue. BC users showed no indications of high fatigue. Among HA and CI users, limited associations were evident between fatigue and audiological factors such as device type, hearing level, or age at first device fit. Greater fatigue reported by adolescents and parents was associated with female gender, greater communication difficulties, and more symptoms of anxiety and depression. The gender effect was no longer significant after accounting for communication and mental well-being. In CI users, greater parent-reported fatigue was associated with more frequent use of school accommodations. In HA users, better expressive language was associated with lower self-reported fatigue. Overall, listening-related fatigue appears to be a concern in adolescents with early-onset HL despite early intervention and long-term device use. Communication and well-being factors appear more informative than audiological variables for addressing fatigue in this group.
This study investigated older and younger adults’ conversational speech understanding using an established topic monitoring task to assess real-time semantic processing. Twenty-three younger adults (mean age = 26.5 years) and 19 older adults (mean age = 71.3 years) monitored for topic targets in 24 dialogues (2-3 minutes each). The impact on monitoring response time and accuracy of two message-level factors (when speakers switch; contextual predictability) was assessed while manipulating two signal-level factors (presentation modality; noise level). Results showed that presentation modality affected response times, with faster correct responses for audio-visual versus audio-only presentation, an effect that interacted with age (older adults had a larger visual speech benefit). Noise (0 dB SNR) increased response times and decreased accuracy compared to presentation in quiet. A speaker switch delayed response time to the first target after the switch compared to non-switch targets. When the target occurred relatively soon after a speaker switch, older adults showed a significant switch cost, whereas younger adults did not. Targets that were more predictable from context (low surprisal) had faster response times. In the most challenging listening condition (Audio-only presentation in noise), the effect of contextual predictability for older adults was marginally larger than for younger adults, suggesting greater use of context. These findings highlight the utility of the topic monitoring task for probing conversational speech comprehension across the lifespan.
Young adults with normal conventional audiograms may nonetheless show early noise-related auditory differences that routine surveillance misses. We evaluated a clinically feasible test battery using 280 occupationally noise-exposed workers and 100 age-matched low-noise controls (20-40 years; ≤20 dB HL at 0.5-8 kHz). Measures included conventional and extended high-frequency (EHF; 9-16 kHz) audiometry, distortion-product otoacoustic emissions (DPOAEs; 0.5-10 kHz), ipsilateral acoustic reflex thresholds (ARTs; 0.5, 1, and 2 kHz; 226-Hz probe tone), click-evoked auditory brainstem responses (ABR; wave I, wave V, and V/I ratio at 50 and 90 dBnHL), and Mandarin Bamford-Kowal-Bench speech-in-noise testing (SiN). Noise-exposed participants showed elevated ARTs at all elicitor frequencies, with the largest elevation at 2 kHz, as well as poorer EHF thresholds and lower DPOAE levels. In contrast, click-evoked ABR metrics did not differ between groups. After controlling for conventional-frequency pure-tone average (PTA 4 ; 0.5, 1, 2, and 4 kHz) and extended high-frequency pure-tone average (PTA EHF ; 9–16 kHz), lower DPOAE amplitudes were associated with higher ARTs. ARTs were not significantly associated with SiN performance, quantified as the signal-to-noise ratio (SNR) required for 50% correct recognition (SNR-50), cumulative noise exposure, or ABR measures. The noise-exposed group showed higher SNR-50 values, indicating poorer speech-in-noise performance. Within the noise-exposed group, SNR-50 was associated with PTA 4 and cumulative noise exposure, whereas its association with PTA EHF was weaker and did not remain significant after correction for multiple comparisons. Together, these findings suggest that cochlear-status measures, conventional ARTs, and functional SiN performance provide complementary information beyond the conventional audiogram in occupational hearing surveillance.
This study addresses the common issue of low externalization in binaural synthesis when listening to virtual sound sources through wearable devices. The improvement of externalization of frontal sound sources (0°–30° azimuth), which are perceived as being closer to the listener, was investigated. A consistent pattern across different rooms was identified using a single-speaker non-individualized binaural room impulse response (NI-BRIR). The binaural level decreases regularly with increasing azimuth up to 45°(front right) but levels off beyond this point. Based on the strong correlation among binaural level, azimuth, and externalization, we developed a two-phase approach. In Phase I, scaling factors were applied to the NI-BRIRs to equalize the binaural level across different frontal azimuths. Our results showed this adjustment significantly improved externalization, with a large effect size for azimuths from 0°–30°. Building on this adjustment, we further enhanced the early reflections of the NI-BRIR for 0° and 15° to simulate the temporal signal conditions of angles with better externalization (Phase II), aiming to modify the ratio between direct sound and main reflections. This modification had a positive and additive effect. The effective enhancement of externalization of frontal virtual sound sources was demonstrated by a consistent set of scaling factors and strategic adjustments to early reflections. These findings suggest a practical corrective method for single-speaker NI-BRIRs, contributing to more realistic binaural synthesis in applications such as wearable audio technology and hearing support systems.
Tonotopy is a fundamental feature of auditory cortical organization, yet its influence on cortical auditory-evoked responses (AERs) remains unclear. Consequently, key properties of cortical AERs-such as their marked amplitude reduction with increasing stimulus frequency-still lack a coherent mechanistic explanation. To address this gap, we combined a meta-analysis of frequency-specific AER amplitudes with forward simulations of AERs informed by current knowledge of auditory cortical tonotopic layout and functional organization. The meta-analysis used a semi-systematic search covering all known automatic cortical AER components-both transient-evoked and steady-state-along with selected subcortical components for comparison. Forward simulations were based on a functional parcellation of the human supratemporal auditory region into subdivisions forming distinct tonotopic maps, and an idealized model of each division's intrinsic tonotopic layout. Parcellation was achieved using a novel, largely automated procedure applied to high-field (3 T) and ultra-high-field (7 T) functional and microstructural MRI mapping data from 30 individual hemispheres. Meta-analytic results revealed that, whilst all cortical AER components consistently show frequency-related amplitude reduction, reduction is greater in steady-state compared to transient-evoked components. Simulations indicated that frequency-related amplitude reduction arising from cortical morphology is confined to the highly myelinated central portion of Heschl's gyrus, suggesting that differences in reduction amount between steady-state and transient-evoked components may reflect differences in the relative strengths of their primary contributions. Our findings provide a new perspective on cortical AER generation. They represent an important step toward explaining morphology-related variability in AER amplitudes and establishing a quantitative link to underlying source strengths.
Understanding speech in noisy environments (SPIN) is an important everyday ability, and engaging in musical activities has been proposed as a factor that may support this ability. However, the cognitive mechanisms underlying a potential musical advantage in SPIN perception remain unclear. Here we investigated whether musical sophistication is associated with better SPIN perception in a large population-based sample, and whether this relationship is mediated by auditory working memory (AWM), verbal working memory (VWM), or non-verbal intelligence. We recruited 203 participants and measured SPIN perception at both word and sentence levels. Musical sophistication was assessed using the Goldsmiths Musical Sophistication Index (Gold-MSI). AWM was measured using delayed matching of tone frequency or the modulation rate of amplitude modulated white noise, VWM was based on backward digit span task, and non-verbal intelligence used matrix reasoning. Mediation analyses revealed that AWM fully mediated the relationship between musical sophistication and SPIN perception, whereas VWM showed no mediation effect. Non-verbal intelligence showed a partial mediating effect. Additional control analyses using structural equation modelling revealed that the indirect effect through AWM remained significant after accounting for age, hearing thresholds, and non-verbal intelligence. Together, these findings suggest that greater musical sophistication is associated with better daily-life listening abilities, and that AWM may be an important cognitive factor linking musical sophistication and SPIN perception.
Binaural fusion occurs when signals arriving separately at the two ears are perceived as one sound. Previously, it was shown that normal-hearing (NH) or hearing-impaired (HI) listeners fused and perceptually averaged dichotic vowels presented through headphones. HI listeners even fused vowels with different fundamental frequencies (F 0 s), indicating a new alternative mechanism underlying difficulties with speech segregation. To extend these findings to a more real-world listening context, the current study investigated if vowel fusion and averaging would also occur in the free field. Synthesized vowels/with F 0 s of 106.9, 151.2, or 201.8 Hz and were presented simultaneously via loudspeakers located on the left and right at ±60° relative to the listener’s head. Listeners indicated which single vowel or two vowels they heard. Similar trends of vowel fusion were observed as in the previous study. For ΔF 0 =0, both groups showed high rates of vowel fusion. For ΔF 0 >0, NH listeners showed decreased fusion, but fusion remained high for HI listeners. Compared to headphone presentation, overall vowel fusion occurred less frequently under free field presentation, with different patterns of perceptual averaging. Results suggest that cues available in the free field may help reduce vowel fusion, especially in NH listeners and for vowels with the same F0. However, HI listeners still experience excessive vowel fusion in the free field, even with ΔF 0 as large as 6–11 semitones. These findings reinforce the importance of reducing both the blending of multiple speech streams and excessive fusion to address difficulties with understanding speech in noise.
Abstract This study investigates the impact of uncalibrated consumer hardware on the outcomes of the German Matrix Sentence Test (GMST) in a remote setting. Speech-in-noise tests like the MST are relevant in audiology, as they better reflect everyday listening conditions compared to pure-tone audiometry. Due to their suprathreshold signal presentation and robustness against calibration offsets, they are particularly suitable for remote testing. It is unclear, however, if a sensitive and efficient test like the MST exhibits the same robustness against uncontrolled conditions in remote testing with uncalibrated consumer devices. This motivates the systematic assessment of the online MST’s reliability and its agreement with laboratory administration, conducted in normal-hearing participants using the browser-based platform “Virtual Hearing Clinic” (VHC). Twenty normal hearing participants completed the test in a randomized sequence: once in a lab with calibrated equipment and twice remotely via the VHC using uncalibrated consumer devices (smartphones, laptops, headphones). Results revealed no significant influence of headphone type, test environment, or pure-tone average. The remote condition demonstrated high test–retest reliability (r = 0.82) and strong correlation with lab-based measurements (r = 0.78). A mean bias of 0.29 dB in speech recognition thresholds was observed with slightly better performance in the lab, though not statistically significant (p = 0.15). Overall, results demonstrate that the MST, when implemented remotely and using uncalibrated consumer hardware, yields results of accuracy and reliability comparable to those obtained in the lab for normal hearing, supporting further applications of remote hearing assessments with optimized test materials.
Designing multi-sensor experiments to study interactive communication is costly, complex, and often poorly documented, particularly in terms of implementation challenges and the rationale behind decision-making. Here, we aim to demystify the process by detailing the development of a custom multi-sensor lab designed to monitor comprehensive behavioural and physiological responses during interactive communication. To illustrate the system architecture, we use an implemented paradigm as a worked example: two adults engaged in natural conversation while listening to realistic background noise scenes delivered via open Sennheiser HD-800 headphones. Multi-sensor data were collected using DPA headset microphones, Tobii Pro Glasses 3, a Vicon motion capture system (with Tobii integration), BIOPAC amplifiers (PPG, EDA, ECG, respiration, temperature), a Fitbit Sense 2 smartwatch, and six Logitech BRIO 4K ultra HD video cameras recording via Open Broadcaster Software (OBS). Due to high bandwidth demands and varying sampling rates (1 Hz to 48 kHz) among the different devices, all systems recorded independently which posed a significant synchronisation challenge. A central RME soundcard delivered trigger pulses to the Tobii, BIOPAC, OBS, and Vicon systems to mark stimulus onset/offset. Redundant synchronisation mechanisms were implemented to mitigate the risk of trigger failure and bespoke processing pipelines were developed to synchronise data streams. We share key challenges and creative solutions encountered in building such a lab, to help researchers understand and anticipate common pain points and make more informed decisions when implementing their own multi-sensor setups.
Cochlear implants (CIs) effectively restore speech understanding in quiet for many recipients. However, bilateral CI listeners continue to perform poorer than acoustic-hearing listeners on spatial-hearing tasks such as sound localization. One explanation is that stimulation strategies and the lack of coordination across sound processors degrade the fidelity of binaural cues. To provide further insight on this issue, this study compared the binaural cues transmitted by two CI stimulation strategies: a spectral peak-picking stimulation strategy and a constantly stimulating strategy. Two clinical CI sound processors were placed on a binaural manikin centered in a horizontal ring of loudspeakers in an anechoic chamber. Speech and non-speech stimuli were presented from the loudspeakers and the resulting electrical pulse trains were recorded. Envelope interaural time differences and interaural level differences were extracted from the recordings. Binaural cues were compared between the two stimulation strategies and against acoustic cues measured using behind-the-ear microphones. Results showed that, compared to the constantly stimulating strategy, the spectral peak-picking strategy led to more variable and less monotonic binaural cues, which could constrain sound-localization abilities. These findings highlight the need for bilateral coordination of CIs in multiple domains, including electrode selection in peak-picking stimulation strategies and pulse timing. (198/250).
Noise-induced and age-related hearing loss involve partially similar pathophysiology, but their combined effects remain poorly understood, especially in under-regulated occupational settings. This study investigated the independent and interaction effects of lifetime noise exposure and age on audiometric, speech-in-noise, and tinnitus outcomes in Palestinian workers. Ninety-eight adult workers aged 18 – 75 years completed standard (0.25 – 8 kHz) and extended high-frequency (10 – 14 kHz) audiometry, Arabic diotic and antiphasic digits-in-noise (DIN) testing, and an Arabic Noise Exposure Structured Interview. Regression models tested age × noise interactions, adjusting for sex (and education for DIN), with Bonferroni–Holm correction applied for multiple comparisons. The protocol was pre-registered (Open Science Framework; https://osf.io/v2c9z ). There was no significant age × noise interaction for standard audiometric thresholds (3 – 8 kHz; β = 0.23 dB/year per log 10 [noise], 95% CI 0.03 – 0.42; p = 0.02), after correction for multiple comparisons. Age-stratified analyses showed greater exposure was associated with worse standard audiometric thresholds in older adults (≥44 years; β = 8.1 dB per 10-fold increase, 95% CI 2.5 – 13.6; p = 0.005) but not younger adults. No interaction or main effects of noise were observed for extended high-frequency thresholds, DIN, or tinnitus ( p > 0.05). Age was significantly associated with standard and extended high-frequency audiometric thresholds and diotic and antiphasic DIN thresholds. Auditory outcomes were primarily age-driven, with limited noise effects. Greater noise-related threshold loss appeared in older adults, but age effects were larger, suggesting ageing as the dominant determinant. Hearing conservation should address age-related vulnerability alongside exposure control.
Voice cloning technology has developed rapidly, and current synthesis techniques can produce humanlike voices. Recent research has shown that cloned voices can now be more intelligible than their human originals in background noise, a signal-extrinsic degradation. Here, we aimed to establish whether this intelligibility benefit for voice clones degraded using noise-vocoding, a signal-intrinsic manipulation that severely degrades fine-grained spectrotemporal and harmonic information in the speech signal. We also investigated whether listeners can perceptually adapt to noise-vocoded voice clones. We compared the intelligibility of ten six-band noise-vocoded voice clones with their human originals. Listeners heard 80 sentences, in two counterbalanced blocks: a block of 40 sentences by voice clones and a block of 40 sentences by human, to compare perceptual adaptation to both speech types. Listeners showed 13.2% higher accuracy for voice clones. Overall, listeners adapted to the noise-vocoded speech and improved by around 16% over the course of 40 sentences in the first block only, regardless of whether voices were human or cloned. The intelligibility benefit for cloned voices thus persisted for noise-vocoded speech, and listeners adapted to noise-vocoded cloned voices as much as for noise-vocoded human voices. Our results have implications for applications of voice clones, in particular in future development of assistive listening technologies.
The purpose of this study was to compare the effects of source-specific (independent) and conventional dynamic range compression (DRC) on sound quality ratings among listeners with hearing loss when ground-truth signals are available to the compressor. Twenty listeners with mild to moderately severe sensorineural hearing loss rated the sound quality on different subscales for two types of signal mixtures: speech in music (Overall Sound Quality, Speech Clarity, and Music Pleasantness) and speech in noise (Overall Sound Quality, Speech Clarity, and Noisiness) at three speech-to-background ratios (SBR: -10, 0, +10 dB). Speech in music was a 10-second-long spontaneous speech excerpt with a duration-matched classical music excerpt. Speech in noise had the same speech signal mixed with speech-shaped noise. Conventional DRC applied a 50-ms release time to the mixed signals, whereas independent DRC applied a 50-ms release time for speech and a 2000-ms release time for music or noise before mixing. The control conditions included linear amplification applied to the signals before and after mixing. Independent DRC resulted in higher overall sound quality, music pleasantness, and lower noisiness ratings than conventional DRC, regardless of SBR. Independent DRC resulted in higher speech clarity ratings than conventional DRC at lower SBRs. The results are generally supported by acoustic metrics, indicating more effective compression and higher output SBRs, especially at lower input SBRs, and reduced across-source modulation correlations with independent DRC. These findings extend the literature on the advantages of independent compression of sound sources by demonstrating sound quality benefits across multiple sub-scales for listeners with hearing loss across music and noise backgrounds.