Environmental sound recognition (ESR) enables listeners to interpret complex acoustic environments, yet the frequency regions that support recognition are poorly understood. This study used deep learning to model ESR in competing speech and estimate frequency band-importance functions (BIFs) underlying recognition performance. Trial-level responses were collected from 46 listeners who identified 25 everyday sounds mixed with speech across a wide range of target-to-masker ratios. Two model variants were evaluated: one trained to mimic human performance, which was trained on soft labels derived from listener responses, and one trained for maximum accuracy, which was trained on ground-truth correct sound labels, enabling a direct comparison between perceptually driven and task-optimal band-importance patterns. The human-trained model closely reproduced key features of human performance, whereas the ground-truth-trained model exceeded human accuracy and showed highly reliable performance across cross-validation folds. BIFs were estimated by bandstop filtering the target signal and quantifying the resulting drop in recognition accuracy. Both model variants yielded reproducible BIFs with five prominent peaks (∼0.43, 0.77, 1.46, 2.6, and 9.7 kHz), largely driven by subsets of sounds having sharply tuned spectral dependence. This convergence across training objectives suggests that human performance closely reflects the task-optimal frequencies for segregating environmental sounds from speech maskers.
Speech perception is essential for successful communication through spoken language, which often involves both speaking and listening by two or more conversation partners. The speech perception abilities of one conversation partner affect the dynamics of the entire conversation, not just the passive speech-recognition performance of that individual. However, speech perception is often evaluated through single-person listening tasks in controlled settings, such as an audiometric booth. While this type of testing can provide insight into an individual's ability to hear and understand speech passively, it is not representative of typical real-world conversations, and it fails to capture the effects of one individual’s speech perception abilities on other conversational participants. This pilot study begins the development and validation of a measure of conversational dynamics for the purpose of evaluating hearing interventions such as hearing aids. Five pairs of conversation partners participated in two experiments. The first involved a series of semi-structured conversation tasks in a controlled laboratory environment. The second involved a naturally occurring conversation in an uncontrolled environment. In both experiments, one conversation partner wore earplugs during half of the experimental trials. Objective and subjective measures of conversational dynamics revealed that the semi-structured laboratory task could serve as an ecologically valid proxy for real-world communication.
Hearing impairment is often characterized by poor speech-in-noise recognition. State-of-the-art laboratory-based noise-reduction technology can eliminate background sounds from a corrupted speech signal and improve intelligibility, but it can also hinder environmental sound recognition (ESR), which is essential for personal independence and safety. This paper presents a time-frequency mask, the ideal compressed mask (ICM), that aims to provide listeners with improved speech intelligibility without substantially reducing ESR. This is accomplished by limiting the maximum attenuation that the mask performs. Speech intelligibility and ESR for hearing-impaired and normal-hearing listeners were measured using stimuli that had been processed by ICMs with various levels of maximum attenuation. This processing resulted in significantly improved intelligibility while retaining high ESR performance for both types of listeners. It was also found that the same level of maximum attenuation provided the optimal balance of intelligibility and ESR for both listener types. It is argued that future deep-learning-based noise reduction algorithms may provide better outcomes by balancing the levels of the target speech and the background environmental sounds, rather than eliminating all signals except for the target speech. The ICM provides one such simple solution for frequency-domain models.
One goal of hearing aids is to increase the intelligibility of speech by improving the signal-to-noise ratio through the use of directional microphones and/or digital noise reduction. With this objective in mind, an ideal hearing aid would output amplified target speech with a signal-to-noise ratio of +∞. However, even in this ideal scenario, vent effects would allow background noise to enter the ear canal via the direct sound path while also allowing amplified sound to escape from the ear canal. In this experiment, a range of hearing aid fittings were simulated to account for these vent effects. Hearing-impaired listeners were tested using the ideal simulated hearing aid, and the results will show the maximum theoretical hearing aid benefit.
OBJECTIVES:This study aimed to determine the speech-to-background ratios (SBRs) at which normal-hearing (NH) and hearing-impaired (HI) listeners can recognize both speech and environmental sounds when the two types of signals are mixed. Also examined were the effect of individual sounds on speech recognition and environmental sound recognition (ESR), and the impact of divided versus selective attention on these tasks. DESIGN:In Experiment 1 (divided attention), 11 NH and 10 HI listeners heard sentences mixed with environmental sounds at various SBRs and performed speech recognition and ESR tasks concurrently in each trial. In Experiment 2 (selective attention), 20 NH listeners performed these tasks in separate trials. Psychometric functions were generated for each task, listener group, and environmental sound. The range over which speech recognition and ESR were both high was determined, as was the optimal SBR for balancing recognition with ESR, defined as the point of intersection between each pair of normalized psychometric functions. RESULTS:The NH listeners achieved greater than 95% accuracy on concurrent speech recognition and ESR over an SBR range of approximately 20 dB or greater. The optimal SBR for maximizing both speech recognition and ESR for NH listeners was approximately +12 dB. For the HI listeners, the range over which 95% performance was observed on both tasks was far smaller (span of 1 dB), with an optimal value of +5 dB. Acoustic analyses indicated that the speech and environmental sound stimuli were similarly audible, regardless of the hearing status of the listener, but that the speech fluctuated more than the environmental sounds. Divided versus selective attention conditions produced differences in performance that were statistically significant yet only modest in magnitude. In all conditions and for both listener groups, recognition was higher for environmental sounds than for speech when presented at equal intensities (i.e., 0 dB SBR), indicating that the environmental sounds were more effective maskers of speech than the converse. Each of the 25 environmental sounds used in this study (with one exception) had a span of SBRs over which speech recognition and ESR were both higher than 95%. These ranges tended to overlap substantially. CONCLUSIONS:A range of SBRs exists over which speech and environmental sounds can be simultaneously recognized with high accuracy by NH and HI listeners, but this range is larger for NH listeners. The single optimal SBR for jointly maximizing speech recognition and ESR also differs between NH and HI listeners. The greater masking effectiveness of the environmental sounds relative to the speech may be related to the lower degree of fluctuation present in the environmental sounds as well as possibly task differences between speech recognition and ESR (open versus closed set). The observed differences between the NH and HI results may possibly be related to the HI listeners' smaller fluctuating masker benefit. As noise-reduction systems become increasingly effective, the current results could potentially guide the design of future systems that provide listeners with highly intelligible speech without depriving them of access to important environmental sounds.
Background London has outperformed smaller towns and rural areas in terms of life expectancy increase. Our aim was to investigate life expectancy change at very-small-area level, and its relationship with house prices and their change. Methods We performed a hyper-resolution spatiotemporal analysis from 2002 to 2019 for 4835 London Lower-layer Super Output Areas (LSOAs). We used population and death counts in a Bayesian hierarchical model to estimate age -and sex-specific death rates for each LSOA, converted to life expectancy at birth using life table methods. We used data from the Land Registry via the real estate website Rightmove (www.rightmove.co.uk), with information on property size, type and land tenure in a hierarchical model to estimate house prices at LSOA level. We used linear regressions to summarise how much life expectancy changed in relation to the combination of house prices in 2002 and their change from 2002 to 2019. We calculated the correlation between change in price and change in sociodemographic characteristics of the resident population of LSOAs and population turnover. Findings In 134 (2.8%) of London's LSOAs for women and 32 (0.7%) for men, life expectancy may have declined from 2002 to 2019, with a posterior probability of a decline >80% in 41 (0.8%, women) and 14 (0.3%, men) LSOAs. The life expectancy increase in other LSOAs ranged from <2 years in 537 (11.1%) LSOAs for women and 214 (4.4%) for men to >10 years in 220 (4.6%) for women and 211 (4.4%) for men. The 2.5th-97.5th-percentile life expectancy difference across LSOAs increased from 11.1 (10.7-11.5) years in 2002 to 19.1 (18.4-19.7) years for women in 2019, and from 11.6 (11.3-12.0) years to 17.2 (16.7-17.8) years for men. In the 20% (men) and 30% (women) of LSOAs where house prices had been lowest in 2002, mainly in east and outer west London, life expectancy increased only in proportion to the rise in house prices. In contrast, in the 30% (men) and 60% (women) most expensive LSOAs in 2002, life expectancy increased solely independently of price change. Except for the 20% of LSOAs that had been most expensive in 2002, LSOAs with larger house price increases experienced larger growth in their population, especially among people of working ages (30-69 years), had a larger share of households who had not lived there in 2002, and improved their rankings in education, poverty and employment. Interpretation Large gains in area life expectancy in London occurred either where house prices were already high, or in areas where house prices grew the most. In the latter group, the increases in life expectancy may be driven, in part, by changing population demographics. Health 2023;27: Published January https://doi.org/10. 1016/j.lanepe.2022. 100580
Recent years have brought considerable advances to our ability to increase intelligibility through deep-learning-based noise reduction, especially for hearing-impaired (HI) listeners. In this study, intelligibility improvements resulting from a current algorithm are assessed. These benefits are compared to those resulting from the initial demonstration of deep-learning-based noise reduction for HI listeners ten years ago in Healy, Yoho, Wang, and Wang [(2013). J. Acoust. Soc. Am. 134, 3029-3038]. The stimuli and procedures were broadly similar across studies. However, whereas the initial study involved highly matched training and test conditions, as well as non-causal operation, preventing its ability to operate in the real world, the current attentive recurrent network employed different noise types, talkers, and speech corpora for training versus test, as required for generalization, and it was fully causal, as required for real-time operation. Significant intelligibility benefit was observed in every condition, which averaged 51% points across conditions for HI listeners. Further, benefit was comparable to that obtained in the initial demonstration, despite the considerable additional demands placed on the current algorithm. The retention of large benefit despite the systematic removal of various constraints as required for real-world operation reflects the substantial advances made to deep-learning-based noise reduction.
Environmental sound recognition is an essential part of the human auditory experience that not only provides a sense of connection to one’s surroundings but also forecasts potential nearby safety hazards. Unfortunately, important environmental sounds can be rendered inaudible or otherwise unrecognizable by modern noise-reduction technology, leading to reduced environmental sound recognition. What is needed is a system that simultaneously provides listeners with access to audible, recognizable environmental sounds and intelligible speech. Many modern noise-reduction systems rely on some form of time-frequency masking, such as the ideal ratio mask. Restricting the output range of this mask by limiting the maximum allowable attenuation of any given time-frequency unit results in a mask that preserves environmental sounds to a certain extent while enhancing speech. In an experiment, subjects with hearing impairment and normal hearing listened to mixtures of sound + speech that had been processed by time-frequency masks with various levels of maximum attenuation, resulting in different amounts of environmental sound preservation. In a dual-task paradigm, environmental sound recognition and speech intelligibility were measured, and it was found that both types of listeners simultaneously attain high levels of performance on both tasks when the attenuation of time-frequency units is limited to 23 dB. [Work supported by NIH F32DC019314, NIH R01DC015521, and The OSU Graduate School.]
BackgroundHigh-resolution data for how mortality and longevity have changed in England, UK are scarce. We aimed to estimate trends from 2002 to 2019 in life expectancy and probabilities of death at different ages for all 6791 middle-layer super output areas (MSOAs) in England.MethodsWe performed a high-resolution spatiotemporal analysis of civil registration data from the UK Small Area Health Statistics Unit research database using de-identified data for all deaths in England from 2002 to 2019, with information on age, sex, and MSOA of residence, and population counts by age, sex, and MSOA. We used a Bayesian hierarchical model to obtain estimates of age-specific death rates by sharing information across age groups, MSOAs, and years. We used life table methods to calculate life expectancy at birth and probabilities of death in different ages by sex and MSOA.FindingsIn 2002–06 and 2006–10, all but a few (0–1%) MSOAs had a life expectancy increase for female and male sexes. In 2010–14, female life expectancy decreased in 351 (5·2%) of 6791 MSOAs. By 2014–19, the number of MSOAs with declining life expectancy was 1270 (18·7%) for women and 784 (11·5%) for men. The life expectancy increase from 2002 to 2019 was smaller in MSOAs where life expectancy had been lower in 2002 (mostly northern urban MSOAs), and larger in MSOAs where life expectancy had been higher in 2002 (mostly MSOAs in and around London). As a result of these trends, the gap between the first and 99th percentiles of MSOA life expectancy for women increased from 10·7 years (95% credible interval 10·4–10·9) in 2002 to reach 14·2 years (13·9–14·5) in 2019, and for men increased from 11·5 years (11·3–11·7) in 2002 to 13·6 years (13·4–13·9) in 2019.InterpretationIn the decade before the COVID-19 pandemic, life expectancy declined in increasing numbers of communities in England. To ensure that this trend does not continue or worsen, there is a need for pro-equity economic and social policies, and greater investment in public health and health care throughout the entire country.FundingWellcome Trust, Imperial College London, Medical Research Council, Health Data Research UK, and National Institutes of Health Research.
The fundamental requirement for real-time operation of a speech-processing algorithm is causality-that it operate without utilizing future time frames. In the present study, the performance of a fully causal deep computational auditory scene analysis algorithm was assessed. Target sentences were isolated from complex interference consisting of an interfering talker and concurrent room reverberation. The talker- and corpus/channel-independent model used Dense-UNet and temporal convolutional networks and estimated both magnitude and phase of the target speech. It was found that mean algorithm benefit was significant in every condition. Mean benefit for hearing-impaired (HI) listeners across all conditions was 46.4 percentage points. The cost of converting the algorithm to causal processing was also assessed by comparing to a prior non-causal version. Intelligibility decrements for HI and normal-hearing listeners from non-causal to causal processing were present in most but not all conditions, and these decrements were statistically significant in half of the conditions tested-those representing the greater levels of complex interference. Although a cost associated with causal processing was present in most conditions, it may be considered modest relative to the overall level of benefit.
The practical efficacy of deep learning based speaker separation and/or dereverberation hinges on its ability to generalize to conditions not employed during neural network training. The current study was designed to assess the ability to generalize across extremely different training versus test environments. Training and testing were performed using different languages having no known common ancestry and correspondingly large linguistic differences-English for training and Mandarin for testing. Additional generalizations included untrained speech corpus/recording channel, target-to-interferer energy ratios, reverberation room impulse responses, and test talkers. A deep computational auditory scene analysis algorithm, employing complex time-frequency masking to estimate both magnitude and phase, was used to segregate two concurrent talkers and simultaneously remove large amounts of room reverberation to increase the intelligibility of a target talker. Significant intelligibility improvements were observed for the normal-hearing listeners in every condition. Benefit averaged 43.5% points across conditions and was comparable to that obtained when training and testing were performed both in English. Benefit is projected to be considerably larger for individuals with hearing impairment. It is concluded that a properly designed and trained deep speaker separation/dereverberation network can be capable of generalization across vastly different acoustic environments that include different languages.
Real-time operation is critical for noise reduction in hearing technology. The essential requirement of real-time operation is causality—that an algorithm does not use future time-frame information and, instead, completes its operation by the end of the current time frame. This requirement is extended currently through the concept of “effectively causal,” in which future time-frame information within the brief delay tolerance of the human speech-perception mechanism is used. Effectively causal deep learning was used to separate speech from background noise and improve intelligibility for hearing-impaired listeners. A single-microphone, gated convolutional recurrent network was used to perform complex spectral mapping. By estimating both the real and imaginary parts of the noise-free speech, both the magnitude and phase of the estimated noise-free speech were obtained. The deep neural network was trained using a large set of noises and tested using complex noises not employed during training. Significant algorithm benefit was observed in every condition, which was largest for those with the greatest hearing loss. Allowable delays across different communication settings are reviewed and assessed. The current work demonstrates that effectively causal deep learning can significantly improve intelligibility for one of the largest populations of need in challenging conditions involving untrained background noises.
Purpose This preliminary investigation compared effects of time compression on intelligibility for male versus female talkers. We hypothesized that time compression would have a greater effect for female talkers. Method Sentence materials from four talkers (two males) were time compressed, and original-speed and time-compressed speech materials were presented in a background of 12-talker babble to young adult listeners with normal hearing. Each talker/processing condition was heard by eight listeners (total N = 64). Generalized linear mixed-effects models were used to determine the effects of and interaction between processing condition and talker sex on keyword intelligibility. Additional post hoc analyses examined whether processing condition effects were related to talker vowel space and word frequency. Results For original-speed sentences, female and male talkers were essentially equally intelligible. Time compression reduced intelligibility for all talkers, but the effect was significantly greater for the female talkers. Supplementary analyses revealed that the effect of time compression depended on both talker vowel space and word frequency: The detrimental effect decreased significantly as word frequency and vowel space increased. Word frequency effects were also greater overall for talkers with larger vowel spaces. Conclusions While the small talker sample limits conclusions about the effects of talker sex, the secondary analyses suggest that intelligibility of talkers with larger vowel spaces is less susceptible to the negative effect of time compression, especially for high-frequency words.
Deep learning based speech separation or noise reduction needs to generalize to voices not encountered during training and to operate under multiple corruptions. The current study provides such a demonstration for hearing-impaired (HI) listeners. Sentence intelligibility was assessed under conditions of a single interfering talker and substantial amounts of room reverberation. A talker-independent deep computational auditory scene analysis (CASA) algorithm was employed, in which talkers were separated and dereverberated in each time frame (simultaneous grouping stage), then the separated frames were organized to form two streams (sequential grouping stage). The deep neural networks consisted of specialized convolutional neural networks, one based on U-Net and the other a temporal convolutional network. It was found that every HI (and normal-hearing, NH) listener received algorithm benefit in every condition. Benefit averaged across all conditions ranged from 52 to 76 percentage points for individual HI listeners and averaged 65 points. Further, processed HI intelligibility significantly exceeded unprocessed NH intelligibility. Although the current utterance-based model was not implemented as a real-time system, a perspective on this important issue is provided. It is concluded that deep CASA represents a powerful framework capable of producing large increases in HI intelligibility for potentially any two voices.
Recently, deep learning based speech segregation has been shown to improve human speech intelligibility in noisy environments. However, one important factor not yet considered is room reverberation, which characterizes typical daily environments. The combination of reverberation and background noise can severely degrade speech intelligibility for hearing-impaired (HI) listeners. In the current study, a deep learning based time-frequency masking algorithm was proposed to address both room reverberation and background noise. Specifically, a deep neural network was trained to estimate the ideal ratio mask, where anechoic-clean speech was considered as the desired signal. Intelligibility testing was conducted under reverberant-noisy conditions with reverberation time T 60 = 0.6 s, plus speech-shaped noise or babble noise at various signal-to-noise ratios. The experiments demonstrated that substantial speech intelligibility improvements were obtained for HI listeners. The algorithm was also somewhat beneficial for normal-hearing (NH) listeners. In addition, sentence intelligibility scores for HI listeners with algorithm processing approached or matched those of young-adult NH listeners without processing. The current study represents a step toward deploying deep learning algorithms to help the speech understanding of HI listeners in everyday conditions.
Abstract Nonnative (L2) English learners are often assumed to exhibit greater speech production variability than native (L1) speakers; however, support for this assumption is primarily limited to secondary observations rather than having been the specific focus of empirical investigations. The present study examined intra-speaker variability associated with L2 English learners’ tense and lax vowel productions to determine whether they showed comparable or greater intra-speaker variability than native English speakers. First and second formants of three tense/lax vowel pairs were measured, and Coefficient of Variation was calculated for 10 native speakers of American English and 30 nonnative speakers. The L2 speakers’ vowel formants were found to be native-like approximately half of the time. Whether their formants were native-like or not, however, they seldom showed greater intra-speaker variability than the L1 speakers.
Previous research has shown that English-speaking learners of Russian, even those with advanced proficiency, often have not acquired the contrast between palatalized and unpalatalized consonants, which is a central feature of the Russian consonant system. The present study examined whether training utilizing electropalatography (EPG) could help a group of Russian learners achieve more native-like productions of this contrast. Although not all subjects showed significant improvements, on average, the Russian learners showed an increase from pre- to post-training in the second formant frequency of vowels preceding palatalized consonants, thus enhancing their contrast between palatalized and unpalatalized consonants. To determine whether these acoustic differences were associated with increased identification accuracy, three native Russian speakers listened to all pre- and post-training productions. A modest increase in identification accuracy was observed. These results suggest that even short-term EPG training can be an effective intervention with adult L2 learners.
Older adults seeking hearing help often complain of particular difficulty understanding female voices. This contrasts with studies using young listeners with normal hearing in which female talkers have been found to be generally more intelligible than male talkers (e.g., Bradlow et al., 1996). Could some factor in addition to talker gender be causing older adults to have increased difficulty understanding female voices? Speech that has been time-compressed has been shown to be less intelligible than unprocessed speech (e.g., Gordon-Salant and Friedman, 2011), but few data exist to show whether an increased presentation rate causes an equal loss of intelligibility for male and female talkers. The present study will explore whether an increased playback rate has a greater negative effect on the intelligibility of speech produced by female versus male talkers. Subjects will listen to sentences produced by two female and two male speakers from the Utah Speaking Style Corpus presented at either their original rate or at an increased rate (1.5 times faster) and type out what they heard. The resulting data will show whether, for young normal-hearing listeners, an increased rate of speech affects the intelligibility of female talkers more than it affects the intelligibility of male talkers.
Purpose To present a novel case of Cryptococcus curvatus (CC) keratitis in a patient with Lattice dystrophy (LD). Methods A 57 year old lady with a family history of LD had a right re-graft for suspected recurrent lattice dystrophy. Her initial penetrating keratoplasty (PK) was in 2004 after which she kept good vision. 5 years later she developed a small ulcer caused by presumable Candida which healed completely following topical miconazole and short course of oral fluconazole. In 6 months she developed gutter, increasing astigmatism, recurrent erosions without inflammation and progressive peripheral opacities leading to a 2nd PK. Results Histology of the 2nd PK revealed amyloid deposits in both host and graft elements confirming recurrent LD. In addition the graft contained numerous yeasts within epithelium and stroma along scars from previous PK. The capsulated fungi were seen on PAS, grocott and mucicarmine stains. No inflammatory response was associated with the fungi. Microbiology confirmed these as (CC). On review of findings from initial infection, it seemed that Cryptococcus instead of Candida was indeed the causative agent of her previous ulcer. Interestingly, she had no history of local trauma prior to her first corneal ulceration and there was never evidence of infection elsewhere. Unfortunately, 2 years after the 2nd PK the graft developed new opacities that lead to a 3rd PK which revealed recurrent (CC) keratitis again without any inflammatory response. The left eye also had a corneal graft in 2009 given her underlying LD and mercifully has never been infected. Conclusion To our knowledge Cryptococcus curvatus keratitis was never described in English literature.In fact, it seems that this is the first description of ocular infection by this species.