Modulation statistics of "natural soundscapes" were estimated by calculating the modulation power spectrum (MPS) of a database of acoustic samples recorded in nine pristine terrestrial habitats for four moments of the day and two contrasting periods, differing in precipitation level. In particular, a set of statistics estimating low-pass quality, starriness, separability, asymmetry, modulation depth, and 1/ftα temporal-modulation power-law relationships were calculated from the MPS of the samples and related to geographical, meteorological factors and diel variations. MPS were found to be generally low-pass in shape in the modulation domain with most of their modulation power restricted to low temporal (<10-20 Hz) and spectral modulations (<0.5-1 cycle/kHz). Modulation statistics were distinguished between habitats irrespective of moment of the day and precipitation period with a greater role of modulation depth and starriness. Separability and starriness were found to be related to the global biodiversity decrease from tropical to polar regions, suggesting that the lack of joint high spectral and fast temporal modulations and MPS complexity are important features that may characterise "biophony," the collective sound produced by animals in a given habitat. These findings may help guide research on monitoring auditory behaviours and underlying mechanisms expected to exploit regularities of natural scenes.
IntroductionThis study explored human auditory capacity to evaluate the number of biological sound sources in natural soundscapes.MethodsThis was achieved by measuring the ability of human participants to judge the number of birds when listening to soundscapes generated by an engineering algorithm that controlled for bird abundance, species richness, level disparities between songs, bird behavior and background noise.Results and discussionAlthough often inaccurate, numerosity judgments were generally affected by the number of birds, demonstrating sub-optimal sensitivity to biodiversity in humans. Numerosity judgments were robust to low-intensity background sounds, and higher when between-species acoustic disparities were introduced, suggesting that grouping mechanisms contribute to biodiversity perception.
Psychophysical work suggested that humans discriminate diel changes in natural soundscapes based on spectral cues. The present study extended this investigation by assessing human discrimination for diel variations in natural soundscapes using a database of scenes recorded in three distinct habitats (two temperate forests and a tropical forest) marginally affected by human activity and processed by a multi-band noise vocoder to degrade selectively spectral and temporal cues. The results confirm that auditory sensitivity to diel variations is high and is only affected by degradations to spectral cues. They also indicate that the resolution of these spectral cues depends on habitat.
Following research on the auditory discrimination of natural soundscapes by human listeners by Apoux, Miller-Viacava, Ferrière, Dai, Krause, Sueur, and Lorenzi [(2023). J. Acoust. Soc. Am. 153, 2706–2723], the present study assessed the ability of normal-hearing listeners to identify and/or discriminate the time of day (Dawn, Midday, Dusk, and Night) for soundscapes recorded in temperate and tropical forests. Identification was assessed using a single-interval four-alternative forced-choice task. Discrimination was assessed using a three-interval oddity task, where the listener was required to pick the odd one out. While identification of time of day was generally poor, discrimination was significantly above chance, indicating that listeners could hear some differences associated with the time of day, but could not effectively use those differences to assign the correct verbal label. The contribution of level, fine spectral, temporal envelope and temporal fine structure cues to discrimination was assessed. None of these cues appeared to be critical for time of day discrimination. Instead, it is suggested that listeners base their decisions on gross spectral cues, including aspects such as those related to ambient-noise spectral distributions.
PURPOSE:The goal was to gain insight into the effects of hearing loss and hearing aids (HAs) on the perception of "natural sounds" and their importance in daily life by documenting the opinions of hearing care professionals (HCPs). METHOD:A questionnaire was designed where HCPs were asked to rate their patients' perception of natural sounds before and after receiving an HA. The online survey was conducted with 301 HCPs in France. RESULTS:According to respondents, the incidence of hearing natural sounds increased substantially at the end of the trial period relative to the start, especially for patients living in remote rural areas. Respondents also indicated an increase in listening accuracy, pleasantness, and importance of natural sounds at the end of the trial period. The majority of respondents indicated (a) that their patients find it important to listen to natural sounds and (b) that they are satisfied with their HAs in that respect. CONCLUSIONS:This study demonstrates the importance of natural sounds for most people with hearing loss. It highlights the effects of HA on patients' awareness of natural sounds and suggests that future research should consider the patients' place of residence.
A previous study reported that human listeners can discriminate natural soundscapes with a high sensitivity and that they appear to base their decision on cues related to biological sounds and environment acoustics when discriminating habitats [Apoux et al., J. Acoust. Soc. Am. 153, 2706 (2023)]. The present study replicated and extended these findings while focusing on the sensitivity to diel acoustical patterns. Twenty-five listeners were asked to discriminate but also identify natural soundscapes recorded during a one-year period in a temperate forest at four distinct and identifiable moments of the day (Dawn, Midday, Dusk, or Night). In one experiment, stimuli were restricted in the audio frequency domain or noise vocoded to, respectively, assess the contribution of spectral and temporal cues to moment of the day discrimination. Overall, identification performance was poor, but discrimination was well above chance, suggesting that human listeners can derive some global properties of soundscapes associated with the time of the day, but struggle to assign these properties to the correct moment. The filtering data indicated that temporal cues and spectral details are not critical for moment of the day discrimination. Like habitat discrimination, listeners may rely primarily on gross spectral cues.
OBJECTIVE:The ability to discriminate natural soundscapes recorded in a temperate terrestrial biome was measured in 15 hearing-impaired (HI) listeners with bilateral, mild to severe sensorineural hearing loss and 15 normal-hearing (NH) controls. DESIGN:Soundscape discrimination was measured using a three-interval oddity paradigm and the method of constant stimuli. On each trial, sequences of 2-second recordings varying the habitat, season and period of the day were presented diotically at a nominal SPL of 60 or 80 dB. RESULTS:Discrimination scores were above chance level for both groups, but they were poorer for HI than NH listeners. On average, the scores of HI listeners were relatively well accounted for by those of NH listeners tested with stimuli spectrally-shaped to match the frequency-dependent reduction in audibility of individual HI listeners. However, the scores of HI listeners were not significantly correlated with pure-tone audiometric thresholds and age. CONCLUSIONS:These results indicate that the ability to discriminate natural soundscapes associated with changes in habitat, season and period of the day is disrupted but it is not abolished. The deficits of the HI listeners are partly accounted for by reduced audibility. Supra-threshold auditory deficits and individual listening strategies may also explain differences between NH and HI listeners.
A previous modelling study reported that spectro-temporal cues perceptually relevant to humans provide enough information to accurately classify "natural soundscapes" recorded in four distinct temperate habitats of a biosphere reserve [Thoret, Varnet, Boubenec, Ferriere, Le Tourneau, Krause, and Lorenzi (2020). J. Acoust. Soc. Am. 147, 3260]. The goal of the present study was to assess this prediction for humans using 2 s samples taken from the same soundscape recordings. Thirty-one listeners were asked to discriminate these recordings based on differences in habitat, season, or period of the day using an oddity task. Listeners' performance was well above chance, demonstrating effective processing of these differences and suggesting a general high sensitivity for natural soundscape discrimination. This performance did not improve with training up to 10 h. Additional results obtained for habitat discrimination indicate that temporal cues play only a minor role; instead, listeners appear to base their decisions primarily on gross spectral cues related to biological sound sources and habitat acoustics. Convolutional neural networks were trained to perform a similar task using spectro-temporal cues extracted by an auditory model as input. The results are consistent with the idea that humans exclude the available temporal information when discriminating short samples of habitats, implying a form of a sub-optimality.
Research in hearing sciences has provided extensive knowledge about how the human auditory system processes speech and assists communication. In contrast, little is known about how this system processes "natural soundscapes," that is the complex arrangements of biological and geophysical sounds shaped by sound propagation through non-anthropogenic habitats [Grinfeder et al. (2022). Frontiers in Ecology and Evolution. 10: 894232]. This is surprising given that, for many species, the capacity to process natural soundscapes determines survival and reproduction through the ability to represent and monitor the immediate environment. Here we propose a framework to encourage research programmes in the field of "human auditory ecology," focusing on the study of human auditory perception of ecological processes at work in natural habitats. Based on large acoustic databases with high ecological validity, these programmes should investigate the extent to which this presumably ancestral monitoring function of the human auditory system is adapted to specific information conveyed by natural soundscapes, whether it operate throughout the life span or whether it emerges through individual learning or cultural transmission. Beyond fundamental knowledge of human hearing, these programmes should yield a better understanding of how normal-hearing and hearing-impaired listeners monitor rural and city green and blue spaces and benefit from them, and whether rehabilitation devices (hearing aids and cochlear implants) restore natural soundscape perception and emotional responses back to normal. Importantly, they should also reveal whether and how humans hear the rapid changes in the environment brought about by human activity.
PURPOSE A dual-task paradigm was implemented to investigate how noise type and sentence context may interact with age and hearing loss to impact word recall during speech recognition. METHOD Three noise types with varying degrees of temporal/spectrotemporal modulation were used: speech-shaped noise, speech-modulated noise, and three-talker babble. Participant groups included younger listeners with normal hearing (NH), older listeners with near-normal hearing, and older listeners with sensorineural hearing loss. An adaptive measure was used to establish the signal-to-noise ratio approximating 70% sentence recognition for each participant in each noise type. A word-recall task was then implemented while matching speech-recognition performance across noise types and participant groups. Random-intercept linear mixed-effects models were used to determine the effects of and interactions between noise type, sentence context, and participant group on word recall. RESULTS The results suggest that noise type does not significantly impact word recall when word-recognition performance is controlled. When data from noise types were pooled and compared with quiet, and recall was assessed: older listeners with near-normal hearing performed well when either quiet backgrounds or high sentence context (or both) were present, but older listeners with hearing loss performed well only when both quiet backgrounds and high sentence context were present. Younger listeners with NH were robust to the detrimental effects of noise and low context. CONCLUSIONS The general presence of noise has the potential to decrease word recall, but type of noise does not appear to significantly impact this observation when overall task difficulty is controlled. The presence of noise as well as deficits related to age and/or hearing loss appear to limit the availability of cognitive processing resources available for working memory during conversation in difficult listening environments. The conversation environments that impact these resources appear to differ depending on age and/or hearing status.
Speech recognition in noise may be viewed as a classification problem, in which the auditory system selects a subset of time-frequency (T-F) units to build a representation of the signal. The present study introduces local signal-to-noise ratio (SNR) importance functions, an approach to determine the contribution to speech intelligibility of T-F units as a function of their SNR. Consistent with previous work from this group on auditory-channel independence, it was hypothesized that T-F units with a relatively equal mixture of signal and noise (around 0 dB SNR) may be the most disruptive for intelligibility because they should be the most difficult to classify. This hypothesis was assessed by measuring the intelligibility of sentences presented in a competing-speech background while discarding a fixed proportion of units all with the same local SNR. The stimuli were either vocoded or left unprocessed to evaluate the influence of fine-structure cues on classification and also help better understand why cochlear implant (CI) users experience greater difficulties in noise. Results were consistent with the ambiguity hypothesis. More importantly, they suggest that discarding the noisiest units may not be the most advantageous way to improve speech intelligibility in noise by CI users.
Purpose: The goal of this study was to examine the role of carrier cues in sound source segregation and the possibility to enhance the intelligibility of 2 sentences presented simultaneously. Dual-carrier (DC) processing (Apoux, Youngdahl, Yoho, & Healy, 2015) was used to introduce synthetic carrier cues in vocoded speech. Method: Listeners with normal hearing heard sentences processed either with a DC or with a traditional single-carrier (SC) vocoder. One group was asked to repeat both sentences in a sentence pair (Experiment 1). The other group was asked to repeat only 1 sentence of the pair and was provided additional segregation cues involving onset asynchrony (Experiment 2). Results: Both experiments showed that not only is the "target" sentence more intelligible in DC compared with SC, but the "background" sentence intelligibility is equally enhanced. The participants did not benefit from the additional segregation cues. Conclusions: The data showed a clear benefit of using a distinct carrier to convey each sentence (i.e., DC processing). Accordingly, the poor speech intelligibility in noise typically observed with SC-vocoded speech may be partly attributed to the envelope of independent sound sources sharing the same carrier. Moreover, this work suggests that noise reduction may not be the only viable option to improve speech intelligibility in noise for users of cochlear implants. Alternative approaches aimed at enhancing sound source segregation such as DC processing may help to improve speech intelligibility while preserving and enhancing the background.
Band-importance functions created using the compound method [Apoux and Healy (2012). J. Acoust. Soc. Am. 132, 1078-1087] provide more detail than those generated using the ANSI technique, necessitating and allowing a re-examination of the influences of speech material and talker on the shape of the band-importance function. More specifically, the detailed functions may reflect, to a larger extent, acoustic idiosyncrasies of the individual talker's voice. Twenty-one band functions were created using standard speech materials and recordings by different talkers. The band-importance functions representing the same speech-material type produced by different talkers were found to be more similar to one another than functions representing the same talker producing different speech-material types. Thus, the primary finding was the relative strength of a speech-material effect and weakness of a talker effect. This speech-material effect extended to other materials in the same broad class (different sentence corpora) despite considerable differences in the specific materials. Characteristics of individual talkers' voices were not readily apparent in the functions, and the talker effect was restricted to more global aspects of talker (i.e., gender). Finally, the use of multiple talkers diminished any residual effect of the talker.
The degrading influence of noise on various critical bands of speech was assessed. A modified version of the compound method [Apoux and Healy (2012) J. Acoust. Soc. Am. 132, 1078-1087] was employed to establish this noise susceptibility for each speech band. Noise was added to the target speech band at various signal-to-noise ratios to determine the amount of noise required to reduce the contribution of that band by 50%. It was found that noise susceptibility is not equal across the speech spectrum, as is commonly assumed and incorporated into modern indexes. Instead, the signal-to-noise ratio required to equivalently impact various speech bands differed by as much as 13 dB. This noise susceptibility formed an irregular pattern across frequency, despite the use of multi-talker speech materials designed to reduce the potential influence of a particular talker's voice. But basic trends in the pattern of noise susceptibility across the spectrum emerged. Further, no systematic relationship was observed between noise susceptibility and speech band importance. It is argued here that susceptibility to noise and band importance are different phenomena, and that this distinction may be underappreciated in previous works.
Purpose: Psychoacoustic data indicate that infants and children are less likely than adults to focus on a spectral region containing an anticipated signal and are more susceptible to remote masking of a signal. These detection tasks suggest that infants and children, unlike adults, do not listen selectively. However, less is known about children's ability to listen selectively during speech recognition. Accordingly, the current study examines remote masking during speech recognition in children and adults. Method: Adults and 7- and 5-year-old children performed sentence recognition in the presence of various spectrally remote maskers. Intelligibility was determined for each remote-masker condition, and performance was compared across age groups. Results: It was found that speech recognition for 5-year-olds was reduced in the presence of spectrally remote noise, whereas the maskers had no effect on the 7-year-olds or adults. Maskers of different bandwidth and remoteness had similar effects. Conclusions: In accord with psychoacoustic data, young children do not appear to focus on a spectral region of interest and ignore other regions during speech recognition. This tendency may help account for their typically poorer speech perception in noise. This study also appears to capture an important developmental stage, during which a substantial refinement in spectral listening occurs.
Speech processed to replace the original temporal fine structure (TFS) with tones or noise carriers (vocoder processing) are generally less intelligible than natural or unprocessed speech, especially if a background noise is present. Moreover, the poorer intelligibility associated with vocoder processing is typically larger if the background fluctuates over time. This deleterious effect of vocoder processing has led to the postulate that TFS cues play a critical role when listening into the dips in the background. Recently, we have proposed a technique to reintroduce synthetic TFS cues in vocoded speech using one carrier for the target and one carrier for the background. This “dual-carrier” approach allows sentence intelligibility with a speech masker to reach a level almost comparable to that of natural speech. The goal of the present study was to investigate the extent to which dual-carrier processing generally improves speech recognition in various noises or if it truly compensates for the loss of TFS cues, therefore engendering masking release, as does natural speech. Results comparing masking release for three processing conditions (single-carrier, dual-carrier, and natural speech) in five backgrounds (speech-shaped noise, speech-modulated noise, and 1, 2, or 8 talkers) will be discussed.
Speech intelligibility generally increases as the difference in fundamental frequency (F0) between two simultaneous talkers increases. Vocoded speech shows no such effect. However, as previous work reported a contribution of envelope periodicity to masking release, an effect of F0 separation should be observed with vocoded speech. In a first experiment, the effect of F0 separation in vocoded speech was directly evaluated by presenting pairs of simultaneous sentences from the same male talker. The background sentence was time-reversed and its F0 was manipulated to range from 0 to 1 octave above that of the target sentence. Although limited, an effect was observed for large F0 separations. In a second experiment, target and background sentences were from different talkers and differed in envelope cutoff rates. The envelope low-pass cut-off of one signal was manipulated independently, ranging from 4 to 400 Hz, while the other was fixed at 400 Hz. As expected, decreasing the cut-off of the target resulted in a decrease in intelligibility. Inversely, decreasing the cut-off of the background improved intelligibility by as much as 50% points at 4 Hz. These findings show a potential contribution of envelope periodicity to speech-on-speech masking but only for very large F0 separations.
It has been long assumed that the corrupting influence of noise on speech is uniform across all frequencies, and that the contribution of each speech frequency decreases at the same rate as noise is added. This assumption is evident in numerous previous works, and is seen clearly in the Speech Intelligibility Index where the contribution of each speech band is scaled in the same way according to the amount of noise present. Here, it is argued that susceptibility to noise may differ across speech bands. To test this hypothesis, the compound method [F. Apoux and E.W. Healy (2012) 132, J. Acoust. Soc. Am.] was modified to evaluate the noise susceptibility of individual critical bands of speech. Noise was added to each target speech band and the signal-to-noise ratio (SNR) required to reduce the contribution of that band by half was estimated. It was found that noise susceptibility varies greatly across speech bands, and that the SNR required to similarly affect each band differed by as much as 12 dB. Interestingly, no obvious systematic relationship appeared to exist between band importance and noise susceptibility.
Normal-hearing adults can isolate frequency regions containing clean speech from surrounding regions containing noise. However, children have been shown to integrate information over a large number of auditory filters, and so they may not be able to isolate frequency regions as well. To assess children’s level of auditory filter independence, words were filtered into 30 contiguous 1-ERB-width bands. Speech was presented in every other band, for a total of 15 speech bands. Speech-shaped noise was then added to: all 30 contiguous bands, the 15 bands not containing speech (OFF), or the 15 bands containing speech (ON). Three age groups were tested: 10 adults, 9 older children (6–7 yr old), and 9 younger children (5 yr old). Consistent with previous findings involving consonant recognition, adults displayed large performance differences between off- and on-frequency noise (OFF vs. ON). The 6- to 7-yr-old group performed similarly to adults. In contrast, the 5-yr-old group displayed equivalent performance in the OFF and ON conditions. This indicates that, for these younger children, even noise that was mostly non-overlapping in frequency interfered with speech recognition as much as noise that was on frequency. [Work supported by NIH.]
Glimpsing models assert that speech recognition in noise relies on the ability to detect and combine time-frequency (T-F) units with “usable” speech information. A factor typically used to determine the usability of a unit is the local signal-to-noise ratio (SNR). Determining the SNR below which a unit stops contributing to overall intelligibility, however, is challenging and most approaches have consisted of varying the local SNR criterion. Here we implemented a new approach in which all the units exhibited the same local SNR. In order to prevent glimpsing within a T-F unit, their time and frequency span was lower than the subjects’ resolution. Stimuli were created by mixing two sentences at a given overall SNR and retaining only those units with the desired local SNR. This procedure was iterated at various overall SNRs with the same two sentences until we were able to reconstruct 50% of the mixture with only T-F units at the desired local SNR. Sentence recognition was measured for mixtures with a variable (VAR) or constant (CON) local SNR. Our results showed a systematic drop in performance in CON of up to 40% points. More importantly, they suggest a limited contribution of T-F units below -3 dB.