Listening to nature soundscapes has been associated with aspects of psychological restoration, including increased cognitive performance and positive affect. Most recordings of these environments, however, are limited to the audible frequency range (i.e., 20 Hz-20 kHz), which leaves open the question of whether frequencies beyond the range of human hearing (ultrasound) influence psychological restoration. Ultrasound has been associated with increased alpha wave activity in the brain, which in turn has been associated with decreased stress and a state of relaxed alertness, suggesting that ultrasound might augment the restorative potential of soundscapes. To investigate this, we conducted a randomized controlled trial (n = 111) comparing different types of soundscapes (nature, recorded in a Borneo rainforest, or urban, recorded in a city in Ontario) with different bandwidths (audible-only or audible plus ultrasound) to quantify their effects on cognitive performance (backward digit span, n-back), self-reported mental fatigue, affect, and perceived restoration. During the sound intervention, participants additionally aesthetically rated the soundscapes. The results revealed no significant effects of ultrasound or environment on cognitive performance. However, ultrasound exposure increased mental fatigue. Additionally, there was no effect of ultrasound on either affect or perceived restoration. Notably, although nature soundscapes facilitated affective and perceived restoration, these effects were driven by participants' aesthetic ratings of the soundscapes.
To determine whether listeners use similar cues in absolute judgments and discrimination of sound locations, data were collected in four tasks used to assess sound-localization ability: a pointing task and three discrimination tasks (single interval, cued single interval, and two-alternative forced choice). Four stimuli were used: broadband noise, scrambled-spectrum noise (noise with levels jittered +/-20 dB in 1/3-octave bands), frozen scrambled-spectrum noise (stimuli with the scrambled noise spectrum fixed across intervals of the two-interval tasks), and a 750-Hz pure tone. Performance was better (higher d') in the discrimination tasks relative to the pointing task. However, performance in the pointing task was brought into better alignment with that of the discrimination tasks through an analysis that estimated and removed slowly fluctuating bias from the pointer responses. Discrimination performance was better for the broadband noise and frozen scrambled-spectrum noise than for the scrambled-spectrum noise and pure tone, indicating the influence of spectral cues in the former two. The results suggest that both localization (as in the present pointing task) and source-location discrimination may be based on the perceived locations of stimuli, but that spectral cues that are not associated with perceived location may influence responses for certain types of stimuli.
PURPOSE:The Connected Speech Test (CST) assesses an individual's ability to understand everyday contextualized running speech amidst competing background babble. To minimize accent effects on speech perception scores and reduce the noise floor of the original recordings, an updated version was developed by Saleh et al. (2020). The updated recordings feature a speaker with a General American accent to replace the Southern U.S. accent in the original test, and modern recording equipment was used to achieve a lower noise floor. The aim of this study was to collect normative data, characterizing performance on the updated CST. Self-reported speech intelligibility and listening effort ratings were collected to examine how subjective perceptions of the task vary across test conditions. METHOD:To evaluate normative performance on this updated test, 40 native English-speaking adults (36 females and four males) with normal hearing were recruited from The University of Western Ontario. Multitalker babble was presented at a fixed level, and speech was presented at fixed signal-to-babble ratios (SBRs) to participants in both co-located and separated loudspeaker conditions. At each SBR, participants were scored based on key words correctly identified. Subjective speech intelligibility and listening effort were evaluated using self-report scales. For each measure, data were fitted with transfer functions to characterize performance on the task. RESULTS:Participants demonstrated significantly better performance in the separated loudspeaker condition, indicating a spatial release from masking. For both conditions, increased SBR was associated with increased performance and subjective speech intelligibility, and decreased listening effort. CONCLUSION:The study provides normative data to characterize expected performance for the updated version of the CST.
Responding to an external stimulus takes ∼200 ms, but this can be shortened to as little as ∼120 ms with the additional presentation of a startling acoustic stimulus. This phenomenon is hypothesized to arise from the involuntary release of a prepared movement (a StartReact effect). However, a startling acoustic stimulus also expedites rapid mid-flight, reactive adjustments to unpredictably displaced targets which could not have been prepared in advance. We surmise that for such rapid visuomotor transformations, intersensory facilitation may occur between auditory signals arising from the startling acoustic stimulus and visual signals relayed along a fast subcortical network. To explore this, we examined how a startling acoustic stimulus shortens reaction times in a task that produces express visuomotor responses, which are brief bursts of muscle activity that arise from a fast tectoreticulospinal network. We measured express visuomotor responses on upper limb muscles in humans as they reached either toward or away from a stimulus in blocks of trials where movements could either be fully prepared or not, occasionally pairing stimulus presentation with a startling acoustic stimulus. The startling acoustic stimulus reliably produced larger but fixed-latency express visuomotor responses in a target-selective manner, and also shortened reaction times, which were equally short for prepared and unprepared movements. Our results provide insights into how a startling acoustic stimulus shortens the latency of reactive movements without full motor preparation. We propose that the reticular formation is the probable node for intersensory convergence during the most rapid transformations of vision into targeted reaching actions. KEY POINTS: A startling acoustic stimulus (SAS) shortens reaction times by releasing fully prepared motor programmes (the StartReact effect), but can also hasten responses in reflexive tasks without any movement preparation. Here we measure the effect of a SAS on reaction times and upper limb muscle recruitment in a reflexive reaching task, focusing on express visuomotor responses that are evoked by visual target presentation and demarcate activity along a subcortical tectoreticulospinal pathway. A SAS robustly increased the magnitude of express visuomotor responses without changing their timing, and this increase was tightly related to the subsequent reaction time even in the absence of motor preparation. Our results attest to intersensory facilitation within the tectoreticulospinal pathway, which provides the shortest pathway mediating visuomotor transformations for reaching. These results reconcile discrepant findings by emphasizing the importance of intersensory facilitation in SAS-induced hastening of reaction times in reflexive tasks.
Hartmann and Rakerd [JASA 94, 2083–2092 (1993)] measured listeners’ ability to identify the source location (front, overhead, or behind) of short (25-μs clicks) and long (880-ms noise bursts) broadband stimuli as a function of intensity. Error rates increased markedly with intensity for clicks—the Negative Level Effect (NLE)—but not for noise bursts. In addition to direct practical applications (e.g., stimulus choice in auditory displays), the result of this simple experiment has inspired a sub-field of its own including psychophysical, physiological, and computational modeling studies. The differing effect of level on vertical-plane localization of short- and long-duration stimuli has been replicated behaviorally in humans and cats and for additional stimulus types. Although degradation of spectral cue representations via saturation of auditory-nerve rate profiles provides the most straightforward explanation of the NLE, attempts to account for NLE-release for long-duration stimuli (via temporal integration, high-threshold fibers, lateral inhibition, or efferent effects) have so far proven inconclusive. Those studies have, however, revealed further complexities of spectral cue processing also requiring explanation. This presentation will survey the research inspired by Hartmann and Rakerd’s experiment and describe why the mechanism of NLE-release observed at long durations is still an open question.
Yost et al. (1974; Percept. Psychophys. 15, 483–487) measured the effect of psychophysical procedure to evaluate movement-related cues for auditory lateralization. A similar approach was used in the current work to assess sound-source localization in a 10-foot x 13-foot semi-anechoic sound field. Performance was measured in four listening conditions, a localization task utilizing a pointer response and three location-discrimination tasks (single-interval, same-different, or 2AFC). Four signals were used: a 750-Hz tone, broadband noise, or two versions of broadband noise with narrowband levels roved + /−20 dB in 1/3-octave bands between intervals, or frozen across intervals with the same rove. Stimuli were generated as phantom sources using tangent-law panning between + /−30° azimuth loudspeaker positions. For each signal type, d′ was lower in the pointer condition than in the three discrimination tasks. Across tasks, performance was best for broadband noise and noise frozen across intervals, relative to the pure-tone and fully roved signals. As suggested by Yost etal., task effects may relate to source movement cues present in only the two-interval discrimination conditions. Alternatively, slow fluctuations in bias may negatively affect only the pointer localization task, with results from single-interval discrimination offering some support for this interpretation.
Higher-order ambisonic rendering is an increasingly common soundscape reproduction technique that, in theory, enables presentation of virtual sounds at nearly any location in three-dimensional (3D) space, relatively unconstrained by the veridical locations of loudspeakers. We evaluated whether 3D sound reproduction through a ninth-order ambisonic loudspeaker array was indeed sufficiently accurate to probe the limits of human spatial perception. We first estimated minimum audible angles for human listeners for a variety of reference points on the horizontal plane. We demonstrated that the system can reproduce sounds with a spatial resolution that is equal or superior to the limits of human acuity, at least on the horizontal plane at the front. Importantly, the resolution of ambisonic reproduction appeared equivalent for regions of the system with high and low loudspeaker density. We also estimated localization cues at the same locations and showed that although localization cues for low-frequency components were well preserved, they were somewhat distorted for components above 4000 Hz. Finally, we provide evidence that these high-frequency distortions can serve elevation cues by human listeners. In summary, we showed that a ninth-order ambisonic system is able to render highly focal sound sources, and that high-frequency reproduction distortions may introduce unwanted localization cues.
Higher-order ambisonic rendering is an increasingly common acoustic field reproduction technique that enables presentation of virtual sounds at nearly any location in 3D space, relatively unconstrained by the veridical locations of loudspeakers. We evaluated whether 3D sound reproduction through a 9th order ambisonic system was sufficiently accurate to probe the limits of human spatial perception. In Experiment 1, we estimated minimum audible angles for human listeners at a variety of reference points on the horizontal plane. Our estimated values are similar to absolute thresholds obtained using single-channel free-field presentation for locations within ±51° on the horizontal plane, including values approaching 1° on the midline, regardless of speaker density. This demonstrates the adequacy of the AudioDome for studies of human auditory spatial perception, at least for displacement in the horizontal plane. In Experiment 2 we estimated monaural and binaural localization cues by presenting linear chirp sweeps (0-22050 Hz) at the horizontal reference points tested in Experiment 1 and recorded ear canal signals through a head and torso simulator. Although localization cues for low-frequency components were well preserved, they were distorted for components above 4000 Hz. In Experiment 3 we provide evidence that these high-frequency distortions are processed as cues to elevation. ### Competing Interest Statement The authors have declared no competing interest.
Auditory source separations of as little as 1 degree are detectable. However, presenting auditory stimuli at small separations presents technical challenges, with loudspeaker separation limited by transducer diameter. An alternative procedure is to utilize phantom sources, with the perceived position of a single source determined by the relative output levels of two spatially separated loudspeakers. Therefore, it is important to determine whether real and phantom sources can be localized with the same precision. In the present experiment, listeners localized real (individual) sources and phantom sources computed using a tangent-law model giving the same nominal azimuthal angles as the real sources. Listeners used a laser pointer to indicate perceived source location. Infrared cameras detected pointer position with responses stored in terms of azimuth. Signals were broadband or narrowband (300-700 Hz and 3800–4200 Hz) noise, 100 or 500 ms in duration. Generally, phantom sources were localized with less precision than real sources, and high-frequency signals were localized with less precision than broadband or low-frequency signals, with no effect of duration. Results show that phantom sources are localized with sufficient accuracy and precision to be useful in assessing auditory spatial acuity, but they are not localized with the same precision as real sources.
Foley et al. [JASA 151, 3189–3196 (2022)] have demonstrated that annoyance of auditory interface sounds can be reduced both by shortening the duration of upper harmonics and by applying percussive rather than flat amplitude envelopes. However, sounds in that study had a maximum frequency of 2400 Hz, which would likely affect localization based on spectral cues. In a setting with multiple interface devices (e.g., a multi-patient hospital ward), localization of interface sounds is a concern. We characterized normally hearing listeners' ability to localize a variety of candidate reduced-annoyance interface sounds (flat or percussive envelopes; durations between 360 and 1600 ms; all but one with uppermost harmonic limited to 2400 Hz) in quiet or at + 4-dB and −11-dB SNR in spatially diffuse multi-talker babble. Listeners stood at the center of a 360-degree loudspeaker array in a darkened anechoic chamber and used a head-pointing response to report the perceived location of each target. Decreasing SNR increased response variability and the rate of front/rear confusions. For all sounds with restricted bandwidth, front/rear confusions were frequent in the absence of head movements, but when head movements were initiated before target offset, confusions were substantially reduced. The results highlight the need to consider localizability when designing improved auditory interface sounds.
In daily life, our heads are in continual motion, but much of our knowledge of spatial hearing is based on experimental paradigms in which this behavior is discouraged or prevented. Head motion can benefit spatial hearing though the creation of dynamic acoustical cues, but also creates challenges such as the need to update head-centered spatial representations or the locus of spatial auditory attention. To utilize acoustic information generated by, or to compensate for, head movements, the auditory system must integrate self-motion information provided by other sensory systems. This presentation will review and contextualize a series of studies from the author's laboratory focused on the psychophysics of dynamic sound localization and the weighting of sensory information from vestibular, proprioceptive, and visual modalities in dynamic localization and maintenance of spatially selective auditory attention during head rotation. Notable findings include a velocity-independent ∼100-ms minimum stimulus duration for disambiguation of front/rear location in dynamic localization and the apparent dominance of vestibular information in the interpretation of dynamic localization cues and in attentional updating. The approach taken provides a step towards understanding the effects of naturalistic behavior on spatial hearing while maintaining significant experimental control and repeatability of stimuli and head movements.
Hearing aids are typically fitted using speech-based prescriptive formulae to make speech more intelligible. Individual preferences may vary from these prescriptions and may also vary with signal type. It is important to consider what motivates listener preferences and how those preferences can inform hearing aid processing so that assistive listening devices can best be tailored for hearing aid users. Therefore, this study explored preferred frequency-gain shaping relative to prescribed gain for speech and music samples. Preferred gain was determined for 22 listeners with mild sloping to moderately severe hearing loss relative to individually prescribed amplification while listening to samples of male speech, female speech, pop music, and classical music across low-, mid-, and high-frequency bands. Samples were amplified using a fast-acting compression hearing aid simulator. Preferences were determined using an adaptive paired comparison procedure. Listeners then rated speech and music samples processed using prescribed and preferred shaping across different sound quality descriptors. On average, low-frequency gain was significantly increased relative to the prescription for all stimuli and most substantially for pop and classical music. High-frequency gain was decreased significantly for pop music and male speech. Gain adjustments, particularly in the mid- and high-frequency bands, varied considerably between listeners. Music preferences were driven by changes in perceived fullness and sharpness, whereas speech preferences were driven by changes in perceived intelligibility and loudness. The results generally support the use of prescribed amplification to optimize speech intelligibility and alternative amplification for music listening for most listeners.
Purpose The original Connected Speech Test (CST; Cox et al., 1987) is a well-regarded and often utilized speech perception test. The aim of this study was to develop a new version of the CST using a neutral North American accent and to assess the use of this updated CST on participants with normal hearing. Method A female English speaker was recruited to read the original CST passages, which were recorded as the new CST stimuli. A study was designed to assess the newly recorded CST passages' equivalence and conduct normalization. The study included 19 Western University students (11 females and eight males) with normal hearing and with English as a first language. Results Raw scores for the 48 tested passages were converted to rationalized arcsine units, and average passage scores more than 1 rationalized arcsine unit standard deviation from the mean were excluded. The internal reliability of the 32 remaining passages was assessed, and the two-way random effects intraclass correlation was .944. Conclusion The aim of our study was to create new CST stimuli with a more general North American accent in order to minimize accent effects on the speech perception scores. The study resulted in 32 passages of equivalent difficulty for listeners with normal hearing.
The dynamic interaural time-difference (ITD) created by listener head rotation is a potent cue for front/rear sound source localization. In principle this cue could also assist in the segregation of simultaneously presented front and rear sources, particularly when robust high-frequency spectral cues are absent. If so, head rotation might provide dynamic spatial release from masking. We assessed this in a spatial auditory attention task in which multiple different equal-intensity sequences of four spoken digits, low-pass filtered at 1500 Hz, were presented simultaneously—the target sequence from 0 or 180 deg azimuth and distractors from lateral angles of ±22.5 and/or ± 45 deg relative to the target, but in the opposite hemisphere. On each trial, listeners either fixated towards 0 azimuth or oscillated their heads at ∼0.5 Hz with an amplitude of ∼±40 deg. Listeners reported the target sequence heard. In a majority of listeners, there was no benefit of head motion, and therefore no evidence of dynamic spatial release from masking for these stimuli. These results are consistent with those of Culling [J. Exp. Psych. 26, 1760–1769 (2000)], who found that in tone complexes with components smoothly changing in ITD, opposite movement direction for one component was not an effective segregation cue.
Objective: To assess the performance of an active transcutaneous implantable-bone conduction device (TI-BCD), and to evaluate the benefit of device digital signal processing (DSP) features in challenging listening environments. Design: Participants were tested at 1- and 3-month post-activation of the TI-BCD. At each session, aided and unaided phoneme perception was assessed using the Ling-6 test. Speech reception thresholds (SRTs) and quality ratings of speech and music samples were collected in noisy and reverberant environments, with and without the DSP features. Self-assessment of the device performance was obtained using the Abbreviated Profile of Hearing Aid Benefit (APHAB) questionnaire. Study sample: Six adults with conductive or mixed hearing loss. Results: Average SRTs were 2.9 and 12.3 dB in low and high reverberation environments, respectively, which improved to -1.7 and 8.7 dB, respectively with the DSP features. In addition, speech quality ratings improved by 23 points with the DSP features when averaged across all environmental conditions. Improvement scores on APHAB scales revealed a statistically significant aided benefit. Conclusions: Noise and reverberation significantly impacted speech recognition performance and perceived sound quality. DSP features (directional microphone processing and adaptive noise reduction) significantly enhanced subjects' performance in these challenging listening environments.
Objectives: We sought to investigate whether children referred to our audiology clinic with a complaint of listening difficulty, that is, suspected of auditory processing disorder (APD), have difficulties localizing sounds in noise and whether they have reduced benefit from spatial release from masking. Design: Forty-seven typically hearing children in the age range of 7 to 17 years took part in the study. Twenty-one typically developing (TD) children served as controls, and the other 26 children, referred to our audiology clinic with listening problems, were the study group: suspected APD (sAPD). The ability to localize a speech target (the word "baseball") was measured in quiet, broadband noise, and speech-babble in a hemi-anechoic chamber. Participants stood at the center of a loudspeaker array that delivered the target in a diffused noise-field created by presenting independent noise from four loudspeakers spaced 90 degrees apart starting at 45 degrees. In the noise conditions, the signal-to-noise ratio was varied between -12 and 0 dB in 6-dB steps by keeping the noise level constant at 66 dB SPL and varying the target level. Localization ability was indexed by two metrics, one assessing variability in lateral plane [lateral scatter (Lscat)] and the other accuracy in the front/back dimension [front/back percent correct (FBpc)]. Spatial release from masking (SRM) was measured using a modified version of the Hearing in Noise Test (HINT). In this HINT paradigm, speech targets were always presented from the loudspeaker at 0 degrees, and a single noise source was presented either at 0 degrees, 90 degrees, or 270 degrees at 65 dB A. The SRM was calculated as the difference between the 50% correct HINT speech reception threshold obtained when both speech and noise were collocated at 0 degrees and when the noise was presented at either 90 degrees or 270 degrees. Results: As expected, in both groups, localization in noise improved as a function of signal-to-noise ratio. Broadband noise caused significantly larger disruption in FBpc than in Lscat when compared with speech babble. There were, however, no group effects or group interactions, suggesting that the children in the sAPD group did not differ significantly from TD children in either localization metric (Lscat and FBpc). While a significant SRM was observed in both groups, there were no group effects or group interactions. Collectively, the data suggest that children in the sAPD group did not differ significantly from the TD group for either binaural measure investigated in the study. Conclusions: As is evident from a few poor performers, some children with listening difficulties may have difficulty in localizing sounds and may not benefit from spatial separation of speech and noise. However, the heterogeneity in APD and the variability in our data do not support the notion that localization is a global APD problem. Future studies that employ a case study design might provide more insights.
Objective: The purpose of this study was to examine developmental trends in spectral ripple discrimination (SRD) and to compare the performance of typically developing children to children with auditory processing disorder (APD). Study design: Cross-sectional study. Study sample: Fifteen children with APD, as well as 17 typically developing children and 14 adults reporting no listening or academic difficulties participated. Results: Typically developing children showed poor SRD thresholds compared to adults, indicating prolonged maturation of spectral shape recognition. Both typically developing children and APD children showed a maturational trend in SRD, but a General Linear Model fit to their thresholds showed that children with APD displayed SRD thresholds that were significantly poorer than those of typically developing children when controlling for age. This suggests that in APD children, SRD maturation lags behind typically developing children. Conclusion: Poor spectral ripple discrimination may explain some of the listening difficulties experienced by children with APD.
Significant differences have been found in hearing aid (HA) performance between laboratory and real world test environments. Virtual sound environments provide a degree of control and reproducibility which is lacking in real world testing but may require an impractical number of loudspeakers. We assessed the accuracy of a simulation approach in which sources’ direct sound is delivered by single loudspeakers while room acoustics are reproduced using low-order Ambisonics and a small number of loudspeakers. In a large office, we recorded binaural hearing aid output in response to sentence targets and babble noise presented at various levels and from various combinations of four loudspeakers surrounding a manikin. We measured the loudspeakers’ room impulse responses (IRs) using a 32-channel spherical microphone array (Eigenmike), and split the IRs into "direct sound" and "room sound" portions. In an anechoic chamber, the original acoustics were simulated using Ambisonics or discrete loudspeakers for each source’s direct portion and Ambisonics for the room portion. Ambisonic order and/or number of playback loudspeakers were also varied. HA output in the simulations was recorded using the manikin and assessed by comparing Hearing-Aid Speech Perception Index (HASPI) values computed on the simulation recordings with those made in the original room.
Purpose A growing body of evidence indicates that treatment of hearing loss by provision of hearing aids leads to improvements in auditory and visual working memory. The purpose of this study was to assess whether similar working memory benefits are observed following provision of cochlear implants (CIs). Method Fifteen adults with postlingually acquired severe bilateral sensorineural hearing loss completed the prospective longitudinal study. Participants were candidates for bilateral cochlear implantation with some aidable hearing in each ear. Implantation surgeries were carried out sequentially, approximately 1 year apart. Working memory was measured with the visual Reading Span Test (Daneman & Carpenter, 1980) at 5 time points: pre-operatively following a 6-month bilateral hearing aid trial, after 6 and 12 months of bimodal (CI plus contralateral hearing aid) listening experience following the 1st CI surgery and activation, and again after 6 and 12 months of bilateral CI listening experience following the 2nd CI surgery and activation. Results Compared to the preoperative baseline, CI listening experience yielded significant improvements in participants' ability to recall test words in the correct serial order after 12 months in the bimodal condition. Individual performance outcomes were variable, but almost all participants showed increases in task performance over the course of the study. Conclusions These results suggest that, similar to appropriate interventions with hearing aids, treatment of hearing loss with CIs can yield working memory benefits. A likely mechanism is the freeing of cognitive resources previously devoted to effortful listening.
The ability to segregate simultaneous speech streams is crucial for successful communication. Recent studies have demonstrated that participants can report 10%-20% more words spoken by naturally familiar (e.g., friends or spouses) than unfamiliar talkers in two-voice mixtures. This benefit is commensurate with one of the largest benefits to speech intelligibility currently known-that which is gained by spatially separating two talkers. However, because of differences in the methods of these previous studies, the relative benefits of spatial separation and voice familiarity are unclear. Here, the familiar-voice benefit and spatial release from masking are directly compared, and it is examined if and how these two cues interact with one another. Talkers were recorded while speaking sentences from a published closed-set "matrix" task, and then listeners were presented with three different sentences played simultaneously. Each target sentence was played at 0° azimuth, and two masker sentences were symmetrically separated about the target. On average, participants reported 10%-30% more words correctly when the target sentence was spoken in a familiar than unfamiliar voice (collapsed over spatial separation conditions); it was found that participants gain a similar benefit from a familiar target as when an unfamiliar voice is separated from two symmetrical maskers by approximately 15° azimuth.