Interaural time and level differences are crucial in sound localization, yet their contributions to sound source segregation and spatial selection remain underspecified. Here, participants completed a spatial auditory selective attention task while we measured hemodynamic activity in the prefrontal cortex and superior temporal gyrus using functional near-infrared spectroscopy. Participants listened to a target sound stream and a simultaneous spatially separated speech stream or white noise masker. Sound streams were spatialized with either 50 μs interaural time differences (ITDs), 500 μs ITDs, naturally occurring interaural level differences (ILDs) from a non-individualized head-related transfer function (HRTF), or broadband 10 dB ILDs. Behavioral results revealed a stronger effect of spatial cues when the masker was speech. Error patterns differed in the two difficult conditions, small ITDs and natural ILDs: Small ITDs produced lower hit rates, while naturally occurring ILDs produced higher false alarm rates. Small ITDs led to greater activity in prefrontal cortex and activity in superior temporal gyrus that was lateralized, greater in the hemisphere contralateral to attentional focus, consistent with previous reports. These results suggest that natural ILDs alone support source segregation even if they are insufficient to cause large shifts in perceived lateralization, explaining high false alarm rates (confusions between target and distractor words). In contrast, small ITDs alone may be insufficient to segregate competing sources, leading to low hit and false alarm rates. Together, these results reveal differences in how ITDs and ILDs contribute to auditory scene analysis and spatial attention.
Listeners face many challenges when trying to maintain attention to a target source in everyday settings; for instance, reverberation distorts acoustic cues and interruptions capture attention. However, little is known about how these challenges affect the ability to maintain selective attention. Here, we measured syllable recall accuracy and pupil dilation during a spatial selective attention task that was sometimes disrupted. Participants heard two competing, temporally interleaved syllable streams presented in pseudo-anechoic or reverberant environments. On randomly selected trials, a sudden interruption occurred mid-sequence. Compared to anechoic trials, reverberant performance was worse overall, and the interrupter disrupted performance. In uninterrupted trials, reverberation reduced peak pupil dilation both when it was consistent across all stimuli in a block and when it was randomized trial to trial, suggesting temporal smearing reduced clarity of the scene and the salience of events in the ongoing streams. Pupil dilations in response to interruptions indicated perceptual salience was strong across reverberant and anechoic conditions. Specifically, baseline pupil size before trials did not vary across room conditions, and mixing or blocking of trials (altering stimulus expectations) had no impact on pupillary responses. Together, these findings highlight that stimulus salience drives cognitive load more strongly than does task performance.
Task-irrelevant features can impact formation of auditory objects and influence the effectiveness of selective attention, including the buildup of attention over time. Using a previously established paradigm exploring the effects of random interruptions on spatial selective attention, this study explores how the task-irrelevant feature of talker identity impacts the buildup of spatial attention and whether it alters the impact of interruptions. Participants performed a sequence recall task in which participants were presented with two competing syllable sequences coming from different spatial directions and were asked to report the syllable sequence coming from the target direction. On half of the trials, an unpredictable, novel interrupting sound occurred, disrupting attentional focus. Two experiments explored how talker identity influenced performance, specifically, whether 1) making the two streams come from different talkers facilitates task performance and reduces the impact of interruption compared to when the streams are spoken by the same talker, and 2) talker discontinuity interferes with attention buildup and harms syllable recall performance compared to when the talker is the same from one syllable to the next. Our results showed that distinct talker features, though task-irrelevant in this spatial task, significantly improved syllable recall performance and reduced the impact of interrupters. Further, irrelevant talker discontinuities damaged attention buildup and reduced syllable recall performance. Post hoc analysis also revealed that repeating syllables in sequence substantially improved recall performance, which should be accounted for in future studies using similar paradigms.
Abstract Intelligible speech disrupts selective auditory attention more than an unintelligible stream. However, low-level acoustic features of intelligible speech are relatively similar to target speech, confounding results. While controlling acoustic similarity and limiting energetic masking, we examined how masker intelligibility affects behavior and electroencephalography (EEG). Normal hearing listeners detected color words within a target stream of randomly timed words while ignoring an ongoing masker. Maskers were either spoken by the same or a different talker and comprised either isochronous sequences of intelligible words or temporally scrambled versions. Scrambled maskers either lacked broadband energy changes over time (Experiment 1) or were amplitude modulated to have the same energy profiles as intelligible, isochronous maskers (Experiment 2). In both experiments, scrambled maskers yielded better performance than intelligible maskers. For intelligible maskers, performance was better for different compared to identical talkers. EEG responses paralleled behavior: target-evoked onset responses were larger for scrambled than for intelligible maskers, particularly for identical talkers. Later target recognition responses were larger for color than other target words but unaffected by masker type or talker. Even when low-level acoustic features were carefully matched, intelligible maskers impaired auditory attention and reduced target-evoked neural responses more than scrambled maskers, implicating early sensory filtering.
Past spatial auditory studies have used various approaches to manipulate perceived source location, including pure interaural time differences (ITDs), pure interaural level differences (ILDs), and more natural and realistic head-related transfer functions (HRTFs). Previously, we showed that the effectiveness of spatial attention is strongest for HRTF simulations and weakest for pure ITDs, but the mechanisms explaining such differences remain unclear. The prefrontal cortex (PFC) supports many higher-order cognitive functions, including working memory. Visual-biased PFC regions show greater activation during auditory tasks that require spatial processing than those that do not, suggesting that visual-biased PFC regions play an important role in auditory spatial cognition. To investigate the interaction between spatial cues and spatial processing, we conducted an fMRI study testing different auditory tasks (spatial, non-spatial, and passive) and different spatialization approaches (ITD, ILD, or HRTF). In spatial tasks, HRTFs yielded the best behavioral performance and strongest activation across PFC; ITDs yielded the lowest performance and weakest activation. Visual-biased PFC regions showed greater activation during spatial than nonspatial tasks. These results provide new insights into how spatial cues interact with PFC regions during auditory tasks.
Bilateral cochlear implant (BiCI) usage makes binaural benefits a possibility for implant users. Yet for BiCI users, limited access to interaural time difference (ITD) cues and reduced saliency of interaural level difference (ILD) cues restricts perceptual benefits of spatially separating a target from masker sounds. The present study explored whether magnifying ILD cues improves intelligibility of masked speech for BiCI listeners in a "symmetrical-masker" configuration, which ensures that neither ear benefits from a long-term positive target-to-masker ratio (TMR) due to naturally occurring ILD cues. ILD magnification estimates moment-to-moment ITDs in octave-wide frequency bands, and applies corresponding ILDs to the target-masker mixtures reaching the two ears at each specific time and frequency band. ILD magnification significantly improved intelligibility in two experiments: one with normal hearing (NH) listeners using vocoded stimuli and one with BiCI users. BiCI listeners showed no benefit of spatial separation between target and maskers with natural ILDs, even for the largest target-masker separation. Because ILD magnification relies on and manipulates only the mixed signals at each ear, the strategy never alters the monaural TMR in either ear at any time. Thus, the observed improvements to masked speech intelligibility come from binaural effects, likely from increased perceptual separation of the competing sources.
Often, spatial attention filters out “distractors,” reducing cortical responses to sounds from an unattended direction. Here, in a nominal “spatial selective attention” task, listeners let through all potential targets before evaluating location. Participants heard a rapid, randomly ordered sequence comprising the words “bash,” “dash,” and “gash.” Individual words were spatialized left or right using either interaural level differences (ILDs) or interaural time differences (ITDs) derived from head-related transfer functions (±5° or ±15°). Listeners were asked to respond to “bash” in the target direction while we measured evoked responses using electroencephalography and hemodynamic responses over prefrontal cortex (PFC) using functional near-infrared spectroscopy. Rather than observing larger onset (P1-N1) responses evoked by words in the target direction, “bash” evoked larger onset responses than “dash” or “gash,” irrespective of direction. All “bashes” elicited a late “target recognition” (P3) response, but it was significantly larger for words from the target direction. PFC responses were larger with ±15° ITDs/ILDs than ±5°, unlike previous studies showing an inverse relationship between difficulty and PFC activity. Rather than relying on spatial selective attention, our listeners filtered words based on phonetic content, then evaluated the location of potential targets, highlighting the influence of task demands on listening strategy.
Task-irrelevant features can impact formation of auditory objects and influence the effectiveness of selective attention, including the buildup of attention over time. Using a previously established paradigm exploring the effects of random interruptions on spatial selective attention, this study explores how the task-irrelevant feature of talker identity impacts the buildup of spatial attention and whether it alters the impact of interruptions. Participants performed a sequence recall task in which participants were presented with two competing syllable sequences coming from different spatial directions and were asked to report the syllable sequence coming from the target direction. On half of the trials, an unpredictable, novel interrupting sound occurred, disrupting attentional focus. Two experiments explored how talker influenced performance, specifically, whether 1) making the two streams come from different talkers facilitates task performance and reduces the impact of interruption compared to when the streams are spoken by the same talker, and 2) talker discontinuity interferes with attention buildup and harms syllable recall performance compared to when the talker is the same from one syllable to the next. Our results showed that distinct talker features, though task-irrelevant in this spatial task, significantly improved syllable recall performance and reduced the impact of interrupters. Further, irrelevant talker discontinuities damaged attention buildup and reduced syllable recall performance.
Top-down selective attention can be disrupted by salient interruptions, leading to decreased accuracy in recalling the contents of whatever source a listener was trying to focus on. This decrease in recall accuracy is referred to as the interruption effect. Here, we explored how the magnitude of the interruption effect varies with (1) novelty of interrupter, (2) working memory load, and (3) timing and predictability of an interrupter while listeners performed an auditory spatial attention task. Three studies, each including two to four online experiments, addressed these questions. In each experiment, listeners attended to a target stream of spoken syllables from a cued location, ignoring a similar distractor stream from the opposite hemifield. At the end of the trial, listeners reported back the attended target syllable sequence. Unpredictable interrupting sounds occurred in half of the trials, slightly before one of the target syllables. We found the following: (1) unfamiliar interrupters are equally disruptive as repeated interrupters, (2) increasing target sequence length increases the impact of the interruption on recall of the target syllable just after the interruption, and (3) unpredictable interrupters are more disruptive than predictable interrupters. These results support the hypothesis that salient interruptions disrupt both top-down attention and working memory maintenance.
Bilateral cochlear implant (BiCI) usage makes binaural benefits a possibility for implant users. Yet for BiCI users, limited access to interaural time difference (ITD) cues and reduced saliency of interaural level difference (ILD) cues restricts perceptual benefits of spatially separating a target from masker sounds. The present study explored whether magnifying ILD cues improves intelligibility of masked speech for BiCI listeners in a "symmetrical-masker" configuration, which ensures that neither ear benefits from a long-term positive target-to-masker ratio (TMR) due to naturally occurring ILD cues. ILD magnification estimates moment-to-moment ITDs in octave-wide frequency bands, and applies corresponding ILDs to the target-masker mixtures reaching the two ears at each specific time and frequency band. ILD magnification significantly improved intelligibility in two experiments: one with NH listeners using vocoded stimuli and one with BiCI users. BiCI listeners showed no benefit of spatial separation between target and maskers with natural ILDs, even for the largest target-masker separation. Because ILD magnification relies on and manipulates only the mixed signals at each ear, the strategy never alters the monaural TMR in either ear at any time. Thus, the observed improvements to masked speech intelligibility come from binaural effects, likely from increased perceptual separation of the competing sources.
Bilateral cochlear implant users struggle in spatial release from masking (SRM) tasks, likely due to restricted access to interaural time difference (ITD) cues. Instead, they must rely on interaural level difference (ILD) cues; however, our previous behavioral experiments suggest that magnification of ILDs can facilitate SRM. Here, we probed the neural mechanisms underlying the benefit of magnified ILDs. We tested 18 normal-hearing subjects in an anechoic chamber. Listeners heard target and masker sequences of object and color words from opposite (left and right) quarterfields and were asked to detect color words in the target stream. Both streams were spatialized using either 50 μS ITD (ITD50), 500 μS ITD (ITD500), a broadband 10 dB ILD (ILD10), or the largest naturally occurring frequency-specific ILD (70 degrees; ILD70n). We recorded task-elicited hemodynamic responses in dorsolateral prefrontal cortex (DLPFC) using functional Near-Infrared Spectroscopy. Subjects performed best in the ITD500 and ILD10 conditions. Hemodynamic response magnitudes were smaller for ITD50 than for all other conditions, consistent with frontal activity increasing when perceptual segregation is possible and spatial attention can be deployed successfully. These data show that magnified ILD cues enhance SRM and that the benefit ILDs confer arises because listeners can engage cognitive attentional processes.
Past research hints that realistic auditory spatial simulations not only “sound better,” but better engage brain mechanisms controlling spatial auditory attention. We sought to replicate this finding by comparing behavior and neural responses for simulations using: (1) individualized head-related transfer functions (HRTFs), (2) filters preserving individualized, frequency-specific interaural level differences (ILDs) but removing interaural time differences (ITDs), and (3) filters preserving ITDs but removing ILDs. Listeners were asked to report back a random stream of consonant-vowel syllables (tokens spoken by a male talker) from left or right while ignoring a different random, temporally inter-digitated consonant-vowel stream from the opposite direction. To extend previous findings, we tested a listener’s ability to ignore a salient distinct sound (a cat MEOW) that occurred in the middle of some randomly selected trials. ITD-only simulations led to worse performance, especially on cat-interrupted trials. Simultaneous electroencephalography showed that in ITD-only simulations, attention evoked no significant lateralized alpha oscillations (a signature of spatially directed attention) and the to-be-ignored cat elicited larger neural responses than for other simulations. These results highlight how different auditory virtual environment simulations can influence perceptual and neural outcomes and suggest that simulations including natural ILDs enhance a listener’s ability to focus spatial attention.
Most human auditory psychophysics research has historically been conducted in carefully controlled environments with calibrated audio equipment, and over potentially hours of repetitive testing with expert listeners. Here, we operationally define such conditions as having high ‘auditory hygiene’. From this perspective, conducting auditory psychophysical paradigms online presents a serious challenge, in that results may hinge on absolute sound presentation level, reliably estimated perceptual thresholds, low and controlled background noise levels, and sustained motivation and attention. We introduce a set of procedures that address these challenges and facilitate auditory hygiene for online auditory psychophysics. First, we establish a simple means of setting sound presentation levels. Across a set of four level-setting conditions conducted in person, we demonstrate the stability and robustness of this level setting procedure in open air and controlled settings. Second, we test participants’ tone-in-noise thresholds using widely adopted online experiment platforms and demonstrate that reliable threshold estimates can be derived online in approximately one minute of testing. Third, using these level and threshold setting procedures to establish participant-specific stimulus conditions, we show that an online implementation of the classic probe-signal paradigm can be used to demonstrate frequency-selective attention on an individual-participant basis, using a third of the trials used in recent in-lab experiments. Finally, we show how threshold and attentional measures relate to well-validated assays of online participants’ in-task motivation, fatigue, and confidence. This demonstrates the promise of online auditory psychophysics for addressing new auditory perception and neuroscience questions quickly, efficiently, and with more diverse samples. Code for the tests is publicly available through Pavlovia and Gorilla.
Non-traumatic noise exposure has been shown in animal models to impact the processing of envelope cues. However, evidence in human studies has been conflicting, possibly because the measures have not been specifically parameterized based on listeners' exposure profiles. The current study examined young dental-school students, whose exposure to high-frequency non-traumatic dental-drill noise during their course of study is systematic and precisely quantifiable. Twenty-five dental students and twenty-seven non-dental participants were recruited. The listeners were asked to recognize unvoiced sentences that were processed to contain only envelope cues useful for recognition and have been filtered to frequency regions inside or outside the dental noise spectrum. The sentences were presented either in quiet or in one of the noise maskers, including a steady-state noise, a 16-Hz or 32-Hz temporally modulated noise, or a spectrally modulated noise. The dental students showed no difference from the control group in demographic information, audiological screening outcomes, extended high-frequency thresholds, or unvoiced speech in quiet, but consistently performed more poorly for unvoiced speech recognition in modulated noise. The group difference in noise depended on the filtering conditions. The dental group's degraded performances were observed in temporally modulated noise for high-pass filtered condition only and in spectrally modulated noise for low-pass filtered condition only. The current findings provide the most direct evidence to date of a link between non-traumatic noise exposure and supra-threshold envelope processing issues in human listeners despite the normal audiological profiles.
Salient interruptions draw attention involuntarily. Here, we explored whether this effect depends on the spatial and temporal relationships between a target stream and interrupter. In a series of online experiments, listeners focused spatial attention on a target stream of spoken syllables in the presence of an otherwise identical distractor stream from the opposite hemifield. On some random trials, an interrupter (a cat “MEOW”) occurred. Experiment 1 established that the interrupter, which occurred randomly in 25% of the trials in the hemifield opposite the target, degraded target recall. Moreover, a majority of participants exhibited this degradation for the first target syllable, which finished before the interrupter began. Experiment 2 showed that the effect of an interrupter was similar whether it occurred in the opposite or the same hemifield as the target. Experiment 3 found that the interrupter degraded performance slightly if it occurred before the target stream began but had no effect if it began after the target stream ended. Experiment 4 showed decreased interruption effects when the interruption frequency increased (50% of the trials). These results demonstrate that a salient interrupter disrupts recall of a target stream, regardless of its direction, especially if it occurs during a target stream.
It is well-known that listeners with hearing impairment have trouble solving the cocktail party problem. Bilateral cochlear implant users have shown little or no perceptual abilities in this regard. This is due in large part to poor encoding of interaural time difference cues by these devices. Interaural level difference cues, which are well-encoded by bilateral cochlear implants, are not thought to facilitate benefit beyond improved signal-to-noise ratio benefits from head-shadow. An approach to mitigating this problem will be presented, that magnifies interaural level difference cues in real time to increase the perceptual distance between sound sources that are separated in the horizontal plane. The effects of ILD magnification on source localization and speech intelligibility in cocktail-party settings will be discussed.
Wolf spider females have the distinctive behaviour of carrying the egg sac attached to the spinnerets. If this egg sac is dislodged, the female will retrieve and reattach it, if possible. However, in some instances, a female may reattach another object of similar size or shape as a replacement for a lost egg sac. Here, I report the first example of an isopod, Armadillidium vulgare, being used as a substitute for a missing egg sac in Pardosa valensBarnes, 1959 from southeastern Arizona.
Autotomy occurs when an animal intentionally sacrifices an appendage to escape predation or free a limb. While immediately beneficial, loss of an appendage can lead to a variety of future costs. In many spiders, leg autotomy is common; previous work has sometimes demonstrated autotomy costs in some behaviors, while other times, no costs of autotomy occur. We examined frequency of autotomy in two riparian zone populations of the wolf spider Pardosa valens Barnes, 1959 and then used both mark–recapture work at these sites and laboratory predation trials to determine whether autotomy affected survival. Autotomy occurred in 31% of spiders; males were more likely than females to have a missing leg, but female reproductive status (carrying an egg sac or not) was unrelated to leg loss status. At both sites, survival over 1 week in the field was significantly higher for intact spiders than for spiders missing a leg, for both sexes and both female reproductive states. Additionally, when we paired intact and autotomized spiders with a predator (the larger wolf spider Rabidosa santrita (Chamberlin & Ivie, 1942)), autotomized spiders were more likely to be attacked and eaten. Our results suggest that leg autotomy in P. valens leads to a significant future survival cost, and we discuss how this cost may affect males and females differently.
OBJECTIVES:"Channel-linked" and "multi-band" front-end automatic gain control (AGC) were examined as alternatives to single-band, channel-unlinked AGC in simulated bilateral cochlear implant (CI) processing. In channel-linked AGC, the same gain control signal was applied to the input signals to both of the two CIs ("channels"). In multi-band AGC, gain control acted independently on each of a number of narrow frequency regions per channel.DESIGN:Speech intelligibility performance was measured with a single target (to the left, at -15 or -30°) and a single, symmetrically-opposed masker (to the right) at a signal-to-noise ratio (SNR) of -2 decibels. Binaural sentence intelligibility was measured as a function of whether channel linking was present and of the number of AGC bands. Analysis of variance was performed to assess condition effects on percent correct across the two spatial arrangements, both at a high and a low AGC threshold. Acoustic analysis was conducted to compare postcompressed better-ear SNR, interaural differences, and monaural within-band envelope levels across processing conditions.RESULTS:Analyses of variance indicated significant main effects of both channel linking and number of bands at low threshold, and of channel linking at high threshold. These improvements were accompanied by several acoustic changes. Linked AGC produced a more favorable better-ear SNR and better preserved broadband interaural level difference statistics, but did not reduce dynamic range as much as unlinked AGC. Multi-band AGC sometimes improved better-ear SNR statistics and always improved broadband interaural level difference statistics whenever the AGC channels were unlinked. Multi-band AGC produced output envelope levels that were higher than single-band AGC.CONCLUSIONS:These results favor strategies that incorporate channel-linked AGC and multi-band AGC for bilateral CIs. Linked AGC aids speech intelligibility in spatially separated speech, but reduces the degree to which dynamic range is compressed. Combining multi-band and channel-linked AGC offsets the potential impact of diminished dynamic range with linked AGC without sacrificing the intelligibility gains observed with linked AGC.