Speech perception is shaped by spectral context effects, where spectral properties of earlier (context) sounds influence perception of later (target) sounds. While studies of a single context effect abound, more recent efforts explore relationships between context effects. The magnitudes of spectral context effects are consistent when measured in the same frequency region but not across different frequency regions. However, these findings are combined across multiple separate samples, which raises questions about replicability. Here, relationships within and across frequency were tested in a single normal-hearing sample. Listeners categorized target words as “dot”, “got” (varying in third formant onset frequency), “sheet”, or “seat” (varying in frication noise spectrum). On each trial, the preceding context sentence was filtered to amplify key frequencies (shifting target word perception toward adjacent frequencies via spectral contrast effects [SCEs]) or attenuate key frequencies (shifting target word perception toward attenuated frequencies via auditory enhancement effects [EEs]). Context effect magnitudes were highly consistent across SCEs and EEs in the same frequency region (dot-got: r = 0.78; sheet-seat: r = 0.79), but neither SCEs nor EEs were reliable across frequency regions (r < 0.25). Results solidify important bounds on the network of context effects in speech perception. [Work supported by NIDCD.]
Listeners face many challenges when trying to maintain attention to a target source in everyday settings; for instance, reverberation distorts acoustic cues and interruptions capture attention. However, little is known about how these challenges affect the ability to maintain selective attention. Here, we measured syllable recall accuracy and pupil dilation during a spatial selective attention task that was sometimes disrupted. Participants heard two competing, temporally interleaved syllable streams presented in pseudo-anechoic or reverberant environments. On randomly selected trials, a sudden interruption occurred mid-sequence. Compared to anechoic trials, reverberant performance was worse overall, and the interrupter disrupted performance. In uninterrupted trials, reverberation reduced peak pupil dilation both when it was consistent across all stimuli in a block and when it was randomized trial to trial, suggesting temporal smearing reduced clarity of the scene and the salience of events in the ongoing streams. Pupil dilations in response to interruptions indicated perceptual salience was strong across reverberant and anechoic conditions. Specifically, baseline pupil size before trials did not vary across room conditions, and mixing or blocking of trials (altering stimulus expectations) had no impact on pupillary responses. Together, these findings highlight that stimulus salience drives cognitive load more strongly than does task performance.
Visual gestures, especially head movements and eyebrow raises, are time-locked to acoustic cues during the expression of spoken prosody. The present study examined the role these visual cues play in prosody perception, particularly for individuals with cochlear implants (CIs), who often experience challenges understanding prosody due to reduced access to pitch cues. A vocal mimicry paradigm was used to obtain granular, objective measures of prosody perception through acoustic analysis of mimicked sentences. Stimuli consisted of audio-visual recordings from one talker that captured naturally occurring variability in the expression of auditory and visual cues to word focus. Participants mimicked these natural recordings, as well as prosody-transplanted stimuli that allowed us to isolate the influence of auditory or visual prosody cues while holding the other modality at a neutral level (broad focus). Participants converted visual prosody cues into acoustic correlates of prosody in their mimicry, repeating acoustically unfocused words with higher F0 and intensity when those words were paired with video containing head and eyebrow gestures. This visual influence was stronger for participants with CIs than an age-matched group of typical-hearing listeners. CI participants who were less successful at acoustic mimicry tended to be more influenced by visual cues. These results indicate that CI listeners compensate for degraded auditory cues by integrating visual gestures into their perception of spoken prosody, potentially highlighting new targets for multisensory counseling or training.
Utterance-final cues like falling pitch contour allows listeners to prepare to respond promptly when their conversation partner finishes speaking. We hypothesized that this ability might be compromised in listeners with poorer pitch perception such as those using cochlear implants. During the experiment, listeners heard 128 lists of 3–4 color-shape phrases (e.g., “a green square and a blue circle…”) where the final phrase was manipulated to have a falling F0 (indicating completion) or a sustained elevated F0 (indicating continuation). Listeners were asked to repeat the last item they heard as soon as the talker finished speaking; thus, response time was the outcome measure for determining the effect of the falling F0 contour. Typical hearing listeners responded more quickly to a list ending with falling pitch, more slowly to noise-vocoded stimuli lacking pitch cues, and even more slowly to lists ending in high plateaued pitch (indicating the task’s sensitivity to perception of continuation and to perturbation in pitch perception). Ongoing data collection with CI listeners is expected to show a reduced within-listener difference in response time to lists ending with the two pitch contours, suggesting these listeners might delay responding in conversation even if the words were perceived correctly.
Auditory sensitivity for spectral peaks is magnified by precursor sounds that have spectrally contrastive peaks. This sensitivity affects the interpretation of speech sounds. For example, a precursor sentence with prominent energy in low-F3 frequencies (1700–2700 Hz) encourages perception of the high-F3 target “da”; precursor context with prominent energy in high-F3 frequencies (2700–3700 Hz) encourages perception of the low-F3 target “ga”. Adaptation of auditory nerve fibers was hypothesized to underlie these perceptual shifts. Here, neural responses to stimuli previously used in behavioral experiments were simulated using an auditory nerve model (Zilany et al., 2014). Center frequencies of high-spontaneous-rate fibers covered tonotopic regions of interest where precursors should affect contrastive components of the target sounds. Rates were averaged over a 50-ms window at target syllable onset and then normalized by the mean of the population response. Then, perception was simulated by categorizing the population response as “da” or “ga” on each trial. Neural responses reflected human listeners’ tendency to magnify spectral contrast: simulations estimated more high-F3 “da” responses following precursor sentences with amplified low-F3 frequencies, and more low-F3 “ga” responses following high-F3-amplified sentences. Results support the hypothesis that neural adaptation in the auditory periphery can explain SCEs in speech perception. [Work supported by NIDCD.]
Identification of speech sounds is influenced by spectral contrast effects (SCEs), the perceptual magnification of spectral differences between successive sounds. SCEs result in the categorization of a target sound being biased away from spectral properties in the preceding acoustic context. Given remarkable consistency in the magnitudes of these contrast effects within the same frequency region [Stilp (2019) J. Acoust. Soc. Am. 146(2), 1503-1517], it was hypothesized that they would also show stable relationships across different frequency regions. In this study, normal-hearing listeners' phoneme categorization and contrast effects were assessed where phonetic contrasts were driven by changes in low-frequency F1 ("big"-"beg" continuum), mid-frequency F3 ("dot"-"got"), or high-frequency frication spectrum regions ("sheet"-"seat"). On each trial, listeners heard a precursor sentence that was filtered to emphasize energy in the lower or higher range within one of these frequency regions, followed by a target word that hinged on the frequency region that was filtered. Results showed that SCEs influenced categorization in each frequency region, as expected. However, effect magnitudes were not correlated with each other across frequency regions within or across two participant samples. This clarifies perception-in-context on a broader scale as the influence of spectral contrast is independent across different frequency regions.
Perception of speech sounds is shaped by the context of the earlier sounds, including speaking rate. Often, acoustic differences between earlier and later sounds are perceptually magnified. In temporal contrast effects (TCEs), a fast-rate precursor sentence can cause the following target word to sound longer (e.g., longer-VOT “tier”), and a slow-rate precursor sentence can cause the target word to sound shorter (e.g., shorter-VOT “deer”). The novel contribution of this study is the exploration of TCEs across different context durations to examine their consistency on a granular level. On each trial, listeners heard one of three contexts (“a”, “the word”, “this time I want you to click on the word”) spoken at a fast or slow rate, then a target word to be identified as “deer” or “tier”. TCEs were observed at all context durations but were surprisingly not consistent in magnitude. While speaking rate forms an important context for speech sound recognition, the degree of its influence may not be reliable across different timescales. This reveals important bounds on how context effects relate to one another. [Work supported by NIDCD.]
Prosody, including speaking rate and intonation, is key for communicating trust and doubt. These attitudes are crucial in healthcare settings, as a provider’s trust in the patient encourages continuity of care and effective communication, while a provider’s doubt diminishes patient satisfaction and willingness to follow recommendations. The present study extends previous research on young typical hearing listeners, comparing perception of doubting and trusting prosody by young typical hearing listeners to older typical hearing listeners and listeners across the adult age span with cochlear implants (CIs). Using a continuous slider scale with endpoints of “doubt” and “trust,” listeners rated attitude in tokens of the word “okay,” in which F0 contour, vowel duration, and closure duration and intensity of /k/ were manipulated independently. Across groups, listeners rated longer duration as more doubting. Young typical hearing listeners consistently identified high-falling F0 as trusting and low-rising F0 as doubting. CI listeners followed the same pattern, but only for tokens with longer durations. Intonation had smaller effects on older listeners’ responses, with some giving responses opposite the expected patterns. These results suggest caution in relying solely on intonation to convey trust and in speaking slowly for intelligibility, as slowed speech is consistently interpreted as doubting.
Perceptual categorization of fricative sounds “sh” and “s” has been used to examine perception of a variety of factors such as coarticulation and talker variability. However, the methods for creating stimuli in these experiments are marked by inconsistency, lack of grounding in vocal tract models, and lack of plausible mechanism of auditory processing for the key contrastive properties. In this presentation, we review the advantages and disadvantages of previous stimulus generation methods and, ultimately, present a new method for synthesizing fricatives that is constrained by and shaped by the vocal tract to which it is linked. The key principle is that the fricative spectrum contains peaks at fixed frequencies with varying amplitudes rather than varying frequencies. Those peaks are fixed to align with real or imputed vocal tract resonances in the adjacent vowel, resulting in natural-sounding sequences. Most importantly, the parameters vary on a psychoacoustic scale that is consistent regardless of vocal tract size or vowel context, enabling the same set of parameters across a wide variety of talkers. Finally, we present effort to increase reproducibility in reporting via a standardized method to describe the parameters and illustrate them for publication.
Speech-language pathologists often use live-voice assessments to identify communication disorders in children. This study investigated the median F0, F0 range, and speaking rate of 12 examiners administering a commonly used sentence repetition assessment to 151 children aged from three to nine years. Results demonstrated that the acoustic characteristics differed between examiners and suggested that variation in these characteristics is associated with the child's age. Further research is needed to understand how variability in examiners' speech characteristics could potentially impact children's performance in language tasks.
OBJECTIVES:Seeing a talker's mouth improves speech intelligibility, particularly for listeners who use cochlear implants (CIs). However, the impacts of visual cues on listening effort for listeners with CIs remain poorly understood, as previous studies have focused on listeners with typical hearing (TH) and featured stimuli that do not invoke effortful cognitive speech perception challenges. This study directly compared the effort of perceiving audiovisual speech between listeners who use CIs and those with TH. Visual cues were hypothesized to yield more relief from listening effort in a cognitively challenging speech perception condition that required listeners to mentally repair a missing word in the auditory stimulus. Eye gaze was simultaneously measured to examine whether the tendency to look toward a talker's mouth would increase during these moments of uncertainty about the speech stimulus. DESIGN:Participants included listeners with CIs and an age-matched group of participants with typical age-adjusted hearing (N = 20 in both groups). The magnitude and time course of listening effort were evaluated using pupillometry. In half of the blocks, phonetic visual cues were severely degraded by selectively blurring the talker's mouth, which preserved stimulus luminance so visual conditions could be compared using pupillometry. Each block included a mixture of trials in which the sentence audio was intact, and trials in which a target word in the auditory stimulus was replaced by noise; the latter required participants to mentally reconstruct the target word upon repeating the sentence. Pupil and gaze data were analyzed using generalized additive mixed-effects models to identify the stretches of time during which effort or gaze strategy differed between conditions. RESULTS:Visual release from effort was greater and lasted longer for listeners with CIs compared with those with TH. Within the CI group, visual cues reduced effort to a greater extent when a missing word needed to be repaired than when the speech was intact. Seeing the talker's mouth also improved speech intelligibility for listeners with CIs, including reducing the number of incoherent verbal responses when repair was required. The two hearing groups deployed different gaze strategies when perceiving audiovisual speech. CI listeners looked more at the mouth overall, even when it was blurred, while TH listeners tended to increase looks to the mouth in the moment following a missing word in the auditory stimulus. CONCLUSIONS:Integrating visual cues from a talker's mouth not only improves speech intelligibility but also reduces listening effort, particularly for listeners with CIs. For listeners with CIs (but not those with TH), these visual benefits are magnified when a missed word needs to be mentally corrected-a common occurrence during everyday speech perception for individuals with hearing loss. These results underscore the importance of including participants with hearing loss in listening effort studies and suggest caution in assuming results from TH listeners will generalize to those with hearing loss. They also highlight the potential clinical relevance of visual speech information, for counseling patients and families and potentially for the development of audiovisual strategies to reduce listening effort.
Understanding vocal prosody is essential to successful communication. However, evaluations of speech recognition have relied heavily on word repetition-type tasks where success does not hinge on prosody perception, or where stimuli do not have enough prosodic variation to even test for this ability. Individuals who use cochlear implants (CIs) are at risk for poorer perception of prosody because of their limited access to pitch perception. This study used a multi-slider visual analog interface to measure perception of contrastive focus prosody in sentence-length stimuli by participants with CIs or with typical hearing (TH). Compared to TH listeners, CI users were more likely to misidentify which word had prosodic focus, as well as having weaker perception of prosodic focus, on average. Whereas TH listeners scaled their perceived strength of prosodic focus based on F0 and vowel intensity features, CI users scaled ratings in accordance with vowel intensity and vowel duration, with no relationship to F0. These results suggest that CI users are at risk of complete misperception of a talker's intended message, even in instances where there was no uncertainty about the words that were spoken.
Correct repetition of speech does not indicate the various effortful processes involved, such as mentally repairing words that were missed, ignoring words because they are already known, or targeting specific words containing key information. Across a series of studies, these abilities were tested using stimuli that gave listeners the opportunity to treat the same speech content with different strategies based on situational needs and cues. Momentary changes in listening effort and sensory gain were revealed by changes in pupil dilation and suppression of microsaccades linked to key stimulus landmarks. Typical-hearing listeners consistently exerted effort at specific times, such as the moment after missing a target word or when hearing repetition of a stimulus previously missed, while also reducing effort during speech that was already heard or irrelevant to the task. However, listeners with cochlear implants instead showed signatures of sustained effort that persisted after stimulus presentation, was not specific to key moments of information, and which did not decrease for irrelevant or redundant speech. These results suggest that listener-driven effort can be situationally dependent, and the ease and quickness of regulating effort is a dimension of success that would not be revealed by measures of word repetition accuracy.
The process of repairing misperceptions has been identified as a contributor to effortful listening in people who use cochlear implants (CIs). The current study was designed to examine the relative cost of repairing misperceptions at earlier or later parts of a sentence that contained contextual information that could be used to infer words both predictively and retroactively. Misperceptions were enforced at specific times by replacing single words with noise. Changes in pupil dilation were analyzed to track differences in the timing and duration of effort, comparing listeners with typical hearing (TH) or with CIs. Increases in pupil dilation were time-locked to the moment of the missing word, with longer-lasting increases when the missing word was earlier in the sentence. Compared to listeners with TH, CI listeners showed elevated pupil dilation for longer periods of time after listening, suggesting a lingering effect of effort after sentence offset. When needing to mentally repair missing words, CI listeners also made more mistakes on words elsewhere in the sentence, even though these words were not masked. Changes in effort based on the position of the missing word were not evident in basic measures like peak pupil dilation and only emerged when the full-time course was analyzed, suggesting the timing analysis adds new information to our understanding of listening effort. These results demonstrate that some mistakes are more costly than others and incur different levels of mental effort to resolve the mistake, underscoring the information lost when characterizing speech perception with simple measures like percent-correct scores.
Perceiving prosody is essential for understanding a talker’s intended meaning. For listeners with typical hearing (TH), the perception of voice pitch can withstand a high amount of background noise. However, listeners with cochlear implants (CI) lack access to harmonic pitch perception, putting them at risk for poor perception of prosody when noise is present, even if they appear to perceive pitch adequately in quiet. The current study tested the hypothesis that prosodic focus would be more heavily affected by noise for CI users compared to listeners with typical hearing. Stimuli were sentences with a contrastive focus on a specific word as if to correct prior information. The sentences were embedded in various levels of speech-shaped noise and presented with accompanying text to test only prosody rather than word identification ability. Participants used a visual analog scale to report the perceived degree of focus on words in each sentence. Results from TH listeners were essentially unaffected by noise regardless of the noise level (even down to −5 dB SNR). Conversely, CI users showed a high probability of misinterpreting which word was emphasized in sentences with any background noise, even with a favorable SNR of +10 dB, compared to their performance in quiet.
People with hearing impairment report listening fatigue as a major barrier to social communication, but most investigations examine momentary effort without establishing its connection to longer-term fatigue. In this study, listeners completed a 60-min sentence-repetition task with an easy condition (intact sentences) or an effortful condition (sentences that demanded mentally repairing missing words). Pre- versus post-listening tasks were used to measure fatigue, including (1) reaction times, (2) verbal creativity, and (3) subjective report. During listening, tonic changes in pupil dilation, verbal reaction times, and repetition accuracy were measured. We hypothesize that repeated moments of effortful listening result in slower decay in pupil size as well as reduced verbal creativity, and increased reaction times in the later parts of the testing block. Conversely, listeners who hear only easy intact sentences are expected to have equivalent performance before and after the testing block. A lack of differences across conditions would contradict the notion that fatigue is a linear product of repeated moments of elevated effort. The value of effects shown in this paradigm will be to demonstrate the impact of fatigue on other concurrent abilities beyond speech perception.
One of the most persistent problems in studying listening effort is distinguishing differences in effort without the confound of differences in speech intelligibility. We address this issue using a design where listening conditions contain similar acoustic content, but where comprehension effort is mediated by prosody. Stimuli were question-and-answer pairs where the answer corrects or affirms information from the question. In most trials, the answer had prosodic focus on the novel information. However, in select trials, the prosodic focus was incorrectly placed on already-known information. We hypothesized that inappropriate prosody would not affect intelligibility, but would elicit lingering listening effort, marked by elevated pupil dilation in the moments after stimulus presentation. We present results from 18 listeners with typical hearing, and discuss potential implications for listeners with cochlear implants who struggle to hear prosodic contours. This experimental design with novel discourse-level stimuli holds promise for distinguishing effortful auditory encoding from effortful comprehension as well as for exploring how the benefits of prosody in speech perception extend beyond improving accuracy, processing speed, and recall, to make listening easier.
Purpose: When words are misperceived, listeners can rely on later context to repair an auditory perception, at the cost of increased effort. The current study examines whether the effort to repair a missing word in a sentence is alleviated when the listener has some advance knowledge of what to expect in the sentence. Method: Sixteen adults with hearing aids and 17 with typical hearing heard sentences with a missing word that was followed by context sufficient to infer what the word was. They repeated the sentences with the missing words repaired. Sentences were preceded by visual text on the screen showing either “XXXX” (unprimed) or a priming word previewing the word that would be masked in the auditory signal. Along with intelligibility measures, pupillometry was used as an index of listening effort over the course of each trial to measure how priming influenced the effort needed to mentally repair a missing word. Results: When listeners were primed for the word that would need to be repaired in an upcoming sentence, listening effort was reduced, as indicated by pupil size returning more quickly toward baseline after the sentence was heard. Priming reduced the lingering cost of mental repair in both listener groups. For the group with hearing loss, priming also reduced the prevalence of errors on target words and words other than the target word in the sentence, suggesting that priming preserves the cognitive resources needed to process the whole sentence. Conclusion: These results suggest that listeners with typical hearing and with hearing loss can benefit from priming (advance cueing) during speech recognition, to accurately repair speech and to process the speech less effortfully.
Speech categorization is influenced by spectral contrast effects, or the perceptual magnification of spectral differences between successive sounds. Spectral contrast effects result in the categorization of a target sound being biased away from spectral properties in the preceding acoustic context. Because of the remarkable consistency in the magnitudes of these contrast effects within the same frequency region, we hypothesized that they would also show a stable relationship across different frequency regions. In this study, normal-hearing listeners' phoneme categorization and contrast effects were assessed where phonetic contrasts were driven by changes in low-frequency F1 (“big”-“beg”), mid-frequency F3 (“dot”-“got”), or high-frequency regions (frication spectrum in “sheet”-“seat”). On each trial, listeners heard a precursor sentence that was filtered to emphasize energy in the lower or higher range for one of these frequency regions, followed by a target word that hinged on that filtered frequency region. Spectral contrast effects influenced categorization in each frequency region, as expected. However, effect magnitudes were not reliably correlated with each other across frequency regions. This further clarifies perception-in-context on a broader scale, as using spectral context during speech categorization may be consistent within a single frequency region but not across different frequency regions.
Purpose: Listening can be effortful for a variety of reasons, including when a person misperceives a word in a sentence and then mentally repairs it using later context. The current study explored whether an external observer (in the role of a tester/clinician) could detect that effort by hearing the listener's voice as they repeat the sentence. Method: Stimuli were audio recordings of 13 adults with cochlear implants repeating sentences that were either intact or with a masked word that could be inferred/repaired using context (the latter of which were previously documented to elicit greater effort). Participants ( n = 171, including 28 audiologists) used a continuous visual analog scale to judge whether the talker heard one type of stimulus or the other. Participants were also surveyed for experiences related to detecting effort or confusion in a talker's voice. Results: Participant judges were unable to discern when the CI users were forced to effortfully infer words from context when repeating a sentence. Ratings indicated a general bias toward assuming the listener heard the original sentence correctly without any need for repair. Acoustic properties of the CI users' voices (hypothesized higher voice pitch and delayed verbal reaction time for stimuli involving repair) did not reliably correlate with ratings of uncertainty. There were also no statistically detectable advantages for audiologists or for people who reported experience or skill in discerning uncertainty in a talker's voice. Conclusions: Despite clear evidence that mental repair incurs extra effort, the process of mental repair gives no reliably perceptible signature in a talker's voice, even for audiologists and others who profess to have experience and skill in conversing with people who have hearing loss. Listening effort is at risk of going unnoticed by conversation partners and by audiologists who might underestimate a patient's effort when listening to speech. Supplemental Material: https://doi.org/10.23641/asha.28688012