The prevalence of cross-lingual speech emotion recognition (SER) modeling has significantly increased due to its wide range of applications. Previous studies have primarily focused on technical strategies to adapt features, domains, and labels across languages, often overlooking the underlying commonalities between the languages. In this study, we address the language adaptation challenge in cross-lingual scenarios by incorporating vowel-phonetic constraints. Our approach is structured in two main parts. First, we investigate the vowel-phonetic commonalities associated with specific emotions across languages, particularly focusing on common vowels that prove to be valuable for SER modeling. Second, we utilize these identified common vowels as anchors to facilitate cross-lingual SER. To demonstrate the effectiveness of our approach, we conduct case studies using American English and Taiwanese Mandarin with two naturalistic emotional speech corpora: the MSP-Podcast and BIIC-Podcast corpora. The approach leverages evidence that certain vowels, including monophthongs and diphthongs, exhibit emotion-specific commonality across languages, serving as phonetic anchors to enhance unsupervised cross-lingual SER learning. The proposed model surpasses baseline performance, highlighting the importance of phonetic similarities for effective language adaptation in cross-lingual SER scenarios.
Modeling cross-lingual speech emotion recognition (SER) has become more prevalent because of its diverse applications. Existing studies have mostly focused on technical approaches that adapt the feature, domain, or label across languages, without considering in detail the similarities between the languages. This study focuses on domain adaptation in cross-lingual scenarios using phonetic constraints. This work is framed in a twofold manner. First, we analyze emotion-specific phonetic commonality across languages by identifying common vowels that are useful for SER modeling. Second, we leverage these common vowels as an anchoring mechanism to facilitate cross-lingual SER. We consider American English and Taiwanese Mandarin as a case study to demonstrate the potential of our approach. This work uses two in-the-wild natural emotional speech corpora: MSP-Podcast (American English), and BIIC-Podcast (Taiwanese Mandarin). The proposed unsupervised cross-lingual SER model using these phonetical anchors outperforms the baselines with a 58.64% of unweighted average recall (UAR).
Primary muscle tension dysphonia (pMTD) is a voice disorder of unknown etiology in which people have reduced volume and become easily fatigued. The tongue has biomechanical linkage to the laryngeal motor system that affects the voice, suggesting that tongue movement variability might be a marker for this disorder. Previous studies of healthy/disordered speech have reduced individual talker differences in physical vocal tract characteristics via normalization procedures. Here, we obtained tongue movement data for diadochokinetic sequences produced by healthy adult talkers and people with pMTD. We used an electromagnetic articulography system to test three healthy adult talkers and a participant with pMTD. We recorded vertical displacement of the tongue dorsum for sustained /a/ and 30 sec of rapidly repeated /pataka/. Displacement (mm) was scaled relative to each talker’s sustained vowel production, while movement time (msec) was normalized by dividing each /ka/ portion by the corresponding /pataka/ utterance length. The results suggest (1) normalization effectively reduced talker variability in tongue movement displacement and timing, and (2) the talker with pMTD showed markedly higher temporal variability, compared to the control talkers. We are currently testing this putative group difference using additional participants.
We describe Opti-Speech-VMT, a prototype tongue tracking system that uses electromagnetic articulography to permit visual feedback during oral movements.Opti-Speech-VMT is specialized for visuomotor tracking (VMT) experiments in which participants follow an oscillating virtual target in the oral cavity using a tongue sensor. The algorithms for linear, curved, and custom trajectories are outlined, and new functionality is briefly presented. Because latency can potentially affect accuracy in VMT tasks, we examined system latency at both the API and total framework levels. Using a video camera, we compared the movement of a sensor (placed on an experimenter’s finger) against an oscillating target displayed on a computer monitor. The average total latency was 87.3 ms, with 69.8 ms attributable to the API, and 17.4 ms to Opti-Speech-VMT. These results indicate minimal reduction in performance due to Opti-Speech-VMT, and suggest the importance of the EMA hardware and signal processing optimizations used.
PURPOSE:This study examined the extent to which prelingual cochlear implant (CI) users show a slowed speaking rate compared with typical-hearing (TH) talkers when repeating various speech stimuli and whether the slowed speech of CI users relates to their immediate verbal memory.METHOD:Participants included 10 prelingually deaf teenagers who received CIs before the age of 5 years and 10 age-matched TH teenagers. Participants repeated nonword syllable strings, word strings, and center-embedded sentences, with conditions balanced for syllable length and metrical structure. Participants' digit span forward and backward scores were collected to measure immediate verbal memory. Speaking rate data were analyzed using a mixed-design, repeated-measures analysis of variance, and the relationships between speaking rate and digit spans were evaluated by Pearson correlation.RESULTS:Participants with CIs spoke more slowly than their TH peers during the sentence repetition task but not in the nonword string and word string repetition tasks. For the CI group, significant correlations emerged between speaking rate and digit span scores (both forward and backward) for the sentence repetition task but not for the nonword string or word string repetition task. For the TH group, no significant correlations were found.CONCLUSIONS:The findings indicate a relation between slowed speech production, reduced immediate verbal memory, and diminished language capabilities of prelingual CI users, particularly for syntactic processing. These results support theories claiming that immediate memory, including components of a central executive, influences the speaking rate of these talkers. Implications for therapies designed to increase speech fluency in CI recipients are discussed.SUPPLEMENTAL MATERIAL:https://doi.org/10.23641/asha.21644795.
Purpose To better understand the role of tongue visibility in speech, this study compared the spatiotemporal patterns of silent versus audible speech for lingual consonants of American English. Kinematic data were obtained for articulatory features assumed to be visually salient, including tongue movement (anterior displacement and midsagittal area), lip aperture, and consonant duration. Method Electromagnetic articulography was used to measure 11 native speakers' productions of five consonants (/ɡ/, /w/, /ɹ/, /l/, and /ð/), selected to represent a continuum of tongue visibility. Nonword consonant–vowel syllables were elicited during a procedure designed to convey a dyadic communication environment. A method of kinematic-based consonant segmentation was developed for data processing, and results were analyzed with repeated-measures analysis of variance. Results Findings indicated increased consonant duration and lip aperture in the silent condition (vs. audible) for all five consonants. Tongue forward displacement was slightly greater in the silent condition, compared to audible, for all consonants except /ɡ/, the only consonant without a visible tongue component. In addition, the extent of tongue forwarding in silent speech corresponded with the degree of tongue visibility. Conclusion During silent speech, talkers increased their lip aperture and consonant duration and tended to shift their tongues forward for the most visible lingual consonants, suggesting that talkers may be aware at some level of the need to increase articulatory visibility of the tongue in the presence of an interlocutor during adverse speech conditions.
To understand how cochlear implant processing affects emotional prosody recognition in tonal languages, how normal-hearing (NH) and cochlear-implanted (CI) adults identify four emotions ("angry," "happy," "sad," and "neutral") in short, semantically neutral, Mandarin sentences are compared. Depending on hearing status (CI, NH), adults heard natural speech and/or noise-vocoded speech conditions (4-, 8-, and 16-spectral channels). Results suggest that Mandarin-speaking adults with CIs recognize emotions with similar accuracy as NH listeners attending to spectrally degraded (4-channel) vocoded speech. The accuracy noted for Mandarin appears to be lower than that described in previous studies of English.
In order to examine the coarticulation resistance hypothesis (CRH) this study investigated patterns of CV and C1C2V overlap in consonant clusters and matching singletons in Greek, including five places of articulation in C2V position (bilabials, labiodentals, interdentals, alveolars, and velars). Stimuli were recorded from three talkers producing eight repetitions of each item, embedded in a carrier phrase. Kinematic and acoustic data were acquired with an electromagnetic articulography (EMA) system. Temporal measures were obtained by determining timing lags for both singletons and clusters, while spatial measures were found by calculating the Euclidean Distance (ED) of the tongue back sensor from the C2V velocity peak to that of the following vowel, for both singletons and clusters. Logarithmic ED ratios for each stimulus triplet (e.g., “ba”-“la”-“bla”) were then computed to test whether C2V allows C1V to exert coarticulatory influence on the vowel. Preliminary findings from both the temporal and spatial analyses suggest pronounced differences in complex versus simple syllable organization as a function of C2V place of articulation. The extent to which these results support the CRH will be further discussed.
In order to investigate the articulatory processes involved in producing Japanese /r/, we obtained speech recordings for native talkers of standard Japanese using an electromagnetic articulography (EMA) system. Each talker produced repetitions of /r/ in a carrier phrase designed to contrast syllable (CV and VCV VCV) and vowel (/a/, /i/, /u/, /e/, and /o/) contexts. Kinematic recordings were made using tongue (tip, TT; dorsum, TD; body, TB; left lateral, TLL; and right lateral, TRL) and lower lip/jaw (LL) sensors. We measured TT vertical displacement, TT duration at maximum position, and tongue blade width for the consonant gestures. In a perceptual experiment, American English listeners decided whether these consonants consisted of `l,' `r,' or `d.' The kinematic results indicate Japanese talkers produced CV consonants with greater stricture and longer closures than consonants in intervocalic positions. CV productions also had narrower tongue blade widths than VCV VCV productions, especially in /i/ and /u/ contexts. The data were modeled with Dirichlet regression in order to determine how strongly tongue width and context (syllable and vowel) factors predict listeners' judgments. The results showed a significant fit for `r' judgments, with the tongue width fit successively increased by the addition of syllable and vowel context information.
Theories of speech production aim to explain how talkers express abstract linguistic forms as audible events that are intelligible to both speaker and listener. The relationship among planned units of speech, their articulatory implementation, and their acoustic consequences is thus a key issue in speech research. The work reported here is part of a larger project designed to investigate the effects of visual acoustic and visual articulatory feedback on second language (L2) learners’ production and perception of non-native speech sounds. L2 talkers from a variety of language backgrounds practiced producing an English vowel, /æ/, while receiving visual feedback on either (1) first and second formant frequencies, provided by a real-time spectrographic display, or (2) tongue back position, shown using a talker-driven tongue avatar. Kinematic data were recorded using an electromagnetic articulograph (EMA) system that tracked tongue midline and lateral movement during vowel productions. Pronunciation accuracy was analyzed by calculating acoustic and kinematic Mahalanobis distances between L2 productions and target (native talker) exemplars. Initial analyses of a single subject’s data showed that both types of visual feedback training improved pronunciation, suggesting that both acoustic and articulatory information are recruited during vowel production.
In a previous study, we asked healthy adult speakers to produce the word head under noise-masked (visual only) conditions and while watching videos of a 3D tongue avatar that gradually morphed from producing head to had. Results indicated that during the visual mismatch phases all participants entrained to the visually presented word, head, without being aware that their vowel quality had changed. Here, we explore whether similar effects occur for individuals with presumed sensorineural processing disorders, patients with Parkinson's disease (PD). We also examine the effects of PD treatment on this entrainment behavior. Participants were 14 individuals with PD, with eight in ongoing speech/language therapy, and six reporting no recent therapy. Participants heard pink noise over headphones and produced the word head under four viewing conditions: First, while viewing repetitions of head (baseline); next, during "morphed" videos shifting gradually from head to had (ramp); then videos of had (maximum hold); and finally videos of head (after effects). Analysis with a linear mixed effects model indicated a significant F1 difference between baseline and maximum hold phases for the productions of the treated PD group, but not for the untreated group. Implications for the causes and treatment of PD speech disorders are discussed.
Talkers of tonal language, such as Mandarin, use the acoustic cues of fundamental frequency (F0), amplitude, and duration to indicate lexical meaning as well as to express linguistic and emotional prosody. It has therefore been hypothesized that tonal language talkers have less prosodic “space” to signal emotional prosody using F0, compared to non-tonal language talkers, and that these differences in prosodic processing should be evident in speech perception and production tasks. In addition, for talkers of both tonal and non-tonal languages, speaking rate interacts with emotional expression, with “sad” mood generally expressed with slower speech, while “happy” and “angry” moods are marked with faster speech rates. Despite the overall importance of speaking rate in signaling particular emotional moods, few data exist for speakers of tonal languages, such as Mandarin, in which lexical tones are typically specified for length. These issues can be addressed by analyzing and modelling data for individuals with cochlear implants, electronic hearing systems that provide relatively good temporal resolution but poor spectral resolution. Findings from our laboratory will be used to address models of prosody perception and production.
OBJECTIVE:To examine the correlations between obstructive sleep apnea (OSA) and psychiatric disorders such as major depressive disorder (MDD), posttraumatic stress disorder (PTSD), or bipolar disorder (BD) and whether comorbid psychiatric diagnosis increases the risk of OSA.METHODS:This retrospective chart review study included all patients (N = 413) seen within a randomly selected 4-month period (August 2014 to November 2014) in a Veterans Administration outpatient psychiatry clinic. Patients were screened for symptoms of OSA with the STOP-BANG Questionnaire. Those with a positive screen were referred to the sleep clinic for confirmation of the diagnosis by polysomnogram (PSG). Frequency of PSG-confirmed OSA was correlated with different psychiatric disorders and comorbid psychiatric diagnoses.RESULTS:The study showed a high prevalence of OSA in psychiatric patients, particularly with MDD (37.8%) and PTSD (35.5%) and less so with BD (16.7%). Among all patients with OSA (n = 155), those with comorbid BD and PTSD had a significantly higher rate of OSA than those with BD alone (χ² = 7.28, P < .05) but not with PTSD alone. We also found a statistically significant higher incidence of OSA in male veterans with either MDD comorbid with PTSD (χ² = 3.869, P < .05) or BD comorbid with PTSD (χ² = 6.631, P < .05) compared with either mood disorder or PTSD alone.CONCLUSIONS:The study showed a high prevalence of OSA in psychiatric patients, particularly in those with PTSD and MDD and less so with BD. There was a statistically significant increase in the incidence of OSA in male veterans with either BD with comorbid PTSD or MDD with comorbid PTSD..
This newsletter is produced and distributed by the CENTER FOR RESEARCH IN LANGUAGE, a research center at the University of California, San Diego that unites the efforts of fields such as Cognitive Science, Linguistics, Psychology, Computer Science, Sociology, and Philosophy, all who share an interest in language. We feature papers related to language and cognition (1-10 pages, sent via email) and welcome response from friends and colleagues at UCSD as well as other institutions. Please forward correspondence to:
Pitch-dominant information is reduced in the spectrally-impoverished signal transmitted by cochlear implants (CIs), leading to potential difficulties in perceiving voice emotion. However, this evidence comes from non-tonal languages such as English, in which pitch information is not required for lexical meaning. In order to better understand how hearing impaired (HI) speakers of a tone language with cochlear implants (CIs) process emotional prosody, an experiment was conducted with healthy normal hearing (NH) Mandarin-speaking adults listening to synthetic stimuli designed to resemble CI input. Listeners heard short sentences from a read-speech database produced by professional actors. Stimuli were selected to express four emotions (“angry,” “happy,” “sad,” and “neutral”), under four conditions which varied the lexical tones of Mandarin. Listeners heard natural speech and three noise-vocoded speech conditions (4-, 8-, and 16-spectral channels) and made a four-alternative, forced-choice decision about the basic emotion underlying each sentence. Preliminary results indicate more accurate emotional prosody recognition for natural speech than for synthesized speech, with greater accuracy for higher channel stimuli than lower channel stimuli. The findings also suggest NH Mandarin-speaking listeners show lower overall vocal emotional prosody accuracy compared with previous studies of non-tonal languages (e.g., English).
Previous kinematic studies on consonant clusters suggest that many factors affect their articulatory timing, including speech rate, frequency of occurrence, and prosody. Other possible factors include the direction of articulatory movement (front-to-back, back-to-front), whether independent articulators (e.g., tongue and lips) or single articulators (e.g., tongue tip and tongue body) are involved, as well as the relative sonority of the consonantal segments. In this study, five talkers of Modern Greek produced the clusters /sp/, /ps/, /sk/, and /ks/ in the carrier phrase “Ipa_______pali” (I said________again). These clusters varied in articulatory direction, articulatory independence, and sonority patterns. We used four methods to determine the degree of articulatory overlap. Each method yielded durations (ms) between gestural landmarks, which were used to compute consonant overlap. The methods differed in the ways gestures were interpreted from the velocity signal (e.g., using extrema or 20% threshold) and how consonant overlap was computed. Preliminary results indicate that articulatory direction and relative sonority exerted the strongest effects on speakers’ productions, while articulator dependence/independence played no apparent role.
This study examined the contributions of the tongue tip (TT), tongue body (TB), and tongue lateral (TL) sensors in the electromagnetic articulography (EMA) measurement of American English alveolar consonants. Thirteen adults produced /ɹ/, /l/, /z/, and /d/ in /ɑCɑ/ syllables while being recorded with an EMA system. According to statistical analysis of sensor movement and the results of a machine classification experiment, the TT sensor contributed most to consonant differences, followed by TB. The TL sensor played a complementary role, particularly for distinguishing /z/.
Visuomotor pursuit tracking (VMT) tasks are used to study the movement of the articulators in order to better understand the processes involved in speech motor planning and execution. Because most of this research has focused on lip/jaw tracking of sinusoids visually presented on a screen, little is known about the tracking capabilities of the tongue, or for visual targets placed in the oral cavity. The present study used a novel technique for measuring tongue (and jaw) VMT based on a 3D electromagnetic articulography (EMA) system. Streaming EMA data from oral sensors were used to construct a real-time avatar of the tongue which was shown to subjects on a computer monitor. Subjects also viewed a (virtual) intra-oral spherical target that they were required to “hit” with the tongue tip sensor. The target was programmed to move in a sinusoidal direction and was varied in frequency (Hz), direction, and predictability. Preliminary results suggest that tracking accuracy is inversely related to target frequency and is higher for vertical and horizontal motion than for lateral motion. The role of target predictability will also be described.
To better understand audiovisual speech processing, we investigated the effects of viewing time-synchronized videos of a 3D tongue avatar on vowel production by healthy individuals. A group of 15 American English-speaking subjects heard pink noise over headphones and produced the word head under four viewing conditions: First, while viewing repetitions of the same vowel, /epsilon/ (baseline phase), then during a series of "morphed" videos shifting gradually from /epsilon/ to /ae/ (ramp phase), followed by repetitions of /ae/ (maximum hold phase), and finally repetitions of /epsilon/ (after effects phase). Results of a formant frequency (F1) analysis indicated that the visual mismatch phases (ramp and maximum hold) caused all subjects to align their productions to the visually-presented vowel, /ae/. No subjects reported being aware that their vowel quality had changed. We conclude that the visual moving tongue stimuli produced entrainment to the viewed vowel category, rather than adaptation in the opposite direction of the perturbation. Further experimentation is needed to determine whether these effects are due to inherent imitation behaviors or subjects' lack of agency with the tongue avatar.