Speech emotion recognition (SER) is a key technology to enable more natural human-machine communication. However, SER has long suffered from a lack of public large-scale labeled datasets. To circumvent this problem, we investigate how unsupervised representation learning on unlabeled datasets can benefit SER. We show that the contrastive predictive coding (CPC) method can learn salient representations from unlabeled datasets, which improves emotion recognition performance. In our experiments, this method achieved state-of-the-art concordance correlation coefficient (CCC) performance for all emotion primitives (activation, valence, and dominance) on IEMOCAP. Additionally, on the MSP-Podcast dataset, our method obtained considerable performance improvements compared to baselines.
Autism spectrum disorder (ASD) is characterized by deficits in social communication, and even children with ASD with preserved language are often perceived as socially awkward. We ask if linguistic patterns are associated with social perceptions of speakers. Twenty-one adolescents with ASD participated in conversations with an adult; each conversation was then rated for the social dimensions of likability, outgoingness, social skilfulness, responsiveness, and fluency. Conversations were analysed for responses to questions, pauses, and acoustic variables. Wide intonation ranges and more pauses within children's own conversational turn were predictors of more positive social ratings while failure to respond to one's conversational partner, faster syllable rate, and smaller quantity of speech were negative predictors of social perceptions.
Individuals with serious mental illness experience changes in their clinical states over time that are difficult to assess and that result in increased disease burden and care utilization. It is not known if features derived from speech can serve as a transdiagnostic marker of these clinical states. This study evaluates the feasibility of collecting speech samples from people with serious mental illness and explores the potential utility for tracking changes in clinical state over time. Patients (n = 47) were recruited from a community-based mental health clinic with diagnoses of bipolar disorder, major depressive disorder, schizophrenia or schizoaffective disorder. Patients used an interactive voice response system for at least 4 months to provide speech samples. Clinic providers (n = 13) reviewed responses and provided global assessment ratings. We computed features of speech and used machine learning to create models of outcome measures trained using either population data or an individual's own data over time. The system was feasible to use, recording 1101 phone calls and 117 hours of speech. Most (92%) of the patients agreed that it was easy to use. The individually-trained models demonstrated the highest correlation with provider ratings (rho = 0.78, p<0.001). Population-level models demonstrated statistically significant correlations with provider global assessment ratings (rho = 0.44, p<0.001), future provider ratings (rho = 0.33, p<0.05), BASIS-24 summary score, depression sub score, and self-harm sub score (rho = 0.25,0.25, and 0.28 respectively; p<0.05), and the SF-12 mental health sub score (rho = 0.25, p<0.05), but not with other BASIS-24 or SF-12 sub scores. This study brings together longitudinal collection of objective behavioral markers along with a transdiagnostic, personalized approach for tracking of mental health clinical state in a community-based clinical setting.
We propose a semi-supervised learning method to improve classification performance in scenarios with limited labeled data. We employ adaptation strategies such as entropy-filtering and self-training, and show that our method achieves up to 17.2% relative improvement in UAR for a multi-class problem. We apply our method to two different tasks: speaker clustering for adult-child interactions during autism assessment sessions, and a variation of the language identification task (LID). We show that in both tasks our method improves classification accuracy while using lesser training data than the baseline and demonstrate the robustness of our setup to the degree of adaptation by controlling the threshold on uncertainty of classification.
Negative emotional arousal during conflict has been related to negative outcomes in romantic relationships and degraded quality of family life. Despite its extensive study in psychology, it is still challenging to quantify emotional arousal in a meaningful way with objective indices beyond traditionally-used self-reported scores. We examine the association of acoustic and physiological arousal between dating couples through speech prosodic patterns and Electrodermal Activity (EDA) features. We use a dynamical systems model (DSM) approach to capture the interplay of arousal indices within and between people. The DSM parameters reflect the amount of self-regulation with respect to the acoustic and physiological cues within a person, the degree of cross-regulation between the two modalities, as well as the within-couple co-regulation. Our results through statistical analysis and classification experiments indicate a significant association between the estimated system parameters and the participants' selfreported relationship satisfaction measures. This is consistent with previous findings and can help towards better understanding regulation mechanisms and escalation effects of emotional arousal during couples' discussions.
Formally, the problem that we present is that of identifying the hidden attributes of the system that modulates the body's signals, uncovered through novel signal processing and machine learning on large-scale multimodal data (Figure 1). Signal processing is the keystone that supports this mapping from data to representations of behaviors and mental states. The pipeline first begins with raw signa...
Social anxiety is a prevalent condition affecting individuals to varying degrees. Research on autism spectrum disorder (ASD), a group of neurodevelopmental disorders marked by impairments in social communication, has found that social anxiety occurs more frequently in this population. Our study aims to further understand the multimodal manifestation of social stress for adolescents with ASD versus neurotypically developing (TD) peers. We investigate this through objective measures of speech behavior and physiology (mean heart rate) acquired during three tasks: a low-stress conversation, a medium-stress interview, and a high-stress presentation. Measurable differences are found to exist for speech behavior and heart rate in relation to task-induced stress. Additionally, we find the acoustic measures are particularly effective for distinguishing between diagnostic groups. Individuals with ASD produced higher prosodic variability, agreeing with previous reports. Moreover, the most informative features captured an individual's vocal changes between low and high social-stress, suggesting an interaction between vocal production and social stressors in ASD.
The mutual influence of participant behavior in a dyadic interaction has been studied for different modalities and quantified by computational models. In this paper, we consider the task of automatic recognition for children's speech, in the context of child-adult spoken interactions during interviews of children suspected to have been maltreated. Our long-term goal is to provide insights within this immensely important, sensitive domain through large-scale lexical and paralinguistic analysis. We demonstrate improvement in child speech recognition accuracy by conditioning on both the domain and the interlocutor's (adult) speech. Specifically, we use information from the automatic speech recognizer outputs of the adult's speech, for which we have more reliable estimates, to modify the recognition system of child's speech in an unsupervised manner. By learning first at session level, and then at the utterance level, we demonstrate an absolute improvement of upto 28% WER and 55% perplexity over the baseline results. We also report results of a parallel human speech recognition (HSR) experiment where annotators are asked to transcribe child's speech under two conditions: with and without contextual speech information. Demonstrated ASR improvements and the HSR experiment illustrate the importance of context in aiding child speech recognition, whether by humans or computers.
The study of speech pathology involves evaluation and treatment of speech production related disorders affecting phonation, fluency, intonation and aeromechanical components of respiration. Recently, speech pathology has garnered special interest amongst machine learning and signal processing (ML-SP) scientists. This growth in interest is led by advances in novel data collection technology, data science, speech processing and computational modeling. These in turn have enabled scientists in better understanding both the causes and effects of pathological speech conditions. In this paper, we review the application of machine learning and signal processing techniques to speech pathology and specifically focus on three different aspects. First, we list challenges such as controlling subjectivity in pathological speech assessments and patient variability in the application of ML-SP tools to the domain. Second, we discuss feature design methods and machine learning algorithms using a combination of domain knowledge and data driven methods. Finally, we present some case studies related to analysis of pathological speech and discuss their design.
Well-being and mental health are directly associated with relationship status particularly in the context of relatedness and support. A key factor in relationship functioning is emotional arousal. We examine the interplay between emotional arousal manifested through acoustic and physiological cues and its association to relationship satisfaction. We propose a dynamical systems model to infer the within- and across-modality as well as the between-partner relations. Our results suggest that increased emotional regulation is negatively associated with relationship satisfaction and indicate that the proposed system consists a viable framework for analyzing such multimodal interrelations within romantic partners.
The very earliest description by Kanner indicated that large head size was associated with autism spectrum disorder. Yet, since then, opinion has ranged from those who argue that head and brain size is a general feature of autism spectrum disorder (ASD) to those who posit that this association is an artifact. We selectively review the literature on macrocephaly and megalencephaly in ASD and come to the conclusion that head and brain enlargement is characteristic of only a subset of approximately 15% of males with ASD. It appears that this is much less common in females. Moreover, the widespread notion that early macrocephalyContents Abstract . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 171 10.1 Introduction. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 172 10.2 Studies of macrocephaly . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 172 10.3 Studies of megalencephaly . . . . . . . . . . . . . . . . . . . . . . . . . . . 173 10.4 Brain enlargement in a subsample of individuals with ASD. . 175 10.5 Sex differences . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 176 10.6 Body size. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 177 10.7 Normalization of brain size . . . . . . . . . . . . . . . . . . . . . . . . . . 178 10.8 Neurobiology of brain enlargement . . . . . . . . . . . . . . . . . . . . 178 10.9 Microcephaly . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 181 10.10 Conclusion. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 181 References . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 182and megalencephaly is followed by normalization of head/brain size is not supported by essential longitudinal studies. Given that an enlarged brain is a feature of one form of ASD, the critical remaining issues are what causes the abnormal brain enlargement and whether this has consequences on the clinical and cognitive outcomes of the affected individuals.
Machine learning (ML) provides novel opportunities for human behavior research and clinical translation, yet its application can have noted pitfalls (Bone et al., 2015). In this work, we fastidiously utilize ML to derive autism spectrum disorder (ASD) instrument algorithms in an attempt to improve upon widely used ASD screening and diagnostic tools. The data consisted of Autism Diagnostic Interview-Revised (ADI-R) and Social Responsiveness Scale (SRS) scores for 1,264 verbal individuals with ASD and 462 verbal individuals with non-ASD developmental or psychiatric disorders, split at age 10. Algorithms were created via a robust ML classifier, support vector machine, while targeting best-estimate clinical diagnosis of ASD versus non-ASD. Parameter settings were tuned in multiple levels of cross-validation. The created algorithms were more effective (higher performing) than the current algorithms, were tunable (sensitivity and specificity can be differentially weighted), and were more efficient (achieving near-peak performance with five or fewer codes). Results from ML-based fusion of ADI-R and SRS are reported. We present a screener algorithm for below (above) age 10 that reached 89.2% (86.7%) sensitivity and 59.0% (53.4%) specificity with only five behavioral codes. ML is useful for creating robust, customizable instrument algorithms. In a unique dataset comprised of controls with other difficulties, our findings highlight the limitations of current caregiver-report instruments and indicate possible avenues for improving ASD screening and diagnostic tools.
BACKGROUND:Machine learning (ML) provides novel opportunities for human behavior research and clinical translation, yet its application can have noted pitfalls (Bone et al., 2015). In this work, we fastidiously utilize ML to derive autism spectrum disorder (ASD) instrument algorithms in an attempt to improve upon widely used ASD screening and diagnostic tools. METHODS:The data consisted of Autism Diagnostic Interview-Revised (ADI-R) and Social Responsiveness Scale (SRS) scores for 1,264 verbal individuals with ASD and 462 verbal individuals with non-ASD developmental or psychiatric disorders, split at age 10. Algorithms were created via a robust ML classifier, support vector machine, while targeting best-estimate clinical diagnosis of ASD versus non-ASD. Parameter settings were tuned in multiple levels of cross-validation. RESULTS:The created algorithms were more effective (higher performing) than the current algorithms, were tunable (sensitivity and specificity can be differentially weighted), and were more efficient (achieving near-peak performance with five or fewer codes). Results from ML-based fusion of ADI-R and SRS are reported. We present a screener algorithm for below (above) age 10 that reached 89.2% (86.7%) sensitivity and 59.0% (53.4%) specificity with only five behavioral codes. CONCLUSIONS:ML is useful for creating robust, customizable instrument algorithms. In a unique dataset comprised of controls with other difficulties, our findings highlight the limitations of current caregiver-report instruments and indicate possible avenues for improving ASD screening and diagnostic tools.
Speech signal processing is being increasingly explored in health domains given both its centrality as a behavioral cue and the promise of robust, automated analysis of data at scale. We discuss general issues in health-related speech research. We further highlight two health applications we have undertaken: addiction counseling and autism spectrum disorder. Methods range from deep supervised learning to knowledge-based signal processing of highly subjective constructs of psychological states and traits. A unique aspect of the research has been to model both health care provider and patient behaviors jointly in clinical encounters, wherein any individual behavior cannot be considered in isolation given the inherent mutual influence.
Child engagement is defined as the interaction of a child with his/her environment in a contextually appropriate manner. Engagement behavior in children is linked to socio-emotional and cognitive state assessment with enhanced engagement identified with improved skills. A vast majority of studies however rely solely, and often implicitly, on subjective perceptual measures of engagement. Access to automatic quantification could assist researchers/clinicians to objectively interpret engagement with respect to a target behavior or condition, and furthermore inform mechanisms for improving engagement in various settings. In this paper, we present an engagement prediction system based exclusively on vocal cues observed during structured interaction between a child and a psychologist involving several tasks. Specifically, we derive prosodic cues that capture engagement levels across the various tasks. Our experiments suggest that a child's engagement is reflected not only in the vocalizations, but also in the speech of the interacting psychologist. Moreover, we show that prosodic cues are informative of the engagement phenomena not only as characterized over the entire task (i.e., global cues), but also in short term patterns (i.e., local cues). We perform a classification experiment assigning the engagement of a child into three discrete levels achieving an unweighted average recall of 55.8% (chance is 33.3%). While the systems using global cues and local level cues are each statistically significant in predicting engagement, we obtain the best results after fusing these two components. We perform further analysis of the cues at local and global levels to achieve insights linking specific prosodic patterns to the engagement phenomenon. We observe that while the performance of our model varies with task setting and interacting psychologist, there exist universal prosodic patterns reflective of engagement.
Lexical planning is an important part of communication and is reflective of a speaker’s internal state that includes aspects of affect, mood, as well as mental health. Within the study of developmental disorders such as autism spectrum disorder (ASD), language acquisition and language use have been studied to assess disorder severity and expressive capability as well as to support diagnosis. In this paper, we perform a language analysis of children focusing on word usage, social and cognitive linguistic word counts, and a few recently proposed psycholinguistic norms. We use data from conversational samples of verbally fluent children obtained during Autism Diagnostic Observation Schedule (ADOS) sessions. We extract the aforementioned lexical cues from transcripts of session recordings and demonstrate their role in differentiating children diagnosed with Autism Spectrum Disorder from the rest. Further, we perform a correlation analysis between the lexical norms and ASD symptom severity. The analysis reveals an increased affinity by the interlocutor towards use of words with greater feminine association and negative valence.
Atypical speech prosody is a hallmark feature of autism spectrum disorder (ASD) that presents across the lifespan, but is difficult to reliably characterize qualitatively. Given the great heterogeneity of symptoms in ASD, an acoustic-based objective measure would be vital for clinical assessment and interventions. In this study, we investigate speech features in child psychologist conversational samples, including: segmental and suprasegmental pitch dynamics, speech rate, coordination of prosodic attributes, and turn-taking. Data consist of 95 children with ASD as well as 81 controls with non-ASD developmental disorders. We demonstrate significant predictive performance using these features as well as interpret feature correlations of both interlocutors. The most robust finding is that segmental and stiprasegmental prosodic variability increases for both participants in interactions with children having higher ASD severity. Recommendations for future research towards a fully automatic quantitative measure of speech prosody in neurodevelopmental disorders are discussed.
The need for reliable, scalable and efficient diagnosis of Parkinson’s Disease (PD) is a major clinical need. Automating the diagnosis can lead to more accurate and objective predictions as well as provide insights regarding the nature of Parkinson’s condition. This paper proposes a fully automated system to rate the severity (UPDRS-III scale) of PD from patients’ speech. Specifically, the system captures atypicalities in an individual’s voice when performing multiple diverse speaking tasks and makes a unified prediction of the PD severity. The performance is tested in a cross-data setting, with different subjects and dissimilar recording conditions. Results indicate that (i) effective features vary depending on the nature of the specific speech task, (ii) additional novel feature sets to detect distortions in Parkinson’s speech significantly improve the prediction accuracy from the Interspeech15 Challenge baseline system and (iii) our fusion system based on an unsupervised clustering technique also improves the accuracy. Our system incorporates ivector and functionals for segmental features, non-linear time series features, speech rhythm and automatic speech recognition decoding based features. By its application on the Interspeech15 eating condition challenge, the system also shows its potential for detecting other sources of speech variability.
Machine learning has immense potential to enhance diagnostic and intervention research in the behavioral sciences, and may be especially useful in investigations involving the highly prevalent and heterogeneous syndrome of autism spectrum disorder. However, use of machine learning in the absence of clinical domain expertise can be tenuous and lead to misinformed conclusions. To illustrate this concern, the current paper critically evaluates and attempts to reproduce results from two studies (Wall et al. in Transl Psychiatry 2(4):e100, 2012a; PloS One 7(8), 2012b) that claim to drastically reduce time to diagnose autism using machine learning. Our failure to generate comparable findings to those reported by Wall and colleagues using larger and more balanced data underscores several conceptual and methodological problems associated with these studies. We conclude with proposed best-practices when using machine learning in autism research, and highlight some especially promising areas for collaborative work at the intersection of computational and behavioral science.
Atypical speech prosody is a primary characteristic of autism spectrum disorders (ASD), yet it is often excluded from diagnostic instrument algorithms due to poor subjective reliability. Robust, objective prosodic cues can enhance our understanding of those aspects which are atypical in autism. In this work, we connect objective signal-derived descriptors of prosody to subjective perceptions of prosodic awkwardness. Subjectively, more awkward speech is less expressive (more monotone) and more often has perceived awkward rate/rhythm, volume, and intonation. We also find expressivity can be quantified through objective intonation variability features, and that speaking rate and rhythm cues are highly predictive of perceived awkwardness. Acoustic-prosodic features are also able to significantly differentiate subjects with ASD from typically developing (TD) subjects in a classification task, emphasizing the potential of automated methods for diagnostic efficiency and clarity.