Recommendations for common outcome measures following pediatric traumatic brain injury (TBI) support the integration of instrumental measurements alongside perceptual assessment in recovery and treatment plans. A comprehensive set of sensitive, robust and non-invasive measurements is therefore essential in assessing variations in speech characteristics over time following pediatric TBI. In this article, we study the changes in the acoustic speech patterns of a pediatric cohort of ten subjects diagnosed with severe TBI. We extract a diverse set of both well-known and novel acoustic features from child speech recorded throughout the year after the child produced intelligible words. These features are analyzed individually and by speech subsystem, within-subject and across the cohort. As a group, older children exhibit highly significant (p<0.01) increases in pitch variation and phoneme diversity, shortened pause length, and steadying articulation rate variability. Younger children exhibit similar steadied rate variability alongside an increase in formant-based articulation complexity. Correlation analysis of the feature set with age and comparisons to normative developmental data confirm that age at injury plays a significant role in framing the recovery trajectory. Nearly all speech features significantly change (p<0.05) for the cohort as a whole, confirming that acoustic measures supplementing perceptual assessment are needed to identify efficacious treatment targets for speech therapy following TBI.
Repeated subconcussive blows to the head during sports or other contact activities may have a cumulative and long lasting effect on cognitive functioning. Unobtrusive measurement and tracking of cognitive functioning is needed to enable preventative interventions for people at elevated risk of concussive injury. The focus of the present study is to investigate the potential for using passive measurements of fine motor movements (smooth pursuit eye tracking and read speech) and resting state brain activity (measured using fMRI) to complement existing diagnostic tools, such as the Immediate Post-concussion Assessment and Cognitive Testing (ImPACT), that are used for this purpose. Thirty-one high school American football and soccer athletes were tracked through the course of a sports season. Hypotheses were that (1) measures of complexity of fine motor coordination and of resting state brain activity are predictive of cognitive functioning measured by the ImPACT test, and (2) within-subject changes in these measures over the course of a sports season are predictive of changes in ImPACT scores. The first principal component of the six ImPACT composite scores was used as a latent factor that represents cognitive functioning. This latent factor was positively correlated with four of the ImPACT composites: verbal memory, visual memory, visual motor speed and reaction speed. Strong correlations, ranging between r = 0.26 and r = 0.49, were found between this latent factor and complexity features derived from each sensor modality. Based on a regression model, the complexity features were combined across sensor modalities and used to predict the latent factor on out-of-sample subjects. The predictions correlated with the true latent factor with r = 0.71. Within-subject changes over time were predicted with r = 0.34. These results indicate the potential to predict cognitive performance from passive monitoring of fine motor movements and brain activity, offering initial support for future application in detection of performance deficits associated with subconcussive events.
Purpose A common way of eliciting speech from individuals is by using passages of written language that are intended to be read aloud. Read passages afford the opportunity for increased control over the phonetic properties of elicited speech, of which phonetic balance is an often-noted example. No comprehensive analysis of the phonetic balance of read passages has been reported in the literature. The present article provides a quantitative comparison of the phonetic balance of widely used passages in English. Method Assessment of phonetic balance is carried out by comparing the distribution of phonemes in several passages to distributions consistent with typical spoken English. Data regarding the distribution of phonemes in spoken American English are aggregated from the published literature and large speech corpora. Phoneme distributions are compared using Spearman rank order correlation coefficient to quantify similarities of phoneme counts in those sources. Results Correlations between phoneme distributions in read passages and aggregated material representative of spoken American English ranged from .70 to .89. Correlations between phoneme counts from all passages, literature sources, and corpus sources ranged from .55 to .99. All correlations were statistically significant at the Bonferroni-adjusted level. Conclusions Passages considered in the present work provide high, but not ideal, phonetic balance. Space exists for the creation of new passages that more closely match the phoneme distributions observed in spoken American English. The Caterpillar provided the best phonetic balance, but phoneme distributions in all considered materials were highly similar to each other.
Autism Spectrum Disorder (ASD) is a developmental disorder characterized by difficulty in communication, which includes a high incidence of speech production errors. We hypothesize that these errors are partly due to underlying deficits in motor coordination and control, which are also manifested in degraded fine motor control of facial expressions and purposeful hand movements. In this pilot study, we computed correlations of acoustic, video, and handwriting time-series derived from five children with ASD and five children with neurotypical development during speech and handwriting tasks. These correlations and eigenvalues derived from the correlations act as a proxy for motor coordination across articulatory, laryngeal, and respiratory speech production systems and for fine motor skills. We utilized features derived from these correlations to discriminate between children with and without ASD. Eigenvalues derived from these correlations highlighted differences in complexity of coordination across speech subsystems and during handwriting, and helped discriminate between the two subject groups. These results suggest differences in coupling within speech production and fine motor skill systems in children with ASD. Our long-term goal is to create a platform assessing motor coordination in children with ASD in order to track progress from speech and motor interventions administered by clinicians.
Between 15% to 40% of mild traumatic brain injury (mTBI) patients experience incomplete recoveries or provide subjective reports of decreased motor abilities, despite a clinically-determined complete recovery. This demonstrates a need for objective measures capable of detecting subclinical residual mTBI, particularly in return-to-duty decisions for warfighters and return-to-play decisions for athletes. In this paper, we utilize features from recordings of directed speech and gait tasks completed by ten healthy controls and eleven subjects with lingering subclinical impairments from an mTBI. We hypothesize that decreased coordination and precision during fine motor movements governing speech production (articulation, phonation, and respiration), as well as during gross motor movements governing gait, can be effective indicators of subclinical mTBI. Decreases in coordination are measured from correlations of vocal acoustic feature time series and torso acceleration time series. We apply eigenspectra derived from these correlations to machine learning models to discriminate between the two subject groups. The fusion of correlation features derived from acoustic and gait time series achieve an AUC of 0.98. This highlights the potential of using the combination of vocal acoustic features from speech tasks and torso acceleration during a simple gait task as a rapid screening tool for subclinical mTBI.1
Objective Military job and training activities place significant demands on service members' (SMs') cognitive resources, increasing risk of injury and degrading performance. Early detection of cognitive fatigue is essential to reduce risk and support optimal function. This paper describes a multimodal approach, based on changes in measures of speech motor coordination and electrodermal activity (EDA), for predicting changes in performance following sustained cognitive effort. Methods Twenty-nine active duty SMs completed computer-based cognitive tasks for 2 h (load period). Measures of speech derived from audio were acquired, along with concurrent measures of EDA, before and after the load period. Cognitive performance was assessed before and during the load period using the Automated Neuropsychological Assessment Metrics Military Battery (ANAM MIL). Subjective assessments of cognitive effort and alertness were obtained intermittently. Results Across the load period, participants' ratings of cognitive workload increased, while alertness ratings declined. Cognitive performance declined significantly during the first half of the load period. Three speech and arousal features predicted cognitive performance changes during this period with statistically significant accuracy: EDA (r = 0.43,p = 0.01), articulator velocity coordination (r = 0.50,p = 0.00), and vocal creak (r = 0.35,p = 0.03). Fusing predictions from these features predicted performance changes withr = 0.68 (p = 0.00). Conclusions Results suggest that speech and arousal measures may be used to predict changes in performance associated with cognitive fatigue. This work supports ongoing efforts to develop reliable, unobtrusive measures for cognitive state assessment aimed at reducing injury risk, informing return to work decisions, and supporting diverse mobile healthcare applications in civilian and military settings.
Lapses in vigilance and slowed reactions due to mental fatigue can increase risk of accidents and injuries and degrade performance. This paper describes a method for rapid, unobtrusive detection of mental fatigue based on changes in electrodermal arousal (EDA), and changes in neuromotor coordination derived from speaking. Twenty-nine Soldiers completed a 2-hour battery of cognitive tasks intended to induce fatigue. Behavioral markers derived from audio and video during speech were acquired before and after the 2hour cognitive load tasks, as was EDA. Exposure to cognitive load produced detectable increases in neuromotor variability in speech and facial measures after load and even after a recovery period. A Gaussian mixture model classifier with cross-validation and fusion across speech, video, and EDA produced an accuracy of AUC=0.99 in detecting a change in cognitive fatigue relative to a personalized baseline.
Recommendations following pediatric traumatic brain injury (TBI) support the integration of instrumental measurement to aid perceptual assessment in recovery and treatment plans. A comprehensive set of sensitive, robust and non-invasive measurements is therefore essential in assessing variations in speech characteristics over time following pediatric TBI. In this paper, we discuss a method for measuring changes in the speech patterns of a pediatric cohort of ten subjects diagnosed with severe TBI. We apply a diverse set of both well-known and novel feature measurements to child speech recorded throughout the year following diagnosis. We analyze these features individually and by speech subsystem for each subject as well as for the entire cohort. In children older than 72 months, we find highly significant (p < 0.01) increases in pitch variation and number of unique phonemes spoken, shortened pause length, and steadying articulation rate variability. Younger children exhibit similar steadied rate variability alongside an increase in articulation complexity. Nearly all speech features significantly change ( p < 0.05) for the cohort as a whole, confirming that acoustic measures expanding upon perceptual assessment are needed to identify efficacious treatment targets for speech therapy following TBI.(1)
Competitive international language recognition evaluations have been hosted by NIST for over two decades. This paper describes the MIT Lincoln Laboratory (MITLL) and Johns Hopkins University (JHU) submission for the recent 2017 NIST language recognition evaluation (LRE17) [1]. The MITLL/JHU LRE17 submission represents a collaboration between researchers at MITLL and JHU with multiple sub-systems reflecting a range of language recognition technologies including traditional MFCC/SDC i-vector systems, deep neural network (DNN) bottleneck feature based i-vector systems, stateof-the-art DNN x-vector systems and a sparse coding system. Each sub-systems uses the same backend processing for domain adaptation and score calibration. Multiple sub-systems were fused using a simple logistic regression ([2]) to create system combinations. The MITLL/JHU submissions were selected based on the top ranking combinations of up to 5 sub-systems using development data provided by NIST. The MITLL/JHU primary submitted systems attained a Cavg of 0.181 and 0.163 for the fixed and open conditions respectively. Post evaluation analysis revealed the importance of carefully partitioning for the development data, using augmented training data and using a condition dependent backend. Addressing these issues including retraining the x-vector system with augmented data yielded gains in performance of over 17%: a Cavg of 0.149 for the fixed condition and 0.132 for the open condition.
In this paper, the NIST 2016 SRE system that resulted from the collaboration between MIT Lincoln Laboratory and the team at Johns Hopkins University is presented. The submissions for the 2016 evaluation consisted of three fixed condition submissions and a single system open condition submission. The primary submission on the fixed (and core) condition resulted in an actual DCF of .618. Details of the submissions are discussed along with some discussion and observations of the 2016 evaluation campaign.
Spoken language recognition requires a series of signal processing steps and learning algorithms to model distinguishing characteristics of different languages. In this paper, we present a sparse discriminative feature learning framework for language recognition. We use sparse coding, an unsupervised method, to compute efficient representations for spectral features from a speech utterance while learning basis vectors for language models. Differentiated from existing approaches in sparse representation classification, we introduce a maximum a posteriori (MAP) adaptation scheme based on online learning that further optimizes the discriminative quality of sparse-coded speech features. We empirically validate the effectiveness of our approach using the NIST LRE 2015 dataset.
: In this paper we describe the most recent MIT Lincoln Laboratory language recognition system developed for the NIST 2015 Language Recognition Evaluation (LRE). The submission features a fusion of five core classifiers, with most systems developed in the context of an i-vector framework. The 2015 evaluation presented new paradigms. First, the evaluation included fixed training and open training tracks for the first time; second, language classification performance was measured across 6 language clusters using 20 language classes instead of an N-way language task; and third, performance was measured across a nominal 3-30 second range. Results are presented for the overall performance across the six language clusters for both the fixed and open training tasks. On the 6-cluster metric the Lincoln system achieved overall costs of 0.173 and 0.168 for the fixed and open tasks respectively.
Unsupervised feature learning methods have proven effective for classification tasks based on a single modality. We present multimodal sparse coding for learning feature representations shared across multiple modalities. The shared representations are applied to multimedia event detection (MED) and evaluated in comparison to unimodal counterparts, as well as other feature learning methods such as GMM supervectors and sparse RBM. We report the cross-validated classification accuracy and mean average precision of the MED system trained on features learned from our unimodal and multimodal settings for a subset of the TRECVID MED 2014 dataset.
: Large unstructured audio data sets have become ubiquitous and present a challenge for organization and search. One logical approach for structuring data is to find common speakers and link occurrences across different recordings. Prior approaches to this problem have focused on basic methodology for the linking task. In this paper, we introduce a novel trainable nonparametric hashing method for indexing large speaker recording data sets. This approach leads to tunable computational complexity methods for speaker linking. We focus on a scalable clustering method based on hashingcanopy-clustering. We apply this method to a large corpus of speaker recordings, demonstrate performance tradeoffs, and compare to other hashing methods.
The goal of this paper is to describe significant corpora available to support speaker recognition research and evaluation, along with details about the corpora collection and design. We describe the attributes of high-quality speaker recognition corpora. Considerations of the application, domain, and performance metrics are also discussed. Additionally, a literature survey of corpora used in speaker recognition research over the last 10 years is presented. Finally we show the most common corpora used in the research community and review them on their success in enabling meaningful speaker recognition research.
This paper presents a description of the MIT Lincoln Laboratory (MITLL) language recognition system developed for the NIST 2011 Language Recognition Evaluation (LRE). The submitted system consisted of a fusion of four core classifiers, three based on spectral similarity and one based on tokenization. Additional system improvements were achieved following the submission deadline. In a major departure from previous evaluations, the 2011 LRE task focused on closed-set pairwise performance so as to emphasize a system’s ability to distinguish confusable language pairs. Results are presented for the 24-language confusable pair task at test utterance durations of 30, 10, and 3 seconds. Results are also shown using the standard detection metrics (DET, minDCF) and it is demonstrated the previous metrics adequately cover difficult pair performance. On the 30 s 24-language confusable pair task, the submitted and post-evaluation systems achieved average costs of 0.079 and 0.070 and standard detection costs of 0.038 and 0.033.
The NIST speaker recognition evaluation (SRE) featured microphone data in the 2005-2010 evaluations. The preprocessing and use of this data has typically been performed with telephone bandwidth and quantization. Although this approach is viable, it ignores the richer properties of the microphone data— multiple channels, high-rate sampling, linear encoding, ambient noise properties, etc. In this paper, we explore alternate choices of preprocessing and examine their effects on speaker recognition performance. Specifically, we consider the effects of quantization, sampling rate, enhancment, and two-channel speech activity detection. Experiments on the NIST 2010 SRE interview microphone corpus demonstrate that performance can be dramatically improved with a different preprocessing chain.
: We present a new perspective on the subspace compensation techniques that currently dominate the field of speaker recognition using Gaussian Mixture Models (GMMs). Rather than the traditional factor analysis approach, we use Gaussian modeling in the sufficient statistic supervector space combined with Probabilistic Principal Component Analysis (PPCA) within-class and shared across class covariance matrices to derive a family of training and testing algorithms. Key to this analysis is the use of two noise terms for each speech cut: a random channel offset and a length dependent observation noise. Using the Wiener filtering perspective, formulas for optimal train and test algorithms for Joint Factor Analysis (JFA) are simple to derive. In addition, we can show that an alternative form of Wiener filtering results in the i-vector approach. thus tying together these two disparate techniques.