The LENA system was designed and validated to provide information about the language environment in children 0 to 4 years of age and its use has been expanded to populations with a number of communication profiles. Its utility in children 5 years of age and older is not yet known. The present study used acoustic data from two samples of children with autism spectrum disorders (ASD) to evaluate the reliability of LENA automated analyses for detecting speech utterances in older, school age children, and adolescents with ASD, in clinic and home environments. Participants between 5 and 18 years old who were minimally verbal (study 1) or had a range of verbal abilities (study 2) completed standardized assessments in the clinic (study 1 and 2) and in the home (study 2) while speech was recorded from a LENA device. We compared LENA segment labels with manual ground truth coding by human transcribers using two different methods. We found that the automated LENA algorithms were not successful (<50% reliable) in detecting vocalizations from older children and adolescents with ASD, and that the proportion of speaker misclassifications by the automated system increased significantly with the target‐child's age. The findings in children and adolescents with ASD suggest possibly misleading results when expanding the use of LENA beyond the age ranges for which it was developed and highlight the need to develop novel automated methods that are more appropriate for older children. Autism Research 2019, 12: 628–635. © 2019 International Society for Autism Research, Wiley Periodicals, Inc.Lay SummaryCurrent commercially available speech detection algorithms (LENA system) were previously validated in toddlers and children up to 48 months of age, and it is not known whether they are reliable in older children and adolescents. Our data suggest that LENA does not adequately capture speech in school age children and adolescents with autism and highlights the need to develop new automated methods for older children.
In this work, we applied Adaptive Neuro-Fuzzy Inference System to three different classification problems: (1) sentence-level subjectivity detection, (2) sentiment analysis of texts, and (3) detecting user intention in natural language call routing system. We used English dataset for the first and second problems, but Azerbaijani dataset for the third problem based on same features. Our feature extraction algorithm calculates a feature vector based on the statistical occurrences of words in a corpus without any lexical knowledge.
The context analysis of customer requests in a natural language call routing problem is investigated in the paper. One of the most significant problems in natural language call routing is a comprehension of client request. With the aim of finding a solution to this issue, the Hybrid HMM and ANFIS models become a subject to an examination. Combining different types of models (ANFIS and HMM) can prevent misunderstanding by the system for identification of user intention in dialogue system. Based on these models, the hybrid system may be employed in various language and call routing domains due to non-usage of lexical or syntactic analysis in classification process.
Children's early language environments are related to later development. Little is known about this association in siblings of children with autism spectrum disorder (ASD), who often experience language delays or have ASD. Fifty-nine 9-month-old infants at high or low familial risk for ASD contributed full-day in-home language recordings. High-risk infants produced more vocalizations than low-risk peers; conversational turns and adult words did not differ by group. Vocalization differences were driven by a subgroup of "hypervocal" infants. Despite more vocalizations overall, these infants engaged in less social babbling during a standardized clinic assessment, and they experienced fewer conversational turns relative to their rate of vocalizations. Two ways in which these individual and environmental differences may relate to subsequent development are discussed.
It has been established that children and adolescents with Autism Spectrum Disorder show a wide range of abilities to use spoken words and establish interactive conversations. Automatic measurement of such abilities in naturalistic environments would greatly facilitate assessment and monitoring of such individuals. If sufficient accuracy could be achieved, studies could be performed whose sample sizes are large enough to draw meaningful conclusions. In the current study, 16-hour home audio recordings using the LENA (Language ENvironment Analysis) device are examined for older children and adolescents. Specific enhancements to the existing LENA analysis platform include the ability to diarize recordings for subjects aged 5 through 13, to detect non-verbal vocalizations such as laughter and whining, to identify child-directed speech, and to determine when questions are posed. Other higher-level descriptors involve extraction of affect, computation of conversational interaction measures, detection of cross-talk and interruption events, and identifying emotional outbursts. Subject-specific diarization based on a small amount of hand-labeled data yields acceptable accuracy. However, a newly developed system based on i-vectors, specifically designed for the environment at hand, requires no such labeling at the onset.
Automated sleep staging is an extensively researched problem with a spread of hand-crafted feature representations. Currently available representations, however, fail to produce sufficiently accurate results. Previously used features tend to struggle with artifacts and inter-patient variability. To address these issues, an aligned time-frequency block structure model was created. This model can be learned by building upon a combination of existing denoising and consensus clustering techniques. Across multiple datasets, this model significantly reduced error rates from raw spectral features and outperformed bandpower features commonly used in commercial tools. For the DREAMS dataset, classic band power features yielded a 30% error rate; raw spectral features had a higher error rate of 37%; and the novel Dense Denoised Spectral (DDS) features resulted in a 17% error rate.
Detecting child and adult vocalizations, and computing their characteristics from audio recorded in natural home environments can be useful in many applications. The current study is interested in monitoring children with autism spectrum disorder to ultimately provide outcome measures that can track the efficacy of clinical treatments. In this paper, we show that it is possible to automate detection of child and adult vocalizations from audio recorded in controlled clinic environments as well as in naturalistic home settings. The results show both high precision and recall for children aged five to fourteen years who have been diagnosed with autism. Further, we describe a highly accurate speaker-independent laughter detector for this age group which will be useful for affect estimation.
PURPOSE:Routine evaluation of basic surgical skills in medical schools requires considerable time and effort from supervising faculty. For each surgical trainee, a supervisor has to observe the trainees in person. Alternatively, supervisors may use training videos, which reduces some of the logistical overhead. All these approaches however are still incredibly time consuming and involve human bias. In this paper, we present an automated system for surgical skills assessment by analyzing video data of surgical activities.METHOD:We compare different techniques for video-based surgical skill evaluation. We use techniques that capture the motion information at a coarser granularity using symbols or words, extract motion dynamics using textural patterns in a frame kernel matrix, and analyze fine-grained motion information using frequency analysis.RESULTS:We were successfully able to classify surgeons into different skill levels with high accuracy. Our results indicate that fine-grained analysis of motion dynamics via frequency analysis is most effective in capturing the skill relevant information in surgical videos.CONCLUSION:Our evaluations show that frequency features perform better than motion texture features, which in-turn perform better than symbol-/word-based features. Put succinctly, skill classification accuracy is positively correlated with motion granularity as demonstrated by our results on two challenging video datasets.
We present an automated framework for visual assessment of the expertise level of surgeons using the OSATS Objective Structured Assessment of Technical Skills criteria. Video analysis techniques for extracting motion quality via frequency coefficients are introduced. The framework is tested on videos of medical students with different expertise levels performing basic surgical tasks in a surgical training lab setting. We demonstrate that transforming the sequential time data into frequency components effectively extracts the useful information differentiating between different skill levels of the surgeons. The results show significant performance improvements using DFT and DCT coefficients over known state-of-the-art techniques.
Earlier studies have shown that certain emotional characteristics are best observed at different analysis-frame lengths. When features of multiple modalities are extracted, it is reasonable to believe that different temporal lengths would better model the underlying characteristics that result from different emotions. In this study, we examine the use of such differing timescales in constructing emotion classifiers. A novel fusion method is introduced that utilizes the outputs of individual classifiers that are trained using multi-dimensional inputs with multiple temporal lengths. We used the IEMOCAP database which contains audiovisual information of 10 subjects in dyadic interaction settings. The classification task was performed over three emotional dimensions: valence, activation, and dominance. The results demonstrate the utility of the multimodal-multitemporal approach. Statistically significant improvements in accuracy are seen for in all three dimensions when compared with unimodal-unitemporal classifiers.
In a previous study, a robust formant-tracking algorithm was introduced to model formant and spectral properties of speech. The algorithm utilizes Gaussian mixtures to estimate spectral parameters, and refines the estimates by using a maximum a posteriori adaptation (MAP) algorithm. In this paper, the formant-tracking algorithm was used to extract the formant-based features for emotion classification. The classification results were compared to a linear predictive coding (LPC) based algorithm for evaluation. On average, the formant features extracted using the algorithm improved the unweighted accuracy by 2.1 percentage points when compared to a LPC-based algorithm. The combination of formant features and other acoustic features statistically significantly improved the unweighted accuracy by 2.7 percentage points, whereas the LPC-based features barely improved it by 1 percentage point. The results clearly indicate that an improved formant-tracking method improved emotion classification accuracy. The effect of formant-based features in emotion classification is also discussed.
Many medications and therapies are available to treat neurological movement disorder symptoms such as tremor, bradykinesia, postural instability, and gait disturbances. However, proper dosages and treatment combinations can be difficult to determine and often must be adjusted over time. Such adjustments are typically based on infrequent clinical evaluations. Mobile, long term monitoring of motor signs will allow better informed treatment adjustments and, as a result, has the potential to improve the quality of life for patients. With this in mind, we built a system for portable monitoring of neurological disorders in humans. The system analyzes movement signals from distributed wirelessly connected gyroscopes and accelerometers using a commercial smartphone, off-the-shelf attachments, and custom-programmed software. With approval from the Georgia Institute of Technology Institutional Review Board, we conducted a study with human subjects to evaluate whether unsupervised remote mobile monitoring might be possible with such a system, and, if so, how accurately the system could assess the signs of Parkinson's disease. We found that it was useful to consider data from multiple sensors in order to discriminate between normal and parkinsonian movement and that it was possible to identify activity types from the accelerometer and gyroscope data with high accuracy (1.0). After the activity type was identified, average discrimination accuracy between parkinsonian and normal conditions was 0.88. Additionally, individual symptoms of the disease could be accurately detected in > 0.8 of cases.
Head and neck cancer can significantly hamper speech production which often reduces speech intelligibility. A method of extracting spectral features is presented. The method uses a multi-resolution sinusoidal transform scheme, which enables better representation of spectral and harmonic characteristics. Regression methods were used to predict interval-scaled intelligibility scores of utterances in the NKI-CCRT speech corpus. The inclusion of these features lowered the mean squared estimation error from 0.43 to 0.39 on a scale from 1 to 7, with a p-value less than 0.001. For binary intelligibility classification, their inclusion resulted in an improvement by 5.0 percentage points when tested on a disjoint set.
Paralinguistic cues in children’s speech convey the child’s affective state and can serve as important markers for the early detection of autism spectrum disorder (ASD). In this paper, we detect paralinguistic events, such as laughter and fussing/crying, along with toddlers’ speech from the Multi-modal Dyadic Behavior Dataset (MMDB). We use both spectral and prosodic acoustic features selected using a combination of filter and wrapperbased methods. The classification accuracy using a support vector machine with a linear kernel for detecting laughter in children’s speech was 77.87% and that for fussing/crying was 79.37%. A tertiary classification scheme for detecting laughter, fussing/crying, and speech yielded an accuracy of 69.73%. To test for the generalization of the approach for detecting fussing/crying, we used recordings from the Strange Situation protocol, which is used to observe attachment behavior between an infant and a parent. Using a cross-corpus testing set for detecting fussing/crying, we obtained a detection accuracy of 71.6%. These results indicate that the selected acoustic features are capable of discriminating children’s laughter, fussing/crying, and speech and the algorithms generalize well to a dataset consisting of paralinguistic cues of a different age group, infants (12 18 months of age), gathered in a different context.
This paper details a novel probabilistic method for automatic neural spike sorting which uses stochastic point process models of neural spike trains and parameterized action potential waveforms. A novel likelihood model for observed firing times as the aggregation of hidden neural spike trains is derived, as well as an iterative procedure for clustering the data and finding the parameters that maximize the likelihood. The method is executed and evaluated on both a fully labeled semiartificial dataset and a partially labeled real dataset of extracellular electric traces from rat hippocampus. In conditions of relatively high difficulty (i.e., with additive noise and with similar action potential waveform shapes for distinct neurons) the method achieves significant improvements in clustering performance over a baseline waveform-only Gaussian mixture model (GMM) clustering on the semiartificial set (1.98% reduction in error rate) and outperforms both the GMM and a state-of-the-art method on the real dataset (5.04% reduction in false positive + false negative errors). Finally, an empirical study of two free parameters for our method is performed on the semiartificial dataset.
Sleep is a key requirement for an individual's health, though currently the options to study sleep rely largely on manual visual classification methods. In this paper we propose a new scheme for automated offline classification based upon cross-frequency-coupling (CFC) and compare it to the traditional band power estimation and the more recent preferential frequency band information estimation. All three approaches allowed sleep stage classification and provided whole-night visualization of sleep stages. Surprisingly, the simple average power in band classification achieved better overall performance than either the preferential frequency band information estimation or the CFC approach. However, combined classification with both average power and CFC features showed improved classification over either approach used singly.
In this study, modulation index (MI) features derived from local field potential (LFP) recordings in the sub-thalamic nucleus (STN) and electroencephalographic recordings (EEGs) from the primary motor cortex are shown to correlate with both the overall motor impairment and motor subscores in a monkey model of parkinsonism. The MI features used are measures of phase-amplitude cross frequency coupling (CFC) between frequency sub-bands. We used complex wavelet transforms to extract six spectral sub-bands within the 3-60 Hz range from LFP and EEG signals. Using the method of canonical correlation, we show that weighted combinations of the MI features in LFP or EEG signals correlate significantly with individual and composite scores on a scale for parkinsonian disability.
James M. Rehg合作论文数Siebel School of Computing and Data Science, The Grainger College of Engineering, University of Illinois Urbana-Champaign3