OBJECTIVES:Classifying strain in the singing voice can help protect professional singers from vocal overuse and support singing training. This study investigates whether machine learning can automatically classify singing voices into two levels of perceived strain. The singing samples represent two genres: classical and contemporary commercial music (CCM). METHODS:A total of 324 singing voice samples from 15 professional normophonic singers (nine female, six male) were analyzed. Nine singers were classical, and six were CCM singers. The samples consisted of syllable strings produced at three to six pitches and three loudness levels. Based on expert auditory-perceptual ratings, the samples were categorized into two strain levels: normal-mild and moderate-severe. Three acoustic feature sets (mel-frequency cepstral coefficients (MFCCs), the extended Geneva Minimalistic Acoustic Parameter Set (eGeMAPS), and wavelet scattering features) were compared using two classifier models [support vector machine (SVM) and multilayer perceptron (MLP)]. Feature selection was performed using recursive feature elimination, and the Mann-Whitney U test was used to assess the discriminative power of the selected features. RESULTS:The highest classification accuracy of 86.1% was achieved using a subset of wavelet scattering features with the MLP classifier. A comparison between individual features showed that the first MFCC coefficient, representing spectral tilt, exhibited the greatest between-class separation. CONCLUSION:This study demonstrates that machine learning models utilizing selected acoustic features can classify perceptual strain of singing voices automatically with high accuracy. These preliminary findings highlight the potential for larger studies involving more diverse singer groups across different genres.
Speech from people with Parkinson's disease (PD) are likely to be degraded on phonation, articulation, and prosody. Motivated to describe articulation deficits comprehensively, we investigated 1) the universal phonological features that model articulation manner and place, also known as speech attributes, and 2) glottal features capturing phonation characteristics. These were further supplemented by, and compared with, prosodic features using a popular compact feature set and standard MFCC. Temporal characteristics of these features were modeled by convolutional neural networks. Besides the features, we were also interested in the speech tasks for collecting data for automatic PD speech assessment, like sustained vowels, text reading, and spontaneous monologue. For this, we utilized a recently collected Finnish PD corpus (PDSTU) as well as a Spanish database (PC-GITA). The experiments were formulated as regression problems against expert ratings of PD-related symptoms, including ratings of speech intelligibility, voice impairment, overall severity of communication disorder on PDSTU, as well as on the Unified Parkinson's Disease Rating Scale (UPDRS) on PC-GITA. The experimental results show: 1) the speech attribute features can well indicate the severity of pathologies in parkinsonian speech; 2) combining phonation features with articulatory features improves the PD assessment performance, but requires high-quality recordings to be applicable; 3) read speech leads to more accurate automatic ratings than the use of sustained vowels, but not if the amount of speech is limited to correspond to the sustained vowels in duration; and 4) jointly using data from several speech tasks can further improve the automatic PD assessment performance.
PurposeThe purpose of this study was to analyse the relationship between automatic vowel articulation index (aVAI) and direct magnitude estimation (DME) among speakers with Parkinson's disease (PD) and healthy controls. We further analysed the potential of aVAI to serve as an objective measure of speech impairment in the clinical setting.MethodSpeech samples from native Finnish speakers were utilised. Expert raters utilised DME to scale the intelligibility of speech samples. aVAI scores for PD speakers and healthy control speakers were analysed in relationship to DME speech intelligibility ratings and, among PD speakers, disease stage utilising nonparametric statistical analysis.ResultMean DME intelligibility ratings were lower among PD speakers compared to healthy controls. Mean aVAI scores were nearly the same between speaker groups. DME intelligibility ratings and aVAI were strongly correlated within the PD speaker group. aVAI and DME intelligibility ratings were moderately correlated with disease stage as measured by the Hoehn and Yahr scale.ConclusionaVAI was observed to be a promising tool for analysing vowel articulation in PD speakers. Further research is warranted on the application of aVAI as an objective measure of severity of speech impairment in the clinical setting, with varying patient populations and speech samples.
Imprecise vowel articulation can be observed in people with Parkinsons disease (PD).Acoustic features measuring vowel articulation have been demonstrated to be effective indicators of PD in its assessment.Standard clinical vowel articulation features of vowel working space area (VSA), vowel articulation index (VAI) and formants centralization ratio (FCR), are derived the first two formants of the three corner vowels /a/, /i/ and /u/.Conventionally, manual annotation of the corner vowels from speech data is required before measuring vowel articulation.This process is time-consuming.The present work aims to reduce human effort in clinical analysis of PD speech by proposing an automatic pipeline for vowel articulation assessment.The method is based on automatic corner vowel detection using a language universal phoneme recognizer, followed by statistical analysis of the formant data.The approach removes the restrictions of prior knowledge of speaking content and the language in question.Experimental results on a Finnish PD speech corpus demonstrate the efficacy and reliability of the proposed automatic method in deriving VAI, VSA, FCR and F2i/F2u (the second formant ratio for vowels /i/ and /u/).The automatically computed parameters are shown to be highly correlated with features computed with manual annotations of corner vowels.In addition, automatically and manually computed vowel articulation features have comparable correlations with experts ratings on speech intelligibility, voice impairment and overall severity of communication disorder.Language-independence of the proposed approach is further validated on a Spanish PD database, PC-GITA, as well as on TORGO corpus of English dysarthric speech.
Traditionally acoustical assessment of voice disorder relies on simple and homogeneous speech samples like sustained vowels. Continuous speech is believed to be more representative of the daily function of voice and more preferable in clinical practice. This paper describes an attempt on automating voice assessment with continuous speech utterances. The proposed system makes use of a novel type of...
For acoustical assessment of pathological speech, naturally spoken sentences are believed to be most suitable from the perspectives of both patients and clinicians. This is a challenging problem, as the extraction of pathology-dependent features is not straightforward. Previous research showed that features derived from lattice posteriors and decoding results of automatic speech recognition (ASR) could be used to quantifying various types of speech impairments. This paper describes a novel feature that can be derived from phone posterior probabilities generated by an ASR system. The Kullback-Leibler (KL) divergence is used to measure the phone-level distortion between unimpaired and impaired speakers. A Cantonese ASR system is trained with a combination of normal and impaired speech corpora. The multi-task learning approach is applied in order to incorporate different speech characteristics. Experimental results show that the proposed KL divergence feature is effective in the continuous speech based assessment of different pathologies, including voice disorder and post-stroke aphasia. The KL divergence feature is found to outperform conventional acoustic features and supra-segmental duration features, and is complementary to text features in quantifying language impairment.
For perceptual and acoustical assessment of voice disorder, a commonly asked question is about what type of speech samples should be examined and analyzed. While sustained vowels are simple, easy to produce, and suitable for direct comparison among patients, continuous speech is ecologically more valid and better represents the daily use of voice. This study makes an attempt to compare the effectiveness and contribution of sustained vowels and continuous speech in predicting voice disorder severity. To measure voice disorder severity for sustained vowels and continuous speech utterances, two three-class (mild, moderate and severe) classifiers are trained separately with different acoustic features sets. The acoustic features include measures on perturbation, prosody, noise, spectrum and cepstrum. Experimental results on a Cantonese speech database of disordered voice show that sustained vowels and continuous speech are complementary to each other. Relating to their different production complexities, these two types of utterances may exhibit significantly different degrees of voice abnormality, and hence be assigned different severity levels. This suggests that when predicting the subject-level severity based on the utterance-level measures, the speech types of individual utterances must be taken into account. Specifically a posterior fusion approach is proposed for generating subject-level prediction result. In the three-class (mild, moderate and severe) subject-level disorder severity prediction task, the proposed method achieves an accuracy of 75% and AUC (Area Under Curve of ROC) of 0.862.
Most previous studies on acoustic assessment of disordered voice were focused on extracting perturbation features from isolated vowels produced with steady-state phonation. Natural speech, however, is considered to be more preferable in the aspects of flexibility, effectiveness and reliability for clinical practice. This paper presents an investigation on applying automatic speech recognition (ASR) technology to disordered voice assessment of Cantonese speakers. A DNN-based ASR system is trained using phonetically-rich continuous utterances from normal speakers. It was found that frame-level phone posteriors obtained from the ASR system are strongly correlated with the severity level of voice disorder. Phone posteriors in utterances with severe disorder exhibit significantly larger variation than those with mild disorder. A set of utterance-level posterior features are computed to quantify such variation for pattern recognition purpose. An SVM based classifier is used to classify an input utterance into the categories of mild, moderate and severe disorder. The two-class classification accuracy for mild and severe disorders is 90.3%, and significant confusion between mild and moderate disorders is observed. For some of the subjects with severe voice disorder, the classification results are highly inconsistent among individual utterances. Furthermore, short utterances tend to have more classification errors.
This paper describes the application of state-of-the-art automatic speech recognition (ASR) systems to objective assessment of voice and speech disorders. Acoustical analysis of speech has long been considered a promising approach to non-invasive and objective assessment of people. In the past the types and amount of speech materials used for acoustical assessment were very limited. With the ASR technology, we are able to perform acoustical and linguistic analyses with a large amount of natural speech from impaired speakers. The present study is focused on Cantonese, which is a major Chinese dialect. Two representative disorders of speech production are investigated: dysphonia and aphasia. ASR experiments are carried out with continuous and spontaneous speech utterances from Cantonese-speaking patients. The results confirm the feasibility and potential of using natural speech for acoustical assessment of voice and speech disorders, and reveal the challenging issues in acoustic modeling and language modeling of pathological speech.
Acoustical analysis of speech is considered a favorable and promising approach to objective assessment of voice disorders. Previous research emphasized on the extraction and classification of voice quality features from sustained vowel sounds. In this paper, an investigation on voice assessment using continuous speech utterances of Cantonese is presented. A DNN-HMM based speech recognition system is trained with speech data of unimpaired voice. The recognition accuracy for pathological utterances is found to decrease significantly with the disorder severity increasing. Average acoustic posterior probabilities are computed for individual phones from the speech recognition output lattices and the DNN soft-max layer. The phone posteriors obtained for continuous speech from the mild, moderate and severe categories are highly distinctive and thus useful to the determination of voice disorder severity. A subset of Cantonese phonemes are identified to be suitable and reliable for voice assessment with continuous speech.