Chronic kidney disease (CKD) is a global health concern characterized by a gradual and irreversible decline in kidney function. Early diagnosis and timely intervention are crucial, yet current methods rely primarily on invasive blood and urine tests. Since CKD affects the respiratory system and alters speech production, vocal characteristics may serve as biomarkers for disease detection. This study proposes a deep learning-based approach that integrates spectrogram and glottal features for CKD diagnosis. Spectrograms capture broad acoustic characteristics, whereas glottal features, known to be influenced by CKD, provide complementary phonatory information. To effectively fuse these features, we employ a transformer-like architecture. The proposed method achieves an accuracy and a macro F1 score of 0.96, demonstrating its potential as an objective, non-invasive diagnostic tool. In addition, we analyze attention weights and gradient-based saliency maps to enhance model interpretability.
This study presents a novel multimodal and multitask learning model for predicting five proficiency scores of L2 English speeches. The proposed approach integrates speech and text embeddings using multimodal transformer blocks with cross-modal attention to refine features dynamically between modalities, capturing complementary information. A joint loss function, combining MSE and a Trait-Aware (TA) loss, enhances the model by leveraging relationships among proficiency traits. Experiments with different combinations of four embeddings (MFCCs, GloVe, wav2vec 2.0, and BERT) revealed that the proposed model with wav2vec 2.0 and BERT embeddings achieved the best performance, with a mean PCC of 0.734 and a standard deviation of 0.0129 across five criteria. This approach significantly outperforms unimodal and baseline multimodal models, demonstrating the potential of advanced multimodal architectures and task-aware optimization in automated speech assessment systems.
This paper presents a novel multilingual speech translation corpus for complex, domain-specific content in Korean, English, Spanish, and Japanese. The corpus contains 4,000 hours of parallel speech, including 1,000 hours of Korean audio with simultaneous sight interpretations in the other three languages by 294 professionals (242 interpreters and 52 Korean voice actors). It also includes transcriptions, translations, and annotations for all languages. The Dewey Decimal Classification was adapted to balance knowledge representation, and speech tasks were conducted in a controlled studio environment to ensure data consistency. Translation, transcription, and annotation workflows were managed through a custom-built platform. The corpus captures nuanced contexts, cultural sensitivities, and domain-specific terminology, addressing linguistic challenges like structural differences between SOV (Korean, Japanese) and SVO languages (English, Spanish). Preliminary evaluations indicate its potential to enhance end-to-end speech translation models, support cross-lingual transfer learning, and tackle real-time translation issues.
Although substantial research has been conducted on automatic speech assessment models leveraging speech representations derived from self-supervised learning models, the underlying mechanisms remain relatively underexplored. This study investigates the acoustic foundations of automatic speech production skill assessment models for children with cochlear implants, which helps enhance model performance and elucidate the basis of assessment outcomes. We analyze the statistical differences in acoustic characteristics as a function of speech scores for articulation and prosody. Using a general probing approach, models are trained with layer-wise embeddings from wav2vec2.0 and probed through simple regression models. The probing model performance is interpreted as an indicator of the information encoded within the speech representations. Experimental results demonstrate that the assessment models capture distinct acoustic features depending on the target of assessment, shedding light on the acoustic basis of the results and revealing the strengths and limitations of the models.
Automatic speech assessment plays a vital role in language learning by providing essential feedback on pronunciation, fluency, and overall speaking ability. However, developing effective multilingual speech assessment systems poses significant challenges with the complexity of modeling multiple languages and limited availability of labeled data, especially for languages other than English. In this study, we propose a multilingual speech assessment system for three languages-English, German, and French, which are produced by Korean learners. Enhanced by cross-attention and multitask learning mechanisms, our model utilizes pre-trained models to capture both language-specific and cross-linguistic features, predicting overall speaking proficiency scores directly from raw speech audio. Experimental results demonstrate that our proposed method, especially with wav2vec 2.0, presents superior performance on both seen and unseen data compared to monolingual models.
Autism Spectrum Disorder (ASD) is a neurodevelopmental condition characterized by deficits in social communication, affecting both language use and speech patterns. Since assessment relies on behavioral observations rather than standardized medical tests, developing an objective evaluation method is essential. Recognizing that ASD impacts both language and speech production, this study proposes a cascaded multi-modal framework for ASD severity assessment. The framework processes raw audio, generates transcriptions via automatic speech recognition, and extracts linguistic and acoustic features using speech-language foundation models. Given the atypical suprasegmental and segmental speech characteristics in ASD, two speech foundation models are employed. A co-attention mechanism then integrates these representations to estimate severity. Achieving a Spearman's correlation of 0.5629 with human ratings, the proposed approach offers a scalable, fully automated ASD assessment tool.
L2 pronunciation is shaped by the interaction of two sound systems, which makes their identity more complex than a single phoneme category. The non-categorical nature demands assessment at a level finer than phonemes. As the granular requirement is highly labor-intensive, unsupervised methods emerged. Nevertheless, they either reverted to categorical diagnosis or used the supervised and phoneme-prescribed feature phonetic posterior-gram (PPG). Alternatively, this study adopts the unprescribed and unsupervised feature, the Wav2Vec2.0 code vector, to locate sub-phonemic variations. We first verify the features’ L2 discernability by comparing their frequency across single-speaker data of L1 (CMU ARCTIC) and L2 (L2 ARCTIC). Clustering is performed on frequency vectors to test their separability on account of nativeness. Subsequently, sub-segmental patterns are analyzed among segmentally identical error samples in L2 Korean English NIA 037 data. After cataloging segmental errors detected by the model finetuned with L1 TIMIT, their corresponding code vector sequences are extracted by referencing the forced alignment result. We then derived dominant patterns of the sequences and compared them against L1 reference materials constructed from TIMIT. Phoneme-code vector co-occurrence probability and code vector clustering were each used to check their attributes and uniqueness. The result confirmed the discernability, followed by linguistically interpretable common traits across patterns. (1) They formed a gradient error continuum along the changed articulatory value, reflecting the non-categorical nuanced understanding. (2) This trait is highlighted by intermediary typology assuming opposite values in two codebooks which was also rare in L1 for being L2 specific. Lastly, (3) distribution skewed towards the most approximate sound in the learner’s L1, from which the patterns’ complexity stems.
Autism Spectrum Disorder (ASD) is a lifelong condition that significantly influencing an individual's communication abilities and their social interactions. Early diagnosis and intervention are critical due to the profound impact of ASD's characteristic behaviors on foundational developmental stages. However, limitations of standardized diagnostic tools necessitate the development of objective and precise diagnostic methodologies. This paper proposes an end-to-end framework for automatically predicting the social communication severity of children with ASD from raw speech data. This framework incorporates an automatic speech recognition model, fine-tuned with speech data from children with ASD, followed by the application of fine-tuned pre-trained language models to generate a final prediction score. Achieving a Pearson Correlation Coefficient of 0.6566 with human-rated scores, the proposed method showcases its potential as an accessible and objective tool for the assessment of ASD.
Despite the growing demand for digital therapeutics for children with Autism Spectrum Disorder (ASD), there is currently no speech corpus available for Korean children with ASD. This paper introduces a speech corpus specifically designed for Korean children with ASD, aiming to advance speech technologies such as pronunciation and severity evaluation. Speech recordings from speech and language evaluation sessions were transcribed, and annotated for articulatory and linguistic characteristics. Three speech and language pathologists rated these recordings for social communication severity (SCS) and pronunciation proficiency (PP) using a 3-point Likert scale. The total number of participants will be 300 for children with ASD and 50 for typically developing (TD) children. The paper also analyzes acoustic and linguistic features extracted from speech data collected and completed for annotation from 73 children with ASD and 9 TD children to investigate the characteristics of children with ASD and identify significant features that correlate with the clinical scores. The results reveal some speech and linguistic characteristics in children with ASD that differ from those in TD children or another subgroup of ASD categorized by clinical scores, demonstrating the potential for developing automatic assessment systems for SCS and PP.
The advancement of Automatic Pronunciation Assessment (APA) systems has been significantly improved by Self-supervised Learning (SSL) models. However, despite these performance gains, there remains a lack of systematic research on effective utilization of SSL models and the explainability of their behavior in APA. This study aims to evaluate pronunciation with high accuracy using SSL models and to provide explanations for the scoring outcomes. To achieve this, we fine-tune various SSL models using multiple strategies, comparing their performance through extrinsic analysis to identify the key factors influencing performance improvements. Furthermore, intrinsic analysis is conducted using Principal Component Analysis (PCA) to gain insights into the model’s scoring patterns. Extrinsic analysis highlights the importance of strategic fine-tuning and acoustic similarity between fine-tuning and pre-training datasets. Intrinsic analysis reveals that different SSL models focus on distinct pronunciation features, with the Wav2Vec2.0 model capturing more advantageous information for APA. This study presents the first in-depth analysis of SSL models in APA, proposing a novel intrinsic analysis method based on feature distribution manifolds. We provide model-specific fine-tuning guidelines for APA tasks and recommend appropriate SSL models based on specific pronunciation assessment goals. This research significantly contributes to the future development of explainable APA systems based on SSL models.
This study introduces an automatic assessment model for speech production skills of children with cochlear implants (CIs) to support home-based speech therapy. The model employs acoustic embeddings from self-supervised models and considers speech traits of both normal hearing (NH) adults and children, which is a novel method for evaluating speech of children with disorders. It combines phoneme embeddings and two acoustic embeddings from Wav2Vec2.0 models, each trained on the speech of NH adults and children, via multi-head attention. Using a speech corpus of Korean-speaking children with CIs, our model outperforms single-embedding methods in a Pearson correlation coefficient between predicted and expert-rated scores, with a relative improvement of 51%. The results highlight the effectiveness of Wav2Vec2.0 acoustic embeddings and the importance of incorporating both of typical speech patterns of NH adults and children in assessing speech production skills in children with CIs.
Children with autism spectrum disorder (ASD) frequently encounter challenges in social communication and interaction, which necessitates continuous, comprehensive interventions to enhance their communication skills. Despite increasing interest in digital therapeutics (DTx), research on speech-utilizing interventions for children with ASD remains limited. This study introduced speech-based technologies integrated into DTx software designed to support the development of communicative skills in children with ASD. We compiled a large speech corpus from both children with ASD and typically developing children, which included clinical scores on social communication severity and speech production, rated by certified speech and language pathologists. Then three speech-based technologies were developed: automatic speech recognition for verbal interaction within the DTx, an automatic assessment model for social communication severity to monitor progress, and an automatic speech production assessment model to facilitate speech production skills. The results were promising, demonstrating a syllable error rate of 12.36
Chronic kidney disease (CKD) causes a continuous decline in kidney function and structural damage to the kidneys. The speech characteristics of CKD speakers will be different from those of non-CKD speakers because the typical characteristics of CKD, which are impairment of respiratory and laryngeal muscles, can affect respiration, the primary source of speech. In this paper, we identify the glottal characteristics of CKD speech and then investigate whether CKD can be automatically detected using the glottal features. Statistical analysis shows significant differences between groups in glottal source features, representing the breathy characteristic of CKD speech. Through the classification experiment, we compare the performance of solely using voice quality features (baseline) against additional glottal and spectral features. When glottal source features and voice quality features are used together, an F1-score of 0.88 with a 76% relative increase compared to the baseline is obtained.
Automatic assessment of dysarthric speech is essential for sustained treatments and rehabilitation. However, obtaining atypical speech is challenging, often leading to data scarcity issues. To tackle the problem, we propose a novel automatic severity assessment method for dysarthric speech, using the self-supervised model in conjunction with multi-task learning. Wav2vec 2.0 XLS-R is jointly trained for two different tasks: severity classification and auxiliary automatic speech recognition (ASR). For the baseline experiments, we employ hand-crafted acoustic features and machine learning classifiers such as SVM, MLP, and XGBoost. Explored on the Korean dysarthric speech QoLT database, our model outperforms the traditional baseline methods, with a relative percentage increase of 1.25% for F1-score. In addition, the proposed model surpasses the model trained without ASR head, achieving 10.61% relative percentage improvements. Furthermore, we present how multi-task learning affects the severity classification performance by analyzing the latent representations and regularization effect.
Empirical studies report a strong correlation between pronunciation proficiency scores and phonetic errors in non-native speech assessments of human evaluators. However, the existing system of computer-assisted pronunciation training (CAPT) regards automatic pronunciation assessment (APA) and mis-pronunciation detection and diagnosis (MDD) as independent and focuses on individual performance improvement. Motivated by the correlation between two tasks, we propose a novel architecture that jointly tackles APA and MDD using CTC and cross-entropy criteria with a multi-task learning scheme to benefit both tasks. To leverage additional knowledge transfer, Wav2Vec2-robust finetuned on TIMIT is used for the joint optimization. The integrated model significantly outperforms single-task learning, with a mean of 0.057 PCC increase for APA and 0.004 F1 increase for MDD on Speechocean762, which reveals that proficiency scores and phonetic errors are correlated for both human and model assessments.
The primary objective of this study is to identify the types of errors made by Korean college students in an oral proficiency interview in relation to specific task topics, and to examine how these errors affect their lexico-grammatical proficiency scores. Ninety-six two-minute-long audio clips of 32 Korean college students on three different topics were transcribed. Lexico-grammatical errors were then coded for statistical analysis and lexico-grammatical scores were estimated using many-facet Rasch measurement analysis with two raters. Friedman tests and Wilcoxon signed-rank tests showed that noun phrase, verb phrase, and prepositional phrase errors were more frequently found with the descriptive tasks than compare-and-contrast or hypothetical prompts. A hierarchical multiple regression analysis showed that noun phrase and verb phrase errors accounted for 22% of the variance in lexico-grammatical scores. Adding utterance length variables to the initial regression model explained an additional 43% of the variance in the lexico-grammatical scores. These findings suggest that noun phrase errors and verb phrase errors should be a priority in English classes, and that it is beneficial to teach English speaking skills in a way that takes into account the task characteristics and contextual factors.