
With the continuous advancement of speech synthesis algorithms, synthetic speech has become increasingly realistic, posing unprecedented challenges to speaker recognition (verification) systems and information security. To address the escalating security risks caused by high-fidelity synthetic speech, this paper proposes a synthetic speech detection model named Group Delay and Latent Attention enhanced XLS-R+Conformer (GDLA-XLS-R+Conformer), based on the previously successful XLS-R+Conformer baseline. The model incorporates group delay to supplement phase features, thereby enhancing the ability to detect phase anomalies in synthetic speech. Furthermore, it employs a Symmetrical Multi-head Latent Cross Attention (SMLCA) and a Learnable Gating Fusion (LGF) to achieve tight integration of waveform and phase information. Experimental results show that the proposed GDLA-XLS-R+Conformer outperforms existing methods on multiple datasets, including ASVspoof 2019 logical access (LA), ASVspoof 2021 LA, and ASVspoof 2021 deepfake (DF), demonstrating its effectiveness and robustness across diverse scenarios.
Purpose This study aimed to identify critical acoustic-prosodic features for emotional expression in Mandarin Chinese using a validated disyllabic corpus and interpretable machine learning techniques. Methods A validated audiometry corpus of 450 Mandarin disyllabic words spanning five emotional states (“Angry,” “Sad,” “Happy,” “Fearful,” and “Neutral”) was analyzed. Speech samples from four male and four female speakers underwent machine learning-driven acoustic feature extraction and classification to determine discriminative parameters for emotional prosody recognition. Results Machine learning classification revealed happiness as the most acoustically challenging emotion to recognize. Frequency-, spectral-, and amplitude-related features (e.g., top and bottom pitch quintiles, pitch range, F1 frequency, spectral slope, loudness, harmony to noise ratio (HNR), Mel-frequency cepstral coefficients (MFCCs)) emerged as primary discriminators, with no temporal features selected. Fear and sadness exhibited distinct pitch range modulation reflected via the 20th and 80th F0 percentiles, while anger demonstrated hybrid acoustic signatures combining multiple characteristics. Conclusion The research demonstrates robust acoustic discriminability exhibited by the CESAP emotional speech corpus, characterizing systematic emotional prosodic patterns. This investigation reveals novel empirical insights through Mandarin disyllabic word materials, highlighting distinctive F0-related signatures due to the interaction between vocalizations of lexical tone and emotions. This study with the CESAP corpus contributes a valuable resource for both clinical applications and experimental investigations in affective vocalization research.
The fusion of Large Language Models (LLMs) with speech capabilities is rapidly reshaping multimodal AI, yet robust systems for low-resource domains like Hindi Automatic Speech Recognition (ASR) remain a formidable challenge. We propose AdaSpeech, a modality adapter-based pipeline extending the SpeechLLM paradigm to connect frozen speech encoders & LLMs for ASR in Hindi. While MLP/transformer adapters are common, our core innovation is a multiresolution convolutional adapter. This design leverages convolutions’ inherent inductive bias for capturing hierarchical, local patterns across varied temporal scales, enabling better cross-modal alignment and effective context capture from speech with significantly fewer parameters. Our extensive analysis also explores various choices for adapter architectures, LLMs, and encoders. Through large-scale experiments, we demonstrate (1) substantial WER improvements, especially when using domain-specific LLMs with shallow adapters; (2) larger LLM sizes help boost accuracy, and (3) ASR-fine-tuned speech encoders paired with our adapter yield the best results. This study provides valuable insights into building scalable, data-efficient ASR systems for Indic languages.
Accurate and early detection of dysarthria severity from speech can enable timely and non-invasive clinical interventions. In this study, we investigate the impact of embedding layer depth on classification performance using two self-supervised speech models—wav2vec2-BASE and DistilALHuBERT. A layer-wise comparison is conducted across four layers (1, 4, 7, and 12) using three datasets: UA Speech, TORGO (English), and the Tamil Dysarthric Speech Corpus (TDSC). Our results show that initial and intermediate layers consistently outperform final-layer embeddings for both dysarthria detection and severity classification. Under the stricter Leave-One-Speaker-Per-Class-Out (LOSPCO) evaluation, DistilALHuBERT achieves 49.9%, 53.9%, and 54.6% accuracy on UA-Speech, TORGO, and TDSC, respectively, demonstrating competitive performance compared to wav2vec2-BASE. These findings suggest that intermediate-layer embeddings better capture articulatory features relevant for dysarthria assessment and highlight the importance of evaluation protocol design for reliable model generalization.
Non-intrusive speech quality assessment aims to automatically predict the perceived quality of speech signals without access to the original reference audio. This study introduces a prompt learning approach, where task-specific textual prompts related to speech quality are constructed to guide the model in learning relevant perceptual features. Based on the Contrastive Language–Audio Pretraining (CLAP) model, three types of prompt strategies — single prompt, opposite prompt, and multiple prompts — are proposed. The model estimates the quality of a given utterance by computing the embedding similarity between the prompt(s) and the input audio. Experiments conducted on multiple public datasets, including NISQA, VCC 2022, VCC 2018, and SOMOS, demonstrate that the multiple prompt strategy consistently outperforms baseline models LDNet and SSL-MOS w/o FT across evaluation metrics such as Pearson’s linear correlation coefficient (LCC), Spearman’s rank correlation coefficient (SRCC), Kendall’s Tau (KTAU), and root mean square error (RMSE). It also exhibits stable performance across various distortion types and acoustic scenarios. To further explore the generalizability of prompt learning, this work also incorporates prompt embeddings into the lightweight speech quality assessment model LDNet, forming the LDNet-prompt framework. Experimental results show that even without large-scale pretraining, the prompt mechanism can significantly enhance model performance, validating the feasibility of prompt learning for non-intrusive speech quality assessment. Overall, the multiple prompt strategy has demonstrated robust and stable performance in models based on CLAP and LDNet, and is regarded as the best prompt scheme for this task.
The study investigates the relationship between lexical complexity and subjective age of acquisition (AoA) for nouns and verbs across six regional dialects of Palestinian Arabic (PA) spoken in northern Israel. Sixty adult native PA speakers (Mage=36.9 years; Range: 25-62) representing six different PA dialects, participated in a picture naming task of 154 nouns and 124 verbs. The target words varied in lexical-phonological complexity measured by syllable length, and the number of intra-syllabic consonant clusters, inter-syllabic consonant clusters (consonant sequences), and consonant doublings. Participants further provided subjective AoA estimates. Our results revealed significant dialectal differences in AoA, but not in complexity. Moreover, significant differences in AoA and complexity were observed between nouns and verbs. Nouns were acquired earlier within all dialects, despite being longer and with more inter-syllabic clusters. Verbs had more consonant doublings. Testing the relationship between complexity and subjective AoA, a regression analysis showed that AoA was predicted by lexical complexity in some but not all of the dialects tested, offering partial support for the hypothesis that greater complexity corresponds to later lexical acquisition. For example, positive correlation analyses showed complexity effects for verbs; as verbs with more intra-syllabic clusters (in Wadi Sallami dialect) were acquired later. Likewise, for nouns, those with more inter-syllabic clusters were acquired later (in Kufur Kanna dialect). An exception is observed in the Msherfi dialect, where a negative correlation indicates that verbs with more inter-syllabic clusters were acquired earlier rather than later. The results have implications that extend beyond Arabic, offering insights into lexical acquisition in dialectal and multilingual contexts.
This study presents the SPEECH-TRACK dataset, developed to evaluate and enhance methods for monitoring speech development in children with special needs, including those diagnosed with autism spectrum disorder (ASD), speech delays, and other communication impairments. The dataset design, structure and validation are addressed and primary evaluations are performed to observe its competence in speech and engagement patterns research. Current speech monitoring tools often lack ecological validity and fail to capture naturalistic speech patterns. To address this gap, data were collected from 30 children using smart watches equipped with voice recognition technology, enabling the continuous recording of speech features such as frequency, pitch, and clarity across a variety of everyday environments. In addition to device-based recordings, naturalistic observations were conducted during play and social interactions to capture spontaneous communicative behaviour. Parental surveys were also administered to provide personalized context regarding each child's speech development and interaction patterns. All speech recordings were annotated by certified speech-language pathologists (SLPs) using two key classifications: 'Clear Speech/Unclear Speech' and 'Engaged/Not Engaged', based on behavioural cues and acoustic quality. Results from technical validation confirmed high-quality audio recordings across settings and strong inter-rater reliability among expert annotators. The dataset demonstrates robustness in capturing diverse speech behaviours and engagement levels, offering both quantitative acoustic data and qualitative interaction markers. In discussion, the SPEECH-TRACK dataset serves as a valuable resource for clinicians, educators, and researchers, facilitating tailored intervention strategies and longitudinal tracking of speech development. Its naturalistic design enhances the ecological validity of assessments and supports the creation of adaptive, context-sensitive speech therapies for children with communication challenges.
Unmanned Aerial Vehicle (UAV)-based audio acquisition integrated with speech enhancement enables high-quality long-range audio capture. However, sound source attenuation and UAV self-noise lead to extremely low SNR conditions, posing a critical challenge for speech enhancement. To address this, this study proposes a dedicated speech enhancement network paired with a multi-spectral joint loss function. These two components synergistically achieve robust speech enhancement under the scenario's inherent extremely low SNR. The proposed network exploits the correlation between noise references and noisy speech for efficient noise-speech separation and high-fidelity speech restoration. The loss function adopts a triple-constraint optimization to ensure spectral matching, mitigate energy fluctuation interference, and adaptively accommodate SNR variations across capture distances. Extensive simulations validate the feasibility and effectiveness of the proposed network and loss function for speech enhancement in UAV noise-contaminated environments.
Objectives This pilot study evaluated a rapid screening protocol integrating acoustic analysis and perceptual assessment to detect early signs of dysphonia among university professors, hypothesizing (1) a high prevalence of subclinical vocal abnormalities, (2) strong correlations between objective and perceptual measures, and (3) acceptable diagnostic accuracy of acoustic parameters for identifying perceptual dysphonia. Method Thirty-two professors (50% female, mean age 44.7 ± 8.3 years) from the University of Tipaza underwent voice assessment (sustained vowels and connected speech in Arabic). Acoustic parameters (Jitter PPQ5, HNR, GNE, CPPS) were analyzed with VOXplot. Two speech-language pathologists performed GRBAS-G perceptual ratings. ROC analyses determined the sensitivity/specificity of each acoustic measure against the dichotomized perceptual rating (G = 0 vs G ≥ 1). Ordinal logistic regression assessed the combination of CPPS and HNR for predicting severity. Results Acoustic abnormalities were prevalent: 75.0% exhibited subnormal HNR, 43.8% reduced CPPS, and 21.9% elevated Jitter. Perceptual ratings identified 37.5% with dysphonia (G ≥ 1). ROC analyses showed that CPPS had the highest diagnostic accuracy (AUC=0.85, 95% CI: 0.66–0.99), with 91.7% sensitivity and 85.0% specificity at a cut-off ≤14.4 dB. HNR (AUC=0.78) and GNE (AUC=0.82) also demonstrated fair to good accuracy, whereas Jitter performed poorly (AUC=0.30). The ordinal logistic regression model combining CPPS (OR=0.59, p = 0.063) and HNR (OR=0.81, p = 0.033) significantly predicted dysphonia severity (χ²(2)=17.0, p < 0.001; Nagelkerke R²=0.49), correctly classifying 78% of cases. Conclusion The protocol showed good convergent validity and promising diagnostic accuracy, particularly for CPPS. CPPS alone or combined with HNR shows potential for screening dysphonia in academic settings. These preliminary findings require validation in larger, independent cohorts.
Gamified biofeedback has the potential to help individuals improve speech motor patterns. For such systems, feedback provided in response to attempted changes must directly characterize performance improvements. This is particularly the case for ultrasound biofeedback therapy (UBT), which is increasingly used to visualize tongue movement for treating misarticulations of speech sounds (e.g., American English /ɹ/). Previous studies on ultrasound imaging of misarticulated /ɹ/ sounds have established that a single, time-dependent measured parameter, δ (relative difference between normalized tongue dorsum and blade displacements) can classify articulatory accuracy, matching human judgments with about 85% agreement. The δ parameter thus is an ideal basis for real-time gamified UBT. However, during ultrasound imaging of a speech production, there are multiple possible image frames from which δ could be automatically evaluated to represent production accuracy. Accordingly, to identify the optimal frame selection for δ evaluation, ultrasound image sequences depicting 2,944 productions of 15 distinct rhotic phrases were analyzed, while perceptual accuracy was judged by trained listeners. Four different selection criteria were compared, with classification performance and optimal thresholds evaluated using receiver-operating-characteristic curve analysis with 8-fold cross validation. For pre-vocalic and post-vocalic rhotic syllables, highest classification success rates resulted from evaluating δ at productions’ end. For pre-vocalic words preceded by the carrier phrase “a”, best classification resulted from evaluating δ at the estimated temporal midpoint of /ɹ/. Results provide a viable basis for gamified ultrasound biofeedback therapy based on real-time ultrasound measurements of tongue movements.
Distributed speech enhancement is of great necessity in applications with a new-generation technology for gathering and processing audio. To achieve distributed noise reduction while reducing data exchange among arrays, the utility-based weighted distributed adaptive node-specific signal estimation (UBW-DANSE) is proposed in unconstrained wireless microphone networks (UWMNs), where each array only performs a weighted summation once on the compressed signals from its neighboring arrays each time. First of all, based on average consensus, the relaxed simultaneous version of topology-independent DANSE with the sub-layer method is introduced for the UWMNs. Then, by deriving the relative utility using the minimized cost function referred to as the utility, we propose a new average consensus weighting matrix. Finally, to minimize data exchange between arrays, the UBW-DANSE method that can be directly applied to UWMNs is presented, which serves as a trade-off between noise reduction performance and energy savings. The simulation experimental results demonstrate the effectiveness of the proposed method.
Flow matching (FM) has emerged as a pivotal paradigm for efficient generative speech enhancement (SE), which models continuous probability paths via ordinary differential equations (ODEs). Existing FM-based SE methods predominantly adopt a unified conditional vector field to jointly conduct deterministic speech recovery and stochastic perturbation generalization within a single framework. This modeling strategy increases the difficulty of model optimization and makes it hard to optimize speech fidelity and generalization at the same time. To mitigate this limitation, we propose FRSE (Flow-Matching with Residual-Decoupled Learning for Speech Enhancement), a novel framework that explicitly decouples these two objectives. Specifically, FRSE decomposes the single velocity field in conventional FM into two independent parallel fields: a residual velocity field dedicated to guiding the transformation of the probability path from noisy speech to clean speech for fine-grained speech structure restoration, and a noise velocity field focused on attenuating Gaussian perturbations and strengthening the model’s generalization ability. Experimental results show that the proposed residual-decoupled flow matching framework enables efficient sampling and high-quality enhancement for SE tasks, and the dual velocity fields perform significantly better than conventional single-field modeling. Evaluations on the VoiceBank-DEMAND and WSJ0-REVERB datasets demonstrate that FRSE achieves competitive performance and efficient low-step inference. With only three sampling steps, it outperforms mainstream flow-matching models and most diffusion models on key metrics including PESQ, ESTOI, and SI-SDR.
During speech perception, listeners extract both linguistic and socio-indexical information. A salient aspect of talker identity relates to their accent, including their language learning status (first- or second-language learner - L1 or L2). Adults are highly sensitive to the presence and strength of an unfamiliar accent. Children are also sensitive to the presence of an L2 accent, but may have less awareness of differences among L1 varieties. Children’s perception of accent strength has received relatively less attention. Furthermore, the acoustic-phonetic cues that children attend to when making accent judgements are not well identified. Here, we address these issues by testing 6- and 12-year-old children and adults on a ladder task in which listeners ranked L1 and L2 talkers of English relative to a local accent baseline. Both groups of children’s rankings were similar to adults’. To interpret the speaker-dependent and -independent stimulus factors driving accent rankings across ages, machine learning was employed. Phonetic differences were the largest contributor to accent rankings; intonation, rhythm, stress, sentence length and complexity also contributed similarly across age groups. This analysis demonstrates the promise of machine learning for identifying the stimulus factors that drive accent perception. This approach could be taken to isolate factors impacting other aspects of L1/L2 processing including intelligibility, comprehension, and listening effort across the lifespan. The results support the primacy of segmental differences in perceptual accent scaling, and elucidate the specific phonological accent components that may drive social decision making starting in the early school age years.
Automatic depression diagnosis is crucial for early detection and intervention but is often hindered by limited sample sizes and severe data imbalance. To address these challenges, this paper presents a novel framework, Population-based Dynamic Audio Augmentation (PDAug), which formulates the data augmentation process as a population evolution optimization problem. This unified framework jointly optimizes the selection of augmentation methods, parameter configurations, and model weights in a closed-loop learning process. PDAug adaptively determines the number and combination of augmentation operations based on the dataset’s Net Imbalance Rate (NIR) and feedback from classification performance. It incorporates six augmentation operations covering both time-domain and time–frequency domain transformations, allowing for the generation of diverse and task-relevant samples tailored to the dataset’s imbalance characteristics. To enable adaptive learning within a continuous augmentation parameter space, the Short-Time Objective Intelligibility (STOI) metric is introduced as a quantitative constraint. This facilitates the construction of a clinically meaningful and continuous parameter selection space, promoting more informed and interpretable augmentation decisions. Extensive experiments on the CMDC, DAIC-WOZ, and EATD-Corpus datasets demonstrate that PDAug achieves state-of-the-art performance on CMDC and yields competitive results on DAIC-WOZ and EATD-Corpus, highlighting its effectiveness in real-world, low-resource, and imbalanced scenarios.
The goal of speech enhancement is to improve the comprehensibility and quality of a speech of interest in noisy background environments. However, speech enhancement in real dynamic acoustic environments is exceptionally challenging. In this paper, an independent component extraction-based speech enhancement algorithm is proposed to enhance a moving source in highly reverberant and noisy environments. First, a new speech enhancement model with continuous separating vectors is constructed for time-varying mixing environments, where both the mixing and separating vectors are time-varying. Second, a contrast function based on the log-likelihood is optimized using the Newton–Raphson method, and the gradient and second-order derivatives of the contrast function are computed explicitly to obtain the updated rules for the separating vectors. In addition, a preprocessing is considered to get a modified update rule. Simulation experiments confirm the validity and competitiveness of the proposed algorithm for speech enhancement of both moving and static sources in highly reverberant and noisy environments, and demonstrate its robustness to real-recorded non-stationary noises.
Smile is a key feature in development and can also be heard in the voice. During development, an improvement of vocal emotion recognition abilities has been described along with global neuroanatomical maturation of auditory cortex. While it has been suggested by the predictive coding framework that internal representation refines with age during development, internal representation of vocal smile has not been characterized in children within this specific age range. The present study aims to bridge this knowledge gap by investigating internal representation of vocal smile in children, along with the robustness of this representation.Vocal smile model was assessed in 46 children aged 8 to 12 years and 42 adults. A reverse correlation approach was implemented allowing to distinguish two components of vocal smile model: the acoustic features underlying a vocal smile and the internal representation robustness, quantified as internal Noise.As a group, children display the same internal representation of vocal smile as adults. However, when comparing children’s individual representations to adults one, we demonstrated that within children group, their representation becomes increasingly closer to adults’ template with age. Furthermore, internal noise estimation allows to distinguish two children subgroup, with one subgroup exhibiting a delay in the acquisition of that internal representation.This study provides key insights into the development and variability of emotional voice processing in school-age children.
Consonants are crucial for effective communication but often pose challenges for individuals with hearing impairments or age-related hearing loss. Enhancing consonant perception improves intelligibility and reduces cognitive load; however, traditional methods like vowel suppression and spectral shaping often compromise speech naturalness. Japanese phonetics, with its predictable vowel-consonant structures, provides a unique context for studying consonant enhancement, particularly in synthetic speech, which frequently suffers from perceptual confusion and acoustic limitations. This study focuses on improving the intelligibility of /bi/ and /gi/, which are commonly misheard as /ri/ in synthetic speech. Using speech generated by Amazon Polly and Google Text-to-Speech, we applied various enhancement techniques, including amplitude amplification and waveform manipulation targeting specific acoustic regions. Through three experiments, critical time segments for consonant intelligibility were identified, and the effects of enhancements were evaluated. Results showed that amplifying the initial consonant region (R1) significantly improved intelligibility, whereas truncating mid-waveform regions (R2) reduced clarity. Platform-specific differences in processing algorithms also influenced the effectiveness of these enhancements. These findings underscore the importance of tailored signal processing strategies for synthetic speech. Future research should explore adaptive enhancements and extend investigations to more complex speech contexts, aiming to optimize synthetic speech systems for individuals with hearing impairments and older adults.