
In professional settings, conversations often involve persons with defined roles (doctor, patient, lawyer, client, etc.), and the intelligibility of a conversational transcript may be improved by annotating conversational turns with the role of the speaker, e.g. "Doctor: How are you feeling? Patient: I sprained my ankle." We propose a novel hybrid architecture that combines an ASR model augmented to label the speaker's role at each speaker change point with a d-vector-based diarization system. This system outperforms modular and fully integrated baselines by 12% and 28%, respectively. We also show that, when an ASR transducer model is trained to predict role or speaker-change tokens as part of the transcript, these token timings can improve diarization more than the adjacent word token timings can, despite there being no explicit training signal conveying precise speaker change points.
In this paper, we propose a novel multi-stream framework for automatic cued speech recognition (ACSR) that directly processes the upper-body region, addressing hand-lip asynchrony without requiring explicit segmentation or synchronization. Our model integrates two distinct modalities: (i) an appearance-based stream leveraging the ResNet18 for feature extraction and (ii) a skeletal stream based on a modulated graph convolutional network (GCN). For graph construction, we incorporate, for the first time in ACSR, 3D pose parameters inferred from the PIXIE model. Both modalities are coupled with temporal convolution for short-range dynamics learning and a BiGRU encoder for long-term sequence modeling. In addition, we introduce an alignment module that combines CTC with two auxiliary losses, improving each modality performance and enabling effective late fusion during inference. Our model achieves state-of-the-art performance across three benchmark datasets, demonstrating its effectiveness.
Understanding speech emotion through artificial intelligence (AI) is crucial for human-computer interaction and mental health monitoring. While audio large language models (ALLMs) excel in speech comprehension, they face challenges in accurately integrating emotional signals from acoustic and semantic features. Moreover, emotions often span dialogues, making sole reliance on current audio insufficient for comprehensive understanding. To address these challenges, we propose a novel emotion-aware audio large language model (EAA). Specifically, we design a dual cross-attention mechanism to fuse acoustic and semantic information for a more comprehensive emotional representation. Furthermore, we use context-aware instruction tuning by incorporating the current and immediately preceding utterances as contextual information, enhancing task understanding and emotion recognition. Our experimental results show that EAA outperforms existing ALLMs on the MELD dataset, improving accuracy by 11.4%.
In this paper, we created the first transcribed, parallel, balanced Farsi dataset (FaVC) that can be used for all tasks of speech synthesis, including both parallel and non-parallel voice conversion. FaVC is a balanced dataset that provides all phonemes of the Farsi language in all possible phonetic combinations. A metadata file is provided, including Farsi transcriptions, normal text, and IPA phoneme transcriptions of all audio files. The first Farsi voice conversion results using FaVC are also reported in this paper. The results indicate that by using FaVC, performance is as good as the chosen baseline methods for voice conversion in English. Moreover, objective evaluations of a voice conversion system that requires a parallel speech corpus can be performed using this dataset.
Alzheimer's disease (AD) is a progressive disorder that gradually affects memory, language, and reasoning, making early detection crucial for timely intervention. Traditional methods, like medical imaging and clinical evaluations, are costly and limit accessibility. To address this, we propose a speech-based AD detection method that leverages a co-attention mechanism to integrate multilevel acoustic and transcribed text features. These acoustic features include spectrograms, MFCC features, and wav2vec2 embeddings. The mechanism dynamically assigns weights to different features, enhancing their interaction and improving fusion. We also compare early and late fusion strategies to optimize integration. Tested on the ADReSSo dataset, our model achieves 83.15% accuracy, demonstrating the effectiveness of a well-structured integration of acoustic and transcribed text features for accessible and cost-efficient AD detection.
The study presents a comprehensive evaluation of the Montreal Forced Aligner (MFA) in aligning phone boundaries of Hong Kong Cantonese (HKC) spontaneous speech. We developed two tailored Cantonese MFA models, designed to address distinct Cantonese phonetic features, such as checked syllables. These models were applied to align the same set of recordings from spontaneous interviews, and their performance was compared against human annotations. Our results reveal that the updated Cantonese MFA models achieved decent alignment accuracy on spontaneous speech, with a satisfactory level of agreement with manually adjusted boundaries in vowels. However, Cantonese-specific features and connected speech process remain major challenges for the current models. This observation allows us to propose specific amendments to the models to improve alignment performance, as well as recommendations on manual boundary adjustments.
Recent fake audio detection methods often leverage large speech models to achieve robust speech representations. These models are typically very deep, providing multiple layer-wise representations. However, current works often rely solely on single layer representation or feature fusion to extract one utterance-level representation for decision making. These methods risk underutilizing rich information from multiple layers and might induce feature collapse. We propose a novel layer-wise decision fusion method that applies fusion after per-layer decision making and achieves the best cross-dataset performance on In-the-Wild dataset (EER 6.90%) compared to other strong baselines. Our model design also makes the model more transparent, allowing us to conduct detailed analysis to reveal the underlying mechanism of decision making.
Although recent text-to-speech (TTS) models based on flow matching have achieved remarkable generation quality, their reliance on numerous sampling steps hinders practical deployment. In this work, we introduce an adversarial post-training strategy for flow matching TTS that significantly reduces the required sampling steps. Our approach treats a pre-trained flow matching model as a few-step generator, optimizing it with reconstruction and adversarial objectives. We integrate this technique into APTTS, our novel latent flow matching framework for zero-shot TTS, and demonstrate its superiority over state-of-the-art baselines with real-time applicability. Furthermore, we validate the scalability of our adversarial post-training approach by applying it to Matcha-TTS, a publicly available flow matching model. Evaluations on a multi-speaker dataset show that our method enhances audio quality while reducing inference time, underscoring its potential as a scalable solution for real-time TTS.
As speech processing systems become more ubiquitous, the need for real-time, efficient speech quality prediction (SQP) is growing. Conventional artificial neural networks (ANNs) offer strong prediction performance but can be computationally demanding, which limits their deployment on mobile and edge devices. Spiking neural networks (SNNs) present a promising alternative for ultra-low-power, streaming inference due to their sparse activity and event-driven processing. However, their potential for SQP remains largely unexplored. This article introduces deep convolutional SNNs for SQP and evaluates their performance against state-of-the-art ANN models. Our results show that SNNs achieve comparable accuracy while significantly reducing computational cost. These findings highlight the potential of SNNs to enable real-time, energy-efficient SQP in resource-constrained settings.
In this study, we propose MH-SENet, which is designed for speech enhancement by extracting the temporal and spectral features of speech signals in parallel. MH-SENet, which is based on the U-Net architecture, has an encoder and decoder consisting of a bi-directional Mamba and processes it more precisely by considering all the context of the input sequence. Furthermore, a cross-domain Mamba-Transformer block is constructed between the encoder and decoder to effectively fuse information between each time and frequency domains. We evaluated the performance of our proposed MH-SENet on the VCTK + DEMAND dataset and thus it outperformed existing methods by achieving the highest PESQ score. Despite being a hybrid model, the proposed MH-SENet has a lower number of parameters compared to the conventional models.
Disentangling distinct types of information in speech representations is crucial for improving speech synthesis and voice conversion systems. In this work, we introduce LombardTokenizer, a neural speech codec able to separate features related to vocal effort from other acoustic (and semantic) information. This model is built on SpeechTokenizer, a model proposed in the literature based on multi-stage quantisation, which focused on isolating semantic content in its first quantisation layer. We show that the level of vocal effort can be effectively captured in the second quantisation layer by conditioning the quantisation layer with neural encoders trained to represent vocal effort. Experimental results demonstrate that the proposed method significantly outperforms existing methods in speech conversion between neutral and Lombard speech, while maintaining excellent speech synthesis quality, offering improved control over vocal effort and naturalness of synthesised speech.
Target-speaker automatic speech recognition (TS-ASR) utilizes speaker embeddings to identify a target speaker in multi-talker environments. While high-performance speaker embedding extractors provide discriminative embeddings, their computational demands limit practical deployment. In this study, we present two novel methods that effectively utilize lightweight extractors to enhance TS-ASR performance. First, we propose a multiple embeddings modulation that effectively transfers comprehensive speaker information to the ASR module, thereby improving overall performance and robustness against embedding variations. Second, we present a virtual speaker embedding augmentation technique that synthesizes embeddings of unseen speakers, reducing dependence on specific extractors while enhancing independent contributions from each extractor. Experimental results on the Libri2Mix dataset demonstrate that our proposed methods achieve significant WER reductions compared to the baseline model.
This study analyzes the power spectral density (PSD) of speech in healthy controls (HC) and Parkinson's disease (PD) patients, focusing on the 0-100 Hz range. These low frequency components are below the fundamental frequency and may reflect both motor and neural mechanisms in speech production. We hypothesize that neural oscillations (NOs) involved in speech perception and production - theta (4-8 Hz), beta (15-35 Hz), and gamma (36-80 Hz) - shape the low-frequency PSD. Since NOs are linked to motor control and cognition, and are altered in PD, we expect systematic differences between HC and PD speech. Using multitaper estimation, we found significant differences in beta power, in line with research on beta oscillations and motor dysfunction in PD. Beyond distinguishing HC from PD speech, our results suggest that sub-fundamental frequency information may reflect neural dynamics in speech production, offering new perspectives for speech pathology and neural oscillation research.
Recent studies have demonstrated the advantage of generative adversarial network (GAN)-based vocoders in high-fidelity speech synthesis and fast inference speed. However, they often suffer from audible artifacts such as aliasing and blurring. In this paper, we propose AF-Vocoder, a novel GAN-based vocoder that can synthesize high-fidelity speech with fewer artifacts. Specifically, we introduce a frequency-domain artifacts filter named GAFilter to achieve artifact removal. GAFilter incorporates a learnable frequency filter, which enforces a desired inductive bias of frequency control for artifact-free speech synthesis. Experimental results show that the proposed AF-Vocoder outperforms other GAN-based vocoders in speech reconstruction quality and artifact suppression on various datasets including out-of-domain speakers.
Speech Emotion Recognition (SER) in naturalistic conditions remains a challenging task due to the variability of emotional expression and class imbalances in the real world. As part of the Interspeech-25 SER challenge, we benchmark state-of-the-art large-scale self-supervised speech models on the MSP-Podcast corpus. To extract rich and expressive representations, we systematically investigate fine-tuning strategies, loss functions tailored to mitigate class imbalance, and pre-trained encoder layer freezing techniques to optimize performance. Our findings highlight the impact of these design choices on model robustness and generalization, offering practical guidance for developing SER systems that excel in real-world scenarios.
The differences in emotional expression present significant challenges for the development of effective emotion recognition systems. Although large language models (LLMs) have demonstrated strong performance, their generalization in emotion recognition tasks is often limited by variations in individual emotional expression, which remain inadequately addressed even by parameter-efficient fine-tuning techniques such as LoRA. To address these challenges, we proposes a multitask memory parameter-efficient fine-tuning method (MMLoRA) that enhances multimodal SER by incorporating gender as an auxiliary task. The method leverages shared LoRA experts to facilitate information exchange between tasks and employs a mixture of LoRA experts to process task-specific information. Additionally, a memory mechanism propagates task-specific information across layers. Experimental results demonstrate that the proposed MMLoRA significantly improves emotion recognition performance compared to vanilla LoRA.
In this paper, we propose Flexible Dynamic Encoder RNN (FDE-RNN), an innovative model capable of seamlessly switching between VAD and PVAD without incurring redundant resource consumption. In static PVAD modeling, performing VAD typically requires either merging categories or omitting speaker embeddings, often resulting in excessively large models that are impractical for VAD tasks. In contrast, FDERNN efficiently adapts by removing the personalization module when functioning as VAD, significantly reducing resource demands. Furthermore, on PVAD tasks, FDE-RNN leverages dynamic neural networks with a gating-based skipping mechanism, enabling it to bypass redundant computations during non-speech segments, further optimizing computational efficiency. Extensive experiments demonstrate that FDE-RNN outperforms all other prior arts on both PVAD and VAD tasks in terms of overall performance. Notably, when functioning as a VAD, FDE-RNN merely utilizes 30% of the parameters required by the competitive models, underscoring its remarkable efficiency and scalability.
Esophageal speech is an alternative speech production method for people who have undergone laryngectomy and often suffer from reduced intelligibility. This paper proposes three lightweight speech enhancement methods trained with the loss given by end-to-end automatic speech recognition models. The proposed methods are based on frame mask (FM), ideal ratio mask (IRM), and voice conversion (VC) techniques. The evaluations show that the enhanced speech produced by the proposed methods led to an improvement over the original esophagus speech, specifically in terms of speech recognition rates by automatic speech recognition systems and human evaluators, naturalness as assessed by mean opinion scores, and more detected voicing segments.
This perception study investigated the role of intonation, specifically pitch accent position and pitch accent type, in the interpretation of utterances as sarcastic. Participants from two regions in Germany, Freiburg and Trier, listened to short utterances such as Das sieht ja umwerfend aus ("That looks stunning") in seven prosodic conditions. The recordings were taken from a production study [1], and participants classified them as sarcastic or sincere. Results show that in both regions, irony perception is driven by (a) the existence of a prenuclear accent and (b) the type of the nuclear accent, particularly L*+H. Reaction times for ironic responses were shorter for L*+H, as well as when a prenuclear accent was present or when the subject carried the nuclear accent. These findings underscore the importance of intonation in irony perception and have implications for the processing of non-canonical meaning in general and the interpretation of sarcasm across varieties in particular.
Recent advancements in time-domain end-to-end neural speech codecs have significantly improved performance. However, existing codecs fail to fully exploit the correlations across different frequency bands in speech, leading to inefficiencies and reduced interpretability. In this paper, we introduce SPCODEC, a time-domain end-to-end neural speech codec featuring a latent split-and-prediction scheme. The model consists of a fully convolutional encoder-decoder and a group residual vector quantization module enhanced with a split-and-prediction mechanism. This mechanism disentangles low- and high-frequency representations and employs prediction to effectively reduce feature redundancy. SPCODEC achieves state-of-the-art MOS-POLQA scores of 4.0 at 6/8 kbps and 4.5 at 10.66/16 kbps for wideband and super-wideband speech, significantly outperforming both neural and traditional codecs.