
The study of speech in different emotional, psychophysiological and cognitive states is an important task for the development of speech systems. Cognitive load is the load on a person's cognitive system when performing a task. This paper analyses speech characteristics and facial features that can serve as the markers of cognitive load. Previous research revealed that cognitive load is associated with increasing fundamental frequency (F0), laryngealization, narrowing F0 range, changing articulation rate. Cognitive load can also be recognized using head pose, eye gaze and facial expressions. Two experiments were conducted in order to study speech and facial movements under cognitive load. During the first experiment, the participants played a driving simulator game and answered general knowledge questions simultaneously. Audio and videosamples were recorded. The information about action units (facial muscle movements) was obtained using Open Face 2.2.0. The results revealed that the most frequent visual characteristics of cognitive load are turning eyes to the right (AU62) and dimpler (AU14, the contraction of the buccinator muscle). In the second experiment, three episodes of a talk show were studied. The interviewer was driving a vehicle and conducting an interview as a dual task. The results showed that the most common visual markers of cognitive load are AU01 (inner brow raising), AU02 (outer brow raising), AU05 (upper lid raiser), AU10 (upper lip raiser), AU15 (lip corner depressor). The findings in both experiments suggest that cognitive load could be recognized by movements in the eye area and lip area.
In the field of automatic speech recognition (ASR), state-of-the-art results are achieved by end-to-end models. These models are sequence-to-sequence models and are trained using pairs of speech and corresponding texts, which implies that additional finetuning of underlying language models is not possible. In this paper we demonstrate that the performance of Serbian Whisperbased ASR can be improved by leveraging data generation with a high quality text-to-speech (TTS) system in Serbian. Synthetic speech is produced based on text extracted from Serbian web-scale text corpus, SrWAC, using data curation and large language model (LLM)-based normalization to mitigate problems in rendering Serbian pronounciation. A total quantity of 1500 h of speech is generated exploiting 9 text-to-speech voices based on deep-neural architectures and neural vocoding. The experiments are conducted on the medium Whisper model. The baseline model is initially finetuned using 1300 h of transcribed data and then additionally finetuned by synthetic speech, during which process the encoder section of the system is kept frozen. The experimental results confirm the character and word-error rate improvements on the CommonVoice database, as well as on real-life recordings.
Recent advancements in text-to-speech (TTS) technology have revolutionised automatic speech recognition (ASR) data augmentation in low-resource settings. In particular, only a few public datasets are available for dysarthric ASR (DASR) and text-to-dysarthric-speech (TTDS) models have addressed data sparsity limitations by increasing training data samples and diversity. In this context, Grad-TTS (G-TTS) has been shown to synthesise speech with accurate dysarthric speech characteristics beneficial for DASR data augmentation; likewise, Matcha-TTS (M-TTS) has recently improved on typical speech synthesis baselines. Recent studies commonly focus on data augmentation (i.e. reference data combined with additional synthetic data). This work analyses Whisper DASR model adaptation performance using reference data and G-TTS & M-TTS generated data, and shows that comparable performance can be achieved using synthesised data only relative to reference data. Additionally, despite growing work on dysarthric data augmentation, the validation of typical TTS metrics for synthetic dysarthric data, and the development of TTDS metrics requires further research. Results of this work show that gold standard metrics for typical TTS and current dysarthric speech assessment metrics lack sensitivity to predict DASR performance and hence a phoneme posteriorgram (PPG) distance based on the Jensen-Shannon divergence (JS) as a metric for dysarthric speech synthesis is introduced, showing correlation with downstream word error rate (WER) scores.
Given the challenges associated with obtaining large volumes of real-world audio, the advancement of Automatic Speech Recognition (ASR) systems increasingly depends on synthetic data. This study focuses on the call center-banking domain, providing a targeted analysis of ASR models in this context. We evaluate real and synthetic datasets, using Word Error Rate (WER) and Character Error Rate (CER) as metrics, and employ a speech quality model to assess the impact on ASR performance. Our research compares Text-to-Speech (TTS) and advanced voice conversion methods, including KNNVC, Seed-VC, and Vec2Wav2, revealing significant improvements in speech quality and ASR accuracy, particularly with Seed-VC. Additionally, our domain-specific experimentation provides insight into the unique challenges and opportunities that arise when applying ASR technologies to industry-relevant settings. This highlights the potential of voice conversion technologies to enhance ASR systems, guiding future research in diverse linguistic scenarios, and paving the way for the broader application of ASR innovations across various fields.
The language-specific domain adaptation problem refers to speech processing in resource-constrained embedded systems when pre-trained large models cannot be applied. This paper investigates language-specific adaptation strategies for automatic text-independent speaker recognition on an open set of speakers for various languages. MobileNetV3 was chosen as the most common model designed for edge applications and achieves a good accuracy-efficiency balance by using depth-wise separable convolutions to reduce the number of parameters and computations. The model was pre-trained in English and investigated for cross-language domain adaptation for German, French, Italian, Russian, Spanish, Dutch, and Chinese. We propose a combination of transfer learning and fine-tuning techniques to successfully adapt speaker verification models to a particular language. The proposed approach is validated using the CommonVoice cross-language dataset. The results demonstrate a notable improvement in the average EER up to 6