
The study of speech in different emotional, psychophysiological and cognitive states is an important task for the development of speech systems. Cognitive load is the load on a person's cognitive system when performing a task. This paper analyses speech characteristics and facial features that can serve as the markers of cognitive load. Previous research revealed that cognitive load is associated with increasing fundamental frequency (F0), laryngealization, narrowing F0 range, changing articulation rate. Cognitive load can also be recognized using head pose, eye gaze and facial expressions. Two experiments were conducted in order to study speech and facial movements under cognitive load. During the first experiment, the participants played a driving simulator game and answered general knowledge questions simultaneously. Audio and videosamples were recorded. The information about action units (facial muscle movements) was obtained using Open Face 2.2.0. The results revealed that the most frequent visual characteristics of cognitive load are turning eyes to the right (AU62) and dimpler (AU14, the contraction of the buccinator muscle). In the second experiment, three episodes of a talk show were studied. The interviewer was driving a vehicle and conducting an interview as a dual task. The results showed that the most common visual markers of cognitive load are AU01 (inner brow raising), AU02 (outer brow raising), AU05 (upper lid raiser), AU10 (upper lip raiser), AU15 (lip corner depressor). The findings in both experiments suggest that cognitive load could be recognized by movements in the eye area and lip area.
In the field of automatic speech recognition (ASR), state-of-the-art results are achieved by end-to-end models. These models are sequence-to-sequence models and are trained using pairs of speech and corresponding texts, which implies that additional finetuning of underlying language models is not possible. In this paper we demonstrate that the performance of Serbian Whisperbased ASR can be improved by leveraging data generation with a high quality text-to-speech (TTS) system in Serbian. Synthetic speech is produced based on text extracted from Serbian web-scale text corpus, SrWAC, using data curation and large language model (LLM)-based normalization to mitigate problems in rendering Serbian pronounciation. A total quantity of 1500 h of speech is generated exploiting 9 text-to-speech voices based on deep-neural architectures and neural vocoding. The experiments are conducted on the medium Whisper model. The baseline model is initially finetuned using 1300 h of transcribed data and then additionally finetuned by synthetic speech, during which process the encoder section of the system is kept frozen. The experimental results confirm the character and word-error rate improvements on the CommonVoice database, as well as on real-life recordings.
Recent advancements in text-to-speech (TTS) technology have revolutionised automatic speech recognition (ASR) data augmentation in low-resource settings. In particular, only a few public datasets are available for dysarthric ASR (DASR) and text-to-dysarthric-speech (TTDS) models have addressed data sparsity limitations by increasing training data samples and diversity. In this context, Grad-TTS (G-TTS) has been shown to synthesise speech with accurate dysarthric speech characteristics beneficial for DASR data augmentation; likewise, Matcha-TTS (M-TTS) has recently improved on typical speech synthesis baselines. Recent studies commonly focus on data augmentation (i.e. reference data combined with additional synthetic data). This work analyses Whisper DASR model adaptation performance using reference data and G-TTS & M-TTS generated data, and shows that comparable performance can be achieved using synthesised data only relative to reference data. Additionally, despite growing work on dysarthric data augmentation, the validation of typical TTS metrics for synthetic dysarthric data, and the development of TTDS metrics requires further research. Results of this work show that gold standard metrics for typical TTS and current dysarthric speech assessment metrics lack sensitivity to predict DASR performance and hence a phoneme posteriorgram (PPG) distance based on the Jensen-Shannon divergence (JS) as a metric for dysarthric speech synthesis is introduced, showing correlation with downstream word error rate (WER) scores.
Given the challenges associated with obtaining large volumes of real-world audio, the advancement of Automatic Speech Recognition (ASR) systems increasingly depends on synthetic data. This study focuses on the call center-banking domain, providing a targeted analysis of ASR models in this context. We evaluate real and synthetic datasets, using Word Error Rate (WER) and Character Error Rate (CER) as metrics, and employ a speech quality model to assess the impact on ASR performance. Our research compares Text-to-Speech (TTS) and advanced voice conversion methods, including KNNVC, Seed-VC, and Vec2Wav2, revealing significant improvements in speech quality and ASR accuracy, particularly with Seed-VC. Additionally, our domain-specific experimentation provides insight into the unique challenges and opportunities that arise when applying ASR technologies to industry-relevant settings. This highlights the potential of voice conversion technologies to enhance ASR systems, guiding future research in diverse linguistic scenarios, and paving the way for the broader application of ASR innovations across various fields.
The language-specific domain adaptation problem refers to speech processing in resource-constrained embedded systems when pre-trained large models cannot be applied. This paper investigates language-specific adaptation strategies for automatic text-independent speaker recognition on an open set of speakers for various languages. MobileNetV3 was chosen as the most common model designed for edge applications and achieves a good accuracy-efficiency balance by using depth-wise separable convolutions to reduce the number of parameters and computations. The model was pre-trained in English and investigated for cross-language domain adaptation for German, French, Italian, Russian, Spanish, Dutch, and Chinese. We propose a combination of transfer learning and fine-tuning techniques to successfully adapt speaker verification models to a particular language. The proposed approach is validated using the CommonVoice cross-language dataset. The results demonstrate a notable improvement in the average EER up to 6
Today’s voice assistants remain fundamentally constrained by their stateless architecture, where each exchange is treated as an isolated incident, precluding meaningful long-term personalization. This limitation results in repetitive, context-blind dialogues that degrade user experience. This paper introduces an architectural blueprint for a lightweight, retention-augmented voice assistant, designed as a proof-of-concept to address this challenge. Our architecture prioritizes user privacy and transparency through on-device Automatic Speech Recognition via Whisper and a human-readable, file-based memory system, using a zero-shot Natural Language Understanding model (Google Gemini) for rapid prototyping. To rigorously test our retention mechanism, we introduce the Personalization Success Rate (PSR) as a novel evaluation metric. In a controlled evaluation with 150 scripted scenarios, our system achieved an 88
In this paper, we describe advances made in transcribing speech from witnesses and judges in Bastarache and Charbonneau commissions in Quebec. This is a rich audio with many speakers speaking fluently in conversational style. We extract features from SSL models to fine tune hybrid HMM/DNN models and also end-to-end models to compare word error rate (WER) on Bastarache and Charbonneau commission test sets. We also try SSL features extracted at both 10 ms and 20 ms temporal resolution for comparison. In a previous paper, we showed that we reduce WER for all the 15 low resource languages in OpenASR21 evaluation (with only 10 h of training audio), when we increase the temporal resolution of feature parameters computed from the speech SSL models from 20 ms to 10 ms. In this paper, we experiment with Quebec French data with 10 ms and 20 ms temporal resolution. For a training set of 472 h, we show that we still benefit from increasing the temporal resolution from 20 ms to 10 ms. Also, hybrid DNN/HMM models give lower word error rate (WER) than end-to-end speech recognition with 472 h of training audio. With a training set over 1000 h of audio, the end-to-end ASR system gives similar WER as the hybrid DNN/HMM system, and there is no significant improvement with increasing temporal resolution. We also compare our results with Whisper, an automatic speech recognition (ASR) system trained on 680,000 h of multilingual and multitask supervised data collected from the web.
This study aims to improve the accuracy of Hungarian automatic speech recognition (ASR) by applying large amounts of Hungarian training data both for self-supervised learning (SSL) and traditional supervised learning methods. In our experiments, the effectiveness of self-supervised pretraining on both smaller public and larger proprietary datasets was tested. Introducing SSL techniques to small Hungarian training sets resulted in noticeable improvements in model accuracy. When fine-tuning on large datasets containing thousands of hours of Hungarian speech, SSL accelerated training convergence, but fine-tuned models pretrained in English in a supervised way could not be outperformed in terms of word error rate. However, models trained or fine-tuned on a larger-than-ever purely Hungarian dataset achieved state-of-the-art accuracy across multiple independent evaluation sets.
Articulatory synthesis generates speech by modeling vocal tract configurations, but estimating articulatory parameters from audio-the acoustic-to-articulatory inversion (AAI) problem-remains challenging due to data scarcity, ambiguity, and the limitations of optimization-based methods. We propose PinkVocalTransformer, a Transformer framework that reformulates AAI as a sequence-to-sequence classification task over 44-dimensional vocal tract diameter sequences derived from the Pink Trombone physical synthesizer. By modeling complete tract shapes rather than higher-level articulatory trajectories, our approach yields a more interpretable and spatially consistent representation. To enable supervised learning, we generated over four million synthetic audio–parameter pairs under controlled static configurations. HuBERT embeddings improve feature extraction and robustness to real audio inputs. Reformulating regression as classification helps mitigate convergence issues arising from multimodal parameter distributions, leading to more stable predictions. Since ground-truth articulatory data are unavailable for real recordings, we regenerate audio from predicted parameters to indirectly evaluate reconstruction quality. Experiments show PinkVocalTransformer outperforms VAE-based and optimization baselines in vowel reconstruction. Objective ViSQOL metrics and ABX listening tests confirm higher perceptual similarity and listener preference for the regenerated audio compared to baselines. While the model performs strongly on static and simple dynamic segments, future work will focus on extending coverage to more diverse articulatory transitions and adapting the framework to more complex vocal tract models. Overall, this approach provides an efficient, data-driven framework for recovering interpretable articulatory parameters from audio, demonstrating both improved reconstruction quality and perceptual similarity compared to existing baselines.
In recent years contrastive learning has become the prevalent approach to training embedding models used in tasks such as information retrieval. Although, due to the nature of contrastive learning, during which negative samples are pushed further apart and positive samples are brought closer together in the embedding space, it imposed several new challenges for effective training. One such challenge lies in the creation of adequately hard negative samples. Most commonly hard-negative samples are created automatically by leveraging an existing retrieval model, which in turn introduces the risk of generating false negative training samples. To combat that, several filtering approaches that rely on ranks or similarity scores have been proposed. The fundamental flaw in those approaches lies in the ambiguity of relevance judgment given by the ranking models. Given only similarity scores, it is hard to accurately determine whether a particular document constitutes a positive or a negative sample. Our proposed method uses LLMs to resolve that issue by determining an optimal cutoff position through generating binary relevance labels. Such an approach also largely mitigates the problems with determining optimal filtration parameters present when dealing with raw relevance scores. The effectiveness of our method has been tested empirically by fine-tuning open-source retrieval models BGE-Reranker-v2-m3 and multilingual-e5-base. Our experiments on publicly available datasets have shown improvements in ranking metrics up to 2,29
The Google Book Ngram corpus is captivating due to its incredible size and availability. It has been widely used in studies of culture, social psychology, language evolution and others. However, as some researchers state, it suffers a number of limitations. Apparently, the most serious limitation of the corpus is its genre imbalance and lack of information on the genre composition of the books included in it. In this paper, we developed an algorithm for estimating the genre composition of the Google Books Ngram corpus. To estimate the percentage of texts of different genres, we used data on relative frequencies of a large range of words. Both linear models and multilayer feedforward neural networks were tested as predictors. To train the predictors, we used random subsamples of texts from the COHA corpus which are marked up by genres. We obtained estimates by using linear predictors and neural network predictors which showed different effectiveness. To assess the achieved accuracy, a cross-validation was performed. The analysis showed that the standard deviation of the neural network estimates obtained from annual data is no worse than 2–2.2
Many real-world NLU deployments require coverage of thousands of dynamically evolving topics beyond what manually curated data can sustain. This paper presents an automated pipeline deployed at Technische Universität Berlin (TU Berlin) to extract and train an intent classification model over 2,600 topics with zero manual annotations. We evaluate the effectiveness and generalizability of this pipeline using both curated and LLM-generated data, showing that our automatically trained models surpasses manually engineered systems in some dimensions. While lacking standardized benchmarks, we use a combination of internal evaluation strategies to provide a well-rounded assessment of model robustness. These findings support the use of scalable, self-maintaining NLU systems in complex and dynamic information environments.
The study investigates the rhythmic organization of speech in two discourse types-reading and spontaneous speech-among native English speakers from Australia and New Zealand, represented by 38 speakers (19 from each country). The audio recordings, totalling 02 h 46 m 01 s, are taken from the IDEA corpus [11], balanced for gender, age, and dialect of English. Applying the set of eleven metrics collected in the Correlatore software [12], we discovered a statistically validated differentiation in rhythmic patterns primarily between discourse types rather than the dialect, which confirmed the existence of 'rhythmic diglossia'. Thus, the rhythm class division of languages into syllable-timed and stress-timed categories, associating English with the stress-timed type, becomes questionable due to the influence of various factors, including dialect and discourse type.
This paper introduces an interpretable, multimodal approach for detecting cognitive disorders, specifically depression and Parkinson’s disease, using non-medical video data from the WSM dataset. Addressing the critical need for explainability in automated health assessment, this study combines interpretable audio, visual, and textual features to bridge the gap between diagnostic accuracy and transparency. Our methodology utilizes acoustic features (eGeMAPS), linguistic and prosodic features (BlaBla), and visual cues (facial landmarks, pose, and personality/emotion traits from OCEAN-AI framework) extracted from spontaneous speech and video recordings. Classical machine learning models, such as Logistic Regression, SVM, Decision Trees, Random Forests, are employed for classification, with performance benchmarked against neural network-based models. Experiments demonstrate that interpretable feature ensembles achieve competitive results, reaching up to 77.8
This study deals with the neural correlates of internet meme perception using EEGs with a focus on prefrontal brain regions associated with complex cognitive and emotional processing. Twenty-one adult subjects watched memes and control stimuli with congruent and incongruent text-image pairs presented sequentially to isolate the periods of cognitive insight. Wavelet analysis revealed a specific pattern of brain activity around the fifth second of stimulus presentation—approximately one second after the image onset—within the alpha (11–14 Hz) and beta (22–28 Hz) frequency bands. This activity was localized in both the orbitofrontal and dorsolateral prefrontal cortices and was significantly noticeably more manifested for memes than for control stimuli. These results suggest that memes trigger neural responses associated with insight and semantic reinterpretation. Additionally, a language-dependent effect was observed: memes in Russian evoked greater dorsolateral prefrontal activity, while English memes entailed stronger orbitofrontal activation. These findings indicate that meme comprehension activates distinct neural mechanisms depending on the linguistic and cultural context of the viewer.
This study presents the first systematic evaluation of automated speech recognition (ASR) systems for assessing the intelligibility of Russian-language esophageal voice (EV) in voice and speech rehabilitation following surgical treatment of laryngeal cancer. EV, produced without vocal folds, poses significant acoustic and articulatory challenges for speech recognition systems. We investigate two ASR platforms-Caesar-R, optimized for Russian, and Google Cloud Speech-to-Text, a multilingual cloud-based system-by comparing their transcription performance across three groups: healthy speakers, patients after oral cancer surgery, and post-laryngectomy patients using EV. Speech material included standard diagnostic phrases, phonetically balanced sentences, and vowel phonations. Recognition quality was measured using the Levenshtein distance to quantify transcription errors. Results show that both systems can process EV, though accuracy decreases proportionally to the extent of surgical intervention. Notably, some EV samples were recognized without errors, demonstrating feasibility for objective assessment. These findings support the use of ASR as a scalable, non-invasive tool for tracking EV intelligibility in Russian-language rehabilitation, including offline applications in settings with limited internet access.
The present study endeavors to identify patterns of cross-modal perception in modern Mongolian, with a particular emphasis on analyzing stable correspondences between the acoustic features of vowel sounds and their associated colour perceptions. To this end, a comprehensive acoustic-perceptive experiment was conducted, incorporating both isolated vowel forms and those used within a contextual consonant-vowel-consonant (CVC) structure. The acoustic analysis (comprising 84 isolated recordings and 1,512 context-based recordings) allowed us to determine the values of the first two formants (F1, F2), which reflect tongue openness and its front-back positioning. Perceptual data collected in both offline and online formats (n = 657; 24,470 responses) revealed statistically significant sound-colour associations that were consistently reproducible across diverse experimental settings. These findings support the hypothesis regarding the correlation between acoustic features and prevailing colour associations: vowels characterized by high F1 values were predominantly linked to warm colours, whereas low F1 corresponded to cold colours. On the contrary, cold shades predominated for high F2 values, while warm shades were more common for low F2 values. Moreover, the study identified the influence of variable factors such as phonetic context (CVC structure), as well as the participants' gender and age, on perceptual outcomes.
This paper addresses issues of modeling Karelian-Russian code-switching for automatic speech recognition, with a focus on intra-word code-switching. Due to grammatical differences between Karelian and Russian, and the lack of automatic translation tools for languages in question, standard augmentation methods relying on parallel translated text corpora are difficult to apply. To address these issues, we developed a set of rules specifically designed for generating words with intra-word code-switching, and then augmented the Karelian text by substituting random words with their corresponding generated counterparts. Besides that, we performed linear interpolation of the Karelian language model with the Russian one. We fine-tuned Wav2Vec2.0-large-uralic-voxpopuli-v2 on both Karelian and Russian speech data with the further integration of the developed language model into the system. An evaluation demonstrates significant accuracy improvement: compared to the baseline system without a language model, we achieved relative WER reductions of 11.3
This paper presents the results of the study of the features of emotion manifestation (joy, neutral state, sadness, anger) in the voice and facial expressions of 12–14 year old adolescents with intellectual disabilities (ID) compared to typically developing (TD) peers. ID are qualitative and quantitative deviations in the development of mental abilities. The methods of perceptual analysis, instrumental acoustic analysis of speech and automatic analysis of facial expressions using FaceReader 8v software were used. It was found that adults better recognize joy and neutral state in the speech of adolescents with ID than TD. Gender differences in the manifestation of emotions were revealed: girls express emotions in speech more accurately than boys. Differences in the acoustic features of speech depending on the type of emotion expressed and gender were shown. Based on automatic analysis of facial expressions, it was found that TD adolescents more often demonstrate a neutral state, while adolescents with ID – joy.
The main goal of the current study was to test the TTS model Tacotron2 for generating intonation models according to N.B.Volskaya's classification. In order to achieve this goal, we consecutively solved various tasks. First, we selected the fully annotated corpus of Russian monological speech (CORPRESS) for the analysis, due to the intonation model markups and a thourough segmentation, as well as to the quality of recorded speech (the corpus consists of professional actors' and speakers' recordings). From this speech corpus we choose 4 male speakers recordings. Then, we modified the architecture of the Tacotron2 model in a way to face the challenge of intonation model classification and made data preprocessing, that included data preliminary statistical analysis and preliminary training, which showed the need of data augmentation for creating a well-equilibrate material in training and validating datasets. After this, we produced an additional training, which showed good results. Two auditory perceptual experiments were conducted. First experiment consisted of MOS evaluation test and resulted at 4.027 points. Second experiment provided data on sentence type recognition. A consecutive comparative acoustic and expert auditory analysis of natural and generated pitch patterns showed that various intonation models can be successfully reproduced, although the most resemblance is noticed for the models with an even tone. The results obtained provide new information on intonation synthesis perspectives and demonstrate a huge potential of using N.B. Volskaya's system for the annotation of the training dataset in order to obtain an effective synthesis of functional intonation models in Russian.