
Session variability caused by channel changes and acoustic conditions often degrades the performance of i-vector-based speaker identification. This study applies data augmentation to mitigate such effects and improve system robustness. We compile a dataset from five Indonesian speech corpora and create six recording conditions: clean, additive noise, two levels of reverberation, gain adjustment, and speed perturbation. Each signal is preprocessed into 39 -dimensional MFCCs, a 128 component GMM-UBM is trained, 200 -dimensional i-vectors are extracted, and similarity is scored using Linear Discriminant Analysis (LDA) with cosine distance and Probabilistic LDA (PLDA). A factorial design yields eight configurations that combine training augmentation on or off with four overlap levels between training and enrollment data. Evaluation across four enrollment scenarios uses accuracy and Equal Error Rate (EER). Results show that augmentation does not always increase accuracy for LDA+cosine distance, with the best accuracy of 98.5 % obtained without augmentation. In contrast, PLDA benefits consistently from augmentation, reaching accuracy up to 94.7 % and reducing EER to 1.74 %, especially under full overlap when augmentation is applied symmetrically to both training and enrollment. These findings indicate that symmetric augmentation paired with PLDA significantly improves the robustness of i -vector speaker identification against session variability.
Tones are pitch variations on syllables that distinguish meaning in Chinese and reflect its phonetic characteristics. Mastery of tones directly impacts learners' Chinese proficiency. This study employs experimental phonetics to analyze the production of neutral tone, T2 and T3 by Chinese learners from five Central Asian countries: Uzbekistan, Kazakhstan, Tajikistan, Kyrgyzstan, and Turkmenistan. Firstly, results show common issues: neutral tones are produced overly stressed, significant deviation and confusion occur between T2 and T3, and tone deviation coexist with relative pitch inaccuracies. Secondly, when learning different tones, learners from different countries also exhibit varied performance, with Uzbekistan and Turkmenistan learners performing worst overall, while Tajikistan learners the best. Based on these findings, the study recommends optimizing tone teaching order, focusing on modal tone exercises, and designing targeted, countryspecific instructional materials.
This project develops a chatbot module for an Intelligent Tutoring System (ITS) in English language learning, powered by a large language model (LLM). The chatbot provides personalized practice, real-time grammar correction, vocabulary support, and question answering, adapting to each student's proficiency. Integrated with student and tutor modules, it tracks progress and tailors feedback dynamically. Leveraging LLM capabilities, the system understands natural language inputs and generates contextually relevant responses, enhancing engagement and learning outcomes. Evaluation focuses on accuracy, informativeness, and user satisfaction, offering a scalable AI-driven solution to improve English proficiency efficiently.
This paper introduces a Tetun text corpus Lafaek-Corpus-1M+. A Large Language Model (LLM) has attracted full attention in natural language processing and speech processing. To build an LLM, huge corpora are basically needed, however, it is quite hard for low-resourced languages. We focus on one Asian language Tetun, and make a text corpus for a Tetun LLM. We collect more than million Tetun sentences from various resources. After that, we applied continual pre-training to a Llama-3.1 model using our corpus to build a Tetun LLM. We conducted machine translation experiments. It is found that our LLM achieved better performance than the original model, and the effectiveness of our corpus is clarified.
Emotional prosody might distort realization of Mandarin tones, challenging young children's tonal perception, especially in noise. However, visual-articulatory cues have been shown to aid tonal perception in auditory-degraded conditions. This study explored emotional prosody's influence on preschoolers' tonal perception accuracy in quiet and noisy environments, and the role of visual-articulatory cues. Ninety-four children aged 4-6 and twenty-six adult controls participated. Stimuli included tones produced in emotional utterances (happy, neutral, and angry) under quiet and noisy environments. The tonal perception task was conducted in audio-only (AO) and audiovisual (AV) conditions. Results showed in AO, child groups generally exhibited high tonal accuracy across emotions, but significantly declined in noise. In AV, 6-year-olds performed comparably in quiet and noisy environments, similar to adults, indicating benefits from audiovisual integration. These findings suggest preschoolers demonstrated robust tonal perception abilities under emotional prosodic influence, and that visual-articulatory cues facilitated Mandarin tone perception in noise for 6-year-olds.
This study compares the acoustic properties of the Standard Chinese /n/ and /l/ in different prevocalic and tonal contexts The speech data of eight Mandarin speakers show that the distinction between /n/ and /l/ primarily lies in the spectral energy distribution along the frequency scale, regardless of the contexts in which two consonants occur. While /l/ exhibits enhanced high-frequency energy, /n/ displays a more skewed spectral shape with energy concentrated in the lower frequency region. The frequency-related features of /n/ and /l/ play a relatively minor role in distinction. Generally /n/ and /l/ share similar formant patterns, but the formant values of /l/ show more variation in different vowel contexts, presumably due to more coarticulation with the following vowel. There is no apparent tonal context effect on the acoustic properties of the two coronal consonants
In this work, we present an acoustic analysis of filled pauses (FPs) produced by 24 Assamese L1 speakers of L2 Hindi, examining their relationship with perceived proficiency. Spontaneous Hindi speech was elicited and rated for proficiency by 20 native Hindi listeners on a $1-5$ scale. A total of 667 FPs were extracted and analysed for type, vowel quality (F1, F2), duration, and speaker gender. By studying speech interaction in L2 speakers through FPs, we aim to answer pertinent questions in the discourse surrounding FPs. Are they speaker-specific, posing as an important tool for forensic science or language-specific? Results indicate a systematic proficiency-linked shift: high-rated speakers predominantly produced uhh, the central hesitation vowel in Hindi, while low-rated speakers employed a wider repertoire, including aah, the Assamese central vowel. Vowel space analysis showed that with increasing proficiency, aah tokens were fronted and centralized, reducing their acoustic distance from uhh, indicating accommodation to L2 sounds. Temporal analysis revealed that high-rated speakers produced significantly shorter and less variable FPs, whereas lower-rated speakers exhibited longer durations. These findings suggest that FPs are structured, language-specific cues that index both linguistic background and proficiency. From an applied perspective, the findings underscore the role of hesitation acoustics in enhancing automatic speech recognition.
Few studies have systematically examined the effect of question intonation on focus in cross-linguistic contexts, particularly in comparing intonation languages and tone languages. This study addresses this gap by analyzing data from 14 native speakers of Tianjin Mandarin (a tonal language) and 4 native speakers of American English (an intonation language) using Autosegmental-Metrical framework to investigate how question intonation modulates focus realization. Key findings include: (1) Question intonation in both languages attenuates focus-induced pitch rise and pitch range expansion, with a stronger weakening effect in American English compared to Tianjin Mandarin. (2) Question intonation reduces post-focus compression, with a stronger effect in American English than in Tianjin Mandarin. (3) Prosodic cues indicate the effect of question intonation on focused targets is primarily pitch-related in Tianjin Mandarin, while mainly durational change in American English.
The present work looks into the acoustics of voiced and voiceless laterals in Hmar. Voiceless laterals are rare crosslinguistically and are not well described. The waveforms show that voiced laterals in Hmar are fully voiced, but the voiceless laterals have a voiced portion just before the vowel and after the aspiration. The data collected from native Hmar speakers were analyzed. Acoustic measures such as duration, voicing, Harmonic-to-Noise Ratio (HNR), Standard Deviation (SD), intensity, the first four formant frequencies, kurtosis, skewness, and the Center of Gravity (CoG) were calculated, and the results showed that there is a significant difference in the voiced and voiceless laterals in Hmar. Further analyses using statistical methods were conducted to validate the findings.
We conducted comprehensive acoustic analyses to identify features influencing speech intelligibility using read-aloud speech from the Japanese super-elderly corpus EARS, including measurements of Fo, vocal quality, formant frequencies, vowel articulation metrics, F2 slope, voice onset time of plosives and following vowel duration, speech and articulation rates, and the frequency of non-speech intervals. Comparisons were made among cognitively healthy speakers with intelligible speech (CH I) or with slightly unintelligible speech (CH SU), and speakers with dementia and slightly unintelligible speech (DM SU). Both CH SU and DM SU exhibited significantly increased frequency of non-speech intervals and reduced speaking rates relative to CH I speakers. Additionally, DM SU speakers demonstrated greater standard deviation in two vowel formants compared to CH I, which could suggest increased articulatory instability associated with dementia. For the plosive-containing syllable, CH I speakers showed longer vowel duration and higher consonant-to-vowel duration ratio than both of SU groups, indicating a potential acoustic correlation of reduced intelligibility.
Measuring how well emotions are preserved in speech-to-speech translation is a difficult task. Previous research focused on measuring the similarity between speech embeddings related to emotion or prosody, which may not accurately capture the emotional contents. This study proposes more direct approaches by evaluating emotion preservation using metrics derived from speech emotion recognition (SER), including balanced accuracy ratio, emotion preservation rate, and Cohen's Kappa. The results show that about half of the original emotion is preserved in the translation process in the MELD-ST dataset, with metrics such as balanced accuracy ratio, valence-arousal similarity, and pause rate serving as reliable indicators of emotion preservation. We also found that high similarities between emotion embeddings do not necessarily mean emotion preservation, since the same acoustic embeddings used for SER lead to distinct performance. The analysis highlights the challenges of maintaining emotional consistency during speech-to-speech translation.
A speech neuroprosthesis is a system that translates neural activity signals into coherent text through several processing stages, from neural feature extraction and phoneme probability decoding to the application of a language model. In recent years, n-gram language models like trigrams and 5grams have been used to guide heuristic search algorithms in converting phoneme probabilities into word sequences. However, Transformer-based language models have demonstrated significant performance improvements in natural language processing and automatic speech recognition, thus offering a promising opportunity to enhance transcription accuracy by reducing error rates. The research dataset includes neural features from 256 ECoG electrodes, which are decoded into phoneme probabilities using a BiRNN with CTC loss. These probabilities are processed through beam search shallow fusion to combine acoustic and language scores. Evaluation using Phone Error Rate (PER), Character Error Rate (CER), Word Error Rate (WER), Words Per Minute (WPM), and Real-Time Factor (RTF) metrics shows that a Transformer based language model fine-tuned on the Open Web Text 2 corpus provides the best trade-off between accuracy and decoding speed. LLaMA 2 achieved the top performance with a WER of 16.9%, CER of 14.5%, WPM of 62.5, and an RTF of 0.98. Conversely, while n-gram models reached a WPM above 74, their WER remained in the 26%-29% range. These findings confirm the superiority of Transformer based language models for speech neuroprosthesis applications.
In the digital era, sentiment analysis represents a significant domain within natural language processing (NLP). Nevertheless, research on Indonesian regional languages, particularly Acehnese, remains limited. The primary challenges involve the unavailability of representative datasets and the absence of BERT-based models optimized through the Masked Language Modeling (MLM) approach. Existing models, such as IndoBERT, are trained on Indonesian corpora and therefore fail to adequately capture the unique linguistic features of Acehnese. This study develops the AcehX Sentiment dataset and introduces the AcehXBERT model through pre-training IndoBERT-base with the MLM approach on the AcehX corpus. The model is then fine-tuned for Acehnese sentiment classification tasks. Experimental results show an F1-macro score of 82.50% on the AcehX Sentiment dataset and 81.89% on the NusaX Sentiment dataset, outperforming NusaBERT. These findings highlight the importance of adapting pretrained models and tokenizers for regional languages, while supporting the preservation and technological integration of the Acehnese language.
Taiwanese part-of-speech (POS) tagging remains highly challenging due to the scarcity of annotated training data. To address this limitation, we propose a large language model (LLM) merging framework that unifies three key functions into a single system: (1) forward translation from Taiwanese into Chinese, (2) POS tagging of the translated Chinese text, and (3) reverse projection of the tags back into Taiwanese. These steps are consolidated within one LLM through parameter-space model merging. Specifically, two LLaMA3-8B-Instruct models were fine-tuned separately: a bidirectional Chinese-Taiwanese translation model trained on parallel corpora, and a Chinese POS tagging model trained on the Sinica Corpus. The two models were then merged into a single LLM capable of direct Taiwanese POS annotation. Without relying on any Taiwanese POS supervision, the merged model achieved an F1score of 75.61 %, a precision of 78.84 %, and a recall of 74.21 % on a 837 -sentences Taiwanese test set, demonstrating the effectiveness of model merging for POS tagging in low-resource languages.
This study uses ERPs (Event-Related Potentials) to examine how Chinese EFL learners process prosodic boundaries and integrate them with syntax in English sentences. Both early closure (EC) and late closure (LC) sentences elicited a Closure Positive Shift (CPS) at prosodic boundaries, with the CPS in EC sentences delayed by approximately 300 ms. When prosodic and syntactic cues conflicted, a P600-like positive deflection was observed at the disambiguation point, indicating syntactic processing difficulty and reanalysis. Notably, EC sentences induced a stronger P600 than LC sentences, reflecting Chinese learners' persistent preference for LC parsing despite the presence of progressive aspect verbs intended to reduce transitive verb bias. Contrary to the Boundary Deletion Hypothesis, the findings suggest that inserting a missing prosodic boundary imposes greater cognitive demand than deleting a superfluous one for Chinese learners. This pattern is attributed to their LC bias and reliance on a “good enough” heuristic, leading to superficial sentence interpretations and challenges in revising initial misanalyses.
While current emotional Text-to-Speech (TTS) models have successfully controlled verbal prosody, they often ignore non-verbal vocalizations (NVs), which are essential for authentic human emotion. Although some non-verbal datasets have recently emerged, they often lack high-quality, fine-grained annotations, which restricts a model's ability to precisely control NV generation. To address this limitation, we propose a novel approach for fine-grained non-verbal expression synthesis. We curate and reprocess female NV utterances from the EARS corpus, develop a new annotation scheme using tags to encode NV types, frequencies, and durations, and build an emotional TTS benchmark to demonstrate its effectiveness. Our evaluation shows that while our NV approach leads to minor trade-offs in perceived naturalness, it significantly improves expressiveness (eMOS 4.20) and emotional recognition accuracy (78.8%). Emotion-specific analysis further reveals that NV cues are highly effective for high-arousal emotions like happy (82.5 %) and fear (82.7 %), and almost perfectly convey sadness (98.3%).
This paper presents analyses of dialectal variation in Assamese and Finnish using utterance-level embeddings extracted from a self-supervised speech representation model fine-tuned for language identification (LID). The languages are represented by two speech corpora substantially differing in their design and composition. Rather than extracting and analyzing specific acoustic features, we apply two linear transformations-principal components analysis and linear discriminant analysis-on the embedding space, enabling a relatively theory-independent investigation of dialectal relationships without the need to define cross-linguistic features for comparison. We evaluate the effects of these transformations using the geographical distances as a proxy of relatedness among varieties. We show that for both languages, our method yields quantifiable and interpretable results in terms of clustering varieties into meaningful dialectal groupings.
Recognizing speech when multiple individuals speak simultaneously remains a significant challenge for Automatic Speech Recognition systems. While modern architectures integrating audio model and large language model show promise, the best way to fine-tune these models is not fully explored. This study investigates whether fine-tuning the language model first, the audio model first, or both components jointly yields the best results. Our results demonstrate that the fine-tuning order is a critical factor. The strategy of adapting the language model first achieves the highest performance with a Word Error Rate of 8.96 %. This surpasses the 9.62 % WER obtained when the audio model is trained first. This performance advantage is also apparent when using LoRA. These results establish a hierarchy of fine-tuning strategies and highlight a key trade-off between transcription accuracy and computational efficiency for practical applications.