
Speech Emotion Recognition (SER) is a significant challenge within the field of Human-Computer Interaction (HCI). Effective HCI design emphasizes the development of systems that are practical, intuitive, and visually appealing, with a focus on real-world applicability and inclusiveness. This study evaluates the usability and user experience of a SER system designed for online teaching, accommodating educators with and without hearing impairments. The objective of this research is to provide a universally accessible solution for online teaching, with a particular focus on educators with late-onset deafness or hearing impairments, a group often overlooked in technological design. This research outlines the methods used for data collection, analysis, and reporting involving both hearing and hearing-impaired educators. The evaluation focuses on the SER system’s real-time emotion detection capabilities and user interface. The findings confirm the system’s high usability and effectiveness in engaging all educators, especially those with hearing impairments.
Recently, speaker verification algorithms have been of great interest because of their wide applicability in several fields such as speech communication, domestic services, security and access control, and intelligent terminals. Today we have interactive devices such as smartphone assistants and smart speakers that are designed to understand basic voice commands. However, the performance of current speaker verification programs degrades significantly when analyzing short utterances, and larger speech segments are needed to achieve high performance. Although CNNs, LSTM networks, MFCCs, CMVN, and PLDA have been thoroughly studied for speaker recognition, their performance is still limited in the case of extremely brief speech, due to the lack of speaker-discriminative information in short utterances. The uniqueness of this work is not in proposing a new neural architecture but in presenting an integrated framework that integrates CMVNC-normalized acoustic data with complementary CNN and LSTM-based temporal modeling for robust short-utterance speaker verification. Unlike prior studies that predominantly utilise traditional MFCC features or assess a sole deep learning architecture, this study systematically explores the contribution of CMVNC features across two deep learning models, and juxtaposes them against the traditional i-vector/PLDA baseline under varying training and testing utterance lengths.
This study presents a comprehensive framework for automatic speech recognition (ASR) tailored to dysarthric speakers, leveraging the newly developed UASPEECH Preprocessed Dataset. The original UASPEECH recordings were enhanced through systematic signal processing techniques, including FFT-based denoising, Hanning windowing, single-sided amplitude doubling, and amplitude normalization, resulting in clean, standardized 16-bit WAV audio at 44.1 kHz. Two modeling pipelines were explored: one using handcrafted acoustic features and another based on mel spectrogram representations. For the acoustic features pipeline, several classical and deep learning models were evaluated, including Multi-Layer Perceptrons (MLP_1–3), LSTM, GRU, BiLSTM, BiGRU, CNN_1D, and an attention-only model. Among these, BiLSTM achieved the best performance, with an accuracy and F1-score of 97.25
This paper introduces a noise-resilient method for segmenting Dravidian-accented Malayalam speech into syllable-like units. These sub-units, derived from acoustic cues that approximate syllables from a linguistic perspective, are referred to as syllable-like units. The proposed approach employs sonority estimation based on the Ramped Autocorrelation Coefficient (RAC), which effectively captures the auditory prominence of syllables even in adverse acoustic conditions. Unlike conventional segmentation techniques, the method exploits the noise-robust property of autocorrelation to detect sonority peaks and valleys, thereby improving segmentation reliability under background noise. The method is evaluated using the Dravidian Accented Malayalam Speech Database (DAMSD), which includes clean and simulated noisy speech generated with White Gaussian, Pink, Red, and Babble noise at three signal-to-noise ratio (SNR) levels: 20 dB, 10 dB, and 0 dB. The manually annotated syllable boundaries from the DAMSD corpus serve as the ground truth for evaluation. Experimental results demonstrate consistent segmentation performance across all noise conditions, with F1-scores ranging from 74.15
Human emotions are essential indicators of mental states and are increasingly utilized in applications ranging from healthcare to intelligent systems. This work starts with a concise review of the field, emphasizing key steps for accurate emotion recognition, including feature extraction from both speech and facial modalities and classification using both traditional machine learning and modern deep learning techniques. Major challenges, such as variability across speakers, cultural influences, and the scarcity of annotated datasets, are also discussed. Building on this foundation, the paper introduces an adaptive multimodal fusion model that moves beyond static integration strategies. By leveraging confidence-based weighting and transformer-inspired attention mechanisms, the model dynamically combines audio and visual features, enhancing resilience in scenarios with noisy audio or partially occluded faces. Overall, this study combines a systematic review with a methodological innovation, offering both a synthesis of current knowledge and a practical framework for developing robust and generalizable emotion recognition systems. Experimental results on the RML and SAVEE datasets under multiple evaluation protocols, including stratified hold-out, 5-fold cross-validation, Leave-One-Speaker-Out (LOSO), and cross-dataset evaluation, demonstrate the effectiveness and stability of the proposed framework. The proposed adaptive cross-attention fusion model consistently outperforms unimodal baselines and conventional fusion strategies, while the cross-dataset experiments reveal the remaining challenges of domain generalization in multimodal emotion recognition.
Deep learning has revolutionized audio processing, with foundation models like Wav2Vec 2.0 (317 M parameters), HuBERT, and Whisper achieving state-of-the-art results across speech recognition, speaker verification, and sound event detection. However, fine-tuning these massive models for domain-specific tasks is computationally prohibitive, requiring updating hundreds of millions of parameters with 16N bytes of memory overhead (where N is parameter count) and substantial GPU/TPU resources. Parameter-efficient fine-tuning (PEFT) techniques address this limitation by updating only a small subset of model parameters while preserving or enhancing performance. This survey systematically reviews PEFT methods that update only 0.1–5
Despite of swift advancements in the field of Artificial Intelligence, an effective immediate assistive solution for visually impaired individuals remains limited. Whereas, Deep Learning and Natural Language Processing (NLP) have attained significant gain in visual sympathetic and language generation but their addition into accessible systems for ecological perception is still underexplored. This dictates the development of intelligent frameworks capable of accurately rendering visual scenes and delivering meaningful descriptions to enhance liberation and situational awareness for visually impaired users. Consequently, we proposed a framework of three levels using the deep learning and NLP approaches to address aforementioned. It also developed a novel approach to learning the better relational features among the objects, scenes, and persons captured in an image and generated an accurate caption. Firstly, object detection algorithms are used to detect the objects in an image. The next level generated the image captions by maximizing the likelihood of the expected captions using deep learning. The third level, the text in the caption, is converted into a voice using NLP algorithms. The proposed model has substantially compared the existing state-of-the-art image captioning and voice conversion methods by experimenting with benchmark datasets such as Flickr 8k, Flickr 30k, COCO Caption, and RyanSpeech. The proposed model outperformed in terms of recall on Flickr 8K dataset @100 images has scored 67.25
Spectral features are highly sensitive to acoustic mismatches, including variations in pitch, speaking rate, and ambient noise. To address these challenges, this paper presents a method for parameterizing subband speech signals in the temporal domain, resulting in a pitch-robust acoustic feature representation. The method involves decomposing the speech signal into subband complementary signals using N equal-bandwidth band-pass filters. The short-term temporal slope of each band-pass signal is calculated by applying a first-order infinite impulse response (IIR) low-pass filter, followed by a non-local differencing operation. These slope values are then logarithmically compressed to generate an N -dimensional feature vector for each analysis frame. The resulting feature, termed the logarithmic compressed subband temporal slope (LC-SBTS), primarily captures the energy transition patterns of sound units while suppressing speaker-specific traits. The robustness of the proposed features is demonstrated through analytical validation using t-distributed stochastic neighbor embedding, evaluation of keyword-spotting performance under pitch-mismatched test conditions, and comparison with existing pitch-normalized feature computation techniques.
Automated Speech Recognition (ASR) systems face significant challenges in accurately transcribing speech, especially in low-resource languages like Tamil, which have a variety of dialects and slang. In such languages, traditional ASR models struggle to adapt to these dynamic variations, often leading to frequent transcription errors when encountering new terms or informal speech patterns. To address these errors, existing systems commonly rely on predefined dictionaries. However, these dictionaries are static and cannot adapt to the different dialects and variations within the Tamil language which leads to inaccuracies. This limitation highlights the need for more flexible solutions which is capable of addressing the complexities of diverse and evolving speech patterns in Tamil Language. In this study, a novel auto-update mechanism called Pattern Matching with Levenshtein Distance for Dictionary-Based Text Correction (PMLD-DTC) is propose, which is integrated with the wav2vec2 ASR model. The PMLD-DTC technique enhances transcription by performing character-level matching with a dynamic vocabulary file, enabling real-time updates. This approach incorporates slang, regional variations, and new terms by auto-updating the dictionary which reduces manual intervention and improving adaptability to linguistic diversity of Tamil. Experimental validation was conducted on CodaLab dataset, and two Hugging face dataset, achieving an accuracy of 92
The extraction of speaker-related features through utterance embeddings has been extensively studied for years. Convolutional Neural Networks (CNNs), particularly deep Residual Networks (ResNets), have been widely used to capture spectral and hierarchical feature representations, enabling rich speaker embeddings. However, the development of robust and generalizable speaker representations remains a central challenge. Tailoring backbone architectures to the unique characteristics of speech is therefore crucial for effective embedding learning, raising a fundamental question: Which architectural principles yield the most robust and generalized speaker representations?. This study investigates CNN scaling dimensions as foundational design factors for speaker embedding backbones. The objective is to systematically investigate how different fundamental scaling dimensions influence Robustness and Generalization (R G) in speaker embedding learning. We conduct a comprehensive evaluation by integrating these dimensions into two widely used ResNet-based baselines, ResNet-34 and ECAPA-TDNN. The experiments span both in-domain (Automatic Speaker Verification (ASV)) and cross-domain (Speech Emotion Recognition (SER)) tasks and are further supported by t-SNE and sharpness-aware generalization analysis through loss landscape visualization. The findings represent a step toward learning more robust and generalizable speaker representations through the integration of scale and cardinality dimensions. This design achieves smoother optimization trajectories, improved stability, and enhanced representational capacity over established backbones, demonstrating its broad applicability and effectiveness as a backbone network in speaker embeddings. Experimental results reveal consistent improvements in R G across diverse conditions, highlighting the ability of multi-scale and multi-branch aggregated transformations to capture rich speech cues—such as pitch, tone, and phonation patterns—by jointly expanding the receptive field and feature diversity. This combination is valuable for capturing speaker traits that span both short- and long-term dependencies.
Parkinson’s disease (PD) is a neurodegenerative disorder that frequently presents with vocal impairments, making speech analysis a potentially effective noninvasive method for early detection. Current techniques reliant on conventional acoustic features frequently fail to accurately represent the nuanced, non-stationary dynamics of pathological speech. This study introduces a novel methodology utilizing time-frequency (TF) analysis through Fourier synchrosqueezing transform to improve the efficacy of PD detection from speech signals. The first speech signal is processed with FSST to produce a TF representation. We used this TF representation to calculate the energy and entropy of each frequency component and used it as a feature for classification. The various machine learning classifiers are used for classification along with the genetic algorithm (GA). The efficacy of the proposed method is assessed utilizing vowels and words from the PC-GITA dataset. The proposed method attained a classification accuracy of 91
Parkinson’s Disease (PD) is a neurological disorder characterized by the gradual degradation of dopamine. It is most prevalent after Alzheimer’s disease particularly seen in old age people. Abnormalities in speech signals have been identified as an indicator of PD. This study introduces a unique method for detecting PD by analyzing speech data using the Ramanujan Fourier Transform (RFT). The projection of the acquired numerical series into a set of fundamental functions composed of Ramanujan sums (RS) is the foundation for the RFT. In this work, RFT-based features are proposed for the diagnosis of PD utilizing speech signals. The proposed features are evaluated using sustained vowels and isolated words from the PC-GITA database. The light gradient boosting machine (LGBM) achieves a maximum classification accuracy of 95
Stemming is an important preprocessing step in Natural Language Processing because it helps handle the rich morphological structure of the Bangla language which increases vocabulary size and affects model performance. This study evaluates different Bangla stemming techniques for text classification tasks using both traditional and deep learning classifiers including Nave Bayes, Support Vector Machines, Random Forest, LSTM, CNN-BiLSTM and BanglaBERT. Experimentation on a Bangla news corpus consisting of 67,564 instances belonging to seven classes demonstrated that stemming reduces vocabulary size by 41
Accurate assessment of service quality is essential in the digital economy; however, conventional methods often fail to capture the emotional and linguistic subtleties of real-world customer interactions, especially in rapidly developing service sectors such as Vietnam. To address this gap, we propose a speech-based multimodal pipeline that integrates audio and textual signals for automated service quality evaluation. The pipeline begins with audio preprocessing using a UNet-based denoising model and WhisperX-based speaker diarization. We then introduce a Dynamic Attention Network, proposed in this study, trained on our self-constructed VNEMOS dataset (250 Vietnamese emotional speech segments across five emotion categories) for speech emotion recognition. For textual analysis, the pipeline incorporates PhoWhisper for transcription and a PhoBERT-CNN model for three-class sentiment analysis (positive, negative, neutral). A probabilistic fusion mechanism, also proposed in this work, leverages a sigmoid-based risk accumulation function to combine emotional and linguistic cues and classify each interaction as either “Good” or “Bad”. The pipeline is evaluated on our second dataset, a Vietnamese logistics customer service corpus (30 minutes, 235 annotated segments), and achieves an F1-score of 81.21
Multilingual Neural Machine Translation (MNMT) plays a vital role in extending language technologies to underrepresented and linguistically diverse regions. However, existing MNMT systems remain largely opaque, particularly when applied to structurally diverse and low-resource languages such as those of Northeast India. In this study, we propose an explainable MNMT framework tailored for bidirectional translation between English and six low-resource Indic languages—Assamese, Bodo, Khasi, Manipuri, Mizo, and Nepali—spanning Indo-Aryan, Tibeto-Burman, and Austroasiatic families. Our architecture employs a Transformer backbone augmented with language-conditioned adapter modules and sparsely activated Mixture-of-Experts (MoE) layers, enabling parameter sharing across languages while preserving family-specific specialization. The model was trained on curated parallel corpora with back-translation augmentation, achieving strong BLEU scores of 29.1 for Assamese and 26.1 for Nepali. To improve transparency, we integrated complementary post-hoc interpretability techniques—attention visualization, SHAP, and LIME—providing token-level and layer-wise explanations of translation decisions. Both quantitative and qualitative analyses show that these interpretability tools effectively diagnose translation challenges, reveal systematic biases, and elucidate failure modes, particularly for data-scarce languages such as Khasi and Mizo. Our results demonstrate that combining adapters, MoE-based specialization, and explainability can advance both the performance and trustworthiness of MNMT for low-resource Indic languages.
The Orani dialect of Arabic is an under-resourced Algerian language variety which also suffers from a lack of systematic evaluation of morphological analyzers. This study attempts to fill a critical resource gap for dialectal Arabic NLP, linguistic research, and educational applications. It presents MADOran, a morphologically annotated corpus for the Orani dialect of Arabic (ORN), together with a systematic evaluation of morphological analyzers across conventional, deep, and transformer-based approaches. The dataset contains 30,919 words drawn from: written texts (41
Spoken Language Identification (SPLID) is crucial for applications like speaker recognition, automatic voice recognition, and multilingual content indexing. Recent advances in deep learning have significantly improved SPLID systems. Multi-Task Learning (MTL) is a successful paradigm that enables a model to learn multiple tasks simultaneously by using shared representations. In MTL, auxiliary tasks provide additional monitoring to support the main work during training. These supplementary activities could be related tasks, which share characteristics with the main objective, or orthogonal tasks, which are unrelated but complementary tasks that enhance the model’s ability to generalize. This work investigates the integration of orthogonal and related auxiliary tasks for spoken language identification of Indian languages within a multi-task learning (MTL) framework. We extend our previous work on related-task learning by using language family classification as the related auxiliary task and real/spoof detection as an orthogonal auxiliary task. The unified integration of both task types within a single MTL framework, which allows the model to learn complementary representations, is what makes the suggested method novel. This design improves robustness and generalization, especially in low-resource Indian language scenarios where phonetic similarities and data scarcity pose major challenges. Three SPLID datasets of Indian languages which pose difficulties because of their phonetic similarity and comparatively low resource availability were used for the experiments. A baseline model for Single-Task Learning (STL) is used to assess performance. The suggested MTL approach outperforms the STL baseline by about 7
Enhancing human-computer interaction requires Speech Emotion Recognition (SER), which allows computers to identify emotions from vocal expressions. Although SER development has advanced in high-resource languages, it has not progressed as far in low-resource languages such as Bangla, particularly when it comes to privacy concerns and dialects. In this study, a safe and dialect-sensitive SER framework designed for Bangla is presented. It can identify five basic emotions: neutral, happy, sad, angry, and surprise. The study examines three hybrid deep learning models: CNN-BiLSTM for spatial-temporal feature extraction, EmoDARTS using differentiable architectural search, and EfficientNet in conjunction with Vision Transformer. The EfficientNet-ViT model ensured decentralized, privacy-preserving training by achieving maximum accuracy (95. 9
Indian Sign Language (ISL) translation and recognition face significant challenges due to a lack of sufficient datasets and difficulties in generalizing across various signs. These complexities make the real-time classification process more critical in improving communication between the Deaf and hearing communities. The latest developments in Machine Learning (ML) and Deep Learning (DL) techniques seem to offer encouraging possibilities for overcoming these challenges. An advanced learning technique is utilized to create more extensive datasets and support cross-linguistic research, which reduces computational time. Recently, different techniques have been analysed to determine effective ISL recognition and translation tasks. Hence, this research is focused on designing a survey by analyzing 75 research papers collected from 2007 to 2024 according to recognition and translation of ISL methods. The main contribution of this survey is to motivate research to enhance the effectiveness and accuracy of ISL recognition and translation, through an examination of performance methods on a range of different measures and datasets. The surveyed papers show attention to the difficulties encountered in previous research, to update these issues, recent works have focused on improving the limited real-time processing capabilities of DL and ML approaches. Exposing these issues and improving them with enhanced datasets has stimulated the ISL to the development of better ways for effective ISL communication.
This paper presents a Transfer Learning framework for developing Automatic Speech Recognition (ASR) systems for resource-poor Indian languages, Odia and Assamese. These languages are widely spoken in eastern India but lack sufficient annotated speech corpora. However, adequate resources are available in other Indian languages, such as Bengali and Hindi. Transfer Learning can help leverage those resources to develop ASR systems in Odia and Assamese. First, a baseline system is created using available resources in the target language using Bidirectional Long Short-Term Memory and Connectionist Temporal Classification. Then, pre-trained models are trained using open ASR resources in Bengali, Hindi, and English. The knowledge accumulated from these languages through the pre-trained models is applied to the target languages. A Modular-Adapter-based transfer learning framework is proposed for this purpose. During experiments, it is found that the proposed transfer learning models substantially outperform the baseline models in both languages.