A soundscape is composed of three types of sound: biophony (sounds made by animals), geophony (natural abiotic sounds) and anthropophony (sounds made by humans). A key research question in the field of soundscape ecology is how these components interact with each other, specifically how biophony responds to geophony and anthropophony. Nevertheless, as of today, there are not many analytical instruments that enable the distinct quantification of these elements. Recent machine learning (ML) approaches aim to support automated analysis but often rely on task-specific or clean data, limiting generalisation to noisy passive acoustic monitoring (PAM) recordings. This study presents a clear and reproducible structure to build ML models for coarse soundscape classification and introduces CoarseSoundNet, a deep learning model trained to distinguish biophony, geophony, and anthropophony under realistic PAM conditions. We systematically investigate model architectures, the influence of an additional training class, data composition, and evaluation strategies. Our findings suggest that model performance improves with additional PAM data, especially when similar to the target domain, and by introducing an explicit silence class during training. Class-specific decision thresholds and duration-based constraints further enhance performance, particularly for anthropophony and geophony. Error analyses exhibit challenges for anthropophony due to masking effects and confusions for silence and insect sounds for geophony and biophony. Finally, we conduct an ecological case study which shows that pre-filtering recordings with CoarseSoundNet yields acoustic index trends comparable to ground-truth filtering, supporting its use as an effective preprocessing tool for ecoacoustic analyses.
Ambulatory assessment studies have proliferated in mental health science, promising unparalleled insights into the dynamic nature of mental health. The high methodological heterogeneity of ambulatory assessment studies calls for the harmonization of approaches and the establishment of research standards. This expert consensus provides an overview of best-practice recommendations to integrate ambulatory assessment into mental health research. We queried 26 experts in the field in a Delphi-adapted process, resulting in recommendations regarding eight topics, encompassing methodological considerations, technological advancements and practical applications. Interdisciplinary collaboration and personalized treatment approaches were recurrent themes. By establishing standardized guidelines and best practices, this expert consensus serves as a resource for mental health scholars seeking to leverage ambulatory assessment tools in research and practice. Ultimately, the integration of ambulatory assessment holds promise for revolutionizing mental health care, fostering precision medicine, and improving outcomes for individuals living with mental disorders. This Consensus Report addresses the methodological variability in ambulatory assessment studies within mental health. Utilizing a Delphi-adapted process with 26 experts, it offers best-practice recommendations across eight topics, emphasizing collaboration, technology and personalized treatment approaches for enhanced research integration.
The convergence of Generative Artificial Intelligence (GenAI) and Affective Computing (AffComp) opens new frontiers in creation of emotionally intelligent systems, with growing relevance for human–machine interaction across domains such as virtual assistants, avatars, gaming, and personalised digital experiences. In this paper, we explore the opportunities arising from the integration of GenAI with AffComp across various application areas—data augmentation and synthesis, emotion generation, personalised content creation, data imputation, and emotion transfer and style manipulation—with a particular focus on the deployment of emotionally intelligent systems through a risk–reward balancing framework. The study adopts an empirical and evaluative approach, leveraging state-of-the-art GenAI models within a multimodal AffComp framework to demonstrate significant improvements in emotional expressiveness, adaptive behavior, and personalized responses, enhancing human–computer interaction in applications such as virtual assistants, avatars, and gaming. However, this convergence introduces privacy, ethical, and security risks. Mitigating these risks requires privacy-preserving techniques, real-time bias detection, transparent and culturally diverse models, strong ethical frameworks, user education, and interdisciplinary collaboration to support responsible deployment, ongoing evaluation, and robust oversight. By balancing the associated risks and rewards, we outline strategies to maximise the benefits of GenAI and AffComp while mitigating the potential harms. This paper invites the research community to collaborate on shaping the future of emotionally intelligent AI systems, prioritising both innovation and the well-being of individuals and society as a whole.
BackgroundIdentification of fatigue in people with Multiple Sclerosis (pwMS) is still mainly based on subjective assessments due to the lack of objective diagnostic tools. We aimed to identify vocal biomarkers to differentiate between pwMS with and without fatigue.MethodsThis COMMITMENT trial was a prospective, observational study recruiting healthy controls (HCs, n = 20) and relapsing MS (RMS) patients (n = 50, EDSS <4.0) at University Hospital of Bern. The primary objective was to predict fatigue in pwMS by means of automated artificial intelligence (AI)-based speech analysis. An exploratory objective was to differentiate between pwMS and HCs.FindingsParticipants had a mean age of 36.0 years, and 73% were female, with no significant group differences. Median EDSS in pwMS was 1.0 (range 0–3.0). Motor fatigue affected 50% and cognitive fatigue 40% of pwMS. Five acoustic features were associated with general fatigue, independently of depression and sleepiness. Further, five features were associated with motor-, and 12 with cognitive fatigue. The best-performing classification and regression models [Leave-One-Speaker-Out (LOSO) paradigm] achieved specificities of 0.68–0.94, whereas sensitivity remained lower (0.38–0.90). Speech biomarkers distinguished pwMS from HCs with a specificity of 0.90 but only a sensitivity of 0.3.InterpretationSpeech in pwMS may serve as a potential biomarker for MS-associated fatigue and might help to differentiate between pwMS and HC. Our findings suggest that AI-assisted speech analysis could complement existing fatigue assessments.
Speech-based automatic estimation of depression levels is essential for enabling early detection and timely intervention, particularly in resource-constrained mental health settings. In recent years, deep learning has demonstrated impressive success across various domains, including affective computing and mental health assessment. Most existing approaches rely on RNN-based architectures (such as LSTM and GRU) to model temporal information for depression estimation. However, the extracted features often emphasize only a few adjacent speech segments, limiting their ability to capture long-range dependencies. To overcome this limitation, we introduce a memory-based feature augmentation method that enhances the representational capacity of GRU-extracted features. Rather than indiscriminately incorporating historical data, our memory bank is designed to selectively integrate two types of components in order to reduce redundancy and irrelevance: (1) historical temporal features that closely resemble the current GRU output, offering complementary contextual information; and (2) dynamic memory features identified based on feature variability, which capture behavioral and emotional fluctuations indicative of depressive symptoms. To effectively fuse the memory-augmented features with GRU outputs, we further design a Hierarchical Attention Fusion (HAF) module. Our method is evaluated on the widely used DAIC-WOZ and E-DAIC datasets, achieving state-of-the-art performance.
Instruction-following audio language models (ALMs) can be augmented with explicit acoustic cues, yet it remains unclear whether such cues are used in a grounded way when the raw audio is already available. We study this question in speech emotion recognition (SER) by deriving six interpretable acoustic concept tokens from the standardised eGeMAPS paralinguistic feature set. These tokens summarise energy, pitch, dynamics, brightness, formants, and voice quality, and are appended to the textual prompt while the audio input is kept unchanged. Across the widely used FAU-Aibo and IEMOCAP benchmarks, aligned tokens improve unweighted average recall (UAR), whereas shuffled, conflicting, or corrupted tokens reduce performance relative to aligned tokens and shift confusions toward neutral. Importantly, predictions do not collapse under strong token perturbations, suggesting that the models are sensitive to the symbolic cue channel but remain partly anchored to the audio signal. We argue that token-only interventions provide a practical way to probe audio-grounded cue use, robustness, and interpretability in ALM-based affective computing.
Machine-generated music (MGM) has become a powerful tool with applications in music therapy, personalised editing, and creative inspiration. However, its unregulated use threatens the entertainment, education, and arts sectors by diminishing the value of high-quality human compositions. Effective detection of machine-generated music (MGMD) is essential, yet progress is hindered by the lack of comprehensive datasets. To address this gap, we introduce M6, a large-scope benchmark dataset designed for MGMD research. M6 is distinguished by its diversity, encompassing multiple generators, domains, languages, cultural contexts, genres, and instruments, all provided in WAV format. We detail the data collection methodology and analysis, alongside baseline performance scores from foundational binary classification models, highlighting the complexity of MGMD and the need for further advancements. M6 serves as a robust resource to support future research in developing effective detection methods. The dataset is available at https://huggingface.co/datasets/yl7622/M6 to promote collaboration and innovation.
The mental wellbeing of seafarers is particularly at risk due to the harsh conditions of their work environment. Speech is a promising modality to assess mental wellbeing in an unobtrusive way. In this study, we outline our approach to collect speech data in the challenging context of an oil tanker to monitor the mental wellbeing of a ship’s crew. In the course of this study, we employed two data collection approaches: 1) “active data collection” through the local (offline) deployment of a web-based platform, which participants accessed through their end devices, and 2) “passive data collection” through the ship’s voyage data recorder (VDR) system from 4 microphones on the bridge and 2 radio communication channels. For the active data, we analysed speech features and self-reported wellbeing measures. For the passive data, we applied denoising, segmented speech, and predicted dimensional emotional expression (valence, arousal, dominance). We incorporated environmental variables such as absolute wind speed and operational context (at sea, in port, cargo operations). We gathered speech data from 25 seafarers over a nine-week period, resulting in 378 actively recorded survey sessions. Analysis of the actively collected data reveals moderate correlations (r 0.3−0.4) between speech features and self-reported wellbeing measures. In the passively collected data, higher wind speeds were associated with lower inferred emotional intensity on the ship’s bridge microphones, particularly at sea (r −0.34). A mediation analysis showed that this ship’s bridge emotion intensity partially mediates the association between wind and concurrent self-assessed stress. Denoising improved signal-to-noise ratio distributions but had only modest effects on predictive performance. Our findings demonstrate the technical feasibility of combined active and passive speech monitoring in maritime settings and provide exploratory evidence for environmental influences on crew affective states. This work lays the foundation for future studies on personalised modelling and large-scale passive monitoring to support seafarer mental health.
Multimodal Emotion Recognition (MER) systems often suffer from missing modalities in real-world scenarios. Existing methods usually generate, align, or distill missing modalities as a whole, overlooking the heterogeneous nature of the information carried by each modality. Such holistic treatment mixes inferable shared semantics with uncertain modality-specific details, yielding unstable representations and degrading robustness. To address this issue, we propose the Primitive Memory Distillation (PriMD) framework. Unlike existing methods, PriMD takes an intra-modal perspective and focuses on how different types of information within a modality differ in recoverability within each modality. PriMD first disentangles cross-modal shared semantics from modality-specific representations, and then discretizes the latter into learnable semantic primitives to construct modality-specific memory banks. When modalities are missing, PriMD is a teacher-student framework that the student model uses the shared semantics of available modalities as queries to dynamically retrieve primitives. It compensates for missing modality-specific information within a constrained memory space and aligns with the teacher model. Extensive experiments on IEMOCAP, CMU-MOSI, and CMU-MOSEI demonstrate that PriMD achieves state-of-the-art performance and consistently stronger robustness across a wide range of missing-modality settings, while mitigating the instability caused by holistic feature inference. Our code and project website are available at https://github.com/JiaqiZhang-Sengoku/PriMD and https://jiaqizhang-sengoku.github.io/PriMD/, respectively.
The expression of affect is integral to spoken communication, yet, its link to underlying articulatory execution remains unclear. Measures of articulatory muscle activity such as EMG could reveal how speech production is modulated by emotion alongside acoustic speech analyses. We investigate affect decoding from facial and neck surface electromyography (sEMG) during phonated and silent speech production. For this purpose, we introduce a dataset comprising 2,780 utterances from 12 participants across 3 tasks, on which we evaluate both intra- and inter-subject decoding using a range of features and model embeddings. Our results reveal that EMG representations reliably discriminate frustration with up to 0.845 AUC, and generalize well across articulation modes. Our ablation study further demonstrates that affective signatures are embedded in facial motor activity and persist in the absence of phonation, highlighting the potential of EMG sensing for affect-aware silent speech interfaces.
The ambiguity of human emotions poses several challenges for machine learning models, as they often overlap and lack clear delineating boundaries. Contrastive language-audio pretraining (CLAP) has emerged as a key technique for generalisable emotion recognition. However, as conventional CLAP enforces a strict one-to-one alignment between paired audio-text samples, it overlooks intra-modal similarity and treats all non-matching pairs as equally negative. This conflicts with the fuzzy boundaries between different emotions. To address this limitation, we propose SmoothCLAP, which introduces softened targets derived from intra-modal similarity and paralinguistic features. By combining these softened targets with conventional contrastive supervision, SmoothCLAP learns embeddings that respect graded emotional relationships, while retaining the same inference pipeline as CLAP. Experiments on eight affective computing tasks across English and German demonstrate that SmoothCLAP is consistently achieving superior performance. Our results highlight that leveraging soft supervision is a promising strategy for building emotion-aware audio-text models.
Emotional intelligence, the capacity of machines to perceive, interpret, and respond to human affective states, demands a holistic understanding of behavior that spans subtle body movements, spoken language, and social interaction. This paper presents the proceedings and position of The First Joint Workshop on Human Behavior Analysis and Interaction for Emotional Intelligence at IJCAI 2026, bringing together with the 4th MiGA Challenge. The competitive MiGA challenge attracted 66 registered participants from worldwide institutions across three tracks. Beyond the competition session, the contributed workshop papers in the other two sessions further highlight recent advances in emotional intelligence. The Explainability in Emotional Intelligence session features papers on interpretability-driven speech emotion recognition and hate speech detection, as well as trustworthiness in speech language models. The Affective Interaction and Social Computing session focuses on socially grounded and safety-critical topics, including richness-aware emotional dialogue generation and clinical safety auditing of implicit sycophancy in mental health dialogue systems. Together, these contributions position the field toward trustworthy and socially aware emotional intelligence, while exposing persistent open challenges in evaluation rigor, domain adaptation, and clinical deployment.
Audio-visual speaker extraction has attracted increasing attention, as it removes the need for pre-registered speech and leverages the visual modality as a complement to audio. Although existing methods have achieved impressive performance, the issue of degraded visual inputs has received relatively little attention, despite being common in real-world scenarios. Previous attempts to address this problem have mainly involved training with degraded visual data. However, visual degradation can occur in many unpredictable ways, making it impractical to simulate all possible cases during training. In this paper, we aim to enhance the robustness of audio-visual speaker extraction against impaired visual inputs without relying on degraded videos during training. Inspired by observations from human perceptual mechanisms, we propose an audio-visual learner that disentangles speaker information, acoustic synchronisation, and semantic synchronisation as distinct cues. Furthermore, we design a dedicated interaction module that effectively integrates these cues to provide a reliable guidance signal for speaker extraction. Extensive experiments demonstrate the strong robustness of the proposed model under various visual degradations and its clear superiority over existing methods.
Acute decompensated heart failure (ADHF) has been shown to affect breathing and speech production. We present a new German speech dataset and conduct a pilot study to show the feasibility of using speech to distinguish between decompensated and recompensated ADHF. We achieve accuracy rates of 77% for this binary classification using read speech, with the use of segmented phrases scoring at 75% and sustained vowels trailing far behind at 52%. Our feature analysis shows a reduction in breathiness and an increase in articulation stability during recompensation. Overall, our results show the promise of speech analysis in the monitoring of ADHF and demonstrate that further efforts should be invested to obtain large-scale data and replicate our results in more heterogeneous populations.
Recent advances in multimodal large language models (MLLMs) have demonstrated remarkable multi- and cross-modal integration capabilities. However, their potential for fine-grained emotion understanding remains systematically underexplored. While open-vocabulary multimodal emotion recognition (MER-OV) has emerged as a promising direction to overcome the limitations of closed emotion sets, no comprehensive evaluation of MLLMs in this context currently exists. To address this, our work presents the first large-scale benchmarking study of MER-OV on the OV-MERD dataset, evaluating 19 mainstream MLLMs, including general-purpose, modality-specialized, and reasoning-enhanced architectures. Through systematic analysis of model reasoning capacity, fusion strategies, contextual utilization, and prompt design, we provide key insights into the capabilities and limitations of current MLLMs for MER-OV. Our evaluation reveals that a two-stage, trimodal (audio, video, and text) fusion approach achieves optimal performance in MER-OV, with video emerging as the most critical modality. We further identify a surprisingly narrow gap between open- and closed-source LLMs. These findings establish essential benchmarks and offer practical guidelines for advancing open-vocabulary and fine-grained affective computing, paving the way for more nuanced and interpretable emotion AI systems. Associated code will be made publicly available upon acceptance.
Emotion is essential in spoken communication, yet most existing frameworks in speech emotion modeling rely on predefined categories or low-dimensional continuous attributes, which offer limited expressive capacity. Recent advances in speech emotion captioning and synthesis have shown that textual descriptions provide a more flexible and interpretable alternative for representing affective characteristics in speech. However, progress in this direction is hindered by the lack of an emotional speech dataset aligned with reliable and fine-grained natural language annotations. To tackle this, we introduce AffectSpeech, a large-scale corpus of human-recorded speech enriched with structured descriptions for fine-grained emotion analysis and generation. Each utterance is characterized across six complementary dimensions, including sentiment polarity, open-vocabulary emotion captions, intensity level, prosodic attributes, prominent segments, and semantic content, enabling multi-granular modeling of vocal expression. To balance annotation quality and scalability, we adopt a human-LLM collaborative annotation pipeline that integrates algorithmic pre-labeling, multi-LLM description generation, and human-in-the-loop verification. Furthermore, these annotations are reformulated into diverse descriptive styles to enhance linguistic diversity and reduce stylistic bias in downstream modeling. Experimental results on speech emotion captioning and synthesis demonstrate that models trained on AffectSpeech consistently achieve superior performance across multiple evaluation settings.
In this work, we introduce CoughPhase-CLR, a self-supervised learning framework designed to leverage the physiological phases of a cough for robust representation learning. Unlike generic contrastive frameworks, CoughPhase-CLR constructs positive pairs based on these specific acoustic phases. We pre-trained our model on approximately 40 hours of public cough audio and evaluated it across five downstream tasks, including COVID-19 detection, chronic obstructive pulmonary disease (COPD) state classification, and smoker status prediction. Our results demonstrate that cough-specific pre-training consistently outperforms standard random-cropping techniques when training on cough recordings. Additionally, we benchmarked a diverse set of state-of-the-art models on COPD state classification, highlighting the difficulty of this task. The best-performing models, pretrained on either general audio or respiratory sounds, achieved a UAR of 57%, failing to outperform the state-of-the-art performance of 84% UAR achieved using speech analysis.
Speech conveys rich emotional information. As Speech Emotion Recognition (SER) is usually deployed in privacy-sensitive and reliability-critical environments, adversarial attacks on SER have attracted increasing attention. Existing sparse attacks control the number of perturbed elements, yet, they often lack explainability guidance and explicit measures of explanation consistency. A unified treatment of sparsity and magnitude constraints is also uncommon. In addition, transferability across attack families and target models remains limited. Hence, we propose a SalIency-Guided sparse Mask Attack (SIGMA). On self-supervised speech features, we use post-hoc explainable artificial intelligence (XAI) techniques to produce saliency maps and identify the scope of the mask, and then restrict magnitude-bounded updates to this mask. The mask is computed once and can be reused across models and different sparsity attacks to amortise cost. We evaluate on the IEMOCAP and TESS datasets. Under matched budgets and across multiple sparse-attack settings, SIGMA maintains competitive attack success rates, navigating a conscious trade-off between attack efficacy and explanation consistency. SIGMA therefore provides an efficient and interpretable framework for analysing the vulnerability and explanation behaviour of SER models under structured perturbations.
Reinforcement learning is a powerful learning paradigm that has spearheaded progress in numerous domains. Its core promise lies in learning through high-level goals without the need for granular labels. However, it still remains elusive in the realm of audio, where it has received substantially less attention than in computer vision or other domains. The key question remains: how can agents learn to listen purely via reward-driven exploration? In this contribution, we present an overview of previous attempts and a new conceptual framework for learning to listen by reward. Our approach depends on the continuous search for novel sound sources. We formulate our framework, discuss open technical challenges, and present a first proof-of-concept implementation that showcases the feasibility of our approach.
Machine-generated music (MGM) has become a groundbreaking innovation with wide-ranging applications, such as music therapy, personalised editing, and creative inspiration within the music industry. However, the unregulated proliferation of MGM presents considerable challenges to the entertainment, education, and arts sectors by potentially undermining the value of high-quality human compositions. Consequently, MGM detection (MGMD) is crucial for preserving the integrity of these fields. Despite its significance, MGMD domain lacks comprehensive systematic evaluation results necessary to drive meaningful progress. To address this gap, we conduct experiments on existing large-scale datasets using a range of foundational models for audio processing, establishing systematic evaluation results tailored to the MGMD task. Our selection includes traditional machine learning models, deep neural networks, Transformer-based architectures, and State space models (SSM). Recognising the inherently multimodal nature of music, which integrates both melody and lyrics, we also explore fundamental multimodal models in our experiments. Beyond providing basic binary classification outcomes, we delve deeper into model behaviour using multiple explainable Artificial Intelligence (XAI) tools, offering insights into their decision-making processes. Our analysis reveals that ResNet18 performs the best according to in-domain and out-of-domain tests. By providing a comprehensive comparison of systematic evaluation results and their interpretability, we propose several directions to inspire future research to develop more robust and effective detection methods for MGM. We provide our codes and some samples on Github repository https://github.com/myxp-lyp/Detecting-Machine-Generated-Music-with-Explainability-A-Challenge-and-Systematic-Evaluation .