
Although the roles of the torso and head-above-torso orientation (HATO) in human spatial hearing have been thoroughly studied in psychoacoustic research, binaural sound source localization (SSL) models generally treat head and torso as a rigid unit, not reflecting natural, independent head-torso postures or their dynamics during rotation. This study is the first to systematically assess HATO in deep-learning-based binaural SSL for both static scenes and under dynamic rotations. Convolutional recurrent networks are trained with and without varying HATO as well as with and without explicit head-orientation input (scalar angle or quaternion) across three configurations: joint head-and-torso rotation with fixed HATO, torso rotation below a fixed head, and head rotation above a fixed torso. Performance is measured via a set of complementary localization metrics including the mean great circle distance, coordinate-isolated angular components and the quadrant error rate on simulated and measured data, with analyses with respect to signal-to-noise ratio, sound-incidence direction, and head-rotation velocity. In static scenarios, matched single-HATO training reduces mean great circle distance and quadrant error rate by about 24
Data augmentation is crucial for automatic speech recognition, especially in low-resource settings. This paper introduces a time-domain augmentation method, FadeOutIn, designed to better capture the dynamics of spontaneous speech and remain compatible with classical techniques. Experimental results show that combining FadeOutIn with SpecAugment yields a relative improvement of 4.24
Music Emotion Recognition (MER) is a computational field of affective computing and audio signal processing. Although previous attempts used traditional machine learning algorithms, e.g., Support Vector Machines and k-Nearest Neighbors, for emotion classification in music, these methods are often challenged by the intricate temporal and spectral nature of sound signals, thereby constraining classification performance. To address this, the work proposes a new model, the Lion-Optimized CNN-BiGRU for Emotion Recognition (LOCBER), leveraging the strengths of Convolutional Neural Networks and Bidirectional Gated Recurrent Units for effective feature extraction and temporal sequence modeling. For additional performance optimization, LOCBER is optimized using the Lion Swarm Optimization (LSO) algorithm, which adjusts model parameters to minimize loss and achieve better accuracy and convergence. This study aims to improve music emotion recognition by integrating a CNN–BiGRU model with Lion Swarm Optimization, achieving 97.7
Domain shifts, such as changes in language, noise types, or recording environments, can significantly degrade the performance of speech enhancement systems. Most current research focuses on domain adaptation methods. While existing research often assumes that domain shifts have already been identified, this study addresses the crucial challenge of detecting them automatically. We propose a novel domain-shift detection method that monitors prediction uncertainty in a speech quality assessment network. Our method includes both supervised and unsupervised approaches. The supervised method requires access to clean speech, whereas the unsupervised method does not. Experimental results across various domain-shift scenarios demonstrate that our method effectively identifies domain mismatches, enabling timely adaptation to improve performance in real-world speech enhancement systems. We use 100 s of noisy speech to detect actionable domain shifts, typically achieving an SI-SDR improvement of 3–6 dB after adaptation. To the best of our knowledge, this is the first method for domain-shift detection in speech enhancement, offering both supervised and unsupervised variants with broad applicability to real-world systems.
Despite advances in cochlear implant (CI) technology, music enjoyment remains a significant challenge for most CI users. To facilitate music listening in CI-mediated hearing, music processing algorithms have been proposed, which reduce the perceived complexity of music pieces. As these algorithms rely on a multitude of signal processing parameters, finding suitable and possibly genre-specific settings can be a tedious process. Therefore, in this study, we propose and evaluate an interactive optimization scheme for parametric music processing algorithms. To minimize the burden on CI users, the process was divided into two phases: first, normal-hearing (NH) participants evaluated a wide range of random parameter combinations using a vocoder-based CI simulation. Then, CI users participated in shorter, focused optimization sessions to refine these parameters. For comparison, we also conducted parallel listening sessions with NH listeners using the vocoder simulation. We tested the approach using a parametric music remixing algorithm that had previously been validated with a hand-crafted parameter setting found through expert knowledge. Starting from completely randomized initial settings, the method was able to find optimized parameter values through interactive user feedback. Both groups, CI listeners and NH listeners using a CI simulation, significantly preferred the optimized remixes over unprocessed music, while for CI listeners, the preference for the optimized settings was on par with the hand-crafted setting. NH listeners using the CI simulation even significantly preferred their optimized remixes. The results of this study demonstrate the utility of the proposed optimization scheme and pave the way for time-efficient, interactive tuning of future music processing algorithms.