
This paper investigates the effectiveness of using synthetic audio, generated from musical scores via virtual instrument software, for training automatic guitar transcription models. Collecting large annotated datasets from real performances is costly and labor-intensive. To overcome this, the use of synthetic data has been explored for developing automatic music transcription (AMT) models. We present a systematic comparison between a high-resolution AMT models trained on synthetic guitar data (SynthTab) and models trained on real data (GAPS), clarifying the effectiveness and limitations of synthetic data in AMT. Our experiments yield four insights: 1) models trained on diverse and well-augmented synthetic guitar data generalize well to real recordings, 2) pre-training with target instrument’s synthetic data is more effective than with non-target instrument’s real data, 3) even small synthetic datasets are valuable for pre-training, and 4) transcription discrepancies arise between models trained on synthetic data and real data, which can be mitigated by fine-tuning. These results demonstrate the utility of synthetic data for training AMT models when real aligned data is scarce.
Speech Activity Detection (SAD) systems often misclassify singing as speech, leading to degraded performance in applications such as dialogue enhancement and automatic speech recognition. We introduce Singing-Robust Speech Activity Detection (SR-SAD), a neural network designed to robustly detect speech in the presence of singing. Our key contributions are: i) a training strategy using controlled ratios of speech and singing samples to improve discrimination, ii) a computationally efficient model that maintains robust performance while reducing inference runtime, and iii) a new evaluation metric tailored to assess SAD robustness in mixed speech-singing scenarios. Experiments on a challenging dataset spanning multiple musical genres show that SR-SAD maintains high speech detection accuracy (AUC = 0.919) while rejecting singing. By explicitly learning to distinguish between speech and singing, SR-SAD enables more reliable SAD in mixed speech-singing scenarios.
In loudspeaker-based sound field reproduction, the perceived sound quality deteriorates significantly when listeners move outside of the sweet spot. Although a substantial increase in the number of loudspeakers enables rendering methods that can mitigate this issue, such a solution is not feasible for most real-life applications. This study aims to extend the listening area by finding a panning strategy that optimises an objective function reflecting the localisation and localisation uncertainty over a listening area. To that end we first introduce a psychoacoustic localisation model that outperforms existing models in the context of multichannel loudspeaker setups. Leveraging this model and an existing model of localisation uncertainty, we optimise inter-channel time and level differences for a stereophonic system. The outcome is a new panning approach that depends on the listening area and the most suitable trade-off between localisation and localisation uncertainty.
We introduce a region-of-interest beamforming approach for audio conferencing that addresses dynamic acoustics and multiple-speaker scenarios. The approach employs a two-stage sparse optimization to select a subset of microphones from dual circular sector arrays: first on the xz plane and then on the xy plane, balancing spatial resolution and efficiency. Using the dual circular layout, we are able to reduce response variability across azimuth and elevation angles. The proposed approach maximizes broadband directivity while ensuring a controlled level of distortion and minimal white noise gain. Compared to existing methods, the mainlobe attained by the resulting beamformer is more accurately aligned with the region of interest. It also achieves a preferable sidelobe and backlobe suppression. Finally, the proposed approach is shown to be superior considering the directivity factor and white noise gain, in particular at medium and high frequencies.
This paper investigates how inference step size in sliding window approaches affects music source separation quality. Through systematic analysis of seven model configurations across five architectures, we demonstrate that increased segment overlap consistently improves separation quality by up to 0.37 dB SDR. We identify a universal pattern where performance improves logarithmically with overlap, with an "elbow point" at 4-8 overlapping segments where efficiency begins to decrease rapidly. Our analysis reveals that: (1) state-of-the-art papers inconsistently report inference overlap settings, making fair comparisons difficult; (2) even modest overlap settings (25%) substantially improve quality through boundary artifact reduction; and (3) higher-performing models show proportionally greater improvements from increased overlap when accounting for the logarithmic nature of SDR. These findings suggest that standardized overlap reporting is essential for meaningful architectural comparisons and that differences attributed to architectural innovations may partly stem from undisclosed inference settings.
Traditional zero-shot voice conversion methods typically extract a speaker embedding from a reference recording first and then generate the source speech content in the target speaker’s voice by conditioning on that embedding. However, this process often overlooks time-dependent speaker characteristics, such as voice dynamics and speaking rates, as well as environmental acoustic properties of the reference recording. To address these limitations, we propose a one-shot voice conversion framework capable of replicating not only voice timbre but also acoustic properties. Our model is built upon Diffusion Transformers (DiT) and conditioned on a designed content representation for acoustic cloning. Besides, we introduce specific augmentations during training to enable accurate speaking rate cloning. Both objective and subjective evaluations demonstrate that our method outperforms existing approaches in terms of audio quality, speaker similarity, and environmental acoustic similarity, while effectively capturing the speaking rate distribution of target speakers. Audio samples are available at: ditvc.github.io.
In this paper we present the TAVA, a novel dataset for advancing research in Speech Emotion Recognition (SER) by disentangling paralinguistic and linguistic information in affective speech signals. The dataset includes 352 audio recordings of emotionally expressive English speech, each paired with a corresponding transformed electroglottographic (tEGG) version—a signal designed to preserve affective cues while systematically suppressing phonetic content. In addition, we provide over 120,000 crowd-sourced ratings of valence, arousal, and dominance for both the original and transformed signals. These ratings support fine-grained comparisons of affect perception across modalities. Building on prior work showing that tEGG signals can effectively isolate vocal affect, this dataset offers a unique resource for evaluating sensitivity to vocal affect in clinical populations with language or communication difficulties, as well as in studies aimed at dissociating linguistic and affective processing in individuals without these impairments. By contributing a phoneme-reduced, affect-rich signal representation to the SER community, we aim to enable more robust modelling of vocal affect and broaden the applicability of SER systems to diverse user populations.
A method for source (secondary loudspeaker) and sensor (control point) placement in sound field control is proposed. Since the placement of secondary loudspeakers and control points has a significant effect on the performance of sound field control, their optimization is of great importance in the development of practical systems. However, most current methods address either source placement or sensor placement. We propose a joint source and sensor placement method based on mean square error, which is applicable to (weighted) pressure matching. Our proposed method can be applied when the area where the sensors can be placed in the target control region is limited. Furthermore, prior information about the desired sound field can be incorporated. Numerical evaluation showed that efficient source and sensor placement can be achieved by the proposed method.
Cross-talk cancellation (CTC) is a well-established technique for delivering binaural audio over loudspeakers. Traditional two-channel CTC systems are commonly implemented using networks of Finite Impulse Response (FIR) filters, with causality ensured through the use of modelling delays and regularisation techniques. Recursive implementations, which rely on a feedback network, have been proposed as an alternative for two-channel systems, offering potential benefits in computational efficiency. This paper investigates whether a similar recursive architecture can be extended to multi-channel CTC systems, particularly those employing linear arrays of loudspeakers. Through theoretical analysis, we demonstrate that recursive multi-channel CTC implementations are intrinsically non-causal under general conditions, making a direct real-time realisation infeasible without significantly compromising the system performance.
In this work, we propose a novel approach towards privacy-preserving audio-visual speech recognition (AV-ASR). We apply feature-wise additive and multiplicative masks to the latent embeddings of a pre-trained AV-ASR model and fine-tune subsequent layers of this model using AdaLoRA. The masks are trained to preserve linguistic content necessary for ASR while degrading speaker-discriminative cues. We derive the masking mechanism directly from the sequential input and optimize it via a contrastive representation learning (CRL) objective. The method closely aggregates slightly perturbed representations of the same utterance while simultaneously increasing the distance to representations of different utterances from the same speaker.Experiments on LRS3 and VoxCeleb2 show that our approach maintains competitive word error rates (WER) while significantly reducing speaker identification performance, as measured by Equal Error Rate (EER) under a strong audio-visual speaker verification attack. These results demonstrate the potential of contrastive fine-tuning for privacy-preserving AV-ASR for edge devices.
Despite advances in hearing technologies, users still face challenges in acoustically complex environments. In particular, social situations with multiple speakers—where a given speaker may be attended to at one moment and become an interfering source at another—are consistently reported as especially difficult. We use a data-driven approach to estimate the engagement level of the user with each sound source in the acoustic foreground based on behavioral cues within dynamic, real-world acoustic scenes. We collect a novel dataset from two experiments where subjects can behave naturally and attended sources can be recorded, covering both induced and free listening intentions. For each source, we engineer features expected to reflect the user’s engagement, without making further assumptions about underlying patterns, which are instead inferred with machine learning. Classic models such as logistic regression and tree-based methods achieve an average precision of 97% when the listening intention is induced and 76% when it is free, surpassing baselines that statically predict the most frequently attended sources. Our results suggest that it is possible to estimate user engagement with individual localized acoustic sources using machine learning, without invasive sensors and with sufficient accuracy for practical application. This paves the way for the development of adaptive hearing systems that can support users in everyday situations.
Six-degrees-of-freedom audio rendering for eXtended Reality (XR) from limited measurements remains a challenging problem, particularly in coupled spaces where anisotropic and multi-slope late reverberation occurs. In this work, we adopt the common slopes model to represent position-dependent, directional late reverberation in a coupled space. Fourier-encoded positional coordinates are used as input to a Multi-Layer Perceptron (MLP) to predict position-dependent common-slope amplitudes in the spherical harmonic (SH) domain. The MLP is trained using a directional energy decay curve (DEDC) loss. At inference, the MLP predicts directional amplitudes of the DEDCs at unseen positions. These amplitudes, combined with the pre-determined common decay times, can drive a modal or shaped white noise reverberator to synthesise directional late reverberation tails at new locations. We demonstrate binaural rendering for 6DoF navigation and evaluate our approach on a simulated three-room dataset. Results show that an SRIR grid spacing of 60 cm is sufficient to obtain a mean EDC error below 1.3 dB. We also compare our method to the Neural Acoustical Field approach and achieve significantly faster inference with comparable EDC mismatch errors.
Recent advancements in deep neural network (DNN)-based hearing aids (HAs) have significantly improved hearing-restoration performance. A recent study introduced a DNN-based framework for designing HA models tailored to the sensorineural hearing loss (SNHL) profile of individual users, improving treatment outcomes for diverse HA users. While these bio-inspired HA models perform well in clean speech conditions, their effectiveness diminishes in noisy environments. Moreover, their high computational complexity makes deployment on resource-constrained devices challenging, particularly when additional noise reduction modules are added. To address these limitations, we propose a low-complexity, real-time individualized noise reduction (INR) model that jointly performs noise suppression and hearing loss compensation. The model is optimized for deployment on embedded systems using quantization to reduce model size and computational load. To mitigate the impact of quantization error, we applied quantization-aware training (QAT), which simulates the quantization effect during the training phase. Experimental results show that the proposed INR model outperforms existing closed-loop bio-inspired HA models in noisy conditions, achieving superior objective measures of speech intelligibility and sound quality. The INT8 quantized version maintains performance comparable to the full-precision model. With a system latency of just 8 ms, the model meets the real-time requirements of hearing aid applications. These findings demonstrate the potential of the proposed INR model for real-time application in hearing aids and other low-power hearing devices.
Mismatch in acoustics between users is an important challenge for interaction in shared XR environments. It can be mitigated through acoustic matching, which traditionally involves dereverberation followed by convolution with a room impulse response (RIR) of the target space. However, the target RIR in such settings is usually unavailable. We propose to tackle this problem in an end-to-end manner using wave-u-net encoder-decoder network with potential for real-time operation. We use FiLM layers to condition this network on the embeddings extracted by a separate reverb encoder to match the acoustic properties between two arbitrarily chosen signals. We demonstrate that this approach outperforms two baseline methods and provides the flexibility to both dereverberate and rereverberate audio signals.
Sound field control using loudspeaker arrays is an important acoustic and audio signal processing applications. In sound field control, least squares (LS) regression based on pressure matching or mode-matching is typically introduced to derive the driving signals of loudspeakers as a closed-form solution. The LS regression is a maximum-likelihood estimation, in which the error is assumed to be Gaussian distribution. Compared with the LS regression, the least absolute deviation (LAD) regression, in which the error is assumed to be Laplace distribution, is robust against outliers. In pressure matching-based sound field methods, outliers appear at higher frequencies according to the spatial Nyquist frequency. To improve the control accuracy for pressure matching-based methods at high frequencies, this paper proposes SFC-L1, pressure matching-based sound field control method with LAD regression instead of LS regression. In the proposed method, the LAD regression combined with L1 regularization is solved with gradient method simply implemented on PyTorch. The results of computer simulations demonstrate that the proposed LAD-based methods can improve the sound field control accuracy at high frequencies compared with the conventional LS-based methods. Additionally, PyTorch-based implementation, Torch-SFC, is open-sourced for accelerating sound field control research.
Separating the individual elements in a musical mixture is an essential process for music analysis and practice. While this is generally addressed using neural networks optimized to mask or transform the time-frequency representation of a mixture to extract the target sources, the flexibility and generalization capabilities of generative diffusion models are giving rise to a novel class of solutions for this complicated task. In this work, we explore singing voice separation from real music recordings using a diffusion model which is trained to generate the solo vocals conditioned on the corresponding mixture. Our approach improves upon prior generative systems and achieves competitive objective scores against non-generative baselines when trained with supplementary data. The iterative nature of diffusion sampling enables the user to control the quality-efficiency trade-off, and also refine the output when needed. We present an ablation study of the sampling algorithm, highlighting the effects of the user-configurable parameters.
Diffusion approaches to speech enhancement gain immense attention due to their improved speech component reconstruction capability. However, they usually suffer from high inference computational complexity due to multiple iterations during the reverse process. In this paper, we propose EffDiffSE, an efficient diffusion-based frequency-domain speech enhancement model with a hybrid discriminative condition DNN and generative score DNN. Our contributions are three-fold. First, we formulate the powerful time-domain Universe++ model in the frequency domain with a combined psychoacoustic loss and score matching loss. Second, we achieve a single-step efficient reverse process both during training and inference with noise-consistent Langevin dynamics. Third, an auxiliary loss is applied to the single-step reverse process output to improve the diffusion performance further. Trained and evaluated on the URGENT 2024 Speech Enhancement Challenge data splits, the proposed EffDiffSE achieves an MOS comparable to the top reported time- and frequency-domain diffusion baseline methods, while excelling them by >0.2 PESQ points, showing significantly less hallucination and inference computational complexity (3.9 GMAC/s vs. 58 ... 7600 GMAC/s).
Because suitable real-world room impulse responses (RIRs) are scarce for many use cases, synthetic RIRs are indispensable to data-driven audio applications; however, any mismatch between them and real-world RIRs can cause significant performance degradation due to domain shift. This study proposes a device-centric RIR augmentation method that reduces this mismatch by incorporating the acoustic characteristics of transducers from a specific audio device along with a reverberation model to generate synthetic RIRs. As an application example, we evaluate the effect of the proposed augmentation technique on the problem of room geometry inference (RGI) using a smart speaker. A neural network is trained with and without augmented RIRs and tested with measured RIRs. Comparing the performance of both training scenarios indicates a reduced domain shift between augmented synthetic RIRs and real-world measurements. The results show a considerable improvement in the estimation accuracy when augmented RIRs are used for training.
We demonstrate that the Joint-Embedding Predictive Architecture is effective for learning representations suitable for Music Information Retrieval tasks. Specifically, we explore its application to multi-instrument automatic music transcription, focusing on multi-pitch estimation and instrument recognition. We evaluate the learned representations across multiple settings: (1) finetuning a pretrained JEPA model with transcription supervision, (2) end-to-end training with transcription supervision, (3) training an instrument-aware transcriber on frozen JEPA embeddings and (4) training an instrument-agnostic transcriber on frozen JEPA embeddings. To assess the structure of the learned representations, we compute Calinski-Harabasz clustering scores with respect to pitch index, pitch class, instrument, and octave. We find that the representations learned by JEPA and its modified version (2), primarily capture instrument identity and pitch height information, rather than pitch class distinctions. Despite this, our results demonstrate promising transcription performance and highlight the potential of non-generative self-supervised learning for multi-instrument music transcription. Code and model configurations are available on GitHub.1
The restoration of degraded audio signals is often performed on complex-valued frequency-domain (FD) representations. This requires manipulation of either magnitudes and phases or real and imaginary parts. In general, these manipulations do not produce consistent representations. The consequence is that the magnitudes and phases (or real and imaginary parts) of the restored time-domain signal (which are always consistent) do not match the generally inconsistent values imposed during FD restoration. In colloquial terms, "What we get is not what we asked for." The enforcement of consistency is always heard in the resulting audio and is known in principle, but it can be better understood. We present two-dimensional FD SNR frameworks (e.g., magnitude/phase or real/imaginary) that visually reveal how consistency enforcement changes the applied FD restorations to arrive at the achieved FD restorations. We also show how extended Griffin-Lim algorithms can reduce and direct, but not eliminate, the changes produced by consistency enforcement. We apply objective estimators to connect this work to estimated speech quality and intelligibility. This work can inform machine learning training and architecture choices that must balance restoration efforts across two dimensions (e.g., magnitude and phase) to arrive at the best possible speech quality.