
Brain-assisted speech enhancement (BASE) aims to extract the attended speech by leveraging electroencephalogram (EEG) signals as assistive clues, which shows great potential in neuro-steered listening devices, e.g., hearing aids. Many models have been proposed for this task but most suffer from severe overfitting and overestimation due to unreasonable dataset splitting and inherent limitations in datasets. In particular, within-trial splitting causes severe attention leakage unless the dataset incorporates within-trial attention switches. Moreover, the underlying principle of BASE and its relation to auditory attention decoding (AAD) have not yet been fully explored. It is expected that the extracted speech should be consistent with the decoded auditory attention. In this work, we record a new EEG-Audio dataset supporting both BASE and AAD, featuring greater trial diversity and frequent attention switching between trials to prevent cross-trial attention leakage. We systematically compare several state-of-the-art end-to-end BASE models using a more rigorous data splitting. For fairness of comparison, we propose to include Pearson correlation coefficient and AAD accuracy to judge the BASE performance. In order to reveal the explicit relation between BASE and AAD, we develop an AAD-driven BASE framework from an alternative perspective. Results on our dataset and the KUL dataset show that current BASE methods cannot perform as well as claimed in publications. The end-to-end BASE models implicitly perform AAD and speech separation, followed by target speech selection based on the AAD decision. The actual bottleneck of BASE lies in the inherent AAD mechanism, and the performance has a large improvement space.
Computational Auditory Scene Analysis (CASA) aims to detect what sounds appear and when they occur in an audio recording, as well as to separate the audio into waveforms corresponding to the identified sound classes. The CASA problem presents several challenges. First, previous source separation systems require training on clean data, which is often scarce. Second, existing conditional source separation systems are designed to separate only a limited number of sound classes and cannot scale to separate hundreds of classes. Third, prior works require users to specify the sounds to be separated and cannot automatically detect which sound classes to separate. Fourth, there is a limited research on building hierarchical source separation systems. The contributions of this work are as follows. First, we propose training universal source separation (USS) systems on large-scale weakly labeled and unlabelled datasets, rather than relying solely on clean data. Second, we develop a large-scale USS system capable of separating up to 527 sound classes, with the potential to scale to unlimited number of classes. Third, we introduce a method to automatically detect and separate sound classes without user specification. Fourth, we present a hierarchical USS system that can separate sound classes at various hierarchical levels. We train the USS systems on AudioSet and evaluate their performance on speech, audio, and music datasets, demonstrating the effectiveness of our approach.
In neural scaling, scaling laws have profoundly shaped our understanding of model performance across easily scalable and expressible parameters. However, when scaling involves multiple parameters that present complex expressions or are difficult to express, it becomes challenging to formulate scaling functions that encompass such parameters. In this work, We found that in the audio domain, the embedding effective rank (RankMe) can serve as a unifying metric that enables label-free, information-theoretic quantification of audio embeddings, encapsulates the impact of diverse variables including masking rate, model size, data volume, computational budget, and architectural configurations on representation quality and demonstrates more significant effectiveness compared to traditional scaling laws. Our empirical findings reveal a consistent power-law relationship between RankMe and representation quality, suggesting that embedding effective rank serves as a reliable proxy for assessing model performance in audio representation learning. Based on RankMe as a label-free proxy metric, our work demonstrates its role as an early-stage monitoring tool. Furthermore, we provide an empirically validated framework for early screening of optimal models from multiple complex hyperparameter settings during pre-training, which saves computational resources.
Generative speech enhancement methods based on generative adversarial networks (GANs) have demonstrated promising performance across various speech enhancement tasks. However, their performance in very low signal-to-noise ratio (SNR) scenarios remains under-explored and limited, as these conditions pose significant challenges to both discriminative and generative state-of-the-art methods. To address this, we propose DisCoGAN, a GAN-based speech enhancement method that leverages latent features extracted from discriminative speech enhancement models as generic conditioning information. By incorporating the proposed discriminative conditioning method, DisCoGAN improves speech quality and intelligibility, particularly in low-SNR scenarios, while maintaining competitive or superior performance in high-SNR conditions and real-world recordings. We also conduct a comprehensive evaluation of conventional GAN-based architectures, including end-to-end GANs, GAN-first, and post-filtering GANs, as well as discriminative models under low-SNR conditions, and show that DisCoGAN consistently outperforms existing methods. Finally, we present ablation studies that highlight the performance gains from discriminative conditioning and demonstrate how DisCoGAN leverages both local and global temporal context, providing insight into the key factors underlying these gains.
Controllable human voice generation, particularly for expressive domains like singing, remains a significant challenge. This paper introduces Vevo2, a unified framework for controllable speech and singing voice generation. To tackle issues like the scarcity of annotated singing data and to enable flexible controllability, Vevo2 introduces two audio tokenizers: (1) a unified music-notation-free prosody tokenizer that captures prosody and melody from speech, singing, and even instrumental sounds, and (2) a unified content-style tokenizer that encodes linguistic content, prosody, and style for both speech and singing, while enabling timbre disentanglement. Vevo2 consists of an auto-regressive content-style modeling stage, which aims to enable controllability over text, prosody, and style, as well as a flow-matching acoustic modeling stage that allows for timbre control. Particularly, during the speech-singing joint training of the AR model, we propose both explicit and implicit prosody learning strategies to bridge speech and singing voice. Moreover, to further enhance the Vevo2's ability to follow text and prosody, we design a multi-objective post-training task that integrates both intelligibility and prosody similarity alignment. Experimental results show that the unified modeling in Vevo2 brings mutual benefits to both speech and singing voice generation. Additionally, Vevo2's effectiveness across a wide range of synthesis, conversion, and editing tasks for both speech and singing further demonstrates its strong generalization ability and versatility.
Audiobook generation aims to create rich, immersive listening experiences from multimodal inputs, but current approaches face three critical challenges: (1) the lack of synergistic generation of diverse audio types (e.g., speech, sound effects, and music) with precise temporal and semantic alignment; (2) the difficulty in conveying expressive, fine-grained emotions, which often results in machine-like vocal outputs; and (3) the absence of automated evaluation frameworks that align with human preferences for complex and diverse audio. To address these issues, we propose Dopamine Audiobook, a unified training-free multi-agent system, where a multimodal large language model (MLLM) serves two specialized roles (i.e., speech designer and audio designer) for emotional, human-like, and immersive audiobook generation and evaluation. Specifically, we firstly propose a stage-wise, context-aware framework for diverse audio generation with word-level semantic and temporal alignment. To enhance expressiveness, we then design word-level paralinguistic augmentation, utterance-level prosody retrieval, and adaptive TTS model selection. Finally, for evaluation, we introduce an MLLM-based evaluation framework incorporating self-critique, perspective-taking, and psychological MagicEmo prompts to ensure human-aligned and self-aligned assessments. Moreover, we build MIA-Bench, the first benchmark for multimodality immersive audiobook generation, comprising 300 data pairs with rich annotations. Experimental results demonstrate that our method achieves state-of-the-art (SOTA) performance on multiple metrics, while our evaluation framework demonstrates superior alignment with human preferences.
Large language models (LLMs) have exhibited impressive multilingual reasoning capabilities, driven by extensive multilingual pre-training corpora and instruction fine-tuning data. However, a performance gap exists between high- and low-resource language reasoning tasks due to the language imbalance in the pre-training corpus, which is exacerbated by evaluation bias in existing reasoning benchmarks lacking low-resource language coverage. To alleviate this issue, we propose LinguaLIFT, a two-stage instruction tuning framework for advancing low-resource language reasoning. LinguaLIFT employs a language alignment layer to capture multilingual alignment in a code-switched tuning way without requiring multilingual instruction or parallel data, thereby transferring the cross-lingual reasoning capabilities to low-resource languages through English-only instruction tuning data. To comprehensively evaluate the multilingual reasoning capabilities, we introduce the Multilingual Math World Problem (MMWP) benchmark, which spans 21 low-resource, 17 medium-resource, and 10 high-resource languages. Experimental results show that LinguaLIFT outperforms several competitive baselines across MMWP and four widely used benchmarks.
We propose a novel dual-path state-space model within an encoder-decoder architecture for multichannel speech enhancement that leverages joint temporal and spectral modeling to significantly improve speech quality in noisy and reverberant environments. At the core of our framework is the S5 state-space model, which efficiently captures complex temporal dependencies across multiple speech channels by modeling both short and long-term dynamics. To effectively integrate spatial and spectral features, our encoder employs S3Conv layers that extract salient characteristics from the raw input, while a dedicated cross-domain interaction mechanism facilitates a dynamic exchange of information between two parallel data streams used for coarse magnitude estimation and complex spectral refinement, respectively. This dual-path design enables the network to jointly enhance amplitude and phase information, resulting in improved perceptual quality and intelligibility. Extensive experiments on public datasets demonstrate that our model outperforms state-of-the-art methods across multiple evaluation metrics. Ablation studies further validate the effectiveness of each component in the overall architecture, confirming that the integration of state-space-based temporal sequence modeling and cross-domain feature fusion is critical for robust, high-quality speech enhancement. Our results also indicate that the proposed framework is well-suited for real-world applications where computational efficiency and superior performance are essential.
In environments such as open offices and co-working spaces, the adverse effects of speech as a primary source of ambient noise extend beyond mere annoyance, contributing to reduced productivity, increased workload and stress. When dealing with speech, active noise control (ANC) systems have difficulties arising from constraints within the ANC system alongside the non-stationary nature of speech. These constraints require the optimal filters to be non-causal, creating a prediction problem. The non-causality is due to the delay incurred by, e.g., digital processing or acoustic propagation paths. To deal with this, we propose prediction-based feedforward ANC for headphone applications with improved voiced speech attenuation. The proposed methods leverage various aspects of speech characteristics and structure, including an optimal linear prediction (LP) scheme based on short- and long-term speech correlations, sparse LP modelling, harmonic and harmonic-chirp model-based prediction to account the non-stationarity of speech, and multiple-frequency ANC with harmonic decomposition of speech. Simulation results demonstrate the superior performance of the proposed methods compared to conventional adaptive feedforward ANC across a wide range of delays, spanning from 1 to 60 samples at a sampling frequency of 48 kHz (0.02 - 1.25 ms). The proposed methods achieve attenuation improvements of up to 8 dB and significantly extend attenuation to higher frequencies. Results from the perceptual study assessing subjective satisfaction with prediction-based ANC reveal that the artifacts from speech prediction have a more significant impact on satisfaction compared to attenuation.
Continual Named Entity Recognition (CNER) is an evolving field that focuses on sequentially updating an existing model to incorporate new entity types. Previous CNER methods primarily utilize Knowledge Distillation (KD) to preserve prior knowledge and overcome catastrophic forgetting, strictly ensuring that the representations of old and new models remain consistent. Consequently, they often impart the model with excessive stability (i.e., retention of old knowledge) but limited plasticity (i.e., acquisition of new knowledge). To address this issue, we propose a Stability-Plasticity Trade-off (SPT) method for CNER that balances these aspects from both representation and weight perspectives. From the representation perspective, we introduce a pooling operation into the original KD, permitting a level of plasticity by consolidating representation dimensions. From the weight perspective, we dynamically merge the weights of old and new models, strengthening old knowledge while maintaining new knowledge. During this fusion, we implement a weight-guided selective mechanism to prioritize significant weights. Moreover, we develop a confidence-based pseudo-labeling approach for the current non-entity type, which predicts entity types using the old model to handle the semantic shift of the non-entity type, a challenge specific to CNER that has largely been ignored by previous methods. Extensive experiments across ten CNER settings on three benchmark datasets demonstrate that our SPT method surpasses previous CNER approaches, highlighting its effectiveness in achieving a suitable stability-plasticity trade-off.
The predominant metric for evaluating speech recognizers, the Word Error Rate (WER) has been extended in different ways to handle transcripts produced by long-form multi-talker speech recognizers. These systems process long transcripts containing multiple speakers and complex speaking patterns so that the classical WER cannot be applied. There are speaker-attributed approaches that count speaker confusion errors, such as the concatenated minimum-permutation WER cpWER and the time-constrained cpWER (tcpWER), and speaker-agnostic approaches, which aim to ignore speaker confusion errors, such as the Optimal Reference Combination WER (ORC-WER) and the MIMO-WER. These WERs evaluate different aspects and error types (e.g., temporal misalignment). A detailed comparison has not been made. We therefore present a unified description of the existing WERs and highlight when to use which metric. To further analyze how many errors are caused by speaker confusion, we propose the Diarization-invariant cpWER (DI-cpWER). It ignores speaker attribution errors and its difference to cpWER reflects the impact of speaker confusions on the WER. Since error types cannot reliably be classified automatically, we discuss ways to visualize sequence alignments between the reference and hypothesis transcripts to facilitate the spotting of errors by a human judge. Since some WER definitions have high computational complexity, we introduce a greedy algorithm to approximate the ORC-WER and DI-cpWER with high precision ($<0.1\%$ deviation in our experiments) and polynomial complexity instead of exponential. To improve the plausibility of the metrics, we also incorporate the time constraint from the tcpWER into ORC-WER and MIMO-WER, also significantly reducing the computational complexity.
Manual sound design with a synthesizer is inherently iterative: an artist compares the synthesized output to a mental target, adjusts parameters, and repeats until satisfied. Iterative sound-matching automates this workflow by continually programming a synthesizer under the guidance of a loss function (or similarity measure) towards a target sound. Prior comparisons of loss functions have typically favored one metric over another, but only within narrow settings: limited synthesis methods, few loss types, often without blind listening tests. This leaves open the question of whether a universally optimal loss exists, or the choice of loss remains a creative decision conditioned on the synthesis method and the sound designer's preference. We propose differentiable iterative sound-matching as the natural extension of the available literature, since it combines the manual approach to sound design with modern advances in machine learning. To analyze the variability of loss function performance across synthesizers, we implemented a mix of four novel and established differentiable loss functions, and paired them with differentiable subtractive, additive, and AM synthesizers. For each of the sixteen synthesizer-loss combinations, we ran 300 randomized sound-matching trials. Performance was measured using parameter differences, spectrogram-distance metrics, and manually assigned listening scores. We observed a moderate level of consistency among the three performance measures. Our post hoc analysis shows that the loss function performance is highly dependent on the synthesizer. These findings underscore the value of expanding the scope of sound-matching experiments, and developing new similarity metrics tailored to specific synthesis techniques, rather than pursuing one-size-fits-all solutions.
The higher audio quality of steganography is directly correlated with the increased resistance to steganalysis tools. The advancement of generative AI technologies, particularly those that decouple style features from content, has shown promising developments by facilitating the creation of superior media content. This paper introduces the concept of audio decoupling and presents HIFI-Stego, an embedding audio steganography technique that aims to improve security while maintaining elevated stego audio quality. HIFI-Stego comprises a generator based on the encoder-decoder architecture and a secret message extractor. The encoder of the generator decouples the original audio, yielding the content vector, while the vocoder WORLD is employed to extract the style vector to preserve high-quality audio related features such as timbre and tone. Subsequently, the decoder embeds the secret vector into the decoupled content vector and then couples it with the style vector to generate high-fidelity stego audio. As embedding is not done in traditional time domain or frequency domain, existing analysis tools targeted at traditional steganographies fail to effectively detect the presence of the hidden message. The secret message extractor reuses the generator encoder and augments it with a single-layer convolutional neural network, resulting in a simplified structure suitable for lightweight deployment. Experimental results demonstrate that HIFI-Stego outperforms traditional generative and embedding steganographies in terms of audio quality, steganographic capacity, anti-analysis ability, and concealment.
In this work, we propose CleanMel, a single-channel Mel-spectrogram denoising and dereverberation network for improving both speech quality and automatic speech recognition (ASR) performance. The proposed network takes as input the noisy and reverberant microphone recording and predicts the corresponding clean Mel-spectrogram. The enhanced Mel-spectrogram can be either transformed to the speech waveform with a neural vocoder or directly used for ASR. The proposed network is composed of interleaved cross-band and narrow-band processing in the Mel-frequency domain, for learning the full-band spectral pattern and the narrow-band properties of signals, respectively. Compared to linear-frequency domain or time-domain speech enhancement, the key advantage of Mel-spectrogram enhancement is that Mel-frequency presents speech in a more compact way and thus is easier to learn, which will benefit both speech quality and ASR. Experimental results on five English and one Chinese datasets demonstrate a significant improvement in both speech quality and ASR performance achieved by the proposed model.
Recent years have witnessed the success of foundation models pre-trained with self-supervised learning (SSL) in various music informatics understanding tasks, including music tagging, instrument classification, key detection, and more. In this paper, we propose a self-supervised music representation learning model for music understanding. Distinguished from previous studies adopting random projection or existing neural codec, the proposed model, named MuQ, is trained to predict tokens generated by Mel Residual Vector Quantization (Mel-RVQ). Our Mel-RVQ utilizes residual linear projection structure for Mel spectrum quantization to enhance the stability and efficiency of target extraction and lead to better performance. Experiments in a large variety of downstream tasks demonstrate that MuQ outperforms previous self-supervised music representation models with only 0.9K hours of open-source pre-training data. Scaling up the data to over 160K hours and adopting iterative training consistently improve the model performance. To further validate the strength of our model, we present MuQ-MuLan, a joint music-text embedding model based on contrastive learning, which achieves state-of-the-art performance in the zero-shot music tagging task on the MagnaTagATune dataset. Code and checkpoints are open source in https://github.com/tencent-ailab/MuQ.
With the emergence of audio-language models, constructing large-scale paired audio-language datasets has become essential yet challenging for model development, primarily due to the time-intensive and labour-heavy demands involved. While large language models (LLMs) have improved the efficiency of synthetic audio caption generation, current approaches struggle to effectively extract and incorporate detailed audio information. In this paper, we propose an automated pipeline that integrates audio-language models for fine-grained content extraction, LLMs for synthetic caption generation, and a contrastive language-audio pretraining (CLAP) model-based refinement process to improve the quality of captions. Specifically, we employ prompt chaining techniques in the content extraction stage to obtain accurate and fine-grained audio information, while we use the refinement process to mitigate potential hallucinations in the generated captions. Leveraging the AudioSet dataset and the proposed approach, we create AudioSetCaps, a dataset comprising 1.9 million audio-caption pairs, the largest audio-caption dataset at the time of writing. The models trained with AudioSetCaps achieve state-of-the-art performance on audio-text retrieval with R@1 scores of 46.3% for text-to-audio and 59.7% for audio-to-text retrieval and automated audio captioning with the CIDEr score of 84.8. As our approach has shown promising results with AudioSetCaps, we create another dataset containing 4.1 million synthetic audio-language pairs based on the Youtube-8 M and VGGSound datasets.
Neural vocoders often struggle with aliasing in latent feature spaces, caused by time-domain nonlinear operations and resampling layers. Aliasing folds high-frequency components into the low-frequency range, making aliased and original frequency components indistinguishable and introducing two practical issues. First, aliasing complicates the waveform generation process, as the subsequent layers must address these aliasing effects, increasing the computational complexity. Second, it limits extrapolation performance, particularly in handling high fundamental frequencies, which degrades the perceptual quality of generated speech waveforms. This paper demonstrates that 1) time-domain nonlinear operations inevitably introduce aliasing but provide a strong inductive bias for harmonic generation, and 2) time-frequency-domain processing can achieve aliasing free waveform synthesis but lacks the inductive bias for effective harmonic generation. Building on this insight, we propose Wave hax, an aliasing-free neural WAVEform generator that integrates 2D convolution and a HArmonic prior for reliable Complex Spectrogram (Wavehax) estimation. Experimental results show that Wavehax achieves speech quality comparable to existing high-fidelity neural vocoders and exhibits exceptional robustness in scenarios requiring high fundamental frequency extrapola tion, where aliasing effects become typically severe. Moreover, Wavehax requires less than 5% of the multiply-accumulate operations and model parameters compared to HiFi-GAN V1, while achieving over four times faster CPU inference speed.
Sound event localization and detection (SELD) has seen substantial advancements through learning-based methods. These systems, typically trained from scratch on specific datasets, have shown considerable generalization capabilities. Recently, deep neural networks trained on large-scale datasets have achieved remarkable success in the sound event classification (SEC) field, prompting an open question of whether these advances can be extended to the development of SELD foundation models. In this paper, leveraging the power of pre-trained SEC models, we propose pre-trained SELD networks (PSELDNets) on a large-scale synthetic dataset. The synthetic dataset, generated by convolving sound events with simulated spatial room impulse responses (SRIRs), contains 1,167 hours of audio clips with an ontology of 170 sound classes. These PSELDNets are applied to various SELD scenarios. When we adapt PSELDNets to specific scenarios, particularly in cases of low-resource data, we introduce a data-efficient fine-tuning method, AdapterBit. PSELDNets are evaluated on synthetic-test-set using collected SRIRs from the TAU Spatial Room Impulse Response Database (TAU-SRIR DB) and achieve satisfactory performance. We also carried out experiments to validate the transferability of PSELDNets to three publicly available datasets and our own real-world recordings. The results demonstrate that PSELDNets surpass state-of-the-art systems across all publicly available datasets. Given the need for direction-of-arrival estimation, SELD generally relies on sufficient multi-channel audio clips. However, incorporating the AdapterBit, PSELDNets show more efficient adaptability to various scenarios using minimal multi-channel or even just monophonic audio clips, outperforming traditional fine-tuning approaches.
Error correction (EC) models play a crucial role in refining Automatic Speech Recognition (ASR) transcriptions, enhancing the readability and quality of transcriptions. Without requiring access to the underlying code or model weights, EC can improve performance and provide domain adaptation for black-box ASR systems. This work investigates the use of large language models (LLMs) for error correction across diverse scenarios. 1-best ASR hypotheses are commonly used as the input to EC models. We propose building high-performance EC models using ASR N-best lists which should provide more contextual information for the correction process. Additionally, the generation process of a standard EC model is unrestricted in the sense that any output sequence can be generated. For some scenarios, such as unseen domains, this flexibility may impact performance. To address this, we introduce a constrained decoding approach based on the N-best list or an ASR lattice. Finally, most EC models are trained for a specific ASR system requiring retraining whenever the underlying ASR system is changed. This paper explores the ability of EC models to operate on the output of different ASR systems. This concept is further extended to zero-shot error correction using LLMs, such as ChatGPT. Experiments on three standard datasets demonstrate the efficacy of our proposed methods for both Transducer and attention-based encoder-decoder ASR systems. In addition, the proposed method can serve as an effective method for model ensembling.
This paper presents an unsupervised method for single-channel blind dereverberation and room impulse response (RIR) estimation, called BUDDy. The algorithm is rooted in Bayesian posterior sampling: it combines a likelihood model enforcing fidelity to the reverberant measurement, and an anechoic speech prior implemented by an unconditional diffusion model. We design a parametric filter representing the RIR, with exponential decay for each frequency subband. Room acoustics estimation and speech dereverberation are jointly carried out, as the filter parameters are iteratively estimated and the speech utterance refined along the reverse diffusion trajectory. In a blind scenario where the RIR is unknown, BUDDy successfully performs speech dereverberation in various acoustic scenarios, significantly outperforming other blind unsupervised baselines. Unlike supervised methods, which often struggle to generalize, BUDDy seamlessly adapts to different acoustic conditions. This paper extends our previous work by offering new experimental results and insights into the algorithm's versatility. We demonstrate the robustness of our proposed method to new acoustic and speaker conditions, as well as its adaptability to highresolution singing voice dereverberation, using both instrumental metrics and subjective listening evaluation. We study BUDDy's performance for RIR estimation and observe it surpasses a state-of-the-art supervised DNN-based estimator on mismatched acoustic conditions. Finally, we investigate the sensitivity of informed dereverberation methods to RIR estimation errors, thereby motivating the joint acoustic estimation and dereverberation design. Audio examples and code can be found online.