Generative speech enhancement methods based on generative adversarial networks (GANs) have demonstrated promising performance across various speech enhancement tasks. However, their performance in very low signal-to-noise ratio (SNR) scenarios remains under-explored and limited, as these conditions pose significant challenges to both discriminative and generative state-of-the-art methods. To address this, we propose DisCoGAN, a GAN-based speech enhancement method that leverages latent features extracted from discriminative speech enhancement models as generic conditioning information. By incorporating the proposed discriminative conditioning method, DisCoGAN improves speech quality and intelligibility, particularly in low-SNR scenarios, while maintaining competitive or superior performance in high-SNR conditions and real-world recordings. We also conduct a comprehensive evaluation of conventional GAN-based architectures, including end-to-end GANs, GAN-first, and post-filtering GANs, as well as discriminative models under low-SNR conditions, and show that DisCoGAN consistently outperforms existing methods. Finally, we present ablation studies that highlight the performance gains from discriminative conditioning and demonstrate how DisCoGAN leverages both local and global temporal context, providing insight into the key factors underlying these gains.
This paper presents an approach for acoustic teleportation by disentangling speech content from acoustic environment characteristics in neural audio codec representations. Acoustic teleportation transfers room characteristics between speech recordings while preserving content and speaker identity, enabling applications in telecommunications and virtual acoustic environments. Extending prior work by adopting the EnCodec architecture, we achieve substantial objective quality improvements with non-intrusive ScoreQ scores of 3.03, compared to 2.44 for previous methods. Our training strategy incorporates five tasks: clean reconstruction, reverberated reconstruction, dereverberation, and two variants of acoustic teleportation. We analyze the trade-off between temporal downsampling of acoustic embeddings and reconstruction quality, demonstrating that even a downsampling factor of two yields a statistically significant degradation in reconstruction quality. The learned acoustic embeddings exhibit a strong correlation with reverberation time. t-SNE analysis reveals that acoustic embeddings cluster predominantly by room while speech embeddings cluster by speaker, confirming substantial disentanglement sufficient for effective acoustic teleportation.
The Real World Computing (RWC) Music Database has been a cornerstone of Music Information Retrieval (MIR) research for over two decades, offering high-quality record- ings across multiple genres, including popular, classical , jazz music. Beyond its extensive audio collection, the dataset is enriched by aligned Musical Instrument Digital Interface (MIDI) encodings , complementary annotations, including beat, structure, and chord labels, making it a valuable resource for music structure analysis, beat track- ing, chord recognition, automatic transcription, and music synchronization. Originally, the RWC audio material was distributed on physical media and made available for purchase at a nominal price. A significant development, announced and initiated with this paper, is the release of the RWC dataset under a Creative Commons license, mak- ing it freely accessible for research purposes. This transition significantly enhances the dataset's usability and supports broader adoption within the MIR research community. We outline the steps taken to enable this release and share a vision for transforming RWC into a community-driven resource that promotes open research and collabora- tion. With the audio recordings now hosted on Zenodo, we also discuss strategies for dataset maintenance, annotation expansion, and reproducibility through collabora- tive platforms such as GitHub. This shift promotes transparency and inclusivity, helping to ensure the dataset's continued relevance for cutting-edge MIR research. We fur- ther revisit the historical significance of the RWC dataset, incorporating insights from an interview with its original creator, Masataka Goto, and provide an overview of its current applications and future potential. In summary, by embracing an open and community-supported approach, we aim not only to renew the dataset's impact and preserve its legacy within the MIR community but also to shed light on broader best practices for open, collaborative, and sustainable research infrastructures.
In this study, we conduct a comprehensive comparative analysis of generative and discriminative deep learning-based speech enhancement methods, specifically in noise reduction tasks. Our investigation focuses on evaluating their effectiveness under high and low signal-to-noise ratio conditions, considering both matched and mismatched training scenarios. We further investigate the impact of training data volume, model convergence speed, and interpret the performance differences in terms of objective results for the considered training paradigms. Additionally, we compare the complexity-performance trade-off and the practical viability of these approaches. To further strengthen the evaluation, we study the hallucination characteristics of generative approaches in terms of word error rate and phoneme similarity. The insights derived from this study provide empirical evidence to assist researchers and practitioners in understanding whether the perceptual gains of different approaches justify their computational cost in practical applications.
The interactive nature of virtual reality (VR) challenges many assumptions of conventional audiovisual quality evaluation approaches, necessitating tools and methods that account for user agency, temporal coupling, and context. We present the Quality and Experience Evaluation (QExE) Tool for interactive VR. Several quality evaluation methods, additional questionnaires, behavioral and interactivity data collection, and the ability to load multiple suitable audio rendering plug-ins are included. The tool streamlines the evaluation process by automatically creating test items that include audiovisual media such as object-based or multichannel audio and three-dimensional (3D) spatial scenes or 360° videos. Utilizing a 3D game engine in tandem with control software, the method, questionnaire, and interactivity data can be saved to subject-specific sub-directories. Prior research utilizing the QExE tool is described, along with a novel case study investigating the effects of a VR training scheme on cognitive load, audio plausibility, and interaction behavior. This study demonstrates that the QExE tool can be effectively employed to collect perceptual, cognitive, and behavioral data for subjective evaluations in interactive VR environments.