• 学术搜索
  • 科研智能体
    • Research Labs
    • AI 阅读
    • AI 文库
    • 深度研究
    • 学者亮点
  • 学术资源
    • AI2000
    • 期刊/会议
    • 学者库
    • 学术API
    • 溯源树
    • 数据集
  • 知识沉淀
    • 学术空间
订阅小程序
旧版功能
aminer vip
开通会员低至0.73元/天
一次搞定AI科研
立即登录
  • English
  • 联系方式
    I

    International Audio Laboratories Erlangen

    EST. 2008
    141论文总数
    2,512引用总数

    论文量&引用量时间轴

    机构学者

    排序
    Emanuël Habets
    Emanuël Habets
    International Audio Laboratories Erlangen, Friedrich-Alexander Universität Erlangen-Nürnberg;Fraunhofer IIS
    论文:55引用:0H-index:0
    Meinard Müller
    Meinard Müller
    International Audio Laboratories Erlangen, Friedrich-Alexander-University of Erlangen-Nürnberg
    论文:42引用:0H-index:0
    Christof Weiß
    Christof Weiß
    Fraunhofer Society
    论文:10引用:0H-index:0
    Christian Dittmar
    Christian Dittmar
    Fraunhofer IIS
    论文:8引用:0H-index:0
    Stefan Balke
    Stefan Balke
    International Audio Laboratories Erlangen
    论文:8引用:0H-index:0
    Oliver Thiergart
    Oliver Thiergart
    Fraunhofer IIS
    论文:7引用:0H-index:0
    Bernd Edler
    Bernd Edler
    Multimedia Commun. Res. Lab, AT&TBell Labs
    论文:7引用:0H-index:0
    Srikanth Raj Chetupally
    Srikanth Raj Chetupally
    Department of Electronics and Communication Engineering, Indian Institute of Science
    论文:7引用:0H-index:0
    Frank Zalkow
    Frank Zalkow
    Int Audio Labs Erlangen, D-91058 Erlangen, Germany
    论文:6引用:0H-index:0

    论文(141)

    年份
    起
    –
    止
    排序
    1Leveraging Discriminative Latent Representations for Conditioning GAN-Based Speech Enhancement
    Shrishti Saha Shetu,Emanuel A. P. Habets, Andreas Brendel

    Generative speech enhancement methods based on generative adversarial networks (GANs) have demonstrated promising performance across various speech enhancement tasks. However, their performance in very low signal-to-noise ratio (SNR) scenarios remains under-explored and limited, as these conditions pose significant challenges to both discriminative and generative state-of-the-art methods. To address this, we propose DisCoGAN, a GAN-based speech enhancement method that leverages latent features extracted from discriminative speech enhancement models as generic conditioning information. By incorporating the proposed discriminative conditioning method, DisCoGAN improves speech quality and intelligibility, particularly in low-SNR scenarios, while maintaining competitive or superior performance in high-SNR conditions and real-world recordings. We also conduct a comprehensive evaluation of conventional GAN-based architectures, including end-to-end GANs, GAN-first, and post-filtering GANs, as well as discriminative models under low-SNR conditions, and show that DisCoGAN consistently outperforms existing methods. Finally, we present ablation studies that highlight the performance gains from discriminative conditioning and demonstrate how DisCoGAN leverages both local and global temporal context, providing insight into the key factors underlying these gains.

    2026IEEE TRANSACTIONS ON AUDIO SPEECH AND LANGUAGE PROCESSING(2026)引用:4
    引用
    AI阅读
    加入学术空间
    2Acoustic Teleportation Via Disentangled Neural Audio Codec Representations
    Philipp Grundhuber,Mhd Modar Halimeh,Emanuël A. P. Habets

    This paper presents an approach for acoustic teleportation by disentangling speech content from acoustic environment characteristics in neural audio codec representations. Acoustic teleportation transfers room characteristics between speech recordings while preserving content and speaker identity, enabling applications in telecommunications and virtual acoustic environments. Extending prior work by adopting the EnCodec architecture, we achieve substantial objective quality improvements with non-intrusive ScoreQ scores of 3.03, compared to 2.44 for previous methods. Our training strategy incorporates five tasks: clean reconstruction, reverberated reconstruction, dereverberation, and two variants of acoustic teleportation. We analyze the trade-off between temporal downsampling of acoustic embeddings and reconstruction quality, demonstrating that even a downsampling factor of two yields a statistically significant degradation in reconstruction quality. The learned acoustic embeddings exhibit a strong correlation with reverberation time. t-SNE analysis reveals that acoustic embeddings cluster predominantly by room while speech embeddings cluster by speaker, confirming substantial disentanglement sufficient for effective acoustic teleportation.

    2026ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)(2026)引用:1
    引用
    AI阅读
    加入学术空间
    3RWC Revisited: Towards a Community-Driven MIR Corpus
    Stefan Balke,Johannes Zeitler,Vlora Arifi-Mueller,Brian Mcfee,Tomoyasu Nakano, Masataka Goto,Meinard Mueller

    The Real World Computing (RWC) Music Database has been a cornerstone of Music Information Retrieval (MIR) research for over two decades, offering high-quality record- ings across multiple genres, including popular, classical , jazz music. Beyond its extensive audio collection, the dataset is enriched by aligned Musical Instrument Digital Interface (MIDI) encodings , complementary annotations, including beat, structure, and chord labels, making it a valuable resource for music structure analysis, beat track- ing, chord recognition, automatic transcription, and music synchronization. Originally, the RWC audio material was distributed on physical media and made available for purchase at a nominal price. A significant development, announced and initiated with this paper, is the release of the RWC dataset under a Creative Commons license, mak- ing it freely accessible for research purposes. This transition significantly enhances the dataset's usability and supports broader adoption within the MIR research community. We outline the steps taken to enable this release and share a vision for transforming RWC into a community-driven resource that promotes open research and collabora- tion. With the audio recordings now hosted on Zenodo, we also discuss strategies for dataset maintenance, annotation expansion, and reproducibility through collabora- tive platforms such as GitHub. This shift promotes transparency and inclusivity, helping to ensure the dataset's continued relevance for cutting-edge MIR research. We fur- ther revisit the historical significance of the RWC dataset, incorporating insights from an interview with its original creator, Masataka Goto, and provide an overview of its current applications and future potential. In summary, by embracing an open and community-supported approach, we aim not only to renew the dataset's impact and preserve its legacy within the MIR community but also to shed light on broader best practices for open, collaborative, and sustainable research infrastructures.

    2026TRANSACTIONS OF THE INTERNATIONAL SOCIETY FOR MUSIC INFORMATION RETRIEVAL(2026)引用:1
    引用
    AI阅读
    加入学术空间
    4A Comparison of Generative and Discriminative Methods for Speech Enhancement: Robustness, Complexity, and Hallucination
    Shrishti Saha Shetu, Emanuël A. P. Habets,Andreas Brendel

    In this study, we conduct a comprehensive comparative analysis of generative and discriminative deep learning-based speech enhancement methods, specifically in noise reduction tasks. Our investigation focuses on evaluating their effectiveness under high and low signal-to-noise ratio conditions, considering both matched and mismatched training scenarios. We further investigate the impact of training data volume, model convergence speed, and interpret the performance differences in terms of objective results for the considered training paradigms. Additionally, we compare the complexity-performance trade-off and the practical viability of these approaches. To further strengthen the evaluation, we study the hallucination characteristics of generative approaches in terms of word error rate and phoneme similarity. The insights derived from this study provide empirical evidence to assist researchers and practitioners in understanding whether the perceptual gains of different approaches justify their computational cost in practical applications.

    2026引用:1
    引用
    AI阅读
    加入学术空间
    5QExE: a Quality and Experience Evaluation Tool for Audiovisual VR Perception, Behavior, and Cognition Research
    Thomas Robotham, Olli S. Rummukainen, Daniela Rebmann,Alexander Raake, Emanuel A. P. Habets

    The interactive nature of virtual reality (VR) challenges many assumptions of conventional audiovisual quality evaluation approaches, necessitating tools and methods that account for user agency, temporal coupling, and context. We present the Quality and Experience Evaluation (QExE) Tool for interactive VR. Several quality evaluation methods, additional questionnaires, behavioral and interactivity data collection, and the ability to load multiple suitable audio rendering plug-ins are included. The tool streamlines the evaluation process by automatically creating test items that include audiovisual media such as object-based or multichannel audio and three-dimensional (3D) spatial scenes or 360° videos. Utilizing a 3D game engine in tandem with control software, the method, questionnaire, and interactivity data can be saved to subject-specific sub-directories. Prior research utilizing the QExE tool is described, along with a novel case study investigating the effects of a VR training scheme on cognitive load, audio plausibility, and interaction behavior. This study demonstrates that the QExE tool can be effectively employed to collect perceptual, cognitive, and behavioral data for subjective evaluations in interactive VR environments.

    2026FRONTIERS IN VIRTUAL REALITY(2026)
    引用
    AI阅读
    加入学术空间
    立即登录,查看全部 141 篇论文

    合作机构(46)

    Fraunhofer Institute for Integrated Circuits,Fraunhofer Society合作论文 25
    巴伊兰大学合作论文 5
    Media Design School合作论文 5
    Fraunhofer Institute for Digital Media Technology,Fraunhofer Society合作论文 4
    魁北克大学合作论文 3
    米兰理工大学合作论文 2
    波茨坦大学合作论文 2
    萨尔大学合作论文 2
    特拉维夫大学合作论文 2
    庞培法布拉大学合作论文 2

    机构统计