• 学术搜索
  • 科研智能体
    • Research Labs
    • AI 阅读
    • AI 文库
    • 深度研究
    • 学者亮点
  • 学术资源
    • AI2000
    • 期刊/会议
    • 学者库
    • 学术API
    • 溯源树
    • 数据集
  • 知识沉淀
    • 学术空间
订阅小程序
旧版功能
aminer vip
开通会员低至0.73元/天
一次搞定AI科研
立即登录
  • English
  • 联系方式
    博

    博士公司

    Bose Corporation
    企业
    317论文总数
    4,344引用总数

    论文量&引用量时间轴

    机构学者

    排序
    Luc De Keersmaeker
    Luc De Keersmaeker
    Research Institute for Nature and Forest
    论文:16引用:0H-index:0
    Kris Vandekerkhove
    Kris Vandekerkhove
    Research Institute for Nature and Forest (INBO), Gaverstraat 4, 9500 Geraardsbergen, Belgium
    论文:16引用:0H-index:0
    Luc Denys
    Luc Denys
    Research Institute for Nature and Forest (INBO), Belgium
    论文:10引用:0H-index:0
    Abhijit Mookerjee
    Abhijit Mookerjee
    S.N. Bose National Centre for Basic Sciences
    论文:9引用:0H-index:0
    Kenneth D. Jacob
    Kenneth D. Jacob
    Fixed Nitr Res Lab, Washington, DC USA
    论文:6引用:0H-index:0
    Marko Stamenovic
    Marko Stamenovic
    Bose Corporation
    论文:6引用:0H-index:0
    Anil Kumar Bajpai
    Anil Kumar Bajpai
    Department of Chemistry, Government Model Science College Jabalpur
    论文:5引用:0H-index:0
    Indra Dasgupta
    Indra Dasgupta
    School of Physical Sciences, Indian Association for Cultivations and Sciences
    论文:5引用:0H-index:0
    Helmut Tributsch
    Helmut Tributsch
    Bio-Mimetics Program, Carinthian University of Applied Sciences
    论文:5引用:0H-index:0

    论文(317)

    年份
    起
    –
    止
    排序
    1FSD50K-Solo: Automated Curation of Single-Source Sound Events
    Ningyuan Yang, Sile Yin, Li-Chia Yang, Bryce Irvin, Xiao Quan,Marko Stamenovic, Shuo Zhang

    High-quality training datasets are essential for the performance of neural networks. However, the audio domain still lacks a large-scale, strongly-labeled, and single-source sound event dataset. The FSD50K dataset, despite being relatively large and open, contains a considerable fraction of multi-source samples where background interference or overlapping events could limit the usefulness of the data. To address this challenge, we introduce a data curation framework designed for large-scale open audio corpora. Our approach leverages a generative diffusion model to synthesize clean single-class events to construct controlled noisy mixtures for supervision. We subsequently employ a pre-trained audio encoder coupled with a discriminative classifier to automatically identify and filter out multi-source samples. Experiments show that our framework achieves strong performance on a human expert-curated test set. Finally, we release FSD50K-Solo, a model-curated subset of FSD50K containing single-source audio samples identified by our method. Beyond FSD50K, our method establishes a scalable paradigm for curating open source audio corpora.

    2026引用:1
    引用
    AI阅读
    加入学术空间
    2StoicLLM: Preference Optimization for Philosophical Alignment in Small Language Models
    Ishmam Khan, Sindhuja Thogarrati, Shuo Zhang

    While large language models excel at factual adaptation, their ability to internalize nuanced philosophical frameworks under severe data constraints remains underexplored. We investigate this by specializing small LLMs on micro-datasets of foundational Stoic texts using preference optimization (ORPO, AlphaPO). Evaluated via a multi-model critic bank, our results show that just 300 high-fidelity examples can induce strong alignment with inward-facing Stoic virtues, closely approaching few-shot prompting while freeing the context window. Critically, however, all models, including few-shot baselines, exhibit a persistent failure on Stoicism's outward-facing cosmopolitan duties, pointing to a representational limitation of small models that micro-dataset adaptation alone cannot overcome.

    2026Proceedings of the 6th International Conference on Natural Language Processing for the Digital Human...(2026)
    引用
    AI阅读
    加入学术空间
    3No Word Left Behind: Mitigating Prefix Bias in Open-Vocabulary Keyword Spotting
    Yi Liu, Chuan-Che Huang, Xiao Quan

    Open-vocabulary keyword spotting (OV-KWS) enables personalized device control via arbitrary voice commands. Recently, researchers have explored using audio-text joint embeddings, allowing users to enroll phrases with text, and proposed techniques to disambiguate similar utterances. We find that existing OV-KWS solutions often overly bias the beginning phonemes of an enrollment, causing false triggers when negative enrollment-query-pairs share a prefix (“turn the volume up” vs. “turn the volume down”). We trace this to two factors: training data bias and position-biased cross-modal scoring. To address these limitations, we introduce the Partial Overlap Benchmark (POB) with two datasets, POB-Spark and POB-LibriPhrase (POB-LP), containing mismatched audio-text pairs with shared prefixes, and propose Equal-weighting Position Scoring (EPS), a lightweight decision layer. Using EPS alone reduces EER on POB-Spark from 64.4% to 29.3% and improves POB-LP accuracy from 87.6% to 96.8%, while maintaining performance on LibriPhrase and Google Speech Commands (GSC). With POB data added in training, our work achieves the best POB benchmark results while incurring the least amount of degradation on prior metrics among baselines. This degradation is most pronounced in GSC, which contains only one-word commands. We surface mitigating this trade-off as future work.

    2026ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)(2026)
    引用
    AI阅读
    加入学术空间
    4MRMAD: A Multi-Round Multi-Audio Benchmark for Evaluating Acoustic Degradation Perception in Large Audio-Language Models
    Yize Li, Ningyuan Yang, Sile Yin, Sindhuja Thogarrati,Sung-En Chang, Andrew C. Singer,Xue Lin, Chuan-Che Huang, Shuo Zhang

    Large audio-language models (LALMs) have shown promising progress in understanding speech, music, and general sound events, yet their ability to reason about how audio signals are degraded remains underexplored. Existing benchmarks primarily evaluate semantic understanding, event recognition, or high-level audio reasoning, leaving a basic question unanswered: Do LALMs understand the differences in audio quality? We introduce MRMAD, a Multi-Round Multi-Audio Degradation benchmark for evaluating audio degradation perception and understanding in LALMs. MRMAD spans speech, music, and sound, and frames evaluation as multi-turn dialogues over multiple audio inputs, requiring models to identify degradation types, compare severity, and perceive corruption changes across turns. Unlike current single-turn audio-language benchmarks, MRMAD evaluates whether LALMs can maintain consistent degradation hypotheses with new evidence and explain low-level acoustic phenomena in natural language. Through a systematic evaluation of 18 representative LALMs from non-thinking to reasoning and Omni models, we find that current models often recognize coarse content while failing to diagnose, compare, or reason about degradations reliably. MRMAD reveals an important yet overlooked aspect of audio-language understanding and provides a diagnostic foundation for building future LALMs that are robust to real-world acoustic conditions.

    2026
    引用
    AI阅读
    加入学术空间
    5Multimodal Artificial Intelligence Based Sensing for Hearables
    Sile Yin,Marko Stamenovic, Shuo Zhang

    With the rapid advancement of deep neural networks that process information from multiple modalities—such as video, audio, and text—wearables can leverage multimodal AI models to solve challenging audio tasks, with potential applications in voice communications and selective hearing with multiple electroacoustic sensors and beyond. In this talk, we discuss recent work from Bose Research which incorporate video, bone-conduction vibration sensors, and ultrasonic sensing to provide improved capabilities on tasks such as audio source separation and keyword spotting. First, we present a system that leverages pretrained visual embeddings to perform audio-visual speech enhancement (AVSE) from the video and audio stream in real time. Second, we highlight work on target speaker extraction (TSE), comparing two strategies for targeting the wearer’s voice: personalized speech enhancement (PSE), which uses enrollment utterances to represent the target, and auxiliary-sensor speech enhancement (AS-SE), which leverages in-ear microphones carrying target-specific cues. Example applications include the recently released Bose SpeechClarity technology for voice communications on earbuds. Finally, we present other work on Whisper Keyword Spotting (WKWS) with ultrasonic sensing, using existing electroacoustics actuator and sensors on a Bose headphone.

    2026
    引用
    AI阅读
    加入学术空间
    立即登录,查看全部 317 篇论文

    合作机构(100)

    捷迈邦美合作论文 8
    S. N. Bose National Centre for Basic Sciences合作论文 8
    明尼苏达大学合作论文 6
    加尔各答大学合作论文 6
    麻省理工学院合作论文 6
    Eurogentec (Belgium)合作论文 5
    密歇根大学合作论文 3
    塔夫茨大学合作论文 3
    Calcutta National Medical College合作论文 3
    加利福尼亚大学圣地亚哥分校合作论文 3

    机构统计