Deep learning-based multi-channel speech enhancement models are often tightly coupled to the microphone array geometries encountered during training, leading to significant performance degradation on unseen configurations. To address this limitation, we propose a new architecture that achieves geometry-agnostic speech enhancement through an eigenbeam-domain input representation. In the front-end, microphone signals from a compact arbitrary planar array are projected onto a canonical, geometry-invariant feature space via a spatial filter bank designed to approximate the directivity patterns of circular harmonics up to a given order. The core of the proposed learning method is the novel Multi-order Spatial Encoder Network (MSEN), a dedicated module that exploits the structured nature of the eigenbeam features by processing zeroth, first, and second-order components through parallel encoders before late fusion. Experiments conducted on fully randomized planar array configurations demonstrate that our method outperforms state-of-the-art baselines across key speech enhancement metrics, including perceptual quality (PESQ, CSIG, CBAK, COVL) and intelligibility (STOI), while substantially reducing computational complexity compared to existing methods.
Accurate spatial acoustic characterization is crucial for immersive audio applications, such as virtual and augmented reality, which require low-latency, high-fidelity rendering of sound fields. Traditional approaches, including physics-based simulations and centralized deep learning models, face significant challenges in scalability, data efficiency, and robustness under sparse measurement conditions, and typically require retraining or fine-tuning whenever new measurements are acquired. This paper introduces DAME, a Distributed Acoustic Mixture-of-Experts framework for large-area spatial acoustic modeling. DAME decomposes the global sound field reconstruction task into localized sub-problems, each handled by a compact neural expert trained on region-specific data, while a lightweight gating network combines the experts' predictions based on spatial queries. Unlike state-of-the-art methods, the framework supports incremental expansion on the receiver side. Specifically, new experts can be added to cover previously unsampled regions, and only the gating network is updated, avoiding full retraining of the model and preserving previously learned behavior. Evaluated on both simulated and real-world higher-order Ambisonics room impulse responses, DAME consistently outperforms state-of-the-art parametric, kernel-based, and physics-informed baselines in terms of reconstruction accuracy, directional estimation, and energy decay preservation. Computational analysis demonstrates its suitability for low-latency operation, and a perceptual evaluation confirms superior audio quality and spatial fidelity. The framework therefore provides a scalable, data-efficient alternative for robust spatial audio reconstruction in challenging sparse-data regimes.
A persistent challenge in generative audio models is data replication, where the model unintentionally generates parts of its training data during inference. In this work, we address this issue in text-to-audio diffusion models by exploring the use of anti-memorization strategies. We adopt Anti-Memorization Guidance (AMG), a technique that modifies the sampling process of pre-trained diffusion models to discourage memorization. Our study explores three types of guidance within AMG, each designed to reduce replication while preserving generation quality. We use Stable Audio Open as our backbone, leveraging its fully open-source architecture and training dataset. Our comprehensive experimental analysis suggests that AMG significantly mitigates memorization in diffusion-based text-to-audio generation without compromising audio fidelity or semantic alignment.
Recently, differentiable Feedback Delay Networks (FDNs) have been successfully employed to model the impulse response of higher-order microphones thanks to data-driven optimization routines. In this work, we aim to shed light on the impact of FDN optimization on its capabilities to model the distribution of energy across time, frequency, and space. In particular, we investigate which parameters are crucial for obtaining coherent and accurate results, and which parameters can instead be frozen without significantly degrading the performance. This analysis paves the way for a more principled and efficient use of differentiable FDNs in spatial audio applications, providing guidelines for future research and practical implementations.
Room Impulse Responses estimation is a fundamental problem in spatial audio processing and speech enhancement. In this paper, we build upon our previously introduced diffusion-based inpainting framework for Room Impulse Response interpolation and demonstrate its applicability to enhancing the performance of practical multi-microphone array processing tasks. Furthermore, we validate the robustness of this method in interpolating real-world Room Impulse Responses.
BACKGROUND:Pediatric cardiac tele-auscultation is often limited by poor signal quality due to respiratory sounds, motion artifacts, and environmental noise. Distinguishing pathological from innocent murmurs remains challenging and operator-dependent, often leading to unnecessary referrals. Advanced signal processing and machine learning may improve the reliability of remote auscultation. METHODS:A total of 135 children underwent cardiac auscultation using a digital stethoscope and echocardiography; 90 patients with confirmed absence or presence of pathological murmur were included. Phonocardiograms were segmented and denoised using permutation-enhanced Non-negative Matrix Factorization. Time-frequency features were extracted and used to train Support Vector Machine classifiers for each auscultation site. RESULTS:Across five sites, specificity ranged from 70.4 to 100.0%, sensitivity from 25.0 to 75.0%, and accuracy from 71.0 to 95.7%. Specificity was ≥88.2% at all sites except the upper right sternal border. Sensitivity reached 75.0% at three sites but was lower at the apex. Combined results yielded specificity, sensitivity, and accuracy of 81.8, 66.7, and 77.4%, respectively. CONCLUSION:Improving signal quality is crucial for reliable automated murmur detection in children. The combination of advanced denoising and machine learning can enhance tele-auscultation, support primary care physicians, and reduce unnecessary referrals. IMPACT:Advanced denoising combined with data-driven classification improves the reliability of pediatric cardiac tele-auscultation in real-world noisy conditions. The study provides clinical evidence that signal quality enhancement is a critical prerequisite for accurate automated murmur detection in children. A multi-site machine-learning approach using digital stethoscope recordings is feasible in a pediatric population. This approach can support primary care physicians in clinical decision-making and help reduce unnecessary referrals to pediatric cardiology specialists.
Sound field reconstruction aims to estimate pressure fields in areas lacking direct measurements. Existing techniques often rely on strong assumptions or face challenges related to data availability or the explicit modeling of physical properties. To bridge these gaps, this study introduces a zero-shot, physics-informed dictionary learning approach to perform sound field reconstruction. Our method relies only on a few sparse measurements to learn a dictionary, without the need for additional training data. Moreover, by enforcing the Helmholtz equation during the optimization process, the proposed approach ensures that the reconstructed sound field is represented as a linear combination of a few physically meaningful atoms. Evaluations on real-world data show that our approach achieves comparable performance to state-of-the-art dictionary learning techniques, with the advantage of requiring only a few observations of the sound field and no training on a dataset.
Advanced remote applications such as Networked Music Performance (NMP) require solutions to guarantee immersive real-world-like interaction among users. Therefore, the adoption of spatial audio formats, such as Ambisonics, is fundamental to let the user experience an immersive acoustic scene. The accuracy of the sound scene reproduction increases with the order of the Ambisonics enconding, resulting in an improved immersivity at the cost of a greater number of audio channels, which in turn escalates both bandwidth requirements and susceptibility to network impairments (e.g., latency, jitter, and packet loss). These factors pose a significant challenge for interactive music sessions, which demand high spatial fidelity and low end-to-end delay. We propose a real-time adaptive higher-order Ambisonics strategy that continuously monitors network throughput and dynamically scales the Ambisonics order. When available bandwidth drops below a preset threshold, the order is lowered to prevent audio dropouts; it then reverts to higher orders once conditions recover, thus balancing immersion and reliability. A MUSHRA-based evaluation indicates that this adaptive approach is promising to guarantee user experience in bandwidth-limited NMP scenarios.
The perceptual evaluation of spatial audio algorithms is an important step in the development of immersive audio applications, as it ensures that synthesized sound fields meet quality standards in terms of listening experience, spatial perception and auditory realism. To support these evaluations, virtual reality can offer a powerful platform by providing immersive and interactive testing environments. In this paper, we present VR-PTOLEMAIC, a virtual reality evaluation system designed for assessing spatial audio algorithms. The system implements the MUSHRA (MUlti-Stimulus test with Hidden Reference and Anchor) evaluation methodology into a virtual environment. In particular, users can position themselves in each of the 25 simulated listening positions of a virtually recreated seminar room and evaluate simulated acoustic responses with respect to the actually recorded second-order ambisonic room impulse responses, all convolved with various source signals. We evaluated the usability of the proposed framework through an extensive testing campaign in which assessors were asked to compare the reconstruction capabilities of various sound field reconstruction algorithms. Results show that the VR platform effectively supports the assessment of spatial audio algorithms, with generally positive feedback on user experience and immersivity.
We propose a transfer learning framework for sound source reconstruction in Near-field Acoustic Holography (NAH), which adapts a well-trained data-driven model from one type of sound source to another using a physics-informed procedure. The framework comprises two stages: (1) supervised pre-training of a complex-valued convolutional neural network (CV-CNN) on a large dataset, and (2) purely physics-informed fine-tuning on a single data sample based on the Kirchhoff-Helmholtz integral. This method follows the principles of transfer learning by enabling generalization across different datasets through physics-informed adaptation. The effectiveness of the approach is validated by transferring a pre-trained model from a rectangular plate dataset to a violin top plate dataset, where it shows improved reconstruction accuracy compared to the pre-trained model and delivers performance comparable to that of Compressive-Equivalent Source Method (C-ESM). Furthermore, for successful modes, the fine-tuned model outperforms both the pre-trained model and C-ESM in accuracy.
The rapid evolution of the Metaverse demands integrated solutions that merge advanced telecommunication networks with immersive audio technologies. High-fidelity, low-latency audio transmission relies on robust network infrastructures such as 5G, emerging 6G, network slicing, and edge computing that deliver high throughput and ultra-reliable connectivity. This paper investigates how these advanced networking technologies can be synergistically combined with state-of-the-art acoustic rendering techniques, including spatial audio and binaural processing, to enable immersive audio applications in the Metaverse. By analyzing key application scenarios and mapping critical performance indicators for both network and acoustic quality, we highlight the pivotal role of telecommunication advancements in meeting the stringent requirements for real-time, immersive experiences. Open challenges and future research directions are discussed, emphasizing the convergence between next-generation networks and immersive audio to shape the future of interactive virtual environments.
The Deep Prior framework has emerged as a powerful generative tool which can be used for reconstructing sound fields in an environment from few sparse pressure measurements. It employs a neural network that is trained solely on a limited set of available data and acts as an implicit prior which guides the solution of the underlying optimization problem. However, a significant limitation of the Deep Prior approach is its inability to generalize to new acoustic configurations, such as changes in the position of a sound source. As a consequence, the network must be retrained from scratch for every new setup, which is both computationally intensive and time-consuming. To address this, we investigate transfer learning in Deep Prior via Low-Rank Adaptation (LoRA), which enables efficient fine-tuning of a pre-trained neural network by introducing a low-rank decomposition of trainable parameters, thus allowing the network to adapt to new measurement sets with minimal computational overhead. We embed LoRA into a MultiResUNet-based Deep Prior model and compare its adaptation performance against full fine-tuning of all parameters as well as classical retraining, particularly in scenarios where only a limited number of microphones are used. The results indicate that fine-tuning, whether done completely or via LoRA, is especially advantageous when the source location is the sole changing parameter, preserving high physical fidelity, and highlighting the value of transfer learning for acoustics applications.
Text-to-music models have revolutionized the creative landscape, offering new possibilities for music creation. Yet their integration into musicians workflows remains underexplored. This paper presents a case study on how TTM models impact music production, based on a user study of their effect on producers creative workflows. Participants produce tracks using a custom tool combining TTM and source separation models. Semi-structured interviews and thematic analysis reveal key challenges, opportunities, and ethical considerations. The findings offer insights into the transformative potential of TTMs in music production, as well as challenges in their real-world integration.
Text-to-audio models have recently emerged as a powerful technology for generating sound from textual descriptions. However, their high computational demands raise concerns about energy consumption and environmental impact. In this paper, we conduct an analysis of the energy usage of 7 state-of-the-art text-to-audio diffusion-based generative models, evaluating to what extent variations in generation parameters affect energy consumption at inference time. We also aim to identify an optimal balance between audio quality and energy consumption by considering Pareto-optimal solutions across all selected models. Our findings provide insights into the trade-offs between performance and environmental impact, contributing to the development of more efficient generative audio models.
Immersive Networked Music Performances (INMP) rely heavily on low-latency transmission and accurate preservation of spatial audio to guarantee satisfying experiences for performers and audiences alike. This study investigates the implications of streaming higher-order Ambisonics spatial audio over the internet, between two locations, and their impact on the musicians’ performances. We extensively evaluate subjective Quality of Experience (QoE) under varying network conditions. Specifically, the research addresses the following question: which features of the transmission channel mostly impact on the QoE? Results highlight the significant impact of latency on mutual engagement, whereas packet loss has a more limited effect. However, packet loss does produce noticeable effects on sound quality, which are not observed in the case of latency. The outcomes will contribute to defining optimal strategies for immersive spatial audio transmission in distributed musical performance scenarios.
Head-Related Transfer Functions (HRTFs) have fundamental applications for realistic rendering in immersive audio scenarios. However, they are strongly subject-dependent as they vary considerably depending on the shape of the ears, head and torso. Thus, personalization procedures are required for accurate binaural rendering. Recently, Denoising Diffusion Probabilistic Models (DDPMs), a class of generative learning techniques, have been applied to solve a variety of signal processing-related problems. In this paper, we propose a first approach for using DDPM conditioned on anthropometric measurements to generate personalized Head-Related Impulse Response (HRIR), the time-domain representation of HRTF. The results show the feasibility of DDPMs for HRTF personalization obtaining performance in line with state-of-the-art models.
We propose the Physics-Informed Neural Network-driven Sparse Field Discretization method (PINN-SFD), a novel self-supervised, physics-informed deep learning approach for addressing the Near-Field Acoustic Holography (NAH) problem. Unlike existing deep learning methods for NAH, which are predominantly supervised by large datasets, our approach does not require a training phase and it is physics-informed. The wave propagation field is discretized into sparse regions, a process referred to as field discretization, which includes a series of set of source planes, to address the inverse problem. Our method employs the discretized Kirchhoff-Helmholtz integral as the wave propagation model. By incorporating virtual planes, additional constraints are enforced near the actual sound source, improving the reconstruction process. Sparse optimization is carried out using Physics-Informed Neural Networks (PINNs), where physics-based constraints are integrated into the loss functions to account for both direct (from equivalent source plane to hologram plane) and additional (from virtual planes to hologram plane) wave propagation paths. Our comprehensive validation across various rectangular and violin top plates, covering a wide range of vibrational modes, demonstrates that PINN-SFD consistently outperforms the conventional Compressive-Equivalent Source Method (C-ESM), particularly in terms of reconstruction accuracy for complex vibrational patterns.
Spatial audio and virtual acoustics research heavily rely on measurement-based datasets of impulse responses to model and simulate real-world environments. Churches, in particular, are of significant interest due to their large, reverberant spaces and complex sound propagation characteristics. To this aim, we introduce ChurchIR, a dataset of church impulse responses acquired for spatial audio applications. The dataset has been captured using a 4th-order microphone and an omnidirectional microphone at thirty receiver positions for three different source locations. Key acoustic metrics and descriptors, including octave-band reverberation times and pseudo-spectra for source localization, are presented to characterize the dataset. ChurchIR offers researchers and developers a resource for spatial audio reproduction and analysis, providing them with data crucial for applications in immersive audio rendering, virtual reality, and architectural acoustics.
Churches are spaces designed with a unique acoustic identity, which is intimately connected to the oratory and musical needs of the historical period in which they were built. For instance, their typically long reverberation time is appropriate to specific uses, such as liturgical functions and choral music performances, but it may impair the repurposing of the space for other functions. Indeed, an acoustic environment suitable for choral or sacred music may not be compatible with other musical genres such as chamber music, solo performances, or small instrumental ensembles, which require greater clarity and frequency-balanced acoustic properties. In such cases, careful analysis of the environment and specific acoustic conditioning become essential steps to enable the space to be used for novel purposes, without compromising its artistic and historical integrity. In this work, we analyze and improve the acoustics of the church of Saints Marcellino and Pietro through space-time acoustic measurements and simulations. After developing and validating our model, we propose various solutions to optimize the church acoustics, transforming it into a functional concert hall while preserving its original identity and artistic grandeur.