Room impulse responses (RIRs) play a central role in sound field control (SFC), yet the common assumption that RIRs remain time-invariant in SFC methods is often unrealistic in real-world scenarios. As a result, real-time RIR tracking becomes essential for providing timely feedback to the control algorithm and maintaining system performance under dynamic acoustic conditions. In this work, we propose a multiband structured subband adaptive filtering approach to RIR tracking, which effectively reduces the impact of input signal correlation due to colored excitation. Additionally, we introduce a decorrelation-enhanced variant to further improve convergence speed. Simulation results support the theoretical analysis, demonstrating that the proposed subband-based methods consistently achieve up to a 10 dB reduction in steady-state error compared to the time-domain NLMS algorithm, while maintaining robust performance in rapidly changing acoustic environments.
This work introduces a passive method for measuring the speed of sound in indoor reverberant environments using speech signals captured by first-order Ambisonic (FOA) microphones. This estimation is a key step in enhancing array signal processing performance, especially in environments where acoustic properties vary due to temperature and humidity changes. By leveraging the pseudo-intensity vector (PIV) derived from FOA signals, we can more accurately estimate the direction-of-arrival (DOA) of a source, independent of the sound speed. Additionally, we present a geometric model to compute the instantaneous sound speed based on the DOA and time delays estimated from signals recorded by two FOA microphones. The study further examines how factors such as source DOA, time delay, and FOA microphone spacing impact the accuracy of the estimate. Both simulations and experiments are conducted to validate the proposed method. Moreover, we demonstrate how instantaneous sound speed estimation can be used for beamforming calibration, facilitating robust beamformer design in time-varying acoustic environments.
Ambisonics delivers compact scene based spatial audio representation, yet higher order Ambisonic encoding poses difficulties for wearables and embedded hardware. Their microphone arrays are often sparse, irregular, and constrained by device specific boundary conditions. These factors make the spherical-harmonic (SH) domain encoding ill conditioned: inverse filtering amplifies noise, while deterministic neural encoders may overfit to array-specific responses or smooth ambiguous higher-order components. This paper presents DiffM2A, a geometry-adaptive conditional diffusion framework for robust Ambisonic encoding from sparse MAs with variable topologies. Its Geometry-Adaptive Spherical Harmonic Projection (GASHP) front-end constructs boundary-aware SH steering functions and applies an energy-normalized modal projection, mapping array-dependent observations to a common modal representation without explicit pseudo-inverse computation. A dual-branch Elucidated Diffusion Model then estimates complex Ambisonic coefficients, conditioned on both the raw microphone spectra and GASHP features. Sound intensity and rotational equivariance losses further enhance inter-channel phase consistency and structured behavior across SH subspaces. Evaluations on both first- and second-order Ambisonic encoding tasks, using simulated room-acoustics and real-world LOCATA recordings, demonstrate that DiffM2A outperforms conventional and neural baseline methods on signal fidelity, spectral accuracy, spatial coherence, and binaural cue preservation. Additional experiments show that these gains are largely retained across unseen five-microphone layouts and under mismatched open-array and rigid-sphere boundary models.
The existing Generative Fixed-Filter Active Noise Control (GFANC) method generates a suitable control filter based on the current noise frame. This reactive design aims to estimate a control filter that is optimal for the present frame rather than the upcoming one. Consequently, it suffers from an inherent tracking lag and lacks the predictive capability to handle rapidly varying noises. To address this limitation, we propose the Predictive Fixed-Filter Active Noise Control (PFANC) method with a proactive control paradigm in this paper. In the PFANC method, multiple consecutive noise frames are processed by a Convolutional Recurrent Neural Network (CRNN) to predict the next-frame control filter. By utilizing temporal correlations across noise frames to anticipate the control filter in advance, the PFANC method can effectively track dynamic noise changes. Furthermore, the theoretical analysis based on a high-order Markov chain shows that incorporating multiple noise frames enhances the prediction of the control filter. Numerical simulations with linear and logarithmic chirp signals, as well as real-world dynamic noises, validate the effectiveness of the PFANC method and its superiority over GFANC and its variations. The PFANC method also exhibits good transferability across different acoustic paths.
Modeling the relationships that may connect optimal parameter vectors is essential for the performance of parameter estimation methods in distributed networks. In this paper, we consider a low-rank relationship and introduce matrix factorization to promote this low-rank property. To devise a distributed algorithm that does not require any prior knowledge about the low-rank space, we first formulate local optimization problems at each node, which are subsequently addressed using the Alternating Direction Method of Multipliers (ADMM). Three subproblems naturally arise from ADMM, each resolved in an online manner with low computational costs. Specifically, the first one is solved using stochastic gradient descent (SGD), while the other two are handled using the conjugate gradient descent method to avoid matrix inversion operations. To further enhance performance, a variance reduction algorithm is incorporated into the SGD. Simulation results validate the effectiveness of the proposed algorithm.
Due to its ability to handle strict constraints on feasible domains, distributed optimization over a Riemannian manifold offers an attractive solution for many practical applications. To develop such an algorithm for scenarios where the explicit expression of the cost function is unavailable, we introduce the zeroth-order (ZO) Riemannian stochastic gradient into distributed optimization on a Riemannian manifold. Specifically, an intermediate estimate is first obtained through a local update step using the ZO Riemannian stochastic gradient, which is approximated based on two function evaluations. Subsequently, an improved estimate is derived by minimizing the weighted Fr & eacute;chet mean over the manifold using information from neighboring nodes. To further enhance performance, a mini-batch strategy is incorporated into the gradient estimation process. Finally, simulation results are presented to validate the effectiveness of the proposed algorithm.
Sound field control (SFC) aims to accurately reproduce a desired sound field within a specified region, which requires both adaptation to input signal characteristics and precise estimation of acoustic paths between loudspeakers and microphones. To meet these demands, two adaptive algorithms are proposed. The first is a signal-adaptive multichannel filteredx least-mean-square (MCFxLMS) filter designed to handle nonstationary input signals such as speech and music. The second is an acoustic path tracking algorithm that incorporates an input signal decorrelation strategy, enabling robust tracking of multichannel room impulse responses (RIRs) even under highly correlated excitation. Additionally, an alternating modeswitching mechanism is introduced to dynamically activate each algorithm based on predefined criteria. This approach can reduce computational complexity in large-scale multichannel systems while preserving sound field fidelity within the control region.
Modeling multitask relations in distributed networks has garnered considerable interest in recent years. In this paper, we present a novel rank-one model, where all the optimal vectors to be estimated are scaled versions of an unknown vector to be determined. By considering the rank-one relation, we develop a constrained centralized optimization problem, and after a decoupling process, it is solved in a distributed way by using the projected gradient descent method. To perform an efficient calculation of this projection, we suggest substituting the intensive singular value decomposition with the computationally efficient power method. Additionally, local estimates targeting the same optimal vector are combined within a neighborhood to further improve their accuracy. Theoretical analyses of the proposed algorithm are conducted for star topologies, and conditions are derived to guarantee its stability in both the mean and mean-square senses. Finally, simulation results are presented to demonstrate the effectiveness of the proposed algorithms.
Conventional online secondary path modeling (SPM) for multi-channel active noise control (ANC) is often computationally intensive and unstable, particularly in large-scale systems. To address this, we propose a novel hybrid ANC method that leverages an adaptive kernel for SPM. The approach models a small subset of secondary paths in real-time and uses an adaptive kernel, whose hyperparameters are optimized via gradient descent, to accurately interpolate the remainder. Simulations in a multi-loudspeaker, multi-microphone setup demonstrate that the proposed method achieves steady-state performance comparable to conventional full-path modeling while exhibiting significantly faster and more stable convergence performance in dynamic acoustic environments.
Accurate modeling and analysis of a radiator mounted on an infinite baffle are crucial for understanding its acoustic radiation characteristics. This paper investigates the radiation behavior of a convex dome-shaped radiator in such a condition, showing that, under the far-field approximation, the pressure field is the three-dimensional Fourier transform of the axisymmetric surface velocity distribution. The study includes a comparison of various typical velocity distributions in terms of their directivity factor (DF) and radiated sound power. Current velocity distributions often suffer from nulls in the DF, so we provide a detailed analysis to uncover the causes of this issue. To address this, we propose an equalization filter designed to smooth the DF across the entire frequency range. Simulations are performed to validate the theoretical findings and to showcase the improved performance of the proposed approach.
Accurate modeling and analysis of baffled radiators are crucial for understanding their acoustic radiation characteristics, as these models serve as the foundation for evaluating the performance of real-world loudspeakers. This paper focuses on the radiation behavior of a dome -shaped radiator in an infinite baffle. A key theoretical insight is that the far-field sound pressure of a dome radiator can be expressed as a three-dimensional Fourier transform (3D-FT) integral of its surface velocity. Building on this foundation, we derive theoretical expressions for several typical acoustical quantities, including beampattern, directivity factor (DF), specific radiation impedance, and radiated sound power, under the far-field approximation and axisymmetric velocity distribution conditions. Next, we address the null problems observed in the DF of current velocity distributions, which arise due to the complex exponential integral term in the DF's numerator. A detailed analysis is performed to identify and understand the root causes of this issue. As a solution, we design an equalization filter based on the frequency spectral characteristics of an actual vibrating dome to smooth the DF across the entire frequency range. A performance comparison between several typical vibrating velocity distributions and the proposed distribution is presented, with evaluation results demonstrating that the proposed velocity distribution achieves a higher and smoother DF than existing models.
Determining an accurate rank is essential for parameter estimation in low-rank distributed networks. To address this challenge, this paper proposes a rank-adaptive learning algorithm that ensures the estimated local matrices match the true rank. Given a strongly convex cost function at each node in the network, matrix factorization is firstly employed to formulate the optimization problem, decomposing a matrix into the product of two low-rank matrices to maintain a potentially low-rank structure. To promote row sparsity, a weighted sparse norm is imposed on one of the factorized matrices as a regularization term. Since the choice of weights critically affects rank estimation, an adaptive strategy is introduced to set weights. The local optimization problem is then solved using the half-quadratic splitting (HQS) algorithm. Finally, simulation results demonstrate the effectiveness of the proposed algorithm.
Image composition has advanced significantly with large-scale pre-trained T2I diffusion models. Despite progress in same-domain composition, cross-domain composition remains under-explored. The main challenges are the stochastic nature of diffusion models and the style gap between input images, leading to failures and artifacts. Additionally, heavy reliance on text prompts limits practical applications. This paper presents the first cross-domain image composition method that does not require text prompts, allowing natural stylization and seamless compositions. Our method is efficient and robust, preserving the diffusion prior, as it involves minor steps for backward inversion and forward denoising without training the diffuser. Our method also uses a simple multilayer perceptron network to integrate CLIP features from foreground and background, manipulating diffusion with a local cross-attention strategy. It effectively preserves foreground content while enabling stable stylization without a pre-stylization network. Finally, we create a benchmark dataset with diverse contents and styles for fair evaluation, addressing the lack of testing datasets for cross-domain image composition. Our method outperforms state-of-the-art techniques in both qualitative and quantitative evaluations, significantly improving the LPIPS score by 30.5
This study investigates methods for modeling surface velocity on a vibrating spherical cap, offering valuable insights for loudspeaker design. While various methods exist, the spherical cap on a rigid surface stands out for its modeling accuracy and theoretical comprehensiveness. Despite extensive discussions on this model, there still exists a scarcity of in-depth theoretical analysis and performance comparisons. To address this gap, the paper first derives the theoretical directivity metric in the spherical harmonic (SH) domain for the spherical cap, enhancing understanding of its directional radiation characteristics. The maximum achievable bound of directivity factor (DF) is then determined through a solution of a generalized Rayleigh quotient. Subsequently, a modified spherical cap model incorporating frequency-dependent and radial equalizing velocity is proposed, aiming to enhance directional radiation performance especially at lower frequencies. Finally, a comparison of different spherical cap surface velocity models is provided, revealing through simulation that the proposed model outperforms others in achieving better directivity within the frequency range from 0 to 3000 Hz.
Nonlinear acoustic echo cancellation (NAEC) is of significant importance in acoustic telecommunication. To improve NAEC performance in the double-talk case, semi-blind source separation-based NAEC (SBSS-NAEC) algorithms have been proposed. However, to deal with reverberation and loudspeaker nonlinearities, convolutive transfer function (CTF) models and power series expansions are employed, which significantly increase the number of free parameters and consequently lead to slow convergence speed and, hence, limited performance. In this paper, we introduce the data-reuse strategy, well-known in the adaptive filter literature, into an SBSS-NAEC framework and propose two algorithms: data-reuse iteration projection (DR-IP) and data-reuse element-wise iterative source steering (DR-EISS). Several simulations demonstrate the superiority of the proposed methods, especially the tracking capability when the impulse response changes.
Accurately representing the sound field with high spatial resolution is crucial for immersive and interactive sound field reproduction technology. In recent studies, there has been a notable emphasis on efficiently estimating sound fields from a limited number of discrete observations. In particular, kernel-based methods using Gaussian processes (GPs) with a covariance function to model spatial correlations have been proposed. However, the current methods rely on pre-defined kernels for modeling, requiring the manual identification of optimal kernels and their parameters for different sound fields. In this work, we propose a novel approach that parameterizes GPs using a deep neural network based on neural processes (NPs) to reconstruct the magnitude of the sound field. This method has the advantage of dynamically learning kernels from data using an attention mechanism, allowing for greater flexibility and adaptability to the acoustic properties of the sound field. Numerical experiments demonstrate that our proposed approach outperforms current methods in reconstructing accuracy, providing a promising alternative for sound field reconstruction.
Direction-of-arrival (DOA) estimation in environments with multiple sources and strong reverberation remains a great challenge. In this letter, we present a novel feature, the higher-order pseudo-intensity vector (HOPIV), derived from recordings obtained with a spherical microphone array. By exploiting the unique properties of the reactive intensity vector, which is derived from the HOPIV, we present a method to identify time-frequency points that are dominated by the direct path. We then propose a DOA estimation method that leverages the HOPIV's high spatial resolution for improving DOA estimation performance. Simulations and experiments show that the proposed method is able to yield superior performance compared to state-of-the-art techniques, even in highly reverberant environments.
The localization of the sound source is crucial in audio signal processing, especially when using acoustic vector sensors. Traditional methods typically utilize active intensity vector information to estimate the direction of arrival (DOA). However, these methods often fail to estimate the DOA of small-amplitude sound sources in the presence of high interference. This study introduces an innovative approach that uses the reactive intensity vector as a clue to enhance the estimation of DOA for sound sources. The simulation results confirm that our proposed method consistently outperforms traditional techniques in estimating DOA for small-amplitude sound sources.
This work proposes an active road noise control (RNC) system that effectively controls non-stationary noise inside the cabin by using a data-driven model to predict primary noise in passenger’s ears from sensors on the car chassis, with a fixed filter design integrated into headrest speakers. We develop STFNet, a novel network that fully exploits spatial, temporal, and spectral information between reference noise recorded by the multiple accelerometers and primary noise present in the passenger’s ears. Extensive testing with real-world recordings at various driving speeds demonstrates that our system not only achieves higher noise reduction, but also demonstrates significantly faster convergence performance compared with traditional multi-channel RNC system.