Room impulse responses (RIRs) play a central role in sound field control (SFC), yet the common assumption that RIRs remain time-invariant in SFC methods is often unrealistic in real-world scenarios. As a result, real-time RIR tracking becomes essential for providing timely feedback to the control algorithm and maintaining system performance under dynamic acoustic conditions. In this work, we propose a multiband structured subband adaptive filtering approach to RIR tracking, which effectively reduces the impact of input signal correlation due to colored excitation. Additionally, we introduce a decorrelation-enhanced variant to further improve convergence speed. Simulation results support the theoretical analysis, demonstrating that the proposed subband-based methods consistently achieve up to a 10 dB reduction in steady-state error compared to the time-domain NLMS algorithm, while maintaining robust performance in rapidly changing acoustic environments.
Microphone array observations with a large number of microphones often exhibit low-rank characteristics. Effectively recognizing and leveraging this low-rank structure is essential for advanced microphone array processing, which motivates the development of low-rank beamforming techniques. Such beamformers are based on Kronecker product decomposition, and traditionally, the two sets of sub-filters are estimated through iterative algorithms. In this paper, we propose a neural low-rank beamformer that employs two neural networks to directly estimate the two sets of sub-filters. Each set is predicted by a Conformer-based network, which combines the strengths of Transformers and convolutional neural networks to capture both long-range dependencies and local features. Unlike prior neural Kronecker beamformers, which are limited to standard rectangular array topologies and first-order designs, the proposed method supports arbitrary array topologies and high-order low-rank beamformers. Experimental results demonstrate that the proposed low-rank neural minimum variance distortionless response (MVDR) beamformer consistently outperforms both conventional and Kronecker decomposition-based MVDR beamformers.
Superarrays refer to microphone arrays that combine both omnidirectional and directional sensors, enabling higher array gain or additional functionalities that traditional arrays with the same geometry cannot achieve. In our previous work, we demonstrated that concentric circular superarrays (CCSAs) can achieve directivity and three-dimensional (3D) steering capabilities comparable to volumetric arrays, but with a more compact, two-dimensional geometry. This makes CCSAs particularly promising for high-fidelity speech signal in compact devices. However, prior CCSA designs have been limited to omnidirectional and bidirectional sensors, restricting their flexibility and broader applicability. Additionally, the theoretical conditions required for effective array configuration remain unclear. This paper introduces a generalized framework for designing CCSAs and their corresponding fully steerable, frequency-invariant beamformers in 3D space. The contributions are twofold: 1) a method is proposed for designing 3D steerable, frequency-invariant beamformers using CCSAs equipped with directional microphones of arbitrary types; and 2) the necessary and sufficient conditions for array configurations are derived to enable effective design of first-, second-, and third-order frequency-invariant beamformers. We further summarize the construction of first-, second-, and third-order beamformers. Extensive simulations and real experiments validate the proposed method and demonstrate the practical efficacy and theoretical novelty of our approach.
Dereverberation, a crucial technique aimed at alleviating the adverse effects of reverberation, plays a central role in improving speech quality and intelligibility in speech communication and human-machine interface systems. Among many dereverberation methods, the approach utilizing delayed linear prediction is notably recognized for its effectiveness, but the resulting algorithms are in general computationally demanding. In this work, we introduce an approach that relies on the utilization of a third-order tensor decomposition to efficiently decrease computational complexity while retaining a competitive level of dereverberation performance compared to multichannel linear prediction-based algorithms. In this method, the global linear prediction filter is first represented as a composition of shorter filters through the Kronecker product (KP) decomposition. Adaptive algorithms are then derived to estimate those short filters. Complexity analysis and results from simulations showcase the advantages of the presented approach over traditional techniques.
Accurate estimation of the noise covariance matrix is critical yet challenging in multichannel speech enhancement. In this work, we propose a spatial covariance matrix (SCM) reconstruction method for speech enhancement using compact microphone arrays in reverberant, multi-source environments. At each time–frequency bin, the normalized SCM of the array observations is modeled as a linear combination of predefined coherence matrices representing individual sources, late reverberation, and ambient noise. The combination coefficients, termed variance ratios, are estimated by minimizing the Frobenius norm between the modeled and observed normalized SCMs, subject to non-negativity and unity-sum constraints. An adaptive algorithm is introduced to efficiently estimate these ratios, and the reconstructed SCMs are subsequently used in the multichannel Wiener filter. Simulation and experimental results show that the proposed SCM estimation method enables the multichannel Wiener filter to achieve robust and effective speech enhancement.
Microphone arrays are commonly used to extract source signals from noisy observations. Traditional methods rely on a priori information about the source to define the target of the array and derive optimal filters for extraction. However, in challenging acoustic environments-such as those with moving sources, multiple sources, or sporadic sound events-obtaining this information is often difficult. This can lead to performance degradation or even failure in extracting the desired source. In contrast, we propose a novel framework for array processing that defines the unwanted interferences and background noise to extract the source signal. Our approach enables real-time detection of new, unseen sources, allowing the array to possess self-awareness regarding emerging signals. This is achieved by calculating a residual model of the covariance matrix of the array observations, given the coherence matrices of the interferences and background noise. We introduce a two-stage method for calculating this residual model: the first stage computes the total coherence matrix of the interferences and noise, and the second stage determines the coherence matrix of of new source through the proposed matrix decomposition method, and further outputs the array's self-awareness, represented by the Wiener filter for extracting the new source. Through simulations, we demonstrate that our approach can automatically extract new emerging sources, regardless of the presence of moving sources, multiple sources, or sporadic sound events, provided that the coherence matrix of the unwanted interferences and noise is known. This work presents a significant advancement in source extraction techniques for adverse acoustic environments.
This paper proposes a robust low-complexity adaptive filtering algorithm for identifying impulse responses in acoustic echo cancellation within practical acoustic environments. Under a low-rank model framework of the filter coefficient vector, a Cauchy estimate function is employed to define a set of robust cost functions, from which the adaptive algorithm is established with a dichotomous coordinate descent scheme. The computational efficiency of the proposed algorithm is significantly improved, and its effectiveness is validated through simulations.
The increasing demand for high-fidelity acoustic signal acquisition, driven by applications such as smart-home and human-machine interface systems, has spurred the development of compact microphone arrays equipped with highly directive (HD) beamformers. In practical settings, the design of HD beamformers, including superdirective and differential beamformers, is typically framed as a constrained optimization problem to improve noise suppression and robustness. Although several methods have been proposed, most are based on convex quadratic constrained quadratic programming. However, designing HD beamformers under general signal-to-noise ratio (SNR) gain constraints presents a fundamental nonconvex challenge, requiring a more comprehensive approach. In this work, we propose a novel method for designing fixed HD beamformers under general SNR gain constraints by converting the nonconvex optimization problems into generalized eigenvalue problems. A detailed filter design methodology is then developed. The major contributions of this work are twofold: 1) the introduction of an innovative approach to fixed HD beamformer design under SNR gain constraints, and 2) the provision of a versatile solution applicable to a broad range of related problems. Simulations validate the proposed method and demonstrate its advantages over existing approaches.
Fixed beamforming is widely used in practice since it does not depend on the estimation of noise statistics and provides relatively stable performance. However, a single beamformer cannot adapt to varying acoustic conditions, which limits its interference suppression capability. To address this, adaptive convex combination (ACC) algorithms have been introduced, where the outputs of multiple fixed beamformers are linearly combined to improve robustness. Nevertheless, ACC often fails in highly non-stationary scenarios, such as rapidly moving interference, since its adaptive updates cannot reliably track rapid changes. To overcome this limitation, we propose a frame-online neural fusion framework for multiple distortionless differential beamformers, which estimates the combination weights through a neural network. Compared with conventional ACC, the proposed method adapts more effectively to dynamic acoustic environments, achieving stronger interference suppression while maintaining the distortionless constraint.
The recursive least-squares (RLS) algorithm is recognized as a powerful tool in adaptive filtering applications, being characterized by a fast convergence rate, even for highly correlated input signals. Its overall performance is mainly influenced by the forgetting factor, which is a memory-related parameter that is influenced by the filter length. On the other hand, its robustness (in noisy conditions) can be controlled by using an appropriate regularization term. In order to achieve a proper balance between the main performance criteria, a regularized RLS algorithm with variable (i.e., time-dependent) forgetting factor is designed in this paper. The proposed solution is tested in the framework of echo cancellation, being supported by several simulation results obtained in different noisy scenarios.
Recursive least-squares (RLS) adaptive algorithms are capable of outperforming least-mean-square (LMS) methods for the identification of long-length impulse responses due to their ability to mitigate the high correlation properties of input signals, such as speech. Despite the encouraging results obtained in terms of tracking speed and accuracy, with respect to LMS methods, most RLS algorithms manifest numerical stability issues. Moreover, when an unknown system changes, the identification process needs to adapt to the new impulse response as soon as possible. The algorithm can require a significant amount of time to generate new accurate results in acoustic echo cancellation (AEC) scenarios. Due to the slow propagation speed of sound, acoustic echo paths are usually modeled using thousands of numerical coefficients, and adaptation energy remains relatively limited. A compromise is usually made between tracking capabilities and steady-state accuracy when choosing the forgetting factor (the most important parameter of the RLS algorithm). This paper analyzes a variable forgetting factor (VFF) RLS type of adaptive filter combined with the conjugate gradient (CG) line search method, which is designed to avoid the classical matrix inversion approach. This VFF-RLS-CG adaptive method is not susceptible to numerical stability issues and is designed to adapt its statistical estimates by determining whether a tracking situation occurs or whether the unknown system is not significantly different. Correspondingly, when necessary, the forgetting factor is decreased for faster adaptation to changes in the working environment. When the filter is estimated to work at steady-state, the above-mentioned parameter’s value is increased in order to boost the accuracy of the adaptive filter. The theoretical model is validated using simulations in AEC scenarios with tracking occurrences and relevant steady-state intervals.
Multichannel linear prediction (MCLP) is widely used for speech dereverberation, with recursive least-squares (RLS)-like algorithms commonly applied to update the linear prediction coefficients. However, these algorithms tend to be computationally intensive, making it necessary in practical implementations to reduce complexity while improving numerical robustness for better dereverberation performance. In this paper, we introduce a more efficient MCLP-based adaptive dereverberation method that combines dichotomous coordinate descent (DCD) with a data-reuse (DR) technique. Compared to the traditional RLS-based approach, the proposed method offers two major benefits. First, it significantly lowers computational demands by replacing most multiplications with bitshifts during DCD iterations, making it more suitable for real-world applications. Second, by avoiding the propagation of the inverse covariance matrix via the Riccati equation, the method ensures numerical stability, making it more suitable for processing long-duration speech signals. Additionally, the DR technique improves dereverberation performance by more efficiently utilizing available observed data. Simulation results show that the proposed methods outperform the conventional RLS-based approach in terms of both numerical stability and computational efficiency, while delivering comparable dereverberation performance.
Binaural audio is essential for delivering immersive spatial auditory experiences through headsets. However, due to the high cost and complexity of binaural recording, there has been growing research interest in binaural audio synthesis (BAS) from monaural inputs. In natural listening environments, humans typically perceive multiple concurrent sound sources, yet most existing BAS approaches render each source independently, relying on perfect source signal separation, a condition rarely achievable in practice and often leading to perceptual quality degradation. To address this limitation, this paper proposes MixBAS, a transformer based end-to-end multi-source mono-to-binaural synthesis framework that eliminates the need for explicit source separation. We design an asymmetric transformer that spatializes a mono mixture, which comprises both speech and non-speech components, into its binaural counterpart by incorporating a user-defined positional prompt for the non-speech source. When reproduced over headphones, the generated binaural audio enables listeners to perceive a high-quality speech signal along with a non-speech source rendered at a user-specified spatial location. Experimental results demonstrate that MixBAS significantly outperforms existing BAS baselines relying on source separation in both objective metrics and perceptual quality.
Direction-of-arrival (DOA) estimation and localization of acoustic sources following major natural disasters, such as devastating earthquakes, are crucial for responding to immediate impacts and conducting search-and-rescue operations. With the rapid advancement of unmanned-aerial-vehicle (UAV) technologies, UAVs have become an excellent choice for carrying sensing and detection systems, as they offer better accessibility to disaster-stricken areas that are difficult for rescue teams to reach. However, a major challenge is that the acoustic sensing systems on UAVs are often affected by strong ego and environmental noise, leading to extremely low signal-to-noise ratios (SNRs), typically well below 0 dB, which makes most DOA estimation and source localization algorithms ineffective. To address this challenge, this work explores the design of dipping microphone arrays carried by UAVs and the associated DOA estimation algorithms. The major contributions are threefold: 1) A dipping microphone array with a reconfigurable topology is designed, which significantly improves the SNR by adjusting the dipping length and enhances DOA estimation by configuring the array topology; 2) A maximum front-to-back ratio (MFBR) beamformer is developed to further mitigate the impact of UAV ego noise, further improving the SNR; 3) Building on the use of the dipping array and MFBR beamformer, a multiple-signal-classification (MUSIC) like algorithm is proposed to achieve accurate DOA estimation in challenging acoustic environments. Simulations and experiments are conducted to validate the effectiveness of the proposed design and algorithms.
Online blind source separation is essential for both speech communication and human-machine interaction. Among existing approaches, overdetermined independent vector analysis (OverIVA) delivers strong performance by exploiting the statistical independence of source signals and the orthogonality between source and noise subspaces. However, when applied to large microphone arrays, the number of parameters grows rapidly, which can degrade online estimation accuracy. To overcome this challenge, we propose decomposing each long separation filter into a bilinear form of two shorter filters, thereby reducing the number of parameters. Because the two filters are closely coupled, we design an alternating iterative projection algorithm to update them in turn. Simulation results show that, with far fewer parameters, the proposed method achieves improved performance and robustness.
Adaptive filtering algorithms based on tensor decomposition represent appealing choices for system identification problems, especially when dealing with the estimation of long-length impulse responses, like in acoustic echo cancellation. The topic has recently been addressed in the literature, showing that the gain (compared to the conventional approach) is twofold in terms of both better performance and lower complexity. The main idea is that a system identification problem with a large parameter space (i.e., a long-length filter) is reformulated based on a group of shorter filters, while their coefficients are combined using the Kronecker product. Nevertheless, one of the main challenges is related to handling the tensor rank, which is particularly addressed for each specific decomposition order. Previous solutions have been designed for second-order (matrix case) and third-order tensorial decompositions. In this paper, we develop a recursive least-squares adaptive filtering algorithm that exploits a fourth-order tensor decomposition, aiming for further performance improvements compared to the existing solutions. In this framework, the influence of the decomposition setup is investigated, which is also related to the main parameters of the algorithm, i.e., the forgetting factors. Simulations performed in the context of acoustic echo cancellation support the theoretical findings and indicate the good performance of the proposed algorithm.
Linear superarrays (LSAs) have been proposed to address the limited steering capability of conventional linear differential microphone arrays (LDMAs) by integrating omnidirectional and directional microphones, enabling more flexible beamformer designs. However, existing approaches remain limited because array geometry and element directivity, both critical to beamforming performance, are not jointly optimized. This paper presents a generalized LSA optimization framework that simultaneously optimizes array geometry, element directivity, and the beamforming filter to minimize the approximation error between the designed beampattern and an ideal directivity pattern (IDP) over the full frequency band and all steering directions within the region of interest. The beamformer is derived by approximating the IDP using a Jacobi-Anger series expansion, while the array geometry and element directivity are optimized via a genetic algorithm. Simulation results show that the proposed optimized array achieves lower approximation error than conventional LSAs across the target frequency band and steering range. Additionally, its directivity factor and white noise gain demonstrate more stable and improved performance across frequencies and steering angles.
The traditional minimum variance distortionless response (MVDR) beamformer requires the estimation of the noise covariance matrix, which poses significant challenges in complex acoustic environments. Recently, model-based approaches have been introduced to address this issue, showing promising results. However, a major drawback of these methods is their high computational complexity, particularly in estimating model parameters, such as the direction of interference. To address this challenge, we propose a low-complexity, model-based beamforming technique tailored for small-spacing linear microphone arrays, commonly found in consumer devices. Drawing inspiration from a useful decomposition, we derive an approximate MVDR beamformer that combines a regularized superdirective beamformer with a first-order adaptive differential beamformer. This approach eliminates the need for explicit estimation of noise statistics and interference direction. Instead, the method relies on two key parameters: one related to the spectral characteristics of the noise, and the other to the direction of the noise source. An alternating iterative optimization process is introduced to determine the optimal values of these parameters by minimizing the variance of the array output. Simulation results show that the proposed beamformer significantly reduces computational complexity compared to two leading model-based MVDR beamformers, while also outperforming other baseline algorithms in terms of speech enhancement performance.
Differential microphone arrays (DMAs) are recognized for their highly directive broadband beampatterns and have attracted significant interest in the design of compact microphone arrays. It has been shown that increasing the number of microphones in a DMA can improve array performance. However, when applying DMAs to embedded systems, this creates challenges due to the increased number of parameters, higher computational complexity, and the need to maintain the array’s robustness. To address these challenges, this paper presents a method for designing robust low-rank (LR) differential beamformers. Initially, we extend traditional differential beamforming by introducing an LR differential beamforming framework, which represents a long filter as the Kronecker product of two sets of shorter filters, significantly reducing both the number of parameters and computational complexity. Next, we derive robust designs for the two sets of shorter filters by maximizing the directivity factor (DF) subject to a white noise gain (WNG) constraint, or by maximizing the WNG subject to a DF constraint. This results in two types of LR differential beamformers that achieve the desired DF or WNG levels. The optimization problems are formulated and transformed into quadratic eigenvalue problems (QEPs), leading to closed-form solutions for both the WNG-constrained and DF-constrained LR differential beamformers. Simulation results demonstrate the effectiveness of the proposed method, confirming its robustness and enhanced computational efficiency.
In this work, we present a new perspective on the origin and interpretation of adaptive filters. By applying Bayesian principles of recursive inference from the state-space model and using a series of simplifications regarding the structure of the solution, we can present, in a unified framework, derivations of many adaptive filters that depend on the probabilistic model of the measurement noise. In particular, under a Gaussian model, we obtain solutions well-known in the literature (such as LMS, NLMS, or Kalman filter), while using non-Gaussian noise, we derive new adaptive algorithms. Notably, under the assumption of Laplacian noise, we obtain a family of robust filters of which the sign-error algorithm is a well-known member, while other algorithms, derived effortlessly in the proposed framework, are entirely new. Numerical examples are shown to illustrate the properties and provide a better insight into the performance of the derived adaptive filters.