In dynamic acoustic environments characterized by time-varying interferers and moving sources, effective beamforming requires accurately identifying stationary regions over time. Traditional Capon beamformers rely on the instantaneous ensemble covariance matrix, which is inaccessible in practice. Practical implementations overcome this by estimating the sample covariance matrix (SCM) through averaging over a block of temporal samples. However, in non-stationary settings, a naive batch approach fails. Moving interferers smear the SCM, causing the beamformer to place nulls in outdated locations while failing to track newly active interferers, thereby degrading its nulling capabilities. To address this fundamental limitation, an Online Segmented Beamformer is proposed. This algorithm incorporates data-driven temporal segmentation to causally minimize output power while dynamically adapting the SCM estimation windows to local stationarity. By framing the problem through the lens of dynamic programming, the proposed method tracks abrupt environmental changes and resets covariance estimates in real-time. We validate the performance of this framework in a complex, reverberant simulated acoustic environment and in highly reverberant real world experiments, demonstrating its superiority over fixed-window adaptive methods.
Reliable adaptive beamforming is critical for large microphone arrays operating in highly dynamic acoustic environments. In scenarios characterized by fast-moving talkers and interferers, the available sample support for estimating the spatial correlation matrix is often snapshot-deficient. This deficiency degrades the White Noise Gain (WNG), leading to severe target signal cancellation. To ensure stable and robust beamforming, we previously proposed an adaptive diagonal loading method that leverages the Kantorovich inequality to guarantee the WNG remains strictly within specified bounds. However, accurately determining the smallest necessary loading level requires calculating the extreme eigenvalues of the spatial correlation matrix, a computationally expensive 𝒪(M^3) operation for large arrays. In this paper, we introduce a highly efficient 𝒪(kM^2) estimation technique using Lanczos iterations to build a small Krylov subspace. By projecting the correlation matrix onto a tridiagonal matrix of dimension k ≪ M, we extract Ritz values that rapidly converge to the exact extreme eigenvalues. Our evaluations demonstrate that this Lanczos-accelerated approach achieves performance identical to exact Eigenvalue Decomposition (EVD), ensuring optimal interference suppression and strict WNG adherence at a fraction of the computational cost.
In dynamic acoustic environments with time-varying interferers, effective beamforming requires identifying stationary regions over time. The Capon beamformer, a whitened matched filter constrained to maintain unity gain in the desired direction, theoretically relies on the instantaneous ensemble covariance matrix. Practical implementations rely on the batch Capon (or Sample Matrix Inversion), which estimates the sample covariance matrix (SCM) by averaging over a block of snapshots. This practical approach implicitly assumes that the data within the batch window is stationary and can be coherently combined. In non-stationary settings, a batch approach that averages over fixed or excessively long windows fails, as moving interferers smear the SCM and degrade the beamformer's nulling capabilities. To address this, this paper introduces a temporally segmented distortionless response beamformer. Inspired by the segmented least squares method, which fits piecewise polynomials to data while penalizing excessive segmentation to prevent overfitting, the framework extends practical Capon beamforming by incorporating data-driven temporal segmentation. This formulation minimizes output power while dynamically adapting the SCM estimation windows to local stationarity, offering a principled approach to tracking time-varying interferers.
Adaptive beamforming is a cornerstone of array signal processing, yet its performance often collapses in the face of complex, rapidly changing interference. When interferers appear or move unpredictably, conventional estimators encounter a fundamental memory trade-off: short windows enable rapid tracking but suffer from high estimation variance, while long windows provide stable rejection but fail to adapt to shifts. This challenge is resolved by introducing the Universal Switching Beamformer (USB), which integrates competitive sequential prediction into the beamforming architecture. By employing a linear transition diagram, the USB implicitly maintains an exponentially large family of candidate covariance histories and dynamically re-weights them based on their cumulative output power. This mechanism allows the beamformer to automatically vary its effective memory length without explicit change detection or heuristic parameter tuning. A theoretical upper bound is proven on the regret relative to an omniscient oracle that selects the best piecewise-stationary covariance model in hindsight. Extensive simulations and experiments on the SwellEx-96 dataset demonstrate that the USB achieves the agility of short-window estimators and the precision of long-term integration, providing a principled solution for tracking highly non-stationary scenes.
Reliable adaptive beamforming is critical for large microphone arrays operating in highly dynamic acoustic environments. In scenarios characterized by fast-moving talkers and interferers, the available sample support for estimating the spatial correlation matrix is often snapshot-deficient. This deficiency, coupled with array imperfections, degrades the White Noise Gain (WNG), leading to severe target signal cancellation. To ensure stable and robust beamforming, we propose a novel adaptive diagonal loading method that guarantees the WNG remains strictly within specified bounds. By leveraging the Kantorovich inequality, we map the desired WNG to a strict upper bound on the condition number of the correlation matrix. Furthermore, we present three estimation techniques for the adaptive loading level, ranging from trace-based bounding to exact eigenvalue decomposition, offering scalable computational complexities of 𝒪(M), 𝒪(M^2), and 𝒪(M^3). Our approach demonstrates highly stable beamforming under fast-changing interference.
Traditional volumetric noise control typically relies on multipoint error minimization to suppress sound energy across a region, but offers limited flexibility in shaping spatial responses. This paper introduces a time domain formulation for linearly constrained minimum variance active noise control (LCMV ANC) for spatial control filter design. We demonstrate how the LCMV ANC optimization framework allows system designers to prioritize noise reduction at specific spatial locations through strategically defined linear constraints, providing a more flexible alternative to uniformly weighted multi point error minimization. An adaptive algorithm based of filtered X least mean squares (FxLMS) is derived for online adaptation of filter coefficients. Simulation and experimental results validate the proposed method's noise reduction and constraint adherence, demonstrating effective, spatially selective and broadband noise control compared to multipoint volumetric noise control.
An adaptive beamformer suppresses interferers and provides spatial filtering gains by making use of the sample covariance matrix. Updates to the sample covariance matrix reflect changes in the environment to which the beamformer must adapt. In environments with intermittent interferers, it is beneficial to remember the “state” that represents a specific pattern of interferer activity. In such cases, an adaptive beamformer that simply averages all the snapshots may result in reduced performance (with respect to an omniscient, context-aware beamformer that is aware of the interferer state) as the beamformer wastes degrees of freedom suppressing interferers that are always not active. By using the directional cosine of the peak of the beamformer scanned response as an information-bearing sequence, we partition the space into angular sectors that represent beamformers averaging a different set of snapshots. To represent and efficiently mix the output of all beamformers represented by such partitions, we employ a context tree that has been previously used for data compression and piecewise linear prediction. We use the context tree to achieve the signal estimation error of the best piecewise adaptive beamformer that can choose the partition of the directional cosine space.
Beamformers often trade off white noise gain against the ability to suppress interferers. With distributed microphone arrays, this trade-off becomes crucial as different arrays capture vastly different magnitude and phase differences for each source. We propose the use of multiple random projections as a first-stage preprocessing scheme in a data-driven approach to dimensionality reduction and beamforming. We show that a mixture beamformer derived from the use of multiple such random projections can effectively outperform the minimum variance distortionless response (MVDR) beamformer in terms of signal-to-noise ratio (SNR) and signal-to-interferer-and-noise ratio (SINR) gain. Moreover, our method introduces computational complexity as a trade-off in the design of adaptive beamformers, alongside noise gain and interferer suppression. This added degree of freedom allows the algorithm to better exploit the inherent structure of the received signal and achieve better real-time performance while requiring fewer computations. Finally, we derive upper and lower bounds for the output power of the compressed beamformer when compared to the full complexity MVDR beamformer.
An adaptive beamformer may be thought of as trading white noise gain for interferer suppression. The beamformer can respond to changing environmental statistics through updates to the sample covariance matrix. In time-varying environments, adaptive beamformers are frequently used with pre-determined sliding windows or forgetting factors for such sample covariance estimation. Thus, an adaptive beamformer must a priori select the regions over which the data are assumed stationary. Such methods perform poorly when the environment suddenly changes, such as strong interferers entering or exiting the acoustic scene. Many real-world environments have intermittent interferers, and such a beamformer may waste degrees of freedom suppressing an interferer that is no longer active or neglecting to suppress one that is. We propose the use of universal methods over a class of time-partitioned beamformers. While there are an exponential number of possible partitions of a block of data into locally stationary regions, methods from universal data compression and prediction for piece-wise stationary sources provide a path for implicitly implementing, and mixing over them all, with only polynomial complexity. We employ a linear transition diagram from this literature to enable efficient performance-weighted mixing of beamformers of all possible such partitions.
Filtered-X LMS (FxLMS) is commonly used for active noise control (ANC), wherein the soundfield is minimized at a desired location. Given prior knowledge of the spatial region of the noise or control sources, we could improve FxLMS by adapting along the low-dimensional manifold of possible adaptive filter weights. We train an auto-encoder on the filter coefficients of the steady-state adaptive filter for each primary source location sampled from a given spatial region and constrain the weights of the adaptive filter to be the output of the decoder for a given state of latent variables. Then, we perform updates in the latent space and use the decoder to generate the cancellation filter. We evaluate how various neural network constraints and normalization techniques impact the convergence speed and steady-state mean squared error. Under certain conditions, our Latent FxLMS model converges in fewer steps with comparable steady-state error to the standard FxLMS.
Interpolating Room Impulse Responses (RIRs) at unmeasured locations within a space is a fundamental challenge in room acoustics, critical for applications such as volumetric noise cancellation, auralization, and spatial audio rendering. Traditional methods including linear interpolation in the time or frequency domains and basis decomposition techniques such as plane wave and spherical harmonic decomposition often fail to capture the non-linear variations in arrival times and energy decay caused by complex propagation paths and occlusions. In this work, we propose a neural network that manipulates the time shifts between RIRs using receiver coordinates as inputs. These time axis manipulations capture time domain misalignments across spatial locations, enabling the network to model spatially varying acoustic delays and reflections. By aligning and blending neighboring RIRs using the modeled time shifts, we interpolate RIRs at unseen positions with higher temporal and spectral fidelity. This approach focuses particularly on preserving early reflection structure and perceptual similarity while offering a data-driven alternative for accurate RIR field reconstruction in complex acoustic environments.
In real-time spatial audio algorithms, adaptive filters are employed to learn and track spatial and acoustic filters. When the space of possible filter weights is close to a low-dimensional manifold, we can improve the convergence rate of adaptive filters by adapting along this manifold. This can be achieved using latent adaptive filters, which constrain the weights to remain within the range space of an auto-encoder's decoder by updating in its latent space. Although previous studies have explored various latent adaptive filters for acoustic system identification, a comparison of the convergence rates among different latent adaptive filter structures has not yet been conducted. Additionally, it is well-established that traditional adaptive filters often struggle to track changes in acoustic impulse responses caused by the continuous movement of the source and receiver. However, the tracking performance of latent adaptive filters has not been investigated. In this study, we empirically evaluate the performance of acoustic impulse response identification and tracking across different variants of latent adaptive filters.
Large aperture arrays improve detection performance with higher gain, especially in low signal-to-noise ratio (SNR) applications such as underwater acoustic (UWA) source detection. However, large arrays are susceptible to phase errors because of limited spatial coherence of signals or a mismatch between the assumed signal models and the true signal models. To mitigate this issue, the array can be partitioned into smaller segments of sensors known as subapertures. The subapertures are processed coherently, and then the power outputs of the subapertures are combined. This processing is the spatial analog of the classic Welch power spectrum estimator which averages periodograms across time windows of a recording. Identifying the subaperture size which optimizes detection performance remains an open problem. We proposed a score function that indicates the detection performance of different array partitions without access to the ground truth. Numerical experiments using the SWellEx-96 data corrupted by additional noise show that the subaperture maximizing this detection score function achieves a better receiver operating characteristic (ROC) in low SNR cases compared to any fixed array partition candidate. [Work supported by ONR Code 321US.]
The Filtered-x LMS (FxLMS) algorithm is a linear adaptive filtering approach in active noise control that estimates the signal played from a secondary speaker to cancel the presence of a noise source in a microphone. One of the requirements for this algorithm is the filter coefficients of the secondary acoustic path between the secondary speaker and the error microphone. For scenarios where the secondary path is time-varying, we would require external information to determine the current path at each time step. If the secondary path at any time lies in a finite set of possible transfer functions, we present a performance-weighted blended FxLMS algorithm that is robust to rapid changes within the set. We compare the effects of blending paths with the online selection of a single path in the set. We also compare the blended approach with the noise injection method commonly used in online secondary-path estimation.
Conventional active noise control systems typically rely on fixed tap lengths in control filters. Determining an appropriate tap length is challenging, even within a straightforward time-invariant environment. This challenge is exacerbated when essential system responses, such as primary noise characteristics or primary path responses, remain partially unknown and necessitate the use of an adaptive filter. A trade-off emerges between convergence rate and steady-state performance when selecting a fixed tap length in adaptive filters. In time-varying environments, the optimal tap length may even dynamically shift over time. Thus, prior attempts have been made to introduce variable tap length algorithms to dynamically adapt tap length in real time. However, the performance of variable tap length algorithms can be sensitive to the choice of additional parameters and noise level. This paper exploits a model order weighting approach by combining the outputs of filters of different tap lengths based on their predicted noise control performance. This proposed method demands minimal additional prior information compared to the variable tap length methods. The noise control performance, as demonstrated by specific examples of active noise control, is presented and analyzed.
Many rooms are equipped with conferencing systems with a large number of microphones. However, these systems are often limited to a fixed number of delay-and-sum beamformers and typical automix software will pick the beamformer output with the most energy. A more constructive combination is possible if we have access to the output of multiple beams constructed by such commercial systems. This article proposes a beamspace projection that may effectively view such a commercial conferencing system as a low-latency dimensionality reduction operation. Such a projection can be formulated as a plane wave decomposition of the received signals. Experiments conducted in simulation show that beamspace projection can rapidly approach the SINR gain achievable with access to all the microphones with lower computational complexity. Finally, we use commercial hardware and its beamformer outputs to separate numerous talkers in a real world environment.
Sound field estimation is the process of analyzing and characterizing the distribution of sound waves in a particular physical space. The applications of sound field estimation extend to various areas, including the visualization of acoustic fields, interpolation of room impulse responses, identification of sound sources, capturing sound fields for spatial audio, and spatial active noise control, among other potential uses. In a previous study, a directionally weighted kernel has been used to estimate the sound field, where the priori information of source directions is employed to improve the estimation accuracy. In another separate study, spherical harmonics have been used to represent the sound field. However, the order of spherical harmonic coefficients was limited due to the limited number of microphones. This research introduces a novel method for sound field estimation using multiple microphones to sample a source-free volume. A physics-informed neural network is used to predict the spherical harmonics coefficients and locations of unknown sources to estimate the sound field.
Recording on moving devices leads to constant deformations in microphone array structure, which interferes with spatial audio processing. Previously, we have established that there is a locally linear mapping between the manifolds of relative transfer functions (RTF) and array geometries. This mapping also applies to trajectories of array deformations and their corresponding RTFs. We run segmented least squares (SLS) on the RTF trajectories to discern potential states of motion in an unsupervised manner. We then examine an implementation of SLS for adaptive beamforming in a deformable setting.
Reverberation poses a challenge for speech processing systems and is unavoidable in real environments. As such, acoustic signal processing researchers are interested in robust algorithms that can perform well regardless of reverb severity. Measuring reverberant speech requires access to rooms with the desired dimensions, which may not be readily available. Furthermore, for heavily instrumented recording setups such as our Mechatronic Acoustic Research System (MARS), moving equipment between locations is not practical. We propose a system of actuated multi-textured panels which can significantly alter the reverb properties of a static room. By changing the angle of these panels, we can smoothly transition from high to low reverberation times. We show this empirically and develop a system that provides a requested reverb level automatically, once configured for a fixed size installation. This tool can allow researchers to use a single room to take measurements applicable to a more diverse range of environments. We integrate our system into the MARS project, an automated tool for generating spatial audio datasets.