Mask-based beamforming is a popular geometry-agnostic approach for speech enhancement, typically applying a single mask across all microphones to estimate the required covariance matrices. While effective for compact arrays, this strategy may be suboptimal for spatially distributed microphones, where signal characteristics may vary strongly across microphones. To effectively capture the spatial diversity across microphones, we extend the mask-based beamformer to a multi-channel formulation, where each microphone is pre-filtered by a separate mask before covariance estimation. To address time-varying acoustic scenes, caused by spectro-temporal nonstationarity, we adopt a frame-causal online implementation with a sliding window. Experiments with simulated compact arrays and distributed microphones show that multi-channel masking yields a benefit over using a single mask when microphone signals differ substantially, while retaining similar performance in compact arrays. We further demonstrate the robustness of the multi-channel masking approach by comparing oracle ideal ratio masks to blind DNN-based mask estimation.
This paper introduces the topology-independent distributed multichannel Wiener filter (TI-dMWF), a novel algorithm for distributed node-specific signal estimation in wireless acoustic sensor networks (WASNs) with unconstrained topologies. The TI-dMWF enables each node in the network to compute its centralized multichannel Wiener filter solution by exchanging only low-dimensional fused signals, without requiring iterative estimation, unlike state-of-the-art approaches such as the topology-independent distributed adaptive node-specific signal estimation (TI-DANSE) algorithm. The TI-dMWF is proven optimal when each source is observed by either all nodes or only one node. Theoretical analysis and numerical simulations confirm that it achieves centralized estimation performance in a single run. Its latency as a function of the pruned-tree depth and its computational complexity are also analyzed. Its robustness is assessed in reverberant-room simulations under estimated second-order statistics, various network topologies, and deviations from the assumed observability model.
Recently, a complex variational autoencoder (VAE)-based single-channel speech enhancement system based on the DCCRN architecture has been proposed. In this system, a noise suppression VAE (NSVAE) learns to extract clean speech representations from noisy speech using pretrained clean speech and noise VAEs with skip connections. In this paper, we improve DCCRN-VAE by incorporating three key modifications: 1) removing the skip connections in the pretrained VAEs to encourage more informative speech and noise latent representations; 2) using β-VAE in pretraining to better balance reconstruction and latent space regularization; and 3) a NSVAE generating both speech and noise latent representations. Experiments show that the proposed system achieves comparable performance as the DCCRN and DCCRN-VAE baselines on the matched DNS3 dataset but outperforms the baselines on mismatched datasets (WSJ0-QUT, Voicebank-DEMEND), demonstrating improved generalization ability. In addition, an ablation study shows that a similar performance can be achieved with classical fine-tuning instead of adversarial training, resulting in a simpler training pipeline.
Acoustic feedback limits the maximum gain in hearing aids. In addition to several approaches based on adaptive filtering, recently a deep-neural-network-based feedback cancellation (DFC) approach has been proposed, which is trained via an open-loop framework. Since open-loop-trained DFC (DFC-OL) can become unstable during inference at high gains, in this paper we propose an in-the-loop-trained DFC (DFC-IL) that integrates the DFC directly into the optimisation loop. This allows the model to be exposed to unstable conditions during training. A two-stage training strategy involving pre-training on stable systems and fine-tuning on a wider gain range enables DFC-IL to learn robust howling reduction. Experimental results on measured feedback paths demonstrate that in scenarios with small gains, the proposed DFC-IL performs similarly to DFC-OL, and both exceed the performance of adaptive filters. In scenarios with high amplification gains, DFC-IL clearly outperforms DFC-OL by maintaining system stability.
This paper focuses on distributed node-specific signal estimation in topology-unconstrained wireless acoustic sensor networks (WASNs) where sensor nodes only transmit fused versions of their local sensor signals. For this task, the topology-independent (TI) distributed adaptive node-specific signal estimation (DANSE) algorithm (TI-DANSE) has previously been proposed. It converges towards the centralized signal estimation solution in non-fully connected and time-varying network topologies. However, the applicability of TI-DANSE in real-world scenarios is limited due to its slow convergence. The latter results from the fact that, in TI-DANSE, nodes only have access to the in-network sum of all fused signals in the WASN. We address this low convergence speed issue by introducing an improved TI-DANSE algorithm, referred to as TI-DANSE$<^>+$. The TI-DANSE$<^>+$ algorithm outperforms TI-DANSE in terms of convergence speed by letting the updating node use each partial in-network sum of fused signals (coming from its neighbors) separately, when updating its estimation parameters. In this way, the number of available degrees of freedom in the optimization problem at the updating node is increased, leading to faster convergence. This separate use of incoming partial in-network sums is further exploited by combining TI-DANSE$<^>+$ with a tree-pruning strategy that maximizes the number of neighbors at the updating node. In fully connected WASNs, it is observed that TI-DANSE$<^>+$ converges as fast as the original DANSE algorithm (the latter only defined for fully connected WASNs) while using peer-to-peer data transmission instead of broadcasting and thus saving communication bandwidth. If link failures occur, the convergence of TI-DANSE$<^>+$ towards the centralized solution is preserved without any change in its formulation. Altogether, the proposed TI-DANSE$<^>+$ algorithm can be viewed as an all-round alternative to DANSE and TI-DANSE which (i) merges the advantages of both, (ii) reconciliates their differences into a single formulation, and (iii) shows advantages of its own in terms of communication bandwidth usage. The convergence properties and signal estimation performance of TI-DANSE$<^>+$ are demonstrated through speech enhancement experiments in simulated topology-unconstrained WASNs.
Relative transfer functions (RTFs) of sound sources play a crucial role in beamforming, enabling effective noise and interference suppression. This paper addresses the challenge of online estimating the relative transfer function (RTF) vectors of multiple sound sources in noisy and reverberant environments, both addressing scenarios with source activations and deactivations. To estimate the RTF vector of a newly activating source, the conventional blind oblique projection (BOP) method relies on computationally expensive gradient descent, random additional vectors, and high signal-to-noise ratio (SNR) conditions. We propose a generalized method that replaces iterative optimization with a closed-form solution, utilizes orthogonal additional vectors to improve accuracy, and integrates noise whitening and subtraction to ensure robustness in low SNR scenarios. Simulations are performed using real-world reverberant noisy recordings, featuring three successively activating speakers. In addition, simulations are performed for highly dynamic scenarios with random source activations and deactivations. Simulation results with and without a-priori knowledge of the source activity pattern demonstrate that the proposed BOP method with orthogonal additional vectors and noise whitening outperforms the conventional BOP method and other baseline methods in terms of computational efficiency and signal-to-interferer-and-noise ratio improvement, when applying the estimated RTF vectors in a linearly constrained minimum variance beamformer.
This paper presents a sequential approach to simultaneously optimize a microphone array geometry and its corresponding beamformer weights. The primary objective is to address the problem of an unknown direction of arrival of a desired source within a given region of interest while maximizing the broadband array directivity criterion. Optimization constraints are developed and applied to guarantee minimal distortion of the desired source and high white noise gain. The proposed approach outperforms recent and traditional approaches, considering the average white noise gain in the entire region of interest and frequency spectrum. It is also preferable in terms of the directivity factor when the direction of arrival of the desired source significantly deviates from its nominal value.
In a wireless acoustic sensor network (WASN), devices (i.e., nodes) can collaborate through distributed algorithms to collectively perform audio signal processing tasks. This paper focuses on the distributed estimation of node-specific desired speech signals using network-wide Wiener filtering. The objective is to match the performance of a centralized system that would have access to all microphone signals, while reducing the communication bandwidth usage of the algorithm. Existing solutions, such as the distributed adaptive node-specific signal estimation (DANSE) algorithm, converge towards the multichannel Wiener filter (MWF) which solves a centralized linear minimum mean square error (LMMSE) signal estimation problem. However, they do so iteratively, which can be slow and impractical. Many solutions also assume that all nodes observe the same set of sources of interest, which is often not the case in practice. To overcome these limitations, we propose the distributed multichannel Wiener filter (dMWF) for fully connected WASNs. The dMWF is non-iterative and optimal even when nodes observe different sets of sources. In this algorithm, nodes exchange neighbor-pair-specific, low-dimensional (fused) signals estimating the contribution of sources observed by both nodes in the pair. We formally prove the optimality of dMWF and demonstrate its performance in simulated speech enhancement experiments. The proposed algorithm is shown to outperform DANSE in terms of objective metrics after short operation times, highlighting the benefit of its iterationless design.
Guided Source Separation (GSS) is a popular front-end for distant automatic speech recognition (ASR) systems using spatially distributed microphones. When considering spatially distributed microphones, the choice of reference microphone may have a large influence on the quality of the output signal and the downstream ASR performance. In GSS-based speech enhancement, reference microphone selection is typically performed using the signal-to-noise ratio (SNR), which is optimal for noise reduction but may neglect differences in early-to-late reverberation ratio (ELR) across microphones. In this paper, we propose two reference microphone selection methods for GSS-based speech enhancement that are based on the normalized ℓp-norm, either using only the normalized ℓp-norm or combining the normalized ℓp-norm and the SNR to account for both differences in SNR and ELR across microphones. Experimental evaluation using a CHiME-8 distant ASR system shows that the proposed ℓp-norm-based methods outperform the baseline method, reducing the macro-average word error rate.
The number of active sound sources is a key parameter in many acoustic signal processing tasks, such as source localization, source separation, and multi-microphone speech enhancement. This paper proposes a novel method for online source counting by detecting changes in the number of active sources based on spatial coherence. The proposed method exploits the fact that a single coherent source in spatially white background noise yields high spatial coherence, whereas only noise results in low spatial coherence. By applying a spatial whitening operation, the source counting problem is reformulated as a change detection task, aiming to identify the time frames when the number of active sources changes. The method leverages the generalized magnitude-squared coherence as a measure to quantify spatial coherence, providing features for a compact neural network trained to detect source count changes framewise. Simulation results with binaural hearing aids in reverberant acoustic scenes with up to 4 speakers and background noise demonstrate the effectiveness of the proposed method for online source counting.
Own voice pickup technology for hearable devices facilitates communication in noisy environments. Own voice reconstruction (OVR) systems enhance the quality and intelligibility of the recorded noisy own voice signals. Since disturbances affecting the recorded own voice signals depend on individual factors, personalized OVR systems have the potential to outperform generic OVR systems. In this paper, we propose personalizing OVR systems through data augmentation and fine-tuning, comparing them to their generic counterparts. We investigate the influence of personalization on speech quality assessed by objective metrics and conduct a subjective listening test to evaluate quality under various conditions. In addition, we assess the prediction accuracy of the objective metrics by comparing predicted quality with subjectively measured quality. Our findings suggest that personalized OVR provides benefits over generic OVR for some talkers only. Our results also indicate that performance comparisons between systems are not always accurately predicted by objective metrics. In particular, certain disturbances lead to a consistent overestimation of quality compared to actual subjective ratings.
A popular method to estimate the positions or directions-of-arrival (DOAs) of multiple sound sources using an array of microphones is based on steered-response power (SRP) beamforming. For a three-dimensional scenario, SRP-based methods need to jointly optimize three continuous variables for position estimation or two continuous variables for DOA estimation, which can be computationally expensive. In this paper, we propose novel methods for multi-source position and DOA estimation by exploiting properties of Euclidean distance matrices (EDMs) and their respective Gram matrices. In the proposed multi-source position estimation method only a single continuous variable, representing the distance between each source and a reference microphone, needs to be optimized. For each source, the optimal continuous distance variable and set of candidate time-difference of arrival (TDOA) estimates are determined by minimizing a cost function that is defined using the eigenvalues of the Gram matrix. The estimated relative source positions are then mapped to estimated absolute source positions by solving an orthogonal Procrustes problem for each source. The proposed multi-source DOA estimation method entirely eliminates the need for continuous variable optimization by defining a relative coordinate system per source such that one of its coordinate axes is aligned with the respective source DOA. The optimal set of candidate TDOA estimates is determined by minimizing a cost function that is defined using the eigenvalues of a rank-reduced Gram matrix. The computational cost of the proposed EDM-based methods is significantly reduced compared to the SRP-based methods. Experimental results for different source and microphone configurations show that the proposed EDM-based method consistently outperforms the SRP-based method in terms of two-source position and DOA estimation accuracy.
Spatially selective active noise control (SSANC) hearables aim to attenuate noise from certain directions at the eardrum while preserving desired speech arriving from selected directions. Existing SSANC systems typically assume an accurate estimate of the secondary path from the loudspeaker to the inner error microphone. In practice, however, this path varies across users and device fits, which can degrade performance and compromise system stability. This paper proposes a robust soft-constrained optimization framework that computes a single control filter by minimizing the average cost over a set of secondary path estimates derived from human measurements. Simulations and experiments on a real-time control platform show that the proposed approach slightly reduces mean performance relative to the matched case but substantially narrows the performance spread under secondary path mismatch. The proposed framework therefore provides a practical design strategy when accurate secondary path estimates are unavailable.
Recently, a spatially selective non-linear filter (SSF) has been proposed for target speaker extraction, using the target direction-of-arrival (DOA) as a spatial cue. Since learned intermediate features are tied to the microphone geometry, the performance of the SSF degrades significantly when evaluated on mismatched array geometries. In this paper, we propose a geometry-conditioned SSF (GC-SSF), which incorporates a geometry-conditioning branch based on FiLM layers. Furthermore, we propose a feature that jointly encodes the DOA and the microphone positions (DOA-MPE). The conditioning branch modulates the intermediate feature maps of the SSF using the DOA-MPE feature to capture the spatial relationship between the microphone positions and the target speaker. Experimental results across circular, uniform linear, and random microphone arrays show that the proposed GC-SSF generalizes better to mismatched geometries while maintaining high spatial selectivity, demonstrating its ability to effectively adapt the filtering process to different array geometries
We introduce a region-of-interest beamforming approach for audio conferencing that addresses dynamic acoustics and multiple-speaker scenarios. The approach employs a two-stage sparse optimization to select a subset of microphones from dual circular sector arrays: first on the xz plane and then on the xy plane, balancing spatial resolution and efficiency. Using the dual circular layout, we are able to reduce response variability across azimuth and elevation angles. The proposed approach maximizes broadband directivity while ensuring a controlled level of distortion and minimal white noise gain. Compared to existing methods, the mainlobe attained by the resulting beamformer is more accurately aligned with the region of interest. It also achieves a preferable sidelobe and backlobe suppression. Finally, the proposed approach is shown to be superior considering the directivity factor and white noise gain, in particular at medium and high frequencies.
Hearable devices, equipped with one or more microphones, are commonly used for speech communication. Here, we consider the scenario where a hearable is used to capture the user's own voice in a noisy environment. In this scenario, own voice reconstruction (OVR) is essential for enhancing the quality and intelligibility of the recorded noisy own voice signals for telephony applications. In previous work, we developed a deep learning-based OVR system, aiming to reduce the amount of device-specific recorded signals for training by using data augmentation with phoneme-dependent models of own voice transfer characteristics. Given the limited computational resources available on hearables, in this paper we propose low-complexity variants of an OVR system based on the frequency and time joint non-linear filter (FT-JNF) architecture and investigate the required amount of device-specific recorded signals for effective data augmentation and fine-tuning. Simulation results show that the proposed OVR system considerably improves speech quality, even under constraints of low complexity and a limited amount of device-specific recorded signals.
This paper addresses the challenge of topology-independent (TI) distributed adaptive node-specific signal estimation (DANSE) in wireless acoustic sensor networks (WASNs) where sensor nodes exchange only fused versions of their local signals. An algorithm named TI-DANSE has previously been presented to handle non-fully connected WASNs. However, its slow iterative convergence towards the optimal solution limits its applicability. To address this, we propose in this paper the TI-DANSE+ algorithm. At each iteration in TI-DANSE+, the node set to update its local parameters is allowed to exploit each individual partial in-network sums transmitted by its neighbors in its local estimation problem, increasing the available degrees of freedom and accelerating convergence with respect to TI-DANSE. Additionally, a tree-pruning strategy is proposed to further increase convergence speed. TI-DANSE+ converges as fast as the DANSE algorithm in fully connected WASNs while reducing transmit power usage. The convergence properties of TI-DANSE+ are demonstrated in numerical simulations.
Hearable devices, equipped with one or more microphones, can be used to capture the user’s own voice in noisy environments. In such environments, an own voice reconstruction (OVR) system is needed to enhance the quality and intelligibility of the recorded own voice. In this work, we aim to estimate clean broadband speech from a microphone at the outer face of the hearable and an in-ear microphone, which captures the own voice at a higher signal-to-noise ratio than the outer microphone, but with a limited bandwidth and additive body-produced noise. Training a supervised deep learning-based OVR system requires a substantial amount of own voice signals as training data. Such training data can be collected by recording many utterances from different talkers wearing the hearable, which is costly, or generated by augmenting existing clean speech datasets. In this paper, we investigate several data augmentation techniques to simulate a large amount of in-ear own voice signals from a limited amount of recorded own voice signals. More specifically, we consider different models for the own voice transfer characteristics between the outer microphone and the in-ear microphone, ranging from a fixed talker-averaged relative transfer function to a phoneme-dependent individual model. We investigate the influence of the amount of recorded own voice signals on the performance of an OVR system based on the FT-JNF architecture, either by directly using the recorded signals for training or by using the recorded signals to generate augmented data for training (with and without fine-tuning with recorded signals). Experimental results show that training using the proposed speech-dependent individual data augmentation technique and additional fine-tuning with recorded signals yields the best performance in terms of objective metrics, even when only few recorded own voice signals are available.
To improve speech quality and intelligibility in environments with noise and interfering sounds, binaural speech enhancement algorithms use the microphone signals from both the left and the right hearing device to generate an enhanced output signal for each ear. As a multi-frame extension of the binaural multi-channel Wiener filter, in this paper we consider the binaural spatio-temporal Wiener filter (STWF) in the short-time Fourier transform domain, which requires estimates of the highly time-varying spatio-temporal correlations of the speech and interference components. To this end, the binaural STWF is embedded into an end-to-end supervised learning framework, where temporal convolutional networks estimate the required quantities, i.e., the inverse spatio-temporal correlation matrices of the interference component and the spatio-temporal correlation vectors and power spectral densities of the speech components. In this paper, we impose spatio-temporal correlation structure on these quantities and relate them between the left and the right hearing device, aiming to reduce computational complexity while maintaining speech enhancement and interaural cue preservation performance. Assuming that the spatial correlation of the speech component is stationary over a small number of frames, we propose to decompose the spatio-temporal correlation vectors as the Kronecker product of a relative transfer function vector and a temporal correlation vector, either considering a global reference microphone or a reference microphone for each hearing device. In addition, we consider a deep bilateral STWF by neglecting the spatio-temporal correlations of the speech and interference components between both devices. The imposed spatio-temporal correlation structures greatly differ in the number of parameters that need to be estimated. The performance of causal versions of the deep binaural and bilateral STWF algorithms is evaluated based on both simulated and measured binaural room impulse responses (BRIRs) as well as diverse speech and noise sources. The simulation results demonstrate that the proposed spatio-temporal correlation structures significantly reduce the computational complexity of the binaural STWF while yielding a similar speech enhancement and interaural cue preservation performance compared to not imposing any spatio-temporal correlation structure. Furthermore, the results confirm that the deep binaural STWF outperforms the binaural Conv-TasNet algorithm as well as an algorithm that directly estimates the binaural multi-frame filter coefficients, while approaching the performance of the non-causal binaural complex convolutional transformer network (BCCTN) algorithm.
Recent advances in active noise control have enabled the development of hearables with spatial selectivity, which actively suppress undesired noise while preserving desired sound from specific directions. In this work, we propose an improved approach to spatially selective active noise control that incorporates acausal relative impulse responses into the optimization process, resulting in significantly improved performance over the causal design. We evaluate the system through simulations using a pair of open-fitting hearables with spatially localized speech and noise sources in an anechoic environment. Performance is evaluated in terms of speech distortion, noise reduction, and signal-to-noise ratio improvement across different delays and degrees of acausality. Results show that the proposed acausal optimization consistently outperforms the causal approach across all metrics and scenarios, as acausal filters more effectively characterize the response of the desired source.