This paper introduces the topology-independent distributed multichannel Wiener filter (TI-dMWF), a novel algorithm for distributed node-specific signal estimation in wireless acoustic sensor networks (WASNs) with unconstrained topologies. The TI-dMWF enables each node in the network to compute its centralized multichannel Wiener filter solution by exchanging only low-dimensional fused signals, without requiring iterative estimation, unlike state-of-the-art approaches such as the topology-independent distributed adaptive node-specific signal estimation (TI-DANSE) algorithm. The TI-dMWF is proven optimal when each source is observed by either all nodes or only one node. Theoretical analysis and numerical simulations confirm that it achieves centralized estimation performance in a single run. Its latency as a function of the pruned-tree depth and its computational complexity are also analyzed. Its robustness is assessed in reverberant-room simulations under estimated second-order statistics, various network topologies, and deviations from the assumed observability model.
Cell-free massive multi-input-multi-output (CFmMIMO) communication networks aim to provide uniform quality of service by distributing access points (APs) across a coverage area. In user-centric variants, each user equipment (UE) can choose a cluster of APs with the best channel conditions (e.g., the closest APs) for accessing service. This approach eliminates the notion of cells with dedicated regions and APs, as found in cellular mMIMO communication networks. Estimating uplink channels between UEs and APs is a crucial step in CFmMIMO communication networks; however, existing channel estimation (CE) approaches typically originate from mMIMO systems without considering the unique properties of CFmMIMO communication networks. For instance, shorter AP-UE distances in CFmMIMO systems result in Rician channel models with prominent line of sight (LoS) components between APs and UEs, motivating cooperation between APs for improved performance. In this paper, we propose a cooperative minimum-mean-squared-error (MMSE)-based uplink CE approach where APs share their linearly compressed signals as fused signals with other APs in the same cluster. The proposed approach is optimal, i.e., its performance is equivalent to that of the centralized CE approach, where APs share their uncompressed raw signals. Notably, this optimality is achieved in one shot; that is, given the required correlation matrices, the optimal fusion filters and estimators are derived non-iteratively. Consequently, the proposed approach guarantees lower communication overhead for cooperative CE compared to the centralized approach. Numerical experiments corroborate the superior performance of the proposed cooperative CE approaches in terms of CE accuracy and convergence rate.
Sound field estimation methods based on kernel ridge regression have proven effective, allowing for strict enforcement of physical properties, in addition to the inclusion of prior knowledge such as directionality of the sound field. These methods have been formulated for single-frequency sound fields, restricting the types of data and prior knowledge that can be used. In this paper, the kernel ridge regression approach is generalized to consider discrete-time sound fields. The proposed method provides time-domain sound field estimates that can be computed in closed form, are guaranteed to be physically realizable, and for which time-domain properties of the sound fields can be exploited to improve estimation performance. Exploiting prior information on the time-domain behaviour of room impulse responses, the estimation performance of the proposed method is shown to be improved using a time-domain data weighting, demonstrating the usefulness of the proposed approach. It is further shown using both simulated and real data that the time-domain data weighting can be combined with a directional weighting, exploiting prior knowledge of both spatial and temporal properties of the room impulse responses. The theoretical framework of the proposed method enables solving a broader class of sound field estimation problems using kernel ridge regression where it would be required to consider the time-domain response rather than the frequency-domain response of each frequency separately.
In many speech recording applications, noise and acoustic echo corrupt the desired speech. Consequently, combined noise reduction (NR) and acoustic echo cancellation (AEC) is required. Generally, a cascade approach is followed, i.e., the AEC and NR are designed in isolation by selecting a separate signal model, separate cost function, and separate solution strategy. The AEC and NR are then cascaded one after the other, not accounting for their interaction. In this paper, an integrated approach is proposed to consider this interaction in a general multi-microphone/multi-loudspeaker setup. Therefore, a single signal model of either the microphone signal vector or the extended signal vector, obtained by stacking microphone and loudspeaker signals, is selected, a single mean squared error cost function is formulated, and a common solution strategy is used. Using this microphone signal model, a multi-channel Wiener filter (MWF) is derived. Using the extended signal model, it is shown that an extended MWF (MWFext) can be derived, and several equivalent expressions can be found, which are nevertheless shown to be interpretable as cascade algorithms. Specifically, the MWFext is shown to be equivalent to algorithms where the AEC precedes the NR (AEC-NR), the NR precedes the AEC (NR-AEC), and the extended NR (NRext) precedes the AEC and post-filter (PF) (NRext-AEC-PF). Under rank-deficiency conditions the MWFext is non-unique. Equivalence then amounts to the expressions being specific, not necessarily minimum-norm solutions, for this MWFext. The practical performances differ due to non-stationarities and imperfect correlation matrix estimation, with the AEC-NR and NRext-AEC-PF attaining best overall performance.
This paper focuses on distributed node-specific signal estimation in topology-unconstrained wireless acoustic sensor networks (WASNs) where sensor nodes only transmit fused versions of their local sensor signals. For this task, the topology-independent (TI) distributed adaptive node-specific signal estimation (DANSE) algorithm (TI-DANSE) has previously been proposed. It converges towards the centralized signal estimation solution in non-fully connected and time-varying network topologies. However, the applicability of TI-DANSE in real-world scenarios is limited due to its slow convergence. The latter results from the fact that, in TI-DANSE, nodes only have access to the in-network sum of all fused signals in the WASN. We address this low convergence speed issue by introducing an improved TI-DANSE algorithm, referred to as TI-DANSE$<^>+$. The TI-DANSE$<^>+$ algorithm outperforms TI-DANSE in terms of convergence speed by letting the updating node use each partial in-network sum of fused signals (coming from its neighbors) separately, when updating its estimation parameters. In this way, the number of available degrees of freedom in the optimization problem at the updating node is increased, leading to faster convergence. This separate use of incoming partial in-network sums is further exploited by combining TI-DANSE$<^>+$ with a tree-pruning strategy that maximizes the number of neighbors at the updating node. In fully connected WASNs, it is observed that TI-DANSE$<^>+$ converges as fast as the original DANSE algorithm (the latter only defined for fully connected WASNs) while using peer-to-peer data transmission instead of broadcasting and thus saving communication bandwidth. If link failures occur, the convergence of TI-DANSE$<^>+$ towards the centralized solution is preserved without any change in its formulation. Altogether, the proposed TI-DANSE$<^>+$ algorithm can be viewed as an all-round alternative to DANSE and TI-DANSE which (i) merges the advantages of both, (ii) reconciliates their differences into a single formulation, and (iii) shows advantages of its own in terms of communication bandwidth usage. The convergence properties and signal estimation performance of TI-DANSE$<^>+$ are demonstrated through speech enhancement experiments in simulated topology-unconstrained WASNs.
In audio signal processing applications with a microphone and a loudspeaker within the same acoustic environment, the loudspeaker signals can feed back into the microphone, thereby creating a closed-loop system that potentially leads to system instability. To remove this acoustic coupling, prediction error method (PEM) feedback cancellation algorithms aim to identify the feedback path between the loudspeaker and the microphone by assuming that the input signal can be modelled by means of an autoregressive (AR) model. It has previously been shown that this PEM framework and resulting algorithms can identify the feedback path correctly in cases where the forward path from microphone to loudspeaker is sufficiently time-varying or non-linear, or when the forward path delay equals or exceeds the order of the AR model. In this paper, it is shown that this delay-based condition can be generalised for one particular PEM-based algorithm, the so-called two-channel adaptive feedback canceller (2ch-AFC), to an invertibility-based condition, for which it is shown that identifiability can be achieved when the order of the forward path feedforward filter exceeds the order of the AR model. Additionally, the condition number of inversion of the correlation matrix as used in the 2ch-AFC algorithm can serve as a measure for monitoring the identifiability.
In a wireless acoustic sensor network (WASN), devices (i.e., nodes) can collaborate through distributed algorithms to collectively perform audio signal processing tasks. This paper focuses on the distributed estimation of node-specific desired speech signals using network-wide Wiener filtering. The objective is to match the performance of a centralized system that would have access to all microphone signals, while reducing the communication bandwidth usage of the algorithm. Existing solutions, such as the distributed adaptive node-specific signal estimation (DANSE) algorithm, converge towards the multichannel Wiener filter (MWF) which solves a centralized linear minimum mean square error (LMMSE) signal estimation problem. However, they do so iteratively, which can be slow and impractical. Many solutions also assume that all nodes observe the same set of sources of interest, which is often not the case in practice. To overcome these limitations, we propose the distributed multichannel Wiener filter (dMWF) for fully connected WASNs. The dMWF is non-iterative and optimal even when nodes observe different sets of sources. In this algorithm, nodes exchange neighbor-pair-specific, low-dimensional (fused) signals estimating the contribution of sources observed by both nodes in the pair. We formally prove the optimality of dMWF and demonstrate its performance in simulated speech enhancement experiments. The proposed algorithm is shown to outperform DANSE in terms of objective metrics after short operation times, highlighting the benefit of its iterationless design.
This paper addresses the analytic Procrustes problem, which aims to find the best least-squares paraunitary approximation of a square matrix of analytic transfer functions, or the best paraunitary transformation between two rectangular analytic matrices. This is accomplished by generalising the Procrustes solution from ordinary matrices to the case of matrices of analytic functions via their analytic singular value decomposition (SVD). Different from the ordinary matrix case, the analytic SVD is not restricted to singular values being nonnegative. In the case that singular values do not possess any zero crossings, we can find an analytic paraunitary matrix analogously to the standard Procrustes approach. In the case that singular values exhibit any zero crossings, the solution does not only depend on the left- and right-singular vectors, but also on a discontinuous and hence non-analytic switching function that forces those analytic singular values to become nonnegative real. We show that a close approximation of this switching function can be achieved via a complex-valued allpass filter, for which we suggest a new suitable design to minimise the overall least squares error of the fit. In addition, we propose a DFT domain algorithm to approximate this polynomial Procrustes solution, which avoids ambiguities in the analytic SVD, and possesses proven convergence. Generally, this solution requires a delay for causality, and this delay grows with the approximation order. Examples and simulations demonstrate our proposed method.
Cell-free massive-multiple-input-multiple-output (CFmMIMO) is a key enabler for sixth-generation (6G) wireless communication networks, where distributed access points (APs) jointly serve user equipments (UEs). In commonly adopted channel models for CFmMIMO networks, inter-AP channel correlation is assumed to be absent, thereby eliminating the potential benefits of centralized processing. However, by carefully designing the pilot transmission phase, the AP received signals during pilot transmission can become correlated, and thus, centralization can improve channel estimation performance, despite the absence of inter-AP channel correlation. In this paper, we propose a channel estimation scheme, termed master-assisted channel estimation (MACE), that aims to leverage inter-AP signal correlation by means of partially centralized processing and hence improve channel estimation performance. In MACE, a subset of APs fuse and forward their received pilot signals to a master AP, which then performs channel estimation using the fused signals together with its locally received signals. This scheme strikes a balance between local and fully centralized processing by leveraging inter-AP signal correlation, while reducing fronthaul signaling and computational complexity. Numerical experiments demonstrate that MACE consistently outperforms local channel estimation, where inter-AP signal correlation is neglected.
Cochlear implants (CIs) restore hearing in individuals with severe sensorineural hearing loss. In recent years, electrically evoked auditory steady-state responses (EASSRs) to amplitude modulated (AM) signals have been studied as an objective measure. EASSRs can be objectively detected in electroencephalography (EEG) recordings at the modulation frequency using statistical tests. However, the presence of electrical stimulation artifacts from the CI itself hinders the EASSR detection. Whereas previous research has focused on an experimental characterization of these artifacts, this study presents a theoretical analysis of the stimulation signal together with an experimental analysis of the resulting artifacts to characterize their properties, origins and the effects of system nonlinearities. A stimulation signal model is presented and analyzed. The effects of pulse asymmetry and nonlinearity are examined. The theoretical statements are experimentally validated using an experimental setup containing a head phantom. The analysis shows that the stimulation artifact at the modulation frequency is inherent to the stimulation signal, even in the absence of system nonlinearities. Moreover, when the pulse asymmetry is taken into account, second and higher order polynomial nonlinearities are found to contribute negligibly to the spectral component at the modulation frequency. The experimental analyses indicate the proposed signal model is a more accurate model for the stimulation signal and the resulting stimulation artifact at the modulation frequency. The model may form an important step in determining artifact contamination in EEG recordings of EASSRs and other envelope-following responses in CI recipients, enabling improved response detection.
In public address systems and hearing aids, the maximally achievable amplification or gain is limited by acoustic feedback. Therefore, in order to be able to apply a higher gain, feedback cancellation methods are required. In addition, it is oftentimes also desirable to dereverberate a recorded signal, that is, remove the late reverberation component of the signal, before playing it back. In this paper, it is shown that under two mild conditions, the acoustic feedback signal can be written as a reverberant version of the source signal. Therefore, it is possible to treat the joint dereverberation and acoustic feedback cancellation problem as a dereverberation-only problem, meaning that dereverberation algorithms can be applied to the joint problem. Simulations corroborate this finding
Sound field estimation with moving microphones can increase flexibility, decrease measurement time, and reduce equipment constraints compared to using stationary microphones. In this paper a sound field estimation method based on kernel ridge regression (KRR) is proposed for moving microphones. The proposed KRR method is constructed using a discrete time continuous space sound field model based on the discrete Fourier transform and the Herglotz wave function. The proposed method allows for the inclusion of prior knowledge as a regularization penalty, similar to kernel-based methods with stationary microphones, which is novel for moving microphones. Using a directional weighting for the proposed method, the sound field estimates are improved, which is demonstrated on both simulated and real data. Due to the high computational cost of sound field estimation with moving microphones, an approximate KRR method is proposed, using random Fourier features (RFF) to approximate the kernel. The RFF method is shown to decrease computational cost while obtaining less accurate estimates compared to KRR, providing a trade-off between cost and performance.
In a cell-free massive MIMO (CFmMIMO) network with a daisy-chain fronthaul, the amount of information that each access point (AP) needs to communicate with the next AP in the chain is determined by the location of the AP in the sequential fronthaul. Therefore, we propose two sequential processing strategies to combat the adverse effect of fronthaul compression on the sum of users' spectral efficiency (SE): 1) linearly increasing fronthaul capacity allocation among APs and 2) Two-Path users' signal estimation. The two strategies show superior performance in terms of sum SE compared to the equal fronthaul capacity allocation and Single-Path sequential signal estimation.
Deep learning (DL) based resource allocation (RA) has recently gained significant attention due to its performance efficiency. However, most related studies assume an ideal case where the number of users and their utility demands, e.g., data rate constraints, are fixed, and the designed DL-based RA scheme exploits a policy trained only for these fixed parameters. Consequently, computationally complex policy retraining is required whenever these parameters change. In this paper, we introduce a DL-based resource allocator (ALCOR) that allows users to adjust their utility demands freely, such as based on their application layer requirements. ALCOR employs deep neural networks (DNNs) as the policy in a time-sharing problem. The underlying optimization algorithm iteratively optimizes the on-off status of users to satisfy their utility demands in expectation. The policy performs unconstrained RA (URA) – RA without considering user utility demands – among active users to maximize the sum utility (SU) at each time instant. Depending on the chosen URA scheme, ALCOR can perform RA in either a centralized or distributed scenario. The derived convergence analyses provide theoretical guarantees for ALCOR's convergence, and numerical experiments corroborate its effectiveness compared to meta-learning and reinforcement learning approaches.
The concurrent nature of multiuser (MU) simultaneous wireless information and power transfer (SWIPT), coupled with the complexity of orthogonal frequency division multiplexing (OFDM) and precoding, poses a challenging non-convex resource allocation problem. While conventional methods like subcarrier assignment or interference suppression can enhance tractability, they are not always optimal. Recent work has proposed leveraging hidden convexity in multicarrier systems to bypass these suboptimal methods, instead utilizing a multiple access channel (MAC)-broadcast channel (BC) duality for a near-optimal linear precoder design. However, this novel strategy relies on a linear power harvesting model, disregarding the nonlinear character of power harvesting in SWIPT networks. This paper addresses this issue by incorporating nonlinear power harvesting effects through a power harvesting model based on sigmoidal-like functions. Sigmoidal-like functions, being neither convex nor concave, typically necessitate transformation for tractability, a challenge compounded by the MAC-BC duality. We propose an alternate approach in which a parameterized class of utility functions known as α -fairness is used to generalize the SWIPT resource allocation problem and concavify the nonlinear power harvesting model. This methodology simplifies optimization and facilitates the integration of nonlinear effects across a broad spectrum of fairness values.
This paper addresses the challenge of topology-independent (TI) distributed adaptive node-specific signal estimation (DANSE) in wireless acoustic sensor networks (WASNs) where sensor nodes exchange only fused versions of their local signals. An algorithm named TI-DANSE has previously been presented to handle non-fully connected WASNs. However, its slow iterative convergence towards the optimal solution limits its applicability. To address this, we propose in this paper the TI-DANSE+ algorithm. At each iteration in TI-DANSE+, the node set to update its local parameters is allowed to exploit each individual partial in-network sums transmitted by its neighbors in its local estimation problem, increasing the available degrees of freedom and accelerating convergence with respect to TI-DANSE. Additionally, a tree-pruning strategy is proposed to further increase convergence speed. TI-DANSE+ converges as fast as the DANSE algorithm in fully connected WASNs while reducing transmit power usage. The convergence properties of TI-DANSE+ are demonstrated in numerical simulations.
Two algorithms for combined acoustic echo cancellation (AEC) and noise reduction (NR) are analysed, namely the generalised echo and interference canceller (GEIC) and the extended multichannel Wiener filter (MWFext). Previously, these algorithms have been examined for linear echo paths, and assuming access to voice activity detectors (VADs) that separately detect desired speech and echo activity. However, algorithms implementing VADs may introduce detection errors. Therefore, in this paper, the previous analyses are extended by 1) modelling general nonlinear echo paths by means of the generalised Bussgang decomposition, and 2) modelling VAD error effects in each specific algorithm, thereby also allowing to model specific VAD assumptions. It is found and verified with simulations that, generally, the MWFext achieves a higher NR performance, while the GEIC achieves a more robust AEC performance.
Objective. Electrically evoked auditory steady-state responses (EASSRs) are potential neural responses for objectively determining stimulation parameters of cochlear implants (CIs). Unfortunately, they are difficult to detect in electroencephalography (EEG) recordings due to the electrical stimulation artifacts of the CI. This study investigates a novel stimulation paradigm hypothesized to improve artifact removal efficacy via system identification (SI), and therefore to improve response detection and clinical applicability.Approach. An amplitude-modulated (AM) CI stimulation pulse train with a step-wise increase in modulation frequency is created (referred to as SWEEP stimulation). Another stimulation is created by randomly shuffling modulation frequencies of the SWEEP stimulation (referred to as Shuffled-SWEEP stimulation). AM pulse trains with fixed modulation frequency (referred to as conventional AM stimulation), which elicit EASSRs, are also created for comparison. EEG data is collected from four CI users. A supra-threshold stimulation condition is used to investigate whether the SWEEP and Shuffled-SWEEP stimulation can elicit envelope-following responses (EFRs). A sub-threshold stimulation condition allows the collection of artifact-only EEG data, which is used to compare the SI accuracy on recordings from the SWEEP and the conventional AM stimulation.Main results. In all CI users, neural responses, following the SWEEP, Shuffled-SWEEP, and conventional AM stimulation are detected after artifact removal with SI. The validation with artifact-only EEG data shows higherF1 scores when comparing recordings with SWEEP stimulation (F1 = 0.9) to recordings with conventional AM stimulation (F1 = 0.82).Significance. Being able to accurately identify the response within one EEG recording enables the development of effective, online, objective fitting protocols. The increased neural response detection sensitivity with SWEEP stimulation reduces clinical recording time on average by a factor of 2.07. Detecting EFRs following complex stimulation paradigms offers a potential advancement in the systematic assessment of the temporal envelope processing in CI users.
We want to recover paraunitary matrices under small random perturbations. The polynomial Procrustes method, based on the analytic singular value decomposition of the perturbed system, in principle solves this. For small random perturbations, where the analytic singular values are close to unity, we propose a simplified polynomial Procrustes method that exploits this property, but show that the support of the solution is generally increased compared to the perturbed matrix. We therefore embed the simplified Procrustes method into an iterative truncation scheme, which can reduce the support while ensuring that a paraunitary approximation remains within a perimeter that is equivalent to the level of perturbation.