With the advent of new sequence models like Mamba and xLSTM, several studies have shown that these models match or outperform the state-of-the-art in single-channel speech enhancement and self-supervised audio representation learning. However, prior research has demonstrated that sequence models like LSTM and Mamba tend to overfit to the training set. To address this issue, previous works have shown that adding self-attention to LSTMs substantially improves generalization performance for single-channel speech enhancement. Nevertheless, neither the concept of hybrid Mamba and time-frequency attention models nor their generalization performance have been explored for speech enhancement. In this paper, we propose a novel hybrid architecture, MambAttention, which combines Mamba and shared time- and frequency-multi-head attention modules for generalizable single-channel speech enhancement. To train our model, we introduce VB-DemandEx, a dataset inspired by VoiceBank+Demand but with more challenging noise types and lower signal-to-noise ratios. Trained on VB-DemandEx, MambAttention significantly outperforms existing state-of-the-art discriminative LSTM-, xLSTM-, Mamba-, and Conformer-based systems of similar complexity across all reported metrics on two out-of-domain datasets: DNS 2020 without reverberation and EARS-WHAM_v2. MambAttention also matches or outperforms generative models such as diffusion models in generalization performance while being competitive with language model baselines. Ablation studies highlight the importance of weight sharing between time- and frequency-multi-head attention modules for generalization performance. Finally, we explore integrating the shared time- and frequency-multi-head attention modules with LSTM and xLSTM, which yields a notable performance improvement on the out-of-domain datasets. However, MambAttention remains superior for cross-corpus generalization across all reported evaluation metrics.
This study addresses potential challenges in evaluating reproduction systems with complex, spatially dynamic audio material because current standardized methods may lead to biased or unreliable results. To investigate this, a listening experiment was conducted comparing two assessment methods-continuous and overall evaluation-applied to two attributes, basic audio quality and surrounding, using spatially dynamic content reproduced on two reproduction systems (a stereo and a 3D surround configuration). To enable comparison with overall evaluations, continuous ratings were summarized using different strategies. Although ratings were mainly influenced by the reproduction system, spatial variation, and program item, the choice of evaluation method also had a significant effect on the scores. This effect varied depending on the attribute and the metric used to summarize the continuous data. In cases where differences were observed, continuous evaluations consistently produced higher scores than overall ratings, regardless of the summary metric. These findings indicate that continuous evaluation can capture perceptual variations over time that are lost in overall ratings, suggesting it can be a useful approach when assessing attributes influenced by spatially dynamic changes.
Sound field estimation methods based on kernel ridge regression have proven effective, allowing for strict enforcement of physical properties, in addition to the inclusion of prior knowledge such as directionality of the sound field. These methods have been formulated for single-frequency sound fields, restricting the types of data and prior knowledge that can be used. In this paper, the kernel ridge regression approach is generalized to consider discrete-time sound fields. The proposed method provides time-domain sound field estimates that can be computed in closed form, are guaranteed to be physically realizable, and for which time-domain properties of the sound fields can be exploited to improve estimation performance. Exploiting prior information on the time-domain behaviour of room impulse responses, the estimation performance of the proposed method is shown to be improved using a time-domain data weighting, demonstrating the usefulness of the proposed approach. It is further shown using both simulated and real data that the time-domain data weighting can be combined with a directional weighting, exploiting prior knowledge of both spatial and temporal properties of the room impulse responses. The theoretical framework of the proposed method enables solving a broader class of sound field estimation problems using kernel ridge regression where it would be required to consider the time-domain response rather than the frequency-domain response of each frequency separately.
Traditionally, hearing-aid speech enhancement (SE) algorithms rely on input-based feature estimation, often derived by a voice activity detector (VAD), to configure beamformers. Yet features extracted from noisy microphone signals can become unreliable in challenging acoustic scenes where users most need help. We introduce a novel paradigm in which the settings of a sound processing system are determined by evaluating characteristics of its output. To demonstrate this idea, we employ an output-based system that selects among a set of minimum power distortionless response (MPDR) beamformers. Although MPDR beamformers are typically avoided due to their sensitivity to steering errors, we show that they become effective within an output-based framework. We compare the proposed system to a conventional input-based minimum variance distortionless response (MVDR) baseline. Experimental results show that the proposed system consistently outperforms the MVDR baseline, particularly at low SNRs, in terms of SNR, ESTOI and PESQ.
Recent advances in speech enhancement have shown that models combining Mamba and attention mechanisms yield superior cross-corpus generalization performance. At the same time, integrating Mamba in a U-Net structure has yielded state-of-the-art enhancement performance, while reducing both model size and computational complexity. Inspired by these insights, we propose RWSA-MambaUNet, a novel and efficient hybrid model combining Mamba and multi-head attention in a U-Net structure for improved cross-corpus performance. Resolution-wise shared attention (RWSA) refers to layerwise attention-sharing across corresponding time- and frequency resolutions. Our best-performing RWSA-MambaUNet model achieves state-of-the-art generalization performance on two out-of-domain test sets. Notably, our smallest model surpasses all baselines on the out-of-domain DNS 2020 test set in terms of PESQ, SSNR, and ESTOI, and on the out-of-domain EARS-WHAM_v2 test set in terms of SSNR, ESTOI, and SI-SDR, while using less than half the model parameters and a fraction of the FLOPs.
Room compensation aims to improve the accuracy of loudspeaker reproduction in reverberant environments. Traditional methods, however, are limited to improving only spectral (timbral) and temporal accuracy, neglecting the spatial accuracy of loudspeaker reproduction. Proposed is a method that compensates for both spectral and spatial properties of loudspeaker reproduction, by adding energy to the perceived reverberant sound field in a frequency-selective manner using a delayed secondary supporting source. This approach allows for the modification of the direct to reverberant ratio as a function of frequency, altering spatial and spectral reproduction. The proposed method is perceptually evaluated, demonstrating its ability to alter the perception of a primary loudspeaker without the listener perceiving the supporting source. The results show that the proposed method performs comparably to a well-established commercial room compensation algorithm and has several advantages over traditional room compensation methods.
We study compression strategies for multipartite entanglement distribution under uncertainty in the partitioning of the quantum state. When the partition is not known at the time of state preparation, we show that a joint design of the resource state and a family of compression schemes can increase the entanglement across partitions under a fixed transmission budget. We formulate this as a source coding problem and derive non-asymptotic upper and lower bounds on the achievable average entanglement subject to an average coding rate. We furthermore design an efficient method for jointly optimizing states and lossless compression maps by exploiting the inherent symmetry of weighted Dicke states. In the bipartite case, we propose practical constructions that closely approach the derived upper bound, and more generally we provide practical constructions for multipartite settings.
Sound field estimation with moving microphones can increase flexibility, decrease measurement time, and reduce equipment constraints compared to using stationary microphones. In this paper a sound field estimation method based on kernel ridge regression (KRR) is proposed for moving microphones. The proposed KRR method is constructed using a discrete time continuous space sound field model based on the discrete Fourier transform and the Herglotz wave function. The proposed method allows for the inclusion of prior knowledge as a regularization penalty, similar to kernel-based methods with stationary microphones, which is novel for moving microphones. Using a directional weighting for the proposed method, the sound field estimates are improved, which is demonstrated on both simulated and real data. Due to the high computational cost of sound field estimation with moving microphones, an approximate KRR method is proposed, using random Fourier features (RFF) to approximate the kernel. The RFF method is shown to decrease computational cost while obtaining less accurate estimates compared to KRR, providing a trade-off between cost and performance.
We systematically investigate neural speech enhancement systems, ranging from very small (∼10 k parameters) to medium-large (∼2-5 M parameters), which specialize to acoustic conditions using contextual information such as speaker identity, noise type, speaker gender, spoken language, and SNR. By fine-tuning generalist models on specific data subsets, we find that specializing to a speaker's identity consistently yields the largest gains in estimated speech intelligibility and quality. In contrast, specializing to SNR, noise type, or gender offers only marginal benefits. Crucially, we show that a small model specialized to both a specific speaker and a specific noise type can match or exceed the performance of a generalist model ten times its size. Further, cross-lingual tests reveal that models specialized to a target language outperform multilingual generalists, suggesting that language is a salient feature for specialization. These findings highlight the potential of small, adaptive models for resource-constrained applications like hearing aids, which specialize on-the-fly to contextual information.
Sentence-level syntactic structure, or syntax, refers to the principles by which grammatically correct sentences are formed. In this paper, we investigate the impact of the absence of syntax on human auditory recognition performance. Specifically, we analyse the word recognition accuracy of human participants listening to short segments of Danish speech-in-noise, with and without syntax. In sufficiently noise-free conditions, information about a spoken word is transmitted to the brain, allowing near-perfect decoding, whereas noise induces information loss. To quantify the information loss, we compute the difference in mutual information between the transmitted and decoded information for speech material with and without syntax, leveraging the data processing inequality applied to a Markov chain representing the communication model comprising the speaker, the noisy channel, and the listener. Our results indicate that the absence of syntax accounts for a loss of up to approximately one-third of the total transmitted information, elucidating the importance of syntax in facilitating effective speech processing.
This paper addresses the near-end listening enhancement (NELE) problem, i.e., how to modify a clean speech signal prior to presentation in noise to improve perceptual aspects, particularly speech intelligibility. We introduce a general optimization problem that generalizes and extends those used in our previous NELE studies. The problem is concave and we derive its closed-form solution. Within this framework, we first revisit our previously proposed OptFractASII method, which redistributes speech energy across frequency to optimize a modified approximation of the Speech Intelligibility Index. We then derive a novel NELE method, TemporalASII, which reallocates energy across time frames to boost weak speech components. We propose to combine these methods into a cascaded system, ftASII, which redistributes speech energy across both time and frequency. Extensive objective speech intelligibility predictions and a subjective listening test show that ftASII outperforms state-of-the-art NELE methods. In particular, ftASII achieves intelligibility improvements from 5.4 dB to 10.0 dB in challenging real-world noise conditions.
Sound zone techniques allow processing audio signals to control a set of loudspeakers and playback independent audio content in specific areas in a room, typically sampled through a microphone array. This task comprises two processes: acquiring room impulse responses (RIRs) between all loudspeakers and microphones and, based on these, calculating the control filters. Recent adaptive filtering methods allow performing both processes simultaneously, resulting in sound zones able to adapt to changes in the system. However, existing sample-based implementations, processing one input sample per iteration, are computationally very expensive. Alternatively, a block-based implementation is proposed, which, processing several samples per iteration, allows reducing the computational demands of such dynamic sound zones. Further reductions are achieved by truncating the RIR estimates and using less computationally demanding adaptive filter update algorithms. With respect to sample-based approaches, the proposed block-based processing can reduce the computational complexity by more than 90%, in that case, at the expense of increasing the time required to reach a certain acoustic performance. Furthermore, the proposed block-based scheme successfully adapts to changes, and inaccurate RIR estimations do not hinder sound zones rendering. The method was experimentally validated, and further reductions in the computational complexity can be made through frequency domain implementations.
This study examines how the signal-to-noise-interference ratio (SNIR) influences auditory performance and neural responses associated with listening effort (LE). A new dataset was collected from individuals with moderate hearing loss, all fitted with hearing aids (HAs). Participants listened to two competing audiobooks presented via front-facing loudspeakers, while 16-talker babble noise was delivered from background speakers. Six SNIR levels (5.47, 3.55, 2.13, 1.19, 0.64, and 0.27 dB) were tested. Participants were instructed to attend to one audiobook while ignoring the competing speech and background noise and were subsequently assessed on content of the attended speech and perceived LE. The performance results revealed a significant linear effect of SNIR on subjective ratings of LE and a primarily quadratic effect on comprehension questionnaire accuracy, suggesting that perceived effort decreases steadily with improving SNIR, while comprehension questionnaire performance exhibits a plateau at higher SNIR levels. The EEG analyses demonstrated a significant relationship between SNIR and local connectivity, specifically in the parietal electrodes and in the alpha frequency band. Further analysis confirmed that parietal local connectivity correlates linearly with subjective LE ratings. Moreover, spectral power analysis showed that parietal alpha power is not significantly related to SNIR, indicating that local connectivity may serve as a more sensitive neural marker. While local connectivity and alpha power may share some neural underpinnings, they offer complementary, yet non-identical insights. These findings highlight the potential of local EEG connectivity as a reliable estimate of LE in acoustically challenging environments.
Superconducting magnet impedance measurements are vital for assessing magnet health and electrical integrity. To further enhance monitoring capabilities beyond contemporary methods that largely rely on manual intervention, recent efforts have focused on enabling in situ and continuous measurements during magnet operation. This evolution is becoming increasingly relevant due to the growing complexity and aging of modern particle accelerator facilities. However, implementing such measurements presents challenges, particularly due to operational constraints and interference from the power converter. This article focuses on noise reduction techniques aimed at reducing the variance of impedance estimates derived from samples collected by a differential probing measurement system. A key contribution is the analysis of a unique, high-resolution dataset comprising uninterrupted impedance measurements across both steady-state magnet current plateaus and current ramping stages. This dataset enables inspection of the interaction between injected stimuli and power converter noise throughout key stages of a magnet's powering cycle, an aspect not previously explored and reported in the literature. Using a differential measurement configuration, we extract a reference of the power converter noise and apply Wiener filtering to reduce the variance of impedance estimates. We evaluate two denoising strategies, a static approach with fixed filter coefficients and an adaptive method with periodically updated coefficients. For long estimation windows (1 s), neither approach yields significant improvements. However, for short windows (10 ms), both methods achieve substantial variance reduction of up to two orders of magnitude. Under certain operating conditions, the adaptive method provides a further improvement of approximately one order of magnitude over the static approach, highlighting the potential advantage of adaptivity for real-time impedance monitoring.
We propose a notion of compression-aware entanglement efficiency that accounts for the communication cost associated with distributing entangled qubits. Our analysis focuses on a class of pure bipartite quantum states constructed as amplitude-weighted superpositions of Dicke states. By optimizing the amplitude distribution and the choice of Hamming weights, we demonstrate that it is possible to maximize the ratio of entanglement entropy to the number of qubits that must be transmitted. This approach offers a principled way to quantify and enhance the efficiency of entanglement distribution in quantum communication protocols. The proposed symmetric states are highly compressible, and for $n=2,3,4$, and 6 qubits, they achieve the optimal entanglement efficiency rates.
This study explores the relationship between auditory attention decoding (AAD), multivariate phase synchrony, and energy (alpha power) in a scenario with competing speech and music, where subjects were asked to attend to a target source in the presence of a distractor. We use an end-to-end deep learning pipeline to extract AAD scores for the target and distracting audio sources and use the circular omega complexity (COC) index, (a recently proposed correlate of listening effort), to extract multivariate phase synchrony from multi-channel electroencephalogram (EEG) signals. The audio sources consist of speech and music combined in four different scenarios for the target and distractor: speech vs. speech, speech vs. music, music vs. speech, and music vs. music. Our results indicate that as the AAD performance decreases, both the COC index and alpha power increase, suggesting increased recruitment of brain regions in conditions characterized by low speech tracking. We use linear mixed models (LMMs) to test the significance of the effect that AAD scores have on COC and alpha power. Our results indicate a significant relationship between target AAD score and COC as well as alpha power, while no significance is seen between distractor AAD score and COC or alpha power.
Using a high signal-to-noise ratio remote microphone (RM) with hearing aids (HAs) is advantageous for HA users. However, the benefit depends significantly on the properties of the wireless channel. While existing literature often assumes an error-free and instantaneous wireless transmission channel, in reality, wireless transmission of audio typically undergoes latency, which introduces a time difference of arrival (TDOA) between the local HA microphone signals and the wirelessly transmitted RM signal received at the HA. We first observe that, as the distance between the HAs and the RM increases, the TDOAs decrease, making the signals received from a distant RM more beneficial than those from a nearby RM. Next, we propose two methods to combine a RM signal with the HA, under realistic TDOAs. Acting in either the complex spectral domain or the magnitude spectral domain, the methods make a binary selection in each time-frequency component between minimum mean-square-error estimates determined from the HA microphone signals and RM signal, respectively. We evaluate the performance of the resulting estimate in terms of their signal quality and speech intelligibility using objective and subjective tests. We show that the proposed methods outperform their baselines that use only HA signals, for TDOAs in the range of 30-60 ms.
While attention-based architectures, such as Conformers, excel in speech enhancement, they face challenges such as scalability with respect to input sequence length. In contrast, the recently proposed Extended Long Short-Term Memory (xLSTM) architecture offers linear scalability. However, xLSTM-based models remain unexplored for speech enhancement. This paper introduces xLSTM-SENet, the first xLSTM-based single-channel speech enhancement system. A direct comparative analysis reveals that xLSTM-and notably, even LSTM-can match or outperform state-of-the-art Mamba- and Conformer-based systems across various model sizes in speech enhancement on the VoiceBank+Demand dataset. Through ablation studies, we identify key architectural design choices such as exponential gating and bidirectionality contributing to its effectiveness. Our best xLSTM-based model, xLSTM-SENet2, outperforms state-of-the-art Mamba- and Conformer-based systems of similar complexity on the Voicebank+DEMAND dataset.
The data acquired at different scalp EEG electrodes when human subjects are exposed to speech stimuli are highly redundant. The redundancy is partly due to volume conduction effects and partly due to localized regions of the brain synchronizing their activity in response to the stimuli. In a competing talker scenario, we use a recent measure of directed redundancy to assess the amount of redundant information that is causally conveyed from the attended stimuli to the left temporal region of the brain. We observe that for the attended stimuli, the transfer entropy as well as the directed redundancy is proportional to the correlation between the speech stimuli and the reconstructed signal from the EEG signals. This demonstrates that both the rate as well as the rate-redundancy are inversely proportional to the distortion in neural speech tracking. Thus, a greater rate indicates a greater redundancy between the electrode signals, and a greater correlation between the reconstructed signal and the attended stimuli. A similar relationship is not observed for the distracting stimuli.
Hearing aid (HA) users often experience increased listening effort, particularly in noisy environments. While noise reduction (NR) algorithms aim to alleviate this, traditional electroencephalography (EEG) methods based on power analysis have limited success in assessing the listening effort in this population. This study proposes a novel method using a whole-head synchronization map analysis that uses local connectivity, a measure of statistical dependencies within localized brain regions. We use EEG electrodes to define a region based on the surrounding electrodes in the first-order neighborhood. This approach was tested using EEG data from 22 HA users with active or inactive NR engaged in a continuous speech-in-noise (SiN) task at low (3dB) and high (8dB) signal-to-noise ratio (SNR) levels. Whole-head synchronization was quantified using circular omega complexity (COC), a multivariate phase synchrony measure. Results showed increased local connectivity in the alpha band (8-12 Hz) within frontal and occipital regions during SiN condition compared to the background noise-only (NO) condition. Furthermore, NR activation impacted the synchronization map differently at the two SNRs of the experiment, with greater effect observed at low SNR, primarily in the left parietal region and alpha band. This behavior is in line with that of existing measures for listening effort, and therefore suggests that EEG local connectivity analysis holds promise as a tool for objectively assessing listening effort in HA users, especially in challenging listening environments.