
This paper proposes the spherical harmonic and the cylindrical harmonic decompositions of the acoustic velocity vectors of the outgoing sound field. The spherical harmonic coefficients and the cylindrical harmonic coefficients of the pressure of the outgoing sound field can be measured by a spherical microphone array and a cylindrical microphone array surrounding the source, respectively. By using the sound field translation formula, the spherical harmonic coefficients and the cylindrical harmonic coefficients of the acoustic velocity vectors of the outgoing sound field are obtained. Simulations show the proposed spherical harmonic and cylindrical harmonic decompositions of the acoustic velocity vectors can accurately model a monopole, a dipole and a quadrupole. The proposed decompositions can be used in exterior noise control, where an array of loudspeakers controls the acoustic velocity vectors of the outgoing sound field.
In environments with competing sound sources, speech intelligibility can be significantly compromised. This paper addresses the near-end listening enhancement (NELE) problem, i.e., the problem of processing an available clean speech signal in order to maximize its intelligibility when it is subsequently presented to a human listener in an adverse acoustic situation. We propose a time-invariant and low-complexity NELE algorithm that maximizes an approximation of the Speech Intelligibility Index by redistributing speech energy across frequency bands. Unlike existing algorithms, the proposed algorithm incorporates a mechanism that allows it to distinguish between temporally fluctuating and non-fluctuating noise maskers by using only long-term speech and noise statistics. Simulation results show that the proposed method outperforms baseline algorithms, whether time-invariant or time-varying, in a wide range of noise conditions.
In this paper, we introduce a post-processing method to minimize the algorithmic latency in traditional blind source separation (BSS) techniques. Our proposed approach involves the incorporation of a minimum variance distortion response method with the spatial covariance matrix, which is derived from conventional BSS methods, to effectively compute short demixing filters. The performance of source separation can be improved either by increasing the number of microphones or by integrating a dereverberation technique as a pre-processing step, or even both. The experimental results confirm the effectiveness and consistency of the proposed approaches on diverse speech databases.
The perceptual performance of active noise control (ANC) head-phones depends on various sound attributes that collectively influence end-user satisfaction. In the speech prediction-based ANC for headphone applications, two key attributes, the amount of speech and the amount of artifacts, significantly affect subjective satisfaction with the ANC experience. This study aims to identify objective metrics correlating with the perceptual assessment of these attributes to predict subjective satisfaction with the considered ANC systems. Such prediction can significantly simplify the design and evaluation process of the ANC systems, avoiding expensive and time-consuming subjective tests. By establishing the perceptual relevance of attenuation, a commonly used objective performance metric for ANC, to the amount of speech and utilizing conventional quality metrics, i.e., PEAQ, POLQA, and PESQ, for the amount of artifacts, we demonstrate high-quality fits when predicting these attributes. Our findings indicate that these metrics effectively predict subjective satisfaction with the considered ANC systems, as evidenced by a high explained variance, with an R-squared of 0.98 and a relatively low root mean square error ranging from 0.09 to 0.19 on a five-point scale.
Hearing aids, and more generally embedded devices, have undergone significant evolution, transitioning into devices capable of significant compute on-device. However, next generation speech processing also utilises computationally expensive deep neural networks. Therefore, it is often preferable to execute the high demanding parts at a central, more powerful device. Certainly for localisation, which has less strict latency requirements. This however requires wireless data transmission between the on ear devices and the central unit, which can adversely impact battery life. In this work, we compare different strategies for bandwidth efficient deep localisation. One strategy is to send the audio signals directly to the central device, and use audio codecs, like the LC3plus codec to minimise the bandwidth. An alternative method is to adapt the co-operative localisation method to binaural hearing and investigate methods to reduce its bandwidth. The co-operative model first processes the microphone signals locally, before transmitting features to the central processor for further analysis. We investigate quantisation, time compression and lowering the dimension of these features. The cooperative model proved slightly better at high SNR scenarios, while the audio transmission model at low SNR cases.
Differential microphone arrays (DMAs) have attracted considerable attention for their high spatial gains and frequency-invariant spatial responses. However, they often face significant white noise amplification at low frequencies. One approach to mitigating this challenge is by increasing the number of microphones while fixing the order of the DMA, leveraging additional degrees of freedom to optimize the white noise gain (WNG). But this compensation of WNG can lead to beampattern distortion at mid and high frequencies. To address this issue, we recently explored an approach to designing differential beamformers with predefined WNG levels. This involves formulating beamforming as a quadratic eigenvalue problem (QEP) to efficiently derive optimal solutions without iterative processes, leading to the QEP-based differential beamformers. While it is successful in controlling WNG, this method is found to exhibit great and atypical performance degradation at lower frequencies in certain scenarios. In this paper, we illustrate this phenomenon using a two-stage structured beamformer as a case study and offer insights into why this occurs, along with proposing a solution.
This paper focuses on two key aspects: region-of-interest beamforming and optimal sparse circular sector array design. The aim is to address the problem of the unknown direction of arrival within a given region of interest while optimizing the array geometry layout and beamformer taps. This is done while meeting the maximum broadband array directivity criterion. To ensure that the desired signal is not distorted, we apply appropriate optimization constraints while maintaining a sufficiently high white noise gain. Our proposed approach outperforms a recently suggested approach in terms of the directivity factor, especially when the direction of arrival of the desired source significantly deviates from its nominal value.
Acoustic echo cancellation (AEC) in multi-device scenarios is a challenging problem due to sample rate offset (SRO) between devices. The SRO hinders the convergence of the AEC filter, diminishing its performance. To address this, we approach the multi-device AEC scenario as a multi-channel AEC problem involving a multi-channel Kalman filter, SRO estimation, and resampling of far-end signals. Experiments in a two-device scenario show that our system mitigates the divergence of the multi-channel Kalman filter in the presence of SRO for both correlated and uncorrelated playback signals during echo-only and double-talk. Additionally, for devices with correlated playback signals, an independent single-channel AEC filter is crucial to ensure fast convergence of SRO estimation.
Learning-based a posteriori speech presence probability (SPP) estimation has shown high accuracy in non-stationary noise environments. In this work, we propose using multi-channel information to estimate multi-channel SPP (MC-SPP) based on deep neural networks (DNNs), which contribute to improving multi-channel speech enhancement performance. Firstly, with the observed signal and the MC-SPP as the training data pairs, one low-parameter DNN model is trained to estimate the MC-SPP. Based on the MC-SPP estimate, the noise power density (PSD) and clean speech PSD matrices are updated recursively. With the clean speech PSD matrix, the steering vector is computed using the covariance subtraction method. Subsequently, the minimum variance distortionless response (MVDR) weight is computed with the clean and noise matrices. To further improve multi-channel speech enhancement performance, a new MVDR modification guided by the MC-SPP estimate is proposed. Finally, spatial filtering is performed by integrating the MVDR beamforming. For experiments, we spatially synthesize a real speech dataset in the isotropic noise fields for training and testing. The PESQ, STOI, and DNSMOS scores are used to evaluate speech quality. The experimental results show that, compared with a recently proposed DNN-guided approach, our proposed method provides an effective statistics estimation approach that can further improve multi-channel speech enhancement performance.
The compression of audio signals plays a crucial role in audio storage and transmission, particularly within the context of streaming media applications, where bandwidth utilization is a dominant factor of cost. Motivated by this challenge, our objective is to continuously enhance the compression rate while simultaneously ensuring the retention of audio quality. In this paper, we present a diffusion-based codec sDiff-Codec, which is a state-of-the-art, high-fidelity neural audio codec. The condition module and generator module serve the role of encoder and decoder in sDiff-Codec, the sound quality enhancement task becomes audio compression task. Additionally, we employed a hybrid quantizer to quantize the latent information using a hyper-prior model, the hyper-prior model is to generate prior auxiliary information of the entropy model. The experiment results show that sDiff-Codec is superior compared with the baseline methods under scenarios when monophonic audio signal bitrate ranges from 16 kbps to 192 kbps.
Speech processing algorithms often rely on statistical knowledge of the underlying process. Despite many years of research, however, the debate on the most appropriate statistical model for speech still continues. Speech is commonly modeled as a wide-sense stationary (WSS) process. However, the use of the WSS model for spectrally correlated processes is fundamentally wrong, as WSS implies spectral uncorrelation. In this paper, we demonstrate that voiced speech can be more accurately represented as a cyclostationary (CS) process. By employing the CS rather than the WSS model for processes that are inherently correlated across frequency, it is possible to improve the estimation of cross-power spectral densities (PSDs), source separation, and beamforming. We illustrate how the correlation between harmonic frequencies of CS processes can enhance system identification, and validate our findings using both simulated and real speech data.
Accurate sound field estimates are often cumbersome to obtain, since they generally rely on microphone measurements at several spatial positions within a room. In addition, room acoustics are often non-stationary, in which case the sound field has to be repeatedly re-estimated with new measurements. However, since the new and old measurements are recorded in different acoustic environments, it is not straight-forward to fully exploit the combined measurements. In this paper, a Bayesian approach is taken where older measured data are considered to be more uncertain than newer data. The proposed method allows for the use of data captured in different acoustic environments. For each set of measurements, the position, directivity, and number of microphones are allowed to differ. It is demonstrated on real sound field measurements that the proposed approach is effective, being able to better account for different levels of uncertainty in the data.
There have been a plethora of methods developed to tackle diverse audio reconstruction problems. Recently, deep generative models have affected this field strongly, some of them allowing to solve multiple problems with only a minimal need for adaptation. However, long inference times still represent a barrier to their real-world deployment. We propose a plug-and-play approach to audio reconstruction enabling a shorter duration of signal generation. We present our approach on a number of inverse problems, all evaluated on a piano sound dataset. Subjectively, the proposed strategy performs competitively with recent methods, however, this is rarely reflected by objective metrics.
Acoustic echo cancellation (AEC) in the frequency domain has been a de facto standard in systems for acoustic echo control, but its robustness against double-talk and its agility during rapid echo path changes all the time had to be carefully managed by hand. This paper, therefore, connects theories of model-based and data-driven optimization in order to accomplish a model-based system design with a data-driven replacement of the former hand-tuning. We can demonstrate that the trainable elements of the echo cancellation algorithm may then use very simple architectures with a pronounced minimum of trainable parameters. Experimental results are depicted using the linear subset of the ICASSP-21 AEC challenge data set.
Reverberation is one of the major causes of speech degradation. The popular weighted prediction error (WPE) technique performs dereverberation by estimating the late room reflections using a multi-channel prediction filter. However, the length of the prediction filter in each short-time-Fourier-transform (STFT) band must be sufficiently long to model the late reverberation component accurately. This leads to inverting a large matrix in every frequency bin, making the WPE method computationally expensive. The WPE method is also vulnerable to additive noise. To tackle these issues, we present a computationally efficient dereverberation technique in this work. We decompose the long prediction filter into three smaller sub-filters using third-order tensor decomposition. One sub-filter acts as a spatial filter, while the other two act as temporal prediction filters. We then develop an iterative algorithm to get optimal solutions for all three sub-filters. The spatial filter is optimized as a weighted distortionless beamformer to deal with noise, while the temporal filters are optimized as weighted Wiener filters. Since the lengths of the sub-filters are smaller, the respective covariance matrices are computationally easier to invert, leading to an efficient algorithm. Simulation results show that the proposed algorithm is robust to noise and outperforms the current WPE based algorithms in terms of dereverberation.
Feedback active noise cancellation (ANC) has become a common tool to reduce unwanted noise in hearing devices. We discuss controller design approaches, which involve the approximation of high-order prototype controllers by low-order ones for resource-constrained applications. A drawback of conventional approaches is that the error introduced by the low-order approximation is not considered in the design of the prototype controller. Our concept is to mitigate this drawback using a novel algorithm which promotes prototype controllers that can be approximated with a low error. For this, we use the well-known nuclear norm heuristic. We describe the design process and provide empirical evidence of a performance improvement in applications where a low-order controller is required.
Nonlinear acoustic echo cancellation (NAEC) is of significant importance in acoustic telecommunication. To improve NAEC performance in the double-talk case, semi-blind source separation-based NAEC (SBSS-NAEC) algorithms have been proposed. However, to deal with reverberation and loudspeaker nonlinearities, convolutive transfer function (CTF) models and power series expansions are employed, which significantly increase the number of free parameters and consequently lead to slow convergence speed and, hence, limited performance. In this paper, we introduce the data-reuse strategy, well-known in the adaptive filter literature, into an SBSS-NAEC framework and propose two algorithms: data-reuse iteration projection (DR-IP) and data-reuse element-wise iterative source steering (DR-EISS). Several simulations demonstrate the superiority of the proposed methods, especially the tracking capability when the impulse response changes.
The image source method (ISM) is often relied upon to simulate room acoustical fields due to its computational efficiency. While it is common to find implementations of the ISM for shoebox-shaped rooms, the ISM can be used to simulate more general room shapes. However, including the directivities of sources and receivers has not been well resolved for arbitrary room geometry in the literature. In this work, we extend the diffraction-enhanced image source method, which uses spherical harmonic descriptions to include directional sources and receivers, to arbitrary room geometries. The proposed extension is evaluated using finite element method simulations for different room configurations and is shown to capture many of the relevant features of the room transfer functions. In addition, discrepancies in the evaluation results indicate possible missing acoustic effects that are difficult to capture with ISM-based methods and need to be addressed.
Virtual navigation through a large listening region of interest (ROI) has broad applications. Obtaining this from limited recordings in reverberant environments for broadband source distributions poses many challenges. In this paper, we address this problem by proposing an efficient method for sparse sound field representation and reconstruction. Our method uses the Iterative Group Complex Orthogonal Matching Pursuit (IG-COMP) as a novel approach to iteratively and sparsely decompose the recorded sound field into a selected set of virtual sources (ViSrcs) and reconstruct it over the ROI. IG-COMP builds upon the Complex Orthogonal Matching Pursuit (COMP) to address broadband sources. It iteratively refines the virtual source selection while ensuring reduced computational complexity. We demonstrate the effectiveness of IG-COMP through simulations of several broadband sources distributed in a reverberant room. These simulations showcase the method’s ability to achieve high-fidelity sound field reproduction across a large ROI using a compact equivalent representation.
When multiple microphones are available, speech enhancement approaches can utilize spatial information to improve speech perception for hearing aid users. In this study, we optimize and extend an existing DNN-based speech enhancement approach to investigate the utility of head rotation (HR) information in several ways. First, optimization of the DNN indicates that especially input features with inter-channel phase and inter-channel level difference information enable improved processing of dynamic spatial cues. Second, introducing encoded HR data as an additional input to the DNN model leads to notable enhancements, particularly when target direction is included. Nevertheless, perceptible improvements are still observed when relative HR data is utilized, that is without information on the absolute target direction. Given that relative HR data can be accurately captured using angular acceleration sensors and integrated with minimal algorithmic complexity, leveraging such data remains an important objective in improving speech enhancement algorithms.