The steered response power (SRP) method is one of the most popular approaches for acoustic source localization with microphone arrays. It is often based on simplifying acoustic assumptions, such as an omnidirectional sound source in the far field of the microphone array(s), free field propagation, and spatially uncorrelated noise. In reality, however, there are many acoustic scenarios where such assumptions are violated. This paper proposes a generalization of the conventional SRP method that allows to apply generic acoustic models for localization with arbitrary microphone constellations. These models may consider, for instance, level differences in distributed microphones, the directivity of sources and receivers, or acoustic shadowing effects. Moreover, also measured acoustic transfer functions may be applied as acoustic model. We show that the delay-and-sum beamforming of the conventional SRP is not optimal for localization with generic acoustic models. To this end, we propose a generalized SRP beamforming criterion that considers generic acoustic models and spatially correlated noise, and derive an optimal SRP beamformer. Furthermore, we propose and analyze appropriate frequency weightings. Unlike the conventional SRP, the proposed method can jointly exploit observed level and time differences between the microphone signals to infer the source location. Realistic simulations of three different microphone setups with speech under various noise conditions indicate that the proposed method can significantly reduce the mean localization error compared to the conventional SRP and, in particular, a reduction of more than 60% can be archived in noisy conditions.
In hands-free telephony and other distant-talking applications, an acoustic echo cancellation system is typically required, where a short adaptive filter is often used in practice to achieve fast convergence at low computational cost. This may result in late residual echo (LRE) remaining due to under-modeling of the echo path and early residual echo (ERE) due to filter misalignment. Both residual echo components can be suppressed using a postfilter in the subband domain, which requires accurate estimates of the power spectral density (PSD) of the ERE and LRE components. State-of-the-art methods estimate the ERE and LRE PSDs independently of each other, where the ERE PSD is estimated by simply multiplying the loudspeaker PSD with a frequency-dependent scalar and the LRE PSD is estimated using a recursive estimator based on frequency-dependent reverberation scaling and decay parameters. In this paper, we propose to extend the ERE PSD estimator from a scalar to a moving average filter on the loudspeaker PSD. In addition, we propose a signal-based method to jointly estimate all model parameters for the ERE and LRE PSD estimators in online mode, and derive two gradient-descent-based algorithms to simultaneously update the model parameters by minimizing the mean squared log error. The proposed method is compared with state-of-the-art methods in terms of estimation accuracy of the model parameters as well as the residual echo PSDs. Simulation results using both artificially generated as well as measured impulse responses show that the proposed method outperforms state-of-the-art methods for all considered scenarios.
The aim of artificial bandwidth extension is to recreate wideband speech (0 - 8 kHz) from a narrowband speech signal (0 - 4 kHz). State-of-the-art approaches use neural networks for this task. As a loss function during training, they employ the mean squared error between true and estimated wideband spectra. This, however, comes with the drawback of over-smoothing, which expresses itself in strongly underestimated dynamics of the upper frequency band. We previously proposed to tackle this problem by discriminative training, i.e., a modification of the loss function that is designed to improve the separation between fricatives and vowels. Other authors instead took a generative adversarial network (GAN) approach. This was motivated by the fact that GANs demonstrated big reductions of over-smoothing in speech synthesis. In this work, we combine these two approaches. In particular, we show that conditional GANs improve the speech quality by a CMOS score of 0.28 compared to GANs while the combined approach yields an improvement of 0.84.
In hands-free telephony and other distant-talk applications, often a short AEC filter is used to achieve fast convergence at low computational cost. As a result, a significant amount of late residual echo (LRE) may remain, especially in highly reverberant environments. This LRE can be suppressed using a postfilter in the subband domain, which requires an estimate of the power spectral density (PSD) of the LRE. To estimate the LRE PSD, an exponentially decaying model with frequency-dependent reverberation scaling and decay parameters has frequently been assumed. State-of-the-art methods estimate both reverberation parameters independently of each other, either in offline or in online mode. In this article, we propose two signal-based methods (i.e. output error and equation error) to jointly estimate both reverberation parameters in online mode. The estimated parameters are then used to generate an estimate for the LRE PSD, which is fed into a postfilter for the purpose of late residual echo suppression. We derive several gradient-descent-based algorithms to simultaneously update both reverberation parameters, minimizing either the mean squared error or the mean squared log error cost function. The proposed methods are compared with state-of-the-art methods in terms of the accuracy of the estimated reverberation parameters and the corresponding LRE PSD estimate. Extensive simulation results using both artificial as well as measured room impulse responses show that the proposed output error method with mean squared log error minimization outperforms state-of-the-art methods in all considered scenarios.
Artificial bandwidth extension reconstructs a 16 kHz wide-band signal from a given 8 kHz narrowband signal. State-of-the-art approaches use regression deep neural networks (DNNs) for extending the spectral envelope. As a cost function during training, they use the mean squared error (MSE) between true and estimated wideband spectral envelopes. With pure MSE training, the extension for fricatives and vowels is not distinctive enough compared to the true WB data. In this work, we propose to add a discriminative term to the cost function that forces the DNN to extend the energy more distinctively for different phoneme classes. The proposed cost function improves the separation of fricatives and vowels in the DNN. It also results in a higher speech quality, which was shown in subjective listening tests.
Artificial bandwidth extension (ABWE) for speech signals is still an important topic in mobile telephony, especially when a 16 kHz wideband (WB) call suddenly falls back to an 8 kHz GSM connection. The aim of ABWE is to bridge the arising voice quality gap by reconstructing the WB signal. In order to achieve this, the speech signal is typically decomposed into a spectral envelope and an excitation signal, both of which are then extended separately. While the algorithms for envelope extension are getting increasingly more sophisticated, excitation generation is still often performed with rudimentary methods such as spectral folding (SF) or spectral shifting (SS). But this can introduce audible artifacts, especially for speech signals where the pitch frequency varies a lot. To reduce these artifacts, we introduce an algorithm that shifts parts of the spectrum multiple times by a smaller frequency shift. Additionally, we investigate if the speech quality can be further improved by interpolating the extended excitation signal with white noise. This is motivated by the fact that the SNR of the harmonic excitation decreases towards higher frequencies for real WB signals. The performance of the proposed algorithm is evaluated and compared to spectral folding and spectral shifting.
Speech enhancement algorithms are employed in many applications, such as hands-free telephones, or speech recognizers, to recover a speech signal that is recorded in a noisy environment. In automotive environments, the noise particularly affects the low frequencies that are relevant for voiced speech. Detection of voiced speech sections and estimation of the pitch frequency help to reconstruct the harmonic structure of voiced speech and to enhance the speech signal. Many algorithms were introduced to detect voiced speech and to estimate the pitch. Most of them rely on a high spectral resolution that is achieved by employing long window lengths. However, some applications, such as in-car-communication (ICC) systems, have to deal with short windows in order to reduce computational costs and to ensure low system latencies. Resolving the pitch is difficult in this case. Spectral refinement techniques have been introduced to increase the spectral resolution by combining multiple consecutive low-resolution spectra. Using these techniques, standard pitch estimation algorithms can be applied even though the resolution of the original spectrum was too low. In this paper, we analyze the performance of pitch estimation using spectral refinement techniques and introduce an alternative approach that explicitly takes into account the short windows of ICC applications. Introduction Speech is an intuitive way for human communication that is employed in more and more applications. Devices, such as the car navigation system or smartphones, can be controlled conveniently via voice commands. Other applications facilitate the voice communication between humans, e.g., via hands-free telephone. In particular, in-car-communication systems amplify the driver’s voice and support the communication with passengers on the backseat. By employing these systems, conversations are possible even in noisy conditions at higher velocities [1]. Voiced speech portions, e.g., vowels are important for correct recognition of human speech. However, the background noise in automotive environments masks especially these low-frequent components. The unvoiced speech portions in higher frequencies are masked less but are also less important for recognition. Therefore, robust detection of voiced speech and estimation of the pitch frequency are important problems in speech enhancement algorithms [2]. Detection of voiced speech can be used to distinguish speech from noise, e.g., for robust noise estimation. The pitch frequency can be employed to reconstruct speech that is masked by noise. To capture the pitch information, long window lengths are required that exceed the pitch period. Some applications, however, need shorter windows in order to reduce the processing delay and the computational complexity. To overcome these contradicting requirements, techniques that approximate a long window by a combination of multiple shorter windows have been introduced in literature. In this paper, two approaches will be discussed in more detail: • Spectral refinement [3] combines multiple complexvalued spectra in order to recreate a spectrum with a higher frequency resolution. • Extended ACF [3] combines multiple crosscorrelations between short frames to approximate a longer auto-correlation function (ACF). Both techniques gain information from some previous frames in addition to the current frame. By employing this temporal context, pitch information can be extracted even for very short windows. In this contribution, the detection of harmonic components, as well as pitch estimation will be summarized. A conventional approach based on the auto-correlation function is employed. Afterwards, we will consider shorter windows and discuss the two approaches to deal with this challenge. We will briefly summarize spectral refinement and provide a more detailed description of the extended ACF. Our analyses focus on the comparison of the different approaches. In particular, the detection performance of voiced speech and the estimated pitch are assessed. Pitch Estimation using ACF First, we describe the basic principle of ACF-based pitch estimation. Based on a frame of an audio signal x̃(l) = [x(lR− Ñ + 1), · · · , x(lR−N + 1), · · · , x(lR)] , (1) the ACF is determined. Here, the number of samples Ñ that are taken into account is chosen much longer than the expected pitch periods. The shift between two succeeding frames is denoted by R and the frame index by DAGA 2017 Kiel
Detection of voiced speech and estimation of the pitch frequency are important tasks for many speech processing algorithms. Pitch information can be used, e.g., to reconstruct voiced speech corrupted by noise.In automotive environments, driving noise especially affects voiced speech portions in the lower frequencies. Pitch estimation is therefore important, e.g., for in-car-communication systems. Such systems amplify the driver's voice and allow for convenient conversations with backseat passengers. Low latency is required for this application, which requires the use of short window lengths and short frame shifts between consecutive frames. Conventional pitch estimation techniques, however. rely on long windows that exceed the pitch period of human speech. In particular, male speakers' low pitch frequencies are difficult to resolve.In this publication. we introduce a technique that approaches pitch estimation from a different perspective. The pitch information is extracted based on phase differences between multiple low-resolution spectra instead of a single long window. The technique benefits from the high temporal resolution provided by the short frame shift and is capable to deal with the low spectral resolution caused by short window lengths. Using the new approach. even very low pitch frequencies can be estimated very efficiently.
This work shows a novel method to suppress linear and nonlinear residual echo components after application of a linear echo canceler. The main idea is to separately treat linear and nonlinear residual echo components, as linear echo is reduced by AEC, while nonlinear echo passes the AEC unchanged. In particular, it is shown that a very simple model, such as a hard clipping function, is sufficient to approximate the nonlinear residual echo power; and that the clipping threshold can be estimated by comparing the broad-band predicted nonlinear residual echo power (produced with the current clipping threshold estimate) to the broad-band observed nonlinear residual echo power (obtained through linear AEC and subtraction of the linear residual echo power, as determined with linear coupling factors). Experimental evaluations show ERLE improvements by up to 14.9 dB compared to linear echo cancellation and suppression at a negligible decrease in speech quality during double talk.
When a speech application is employed in a crowded environment, the user's voice superposes with many interfering voices. This babble noise is a challenge for many speech processing algorithms since assumptions like stationarity of the noise or a good SNR may not be valid. In this contribution, characteristics of babble noise are discussed and distinctive features are summarized that allow for distinguishing the desired foreground speech from background noise. In particular, the kurtosis of a signal is identified as a good measure to detect the presence of speech in babble noise. Furthermore, a babble noise suppression system is introduced that employs a speech detector to dynamically control the aggressiveness of noise suppression. It is desirable that the residual noise after processing the signal is perceived as more pleasant by human listeners. To evaluate the improvements, a subjective listening test is conducted. Speech distortions are quantified using an objective measure.
Acoustic echo control is still challenged with finding computationally efficient and accurate methods for handling nonlinear echo components introduced by cheap loudspeaker components and high signal levels. We propose a robust echo control scheme that utilizes a common linear acoustic echo canceler as well as a combined linear and nonlinear residual echo suppression unit. Linear and nonlinear residual echo powers are estimated separately making use of the same room impulse response estimate provided by the linear canceler. Frequency-dependent coupling factors between the observed and the respective estimated residual echo power are used to account for deviations due to varying convergence states of the linear filter or the nonlinear model. A weighted sum of the two power estimates is used in a sub-band-domain suppression filter. This contribution compares the potential of the proposed concept to a Hammerstein-type nonlinear acoustic echo canceler under fair simulation conditions.
Many speech processing algorithms rely on voice activity detection (VAD) that separates speech from noise. For this task, several features have been introduced that employ different characteristic properties of speech. In this contribution, we introduce a new feature that is robust against various types of noise. By considering an alternating excitation structure of low and high frequencies, speech is detected with a high confidence. The computationally low complex feature can cope even with the limited spectral resolution that is typical for in-car-communication systems. By combining the feature with a conventional modulation feature, the performance can be improved. Our simulations confirm the robustness of the feature and show the increasing performance compared to established VAD features.
When deploying acoustic echo cancellation systems in large rooms, using short filters may result in significant amount of residual echo caused by room reverberation. In this paper, we model the late residual echo as exponentially decaying and use a parametric IIR filter to estimate its power in the subband domain for application in residual echo suppression. Working in an offline system identification setup, the problem of finding the optimal parameters of the IIR filter is addressed, with an analysis conducted on the performance of two parameter estimation methods: output error and equation error. The late residual echo power estimates obtained using the two methods are furthermore judged using the mean squared error and the mean squared log error cost functions. Results indicate that minimizing the mean squared log error for the output error method provides accurate estimates for the late residual echo power and the reverberation decay parameter.
In many speech signal processing applications, voice activity detection (VAD) plays an essential role for separating an audio stream into time intervals that contain speech activity and time intervals where speech is absent. Many features that reflect the presence of speech were introduced in literature. However, to our knowledge, no extensive comparison has been provided yet. In this article, we therefore present a structured overview of several established VAD features that target at different properties of speech. We categorize the features with respect to properties that are exploited, such as power, harmonicity, or modulation, and evaluate the performance of some dedicated features. The importance of temporal context is discussed in relation to latency restrictions imposed by different applications. Our analyses allow for selecting promising VAD features and finding a reasonable trade-off between performance and complexity.
Hans-Jörg Pfleiderer合作论文数microelectronics department in Ulm2