This paper presents an alternative approach to acoustic source localization which modifies the traditional two-step localization procedure to not require explicit time-delay estimates. Instead, the cross-correlation functions derived from various microphone pairs are simultaneously maximized over a set of potential delay combinations consistent with candidate source locations. The result is a procedure that combines the advantages offered by the phase transform (PHAT) weighting (or any reasonable cross-correlation-type function) and a more robust localization procedure without dramatically increasing computational load. Simulations are performed across a range of reverberation conditions to illustrate the utility of the proposed method relative to conventional generalized cross-correlation (GCC) filtering approaches and a more modern eigenvalue-based technique.
This paper presents a model-based method for the enhancement of multichannel speech acquired under reverberant conditions. A very coarse estimate of the channel responses associated with each source-microphone pair is derived directly from the received data on a short-term basis. These estimates are employed to modify the LPC residuals of the channel data in an effort to deemphasize the effects of reverberant energy in the resulting synthesized signal. The approach is robust to conditions of partial and approximate channel information. Specifically, the incorporated channel model requires only approximate times and amplitudes of the initial multipath reflections. In practice these impulses are responsible for the bulk of reverberant energy in the received speech signal and can be estimated to a sufficient degree on a time-varying basis
This paper introduces a microphone array processing method that possesses the robustness of fixed beamforming along with the ability to be dynamically reconfigured to limit interference and reverberation. The basic approach is to partition the environment into two regions: an interior region (containing sources that are physically present within the room enclosure), and an exterior region (containing virtual sources of reverberation). The interior region is further divided into cells, and standard source localization techniques are used to identify those cells containing the desired source as well as sources of interference (e.g., competing talkers). Beamforming weights are then found to pass the desired signal, while simultaneously minimizing a weighted combination of interior interference and exterior reverberation. Simulation results are presented to demonstrate the effectiveness of the proposed technique when compared with conventional beamforming methods.
This paper presents the multi-channel multi-pulse (MCMP) algorithm for the enhancement of speech degraded by reverberations and additive noise. The enhanced speech is synthesized from a sequence of impulses exciting a linear predictive filter. The excitation signal is computed from a nonlinear process which uses impulse clustering of the multi-channel speech data to discriminate portions of the linear prediction residual produced by the desired speech signal from those due to multipath effects and uncorrelated noise. The MCMP algorithm is shown to be capable of identifying and attenuating reverberant portions of the speech signal as well as reducing the effects of additive noise.
The relative time delay associated with a speech signal received at a pair of spatially separated microphones is a key component in talker localization and microphone array beamforming procedures. The traditional method for estimating this parameter utilizes the generalized cross correlation (GCC), the performance of which is compromised by the presence of room reverberations and background noise. Typically, the GCC filtering criteria used are either focused on the signal degradations due to additive noise or those due exclusively to multipath channel effects. There has been relatively little success at applying GCC weighting schemes which are robust to both of these conditions. This paper details an alternative approach which attempts to employ a signal-dependent criterion, namely, the estimated periodicity of the speech signal, to design a GCC filter appropriate for the combination of noise and multipath distortions. Simulations are performed across a range of room conditions to illustrate the utility of the proposed time-delay estimation method relative to conventional GCC filtering approaches.
This paper addresses the limitations of current approaches to distant-talker speech acquisition and advocates the development of techniques which explicitly incorporate the nature of the speech signal (e.g. statistical non-stationarity, method of production, pitch, voicing, formant structure, and source radiator model) into a multi-channel context. The goal is to combine the advantages of spatial filtering achieved through beamforming with knowledge of the desired time-series attributes. The potential utility of such an approach is demonstrated through the application of a multi-channel version of the dual excitation speech model.
Generalized cross-correlation (GCC) has been the traditional method for estimating the relative time-delay associated with speech signals received by a pair of microphones in a reverberant, noisy environment. The filtering criterion employed is either focussed on the signal degradations due to additive noise or those due exclusively to multipath channel effects. There has been relatively little success at applying GCC weighting schemes which are robust to both of these conditions. This paper details an alternative approach which attempts to employ a signal dependent criterion, namely the estimated periodicity of harmonic spectral intervals, to design a GCC filter appropriate for the combination of noise and multipath signal distortions. Simulations are performed across a range of room conditions to illustrate the utility of the proposed time-delay estimation method relative to conventional GCC filtering approaches.
A method for tracking the positional estimates of multiple talkers in the operating region of an acoustic microphone array is presented. Initial talker location estimates are provided by a time-delay-based localization algorithm. These raw estimates are spatially smoothed by a Kalman filter derived from a set of potential source motion models. Data association techniques based on the estimate clusterings and source trajectories are incorporated to match location observations with individual talkers. Experimental results are presented for array recorded data using multiple talkers in a variety of scenarios.
This paper presents a means for predicting the error region associated with a speech-source location estimate obtained from a set of microphones in a room environment. The error predictor presented is derived assuming a specific source-sensor geometry consisting of pairs of closely-spaced sensors for which a delay estimate associated with the potential source has been evaluated. The accuracy of the predictor is evaluated through a set of Monte Carlo simulations and an application of the predictor to microphone-array design in the context of a video-teleconferencing scenario is presented.
The linear intersection (LI) estimator, a closed-form method for the localization of source positions given only the sensor array time-delay estimate information, is presented. The array is constrained to be composed of 4-element sub-arrays configured in 2 centered orthogonal pairs. A bearing line in 3-space is estimated from each sub-array and potential source locations are found via closest intersection of bearing line pairs. The final location estimate is determined by a probabilistic weighting of these potential locations. The LI estimator is shown to be robust and accurate, to closely model the ML estimator, and to outperform a representative algorithm. The computational complexity of the LI estimator is suitable for use in real-time microphone-array applications
A frequency-domain-based delay estimator is described, designed specifically for speech signals in a microphone-array environment. It is shown to be capable of obtaining precision delay estimates over a wide range of signal-to-noise ratio conditions and is computationally simple enough to make it practical for real-time systems. A location algorithm based upon the delay estimator is then developed. With this algorithm it is possible to localize talker positions to a region only a few centimetres in diameter (not very different from the size of the source), and to track a moving source. Experimental results using data from a real 16-element array are presented to indicate the true performance of the algorithms.
A real-time, single digital signal processing (DSP) chip implementation of a 2.4-, 4.8-, and 8.0-kb/s improved multiband excitation (IMBE) vocoder is presented. The IMBE vocoder is based on the MBE speech model, and it is shown to generate high-quality speech under both clean and noisy conditions. In addition, the IMBE vocoder is well suited for real-time implementation since it does not require excessive computation or storage. Full-duplex operation is demonstrated using a single AT&T WE DSP 32. Aspects of the hardware architecture, algorithm implementation, and system performance are addressed.< >