Cross-correlation-based methods have been used extensively in the task of locating multiple, simultaneous sound sources in adverse environments. In this paper, we present a low-cost prealignment enhancement to fix the temporal drawback of the cross-correlation functionals by aligning all the microphone signals to ensure they correlate to the same temporal event. We further introduce a new functional, the steered-response power of the minimum-variance distortionless response using the phase transform (MVDR-PHAT), for multiple-source detection and localization. Experimental results using real data of a 10-talker recording in an adverse room show the improvements of the proposed functional and the prealignment enhancement over traditional techniques in detecting and locating simultaneously active talkers.
Speaker identification is a well-established research problem but has not been a major application used in gaming scenarios. In this paper, we propose a new algorithm for the open-set, text-independent, speaker ID problem, applied as an important component (among other cues) of a game player identification system. This scenario poses new challenges: far-field, limited training and very short test data, and almost real-time processing. To tackle this, we introduce new and more informative feature sets. The scores given by these feature sets are then combined in an optimal way to construct the final score. Experimental results on the gaming device's processed reverberated-speech show the effectiveness of the new features, and that reliable decisions can be made after very short (2 - 5 second) test utterances required by the gaming scheme.
Extracting a high-quality speech signal of a single source from a multiple-source input in an adverse environment has always been a challenge for microphone-array processing. Three major approaches have been proposed to tackle this problem: blind-source separation (BSS), beamforming (BF), and computational auditory scene analysis (CASA). Combinations of the CASA and BF, BSS and BF also have been introduced. In this paper, we propose a new algorithm which utilizes the null-steering beamformer minimum-variance distortionless response (MVDR) using the proven-robust phase transform (MVDR-PHAT) and the CASA framework that closely mimics human hearing perception. Experimental results using real data recorded in a room with high background and reverberation noise indicated the improved performance of the proposed algorithm compared to those of traditional beamforming algorithms and an SRP-PHAT-based source-separation algorithm recently described at ICASSP 2010.
Two new methods for locating multiple sound sources using a single segment of data from a large-aperture microphone array are presented. Both methods employ the proven-robust steered response power using the phase transform (SRP-PHAT) as a functional. To cluster the data points into highly probable regions containing global peaks, the first method fits a Gaussian mixture model (GMM), whereas the second one sequentially finds the points with highest SRP-PHAT values that most likely represent different clusters. Then the low-cost global optimization method, stochastic region contraction (SRC), is applied to each cluster to find the global peaks. We test the two methods using real data from five simultaneous talkers in a room with high noise and reverberation. Results are presented and discussed.
Computational cost has been an issue for the proven robust source localization algorithm, steered response power (SRP) using the phase transform (SRP-PHAT). Some proposed computation reduction algorithms degrade under high noise and reverberant conditions. Some require at least 10% the cost of a full SRP-PHAT gridsearch. In ICASSP 2007, we introduced a robust, low-cost global optimization technique, stochastic region contraction (SRC). In this paper, we present another algorithm, stochastic particle filtering (SPF), which uses SRC's initialization and is a kind of Importance Sampling technique. In this paper, the SRP is computed using a modification to the conventional PHAT, namely Ã-PHAT. Extensive experiments using real data and simulated data are shown. The results indicate that, while maintaining the desirable accuracy of the full search, this method reduces the cost to about half the cost of SRC (0.03% the cost of full search), thus making SRP-PHAT more practical for real-time applications.
In this paper we present a new method for locating multiple sound sources using only a local segment of data from a large-aperture microphone array. The result of this work may be used directly or as an open-loop input to a tracking algorithm. The proposed method employs the proven-robust steered response power using the phase transform as a functional, agglomerative clustering, and low-cost global optimization (stochastic region contraction). We test the algorithm with five simultaneous "talkers" under very difficult conditions in a real room but using electrical speakers instead of human talkers to have a controlled experiment. Results are presented and discussed.
Most real microphone-array applications require sound sources to be localized in a noisy, reverberant environment. In such conditions, the steered response power using the phase transform (SRP-PHAT) has been shown to be more robust than faster, two-stage, time-difference of arrival methods. The complication is that the SRP-PHAT space has many local extrema which has required computationally costly grid-search methods. In this paper, we introduce the use of coarse-to-fine region contraction (CFRC) to make computing the SRP practical. We compare CFRC cost and performance to that of using stochastic region contraction (SRC), a method we presented recently at ICASSP 2007, which showed the computation for SRC was reduced by about 3 orders of magnitude from a comparatively fine grid-search. Results here from real data from human talkers show that CFRC costs about the same as SRC overall, but requires only about 63% of SRC's cost under very noisy conditions.
In most microphone array applications, it is essential to localize sources in a noisy, reverberant environment. It has been shown that computing the steered response power(SRP) is more robust than faster, two-stage, direct time-difference of arrival methods. The problem with computing SRP is that the SRP space has many local maxima and thus computationally-intensive grid-search methods are used to find a global maximum. Grid search is too expensive for a real-time system. Several papers have addressed this issue. In this paper we propose using stochastic region contraction(SRC) to make computing the SRP practical. We discuss one important SRP method, computing it from the phase transform (SRP-PHAT), review SRC, and show the computational saving. Using real data from human talkers, we show that SRC saves computation by more than two orders of magnitude with almost no loss in accuracy.