Automatic speech recognition in mobile devices has to cope with varying acoustical background noises in potentially low SNR situations. Therefore, techniques such as noise reduction are required to ensure the accuracy of the speech recognition process. This paper compares different kinds of environment compensation techniques, all operating frame-wise, either in the spectral or in the log-spectral domain. We report on word recognition rate as well as on word accuracy, the later also being a performance measure in absence of speech (i.e. only background noise) cases. As a classical technique, we first investigate Wiener filtering using a voice-activity-driven noise power spectral density (psd) estimation. Then we perform a comparison with the more advanced recursive least squares (RLS) weighting rule for speech enhancement, as well as with the use of minimum statistics as noise psd estimation. Finally, simulation results with the S-IMM (sequential interacting multiple models) approach are shown. It turns out that approaches known from speech enhancement perform very well also for speech recognition. This allows the use of the specific noise reduction function in a mobile device for speech telephony on the one hand, and for robust speech recognition on the other hand.
In this paper, we present a training-based approach to speech enhancement that exploits the spectral statistical characteristics of clean speech and noise in a specific environment. In contrast to many state-of-the-art approaches, we do not model the probability density function (pdf) of the clean speech and the noise spectra. Instead, subband-individual weighting rules for noisy speech spectral amplitudes are separately trained for speech presence and speech absence from noise recordings in the environment of interest. Weighting rules for a variety of cost functions are given; they are parameterized and stored as a table look-up. The speech enhancement system simply works by computing the weighting rules from the table look-up indexed by the a posteriori signal-to-noise ratio (SNR) and the a priori SNR for each subband computed on a Bark scale. Optimized for an automotive environment, our approach outperforms known-environment-independent-speech enhancement techniques, namely the a priori SNR-driven Wiener filter and the minimum mean square error (MMSE) log-spectral amplitude estimator, both in terms of speech distortion and noise attenuation.
Data-driven speech enhancement (Fingscheidt and Suhadi [1]) aims at improving speech quality for voice calls in a specific noise environment. The essence of the method are a set of frequency-dependent weighting rules, indexed by a priori and a posteriori SNRs, which are learned from clean speech and background noise training data. The weighting rules must be stored for each frequency bin separately and take up about 400 kBytes memory, which makes DSP implementations relatively expensive.In this paper we propose an alternative definition of the weighting rules which requires only 27 kBytes memory. That is 6.7% of the memory consumption of the original algorithm, with virtually no loss in performance measured in terms of speech distortion and noise attenuation. Our approach is to redefine the weighting rules on the Bark scale and store their parametric representation obtained by polynomial curve fitting.
Data-driven speech enhancement (Fingscheidt and Suhadi [1]) aims at improving speech quality for voice calls in a specific noise environment. The essence of the method are a set of frequencydependent weighting rules, indexed by a priori and a posteriori SNRs, which are learned from clean speech and background noise training data. The weighting rules must be stored for each frequency bin separately and take up about 400 kBytes memory, which makes DSP implementations relatively expensive. In this paper we propose an alternative definition of the weighting rules which requires only 27 kBytes memory. That is 6.7% of the memory consumption of the original algorithm, with virtually no loss in performance measured in terms of speech distortion and noise attenuation. Our approach is to redefine the weighting rules on the Bark scale and store their parametric representation obtained by polynomial curve fitting.
Automatic speech recognition in mobile devices has to cope with varying acoustical background noises in potentially low SNR situations. Its performance in car noise environments is of our particular interest. We put focus on noise reduction techniques as applicable for speech enhancement to ensure the accuracy of the speech recognition process. We report on word recognition rate as well as on word accuracy, the latter also being a performance measure in the absence of speech (i.e. only background noise) cases. As a classical technique, we first investigate Wiener filtering using a voice-activity-driven noise power spectral density (psd) estimation. Then we perform a comparison with the more advanced recursive least-squares (RLS) weighting rule for speech enhancement, as well as with the use of minimum statistics as noise psd estimation. Mel based root-cepstral coefficients has been taken as an alternative to the conventional Mel-frequency cepstral coefficients (MFCCs). The a-priori SNR based Wiener filtering with the minimum statistics and Mel based root-cepstral coefficients achieves 33.16% relative improvement in word accuracy over the classical technique.
In this paper we evaluate on a forensic task our text and language independent speaker recognition system, characterized by modest memory requirements and robustness to environment noise. Noise robustness is achieved by employing a Kalman filter-based sequential interacting multiple models (SIMM) algorithm. The evaluation data was provided by the Netherlands Forensic Institute (NFI) and consisted of telephone conversations in four different languages gathered in real police investigations. The results of NFI evaluation show that our small-footprint system provides competitive equal error rates (EER) for the class of text independent systems operating on telephone speech with strong channel mismatch.
In this paper we evaluate some model-based and data-driven algorithms for robust speech recognition in noise, using the experimental framework provided by ETSI Aurora 2. Specifically, we focus on statistical linear approximation (SLA), sequential interacting multiple models (S-IMM), and histogram normalization (HN). As the baseline for the feature extraction scheme we use the ETSI front-end. Recognition tests on a subset of Aurora 2 show that SLA is approximately 4 % better than HN and that S-IMM is worse than HN by almost 3 % in terms of absolute word accuracy. A comparison with the ETSI advanced front-end (AFE) is also presented. While none of these algorithms outperforms AFE, we identify the reasons why this might have happened and point out potential directions for improvement.
The invention relates to method for processing a noisy speech signal (S) for a subsequent speech recognition (SR), wherein the speech signal (S) representing at least a voice command, comprising the steps of: a) detecting the noisy speech signal (S); b) application of a noise reduction (NR) to the speech signal (S) for generating a noise-suppressed speech signal (S '); c) normalizing the noise-suppressed speech signal (S ') by means of a normalization factor to a desired value signal to generate a noise-suppressed, normalized speech signal (S' ').
L'invention concerne des procedes servant a traiter un signal vocal (S) empreint de bruit pour une reconnaissance vocale consecutive (SR), le signal vocal (S) representant au moins une commande vocale. Les procedes selon l'invention comprennent les etapes suivantes : a) detection du signal vocal (S) empreint de bruit ; b) application au signal vocal (S) d'une reduction de bruit (NR) afin de generer un signal vocal (S') a bruit reduit ; c) normalisation a une valeur de signal de consigne du signal vocal (S1) a bruit reduit au moyen d'un facteur de normalisation afin de generer un signal vocal normalise (S'') a bruit reduit.
A method for noise reduction (NC) in a speech input signal (SS) of a speaker comprising the steps of: detecting the speech input signal (SS); Accessing a predetermined voice characteristic (GMM-L-SD, XST-L-SD) of the speaker; Reducing a noise component in the voice input signal (SS) on the basis by means of the determined voice characteristics (GMM-L-SD, XST-L-SD) of the speaker.
This paper presents the Siemens speech recognizer for mobile phones, VSR. VSR employs HMM technology and uses general-purpose phoneme-based acoustic models which make it speaker and vocabulary independent. The system can be easily reconfigured to work with arbitrary vocabularies. This provides full flexibility for the design of the user interface which contrasts with the capabilities of other low-resource recognizers. The system requirements of VSR are very low. The emission probability calculation and the Viterbi search with a vocabulary of 30 words need only 16 MHz for real-time operation on an ARM microcontroller. The HMM acoustic models take up about 12 kilobytes of permanent storage. The most significant algorithmic improvement is the newly developed 3-D stream-based coding of the HMMs. Despite low requirements in terms of system resources VSR achieves an outstanding recognition performance. The word error rate (WER) for a recognition task with 62 German isolated words including highly confusable digits is 7.0%.
The performance of speaker verification (SV) systems degrades rapidly in noise rendering them unsuitable for security-critical applications in mobile phones, where false acceptance rates (FAR) of ∼ 10 − 4 are required. However, less demanding applications for which equal error rates (EER) comparable to word error rates (WER) of speech recognizers are acceptable could benefit from the SV technology. In this paper we evaluate two feature-based noise compensation algorithms in the context of SV: vector Taylor series (VTS) combined with statistical linear approximation (SLA), and Kalman filter-based interacting multiple models (IMM). Tests with theYOHO database and the NTT-AT ambient noises show that EERs as low as 5%–10% in medium to high noise conditions can be achieved for a text-independent SV system.
Hands-free operation of a mobile phone in car raises major challenges for acoustic enhancement algorithms and speech recognition engines. This is due to a degradation of the speech signal caused by reverberation effects and engine noise. In a typical mobile phone/carkit configuration only the car-kit microphone is used. A legitimate question is whether it is possible to improve the useful signal using the input from the second microphone, namely the microphone of the mobile terminal. In this paper we show that a speech enhancement algorithm specifically developed for two input channels significantly increases the word recognition rates in comparison with singlechannel noise reduction techniques.
Distributed speech recognition (DSR) is motivated by the fact that codecs used in speech transmission usually reveal a degrading voice quality below some channel quality (carrier-to-interferer ratio C/I), which justifies efficient coding of features with an appropriate channel coding in the mobile terminal. The Adaptive Multi-Rate (AMR) speech codec standardized for GSM and UMTS however delivers an acceptable speech quality way down to C/I ratios of about 4 dB in the GSM full-rate speech channel. In this paper we investigate network-based speech recognition (NSR) using a conventional speech channel with AMR coding as an alternative to a DSR system. This approach is natural and attractive since information services usually require a duplex channel for conversation anyway, furthermore no change to existing mobiles is required. For a GSM full-rate channel it turns out that an NSR system based on AMR coding indeed is comparable to DSR approaches.