Major progress is being recorded regularly on both the technology and exploitation of automatic speech recognition (ASR) and spoken language systems. However, there are still technological barriers to flexible solutions and user satisfaction under some circumstances. This is related to several factors, such as the sensitivity to the environment (background noise), or the weak representation of grammatical and semantic knowledge. Current research is also emphasizing deficiencies in dealing with variation naturally present in speech. For instance, the lack of robustness to foreign accents precludes the use by specific populations. Also, some applications, like directory assistance, particularly stress the core recognition technology due to the very high active vocabulary (application perplexity). There are actually many factors affecting the speech realization: regional, sociolinguistic, or related to the environment or the speaker herself. These create a wide range of variations that may not be modeled correctly (speaker, gender, speaking rate, vocal effort, regional accent, speaking style, non-stationarity, etc.), especially when resources for system training are scarce. This paper outlines current advances related to these topics.
In this paper we report the results of experiments we carried out on our hybrid HMM/ANN systems which aims at combining Artiicial Neural Networks (ANN) and Hidden Markov Models (HMMs) for speech recognition of a French continuous speech database : BREF-80. As this database is not manually labelled, we describe a new method based on the temporal alignment of the speech signal on a high quality synthetic speech pattern to generate a rst segmentation in order to bootstrap the training procedure. A phone recognition experiment with our baseline system achieved a phone accuracy of about 75% which is very similar to the best results reported in the litterature 5]. Preliminary experiments on continuous speech recognition have set a baseline performance for our hybrid HMM/ANN system on BREF using 1K, 3K, 13K and 64 K word lexicons. All the experiments were carried out with the STRUT (Speech Training and Recognition Uniied Toolkit) software 11] and the NOWAY large vocabulary decoder 2]
In the paper, we expose a formalism that allows to make use of features representing both short-term and long-term spe ech behavior. This amounts to using multiple (specific, compensated or adapted) acoustic models which are defined according to additional hidden variables not pertaining to the pho netic sequence, but rather to long-term stable structures i n the speech signal, like the speaker identity or the speaking rat e. This formalism has been evaluated for recognition using vocal tract length (VTL) normalization. Features based on long-term pitch and formant measures, as well as PCA reductions of these, have been investigated and show significant correlation with the VTL. Speech recognition experiments performed on the children portion of the TI-DIGITS database show the improved accuracy obtained using this technique compared to VTL selection based on the traditional Maximum Likelihood criterion.
This paper briefly reviews state of the art related to the topic of speech variability sources in automatic speech recognition systems. It focuses on some variations within the speech signal that make the ASR task difficult. The variations detailed in the paper are intrinsic to the speech and affect the different levels of the ASR processing chain. For different sources of speech variation, the paper summarizes the current knowledge and highlights specific feature extraction or modeling weaknesses and current trends
Major progress is being recorded regularly on both the technology and exploitation of Automatic Speech Recognition (ASR) and spoken language systems. However, there are still technological barriers to flexible solutions and user satisfaction under some circumstances. This is related to several factors, such as the sensitivity to the environment (background noise or channel variability), or the weak representation of grammatical and semantic knowledge. Current research is also emphasizing deficiencies in dealing with variation naturally present in speech. For instance, the lack of robustness to foreign accents precludes the use by specific populations. There are actually many factors affecting the speech realization: regional, sociolinguistic, or related to the environment or the speaker itself. These create a wide range of variations that may not be modeled correctly (speaker, gender, speech rate, vocal effort, regional accents, speaking style, non stationarity...), especially when resources for system training are scarce. This paper outlines some current advances related to variabilities in ASR.
In this paper, a new acoustic confidence measure of automatic speech recognition hypothesis is proposed and it is compared to approaches proposed in the literature. This approach takes into account prior information on the acoustic model performance specific to each phoneme. The new method is tested on two types of recognition errors: the out-of-vocabulary words and the errors due to additive noise. An efficient way to interpret the raw confidence measure as a correctness prior probability is also proposed in the paper.
In this paper, we focus on the modeling of coarticulation and pronunciation variation in Automatic Speech Recognition systems (ASR). Most ASR systems explicitly describe these production phenomena through context-dependent phoneme models and multiple pronunciation lexicons. Here, we explore the potential benefit of using feature spaces covering longer time segments in terms of implicit modeling of coarticulation and pronunciation variants. The study is based on the analysis at the phonetic level of the performance of context-independent and context-dependent acoustic models, and more particularly the impact of modeling different time context going from 70 ms up to 310 ms on typical cases of pronunciation variants. Results, confirmed by word recognition experiment, put into light some ability of generic acoustic models to implicitly handle pronunciation variation.
The paper proposes a solution that brings some advances to the genericity of the ASR technology towards tasks and languages. A non-linear discriminant model is built from multi-lingual, multi-task speech material in order to classify the acoustic signal into language independent phonetic units. Instead of considering this model for direct HMM state likelihood estimation, it rather operates as a first stage to produce discriminant features that can be further used in cascade with a traditional task/language specific ASR system. This first stage structure is expected to achieve a strong modeling of the cross-language variability of speech that can better handle pronunciation variations due for instance to regional and non-native accents. Moreover, the flexibility of this architecture still allow the development of small task/language dedicated ASR systems as a second stage structure, possibly with small amount of data. The benefit of this architecture is demonstrated through a fine analysis of modeling performance at the phoneme level and on two different isolated word recognition tasks featuring accent variabilities
This paper intends to summarize recent developments and experimental results related to Automatic Speech Recognition (ASR) using signals captured with a throat-microphone. Due to the proximity of the sensor to the voice source, the signal is naturally less subject to background noise. This however yields speech sounds that have different frequency contents than with traditional microphones, and requires having specific acoustic models. We propose to use the information from both signals by combining the probability vectors provided by both acoustic models. The systems are evaluated on a connected digit recognition task in French. A database has been recorded for both training the acoustic models and for testing the whole setup. It contains both throat and “ordinary” close-talk signals. To avoid any possibly unrealistic assumption on the effect of noise on each signal, the test portion has been acquired using a background noise played back through loudspeakers. The ASR experiments that we achieved demonstrate the benefit of using alternative microphones. Relative recognition improvements as high as 80% were obtained on sequences of digits recorded in loud musical environment.
Automatic Speech Recognition (ASR) on Personal Digital Assistant (PDA) suffers from the intrinsic hardware characteristics of the audio interface, for example, low quality microphones and device internal noises. In this paper, we propose to compensate for these weaknesses by contaminating clean training data with the distortion sources that are specific to the target device. We present a method to estimate both the frequency response of the audio acquisition channel and the internal additive noise from a few tens of minutes of recordings on PDA. The channel characteristics are estimated from the longterm power spectra of clean speech and PDA recordings, while the noise power spectrum is estimated during silence segments in these recordings. All the recordings are performed in a controlled way, i.e. quiet environnement and no reverberation, in order to ensure that we measure only the internal device characteristics. The PDA-specific training data are then obtained by filtering the clean training data with the audio channel frequency response and contaminating them with internal noise, and a specific acoustic model is eventually trained for the target device. Recognition tests have been performed on digit sequences on three different PDA’s. Our approach has been compared to other channel and noise robust methods and presents very competitive performance.
In this paper we compare two different methods for automatically phonetically labeling a continuous speech data-base, as usually required for designing a speech recognition or speech synthesis system. The first method is based on temporal alignment of speech on a synthetic speech pattern; the second method uses either a continuous density hidden Markov models (HMM) or a hybrid HMM/ANN (artificial neural network) system in forced alignment mode. Both systems have been evaluated on read utterances not part of the training set of the HMM systems, and compared to manual segmentation. This study outlines the advantages and drawbacks of both methods. The speech synthetic system has the great advantage that no training stage (hence no large labeled database) is needed, while HMM Systems easily handle multiple phonetic transcriptions (phonetic lattice). We deduce a method for the automatic creation of large phonetically labeled speech databases, based on using the synthetic speech segmentation tool to bootstrap the training process of either a HMM or a hybrid HMM/ANN system. The importance of such segmentation tools is a key point for the development of improved multilingual speech synthesis and recognition systems.
This paper intends to summarize some of the robust feature extraction and acoustic modeling technologies used at Multitel, together with their assessment on some of the ETSI Aurora reference tasks. Ongoing work and directions for further research are also presented. For feature extraction (FE), we are using PLP coefficients. Additive and convolutional noise are addressed using a cascade of spectral subtraction and temporal trajectory filtering. For acoustic modeling (AM), artificial neural networks (ANNs) are used for estimating the HMM state probabilities. At the junction of FE and AM, the multi-band structure provides a way to address the needs of robustness by targeting both processing levels. Robust features within sub-bands can be extracted using a form of discriminant analysis. In this work, this is obtained using sub-band ANN acoustic models. Therobust sub-band features are then used for the estimation of state probabilities. These systems are evaluated on the Aurora tasks in comparison to the existing ETSI features. Our baseline system has similar performance than the ETSI advanced features coupled with the HTK back-end. On the Aurora 3 tasks, the multi-band system outperforms the best ETSI results with an average reduction of the word error rate of about 62% with respect to the baseline ETSI system and of about 18% with respect to the advanced ETSI system. This confirm previous positive experience with the multi-band architecture on other databases.
While wireless networks provide a greater autonomy to working people, interacting with a distant system is still constraining in terms of focus and abilities. Indeed, using a mobile terminal generally involves to look at it and use it manually. On another hand, last decade’s improvements in the field of speech processing techniques have allowed to realize voice enabled interfaces which represent a more natural way to interact for human operators and free their visual focus as their hands. Yet, low cost and low consuming mobile devices are not suitable for the implementation of complete speech processing algorithms which are often CPU and memory consuming. Moreover, inherent problems of wireless networking, such as connection loss, have to be taken into account in the design of a voice enabled interface. In this paper we propose an architecture of a mobile and distributed voice enabled interface for accessing to customer applications.
Lots of industrial tasks need contacts between operators and a central Information Management System. Permanent contact creates a more effective and efficient workforce with a direct update in the central Information Management System. In the case of mobile operators (e.g. warehouses or picking operators), this can be realized thanks to electronic remote devices connected to a private wireless network. Especially in mobile environment, voice often appears as a natural way to interface with computers. Indeed, a voice-enabled interface (understand voice recognition, dialogue management and speech synthesis) can improve user’s ergonomics, as hands and eyes can still be focused on the current task therefore improving the productivity by avoiding the necessary amount of time of extra manipulation as form-filling, data encoding, etc. In this paper, we propose a generic framework for the implementation of efficient mobile and distributed voice-enabled interfaces.
In this communication, we present a method for noise-robust multi-microphone automatic speech recognition (ASR). It is assumed that the speech source to be recognized is recorded with several microphones in a noisy acoustic environment. The proposed method estimates the short-term subband energies (as they are needed for computing the ASR front-end) of the clean speech source from the ones of the microphone noisy signals. The estimation procedure is based on the concept of Independent Component Analysis (ICA) and it is driven by the acoustic model used by the ASR de-coder. The method is shown to be highly robust for a connected digit recognition task in high noise conditions, improving word error rates by more than 50% relatively to the performance of the baseline single-microphone ASR system.
We present a fast method, i.e. requiring little data, for adapting a hybrid Hidden Markov Model / Multi Layer Perceptron speech recognizer to reverberant environments. Adaptation is per- formed by a linear transformation of the acoustic feature space. A dimensionality reduction technique similar to the eigenvoice approach is also investigated. A pool of adaptation transfor- mations are estimated a priori for various reverberant environ- ments. Then, the principal directions of the pool are extracted, the so-called eigenrooms. The adaptation transformation for ev- ery new reverberant environment is constrained to lay on the subspace spanned by the most significant eigenrooms. Conse- quently, the adaptation procedure involves estimating only the projection coefficients on the selected eigenrooms, which re- quires less data than direct estimation of the adaptation trans- formation. Supervised adaptation experiments for recognition of connected digit sequences (AURORA database) in reverberant environments are carried out. Standard adaptation demonstrates improvements in word error rate higher than 30% for typical re- verberation levels. The eigenroom-based adaptation technique implemented so far allows at most 50% reduction of adaptation data for the same improvement.
In this paper, we assess and compare four methods for the local estimation of noise spectra, namely the energy clustering, the Hirsch histograms, the weighted average method and the low-energy envelope tracking. Moreover we introduce, for these four approaches, the harmonic filtering strategy, a new pre-processing technique, expected to better track fast modulations of the noise energy. The speech periodicity property is used to update the noise level estimate during voiced parts of speech, without explicit detection of voiced portions. Our evaluation is performed with six different kinds of noises (both artificial and real noises) added to clean speech. The best noise level estimation method is then applied to noise robust speech recognition based on techniques requiring a dynamic estimation of the noise spectra, namely spectral subtraction and missing data compensation. (C) 2001 Elsevier Science B.V. All rights reserved.
In this paper, we present a new approach for improving the robustness of automatic speech recognition systems to additive noise. This approach lies in the use of a particular training procedure (based on data contamination) in a particular architecture (the multi-band paradigm). With this framework, we expect to remove the drawbacks of both the corpus contamination approach which is the dependency to noise spectral characteristics, and the multi-band architecture which is its relative inefficiency in the case of wideband noise. This method has been tested on the AURORA 2 continuous digits task and compared to other robust methods such as spectral subtraction, J-RASTA filtering and missing data compensation. It yields very good performance on different kinds of additive noise, without any a priori knowledge of the noise power spectrum.
In this paper, we propose a new acoustic confidence measure of ASR hypothesis and compare it to approaches proposed in the literature. This approach takes into account prior information on the acoustic model performance specific to each phoneme. The new method is tested on two types of recognition errors: the out-of-vocabulary words and the errors due to additive noise. We then propose an efficient way to interpret the raw confidence measure as a correctness prior probability.
This paper presents a method for blind estimation of reverberation times in reverberant enclosures. The proposed algorithm is based on a statistical model of short-term log-energy sequences for echo-free speech. Given a speech utterance recorded in a reverberant room, it computes a Maximum Likelihood estimate of the room full-band reverberation time. The estimation method is shown to require little data and to perform satisfactorily. The method has been successfully applied to robust automatic speech recognition in reverberant environments by model selection. For this application, the reverberation time is first estimated from the reverberated speech utterance to be recognized. The estimation is then used to select the best acoustic model out of a library of models trained in various artificial reverberant conditions. Speech recognition experiments in simulated and real reverberant environments show the efficiency of our approach which outperforms standard channel normalization techniques.
Nelson Morgan合作论文数Department of Electrical Engineering and Computer Sciences at UC Berkeley2
Jean Hennebert合作论文数;Software Engineering Unit
Business Information System Institute
University of Applied Science - HES-SO ; Wallis1