A smart home controller that responds to natural language input is demonstrated on an Intel embedded processor. This device contains two DSP cores and a neural network co-processor which share 4MB SRAM. An embedded configuration of the Intel RealSpeechTM speech recognizer and intent extraction engine runs on the DSP cores with neural network operations offloaded to the co-processor. The prototype demonstrates that continuous speech recognition and understanding is possible on hardware with very low power consumption. As an example application, control of lights in a home via natural language is shown. An Intel development kit is demonstrated together with a set of tools. Conference attendees are encouraged to interact with the demo and development system.
The problem of the effect of accent on the performance of Automatic Speech Recognition (ASR) systems is well known. In this paper, we study the effect of accent variability on the performance of the Indian English ASR task. We evaluate the test vocabularies on HMMs trained on (a) Accent specific training data (b) Accent pooled training data which combines all the accent specific training data (c) Accent pooled training data of reduced size matching the size of the accent specific training data. We demonstrate that the accent pooled training set performs the best on phonetically rich isolated word recognition task. But the accent specific HMMs perform better than the reduced accent pooled HMMs, indicating a possible approach of using a first stage accent identification to choose the correct accent trained HMMs for further recognition.
This paper compares two approaches of automatic age and gender classification with 7 classes. The first approach are Gaussian mixture models (GMMs) with universal background models (UBMs), which is well known for the task of speaker identification/verification. The training is performed by the EM algorithm or MAP adaptation respectively. For the second approach for each speaker of the test and training set a GMM model is trained. The means of each model are extracted and concatenated, which results in a GMM supervector for each speaker. These supervectors are then used in a support vector machine (SVM). Three different kernels were employed for the SVM approach: a polynomial kernel (with different polynomials), an RBF kernel and a linear GMM distance kernel, based on the KL divergence. With the SVM approach we improved the recognition rate to 74% (p < 0.001) and are in the same range as humans.
This paper presents a comparative study of four different approaches to automatic age and gender classification using seven classes on a telephony speech task and also compares the results with Human performance on the same data. The automatic approaches compared are based on (1) a parallel phone recognizer, derived from an automatic language identification system; (2) a system using dynamic Bayesian networks to combine several prosodic features; (3) a system based solely on linear prediction analysis; and (4) Gaussian mixture models based on MFCCs for separate recognition of age and gender. On average, the parallel phone recognizer performs as well as Human listeners do, while loosing performance on short utterances. The system based on prosodic features however shows very little dependence on the length of the utterance.
Recently an unsupervised learning scheme for Hidden Markov Models (HMMs) used in acoustical Language Identification (LID) based on Parallel Phoneme Recognizers (PPR) was proposed. This avoids the high costs for orthographically transcribed speech data and phonetic lexica but was found to intro-duce a considerable increase of classification errors. Also very recently discriminative Minimum Language Identification Error (MLIDE) optimization of HMMs for PPR based LID was introduced that again only requires language tagged speech data and an initial HMM. The described work shows how to combine both approaches to an unsupervised and discriminative learning scheme. Experimental results on large telephone speech databases show that using MLIDE the relative increase in error rate introduced by unsupervised learning can be reduced from 61% to 26%. The absolute difference in LID error rate due to the supervised learning step is reduced from 4.1% to 0.8%.
The goal of acoustic Language Identification (LID) is to identify the language of spoken utterances. The described system is based on parallel Hidden Markov Model (HMM) phoneme recognizers. The standard approach for parameter learning of Hidden Markov Model parameters is Maximum Likelihood (ML) estimation which is not directly related to the classification error rate. Based on the Minimum Classification Error (MCE) parameter estimation scheme we introduce Minimum Language Identification Error (MLIDE) training that results in HMM model parameters (mean vectors) that give minimum classification error on the training data. Using a large telephone speech corpus with 7 languages achieve a language classification error rate of 4.7% which is a 40% reduction of error rate compared with a baseline system using ML trained HMMs. Even if the system trained on fixed network telephone speech is applied to mobile network speech data MLIDE can greatly improve the system performance.
The goal of acoustic Language Identification (LID) is to identify the language of spoken utterances. The described system is based on parallel Hidden Markov Model (HMM) phoneme recognizers. The standard approach for parameter learning of Hidden Markov Model parameters is Maximum Likelihood (ML) estimation which is not directly related to the classification error rate. Based on the Minimum Classification Error (MCE) parameter estimation scheme we introduce Minimum Language Identification Error (MLIDE) training that results in HMM model parameters (mean vectors) that give minimum classification error on the training data. Using a large telephone speech corpus with 7 languages achieve a language classification error rate of 4.7% which is a 40% reduction of error rate compared with a baseline system using ML trained HMMs. Even if the system trained on fixed network telephone speech is applied to mobile network speech data MLIDE can greatly improve the system performance.
The offline HMM adaptation of a generic car speech recognizer to a specific car environment is investigated. For the generation of the adaptation database the approach of environment adapted databases (EADB) is applied that avoids real speech recordings in the target environment and therefore reduces the effort significantly. With MLLR adaptation using such an EADB a relative reduction of the word error rate of more than 10% can be achieved on a British city names task. It is proven by adaptation on real speech recordings from the target environment that the improvement with EADBs fully exploits the potential of HMM adaptation for the given car. Additionally it can be shown that if task matching material is available for adaptation a performance improvement of more than 30% can be reached with an additional maximum likelihood training iteration
Our system for automatic language identification (LID) of spoken utterances is performed with language dependent parallel phoneme recognition (PPR) using Hidden Markov Model (HMM) phoneme recognizers and optional phoneme language models (LMs). Such a LID system for continuous speech requires many hours of orthographically transcribed data for training of language dependent HMMs and LMs as well as phonetic lexica for every considered language (supervised training). To avoid the time consuming process of obtaining the orthographically transcribed training material we propose an algorithm for automatic unsupervised adaptation that requires only raw audio data as training material covering the requested language and acoustic environment. The LID system was trained and evaluated using fixed and mobile network databases (DBs) from the SpeechDat II corpus. The baseline system - based on supervised training using fixed network databases and covering 4 languages - achieved a LID error rate of 6.7 % for fixed data and 19.5 % for mobile data. Using unsupervised adaptation of the HMMs trained on fixed network data the error rate for mobile DBs database mismatch is reduced to 10.6 %. Exploring a situation when orthographically transcribed training data is not available at all multilingual HMMs were unsupervised adapted to fixed and mobile DBs and perform at 10.8 % and 12.4 % error rate respectively.
Speech recognition on embedded systems requires components of low memory footprint and low computational complexity. In this paper a POS-based (part of speech based) language modeling approach is presented which decreases the number of language model parameters combined with a method for reducing memory consumptions via quantization of language model penalties. For the application of short message dictation a language model with about 10,000 words of vocabulary is generated. Using the POS-based language modeling approach the number of parameters comprises 70,058 penalties. The memory consumptions for storing those penalties are reduced about 50% using the presented coding method. Experiments show that the POS-based language model is able to reduce the WER up to 65% for n-best isolated word recognition in comparison to the case without language model. Moreover the increase of WER caused by coding of the language model penalties is not significant.
This paper describes an extension to an automatic speech recognition system that improves the robustness concerning varying environments. A dedicated control unit tries to derive an optimal set of parameters for a Wiener Filter based noise reduction unit aiming at maximum recognition performances in different environments. The input measure for the control unit is derived from the speech signal. Apart from the SNR level, several other measures are investigated. The controlled parameters are closely related to the strength of the noise reduction. Several non-linear methods such as the Tabulated References and Neural Networks serve as the core of the control unit. Experiments on realistic handsfree as well as non-handsfree speech data show that the word error rate can be reduced by as much as 31% through the proposed methods. An already optimized static configuration of the applied noise reduction hereby serves as the baseline level.
One of the most commonly used discriminative approaches in parameter estimation for Hidden Markov Models is the Minimum Classification Error (MCE) method ([1]). This paper studies possible choices for the classes (i.e. basic speech units) in MCE training and their application for several tasks suitable for speech driven dialog systems in the telephone environment. The considered choices of classes are HMM states, phonemes, words and sequences of words. The theoretical suitability and practical considerations for the different criteria are discussed. Using the different training criteria consistent experimental results are given for four tasks: non–task–specific training, training for small vocabulary isolated word recognition, training for connected digit recognition and for letter recognition. In all experiments not only the objective of the optimization but also the resulting word recognition performance is investigated. It shows that for the given setup only word and word string based criteria are capable to reduce the word error rate.
Tying of Hidden Markov Model states is an important issue for the use of triphones as modeling units in automatic speech recognition systems. This paper studies the application of a–priori rules for tying in combination with data driven methods. The baseline method features a combination of a–priori rules that reduce the theoretical number of units by an oder of magnitude and a simple back–off tying. Back–off tying is based on the frequency of units appearing in the training material. The use of the a–priori rules has practical advantages especially for the implementation of continuous phoneme recognition. This method is compared to the widely used decision tree based clustering that makes no use of a–priori rules. A third method is proposed that combines a–priori rules with decision tree based clustering. Experiments on telephone data show that the combined method outperforms both other methods preserving the advantages of apply-ing a–priori rules.
Entering city names is an important issue for various speech driven applications such as telephone directory assistance. This paper proposes a system that combines word recognition with utterance verification and spelling as a fall back strategy. Word recognition experiments show that the use of a medium size vocabulary yields the lowest error rate when only permitting very few false acceptances. For letter recognition discriminatively trained letter models are applied. In order to shorten the spelling procedure two abort conditions are introduced which reduce the number of letters that have to be spelled. The system handles 96.5% of all calls correctly while less than 45% of all callers must spell the name.
The paper studies the use of discriminative techniques for a telephone based isolated digit recognizer with respect to a reduced system complexity. The combination of linear discriminant analysis (LDA) and minimum error classification (MEC) training provides improved system performance at reduced costs for the training process and for the application. Experiments are performed on an isolated digit database recorded over public lines including approximately 700 speakers. The use of a single linear transformation matrix based on LDA allows the use of density modeling, that doesn't consider variances explicitly at a high recognition rate. Minimum classification error training is found to perform best in case of a small amount of system parameters. A reduction of error rate up to 80% was achieved by the combination of the two methods for such a system configuration.
A system for understanding time utterances spoken in German language is presented. Stochastic models contain the knowledge in the semantic, syntactic and acousticphonetic levels. An adequate semantic representation allows the integration of these models within a one-pass Viterbi search. The simultaneous use of all knowledge sources for the search procedure results in the smallest possible search space for the determination of the most probable semantic content accurately following the Bayes classification rule. Both the recognition accuracy and the computing speed facilitate a realistic application.