We study two key issues in task independent training, namely selection of a universal set of subword units and modeling of the selected units. Since no a priori knowledge about the application vocabulary and syntax was used in the collection of the training corpus and the recognition task is frequently changing, the conventional strategy can no longer provide the best performance across many different tasks. We present an approach that uses the complete sets of right and left context dependent units as the basis phone sets. Training of these models is accomplished by a new training criterion that maximizes phone separation between competing models. The proposed phone selection and modeling approach was evaluated across different tasks in American English. Good recognition results were obtained for both context independent and context dependent phone models even for unseen tasks. The same strategy has also been applied to two other languages, Mandarin Chinese and Spanish, with similar success
In this paper, a minimum error rate pattern recognition approach to speech recognition is studied with particular emphasis on the speech recognizer designs based on hidden Markov models (HMMs) and Viterbi decoding. This approach differs from the traditional maximum likelihood based approach in that the objective of the recognition error rate minimization is established through a specially designed loss function, and is not based on the assumptions made about the speech generation process. Various theoretical and practical issues concerning this minimum error rate pattern recognition approach in speech recognition are investigated. The formulation and the algorithmic structures of several minimum error rate training algorithms for an HMM-based speech recognizer are discussed. The tree-trellis based N-best decoding method and a robust speech recognition scheme based on the combined string models are described. This approach can be applied to large vocabulary, continuous speech recognition tasks and to speech recognizers using word or subword based speech recognition units. Various experimental results have shown that significant error rate reduction can be achieved through the proposed approach.
Accurate and robust connected digit recognition is essential for a wide range of telecommunication services. Based on training and testing using only clean network digit data, and using the same whole-word model architecture as in the TI/NIST connected digit testing, the string error rate increased from less than 1% to more than 5%. The performance degraded even further when evaluated on data collected with different network conditions. Most of the observed errors were caused by changing channel characteristics, highly variable digit pronunciations, and inadequate modeling of cross-digit coarticulation. Results are presented for a number of context-dependent whole-word and subword modeling techniques developed to overcome some of the above problems. The most effective one is a new acoustic subword modeling approach that assumes that each digit model consists of three parts, namely, head, body, and tail subword units. Multiple heads and tails are also allowed, one for each of the 11 possible preceding and following digits and the background. Cross-digit coarticulation is modeled by connecting the pair of digits through the corresponding tail and head units. Testing on about 12 000 digit strings, collected from five regions, this new model architecture reduced the string error rate to under 2%.
Connected digit recognition is a problem that has received a lot of attention over the past several years because of its importance in providing speech recognition services (e.g., catalog ordering, credit card entry, all digit dialing of telephone numbers, etc.). Although a number of systems have been described that provide very high string accuracy on a standard database of connected digits (i.e., the TI database), most of these systems require a great deal of computation to provide high performance. Most recently, there has been a renewed interest in connected digit recognition systems based on discrete density models using VQ codebooks (e.g., the work of Normandin and colleagues at CRIM in Montreal) where the computation is significantly lower than that required for continuous density models, and the robustness to variations in talkers, background, microphones, etc., has the potential to be high. In this study, the effect of multiple codebooks, multiple models, and codebook weighting on the performance of a standard hidden Markov model recognizer using the TI connected digits database is examined. It is shown that a single codebook can provide high string recognition accuracy, and that multiple codebooks provide accuracy comparable to that of continuous density models.