We report our recent work on applications of the MAP approach to estimating the time-varying polynomial Gaussian mean functions in the nonstationary-state or trended HMM. Assuming uncorrelatedness among the polynomial coefficients in the trended HMM, we have obtained analytical results for the MAP estimates of the time-varying mean and precision parameters. We have implemented a speech recognizer based on these results in speaker adaptation experiments using the T146 corpora. Experimental results show that the trended HMM always outperforms the standard, stationary-state HMM and that adaptation of polynomial coefficients only is better than adapting both polynomial coefficients and precision matrices when fewer than four adaptation tokens are used.
In this paper, we report our recent work on applications of the MAP approach to estimating the time-varying polynomial Gaussian mean functions in the nonstationary-state or trended HMM. Assuming uncorrelatedness among the polynomial coefficients in the trended HMM, we have obtained analytical results for the MAP estimates of the time-varying mean and precision parameters. We have implemented a speech recognizer based on these results in speaker adaptation experiments using TI46 corpora. Experimental results show that the trended HMM always outperforms the standard, stationary-state HMM and that adaptation of polynomial coefficients only is better than adapting both polynomial coefficients and precision matrices when fewer than four adaptation tokens are used.
The authors extend the maximum likelihood (ML) training algorithm to the minimum classification error (MCE) training algorithm for optimal estimation of the state-dependent polynomial coefficients in the trended HMM. The problem of automatic speech recognition is viewed as a discriminative dynamic data-fitting problem, where relative (not absolute) closeness in fitting an array of dynamic speech models to the unknown speech data sequence provides the recognition decision. In this view, the properties of the MCE formulation for training the trended HMM are analyzed by fitting raw speech data using MCE-trained trended HMMs, contrasting the poor discriminative fitting using the ML-trained models. Comparisons between the phonetic classification as well as data-fitting results obtained with ML and with MCE training algorithms demonstrate the effectiveness of the discriminatively trained trended HMMs.
We present in this paper an integrated view on the speech preprocessing and speech modeling problems in the design of a hidden Markov model (HMM) based speech recognizer. The integrated model we developed in this study generalizes the conventional, currently widely used delta-parameter technique, which has been confined strictly to the preprocessing domain only, in mio significant ways. First, the new model contains state-dependent weighting functions responsible for transforming static speech features into the dynamic ones in a slowly time-varying manner. Second, a novel maximum-likelihood based learning algorithm is developed for the model that allows joint optimization of the state-dependent weighting functions and the remaining conventional HMM parameters. The experimental results obtained from a standard TIMIT phonetic classification task provide preliminary evidence for the effectiveness of our new, general approach to the use of the dynamic characteristics of speech spectra. The results demonstrate that the new approach is most effective for discrimination of stop consonants exhibiting the fastest and most conspicuous dynamic patterns.
We investigate the interactions of front-end feature extraction and back-end classification techniques in HMM based speech recognizer. This work concentrates on finding the optimal linear transformation of Mel-warped short-time DFT information according to the minimum classification error criterion. These transformations, along with the HMM parameters, are automatically trained using the gradient descent method to minimize a measure of overall empirical error count. The discriminatively derived state-dependent transformations on the DFT data are then combined with their first time derivatives to produce a basic feature set. Experimental results show that Mel-warped DFT features, subject to appropriate transformation in a state-dependent manner, are more effective than the Mel-frequency cepstral coefficients that have dominated current speech recognition technology. The best error rate reduction of 9% is obtained using the new model, tested on a TIMIT phone classification task, relative to conventional HMM
In this study we implemented a speech recognizer based on the integrated view, proposed first by Deng (see IEEE Signal Processing Letters, vol.1, no.4, p.66-69, 1994), on the speech preprocessing and speech modeling problems in the recognizer design. The integrated model we developed generalizes the conventional, currently widely used delta-parameter technique, which has been confined strictly to the preprocessing domain only, in two significant ways. First, the new model contains state-dependent weighting functions responsible for transforming static speech features into the dynamic ones in a slowly time-varying manner. Second, novel maximum-likelihood and minimum-classification-error based learning algorithms are developed for the model that allows joint optimization of the state-dependent weighting functions and the remaining conventional HMM parameters. The experimental results obtained from a standard TIMIT phonetic classification task provide preliminary evidence for the effectiveness of our new, general approaches to the use of the dynamic characteristics of speech spectra.
In this paper we report our development of a new class of hidden Markov models (HMMs) with each state characterized by a time series model which is non-stationary up to the second order. A close-form solution for the model parameter estimation is obtained based on the EM algorithm and on the matrix-calculus implementation technique. In the first set of evaluation experiments, we adopt the residual square sum, over states and over time frames within state bounds, as a quantitative measure for goodness of fit between the model and the speech data. It is observed that inclusion of state-conditioned second-order non-stationarity, implemented by use of time-varying regression coefficients, has substantially greater effects on reducing data-fitting error than increase of the regression terms while maintaining the coefficients of each term constant. In the second set, isolated-word recognition experiments, it is found that use of mix of first-order and second-order non-stationarities consistently produces higher recognition accuracy than the conventional, stationary-state HMMs.