Current state-of-the-art automatic, continuous speech recognition systems have enjoyed huge leaps in accuracy using speech features that assumes stationarity in the signals that are being processed. However, the said performance can often be attributed to the inclusion of lexical information. For the case of isolated word tasks, without the use of a priori models for the expected words, the static speech representation breaks down. For example, words that only differ in one unvoiced part are often misrecognized. Thus, time-varying speech representations has become an interest in the field. This paper is concerned with the use of simple time-varying features based on an autoregressive modeling of speech that provides high resolution features. In particular, how the said high resolution features fit into a finite-length Hidden Markov Model-based acoustic model that was originally used for static features. Its performance is compared with the best performing static features (Mel-Frequency Cepstral Coefficients) and while it is currently viewed as suboptimal, ample rooms for improvement are also emphasized.
Isolated word speech recognition for small vocabulary tasks has found great success with Mel-frequency cepstral coefficients as the speech feature of choice. Voice-controlled embedded systems, using word models as the basic units of speech, have found their way in a variety of commercial products. While the recognition rates for these products can be considered commercially acceptable under clean environments, channel noise and other external factors can still degrade recognition performance in practice. We propose the use of cepstral features derived from time-varying linear predictive coding, where the autoregressive model of the speech signal is represented by coefficients that are linear combinations of some simple basis functions. Variations in the usage of the features are investigated, such as skipping adjacent features, averaging and hybrid features with the goal of improving the performance of a 142 vocabulary, isolated words Japanese speech recognition task.
In this paper, a new database suitable for HMM-based automatic Filipino speech recognition is described for the purpose of training a domain-independent, large-vocabulary continuous speech recognition system. Although it is known that high-performance speech recognition systems depend on a superior speech database used in the training stage, due to the lack of such an appropriate database, previous reports on Filipino speech recognition had to contend with serious data sparsity issues. In this paper we alleviate such sparsity through appropriate data analysis that makes the evaluation results more reliable. The best system is identified through its low word-error rate to a cross-validation set containing almost three hours of unknown speech data. Language-dependent problems are discussed, and their impact on accuracy was analyzed. The approach is currently data driven, however it serves as a competent baseline model for succeeding future developments.
It is widely known that database quality has a huge impact on speech recognition system performance, most especially when the expected domain is well represented. In this paper, we use this idea as leverage for a data-driven solution to the problem of code-switching in Filipino. Practical Filipino conversations often contain English and other loan words in varying frequencies, demanding better training of parameters and models for its speech recognition system. We alleviate the underrepresentation of loan words through the development of a new speech database for training, and applied appropriate data analysis to make reliable evaluation results. The best system was searched via lattice rescoring from a cross-validation set containing almost three hours of unknown speech data. The description and results of our experiments serve as a new and competent baseline model for succeeding future developments.
It is widely known that smoothing techniques are essential for n-gram-based statistical language modeling, especially in large vocabulary continuous speech recognition (LVCSR) tasks. The goal in this paper is to investigate several smoothing algorithms for n-gram models in Filipino LVCSR. The automatic speech recognition system was developed using the Janus Speech Recognition Toolkit (JRTk) of Carnegie Mellon University and Karlsruhe Institute of Technology. The language models were generated using Stanford's language modeling toolkit, SRILM. The data consisted of approximately 60 hours of transcribed recordings of Filipino speech from several domains spoken by 156 speakers. A total of 24 systems employing different language models were fine-tuned and tested for improved performance at a base metric. An instance of the Kneser-Ney algorithm with modified-at-end counts applied to an n-gram of order 5 registered the highest word recognition accuracy at 80.9% and 81.3% for the development and evaluation tests, respectively.
In this paper, the development of a closed captioning system for Filipino TV news programs is discussed. The researchers tested the system for offline captioning and evaluated the performance of the system based on word error rate (WER). Carnegie Mellon University's open-source speech recognition system, Sphinx-III, was used as the primary training and recognition engine. A Filipino News Corpus was built consisting of speech and text data obtained from Filipino news videos. Training and testing sets were generated and from this, different training and decoding parameters of Sphinx were evaluated. Using the word error rate (WER) computation, the highest average recognition accuracy achieved in developing for the test set was 57.36% using flat start context-dependent models and a language model with absolute discounting applied. This project is a first step towards establishing the baseline accuracy for future development of the system.
In this paper, the use of Source-Controlled Channel Decoding (SCCD) was investigated in estimating parameter values of parametric wideband speech codecs subjected to noise and packet losses. The technique was applied to LP-based and sinusoidal transform-based wideband speech codecs assuming a first order Markov chain to model interframe parameter distribution. Fading and non-fading channels were considered with AWGN using BPSK at the bit level and specified Packet Loss Rates (PLR) at the packet level. An approximation of the 2-state Gilbert model was used in erasing packets. The performance of the scheme was compared to the frame replication scheme of the GSM standard. The scheme gives favorable results in objective and subjective quality measures, and is very robust even at a high PLR of 30%.