The combined use of different set of features extracted from the speech signal with different processing algorithms is a promising approach to improve speech recognition performances. Artificial Neural Networks are well suited to this task since they are able to use directly multiple heterogeneous input features to estimate a near optimal combination of them for classification, without being constrained by a priori assumptions on the stochastic independence of the input sources. This work shows how we have taken advantage of these characteristics of Neural Networks to improve the recognition accuracy of our systems. In particular, three set of input features have been considered as sources in this work: Mel based Cepstral Coefficients derived from the FFT spectrum, RASTA-PLP Cepstral Coefficients, and a set of features that describe the dynamics of the FFT power spectrum along the frequency dimension, instead of the usual time dimension. The experimental results confirm the usefulness of the proposed approach of feature integration that leads to a significant error reduction both on isolated and continuous speech recognition tasks on a large telephone speech test set. ,
This paper exploits the broad concept of dialogue predictions by linking a point in a human-machine dialogue with a speci(cid:12)c language model which is used during the recognition of the next user utterance. The idea is to cluster several dialogue contexts into a class and to create a speci(cid:12)c language model for each class. We present an automatic algorithm based on the minimal decrease of mutual information which clusters the dialogue contexts. Moreover the algorithm is able to guess an appropriate number of classes, that gives a good trade o(cid:11) between the mutual information and the amount of training data. This automatic classi(cid:12)cation procedure allows the full automatic creation of context-dependent language models for a spoken dialogue system.
Interactions with spoken language systems may present breakdowns that are due to errors in the acoustic decoding of user utterances. Some of these errors have important consequences in reducing the naturalness of human-machine dialogues. In this paper we identify some typologies of recognition errors that cannot be recovered during the syntactico-semantic analysis, but that may be effectively approached at the dialogue level. We will describe how non-understanding and the effects of misrecognition are dealt with by Dialogos, a realtime spoken dialogue system that allows users to access a database of railway information by telephone. We will discuss the importance of supporting confirmation turns, and clarification and correction sub-dialogues. We will show the positive effects of robust dialogue management and dialogue state dependent language modeling, by taking into account both the recognition and understanding performance, and the success rate of dialogue transactions.
In this paper we explain how contextual expectations are generated and used in the task-oriented spoken language understanding system Dialogos. The hard task of recognizing spontaneous speech on the telephone may greatly benefit from the use of specific language models during the recognition of callers' utterances. By 'specific language models' we mean a set of language models that are trained on contextually appropriated data, and that are used during different states of the dialogue on the basis of the information sent to the acoustic level by the dialogue management module. In this paper we describe how the specific language models are obtained on the basis of contextual information. The experimental result we report show that recognition and understanding performance are improved thanks to the use of specific language models.