Most voice driven applications are based on recognition grammars. In complex applications it is difficult to exactly predict how the users will formulate their requests even if a careful study of the user's behavior has been performed. Moreover, it is possible that a speaker's word pronunciation does not match the phonetic transcription of the system, mainly in the case of foreign words.Loquendo has developed a tool that collects field data, detects the most significant weaknesses of the application due to pronunciation of formulation mismatches, and filters the collected field corpora. This permits the application designers to perform their analysis only on a reasonable amount of preprocessed and automatically labeled data.This paper presents the approaches that have been devised to detect pronunciation variants of vocabulary words and linguistic formulations not covered by the recognition grammar. Results showing the improvements that have been obtained including automatically detected formulations in three grammars for two languages are also detailed.
Telecom Italia has deployed since the beginning of year 2001 a nationwide automatic Directory Assistance (DA) system that routinely serves customers asking for residential and business listings.
One of the main problems in automatic directory assistance (DA) for business listings is that customers formulate their requests for the same listing with a great variability. We show that an automatic approach allows the detection, from field data, of user formulations that were not foreseen by the designers, and that they can be added, as variants, to the denominations already included in the system to reduce its failures.
This paper exploits the broad concept of dialogue predictions by linking a point in a human-machine dialogue with a speci(cid:12)c language model which is used during the recognition of the next user utterance. The idea is to cluster several dialogue contexts into a class and to create a speci(cid:12)c language model for each class. We present an automatic algorithm based on the minimal decrease of mutual information which clusters the dialogue contexts. Moreover the algorithm is able to guess an appropriate number of classes, that gives a good trade o(cid:11) between the mutual information and the amount of training data. This automatic classi(cid:12)cation procedure allows the full automatic creation of context-dependent language models for a spoken dialogue system.
In this paper we explain how contextual expectations are generated and used in the task-oriented spoken language understanding system Dialogos. The hard task of recognizing spontaneous speech on the telephone may greatly benefit from the use of specific language models during the recognition of callers' utterances. By 'specific language models' we mean a set of language models that are trained on contextually appropriated data, and that are used during different states of the dialogue on the basis of the information sent to the acoustic level by the dialogue management module. In this paper we describe how the specific language models are obtained on the basis of contextual information. The experimental result we report show that recognition and understanding performance are improved thanks to the use of specific language models.
Analyses language modeling in spoken dialogue systems for accessing a database. The use of several language models obtained by exploiting dialogue predictions gives better results than the use of a single model for the whole dialogue interaction. For this reason, several models have been created, each one for a specific system question, such as the request for or the confirmation of a parameter. The use of dialogue-dependent language models increases the performance both at the recognition level and at the understanding level, especially on answers to system requests. Moreover, using other methods to increase the performance, like the automatic clustering of vocabulary words or the use of better acoustic models during recognition, does not affect the improvements given by dialogue-dependent language models. The system used in our experiments is Dialogos, the Italian spoken dialogue system used for accessing railway timetable information over the telephone. The experiments were carried out on a large corpus of dialogues collected using Dialogos.
This paper is focused on the language modelling for task-oriented domains and presents an accurate analysis of the utterances acquired by the Dialogos spoken dialogue system. Dialogos allows access to the Italian Railways timetable by using the telephone over the public network. The language modelling aspects of specificity and behaviour to rare events are studied. A technique for getting a language model more robust, based on sentences generated by grammars, is presented. Experimental results show the benefit of the proposed technique. The increment of performance between language models created using grammars and usual ones, is higher when the amount of training material is limited. Therefore this technique can give an advantage especially for the development of language models in a new domain.
Toward Automatic Adaptationof the Acoustic Models and of the Formulation Variants in a Directory Assistance Application / M. ANDORNO; LAFACE P.; C. POPOVICI; L. FISSORE; C. VAIR. (2001), pp. 175-178. ((Intervento presentato al convegno ISCA ITR-Workshop 2001 on Adaptation Methods for Speech Recognition nel August. Original Toward Automatic Adaptationof the Acoustic Models and of the Formulation Variants in a Directory Assistance Application