This paper presents the results of an elaborate study on pen and speech-based multimodal interaction systems. The performance o f the “COM IC” system is assessed through human factors analyses and evaluation o f the acquired multimodal data. The latter requires tools that are able to m onitor user input, system feedback, and performance of the multimodal system components. Such tools can bridge the gap between observational data and the complex process o f the design and evaluation of multimodal systems. The evaluation tool presented here is validated in a human factors study on the usability o f COMIC for design applications and can be used for semi automatic transcription of multimodal data.
The design of a Spoken Dialogue System (SDS) is a long, iterative and costly process. Especially, it requires test phases on actual users either for assessment of performance or optimization. The number of test phases should be minimized, yet without degrading the final performance of the system. For these reasons, there has been an increasing interest for dialogue simulation during the last decade. Dialogue simulation requires simulating the behavior of users and therefore requires user modeling. User simulation is often done by statistical systems that have to be tuned or trained on data. Yet data are generally incomplete with regard to the necessary information for simulating the user decision making process. For example, the internal knowledge the user builds along the conversation about the information exchanged while interacting is difficult to annotate. In this contribution, we propose the use of a previously developed user simulation system based on Bayesian Networks (BN) and the training of this model using algorithms dealing with missing data. Experiments show that this training method increases the simulation performance in terms of similarity with real dialogues.
Cet article presente une mesure de voisement basee sur le calcul du signal analytique. Cette mesure peut etre utile pour plusieurs applications concernant le traitement de la parole. Par exemple, considerant la reconnaissance automatique de la parole, elle pourrait etre incorporee dans les vecteurs acoustiques communement utilises, comme par exemple les Mel Frequency Cepstral Coefficients (MFCC) et leurs deux premieres derivees, ceci pour ameliorer les performances du systeme. La base de donnees TIMIT est segmentee manuellement en phonemes : c'est pourquoi l'evaluation de la mesure developpee est effectuee sur cette base. L'information de voisement est deduite de cette segmentation. Il est montre dans cet article que la segmentation automatique voise/non-voise obtenue en utilisant la methode decrite dans les sections suivantes et la segmentation manuelle voise/non-voise fournie dans TIMIT sont tres similaires.
This document is a short report to accompany the prototype deliverable 3.6, due at month 30 of the CLASSiC project. The prototype focuses on user simulations which can represent different types of user, in both the Appointment Scheduling and Self-Help domains. This deliverable has been delayed to month 33 due to the time required to analyse data from WP 5 which was delivered later than planned. Publications related to this deliverable are presented in the Appendix. They are available at www.classic-project.org
User simulation has become an important trend of research in the field of spoken dialogue systems because collecting and annotating real interactions with users is often expensive and time consuming. Yet, such data are generally required for designing and assessing efficient dialogue systems. The general problem of user simulation is thus to produce as many as necessary natural, various and consistent interactions from as few data as possible. In this paper, we propose a user simulation method based on Bayesian Networks (BN) that is able to produce consistent interactions in terms of user goal and dialogue history but also to simulate the grounding process that often appears in human-human interactions. The BN is trained on a database of 1234 human-machine dialogues in the TownInfo domain (a tourist information application). Experiments with a state-of-the-art dialogue system (REALL-DUDE/DIPPER/OAA) have been realized and promising results are presented.
The demand for content-based management and real-time manipulation of audio data is constantly increasing. This paper presents a method to identify temporal regions, in a segment of co-channel speech, as being either single-speaker or multispeaker speech. The state of the art approach for this purpose is the kurtosis. In this paper, a set of complementary time-domain and frequency-domain features is studied. The employed classification scheme is the one-class SVM classifier. A recognition rate of 94.75 % is reached. The set of features providing the best performance is determined. Index Terms: Speech segmentation, Speaker characterization and recognition
User simulation has become an important trend of research in the field of spoken dialog systems because collecting and annotating real man-machine interactions with users is often expensive and time consuming. Yet, such data are generally required for designing and assessing efficient dialog systems. The general problem of user simulation is thus to produce as many as necessary natural, various and consistent interactions from as few data as possible. In this paper, is proposed a user simulation method based on Bayesian Networks (BN) that is able to produce consistent interactions in terms of user goal and dialog history but also to simulate the grounding process that often appears in human-human interactions. The BN is trained on a database of 1234 human-machine dialogs in the TownInfo domain (a tourist information application). Experiments with a state-of-the-art dialog system (REALL-DUDE/DIPPER/OAA) have been realized and promising results are presented.
User simulation has become an important trend of research in the field of spoken dialog systems because collecting and annotating real man-machine interactions with users is often expensive and time consuming. Yet, such data are generally required for designing and assessing efficient dialog systems. The general problem of user simulation is thus to produce as many as necessary natural, various and consistent interactions from as few data as possible. In this paper, is proposed a user simulation method based on Bayesian Networks (BN) that is able to produce consistent interactions in terms of user goal and dialog history but also to simulate the grounding process that often appears in human-human interactions. The BN is trained on a database of 1234 human-machine dialogs in the TownInfo domain (a tourist information application). Experiments with a stateof-the-art dialog system (REALL-DUDE/DIPPER/OAA) have been realized and promising results are presented.
This paper presents a method aimed at recognizing environmental sounds for surveillance and security applications. We propose to apply one-class support vector machines (1-SVMs) together with a sophisticated dissimilarity measure in order to address audio classification, and more specifically, sound recognition. We illustrate the performance of this method on an audio database, which consists of 1015 sounds belonging to nine classes. The database used presents high intraclass diversity in temps of signal properties and some kind of interclass similarities. A large discrepancy in the number of items in each class implies nonuniform probability of sound appearances. The method proceeds as follows: first, the use of a set of state-of-the-art audio features is studied. Then, we introduce a set of novel features obtained by combining elementary features. Experiments conducted on a nine-class classification problem show the superiority of this novel sound recognition method. The best recognition accuracy (96.89%) is obtained when combining wavelet-based features, MFCCs, and individual temporal and frequency features. Our 1-SVM-based multiclass classification approach overperforms the conventional hidden Markov model-based system in the experiments conducted, the improvement in the error rate can reach 50%. Besides, we provide empirical results showing that the single-class SVM outperforms a combination of binary SVMs. Additional experiments demonstrate our method is robust to environmental noise.
This paper deals with the voiced/unvoiced segmentation of natural speech signals and, inside the voiced intervals of these signals, with the determination of the fundamental frequency f0. Our primary motivation of developing the method presented in this paper is to obtain precise f0 information, this both in terms of frequency for relatively fast evolving events, and in terms of time location of these frequential events. The use of a first partial enhancement and of the Analytic Signal extraction techniques allows this. Comparisons with a standard FFT-based method are provided.
This paper proposes a voiced - unvoiced measure based on the Analytic Signal computation. This voiced - unvoiced feature can be useful for many speech processing applications. For instance, considering speech recognition, it could be incorporated into commonly used acoustic feature vectors, such as for example the Mel Frequency Cepstral Coefficients (MFCC) and their first two derivatives, in order to improve the performance of the overall system. The evaluation of the developed measure has been performed on the TIMIT database. TIMIT has been manually segmented into phones. The voicing information can easily be derived from this segmentation. It is shown in this paper that the automatic voiced - unvoiced segmentation obtained using the method described in the next sections and the manual voiced - unvoiced segmentation provided by TIMIT are very similar.
We deal with segmentation into note and/or into phone (according to the nature of the sound: instrumental part or singing voice excerpt or speech) or more generally into “stable” parts. There should have no restriction (see [9]) on the signal to be segmented: the sound can be monophonic music, but it can also be speech or polyphonic music; it can be harmonic or inharmonic (castanets or drums for example); fast notes (less than 100 milliseconds duration for example) and vibrato should be taken into account. The system presented in [2] and [6] was a work in progress towards this goal. The segmentation and feature extraction system described here is another step towards the more general system. It aims also to improve the robustness of the system described in [2] and [6].
In this paper we investigate the usability of speech-centric multimodal interaction by comparing two systems that support the same unfamiliar task, viz. bathroom design. One version implements a conversational agent (CA) metaphor, while the alternative one is based on direct manipulation (DM). Twenty subjects, 10 males and 10 females, none of whom had recent experience with bathroom (re-)design completed the same task with both systems. After each task we collected objective measures (task completion time, task completion rate, number of actions performed, speech and pen recognition errors) and subjective measures in the form of Likert Scale ratings.We found that the task completion rate for the CA system is higher than for the DM system. Nevertheless, subjects did not agree on their preference for one of the systems: those subjects who were able to use the DM system effectively preferred that system, mainly because it was faster for them, and they felt more in control.We conclude that for multimodal CA systems to become widely accepted substantial improvements in system architecture and in the performance of almost all individual modules are needed. (C) 2005 Elsevier B.V. All rights reserved.
This paper presents the results of an elaborate study on pen and speech-based multimodal interaction systems. The performance o f the “COM IC” system is assessed through human factors analyses and evaluation o f the acquired multimodal data. The latter requires tools that are able to m onitor user input, system feedback, and performance of the multimodal system components. Such tools can bridge the gap between observational data and the complex process o f the design and evaluation of multimodal systems. The evaluation tool presented here is validated in a human factors study on the usability o f COMIC for design applications and can be used for semi automatic transcription of multimodal data.
On-line pen input benefits greatly from mode detection when the user is in a free writing situation, where he is allowed to write, to draw, and to generate gestures. Mode detection is performed before recognition to restrict the classes that a classifier has to consider, thereby increasing the performance of the overall recognition. In this paper we present a hybrid system which is able to achieve a mode detection performance of 95.6% on seven classes; handwriting, lines, arrows, ellipses, rectangles, triangles, and diamonds. The system consists of three kNN classifiers which use global and structural features of the pen trajectory and a fitting algorithm for verifying the different geometrical objects. Results are presented on a significant amount of data, acquired in different contexts like scribble matching and design applications.
When the user is free to write anything, like handwriting, drawings, or gestures, techniques are required to distinguish between modes. Mode detection, preceding recognition, can be an important aid in applications that invite natural pen input. In this paper, a large amount of data, acquired in dierent contexts, is used to assess eight features on their suitability for mode detection. Six global features: length, area, compactness, eccentricity, circular variance, and closure, and two structural features: curvature, and perpendicularity, have shown to be particularly useful for determining whether a pen trajectory contains handwriting, lines, arrows, or geometric shapes. Using these eight features an overall performance on unseen data was achieved of 98.7%, using a KNN classifier. According to the principal components analysis of the data, the most important features were closure, curvature, perpendicularity, and eccentricity. The results of this study are employed in two large research projects on natural multi-modal interaction that pursue design, route map, and map annotation scenarios.
In this paper, ongoing research pursuing the distinction of online handwriting into textual and different drawing classes is described. In the context of natural pen-based interactions, users will seamlessly switch between such different input modes. Therefore, it is vital for pen input recognition systems to be able to distinguish between these cases, preferably in an early stage of processing. The method described in this paper is tested on data acquired in a multi-modal task setting where users are requested to specify shape and dimensions of bathrooms, using pen and speech. Mode detection in this context yields comparable outcomes to recent findings from the literature. The results presented here elaborate on these findings by examining the possibility to perform early recognition of input modes, so-called incremental recognition. To this end, PENDOWN as well as PENUP trajectories are being explored.