We present a method to score automatic speech recognition (ASR) hypotheses. Potential candidate hypotheses are scored in terms of their phonetic confusability with all competing hypotheses in an n-best list. The scores are computed with a probabilistic phoneme error model of the ASR process that is taken as a black box. One of the applications of the scoring technique is nbest list re-ranking. In this paper, we evaluate and compare different phoneme error models for this task. We obtained significant improvements of the speech recognition accuracy for both spontaneous and read speech in conjunction with a decision tree classifier that can predict in which cases the re-ranking should be applied.
Current spoken dialogue systems (SDS) often behave inappropriately as they do not feature the same capabilities to detect speech recognition errors and handle them adequately as is achieved in human conversation. Adopting human abilities to identify perception problems and strategies to recover from them would enable SDS to show more constructive and naturalistic behavior.We investigated human error detection and error handling strategies within the context of a SDS for pedestrian assistance. The human behavior serves as a model for future algorithms that could yield reduced error rates in speech processing.The results contribute to a better understanding which knowledge humans employ to build up interpretations from perceived words and establish their confidence in perception and interpretation. The findings provide useful input for SDS developers and enable researchers to estimate the potential benefit of future research avenues.
We present a mobile tourist guide for planning and conducting sightseeing day trips. The system combines a hybrid recommender system for sights, events and other points of interest with a tour planner for time-constrained activities taking additional constraints for public transport connections into account. A novelty of the implemented approach compared to existing solutions for tourists is that the user retains full control over the tour by being able to directly edit any detail at any time, even if there are existing constraints that hinder the direct execution of an edit operation. Moreover, recommender and planner are closely interconnected by regarding the reachability of recommended items with respect to the current selection as well as filling unavoidable gaps with fitting recommendations. The application is tailored to the city of Nuremberg, Germany, but can be extended by additional data for other destinations as well.
We present a new method to augment the correct transcript from automatic speech recognition (ASR) output containing multiple hypotheses. The error-prone ASR process is taken as black box and modeled as a noisy channel on phoneme level. The probabilities of the individual phoneme errors are assigned according to phonetic confusability. We score potential candidate hypotheses by their posterior probability of being the channel input given the competing ASR hypotheses as observed output. The resulting scores provide useful information not included in traditional confidence measures. We investigated the usefulness of the method for rescoring, re-ranking and word error detection. The method alone is not powerful enough to improve the recognition results, but by employing a decision tree classifier it is possible to isolate cases where the method works very well. Our results show that the combination with other knowledge sources and postprocessing techniques can lead to promising improvements.
. Automatic speech recognition lacks the ability to integrate contextual knowledge into its optimization process – a resource that human speech perception makes extensive use of. We discuss shortcomings of current approaches to solve this problem, formalize the problem of context-aware speech recognition and understanding and introduce a robot navigation game that can be used to demonstrate and evaluate the impact of context on speech processing.
In this paper we present a longitudinal, naturalistic study of email behavior (n=47) and describe our efforts at isolating re-finding behavior in the logs through various qualitative and quantitative analyses.The presented work underlines the methodological challenges faced with this kind of research, but demonstrates that it is possible to isolate refinding behavior from email interaction logs with reasonable accuracy.Using the approaches developed we uncover interesting aspects of email re-finding behavior that have so far been impossible to study, such as how various features of email-clients are used in re-finding and the difficulties people encounter when using these.We explain how our findings could influence the design of email-clients and outline our thoughts on how future, more in depth analyses, can build on the work presented here to achieve a fuller understanding of email behavior and the support that people need.
Providing navigation assistance to users is a complex task generally consisting of two phases: planning a tour (phase one) and supporting the user during the tour (phase two). In the first phase, users interface to databases via constrained or natural language interaction to acquire prior knowledge such as bus schedules etc. In the second phase, often unexpected external events, such as delays or accidents, happen, user preferences change, or new needs arise. This requires machine intelligence to support users in the navigation real-time task, update information and trip replanning. To provide assistance in phase two, a navigation system must monitor external events, detect anomalies of the current situation compared to the plan built in the first phase, and provide assistance when the plan has become unfeasible. In this paper we present a prototypical mobile speech-controlled navigation system that provides assistance in both phases. The system was designed based on implications from an analysis of real user assistance needs investigated in a diary study that underlines the vital importance of assistance in phase two.
This paper presents our initial efforts at visualising personal information behaviour using Markov Chains. We describe a laboratory-based study of email re-finding and use Markov Chains, created from captured user interactions, as a means of understanding the behaviour exhibited. The models we generate not only provide an excellent overview of how the participants interacted with the experimental interface, but, by forcing the experimenters to ask questions they would not normally ask in order to comprehend the models, they also offer a starting point from which a fuller understanding of the exhibited behaviour can be attained. We illustrate this through examples, discuss the advantages and limitations of the approach and outline how we will expand on the work in future research.
We describe the linguistic richness and the technical aspects of an incremental finite-state parser for Icelandic. We argue that our parser outputs a linguistically rich annotation which in many simple sentences amounts to full parsing. Additionally, we provide arguments for various technical design and implementation decisions regarding the parser. Our description may be used as guidelines for other researchers developing similar parsers.
A major problem of understanding language in spoken dialog systems is to detect recognition errors in the output of a speech recognizer. Such a capability is the basis of implementing repair strategies that allow a dialog system to handle communication about misunderstandings similarly to other clarifications. In this paper we present a two-phase approach that combines chunk and dependency parsing and takes the global syntactic structure of recognizer output into account. This enables us to identify dependencies between chunks and detect syntactical errors caused by word confusions in case dependency constraints are violated. Finally, we apply these diagnostics to dialog modeling and discuss how the resulting error information can be used by clarification strategies.
Dominik Kuropka合作论文数alfabet AG1
Hilmar Schuschel合作论文数Hasso Plattner Institute for IT-Systems Engineering
at the University of Potsdam1