This report describes how spoken language turns are segmented into utterances in the framework of the verbmobil project. The problem of segmenting turns is directly related to the task of annotating a discourse with dialogue act information: an utterance can be characterized as a stretch of dialogue that is attributed one dialogue act. Unfortunately, this rule in many cases is insufficient and many doubtful cases remain. We tried to at least reduce the number of unclear cases by providing a number of hands-on rules for segmenting dialogues. In that sense, we hope that this document is helpful for labeling dialogues with discourse information. This report has to be seen as an extension of the verbmobil Report No.65 that includes a detailed discussion of the various dialogue act types employed in verbmobil. All dialogue act types used throughout this report are taken from the set defined there.
In an era where tourism is dominated by requests for tailored experiences, SMEs play a key role in providing adequate products and services to tourists by responding to their most specific requirements.This paper uses network and clusters as a framework providing SMEs with innovative opportunities to operate in a competitive tourism environment. A review of relevant literature on clusters, networks and tourism business innovation is undertaken, then focusing on the specific issues of Healthy Lifestyle Tourism.The UK 'Healthy Lifestyle Tourism Cluster' experience is employed to discuss the process and the implication of network and cluster development in tourism. However, the development of clusters should not be seen as a simple and spontaneous process due to the nature of businesses involved, but as a very complex process linked to strong stakeholder collaboration. (c) 2006 Elsevier Ltd. All rights reserved.
This report describes the dialogue phases and the second edition of dialogue actswhich are used in the Verbmobil 2 project(see [1] for the first edition). Whilein the first project phase the scenario was restricted to appointment schedulingdialogues, it has been extended to travel planning in the second phase with appointmentscheduling being only a part of the new scenario [7].As a consequence, the range of tasks to be solved during the dialogue hasgrown: formerly, we only had to cope...
This paper focuses on a particular problem of Automatic Dialogue Interpreting. As the input is unplanned spoken language, it contains certain traces of the effort to formulate the utterances (such as false starts, filled pauses, self-repairs or repetitions). In an adequate translation, these traces are obviously to be eliminated. I introduce a pragmatic-based approach to Automatic Dialogue Interpreting that handles these traces adequately. This approach is characterized by the fact that dialogue acts representing both the propositional content and the communicative intention of an utterance are used as a translation invariant.
This paper starts from the assumption that collaboration is a fundamental prerequisite of dialogues. This view is applied to a special ins ta ce of human-human dialogues that are mediated by an interpreter. I demonstrat e how this situation differs from a normal human-human-dialogue situation: The res pon ibility for mutual understanding and the control over the dialogue lies wi th the interpreter. The paper then focuses on the question how a human-human dialogu e mediated (i.e. interpreted) by a machine can be adequately modeled. Based o n the assumption that the machine cannot bear the responsibility for underst anding, I propose certain strategies to compensate for the machine’s incapabili ty to take the control.
In this paper we demonstrate that for an adequate translation of an utterance spoken in a dialogue the dialogue act it performs has to be determined. We introduce an approach that automatically assigns types of dialogue acts to utterances on the basis of both micro- and macro-structural information. Technically, this assignment is realized by modeling preference rules as weighted defaults in the Description Logic system FLEX. The dialogue-act type of an utterance is determined by qualitatively minimizing the exceptions to these defaults. The results described here have been developed within the VERBMOBIL project, a project concerned with face-to-face dialogue interpreting funded by the German Federal Ministry of Education, Science, Research and Technology (BMBF). We present the rather positive results of a first evaluation of this implementation showing the accuracy of dialogue act assignment.
The resolution of ambiguities is one of the central problems for Machine Translation. In this paper we propose a knowledge-based approach to disambiguation which uses Description Logics (dl) as representation formalism. We present the process of anaphora resolution implemented in the Machine Translation systemfast and show how thedl systemback is used to support disambiguation. The disambiguation strategy uses factors representing syntactic, semantic, and conceptual constraints with different weights to choose the most adequate antecedent candidate. We show how these factors can be declaratively represented as defaults inback. Disambiguation is then achieved by determining the interpretation that yields a qualitatively minimal number of exceptions to the defaults, and can thus be formalized as exception minimization.
On the occasion of the annual meeting of the Deutsche Gesellschaft fuer Sprachwissenschaft in March 94 we organized a workshop on ambiguity and disambiguation strategies. This report contains seven of the contributions presented at the workshop (mostly written in German). In organizing the workshop we aimed at bringing together researchers from such different fields as computational linguistics, artificial intelligence, machine translation, theoretical linguistics, psychology, and translation theory. Birgit Apfelbaum uses a corpus of German-French interpreting situations as a basis for investigating human strategies towards disambiguation. Michael Grabski discusses a specific aspect of polysemy in the context of translation, namely cases where a lexical distinction in the source language is not made in the target language. Klaus von Heusinger proposes a formal framework in which definite and anaphoric noun phrases as well as anaphoric pronouns can be uniformly represented. Bernhard Kipper discusses ambiguities arising in the interpretation of modal verbs. Lars Konieczny, Barbara Hemforth, and Nicole Voelker investigate human strategies to resolve ambiguities stemming from PP-attachment. Birte Prahl investigates human translation strategies in a situation where only limited contextual information is available. Manfred Stede describes an ambiguity problem, namely the lexicalization of the ``substitution'''' relation (e.g. `but'', `instead'', `rather'') from the perspective of language generation. In our introduction we briefly sketch our own perspective on the problem of ambiguity, distinguish various aspects of ambiguity, indicate which aspects are addressed by the articles, and finally summarize their contents.
The results described here have been developed within the field of automatic interpreting. We focus on the analysis of spoken language with dialogue processing methods. We demonstrate that for an adequate translation of most of the verbs occurring in appointment scheduling dialogues the speech-event types of their utterance is relevant. We propose an approach for the automatic assignment of speech-event types to utterances. This approach applies weighted preference rules that are based on micro as well as macro-structural information.
This report describes the domain model used in the German Dialogue Interpreting project VERBMOBIL. In order to make the design principles underlying the modeling explicit, we begin with a brief sketch of the VERBMOBIL demonstrator architecture from the perspective of the domain model. We then present some rather general considerations on the nature of domain modeling and its relationship to semantics. We claim that the semantic information contained in the model mainly serves two tasks. For one thing, it provides the basis for a conceptual transfer from German to English; on the other hand, it provides information needed for disambiguation. We argue that these tasks pose different requirements, and that domain modeling in general is highly task-dependent. A brief survey of domain models or ontologies used in existing NLP systems confirms this position. We finally describe the different parts of the domain model, explain our design decisions, and present examples of how the information contained in the model can be actually used in the VERBMOBIL demonstrator. In doing so, we also point out the main functionality of FLEX, the Description Logic system used for the modeling.
In this paper we try to combine experiences gained in the Machine Translation project FAST with research concerning the default extension of the terminological representation system BACK. We analyze various types of ambiguities that are problematic for translation and show that disambiguation is possible on the basis of default assumptions. Some of these defaults form preference rule systems, while others are merely an elegant way of modeling exceptions. We show how the ideas underlying the FAST implementation can be reformulated in a declarative framework by using terminological logics. In particular the implementation of anaphora resolution can only be expressed in this framework with resort to defaults. Whereas default reasoning is usually concerned with the problem of deducing valid extensions of a given set of formulas, preference rule systems pose slightly different requirements. A preference ordering of the possible extensions can be computed by taking into account which defaults are violated in them. Since the extension with the qualitatively minimal number of exceptions is the preferred one, disambiguation can be modeled by exception minimization.
The project dealt with the resolution of anaphoric expressions with respect to the need of Machine Translation. This involves several aspects that are crucial for Machine Translation: Translation of texts instead of single sentences, treatment of ambigous expressions, integration of background knowledge and strategies to deal with uncertain knowlegde. The project concentrated on personal and possessive pronouns that refer to objects and are referentially identical with their antecedents. A resolution procedure was developed that integrates different criteria (morphological, syntactic, semantic and conceptual) to determine the antecedent of a pronominal anaphor. The resolution procedure is based on a dual representation of the text: The structural aspects are represented by means of the Functor Argument Structure that has been developed in the preceeding project. The referential aspects are represented by the knowlegde representation system BACK that has been integrated into the Berlin MT-System. On the basis of the criteria so far developed remarkably good results were achieved. Moreover, the treatment of personal and possessive pronouns was unified. It is possible to apply the resolution procedure to other types of ambiguity as well, but then one has to take interdependencies between ambiguities into account. Besides anaphoric resolution other interfering subjects were dealt with, such as problems of consistency of the lexicon, of logical foundation of the MT system and of the term rewriting procedure that is used for various mappings within the system. A detailed description of the results can be found in the final project report.
In this paper we give an overview of an approach to anaphora resolution that takes a whole variety of different factors into account. These factors concern on the one hand structural information like agreement, proximity, binding principles, subject preference, topic preference, negative preference for free adjuncts and on the other hand information about the contents of a text like conceptual consistency. This led to a twofold text representation: a structural text representation and a referential text representation. The factors are implemented as preference rules with different weights that express the influence each of the factors has in the process of anaphora resolution. For each factor a linguistic motivation as well as a formal representation of the information it works on is given. We show how we integrated the anaphora resolution component into the existing experimental MT system.
In this paper we want to show how representation of textual and background knowledge can be used fo Machine Translation (MT). As a basis we took the result of the KIT FAST Project, which started from the problem of anaphora resolution. We argue that two kinds of text representation are useful in order to deal with intersentential phenomena: one is to represent structural characteristics of the text, and the other is to represent aspects of the contents. We discuss a concept of modeling background knowledge for MT, that is based on task-oriented criteria. On balance the result is that MT and especially anaphora resolution require a language particular modeling of conceptual knowledge. As a means to represent text contents and background knowledge the project KIT FAST uses the BACK-System, that is based on terminological logic.
This paper investigates the role knowledge representation can play in machine translation (MT). After a short survey of the methods that have been used in MT to desambiguate an expression in the source language, I formulate the requirements for a knowledge representation formalism concerning syntax and semantics and demonstrate how KL-ONE fulfills these requirements. I am concerned with the question to what degree KL-ONE fits the task of natural language processing. Several MT-systems that use knowledge representation systems are represented. From the advantages and disadvantages of these systems I infer some proposals for an ``ideal'''' MT-system.
The research of the Berlin project of the complementary research for EUROTRA-D is based on a model of a translation system, which assigns for four levels of representation: a syntactic and a sentence semantic level, a level for the representation of texts and as an invariant the thematic and argumentative structure of text. In the preceding project ''''New algorithms for analysis and synthesis for machine translationthe Generalized Phrase Structure Grammar (GPSG) has been investigated for its applicability for the syntactic level of representation and has been implemented to verify the hypotheses, which have been developed.