Otto is a call center agent designed to help cell phone and mobile provider customers. It has the twin goals of automating call center operations while maintaining a high level of customer satisfaction. It is intended to engage with human users in mixed-initiative dialogs to answer questions, explain procedures, and help with diagnostic troubleshooting. Within the domain of mobile devices, it collaborates with customers on specific tasks but maintains a degree of flexibility and naturalness in the interaction. The development of Otto incorporated aspects of human conversation to improve the quality and likelihood of the success of its interactions.
Although parsing with ontext-free phrase stru ture rules an be done in worstase ubi time, it is well-known that adding feature onstraints an make the resulting system exponential in the worst ase, if not undeidable. In this paper we observe that the standard algorithms for ombining feature onstraints with ontext-free phrase stru ture rules an be exponential even when the ombination is ontext-free equivalent. We then introdu e two new algorithms based on disjun tive lazy opy links that automati ally take advantage of simple ontext-freeness in su h grammars and parse in ubi time. The algorithms des ribed do not require that the grammar be pre-analyzed or ompiled; instead, the omputational properties are a by-produ t of how the algorithms work. Consequently, the algorithms improve performan e even for grammars that are not altogether ontext-free equivalent, sin e they will automati ally take advantage of situations where the amount of information owing upward from a onstituent to any of its onsumers is bounded.
Bogel et al. (2009) outlined a new architecture for modeling the interaction between prosody and syntax: we proposed an arrangement of interacting components in which prosodic information is developed in a module that operates independently of the syntax while still allowing for syntactic rules and preferences to be conditioned on prosodic boundaries and other features. This architecture allows for misalignments between prosodic units and syntactic constituency, but it incorporates a Principle of Prosodic Preference that causes syntactic structures that do not coincide with prosodic boundaries to be dispreferred. In this paper, we extend the proposal to account for so-called second position clitics. These are clitics that are interpreted syntactically as if they are immediate constituents of a clause, but their appearance after the first prosodic word may embed them in lower constituents and thus insulate them from normal clausal interpretation. We meet this theoretical challenge by adding to the architecture a mathematically restricted “interface mapping” in the form of a regular relation that mediates between the divergent syntactic and prosodic requirements that clitics must jointly satisfy.
In this paper we outline a new architecture for modeling the interaction between syntax and prosody. This architecture does not make use of correspondences between separate projections, but it is still consonant with the overall framework of LFG. We propose that prosodic information is developed in a component that operates independently of the syntax, thus allowing easy description of misalignment phenomena. We also propose a simple way of making prosodic information accessible to syntax, so that it is possible to condition syntactic rules and preferences on prosodic boundaries. We place the prosodic and syntactic components of the grammar in a pipeline configuration such that the terminal string of the syntactic tree is a sequence of lexical formatives intermixed with features inserted by the prosodic component. Depending on how they are distributed with respect to syntactic groupings, those features may or may not have an impact on the syntactic analysis.
In this paper we present a method for greatly reducing parse times in LFG parsing, while at the same time maintaining parse accuracy. We evaluate the methodology on data from English, German and Norwegian and show that the same patterns hold across languages. We achieve a speedup of 67% on the English data and 49% on the German data. On a small amount of data for Norwegian, we achieve a speedup of 40%, although with more training data we expect this figure to increase.
The demand for deep linguistic analysis for huge volumes of data means that it is increasingly important that the time taken to parse such data is minimized. In the XLE parsing model which is a hand-crafted, unification-based parsing system, most of the time is spent on unification, searching for valid f-structures (dependency attribute-value matrices) within the space of the many valid c-structures (phrase structure trees). We carried out an experiment to determine whether pruning the search space at an earlier stage of the parsing process results in an improvement in the overall time taken to parse, while maintaining the quality of the f-structures produced. We retrained a state-of-the-art probabilistic parser and used it to pre-bracket input to the XLE, constraining the valid c-structure space for each sentence. We evaluated against the PARC 700 Dependency Bank and show that it is possible to decrease the time taken to parse by ~18% while maintaining accuracy.
AbstractDeep grammars that include tokenization, morphology, syntax, and se-mantic layers have obtained broad coverage in conjunction with high effi-ciency. This allows them to play a crucial role in applications. However,these grammars are often developed as a general purpose grammar, expect-ing “standard” input, and have to be specialized for the application domain.This paper discusses some engineering tools that are used in the XLE gram-mar development platform to allow for domain specialization. It provides ex-amples of techniques used toallowspecialization viaoverlay grammarsat thelevel of tokenization, morphology, syntax, the lexicon, and semantics. As anexample, the paper focuses on the use of the broad coverage, general purposeParGramEnglishgrammarandsemanticsinthecontext ofan IntelligentDoc-ument Security Solutions (IDSS) system. Within this system, the grammar isused to automatically identify sensitive entities and relations among entities,which can then be redacted to protect the content.
We present an approach to statistical machine translation that combines ideas from phrase-based SMT and traditional grammar-based MT. Our system incorporates the concept of multi-word translation units into transfer of dependency structure snippets, and models and trains statistical components according to state-of-the-art SMT systems. Compliant with classical transfer-based MT, target dependency structure snippets are input to a grammar-based generator. An experimental evaluation shows that the incorporation of a grammar-based generator into an SMT framework provides improved grammaticality while achieving state-of-the-art quality on in-coverage examples, suggesting a possible hybrid framework.
We investigate some pitfalls regarding the discriminatory power of MT evaluation metrics and the accuracy of statistical significance tests. In a discriminative reranking experiment for phrase-based SMT we show that the NIST metric is more sensitive than BLEU or F-score despite their incorporation of aspects of fluency or meaning adequacy into MT evaluation. In an experimental comparison of two statistical significance tests we show that p-values are estimated more conservatively by approximate randomization than by bootstrap tests, thus increasing the likelihood of type-I error for the latter. We point out a pitfall of randomly assessing significance in multiple pairwise comparisons, and conclude with a recommendation to combine NIST with approximate randomization, at more stringent rejection levels than is currently standard.
This paper presents a strategy for a syntax based ranking of documents specifically oriented to Question Answering (QA). This strategy should limit the number of documents, processed by an answer extraction module of an syntax oriented QA system. Several measures for statistical scoring of expressions are presented and evaluated on 400 factoid questions from the TREC-12 competition. We prove that syntax based document filtering can outperform classical inverse document frequency approaches (idf).
When linguistically motivated grammars are implemented on a largerscale and applied to real-life corpora, keeping track ofambiguity sources becomes a difficult task. Yet it is of greatimportance, since unintended ambiguities arising fromunderrestricted rules or interactions have to be distinguishedfrom linguistically warranted ambiguities. In this paper wereport on various tools in the XLE grammar development platformwhich can be used for ambiguity management in grammar writing. Inparticular, we look at packed representations of ambiguities thatallow the grammar writer to view sorted descriptions of ambiguitysources. Also discussed are tools for specifying desired treestructures.
In this paper, we describe the types of sentence condensation rules used in the sentence condensation of Riezler et al. 2003 in detail. We show how the distinctions made in LFG f-structures as to grammatical functions and features make it possible to state simple but accurate rules to create smaller, well-formed f-structures from which the condensed sentence can be generated.
Complex Predicates are a crosslinguistically general phenomenon, but are more pervasive in South Asian than in European languages. This paper describes an LFG solution for Urdu/Hindi complex predication in terms of a RESTRICTION OPERATOR. The solution is theoretically well motivated and can be extended straightforwardly to related phenomena in European languages such as German, Norwegian, and French. 1 The ParGram Project In this paper, we report on the implementation of complex predicates (CP) for Urdu in the Parallel Grammar (ParGram) project (Butt et al., 1999; Butt et al., 2002). The ParGram project originally focused on three European languages: English, French, and German. Three other languages were added later: Japanese, Norwegian, and Urdu. The ParGram project uses the XLE parser and grammar development platform (Maxwell and Kaplan, 1993) to develop deep grammars, i.e., grammars which provide an in-depth analysis of a given sentence (as opposed to shallow parsing or chunk parsing, where a relatively rough analysis of a given sentence is returned). All of the grammars in the ParGram project use the Lexical-Functional Grammar (LFG) formalism, which produces c(onstituent)-structures (trees) and f(unctional)-structures (attribute-value matrices) as syntactic analyses. LFG assumes a version of Chomsky’s Universal Grammar hypothesis, namely that all languages are governed by similar underlying structures. Within LFG, f-structures encode a language universal level of analysis, allowing for crosslinguistic parallelism. ParGram aims to see how far parallelism can be maintained across languages. In the project, analyses for similar constructions across languages are held as similar as possible. This parallelism requires the formulation of a rigid standard for linguistic analysis. This standardization has the computational advantage that the grammars can be used in similar applications, and it can simplify cross-language applications such as machine translation (Frank, 1999). The conventions developed within the ParGram grammars are extensive. The ParGram project dictates not only the form of the features used in the grammars, but also the types of analyses chosen for constructions. The integration of new languages into the project has so far proven successful, including the adoption of the standards that were originally designed for the European languages (Butt and King, 2002b). As the new languages also contain constructions not necessarily found in the original European languages, the integration of new languages has contributed to the formulation of new standards of analysis. One such example is furnished by complex predicates in Urdu. 2 South Asian Complex Predicates South Asian languages are known for the extensive and productive use of CPs. CPs combine a light verb with a verb, noun or adjective to produce a new verb. For example, Urdu has a large class of “aspectual” CPs which combine with verbs to change the aktionsart properties of the event. Examples are shown in (1b,c), cf. (1a). (1) a. nAdyA AyI Nadya-NOM came ‘Nadya came.’ b. nAdyA A gayI Nadya-NOM come went ‘Nadya arrived.’ c. nAdyA A paRI Nadya-NOM come fell ‘Nadya came (suddenly, unexpectedly).’ The addition of a light verb modulates the event predication in subtle ways: beyond expressing defeasible meanings such as benefaction, suddenness, inception, or responsibility, the CP expresses a different aktionsart in comparison to the simple main verb. For example, in (1b) Nadya is in the result state of having arrived. The aktionsart effects of the light verbs on the event predication are quite complex and continue to be the subject of on-going theoretical research (Butt and Ramchand, 2003). The general effect is the encoding of a result state (a song is in the state of having been sung, a person is in the state of having arrived). However, a result state can be interpreted in two differing ways depending on whether one wants to consider the event to come (inception), or the event that has passed (completion). The precise interpretation is lexically determined by the light verbs. For the purposes of the Urdu grammar, we mark light verbs like ‘go’ as signifying completion of an action, whereas light verbs like ‘fall’ signify inception. Although these aspectual CPs do not alter the subcategorization frame of the verb, they change the resulting functional structure of the sentence, providing new information about the kind of event/action that is being described. The light verb also determines case marking on the subject: light verbs based on intransitive main verbs like paR ‘fall’ require a nominative subject. Light verbs like lE ‘take’ or dE ‘give’, which are based on (di)transitives main verbs, require an ergative subject. For example, transitive main verbs in the perfect tense usually require an ergative subject, as in (2a). When combined with a light verb like paR ‘fall’, the subject must be nominative as in (2b). Case marking in Urdu is governed by a combination of structural and semantic factors which we do not go into here (Butt and King, 2001). The light verb facts present an extension of the basic pattern. (2) a. nAdyA nE gAnA gayA Nadya-ERG song sang ‘Nadya sang a song.’ b. nAdyA gAnA gA paRI Nadya-NOM song sing fell ‘Nadya burst into song. c. nAdyA nE gAnA gA lIyA Nadya-ERG song sing took ‘Nadya sang a song (completely).’ As already mentioned, these CPs are extraordinarily productive in Urdu: most verbal predication involves complex predicate formation of the kind in (1) and (2). A light verb is in principle compatible with any main verb; however, (mostly semantic) selectional restrictions do apply so that some combinations are ruled out completely, whereas others are subject to considerable dialectal variation. Furthermore, the CPs are not formed within the lexicon, but are the result of the syntactic composition of two predicational elements (Alsina, 1996; Butt, 1995). Within LFG (as well as other syntactic frameworks), predicational elements play a special role: it is over these that argument saturation is checked. The difficulties involved with CP formation are better illustrated by means of another type of CP, the Urdu permissive, which alters the argument structure of the verb (Butt, 1995). The permissive light verb adds a new subject and “demotes” the other verb’s subject to a dative-marked indirect object, as in (3b), cf. (3a). (3) a. nAdyA sOyI Nadya-NOM slept ‘Nadya slept.’ b. yassin nE nAdyA kO sOnE dIA Yassin-ERG Nadya-DAT sleep-INF gave ‘Yassin let Nadya sleep.’ Since CPs are productive and occur frequently, an implementation that is both scalable and efficient is necessary. Most verbs can occur with several light verbs, and a given light verb can in principle occur with any verb of a given class (e.g., agentive verbs). So, it is not feasible to have multiple lexical entries for each verb depending on which light verb they occur with. This is especially true since the CPs combine with auxiliaries and other light verbs in predictable ways.
We report on the XLE parser and grammar development platform (Maxwell and Kaplan, 1993) and describe how a basic Lexical Functional Grammar for English has been adapted to two different corpora (newspaper text and copier repair tips).
Jonas Kuhn合作论文数Institute for Natural Language Processing, University of Stuttgart1