The idea for this special issue came up during the preparations of the International Workshop on Finite-State Methods in Natural Language Processing, that was held at Bilkent University in Ankara, Turkey in the summer of 1998. The number of the submissions had exceeded our initial expectations and we were able to select quite a good set of papers from those submitted. Further, the workshop and the preceding tutorial by Kenneth Beesley, on finite-state methods, was attended by quite a large number of participants. This led us to believe that interest in the theory and applications of finitestate machinery was alive and well, and that some of the papers from this workshop along with further additional submissions could make a very good special issue for this journal. The five papers in this issue are the result of this process. The last decade has seen a quite a substantial surge in the use of finite-state methods in all aspects of natural language applications. Fueled by the theoretical contributions of Kaplan and Kay (1994), Mohri’s recent contributions on the use of finite-state techniques in various NLP problems (Mohri 1996, 1997), the success of finite-state approaches especially in computational morphology, for example, Koskenniemi (1983), Karttunen (1983), and Karttunen, Kaplan, and Zaenen (1992), and, finally, the availability of state-of-the-art tools for building and manipulating large-scale finite-state systems (Karttunen 1993; Karttunen and Beesley 1992; Karttunen et al. 1996; Mohri, Pereira, and Riley 1998; van Noord 1999), recent years have seen many successful applications of finite-state approaches in tagging, spell checking, information extraction, parsing, speech recognition, and text-to-speech applications. This is a remarkable comeback considering that in the dawn of modern linguistics (Chomsky 1957), finitestate grammars were dismissed as fundamentally inadequate. As a result, most of the work in computational linguistics in the past few decades has been focused on far more powerful formalisms. Recent publications on finite-state technology include two collections of papers (Roche and Schabes 1997; Kornai 1999) with contributions covering a wide range of these topics. This special issue, we hope, will add to these contributions. The five papers in this collection cover many aspects of finite-state theory and applications. The papers Treatment of Epsilon Moves in Subset Construction by van Noord and Incremental Construction of Minimal Acyclic Finite-State Automata and Transducers by Daciuk, Watson, Watson, and Mihov, address two fundamental aspects in the construction of finite-state recognizers. Van Noord presents results for various methods for producing a deterministic automaton with no epsilon transitions from a nondeterministic automaton with a large number of epsilon transitions, especially those resulting from finite-state approximations of context-free and more powerful formalisms. Daciuk et al. present a new method for constructing minimal, deterministic, acyclic
Routing and Recursive Routing Networks (RRNs) are highly expressive neural networks with modular self-assembling architectures that have proven successful in complex NLU tasks. However, this expressive power can at times be both a blessing and a curse. For many practical problems, vanilla RRNs tend to overfit to the limited data available. However, recent work has shown that high-quality meta-information can be extremely useful as a guide for routing in settings that require sample efficiency. Unfortunately, this meta-information is highly problem-dependent, and oftentimes it is not available at test-time. To compensate, we introduce an additional network that is trained jointly with the routing network to map samples to meta-information: a dispatcher. The dispatcher’s goal is finding groups of samples that are as useful to the router as the groups defined by meta-information. We find that RRNs augmented with an end-to-end dispatcher achieve strong performance in multi-task learning scenarios while exhibiting high levels of generalization when adapting to new, unseen tasks.
The recent years have seen an exponential increase in the amount of information available through the Internet on any given topic. Information retrieval techniques have been steadily improving and can provide a mass of relevant results, but those results still have to be processed and digested by a human reader. Computer displays of traditional documents often simply model the way a paper rendition of the document would work. But more effective display techniques may be possible by exploiting the dynamic display properties of the computer screen. In this paper, we describe a system called Animated Dynamic Highlighting (ADH), which has been added to the ReadUp reader, part of the UpLib personal digital library system. It is an interactive, user-controlled technique that enhances the presentational aspects of the reading task. We ran a pilot study to get an initial impression of user acceptance of active display techniques, comparing ADH to two other display techniques. The results of the study are discussed here.
Deep learning models for semantics are generally evaluated using naturalistic corpora. Adversarial methods, in which models are evaluated on new examples with known semantic properties, have begun to reveal that good performance at these naturalistic tasks can hide serious shortcomings. However, we should insist that these evaluations be fair - that the models are given data sufficient to support the requisite kinds of generalization. In this paper, we define and motivate a formal notion of fairness in this sense. We then apply these ideas to natural language inference by constructing very challenging but provably fair artificial datasets and showing that standard neural models fail to generalize in the required ways; only task-specific models that jointly compose the premise and hypothesis are able to achieve high performance, and even these models do not solve the task perfectly.
We introduce Recursive Routing Networks (RRNs), which are modular, adaptable models that learn effectively in diverse environments. RRNs consist of a set of functions, typically organized into a grid, and a meta-learner decision-making component called the router. The model jointly optimizes the parameters of the functions and the meta-learner’s policy for routing inputs through those functions. RRNs can be incorporated into existing architectures in a number of ways; we explore adding them to word representation layers, recurrent network hidden layers, and classifier layers. Our evaluation task is natural language inference (NLI). Using the MultiNLI corpus, we show that an RRN’s routing decisions reflect the high-level genre structure of that corpus. To show that RRNs can learn to specialize to more fine-grained semantic distinctions, we introduce a new corpus of NLI examples involving implicative predicates, and show that the model components become fine-tuned to the inferential signatures that are characteristic of these predicates.
Standard evaluations of deep learning models for semantics using naturalistic corpora are limited in what they can tell us about the fidelity of the learned representations, because the corpora rarely come with good measures of semantic complexity. To overcome this limitation, we present a method for generating data sets of multiply-quantified natural language inference (NLI) examples in which semantic complexity can be precisely characterized, and we use this method to show that a variety of common architectures for NLI inevitably fail to encode crucial information; only a model with forced lexical alignments avoids this damaging information loss.
When the first generation of generative linguists discovered presuppositions in the late 1960s and early 1970s, the initial set of examples was quite small. Aspectual verbs like stop were discussed already by Greek philosophers, proper names, Kepler, and definite descriptions, the present king of France, go back to Gottlob Frege and Bertrand Russell by the turn of the century. Just in the span of a few years my generation of semanticists assembled a veritable zoo of ‘presupposition triggers’ under the assumption that they were all of the same species. Generations of students have learned about presuppositions from Stephen Levinson’s 1983 book on Pragmatics that contains a list of 13 types of presupposition triggers, an excerpt of an even longer unpublished list attributed to a certain Lauri Karttunen. My task in this presentation is to come clean and show why the items on Levinson’s list should not have been lumped together. In retrospect it is strange that the early writings about presupposition by linguists and even by philosophers like Robert Stalnaker or Scott Soames do not make any reference to the rich palette of semantic relations they could have learned from Frege and later from Paul Grice. If we had known Frege’s concepts of Andeutung – Grice’s conventional implicature – and Nebengedanke, it would have been easy to see that there are types of author commitment that are neither entailments nor presuppositions.
Natural Logic attempts to do formal reasoning in natural language in a proof-theoretic way making use of the syntactic structure and the semantic properties of lexical items and constructions. It goes beyond trying to establish who did what to whom to yield further inferences. Natural Logic contrasts with approaches that involve a translation from a natural into a formal language such as predicate calculus or a higher-order logic that is sometimes deemed a more appropriate basis for reasoning although difficult to implement efficiently. Both approaches have been used in current work on computational semantics.
This paper starts with a brief history of Natural Logic from its origins to the most recent work on implicatives. It then describes on-going attempts to represent the meanings of so-called 'evaluative adjectives' in these terms based on what linguists have traditionally assumed about constructions such as NP was stupid to VP, NP was not lucky to VP that have been described as factive. It turns out that the account cannot be based solely on lexical classification as the existing framework of Natural Logic assumes.The conclusion we draw from this ongoing work is that Natural Logic of the classical type must be grounded in a more inclusive theory of Natural Reasoning that takes into account pragmatic factors in the context of use such as the assumed relation between the evaluative adjective and even the perceived communicative intent of the speaker.
This is is an experimental study of the semantics of the construction NP was (not) Adj to VP where Adj is an evaluative adjective such as stupid. We show that in the simple past tense this construction is predominantly factive for most people but implicative for some. We also demonstrate that the interpretations are sensitive to preconceptions about how suitable the adjective is as a characterization of the event described by the in nitival clause. This Consonance/Dissonance e ect gives the construction its chameleon-like characteristics.
Proceedings of the First Annual Meeting of the Berkeley Linguistics Society (1975), pp. 266-278
In this note, we look at the factors that influence veridicity judgments with factive predicates. We show that more context factors play a role than is generally assumed. We propose to use crowd sourcing techniques to understand these factors better and briefly discuss the consequences for the association of lexical signatures with items in the lexicon. 1 Veridicity: what and why Recognizing the inferential properties of constructions and of lexical items is important for NLU (Natural Language Understanding) systems. In this paper we look at FACTUAL INFERENCES, inferences that allow the reader to conclude that an event has happened or will happen or that a state of affairs pertains or will pertain. We will refer to events and states together as SOAs. Factuality is in the world and outside of the text. In cases where the reader has no direct perceptual knowledge about the SOAs, she has to evaluate the factuality of a SOA referred to in a text based on her decoding of the author’s representation of the factuality of the SOA and on her knowledge about the world and about the author’s reliability. Authors have a plethora of means to signal whether they want to present SOAs as factual, as having happened or going to happen or as being more or less probable, possible, unlikely or not factual at all. We will call this presentation of a SOA the VERIDICITY of a SOA. We will call the reader’s interpretation of the author’s intention, the RIV (READER INFERRED VERIDICITY) and the reader judgment about the factuality of a SOA, RIF (READER INFERRED FACTUALITY). Annotation can, at its best, only provide us with RIVs as the author is typically not available for consultation. This leads to a methodological problem. A reader will in his interpretation of a sentence be sensitive, not only to the way an author signals her intentions but also to what he knows about the world. To circumvent this problem as much as possible, corpus annotation for veridicity is typically done by trained annotators with extensive guidelines (see e.g. (Sauri, 2008), (Sauri and Pustejovsky, 2012)) but corpus annotation by trained annotators is an expensive enterprise, hence looks at a limited number of cases. For instance, to anticipate on a case we will discuss later in the paper, lucky occurs only once in the FactBank ((Sauri and Pustejovsky, 2009). Given that annotation is done on running text, it is also difficult to avoid that the reader’s evaluation of the wider extralinguistic context might still play a role. We propose to supplement corpus annotation with crowd sourcing experiments. In these, sentences are presented to Mechanical Turk workers in limited contexts, very similar to the contexts in which linguists judge the effect of the contribution of a lexical item or a construction. But contrary to linguistic practice, we derive our examples from really occurring ones culled from the web and, more importantly, present them to many native speakers (typically 100) and in different variations to explore factors that can influence the interpretation. This kind of variation is very difficult to find in naturally occurring corpora of the type that are used for annotations (e.g. FactBank). This type of study comple-
Twenty years ago morphological analysis of natural language was a challenge to computational linguists. Simple cut-and-paste programs could be and were written to analyze strings in particular languages, but there was no general language-independent method available. Furthermore, cut-and-paste programs for analysis were not reversible, they could not be used to generate words. Generative phonologists of that time described morphological alternations by means of ordered rewrite rules, but it was not understood how such rules could be used for analysis. This was the situation in the spring of 1981 when Kimmo Koskenniemi came to a conference on parsing that Lauri Karttunen had organized at the University of Texas at Austin. Also at the same conference were two Xerox researchers from Palo Alto, Ronald M. Kaplan and Martin Kay. The four Ks discovered that all of them were interested and had been working on the problem of morphological analysis. Koskenniemi went on to Palo Alto to visit Kay and Kaplan at PARC. This was the beginning of Two-Level Morphology, the first general model in the history of computational linguistics for the analysis and generation of morphologically complex languages. The language-specific components, the lexicon and the rules, were combined with a runtime engine applicable to all languages. In this article we trace the development of the finite-state technology that Two-Level Morphology is based on.
This paper complements a series of works on implicative verbs such as manage to and fail to. It extends the description of simple implicative verbs to phrasal implicatives as take the time to and waste the chance to. It shows that the implicative signatures of over 300 verb-noun collocations depend both on the semantic type of the verb and the semantic type of the noun in a systematic way.
fst stands for Finite-State Toolkit. It is an enhanced version of the xfst tool described in the 2003 Beesley and Karttunen book Finite State Morphology. Like xfst, fst serves two purposes. It is a development tool for compiling finite-state networks and a runtime tool that applies networks to input strings or files. xfst is limited to morphological analysis and generation. fst can also be used for other applications. This paper describes the new features of the fst regular expression formalism and illustrates their use for named-entity recognition, relation extraction, tokenization and parsing. The fst pattern matching algorithm (pmatch) operates on a single pattern network but the network can be the union of any number of distinct pattern definitions. Many patterns can be matched simultaneously in one pass over a text. This is a distinct fst advantage over pattern matching facilities in languages such as Perl and Python.
To make learning by reading easier, one strategy is to map language that is read onto a normalized knowledge representation. The normalization maps alternative expressions of the same content onto essentially identical representations. This brief position paper illustrates the power of this approach in PARC’s Bridge system through examples of textual inference.
André Kempe合作论文数Yahoo! Search Technologies in Paris3
Juhani Karhumäki合作论文数Department of Mathematics, University of Turku2
Valeria De Paiva合作论文数School of Computer Science University of Birmingham, Birmingham, UK2