
Information extraction (IE) is becoming increasingly useful as a form of shallow semantic analysis. Learning relational facts from text is one of the core tasks of IE and has applications in a variety of fields including summarization, question answering, and information retrieval. Previous work has traditionally relied on extensive human involvement (e.g., hand-annotated training instances, manual pattern extraction rules, hand-picked seeds). Standard supervised techniques can yield high performance when large amounts of hand-labeled data are available for a fixed inventory of relation types, however, extraction systems do not easily generalize beyond their training domains and often must be re-engineered for each application. In this talk I will present an unsupervised approach to relational information extraction which could lead to significant resource savings and more portable extraction systems that require less engineering effort. The proposed model partitions tuples representing an observed syntactic relationship between two named entities (e.g., “X was born in Y” and “X is from Y”) into clusters corresponding to underlying semantic relation types (e.g., BornIn, Located). Our approach incorporates general domain knowledge which we encode as First Order Logic rules. Specifically and automatically combine with we combine a topic model developed for the relation extraction task with automatically extracted domain relevant rules, and present an algorithm that estimates the parameters of this model. Evaluation results on the ACE 2007 English Relation Detection and Categorization (RDC) task show that our model outperforms competitive unsupervised approaches by a wide margin and is able to produce clusters shaped by both the data and the rules.
The Textual Entailment task has become influential in NLP and many researchers have become interested in applying it to other tasks. However, the two major issues emerging from this body of work are the fact that NLP applications need systems that (1) attain results which are not corpus dependent and (2) assume that the text for entailment cannot be incorrect or even contradictory. In this paper we propose a system which decomposes the text into chunks via a shallow text analysis, and determines the entailment relationship by matching the information contained in the is − a pattern. The results show that the method is able to cope with the two requirements above.
In this paper, we describe an ongoing experiment which aims to extend Hungarian WordNet with new verb-noun relations that specify selectional restrictions for various argument positions. We present an algorithm that uses frequency data from a representative corpus and information from a verb frame description database to generate sets of semantic classes, represented by WN hypernym sub-networks. The method intends to cover all possible argument positions of verbs found in the corpus which are marked by various case inflections or postposition particles. The new links in HuWN are assigned corpus-based probabilities. We present some preliminary results and discuss some of the arising issues.
Taxonomy-based representations are widely used to model compactly large amounts of textual data. While current methods allow organizing knowledge at the lexical level (keywords/concepts/topics), there is an increasing demand to move towards more informative representations, which express properties of concepts and relations among them. This demand triggered our research on statement entailment graphs. In these graphs, nodes are natural language statements (propositions), comprising of predicates with their arguments and modifiers, while edges represent entailment relations between nodes. In this talk we report initial research that defines the properties of entailment graphs and their potential applications. Particularly, we show how entailment graphs can be profitably used for both knowledge acquisition and text exploration. Beyond providing a rich and informative representation, statement entailment graphs allow integrating multiple semantic inferences. So far, textual inference research focused on single, mutually independent, entailment judgments. However, in many scenarios there are dependencies among Text/Hypothesis pairs, which need to be captured consistently. This calls for global optimization algorithms for inter-dependent entailment judgments, taking advantage of the overall entailment graph structure (e.g. ensuring entailment graph transitivity). From the applied perspective, we are experimenting with entailment graphs in the context of the EXCITEMENT project industrial scenarios. We focus on the text analytics domain, and particularly on the analysis of customer interactions across multiple channels, including speech, email, chat and social media, and multiple languages (English, German, Italian). For example, we would like to recognize that the complaint they charge too much for sandwiches entails food is too expensive, and allow an analyst to compactly navigate through an entailment graph that consolidates the information structure of a large number of customer statements. Our eventual applied goal is to develop a new generation of inference-based text exploration applications, which will enable businesses to better analyze their diverse and often unpredicted client content. This task will be exemplified with data collected from real customer interactions, while referring to the EXCITEMENT Open Platform that we developed as a generic open source framework for textual inferences.
By applying the Corpus Pattern Analysis procedure (CPA, Hanks 2004) to the analysis of concordances for ca 1000 English, Italian and Spanish verbs conducted with the aim of acquiring their most recurrent patterns, intended as corpus-derived argument structures with specification of the expected semantic type for each argument position (i.e. [[Human]] attends [[Event]]), we compiled a list of about 220 semantic types obtained from manual clustering and generalization over sets of lexical items found in the argument positions in the corpus (details of the Italian project in Jezek 2012). These types look very much like conceptual / ontological categories for nouns but should instead be conceived as semantic classes, as they are induced by the analysis of selectional properties of verbs. They are language-driven, and reflect how we talk about entities in the world. As such, despite the obvious correlations, they differ from categories of entities defined on the basis of ontological axioms, such as those of DOLCE (Descriptive Ontology for Linguistic and Cognitive Engineering), which, despite “aiming at capturing the ontological categories underlying natural language and human common sense” (Gangemi et al. 2002) does not base category distinctions on systematic observation and clustering of language data. In my presentation, I will report the preliminary results of the experiment of aligning the type inventory to the categories of DOLCE, with the aim of verifying how semantic classes obtained through pattern-based corpus analysis differ from categories which are defined on the basis of axiomatization. Also, I will discuss the opportunity to enhance the taxonomic structuring of our list using the OntoClean methodology (Guarino and Welty, 2009), which was also exploited for the development of DOLCE. Finally, I will highlight the mutual benefit of the experiment, and confirm the advantages of keeping the lexical level separated from the ontological level in language resource building (Oltramari et al. 2013).
It is a truism that meaning depends on context. Corpus evidence now shows us that normal contexts can be summarised and indeed quantified, while the creative exploitations of normal contexts by ordinary language users far exceed anything dreamed up in speculative linguistic theory. Human linguistic behaviour is indeed rule-governed, but in recent years, corpus analysis (e.g. Hanks 2013) has shown that there is not just a single monolithic system of rules: instead, language use is governed by two interlinked systems: one set of rules governing normal, idiomatic uses of words and another set of rules governing how we exploit those norms creatively. Types of creative exploitation include (among others):
English. Several textual inference tasks rely on kernel-based learning. In particular Tree Kernels (TKs) proved to be suitable to the modeling of syntactic and semantic similarity between linguistic instances. In order to generalize the meaning of linguistic phrases, Distributional Compositional Semantics (DCS) methods have been defined to compositionally combine the meaning of words in semantic spaces. However, TKs still do not account for compositionality. A novel kernel, i.e. the Compositional Tree Kernel, is presented integrating DCS operators in the TK estimation. The evaluation over Question Classification and Metaphor Detection shows the contribution of semantic compositions w.r.t. traditional TKs. Italiano. Sono numerosi i problemi di interpretazione del testo che beneficiano dall’applicazione di metodi di apprendimento automatico basato su funzioni kernel. In particolare, i Tree Kernel (TK) sono applicati alla modellazione di metriche di similarita sintattica e semantica tra espressioni linguistiche. Allo scopo di generalizzare i significati legati a sintagmi complessi, i metodi di Distributional Compositional Semantics combinano algebricamente i vettori associati agli elementi lessicali costituenti. Ad oggi i modelli di TK non esprimono criteri di composizionalita. In questo lavoro dimostriamo il beneficio di modelli di composizionalita applicati ai TK, in problemi di Question Classification e Metaphor Detection.
Comparisons are phrases that express the likeness of two entities. They are usually marked by linguistic patterns, among which the most discussed are X is like Y and as X as Y . We propose a simple slot-based dependency-driven description of such patterns that refines the phrase structure approach of Niculae and Yaneva (2013). We introduce a simple similarity-based approach that proves useful for measuring the degree of figurativeness of a comparison and therefore in simile (figu-rative comparison) identification. We pro-pose an evaluation method for this task on the VUAMC metaphor corpus.
Almost any word in natural language has a great potential of expressing different meanings. However, in certain contexts, this potential is limited up to the point that one and only one sense is possible. When this happens, we are not dealing with an individual phenomenon, but, rather, all the words in that context have their own meaning potential limited. In this talk, we properly define such meaning restricting contexts, analyze their properties and propose an automatic procedure for their identification in large corpora. We show that these contexts are patternable and that the words are completely disambiguated. We therefore call such contexts sense discriminative patterns (SDP). By comparing minimally different SDPs, we discover a set of lexical semantic features that are used in devising a learning algorithm. The form of patterns is regular, they are generated by a finite state automaton. Inducing the form of the grammar from annotated examples and finding the right generalization level is done using Angluin Algorithm. The patterns contain the syntactic and lexical information which is relevant for sense disambiguation, so they are SDPs. The patterns are minimally self-sufficient, thus the senses of the words matched by a pattern are in mutual disambiguation relationship. The disambiguation process of the meanings of all slots is sequential, identifying the meaning of one slots leads to the identification of the meaning of all slots. We call this relationship between the senses of the words which are caught in a pattern, chain clarifying relationship, CCR. The main problem that needs to be addressed is the fact that pattern acquisition is very sensitive to errors. On the basis of the PAC-learning technique, we have developed a technique that produces an approximately correct grammar, having a high probability to be correct in spite of the noisy examples. We restrict the type of patterns that could be learned and we construct hypotheses which are statistically tested against large sample using the statistical query model for learning new patterns. We will also present the applications of SDPs to various meaning related natural language processing tasks, like word sense disambiguation, textual entailment and meaning preserving translation.
Distributional models assume that the contexts of a linguistic unit (such as a word, a multi-word expression, a phrase, a sentence, etc.) provide information about the meaning of the linguistic unit (Firth, 1957; Harris, 1968). They have been widely applied in data-intensive lexical semantics (among other areas), and proven successful in diverse research issues, such as the representation and disambiguation of word senses (Schütze, 1998; McCarthy et al., 2004; Springorum et al., 2013), selectional preference modelling (Herdagdelen and Baroni, 2009; Erk et al., 2010; Schulte im Walde, 2010), the compositionality of compounds and phrases (McCarthy et al., 2003; Reddy et al., 2011; Boleda et al., 2013), or as a general framework across semantic tasks (’distributional memory’, cf. Baroni and Lenci, 2010; Pado and Utt, 2012), to name just a few examples. While it is clear that distributional knowledge does not cover all the cognitive knowledge humans possess with respect to word meaning (Marconi, 1997; Lenci, 2008), distributional models are very attractive, as the underlying parameters are accessible from even low-level annotated corpus data. We are thus interested in maximising the benefit of distributional information for lexical semantics, by exploring the meaning and the potential of comparatively simple distributional models. In this respect, this talk will present four case studies on semantic relatedness tasks that demonstrate the potential and the limits of distributional models.
Similarity is at the core of scientific inquiry in general and is one of the basic functionalities in Natural Language Processing (NLP) in particular. To arrive at generalizations across different phenomena, we need to recognize patterns of similarity, or divergence, to make scientific claims. Semantic textual similarity plays a significant role in NLP research both directly and indirectly. For example, for document summarization, we need to compress redundant information which requires identifying where the text is similar; for question answering, we need to recognize the similarity between the questions and the answers; textual similarity is an important component of an entailment system; evaluating machine translation (MT) output relies on calculating the similarity between the system’s output and some reference gold translations; textual generation technology benefits from sentence similarity by generating different expressions. In this talk, I will address the problem of textual semantic similarity. We have run 2 major tasks of STS over the span of two years within the context of Semeval in 2012 and *SEM shared task in 2013. The task to date is one of the most successful to be carried out within our community by virtue of being quite popular. I will share with you the details of the task, some interesting insights into the scientific merits of this enterprise and lessons learned. Finally I will share some thoughts on the future.
This work describes the evaluations of three different approaches, Lexical Match, Sense Similarity based on Personalized Page Rank, and Semantic Match based on Shallow Frame Structures, for word sense alignment of verbs between two Italian lexical-semantic resources, MultiWordNet and the Senso Comune Lexicon. The results obtained are quite satisfying with a final F1 score of 0.47 when merging together Lexical Match and Sense Similarity.
This paper defines a representation of universal quantification within distributional semantic space. We propose a discourse-internal approach to the meaning of limited instances of every, highlighting the possibilities and limitations of doing textual logic in a purely distributional framework.
1 We propose and implement an alternative source of contextual features for word similarity detection based on the notion of lexicogrammatical construction. On the assumption that selectional restrictions provide indicators of the semantic similarity of words attested in selected positions, we extend the notion of selection beyond that of single selecting heads to multiword constructions exerting selectional preferences. Our model of 92 million cross-indexed hybrid n-grams (serving as our machine-tractable proxy for constructions) extracted from BNC provides the source of contextual features. We compare results with those of a grammatical dependency approach (Lin 1998), testing both against WordNetbased similarity rankings (Lin 1998; Resnik 1995). Averaged over the entire set of target nouns and 10-best candidate similar words, Lin’s approach gives overall similarity results closer to WordNet rankings than the constructional approach does, while the constructional approach overtakes Lin’s in approximating WordNet similarity for target nouns with a frequency over 3000. While this suggests feature sparseness for constructions that resolves with higher frequency nouns, constructions as shared contextual features render a much higher yield in similarity performance in approximating WordNet similarity than grammatical relations do. We examine some cases in detail showing the sorts of similarity detected by a constructional approach that are undetected by a grammatical relations approach or by WordNet or both and thus overlooked in benchmark eval-
In the knowledge representation and reasoning research area, argumentation theory aims at representing and reasoning over information items called arguments. In everyday life, arguments are reasons to believe and reasons to act, and they are usually expressed in natural language. Even if ad-hoc natural language examples are often provided in argumentation theory works, no automated processing of such natural language arguments is carried out, making it impossible to exploit the results of this research area in real world scenarios. In this paper, we propose to adopt textual entailment to address this issue. In particular, we discuss and evaluate, on a sample of natural language arguments extracted from Debatepedia, the support and attack relations among arguments in bipo- lar abstract argumentation with respect to the more specific notions of textual entailment and contradiction.
Textual Inference requires the analyzing text at multiple levels as well as to disambiguating it and grounding it in knowledge resources to facilitate knowledge driven reasoning. Computational approaches to these problems in Natural Language Understanding and Information Extraction are often modeled as structured predictions predictions that involve assigning values to sets of interdependent variables. Over the last few years, one of the most successful approaches to studying these problems involves Constrained Conditional Models (CCMs), an Integer Learning Programming formulation that augments probabilistic models with declarative constraints as a way to support such decisions. I will focus on exemplifying this framework in the context of developing better semantic analysis of sentences Extended Semantic Role Labeling and the task of Wikification identifying concepts and entities in text and disambiguating them into Wikipedia or other knowledge bases.
This article proposes a new approach to verb classification based on Semantic Types selected in corpus-based verb patterns. This work !"#$%& '(& )#(*%+%& ,-.'"/& '0& 1'"2%& #(!& 3xploitations (Hanks 2013) and applies Corpus Pattern Analysis to a subset of verbs from Lev4(+%&56'4%'(+&78#%%9&4(78:!4(;&<."=%&%:7-&#%& hang and stab. These patterns are taken from the Pattern Dictionary of English Verbs, which aims at recording prototypical phraseological patterns for the most frequent verbs of English using the British National Corpus.