This is the data required to build the paper "An Ensemble Method to Build High-Quality Word Embeddings", by Robyn Speer and Joshua Chin. The input data itself comes from: ConceptNet 5.4, which contains data from Wiktionary, WordNet, and many contributors to Open Mind Common Sense projects, edited by Robyn Speer GloVe, by Jeffrey Pennington, Richard Socher, and Christopher Manning word2vec, by Tomas Mikolov and Google Research PPDB, by Juri Ganitkevitch, Benjamin Van Durme, and Chris Callison-Burch
Machine learning about language can be improved by supplying it with specific knowledge and sources of external information. We present here a new version of the linked open data resource ConceptNet that is particularly well suited to be used with modern NLP techniques such as word embeddings. ConceptNet is a knowledge graph that connects words and phrases of natural language with labeled edges. Its knowledge is collected from many sources that include expert-created resources, crowd-sourcing, and games with a purpose. It is designed to represent the general knowledge involved in understanding language, improving natural language applications by allowing the application to better understand the meanings behind the words people use. When ConceptNet is combined with word embeddings acquired from distributional semantics (such as word2vec), it provides applications with understanding that they would not acquire from distributional semantics alone, nor from narrower resources such as WordNet or DBPedia. We demonstrate this with state-of-the-art results on intrinsic evaluations of word relatedness that translate into improvements on applications of word vectors, including solving SAT-style analogies.
A currently successful approach to computational semantics is to represent words as embeddings in a machine-learned vector space. We present an ensemble method that combines embeddings produced by GloVe (Pennington et al., 2014) and word2vec (Mikolov et al., 2013) with structured knowledge from the semantic networks ConceptNet (Speer and Havasi, 2012) and PPDB (Ganitkevitch et al., 2013), merging their information into a common representation with a large, multilingual vocabulary. The embeddings it produces achieve state-of-the-art performance on many word-similarity evaluations. Its score of $\rho = .596$ on an evaluation of rare words (Luong et al., 2013) is 16% higher than the previous best known system.
The Narratarium Colorizer device receives either keyboard input or speech recognition input and uses natural language processing to extract key terms. The terms are queried for in a knowledge base of words and associated colors, created by leveraging the Open Mind Common Sense database and ConceptNet. The system outputs a continually changing color display, which is projected uniformly throughout the room using a custom designed curved mirror projection system.
ConceptNet is a knowledge representation project, providing a large semantic graph that describes general human knowledge and how it is expressed in natural language. Here we present the latest iteration, ConceptNet 5, with a focus on its fundamental design decisions and ways to interoperate with it.
Modeling Verb Lexicalization Biases using Hierarchical Bayesian Models Catherine Havasi (havasi@media.mit.edu) MIT Media Lab, 20 Ames Street Cambridge, MA 02139 USA Robert Speer (rspeer@mit.edu) MIT Media Lab, 20 Ames Street Cambridge, MA 02139 USA Abstract a single example (Gentner & Boroditsky, 2001), and they can even learn words for events they are unable to observe (Landau & Gleitman, 1985). Two faster and more noise-resistant strategies have been hypothesized by researchers. One is syntactic bootstrapping (Gleitman, 1990). In this theory, the syntactic frame of the verb is used to constrain hypotheses to those which makes sense in the given frame and are similar to known verbs with similar frames. In the manner/path example given earlier, you would be more likely to think the meaning of the novel verb was related to its motion if you had heard the semantically rich fame “Jesse gorped the frisbee to Edison.” Another hypothesis is that we are able to quickly learn words from few examples because we rely on our learned lexicalization biases about the meanings of words (Gentner & Boroditsky, 2001). Learners select word meanings that align with the features that are dominant in the learner’s native language (Naigles, 1990), indicating that language learners observe general features of the meanings of other words and apply them to new words as well. Modern evidence suggests that children use a combination of these strategies (Papafragou & Selimis, 2010). But how are these biases learned and regulated? In this paper, we explore the possibility that biases for certain components of meaning are associated with language and semantic frame. These biases represent examples of Bayesian overhypotheses about what a word is likely to mean, and these overhypotheses can themselves be learned from examples (Kemp, Perfors, & Tenenbaum, 2007). The overhypotheses can depend on observable features such as whether the referent is animate (Smith, Jones, Landau, Gershkoff-Stowe, & Samuelson, 2002), the syntactic patterns in which the word appears (Cifuentes- F´erez & Gentner, 2006), or known lexical relations to other words (Pustejovsky, 1998). Return momentarily to the analogous results for shape bi- ases in nouns — that early nouns that children learn tend to be easily clustered by the shape of their referents. In order to model this bias, it was postulated that children learn a “second- order generalization” that objects are often categorized by their shape (Samuelson & Smith, 1999). Smith et al. demonstrated this generalization by teaching 17-19 month old children a precocious shape bias (Smith et al., 2002). Kemp, Perfors, and Tenenbaum explained this kind of learning using a hierarchical Bayesian model, which could learn both base meanings and overhypotheses simultaneously (Kemp et al., 2007) and cases The expression of motion verbs differs between languages. The path of motion, such as crossing or entering, is more promi- nently featured in path-based languages such as Spanish than in manner-based languages such as English. Here, we revisit the data from a study on manner and path biases in verb lexi- calization (Havasi & Snedeker, 2004), and create a hierarchical Baysian computational model to further explore, verify, and define these biases. With this model, we can discover the large differences in subjects’ pre-existing manner and path biases that depend on the syntactic frame in which new verbs appear, as well as a difference in the learning rate between English speakers taking the experiment in English and bilingual Span- ish speakers taking the experiment in Spanish. We can also use the model to predict the responses of subjects in the experiment with more accuracy than before. Keywords: verb learning; bayesian modeling; hierarchical Bayes modeling; manner and path verbs Linguistic lexicalization biases People have the ability to intuit the meaning of a new verb after hearing it used to describe just a single event. In the case of a novel verb, there are many potential hypotheses of the verb’s meaning which may be consistent with the event witnessed. Suppose you hear a novel verb, such as “gorp”, being used to describe an event in which Jesse throws a frisbee across a field to , her dog. The verb could refer to Jesse throwing the frisbee, the frisbee’s motion as it glides across the field, the frisbee’s traverse of the field, or Edison’s act of catching the frisbee. To understand which aspect of the action the verb refers to, you must use situational clues and background knowledge. When one encounters a new object noun, one encounters the same ambiguity in meaning. In practice, languages sys- tematically favor a few different characteristics such as com- mon ancestry or base level category (Nelson, 1973) for noun meanings which is often indicated by shape. However, event categorization tends to be flexible across languages and even with a language (Talmy, 1975). A motion verb, for example, could easily refer to the manner, cause, or path of the motion with no universal preference across languages (Aske, 1989; Berman & Slobin, 1994; Jackendoff, 1990). Given the plethora of possible referents for a novel verb, how do children learn verb meanings? One solution would be to observe, over several examples, that certain semantic features seem to always be present and are thus associated with the verb’s meaning. However, this would require too much data to match the way that children learn words; children can often determine the relevant aspect of a word’s meaning from
We present Luminoso, a tool that helps researchers to visualize and understand a dimensionality-reduced semantic space based on textual information by exploring it interactively. It streamlines the process of creating such a space by taking input from a directory of text documents, and optionally including common-sense background information. This interface is useful for interactively discovering trends in a text corpus, such as free-text responses to a survey. We discuss a case study about restaurant reviews to show how Luminoso can be used for opinion mining.
Singular value decomposition (SVD) is a powerful technique for finding similarities and patterns in large data sets. SVD has applications in text analysis, bioinformatics, and recommender systems, and in particular was used in many of the top entries to the Netflix Challenge. It can also help generalize and learn from knowledge represented in a sparse semantic network.
Open Mind Common Sense (OMCS) is a freely available crowd-sourced knowledge base of natural language statements about the world. The goal of Open Mind Common Sense is to provide intuition to AI systems and applications by giving them access to a broad collection of basic information and the computational tools to work with this data. For our system demo, we will be presenting three aspects of the OMCS project: the OMCS knowledge base, the Concept-Net semantic network (Liu and Singh 2004) (Havasi, Speer, and Alonso 2007), and the AnalogySpace algorithm (Speer, Havasi, and Lieberman 2008) which deals well with noisy, user-contributed data.
Increasingly, we need to computationally understand real-time streams of information in places such as news feeds, speech streams, and social networks. We present Streaming AnalogySpace, an efficient technique that discovers correlations in and makes predictions about sparse natural-language data that arrives in a real-time stream. AnalogySpace is a noise-resistant PCA-based inference technique designed for use with collaboratively collected common sense knowledge and semantic networks. Streaming AnalogySpace advances this work by computing it incrementally using CCIPCA, and keeping a dense cache of recently-used features to efficiently represent a sparse and open domain. We show that Streaming AnalogySpace converges to the results of standard AnalogySpace, and verify this by evaluating its accuracy empirically on common-sense predictions against standard AnalogySpace.
Verbosity, a ``game with a purpose'', uses the collective activity of people playing an Internet word game as a body of common sense knowledge. One purpose of Verbosity has always been to provide large quantities of input for a common sense knowledge base, and we have now achieved this purpose by connecting it to the Open Mind Common Sense (OMCS) project. Verbosity now serves as a way to contribute to OMCS in addition to being an entertaining game in its own right. Here, we explain the process of filtering and adapting Verbosity's data for use in OMCS, showing that the results are of a quality comparable to OMCS's existing data, and discuss how this informs the future development of games for common sense.
Today millions of web-users express their opinions about many topics through blogs, wikis, fora, chats and social networks. For sectors such as e-commerce and e-tourism, it is very useful to automatically analyze the huge amount of social information available on the Web, but the extremely unstructured nature of these contents makes it a difficult task. SenticNet is a publicly available resource for opinion mining built exploiting AI and Semantic Web techniques. It uses dimensionality reduction to infer the polarity of common sense concepts and hence provide a public resource for mining opinions from natural language text at a semantic, rather than just syntactic, level.
Coarse word sense disambiguation (WSD) is an NLP task that is both important and practical: it aims to distinguish senses of a word that have very different meanings, while avoiding the complexity that comes from trying to finely distinguish every possible word sense. Reasoning techniques that make use of common sense information can help to solve the WSD problem by taking word meaning and context into account. We have created a system for coarse word sense disambiguation using blending, a common sense reasoning technique, to combine information from SemCor, WordNet, ConceptNet and Extended WordNet. Within that space, a correct sense is suggested based on the similarity of the ambiguous word to each of its possible word senses. The general blending-based system performed well at the task, achieving an f-score of 80.8\% on the 2007 SemEval Coarse Word Sense Disambiguation task.
We present a game-based interface for acquiring common sense knowledge. In addition to being interactive and entertaining, our interface guides the knowledge acquisition process to learn about the most salient characteristics of a particular concept. We use statistical classification methods to discover the most informative characteristics in the Open Mind Common Sense knowledge base, and use these characteristics to play a game of 20 Questions with the user. Our interface also allows users to enter knowledge more quickly than a more traditional knowledge-acquisition interface. An evaluation showed that users enjoyed the game and that it increased the speed of knowledge acquisition.
We are interested in the problem of reasoning over very large common sense knowledge bases. When such a knowledge base contains noisy and subjective data, it is important to have a method for making rough conclusions based on similarities and tendencies, rather than absolute truth. We present Analogy Space, which accomplishes this by forming the analogical closure of a semantic network through dimensionality reduction. It self-organizes concepts around dimensions that can be seen as making distinctions such as "good vs. bad" or "easy vs. hard", and generalizes its knowledge by judging where concepts lie along these dimensions. An evaluation demonstrates that users often agree with the predicted knowledge, and that its accuracy is an improvement over previous techniques.