While the celebrated Word2Vec technique yields semantically rich representations for individual words, there has been relatively less success in extending to generate unsupervised sentences or documents embeddings. Recent work has demonstrated that a distance measure between documents called Word Mover's Distance (WMD) that aligns semantically similar words, yields unprecedented KNN classification accuracy. However, WMD is expensive to compute, and it is hard to extend its use beyond a KNN classifier. In this paper, we propose the Word Mover's Embedding (WME), a novel approach to building an unsupervised document (sentence) embedding from pre-trained word embeddings. In our experiments on 9 benchmark text classification datasets and 22 textual similarity tasks, the proposed technique consistently matches or outperforms state-of-the-art techniques, with significantly higher accuracy on problems of short length.
In this paper, we discuss Semantic Construction Grammar (SCG), a system developed over the past several years to facilitate translation between natural language and logical representations. Crucially, SCG is designed to support a variety of different methods of representation, ranging from those that are fairly close to the NL structure (e.g. so-called "logical forms"), to those that are quite different from the NL structure, with higher-order and high-arity relations. Semantic constraints and checks on representations are integral to the process of NL understanding with SCG, and are easily carried out due to the SCG's integration with Cyc's Knowledge Base and inference engine [1] , [2].
The true promise of the Web canpsilat be realised by the lone programmers and simple applications of Web 1.0; it canpsilat even be realised by the advanced interfaces and hordes of contributing users of Web 2.0 - it can be realised by individuals and groups of humans collaborating with individual and cloud connected computers. Thatpsilas Web 3.0. What will it take to make computers into effective collaborators? It will take heterogeneous, ubiquitous reasoning at massive scale - the aim of the EU funded LarKC project; It will take semantically rich shared representations - the aim of OpenCyc (and other, linked projects); it will take more sophisticated reasoning, including probabilistic and contextual reasoning, and it will take sophisticated, social, human-computer and computer-interfaces. In this talk, Ipsilall focus on Cycorp Europe and our effort in the LarKC project, and describe how we hope it will start to tie our work, and the work of others, together to produce a truly knowledgeable, collaborative, intelligent Web.
The Cyc project is predicated on the idea that effective machine learning depends on having a core of knowledge that provides a context for novel learned information - what is known informally as "common sense." Over the last twenty years, a sufficient core of common sense knowledge has been entered into Cyc to allow it to begin effectively and flexibly supporting its most important task: increasing its own store of world knowledge. In this paper, we present initial work on a method of using a combination of Cyc and the World Wide Web, accessed via Google, to assist in entering knowledge into Cyc. The long-term goal is automating the process of building a consistent, formalized representation of the world in the Cyc knowledge base via machine learning. We present preliminary results of this work and describe how we expect the knowledge acquisition process to become more accurate, faster, and more automated in the future.
CycSecure™ is a network risk assessment and network monitoring application that relies on knowledge-based artificial intelligence technologies to improve on traditional network vulnerability assessment. CycSecure integrates public reports of software faults from online databases, data gathered automatically from computers on a network and hand-ontologized information about computers and computer networks. This information is stored in the Cyc® knowledge base (KB) and reasoned about by the Cyc inference engine and planner to provide detailed analyses of the security (and vulnerability) of networks.
In theory, speech recognition technology can make any spoken words in video or audio media usable for text indexing, search and retrieval. This article describes the News-on-Demand application created within the InformediaTM Digital Video Library project and discusses how speech recognition is used in transcript creation from video, alignment with closed-captioned transcripts, audio paragraph segmentation and a spoken query interface. Speech recognition accuracy varies dramatically depending on the quality and type of data used. Informal information retrieval test show that reasonable recall and precision can be obtained with only moderate speech recognition accuracy. 1. INFORMEDIA: NEWS-ON-DEMAND The InformediaTM digital video library project [1,2,3,4] at Carnegie Mellon University is creating a digital library of text, images, videos and audio data available for full content search and retrieval. News-on-Demand is an application within Informedia that monitors news from TV, radio and text sources and allows the user to retrieve news stories of interest. A compelling application of the Informedia project is the indexing and retrieval of television, radio and text news. The Informedia: News-on-Demand application [5,6] is an innovative example of indexing and searching broadcast news video and news radio material by text content. News-on-Demand is a fullyautomatic system that monitors TV, radio and text news and allows selective retrieval of news stories based on spoken queries. The user may choose among the retrieved stories and play back news stories of interest. The system runs on a Pentium PC using MPEGI video compression. Speech recognition is done on a separate platform using the Sphinx-II continuous speech recognition system [7]. The News-on-Demand application forces us to consider the limits of what can be done automatically and in limited time. News events happen daily and it is not feasible to process, segment and label news through manual or “human-assisted” methods. Timeliness of the library information is important, as is the ability to continuously update the contents. Thus we are forced to fully exploit the potential of computer speech recognition without the benefit of human corrections and editing. INFORMEDIATM: NEWS-ON-DEMAND EXPERIMENTS IN SPEECH RECOGNITION Howard D. Wactlar, Alexander G. Hauptmann and Michael J. Witbrock School of Computer Science Carnegie Mellon University Pittsburgh, PA 15213-3890 Even though our work is centered around processing news stories from TV broadcasts, the system exemplifies an approach that can make any video, audio or text data accessible. Similar methods can help to index and search other streamed multi-media data by content in other Informedia applications. Other attempts at solutions have been obliged to restrict the data to only text material, as found in most news databases. Video-ondemand allows a user to select (and pay for) a complete program, but does not allow selective retrieval. The closest approximation to News-on-Demand can be found in the “CNN-AT-WORK” system offered to businesses by a CNN/Intel venture. At the heart of the CNN-AT-WORK solution is a digitizer that encodes the video into the INDEO compression format and transmits it to workstations over a local area network. Users can store headlines together with video clips and retrieve them at a later date. However, this service depends entirely on the separately transmitted “headlines” and does not include other news sources. In addition, CNN-AT-WORK does not feature an integrated multimodal query interface [8]. Preliminary investigations on the use of speech recognition to analyze a news story were made by [9]. Without a powerful speech recognizer, their approach used a phonetic engine that transformed the spoken text into an (errorful) phoneme string. The query was also transformed into a phoneme string and the database searched for the best approximate match. Errors in recognition, as well as word prefix and suffix differences did not severely affect the system since they scattered equally over all documents and wellmatching search scores dominate the retrieval. Another news processing systems that includes video materials is the MEDUSA system [10]. The MEDUSA news broadcast application can digitize and record news video and teletext, which is equivalent to closed-captions. Instead of segmenting the news into stories, the system uses overlapping windows of adjacent text lines for indexing and retrieval. During retrieval the system responds to typed requests returning an ordered list of the most relevant news broadcasts. Query words are stripped of suffixes before search and the relevance ranking takes word frequency in the segment and over all the corpus into account, as well as the ability of words to discriminate between stories. Within a news broadcast, it is up to the user to select and play a region using information given by the system about the location of the matched keywords. The focus of MEDUSA is in the system architecture and the information retrieval component. No image processing and no speech recognition is performed. 1.1. Component Technologies There are three broad categories of technologies we can bring This material is based upon work supported by the National Science Foundation under Cooperative Agreement No. IRI-9411299. Any opinions, findings, and conclusions or recommendations expressed in this material are those of the author(s) and do not necessarily reflect the views of the National Science Foundation. to bear to create and search a digital video library from broadcast video and audio materials [11]: Text processing looks at the textual (ASCII) representation of the words that were spoken, as well as other text annotations. These may be derived from the transcript, from the production notes or from the closed-captioning that might be available. Text analysis can work on an existing transcript to help segment the text into paragraphs [12]. An analysis of keyword prominence allows us to identify important sections in the transcript [13]. Other more sophisticated language based criteria are under investigation. We currently use two main techniques for text analysis: 1. If we have a complete time aligned transcript available from the closed-captioning or through a human-generated transcription, we can exploit natural “structural” text markers such as punctuation to identify news story boundaries 2. To identify and rank the contents of one news segment, we use the well-known technique of TF/IDF (term frequency/ inverse document frequency) to identify critical keywords and their relative importance for the video document[13]. Image analysis looks at the images in the video portion of the MPEG stream. This analysis is primarily used for the identification of scene breaks and to select static frame icons that are representative of a scene. Image statistics are computed for Library Creation Library Exploration Offline Online TV News Radio News Text News
The Informedia Digital Video Library Project at Carnegie Mellon University is making large corpora of video and audio data available for full content retrieval by integrating natural language understanding, image processing, speech recognition and information retrieval. Information retrieval of from corpora of speech recognition output is critical to the project’s success. In this paper, we outline how this output is combined information from other modalities to produce a successful interface. We then describe experiments that compare retrieval effectiveness on spoken and text documents and investigate the sources of retrieval errors on the former. Finally we investigate how improvements in speech recognizer accuracy may affect retrieval, and whether retrieval will still be effective when larger spoken corpora are indexed.
For the huge amounts of audio and video material that could usefully be included in digital libraries, the cost of producing human-generated annotations and meta-data is prohibitive. In the Informedia Digital Video Library, the production of meta-data supporting the library interface is automated using techniques from Artificial Intelligence (AI). By applying speech recognition, natural language processing and image analysis, the interface helps users locate the information they want and navigate or browse the digital video library more effectively. Specific AI-based interface components include automatic titles, filmstrips, video skims, word location marking and representative frames for shots.
We describe the Informediatm News-on-Demand system. News-on-Demand is an innovative example of indexing and searching broadcast video and audio material by text content. The fully-automatic system monitors TV news and allows selective retrieval of news items based on spoken queries. The user then plays the appropriate video "paragraph". The system runs on a Pentium PC using MPEG-I video compression and the Sphinx-II continuous speech recognition system [6].
We describe a method for automatic discovery of generalizations about data residing in the Cyc Knowledge Base (KB). Data is exported from the Cyc KB into a feature representation, a decision tree classifier is used to discover generalizations in that data, and the learned generalizations are imported back into the KB using probabilistic vocabulary. Our case study is the Whodunit problem: guessing the perpetrator of a terrorism event. We compare two methods of building decision-tree models from which rules can be imported. A single-multiclass model involves a single decision tree with multiple possible values for the predicted variable. A multiple-binary model involves multiple decision trees with binary predicted values, predicting, for a given instance i and class c, whether i has class c. We show that the single multiclass model is more prone to overgeneralization on larger datasets, due to its greater reliance on the closed world assumption. We conclude that for the purpose of rule extraction into a large knowledge base, it is better to use multiple binary decision trees.
Abstract In theory, speech recognition technology can make any spoken words in video or audio media usable for text indexing, search and retrieval. This article describes the News-on-Demand application created within the Informedia ,Digital Video Library project and discusses how speech recognition is used in transcript creation from video, alignment with closed-captionedtranscripts, audio paragraph segmentation and a spoken query interface. Speech recognition accuracy,varies dramatically depending,on the quality and type of data used. Informal information retrieval tests show that reasonable recall and precision can be obtained with only moderate,speech recognition accuracy. This paper is based on work supported by the National Science Foundation, DARPA and NASA under NSF Cooperative
Pierre Flener合作论文数Uppsala University ;Computing Science Division;Department of Information Technology 1