IsoQuest used its commercial software product, NetOwl Extractor, for the MUC-7 Named Entity task. The product consists of a high-speed C engine that analyzes text based on a configuration file containing a pattern rule base and lexicon. IsoQuest used the NameTag Configuration to recognize proper names and other key phrases in text, and mapped the product’s extraction tags to the MUC-7 NE tags. NetOwl Extractor provides access to the extracted information through a flexible API, and IsoQuest used a small application program to process the documents and write the SGML output.
SRA used the combination of two systems for the MUC-6 tasks: NameTag ™, a commercial software product that recognizes proper names and other key phrases in text; and HASTEN , an experimental text extraction system that has been under development for only one year. For the Named Entity task, SRA adapted a subset of NameTag's capabilities to the MUC-6 specification. For the Template Element task, SRA fed the full results of NameTag into HASTEN , which performed additional processing to extract and generate the organization and person templates. For the Scenario Template task, SRA fed NameTag's results into HASTEN , which used its full extraction capabilities to extract and generate the management succession templates. Figure 1 illustrates the contribution of each system to the MUC-6 tasks. Due to the relative complexity of the scenario template task, this paper will focus on HASTEN and the experimental results for scenario extraction, and will provide brief descriptions of NameTag and the other tasks.
The GE-CMU team is developing the TIPSTER/SHOGUN system under the government-sponsored TIPSTER program, which aims to advance coverage, accuracy, and portability in text interpretation. The system will soon be tested on Japanese and English news stories in two new domains. MUC-4 served as the first substantial test of the combined system. Because the SHOGUN system takes advantage of most of the components of the GE NLTOOLSET except for the parser, this paper supplements the NLTOOLSET system description by explaining the relationship between the two systems and comparing their performance on the examples from MUC-4.
This paper reports on the results of the adjunct test performed by GE for the MUC-4 evaluation of text processing systems. In this test, we evaluated the effect of an object-oriented template design and associated matching conditions on the scores. The results indicate that the current MUC-4 "flat" templade design with cross-references closely approximates a true object-oriented design. However the object-oriented design allows for additional performance data to be calculated, facilitating diagnosis.
Discourse comprises those phenomena that usually do not arise when processing a single sentence. It appears to be the most difficult and probably the least understood aspect of automated message understanding. Five out of fifteen sites on a MUC-3 survey listed discourse as their main weakness and an area in which to concentrate future research. Virtually all systems presented here take a sentence-by-sentence approach to text understanding. Parsing and domain-dependent interpretation of sentences or sentence fragments (usually the latter) are followed by modules that attempt to connect these interpretations into a coherent whole. This paper gives an overview of the modules that make the transition from the interpretation of sentences to the interpretation of the text that contains these sentences. Systems presented in this paper exhibit various degrees of the following discourse understanding capabilities:• identifying portions of text that describe different domain events; this includes the capability of recognizing a single event and the capability of distinguishing multiple events;• resolving references:- pronoun references, e.g., finding the referent of It in the sentence It took place this morning,- proper name references, e.g., understanding that Luis Galan may be referred to as Senator Galan;- definite references, e.g., deciding what is the referent for The attack in the sentence The attack look us by surprise.• discourse representation : representation at the message level.
This paper reports on the GE NLTooLSET customization effort for MUC-3, and analyzes the results of the TST2 run. Although our own tests had shown steady improvement between TST1 and TST2, our official scores on TST2 were lower than on TST1. The analysis of this unexpected result explains some of the details of the MUC-3 test, and we propose ways of looking at the scores to distinguish different aspects of system performance.
The GE NLTooLSET is a set of text interpretation tools designed to be easily adapted to new domains. This report summarizes the system and its performance on the MUC-3 task.
A generic natural language system, without modification, can effectively analyze an arbitrary input at least to the level of word sense tagging. Considerable research has addressed the transportability of natural language systems, but not generic text processing capabilities. For example, previous DARPA-sponsored work [1, 2] produced transportable interfaces to database systems. Each new application of these interfaces generally required modifications to lexicons, new semantic knowledge bases, and other specialized features. The most that natural language text processing systems have accomplished has been the parsing of arbitrary text, without any real semantic analysis.
Douglas E. Appelt合作论文数Artificial Intelligence Center1