As evaluations of computational linguistics technology progress toward higher-level interpretation tasks, the problem of determining alignments between system responses and answer key entries may become less straightforward. We present an extensive analysis of the alignment procedure used in the MUC-6 evaluation of information extraction technology, which reveals effects that interfere with the stated goals of the evaluation. These effects are shown to be pervasive enough that they have the potential to adversely impact the technology development process. These results argue strongly for the use of accurate alignment criteria in natural language evaluations, and for maintaining the independence of alignment criteria and mechanisms used to calculate scores.
Collaboration is at the heart of many activities required for effective homeland security, from intelligence analysis to policy formation. We are exploring new approaches to facilitating effective collaboration that remove or reduce common barriers and that exploit opportunities to encourage more effective collaboration, including transcending the cognitive biases of the participants. In order to evaluate our approaches we are developing Angler, a web-services tool that supports collaboration among participants on some focus topic. Several challenges arise in helping participants manage their contributions. A semantic index over the participant contributions is used to address these challenges. Angler as a Collaboration Tool Collaboration is at the heart of many activities required for effective homeland security, from intelligence analysis to policy formation. We are exploring new approaches to facilitating effective collaboration that remove or reduce its common barriers and that exploit opportunities to encourage more effective collaboration, including transcending the cognitive biases of the collaboration participants and expanding their joint cognitive vision. To such effect we are building Angler, a web-services tool for supporting collaboration on some focus topic (Rodriguez et al. 2005). Angler facilitates effective collaboration by overcoming some of the common barriers to collaboration. It does so allowing for participants to be geographically distributed and allowing for asynchronous collaboration. This differentiates the Angler process from similar business processes that normally take place in a synchronous, face-to-face manner. Angler facilitates cognitive expansion during collaboration by using divergent (such as brainstorming) and convergent (such as clustering or ranking) thinking techniques. Angler is organized around the idea of virtual (possibly hierarchical) workshops. Workshops are organized by a facilitator that brings a group of people together to accomplish a knowledge task. Simple workshops usually start in a brainstorming phase and then go through other phases, such as clustering and ranking. The facilitator for the workshop will enforce and manage the timelines. As more and more workshops are stored in the Knowledge Base, a Corporate Memory is formed. In the brainstorming phase, collaboration participants are asked to contribute thoughts (or ideas) to answer a focused request from a facilitator. Participants submit contributions as brief statements on some aspect of the overall focus topic being considered; each such contributed thought is authored as a small textual document that can be reviewed by other participants. Other participants’ thoughts will be incrementally disclosed to a participant, so that he or she can think independently and can benefit from others’ ideas. Angler provides a convenient interface for participants to author such thoughts, review and respond to the thoughts of other participants, and organize the growing set of shared contributions in meaningful ways After sufficient input and review, the facilitator moves the workshop into the clustering phase. The participants begin to organize the thoughts into various clusters. Each cluster suggests one or more candidate ways to coalesce individual thoughts into a coherent theme, and the participants benefit by considering the emerging themes, for example, for scenario planning (Schwartz 1991). This process promotes a rich interchange of ideas and perspectives while removing some of the classic barriers to conventional collaboration, such as the requirement that all participants must be in the same place at the same time. It also allows each participant to express his or her own abstractions and themes (as opposed to group clustering in the same room). Angler then calculates a consensus clustering based on agglomerative clustering (King 1967) techniques and calculates how aligned is the group vision. Finally, there is a consensus and ranking phase, where participants come to understand their differing views and vote on the final names of consensus clusters. Angler has been used within several workshops (e.g., for scenario planning) and has been shown to facilitate collaboration among participants of those workshops. Table 1 presents (slightly obfuscated versions of) some thoughts contributed during a workshop considering the consequences of a hypothetical political regime change in a country (Country-X) that has nuclear weapons. These are the 23 thoughts related to nuclear security issues. A total of 140 thought contributions were made by the seven participants during this phase of the workshop. Each thought includes a pithy summarizing “Catch Phrase” as well as a concise description of the issue to be considered. ID Catch Phrase Author Description 1 Amount of fissile material G How much fissile material is in Country-X? In what form is it stored? 2 C2 apparatus D The nature of the nuclear command and control structure. 3 configuration B configuration of arsenal assembled, disassembled, state of deployment 4 control over nukes C Does the new regime have effective control over Country-X's nuclear arsenal and facilities? 5 domestic audience B domestic public opinion 6 International Support E Level of US and international support for maintaining Country-X’s nuclear security 7 new regime nuke policy C Will the new regime alter Country-X's existing nuclear policies concerning deterrence, etc.? 8 Nuclear Command and Control G Specific structures and systems; safeguards, command authority, designation 9 nuclear labs and military B domestic nuclear and military corporate interests 10 Personnel Reliability F Extent of radical Religion-Z leanings of nuclear C2 personnel, and/or level of corruption. 11 Physical security of nuclear arsenal G Permissive Action Links, Hardened sites etc. 12 quality and reliability of personnel B nature of Religion-Z among nucledar personnel 13 Regional Developments A Country-X, Country-Y and stability 14 Resource allocation D The inclination of the new leadership to allocate resources to nuclear safety and security measures (or alternatively, force expansion/enhancement). 15 Successor regime A Nature of the successor regime and its willingness to maintain strong positive control 16 Technical Reliability F Functionality of nuclear safeguard devices. 17 Threat level E Who wants Country-X's nukes, either to use or to sabotage? 18 Threat perceptions D The nature of the new leadership's threat perceptions and how it impacts nuclear decision making. 19 US contingency plans C Does the US have effective contingency plans to seize control of Country-X's nuclear arsenal to prevent it from falling into hostile hands? 20 US inputs B US inputs into safety and security of Country-X’s arsenal 21 US intervention A US willingness to preempt should instability occur 22 Vulnerability to preemption F Extent of device core separation from firing technology given events surrounding succession. 23 Willingness to divert funds E State of Country-X's economy and willingness of govt to divert funds/budgetary allocation towards nuclear security Table 1: Example Angler “thought” Contributions Technical Challenges in Supporting Collaboration One challenge that arises in collaborative contexts such as Angler involves helping participants manage the often large number of thoughts that are contributed. A rich exchange of ideas among even a relatively small group of collaborators can quickly produce a volume of thoughts that may overwhelm some participants and impair their continuing participation, obscuring “the forest for the trees.” An important observation about such sets of thought contributions is that they are not independent from each other, and semantic relations can be defined among them and used to organize them. Sometimes a thought contributed by one participant may be redundant with a thought contributed by another participant (e.g., from Table 1, thought 8 subsumes thought 2). Sometimes two thoughts share a significant amount of semantic content, and although one is not subsumed by the other the two are good candidates for merging (e.g., from Table 1 thoughts 2, 4, and 8 or thoughts 19 and 21). Establishing such semantic relations among thought contributions within a collaboration thread helps to organize the shared thoughts and helps the participants in reviewing the contributions of others and grasping how some thought contributions may converge with or diverge from other thought contributions. A second challenge that arises in collaborative contexts such as Angler involves helping participants appreciate how extensive is the coverage of their collective contributions. Sometimes two or more thoughts may be complementary in nature, so that together they cover a natural set of possible considerations. For example, some thoughts may focus on international considerations (e.g., thoughts 6 and 13 from Table 1) while other thoughts focus on domestic considerations (e.g., thoughts 5 and 9 from Table 1). Sometimes a set of thoughts will partially cover a natural set of considerations. Sometimes a set of thoughts will fail to cover any of the considerations within some natural set (e.g., no thought in Table 1 considers treaties). Establishing such properties of the coverage of a set of contributions can reveal to the participants areas that are relatively under-considered and remain candidates for additional attention and may help collaborators overcome their personal and collective cognitive biases. Through spurring considerations of otherwise unexamined horizons during collaborative brainstorming, Angler may address one of the primary deficiencies cited by the 9/11 Commission; that is, “a failure of imagination” (9/11 Commission 2004). A third challenge that arises in collaborative contexts such as Angler involves helping participants to perceive shared interests or complementing expertise with other participants. The c
Questions in natural language are answered by consulting multiple sources and inferring answers from information they provide. An automated deduction system, equipped with an axiomatic application-domain theory, serves as the coordinator for the process. Sources include data bases, Web pages, programs, and unstructured text. Answers may contain text or visualizations. Although the approach is domain-independent, many of our experiments have dealt with geographic questions.
State-of-the-art pronoun interpretation systems rely predominantly on morphosyntactic contextual features. While the use of deep knowledge and inference to improve these models would appear technically infeasible, previous work has suggested that predicate-argument statistics mined from naturally-occurring data could provide a useful approximation to such knowledge. We test this idea in several system configurations, and conclude from our results and subsequent error analysis that such statistics offer little or no predictive information above that provided by morphosyntax.
This paper describes the application of the TextPro system to the task of recognition of named entities in speech. TextPro is a lightweight engine for interpreting cascaded finite-state transducers. Although originally intended for processing text, the experience of this evaluation demonstrates the system can easily be adapted to processing transcripts generated by a speech recognizer as well. 1. THE TEXTPRO EXTRACTION SYSTEM For its participation in the Hub4 named-entity identification task, SRI International employed a newly developed information extraction system called TextPro. TextPro is a lightweight interpreter of cascaded finite-state transducers that is based on the TIPSTER Document Manager architecture [Grishman et al., 1996] and the TIPSTER Common Pattern Specification Language1 (CPSL). TextPro finite-state transducers accept and produce sequences of annotations on the document conforming to the structure specified by the TIPSTER document manager architecture. The transducers themselves are expressed by finite-state rules written in CPSL. The grammars employed by the Hub4 name recognizer specified the creation of ENAMEX, NUMEX, and TIMEX annotations, as well as other annotations used by the system internally. After having run each of the cascaded transducers over an input text, a postprocessor would insert SGML markup as required by the rules of the named-entity task. TextPro was originally developed to process text documents, and to test alternative specifications for CPSL. The first author participated in the design committee for CPSL under the TIPSTER program. The program runs on PowerPC Macintosh computers, and is freely downloadable from the World Wide Web.2 Although originally developed for limited objectives, experience led us to conclude that TextPro was a very useful 1 Because of the premature end of the TIPSTER program, the specifications for the Common Pattern Specification Language were never finalized or published. Further information is obtainable from the authors. 2 The URL for obta in ing TextPro is http://www.ai.sri.com/~appelt/TextPro/. system for performing document annotation tasks that do not involve the construction and merging of template structures such as those typical of the MUC scenario template tasks [ref Muc6]. For this reason, and because it is small and extremely fast, we felt that TextPro was a superior alternative to the more well known FASTUS system [Hobbs et al., 1996], which SRI has employed in various MUC evaluations. 1.1. Adapting TextPro to the Hub4 task Although TextPro was originally intended to process newspaper texts, it proved to be very straightforward to process speech transcriptions in Universal Transcription Format, whether human or machine generated. The adaptation process began by translating the FASTUS grammar used for SRI International’s participation in MUC-6 into CPSL. The MUC-6 grammar provided a high-performance baseline to start from; the SRI MUC-6 FASTUS system performed well on the named-entity task, achieving an F-measure of 94. The MUC-6 name recognizer, however, was optimized for mixed case texts, and typical Wall Street Journal articles, and therefore its performance on the Hub4 task was considerably short of optimal. Adapting the grammar to work well with monocase texts, absent any information provided by capitalization in ordinary texts, required the use of large lexicons to indicate which words were likely to be parts of names. The TextPro Hub4 system uses four large lexicons in addition to the lexicons used by the MUC-6 system: 1. A large lexicon of United States place names that was originally distributed with the place-name gazetteer for MUC-5, supplemented by a manually-culled set of foreign place names from the same source. 2 . A proprietary list of person names of many national i t ies obtained from Nuance Communications Corp. 3 3 This proprietary list cannot be given out in a public distribution. It is possible to replace this list with a list of American first and last names culled from census data publicly available on the Web. However, because of the 3. A list of prominent American and multinational corporations. This lexicon was the same one used by SRI for MUC-5 and MUC-6 participation, with the addition of some recently founded corporations. 4. A list of United States government agencies and departments. The lexicon used in this evaluation was expanded considerably over previous versions by using names appearing in the Hub4 training data. After the initial grammars and large lexicons were in place, the next step was to do iterative testing and debugging to raise the level of the system’s performance by refining the rules and lexical entries. Given the high speed of the TextPro system, it was very easy to do runs over the entire set of training data. Ten megabytes of training data could be processed in about two hours on our available hardware. The process of hill climbing on training data was not much different for the Hub4 task than other information extraction tasks in which SRI has participated. A few innovations were necessary to process speech transcription data successfully: · Important discourse contexts, in particular sports report and weather report contexts, were recognized. Sports report contexts were recognized so that names referring to sports teams that were ambiguous with ordinary English words (e.g. “Indians”) would be properly identified. In weather report contexts, it was important to recognize that phrases like “in the sixties” are not temporal expressions. census data’s weaker coverage of foreign names, its performance on the Hub4 task is noticeably lower. · Rules were needed to decompose lists of person-name words into likely combinations of first, middle, and last names. This was very important for lists or conjunctions of person names, and for the frequent situations where names appeared adjacent to a sentence boundary that would not be marked in the speech transcript. · Successful name recognition requires recognizing subsequent references to the same person, particularly when such references involve only the first name or the last name of a previously mentioned person that would not otherwise be tagged as names because of ambiguity with ordinary English words. This strategy works very well, except in the relatively frequent situations in which the speaker utters a fragment, or a repair. These fragments, if incorrectly recognized as names, can cause the erroneous tagging of many words in a text when the fragments are recognized as common one-syllable words. The TextPro Hub4 system used frequency data gleaned from the Penn Treebank Wall Street Journal corpus to limit the recognition of very common words as name parts to those contexts in which their status as names was unambiguous. 1.2. Evaluation Results The table in Figure 1 illustrates the results obtained by the TextPro Hub4 system in the recent evaluation for TextPro applied to the reference transcripts, and for TextPro applied to the output of SRI’s own speech recognition system. For the reference transcripts, and the baseline recognizer output, the TextPro results are very close to the best reported in each category. These evaluation results are quite consistent with the results obtained by SRI during our development testing. For development and testing, we divided the available 10 megabytes of training data furnished by Mitre and BBN into an eightEvaluation Task Content Extent Type Average Reference transcript, Segment 1 0.93 0.87 0.90 0.90 Reference transcript, Segment 2 0.93 0.88 0.91 0.91 SRI recognizer, Segment 1 0.76 0.75 0.79 0.77 SRI recognizer, Segment 2 0.80 0.76 0.81 0.79 SRI <10X Real Time Recognizer 0.76 0.74 0.78 0.76 Figure 1: SRI International’s TextPro tagging results (F-measures) on the Hub4 named entity recognition task megabyte training corpus and a two-megabyte test corpus, which was kept blind. In a final run before the official test, we recorded an average F-measure of 92 for the training data, and 89 for the blind development test data.
In recent years, analysts have been confronted with the increasing availability of on-line sources of information in the form of natural-language texts. This increased accessibility of textual information has led to a corresponding interest in technology for processing this text automatically to extract task-relevant information. This demand for a technological solution to the need to deal with the often-overwhelming quantity of available information has stimulated the development of the field of Information Extraction. This article provides an overview of the problems addressed, current approaches toward solutions, and assesses the state of the art and its potential for future progress.
FASTUS is a system for extracting information from natural language text for entry into a database and for other applications. It works essentially as a cascaded, nondeterministic finite-state automaton. There are five stages in the operation of FASTUS. In Stage 1, names and other fixed form expressions are recognized. In Stage 2, basic noun groups, verb groups, and prepositions and some other particles are recognized. In Stage 3, certain complex noun groups and verb groups are constructed. Patterns for events of interest are identified in Stage 4 and corresponding ``event structures'' are built. In Stage 5, distinct event structures that describe the same event are identified and merged, and these are used in generating database entries. This decomposition of language processing enables the system to do exactly the right amount of domain-independent syntax, so that domain-dependent semantic and pragmatic processing can be applied to the right larger-scale structures. FASTUS is very efficient and effective, and has been used successfully in a number of applications.
SRI International participated in the MUC-6 evaluation using the latest version of SRI's FASTUS system [1]. The FASTUS system was originally developed for participation in the MUC-4 evaluation [3] in 1992, and the performance of FASTUS in MUC-4 helped demonstrate the viability of finite state technologies in constrained natural-language understanding tasks. The system has undergone significant revision since MUC-4, and it is safe to say that the current system does not share a single line of code with the original. The fundamental ideas behind FASTUS, however, are retained in the current system: an architecture consisting of cascaded finite state transducers, each providing an additional level of analysis of the input, together with merging of the final results.
SRI International developed an information extraction system called FASTUS, a permuted acronym standing for "Finite State Automata-based Text Understanding System. The choice of acronym is some-what misleading, however, because FASTUS is a system for information extraction , not text understanding. The former problem is much simpler and more tractable, characterized by a relatively straightforward specification of information to be extracted from the text, only a fraction of which is relevant to the extraction task, and with the author's underlying goals and nuances of meaning of little interest. In contrast, a text understanding task is to recover all of the information in a text, including that which is only implicit in what is actually written. All the richness of natural language becomes fair game, including metaphor, metonymy, discourse structure, and the recognition of the author's underlying intentions, and the full interplay between language and world knowledge becomes central to the task.
SRI International developed an information extraction system called FASTUS, a permuted acronym standing for “Finite State Automata-based Text Understanding Ssystem for application to general information extraction tasks. The choice of acronym is, however, unfortunately somewhat misleading, because FASTUS is an information extraction system, not a text understanding system. The former problem is a much simpler, more tractable problem that is characterised by a relatively straightforward specification of information to be extracted from the text that changes slowly over time, if at all, with only a fraction of the text being relevant to the extraction task, and with the author’s underlying goals and nuances of meaning of little interest. In contrast, a text understanding task is to recover all of the information that there is in a text, including that which is only implicit in what is actually written. All the richness of natural language becomes fair game, including metaphor, metonymy, discourse structure, and the recognition of the author’s underlying intentions, and the full interplay between language and world knowledge becomes central to the task. Text understanding is extremely difficult, and presents a number of research problems that have not yet been adequately solved. On the other hand, the relative simplicity of the information extraction task means that the full complexity of natural language need not be confronted headon. In fact, much simpler mechanisms can be successfully employed to solve the more constrained problem, and do so in a computationally efficient and conceptually elegant way. It was this insight that led to the development of the FASTUS system that was applied to the task of extracting information from articles about terrorism in Latin America for the MUC-4 evaluation [Hobbs et al., 1992; Appelt et al., 1993]. In contrast to NL-processing systems designed for text understanding applications, FASTUS does not do a complete syntactic and semantic analysis of each sentence. Instead, sentences are processed by a sequence of nondeterministic finite-state transducers. The output of each level of transducers
Approaches to text processing that rely on parsing the text with a context-free grammar tend to be slow and error-prone because of the massive ambiguity of long sentences. In contrast, FASTUS employs a nondeterministic finite-state language model that produces a phrasal decomposition of a sentence into noun groups, verb groups and particles. Another finite-state machine recognizes domain-specific phrases based on combinations of the heads of the constituents found in the first pass. FASTUS has been evaluated on several blind tests that demonstrate that state-of-the-art performance on information-extraction tasks is obtainable with surprisingly little computational effort.
We describe the results that SRI International achieved on the February 1992 ATIS Speech and Natural Language System Test. The basic architecture of the system is described, including a set of parameters capable of altering the system's behavior and processing strategy. We report on several experiments that were run on the February test set to evaluate several processing strategies for both natural-language only and full spoken-language system tests.
The system that SRI used for the MUC-4 evaluation represents a significant departure from system architectures that have been employed in the past. In MUC-2 and MUC-3, SRI used the TACITUS text processing system [1], which was based on the DIALOGIC parser and grammar, and an abudctive reasoner for horn-clause logic. In MUC-4, SRI designed a new system called FASTUS (a permutation of the initial letters in Finite State Automata-based Text Understanding System ) which we feel represents a significant advance in the state of the art of text processing. The system shares certain modules with the earlier TACITUS system, namely modules for text preprocessing and standardization, spelling correction, Hispanic name recognition, and the core lexicon. However, the DIALOGIC system and abductive reasoner, which were the heart and soul of the previous system, were replaced by a system whose architecture is based on cascaded finite-state automata. Using this system we were capable of achieving a significant level of performance on the MUC-4 task with less than one month devoted to domain-specific development. In addition, the system is extremely fast, and is capable of processing texts at the rate of approximately 3,200 words per minute, measured in CPU time on a Sun SPARC-2 processor. (Measured according to elapsed real time, the system about 50% slower, but the observed time depends on the particular hardware configuration involved.)
It is often assumed that when natural language processing meets the real world, the ideal of aiming for complete and correct interpretations has to be abandoned. However, our experience with TACITUS; especially in the MUC-3 evaluation, has shown that principled techniques for syntactic and pragmatic analysis can be bolstered with methods for achieving robustness. We describe three techniques for making syntactic analysis more robust-an agenda-based scheduling parser, a recovery technique for failed parses, and a new technique called terminal substring parsing. For pragmatics processing, we describe how the method of abductive inference is inherently robust, in that an interpretation is always possible, so that in the absence of the required world knowledge, performance degrades gracefully. Each of these techniques have been evaluated and the results of the evaluations are presented.
Discourse comprises those phenomena that usually do not arise when processing a single sentence. It appears to be the most difficult and probably the least understood aspect of automated message understanding. Five out of fifteen sites on a MUC-3 survey listed discourse as their main weakness and an area in which to concentrate future research. Virtually all systems presented here take a sentence-by-sentence approach to text understanding. Parsing and domain-dependent interpretation of sentences or sentence fragments (usually the latter) are followed by modules that attempt to connect these interpretations into a coherent whole. This paper gives an overview of the modules that make the transition from the interpretation of sentences to the interpretation of the text that contains these sentences. Systems presented in this paper exhibit various degrees of the following discourse understanding capabilities:• identifying portions of text that describe different domain events; this includes the capability of recognizing a single event and the capability of distinguishing multiple events;• resolving references:- pronoun references, e.g., finding the referent of It in the sentence It took place this morning,- proper name references, e.g., understanding that Luis Galan may be referred to as Senator Galan;- definite references, e.g., deciding what is the referent for The attack in the sentence The attack look us by surprise.• discourse representation : representation at the message level.
Abstract : SRI International has been engaged in research on text understanding for a number of years. The Naval Ocean Systems Center (NOSC) has sponsored three workshops in recent years for evaluating text understanding systems. SRI participated in the first Message Understanding Conference (MUC-1) in June 1987 as an observer, and subsequently as a participant. Our system was evaluated in the second and third workshops, MUC-2 and MUC-3. For MUC-2, the task that the systems had to perform was to extract information for database entries saying who did what to whom, when, where, and with what result. The application domain for MUC-3 was news articles on terrorist activities in Latin America. The task was similar to that in MUC-2, though somewhat more information had to be extracted. The principal measures in the MUC-3 evaluation were recall and precision. Recall is the number of answers the system got right divided by the number of possible right answers. It measures how comprehensive the system is in its extraction of relevant information. Precision is the number of answers the system got right divided by the number of answers the system gave. It measures the system's accuracy. The system SRI used for these evaluations is called TACITUS. TACITUS is a system for interpreting natural language texts that has been under development since 1985. It has a preprocessor and postprocessor currently tailored to the MUC-3 application. It performs a syntactic analysis of the sentences in the text, using a fairly complete grammar of English, producing a logical form in first-order predicate calculus. Pragmatics problems are solved by abductive inference in a pragmatics, or interpretation, component.
Megumi Kameyama合作论文数4