
Current question answering systems succeed in many respects regarding questions about textual documents. However, information exists in other media, which provides both opportunities and challenges for question answering. We describe our efforts in extending question answering capabilities to video data: our implemented prototype, Spot, can answer questions about moving objects in a surveillance setting. This novel application of vision and language technology is situated within a larger framework designed to integrate knowledge from multiple domains under a common representation. We believe that our framework will support the next generation of multimodal natural language information access systems.
In this paper we present a method for answering relationship questions, as posed for example in the spring of 2003 evaluation exercise of the AQUAINT program, which has funded this research. The goal of the exercise was to provide answers to questions requesting an account of the relationship between two or more entities. No restriction on the format of the answer was imposed, except that it should not consist of entire documents and that it should list document identifiers for the passages from which the answer was drawn or generated.
This paper illustrates ongoing research and issues faced when dealing with real-time questions in the domain of Reusable Launch Vehicles (aerospace engineering). The question-answering system described in this paper is used in a collaborative learning environment with real users and live questions. The paper describes an analysis of these more complex questions as well as research to include the user in the question-answering process by implementing a question negotiation module based on the traditional reference interview.
People query search engines to find answers to a variety of questions on the Internet. Search cost would have been greatly reduced if search engines could accept natural language questions as queries, and provide summaries that contain the answers to these questions. We introduce the notion of Answer-Focused Summarization, which is to combine summarization and question answering. We develop a set of criteria and performance metrics, to evaluate Answer-Focused Summarization. We demonstrate that the summaries produced by Google, the most popular search engine nowadays, can be largely improved for question answering. We develop a proximity-based summary extraction system, and then utilize question types, i.e. whether the question is a "person" or a "place" question, to improve the performance. We suggest that there is a great research potential for Answer-Focused Summarization. 2 Chapter 18
Questions in natural language are answered by consulting multiple sources and inferring answers from information they provide. An automated deduction system, equipped with an axiomatic application-domain theory, serves as the coordinator for the process. Sources include data bases, Web pages, programs, and unstructured text. Answers may contain text or visualizations. Although the approach is domain-independent, many of our experiments have dealt with geographic questions.
Experiences from combining dialogue system development with information extraction techniques
People when asked a number of questions about a particular topic begin to become knowledgeable about the topic as they look for and find answers to the questions. A question answering system should also possess this ability to “reuse” information used in answering previous questions. This article defines and exemplifies a dozen categories of reuse that were found in usergenerated question sets. The corpus of question sets is also discussed.
Advanced reasoning for Question Answering (QA) systems, as described in a recent roadmap (http://www-nlpir.nist.gov/projects/duc/papers/qa.Roadmappaper_v2.doc), raises new challenges since answers are not only directly extracted from written texts or structured databases but also constructed via several forms of reasoning in order to generate answer explanations and justifications. These systems require the integration of reasoning components operating over a variety of knowledge bases, encoding common sense knowledge as well as knowledge specific to a variety of domains by means, for example, of conceptual ontologies. These kinds of QA systems can be viewed as an enhancement, rather than a rival to retrieval based approaches. Integrating knowledge representation and reasoning mechanisms allow, for example, to respond to unanticipated questions and to resolve situations in which no answer is found in the data sources. Cooperative answering systems are typically designed to deal with such situations by providing useful and informative answers. These systems can e.g. identify and explain false presuppositions or various types of misunderstandings found in questions. Constraints relaxation in questions occur when the system cannot find any response. Intensional responses are provided instead of a large unstructured set of extensional answers. Cooperative answering systems can also provide summaries or conditional responses. An overview of these aspects is given in (Gaasterland et al, 94).
Next generation question answering systems are challenged on many fronts including but not limited to massive, heterogeneous and sometimes streaming collections, diverse and challenging users, and the need to be sensitive to context, ambiguity, and even deception. This chapter describes new directions in question answering (QA) including enhanced question processing, source selection, document retrieval, answer determination, and answer presentation generation. We consider important directions such as answering questions in context (e.g., previous queries, day or time, the data, the task, location of the interactive device), scenario based QA, event and temporal QA, spatial QA, opinionoid QA, multimodal QA, multilingual QA, user centered and collaborative QA, explanation, interactive QA, QA reuse, and novel architectures for QA. The chapter concludes by outlining a roadmap of the future of question answering, articulating necessary resources for, impediments to, and planned or possible future capabilities.
The current tendency in Question Answering is towards the processing of large volumes of open-domain text. This tendency is spurred by the creation of the Question Answering track in TREC, and the recent increase of systems that use the Web to extract the answers to the questions. This has undoubtly the advantage that narrow, application-specific concerns can be overlooked in favor of more general approaches. However the unconstrained nature of the and questions does not necessarily lead to systems that are better at specific tasks, as they might be required in a deployed application. It has been already been observed in other competitions (notably the Information Extraction competitions organized under the name of Message Understanding Conferences) that the nature of the competitive process tends to select a type of system that better adapts to the evaluation itself, rather than systems that deal in an optimal way with the problem. [To use a comparison from evolution theory, a too severe selection in a given local environment leads to a converge of the population to a very limited genetic pool, which is then uncapable of coping with even a minor change in the environment.] In restricted domains, systems cannot take advantage of the so-called Zipf's law of Questions [Prager], which states that there is an inverse relation between the frequency of certain types of questions and their complexity. In other words, the questions most frequently asked are those that can be solved with simpler techniques. By targeting a smaller set of frequent questions types, system can achieve good results with limited effort. By contrast, the non-redundant nature of most technical documentation, and the use of specific sublanguage and terminology, makes them unsuitable to (some of) the approaches seen in the TREC QA competition. In the proposed contribution We will discuss the specific nature of technical documentation, with examples from real domains (e.g. the Maintenance Manual of a large commercial aircraft) and illustrate solutions that have been adopted in a deployed system. An example of the difference between technical documents and open texts is the focus on specific types of entities. While in Open Domain systems Named Entities play a major role, in Technical Documentation they are almost irrelevant, by contrast a far greater role is played by terminology. Technical domains present the additional problem of domain navigation. By assuming that users are familiar with concepts, inexpert users are presented with a barrier separating questions from answers. Unfamiliarity with terminology might lead to questions which contain imperfect formulations of terms. A question answering system for junior doctors or training technicians needs therefore to use whatever scarce knowledge is contained in a query to extract relevant answers. Detecting terminological variants and exploiting the relations between terms (like synonymy, meronymy, antonymy) is vital to this task. Another idiosyncrasy of technical domains is the tendency towards definitional questions (what is the ANT connection?), which prove tricky to answer precisely in a generic document collection (and for this reason they have been deliberately left out of the recent TREC 2002). In Technical Domains it can be expected that such type of question would play a major role, and therefore systems must be capable of coping with them. In this book chapter we aim to explain the above concepts and illustrate them with examples taken from text from technical domains. We will also illustrate why techniques that are typically used in data-intensive open-domain question-answering systems would not work effectively in technical domains that have less data redundancy. In sum, we will show that question-answering of technical domains present a better opportunity to explore content-based approaches to question-answering, while at the same time bringing the possibility of producing commercially viable systems in the short term.
Although the World Wide Web contains a tremendous amount of information, the lack of intuitive information access methods and the paucity of uniform structure make finding the right knowledge difficult. Our solution is to turn the Web into a “virtual database” and to access it through natural language. We have accomplished this by developing a stylized relational framework, called the object-property-value model, which captures the regularity found in both natural language questions and Web resources. We have adopted this framework in START and Omnibase, two components of a system that understands natural language questions and responds with answers extracted on the fly from heterogeneous and semistructured Web sources. Our system can answer millions of questions from hundreds of Web resources with high precision.
The idea of compiling knowledge into FAQ files has existed for some time. The Usenet became an early repository of on-line FAQ files, and currently the Internet FAQ Archives web site (http://www.faqs.org) has 2490 " popular " FAQs archived. There are many other sources of FAQ files. Call-center manuals are also often structured as FAQ files. As the world wide web has become widely accessible, FAQ files have also become a popular way for web sites to store knowledge and convey answers to customers/site users about questions that these users would commonly ask. Thus, answers to a very wide variety of questions can be found in FAQ files, and developing applications and techniques for question answering tailored to FAQ files is an important thrust in question answering. An FAQ file typically contains several question-and-answer (Q and A) pairs where questions are pre-answered and compiled by domain experts. Thus, finding answers from FAQ files essentially is to reuse previously answered questions instead of finding answers from scratch every time—an economical solution. Due to this semistructured format, retrieval models for FAQ files place emphasis on different techniques than retrieval models for unstructured documents. First, the primary focus is on finding an FAQ question (not answer) which is similar to the user query/question, that is, a Q-to-Q match. In essence, this is the task of recognizing question paraphrases-two or more questions which ask the same thing(s) but are formulated in different ways. Recently, the issue of paraphrase recognition has been receiving attention in question-answering research as a way to fill the gap between words in a question and those in an answer. Most approaches try to enumerate paraphrase patterns, for instance " How much does X cost? " ⇒ " X costs Y " ≡ " the price of