What is so difficult about building a lexicon? Although the power of a lexicon comes from its information content, the size of a lexicon in terms of number of words alone is not that important. Hand-held devices have advertised vocabularies of 100,000 words or more. What they are missing is any notion of what the words mean or how to use them. Word senses and word meanings are at the core of the lexical acquisition problem. Most natural language applications require some sort of word sense discrimination using knowledge about word meanings, knowledge not found in existing lexicons and dictionaries. This chapter covers several text processing applications, the knowledge sources that can help in word sense discrimination for these applications, and some traps to avoid in acquiring word sense information.
The GE-CMU team is developing the TIPSTER/SHOGUN system under the governmentsponsored TIPSTER program, which aims to advance coverage, accuracy, and portability in tex t interpretation . The system will soon be tested on Japanese and English news stories in tw o new domains . MUC-4 served as the first substantial test of the combined system. Because th e SHOGUN system takes advantage of most of .the components of the GE NLTooLsET excep t for the parser, this paper supplements the NLTOOLSET system description by explaining th e relationship between the two systems and comparing their performance on the examples from MUC-4 . INTRODUCTIO N Work on the GE-CMU TIPSTER-SHOGUN system began in the fall of 1991 . The first stage of the projec t involved integrating the resources and algorithms of the existing GE and CMU systems, in order that the combined team could effectively explore new data extraction methods . This integration was barely complete d in time for MUC-4: There are still some loose ends, and the combined system is still an infant . The syste m performed well on MUC-4 because it takes advantage of most of the capabilities of the GE system, an d because of the overall system architecture in which different parsers can be easily interchanged . This paper describes this architecture and compares the results of the two parsers on MUC-4 . SYSTEM OVERVIEW The TIPSTER-SHOGUN system as configured for MUC-4 uses exactly the same architecture as the G E system, except that it uses the CMU generalized LR parser instead of the TRUMP analyzer . The core lexicon and grammar of the GE system, as well as MUC-specific restrictions and additions to the knowledge base , have been converted to a form that the LR parser can use, and the system interfaces allow the two parser s to be interchanged at the flip of a switch . In fact, it occasionally proved useful in MUC to test the syste m by toggling between parsers . 'This research was sponsored (in part) by the Defense Advanced Research Project Agency (DOD) and other governmen t agencies . The views and conclusions contained in this document are those of the authors and should not be interpreted as representing the official policies, either expressed or implied, of the Defense Advanced Research Project Agency or the U S Government.
Nanoscaled interdigitated electrodes (IDEs) were developed for the purpose Of being used as miniature and sensitive affinity biosensors. IDEs with palladium as the electrode material on a SiO2 substrate were made with electrode widths and spacings ranging from 550 nm down to 250 nm. These sensors were theoretically modeled, enabling the prediction of their impedimetric response in KCl solutions. The strength of the theoretical model was proven by a characterization of IDEs of various dimensions in KCl solutions of different concentrations. The results demonstrate a quite reproducible behavior and a correlation between the impedimetric response in KCl solutions and the electrode dimensions.
Nanoscaled interdigitated electrodes (IDEs) are being developed for the realization of miniaturized and highly-sensitive affinity biosensors. Until now, nanoscaled IDEs have been realized on silicon wafers using deep UV lithography or e-beam patterning. However, for many applications in the biochemical field there is a strong need for cheap and/or disposable sensor devices. Therefore, a new, cost effective fabrication method for nanoscaled IDEs, which can also be applied on cheap, micro-moulded plastic substrates, has been developed. The method is based on the directionality of a vacuum evaporation process and omits expensive lithography steps completely. The feasibility of this electrode deposition technique has been proven by realizing IDEs on silicon substrates. Future work is focused on the realization of IDEs on injection moulded plastic substrates.
A novel design technique for planar conductometric sensors is proposed. This procedure is based on the Schwartz-Christoffel conformal transform and keeps under control the design parameters for planar conductometric devices. The new design method is exemplified for the case of a new planar haematocrit sensor. A basic three electrode structure and a semipermeable membrane is considered as a starting point for its development. The proposed device allows simultaneous evaluation of plasma and blood conductivity at one single frequency in the low range of the spectrum. The haematocrit sensor is to be further used as an ultrafiltration monitoring device in the haemodialysis process.
Planar conductivity sensors are the subject of increasing interest as basic transducers for biosensors. The high degree of control of the performance characteristics undoubtedly forms an important argument in favour of conductivity-based sensing. The paper provides an outline of the design rules to be followed if an optimal design of a planar conductivity cell is required. Based on a simplified model, it is shown that the required accuracy establishes a lower limit to the overall sensor dimensions. This lower limit is expressed as a minimum longitudinal path length necessary to obtain the desired accuracy. Given an available area, the optimum ratio of electrode-width over inter-electrode spacing for a basic two-electrode structure is shown to be close to unity. Furthermore, it is shown that the decomposition of the two electrodes into an interdigitated structure decreases the accuracy of the device if all other parameters are considered constant. If the sensing region has to be limited to within a thin sensitive layer, the splitting is proposed of one of the electrodes into a compound electrode. The optimum lay-out of this compound structure is calculated as a function of the layer thickness.
We present a method of enzyme immobilization on planar sensors. The method combines covalent enzyme bonding on magnetic beads with physical entrapment on the sensor surface. This procedure is suitable for batch production of planar biosensors, with the facility of enzyme patterning on the wafer. The results for glucose oxidase (GOD) enzymatic layers are presented and commented.
Many of the more ambitious goals of artificial intelligence have proved unattainable because of the failure of the many small, successful systems to scale up. The general use of technologies such as natural language interfaces and expert systems has done little to alleviate the basic difficulties and overwhelming cost of knowledge engineering. At the same time, emerging text processing techniques, including data extraction from text and new text retrieval methods, offer a means of accessing stores of information many times larger than any organized knowledge base or database. Although knowledge acquisition from text is at the heart of the information management problem, interpreting text, paradoxically, requires large amounts of knowledge, mainly about the way words are used in context. In other words, before intelligent text processing systems can be trained to mine for useful knowledge, they must already have enough knowledge to interpret what they read. The point at which there is “enough”, is still a matter of debate, as no real program seems close to having enough knowledge to achieve general human-like understanding. Current research in large-scale natural language processing has come, rightly, to focus on lexical acquisition as the key to future progress. Unfortunately, the current state of the art is quite far from the recipe for acquiring knowledge about words, because it leans too heavily on resources that are available, without consideration for what is needed
A detailed study of the performance of a planar thin-film differential-conductivity sensor for urea is described. The basic transducer consists of two identical interdigitated conductivity cells forming a differential pair. Special attention has been paid to the influence of the design parameters on the characteristics of the conductivity cells. The performance of the sensor is characterized by its differential gain, which is proportional to the cell constant of the individual cells. A common-mode rejection ratio (CMRR) is defined to quantify the capability of the differential sensor to suppress the influence of background conductivity changes on the output signal. This common-mode rejection reflects the selectivity of the transducer. CMRR values of over 40 dB can easily be obtained, comparable with a selectivity coefficient of 100 and more. Baseline stability is shown to be closely related to the temperature behaviour of the cells. In combination with the urease-containing protein membrane, a urea-specific sensor is obtained showing a linear response between 0.25 and 10 mM, with a typical sensitivity of 34 mV mM(-1) of urea.
The papers in this group cover a broad range of topics from different perspectives. They have in common an emphasis on the handling of frequently-occurring phenomena in real data sets of spoken and written language, phenomena which are in some sense outside of the scope of some of the core problems in human language technologies. We can view problems such as acoustic processing, word recognition, sentence parsing, and word sense disambiguation as core problems because they have a wealth of published literature and a set of broadly applied techniques. By contrast, the natural language papers in this session hit upon issues like recognizing speech repairs and designing "templates" to capture information and test text understanding. These issues are also central to HLT work, but have certainly not evolved into mature practices.
We discuss a method for using automated corpus analysis to acquire word sense information for multilingual text interpretation. Our system, SHOGUN, extracts data from news stories with broad coverage in Japanese and English. Our approach focuses on tying together word senses, using a combination of world knowledge (oatlogy) with word knowledge (corpus data). We explain the approach and its results in SHOGUN.
This report describes a few experiments aimed at producing high accuracy routing and retrieval with a simple Boolean engine. There are several motivations for this work, including: (1) using Boolean term combinatinations as a filter for advanced extraction systems, (2) improving legacy Boolean retrieval system by helping to automate the generation of Boolean queries, and (3) focusing on query content, rather than retrieval or ranking, as the key to system performance
NLDB, a knowledge-based system that automatically categorizes news stories for dissemination, retrieval, and browsing, is discussed. The major knowledge-based component of NLDB is a lexicosemantic pattern matcher that identifies combinations of words and phrases, as well as more complex patterns. These include word roots, grammatical categories, and semantic structures, such as verbs describing classes of events. It is shown that this linguistic analysis outperforms statistical methods. Because building lexicosemantic patterns can be a laborious process, a set of statistical methods that automate pattern acquisition while preserving the benefits of a knowledge-based approach are developed.< >