Bibliographies and their knowledge organization systems commonly cover broad topical areas. Indexes to knowledge organization systems, such as the Subject Index to the Dewey Decimal Classification, provide a general index to the entirety. However, every community and every specialty develops its own specialized vocabulary. An index derived from the specialized use of language within a single subdomain could well be different from a general-purpose index for all domains and preferable for that subdomain. Statistical association techniques can be used to create indexes to knowledge systems. A preliminary analysis based on the INSPEC database shows that subdomain indexes differ significantly from each other and from a general index, The greater the polysemy of individual words the greater difference in the indexes. 1. Indexes to Knowledge Organization Systems The Subject Index (or "Relative Index") to Dewey's Decimal Classification is an obvious example of an index to a knowledge organization system, Library users could not be expected to know what numbers in the classification scheme stood for any given topic, so the Subject Index provided a map from words and phrases in English to the corresponding numbers, in the following fonn: "If you are interested in [noun], look for this [number]," Dewey himself considered that the index was of great importance because it would lead the user from the right topic (in their own words) to the right number in the classified arrangement. We refer to such indexes as "entry vocabulary" indexes, because they start with the vocabulary of the language with which a user approaches the knowledge organization system. An English language index to a classification scheme that uses a decimal numeric notation serves an obvious need. It is, however, only one case of a broader class of mappings in the multiplicity of vocabularies which permeate infonnation services (Buckland, 1999), There is likely to be a need for indexes in French or other natural languages to the same classification, Also. knowledge organization systems based on natural language, such as lists of subject headings and thesauri, are also likely to need indexes because these controlled vocabularies are invariably more or less stylized, It cannot be assumed that the provision of USE cross-references will meet all users' needs. Another category is the creation of indexes to knowledge organization systems not in natural language vocabulary but in the "vocabulary" of another knowledge organization system, such a mapping between the U.S, Patent & Trademark Office's Patent Classification and the International Patent Classification of the World Intellectual Property Office. Unfortunately, such indexes (or mappings or "crosswalks") are expensive and time consuming to create, and they are liable to become obsolescent as terminology evolves, and knowledge organization systems are revised. Creating such mappings requires a large investment of effort by people highly skilled in knowledge organization and in the domain concemed, The development of the Unified Medical Language System b y the U.S. National Library of Medicine is an outstanding example, A team at the School of Information Management and Systems, University of California, Berkeley, has been developing an alternative approach to the creation of indexes to knowledge organization systems, and, more generally, to mappings between pairs of vocabularies. We have been using statistical association techniques to establish relationships
This paper reports on an integrated energy harvesting prototype that consists of dispenser-printed thermoelectric energy harvesting and electrochemical energy storage devices. Parallel-connected thermoelectric devices with low internal resistances were designed, fabricated and characterized. The use of a commercially available dc-to-dc converter was explored to step-up a 27.1 mV input voltage from a printed thermoelectric device to a regulated 2.34 V output at a maximum of 34% conversion efficiency. The regulated power succeeds in charging dispenser-printed, zinc-based micro-batteries with charging efficiencies of up to 67%. The prototype presented in this work demonstrates the feasibility of deploying a printable, cost-effective and perpetual power solution for practical wireless sensor network applications.
This work presents advancements in dispenser-printed thick film thermoelectric materials for the fabrication of planar and printable thermoelectric energy generators. The thermoelectric properties of the printed thermoelectric materials were measured as a function of temperature. The maximum dimensionless figures of merit (ZTs) at 302 K for the n-type Bi2Te3-epoxy composite and the p-type Sb2Te3-epoxy composite are 0.18 and 0.19, respectively. A 50-couple prototype with 5 mm × 640 µm × 90 µm printed element dimensions was fabricated on a polyimide substrate with evaporated metal contacts. The prototype device produced a power output of 10.5 µW at 61.3 µA and 171.6 mV for a temperature difference of 20 K resulting in a device areal power density of 75 µW cm−2.
The ALMA front-end archive system has to capture up to 64 IVIB/s for a period of several days plus the data of about 100,000 monitor points from all 66 antennas and the correlators. The main science data is delivered through corba based audio/video streams and finally stored on SATA disk arrays hosted on 6 computers and controlled by 12 daemons. All data is collected by software components running on computers in the antennas and then sent through dedicated fiber links to the Array Operations Site at 5000 m and from there to the Operations Support Facility (OSF) at 3000 m elevation. The various hardware and software components have been tuned and tested to be able to meet the performance requirements. This paper describes the setup and the various components in more detail and gives results of various test runs.
This paper describes a Chinese part-ofspeech tagging system based on the maximum entropy model. It presents a novel two-stage approach to using the part-ofspeech tags of the words on both sides of the current word in Chinese part-of-speech tagging. The system is evaluated on four corpora at the Fourth SIGHAN Bakeoff in the close track of the Chinese part-ofspeech tagging task.
Libraries need to support geographic search. The traditional reliance on place names and political jurisdictions needs to be complemented by greater attention to space, using latitude and longitude. If place name authority files are linked to (or developed into) place name gazetteers, spatial coordinates can be added, places can be located in space, similar and multiple place names can be disambiguated, additional spatial relationships can be established (for example, near, between). Map visualizations used to display geographic aspects of retrieved sets can also provide a more flexible way in to specify the geographic facet in search queries. Analyses show that library catalog records contain geographic data that remains unused. Recommendations and prototype interfaces are presented.
This paper describes the work on Chinese named entity recognition performed by Yahoo team at the third International Chinese Language Processing Bakeoff. We used two conditional probabilistic models for this task, including conditional random fields (CRFs) and maximum entropy models. In particular, we trained two conditional random field recognizers and one maximum entropy recognizer for identifying names of people, places, and organizations in unsegmented Chinese texts. Our best performance is 86.2% F-score on MSRA dataset, and 88.53% on CITYU dataset.
Digital technology encourages the hope of searching across and between different media forms (text, sound, image, numeric data). Topic searches are described in two different media: text files and socioeconomic numeric databases and also for transverse searching, whereby retrieved text is used to find topically related numeric data and vice versa. Direct transverse searching across different media is impossible. Descriptive metadata provide enabling infrastructure, but usually require mappings between different vocabularies and a search-term recommender system. Statistical association techniques and natural-language processing can help. Searches in socioeconomic numeric databases ordinarily require that place and time be specified.
This paper describes a Chinese word segmentation system based on unigram language model for resolving segmentation ambiguities. The system is augmented with a set of pre-processors and post-processors to extract new words in
This demonstration utilizes a geographic information system interface to display multilingual news documents in time and space by extracting place names from text and matching them to a multilingual multi-script gazetteer which identifies the latitude and longitude of the location.
This paper describes monolingual, bilingual, and multilingual retrieval experiments using the CLEF 2003 test collection. The paper compares query translation-based multilingual retrieval with document translation-based multilingual retrieval where the documents are translated into the query language by translating the document words individually using machine translation systems or statistical translation lexicons derived from parallel texts. The multilingual retrieval results show that document translation-based retrieval is slightly better than the query translation-based retrieval on the CLEF 2003 test collection. Furthermore, combining query translation and document translation in multilingual retrieval achieves even better performance.
Gordon Sun合作论文数Microsoft;Microsoft China;Tencent Technology, China;Yahoo Search3