
Thesauri have been proven means to identify documents in libraries for centuries. In this paper, we show how this approach can be combined with most recent Internet technologies. The Java-based general thesaurus browser GenThes is able to handle several heterogeneous, multilingual thesauri. With the General European Multilingual Environmental Thesaurus (GEMET), GenThes is currently being used with several environmental catalogue systems. Feedback from users reveal that this approach greatly facilitates search and retrieval as compared to free-text only search. Performance which is crucial in Web applications as well has been improved by reducing the transfer volume of data and code. The software architecture of GenThes supports easy configuration and adaption to the individual needs of different systems that it is connected to. Ongoing work as well addresses to use GenThes as a query expansion module for multilingual search in distributed document collections.
The paper gives an overview on the meta-data specification for administrating and enforcing enterprisewide security for heterogeneous and distributed information systems. The meta-data serves as a basis to maintain enterprise-wide security information centrally, to integrate isolated security specifications, to keep the consistency between different security policies, and to perform access controls. The meta-data specifies all the information necessary for retaining the security concepts of an interoperable environment as well as all corresponding security information. Since several security systems have to be integrated within an interoperable environment the meta-data also contains the specification of mappings between security concepts and concrete security information.
The ability for knowledge workers to customize their information space, to personalize the information items with which they work, is an important capability, one that is valuable in the performance of complex tasks. Ideally all of the information systems with which users interact would support the capability to personalize information by allowing the individualized customization of the data objects they contain. However, this is a difficult capability to provide and is not supported by most current systems. This is especially true for networked based information systems with a large number of distributed users, a type of system that represents an increasingly important part of the information needs of knowledge workers. The research described in this paper addresses the information customization issue. A metadatabased architecture for supporting the customization of data objects is presented along with usage scenarios describing how the functionality it provides can be utilized to support the customization and personalization process in existing systems.
We present the Smart Object, Dumb Archive (SODA) model for digital libraries (DLs), and discuss the role of metadata in SODA. The premise of the SODA model is to "push down" many of the functionalities generally associated with archives into the data objects themselves. Thus the data objects become "smarter", and the archives "dumber". In the SODA model, archives become primarily set managers, and the objects themselves negotiate and handle presentation, enforce terms and conditions, and perform data content management. Buckets are our implementation of smart objects, and da is our reference implementation for dumb archives. We also present our approach to metadata translation for buckets.
The popularity and growth of the “Information SuperHighway” (e.g., the Web) have dramatically increased the number of information sources available for use and the opportunity for important new information-intensive applications (e.g., massive data warehouses, integrated supply chain management, global risk management, in-transit visibility). Unfortunately, there are significant challenges to be overcome regarding data extraction and data interpretation in order for this opportunity to be realized. Data Extraction: One problem is the difficulty in easily and automatically extracting very specific data elements from Web sites for use by operational systems. New technologies, such as XML and Web Querying/Wrapping, offer possible solutions to this problem. Data Interpretation: Another serious problem is the existence of heterogeneous contexts, whereby each SOURCE of information and potential RECEIVER of that information may operate with a different context, leading to large-scale semantic heterogeneity. A context is the collection of implicit assumptions about the context definition (i.e., meaning) and context characteristics (i.e., quality) of the information. As a simple example, whereas most US universities grade on a 4.0 scale, MIT uses a 5.0 scale – posing a problem if one is comparing student GPA’s. Another typical example might be the extraction of price information from the Web: but is the price in Dollars or Yen (If dollars, is it US dollars or Hong Kong dollars), does it include taxes, does it include shipping, etc. – and does that match the receiver’s assumptions? In this paper, examples of important context challenges will be presented and the critical role of metadata, in the form of context knowledge, will be discussed. Preamble The Bible tells the tale of the Tower of Babel where mankind endeavored to build a tower to reach to the Heavens. According to the Bible, God introduced a multiplicity of languages – the resulting confusion made it impossible for such large-scale coordination and communication and led to the termination of the tower’s construction. Today we are attempting to build “information superhighways” to access information from around the organization and around the world. Will this current great endeavor succeed or will it also be overcome by a “confusion of tongues”? The effective use of metadata can provide an approach to overcoming the challenges. Motivation There have been significant research efforts focused on physical information infrastructure, such as establishing high-speed data links to access information distributed throughout the world. It is increasingly obvious, however, that this kind of “physical
The economic advantages of electronic data interchange (EDI) are widely recognized. Nevertheless, the number of organisations and companies employing EDI is relatively small compared to the total number of businesses worldwide. The huge difference is caused by the fact that current EDI standards include a lot of complexity and their integration into existing applications is too expensive. This is due to the fact that current EDI standard messages are based on data models intended to capture all data that may appear in any business document of the corresponding business transaction. Business partners have to specify within a trading partner agreement a subset of the standard message which reflects their actual need before they are able to run an EDI partnership. Therefor, EDI should move to the meta level. The basic idea is that the concept of EDI is used for an agreement on the subset. Therefore, a meta message that is able to capture all the semantics in a trading partner agreement is needed. We demonstrate feasibility and advantages of a meta message approach by the example of the EDIFACT (Electronic Data Interchange For Administration, Commerce and Transport) standard.
this paper,we propose a meta-data representation model based on Tuple-Generating Dependencies. Thisrepresentation uses concepts extracted from a thesaurus. In section 1 we state the problem ofdocument retrieval, and present the Rameau thesaurus which defines the basic vocabulary usedthat can be used to build meta-data on documents. Depending on the use of meta-data for ourapplication field, we propose in section 2 a classification in three levels. We formally describe insection 3 the...
Searching of databases, textual or numeric, is likely to be effective and efficient only if the user is familiar with the classification, categorizing, and indexing schemes (metadata vocabularies) being searched. Therefore, it is obviously beneficial to provide a bridge between the user’s ordinary language and the metadata vocabularies of the unfamiliar database in order to compensate for abbreviated, cryptic, or specialized terminologies. Advanced search technologies would utilize customized Entry Vocabulary Modules (EVM) which respond adaptively to the searcher’s ordinary language query with a ranked list of search terms in the target metadata vocabularies that may more accurately represent what is sought in the unfamiliar database [l]. Experienced searchers know that familiarity with the source being searched is critical for effective searching. Each source has its own special characteristics, and familiarity comes from frequent use. The rapid increase in network accessible repositories increases the number and proportion of information resources that are unfamiliar. Entry Vocabulary Modules axe designed to respond to a searcher’s query with a ranked list of terms from a system vocabulary to help the searcher to deal with unfamiliar metadata. In order to accomplish this goal, EVMs make use of association dictionaries that map ordinary language terms to metadata vocabularies based on a co-occurrence measure. The source of these ordinary language terms is the titles, abstracts, and sometimes full text from documents indexed with metadata terms. Associations are recorded between ordinary language and the domain-specific, and often technical, metadata vocabularies that are typically used to describe databases. The technique of mapping terms in the target metadata vocabulary to ordinary language terms is a twostage lexical collocation process. In the creation of an Entry Vocabulary Module a “dictionary” of associations between the lexical items found in the titles, authors, and abstracts and the metadata terms (e.g., classification numbers or thesaural terms) assigned using a likelihood ratio statistic as a measure of association [2, 31. The terms can be words or noun phrases extracted using natural language parsing software. The dictionary is used to predict which of the metadata terms best represent the topic being searched. We have developed several prototype EVMs to map from ordinary language to the United States Patent Classification and the International Patent Classification, ordinary language to BIOSIS (biological abstracts) concept codes, and ordinary language to INSPEC thesaurus terms’. Entry vocabulary functionality has very extensive implementation potential and can be usefully positioned in several ways: As an aid on the searcher’s desktop to provide assistance when accessing an unfamiliar remote repository or as an aid on a repository server to help remote searchers unfamiliar with the local metadata scheme of that repository. An Entry Vocabulary Module can also be used for computer-assisted categorization. Because it can provide a ranked list of probably relevant metadata terms for any fragment of text, it can also be used for automating the process of categorization. For example, an Entry Vocabulary Module for the Patent Classification might be a useful support for the assignment of classification numbers as well as for patent searching.