We describe an approach to pattern-based expertise topic extraction from publicly available scientific publications using Google Scholar. The approach is based on the observation that in the scientific text genre expertise topics will occur frequently in the context of particular phrasings that introduce them, such as 'method for', 'approach to', etc. The extracted knowledge can be used to analyze research structure in terms of expertise topics, researchers associated with these and relations between them.
To manage the flood of information that threatens to engulf (life-)scientists, an abundance of computer-aided tools are being developed. These tools aim to provide access to the knowledge conveyed within a collection of research papers, without actually having to read the papers. Many of these tools focus on text mining, by looking for specific named-entities that have scientific meaning, and relationships between these. An overview of the current state of the art is given in Rebholz-Schuhmann et al. (2005) and Couto et al. (2003). Typically, these tools identify a list of sentences containing relationships between two specific named-entities that can be found using rules or a thesaurus of synonyms. These sentences represent an overview of the interactions that are known with a specific entity, thus precluding the need for an exhaustive literature study. For example, the following are a few sentences that have been found using a typical text mining tool for the relationship 'p53 activates*': 1. The p53 tumor suppressor protein exerts most of its anti-tumorigenic activity by transcriptionally activating several pro-apoptotic genes. 2. We found that p53 ... activates[,] the promoter of the myosin VI gene.
We describe an approach towards automatic, dynamic and timecritical support for competency management and expertise search through topic extraction from scientific publications. In the use case we present, we focus on the automatic extraction of scientific topics and technologies from publicly available publications using web sites like Google Scholar. We discuss an experiment for our own organization, DFKI, as example of a knowledge organization. The paper presents evaluation results over a sample of 48 DFKI researchers that responded to our request for a-posteriori evaluation of automatically extracted topics. The results of this evaluation are encouraging and provided us with useful feedback for further improving our methods. The extracted topics can be organized in an association network that can be used further to analyze how competencies are interconnected, thereby enabling also a better exchange of expertise and competence between researchers.
OntoSelect is a dynamic web-based ontology library that harvests, analyzes and organizes ontologies published on the Semantic Web. OntoSelect allows searching as well as browsing of ontologies according to size (number of classes, properties), representation format (DAML, RDFS, OWL), connectedness (score over the number of included and referring ontologies) and human languages used for class-and object property-labels. Ontology search in OntoSelect is based on a combined measure of coverage, structure and connectedness. Further, and in contrast to other ontology search engines, OntoSelect provides ontology search based on a complete web document instead of one or more keywords only.
Watson is a gateway to the Semantic Web: it collects, analyzes and gives access to ontologies and semantic data available online with the objective of supporting their dynamic exploitation by semantic applications. We report on the analysis of 25 500 ontologies and semantic documents collected by Watson, giving an account about the way semantic technologies are used to publish knowledge on the Web, about the characteristics of the published knowledge, and about the networked aspects of the Semantic Web. Our main conclusions are 1that the Semantic Web is characterized by a large number of small, lightweight ontologies and a small number of large-scale, heavyweight ontologies, and 2that important efforts still need to be spent on improving the published ontologies (coverage of different topic domains, connectedness of the semantic data, etc.) and the tools that produce and manipulate
As more and more ontologies are being published on the Semantic Web, selecting the most appropriate ontology will become an increasingly important subtask in Semantic Web applications. Here we present an approach towards ontology search in the context of OntoSelect, a dynamic web-based ontology library. In OntoSelect, ontologies can be searched by keyword or by document. In keyword-based search only the keyword(s) provided by the user will be used for the search. In document-based search the user can provide either a URL for a web document that represents a specific topic or the user simply provides a keyword as the topic which is then automatically linked to a corresponding Wikipedia page from which a linguistically/statistically derived set of most relevant keywords will be extracted and used for the search. In this paper we describe an experiment in evaluating the document-based ontology search strategy based on an evaluation data set that we constructed specifically for this task.
This demo abstract describes the SmartWeb Ontology-based Annotation system (SOBA). A key feature of SOBA is that all information is extracted and stored with respect to the SmartWeb Integrated Ontology (SWIntO). In this way, other components of the systems, which use the same ontology, can access this information in a straightforward way. We will show how information extracted by SOBA is visualized within its original context, thus enhancing the browsing experience of the end user.
The paper describes VIeWs, a system that combines ontologies, web-based in- formation extraction, and automatic hyperlinking to enrich web documents with additional relevant background information. The central idea behind VIeWs is to demonstrate how web portals can be dynamically tailored to special interest groups by use of corresponding ontologies. As a particular use case we devel- oped an application for the "saarland.de" web portal of the Saarland region in Germany, which we present here in some detail. The paper describes the ideas behind the system and the Saarland.de application and provides an overview of the system architecture and components. Additionally, next to a comparison with related work, also some discussion on end user aspects of the application and its connection to the Semantic Web is given. It is argued that VIeWs is a typical end user application that depends on ontologies as semantic models for different scenarios, but that the need for Semantic Web technology beyond this has not been proven yet.
OntoSelect provides an access point for ontologies on any possible topic or domain that will be updated continuously, organized in a meaningful way and with automatic support for ontology selection in knowledge markup. Unlike the DAML1 and SchemaWeb2 ontology libraries, OntoSelect is not based primarily on a static registration of published ontologies, but includes a crawling procedure that monitors the web for any newly published ontologies in the following representation formats: RDF/S, DAML or OWL. Collected ontologies are analyzed using the OWL API (Bechhofer et al., 2003) that allows for the extraction of structure and content of any RDF/S, DAML or OWL ontology. There are currently around 745 ontologies in the OntoSelect library, covering a wide range of topics and domains. Ontologies are stored in a database and are organized according to: format; ontology-, classand property-names; classand property-labels. In the following two tables we present some statistics for the ontologies collected so far. Table 1 gives an indication of the distribution of the three representation formats used. Here, it is interesting to see that the OWL format already shows a clear advance over the other two formats, even quite shortly after the finalization of its definition3.
Guus Schreiber合作论文数VU University Amsterdam, Faculty of Sciences, Computer Science1