Author name disambiguation is becoming increasingly important due to the prevalent availability of publications in digital libraries. Various approaches for author name disambiguation are available, utilising a variety of information, e.g., author name, affiliation, title, journal and conference name or venue, citation, co-author, and topic information (Ferreira et al., SIGMOD Rec 41(2), 2012). Topics can be obtained, e.g., using subject information captured in various controlled vocabularies, classifications and mappings between them used to index publications (Torvik et al., J Am Soc Inf Sci Technol 56(2):140–158, 2005). Research interests of authors, evident in topics, might change over time though (Ferreira et al., SIGMOD Rec 41(2), 2012), and thus limit their usefulness for author name disambiguation. Here we present a longitudinal analysis of topics with respect to their suitability for author name disambiguation. We analyse the distribution of subject headings and classification notations taken from the Thesaurus (TSS) and the Classification for the Social Sciences (CSS) (http://www.gesis.org/en/services/research/thesauri-und-klassifikationen/) for research projects and literature (available in sowiport—http://sowiport.gesis.org maintained by GESIS) and the changes in distribution over time. To assess the suitability of subject information for author name disambiguation more closely, we then analyse the changes in the annotation over time for a selection of authors and author groups at different stages in their career, also taking into account the hierarchical organisation of the applied controlled vocabularies.
It has become widely recognized that user feedback can play a fundamental role in facilitating information integration tasks, e.g., the construction of integration schema and the specification of schema mappings. While promising, existing proposals make the assumption that the users providing feedback expect the same results from the integration system. In practice, however, different users may anticipate different results, due, e.g., to their preferences or application of interest, in which case the feedback they provide may be conflicting, thereby deteriorating the quality of the services provided by the integration system. In this paper, we present clustering strategies for grouping information integration users into groups of users with similar expectations as to the results delivered by the integration system. As well as grouping information integration users, we show that clustering results can be used as inputs to a wide range of functionalities that are relevant in the context of crowd-driven information integration. Specifically, we show that clustering can be used to identify feedback of relevance to a given user by exploiting the feedback provided by other users in the same cluster. We report on evaluation exercises that assess the effectiveness of the clustering strategies we propose, and showcase the benefits community- and crowd-driven information integration can derive from clustering.
The quality, and therefore, the usability and reliability of data in digital libraries depends on author disambiguation, i.e., the correct assignment of publications to a particular person. Author disambiguation aims to resolve name ambiguity, i.e., synonyms (the same author publishing under different names), and polysemes (different authors with the same name), and assign publications to the correct person. However, author disambiguation is difficult given that the information available in digital libraries is sparse and, when integrated from multiple data sources, contain inconsistencies in representation, e.g., of person names, or venue titles. Here we analyse and evaluate the usability of person-centred reference data available as linked data to complement the information present in digital libraries and aid author disambiguation.
Medical vocabulary is complex, and convoluted not least because of large numbers of compound terms. Formalized medical terminologies such as SNOMED-CT and ICD-10 take one of two strategies when representing medical language: so-called pre-coordination where valid compound terms are included explicitly and post-coordination where the terminology consists of a basis and a generative function from which the compound terms may be derived. However, these notions are not used with particular precision. Here we provide a formalization of the notion of coordination, a technique for estimating the degree of coordination in a given system, and an examination, based on our technique, of the coordination level of a number of terminologies.
Schema matching provides an important foundation for both manual and semi-automatic derivation of mappings between sources. However, schema matchers typically return large numbers of potentially inconsistent matches that are neither conducive to automatic mapping generation nor readily digested by mapping developers. This paper presents a method, EvoMatch, for automatically inferring schematic correspondences, from which mappings can be generated directly. It aims to offer a more expressive characterization of the relationships between sources than matches identified by existing schema matching methods. In particular, the paper contributes: i) an evolutionary search method for inferring schematic correspondences; ii) an objective function for calculating the fitness value of a solution within the search space; and iii) an empirical evaluation assessing the effectiveness of EvoMatch for inferring schematic correspondences in comparison with well established existing techniques. In doing so, EvoMatch automatically identifies correspondences that can be used directly to bootstrap information integration systems, or to inform the manual refinement of mappings.
One aspect of the vision of dataspaces has been articulated as providing various benefits of classical data integration with reduced up-front costs. In this paper, we present techniques that aim to support schema mapping specification through interaction with end users in a pay-as-you-go fashion. In particular, we show how schema mappings, that are obtained automatically using existing matching and mapping generation techniques, can be annotated with metrics estimating their fitness to user requirements using feedback on query results obtained from end users.
Schema matching algorithms aim to identify relationships between database schemas, which are useful in many data integration tasks. However, the results of most matching algorithms are expressed as semantically inexpressive, 1-to-1 associations between pairs of attributes or entities, rather than semantically-rich characterisations of relationships. This paper presents a benchmark for evaluating schema matching algorithms in terms of their semantic expressiveness. The definition of such semantics is based on the classification of schematic heterogeneities of Kim et al.. The benchmark explores the extent to which matching algorithms are effective at diagnosing schematic heterogeneities. The paper contributes: (i) a wide range of scenarios that are designed to systematically cover several reconcilable types of schematic heterogeneities; (ii) a collection of experiments over the scenarios that can be used to investigate the effectiveness of different matching algorithms; and (iii) an application of the experiments for the evaluation of matchers from three well-known and publicly available schema matching systems, namely COMA++, Similarity Flooding and Harmony.
Dataspace management systems (DSMSs) hold the promise of pay-as-you-go data integration. We describe a comprehensive model of DSMS functionality using an algebraic style. We begin by characterizing a dataspace life cycle and highlighting opportunities for both automation and user-driven improvement techniques. Building on the observation that many of the techniques developed in model management are of use in data integration contexts as well, we briefly introduce the model management area and explain how previous work on both data integration and model management needs extending if the full dataspace life cycle is to be supported.We show that many model management operators already enable important functionalities (e.g., the merging of schemas, the composition of mappings, etc.) and formulate these capabilities in an algebraic structure, thereby giving rise to the notion of the core functionality of a DSMS as a many-sorted algebra. Given this view, we show how core tasks in the dataspace life cycle can be enacted by means of algebraic programs. An extended case study illustrates how such algebraic programs capture a challenging, practical scenario.
Linked Data (LD) provides principles for publishing data that underpin the development of an emerging web of data. LD follows the web in providing low barriers to entry: publishers can make their data available using a small set of standard technologies, and consumers can search for and browse published data using generic tools. Like the web, consumers frequently consume data in broadly the form in which it was published; this will be satisfactory in some cases, but the diversity of publishers means that the data required to support a task may be stored in many different sources, and described in many different ways. As such, although RDF provides a syntactically homogeneous language for describing data, sources typically manifest a wide range of heterogeneities, in terms of how data on a concept is represented. This paper makes the case that many aspects of both publication and consumption of LD stand to benefit from a pay-as-you-go approach to data integration. Specifically, the paper: (i) identifies a collection of opportunities for applying pay-as-you-go techniques to LD; (ii) describes some preliminary experiences applying a pay-as-you-go data integration system to LD; and (iii) presents some open issues that need to be addressed to enable the full benefits of pay-as-you go integration to be realised.
The vision of dataspaces is to provide various of the benefits of classical data integration, but with reduced up-front costs. Combining this with opportunities for incremental refinement enables a ‘pay-as-you-go' approach to data integration, resulting in simplified integrated access to distributed data. It has been speculated that model management could provide the basis for Dataspace Management, however, this has not been investigated until now. Here, we present DSToolkit, the first dataspace management system that is based on model management, and therefore, benefits from the flexibility provided by the approach for the management of schemas represented in heterogeneous models, supports the complete dataspace lifecycle, which includes automatic initialisation, maintenance and improvement of a dataspace, and allows the user to provide feedback by annotating result tuples returned as a result of queries the user has posed. The user feedback gathered is utilised for improvement by annotating, selecting and refining mappings. Without the need for additional feedback on a new data source, these techniques can also be applied to determine its perceived quality with respect to already gathered feedback and to identify the best mappings over all sources including the new one.
User feedback is gaining momentum as a means of addressing the diculties underlying information integration tasks. It can be used to assist users in building information integration systems and to improve the quality of existing systems, e.g., in dataspaces. Existing proposals in the area are conned to specic integration sub-problems considering a specic kind of feedback sought, in most cases, from a single user. We argue in this paper that, in order to maximize the benets that can be drawn from user feedback, it should be considered and managed as a rst class citizen. Accordingly, we present generic operations that underpin the management of feedback within information integration systems, and that are applicable to feedback of dierent kinds, potentially supplied by multiple users with dierent expectations. We present preliminary solutions that can be adopted for realizing such operations, and sketch a research agenda for the information integration community.
The vision of dataspaces proposes an alternative to classical data integration approaches with reduced up-front costs followed by incremental improvement on a pay-as-you-go basis. In this paper, we demonstrate DSToolkit, a system that allows users to provide feedback on results of queries posed over an integration schema. Such feedback is then used to annotate the mappings with their respective precision and recall. The system then allows a user to state the expected levels of precision (or recall) that the query results should exhibit and, in order to produce those results, the system selects those mappings that are predicted to meet the stated constraints.
Model Management, and its associated operators, provides generic means for dealing with multiple schemas and the mappings between them, for example, in the context of multiple heterogeneous data sources that need to be integrated. One example of a Model Management framework is the 'Model Independent SchemaManagement'(MISM) platform. In the context of MISM, algorithms and implementations of various operators have been proposed that act on a source-model independent metamodel. However, although the results on MISM indicate how to import and manipulate data from heterogeneous source types, to date no approach has been proposed to utilise MISM for querying across the multiple data sources. This paper presents SMql, a query language over the source-model independent supermodel, presents an algebra into which the query is translated and presents an approach for rewriting SMql queries into source-model-specific queries posed over the corresponding relational or XSD models of the data source to be queried. Thus this paper helps to complete the collection of problems that need to be addressed to allow source model-independent model management using universal models in the context of MISM.
The vision of dataspaces is to provide various of the benefits of classical data integration, but with reduced up-front costs, combined with opportunities for incremental refinement, enabling a "pay as you go" approach. As such, dataspaces join a long stream of research activities that aim to build tools that simplify integrated access to distributed data. To address dataspace challenges, many different techniques may need to be considered: data integration from multiple sources, machine learning approaches to resolving schema heterogeneity, integration of structured and unstructured data, management of uncertainty, and query processing and optimization. Results that seek to realize the different visions exhibit considerable variety in their contexts, priorities and techniques. This chapter presents a classification of the key concepts in the area, encouraging the use of consistent terminology, and enabling a systematic comparison of proposals. This chapter also seeks to identify common and complementary ideas in the dataspace and search computing literatures, in so doing identifying opportunities for both areas and open issues for further research.
The specification of schema mappings has proved to be time and resource consuming, and has been recognized as a critical bottleneck to the large scale deployment of data integration systems. In an attempt to address this issue, dataspaces have been proposed as a data management abstraction that aims to reduce the up-front cost required to setup a data integration system by gradually specifying schema mappings through interaction with end users in a pay-as-you-go fashion. As a step in this direction, we explore an approach for incrementally annotating schema mappings using feedback obtained from end users. In doing so, we do not expect users to examine mapping specifications; rather, they comment on results to queries evaluated using the mappings. Using annotations computed on the basis of user feedback, we present a method for selecting from the set of candidate mappings, those to be used for query evaluation considering user requirements in terms of precision and recall. In doing so, we cast mapping selection as an optimization problem. Mapping annotations may reveal that the quality of schema mappings is poor. We also show how feedback can be used to support the derivation of better quality mappings from existing mappings through refinement. An evolutionary algorithm is used to efficiently and effectively explore the large space of mappings that can be obtained through refinement. The results of evaluation exercises show the effectiveness of our solution for annotating, selecting and refining schema mappings.
The vision of dataspaces has been articulated as providing various of the benefits of classical data integration but with reduced up-front costs, which, combined with opportunities for incremental refinement, enables a “pay as you go” approach to the data integration problem. However, results that seek to realise the vision tend to make design commitments, often to meet quite specific application assumptions, that are likely to restrict their wider use. Instead of precommitting to a specific solution, we build on research in model management and present a generic framework consisting of a collection of types and operations for dataspace management systems that can be instantiated in various ways. The key extension for dataspaces is the integration of user feedback as annotations to model management constructs, and the development of operations that take account of these annotations. The flexibility of the framework is demonstrated through various case studies that meet differing requirements.
The vision of dataspaces has been articulated as providing various of the benefits of classical data integration, but with reduced up-front costs, combined with opportunities for incremental refinement, enabling a "pay as you go" approach. However, results that seek to realise the vision exhibit considerable variety in their contexts, priorities and techniques, to the extent that the definitional characteristics of dataspaces are not necessarily becoming clearer over time. With a view to clarifying the key concepts in the area, encouraging the use of consistent terminology, and enabling systematic comparison of proposals, this paper defines a collection of dimensions that capture both the components that a dataspace management system may contain and the lifecycle it may support, and uses these dimensions to characterise representative proposals.
Data integration in the life sciences continues to be important but challe- ing. The ongoing development of new experimental methods gives rise to an increasingly wide range of data sets, which in tur
Proteomics, the study of the protein complement of a biological system, is generating increasing quantities of data from rapidly developing technologies employed in a variety of different experimental workflows. Experimental processes, e.g. for comparative 2D gel studies or LC-MS/MS analyses of complex protein mixtures, involve a number of steps: from experimental design, through wet and dry lab operations, to publication of data in repositories and finally to data annotation and maintenance. The presence of inaccuracies throughout the processing pipeline, however, results in data that can be untrustworthy, thus offsetting the benefits of high-throughput technology. While researchers and practitioners are generally aware of some of the information quality issues associated with public proteomics data, there are few accepted criteria and guidelines for dealing with them. In this article, we highlight factors that impact on the quality of experimental data and review current approaches to information quality management in proteomics. Data quality issues are considered throughout the lifecycle of a proteomics experiment, from experiment design and technique selection, through data analysis, to archiving and sharing.
Bijan Parsia合作论文数Department of Computer Science, School of Engineering, The University of Manchester4
Suzanne M. Embury合作论文数Manchester University;Department of Computer Science2