The Context Interchange strategy presents a novel approach for mediated data access in which semantic conflicts among heterogeneous systems are not identified a priori, but are detected and reconciled by a context mediator through comparison of contexts . This paper reports on the implementation of a Context Interchange Prototype which provides a concrete demonstration of the features and benefits of this integration strategy.
A quality perspective in data resource management is critical. Because users have different criteria for determining the quality of data, we propose tagging data at the cell level with quality indicators, which are objective characteristics of the data and its manufacturing process. Based on these indicators, the user may assess the data's quality for the intended application. This paper investigates how such quality indicators may be specified, stored, retrieved, and processed. We propose an attribute-based data model, query algebra, and integrity rules that facilitate cell-level tagging as well as the processing of application data that is augmented with quality indicators. An ER-based data quality requirements analysis methodology is proposed for specification of the kinds of quality indicator to be modeled.
Data error is an obstacle to effective application, integration, and sharing of data. Error reduction, although desirable, is not always necessary or feasible. Error measurement is a natural alternative. In this paper, we outline an error propagation calculus which models the propagation of an error representation through queries. A closed set of three error types is defined: attribute value inaccuracy (and nulls), object mismembership in a class, and class incompleteness. Error measures are probability distributions over these error types. Given measures of error in query inputs, the calculus both computes and "explains" error in query outputs, so that users and administrators better understand data error. Error propagation is non-trivial as error may be amplified or diminished through query partitions and aggregations. As a theoretical foundation, this work suggests managing error in practice by instituting measurement of persistent tables and extending database output to include a quantitative error term, akin to the confidence interval of a statistical estimate. Two theorems assert the completeness of our error representation.
The Productivity From Information Technology (PROFIT) Initiative was established on October 23, 1992 by MIT President Charles Vest and Provost Mark Wrighton "to study the use of information technology in both the private and public sectors and to enhance productivity in areas ranging from finance to transportation, and from manufacturing to telecommunications." At the time of its inception, PROFIT took over the Composite Information Systems Laboratory and Handwritten Character Recognition Laboratory. These two laboratories are now involved in research related to context mediation and imaging respectively.
A set or premises, terms, and definitions for data quality management are established, and a step-by-step methodology for defining and documenting data quality parameters important to users is developed. These quality parameters are used to determine quality indicators about the data manufacturing process, such as data source creation time, and collection method, that are tagged to data items. Given such tags, and the ability to query over them, users can filter out data having undesirable characteristics. The methodology provides a concrete approach to data quality requirements collection and documentation. It demonstrates that data quality can be an integral part of the database design process. A perspective on the migration towards quality management of data in a database environment is given.< >