
Label noise remains a fundamental challenge in supervised learning, often degrading model performance and data reliability. While many existing approaches assume that a single relabeling action by a human annotator suffices to correct errors, this assumption overlooks the inherent fallibility of human judgment. In this article, we propose a general-purpose framework that models relabeling as an iterative and imperfect process, drawing inspiration from industrial sample inspection techniques. Our method integrates uncertainty sampling with targeted inspection to identify and refine mislabeled data through repeated annotation and aggregation. Crucially, this framework extends label correction beyond categorical classification to encompass real-valued regression tasks. For categorical data, majority voting is used to resolve annotation conflicts, while for continuous labels, we introduce an averaging-based correction strategy that incrementally approximates true values over successive relabeling rounds. This design enables systematic refinement of both discrete and continuous labels under noisy supervision. We validate our approach through a two-stage evaluation. First, on clean benchmark datasets with synthetic noise, we demonstrate significant improvements in relabeling efficiency under various uncertainty metrics and model capacities. Second, on a real-world image dataset with inherently noisy labels, our method continues to outperform baseline strategies, especially when paired with strong learners like XGBoost. Our findings show that sample inspection is a scalable and cost-effective mechanism for robust label correction across both classification and regression domains, even under high error rates or limited computational resources.
Disinformation, although an ancient phenomenon, has gained unprecedented reach and speed with the rise of the internet and social media platforms. While traditional fact-checking approaches focus on the semantic content of information, this article proposes a quantitative analysis based on metadata and formal textual features to investigate disinformation from a quality dimension perspective, assuming that false or misleading information often fails to meet informational quality criteria. Using an experimental approach, we analyzed two datasets of news from reliable and unreliable sources and applied statistical methods, including the Mann–Whitney U test, Cliff’s Delta, and Rosenthal’s r, to measure differences and effect size in the quality dimensions of accuracy, currency, readability, consistency, and reliability. The results show that lexical cohesion and lexical diversity are the strongest discriminators of source reliability, followed by structural error rates, while currency and readability display only weak discriminative power. The proposed News Reliability Index (NRI) emerges as a moderate but complementary indicator. Overall, reliable sources consistently demonstrate higher information quality, but structural differences alone are insufficient to detect disinformation, especially considering the capacity of generative AI to produce syntactically coherent texts. We conclude that semantic content analysis remains essential for identifying disinformation, with structural features best applied as supporting signals in detection models. Finally, we highlight future challenges, such as the growing use of artificial intelligence in generating high-quality disinformation, which may reduce the effectiveness of structural metrics and complicate automation in verification processes.
Data cleaning is a critical component of modern data intelligence systems. However, most existing approaches have been proposed for tabular data, despite the fact that much real-world data is unstructured. We propose to leverage information extraction programs to disclose where to apply data cleaning processes for documents. After using information extraction to produce a table of values extracted from the document (an “extracted view” of the documents, in the sense of a database view ), the resulting table is subject to any desired tabular-based data cleaning protocol. However, in order to reflect such cleaning back in the source documents, we require that updates to the tabular views can be translated into document updates without introducing any unintended changes to the contents in those views. With this approach, we consider a document to be clean if all its extracted views are clean. In this article, we characterize and verify a set of sufficient conditions for rule-based extraction programs that ensure document updates do not introduce unintended changes to the extracted views, thus qualifying them for inclusion in a document cleaning pipeline. Through experiments conducted on medical records and on math-intensive documents, we demonstrate that our approach provides an effective, practical pipeline for correcting data quality problems in documents. The second of these experiments demonstrates that a math retrieval system can perform at least twice as well on the cleaned documents as it does when applied to the documents without cleaning, matching the improvements achieved by specially tailored cleaning methods applied to those documents during the search pipeline but preserving the cleaned documents for use in other applications as well.
Anomaly-management tools monitor data streams in complex systems, e.g., tracking weather parameters collected from sensors in a region or inspecting resource usage of machines in a data center. Their goal is to detect outliers and abnormal behavior of the system, including a burst of outliers, which is often the result of a major event, like a hurricane or a crash of machines in a data center. An interval alert is an alert raised for such a burst of outliers. In data pipelines, such alerts indicate time intervals during which major data quality issues may occur. This allows alerting downstream applications on potential events and data quality risks. In this paper we present Ourea , a novel unsupervised learning method that computes interval alerts by identifying anomalous time intervals of variable lengths in data streams. Ourea raises interval alerts on significant events, based on a concentration of outliers detected in the raw data stream and in high-level aggregate views. It uses Kernel Density Estimate (KDE) to identify the most significant time intervals and their boundaries, can issue preliminary alerts early before intervals end, and minimizes the number of raised alerts by alerting only on significant ones. Extensive experiments over real and synthetic data demonstrate the effectiveness and scalability of Ourea and show that for detecting variable-length anomalous intervals in data streams, our algorithm is more accurate and more efficient than state-of-the-art methods.
An important goal of Chief Data Officers (CDOs) and data quality efforts is to increase the value of an organization's data. But that increased value makes the data an even more desirable target for cyberattacks, which have become more frequent, sophisticated, and impactful. In addition to the efforts that individual companies have made, the governments worldwide are responding by introducing or proposing new cybersecurity regulations to help protect that data, making security an important aspect of data quality. This study offers a novel perspective on the evolving global cybersecurity regulatory environment. Drawing on a comprehensive comparative analysis of nearly 200 regulatory frameworks from a wide array of international and national jurisdictions, the research identifies a core group of regulatory features that are systematically organized into five principal thematic categories. In particular, this research employs an integrated classification schema and a multidimensional taxonomy to facilitate more precise navigation of the complex regulatory landscape. Using a structured qualitative synthesis approach combining elements of review and cross-jurisdictional mapping, the analysis highlights notable disparities in regional regulatory focus, with Data Privacy, Incident Reporting, and Security by Design standing out as the most recurrent regulatory priorities. A significant outcome of the study is the identification of varying synergy levels between regulatory features, with high integration observed in Data Privacy and Cross-Border Data Transfer, medium synergy in areas such as Incident Reporting and Risk Management, and low synergy between Security by Design and Emerging Technologies. Therefore, by examining how various regulatory features align—or fail to align—across jurisdictions, the study provides critical insights for legislators, regulators, and data quality leaders and researchers. The article concludes with targeted recommendations to support more consistent, adaptive, and future-resilient cybersecurity governance worldwide to further improve data quality.
Data quality can be assessed across multiple high-level concepts called dimensions, such as accuracy, completeness, consistency, and timeliness. While extensive research and several attempts for standardization (e.g., ISO/IEC 25012) exist for data quality dimensions, their practical application often remains unclear. In parallel to research endeavors, a large number of tools have been developed that implement functionalities for the detection and mitigation of specific data quality issues, such as missing values or outliers. With this article, we aim to bridge this gap between data quality theory and practice by systematically connecting low-level functionalities offered by data quality tools with high-level dimensions, revealing their many-to-many relationships. Through an examination of seven open-source data quality tools, we provide a comprehensive mapping between their functionalities and the data quality dimensions, demonstrating how individual functionalities and their variants partially contribute to the assessment of single dimensions. This systematic survey provides both practitioners and researchers with a unified view on the fragmented landscape of data quality checks, offering actionable insights for quality assessment across multiple dimensions.
Since 2007, the Linked Open Data (LOD) Cloud has served as a central hub for datasets following Linked Data (LD) principles, offering a large repository of interconnected information. Over time, it has undergone multiple quality assessments to ensure datasets are accessible, well-maintained, and meet standards. Grounded on metadata assessment performed over time, this paper examines the current quality of the LOD Cloud by analyzing 1,658 datasets from the December 2024 snapshot, evaluated against 52 quality metrics. By proposing a reproducible methodology, it reports about the quality assessment and the trend analysis assessing progress, identifying persistent problems, and verifying how datasets registered in the LOD Cloud evolve over time. According to results, many earlier issues persist. Datasets still lack consistency in metadata structure, licenses, and distribution format. Moreover, they mainly remain in archived versions, with real-time access often poorly maintained. Well-curated, up-to-date datasets are exceptions rather than the rule.
Record linkage is the process of identifying records that refer to the same real-world entity within or across databases. If training data in the form of true matches (two records referring to the same entity) and true non-matches (two records referring to different entities) are available, then record linkage can be viewed as a supervised classification problem. Performance measures such as precision, recall, and the F-measure, are commonly used to evaluate the linkage quality obtained with a trained classifier. However, as we show in this article, comparing multiple classifiers using such measures can lead to inconsistent evaluation because for a given measure the same result can be obtained from different classification outcomes. This can cause a suboptimal classifier being selected, which can result in linked data sets of poor quality and possibly wrong decisions being made. To overcome this problem, we propose the Consistent Record Linkage (CRL) measure, an application focused evaluation method that ensures record linkage classifiers are assessed in a fair and transparent way. The CRL-measure allows a user to define the maximum error rates that are acceptable for their linkage application, and it provides practically useful information about the robustness of a classifier with respect to the range of classification thresholds obtained with these error rates. We evaluate the CRL-measure on both synthetic and real-world data sets using multiple classifiers to show its advantage over standard performance measures.
Collaborative content generation (CCG) enables collective creation of artifacts like scientific articles. Quality is a paramount concern in CCG, and a multitude of methods have been proposed to evaluate the quality of artifacts. Nevertheless, the majority of these methods are reliant on centralized architectures, which present challenges pertaining to security, privacy, and availability. Blockchain technology proffers a potential resolution to these challenges, by furnishing a decentralized and immutable ledger of quality scores. In this manuscript, we introduce a blockchain-based quality control model for CCG that uses a semi-iterative algorithm to interdependently compute quality scores of artifacts and reputation of nodes. Our model addresses critical challenges in academic informetrics, such as citation manipulation, transparency in collaborative scholarship, and decentralized trust in metric computation. Our model also exhibits sensitivity to processing latency, rendering it more agile in the presence of delays. Our model's quality scores, evaluated against PageRank and HITS baselines, show comparable performance, with additional assessments of throughput, latency, and robustness against malicious nodes confirming its reliability. A theoretical comparison with recent studies validates its feasibility for real world informetric application.
This editorial summarizes the content of the Special Issue on Data quality dimensions in Data FAIRification design and processes of the Journal of Data and Information Quality (JDIQ).
Data preparation is crucial for achieving good data management following the four foundational FAIR principles-Findability, Accessibility, Interoperability, and Reusability. Processing datasets to achieve high data (and metadata) quality is mandatory in modern applications. However, the data preparation activities that are needed to reach such levels may easily become unsustainable due to, for example, resource intensity or scalability challenges. Moreover, some preparation efforts may become unnecessary if they result in negligible improvements or duplicate actions. This article examines the sustainability aspects of data preparation through the lens of a circular economy. Within the data landscape, this perspective encourages practices that minimize waste, extend the data life cycle, and maximize reuse in alignment with the FAIR principles. We explore these practices and their impact on selecting and configuring effective data preparation strategies to design sustainable, high-quality pipelines. To this end, we propose an evaluation model that integrates data quality metrics with sustainability parameters for human and computational tasks. Finally, we apply the model in a comparative analysis of key data preparation methods, demonstrating its effectiveness in assessing sustainability and quality tradeoffs.
The quality of metadata plays a crucial role in many data FAIRification processes. So much so, in fact, that allthe four main principles of data FAIRification prescribe the use of high-quality metadata.One of the main data management paradigms where metadata is a first-class citizen is Ontology-Based DataManagement (OBDM). The goal of OBDM is to provide users with a reconciled view of a set of heterogeneousdata sources by means of a semantic metadata layer comprising an ontology and a mapping. The former is ahigh-level, declarative representation of the domain of interest written in terms of a logical theory, and thelatter is a formal description of the relation between the symbols in the ontology and the data at the sources.In this article, we introduce a novel data quality framework based on OBDM and specifically tailored formetadata analysis. The target of this framework is one of the most common forms of metadata currentlyin circulation, i.e., the integrity constraints defined by a database schema. Specifically, we will focus on thedata quality dimension known as Consistency, i.e., the property of data that is free of contradictions andincoherence. In this context, our techniques provide a set of tools to compare the integrity constraints definedby a database schema against the knowledge encoded in an ontology and check whether these constraints arestrict enough (i.e., protect) and are not too strict (i.e., are faithful to) for such knowledge.The contribution of the article is the presentation of the framework and the study of the related computationalproblems. We will present a detailed computational complexity analysis of such problems and show that theyare decidable for classes of OBDM specifications and integrity constraints that are very popular in practice.
The Linguistic Linked Open Data (LLOD) Cloud has emerged as a cornerstone of linguistic research, fostering dataset sharing and data reuse. Leveraging Semantic Web technologies, LLOD provides a rich tapestry of interconnected linguistic datasets that underpin advancements in both linguistics and Natural Language Processing. However, the ecosystem faces challenges related to data accessibility, interoperability, and reuse. This article evaluates the compliance of LLOD datasets with the FAIR principles, i.e., Findability, Accessibility, Interoperability, and Reusability, to assess their quality. A systematic literature review was conducted, identifying 69 linguistic datasets published over the last decade (2014-2024) using Semantic Web technologies. The datasets were evaluated through KGHeartBeat, an automated framework that assesses linked data quality. The analysis focused on an alignment between FAIR principles and Quality dimensions, including Accessibility and Trust, revealing that LLOD datasets are only partially findable and accessible, with scarce interlinking and a limited use of open licenses, which inhibits broader reuse. More in detail, the mapping proposed in this article is a novel and actionable alignment between quality dimensions and the FAIR principles, providing a structured framework for improving dataset compliance. The findings emphasize the need for enhanced accessibility, improved interlinking, and more widespread adoption of open licensing to maximize the value of LLOD for research and applications.
The FAIR (Findable, Accessible, Interoperable, and Reusable) data principles are crucial for data discoverability, sharing, and exploitation across diverse contexts, where evaluating data FAIRness reliably is essential for quality assessment and continuous improvement of data assets. Organizations can identify issues, implement targeted enhancements, and increase data value and trustworthiness by leveraging actionable insights bluefrom systematic data FAIRness assessment. However, this often requires heterogeneous metrics and guidelines, particularly when domain-specific challenges are involved. To bridge the gap between theoretical FAIR principles and their practical implementation, this article introduces xFAIR, a multi-layer platform architecture for assessing and enhancing data FAIRness, and for achieving data FAIRification in multiple domains. xFAIR incorporates modules for data acquisition, FAIRness evaluation, and ontology support. Its versatility is demonstrated via three real-world use cases: i. improving open data portals for Public Administrations, ii. extending FAIR assessment to multi-level European data portals, and iii. tackling metadata quality in news media by applying FAIR data principles to examine how source trustworthiness and metadata richness are linked. Each use case highlights specific aspects (e.g., domain-dependent metadata validation, or trust scores integration) to enhance quality assurance. Additional components support user feedback and media literacy. The obtained research outcomes underscore the importance of combining automated metadata validation with
Federated Search enables data accessibility under privacy-preserving regulations. One of the strategic objectives of the BBMRI-ERIC biobanking infrastructure is to make high-quality samples findable via Federated Search. The main prerequisites for biobanks to join the Federated Network are the conversion of a local database into a Common Data Model and the setup of a server node in which the database is loaded and made accessible for external queries. Data conversion is often the most critical step for institutions, as many of them lack the technical expertise needed to improve the FAIRness of their data. This is achieved through a local Extraction, Transformation and Loading process of data, usually extracted from a Biobank Information Management System. We hereby present a framework for the conversion of minimal information datasets into HL7-FHIR transaction bundles, allowing basic Biobank Interoperability, and enabling biobanks to be connected to the BBMRI-ERIC European Federated Platform. The toolkit consists of several Python modules, creating JSON files, ready to be uploaded to an internal FHIR server connected to the federated network, enabling data sharing and query execution. This tool has been successfully integrated in three BBMRI.it biobanks, allowing them to share their data correctly. In general, this tool will enforce data harmonization and standardization among research infrastructures, integrating the current pipeline into local information systems. The framework is available at https://github.com/bbdataeng/a-small-fire.
Data Management Plans (DMPs) describe how research data is managed, stored, and preserved. While machine-actionable DMPs (maDMPs) enable structured metadata for applications, their review remains a manual, labor-intensive process. This article introduces a conceptual framework for the automated evaluation of DMPs, focusing on FAIR assessment and funder compliance. Key contributions include the collection of requirements for automated DMP evaluation based on prior work on maDMPs and community input, as well as the development of a taxonomy of evaluation goals and dimensions with corresponding metrics. The article proposes the DMP Quality Vocabulary (DMPQV) for standardized communication of DMP quality measurements and introduces a mechanism for representing contextual information for maDMPs. A prototype implementation, validated through a case study using the Science Europe Practical Guide, extends the Research Data Alliance maDMP standard and supports funder-independent metrics such as completeness, feasibility, quality of actions, and compliance. Results demonstrate the framework's ability to generate standardized quality measurements and reports, aligning with manual assessments, while emphasizing the importance of clear guidelines for accurate evaluation.
Errors in data are a key challenge in modern data management and processing systems. Monitoring and mitigating risks associated with errors in data transformations and downstream applications, such as Machine Learning (ML) model training, requires a profound understanding of error generation and impact of errors on data pipelines. Unfortunately, scientific progress in the field is facing two main challenges: For one, research on data errors often does not adhere to the FAIR (Findable, Accessible, Interoperable, and Reusable) principles, which impedes reproducibility and comparisons. Second, existing data error models are oversimplified and fail to capture the complex statistical dependencies underlying the types and distributions of errors observed in real-world data. Building on prior work in the database management systems and statistics literature, we extend the theory on missing values to encompass a broader range of errors in tables and provide an overview of relevant error types. Combining error sampling mechanisms often observed in real data with a comprehensive categorization of errors, we introduce a latent factor model for tabular data errors that is simple to implement and can effectively model realistic error dependencies. Error sampling is decoupled from error types, which allows for simple extensions with more error types or sampling mechanisms. Using established benchmarks, we evaluate our model in two application scenarios, data cleaning and tabular ML tasks. In a comprehensive suite of experiments we demonstrate the impact of realistic error models on data cleaning benchmarks. Our results also show that a simple generative error model captures a wide range of error mechanisms and offers a convenient formalization of data perturbations to improve the generalizability, robustness and reproducibility of data cleaning research.
Collaboration is essential for scientific research. This is the foundation behind Open Science and the FAIR Principles, aimed at standardizing the development of scientific data sharing repositories. However, developing FAIR-compliant repositories can be a challenge, mostly due to managing a huge volume and variety of research data and metadata generated at a high velocity. We address these challenges by proposing BigFAIR, a novel FAIR-compliant architecture capable of managing this type of information in a massive scale. BigFAIR leverages existing local repositories, using separate infrastructures to handle scientific data and metadata. With this separation, it can reduce development and maintenance efforts, support data ownership, and increase flexibility. We define pipelines to demonstrate how BigFAIR answers queries in different contexts, introduce guidelines to support its instantiation, and propose a generic metadata warehouse model to support analytical query processing. We also demonstrate the applicability of BigFAIR through a case study in the context of two real-world datasets, detailing different types of queries and highlighting their importance to big data analytics.
This editorial summarizes the content of the Special Issue on Advanced Artificial Intelligence Technologies for Multimedia Big Data Quality of the Journal of Data and Information Quality (JDIQ).
Entity resolution is the problem of identifying records that refer to the same entity from one or multiple databases. Applications of entity resolution range from health and social science research to national security and online commerce. Entity resolution can be viewed as a classification task where pairs of records are classified as matches (referring to the same entity) or non-matches (referring to different entities). Alternatively, clustering-based entity resolution methods generate clusters of records such that each cluster refers to one entity, and each entity is represented by one cluster. If ground truth data in the form of known matches and non-matches are available, then performance measures such as precision, recall, and the F-measure, are commonly used to evaluate the quality of entity resolution methods. In practical applications, however, ground truth data are often not available, or they can be incomplete or biased, making quality evaluation challenging. To overcome this gap, we develop multiple methods to evaluate the quality of an entity resolution result without the need of ground truth data by calculating estimated numbers of true and false matches, as well as missed matches. These allow the calculation of estimates for precision, recall, and the F-measure. Our methods are either based on analysing links (classified record pairs) or the clustering structure provided by an entity resolution method. We validate our methods on multiple data sets from diverse domains, showing they can obtain precision and recall estimates close to their true values.