
Science increasingly requires access to and use of accurate and secure data across the full spectrum of scientific fields. “Trustworthy” scientific data repositories are essential to meet this goal, having the structure and commitment to provide reliability, transparency, and sustainability of the data they hold. Certification frameworks for trustworthy data repositories serve as benchmarks for rigorous data stewardship, playing a pivotal role in establishing and maintaining the integrity of trustworthy repositories. Certification processes are highly valued, but navigating the complexity of procedures to attain certification can be a significant challenge, even for repositories that meet the requirements. This is particularly problematic for institutions with limited resources or unique local contexts, underscoring the need for greater support and flexibility to help scientific data repositories manage the procedures toward certification. This paper is the result of research conducted within World Data System (WDS), a global organization of science data repositories focused on strengthening the scientific enterprise throughout the entire data lifecycle in support of first-class research outputs. Through an analysis of current practices, challenges, and opportunities in repository certification, actionable recommendations are offered to reduce barriers, provide enhanced guidance, and foster equitable participation in trustworthy data stewardship, thereby advancing best practices in scientific data management.
Darwin Core (DwC) is an essential standard for sharing biodiversity data. However, the term dwc:habitat suffers from an inherent ambiguity due to its free-text format. This ambiguity severely compromises the interoperability and reusability of habitat data, hindering large-scale comparative analyses and impeding the formulation of effective conservation policies. As a solution to this problem, we propose adopting controlled vocabularies and ontologies. The NATURA2000 and EUNIS habitat classifications emerge as ideal candidates to standardize dwc:habitat. NATURA2000 offers a consolidated regulatory framework and habitat type definitions with direct implications for European conservation, while EUNIS provides a more comprehensive, hierarchical, and scientifically grounded system with the ability to cross-map with other standards. The implementation of such vocabularies would significantly improve the quality, consistency, interoperability, and reusability of habitat data, more robustly supporting scientific research and conservation policies.
This publication focuses on the implementation of a FAIR (findable, accessible, interoperable, and reusable) data workflow for the analytical method ion-exchange chromatography and its impact on the FAIR compliance of the resulting measurement data. The workflow includes the structured collection of measurement- and method-relevant metadata, their storage in an electronic laboratory notebook, and the representation of these metadata in a graph database with the help of an ontology. For the implementation of this workflow, metadata schemas, ontologies, and programming scripts for handling the data need to be designed to meet the requirements of the target group. The current implementation status is evaluated with regard to the fulfillment of the FAIR guiding principles according to Wilkinson et al. In addition, two structured FAIR assessments were carried out using the Australian Research Data Commons FAIR self-assessment and the Data Stewardship Wizard software. The comparison of the assessments for the current implementation and its advancement (e.g., by including an ontology for ion chromatography) enables a better understanding of which FAIR compliance criteria are already realized and which criteria need more focus during further workflow implementation. In the spirit of open science and FAIR data, the individual research artifacts (e.g., metadata schemas and the ontology) will be made available in the domains of analytics and plasma science.
The FAIR Principles—Findable, Accessible, Interoperable, and Reusable—offer a widely accepted framework for improving the sharing and reuse of digital scientific data by both human and machine users. Following these principles is critical for effective scientific data stewardship, broader scientific collaboration, and compliance with federal and agency data policies. This paper, based on the work of NASA’s Open, Free, and FAIR Working Group (O’FAIR WG) under the Earth Science Data Systems Program, presents an overview of how FAIR Principles are being applied within NASA’s Earth science data landscape. It highlights ongoing progress and challenges, identifies FAIR-enabling resources, and offers recommendations and strategic actions to enhance the FAIRness of NASA-funded open and free Earth science data products. The FAIR-enabling resources identified underscore the vital role of NASA’s existing enterprise processes, standards, tools, and infrastructures in supporting FAIR implementation. Our findings show strong performance in making NASA Earth science data findable and accessible. However, further work is needed—especially in enhancing interoperability, so that different systems and tools can better understand and exchange data. This is especially important for enabling machine-driven discovery and analysis. We emphasize the importance of a balanced strategy that combines a centralized, top-down approach—focused on building enterprise-level capabilities and processes—with a decentralized, bottom-up approach driven by discipline-specific needs and community practices. We advocate for coordinated efforts to enhance (meta)data interoperability to facilitate seamless data and information sharing and exchange of Earth science data both within NASA and across other agencies managing Earth science data.
Despite substantial investment in research data infrastructure, data discovery remains a fundamental challenge in the era of open science. The proliferation of repositories and the rapid growth of deposited data have not resulted in a corresponding improvement in data findability. Researchers continue to struggle to find data that are relevant to their work, revealing a persistent gap between data availability and data discoverability. Without rich, high-quality metadata, robust and user-centred data discovery systems, and a deeper understanding of how different researchers seek and evaluate data, much of the potential value of open data remains unrealised. This paper presents a set of practical, evidence-based recommendations for data repositories and discovery service providers aimed at improving data discoverability for both human and machine users. These recommendations emphasise the importance of 1) understanding the search needs and contexts of data users, 2) addressing the roles that data repositories play in enhancing metadata quality to meet users’ data search needs, and 3) designing discovery interfaces that support effective and diverse search behaviours. By bridging the gap between data curation practices, discovery system design, and user-centred approaches, this paper argues for a more integrated and strategic approach to data discovery.
As communities increasingly aim to understand and demonstrate how their data repositories support the FAIR Principles, decision-makers need tools that provide a clear, actionable snapshot. This paper introduces FIP Check, a rubric-based tool for assessing FAIR Implementation Profiles (FIPs) through a structured assessment of FAIR Enabling Resources (FERs). Built on the FIP ontology, the FIP Check brings three key innovations to FAIR assessment. First, it enables granular and practical evaluation by using a FAIR Principle-aligned rubric to assess individual FERs. Second, it measures partial FAIRness through a progressive scoring scale that captures varying levels of FAIR alignment, offering constructive, context-aware feedback rather than purely binary results. Additionally, this tool complements existing FAIR assessment tools by focusing on FERs as the units of assessment rather than on entire repositories or datasets. Third, it ensures transparent and inspectable balance by embedding expert-informed assessments within a standardized rubric and making all scoring decisions directly visible. FIP Check was piloted across seven biomedical data repositories, testing its utility in identifying strengths and actionable gaps, and supporting more informed FAIR improvement efforts. It enables communities to assess their current practices and plan targeted enhancements. By providing detailed insights into how individual FERs contribute to FIP’s FAIR alignment, the FIP Check transforms the FAIR Principles from abstract ideals into practical guidance that support strategic alignment, foster shared understanding, and encourage repository owners to see their resources not only as providers of data, but as infrastructure components.
Background: In poly crisis, where multiple, interconnected crises such as climate change, pandemics, and socio-economic inequalities converge, the integration of open science (OS), artificial intelligence (AI), geoinformatics, virtual reality (VR), and augmented reality (AR) offers a transformative approach to addressing these complex challenges. This article explores how these technologies can synergize to create resilient, adaptive, and inclusive solutions. Methodology: This paper uses a multifaceted technique, including structured review and technology convergence assessment analysis, to raise awareness of the abovementioned technologies to create resilient and adaptive solutions. Semantic network analysis is used to identify the current gaps and challenges in poly crisis. Findings: This work revealed that integrating these technologies can lead to more and better-informed decision-making, enhanced public engagement, and improved crisis response. The work also highlighted significant challenges, including disparities and inequities in access to data and technology and the lack of integration of science with policy and societal values in places. This integrated approach is essential for building resilient systems that address the current and future poly crisis. Policymakers, researchers, and technologists must work together to address these challenges and maximize the benefits of this integrated approach. Value/Originality: The key contribution of this work is the development of an assessment framework for managing poly crisis. Additionally, the work identifies gaps and proposes future research directions to guide the development and implementation of poly crisis solutions.
This article details a correction to: Silva, F.C.C., de Albuquerque Siebra, S., Rezende, L.V.R., de Oliveira, A.F. and de Araújo, D.O. (2026) ‘Essential Aspects of Tools for Developing Scientific Data Management Plans’, Data Science Journal, 25(1). Available at: https://doi.org/10.5334/dsj-2026-005.
The increasing use of AI-based approaches such as machine learning (ML) across diverse scientific fields presents challenges for reproducibly disseminating and assessing research. As ML becomes integral to a growing range of computationally intensive applications (e.g. clinical research), there is a critical need for transparent reporting methods to ensure both comprehensibility and the reproducibility of the supporting studies. There are a growing number of standards, checklists and guidelines enabling more standardized reporting of ML research, but the proliferation and complexity of these make them challenging to use. Particularly in assessment and peer review, which has to date, been an ad hoc process that has struggled to throw light on increasingly complicated computational supporting methods that are otherwise unintelligible to other researchers. Taking the publication process beyond these black boxes, GigaScience Press has experimented with integrating many of these ML-standards into the publication process. Having a broad-scope that necessitated looking at more generalist and automated approaches. Here, we map the current landscape of artificial intelligence (AI) standards, and outline our adoption of the DOME recommendations for Machine Learning in biology. We developed a publishing workflow that integrates the DOME Data Stewardship Wizard and DOME Registry tools into the peer-review and publication process. From this case study we provide journal authors, reviewers and Editors examples of approaches, workflows and strategies to more logically disseminate and review ML research. Demonstrating the need for continued dialogue and collaboration among various ML communities to create unified, comprehensive standards, to enhance the credibility, sustainability and impact of ML-based scientific research.
High-quality, "rich" metadata are essential for making research data findable, interoperable, and reusable. The Center for Expanded Data Annotation and Retrieval (CEDAR) has long addressed this need by providing tools to design machine-actionable metadata templates that encode community standards in a computable form. To make these capabilities more accessible within real-world research workflows, we have developed the CEDAR Embeddable Editor (CEE)-a lightweight, interoperable Web Component that brings structured, standards-based metadata authoring directly into third-party platforms. The CEE dynamically renders metadata forms from machine-actionable templates and produces semantically rich metadata in JSON-LD format. It supports ontology-based value selection via the BioPortal ontology repository, and it includes external authority resolution for persistent identifiers such as ORCIDs for individuals and RORs for research organizations. Crucially, the CEE requires no custom user-interface development, allowing deployment across diverse platforms. The CEE has been successfully integrated into generalist scientific data repositories such as Dryad and the Open Science Framework, demonstrating its ability to support discipline-specific metadata creation. By supporting the embedding of metadata authoring within existing research environments, the CEE can facilitate the adoption of community standards and help improve metadata quality across scientific disciplines.
Inspired by a proposal made almost ten years ago, this paper presents a model for classifying per-sonal data for research to inform researchers on how to manage them. The classification is based on the principles of the European General Data Protection Regulation and its implementation under the Spanish Law. The paper also describes in which conditions personal data may be stored and can be accessed ensuring compliance with data protection regulations and safeguarding privacy. The work has been developed collaboratively by the Library and the Data Protection Office. The outcomes of this collaboration are a decision tree for researchers and a list of requirements for research data re-positories to store and grant access to personal data securely. This proposal is aligned with the FAIR principles and the commitment for responsible open science practices.
This article details a correction to: Gandhi, S., Diggs, S., Córdoba, M.A., Bezuidenhout, L., Cobe, R., El Jadid, S., Peterson, B., Quick, R., Shanahan, H., Venkataraman, S., Okorafor, E. and Van den Eynden, V. (2026) ‘Building Responsible and Sustainable Open Data Literacy Skills for Early Career Researchers: A Decade of the SoRDS Programme’, Data Science Journal, 25(1), p. 12. https://doi.org/10.5334/dsj-2026-012.
FAIRness of research data, meaning that data are managed according to the principles of being Findable, Accessible, Interoperable, and Reusable, has become a ubiquitous requirement in research data policies as well as in general guidelines for research data management. Meeting this requirement largely depends on the availability of rich and standardized DDI-metadata—based on the Data Documentation Initiative family of metadata standards—which is of particular importance for tabular data resulting from surveys and other structured observations, is often lacking (Wenzig and Han, 2024). The lack of such metadata can largely be attributed to the absence of lightweight approaches that integrate its creation into existing data preparation workflows. Against this background, a project funded by KonsortSWD-NFDI4Society brought together three research data centers (SOEP, LIfBi, and FDZ-DZHW) to investigate the requirements for converting existing metadata into a standardized DDI format and publishing it using a common protocol (OAI-PMH). The results demonstrate that generating fine-grained DDI-metadata is easy to implement, even with limited resources. Contrary to expectations, OAI-PMH did not prove to be a straightforward approach for publishing metadata. Based on these findings, the authors evaluate FAIR signposting as an alternative approach, which shows considerable potential. The results of this study may therefore serve as a best-practice example for institutions, especially from survey-based research domains seeking to implement fine-grained standardized DDI metadata, as well as a starting point for further research on approaches to publishing fine-grained standardized metadata.
Synthetic data is a useful solution when data is scarce or private, as it supports reproducible experimentation, privacy-preserving data sharing, data re-purposing, and robust evaluation of data systems. This study presents a benchmark for tabular data synthesis (TDS) tools, evaluating their performance across six critical dimensions: handling dataset imbalance, dataset augmentation, handling missing values, privacy, machine learning (ML) utility, and computational performance. Our findings provide practical insights to guide tool selection based on specific use cases and constraints. We assessed 13 tools across 15 datasets from different use cases, focusing on prosumer hardware configurations for end-users and highlight the trade-offs among various TDS models. Sampling-based tools like SMOTE excelled in handling imbalance and efficiency but lacked privacy and variability. Hybrid and Transformer models demonstrated strong results across most dimensions but required substantial computational resources. Diffusion models achieved high scores but were complex to configure, while Bayesian Networks offered efficiency and privacy with limitations in utility. The study also emphasizes non-functional considerations such as runtime, resource efficiency, and configuration challenges. The source code and data have been made available at the Github Repository.