Rich, standardised metadata is essential for improving the findability, accessibility, interoperability, and reusability (FAIR) of health research resources. The European Health Data Space (EHDS) requires harmonised catalogue metadata across countries, yet national implementations demonstrating how DCAT‑based standards can be operationalised in practice are still limited. To develop, document, and demonstrate a DCAT‑based metadata schema for a National Health Data Catalogue, using the Dutch Health‑RI National Health Data Infrastructure as a concrete implementation example, and to assess how this approach supports interoperability and FAIR‑aligned metadata publication. Metadata requirements were gathered from Dutch health institutions and mapped to DCAT, DCAT‑AP, HealthDCAT‑AP and DCAT‑AP‑NL. A multi‑stakeholder modelling process involving semantic experts, ontology engineers and data stewards produced a core schema and domain‑specific extensions. RDF and SHACL were used for validation, and FAIR Data Points enabled decentralised metadata publication. Several pilot use cases were onboarded to evaluate applicability, interoperability and usability in real‑world settings. The resulting schema comprises DCAT‑aligned classes and expanded mandatory and recommended fields aligned with HealthDCAT‑AP, the metadata model supported by EHDS. The model supports both general catalogue metadata and evolving domain‑specific extensions. Pilot implementations across Dutch institutions demonstrated improved metadata consistency, enhanced resource discoverability and successful interoperability between local FAIR Data Points and the National Health Data Catalogue. The developed DCAT‑based schema provides a scalable, standards‑aligned foundation for a National Health Data Catalogue and supports cross‑infrastructure interoperability mandated by the EHDS. The Dutch implementation shows how structured metadata and coordinated national governance can enhance FAIRness and improve access to health research resources. The approach offers a practical template for national Health Data Access Bodies across Europe. Future work includes completing domain‑specific extensions, increasing automation in metadata generation and validation, and strengthening documentation and training to support sustainable, community‑driven adoption.
Use of the FAIR principles (Findable, Accessible, Interoperable and Reusable) allows the rapidly growing number of biomedical datasets to be optimally (re)used. An important aspect of the FAIR principles is metadata. The FAIR Data Point specifications and reference implementation have been designed as an example on how to publish metadata according to the FAIR principles. Metadata can be added to a FAIR Data Point with the FDP’s web interface or through its API. However, these methods are either limited in scalability or only usable by users with a background in programming. We aim to provide a new tool for populating FDPs with metadata that addresses these limitations with the FAIR Data Point Populator. The FAIR Data Point Populator consists of a GitHub workflow together with Excel templates that have tooltips, validation and documentation. The Excel templates are targeted towards non-technical users, and can be used collaboratively in online spreadsheet software. A more technical user then uses the GitHub workflow to read multiple entries in the Excel sheets, and transform it into machine readable metadata. This metadata is then automatically uploaded to a connected FAIR Data Point. We applied the FAIR Data Point Populator on the metadata of two datasets, and a patient registry. We were then able to run a query on the FAIR Data Point Index, in order to retrieve one of the datasets. The FAIR Data Point Populator addresses the limitations of the other metadata publication methods by allowing the bulk creation of metadata entries while remaining accessible for users without a background in programming. Additionally, it allows efficient collaboration. As a result of this, the barrier of entry for FAIRification is lower, which allows the creation of FAIR data by more people.
Since 2014, “Bring Your Own Data” workshops (BYODs) have been organised to inform people about the process and benefits of making resources Findable, Accessible, Interoperable, and Reusable (FAIR, and the FAIRification process). The BYOD workshops’ content and format differ depending on their goal, context, and the background and needs of participants. Data-focused BYODs educate domain experts on how to make their data FAIR to find new answers to research questions. Management-focused BYODs promote the benefits of making data FAIR and instruct project managers and policy-makers on the characteristics of FAIRification projects. Software-focused BYODs gather software developers and experts on FAIR to implement or improve software resources that are used to support FAIRification. Overall, these BYODs intend to foster collaboration between different types of stakeholders involved in data management, curation, and reuse (e.g. domain experts, trainers, developers, data owners, data analysts, FAIR experts). The BYODs also serve as an opportunity to learn what kind of support for FAIRification is needed from different communities and to develop teaching materials based on practical examples and experience. In this paper, we detail the three different structures of the BYODs and describe examples of early BYODs related to plant breeding data, and rare disease registries and biobanks, which have shaped the structure of the workshops. We discuss the latest insights into making BYODs more productive by leveraging our almost ten years of training experience in these workshops, including successes and encountered challenges. Finally, we examine how the participants’ feedback has motivated the research on FAIR, including the development of workflows and software.
ABSTRACTWhile the FAIR Principles do not specify a technical solution for ‘FAIRness’, it was clear from the outset of the FAIR initiative that it would be useful to have commodity software and tooling that would simplify the creation of FAIR-compliant resources. The FAIR Data Point is a metadata repository that follows the DCAT(2) schema, and utilizes the Linked Data Platform to manage the hierarchical metadata layers as LDP Containers. There has been a recent flurry of development activity around the FAIR Data Point that has significantly improved its power and ease-of-use. Here we describe five specific tools—an installer, a loader, two Web-based interfaces, and an indexer—aimed at maximizing the uptake and utility of the FAIR Data Point.
ABSTRACTMetadata, data about other digital objects, play an important role in FAIR with a direct relation to all FAIR principles. In this paper we present and discuss the FAIR Data Point (FDP), a software architecture aiming to define a common approach to publish semantically-rich and machine-actionable metadata according to the FAIR principles. We present the core components and features of the FDP, its approach to metadata provision, the criteria to evaluate whether an application adheres to the FDP specifications and the service to register, index and allow users to search for metadata content of available FDPs.
Background The COVID-19 pandemic has challenged healthcare systems and research worldwide. Data is collected all over the world and needs to be integrated and made available to other researchers quickly. However, the various heterogeneous information systems that are used in hospitals can result in fragmentation of health data over multiple data ‘silos’ that are not interoperable for analysis. Consequently, clinical observations in hospitalised patients are not prepared to be reused efficiently and timely. There is a need to adapt the research data management in hospitals to make COVID-19 observational patient data machine actionable, i.e. more Findable, Accessible, Interoperable and Reusable (FAIR) for humans and machines. We therefore applied the FAIR principles in the hospital to make patient data more FAIR. Results In this paper, we present our FAIR approach to transform COVID-19 observational patient data collected in the hospital into machine actionable digital objects to answer medical doctors’ research questions. With this objective, we conducted a coordinated FAIRification among stakeholders based on ontological models for data and metadata, and a FAIR based architecture that complements the existing data management. We applied FAIR Data Points for metadata exposure, turning investigational parameters into a FAIR dataset. We demonstrated that this dataset is machine actionable by means of three different computational activities: federated query of patient data along open existing knowledge sources across the world through the Semantic Web, implementing Web APIs for data query interoperability, and building applications on top of these FAIR patient data for FAIR data analytics in the hospital. Conclusions Our work demonstrates that a FAIR research data management plan based on ontological models for data and metadata, open Science, Semantic Web technologies, and FAIR Data Points is providing data infrastructure in the hospital for machine actionable FAIR digital objects. This FAIR data is prepared to be reused for federated analysis, linkable to other FAIR data such as Linked Open Data, and reusable to develop software applications on top of them for hypothesis generation and knowledge discovery.
This document is the second report from Task 2.2 of the FAIRsFAIR project. It demonstrates how the features for FAIR enabling repositories (Behnke et al., 2020) are put into practice by building a FAIR Data Point (FDP) prototype 1 based on the DCAT2 data model, which is an RDF vocabulary designed to facilitate interoperability between data catalogues published on the Web. The root of the FAIRsFAIR reference implementation is a FAIR Data Point. In general, it serves three goals: It allows a repository or any other holder or editor2 of data to expose metadata in a FAIR manner, with a strong focus on the F, A, and R as described in the FAIR Principles (Wilkinson et al., 2016) A consumer or viewer of the data can discover information that is stored in it. It is optimised to interact with humans and machines. The first chapter gives a recap on the motivation and maps requirements from Behnke et al. (2020) to the prototype. The second chapter explains the reference implementation’s technical details, while the third describes the DCAT2 data model. In the fourth chapter, the authors analyse the community uptake, and in the fifth and last chapter, an outlook is presented.
BACKGROUND:Integration of heterogenous resources is key for Rare Disease research. Within the EJP RD, common Application Programming Interface specifications are proposed for discovery of resources and data records. This is not sufficient for automated processing between RD resources and meeting the FAIR principles.OBJECTIVE:To design a solution to improve FAIR for machines for the EJP RD API specification.METHODS:A FAIR Data Point is used to expose machine-actionable metadata of digital resources and it is configured to store its content to a semantic database to be FAIR at the source.RESULTS:A solution was designed based on grlc server as middleware to implement the EJP RD API specification on top of the FDP.CONCLUSION:grlc reduces potential API implementation overhead faced by maintainers who use FAIR at the source.
Since their publication in 2016 we have seen a rapid adoption of the FAIR principles in many scientific disciplines where the inherent value of research data and, therefore, the importance of good data management and data stewardship, is recognized. This has led to many communities asking "What is FAIR?" and "How FAIR are we currently?", questions which were addressed respectively by a publication revisiting the principles and the emergence of FAIR metrics. However, early adopters of the FAIR principles have already run into the next question: "How can we become (more) FAIR?" This question is more difficult to answer, as the principles do not prescribe any specific standard or implementation. Moreover, there does not yet exist a mature ecosystem of tools, platforms and standards to support human and machine agents to manage, produce, publish and consume FAIR data in a user-friendly and efficient (i.e., "easy") way. In this paper we will show, however, that there are already many emerging examples of FAIR tools under development. This paper puts forward the position that we are likely already in a creolization phase where FAIR tools and technologies are merging and combining, before converging in a subsequent phase to solutions that make FAIR feasible in daily practice.
The synthetic dataset that models COVID-19 real world observations from WHO COVID-19 RAPID Version CRFs of hospitalized patients for the hypothesis under study, originally created by the TWOC project.
We report on the activities of the 2015 edition of the BioHackathon, an annual event that brings together researchers and developers from around the world to develop tools and technologies that promote the reusability of biological data. We discuss issues surrounding the representation, publication, integration, mining and reuse of biological data and metadata across a wide range of biomedical data types of relevance for the life sciences, including chemistry, genotypes and phenotypes, orthology and phylogeny, proteomics, genomics, glycomics, and metabolomics. We describe our progress to address ongoing challenges to the reusability and reproducibility of research results, and identify outstanding issues that continue to impede the progress of bioinformatics research. We share our perspective on the state of the art, continued challenges, and goals for future research and development for the life sciences Semantic Web.
We report on the activities of the 2015 edition of the BioHackathon, an annual event that brings together researchers and developers from around the world to develop tools and technologies that promote the reusability of biological data. We discuss issues surrounding the representation, publication, integration, mining and reuse of biological data and metadata across a wide range of biomedical data types of relevance for the life sciences, including chemistry, genotypes and phenotypes, orthology and phylogeny, proteomics, genomics, glycomics, and metabolomics. We describe our progress to address ongoing challenges to the reusability and reproducibility of research results, and identify outstanding issues that continue to impede the progress of bioinformatics research. We share our perspective on the state of the art, continued challenges, and goals for future research and development for the life sciences Semantic Web.
1 Leiden University Medical Centre, The Netherlands {m.roos,m.thompson,r.kaliyaperumal,a.jacobsen}@lumc.nl 2 Universidad Politcnica de Madrid, Spain markw@illuminae.com 3 Istituto Superiore di Sanitá, Italy claudio.carta@iss.it 4 University Medical Center Groningen, The Netherlands david.van.enckevort@umcg.nl 5 Wageningen Plant Research, The Netherlands richard.finkers@wur.nl 6 Dutch Techcentre for Life Sciences, The Netherlands {luiz.bonino,erik.schultes,mascha.jansen}@dtls.nl
Principles of Findable, Accessible, Interoperable, and Reusable data for humans and computers (FAIR)1are widely endorsed by organizations such as the Euro- pean Open Science Cloud, the life science data infrastructure ELIXIR, the NIH via its commons program, the biobanking infrastructure consortium BBMRI- ERIC, the G20 and the G7. Implementing a data ecosystem based on FAIR principles requires guidelines, tools, and training, and FAIR data stewards to help apply them. The principles as such do not recommend any particular im- plementation: user communities will have to decide the most appropriate imple- mentation for their domain. Here, we demonstrate the use of a suite of Semantic Web-based middle-ware services that help communities implement FAIR data principles2. Aiming to facilitate adoption, the services are made to complement existing data infrastructures, including local and centralised data resources, and thus establish a robust, federated ecosystem of FAIR resources. The services are also particularly suited for training data stewards. We demonstrate the ap- plication of the services by rare disease and plant breeding communities where the combination of Ontologies, Linked Data, and light-weight FAIR services are being explored as the means to implement FAIR data principles.
The discovery of new medicines requires pharmacologists to interact with a number of information sources ranging from tabular data to scientific papers, and other specialized formats. In this application report, we describe a linked data platform for integrating multiple pharmacology datasets that form the basis for several drug discovery applications. The functionality offered by the platform has been drawn from a collection of prioritised drug discovery business questions created as part of the Open PHACTS project, a collaboration of research institutions and major pharmaceutical companies. We describe the architecture of the platform focusing on seven design decisions that drove its development with the aim of informing others developing similar software in this or other domains. The utility of the platform is demonstrated by the variety of drug discovery applications being built to access the integrated data.An alpha version of the OPS platform is currently available to the Open PHACTS consortium and a first public release will be made in late 2012, see http://www.openphacts.org/ for details.
We present the Open PHACTS linked data platform that is being developed to address a set of example drug discovery research questions and which supports several drug discovery applications. The platform retrieves data from many complementary, but overlapping, data sources to present an integrated view of the data. The platform exploits two entity resolution services: respectively for transforming text and chemical structures to a concept. The single concept URI provided by the resolution service is then expanded to a set of equivalent URIs used by the data sources.Availability. An alpha version is currently available to the Open PHACTS consortium. A first public release of the platform will be made in late 2012, see http://www.openphacts.org/.
Barend Mons合作论文数University of Rotterdam and7