Results of scientific work in chemistry can usually be obtained in the form of materials and data. A big step towards transparency and reproducibility of the scientific work can be gained if scientists publish their data in research data repositories in a FAIR manner. Nevertheless, in order to make chemistry a sustainable discipline, obtaining FAIR data is insufficient and a comprehensive concept that includes preservation of materials is needed. In order to offer a comprehensive infrastructure to find and access data and materials that were generated in chemistry projects, we combined the infrastructure Chemotion repository with an archive for chemical compounds. Samples play a key role in this concept: we describe how FAIR metadata of a virtual sample representation can be used to refer to a physically available sample in a materials’ archive and to link it with the FAIR research data gained using the said sample. We further describe the measures to make the physically available samples not only FAIR through their metadata but also findable, accessible and reusable.
Publishing research data aims to improve the transparency of research results and facilitate the reuse of datasets. In both cases, referencing the datasets that were used is recommended. Research data repositories can support data referencing through various measures and also benefit from it, for example using this information to demonstrate their impact. However, the literature shows that the practice of formally citing research data is not widespread, data metrics are not yet established, and effective incentive structures are lacking. This article examines how often and in what form datasets published via the research data repository RADAR are referenced. For this purpose, the data sources Google Scholar, DataCite Event Data and the Data Citation Corpus were analyzed. The analysis shows that 27.9 were referenced at least once. 21.4 in the reference lists and are therefore considered data citations. Datasets were referenced often in data availability statements. A comparison of the three data sources showed that there was little overlap in the coverage of references. In most cases (75.8 in the same year. Two definition approaches were considered to investigate data reuse. 118 RADAR datasets were referenced more than once. Only 21 references had no overlaps in the authorship information – these datasets were referenced by researchers that were not involved in data collection.
The first funding period of NFDI4Chem established a robust foundation for research data management (RDM) in chemistry by promoting FAIR data principles and creating a cohesive infrastructure to capture well-annotated data early in the lab through electronic lab notebooks (ELNs) and making this data available in public repositories. Key achievements include standardised data formats and metadata, a federated repository environment, and improved data visibility and accessibility. Training programs and outreach have significantly increased awareness and adoption of best RDM practices. In the second funding period, the consortium aims to advance these achievements by consolidating this infrastructure, developing a model for its sustainable maintenance and operation, and fostering cultural change for its widespread adoption. Goals include ensuring seamless data workflows from laboratories to open repositories, enhancing interoperability, and supporting innovative research through AI-ready data. The work plan is organised into six task areas (TAs). TA1 (Management) provides leadership and supports all other TAs in achieving their objectives. TA2 (Smart Lab) aims to develop a fully digital research environment, including an ELN as a modular platform. This environment will support data collection, management, storage, analysis, and sharing. Integrating devices and external resources will enable seamless data transfer to repositories. TA3 (Repositories) will consolidate the repository ecosystem. The goal is to integrate repositories into a federated system for better accessibility and interoperability, ensuring long-term data availability and sustainability. TA4 (Metadata, Data Standards, and Publication Standards) focuses on developing and promoting new data and metadata standards in an international community process. This includes applying ontologies to create a semantic foundation for linking research data, making it machine-readable and enabling knowledge graphs. TA5 (Community and Training) is dedicated to fostering a cultural shift towards digital chemistry through continuous engagement, collecting requirements, and providing extensive training and support through workshops and open education resources. It will promote FAIR-compliant machine learning applications, embedding RDM into academic curricula to ensure future scientists are well-versed in these practices. TA6 (Synergies and Cross-Cutting Topics) aims to enhance collaboration across NFDI consortia and beyond. This includes developing ontologies, terminology services, the search service, and other cross-cutting solutions, integrating these developments into existing infrastructure, enabling interdisciplinary data harmonisation and fostering machine learning applications.
In the rapidly advancing domain of environmental research, the deployment of a comprehensive, state-of-the-art Research Data Management (RDM) framework is increasingly pivotal. Such a framework is key to ensure FAIR data, laying the groundwork for transparent and reproducible earth system sciences.Today, datasets associated with research articles are commonly published via prominent data repositories like Pangaea or Zenodo. Conversely, data used in actual day-to-day research and inter-institutional projects tends to be shared through basic cloud storage solutions or, even worse, via email. This practice, however, often conflicts with the FAIR principles, as much of this data ends up in private, restricted systems and local storage, limiting its broader accessibility and use.In response to this challenge, our research project Cat4KIT aims to establish a cross-institutional catalog and Research Data Management framework. The Cat4KIT framework is, hence, an important building block towards the FAIRification of environmental data. It not only streamlines the process of ensuring availability and accessibility of large-scale environmental datasets but also significantly enhances their value for interdisciplinary research and informed decision-making in environmental policy.The Cat4KIT system comprises four essential elements: data service provision, meta(data) harvesting, catalogue service, and user-friendly data presentation. The data service provision module is tailored to facilitate access to data within typical storage systems by using well-defined and standardized community interfaces via tools like the Thredds data server, Intake Catalogues, and the OGC SensorThings API. By this, we ensure seamless data retrieval and management for typical use-casers in environmental sciences.(Meta)data harvesting via our so-called DS2STAC-package entails collecting metadata from various data services, followed by creating STAC-metadata and integrating it into our STAC-API-based catalog service.This catalog service module synergizes diverse datasets into a cohesive, searchable spatial catalog, enhancing data discoverability and utility via our Cat4KIT UI.Finally, our framework's data portal is tailored to elevate data accessibility and comprehensibility for a wide audience, including researchers, enabling them to efficiently search, filter, and navigate through data from decentralized research data infrastructures.One notable characteristic of Cat4KIT is its dependence on open-source solutions and strict adherence to community standards. This guarantees not just the framework's ability to function well with current data systems but also its simple adaption and expansion to meet future needs. Our presentation demonstrates the technical structure of Cat4KIT, examining the development and integration of each module to adhere to the FAIR principles. Additionally, it showcases examples to illustrate the practical use of the framework in real-life situations, emphasizing its efficacy in enhancing data management practices within KIT and its potential relevance in other research organizations.
The progress of the DFG-funded NFDI4Chem consortium (NFDI 4/1 - project number 441958208) in data management in chemistry is outlined in our latest report, highlighting the steps we have taken to integrate a data-centric approach within the chemistry community. This interim report offers a comprehensive overview of our data management activities, covering the reporting period from October 2020 to August 2023. The shift to digital tools in research documentation is driven by our work with Electronic Laboratory Notebooks (ELNs), such as Chemotion ELN, offering systematic data storage for easy retrieval and sharing. Additionally, we focus on developing repositories, such as Chemotion repository and RADAR4Chem, which fulfil the needs for the storage of chemical data. The NFDI4Chem Search Service ensures easy data access from our repositories. Our efforts extend to community engagement through conference visits and online presence, aimed at creating awareness for (digital) research data management and connecting to chemistry students and researchers. Our training programs have reached over 600 participants to date. Initiatives like the FAIR4Chem award and the Chemistry Data Days promote cultural change towards FAIR data. Our Editors4Chem initiative collaborates with publishers for standardised data management and the Ontologies4Chem workshops organised by our consortium promote the ontology development in the field. Apart from the consortium's engagement for chemists, NFDI4Chem members played key roles in the development of the NFDI as a whole. Being actively involved in the sections and task forces, NFDI4Chem promotes collaborative solutions across NFDI consortia.
The Chemistry consortium NFDI4Chem aims to digitalise key steps in chemical research, supporting scientists in managing research data throughout its life cycle. The SmartLab, embedded in a federation of services, integrates various tools such as electronic lab notebooks, data repositories, and search services, to create a smart lab environment for structured data gathering. Utilizing terminology services and adhering to data format standards, NFDI4Chem promotes secure and FAIR data sharing, fostering collaboration and expediting scientific discoveries. This development is supported by community building measures, workshops, and training initiatives, along with collaboration on international minimum information standards.
Abstract Research data provide evidence for the validation of scientific hypotheses in most areas of science. Open access to them is the basis for true peer review of scientific results and publications. Hence, research data are at the heart of the scientific method as a whole. The value of openly sharing research data has by now been recognized by scientists, funders and politicians. Today, new research results are increasingly obtained by drawing on existing data. Many organisations such as the Research Data Alliance (RDA), the goFAIR initiative, and not least IUPAC are supporting and promoting the collection and curation of research data. One of the remaining challenges is to find matching data sets, to understand them and to reuse them for your own purpose. As a consequence, we urgently need better research data management.
A contemporary and flexible Research Data Management (RDM) framework is required to make environmental research data Findable, Accessible, Interoperable, and Reusable (FAIR) and, hence, provide the foundation for open and reproducible earth system sciences. While data-sets that accompany scientific articles are typically published via large data repositories like Pangaea or Zenodo, intermediate, day-to-day, or actively-used data (e.g., data from research projects or prototypical data) is still exchanged via simple cloud storage services and email. And while the FAIR principles require data to be openly findable and accessible, it is often only available within closed and restricted infrastructures and local file systems.Our research project Cat4KIT hence aims to develop a cross-institutional catalog and RDM framework for the FAIRification of such day-to-day research data. This framework is comprised of four modules / services for providing access to data on storage systems through well-defined and standardized interfaces harvesting and transforming (meta)data into standardized formats making (meta)data accessible to the public using well-defined and standardized catalog services and interfaces enabling users to search, filter, and explore data from decentralized research data infrastructures. We develop, implement and evaluate each of these four modules within an inter-institutional consortium consisting of scientists, software developers and potential end-users. This allows us to include a wide-range of research data from multi-dimensional climate model outputs to high-frequency in-situ measurements. We emphasize the application of existing open-source solutions and community standards for data interfaces (THREDDS, STA, S3), (meta)data schemes, and catalog services (Spatio-Temporal Assets Catalog - STAC) in order to ensure an easy integration of research data into the Cat4KIT-framework and a straightforward extension to further research data infrastructures.In our presentation, we demonstrate the current status of our Cat4KIT-framework as an inter-institutional research data management and catalog platform for the FAIRification of day-to-day research data.
The collection of metadata for research data is an important aspect in the FAIR principles. The schema.org and Bioschemas initiatives created a vocabulary to embed markup for many different types, including BioChemEntity, ChemicalSubstance, Gene, MolecularEntity, Protein, and others relevant in the Natural and Life Sciences with immediate benefits for findability of data packages. To bridge the gap between the worlds of semantic-web-driven JSON+LD metadata on the one hand, and established but separately developed interface services in libraries, we have designed an architecture for harmonising, federating and harvesting metadata from several resources. Our approach is to serve JSON+LD embedded in an XML container through a central OAI-Provider. Several resources in NFDI4Chem provide such domain-specific metadata. The CKAN-based NFDI4Chem search service can harvest this metadata using an OAI-PMH harvester extension that can extract the XML-encapsulated JSON+LD metadata, and has search capabilities relevant in the chemistry domain. We invite the community to collaborate and reach a critical mass of providers and consumers in the NFDI.
The research data repository RADAR is designed to support the secure management, archiving, publication and dissemination of digital research data from completed scientific studies and projects. Developed as a collaborative project funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) (2013-2016), the system is operated by FIZ Karlsruhe - Leibniz Institute for Information Infrastructure - and currently serves as a generic cloud service for about 20 universities and non-university research institutions. Since its launch, RADAR has witnessed significant changes in the landscape of research data repositories and the evolving needs of researchers, research communities and institutions. In our presentation within the “Enabling RDM” Track, we will show how RADAR is responding to these dynamic changes. In order to create a sufficiently large user base for the sustainable operation of the system, we have moved RADAR away from its previous single focus on a discipline-agnostic cloud service and towards a demand-driven functional optimisation. In 2021, we introduced an additional operating model for institutions (RADAR Local), where we operate a separate RADAR instance locally at the institution site exclusively using the institutional IT-infrastructure. In 2022 we opened up RADAR to new target groups with community-specific service offerings, in particular in the context of the National Research Data Infrastructure (NFDI). Beside the expansion of the functional scope, our ongoing development work focuses also on strengthening the system's support for the FAIR principles [1] and the concepts of FAIR Digital Objects (FDO) [2] and Schema.org. Our presentation will outline recent RADAR developments and achievements as well as future plans thus providing solutions and synergy potential for the scientific community and for other service providers.
The Chemistry consortium NFDI4Chem aims to digitalise key steps in chemical research, supporting scientists in managing research data throughout its life cycle. The SmartLab, embedded in a federation of services, integrates various tools such as electronic lab notebooks, data repositories, and search services, to create a smart lab environment for structured data gathering. Utilizing terminology services and adhering to data format standards, NFDI4Chem promotes secure and FAIR data sharing, fostering collaboration and expediting scientific discoveries. This development is supported by community building measures, workshops, and training initiatives, along with collaboration on international minimum information standards.
Results of scientific work in chemistry can usually be obtained in the form of materials and data. A big step towards transparency and reproducibility of the scientific work can be gained if scientists publish their data in a FAIR (Findable, Accessible, Interoperable, Reusable) manner in research data repositories. Nevertheless, in order to make chemistry as a discipline sustainable, obtaining FAIR data is insufficient and a comprehensive concept including the preservation of materials is needed. We describe in this article how we combined two infrastructures, a repository for research data (Chemotion repository) and an archive for chemical compounds (Molecule Archive), in order to offer a comprehensive infrastructure to find and access data and materials that were generated in chemistry projects. Samples play a key role in this concept: we describe how FAIR metadata of a virtual sample representation can be used to refer to the physically available sample stored in a materials’ archive and to link FAIR research data gained with the sample. We further describe the measures to make the physically available samples not only FAIR through the sample’s metadata but also accessible and reusable in the form of their material for others.
The Chemistry consortium NFDI4Chem aims to digitalise key steps in chemical research, supporting scientists in managing research data throughout its life cycle. The SmartLab, embedded in a federation of services, integrates various tools such as electronic lab notebooks, data repositories, and search services, to create a smart lab environment for structured data gathering. Utilizing terminology services and adhering to data format standards, NFDI4Chem promotes secure and FAIR data sharing, fostering collaboration and expediting scientific discoveries. This development is supported by community building measures, workshops, and training initiatives, along with collaboration on international minimum information standards.
This work describes the setup of an advanced technical infrastructure for collaborative software development (CDE) in large, distributed projects based on GitLab. We present its customization and extension, additional features and processes like code review, continuous automated testing, DevOps practices, and sustainable life-cycle management including long-term preservation and citable publishing of software releases along with relevant metadata. The environment is currently used for developing the open cardiac simulation software openCARP and an evaluation showcases its capability and utility for collaboration and coordination of sizeable heterogeneous teams. As such, it could be a suitable and sustainable infrastructure solution for a wide range of research software projects.
Research data management (RDM) is needed to assist experimental advances and data collection in the chemical sciences. Many funders require RDM because experiments are often paid for by taxpayers and the resulting data should be deposited sustainably for posterity. However, paper notebooks are still common in laboratories and research data is often stored in proprietary and/or dead-end file formats without experimental context. Data must mature beyond a mere supplement to a research paper. Electronic lab notebooks (ELN) and laboratory information management systems (LIMS) allow researchers to manage data better and they simplify research and publication. Thus, an agreement is needed on minimum information standards for data handling to support structured approaches to data reporting. As digitalization becomes part of curricular teaching, future generations of digital native chemists will embrace RDM and ELN as an organic part of their research.
RADAR is a cross-disciplinary internet-based service for long-term and format-independent archiving and publishing of digital research data from scientific studies and projects. The focus is on data from disciplines that are not yet supported by specific research data management infrastructures. The repository aims to ensure access and long-term availability of deposited datasets according to FAIR criteriaWilkinson et al. 2016 for the benefit of the scientific community. Published datasets are retained for at least 25 years; for archived datasets, the retention period can be flexibly selected up to 15 years. The RADAR Cloud service was developed as a cooperation project funded by the DFG (2013-2016) and started operations in 2017. It is operated by FIZ Karlsruhe - Leibniz-Institute for Information Infrastructure.As a distributed, multilayer application, RADAR is structured into a multitude of services and interfaces. The system architecture is modular and consists of a user interface (frontend), management layer (backend) and storage layer (archive), which communicate with each other via application programming interfaces (API). This open structure and the access to the APIs from outside allows integrating RADAR into existing systems and work processes, e. g. for automated upload of metadata from other applications using the RADAR API. RADAR's storage layer is encapsulated via the Data Center API. This approach guarantees independence from a specific storage technology and makes it possible to integrate alternative archives for the bitstream preservation of the research data.The data transfer to RADAR takes place in two steps: In the first step, the data is transferred to a temporary work storage. The ingest service accepts individual files and packed archives, optionally unpacks them while retaining the original directory structure and creates a dataset. For each file found, the MIME Type (see Multipurpose Internet Mail Extensions specification)) is analysed and stored in the technical metadata. When archiving and publishing, a dataset is created in the second step. The structure of this dataset - the AIP (archival information package) in the sense of the OAIS standard - corresponds to the BagIt standard. It contains, in addition to the actual research data in original order, technical and descriptive metadata (if created) for each file or directory as well as a manifest within one single TAR ("tape archive", a unix archiving format and utility) file as an entity in one place. This TAR file is stored permanently on magnetic tapes redundantly in three copies at different locations in two academic computing centres.The FAIR Principles are currently being given special importance in the research community. They define measures that ensure the optimal processing of research data, accessibility for both humans and machines, as well as reusability for further research. RADAR also promotes the implementation of the FAIR Principles with different measures and functional features, amongst others:Descriptive metadata are recorded using the internal RADAR Metadata Schema (based on DataCite Metadata Schema 4.0), which supports 10 mandatory and 13 optional metadata fields. Annotations can be made on the dataset level and on the individual files and folders level. A user licence which rules re-use of the data, must be defined for each dataset. Each published dataset receives a DOI which is registered with DataCite. RADAR metadata uses a combination of controlled lists and free text entries. Author identification is ensured by using an ORCID ID and funder identification by CrossRef Open Funder Registry. More interfacing options, e.g. ROR and the Integrated Authority File (GND) are currently implemented. Datasets can be easily linked with other digital resources (e.g. text publications) via a “related identifier”. To maximise data dissemination and discoverability, the metadata of published datasets are indexed in various formats (e.g. DataCite and DublinCore) and offered for public metadata harvesting e.g. via an OAI-provider.These measures are - to our minds - undoubtedly already significant, but not yet sufficient in the medium to long term. Especially in terms of interoperability, we see development potential for RADAR. The FAIR Digital Object (FDO) Framework seems to offer a promising concept, especially to further promote data interoperability and to close respective gaps in the current infrastructure and repository landscape.RADAR aims to participate in this community driven approach also in its role within the National Research Data Infrastructure (NFDI). As part of the NFDI, RADAR already plays a relevant role as a generic infrastructure service in several NFDI consortia (e.g. NFDI4Culture and NFDI4Chem). With RADAR4Chem and RADAR4Culture, FIZ Karlsruhe for example offers researchers from chemistry and the cultural sciences low-threshold data publication services based on RADAR. We successively develop these services further according to the needs of the communities, e.g. by integrating and linking them with subject-specific terminologies, by providing annotation options with subject-specific metadata or by enabling selective reading or previewing options for individual files in existing datasets.In our presentation, we would like to describe the present and future functionality of RADAR and its current level of FAIRness as possible starting points for further discussion with the FDO community with regard to the implementation of the FDO framework for our service.
This work describes the setup of an advanced technical infrastructure for collaborative software development (CDE) in large, distributed projects based on GitLab. We present its customization and extension, additional features and processes like code review, continuous automated testing, DevOps practices, and sustainable life-cycle management including long-term preservation and citable publishing of software releases along with relevant metadata. The environment is currently used for developing the open cardiac simulation software openCARP and an evaluation showcases its capability and utility for collaboration and coordination of sizeable heterogeneous teams. As such, it could be a suitable and sustainable infrastructure solution for a wide range of research software projects.
AbstractForschungsdatenmanagement (FDM) ist erforderlich, um wissenschaftlichen Fortschritt und das Sammeln von Daten zu fördern. Viele Fördergeldgeber verlangen FDM, da Forschung oft durch Steuergelder finanziert wird und daraus resultierende Daten nachhaltig für kommende Generationen hinterlegt werden sollten. In heutigen Laboren sind Papier‐Laborbücher allerdings noch immer gängige Praxis, und die Forschungsdaten werden oft in proprietären Dateiformaten ohne experimentellen Kontext gespeichert. Daten müssen über eine Rolle als bloßes Anhängsel von Publikationen hinauswachsen. Elektronische Laborbücher (ELN) und Laborinformationsmanagementsysteme (LIMS) ermöglichen es, Daten besser zu verwalten sowie Forschung und Veröffentlichung zu vereinfachen. Die in gutes FDM investierte Zeit zahlt sich später in vielerlei Hinsicht aus. Daher werden Mindestinformationsstandards (MI) für die Handhabung von Daten benötigt. Die Digitalisierung hält gerade Einzug in die Lehrpläne, sodass zukünftige Generationen von ChemikerInnen mit FDM und ELN‐Nutzung vertraut sein werden.
Christoph Steinbeck合作论文数EMBL Outstation - Hinxton,
European Bioinformatics Institute,
Wellcome Trust Genome Campus11