Key Performance Indicators (KPIs) are essential for evaluating project success and establishing control mechanisms to monitor development, performance, and user acceptance of services in joint projects. However, the absence of standardized frameworks and effective monitoring tools, combined with service providers' reluctance due to fears of comparability, has limited their adoption in scientific contexts. To address this gap, we developed Scorpion, a flexible tool for KPI monitoring in project management. Scorpion enables service providers to retain control over their metrics while supporting centralized reporting. It offers both web-based and programmatic access, with features for KPI submission, visualization, and user and service management. Initially created for bioinformatics and biodiversity projects, Scorpion is applicable across diverse domains. It is particularly valuable for initiatives like the German National Research Data Infrastructure (NFDI), where funding agencies require KPI reporting for evaluation. We present the Scorpion framework, highlighting its design principles, features, and potential to improve project management practices. Use cases illustrate how Scorpion enhances KPI monitoring efficiency and accuracy, contributing to better impact evaluation, quality assurance, and informed decision-making in project and service management.
The current landscape of animal phenomics is characterised by a substantial lack of standardisation, hindering data reuse, reproducibility, and interoperability across studies, all of which are particularly important in light of the 3Rs principles for animal experiments (replace, reduce, refine). Within ELIXIR, the Domestic Animals Genome and Phenome Focus Group emerged to establish standardised practices that enhance the quality and interoperability of animal research data. In this context, the ISA model presents a robust, domain-agnostic framework well-established in the life sciences for describing experimental metadata. Notably, other scientific communities, such as the ELIXIR Plant and Metabolomics Communities (MIAPPE, PhenoMeNal), have successfully leveraged the ISA model to improve the consistency and usability of their metadata. Our project aims to develop a minimal information checklist tailored specifically for phenomics, facilitating the integration of diverse datasets, including recirculation systems in agriculture, and fostering collaborative research efforts. We will focus on various goals. Identifying essential aspects of animal phenotyping, informed by existing frameworks and community input. We aim to produce a concise and practical checklist that can be readily adopted by researchers, and promote a culture of standardisation. Mapping the checklist to the ISA model ensures alignment with established standards, promotes interoperability and facilitates data reuse while improving the overall quality of research outputs. Adopting existing ISA tools streamlines the implementation of our metadata checklist, providing user-friendly interfaces for researchers to manage, document, and share animal phenotyping data efficiently.
Robust validation of both research data and its accompanying metadata is essential for ensuring adherence to FAIR principles. Current approaches often handle these aspects separately, hindering a holistic quality assessment. Building upon previous BioHackathon work establishing ARCs (Annotated Research Context) as RO-Crates (ARC RO-Crate), we aim to develop and demonstrate an integrated validation strategy for FAIR digital objects. It distinguishes between validating the metadata descriptor and the payload data files. For the metadata descriptor, validation will ensure structural and semantic compliance to the base RO-Crate specification and the ARC-ISA family of RO-Crate profiles, using and extending the RO-Crate validator tool. For the payload data files, validation targets the actual content, since data files often require domain-specific structural and value constraints, which requires explicit schema definitions. For this, we will integrate Frictionless for checking data content against community standards (e.g. MIAPPE, as demonstrated in the HORIZON project AGENT). Crucially, this project will also explore mechanisms for specifying expected data structures’ requirements within the ARC RO-Crate itself. This aims to provide a more self-contained description of data, investigating how such internal requirements can be linked to data validation frameworks, complementing the crate’s metadata validation. The overall goal is to provide a powerful, holistic validation mechanism for ARC RO-Crates, enhancing their reliability, trustworthiness, and FAIRness. A MIAPPE-compliant plant phenomics dataset will serve as a use case. This integrated validation approach aims to streamline quality control for researchers and will be packaged as a deployable microservice, offering broad applicability across diverse research workflows.
As part of the BioHackathon Germany 2025, we report here about the progress of Project 8 - Enhancing e!DAL-PGP: A Modern Data Submission Platform for Plant Science Research Data during the event. The increasing volume of data generated in plant research underscores the necessity for efficient data management and sharing solutions. The de.NBI Service e!DAL-PGP (Arend et al., 2016, p. Arend2020) serves as a critical research data repository, facilitating the storage, management, and dissemination of plant research data. However, the current implementation faces significant challenges concerning the submission process and the provision of a submission tool for different operating systems, which complicate user interactions and hinder data contribution. A primary issue with the existing e!DAL-PGP service is the cumbersome nature of maintaining and deploying a submission tool across various OS environments. This requirement necessitates extensive effort to build, test, and provide the application for each platform. Consequently, this fragmentation can lead to delays and inconsistencies in the submission process, ultimately hindering researchers from effectively submitting their valuable data to the repository. To address these challenges, this project proposes the development of a unified and user-friendly web submission tool that streamlines the data submission process to eliminate the complexities associated with OS-specific requirements and to ensure that all users can submit their data seamlessly. This simplifies the submission process and enhances usability by focussing on improving the design and functionality. A well-structured and user-centric form is essential for facilitating accurate and complete data submissions. The current interface lacks features that enhance user experience, such as lookup services, contextual help, and clear instructions. By incorporating these elements, we aim to create a more efficient and engaging submission experience, encouraging researchers to contribute their valuable data without unnecessary complexity. This initiative aligns closely with the goals of de.NBI, which emphasizes the provision of high-quality bioinformatics services and the facilitation of FAIR Research Data Management (RDM). Enhancing the e!DAL-PGP service will streamline the data submission process and promote a culture of collaboration and data sharing within the plant research community.
Abstract The NFDI-consortium FAIRagro has established a systematic framework for evaluating the FAIRness of Research Data Infrastructures (RDI) within the German agrosystem research landscape. While FAIR principles are widely accepted, their practical implementation by RDIs remains challenging. By operationalizing the FAIR principles into a reproducible multi-dimensional scoring methodology, this initiative addresses the critical need for a transparent and citable benchmark of RDIs that moves beyond simple compliance. This paper details the underlying assessment criteria, comprising 20 aggregated core metrics, the iterative community-driven validation process, and the integration of these metrics into the FAIRagro Search Hub. This framework evaluates RDIs, like repositories or databases, instead of sampling hosted data sets, across the four distinct categories of FAIR independently, yielding granular, pillar-specific ratings. Unlike aggregate scoring models, which can inadvertently mask technical deficiencies by averaging performance across categories, this multi-dimensional approach ensures that a repository’s distinct strengths and bottlenecks remain fully visible. Our findings demonstrate that standardized scoring not only clarifies data accessibility for users but also highlights specific operational gaps, allowing repository providers to identify precisely where the service implementation can be enhanced. By establishing this data-driven service in the agronomy domain, we provide a scalable template for the broader NFDI and EOSC ecosystems to foster a culture of excellence in research data stewardship.
The ISA framework is well-established in Life Sciences domains as it has a generic and hierarchical structure, which allows describing experimental workflows in a uniform way. Nevertheless, the manual creation of ISA files requires a deep understanding of the ISA data model and knowledge about the specific data domain particularities, which makes it only suitable for data experts. Existing frameworks such as isa-tools or isa4j are powerful, but mainly made for developers and not for common data producers. We develop a user-friendly tool enabling scientists who are not experts for ISA to annotate their datasets and to export them in an ISA-compatible format.
Robust validation of both research data and its accompanying metadata is essential for ensuring adherence to FAIR principles. Current approaches often handle these aspects separately, hindering a holistic quality assessment. Building upon previous BioHackathon work establishing ARCs (Annotated Research Context) as RO-Crates (ARC RO-Crate), we aim to develop and demonstrate an integrated validation strategy for FAIR digital objects. It distinguishes between validating the metadata descriptor and the payload data files.For the metadata descriptor, validation will ensure structural and semantic compliance to the base RO-Crate specification and the ARC-ISA family of RO-Crate profiles, using and extending the RO-Crate validator tool.For the payload data files, validation targets the actual content, since data files often require domain-specific structural and value constraints, which requires explicit schema definitions. For this, we will integrate Frictionless for checking data content against community standards (e.g. MIAPPE, as demonstrated in the HORIZON project AGENT). Crucially, this project will also explore mechanisms for specifying expected data structures’ requirements within the ARC RO-Crate itself. This aims to provide a more self-contained description of data, investigating how such internal requirements can be linked to data validation frameworks, complementing the crate’s metadata validation.The overall goal is to provide a powerful, holistic validation mechanism for ARC RO-Crates, enhancing their reliability, trustworthiness, and FAIRness. A MIAPPE-compliant plant phenomics dataset will serve as a use case. This integrated validation approach aims to streamline quality control for researchers and will be packaged as a deployable microservice, offering broad applicability across diverse research workflows.
The primary goal of this project is to develop a web application for creating data usage agreements (DUA) in a way that allows automated evaluation of access permissions. Specifically, we want to adhere to the Open Digital Rights Language (ODRL) standard [1] and model permissions and prohibitions for the use of digital objects. ODRL is a policy expression language developed and adopted by the W3C. It provides a flexible and interoperable data model and vocabulary to enable fine-grained statements about the use of digital content and services. Recently, the Data Governance Act (DGA) was published as an implementation of the EU Data Act and defined roles for data intermediaries such as data trustees with certain prohibitions and obligations. The high expectations placed on the data trustee require that they have technical measures in place to facilitate the negotiation and enforcement of data use agreements. The “Ethical, Legal & Social Aspects” section of the NFDI (ELSA), has also issued a statement to the DGA [2], demonstrating the importance of this issue.Usually, DUAs are negotiated individually between parties and are not stored in a machine-readable format, which prevents automated modeling and verification of access rights for digital objects. Our web application will allow to create a DUA step by step via a configurable graphical user interface using ODRL data model in the background. This enables legal laymen to create data use agreements without much effort. The use of ODRL allows to programmatically query the data use agreements and to answer access requests automatically. To this end, the resulting DUAs can be persisted and queried through an API e.g. according to the GAIA-X specifications of Eclipse Dataspace Components (EDC) [3, 4] to exchange data compliant to rules and policies. Additionally, we want to address the integration of ODRLs in FDOs, such as ARC-RO-Crate of the DataPlant consortia and discuss extensions of the RO-Crate profiles. At last, for legal review and formal signing, negotiated DUAs can be rendered as PDFs.In summary, we will simplify the process of creating DUAs by adhering to international standards and will contribute to efforts to harmonize technical solutions, as the EOSC describes ODRL as a core metadata schema for legal interoperability [5]. DUAs can serve as a platform to gain the trust of data owners with protected, sensitive data and thus enable access to such resources. The project is in line with ELIXIR-DE/de.NBI's objective to improve the accessibility of resources and to ensure efficient, interoperable and secure resource sharing. It also aligns with the goals of the NFDI by handling sensitive data and enabling data protection. Especially when dealing with data from the health sector, but also handling agronomic data like land survey data or data from breeding programs.The project is a joint activity of the Leibniz IPK in Gatersleben (ELIXIR-DE/de.NBI Service Center GCBN), and the Justus-Liebig-University Giessen (ELIXIR-DE/de.NBI Service Center BiGi) and contributes to NFDI4Biodiversity, FAIRAgro, NFDI4Microbiota, DataPlant and FAIR-DS as well as to European initiatives such as EOSC and Gaia-X.
The Biodiversity, Food Security and Pathogens (BFSP) priority area is developing nicely as we are approaching the mid-point of the 2024-2026 work programme. Four Work Packages, equivalent to the open call projects, selected during 2024, were added and started in January 2025. During this Mini Symposium these four projects will be presented, along with a series of Node presentations showcasing their national BFSP-related activities. The second part of the session will be used to present the BFSP strategy and speak about the planned second BFSP open call. Focus areas will be presented and discussed. There will be time to think of future project submissions, with a specific emphasis on cross-Community/entity activities working towards advancing ELIXIR within the space of BFSP, as described in the Strategy.
Lack of interoperable datasets in plant breeding research creates an innovation bottleneck, requiring additional effort to integrate diverse datasets-if access is possible at all. Handling of plant breeding data and metadata must, therefore, change toward adopting practices that promote openness, collaboration, standardization, ethical data sharing, sustainability, and transparency of provenance and methodology. FAIR Digital Objects, which build on research data infrastructures and FAIR principles, offer a path to address this interoperability crisis, yet their adoption remains in its infancy. In the present work, we identify data sharing practices in the plant breeding domain as Data Cohorts and establish their connection to FAIR Digital Objects. We further link these cohorts to broader research infrastructures and propose a Data Trustee model for federated data sharing. With this we aim to push the boundaries of data management, often viewed as the last step in plant breeding research, to an ongoing process to enable future innovations in the field.
As part of the BioHackathon Europe 2024, we here report on the progress that both project 19 and project 24 have made during the event. For the purpose of this report we will present the abstract of both projects and then dive deeper on what work was done during the BioHackathon.
ELIXIR Plant Sciences Community is an interdisciplinary group of researchers, including computer scientists and plant biologists, with a common goal of developing infrastructure supporting the integration of different plant-related data types, ranging from omics, such as phenomics, genomics, transcriptomics, and metabolomics, to networks and systems biology models. Community activities encompass tool development, data standards development and FAIRification, establishing connections between data repositories and analysis environments, promoting of existing workflows and best practices, and interoperability of resources. Within the recently completed ELIXIR implementation study INCREASING (2021-2023) the community prepared three service bundles. Each bundle is a systematic organisation of tools and resources for dealing with various plant-derived data. Bundles are designed for better accessibility to the end user and are equipped with detailed instructions for use. The three bundles are: Plant Orthology Bundle for comparative genomics approaches, including tools such as Ensembl Plants, Plaza, Mercator, MapMan, and GoMapMan, Plant Data Analysis & Visualisation Bundle intended for researchers that work with complex datasets generated at multiple molecular levels. It includes a federated search engine (FAIDARE) across three knowledge graph resources: KnetMiner, AgroLD, and Stress Knowledge Map (SKM), each able to contextualise data within prior knowledge, and data analysis and visualisation tools such as DiNAR, MapMan, and SKM-tools, Phenotyping Data Publication Bundle, based on the guidelines published in RDMkit and the FAIR Cookbook, defines the best practices for data standardisation, data submission and publication, and data findability Here we present current approaches in plant sciences and highlight the applications of systems biology.
The Leibniz Institute of Plant Genetics and Crop Plant Research (IPK) Gatersleben is a leading international plant science institute specializing in biodiversity and crop plant performance research. Over the last decade, all phases of the research data lifecycle were implemented as a continuous process in conjunction with information technology, standardization, and sustainable research data management (RDM) processes. Under the leadership of a team of data stewards, a research data infrastructure, process landscape, capacity building, and governance structures were successfully established. As a result, a generic research data infrastructure was created to serve the principles of good scientific practice, archiving research data in an accessible and sustainable manner, even before the FAIR criteria were formulated. In this paper, we discuss success stories as well as pitfalls and summarize the experiences from 15 years of operating a central RDM infrastructure. We present measures for agile requirements engineering, technical and organizational implementation, governance, training, and roll-out. We show the benefits of a participatory approach across all departments, personnel roles, and researcher profiles through pilot working groups and data management champions. As a result, an ambidextrous approach to data management was implemented, referring to the ability to efficiently combine operational needs, support daily tasks in compliance with the FAIR criteria, while remaining open to adopting technical innovations in an agile manner.
As part of the de.NBI BioHackathon 2023, we here report about our progress on increasing FAIR-compliance in agrosystem sciences and plant phenomics. Through the collaborative efforts of the agrosystem and plant sciences communities, research data are available through various data repositories and infrastructures. To foster these developments and increase the value for the communities, enabling FAIR-compliance for scientific datasets is one top priority strategic aim. Due to the heterogeneity of the sub-domains and their requirements, we addressed three challenges with direct relation to specific FAIR principles: Increasing findability of digital agrosystem resources by extending Schema.org, Enabling easy creation of MIAPPE-compliant ISA metadata for Plant Phenotyping Experiments, and Increasing Plant Data Accessibility and Collaboration with FAIDARE.
Agriculture is confronted with several challenges such as climate change, the loss of biodiversity and stagnating productivity. The massive increasing amount of data and new digital technologies promise to overcome them, but they necessitate careful data integration and data management to make them usable. The FAIRagro consortium is part of the National Research Data Infrastructure (NFDI) in Germany and will develop FAIR compliant infrastructure services for the agrosystems science community, which will be integrated in the existing research data infrastructure service landscape. Here we present the initial steps of designing and implementing the FAIRagro middleware infrastructure to connect existing data infrastructures. The middleware will feature services for the seamless data integration across diverse infrastructures. Data and metadata are streamlined for research in agrosystems science by downstream processing in the central FAIRagro Search and Inventory Portal and the data integration and analysis workflow system "SciWIn".
A prevailing paradigm in Research Data Management (RDM) is to publish research datasets in designated archives upon conclusion of a research process. However, it is beneficial to abandon the notion of final or static data artifacts and instead adopt a continuous approach towards working with research data, where data is constantly shared, versioned, and updated. This immutable yet evolving perspective allows for the application of existing technologies and processes from software engineering, such as continuous integration, release practices, and version management backed by decades of experience, and adaptable to RDM.To facilitate this, we propose the Annotated Research Context (ARC), a data and metadata layout convention based on the well-established ISA model for metadata annotation and implemented using Git repositories. ARCs are amenable towards frequent, lightweight data management operations, such as (meta)data validation and transformation. The Omnipy Python library is designed to help develop stepwise validated (meta)data transformations as scalable data flows that can be incrementally designed, updated, and rerun as requirements or data evolve.To demonstrate the concept of continuous RDM we will use Omnipy to define and orchestrate Git-backed CI/CD (Continuous Integration/Continuous Delivery) data flows to convert ISA metadata present in ARCs into validated RO-Crate representations adhering to the Bioschemas convention. A RO-Crate package combines the actual research data with its metadata description. Downstream, this allows semantic interpretation by Galaxy for e.g. workflow execution as well as machine-readable data access and data harvesting for search engines such as FAIDARE.
As part of the BioHackathon Europe 2023, we here report on the progress of the hacking team preparing a resource index and knowledge graph based on the JSON-LD Bioschemas markup from several resources in the life- and natural sciences, predominantly from the fields of plant- and (bio)chemistry research. This preliminary analysis will allow us to better understand how Bioschemas markup is currently used in these two communities, so we can take actions to improve guidelines and validation on the Bioschemas markup and the data providers side. The lessons learnt will be useful for other communities as well. The ultimate goal is facilitating and improving interoperability across resources.
As part of the BioHackathon Germany 2022, we hereby report on the success of the two projects “MIAPPE Wizard: Enabling easy creation of MIAPPE-compliant ISA metadata for Plant Phenotyping Experiments” and “DataPLANT - Facilitating Research Data Management to combat the reproducibility crisis”. Shortly before the actual hackathon, it became apparent to the participants that close coordination between the projects would be very beneficial. Both projects aimed to improve the process of collecting and aggregating metadata on plant experiments, but with different approaches.