
The demand for food is expected to grow substantially in the coming years. To address this challenge, especially in the context of climate change, a deeper understanding of genotype-phenotype relationships is crucial for improving crop yields. Recent advances in high-throughput technologies have transformed the landscape of plant science research. However, there is an urgent need to integrate and consolidate complementary data to understand the biological system. We introduce AgroLD, a knowledge graph that uses Semantic Web technologies to seamlessly integrate plant science data. AgroLD is designed to facilitate hypothesis formulation and validation within the scientific community. With approximately 1.08 billion triples, it integrates and annotates data from more than 151 datasets across 19 distinct sources. The overarching goal is to provide a specialized knowledge platform addressing complex biological questions in the plant sciences, including gene participation in plant disease resistance and adaptive responses to climate change.
The Swiss Personalized Health Network has developed a national framework for enabling the semantic representation of health data within a Knowledge Graph. This framework has been implemented in all Swiss university hospitals, promoting seamless sharing and integration of clinical routine data with other health-related data, including omics and clinical research data. While research projects often have flexibility in selecting terminologies and specific versions, historical clinical routine data are typically coded for billing or administrative purposes using predefined terminologies in various (sometimes even unknown) versions over time. Some of these terminologies do not adhere to best practices for ontology design, presenting significant challenges to the retrospective re-use of such coded data for research. Common issues with these terminologies include the lack of machine-readable traceability across versions and non-adherence to FAIR principles. Terms from older versions frequently disappear in newer ones, making it challenging to distinguish outdated from invalid terms. Additionally, 'semantic drift' occurs, where the meaning of terms changes across versions. To address these challenges, we have implemented FAIR and historized versions of ATC, CHOP, and ICD-10-GM. We represent each version in RDF using versioned URIs and track meaning changes between versions in a machine-readable way using OWL and RDFS. The integration of these historized terminologies into our quality control framework, based on SHACLs, enables comprehensive data quality control in hospitals and empowers researchers to effectively utilize this data. Our work aims to bridge the gap between health data coded in different terminology versions, ensuring a consistent and reliable semantic representation.
The Swiss Personalized Health Network (SPHN) is a Swiss research infrastructure initiative that aims to facilitate the exchange of health-related data in a FAIR manner. The SPHN Dataset and SPHN RDF Schema form an essential part of the SPHN Semantic Interoperability Framework, which currently covers mostly clinical routine data. To facilitate the integration of omics data produced by the SPHN National Data Streams, a genomics extension was developed. This was done in close collaboration with clinicians, researchers, bioinformaticians, and data managers, from Swiss university hospitals, academic research groups and the omics platforms. Here, we present the genomics extension of the SPHN RDF Schema, which can be used to semantically describe genomics experiments and covers both clinical and research domains. The schema centers around the general omics process flow, with concepts that denote the individual steps, such as sample processing, assay, and data processing. Genomics-specific specializations are provided, such as library preparation, sequencing assay, and sequencing analysis. The schema also facilitates in capturing other important omics metadata, such as information about the sequencing instrument, standard operating procedure, and quality control metrics. The extension aligns with existing semantic data models and reuses common biomedical vocabularies, such as EDAM, OBI and FAIR genomes, as value sets, thereby facilitating semantic interoperability. It will be used to FAIRify data that is produced within the Swiss network and to facilitate sharing this data as one knowledge graph for reuse among its participants.
Plants have a complex chemo-diversity and represent a reservoir of potential new therapeutic agents. Within a Swiss research project, six scientific research groups from different disciplines are collaborating to investigate a collection of more than 17’000 unique dried plant extracts. It aims to find new bioactive molecules and their modes of action, with for example anti-infective or pro-metabolic activities. One of the main challenges of this enterprise is the management, integration and sharing of the highly heterogeneous data that are produced by the different research groups. Among these we find (i) massive high-resolution mass spectrometry data, (ii) the numerical results of innovative chemo-informatics methods, (iii) bioassay results from experimental models of tuberculosis and obesity, and (iv) organic synthetic chemistry. Additionally, requirements for data management plan and open-source science with the FAIR principles must be met. We have established an agile pipeline to capture and structure this heterogeneous data into an RDF graph. The data content's gradual expansion and evolution throughout the project presented considerable challenges, particularly in terms of data modeling. Additionally, despite many collaborators not being RDF experts, most were technically adept at producing RDF triples relevant to their contributions. We have deployed multiple instances of a triplestore and developed an in-house custom tool (i.e. KGSteward) to synchronize their content, based on a configuration file, which is centrally managed and version-controlled using Git. This strategy gave us the flexibility required to address global project challenges in common data management effectively.