Heterogeneous and multidisciplinary data generated by research on sustainable global agriculture and agrifood systems requires quality data labeling or annotation in order to be interoperable. As recommended by the FAIR principles, data, labels, and metadata must use controlled vocabularies and ontologies that are popular in the knowledge domain and commonly used by the community. Despite the existence of robust ontologies in the Life Sciences, there is currently no comprehensive full set of ontologies recommended for data annotation across agricultural research disciplines. In this paper, we discuss the added value of the Ontologies Community of Practice (CoP) of the CGIAR Platform for Big Data in Agriculture for harnessing relevant expertise in ontology development and identifying innovative solutions that support quality data annotation. The Ontologies CoP stimulates knowledge sharing among stakeholders, such as researchers, data managers, domain experts, experts in ontology design, and platform development teams.
One prerequisite for the success of community‐based WG is to have at least a sponsor and ideally a sponsor that has legitimacy with regards the community. WDI was supported by the Wheat Initiative (https://www.wheatinitiative.org) and several well recognized and key players of domain of the wheat research.
Heterogeneous and multidisciplinary data generated by research on sustainable global agriculture and agrifood systems requires quality data labelling to be interoperable. As recommended by the FAIR principles, data, labels and metadata must use controlled vocabularies and ontologies that are popular in the knowledge domain and commonly used by the community. Despite the existence of robust ontologies in the Life Sciences, there is currently no agreed full set of ontologies recommended for data annotation across agricultural research disciplines, which may span genetics, environment, agroecology, biology and socioeconomics. In this paper, we discuss the added value of the Ontologies Community of Practice (CoP) of the CGIAR Platform for Big Data in Agriculture for harnessing relevant ontology expertise. This CoP aims to stimulate knowledge sharing and directly support platform development teams by producing ontologies or contributing missing concepts, recommending best practices and identifying mitigation solutions when gold standard datasets are difficult to attain.
Agronomy/agriculture and biodiversity (ag & biodiv) communities face several major societal, economic, and environmental challenges that data science approaches will help address. To achieve their goals, researchers of these communities must be able to rapidly discover, aggregate, integrate, and analyse different types of data and information sources. Semantic technologies, combined to open, FAIR data and services, is one of the answers to fully knowledge-driven, and transparent science and innovation. The D2KAB project (www.d2kab.org) aims to create a framework to turn agronomy and biodiversity data into knowledge – semantically described, interoperable, actionable, open – and investigate the scientific methods and tools to exploit this knowledge for applications in agriculture and biodiversity sciences. This project, funded by French ANR (2019-2023), will provide the means –ontologies and linked open data– for ag & biodiv to embrace semantic Web technologies in order to produce and exploit FAIR data and services. To do so, D2KAB will develop new original methods and algorithms in the following areas: data integration, text mining, semantic annotation, ontology alignment and linked data exploitation and visualization. D2KAB project brings together a unique multidisciplinary consortium of 12 partners to achieve this objective: 2 informatics research units (LIRMM, I3S); 6 INRA/IRSTEA/IRD research units at the interface of computer science and ag & biodiv (URGI, MaIAGE, IATE, DIST, TSCF, DIADE) specialized in agronomy or agriculture; 2 labs in biodiversity and ecosystem research (CEFE, URFM); 1 association of agriculture stakeholders (ACTA); and 1 partnership with Stanford BMIR department. Three main goals drive D2KAB’s roadmap: 1. To develop state-of-the-art methods and technologies for ontology lifecycle and alignment. 2. To build the agronomy, agriculture and biodiversity Linked Open Data cloud. 3. To enable new semantically driven agronomy and biodiversity science. The work is starting from the recommendations of several RDA WG and IG already published or in progress (e.g. Agrisemantic WG, Vocabulary Services IG, Wheat and Rice Data Interoperability WGs, Agricultural Data IG, SHARC IG). Some of the key technological building blocks of D2KAB are AgroPortal, a reference repository for ontologies and vocabularies in agronomy; AgroLD, a semantic Web knowledge base that integrates agronomic data from public databases including GO associations, Gramene, UniprotKB, and OryGenesDB ; Corese, a semantic Web factory that implements the W3C standards RDF, RDFS, OWL-RL and SPARQL, and LDScript, a Linked Data Script Language, and STTL, the SPARQL Template Transformation Language for RDF; and Alvis, a text mining for semantic normalisation of free text by ontologies. D2KAB will allow the valorization of ag & biodiv data into real world applications leading to economic impact, smart agriculture and ecological preservation. Five driving scenarios are planned: development of an ontology-based expert system to select food packaging solutions; creation of an augmented semantic reader for Plant Health Bulletins; advanced integration of textual and experimental data on wheat phenotypes; development of new ontologies on plant root traits and extension of the Thesaurus Of Plant Characteristics; integration of plant functional biogeography data related to the Mediterranean Basin. Each of the project scenarios will have a significant impact and produce concrete outcomes for ag & biodiv scientific communities and socio-economic stakeholders in agriculture.
In this article, we present a joint effort of the wheat research community, along with data and ontology experts, to develop wheat data interoperability guidelines. Interoperability is the ability of two or more systems and devices to cooperate and exchange data, and interpret that shared information. Interoperability is a growing concern to the wheat scientific community, and agriculture in general, as the need to interpret the deluge of data obtained through high-throughput technologies grows. Agreeing on common data formats, metadata, and vocabulary standards is an important step to obtain the required data interoperability level in order to add value by encouraging data sharing, and subsequently facilitate the extraction of new information from existing and new datasets. During a period of more than 18 months, the RDA Wheat Data Interoperability Working Group (WDI-WG) surveyed the wheat research community about the use of data standards, then discussed and selected a set of recommendations based on consensual criteria. The recommendations promote standards for data types identified by the wheat research community as the most important for the coming years: nucleotide sequence variants, genome annotations, phenotypes, germplasm data, gene expression experiments, and physical maps. For each of these data types, the guidelines recommend best practices in terms of use of data formats, metadata standards and ontologies. In addition to the best practices, the guidelines provide examples of tools and implementations that are likely to facilitate the adoption of the recommendations. To maximize the adoption of the recommendations, the WDI-WG used a community-driven approach that involved the wheat research community from the start, took into account their needs and practices, and provided them with a framework to keep the recommendations up to date. We also report this approach’s potential to be generalizable to other (agricultural) domains.
Many vocabularies and ontologies are produced to represent and annotate agronomic data. However, those ontologies are spread out, in different formats, of different size, with different structures and from overlapping domains. Therefore, there is need for a common platform to receive and host them, align them, and enabling their use in agro-informatics applications. By reusing the National Center for Biomedical Ontologies (NCBO) BioPortal technology, we have designed AgroPortal, an ontology repository for the agronomy domain. The AgroPortal project re-uses the biomedical domain’s semantic tools and insights to serve agronomy, but also food, plant, and biodiversity sciences. We offer a portal that features ontology hosting, search, versioning, visualization, comment, and recommendation; enables semantic annotation; stores and exploits ontology alignments; and enables interoperation with the semantic web. The AgroPortal specifically satisfies requirements of the agronomy community in terms of ontology formats (e.g., SKOS vocabularies and trait dictionaries) and supported features (offering detailed metadata and advanced annotation capabilities). In this paper, we present our platform’s content and features, including the additions to the original technology, as well as preliminary outputs of five driving agronomic use cases that participated in the design and orientation of the project to anchor it in the community. By building on the experience and existing technology acquired from the biomedical domain, we can present in AgroPortal a robust and feature-rich repository of great value for the agronomic domain.
The wheat data interoperability guidelines ( http://ist.blogs.inra.fr/wdi/ ) result from a joint effort of the wheat research community and data experts, who wish to make their research data more accessible, interoperable and reusable. During more than 18 months, the Wheat Data Interoperability working group questioned the wheat research community about their usage of data standards, discussed and selected a set of recommendations based on some consensual criterias. This presentation will give you some feedback about the work of the WDI working group, update you with the latest news and discuss possible future developments.
Discussion subjects: How to promote the guidelines? Training the users to increase adoption Which deliverables could we share with the other crop specific groups? The future of the WDI Action plan