Face aux défis de la transition agroalimentaire et aux exigences de durabilité, il est essentiel de combiner des données portant sur de multiples dimensions. Aujourd’hui, ces données sont souvent dispersées et difficiles à interconnecter, ce qui limite leur exploitation et leur réutilisation. La gestion et l’ouverture des données scientifiques deviennent essentielles afin de structurer et partager efficacement les connaissances. La construction de l’ontologie PO²/TransformON répond à ce besoin en proposant une approche fondée sur les principes FAIR (Facile à trouver, Accessible, Interopérable et Réutilisable) et le Web sémantique. L’ontologie vise à organiser et modéliser les données liées aux procédés de transformation et à la caractérisation des produits alimentaires et non alimentaires. L’objectif est de faciliter leur partage et leur exploitation par la communauté scientifique. Dans cette optique, le processus de FAIRification s’appuie sur le développement des solutions logicielles PO² Manager et SPO²Q. Ces solutions permettent de structurer les informations, dès leur acquisition, pour les représenter sous forme de graphes et de les requêter à l’aide des standards du Web sémantique RDF, OWL et SPARQL. L’intégration de ces données qui répondent aux principes FAIR dans des entrepôts ouverts, comme Recherche Data Gouv, favorise leur réutilisation et leur valorisation. À terme, cette démarche ouvre de nouvelles perspectives, notamment en intelligence artificielle et en modélisation prédictive pour l’aide à la décision. L’enrichissement des ontologies et l’interopérabilité accrue des systèmes d’information faciliteront la gestion durable des ressources et l’innovation dans les domaines agroalimentaire et environnemental.
Background The Wheat Crop ontology was created to annotate phenotypic experimental data (i.e. field and greenhouse measurements standardized and integrated in databases). The Wheat Trait and Phenotype ontology was created to annotate information on wheat traits from the literature (i.e. text found in the abstract, results and discussion of scholarly articles). To enable seamless data retrieval on wheat traits from these complementary sources, the classes in the two ontologies have been aligned. Methods All pairs of ontology classes were examined and categorized in nine groups based on the nature of their relationships (e.g. equivalence, subsumption). General principles emerged from this process which were formalized into rules. The Simple Standard for Sharing Ontological Mappings (SSSOM) representation was chosen to represent the mappings in RDF (Resource Description Framework), including their metadata such as creators, reviewers, and justification (including rules). Results The mapping dataset is publicly available. It covers 77% of the ontology classes. Most labels of the aligned classes differed significantly and required domain expertise for decisions, especially for traits related to biotic stress. Consequently, most mappings are close mappings rather than exact equivalents. Conclusions We present the end-to-end manual process used to select and represent mappings in SSSOM within the specific domain of wheat traits. We derive general lessons from the complex alignment process that extend beyond the specific case of these two ontologies and more generally apply to alignments of specialized ontologies for information retrieval purposes. This work demonstrates the relevance of SSSOM for representing these mappings.
There is a growing interest in milk oligosaccharides (MOs) because of their numerous benefits for newborns’ and long-term health. A large number of MO structures have been identified in mammalian milk. Mostly described in human milk, the oligosaccharide richness, although less broad, has also been reported for a wide range of mammalian species. The structure of MOs is particularly difficult to report as it results from the combination of 5 monosaccharides linked by various glycosidic bonds forming structurally diverse and complex matrices of linear and branched oligosaccharides. Exploring the literature and extracting relevant information on MO diversity within or across species appears promising to elucidate structure-function role of MOs. Currently, given the complexity of these molecules, the main issues in exploring literature to extract relevant information on MO diversity within or across species relate to the heterogeneity in the way authors refer to these molecules. Herein, we provide a thesaurus (MilkOligoThesaurus) including the names and synonyms of MOs collected from key selected articles on mammalian milk analyses. MilkOligoThesaurus gathers the names of the MOs with a complete description of their monosaccharide composition and structures. When available, each unique MO molecule is linked to its ID from the NCBI PubChem and ChEBI databases. MilkOligoThesaurus is provided in a tabular format. It gathers 245 unique oligosaccharide structures described by 22 features (columns) including the name of the molecule, its abbreviation, the chemical database IDs if available, the monosaccharide composition, chemical information (molecular formula, monoisotopic mass), synonyms, its formula in condensed form, and in abbreviated condensed form, the abbreviated systematic name, the systematic name, the isomer group, and scientific article sources. MilkOligoThesaurus is also provided in the SKOS (Simple Knowledge Organization System) format. This thesaurus is a valuable resource gathering MO naming variations that are not found elsewhere for (i) Text and Data Mining to enable automatic annotation and rapid extraction of milk oligosaccharide data from scientific papers; (ii) biology researchers aiming to search for or decipher the structure of milk oligosaccharides based on any of their names, abbreviations or monosaccharide compositions and linkages.
This article describes our study on the alignment of two complementary knowledge graphs useful in agriculture: the thesaurus of cultivated plants in France named French Crop Usage (FCU) and the French national taxonomic repository TAXREF for fauna, flora, and fungi. FCU describes the usages of plants in agriculture: “tomatoes” are crops used for human food, and “grapevines” are crops used for human beverage. TAXREF describes biological taxa and associated scientific names: for example, a tomato species may be “Solanum lycopersicum” or a grapevine species may be “Vitis vinifera”. Both knowledge graphs contain vernacular names of plants but those names are ambiguous. Thus, a group of agricultural experts produced some mappings from FCU crops to TAXREF taxa. Moreover, new RDF properties have been defined to declare those new types of mapping relations between plant descriptions. The metadata for the mappings and the mapping set are encoded with the Simple Standard for Sharing Ontological Mappings (SSSOM), a new model which, among other qualities, offers means to report on provenance of particular interest for this study. The produced mappings are available for download in Recherche Data Gouv, the federated national platform for research data in France.
In this article, we present a joint effort of the wheat research community, along with data and ontology experts, to develop wheat data interoperability guidelines. Interoperability is the ability of two or more systems and devices to cooperate and exchange data
Nowadays, it is important to make the results of scientific research accessible in a simple and understandable way according to the Open Science policy. This movement uses tools to enhance findability and interoperability of data. This paper describes the transformation of the meat dictionary published by the French Meat Academy as a book into a machine actionable and freely accessible terminological resource based on the SKOS standard format. This thesaurus contains 1567 concepts describing the meat production chain. This work was carried out by experts in semantic web, meat biology and meat vocabulary. This thesaurus can be used to index articles, journals and datasets, thus facilitating consultation; it can also be used to facilitate interoperability of the indexed datasets and provide contextual definitions for building ontologies, i.e. formal descriptions of knowledge for reasoning on data. The thesaurus can be useful to enrich other vocabularies with new knowledge, such as French specificities in terms of meat cuts or definitions.
Making data compliant with the FAIR Data principles (Findable, Accessible, Interoperable, Reusable) is still a challenge for many researchers, who are not sure which criteria should be met first and how. Illustrated from experimental data tables associated with a Design of Experiments, we propose an approach that can serve as a model for a research data management that allows researchers to disseminate their data by satisfying the main FAIR criteria without insurmountable efforts. More importantly, this approach aims to facilitate the FAIRification process by providing researchers with tools to improve their data management practices.
This study aims to examine, through stakeholders consultation, the widely used definitions of four terms related to plastics sustainability: ‘bio-based plastics', ‘bioplastics’, ‘biodegradable plastics’ and ‘plastics recycling’ and to mitigate their potential ambiguity for diverse scientific communities and sectors of activity. For the three terms ‘bio-based plastics', ‘biodegradable’ and ‘recycling’, consolidated definitions were elaborated based on the feedback of online survey and analysis of the pro and con arguments given by face-to-face interviews with 18 experts followed by an online survey of 122 stakeholders. Acceptance of the consolidated definitions was higher than the official ones with an increase of acceptance from 43% to 81% for bio-based plastics, from 47% to 61% for biodegradable plastics, and from 28% to 60% for plastics recycling. The terms ‘biodegradable’ and ‘recycling’ remain ambiguous even after consolidation of the definition. This highlights that more discussions are necessary to achieve a consensual and fair definition of such complex properties and mechanisms. In the term ‘Bioplastics’ the prefix ‘bio’, referring either to the origin of the resources or the end of life of the material, remains difficult to understand and we prefer to advise against its use, especially with non-expert people (e.g. consumers and the public at large), in favour of the use of ‘bio-based plastics’ or ‘biodegradable plastics’. The issue of this study is to help consolidate wider efforts to develop new strategies for replacing oil-based plastics and improving end-of-life options.
In this article, we present a joint effort of the wheat research community, along with data and ontology experts, to develop wheat data interoperability guidelines. Interoperability is the ability of two or more systems and devices to cooperate and exchange data, and interpret that shared information. Interoperability is a growing concern to the wheat scientific community, and agriculture in general, as the need to interpret the deluge of data obtained through high-throughput technologies grows. Agreeing on common data formats,
De TermSciences a Loterre : comment l'Inist-CNRS a rendu les terminologies ouvertes plus conformes aux principes FAIR Cet article a ete publie le 1 er avril 2021 sur le site du projet FooSIN : https://foosin.fr/determsciences-a-loterre-comment-linist-cnrs-a-rendu-les-terminologies-ouvertes-plusconformes-aux-principes-fair/ Les membres du projet FooSIN ont interviewe l'Inist-CNRS a propos du passage de TermSciences a Loterre, deux portails publics exposant des terminologies d'interet notamment pour les acteurs de l'agriculture, de l'alimentation et de l'environnement. En effet, une dizaine de ces terminologies ont ete publiees par INRAE. Les objectifs de cet entretien est de montrer en quoi Loterre a progresse par rapport a TermSciences et quel est l'impact de ce changement sur les ressource hebergees en termes de satisfaction des principes FAIR. Suite aux echanges et en s'appuyant sur l'utilisation de la grille d'evaluation SHARC, le projet FooSIN suggere quelques pistes pour aller plus loin dans la demarche. Enfin, vous trouverez dans cet article des pointeurs vers les outils, standards et ressources utilisees pour la FAIRification.
Making data compliant with the FAIR Data principles (Findable, Accessible, Interoperable, Reusable) is still a challenge for many researchers, who are not sure which criteria should be met first and how. Illustrated with experimental data tables associated with a Design of Experiments, we propose an approach that can serve as a model for research data management that allows researchers to disseminate their data by satisfying the main FAIR criteria without insurmountable efforts. More importantly, this approach aims to facilitate the FAIR compliance process by providing researchers with tools to improve their data management practices.
En l’absence de thésaurus spécialisé dans le domaine de l’agroécologie, un groupe métier aux compétences complémentaires (experts scientifiques et spécialistes de l’information scientifique et technique) a construit un « thésaurus d’agroécologie ». Ce thésaurus est issu de la valorisation de l’ensemble des termes capitalisés par le dispositif de veille territoriale Agroécologie conduit à l’échelle de la région Midi-Pyrénées sur la période 2013–2017. L’ensemble des données constitutives de ce thésaurus est accessible sous Licence Ouverte et dans un format standard. Exposé sur différents portails de vocabulaires généralistes ou thématiques, ce thésaurus pourra être réutilisé à d’autres fins. Cet article décrit la méthodologie de constitution, le contenu et la structuration de ce thésaurus ainsi que son potentiel de réutilisation. L’originalité de ce travail est de reposer sur une expertise scientifique conduite à partir des termes d’usage des opérateurs du monde agricole et agro-alimentaire. La méthode proposée peut être réemployée pour construire des thésaurus sur d’autres domaines émergents.