Ontologies are widely used in databases to standardize data, improving data quality, integration, and ease of comparison. Within ontologies tailored to diverse use cases, post-composing user-defined terms reconciles the demands for standardization on the one hand and flexibility on the other. In many instances of Breedbase, a digital ecosystem for plant breeding designed for genomic selection, the goal is to capture phenotypic data using highly curated and rigorous crop ontologies, while adapting to the specific requirements of plant breeders to record data quickly and efficiently. For example, post-composing enables users to tailor ontology terms to suit specific and granular use cases such as repeated measurements on different plant parts and special sample preparation techniques. To achieve this, we have implemented a post-composing tool based on orthogonal ontologies providing users with the ability to introduce additional levels of phenotyping granularity tailored to unique experimental designs. Post-composed terms are designed to be reused by all breeding programs within a Breedbase instance but are not exported to the crop reference ontologies. Breedbase users can post-compose terms across various categories, such as plant anatomy, treatments, temporal events, and breeding cycles, and, as a result, generate highly specific terms for more accurate phenotyping.
Genome-wide association studies (GWAS) are widely used to infer the genetic basis of traits in organisms; however, selecting appropriate thresholds for analysis remains a significant challenge. In this study, we introduce the Sequential SNP Prioritization Algorithm (SSPA) to investigate the genetic underpinnings of two key phenotypes in Sorghum bicolor: maximum canopy height and maximum growth rate. Using a subset of the Sorghum Bioenergy Association Panel cultivated at the Maricopa Agricultural Center in Arizona, we performed GWAS with specific permissive-filtered thresholds to identify genetic markers associated with these traits, enabling the identification of a broader range of explanatory candidate genes. Building on this, our proposed method employed a feature engineering approach leveraging statistical correlation coefficients to unravel patterns between phenotypic similarity and genetic proximity across 274 accessions. This approach helps prioritize Single Nucleotide Polymorphisms (SNPs) that are likely to be associated with the studied phenotype. Additionally, we conducted a complementary analysis to evaluate the impact of SSPA by including all variants (SNPs) as inputs, without applying GWAS. Empirical evidence, including ontology-based gene function, spatial and temporal expression, and similarity to known homologs demonstrates that SSPA effectively prioritizes SNPs and genes influencing the phenotype of interest, providing valuable insights for functional genetics research.
The Planteome project (https://planteome.org/) provides a suite of reference and crop-specific ontologies and an integrated knowledgebase of plant genomics data. The plant genomics data in the Planteome has been obtained through manual and automated curation and sourced from more than 40 partner databases and resources. Here, we report on updates to the Planteome reference ontologies, namely, the Plant Ontology (PO), Trait Ontology (TO), the Plant Experimental Conditions Ontology (PECO), and integration of species/crop-specific vocabularies from our partners, the Crop Ontology (CO) into the TO ontology graph. Currently, 11 CO vocabularies are integrated into the Planteome with the addition of yam, sorghum, and potato since 2018. In addition, the size of the annotation database has increased by 34%, and the number of bioentities (genes, proteins, etc.) from 125 plant taxa has increased by 72%. We developed new tools to facilitate user requests and improvements to the CO vocabularies, and to allow fast searching and browsing of PO terms and definitions. These enhancements and future changes to automate the TO-CO mappings and knowledge discovery tools ensure that the Planteome will continue to be a valuable resource for plant biology.
Global climate change and associated environmental extremes present a pressing need to understand and predict social–environmental impacts while identifying opportunities for mitigation and adaptation. In support of informing a more resilient future, emerging data analytics technologies can leverage the growing availability of Earth observations from diverse data sources ranging from satellites to sensors to social media. Yet, there remains a need to transition from research for knowledge gain to sustained operational deployment. In this paper, we present a research-to-commercialization (R2C) model and conduct a case study using it to address the wicked wildfire problem through an industry–university partnership. We systematically evaluated 39 different user stories across eight user personas and identified information gaps in public perception and dynamic risk. We discuss utility and challenges in deploying such a model as well as the relevance of the findings from this use case. We find that research-to-commercialization is non-trivial and that academic–industry partnerships can facilitate this process provided there is a clear delineation of (i) intellectual property rights; (ii) technical deliverables that help overcome cultural differences in working styles and reward systems; and (iii) a method to both satisfy open science and protect proprietary information and strategy. The R2C model presented provides a basis for directing solutions-oriented science in support of value-added analytics that can inform a more resilient future.
IntroductionClimate change is already affecting ecosystems around the world and forcing us to adapt to meet societal needs. The speed with which climate change is progressing necessitates a massive scaling up of the number of species with understood genotype-environment-phenotype (G×E×P) dynamics in order to increase ecosystem and agriculture resilience. An important part of predicting phenotype is understanding the complex gene regulatory networks present in organisms. Previous work has demonstrated that knowledge about one species can be applied to another using ontologically-supported knowledge bases that exploit homologous structures and homologous genes. These types of structures that can apply knowledge about one species to another have the potential to enable the massive scaling up that is needed through in silico experimentation.MethodsWe developed one such structure, a knowledge graph (KG) using information from Planteome and the EMBL-EBI Expression Atlas that connects gene expression, molecular interactions, functions, and pathways to homology-based gene annotations. Our preliminary analysis uses data from gene expression studies in Arabidopsis thaliana and Populus trichocarpa plants exposed to drought conditions.ResultsA graph query identified 16 pairs of homologous genes in these two taxa, some of which show opposite patterns of gene expression in response to drought. As expected, analysis of the upstream cis-regulatory region of these genes revealed that homologs with similar expression behavior had conserved cis-regulatory regions and potential interaction with similar trans-elements, unlike homologs that changed their expression in opposite ways.DiscussionThis suggests that even though the homologous pairs share common ancestry and functional roles, predicting expression and phenotype through homology inference needs careful consideration of integrating cis and trans-regulatory components in the curated and inferred knowledge graph.
Over the last couple of decades, there has been a rapid growth in the number and scope of agricultural genetics, genomics and breeding databases and resources. The AgBioData Consortium (https://www.agbiodata.org/) currently represents 44 databases and resources (https://www.agbiodata.org/databases) covering model or crop plant and animal GGB data, ontologies, pathways, genetic variation and breeding platforms (referred to as 'databases' throughout). One of the goals of the Consortium is to facilitate FAIR (Findable, Accessible, Interoperable, and Reusable) data management and the integration of datasets which requires data sharing, along with structured vocabularies and/or ontologies. Two AgBioData working groups, focused on Data Sharing and Ontologies, respectively, conducted a Consortium-wide survey to assess the current status and future needs of the members in those areas. A total of 33 researchers responded to the survey, representing 37 databases. Results suggest that data-sharing practices by AgBioData databases are in a fairly healthy state, but it is not clear whether this is true for all metadata and data types across all databases; and that, ontology use has not substantially changed since a similar survey was conducted in 2017. Based on our evaluation of the survey results, we recommend (i) providing training for database personnel in a specific data-sharing techniques, as well as in ontology use; (ii) further study on what metadata is shared, and how well it is shared among databases; (iii) promoting an understanding of data sharing and ontologies in the stakeholder community; (iv) improving data sharing and ontologies for specific phenotypic data types and formats; and (v) lowering specific barriers to data sharing and ontology use, by identifying sustainability solutions, and the identification, promotion, or development of data standards. Combined, these improvements are likely to help AgBioData databases increase development efforts towards improved ontology use, and data sharing via programmatic means. Database URL https://www.agbiodata.org/databases.
The Gene Ontology (GO) knowledgebase (http://geneontology.org) is a comprehensive resource concerning the functions of genes and gene products (proteins and noncoding RNAs). GO annotations cover genes from organisms across the tree of life as well as viruses, though most gene function knowledge currently derives from experiments carried out in a relatively small number of model organisms. Here, we provide an updated overview of the GO knowledgebase, as well as the efforts of the broad, international consortium of scientists that develops, maintains, and updates the GO knowledgebase. The GO knowledgebase consists of three components: (1) the GO-a computational knowledge structure describing the functional characteristics of genes; (2) GO annotations-evidence-supported statements asserting that a specific gene product has a particular functional characteristic; and (3) GO Causal Activity Models (GO-CAMs)-mechanistic models of molecular "pathways" (GO biological processes) created by linking multiple GO annotations using defined relations. Each of these components is continually expanded, revised, and updated in response to newly published discoveries and receives extensive QA checks, reviews, and user feedback. For each of these components, we provide a description of the current contents, recent developments to keep the knowledgebase up to date with new discoveries, and guidance on how users can best make use of the data that we provide. We conclude with future directions for the project.
As one of the US Department of Agriculture—Agricultural Research Service flagship databases, GrainGenes (https://wheat.pw.usda.gov) serves the data and community needs of globally distributed small grains researchers for the genetic improvement of the Triticeae family and Avena species that include wheat, barley, rye and oat. GrainGenes accomplishes its mission by continually enriching its cross-linked data content following the findable, accessible, interoperable and reusable principles, enhancing and maintaining an intuitive web interface, creating tools to enable easy data access and establishing data connections within and between GrainGenes and other biological databases to facilitate knowledge discovery. GrainGenes operates within the biological database community, collaborates with curators and genome sequencing groups and contributes to the AgBioData Consortium and the International Wheat Initiative through the Wheat Information System (WheatIS). Interactive and linked content is paramount for successful biological databases and GrainGenes now has 2917 manually curated gene records, including 289 genes and 254 alleles from the Wheat Gene Catalogue (WGC). There are >4.8 million gene models in 51 genome browser assemblies, 6273 quantitative trait loci and >1.4 million genetic loci on 4756 genetic and physical maps contained within 443 mapping sets, complete with standardized metadata. Most notably, 50 new genome browsers that include outputs from the Wheat and Barley PanGenome projects have been created. We provide an example of an expression quantitative trait loci track on the International Wheat Genome Sequencing Consortium Chinese Spring wheat browser to demonstrate how genome browser tracks can be adapted for different data types. To help users benefit more from its data, GrainGenes created four tutorials available on YouTube. GrainGenes is executing its vision of service by continuously responding to the needs of the global small grains community by creating a centralized, long-term, interconnected data repository. Database URL:https://wheat.pw.usda.gov
In this article, we present a joint effort of the wheat research community, along with data and ontology experts, to develop wheat data interoperability guidelines. Interoperability is the ability of two or more systems and devices to cooperate and exchange data
In this article, we present a joint effort of the wheat research community, along with data and ontology experts, to develop wheat data interoperability guidelines. Interoperability is the ability of two or more systems and devices to cooperate and exchange data, and interpret that shared information. Interoperability is a growing concern to the wheat scientific community, and agriculture in general, as the need to interpret the deluge of data obtained through high-throughput technologies grows. Agreeing on common data formats,
Heterogeneous and multidisciplinary data generated by research on sustainable global agriculture and agrifood systems requires quality data labeling or annotation in order to be interoperable. As recommended by the FAIR principles, data, labels, and metadata must use controlled vocabularies and ontologies that are popular in the knowledge domain and commonly used by the community. Despite the existence of robust ontologies in the Life Sciences, there is currently no comprehensive full set of ontologies recommended for data annotation across agricultural research disciplines. In this paper, we discuss the added value of the Ontologies Community of Practice (CoP) of the CGIAR Platform for Big Data in Agriculture for harnessing relevant expertise in ontology development and identifying innovative solutions that support quality data annotation. The Ontologies CoP stimulates knowledge sharing among stakeholders, such as researchers, data managers, domain experts, experts in ontology design, and platform development teams.
Heterogeneous and multidisciplinary data generated by research on sustainable global agriculture and agrifood systems requires quality data labelling to be interoperable. As recommended by the FAIR principles, data, labels and metadata must use controlled vocabularies and ontologies that are popular in the knowledge domain and commonly used by the community. Despite the existence of robust ontologies in the Life Sciences, there is currently no agreed full set of ontologies recommended for data annotation across agricultural research disciplines, which may span genetics, environment, agroecology, biology and socioeconomics. In this paper, we discuss the added value of the Ontologies Community of Practice (CoP) of the CGIAR Platform for Big Data in Agriculture for harnessing relevant ontology expertise. This CoP aims to stimulate knowledge sharing and directly support platform development teams by producing ontologies or contributing missing concepts, recommending best practices and identifying mitigation solutions when gold standard datasets are difficult to attain.
Global population growth and climate change result in the need to develop new and better-adapted wheat varieties. Three integrated resources, the Plant Trait Ontology (TO), the Crop Ontology (CO) and the GrainGenes1 database provide researchers and plant breeders interconnected tools and resources to utilize the large amounts of genetic and genomic data that are available for plant genomics and crop improvement. The TO [1–3] is a reference-level ontology for plant traits, and is a key part of the Planteome Project2 [4]. In the current Planteome Release Version 4.0, the TO consists of more than 1500 plant traits organized into nine upper-level categories: biochemical, biological process, plant growth and development, plant morphology, plant quality, plant vigor, plant stress and yield. As part of the Planteome, the TO is integrated with the Plant Ontology [5,6], and is used to annotate, or link to plant genomics and genetics data objects (e.g. germplasm, QTLs, genes, and proteins) from a wide variety of plant taxa, including important world crops and model plant species. In this Release, there are more than four million annotations linking TO terms to about 165,000 data objects in the Planteome database. Users are encouraged to submit requests for new TO terms or comments through the TO GitHub Issue Tracker 3.
Eating Disorders have the highest mortality rates among psychiatric disorders. In 2014, a national working group, MARSIPAN (The Management of Really Sick Patients with Anorexia Nervosa) arose out of concerns that a number of patients with severe anorexia nervosa were being admitted to general medical units and sometimes deteriorating and dying on those units. After the death of a patient with an eating disorder at the Royal Surrey, there was a realisation that there was a need for improvement in the management of these patients. We reviewed the MARSIPAN guidelines and attended their course, before making links with the local eating disorders service and have built up our knowledge and expertise since 2016. Here we review the adult patients who have been admitted to the Royal Surrey in the last 21 months, looked after by two consultant gastroenterologists and their team of junior doctors and a specialist gastroenterology dietician.
The Plant Ontology (PO) is a community resource consisting of standardized terms, definitions, and logical relations describing plant structures and development stages, augmented by a large database of annotations from genomic and phenomic studies. This paper describes the structure of the ontology and the design principles we used in constructing PO terms for plant development stages. It also provides details of the methodology and rationale behind our revision and expansion of the PO to cover development stages for all plants, particularly the land plants (bryophytes through angiosperms). As a case study to illustrate the general approach, we examine variation in gene expression across embryo development stages in Arabidopsis and maize, demonstrating how the PO can be used to compare patterns of expression across stages and in developmentally different species. Although many genes appear to be active throughout embryo development, we identified a small set of uniquely expressed genes for each stage of embryo development and also between the two species. Evaluating the different sets of genes expressed during embryo development in Arabidopsis or maize may inform future studies of the divergent developmental pathways observed in monocotyledonous versus dicotyledonous species. The PO and its annotation database (http://www.planteome.org) make plant data for any species more discoverable and accessible through common formats, thus providing support for applications in plant pathology, image analysis, and comparative development and evolution.
Assembling large-scale phenotypic datasets for evolutionary and biodiversity studies of plants can be extremely difficult and time consuming. New semi-automated Natural Language Processing (NLP) pipelines can extract phenotypic data from taxonomic descriptions, and their performance can be enhanced by incorporating information from ontologies, like the Plant Ontology (PO) and the Plant Trait Ontology (TO). These ontologies are powerful tools for comparing phenotypes across taxa for large-scale evolutionary and ecological analyses, but they are largely focused on terms associated with flowering plants. We describe a bottom-up approach to identify terms from flagellate plants (including bryophytes, lycophytes, ferns, and gymnosperms) that can be added to existing plant ontologies. We first parsed a large corpus of electronic taxonomic descriptions using the Explorer of Taxon Concepts tool (http://taxonconceptexplorer.org/) and identified flagellate plant specific terms that were missing from the existing ontologies. We extracted new structure and trait terms, and we are currently incorporating the missing structure terms to the PO and modifying the definitions of existing terms to expand their coverage to flagellate plants. We will incorporate trait terms to the TO in the near future. Keywords—Natural Language Processing; Plant Ontology; Plant Trait Ontology; taxonomic descriptions; flagellate plants; phenotypic traits; matrices; phylogeny
Plant stress traits are important breeding targets for all crop species. Massive amounts of research dollars are spent generating data to combat plant diseases and environmental stress. Often this data is used to achieve a single goal, and then left in a repository to never be used again. As a scientific community, we should be striving to make all publicly funded data reusable, and interoperable. This goal is achievable only through careful annotation using universal data and metadata standards. One such standard is the use of a standardized vocabulary, or ontology. This paper presents a semi-automated method to define and label plant stresses using a combination of web scraping and ontology design patterns. Standardizing the definitions and linking plant stress with established hierarchies leverages previous work of developed knowledge bases such as taxonomic classifications and other ontologies. Keywords—ontology; plant pathology; nutrient deficiency; data standards; Planteome; automation; web scraping.
The USDA–ARS has developed and released a transposon‐tagging barley (Hordeum vulgare L.) population comprising 61 lines: TNPs 1, 3, 6, 11, 13, 18, 19, 22, 24, 26–35, 32B, 39–41, 48, 52, 53, 66–69, 71, 74, 79–83, 90, 92–94, 99, 100, 103, 112, 115, 116A, 116B, 117–120, 125, 130, 132, 141, 142, 144, 230, 231, and 234 (Reg. Nos. GS‐83–GS‐143; PIs 687803–687863). Each line is homozygous for a single modified Ds element encoding bar (Ds‐bar) insertion at a different genomic location, except for TNP 120, which has two insertions. The ∼3.6 kb Ds‐bar element is a modified maize (Zea mays L.) Dissociation (Ds) element containing a bar expression cassette. Two of lines (TNP 32B and TNP 48) derive from primary Ds (PDS) transpositions resulting from biolistic co‐introduction of Ds‐bar and maize Ac transposase (AcT) into ‘Golden Promise’. The rest contain either secondary or tertiary transpositions generated by remobilizing Ds‐bar via hybridization of AcT‐expressing Golden Promise plants with, respectively, PDS lines containing primary transpositions or TNP lines containing secondary transpositions. Each line was characterized for the 5′ and/or 3′ sequences flanking the Ds‐bar insertion. Insertions that interrupt native genes or regulatory elements likely will alter or abolish their function, and these lines are expected to be useful for the study of the biological roles of such genes or regulatory sequences. Most of these lines can also be used as parents for the creation of additional lines with novel transpositions of Ds‐bar.
The future of agricultural research depends on data. The sheer volume of agricultural biological data being produced today makes excellent data management essential. Governmental agencies, publishers and science funders require datamanagement plans for publicly funded research. Furthermore, the value of data increases exponentially when they are properly stored, described, integrated and shared, so that they can be easily utilized in future analyses. AgBioData (https://www.agbiodata.org) is a consortium of people working at agricultural biological databases, data archives and knowledgbases who strive to identify common issues in database development, curation and management, with the goal of creating database products that are more Findable, Accessible, Interoperable and Reusable. We strive to promote authentic, detailed, accurate and explicit communication between all parties involved in scientific data. As a step toward this goal, we present the current state of biocuration, ontologies, metadata and persistence, database platforms, programmatic (machine) access to data, communication and sustainability with regard to data curation. Each section describes challenges and opportunities for these topics, along with recommendations and best practices.