AbstractThe biomedical domain has shown that in silico analyses over vast data pools enhances the speed and scale of scientific innovation. This can hold true in agricultural research and guide similar multi-stakeholder action in service of global food security as well (Streich et al. Curr Opin Biotechnol 61:217–225. Retrieved from https://doi.org/10.1016/j.copbio.2020.01.010, 2020). However, entrenched research culture and data and standards governance issues to enable data interoperability and ease of reuse continue to be roadblocks in the agricultural research for development sector. Effective operationalization of the FAIR Data Principles towards Findable, Accessible, Interoperable, and Reusable data requires that agricultural researchers accept that their responsibilities in a digital age include the stewardship of data assets to assure long-term preservation, access and reuse. The development and adoption of common agricultural data standards are key to assuring good stewardship, but face several challenges, including limited awareness about standards compliance; lagging data science capacity; emphasis on data collection rather than reuse; and limited fund allocation for data and standards management. Community-based hurdles around the development and governance of standards and fostering their adoption also abound. This chapter discusses challenges and possible solutions to making FAIR agricultural data assets the norm rather than the exception to catalyze a much-needed revolution towards “translational agriculture”.
Global agriculture is poised to benefit from the rapid advance and diffusion of artificial intelligence (AI) technologies. AI in agriculture could improve crop management and agricultural productivity through plant phenotyping, rapid diagnosis of plant disease, efficient application of agrochemicals and assistance for growers with location-relevant agronomic advice. However, the ramifications of machine learning (ML) models, expert systems and autonomous machines for farms, farmers and food security are poorly understood and under-appreciated. Here, we consider systemic risk factors of AI in agriculture. Namely, we review risks relating to interoperability, reliability and relevance of agricultural data, unintended socio-ecological consequences resulting from ML models optimized for yields, and safety and security concerns associated with deployment of ML platforms at scale. As a response, we suggest risk-mitigation measures, including inviting rural anthropologists and applied ecologists into the technology design process, applying frameworks for responsible and human-centred innovation, setting data cooperatives for improved data transparency and ownership rights, and initial deployment of agricultural AI in digital sandboxes. Machine learning applications in agriculture can bring many benefits in crop management and productivity. However, to avoid harmful effects of a new round of technological modernization, fuelled by AI, a thorough risk assessment is required, to review and mitigate risks such as unintended socio-ecological consequences and security concerns associated with applying machine learning models at scale.
Agricultural research has been traditionally driven by linear approaches dictated by hypothesis-testing. With the advent of powerful data science capabilities, predictive, empirical approaches are possible that operate over large data pools to discern patterns. Such data pools need to contain well-described, machine-interpretable, and openly available data (represented by high-scoring Findable, Accessible, Interoperable, and Reusable—or FAIR—resources). CGIAR's Platform for Big Data in Agriculture has developed several solutions to help researchers generate open and FAIR outputs, determine their FAIRness in quantitative terms 1 , and to create high-value data products drawing on these outputs. By accelerating the speed and efficiency of research, these approaches facilitate innovation, allowing the agricultural sector to respond agilely to farmer challenges. In this paper, we describe the Agronomy Field Information Management System or AgroFIMS, a web-based, open-source tool that helps generate data that is “born FAIRer” by addressing data interoperability to enable aggregation and easier value derivation from data. Although license choice to determine accessibility is at the discretion of the user, AgroFIMS provides consistent and rich metadata helping users more easily comply with institutional, founder and publisher FAIR mandates. The tool enables the creation of fieldbooks through a user-friendly interface that allows the entry of metadata tied to the Dublin Core standard schema, and trial details via picklists or autocomplete that are based on semantic standards like the Agronomy Ontology (AgrO). Choices are organized by field operations or measurements of relevance to an agronomist, with specific terms drawn from ontologies. Once the user has stepped through required fields and desired modules to describe their trial management practices and measurement parameters, they can download the fieldbook to use as a standalone Excel-driven file, or employ via free Android-based KDSmart, Fieldbook, or ODK applications for digital data collection. Collected data can be imported back to AgroFIMS for statistical analysis and reports. Development plans for 2021 include new features such ability to clone fieldbooks and the creation of agronomic questionnaires. AgroFIMS will also allow archiving of FAIR data after collection and analysis from a database and to repository platforms for wider sharing.
Heterogeneous and multidisciplinary data generated by research on sustainable global agriculture and agrifood systems requires quality data labeling or annotation in order to be interoperable. As recommended by the FAIR principles, data, labels, and metadata must use controlled vocabularies and ontologies that are popular in the knowledge domain and commonly used by the community. Despite the existence of robust ontologies in the Life Sciences, there is currently no comprehensive full set of ontologies recommended for data annotation across agricultural research disciplines. In this paper, we discuss the added value of the Ontologies Community of Practice (CoP) of the CGIAR Platform for Big Data in Agriculture for harnessing relevant expertise in ontology development and identifying innovative solutions that support quality data annotation. The Ontologies CoP stimulates knowledge sharing among stakeholders, such as researchers, data managers, domain experts, experts in ontology design, and platform development teams.
COVID-19 has shown that Findable, Accessible, Interoperable, and Reusable (FAIR) data assets are the building blocks for data-driven, collaborative and agile crisis response and resilience over the long term. Responding rapidly and effectively to disruptions in agriculture similarly hinges around enabling discovery of and access to publications, data, and data products that are interpretable and interoperable – for humans and machines. This is also a critical piece of the strategy for enhancing the impact of research and development in the agricultural domain in general, and for catalyzing innovation in our efforts to fuel a “translational agriculture” revolution. There is a strong model in the biomedical sector to achieve seamless discoverability, interlinkages and interoperability across their data resources. Guided by this, CGIAR’s Big Data Platform has developed and/or curated several openly available solutions to enable good data management practices leading to the generation of open and FAIR data assets, along with tools to access them, and an analytical environment and other services to more seamlessly and ethically process and derive insight from data. This talk will provide an overview of these as a response to a potential data gap in this moment of crisis coupled with an increased recognition of the importance of open, human and machine interoperable data through the systematic use of these tools and services.
Heterogeneous and multidisciplinary data generated by research on sustainable global agriculture and agrifood systems requires quality data labelling to be interoperable. As recommended by the FAIR principles, data, labels and metadata must use controlled vocabularies and ontologies that are popular in the knowledge domain and commonly used by the community. Despite the existence of robust ontologies in the Life Sciences, there is currently no agreed full set of ontologies recommended for data annotation across agricultural research disciplines, which may span genetics, environment, agroecology, biology and socioeconomics. In this paper, we discuss the added value of the Ontologies Community of Practice (CoP) of the CGIAR Platform for Big Data in Agriculture for harnessing relevant ontology expertise. This CoP aims to stimulate knowledge sharing and directly support platform development teams by producing ontologies or contributing missing concepts, recommending best practices and identifying mitigation solutions when gold standard datasets are difficult to attain.
The CGIAR International Research Centers collect large amounts of data through on-station and on-farm experiments, surveys and others means of data collection. These data are extremely valuable and application of the FAIR (Findable, Accessible, Interoperable, Reusable) principles to these data would increase the value, particularly for data which are suitable for quantitative analyses. Through proper metadata description and data annotation, it is possible to allow datasets of interest to researchers to be made interoperable. The Big Data Platform initiative of CGIAR is committed to this effort and the AgMIP data interoperability tools will be a useful tool in the process. The Global Agricultural Research Data Innovation & Acceleration Network (GARDIAN; http://gardian.bigdata.cgiar.org) web interface already provides discovery and access to selected CGIAR data and publications. Two main efforts are taking place now to address these issues of data interoperability of CGIAR data. The first deals with collection of new data. The electronic Field Book, currently under development by the CGIAR, will allow new data collected in field experiments to be automatically annotated with descriptive metadata which will allow data to be discovered and reused. The second focus for CGIAR data interoperability is on capturing old datasets, which are in distributed databases, diverse formats, and do not use a consistent vocabulary. A project is underway to allow data to be annotated with standardized metadata to allow automated discovery and translation to standardized formats useful in models and other types of quantitative analyses. The first phase of this project included the design of the dataset annotations and manual implementation of these annotations for a limited number of datasets as a proof of concept. Annotation of datasets was accomplished with three “Sidecar files” which contain additional metadata beyond the core GARDIAN metadata. Regardless of the physical location and format of the raw data in this distributed system, these metadata files will be readily available for rapid data searches for specific analytics. Phase 2 will allow these sidecar files to be created and used in data translation in a more automated way using an expanded set of AgMIP data translators.
Scientific innovation is increasingly reliant on data and computational resources. Much of today’s life science research involves generating, processing, and reusing heterogeneous datasets that are growing exponentially in size. Demand for technical experts (data scientists and bioinformaticians) to process these data is at an all-time high, but these are not typically trained in good data management practices. That said, we have come a long way in the last decade, with funders, publishers, and researchers themselves making the case for open, interoperable data as a key component of an open science philosophy. In response, recognition of the FAIR Principles (that data should be Findable, Accessible, Interoperable and Reusable) has become commonplace. However, both technical and cultural challenges for the implementation of these principles still exist when storing, managing, analysing and disseminating both legacy and new data. COPO is a computational system that attempts to address some of these challenges by enabling scientists to describe their research objects (raw or processed data, publications, samples, images, etc.) using community-sanctioned metadata sets and vocabularies, and then use public or institutional repositories to share them with the wider scientific community. COPO encourages data generators to adhere to appropriate metadata standards when publishing research objects, using semantic terms to add meaning to them and specify relationships between them. This allows data consumers, be they people or machines, to find, aggregate, and analyse data which would otherwise be private or invisible, building upon existing standards to push the state of the art in scientific data dissemination whilst minimising the burden of data publication and sharing.
Agricultural systems models are widely used in agricultural research. Most such models are data intensive but having been developed by independent groups or individuals, their required data formats and semantics vary greatly. The Agricultural Model Intercomparison and Improvement Project (AgMIP) promotes comparisons among cropping system models through use of common datasets. Interoperability tools were developed to allow multiple models to access consistent input data and to harmonize model outputs regardless of internal model requirements. Simulated responses could then be compared for a target set of simulation protocols. AgMIP crop model datasets were harmonized using the vocabularies and standards developed by the International Consortium for Agricultural Systems Application (ICASA). The ICASA Data Dictionary describes terms and units for data related to field crop experiments including weather, soils, crop management, and field measurements. One use case for the AgMIP data interoperability tools involved intercomparison of over 30 wheat models using detailed field experimental datasets. Standardizing the data formats reduced ambiguity in disseminating data to this large group of modeling teams working independently. Other AgMIP activities using the data interoperability tools were the Regional Integrated Assessments, consisting of eight research teams in Sub-Saharan Africa and South Asia performing climate impact and adaptation assessments for their regions, using multiple models. The AgMIP data interoperability tools are also used by non-AgMIP researchers, who recognize the need for harmonizing agricultural data from diverse sources for quantitative analyses. Applications include the PSIMs global gridded modeling platform, the CGIAR CCAFS Regional Agricultural Forecast Tool (CRAFT), S-World global soil mapping software, and others. Additionally, the ICASA Data Dictionary has been aligned with the Crop Ontology and the Agronomy Ontology (AgrO) so that AgMIP data can be annotated and made available using semantic technologies.
CGIAR’s Big Data Platform will harness the capacity of Big Data to accelerate and enhance the impact of international agricultural research by providing opportunities for researchers to discover, share, analyze and visualize agricultural data to generate rapid, actionable insights. The platform encompasses three modules: Module 1 addresses the necessity to enable a culture that values data as a product with global public good potential, and to develop, support, and apply best practices to managing and making it widely available following FAIR (Findable, Accessible, Interoperable, Reusable) Principles. Module 2 fosters collaboration and convening around big data and agricultural development via ambitious external partnerships to deliver the potential of big data to smallholder agriculture. Module 3 inspires big data approaches that use big data analytics and ICTs to provide multi-disciplinary data to researchers, deliver novel information to farmers, monitor the state of agriculture and food security in real time and inform critical national, regional and global policies and decisions. The Platform has developed a prototype infrastructure to harvest research data and publications from interoperable repositories and M+E platforms, which will be demonstrated at this meeting. Tools for seamless analysis and visualization, and to enable secure data sharing, storage and computation will soon be available as part of this infrastructure.
CGIAR is a global research partnership of 15 geographically and scientifically diverse Centers dedicated to reducing poverty, enhancing food and nutrition security, and improving natural resource management. The Centers are charged with accelerating innovation to tackle challenges at a variety of scales from the local to the global. This requires data and other research outputs to be findable, accessible, interoperable, and reusable – that is, open via FAIR principles, and inter-linked where relevant. CGIAR Centers have made strong progress in implementing publication and data repositories; however, many of these still represent silos whose contents are not generally easily discoverable or inter-linked (e.g., agronomic trial data with socioeconomic or adoption data in the same geographies). In the absence of such interoperability-mediated discovery, “open” is of limited utility. The overall goal is for CGIAR’s trove of research data and associated information to be indexed and interlinked through a demand-driven cyberinfrastructure for agriculture, ensuring that research outputs are discoverable by humans and machines, and reusable via appropriate licensing to enhance innovation, uptake and impact. There are challenges to achieving this goal, not only across CGIAR, but for the agricultural domain in general. Among the foremost hurdles is that “open” tends to remain an unfunded mandate, making it difficult to operationalize effectively. Further, there is still significant concern on the part of scientists about making data open – largely centered around issues of trust, time, and quality – resulting in repositories frequently exposing metadata rather than the data sets themselves. While the ability to find metadata about resources qualifies as improvement, it continues to impose barriers to data access, discoverability, integration, and analysis, without which complex challenges to global agriculture development cannot be effectively addressed. CGIAR is addressing the urgent need to create a data sharing culture and enabling environment for Open Access and Open Data (OA/OD) that includes projects planning for OA/OD and allocating funds to support it, in parallel with the technical infrastructure mentioned above. While the technology necessary to enable FAIR outputs exists, achieving success implies data provider and consumer trust and buy-in, agreement and adherence to interoperability standards and/or mapping across varied approaches, and compliance with guidelines (including those on citation and licensing governing content reuse). Agricultural institutions, including CGIAR, are only now beginning to address these issues systematically, to agree on and adopt standards-based systems and processes, and to build cross-walks across differing schemas. Through its Open Access and Open Data initiative funded by the Bill and Melinda Gates Foundation, and via plans for an ambitious Big Data and ICT Platform , CGIAR is developing technical and cultural approaches that will enable research content to be consistently and seamlessly discovered, interlinked, and analyzed across its Centers. This paper describes the strategy used to identify the specific contexts and challenges faced by Centers in building an infrastructure and culture for OA/OD across CGIAR, with the ultimate goal of achieving greater impact in agricultural research for development.
CGIAR’s Big Data Platform aims to accelerate and enhance the impact of international agricultural research by enabling researchers to discover, analyze and visualize agricultural data on a large scale to generate rapid, actionable insights. A key aspiration is to build and nurture partnerships across CGIAR Centers and a variety of external partners to deliver effective data management, analytics and ICT-focused solutions to target geographies and communities. The platform encompasses three modules: Module 1 addresses the necessity to enable a culture that values data as a product in itself with global public good potential, and applies best practices to managing and making it widely available following FAIR (Findable, Accessible, Interoperable, Reusable) data principles. An infrastructure is also envisioned to harvest and make discoverable research resources from interoperable repositories and M+E platforms, coupled with tools for seamless analysis and visualization, leading to data-based decision support and foresight. Module 2 fosters collaboration and convening around big data and agricultural development via ambitious external partnerships to deliver the potential of big data to smallholder agriculture. Module 3 inspires big data approaches that deliver development outcomes through projects that solve key challenges. These may include projects that use big data analytics and ICTs to provide multi-disciplinary data to researchers, deliver novel information to farmers, monitor the state of agriculture and food security in real time and inform critical national, regional and global policies and decisions.
Background: Opportunities to use data and information to address challenges in international agricultural research and development are expanding rapidly. The use of agricultural trial and evaluation data has enormous potential to improve crops and management practices. However, for a number of reasons, this potential has yet to be realized. This paper reports on the experience of the AgTrials initiative, an effort to build an online database of agricultural trials applying principles of interoperability and open access. Methods: Our analysis evaluates what worked and what did not work in the development of the AgTrials information resource. We analyzed data on our users and their interaction with the platform. We also surveyed our users to gauge their perceptions of the utility of the online database. Results: The study revealed barriers to participation and impediments to interaction, opportunities for improving agricultural knowledge management and a large potential for the use of trial and evaluation data. Conclusions: Technical and logistical mechanisms for developing interoperable online databases are well advanced. More effort will be needed to advance organizational and institutional work for these types of databases to realize their potential.
The Crop Ontology (CO), as service of the Integrated Breeding Platform ( www.integratedbreeding.net ) and provider of controlled trait description for the Breeding Management System, is expending to new crops and will be completed by an Agronomy Ontology (AgrO). The Crop Ontology ( www.cropontology.org ) provides harmonized and validated breeders’ trait names, measurement methods, scales and variables for currently 20 crops namely : cassava, banana, barley, chickpea, common bean, cowpea, groundnut, lentil, maize, oat, pearl millet, pigeon pea, potato, rice, sorghum, soybean, sweet potato, vitis, wheat and yam. The NextGeneration Breeding Databases developed by Boyce Thompson Institute for banana, cassava, potato, sweet potato also embed the Crop Ontology traits. Combining results of field management practices with crop traits is important to fully understand the dynamic of varying factors within any cropping system. Therefore, an Agronomy Ontology (AgrO) is under development and currently compiles 350 variables selected out of the set of the variables produced by the International Consortium for Agricultural Systems Applications (ICASA). All variables were documented with a method and a scale to follow the Crop Ontology model. A new Trait Dictionary Template was released in 2015 that now includes the ‘standard variable’ composed by a property, a method and a scale and needed to accurately annotate the measurements stored in the databases and asupport the creation of standard electronic fieldbooks. This template is currently tested by INRA to describe the Wheat traits of Ephesis database, by The Crop Ontology project is a partner of the NSF-awarded project Planteome to contribute improving the reference ontologies for plants by mapping the crop traits of CO to reference ontologies ( http://www.planteome.org/ ) which is also a DivSeek project ( http://www.divseek.org/ ). Additionally, the Crop Ontology is a partner of the pilot project Agroportal ( http://agroportal.lirmm.fr/ ) developed by LIRRM, IRD and Bioversity.
The possibly higher lignin contents or altered carbon (C) allocation patterns in Bt corn hybrids, compared to their non-transgenic parental varieties, may alter the quality and quantity of plant residues incorporated into soils. In this study, we conducted a greenhouse experiment to investigate C allocation and lignin contents in Cry3Bb Bt and NonBt corn as affected by corn rootworm (CRW, Diabrotica virgifera virgifera) infestation. The partitioning of photosynthate C to various plant components was measured as short-term C allocation by a (CO2)-C-13 pulse-labeling system, and the lignin content or concentration was measured by the acid detergent method. Results showed that NonBt corn was significantly taller than Bt corn at all measured stages, likely resulting from inherent variability in the parental lines used in this study. However, there was no significant genotype effect on C-13 allocation, total C and lignin content or concentration in plant tissues without CRW infestation. With CRW, the percentage of fixed C-13 during labeling allocated to roots was significantly lower in NonBt than in Bt corn, likely caused by CRW damage in NonBt roots. The lignin content in NonBt roots was significantly higher with than without the CRW infestation, implying the stimulating effect of CRW possibly due to the triggered reaction of induced systemic resistance. Overall, the transgenic Cry3Bb event in MON863 corn did not affect measured variables, but CRW resistance in Bt corn affected the pattern of short-term C allocation and root lignin content compared to NonBt corn in the CRW presence, and has implications for soil C dynamics. (C) 2012 Elsevier B.V. All rights reserved.
ABSTRACT Despite the rapid adoption of crops expressing the insecticidal Cry protein(s) from Bacillus thuringiensis (Bt), public concern continues to mount over the potential environmental impacts. Reduced residue decomposition rates and increased tissue lignin concentrations reported for some Bt corn hybrids have been highlighted recently as they may influence soil carbon dynamics. We assessed the effects of MON863 Bt corn, producing the Cry3Bb protein against the corn rootworm complex, on these aspects and associated decomposer communities by terminal restriction fragment length polymorphism (T-RFLP) analysis. Litterbags containing cobs, roots, or stalks plus leaves from Bt and unmodified corn with (non-Bt+I) or without (non-Bt) insecticide applied were placed on the soil surface and at a 10-cm depth in field plots planted with these crop treatments. The litterbags were recovered and analyzed after 3.5, 15.5, and 25 months. No significant effect of treatment (Bt, non-Bt, and non-Bt+I) was observed on initial tissue lignin concentrations, litter decomposition rate, or bacterial decomposer communities. The effect of treatment on fungal decomposer communities was minor, with only 1 of 16 comparisons yielding separation by treatment. Environmental factors (litterbag recovery year, litterbag placement, and plot history) led to significant differences for most measured variables. Combined, these results indicate that the differences detected were driven primarily by environmental factors rather than by any differences between the corn hybrids or the use of tefluthrin. We conclude that the Cry3Bb corn tested in this study is unlikely to affect carbon residence time or turnover in soils receiving these crop residues.