Genomic prediction has become central to human, animal and plant biology, enabling quantitative inference of how genetic variation shapes complex traits. Although these domains share statistical foundations, such as linear mixed models, Bayesian regression and deep-learning frameworks, they have advanced largely in parallel. Here we synthesize their methodological evolution and highlight opportunities for integration and deeper collaborations. Agricultural genetics contributed to the mixed-model and Bayesian frameworks underlying modern polygenic scores, while human genomics has driven advances in nonlinear modeling, federated learning and biology-informed artificial intelligence. We propose a roadmap centered on interoperable data standards, shared benchmarks and cross-disciplinary training to unify predictive genomics across species. Together, these efforts establish genomic prediction as a comparative science capable of explaining how genetic information drives form and function across the diversity of life. We emphasize that shared biological architectures and knowledge transfer across species can directly improve the robustness, interpretability and generalizability of predictive models.
Plant breeding is essential for crop improvement, yet progress is often hindered by slow, laborious, and subjective field phenotyping methods. High-throughput phenotyping (HTP), particularly image-based methodologies powered by machine learning, offers a pathway to overcome these limitations. However, achieving robustness and generalization when analyzing diverse genotypes within a crop and across reproductive stages remains challenging and can affect model performance and the accurate extraction of phenotypic features. This study evaluated the performance of semantic segmentation models across a diverse panel of genotypes and distinct crop reproductive stages, using wheat (Triticum aestivum L.), sorghum (Sorghum bicolor L.), and corn (Zea mays L.) as case studies. The primary objectives were to analyze (i) the overall prediction performance on the aggregated dataset for each crop, (ii) the stratified performance by genotype and collection date, and (iii) the temporal and genotypic transferability across growth stages and unseen genotypes. Four distinct smartphone cameras were used to collect images of the reproductive structure across crop growth stages (different collection dates) from 160 corn, 80 sorghum, and 40 wheat genotypes. The total number of images per crop was 2000 for wheat, 4000 for sorghum, and 3840 for corn. Five semantic segmentation models were tested in this study-DeepLabv3+, MaskFormer, SegFormer, SegNet, and U-Net-using the images and respective binary masks for training and testing. The SegFormer model achieved the highest intersection over union (IoU) values for corn (0.90) and sorghum (0.92), while the U-Net model performed best for wheat (0.89). A minor performance decline, with IoU differences up to 0.1, was observed when testing the same model across different genotypes. However, the temporal transferability drops up to 0.5 IoU when training and inferring on different crop growth stages. The main reason for those changes may lie in the natural color and organ architecture temporal changes between the trained and tested datasets when transferring the models across growth stages. These results highlight the urgent need to prioritize robustness and transferability when developing reliable in-field HTP methodologies.
OBJECTIVES:The Genomes to Fields (G2F) 2022 Maize Genotype by Environment (GxE) Prediction Competition aimed to develop models for predicting grain yield for the 2022 Maize GxE project field trials, leveraging the datasets previously generated by this project and other publicly available data.DATA DESCRIPTION:This resource used data from the Maize GxE project within the G2F Initiative [1]. The dataset included phenotypic and genotypic data of the hybrids evaluated in 45 locations from 2014 to 2022. Also, soil, weather, environmental covariates data and metadata information for all environments (combination of year and location). Competitors also had access to ReadMe files which described all the files provided. The Maize GxE is a collaborative project and all the data generated becomes publicly available [2]. The dataset used in the 2022 Prediction Competition was curated and lightly filtered for quality and to ensure naming uniformity across years.
ObjectivesThis release note describes the Maize GxE project datasets within the Genomes to Fields (G2F) Initiative. The Maize GxE project aims to understand genotype by environment (GxE) interactions and use the information collected to improve resource allocation efficiency and increase genotype predictability and stability, particularly in scenarios of variable environmental patterns. Hybrids and inbreds are evaluated across multiple environments and phenotypic, genotypic, environmental, and metadata information are made publicly available.Data descriptionThe datasets include phenotypic data of the hybrids and inbreds evaluated in 30 locations across the US and one location in Germany in 2020 and 2021, soil and climatic measurements and metadata information for all environments (combination of year and location), ReadMe, and description files for each data type. A set of common hybrids is present in each environment to connect with previous evaluations. Each environment had a collaborator responsible for collecting and submitting the data, the GxE coordination team combined all the collected information and removed obvious erroneous data. Collaborators received the combined data to use, verify and declare that the data generated in their own environments was accurate. Combined data is released to the public with minimal filtering to maintain fidelity to the original data.
Objectives This report provides information about the public release of the 2018–2019 Maize G X E project of the Genomes to Fields (G2F) Initiative datasets. G2F is an umbrella initiative that evaluates maize hybrids and inbred lines across multiple environments and makes available phenotypic, genotypic, environmental, and metadata information. The initiative understands the necessity to characterize and deploy public sources of genetic diversity to face the challenges for more sustainable agriculture in the context of variable environmental conditions. Data description Datasets include phenotypic, climatic, and soil measurements, metadata information, and inbred genotypic information for each combination of location and year. Collaborators in the G2F initiative collected data for each location and year; members of the group responsible for coordination and data processing combined all the collected information and removed obvious erroneous data. The collaborators received the data before the DOI release to verify and declare that the data generated in their own locations was accurate. ReadMe and description files are available for each dataset. Previous years of evaluation are already publicly available, with common hybrids present to connect across all locations and years evaluated since this project’s inception.
Advances in phenotyping tools, genomic methodologies, and analytics strategies provide new tools to assess germplasm merit; however, more work is needed to integrate these systems into modern plant breeding approaches. The objective of this work is to integrate genomics and phenomics for yield prediction in maize. A panel of 830 temperate and tropical inbred lines were evaluated for their testcross performance in 2018, and a subset of 400 testcross hybrids were evaluated in 2021 and 2022. These experiments were performed in West Lafayette, IN in a randomized complete block design with two replications. Remote sensing data was collected on a near weekly basis throughout each growing season for RGB (red-green-blue), LiDAR (light detection and ranging), and VNIR (visible near infrared) hyperspectral data and grain yield was harvested with a plot combine. Remote sensing traits extracted include canopy cover, plot volume, plant height, and NDVI. A GBLUP genomic prediction model was used to estimate yield performance in 2018 using data collected in 2021 and 2022. Remote sensing traits were estimated at regular intervals throughout each growing season using random regression modelling. Grain yield was estimated using the genomic estimated yield and the remote sensing traits in a machine learning model. Preliminary results indicate remote sensing can improve prediction accuracy of grain yield compared to genomic prediction alone even with data only collected before flowering. Improved prediction accuracy could benefit hybrid selection, increase genetic gain, and reduce cost in a breeding program.
Lack of high-throughput phenotyping is a bottleneck to breeding for abiotic stress tolerance in crop plants. Efficient and non-destructive hyperspectral imaging can quantify plant physiological traits under abiotic stresses; however, prediction models generally are developed for few genotypes of one species, limiting the broader applications of this technology. Therefore, the objective of this research was to explore the possibility of developing cross-species models to predict physiological traits (relative water content and nitrogen content) based on hyperspectral reflectance through partial least square regression for three genotypes of sorghum (Sorghum bicolor (L.) Moench) and six genotypes of corn (Zea mays L.) under varying water and nitrogen treatments. Multi-species models were predictive for the relative water content of sorghum and corn (R-2 = 0.809), as well as for the nitrogen content of sorghum and corn (R-2 = 0.637). Reflectances at 506, 535, 583, 627, 652, 694, 722, and 964 nm were responsive to changes in the relative water content, while the reflectances at 486, 521, 625, 680, 699, and 754 nm were responsive to changes in the nitrogen content. High-throughput hyperspectral imaging can be used to predict physiological status of plants across genotypes and some similar species with acceptable accuracy.
Objectives Advanced tools and resources are needed to efficiently and sustainably produce food for an increasing world population in the context of variable environmental conditions. The maize genomes to fields (G2F) initiative is a multi-institutional initiative effort that seeks to approach this challenge by developing a flexible and distributed infrastructure addressing emerging problems. G2F has generated large-scale phenotypic, genotypic, and environmental datasets using publicly available inbred lines and hybrids evaluated through a network of collaborators that are part of the G2F’s genotype-by-environment (G × E) project. This report covers the public release of datasets for 2014–2017. Data description Datasets include inbred genotypic information; phenotypic, climatic, and soil measurements and metadata information for each testing location across years. For a subset of inbreds in 2014 and 2015, yield component phenotypes were quantified by image analysis. Data released are accompanied by README descriptions. For genotypic and phenotypic data, both raw data and a version without outliers are reported. For climatic data, a version calibrated to the nearest airport weather station and a version without outliers are reported. The 2014 and 2015 datasets are updated versions from the previously released files [ 1 ] while 2016 and 2017 datasets are newly available to the public.
Sorghum is a staple food for over 500 million people in Sub-Saharan Africa and Asia, however, sorghum proteins are poorly digested when wet-cooked. Three sorghum mutants were identified in a mutagenized population of the inbred line BTx623 that showed a 23-37% increase in wet-cooked protein digestibility compared to their unmutagenized parent. Furthermore, in comparison to the known high lysine, highly digestible sorghum mutant, P721Q, these mutants had 9% more protein overall that was 10% more digestible, had 12% more lysine, as well as better seed hardness. Using bulked segregant analysis based on whole genome sequencing data, we identified unique genomic regions on chromosome 5 of each EMS mutant that are associated with the increase in protein digestibility. Analyzing shared mutations in candidate genes, the high protein digestibility phenotype in one mutant is linked to a point mutation in a novel, ankyrin repeat protein. In another, the increase is associated with a mutation in a kafirin gene and suggests novel genetic modifiers. This study provides material and molecular markers that can be used to enhance sorghum nutritional value, contribute to fighting malnutrition and elucidate new roles for ankyrin-repeat proteins in plants. One sentence summary Mutations in a novel, ankyrin domain protein and genetic modifiers of a known mutation in a seed storage protein lead to increased digestibility of seed proteins in sorghum after wet cooking.
Objectives Crop improvement relies on analysis of phenotypic, genotypic, and environmental data. Given large, well-integrated, multi-year datasets, diverse queries can be made: Which lines perform best in hot, dry environments? Which alleles of specific genes are required for optimal performance in each environment? Such datasets also can be leveraged to predict cultivar performance, even in uncharacterized environments. The maize Genomes to Fields (G2F) Initiative is a multi-institutional organization of scientists working to generate and analyze such datasets from existing, publicly available inbred lines and hybrids. G2F’s genotype by environment project has released 2014 and 2015 datasets to the public, with 2016 and 2017 collected and soon to be made available. Data description Datasets include DNA sequences; traditional phenotype descriptions, as well as detailed ear, cob, and kernel phenotypes quantified by image analysis; weather station measurements; and soil characterizations by site. Data are released as comma separated value spreadsheets accompanied by extensive README text descriptions. For genotypic and phenotypic data, both raw data and a version with outliers removed are reported. For weather data, two versions are reported: a full dataset calibrated against nearby National Weather Service sites and a second calibrated set with outliers and apparent artifacts removed.