Plant digital twins (DTs) are dynamic, data-aware virtual representations that integrate plant structure, physiology, and environmental inputs to simulate plant function, growth, and performance. By combining process-based models, phenomics data, and environmental states in a continuous feedback loop linking physical and virtual plants, DTs offer a novel approach for investigating complex biological systems. In this opinion article, we outline the conceptual foundations of DTs and how they offer new opportunities to bridge the genotype-to-phenotype gap and aid in crop improvement. Despite their promise, major challenges remain in DT development, scalability, and accessibility. Addressing these challenges will be critical for enabling the broad adoption of DTs as transformative tools for accelerating plant science research and crop improvement.
IntroductionPhotosynthesis is fundamental to agricultural productivity, but its relatively low light-to-biomass conversion efficiency represents an opportunity for enhancement. High-throughput phenotyping is crucial for unraveling the genetic basis of variation in photosynthetic activity. However, the heritability of chlorophyll fluorescence parameters measured during the day is often low as a result of high levels of variation introduced by environmental fluctuations.MethodsTo address these limitations, we measured fluorescence phenotypes at night, leveraging natural dark adaptation to minimize environmental noise.ResultsNight measurement significantly increased the heritability of fluorescence traits compared to daytime measurements, with the maximum quantum yield of photosystem II (Fv/Fm) showing an increase in heritability from 0.32 to 0.72. Genome-wide association studies (GWAS) conducted using three photosynthetic fluorescence traits measured at night across two growing seasons identified several significant single nucleotide polymorphisms (SNPs). Notably, two candidate genes near SNPs linked to multiple fluorescence traits, Zm00001eb271820 and Zm00001eb012130, have known roles in photosynthesis regulation. Four of the significant signal nucleotide polymorphisms identified in GWAS conducted using nighttime collected data also exhibited statistically significant associations with the same phenotypes during the day. In a majority of other cases, direction of effect was consistent but greater variance in day measured data relative to night measured data resulted in the differences not being statistically significant.DiscussionThese results highlight the effectiveness of phenotyping photosynthetic traits at night in reducing environmental noise and enhancing the discovery of genomic intervals related to photosynthesis. While nighttime data collection may not be applicable for all photosynthetic traits, it offers a promising avenue for advancing our understanding of the genetic variation of photosynthesis in modern crop species.
Improving crop resilience in the face of increasingly extreme and unpredictable weather and reduced access to agricultural inputs such as nitrogen fertilizer and water will require an improved understanding of phenotypic plasticity in crops. To understand the roles of different component traits in determining overall plasticity for grain yield, we generated data from a panel of 122 maize (Zea mays) hybrids grown in replicated field trials in 34 environments spanning 1126 km (700 miles) of the US Corn Belt. We observed that the levels of genetic versus environmental control and the relationships between mean parent release year, overall performance, and linear plasticity were trait-dependent across the 18 agronomic and yield components studied. Importantly and unexpectedly, we observed no clear tradeoff between linear plasticity and mean performance and found only rare examples where genotype-by-environment interactions would alter selection decisions based on the environments tested in our dataset. Furthermore, we showed that overall plasticity was repeatable and that plasticity in response to nitrogen fertilization was not, which may help explain the limited success in breeding for nitrogen use efficiency. Together, these findings improve our understanding of phenotypic plasticity, with implications for maize breeding.
Prebiotics are useful tools for shifting the gut microbiome and metabolome to confer immune, metabolic, and preventive benefits for human disease. However, there are emerging concerns about methods for discovering optimal candidate compounds and their sustainable production for long-term success in targeting mechanisms of disease. In this perspective, we review the current state of prebiotics moving from nutrition to function in health, highlighting the opportunities for using food as medicine in treating a model disease, inflammatory bowel disease, discussing the ways crop breeding can be used to identify and improve the functional (beyond nutritional) value of a prebiotic and the critical role that food processing plays in sustaining the integrity and scalability of the prebiotic compound while influencing consumer adoption of these agents. This information supports a trajectory for the future of food as medicine to be moving from population-scale dietary guidelines to personalized or precision nutrition guidelines, from "eat more fiber" to "eat this specific type of fiber that has been enhanced in and processed from this specific genotype of food crop". This could include developing designer fibers by crop breeding to develop panels of genotypes with different fiber structures that deliver reliable shifts in the microbiome and metabolome for most patients with the targeted disease. Finally, there needs to be a commitment to sustainability of the plant-derived product so that future generations will have access to the same benefits of the food as medicine initially developed for disease prevention and treatment.
Understanding genotype-by-environment (G × E) interactions that underlie phenotypic variation, when observed for complex traits in multi-environment trials, is important for biological discovery and for crop improvement. The regression-on-the-mean model is an approach to observe G × E trends for complex traits across a gradient of environmental inputs. Biologically relevant environmental index values can be utilized to quantify phenotypic plasticity of individuals by correlating environmental means and environmental parameters within specific time windows. By accounting for trait stability, improvements can be made in genome-wide association studies and genomic prediction models involving data with high volumes of environments, genotypes, and their interaction effects. Here, field data collected through the national hybrid maize (Zea mays L.) Genomes-to-Fields project was analyzed. Reaction norm parameters were obtained from photothermal ratio (PTR) indices for hybrid grain yield (GY) using three separate tester populations across 29 diverse environments. The PTR time windows most correlated with the average GY were discovered to differ by tester but were confounded by region. Using 100,000 single-nucleotide polymorphisms (SNPs), we discovered 96 quantitative trait loci (QTLs) significantly associated with GY and six QTLs significantly associated with GY stability. The modified, regression-on-the-mean genomic prediction model using PTR-estimated reaction norm parameters of each hybrid worked nearly as well as a traditional, additive genomic prediction model using the G × E interaction terms but took 192× less time. The PTR genomic prediction model predicted untested environment performance (0.57-0.71) better than untested hybrid performance (0.26-0.37). This study suggests improved potential for multi-environment genomic predictions by incorporating environmental measures to dissect the complexities of differential performance of genotypes across environments.
Differences in canopy architecture play a role in determining both the light and water use efficiency. Canopy architecture is determined by several component traits, including leaf length, width, number, angle, and phyllotaxy. Phyllotaxy may be among the most difficult of the leaf canopy traits to measure accurately across large numbers of individual plants. As a result, in simulations of the leaf canopies of grain crops such as maize and sorghum, this trait is frequently approximated as alternating 180° angles between sequential leaves. We explore the feasibility of extracting direct measurements of the phyllotaxy of sequential leaves from 3D reconstructions of individual sorghum plants generated from 2D calibrated images and test the assumption of consistently alternating phyllotaxy across a diverse set of sorghum genotypes. Using a voxel-carving-based approach, we generate 3D reconstructions from multiple calibrated 2D images of 366 sorghum plants representing 236 sorghum genotypes from the sorghum association panel. The correlation between automated and manual measurements of phyllotaxy is only modestly lower than the correlation between manual measurements of phyllotaxy generated by two different individuals. Automated phyllotaxy measurements exhibited a repeatability of R2 = 0.41 across imaging timepoints separated by a period of two days. A resampling based genome wide association study (GWAS) identified several putative genetic associations with lower-canopy phyllotaxy in sorghum. This study demonstrates the potential of 3D reconstruction to enable both quantitative genetic investigation and breeding for phyllotaxy in sorghum and other grain crops with similar plant architectures.
Maize (Zea mays L.) is the world's most productive grain crop and a cornerstone of global food supply. However, in temperate agricultural systems, maize exhibits 2 key anomalies. First, as a tropical species, maize cannot be planted in the cold conditions of early spring when light and natural soil nitrogen are available, resulting in a shorter growing season and creating a seasonal mismatch between nitrogen accessibility and demand. Second, maize kernel protein is a major nitrogen sink, driving fertilizer demand because of the scale of cultivation. This inefficient mismatch stems from modern maize's uses and the modest nutritional value of storage proteins. To address these anomalies, we established the Circular Economy that Reimagines Corn Agriculture initiative. Our vision requires advances in 3 research areas: (ⅰ) developing cold and frost tolerance during germination and early growth to enable the use of spring nitrogen and light resources; (ⅱ) reducing nitrogen allocation to grain by reducing low-quality storage proteins and developing alternative nitrogen sinks; and (ⅲ) stabilizing soil nitrogen by enhancing biological nitrification inhibition. We present blueprints for a nitrogen-efficient, cold-tolerant maize designed to utilize the full growing season, enabling farmers in temperate regions to fully leverage maize's C4 photosynthesis, reduce fertilizer inputs, increase yields, and minimize environmental impact.
Accurately predicting yield during the growing season enables improved crop management and better resource allocation for both breeders and growers. Existing yield prediction models for an entire field or individual plots are based on satellite-derived vegetation indices (VIs) and widely used machine learning-based feature extraction models, including principal component analysis (PCA) and autoencoders (AE). Here, we significantly enhance pre-harvest yield prediction at plot-scale using Compositional Autoencoders (CAE) — a deep-learning-based feature extraction approach designed to disentangle genotype (G) and environment (E) features — on high-resolution, plot-level satellite imagery. Our approach uses a dataset of approximately 4,000 satellite images collected from replicated plots of 84 hybrid maize varieties grown at five distinct locations across the U.S. Corn Belt. By deploying the CAE model, we improve the separation of genotype and environment effects, enabling more accurate incorporation of genotype-by-environment (GxE) interactions for downstream prediction tasks. Results show that the CAE-based features improve early-stage yield predictions by up to 10% compared to traditional autoencoder-based features and outperform vegetation indices (VIs) by 9% across various growth stages. The CAE model also excels in separating environmental factors, achieving a high silhouette score of 0.919, indicating effective clustering of environmental features. Moreover, the CAE consistently outperforms standard models in unseen environments and unseen genotypes yield predictions, demonstrating strong generalizability. This study demonstrates the value of disentangling G and E effects for providing more accurate and early yield predictions that support informed decision-making in precision agriculture and plant breeding.
Quantifying the variation in yield component traits of maize (Zea mays L.), which together determine the overall productivity of this globally important crop, plays a critical role in plant genetics research, plant breeding, and the development of improved farming practices. Grain yield per acre is calculated by multiplying the number of plants per acre, ears per plant, number of kernels per ear, and the average kernel weight. The number of kernels per ear is determined by the number of kernel rows per ear multiplied by the number of kernels per row. Traditional manual methods for measuring these two traits are time-consuming, limiting large-scale data collection. Recent automation efforts using image processing and deep learning encounter challenges such as high annotation costs and uncertain generalizability. We tackle these issues by exploring Large Vision Models for zero-shot, annotation-free maize kernel segmentation. By using an open-source large vision model, the Segment Anything Model (SAM), we segment individual kernels in RGB images of maize ears and apply a graph-based algorithm to calculate the number of kernels per row. Our approach successfully identifies the number of kernels per row across a wide range of maize ears, showing the potential of zero-shot learning with foundation vision models combined with image processing techniques to improve automation and reduce subjectivity in agronomic data collection. All our code is open-sourced to make these affordable phenotyping methods accessible to everyone.
AbstractSeed color is a complex phenotype linked to both the impact of grains on human health and consumer acceptance of new crop varieties. Today, seed color is often quantified via qualitative human assessment or biochemical assays for specific colored metabolites. Imaging‐based approaches have the potential to be more quantitative than human scoring while being lower cost than biochemical assays. We assessed the feasibility of employing image analysis tools trained on rice (Oryza sativa) or wheat (Triticum aestivum) seeds to quantify seed color in sorghum (Sorghum bicolor) using a dataset of 1,500 images. Quantitative measurements of seed color from images were substantially more consistent across biological replicates than human assessment. Genome‐wide association studies conducted using color phenotypes for 682 sorghum genotypes identified more signals near known seed color genes in sorghum with stronger support than manually scored seed color for the same experiment. Previously unreported genomic intervals linked to variation in seed color in our study co‐localized with a gene encoding an enzyme in the biosynthetic pathway leading to anthocyanins, tannins, and phlobaphenes—colored metabolites in sorghum seeds—and with the sorghum ortholog of a transcription factor shown to regulate several enzymes in the same pathway in rice. The cross‐species transferability of image analysis tools, without the retraining, may aid efforts to develop higher value and health‐promoting crop varieties in sorghum and other specialty and orphan grain crops.
Comprehensive maps of functional variation at transcription factor (TF) binding sites (cis-elements) are crucial for elucidating how genotype shapes phenotype. Here, we report the construction of a pan-cistrome of the maize leaf under well-watered and drought conditions. We quantified haplotype-specific TF footprints across a pan-genome of 25 maize hybrids and mapped over 200,000 variants, genetic, epigenetic, or both (termed binding quantitative trait loci (bQTL)), linked to cis-element occupancy. Three lines of evidence support the functional significance of bQTL: (1) coincidence with causative loci that regulate traits, including vgt1, ZmTRE1 and the MITE transposon near ZmNAC111 under drought; (2) bQTL allelic bias is shared between inbred parents and matches chromatin immunoprecipitation sequencing results; and (3) partitioning genetic variation across genomic regions demonstrates that bQTL capture the majority of heritable trait variation across ~72% of 143 phenotypes. Our study provides an auspicious approach to make functional cis-variation accessible at scale for genetic studies and targeted engineering of complex traits.
Natural genetic variation in photosynthesis-related traits can help both to identify genes involved in regulating photosynthetic processes and to develop crops with improved productivity and photosynthetic efficiency. However, rapidly fluctuating environmental parameters create challenges for measuring photosynthetic parameters in large populations under field conditions. We measured chlorophyll fluorescence and absorbance-based photosynthetic traits in a maize diversity panel in the field using an experimental design that allowed us to estimate and control multiple confounding factors. Controlling the impact of day of measurement and light intensity as well as patterns of two-dimensional spatial variation in the field increased heritability for 11 out of 14 traits measured. We were able to identify high-confidence genome-wide association study (GWAS) signals associated with variation in four spatially corrected traits (the quantum yield of PSII, non-photochemical quenching, redox state of QA, and relative chlorophyll content). Insertion alleles for Arabidopsis orthologs of three candidate genes exhibited phenotypes consistent with our GWAS results. Collectively these results illustrate the potential of applying best practices from quantitative genetics research to address outstanding questions in plant physiology and to understand natural variation in photosynthesis.
Current efforts to detect and evaluate crop resistance to insect pests are limited by traditional phenotyping methods, which are time-consuming and highly variable. Sugarcane aphid (SCA; Melanaphis sacchari) is a major pest of sorghum in North America that has emerged over the last decade and negatively impacts plant growth and development. The spectral reflectance data in visible, near infrared and shortwave infrared range (VIS-NIR-SWIR; 400-2500 nm) have been used to measure plant traits related to stress responses, nutrient dynamics, and physiological status. We examined the potential of spectral features (VIS-NIR-SWIR) to improve the current phenotyping methods in monitoring sorghum resistance mechanisms to SCA. We used eight sorghum lines that displayed varied levels of resistance to SCA and collected data from control and aphid-infested plants. Spectral feature data were collected using a leaf spectrometer, while plant physiological and chlorophyll fluorescence parameters were measured with LICOR and MultispeQ devices. The random forest classifier model differentiated the control and aphid-infested plants with a high accuracy of 87.4% with important spectral features in the VIS-NIR spectral range, particularly from 508 to 573 nm and 715 to 728 nm. The spectral indices exhibit significant difference in Greenness Index and Plant Senescence Reflectance Index in aphid-infested susceptible lines (BTx623, SC1345) compared with control plants. In addition, plant physiological parameters, such as stomatal conductance and chlorophyll fluorescence, showed significantly higher value for aphid-infested resistant line (Tx2783) compared with susceptible line (BTx623) in both treatments. Further, a partial least square regression model demonstrated medium predictive capability for plant physiological parameters related to fluorescence. In summary, spectral features at VIS-NIR range demonstrated promising results in differentiating aphid-infested sorghum plants. This is a proof-of-concept study on potential of spectral sensing to develop an effective monitoring and phenotyping plant resistance to aphids.
Sorghum is emerging as an ideal genetic model for designing high-biomass bioenergy crops. Biomass yield, a complex trait influenced by various plant architectural features, is typically regulated by numerous genes. This study aims to dissect the genetic mechanisms underlying fourteen plant architectural and ten biomass yield traits in a sorghum association panel (SAP) across two growing seasons. We identified 321 associated loci via genome-wide association studies involving 234,264 single nucleotide polymorphisms (SNPs). These loci encompass both genes with a priori links to biomass traits, such as maturity, dwarfing (Dw), leafbladeless1, cryptochrome, and several loci not previously linked to roles in determining these traits. We identified 22 pleiotropic loci associated with variation in multiple phenotypes. Three of these loci, located on chromosomes 3 (S03_15463061), 6 (S06_42790178; Dw2), and 9 (S09_57005346; Dw1), exert significant and consistent effects on multiple traits. Additionally, we identified three genomic hotspots on chromosomes 6, 7, and 9, containing multiple SNPs associated with variation in plant architecture and biomass yield traits. Positive correlations were observed among linked SNPs close to or within the same genomic regions. Thirteen haplotypes were identified from these positively correlated SNPs on chr 6, with haplotypes 8 and 11 emerging as optimal combinations, exhibiting pronounced effects on the traits. Lastly, network analysis revealed that loci associated with flowering, plant heights, leaf characteristics, plant number, and tiller number per plant were highly interconnected with other genetic loci linked to plant architecture and biomass yield traits. The pyramiding of favorable alleles related to these traits holds promise for enhancing the future development of bioenergy sorghum crops.
In‐context promoter bashing via genome editing is a route to identify and characterize critical regulatory regions that govern expression of genes of interest. The outcomes of in‐context promoter bashing can be used to inform editing strategies to modulate the expression of selected gene models in a desired fashion. Here, we employed in‐context promoter bashing to characterize the proximal upstream regulatory regions of sorghum genes encoding phosphoenolpyruvate carboxykinase bundle sheath ( Sb PEPCK.BS, SbiTx430.01G455400) and alanine aminotransferase bundle sheath ( Sb AlaAT.BS, SbiTx430.02G006600), two proteins involved in the PCK C 4 pathway. Characterized germinal edits within the targeted regions upstream of these two genes ranged in size from 138 up to 1790 bp. A 138 bp within the Sb PEPCK.BS upstream region and a 1643 bp element within the Sb AlaAT.BS upstream region were determined to be important for maintenance of transcription levels. No change in development or various physiological parameters was observed in characterized lineages carrying promoter edits. However, significant changes in seed reserves and a reduction in 100‐seed weight were consistently observed, under both greenhouse and field environments, in plants carrying an edit in the promoter of Sb PEPCK.BS gene, which were significantly reduced in transcript accumulation for this gene.
Inorganic nitrogen (N) fertilizer has emerged as one of the key factors driving increased crop yields in the past several decades; however, the overuse of chemical N fertilizer has led to severe ecological and environmental burdens. Understanding how crops respond to N fertilizer has become a central topic in plant science and plant genetics, with the ultimate goal of enhancing N use efficiency (NUE) in crop production. N, being one of the most essential macronutrients, significantly impacts crop performance across different stages of plant development, as the phenotypic traits result from the accumulative effects of genetic factors, environmental conditions (i.e., N availability), and their interactions. To characterize the N response and growth trajectory, we generated sorghum mutants using CRISPR technology. Using a LemnaTec plant imaging system, we obtained time series imagery data for these CRISPR-edited mutants under high N and low N greenhouse conditions. After imagery data analysis, we extracted a number of morphological and greenness index traits as a proxy of plant growth and N responses. Subsequently, we employed two different methods to model the temporal N-responsive traits, allowing us to estimate seven key parameters from the growth curve. Results revealed that the wildtype and the edited sorghum lines exhibited differences in N responses for several of the key growth-related parameters. The high-throughput N phenotyping pipeline paves the way for a better understanding of the N responses of edited lines in a dynamic manner and sheds light on further improvements in crop NUE. ### Competing Interest Statement J.C.S. has equity interests in Data2Bio, LLC; Dryland Genetics LLC and EnGeniousAg LLC. He is a member of the scientific advisory board of GeneSeek and currently serves as a guest editor for The Plant Cell. The authors declare no other conflicts of interest associated with this work.
The timing of flowering is determined by a complex genetic architecture integrating signals from a diverse set of external and internal stimuli and plays a key role in determining plant fitness and adaptation. However, significant divergence in the identities and functions of many flowering time pathway components has been reported among plant species. Here, we employ a combination of genome and transcriptome wide association studies to identify genetic determinants of variation in flowering time across multiple environments in a large panel of primarily photoperiod-insensitive sorghum (Sorghum bicolor), a major crop that has, to date, been the subject of substantially less genetic investigation than its relatives. Gene families that form core components of the flowering time pathway in other species, FT-like and SOC1-like genes, appear to play similar roles in sorghum, but the genes identified are not orthologous to the primary FT-like or SOC1-like genes that play similar roles in related species. The ageing pathway appears to play a role in determining non-photoperiod determined variation in flowering time in sorghum. Two components of this pathway were identified in a transcriptome wide association study, while a third was identified via genome-wide association. Our results demonstrate that while the functions of larger gene families are conserved, functional data from even closely related species is not a reliable guide to which gene copies will play roles in determining natural variation in flowering time.
Societal Impact StatementMaize plays a key role in agricultural profitability and food security on six continents. Successful efforts to breed higher‐yielding maize varieties depend on replicated yield trials in many environments. Capturing in‐season data can help improve and accelerate the development of regionally adapted hybrids, but collecting these data can be impractical in many locations. We demonstrate that satellite remote sensing could play a similar role in crop performance assessment as unmanned aerial vehicles (UAVs) but with far lower labor costs and the greatest ease of collection at remote field sites. This dataset and benchmarks have the potential to enable predictive models that could guide farmers and crop breeders in decision‐making.Summary Accurate early yield estimates at fields and plots offer potential benefits to farmers in optimizing their agronomic practices and breeders in screening thousands of varieties contributing to improving agriculture and food production systems. Effective approaches to track plant growth and predict yield require large datasets of remote sensing and ground truth data collected across multiple environments. Low‐altitude drone flights are increasingly being used to collect data from field evaluations of new crop varieties, while satellite imagery is being explored to track yield and management practices at regional scales. Satellite platforms exhibit logistical and technical advantages in scalability and accessibility and could facilitate plot‐level predictions, especially with steadily improving spatial resolution. However, plot‐level, high‐resolution satellite images capturing differences in genotypes from multiple environments with ground truth measurements are not publicly available. Here, we generated, described, and evaluated over 20,000 plot‐level images of over 80 hybrid maize varieties grown across the US corn belt under various management practices collected from (near simultaneous) satellite and drone (synonym UAVs, UASs) flights integrated with ground truth yield measurement. Of the six baseline models examined, models employing data collected from satellite images often matched the performance of models employing drone images for both within and cross‐environment yield prediction.