Crop adaptation to the mixture of environments that defines the target population of environments is the result of balanced resource allocation between roots, shoots, and reproductive organs. Root growth plays a critical role in the determination of this delicate balance. The responses of root growth and function to temperature can determine the strength of roots as sinks but also influence a crop's ability to uptake water and nutrients. Surprisingly, this behavior has not been studied in maize (Zea mays) since the middle of the last century, and the genetic determinants are unknown. Low temperatures recorded frequently in deep soil layers limit root growth and soil exploration and may constitute a bottleneck for increasing drought tolerance, nitrogen recovery, sequestration of carbon, and productivity in maize. We developed high-throughput phenotyping systems to investigate these responses and to examine genetic variability therein across diverse maize germplasm. Here, we show that there is (i) genetic variation in root growth under low temperature below a previously set threshold of 10 °C and (ii) genotypic variation in water transport under low temperature. The trait set examined herein and the high-throughput phenotyping platform developed for its characterization provide a unique opportunity for removing a major bottleneck for crop improvement and adaptation to climate change.
The modern era of crop improvement is characterized by two distinct paths: (i) biology-driven molecular genetics focused on dissecting the mechanisms underlying complex traits by reductionist approach, and (ii) data-driven quantitative genetics, which prioritizes accurate prediction of final trait performance, such as yield, without necessarily clarifying the underlying biological mechanisms. These approaches, one rooted in functional discovery and the other in statistical association, have each made significant strides, yet with limited overlap. However, their conjunction represents an emerging opportunity; integrating the functional insights from pangenetics, encompassing causal variants and haplotype diversity into artificial intelligence (AI) based predictive frameworks, can offer a more holistic perspective for customizing and testing crop genomic ideotypes for specific agroecologies, dynamic market demands, and increasing climate uncertainties.
Accurate selection of favourable crop genotypes has motivated the exploration of diverse prediction algorithms for crop breeding applications. One genomic prediction method that has not been fully explored is graph attention networks (GAT). By directly analysing graphical data with the attention mechanism, GAT can incorporate the genotype-to-phenotype (G2P) structure to regularise predictions. As one potential G2P structure, a gene network can be inferred from interpretable machine learning models to effectively learn key features of prediction patterns, potentially improving prediction performance. Here, we investigated whether incorporating such data-driven prior knowledge into GAT improved prediction performance compared to GAT models representing a continuum of G2P structures, ranging from infinitesimal to fully connected. Applying the Diversity Prediction Theorem, we also combined these diverse G2P structures into an ensemble of GAT genomic prediction models to integrate complementary strengths of multiple models. The results for flowering time traits in two maize nested association mapping datasets showed a lack of consistent performance improvement in the data-driven prior knowledge GAT model. However, consistent outperformance was observed for the ensemble of GAT models. Improved predictions from the ensemble model may be driven by its ability to capture a more complete representation of the inferred gene network through the integration of information from diverse G2P structures. The observed results using the GAT methodology provided the foundation for potential performance improvement using GAT by integrating biological prior knowledge derived from omics data and empirically verified gene interactions in future research, thereby potentially enhancing the GAT ensemble performance.
Plant breeding operates within a complex genetic landscape determined by genes interacting within biological networks and with the environment. This environment is not constant but subject to short-term fluctuations and long-term shifts. This makes finding a balance between adapting germplasm for short- and long-term objectives challenging. We previously investigated the implications of genetic complexity on breeding program design. Here, we build on this work by adding an environmental dimension in the form of the E(NK) model to the simulation framework. We found that the addition of environmental interactivity and change creates greater uncertainty associated with pursuing any specific selection trajectory. This advantages preserving genetic variability and genetic landscape exploration over quickly exposing additive variation by constraining genetic space around a particular and temporary local optimum. Nonetheless, also in a dynamically changing environment, a distributed breeding program structure finds the best balance between short- and long-term objectives. In this structure, several breeding programs explore genetic space while maintaining constant germplasm exchange. This is in contrast to isolated programs or one large undifferentiated program, which exclusively emphasize short respectively long-term objectives. We furthermore highlight the difficulty of exchanging germplasm to restore genetic variability with nonstationary and germplasm context dependent genetic effects. In summary, also under environmental complexity and change, the structural features that characterized breeding operations hitherto and allowed them to navigating biological complexity apply. Namely, the necessity to constrain genetic space in order for heritable additive variation to emerge. We end by arguing that optimal breeding program design depends on the level of genetic and environmental complexity. This complexity should be appropriately reflected when modeling the long-term behavior of selection programs and the implications of specific interventions into these.
Ensembles of multiple genomic prediction models have demonstrated improved prediction performance over the individual models contributing to the ensemble. The outperformance of ensemble models is expected from the Diversity Prediction Theorem, which states that for ensembles constructed with diverse prediction models, the ensemble prediction error becomes lower than the mean prediction error of the individual models. While a na & iuml;ve ensemble-average model provides baseline performance improvement by aggregating all individual prediction models with equal weights, optimizing weights for each individual model could further enhance ensemble prediction performance. The weights can be optimized based on their level of informativeness regarding prediction error and diversity. Here, we evaluated weighted ensemble-average models with three possible weight optimization approaches (linear transformation, Nelder-Mead and Bayesian) using flowering time and tillering traits from two maize nested associated mapping (NAM) datasets: TeoNAM and MaizeNAM. The three proposed weighted ensemble-average approaches improved prediction performance in several of the prediction scenarios investigated. In particular, the weighted ensemble models enhanced prediction performance when the adjusted weights differed substantially from the equal weights used by the na & iuml;ve ensemble models. For performance comparisons among the weighted ensembles, there was no clear superiority among the proposed approaches in both prediction accuracy and error across the prediction scenarios. Weight optimization for ensembles warrants further investigation to explore the opportunities to improve their prediction performance; for example, integration of a weighted ensemble with a simultaneous hyperparameter tuning process may offer a promising direction for further research.
Thousands of years of breeding and agronomic change have pushed genetic change into increasingly narrow corridors. The approach has worked effectively, but gains are slowing and much genetic variance expected from heritability estimates is not readily found. We propose that breeding constraints, the practical demands that make crops useful, force genetic exploration onto curved, lower-dimensional surfaces within the larger landscape of possibility. In crops, these constraints arise jointly from genetic choices and management decisions about fertilizer, sowing density, weed and pest control, irrigation and harvest logistics, which together define the range of environments in which genotypes are routinely grown. These constraints may hide certain classes of gene interactions from breeding programmes, creating an additive appearance where the underlying biology may remain epistatic. Selection moves by additive steps, but on a curved, constrained surface: the speed looks additive, but the long-term path follows the geometry of the constraints. This framework particularly applies to major crop species subjected to intensive directional selection over many generations for stable agricultural requirements, the context where constraint-based filtering of genetic variance is most pronounced. When selection has had less opportunity to impose its effects and gene interactions are weak, this geometric filtering may be less consequential. But when epistatic effects shape fitness, constraints could hide substantial genetic potential behind these boundaries. This Perspective suggests potential escape routes from current plateaus: strategic wide crosses, transgene combinations and targeted edits that access genetic variance currently excluded by constraints.
While various genomic prediction models have been evaluated for their potential to accelerate genetic gain for multiple traits, no individual genomic prediction model has outperformed all others across all applications. As an alternative approach, ensembles of multiple individual genomic prediction models can be applied to utilize the complementary strengths of individual prediction models and offset the prediction errors of each. We used the EasiGP (Ensemble AnalySis with Interpretable Genomic Prediction) pipeline to investigate the performance of an ensemble approach, targeting flowering-time traits measured in 2 maize nested association mapping datasets. For both datasets, the ensemble-based prediction approach achieved higher prediction accuracy and lower prediction error across the flowering-time traits compared to each individual model. Multiple genomic regions known to contain key flowering-time-related genes were repeatedly included as features across individual genomic prediction models, indicating the models successfully captured SNPs as features that are associated with genomic regions known to contain flowering-time genes. Although repeatability was high for some genomic regions, estimated marker effects varied across many genomic regions, suggesting that the models might also have captured different aspects of the genetic variation underlying the traits. The ensemble combination of the diverse views likely contributed to the improvement of prediction performance by the ensemble-based approach over the individual prediction models. Ensemble-based prediction can be applied to overcome limitations observed in the continuous exploration for the best individual genomic prediction models that can consistently achieve the highest prediction performance, thereby potentially contributing to improved prediction accuracy for applications in crop breeding.
Lodging in sorghum presents a significant challenge for plant breeders due to the trade-off between lodging resistance and grain yield. Manually measuring lodging across thousands of plots is time-consuming, expensive, and error-prone, making selection for lodging resistance challenging in breeding programs. Unmanned aerial vehicle (UAV)-derived metrics provide a potential high-throughput alternative; however, it remains unclear whether photogrammetric heights derived from UAV imagery can estimate plot-level lodging severity in large sorghum breeding trials. This study developed a framework for predicting plot-level lodging from UAV imagery across 2,675 sorghum breeding plots. Multi-temporal canopy height data were collected at two critical time points: maximum crop height and at manual lodging assessment. Height percentiles were extracted from UAV-derived point clouds generated using photogrammetric algorithms. These data were used to develop parametric, non-parametric, and ensemble prediction models, which were evaluated using three statistical metrics. The ensemble model, averaging predictions from all models, achieved the highest accuracy with Pearson correlations of r = 0.80-0.84 and lowest root mean square error (RMSE=16-18%), explaining 64-70% of variation in manual lodging counts. Model diagnostics and iterative refinement, including inspection of UAV imagery and dataset curation, had minimal impact on model performance, demonstrating the robustness of the approach. Model performance was consistent across sites, with minimal effects of stratified sampling on accuracy, confirming the ensemble approach as optimal for plot-level lodging assessment. This study demonstrates that integrated multi-temporal UAV imagery offers a practical alternative to labor-intensive manual evaluation methods by enabling high-throughput lodging assessment suitable for implementation in sorghum breeding programs.
Abstract Water deficit is ubiquitous in maize ( Zea mays L.) cropping systems worldwide. Ethylene insensitivity in maize has been implicated in improving kernel set and yield under drought, and ARGOS genes modulate ethylene signal transduction by reducing ethylene sensitivity. Given the strong water sensitivity of silk elongation, ARGOS8 overexpression is expected to alter silk growth responses to drought. Experiments were conducted under controlled and field conditions to test the effect of ARGOS8 gene overexpression on silk growth under water deficit. Silk lengths and water use were continuously monitored, and silk elongation rate (SER) response to the fraction of transpirable soil water (FTSW) was evaluated. Silk emergence dynamics were measured in the field by daily counting the silks under contrasting water regimes. ARGOS8 transgenics maintained SER at lower FTSW than controls; however, responses varied among hybrids. Higher SER under water stress translated into a faster silk exertion rate in ARGOS8 transgenics than controls (45 vs. 25 silks d -1 ), leading to a greater number of exerted silks at three days post-silking (424 vs. 377, ∼90% vs. 84% of total silks, p < 0.01). Together, these results help clarify the mechanism underlying the ectopically expressed ARGOS8 effect on maize yield improvement under water stress. Highlight ARGOS8 transgenic expression sustains silk elongation rates and increases silk emergence under water deficit, improving the reproductive performance of maize in water-limited environments.
Climate-driven variability is reducing our ability to accurately predict crop performance across environments, limiting genetic gain in breeding programs. Sustained progress requires predictive frameworks that capture plant-environment interactions across diverse genetics and management conditions. Integrating mechanistic insights from plant science into predictive models offers a path to improve the accuracy, precision, and interpretability of breeding decisions under changing environments. We present emerging hierarchical genome-phenome frameworks and outline how they can be leveraged within breeding programs to evaluate how biological knowledge informs predictions across target environments and supports long-term genetic gain.
Ensembles of multiple genomic prediction models have demonstrated improved prediction performance over the individual models contributing to the ensemble. The outperformance of ensemble models is expected from the Diversity Prediction Theorem, which states that for ensembles constructed with diverse prediction models, the ensemble prediction error becomes lower than the mean prediction error of the individual models. While a naïve ensemble-average model provides baseline performance improvement by aggregating all individual prediction models with equal weights, optimising weights for each individual model could further enhance ensemble prediction performance. The weights can be optimised based on their level of informativeness regarding prediction error and diversity. Here, we evaluated weighted ensemble-average models with three possible weight optimisation approaches (linear transformation, Nelder-Mead and Bayesian) using flowering time traits from two maize nested associated mapping (NAM) datasets; TeoNAM and MaizeNAM. The three proposed weighted ensemble-average approaches improved prediction performance in several of the prediction scenarios investigated. In particular, the weighted ensemble models enhanced prediction performance when the adjusted weights differed substantially from the equal weights used by the naïve ensemble models. For performance comparisons within the weighted ensembles, there was no clear superiority among the proposed approaches in both prediction accuracy and error across the prediction scenarios. Weight optimisation in ensembles warrants further investigation to explore the opportunities to improve their prediction performance; for example, integration of a weighted ensemble with a simultaneous hyperparameter tuning process may offer a promising direction for further research. ### Competing Interest Statement The authors have declared no competing interest. The Australian Research Council Centre of Excellence for Plant Success in Nature and Agriculture, CE200100015
Abstract Additive models of inheritance predict the short-term response to selection remarkably well, even when the underlying biology involves widespread dominance and gene–gene interaction. We argue that this success reflects a property of how fitness varies with the additive genetic component of a trait, not a property of the molecular architecture beneath. We make this quantitative through an additivity index, A g , that measures the fraction of local log-fitness variance captured by a linear approximation, with the remainder attributable to local curvature of the fitness surface. Under Gaussian/quadratic assumptions, A g equals the squared correlation between the linear approximation and local log-fitness, so 1− A g is the fraction of local log-fitness variance not captured by a purely linear predictor. We call regions of breeding-value space where A g is high additive channels and develop a coupled selection–inheritance framework that identifies when populations enter, persist in, and leave them. The framework predicts that stable, well-adapted populations can be those for which additive prediction is least informative for directional-response prediction. Article summary Gene interactions are common, yet additive genetic models often predict short-term evolution. We propose that additivity is a local property of where populations sit in breeding-value space on a curved fitness landscape, not a global property of the biology beneath. We define an index, A g , that compares slope variance to curvature variance and identifies populations in additive channels for which linear prediction performs well. The framework predicts that elite breeding pools enter such channels under sustained selection, while natural populations near a fitness optimum may leave them because the directional signal collapses.
The Big Breeding Innovation Team (Big BIT) maize (Zea mays L.) experiment was one of the largest genomic data-informed predictive breeding validation studies ever conducted. The experiment was a multi-location, multi-year, multi-tester, multi-population study involving F1 maize hybrids created by crossing individual doubled haploids to inbred testers. The purpose of the study, performed by DuPont Pioneer/Corteva Agriscience in 2017, 2018, and 2019, was to build comprehensive datasets to help answer a wide range of practical questions focused on optimizing predictive breeding strategies in maize. The purpose of our study is to (1) describe the design and unique features of our study and (2) discuss learnings with practical implications for plant breeders. Since the same F1 maize hybrids were grown across three distinct years, we use basic descriptive summary statistics to discuss our learnings. We provide a technical justification for the use of basic statistics and discuss the expected theoretical prediction accuracy of genomic estimated breeding values (GEBVs) of Big BIT individuals and families, and predictive abilities obtained by performing large-scale cross-validations. Our study provides multi-year field data-based evidence that, for inbred/variety development focused plant improvement efforts, early-stage genetic evaluation should be based on GEBVs generated from wide-area testing training datasets. This holds true for candidates for selection with or without own phenotypic records.
An ensemble of multiple genomic prediction models has grown in popularity due to consistent prediction performance improvements in crop breeding. However, technical tools that analyze the predictive behavior at the genome level are lacking. Here, we develop a computational tool called Ensemble AnalySis with Interpretable Genomic Prediction (EasiGP) that uses circos plots to visualize how different genomic prediction models quantify contributions of marker effects to trait phenotypes. As a demonstration of EasiGP, multiple genomic prediction models, spanning conventional statistical and machine learning algorithms, were used to infer the genetic architecture of days to anthesis (DTA) in a maize mapping population. The results indicate that genomic prediction models can capture different views of trait genetic architecture, even when their overall profiles of prediction accuracy are similar. Combinations of diverse views of the genetic architecture for the DTA trait in the teosinte nested association mapping study might explain the improved prediction performance achieved by ensembles, aligned with the implication of the Diversity Prediction Theorem. In addition to identifying well-known genomic regions contributing to the genetic architecture of DTA in maize, the ensemble of genomic prediction models highlighted several new genomic regions that have not been previously reported for DTA. Finally, different views of trait genetic architecture were observed across subpopulations, highlighting challenges for between-population genomic prediction. A deeper understanding of genomic prediction models with enhanced interpretability using EasiGP can reveal several critical findings at the genome level from the inferred genetic architecture, providing insights into the improvement of genomic prediction for crop breeding programs.
Trait Genome-to-Phenome (G2P) dimensionality and “breeding context” combine to influence the realised prediction skill of different whole genome prediction (WGP) methods. Theory and empirical evidence both suggest there is likely to be “No Free Lunch” for prediction-based breeding. Ensembles of diverse sets of G2P models provide a framework to expose and investigate the high G2P dimensionality of trait genetic architecture for WGP applications. Artificial Intelligence and Machine Learning (AI-ML) prediction algorithms contribute novel trait G2P model diversity to ensemble-based WGP. Prediction-based breeding leveraging ensembles of G2P models creates new opportunities to identify and design novel paths for genetic gain. Improving our understanding of trait genetic architecture is motivated by creating new opportunities to enhance breeding methodology, create new selection trajectories for crop improvement, and accelerate rates of genetic gain. With access to high-throughput sequencing, phenotyping and envirotyping technologies we can model the complex multidimensional relationships between sequence variation and trait phenotypic variation that are under the influences of selection. Using the framework of the diversity prediction theorem, we consider applications of ensembles of diverse trait genome-to-phenome (G2P) models. Crop growth models (CGM) are an example of a hierarchical framework for studying the influences of quantitative trait loci (QTL) within trait networks and their interactions with different environments to determine yield. Hybrid CGM-G2P models combine elements of CGMs, to understand how trait networks influence crop yield performance, with trait G2P models, to understand influences of trait genetic architecture on selection trajectories. We discuss hybrid CGM-G2P models and their potential applications to enhance ensemble-based prediction. Multi-environment trials conducted across breeding cycles can be designed to include contrasting environments to expose the different CGM-G2P dimensions of the trait by environment interactions that are influential on selection trajectories. Artificial intelligence and machine learning (AI-ML) algorithms can be applied as components of ensembles to improve gene discovery and quantification of allele effects for traits to enhance G2P prediction applications. We use the trait flowering time in the maize TeoNAM experiment to illustrate and motivate further investigations of how to leverage ensembles of G2P models for prediction-based breeding.
Plant breeding operates within a highly complex genetic landscape determined by gene effects emerging through biological networks and their interactions with the environment. This environment is not constant but subject to short-term fluctuations and long-term shifts. This significantly complicates the task of plant breeders in finding a balance between adapting their germplasm for the short- and long-term. Here we build on previous work of us that investigated the implications of genetic complexity on breeding program design, by adding an environmental dimension in the form of the E(NK) model to the simulation framework. We found that the addition of environmental interactivity and change creates greater uncertainty associated with pursuing any specific selection trajectory, as compared to a static environment. This advantages preserving genetic variability and genetic landscape exploration over quickly exposing additive variation by constraining genetic space around a particular and temporary local optimum. Nonetheless, we found that also in a dynamically changing environment, a structure in which several breeding programs explore genetic space while maintaining constant germplasm exchange, finds the best balance between short and long-term objectives, as opposed to isolated programs or one large undifferentiated program, which exclusively emphasize short respectively long-term objectives. We furthermore highlight the difficulty of exchanging germplasm to restore genetic variability with non-stationary and germplasm context dependent genetic effects. In summary we found that also with addition of environmental complexity and change, the structural features that characterized breeding operations hitherto and allowed them to navigating biological complexity apply. Namely the necessity to constraining genetic space in order for heritable additive variation to emerge. We end by arguing that optimal breeding program design depends on the level of genetic and environmental complexity, which should be appropriately reflected when modeling the long-term behavior of selection programs and the implications of specific interventions into these. ### Competing Interest Statement Frank Technow and Dean Podlich were employed by the company Corteva Agriscience. Mark Cooper was supported by the Australian Research Council Centre of Excellence for Plant Success in Nature and Agriculture (CE200100015).
Better understanding genotype by environment interaction (GxE) can help breeding for better adapted varieties. Envirotyping for environmental water status was applied to assist interpretation of GxE interactions for wheat yield in multi-environment trials conducted in drought-prone Australian environments. Genotypes from a multi-reference parent nested association mapping (MR-NAM) population were tested in 10 trials across the Australian wheatbelt. Genotype yield and phenology were measured in all trials, while traits associated with the stay-green phenotype were assessed for a subset of 5 trials. Envirotyping was conducted by characterizing water stress experienced by genotypes at each trial using crop modelling. Envirotyping facilitated the understanding of GxE interactions by explaining 75, 67, and 66 % of the genotypic variance for yield in severe water-limited (ET3), mild terminal water-stress (ET2), and water-sufficient (ET1) environments, respectively. Yield and stay-green were negatively correlated with flowering time in most trials. However, when focusing on genotypes flowering at similar times within a trial, no significant correlation was found between yield and flowering. Importantly stay-green traits remained significantly correlated with yield. Stay-green traits such as delayed onset of senescence and slower senescence rate benefited yield by 0.2-1.1 t ha-1 across environments, highlighting the breeding potential for stay-green traits in both water-sufficient and water-limited environments. Hence, sustaining green leaf area during grain filling helped to enhance yield. Envirotyping to better understand GxE interactions for yield, coupled with screening for traits exhibiting superior adaptive mechanisms, are powerful assets in assisting plant breeders to select more effectively drought adapted genotypes.
Context: Australia contributes similar to 3 % of the global grain sorghum production despite relatively large gaps between water-limited potential (PYw) and attainable on-farm (AYw) yields. With sorghum yield gaps typically 59 % of the PYw in Australia, it is important to identify the drivers of these yield gaps and what are the optimal genotype (G) x environment (E) x management (M) combinations needed for sustainable improvement of sorghum productivity. Objectives: APSIM farming systems modelling framework was used to simulate comprehensive G x E x M scenarios for sorghum in Australia using modern hybrids and current management to (1) define the PYw fronts and identify key drivers of spatial sorghum yield variability, and (2) determine desirable G x M combinations that can improve sorghum productivity and stability across the expected range of seasonal evapotranspiration (ET) for the Australian Target Population of Environments (TPE). Methods: We use a set of 70 comprehensive national variety trials (NVT) data from 2017 - 2021 to parameterize and evaluate APSIM and then subsequently apply the model to define the range of G x E x M dimensions for sorghum TPEs in Australia. From these, we systematically simulate 93 sites-soil combinations for 103 years to extend the NVT to other fields beyond the NVT sites directly sampled. Using yield frontier analysis, we then define the expected PYw fronts, yield variability and the drivers of this variability. Results: A non-linear relationship was observed between simulated grain yields and seasonal ET with most yields between 2.6 and 4.5 t ha(-1) associated with an ET of 167 - 420 mm. The yield opportunity frontiers were estimated to fall between 9.4 t ha(-1) (Q80 %) and 12.0 t ha(-1) (Q99 %). Nitrogen application rate at sowing explained the greatest component of the grain yield variation across most subregions. We found that most crop failures (defined as yield at or below Q10 %; 1.1 t ha(-1) ) occurred under low N and high plant density with greater risks in North-East and North-Central Queensland. Yield failure risks were higher at higher densities, while earlier sowing buffered against these crop failures. Conclusions: This study estimated PYw fronts and yield gap distributions for sorghum, elucidating the drivers of spatial sorghum yield variability for the Australian sorghum TPE. This finding highlights potential opportunities to close the on-farm yield gaps through optimized region-specific agronomic recommendations.
While many genomic prediction models have been evaluated for their potential to accelerate genetic gain for multiple traits, no individual genomic prediction model has outperformed others across all applications. This problem aligns with the implications of the No Free Lunch Theorem, stating that the average performance of individual prediction models becomes equivalent across the state space of diverse prediction scenarios. Ensembles of multiple individual genomic prediction models can be a potential alternative approach. The framework of the Diversity Prediction Theorem suggests the potential for a reduction in prediction error with the inclusion of diverse prediction models in the ensemble. We investigated the performance of an ensemble approach that combines multiple genomic prediction models. We demonstrate the results using flowering time traits measured in two maize Nested Association Mapping datasets. For both datasets, the ensemble-based prediction approach achieved the highest prediction accuracy and lowest prediction error across traits. Multiple genomic regions containing key flowering time-related genes were captured by the different genomic prediction models with diverse weights, demonstrating different views of the trait genetic architecture. The combination of such diverse views contributed to the improvement of prediction performance by the ensemble-based approach over the individual prediction models. Exploiting the expectations of the Diversity Prediction Theorem, the ensemble can overcome some limitations proposed by the No Free Lunch Theorem when applying individual genomic prediction models. Key message Applying the Diversity Prediction Theorem, an ensemble-based prediction leveraging multiple individual genomic prediction models improved the prediction performance over the individual models by combining multiple views of trait genetic architecture ### Competing Interest Statement The authors have declared no competing interest. The Australian Research Council Centre of Excellence for Plant Success in Nature and Agriculture, CE200100015
The modern era of crop improvement is characterized by two distinct paths: (i) biology-driven molecular genetics focused on dissecting the mechanisms underlying complex traits by reductionist approach, and (ii) data-driven quantitative genetics, which prioritizes accurate prediction of final trait performance, such as yield, without necessarily clarifying the underlying biological mechanisms. These approaches, one rooted in functional discovery and the other in statistical association, have each made significant strides, yet with limited overlap. However, their conjunction represents an untapped opportunity; integrating the functional insights from pangenetics, encompassing causal variants and haplotype diversity into artificial intelligence (AI) based predictive frameworks can offer a more holistic perspective for customizing and testing crop genomic ideotypes for specific agroecologies, dynamic market demands, and increasing climate uncertainties.