ABSTRACT Chromosomal inversions facilitate local adaptation by maintaining co-adapted allele complexes as a single inherited unit, but the genes driving their phenotypic effects are rarely identified. Inv4m is an inversion from the highland teosinte Zea mays ssp. mexicana that is nearly fixed at 2500 masl in Mexican traditional varieties but absent in temperate maize. While association studies link Inv4m to flowering time, its key adaptive trait, the mechanisms connecting it to this phenotype remain unknown. Here, we generate near isogenic lines by introgressing the highland Michoacán 21 (Mi21) Inv4m haplotype into temperate B73 through eight backcross generations. We assemble the Inv4m -Mi21 NIL genome and, aligning to three Zea reference genomes, provide the first precise delimitation of Inv4m , a 15.2 Mb region with breakpoints overlapping knob repeat arrays. In field trials across Pennsylvania and North Carolina, Inv4m consistently accelerated flowering, whereas its plant height effect reversed sign between regions, a genotype by environment interaction consistent with environment-dependent fitness. Using RNA-seq, we identify 465 differentially expressed genes, the strongest from a cluster of JUMONJI histone demethylases (JMJ). Comparing five Zea genomes, the B73 reference carries a lineage-specific tandem expansion of five JMJ paralogs while highland genotypes carrying Inv4m retain a single ancestral copy, a difference that accounts for most of the cluster’s differential expression. We then show Inv4m disrupts growth related coexpression modules, with the JMJ cluster losing connectivity, and traced a trans regulatory network linking it to cell proliferation genes including pcna2 and sec6 . In summary, we identify candidate genes and networks underlying Inv4m ’s effects and propose that part of this inversion’s adaptive value may reside in a dosage-sensitive regulator whose action propagates through a downstream network of genes that are involved in growth and flowering and are important for highland adaptation. Recombination suppression may thus protect a co-adapted regulatory architecture rather than independent alleles alone. Graphical Abstract
ABSTRACT Local adaptation of a species involves the selection of adaptive alleles that confer a fitness advantage in their local environment. Inversions prevent recombination between the standard and inverted heterozygous hybrids. Inversions can play a crucial role in local adaptation by locking together a set of co-adapted alleles, acting as supergenes. Inv4m is a 13 Mb inversion in maize prevalent in highland maize and highland wild relatives from México. Maize from the highlands of the Trans-Mexican volcanic belt has been shown to be well-adapted to volcanic, acidic soils with low phosphorus availability. Inv4m carries several genes involved in P acquisition and utilization. We therefore tested the hypothesis that Inv4m contributes to maize adaptation to these environments through enhanced phosphorus acquisition or utilization. Alternatively, Inv4m possible adaptive value may operate through constitutive developmental effects independent of nutrient stress responses. To test this hypothesis, we introgressed a highland maize variety from the highlands of Michoacán, México, carrying Inv4m into the temperate line B73 and developed Near-Introgression Lines (NILs) carrying Inv4m . We then grew NILs carrying the inversion and controls without it in soils with different phosphorus levels and evaluated the fitness effects of the inversion, as well as changes in gene expression using RNA-Seq. Our results show that P starvation elicits highly conserved transcriptomic, lipidomic, and ionomic responses, independently of the Inv4m inversion genotype. Therefore, phosphorus deficiency does not seem to be driving the adaptive value of Inv4m . Additionally, we observed a phosphorus modulated transcriptional gradient from the collar leaf downward, characterized by a decrease in the expression of photosynthesis genes and an increase in the expression of senescence-associated genes, corresponding to the positional onset and initial stages of sequential leaf senescence. Although the magnitude of the phosphorus response increased with leaf age, we did not observe significant interactions with Inv4m . Our multi-omics analysis of the maize phosphorus starvation response identified and characterized two coordinately regulated molecular programs, light harvesting shutdown and accelerated senescence, whose deployment depends on leaf developmental stage, with older leaves below the collar integrating nutrient limitation into the natural progression toward senescence.
Integrating innovative technologies into plant breeding is critical to bolster food and nutritional security under biotic and abiotic stresses in changing climates. While breeding efforts have focused primarily on yield and stress tolerance, emerging evidence highlights the need to also prioritize nutritional quality. Advanced molecular breeding approaches have enhanced our ability to develop improved crop varieties and could be substantially informed by the routine integration of crop modeling and remote sensing technologies. This review article discusses the potential of combining crop modeling and sensing with molecular breeding to address the dual challenge of nutritional quality and stress tolerance. We provide overviews of stress response strategies, challenges in breeding for quality traits, and the use of environmental data in genomic prediction. We also describe the status of crop modeling and sensing technologies in grain legumes, rice, and leafy greens, alongside the status of -omics tools in these crops and the use of AI with directed evolution to identify novel resistance genes. We describe the pairwise and three-way integration of AI-enabled sensing and biophysically and empirically constrained crop modeling into breeding to enable prediction of phenotypic and breeding values and dissection of genotype-by-environment-by-management interactions with increasing fidelity, efficiency, and temporal/spatial resolution to inform selection decisions. This article highlights current initiatives and future trends that focus on leveraging these advancements to develop more climate-resilient and nutritionally dense crops, ultimately enhancing the effectiveness of molecular breeding.
Phenomic Selection is a new paradigm in plant breeding that uses high-throughput phenotyping technologies and machine learning models to predict traits of new individuals and make selections. This can allow breeders to evaluate more plants in higher throughput more accurately, resulting in faster rates of gain and reduced labor costs. However, Phenomic Prediction models are frequently benchmarked against Genomic Prediction models using cross-validation to demonstrate their usefulness to breeders. We argue that this is inappropriate for two reasons: 1) Differences in the accuracy statistic measured by cross-validation do not reliably indicate differences in the accuracy parameter of the breeder's equation, and 2) Accuracy alone is insufficient to compare breeding schemes using Phenomic vs. Genomic Prediction because these tools differentially influence other parameters of the breeder's equation. We show analytically and through re-analysis of data from three representative Phenomic Prediction studies that conclusions about the superiority of Phenomic Prediction over Genomic Prediction change if compared using consistent methods. We conclude that Phenotypic Selection may be useful, but comparisons of accuracy between Genomic Prediction and Phenotypic Prediction models are not. ### Competing Interest Statement The authors have declared no competing interest.
Phenomic selection is a new paradigm in plant breeding that uses high‐throughput phenotyping technologies and machine learning models to predict traits of new individuals and make selections. This can allow breeders to evaluate more plants in higher throughput more accurately, resulting in faster rates of gain and reduced labor costs. However, phenomic prediction models are frequently benchmarked against genomic prediction models using cross‐validation to demonstrate their usefulness to breeders. We argue that this is inappropriate for two reasons: (1) differences in the accuracy statistic measured by cross‐validation do not reliably indicate differences in the accuracy parameter of the breeder's equation, which we show analytically and through reanalysis of data from three representative phenomic prediction studies and (2) phenomic and genomic selection tools influence other parameters of the breeder's equation, so comparing accuracy, even if done properly, is insufficient to advocate for one approach over the other. We conclude that phenomic selection may be useful, but comparisons of accuracy between genomic prediction and phenomic prediction models are not.
With growing evidence that genomic selection (GS) improves genetic gains in plant breeding, it is timely to review the key factors that improve its efficiency. In this feature review, we focus on the statistical machine learning (ML) methods and software that are democratizing GS methodology. We outline the principles of genomic-enabled prediction and discuss how statistical ML tools enhance GS efficiency with big data. Additionally, we examine various statistical ML tools developed in recent years for predicting traits across continuous, binary, categorical, and count phenotypes. We highlight the unique advantages of deep learning (DL) models used in genomic prediction (GP). Finally, we review software developed to democratize the use of GP models and recent data management tools that support the adoption of GS methodology.
Abstract Mutations fuel evolution while also causing diseases like cancer. Epigenome-targeted DNA repair can help organisms protect important genomic regions from mutation. However, the adaptive value, mechanistic diversity, and evolution of epigenome-targeted DNA repair systems across the tree of life remain unresolved. Here, we investigated the evolution of histone reader domains fused to the DNA repair protein MSH6 (MutS Homolog 6) across over 4,000 eukaryotes. We uncovered a paradigmatic example of convergent evolution: MSH6 has independently acquired distinct histone reader domains; PWWP (metazoa) and Tudor (plants), previously shown to target histone modifications in active genes in humans (H3K36me3) and Arabidopsis (H3K4me1). Conservation in MSH6 histone reader domains shows signatures of natural selection, particularly for amino acids that bind specific histone modifications. Species that have gained or retained MSH6 histone readers tend to have larger genome sizes, especially marked by significantly more introns in genic regions. These patterns support previous theoretical predictions about the co-evolution of genome architectures and mutation rate heterogeneity. The evolution of epigenome-targeted DNA repair has implications for genome evolution, health, and the mutational origins of genetic diversity across the tree of life.
Polyploidy is ubiquitous across North American prairies, which provide essential ecosystem services and rich soil for agriculture. Yet the mechanism driving polyploid abundance is unclear. Multiple hypotheses have been proposed including polyploid abundance is proportional to the opportunity for whole genome duplication (WGD), and WGD alters phenotypes that may increase fitness. We tested these two hypotheses together in the mixed-ploidy species Andropogon gerardi , a dominant grass species in endangered North American tallgrass prairies. Leveraging a novel, phased allopolyploid reference genome, we found the A. gerardi hexaploid arose after the C 4 grassland expansion in the early Pleistocene, when glacial cycles likely increased secondary contact between the diploid progenitors. We sequenced A. gerardi from 25 popula-tions and examined cytotype performance and morphology in a controlled environment to investigate the consequences of the contemporary mixed-ploidy populations. We found the 9 x A. gerardi cytotype is a neopolyploid and a result of recurrent WGD events. Further, we demonstrate the 9 x neopolyploids have greater growth and a decreased stomatal pore index, which is adaptive in xeric climates where the 9 x cy-totype is most common. Together, our results support both hypotheses for polyploid abundance in North America: WGD is a product of opportunity and can have immediate fitness consequences. Although the changes to fitness may provide an advantage to 9 x A. gerardi , the establishment of 9 x may lower overall population fitness due to the lower reproductive viability of 9 x individuals. Polyploid species are abundant in North American prairies and make up many of the dominant species in the ecosystem. This prominence could be a result of whole genome duplication conferring an advantage that increases the frequency of polyploids or could simply indicate that the opportunity for whole genome duplication is higher in this ecosystem, or both. Through examining three polyploidization events in A. gerardi , the dominant species in endangered tallgrass North American prairies, we found whole genome duplication is both surprisingly common and confers traits that are beneficial in some environments.
Mate selection plays an important role in breeding programs. The Usefulness Criterion was proposed to improve mate selection, combining information on both the mean and standard deviation of the potential offspring of a cross, particularly in clonally propagated species where large family sizes are possible. Predicting the mean value of a cross is generally easier than predicting the standard deviation, especially in outbred species when the linkage of alleles is unknown and phasing is required. In this study, we developed a method for estimating phasing accuracy from unphased genotype data on possible parental lines and evaluated whether the accuracy was sufficient to predict family standard deviations of possible crosses. We used simulations spanning a wide range of genetic architectures and used genotypes from a real strawberry breeding population to evaluate the conditions when usefulness could be accurately predicted. We found that with highly accurate computational phasing, predicting family standard deviations and usefulness criteria for potential crosses yields benefit over simply selecting crosses based on predicted family means only at high selection intensity and high heritability and with small numbers of QTL. However, even then the gain from using the family usefulness is small.
Predicting phenotypes from a combination of genetic and environmental factors is a grand challenge of modern biology. Slight improvements in this area have the potential to save lives, improve food and fuel security, permit better care of the planet, and create other positive outcomes. In 2022 and 2023 the first open-to-the-public Genomes to Fields (G2F) initiative Genotype by Environment (GxE) prediction competition was held using a large dataset including genomic variation, phenotype and weather measurements and field management notes, gathered by the project over nine years. The competition attracted registrants from around the world with representation from academic, government, industry, and non-profit institutions as well as unaffiliated. These participants came from diverse disciplines include plant science, animal science, breeding, statistics, computational biology and others. Some participants had no formal genetics or plant-related training, and some were just beginning their graduate education. The teams applied varied methods and strategies, providing a wealth of modeling knowledge based on a common dataset. The winners strategy involved two models combining machine learning and traditional breeding tools: one model emphasized environment using features extracted by Random Forest, Ridge Regression and Least-squares, and one focused on genetics. Other high-performing teams methods included quantitative genetics, classical machine learning/deep learning, mechanistic models, and model ensembles. The dataset factors used, such as genetics; weather; and management data, were also diverse, demonstrating that no single model or strategy is far superior to all others within the context of this competition.
Multienvironment trials (METs) are crucial for identifying varieties that perform well across a target population of environments. However, METs are typically too small to sufficiently represent all relevant environment-types, and face challenges from changing environment-types due to climate change. Statistical methods that enable prediction of variety performance for new environments beyond the METs are needed. We recently developed MegaLMM, a statistical model that can leverage hundreds of trials to significantly improve genetic value prediction accuracy within METs. Here, we extend MegaLMM to enable genomic prediction in new environments by learning regressions of latent factor loadings on Environmental Covariates (ECs) across trials. We evaluated the extended MegaLMM using the maize Genome-To-Fields dataset, consisting of 4,402 varieties cultivated in 195 trials with 87.1% of phenotypic values missing, and demonstrated its high accuracy in genomic prediction under various breeding scenarios. Furthermore, we showcased MegaLMM's superiority over univariate GBLUP in predicting trait performance of experimental genotypes in new environments. Finally, we explored the use of higher-dimensional quantitative ECs and discussed when and how detailed environmental data can be leveraged for genomic prediction from METs. We propose that MegaLMM can be applied to plant breeding of diverse crops and different fields of genetics where large-scale linear mixed models are utilized.
Comprehensive maps of functional variation at transcription factor (TF) binding sites (cis-elements) are crucial for elucidating how genotype shapes phenotype. Here, we report the construction of a pan-cistrome of the maize leaf under well-watered and drought conditions. We quantified haplotype-specific TF footprints across a pan-genome of 25 maize hybrids and mapped over 200,000 variants, genetic, epigenetic, or both (termed binding quantitative trait loci (bQTL)), linked to cis-element occupancy. Three lines of evidence support the functional significance of bQTL: (1) coincidence with causative loci that regulate traits, including vgt1, ZmTRE1 and the MITE transposon near ZmNAC111 under drought; (2) bQTL allelic bias is shared between inbred parents and matches chromatin immunoprecipitation sequencing results; and (3) partitioning genetic variation across genomic regions demonstrates that bQTL capture the majority of heritable trait variation across ~72% of 143 phenotypes. Our study provides an auspicious approach to make functional cis-variation accessible at scale for genetic studies and targeted engineering of complex traits.
Climate change poses a major challenge for both natural and cultivated species. Genomic tools are increasingly used in both conservation and breeding to identify adaptive loci that can be used to guide management in future climates. Here, we study the utility of climate and genomic data for identifying promising alleles using common gardens of a large, geographically diverse sample of traditional maize varieties to evaluate multiple approaches. First, we used genotype data to predict environmental characteristics of germplasm collections to identify varieties that may be pre-adapted to target environments. Second, we used environmental GWAS (envGWAS) to identify loci associated with historical divergence along climatic gradients. Finally, we compared the value of environmental data and envGWAS-prioritized loci to genomic data for prioritizing traditional varieties. We find that maize yield traits are best predicted by genome-wide relatedness and population structure, and that incorporating envGWAS-identified variants or environment-of-origin data provide little additional predictive information. While our results suggest that environmental data provide limited benefit in predicting fitness-related phenotypes, environmental GWAS is nonetheless a potentially powerful approach to identify individual novel loci associated with adaptation, especially when coupled with high density genotyping.
Genomic prediction models that capture genotype-by-environment (GxE) interaction are useful for predicting site-specific performance by leveraging information among related individuals and correlated environments, but implementing such models is computationally challenging. This study describes the algorithm of these scalable approaches, including 2 models with latent representations of GxE interactions, namely MegaLMM and MegaSEM, and an efficient multivariate mixed-model solver, namely Pseudo-expectation Gauss-Seidel (PEGS), fitting different covariance structures [unstructured, extended factor analytic (XFA), Heteroskedastic compound symmetry (HCS)]. Accuracy and runtime are benchmarked on simulated scenarios with varying numbers of genotypes and environments. MegaLMM and PEGS-based XFA and HCS models provided the highest accuracy under sparse testing with 100 testing environments. PEGS-based unstructured model was orders of magnitude faster than restricted maximum likelihood (REML) based multivariate genomic best linear unbiased predictions (GBLUP) while providing the same accuracy. MegaSEM provided the lowest runtime, fitting a model with 200 traits and 20,000 individuals in ∼5 min, and a model with 2,000 traits and 2,000 individuals in less than 3 min. With the genomes-to-fields data, the most accurate predictions were attained with the univariate model fitted across environments and by averaging environment-level genomic estimated breeding values (GEBVs) from models with HCS and XFA covariance structures.
Transcriptomics and proteomics information collected on a platform can predict additive and non-additive effects for platform traits and additive effects for field traits. The effects of climate change in the form of drought, heat stress, and irregular seasonal changes threaten global crop production. The ability of multi-omics data, such as transcripts and proteins, to reflect a plant’s response to such climatic factors can be capitalized in prediction models to maximize crop improvement. Implementing multi-omics characterization in field evaluations is challenging due to high costs. It is, however, possible to do it on reference genotypes in controlled conditions. Using omics measured on a platform, we tested different multi-omics-based prediction approaches, using a high dimensional linear mixed model (MegaLMM) to predict genotypes for platform traits and agronomic field traits in a panel of 244 maize hybrids. We considered two prediction scenarios: in the first one, new hybrids are predicted (CV-NH), and in the second one, partially observed hybrids are predicted (CV-POH). For both scenarios, all hybrids were characterized for omics on the platform. We observed that omics can predict both additive and non-additive genetic effects for the platform traits, resulting in much higher predictive abilities than GBLUP. It highlights their efficiency in capturing regulatory processes in relation to growth conditions. For the field traits, we observed that the additive components of omics only slightly improved predictive abilities for predicting new hybrids (CV-NH, model MegaGAO) and for predicting partially observed hybrids (CV-POH, model GAOxW-BLUP) in comparison to GBLUP. We conclude that measuring the omics in the fields would be of considerable interest in predicting productivity if the costs of omics drop significantly.
Many plant populations exhibit synchronous flowering, which can be advantageous in plant reproduction. However, molecular mechanisms underlying flowering synchrony remain poorly understood. We studied the role of known vernalization-response and flower-promoting pathways in facilitating synchronized flowering in Arabidopsis thaliana. Using the vernalization-responsive Col-FRI genotype, we experimentally varied germination dates and daylength among individuals to test flowering synchrony in field and controlled environments. We assessed the activity of flowering regulation pathways by measuring gene expression across leaves produced at different time points during development and through a mutant analysis. We observed flowering synchrony across germination cohorts in both environments and discovered a previously unknown process where flower-promoting and repressing signals are differentially regulated between leaves that developed under different environmental conditions. We hypothesized this mechanism may underlie synchronization. However, our experiments demonstrated that signals originating from sources other than leaves must also play a pivotal role in synchronizing flowering time, especially in germination cohorts with prolonged growth before vernalization. Our results suggest flowering synchrony is promoted by a plant-wide integration of flowering signals across leaves and among organs. To summarize our findings, we propose a new conceptual model of vernalization-induced flowering synchrony and provide suggestions for future research in this field.
With the rapid development of animal phenomics and deep phenotyping, we can obtain thousands of traditional (but also molecular) phenotypes per individual. However, there is still a lack of exploration regarding how to handle this huge amount of data in the context of animal breeding, presenting a challenge that we are likely to encounter more and more in the future. This study aimed to (1) explore the use of the mega-scale linear mixed model (MegaLMM), a factor model-based approach that is able to simultaneously estimate (co)variance components and genetic parameters in the context of thousands of milk traits, hereafter called thousand-trait (TT) models; (2) compare the phenotype values and genomic breeding value (u) predictions for focal traits (i.e., traits that are targeted for prediction, compared with secondary traits that are helping to evaluate), from single-trait (ST) and TT models, respectively; (3) propose a new approximate method of GEBV (U) prediction with TT models and MegaLMM. We used a total of 3,421 milk mid-infrared (MIR) spectra wavepoints (called secondary traits) and 3 focal traits (average fat percentage [AFP], average methane production [ACH4], and average SCS [ASCS]) collected on 3,302 first-parity Holstein cows. The 3,421 milk MIR wavepoint traits were composed of 311 wave- points in 11 classes (months in lactation). Genotyping information of 564,439 SNPs was available for all animals and was used to calculate the genomic relationship matrix. The MegaLMM was implemented in the framework of the Bayesian sparse factor model and solved through Gibbs sampling (Markov chain Monte Carlo). The heritabilities of the studied 3,421 milk MIR wave- points gradually increased and then decreased in units of 311 wavepoints throughout the lactation. The genetic and phenotypic correlations between the first 311 wavepoints and the other 3,110 wavepoints were low. The accuracies of phenotype predictions from the ST model were lower than those from the TT model for AFP (0.51 vs. 0.93), ACH4 (0.30 vs. 0.86), and ASCS (0.14 vs. 0.33). The same trend was observed for the accuracies of u predictions for AFP (0.59 vs. 0.86), ACH4 (0.47 vs. 0.78), and ASCS (0.39 vs. 0.59). The average correlation between U predicted from the TT model and the new approximate method was 0.90. The new approximate method used for estimating U in MegaLMM will enhance the suitability of MegaLMM for applications in animal breeding. This study conducted an initial investigation into the application of thousands of traits in animal breeding and showed that the TT model is beneficial for the prediction of focal traits (phenotype and breeding values), especially for difficult-to-measure traits (e.g., ACH4).
Maintaining crop yields in the face of climate change is a major challenge facing plant breeding today. Considerable genetic variation exists in ex-situ collections of traditional crop varieties, but identifying adaptive loci and testing their agronomic performance in large populations in field trials is costly. Here, we study the utility of climate and genomic data for identifying promising traditional varieties to incorporate into maize breeding programs. To do so, we use phenotypic data from more than 4,000 traditional maize varieties grown in 13 trial environments. First, we used genotype data to predict environmental characteristics of germplasm collections to identify varieties that may be locally adapted to target environments. Second, we used environmental GWAS (envGWAS) to identify genetic loci associated with historical divergence along climatic gradients, such as the putative heat shock protein Hsftf9 and the large-scale adaptive inversion Inv4m. Finally, we compared the value of environmental data and envGWAS-prioritized loci to genomic data for prioritizing traditional varieties. We find that maize yield traits are best predicted by genomic data, and that envGWAS-identified variants provide little direct predictive information over patterns of population structure. We also find that adding environment-of-origin variables does not improve yield component prediction over kinship or population structure alone, but could be a useful selection proxy in the absence of sequencing data. While our results suggest little utility of environmental data for selecting traditional varieties to incorporate in breeding programs, environmental GWAS is nonetheless a potentially powerful approach to identify individual novel loci for maize improvement, especially when coupled with high density genotyping. ### Competing Interest Statement The authors have declared no competing interest.
Statistical machine learning (ML) extracts patterns from extensive genomic, phenotypic, and environmental data. ML algorithms automatically identify relevant features and use cross-validation to ensure robust models and improve prediction reliability in new lines. Furthermore, ML analyses of genotype-by-environment (G×E) interactions can offer insights into the genetic factors that affect performance in specific environments. By leveraging historical breeding data, ML streamlines strategies and automates analyses to reveal genomic patterns. In this review we examine the transformative impact of big data, including multi-trait genomics, phenomics, and environmental covariables, on genomic-enabled prediction in plant breeding. We discuss how big data and ML are revolutionizing the field by enhancing prediction accuracy, deepening our understanding of G×E interactions, and optimizing breeding strategies through the analysis of extensive and diverse datasets.