European maize landraces encompass a large amount of genetic diversity, allowing them to be well-adapted to their local environments. This diversity can be exploited to improve the fitness of elite material in the face of a changing climate. We characterized the genetic diversity of 333 individual plants from 40 European maize landrace populations (EMLPs). We identified five genetic groups that mirrored the proximities of their geographical origins. Fixation indices showed moderate differentiation among genetic groups (0.034 to 0.093). More than half of the genetic variance was observed to be partitioned among individuals. Nucleotide diversity of EMLPs decreased significantly as latitude increased (from 0.16 to 0.04), suggesting serial founder events during maize expansion in Europe. GWAS with latitude, longitude, and elevation as response variables identified 28, 347, and 68 significant SNP positions, respectively. We pinpointed significant SNPs near dwarf8, tb1, ZCN7, ZCN8, and ZmMADS69 and identified 126 candidate genes with ontology terms indicative of local adaptation in maize, regulating adaptation to diverse abiotic and biotic environmental stresses. This study suggests a quick and cost-efficient approach to identifying genes involved in local adaptation without requiring field data. The EMLPs used in this study have been assembled to serve as a continuing resource of genetic diversity for further research aimed at improving agronomically relevant adaptation traits.
Forest tree breeding is an extremely long and tedious process. To study the genetic architecture of polygenic traits in long-lived species such as forest trees, costly field experiments are implemented. Phenotypic data on traits, that are measured at maturity, are only available after a long time and juvenile-mature correlations are often unknown. Genome-wide association studies (GWAS) aim to identify loci associated with drought stress, tree physiology, growth or wood quality traits, which could be prioritized in breeding programs. Genotypic and phenotypic data were collected from approximately 100 adult beech trees per stand in five locations in the South-Eastern Romanian Carpathians along an altitudinal gradient associated with precipitation and temperature. We performed GWAS using PLINK to identify SNP markers associated with traits related to drought stress. A total of 121 markers on eight chromosomes were identified as being associated with stomatal density. Sixty-four markers are located on chromosome 10 in a region spreading from ∼4.89 to 13.67 Mb. There are five genes in this region that are thought to play a role in controlling stomatal density. All markers within this region exhibit similar allele frequencies, which are correlated with stomatal density and the altitudinal gradient of the stands. We assume this entire region is jointly involved in local adaptation to drought stress. We identified one interesting candidate SNP associated with leaf nitrogen content. Two SNP markers were identified as being significantly associated with δ13C as measure of intrinsic water use efficiency. Additionally, signals of significant polygenic selection for δ13C were observed. ### Competing Interest Statement The authors have declared no competing interest. Federal Ministry of Food and AgricultureFederal Ministry of Food and Agriculture, https://ror.org/04jw21793, 2218WK43B4, 2218WK43A4
Experimental evolution studies are common in agricultural research, where they are often deemed “long-term selection.” These are often used to perform selection mapping, which involves identifying markers that were putatively under selection based on finding signals of selection left in the genome. A challenge of previous selection mapping studies, especially in agricultural research, has been the specification of robust significance thresholds. This is in large part because long-term selection studies in crops have rarely included replication. Usually, significance thresholds in long-term selection experiments are based on outliers from an empirical distribution. This approach is prone to missing true positives or including false positives. Under laboratory conditions with model species, replicated selection has been shown to be a powerful tool, especially for the specification of significance thresholds. Another challenge is that commonly used single-marker-based statistics may identify neutral linked loci which have hitchhiked along with regions that are actually under selection. In this study, we conducted divergent, replicated selection for short and tall plant height in a random-mating maize population under real field conditions. Selection of the 5% tallest and shortest plants was conducted for 3 generations. Significance thresholds were specified using the false discovery rate for selection (FDRfS) based on a window-based statistic applied to a statistic leveraging replicated selection (FSTSum). Overall, we found 2 significant regions putatively under selection. One region was located on chromosome 3 close to the plant-height genes Dwarf1 and iAA8. We applied a haplotype block analysis to further dissect the pattern of selection in significant regions of the genome. We observed patterns of strong selection in the subpopulations selected for short plant height on chromosome 3.
OBJECTIVES:The Genomes to Fields (G2F) 2022 Maize Genotype by Environment (GxE) Prediction Competition aimed to develop models for predicting grain yield for the 2022 Maize GxE project field trials, leveraging the datasets previously generated by this project and other publicly available data.DATA DESCRIPTION:This resource used data from the Maize GxE project within the G2F Initiative [1]. The dataset included phenotypic and genotypic data of the hybrids evaluated in 45 locations from 2014 to 2022. Also, soil, weather, environmental covariates data and metadata information for all environments (combination of year and location). Competitors also had access to ReadMe files which described all the files provided. The Maize GxE is a collaborative project and all the data generated becomes publicly available [2]. The dataset used in the 2022 Prediction Competition was curated and lightly filtered for quality and to ensure naming uniformity across years.
Accurate prediction of the phenotypic outcomes produced by different combinations of genotypes, environments, and management interventions remains a key goal in biology with direct applications to agriculture, research, and conservation. The past decades have seen an expansion of new methods applied toward this goal. Here we predict maize yield using deep neural networks, compare the efficacy of 2 model development methods, and contextualize model performance using conventional linear and machine learning models. We examine the usefulness of incorporating interactions between disparate data types. We find deep learning and best linear unbiased predictor (BLUP) models with interactions had the best overall performance. BLUP models achieved the lowest average error, but deep learning models performed more consistently with similar average error. Optimizing deep neural network submodules for each data type improved model performance relative to optimizing the whole model for all data types at once. Examining the effect of interactions in the best-performing model revealed that including interactions altered the model's sensitivity to weather and management features, including a reduction of the importance scores for timepoints expected to have a limited physiological basis for influencing yield-those at the extreme end of the season, nearly 200 days post planting. Based on these results, deep learning provides a promising avenue for the phenotypic prediction of complex traits in complex environments and a potential mechanism to better understand the influence of environmental and genetic factors.
Identifying selection on polygenic complex traits in crops and livestock is important for understanding evolution and helps prioritize important characteristics for breeding. Quantitative trait loci (QTL) that contribute to polygenic trait variation often exhibit small or infinitesimal effects. This hinders the ability to detect QTL-controlling polygenic traits because enormously high statistical power is needed for their detection. Recently, we circumvented this challenge by introducing a method to identify selection on complex traits by evaluating the relationship between genome-wide changes in allele frequency and estimates of effect size. The approach involves calculating a composite statistic across all markers that capture this relationship, followed by implementing a linkage disequilibrium-aware permutation test to evaluate if the observed pattern differs from that expected due to drift during evolution and population stratification. In this manuscript, we describe "Ghat," an R package developed to implement this method to test for selection on polygenic traits. We demonstrate the package by applying it to test for polygenic selection on 15 published European wheat traits including yield, biomass, quality, morphological characteristics, and disease resistance traits. Moreover, we applied Ghat to different simulated populations with different breeding histories and genetic architectures. The results highlight the power of Ghat to identify selection on complex traits. The Ghat package is accessible on CRAN, the Comprehensive R Archival Network, and on GitHub.
A deeper understanding of the mechanisms underlying the impacts of multiple stresses in crops is direly needed given the climate change-induced risks to achieving food security for a growing world population. Global warming has already led to a higher frequency of multiple stresses occurring concurrently or subsequently and will continue to do so for the next decades. Plant-stress interactions are commonly subdivided into abiotic and biotic stresses and studied separately. Under field conditions, these stress interactions are usually multiple and interactive in character.To date, the mechanisms determining interactions between abiotic and biotic stresses and their effects on crop performance are unknown for most crops and stress combinations. Field data are particularly scarce as most studies have focused on laboratory model systems using few environmental parameters in controlled conditions, which cannot reflect the dynamics in the field. Adequate modelling approaches capable of describing basic crop growth processes and simultaneously capturing response to abiotic and biotic stress interactions and their impacts on crop yield and quality do not exist so far.The aim of this paper is to present the design of a joint experimental and modelling platform (MultiStress) capable of creating a deeper understanding of the overall impact of combined (abiotic+biotic) stresses on crop physiology and productivity (grain yield, biomass, grain and stover quality, nutrient/water use efficiency, etc.) using the cereal maize as one of the most important crops globally as a model.The empirical knowledge gained from the experimental set-up and formalized in an associated modelling platform is utilized to define traits for stress tolerant breeding to be considered in ideotyping cereal cultivars for future target environments. In our example, in a research Pillar I, we describe a field experimental platform (with rainout shelters) applicable under temperate and tropical climate conditions to investigate the interactions of drought and nitrogen deficiency with the foliar disease Northern Corn Leaf Blight caused by Setosphaeria turcica on the one hand, and stem borer caterpillars on the other. Pillar II is an associated process-based modelling platform enabling integration of new genetic and ecophysiological knowledge and extrapolate the findings in time and space.Applying a systems approach in conjunction with this platform we can test the following hypotheses: (i) the impact of combined abiotic and biotic stress interactions on crop growth and yield formation and quality is non-additive and thus differs from the sum of individual stress impacts; (ii) while the mechanisms underlying the abiotic and biotic stress interactions are of universal validity, their impacts are modulated by certain environmental conditions (such as temperature, light conditions and soil properties).Realization and evaluation of such platform will allow consideration of interactions between abiotic and biotic stresses and hence improve the predictive skill of crop growth models.
ObjectivesThis release note describes the Maize GxE project datasets within the Genomes to Fields (G2F) Initiative. The Maize GxE project aims to understand genotype by environment (GxE) interactions and use the information collected to improve resource allocation efficiency and increase genotype predictability and stability, particularly in scenarios of variable environmental patterns. Hybrids and inbreds are evaluated across multiple environments and phenotypic, genotypic, environmental, and metadata information are made publicly available.Data descriptionThe datasets include phenotypic data of the hybrids and inbreds evaluated in 30 locations across the US and one location in Germany in 2020 and 2021, soil and climatic measurements and metadata information for all environments (combination of year and location), ReadMe, and description files for each data type. A set of common hybrids is present in each environment to connect with previous evaluations. Each environment had a collaborator responsible for collecting and submitting the data, the GxE coordination team combined all the collected information and removed obvious erroneous data. Collaborators received the combined data to use, verify and declare that the data generated in their own environments was accurate. Combined data is released to the public with minimal filtering to maintain fidelity to the original data.
Objectives This report provides information about the public release of the 2018–2019 Maize G X E project of the Genomes to Fields (G2F) Initiative datasets. G2F is an umbrella initiative that evaluates maize hybrids and inbred lines across multiple environments and makes available phenotypic, genotypic, environmental, and metadata information. The initiative understands the necessity to characterize and deploy public sources of genetic diversity to face the challenges for more sustainable agriculture in the context of variable environmental conditions. Data description Datasets include phenotypic, climatic, and soil measurements, metadata information, and inbred genotypic information for each combination of location and year. Collaborators in the G2F initiative collected data for each location and year; members of the group responsible for coordination and data processing combined all the collected information and removed obvious erroneous data. The collaborators received the data before the DOI release to verify and declare that the data generated in their own locations was accurate. ReadMe and description files are available for each dataset. Previous years of evaluation are already publicly available, with common hybrids present to connect across all locations and years evaluated since this project’s inception.
Low-density genotyping followed by imputation reduces genotyping costs while still providing high-density marker information. An increased marker density has the potential to improve the outcome of all applications that are based on genomic data. This study investigates techniques for 1k to 20k genomic marker imputation for plant breeding programs with sugar beet as an example crop, where these are realistic marker numbers for modern breeding applications. The generally accepted ‘gold standard’ for imputation, Beagle 5.1, was compared to the recently developed software AlphaPlantImpute2 which is designed specifically for plant breeding. For Beagle 5.1 and AlphaPlantImpute2, the imputation strategy as well as the imputation parameters were optimized in this study. We found that the imputation accuracy of Beagle could be tremendously improved (0.22 to 0.67) by tuning parameters, mainly by lowering the values for the parameter for the effective population size and increasing the number of iterations performed. Separating the phasing and imputation steps also improved accuracies when optimized parameters were used (0.67 to 0.82). We also found that the imputation accuracy of Beagle decreased when more low-density lines were included for imputation. AlphaPlantImpute2 produced very high accuracies without optimization (0.89) and was generally less responsive to optimization. Overall, AlphaPlantImpute2 performed relatively better for imputation while Beagle was better for phasing. Combining both tools yielded the highest accuracies. Summary Genotype marker information allows the prediction of an individual’s breeding value without the need to observe its actual phenotype which can accelerate the breeding progress. The more markers are genotyped, the better the genomic prediction may be. However, analyzing many markers is costly, particularly in commercial breeding programs where thousands of new individuals are genotyped. A solution to obtain information for all markers, while spending comparatively little on genotyping, is to genotype only a small fraction of markers in most individuals. Together with high-density information on other individuals, the low-density individuals can be imputed to high-density. High-density individuals are typically parents or highly influential individuals. In this study, we compare the widely used software Beagle with the recently developed software AlphaPlantImpute2 on plant breeding data. To allow a fair comparison, we first optimized existing methods and developed new approaches. This was done to avoid comparing results of a less ideal version of one software to optimized settings of another software. After optimization, the software were evaluated in different scenarios with regards to genotyping errors, population types and number of markers based on simulated data. Simulated data were based on real marker data from a sugar beet population as input to mimic the population history of a commercial breeding population. AlphaPlantImpute2 performs well with default parameters, while much optimization with regards to parameters and strategy was needed to boost accuracies of Beagle. A pipeline is presented which uses Beagle for phasing and AlphaPlantImpute2 for imputation. This pipeline yielded the highest accuracies and shortest run time. Core Ideas Beagle is sensitive to parameter tuning Best imputation accuracies could be achieved by using a combination of Beagle and AlphaPlantImpute2 The population structure influence imputation accuracy
A collection of 46 pea (Pisum sativum L.) accessions, mostly from Europe, were analysed for genetic diversity using the GenoPea 13.2 K SNP Array chip. Of these accessions were 24 nomal-leaved and 22 semi-leafless. Principal components analysis (PCA) separated the peas into two groups characterized by the two different leaf types, although some genotypes were exceptions and appeared in the opposite group. Cluster analysis confirmed the two groups. A dendrogram showed larger genetic distances between genotypes in the normal-leafed group compared to semi-leafless genotypes. Both PCA and cluster analysis show that the two leave types are genetically divergent. So normal-leaved peas are an interesting genetic resource, even if the breeding goal is to develop semi-leafless varieties.
We introduce the R-package learnMET, developed as a flexible framework to enable a collection of analyses on multi-environment trial breeding data with machine learning-based models. learnMET allows the combination of genomic information with environmental data such as climate and/or soil characteristics. Notably, the package offers the possibility of incorporating weather data from field weather stations, or to retrieve global meteorological datasets from a NASA database. Daily weather data can be aggregated over specific periods of time based on naive (for instance, nonoverlapping 10-day windows) or phenological approaches. Different machine learning methods for genomic prediction are implemented, including gradient-boosted decision trees, random forests, stacked ensemble models, and multilayer perceptrons. These prediction models can be evaluated via a collection of cross-validation schemes that mimic typical scenarios encountered by plant breeders working with multi-environment trial experimental data in a user-friendly way. The package is published under an MIT license and accessible on GitHub.
Abstract Background Genomic selection is a powerful tool in plant breeding. By building a prediction model using a training set with markers and phenotypes, genomic estimated breeding values (GEBVs) can be used as predictions of breeding values in a target set with only genotype data. There is, however, limited information on how prediction accuracy of genomic prediction can be optimized. The objective of this study was to evaluate the performance of 11 genomic prediction models across species in terms of prediction accuracy for two traits with different heritabilities using several subsets of markers and training population proportions. Species studied were maize (Zea mays, L.), soybean (Glycine max, L.), and rice (Oryza sativa, L.), which vary in linkage disequilibrium (LD) decay rates and have contrasting genetic architectures. Results Correlations between observed and predicted GEBVs were determined via cross validation for three training-to-testing proportions (90:10, 70:30, and 50:50). Maize, which has the shortest extent of LD, showed the highest prediction accuracy. Amongst all the models tested, Bayes B performed better than or equal to all other models for each trait in all the three crops. Traits with higher broad-sense and narrow-sense heritabilities were associated with higher prediction accuracy. When subsets of markers were selected based on LD, the accuracy was similar to that observed from the complete set of markers. However, prediction accuracies were significantly improved when using a subset of total markers that were significant at P ≤ 0.05 or P ≤ 0.10. As expected, exclusion of QTL-associated markers in the model reduced prediction accuracy. Prediction accuracy varied among different training population proportions. Conclusions We conclude that prediction accuracy for genomic selection can be improved by using the Bayes B model with a subset of significant markers and by selecting the training population based on narrow sense heritability.
Engaging students in an international online setting that is interdisciplinary and culturally diverse is a challenge. A joint classroom between German and Ugandan universities used a formative assessment approach paired with active learning elements to foster individual and peer learning in an international virtual setting. A survey at three different times across the semester explored students’ perceptions towards the value of the active learning activities and evaluated how perceptions changed over time. Overall, students enjoyed the diverse active learning activities and perceived value toward their success in class. This was more pronounced and unidirectional for individual tasks than it was for group work. In addition to the findings of the structured survey, observation and feedback indicated that other elements contributed to effective course delivery. These included clear and frequent communication to the students from the primary instructor, prompt feedback from the instructor on graded exercises, such as a reflective learning diary and ungraded quizzes, and student confidence that sincere effort would achieve a good grade.
For decades, practical limitations and cost have created a bias in plant genetic studies toward species that tolerate self pollination.An important work by Chen et al.(2021) in this issue provides key tools to overcome this bias and paves the way toward more efficient genetics and breeding in outcrossing crops. Plants that tolerate full or partial inbreeding can be maintained in perpetuity and increased infinitely.Conversely, outcrossing species that do not tolerate inbreeding can be difficult to study due to their genetic makeup.For instance, heterozygous individuals segregating in a population cannot be grown as replicated "genotypes," but instead must be grown and phenotyped as individual plants.Due to the heavy influence of macro-and microenvironmental factors on the performance of most plants, studies that use phenotypes from single individuals are rarely implemented except when absolutely necessary due to time or mating system (e.g., Müller et al., 2019), with a few notable exceptions in crops (e.g., Gyawali et al., 2019).In addition to challenges surrounding phenotyping, genotyping challenges must also be overcome when studying heterozygous populations of outcrossers.For instance, quantitative trait locus (QTL) mapping requires genomic segments within each individual in the mapping population to be assigned to a corresponding parent.In double haploid, recombinant-inbred, or F2 mapping populations, the mathematics behind assigning genotypes to "parent 1" or "parent 2" is trivial.But, when these parents are themselves segregating for various alleles across loci, the math becomes difficult or impossible, especially for the case of polyploids.
The strength of the stalk rind, measured as rind penetrometer resistance (RPR), is an important contributor to stalk lodging resistance. To enhance the genetic architecture of RPR, we combined selection mapping on populations developed by 15 cycles of divergent selection for high and low RPR with time-course transcriptomic and metabolic analyses of the stalks. Divergent selection significantly altered allele frequencies of 3,656 and 3,412 single- nucleotide polymorphisms (SNPs) in the high and low RPR populations, respectively. Surprisingly, only 110 (1.56%) SNPs under selection were common in both populations, while the majority (98.4%) were unique to each population. This result indicated that high and low RPR phenotypes are produced by biologically distinct mechanisms. Remarkably, regions harboring lignin and polysaccharide genes were preferentially selected in high and low RPR populations, respectively. The preferential selection was manifested as higher lignification and increased saccharification of the high and low RPR stalks, respectively. The evolution of distinct gene classes according to the direction of selection was unexpected in the context of parallel evolution and demonstrated that selection for a trait, albeit in different directions, does not necessarily act on the same genes. Tricin, a grass-specific monolignol that initiates the incorporation of lignin in the cell walls, emerged as a key determinant of RPR. Integration of selection mapping and transcriptomic analyses with published genetic studies of RPR identified several candidate genes including ZmMYB31, ZmNAC25, ZmMADS1, ZmEXPA2, ZmIAA41 and hk5. These findings provide a foundation for an enhanced understanding of RPR and the improvement of stalk lodging resistance.
The development of crop varieties with stable performance in future environmental conditions represents a critical challenge in the context of climate change. Environmental data collected at the field level, such as soil and climatic information, can be relevant to improve predictive ability in genomic prediction models by describing more precisely genotype-by-environment interactions, which represent a key component of the phenotypic response for complex crop agronomic traits. Modern predictive modeling approaches can efficiently handle various data types and are able to capture complex nonlinear relationships in large datasets. In particular, machine learning techniques have gained substantial interest in recent years. Here we examined the predictive ability of machine learning-based models for two phenotypic traits in maize using data collected by the Maize Genomes to Fields (G2F) Initiative. The data we analyzed consisted of multi-environment trials (METs) dispersed across the United States and Canada from 2014 to 2017. An assortment of soil- and weather-related variables was derived and used in prediction models alongside genotypic data. Linear random effects models were compared to a linear regularized regression method (elastic net) and to two nonlinear gradient boosting methods based on decision tree algorithms (XGBoost, LightGBM). These models were evaluated under four prediction problems: (1) tested and new genotypes in a new year; (2) only unobserved genotypes in a new year; (3) tested and new genotypes in a new site; (4) only unobserved genotypes in a new site. Accuracy in forecasting grain yield performance of new genotypes in a new year was improved by up to 20% over the baseline model by including environmental predictors with gradient boosting methods. For plant height, an enhancement of predictive ability could neither be observed by using machine learning-based methods nor by using detailed environmental information. An investigation of key environmental factors using gradient boosting frameworks also revealed that temperature at flowering stage, frequency and amount of water received during the vegetative and grain filling stage, and soil organic matter content appeared as important predictors for grain yield in our panel of environments.
Understanding the evolutionary history of crops, including identifying wild relatives, helps to provide insight for conservation and crop breeding efforts. Cultivated Brassica oleracea has intrigued researchers for centuries due to its wide diversity in forms, which include cabbage, broccoli, cauliflower, kale, kohlrabi, and Brussels sprouts. Yet, the evolutionary history of this species remains understudied. With such different vegetables produced from a single species, B. oleracea is a model organism for understanding the power of artificial selection. Persistent challenges in the study of B. oleracea include conflicting hypotheses regarding domestication and the identity of the closest living wild relative. Using newly generated RNA-seq data for a diversity panel of 224 accessions, which represents 14 different B. oleracea crop types and nine potential wild progenitor species, we integrate phylogenetic and population genetic techniques with ecological niche modeling, archaeological, and literary evidence to examine relationships among cultivars and wild relatives to clarify the origin of this horticulturally important species. Our analyses point to the Aegean endemic B. cretica as the closest living relative of cultivated B. oleracea, supporting an origin of cultivation in the Eastern Mediterranean region. Additionally, we identify several feral lineages, suggesting that cultivated plants of this species can revert to a wild-like state with relative ease. By expanding our understanding of the evolutionary history in B. oleracea, these results contribute to a growing body of knowledge on crop domestication that will facilitate continued breeding efforts including adaptation to changing environmental conditions.
Crop domestication is a fascinating area of study, as shown by a multitude of recent reviews. Coupled with the increasing availability of genomic and phenomic resources in numerous crop species, insights from evolutionary biology will enable a deeper understanding of the genetic architecture and short-term evolution of complex traits, which can be used to inform selection strategies. Future advances in crop improvement will rely on the integration of population genetics with plant breeding methodology, and the development of community resources to support research in a variety of crop life histories and reproductive strategies. We highlight recent advances related to the role of selective sweeps and demographic history in shaping genetic architecture, how these breakthroughs can inform selection strategies, and the application of precision gene editing to leverage these connections.