1AbstractSoybean (Glycine max(L.) Merr.) provides plant-based protein for global food production and is extensively bred to create cultivars with greater productivity in distinct environments. Plant breeders evaluate new soybean genotypes using multi-environment trials (MET). The application of MET assumes that trial locations provide representative environmental conditions that cultivars are likely to encounter when grown by farmers. In addition, MET are important to depict the patterns of genotype by environment interactions (GEI). To evaluate GEI for soybean seed yield and identify mega-environments (ME), a retrospective analysis of 39,006 data points from experimental soybean genotypes evaluated in preliminary and uniform field trials conducted by public plant breeders from 1989-2019 was considered. ME were identified from phenotypic information from the annual trials, geographic, soil, and meteorological records at the trial locations. Results indicate that yield variation was mostly explained by location and location by year interactions. The static portion of the GEI represented 26.30% of the total yield variance. Estimates of variance components derived from linear mixed models demonstrated that the phenotypic variation due to genotype by location interaction effects was greater than genotype by year interaction effects. A trend analysis indicated a two-fold increase in the genotypic variance between 1989-1995 and 1996-2019. Furthermore, the heterogeneous estimates of genotypic, genotype by location, genotype by year, and genotype by location by year variances, were encapsulated by distinct probability distributions. The observed target population of environments can be divided into at least two and at most three ME, thereby suggesting improvements in the response to selection can be achieved when selecting directly for clustered (i.e., regions, ME) versus selecting across regions. Clusters obtained using phenotypic data, latitude, and soil variables plus elevation, were the most effective. In addition, we published the R package SoyURT which contains the data sets used in this work.2HighlightsMega-environments can be identified with phenotypic, geographic, and meteorological data.Reliable estimates of variances can be obtained with proper analyses of historical data.Genotype by location was more important than genotype by year variation for seed yield.The trend in genotype by environment variances was captured in probability distributions.
1 Abstract Genetic improvements of discrete characteristics such as flower color, the genetic improvements are obvious and easy to demonstrate; however, for characteristics that are measured on continuous scales, the genetic contributions are incremental and less obvious. Reliable and accurate methods are required to disentangle the confounding genetic and non-genetic components of quantitative traits. Stochastic simulations of soybean (Glycine max (L.) Merr.) breeding programs were performed to evaluate models to estimate the realized genetic gain (RGG) from 30 years of multi-environment trials (MET). True breeding values were simulated under an infinitesimal model to represent the genetic contributions to soybean seed yield under various MET conditions. Estimators were evaluated using objective criteria of bias and linearity. Results indicated all estimation models were biased. Covariance modeling as well as direct versus indirect estimation resulted in substantial differences in RGG estimation. Although there were no unbiased models, the three best-performing models resulted in an average bias of ±7.41 kg/ha −1 /yr −1 (±0.11 bu/ac −1 /yr −1 ). Rather than relying on a single model to estimate RGG, we recommend the application of multiple models and consider the range of the estimated values. Further, based on our simulations parameters, we do not think it is appropriate to use any single models to compare breeding programs or quantify the efficiency of proposed new breeding strategies. Lastly, for public soybean programs breeding for maturity groups II and III in North America from 1989 to 2019, the range of estimated RGG values was from 18.16 to 39.68 kg/ha −1 /yr −1 (0.27 to 0.59 bu/ac −1 /yr −1 ).
In maize, doubled haploid (DH) lines are created in vivo through crosses with maternal haploid inducers. Their induction ability, usually expressed as haploid induction rate (HIR), is known to be under polygenic control. Although two major genes (MTL and ZmDMP) affecting this trait were recently described, many others remain unknown. To identify them, we designed and performed a SNP based (~9007) genome-wide association study using a large and diverse panel of 159 maternal haploid inducers. Our analyses identified a major gene near MTL, which is present in all inducers and necessary to disrupt haploid induction. We also found a significant quantitative trait loci (QTL) on chromosome 10 using a case-control mapping approach, in which 793 noninducers were used as controls. This QTL harbors a kokopelli ortholog, whose role in maternal haploid induction was recently described in Arabidopsis. QTL with smaller effects were identified on six of the ten maize chromosomes, confirming the polygenic nature of this trait. These QTL could be incorporated into inducer breeding programs through marker-assisted selection approaches. Further improving HIR is important to reduce the cost of DH line production.
Soybean is grown primarily for the protein and oil extracted from its seed and its value is influenced by these components. The objective of this study was to map marker‐trait associations (MTAs) for the concentration of seed protein, oil, and meal protein using the soybean nested association mapping (SoyNAM) population. The composition traits were evaluated on seed harvested from over 5000 inbred lines of the SoyNAM population grown in 10 field locations across 3 years. Estimated heritabilities were at least 0.85 for all three traits. The genotyping of lines with single nucleotide polymorphism markers resulted in the identification of 107 MTAs for the three traits. When MTAs for the three traits that mapped within 5 cM intervals were binned together, the MTAs were mapped to 64 intervals on 19 of the 20 soybean chromosomes. The majority of the MTA effects were small and of the 107 MTAs, 37 were for protein content, 39 for meal protein, and 31 for oil content. For cases where a protein and oil MTAs mapped to the same interval, most (94%) significant effects were opposite for the two traits, consistent with the negative correlation between these traits. A coexpression analysis identified candidate genes linked to MTAs and 18 candidate genes were identified. The large number of small effect MTAs for the composition traits suggest that genomic prediction would be more effective in improving these traits than marker‐assisted selection.
Accounting for field variation patterns plays a crucial role in interpreting phenotype data and, thus, in plant breeding. Several spatial models have been developed to account for field variation. Spatial analyses show that spatial models can successfully increase the quality of phenotype measurements and subsequent selection accuracy for continuous data types such as grain yield and plant height. The phenotypic data for stress traits are usually recorded in ordinal data scores but are traditionally treated as numerical values with normal distribution, such as iron deficiency chlorosis (IDC). The effectiveness of spatial adjustment for ordinal data has not been systematically compared. The research objective described here is to evaluate methods for spatial adjustment of ordinal data, using soybean IDC as an example. Comparisons of adjustment effectiveness for spatial autocorrelation were conducted among eight different models. The models were divided into three groups: Group I, moving average grid adjustment; group II, geospatial autoregressive regression (SAR) models; and Group III, tensor product penalized P-splines. Results from the model comparison show that the effectiveness of the models depends on the severity of field variation, the irregularity of the variation pattern, and the model used. The geospatial SAR models outperform the other models for ordinal IDC data. Prediction accuracy for the lines planted in the IDC high-pressure area is 11.9% higher than those planted in low-IDC-pressure regions. The relative efficiency of the mixed SAR model is 175%, relative to the baseline ordinary least squares model. Even though the geospatial SAR model is the best among all the compared models, the efficiency is not as good for ordinal data types as for numeric data.
Photosynthesis is a key target to improve crop production in many species including soybean [Glycine max (L.) Merr.]. A challenge is that phenotyping photosynthetic traits by traditional approaches is slow and destructive. There is proof-of-concept for leaf hyperspectral reflectance as a rapid method to model photosynthetic traits. However, the crucial step of demonstrating that hyperspectral approaches can be used to advance understanding of the genetic architecture of photosynthetic traits is untested. To address this challenge, we used full-range (500-2,400 nm) leaf reflectance spectroscopy to build partial least squares regression models to estimate leaf traits, including the rate-limiting processes of photosynthesis, maximum Rubisco carboxylation rate, and maximum electron transport. In total, 11 models were produced from a diverse population of soybean sampled over multiple field seasons to estimate photosynthetic parameters, chlorophyll content, leaf carbon and leaf nitrogen percentage, and specific leaf area (with R2 from 0.56 to 0.96 and root mean square error approximately <10% of the range of calibration data). We explore the utility of these models by applying them to the soybean nested association mapping population, which showed variability in photosynthetic and leaf traits. Genetic mapping provided insights into the underlying genetic architecture of photosynthetic traits and potential improvement in soybean. Notably, the maximum Rubisco carboxylation rate mapped to a region of chromosome 19 containing genes encoding multiple small subunits of Rubisco. We also mapped the maximum electron transport rate to a region of chromosome 10 containing a fructose 1,6-bisphosphatase gene, encoding an important enzyme in the regeneration of ribulose 1,5-bisphosphate and the sucrose biosynthetic pathway. The estimated rate-limiting steps of photosynthesis were low or negatively correlated with yield suggesting that these traits are not influenced by the same genetic mechanisms and are not limiting yield in the soybean NAM population. Leaf carbon percentage, leaf nitrogen percentage, and specific leaf area showed strong correlations with yield and may be of interest in breeding programs as a proxy for yield. This work is among the first to use hyperspectral reflectance to model and map the genetic architecture of the rate-limiting steps of photosynthesis.
Plant breeding is a decision-making discipline based on understanding project objectives. Genetic improvement projects can have two competing objectives: maximize the rate of genetic improvement and minimize the loss of useful genetic variance. For commercial plant breeders, competition in the marketplace forces greater emphasis on maximizing immediate genetic improvements. In contrast, public plant breeders have an opportunity, perhaps an obligation, to place greater emphasis on minimizing the loss of useful genetic variance while realizing genetic improvements. Considerable research indicates that short-term genetic gains from genomic selection are much greater than phenotypic selection, while phenotypic selection provides better long-term genetic gains because it retains useful genetic diversity during the early cycles of selection. With limited resources, must a soybean breeder choose between the two extreme responses provided by genomic selection or phenotypic selection? Or is it possible to develop novel breeding strategies that will provide a desirable compromise between the competing objectives? To address these questions, we decomposed breeding strategies into decisions about selection methods, mating designs, and whether the breeding population should be organized as family islands. For breeding populations organized into islands, decisions about possible migration rules among family islands were included. From among 60 possible strategies, genetic improvement is maximized for the first five to 10 cycles using genomic selection and a hub network mating design, where the hub parents with the largest selection metric make large parental contributions. It also requires that the breeding populations be organized as fully connected family islands, where every island is connected to every other island, and migration rules allow the exchange of two lines among islands every other cycle of selection. If the objectives are to maximize both short-term and long-term gains, then the best compromise strategy is similar except that the mating design could be hub network, chain rule, or a multi-objective optimization method-based mating design. Weighted genomic selection applied to centralized populations also resulted in the realization of the greatest proportion of the genetic potential of the founders but required more cycles than the best compromise strategy.
Trait introgression is a complex process that plant breeders use to introduce desirable alleles from one variety or species to another. Two of the major types of decisions that must be made during this sophisticated and uncertain workflow are: parental selection and resource allocation. We formulated the trait introgression problem as an engineering process and proposed a Markov Decision Processes (MDP) model to optimize the resource allocation procedure. The efficiency of the MDP model was compared with static resource allocation strategies and their trade-offs among budget, deadline, and probability of success are demonstrated. Simulation results suggest that dynamic resource allocation strategies from the MDP model significantly improve the efficiency of the trait introgression by allocating the right amount of resources according to the genetic outcome of previous generations.
The Beavis effect in quantitative trait locus (QTL) mapping describes a phenomenon that the estimated effect size of a statistically significant QTL (measured by the QTL variance) is greater than the true effect size of the QTL if the sample size is not sufficiently large. This is a typical example of the Winners' curse applied to molecular quantitative genetics. Theoretical evaluation and correction for the Winners' curse have been studied for interval mapping. However, similar technologies have not been available for current models of QTL mapping and genome-wide association studies where a polygene is often included in the linear mixed models to control the genetic background effect. In this study, we developed the theory of the Beavis effect in a linear mixed model using a truncated noncentral Chi-square distribution. We equated the observed Wald test statistic of a significant QTL to the expectation of a truncated noncentral Chi-square distribution to obtain a bias-corrected estimate of the QTL variance. The results are validated from replicated Monte Carlo simulation experiments. We applied the new method to the grain width (GW) trait of a rice population consisting of 524 homozygous varieties with over 300 k single nucleotide polymorphism markers. Two loci were identified and the estimated QTL heritability were corrected for the Beavis effect. Bias correction for the larger QTL on chromosome 5 (GW5) with an estimated heritability of 12% did not change the QTL heritability due to the extremely large test score and estimated QTL effect. The smaller QTL on chromosome 9 (GW9) had an estimated QTL heritability of 9% reduced to 6% after the bias-correction.
In soybean variety development and genetic improvement projects, iron deficiency chlorosis (IDC) is visually assessed as an ordinal response variable. Linear Mixed Models for Genomic Prediction (GP) have been developed, compared, and used to select continuous plant traits such as yield, height, and maturity, but can be inappropriate for ordinal traits. Generalized Linear Mixed Models have been developed for GP of ordinal response variables. However, neither approach addresses the most important questions for cultivar development and genetic improvement: How frequently are the ‘wrong’ genotypes retained, and how often are the ‘correct’ genotypes discarded? The research objective reported herein was to compare outcomes from four data modeling and six algorithmic modeling GP methods applied to IDC using decision metrics appropriate for variety development and genetic improvement projects. Appropriate metrics for decision making consist of specificity, sensitivity, precision, decision accuracy, and area under the receiver operating characteristic curve. Data modeling methods for GP included ridge regression, logistic regression, penalized logistic regression, and Bayesian generalized linear regression. Algorithmic modeling methods include Random Forest, Gradient Boosting Machine, Support Vector Machine, K-Nearest Neighbors, Naïve Bayes, and Artificial Neural Network. We found that a Support Vector Machine model provided the most specific decisions of correctly discarding IDC susceptible genotypes, while a Random Forest model resulted in the best decisions of retaining IDC tolerant genotypes, as well as the best outcomes when considering all decision metrics. Overall, the predictions from algorithmic modeling result in better decisions than from data modeling methods applied to soybean IDC.
Models have been developed to account for heterogeneous spatial variation in field trials. These spatial models have been shown to successfully increase the quality of phenotypic data resulting in improved effectiveness of selection by plant breeders. The models were developed for continuous data types such as grain yield and plant height, but data for most traits, such as in iron deficiency chlorosis (IDC), are recorded on ordinal scales. Is it reasonable to make spatial adjustments to ordinal data by simply applying methods developed for continuous data? The objective of the research described herein is to evaluate methods for spatial adjustment on ordinal data, using soybean IDC as an example. Spatial adjustment models are classified into three different groups: group I, moving average grid adjustment; group II, geospatial autoregressive regression (SAR) models; and group III, tensor product penalized P-splines. Comparisons of eight models sampled from these three classes demonstrate that spatial adjustments depend on severity of field heterogeneity, the irregularity of the spatial patterns, and the model used. SAR models generally produce better performance metrics than other classes of models. However, none of the eight evaluated models fully removed spatial patterns indicating that there is a need to either adjust existing models or develop novel models for spatial adjustments of ordinal data collected in fields exhibiting discontinuous transitions between heterogeneous patches.
Herein we report the impacts of applying five selection methods across 40 cycles of recurrent selection and identify interactions among factors that affect genetic responses in sets of simulated families of recombinant inbred lines derived from 21 homozygous soybean lines. Our use of recurrence equation to model response from recurrent selection allowed us to estimate the half-lives, asymptotic limits to recurrent selection for purposes of assessing the rates of response and future genetic potential of populations under selection. The simulated factors include selection methods, training sets, and selection intensity that are under the control of the plant breeder as well as genetic architecture and heritability. A factorial design to examine and analyze the main and interaction effects of these factors showed that both the rates of genetic improvement in the early cycles and limits to genetic improvement in the later cycles are significantly affected by interactions among all factors. Some consistent trends are that genomic selection methods provide greater initial rates of genetic improvement (per cycle) than phenotypic selection, but phenotypic selection provides the greatest long term responses in these closed genotypic systems. Model updating with training sets consisting of data from prior cycles of selection significantly improved prediction accuracy and genetic response with three parametric genomic prediction models. Ridge Regression, if updated with training sets consisting of data from prior cycles, achieved better rates of response than BayesB and Bayes LASSO models. A Support Vector Machine method, with a radial basis kernel, had the worst estimated prediction accuracies and the least long term genetic response. Application of genomic selection in a closed breeding population of a self-pollinated crop such as soybean will need to consider the impact of these factors on trade-offs between short term gains and conserving useful genetic diversity in the context of the goals for the breeding program.
Models have been developed to account for heterogeneous spatial variation in field trials. These spatial models have been shown to successfully increase the quality of phenotypic data resulting in improved effectiveness of selection by plant breeders. The models were developed for continuous data types such as grain yield and plant height, but data for most traits, such as in iron deficiency chlorosis (IDC), are recorded on ordinal scales. Is it reasonable to make spatial adjustments to ordinal data by simply applying methods developed for continuous data? The objective of the research described herein is to evaluate methods for spatial adjustment on ordinal data, using soybean IDC as an example. Spatial adjustment models are classified into three different groups: group I, moving average grid adjustment; group II, geospatial autoregressive regression (SAR) models; and group III, tensor product penalized splines. Comparisons of eight models sampled from these three classes demonstrate that spatial adjustments depend on severity of field heterogeneity, the irregularity of the spatial patterns, and the model used. SAR models generally produce better performance metrics than other classes of models. However, none of the eight evaluated models fully removed spatial patterns indicating that there is a need to either adjust existing models or develop novel models for spatial adjustments of ordinal data collected in fields exhibiting discontinuous transitions between heterogeneous patches.
Selection of markers linked to alleles at quantitative trait loci (QTL) for tolerance to Iron Deficiency Chlorosis (IDC) has not been successful. Genomic selection has been advocated for continuous numeric traits such as yield and plant height. For ordinal data types such as IDC, genomic prediction models have not been systematically compared. The objectives of research reported in this manuscript were to evaluate the most commonly used genomic prediction method, ridge regression and it’s equivalent logistic ridge regression method, with algorithmic modeling methods including random forest, gradient boosting, support vector machine, K-nearest neighbors, Naïve Bayes, and artificial neural network using the usual comparator metric of prediction accuracy. In addition we compared the methods using metrics of greater importance for decisions about selecting and culling lines for use in variety development and genetic improvement projects. These metrics include specificity, sensitivity, precision, decision accuracy, and area under the receiver operating characteristic curve. We found that Support Vector Machine provided the best specificity for culling IDC susceptible lines, while Random Forest GP models provided the best combined set of decision metrics for retaining IDC tolerant and culling IDC susceptible lines.
Soybean, Glycine max (L.) Merr., has been grown as a forage and as an important protein and oil crop for thousands of years. Domestication, breeding improvements and enhanced cropping systems have made soybeans the most cultivated and utilized oilseed crop globally. Soybeans provide a high-quality protein source for livestock and aquaculture, oil for industrial uses and a valued component of human diets. Originating in China and Eastern Asia, today 80-85% of the world's soybeans, approximately 88 million ha, are grown in the Western Hemisphere. United States soybean breeding and development efforts for over 80 years have transitioned from primarily universities and United States Department of Agriculture (USDA) programs to private company-led investments in commercial cultivar development. Soybean breeders continuously adapt tools and technologies that encompass classical breeding, mutation breeding and marker-assisted selection, biotechnology and transgenic approaches, gene silencing, and genome editing. In addition to breeding technologies, improved agronomics, precision agriculture and digital agriculture have advanced soybean production and profitability. The primary goals of soybean breeding and cropping systems advances include yield improvement, increased seed protein and oil composition and quality, and yield preservation through weed, pathogen, insect pest and abiotic stress resistance and management. This chapter primarily describes the introduction and improvement of soybeans in the United States. Contributing authors describe classical and molecular breeding, biotechnology, biotic and abiotic stress management, and soybean agronomics and cropping systems improvements that maximize soybean productivity, profitability and sustainability to supply a continually increasing world demand for protein and oil for feed, fuel and food.
The development of ubiquitous polymorphic genetic markers that span the genome have made it possible for quantitative and molecular geneticists to investigate what M. D. Edwards et al. referred to as the numbers, magnitudes, and distributions of quantitative trait loci (QTL). QTL are identified as significant statistical associations between genotypic values and phenotypic variability among the segregating progeny. Magnitude of QTL effects are reported as the minimum and maximum percent of phenotypic variability explained by the significant QTL. Genomic locations of the QTL with estimated large effects mapped to syntenic regions across all three genera. Identification of QTL for agronomically important traits has been pursued through progeny derived from intraspecific crosses of adapted inbred lines. There is a tendency to make the inferential leap from QTL to physiology of gene effects at genetic loci, but QTL are, by definition, merely significant statistical associations.
Maize has for many decades been both one of the most important crops worldwide and one of the primary genetic model organisms. More recently, maize breeding has been impacted by rapid technological advances in sequencing and genotyping technology, transformation including genome editing, doubled haploid technology, parallelled by progress in data sciences and the development of novel breeding approaches utilizing genomic information. Herein, we report on past, current and future developments relevant for maize breeding with regard to (1) genome analysis, (2) germplasm diversity characterization and utilization, (3) manipulation of genetic diversity by transformation and genome editing, (4) inbred line development and hybrid seed production, (5) understanding and prediction of hybrid performance, (6) breeding methodology and (7) synthesis of opportunities and challenges for future maize breeding.
Multi-objective optimization is an emerging field in mathematical optimization which involves optimization a set of objective functions simultaneously. The purpose of most plant and animal breeding programs is to make decisions that will lead to sustainable genetic gains in more than one traits while controlling the amount of co-ancestry in the breeding population. The decisions at each cycle in a breeding program involve multiple, usually competing, objectives; these complex decisions can be supported by the insights that are gained by using the multi-objective optimization principles in breeding. The discussion here includes the definition of several multi-objective optimized breeding approaches and the comparison of these approaches with the standard multi-trait breeding schemes such as tandem selection, culling and index selection. We have illustrated the newly proposed methods with two empirical data sets and with simulations.
Soybean is the world's leading source of vegetable protein and demand for its seed continues to grow. Breeders have successfully increased soybean yield, but the genetic architecture of yield and key agronomic traits is poorly understood. We developed a 40-mating soybean nested association mapping (NAM) population of 5,600 inbred lines that were characterized by single nucleotide polymorphism (SNP) markers and six agronomic traits in field trials in 22 environments. Analysis of the yield, agronomic, and SNP data revealed 23 significant marker-trait associations for yield, 19 for maturity, 15 for plant height, 17 for plant lodging, and 29 for seed mass. A higher frequency of estimated positive yield alleles was evident from elite founder parents than from exotic founders, although unique desirable alleles from the exotic group were identified, demonstrating the value of expanding the genetic base of US soybean breeding.