Characterizing population structure and admixture events between ancestral groups plays a key role in understanding the evolutionary history of species and crops. Most tools for inferring admixture have been developed for diploids and are not suitable for polyploids, in particular those with high and mixed ploidy such as Saccharum. Here we present AdmixPoly, an R-package designed to infer admixture in polyploid species both at the genome-wide scale and locally along chromosomes. We compare AdmixPoly with state-of-the-art methods using simulations, demonstrating its precision and computational efficiency. Notably, local admixture inference in complex scenarios, such as high ploidy levels, large numbers of ancestral groups and alleles per marker is enabled through efficient approximations of emission and transition probabilities within a hidden Markov model framework. We apply this approach to characterize the contributions of wild Saccharum species to the complex polyploid genome of modern sugarcane cultivars. A panel of wild and cultivated Saccharum accessions is genotyped for 80K genomic regions, each revealing approximately 50 read-scale haplotypes. The results reveal that most of the approximately 12 copies of each basic chromosome in modern cultivars are derived from the domesticated species Saccharum officinarum, with one to four copies typically contributed by distinct subgroups of the wild species Saccharum spontaneum. In addition, contributions from an unknown wild Saccharum group originating from the Pacific were identified in most cultivars. The conserved pattern of these introgressions suggests that they can be traced back to the early stages of sugarcane breeding approximately a century ago.
Banana breeding is hampered by the very low fertility of domesticated bananas and the lack of knowledge about the genetic determinism of agronomic traits. We analysed a breeding population of 2723 triploid hybrids resulting from crosses between diploid and tetraploid Musa acuminata parents, which was evaluated over three successive crop cycles for 24 traits relating to yield components and plant, bunch, and fruit architectures. A subset of 1129 individuals was genotyped by sequencing, revealing 205 612 single-nucleotide polymorphisms (SNPs). Most parents were heterozygous for one or several large reciprocal chromosomal translocations, which are known to impact recombination and chromosomal segregation. We applied two linear mixed models to detect associations between markers and traits: (i) a standard model with a kinship calculated using all SNPs and (ii) a model with chromosome-specific kinships that aims at recovering statistical power at alleles carried by long non-recombined haplotypic segments. For 23 of the 24 traits, we identified one to five significant quantitative trait loci (QTLs) for which the origin of favourable alleles could often be determined amongst the main ancestral contributors to banana cultivars. Several QTLs, located in the rearranged regions, were only detected using the second model. The resulting QTL landscape represents an important resource to support breeding programmes. The proposed strategy for recovering power at SNPs carried by long non-recombined rearranged haplotypic segments is an important methodological advance for future association studies in banana and other species affected by chromosomal rearrangements.
The advent of next-generation sequencing (NGS) has revolutionized the study of single nucleotide polymorphisms (SNPs), making it increasingly cost-effective. Haplotypes, which combine alleles from adjacent variants, offer several advantages over bi-allelic SNPs, including enhanced information content, reduced dimensionality, and improved statistical power in genomic studies. These benefits are particularly significant for polyploid species, where distinguishing all homologous copies using SNP markers alone can be challenging. This article introduces HaploCharmer, a flexible workflow designed for read-scale haplotype calling from NGS data. HaploCharmer identifies haplotypes within preconfigured genomic regions smaller than a sequencing read, ensuring direct comparability across individuals. It integrates a series of processing steps including mapping, haplotype identification, filtration, and reporting of haplotype sequences, as presence-absence, in the panel of accessions analyzed. The performance of HaploCharmer was validated by building a genetic map using whole-genome sequencing data from a highly polyploid sugarcane cultivar (R570) and its self-progeny, and performing a diversity analysis in the polyploid Saccharum genus using targeted sequencing data. The workflow successfully identified a large number of high-quality haplotypes, with less than 1% of false positives. The dense genetic map obtained using single-dose haplotypes accurately depicted the known genome architecture of the R570 cultivar, including large chromosome rearrangements. The diversity analysis accurately reflected the known genetic structure within this genus. It also allowed inferring ancestral origins to mapped haplotypes and the corresponding chromosome segments in the R570 genetic map. HaploCharmer provides a robust method for diversity, genetic mapping, and quantitative genetics studies in both diploid and polyploid species.
Sugarcane is a major crop of unclear origins due to its complex polyploid interspecific genome. We analyzed genome ancestries using whole-genome sequence data from 390 representative accessions based on repeated k-mers and chloroplast phylogeny. The results provided evidence that Saccharum officinarum was domesticated in the New Guinea region from the S. robustum wild species and revealed that its genome is a mosaic involving different S. robustum subgroups. We discovered a wild Saccharum contributor to most modern cultivars, likely originating from East Melanesia. We highlighted two early centers of sugarcane diversification associated with human transport, one in continental Asia through hybridization with different S. spontaneum subgroups and one in the Melanesian and Polynesian islands via hybridization with the discovered ancestor and Miscanthus. Finally, we revealed the genome ancestry of modern cultivars, highlighting untapped wild Saccharum diversity as a source of alleles for breeding programs.
Breeding disease-resistant cultivars that meet commercial criteria is essential to sustain banana production threatened by major diseases. Edible bananas are seedless triploid hybrids that represent end-breeding products. Hence, the crucial step in banana breeding is to improve and combine the parents. Currently, little information is available on parental combining abilities and on the inheritance of major traits to effectively guide banana breeding strategies. In this study, a breeding population of 2,723 triploid individuals resulting from multiparental diploid-tetraploid crosses was characterized during three crop cycles for 23 traits relating to plant and fruit architecture and bunch yield components. The phenotypic variance was partitioned between non-genetic and genetic effects, the latter including the general combining ability of diploid and tetraploid parents, their specific combining ability, and additional variance due to the within-cross genetic variability. Heritability was moderate to high depending on the trait and revealed the predominance of the tetraploid parent's contribution to hybrid performance for most traits. The use of parental genomic information enabled cross-mean performance prediction through genomic relationship matrices of general and specific combining abilities, the latter being partitioned into dominance and across-population epistasis contributions. Predictive abilities often greater than 0.5 were obtained, particularly when the tetraploid parent was observed in other crosses and, for some traits, when neither parent was observed. Information on trait inheritance and genomic prediction of cross-mean performance will help in selecting and combining parents, facilitating the identification of promising hybrids.
Resistance to most sugarcane diseases in modern interspecific hybrids (Saccharum spp.) is often intuitively attributed to resistance alleles that may have been derived from the wild species S. spontaneum. This intuitive breeders’ opinion stems from the fact that, for many diseases, the noble species S. officinarum is often relatively susceptible, while the wild species shows good levels of resistance and is therefore thought to be the main species that have provided improved resistance. Genomic analysis using dense SNP polymorphism mapped on a reference sugarcane genome allows the investigation of this widely held view. Genome-wide association studies (GWAS) were conducted to dissect the genetic basis of resistance to orange rust (caused by Puccinia kuehnii) and yellow leaf (caused by Sugarcane yellow leaf virus) and to detect quantitative trait loci (QTL) of resistance for both diseases. GWAS experiments were conducted on two large panels of modern interspecific clones (n = 526 and 673) using a large number of SNPs (> 180,000) distributed across the sugarcane genome. The experiments revealed four major-effect resistance QTLs for yellow leaf and five major-effect resistance QTLs for orange rust. The comparative frequency of all these resistance QTL alleles in 10 S. officinarum and in nine S. spontaneum accessions clearly support a S. spontaneum origin for most resistance alleles. Taken together, these small sets of major-effect resistance alleles are sufficient to predict the resistance levels of the surveyed clones with an interesting accuracy (0.50 for yellow leaf and 0.60 for orange rust). These novel QTL results pave the way for effective marker-assisted breeding approaches in the view of improving genetic resistance to these two important sugarcane diseases.
Six QTLs of resistance to sugarcane orange rust were identified in modern interspecific hybrids by GWAS. For five of them, the resistance alleles originated from S. spontaneum. Altogether, they efficiently predict disease resistance. Sugarcane orange rust (SOR) is a threatening emerging disease in many sugarcane industries worldwide. Improving the genetic resistance of commercial cultivars remains the most promising solution to control this disease. In this study, an association panel of 568 modern interspecific sugarcane hybrids (Saccharum officinarum x S. spontaneum) from Réunion’s breeding program was evaluated for its resistance to SOR under natural conditions of infection. Two genome-wide association studies (GWAS) were conducted between disease reactions and 183,842 single nucleotide polymorphism (SNP) markers obtained by targeted genotyping-by-sequencing. Five resistance quantitative trait loci (QTLs), named Oru1, Oru2, Oru3, Oru4 and Oru5, were identified using a single-locus GWAS (SL-GWAS). These five QTLs all originated from the species S. spontaneum. A multi-locus GWAS (ML-GWAS) uncovered an additional but less significant resistance QTL named Oru6, which originated from S. officinarum. All six QTLs had a moderate to major phenotypic effect on disease resistance. Prediction accuracy estimated with linear regression models based on each of the five QTLs identified by SL-GWAS was between 0.16-0.41. Altogether, these five QTLs provided a relatively high prediction accuracy of 0.60. In comparison, accuracies obtained with six genome-wide prediction models (i.e., GBLUP, Bayes-A, Bayes-B, Bayes-C, Bayesian Lasso and RKHS) reached only 0.65. The good prediction accuracy of disease resistance provided by the QTLs and the predominant S. spontaneum origin of their resistance alleles pave the way for effective marker-assisted breeding strategies.
The selection of highly productive genotypes with stable performance across environments is a major challenge of plant breeding programs due to genotype-by-environment (GE) interactions. Over the years, different metrics have been proposed that aim at characterizing the superiority and/or stability of genotype performance across environments. However, these metrics are traditionally estimated using phenotypic values only and are not well suited to an unbalanced design in which genotypes are not observed in all environments. The objective of this research was to propose and evaluate new estimators of the following GE metrics: Ecovalence, Environmental Variance, Finlay-Wilkinson regression coefficient, and Lin-Binns superiority measure. Drawing from a multi-environment genomic prediction model, we derived the best linear unbiased prediction for each GE metric. These derivations included both a squared expectation and a variance term. To assess the effectiveness of our new estimators, we conducted simulations that varied in traits and environment parameters. In our results, new estimators consistently outperformed traditional phenotype-based estimators in terms of accuracy. By incorporating a variance term into our new estimators, in addition to the squared expectation term, we were able to improve the precision of our estimates, particularly for Ecovalence in situations where heritability was low and/or sparseness was high. All methods are implemented in a new R-package: GEmetrics. These genomic-based estimators enable estimating GE metrics in unbalanced designs and predicting GE metrics for new genotypes, which should help improve the selection efficiency of high-performance and stable genotypes across environments.
Epistasis, commonly defined as interaction effects between alleles of different loci, is an important genetic component of the variation of phenotypic traits in natural and breeding populations. In addition to its impact on variance, epistasis can also affect the expected performance of a population and is then referred to as directional epistasis. Before the advent of genomic data, the existence of epistasis (both directional and non-directional) was investigated based on complex and expensive mating schemes involving several generations evaluated for a trait of interest. In this study, we propose a methodology to detect the presence of epistasis based on simple inbred biparental populations, both genotyped and phenotyped, ideally along with their parents. Thanks to genomic data, parental proportions as well as shared parental proportions between inbred individuals can be estimated. They allow the evaluation of epistasis through a test of the expected performance for directional epistasis or the variance of genetic values. This methodology was applied to two large multiparental populations, i.e. the American maize and soybean nested association mapping populations, evaluated for different traits. Results showed significant epistasis, especially for the test of directional epistasis, e.g. the increase in anthesis to silking interval observed in most maize inbred progenies or the decrease in grain yield observed in several soybean inbred progenies. In general, the effects detected suggested that shuffling allelic associations of both elite parents had a detrimental effect on the performance of their progeny. This methodology is implemented in the EpiTest R-package and can be applied to any bi/multiparental inbred population evaluated for a trait of interest.
The efficiency of genomic selection strongly depends on the prediction accuracy of the genetic merit of candidates. Numerous papers have shown that the composition of the calibration set is a key contributor to prediction accuracy. A poorly defined calibration set can result in low accuracies, whereas an optimized one can considerably increase accuracy compared to random sampling, for a same size. Alternatively, optimizing the calibration set can be a way of decreasing the costs of phenotyping by enabling similar levels of accuracy compared to random sampling but with fewer phenotypic units. We present here the different factors that have to be considered when designing a calibration set, and review the different criteria proposed in the literature. We classified these criteria into two groups: model-free criteria based on relatedness, and criteria derived from the linear mixed model. We introduce criteria targeting specific prediction objectives including the prediction of highly diverse panels, biparental families, or hybrids. We also review different ways of updating the calibration set, and different procedures for optimizing phenotyping experimental designs.
Spring barley (Hordeum vulgare L.) is the most important cereal in Iceland and its national breeding program aims to select barley genotypes adapted to its environment. A critical step to understand the adaptation of Nordic barley material to a cool maritime climate is to assess the genotype by environment interaction (GxE). In this study, we evaluated the yield and thousand-kernel weight (TKW) of 32 spring barley genotypes in seven Icelandic environments. We applied three methods to analyze GxE: the additive main effects and multiplicative interaction model, a factorial model, and a linear mixed model. For yield, GxE was mainly caused by a better response of six-row genotypes compared to two-row genotypes in high fertility soils. For TKW, GxE showed a pattern along a gradient of daily mean temperatures. This pattern translated into a divergent TKW response between the 2-row and 6-row genotypes, with substantial crossovers along the temperature gradient. This GxE pattern was disentangled using all three methods, illustrating the value of cross-analysis. As yield is the main trait of interest for barley cultivation in Iceland, and few crossovers of genotype performance have been observed between environments, the definition of one mega-environment was recommended for Icelandic cultivation and breeding. We identified promising genetic material for both traits and highlighted the superiority of six-row genotypes for yield.
New forms of the coefficient of determination can help to forecast the accuracy of genomic prediction and optimize experimental designs in multi-environment trials with genotype-by-environment interactions. In multi-environment trials, the relative performance of genotypes may vary depending on the environmental conditions, and this phenomenon is commonly referred to as genotype-by-environment interaction (G $$\times$$ E). With genomic prediction, G $$\times$$ E can be accounted for by modeling the genetic covariance between trials, even when the overall experimental design is highly unbalanced between trials, thanks to the genomic relationship between genotypes. In this study, we propose new forms of the coefficient of determination (CD, i.e., the expected model-based square correlation between a genetic value and its corresponding prediction) that can be used to forecast the genomic prediction reliability of genotypes, both for their trial-specific performance and their mean performance. As the expected prediction reliability based on these new CD criteria is generally a good approximation of the observed reliability, we demonstrate that they can be used to optimize multi-environment trials in the presence of G $$\times$$ E. In addition, this reliability may be highly variable between genotypes, especially in unbalanced designs with complex pedigree relationships between genotypes. Therefore, it can be useful for breeders to assess it before selecting genotypes based on their predicted genetic values. Using a wheat population evaluated both for simulated and phenology traits, and two maize populations evaluated for grain yield, we illustrate this approach and confirm the value of our new CD criteria.
The strong genetic structure observed in Mediterranean oats affects the predictive ability of genomic prediction as well as the performance of training set optimization methods. In this study, we investigated the efficiency of genomic prediction and training set optimization in a highly structured population of cultivars and landraces of cultivated oat (Avena sativa) from the Mediterranean basin, including white (subsp. sativa) and red (subsp. byzantina) oats, genotyped using genotype-by-sequencing markers and evaluated for agronomic traits in Southern Spain. For most traits, the predictive abilities were moderate to high with little differences between models, except for biomass for which Bayes-B showed a substantial gain compared to other models. The consistency between the structure of the training population and the population to be predicted was key to the predictive ability of genomic predictions. The predictive ability of inter-subspecies predictions was indeed much lower than that of intra-subspecies predictions for all traits. Regarding training set optimization, the linear mixed model optimization criteria (prediction error variance (PEVmean) and coefficient of determination (CDmean)) performed better than the heuristic approach “partitioning around medoids,” even under high population structure. The superiority of CDmean and PEVmean could be explained by their ability to adapt the representation of each genetic group according to those represented in the population to be predicted. These results represent an important step towards the implementation of genomic prediction in oat breeding programs and address important issues faced by the genomic prediction community regarding population structure and training set optimization.
A major barrier to the wider use of supervised learning in emerging applications, such as genomic selection, is the lack of sufficient and representative labeled data to train prediction models. The amount and quality of labeled training data in many applications is usually limited and therefore careful selection of the training examples to be labeled can be useful for improving the accuracies in predictive learning tasks. In this paper, we present an R package, TrainSel, which provides flexible, efficient, and easy-to-use tools that can be used for the selection of training populations (STP). We illustrate its use, performance, and potentials in four different supervised learning applications within and outside of the plant breeding area.
When handling a structured population in association mapping, group-specific allele effects may be observed at quantitative trait loci (QTLs) for several reasons: (i) a different linkage disequilibrium (LD) between SNPs and QTLs across groups, (ii) group-specific genetic mutations in QTL regions, and/or (iii) epistatic interactions between QTLs and other loci that have differentiated allele frequencies between groups. We present here a new genome-wide association (GWAS) approach to identify QTLs exhibiting such group-specific allele effects. We developed genetic materials including admixed progeny from different genetic groups with known genome-wide ancestries (local admixture). A dedicated statistical methodology was developed to analyze pure and admixed individuals jointly, allowing one to disentangle the factors causing the heterogeneity of allele effects across groups. This approach was applied to maize by developing an inbred "Flint-Dent" panel including admixed individuals that was evaluated for flowering time. Several associations were detected revealing a wide range of configurations of allele effects, both at known flowering QTLs (Vgt1, Vgt2 and Vgt3) and new loci. We found several QTLs whose effect depended on the group ancestry of alleles while others interacted with the genetic background. Our GWAS approach provides useful information on the stability of QTL effects across genetic groups and can be applied to a wide range of species.