Brassica oleracea forms a diverse and economically significant crop group. Improvement efforts are often hindered by limited knowledge of diversity contained within available germplasm. Here, we employ genotyping-by-sequencing to investigate a diverse panel of 85 landrace and improved B. oleracea broccoli, cauliflower, and Chinese kale entries. Ultimately, 21,680 high-quality SNPs were used to reveal a complex and admixed population structure and clarify phylogenetic relationships among B. oleracea groups. Each broccoli landrace contained, on average, 8.4 times as many unique alleles as an improved broccoli and landraces collectively represented 81% of all broccoli-specific alleles. Commercial broccoli hybrids were largely represented by a single subpopulation identified within a complex population structure. Greater allelic diversity in landrace broccoli and 96.1% of SNPs differentiating improved cauliflower from landrace cauliflower were common to the larger pool of broccoli germplasm, supporting a parallel or later development of cauliflower due to introgression events from broccoli. Chinese kale was readily distinguished by principal coordinate analysis. Genotyping was accomplished with and without reliance upon a reference genome producing 141,317 and 20,815 filtered SNPs, respectively, supporting robust SNP discovery methods in neglected or unimproved crop groups that lack a reference genome. This work clarifies the population structure, phylogeny, and domestication footprints of landrace and improved B. oleracea broccoli using many genotyping-by-sequencing markers. Additionally, a large pool of genetic diversity contained in broccoli landraces is described which may enhance future breeding efforts.
A comprehensive and meaningful phylogenetic hypothesis for the commercially important coffee genus (Coffea) has long been a key objective for coffee researchers. For molecular studies, progress has been limited by low levels of sequence divergence, leading to insufficient topological resolution and statistical support in phylogenetic trees, particularly for the major lineages and for the numerous species occurring in Madagascar. We report here the first almost fully resolved, broadly sampled phylogenetic hypothesis for coffee, the result of combining genotyping-by-sequencing (GBS) technology with a newly developed, lab-based workflow to integrate short read next-generation sequencing for low numbers of additional samples. Biogeographic patterns indicate either Africa or Asia (or possibly the Arabian Peninsula) as the most likely ancestral locality for the origin of the coffee genus, with independent radiations across Africa, Asia, and the Western Indian Ocean Islands (including Madagascar and Mauritius). The evolution of caffeine, an important trait for commerce and society, was evaluated in light of our phylogeny. High and consistent caffeine content is found only in species from the equatorial, fully humid environments of West and Central Africa, possibly as an adaptive response to increased levels of pest predation. Moderate caffeine production, however, evolved at least one additional time recently (between 2 and 4Mya) in a Madagascan lineage, which suggests that either the biosynthetic pathway was already in place during the early evolutionary history of coffee, or that caffeine synthesis within the genus is subject to convergent evolution, as is also the case for caffeine synthesis in coffee versus tea and chocolate.
Genotyping-by-sequencing (GBS) refers to a suite of related methods that obtain genotype data from samples by using restriction enzyme digestion followed by high-throughput sequencing. GBS is a refinement of restriction site-associated DNA sequencing (RADseq) methods, with a goal of being able to perform library preparations quickly, cost-effectively, and in a high-throughput manner. This protocol contains the steps necessary to go from purified DNA to Illumina-ready libraries. It also covers the considerations that go into planning a GBS experiment. © 2017 by John Wiley & Sons, Inc.
BACKGROUND:Sorghum is an important C4 crop which relies on applied Nitrogen fertilizers (N) for optimal yields, of which substantial amounts are lost into the atmosphere. Understanding the genetic variation of sorghum in response to limited nitrogen supply is important for elucidating the underlying genetic mechanisms of nitrogen utilization.RESULTS:A bi-parental mapping population consisting of 131 recombinant inbred lines (RILs) was used to map quantitative trait loci (QTLs) influencing different agronomic traits evaluated under normal N (100 kg.ha(-1) fertilizer) and low N (0 kg.ha(-1) fertilizer) conditions. A linkage map spanning 1614 cM was developed using 642 polymorphic single nucleotide polymorphisms (SNPs) detected in the population using Genotyping-By-Sequencing (GBS) technology. Composite interval mapping detected a total of 38 QTLs for 11 agronomic traits tested under different nitrogen levels. The phenotypic variation explained by individual QTL ranged from 6.2 to 50.8%. Illumina RNA sequencing data generated on seedling root tissues revealed 726 differentially expressed gene (DEG) transcripts between parents, of which 108 were mapped close to the QTL regions.CONCLUSIONS:Co-localized regions affecting multiple traits were detected on chromosomes 1, 5, 6, 7 and 9. These potentially pleiotropic regions were coincident with the genomic regions of cloned QTLs, including genes associated with flowering time, Ma3 on chromosome 1 and Ma1 on chromosome 6, gene associated with plant height, Dw2 on chromosome 6. In these regions, RNA sequencing data showed differential expression of transcripts related to nitrogen metabolism (Ferredoxin-nitrate reductase), glycolysis (Phosphofructo-2-kinase), seed storage proteins, plant hormone metabolism and membrane transport. The differentially expressed transcripts underlying the pleiotropic QTL regions could be potential targets for improving sorghum performance under limited N fertilizer through marker assisted selection.
The 7.4 million plant accessions in gene banks are largely underutilized due to various resource constraints, but current genomic and analytic technologies are enabling us to mine this natural heritage. Here we report a proof-of-concept study to integrate genomic prediction into a broad germplasm evaluation process. First, a set of 962 biomass sorghum accessions were chosen as a reference set by germplasm curators. With high throughput genotyping-by-sequencing (GBS), we genetically characterized this reference set with 340,496 single nucleotide polymorphisms (SNPs). A set of 299 accessions was selected as the training set to represent the overall diversity of the reference set, and we phenotypically characterized the training set for biomass yield and other related traits. Cross-validation with multiple analytical methods using the data of this training set indicated high prediction accuracy for biomass yield. Empirical experiments with a 200-accession validation set chosen from the reference set confirmed high prediction accuracy. The potential to apply the prediction model to broader genetic contexts was also examined with an independent population. Detailed analyses on prediction reliability provided new insights into strategy optimization. The success of this project illustrates that a global, cost-effective strategy may be designed to assess the vast amount of valuable germplasm archived in 1,750 gene banks.
Pearl millet [ (L.) R. Br; also (L.) Morrone] is an important crop throughout the world but better genomic resources for this species are needed to facilitate crop improvement. Genome mapping studies are a prerequisite for tagging agronomically important traits. Genotyping-by-sequencing (GBS) markers can be used to build high-density linkage maps, even in species lacking a reference genome. A recombinant inbred line (RIL) mapping population was developed from a cross between the lines 'Tift 99DB' and 'Tift 454'. DNA from 186 RILs, the parents, and the F was used for 96-plex KI GBS library development, which was further used for sequencing. The sequencing results showed that the average number of good reads per individual was 2.2 million, the pass filter rate was 88%, and the CV was 43%. High-quality GBS markers were developed with stringent filtering on sequence data from 179 RILs. The reference genetic map developed using 150 RILs contained 16,650 single-nucleotide polymorphisms (SNPs) and 333,567 sequence tags spread across all seven chromosomes. The overall average density of SNP markers was 23.23 SNP/cM in the final map and 1.66 unique linkage bins per cM covering a total genetic distance of 716.7 cM. The linkage map was further validated for its utility by using it in mapping quantitative trait loci (QTLs) for flowering time and resistance to leaf spot [ (Cke.) Sacc.]. This map is the densest yet reported for this crop and will be a valuable resource for the pearl millet community.
Molecular breeding can complement traditional breeding approaches to achieve genetic gains in a more efficient way. In the present study, genetic mapping was conducted in a sorghum recombinant inbred line (RIL) population developed from Tx436 (a non‐stay‐green high food quality inbred) × 00MN7645 (a stay‐green high yield inbred) and evaluated in eight environments (location and year combination) in a hybrid background of Tx3042 (a non‐stay‐green A‐line). Phenotyping was conducted for agronomic traits (grain yield and flowering time), physiological traits of stay‐green (chlorophyll content [SPAD] and chlorophyll fluorescence [Fv/Fm] measured on the leaves), and green leaf area visual score (GLAVS). This population was genotyped with genotyping‐by‐sequencing (GBS) technology. Data processing resulted in 7144 high quality single nucleotide polymorphisms (SNPs) that were used in a genome‐wide single marker scan with physical distance. A selected subset of 1414 SNPs was used for composite interval mapping (CIM) with genetic distance. These complementary methods revealed fifteen QTLs for the traits studied. In addition, QTL mapping for individual environments and year‐wise combinations revealed 42 QTLs. A consistent QTL for grain yield under normal and stressed conditions was identified in chromosome 1 that explained 8 to 16% of the phenotypic variation. QTLs for flowering time were identified in chromosomes 2, 6, and 9 that explained 6 to 11% of the phenotypic variation. Stay‐green QTLs in chromosomes 3 and 4 explained 8 to 24% of the phenotypic variation. These identified QTLs with flanking SNPs of known genomic positions could be used to improve grain yield, flowering time, and stay‐green in sorghum molecular breeding programs.
Repetitive sequences have been used for DNA fingerprinting and genotyping for more than a quarter century. Now, with our knowledge of whole genome sequences, repetitive sequences can be used to identify polymorphisms that can be mapped and scored in a systematic manner. We have developed a simple, robust platform for designing primers, PCR amplification, and high throughput cloning that allows hundreds to thousands of markers to be scored for less than $5 per sample. Conserved regions were used to design PCR primers for amplifying thousands of middle repetitive regions of the maize ( Zea mays ssp. mays ) genome. Bioinformatic scans were then used to identify DNA sequence polymorphisms in the low copy intervening sequences. When used in conjunction with simple DNA preps, optimized PCR conditions, high multiplex Illumina indexing and a bioinformatic marker calling platform tailored for repetitive sequences, this methodology provides a cost effective genotyping strategy for large-scale genomic selection projects. We show detailed results from four maize primer sets that produced between 1,335-3,225 good coverage loci with 1056 that segregated appropriately in a bi-parental family. This approach could have wide applicability to breeding and conservation biology, where hundreds of thousands of samples need to be genotyped for very minimal cost.
The silver fox (Vulpes vulpes) offers a novel model for studying the genetics of social behavior and animal domestication. Selection of foxes, separately, for tame and for aggressive behavior has yielded two strains with markedly different, genetically determined, behavioral phenotypes. Tame strain foxes are eager to establish human contact while foxes from the aggressive strain are aggressive and difficult to handle. These strains have been maintained as separate outbred lines for over 40 generations but their genetic structure has not been previously investigated. We applied a genotyping-by-sequencing (GBS) approach to provide insights into the genetic composition of these fox populations. Sequence analysis of EcoT22I genomic libraries of tame and aggressive foxes identified 48,294 high quality SNPs. Population structure analysis revealed genetic divergence between the two strains and more diversity in the aggressive strain than in the tame one. Significant differences in allele frequency between the strains were identified for 68 SNPs. Three of these SNPs were located on fox chromosome 14 within an interval of a previously identified behavioral QTL, further supporting the importance of this region for behavior. The GBS SNP data confirmed that significant genetic diversity has been preserved in both fox populations despite many years of selective breeding. Analysis of SNP allele frequencies in the two populations identified several regions of genetic divergence between the tame and aggressive foxes, some of which may represent targets of selection for behavior. The GBS protocol used in this study significantly expanded genomic resources for the fox, and can be adapted for SNP discovery and genotyping in other canid species.
Among the fundamental evolutionary forces, recombination arguably has the largest impact on the practical work of plant breeders. Varying over 1,000-fold across the maize genome, the local meiotic recombination rate limits the resolving power of quantitative trait mapping and the precision of favorable allele introgression. The consequences of low recombination also theoretically extend to the species-wide scale by decreasing the power of selection relative to genetic drift, and thereby hindering the purging of deleterious mutations. In this study, we used genotyping-by-sequencing (GBS) to identify 136,000 recombination breakpoints at high resolution within US and Chinese maize nested association mapping populations. We find that the pattern of cross-overs is highly predictable on the broad scale, following the distribution of gene density and CpG methylation. Several large inversions also suppress recombination in distinct regions of several families. We also identify recombination hotspots ranging in size from 1 kb to 30 kb. We find these hotspots to be historically stable and, compared with similar regions with low recombination, to have strongly differentiated patterns of DNA methylation and GC content. We also provide evidence for the historical action of GC-biased gene conversion in recombination hotspots. Finally, using genomic evolutionary rate profiling (GERP) to identify putative deleterious polymorphisms, we find evidence for reduced genetic load in hotspot regions, a phenomenon that may have considerable practical importance for breeding programs worldwide.
Genotyping by sequencing (GBS) provides opportunities to generate high-resolution genetic maps at a low genotyping cost, but for highly heterozygous species, missing data and heterozygote undercalling complicate the creation of GBS genetic maps. To overcome these issues, we developed a publicly available, modular approach called HetMappS, which functions independently of parental genotypes and corrects for genotyping errors associated with heterozygosity. For linkage group formation, HetMappS includes both a reference-guided synteny pipeline and a reference-independent de novo pipeline. The de novo pipeline can be utilized for under-characterized or high diversity families that lack an appropriate reference. We applied both HetMappS pipelines in five half-sib F1 families involving genetically diverse Vitis spp. Starting with at least 116,466 putative SNPs per family, the HetMappS pipelines identified 10,440 to 17,267 phased pseudo-testcross (Pt) markers and generated high-confidence maps. Pt marker density exceeded crossover resolution in all cases; up to 5,560 non-redundant markers were used to generate parental maps ranging from 1,047 cM to 1,696 cM. The number of markers used was strongly correlated with family size in both de novo and synteny maps (r = 0.92 and 0.91, respectively). Comparisons between allele and tag frequencies suggested that many markers were in tandem repeats and mapped as single loci, while markers in regions of more than two repeats were removed during map curation. Both pipelines generated similar genetic maps, and genetic order was strongly correlated with the reference genome physical order in all cases. Independently created genetic maps from shared parents exhibited nearly identical results. Flower sex was mapped in three families and correctly localized to the known sex locus in all cases. The HetMappS pipeline could have wide application for genetic mapping in highly heterozygous species, and its modularity provides opportunities to adapt portions of the pipeline to other family types, genotyping technologies or applications.
Improving environmental adaptation in crops is essential for food security under global change, but phenotyping adaptive traits remains a major bottleneck. If associations between single-nucleotide polymorphism (SNP) alleles and environment of origin in crop landraces reflect adaptation, then these could be used to predict phenotypic variation for adaptive traits. We tested this proposition in the global food crop Sorghum bicolor , characterizing 1943 georeferenced landraces at 404,627 SNPs and quantifying allelic associations with bioclimatic and soil gradients. Environment explained a substantial portion of SNP variation, independent of geographical distance, and genic SNPs were enriched for environmental associations. Further, environment-associated SNPs predicted genotype-by-environment interactions under experimental drought stress and aluminum toxicity. Our results suggest that genomic signatures of environmental adaptation may be useful for crop improvement, enhancing germplasm identification and marker-assisted selection. Together, genome-environment associations and phenotypic analyses may reveal the basis of environmental adaptation.
BACKGROUND:Genetic diversity provides the capacity for plants to meet changing environments. It is fundamentally important in crop improvement. Fifty-nine local maize lines developed at INERA and 41 exotic (temperate and tropical) inbred lines were characterized using 1057 SNP markers to (1) analyse the genetic diversity in a diverse set of maize inbred lines; (2) determine the level of genetic diversity in INERA inbred lines and patterns of relationships of these inbred lines developed from two sources; and (3) examine the genetic differences between local and exotic germplasms.RESULTS:Roger's genetic distance for about 64% of the pairs of lines fell between 0.300 and 0.400. Sixty one per cent of the pairs of lines also showed relative kinship values of zero. Model-based population structure analysis and principal component analysis revealed the presence of 5 groups that agree, to some extent, with the origin of the germplasm. There was genetic diversity among INERA inbred lines, which were genetically less closely related and showed a low level of heterozygosity. These lines could be divided into 3 major distinct groups and a mixed group consistent with the source population of the lines. Pairwise comparisons between local and exotic germplasms showed that the temperate and some IITA lines were differentiated from INERA lines. There appeared to be substantial levels of genetic variation between local and exotic germplasms as revealed by missing and unique alleles.CONCLUSIONS:Allelic frequency differences observed between the germplasms, together with unique alleles identified within each germplasm, shows the potential for a mutual improvement between the sets of germplasm. The results from this study will be useful to breeders in designing inbred-hybrid breeding programs, association mapping population studies and marker assisted breeding.
Next‐generation sequencing technology such as genotyping‐by‐sequencing (GBS) made low‐cost, but often low‐coverage, whole‐genome sequencing widely available. Extensive inbreeding in crop plants provides an untapped, high quality source of phased haplotypes for imputing missing genotypes. We introduce Full‐Sib Family Haplotype Imputation (FSFHap), optimized for full‐sib populations, and a generalized method, Fast Inbred Line Library ImputatioN (FILLIN), to rapidly and accurately impute missing genotypes in GBS‐type data with ordered markers. FSFHap and FILLIN impute missing genotypes with high accuracy in GBS‐genotyped maize (Zea mays L.) inbred lines and breeding populations, while Beagle v. 4 is still preferable for diverse heterozygous populations. FILLIN and FSFHap are implemented in TASSEL 5.0.
Genotyping by sequencing (GBS) is used to understand the origin and domestication of guinea yams, including the contribution of wild relatives and polyploidy events to the cultivated guinea yams.
Background A large single nucleotide polymorphism (SNP) dataset was used to analyze genome-wide diversity in a diverse collection of watermelon cultivars representing globally cultivated, watermelon genetic diversity. The marker density required for conducting successful association mapping depends on the extent of linkage disequilibrium (LD) within a population. Use of genotyping by sequencing reveals large numbers of SNPs that in turn generate opportunities in genome-wide association mapping and marker-assisted selection, even in crops such as watermelon for which few genomic resources are available. In this paper, we used genome-wide genetic diversity to study LD, selective sweeps, and pairwise F ST distributions among worldwide cultivated watermelons to track signals of domestication. Results We examined 183 Citrullus lanatus var. lanatus accessions representing domesticated watermelon and generated a set of 11,485 SNP markers using genotyping by sequencing. With a diverse panel of worldwide cultivated watermelons, we identified a set of 5,254 SNPs with a minor allele frequency of ≥ 0.05, distributed across the genome. All ancestries were traced to Africa and an admixture of various ancestries constituted secondary gene pools across various continents. A sliding window analysis using pairwise F ST values was used to resolve selective sweeps. We identified strong selection on chromosomes 3 and 9 that might have contributed to the domestication process. Pairwise analysis of adjacent SNPs within a chromosome as well as within a haplotype allowed us to estimate genome-wide LD decay. LD was also detected within individual genes on various chromosomes. Principal component and ancestry analyses were used to account for population structure in a genome-wide association study. We further mapped important genes for soluble solid content using a mixed linear model. Conclusions Information concerning the SNP resources, population structure, and LD developed in this study will help in identifying agronomically important candidate genes from the genomic regions underlying selection and for mapping quantitative trait loci using a genome-wide association study in sweet watermelon.
Genome-wide association studies are a powerful method to dissect the genetic basis of traits, although in practice the effects of complex genetic architecture and population structure remain poorly understood. To compare mapping strategies we dissected the genetic control of flavonoid pigmentation traits in the cereal grass sorghum by using high-resolution genotyping-by-sequencing single-nucleotide polymorphism markers. Studying the grain tannin trait, we find that general linear models (GLMs) are not able to precisely map tan1-a, a known loss-of-function allele of the Tannin1 gene, with either a small panel (n = 142) or large association panel (n = 336), and that indirect associations limit the mapping of the Tannin1 locus to Mb-resolution. A GLM that accounts for population structure (Q) or standard mixed linear model that accounts for kinship (K) can identify tan1-a, whereas a compressed mixed linear model performs worse than the naive GLM. Interestingly, a simple loss-of-function genome scan, for genotype-phenotype covariation only in the putative loss-of-function allele, is able to precisely identify the Tannin1 gene without considering relatedness. We also find that the tan1-a allele can be mapped with gene resolution in a biparental recombinant inbred line family (n = 263) using genotyping-by-sequencing markers but lower precision in the mapping of vegetative pigmentation traits suggest that consistent gene-level resolution will likely require larger families or multiple recombinant inbred lines. These findings highlight that complex association signals can emerge from even the simplest traits given epistasis and structured alleles, but that gene-resolution mapping of these traits is possible with high marker density and appropriate models.
BACKGROUND:Genotyping by sequencing, a new low-cost, high-throughput sequencing technology was used to genotype 2,815 maize inbred accessions, preserved mostly at the National Plant Germplasm System in the USA. The collection includes inbred lines from breeding programs all over the world.RESULTS:The method produced 681,257 single-nucleotide polymorphism (SNP) markers distributed across the entire genome, with the ability to detect rare alleles at high confidence levels. More than half of the SNPs in the collection are rare. Although most rare alleles have been incorporated into public temperate breeding programs, only a modest amount of the available diversity is present in the commercial germplasm. Analysis of genetic distances shows population stratification, including a small number of large clusters centered on key lines. Nevertheless, an average fixation index of 0.06 indicates moderate differentiation between the three major maize subpopulations. Linkage disequilibrium (LD) decays very rapidly, but the extent of LD is highly dependent on the particular group of germplasm and region of the genome. The utility of these data for performing genome-wide association studies was tested with two simply inherited traits and one complex trait. We identified trait associations at SNPs very close to known candidate genes for kernel color, sweet corn, and flowering time; however, results suggest that more SNPs are needed to better explore the genetic architecture of complex traits.CONCLUSIONS:The genotypic information described here allows this publicly available panel to be exploited by researchers facing the challenges of sustainable agriculture through better knowledge of the nature of genetic diversity.
High-throughput genotyping methods have increased the analytical power to study complex traits but high cost has remained a barrier for large scale use in animal improvement. We have adapted genotyping-by-sequencing (GBS) used in plants for genotyping 47 animals representing 7 taurine and indicine breeds of cattle from the US and Africa. Genomic DNA was digested with different enzymes, ligated to adapters containing one of 48 unique bar codes and sequenced by the Illumina HiSeq 2000. PstI was the best enzyme producing 1.4 million unique reads per animal and initially identifying a total of 63,697 SNPs. After removal of SNPs with call rates of less than 70%, 51,414 SNPs were detected throughout all autosomes with an average distance of 48.1 kb, and 1,143 SNPs on the X chromosome at an average distance of 130.3 kb, as well as 191 on unmapped contigs. If we consider only the SNPs with call rates of 90% and over, we identified 39,751 on autosomes, 850 on the X chromosome and 124 on unmapped contigs. Of these SNPs, 28,843 were not tightly linked to other SNPs. Average marker density per autosome was highly correlated with chromosome size (coefficient of correlation = -0.798, r(2) = 0.637) with higher density in smaller chromosomes. Average SNP call rate was 86.5% for all loci, with 53.0% of the loci having call rates >90% and the average minor allele frequency being 0.212. Average observed heterozygosity ranged from 0.046-0.294 among individuals, and from 0.064-0.197 among breeds, with Brangus showing the highest diversity as expected. GBS technique is novel, flexible, sufficiently high-throughput, and capable of providing acceptable marker density for genomic selection or genome-wide association studies at roughly one third of the cost of currently available genotyping technologies.