Motivation:Genotyping datasets generated via the Thermo Fisher Axiom® array are generally big, as they comprise tens of thousands of markers and hundreds of individuals, and currently, no automatic data curation pipelines are available for this kind of data. This leaves researchers with only time-consuming manual analysis as the current standard for processing these complex genotyping datasets. There is a clear need for a more efficient, streamlined approach to handle the specific quality control challenges inherent in this platform. Results:AxioSAFE (Axiom SNP Assessment and Filtering Engine) is a semi-automatic computer tool for the curation of single nucleotide polymorphism (SNP) genotyping datasets generated via Thermo Fisher Axiom® array experiments. AxioSAFE provides an alternative methodology to cover a set of data curation operations, including steps such as a ploidy check, SNP filtering, Mendelian error analysis, and phasing. AxioSAFE identifies major occurrences of problematic SNPs and samples, including those not caught by the Axiom array default QC filters. Further functionality is included to let the user review identified problematic SNP classes. Availability and implementation:AxioSAFE is a Python program that can be either used via the command line interface or through a graphical user interface (GUI) and is provided as a Docker container available on DockerHub at https://hub.docker.com/r/lzspin/axiosafe, which includes all required libraries, software, and a tutorial dataset. The source code and documentation are available at https://bitbucket.org/lzspin/axiosafe/. The apple dataset used for the development of AxioSAFE is available at DOI: https://doi.org/10.5281/zenodo.18034024.
Understanding the quality of a whole genome sequence (WGS) is important for its further use. Most WGS quality evaluations are based on bioinformatic quality metrics such as the N50 score, BUSCO score, and number of contigs and scaffolds present, yet genetic information considering principles of inheritance could be used to evaluate and improve assembly and phasing. Furthermore, WGS and genome resequencing data of related individuals could provide useful information when large chromosomal segments are shared with the target individual through common ancestry. Here, we show how high-quality, phased, genome-wide genotypic information is useful to evaluate the quality of a WGS. We provide an R-tool to routinely conduct such quality evaluations. The script also provides a method to accurately determine the WGS positions of reference SNP markers, which is needed for integration of SNP array-based genotypic data sets with WGS data, and the identification and comparison of segments across WGSs that are shared by descent. Finally, we provide suggestions on how such sharing can be used to evaluate and improve new WGSs. The approach is demonstrated in apple, for which improvements in WGS quality are evident from the first collapsed WGS with many inconsistencies in genetic marker order and genotype scores, through well-assembled haploid WGSs, to incorrectly and correctly phased diploid WGSs. This study shows that homozygous regions might need extra attention in phased WGSs and that further improvements to phased WGSs can be achieved by grouping chromosomes of single parental origin into the same haplome. ### Competing Interest Statement The authors have declared no competing interest.
Heirloom Danish apple cultivars are historically and pomologically important, part of the cultural heritage, and have valuable adaptation to regional climate conditions. However, lack of information about their genetic identity and pedigree relatedness with other cultivars hampers proper cultivar identification, germplasm curation, genebank management, and future regional breeding efforts. Many Danish apple cultivars are maintained in the national collection “The Pometum”, maintaining around 850 apple accessions. Additional material is maintained in public or private Danish collections. However, no information exists regarding genotypic duplicates between these collections and germplasm collections in other countries, pedigree inferences across collections, and genotypically unique accessions at the genebank level. To provide such information, 976 accessions from Denmark were genotyped with simple sequence repeat (SSR) markers and the Illumina Infinium 20K single nucleotide polymorphism (SNP) array. The resulting genotypic data were compared to large databases of genotypic data from germplasm collections in multiple countries to identify genotypic duplicates and conduct pedigree reconstruction. The germplasm maintains 305 unique genotypic profiles which were not found in other germplasm collections. The study exposed previously unknown synonyms, accessions not true-to-type, and novel pedigree relationships involving accessions from multiple collection sites. The most frequent parents of Danish germplasm were ‘Hvid Vinter Pigeon’ and ‘Cox’s Orange Pippin’ whereas ‘Reinette Franche’ was the most common grandparent. The accession-level information will benefit germplasm curation, cultivar identification, genebank management, and future breeding efforts, and shed new light on cultivar history and origin.
Resistance to European canker ( Neonectria ditissima ) in apple is currently one of the most important breeding targets for commercial production in Sweden. Previous research has identified significant genetic variation in susceptibility to the disease, with the local Swedish cultivar ‘Aroma’ considered as one of the most resistant cultivars. Identification of genetic regions underlying the resistance of this cultivar would be a valuable tool for future breeding. Thus, we performed Bayesian quantitative trait loci (QTL) mapping for resistance to European canker in a full-sib family of ‘Aroma’ × ‘Discovery’. Mapping was performed with the area under the disease progression curves (AUDPCs) from all seven (AUDPC_All7) and the first four assessments (AUDPC_First4), and three parameters of a sigmoid growth model for lesion length. As a scale for the effect of the different parameters, historic phenotypic data from screenings of a genetically diverse germplasm was compiled and re-analyzed. The parametrization of the data on lesion growth increased the number of QTL that could be identified with high statistical power, and provided some insight into their roles during different stages of disease development in the current experimental setup. Five QTL regions with strong or decisive evidence were identified on linkage groups 1, 8, 15, and 16. The QTL regions could be assigned to either of the parameters lesion length at the first assessment (‘LL_A1’), the maximal lesion growth rate (lesion length doubling time, ‘t_gen’), and the lesion length at girdling (‘LL_G’). Three of these QTL were traced along the pedigrees of some known relatives of the FS family, and discussed in relation to future crosses for breeding and genetic research.
High-throughput and high-density (HD) genetic marker genotyping systems are critical to optimize the efficiency of oil palm breeding and improvement programmes. This study reports the development of the 78 K Infinium (R) HD customized SNP array, which was used to genotype a thousand palms of a commercial Deli dura x AVROS pisifera family. A total 64,108 (similar to 82.0%) polymorphic SNP markers were identified of which, 57, 465 (89.6%) were mapped onto the genetic map that has the largest number of markers published so far in oil palm and holds 14,781 SNPs on 2,363 orphan scaffolds (whose chromosomal locations were unknown), which will improve the existing oil palm reference genome (EG5.1). The SNPs were highly informative based on the parent-to-progeny allelic inheritance analysis. The data demonstrated that 4.3% of the progeny resulted from unintentional self-fertilization of the dura female parent. These unintended 'selfs' are highly inbred, which will affect their yield. This study also for the first time, describes the homozygosity in the Deli dura and AVROS pisifera, two important parental lines widely used in commercial seed production. As expected, both the parental palms were highly homozygous, having 138 Mb homozygous regions in common, with 70.3% identical alleles. Such a detailed genetic analysis of the individual palm has been made possible with this customized HD SNP array, which will be a valuable tool for routine application in oil palm improvement programmes. The strategy used to design and apply the array will also be of interest for wider scientific research.
Linkage mapping is an approach to order markers based on recombination events. Mapping algorithms cannot easily handle genotyping errors, which are common in high-throughput genotyping data. To solve this issue, strategies have been developed, aimed mostly at identifying and eliminating these errors. One such strategy is SMOOTH, an iterative algorithm to detect genotyping errors. Unlike other approaches, SMOOTH can also be used to impute the most probable alternative genotypes, but its application is limited to diploid species and to markers heterozygous in only one of the parents. In this study we adapted SMOOTH to expand its use to any marker type and to autopolyploids with the use of identity-by-descent probabilities, naming the updated algorithm Smooth Descent (SD). We applied SD to real and simulated data, showing that in the presence of genotyping errors this method produces better genetic maps in terms of marker order and map length. SD is particularly useful for error rates between 5% and 20% and when error rates are not homogeneous among markers or individuals. With a starting error rate of 10%, SD reduced it to ∼5% in diploids, ∼7% in tetraploids and ∼8.5% in hexaploids. Conversely, the correlation between true and estimated genetic maps increased by 0.03 in tetraploids and by 0.2 in hexaploids, while worsening slightly in diploids (∼0.0011). We also show that the combination of genotype curation and map re-estimation allowed us to obtain better genetic maps while correcting wrong genotypes. We have implemented this algorithm in the R package Smooth Descent.
Quantitative trait locus (QTL) analysis allows to identify regions responsible for a trait and to associate alleles with their effect on phenotypes. When using biallelic markers to find these QTL regions, two alleles per QTL are modelled. This assumption might be close to reality in specific biparental crosses but is unrealistic in situations where broader genetic diversity is studied. Diversity panels used in genome-wide association studies or multi-parental populations can easily harbour multiple QTL alleles at each locus, more so in the case of polyploids that carry more than two alleles per individual. In such situations a multiallelic model would be closer to reality, allowing for different genetic effects for each potential allele in the population. To obtain such multiallelic markers we propose the usage of haplotypes, concatenations of nearby SNPs. We developed “mpQTL” an R package that can perform a QTL analysis at any ploidy level under biallelic and multiallelic models, depending on the marker type given. We tested the effect of genetic diversity on the power and accuracy difference between bi-allelic and multiallelic models using a set of simulated multiparental autotetraploid, outbreeding populations. Multiallelic models had higher detection power and were more precise than biallelic, SNP-based models, particularly when genetic diversity was higher. This confirms that moving to multi-allelic QTL models can lead to improved detection and characterization of QTLs. Key message QTL detection in populations with more than two functional QTL alleles (which is likely in multiparental and/or polyploid populations) is more powerful when using multiallelic models, rather than biallelic models.
Unordered parent-offspring(PO)relationships are an outstanding issue in pedigree reconstruction studies.Resolution of the order of these relationships would expand the results,conclusions,and usefulness of such studies;however,no such PO order resolution(POR)tests currently exist.This study describes two such tests,demonstrated using SNP array data in the outcrossing species apple(Malus×domestica)on a PO relationship of known order('Keepsake'as a parent of'Honeycrisp')and two PO relationships previously ordered only via provenance information.The first test,POR-1,tests whether some of the extended haplotypes deduced from homozygous SNP calls from one individual in an unordered PO duo are composed of recombinant haplotypes from accurately phased SNP genotypes from the second individual.If so,the first individual would be the offspring of the second individual,otherwise the opposite relationship would be present.The second test,POR-2,does not require phased SNP genotypes and uses similar logic as the POR-1 test,albeit in a different approach.The POR-1 and POR-2 tests determined the correct relationship between'Keepsake'and'Honeycrisp'.The POR-2 test confirmed'Reinette Franche'as a parent of'Nonpareil'and'Brabant Bellefleur'as a parent of'Court Pendu Plat'.The latter finding conflicted with the recorded provenance information,demonstrating the need for these tests.The successful demonstration of these tests suggests they can add insights to future pedigree reconstruction studies,though caveats,like extreme inbreeding or selfing,would need to be considered where relevant.
Pedigree information is of fundamental importance in breeding programs and related genetics efforts. However, many individuals have unknown pedigrees. While methods to identify and confirm direct parent–offspring relationships are routine, those for other types of close relationships have yet to be effectively and widely implemented with plants, due to complications such as asexual propagation and extensive inbreeding. The objective of this study was to develop and demonstrate methods that support complex pedigree reconstruction via the total length of identical by state haplotypes (referred to in this study as “summed potential lengths of shared haplotypes”, SPLoSH). A custom Python script, HapShared, was developed to generate SPLoSH data in apple and sweet cherry. HapShared was used to establish empirical distributions of SPLoSH data for known relationships in these crops. These distributions were then used to estimate previously unknown relationships. Case studies in each crop demonstrated various pedigree reconstruction scenarios using SPLoSH data. For cherry, a full-sib relationship was deduced for ‘Emperor Francis, and ‘Schmidt’, a half-sib relationship for ‘Van’ and ‘Windsor’, and the paternal grandparents of ‘Stella’ were confirmed. For apple, 29 cultivars were found to share an unknown parent, the pedigree of the unknown parent of ‘Cox’s Pomona’ was reconstructed, and ‘Fameuse’ was deduced to be a likely grandparent of ‘McIntosh’. Key genetic resources that enabled this empirical study were large genome-wide SNP array datasets, integrated genetic maps, and previously identified pedigree relationships. Crops with similar resources are also expected to benefit from using HapShared for empowering pedigree reconstruction.
Background Environmental adaptation and expanding harvest seasons are primary goals of most peach [ Prunus persica (L.) Batsch] breeding programs. Breeding perennial crops is a challenging task due to their long breeding cycles and large tree size. Pedigree-based analysis using pedigreed families followed by haplotype construction creates a platform for QTL and marker identification, validation, and the use of marker-assisted selection in breeding programs. Results Phenotypic data of seven F 1 low to medium chill full-sib families were collected over 2 years at two locations and genotyped using the 9 K SNP Illumina array. Three QTLs were discovered for bloom date (BD) and mapped on linkage group 1 (LG1) (172–182 cM), LG4 (48–54 cM), and LG7 (62–70 cM), explaining 17–54%, 11–55%, and 11–18% of the phenotypic variance, respectively. The QTL for ripening date (RD) and fruit development period (FDP) on LG4 was co-localized at the central part of LG4 (40–46 cM) and explained between 40 and 75% of the phenotypic variance. Haplotype analyses revealed SNP haplotypes and predictive SNP marker(s) associated with desired QTL alleles and the presence of multiple functional alleles with different effects for a single locus for RD and FDP. Conclusions A multiple pedigree-linked families approach validated major QTLs for the three key phenological traits which were reported in previous studies across diverse materials, geographical distributions, and QTL mapping methods. Haplotype characterization of these genomic regions differentiates this study from the previous QTL studies. Our results will provide the peach breeder with the haplotypes for three BD QTLs and one RD/FDP QTL to create predictive DNA-based molecular marker tests to select parents and/or seedlings that have desired QTL alleles and cull unwanted genotypes in early seedling stages.
Background Single nucleotide polymorphism (SNP) array technology has been increasingly used to generate large quantities of SNP data for use in genetic studies. As new arrays are developed to take advantage of new technology and of improved probe design using new genome sequence and panel data, a need to integrate data from different arrays and array platforms has arisen. This study was undertaken in view of our need for an integrated high-quality dataset of Illumina Infinium® 20 K and Affymetrix Axiom® 480 K SNP array data in apple ( Malus × domestica ). In this study, we qualify and quantify the compatibility of SNP calling, defined as SNP calls that are both accurate and concordant, across both arrays by two approaches. First, the concordance of SNP calls was evaluated using a set of 417 duplicate individuals genotyped on both arrays starting from a set of 10,295 robust SNPs on the Infinium array. Next, the accuracy of the SNP calls was evaluated on additional germplasm ( n = 3141) from both arrays using Mendelian inconsistent and consistent errors across thousands of pedigree links. While performing this work, we took the opportunity to evaluate reasons for probe failure and observed discordant SNP calls. Results Concordance among the duplicate individuals was on average of 97.1% across 10,295 SNPs. Of these SNPs, 35% had discordant call(s) that were further curated, leading to a final set of 8412 (81.7%) SNPs that were deemed compatible. Compatibility was highly influenced by the presence of alternate probe binding locations and secondary polymorphisms. The impact of the latter was highly influenced by their number and proximity to the 3′ end of the probe. Conclusions The Infinium and Axiom SNP array data were mostly compatible. However, data integration required intense data filtering and curation. This work resulted in a workflow and information that may be of use in other data integration efforts. Such an in-depth analysis of array concordance and accuracy as ours has not been previously described in the literature and will be useful in future work on SNP array data integration and interpretation, and in probe/platform development.
Karyotyping using high-density genome-wide SNP markers identified various chromosomal aberrations in oil palm (Elaeis guineensis Jacq.) with supporting evidence from the 2C DNA content measurements (determined using FCM) and chromosome counts. Oil palm produces a quarter of the world’s total vegetable oil. In line with its global importance, an initiative to sequence the oil palm genome was carried out successfully, producing huge amounts of sequence information, allowing SNP discovery. High-capacity SNP genotyping platforms have been widely used for marker–trait association studies in oil palm. Besides genotyping, a SNP array is also an attractive tool for understanding aberrations in chromosome inheritance. Exploiting this, the present study utilized chromosome-wide SNP allelic distributions to determine the ploidy composition of over 1,000 oil palms from a commercial F1 family, including 197 derived from twin-embryo seeds. Our method consisted of an inspection of the allelic intensity ratio using SNP markers. For palms with a shifted or abnormal distribution ratio, the SNP allelic frequencies were plotted along the pseudo-chromosomes. This method proved to be efficient in identifying whole genome duplication (triploids) and aneuploidy. We also detected several loss of heterozygosity regions which may indicate small chromosomal deletions and/or inheritance of identical by descent regions from both parents. The SNP analysis was validated by flow cytometry and chromosome counts. The triploids were all derived from twin-embryo seeds. This is the first report on the efficiency and reliability of SNP array data for karyotyping oil palm chromosomes, as an alternative to the conventional cytogenetic technique. Information on the ploidy composition and chromosomal structural variation can help to better understand the genetic makeup of samples and lead to a more robust interpretation of the genomic data in marker–trait association analyses.
A wealth of previously unknown pedigree information has been generated for apple (Malus x domestica) through recent studies and an ongoing pedigree identification project using SNP array data. This information has been postulated to be useful in a number of ways for germplasm collections and breeding programs. For example, pedigree information is an important part of genetic characterization of material in germplasm collections. Germplasm collections often seek to preserve cultivars that currently lack commercial relevance but have regional or historical significance or are phenotypically interesting. Pedigree information is often limited on such material. With pedigree information, an accurate genetic structure of these collections can be conveyed to breeders who wish to incorporate novel germplasm into new cultivars. To illustrate ways in which this newly generated pedigree information can be used, we provide examples involving accessions from the germplasm collection Okowerk and the breeding association "apfel:gut e.V.", which focus in the development of apples for regional organic production. Both organizations are located in northern Germany. Despite focusing on regional apple cultivars and genetic diversity, many of their accessions and breeding lines are common, cosmopolitan, and foundational cultivars. However, the SNP-based analyses also reveal regionally specific founders, not true-to-type accessions, and genetically unique historical cultivars. This new information has benefitted these organizations by providing genetic characterization that was previously lacking, enabling them to better utilize these genetic resources. Additionally, this information will also be useful in the context of an ongoing larger scale apple pedigree reconstruction project.
The Netherlands' field genebank collection of European wild apple (Malus sylvestris), consisting of 115 accessions, was studied in order to determine whether duplicates and mistakes had been introduced, and to develop a strategy to optimize the planting design of the collection as a seed orchard. We used the apple 20K Infinium single nucleotide polymorphism (SNP) array, developed in M. domestica, for the first time for genotyping in M. sylvestris. We could readily detect the clonal copies and unexpected duplicates. Thirty-two M. sylvestris accessions (29%) showed a close genetic relationship (parent-child, full-sib, or half-sib) to another accession, which reflects the small effective population size of the in situ populations. Traces of introgression from M. domestica were only found in 7 individuals. This indicates that pollination preferentially took place among the M. sylvestris trees. We conclude that the collection can be considered as mainly pure M. sylvestris accessions. The results imply that it should be managed as one unit when used for seed production. A bias in allele frequencies in the seeds may be prevented by not harvesting all accessions with a close genetic relationship to the others in the seed orchard. We discuss the value of using the SNP array to elaborate the M. sylvestris genetic resources more in depth, including for phasing the markers in a subset of the accessions, as a first step towards genetic resources management at the level of haplotypes.
The Rosaceae crop family (including almond, apple, apricot, blackberry, peach, pear, plum, raspberry, rose, strawberry, sweet cherry, and sour cherry) provides vital contributions to human well-being and is economically significant across the U.S. In 2003, industry stakeholder initiatives prioritized the utilization of genomics, genetics, and breeding to develop new cultivars exhibiting both disease resistance and superior horticultural quality. However, rosaceous crop breeders lacked certain knowledge and tools to fully implement DNA-informed breeding-a "chasm" existed between existing genomics and genetic information and the application of this knowledge in breeding. The RosBREED project ("Ros" signifying a Rosaceae genomics, genetics, and breeding community initiative, and "BREED", indicating the core focus on breeding programs), addressed this challenge through a comprehensive and coordinated 10-year effort funded by the USDA-NIFA Specialty Crop Research Initiative. RosBREED was designed to enable the routine application of modern genomics and genetics technologies in U.S. rosaceous crop breeding programs, thereby enhancing their efficiency and effectiveness in delivering cultivars with producer-required disease resistances and market-essential horticultural quality. This review presents a synopsis of the approach, deliverables, and impacts of RosBREED, highlighting synergistic global collaborations and future needs. Enabling technologies and tools developed are described, including genome-wide scanning platforms and DNA diagnostic tests. Examples of DNA-informed breeding use by project participants are presented for all breeding stages, including pre-breeding for disease resistance, parental and seedling selection, and elite selection advancement. The chasm is now bridged, accelerating rosaceous crop genetic improvement.
Acidity and sweetness are important qualities for apple breeders and understanding their genetic regulation can improve the breeding process. In previous QTL studies, fruit quality assessments were performed using instrumental measurements, leading to the identification of two major QTL Ma and Ma3. Here, we use sensorial data to investigate the role of known and unknown genetic factors in the perception of acidity and sweetness in three pedigreed full-sib families. For 2 years, these families were phenotyped by a trained panel, at harvest and after 2 months of cold storage, and genotyped with a new 50 K SNP array. FlexQTLTM analyses using both an additive and an additive + dominance model resulted in the identification of Ma and Ma3 for acidity as well as sweetness, whereas the use of the additive model yielded decisive evidence for the discovery of two additional QTL on LG1 and LG6 for acidity. QTL genotypes were qualified as the inverse of each other, which indicates that individuals with the less-acidity alleles of Ma and Ma3 are perceived sweeter. Due to the genetic configuration in the families studied, resulting from a link with the Pale Green Disorder locus, no incomplete dominance effect could be detected for the Ma locus, although previously reported in the literature. For the Ma3 locus, however, an incomplete dominance effect (58%) is reported here for the first time. The Ma3 locus was also further confined to a 2–4-cM region and a predictive marker for this locus was identified.
Background: Fruit quality traits have a significant effect on consumer acceptance and subsequently on peach (Prunus persica (L.) Batsch) consumption. Determining the genetic bases of key fruit quality traits is essential for industry to improve fruit quality and increase consumption. A Bayesian approach embedded in the FlexQTL software increases the accuracy of QTL mapping and the probability of identifying new and validating known QTLs across a wide range of genetic backgrounds.Results: Phenotypic data of seven F1 low to medium chill full-sib families were collected over two years at two locations and genotyped using the 9K SNP Illumina array. One major QTL for fruit blush was found on linkage group 4 (LG4) at 40–46 cM that explained from 20 to 32% of the total phenotypic variance and showed three QTL alleles of different effects. For SSC, one QTL was mapped on LG5 at 60-72cM and explained from 17 to 39% of the phenotypic variance. A major QTL for TA that co-localized with the major locus for low-acid fruit (D-locus) was mapped at the proximal end of LG5 and explained 35 to 80% of the phenotypic variance. The new QTL for TA on the distal end of LG5 explained 14 to 22% of the phenotypic variance. This QTL co-localized with the QTL for SSC and affected TA only when the first QTL is homozygous for high acidity (epistasis). Haplotype analyses revealed SNP haplotypes and predictive SNP marker(s) associated with desired QTL alleles.Conclusions: A multi-family-based QTL discovery approach enhanced the ability to discover a new TA QTL and validated other QTLs which were reported in previous studies. Identified predictive SNPs and their original sources will facilitate the selection of parents and/or seedlings that have desired haplotype alleles. Our findings will help peach breeders develop new predictive, DNA-based molecular marker tests for routine use in marker-assisted breeding (MAB).
Background Fruit quality traits have a significant effect on consumer acceptance and subsequently on peach ( Prunus persica (L.) Batsch) consumption. Determining the genetic bases of key fruit quality traits is essential for the industry to improve fruit quality and increase consumption. Pedigree-based analysis across multiple peach pedigrees can identify the genomic basis of complex traits for direct implementation in marker-assisted selection. This strategy provides breeders with better-informed decisions and improves selection efficiency and, subsequently, saves resources and time. Results Phenotypic data of seven F 1 low to medium chill full-sib families were collected over 2 years at two locations and genotyped using the 9 K SNP Illumina array. One major QTL for fruit blush was found on linkage group 4 (LG4) at 40–46 cM that explained from 20 to 32% of the total phenotypic variance and showed three QTL alleles of different effects. For soluble solids concentration (SSC), one QTL was mapped on LG5 at 60-72 cM and explained from 17 to 39% of the phenotypic variance. A major QTL for titratable acidity (TA) co-localized with the major locus for low-acid fruit ( D -locus). It was mapped at the proximal end of LG5 and explained 35 to 80% of the phenotypic variance. The new QTL for TA on the distal end of LG5 explained 14 to 22% of the phenotypic variance. This QTL co-localized with the QTL for SSC and affected TA only when the first QTL is homozygous for high acidity (epistasis). Haplotype analyses revealed SNP haplotypes and predictive SNP marker(s) associated with desired QTL alleles. Conclusions A multi-family-based QTL discovery approach enhanced the ability to discover a new TA QTL at the distal end of LG5 and validated other QTLs which were reported in previous studies. Haplotype characterization of the mapped QTLs distinguishes this work from the previous QTL studies. Identified predictive SNPs and their original sources will facilitate the selection of parents and/or seedlings that have desired QTL alleles. Our findings will help peach breeders develop new predictive, DNA-based molecular marker tests for routine use in marker-assisted breeding.
Many ornamental crops are polyploid or even exist at different ploidy levels. Polyploid QTL analysis tools have been developed in recent years, yet they are limited in the population types they accept. Biparental populations are nowadays being regarded as a limited tool for QTL discovery, as only a limited number of QTLs occurs in an experimental cross and their effects might not be stable across genetic backgrounds. Genome-Wide Association Studies include more genetic diversity but suffer from (hidden) genetic structure and low frequency of QTL alleles. Both factors influence QTL detection and effect estimation, decreasing the sensitivity of QTL analysis. Alternatively, multiparental populations (MPP) can be used, potentially combining multiple QTLs and QTL alleles with known population structure and balanced allele frequencies. Breeding populations of interconnected crosses also constitute a form of MPP and QTLs identified in them might be more applicable to commercial cultivars. To perform QTL analysis in polyploids, mixed models or Bayesian approaches that consider pedigree information are recommended. During the analysis, QTL effects are ideally estimated using Identity by Descent (IBD) alleles (genomic regions that originate from the same ancestor) which can be obtained through haplotype estimation. Although MPPs could thus be a powerful set-up to estimate polyploid haplotypes, a software gap was identified as no current polyploid haplotyping tools are able to utilize MPP pedigree information to obtain haplotypes across an MPP. In order to utilize MPPs to their full extent and expand polyploid QTL analyses to encompass typical breeding populations, new haplotyping tools must be developed.