In recent years many advances have been made towards developing cost-efficient low-density genomic tools for a wider implementation of genomic selection in aquaculture breeding programmes. Genotype imputation from low-density (LD) SNP panels of just hundreds of markers to high-density (HD) SNP panels has become a promising strategy to reduce the cost of genotyping while maintaining accurate genomic prediction. In this study we assessed the impact of the makeup of HD-genotyped reference populations on i) the accuracy of imputation for LD-genotyped individuals and ii) the accuracy of genomic prediction for three traits of importance in Atlantic salmon production: growth, resistance to cardiomyopathy syndrome and resistance to pancreas disease. An Atlantic salmon population genotyped with a 47 K SNP array was used for the study, along with an in silico LD panel of 554 SNPs. Five reference population scenarios for imputation were tested, which included only the parents of the candidates for selection, a combination of parents and candidates, or just candidates. All scenarios resulted in highly accurate imputation rates except when the HD reference population was only composed of selection candidates which also resulted in slightly under-dispersed estimated breeding values. The accuracy of imputation barely had an impact on the accuracy of genomic prediction, with imputed datasets reaching similar accuracy to the HD-panel. When HD-genotyped parents were used, adding a proportion of the offspring to the imputation reference population did not result in any benefit in terms of genomic prediction. Imputation is a cost-effective and robust option for genomic selection in aquaculture.
Rising temperatures due to anthropogenic climate change pose a threat to wild Atlantic salmon (Salmo salar) populations in their natural habitat and to farmed populations during their major growth phase in coastal (seawater) net pens. Given that tremendous gains have been made in farmed salmon production through artificial selection programs for traits such as growth rate and disease resistance, we therefore examined variation among 105 families in post-smolt seawater thermal tolerance to assess whether this trait warrants inclusion in a selective breeding program. We used two established thermal challenge protocols for this: a rapid temperature increase using loss of equilibrium as the endpoint (critical thermal maximum; CTmax [1506 fish]) and a slower increase with mortality or morbidity as the endpoint (incremental thermal maximum; ITmax [936 fish]). High estimated heritability values were obtained for both (h² = 0.47 and 0.40, respectively), suggesting that improved acute and/or chronic high-temperature tolerance may be attainable for farmed salmon through artificial selection. Furthermore, given that farmed salmon are not many generations removed from wild, wild populations may also have some capacity to adapt to increasing temperatures brought about by climate change. However, we found no genetic correlation between CTmax and ITmax. Genetic correlations between these indices and other traits that might influence thermal tolerance (body size, condition factor, ventricle size, and hematocrit) were absent or, at most, weak.
The yield from fisheries and aquaculture production has risen dramatically over the last few decades, with rising demands to meet the needs of a growing world population that is also consuming more seafood per capita. Aquaculture products have, in recent years, also overtaken tonnage obtained from wild fisheries. This is particularly true for finfish species such as Atlantic salmon. However, the marine lifecycle stage for Atlantic salmon aquaculture requires ocean-based net pens, which introduces environmental vulnerabilities that can negatively affect production. Of particular concern is the ability of the fish to tolerate rising seawater temperatures in the face of climate change. One way of tackling this challenge is to incorporate selective breeding for Atlantic salmon families that display greater tolerance to elevated seawater temperatures. The goal of this study was to determine the genomic heritability of seawater temperature tolerance in Atlantic salmon and to investigate the genetic architecture of this trait for more effective broodstock management. We report a moderate genomic-based heritability for seawater temperature at mortality in Atlantic salmon (0.208 +/- 0.041) with a favourable correlation between growth traits (body mass and fork length) and seawater temperature at mortality. Our results also show that seawater temperature at mortality is a polygenic trait, with no significant quantitative trait loci identified. The moderate genomic heritability for seawater temperature at mortality suggests that this trait could be included in Atlantic salmon breeding programs. Importantly, the favourable correlation between growth traits and seawater temperature at mortality suggests that selection for seawater temperature tolerance would not have a negative impact on growth in this population. No single quantitative trait locus for seawater temperature at mortality was identified, suggesting that genomic selection will be required to achieve an increased tolerance to seawater temperature in this population.
Atlantic salmon is an important aquaculture species farmed in ocean net-pens and therefore subjected to changing environmental conditions, including rising temperatures. This creates a need for research on the thermal tolerance of this species for the future of sustainable aquaculture. We investigated the thermal tolerance of individually tagged Atlantic salmon post-smolts subjected sequentially to two common high-temperature challenges: critical thermal maximum (CTmax) followed by incremental thermal maximum (ITmax). Our goals were (1) to determine whether CTmax can predict ITmax for individual fish, and (2) to examine connections between various body size (mass, length, condition factor), cardiac (absolute and relative ventricle mass) and blood (hematocrit) metrics and thermal tolerance. We found no relationship between CTmax and ITmax. This is of concern because CTmax, which is a quick and easy test, is often used to predict upper lethal limits in fish despite not using real-world rates of temperature increase and not using death as the experimental endpoint (unlike ITmax). Also, some metrics which correlated in one direction with CTmax had the opposite correlation with ITmax. For instance, smaller fish or fish with smaller ventricles had a higher CTmax but a lower ITmax than larger fish or fish with larger ventricles. Taken together, these results highlight the need to take care when using acute thermal tolerance tests to predict real-world responses to rising temperatures.
The objective of this study was to fine map the quantitative trait loci region of chromosome 27 which was previously reported as a major contributor of total genetic variation for resistance against cardiomyopathy syndrome in Atlantic salmon. The fine mapping using whole genome resequencing of individuals increased resolution of the targeted genomic region (Ssa27:10-12 Mbps), resulting in uncovering the two immune related genes (HA1F and MR1). Transcriptome analysis revealed only two differentially expressed genes, MR1 and HA1F across resistant vs susceptible individuals. The gene MR1 was highly expressed in the susceptible group with log2 fold change of 6.405 compared to resistant individuals. Contrary to MR1 expression, HA1F gene was lowly expressed (log2 fold change of 2.099) in susceptible group. Moreover, the proteomics results partially validated the results of transcriptomics with the product of MR1 being differentially produced across resistant vs susceptible individuals.
Rising temperatures due to anthropogenic climate change pose a threat to wild Atlantic salmon populations in their natural habitat and to farmed populations during their major growth phase in coastal (seawater) net pens. While tremendous gains have been made in farmed salmon production through artificial selection programs, these have not considered the improvement of high-temperature tolerance. We therefore used two thermal challenge protocols to examine variation among families in a commercial salmon breeding program: a rapid temperature increase using loss of equilibrium as the endpoint (critical thermal maximum; CTmax) and a slower increase with mortality or morbidity as the endpoint (incremental thermal maximum; ITmax). High estimated heritability values were obtained for both (h² = 0.47 and 0.40, respectively), meaning that improved high-temperature tolerance should be attainable for farmed salmon through artificial selection. Furthermore, given that farmed salmon are not many generations removed from wild, this also suggested that wild populations may have some capacity to adapt to increasing temperatures brought about by climate change. However, we found no genetic or phenotypic correlations between CTmax and ITmax, and only weak (or no) genetic and phenotypic correlations between them and other traits predicted to influence thermal tolerance (body size and condition factor, ventricle size, and hematocrit).
A controlled Saprolegnia parasitica infection model was used to challenge 1158 fish representing 105 pedigreed Atlantic salmon families to evaluate the possibility of selecting for Saprolegnia resistance in a commercial breeding programme. Fish were infected in five study tanks and observed for 40 days post-infection for lesion score and survival. Survival analysis of the top 10 resistant and bottom 10 susceptible families indicated that the hazard of dying following Saprolegnia infection was 1509% higher in susceptible families. In all fish, a 10 g increase in weight correlated with a 7.8% increase in the hazard of dying while sex did not affect mortality. Resistance to Saprolegnia was estimated to have a heritability of 0.25, indicating that selection is possible. Genetic and phenotypic correlations indicated that the 11-point scoring system, developed in this study to quantify Saprolegnia infection severity, had a high negative correlation with survival as days to mortality at ≥-0.922(±0.005), suggesting that the scoring method could help assess lesion development in studies where mortality is not the primary biological endpoint.
This paper presents an extension to a heuristic method for phasing and imputation of genotypes of descendants in bi-parental populations so that it can phase and impute genotypes of parents of bi-parental populations that are fully ungenotyped or partially genotyped. The imputed genotypes of the parent are then used to impute low-density genotyped descendants of the bi-parental population to high-density. The extension works in three steps. First, it identifies whether a parent has no or low-density genotypes available and it identifies all of its relatives that have high-density genotypes. Second, using the high-density information of relatives, it determines whether the parent is homozygous or heterozygous for a given locus. Third, it phases heterozygous positions of the parent by matching haplotypes to its relatives. We implemented the new algorithm in an extension of the AlphaPlantImptue software and tested its accuracy of imputing missing parent genotypes in simulated bi-parental populations from different scenarios. We also tested the accuracy of imputation of the missing parent’s descendants using the true genotype of the parent and compared this to using the imputed genotypes of the parent. Our results show that across all scenarios, the accuracy of imputation of a parent, measured as the correlation between true and imputed genotypes, was > 0.98 and did not drop below ∼ 0.96. The imputation accuracy of a parent was always higher when it was inbred than when it was outbred and when it had low-density genotypes. Including ancestors of the parent at HD, increasing the number of crosses and the number of high-density descendants all increased the accuracy of imputation. The high imputation accuracy achieved for the parent across all scenarios translated to little or no impact on the accuracy of imputation of its descendants at low-density. Key Message New fast and accurate method for phasing and imputation of SNP chip genotypes within diploid bi-parental plant populations.
This paper describes a family-based phasing algorithm, for variable-coverage sequence data, that first minimises phasing errors and then maximises the proportion of alleles phased. This algorithm is one of the essential tools that underpin an overall strategy for generating highly accurate sequence data on whole populations at low cost. The algorithm is called AlphaFamSeq. It uses sequence data on the focal individual and at least two generations of ancestors to phase alleles. In the first step, AlphaFamSeq calculates allele probabilities using iterative peeling. In subsequent steps, the alleles are phased using heuristics deriving information from the sequence data of parents, grandparents and progenies and, if available, from other families in the pedigree. AlphaFamSeq was tested on a range of simulated data sets. AlphaFamSeq gives low phasing error rates and, if there is sufficient sequence information and haplotype sharing amongst individuals, it can give a high yield of correctly phased alleles. The allele threshold had a large effect and window size had a small effect on performance. When all individuals in a single family were sequenced at different coverages the highest correctly phased alleles reached 90% of the possible maximum (98.9%) at ~1/6 of the maximum aggregate coverage. Adding sequence information from other related individuals increased the percentage of correctly phased alleles. Imputation performance was high across all allele frequencies (average correlation by marker of 0.94), except for a slight decrease at very low frequencies (≤0.01 MAF). Within an overall strategy for generating highly accurate sequence data on whole populations at low cost the role of AlphaFamSeq is to provide very accurately phased haplotypes on focal individuals, who are individuals whose haplotypes are very common in the population.
Background: This study uses simulation to explore and quantify the potential effect of shifting recombination hotspots on genetic gain in livestock breeding programs.Methods: We simulated three scenarios that differed in the locations of quantitative trait nucleotides (QTN) and recombination hotspots in the genome. In scenario 1, QTN were randomly distributed along the chromosomes and recombination was restricted to occur within specific genomic regions (i.e. recombination hotspots). In the other two scenarios, both QTN and recombination hotspots were located in specific regions, but differed in whether the QTN occurred outside of (scenario 2) or inside (scenario 3) recombination hotspots. We split each chromosome into 250, 500 or 1000 regions per chromosome of which 10% were recombination hotspots and/or contained QTN. The breeding program was run for 21 generations of selection, after which recombination hotspot regions were kept the same or were shifted to adjacent regions for a further 80 generations of selection. We evaluated the effect of shifting recombination hotspots on genetic gain, genetic variance and genic variance.Results: Our results show that shifting recombination hotspots reduced the decline of genetic and genic variance by releasing standing allelic variation in the form of new allele combinations. This in turn resulted in larger increases in genetic gain. However, the benefit of shifting recombination hotspots for increased genetic gain was only observed when QTN were initially outside recombination hotspots. If QTN were initially inside recombination hotspots then shifting them decreased genetic gain. Discussion: Shifting recombination hotspots to regions of the genome where recombination had not occurred for 21 generations of selection (i.e. recombination deserts) released more of the standing allelic variation available in each generation and thus increased genetic gain. However, whether and how much increase in genetic gain was achieved by shifting recombination hotspots depended on the distribution of QTN in the genome, the number of recombination hotspots and whether QTN were initially inside or outside recombination hotspots.Conclusions: Our findings show future scope for targeted modification of recombination hotspots e.g. through changes in zinc-finger motifs of the PRDM9 protein to increase genetic gain in production species.
Background: This paper describes a heuristic method for allocating low-coverage sequencing resources by targeting haplotypes rather than individuals. Low-coverage sequencing assembles high-coverage sequence information for every individual by accumulating data from the genome segments that they share with many other individuals into consensus haplotypes. Deriving the consensus haplotypes accurately is critical for achieving a high phasing and imputation accuracy. In order to enable accurate phasing and imputation of sequence information for the whole population, we allocate the available sequencing resources among individuals with existing phased genomic data by targeting the sequencing coverage of their haplotypes.Results: Our method, called AlphaSeqOpt, prioritizes haplotypes using a score function that is based on the frequency of the haplotypes in the sequencing set relative to the target coverage. AlphaSeqOpt has two steps: (1) selection of an initial set of individuals by iteratively choosing the individuals that have the maximum score conditional on the current set, and (2) refinement of the set through several rounds of exchanges of individuals. AlphaSeqOpt is very effective for distributing a fixed amount of sequencing resources evenly across haplotypes, which results in a reduction of the proportion of haplotypes that are sequenced below the target coverage. AlphaSeqOpt can provide a greater proportion of haplotypes sequenced at the target coverage by sequencing less individuals, as compared with other methods that use a score function based on haplotype frequencies in the population. A refinement of the initially selected set can provide a larger more diverse set with more unique individuals, which is beneficial in the context of low-coverage sequencing. We extend the method with an approach for filtering rare haplotypes based on their flanking haplotypes, so that only those that are likely to derive from a recombination event are targeted.Conclusions: We present a method for allocating sequencing resources so that a greater proportion of haplotypes are sequenced at a coverage that is sufficiently high for population-based imputation with low-coverage sequencing. The haplotype score function, the refinement step, and the new approach for filtering rare haplotypes make AlphaSeqOpt more effective for that purpose than previously reported methods for reducing sequencing redundancy.
Genotyping‐by‐sequencing (GBS) is an alternative genotyping method to single‐nucleotide polymorphism (SNP) arrays that has received considerable attention in the plant breeding community. In this study we use simulation to quantify the potential of low‐coverage GBS and imputation for cost‐effective genomic selection in biparental segregating populations. The simulations comprised a range of scenarios where SNP array or GBS data were used to train the genomic selection model, to predict breeding values, or both. The GBS data were generated with sequencing coverages (x) from 4x to 0.01x. The data were used either nonimputed or imputed by the AlphaImpute program. The size of the training and prediction sets was either held fixed or was increased by reducing sequencing coverage per individual. The results show that nonimputed 1x GBS data provided comparable prediction accuracy and bias, and for the used measurement of return on investment, outperformed the SNP array data. Imputation allowed for further reduction in sequencing coverage, to as low as 0.1x with 10,000 markers or 0.01x with 100,000 markers. The results suggest that using such data in biparental families gave up to 5.63 times higher return on investment than using the SNP array data. Reduction of sequencing coverage per individual and imputation can be leveraged to genotype larger training sets to increase prediction accuracy and larger prediction sets to increase selection intensity, which both allow for higher response to selection and higher return on investment.
This paper uses simulation to explore how gene drives can increase genetic gain in livestock breeding programs. Gene drives are naturally occurring phenomena that cause a mutation on one chromosome to copy itself onto its homologous chromosome.
Background This paper describes a method, called AlphaSeqOpt, for the allocation of sequencing resources in livestock populations with existing phased genomic data to maximise the ability to phase and impute sequenced haplotypes into the whole population. Methods We present two algorithms. The first selects focal individuals that collectively represent the maximum possible portion of the haplotype diversity in the population. The second allocates a fixed sequencing budget among the families of focal individuals to enable phasing of their haplotypes at the sequence level. We tested the performance of the two algorithms in simulated pedigrees. For each pedigree, we evaluated the proportion of population haplotypes that are carried by the focal individuals and compared our results to a variant of the widely-used key ancestors approach and to two haplotype-based approaches. We calculated the expected phasing accuracy of the haplotypes of a focal individual at the sequence level given the proportion of the fixed sequencing budget allocated to its family. Results AlphaSeqOpt maximises the ability to capture and phase the most frequent haplotypes in a population in three ways. First, it selects focal individuals that collectively represent a larger portion of the population haplotype diversity than existing methods. Second, it selects focal individuals from across the pedigree whose haplotypes can be easily phased using family-based phasing and imputation algorithms, thus maximises the ability to impute sequence into the rest of the population. Third, it allocates more of the fixed sequencing budget to focal individuals whose haplotypes are more frequent in the population than to focal individuals whose haplotypes are less frequent. Unlike existing methods, we additionally present an algorithm to allocate part of the sequencing budget to the families (i.e. immediate ancestors) of focal individuals to ensure that their haplotypes can be phased at the sequence level, which is essential for enabling and maximising subsequent sequence imputation. Conclusions We present a new method for the allocation of a fixed sequencing budget to focal individuals and their families such that the final sequenced haplotypes, when phased at the sequence level, represent the maximum possible portion of the haplotype diversity in the population that can be sequenced and phased at that budget.
This paper describes AlphaSim, a software package for simulating plant and animal breeding programs. AlphaSim enables the simulation of multiple aspects of breeding programs with a high degree of flexibility. AlphaSim simulates breeding programs in a series of steps: (i) simulate haplotype sequences and pedigree; (ii) drop haplotypes into the base generation of the pedigree and select single-nucleotide polymorphism (SNP) and quantitative trait nucleotide (QTN); (iii) assign QTN effects, calculate genetic values, and simulate phenotypes; (iv) drop haplotypes into the burn-in generations; and (v) perform selection and simulate new generations. The program is flexible in terms of historical population structure and diversity, recent pedigree structure, trait architecture, and selection strategy. It integrates biotechnologies such as doubled-haploids (DHs) and gene editing and allows the user to simulate multiple traits and multiple environments, specify recombination hot spots and cold spots, specify gene jungles and deserts, perform genomic predictions, and apply optimal contribution selection. AlphaSim also includes restart functionalities, which increase its flexibility by allowing the simulation process to be paused so that the parameters can be changed or to import an externally created pedigree, trial design, or results of an analysis of previously simulated data. By combining the options, a user can simulate simple or complex breeding programs with several generations, variable population structures and variable breeding decisions over time. In conclusion, AlphaSim is a flexible and computationally efficient software package to simulate biotechnology enhanced breeding programs with the aim of performing rapid, low-cost, and objective in silico comparison of breeding technologies.
Pancreas disease (PD), caused by a salmonid alphavirus (SAV), has a large negative economic and animal welfare impact on Atlantic salmon aquaculture. Evidence for genetic variation in host resistance to this disease has been reported, suggesting that selective breeding may potentially form an important component of disease control. The aim of this study was to explore the genetic architecture of resistance to PD, using survival data collected from two unrelated populations of Atlantic salmon; one challenged with SAV as fry in freshwater (POP 1) and one challenged with SAV as post-smolts in sea water (POP 2). Analyses of the binary survival data revealed a moderate-to-high heritability for host resistance to PD in both populations (fry POP 1 h 2 ~0.5; post-smolt POP 2 h 2 ~0.4). Subsets of both populations were genotyped for single nucleotide polymorphism markers, and six putative resistance quantitative trait loci (QTL) were identified. One of these QTL was mapped to the same location on chromosome 3 in both populations, reaching chromosome-wide significance in both the sire- and dam-based analyses in POP 1, and genome-wide significance in a combined analysis in POP 2. This independently verified QTL explains a significant proportion of host genetic variation in resistance to PD in both populations, suggesting a common underlying mechanism for genetic resistance across lifecycle stages. Markers associated with this QTL are being incorporated into selective breeding programs to improve PD resistance.
Restriction site-Associated DNA sequencing (RAD-Seq) is widely applied to generate genome-wide sequence and genetic marker datasets. RAD-Seq has been extensively utilised, both at the population level and across species, for example in the construction of phylogenetic trees. However, the consistency of RAD-Seq data generated in different laboratories, and the potential use of cross-species orthologous RAD loci in the estimation of genetic relationships, have not been widely investigated. This study describes the use of SbfI RAD-Seq data for the estimation of evolutionary relationships amongst ten teleost fish species, using previously established phylogeny as a benchmark.