We assessed the usefulness of defining kinship based on the combinations of alleles (haplotypes) instead of individual alleles for genomic predictions of milk traits. The accuracy and bias of predictions were calculated when pure Holstein (H), pure Jersey (J) and/or H×J crossbred cattle were included in the reference set. For within-breed predictions, using haplotypes marginally improved the accuracy by 1%. For across-breed predictions, using haplotypes of 5 adjacent markers increased the accuracy of prediction for fat yield for H and J by 4 and 2%, respectively. The most consistent gains for H (1.8%) and J (1.5%) were achieved when all pure and crossbred animals were included in the reference. The change in accuracy of predictions for crossbred animals was not consistent, which could be associated with shorter haplotypes and lower accuracy of phasing. Generally, the bias of predictions was lower when haplotypes were used in the relationship matrix.
Increasing the reliability of genomic prediction (GP) of economic traits in the pasture-based dairy production systems of New Zealand (NZ) and Australia (AU) is important to both countries. This study assessed if sharing cow phenotype and genotype data of NZ and AU improves the reliability of GP for NZ bulls. Data from approximately 32,000 NZ genotyped cows and their contemporaries were included in the May 2018 routine genetic evaluation of the Australian Dairy cattle in an attempt to provide consistent phenotypes for both countries. After the genetic evaluation, deregressed proofs of cows were calculated for milk yield traits. The April 2018 multiple across-country evaluation of Interbull was also used to calculate deregressed proofs for bulls on the NZ scale. Approximately 1,178 Jersey (Jer) and 6,422 Holstein (Hol) bulls had genotype and phenotype data. In addition to NZ cows, phenotype data of close to 60,000 genotyped Australian (AU) cows from the same genetic evaluation run as NZ cows were used. All AU and NZ females were genotyped using low-density SNP chips (<10K SNP) and were imputed first to 50K and then to ∼600K (referred to as high density; HD). We used up to 98,000 animals in the reference populations, both by expanding the NZ reference set (cow, bull, single breed to multi-breed set) and by adding AU cows. Reliabilities of GP were calculated for 508 Jer and 1,251 Hol bulls whose sires are not included in the reference set (RS) to ensure that real differences are not masked by close relationships. The GP was tested using 50K or high-density SNP chip using genomic BLUP in bivariate (considering country as a trait) or single trait models. The RS that gave the highest reliability for each breed were also tested using a hybrid GP method that combines expectation maximization with Bayes R. The addition of the AU cows to an NZ RS that included either NZ cows only, or cows and bulls, improved the reliability of GP for both NZ Hol and Jer validation bulls for all traits. Using single breed reference populations also increased reliability when NZ crossbred cows were added to reference populations that included only purebred NZ bulls and cows and AU cows. The full multi-breed RS (all NZ cows and bulls and AU cows) provided similar reliabilities in NZ Hol bulls, when compared with the single breed reference with crossbred NZ cows. For Jer validation bulls, the RS that included Jer cows and bulls and crossbred cows from NZ and Jer cows from AU was marginally better than the all-breed, all-country RS. In terms of reliability, the advantage of the HD SNP chip was small but captured more of the genomic variance than the 50K, particularly for Hol. The expectation maximization Bayes R GP method was slightly (up to 3 percentage points) better than genomic BLUP. We conclude that GP of milk production traits in NZ bulls improves by up to 7 percentage points in reliability by expanding the NZ reference population to include AU cows.
Genomic prediction (GP) in numerically small breeds is limited due to the requirement for a large reference set. Across breed prediction has not been very successful either. Our objective was to test alternative models for across breed and multi-breed GP in a small Jersey population, utilizing prior information on marker causality. We used data on 596 Jersey bulls from new Zealand and 5503 Holstein bulls from the Netherlands, all of which had deregressed proofs for stature. Two sets of genotype data were used, one containing 357 potential causal markers identified from a multi-breed meta-GWAS on stature (top markers), while the other contained 48,912 markers on the custom 50k chip, excluding the top markers. We used models in which only one GRM (either top markers, 50k, or top plus 50k markers combined) was fitted, and models in which two GRMs (both the top and 50k) were fitted simultaneously, however with different variance components to weight the GRMs differently. Moreover, we estimated the genetic correlation(s) between the breeds (for each GRM) using a multi-trait GP model, which implicitly weights the contribution of one breed’s information to another. Across breed, we observed low accuracies of GP when the 50k markers were fitted alone (0.06) or when the top markers were added to 50k (0.15). Higher accuracy was obtained when only the top markers were fitted (0.21), whereas the highest accuracy was obtained when fitting 50k and top markers simultaneously as two independent GRMs (0.25). Multi-breed prediction outperformed both within and across breed prediction with accuracies ranging from 0.34 to 0.45, with the same trend as in across breed prediction. Based on our results, the best approach for across and multi-breed GP is to fit models that are able to isolate and differentially weight the most important markers for the trait. Keywords: Across breed genomic prediction, marker pre-selection, multi-trait model, sequence data.
For countries without large-scale genotyping, it is a challenge to implement an effective genomic evaluation system. Breeding company CRV has merged its bull training population with two such countries to build a training population of significance, and provide genomic software. Results show 6-20 times increase in training population adding the CRV bulls. The b-factors of the model(s) improve for most traits for both countries, with an exception for Israeli calving traits. Finally, the added reliability of genomic information, measured as equivalent daughter contributions, is on average four times bigger in the training populations including the CRV bulls, as compared to using only bulls genotyped by the country itself. Validation success may depend on the heritability of the trait, the estimated between-country correlation for the trait, the kinship between the two populations, as well as between the training population and the validation population, and -implicit in that between-country kinship- the past selection criteria as reflected in the similarities in the total merit indices over time.
Genomic prediction (GP) in farm livestock generally exploits SNP array genotypes. Now it is possible to impute from SNP chip genotypes to whole genome sequence. However, in an industry setting it is impractical to implement GP using millions of sequence variants. Livestock industries are therefore keen to leverage sequence data by selecting subsets of variants to develop custom SNP arrays. In this study we demonstrate that there are potential pitfalls in this approach that can lead to considerable bias in GP and can underestimate the potential advantages of sequence.
Background Dense SNP genotypes are often combined with complex trait phenotypes to map causal variants, study genetic architecture and provide genomic predictions for individuals with genotypes but no phenotype. A single method of analysis that jointly fits all genotypes in a Bayesian mixture model (BayesR) has been shown to competitively address all 3 purposes simultaneously. However, BayesR and other similar methods ignore prior biological knowledge and assume all genotypes are equally likely to affect the trait. While this assumption is reasonable for SNP array genotypes, it is less sensible if genotypes are whole-genome sequence variants which should include causal variants. Results We introduce a new method (BayesRC) based on BayesR that incorporates prior biological information in the analysis by defining classes of variants likely to be enriched for causal mutations. The information can be derived from a range of sources, including variant annotation, candidate gene lists and known causal variants. This information is then incorporated objectively in the analysis based on evidence of enrichment in the data. We demonstrate the increased power of BayesRC compared to BayesR using real dairy cattle genotypes with simulated phenotypes. The genotypes were imputed whole-genome sequence variants in coding regions combined with dense SNP markers. BayesRC increased the power to detect causal variants and increased the accuracy of genomic prediction. The relative improvement for genomic prediction was most apparent in validation populations that were not closely related to the reference population. We also applied BayesRC to real milk production phenotypes in dairy cattle using independent biological priors from gene expression analyses. Although current biological knowledge of which genes and variants affect milk production is still very incomplete, our results suggest that the new BayesRC method was equal to or more powerful than BayesR for detecting candidate causal variants and for genomic prediction of milk traits. Conclusions BayesRC provides a novel and flexible approach to simultaneously improving the accuracy of QTL discovery and genomic prediction by taking advantage of prior biological knowledge. Approaches such as BayesRC will become increasing useful as biological knowledge accumulates regarding functional regions of the genome for a range of traits and species.
Since April 2015, the Dutch-Flemish evaluation has been publishing genomically enhanced breeding values (GEBV) for a total of 65 traits and indices. These include traditional traits like milk production, conformation, longevity and calving traits; health traits like udder health, claw health, and fertlity traits; as well as novel traits that were recently introduced in the Dutch-Flemish (genomic) evaluation: lactose, urea, conception rate (to replace non-return 56), heifer fertility traits, calf survival, ketosis, and automatic milking system (AMS) traits. In the Dutch-Flemish genomic evaluation system bulls are considered for the training population when their breeding value for the trait of interest exceeds 50 percent reliability. However, additional edits might be advisable when correlations between countries are low or for traits where the reliability is highly dependent on correlated traits, rather than true (daughter) observations.
In dairy cattle, the rate of genetic gain from genomic selection depends on reliability of direct genomic values (DGV). One option to increase reliabilities could be to increase the size of the reference set used for prediction, by using genotyped bulls with daughter information in countries other than the evaluating country. The increase in reliabilities of DGV from using this information will depend on the extent of genotype by environment interaction between the evaluating country and countries contributing information, and whether this is correctly accounted for in the prediction method. As the genotype by environment interaction between Australia and Europe or North America is greater than between Europe and North America for most dairy traits, ways of including information from other countries in Australian genomic evaluations were examined. Thus, alternative approaches for including information from other countries and their effect on the reliability and bias of DGV of selection candidates were assessed. We also investigated the effect of including overseas (OS) information on reliabilities of DGV for selection candidates that had weaker relationships to the current Australian reference set. The DGV were predicted either using daughter trait deviations (DTD) for the bulls with daughters in Australia, or using this information as well as OS information by including deregressed proofs (DRP) from Interbull for bulls with only OS daughters in either single trait or bivariate models. In the bivariate models, DTD and DRP were considered as different traits. Analyses were performed for Holstein and Jersey bulls for milk yield traits, fertility, cell count, survival, and some type traits. For Holsteins, the data used included up to 3,580 bulls with DTD and up to 5,720 bulls with only DRP. For Jersey, about 900 bulls with DTD and 1,820 bulls with DRP were used. Bulls born after 2003 and genotyped cows that were not dams of genotyped bulls were used for validation. The results showed that the combined use of DRP on bulls with OS daughters only and DTD for Australian bulls in either the single trait or bivariate model increased the coefficient of determination [(R(2)) (DGV,DTD)] in the validation set, averaged across 6 main traits, by 3% in Holstein and by 5% in Jersey validation bulls relative to the use of DTD only. Gains in reliability and unbiasedness of DGV were similar for the single trait and bivariate models for production traits, whereas the bivariate model performed slightly better for somatic cell count in Holstein. The increase in R(2) (DGV,DTD) as a result of using bulls with OS daughters was relatively higher for those bulls and cows in the validation sets that were less related to the current reference set. For example, in Holstein, the average increase in R(2) for milk yield traits when DTD and DRP were used in a single trait model was 23% in the least-related cow group, but only 3% in the most-related cow group. In general, for both breeds the use of DTD from domestic sources and DRP from Interbull in a single trait or bivariate model can increase reliability of DGV for selection candidates.
INTRODUCTION Dairy cow fertility has caused much concern over the past three decades in many countries because it had been in decline, partly due to an unfavourable genetic correlation between fertility and milk traits. Accurate genomic predictions of fertility would be of great benefit to the dairy industry because it is measured only in mature females and has low heritability. It is also important to identify genes affecting fertility to better understand genetic factors that underpin the trait. BayesR, a Bayesian genomic prediction method, can achieve higher accuracy of genomic prediction compared to genomic best linear unbiased prediction (GBLUP), particularly for traits affected by many small effect genes as well as some of much larger effect (Erbe et al. 2012, Kemper et al. 2015). This occurs because BayesR models the single nucleotide polymorphism (SNP) effects as a mixture of four normal distributions, including a null distribution and one distribution with moderate to large variance. BayesR should also be a more precise method for QTL discovery than GWAS (genome-wide association analysis) or GBLUP. GWAS fits SNP individually which often results in one QTL being predicted by a large number of SNP in LD. GBLUP fits all SNP simultaneously but effects are distributed as a single normal distribution so are smeared across many adjacent SNP with strong shrinkage of larger effects. Furthermore, BayesR provides a well calibrated test of the likelihood that a SNP predicts a real QTL effect (posterior probability). Using BayesR we compare accuracy of genomic prediction for dairy cow fertility using high density SNP markers and imputed sequence variants in and close to genes. We also identify candidate genes and potential causal mutations associated with fertility.
The objectives were to investigate the accuracy of genomic evaluations obtained for a small dairy cattle population (Israeli Holsteins) via joint evaluation with a larger population (Dutch Holsteins), and to evaluate the use of pedigree data from foreign bulls computed by Interbull without daughter records in Israel. The training population included 4,010 Dutch bulls and 713 Israeli bulls. The validation population included 185 Israeli bulls with daughter records for milk production traits and slightly fewer bulls for the nonproduction traits. Milk, fat, and protein yields, somatic cell score, longevity, female fertility, direct and maternal calving ease, direct and maternal stillbirth, and the Israeli breeding index were analyzed. The genomic prediction model was based on the Bayesian multi-QTL model of Meuwissen and Goddard, where the effects of dense single nucleotide polymorphisms across the whole genome are fitted directly, without the use of haplotypes or identical-by-descent probabilities. Correlations of May 2014 estimated breeding values (EBV14) with genomic EBV (GEBV) were higher than the correlations of EBV14 with parent averages (PA) computed from the June 2009 evaluation for all traits. For the Israel selection index, the difference between EBV14 and GEBV correlation on the one hand and EBV14 and PA computed using Interbull data on the other hand was 15 percentage points. For protein, the difference between the corresponding correlations was 14 percentage points. Generally, correlations of EBV14 with PA based on Israeli EBV only were similar to correlations of EBV14 with PA including Interbull evaluations. Relative to EBV14, milk production traits were biased upwards for both GEBV and PA, but the bias was greater for PA. The Y-intercepts of regressions of EBV14 were significantly different from zero for regression on GEBV for all 3 milk production traits and the Israeli selection index. This was not the case for regression of EBV14 on PA. The regression line intersected with the line of unbiased estimation near the EBV of the bulls with highest values. Because only bulls with high evaluations are of interest for selection, GEBV for these bulls were less biased compared with that of bulls with lower evaluations. The difference in mean EBV14 between bulls born during 2007-2008 selected by GEBV and PA was 65 units. If half of all inseminations are by young bulls, then the annual genetic gain obtained by implementation of genomic evaluation will be 8 units per year (65/8). Because annual gain is currently 107 units, this is a gain of 7%.
Past attempts to find mutations causing variation in quantitative traits important in livestock production have largely been thwarted by the small effects that individual mutations have. However, we now have more powerful tools, such as genome sequence on individual animals, that increase our power to identify the mutations underlying quantitative trait loci (QTL). Here we review the statistical methods used to analyze association studies and the biological information that might be used to discriminate between alternative sites as the causal mutation. As well as the small effect of most QTL, it is also difficult to identify causal mutations because they are typically in LD with other polymorphic sites and so one cannot decide which is the causal mutation. The problem caused by LD can be reduced by using multiple breeds and by fitting all sequence variants simultaneously with a Bayesian model in which many variants are expected to have no effect. The effect size of a QTL may be increased by using traits, such as gene expression, that are close to the primary effect of the mutations. Biological information, such as that provided for the human genome by the ENCODE project, can be incorporated into the statistical analysis of the association study so that objective estimation of the probability that each sequence variant has an effect on the trait can be made.