Genomic prediction is applicable to individuals of different breeds. Empirical results to date, however, show limited benefits in using information on multiple breeds in the context of genomic prediction. We investigated a multitask Bayesian model, presented previously by others, implemented in a Bayesian stochastic search variable selection (BSSVS) model. This model allowed for evidence of quantitative trait loci (QTL) to be accumulated across breeds or for both QTL that segregate across breeds and breed-specific QTL. In both cases, single nucleotide polymorphism effects were estimated with information from a single breed. Other models considered were a single-trait and multitrait genomic residual maximum likelihood (GREML) model, with breeds considered as different traits, and a single-trait BSSVS model. All single-trait models were applied to each of the 2 breeds separately and to the pooled data of both breeds. The data used included a training data set of 6,278 Holstein and 722 Jersey bulls, as well as 374 Jersey validation bulls. All animals had genotypes for 474,773 single nucleotide polymorphisms after editing and phenotypes for milk, fat, and protein yields. Using the same training data, BSSVS consistently outperformed GREML. The multitask BSSVS, however, did not outperform single-trait BSSVS, which used pooled Holstein and Jersey data for training. Thus, the rigorous assumption that the traits are the same in both breeds yielded a slightly better prediction than a model that had to estimate the correlation between the breeds from the data. Adding the Holstein data significantly increased the accuracy of the single-trait GREML and BSSVS in predicting the Jerseys for milk and protein, in line with estimated correlations between the breeds of 0.66 and 0.47 for milk and protein yields, whereas only the BSSVS model significantly improved the accuracy for fat yield with an estimated correlation between breeds of only 0.05. The relatively high genetic correlations for milk and protein yields, and the superiority of the pooling strategy, is likely the result of the observed admixture between both breeds in our data. The Bayesian model was able to detect several QTL in Holsteins, which likely enabled it to outperform GREML. The inability of the multitask Bayesian models to outperform a simple pooling strategy may be explained by the fact that the pooling strategy assumes equal effects in both breeds; furthermore, this assumption may be valid for moderate-to large-sized QTL, which are important for multibreed genomic prediction.
Sequence data may potentially increase prediction accuracy compared to medium or high density (HD) SNP markers, by containing causative mutations directly rather than relying on linkage disequilibrium between markers and causative mutations. Besides causative mutations, sequence data contains a much larger number of variants that have no effect on the analysed trait. A Bayesian variable selection model could be used to assign large effects only to the causative mutations. In practice, however, analysing millions of sequence variants is computationally challenging. Therefore, we tested an approach to split up the analysis per chromosome, correcting for all other chromosomes using HD estimates, in a simulation study, using a faster, hybrid version of the Bayes R variable selection model. While directly computing breeding values based on effects estimated per chromosome resulted in a reduced accuracy, reanalysing all variants that were selected per chromosome resulted in a similar accuracy to analysing all variants simultaneously, especially when HD variants were included.
Background Dense SNP genotypes are often combined with complex trait phenotypes to map causal variants, study genetic architecture and provide genomic predictions for individuals with genotypes but no phenotype. A single method of analysis that jointly fits all genotypes in a Bayesian mixture model (BayesR) has been shown to competitively address all 3 purposes simultaneously. However, BayesR and other similar methods ignore prior biological knowledge and assume all genotypes are equally likely to affect the trait. While this assumption is reasonable for SNP array genotypes, it is less sensible if genotypes are whole-genome sequence variants which should include causal variants. Results We introduce a new method (BayesRC) based on BayesR that incorporates prior biological information in the analysis by defining classes of variants likely to be enriched for causal mutations. The information can be derived from a range of sources, including variant annotation, candidate gene lists and known causal variants. This information is then incorporated objectively in the analysis based on evidence of enrichment in the data. We demonstrate the increased power of BayesRC compared to BayesR using real dairy cattle genotypes with simulated phenotypes. The genotypes were imputed whole-genome sequence variants in coding regions combined with dense SNP markers. BayesRC increased the power to detect causal variants and increased the accuracy of genomic prediction. The relative improvement for genomic prediction was most apparent in validation populations that were not closely related to the reference population. We also applied BayesRC to real milk production phenotypes in dairy cattle using independent biological priors from gene expression analyses. Although current biological knowledge of which genes and variants affect milk production is still very incomplete, our results suggest that the new BayesRC method was equal to or more powerful than BayesR for detecting candidate causal variants and for genomic prediction of milk traits. Conclusions BayesRC provides a novel and flexible approach to simultaneously improving the accuracy of QTL discovery and genomic prediction by taking advantage of prior biological knowledge. Approaches such as BayesRC will become increasing useful as biological knowledge accumulates regarding functional regions of the genome for a range of traits and species.
In this study, we aimed to develop genomic estimated breeding values for heat tolerance in Australian dairy cattle. We combined test-day herd recording data with temperature and humidity measurements (in the form of temperature-humidity index or THI) from weather stations that were closest to the herds for test days between 2003 and 2013. Tolerance to heat stress was then estimated for each cow using random regression (intercept and slope) to model the rate of decline in production with increasing THI accumulated over the four days prior to the day of milking, for milk, fat and protein yields. The cow slopes from this model were used to define daughter trait deviations (DTD) for their sires. Data were analysed separately for Holsteins and Jerseys. The reference population for genomic prediction was 2,300 Holstein and 575 Jersey genotyped sires with DTD for response to heat stress for milk, fat and protein yield. With this reference, and using GBLUP, the range in accuracy of genomic predictions for heat tolerance across traits were 0.38 – 0.53 and 0.49 – 0.63 for 435 Holstein and 135 Jersey validation sires, respectively. When 2,191 Holstein and 1,190 Jersey cows were added in the reference populations, no substantial improvements in accuracy were observed. Genomic selection appears to be a useful tool to enable farmers to improve milk production in environments with higher heat load.
QTL affecting milk production were mapped using a Bayesian method that fits all 632,002 SNPs simultaneously in a population of 11,527 Holstein and 4687 Jersey animals. We identified 265 windows of 250 kb where the variance of predicted genetic value was > 50 times the variance of an average window. Windows were clustered into 42 regions across the genome. Based on their pattern of pleiotropic effects, we grouped these regions into 9 possible groups. For example, one group had antagonistic effects on fat and milk yield whilst another group had positive effects on both traits. We confirmed many known loci and highlight strong candidates in other regions. The observed pattern of QTL effects matched the known function in milk synthesis for some loci. Using a multibreed population, fitting all SNPs simultaneously and studying QTL effects across multiple traits allowed separation of closely linked QTL and more precise positioning than previous single SNP regression methods.
An error appeared in the equation and description of the log likelihood of the method BayesR beginning on page 4119. The corrected version is as follows: The likelihood that SNP i is in distribution k isLogLi,k=−0.5logV−0.5y*′y*−y*Z*uk*σe2+logprk,where y* is the vector of phenotypes corrected for all marker effects other than marker i, the overall mean, and the polygenic effects; uk* is the mean of the posterior distribution of the SNP effect when assumed to be in the kth distribution; Z* is a column vector containing the SNP genotypes of all animals for SNP i; V is the variance-covariance structure of a reduced model including only the effect of the respective SNP and a residual effect; and logV was calculated asnlogσe2+logσk2Z*′Z*σe2+1,where Z* contains only the information for the current SNP effect. The authors regret the error. Improving accuracy of genomic predictions within and between dairy cattle breeds with imputed high-density single nucleotide polymorphism panelsJournal of Dairy ScienceVol. 95Issue 7PreviewAchieving accurate genomic estimated breeding values for dairy cattle requires a very large reference population of genotyped and phenotyped individuals. Assembling such reference populations has been achieved for breeds such as Holstein, but is challenging for breeds with fewer individuals. An alternative is to use a multi-breed reference population, such that smaller breeds gain some advantage in accuracy of genomic estimated breeding values (GEBV) from information from larger breeds. However, this requires that marker-quantitative trait loci associations persist across breeds. Full-Text PDF Open Access
Genetic parameters were estimated with the aim of identifying useful predictor traits for the genetic evaluation of fertility. For this study, data included calving interval (CI), days from calving to first service (CFS), pregnancy diagnosis; lactation length (LL), daily milk yield close to 90 d of lactation (milk yield), and survival to second lactation on Australian Holstein and Jersey cows. The effect of level of fertility, measured here as CI, on correlations among traits was investigated by dividing the Holstein herds into those that managed short CI (proxy for seasonal-calving herds) and long CI (proxy for herds that practice extended lactations). In all cases, genetic correlations of CI with CFS, pregnancy, and LL were high (>0.7). Genetic correlations between fertility and predictor traits were generally similar in the 2 Holstein herd groups and in Jerseys. However, some differences in both the direction and strength of correlations were observed. In Jerseys, the genetic correlation between CI and survival was positive; but in Holstein herds, this correlation was negative. Particularly in low mean CI herds, the correlation suggests that cows with a genetic potential for longer CI were more likely to be culled. The genetic correlation of CI with survival was intermediate in high mean CI Holstein herds. Furthermore. Jersey cows with a high genetic potential for milk yield had a higher chance of surviving than those with low genetic potential. In contrast, the genetic correlation between milk yield and survival in low mean CI Holstein herds was near zero. The high genetic correlation between CI and LL suggests that LL could be used as proxy for CI in cows that do not calve again. Although the phenotypic variance for CI in high mean CI herds was nearly twice that in Jerseys and low mean CI herds, we found no bull reranking for CI due to having daughters in low or high mean CI herds. However, the ranges in estimated breeding values (EBV) were narrower in low mean CI herds than in high mean CI herds. The genetic trend in cows and bulls showed that CI EBV were increasing by 0.3 to 0.8 d/yr in both Holstein and Jersey. Phenotypically, CI was increasing by 2 d/yr in high mean CI Holstein herds and by 1 d/yr in Jersey and low mean CI Holstein herds. However, in recent years, both phenotypic and genetic trends have stabilized. In summary, if the main trait for genetic evaluation of fertility is CI, predictor traits such as milk yield, survival, LL, and other fertility traits can be used in joint analyses to increase reliability of bull EBV. If the genetic evaluation is to be carried out simultaneously for Holstein and Jersey using the same variance-covariance matrix, survival should not be used as a predictor because its correlation with CI is different in Jersey than in Holstein. On the other hand, LL could be used instead of CI for cows that do not calve again in both breeds and herd groups.
Single nucleotide polymorphism (SNP) associations with milk production traits found to be significant in different screening experiments, including SNP in genes hypothesized to be in gene pathways affecting milk production, were tested in a validation population to confirm their association. In total, 423 SNP were genotyped across 411 Holstein bulls, and their association with 6 milk production traits--Australian Selection Index (indicating the profitability of an animal's milk production), protein, fat, and milk yields, and protein and fat composition--were tested using single SNP regressions. Seventy-two SNP were significantly associated with one or more of the traits; their effects were in the same direction as in the screening experiment and therefore their association was considered validated. An over-representation of SNP (43 of the 423) on chromosome 20 was observed, including a SNP in the growth hormone receptor gene previously published as having an association with protein composition and protein and milk yields. The association with protein composition was confirmed in this experiment, but not the association with protein and milk yields. A multiple SNP regression analysis for all SNP on chromosome 20 was performed for all 6 traits, which revealed that this mutation was not significantly associated with any of the milk production traits and that at least 2 other quantitative trait loci were present on chromosome 20.
Achieving accurate genomic estimated breeding values for dairy cattle requires a very large reference population of genotyped and phenotyped individuals. Assembling such reference populations has been achieved for breeds such as Holstein, but is challenging for breeds with fewer individuals. An alternative is to use a multi-breed reference population, such that smaller breeds gain some advantage in accuracy of genomic estimated breeding values (GEBV) from information from larger breeds. However, this requires that marker-quantitative trait loci associations persist across breeds. Here, we assessed the gain in accuracy of GEBV in Jersey cattle as a result of using a combined Holstein and Jersey reference population, with either 39,745 or 624,213 single nucleotide polymorphism (SNP) markers. The surrogate used for accuracy was the correlation of GEBV with daughter trait deviations in a validation population. Two methods were used to predict breeding values, either a genomic BLUP (GBLUP_mod), or a new method, BayesR, which used a mixture of normal distributions as the prior for SNP effects, including one distribution that set SNP effects to zero. The GBLUP_mod method scaled both the genomic relationship matrix and the additive relationship matrix to a base at the time the breeds diverged, and regressed the genomic relationship matrix to account for sampling errors in estimating relationship coefficients due to a finite number of markers, before combining the 2 matrices. Although these modifications did result in less biased breeding values for Jerseys compared with an unmodified genomic relationship matrix, BayesR gave the highest accuracies of GEBV for the 3 traits investigated (milk yield, fat yield, and protein yield), with an average increase in accuracy compared with GBLUP_mod across the 3 traits of 0.05 for both Jerseys and Holsteins. The advantage was limited for either Jerseys or Holsteins in using 624,213 SNP rather than 39,745 SNP (0.01 for Holsteins and 0.03 for Jerseys, averaged across traits). Even this limited and nonsignificant advantage was only observed when BayesR was used. An alternative panel, which extracted the SNP in the transcribed part of the bovine genome from the 624,213 SNP panel (to give 58,532 SNP), performed better, with an increase in accuracy of 0.03 for Jerseys across traits. This panel captures much of the increased genomic content of the 624,213 SNP panel, with the advantage of a greatly reduced number of SNP effects to estimate. Taken together, using this panel, a combined breed reference and using BayesR rather than GBLUP_mod increased the accuracy of GEBV in Jerseys from 0.43 to 0.52, averaged across the 3 traits.
Feed makes up a large proportion of variable costs in dairying. For this reason, selection for traits associated with feed conversion efficiency should lead to greater profitability of dairying. Residual feed intake (RFI) is the difference between actual and predicted feed intakes and is a useful selection criterion for greater feed efficiency. However, measuring individual feed intakes on a large scale is prohibitively expensive. A panel of DNA markers explaining genetic variation in this trait would enable cost-effective genomic selection for this trait. With the aim of enabling genomic selection for RFI, we used data from almost 2,000 heifers measured for growth rate and feed intake in Australia (AU) and New Zealand (NZ) genotyped for 625,000 single nucleotide polymorphism (SNP) markers. Substantial variation in RFI and 250-d body weight (BW250) was demonstrated. Heritabilities of RFI and BW250 estimated using genomic relationships among the heifers were 0.22 and 0.28 in AU heifers and 0.38 and 0.44 in NZ heifers, respectively. Genomic breeding values for RFI and BW250 were derived using genomic BLUP and 2 bayesian methods (BayesA, BayesMulti). The accuracies of genomic breeding values for RFI were evaluated using cross-validation. When 624,930 SNP were used to derive the prediction equation, the accuracies averaged 0.37 and 0.31 for RFI in AU and NZ validation data sets, respectively, and 0.40 and 0.25 for BW250 in AU and NZ, respectively. The greatest advantage of using the full 624,930 SNP over a reduced panel of 36,673 SNP (the widely used BovineSNP50 array) was when the reference population included only animals from either the AU or the NZ experiment. Finally, the bayesian methods were also used for quantitative trait loci detection. On chromosome 14 at around 25 Mb, several SNP closest to PLAG1 (a gene believed to affect stature in humans and cattle) had an effect on BW250 in both AU and NZ populations. In addition, 8 SNP with large effects on RFI were located on chromosome 14 at around 35.7 Mb. These SNP may be associated with the gene NCOA2, which has a role in controlling energy metabolism.
Using data from almost 2,000 Holstein-Friesian dairy heifers measured for growth rate and feed intake in both Australia and New Zealand (NZ), we demonstrated substantial variation in residual feed intake (RFI) and 250-day-liveweight (LWT250d). The respective heritabilities of RFI and LWT250d were 0.25 and 0.31 in Australian data and 0.41 and 0.25 in NZ data. Further, using around 630,000 SNP markers, genomic breeding values for RFI and LWT250d could be predicted with a moderate degree of accuracy (RFI: 0.41 and 0.31 in Australian and NZ data respectively; LWT250d: 0.41 and 0.25 in Australian and NZ data respectively).
Three breeds (Fleckvieh, Holstein, and Jersey) were included in a reference population, separately and together, to assess the accuracy of prediction of genomic breeding values in single-breed validation populations. The accuracy of genomic selection was defined as the correlation between estimated breeding values, calculated using phenotypic data, and genomic breeding values. The Holstein and Jersey populations were from Australia, whereas the Fleckvieh population (dual-purpose Simmental) was from Austria and Germany. Both a BLUP with a multi-breed genomic relationship matrix (GBLUP) and a Bayesian method (BayesA) were used to derive the prediction equations. The hypothesis tested was that having a multi-breed reference population increased the accuracy of genomic selection. Minimal advantage existed of either GBLUP or BayesA multi-breed genomic evaluations over single-breed evaluations. However, when the goal was to predict genomic breeding values for a breed with no individuals in the reference population, using 2 other breeds in the reference was generally better than only 1 breed.
Although genomic selection offers the prospect of improving the rate of genetic gain in meat, wool and dairy sheep breeding programs, the key constraint is likely to be the cost of genotyping. Potentially, this constraint can be overcome by genotyping selection candidates for a low density (low cost) panel of SNPs with sparse genotype coverage, imputing a much higher density of SNP genotypes using a densely genotyped reference population. These imputed genotypes would then be used with a prediction equation to produce genomic estimated breeding values. In the future, it may also be desirable to impute very dense marker genotypes or even whole genome re-sequence data from moderate density SNP panels. Such a strategy could lead to an accurate prediction of genomic estimated breeding values across breeds, for example. We used genotypes from 48 640 (50K) SNPs genotyped in four sheep breeds to investigate both the accuracy of imputation of the 50K SNPs from low density SNP panels, as well as prospects for imputing very dense or whole genome re-sequence data from the 50K SNPs (by leaving out a small number of the 50K SNPs at random). Accuracy of imputation was low if the sparse panel had less than 5000 (5K) markers. Across breeds, it was clear that the accuracy of imputing from sparse marker panels to 50K was higher if the genetic diversity within a breed was lower, such that relationships among animals in that breed were higher. The accuracy of imputation from sparse genotypes to 50K genotypes was higher when the imputation was performed within breed rather than when pooling all the data, despite the fact that the pooled reference set was much larger. For Border Leicesters, Poll Dorsets and White Suffolks, 5K sparse genotypes were sufficient to impute 50K with 80% accuracy. For Merinos, the accuracy of imputing 50K from 5K was lower at 71%, despite a large number of animals with full genotypes (2215) being used as a reference. For all breeds, the relationship of individuals to the reference explained up to 64% of the variation in accuracy of imputation, demonstrating that accuracy of imputation can be increased if sires and other ancestors of the individuals to be imputed are included in the reference population. The accuracy of imputation could also be increased if pedigree information was available and was used in tracking inheritance of large chromosome segments within families. In our study, we only considered methods of imputation based on population-wide linkage disequilibrium (largely because the pedigree for some of the populations was incomplete). Finally, in the scenarios designed to mimic imputation of high density or whole genome re-sequence data from the 50K panel, the accuracy of imputation was much higher (86-96%). This is promising, suggesting that in silico genome re-sequencing is possible in sheep if a suitable pool of key ancestors is sequenced for each breed.
Genome-wide association studies (GWAS) were used to discover genomic regions explaining variation in dairy production and fertility traits. Associations were detected with either single nucleotide polymorphism (SNP) markers or haplotypes of SNP alleles. An across-breed validation strategy was used to narrow the genomic interval containing causative mutations. There were 39,048 SNP tested in a discovery population of 780 Holstein sires and validated in 386 Holsteins and 364 Jersey sires. Previously identified mutations affecting milk production traits were confirmed. In addition, several novel regions were identified, including a putative quantitative trait loci for fertility on chromosome 18 that was detected only using haplotypes greater than 3 SNP long. It was found that the precision of quantitative trait loci mapping increased with haplotype length as did the number of validated haplotypes discovered, especially across breed. Promising candidate genes have been identified in several of the validated regions.