BACKGROUND:Negative energy balance (NEB) during the transition period is associated with profound changes in the body condition and metabolic dynamics of dairy cows. However, the detailed lipidomic changes in milk induced by NEB are unclear, and lipid biomarkers that indicate the energy status of cows remain to be established. METHODS:Using a combination of GC-FID, HILIC-MS and RP-LC-MS, we performed a systematic comparison of lipid composition between early lactating (DIM: 5-14) and mid-lactating (DIM: 65-80) milk. RESULTS:We found that NEB in cows caused a profound modification in the profile of all the lipid classes surveyed, including phosphatidylcholine (PC), phosphatidylethanolamine (PE), phosphatidylserine (PS), phosphatidylinositol (PI), sphingomyelin (SM), lysophosphatidylcholine (LPC), PC-plasmalogen (PCP), PE-plasmalogen (PEP), lactosylceramide (LacCer), acylcarnitine (AcylCar) and triglycerides (TAGs). Except for LPC and AcylCar, which were reduced and increased, respectively, by NEB, the responses of other lipid classes varied across different species. For phospholipids and TAGs, species containing de novo FAs (C4:0-C16:0) and odd-chain FAs (C15:0 and C17:0) were markedly downregulated, whereas those comprising long-chain preformed FAs were upregulated by NEB. CONCLUSIONS:Comprehensive lipidomic profiling of early and mid-lactating milk from two large cohorts of cows allowed us to identify nine lipids (PE 33:1, LacCer 32:1, LacCer 39:1, LacCer 41:1, SM 36:1, SM 36:2, SM 37:1, PEP 38:4 and PEP 38:5) as potential biomarkers of NEB in dairy cows.
The current clinical test for ketosis in dairy cows, which is based on measuring blood β-hydroxybutyrate (BHB) concentrations, is an invasive procedure that impacts herd welfare. Alternative biomarkers, especially non-invasive ones, are urgently needed for subclinical ketosis (SCK) diagnosis. To address this, we examined the milk lipidomic profiles of healthy and SCK cows from a research dairy farm over two consecutive years. A total of 125 polar lipid species were quantified in milk samples using a targeted liquid chromatography–mass spectrometry (LC-MS) approach. Chemometric analysis identified four potential biomarkers (PC 28:0, PC 29:0, PC 30:1 and PC 31:0), which met the thresholds for p value (<0.0004) and fold change (FC > 2). Global fatty acid (FA) profiling by gas chromatography–flame ionization detection (GC-FID) after transesterification of all lipids revealed that the levels of C8:0, C10:0, C12:0, C14:0, C15:0 and C18:1t11 were significantly suppressed in the SCK group (15–52% lower compared to the control group, p < 0.01), and C15:0 and C18:1t11 were the most promising markers for SCK in dairy cows at the FA level (43–50% reduction compared to the control group, p < 0.0001). Tandem LC-MS and FA data generated by GC demonstrated that SCK also caused significant changes in regio-isomer and double bond isomer composition of a model phospholipid molecule (PC 34:1) and a model triglyceride molecule TAG 52:2; these fine-structural alterations could also be explored for discovering novel SCK biomarkers in dairy cows. The PLS-DA predictive model built from lipidomic data showed excellent accuracy (>95%) in predicting cows with SCK. All of the promising biomarkers identified and the predictive model built in this study need further validation using larger SCK cohorts from different herds.
Background: The presence and concentration of lipids in serum of dairy cows have significant implications for both animal health and productivity and are potential biomarkers for several common diseases. However, information on serum lipid composition is rather fragmented, and lipid remodelling during the transition period is only partially understood. Methods: Using a combination of reversed-phase liquid chromatography-mass spectrometry (RP-LC-MS), hydrophilic interaction-mass spectrometry (HILIC-MS), and lipid annotation software, we performed a comprehensive identification and quantification of serum of dairy cows in pasture-based Holstein-Friesian cows. The lipid remodelling induced by negative energy balance was investigated by comparing the levels of all identified lipids between the fresh lactation (5–14 days in milk, DIM) and full lactation (65–80 DIM) stages. Results: We identified 535 lipid molecular species belonging to 19 classes. The most abundant lipid class was cholesteryl ester (CE), followed by phosphatidylcholine (PC), sphingomyelin (SM), and free fatty acid (FFA), whereas the least abundant lipids included phosphatidylserine (PS), phosphatidic acid (PA), phosphatidylglycerol (PG), acylcarnitine (AcylCar), ceramide (Cer), glucosylceramide (GluCer), and lactosylceramide (LacCer). Conclusions: A remarkable increase in most lipids and a dramatic decrease in FFAs, AcylCar, and DHA-containing species were observed at the full lactation compared to fresh lactation stage. Several serum lipid biomarkers for detecting negative energy balance in cows were also identified.
Background Meta-analysis describes a category of statistical methods that aim at combining the results of multiple studies to increase statistical power by exploiting summary statistics. Different industries that use genomic prediction do not share their raw data due to logistic or privacy restrictions, which can limit the size of their reference populations and creates a need for a practical meta-analysis method. Results We developed a meta-analysis, named MetaGS, that duplicates the results of multi-trait best linear unbiased prediction (mBLUP) analysis without accessing raw data. MetaGS exploits the correlations among different populations to produce more accurate population-specific single nucleotide polymorphism (SNP) effects. The method improves SNP effect estimations for a given population depending on its relations to other populations. MetaGS was tested on milk, fat and protein yield data of Australian Holstein and Jersey cattle and it generated very similar genomic estimated breeding values to those produced using the mBLUP method for all traits in both breeds. One of the major difficulties when combining SNP effects across populations is the use of different variants for the populations, which limits the applications of meta-analysis in practice. We solved this issue by developing a method to impute missing summary statistics without using raw data. Our results showed that imputing summary statistics can be done with high accuracy ( r > 0.9) even when more than 70% of the SNPs were missing with a minimal effect on prediction accuracy. Conclusions We demonstrated that MetaGS can replace the mBLUP model when raw data cannot be shared, which can lead to more flexible collaborations compared to the single-trait BLUP model.
Abstract Background Sequence-based genome-wide association studies (GWAS) provide high statistical power to identify candidate causal mutations when a large number of individuals with both sequence variant genotypes and phenotypes is available. A meta-analysis combines summary statistics from multiple GWAS and increases the power to detect trait-associated variants without requiring access to data at the individual level of the GWAS mapping cohorts. Because linkage disequilibrium between adjacent markers is conserved only over short distances across breeds, a multi-breed meta-analysis can improve mapping precision. Results To maximise the power to identify quantitative trait loci (QTL), we combined the results of nine within-population GWAS that used imputed sequence variant genotypes of 94,321 cattle from eight breeds, to perform a large-scale meta-analysis for fat and protein percentage in cattle. The meta-analysis detected (p ≤ 10−8) 138 QTL for fat percentage and 176 QTL for protein percentage. This was more than the number of QTL detected in all within-population GWAS together (124 QTL for fat percentage and 104 QTL for protein percentage). Among all the lead variants, 100 QTL for fat percentage and 114 QTL for protein percentage had the same direction of effect in all within-population GWAS. This indicates either persistence of the linkage phase between the causal variant and the lead variant across breeds or that some of the lead variants might indeed be causal or tightly linked with causal variants. The percentage of intergenic variants was substantially lower for significant variants than for non-significant variants, and significant variants had mostly moderate to high minor allele frequencies. Significant variants were also clustered in genes that are known to be relevant for fat and protein percentages in milk. Conclusions Our study identified a large number of QTL associated with fat and protein percentage in dairy cattle. We demonstrated that large-scale multi-breed meta-analysis reveals more QTL at the nucleotide resolution than within-population GWAS. Significant variants were more often located in genic regions than non-significant variants and a large part of them was located in potentially regulatory regions.
Additional file 2: Table S1. QTL detected for fat percentage in the meta-analysis. chr = chromosome, pos = position in base pair on UMD3.1 assembly, ARS-UCD2.1 = position in base pair on ASR-UCD2.1 assembly, p = p-value in the meta-analysis, dir = direction of effect in each of the GWAS (from left to right: Braunvieh, Fleckvieh, German Holstein, Norwegian Red, Australian bull dataset, Australian cow dataset, Montbéliarde, Normande, French Holstein), start = pos – 250 kb, end = pos + 250 kb, nSig = number of variants with a p-value ≤ 10−8 in the interval, nGenic = nSig associated with a gene, genes = genes in the interval with significant variants, cojo = p-value in COJO analysis for retained variants, or discarded/not present to indicate variants that were discarded or not included in COJO analysis), sameDirVal = indicates whether a variant had the same direction of Z-score in the meta-analysis and the validation analysis (- = not included in validation analysis, 0 = opposite direction, 1 = same direction), pVal = p-value in the validation analysis.
MicroRNAs regulate many eukaryotic biological processes in a temporal- and spatial-specific manner. Yet in cattle it is not fully known which microRNAs are expressed in each tissue, which genes they regulate, or which sites a given microRNA bind to within messenger RNAs. An improved annotation of tissue-specific microRNA network may in the future assist with the identification of causal variants affecting complex traits. Here, we report findings from analysing short RNA sequence from 17 tissues from a single lactating dairy cow. Using miRDeep2, we identified 699 expressed mature microRNA sequences. Using TargetScan, known (60%) and novel (40%) microRNAs were predicted to interact with 780,481 sites in bovine messenger RNAs homologous with human. Putative interactions between microRNA families and targets were significantly enriched for interactions from previous experimental and computational identification. Characterizing features of microRNAs and targets, we showed that (1) mature microRNAs derived from different arms of the same precursor targeted different genes in different tissues; (2) miRNA target sites preferentially occurred within gene regions marked with active histone modification; (3) variants within microRNAs and targets had lower allele frequencies than variants across the genome, as identified from 65 million whole genome sequence variants; (4) no significant correlation was found between the abundance of microRNAs and messenger RNAs differentially expressed in the same tissue; (5) microRNAs and target sites weren’t significantly associated with allelic imbalance of gene targets. This study contributes to the goals of Functional Annotation of Animal Genomes consortium to improve the annotation of genomes of domestic animals.
Background Two distinct populations have been extensively studied in Atlantic cod ( Gadus morhua L.): the Northeast Arctic cod (NEAC) population and the coastal cod (CC) population. The objectives of the current study were to identify genomic islands of divergence and to propose an approach to quantify the strength of selection pressures using whole-genome single nucleotide polymorphism (SNP) data. After applying filtering criteria, information on 93 animals (9 CC individuals, 50 NEAC animals and 34 CC × NEAC crossbred individuals) and 3,123,434 autosomal SNPs were used. Results Four genomic islands of divergence were identified on chromosomes 1, 2, 7 and 12, which were mapped accurately based on SNP data and which extended in size from 11 to 18 Mb. These regions differed considerably between the two populations although the differences in the rest of the genome were small due to considerable gene flow between the populations. The estimates of selection pressures showed that natural selection was substantially more important than genetic drift in shaping these genomic islands. Our data confirmed results from earlier publications that suggested that genomic islands are due to chromosomal rearrangements that are under strong selection and reduce recombination between rearranged and non-rearranged segments. Conclusions Our findings further support the hypothesis that selection and reduced recombination in genomic islands may promote speciation between these two populations although their habitats overlap considerably and migrations occur between them.
Stature is affected by many polymorphisms of small effect in humans 1 . In contrast, variation in dogs, even within breeds, has been suggested to be largely due to variants in a small number of genes2,3. Here we use data from cattle to compare the genetic architecture of stature to those in humans and dogs. We conducted a meta-analysis for stature using 58,265 cattle from 17 populations with 25.4 million imputed whole-genome sequence variants. Results showed that the genetic architecture of stature in cattle is similar to that in humans, as the lead variants in 163 significantly associated genomic regions (P < 5 × 10−8) explained at most 13.8% of the phenotypic variance. Most of these variants were noncoding, including variants that were also expression quantitative trait loci (eQTLs) and in ChIP–seq peaks. There was significant overlap in loci for stature with humans and dogs, suggesting that a set of common genes regulates body size in mammals. Meta-analysis of data from 58,265 cattle shows that the genetic architecture underlying stature is similar to that in humans, where many genomic regions individually explain only a small amount of phenotypic variance.
Background Topological association domains (TADs) are chromosomal domains characterised by frequent internal DNA-DNA interactions. The transcription factor CTCF binds to conserved DNA sequence patterns called CTCF binding motifs to either prohibit or facilitate chromosomal interactions. TADs and CTCF binding motifs control gene expression, but they are not yet well defined in the bovine genome. In this paper, we sought to improve the annotation of bovine TADs and CTCF binding motifs, and assess whether the new annotation can reduce the search space for cis- regulatory variants. Results We used genomic synteny to map TADs and CTCF binding motifs from humans, mice, dogs and macaques to the bovine genome. We found that our mapped TADs exhibited the same hallmark properties of those sourced from experimental data, such as housekeeping genes, transfer RNA genes, CTCF binding motifs, short interspersed elements, H3K4me3 and H3K27ac. We showed that runs of genes with the same pattern of allele-specific expression (ASE) (either favouring paternal or maternal allele) were often located in the same TAD or between the same conserved CTCF binding motifs. Analyses of variance showed that when averaged across all bovine tissues tested, TADs explained 14% of ASE variation (standard deviation, SD: 0.056), while CTCF explained 27% (SD: 0.078). Furthermore, we showed that the quantitative trait loci (QTLs) associated with gene expression variation (eQTLs) or ASE variation (aseQTLs), which were identified from mRNA transcripts from 141 lactating cows’ white blood and milk cells, were highly enriched at putative bovine CTCF binding motifs. The linearly-furthermost, and most-significant aseQTL and eQTL for each genic target were located within the same TAD as the gene more often than expected (Chi-Squared test P -value < 0.001). Conclusions Our results suggest that genomic synteny can be used to functionally annotate conserved transcriptional components, and provides a tool to reduce the search space for causative regulatory variants in the bovine genome.
BACKGROUND:The increasing availability of whole-genome sequence data is expected to increase the accuracy of genomic prediction. However, results from simulation studies and analysis of real data do not always show an increase in accuracy from sequence data compared to high-density (HD) single nucleotide polymorphism (SNP) chip genotypes. In addition, the sheer number of variants makes analysis of all variants and accurate estimation of all effects computationally challenging. Our objective was to find a strategy to approximate the analysis of whole-sequence data with a Bayesian variable selection model. Using a simulated dataset, we applied a Bayes R hybrid model to analyse whole-sequence data, test the effect of dropping a proportion of variants during the analysis, and test how the analysis can be split into separate analyses per chromosome to reduce the elapsed computing time. We also investigated the effect of imputation errors on prediction accuracy. Subsequently, we applied the approach to a dataset that contained imputed sequences and records for production and fertility traits for 38,492 Holstein, Jersey, Australian Red and crossbred bulls and cows.RESULTS:With the simulated dataset, we found that prediction accuracy was highly increased for a breed that was not represented in the training population for sequence data compared to HD SNP data. Either dropping part of the variants during the analysis or splitting the analysis into separate analyses per chromosome decreased accuracy compared to analysing whole-sequence data. First, dropping variants from each chromosome and reanalysing the retained variants together resulted in an accuracy similar to that obtained when analysing whole-sequence data. Adding imputation errors decreased prediction accuracy, especially for errors in the validation population. With real data, using sequence variants resulted in accuracies that were similar to those obtained with the HD SNPs.CONCLUSIONS:We present an efficient approach to approximate analysis of whole-sequence data with a Bayesian variable selection model. The lack of increase in prediction accuracy when applied to real data could be due to imputation errors, which demonstrates the importance of developing more accurate methods of imputation or directly genotyping sequence variants that have a major effect in the prediction equation.
In livestock, genotype × environment interaction (G × E) has been widely investigated, with genotype defined at the level of subspecies, breeds, individual animals within a breed (for example performance of offspring of elite sires across environments), and genotypes at single‐nucleotide polymorphisms (SNPs). Environments can be described by category (e.g., tropical vs. temperate, high vs. low farm input levels, countries) and by continuous variables such as temperature. To predict breeding values of genotypes in environments described by categories, multitrait models with each category a different trait are used. The models are now being used to predict genomic estimated breeding values (GEBV) for different environments such as the value of a bull's genetics for his daughter's milk production in different countries. The multitrait genomic model has also been used to enable reference populations to be merged across environments and across countries, leading to more accurate GEBV. When the environment can be described by a continuous variable, random regression models have been used to predict response of genotypes to the environment. For example, these models have been used to determine if there are SNP genotypes associated with less sensitivity of milk production to increasing temperature. In both livestock and plant breeding, methods that use genomic information can better cope with a reduced degree of replication of individuals across environments, as it is actually the alleles that must be replicated across environments. More accurate estimates of G × E with the genomic approach may therefore be achievable than was possible in the past.
Whole genome sequence data (BAM format) of 234 bovine individuals aligned to UMD3.1. The aim of the study was to identify genetic variants (SNPs and indels) for downstream analysis such as imputation, GWAS, and detection of lethal recessives. Additional sequences for later 1000 bull genomes runs can be found at partners individual projects including PRJEB9343, PRJNA176557, PRJEB18113, PRNJA343262, PRJNA324822, PRJNA324270, PRJNA277147, PRJEB5462.
The 1000 bull genomes project supports the goal of accelerating the rates of genetic gain in domestic cattle while at the same time considering animal health and welfare by providing the annotated sequence variants and genotypes of key ancestor bulls. In the first phase of the 1000 bull genomes project, we sequenced the whole genomes of 234 cattle to an average of 8.3-fold coverage. This sequencing includes data for 129 individuals from the global Holstein-Friesian population, 43 individuals from the Fleckvieh breed and 15 individuals from the Jersey breed. We identified a total of 28.3 million variants, with an average of 1.44 heterozygous sites per kilobase for each individual. We demonstrate the use of this database in identifying a recessive mutation underlying embryonic death and a dominant mutation underlying lethal chrondrodysplasia. We also performed genome-wide association studies for milk production and curly coat, using imputed sequence variants, and identified variants associated with these traits in cattle.
Whole-genome sequence is potentially the richest source of genetic data for inferring ancestral demography. However, full sequence also presents significant challenges to fully utilize such large data sets and to ensure that sequencing errors do not introduce bias into the inferred demography. Using whole-genome sequence data from two Holstein cattle, we demonstrate a new method to correct for bias caused by hidden errors and then infer stepwise changes in ancestral demography up to present. There was a strong upward bias in estimates of recent effective population size (Ne) if the correction method was not applied to the data, both for our method and the Li and Durbin (Inference of human population history from individual whole-genome sequences. Nature 475:493-496) pairwise sequentially Markovian coalescent method. To infer demography, we use an analytical predictor of multiloci linkage disequilibrium (LD) based on a simple coalescent model that allows for changes in Ne. The LD statistic summarizes the distribution of runs of homozygosity for any given demography. We infer a best fit demography as one that predicts a match with the observed distribution of runs of homozygosity in the corrected sequence data. We use multiloci LD because it potentially holds more information about ancestral demography than pairwise LD. The inferred demography indicates a strong reduction in the Ne around 170,000 years ago, possibly related to the divergence of African and European Bos taurus cattle. This is followed by a further reduction coinciding with the period of cattle domestication, with Ne of between 3,500 and 6,000. The most recent reduction of Ne to approximately 100 in the Holstein breed agrees well with estimates from pedigrees. Our approach can be applied to whole-genome sequence from any diploid species and can be scaled up to use sequence from multiple individuals.
Three recent breakthroughs have resulted in the current widespread use of DNA information: the genomic selection (GS) methodology, which is a form of marker-assisted selection on a genome-wide scale, and the discovery of large numbers of single-nucleotide markers and cost effective methods to genotype them. GS estimates the effect of thousands of DNA markers simultaneously. Nonlinear estimation methods yield higher accuracy, especially for traits with major genes. The marker effects are estimated in a genotyped and phenotyped training population and are used for the estimation of breeding values of selection candidates by combining their genotypes with the estimated marker effects. The benefits of GS are greatest when selection is for traits that are not themselves recorded on the selection candidates before they can be selected. In the future, genome sequence data may replace SNP genotypes as markers. This could increase GS accuracy because the causative mutations should be included in the data.
A major aim of the research program known as SheepGENOMICS was to deliver DNA markers for commercial breeding programs. To that end, a resource flock was established, comprehensively phenotyped and genotyped with DNA markers. The flock of nearly 5000 sheep, born over two consecutive years, was extensively phenotyped, with more than 100 recorded observations being made on most of the animals. This generated more than 460000 records over 17 months of gathering information on each animal. Here, we describe the experimental design and sample-collection procedures, and provide a summary of the basic measurements taken. Data from this project are being used to identify collections of genome markers for estimating genomic breeding values for new sheep industry traits.
The genetic architecture of complex traits in cattle includes very large numbers of loci affecting any given trait. Most of these loci have small effects but occasionally there are loci with moderate-to-large effects segregating due to recent selection for the mutant allele. Genomic markers capture most but not all of the additive genetic variance for traits, probably because there are causal mutations with low allele frequency and therefore in incomplete linkage disequilibrium with the markers. The prediction of genetic value from genomic markers can achieve high accuracy by using statistical models that include all markers and assuming that marker effects are random variables drawn from a specified prior distribution. Recent effective population size is in the order of 100 within cattle breeds and ≈ 2500 animals with genotypes and phenotypes are sufficient to predict the genetic value of animals with an accuracy of 0.65. Recent effective population size for humans is much larger, in the order of 10,000-15,000, and more than 145,000 records would be required to reach a similar accuracy for people. However, our calculations assume that genomic markers capture all the genetic variance. This may be possible in the future as causal polymorphisms are genotyped using genome sequence data.
Related individuals share potentially long chromosome segments that trace to a common ancestor. We describe a phasing algorithm (ChromoPhase) that utilizes this characteristic of finite populations to phase large sections of a chromosome. In addition to phasing, our method imputes missing genotypes in individuals genotyped at lower marker density when more densely genotyped relatives are available. ChromoPhase uses a pedigree to collect an individual’s (the proband) surrogate parents and offspring and uses genotypic similarity to identify its genomic surrogates. The algorithm then cycles through the relatives and genomic surrogates one at a time to find shared chromosome segments. Once a segment has been identified, any missing information in the proband is filled in with information from the relative. We tested ChromoPhase in a simulated population consisting of 400 individuals at a marker density of 1500/M, which is approximately equivalent to a 50K bovine single nucleotide polymorphism chip. In simulated data, 99.9% loci were correctly phased and, when imputing from 100 to 1500 markers, more than 87% of missing genotypes were correctly imputed. Performance increased when the number of generations available in the pedigree increased, but was reduced when the sparse genotype contained fewer loci. However, in simulated data, ChromoPhase correctly imputed at least 12% more genotypes than fastPHASE, depending on sparse marker density. We also tested the algorithm in a real Holstein cattle data set to impute 50K genotypes in animals with a sparse 3K genotype. In these data 92% of genotypes were correctly imputed in animals with a genotyped sire. We evaluated the accuracy of genomic predictions with the dense, sparse, and imputed simulated data sets and show that the reduction in genomic evaluation accuracy is modest even with imperfectly imputed genotype data. Our results demonstrate that imputation of missing genotypes, and potentially full genome sequence, using long-range phasing is feasible.