Genotyping by sequencing (GBS) is widely employed in aquaculture selective breeding programs to acquire SNP genotypes for various applications, including pedigree reconstruction, GWAS, and in the estimation of genetic parameters, genomic relationships and breeding values. Because sequencing libraries encompass all DNA present in fin-clip tissue, they also recover the non-host metagenomic fraction, including pathogens. We leveraged this property to survey for the presence of scale-drop disease virus (SDDV), responsible for 40-90 % mortality in farmed barramundi, while simultaneously genotyping the host. Raw reads from 4484 fish of four commercial cohorts (2239 moribund, 2245 asymptomatic) were aligned to the SDDV reference genome, and viral reads were normalised as reads per million (RPM) per individual. SDDV prevalence and load were tightly associated with clinical status: 88.9 % of moribund fish carried SDDV at 21.8 f 0.6 RPM, whereas only 0.2 % of healthy fish were positive with 0.002 f 0.001 RPM. Independent validation by quantitative PCR on fin and spleen from a nested subset of 172 fish (81 moribund, 91 healthy) yielded viral copy numbers strongly correlated with ddRADseq RPM (Spearman's rho = 0.84 for fin; rho = 0.76 for spleen; both P < 0.0001). Viral load was consistently higher in fin (mean 281 f 45 copies ng-1 DNA) than spleen (147 f 32 copies ng-1), corroborating the suitability of non-lethal fin tissue for surveillance. Prevalence and load distributions were homogeneous across cohorts, and no qPCRpositive individuals escaped detection by ddRADseq. These findings show that routine ddRADseq datasets in barramundi can also be repurposed into a sensitive epidemiological assay that unites SDDV monitoring with genomic improvement. Further, these findings suggest that breeding programs generating large ddRADseq GBS datasets may also serve pathogen surveillance purposes where the target pathogen infects the host genotyped tissue.
Annotation of regulatory elements is essential for understanding mechanisms underlying gene regulation, particularly tissue-specific regulation in human and animals. Here, we characterize 274,682 enhancers and 25,975 promoters across 24 tissues from an adult female sheep using ChIP-seq, ATAC-seq, CAGE-seq, RRBS, WGBS, and RNA-seq. We identify seven neural development-related genes with over 10 enhancers in brain tissues, highlighting the role of tissue-specific regulation. Cis-regulatory enhancer-promoter combinations provide insights into tissue-specific enhancers, such as the cerebellum-specific enhancer (chr15: 57390520-57390685) regulating BDNF, which is expressed in both the cerebellum and cerebral cortex. Comparative analysis of enhancer-promoter combinations in human, mouse, pig, cattle, and sheep reveals ruminant-specific pathways, including pentose catabolism and long-chain fatty acid import regulation. A milk fat yield quantitative trait locus (QTL) identified within an enhancer interacts with the fat metabolism-related gene COMMD1, and a birth weight-associated QTL detected within a cerebellum-specific enhancer regulates XKR4. This study provides a robust framework for exploring cis-regulatory mechanisms and tissue-specific regulation, advancing the functional annotation of the sheep reference genome.
Conservation management of endangered species increasingly relies on genomic approaches to understand how long-term small population sizes affect the fitness of extant individuals. However, despite the growing investment in genomic resources by conservation programmes, the impact that sequencing methods have on the ability to detect inbreeding-related phenomena has been largely overlooked. Here, we compare the use of whole-genome and reduced-representation sequencing approaches in 148 individuals of the critically endangered parrot, the kākāpō ( Strigops habroptilus ), to assess inbreeding and its effects on female reproductive success. We explore how sequencing choice influences the identification of long stretches of homozygosity across the genome (runs of homozygosity, ROH), and compare the conservation implications of the results produced by each method. Both whole-genome and reduced-representation sequencing approaches provided comparable estimates of genome-wide inbreeding ( F ROH ) and revealed consistent effects on egg hatching success, suggesting that reduced-representation sequencing is capable of detecting inbreeding depression in wild populations under certain conditions. Whole-genome sequencing enabled chromosome-level inbreeding analyses, which revealed no strong evidence of chromosome-specific effects beyond the genome-wide signal. These results suggest that inbreeding depression in kākāpō reflects small effects across many chromosomes rather than strong effects on only a few, and that the primary benefit of whole-genome sequencing lies in improving the precision of genome-wide inbreeding estimates rather than identifying chromosome-specific effects. Our findings highlight the distinct benefits of each sequencing approach in conservation, particularly within the context of resource limitations.
Peruvian alpacas represent 85% of the global population and are mainly bred for their fiber and adaptation to high-altitude Andean environments. Two distinct phenotypes exist: Huacaya and Suri. This study aimed to investigate the genetic population structure of alpacas (Vicugna pacos) using genotyping by sequencing (GBS), also known as restriction enzyme-reduced representation sequencing, and to evaluate its application as a breeding tool. We sampled 838 Huacaya and Suri alpacas of varying coat colors and fiber qualities from ten locations in Cusco and Puno, Peru. After quality filtering, 68,641 high-confidence SNPs were obtained. These were used to estimate inbreeding, population differentiation (Fst), and observed and expected heterozygosity. Mean values of inbreeding, observed heterozygosity, and expected heterozygosity were 0.05, 0.30, and 0.34, respectively. Several animals across all sampling sites exhibited high inbreeding levels (≥0.20). Low genetic differentiation between Huacaya and Suri suggested close genetic relationships and continuous gene flow despite marked phenotypic differences. Overall, the results demonstrate that GBS provides a powerful and cost-effective approach for SNP discovery and genotyping in alpacas, offering valuable insights for genomic characterization and the potential identification of markers associated with productive traits in South American camelids managed by smallholder breeders.
The bighorn sheep (Ovis canadensis), despite its close relation to domestic sheep, suffer higher morbidity and mortality from respiratory disease complexes, likely due to genetic differences in immune responses. Unraveling highly repetitive regions such as immune loci and genetic differences was problematic until now. We generated a bighorn sheep telomere-to-telomere assembly, adding 14.28% of novel sequence compared to the previous reference. This enabled the first complete immune loci annotation revealing the IGL and TR loci are significantly short in bighorn sheep. Importantly, a critical immune gene GBP5 and ZNF501, involved in Golgi-mediated immune response, are lacking in bighorn but present in domestic sheep. Re-analysis of a Mycoplasma ovipneumoniae carriage study, using this assembly, identified the immune gene CAPN2 as a key genetic marker for disease carriage, not observable in the original study. This work provides a critical resource for identifying phenotype-linked genetic variation and exploring evolutionary adaptations of bighorn sheep.
BACKGROUND: Genotyping-by-sequencing (GBS) offers a cost-effective solution to access genomic information on many samples. The advent of GBS has promoted large scale genomic analyses, including linkage analysis, heritability estimation, inference of genetic relatedness and genome-wide prediction and association studies, in many non-model species. Low-depth GBS can be used as a cost-effective method to obtain genotypes; however, not all alleles will be captured during the sequencing process, which can lead to high levels of genotype uncertainty. New methods that take genotype uncertainty into account are needed, but these methods can be difficult to develop and benchmark on real data. Simulations provide a convenient way to benchmark different approaches, and therefore, are essential for assessing existing methods and guiding future method development. RESULTS: SimGBS is a method for simulating large-scale GBS data from any real (e.g., reference) or simulated diploid genome. It is implemented in Julia but can also be accessed from R through the JuliaCall interface. Users can define populations with distinct demographic histories by modifying population size over a specified number of generations or providing a pedigree to generate structured populations. Most GBS parameters are highly customisable, including the choice of restriction enzyme and sequencing depth. SimGBS outputs both true genotype calls and DNA sequences in FASTQ format. The method is computationally efficient, capable of generating datasets for thousands of individuals; for example, simulating 172 samples required less than two hours while successfully reproducing the complexities observed in empirical GBS data. CONCLUSION: SimGBS is a valuable tool to assist with designing GBS experiments or evaluating bioinformatic pipelines and statistical methods for GBS data analysis in large scale population genetics studies.
Molecular research for genetic variants underlying body weight (BW) provides crucial information for this important selected trait when developing productive poultry breeds, lines and crosses. We searched for molecular markers-single nucleotide polymorphisms (SNPs)-and candidate genes associated with this trait in 240 F2 resource population Japanese quails (Coturnix japonica). This population was produced by crossing two breeds with contrasting growth phenotypes, i.e., Japanese (with lower growth) and Texas White (with higher growth). The birds were genotyped using the genotyping-by-sequencing method followed by a genome-wide association study (GWAS). Using 74,387 SNPs, GWAS resulted in 142 significant SNPs and 42 candidate genes associated with BW at the age of 1, 14, 28, 35, 42, 49 and 56 days. Hereby, 25 SNPs simultaneously associated with BW at more than one age were established that colocalized with nine prioritized candidate genes (PCGs), including ITM2B, SLC35F3, ADAM33, UNC79, LEPR, RPP14, MVK, ASTN2, and ZBTB16. Twelve PCGs were identified in the regions of two or more significant SNPs, including MARCHF6, EGFR, ADGRL3, ADAM33, NPC2, LTBP2, ZC2HC1C, SATB2, ASTN2, ZBTB16, ADAR, and LGR6. These SNPs and PCGs can serve as molecular genetic markers for the genomic selection of quails with desirable BW phenotypes to enhance growth rates and meat productivity.
The deer industry in New Zealand has made notable genetic progress in recent decades. Initially based around live weight and velvet antler, deer are now selected on several traits including carcass composition, reproduction, and disease resistance within DEERSelect, the industry performance recording system in New Zealand. Due to its low cost and manageable logistics for deer, the genotyping-by-sequencing technology has replaced DNA microsatellites for parentage assignment. Since 2015 more than 60,000 animals in the national herd have been genotyped using this technology. Genomic information, however, is not yet fully exploited as evaluations currently only use pedigree information to estimate genetic merit. To assess the benefits of using genomic information, we compare pedigree, genomic and single-step genomic BLUP approaches for weaning weight, yearling weight and velvet weight at 2 years of age, in red deer from New Zealand. Using forward validation, we estimate the prediction accuracy for 56,178 animals (red and red x wapiti crossbreds) born between 1995 and 2022. We show that incorporating genomic information explicitly improves prediction accuracy by at least 15% in three key industry traits, allowing better ranking of animals across birth years. We recommend the incorporation of genomic information in the industry evaluations performed by DEERSelect.
Sequencing-based methods are increasingly being used for genetic studies. Genotypes derived from the sequence data are sometimes assigned incorrectly, due to having no reads for one of the alleles. This is particularly the case when the data has low sequencing depth. Here, we develop a correction that can be applied to both a measure and a test for the population differentiation metric F ST when considering populations as a fixed effect, to account for this incomplete genotyping. It is shown that the measure of differentiation is not influenced much by this correction in reasonably sized studies but that significance testing is too liberal without the correction. This correction will allow appropriate inference in population studies.
While conducting a landscape genomics study of invasive tammar wallabies (Notamacropus eugenii) in Aotearoa New Zealand we discovered that parma wallabies (N. parma) are also present in the North Island. This population has gone undetected for at least 30 years (and potentially for over a century), hidden amongst the morphologically similar tammar wallabies. The fact that an invasive wallaby species could remain undetected for so long, highlights the need for greater monitoring efforts for invasive species including genomic species identification.
The search for SNPs and candidate genes that determine the manifestation of major selected traits is one crucial objective for genomic selection aimed at increasing poultry production efficiency. Here, we report a genome-wide association study (GWAS) for traits characterizing meat performance in the domestic quail. A total of 146 males from an F2 reference population resulting from crossing a fast (Japanese) and a slow (Texas White) growing breed were examined. Using the genotyping-by-sequencing technique, genomic data were obtained for 115,743 SNPs (92,618 SNPs after quality control) that were employed in this GWAS. The results identified significant SNPs associated with the following traits at 8 weeks of age: body weight (nine SNPs), daily body weight gain (eight SNPs), dressed weight (33 SNPs), and weights of breast (18 SNPs), thigh (eight SNPs), and drumstick (three SNPs). Also, 12 SNPs and five candidate genes (GNAL, DNAJC6, LEPR, SPAG9, and SLC27A4) shared associations with three or more traits. These findings are consistent with the understanding of the genetic complexity of body weight-related traits in quail. The identified SNPs and genes can be used in effective quail breeding as molecular genetic markers for growth and meat characteristics for the purpose of genetic improvement.
BACKGROUND:Producing animal protein while reducing the animal's impact on the environment, e.g., through improved feed efficiency and lowered methane emissions, has gained interest in recent years. Genetic selection is one possible path to reduce the environmental impact of livestock production, but these traits are difficult and expensive to measure on many animals. The rumen microbiome may serve as a proxy for these traits due to its role in feed digestion. Restriction enzyme-reduced representation sequencing (RE-RRS) is a high-throughput and cost-effective approach to rumen metagenome profiling, but the systematic (e.g., sequencing) and biological factors influencing the resulting reference based (RB) and reference free (RF) profiles need to be explored before widespread industry adoption is possible.RESULTS:Metagenome profiles were generated by RE-RRS of 4,479 rumen samples collected from 1,708 sheep, and assigned to eight groups based on diet, age, time off feed, and country (New Zealand or Australia) at the time of sample collection. Systematic effects were found to have minimal influence on metagenome profiles. Diet was a major driver of differences between samples, followed by time off feed, then age of the sheep. The RF approach resulted in more reads being assigned per sample and afforded greater resolution when distinguishing between groups than the RB approach. Normalizing relative abundances within the sampling Cohort abolished structures related to age, diet, and time off feed, allowing a clear signal based on methane emissions to be elucidated. Genus-level abundances of rumen microbes showed low-to-moderate heritability and repeatability and were consistent between diets.CONCLUSIONS:Variation in rumen metagenomic profiles was influenced by diet, age, time off feed and genetics. Not accounting for environmental factors may limit the ability to associate the profile with traits of interest. However, these differences can be accounted for by adjusting for Cohort effects, revealing robust biological signals. The abundances of some genera were consistently heritable and repeatable across different environments, suggesting that metagenomic profiles could be used to predict an individual's future performance, or performance of its offspring, in a range of environments. These results highlight the potential of using rumen metagenomic profiles for selection purposes in a practical, agricultural setting.
The Aotearoa Genomic Data Repository (AGDR) is an initiative to provide a secure within-nation option for the storage, management and sharing of non-human genomic data generated from biological and environmental samples originating in Aotearoa New Zealand. This resource has been developed to follow the principles of Māori Data Sovereignty, and to enable the right of kaitiakitanga (guardianship), so that iwi, hapū and whānau (tribes, kinship groups and families) can effectively exercise their responsibilities as guardians over biological entities that they regard as taonga (precious or treasured). While the repository is designed to facilitate the sharing of data-making it findable by researchers and interoperable with data held in other genomic repositories-the decision-making process regarding who can access the data is entirely in the hands of those holding kaitiakitanga over each data set. No data are made available to the requesting researcher until the request has been approved, and the conditions for access (which can vary by data set) have been agreed to. Here we describe the development of the AGDR, from both a cultural perspective, and a technical one, and outline the processes that underpin its operation.
In developing countries, the use of simple and cost-efficient molecular technology is crucial for genetic characterization of local animal resources and better development of conservation strategies. The genotyping by sequencing (GBS) technique, also called restriction enzyme- reduced representational sequencing, is an efficient, cost-effective method for simultaneous discovery and genotyping of many markers. In the present study, we applied a two-enzyme GBS protocol (PstI/MspI) to discover and genotype SNP markers among 197 Tunisian sheep samples. A total of 100 333 bi-allelic SNPs were discovered and genotyped with an SNP call rate of 0.69 and mean sample depth 3.33. The genomic relatedness between 183 samples grouped the samples perfectly to their populations and pointed out a high genetic relatedness of inbred subpopulation reflecting the current adopted reproductive strategies. The genome-wide association study contrasting fat vs. thin-tailed breeds detected 41 significant variants including a peak positioned on OAR20. We identified FOXC1, GMDS, VEGFA, OXCT1, VRTN and BMP2 as the most promising for sheep tail-type trait. The GBS data have been useful to assess the population structure and improve our understanding of the genomic architecture of distinctive characteristics shaped by selection pressure in local sheep breeds. This study successfully investigates a cost-efficient method to discover genotypes, assign populations and understand insights into sheep adaptation to arid area. GBS could be of potential utility in livestock species in developing/emerging countries.
Traces of long-term artificial selection can be detected in genomes of domesticated birds via whole-genome screening using single-nucleotide polymorphism (SNP) markers. This study thus examined putative genomic regions under selection that are relevant to the development history, divergence and phylogeny among Japanese quails of various breeds and utility types. We sampled 99 birds from eight breeds (11% of the global gene pool) of egg (Japanese, English White, English Black, Tuxedo and Manchurian Golden), meat (Texas White and Pharaoh) and dual-purpose (Estonian) types. The genotyping-by-sequencing analysis was performed for the first time in domestic quails, providing 62,935 SNPs. Using principal component analysis, Neighbor-Net and Admixture algorithms, the studied breeds were characterized according to their genomic architecture, ancestry and direction of selective breeding. Japanese and Pharaoh breeds had the smallest number and length of homozygous segments indicating a lower selective pressure. Tuxedo and Texas White breeds showed the highest values of these indicators and genomic inbreeding suggesting a greater homozygosity. We revealed evidence for the integration of genomic and performance data, and our findings are applicable for elucidating the history of creation and genomic variability in quail breeds that, in turn, will be useful for future breeding improvement strategies.
In this study, we analysed the effect of human-mediated selection on the gene pool of wild and farmed red deer populations based on genotyping-by-sequencing data. The farmed red deer sample covered populations spread across seven countries and two continents (France, Germany, Hungary, Latvia, New Zealand, Poland, and Slovakia). The Slovak and Spain wild red deer populations (the latter one in a large game estate) were used as control outgroups. The gene flow intensity, relationship and admixture among populations were tested by the Bayesian approach and discriminant analysis of principal components (DAPC). The highest gene diversity (He = 0.19) and the lowest genomic inbreeding (FHOM = 0.04) found in Slovak wild population confirmed our hypothesis that artificial selection accompanied by bottlenecks has led to the increase in overall genomic homozygosity. The Bayesian approach and DAPC consistently identified three separate genetic groups. As expected, the farmed populations were clustered together, while the Slovak and Spanish populations formed two separate clusters. Identified traces of genetic admixture in the gene pool of farmed populations reflected a strong contemporary migration rate between them. This study suggests that even if the history of deer farming has been shorter than traditional livestock species, it may leave significant traces in the genome structure.
Background Rumen microbes break down complex dietary carbohydrates into energy sources for the host and are increasingly shown to be a key aspect of animal performance. Host genotypes can be combined with microbial DNA sequencing to predict performance traits or traits related to environmental impact, such as enteric methane emissions. Metagenome profiles were generated from 3139 rumen samples, collected from 1200 dual purpose ewes, using restriction enzyme-reduced representation sequencing (RE-RRS). Phenotypes were available for methane (CH4) and carbon dioxide (CO2) emissions, the ratio of CH4 to CH4 plus CO2 (CH4Ratio), feed efficiency (residual feed intake: RFI), liveweight at the time of methane collection (LW), liveweight at 8 months (LW8), fleece weight at 12 months (FW12) and parasite resistance measured by faecal egg count (FEC1). We estimated the proportion of phenotypic variance explained by host genetics and the rumen microbiome, as well as prediction accuracies for each of these traits. Results Incorporating metagenome profiles increased the variance explained and prediction accuracy compared to fitting only genomics for all traits except for CO2 emissions when animals were on a grass diet. Combining the metagenome profile with host genotype from lambs explained more than 70% of the variation in methane emissions and residual feed intake. Predictions were generally more accurate when incorporating metagenome profiles compared to genetics alone, even when considering profiles collected at different ages (lamb vs adult), or on different feeds (grass vs lucerne pellet). A reference-free approach to metagenome profiling performed better than metagenome profiles that were restricted to capturing genera from a reference database. We hypothesise that our reference-free approach is likely to outperform other reference-based approaches such as 16S rRNA gene sequencing for use in prediction of individual animal performance. Conclusions This paper shows the potential of using RE-RRS as a low-cost, high-throughput approach for generating metagenome profiles on thousands of animals for improved prediction of economically and environmentally important traits. A reference-free approach using a microbial relationship matrix from log 10 proportions of each tag normalized within cohort (i.e., the group of animals sampled at the same time) is recommended for future predictions using RE-RRS metagenome profiles.
OBJECTIVE:Growth performance and growth-related traits have a crucial role in livestock due to their influence on productivity. This genome-wide association study (GWAS) in Pakistani dromedary camels was conducted to identify single nucleotide polymorphisms (SNPs) associated with growth at specific camel ages, and for selected SNPs, to investigate in detail how their effects change with increasing camel age. This is the first GWAS conducted on dromedary camels in this region.METHODS:Two Pakistani breeds, Marecha and Lassi, were selected for this study. A genotypingby-sequencing method was used, and a total of 65,644 SNPs were identified. For GWAS, weight records data with several body weight traits, namely, birthweight, weaning weight, and weights of camels at 1, 2, 4, and 6 years of age were analysed by using model-based growth curve analysis. Age-specific weight data were analysed with a linear mixed model that included fixed effects of SNP genotype as well as sex.RESULTS:Based on the q-value method for false discovery control, for Marecha camels, five SNPs at q<0.01 and 96 at q<0.05 were significantly associated with the weight traits considered, while three (q<0.01) and seven (q<0.05) SNP associations were identified for Lassi camels. Several candidate genes harbouring these SNP were discovered.CONCLUSION:These results will help to better understand the genetic architecture of growth including how these genes are expressed at different phases of their life. This will serve to lay the foundations for applied breeding programs of camels by allowing the genetic selection of superior animals.
Characterizing the locations of genetic regulatory elements is critical for understanding the regulatory mechanisms of complex phenotypic traits related to production traits and health in livestock species. The Ovine Functional Annotation of Animal Genomes (FAANG) Project aims to characterize transcriptional regulatory elements across the sheep genome to facilitate a better understanding of the biological mechanisms influencing phenotypic traits in sheep. Assays including sequencing of messenger RNA (mRNA-seq), cap analysis of gene expression (CAGE), chromatin immunoprecipitation of histones (ChIP-seq), assay for transposase-accessible chromatin (ATAC-seq), whole genome bisulfite sequencing (WGBS) and reduced representation bisulfite sequencing (RRBS) were performed on tissues collected from the Rambouillet ewe used to assemble the reference genome ARS-UI_Ramb_v2.0. Histone modifications were used to define nine chromatin states for tissues across the genome depicting promoters and enhancers (active, poised, and repressed) using ChromHMM. Chromatin states were overlayed with RNA-seq, ATAC-seq and DNA methylation. These data suggest that active promoter and enhancer states reside in open chromatin regions with a greater transcriptional activity and hypomethylated regions than other states. Further, poised and repressed enhancers did not primarily reside in open chromatin and had less transcriptional activity and more hypermethylated sites compared with active states. Collectively these data define transcriptional regulatory regions throughout the ovine genome which provides a valuable resource to better understand regulatory regions in the genome and how these influence economically important traits in sheep and other livestock species.