Background: The Siberian Husky has evolved as a versatile dog capable of traversing over 1600 km in extreme Arctic conditions, being a competitive show dog in the American Kennel Club, or a favorite pet for companionship. Modern genomics provides an opportunity to explore the biological implications of selection within the Siberian Husky breed for the purpose of sledding, show, or pet. Methods: We identified regions of genetic selection associated with sledding, show, or pet purposes using a whole-genome panel of 234 K SNPs from 237 Siberian Huskies. We assessed allelic variation using Wright’s FST and selective sweeps with runs of homozygosity (ROH). Results: Genomic and morphometric measurement principal component analyses identified population structure aligning with breeding purpose. In total, 118 SNPs demonstrated significant allelic variation (FST ≥ 0.6) and 22,598 ROH segments were identified within the Siberian Husky breed. ROH islands (n = 91) highlighted selective sweeps, whereas homozygosity association tests characterized regions of the genome under differential selection between populations. Genes within regions were assessed using GO and KEGG pathway analysis for biological insight. Pet dogs showed selection for olfactory performance genes, whereas show dogs were selected for immune function, tissue and nervous system development, and cytoskeletal motor activity. Sledding Siberian Huskies were selected for the development of muscle organs, lung vasculature, limbs, bones, eye structure, and pigmentation, plus genes influencing lipid metabolism and glucose transport. Conclusions: In all, this provides the first evidence of the biological impact of genetic selection within a breed for the distinct sledding, show, and pet purposes while simultaneously maintaining overall population uniformity to meet breed standards.
This study investigates the efficacy of various genomic prediction models—Genomic Best Linear Unbiased Prediction (GBLUP), Random Forest (RF), Support Vector Machine (SVM), Extreme Gradient Boosting (XGB), and Multilayer Perceptron (MLP)—in predicting genomic breeding values (gEBVs). The phenotypic data include three binary health traits (anodontia, distichiasis, oral papillomatosis) and one behavioral trait (distraction) in a population of guide dogs. These traits impact the potential for success in guide dogs and are therefore routinely characterized but were chosen based on differences in heritability and case counts specifically to assess gEBV model performance. Utilizing a dataset from The Seeing Eye organization, which includes German Shepherds (n = 482), Golden Retrievers (n = 239), Labrador Retrievers (n = 1188), and Labrador and Golden Retriever crosses (n = 111), we assessed model performance within and across different breeds, trait heritability, case counts, and SNP marker densities. Our results indicate that no significant differences were found in model performance across varying heritabilities, case counts, or SNP densities, with all models performing similarly. Given its lack of need for parameter optimization, GBLUP was the most efficient model. Distichiasis showed the highest overall predictive performance, likely due to its higher heritability, while anodontia and distraction exhibited moderate accuracy, and oral papillomatosis had the lowest accuracy, correlating with its low heritability. These findings underscore that lower density SNP datasets can effectively construct gEBVs, suggesting that high-cost, high-density genotyping may not always be necessary. Additionally, the similar performance of all models indicates that simpler models like GBLUP, which requires less fine tuning, may be sufficient for genomic prediction in canine breeding programs. The research highlights the importance of standardized phenotypic assessments and carefully constructed reference populations to optimize the utility of genomic selection in canine breeding programs.
Introduction:Genomic breeding values and multi-trait selection indices have significantly advanced genetic improvement in livestock but remain underutilized in guide dog breeding. This study developed a genomically informed selection framework for a population of Labrador Retrievers by integrating health (e.g., dental, ocular, and dermatological conditions) and behavioral (e.g., trainability, distraction level, pace) traits into a "Behavior Score," "Health Score," and "Total Score" index by applying Genomic Best Linear Unbiased Prediction (GBLUP) to estimate breeding values. Results:Phenotypic and genotypic data were collected from 844 dogs over 26 years at The Seeing Eye guide dog school. Predictive performance was evaluated via five-fold cross-validation and correlation-based metrics. Results showed that some dentition related health traits exhibited moderate to high Area Under Receiving Operating Characteristic (AUROC) values (0.79-0.87), indicating potential for immediate use for genetic improvement. In contrast, most other health traits demonstrated weak to moderate predictive accuracy. Behavioral traits exhibited lower predictive accuracy but showed a stronger association with training success. Models were commonly unable to correctly classify individuals for binary or ordinal traits yet performed well in ranking individuals, likely due to lower heritability or strong environmental influences of traits or limitations of the dataset itself. The behavior-focused Total Score (AUROC ~0.72) outperformed health-based indices as a fixed effect in predicting breeding success despite the weaker predictive ability of individual behavioral traits. Incorporating parental scores as fixed effects modestly improved breeding values for success, indicating the importance of integrating additional data sources where available. Discussion:While these findings underscore the utility of genomic selection for guide dog breeding, they also highlight constraints stemming from small, genetically homogeneous populations and variable phenotyping. Ultimately, we provide the first usable individual and multi-trait genomic approaches to enhance both health and performance outcomes in working dog programs and a foundation to expand upon the reference population and behavioral trait assessment to improve prediction accuracy in the future.
Recent evidence demonstrates genomic and morphological continuity in the Arctic ancestral lineage of dogs. Here, we use the Siberian Husky to investigate the genomic legacy of the northeast Eurasian Arctic lineage and model the deep population history using genome-wide single nucleotide polymorphisms. Utilizing ancient dog-calibrated molecular clocks, we found that at least two distinct lineages of Arctic dogs existed in ancient Eurasia at the end of the Pleistocene. This pushes back the origin of sled dogs in the northeast Siberian Arctic with humans likely intentionally selecting dogs to perform different functions and keeping breeding populations that overlap in time and space relatively reproductively isolated. In modern Siberian Huskies, we found significant population structure based on how they are used by humans, recent European breed introgression in about half of the dogs that participate in races, moderate levels of inbreeding, and fewer potentially harmful variants in populations under strong selection for form and function (show, sled show, and racing populations of Siberian Huskies). As the struggle to preserve unique evolutionary lineages while maintaining genetic health intensifies across pedigreed dogs, understanding the genomic history to guide policies and best practices for breed management is crucial to sustain these ancient lineages and their unique evolutionary identity.
Global warming is expected to result in larger temperature fluctuations by which heat stress may become an important stressor for animals, affecting health and productivity. Animals can cope with and adapt to heat stress by changing their physiology. To investigate general physiological reactions to heat stress in muscle and heart tissues of chickens we combined results from three independent experiments. Two experiments studied the transcriptome profiles of heart and muscle tissues of mature chickens using heat stress adapted and nonadapted chickens. One experiment studied the epigenome changes of heat stress during chicken egg incubation. In all three datasets epigenome changes were important biological response mechanisms, which may underlie genome-wide regulation of the affected biological mechanisms. Pre and postnatal heat stress reaction showed changed expression of genes related to metabolic rate, energy, and protein metabolism. Furthermore, tissue integrity may be affected due to changed cell−cell contacts, vascularization, and growth reduction.
We reconstruct the phenotype of Balto, the heroic sled dog renowned for transporting diphtheria antitoxin to Nome, Alaska, in 1925, using evolutionary constraint estimates from the Zoonomia alignment of 240 mammals and 682 genomes from dogs and wolves of the 21st century. Balto shares just part of his diverse ancestry with the eponymous Siberian husky breed. Balto’s genotype predicts a combination of coat features atypical for modern sled dog breeds, and a slightly smaller stature. He had enhanced starch digestion compared with Greenland sled dogs and a compendium of derived homozygous coding variants at constrained positions in genes connected to bone and skin development. We propose that Balto’s population of origin, which was less inbred and genetically healthier than that of modern breeds, was adapted to the extreme environment of 1920s Alaska.
Egg production is an important economic trait and a key indicator of reproductive performance in ducks. Egg production is regulated by several factors including genes. However the genes involved in egg production in duck remain unclear. In this study, we compared the ovarian transcriptome of high egg laying (HEL) and low egg laying (LEL) ducks using RNA-Seq to identify the genes involved in egg production. The HEL ducks laid on average 433 eggs while the LEL ducks laid 221 eggs over 93 weeks. A total of 489 genes were found to be significantly differentially expressed out of which 310 and 179 genes were up and downregulated, respectively, in the HEL group. Thirty-eight differentially expressed genes (DEGs), including LHX9, GRIA1, DBH, SYCP2L, HSD17B2, PAR6, CAPRIN2, STC2, and RAB27B were found to be potentially related to egg production and folliculogenesis. Gene ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway analysis suggested that DEGs were enriched for functions related to glutamate receptor activity, serine-type endopeptidase activity, immune function, progesterone mediated oocyte maturation and MAPK signaling. Protein–protein interaction network analysis (PPI) showed strong interaction between 32 DEGs in two distinct clusters. Together, these findings suggest a mix of genetic and immunological factors affect egg production, and highlights candidate genes and pathways, that provides an understanding of the molecular mechanisms regulating egg production in ducks and in birds more broadly.
Background Molecular studies on egg production in ducks were mostly focused on brain and ovaries as they are directly involved in egg production. Liver plays a vital role in cellular lipid metabolism. It also plays a decisive role in reproductive organ development, including yolk generation in laying ducks at sexual maturity. However, the precise molecular mechanism involved in the liver-blood-ovary axis in ducks remains elusive. Methods and Results In this study, we analysed the liver transcriptome of laying (LA), immature (IM) and broody (BR) ducks using RNA sequencing to understand the role of genes expressed in the liver. The comparative transcriptome analysis revealed 82 DEGs between LA and IM ducks, 47 DEGs between LA and BR ducks and 51 DEGs between IM and BR ducks. GO analysis of DEGs, showed that DEGs were mainly involved in cellular anatomical entity, intracellular, metabolic process, and binding. Furthermore, pathway analysis indicated the important role of Wnt signaling pathway in egg formation and embryo development. Our study showed several candidate genes including vitellogenin-1, vitellogenin-2, riboflavin binding protein, G protein subunit gamma 4, and fatty acid binding protein 3 that are potentially related to egg production in ducks. Conclusions The study provides valuable information on the genes responsible for egg production and thus, pave the way for further investigation on the molecular mechanisms of egg production in duck.
Congenital laryngeal paralysis (CLP) is an inherited disorder that affects the ability of the dog to exercise and precludes it from functioning as a working sled dog. Though CLP is known to occur in Alaskan sled dogs (ASDs) since 1986, the genetic mutation underlying the disease has not been reported. Using a genome-wide association study (GWAS), we identified a 708 kb region on CFA 18 harboring 226 SNPs to be significantly associated with CLP. The significant SNPs explained 47.06% of the heritability of CLP. We narrowed the region to 431 kb through autozygosity mapping and found 18 of the 20 cases to be homozygous for the risk haplotype. Whole genome sequencing of two cases and a control ASD, and comparison with the genome of 657 dogs from various breeds, confirmed the homozygous status of the risk haplotype to be unique to the CLP cases. Most of the dogs that were homozygous for the risk allele had blue eyes. Gene annotation and a gene-based association study showed that the risk haplotype encompasses genes implicated in developmental and neurodegenerative disorders. Pathway analysis showed enrichment of glycoproteins and glycosaminoglycans biosynthesis, which play a key role in repairing damaged nerves. In conclusion, our results suggest an important role for the identified candidate region in CLP.
[This corrects the article DOI: 10.5187/jast.2020.62.6.765.].
Pig as a food source serves daily dietary demand to a wide population around the world. Preference of meat depends on various factors with muscle play the central role. In this regards, selective breeding abled us to develop “Nanchukmacdon” a pig breeds with an enhanced variety of meat and high fertility rate. To identify genomic regions under selection we performed whole-genome resequencing, transcriptome, and whole-genome bisulfite sequencing from Nanchukmacdon muscles samples and used published data for three other breeds such as Landrace, Duroc, Jeju native pig and analyzed the functional characterization of candidate genes. In this study, we present a comprehensive approach to identify candidate genes by using multi-omics approaches. We performed two different methods XP-EHH, XP-CLR to identify traces of artificial selection for traits of economic importance. Moreover, RNAseq analysis was done to identify differentially expressed genes in the crossed breed population. Several genes (UGT8, ZGRF1, NDUFA10, EBF3, ELN, UBE2L6, NCALD, MELK, SERP2, GDPD5, and FHL2) were identified as selective sweep and differentially expressed in muscles related pathways. Furthermore, nucleotide diversity analysis revealed low genetic diversity in Nanchukmacdon for identified genes in comparison to related breeds and whole-genome bisulfite sequencing data shows the critical role of DNA methylation pattern in identified genes that leads to enhanced variety of meat. This work demonstrates a way to identify the molecular signature and lays a foundation for future genomic enabled pig breeding.
Meat from Korean native chickens (KNCs) has high consumer demand; however, slow growth performance and high variation in body weight (BW) of KNCs remain an issue. Genome-wide association study (GWAS) is a powerful method to identify quantitative trait-associated genomic loci. A GWAS, based on a large-scale KNC population, is needed to identify underlying genetic mechanisms related to its growth traits. To identify BW-associated genomic regions, we performed a GWAS using the chicken 60K single nucleotide polymorphism (SNP) panel for 1328 KNCs. BW was measured at 8 weeks of age, from 2018 to 2020. Twelve SNPs were associated with BW at the suggestive significance level (p < 2.95 × 10−5) and located near or within 11 candidate genes, including WDR37, KCNIP4, SLIT2, PPARGC1A, MYOCD and ADGRA3. Gene set enrichment analysis based on the GWAS results at p < 0.05 (1680 SNPs) showed that 32 Gene Ontology terms and two Kyoto Encyclopedia of Genes and Genomes pathways, including regulation of transcription, motor activity, the mitogen-activated protein kinase signaling pathway, and tight junction, were significantly enriched (p < 0.05) for BW-associated genes. These pathways are involved in cell growth and development, related to BW gain. The identified SNPs are potential biomarkers in KNC breeding.
Whole-genome sequence (WGS) data are increasingly being applied into genomic predictions, offering a higher predictive ability by including causal mutations or single-nucleotide polymorphisms (SNPs) putatively in strong linkage disequilibrium with causal mutations affecting the trait. This study aimed to improve the predictive performance of the customized Hanwoo 50 k SNP panel for four carcass traits in commercial Hanwoo population by adding highly predictive variants from sequence data. A total of 16,892 Hanwoo cattle with phenotypes (i.e., backfat thickness, carcass weight, longissimus muscle area, and marbling score), 50 k genotypes, and WGS imputed genotypes were used. We partitioned imputed WGS data according to functional annotation [intergenic (IGR), intron (ITR), regulatory (REG), synonymous (SYN), and non-synonymous (NSY)] to characterize the genomic regions that will deliver higher predictive power for the traits investigated. Animals were assigned into two groups, the discovery set (7324 animals) used for predictive variant detection and the cross-validation set for genomic prediction. Genome-wide association studies were performed by trait to every genomic region and entire WGS data for the pre-selection of variants. Each set of pre-selected SNPs with different density (1000, 3000, 5000, or 10,000) were added to the 50 k genotypes separately and the predictive performance of each set of genotypes was assessed using the genomic best linear unbiased prediction (GBLUP). Results showed that the predictive performance of the customized Hanwoo 50 k SNP panel can be improved by the addition of pre-selected variants from the WGS data, particularly 3000 variants from each trait, which is then sufficient to improve the prediction accuracy for all traits. When 12,000 pre-selected variants (3000 variants from each trait) were added to the 50 k genotypes, the prediction accuracies increased by 9.9, 9.2, 6.4, and 4.7% for backfat thickness, carcass weight, longissimus muscle area, and marbling score compared to the regular 50 k SNP panel, respectively. In terms of prediction bias, regression coefficients for all sets of genotypes in all traits were close to 1, indicating an unbiased prediction. The strategy used to select variants based on functional annotation did not show a clear advantage compared to using whole-genome. Nonetheless, such pre-selected SNPs from the IGR region gave the highest improvement in prediction accuracy among genomic regions and the values were close to those obtained using the WGS data for all traits. We concluded that additional gain in prediction accuracy when using pre-selected variants appears to be trait-dependent, and using WGS data remained more accurate compared to using a specific genomic region.
The estrous cycle is a complex process regulated by several hormones. To understand the dynamic changes in gene expression that takes place in the swine endometrium during the estrous cycle relative to the day of estrus onset, we performed RNA-sequencing analysis on days 0, 3, 6, 9, 12, 15, and 18, resulting in the identification of 4495 differentially expressed genes (DEGs; Q ≤ 0.05 and |log2FC| ≥ 1) at various phases in the estrous cycle. These DEGs were integrated into multiple gene co-expression networks based on different fold changes and correlation coefficient (R2) thresholds and a suitable network, which included 899 genes (|log2FC| ≥ 2 and R2 ≥ 0.99), was identified for downstream analyses based on the biological relevance of the Gene Ontology (GO) terms enriched. The genes in this network were partitioned into 6 clusters based on the expression pattern. Several GO terms including cell cycle, apoptosis, hormone signaling, and lipid biosynthetic process were found to be enriched. Furthermore, we found 15 significant KEGG pathways, including cell adhesion molecules, cytokine-cytokine receptor signaling, steroid biosynthesis, and estrogen signaling pathways. We identified several genes and GO terms to be stage-specific. Moreover, the identified genes and pathways extend our understanding of porcine endometrial regulation during estrous cycle and will serve as a good resource for future studies.
Transcriptome expression reflects genetic response in diverse conditions. In this study, RNA sequencing was utilized to profile multiple tissues such as liver, breast, caecum, and gizzard of Korean commercial chicken raised in Korea and Kyrgyzstan. We analyzed ten samples per tissue from each location to identify candidate genes which are involved in the adaptation of Korean commercial chicken to Kyrgyzstan. At false discovery rate (FDR) < 0.05 and fold change (FC) > 2, we found 315, 196, 167 and 198 genes in liver, breast, cecum, and gizzard respectively as differentially expressed between the two locations. GO enrichment analysis showed that these genes were highly enriched for cellular and metabolic processes, catalytic activity, and biological regulations. Similarly, KEGG pathways analysis indicated metabolic, PPAR signaling, FoxO, glycolysis/gluconeogenesis, biosynthesis, MAPK signaling, CAMs, citrate cycles pathways were differentially enriched. Enriched genes like TSKU, VTG1, SGK, CDK2 etc. in these pathways might be involved in acclimation of organisms into diverse climatic conditions. The qRT-PCR result also corroborated the RNA-Seq findings with R 2 of 0.76, 0.80, 0.81, and 0.93 for liver, breast, caecum, and gizzard respectively. Our findings can improve the understanding of environmental acclimation process in chicken.
Until recently, genome-scale phasing was limited due to the short read sizes of sequence data. Though the use of long-read sequencing can overcome this limitation, they require extensive error correction. The emergence of technologies such as 10X genomics linked read sequencing and Hi-C which uses short-read sequencers along with library preparation protocols that facilitates long-read assemblies have greatly reduced the complexities of genome scale phasing. Moreover, it is possible to accurately assemble phased genome of individual samples using these methods. Therefore, in this study, we compared three phasing strategies which included two sample preparation methods along with the Long Ranger pipeline of 10X genomics and HapCut2 software, namely 10X-LG, 10X-HapCut2, and HiC-HapCut2 and assessed their performance and accuracy. We found that the 10X-LG had the best phasing performance amongst the method analyzed. They had the highest phasing rate (89.6%), longest adjusted N50 (1.24 Mb), and lowest switch error rate (0.07%). Moreover, the phasing accuracy and yield of the 10X-LG stayed over 90% for distances up to 4 Mb and 550 Kb respectively, which were considerably higher than 10X-HapCut2 and Hi-C Hapcut2. The results of this study will serve as a good reference for future benchmarking studies and also for reference-based imputation in Hanwoo.
Hanwoo, is the most popular native beef cattle in South Korea. Due to its extensive popularity, research is ongoing to enhance its carcass quality and marbling traits. In this study we conducted a haplotype-based genome-wide association study (GWAS) by constructing haplotype blocks by three methods: number of single nucleotide polymorphisms (SNPs) in a haplotype block (nsnp), length of genomic region in kb (Len) and linkage disequilibrium (LD). Significant haplotype blocks and genes associated with them were identified for carcass traits such as BFT (back fat thickness), EMA (eye Muscle area), CWT (carcass weight) and MS (marbling score). Gene-set enrichment analysis and functional annotation of genes in the significantly-associated loci revealed candidate genes, including PLCB1 and PLCB4 present on BTA13, coding for phospholipases, which might be important candidates for increasing fat deposition due to their role in lipid metabolism and adipogenesis. CEL (carboxyl ester lipase), a bile-salt activated lipase, responsible for lipid catabolic process was also identified within the significantly-associated haplotype block on BTA11. The results were validated in a different Hanwoo population. The genes and pathways identified in this study may serve as good candidates for improving carcass traits in Hanwoo cattle.
The present study aimed to identify causative loci and genes enriched in pathways associated with canine obesity using a genome-wide association study (GWAS). The GWAS was first performed to identify candidate single-nucleotide polymorphisms (SNPs) associated with obesity and obesity-related traits including body weight and blood sugar in 18 different breeds of 153 dogs. A total of 10 and 2 SNPs were found to be significantly (p < 3.74 × 10-7) associated with body weight and blood sugar, respectively. None of the SNPs were identified to be significantly associated with obesity trait. We subsequently followed up the GWAS analysis with gene-set enrichment and pathway analyses. A gene-set with 1057, 1409, and 1243 SNPs annotated to 449, 933 and 820 genes for obesity, body weight, and blood sugar, respectively was created by sub-setting the GWAS result at a threshold of p < 0.01 for the gene-set enrichment analysis. In total, 84 GO and 21 KEGG pathways for obesity, 114 GO and 44 KEGG pathways for blood sugar, 120 GO and 24 KEGG pathways for body weight were found to be enriched. Among the pathways and GO terms, we highlighted five enriched pathways (Wnt signaling pathway, adherens junction, pathways in cancer, axon guidance, and insulin secretion) and seven GO terms (fat cell differentiation, calcium ion binding, cytoplasm, nucleus, phospholipid transport, central nervous system development, and cell surface) that were found to be shared among all the traits. Our data provide insights into the genes and pathways associated with obesity and obesity-related traits.
An amendment to this paper has been published and can be accessed via a link at the top of the paper.