Allele-specific expression (ASE) is the imbalanced expression of two alleles of the same locus. It is quite pervasive among many species and is associated with health and economically relevant traits. ASE is often used to support the identification of variants related to gene expression (cis-eQTL). Thus, profiling allele-specific expression represents a significant step in elucidating the mechanism underlying gene expression regulation. In this study, we developed an ASE pipeline using publicly available RNA-seq data and open-source software. Using this pipeline, we are able to profile pervasive allelic imbalance across 42 tissues and 34 breeds from the Farm-GTEx-pig consortium at both SNP and gene levels without the need for parental genotypes or whole genome sequence data. We find that ASE is widely, but not evenly, spread across the genome. We also observe considerable variation in ASE profiles across various tissues, where the site fraction ranged from 1.3
Biological control is a sustainable strategy to combat agricultural pests. Yet, legislation increasingly restricts importing nonnative biocontrol agents. Thus, selective breeding of biocontrol traits is suggested to enhance performance of existing biocontrol agents. Genomic prediction, where genomic data are used to estimate the genetic merit of an individual for specific traits, is an alternative to exploit genetic variation for the improvement of native biocontrol agents. This study aims to establish a proof of principle for genomic prediction in insect biocontrol agents, using wing morphology traits in the model parasitoid Nasonia vitripennis Walker (Pteromalidae). We performed genomic prediction using a genomic best linear unbiased prediction (GBLUP) model, using 1,230 individuals with 8,639 SNPs generated by genotyping-by-sequencing (GBS). We used individuals from 2 generations from the outbred HVRx population, 717 individuals from generation 169 (G169) and 513 individuals from generation 172 (G172). To assess genomic prediction accuracy, we used across generation validation (forward validation for G172 from G169 and backward validation for G169 from G172) and also 5-fold cross-validation. For size-related traits, including tibia length, wing length, wing width, and second moment area, the accuracy of genomic prediction was close to 0 in both across generation validations but much higher in 5-fold cross-validation (ranging from 0.54 to 0.68). For the shape-related trait wing aspect ratio, a high accuracy was found for all 3 validation strategies, with 0.47 for across generation forward validation (AGFV), 0.65 for across generation backward validation (AGBV), and 0.54 for 5-fold cross-validation. Overall, genomic selection in insect biocontrol agents with a relative small effective population size seems promising. However, factors such as the biology of insects, phenotyping techniques, and large-scale genotyping costs still challenge the application of genomic selection to biocontrol agents.
Combined Annotation Dependent Depletion (CADD) is a machine learning approach used to predict the deleteriousness of genetic variants across a genome. By integrating diverse genomic features, CADD assigns a PHRED-like rank score to each potential variant. Unlike other methods, CADD does not rely on limited datasets of known pathogenic or benign variants but uses larger and less biased training sets. The rapid increase in high-quality genomes and functional annotations across species highlights the need for an automated, non-species-specific pipeline to generate CADD scores. Here, we introduce such a pipeline, facilitating the generation of CADD scores for various species using only a high-quality genome with gene annotation and a multi-species alignment. Additionally, we present updated chickenCADD scores and newly generated turkeyCADD scores, both generated with the pipeline.
Salinity tolerance varies among tilapia species and salinity-tolerant strains can be produced through selective breeding. In Indonesia, breeding programs have produced Sukamandi, a tilapia strain with improved growth in brackish water. Here, we investigated the salinity tolerance of this strain by performing a transcriptomic analysis using offsprings from the fifth generation of a breeding population. Based on estimated breeding values, we formed three genetic groups: brackish water-specialized tilapia, freshwater-specialized tilapia, and generalist tilapia. Progeny (average weight of similar to 100 g) from these groups were stocked in both brackish water (25 ppt) and freshwater (0 ppt) tanks for short (3 days) and long term (28 days) exposure periods. Our results showed that brackish water reared fish had higher Na+ and Cl- concentrations in blood plasma while haematocrit levels were not affected by salinity across all genetic groups. For Na+/K+-ATPase concentration, the interaction between genetic groups and salinity environment was significant. Ion homeostasis and sodium-potassium-chloride symporter activities were key processes exhibited by brackish water relative to freshwater fish. Short-term brackish water acclimation uniquely triggered the activation of prolactin receptor activity and calcium signaling pathway, whereas long-term exposure was specifically associated with immune responses, DNA damage, and apoptosis. Upon checking the Differentially Expressed Genes within osmoregulation-related terms common in all tested groups, six candidate responsive genes to osmotic stress were identified: aqp3a, atp1a1b, slc12a2, slc45a3, slc12a10.1 and slc1210.3. Overall, this study sheds light into the genetic factors influencing the adaptation of Sukamandi strain to salinity stress and identify candidate genes that can be potentially used as markers for salinity tolerance breeding in the future.
The Holstein Friesian (HF) cattle breed is the most dominant breed in commercial dairy farming worldwide and managed in more than 150 countries. These countries span diverse agro-climatic zones, ranging from tropical to cold regions. The introduction of HF animals in these regions occurred at different moments in the past which are poorly recorded and continued through importation of live animal and frozen semen. We hypothesize that the HF cattle populations in these regions underwent early forms of adaptation to these specific local environments. However, the detection of genetic variation associated with this adaptation remains poorly documented. This study investigates genetic relationship and potential early selection signatures in HF populations from three African countries (Egypt, South Africa, Uganda) and three European countries (Finland, Portugal, The Netherlands), considering five animals per country. Approximately 16.0 million single nucleotide polymorphisms (SNPs) were detected in the 30 HF animals and used for further analyses. Across all countries, we identified dispersed regions totaling 3.3 megabase of ecosystem-specific genomic regions (43 genes), indicative of early selection signatures based on fixation indices (F-statistic, Fst). Furthermore, comparing variants between tropical (Egypt and Uganda) and cold regions (Finland and The Netherlands) by Fst, nucleotide diversity (θπ ratio), and extended haplotype homozygosity (XP-EHH), we identified a total of 10 candidate regions, comprising 12 genes within a 0.57 megabase size. The regions were enriched with genes involved in signaling pathways associated directly or indirectly with adaptation, including the immune system (PGLYRP4,PGLYRP3, PAG1, CD48, SLAMF1, DYSF,and LOC615223), organ development and reproduction (LDB3, ADAMTSL4, TPRN, CCDC40, OR2AG1G, and OR8B3), thermogenic activation (TBC1D16), phospholipid metabolism (PLPPR4 and PITPNB), thermos-tolerance (ZNF423), and stimulus response (NCOA7, CYP2C85, and ARFGEF3). This study provides new insights into early forms of genetic plasticity of animals adapted to very diverse ecosystems. Our findings highlight candidate genes related to immune response, organ development, reproduction, metabolism, and thermo-tolerance, hypothesizing their role in facilitating adaptation to different environments.
Genetic mutation and drift, coupled with natural and human-mediated selection and migration, have produced a wide variety of genotypes and phenotypes in farmed animals. We here introduce the Farm Animal Genotype-Tissue Expression (FarmGTEx) Project, which aims to elucidate the genetic determinants of gene expression across 16 terrestrial and aquatic domestic species under diverse biological and environmental contexts. For each species, we aim to collect multiomics data, particularly genomics and transcriptomics, from 50 tissues of 1,000 healthy adults and 200 additional animals representing a specific context. This Perspective provides an overview of the priorities of FarmGTEx and advocates for coordinated strategies of data analysis and resource-sharing initiatives. FarmGTEx aims to serve as a platform for investigating context-specific regulatory effects, which will deepen our understanding of molecular mechanisms underlying complex phenotypes. The knowledge and insights provided by FarmGTEx will contribute to improving sustainable agriculture-based food systems, comparative biology and eventual human biomedicine.
BACKGROUND:Investigating the functional impact of genomic variants is essential to uncover the molecular mechanisms behind complex traits. This study compiled a comprehensive dataset of 1,817 whole-genome sequences from diverse pig breeds and populations, capturing the global pig genetic diversity. RESULTS:Our analyses first revealed 27,167 loss-of-function variants (LoFs), the majority of which also influenced gene expression and splicing, and enriched in genomic regions associated with complex traits in pigs. We further genome-wide annotated non-coding variants, and then focused on these resided in 5' untranslated region (5'UTR). Although they had lower deleterious impact on protein sequence compared to coding variants, they enriched in promoters and exhibited functional consequences on gene expression and splicing and finally complex traits. We employed the Basenji deep learning model and ATAC-seq to predict the impact of these SNPs on chromatin accessibility in 13 pig tissues. SNPs with higher predicted scores demonstrated stronger effects on gene expression/splicing and complex traits-particularly average backfat thickness-compared to variants with lower scores. CONCLUSIONS:In summary, our study provides a comprehensive catalog of genomic variants in both protein-coding and non-coding regions, and elucidated their functional consequences on epigenome, transcriptome, and complex traits in pigs.
In horses, genetic diversity is predominantly observed between breeds, with little variation within breeds. The studbooks of the two largest horse populations in the Netherlands, the Dutch Warmblood horse and Friesian horse population, have ongoing conservation projects including collecting large-scale genotype and sequence data. The current reference genome, derived from a Thoroughbred horse can lead to bias in genetic analyses of other horse breeds. Therefore, the aim of this study was to create high-quality breed-specific reference genomes of Dutch Warmblood and Friesian horses. We performed nanopore long-read sequencing (R10.4, Q20+) of an F1 cross between a Dutch Warmblood horse and a Friesian horse to create two breed-specific reference genomes by trio binning. This resulted in high-quality, haplotype-resolved reference genomes with contig N50 of 37 and 35 Mb and single copy gene completeness of 99.2 and 99.3
Direct estimates of mutation rates in humans have changed our understanding of evolutionary timing and de novo mutations (DNM) have been associated with several developmental disorders in humans. Livestock species, including pigs, can contribute to the study of DNM because of their ideal population structure and routine phenotype collection. In principle, there is the potential for livestock populations to quickly accumulate new genetic variants because of short generation intervals and high selection intensity. However, the impact of DNM on the fitness of individuals is not known and with current genomic selection programs they cannot contribute to estimated breeding values. The aims of our project were to detect and validate single nucleotide DNM in two commercial pig breeding lines, estimate the single nucleotide mutation rate, and characterise DNM. We sequenced (150 bp paired end reads, 30X coverage) 46 pig trios from two commercial lines. Single nucleotide DNM were detected using a trio-aware method. We defined candidate DNM as single nucleotide variants (SNVs) found in heterozygous state in trio-offspring with both trio-parents homozygous for the reference allele. In this study, we estimate a lower threshold of the DNM rate in pigs of 6.3 × 10–9 per site per gamete. Our findings are consistent with those from other mammals and those published for a small number of livestock species. Most DNM we detected were in introns (47
Structural variants (SVs) are one of the main sources of genetic variants and have a significant impact on phenotype evolution, disease susceptibility, and environmental adaptations. We used 73 whole genome sequencing (12x) to apply a mapping approach to identify SVs in five turkey populations. A notable degree of genetic isolation was observed between the Basilicata and Apulian populations, as indicated by principal component analysis and admixture results. A total of 11,733 SVs were detected, including 6712 deletions, 2671 duplications, 1430 inversions, and 920 translocations. The Variant Effect Predictor (VEP) analysis predicted various consequences of filtered SVs as follows: intron variants (35.8%), intergenic variants (9.6%), coding sequence variants (8.3%), downstream gene variants (7.5%), and transcript ablations (7.3%). Our functional annotation of genes overlapping with SVs was mainly enriched in recognized pathways governing positive regulation of nucleoplasm, protein binding, mitochondrion, negative regulation of cell population proliferation, identical protein binding, and calcium signaling. We produced a comprehensive SV catalog utilizing unique whole-genome turkey data. This SV catalog not only increases our understanding of genetic diversity in turkeys but also enhances our knowledge of the role of SVs in their phenotypic traits.
Basilicata and Apulian (BAS-APU) turkeys, a native population in the Basilicata and Puglia regions of southern Italy, are known for their high meat quality and tolerance to local conditions. Understanding the genomic patterns of BAS-APU turkeys is critical for effective breeding and preservation strategies. In this study, we characterized runs of homozygosity (ROH), and selection signatures using the integrated haplotype score (iHS) and ROH approaches. A total of 73 BAS-APU turkeys from five populations were sequenced (12X). The inbreeding coefficients based on ROH ranged from 0.177 to 0.405. A total of 120,956 ROH were detected in BAS-APU populations. We identified 27 genomic regions that harbor 61 candidate genes in ROH islands in which single nucleotide polymorphisms (SNPs) occur in more than 90% of individuals. In addition, we detected 608 genomic regions under positive selection using the iHS method being 104, 98, 130, 102, and 174 for BAS, APU_C, APU_M, APU_PN, and APU_PS, respectively. For both methods, most of the genes within these regions are related to production performance, reproduction, immune responses, and adaptation. This study contributes significantly to our understanding of the genetic makeup of native turkey populations in southern Italy. The identified genes under selection can aid future breeding and conservations programs for southern Italian native turkeys. The results of inbreeding levels, especially in the absence of complete pedigrees or when only a few samples are available, which is often the case for local breeds, will help to avoid genetic relatedness in the mating plan in breeding and conservation plans for BAS-APU populations. Also, the detected genes in the selective sweep regions could be used as a marker-assisted selection to improve productive traits and adaptation of BAS-APU local populations.
This study introduces, for the first time, whole-genome sequencing (WGS) data from predominantly wild-born Asian elephants currently housed in European zoos, covering the distribution range of Asian elephants. With this WGS data, we aim to validate the current designation of Asian elephant subspecies and address currently discussed ambiguities about their origin, particularly concerning Bornean and Sri Lankan elephants by analyzing population structure, determining divergence times, and exploring ancient and recent bottlenecks. Understanding the evolutionary history of the Asian elephant subspecies is essential for developing targeted conservation strategies and mitigating risks to their survival. Analysis reveals a clear population structure with relatively recent splits, delineating three distinct genetic clusters: Borneo, Sumatra, and Asian Mainland, with Sri Lanka forming an additional group. We estimated the divergence time between Bornean and Sumatran elephants to be around 170,000 years ago. The divergence of the Sri Lankan elephant from the Mainland is estimated to have occurred around 48,000 years ago, with Sri Lankan elephants predominantly clustering with those from Myanmar, possibly due to historical trade networks. The genome of the Bornean elephant exhibited signatures of severe bottlenecks as recently as 8 and 38 generations ago, further supporting hypotheses of their introduction. Our data reflect the current Asian elephant subspecies designation. Additionally, for the first time, the Sumatra elephant is confirmed as a distinct subspecies with genomic data. Furthermore, the study discusses genetic management strategies for ex-situ populations, emphasizing the importance of implementing cluster-specific conservation measures.
Wild boars exhibit genetic and phenotypic diversity shaped by migrations and local adaptations. Their expansion across Eurasia, especially in Central Asia, remains underexplored. Here, we present newly sequenced whole-genome data of 47 wild boars from Eastern Asia, Central Asia, and Europe, combined with 49 existing genomes, creating a comprehensive dataset of 96 individuals. Our analyses show that Asian wild boars and Southeast Asian Suids split ∼3.6 million years ago (mya), with Central Asian and Southern Chinese ancestors diverging ∼1.8 mya. The split between Central Asian and European-Near East ancestors occurred ∼0.9 mya, followed by a European-Near East divergence ∼0.6 mya. We identify signatures of local adaptation in Central Asian populations, including two positively selected variants in LPIN1, associated with lipid metabolism, and a missense mutation in ALPK2, linked to meat traits. These findings provide insights into wild boar dispersal and adaptation and shed light on domestic pig breeding.
The Asian elephant (Elephas maximus L.) is classified as an endangered species, comprising four recognized subspecies: Indian (E. m. indicus), Sri Lankan (E. m. maximus), Sumatran (E. m. sumatranus), and Borneo (E. m. borneensis). Elephant populations in Southeast Asia, though small and fragmented, face high risks of extirpation due to habitat loss, poaching, human-wildlife conflict, and climate change. These factors jeopardize their survival and highlight the urgent need for targeted conservation efforts. Despite these challenges, Asian elephants possess crucial genetic diversity that needs to be maintained for future adaptive potential, making their conservation a high priority. Genetic studies are essential for informing conservation strategies. This review aims to compile and summarize the relevant literature on the genetic data of Asian elephants, specifically focusing on their phylogenetic relationships, historical biogeography, and phylogeography, while emphasizing the need for acquiring genomic data. In addition, we explore how important captive populations have been in acquiring genomic data for this endangered species. It also highlights the importance of genetically monitoring captive populations to maintain sufficient genetic variation for conservation and research purposes. We discuss how understanding the elephants’ evolutionary history from a genomic perspective can offer insights into subspecies recognition and provide a data-driven foundation for planning management strategies, such as reintroduction, translocation, and captive breeding. Ultimately, these efforts will enhance conservation strategies and secure the survival of this iconic species in the face of ongoing anthropogenic and environmental challenges.
BACKGROUND:Long non-coding RNAs (lncRNAs), a type of non-coding RNA molecules, are known to play critical regulatory roles in various biological processes. However, the functions of the majority of lncRNAs remain largely unknown, and little is understood about the regulation of lncRNA expression. In this study, high-throughput DNA genotyping and RNA sequencing were applied to investigate genomic regions associated with lncRNA expression, commonly referred to as lncRNA expression quantitative trait loci (eQTLs). We analyzed the liver, lung, spleen, and muscle transcriptomes of 100 three-way crossbred sows to identify lncRNA transcripts, explore genomic regions that might influence lncRNA expression, and identify potential regulators interacting with these regions. RESULT:We identified 6380 lncRNA transcripts and 3733 lncRNA genes. Correlation tests between the expression of lncRNAs and protein-coding genes were performed. Subsequently, functional enrichment analyses were carried out on protein-coding genes highly correlated with lncRNAs. Our correlation results of these protein-coding genes uncovered terms that are related to tissue specific functions. Additionally, heatmaps of lncRNAs and protein-coding genes at different correlation levels revealed several distinct clusters. An expression genome-wide association study (eGWAS) was conducted using 535,896 genotypes and 1829, 1944, 2089, and 2074 expressed lncRNA genes for liver, spleen, lung, and muscle, respectively. This analysis identified 520,562 significant associations and 6654, 4525, 4842, and 7125 eQTLs for the respective tissues. Only a small portion of these eQTLs were classified as cis-eQTLs. Fifteen regions with the highest eQTL density were selected as eGWAS hotspots and potential mechanisms of lncRNA regulation in these hotspots were explored. However, we did not identify any interactions between the transcription factors or miRNAs in the hotspots and the lncRNAs, nor did we observe a significant enrichment of regulatory elements in these hotspots. While we could not pinpoint the key factors regulating lncRNA expression, our results suggest that the regulation of lncRNAs involves more complex mechanisms. CONCLUSION:Our findings provide insights into several features and potential functions of lncRNAs in various tissues. However, the mechanisms by which lncRNA eQTLs regulate lncRNA expression remain unclear. Further research is needed to explore the regulation of lncRNA expression and the mechanisms underlying lncRNA interactions with small molecules and regulatory proteins.
The recognition that climate change is occurring at an unprecedented rate means that there is increased urgency in understanding how organisms can adapt to a changing environment. Wild great tit (Parus major) populations represent an attractive ecological model system to understand the genomics of climate adaptation. They are widely distributed across Eurasia and they have been documented to respond to climate change. We performed a Bayesian genome-environment analysis, by combining local climate data with single nucleotide polymorphisms genotype data from 20 European populations (broadly spanning the species’ continental range). We found 36 genes putatively linked to adaptation to climate. Following an enrichment analysis of biological process Gene Ontology (GO) terms, we identified over-represented terms and pathways among the candidate genes. Because many different genes and GO terms are associated with climate variables, it seems likely that climate adaptation is polygenic and genetically complex. Our findings also suggest that geographical climate adaptation has been occurring since great tits left their Southern European refugia at the end of the last ice age. Finally, we show that substantial climate-associated genetic variation remains, which will be essential for adaptation to future changes.
A major aim of evolutionary biology is to understand why patterns of genomic diversity vary within taxa and space. Large-scale genomic studies of widespread species are useful for studying how environment and demography shape patterns of genomic divergence. Here, we describe one of the most geographically comprehensive surveys of genomic variation in a wild vertebrate to date; the great tit (Parus major) HapMap project. We screened ca 500,000 SNP markers across 647 individuals from 29 populations, spanning similar to 30 degrees of latitude and 40 degrees of longitude - almost the entire geographical range of the European subspecies. Genome-wide variation was consistent with a recent colonisation across Europe from a South-East European refugium, with bottlenecks and reduced genetic diversity in island populations. Differentiation across the genome was highly heterogeneous, with clear 'islands of differentiation', even among populations with very low levels of genome-wide differentiation. Low local recombination rates were a strong predictor of high local genomic differentiation (FST), especially in island and peripheral mainland populations, suggesting that the interplay between genetic drift and recombination causes highly heterogeneous differentiation landscapes. We also detected genomic outlier regions that were confined to one or more peripheral great tit populations, probably as a result of recent directional selection at the species' range edges. Haplotype-based measures of selection were related to recombination rate, albeit less strongly, and highlighted population-specific sweeps that likely resulted from positive selection. Our study highlights how comprehensive screens of genomic variation in wild organisms can provide unique insights into spatio-temporal evolutionary dynamics.
The development of a comprehensive pig graph pangenome assembly encompassing 27 genomes represents the most extensive collection of pig genomic data to date. Analysis of this pangenome reveals the critical role of structural variations in driving adaptation and defining breed-specific traits. Notably, the study identifies BTF3 as a key candidate gene governing intramuscular fat deposition and meat quality in pigs. These findings underscore the power of pangenome approaches in uncovering novel genomic features underlying economically important agricultural traits. Collectively, these results demonstrate the value of leveraging large-scale, multi-genome analyses for advancing our understanding of livestock genomes and accelerating genetic improvement.
The genetic basis of complex traits and phenotypic differentiation remains unclear in pigs. Using nine genomes-seven of which were newly generated, high-quality de novo assembled genomes-and 1081 resequencing genomes, we built a pan-genome and identified 134.24 Mb nonredundant nonreference sequences, 1099 novel protein-coding genes, 187,927 structural variations (SVs) and 30,143,962 single-nucleotide polymorphisms (SNPs). Analysis of selective domestication revealed BRCA1 associated with enhanced adipocyte growth and fat deposition, and ABCA3 linked to an alleviated immune response and reduced lung injury. Integrating 162 transcriptomes and 162 methylomes of skeletal muscle across 27 developmental stages revealed the regulatory mechanism of phenotypic differentiation between Eastern and Western breeds. Artificial selection reshaped local DNA methylation status and imparted regulatory effects on the progression patterns of heterochronic genes such as GHSR and BDH1, particularly during embryonic development. Altogether, our work provides valuable resources for understanding molecular mechanisms behind phenotypic variations and enhancing the genetic improvement programs in pigs.
BackgroundIntegration of high throughput DNA genotyping and RNA-sequencing data enables the discovery of genomic regions that regulate gene expression, known as expression quantitative trait loci (eQTL). In pigs, efforts to date have been mainly focused on purebred lines for traits with commercial relevance as such growth and meat quality. However, little is known on genetic variants and mechanisms associated with the robustness of an animal, thus its overall health status. Here, the liver, lung, spleen, and muscle transcriptomes of 100 three-way crossbred female finishers were studied, with the aim of identifying novel eQTL regulatory regions and transcription factors (TFs) associated with regulation of porcine metabolism and health-related traits.ResultsAn expression genome-wide association study with 535,896 genotypes and the expression of 12,680 genes in liver, 13,310 genes in lung, 12,650 genes in spleen, and 12,595 genes in muscle resulted in 4,293, 10,630, 4,533, and 6,871 eQTL regions for each of these tissues, respectively. Although only a small fraction of the eQTLs were annotated as cis-eQTLs, these presented a higher number of polymorphisms per region and significantly stronger associations with their target gene compared to trans-eQTLs. Between 20 and 115 eQTL hotspots were identified across the four tissues. Interestingly, these were all enriched for immune-related biological processes. In spleen, two TFs were identified: ERF and ZNF45, with key roles in regulation of gene expression.ConclusionsThis study provides a comprehensive analysis with more than 26,000 eQTL regions identified that are now publicly available. The genomic regions and their variants were mostly associated with tissue-specific regulatory roles. However, some shared regions provide new insights into the complex regulation of genes and their interactions that are involved with important traits related to metabolism and immunity.