The Farm Animal Genotype-Tissue Expression (FarmGTEx) project has been established to develop a public resource of genetic regulatory variants in livestock, which is essential for linking genetic polymorphisms to variation in phenotypes, helping fundamental biological discovery and exploitation in animal breeding and human biomedicine. Here we show results from the pilot phase of PigGTEx by processing 5,457 RNA-sequencing and 1,602 whole-genome sequencing samples passing quality control from pigs. We build a pig genotype imputation panel and associate millions of genetic variants with five types of transcriptomic phenotypes in 34 tissues. We evaluate tissue specificity of regulatory effects and elucidate molecular mechanisms of their action using multi-omics data. Leveraging this resource, we decipher regulatory mechanisms underlying 207 pig complex phenotypes and demonstrate the similarity of pigs to humans in gene expression and the genetic regulation behind complex phenotypes, supporting the importance of pigs as a human biomedical model.
Abstract Background Structural variations (SVs) have significant impacts on complex phenotypes by rearranging large amounts of DNA sequence. Results We present a comprehensive SV catalog based on the whole-genome sequence of 1060 pigs (Sus scrofa) representing 101 breeds, covering 9.6% of the pig genome. This catalog includes 42,487 deletions, 37,913 mobile element insertions, 3308 duplications, 1664 inversions, and 45,184 break ends. Estimates of breed ancestry and hybridization using genotyped SVs align well with those from single nucleotide polymorphisms. Geographically stratified deletions are observed, along with known duplications of the KIT gene, responsible for white coat color in European pigs. Additionally, we identify a recent SINE element insertion in MYO5A transcripts of European pigs, potentially influencing alternative splicing patterns and coat color alterations. Furthermore, a Yorkshire-specific copy number gain within ABCG2 is found, impacting chromatin interactions and gene expression across multiple tissues over a stretch of genomic region of ~200 kb. Preliminary investigations into SV’s impact on gene expression and traits using the Pig Genotype-Tissue Expression (PigGTEx) data reveal SV associations with regulatory variants and gene-trait pairs. For instance, a 51-bp deletion is linked to the lead eQTL of the lipid metabolism regulating gene FADS3, whose expression in embryo may affect loin muscle area, as revealed by our transcriptome-wide association studies. Conclusions This SV catalog serves as a valuable resource for studying diversity, evolutionary history, and functional shaping of the pig genome by processes like domestication, trait-based breeding, and adaptive evolution.
Repressor elements significantly influence economically relevant phenotypes in pigs; however, their precise roles and characteristics are inadequately understood. In the present study, we employed H3K27me3 profiling, assay for transposase-accessible chromatin with highthroughput sequencing (ATAC-seq), and RNA sequencing (RNA-seq) data across six tissues derived from three embryonic layers to identify and map 2 034 super repressor elements (SREs) and 22 223 typical repressor elements (TREs) in the pig genome. Notably, many repressor elements were conserved across mesodermal and ectodermal tissues. SREs exhibited tight regulation of their target genes, affecting a limited number of genes within a specific genomic region with pronounced effects, while TREs exerted broader but weaker regulation over a wider range of target genes. Furthermore, in neuronal tissues, genes regulated by repressor elements started to be repressed during the differentiation of stem cells into progenitor cells. Notably, analysis showed that many repressor elements exhibited cooperative and additive effects on the modulation of KLF4 expression. This research provides the first comprehensive map of pig repressor elements, serving as an essential reference for future studies on repressor elements.
Social isolation (SI) exerts diverse adverse effects on brain structure and function in humans. To gain an insight into the mechanisms underlying these effects, we conducted a systematic analysis of multiple brain regions from socially isolated and group-housed dogs, whose brain and behavior are similar to humans. Our transcriptomic analysis revealed reduced expression of myelin-related genes specifically in the white matter of prefrontal cortex (PFC) after SI during the juvenile stage. Despite these gene expression changes, myelin fiber organization in PFC remained unchanged. Surprisingly, we observed more mature oligodendrocytes and thicker myelin bundles in the somatosensory parietal cortex in socially isolated dogs, which may be linked to an increased expression of ADORA2A, a gene known to promote oligodendrocyte maturation. Additionally, we found a reduced expression of blood-brain barrier (BBB) structural components Aquaporin-4, Occludin, and Claudin1 in both PFC and parietal cortices, indicating BBB disruption after SI. In agreement with BBB disruption, myelin-related sphingolipids were increased in cerebrospinal fluid in the socially isolated group. These unexpected findings show that SI induces distinct alterations in oligodendrocyte development and shared disruption in BBB integrity in different cortices, demonstrating the value of dogs as a complementary animal model to uncover molecular mechanisms underlying SI-induced brain dysfunction.
Understanding the molecular and cellular mechanisms that underlie complex traits in pigs is crucial for enhancing their genetic improvement program and unleashing their substantial potentials in human biomedicine research. Here, we conducted a meta-GWAS analysis for 232 complex traits with 28.3 million imputed whole-genome sequence variants in 70,328 individuals from 14 pig breeds. We identified a total of 6,878 genomic regions associated with 139 complex traits. By integrating with the Pig Genotype-Tissue Expression (PigGTEx) resource, we systemically explored the biological context and regulatory circuits through which these trait-associated variants act and finally prioritized 16,664 variant-gene-tissue-trait circuits. For instance, rs344053754 regulates the expression of UGT2B31 in the liver by affecting the activity of regulatory elements and ultimately influences litter weight at weaning. Furthermore, we investigated the conservation of genetic and regulatory mechanisms underlying 136 human traits and 232 pig traits. Overall, our multi-breed meta-GWAS in pigs provides invaluable resources and novel insights for understanding the regulatory and evolutionary mechanisms of complex traits in both pigs and humans.
BACKGROUND:The pig is an economically important livestock species and is a widely applied large animal model in medical research. Enhancers are critical regulatory elements that have fundamental functions in evolution, development and disease. Genome-wide quantification of functional enhancers in the pig is needed.RESULTS:We performed self-transcribing active regulatory region sequencing (STARR-seq) in the porcine kidney epithelial PK15 and testicular ST cell lines, and reliably identified 2576 functional enhancers. Most of these enhancers were located in repetitive sequences and were enriched within silent and lowly expressed genes. Enhancers poorly overlapped with chromatin accessibility regions and were highly enriched in chromatin with the repressive histone modification H3K9me3, which is different from predicted pig enhancers detected using ChIP-seq for H3K27ac or/and H3K4me1 modified histones. This suggests that most pig enhancers identified with STARR-seq are endogenously repressed at the chromatin level and may function during cell type-specific development or at specific developmental stages. Additionally, the PPP3CA gene is associated with the loin muscle area trait and the QKI gene is associated with alkaline phosphatase activity that may be regulated by distal functional enhancers.CONCLUSIONS:In summary, we generated the first functional enhancer map in PK15 and ST cells for the pig genome and highlight its potential roles in pig breeding.
BackgroundIntramuscular fat (IMF) is associated with meat quality and insulin resistance in animals. Research on genetic mechanism of IMF decomposition has positive meaning to pork quality and diseases such as obesity and type 2 diabetes treatment. In this study, an IMF trait segregation population was used to perform RNA sequencing and to analyze the joint or independent effects of genes and long intergenic non-coding RNAs (lincRNAs) on IMF.ResultsA total of 26 genes including six lincRNA genes show significantly different expression between high- and low-IMF pigs. Interesting, one lincRNA gene, named IMF related lincRNA (IRLnc) not only has a 292-bp conserved region in 100 vertebrates but also has conserved up and down stream genes (<10 kb) in pig and humans. Real-time quantitative polymerase chain reaction (RT-qPCR) validation study indicated that nuclear receptor subfamily 4 group A member 3 (NR4A3) which located at the downstream of IRLnc has similar expression pattern with IRLnc. RNAi-mediated loss of function screens identified that IRLnc silencing could inhibit both of the RNA and protein expression of NR4A3. And the in-situ hybridization co-expression experiment indicates that IRLnc may directly binding to NR4A3. As the NR4A3 could regulate the catecholamine catabolism, which could affect insulin sensitivity, we inferred that IRLnc influence IMF decomposition by regulating the expression of NR4A3.ConclusionsIn conclusion, a novel functional noncoding variation named IRLnc has been found contribute to IMF by regulating the expression of NR4A3. These findings suggest novel mechanistic approach for treatment of insulin resistance in human beings and meat quality improvement in animal.
Understanding the zoonotic origin and evolution history of SARS-CoV-2 will provide critical insights for alerting and preventing future outbreaks. A significant gap remains for the possible role of pangolins as a reservoir of SARS-CoV-2 related coronaviruses (SC2r-CoVs). Here, we screened SC2r-CoVs in 172 samples from 163 pangolin individuals of four species, and detected positive signals in muscles of four Manis javanica and, for the first time, one M. pentadactyla. Phylogeographic analysis of pangolin mitochondrial DNA traced their origins from Southeast Asia. Using in-solution hybridization capture sequencing, we assembled a partial pangolin SC2r-CoV (pangolin-CoV) genome sequence of 22895 bp (MP20) from the M. pentadactyla sample. Phylogenetic analyses revealed MP20 was very closely related to pangolin-CoVs that were identified in M. javanica seized by Guangxi Customs. A genetic contribution of bat coronavirus to pangolin-CoVs via recombination was indicated. Our analysis revealed that the genetic diversity of pangolin-CoVs is substantially higher than previously anticipated. Given the potential infectivity of pangolin-CoVs, the high genetic diversity of pangolin-CoVs alerts the ecological risk of zoonotic evolution and transmission of pathogenic SC2r-CoVs.
Genomic imprinting often results in parent-of-origin specific differential expression of maternally and paternally inherited alleles and plays an essential role in mammalian development and growth. Mammalian genomic imprinting has primarily been studied in mice and humans, with only limited information available for pigs. To systematically characterize this phenomenon and evaluate imprinting status between different species, we investigated imprinted genes on a genome-wide scale in pig brain tissues. Specifically, we performed bioinformatics analysis of high-throughput sequencing results from parental genomes and offspring transcriptomes of hybrid crosses between Duroc and Diannan small-ear pigs. We identified 11 paternally and five maternally expressed imprinted genes in pigs with highly stringent selection criteria. Additionally, we found that the KCNQ1 and IGF2R genes, which are related to development, displayed a different imprinting status in pigs compared with that in mice and humans. This comprehensive research should help improve our knowledge on genomic imprinting in pigs and highlight the potential use of imprinted genes in the pig breeding field.
Understanding the mutational and evolutionary dynamics of SARS-CoV-2 is essential for treating COVID-19 and the development of a vaccine. Here, we analyzed publicly available 15,818 assembled SARS-CoV-2 genome sequences, along with 2,350 raw sequence datasets sampled worldwide. We investigated the distribution of inter-host single nucleotide polymorphisms (inter-host SNPs) and intra-host single nucleotide variations (iSNVs). Mutations have been observed at 35.6% (10,649/29,903) of the bases in the genome. The substitution rate in some protein coding regions is higher than the average in SARS-CoV-2 viruses, and the high substitution rate in some regions might be driven to escape immune recognition by diversifying selection. Both recurrent mutations and human-to-human transmission are mechanisms that generate fitness advantageous mutations. Furthermore, the frequency of three mutations (S protein, F400L; ORF3a protein, T164I; and ORF1a protein, Q6383H) has gradual increased over time on lineages, which provides new clues for the early detection of fitness advantageous mutations. Our study provides theoretical support for vaccine development and the optimization of treatment for COVID-19. We call researchers to submit raw sequence data to public databases.
RNA editing is one of the most common RNA level modifications that potentially generate amino acid changes similar to those resulting from genomic nonsynonymous mutations. However, unlike DNA level allele-specific modifications such as DNA methylation, it is currently unknown whether RNA editing displays allele-specificity across tissues and species. Here, we analyzed allele-specific RNA editing in human tissues and from brain tissues of heterozygous mice generated by crosses between divergent mouse strains and found a high proportion of overlap of allele-specific RNA editing sites between different samples. We identified three allele-specific RNA editing sites cause amino acid changes in coding regions of human and mouse genes, whereas their associated SNPs yielded synonymous differences. In vitro cellular experiments confirmed that sequences differing at a synonymous SNP can have differences in a linked allele-specific RNA editing site with nonsynonymous implications. Further, we demonstrate that allele-specific RNA editing is influenced by differences in local RNA secondary structure generated by SNPs. Our study provides new insights towards a better comprehension of the molecular mechanism that link SNPs with human diseases and traits.
Long intergenic noncoding RNAs (lincRNAs) play a crucial role in many biological processes. The rat is an important model organism in biomedical research. Recent studies have detected rat lincRNA genes from several samples. However, identification of rat lincRNAs using large-scale RNA-seq datasets remains unreported. Herein, using more than 100 billion RNA-seq reads from 59 publications together with RefSeq and UniGene annotated RNAs, we report 39,154 lincRNA transcripts encoded by 19,162 lincRNA genes in the rat. We reveal sequence and expression similarities in lincRNAs of rat, mouse and human. DNA methylation level of lincRNAs is higher than that of protein-coding genes across the transcription start sites (TSSs). And, three lincRNA genes overlap with differential methylation regions (DMRs) which associate with spontaneously hypertensive disease. In addition, there are similar binding trends for three transcription factors (HNF4A, CEBPA and FOXA1) between lincRNA genes and protein-coding genes, indicating that they harbour similar transcription regulatory mechanisms. To date, this is the most comprehensive assessment of lincRNAs in the rat genome. We provide valuable data that will advance lincRNA research using rat as a model.
Domestic dogs have an ancient origin and a long history in Africa. Nevertheless, the timing and sources of their introduction into Africa remain enigmatic. Herein, we analyse variation in mitochondrial DNA (mtDNA) D-loop sequences from 345 Nigerian and 37 Kenyan village dogs plus 1530 published sequences of dogs from other parts of Africa, Europe and West Asia. All Kenyan dogs can be assigned to one of three haplogroups (matrilines; clades): A, B, and C, while Nigerian dogs can be assigned to one of four haplogroups A, B, C, and D. None of the African dogs exhibits a matrilineal contribution from the African wolf (Canis lupus lupaster). The genetic signal of a recent demographic expansion is detected in Nigerian dogs from West Africa. The analyses of mitochondrial genomes reveal a maternal genetic link between modern West African and North European dogs indicated by sub-haplogroup D1 (but not the entire haplogroup D) coalescing around 12,000 years ago. Incorporating molecular anthropological evidence, we propose that sub-haplogroup D1 in West African dogs could be traced back to the late-glacial dispersals, potentially associated with human hunter-gatherer migration from southwestern Europe.
Abstract Pigs are excellent large-animal models for medical research and a promising organ donor source for transplant patients. Next-generation sequencing technology has yielded a dramatic increase in the volume of genomic data for pigs. However, the limited amount of variation data provided by dbSNP, and non-congruent criteria used for calling variation, present considerable hindrances to the utility of this data. We used a uniform pipeline, based on GATK, to identify non-redundant, high-quality, whole-genome SNPs from 280 pigs and 6 outgroup species. A total of 64.6 million SNPs were identified in 280 pigs and 36.8 million in the outgroups. We then used LUMPY to identify a total of 7 236 813 structural variations (SVs) in 211 pigs. Positively selected loci were identified through five statistical tests of different evolutionary attributes of the SNPs. Combining the non-redundant variations and the evolutionary selective scores, we built the first pig-specific variation database, PigVar (http://www.ibiomedical.net/pigvar/), which is a web-based open-access resource. PigVar collects parameters of the variations including summary lists of the locations of the variations within protein-coding and long intergenic non-coding RNA (lincRNA) genes, whether the SNPs are synonymous or non-synonymous, their ancestral and derived states, geographic sampling locations, as well as breed information. The PigVar database will be kept operational and updated to facilitate medical research using the pig as model and agricultural research including pig breeding. Database URL: http://www.ibiomedical.net/pigvar/
BACKGROUND:Long noncoding RNAs (lncRNAs) have attracted significant attention in recent years due to their important roles in many biological processes. Domestic animals constitute a unique resource for understanding the genetic basis of phenotypic variation and are ideal models relevant to diverse areas of biomedical research. With improving sequencing technologies, numerous domestic-animal lncRNAs are now available. Thus, there is an immediate need for a database resource that can assist researchers to store, organize, analyze and visualize domestic-animal lncRNAs.RESULTS:The domestic-animal lncRNA database, named ALDB, is the first comprehensive database with a focus on the domestic-animal lncRNAs. It currently archives 12,103 pig intergenic lncRNAs (lincRNAs), 8,923 chicken lincRNAs and 8,250 cow lincRNAs. In addition to the annotations of lincRNAs, it offers related data that is not available yet in existing lncRNA databases (lncRNAdb and NONCODE), such as genome-wide expression profiles and animal quantitative trait loci (QTLs) of domestic animals. Moreover, a collection of interfaces and applications, such as the Basic Local Alignment Search Tool (BLAST), the Generic Genome Browser (GBrowse) and flexible search functionalities, are available to help users effectively explore, analyze and download data related to domestic-animal lncRNAs.CONCLUSIONS:ALDB enables the exploration and comparative analysis of lncRNAs in domestic animals. A user-friendly web interface, integrated information and tools make it valuable to researchers in their studies. ALDB is freely available from http://res.xaut.edu.cn/aldb/index.jsp.
Long intergenic noncoding RNAs (lincRNAs) are one of the major unexplored components of genomes. Here we re-analyzed a published methylated DNA immunoprecipitation sequencing (MeDIP-seq) dataset to characterize the DNA methylation pattern of pig lincRNA genes in adipose and muscle tissues. Our study showed that the methylation level of lincRNA genes was higher than that of mRNA genes, with similar trends observed in comparisons of the promoter, exon or intron regions. Different methylation pattern were observed across the transcription start sites (TSS) of lincRNA and protein-coding genes. Furthermore, an overlap was observed between many lincRNA genes and differentially methylated regions (DMRs) identified among different breeds of pigs, which show different fat contents, sexes and anatomic locations of tissues. We identify a lincRNA gene, linc-sscg3623, that displayed differential methylation levels in backfat between Min and Large White pigs at 60 and 120 days of age. We found that a demethylation process occurred between days 150 and 180 in the Min and Large White pigs, which was followed by remethylation between days 180 and 210. These results contribute to our understanding of the domestication of domestic animals and identify lincRNA genes involved in adipogenesis and muscle development.
Thousands of long intergenic noncoding RNAs (lincRNAs) have been identified in the human and mouse genomes, some of which play important roles in fundamental biological processes. The pig is an important domesticated animal, however, pig lincRNAs remain poorly characterized and it is unknown if they were involved in the domestication of the pig. Here, we used available RNA-seq resources derived from 93 samples and expressed sequence tag data sets, and identified 6,621 lincRNA transcripts from 4,515 gene loci. Among the identified lincRNAs, some lincRNA genes exhibit synteny and sequence conservation, including linc-sscg2561, whose gene neighbor Dnmt3a is associated with emotional behaviors. Both linc-sscg2561 and Dnmt3a show differential expression in the frontal cortex between domesticated pigs and wild boars, suggesting a possible role in pig domestication. This study provides the first comprehensive genome-wide analysis of pig lincRNAs.
Background High-throughput transcriptome sequencing (RNA-seq) technology promises to discover novel protein-coding and non-coding transcripts, particularly the identification of long non-coding RNAs (lncRNAs) from de novo sequencing data. This requires tools that are not restricted by prior gene annotations, genomic sequences and high-quality sequencing. Results We present an alignment-free tool called PLEK ( p redictor of l ong non-coding RNAs and m e ssenger RNAs based on an improved k -mer scheme), which uses a computational pipeline based on an improved k -mer scheme and a support vector machine (SVM) algorithm to distinguish lncRNAs from messenger RNAs (mRNAs), in the absence of genomic sequences or annotations. The performance of PLEK was evaluated on well-annotated mRNA and lncRNA transcripts. 10-fold cross-validation tests on human RefSeq mRNAs and GENCODE lncRNAs indicated that our tool could achieve accuracy of up to 95.6%. We demonstrated the utility of PLEK on transcripts from other vertebrates using the model built from human datasets. PLEK attained >90% accuracy on most of these datasets. PLEK also performed well using a simulated dataset and two real de novo assembled transcriptome datasets (sequenced by PacBio and 454 platforms) with relatively high indel sequencing errors. In addition, PLEK is approximately eightfold faster than a newly developed alignment-free tool, named Coding-Non-Coding Index (CNCI), and 244 times faster than the most popular alignment-based tool, Coding Potential Calculator (CPC), in a single-threading running manner. Conclusions PLEK is an efficient alignment-free computational tool to distinguish lncRNAs from mRNAs in RNA-seq transcriptomes of species lacking reference genomes. PLEK is especially suitable for PacBio or 454 sequencing data and large-scale transcriptome data. Its open-source software can be freely downloaded from https://sourceforge.net/projects/plek/files/ .