Transposable elements (TEs) are DNA sequences able to create copies of themselves within the genome. TEs have been shown to act as cis-regulatory elements and be co-opted in the human genome. Thus, their impact might come from their relationship with the epigenome. However, a systematic analysis that relates TEs with chromatin histone marks across human cell types remains lacking. Here we leverage a dataset from the International Human Epigenome Consortium featuring 4867 uniformly processed ChIP-seq experiments for 6 histone marks across 47 cell types and show that TEs have drastically different enrichments levels across histone marks. We find that TEs are generally depleted but enriched in select contexts such as L1s in H3K9me3 histone mark. Notably, we identify 456 cell type-histone-TE triplets with strong cell-type specific enrichments and show that many of these triplets are associated with relevant biological processes.
Variation in short tandem repeats (STRs) is implicated in Mendelian disease and complex traits but can be difficult to resolve with short-read genome sequencing. We present STRkit, a software package for genotyping STRs using long-read sequencing (LRS) that uses proximate single-nucleotide variants to improve genotyping accuracy without a priori haplotype information. We show that STRkit has unique strengths versus other methods: It can use data from both major LRS technologies (Pacific Biosciences HiFi [PacBio] and Oxford Nanopore Technologies [ONT]) to output both allele- and read-level copy number and sequence; it performs best in benchmarking with F1 scores of 0.9631 and 0.9544 with PacBio and ONT data, respectively; it achieves higher rates of Mendelian consistency than other genotyping tools; and it is open source software. STRkit's features open up new possibilities for association testing, assessing patterns of STR inheritance and better understanding the functional effects of these notable repeat elements.
Using various biochemical assays that identify transcription factor (TF) binding and histone modifications, cis-regulatory elements (CREs) can be annotated in a genome-wide manner. However, these assays are descriptive and require functional validation. To the best of our knowledge, no technology can simultaneously analyze the regulatory function and epigenomic modifications of a specific sequence. Here, we develop an enrichment followed by epigenomic profiling massively parallel reporter assay (e2MPRA). This technique uses lentivirus to enrich for the integration of specific CREs into the genome and applies MPRA, Cut&Tag or ATAC-seq on them enabling simultaneous, high-throughput analysis of regulatory activity, protein binding, and epigenetic modification. We demonstrate that e2MPRA can dissect the epigenetic functions of TF motifs arranged within synthetic enhancers and evaluate the effects of sequence perturbation on epigenetic states. In summary, e2MPRA advances our understanding of the regulatory code, its effect on the epigenome and how its alteration leads to phenotypic effects.
Structural variants (SVs) are omnipresent in human DNA, yet their genotype and methylation statuses are rarely characterized due to previous limitations in genome assembly and detection of modified nucleotides. Also, the extent to which SVs act as methylation quantitative trait loci (SV-mQTLs) is largely unknown. Here, we generated a pangenome graph summarizing SVs in 782 de novo assemblies obtained from Genomic Answers for Kids, capturing 14.6 million CpG dinucleotides that are absent from the CHM13v2 reference (SV-CpGs), thus expanding their number by 43.6%. Using 435 methylomes, we genotyped 4.06 million SV-CpGs, of which 3.93 million (96.8%) are methylated at least once. Nonrepeat sequences contribute 1.59 × 106 novel SV-CpGs, followed by centromeric satellites (6.57 × 105), simple repeats (5.40 × 105), Alu elements (5.07 × 105), satellites (2.17 × 105), LINE-1s (1.83 × 105), and SVA (SINE-VNTR-Alu) elements (1.50 × 105). Centromeric satellites, simple repeats, and SVAs are overrepresented in SV-CpGs versus reference CpGs. Similarly, methylation levels in SV-CpGs are more variable than in reference CpGs. To explore if SVs are potentially causal for functional variation, we measured SV-mQTLs. This revealed over 230,464 methylation bins where the methylation is associated with common SVs within 100 kbp. Finally, we identified 65,659 methylation bins (28.5%) where the leading QTL variant is an SV. In conclusion, we demonstrate that graph pangenomes provide full SV structures, the associated methylation variation, and reveal tens of thousands of SV-mQTLs, underscoring the importance of assembly based analyses of human traits.
Understanding the mechanisms for human ovarian folliculogenesis is fundamental to reproductive biology and medicine. Here, we investigated transcriptomic dynamics in individual oocytes and their associated granulosa cells (GCs) during folliculogenesis in mice, monkeys and humans. Unlike mouse oocytes, which exhibited stage-specific, stepwise transcriptomic maturation, monkey and human oocytes showed minimal transcriptomic changes until the secondary follicle stage and could be broadly categorized as either immature or mature. In all three species, most highly variable genes during oocyte growth displayed monotonic up-or downregulation, with limited overlap in highly variable genes across species. GCs exhibited similarly species-specific transcriptomic trajectories. Correspondingly, intercellular communication pathways, including ligand-receptor signaling, gap junctions, and metabolic coupling between oocytes and GCs, demonstrated substantial species-specific differences. Nonetheless, X-chromosome dosage compensation and repression of evolutionarily young transposons were conserved across species. We established in vitro culture systems supporting preantral to antral follicle development in monkeys and humans, revealing relatively normal oocyte transcriptome maturation but aberrant GC profiles. By delineating interspecies differences in folliculogenesis, this study provides a framework for understanding human ovarian development and advancing its in vitro reconstruction.
Approximately half of the human genome is derived from transposable elements (TEs) and several studies support the involvement of TEs in genome regulation in development, immunity and disease. We previously leveraged 4614 ChIP-seq samples from the International Human Epigenome Consortium (IHEC) EpiATLAS dataset and did a comprehensive analysis of the relationship between TEs and 6 histone marks across 57 human cell types. However, with over 6 million measurements of TE / histone mark / cell type enrichment, it was challenging to navigate the results and it was not possible to integrate them with user data. To address this, we developed a web tool, TEExplorer, which makes available TE overlaps and enrichments in an accessible and intuitive manner. The tool presents an interactive view of TE families and subfamilies, with their overlap and enrichments across histone marks and cell types. Finally, the tool allows users to upload their own ChIP-seq BED file to obtain the TE overlap and enrichment relative to random controls and compare their data with the EpiATLAS dataset. With TEExplorer, researchers with an interest in a particular TE family or subfamily, histone mark, or cell type, or those bringing their own ChIP-seq dataset, can dynamically explore and contrast hundreds of associations found within the large EpiATLAS dataset.
The sequence of the human genome provides a foundation for understanding cellular processes in health and disease1. The organisation of this primary genetic information into cell-specific structure and function is critical to understanding the cell type-specific interpretation and execution of the genome. Epigenetic processes are essential for packaging and higher-level functional organisation of the genome, and changes therein are increasingly recognised as contributors to human disease. Building on primary data generated by multinational consortia, the International Human Epigenome Consortium2 (IHEC) has uniformly processed a collection of more than 2000 comprehensive human reference epigenomes, collectively referred to as EpiATLAS. This effort involved the development of standardised molecular and bioinformatics protocols, metadata models, and analytical tools to manage, integrate, display, and share vast amounts of epigenomic data. This includes the creation of a publicly available Epigenome Reference Registry, which provides a system for accessing protected human subject datasets and facilitates open searching of de-identified samples and experimental data. The integrated EpiATLAS ecosystem and its comprehensive human reference epigenome maps provide an unprecedented resource for the biosciences, expanding the annotated epigenomic landscape while uncovering previously unappreciated relationships among regulatory layers and revealing how epigenetic inputs underpin fundamental cellular functions and disease associations.
Advances in sequencing technologies continue to improve the resolution and completeness with which human genetic variation can be characterized. Short-read sequencing remains widely used due to its high base accuracy, throughput, and cost efficiency; however, its limited ability to resolve repetitive and structurally complex regions has accelerated adoption of long-read sequencing platforms, including those from Pacific Biosciences (PacBio) and Oxford Nanopore Technologies (ONT). We systematically compared sequencing technologies and variant calling pipelines for small variants and structural variants across diverse genomic contexts and sequencing depths. Short-read sequencing combined with DRAGEN achieved high accuracy for single-nucleotide variants (SNVs) and indels in well-mapped and moderately complex regions but showed reduced sensitivity and completeness for structural variant detection. In contrast, long-read sequencing platforms demonstrated clear advantages in detecting structural variants and resolving small variants in difficult genomic regions, although challenges remain in specific indel-prone sequence contexts. Among long-read pipelines, PacBio Revio with DeepVariant achieved the highest SNV and indel accuracy genome-wide, while ONT R10 with DeepVariant performed particularly well in clinically relevant loci. Structural variant detection was dominated by long-read optimized callers, with SVIM and Sawfish performing best for PacBio, and Sniffles2 and CuteSV2 for ONT, consistently outperforming short-read-based methods across variant classes and sizes. Coverage analyses indicated that long-read sequencing reached accuracy saturation between 20 × and 45 × , whereas short-read sequencing required more than 60 × coverage to approach maximal genome completeness. These results provide practical guidance for platform and pipeline selection. Long-read sequencing enables more comprehensive detection and resolution of structural variants and variation in complex genomic regions, while short-read sequencing remains a cost-effective and scalable solution for high-throughput genotyping and clinically focused applications.
A pangenome reference overcomes the inherent limitation of any individual reference genome by integrating the variation present in a population. We present the Human Pangenome Reference Consortium's (HPRC) Release 2 (HPRC2), an openly available, second phase pangenome that is an approximately fivefold expansion in genome number over HPRC Release 1 (HPRC1) and measurable improvement in genome completeness, contiguity, and accuracy. Selecting samples with a principled algorithm prioritising common variant coverage, HPRC2 contributes 460 haplotypes that together capture over 99% of common variation observed in the All of Us Research Program v8 cohort. Combining high-coverage long and ultra-long reads with modern assemblers and polishers, we produce thousands of telomere-to-telomere (T2T) chromosomes, and relative to HPRC1 halve the number of structurally unreliable regions as well as individual base errors per haplotype. We complement the assemblies with whole genome multiple alignments and gene annotations, and derive formal pangenome coordinate systems for addressing off-reference variation, demonstrating that individual human genomes contain more than one hundred thousand variants not succinctly described with respect to existing reference genomes. We also present the first matched long-read backed pantranscriptome and panepigenome at this scale, provide continuous local-ancestry estimates spanning every genome, and outline a host of new tools and applications that leverage the pangenome resource for improved genomics analysis.
Advances in sequencing technology have enabled population-level Whole Genome Sequencing (WGS) efforts to be undertaken in many countries. Often, this requires collaboration across a distributed network of sequencing centres to allow efficient use of existing resources. Previously we tested the robustness of short-read sequencing technology and analysis pipelines across three established sequencing centres located in Montreal, Toronto, and Vancouver, constituting CGEn, Canada's national platform for genome sequencing and analysis (www.cgen.ca). In this work, we extend the study to cover Oxford Nanopore Technologies (ONT) long read-based WGS technology which is increasingly being used for large-scale genomics studies. Thus, we performed ONT WGS of the HG002 cell line, a well-characterized standard obtained directly from the Coriell Institute, aiming for a minimum of 30× coverage using one R10.4 PromethION flowcell at each centre. The sequencing datasets were analyzed using commonly developed pipelines for SNVs, Indels, SV, and CpG methylation detection and then compared to the relevant publicly available GIAB benchmark datasets. As a result, we tested the robustness of the laboratory protocols as well as the effectiveness of the analytical pipelines for simultaneous analysis of genomic variation and CpG methylation. Key findings include: SNV detection with higher F1-scores in RefSeq Coding regions for the ONT datasets (99.1%-99.5%) compared to Illumina NovaSeqX data (96.5%); additionally, there was high correlation of CpG methylation across all the sequencing centres (R = 0.97), as well as with publicly available WGBS (R = 0.88) and EM-Seq (R = 0.93) data from the EpiQC study.
Sex-biased gene regulation is the basis of sexual dimorphism in phenotypes and has been studied across different cell types and different developmental stages. However, sex-biased expression of transposable elements (TEs), which represent nearly half of the mammalian genome and have the potential of influencing genome integrity and regulation, remains underexplored. We report a survey of gene, lncRNA, and TE expression in four organs from mice with different combinations of gonadal and genetic sex. The data show remarkable variability among organs with respect to the impact of gonadal sex on transcription with the strongest effects observed in the liver. In contrast, the X-chromosome dosage alone had a modest influence on sex-biased transcription across organs, albeit interaction between X-dosage and gonadal sex cannot be ruled out. The presence of the Y-chromosome influences TE, but not gene or lncRNA, expression in the liver. Notably, 90
Cell fate and identity require timely activation of lineage-specific and concomitant repression of alternate-lineage genes. How this process is epigenetically encoded remains largely unknown. In skeletal muscle stem cells, the myogenic regulatory factors are well-established drivers of muscle gene activation but less is known about how non-muscle gene repression is achieved. Here, we show that the master epigenetic regulator, Repressor Element 1-Silencing Transcription factor (REST), also known as Neuron-Restrictive Silencer Factor (NRSF), is a key regulator of this process. We show that many non-lineage genes retain permissive chromatin state but are actively repressed by REST. Loss of functional REST in muscle stem cells and progenitors disrupts muscle specific epigenetic and transcriptional signatures, impairs differentiation, and triggers apoptosis in progenitor cells, leading to depletion of the stem cell pool. Consequently, REST-deficient skeletal muscle exhibits impaired regeneration and reduced myofiber growth postnatally. Collectively, our data suggests that REST plays a key role in safeguarding muscle stem cell identity by repressing multiple non-muscle lineage and developmentally regulated genes in adult mice.
Telomeres are essential for maintaining genomic integrity and are associated with cellular aging and disease, yet the factors influencing their inheritance across generations remain poorly understood. Leveraging PacBio HiFi long-read sequencing and 75 parent-offspring trios (n = 225) from the Genomic Answers for Kids program, we analyzed individual telomeres across chromosomes and their inheritance. Telomere length (TL) varied between chromosome arms in a way that was consistent in parents and offsprings, with average values ranging from 5000 to 8000 base pairs. Maternal and paternal TL together were a strong predictor of child TL (R[2][1] = 0.59). Notably, using telomeric variant repeats, we developed a tool that enabled allelic tracing for 53.3% of maternally and 49.9% of paternally inherited telomeres. In the child, paternally transmitted alleles were significantly longer than age-matched maternal ones (Δmean = 409 bp, p = 2.6e-05), particularly when from older parents (Δmean = 698 bp, p = 8.9e-05) and at chromosome arms with shorter average TL (Δmean = 752 bp, p = 1.6e-06). These findings reveal parent-of-origin effects and heritable influences on TL, providing novel insights into telomere dynamics and their potential implications in age-related disease susceptibility. ### Competing Interest Statement The authors have declared no competing interest. Natural Sciences and Engineering Research Council Canada Research Chairs, https://ror.org/0517h6h17 Children’s Mercy Research Institute, https://ror.org/0169kb131 [1]: #ref-2
Variation in short tandem repeats (STRs) is implicated in Mendelian disease and complex traits, but can be difficult to resolve with short-read genome sequencing. We present STRkit , a software package for genotyping STRs using long read sequencing (LRS) that uses nearby single-nucleotide variants to improve genotyping accuracy without a priori haplotype information. We show that STRkit has unique strengths versus other methods: it can use data from both major LRS technologies (Pacific Biosciences HiFi [PB] and Oxford Nanopore [ONT]) to output both allele and read-level copy number and sequence, performs best in benchmarking with F1 scores of 0.9633 and 0.9056 with PB and ONT data respectively, achieves a Mendelian inheritance rate of 97.86% with PB data, and is open source software. STRkit 's features open up new possibilities for association testing, assessing patterns of STR inheritance, and better understanding the functional effects of these notable repeat elements. ### Competing Interest Statement The authors have declared no competing interest.
Brain regions drive multiple physiological functions through specific gene expression patterns that adapt to environmental influences, drug treatments and disease conditions. To generate a detailed atlas of the brain transcriptome in the context of diabetes, we carried out RNA sequencing in hypothalamus, hippocampus, brainstem and striatum of the Goto-Kakizaki (GK) rat model of spontaneous type 2 diabetes, which was applied to identify gene transcription adaptation to improved glycemic control following vertical sleeve gastrectomy (VSG) in the GK. Over 19,000 distinct transcripts were detected in the rat brain, including 2794 which were consistently expressed in the four brain regions. Region-specific gene expression was identified in hypothalamus (n = 477), hippocampus (n = 468), brainstem (n = 1173) and striatum (n = 791), resulting in differential regulation of biological processes between regions. Differentially expressed genes between VSG and sham operated rats were only found in the hypothalamus and were predominantly involved in the regulation of endothelium and extracellular matrix. These results provide a detailed atlas of regional gene expression in the diabetic rat brain and suggest that the long term effects of gastrectomy-promoted diabetes remission involve functional changes in the hypothalamus endothelium.
Whole-genome duplication (WGD) events are common across various organisms; however, the retention and evolution of WGD paralogs is not fully understood. Quantitative measure of protein redistribution in response to the deletion of their WGD paralog provides insight into sources of gene retention. Here, we describe PARPAL (PARalog Protein Redistribution using Abundance and Localization in Yeast), a web database that houses results of high-content screening and deep learning neural network analysis of the redistribution of 164 proteins reflecting how their subcellular localization and protein abundance change in response to their paralog deletion in the budding yeast, Saccharomyces cerevisiae. We interrogated a total of 82 paralog pairs in 2 genetic backgrounds for a total of ∼3,500 micrographs of ∼460,000 cells. For example, Skn7-Hms2 exhibited dependent redistribution, and Cue1-Cue4 showed compensatory redistribution response. PARPAL also links to other studies on trigenic interactions, protein-protein interactions and protein abundance. PARPAL is available at https://parpal.c3g-app.sd4h.ca and is a valuable resource for the yeast community interested in understanding the retention and evolution of paralogs and can help researchers to investigate protein dynamics of paralogs in other organisms.
Disease-Free Survival outcomes amongst genomic groups when applied to the 247 VHL mutated ccRCCs from the TCGA dataset
Advances in DNA sequencing have transformed genomics, enabling comprehensive insights into human genetic variation. While short-read sequencing (SRS) remains dominant due to its high accuracy and affordability, its limitations in complex genomic regions have spurred the adoption of long-read sequencing (LRS) platforms, such as those from Pacific Biosciences (PacBio) and Oxford Nanopore Technologies (ONT). Despite these advances, there is still a lack of systematic, large-scale benchmarking of variant calling performance across diverse platforms, variant types, sequencing depths, and genomic contexts. Here, we present a comprehensive benchmark of sequencing technologies and variant calling algorithms, evaluating their performance in detecting single nucleotide polymorphisms (SNPs), insertions/deletions (indels), and structural variants (SVs). We show that while SRS combined with DeepVariant or DRAGEN offers excellent small variant detection in well-mapped regions, LRS technologies significantly outperform SRS in complex regions and SV detection. PacBio achieves high SNP and smaller SV accuracy even at moderate coverage, while ONT excels in detecting large SVs. The SV callers Dysgu and SVIM emerged as top performers across LRS datasets. Our results highlight that no single platform is optimal for all variant types or regions: SRS remains optimal for high-throughput small variant detection in accessible regions, whereas LRS is critical for capturing SVs and resolving difficult-to-map loci. These findings offer practical guidance for selecting sequencing technologies, coverage and variant calling strategies tailored to specific research or clinical goals, contributing to more accurate and cost-effective genomic analyses. Highlights ● State-of-the-art short variant calling algorithms are highly comparable but focus on precision and sensitivity differently ● Short-read technologies outperform long-read technologies for small SNP and InDels, with the exception of difficult-to-map variants. ● To capture the majority of variants, a minimum coverage of 15x for PacBio, 20x for SRS, or 30x for ONT is required. However, optimal coverage depends on zygosity, variant type, and the region of interest. ● Long-read technologies outperform short-read technologies for all validation sets tested for deletions and insertions in all size categories. ### Competing Interest Statement The authors have declared no competing interest. Canada Institute of Health Research (CIHR) project grant, PJT-191707 Genome Canada Genome Technology Platform grant
Wing-Kin Sung合作论文数Department of Computer Science, School of Computing, National University of Singapore11