Centromeres are essential chromosome components yet remain poorly understood due to their highly repetitive sequence architecture. Using fully-phased telomere-to-telomere diploid assemblies from a three-generation pedigree integrated with long-read epigenomes from matched peripheral blood mononuclear cells, induced pluripotent stem cells, and neural progenitor cells, we generate allele-resolved single basepair resolution maps of centromere genetic and epigenetic dynamics across inheritance, reprogramming, and differentiation. We show that centromeric dip regions (CDRs), which define the functional core of centromeres, are positionally stable across generations and cell-fate transitions. In contrast, CDR epigenetic architecture is highly dynamic. Reprogramming markedly attenuates CDR hypomethylation, which is partially restored during differentiation in parallel with global hypomethylation of active alpha-satellite arrays and coordinated changes in nucleosome organization and protein occupancy. Centromeric remodeling is insulated from X-chromosome status, including Xa, Xi, and erosion. Finally, de novo mutations arising during reprogramming are enriched in centromeric regions but depleted within functional centromeric cores.
Long-read sequencing (LRS) and diploid genome assembly have enabled nearly complete structural variant (SV) discovery. Using 293 nearly complete genomes, we characterize the full spectrum of genetic variation and show that while 99% of the variants between any two genomes are single base-pair substitutions, 88% of the euchromatic variant base pairs are SVs, including insertions, deletions, duplications, and inversions. We identify 24 gene-rich regions subject to megabase-scale variation, 2,293 potentially unstable tandem repeats, and 890 novel expression quantitative trait loci associated with SVs in humans. Expanding to 1,218 LRS samples from the 1000 Genomes Project and applying a newly developed cross-platform breakpoint evaluation tool, BoostSV, we construct a nonredundant callset comprising 614,522 SVs. We demonstrate the utility of this population-level SV reference callset by filtering >99% of the common variation from 44 unsolved LRS probands from the Undiagnosed Diseases Network to discover likely disease-causing SVs. Second, we genotype 1,053 high-impact biallelic SVs from the pangenome callset in 232,090 samples from All of Us and discover 105 SVs with significant associations, including 26% where the SV is the lead variant. This publicly available pangenome SV resource will drive new disease associations and further our understanding of the missing heritability of human genetic disease.
When aligning next-generation sequencing (NGS) reads to a reference genome, differences between the true genome of the individual under study and the reference result in a biased interpretation of aligned data through systematic errors known as reference alignment bias (RAB). The degree to which RAB impacts functional readouts has not been thoroughly quantified. Leveraging resources from the Human Pangenome Reference Consortium, here we quantify RAB in functional genomics assays. Our results indicate that, on average, 0.2% of the genome is susceptible to bias in RNA sequencing (RNA-seq) studies, 1% in ATAC-seq, and 3% in WGBS when using the human reference hg38. Our study quantifies the effect of RAB on functional assays and highlights the importance of using an adequately representative reference genome.
Cancer in younger adults is rising globally, with notable birth-cohort effects. This epidemiological shift underscores the urgent need to accelerate the identification of novel causes and underlying biological networks, with the aim of translating these insights into prevention and interception strategies. In this perspective, we revisit the major milestones in the discovery of cancer causes and outline challenges that hinder progress. To address these challenges, we advocate closer integration of epidemiologic and mechanistic studies and propose three interconnected frameworks that extend current epidemiologic approaches: a tissue ecosystem-anchored framework for cancer cause discovery, a biological state-based framework for precision cancer risk assessment, and a dynamic framework to characterize cancer preventability. This roadmap aims to stimulate conceptual, resource, and methodological advances to accelerate cancer etiology research and prevention in the era of rising early-onset cancers.
The Toxicant Exposures and Responses by Genomic and Epigenomic Regulators of Transcription (TaRGET) program is a multiphase program that aims to understand how environmental factors contribute to disease susceptibility using toxicant-exposed mouse models. Here, we present the TaRGET II Data Portal ( https://data.targetepigenomics.org/ ), a comprehensive repository of multi-omics datasets generated from mice exposed to a range of environmental toxicants, including arsenic (As), lead (Pb), bisphenol A (BPA), tributyltin (TBT), di(2-ethylhexyl) phthalate (DEHP), tetrachlorodibenzo-p-dioxin (TCDD), and ambient air pollution (PM2.5). Epigenomic profiling assays capturing alterations in chromatin accessibility, DNA methylation, gene expression, and histone modifications were produced across multiple research centers, subjected to rigorous quality control, and processed using standardized pipelines. The portal currently hosts 3,612 datasets spanning multiple tissues and genomic assays, collected at four developmental and exposure time points: 3, 5, 20, and 40 weeks. The portal offers an efficient way to browse, search, visualize, and download relevant datasets and associated metadata, serving as a key resource for studying the impact of environmental toxicant exposures on disease susceptibility for the broader scientific community.
The basal ganglia are a group of forebrain nuclei critical for motor control and reward processing, and their dysfunction contributes to neurological and neuropsychiatric disorders. Here, we present the first multimodal single-cell epigenomic atlas of the human basal ganglia across major subregions and cell types. We jointly profiled DNA methylation and 3D chromatin conformation in 197,003 nuclei from eight basal ganglia subregions using multi-omic sequencing (snm3C-seq), and integrated these data with existing DNA methylation and chromatin conformation sequencing datasets to build a unified atlas of 261,331 cells spanning 31 subclasses and 59 groups. This atlas reveals extensive cell-type- and region-specific differential methylation, enriched for distinct transcription factor motifs, and validated by MERFISH spatial transcriptomics, which uncovered epigenetic gradients linked to transcriptional output. Compared to neuronal cells, non-neuronal cells exhibit distinct 3D genome organization including smaller chromatin compartments, increased long-range inter-compartment contacts, shorter loops, and stronger CG hypomethylation in A compartments. We further identified genes that display compartment switches, are strongly correlated with compartment scores, and exhibit differential domain boundaries and chromatin looping across basal ganglia cell types. We identified multiple medium spiny neuron subtypes defined by distinct hypomethylated signature genes, with 3D genome embeddings emphasizing dorsal, ventral, and hybrid populations. By integrating chromatin accessibility and histone modification profiles, we reconstructed cell-type-resolved enhancer-promoter links and gene regulatory networks, providing a comprehensive epigenomic framework for interpreting genetic risk loci and regulatory architecture in the human basal ganglia.
Abstract Epigenomic remodeling, such as changes in chromatin state, plays a crucial role in cancer biology by regulating gene expression changes that drive tumor progression, immune evasion, and the evolution of therapeutic resistance. Popular methods to study the epigenome include assay for transposase-accessible chromatin with sequencing (ATAC-seq) and chromatin immunoprecipitation sequencing (ChIP-seq). ATAC-seq only profiles open chromatin, missing critical details about the nature of accessible regions or silenced chromatin. ChIP-seq and its newer relative, cleavage under targets and tagmentation (CUT&Tag), provide more detailed information on chromatin states by targeting histone modifications with specific antibodies. ChIP-seq requires 106 cells to generate meaningful data—too high for precious samples. CUT&Tag offers higher sensitivity with 10-100x lower input and a significantly simplified workflow. When studying epigenetic changes in cancer, it is widely accepted that single-cell resolution is essential, given the complexity and heterogeneity of tumors and their microenvironment. Hence, there is a high demand to convert current epigenetic assays from bulk to single-cell resolution. In CUT&Tag, DNA is tagmented with a protein A/G-Tn5 fusion enzyme, which inserts sequencing adapters in situ for cell-specific labeling—enabling single-cell resolution. Some labs have experimented with single-cell CUT&Tag (scCUT&Tag), but here we present a novel, ready-to-use, validated method for automated, high-throughput scCUT&Tag. To assess drug induced changes in acetylation, we profiled H3K27ac patterns in thousands of single cells from a lung cancer cell line (A549 WT and A549 p53 KO) before and after epigenetic treatment (decitabine + panabinistat vs DMSO control). We observed global cell-type-specific acetylation in response to epigenetic therapy in both A549 WT and A549 p53 KO cells. When comparing scCUT&Tag data with single-cell total RNA-seq data we found that increased acetylation at gene promoters and enhancers correlated with increased gene expression, corroborating the observed changes in histone modification. In summary, the validated scCUT&Tag method provides a high-throughput, automated approach to identify genes regulated in response to epigenetic drug treatment of tumor cells at the single-cell level. Integrating this single-cell expression data provides deeper insight into cellular regulation, particularly in the context of drug responses in tumor cells. Citation Format: Shuwen Chen, Lisa Welter, Gilma Sevilla, Ploy Setthasap, Yana Ryan, Shiyi Yin, Alan Du, Jackson Peterson, Mike Covington, Mohammad Fallahi, Kazuo Tori, Bryan Bell, Bria Graham, Matt J. Meiners, Andrea L. Johnstone, Keith E. Maier, Martis W. Cowles, Bryan J. Venters, Michael Keogh, Xuan Qu, Colin McCornack, Ting Wang, Yue Yun, Andrew Farmer. Profiling histone modifications in single cells to gain insight into the effects of epigenetic drug treatment on tumor cells [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2026; Part 1 (Regular Abstracts); 2026 Apr 17-22; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2026;86(7 Suppl):Abstract nr 3221.
The basal ganglia regulate motor, cognitive, and affective behaviors, and their dysfunction underlies diverse neurological and psychiatric disorders. Comprehensive, accessible multi-omics resources are needed to understand the regulatory mechanisms governing basal ganglia cell types. Here we present an open, interactive web-based platform for exploring single-cell multi-omics datasets from basal ganglia, generated using 10X Multiome, snm3C-seq, and Paired-Tag technologies from the BICAN (NIH BRAIN Initiative Cell Atlas Network) consortium. The platform is available at https://basalganglia.epigenomes.net/ and enables integrated visualization of gene expression, chromatin accessibility, DNA methylation, histone modifications, and chromatin conformation across cell types and human, macaque, marmoset, and mouse species, with direct genome browser support and comparative epigenomic functionality. Representative analyses demonstrate cell-type-specific regulatory landscapes, conserved and species-specific regulatory elements, and links between epigenomic regulation and transcription. This resource provides a scalable, community-oriented foundation for advancing basal ganglia biology and interpreting regulatory mechanisms relevant to brain function and disease.
Somatic variant detection is technically challenging due to low variant allele fractions, the confounding presence of germline variation, and reference bias. Linear references such as GRCh38 miss sample-specific variation, causing misalignments and incorrect variant calls. Although telomere-to-telomere donor-specific assemblies (DSAs) accurately represent individual genomes, their application is limited by cost and technical barriers. Alternatively, the graph-based human pangenome provides a scalable framework to improve read alignment and perform genome inference. Here, we benchmarked somatic variant detection using GRCh38, graph-based pangenomes, and pangenome-inferred DSAs with a HapMap mixture dataset and the COLO829 melanoma cell line. Pangenome-guided alignment improves read mapping and somatic variant calling accuracy. Furthermore, personalized pangenomes partially reconstruct donor-specific genomic content, improving accuracy, reducing germline contamination, and enabling detection of events in loci absent or poorly represented in GRCh38. These findings demonstrate that graph-based and personalized pangenomes are effective strategies for enhancing somatic variant detection compared with GRCh38.
Cytosine methylation, a crucial epigenetic modification, plays a vital role in genomic regulation. Leveraging the advancements in long-read sequencing, we investigate the methylation patterns of polymorphic transposable element (TE) insertions of human lymphoblastoid cell lines (LCLs). We validate the high concordance between long-read methylation calls and the conventional whole-genome bisulfite sequencing (WGBS) method. We then aim to establish general rules of TE methylation with our data by addressing three key questions: (1) what is the methylation profile of each insertion; (2) do newly inserted TEs adopt the methylation pattern of their genomic context; and (3) do new TE insertions affect the methylation of their flanking regions. Although most non-TE insertions exhibit DNA methylation patterns consistent with their genomic context, TE insertions are generally highly methylated, exhibiting distinct, class-specific patterns with some variation within TE bodies. A small percentage of Alu insertions are hypomethylated, particularly those inserted within hypomethylated CpG islands. We also reveal that majority of TEs exhibited minimal impact on nearby regions, although numerous exceptions exist in which the methylation status of TEs spread into nearby regions. In conclusion, although TE insertions primarily exhibit methylation patterns restricted within their boundaries, some TEs are able to affect the methylation level of their genomic neighborhoods. At last, our findings are limited to human LCLs, and more comprehensive analysis would be needed to test the rules we found here on a broader spectrum of cell types and developmental stages.
Most genetic risk variants linked to ocular diseases are nonprotein coding and presumably contribute to disease through dysregulation of gene expression; however, understanding their mechanisms has been impeded by incomplete annotation of transcriptional regulatory elements across retinal cell types. To address this, we carried out single-cell multiomics assays to investigate gene expression, chromatin accessibility, DNA methylome, and three-dimensional (3D) chromatin architecture in human retina, macula, and retinal pigment epithelium/choroid. We identified 420,824 unique candidate regulatory elements and characterized their chromatin states in 23 retinal cell types. Comparative analysis of chromatin landscapes between human and mouse retina cells further revealed both evolutionarily conserved and divergent retinal gene-regulatory programs. Leveraging the advancements in deep-learning techniques, we developed sequence-based predictors to interpret noncoding risk variants of retinal diseases. Our study establishes retina-wide, single-cell transcriptome, epigenome, and 3D genome atlases and provides a resource for studying the gene regulatory programs of the human retina and ocular diseases.
Abstract The human pangenome captures genetic diversity beyond a single linear reference, yet conventional DNA methylome maps largely assume that the underlying CpG substrate is fixed. Here we present a first draft of the human panepigenome on DNA methylation, generated from 440 haplotype-resolved long-read methylomes spanning 26 globally distributed populations. This map jointly captures CpG presence, absence and methylation across diverse human haplotypes, revealing 12.8 million CpGs absent from GRCh38 and an average of 1.2 million variant-associated CpGs (var-CpGs) per methylome. Var-CpGs captured population-associated epigenomic variation beyond that represented by shared reference CpGs, and structural variants could create or remove entire CpG islands, frequently through mobile element insertions. Promoter var-CpG methylation was associated with transcript expression, with most associations persisting after adjustment for the underlying CpG-altering variant, whereas a subset showed evidence of mediation or variant-by-methylation interactions. CpG-altering variants were also enriched among molecular QTLs in HPRC2 and across GTEx tissues, and among clinically and pharmacogenomically annotated loci. Together, these results establish a haplotype-resolved framework in which human epigenomic diversity reflects not only variation in methylation level but also genetic variation in presence or absence of CpG substrates, providing a foundation for panepigenomic studies across tissues, populations and disease contexts.
Heart failure is a leading cause of morbidity and mortality, yet gene-regulatory mechanisms driving cell type-specific pathologic responses remain undefined. Here, we present the cell type-resolved transcriptomes, chromatin accessibility, histone modifications, and chromatin organization of 13 nonfailing and 23 failing human hearts across all cardiac chambers. Integrative analyses revealed dynamic changes in cell type composition, gene-regulatory programs, and chromatin organization, particularly in cardiomyocytes and fibroblasts. Mapping cell type-specific enhancer-gene interactions from these analyses enabled the illumination of likely causal genetic contributors to heart failure from genetic association data. Together, these findings provide multimodal gene-regulatory maps of the human heart in health and disease, offering a framework for designing precise, cell type-targeted therapies for treating heart failure.
Environmental toxicant exposures can induce widespread alterations in both the transcriptome and epigenome of mammals, and directly contribute to the increased risk of various diseases, including cardiovascular disorders, cancer, and neurological disorders. To evaluate how early-life toxicants produce long-term impacts on the transcriptome and epigenome in mice, the Toxicant Exposures and Responses by Genomic and Epigenomic Regulators of Transcription II (TaRGET II) Consortium generated a landmark resource comprising 3607 multi-omics datasets from longitudinal studies in mice. The molecular changes in responding to distinct environmental toxicants, including arsenic (As), lead (Pb), bisphenol A (BPA), tributyltin (TBT), di-2-ethylhexyl phthalate (DEHP), dioxin (TCDD), and fine particulate matter (PM2.5), were systematically identified and visualized on an integrative platform, ToxiTaRGET, to allow quickly search and browse by researchers. ToxiTaRGET houses a rich repository of molecular signatures, including gene expression, chromatin accessibility, and DNA methylation profiles, in response to early-life toxicant exposures. These molecular signatures span multiple biologically important tissues in both male and female mice at three distinct life stages, offering a valuable resource for the environmental health and toxicogenomic research communities.
Background The detection of estrogen receptor 1 (ESR1) ligandu2010binding domain mutations in circulating tumor DNA (ctDNA) is crucial for guiding therapy in estrogen receptoru2010positive metastatic breast cancer. However, widespread clinical adoption of approaches for monitoring drug resistance and guiding treatment decisions is hindered by limitations of current methods regarding sensitivity, cost, and multiplexing capability. Methods The application of switchu2010blocker technology, which has been patented for detecting ESR1 hotspot mutations (Y537S/C, D538G, E380Q, and L536H/P), suppresses the amplification of wildu2010type alleles while allowing specific amplification of lowu2010frequency mutant alleles. We used a switchu2010blocker to inhibit the amplification of a DNA target approximately 10 base pairs in length (e.g., the switchu2010blocker covering codon 536 of ESR1 targets various variants at positions 536, 537, and 538). Targeted enrichment was achieved by quantitative polymerase chain reaction, followed by pyrosequencing to confirm mutation components. Nextu2010generation sequencing and Sanger sequencing served as supplementary methods for the verification of results. Results The ESR1u2010targeted DNA assay was validated for feasibility on plasmid circular templates and ctDNA linear templates. In tests using gradientu2010diluted ESR1 plasmid templates, the proportion of L536H mutant copies increased from 0.0015% to 16.89% after targeted amplification, while the proportion of E380Q mutant copies increased from 0.0015% to 1.35%. In ctDNA samples previously analyzed by nextu2010generation sequencing, the switchu2010blocker considerably enriched other mutant copies within the coverage range of the switch element. Conclusions This switchu2010blockeru2013enhanced pyrosequencing assay presents a targeted, multiplexed, and accessible approach for detecting ESR1 hotspot mutations in liquid biopsies. This assay has potential for dynamic monitoring of therapeutic resistance, facilitating timely treatment decisions in advanced breast cancer management.
Somatic variant calling, the identification of mutations in non-germline cells acquired over an individual's lifetime, is critical for studying diseases, including cancer, and for developing precision oncology strategies. Traditional somatic variant calling methods rely on linear reference genomes, which do not adequately capture human genetic diversity and result in reference bias, compromising the accuracy of somatic variant detection. Recently developed graph-based human pangenome reference represents diverse genetic variants across human populations and has promised to drive advances in many genetics and genomics studies. In this study, we introduced Pansoma, a novel pangenome-native and machine learning-based tool specifically designed for somatic variant calling using a pangenome graph reference. Pansoma performs somatic variant detection from both short- and long-read sequencing data by learning tensor representations of alignment on graph nodes rather than on a linear reference. Pansoma outputs variant representations anchored to the pangenome graph paths and conventional somatic variant calls remapped to the linear reference. Additionally, we provide accompanying bioinformatics tools tailored for graph-based genomic data management and variant calling results analysis. Benchmarking shows that Pansoma not only improves tumor-only somatic variant detection but also preserves graph-specific variant representations that are not directly recoverable from linear-reference outputs.
Background:The detection of estrogen receptor 1 (ESR1) ligand-binding domain mutations in circulating tumor DNA (ctDNA) is crucial for guiding therapy in estrogen receptor-positive metastatic breast cancer. However, widespread clinical adoption of approaches for monitoring drug resistance and guiding treatment decisions is hindered by limitations of current methods regarding sensitivity, cost, and multiplexing capability. Methods:The application of switch-blocker technology, which has been patented for detecting ESR1 hotspot mutations (Y537S/C, D538G, E380Q, and L536H/P), suppresses the amplification of wild-type alleles while allowing specific amplification of low-frequency mutant alleles. We used a switch-blocker to inhibit the amplification of a DNA target approximately 10 base pairs in length (e.g., the switch-blocker covering codon 536 of ESR1 targets various variants at positions 536, 537, and 538). Targeted enrichment was achieved by quantitative polymerase chain reaction, followed by pyrosequencing to confirm mutation components. Next-generation sequencing and Sanger sequencing served as supplementary methods for the verification of results. Results:The ESR1-targeted DNA assay was validated for feasibility on plasmid circular templates and ctDNA linear templates. In tests using gradient-diluted ESR1 plasmid templates, the proportion of L536H mutant copies increased from 0.0015% to 16.89% after targeted amplification, while the proportion of E380Q mutant copies increased from 0.0015% to 1.35%. In ctDNA samples previously analyzed by next-generation sequencing, the switch-blocker considerably enriched other mutant copies within the coverage range of the switch element. Conclusions:This switch-blocker-enhanced pyrosequencing assay presents a targeted, multiplexed, and accessible approach for detecting ESR1 hotspot mutations in liquid biopsies. This assay has potential for dynamic monitoring of therapeutic resistance, facilitating timely treatment decisions in advanced breast cancer management.
Directly measuring chromatin states alongside transcription is essential for understanding how cell-type-specific regulatory programs are established and maintained in the adult human brain. We present a large-scale single-cell multimodal atlas generated by jointly profiling transcriptome with active (H3K27ac) and repressive (H3K27me3) histone modifications across 18 brain regions. We profile >750,000 nuclei spanning 160 cell types and integrate these data with chromatin accessibility, DNA methylation, 3D genome architecture, and spatial transcriptome. This framework annotates >500,000 regulatory elements and resolves cell-type-specific chromatin states. We link enhancers to target genes, infer gene regulatory networks, and classify chromatin interactions, revealing neuron-enriched long-range Polycomb repression of developmental genes. Integrating these maps with GWAS data and sequence-based model prioritizes noncoding variants, effector genes, and vulnerable cell types for neuropsychiatric disorders. Finally, cross-species comparisons show conserved activation but more divergent repression. Together, this study provides a functional reference for interpreting noncoding variants, epigenetic memory, and brain organization.
The basal ganglia play essential roles in motor control, emotion, learning and reward processing. Their dysfunction contributes to many neurological and psychiatric disorders. However, the gene regulatory programs defining basal ganglia cell-type identity and function remain poorly understood, limiting interpretation of disease-associated non-coding variants. Here, we present the first single-cell multiome atlas of histone modifications and transcriptomes across eight basal ganglia regions from neurotypical adult human donors. Joint profiling reveals cell-type-specific deployment of active and repressive cis-regulatory elements and gene regulatory networks, and suggests a combinatorial homeobox transcription factor code underlying cell identity. Integration with matched spatial transcriptomic MERFISH data uncovers regional heterogeneity of epigenomic landscapes. Comparative analysis between human and mouse medium spiny neurons uncovers conservation of core gene regulatory features. This atlas interprets non-coding risk variants of neuropsychiatric disorders and supports the development of a deep learning model to predict gene regulation and functional effects of disease-associated variants.