The Toxicant Exposures and Responses by Genomic and Epigenomic Regulators of Transcription (TaRGET) program is a multiphase program that aims to understand how environmental factors contribute to disease susceptibility using toxicant-exposed mouse models. Here, we present the TaRGET II Data Portal ( https://data.targetepigenomics.org/ ), a comprehensive repository of multi-omics datasets generated from mice exposed to a range of environmental toxicants, including arsenic (As), lead (Pb), bisphenol A (BPA), tributyltin (TBT), di(2-ethylhexyl) phthalate (DEHP), tetrachlorodibenzo-p-dioxin (TCDD), and ambient air pollution (PM2.5). Epigenomic profiling assays capturing alterations in chromatin accessibility, DNA methylation, gene expression, and histone modifications were produced across multiple research centers, subjected to rigorous quality control, and processed using standardized pipelines. The portal currently hosts 3,612 datasets spanning multiple tissues and genomic assays, collected at four developmental and exposure time points: 3, 5, 20, and 40 weeks. The portal offers an efficient way to browse, search, visualize, and download relevant datasets and associated metadata, serving as a key resource for studying the impact of environmental toxicant exposures on disease susceptibility for the broader scientific community.
The basal ganglia are a group of forebrain nuclei critical for motor control and reward processing, and their dysfunction contributes to neurological and neuropsychiatric disorders. Here, we present the first multimodal single-cell epigenomic atlas of the human basal ganglia across major subregions and cell types. We jointly profiled DNA methylation and 3D chromatin conformation in 197,003 nuclei from eight basal ganglia subregions using multi-omic sequencing (snm3C-seq), and integrated these data with existing DNA methylation and chromatin conformation sequencing datasets to build a unified atlas of 261,331 cells spanning 31 subclasses and 59 groups. This atlas reveals extensive cell-type- and region-specific differential methylation, enriched for distinct transcription factor motifs, and validated by MERFISH spatial transcriptomics, which uncovered epigenetic gradients linked to transcriptional output. Compared to neuronal cells, non-neuronal cells exhibit distinct 3D genome organization including smaller chromatin compartments, increased long-range inter-compartment contacts, shorter loops, and stronger CG hypomethylation in A compartments. We further identified genes that display compartment switches, are strongly correlated with compartment scores, and exhibit differential domain boundaries and chromatin looping across basal ganglia cell types. We identified multiple medium spiny neuron subtypes defined by distinct hypomethylated signature genes, with 3D genome embeddings emphasizing dorsal, ventral, and hybrid populations. By integrating chromatin accessibility and histone modification profiles, we reconstructed cell-type-resolved enhancer-promoter links and gene regulatory networks, providing a comprehensive epigenomic framework for interpreting genetic risk loci and regulatory architecture in the human basal ganglia.
The basal ganglia regulate motor, cognitive, and affective behaviors, and their dysfunction underlies diverse neurological and psychiatric disorders. Comprehensive, accessible multi-omics resources are needed to understand the regulatory mechanisms governing basal ganglia cell types. Here we present an open, interactive web-based platform for exploring single-cell multi-omics datasets from basal ganglia, generated using 10X Multiome, snm3C-seq, and Paired-Tag technologies from the BICAN (NIH BRAIN Initiative Cell Atlas Network) consortium. The platform is available at https://basalganglia.epigenomes.net/ and enables integrated visualization of gene expression, chromatin accessibility, DNA methylation, histone modifications, and chromatin conformation across cell types and human, macaque, marmoset, and mouse species, with direct genome browser support and comparative epigenomic functionality. Representative analyses demonstrate cell-type-specific regulatory landscapes, conserved and species-specific regulatory elements, and links between epigenomic regulation and transcription. This resource provides a scalable, community-oriented foundation for advancing basal ganglia biology and interpreting regulatory mechanisms relevant to brain function and disease.
Somatic variant detection is technically challenging due to low variant allele fractions, the confounding presence of germline variation, and reference bias. Linear references such as GRCh38 miss sample-specific variation, causing misalignments and incorrect variant calls. Although telomere-to-telomere donor-specific assemblies (DSAs) accurately represent individual genomes, their application is limited by cost and technical barriers. Alternatively, the graph-based human pangenome provides a scalable framework to improve read alignment and perform genome inference. Here, we benchmarked somatic variant detection using GRCh38, graph-based pangenomes, and pangenome-inferred DSAs with a HapMap mixture dataset and the COLO829 melanoma cell line. Pangenome-guided alignment improves read mapping and somatic variant calling accuracy. Furthermore, personalized pangenomes partially reconstruct donor-specific genomic content, improving accuracy, reducing germline contamination, and enabling detection of events in loci absent or poorly represented in GRCh38. These findings demonstrate that graph-based and personalized pangenomes are effective strategies for enhancing somatic variant detection compared with GRCh38.
Most genetic risk variants linked to ocular diseases are nonprotein coding and presumably contribute to disease through dysregulation of gene expression; however, understanding their mechanisms has been impeded by incomplete annotation of transcriptional regulatory elements across retinal cell types. To address this, we carried out single-cell multiomics assays to investigate gene expression, chromatin accessibility, DNA methylome, and three-dimensional (3D) chromatin architecture in human retina, macula, and retinal pigment epithelium/choroid. We identified 420,824 unique candidate regulatory elements and characterized their chromatin states in 23 retinal cell types. Comparative analysis of chromatin landscapes between human and mouse retina cells further revealed both evolutionarily conserved and divergent retinal gene-regulatory programs. Leveraging the advancements in deep-learning techniques, we developed sequence-based predictors to interpret noncoding risk variants of retinal diseases. Our study establishes retina-wide, single-cell transcriptome, epigenome, and 3D genome atlases and provides a resource for studying the gene regulatory programs of the human retina and ocular diseases.
Abstract The human pangenome captures genetic diversity beyond a single linear reference, yet conventional DNA methylome maps largely assume that the underlying CpG substrate is fixed. Here we present a first draft of the human panepigenome on DNA methylation, generated from 440 haplotype-resolved long-read methylomes spanning 26 globally distributed populations. This map jointly captures CpG presence, absence and methylation across diverse human haplotypes, revealing 12.8 million CpGs absent from GRCh38 and an average of 1.2 million variant-associated CpGs (var-CpGs) per methylome. Var-CpGs captured population-associated epigenomic variation beyond that represented by shared reference CpGs, and structural variants could create or remove entire CpG islands, frequently through mobile element insertions. Promoter var-CpG methylation was associated with transcript expression, with most associations persisting after adjustment for the underlying CpG-altering variant, whereas a subset showed evidence of mediation or variant-by-methylation interactions. CpG-altering variants were also enriched among molecular QTLs in HPRC2 and across GTEx tissues, and among clinically and pharmacogenomically annotated loci. Together, these results establish a haplotype-resolved framework in which human epigenomic diversity reflects not only variation in methylation level but also genetic variation in presence or absence of CpG substrates, providing a foundation for panepigenomic studies across tissues, populations and disease contexts.
Heart failure is a leading cause of morbidity and mortality, yet gene-regulatory mechanisms driving cell type-specific pathologic responses remain undefined. Here, we present the cell type-resolved transcriptomes, chromatin accessibility, histone modifications, and chromatin organization of 13 nonfailing and 23 failing human hearts across all cardiac chambers. Integrative analyses revealed dynamic changes in cell type composition, gene-regulatory programs, and chromatin organization, particularly in cardiomyocytes and fibroblasts. Mapping cell type-specific enhancer-gene interactions from these analyses enabled the illumination of likely causal genetic contributors to heart failure from genetic association data. Together, these findings provide multimodal gene-regulatory maps of the human heart in health and disease, offering a framework for designing precise, cell type-targeted therapies for treating heart failure.
Environmental toxicant exposures can induce widespread alterations in both the transcriptome and epigenome of mammals, and directly contribute to the increased risk of various diseases, including cardiovascular disorders, cancer, and neurological disorders. To evaluate how early-life toxicants produce long-term impacts on the transcriptome and epigenome in mice, the Toxicant Exposures and Responses by Genomic and Epigenomic Regulators of Transcription II (TaRGET II) Consortium generated a landmark resource comprising 3607 multi-omics datasets from longitudinal studies in mice. The molecular changes in responding to distinct environmental toxicants, including arsenic (As), lead (Pb), bisphenol A (BPA), tributyltin (TBT), di-2-ethylhexyl phthalate (DEHP), dioxin (TCDD), and fine particulate matter (PM2.5), were systematically identified and visualized on an integrative platform, ToxiTaRGET, to allow quickly search and browse by researchers. ToxiTaRGET houses a rich repository of molecular signatures, including gene expression, chromatin accessibility, and DNA methylation profiles, in response to early-life toxicant exposures. These molecular signatures span multiple biologically important tissues in both male and female mice at three distinct life stages, offering a valuable resource for the environmental health and toxicogenomic research communities.
Somatic variant calling, the identification of mutations in non-germline cells acquired over an individual's lifetime, is critical for studying diseases, including cancer, and for developing precision oncology strategies. Traditional somatic variant calling methods rely on linear reference genomes, which do not adequately capture human genetic diversity and result in reference bias, compromising the accuracy of somatic variant detection. Recently developed graph-based human pangenome reference represents diverse genetic variants across human populations and has promised to drive advances in many genetics and genomics studies. In this study, we introduced Pansoma, a novel pangenome-native and machine learning-based tool specifically designed for somatic variant calling using a pangenome graph reference. Pansoma performs somatic variant detection from both short- and long-read sequencing data by learning tensor representations of alignment on graph nodes rather than on a linear reference. Pansoma outputs variant representations anchored to the pangenome graph paths and conventional somatic variant calls remapped to the linear reference. Additionally, we provide accompanying bioinformatics tools tailored for graph-based genomic data management and variant calling results analysis. Benchmarking shows that Pansoma not only improves tumor-only somatic variant detection but also preserves graph-specific variant representations that are not directly recoverable from linear-reference outputs.
Directly measuring chromatin states alongside transcription is essential for understanding how cell-type-specific regulatory programs are established and maintained in the adult human brain. We present a large-scale single-cell multimodal atlas generated by jointly profiling transcriptome with active (H3K27ac) and repressive (H3K27me3) histone modifications across 18 brain regions. We profile >750,000 nuclei spanning 160 cell types and integrate these data with chromatin accessibility, DNA methylation, 3D genome architecture, and spatial transcriptome. This framework annotates >500,000 regulatory elements and resolves cell-type-specific chromatin states. We link enhancers to target genes, infer gene regulatory networks, and classify chromatin interactions, revealing neuron-enriched long-range Polycomb repression of developmental genes. Integrating these maps with GWAS data and sequence-based model prioritizes noncoding variants, effector genes, and vulnerable cell types for neuropsychiatric disorders. Finally, cross-species comparisons show conserved activation but more divergent repression. Together, this study provides a functional reference for interpreting noncoding variants, epigenetic memory, and brain organization.
The basal ganglia play essential roles in motor control, emotion, learning and reward processing. Their dysfunction contributes to many neurological and psychiatric disorders. However, the gene regulatory programs defining basal ganglia cell-type identity and function remain poorly understood, limiting interpretation of disease-associated non-coding variants. Here, we present the first single-cell multiome atlas of histone modifications and transcriptomes across eight basal ganglia regions from neurotypical adult human donors. Joint profiling reveals cell-type-specific deployment of active and repressive cis-regulatory elements and gene regulatory networks, and suggests a combinatorial homeobox transcription factor code underlying cell identity. Integration with matched spatial transcriptomic MERFISH data uncovers regional heterogeneity of epigenomic landscapes. Comparative analysis between human and mouse medium spiny neurons uncovers conservation of core gene regulatory features. This atlas interprets non-coding risk variants of neuropsychiatric disorders and supports the development of a deep learning model to predict gene regulation and functional effects of disease-associated variants.
A pangenome reference overcomes the inherent limitation of any individual reference genome by integrating the variation present in a population. We present the Human Pangenome Reference Consortium's (HPRC) Release 2 (HPRC2), an openly available, second phase pangenome that is an approximately fivefold expansion in genome number over HPRC Release 1 (HPRC1) and measurable improvement in genome completeness, contiguity, and accuracy. Selecting samples with a principled algorithm prioritising common variant coverage, HPRC2 contributes 460 haplotypes that together capture over 99% of common variation observed in the All of Us Research Program v8 cohort. Combining high-coverage long and ultra-long reads with modern assemblers and polishers, we produce thousands of telomere-to-telomere (T2T) chromosomes, and relative to HPRC1 halve the number of structurally unreliable regions as well as individual base errors per haplotype. We complement the assemblies with whole genome multiple alignments and gene annotations, and derive formal pangenome coordinate systems for addressing off-reference variation, demonstrating that individual human genomes contain more than one hundred thousand variants not succinctly described with respect to existing reference genomes. We also present the first matched long-read backed pantranscriptome and panepigenome at this scale, provide continuous local-ancestry estimates spanning every genome, and outline a host of new tools and applications that leverage the pangenome resource for improved genomics analysis.
Genomic variation between individuals is essential for understanding how differences in the genome sequence affect molecular and cellular processes. The Impact of Genomic Variation on Function (IGVF) Consortium aims to uncover the relationships among genomic variation, genome function, and phenotypes by combining experimental techniques, such as single-cell mapping and genomic perturbation assays, with computational approaches such as machine learning-based predictive modeling. The IGVF Data and Administrative Coordinating Centers collect, analyze, and disseminate data and results from across the consortium through an open-source platform called the IGVF Catalog. This resource includes, but is not limited to, data on the effects of coding variants on protein abundance and function, noncoding variants on enhancer activity (measured by MPRA or predicted computationally), and associations between variants and quantitative traits. All data are organized within a graph database comprising over 50 types of data collections with nearly 3 billion nodes and over 7.5 billion edges. The Catalog offers public API endpoints (https://api.catalogkg.igvf.org/) and a user-friendly interface for exploring, querying, and visualizing the data at https://catalog.igvf.org. We expect that this open-access platform will support the broader scientific community to advance our understanding of how genomic variation influences biology and disease.
Changes in gene expression have been observed in the aging human brain, but our understanding of the underlying regulatory mechanisms remains limited. To unravel these complexities, we analyzed single-nucleus gene expression, chromatin accessibility, DNA methylation, and three-dimensional (3D) chromatin architecture from human hippocampal tissues spanning the adult lifespan. We identified both linear and nonlinear dynamic gene regulatory programs during aging. Between the ages of 50 to 75, embryonic yolk sac-derived microglia were depleted and replaced by cells resembling peripheral blood monocyte-derived microglia. Hippocampal astrocytes decreased substantially with age, including those regulating synaptic transmission. Across cell types, 3D genome architecture underwent global erosion. Our analysis provides insights for how altered gene regulatory programs promote cell type-specific aging phenotypes in the human brain.
The human genome reference established a shared coordinate system for genome function, but it is incomplete and not fully representative of human diversity. Here, we benchmark how genome representation and corresponding analytical frameworks for each representation shape functional genomics using chromatin accessibility sequencing (ATAC-seq), RNA sequencing, whole-genome bisulfite sequencing, and chromosome conformation capture (Hi-C) data from lymphoblastoid cell lines derived from five individuals with fully phased genome assemblies. We compare results across hg38, CHM13, the draft human pangenome, and each individual's maternal and paternal assemblies. Because current pipelines and quality control conventions are tuned to hg38, several of these comparisons reflect genome representation in the context of available methods, rather than sequence alone. Individual identity accounts for 57.52-78.47% of total variance in functional estimates, whereas genome choice contributes 0.002-7.85% and sample-by-genome interactions contribute 0.63-5.43%. About 2% of biological signals are detectable only with personal assemblies. Although these effects are modest overall, some biologically important features remain inaccessible to linear references. Consistent with this, graph-based DNA methylation analysis in the human pangenome reveals a non-reference AluY5a insertion within a putative TNKS enhancer at chromosome 8p23.1 that becomes visible and hypermethylated only in the pangenome.
Heart failure is a leading cause of morbidity and mortality; yet gene regulatory mechanisms driving cell type-specific pathologic responses remain undefined. Here, we present the cell type-resolved transcriptomes, chromatin accessibility, histone modifications and chromatin organization of 36 non-failing and failing human hearts profiled from 776,479 cells spanning all cardiac chambers. Integrative analyses revealed dynamic changes in cell type composition, gene regulatory programs and chromatin organization, which expanded the annotation of cardiac cis-regulatory sequences by ten-fold and mapped cell type-specific enhancer-gene interactions. Cardiomyocytes and fibroblasts particularly exhibited complex disease-associated cellular states, gene regulatory programs and global chromatin reorganization. Mapping genetic association data onto cell type-specific regulatory programs revealed likely causal genetic contributors to heart failure. Together, these findings provide comprehensive, multimodal gene regulatory maps of the human heart in health and disease, offering a valuable framework for designing precise cell type-targeted therapies for treating heart failure.
Metabolic dysfunction-associated steatotic liver disease (MASLD) has limited treatments, and cell type-specific regulatory networks driving MASLD represent therapeutic avenues. We assayed five transcriptomic and epigenomic modalities in 2.4M cells from 86 livers across MASLD stages. Integrating modalities increased annotation of the genome in liver cell types several-fold over previous catalogs. We identified cell type regulatory networks of MASLD progression, including distinct hepatocyte networks driving MASL and mild and severe fibrosis MASH. Our single cell atlas annotated 88% of MASH-associated loci, including a third affecting hepatocyte regulation which we linked to distal target genes. Finally, we characterized hepatocyte heterogeneity, including MASH-enriched populations with altered repression, localization, and signaling. Overall, our results provide high-resolution maps of liver cell types and revealed novel targets for anti-MASH therapy.
The Toxicant Exposures and Responses by Genomic and Epigenomic Regulators of Transcription (TaRGET) is a multiphase program that aims to understand how environmental factors contribute to disease susceptibility using toxicant-exposed mouse models. Here, we introduce the TaRGET II Data Portal (), a repository of exposomes from mice exposed to environmental toxicants, including arsenic (As), lead (Pb), bisphenol A (BPA), tributyltin (TBT), di(2-ethylhexyl) phthalate (DEHP), tetrachlorodibenzo-p-dioxin (TCDD), and air pollution (PM2.5). Sequencing assays capturing changes in chromatin accessibility, DNA methylation, gene expression, and post-translational histone modifications from multiple centers were quality-controlled and uniformly processed. The datasets cover multiple tissues collected at four time points: 3 weeks, 5 weeks, 20 weeks, and 40 weeks. The TaRGET II Data Portal offers an efficient way to browse, search, visualize, and download relevant datasets and associated metadata, serving as a key resource for studying the impact of environmental toxicant exposures on disease susceptibility for the broader scientific community. ### Competing Interest Statement The authors have declared no competing interest. National Institute Health, U24ES026699
The WashU Epigenome Browser (https://epigenomegateway.wustl.edu/) is a web-based tool for exploring genomic data and providing visualization, investigation, and analysis of epigenomic datasets. Since its 2018 update, the redesigned user interface and newly developed features have enhanced how investigators interact with both the Browser and the extensive genomic data it hosts. The rapid evolution of the JavaScript ecosystem has presented new challenges and opportunities in maintaining and developing the WashU Epigenome Browser. In this update, we present a completely rewritten codebase. This new codebase minimizes the use of external libraries whenever possible, resulting in a significantly smaller code bundle size after production compilation. The reduced code size improves loading efficiency and boosts the Browser's performance, with improved scripting, graphics rendering, and painting performance. Lowering external dependencies also allows for faster and more straightforward installation. Additionally, the update includes a redesign of the user interface to further enhance user experience and features a new modular design in the codebase that enables the Browser to be exported as stand-alone modules for use in other web applications. Several novel track types for long-read methylation data and single-cell methylation data visualization have been added, and we continue to update and expand the data hubs we host for major consortia. We constructed the first data hub to systematically compare genomic data mapped to different genome assemblies, focusing on comparisons between hg38 and the first human T2T genome, chm13, using our new comparative genomics track function. The WashU Epigenome Browser also serves as a foundation for other genomics platforms, such as the WashU Virus Genome Browser, developed for SARS-COV-2 research, the WashU Comparative Epigenome Browser, and the WashU Repeat Browser.
We generated a multi-region, subcellular-resolution spatial transcriptomic atlas of the human basal ganglia by integrating MERFISH+ and Stereo-seq across four neurotypical donors. These datasets profiled ~7 million cells spanning the caudate, putamen, nucleus accumbens, and globus pallidus, resolving 60 transcriptionally distinct cell types. We show region-selective, molecular and spatial diversification of medium-spiny-neuron cell types and multiple non-neuronal populations with distinct molecular identities and spatial localizations. Subcellular RNA localization captures somatic size and projection-inferred signatures that reflect direct and indirect pathway topology. Cellular community analyses reveal the enrichment of sub-clusters of astrocytes and oligodendrocytes at striosome-matrix borders, while primate-expanded interneurons are confined to matrix territories. Cross-species mapping uncovers orthologous striosome-matrix organization and conserved dorsolateral-ventromedial gene expression gradients. This atlas provides a foundational molecular and spatial framework for studying human basal ganglia architecture, offering a multi-centimeter scale resource that links cell types, spatial architecture, and subcellular transcript topography across multiple nuclei. ### Competing Interest Statement B.B. and Q.Z are co-inventors on a patent application for MERFISH+ filed by the University of California, San Diego. To date, two US provisional patent applications have been filed. The other authors declare no competing interests. National Institutes of Health, UM1MH130994