Centromeres are essential chromosome components yet remain poorly understood due to their highly repetitive sequence architecture. Using fully-phased telomere-to-telomere diploid assemblies from a three-generation pedigree integrated with long-read epigenomes from matched peripheral blood mononuclear cells, induced pluripotent stem cells, and neural progenitor cells, we generate allele-resolved single basepair resolution maps of centromere genetic and epigenetic dynamics across inheritance, reprogramming, and differentiation. We show that centromeric dip regions (CDRs), which define the functional core of centromeres, are positionally stable across generations and cell-fate transitions. In contrast, CDR epigenetic architecture is highly dynamic. Reprogramming markedly attenuates CDR hypomethylation, which is partially restored during differentiation in parallel with global hypomethylation of active alpha-satellite arrays and coordinated changes in nucleosome organization and protein occupancy. Centromeric remodeling is insulated from X-chromosome status, including Xa, Xi, and erosion. Finally, de novo mutations arising during reprogramming are enriched in centromeric regions but depleted within functional centromeric cores.
BACKGROUND: Transposable elements (TEs) are vital components of eukaryotic genomes and have played a critical role in genome evolution. Although most TEs are silenced in the mammalian genome, increasing evidence suggests that certain TEs are actively involved in gene regulation during early developmental stages. However, the extent to which human TEs drive gene transcription in adult tissues remains largely unexplored. RESULTS: In this study, we systematically analyze 17,329 human transcriptomes to investigate how TEs influence gene transcription across 47 adult tissues. Our findings reveal that TE-driven transcripts are broadly expressed in human tissues, contributing to both housekeeping functions and tissue-specific gene regulation. We identify sex-specific expression of TE-driven transcripts regulated by sex hormones in breast tissue between females and males. Our results suggest that TE-driven alternative transcription initiation significantly enhances the variety of translated protein products, e.g., changes in the N-terminal peptide length of WNT2B caused by TE-driven transcription result in isoform-specific subcellular localization. Additionally, we identify 68 human-specific TE-driven transcripts of genes, which are involved in distinct biological processes. CONCLUSIONS: These findings indicate the regulatory role of TEs in human genome evolution and highlight their substantial contribution to the diversification of transcriptional and translational outputs.
The Toxicant Exposures and Responses by Genomic and Epigenomic Regulators of Transcription (TaRGET) program is a multiphase program that aims to understand how environmental factors contribute to disease susceptibility using toxicant-exposed mouse models. Here, we present the TaRGET II Data Portal ( https://data.targetepigenomics.org/ ), a comprehensive repository of multi-omics datasets generated from mice exposed to a range of environmental toxicants, including arsenic (As), lead (Pb), bisphenol A (BPA), tributyltin (TBT), di(2-ethylhexyl) phthalate (DEHP), tetrachlorodibenzo-p-dioxin (TCDD), and ambient air pollution (PM2.5). Epigenomic profiling assays capturing alterations in chromatin accessibility, DNA methylation, gene expression, and histone modifications were produced across multiple research centers, subjected to rigorous quality control, and processed using standardized pipelines. The portal currently hosts 3,612 datasets spanning multiple tissues and genomic assays, collected at four developmental and exposure time points: 3, 5, 20, and 40 weeks. The portal offers an efficient way to browse, search, visualize, and download relevant datasets and associated metadata, serving as a key resource for studying the impact of environmental toxicant exposures on disease susceptibility for the broader scientific community.
The basal ganglia regulate motor, cognitive, and affective behaviors, and their dysfunction underlies diverse neurological and psychiatric disorders. Comprehensive, accessible multi-omics resources are needed to understand the regulatory mechanisms governing basal ganglia cell types. Here we present an open, interactive web-based platform for exploring single-cell multi-omics datasets from basal ganglia, generated using 10X Multiome, snm3C-seq, and Paired-Tag technologies from the BICAN (NIH BRAIN Initiative Cell Atlas Network) consortium. The platform is available at https://basalganglia.epigenomes.net/ and enables integrated visualization of gene expression, chromatin accessibility, DNA methylation, histone modifications, and chromatin conformation across cell types and human, macaque, marmoset, and mouse species, with direct genome browser support and comparative epigenomic functionality. Representative analyses demonstrate cell-type-specific regulatory landscapes, conserved and species-specific regulatory elements, and links between epigenomic regulation and transcription. This resource provides a scalable, community-oriented foundation for advancing basal ganglia biology and interpreting regulatory mechanisms relevant to brain function and disease.
Somatic variant detection is technically challenging due to low variant allele fractions, the confounding presence of germline variation, and reference bias. Linear references such as GRCh38 miss sample-specific variation, causing misalignments and incorrect variant calls. Although telomere-to-telomere donor-specific assemblies (DSAs) accurately represent individual genomes, their application is limited by cost and technical barriers. Alternatively, the graph-based human pangenome provides a scalable framework to improve read alignment and perform genome inference. Here, we benchmarked somatic variant detection using GRCh38, graph-based pangenomes, and pangenome-inferred DSAs with a HapMap mixture dataset and the COLO829 melanoma cell line. Pangenome-guided alignment improves read mapping and somatic variant calling accuracy. Furthermore, personalized pangenomes partially reconstruct donor-specific genomic content, improving accuracy, reducing germline contamination, and enabling detection of events in loci absent or poorly represented in GRCh38. These findings demonstrate that graph-based and personalized pangenomes are effective strategies for enhancing somatic variant detection compared with GRCh38.
Environmental toxicant exposures can induce widespread alterations in both the transcriptome and epigenome of mammals, and directly contribute to the increased risk of various diseases, including cardiovascular disorders, cancer, and neurological disorders. To evaluate how early-life toxicants produce long-term impacts on the transcriptome and epigenome in mice, the Toxicant Exposures and Responses by Genomic and Epigenomic Regulators of Transcription II (TaRGET II) Consortium generated a landmark resource comprising 3607 multi-omics datasets from longitudinal studies in mice. The molecular changes in responding to distinct environmental toxicants, including arsenic (As), lead (Pb), bisphenol A (BPA), tributyltin (TBT), di-2-ethylhexyl phthalate (DEHP), dioxin (TCDD), and fine particulate matter (PM2.5), were systematically identified and visualized on an integrative platform, ToxiTaRGET, to allow quickly search and browse by researchers. ToxiTaRGET houses a rich repository of molecular signatures, including gene expression, chromatin accessibility, and DNA methylation profiles, in response to early-life toxicant exposures. These molecular signatures span multiple biologically important tissues in both male and female mice at three distinct life stages, offering a valuable resource for the environmental health and toxicogenomic research communities.
The Toxicant Exposures and Responses by Genomic and Epigenomic Regulators of Transcription (TaRGET) is a multiphase program that aims to understand how environmental factors contribute to disease susceptibility using toxicant-exposed mouse models. Here, we introduce the TaRGET II Data Portal (), a repository of exposomes from mice exposed to environmental toxicants, including arsenic (As), lead (Pb), bisphenol A (BPA), tributyltin (TBT), di(2-ethylhexyl) phthalate (DEHP), tetrachlorodibenzo-p-dioxin (TCDD), and air pollution (PM2.5). Sequencing assays capturing changes in chromatin accessibility, DNA methylation, gene expression, and post-translational histone modifications from multiple centers were quality-controlled and uniformly processed. The datasets cover multiple tissues collected at four time points: 3 weeks, 5 weeks, 20 weeks, and 40 weeks. The TaRGET II Data Portal offers an efficient way to browse, search, visualize, and download relevant datasets and associated metadata, serving as a key resource for studying the impact of environmental toxicant exposures on disease susceptibility for the broader scientific community. ### Competing Interest Statement The authors have declared no competing interest. National Institute Health, U24ES026699
Exposure to toxic substances, particularly early in life, can perturb epigenomic marks linked to disease susceptibility. Human studies of environmental exposures often rely on surrogate tissues such as blood, but toxicant accumulation differs across organs and results in tissue-specific responses. Thus, understanding whether exposure-induced epigenomic alterations in surrogate tissues such as blood reflect changes in toxicant target tissues, such as liver, is essential for designing and interpreting environmental epigenetic studies. To address this knowledge gap, we systematically analyzed 1,013 multi-omics data from the TaRGET II Consortium, comparing molecular responses in mouse liver and blood following perinatal exposure to arsenic, lead, bisphenol A, tributyltin, di-2-ethylhexyl phthalate, tetrachlorodibenzo-p-dioxin, or air pollution in the form of particulate matter < 2.5μm (PM2.5). Most toxicant-induced molecular changes were tissue-specific, yet we identified a subset of co-regulated genes and regulatory elements in liver and blood in response to early-life exposure to toxicants. Moreover, we discovered that specific pathways, such as immune-related processes, were commonly affected by exposures in both tissues, and transcription factors, including Klf, Jun, Ets1, and Cebp, emerged as shared regulators. While molecular alterations are infrequently conserved between tissues following toxicant exposure, the shared alterations in transcription factors and biological pathways may provide a strategy to link effects in surrogate tissues to target tissues.
Mobile element insertions (MEI) shape the human genome in both germline and somatic tissues. While inherited MEIs are well characterized, mapping somatic MEIs (sMEI) in non-cancer tissues remains challenging due to their low allelic fraction and repetitive nature. We established an integrative framework for sMEI analysis leveraging modern sequencing technologies and analytical innovations. We first benchmarked sMEI detection and demonstrated advantages of long-read and MEI-targeted sequencing for ultra-low-frequency events using a mixture of well-established cell lines. We then showed that haplotype phasing and donor-specific assemblies refine sMEI detection, effectively distinguishing from germline and false signals in in-silico tumor-normal mixtures. We further developed a source-tracing strategy based on internal sequence variation, expanding the catalogue of active source elements beyond traditional transduction-based methods. Applying this framework to donor tissues, we identified 18 rare somatic L1 insertions, revealing structural and source diversity. Our work provides a foundational framework and biological insight into sMEIs.
Environmental exposures to toxic chemicals can profoundly alter the transcriptome and epigenome in both humans and animals, contributing to disease development across the lifespan. To elucidate how early-life exposure to toxicants exerts such persistent effects, the Toxicant Exposures and Responses by Genomic and Epigenomic Regulators of Transcription II (TaRGET II) Consortium generated a landmark resource comprising 2,570 epigenomes and 1,043 transcriptomes from longitudinal studies in mice. All data are publicly available through the TaRGET II data portal and the WashU Epigenome Browser. This resource from target (liver, brain, lung, heart) and surrogate (blood) tissues at weaning (3 weeks) and two adult time-points (5 and 10 months) characterized the molecular response to arsenic (As), lead (Pb), bisphenol-A (BPA), di-2-ethylhexyl phthalate(DEHP), tributyltin (TBT), tetrachlorodibenzo-p-dioxin (TCDD), and particulate matter with a diameter of <2.5μm (PM2.5). The findings revealed persistent, toxicant-specific, sex-dependent epigenomic and transcriptomic perturbations, resulting in disrupted expression of 14,908 genes, altered chromatin accessibility at 87,409 regulatory elements, DNA methylation changes at 113,186 genomic regions, and chromatin state switching of histone modifications. The resulting high-resolution map of how environmental exposures reprogram the epigenome and transcriptome is broadly accessible via ToxiTaRGET database, offering unparalleled opportunities for the scientific community to investigate the molecular underpinnings of environmental toxicant exposures and their contributions to disease pathogenesis.
Environmental toxicant exposures can induce widespread alterations in both the transcriptome and epigenome of mammals, and directly contribute to the increased risk of various diseases, including cardiovascular disorders, cancer, and neurological disorders. To evaluate how early-life toxicants produce long-term impacts on the transcriptome and epigenome in mice, the Toxicant Exposures and Responses by Genomic and Epigenomic Regulators of Transcription II (TaRGET II) Consortium generated a landmark resource comprising 3,607 multi-omics from longitudinal studies in mice. The molecular changes in responding to distinct environmental toxicants, including arsenic (As), lead (Pb), bisphenol A (BPA), tributyltin (TBT), di-2-ethylhexyl phthalate (DEHP), dioxin (TCDD), and fine particulate matter (PM2.5), were systematically identified and visualized on an integrative platform, ToxiTaRGET, to allow quickly search and browse by researchers. ToxiTaRGET houses a rich repository of molecular signatures, including gene expression, chromatin accessibility, and DNA methylation profiles, in response to early-life toxicant exposures. These molecular signatures span multiple biologically important tissues in both male and female mice at three distinct life stages, offering a valuable resource for the environmental health and toxicogenomic research communities.
Environmental exposures to toxic chemicals can profoundly alter the transcriptome and epigenome in both humans and animals, contributing to disease development across the lifespan. To elucidate how early-life exposure to toxicants exerts such persistent effects, the Toxicant Exposures and Responses by Genomic and Epigenomic Regulators of Transcription II (TaRGET II) Consortium generated a landmark resource comprising 2,564 epigenomes and 1,043 transcriptomes from longitudinal studies in mice. All data are publicly available through the TaRGET II data portal and the WashU Epigenome Browser. This resource from target (liver, brain, lung, heart) and surrogate (blood) tissues at weaning (3 weeks) and two adult time-points (5 and 10 months) characterized the molecular response to arsenic (As), lead (Pb), bisphenol-A (BPA), di-2-ethylhexyl phthalate(DEHP), tributyltin (TBT), tetrachlorodibenzo-p-dioxin (TCDD), and particulate matter with a diameter of <2.5μm (PM2.5). The findings revealed persistent, toxicant-specific, sex-dependent epigenomic and transcriptomic perturbations, resulting in disrupted expression of 14,908 genes, altered chromatin accessibility at 87,409 regulatory elements, DNA methylation changes at 113,186 genomic regions, and chromatin state switching of histone modifications. The resulting high-resolution map of how environmental exposures reprogram the epigenome and transcriptome is broadly accessible via ToxiTaRGET database, offering unparalleled opportunities for the scientific community to investigate the molecular underpinnings of environmental toxicant exposures and their contributions to disease pathogenesis.
Transposable elements (TEs) are mobile DNA sequences that constitute a significant portion of mammalian genomes. While typically silenced by epigenetic mechanisms, mounting evidence indicates TEs can regulate gene expression and chromatin architecture. However, their regulatory roles under various environmental exposures remain largely unexplored. In this study, we investigate the regulatory functions of TEs in mouse liver tissue following early-life exposure to environmental toxicants, including arsenic (As), lead (Pb), bisphenol A (BPA), tributyltin (TBT), di-2-ethylhexyl phthalate (DEHP), tetrachlorodibenzo-p-dioxin (TCDD), and particulate matter less than 2.5 micrometers (PM2.5). These toxicants are linked to various health issues, including neurodevelopmental deficits, metabolic and immune dysfunction, and increased cancer risks. Integrative analysis of 351 multi-omics datasets from liver tissues of 5-month-old mice indicated that early-life environmental exposures significantly altered chromatin accessibility and expression of TEs in later life stage, revealing distinct exposure-specific signatures and sex-dependent responses. 6,699 TEs were identified with altered chromatin accessibility, mostly in non-coding regions, suggesting potential impact on gene regulation. Within these TEs, LINE elements were enriched in genes involved in metabolic pathways, while LTR elements, particularly the ORR1E subfamily, were predominantly associated with immune-related genes. Additionally, we identified 140 TE-gene chimeric transcripts with TE-derived novel transcription start sites, highlighting TE-contributed transcriptional plasticity. Our findings depict a comprehensive landscape of TE regulation under early-life toxicant exposures, offering insights into TEs biology and their impact on health and disease.
Somatic mosaicism is essential in human biology and disease, yet robust benchmarks are scarce. The SMaHT Consortium mixed six HapMap cell lines to create artificial somatic variants spanning 0.25% to 16.5% variant allele fractions. We developed a technology-agnostic method that builds pangenome graphs from individual assemblies to create unified benchmarking sets: > 6M single-nucleotide variants, 1.8M small insertions/deletions, 49K structural variations, and 10K mobile element insertions across autosomes, X, and mitochondrial chromosomes. We validated the variants using ultra-deep simulated reads and developed a binomial-based model to estimate coverage requirements for variant detection. Evaluating multiple callers showed CHM13 alignment improves structural variant detection and offers advantages in difficult-to-map regions compared to GRCh38. Systematic characterization showed regions with low detection rate are enriched in centromeres, satellite sequences, tandem repeats, and falsely duplicated genes. This accurate, versatile resource enables systematic evaluation of somatic variant detection technologies.
Efficient and reliable profiling methods are essential to study epigenetics. Tn5, one of the first identified prokaryotic transposases with high DNA-binding and tagmentation efficiency, is widely adopted in different genomic and epigenomic protocols for high-throughputly exploring the genome and epigenome. Based on Tn5, the Assay for Transposase-Accessible Chromatin using sequencing (ATAC-seq) and the Cleavage Under Targets and Tagmentation (CUT&Tag) were developed to measure chromatin accessibility and detect DNA-protein interactions. These methodologies can be applied to large amounts of biological samples with low-input levels, such as rare tissues, embryos, and sorted single cells. However, fast and proper processing of these epigenomic data has become a bottleneck because massive data production continues to increase quickly. Furthermore, inappropriate data analysis can generate biased or misleading conclusions. Therefore, it is essential to evaluate the performance of Tn5-based ATAC-seq and CUT&Tag data processing bioinformatics tools, many of which were developed mostly for analyzing chromatin immunoprecipitation followed by sequencing (ChIP-seq) data. Here, we conducted a comprehensive benchmarking analysis to evaluate the performance of eight popular software for processing ATAC-seq and CUT&Tag data. We compared the sensitivity, specificity, and peak width distribution for both narrow-type and broad-type peak calling. We also tested the influence of the availability of control IgG input in CUT&Tag data analysis. Finally, we evaluated the differential analysis strategies commonly used for analyzing the CUT&Tag data. Our study provided comprehensive guidance for selecting bioinformatics tools and recommended analysis strategies, which were implemented into Docker/Singularity images for streamlined data analysis.
Polyglutamine expansion in Huntingtin (HTT) causes its aggregation and progressive loss of striatal neurons in Huntington’s disease (HD). HD is a mostly adult-onset neurodegenerative disorder with no disease-modifying therapies. Here we found that human striatal aging is associated with a global upregulation of genes involved in translation, including the translation and proteostasis regulator PPP1R15B (R15B). We used the R15B inhibitor Raphin1 to investigate if the age-associated changes could modify HD pathology. R15B inhibition rescued early learning and late motor deficits in HDYAC128 mice. In striatal medium spiny neurons directly reprogrammed from fibroblasts of symptomatic HD patients (HD-MSNs), Raphin1 reduced the formation of mutant HTT aggregates and neuronal death. Genetic knockdown of R15B also protected HD-MSNs from neurodegeneration whereas its overexpression exacerbated disease phenotypes. Moreover, both human striatum and reprogrammed MSNs exhibited age-dependent decline of miR-196a, a microRNA that directly targets non-conserved sites in human R15B 3’UTR and overexpressing miR-196a lowered mutant HTT aggregation. This work identifies age-dependent alterations in miR-196a and its target R15B and demonstrates the therapeutic potential of reversing these changes in diverse models and readouts of HD. We propose miR-196a and R15B as disease-modifying targets in HD.
Psychostimulant methamphetamine (METH) is neurotoxic to the brain and, therefore, its misuse leads to neurological and psychiatric disorders. The gene regulatory network (GRN) response to neurotoxic METH binge remains unclear in most brain regions. Here we examined the effects of binge METH on the GRN in the nucleus accumbens, dentate gyrus, Ammon’s horn, and subventricular zone in male rats. At 24 h after METH, ~16% of genes displayed altered expression and over a quarter of previously open chromatin regions - parts of the genome where genes are typically active - showed shifts in their accessibility. Intriguingly, most changes were unique to each area studied, and independent regulation between transcriptome and chromatin accessibility was observed. Unexpectedly, METH differentially impacted gene activity and chromatin accessibility within the dentate gyrus and Ammon’s horn. Around 70% of the affected chromatin-accessible regions in the rat brain have conserved DNA sequences in the human genome. These regions frequently act as enhancers, ramping up the activity of nearby genes, and contain mutations linked to various neurological conditions. By sketching out the gene regulatory networks associated with binge METH in specific brain regions, our study offers fresh insights into how METH can trigger profound, region-specific molecular shifts.
Background Methamphetamine (METH) is a highly addictive central nervous system stimulant. Chronic use of METH is associated with multiple neurological and psychiatric disorders. An overdose of METH can cause brain damage and even death. Mounting evidence indicates that epigenetic changes and functional impairment in the brain occur due to addictive drug exposures. However, the responses of different brain regions to a METH overdose remain unclear. Results We investigated the transcriptomic and epigenetic responses to a METH overdose in four regions of the rat brain, including the nucleus accumbens, dentate gyrus, Ammon’s horn, and subventricular zone. We found that 24 hours after METH overdose, 15.6% of genes showed changes in expression and 27.6% of open chromatin regions exhibited altered chromatin accessibility in all four rat brain regions. Interestingly, only a few of those differentially expressed genes and differentially accessible regions were affected simultaneously. Among four rat brain regions analyzed, 149 transcription factors and 31 epigenetic factors were significantly affected by METH overdose. METH overdose also resulted in opposite-direction changes in regulation patterns of both gene and chromatin accessibility between the dentate gyrus and Ammon’s horn. Approximately 70% of chromatin-accessible regions with METH-induced alterations in the rat brain are conserved at the sequence level in the human genome, and they are highly enriched in neurological processes. Many of these conserved regions are active brain-specific enhancers and harbor SNPs associated with human neurological functions and diseases. Conclusion Our results indicate strong region-specific transcriptomic and epigenetic responses to a METH overdose in distinct rat brain regions. We describe the conservation of region-specific gene regulatory networks associated with METH overdose. Overall, our study provides clues toward a better understanding of the molecular responses to METH overdose in the human brain.
BTB domain And CNC Homolog 2 (Bach2) is a transcription repressor that actively participates in T and B lymphocyte development, but it is unknown if Bach2 is also involved in the development of innate immune cells, such as natural killer (NK) cells. Here, we followed the expression of Bach2 during murine NK cell development, finding that it peaked in immature CD27+CD11b+ cells and decreased upon further maturation. Bach2 showed an organ and tissue-specific expression pattern in NK cells. Bach2 expression positively correlated with the expression of transcription factor TCF1 and negatively correlated with genes encoding NK effector molecules and those involved in the cell cycle. Lack of Bach2 expression caused changes in chromatin accessibility of corresponding genes. In the end, Bach2 deficiency resulted in increased proportions of terminally differentiated NK cells with increased production of granzymes and cytokines. NK cell-mediated control of tumor metastasis was also augmented in the absence of Bach2. Therefore, Bach2 is a key checkpoint protein regulating NK terminal maturation.
Assay for transposase-accessible chromatin with high-throughput sequencing (ATAC-seq) is a technique widely used to investigate genome-wide chromatin accessibility. The recently published Omni-ATAC-seq protocol substantially improves the signal/noise ratio and reduces the input cell number. High-quality data are critical to ensure accurate analysis. Several tools have been developed for assessing sequencing quality and insertion size distribution for ATAC-seq data; however, key quality control (QC) metrics have not yet been established to accurately determine the quality of ATAC-seq data. Here, we optimized the analysis strategy for ATAC-seq and defined a series of QC metrics for ATAC-seq data, including reads under peak ratio (RUPr), background (BG), promoter enrichment (ProEn), subsampling enrichment (SubEn), and other measurements. We incorporated these QC tests into our recently developed ATAC-seq Integrative Analysis Package (AIAP) to provide a complete ATAC-seq analysis system, including quality assurance, improved peak calling, and downstream differential analysis. We demonstrated a significant improvement of sensitivity (20%-60%) in both peak calling and differential analysis by processing paired-end ATAC-seq datasets using AIAP. AIAP is compiled into Docker/Singularity, and it can be executed by one command line to generate a comprehensive QC report. We used ENCODE ATAC-seq data to benchmark and generate QC recommendations, and developed qATACViewer for the user-friendly interaction with the QC report. The software, source code, and documentation of AIAP are freely available at https://github.com/Zhang-lab/ATAC-seq_QC_analysis.