Structural variants (SVs) contribute substantially to genomic variation and disease, but detecting somatic SVs (sSVs) remains difficult due to reference bias, mosaicism, and enrichment in repetitive regions. Linear reference genomes, like GRCh38 and CHM13, do not fully capture individual genomic structure, which can obscure true somatic variation. Donor-specific assemblies (DSAs) generated from the same genome where sSVs are being assayed provide a personalized alternative, yet their performance for sSV detection has not been systematically assessed. As part of the Somatic Mosaicism across Human Tissues (SMaHT) Network, we benchmark a DSA for sSV discovery in the COLO829 melanoma cell line with a matched normal sample from the same individual. We compare sSV detection across GRCh38, CHM13, and the COLO829BL_DSA using three different sSV callers (Delly, Severus, and Sniffles2) and sequence data from multiple long-read platforms. The COLO829BL_DSA identifies 1.8-fold more manually validated sSVs than linear references, in regions both shared with GRCh38 and CHM13 and unique to the COLO829BL_DSA. Variants detected only with the COLO829BL_DSA are often found in satellite and other repeat-rich regions that are difficult to resolve using standard references. In addition, several COLO829BL_DSA-specific sSVs are located in genes, some of which are associated with cancer. Overall, these results underscore the utility of DSAs in improving sSV detection.
The mitochondrial genome (mtDNA), rich in repeats and prone to nuclear mitochondrial DNA segments (NUMTs), drives somatic mosaicism implicated in cancer, metabolic syndromes, and neurodegeneration, yet short-read sequencing yields incomplete catalogs, mapping artifacts, and false heteroplasmies. Here, we introduce MitoScope, a scalable long-read workflow to assemble mtDNA, perform high-fidelity variant calling, resolve heteroplasmy, and characterize NUMTs in benchmarking tissues from the Somatic Mosaicism Across Human Tissues (SMaHT) Network. MitoScope shows high sensitivity and precision, determines copy number, and uncovers low-frequency variants. We define an age- and tissue-dependent landscape of mtDNA mosaicism, including low-frequency pathogenic heteroplasmies, a bimodal heteroplasmy spectrum shaped by purifying selection, and age-accumulating deletions enriched for microhomology. Parallel profiling of NUMTs identifies high-confidence events with >2-fold more NUMTs than short-read surveys-with evidence of nonrandom trinucleotide contexts at breakpoints. These findings expose pervasive, tissue-resolved somatic mtDNA and NUMT instability with direct relevance for variant interpretation, aging, and human disease.
Human neocentromeres are functional centromeres demarcated by CENP-A nucleosomes that form ectopically at alpha satellite-free loci. How neocentromeres reshape local chromatin and which features of native centromeric chromatin are preserved are unknown. We generated gapless, haplotype-resolved assemblies of native and neocentromeres from three patient-derived cell lines. Integrating CpG methylation, CENP-A profiling, and single-molecule chromatin fiber sequencing, we reveal chromatin features that define the essential centromeric architecture reconstituted during neocentromere establishment. We find that a deletion within the satellite array encompassing the hypo-CpG methylation centromere dip regions (CDRs) led to native centromere inactivation, that neocentromeres harbor CDRs and a dichromatin architecture, recapitulating features of alpha-satellite centromeres, and that LINEs demarcate neocentromere boundaries, implicating transposable elements in restricting CENP-A domain spreading. Moreover, neocentromeric chromatin is incompatible with promoter-like chromatin states, redefining the regulatory landscape within genic regions. Finally, using haplotype-specific chromatin footprinting, we resolve CENP-A nucleosome chromatin architecture of active centromeres.
The three-dimensional (3D) architecture of the genome plays a crucial role in gene regulation and various human diseases. Short-read sequencing methods for measuring 3D genome organization are powerful, but they lack the ability to resolve individual human haplotypes or structurally complex regions. To address this, we present FiberFold, a deep learning model that combines convolutional neural networks and transformer architectures to accurately predict cell-type-specific and haplotype-specific 3D genome organization using multi-omic data from a single, long-read sequencing assay, Fiber-seq. By applying FiberFold to a cell line with allelic X-inactivation, we show that Topologically Associated Domains (TADs) are attenuated on the inactive chrX. Furthermore, FiberFold predicts significant changes to TADs surrounding a 13;X balanced translocation in a patient with a rare Mendelian disease. FiberFold showcases the power of integrating long-read epigenomic sequencing with deep learning tools to investigate fundamental chromatin biology as well as the molecular basis of human disease.
Diploid human cells contain two non-identical genomes, and differences in their regulation underlie human development and disease. We present Fiber-seq Inferred Regulatory Elements (FIRE) and show that FIRE provides a more comprehensive and quantitative snapshot of the accessible chromatin landscape across the 6 Gbp diploid human genome, overcoming previously known and unknown biases afflicting our existing regulatory element catalog. FIRE provides a comprehensive genome-wide map of haplotype-selective chromatin accessibility (HSCA), exposing novel imprinted elements that lack underlying parent-of-origin CpG methylation differences, common and rare genetic variants that disrupt gene regulatory patterns, gene regulatory modules that enable genes to escape X chromosome inactivation, and autosomal mitotically stable somatic epimutations. We find that the human leukocyte antigen (HLA) locus harbors the most HSCA in immune cells, and we resolve the specific transcription factor (TF) binding events disrupted by disease-associated variants within the HLA locus. Finally, we demonstrate that the regulatory landscape of a cell is littered with autosomal somatic epimutations that are propagated by clonal expansions to create mitotically stable and non-genetically deterministic chromatin alterations.
Resolving whether and how rare noncoding genetic variants cause Mendelian conditions remains challenging owing to the diverse mechanisms by which they cause disease. Here we demonstrate the utility of single-molecule chromatin fiber sequencing (Fiber-seq) for resolving the mechanistic basis of the Mendelian condition autosomal dominant resistance to thyrotropin (RTSH), which had previously been linked to noncoding variants within a short tandem repeat (STR) variant on chromosome 15.
Great apes have maintained a stable karyotype with few large-scale rearrangements; in contrast, gibbons have undergone a high rate of chromosomal rearrangements coincident with rapid centromere turnover. Here, we characterize fully assembled centromeres in the eastern hoolock gibbon, Hoolock leuconedys (HLE), finding a diverse group of transposable elements (TEs) that differ from the canonical alpha-satellites found across centromeres of other apes. We find that HLE centromeres contain a CpG methylation centromere dip region, providing evidence that this epigenetic feature is conserved in the absence of satellite arrays. We uncovered a variety of atypical centromeric features, including protein-coding genes and mismatched replication timing. Further, we identify duplications and deletions in HLE centromeres that distinguish them from other gibbons. Finally, we observed differentially methylated TEs, topologically associated domain boundaries, and segmental duplications at chromosomal breakpoints, and thus propose that a combination of multiple genomic attributes with propensities for chromosome instability shaped gibbon centromere evolution.
The genome is reprogrammed during development to produce diverse cell types, largely through altered expression and activity of key transcription factors. The accessibility and critical functions of epidermal cells have made them a model for connecting transcriptional events to development in a range of model systems. In Arabidopsis thaliana and many other plants, fertilization triggers differentiation of specialized epidermal seed coat cells that have a unique morphology caused by large extracellular deposits of polysaccharides. Here, we used DNase I-seq to generate regulatory landscapes of A. thaliana seeds at two critical time points in seed coat maturation (4 and 7 DPA), enriching for seed coat cells with the INTACT method. We found over 3,000 developmentally dynamic regulatory DNA elements and explored their relationship with nearby gene expression. The dynamic regulatory elements were enriched for motifs for several transcription factors families; most notably the TCP family at the earlier time point and the MYB family at the later one. To assess the extent to which the observed regulatory sites in seeds added to previously known regulatory sites in A. thaliana, we compared our data to 11 other data sets generated with 7-day-old seedlings for diverse tissues and conditions. Surprisingly, over a quarter of the regulatory, i.e. accessible, bases observed in seeds were novel. Notably, plant regulatory landscapes from different tissues, cell types, or developmental stages were more dynamic than those generated from bulk tissue in response to environmental perturbations, highlighting the importance of extending studies of regulatory DNA to single tissues and cell types during development.
The Summary: The Illumina Infinium EPIC BeadChip is a new high-throughput array for DNA methylation analysis, extending the earlier 450k array by over 400 000 new sites. Previously, a method named eFORGE was developed to provide insights into cell type-specific and cell-composition effects for 450k data. Here, we present a significantly updated and improved version of eFORGE that can analyze both EPIC and 450k array data. New features include analysis of chromatin states, transcription factor motifs and DNase I footprints, providing tools for epigenome-wide association study interpretation and epigenome editing.
The genome is reprogrammed during development to produce diverse cell types, largely through altered expression and activity of key transcription factors. The accessibility and critical functions of epidermal cells have made them a model for connecting transcriptional events to development in a range of model systems. In Arabidopsis thaliana and many other plants, fertilization triggers differentiation of specialized epidermal seed coat cells that have a unique morphology caused by large extracellular deposits of pectin. Here, we used DNase I-seq to generate regulatory landscapes of A. thaliana seeds at two critical time points in seed coat maturation, enriching for seed coat cells with the INTACT method. We found over 3000 developmentally dynamic regulatory DNA elements and explored their relationship with nearby gene expression. The dynamic regulatory elements were enriched for motifs for several transcription factors families; most notably the TCP family at the earlier time point and the MYB family at the later one. To assess the extent to which the observed regulatory sites in seeds added to previously known regulatory sites in A. thaliana , we compared our data to 11 other data sets generated with seven-day-old seedlings for diverse tissues and conditions. Surprisingly, over a quarter of the regulatory, i.e. accessible, bases observed in seeds were novel. Notably, in this comparison, development exerted a stronger effect on the plant regulatory landscape than extreme environmental perturbations, highlighting the importance of extending studies of regulatory landscapes to other tissues and cell types during development.
Transcriptional dysregulation drives cancer formation but the underlying mechanisms are still poorly understood. As a model system, we used renal cell carcinoma (RCC), the most common malignant kidney tumor which canonically activates the hypoxia-inducible transcription factor (HIF) pathway. We performed genome-wide chromatin accessibility and transcriptome profiling on paired tumor/normal samples and found that numerous transcription factors with a RCC-selective expression pattern also demonstrated evidence of HIF binding in the vicinity of their gene body. Some of these transcription factors influenced the tumor’s regulatory landscape, notably the stem cell transcription factor POU5F1 ( OCT4 ). Unexpectedly, we discovered a HIF-pathway-responsive cryptic promoter embedded within a human-specific retroviral repeat element that drives POU5F1 expression in RCC via a novel transcript. Elevat POU5F1 expression levels were correlated with advanced tumor stage and poorer overall survival in RCC patients. Thus, integrated transcriptomic and epigenomic analysis of even a small number of primary patient samples revealed remarkably convergent shared regulatory landscapes and a novel mechanism for dysregulated expression of POU5F1 in RCC.
BACKGROUND:Myocardial mass is a key determinant of cardiac muscle function and hypertrophy. Myocardial depolarization leading to cardiac muscle contraction is reflected by the amplitude and duration of the QRS complex on the electrocardiogram (ECG). Abnormal QRS amplitude or duration reflect changes in myocardial mass and conduction, and are associated with increased risk of heart failure and death. OBJECTIVES:This meta-analysis sought to gain insights into the genetic determinants of myocardial mass. METHODS:We carried out a genome-wide association meta-analysis of 4 QRS traits in up to 73,518 individuals of European ancestry, followed by extensive biological and functional assessment. RESULTS:We identified 52 genomic loci, of which 32 are novel, that are reliably associated with 1 or more QRS phenotypes at p < 1 × 10(-8). These loci are enriched in regions of open chromatin, histone modifications, and transcription factor binding, suggesting that they represent regions of the genome that are actively transcribed in the human heart. Pathway analyses provided evidence that these loci play a role in cardiac hypertrophy. We further highlighted 67 candidate genes at the identified loci that are preferentially expressed in cardiac tissue and associated with cardiac abnormalities in Drosophila melanogaster and Mus musculus. We validated the regulatory function of a novel variant in the SCN5A/SCN10A locus in vitro and in vivo. CONCLUSIONS:Taken together, our findings provide new insights into genes and biological pathways controlling myocardial mass and may help identify novel therapeutic targets.
The bulk of modern genomics research includes, in part, analyses of large data sets, such as those derived from high resolution, high-throughput experiments, that make computations challenging. The BEDOPS toolkit offers a broad spectrum of fundamental analysis capabilities to query, operate on, and compare quantitatively genomic data sets of any size and number. The toolkit facilitates the construction of complex analysis pipelines that remain efficient in both memory and time by chaining together combinations of its complementary components. The principal utilities accept raw or compressed data in a flexible format, and they provide built-in features to expedite parallel computations.
The reference human genome sequence set the stage for studies of genetic variation and its association with human disease, but epigenomic studies lack a similar reference. To address this need, the NIH Roadmap Epigenomics Consortium generated the largest collection so far of human epigenomes for primary cells and tissues. Here we describe the integrative analysis of 111 reference human epigenomes generated as part of the programme, profiled for histone modification patterns, DNA accessibility, DNA methylation and RNA expression. We establish global maps of regulatory elements, define regulatory modules of coordinated activity, and their likely activators and repressors. We show that disease- and trait-associated genetic variants are enriched in tissue-specific epigenomic marks, revealing biologically relevant cell types for diverse human traits, and providing a resource for interpreting the molecular basis of human disease. Our results demonstrate the central role of epigenomic information for understanding gene regulation, cellular differentiation and human disease.
Our understanding of gene regulation in plants is constrained by our limited knowledge of plant cis-regulatory DNA and its dynamics. We mapped DNase I hypersensitive sites (DHSs) in A. thaliana seedlings and used genomic footprinting to delineate ∼700,000 sites of in vivo transcription factor (TF) occupancy at nucleotide resolution. We show that variation associated with 72 diverse quantitative phenotypes localizes within DHSs. TF footprints encode an extensive cis-regulatory lexicon subject to recent evolutionary pressures, and widespread TF binding within exons may have shaped codon usage patterns. The architecture of A. thaliana TF regulatory networks is strikingly similar to that of animals in spite of diverged regulatory repertoires. We analyzed regulatory landscape dynamics during heat shock and photomorphogenesis, disclosing thousands of environmentally sensitive elements and enabling mapping of key TF regulatory circuits underlying these fundamental responses. Our results provide an extensive resource for the study of A. thaliana gene regulation and functional biology.
The laboratory mouse shares the majority of its protein-coding genes with humans, making it the premier model organism in biomedical research, yet the two mammals differ in significant ways. To gain greater insights into both shared and species-specific transcriptional and cellular regulatory programs in the mouse, the Mouse ENCODE Consortium has mapped transcription, DNase I hypersensitivity, transcription factor binding, chromatin modifications and replication domains throughout the mouse genome in diverse cell and tissue types. By comparing with the human genome, we not only confirm substantial conservation in the newly annotated potential functional sequences, but also find a large degree of divergence of sequences involved in transcriptional regulation, chromatin state and higher order chromatin organization. Our results illuminate the wide range of evolutionary forces acting on genes and their regulatory regions, and provide a general resource for research into mammalian biology and mechanisms of human diseases.
The basic body plan and major physiological axes have been highly conserved during mammalian evolution, yet only a small fraction of the human genome sequence appears to be subject to evolutionary constraint. To quantify cis- versus trans-acting contributions to mammalian regulatory evolution, we performed genomic DNase I footprinting of the mouse genome across 25 cell and tissue types, collectively defining ∼8.6 million transcription factor (TF) occupancy sites at nucleotide resolution. Here we show that mouse TF footprints conjointly encode a regulatory lexicon that is ∼95% similar with that derived from human TF footprints. However, only ∼20% of mouse TF footprints have human orthologues. Despite substantial turnover of the cis-regulatory landscape, nearly half of all pairwise regulatory interactions connecting mouse TF genes have been maintained in orthologous human cell types through evolutionary innovation of TF recognition sequences. Furthermore, the higher-level organization of mouse TF-to-TF connections into cellular network architectures is nearly identical with human. Our results indicate that evolutionary selection on mammalian gene regulation is targeted chiefly at the level of trans-regulatory circuitry, enabling and potentiating cis-regulatory plasticity.
The precise splicing of genes confers an enormous transcriptional complexity to the human genome. The majority of gene splicing occurs cotranscriptionally, permitting epigenetic modifications to affect splicing outcomes. Here we show that select exonic regions are demarcated within the three-dimensional structure of the human genome. We identify a subset of exons that exhibit DNase I hypersensitivity and are accompanied by 'phantom' signals in chromatin immunoprecipitation and sequencing (ChIP-seq) that result from cross-linking with proximal promoter- or enhancer-bound factors. The capture of structural features by ChIP-seq is confirmed by chromatin interaction analysis that resolves local intragenic loops that fold exons close to cognate promoters while excluding intervening intronic sequences. These interactions of exons with promoters and enhancers are enriched for alternative splicing events, an effect reflected in cell type-specific periexonic DNase I hypersensitivity patterns. Collectively, our results connect local genome topography, chromatin structure and cis-regulatory landscapes with the generation of human transcriptional complexity by cotranscriptional splicing.
Cellular-state information between generations of developing cells may be propagated via regulatory regions. We report consistent patterns of gain and loss of DNase I-hypersensitive sites (DHSs) as cells progress from embryonic stem cells (ESCs) to terminal fates. DHS patterns alone convey rich information about cell fate and lineage relationships distinct from information conveyed by gene expression. Developing cells share a proportion of their DHS landscapes with ESCs; that proportion decreases continuously in each cell type as differentiation progresses, providing a quantitative benchmark of developmental maturity. Developmentally stable DHSs densely encode binding sites for transcription factors involved in autoregulatory feedback circuits. In contrast to normal cells, cancer cells extensively reactivate silenced ESC DHSs and those from developmental programs external to the cell lineage from which the malignancy derives. Our results point to changes in regulatory DNA landscapes as quantitative indicators of cell-fate transitions, lineage relationships, and dysfunction.