
Super-enhancer (SE) hubs have been proposed to coordinate gene expression through 3D genome organization and transcriptional condensates. Using multiplexed imaging, we mapped hundreds of SEs in thousands of mouse embryonic stem cells and also paired SE position with nascent transcription measurements. We found that most SEs are spatially isolated, with multiway SE hubs occurring in only a small fraction of cells. Rare hubs were largely promiscuous, cooperative aggregates shaped by genomic proximity, nuclear speckle association, and general transcription machinery occupancy. Perturbing cohesin, CTCF, BET proteins, or RNA polymerase II showed that normal genome organization generally suppresses SE clustering. Combined RNA and DNA imaging demonstrated that SE hubs were neither necessary nor sufficient for transcriptional bursting, although larger hubs weakly increased burst probability. These results challenge models in which SE hubs are a dominant mechanism of enhancer function and instead suggest rare transcriptional crosstalk.
This month, Cell and Cell Genomics will publish multiple papers from the Telomere-to-Telomere (T2T) Consortium. We spoke to some of the authors involved in the studies and asked them to tell us about their work.
Prime editing could resolve gene variant function at scale, but the editing machinery needs to be delivered efficiently and reproducibly. Langley, Baudrier, et al.1 show that timing of editor delivery is rate limiting and introduce PRIME-VLP, a repeated dosing strategy with engineered virus-like particles that improves editing efficiency and screening performance.
Diverse and globally representative datasets are essential to genomic science and medicine. Here, we analyzed population descriptor metadata from RNA sequencing (RNA-seq) studies in two major public repositories: the Sequence Read Archive (SRA) and the Database of Genotypes and Phenotypes. We examined geographic and economic characteristics of institutions depositing the data and compared SRA-deposited descriptors to empirical estimates of genetic ancestry and to those reported in publications, analyzing trends over time. We found that 55% of RNA-seq samples were deposited by United States (US) institutions and 90% by institutions in high-income countries. Only 3% of SRA samples were associated with population descriptors, and among those with US Census terms, 69% were labeled as White. Among samples with continental descriptors, 56% were labeled as European. Our analyses emphasize widespread bias in the composition of public RNA-seq datasets and, more generally, a lack of consistent and careful reporting of population descriptors needing urgent improvement.
Flowering plants (angiosperms) exhibit extraordinary species diversity, ∼200-fold variation in genome size, and relatively compact coding regions, presenting both a unique challenge and opportunity for DNA language models. Here, we introduce PlantCAD2, an extended-context, plant-specific DNA language model with single-nucleotide resolution, pre-trained on 65 angiosperm genomes, together with a series of public benchmarks for evaluation. Comprehensive zero-shot testing shows that PlantCAD2 (676 million parameters) efficiently captures evolutionary conservation, surpassing the 7-billion-parameter Evo2 in 10 of 12 tasks. With parameter-efficient fine-tuning, PlantCAD2 outperforms the 1-billion-parameter AgroNT across seven cross-species tasks including chromatin accessible region, gene expression, and protein translation. Its 8,192-bp context window substantially improves accessible chromatin prediction in large genomes such as maize (area under the precision-recall curve [AUPRC] increasing from 0.587 to 0.711), underscoring the importance of long-range context for modeling distal regulation. These results establish PlantCAD2 as a powerful and versatile foundation model for plant genome annotation and interpretation across diverse species.
We report a complete rodent telomere-to-telomere genome assembly from the brown rat, Rattus norvegicus. Annotation was enriched with multi-tissue long-read RNA sequencing and uncovered numerous novel genes. Assembly of both sex chromosomes reveals the absence of gene coding in the presumed pseudo-autosomal regions and the presence of centromeric satellite repeats on distal chromosome Y (chrY). We provide evidence of meiotic conjunction between Xp and Yq. The genome assembly reveals several expanded autosomal regions enriched for testis-expressed genes. Finally, we have generated a pangenome from recent high-quality assemblies of 8 distinct inbred rat strain genomes. This allows the strain-specific distribution of structural variation to be examined. Non-allelic homologous recombination has produced multiple copies of several genes with evidence of transcription from duplicated copies.
We present telomere-to-telomere genome assemblies of a Thoroughbred horse and a donkey derived from their mule offspring. Now adopted and annotated by NCBI as reference genomes, these assemblies resolve previously inaccessible regions, including satellite arrays, duplications, and telomeres. Equids are known to exhibit an uncoupling between satellite DNA and centromeric function. The completeness of these assemblies enabled annotation of both satellite-based and satellite-free centromeres, as well as non-centromeric satellite loci, revealing notable centromeric plasticity. They also allowed detailed characterization of the variable binding domains of CENP-A-the epigenetic determinant of centromere identity-and CENP-B, whose association with CENP-A, previously considered typical based on a few model organisms, is absent in equids. Comparative analyses of satellite repeats and centromere positions provide new insights into the accelerated karyotypic reshuffling in equid evolution. These assemblies represent foundational resources for equine genomics and support ongoing initiatives such as the Equine Pangenome Project.
Recent developments have enabled the automated assembly of vertebrate chromosomes from telomere to telomere. However, for long, highly similar repeats, genome assemblers may leave tangles in the assembly graph and gaps in the assembly. In recently published genomes, such gaps are closed by manual graph curation, a process that is labor intensive, error prone, and sometimes infeasible. Consequently, important genomic regions may be misassembled or omitted. Here, we present the trivial tangle traverser (TTT) algorithm that finds optimized resolutions of assembly graph tangles. TTT uses depth of coverage and read-to-graph alignment information in a two-stage process to estimate sequence multiplicities and identify traversals that are consistent with the underlying data. We evaluate TTT traversals on the HG002 human reference genome, compare TTT with a state-of-the-art assembler on the giraffe T2T assembly, and demonstrate its use to characterize a previously unassembled amplified p21-activated serine/threonine kinase 3-like (PAK3L) gene array in the zebra finch genome.
Centromeres ensure chromosome segregation, but their chromatin organization within repetitive alpha-satellite DNA has been difficult to resolve. To address this, we generated haplotype-resolved satellite DNA annotations for the complete diploid T2T-HG002 human genome assembly, then we mapped centromere protein A (CENP-A), H3K9me3, and CpG methylation on ultra-long, adaptively sampled nanopore reads using directed methylation with long-read sequencing (DiMeLo-seq). We find that CENP-A occupies multiple discrete subdomains within hypomethylated centromere dip regions (CDRs), with constrained aggregate size and balanced CENP-A dosage between homologous chromosomes despite extensive satellite array variation. We also show that extended lymphoblastoid cell culture and induced pluripotent stem cell (iPSC) reprogramming remodel DNA methylation and alter CENP-A abundance and CDR subdomain organization. These results define a single-molecule, haplotype-resolved framework for studying human centromere plasticity, epigenetic inheritance, and chromosomal instability in development and disease.
Genome-wide association studies identify cancer susceptibility loci, but downstream protein mechanisms remain incompletely defined. We integrate polygenic risk scores (PRSs) for 21 cancers with 4,955 plasma proteins measured in cancer-free Atherosclerosis Risk in Communities (ARIC) participants to prioritize cancer-related proteins and protein networks. The protein quantitative trait score (pQTS) approach assesses associations between cancer PRS and individual protein levels, while ARCHIE partitions cancer risk variants into trans-regulated protein-network components using sparse canonical correlation analysis. Across cancers, pQTS identifies 90 protein associations, including 53 distal trans associations, and ARCHIE identifies 19 components spanning 433 proteins. Downstream analyses connect prioritized proteins to cancer driver genes, somatic alterations, immune cell populations, CRISPR dependency, and cancer-relevant pathways. Cervical cancer and basal cell carcinoma illustrate immune, human papillomavirus (HPV)-related, pigmentation, and inflammatory mechanisms. These findings show that PRS-proteome integration can reveal circulating protein networks underlying inherited cancer susceptibility.
Large differences in gene dosage are usually associated with opposite phenotypic effects but show a bias toward one direction in aggregate genome wide. Milind et al. suggest that this is explained by differences in regulatory mechanisms by which genes influence phenotypes and by the increased selective pressures acting on a subset of genes.
Rare variant association analyses are typically performed at the single-gene level, overlooking the molecular interactions that organize cellular systems. In this issue, Nazeen et al. introduce NERINE, a probabilistic rare variant burden test that integrates gene and protein networks to improve statistical power and biological interpretability.
Genome-wide association studies (GWASs) have shown that disease-associated variants are concentrated in candidate regulatory elements (cREs) from disease-relevant cell types. Here, we introduce cell-type fine-mapping (CT-FM) and CT-FM-SNP, probabilistic methods that account for cRE sharing across cell types to infer independent causal cell-type sets for complex traits and candidate causal variants. Applying CT-FM to 63 GWASs using 924 cRE annotations, we inferred 79 independent cell-type sets explaining 39.0% ± 1.8% of trait SNP heritability and identified 14 traits with multiple independent cellular mechanisms, including height, schizophrenia, and autoimmune diseases. Applying CT-FM-SNP to 39 UK Biobank traits, we assigned high-confidence causal cell types to 3,091 candidate non-coding variant-trait pairs. Most variants appeared to act through a single cell-type set, whereas pleiotropic variants often acted through different cell types depending on the phenotype context. Together, CT-FM and CT-FM-SNP provide a framework for dissecting the cellular architecture of complex traits.
Autism spectrum disorder (ASD) is a heterogeneous neurodevelopmental condition. Studies of postmortem ASD brain tissue have revealed convergent molecular changes across the cortex. Whether these features are reflected in cell-type-specific epigenetic signatures is unknown. Here, we present a single-cell analysis of DNA methylation (DNAm) coupled with transcriptomics in ASD. Using snmCT-seq, we profiled DNAm and transcript levels from over 60,000 nuclei derived from the prefrontal cortex of 49 donors. We identified over 30,000 differentially methylated regions (DMRs) in ASD that were enriched in promoters and cell-type-specific regulatory elements active across the lifespan. ASD-related methylation changes were uncorrelated with transcript levels and were small in magnitude compared with age-associated effects. Age-DMRs were concentrated in excitatory neurons and revealed distinct roles for CG and non-CG methylation. Age-varying methylation signatures of ASD identified neuron projection development as a key process perturbed in ASD, highlighting the heterogeneous impact of ASD across the lifespan.
Genome-wide association studies identify single-nucleotide polymorphisms (SNPs) associated with disease in a population but do not reliably account for individual environmental effects, despite evidence that environment mediates SNP functional regulatory capacity. Body mass index (BMI) is associated with physiologic processes across disorders but hasn’t been modeled as an environment for disease-associated SNPs. We use an interaction approach to identify SNPs that contextually regulate gene expression across the BMI spectrum, called BMI-dynamic expression quantitative trait loci (BMI-eQTL). We found BMI-eQTL across tissues, including brain and gut, while the main effects of BMI were confined to endocrine tissues. We demonstrate that cell type, putative enhancers, and/or inflammatory cytokines underlie BMI-eQTL. We develop models to predict gene expression using BMI-by-SNP interactions and identify more replicating disease-associated genes than SNP-only models. While neither genetics nor BMI is sufficient as a standalone measure to capture the complexity of downstream cellular consequences, including environment helps power disease gene discovery.
With a vast corpus of findings from nearly two decades of genome-wide association studies (GWASs), many studies now focus on translating these genetic associations into biological insights at multiple scales, from proteins and cells to entire organs. This approach will help build the foundation for the next generation of treatments. In this review, we highlight key recent studies that have informed target prioritization and drug repurposing, linked genetic variants to gene regulation in cellular contexts, and uncovered the genetic architecture of organ structure and function. Nearly 25 years after the initial draft of the human genome, it is clear that genomics is driving tangible advances in therapies and opening new ways to understand multi-scale biology.
In this issue of Cell Genomics, Dr. Jiarui Ding presents his paper, "ProtoCloud: A prototypical self-explaining model for single-cell analysis." Dr. Ding is an assistant professor in Computer Science at the University of British Columbia in Vancouver, Canada. He describes the scientific journey he undertook to publish this study and the current topics his lab is exploring.
Recent advances in single-cell transcriptomics and CRISPR-based genome editing have enabled large-scale perturbation experiments with genome-wide expression readouts. Single-cell CRISPR screens offer the opportunity to move beyond correlation and estimate causal effects of genetic perturbations on gene expression at scale. These approaches promise to substantially deepen insights into cellular functions and disease mechanisms. However, interpreting statistical associations as causal effects requires additional assumptions beyond those needed for standard statistical analyses. In this minireview, we introduce key concepts and principles for causal effect estimation in trans-regulatory single-cell CRISPR studies. We describe a set of assumptions under which estimates from existing statistical methods admit a causal interpretation and provide a concise overview of these approaches. Finally, through an illustrative example, we demonstrate how violations of these assumptions can bias estimated effects.