Centromeres are essential chromosomal regions that ensure accurate chromosome segregation during cell division, yet their highly repetitive sequence has historically hindered their complete assembly and characterization1. Consequently, the full spectrum of centromere diversity across individuals, populations and evolutionary contexts remains largely unexplored. Here we address this gap in knowledge by assembling and characterizing 2,110 centromeres from diverse individuals representing 5 continental and 28 population groups. Using bioinformatic tools tailored for centromeres, we identify variation, including 226 centromere haplotypes and 1,870 α-satellite higher-order repeat variants. While most centromeres have a single kinetochore site, we find that around 6% have di-kinetochores, and less than 1% have tri-kinetochores, which we confirm using long-read chromatin profiling and multigenerational inheritance. We also show that kinetochore position is closely associated with the underlying sequence and structure of the centromere. To understand the nature of evolutionary change, we compared these centromeres to 5,747 centromeres assembled by the Human Pangenome Reference Consortium. We show that centromeres have a 20-fold variation in mutation rate, and a subset of centromeres has evidence of archaic hominin introgression. We validate these mutation rates in a 4-generation, 28-member family and show that the kinetochore site is the most rapidly mutating region in the centromere. We propose a model that reveals an 'arms race' between centromeric sequence and proteins, with frequent mutations within the kinetochore site that lead to changes in genetic and epigenetic landscapes and, ultimately, rapid evolution of these critically important regions.
Abstract Centromeres are essential for chromosome segregation, yet in many genomes they are composed entirely of rapidly evolving repetitive DNA, embedded in other repetitive DNA that forms pericentromeric heterochromatin. Due to the difficulties of manipulating these repeat-rich regions, how the relative size of pericentromeric repeat regions influences chromosome segregation remains an open question. Here, we take advantage of the tractable Schizosaccharomyces pombe system by combining population-level analysis, complete long-read assemblies, and engineered near-isogenic strains to test how pericentromeric repeat copy number affects chromosome biology in its native context. We find that pericentromeric dh/dg arrays on chromosome 3 vary almost tenfold in size among natural S. pombe isolates, ranging from 35 to 265 kb. We converted this natural diversity into an experimental system of nearly isogenic strains that primarily differ in pericentromere size (35 to >350 kb). We found that pericentromere size does not alter baseline growth under standard conditions. However, larger pericentromeres alter transcriptional output and sensitize cells to spindle stress. We show that this spindle-stress phenotype depends on heterochromatin: loss of the H3K9 methyltransferase Clr4 abolishes size-dependent differences, whereas artificial targeting of the Chromosomal Passenger Complex to heterochromatin partially rescues the defect. Thus, we find that larger pericentromeres act as sinks for limiting regulatory factors, weakening their effective concentration at centromeres and compromising faithful chromosome segregation under stress. These results establish that naturally occurring copy-number variation within repetitive pericentromeric DNA is not merely noise, but a functional source of variation in chromosome segregation and gene regulation. Our work provides an experimentally tractable framework for understanding how repeat expansion in centromere-proximal heterochromatin influences chromosome behavior across eukaryotes.
Abstract Long-read sequencing and assembly of human genomes have revealed extreme variation in centromeric alpha-satellite array sizes across chromosomes and individuals. However, assessing the impact of centromeric array size on centromere function remains challenging. In this study, using quantitative FISH and microscopy assays, we established a framework for the systematic functional evaluation of naturally occurring variations in centromeric array size during mitosis. By comparing 8 pairs of homologous chromosomes exhibiting 1.24 – 4.6 –fold variation in centromeric array size, we uncovered size-based differences in function. We found that between a pair of homologous chromosomes, the chromosome harboring the smaller centromeric array is more prone to chromosome missegregation. Centromere size-based differences in chromosome missegregation rates cannot be explained by differential enrichment of molecular factors like centromeric histone CENP-A, CENP-B, or most kinetochore proteins. However, the smaller centromere out of a homologous pair consistently exhibits increased cohesion fatigue, suggesting the involvement of the cohesin complex in size-based differences in centromere cohesion. Consistent with our hypothesis, complete deprotection of the cohesin complex by depletion of shugoshin-1, removes the size-based bias in centromere cohesion. These results suggest that while CENP-A–enriched core centromeres are essential for kinetochore assembly, variation in overall centromeric array size has a significant impact on centromere performance and the fidelity of chromosome segregation during mitosis.
Centromeres ensure proper chromosome segregation during cell division, yet the organization and regulation of centromeric chromatin within satellite DNA arrays remain incompletely understood. Here, we leverage the complete diploid human genome benchmark (T2T-HG002) to provide a detailed study of centromeric sequence and chromatin architecture on individual haplotypes. Using adaptive-sampling-enriched, ultra-long-read DiMeLo-seq, we achieve single-molecule chromatin profiling across all centromeres, revealing that along single chromatin fibers, CENP-A, the histone variant specifying centromere identity, forms multiple discrete subdomains within hypomethylated centromere dip regions (CDRs) that are flanked by H3K9me3-enriched heterochromatin. Despite underlying sequence variation, CDRs localize to sequence-homogeneous domains and maintain relatively balanced CENP-A dosage and aggregate length across all chromosomes and between haplotypes. Further, we show that bidirectional changes to centromeric and pericentromeric DNA methylation are accompanied by changes to centromeric chromatin architecture. In passaged cells with centromeric hypomethylation, subdomain boundaries are eroded, and adjacent CENP-A domains tend to merge and expand. Conversely, in pluripotent stem cells with centromeric hypermethylation, CDRs are fundamentally reorganized, such that discrete hypomethylated domains are frequently consolidated into broader contiguous tracts. These methylation-associated CDR restructuring events suggest that DNA methylation acts as a principal regulator of human centromere organization, with implications for understanding centromere plasticity, epigenetic inheritance, and chromosomal instability in development and disease.
Age-dependent reproductive decline has become a significant global health concern as the average maternal age at first birth increases. Fertility loss associated with reproductive aging is driven in part by alterations to ovarian composition and function, dysregulation of folliculogenesis, and increased inflammatory signaling. Our understanding of the molecular changes underlying ovarian aging has been expanded by single-cell and spatial transcriptomic studies, which identified infiltration of immune cells as a feature of ovarian aging. However, the function of these age-associated immune cells and their potential contributions to the inflammaging phenotype remain unclear. In this study, we integrate single-cell and spatial transcriptomics to define changes in the composition and intercellular signaling in the aging mouse ovary. We identify specific macrophage and T cell subpopulations that increase with age and are key sources of pro-inflammatory signaling in old ovaries. Further, we predict bidirectional signaling between these pro-inflammatory cells and granulosa cell populations that may impair follicular growth and development while promoting immune cell recruitment. These findings provide insights into the mechanisms that drive ovarian inflammaging.
Ribosome biogenesis is a conserved and highly regulated process that starts in the nucleolus, a membrane-less multi-phase organelle. Although the architecture of the nucleolus is known to change due to perturbations, how nucleolar organization is modulated during physiological processes to meet changing translational demands remains unclear. Here, we use zebrafish oogenesis as a developmental context requiring a rapid expansion of translational capacity to investigate the regulation of nucleolar architecture. We show nucleoli undergo coordinated changes in number, size, subnuclear localization, and layering throughout oogenesis. We further demonstrate that nucleoli form around extrachromosomal DNA circles that contain the rDNA locus. Notably, mouse oocytes undergo similar developmental changes in nucleolar layering and phase organization, indicating that remodeling of nucleolar condensates is a conserved feature of oogenesis. These findings reveal previously unexplored regulation of nucleolar architecture as developmental adaptations to changing biosynthetic needs.
Centromeres ensure chromosome segregation, but their chromatin organization within repetitive alpha-satellite DNA has been difficult to resolve. To address this, we generated haplotype-resolved satellite DNA annotations for the complete diploid T2T-HG002 human genome assembly, then we mapped centromere protein A (CENP-A), H3K9me3, and CpG methylation on ultra-long, adaptively sampled nanopore reads using directed methylation with long-read sequencing (DiMeLo-seq). We find that CENP-A occupies multiple discrete subdomains within hypomethylated centromere dip regions (CDRs), with constrained aggregate size and balanced CENP-A dosage between homologous chromosomes despite extensive satellite array variation. We also show that extended lymphoblastoid cell culture and induced pluripotent stem cell (iPSC) reprogramming remodel DNA methylation and alter CENP-A abundance and CDR subdomain organization. These results define a single-molecule, haplotype-resolved framework for studying human centromere plasticity, epigenetic inheritance, and chromosomal instability in development and disease.
Cancer genome sequencing is essential for understanding tumor evolution and advancing precision medicine.1 However, reference gaps and germline variants obscure detection of small and large somatic variants and methylation in repetitive regions.1-3 It is common for tumor cells to gain or lose chromosome arms due to somatic structural changes that occur inside highly repetitive satellite DNA sequences in the centromeres.4 To identify the full spectrum of somatic variants, including complex rearrangements, we construct and curate near-complete, haplotype-resolved assemblies of the most recent common ancestor of an early-passage broadly-consented hypodiploid pancreatic cancer cell line and matched normal tissues. The tumor assembly completely recapitulates all 35 tumor chromosomes observed with karyotyping, with multiple translocation-induced hybrid chromosomes. The hybrid chromosomes contain putative functional dicentric and fused centromeres, nested foldback inversions causing 14 breakpoints with a haplotype switch in a single event, and centromeric satellite tandem duplications up to 136 kbp. Direct comparison of tumor and normal assembly haplotypes uncovers >7,000 variants altering >1 Mbp of sequence in repetitive regions that have been hidden by reference gaps and germline variants. 44 % of somatic small variants change representation because they alter germline variants on GRCh38, impacting mutational signatures and kataegis/omikli clusters. Most somatic LINE insertions originate from two hypomethylated non-reference germline LINE insertions, highlighting their impact on insertion mutation burden. These assemblies demonstrate that centromeric, acrocentric, and telomeric regions conventionally excluded from analysis harbor extensive somatic and epigenetic changes. Resolving complete tumor genomes enables a deeper understanding of cancer structural plasticity and the endpoints of breakage-fusion-bridge cycles. These assembled, curated paired normal-tumor benchmarks will serve as a critical foundation for developing future algorithms to characterize the most intractable regions of cancer genomes.
The common marmoset is a New World monkey widely used to study primate evolution and human disease. We present a telomere-to-telomere (T2T) reference assembly for the species, plus three near-T2T haplotypes. These resolve previously inaccessible regions, including the centromeres, sex chromosomes, subterminal satellites, acrocentric chromosomes, and the major histocompatibility complex (MHC). We find marmoset centromeres carry dimeric alpha satellites with chromosomal specificity, flanked by inactive layers interpreted as ancestral centromere remnants. We assemble gene-poor, satellite-rich short arms of the acrocentrics and find that most can harbor rDNA and all share pseudo-homolog regions (PHRs). PHR-sharing chromosomes also share closely related centromeric satellites, consistent with a model of ongoing rDNA-facilitated recombinational exchange between heterologous chromosomes. We further identify over 500 marmoset-lineage-specific transcribed genes with previously unknown transcript models or expansions. These resources, along with a preliminary pangenome, improve the utility of the marmoset as a model organism and address gaps in primate genome evolution.
The pangenome era is producing long-read sequencing data and complete genome assemblies (1–3) at a pace that current annotation methods cannot match. Existing tools were each built for a single feature class (repeats, centromeric satellites, or genes) and falter precisely where the genome is most variable and harbours clinically important variation: the centromeres, subtelomeres, and acrocentric short arms. Here we present KaryoScope, an alignment-free method to annotate an assembly at base-pair resolution across any desired feature classes in a single pass, completing in minutes on a standard workstation. Applied to the Human Pangenome Reference Consortium Release 2 assemblies (3), KaryoScope identifies the SST1 macrosatellite as the recurrent sequence at Robertsonian translocation fusion points (4, 5), delivers the first pangenome-wide census of D4Z4 macrosatellite structural diversity at the 4q and 10q subtelomeres relevant to facioscapulohumeral muscular dystrophy (6), and reveals previously uncharacterised centromere structural polymorphism, including chromosome-specific satellite loss and megabase-scale rearrangement validated by fluorescence in situ hybridization. A pre-built KaryoScope database for the human genome is distributed alongside the tool, and additional databases can be built for any reference genome or annotation source. Together, these capabilities bring the most variable regions of the genome within reach for comparative, clinical, and pangenome-scale analysis. KaryoScope is available at https://github.com/barthel-lab/KaryoScope .
Eukaryotic genomes frequently contain large arrays of tandem repeats, called satellite DNA. While some satellite DNAs participate in centromere function, others do not. For example, Human Satellite 3 (HSat3) forms the largest satellite DNA arrays in the human genome, but these multi-megabase regions were almost fully excluded from genome assemblies until recently, and their potential functions remain understudied and largely unknown. To address this, we performed a systematic screen for HSat3 binding proteins. Our work revealed that HSat3 contains millions of copies of transcription factor (TF) motifs bound by over a dozen TFs from various signaling pathways, including the growth-regulating transcription effector family TEAD1-4 from the Hippo pathway. Imaging experiments show that TEAD recruits the co-activator YAP to HSat3 regions in a cell-state specific manner. Using synthetic reporter assays, targeted repression of HSat3, inducible degradation of YAP, and super-resolution microscopy, we show that HSat3 arrays can localize YAP/TEAD inside the nucleolus, enhancing RNA Polymerase I activity. Beyond discovering a direct relationship between the Hippo pathway and ribosomal DNA regulation, this work demonstrates that satellite DNA can encode multiple transcription factor binding motifs, defining an important functional role for these enormous genomic elements.
Robertsonian chromosomes are a type of variant chromosome that is commonly found in nature. Present in 1 in 800 humans, these chromosomes can underlie infertility, trisomies and increased cancer incidence1-5. They have been recognized cytogenetically for more than a century6, yet their origins have remained unknown. Here we describe complete assemblies of three human Robertsonian chromosomes. We identified a common breakpoint in SST1, a macrosatellite DNA located on chromosomes 13, 14 and 21, which commonly undergo Robertsonian translocation. SST1 is contained within a larger shared homology domain7 that is inverted on chromosome 14, which enables a meiotic crossover event that fuses the long arms of two chromosomes. Robertsonian chromosomes have two centromeric DNA arrays and have lost all ribosomal DNA. In two cases, we find that only one of the two centromeric arrays is active. In the third case, both arrays can be active but owing to their proximity, they are often encompassed by a single outer kinetochore. Thus a combination of array proximity and epigenetic changes in centromeres facilitates the stable propagation of Robertsonian chromosomes. Investigation of the assembled genomes of chimpanzee and bonobo highlights that the inversion on chromosome 14 is unique to the human genome. Resolving the structural and epigenetic features of human Robertsonian chromosomes at a molecular level provides a foundation for a broader understanding of the molecular mechanisms of structural variation and chromosome evolution.
We synthesize current evidence that granulosa cells possess unique innate immune signaling capabilities. We suggest the novel concept that this serves as a quality control surveillance mechanism by integrating signals from the oocyte and ovarian microenvironment to prevent poor-quality follicles from producing gametes that contribute to the next generation.
Ribosomal RNA (rRNA) genes are organized in tandem arrays known as ribosomal DNA (rDNA) on multiple chromosomes in Hominidae genomes. We measured copy number and transcriptional activity status of rRNA gene arrays across multiple individual genomes, revealing an identifiable fingerprint of rDNA copy number and activity. In some cases, entire arrays were transcriptionally silent, characterized by high DNA methylation across the rRNA gene, inaccessible chromatin, and the absence of transcription factors and transcripts. Silent arrays showed reduced association with the nucleolus and decreased interchromosomal interactions, consistent with the model that nucleolar organizer function depends on transcriptional activity. Removing rDNA methylation activated silent arrays. Array activity status remained stable through induced pluripotent stem cell reprogramming and differentiation into cerebral and intestinal organoids. Haplotype tracing in two unrelated family trios showed paternal transmission of silent arrays. We propose that the epigenetic state buffers rRNA gene dosage, specifies nucleolar organizer function, and can propagate transgenerationally.
The most dynamic and repetitive regions of great ape genomes have traditionally been excluded from comparative studies 1–3 . Consequently, our understanding of the evolution of our species is incomplete. Here we present haplotype-resolved reference genomes and comparative analyses of six ape species: chimpanzee, bonobo, gorilla, Bornean orangutan, Sumatran orangutan and siamang. We achieve chromosome-level contiguity with substantial sequence accuracy (<1 error in 2.7 megabases) and completely sequence 215 gapless chromosomes telomere-to-telomere. We resolve challenging regions, such as the major histocompatibility complex and immunoglobulin loci, to provide in-depth evolutionary insights. Comparative analyses enabled investigations of the evolution and diversity of regions previously uncharacterized or incompletely studied without bias from mapping to the human reference genome. Such regions include newly minted gene families in lineage-specific segmental duplications, centromeric DNA, acrocentric chromosomes and subterminal heterochromatin. This resource serves as a comprehensive baseline for future evolutionary studies of humans and our closest living ape relatives.
Pedigree analysis remains the gold standard for rare disease diagnostics, yet whole genome sequencing studies typically omit critical regions like centromeres, telomeres, and acrocentric chromosome p-arms. Here, we present telomere-to-telomere (T2T) reference genomes for four self-identified African American individuals of admixed ancestry spanning three generations. Our parent-of-origin assigned, chromosome-level assemblies revealed precise meiotic recombination breakpoints in previously inaccessible regions, including recombination events across acrocentric and subtelomeric sequences. Centromeric regions were highly stable, with multi-megabase arrays inherited intact across three generations, while the position of kinetochore assembly sites remained consistent and predominantly associated with the p-arm proximal region. The relative lengths of telomeres on individual chromosomes were maintained across generations. Using a targeted rDNA assembly approach, we reconstructed a complete megabase-scale ribosomal DNA (rDNA) array corresponding to the paternal chromosome 14. This openly available pedigree provides a benchmark dataset for studying recombination and genetic and epigenetic variation across the complete genome.
The short arms of human acrocentric chromosomes are characterized by nucleolar organizer regions essential for ribosome biogenesis, but their highly repetitive nature has hindered genomic analysis. Leveraging the recently completed genomes of all major ape lineages, we identified recurrent features of their acrocentrics, including enriched repeat classes, centromere repositioning by whole-arm inversion, interchromosomal sequence exchange, and birth-and-death evolution of multiple gene families. Together, these processes have enabled the repeated amplification and diversification of the FRG1 gene family over 25 million years of ape evolution, and, in gorilla, the formation and amplification of a novel IGSF3-GGT fusion gene under positive selection. Similar evolutionary events also explain the distribution of segmental duplications and heterochromatin in the modern human genome, predisposing it to karyotypic abnormalities such as Robertsonian translocations. Our findings highlight acrocentric chromosomes as key drivers of evolution in the great apes, with implications for speciation, adaptation, and clinical genomics.
Nucleolar morphology is a well-established indicator of ribosome biogenesis activity that has served as the foundation of many screens investigating ribosome production. Missing from this field of study is a broad-scale investigation of the regulation of ribosomal DNA morphology, despite the essential role of rRNA gene transcription in modulating ribosome output. We hypothesized that the morphology of rDNA arrays reflects ribosome biogenesis activity. We established GapR-GFP, a prokaryotic DNA-binding protein that recognizes transcriptionally-induced overtwisted DNA, as a live visual fluorescent marker for quantitative analysis of rDNA organization in Schizosaccharomyces pombe. We found that the morphology—which we refer to as spatial organization—of the rDNA arrays is dynamic throughout the cell cycle, under glucose starvation, RNA pol I inhibition, and TOR activation. Screening the haploid S. pombe Bioneer deletion collection for spatial organization phenotypes revealed large ribosomal protein (RPL) gene deletions that alter rDNA organization. Further work revealed RPL gene deletion mutants with altered rDNA organization also demonstrate resistance to the TOR inhibitor Torin1. A genetic analysis of signaling pathways essential for this resistance phenotype implicated many factors including a conserved MAPK, Pmk1, previously linked to extracellular stress responses. We propose RPL gene deletion triggers altered rDNA morphology due to compensatory changes in ribosome biogenesis via multiple signaling pathways, and we further suggest compensatory responses may contribute to human diseases such as ribosomopathies. Altogether, GapR-GFP is a powerful tool for live visual reporting on rDNA morphology under myriad conditions.
ABSTRACT Robertsonian chromosomes form by fusion of two chromosomes that have centromeres located near their ends, known as acrocentric or telocentric chromosomes. This fusion creates a new metacentric chromosome and is a major mechanism of karyotype evolution and speciation. Robertsonian chromosomes are common in nature and were first described in grasshoppers by the zoologist W. R. B. Robertson more than 100 years ago. They have since been observed in many species, including catfish, sheep, butterflies, bats, bovids, rodents and humans, and are the most common chromosomal change in mammals. Robertsonian translocations are particularly rampant in the house mouse, Mus musculus domesticus, where they exhibit meiotic drive and create reproductive isolation. Recent progress has been made in understanding how Robertsonian chromosomes form in the human genome, highlighting some of the fundamental principles of how and why these types of fusion events occur so frequently. Consequences of these fusions include infertility and Down's syndrome. In this Hypothesis, I postulate that the conditions that allow these fusions to form are threefold: (1) sequence homology on non-homologous chromosomes, often in the form of repetitive DNA; (2) recombination initiation during meiosis; and (3) physical proximity of the homologous sequences in three-dimensional space. This Hypothesis highlights the latest progress in understanding human Robertsonian translocations within the context of the broader literature on Robertsonian chromosomes.
Ribosomal RNA (rRNA) genes exist in multiple copies arranged in tandem arrays known as ribosomal DNA (rDNA). The total number of gene copies is variable, and the mechanisms buffering this copy number variation remain unresolved. We surveyed the number, distribution, and activity of rDNA arrays at the level of individual chromosomes across multiple human and primate genomes. Each individual possessed a unique fingerprint of copy number distribution and activity of rDNA arrays. In some cases, entire rDNA arrays were transcriptionally silent. Silent rDNA arrays showed reduced association with the nucleolus and decreased interchromosomal interactions, indicating that the nucleolar organizer function of rDNA depends on transcriptional activity. Methyl-sequencing of flow-sorted chromosomes, combined with long read sequencing, showed epigenetic modification of rDNA promoter and coding region by DNA methylation. Silent arrays were in a closed chromatin state, as indicated by the accessibility profiles derived from Fiber-seq. Removing DNA methylation restored the transcriptional activity of silent arrays. Array activity status remained stable through the iPS cell re-programming. Family trio analysis demonstrated that the inactive rDNA haplotype can be traced to one of the parental genomes, suggesting that the epigenetic state of rDNA arrays may be heritable. We propose that the dosage of rRNA genes is epigenetically regulated by DNA methylation, and these methylation patterns specify nucleolar organizer function and can propagate transgenerationally.