Balanced Robertsonian translocation (ROB) is the most common chromosomal rearrangement in humans, with an estimated occurrence of 1 in 800 in newborn studies. Carriers are at increased risk of cancer and often diagnosed at fertility clinics after facing recurrent miscarriages, infertility, or aneuploid offspring. Genotyping carriers with DNA sequencing has been challenging because of gaps and misrepresentation of the translocation fusion site in the human reference genome. Only recently, telomere-to-telomere (T2T) human genomes successfully revealed sequences of the acrocentric short arms, including the most common ROB fusion site. A ROB results in loss of two ribosomal DNA (rDNA) arrays and its adjacent distal sequences, including the highly conserved distal junction (DJ). Here, we present a novel method to type ROB carriers directly from short sequencing reads by estimating DJ copy number. We demonstrate that our method successfully genotypes ROBs using a reference-free approach or alignments to either T2T-CHM13v2 or GRCh38. Applying the method to a cohort of healthy newborns and family members (n=4,172) as well as the UK Biobank (n=490,416), we find candidate ROBs at a frequency consistent with the previously reported 1 in 800 incidence (0.11-0.12%). In addition to ROB carriers, we report the frequency of one DJ loss (9, 2.8-3.4%) or gain (11+, 8.4-9.3%) from the two cohorts and the 1000 Genomes Project (n=3,202), and characterize the underlying structural variation in near-T2T genome assemblies from the Human Pangenome Reference Consortium. Importantly, our method provides the first sequencing-based diagnostic for Robertsonian chromosomes and can be applied to low-coverage sequencing data, enhancing its clinical applicability and enabling new studies of structural variation on the acrocentric chromosomes.
The common marmoset is a New World monkey widely used to study primate evolution and human disease. We present a telomere-to-telomere (T2T) reference assembly for the species, plus three near-T2T haplotypes. These resolve previously inaccessible regions, including the centromeres, sex chromosomes, subterminal satellites, acrocentric chromosomes, and the major histocompatibility complex (MHC). We find marmoset centromeres carry dimeric alpha satellites with chromosomal specificity, flanked by inactive layers interpreted as ancestral centromere remnants. We assemble gene-poor, satellite-rich short arms of the acrocentrics and find that most can harbor rDNA and all share pseudo-homolog regions (PHRs). PHR-sharing chromosomes also share closely related centromeric satellites, consistent with a model of ongoing rDNA-facilitated recombinational exchange between heterologous chromosomes. We further identify over 500 marmoset-lineage-specific transcribed genes with previously unknown transcript models or expansions. These resources, along with a preliminary pangenome, improve the utility of the marmoset as a model organism and address gaps in primate genome evolution.
The pangenome era is producing long-read sequencing data and complete genome assemblies (1–3) at a pace that current annotation methods cannot match. Existing tools were each built for a single feature class (repeats, centromeric satellites, or genes) and falter precisely where the genome is most variable and harbours clinically important variation: the centromeres, subtelomeres, and acrocentric short arms. Here we present KaryoScope, an alignment-free method to annotate an assembly at base-pair resolution across any desired feature classes in a single pass, completing in minutes on a standard workstation. Applied to the Human Pangenome Reference Consortium Release 2 assemblies (3), KaryoScope identifies the SST1 macrosatellite as the recurrent sequence at Robertsonian translocation fusion points (4, 5), delivers the first pangenome-wide census of D4Z4 macrosatellite structural diversity at the 4q and 10q subtelomeres relevant to facioscapulohumeral muscular dystrophy (6), and reveals previously uncharacterised centromere structural polymorphism, including chromosome-specific satellite loss and megabase-scale rearrangement validated by fluorescence in situ hybridization. A pre-built KaryoScope database for the human genome is distributed alongside the tool, and additional databases can be built for any reference genome or annotation source. Together, these capabilities bring the most variable regions of the genome within reach for comparative, clinical, and pangenome-scale analysis. KaryoScope is available at https://github.com/barthel-lab/KaryoScope .
Robertsonian chromosomes are a type of variant chromosome that is commonly found in nature. Present in 1 in 800 humans, these chromosomes can underlie infertility, trisomies and increased cancer incidence1-5. They have been recognized cytogenetically for more than a century6, yet their origins have remained unknown. Here we describe complete assemblies of three human Robertsonian chromosomes. We identified a common breakpoint in SST1, a macrosatellite DNA located on chromosomes 13, 14 and 21, which commonly undergo Robertsonian translocation. SST1 is contained within a larger shared homology domain7 that is inverted on chromosome 14, which enables a meiotic crossover event that fuses the long arms of two chromosomes. Robertsonian chromosomes have two centromeric DNA arrays and have lost all ribosomal DNA. In two cases, we find that only one of the two centromeric arrays is active. In the third case, both arrays can be active but owing to their proximity, they are often encompassed by a single outer kinetochore. Thus a combination of array proximity and epigenetic changes in centromeres facilitates the stable propagation of Robertsonian chromosomes. Investigation of the assembled genomes of chimpanzee and bonobo highlights that the inversion on chromosome 14 is unique to the human genome. Resolving the structural and epigenetic features of human Robertsonian chromosomes at a molecular level provides a foundation for a broader understanding of the molecular mechanisms of structural variation and chromosome evolution.
The most dynamic and repetitive regions of great ape genomes have traditionally been excluded from comparative studies 1–3 . Consequently, our understanding of the evolution of our species is incomplete. Here we present haplotype-resolved reference genomes and comparative analyses of six ape species: chimpanzee, bonobo, gorilla, Bornean orangutan, Sumatran orangutan and siamang. We achieve chromosome-level contiguity with substantial sequence accuracy (<1 error in 2.7 megabases) and completely sequence 215 gapless chromosomes telomere-to-telomere. We resolve challenging regions, such as the major histocompatibility complex and immunoglobulin loci, to provide in-depth evolutionary insights. Comparative analyses enabled investigations of the evolution and diversity of regions previously uncharacterized or incompletely studied without bias from mapping to the human reference genome. Such regions include newly minted gene families in lineage-specific segmental duplications, centromeric DNA, acrocentric chromosomes and subterminal heterochromatin. This resource serves as a comprehensive baseline for future evolutionary studies of humans and our closest living ape relatives.
The short arms of human acrocentric chromosomes are characterized by nucleolar organizer regions essential for ribosome biogenesis, but their highly repetitive nature has hindered genomic analysis. Leveraging the recently completed genomes of all major ape lineages, we identified recurrent features of their acrocentrics, including enriched repeat classes, centromere repositioning by whole-arm inversion, interchromosomal sequence exchange, and birth-and-death evolution of multiple gene families. Together, these processes have enabled the repeated amplification and diversification of the FRG1 gene family over 25 million years of ape evolution, and, in gorilla, the formation and amplification of a novel IGSF3-GGT fusion gene under positive selection. Similar evolutionary events also explain the distribution of segmental duplications and heterochromatin in the modern human genome, predisposing it to karyotypic abnormalities such as Robertsonian translocations. Our findings highlight acrocentric chromosomes as key drivers of evolution in the great apes, with implications for speciation, adaptation, and clinical genomics.
Ribosomal RNA (rRNA) genes exist in multiple copies arranged in tandem arrays known as ribosomal DNA (rDNA). The total number of gene copies is variable, and the mechanisms buffering this copy number variation remain unresolved. We surveyed the number, distribution, and activity of rDNA arrays at the level of individual chromosomes across multiple human and primate genomes. Each individual possessed a unique fingerprint of copy number distribution and activity of rDNA arrays. In some cases, entire rDNA arrays were transcriptionally silent. Silent rDNA arrays showed reduced association with the nucleolus and decreased interchromosomal interactions, indicating that the nucleolar organizer function of rDNA depends on transcriptional activity. Methyl-sequencing of flow-sorted chromosomes, combined with long read sequencing, showed epigenetic modification of rDNA promoter and coding region by DNA methylation. Silent arrays were in a closed chromatin state, as indicated by the accessibility profiles derived from Fiber-seq. Removing DNA methylation restored the transcriptional activity of silent arrays. Array activity status remained stable through the iPS cell re-programming. Family trio analysis demonstrated that the inactive rDNA haplotype can be traced to one of the parental genomes, suggesting that the epigenetic state of rDNA arrays may be heritable. We propose that the dosage of rRNA genes is epigenetically regulated by DNA methylation, and these methylation patterns specify nucleolar organizer function and can propagate transgenerationally.
We present haplotype-resolved reference genomes and comparative analyses of six ape species, namely: chimpanzee, bonobo, gorilla, Bornean orangutan, Sumatran orangutan, and siamang. We achieve chromosome-level contiguity with unparalleled sequence accuracy (<1 error in 500,000 base pairs), completely sequencing 215 gapless chromosomes telomere-to-telomere. We resolve challenging regions, such as the major histocompatibility complex and immunoglobulin loci, providing more in-depth evolutionary insights. Comparative analyses, including human, allow us to investigate the evolution and diversity of regions previously uncharacterized or incompletely studied without bias from mapping to the human reference. This includes newly minted gene families within lineage-specific segmental duplications, centromeric DNA, acrocentric chromosomes, and subterminal heterochromatin. This resource should serve as a definitive baseline for all future evolutionary studies of humans and our closest living ape relatives.
Abstract The publication of the first complete, haploid telomere-to-telomere (T2T) human genome revealed new insights into the structure and function of the heretofore “invisible” parts of the genome including centromeres, tandem repeat arrays, and segmental duplications. Refinement of T2T processes now enables comparative analyses of complete genomes across entire clades to gain a broader understanding of the evolution of chromosome structure and function. The human T2T project involved a unique ad hoc effort involving many researchers and laboratories, serving as a model for collaborative open science. Subsequent generation and analysis of diploid, near T2T assemblies for multiple species represents a substantial increase in scale and would be daunting for any single laboratory. Efforts focused on the primate lineage continue to employ the successful open collaboration strategy and are revealing details of chromosomal evolution, species-specific gene content, and genomic adaptations, which may be general or lineage-specific features. The suborder Ruminantia has a rich history within the field of chromosome biology and includes a broad range of species at varying evolutionary distances with separation of tens of millions of years to subspecies that are still able to interbreed. We propose an open collaborative effort dubbed the “Ruminant T2T Consortium” (RT2T) to generate complete diploid assemblies for species in the Artiodactyla order, focusing on suborder Ruminantia. Here we present the initial near T2T assemblies of cattle, gaur, domestic goat, bighorn sheep, and domestic sheep, and describe the motivation, goals, and proposed comparative analyses to examine chromosomal evolution in the context of natural selection and domestication of species for use as livestock.
Apes possess two sex chromosomes-the male-specific Y and the X shared by males and females. The Y chromosome is crucial for male reproduction, with deletions linked to infertility. The X chromosome carries genes vital for reproduction and cognition. Variation in mating patterns and brain function among great apes suggests corresponding differences in their sex chromosome structure and evolution. However, due to their highly repetitive nature and incomplete reference assemblies, ape sex chromosomes have been challenging to study. Here, using the state-of-the-art experimental and computational methods developed for the telomere-to-telomere (T2T) human genome, we produced gapless, complete assemblies of the X and Y chromosomes for five great apes (chimpanzee, bonobo, gorilla, Bornean and Sumatran orangutans) and a lesser ape, the siamang gibbon. These assemblies completely resolved ampliconic, palindromic, and satellite sequences, including the entire centromeres, allowing us to untangle the intricacies of ape sex chromosome evolution. We found that, compared to the X, ape Y chromosomes vary greatly in size and have low alignability and high levels of structural rearrangements. This divergence on the Y arises from the accumulation of lineage-specific ampliconic regions and palindromes (which are shared more broadly among species on the X) and from the abundance of transposable elements and satellites (which have a lower representation on the X). Our analysis of Y chromosome genes revealed lineage-specific expansions of multi-copy gene families and signatures of purifying selection. In summary, the Y exhibits dynamic evolution, while the X is more stable. Finally, mapping short-read sequencing data from >100 great ape individuals revealed the patterns of diversity and selection on their sex chromosomes, demonstrating the utility of these reference assemblies for studies of great ape evolution. These complete sex chromosome assemblies are expected to further inform conservation genetics of nonhuman apes, all of which are endangered species.
Telomere-to-telomere (T2T) assemblies reveal new insights into the structure and function of the previously ‘invisible’ parts of the genome and allow comparative analyses of complete genomes across entire clades. We present here an open collaborative effort, termed the ‘Ruminant T2T Consortium’ (RT2T), that aims to generate complete diploid assemblies for numerous species of the Artiodactyla suborder Ruminantia to examine chromosomal evolution in the context of natural selection and domestication of species used as livestock. Here we describe an open collaborative effort termed the ‘Ruminant T2T Consortium’. It aims to generate complete diploid assemblies for many species of ruminants to examine chromosomal evolution in the context of natural selection and domestication.
Human centromeres have been traditionally very difficult to sequence and assemble owing to their repetitive nature and large size1. As a result, patterns of human centromeric variation and models for their evolution and function remain incomplete, despite centromeres being among the most rapidly mutating regions2,3. Here, using long-read sequencing, we completely sequenced and assembled all centromeres from a second human genome and compared it to the finished reference genome4,5. We find that the two sets of centromeres show at least a 4.1-fold increase in single-nucleotide variation when compared with their unique flanks and vary up to 3-fold in size. Moreover, we find that 45.8% of centromeric sequence cannot be reliably aligned using standard methods owing to the emergence of new α-satellite higher-order repeats (HORs). DNA methylation and CENP-A chromatin immunoprecipitation experiments show that 26% of the centromeres differ in their kinetochore position by >500 kb. To understand evolutionary change, we selected six chromosomes and sequenced and assembled 31 orthologous centromeres from the common chimpanzee, orangutan and macaque genomes. Comparative analyses reveal a nearly complete turnover of α-satellite HORs, with characteristic idiosyncratic changes in α-satellite HORs for each species. Phylogenetic reconstruction of human haplotypes supports limited to no recombination between the short (p) and long (q) arms across centromeres and reveals that novel α-satellite HORs share a monophyletic origin, providing a strategy to estimate the rate of saltatory amplification and mutation of human centromeric DNA.
We completely sequenced and assembled all centromeres from a second human genome and used two reference sets to benchmark genetic, epigenetic, and evolutionary variation within centromeres from a diversity panel of humans and apes. We find that centromere single-nucleotide variation can increase by up to 4.1-fold relative to other genomic regions, with the caveat that up to 45.8% of centromeric sequence, on average, cannot be reliably aligned with current methods due to the emergence of new α-satellite higher-order repeat (HOR) structures and two to threefold differences in the length of the centromeres. The extent to which this occurs differs depending on the chromosome and haplotype. Comparing the two sets of complete human centromeres, we find that eight harbor distinctly different α-satellite HOR array structures and four contain novel α-satellite HOR variants in high abundance. DNA methylation and CENP-A chromatin immunoprecipitation experiments show that 26% of the centromeres differ in their kinetochore position by at least 500 kbp—a property not readily associated with novel α-satellite HORs. To understand evolutionary change, we selected six chromosomes and sequenced and assembled 31 orthologous centromeres from the common chimpanzee, orangutan, and macaque genomes. Comparative analyses reveal nearly complete turnover of α-satellite HORs, but with idiosyncratic changes in structure characteristic to each species. Phylogenetic reconstruction of human haplotypes supports limited to no recombination between the p- and q-arms of human chromosomes and reveals that novel α-satellite HORs share a monophyletic origin, providing a strategy to estimate the rate of saltatory amplification and mutation of human centromeric DNA.
Ribosome biogenesis is a vital and highly energy-consuming cellular function occurring primarily in the nucleolus. Cancer cells have an elevated demand for ribosomes to sustain continuous proliferation. This study evaluated the impact of existing anticancer drugs on the nucleolus by screening a library of anticancer compounds for drugs that induce nucleolar stress. For a readout, a novel parameter termed ‘nucleolar normality score’ was developed that measures the ratio of the fibrillar center and granular component proteins in the nucleolus and nucleoplasm. Multiple classes of drugs were found to induce nucleolar stress, including DNA intercalators, inhibitors of mTOR/PI3K, heat shock proteins, proteasome, and cyclin-dependent kinases (CDKs). Each class of drugs induced morphologically and molecularly distinct states of nucleolar stress accompanied by changes in nucleolar biophysical properties. In-depth characterization focused on the nucleolar stress induced by inhibition of transcriptional CDKs, particularly CDK9, the main CDK that regulates RNA Pol II. Multiple CDK substrates were identified in the nucleolus, including RNA Pol I– recruiting protein Treacle, which was phosphorylated by CDK9 in vitro. These results revealed a concerted regulation of RNA Pol I and Pol II by transcriptional CDKs. Our findings exposed many classes of chemotherapy compounds that are capable of inducing nucleolar stress, and we recommend considering this in anticancer drug development.
Full text Figures and data Side by side Abstract eLife assessment eLife digest Introduction Results Discussion Materials and methods Data availability References Peer review Author response Article and author information Abstract Ribosome biogenesis is a vital and highly energy-consuming cellular function occurring primarily in the nucleolus. Cancer cells have an elevated demand for ribosomes to sustain continuous proliferation. This study evaluated the impact of existing anticancer drugs on the nucleolus by screening a library of anticancer compounds for drugs that induce nucleolar stress. For a readout, a novel parameter termed 'nucleolar normality score' was developed that measures the ratio of the fibrillar center and granular component proteins in the nucleolus and nucleoplasm. Multiple classes of drugs were found to induce nucleolar stress, including DNA intercalators, inhibitors of mTOR/PI3K, heat shock proteins, proteasome, and cyclin-dependent kinases (CDKs). Each class of drugs induced morphologically and molecularly distinct states of nucleolar stress accompanied by changes in nucleolar biophysical properties. In-depth characterization focused on the nucleolar stress induced by inhibition of transcriptional CDKs, particularly CDK9, the main CDK that regulates RNA Pol II. Multiple CDK substrates were identified in the nucleolus, including RNA Pol I– recruiting protein Treacle, which was phosphorylated by CDK9 in vitro. These results revealed a concerted regulation of RNA Pol I and Pol II by transcriptional CDKs. Our findings exposed many classes of chemotherapy compounds that are capable of inducing nucleolar stress, and we recommend considering this in anticancer drug development. eLife assessment This study and associated data is compelling, novel, important, and well-carried out. The study demonstrates a novel finding that different chemotherapeutic agents can induce nucleolar stress, which manifests with varying cellular and molecular characteristics. The study also proposes a mechanism for how a novel type of nucleolar stress driven by CDK inhibitors may be regulated. The study sheds light on the importance of nucleolar stress in defining the on-target and off-target effects of chemotherapy in normal and cancer cells. https://doi.org/10.7554/eLife.88799.3.sa0 About eLife assessments eLife digest Ribosomes are cell structures within a compartment called the nucleolus that are required to make proteins, which are essential for cell function. Due to their uncontrolled growth and division, cancer cells require many proteins and therefore have a particularly high demand for ribosomes. Due to this, some anti-cancer drugs deliberately target the activities of the nucleolus. However, it was not clear if anti-cancer drugs with other targets also disrupt the nucleolus, which may result in side effects. Previously, it had been difficult to study how nucleoli work, partly because in human cells they vary naturally in shape, size, and number. Potapova et al. used fluorescent microscopy to develop a new way of assessing nucleoli based on the location and ratio of certain proteins. These measurements were used to calculate a "nucleolar normality score". Potapova et al. then tested over a thousand anti-cancer drugs in healthy and cancerous human cells. Around 10% of the tested drugs changed the nucleolar normality score when compared to placebo treatment, indicating that they caused nucleolar stress. For most of these drugs, the nucleolus was not the intended target, suggesting that disrupting it was an unintended side effect. Drugs inhibiting proteins called cyclin-dependent kinases caused the most drastic changes in the size and shape of nucleoli, disrupting them completely. These kinases are known to be involved in activating enzymes required for general transcription. Potapova et al. showed that they also are involved in production of ribosomal RNA, revealing an additional role in coordinating ribosome assembly. Taken together, the findings suggest that evaluating the effect of new anti-cancer drugs on the nucleolus could help to develop future treatments with less toxic side effects. The experiments also reveal new avenues for researching how cyclin-dependent kinases control the production of RNA more generally. Introduction The nucleolus is the most prominent nuclear organelle. Its primary function is the biogenesis of ribosomes – a pivotal housekeeping process essential for the translation of all proteins. Ribosome biogenesis is a major metabolic expense in a cell. This biosynthetic program requires transcription and processing of the most abundant cellular RNA – the ribosomal RNA (rRNA), and the production of 80 ribosomal proteins and hundreds of other nucleolar proteins involved in rRNA processing and assembly of ribosomal subunits (Moss and Stefanovsky, 2002; Granneman and Tollervey, 2007). Rapidly proliferating cancer cells have ribosome biogenesis shifted into overdrive, which may be one of their primary metabolic alterations (Drygin et al., 2010). Several anticancer drugs targeting ribosome biogenesis pathways have been developed (Ferreira et al., 2020), yet anticancer therapies targeting nucleolar function have not been a major focus of new drug development because of the universal role of this pathway in maintaining basic cellular functions. The main objective of this study was to identify the compounds that disrupt normal nucleolar physiology and further explore the new and unconventional agents that induce nucleolar stress. The nucleolus is a membrane-less organelle that assembles around ribosomal RNA genes (rDNA). rRNA genes in eukaryotic cells are present in hundreds of tandemly arranged repetitive copies that are transcribed by RNA polymerase I (Pol I) (reviewed in Potapova and Gerton, 2019). Nucleolar anatomy in animal cells is comprised of three distinct compartments: the fibrillar center (FC), the dense fibrillar component (DFC), and the granular component (GC) (Pederson, 2011). FC is the site of transcription that consists of rDNA and its associated transcription machinery such as transcription factor UBF and RNA Pol I. The DFC is the site of pre-rRNA processing distinguished by early RNA processing factors such as fibrillarin. The GC contains proteins involved in late rRNA processing and assembly of pre-ribosomal particles. It is marked by proteins such as nucleolin and nucleophosmin (NPM1). Changes in nucleolar organization during stress have not been studied extensively, except for the inhibition of RNA Pol I that causes the reorganization of rDNA arrays and associated FC proteins into round nucleoli with peripheral 'stress caps.' Biophysical and biochemical events underlying nucleolar reorganization under stress remain poorly understood. Nucleoli in mammalian cells can be highly polymorphic – different in shape, size, and number. It is difficult to find a single parameter that can quantitatively distinguish normal nucleolar anatomy from abnormal. To quantify the impact of anticancer drugs on nucleoli, we developed a novel imaging-based parameter that we termed 'the nucleolar normality score.' It is based on measuring nucleolar/nucleoplasmic ratios of GC component nucleolin and FC component UBF. Measuring the normality score allowed us to detect distinct states of nucleolar stress in a screen of more than a thousand chemical compounds developed as anticancer agents. The screen was conducted using a noncancer-derived cell line RPE1. This cell line was selected for evaluating the effects of anticancer drugs on normal nucleolar function. The outcome of the screen provided a broad atlas of aberrant nucleolar morphologies and their molecular triggers, where multiple drugs with the same target often produced a similar morphological and functional state. We classify four distinct categories of nucleolar stress: (1) canonical nucleolar stress with the formation of stress caps caused by DNA intercalators, (2) metabolic suppression of function caused by PI3K and mTOR inhibitors, (3) proteotoxicity with or without formation of aggresomes caused by HSP90 and proteasome inhibitors, and (4) nucleolar dissolution with an extended bare rDNA scaffold caused by cyclin-dependent kinase (CDK) inhibitors. An in-depth examination of the nucleolar stress caused by CDK inhibitors uncovered previously unknown regulation of RNA Pol I by CDKs and suggests the possibility of concerted regulation of Pol I and Pol II by transcriptional CDK activity. Finally, our study highlights the fact that many anticancer drugs can cause unintended effects on the nucleolus that can underlie off-target toxicity, which should be considered in the development and use of antineoplastic agents. Results The biological basis for the nucleolar normality score To establish a robust quantitative method for measuring nucleolar stress, we first investigated the properties of nucleolar components during the inhibition of RNA Pol I. Inhibition of Pol I transcription manifests in acute morphological changes referred to as canonical nucleolar stress. Canonical nucleolar stress is well characterized in the instance of antineoplastic agent actinomycin D (dactinomycin) that stalls Pol I transcription by intercalating into G/C-rich rDNA. This causes nucleoli to shrink and round up, with the partial dissolution of some GC proteins into the nucleoplasm and the formation of so-called 'stress caps' at the nucleolar periphery. Stress caps consist of segregated rDNA with bound FC proteins (Shav-Tal et al., 2005; Mangan et al., 2017). In this study, we inhibited Pol I using a small molecule compound CX-5461 (Drygin et al., 2011). This drug has been shown to arrest Pol I at the rDNA promoter, which blocks transcription initiation (Mars et al., 2020). To quantify the effects of CX-5461 on nucleoli by live imaging, we used hTERT immortalized human RPE1 cell lines stably expressing GC component nucleolin tagged with GFP, or FC component UBF tagged with the GFP. Expression of eGFP-nucleolin enabled us to visualize the process of nucleolar shrinking and rounding up, and the formation of small circular remnants within the first hour after RNA Pol I inhibitor treatment (Figure 1A and Video 1). With the first hour after treatment, the average intensity of eGFP-nucleolin decreased in the nucleoli and increased in the nucleoplasm (Figure 1A, right panel), indicating a higher proportion of total nucleolin dissolved in the nucleoplasm. This resulted in a decrease in the fluorescence intensity ratio of the nucleolar pool relative to the nucleoplasmic pool. In cells expressing eGFP-UBF, treatment with CX-5461 induced UBF condensation at the periphery of the nucleolar remnants and the formation of stress caps (Figure 1B and Video 2). The intensity of eGFP-UBF increased in these small stress caps, while the intensity in the nucleoplasm did not change (Figure 1B, right panel). For eGFP-UBF, the average fluorescence intensity ratio of the stress caps relative to the nucleoplasmic pool increased. Figure 1 Download asset Open asset Nucleolar normality score as a parameter for measuring nucleolar stress. (A) Time-lapse images of eGFP-nucleolin expressing cell treated with 2.5 µM Pol I inhibitor CX-5461 at time 0 are shown. Nucleoli shrink and round up forming small circular remnants. Fluorescence intensity, indicated by the heatscale, decreases in nucleolar remnants and increases in the nucleoplasm. The complete video sequence is shown in Video 1. Bar, 10 µm. The plot on the right shows the average intensity of eGFP-nucleolin in nucleoli and in the nucleoplasm normalized to the initial intensity at time 0. The plot is an average of 10 cells, bars denote standard deviation. (B) Time‐lapse images of eGFP‐UBF expressing cell treated with 2.5 μM CX‐5461 at time 0 are shown. UBF condenses on the periphery of nucleolar remnants forming stress caps of high fluorescence intensity. The complete video sequence is shown in Video 2. Bar, 10 μm. The plot on the right shows the average intensity of eGFP‐UBF in stress caps and in the nucleoplasm normalized to the initial intensity at time 0. The plot is an average of 13 cells, bars denote standard deviation. (C) Fluorescence recovery after photobleaching (FRAP) analysis of eGFP-nucleolin in untreated cells and cells treated with 2.5 µM CX-5461 is shown. The plot is an average of normalized fluorescence intensities of 14 and 10 cells. Bars denote standard deviation. The graph on the right shows corresponding individual T1/2 measurements. Asterisk indicates p<0.05 (t-test comparing the drug-treated group to untreated). (D) FRAP analysis of eGFP-UBF in untreated cells and cells treated with 2.5 µM CX-5461. The plot is an average of normalized fluorescence intensities of 11 and 12 cells. Bars denote standard deviation. The graph on the right shows corresponding individual T1/2 measurements. t-test did not detect a significant difference between the two treatments. (E) The immunofluorescence image illustrates the Nucleolar Normality score measurement. RPE1 cells were labeled with antibodies against nucleolin and UBF and counterstained with DAPI. Segmentation of nucleolar regions was performed on UBF, and whole nuclei were segmented on DAPI. Nucleoplasm regions are areas within the nuclei without nucleoli. The nucleoplasmic intensity was calculated by subtracting the integrated intensity of nucleoli from the integrated intensity of the whole nuclei. For both nucleolin and UBF, the integrated intensity of the nucleolar regions of each cell was divided by the integrated intensity of the nucleoplasm of that cell, giving the nucleolar/nucleoplasm ratio. Dividing the nucleolar/nucleoplasm ratio of the nucleolin by the nucleolar/nucleoplasm ratio of the UBF provides a nucleolar normality score for each cell. (F) Normality score measurements of individual cells treated with DMSO (vehicle), 2.5 µM CX-5461, or 5 µM topoisomerase inhibitor camptothecin, normalized to the average value of DMSO-treated cells. More than 40 individual cells were measured for each condition. Asterisks indicate p<0.0001 (unpaired t-test comparing drug-treated groups to DMSO). Figure 1—source data 1 Source data for Figure 1A-F. https://cdn.elifesciences.org/articles/88799/elife-88799-fig1-data1-v1.zip Download elife-88799-fig1-data1-v1.zip Video 1 Download asset This video cannot be played in place because your browser does support HTML5 video. You may still download the video for offline viewing. Download as MPEG-4 Download as WebM Download as Ogg Fluorescence and phase-contrast time-lapse video of a human RPE1 cell stably expressing eGFP-nucleolin that was treated with 2.5 µM RNA Pol I inhibitor CX-5461. Nucleoli shrink and round up forming small circular remnants. Fluorescence intensity decreases in nucleolar remnants and increases in the nucleoplasm. Time is indicated as minutes after drug addition. Bar, 10 μm. Video 2 Download asset This video cannot be played in place because your browser does support HTML5 video. You may still download the video for offline viewing. Download as MPEG-4 Download as WebM Download as Ogg Fluorescence and phase-contrast time-lapse video of a human RPE1 cell stably expressing eGFP-UBF that was treated with 2.5 µM RNA Pol I inhibitor CX-5461. UBF condenses on the periphery of nucleolar remnants forming stress caps of high fluorescence intensity. Time is indicated as minutes after drug addition. Bar, 10 μm. Next, we investigated the mobility of the eGFP-nucleolin and eGFP-UBF by fluorescence recovery after photobleaching (FRAP) before and after nucleolar stress induced with CX-5461. Nucleolin became more mobile in stressed cells (the average half-time recovery T1/2 went down from 4.58 ± 1.88 s to 2.89 ± 0.88 s, Figure 1C), consistent with its redistribution to the nucleoplasm. The T1/2 of UBF did not significantly change with stress and stayed on the average of 12–14 s (Figure 1D), indicating that the rDNA-binding properties of UBF that likely underlie its FRAP behavior were not affected by RNA Pol I inhibition. This is consistent with UBF acting as a stable bookmark of the rDNA during mitosis, when RNA Pol I activity is very low (Roussel et al., 1993; Gébrane-Younès et al., 1997). This contrasting behavior of nucleolin and UBF after Pol I inhibition provided the basis for the nucleolar stress parameter that we termed the nucleolar normality score. The nucleolar normality score is a ratio of the nucleolar fraction of nucleolin relative to the nucleolar fraction of UBF (Figure 1E). Image processing and calculation of the normality score are explained in detail in 'Materials and methods.' This parameter is applicable to fixed cells where both proteins are labeled by immunofluorescence. In a normal, unstressed situation the average normality score has a consistent value that is characteristic for a given experimental system. As nucleolin dissolves in the nucleoplasm and UBF becomes segregated, the normality score decreases. The normality score was very robust at detecting the strong nucleolar stress phenotype caused by CX-5461, but it was also proven to detect more subtle morphological changes, such as the stress caused by topoisomerase inhibitor camptothecin (Figure 1F). This parameter allowed us to detect nucleolar stress phenotypes that are less pronounced and measure the degree of nucleolar perturbations of various origins. High-throughput imaging screen for anticancer drugs that induce nucleolar stress We screened nucleolar normality in cells treated with a chemical library containing 1180 anticancer compounds developed for multiple cancers, some of them FDA-approved and used clinically. The main goal was to broadly identify and categorize distinct states of nucleolar stress and their molecular triggers. For the screen, normal human hTERT-immortalized RPE1 cells were seeded in 384-well plates and treated with the library compounds at 1 µM and 10 µM for 24 hr. Drug treatment was followed by fixation and labeling with antibodies against UBF and nucleolin (Figure 2A). Forty single-plane fields containing hundreds of cells were imaged per well. Compounds were called hits if their normality score was more than 2 standard deviations away from the DMSO (vehicle) control average. Of 1180 compounds present in the library, 12.9% were hits. Also, 7% of the compounds in the library were hits at both 1 and 10 µM, and 5.8% were hits only at 10 µM (Figure 2B). The majority of the hits were validated (Figure 2—figure supplement 1A). The complete list of hits is provided in Supplementary file 1. Figure 2 with 1 supplement see all Download asset Open asset Anticancer drug screen for compounds that induce nucleolar stress. (A) The diagram illustrates the workflow for the library screen for anticancer compounds that induce nucleolar stress. (B) From the total 1180 compounds, 83 were hits at both 1 and 10 µM, and 69 were hits at 10 µM only. The full list of hit compounds with normality scores is provided in Supplementary file 1. (C) Normality score results from cells treated with 10 µM drug are plotted versus cell count. Both parameters were normalized to the average of the DMSO control (black points). Red points denote hits at 1 and 10 µM, purple points 10 µM only. BMH21 is a Pol I inhibitor present in the library and serves as an internal control. (D) Combined 1 and 10 µM hits and 10 µM only hits grouped by the target. (E) Enrichment of hit drug targets relative to their presence in the library is plotted versus the probability of random occurrence (p-value). A low p-value indicates that the probability of a target being enriched at random is low. Gray points indicate targets whose enrichment was not significant, colored points with labels denote significantly enriched targets (p<0.05). (F) Validation of selected hits from different target classes in multiple cell lines is shown. For each cell line, normality scores were normalized to their own DMSO controls. All drugs caused significant (p<0.05) reductions in normality scores in all cell lines. Figure 2—source data 1 Source data for Figure 2C-F. https://cdn.elifesciences.org/articles/88799/elife-88799-fig2-data1-v1.zip Download elife-88799-fig2-data1-v1.zip All hits in the screen had normality scores lower than the control, that is, this parameter only went down, not up, in drug-treated cells. The number of cells in hit wells was typically lower than in control wells, indicating that the majority of drugs that induced nucleolar stress were cytostatic or cytotoxic (Figure 2C). The screening process did not distinguish whether the cytotoxic effects of the identified hits were a result of inhibiting their intended targets, impacting the nucleolus, or a combined effect. It is important to note that a low normality score is not necessarily a consequence of reduced viability because many drugs in the screen were cytostatic/cytotoxic without causing nucleolar stress. Rather, it underscores the fact that inhibition of nucleolar biological processes is overall detrimental to viability and proliferation. One of the internal positive controls for nucleolar stress in the screen was the compound BMH-21 – a well-characterized RNA Pol I inhibitor present in the library. BMH-21 intercalates in the DNA and binds strongly to GC-rich rDNA, repressing RNA Pol I transcription (Colis et al., 2014; Wei et al., 2018). BMH-21 induced a canonical nucleolar stress phenotype with dispersed nucleolin and segregation of UBF into stress caps. Cells treated with BMH-21 showed a 7.7-fold reduction in the normality score and a 2-fold reduction in cell number compared to DMSO control (highlighted in Figure 2C). The anticancer compound library contained chemical inhibitors for various targets, mostly enzymes. Grouping hits by drug target showed that inhibitors of mTOR and PI3 kinase had the highest frequency among all hits. Other frequently hit drug targets were HSP90, Topoisomerases, and CDKs (Figure 2D). However, the overall representation of targets in the library varied: prioritized cancer targets and highly druggable targets were among the most represented. Since the representation of targets in the library was not equivalent, we calculated the enrichment of targets among hits relative to their presence in the library. The most significantly enriched targets (p<0.001) were HSP90, mTOR, PI3K, and topoisomerase inhibitors. Among other significantly enriched targets (p<0.05) were inhibitors of dihydrofolate reductase (DFHR), proteasome, CDKs, and other kinases (Figure 2E). To ensure that the drug responses were not unique to RPE1 cells, validation was performed on additional cell lines with a panel of selected potent hits from different target classes: HSP90 inhibitors – 17-AAG, onalespib, BIIB-21; CDK inhibitors – dinaciclib, flavopiridol, LY2857785; proteasome inhibitors – carfilzomib and oprozomib; mTOR inhibitor sapanisertib; PI3K inhibitor taselicib; and topoisomerase inhibitors camptothecin and doxorubicin. This panel of drugs was validated in four other cell lines: two hTERT-immortalized cell lines – BJ5TA skin fibroblasts and CHON-002 fibroblasts, and two cancer-derived cell lines – DLD1 colon adenocarcinoma and HCT116 colon carcinoma. In all experimental cell lines, raw nucleolar normality scores before the drug treatments were different. Cancer cell lines had lower starting normality scores than hTERT cell lines (Figure 2—figure supplement 1B and C). To compensate for this initial difference, the results of drug treatments from each cell line were normalized to the vehicle control of that cell line. The degree of reduction in nucleolar normality scores varied between cell lines, which could be attributed to differences in baseline normality scores, as well as proteomic and metabolic shifts, alterations in signaling pathways that control ribosome production, and, potentially, variations in intracellular drug levels. Nonetheless, all compounds caused a significant reduction in nucleolar normality scores in all cell lines (Figure 2F). This result ensures that the nucleolar stress induced by these drugs was not specific to a particular cell line. Characterization of nucleolar stress induced by selected inhibitors Canonical nucleolar stress induced by Pol I inhibitors is linked to reduced rRNA production. We measured the effect of the selected drug panel on rRNA synthesis by incorporation of 5-ethynyluridine (5-EU) into nascent RNA (Jao and Salic, 2008). Since ribosomal RNA can account for ~80% of the total cellular RNA (Palazzo and Lee, 2015), the total amount of nascent RNA approximates the synthesis of ribosomal RNA. All drugs in the panel caused a decrease in 5-EU incorporation, but to varying degrees. The level of reduction was similar within the same classes of drugs based on target, but different between classes (Figure 3A). Correlation analysis with normality scores showed that there was a trend for drugs with lower normality scores to have lower rRNA synthesis, but it was not statistically significant (Figure 3B). Furthermore, nucleolar stress phenotypes were distinct by target (Figure 3—figure supplement 1). This lack of significant correlation implied that the normality score may not be explained by a reduction in rDNA transcription alone. Figure 3 with 2 supplements see all Download asset Open asset Characterization of nucleolar stresses in a panel of selected drugs. (A) 5-ethynyluridine (5-EU) incorporation was measured in RPE1 cells treated with the panel of selected drug hits from the screen. All compounds were added for 10 hr followed by 4 hr of 0.5 mM 5-EU incorporation. All drugs were at 10 µM concentration except LY2857785 and CX-5461 were used at 2.5 µM, camptothecin and flavopiridol were used at 5 µM, and doxorubicin and BMH-21 at 1 µM. 5-EU-labeled RNA was detected with fluorescent azide and quantified by imaging. Plots represent means with standard deviations of three or more large fields of view containing hundreds of cells. Raw fluorescent intensity values were normalized to the average of the DMSO controls. All drug treatments caused a significant reduction in 5-EU incorporation compared to DMSO (p<0.01, unpaired t-tests). (B) A correlation plot of average nucleolar normality scores versus average 5-EU fluorescence is shown. Both parameters were normalized to the average of the DMSO controls. The trend for drugs with lower normality scores to have lower 5-EU incorporation was not significant (Pearson's r = 0.33, p=0.23). (C) Fluorescent in situ hybridization with antibody immunolabeling (immuno-FISH) images of drug-treated RPE1 cells labeled with human rDNA probe (green), UBF (red), and nucleolin (magenta) are shown. Nuclei were counterstained with DAPI (blue). Bar, 10 µm. The duration of 2.5 µM CX-5461 and 10 µM flavopiridol treatments was 5 hr, 10 µM 17-AAG and 10 µM carfilzomib 10 hr. Magnified inserts show details of individual nucleoli (bar, 1 µm). Note peripheral stress caps in CX-5461 and unfolded rDNA/UBF in flavopiridol–treated cells. The arrow in the carfilzomib panel indicates the diffuse pool of UBF not associated with rDNA. (D) Immunofluorescence images of RPE1 cells treated as in (C) and labeled with antibodies against UBF (green) and POLR1A (red, antibody C-1). Nuclei were counterstained with DAPI. Bar, 10 µm. Magnified insets show details of individual nucleoli (bar, 1 µm). UBF and POLR1A label the same structures in all treatments except flavopiridol. (E) The quantification of POLR1A immunofluorescence from (D) is plotted. The box plot depicts ratios of POLR1A signal intensity in the nucleolus versus nucleoplasm normalized to the average of DMSO controls. The plot represents the means of 4–5 fields of view containing a total of 80–100 cells. Asterisks indicate a significant reduction in nucleolar POLR1A (p<0.0001, unpaired t-test flavopiridol vs. DMSO). (F) Western blot analysis of POLR1A protein levels in RPE1 cells treated with the indicated drugs for 8 hr. Total POLR1A levels were not altered. Figure 3—source data 1 Source data for Figure 3A,B and E. https://cdn.elifesciences.org/articles/88799/elife-88799-fig3-data1-v1.zip Download elife-88799-fig3-data1-v1.zip Figure 3—source data 2 Source data for Figure 3F. https://cdn.elifesciences.org/articles/88799/elife-88799-fig3-data2-v1.zip Download elife-88799-fig3-data2-v1.zip Inhibitors of mTOR and PI3 kinase had the highest representation among all hits in the anticancer compound library. mTOR and PI3K are metabolic pathways that positively regulate ribosome biogenesis on multiple levels including rDNA transcription (Mayer and Grummt, 2006; Pelletier et al., 2018), so the strong (~60%) reduction in 5-EU incorporation in mTOR inhibitor sapanisertib and PI3K inhibitor taselicib was predictable. The reduction in normality score was likely a consequence of inhibiting upstream activating pathways that stimulate rDNA transcription and ribosome biogenesis. Another major class of drugs that induced low normality scores were inhibitors of topoisomerase II, particularly anthracyclines that intercalate into DNA and act as topoisomerase poisons (doxorubicin, epirubicin, idarubicin, daunorubicin, pirarubicin, mitoxantrone, pixantrone). All DNA intercalating topoisomerase poison hits caused nucleolar shrinkage, rounding, and the canonical stress caps associated with RNA Pol I inhibition. Notably, actinomycin D and CX-5461 can also poison the action of topoisomerases (Trask and Muller, 1988; Bruno et al., 2020). Topoisomerase activity may be needed to resolve topological stress at the rDNA to continue transcription. rDNA transcription may be hypersensitive to DNA intercalators in general (Andrews et al., 2021), and for man
The short arms of the human acrocentric chromosomes 13, 14, 15, 21 and 22 (SAACs) share large homologous regions, including ribosomal DNA repeats and extended segmental duplications 1,2 . Although the resolution of these regions in the first complete assembly of a human genome—the Telomere-to-Telomere Consortium’s CHM13 assembly (T2T-CHM13)—provided a model of their homology 3 , it remained unclear whether these patterns were ancestral or maintained by ongoing recombination exchange. Here we show that acrocentric chromosomes contain pseudo-homologous regions (PHRs) indicative of recombination between non-homologous sequences. Utilizing an all-to-all comparison of the human pangenome from the Human Pangenome Reference Consortium 4 (HPRC), we find that contigs from all of the SAACs form a community. A variation graph 5 constructed from centromere-spanning acrocentric contigs indicates the presence of regions in which most contigs appear nearly identical between heterologous acrocentric chromosomes in T2T-CHM13. Except on chromosome 15, we observe faster decay of linkage disequilibrium in the pseudo-homologous regions than in the corresponding short and long arms, indicating higher rates of recombination 6,7 . The pseudo-homologous regions include sequences that have previously been shown to lie at the breakpoint of Robertsonian translocations 8 , and their arrangement is compatible with crossover in inverted duplications on chromosomes 13, 14 and 21. The ubiquity of signals of recombination between heterologous acrocentric chromosomes seen in the HPRC draft pangenome suggests that these shared sequences form the basis for recurrent Robertsonian translocations, providing sequence and population-based confirmation of hypotheses first developed from cytogenetic studies 50 years ago 9 .
The human Y chromosome has been notoriously difficult to sequence and assemble because of its complex repeat structure that includes long palindromes, tandem repeats and segmental duplications1-3. As a result, more than half of the Y chromosome is missing from the GRCh38 reference sequence and it remains the last human chromosome to be finished4,5. Here, the Telomere-to-Telomere (T2T) consortium presents the complete 62,460,029-base-pair sequence of a human Y chromosome from the HG002 genome (T2T-Y) that corrects multiple errors in GRCh38-Y and adds over 30 million base pairs of sequence to the reference, showing the complete ampliconic structures of gene families TSPY, DAZ and RBMY; 41 additional protein-coding genes, mostly from the TSPY family; and an alternating pattern of human satellite 1 and 3 blocks in the heterochromatic Yq12 region. We have combined T2T-Y with a previous assembly of the CHM13 genome4 and mapped available population variation, clinical variants and functional genomics data to produce a complete and comprehensive reference sequence for all 24 human chromosomes. We present the complete 62,460,029-base-pair sequence of a human Y chromosome from the HG002 genome (T2T-Y) that corrects multiple errors in GRCh38-Y and adds over 30 million base pairs of sequence to the reference.
This protocol describes the fluorescence in situ hybridization (FISH) of DNA probes on mitotic chromosome spreads optimized for two super-resolution microscopy approaches-structured illumination microscopy (SIM) and stimulated emission depletion (STED). It is based on traditional DNA FISH methods that can be combined with immunofluorescence labeling (Immuno-FISH). This technique previously allowed us to visualize ribosomal DNA linkages between human acrocentric chromosomes and provided information about the activity status of linked rDNA loci. Compared to the conventional wide-field and confocal microscopy, the quality of SIM and STED data depends a lot more on the optimal specimen preparation, choice of fluorophores, and quality of the fluorescent labeling. This protocol highlights details that make specimens suitable for super-resolution microscopy and tips for good imaging practices.
Existing human genome assemblies have almost entirely excluded repetitive sequences within and near centromeres, limiting our understanding of their organization, evolution, and functions, which include facilitating proper chromosome segregation. Now, a complete, telomere-to-telomere human genome assembly (T2T-CHM13) has enabled us to comprehensively characterize pericentromeric and centromeric repeats, which constitute 6.2% of the genome (189.9 megabases). Detailed maps of these regions revealed multimegabase structural rearrangements, including in active centromeric repeat arrays. Analysis of centromere-associated sequences uncovered a strong relationship between the position of the centromere and the evolution of the surrounding DNA through layered repeat expansions. Furthermore, comparisons of chromosome X centromeres across a diverse panel of individuals illuminated high degrees of structural, epigenetic, and sequence variation in these complex and rapidly evolving regions.
Since its initial release in 2000, the human reference genome has covered only the euchromatic fraction of the genome, leaving important heterochromatic regions unfinished. Addressing the remaining 8% of the genome, the Telomere-to-Telomere (T2T) Consortium presents a complete 3.055 billion–base pair sequence of a human genome, T2T-CHM13, that includes gapless assemblies for all chromosomes except Y, corrects errors in the prior references, and introduces nearly 200 million base pairs of sequence containing 1956 gene predictions, 99 of which are predicted to be protein coding. The completed regions include all centromeric satellite arrays, recent segmental duplications, and the short arms of all five acrocentric chromosomes, unlocking these complex regions of the genome to variational and functional studies.