I. ABSTRACT We present an analytical framework for modeling eukaryotic DNA replication that, given experimental Replication Fork Directionality (RFD) data, enables Bayesian inference of origin number, activation delay and intrinsic timing ( λ i ), (the mean replication time if each origin were isolated). By deriving closed-form expressions for RFD and Mean Replication Timing (MRT) under exponential and a specific Weibull firing-time distributions as functions of and ( λ i ), we eliminate the need for stochastic simulations. These analytical results reveal that RFD, as a ratio of fork directions, is invariant under joint rescaling of intrinsic timing and fork speed; absolute intrinsic timing can nonetheless be inferred when fork speed is independently measured. We demonstrate that under exponential firing distribution for the origin, the observed efficiency ( E i ), i.e. the probability for an origin to fire which accounts for nearby origin, is simply MRT( x i ) /λ i . The closed-form RFD expressions allow to use a Bayesian method that achieves 0.96–0.99 correlation with yeast RFD profiles and resolves ~780 origins in S. cerevisiae . Our framework identifies about 150 origins with biologically significant delays (≥ 3 minute), revealing regulated activation kinetics undetectable by existing methods. By quantifying how origin intrinsic timing and delays shape replication timing landscapes, this work confirms yeast as a paradigm organism for studying DNA replication control mechanisms.
Current approaches to mapping fork progression in the human genome suffer from drastically low throughput. Here, we introduce ForkML, a nanopore sequencing-based method automatically positioning thousands of individual fork velocities by tracking BrdU incorporation into replicating DNA after double pulse-labelling of asynchronous cells. ForkML recovers known human fork speed, accurately detects replication stress, and, crucially, connects replication dynamics to genomic and chromatin contexts, exposing fork slowdown in early-replicating transcribed regions. Current mapping of fork progression in the human genome suffers from drastically low throughput. Here, the authors introduce ForkML, a nanopore sequencing-based method automatically positioning thousands of individual fork velocities by tracking BrdU incorporation into asynchronously growing cells.
Large vertebrate genomes duplicate by activating tens of thousands of DNA replication origins, irregularly spaced along the genome. The spatial and temporal regulation of the replication process is not yet fully understood. To investigate the DNA replication dynamics, we developed a methodology called RepliCorr, which uses the spatial correlation between replication patterns observed on stretched single-molecule DNA obtained by either DNA combing or high-throughput optical mapping. The analysis revealed two independent spatiotemporal processes that regulate the replication dynamics in the Xenopus model system. These mechanisms are referred to as a fast and a slow replication mode, differing by their opposite replication fork speed and rate of origin firing. We found that Polo-like kinase 1 (Plk1) depletion abolished the spatial separation of these two replication modes. In contrast, neither replication checkpoint inhibition nor Rap1-interacting factor (Rif1) depletion affected the distribution of these replication patterns. These results suggest that Plk1 plays an essential role in the local coordination of the spatial replication program and the initiation-elongation coupling along the chromosomes in Xenopus, ensuring the timely completion of the S phase.
Mammalian DNA replication origins have been historically difficult to identify and their determinants are still unresolved. Here, we first review methods developed over the last decades to map replication initiation sites either directly via initiation intermediates or indirectly via determining replication fork directionality profiles. We also discuss the factors that may specify these sites as replication initiation sites. Second, we address the controversy that has emerged from these results over whether origins are narrowly defined and localized to specific sites or are more dispersed and organized into broad zones. Ample evidence in favor of both scenarios currently creates an impression of unresolved confusion in the field. We attempt to formulate a synthesis of both models and to reconcile discrepant findings. It is evident that not only one approach is sufficient in isolation but that the combination of several is instrumental toward understanding initiation sites in mammalian genomes. We argue that an aggregation of several individual and often inefficient initiation sites into larger initiation zones and the existence of efficient unidirectional initiation sites and fork stalling at the borders of initiation zones can reconcile the different observations.
Current temporal studies of DNA replication are either low-resolution or require complex cell synchronisation and/or sorting procedures. Here we introduce Nanotiming, a single-molecule, nanopore sequencing-based method producing high-resolution, telomere-to-telomere replication timing (RT) profiles of eukaryotic genomes by interrogating changes in intracellular dTTP concentration during S phase through competition with its analogue bromodeoxyuridine triphosphate (BrdUTP) for incorporation into replicating DNA. This solely demands the labelling of asynchronously growing cells with an innocuous dose of BrdU during one doubling time followed by BrdU quantification along nanopore reads. We demonstrate in S. cerevisiae model eukaryote that Nanotiming reproduces RT profiles generated by reference methods both in wild-type and mutant cells inactivated for known RT determinants. Nanotiming is simple, accurate, inexpensive, amenable to large-scale analyses, and has the unique ability to access RT of individual telomeres, revealing that Rif1 iconic telomere regulator selectively delays replication of telomeres associated with specific subtelomeric elements.
Nuclear architecture and chromosome folding are often speculated to influence genome replication. In yeasts, centromeres cluster close to the spindle pole body and telomeres position at the nuclear periphery. This 'Rabl configuration' spatially segregates the most early and late replicating parts of the genome, centromeres and telomeres, respectively, suggesting that origin position along the centromere-telomere axis may influence origin activity. Here, we investigated DNA replication in a wild-type, 16-chromosome Saccharomyces cerevisiae strain and in its single-chromosome counterpart engineered by chromosome fusion and elimination of all but two telomeres and one centromere, which strongly affects genome folding and abrogates the Rabl conformation. Using nanopore sequencing-based methods, we found that the DNA replication program of both strains was virtually indistinguishable, with the exception of origin inactivation next to deleted centromeres and changes in origin efficiency and fork direction at chromosome fusions, as anticipated from the known origin-regulation properties of centromeres and telomeres. Only a handful of replication changes, mostly due to local origin repression, were observed elsewhere. Fork speed was also unaffected except at deleted centromeres. In conclusion, the DNA replication program of budding yeast is remarkably resilient to perturbations of chromosome folding and loss of the Rabl conformation.
Current temporal studies of DNA replication are either low-resolution or require complex cell synchronisation and/or sorting procedures. Here we introduce Nanotiming, a nanopore sequencing-based method producing high-resolution, telomere-to-telomere replication timing (RT) profiles of eukaryotic genomes by interrogating changes in intracellular dTTP concentration during S phase through competition with its analogue bromodeoxyuridine triphosphate (BrdUTP) for incorporation into replicating DNA. Nanotiming solely demands the labelling of asynchronously growing cells with an innocuous dose of BrdU during one doubling time followed by BrdU quantification along nanopore reads. We demonstrate in yeast S. cerevisiae that Nanotiming precisely reproduces RT profiles generated by reference methods in wild-type and mutant cells inactivated for known RT determinants, for one-tenth of the cost. Nanotiming is simple, accurate, inexpensive, amenable to large-scale analyses, and is capable of unveiling RT at individual telomeres, revealing that Rif1 iconic telomere regulator directly delays the replication only of telomeres with specific subtelomeric elements. ### Competing Interest Statement The authors have declared no competing interest.
Studying the dynamics of genome replication in mammalian cells has been historically challenging. To reveal the location of replication initiation and termination in the human genome, we developed Okazaki fragment sequencing (OK-seq), a quantitative approach based on the isolation and strand-specific sequencing of Okazaki fragments, the lagging strand replication intermediates. OK-seq quantitates the proportion of leftward- and rightward-oriented forks at every genomic locus and reveals the location and efficiency of replication initiation and termination events. Here we provide the detailed experimental procedures for performing OK-seq in unperturbed cultured human cells and budding yeast and the bioinformatics pipelines for data processing and computation of replication fork directionality. Furthermore, we present the analytical approach based on a hidden Markov model, which allows automated detection of ascending, descending and flat replication fork directionality segments revealing the zones of replication initiation, termination and unidirectional fork movement across the entire genome. These tools are essential for the accurate interpretation of human and yeast replication programs. The experiments and the data processing can be accomplished within six days. Besides revealing the genome replication program in fine detail, OK-seq has been instrumental in numerous studies unravelling mechanisms of genome stability, epigenome maintenance and genome evolution.
In human and other metazoans, the determinants of replication origin location and strength are still elusive. Origins are licensed in G1 phase and fired in S phase of the cell cycle, respectively. It is debated which of these two temporally separate steps determines origin efficiency. Experiments can independently profile mean replication timing (MRT) and replication fork directionality (RFD) genome-wide. Such profiles contain information on multiple origins' properties and on fork speed. Due to possible origin inactivation by passive replication, however, observed and intrinsic origin efficiencies can markedly differ. Thus, there is a need for methods to infer intrinsic from observed origin efficiency, which is context-dependent. Here, we show that MRT and RFD data are highly consistent with each other but contain information at different spatial scales. Using neural networks, we infer an origin licensing landscape that, when inserted in an appropriate simulation framework, jointly predicts MRT and RFD data with unprecedented precision and underlies the importance of dispersive origin firing. We furthermore uncover an analytical formula that predicts intrinsic from observed origin efficiency combined with MRT data. Comparison of inferred intrinsic origin efficiencies with experimental profiles of licensed origins (ORC, MCM) and actual initiation events (Bubble-seq, SNS-seq, OK-seq, ORM) show that intrinsic origin efficiency is not solely determined by licensing efficiency. Thus, human replication origin efficiency is set at both the origin licensing and firing steps. Author summaryDNA replication is a vital process that produces two identical replicas of DNA from one DNA molecule, ensuring the faithful transmission of genetic information from mother to daughter cells. The synthesis of new DNA strands initiates at multiple sites, termed replication origins, propagates bidirectionally, and terminates by merging of converging strands. Replication initiation continues in unreplicated DNA but is blocked in replicated DNA. Experiments have only given partial information about origin usage. In this work we reveal the exact propensity of any site to initiate replication along human chromosomes. First, we simulate the DNA replication process using approximate origin information, predict the direction and time of replication at each point of the genome, and train a neural network to precisely recover from the predictions the starting origin information. Second, we apply this network to real replication time and direction data, extracting the replication initiation propensity landscape that exactly predicts them. We compare this landscape to independent origin usage data, benchmarking them, and to landscapes of protein factors that mark potential origins. We find that the local abundance of such factors is insufficient to predict replication initiation and we infer to which extent other chromosomal cues locally influence potential origin usage.
Most genome replication mapping methods profile cell populations, masking cell-to-cell heterogeneity. Here, we describe FORK-seq, a nanopore sequencing method to map replication of single DNA molecules at 200 nucleotide resolution using a nanopore current interpretation tool allowing the quantification of BrdU incorporation. Along pulse-chased replication intermediates from Saccharomyces cerevisiae, we can orient replication tracks and reproduce population-based replication directionality profiles. Additionally, we can map individual initiation and termination events. Thus, FORK-seq reveals the full extent of cell-to-cell heterogeneity in DNA replication.
Little is known about replication fork velocity variations along eukaryotic genomes, since reference techniques to determine fork speed either provide no sequence information or suffer from low throughput. Here we present NanoForkSpeed, a nanopore sequencing-based method to map and extract the velocity of individual forks detected as tracks of the thymidine analogue bromodeoxyuridine incorporated during a brief pulse-labelling of asynchronously growing cells. NanoForkSpeed retrieves previous Saccharomyces cerevisiae mean fork speed estimates (≈2 kb/min) in the BT1 strain exhibiting highly efficient bromodeoxyuridine incorporation and wild-type growth, and precisely quantifies speed changes in cells with altered replisome progression or exposed to hydroxyurea. The positioning of >125,000 fork velocities provides a genome-wide map of fork progression based on individual fork rates, showing a uniform fork speed across yeast chromosomes except for a marked slowdown at known pausing sites.
Mutational signatures defined by single base substitution (SBS) patterns in cancer have elucidated potential mutagenic processes that contribute to malignancy. Two prevalent mutational patterns in human cancers are attributed to the APOBEC3 cytidine deaminase enzymes. Among the seven human APOBEC3 proteins, APOBEC3A is a potent deaminase and proposed driver of cancer mutagenesis. In this study, we prospectively examine genome-wide aberrations by expressing human APOBEC3A in avian DT40 cells. From whole-genome sequencing, we detect hundreds to thousands of base substitutions per genome. The APOBEC3A signature includes widespread cytidine mutations and a unique insertion-deletion (indel) signature consisting largely of cytidine deletions. This multi-dimensional APOBEC3A signature is prevalent in human cancer genomes. Our data further reveal replication-associated mutations, the rate of stem-loop and clustered mutations, and deamination of methylated cytidines. This comprehensive signature of APOBEC3A mutagenesis is a tool for future studies and a potential biomarker for APOBEC3 activity in cancer.
The replication strategy of metazoan genomes is still unclear, mainly because definitive maps of replication origins are missing. High-throughput methods are based on population average and thus may exclusively identify efficient initiation sites, whereas inefficient origins go undetected. Single-molecule analyses of specific loci can detect both common and rare initiation events along the targeted regions. However, these usually concentrate on positioning individual events, which only gives an overview of the replication dynamics. Here, we computed the replication fork directionality (RFD) profiles of two large genes in different transcriptional states in chicken DT40 cells, namely untranscribed and transcribed DMD and CCSER1 expressed at WT levels or overexpressed, by aggregating hundreds of oriented replication tracks detected on individual DNA fibres stretched by molecular combing. These profiles reconstituted RFD domains composed of zones of initiation flanking a zone of termination originally observed in mammalian genomes and were highly consistent with independent population-averaging profiles generated by Okazaki fragment sequencing. Importantly, we demonstrate that inefficient origins do not appear as detectable RFD shifts, explaining why dispersed initiation has remained invisible to population-based assays. Our method can both generate quantitative profiles and identify discrete events, thereby constituting a comprehensive approach to study metazoan genome replication.