Abstract Pangenome graphs conveniently represent genetic variation within a population. Several types of such graphs have been proposed, with varying properties and potential applications. Among them, variation graphs (VGs) seem best suited to replace reference genomes in sequencing data processing, while whole genome alignments (WGAs) are particularly practical for comparative genomics applications. For both models, no widely accepted optimization criteria for a graph representing a given set of genomes have been proposed. In the current paper we introduce the concept of homology relation induced by a pangenome graph on the characters of represented genomic sequences and define such relations for both VG and WGA model. Then, we use this concept to propose homology-based metrics for comparing different graphs representing the same genome collection, and to formulate the desired properties of transformations between VG and WGA models. Moreover, we propose several such transformations and examine their properties on pangenome graph data. Finally, we provide implementations of these transformations in a package WGAtools, available at https://github.com/anialisiecka/WGAtools .
The success of pangenome-based approaches to genomics analysis depends largely on the existence of efficient methods for constructing pangenome graphs that are applicable to large genome collections. In the current paper we present AlfaPang, a new pangenome graph building algorithm. AlfaPang is based on a novel alignment-free approach that allows to construct pangenome graphs using significantly less computational resources than state-of-the-art tools. The code of AlfaPang is freely available at https://github.com/AdamCicherski/AlfaPang .
BackgroundMultiple sequence alignment (MSA) has proven extremely useful in computational biology, especially in inferring evolutionary relationships via phylogenetic analysis and providing insight into protein structure and function. An alternative to the standard MSA model is partial order alignment (POA), in which aligned sequences are represented as paths in a graph rather than rows in a matrix. While the POA model has proven useful in several applications (e.g. sequencing reads assembly and pangenome structure exploration), we lack efficient visualization tools that could highlight its advantages.ResultsWe propose Sequence Flow - a web application designed to address the above problem. Sequence Flow presents the POA as a Sankey diagram, a kind of graph visualisation typically used for graphs representing flowcharts. Sequence Flow enables interactive alignment exploration, including fragment selection, highlighting a selected group of sequences, modification of the position of graph nodes, structure simplification etc. After adjustment, the visualization can be saved as a high-quality graphic file. Thanks to the use of SanKEY.js - a JavaScript library for creating Sankey diagrams, designed specifically to visualize POAs, Sequence Flow provides satisfactory performance even with large alignments.ConclusionsWe provide Sankey diagram-based POA visualization tools for both end users (Sequence Flow) and bioinformatic software developers (SanKEY.js). Sequence Flow webservice is available at https://sequenceflow.mimuw.edu.pl/. The source code for SanKEY.js is available at https://github.com/Krzysiekzd/SanKEY.js and for Sequence Flow at https://github.com/Krzysiekzd/SequenceFlow.
Pangenomes serve as a framework for joint analysis of genomes of related organisms. Several pangenome models were proposed, offering different functionalities, applications provided by available tools, their efficiency etc. Among them, two graph-based models are particularly widely used: variation graphs and de Bruijn graphs. In the current paper we propose an axiomatization of the desirable properties of a graph representation of a collection of strings. We show the relationship between variation graphs satisfying these criteria and de Bruijn graphs. This relationship can be used to efficiently build a variation graph representing a given set of genomes, transfer annotations between both models or compare the results of analyzes based on each model.
The need to include the genetic variation within a population into a reference genome led to the concept of a genome sequence graph. Nodes of such a graph are labeled with DNA sequences occurring in represented genomes. Due to double-stranded nature of DNA, each node may be oriented in one of two possible ways, resulting in marking one end of the labeling sequence as in-side and the other as out-side. Edges join pairs of sides and reflect adjacency between node sequences in genomes constituting the graph. Linearization of a sequence graph aims at orienting and ordering graph nodes in a way that makes it more efficient for visualization and further analysis, e.g. access and traversal. We propose a new linearization algorithm, called ALIBI - Algorithm for Linearization by Incremental graph BuIlding. The evaluation shows that ALIBI is computationally very efficient and generates high-quality results.
DNA double-strand breaks (DSBs), are a major threat to genomic stability and may lead to cancer. Several technologies to accurately detect DSBs genome-wide have been developed recently, but still lacking publicly available tools for analysis of the resulting data. Here, we present a step-by-step iSeq package ( http://breakome.utmb.edu/software.html ), custom designed for analysis and interpretation of DSB-sequencing data. iSeq performs barcode trimming and read counting, and identifies DSB-enriched regions by statistical test and annotate them to the desired genomic features. Applying this package, users can identify and annotate DSB-enriched regions from base pair (eg. Cas9 cleavage sites) up to megabase (eg. DNA replication stress-induced) resolution, and if possible quantify DSB frequencies per cell genome-wide by combining with qDSB-Seq. iSeq can be used for any sequencing-based DSB detection techniques. The analysis for Steps 1-19 can be performed within ~4 hours.
Background: The term pan-genome was proposed to denominate collections of genomic sequences jointly analyzed or used as a reference. The constant growth of genomic data intensifies development of data structures and algorithms to investigate pan-genomes efficiently. Results: This work focuses on providing a tool for discovering and visualizing the relationships between the sequences constituting a pan-genome. A new structure to represent such relationships - called affinity tree - is proposed. Each node of this tree has assigned a subset of genomes, as well as their homogeneity level and averaged consensus sequence. Moreover, subsets assigned to sibling nodes form a partition of the genomes assigned to their parent. Conclusions: Functionality of affinity tree is demonstrated on simulated data and on the Ebola virus pan-genome. Furthermore, two software packages are provided: PangTreeBuild constructs affinity tree, while PangTreeVis presents its result.
DNA double-strand breaks (DSBs) are among the most lethal types of DNA damage and frequently cause genome instability. Sequencing-based methods for mapping DSBs have been developed but they allow measurement only of relative frequencies of DSBs between loci, which limits our understanding of the physiological relevance of detected DSBs. Here we propose quantitative DSB sequencing (qDSB-Seq), a method providing both DSB frequencies per cell and their precise genomic coordinates. We induce spike-in DSBs by a site-specific endonuclease and use them to quantify detected DSBs (labeled, e.g., using i-BLESS). Utilizing qDSB-Seq, we determine numbers of DSBs induced by a radiomimetic drug and replication stress, and reveal two orders of magnitude differences in DSB frequencies. We also measure absolute frequencies of Top1-dependent DSBs at natural replication fork barriers. qDSB-Seq is compatible with various DSB labeling methods in different organisms and allows accurate comparisons of absolute DSB frequencies across samples.
Arabidopsis thaliana contains two nuclear XRN2/3 5'-3' exonucleases that are homologs of yeast and human Rat1/Xrn2 proteins involved in the processing and degradation of several classes of nuclear RNAs and in transcription termination of RNA polymerase II. Using strand-specific short read sequencing we show that knockdown of XRN3 leads to an altered expression of hundreds of genes and the accumulation of uncapped and polyadenylated read-through transcripts generated by inefficiently terminated Pol II. Our data support the notion that XRN3-mediated changes in the expression of a subset of genes are caused by upstream read-through transcription and these effects are enhanced by RNA-mRNA chimeras generated in xrn3 plants. In turn, read-through transcripts that are antisense to downstream genes may trigger production of siRNA. Our results highlight the importance of XRN3 exoribonuclease in Pol II transcription termination in plants and show that disturbance in this process may significantly alter gene expression.
Double-strand breaks (DSBs) are extremely detrimental DNA lesions that can lead to cancer-driving mutations and translocations. Non-homologous end joining (NHEJ) and homologous recombination (HR) represent the two main repair pathways operating in the context of chromatin to ensure genome stability. Despite extensive efforts, our knowledge of DSB-induced chromatin still remains fragmented. Here, we describe the distribution of 20 chromatin features at multiple DSBs spread throughout the human genome using ChIP-seq. We provide the most comprehensive picture of the chromatin landscape set up at DSBs and identify NHEJ- and HR-specific chromatin events. This study revealed the existence of a DSB-induced monoubiquitination-to-acetylation switch on histone H2B lysine 120, likely mediated by the SAGA complex, as well as higher-order signaling at HR-repaired DSBs whereby histone H1 is evicted while ubiquitin and 53BP1 accumulate over the entire γH2AX domains.
DNA double-strand breaks (DSBs) can be detected by label-based sequencing or pulsed-field gel electrophoresis (PFGE). Sequencing yields population-average DSB frequencies genome-wide, while PFGE reveals percentages of broken chromosomes. We constructed a mathematical framework to combine advantages of both: high-resolution DSB locations and their population distribution. We also use sequencing read patterns to identify replication-induced DSBs and active replication origins. We describe changes in spatiotemporal replication program upon hydroxyurea-induced replication stress. We found that one-ended DSBs, resulting from collapsed replication forks, are population-representative, while majority of two-ended DSBs (79-100%) are not. To study replication fork collapse, we used strains lacking the checkpoint protein Mec1 and the endonuclease Mus81 and quantified that 19% and 13% of hydroxyurea-induced one-ended DSBs are Mec1-and Mus81-dependent, respectively. We also clarified that Mus81-induced one-ended DSBs are Mec1-dependent.
Emerging genome-wide methods for mapping DNA double-strand breaks (DSBs) by sequencing (e.g. BLESS) are limited to measuring relative frequencies of breaks between loci. Knowing the absolute DSB frequency per cell, however, is key to understanding their physiological relevance. Here, we propose quantitative DSB sequencing (qDSB-Seq), a method to infer the absolute DSB frequency genome-wide. qDSB-Seq relies on inducing spike-in DSBs by a site-specific endonuclease and estimating the efficiency of the endonuclease cleavage by sequencing or PCR. This spike-in frequency is used to quantify DSB sequencing data. We present validation of the qDSB-Seq method and results of its application. We quantify DSBs resulting from replication stress and the collapse of replication forks on natural fork barriers in the ribosomal DNA. The qDSB-Seq approach can be used with any DSB sequencing method and allows accurate comparisons of absolute DSB frequencies across samples and precise quantification of the impact of various DSB-causing agents.
Hematopoietic stem and progenitor cells (HSPCs) are vulnerable to endogenous damage and defects in DNA repair can limit their function. The 2 single-stranded DNA (ssDNA) binding proteins SSB1 and SSB2 are crucial regulators of the DNA damage response; however, their overlapping roles during normal physiology are incompletely understood. We generated mice in which both Ssb1 and Ssb2 were constitutively or conditionally deleted. Constitutive Ssb1/Ssb2 double knockout (DKO) caused early embryonic lethality, whereas conditional Ssb1/Ssb2 double knockout (cDKO) in adult mice resulted in acute lethality due to bone marrow failure and intestinal atrophy featuring stem and progenitor cell depletion, a phenotype unexpected from the previously reported single knockout models of Ssb1 or Ssb2. Mechanistically, cDKO HSPCs showed altered replication fork dynamics, massive accumulation of DNA damage, genome-wide doublestrand breaks enriched at Ssb-binding regions and CpG islands, together with the accumulation of R-loops and cytosolic ssDNA. Transcriptional profiling of cDKO HSPCs revealed the activation of p53 and interferon (IFN) pathways, which enforced cell cycling in quiescent HSPCs, resulting in their apoptotic death. The rapid cell death phenotype was reproducible in in vitro cultured cDKO-hematopoietic stem cells, which were significantly rescued by nucleotide supplementation or after depletion of p53. Collectively, Ssb1 and Ssb2 control crucial aspects of HSPC function, including proliferation and survival in vivo by resolving replicative stress to maintain genomic stability.
Sequencing-based methods for mapping DNA double-strand breaks (DSBs) allow measurement only of relative frequencies of DSBs between loci, which limits our understanding of the physiological relevance of detected DSBs. We propose quantitative DSB sequencing (qDSB-Seq), a method providing both DSB frequencies per cell and their precise genomic coordinates. We induced spike-in DSBs by a site-specific endonuclease and used them to quantify labeled DSBs (e.g. using i-BLESS). Utilizing qDSB-Seq, we determined numbers of DSBs induced by a radiomimetic drug and various forms of replication stress, and revealed several orders of magnitude differences in DSB frequencies. We also measured for the first time Top1-dependent absolute DSB frequencies at replication fork barriers. qDSB-Seq is compatible with various DSB labeling methods in different organisms and allows accurate comparisons of absolute DSB frequencies across samples.
The current paper addresses two problems observed in structure learning applications to computational biology.
BACKGROUND:For many years now, binding preferences of Transcription Factors have been described by so called motifs, usually mathematically defined by position weight matrices or similar models, for the purpose of predicting potential binding sites. However, despite the availability of thousands of motif models in public and commercial databases, a researcher who wants to use them is left with many competing methods of identifying potential binding sites in a genome of interest and there is little published information regarding the optimality of different choices. Thanks to the availability of large number of different motif models as well as a number of experimental datasets describing actual binding of TFs in hundreds of TF-ChIP-seq pairs, we set out to perform a comprehensive analysis of this matter.RESULTS:We focus on the task of identifying potential transcription factor binding sites in the human genome. Firstly, we provide a comprehensive comparison of the coverage and quality of models available in different databases, showing that the public databases have comparable TFs coverage and better motif performance than commercial databases. Secondly, we compare different motif scanners showing that, regardless of the database used, the tools developed by the scientific community outperform the commercial tools. Thirdly, we calculate for each motif a detection threshold optimizing the accuracy of prediction. Finally, we provide an in-depth comparison of different methods of choosing thresholds for all motifs a priori. Surprisingly, we show that selecting a common false-positive rate gives results that are the least biased by the information content of the motif and therefore most uniformly accurate.CONCLUSION:We provide a guide for researchers working with transcription factor motifs. It is supplemented with detailed results of the analysis and the benchmark datasets at http://bioputer.mimuw.edu.pl/papers/motifs/ .
Jerzy Tiuryn合作论文数Faculty of Mathematics, Informatics and Mechanics, University of Warsaw3