RNA-guided CRISPR-Cas9 endonucleases are widely used for genome engineering, but our understanding of Cas9 specificity remains incomplete. Here, we developed a biochemical method (SITE-Seq), using Cas9 programmed with single-guide RNAs (sgRNAs), to identify the sequence of cut sites within genomic DNA. Cells edited with the same Cas9-sgRNA complexes are then assayed for mutations at each cut site using amplicon sequencing. We used SITE-Seq to examine Cas9 specificity with sgRNAs targeting the human genome. The number of sites identified depended on sgRNA sequence and nuclease concentration. Sites identified at lower concentrations showed a higher propensity for off-target mutations in cells. The list of off-target sites showing activity in cells was influenced by sgRNP delivery, cell type and duration of exposure to the nuclease. Collectively, our results underscore the utility of combining comprehensive biochemical identification of off-target sites with independent cell-based measurements of activity at those sites when assessing nuclease activity and specificity.
We describe a microfluidics-based strategy for genome-wide analysis of multiplex chromatin interactions with single-molecule precision. In multiplex chromatin interaction analysis (multi-ChIA), individual chromatin complexes are partitioned into droplets that contain a gel bead with unique DNA barcode, in which tethered chromatin DNA fragments are barcoded and amplified for sequencing and mapping to demarcate chromatin contacts. Thus, multi-ChIA has the unprecedented ability to uncover multiplex chromatin interactions at single-molecule level, which has been impossible using previous methods that rely on analyzing pairwise contacts via proximity ligation. We demonstrate that multiplex chromatin interactions predominantly contribute to topologically associated domains, and clusters of gene promoters and enhancers provide a fundamental topological framework for co-transcriptional regulation.
ChIA-Drop is a new experimental method for mapping multiplex chromatin interactions with single-molecule precision by barcoding chromatin complexes inside microfluidics droplets, followed by pooled DNA sequencing. The chromatin DNA reads with the same droplet-specific barcodes are inferred to be derived from the same chromatin interaction complex. Here, we describe an integrated computational pipeline, named ChIA-DropBox, that is specifically designed for reconstructing chromatin reads in each droplet and refining multiplex chromatin complexes from raw ChIA-Drop sequencing reads, and then visualizing the results. First, ChIA-DropBox maps and filters sequencing reads, and then reconstructs the chromatin droplets by parsing the barcode sequences and grouping together chromatin reads with the same barcode. Based on the concept of chromosome territories that most chromatin interactions take place within the same chromosome, potential mixing up of chromatin complexes derived from different chromosomes could be readily identified and separated. Accordingly, ChIA-DropBox refines these chromatin droplets into purely intra-chromosomal chromatin complexes, ready for downstream analysis. For visualization, ChIA-DropBox converts the ChIA-Drop data to pairwise format and automatically generates input files for viewing 2D contact maps in Juicebox and viewing loops in BASIC Browser. Finally, ChIA-DropBox introduces a new browser, named ChIA-View, for interactive visualization of multiplex chromatin interactions.
The spatial organization of the genome influences cellular function, notably gene regulation. Recent studies have assessed the three-dimensional (3D) co-localization of functional annotations (e.g. centromeres, long terminal repeats) using 3D genome reconstructions from Hi-C (genome-wide chromosome conformation capture) data; however, corresponding assessments for continuous functional genomic data (e.g. chromatin immunoprecipitation-sequencing (ChIP-seq) peak height) are lacking. Here, we demonstrate that applying bump hunting via the patient rule induction method (PRIM) to ChIP-seq data superposed on a Saccharomyces cerevisiae 3D genome reconstruction can discover 'functional 3D hotspots', regions in 3-space for which the mean ChIP-seq peak height is significantly elevated. For the transcription factor Swi6, the top hotspot by P-value contains MSB2 and ERG11 - known Swi6 target genes on different chromosomes. We verify this finding in a number of ways. First, this top hotspot is relatively stable under PRIM across parameter settings. Second, this hotspot is among the top hotspots by mean outcome identified by an alternative algorithm, k-Nearest Neighbor (k-NN) regression. Third, the distance between MSB2 and ERG11 is smaller than expected (by resampling) in two other 3D reconstructions generated via different normalization and reconstruction algorithms. This analytic approach can discover functional 3D hotspots and potentially reveal novel regulatory interactions.
The repair outcomes at site-specific DNA double-strand breaks (DSBs) generated by the RNA-guided DNA endonuclease Cas9 determine how gene function is altered. Despite the widespread adoption of CRISPR-Cas9 technology to induce DSBs for genome engineering, the resulting repair products have not been examined in depth. Here, the DNA repair profiles of 223 sites in the human genome demonstrate that the pattern of DNA repair following Cas9 cutting at each site is nonrandom and consistent across experimental replicates, cell lines, and reagent delivery methods. Furthermore, the repair outcomes are determined by the protospacer sequence rather than genomic context, indicating that DNA repair profiling in cell lines can be used to anticipate repair outcomes in primary cells. Chemical inhibition of DNA-PK enabled dissection of the DNA repair profiles into contributions from c-NHEJ and MMEJ. Finally, this work elucidates a strategy for using "error-prone" DNA-repair machinery to generate precise edits.
Repair of a chromosomal double-strand break (DSB) by gene conversion depends on the ability of the broken ends to encounter a donor sequence. To understand how chromosomal location of a target sequence affects DSB repair, we took advantage of genome-wide Hi-C analysis of yeast chromosomes to create a series of strains in which an induced site-specific DSB in budding yeast is repaired by a 2-kb donor sequence inserted at different locations. The efficiency of repair, measured by cell viability or competition between each donor and a reference site, showed a strong correlation (r = 0.85 and 0.79) with the contact frequencies of each donor with the DSB repair site. Repair efficiency depends on the distance between donor and recipient rather than any intrinsic limitation of a particular donor site. These results further demonstrate that the search for homology is the rate-limiting step in DSB repair and suggest that cells often fail to repair a DSB because they cannot locate a donor before other, apparently lethal, processes arise. The repair efficiency of a donor locus can be improved by four factors: slower 5' to 3' resection of the DSB ends, increased abundance of replication protein factor A (RPA), longer shared homology, or presence of a recombination enhancer element adjacent to a donor.
Background Recent studies used the contact data or three-dimensional (3D) genome reconstructions from Hi-C (chromosome conformation capture with next-generation sequencing) to assess the co-localization of functional genomic annotations in the nucleus. These analyses dichotomized data point pairs belonging to a functional annotation as “close” or “far” based on some threshold and then tested for enrichment of “close” pairs. We propose an alternative approach that avoids dichotomization of the data and instead directly estimates the significance of distances within the 3D reconstruction. Results We applied this approach to 3D genome reconstructions for Plasmodium falciparum , the causative agent of malaria, and Saccharomyces cerevisiae and compared the results to previous approaches. We found significant 3D co-localization of centromeres, telomeres, virulence genes, and several sets of genes with developmentally regulated expression in P. falciparum ; and significant 3D co-localization of centromeres and long terminal repeats in S. cerevisiae . Additionally, we tested the experimental observation that telomeres form three to seven clusters in P. falciparum and S. cerevisiae . Applying affinity propagation clustering to telomere coordinates in the 3D reconstructions yielded six telomere clusters for both organisms. Conclusions Distance-based assessment replicated key findings, while avoiding dichotomization of the data (which previously yielded threshold-sensitive results).
It is widely recognized that the three-dimensional (3D) architecture of eukaryotic chromatin plays an important role in processes such as gene regulation and cancer-driving gene fusions. Observing or inferring this 3D structure at even modest resolutions had been problematic, since genomes are highly condensed and traditional assays are coarse. However, recently devised high-throughput molecular techniques have changed this situation. Notably, the development of a suite of chromatin conformation capture (CCC) assays has enabled elicitation of contacts—spatially close chromosomal loci—which have provided insights into chromatin architecture. Most analysis of CCC data has focused on the contact level, with less effort directed toward obtaining 3D reconstructions and evaluating the accuracy and reproducibility thereof. While questions of accuracy must be addressed experimentally, questions of reproducibility can be addressed statistically—the purpose of this paper. We use a constrained optimization technique to reconstruct chromatin configurations for a number of closely related yeast datasets and assess reproducibility using four metrics that measure the distance between 3D configurations. The first of these, Procrustes fitting, measures configuration closeness after applying reflection, rotation, translation, and scaling-based alignment of the structures. The others base comparisons on the within-configuration inter-point distance matrix. Inferential results for these metrics rely on suitable permutation approaches. Results indicate that distance matrix-based approaches are preferable to Procrustes analysis, not because of the metrics per se but rather on account of the ability to customize permutation schemes to handle within-chromosome contiguity. It has recently been emphasized that the use of constrained optimization approaches to 3D architecture reconstruction are prone to being trapped in local minima. Our methods of reproducibility assessment provide a means for comparing 3D reconstruction solutions so that we can discern between local and global optima by contrasting solutions under perturbed inputs.
BACKGROUND:Applying supervised learning/classification techniques to epigenomic data may reveal properties that differentiate histone modifications. Previous analyses sought to classify nucleosomes containing histone H2A/H4 arginine 3 symmetric dimethylation (H2A/H4R3me2s) or H2A.Z using human CD4+ T-cell chromatin immunoprecipitation sequencing (ChIP-Seq) data. However, these efforts only achieved modest accuracy with limited biological interpretation. Here, we investigate the impact of using appropriate data pre-processing -deduplication, normalization, and position- (peak-) finding to identify stable nucleosome positions - in conjunction with advanced classification algorithms, notably discriminatory motif feature selection and random forests. Performance assessments are based on accuracy and interpretative yield.RESULTS:We achieved dramatically improved accuracy using histone modification features (99.0%; previous attempts, 68.3%) and DNA sequence features (94.1%; previous attempts, <60%). Furthermore, the algorithms elicited interpretable features that withstand permutation testing, including: the histone modifications H4K20me3 and H3K9me3, which are components of heterochromatin; and the motif TCCATT, which is part of the consensus sequence of satellite II and III DNA. Downstream analysis demonstrates that satellite II and III DNA in the human genome is occupied by stable nucleosomes containing H2A/H4R3me2s, H4K20me3, and/or H3K9me3, but not 18 other histone methylations. These results are consistent with the recent biochemical finding that H4R3me2s provides a binding site for the DNA methyltransferase (Dnmt3a) that methylates satellite II and III DNA.CONCLUSIONS:Classification algorithms applied to appropriately pre-processed ChIP-Seq data can accurately discriminate between histone modifications. Algorithms that facilitate interpretation, such as discriminatory motif feature selection, have the added potential to impart information about underlying biological mechanism.
Most existing methods for sequence-based classification use exhaustive feature generation, employing, for example, all k-mer patterns. The motivation behind such (enumerative) approaches is to minimize the potential for overlooking important features. However, there are shortcomings to this strategy. First, practical constraints limit the scope of exhaustive feature generation to patterns of length ≤ k, such that potentially important, longer (> k) predictors are not considered. Second, features so generated exhibit strong dependencies, which can complicate understanding of derived classification rules. Third, and most importantly, numerous irrelevant features are created. These concerns can compromise prediction and interpretation. While remedies have been proposed, they tend to be problem-specific and not broadly applicable. Here, we develop a generally applicable methodology, and an attendant software pipeline, that is predicated on discriminatory motif finding. In addition to the traditional training and validation partitions, our framework entails a third level of data partitioning, a discovery partition. A discriminatory motif finder is used on sequences and associated class labels in the discovery partition to yield a (small) set of features. These features are then used as inputs to a classifier in the training partition. Finally, performance assessment occurs on the validation partition. Important attributes of our approach are its modularity (any discriminatory motif finder and any classifier can be deployed) and its universality (all data, including sequences that are unaligned and/or of unequal length, can be accommodated). We illustrate our approach on two nucleosome occupancy datasets and a protein solubility dataset, previously analyzed using enumerative feature generation. Our method achieves excellent performance results, with and without optimization of classifier tuning parameters. A Python pipeline implementing the approach is available at http://www.epibiostat.ucsf.edu/biostat/sen/dmfs/.