Abstract Cancers maintain their telomeres through two telomere maintenance mechanisms: 85-90% of cancers rely on telomerase (TEL+), while 10-15% of cancers adopt the Alternative Lengthening of Telomeres (ALT) pathway. The Break-Induced Replication (BIR) pathway plays a critical role in maintaining telomere length in the ALT+ cells. In both yeast and human, PIF1, a 5’ to 3’ helicase, is required for the robust activity of BIR. However, the extent of human PIF1 (hPIF1) involvement in the ALT pathway remains unknown. Here we showed that hPIF1 can be recruited to damaged telomeres in ALT+ cells. In addition, we demonstrated that inhibition of hPIF1 induced DNA damage and G quadruplex (G4) accumulation at ALT telomeres, leading to a moderate reduction of the mean telomere length. Most interestingly, we demonstrated that inhibition of hPIF1 also attenuates checkpoint activation, BLM recruitment, single-stranded DNA (ssDNA) formation, DNA damage, and G4s at telomeres in the FANCM deficient ALT+ cells. Finally, we showed that inactivation of hPIF1 affects the viability of both ALT+ and TEL+ cancers, suggesting that hPIF1 is a potential drug target for cancer therapy.
Cancers maintain their telomeres through two main telomere maintenance mechanisms (TMMs): 85-90% of cancers rely on telomerase, while 10-15% of cancers adopt the Alternative Lengthening of Telomere (ALT) pathway. Previously, we and others reported that FANCM, one of the Fanconi Anemia proteins, plays a critical role in suppressing replication stress and DNA damage at ALT telomeres by actively disrupting TERRA R-loops [1-4]. Here, we showed that inactivation of DNA2 in ALT-positive (ALT+) cells, but not in telomerase-positive (TEL+) cells, induces a robust increase of replication stress and DNA damage at telomeres, which leads to a pronounced increase of many ALT properties, including telomere dysfunction-induced foci (TIFs), ALT-associated PML bodies (APBs), and C-circles. We further demonstrated that depletion of DNA2 induces a pronounced increase of TERRA R-loops and a decrease in replication efficiency at ALT telomeres. Most importantly, we uncovered a strong additive genetic interaction between DNA2 and FANCM in the ALT pathway. Furthermore, co-depletion of DNA2 and FANCM causes synthetic lethality in ALT+ cells, but not in TEL+ cells, suggesting that targeting DNA2 and FANCM could be a viable strategy to treat ALT+ cancers. Finally, utilizing the single-molecule telomere assay via optical mapping (SMTA-OM) technology, we thoroughly characterized genome-wide changes in DNA2 deficient cells and FANCM deficient cells and found that most chromosome arms manifested increased telomere length. Unexpectedly, we uncovered many chromosome arm-specific telomere changes in those cells, suggesting that telomeres at different chromosome arms may regulate and respond to replication stress differently. Collectively, our study not only shed new light on the molecular mechanisms of the ALT pathway, but also discovered a new strategy for targeting ALT+ cancer.
Chromosome instability (CIN) is frequently observed in many tumors. The breakage-fusion-bridge (BFB) cycle has been proposed to be one of the main drivers of CIN during tumorigenesis and tumor evolution. However, the detailed mechanisms for the individual steps of the BFB cycle warrants further investigation. Here, we demonstrated that a nuclease-dead Cas9 (dCas9) coupled with a telomere-specific single-guide RNA (sgTelo) can be used to model the BFB cycle. First, we showed that targeting dCas9 to telomeres using sgTelo impeded DNA replication at telomeres and induced a pronounced increase of replication stress and DNA damage. Using Single-Molecule Telomere Assay via Optical Mapping (SMTA-OM), we investigated the genome-wide features of telomeres in the dCas9/sgTelo cells and observed a dramatic increase of chromosome end fusions, including fusion/ITS+ and fusion/ITS-.Consistently, we also observed an increase in the formation of dicentric chromosomes, anaphase bridges, and intercellular telomeric chromosome bridges (ITCBs). Utilizing the dCas9/sgTelo system, we uncovered many novel molecular and structural features of the ITCB and demonstrated that multiple DNA repair pathways are implicated in the formation of ITCBs. Our studies shed new light on the molecular mechanisms of the BFB cycle, which will advance our understanding of tumorigenesis, tumor evolution, and drug resistance.
In this report, we present OLAF-Seq, a novel strategy to construct a long-read sequencing library such that adjacent fragments are linked with end-terminal duplications. We use the CRISPR-Cas9 nickase enzyme and a pool of multiple sgRNAs to perform non-random fragmentation of targeted long DNA molecules (> 300kb) into smaller library-sized fragments (about 20 kbp) in a manner so as to retain physical linkage information (up to 1000 bp) between adjacent fragments. DNA molecules targeted for fragmentation are preferentially ligated with adaptors for sequencing, so this method can enrich targeted regions while taking advantage of the long-read sequencing platforms. This enables the sequencing of target regions with significantly lower total coverage, and the genome sequence within linker regions provides information for assembly and phasing. We demonstrated the validity and efficacy of the method first using phage and then by sequencing a panel of 100 full-length cancer-related genes (including both exons and introns) in the human genome. When the designed linkers contained heterozygous genetic variants, long haplotypes could be established. This sequencing strategy can be readily applied in both PacBio and Oxford Nanopore platforms for both long and short genes with an easy protocol. This economically viable approach is useful for targeted enrichment of hundreds of target genomic regions and where long no-gap contigs need deep sequencing.
Telomeres play an essential role in protecting the ends of linear chromosomes and maintaining the integrity of the human genome. One of the key hallmarks of cancers is their replicative immortality. As many as 85-90% of cancers activate the expression of telomerase (TEL+) as the telomere maintenance mechanism (TMM), and 10-15% of cancers utilize the homology-dependent repair (HDR)-based Alternative Lengthening of Telomere (ALT+) pathway. Here, we performed statistical analysis of our previously reported telomere profiling results from Single Molecule Telomere Assay via Optical Mapping (SMTA-OM), which is capable of quantifying individual telomeres from single molecules across all chromosomes. By comparing the telomeric features from SMTA-OM in TEL+ and ALT+ cancer cells, we demonstrated that ALT+ cancer cells display certain unique telomeric profiles, including increased fusions/internal telomere-like sequence (ITS+), fusions/internal telomere-like sequence loss (ITS-), telomere-free ends (TFE), super-long telomeres, and telomere length heterogeneity, compared to TEL+ cancer cells. Therefore, we propose that ALT+ cancer cells can be differentiated from TEL+ cancer cells using the SMTA-OM readouts as biomarkers. In addition, we observed variations in SMTA-OM readouts between different ALT+ cell lines that may potentially be used as biomarkers for discerning subtypes of ALT+ cancer and monitoring the response to cancer therapy.
ABSTRACT In this report, we present linked-pair sequencing, a novel strategy to construct a long-read sequencing library such that adjacent fragments are linked with end-terminal duplications. We use the CRISPR-Cas9 nickase enzyme and a pool of multiple sgRNAs to perform non-random fragmentation of targeted long DNA molecules (>300kb) into smaller library-sized fragments (about 20 kbp) in a manner so as to retain physical linkage information (up to 1000 bp) between adjacent fragments. DNA molecules targeted for fragmentation are preferentially ligated with adaptors for sequencing, so this method can enrich targeted regions while taking advantage of the long-read sequencing platforms. This enables the sequencing of target regions with significantly lower total coverage, and the genome sequence within linker regions provides information for assembly and phasing. We demonstrated the validity and efficacy of the method first using phage and then by sequencing a panel of 100 full-length cancer-related genes (including both exons and introns) in the human genome. When the designed linkers contained heterozygous genetic variants, long haplotypes could be established. This sequencing strategy can be readily applied in both PacBio and Oxford Nanopore platforms. This economically viable approach is useful for targeted enrichment of hundreds of target genomic regions and where long no-gap contigs need deep sequencing.
Identification of structural variants (SVs) breakpoints is important in studying mutations, mutagenic causes, and functional impacts. Next-generation sequencing and whole-genome optical mapping are extensively used in SV discovery and characterization. However, multiple platforms and computational approaches are needed for comprehensive analysis, making it resource-intensive and expensive. Here, we propose a strategy combining optical mapping and cas9-assisted targeted nanopore sequencing to analyze SVs. Optical mapping can economically and quickly detect SVs across a whole genome but does not provide sequence-level information or precisely resolve breakpoints. Furthermore, since only a subset of all SVs is known to affect biology, we attempted to type a subset of all SVs using targeted nanopore sequencing. Using our approach, we resolved the breakpoints of five deletions, five insertions, and an inversion, in a single experiment.
Long interspersed nuclear elements (LINEs) are retrotransposons that contribute to genetic variation in the human genome. LINE-1 elements in larger-scale studies are challenging to identify using sequencing technologies due to cost and scalability. We developed an approach using optical mapping for detection of full-length LINE-1 insertions and 10× sequencing for confirmation. We found 51 true positive full-length LINE-1 insertions, of which 4 are novel insertions, in NA12878. Repeating our analysis on a larger sample set representing 26 populations, we identified 329 full-length LINE-1 elements, of which 123 are novel. 24.8% of these 329 LINE-1 insertions were shared amongst all 5 superpopulations (AFR, AMR, EUR, EAS, SAS). The African superpopulation has a higher percentage of population-specific LINE-1 insertions than any other superpopulation. These data indicate that our approach can provide high-speed, cost-effective, and increased accuracy for LINE-1 detection. These data also provide an insight into variations of LINE-1 elements between different populations.
CRISPR Cas9 has been widely applied as a molecular tool to produce information from specific regions of the human genome. Combining Cas9 with cutting-edge whole genome techniques is helping tackle complex genomic regions. Most extant DNA sequencing and mapping methods are agnostic to large repetitive regions and structural variations requiring specifically designed protocols to extract sequence information in these cases. Cas9 is being used to solve many of these problems giving rise to methods that can efficiently target specific genomic regions. In whole genome mapping, conventional fluorescent labeling of DNA molecules is limited to specific repeat sequences in the human genome. In complex regions, high-resolution DNA mapping becomes necessary, and this becomes challenging given the lack of control in the targeting. Cas9-based mapping enables efficient targeting and complements conventional labeling techniques. This chapter will review and discuss applications of Cas9 in genome research, but more specifically in whole genome mapping, telomere characterization, and sequencing. The Cas9-enabled telomere characterization method is described in detail and the results from studying aging, telomerase-positive, and ALT-positive cells are enumerated and discussed. The chapter also summarizes recent progresses in adapting Cas9 to DNA sequencing methods.
Analysis of structural variations (SVs) is important to understand mutations underlying genetic disorders and pathogenic conditions. However, characterizing SVs using short-read, high-throughput sequencing technology is difficult. Although long-read sequencing technologies are being increasingly employed in characterizing SVs, their low throughput and high costs discourage widespread adoption. Sequence motif-based optical mapping in nanochannels is useful in whole-genome mapping and SV detection, but it is not possible to precisely locate the breakpoints or estimate the copy numbers. We present here a universal multicolor mapping strategy in nanochannels combining conventional sequence-motif labeling system with Cas9-mediated target-specific labeling of any 20-base sequences (20mers) to create custom labels and detect new features. The sequence motifs are labeled with green fluorophores and the 20mers are labeled with red fluorophores. Using this strategy, it is possible to not only detect the SVs but also utilize custom labels to interrogate the features not accessible to motif-labeling, locate breakpoints, and precisely estimate copy numbers of genomic repeats. We validated our approach by quantifying the D4Z4 copy numbers, a known biomarker for facioscapulohumeral muscular dystrophy (FSHD) and estimating the telomere length, a clinical biomarker for assessing disease risk factors in aging-related diseases and malignant cancers. We also demonstrate the application of our methodology in discovering transposable long non-interspersed Elements 1 (LINE-1) insertions across the whole genome.
Genomic regions of high segmental duplication content and/or structural variation have led to gaps and misassemblies in the human reference sequence, and are refractory to assembly from whole-genome short-read datasets. Human subtelomere regions are highly enriched in both segmental duplication content and structural variations, and as a consequence are both impossible to assemble accurately and highly variable from individual to individual. Recently, we developed a pipeline for improved region-specific assembly called Regional Extension of Assemblies Using Linked-Reads (REXTAL). In this study, we evaluate REXTAL and genome-wide assembly (Supernova) approaches on 10X Genomics linked-reads data sets partitioned and barcoded using the Gel Bead in Emulsion (GEM) microfluidic method. Our results describe the accuracy and relative performance of these two approaches using the reference-based assessment module of QUAST. We show that REXTAL dramatically outperforms the Supernova whole genome assembler in subtelomeric segmental duplication regions, and results in highly accurate assemblies. Nearly all of the REXTAL "misassemblies" identified using default QUAST parameters simply pinpoint locations of tandem repeat arrays in the reference sequence where the repeat array length differs from that in the cognate REXTAL assembly by 1000 bp.
Membrane Attack Complex and Perforin (MACPF) proteins play crucial roles in plant development and plant responses to environmental stresses. To date, only fourMACPFgenes have been identified inArabidopsis thaliana, and the functions of theMACPFgene family members in other plants, especially in important crop plants, such as the Poaceae family, remain largely unknown. In this study, we identified and analyzed 42MACPFgenes from six completely sequenced and well annotated species representing the major Poaceae clades. A phylogenetic analysis ofMACPFgenes resolved four groups, characterized by shared motif organizations and gene structures within each group.MACPFgenes were unevenly distributed along the Poaceae chromosomes. Moreover, segmental duplications and dispersed duplication events may have played significant roles duringMACPFgene family expansion and functional diversification in the Poaceae. In addition, phylogenomic synteny analysis revealed a high degree of conservation among the PoaceaeMACPFgenes. In particular, Group I, II, and IIIMACPFgenes were exposed to strong purifying selection with different evolutionary rates. Temporal and spatial expression analyses suggested that Group IIIMACPFgenes were highly expressed relative to the other groups. In addition, mostMACPFgenes were highly expressed in vegetative tissues and up-regulated by several biotic and abiotic stresses. Taken together, these findings provide valuable information for further functional characterization and phenotypic validation of the PoaceaeMACPFgene family.
In humans, the telomere consists of tandem 5 ' TTAGGG3 ' DNA repeats on both ends of all 46 chromosomes. Telomere shortening has been linked to aging and age-related diseases. Similarly, telomere length changes have been associated with chemical exposure, molecular-level DNA damage, and tumor development. Telomere elongation has been associated to tumor development, caused due to chemical exposure and molecular-level DNA damage. The methods used to study these effects mostly rely on average telomere length as a biomarker. The mechanisms regulating subtelomere-specific and haplotype-specific telomere lengths in humans remain understudied and poorly understood, primarily because of technical limitations in obtaining these data for all chromosomes. Recent studies have shown that it is the short telomeres that are crucial in preserving chromosome stability. The identity and frequency of specific critically short telomeres potentially is a useful biomarker for studying aging, age-related diseases, and cancer. Here, we will briefly review the role of telomere length, its measurement, and our recent single-molecule telomere length measurement assay. With this assay, one can measure individual telomere lengths as well as identify their physically linked subtelomeric DNA. This assay can also positively detect telomere loss, characterize novel subtelomeric variants, haplotypes, and previously uncharacterized recombined subtelomeres. We will also discuss its applications in aging cells and cancer cells, highlighting the utility of the single molecule telomere length assay.
Background Telomeric DNA is typically comprised of G-rich tandem repeat motifs and maintained by telomerase (Greider CW, Blackburn EH; Cell 51:887–898; 1987). In eukaryotes lacking telomerase, a variety of DNA repair and DNA recombination based pathways for telomere maintenance have evolved in organisms normally dependent upon telomerase for telomere elongation (Webb CJ, Wu Y, Zakian VA; Cold Spring Harb Perspect Biol 5:a012666; 2013); collectively called Alternative Lengthening of Telomeres (ALT) pathways. By measuring (TTAGGG) n tract lengths from the same large DNA molecules that were optically mapped, we simultaneously analyzed telomere length dynamics and subtelomere-linked structural changes at a large number of specific subtelomeric loci in the ALT-positive cell lines U2OS, SK-MEL-2 and Saos-2. Results Our results revealed loci-specific ALT telomere features. For example, while each subtelomere included examples of single molecules with terminal (TTAGGG) n tracts as well as examples of recombinant telomeric single molecules, the ratio of these molecules was subtelomere-specific, ranging from 33:1 (19p) to 1:25 (19q) in U2OS. The Saos-2 cell line shows a similar percentage of recombinant telomeres. The frequency of recombinant subtelomeres of SK-MEL-2 (11%) is about half that of U2OS and Saos-2 (24 and 19% respectively). Terminal (TTAGGG) n tract lengths and heterogeneity levels, the frequencies of telomere signal-free ends, and the frequency and size of retained internal telomere-like sequences (ITSs) at recombinant telomere fusion junctions all varied according to the specific subtelomere involved in a particular cell line. Very large linear extrachromosomal telomere repeat (ECTR) DNA molecules were found in all three cell lines; these are in principle capable of templating synthesis of new long telomere tracts via break-induced repair (BIR) long-tract DNA synthesis mechanisms and contributing to the very long telomere tract length and heterogeneity characteristic of ALT cells. Many of longest telomere tracts (both end-telomeres and linear ECTRs) displayed punctate CRISPR/Cas9-dependent (TTAGGG) n labeling patterns indicative of interspersion of stretches of non-canonical telomere repeats. Conclusion Identifying individual subtelomeres and characterizing linked telomere (TTAGGG) n tract lengths and structural changes using our new single-molecule methodologies reveals the structural consequences of telomere damage, repair and recombination mechanisms in human ALT cells in unprecedented molecular detail and significant differences in different ALT-positive cell lines.
The current human reference genome is predominantly derived from a single individual and it does not adequately reflect human genetic diversity. Here, we analyze 338 high-quality human assemblies of genetically divergent human populations to identify missing sequences in the human reference genome with breakpoint resolution. We identify 127,727 recurrent non-reference unique insertions spanning 18,048,877 bp, some of which disrupt exons and known regulatory elements. To improve genome annotations, we linearly integrate these sequences into the chromosomal assemblies and construct a Human Diversity Reference. Leveraging this reference, an average of 402,573 previously unmapped reads can be recovered for a given genome sequenced to ~40X coverage. Transcriptomic diversity among these non-reference sequences can also be directly assessed. We successfully map tens of thousands of previously discarded RNA-Seq reads to this reference and identify transcription evidence in 4781 gene loci, underlining the importance of these non-reference sequences in functional genomics. Our extensive datasets are important advances toward a comprehensive reference representation of global human genetic diversity.
Multidrug and Toxic Compound Extrusion (MATE) proteins are essential transporters that extrude metabolites and participate in plant development and the detoxification of toxins. Little is known about the MATE gene family in the Solanaceae, which includes species that produce a broad range of specialized metabolites. Here, we identified and analyzed the complement of MATE genes in pepper (Capsicum annuum) and potato (Solanum tuberosum). We classified all MATE genes into five groups based on their phylogenetic relationships and their gene and protein structures. Moreover, we discovered that tandem duplication contributed significantly to the expansion of the pepper MATE family, while both tandem and segmental duplications contributed to the expansion of the potato MATE family, indicating that MATEs took distinct evolutionary paths in these two Solanaceous species. Analysis of ω values showed that all potato and pepper MATE genes experienced purifying selection during evolution. In addition, collinearity analysis showed that MATE genes were highly conserved between pepper and potato. Analysis of cis-elements in MATE promoters and MATE expression patterns revealed that MATE proteins likely function in many stages of plant development, especially during fruit ripening, and when exposed to multiple stresses, consistent with the existence of functional differentiation between duplicated MATE genes. Together, our results lay the foundation for further characterization of pepper and potato MATE gene family members.
Detailed comprehensive knowledge of the structures of individual long-range telomere-terminal haplotypes are needed to understand their impact on telomere function, and to delineate the population structure and evolution of subtelomere regions. However, the abundance of large evolutionarily recent segmental duplications and high levels of large structural variations have complicated both the mapping and sequence characterization of human subtelomere regions. Here, we use high throughput optical mapping of large single DNA molecules in nanochannel arrays for 154 human genomes from 26 populations to present a comprehensive look at human subtelomere structure and variation. The results catalog many novel long-range subtelomere haplotypes and determine the frequencies and contexts of specific subtelomeric duplicons on each chromosome arm, helping to clarify the currently ambiguous nature of many specific subtelomere structures as represented in the current reference sequence (HG38). The organization and content of some duplicons in subtelomeres appear to show both chromosome arm and population-specific trends. Based upon these trends we estimate a timeline for the spread of these duplication blocks. Author Summary The ends of human chromosomes have caps called telomeres that are essential. These telomeres are influenced by the portions of DNA next to them, a region known as the subtelomere. We need to better understand the subtelomeric region to understand how it impacts the telomeres. This subtelomeric region is not well described in the current references. This is due to large variations in this region and portions that are repeated many times, making current sequencing technologies struggle to capture these regions. Many of these variations are evolutionary recent. Here we use 154 different samples from the 26 geographic regions of the world to gain a better understanding of the variation in these regions. We found many new haplotypes and clarified the haplotypes existing in the current reference. We then examined population and chromosome specific trends.
The most prevalent microdeletion in humans occurs at 22q11.2, a region rich in chromosome-specific low copy repeats (LCR22s). The structure of this region has defied elucidation due to its size, regional complexity, and haplotype diversity, and is not well represented in the human genome reference. Most individuals with 22q11.2 deletion syndrome (22q11.2DS) carry a de novo hemizygous deletion of ~ 3 Mbp occurring by non-allelic homologous recombination (NAHR) mediated by LCR22s. In this study, optical mapping has been used to elucidate LCR22 structure and variation in 88 individuals in thirty 22q11.2DS families to uncover potential risk factors for germline rearrangements leading to 22q11.2DS offspring. Families were optically mapped to characterize LCR22 structures, NAHR locations, and genomic signatures associated with the deletion. Bioinformatics analyses revealed clear delineations between LCR22 structures in normal and deletion-containing haplotypes. Despite no explicit whole-haplotype predisposing configurations being identified, all NAHR events contain a segmental duplication encompassing FAM230 gene members suggesting preferred recombination sequences. Analysis of deletion breakpoints indicates that preferred recombinations occur between FAM230 and specific segmental duplication orientations within LCR22A and LCR22D, ultimately leading to NAHR. This work represents the most comprehensive analysis of 22q11.2DS NAHR events demonstrating completely contiguous LCR22 structures surrounding and within deletion breakpoints.
Olduvai (formerly DUF1220) protein domains have undergone the largest human-specific increase in copy number of any coding region in the genome (∼300 copies of which 165 are human-specific) and have been implicated in human brain evolution... Sequences encoding Olduvai protein domains (formerly DUF1220) show the greatest human lineage-specific increase in copy number of any coding region in the genome and have been associated, in a dosage-dependent manner, with brain size, cognitive aptitude, autism, and schizophrenia. Tandem intragenic duplications of a three-domain block, termed the Olduvai triplet, in four NBPF genes in the chromosomal 1q21.1-0.2 region, are primarily responsible for the striking human-specific copy number increase. Interestingly, most of the Olduvai triplets are adjacent to, and transcriptionally coregulated with, three human-specific NOTCH2NL genes that have been shown to promote cortical neurogenesis. Until now, the underlying genomic events that drove the Olduvai hyperamplification in humans have remained unexplained. Here, we show that the presence or absence of an alternative first exon of the Olduvai triplet perfectly discriminates between amplified (58/58) and unamplified (0/12) triplets. We provide sequence and breakpoint analyses that suggest the alternative exon was produced by an nonallelic homologous recombination-based mechanism involving the duplicative transposition of an existing Olduvai exon found in the CON3 domain, which typically occurs at the C-terminal end of NBPF genes. We also provide suggestive in vitro evidence that the alternative exon may promote instability through a putative G-quadraplex (pG4)-based mechanism. Lastly, we use single-molecule optical mapping to characterize the intragenic structural variation observed in NBPF genes in 154 unrelated individuals and 52 related individuals from 16 families and show that the presence of pG4-containing Olduvai triplets is strongly correlated with high levels of Olduvai copy number variation. These results suggest that the same driver of genomic instability that allowed the evolutionarily recent, rapid, and extreme human-specific Olduvai expansion remains highly active in the human genome.