Genomic DNA is under constant oxidative damage, with 8-oxo-7,8-dihydro-2'-deoxyguanosine (8-oxo-dG) being the prominent lesion linked to mutagenesis, epigenetics, and gene regulation. Existing methods to detect 8-oxo-dG rely on indirect approaches, while nanopore sequencing enables direct detection of base modifications. A model for 8-oxo-dG detection is currently missing due to the lack of training data. Here, we develop a strategy using synthetic oligos to generate long, 8-oxo-dG context-variable DNA molecules for deep learning and nanopore sequencing. Our training approach addresses the rarity of 8-oxo-dG relative to guanine, enabling specific detection. Applied to a tissue culture model of oxidative damage, our method reveals uneven genomic 8-oxo-dG distribution, dissimilar context pattern to C>A mutations, and local 5-mC depletion. This dual measurement of 5-mC and 8-oxo-dG at single-molecule resolution uncovers new insights into their interplay. Our approach also provides a general framework for detecting other rare DNA modifications using synthetic DNA and nanopore sequencing.
Shallow genome-wide cell-free DNA sequencing holds great promise for noninvasive cancer monitoring by providing reliable copy number alteration (CNA) and fragmentomic profiles. Single-nucleotide variations (SNVs) are, however, much harder to identify with low sequencing depth due to sequencing errors. Here, we present Nanopore Rolling Circle Amplification (RCA)-enhanced Consensus Sequencing (NanoRCS), which leverages RCA and consensus calling based on genome-wide long-read nanopore sequencing to enable simultaneous multimodal tumor fraction (TF) estimation through SNVs, CNAs, and fragmentomics. The efficacy of NanoRCS is tested on 18 cancer patient samples and seven healthy controls, demonstrating its ability to reliably detect TFs as low as 0.24%. In vitro experiments confirm that SNV measurements are essential for detecting TFs below 3%. NanoRCS provides an opportunity for cost-effective and rapid sample processing, which aligns well with clinical needs, particularly in settings where quick and accurate cancer monitoring is essential for personalized treatment strategies.
Sepsis is the leading cause of death in neonatal foals, yet current diagnostics lack sufficient sensitivity and specificity. Here, we present a foal cell-free DNA (cfDNA) sequencing for bacterial identification (cfFBI) workflow that integrates wet-lab and computational protocols, enabling direct bacterial profiling through enrichment of the bacterial cfDNA and minimization of false-positive detections. We applied cfFBI to blood from 25 hospitalized foals and 7 healthy foals (H). Sepsis-associated bacterial genera were elevated in all 11 nSIRSpositive (S+) foals compared to H, and in 8/11 when compared to both nSIRS-negative (nS-) and H, with multiple genera elevated in nearly half (45.5%). While total cfDNA concentration, bacterial fraction, and microbial diversity did not differ between groups, S+ foals showed distinct cfDNA end-motif patterns and reduced mitochondrial cfDNA fractions. These findings indicate that cfDNA sequencing enables the detection of pathogenic bacteria and can help identify additional (host-related) sepsis biomarkers.
DNA methylation is important for establishing and maintaining cell identity and for genomic stability. This is achieved by regulating the accessibility of regulatory and transcriptional elements and the compaction of subtelomeric, centromeric, and other inactive genomic regions. Carcinogenesis is accompanied by a global loss in DNA methylation, which facilitates the transformation of cells. Cancer hypomethylation may also cause genomic instability, for example through interference with the protective function of telomeres and centromeres. However, understanding the role(s) of hypomethylation in tumor evolution is incomplete because the precise mutational consequences of global hypomethylation have thus far not been systematically assessed. Here we made genome-wide inventories of all possible genetic variation that accumulates in single cells upon the long-term global hypomethylation by CRISPR interference-mediated conditional knockdown of DNMT1 . Depletion of DNMT1 resulted in a genomewide reduction in DNA methylation. The degree of DNA methylation loss was similar to that observed in many cancer types. Hypomethylated cells showed reduced proliferation rates, increased transcription of genes, reactivation of the inactive X-chromosome and abnormal nuclear morphologies. Prolonged hypomethylation was accompanied by increased chromosomal instability. However, there was no increase in mutational burden, enrichment for certain mutational signatures or accumulation of structural variation to the genome. In conclusion, the primary consequence of hypomethylation is genomic instability, which in cancer leads to increased tumor heterogeneity and thereby fuels cancer evolution.
Bladder cancer has a high recurrence rate and low survival of advanced stage patients. Few genetic drivers of bladder cancer have thus far been identified. We performed in-depth structural variant analysis on whole-genome sequencing data of 206 metastasized urinary tract cancers. In ~ 10% of the patients, we identified recurrent in-frame deletions of exons 8 and 9 in the aryl hydrocarbon receptor gene ( AHR Δe8-9 ), which codes for a ligand-activated transcription factor. Pan-cancer analyses show that AHR Δe8-9 is highly specific to urinary tract cancer and mutually exclusive with other bladder cancer drivers. The ligand-binding domain of the AHR Δe8-9 protein is disrupted and we show that this results in ligand-independent AHR-pathway activation. In bladder organoids, AHR Δe8-9 induces a transformed phenotype that is characterized by upregulation of AHR target genes, downregulation of differentiation markers and upregulation of genes associated with stemness and urothelial cancer. Furthermore, AHR Δe8-9 expression results in anchorage independent growth of bladder organoids, indicating tumorigenic potential. DNA-binding deficient AHR Δe8-9 fails to induce transformation, suggesting a role for AHR target genes in the acquisition of the oncogenic phenotype. In conclusion, we show that AHR Δe8-9 is a novel driver of urinary tract cancer and that the AHR pathway could be an interesting therapeutic target.
Precision cancer medicine and tailoring therapies to specific molecular targets are based mainly on driver mutations in protein-coding regions of the genome. Complex computational tools were developed to characterize coding driver mutations in 30+ cancer types. However, genomic and functional characterization of noncoding drivers outside coding regions (~98% of the genome) and a systematic understanding of how these events contribute to tumor development is still lacking. While driver mutations in coding regions directly change protein functions (e.g., tonic activation of a receptor), drivers in the noncoding genome can activate genes not expressed in normal tissue, thereby recruiting them for oncogenesis. This suggests that the biology of drivers in coding and noncoding regions differs substantially, and that specific tools will be needed for the accurate interpretation of noncoding mutations in whole-genome sequencing data. Here, we developed a new statistical test that accounts for the biology of noncoding regions, including epigenetic structure, fluctuations of mutation rates, and positional clustering, for the genome-wide analysis of somatic noncoding mutations in more than 3,500 whole cancer genomes across 19 human cancer types. On average, we identified ~5 novel noncoding events per cancer type, including many of the noncoding findings from previous studies (e.g., PCAWG consortium) along with novel observations. Many significant results fell into two categories: (i) highly localized mutagenic processes not observed in the rest of the genome and (ii) significantly mutated promoter and enhancer regions of genes associated with tumor signaling. We observed that findings associated with localized processes translated into distinct transcriptional states in bulk RNA-seq data and single-cell expression profiles. Moreover, they enabled the prediction of tumor cells-of-origin. Many significantly mutated regulatory regions exhibited differential expression of their associated target gene. We validated a subset of these findings, including noncoding mutations altering the activity of the estrogen pathway in breast cancer, by performing CRISPR-interference and luciferase reporter experiments in cancer cell lines. Mutated and non-mutated samples further harbored differential epigenomic profiles, suggesting noncoding mutations directly affected the regulatory activity of their target regions. Many significantly mutated regulatory regions involved known cancer genes, while others targeted genes that had been functionally linked to cancer before but were not considered canonical cancer genes. Broadly, our work reveals that interpreting whole cancer genomes involves challenges specific to noncoding regions. Our extensive catalog of testable hypotheses provides a blueprint for prospective experimental and computational follow-up studies that build on the concepts of our work. Citation Format: Felix Dietlein, Alex B. Wang, Christian Fagre, Anran Tang, Nicolle Besselink, Edwin Cuppen, Chunliang Li, Shamil R. Sunyaev, James T. Neal, Eliezer M. Van Allen. Automated analysis of significant noncoding mutations in somatic whole cancer genomes [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2022; 2022 Apr 8-13. Philadelphia (PA): AACR; Cancer Res 2022;82(12_Suppl):Abstract nr 2170.
We established a genome-wide compendium of somatic mutation events in 3949 whole cancer genomes representing 19 tumor types. Protein-coding events captured well-established drivers. Noncoding events near tissue-specific genes, such as ALB in the liver or KLK3 in the prostate, characterized localized passenger mutation patterns and may reflect tumor-cell-of-origin imprinting. Noncoding events in regulatory promoter and enhancer regions frequently involved cancer-relevant genes such as BCL6 , FGFR2 , RAD51B , SMC6 , TERT , and XBP1 and represent possible drivers. Unlike most noncoding regulatory events, XBP1 mutations primarily accumulated outside the gene’s promoter, and we validated their effect on gene expression using CRISPR-interference screening and luciferase reporter assays. Broadly, our study provides a blueprint for capturing mutation events across the entire genome to guide advances in biological discovery, therapies, and diagnostics.
Accurate detection of somatic structural variation (SV) in cancer genomes remains a challenging problem. This is in part due to the lack of high-quality, gold-standard datasets that enable the benchmarking of experimental approaches and bioinformatic analysis pipelines. Here, we performed somatic SV analysis of the paired melanoma and normal lymphoblastoid COLO829 cell lines using four different sequencing technologies. Based on the evidence from multiple technologies combined with extensive experimental validation, we compiled a comprehensive set of carefully curated and validated somatic SVs, comprising all SV types. We demonstrate the utility of this resource by determining the SV detection performance as a function of tumor purity and sequence depth, highlighting the importance of assessing these parameters in cancer genomics projects. The truth somatic SV dataset as well as the underlying raw multi-platform sequencing data are freely available and are an important resource for community somatic benchmarking efforts.
GRIDSS2 is the first structural variant caller to explicitly report single breakends—breakpoints in which only one side can be unambiguously determined. By treating single breakends as a fundamental genomic rearrangement signal on par with breakpoints, GRIDSS2 can explain 47% of somatic centromere copy number changes using single breakends to non-centromere sequence. On a cohort of 3782 deeply sequenced metastatic cancers, GRIDSS2 achieves an unprecedented 3.1% false negative rate and 3.3% false discovery rate and identifies a novel 32–100 bp duplication signature. GRIDSS2 simplifies complex rearrangement interpretation through phasing of structural variants with 16% of somatic calls phasable using paired-end sequencing.
Inflammatory liver disease increases the risk of developing primary liver cancer. The mechanism through which liver disease induces tumorigenesis remains unclear, but is thought to occur via increased mutagenesis. Here, we performed whole-genome sequencing on clonally expanded single liver stem cells cultured as intrahepatic cholangiocyte organoids (ICOs) from patients with alcoholic cirrhosis, non-alcoholic steatohepatitis (NASH), and primary sclerosing cholangitis (PSC). Surprisingly, we find that these precancerous liver disease conditions do not result in a detectable increased accumulation of mutations, nor altered mutation types in individual liver stem cells. This finding contrasts with the mutational load and typical mutational signatures reported for liver tumors, and argues against the hypothesis that liver disease drives tumorigenesis via a direct mechanism of induced mutagenesis. Disease conditions in the liver may thus act through indirect mechanisms to drive the transition from healthy to cancerous cells, such as changes to the microenvironment that favor the outgrowth of precancerous cells.
Central to tumor evolution is the generation of genetic diversity. However, the extent and patterns by which de novo karyotype alterations emerge and propagate within human tumors are not well understood, especially at single-cell resolution. Here, we present 3D Live-Seq-a protocol that integrates live-cell imaging of tumor organoid outgrowth and whole-genome sequencing of each imaged cell to reconstruct evolving tumor cell karyotypes across consecutive cell generations. Using patient-derived colorectal cancer organoids and fresh tumor biopsies, we demonstrate that karyotype alterations of varying complexity are prevalent and can arise within a few cell generations. Sub-chromosomal acentric fragments were prone to replication and collective missegregation across consecutive cell divisions. In contrast, gross genome-wide karyotype alterations were generated in a single erroneous cell division, providing support that aneuploid tumor genomes can evolve via punctuated evolution. Mapping the temporal dynamics and patterns of karyotype diversification in cancer enables reconstructions of evolutionary paths to malignant fitness.
An amendment to this paper has been published and can be accessed via a link at the top of the paper.
We have developed a novel, integrated and comprehensive purity, ploidy, structural variant and copy number somatic analysis toolkit for whole genome sequencing data of paired tumor/normal samples. We show that the combination of using GRIDSS for somatic structural variant calling and PURPLE for somatic copy number alteration calling allows highly sensitive, precise and consistent copy number and structural variant determination, as well as providing novel insights for short structural variants and regions of complex local topology. LINX, an interpretation tool, leverages the integrated structural variant and copy number calling to cluster individual structural variants into higher order events and chains them together to predict local derivative chromosome structure. LINX classifies and extensively annotates genomic rearrangements including simple and reciprocal breaks, LINE, viral and pseudogene insertions, and complex events such as chromothripsis. LINX also comprehensively calls genic fusions including chained fusions. Finally, our toolkit provides novel visualisation methods providing insight into complex genomic rearrangements.
A developing human fetus needs to balance rapid cellular expansion with maintaining genomic stability. Here, we accurately quantified and characterized somatic mutation accumulation in fetal tissues by analyzing individual stem cells from human fetal liver and intestine. Fetal mutation rates were about fivefold higher than in tissue-matched adult stem cells. The mutational landscape of fetal intestinal stem cells resembled that of adult intestinal stem cells, while the mutation spectrum of fetal liver stem cells is distinct from stem cells of the fetal intestine and the adult liver. Our analyses indicate that variation in mutational mechanisms, including oxidative stress and spontaneous deamination of methylated cytosines, contributes to the observed divergence in mutation accumulation patterns and drives genetic mosaicism in humans.
5-Fluorouracil (5-FU) is a chemotherapeutic drug commonly used for the treatment of solid cancers. It is proposed that 5-FU interferes with nucleotide synthesis and incorporates into DNA, which may have a mutational impact on both surviving tumor and healthy cells. Here, we treat intestinal organoids with 5-FU and find a highly characteristic mutational pattern that is dominated by T>G substitutions in a CTT context. Tumor whole genome sequencing data confirms that this signature is also identified in vivo in colorectal and breast cancer patients who have received 5-FU treatment. Taken together, our results demonstrate that 5-FU is mutagenic and may drive tumor evolution and increase the risk of secondary malignancies. Furthermore, the identified signature shows a strong resemblance to COSMIC signature 17, the hallmark signature of treatment-naive esophageal and gastric tumors, which indicates that distinct endogenous and exogenous triggers can converge onto highly similar mutational signatures.
Male breast cancer (MBC) is extremely rare and accounts for less than 1% of all breast malignancies. Therefore, clinical management of MBC is currently guided by research on the disease in females. In this study, DNA obtained from 45 formalin-fixed paraffin-embedded (FFPE) MBCs with and 90 MBCs (52 FFPE and 38 fresh-frozen) without matched normal tissues was subjected to massively parallel sequencing targeting all exons of 1943 cancer-related genes. The landscape of mutations and copy number alterations was compared to that of publicly available estrogen receptor (ER)-positive female breast cancers (smFBCs) and correlated to prognosis. From the 135 MBCs, 90% showed ductal histology, 96% were ER-positive, 66% were progesterone receptor (PR)-positive, and 2% HER2-positive, resulting in 50, 46 and 4% luminal A-like, luminal B-like and basal-like cases, respectively. Five patients had Klinefelter syndrome (4%) and 11% of patients harbored pathogenic BRCA2 germline mutations. The genomic landscape of MBC to some extent recapitulated that of smFBC, with recurrent PIK3CA (36%) and GATA3 (15%) somatic mutations, and with 40% of the most frequently amplified genes overlapping between both sexes. TP53 (3%) somatic mutations were significantly less frequent in MBC compared to smFBC, whereas somatic mutations in genes regulating chromatin function and homologous recombination deficiency-related signatures were more prevalent. MDM2 amplifications were frequent (13%), correlated with protein overexpression (P = 0.001) and predicted poor outcome (P = 0.007). In conclusion, despite similarities in the genomic landscape between MBC and smFBC, MBC is a molecularly unique and heterogeneous disease requiring its own clinical trials and treatment guidelines.
The FGF receptor signaling pathway is recurrently involved in the leukemogenic processes. Oncogenic fusions of FGFR1 with various fusion partners were described in myeloid proliferative neoplasms, and overexpression and mutations of FGFR3 are common in multiple myeloma. In addition, fibroblast growth factors are abundant in the bone marrow, and they were shown to enhance the survival of acute myeloid leukemia cells. Here we investigate the effect of FGFR stimulation on pediatric BCP-ALL cells in vitro, and search for mutations with deep targeted next-generation sequencing of mutational hotspots in FGFR1, FGFR2, and FGFR3. In 481 primary BCP-ALL cases, 28 samples from 19 unique relapsed BCP-ALL cases, and twelve BCP-ALL cell lines we found that mutations are rare (4/481 = 0.8%, 0/28 and 0/12) and do not affect codons which are frequently mutated in other malignancies. However, recombinant ligand FGF2 reduced the response to prednisolone in several BCP-ALL cell lines in vitro. We therefore conclude that FGFR signaling can contribute to prednisolone resistance in BCP-ALL cells, but that activating mutations in this receptor tyrosine kinase family are very rare.
ABSTRACT Excessive alcohol consumption increases the risk of developing liver cancer, but the mechanism through which alcohol drives carcinogenesis is as yet unknown. Here, we determined the mutational consequences of chronic alcohol use on the genome of human liver stem cells prior to cancer development. No change in base substitution rate or spectrum could be detected. Analysis of the trunk mutations in an alcohol-related liver tumor by multi-site whole-genome sequencing confirms the absence of specific alcohol-induced mutational signatures driving the development of liver cancer. However, we did identify an enrichment of nonsynonymous base substitutions in cancer genes in stem cells of the cirrhotic livers, such as recurrent nonsense mutations in PTPRK that disturb Epidermal Growth Factor (EGF)-signaling. Our results thus suggest that chronic alcohol use does not contribute to carcinogenesis through altered mutagenicity, but instead induces microenvironment changes which provide a ‘fertile ground’ for selection of cells with oncogenic mutations.
Nucleotide excision repair (NER) is one of the main DNA repair pathways that protect cells against genomic damage. Disruption of this pathway can contribute to the development of cancer and accelerate aging. Mutational characteristics of NER-deficiency may reveal important diagnostic opportunities, as tumors deficient in NER are more sensitive to certain treatments. Here, we analyzed the genome-wide somatic mutational profiles of adult stem cells (ASCs) from NER-deficient Ercc1−/Δ mice. Our results indicate that NER-deficiency increases the base substitution load twofold in liver but not in small intestinal ASCs, which coincides with the tissue-specific aging pathology observed in these mice. Moreover, NER-deficient ASCs of both tissues show an increased contribution of Signature 8 mutations, which is a mutational pattern with unknown etiology that is recurrently observed in various cancer types. The scattered genomic distribution of the base substitutions indicates that deficiency of global-genome NER (GG-NER) underlies the observed mutational consequences. In line with this, we observe increased Signature 8 mutations in a GG-NER-deficient human organoid culture, in which XPC was deleted using CRISPR-Cas9 gene-editing. Furthermore, genomes of NER-deficient breast tumors show an increased contribution of Signature 8 mutations compared with NER-proficient tumors. Elevated levels of Signature 8 mutations could therefore contribute to a predictor of NER-deficiency based on a patient's mutational profile.
Background: Genomic structural variants (SVs) can affect many genes and regulatory elements. Therefore, the molecular mechanisms driving the phenotypes of patients carrying de novo SVs are frequently unknown. Methods: We applied a combination of systematic experimental and bioinformatic methods to improve the molecular diagnosis of 39 patients with multiple congenital abnormalities and/or intellectual disability harboring apparent de novo SVs, most with an inconclusive diagnosis after regular genetic testing. Results: In 7 of these cases (18%), whole-genome sequencing analysis revealed disease-relevant complexities of the SVs missed in routine microarray-based analyses. We developed a computational tool to predict the effects on genes directly affected by SVs and on genes indirectly affected likely due to the changes in chromatin organization and impact on regulatory mechanisms. By combining these functional predictions with extensive phenotype information, candidate driver genes were identified in 16/39 (41%) patients. In 8 cases, evidence was found for the involvement of multiple candidate drivers contributing to different parts of the phenotypes. Subsequently, we applied this computational method to two cohorts containing a total of 379 patients with previously detected and classified de novo SVs and identified candidate driver genes in 189 cases (50%), including 40 cases whose SVs were previously not classified as pathogenic. Pathogenic position effects were predicted in 28% of all studied cases with balanced SVs and in 11% of the cases with copy number variants. Conclusions: These results demonstrate an integrated computational and experimental approach to predict driver genes based on analyses of WGS data with phenotype association and chromatin organization datasets. These analyses nominate new pathogenic loci and have strong potential to improve the molecular diagnosis of patients with de novo SVs.