Nat. Genet. 48, 827–837 (2016); published online 13 June 2016; corrected after print 20 March 2017 In the version of this article initially published, residue R384 was incorrectly highlighted in the protein model depicted in Figure 7e. The correct residue is R394. The error has been corrected in theHTML and PDF versions of the article.
Recent progress in next-generation sequencing has greatly facilitated our study of genomic structural variation. Unlike single nucleotide variants and small indels, many structural variants have not been completely characterized at nucleotide resolution. Deriving the complete sequences underlying such breakpoints is crucial for not only accurate discovery, but also for the functional characterization of altered alleles. However, our current ability to determine such breakpoint sequences is limited because of challenges in aligning and assembling short reads. To address this issue, we developed a targeted iterative graph routing assembler, TIGRA, which implements a set of novel data analysis routines to achieve effective breakpoint assembly from next-generation sequencing data. In our assessment using data from the 1000 Genomes Project, TIGRA was able to accurately assemble the majority of deletion and mobile element insertion breakpoints, with a substantively better success rate and accuracy than other algorithms. TIGRA has been applied in the 1000 Genomes Project and other projects and is freely available for academic use.
Material Supplemental http://genome.cshlp.org/content/suppl/2013/01/02/gr.142604.112.DC1.html References http://genome.cshlp.org/content/23/3/431.full.html#ref-list-1 This article cites 66 articles, 18 of which can be accessed free at: License Commons Creative . http://creativecommons.org/licenses/by-nc/3.0/ described at a Creative Commons License (Attribution-NonCommercial 3.0 Unported License), as ). After six months, it is available under http://genome.cshlp.org/site/misc/terms.xhtml first six months after the full-issue publication date (see This article is distributed exclusively by Cold Spring Harbor Laboratory Press for the
Producing gene fusions through genomic structural rearrangements is a major mechanism for tumor evolution. Therefore, accurately detecting gene fusions and the originating rearrangements is of great importance for personalized cancer diagnosis and targeted therapy. We present a tool, BreakTrans, that systematically maps predicted gene fusions to structural rearrangements. Thus, BreakTrans not only validates both types of predictions, but also provides mechanistic interpretations. BreakTrans effectively validates known fusions and discovers novel events in a breast cancer cell line. Applying BreakTrans to 43 breast cancer samples in The Cancer Genome Atlas identifies 90 genomically validated gene fusions. BreakTrans is available at http://bioinformatics.mdanderson.org/main/BreakTrans
The sequencing of AML genomes of eight patients before and after relapse reveals two major patterns of clonal evolution, with chemotherapy appearing to have a role in both patterns. Many patients with acute myeloid leukaemia (AML) achieve remission, but it is often short-lived and the returned disease is usually refractory to therapy. Genome sequencing of eight patients with AML before and after relapse reveals two major patterns of tumour cell evolution. The founding clone survives chemotherapy in all patients, and, in one clonal pattern, it acquires new mutations and expands at relapse. In the other, a subclone surviving from the original tumour expands and then acquires new mutations. Comparisons of relapse-specific and primary tumour mutations point to an increase in transversions, implying DNA damage caused by cytotoxic chemotherapy. This work demonstrates that the AML genome in an individual patient presents a moving target, and highlights the importance of striving to eradicate both the founding clone and all of its subclones. Most patients with acute myeloid leukaemia (AML) die from progressive disease after relapse, which is associated with clonal evolution at the cytogenetic level1,2 . To determine the mutational spectrum associated with relapse, we sequenced the primary tumour and relapse genomes from eight AML patients, and validated hundreds of somatic mutations using deep sequencing; this allowed us to define clonality and clonal evolution patterns precisely at relapse. In addition to discovering novel, recurrently mutated genes (for example, WAC, SMC3, DIS3, DDX41 and DAXX) in AML, we also found two major clonal evolution patterns during AML relapse: (1) the founding clone in the primary tumour gained mutations and evolved into the relapse clone, or (2) a subclone of the founding clone survived initial therapy, gained additional mutations and expanded at relapse. In all cases, chemotherapy failed to eradicate the founding clone. The comparison of relapse-specific versus primary tumour mutations in all eight cases revealed an increase in transversions, probably due to DNA damage caused by cytotoxic chemotherapy. These data demonstrate that AML relapse is associated with the addition of new mutations and clonal evolution, which is shaped, in part, by the chemotherapy that the patients receive to establish and maintain remissions.
503 Background: To correlate clinical features of estrogen receptor positive breast cancer with somatic mutations, massively parallel sequencing (MPS) was applied to tumor and normal DNAs accrued from patients treated with neoadjuvant aromatase inhibitors (AI). Methods: MPS was applied to 77 baseline tumor biopsy samples from the preoperative letrozole trial (JACS 2009: 208, 906) and the Z1031 trial (JCO 2011: 29, 2342) followed by targeted sequencing in another 240 trial samples. Standard statistical approaches were used to compare mutation status and clinical parameters and pathway-based correlation was used to assess interactions between signaling perturbations induced by gene mutations and response to neoadjuvant AI. Results: Eighteen genes were significantly mutated above background. Aside from PIK3CA mutations, the list is dominated by loss-of-function mutations in tumor suppressor genes. Five (RUNX1, CBFP, MYH9, MLL3 and SF3B1) have been previously linked to benign and malignant hematopoietic disorders. Clinical correlation revealed that TP53 mutation was associated with PAM50 LumB status, high-grade histology and high proliferation rates whereas loss-of-function mutations in MAP3K1 associate with PAM50 lumA status, low proliferation rates, and low grade histology. Mutations in GATA3 were associated with greater suppression of proliferation upon AI treatment suggesting mutGATA3 may predict endocrine response. Notably, mutations in MAP3K1 were more common in PIK3CA mutant cases, suggesting cooperation. Pathway analysis demonstrated that rare MAP2K4 mutations produce similar pathway perturbations as MAP3K1 mutation, a logical finding since MAP2K4 is a substrate for MAP3K1. Signaling network patterns driven by lncRNA MALAT1 mutations were associated with multiple poor clinical outcome features. Rare mutations in druggable kinases included two in the kinase domain of HER2. Conclusions: Tumor heterogeneity in luminal-type breast cancer is driven by specific patterns of somatic mutations, however most druggable or potentially prognostic mutations are infrequent. Prospective clinical trials based on these findings will require comprehensive genome sequencing approaches and large scale investigations.
UNLABELLED:Despite recent progress, computational tools that identify gene fusions from next-generation whole transcriptome sequencing data are often limited in accuracy and scalability. Here, we present a software package, BreakFusion that combines the strength of reference alignment followed by read-pair analysis and de novo assembly to achieve a good balance in sensitivity, specificity and computational efficiency.AVAILABILITY:http://bioinformatics.mdanderson.org/main/BreakFusion
Abstract Background: Estrogen receptors are over-expressed in around 70% of breast cancer cases. The genetic changes that occur during aromatase inhibitor (AI) treatment are not well understood and may differ depending upon the patient's response phenotype. Methods: We performed whole genome sequencing (WGS) of matched blood, pre-treatment, and post-treatment biopsy samples from 22 estrogen receptor positive breast cancer patients treated with neoadjuvant aromatase inhibitors. For 5 cases, we performed the whole genome sequencing (WGS) on patients’ matched normal, two pre AI-treatment, and two post AI-treatment DNA isolates from biopsy samples. We validated all putative coding and non-coding somatic mutations using deep sequencing. By comparing the validated somatic mutations from pre- and post- AI treatment biopsy samples, we were able to determine the alterations in the tumor genomes. In every case we defined the clonal architecture of each pair of pre-treatment and post-treatment biopsy samples by comparing the variant allele frequencies from thousands of validated somatic mutations. Results: Comparisons of the two pre AI-treatment biopsy samples from the same patient indicates that the variant allele frequencies of mutations showed high concordances in all 5 cases, 0.74 to 0.95 range of correlation coefficient. Only a small percentage of somatic mutations were detected in one pre-treatment sample and not the other (4.65% overall). In comparing the somatic variations between pre-treatment and matched post-treatment biopsy samples in 22 cases, we found that patients with good clinical response to AI treatment retained known driver mutations only in their pre-treatment tumors. Conversely, those patients with poor clinical response presented new driver mutations in their post-treatment samples. Furthermore, the variant allele frequency for most mutated genes decreased in post AI treatment samples for patients with good AI treatment response; on the contrary, the variant allele frequency increased for patients with poor clinical response. Conclusions: From WGS of matched normal, pre-treatment, and post-treatment biopsy samples, we identified new driver genes mutated in patients with poor clinical response, while patients with good clinical response had lost mutated driver genes in their post-treatment biopsy samples. The genetic landscape revealed by WGS of pre-treatment and post-treatment biopsy samples reveals mutational repertoires are remodeled by AI therapy. This finding suggests deep sequencing of AI treated samples will be necessary to reveal the complete complement of mutations present in a patient's tumor. Citation Format: {Authors}. {Abstract title} [abstract]. In: Proceedings of the 103rd Annual Meeting of the American Association for Cancer Research; 2012 Mar 31-Apr 4; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2012;72(8 Suppl):Abstract nr LB-423. doi:1538-7445.AM2012-LB-423
Most mutations in cancer genomes are thought to be acquired after the initiating event, which may cause genomic instability and drive clonal evolution. However, for acute myeloid leukemia (AML), normal karyotypes are common, and genomic instability is unusual. To better understand clonal evolution in AML, we sequenced the genomes of M3-AML samples with a known initiating event (PML-RARA) versus the genomes of normal karyotype M1-AML samples and the exomes of hematopoietic stem/progenitor cells (HSPCs) from healthy people. Collectively, the data suggest that most of the mutations found in AML genomes are actually random events that occurred in HSPCs before they acquired the initiating mutation; the mutational history of that cell is "captured" as the clone expands. In many cases, only one or two additional, cooperating mutations are needed to generate the malignant founding clone. Cells from the founding clone can acquire additional cooperating mutations, yielding subclones that can contribute to disease progression and/or relapse.
Low-grade brain tumors (pilocytic astrocytomas) arising in the neurofibromatosis type 1 (NF1) inherited cancer predisposition syndrome are hypothesized to result from a combination of germline and acquired somatic NF1 tumor suppressor gene mutations. However, genetically engineered mice (GEM) in which mono-allelic germline Nf1 gene loss is coupled with bi-allelic somatic (glial progenitor cell) Nf1 gene inactivation develop brain tumors that do not fully recapitulate the neuropathological features of the human condition. These observations raise the intriguing possibility that, while loss of neurofibromin function is necessary for NF1-associated low-grade astrocytoma development, additional genetic changes may be required for full penetrance of the human brain tumor phenotype. To identify these potential cooperating genetic mutations, we performed whole-genome sequencing (WGS) analysis of three NF1-associated pilocytic astrocytoma (PA) tumors. We found that the mechanism of somatic NF1 loss was different in each tumor (frameshift mutation, loss of heterozygosity, and methylation). In addition, tumor purity analysis revealed that these tumors had a high proportion of stromal cells, such that only 50%-60% of cells in the tumor mass exhibited somatic NF1 loss. Importantly, we identified no additional recurrent pathogenic somatic mutations, supporting a model in which neuroglial progenitor cell NF1 loss is likely sufficient for PA formation in cooperation with a proper stromal environment.
To correlate the variable clinical features of oestrogen-receptor-positive breast cancer with somatic alterations, we studied pretreatment tumour biopsies accrued from patients in two studies of neoadjuvant aromatase inhibitor therapy by massively parallel sequencing and analysis. Eighteen significantly mutated genes were identified, including five genes (RUNX1, CBFB, MYH9, MLL3 and SF3B1) previously linked to haematopoietic disorders. Mutant MAP3K1 was associated with luminal A status, low-grade histology and low proliferation rates, whereas mutant TP53 was associated with the opposite pattern. Moreover, mutant GATA3 correlated with suppression of proliferation upon aromatase inhibitor treatment. Pathway analysis demonstrated that mutations in MAP2K4, a MAP3K1 substrate, produced similar perturbations as MAP3K1 loss. Distinct phenotypes in oestrogen-receptor-positive breast cancer are associated with specific patterns of somatic mutations that map into cellular pathways linked to tumour biology, but most recurrent mutations are relatively infrequent. Prospective clinical trials based on these findings will require comprehensive genome sequencing.
Abstract Abstract 320 Multiple myeloma (MM) is an incurable malignancy of antibody secreting plasma B-cells whose etiology is still poorly understood. We used whole genome sequencing (WGS) and deep-read capture validation to thoroughly characterize the mutation landscape in four newly diagnosed MM patients and further examined mutation recurrence in 89 additional MM patients. WGS cases were selected to be racially diverse and represent both hyperdiploid and non-hyperdiploid MM with otherwise “simple” karyotypes. All studies to date have used peripheral blood as controls, however, abnormal B cells and circulating tumor cells frequently contaminate the peripheral blood of MM patients and these studies may have missed the early genetic events potentially important in disease pathogenesis. We therefore chose to use matched skin samples as normal controls. The use of skin DNA controls and deep read count capture validation of single nucleotide variants (SNV) allowed us, for the first time, to demonstrate clear separation between normal and malignant cell populations. Our analysis pipeline detected both somatic SNVs and structural variants (SV, ie. translocations, deletions, insertions). In all four patients, we observed chromosomal translocations at the Ig heavy chain locus, VDJ recombination, and Ig locus somatic hypermuation. We found somatic mutations affecting the E3 ubiquitin ligase HUWE1 in 4% (4/94) of cases, as well as recurring deletions affecting the Rho pathway regulator DIAPH2 (recurrence data will be presented). RNA sequencing was preformed for 3 of the WGS patients to explore transcriptional effects of mutations observed. Statistical analysis identified significantly mutated genes, i.e. mutation hotspots, including ROBO2, BCL6 and cadherin/catenin genes. Genes involved in cell polarity, cell adhesion and axon guidance were affected in all four WGS patients. Chromosomal translocations present in all four WGS patients involved both known oncogenes (CCND1, MYC and MAFB) and putative oncogenes. A novel translocation implicated KCNT2, encoding a sodium-activated potassium channel, as a potential oncogene in MM. Recurrent point mutations were relatively rare in MM, suggesting that SVs are more common driver mutations and/or that SNVs in MM disrupt a diverse array of genes to affect key pathways. For genomic studies in MM moving forward, our results emphasize the importance of matched normal controls uncontaminated by tumor cells, and suggest that careful analysis of SVs be included to find novel oncogenes and important clinical correlations. Disclosures: No relevant conflicts of interest to declare.
We report the results of whole-genome and transcriptome sequencing of tumor and adjacent normal tissue samples from 17 patients with non-small cell lung carcinoma (NSCLC). We identified 3,726 point mutations and more than 90 indels in the coding sequence, with an average mutation frequency more than 10-fold higher in smokers than in never-smokers. Novel alterations in genes involved in chromatin modification and DNA repair pathways were identified, along with DACH1, CFTR, RELN, ABCB5, and HGF. Deep digital sequencing revealed diverse clonality patterns in both never-smokers and smokers. All validated EFGR and KRAS mutations were present in the founder clones, suggesting possible roles in cancer initiation. Analysis revealed 14 fusions, including ROS1 and ALK, as well as novel metabolic enzymes. Cell-cycle and JAK-STAT pathways are significantly altered in lung cancer, along with perturbations in 54 genes that are potentially targetable with currently available drugs.
MOTIVATION The expansion of cancer genome sequencing continues to stimulate development of analytical tools for inferring relationships between somatic changes and tumor development. Pathway associations are especially consequential, but existing algorithms are demonstrably inadequate. METHODS Here, we propose the PathScan significance test for the scenario where pathway mutations collectively contribute to tumor development. Its design addresses two aspects that established methods neglect. First, we account for variations in gene length and the consequent differences in their mutation probabilities under the standard null hypothesis of random mutation. The associated spike in computational effort is mitigated by accurate convolution-based approximation. Second, we combine individual probabilities into a multiple-sample value using Fisher-Lancaster theory, thereby improving differentiation between a few highly mutated genes and many genes having only a few mutations apiece. We investigate accuracy, computational effort and power, reporting acceptable performance for each. RESULTS As an example calculation, we re-analyze KEGG-based lung adenocarcinoma pathway mutations from the Tumor Sequencing Project. Our test recapitulates the most significant pathways and finds that others for which the original test battery was inconclusive are not actually significant. It also identifies the focal adhesion pathway as being significantly mutated, a finding consistent with earlier studies. We also expand this analysis to other databases: Reactome, BioCarta, Pfam, PID and SMART, finding additional hits in ErbB and EPHA signaling pathways and regulation of telomerase. All have implications and plausible mechanistic roles in cancer. Finally, we discuss aspects of extending the method to integrate gene-specific background rates and other types of genetic anomalies. AVAILABILITY PathScan is implemented in Perl and is available from the Genome Institute at: http://genome.wustl.edu/software/pathscan.
CONTEXT Whole-genome sequencing is becoming increasingly available for research purposes, but it has not yet been routinely used for clinical diagnosis. OBJECTIVE To determine whether whole-genome sequencing can identify cryptic, actionable mutations in a clinically relevant time frame. DESIGN, SETTING, AND PATIENT We were referred a difficult diagnostic case of acute promyelocytic leukemia with no pathogenic X-RARA fusion identified by routine metaphase cytogenetics or interphase fluorescence in situ hybridization (FISH). The case patient was enrolled in an institutional review board-approved protocol, with consent specifically tailored to the implications of whole-genome sequencing. The protocol uses a "movable firewall" that maintains patient anonymity within the entire research team but allows the research team to communicate medically relevant information to the treating physician. MAIN OUTCOME MEASURES Clinical relevance of whole-genome sequencing and time to communicate validated results to the treating physician. RESULTS Massively parallel paired-end sequencing allowed identification of a cytogenetically cryptic event: a 77-kilobase segment from chromosome 15 was inserted en bloc into the second intron of the RARA gene on chromosome 17, resulting in a classic bcr3 PML-RARA fusion gene. Reverse transcription polymerase chain reaction sequencing subsequently validated the expression of the fusion transcript. Novel FISH probes identified 2 additional cases of t(15;17)-negative acute promyelocytic leukemia that had cytogenetically invisible insertions. Whole-genome sequencing and validation were completed in 7 weeks and changed the treatment plan for the patient. CONCLUSION Whole-genome sequencing can identify cytogenetically invisible oncogenes in a clinically relevant time frame.
Abstract Background: Endocrine-therapy resistant HER2 negative luminal-type breast cancer is the commonest cause of breast cancer death but the molecular basis for aggressive clinical behavior is poorly understood. Endocrine therapy resistant cases can be identified early in the course of the disease though the persistant expression of the cell cycle biomarker Ki67 despite neoadjuvant treatment with a potent aromatase inhibitor (JNCI 2008:100;159–61). Methods: Massively parallel DNA sequencing (Nature 2010; 464:999–1005) was applied to samples from two neoadjuvant endocrine trials (“POL” J Am Coll Surg. 2009; 208: 906–14 and “ACOSOG Z1031” JCO in press) to define the whole genome sequence of 50 luminal-type breast cancers (PAM50 definition, JCO 2009;27:1160–7) with an average 30 fold coverage. Of these, 24 were defined as resistant (surgical specimen Ki67 > 10%) and 26 sensitive (Ki67 ≤ 10%). Approximately 10 trillion base pairs of sequence were analyzed. Putative somatic mutations were confirmed through targeted sequence analysis and the absence of the variant in matched normal DNA. Results: The mean number of single nucleotide non-silent variants (SNV) and small deletions (indels) was greater in the resistant group (49 per case) than sensitive cases (23) (p=0.02). The significantly mutated gene list included PIK3CA, TP53, ATR, RUNX1, MYST3, PRSS8, ZNHIT2 and MAP3K1. A large number of structural variations have also been identified with greater than 800 events from 28 patients orthogonally validated to date, mostly deletions but also translocations, although no recurrent translocations have been seen. About 25% of these deletions were paired with an SNV suggesting multiple novel double-hit tumor suppressor events. Extension analysis in 121 POL/Z1031 samples produced the following mutation incidence data PIK3CA (43%), TP53 (15.2%) and MAP3K1 (9.3%). The MAP3K1 mutations were most frequently frame-shift or nonsense, disrupting the C terminal kinase domain, a logical finding in a kinase activated by caspase 7 to promote apoptosis. Only 1 MAP3K1 mutation was observed in 60 basal-like breast cancers (luminal versus basal MAP3K1 mutation frequency p=0.04). Conclusions: MAP3K1is a novel breast cancer tumor suppressor gene more frequently lost in luminal-type than basal-like breast cancer. Further data analysis will be presented including pathway analyses that place less common mutations into a variety of cancer pathways, including phosphoinositol-3-kinase signaling, cell cycle regulation, metabolism, mitotic spindle regulation and DNA mismatch repair. These analyses suggest a classification of the disease that is not based on individual gene abnormalities but common cancer pathway activation events that ultimately determine outcome for patients with luminal-type disease. Citation Format: {Authors}. {Abstract title} [abstract]. In: Proceedings of the 102nd Annual Meeting of the American Association for Cancer Research; 2011 Apr 2-6; Orlando, FL. Philadelphia (PA): AACR; Cancer Res 2011;71(8 Suppl):Abstract nr LB-87. doi:10.1158/1538-7445.AM2011-LB-87
Abstract The ability to study genome-wide sequence alterations in cancer genomes has transformed our understanding of the intricacies of this disease. We have applied whole genome sequencing and analysis to tease apart the genomic predictors of aromatase inhibitor (AI) response, by selecting luminal breast cancer cases obtained within a clinical trial of aromatase inhibitor treatment (ACOSOG Z1031). Our study compared the constitutional genome to the cancer genome for patients that exhibit either a response or a resistance phenotype, based on their Ki67 IHC score following 4 month AI treatment and resection. My talk will highlight the process by which these genomes are sequenced and analyzed, and will provide our most recent insights into how the combination of the constitutional and somatic genomes determine whether a patient will respond to AI treatment. Citation Information: Cancer Res 2010;70(24 Suppl):Abstract nr ES7-1.