In this work, we present the Genome Modeling System (GMS), an analysis information management system capable of executing automated genome analysis pipelines at a massive scale. The GMS framework provides detailed tracking of samples and data coupled with reliable and repeatable analysis pipelines. The GMS also serves as a platform for bioinformatics development, allowing a large team to collaborate on data analysis, or an individual researcher to leverage the work of others effectively within its data management system. Rather than separating ad-hoc analysis from rigorous, reproducible pipelines, the GMS promotes systematic integration between the two. As a demonstration of the GMS, we performed an integrated analysis of whole genome, exome and transcriptome sequencing data from a breast cancer cell line (HCC1395) and matched lymphoblastoid line (HCC1395BL). These data are available for users to test the software, complete tutorials and develop novel GMS pipeline configurations. The GMS is available at https://github.com/genome/gms.
Retinoblastoma is a rare childhood cancer of the developing retina. Most retinoblastomas initiate with biallelic inactivation of the RB1 gene through diverse mechanisms including point mutations, nucleotide insertions, deletions, loss of heterozygosity and promoter hypermethylation. Recently, a novel mechanism of retinoblastoma initiation was proposed. Gallie and colleagues discovered that a small proportion of retinoblastomas lack RB1 mutations and had MYCN amplification [1]. In this study, we identified recurrent chromosomal, regional and focal genomic lesions in 94 primary retinoblastomas with their matched normal DNA using SNP 6.0 chips. We also analyzed the RB1 gene mutations and compared the mechanism of RB1 inactivation to the recurrent copy number variations in the retinoblastoma genome. In addition to the previously described focal amplification of MYCN and deletions in RB1 and BCOR, we also identified recurrent focal amplification of OTX2, a transcription factor required for retinal photoreceptor development. We identified 10 retinoblastomas in our cohort that lacked RB1 point mutations or indels. We performed whole genome sequencing on those 10 tumors and their corresponding germline DNA. In one of the tumors, the RB1 gene was unaltered, the MYCN gene was amplified and RB1 protein was expressed in the nuclei of the tumor cells. In addition, several tumors had complex patterns of structural variations and we identified 3 tumors with chromothripsis at the RB1 locus. This is the first report of chromothripsis as a mechanism for RB1 gene inactivation in cancer.
The genetic basis of hypodiploid acute lymphoblastic leukemia (ALL), a subtype of ALL characterized by aneuploidy and poor outcome, is unknown. Genomic profiling of 124 hypodiploid ALL cases, including whole-genome and exome sequencing of 40 cases, identified two subtypes that differ in the severity of aneuploidy, transcriptional profiles and submicroscopic genetic alterations. Near-haploid ALL with 24-31 chromosomes harbor alterations targeting receptor tyrosine kinase signaling and Ras signaling (71%) and the lymphoid transcription factor gene IKZF3 (encoding AIOLOS; 13%). In contrast, low-hypodiploid ALL with 32-39 chromosomes are characterized by alterations in TP53 (91.2%) that are commonly present in nontumor cells, IKZF2 (encoding HELIOS; 53%) and RB1 (41%). Both near-haploid and low-hypodiploid leukemic cells show activation of Ras-signaling and phosphoinositide 3-kinase (PI3K)-signaling pathways and are sensitive to PI3K inhibitors, indicating that these drugs should be explored as a new therapeutic strategy for this aggressive form of leukemia.
The Drug-Gene Interaction database (DGIdb) mines existing resources that generate hypotheses about how mutated genes might be targeted therapeutically or prioritized for drug development. It provides an interface for searching lists of genes against a compendium of drug-gene interactions and potentially 'druggable' genes. DGIdb can be accessed at http://dgidb.org/.
The sequencing of AML genomes of eight patients before and after relapse reveals two major patterns of clonal evolution, with chemotherapy appearing to have a role in both patterns. Many patients with acute myeloid leukaemia (AML) achieve remission, but it is often short-lived and the returned disease is usually refractory to therapy. Genome sequencing of eight patients with AML before and after relapse reveals two major patterns of tumour cell evolution. The founding clone survives chemotherapy in all patients, and, in one clonal pattern, it acquires new mutations and expands at relapse. In the other, a subclone surviving from the original tumour expands and then acquires new mutations. Comparisons of relapse-specific and primary tumour mutations point to an increase in transversions, implying DNA damage caused by cytotoxic chemotherapy. This work demonstrates that the AML genome in an individual patient presents a moving target, and highlights the importance of striving to eradicate both the founding clone and all of its subclones. Most patients with acute myeloid leukaemia (AML) die from progressive disease after relapse, which is associated with clonal evolution at the cytogenetic level1,2 . To determine the mutational spectrum associated with relapse, we sequenced the primary tumour and relapse genomes from eight AML patients, and validated hundreds of somatic mutations using deep sequencing; this allowed us to define clonality and clonal evolution patterns precisely at relapse. In addition to discovering novel, recurrently mutated genes (for example, WAC, SMC3, DIS3, DDX41 and DAXX) in AML, we also found two major clonal evolution patterns during AML relapse: (1) the founding clone in the primary tumour gained mutations and evolved into the relapse clone, or (2) a subclone of the founding clone survived initial therapy, gained additional mutations and expanded at relapse. In all cases, chemotherapy failed to eradicate the founding clone. The comparison of relapse-specific versus primary tumour mutations in all eight cases revealed an increase in transversions, probably due to DNA damage caused by cytotoxic chemotherapy. These data demonstrate that AML relapse is associated with the addition of new mutations and clonal evolution, which is shaped, in part, by the chemotherapy that the patients receive to establish and maintain remissions.
Early T-cell precursor acute lymphoblastic leukaemia (ETP ALL) is an aggressive malignancy of unknown genetic basis. We performed whole-genome sequencing of 12 ETP ALL cases and assessed the frequency of the identified somatic mutations in 94 T-cell acute lymphoblastic leukaemia cases. ETP ALL was characterized by activating mutations in genes regulating cytokine receptor and RAS signalling (67% of cases; NRAS, KRAS, FLT3, IL7R, JAK3, JAK1, SH2B3 and BRAF), inactivating lesions disrupting haematopoietic development (58%; GATA3, ETV6, RUNX1, IKZF1 and EP300) and histone-modifying genes (48%; EZH2, EED, SUZ12, SETD2 and EP300). We also identified new targets of recurrent mutation including DNM2, ECT2L and RELN. The mutational spectrum is similar to myeloid tumours, and moreover, the global transcriptional profile of ETP ALL was similar to that of normal and myeloid leukaemia haematopoietic stem cells. These findings suggest that addition of myeloid-directed therapies might improve the poor outcome of ETP ALL.
Abstract Background: Estrogen receptors are over-expressed in around 70% of breast cancer cases. The genetic changes that occur during aromatase inhibitor (AI) treatment are not well understood and may differ depending upon the patient's response phenotype. Methods: We performed whole genome sequencing (WGS) of matched blood, pre-treatment, and post-treatment biopsy samples from 22 estrogen receptor positive breast cancer patients treated with neoadjuvant aromatase inhibitors. For 5 cases, we performed the whole genome sequencing (WGS) on patients’ matched normal, two pre AI-treatment, and two post AI-treatment DNA isolates from biopsy samples. We validated all putative coding and non-coding somatic mutations using deep sequencing. By comparing the validated somatic mutations from pre- and post- AI treatment biopsy samples, we were able to determine the alterations in the tumor genomes. In every case we defined the clonal architecture of each pair of pre-treatment and post-treatment biopsy samples by comparing the variant allele frequencies from thousands of validated somatic mutations. Results: Comparisons of the two pre AI-treatment biopsy samples from the same patient indicates that the variant allele frequencies of mutations showed high concordances in all 5 cases, 0.74 to 0.95 range of correlation coefficient. Only a small percentage of somatic mutations were detected in one pre-treatment sample and not the other (4.65% overall). In comparing the somatic variations between pre-treatment and matched post-treatment biopsy samples in 22 cases, we found that patients with good clinical response to AI treatment retained known driver mutations only in their pre-treatment tumors. Conversely, those patients with poor clinical response presented new driver mutations in their post-treatment samples. Furthermore, the variant allele frequency for most mutated genes decreased in post AI treatment samples for patients with good AI treatment response; on the contrary, the variant allele frequency increased for patients with poor clinical response. Conclusions: From WGS of matched normal, pre-treatment, and post-treatment biopsy samples, we identified new driver genes mutated in patients with poor clinical response, while patients with good clinical response had lost mutated driver genes in their post-treatment biopsy samples. The genetic landscape revealed by WGS of pre-treatment and post-treatment biopsy samples reveals mutational repertoires are remodeled by AI therapy. This finding suggests deep sequencing of AI treated samples will be necessary to reveal the complete complement of mutations present in a patient's tumor. Citation Format: {Authors}. {Abstract title} [abstract]. In: Proceedings of the 103rd Annual Meeting of the American Association for Cancer Research; 2012 Mar 31-Apr 4; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2012;72(8 Suppl):Abstract nr LB-423. doi:1538-7445.AM2012-LB-423
Most mutations in cancer genomes are thought to be acquired after the initiating event, which may cause genomic instability and drive clonal evolution. However, for acute myeloid leukemia (AML), normal karyotypes are common, and genomic instability is unusual. To better understand clonal evolution in AML, we sequenced the genomes of M3-AML samples with a known initiating event (PML-RARA) versus the genomes of normal karyotype M1-AML samples and the exomes of hematopoietic stem/progenitor cells (HSPCs) from healthy people. Collectively, the data suggest that most of the mutations found in AML genomes are actually random events that occurred in HSPCs before they acquired the initiating mutation; the mutational history of that cell is "captured" as the clone expands. In many cases, only one or two additional, cooperating mutations are needed to generate the malignant founding clone. Cells from the founding clone can acquire additional cooperating mutations, yielding subclones that can contribute to disease progression and/or relapse.
Retinoblastoma is an aggressive childhood cancer of the developing retina that is initiated by the biallelic loss of RB1. Tumours progress very quickly following RB1 inactivation but the underlying mechanism is not known. Here we show that the retinoblastoma genome is stable, but that multiple cancer pathways can be epigenetically deregulated. To identify the mutations that cooperate with RB1 loss, we performed whole-genome sequencing of retinoblastomas. The overall mutational rate was very low; RB1 was the only known cancer gene mutated. We then evaluated the role of RB1 in genome stability and considered non-genetic mechanisms of cancer pathway deregulation. For example, the proto-oncogene SYK is upregulated in retinoblastoma and is required for tumour cell survival. Targeting SYK with a small-molecule inhibitor induced retinoblastoma tumour cell death in vitro and in vivo. Thus, retinoblastomas may develop quickly as a result of the epigenetic deregulation of key cancer pathways as a direct or indirect result of RB1 loss.
To correlate the variable clinical features of oestrogen-receptor-positive breast cancer with somatic alterations, we studied pretreatment tumour biopsies accrued from patients in two studies of neoadjuvant aromatase inhibitor therapy by massively parallel sequencing and analysis. Eighteen significantly mutated genes were identified, including five genes (RUNX1, CBFB, MYH9, MLL3 and SF3B1) previously linked to haematopoietic disorders. Mutant MAP3K1 was associated with luminal A status, low-grade histology and low proliferation rates, whereas mutant TP53 was associated with the opposite pattern. Moreover, mutant GATA3 correlated with suppression of proliferation upon aromatase inhibitor treatment. Pathway analysis demonstrated that mutations in MAP2K4, a MAP3K1 substrate, produced similar perturbations as MAP3K1 loss. Distinct phenotypes in oestrogen-receptor-positive breast cancer are associated with specific patterns of somatic mutations that map into cellular pathways linked to tumour biology, but most recurrent mutations are relatively infrequent. Prospective clinical trials based on these findings will require comprehensive genome sequencing.
To characterize somatic alterations in colorectal carcinoma, we conducted a genome-scale analysis of 276 samples, analysing exome sequence, DNA copy number, promoter methylation and messenger RNA and microRNA expression. A subset of these samples (97) underwent low-depth-of-coverage whole-genome sequencing. In total, 16% of colorectal carcinomas were found to be hypermutated: three-quarters of these had the expected high microsatellite instability, usually with hypermethylation and MLH1 silencing, and one-quarter had somatic mismatch-repair gene and polymerase e (POLE) mutations. Excluding the hypermutated cancers, colon and rectum cancers were found to have considerably similar patterns of genomic alteration. Twenty-four genes were significantly mutated, and in addition to the expected APC, TP53, SMAD4, PIK3CA and KRAS mutations, we found frequent mutations in ARID1A, SOX9 and FAM123B. Recurrent copy-number alterations include potentially drug-targetable amplifications of ERBB2 and newly discovered amplification of IGF2. Recurrent chromosomal translocations include the fusion of NAV2 and WNT pathway member TCF7L1. Integrative analyses suggest new markers for aggressive colorectal carcinoma and an important role for MYC-directed transcriptional activation and repression.
Medulloblastoma is a malignant childhood brain tumour comprising four discrete subgroups. Here, to identify mutations that drive medulloblastoma, we sequenced the entire genomes of 37 tumours and matched normal blood. One-hundred and thirty-six genes harbouring somatic mutations in this discovery set were sequenced in an additional 56 medulloblastomas. Recurrent mutations were detected in 41 genes not yet implicated in medulloblastoma; several target distinct components of the epigenetic machinery in different disease subgroups, such as regulators of H3K27 and H3K4 trimethylation in subgroups 3 and 4 (for example, KDM6A and ZMYM3), and CTNNB1-associated chromatin re-modellers in WNT-subgroup tumours (for example, SMARCA4 and CREBBP). Modelling of mutations in mouse lower rhombic lip progenitors that generate WNT-subgroup tumours identified genes that maintain this cell lineage (DDX3X), as well as mutated genes that initiate (CDH1) or cooperate (PIK3CA) in tumorigenesis. These data provide important new insights into the pathogenesis of medulloblastoma subgroups and highlight targets for therapeutic development.
BACKGROUND:The myelodysplastic syndromes are a group of hematologic disorders that often evolve into secondary acute myeloid leukemia (AML). The genetic changes that underlie progression from the myelodysplastic syndromes to secondary AML are not well understood. METHODS:We performed whole-genome sequencing of seven paired samples of skin and bone marrow in seven subjects with secondary AML to identify somatic mutations specific to secondary AML. We then genotyped a bone marrow sample obtained during the antecedent myelodysplastic-syndrome stage from each subject to determine the presence or absence of the specific somatic mutations. We identified recurrent mutations in coding genes and defined the clonal architecture of each pair of samples from the myelodysplastic-syndrome stage and the secondary-AML stage, using the allele burden of hundreds of mutations. RESULTS:Approximately 85% of bone marrow cells were clonal in the myelodysplastic-syndrome and secondary-AML samples, regardless of the myeloblast count. The secondary-AML samples contained mutations in 11 recurrently mutated genes, including 4 genes that have not been previously implicated in the myelodysplastic syndromes or AML. In every case, progression to acute leukemia was defined by the persistence of an antecedent founding clone containing 182 to 660 somatic mutations and the outgrowth or emergence of at least one subclone, harboring dozens to hundreds of new mutations. All founding clones and subclones contained at least one mutation in a coding gene. CONCLUSIONS:Nearly all the bone marrow cells in patients with myelodysplastic syndromes and secondary AML are clonally derived. Genetic evolution of secondary AML is a dynamic process shaped by multiple cycles of mutation acquisition and clonal selection. Recurrent gene mutations are found in both founding clones and daughter subclones. (Funded by the National Institutes of Health and others.).
The Human Microbiome Project will establish a reference data set for analysis of the microbiome of healthy adults by surveying multiple body sites from 300 people and generating data from over 12,000 samples. To characterize these samples, the participating sequencing centers evaluated and adopted 16S rDNA community profiling protocols for ABI 3730 and 454 FLX Titanium sequencing. In the course of establishing protocols, we examined the performance and error characteristics of each technology, and the relationship of sequence error to the utility of 16S rDNA regions for classification- and OTU-based analysis of community structure. The data production protocols used for this work are those used by the participating centers to produce 16S rDNA sequence for the Human Microbiome Project. Thus, these results can be informative for interpreting the large body of clinical 16S rDNA data produced for this project.
We report the results of whole-genome and transcriptome sequencing of tumor and adjacent normal tissue samples from 17 patients with non-small cell lung carcinoma (NSCLC). We identified 3,726 point mutations and more than 90 indels in the coding sequence, with an average mutation frequency more than 10-fold higher in smokers than in never-smokers. Novel alterations in genes involved in chromatin modification and DNA repair pathways were identified, along with DACH1, CFTR, RELN, ABCB5, and HGF. Deep digital sequencing revealed diverse clonality patterns in both never-smokers and smokers. All validated EFGR and KRAS mutations were present in the founder clones, suggesting possible roles in cancer initiation. Analysis revealed 14 fusions, including ROS1 and ALK, as well as novel metabolic enzymes. Cell-cycle and JAK-STAT pathways are significantly altered in lung cancer, along with perturbations in 54 genes that are potentially targetable with currently available drugs.
Massively parallel sequencing technology and the associated rapidly decreasing sequencing costs have enabled systemic analyses of somatic mutations in large cohorts of cancer cases. Here we introduce a comprehensive mutational analysis pipeline that uses standardized sequence-based inputs along with multiple types of clinical data to establish correlations among mutation sites, affected genes and pathways, and to ultimately separate the commonly abundant passenger mutations from the truly significant events. In other words, we aim to determine the Mutational Significance in Cancer (MuSiC) for these large data sets. The integration of analytical operations in the MuSiC framework is widely applicable to a broad set of tumor types and offers the benefits of automation as well as standardization. Herein, we describe the computational structure and statistical underpinnings of the MuSiC pipeline and demonstrate its performance using 316 ovarian cancer samples from the TCGA ovarian cancer project. MuSiC correctly confirms many expected results, and identifies several potentially novel avenues for discovery.
Abstract Abstract 404 To characterize the genomic events associated with distinct subtypes of AML, we used whole genome sequencing to compare 24 tumor/normal sample pairs from patients with normal karyotype (NK) M1-AML (12 cases) and t(15;17)-positive M3-AML (12 cases). All single nucleotide variants (SNVs), small insertions and deletions (indels), and cryptic structural variants (SVs) identified by whole genome sequencing (average coverage 28x) were validated using sample-specific custom Nimblegen capture arrays, followed by Illumina sequencing; an average coverage of 972 reads per somatic variant yielded 10,597 validated somatic variants (average 421/genome). Of these somatic mutations, 308 occurred in 286 unique genes; on average, 9.4 somatic mutations per genome had translational consequences. Several important themes emerged: 1) AML genomes contain a diverse range of recurrent mutations. We assessed the 286 mutated genes for recurrency in an additional 34 NK M1-AML cases and 9 M3-AML cases. We identified 51 recurrently mutated genes, including 37 that had not previously been described in AML; on average, each genome had 3 recurrently mutated genes (M1 = 3.2; M3 = 2.8, p = 0.32). 2) Many recurring mutations cluster in mutually exclusive pathways, suggesting pathophysiologic importance. The most commonly mutated genes were: FLT3 (36%), NPM1 (25%), DNMT3A (21%), IDH1 (18%), IDH2 (10%), TET2 (10%), ASXL1 (6%), NRAS (6%), TTN (6%), and WT1 (6%). In total, 3 genes (excluding PML-RARA) were mutated exclusively in M3 cases. 22 genes were found only in M1 cases (suggestive of alternative initiating mutations which occurred in methylation, signal transduction, and cohesin complex genes). 25 genes were mutated in both M1 and M3 genomes (suggestive of common progression mutations relevant for both subtypes). A single mutation in a cell growth/signaling gene occurred in 38 of 67 cases (FLT3, NRAS, RUNX1, KIT, CACNA1E, CADM2, CSMD1); these mutations were mutually exclusive of one another, and many of them occurred in genomes with PML-RARA, suggesting that they are progression mutations. We also identified a new leukemic pathway: mutations were observed in all four genes that encode members of the cohesin complex (STAG2, SMC1A, SMC3, RAD21), which is involved in mitotic checkpoints and chromatid separation. The cohesin mutations were mutually exclusive of each other, and collectively occur in 10% of non-M3 AML patients. 3) AML genomes also contain hundreds of benign “passenger” mutations. On average 412 somatic mutations per genome were translationally silent or occurred outside of annotated genes. Both M1 and M3 cases had similar total numbers of mutations per genome, similar mutation types (which favored C>T/G>A transitions), and a similar random distribution of variants throughout the genome (which was affected neither by coding regions nor expression levels). This is consistent with our recent observations of random “passenger” mutations in hematopoietic stem cell (HSC) clones derived from normal patients (Ley et al manuscript in preparation), and suggests that most AML-associated mutations are not pathologic, but pre-existed in the HSC at the time of initial transformation. In both studies, the total number of SNVs per genome correlated positively with the age of the patient (R2 = 0.48, p = 0.001), providing a possible explanation for the increasing incidence of AML in elderly patients. 4) NK M1 and M3 AML samples are mono- or oligo-clonal. By comparing the frequency of all somatic mutations within each sample, we could identify clusters of mutations with similar frequencies (leukemic clones) and determined that the average number of clones per genome was 1.8 (M1 = 1.5; M3 = 2.2; p = 0.04). 5) t(15;17) is resolved by a non-homologous end-joining repair pathway, since nucleotide resolution of all 12 t(15;17) breakpoints revealed inconsistent micro-homologies (0 – 7 bp). Summary: These data provide a genome-wide overview of NK and t(15;17) AML and provide important new insights into AML pathogenesis. AML genomes typically contain hundreds of random, non-genic mutations, but only a handful of recurring mutated genes that are likely to be pathogenic because they cluster in mutually exclusive pathways; specific combinations of recurring mutations, as well as rare and private mutations, shape the leukemia phenotype in an individual patient, and help to explain the clinical heterogeneity of this disease. Disclosures: Westervelt: Novartis: Speakers Bureau.
CONTEXT:The identification of patients with inherited cancer susceptibility syndromes facilitates early diagnosis, prevention, and treatment. However, in many cases of suspected cancer susceptibility, the family history is unclear and genetic testing of common cancer susceptibility genes is unrevealing. OBJECTIVE:To apply whole-genome sequencing to a patient without any significant family history of cancer but with suspected increased cancer susceptibility because of multiple primary tumors to identify rare or novel germline variants in cancer susceptibility genes. DESIGN, SETTING, AND PARTICIPANT: Skin (normal) and bone marrow (leukemia) DNA were obtained from a patient with early-onset breast and ovarian cancer (negative for BRCA1 and BRCA2 mutations) and therapy-related acute myeloid leukemia (t-AML) and analyzed with the following: whole-genome sequencing using paired-end reads, single-nucleotide polymorphism (SNP) genotyping, RNA expression profiling, and spectral karyotyping. MAIN OUTCOME MEASURES:Structural variants, copy number alterations, single-nucleotide variants, and small insertions and deletions (indels) were detected and validated using the described platforms. RESULTS; Whole-genome sequencing revealed a novel, heterozygous 3-kilobase deletion removing exons 7-9 of TP53 in the patient's normal skin DNA, which was homozygous in the leukemia DNA as a result of uniparental disomy. In addition, a total of 28 validated somatic single-nucleotide variations or indels in coding genes, 8 somatic structural variants, and 12 somatic copy number alterations were detected in the patient's leukemia genome. CONCLUSION:Whole-genome sequencing can identify novel, cryptic variants in cancer susceptibility genes in addition to providing unbiased information on the spectrum of mutations in a cancer genome.