Additional file 5 Supplementary Table S5.
Additional file 3 Supplementary Table S3.
Large-scale cancer genome studies suggest that tumors are driven by somatic copy number alterations (SCNAs) or single-nucleotide variants (SNVs). Due to the low-cost, the clinical use of genomics assays is biased towards targeted gene panels, which identify SNVs. There is a need for a comparably low-cost and simple assay for high-resolution SCNA profiling. Here we present our method, conliga, which infers SCNA profiles from a low-cost and simple assay.
Additional file 2 Supplementary Table S2.
The poor outcomes in esophageal adenocarcinoma (EAC) prompted us to interrogate the pattern and timing of metastatic spread. Whole-genome sequencing and phylogenetic analysis of 388 samples across 18 individuals with EAC showed, in 90% of patients, that multiple subclones from the primary tumor spread very rapidly from the primary site to form multiple metastases, including lymph nodes and distant tissues—a mode of dissemination that we term ‘clonal diaspora’. Metastatic subclones at autopsy were present in tissue and blood samples from earlier time points. These findings have implications for our understanding and clinical evaluation of EAC. Whole-genome sequencing and phylogenetic analysis of nearly 400 esophageal adenocarcinoma samples suggest a model of rapid subclone spreading from the primary tumor to multiple metastatic sites, which the authors term ‘clonal diaspora’.
Cancers occurring at the gastroesophageal junction (GEJ) are classified as predominantly esophageal or gastric, which is often difficult to decipher. We hypothesized that the transcriptomic profile might reveal molecular subgroups which could help to define the tumor origin and behavior beyond anatomical location. The gene expression profiles of 107 treatment‐naïve, intestinal type, gastroesophageal adenocarcinomas were assessed by the Illumina‐HTv4.0 beadchip. Differential gene expression (limma), unsupervised subgroup assignment (mclust) and pathway analysis (gage) were undertaken in R statistical computing and results were related to demographic and clinical parameters. Unsupervised assignment of the gene expression profiles revealed three distinct molecular subgroups, which were not associated with anatomical location, tumor stage or grade (p > 0.05). Group 1 was enriched for pathways involved in cell turnover, Group 2 was enriched for metabolic processes and Group 3 for immune‐response pathways. Patients in group 1 showed the worst overall survival (p = 0.019). Key genes for the three subtypes were confirmed by immunohistochemistry. The newly defined intrinsic subtypes were analyzed in four independent datasets of gastric and esophageal adenocarcinomas with transcriptomic data available (RNAseq data: OCCAMS cohort, n = 158; gene expression arrays: Belfast, n = 63; Singapore, n = 191; Asian Cancer Research Group, n = 300). The subgroups were represented in the independent cohorts and pooled analysis confirmed the prognostic effect of the new subtypes. In conclusion, adenocarcinomas at the GEJ comprise three distinct molecular phenotypes which do not reflect anatomical location but rather inform our understanding of the key pathways expressed.
Esophageal adenocarcinoma (EAC) is a poor-prognosis cancer type with rapidly rising incidence. Understanding of the genetic events driving EAC development is limited, and there are few molecular biomarkers for prognostication or therapeutics. Using a cohort of 551 genomically characterized EACs with matched RNA sequencing data, we discovered 77 EAC driver genes and 21 noncoding driver elements. We identified a mean of 4.4 driver events per tumor, which were derived more commonly from mutations than copy number alterations, and compared the prevelence of these mutations to the exome-wide mutational excess calculated using non-synonymous to synonymous mutation ratios ( dN / dS ). We observed mutual exclusivity or co-occurrence of events within and between several dysregulated EAC pathways, a result suggestive of strong functional relationships. Indicators of poor prognosis ( SMAD4 and GATA4 ) were verified in independent cohorts with significant predictive value. Over 50% of EACs contained sensitizing events for CDK4 and CDK6 inhibitors, which were highly correlated with clinically relevant sensitivity in a panel of EAC cell lines and organoids.
Esophageal adenocarcinoma (EAC) incidence is increasing while 5-year survival rates remain less than 15%. A lack of experimental models has hampered progress. We have generated clinically annotated EAC organoid cultures that recapitulate the morphology, genomic, and transcriptomic landscape of the primary tumor including point mutations, copy number alterations, and mutational signatures. Karyotyping of organoid cultures has confirmed polyclonality reflecting the clonal architecture of the primary tumor. Furthermore, subclones underwent clonal selection associated with driver gene status. Medium throughput drug sensitivity testing demonstrates the potential of targeting receptor tyrosine kinases and downstream mediators. EAC organoid cultures provide a pre-clinical tool for studies of clonal evolution and precision therapeutics.
Continual evolution of cancer on a microscopic scale makes accurate clinical assessment challenging. The highly varied and unpredictable patient outcomes in esophageal adenocarcinoma (EAC) prompted us to question conventional modes of metastatic spread. Whole genome sequencing and phylogenetic analysis of 396 samples across 18 EAC cases demonstrated contemporaneous polyclonal seeding from the primary to multiple disparate sites at a single time-point in 90% cases- so-called clonal diaspora. The age-dependent trinucleotide signature was rarely observed post diaspora supporting a near-synchronous metastatic seeding process. Clustering of lymph nodes and distant metastases (n=250) demonstrated that samples sharing a common clonal origin were widely dispersed spatially. Metastatic subclones at autopsy were present in the primary tumor at diagnosis, and in blood samples at earlier time-points. This is consistent with our hypothesis that metastasis occurs at a single time-point and has far-reaching implications of our understanding of metastatic progression, clinical staging and patient management.
Introduction: Epithelial-mesenchymal transition (EMT) may occur early in the malignant transformation of Barrett's esophagus.In previous studies, we found that Barrett's epithelial cells exposed to acidic bile salts increased their expression of EMT markers, and exhibited increased nuclear levels of hypoxia inducible factor (HIF)-1α, a transcription factor that can induce EMT.Conversely, members of the microRNA (miR)-200 family can prevent cells from undergoing EMT, and our earlier studies showed that acidic bile salts decreased miR-200a&b expression in Barrett's cells by suppressing activity of the miR-200b/a/429 promoter independent of HIF-1α binding a canonical HIF-DNA binding element.Now, we have explored mechanisms whereby HIF-1α that is induced by acidic bile salts suppresses miR-200 promoter activity in Barrett's cells.Methods: Telomerase-immortalized, non-neoplastic Barrett's epithelial cells (BAR-10T) were treated with acidic medium (pH 5.5) containing conjugated bile salts (400 µM) for 5-30 minutes.Cells were transfected with: 1) HIF-1α shRNA (to knockdown HIF-1α) or a constitutively-active ODD mutant (to overexpress HIF-1α), 2) a series of miR-200b/a/429 promoter deletion constructs, 3) miR-200b/a/429 promoter construct (-871 to -834 bp) with or without a point mutation at the binding site for specificity protein (SP)1 (-862 to -852 bp) or Myc-associated zinc finger protein (MAZ) (-850 to -835 bp), 3) SP1 siRNA, or 4) MAZ shRNA.We determined: 1) promoter activity by a luciferase reporter assay, 2) HIF-1α, SP1, and MAZ interactions by nuclear co-IP, and 3) DNA binding of SP1 or MAZ by ChIP.Results: We identified a minimal acidic bile saltresponsive region of the miR-200b/a/429 promoter at -871 to -834 bp upstream of the transcription initiation site.HIF-1α overexpression significantly decreased activity of this promoter region, as did acidic bile salts, and this effect of acidic bile salts was blocked by HIF-1α knockdown.Within the acidic bile salt-responsive region, we identified two putative transcription factor binding sites: SP1 and MAZ.Acidic bile salts increased: 1) nuclear levels of SP1 and MAZ proteins, 2) binding of SP1 and MAZ proteins to HIF-1α in the nucleus, and 3) binding of SP1 and MAZ to DNA in the miR-200b/a/429 promoter.A point mutation in either the SP1 or MAZ binding site prevented acidic bile salts and HIF-1α overexpression from suppressing promoter activity.Knockdown of SP1 or MAZ eliminated the decrease in miR-200b/a/429 promoter activity induced by acidic bile salts.Conclusions: In Barrett's epithelial cells, acidic bile salts decrease expression of miR-200a&b through HIF-1α, which complexes with SP1 and MAZ proteins to suppress activity of the miR-200b/a/429 promoter.These findings elucidate molecular mechanisms whereby refluxed acid and bile can induce EMT in Barrett's esophagus.
Esophageal adenocarcinoma (EAC) develops in an inflammatory microenvironment with reduced microbial diversity, but mechanisms for these influences remain poorly characterized. We hypothesized that mutations targeting the Toll-like receptor (TLR) pathway could disrupt innate immune signaling and promote a microenvironment that favors tumorigenesis. Through interrogating whole genome sequencing data from 171 EAC patients, we showed that non-synonymous mutations collectively affect the TLR pathway in 25/171 (14.6%, PathScan p = 8.7x10-5) tumors. TLR mutant cases were associated with more proximal tumors and metastatic disease, indicating possible clinical significance of these mutations. Only rare mutations were identified in adjacent Barrett's esophagus samples. We validated our findings in an external EAC dataset with non-synonymous TLR pathway mutations in 33/149 (22.1%, PathScan p = 0.05) tumors, and in other solid tumor types exposed to microbiomes in the COSMIC database (10,318 samples), including uterine endometrioid carcinoma (188/320, 58.8%), cutaneous melanoma (377/988, 38.2%), colorectal adenocarcinoma (402/1519, 26.5%), and stomach adenocarcinoma (151/579, 26.1%). TLR4 was the most frequently mutated gene with eleven mutations in 10/171 (5.8%) of EAC tumors. The TLR4 mutants E439G, S570I, F703C and R787H were confirmed to have impaired reactivity to bacterial lipopolysaccharide with marked reductions in signaling by luciferase reporter assays. Overall, our findings show that TLR pathway genes are recurrently mutated in EAC, and TLR4 mutations have decreased responsiveness to bacterial lipopolysaccharide and may play a role in disease pathogenesis in a subset of patients.
Nat. Genet.; 10.1038/ng.3659; corrected online 19 September 2016 In the version of this article initially published online, the mutation signature illustrations for S1 and S2 in Figure 3a were switched. Additionally, in the Online Methods, the text originally stated that structural variants were called using BWA-MEM, when it should have stated that these were called using BWA.
The scientific community has avoided using tissue samples from patients that have been exposed to systemic chemotherapy to infer the genomic landscape of a given cancer. Esophageal adenocarcinoma is a heterogeneous, chemoresistant tumor for which the availability and size of pretreatment endoscopic samples are limiting. This study compares whole-genome sequencing data obtained from chemo-naive and chemo-treated samples. The quality of whole-genomic sequencing data is comparable across all samples regardless of chemotherapy status. Inclusion of samples collected post-chemotherapy increased the proportion of late-stage tumors. When comparing matched pre- and post-chemotherapy samples from 10 cases, the mutational signatures, copy number, and SNV mutational profiles reflect the expected heterogeneity in this disease. Analysis of SNVs in relation to allele-specific copy-number changes pinpoints the common ancestor to a point prior to chemotherapy. For cases in which pre- and post-chemotherapy samples do show substantial differences, the timing of the divergence is near-synchronous with endoreduplication. Comparison across a large prospective cohort (62 treatment-naive, 58 chemotherapy-treated samples) reveals no significant differences in the overall mutation rate, mutation signatures, specific recurrent point mutations, or copy-number events in respect to chemotherapy status. In conclusion, whole-genome sequencing of samples obtained following neoadjuvant chemotherapy is representative of the genomic landscape of esophageal adenocarcinoma. Excluding these samples reduces the material available for cataloging and introduces a bias toward the earlier stages of cancer.
Esophageal adenocarcinoma (EAC) is highly mutated and molecularly heterogeneous. The number of cell lines available for study is limited and their genome has been only partially characterized. The availability of an accurate annotation of their mutational landscape is crucial for accurate experimental design and correct interpretation of genotype-phenotype findings. We performed high coverage, paired end whole genome sequencing on eight EAC cell lines—ESO26, ESO51, FLO-1, JH-EsoAd1, OACM5.1 C, OACP4 C, OE33, SK-GT-4—all verified against original patient material, and one esophageal high grade dysplasia cell line, CP-D. We have made available the aligned sequence data and report single nucleotide variants (SNVs), small insertions and deletions (indels), and copy number alterations, identified by comparison with the human reference genome and known single nucleotide polymorphisms (SNPs). We compare these putative mutations to mutations found in primary tissue EAC samples, to inform the use of these cell lines as a model of EAC.
As whole-genome sequencing for cancer genome analysis becomes a clinical tool, a full understanding of the variables affecting sequencing analysis output is required. Here using tumour-normal sample pairs from two different types of cancer, chronic lymphocytic leukaemia and medulloblastoma, we conduct a benchmarking exercise within the context of the International Cancer Genome Consortium. We compare sequencing methods, analysis pipelines and validation methods. We show that using PCR-free methods and increasing sequencing depth to ∼100 × shows benefits, as long as the tumour:control coverage ratio remains balanced. We observe widely varying mutation call rates and low concordance among analysis pipelines, reflecting the artefact-prone nature of the raw data and lack of standards for dealing with the artefacts. However, we show that, using the benchmark mutation set we have created, many issues are in fact easy to remedy and have an immediate positive impact on mutation detection accuracy.
The emergence of next generation DNA sequencing technology is enabling high-resolution cancer genome analysis. Large-scale projects like the International Cancer Genome Consortium (ICGC) are systematically scanning cancer genomes to identify recurrent somatic mutations. Second generation DNA sequencing, however, is still an evolving technology and procedures, both experimental and analytical, are constantly changing. Thus the research community is still defining a set of best practices for cancer genome data analysis, with no single protocol emerging to fulfil this role. Here we describe an extensive benchmark exercise to identify and resolve issues of somatic mutation calling. Whole genome sequence datasets comprising tumor-normal pairs from two different types of cancer, chronic lymphocytic leukaemia and medulloblastoma, were shared within the ICGC and submissions of somatic mutation calls were compared to verified mutations and to each other. Varying strategies to call mutations, incomplete awareness of sources of artefacts, and even lack of agreement on what constitutes an artefact or real mutation manifested in widely varying mutation call rates and somewhat low concordance among submissions. We conclude that somatic mutation calling remains an unsolved problem. However, we have identified many issues that are easy to remedy that are presented here. Our study highlights critical issues that need to be addressed before this valuable technology can be routinely used to inform clinical decision-making. SSM : Somatic Single-base Mutations or Simple Somatic Mutations, refers to a somatic single base change SIM : Somatic Insertion/deletion Mutation CNV : Copy Number Variant SV : Structural Variant SNP : Single Nucleotide Polymorphisms, refers to a single base variable position in the germline with a frequency of > 1% in the general population CLL : Chronic Lymphocytic Leukaemia MB : Medulloblastoma ICGC : International Cancer Genome Consortium BM : Benchmark Abbreviations and Definitions aligner = mapper, these terms are used interchangeably
The European Nucleotide Archive (ENA; http://www.ebi.ac.uk/ena/) collects, maintains and presents comprehensive nucleic acid sequence and related information as part of the permanent public scientific record. Here, we provide brief updates on ENA content developments and major service enhancements in 2012 and describe in more detail two important areas of development and policy that are driven by ongoing growth in sequencing technologies. First, we describe the ENA data warehouse, a resource for which we provide a programmatic entry point to integrated content across the breadth of ENA. Second, we detail our plans for the deployment of CRAM data compression technology in ENA.