Traditional gene expression deconvolution methods assess a limited number of cell types, therefore do not capture the full complexity of the tumor microenvironment (TME). Here, we integrate nine deconvolution tools to assess 79 TME cell types in 10,592 tumors across 33 different cancer types, creating the most comprehensive analysis of the TME. In total, we found 41 patterns of immune infiltration and stroma profiles, identifying heterogeneous yet unique TME portraits for each cancer and several new findings. Our findings indicate that leukocytes play a major role in distinguishing various tumor types, and that a shared immune-rich TME cluster predicts better survival in bladder cancer for luminal and basal squamous subtypes, as well as in melanoma for RAS-hotspot subtypes. Our detailed deconvolution and mutational correlation analyses uncover 35 therapeutic target and candidate response biomarkers hypotheses (including CASP8 and RAS pathway genes).
In solid tumor oncology, circulating tumor DNA (ctDNA) is poised to transform care through accurate assessment of minimal residual disease (MRD) and therapeutic response monitoring. To overcome the sparsity of ctDNA fragments in low tumor fraction (TF) settings and increase MRD sensitivity, we previously leveraged genome-wide mutational integration through plasma whole-genome sequencing (WGS). Here we now introduce MRD-EDGE, a machine-learning-guided WGS ctDNA single-nucleotide variant (SNV) and copy-number variant (CNV) detection platform designed to increase signal enrichment. MRD-EDGESNV uses deep learning and a ctDNA-specific feature space to increase SNV signal-to-noise enrichment in WGS by ~300× compared to previous WGS error suppression. MRD-EDGECNV also reduces the degree of aneuploidy needed for ultrasensitive CNV detection through WGS from 1 Gb to 200 Mb, vastly expanding its applicability within solid tumors. We harness the improved performance to identify MRD following surgery in multiple cancer types, track changes in TF in response to neoadjuvant immunotherapy in lung cancer and demonstrate ctDNA shedding in precancerous colorectal adenomas. Finally, the radical signal-to-noise enrichment in MRD-EDGESNV enables plasma-only (non-tumor-informed) disease monitoring in advanced melanoma and lung cancer, yielding clinically informative TF monitoring for patients on immune-checkpoint inhibition. Detection of circulating tumor DNA using MRD-EDGE, a machine-learning-guided single-nucleotide variant and copy-number variant detection platform for signal enrichment, enables monitoring of minimal residual disease and immunotherapy response in settings of low tumor burden.
Immune checkpoint blockade (ICB) has transformed the treatment of metastatic cancer but is hindered by variable response rates. A key unmet need is the identification of biomarkers that predict treatment response. To address this, we analyzed six whole exome sequencing cohorts with matched disease outcomes to identify genes and pathways predictive of ICB response. To increase detection power, we focus on genes and pathways that are significantly mutated following correction for epigenetic, replication timing, and sequence-based covariates. Using this technique, we identify several genes ( BCLAF1, KRAS, BRAF , and TP53) and pathways (MAPK signaling, p53 associated, and immunomodulatory) as predictors of ICB response and develop the Cancer Immunotherapy Response CLassifiEr (CIRCLE). Compared to tumor mutational burden alone, CIRCLE led to superior prediction of ICB response with a 10.5% increase in sensitivity and a 11% increase in specificity. We envision that CIRCLE and more broadly the analysis of recurrently mutated cancer genes will pave the way for better prognostic tools for cancer immunotherapy.
High-order three-dimensional (3D) interactions between more than two genomic loci are common in human chromatin, but their role in gene regulation is unclear. Previous high-order 3D chromatin assays either measure distant interactions across the genome or proximal interactions at selected targets. To address this gap, we developed Pore-C, which combines chromatin conformation capture with nanopore sequencing of concatemers to profile proximal high-order chromatin contacts at the genome scale. We also developed the statistical method Chromunity to identify sets of genomic loci with frequencies of high-order contacts significantly higher than background (‘synergies’). Applying these methods to human cell lines, we found that synergies were enriched in enhancers and promoters in active chromatin and in highly transcribed and lineage-defining genes. In prostate cancer cells, these included binding sites of androgen-driven transcription factors and the promoters of androgen-regulated genes. Concatemers of high-order contacts in highly expressed genes were demethylated relative to pairwise contacts at the same loci. Synergies in breast cancer cells were associated with tyfonas, a class of complex DNA amplicons. These results rigorously link genome-wide high-order 3D interactions to lineage-defining transcriptional programs and establish Pore-C and Chromunity as scalable approaches to assess high-order genome structure.
Uterine leiomyosarcomas (uLMS) are aggressive tumors arising from the smooth muscle layer of the uterus. We analyzed 83 uLMS sample genetics, including 56 from Yale and 27 from The Cancer Genome Atlas (TCGA). Among them, a total of 55 Yale samples including two patient-derived xenografts (PDXs) and 27 TCGA samples have whole-exome sequencing (WES) data; 10 Yale and 27 TCGA samples have RNA-sequencing (RNA-Seq) data; and 11 Yale and 10 TCGA samples have whole-genome sequencing (WGS) data. We found recurrent somatic mutations in TP53, MED12, and PTEN genes. Top somatic mutated genes included TP53, ATRX, PTEN, and MEN1 genes. Somatic copy number variation (CNV) analysis identified 8 copy-number gains, including 5p15.33 (TERT), 8q24.21 (C-MYC), and 17p11.2 (MYOCD, MAP2K4) amplifications and 29 copy-number losses. Fusions involving tumor suppressors or on-cogenes were deetected, with most fusions disrupting RB1, TP53, and ATRX/DAXX, and one fusion (ACTG2-ALK) being po-tentially targetable. WGS results demonstrated that 76% (16 of 21) of the samples harbored chromoplexy and/or chromothrip-sis. Clinically actionable mutational signatures of homologous-recombination DNA-repair deficiency (HRD) and microsatellite instability (MSI) were identified in 25% (12 of 48) and 2% (1 of 48) of fresh frozen uLMS, respectively. Finally, we found olaparib (PARPi; P = 0.002), GS-62 6510 (C-MYC/BETi; P < 0.000001 and P = 0.0005), and copanlisib (PIK3CAi; P = 0.0001) monotherapy to significantly inhibit uLMS-PDXs harboring de-rangements in C-MYC and PTEN/PIK3CA/AKT genes (LEY11) and/ or HRD signatures (LEY16) compared to vehicle-treated mice. These findings define the genetic landscape of uLMS and sug-gest that a subset of uLMS may benefit from existing PARP-, PIK3CA-, and C-MYC/BET-targeted drugs.
Summary Recent pan-cancer studies have delineated patterns of structural genomic variation across thousands of tumor whole genome sequences. It is not known to what extent the shortcomings of short read (≤ 150 bp) whole genome sequencing (WGS) used for structural variant analysis has limited our understanding of cancer genome structure. To formally address this, we introduce the concept of “loose ends” - copy number alterations that cannot be mapped to a rearrangement by WGS but can be indirectly detected through the analysis of junction-balanced genome graphs. Analyzing 2,319 pan-cancer WGS cases across 31 tumor types, we found loose ends were enriched in reference repeats and fusions of the mappable genome to repetitive or foreign sequences. Among these we found genomic footprints of neotelomeres, which were surprisingly enriched in cancers with low telomerase expression and alternate lengthening of telomeres phenotype. Our results also provide a rigorous upper bound on the role of non-allelic homologous recombination (NAHR) in large-scale cancer structural variation, while nominating INO80 , FANCA , and ARID1A as positive modulators of somatic NAHR. Taken together, we estimate that short read WGS maps >97% of all large-scale (>10 kbp) cancer structural variation; the rest represent loose ends that require long molecule profiling to unambiguously resolve. Our results have broad relevance for future research and clinical applications of short read WGS and delineate precise directions where long molecule studies might provide transformative insight into cancer genome structure.
Abstract Lung adenocarcinomas (LUAD) are typically characterized by genetic activation of the receptor tyrosine kinase (RTK)/RAS/RAF/MAP kinase (MAPK) pathway. A minority of LUAD cases (20-25%) lack apparent genetic alterations in this pathway, and thus are ineligible for most targeted therapies. These candidate “oncogene negative” LUADs may harbor novel classes of oncogenic drivers or represent a biologically distinct class of tumors. To characterize the genomic landscape of oncogene-negative LUADs, we nominated 98 cases that were found to lack an activating RTK/RAS/RAF/MAPK pathway alteration in a TCGA study utilizing whole exome sequencing, microarray, and transcriptome data. We profiled these tumors with high-depth whole genome sequencing (WGS), with the goal of identifying noncoding and structural variant driver DNA alterations in both known and novel loci. Of the 98 cases, 20 harbored somatic KRAS mutations that had been missed in the prior WES and transcriptome studies because of insufficient coverage, including 8 cases with the recently targetable p.G12C mutation. 16 samples harbored oncogenic or loss-of-function structural variants in FGFR1, MAPK1, EGFR, NF1, RASA1, ARAF, NTRK2 and NRG1. 5 other samples with SNV or indels in EGFR, ERBB2 and SOS1 were reclassified as oncogene positive. Thus via comprehensive genomic analysis, we confirmed that 57 of the 98 WGS cases did not harbor any detectable alterations in genes encoding any known RTK/RAS/RAF/MAPK members, representing 13% cases chosen as “lung adenocarcinomas” for the TCGA study. Among the 57 confirmed oncogene-negative LUADs, we identified focal deletions targeting the promoter and transcription start site of tumor suppressor genes STK11, KEAP1 and SMARCA4 in 10 samples. Expression and methylation profiling suggested an enrichment of the TP53-deficient phenotype, including cell cycle and FOXM1 deregulation, among the oncogene-negative samples. Moreover, novel promoter mutations associated with increased expression were identified in ILF2, which regulates DNA damage response pathways. Finally, a subset of confirmed oncogene-negative LUADs harbored increased expression of neuroendocrine markers, suggesting that these oncogene-negative samples may either be mis-diagnosed as LUAD or represent LUAD with mixed features of other subtypes of lung cancer; indeed, 14 of the 57 confirmed oncogene-negative cases show histological features of large cell neuroendocrine lung carcinoma. This would suggest that 10% of the cases in this study are both lung adenocarcinoma and “oncogene-negative” to date. Our results provide some of the first comprehensive genomic characterization of oncogene-negative LUADs, implicating TP53 and structural variants in the pathogenesis of this common and difficult to treat entity. Citation Format: Jian Carrot-Zhang, Siddhartha Devarakonda, Nicolas Robine, Xiaotong Yao, Tiago C. Silva, Jeff Damrauer, Aditya Deshpande, Ming-Sound Tsao, Christina Yao, Chris Wong, Lisui Bao, Hyo Young Choi, Ina Felau, Jean C. Zenklusen, Gordon Robertson, Tuan Trieua, Wei-Wei Liang, Meng Zhou, Esther Rheinbay, Neil Hayes, Ekta Khurana, Li Ding, Peter Laird, Olivier Elemento, John Weinstein, David Kwiatkowski, Chris Benz, Josh Stuart, Lixing Yang, Mauro Castro, William Travis, Katherine Hoadley, Ben Berman, TCGA Analysis Network, Matthew Meyerson, Ramaswamy Govindan, Marcin Imielinski. Whole-genome characterization of lung adenocarcinomas lacking alterations in RTK/RAS/RAF/MAPK pathway [abstract]. In: Proceedings of the Annual Meeting of the American Association for Cancer Research 2020; 2020 Apr 27-28 and Jun 22-24. Philadelphia (PA): AACR; Cancer Res 2020;80(16 Suppl):Abstract nr 5895.
Cancer genomes often harbor hundreds of somatic DNA rearrangement junctions, many of which cannot be easily classified into simple (e.g., deletion) or complex (e.g., chromothripsis) structural variant classes. Applying a novel genome graph computational paradigm to analyze the topology of junction copy number (JCN) across 2,778 tumor whole-genome sequences, we uncovered three novel complex rearrangement phenomena: pyrgo, rigma, and tyfonas. Pyrgo are "towers" of low-JCN duplications associated with early-replicating regions, superenhancers, and breast or ovarian cancers. Rigma comprise "chasms" of low-JCN deletions enriched in late-replicating fragile sites and gastrointestinal carcinomas. Tyfonas are "typhoons" of high-JCN junctions and fold-back inversions associated with expressed protein-coding fusions, breakend hypermutation, and acral, but not cutaneous, melanomas. Clustering of tumors according to genome graph-derived features identified subgroups associated with DNA repair defects and poor prognosis.
Higher-order chromatin structure arises from the combinatorial physical interactions of many genomic loci. To investigate this aspect of genome architecture we developed Pore-C, which couples chromatin conformation capture with Oxford Nanopore Technologies (ONT) long reads to directly sequence multi-way chromatin contacts without amplification. In GM12878, we demonstrate that the pairwise interaction patterns implicit in Pore-C multi-way contacts are consistent with gold standard Hi-C pairwise contact maps at the compartment, TAD, and loop scales. In addition, Pore-C also detects higher-order chromatin structure at 18.5-fold higher efficiency and greater fidelity than SPRITE, a previously published higher-order chromatin profiling technology. We demonstrate Pore-C’s ability to detect and visualize multi-locus hubs associated with histone locus bodies and active / inactive nuclear compartments in GM12878. In the breast cancer cell line HCC1954, Pore-C contacts enable the reconstruction of complex and aneuploid rearranged alleles spanning multiple megabases and chromosomes. Finally, we apply Pore-C to generate a chromosome scale de novo assembly of the HG002 genome. Our results establish Pore-C as the most simple and scalable assay for the genome-wide assessment of combinatorial chromatin interactions, with additional applications for cancer rearrangement reconstruction and de novo genome assembly.
Sensitive detection of somatic copy number alterations (SCNA) in cancer genomes is confounded by “waviness” in read depth data. We present dryclean , a signal processing algorithm to optimize SCNA detection in whole genome (WGS) and targeted sequencing platforms through foreground detection and background subtraction of read depth data. Application of dryclean to WGS demonstrates that WGS waviness is driven by replication timing. Re-analysis of thousands of tumor profiles reveals that dryclean provides superior detection of biologically relevant SCNAs relative to state-of-the-art algorithms. Applied to in silico tumor dilutions, dryclean improves the sensitivity of relapse detection 10-fold relative to current standards. dryclean is available as an R package in the GitHub repository https://github.com/mskilab/dryclean
Lung adenocarcinoma (LUAD) and squamous cell carcinoma (LUSC) are the most prevalent types of non-all cell lung cancer (NSCLC), a leading cause of cancer death worldwide. In this study, we analyzed the transcriptomes of ~45,000 single cells (scRNA) from 13 NSCLC patients, including 5 LUAD cases which were collected and profiled at our institution. To correlate genomes and transcriptomes we performed Whole-Genome Sequencing (WGS) on 3 of these 5 LUAD cases. By comparing tumor tissue with matched adjacent non-malignant lung tissue we are able to confidently distinguish 13 cell-type specific clusters that unambiguously match previously characterized lineages. We developed algorithms for the identification of malignant cells derived from tumor tissue through scRNA analysis of copy number alterations and single nucleotide variants (SNV). Joint analysis of WGS and scRNA confirmed an enrichment of tobacco-associated SNVs among malignant cells of the tumor. Stromal cell types demonstrated consistent expression patterns across cases, while malignant cells demonstrated both inter- and intra-tumoral heterogeneity in their expression of signatures related to GPCR signaling, 3’ UTR mediated translational regulation, and cell-cell junction organization. In particular, one case displayed a unique pattern of intra-tumoral heterogeneity, as a subset of malignant cells robustly express a marker of pulmonary neuroendocrine cells, CGRP. Employing immunohistochemistry, the spatial organization of these malignant cells is revealed to be mutually exclusive within the tumor microenvironment and overlapping in expression of clinical markers of small-cell lung cancer. Finally, we deconvolved bulk TCGA LUAD and LUSC gene expression samples and analyzed the relationship between cell type specific gene expression in cell types of the lung and passenger mutation topographies. Our results provide insight into the molecular and clinical correlates of deconvolved NSCLC transcriptomes and provide a novel methodology with which to explore genomic variation at a single cell resolution. Furthermore, our dataset provides a resource for illuminating cancer-cell transcriptional changes and revealing key molecular drivers of tumor-stromal interactions in lung cancer. Citation Format: Kofi E. Gyan, Aditya Deshpande, Shaham Beg, Huasong Tian, Joel Rosiene, Marlon Stoeckius, Peter Smibert, Davide Risso, Juan Miguel Mosquera, Marcin Imielinski. Single-cell transcriptomic profiling of non-small cell lung cancer uncovers inter- and intracell population structure across TCGA lung adenocarcinoma and lung squamous cancer subtypes [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2019; 2019 Mar 29-Apr 3; Atlanta, GA. Philadelphia (PA): AACR; Cancer Res 2019;79(13 Suppl):Abstract nr 909.
Cancer genomes often harbor hundreds of somatic DNA rearrangement junctions, many of which cannot be easily classified into simple (e.g. deletion, translocation) or complex (e.g. chromothripsis, chromoplexy) structural variant classes. Applying a novel genome graph computational paradigm to analyze the topology of junction copy number (JCN) across 2,833 tumor whole genome sequences (WGS), we introduce three complex rearrangement phenomena: pyrgo, rigma , and tyfonas . Pyrgo are “towers” of low-JCN duplications associated with early replicating regions and superenhancers, and are enriched in breast and ovarian cancers. Rigma comprise “chasms” of low-JCN deletions at late-replicating fragile sites in esophageal and other gastrointestinal (GI) adenocarcinomas. Tyfonas are “typhoons” of high-JCN junctions and fold back inversions that are enriched in acral but not cutaneous melanoma and associated with a previously uncharacterized mutational process of non-APOBEC kataegis. Clustering of tumors according to genome graph-derived features identifies subgroups associated with DNA repair defects and poor prognosis.
BACKGROUND:'Next-generation' (NGS) sequencing has wide application in medical genetics, including the detection of somatic variation in cancer. The Ion Torrent-based (IONT) platform is among NGS technologies employed in clinical, research and diagnostic settings. However, identifying mutations from IONT deep sequencing with high confidence has remained a challenge. We compared various computational variant-calling methods to derive a variant identification pipeline that may improve the molecular diagnostic and research utility of IONT.RESULTS:Using IONT, we surveyed variants from the 409-gene Comprehensive Cancer Panel in whole-section tumors, intra-tumoral biopsies and matched normal samples obtained from frozen tissues and blood from four early-stage non-small cell lung cancer (NSCLC) patients. We used MuTect, Varscan2, IONT's proprietary Ion Reporter, and a simple subtraction we called "Poor Man's Caller." Together these produced calls at 637 loci across all samples. Visual validation of 434 called variants was performed, and performance of the methods assessed individually and in combination. Of the subset of inspected putative variant calls (n=223) in genomic regions that were not intronic or intergenic, 68 variants (30%) were deemed valid after visual inspection. Among the individual methods, the Ion Reporter method offered perhaps the most reasonable tradeoffs. Ion Reporter captured 83% of all discovered variants; 50% of its variants were visually validated. Aggregating results from multiple packages offered varied improvements in performance.CONCLUSIONS:Overall, Ion Reporter offered the most attractive performance among the individual callers. This study suggests combined strategies to maximize sensitivity and positive predictive value in variant calling using IONT deep sequencing.
OBJECTIVE:Uterine carcinosarcoma (UCS) is a rare and aggressive form of uterine cancer. It is bi-phasic, exhibiting histological features of both malignant epithelial (carcinoma) and mesenchymal (sarcoma) elements, reflected in ambiguity in accepted treatment guidelines. We sought to study the genomic and transcriptomic profiles of these elements individually to gain further insights into the development of these tumors.METHODS:We macro-dissected carcinomatous, sarcomatous, and normal tissues from formalin fixed paraffin embedded uterine samples of 10 UCS patients. Single nucleotide polymorphism microarrays, targeted DNA sequencing and whole-transcriptome RNA-sequencing were performed. Somatic chromosomal alterations (SCAs), point mutation and gene expression profiles were compared between carcinomatous and sarcomatous components.RESULTS:In addition to TP53, other recurrently mutated genes harboring putative driver or loss-of-function mutations included PTEN, FBXW7, FGFR2, KRAS, PIK3CA and CTNNB1, genes known to be involved in UCS. Intra-patient somatic mutation and SCA profiles were highly similar between paired carcinoma and sarcoma samples. An epithelial-mesenchymal transition (EMT) signature tended to differentiate components, with EMT-like status more common in advanced-stage patients exhibiting higher inter-component SCA heterogeneity.CONCLUSIONS:From DNA analysis, our results indicate a monoclonal disease origin for this cohort. Yet expression-derived EMT statuses of the carcinomatous and sarcomatous components were often discrepant, and advanced cases displayed greater genomic heterogeneity. Therefore, separately-profiled components of UCS tumors may better inform disease progression or potential.
Hispanics with acute leukemias have poorer outcomes than non-Hispanic whites (NHWs), despite an increased likelihood of favorable prognostic features. We reviewed medical records from 167 children ages 0-18 years diagnosed with de novo AML over an 18-year period at Texas Children's Cancer Center, among whom 129 self-identified as Hispanic or NHW. Although Hispanics were significantly more likely to have the favorable prognostic cytogenetic feature t(8;21) (P = 0.04), the expected survival benefit was not observed. This lack of survival benefit was primarily due to significantly poorer event-free and overall survival among Hispanics treated with upfront stem cell transplantation after achieving first clinical remission (P = 0.008).
Abstract Uterine carcinosarcoma (UCS) is a rare and aggressive form of uterine cancer. It is bi-phasic, exhibiting histological features of both malignant epithelial (carcinomatous) and mesenchymal (sarcomatous) elements. Studies have indicated that UCS arises from sarcomatous differentiation of high-grade carcinoma while others have suggested a bi-clonal nature. Given these differences, we sought to separate the carcinoma and sarcoma elements of UCS to try to understand their molecular differences and gain further insights into how these tumors develop. We macrodissected carcinomatous, sarcomatous, and normal cells from formalin fixed paraffin embedded (FFPE) uterine samples of 10 UCS patients. DNA and RNA were isolated and extracted using the Qiagen AllPrep DNA/RNA FFPE kit. Whole-genome SNP microarrays and deep sequencing of 26 cancer genes was performed, using the Illumina Infinium OmniExpressExome array and the TruSight Tumor panel respectively. Illumina HiSeq mRNA sequencing was also performed to quantify gene expression. The genomic allelic imbalance (AI) profiling, called from the SNP data by hapLOH, showed that sarcoma samples were more aberrant than their carcinoma counterparts (abstract 131, AACR 2016). From the targeted sequencing, the Illumina Amplicon-DS Somatic Variant Caller was employed to call somatic mutations. Mutations were identified in TP53 in both the sarcoma and carcinoma samples of all 10 patients. Frequently mutated genes included APC, EGFR, MET and MSH6 which were found in 60-80% of the patients. Genes mutated in less than 50% of the patients included PTEN, KRAS, KIT, FBXW7, PIK3CA, FGFR2, and CTNNB1. Current results showed no association of a mutated gene to either the sarcoma or carcinoma component of UCS. RSEM, STAR and EBSeq were applied to the RNA-seq data for gene expression quantification. Approximately 2500 genes were identified as being differentially expressed (DE) between normal and carcinoma samples. Just over 4000 genes were identified as being DE between normal and sarcoma samples. 75% of the DE genes in the carcinoma were also identified in the sarcoma. Using DAVID functional annotation tool, we characterized these gene sets with KEGG pathways. Deregulated pathways identified in both carcinoma and sarcoma include: cell cycle, transcriptional regulation, Ras and p53 signaling. Some additional pathways are putatively associated with sarcoma only, including MAPK and PI3K-Akt signaling. We report here the differences between sarcoma and carcinoma components of UCS from multiple molecular perspectives. From the genomic AI and DE analysis, the carcinoma aberrations appear to be mostly a subset of the sarcoma tumor profiles, where sarcoma samples appear to be more highly aberrant compared to the carcinoma samples. One possible inference from this is that the sarcoma originated and evolved from the carcinoma cells. Citation Format: Yihua Liu, Zachary Weber, F. Anthony San Lucas, Aditya Deshpande, Raed Sulaiman, Mary Fagerness, Natasha Flier, Joseph Sulaiman, Christel M. Davis, Jerry Fowler, Gareth E. Davies, David Starks, Luis Rojas-Espaillat, Paul Scheet, Erik A. Ehli. Tumor profiling of separated carcinomatous and sarcomatous components from uterine carcinosarcoma biopsies provides insights into their development [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2017; 2017 Apr 1-5; Washington, DC. Philadelphia (PA): AACR; Cancer Res 2017;77(13 Suppl):Abstract nr 2463. doi:10.1158/1538-7445.AM2017-2463
Abstract Reproducing results is a major issue in cancer biology, whose “work bench” is dynamic and complex, with frequently updated algorithms and software. The better to manage our work in this environment we have developed SyQADA, a System for Quality-Assured Data Analysis – a workflow automation system designed to simplify common sequential analysis processes on the same or different data. SyQADA manages many of the details of procedural bookkeeping involved in bioinformatics workflows: What samples are we using? Where are the raw data? Were all the samples processed? Did every job complete satisfactorily? Is there as much output as expected? Where are the input files for the next step? How long does a typical job take to run? Which program versions did we use? Can we easily compare these results with the output of a different version of a program, or with different input data? Using SyQADA, we have found ourselves better able to reproduce results while at the same time reducing the human effort required to manage our upstream data analyses. Here, we briefly describe how our lung cancer studies have benefitted from the use of SyQADA. To understand the effect of different variant callers for Ion Torrent deep sequencing data in a lung cancer genomics study, we created a work protocol that allowed us to compare the different sets of variants called on 34 distinct somatic DNA samples from 4 patients. This complex processing framework involved running multiple variant callers, annotating variants, filtering germline variants using quality control metrics, and collating results across samples and callers. With SyQADA, we were able to re-run individual processes changing parameters with trivial changes to our configuration, yielding improved output. We then applied that unmodified protocol to the 500 samples from 48 individuals in our study, and rapidly produced data from which we could perform biological analysis. We then applied the protocol to a study of pre-malignant lesions in 25 lung cancer patients. In both studies, our workflow allowed us to generate comparable results in a matter of hours rather than days. SyQADA has been used by individuals with backgrounds ranging from expert programmer to Unix novice, to perform and repeat dozens of diverse analytical workflows. Projects to which SyQADA has been applied include allelic imbalance studies of TCGA samples for cancers of the breast, pancreas, lung, and colon, processing roughly 6000 samples through a dozen steps. A zipfile containing the SyQADA executable source code, documentation, tutorial examples, and workflows used in our lab will be available. Citation Format: Jerry Fowler, F. Anthony San Lucas, Smruthy Sivakumar, Aditya Deshpande, Humam Kadara, Paul A. Scheet. Optimizing the replication of cancer genomics workflows: case studies [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2017; 2017 Apr 1-5; Washington, DC. Philadelphia (PA): AACR; Cancer Res 2017;77(13 Suppl):Abstract nr 2594. doi:10.1158/1538-7445.AM2017-2594