The UK Biobank Exome Sequencing Consortium (UKB-ESC) is a private–public partnership between the UK Biobank (UKB) and eight biopharmaceutical companies that will complete the sequencing of exomes for all ~500,000 UKB participants. Here, we describe the early results from ~200,000 UKB participants and the features of this project that enabled its success. The biopharmaceutical industry has increasingly used human genetics to improve success in drug discovery. Recognizing the need for large-scale human genetics data, as well as the unique value of the data access and contribution terms of the UKB, the UKB-ESC was formed. As a result, exome data from 200,643 UKB enrollees are now available. These data include ~10 million exonic variants—a rich resource of rare coding variation that is particularly valuable for drug discovery. The UKB-ESC precompetitive collaboration has further strengthened academic and industry ties and has provided teams with an opportunity to interact with and learn from the wider research community. The UK Biobank Exome Sequencing Consortium aims to sequence all the exomes of approximately 500,000 UK Biobank participants. This Perspective describes the results from approximately 200,000 exomes and discusses the lessons learned from this UK Biobank–biopharmaceutical company collaboration.
Abstract Although next-generation sequencing is widely used in cancer to profile tumors and detect variants, most somatic variant callers used in these pipelines identify variants at the lowest possible granularity, single-nucleotide variants (SNV). As a result, multiple adjacent SNVs are called individually instead of as a multi-nucleotide variants (MNV). With this approach, the amino acid change from the individual SNV within a codon could be different from the amino acid change based on the MNV that results from combining SNV, leading to incorrect conclusions about the downstream effects of the variants. Here, we analyzed 10,383 variant call files (VCF) from the Cancer Genome Atlas (TCGA) and found 12,141 incorrectly annotated MNVs. Analysis of seven commonly mutated genes from 178 studies in cBioPortal revealed that MNVs were consistently missed in 20 of these studies, whereas they were correctly annotated in 15 more recent studies. At the BRAF V600 locus, the most common example of MNV, several public datasets reported separate BRAF V600E and BRAF V600M variants instead of a single merged V600K variant. VCFs from the TCGA Mutect2 caller were used to develop a solution to merge SNV to MNV. Our custom script used the phasing information from the SNV VCF and determined whether SNVs were at the same codon and needed to be merged into MNV before variant annotation. This study shows that institutions performing NGS sequencing for cancer genomics should incorporate the step of merging MNV as a best practice in their pipelines. Significance: Identification of incorrect mutation calls in TCGA, including clinically relevant BRAF V600 and KRAS G12, will influence research and potentially clinical decisions.
Although next-generation sequencing assays are routinely carried out using samples from cancer trials, the sequencing data are not always of the required quality. There is a need to evaluate the performance of tissue collection sites and provide feedback about the quality of next-generation sequencing data. This study used a modeling approach based on whole exome sequencing quality control (QC) metrics to evaluate the relative performance of sites participating in the Bristol Myers Squibb Immuno-Oncology clinical trials sample collection. We identified several events for the sample swap. Overall, most sites performed well and few showed poor performance. These findings can increase awareness of sample failure and improve the quality of samples.
Tumor mutational burden (TMB) has emerged as a clinically relevant biomarker that may be associated with immune checkpoint inhibitor efficacy. Standardization of TMB measurement is essential for implementing diagnostic tools to guide treatment. Here we describe the in-depth evaluation of bioinformatic TMB analysis by whole exome sequencing (WES) in formalin-fixed, paraffin-embedded samples from a phase III clinical trial. In the CheckMate 026 clinical trial, TMB was retrospectively assessed in 312 patients with non-small-cell lung cancer (58% of the intent-to-treat population) who received first-line nivolumab treatment or standard-of-care chemotherapy. We examined the sensitivity of TMB assessment to bioinformatic filtering methods and assessed concordance between TMB data derived by WES and the FoundationOne® CDx assay. TMB scores comprising synonymous, indel, frameshift, and nonsense mutations (all mutations) were 3.1-fold higher than data including missense mutations only, but values were highly correlated (Spearman’s r = 0.99). Scores from CheckMate 026 samples including missense mutations only were similar to those generated from data in The Cancer Genome Atlas, but those including all mutations were generally higher. Using databases for germline subtraction (instead of matched controls) showed a trend for race-dependent increases in TMB scores. WES and FoundationOne CDx outputs were highly correlated (Spearman’s r = 0.90). Parameter variation can impact TMB calculations, highlighting the need for standardization. Encouragingly, differences between assays could be accounted for by empirical calibration, suggesting that reliable TMB assessment across assays, platforms, and centers is achievable.
Background: Metastatic or surgically unresectable urothelial cancer (UC) is a disease with high unmet need. Nivolumab monotherapy demonstrated clinical benefit in patients with platinum-resistant UC (CheckMate 275; NCT02387996). Because not all patients respond to nivolumab therapy, identifying biomarkers for response is of critical importance. In the CheckMate 275 cohort, higher tumor mutational burden (TMB) was associated with longer overall survival (OS) (1). Previous studies have shown that certain mutational signatures, combinations of mutation types arising from specific mutagenesis processes, were associated with mutational burden and clinical outcome in UC (2,3). We hypothesized that specific mutational signatures may be associated with response to nivolumab in CheckMate 275. Methods: Whole exome sequencing data and clinical annotations for 139 archival, tumor-normal paired samples from patients enrolled in the CheckMate 275 cohort were collected and analyzed. Sequencing reads were processed and somatic variants called as described previously (4). Percentage of mutations across 30 mutational signatures were generated using deconstructSigs, which infers signature activity with given known signatures (5). Association of the most prevalent mutational signatures with TMB, previously known clinical biomarkers, and clinical efficacy (OS, progression free survival [PFS], and objective response [OR]) were examined. Results: We identified age-related (signature 1), endogenous mutagenesis-related APOBEC (signatures 2 and 13) and UV-induced (signature 7) as the most prevalent mutational signatures in UC. When examining their correlations with TMB, signature 1 showed negative correlation, signatures 2 and 13 showed high positive correlation, and signature 7 showed low correlation. Higher signature 2 scores were associated with better OS, but were poor predictors for PFS and OR. Overall, while signature 2 added predictive value over previously known clinical biomarkers such as PD-L1 expression, serum hemoglobin level, and presence of liver metastasis, it did not improve on the predictive performance of TMB. Conclusions: An APOBEC-related mutational signature (signature 2) was associated with OS in patients with advanced UC treated with nivolumab. This signature was correlated with TMB and did not improve upon TMB as a predictive biomarker. Our work demonstrates the importance of including TMB as a covariate when analyzing mutational signatures. References: 1. Galsky MD et al. Ann Oncol 2017;28(suppl 5). Abstract 848PD 2. Glaser AP et al. Oncotarget 2018;9:4537-48 3. Roberts SA et al. Nat Genet 2013;45:970-6 4. Carbone DP et al. NEJM 2017;376:2415-26 5. Forbes SA et al. Nucleic Acids Res 2017;45:D777-83 Citation Format: G. Celine Han, Peter M. Szabo, Abdel Saci, Alice M. Walsh, Natallia Kalinava, Ariella Sasson, Bruce Fischer, Matthew D. Galsky, Padmanee Sharma. Mutational signatures as biomarkers of response to nivolumab in metastatic bladder cancer [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2019; 2019 Mar 29-Apr 3; Atlanta, GA. Philadelphia (PA): AACR; Cancer Res 2019;79(13 Suppl):Abstract nr 4890.
Durable responses and encouraging survival have been demonstrated with immune checkpoint inhibitors in small-cell lung cancer (SCLC), but predictive markers are unknown. We used whole exome sequencing to evaluate the impact of tumor mutational burden on efficacy of nivolumab monotherapy or combined with ipilimumab in patients with SCLC from the nonrandomized or randomized cohorts of CheckMate 032. Patients received nivolumab (3 mg/kg every 2 weeks) or nivolumab plus ipilimumab (1 mg/kg plus 3 mg/kg every 3 weeks for four cycles, followed by nivolumab 3 mg/kg every 2 weeks). Efficacy of nivolumab ± ipilimumab was enhanced in patients with high tumor mutational burden. Nivolumab plus ipilimumab appeared to provide a greater clinical benefit than nivolumab monotherapy in the high tumor mutational burden tertile.
The scientific value of re-analyzing existing datasets is often proportional to the complexity of the data. Proteomics data are inherently complex and can be analyzed at many levels, including proteins, peptides, and post-translational modifications to verify and/or develop new hypotheses. In this paper, we present our re-analysis of a previously published study comparing colon biopsy samples from ulcerative colitis (UC) patients to non-affected controls. We used a different statistical approach, employing a linear mixed-effects regression model and analyzed the data both on the protein and peptide level. In addition to confirming and reinforcing the original finding of upregulation of neutrophil extracellular traps (NETs), we report novel findings, including that Extracellular Matrix (ECM) degradation and neutrophil maturation are involved in the pathology of UC. The pharmaceutically most relevant differential protein expressions were confirmed using immunohistochemistry as an orthogonal method. As part of this study, we also compared proteomics data to previously published mRNA expression data. These comparisons indicated compensatory regulation at transcription levels of the ECM proteins we identified and open possible new avenues for drug discovery.
Copy-number variation (CNV) plays an important role in cancer development since activation of an oncogene may be linked to a genomic copy-number amplification and inactivation of a tumor suppressor gene can be caused by a deletion. Benchmarking of CNV callers is hampered by the lack of appropriate truth sets, necessitating the use of simulated data, especially in studies focusing on somatic CNVs. In this study, we have (i) tested multiple commonly used CNV data simulators, (ii) used them to create several synthetic CNV datasets, and (iii) benchmarked seven CNV callers using both synthetic and real data. Seven CNV callers (CNVkit, Control-FREEC, Sequenza, FACETS, PureCN, CNVnator and GATK4 CNV) were benchmarked under three scenarios which included both whole genome sequencing (WGS) and whole exome sequencing (WES) data: matched tumor-normal sample pairs, tumor sample against a panel of normal samples and tumor-only sample CNV calling. For each dataset two complementary evaluation methods were applied, resulting in two distinct, but positively correlated scores. Analysis of concordance between calls made by top performing tools shows potential for consensus calling which would generate high-confidence calls, as well as mark regions where different callers generate contradictory calls as high complexity regions.
Purpose Hearing loss (HL) is the most common sensory disorder in children. Prompt molecular diagnosis may guide screening and management, especially in syndromic cases when HL is the single presenting feature. Exome sequencing (ES) is an appealing diagnostic tool for HL as the genetic causes are highly heterogeneous. Methods ES was performed on a prospective cohort of 43 probands with HL. Sequence data were analyzed for primary and secondary findings. Capture and coverage analysis was performed for genes and variants associated with HL. Results The diagnostic rate using ES was 37.2%, compared with 15.8% for the clinical HL panel. Secondary findings were discovered in three patients. For 247 genes associated with HL, 94.7% of the exons were targeted for capture and 81.7% of these exons were covered at 20× or greater. Further analysis of 454 randomly selected HL-associated variants showed that 89% were targeted for capture and 75% were covered at a read depth of at least 20×. Conclusion ES has an improved yield compared with clinical testing and may capture diagnoses not initially considered due to subtle clinical phenotypes. Technical challenges were identified, including inadequate capture and coverage of HL genes. Additional considerations of ES include secondary findings, cost, and turnaround time.
KRAS is the most common oncogenic driver in lung adenocarcinoma (LUAC). We previously reported that STK11/LKB1 (KL) or TP53 (KP) comutations define distinct subgroups of KRAS-mutant LUAC. Here, we examine the efficacy of PD-1 inhibitors in these subgroups. Objective response rates to PD-1 blockade differed significantly among KL (7.4%), KP (35.7%), and K-only (28.6%) subgroups (P < 0.001) in the Stand Up To Cancer (SU2C) cohort (174 patients) with KRAS-mutant LUAC and in patients treated with nivolumab in the CheckMate-057 phase III trial (0% vs. 57.1% vs. 18.2%; P = 0.047). In the SU2C cohort, KL LUAC exhibited shorter progression-free (P < 0.001) and overall (P = 0.0015) survival compared with KRASMUT;STK11/LKB1WT LUAC. Among 924 LUACs, STK11/LKB1 alterations were the only marker significantly associated with PD-L1 negativity in TMBIntermediate/High LUAC. The impact of STK11/LKB1 alterations on clinical outcomes with PD-1/PD-L1 inhibitors extended to PD-L1-positive non-small cell lung cancer. In Kras-mutant murine LUAC models, Stk11/Lkb1 loss promoted PD-1/PD-L1 inhibitor resistance, suggesting a causal role. Our results identify STK11/LKB1 alterations as a major driver of primary resistance to PD-1 blockade in KRAS-mutant LUAC.Significance: This work identifies STK11/LKB1 alterations as the most prevalent genomic driver of primary resistance to PD-1 axis inhibitors in KRAS-mutant lung adenocarcinoma. Genomic profiling may enhance the predictive utility of PD-L1 expression and tumor mutation burden and facilitate establishment of personalized combination immunotherapy approaches for genomically defined LUAC subsets. Cancer Discov; 8(7); 822-35. ©2018 AACR.See related commentary by Etxeberria et al., p. 794This article is highlighted in the In This Issue feature, p. 781.
TMB has emerged as a predictive biomarker of response to immune checkpoint inhibitors. CheckMate 227 demonstrated that patients with non-small cell lung cancer (NSCLC) with TMB ≥10 mutations/megabase derived enhanced benefit from first-line treatment with nivolumab + ipilimumab vs chemotherapy (Hellmann et al. NEJM 2018). Standardized approaches for the measurement and reporting of TMB are essential for the real-world implementation of TMB. This study aimed to refine a bioinformatic pipeline for mutation calling and annotation of whole exome sequencing (WES) data for TMB assessment. In CheckMate 026, TMB was assessed by WES on formalin-fixed, paraffin-embedded tumor samples and matched blood from 312 patients with NSCLC (Carbone et al. NEJM 2017). Data from each sample were aligned to a reference human genome and somatic mutations were called by comparing matched tumor and blood samples using the TNsnv and Strelka algorithms. The somatic mutations were additionally filtered for germline variants in public databases. TMB scores were compared with data from 710 NSCLC samples in The Cancer Genome Atlas (TCGA) dataset. We examined the concordance of TMB estimates using several mutation filtering schemes, with and without matched germline controls. TMB scores including synonymous, indel, frameshift, and nonsense mutations (all mutations) were ∼3-fold higher than matched data filtered for missense mutations only, but values were highly correlated (Spearman's r=0.99). Scores including missense mutations only were similar to those generated from TCGA, but those including all mutations were on average higher. Using public databases for germline subtraction showed a trend for race-dependent increases in TMB scores. Standardization of bioinformatic analyses is critical to the clinical implementation of TMB assessment. TMB assessment is sensitive to variations in bioinformatic parameters (eg, which type of mutation to include), which may affect the identification of patients likely to respond to immunotherapy. These results show that data from different pipelines are highly correlated, suggesting that reliable assessment of TMB across different centers and platforms is achievable. Chang H et al., ESMO 2018, Annals of Oncology, Volume 29, 2018 Supplement 6.
Abstract Background: Integrated analysis of multi-omic data is important in understanding oncogenesis, pathology, and drug mechanism of action. Formalin-fixed, paraffin-embedded (FFPE) tissues comprise the bulk of archival specimens in hospitals and clinical trials. Commercial kits are commonly used to extract RNA, DNA, or protein for biomarker studies; however, high-quality extractions are challenging due to crosslinking. We explored the feasibility of performing proteomic, transcriptomic, and genetic analysis from limited FFPE samples collected in clinical trials through a pilot study in colorectal cancer (CRC) and normal tissue. Methods: Qiagen kits were employed, with some modifications, to extract nucleic acid and protein from FFPE sample lysates that would be amenable to next-generation sequencing and liquid chromatography/mass-spectrometry (LC/MS) proteomic profiling. Commercially sourced FFPE slides from CRC biopsies and normal colon, heart, and skin samples were processed (n = 10 each; ~600-mm2 by 5-μm/slice). RNA-seq library preparation was performed using the Illumina® TruSeq Access kit. Label-free LC/MS proteomic analysis was carried out using trypsin/Lys-C-digested protein fragments on an EASY-nLC™ 1200 coupled to a Q Exactive™ Plus (Thermo Fisher). MS data were processed with MaxQuant to estimate protein abundance (Cox and Mann, Nat Biotech 2008). Results: Of 40 samples tested, 38 yielded quantifiable nucleic acid and protein of sufficient quality for transcriptional and proteomic analyses. Mean RNA yield was ~300 ng (150 ng-1.2 µg). RNA degradation was significant, with a mean DV>200 of 45% (Agilent Bioanalyzer). Mean DNA yield was 530 ng, with a mean fragment size of 2 kb. Mean protein yield was 50 µg (1 µg-300 µg). RNA libraries had a mapping rate range of 85%-90% and a coding rate range of 70%-80%; ~16,000 genes were reliable detected (log2 TPM > 1). The number of unique proteins detected ranged from 3,000 to 4,000 for CRC, normal colon, and heart samples; skin samples averaged 1,000 proteins and had the lowest overall yield. Correlation of RNA and protein expression for the same genes was weak (Spearman's rho ~0.3-0.4) at the per-sample level but statistically significant. Using linear models, we identified differentially expressed transcripts and proteins between the CRC and normal colon samples. Pathway enrichment analysis of both modalities indicated changes in cell cycle, which is consistent with rapid growth of tumor cells. Changes in nucleoside metabolism, extracellular matrix remodeling, and innate immune response were more apparent at the protein level. Conclusion: We found that multi-omic analysis was feasible with FFPE samples, and proteomics can be used to validate RNA results. Additionally, proteomics reveal post-translational events, such as extracellular matrix remodeling, that provide unique insights into cancer pathology. Citation Format: Vishal Patel, Ji Gao, Mingyi Liu, Aiqing He, Xi-Tao Wang, Kathryn Vanderlaag, Ariella Sasson, Stefan Kirov, Sunil Kuppasani, Omar J. Jabado, Kandasamy Ravi, Ashok Dongre, Julie Carman, Heidi LeBlanc. Integrated analysis of colorectal carcinoma by co-extraction of RNA, DNA and protein from FFPE tumor samples [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2018; 2018 Apr 14-18; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2018;78(13 Suppl):Abstract nr 2707.
Purpose As massively parallel sequencing is increasingly being used for clinical decision making, it has become critical to understand parameters that affect sequencing quality and to establish methods for measuring and reporting clinical sequencing standards. In this report, we propose a definition for reduced coverage regions and describe a set of standards for variant calling in clinical sequencing applications. Methods To enable sequencing centers to assess the regions of poor sequencing quality in their own data, we optimized and used a tool (ExCID) to identify reduced coverage loci within genes or regions of particular interest. We used this framework to examine sequencing data from 500 patients generated in 10 projects at sequencing centers in the National Human Genome Research Institute/National Cancer Institute Clinical Sequencing Exploratory Research Consortium. Results This approach identified reduced coverage regions in clinically relevant genes, including known clinically relevant loci that were uniquely missed at individual centers, in multiple centers, and in all centers. Conclusion This report provides a process road map for clinical sequencing centers looking to perform similar analyses on their data.
Background The ability to capture and sequence large contiguous DNA fragments represents a significant advancement towards the comprehensive characterization of complex genomic regions. While emerging sequencing platforms are capable of producing several kilobases-long reads, the fragment sizes generated by current DNA target enrichment technologies remain a limiting factor, producing DNA fragments generally shorter than 1 kbp. The DNA enrichment methodology described herein, Region-Specific Extraction (RSE), produces DNA segments in excess of 20 kbp in length. Coupling this enrichment method to appropriate sequencing platforms will significantly enhance the ability to generate complete and accurate sequence characterization of any genomic region without the need for reference-based assembly. Results RSE is a long-range DNA target capture methodology that relies on the specific hybridization of short (20-25 base) oligonucleotide primers to selected sequence motifs within the DNA target region. These capture primers are then enzymatically extended on the 3’-end, incorporating biotinylated nucleotides into the DNA. Streptavidin-coated beads are subsequently used to pull-down the original, long DNA template molecules via the newly synthesized, biotinylated DNA that is bound to them. We demonstrate the accuracy, simplicity and utility of the RSE method by capturing and sequencing a 4 Mbp stretch of the major histocompatibility complex (MHC). Our results show an average depth of coverage of 164X for the entire MHC. This depth of coverage contributes significantly to a 99.94 % total coverage of the targeted region and to an accuracy that is over 99.99 %. Conclusions RSE represents a cost-effective target enrichment method capable of producing sequencing templates in excess of 20 kbp in length. The utility of our method has been proven to generate superior coverage across the MHC as compared to other commercially available methodologies, with the added advantage of producing longer sequencing templates amenable to DNA sequencing on recently developed platforms. Although our demonstration of the method does not utilize these DNA sequencing platforms directly, our results indicate that the capture of long DNA fragments produce superior coverage of the targeted region.