High-dimensional, localized ribonucleic acid (RNA) sequencing is now possible owing to recent developments in spatial transcriptomics (ST). ST is based on highly multiplexed sequence analysis and uses barcodes to match the sequenced reads to their respective tissue locations. ST expression data suffer from high noise and dropout events; however, smoothing techniques have the promise to improve the data interpretability prior to performing downstream analyses. Single-cell RNA sequencing (scRNA-seq) data similarly suffer from these limitations, and smoothing methods developed for scRNA-seq can only utilize associations in transcriptome space (also known as one-factor smoothing methods). Since they do not account for spatial relationships, these one-factor smoothing methods cannot take full advantage of ST data. In this study, we present a novel two-factor smoothing technique, spatial and pattern combined smoothing (SPCS), that employs the k-nearest neighbor (kNN) technique to utilize information from transcriptome and spatial relationships. By performing SPCS on multiple ST slides from pancreatic ductal adenocarcinoma (PDAC), dorsolateral prefrontal cortex (DLPFC) and simulated high-grade serous ovarian cancer (HGSOC) datasets, smoothed ST slides have better separability, partition accuracy and biological interpretability than the ones smoothed by preexisting one-factor methods. Source code of SPCS is provided in Github (https://github.com/Usos/SPCS).
Motivation: Tumor-specific antigen (TSA) identification in human cancer predicts response to immunotherapy and provides targets for cancer vaccine and adoptive T-cell therapies with curative potential, and TSAs that are highly expressed at the RNA level are more likely to be presented on major histocompatibility complex (MHC)-I. Direct measurements of the RNA expression of peptides would allow for generalized prediction of TSAs. Human leukocyte antigen (HLA)-I genotypes were predicted with seq2HLA. RNA sequencing (RNAseq) fastq files were translated into all possible peptides of length 8-11, and peptides with high and low expressions in the tumor and control samples, respectively, were tested for their MHC-I binding potential with netMHCpan-4.0. Results: A novel pipeline for TSA prediction from RNAseq was used to predict all possible unique peptides size 8-11 on previously published murine and human lung and lymphoma tumors and validated on matched tumor and control lung adenocarcinoma (LUAD) samples. We show that neoantigens predicted by exomeSeq are typically poorly expressed at the RNA level, and a fraction is expressed in matched normal samples. TSAs presented in the proteomics data have higher RNA abundance and lower MHC-I binding percentile, and these attributes are used to discover high confidence TSAs within the validation cohort. Finally, a subset of these high confidence TSAs is expressed in a majority of LUAD tumors and represents attractive vaccine targets.
Abstract STK11 (liver kinase B1, LKB1) is the fourth most frequently mutated gene in lung adenocarcinoma, with loss of function observed in up to 30% of all cases. Our previous work identified a 16-gene signature for LKB1 loss of function through mutational and nonmutational mechanisms. In this study, we applied this genetic signature to The Cancer Genome Atlas (TCGA) lung adenocarcinoma samples and discovered a novel association between LKB1 loss and widespread DNA demethylation. LKB1-deficient tumors showed depletion of S-adenosyl-methionine (SAM-e), which is the primary substrate for DNMT1 activity. Lower methylation following LKB1 loss involved repetitive elements (RE) and altered RE transcription, as well as decreased sensitivity to azacytidine. Demethylated CpGs were enriched for FOXA family consensus binding sites, and nuclear expression, localization, and turnover of FOXA was dependent upon LKB1. Overall, these findings demonstrate that a large number of lung adenocarcinomas exhibit global hypomethylation driven by LKB1 loss, which has implications for both epigenetic therapy and immunotherapy in these cancers. Significance: Lung adenocarcinomas with LKB1 loss demonstrate global genomic hypomethylation associated with depletion of SAM-e, reduced expression of DNMT1, and increased transcription of repetitive elements.
Background: Tumor specific antigen (TSA) identification in human cancer predicts response to immunotherapy and provides vaccine targets for precision medicine. In addition to neoantigens from somatic coding mutations, numerous non-mutated TSAs can elicit T-cell responses but are often overlooked by current methods. We present a method that accurately and comprehensively predict TSAs from RNAseq data regardless of mutation status. Methods: HLA-I genotypes were predicted with seq2HLA. RNAseq fastq files were translated into all possible peptides of length 8-11, and peptides with high expression in the tumor and comparatively low expression in normal were tested for their MHC-I binding potential with netMHCpan-4.0. We defined our predicted TSA by i) high expression in tumor samples, ii) low expression in normal samples, and iii) high predicted patient-specific MHC-I binding affinity. Results: We developed a novel pipeline for TSA prediction from RNAseq that is not limited to mutation-derived TSAs. This pipeline was used to predict all possible unique peptides size 8-11 on previously published murine and human lung and lymphoma tumors then validated on matched tumor and control lung adenocarcinoma (LUAD) samples. This pipeline is able to predict TSAs in MHC-I ligand-purified proteomics data with favorable performance to existing methods. Furthermore, neoantigens predicted by exomeSeq are typically poorly expressed at the RNA level, (28% of predicted neoantigens with >0 expression, mean of 15.6 reads/sample) and a fraction of them (47/6,928, 0.68%) are expressed in matched normal samples. Finally, a set of 6 TSAs are expressed in 22/39 (56%) of LUAD tumors and represent attractive vaccine targets. Conclusion: Direct quantification of RNAseq evidence of the potential peptidome in matched tumor and control RNAseq samples, via our novel pipeline, allows for exhaustive detection of TSAs. Citation Format: Michael Sharpnack, Travis Johnson, Robert Chalkley, Zhi Han, David Carbone, Kun Huang, Kai He. Exhaustive tumor specific antigen detection with RNAseq [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2021; 2021 Apr 10-15 and May 17-21. Philadelphia (PA): AACR; Cancer Res 2021;81(13_Suppl):Abstract nr 238.
Integration of transcriptomic and proteomic data should reveal multi-layered regulatory processes governing cancer cell behaviors. Traditional correlation-based analyses have demonstrated limited ability to identify the post-transcriptional regulatory (PTR) processes that drive the non-linear relationship between transcript and protein abundances. In this work, we ideate an integrative approach to explore the variety of post-transcriptional mechanisms that dictate relationships between genes and corresponding proteins. The proposed workflow utilizes the intuitive technique of scatterplot diagnostics or scagnostics, to characterize and examine the diverse scatterplots built from transcript and protein abundances in a proteogenomic experiment. The workflow includes representing gene-protein relationships as scatterplots, clustering on geometric scagnostic features of these scatterplots, and finally identifying and grouping the potential gene-protein relationships according to their disposition to various PTR mechanisms. Our study verifies the efficacy of the implemented approach to excavate possible regulatory mechanisms by utilizing comprehensive tests on a synthetic dataset. We also propose a variety of 2D pattern-specific downstream analyses methodologies such as mixture modeling, and mapping miRNA post-transcriptional effects to explore each mechanism further. This work suggests that the proposed methodology has the potential for discovering and categorizing post-transcriptional regulatory mechanisms, manifesting in proteogenomic trends. These trends subsequently provide evidence for cancer specificity, miRNA targeting, and identification of regulation impacted by biological functionality and different types of degradation. (Supplementary Material - https://github.com/arunima2/PTRE_PSB_2020).
INTRODUCTION:Recent clinical studies have identified tumor mutation burden (TMB) as a promising therapeutic biomarker of anti-tumor immune checkpoint blockade. However, given the relatively slow turnaround time and high expense in measuring TMB, tobacco smoking history (TSH) is an attractive replacement biomarker. The carcinogenic effects of tobacco smoking may be modified by the protective effects of genome stability genes. This study aims to test the associations between tobacco smoking, genome stability gene inactivation, and TMB.METHODS:Publicly available TSH and DNA somatic alteration data from NSCLC were downloaded from The Cancer Genome Atlas. Correlations and enrichments were calculated with Spearman and Fisher's exact test methods, respectively. Multivariate modeling of TMB was performed with penalized linear regression.RESULTS:85% of never smokers in adenocarcinomas (LUAD) had low TMB, but a positive TSH was not predictive of hypermutancy. The limited utility of TSH in predicting TMB was reproduced on an independent LUAD dataset. To expand our search for predictors of TMB, we further investigated the contributions of genome stability related genes (GSGs) to TMB. 242/461 (52%) and 300/465 (65%) patients with LUAD and squamous carcinomas (LUSC), respectively, showed evidence of loss of function in at least one of the 182 GSGs. 182 GSGs from 16 pathways were assessed for associations with TMB high tumor status using Fisher's exact test. We performed univariate gene and pathway enrichments in TMB high tumors and found roles forPOLE, REV3L, and FANCE genes, as well as several key GSG pathways.CONCLUSIONS:This study comprehensively tested the association between GSG, tobacco smoking, and TMB in NSCLC. In LUAD, never-smoking status was predictive of low TMB, but overall TSH was not an adequate surrogate biomarker for TMB in NSCLC. Furthermore, we identified an association between GSG inactivation and TMB.
Recent clinical studies suggest tumor mutation burden (TMB) as a promising therapeutic biomarker of anti-tumor immune checkpoint blockade (ICB). Given the causal link between cancer-causing mutations and tobacco smoking, patients with a significant smoking history may respond better to ICB. However, it is not clear if smoking history is an adequate surrogate biomarker for TMB. Here, we sought assess the clinical utility of smoking history in predicting tumor mutation burden. Publicly available smoking history and DNA somatic alteration data from NSCLC were downloaded from The Cancer Genome Atlas and a large dataset of lung adenocarcinoma tumors published by Imielinski, et al (Cell, 2012) Tumor mutation burden was calculated as the sum of all somatic mutations divided by the exome sequencing coverage. Smoking history was analyzed both as categorical (ever, never, former) and semi-continuous variables (pack years). Hypermutancy was defined as greater than or equal to 10 mutations per megabase. A total of 395 LUAD and 419 LUSC patients were included in this analysis. Smokers had significantly higher tumor mutation burdens than non-smokers; however, in both LUAD and LUSC, there were smokers with low TMB and non-smokers with high TMB. Smoking pack year history (SPY) was weakly positively correlated (Spearman ρ = 0.20, p = 2.5x10-4) in LUAD but uncorrelated (Spearman ρ = -0.026, p = 0.61) in LUSC. Non-smokers and patients without a recorded SPY were excluded from the SPY analysis. We calculated AUCs for predicting hypermutancy in tumors, using variable thresholds of SPY. In LUAD and LUSC, SPY had an AUC of 0.38 and 0.47 in predicting TMB, showing that SPY was not better than random prediction. We also sought to predict TMB from smoking as a binary variable. In LUSC, 8/18 (44%) non-smokers and 253/447 (57%) smokers were hypermutant. In LUAD, 9/61 (15%) non-smokers and 219/391 (56%) smokers were hypermutant. Additionally, we repeated this analysis on matched smoking history and TMB from an independent cohort of 162 LUAD tumors published by Imielinski, et al. Similarly, we found that 1/27 (4%) of nonsmokers and 66/135 (49%) of smokers were hypermutant. In this cohort, the AUC in predicting TMB with SPY was 0·21. In this study, we investigated the relationship between tobacco smoking and TMB. While the average lung cancer patient with a history of tobacco smoking has a higher TMB than the average never-smoker, there is not a clear relationship between the extent of exposure in pack years and TMB. In general, smoking is not an informative biomarker for TMB, however, non-smokers who develop LUAD are unlikely to have high TMB. This study highlights the value of next generation sequencing for TMB in predicting therapeutic response to ICB.
12072 Background: Tumor somatic mutation burden (TMB) and neoantigen burden (NB), are emerging therapeutic biomarkers for immune checkpoint blockade in NSCLC; however, these biomarkers are costly and require extensive expertise to measure. Cheaper, simpler surrogate biomarkers are necessary to keep up with current translational research findings. Methods: We curated RNAseq, NB, and clinical data from The Cancer Genome Atlas (TCGA) for lung adenocarcinoma (LUAD, n = 461) tumors. Gene-level RNAseq data was correlated with TMB and NB, and ontological enrichment of genes associated with these quantities are discovered using the enrichR tool. In addition, RNA expression and TMB data from lung adenocarcinoma cell lines were curated from the Cancer Cell Line Encyclopedia. Neural network and penalized linear regression methods were used to create and test RNA expression signatures of somatic mutation and neoantigen burden. The code for these methods was implemented in R. Results: 5199 genes' RNA expression are significantly correlated with NB (Spearman correlation, Benjamin hochberg q-value < 0.01). Genes positively associated with NB are highly enriched in cell cycle related genes (pathway database, corrected p-value < 0.01). Further, we correlated RNA expression with mutation burden in lung adenocarcinoma cell lines and found a similar enrichment of cell cycle related genes. We created RNA expression signatures of NB using two separate methods and tested their performance on TCGA LUAD tumors using a random cross validation approach. Neural network and penalized linear regression cross-validation experiments had mean AUCs of 0.74 and 0.68, respectively. The final signatures were selected based on consistent and accurate performance on LUAD tumors. 50 genes were selected for the final signature, including YBX2, TDRKH, HDGF and others. Conclusions: Cell cycle-related RNA abundances are strongly associated with NB in LUAD. We show that these genes can be incorporated into a biomarker of NB. We are in the process of validating this RNA signature's ability to predict NB with a NanoString RNA panel on a separate LUAD cohort and investigating its implication in lung cancer immunotherapy.
Introduction: Despite apparently complete surgical resection, approximately half of resected early-stage lung cancer patients relapse and die of their disease. Adjuvant chemotherapy reduces this risk by only 5% to 8%. Thus, there is a need for better identifying who benefits from adjuvant therapy, the drivers of relapse, and novel targets in this setting. Methods: RNA sequencing and liquid chromatography/liquid chromatography-mass spectrometry proteomics data were generated from 51 surgically resected non-small cell lung tumors with known recurrence status. Results: We present a rationale and framework for the incorporation of high-content RNA and protein measurements into integrative biomarkers and show the potential of this approach for predicting risk of recurrence in a group of lung adenocarcinomas. In addition, we characterize the relationship between mRNA and protein measurements in lung adenocarcinoma and show that it is outcome specific. Conclusions: Our results suggest that mRNA and protein data possess independent biological and clinical importance, which can be leveraged to create higher-powered expression biomarkers. (C) 2018 International Association for the Study of Lung Cancer. Published by Elsevier Inc. All rights reserved.
Background: NSCLC remains a challenging disease to treat, especially because the majority of patients present with advanced disease [1].For many of these patients the standard first-line therapy is platinumbased chemotherapy, which can prolong survival by 8-12 months in some patients and also improve disease-related symptoms [2].The chemotherapy treatments are frequently poorly tolerated.Treatment choices following chemotherapy are limited, approved options include docetaxel, pemetrexed and erlotinib.Recent developments with immunotherapies and approval of agents such as nivolumab and pembrolizumab [3, 4] are exciting.Other immunotherapy agents such as darvolumab, atezolizumab and avelumab are also investigated and show good results.Materials and methods: Literature was reviewed about checkpoint inhibitors in metastatic NSCLC.Both the efficacy, patient reported outcomes and adverse events will be reported, with emphasis on the results from phase 3 trials.Results: Checkpoint inhibitors in development are ipilimumab and tremelimumab, they block CTLA4.Nivolumab and pembrolizumab are blocking binding of PD1 to PDL1 and PDL2.Atezolizumab and durvalumab are blocking binding of PDL1 to PD1 and CD 80. Avelumab is in early stages of research.Nivolumab trials CheckMate 017 [5] and 057 in second-line phase 3 trials and ChecMate 026 in 1st-line monotharapy vs SOC phase 3 trial were reported.In second-line nivolumab showed superiority over docetaxel in progression free survival (PFS) 3.5 vs 2.8 months, HR = 0.62, p-0.004, mOS = 9.2 vs 6 months independent to PDL1 expression in squamous histology.Adverse events were rare and manageable.Combination of nivolumab with chemotherapy and with erlotinib are being investigated.Pembrolizumab monotherapy showed responses which exceeded one year, median PFS = 6.3 months, especially high tumor proportion score (TPS) ≥50%, representing 23% of the screened NSCLC population.In pembrolizumab vs docetaxel, Keynote 010, median OS with pembrolizumab 2 mg/kg and at least 50% PDL1 TPS was 14.9 months vs 8.2 months on docetaxel.Adverse events were less common on pembrolizumab.PDL1 expression ≥50% correlated with improved RR,PFS and OS regardless of histology.Smoking status was associated with increased RR.Keynote 024 = pembrolizumab vs chemotherapy showed superiority in phase 3, chemonaive patients of pembrolizumab in OS. (=Press release only) Atezolizumab vs docetaxel, phase 2 randomized trial [6] POPLAR showed improvement in OS (HR 0.69 vs 0.73) in ITT population.Duration of response of 18.6 months compared to 7.2 months with docetaxel and well tolerated safety profile.OAK trial = phase 3 is ongoing.Durvalumab with trmelimumab phase 3 trial results are pending.Conclusion: The results of immunotherapy trials are encouraging, both from the point of improved efficacy and their tolerability.Combination treatments with immunotherapy may further improve the efficacy and prolong the lives of the patients with metastatic NSCLC.These combination treatments will replace eventually the platinum doublets in first-line treatment of metastatic NSCLC Immunotherapy in oncology: data from clinical trial K2 Immunotherapy in ovarian and endometrial cancer
Background: Biomarkers predictive of response to chemotherapy are critically needed for the precise selection of treatment protocols in lung adenocarcinoma (LUAD). Aquaporins (AQPs), transmembrane water channels, are emerging targets in cancer. Recently discovered, non-ubiquitous family member aquaporin 11 (AQP11), a tissue-specific endoplasmic reticulum (ER) resident, was identified as a cellular pro-survival factor implicated in the maintenance of ER homeostasis. AQP11 is mapped to 11q13-q14 amplicon harboring oncogenic drivers and associated with poor prognosis in cancer patients. We recently showed that high AQP11 expression is an in-vitro therapeutic biomarker of cisplatin therapy, which directly interferes with AQP11 functional structure. In addition, silencing of AQP11 expression in human lung cancer cell lines significantly increased response to cisplatin treatment. We hypothesized that LUAD tumors expressing high levels of AQP11 depend on AQP11-mediated cytoprotection and would be more responsive to platinum-based chemotherapy targeting AQP11. Method: We downloaded and curated matched mRNA expression, survival, and drug response data from The Cancer Genome Atlas (TCGA) LUAD dataset (N=369). Results: Analysis of PRECOG and TCGA databases showed that high AQP11 mRNA expression is negatively prognostic in patients with LUAD. TCGA LUAD cases were categorized by AQP11 mRNA levels into high and low expression (high = AQP11 mRNA expression mean + 1 standard deviation). Patients with high tumor AQP11 mRNA expression (10.3%) had significantly worse overall survival (OS) compared to patients with low AQP11 expression (p=0.0015). An analysis of patients treated with platinum-based chemotherapy (N=74) showed that patients with high AQP11 mRNA expression (7%) had significantly higher OS then patients with LUAD expressing low levels of AQP11 (p=0.0263). Conclusions: This study identifies AQP11 as a new biomarker of OS and chemotherapy in LUAD patients. It is conceivable that lung tumors expressing high levels of AQP11 are dependent on its function and elevated AQP11 expression renders tumor resistance to microenvironment and therapy-induced stress and associates with lower OS of patients. At the same time, as cisplatin efficiently targets AQP11 functional multimeric structure, these AQP11-depending tumors are more prone to platinum based chemotherapy. This study provides a rationale for combination anti-AQP11 and chemotherapy in LUAD tumors with high AQP11 expression. Citation Format: Michael Sharpnack, David P. Carbone, Mikhail M. Dikov, Elena E. Tchekneva. Aquaporin 11 as a new predictive biomarker of overall survival and platinum-based chemotherapy response in lung adenocarcinoma patients [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2018; 2018 Apr 14-18; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2018;78(13 Suppl):Abstract nr 2620.
MOTIVATIONTechnologies that generate high-throughput omics data are flourishing, creating enormous, publicly available repositories of multi-omics data. As many data repositories continue to grow, there is an urgent need for computational methods that can leverage these data to create comprehensive clusters of patients with a given disease.RESULTSOur proposed approach creates a patient-to-patient similarity graph for each data type as an intermediate representation of each omics data type and merges the graphs through subspace analysis on a Grassmann manifold. We hypothesize that this approach generates more informative clusters by preserving the complementary information from each level of omics data. We applied our approach to The Cancer Genome Atlas (TCGA) breast cancer dataset and show that by integrating gene expression, microRNA and DNA methylation data, our proposed method can produce clinically useful subtypes of breast cancer. We then investigate the molecular characteristics underlying these subtypes. We discover a highly expressed cluster of genes on chromosome 19p13 that strongly correlates with survival in TCGA breast cancer patients and validate these results in three additional breast cancer datasets. We also compare our approach with previous integrative clustering approaches and obtain comparable or superior results.AVAILABILITY AND IMPLEMENTATIONhttps://github.com/michaelsharpnack/GrassmannCluster.SUPPLEMENTARY INFORMATIONSupplementary data are available at Bioinformatics online.
Abstract Immunotherapy approaches targeting the PD-1 pathway have shown some clinical benefits in a fraction of patients with lung cancer, but expression of PD-L1 has proved to be an imperfect biomarker of efficacy. Recent studies have shown that tumor mutation burden (TMB) is also correlated with outcome and that it appears to be independent of PD-L1. TMB, however, only indirectly measures the number of neoantigenic peptides presented on tumor cell surface class I MHC, and predicted MHC matches may be an even better predictor of benefit. In addition, mutations in genes such as LKB1 may modulate the immune response, and mutations in the antigen presentation pathway may block it altogether. Mutations in DNA repair pathway genes may increase the number of potential neoantigens. Analysis of non-PD-1 pathway immunomodulators, immune cell infiltration, microenvironmental and microbiomic context, together with in-depth analysis of tumor somatic genomics, could lead to better patient selection for immunotherapy. Citation Format: David P. Carbone, Michael Sharpnack, Kai He. Immunotherapy: Biomarkers and checkpoint blockade in NSCLC [abstract]. In: Proceedings of the Fifth AACR-IASLC International Joint Conference: Lung Cancer Translational Science from the Bench to the Clinic; Jan 8-11, 2018; San Diego, CA. Philadelphia (PA): AACR; Clin Cancer Res 2018;24(17_Suppl):Abstract nr IA20.
Despite tremendous advances in targeted therapies against lung adenocarcinoma, the majority of patients do not benefit from personalized treatments. A deeper understanding of potential therapeutic targets is crucial to increase the survival of patients. One promising target, ADAR, is amplified in 13% of lung adenocarcinomas and in-vitro studies have demonstrated the potential of its therapeutic inhibition to inhibit tumor growth. ADAR edits millions of adenosines to inosines within the transcriptome, and while previous studies of ADAR in cancer have solely focused on protein-coding edits, >99% of edits occur in non-protein coding regions. Here, we develop a pipeline to discover the regulatory potential of RNA editing sites across the entire transcriptome and apply it to lung adenocarcinoma tumors from The Cancer Genome Atlas. This method predicts that 1413 genes contain regulatory edits, predominantly in non-coding regions. Genes with the largest numbers of regulatory edits are enriched in both apoptotic and innate immune pathways, providing a link between these known functions of ADAR and its role in cancer. We further show that despite a positive association between ADAR RNA expression and apoptotic and immune pathways, ADAR copy number is negatively associated with apoptosis and several immune cell types' signatures.
1545 Background: Comprehensive NGS panel based genetic testing is becoming more common to help clinicians provide personalized cancer care. Matched tumor-normal sequencing is recommended primarily to detect tumor-specific variants. Previously under-explored, it could also detect pathogenic germline alterations in cancer patients. Using targeted matched tumor-normal NGS, we identified and characterized germline variants in a large pan-cancer patient cohort in China. Methods: We surveyed the germline variants in 7363 Chinese patients across more than 18 diverse cancer types. Germline variants in 62 cancer-susceptibility genes were called from a 1021 gene NGS panel analyzing matched normal DNA. Following AMCG guidelines, variants were classified into pathogenic, likely pathogenic, variant of unknown significance, likely benign, or benign. Results: 385 germline pathogenic and likely pathogenic variants (GPVs) were identified in 374/7363 (5.1%) patients. Ovarian cancer (27.6%, 37/134) represented the highest prevalence. Breast cancer (11.3%, 92/813), colorectal cancer (8.3%, 66/791), pancreatic cancer (6.2%, 8/130), renal cell cancer (6%, 5/84), and gastric cancer (5.1%, 14/273) displayed relatively high rates of GPVs in line with expectations. Interestingly, NSCLC (2.5%, 88/3572) and hepatocellular cancer (2.3%, 5/214) also showed such events. In total, only 192/385 (49.9%) participants presented with GPVs in cancer-susceptibility genes in the expected cancer types. BRCA2 and BRCA1 were the top two common genes, which were found in 146 patients across 15 cancer types including 27/3572 (1%) NSCLC and 3/214 (1.4%) hepatocellular cancer patients. 285/385 (74%) of the GPVs were actionable for targeted therapy. Conclusions: Germline variants can be identified on routine targeted matched tumor-normal NGS and commonly exist in patients with cancers of diverse tissue origin. Recognition of germline variants may be valuable in therapeutic interventions and genetic risk analysis. We are currently performing retrospective family history analysis and genetic counseling for those patients with unreported GPVs.
A key hallmark of cancer cells is the ability to proliferate despite remarkable levels of DNA damage. Non-small cell lung cancers (NSCLC) tend to have high mutation rates, frequently related to smoking. While many genes have been functionally implicated in maintaining the integrity of the genome, for the majority of these genes there remains a lack of evidence of a direct relationship between loss-of-function and increased tumor mutation burden (TMB). Recent studies suggested an association between high TMB and cancer response to immunotherapies. The aim of this study was to comprehensively analyze the relationship between DNA integrity-related genes and TMB in NSCLC. Whole exome DNA sequencing and copy number array data were downloaded from TCGA lung adenocarcinoma (LAC) and squamous cell carcinoma (LSCC) datasets, and mutation burdens were calculated for each of 974 tumors. We identified 150 genes across 7 pathways and 9 groups known to be involved in repairing or compensating for DNA damage. To test each gene, tumors were placed into one of three groups according to the gene's mutation status; wild-type, homozygous deleted or mutated with loss-of-function, and non-synonymous missense mutated. We then compared the average mutation burden in each of these groups. This workflow was then repeated with pathways instead of genes. Our comprehensive analysis demonstrates a landscape of significant alterations to genes and pathways responsible for maintaining DNA integrity in NSCLC. A loss of function mutation or homozygous deletion in at least one signature gene occurred in 49% of LAC and 59% of LSCC. We searched for genes in this signature associated with significantly higher tumor mutation burdens (one sided t-test, p <0.05) and found 4 in LAC (RRM1 1%, TP53 17%, FANCE 1%, and MLH1 2%) and 8 in LSCC (NEIL1 0.5%, POLE 4%, POLG 0.5%, FANCE 3%, GEN1 1%, MLH1 4%, MSH6 1%, and RPA1 2%) datasets. Of note, tumors with nonsense mutations, indels, or homozygous deletions in the FANCE or MLH1 genes have significantly higher TMB in both LAC and LSCC. We repeat this process to find pathways significantly associated with increased TMB. We present a comprehensive study of the association between genes responsible for maintaining DNA integrity and TMB in NSCLC. These findings are important to the search for potential predictive biomarkers for immunotherapy.
The delivery of personalized healthcare is predicated on the application of the best available scientific knowledge to the practice of medicine in order to promote health, improve outcomes and enhance patient safety [1-3]. Unfortunately, current approaches to basic science research and clinical care are poorly integrated, yielding clinical decision-making processes that do not take advantage of up-to-date scientific knowledge [2-4]. Basic scientists investigating the biological basis for a given disease may regularly encounter synergistic effects spanning two or more bio-molecular entities or processes that can contribute to our understanding of the mechanisms underlying phenomena such as the etiologic basis of the targeted disease state or potential response to therapeutic agents [5]. However, systematic approaches to the use of that knowledge in order to directly inform the selection of targeted molecular therapies for “real world” patients are extremely limited [1, 3, 6-9]. There are an increasing number of multi-modelling and in-silico knowledge synthesis techniques that can provide investigators with the tools to quickly generate hypotheses concerning the relationships between entities found in heterogeneous collections of scientific data — for example, exploring potential linkages among genes, phenotypes and molecularly targeted therapeutic agents, thus enabling the “forward engineering” of treatment strategies based on knowledge generated via basic science studies [1, 4, 6, 10, 11]. Ultimately, the goal of such methodologies is to accelerate the identification of actionable research questions that can make direct contributions to clinical practice. Given increasing concerns over the barriers to the timely translation of discoveries from the laboratory to the clinic or broader population settings, such high-throughput hypothesis generation and testing is highly desirable [1, 4, 6, 8, 12]. These needs are particularly critical in numerous disease areas where the availability of new therapeutic agents is constrained, thus calling for the re-use and repositioning of existing treatments [13, 14]. In response to the challenges and opportunities enumerated above, there exits an emerging body of research and development focusing on multi-modeling approaches to the discovery of molecularly targeted therapies, including experimental paradigms spanning a spectrum from the identification of molecular targets for drugs, to the repurposing or repositioning of existing agents that utilize such targets, to the systematic identification of novel combination therapy regimens that amplify or enhance the effectiveness of their constituent components. This focus is motivated by recent and significant advances in the state of systems biology and medicine that have demonstrated that the ability to generate and reason across complex and scalar models is essential to the discovery of high-impact biologically and clinically actionable knowledge [1, 4, 12]. Such approaches are designed to overcome the limitations of reductionist approaches to scientific discovery, replacing decomposition-focused problem-solving with integrative network-based modeling and analysis techniques [4, 8]. Systems-level analysis of complex problem domains ultimately enables the study of critical interactions that influence health and wellness across a scale from molecules to populations, and are not observable when such systems are broken down into constituent components. The use of systems-level analysis methodologies is well supported by the foundational theory of vertical reasoning first proposed by Blois [15]. This theory holds that effective decision-making in the biomedical domain is predicated on the vertical integration of multiple, scalar levels of reasoning. This fundamental premise is the basis for a correlative framework put forth by Tsafnat and colleagues, which states that the ability to replicate expert reasoning relative to complex biomedical problems using computational agents (e.g., in-silico knowledge synthesis) requires the replication of such multi-scalar and integrative decision-making [16]. In order to achieve such an outcome, Tsafnat posits that multi-scalar decision-making in an in-silico context requires both: 1) the generation of component decision-making models at multiple scales; and 2) the similar generation of interchange layers that define important pair-wise connections between entities situated in two or more component models, often referred to as vertical linkages [16]. When such component models and interchange layers are combined in a computationally actionable format, they yield what can be referred to as a multi-model for a given domain that is able to satisfy the premises of Blois’ vertical reasoning axiom, and therefore facilitate the replication of expert performance in a high-throughput manner [16]. Of note, this type of approach is extremely reliant upon graph-theoretic reasoning and representational models, using a network paradigm that allows for the application of logical reasoning operations spanning the entities and relationships that make up a multi-model [8]. Network paradigms have been regularly shown to be the ideal representational model for naturally occurring systems, such as the ‘scale-free’ networks encountered in biological and clinical phenomena [8]. At the most basic level, network-based multi-modeling across scales presents an elegant and computationally tractable approach to understanding and evaluating complex biological and clinical systems in order to discover the knowledge incumbent to such constructs. This type of approach benefits from a robust set of foundational theories and frameworks that can inform and shape the application of multi-modeling techniques to a variety of knowledge discovery use cases. As such, there is a growing body of evidence concerning the application of network-based approaches to multi-modeling with an emphasis on therapeutic agent discovery, re-positioning and molecular targeting. Examples of such evidence include reports and perspectives published by Hood and Perlmutter [1], Butcher and colleagues [12], and Lussier and Chen [13].
Biological pathway regulation is complex, yet it underlies the functional coordination in a cell. Cancer is a disease that is characterized by unregulated growth, driven by underlying pathway deregulation. This pathway deregulation is both within pathways and between pathways. Here, we propose a method to detect inter-pathway coordination using distance correlation. Utilizing data generated from microarray experiments, we separate the genes into pathways and calculate the pairwise distance correlation between them. The result is intuitively viewed as a network of differentially dependent pathways. We find intuitive, yet surprising significant hub pathways, including glycophosphatidylinositol anchor synthesis in lung cancer.