Introduction:Cancers presenting at advanced stages inherently have poor prognosis. High grade serous carcinoma (HGSC) is the most common and aggressive form of tubo-ovarian cancer. Clinical tests to accurately diagnose and monitor this condition are lacking. Hence, development of disease-specific tests are urgently required.Methods:The molecular profile of HGSC during disease progression was investigated in a unique patient cohort. A bespoke data browser was developed to analyse gene expression and DNA methylation datasets for biomarker discovery. The Ovarian Cancer Data Browser (OCDB) is built in C# with a.NET framework using an integrated development environment of Microsoft Visual Studio and fast access files (.faf). The graphical user interface is easy to navigate between four analytical modes (gene expression; methylation; combined gene expression and methylation data; methylation clusters), with a rapid query response time. A user should first define a disease progression trend for prioritising results. Single or multiomics data are then mined to identify probes, genes and methylation clusters that exhibit the desired trend. A unique scoring system based on the percentage change in expression/methylation between disease stages is used. Results are filtered and ranked using weighting and penalties.Results:The OCDB's utility for biomarker discovery is demonstrated with the identified target OSR2. Trends in OSR2 repression and hypermethylation with HGSC disease progression were confirmed in the browser samples and an independent cohort using bioassays. The OSR2 methylation biomarker could discriminate HGSC with high specificity (95%) and sensitivity (93.18%).Conclusions:The OCDB has been refined and validated to be an integral part of a unique biomarker discovery pipeline. It may also be used independently to aid identification of novel targets. It carries the potential to identify further biomarker assays that can reduce type I and II errors within clinical diagnostics.
Adult brain tumors (glioma) represent a cancer of unmet need where standard-of-care is non-curative; thus, new therapies are urgently needed. It is unclear whether isocitrate dehydrogenases (IDH1/2) when not mutated have any role in gliomagenesis or tumor growth. Nevertheless, IDH1 is overexpressed in glioblastoma (GBM), which could impact upon cellular metabolism and epigenetic reprogramming. This study characterizes IDH1 expression and associated genes and pathways. A novel biomarker discovery pipeline using artificial intelligence (evolutionary algorithms) was employed to analyze IDH-wildtype adult gliomas from the TCGA LGG-GBM cohort. Ninety genes whose expression correlated with IDH1 expression were identified from: (1) All gliomas, (2) primary GBM, and (3) recurrent GBM tumors. Genes were overrepresented in ubiquitin-mediated proteolysis, focal adhesion, mTOR signaling, and pyruvate metabolism pathways. Other non-enriched pathways included O-glycan biosynthesis, notch signaling, and signaling regulating stem cell pluripotency (PCGF3). Potential prognostic (TSPYL2, JAKMIP1, CIT, TMTC1) and two diagnostic (MINK1, PLEKHM3) biomarkers were downregulated in GBM. Their gene expression and methylation were negatively and positively correlated with IDH1 expression, respectively. Two diagnostic biomarkers (BZW1, RCF2) showed the opposite trend. Prognostic genes were not impacted by high frequencies of molecular alterations and only one (TMTC1) could be validated in another cohort. Genes with mechanistic links to IDH1 were involved in brain neuronal development, cell proliferation, cytokinesis, and O-mannosylation as well as tumor suppression and anaplerosis. Results highlight metabolic vulnerabilities and therapeutic targets for use in future clinical trials.
The development of gene signatures is key for delivering personalized medicine, despite only a few signatures being available for use in the clinic for cancer patients. Gene signature discovery tends to revolve around identifying a single signature. However, it has been shown that various highly predictive signatures can be produced from the same dataset. This study assumes that the presentation of top ranked signatures will allow greater efforts in the selection of gene signatures for validation on external datasets and for their clinical translation. Particle swarm optimization (PSO) is an evolutionary algorithm often used as a search strategy and largely represented as binary PSO (BPSO) in this domain. BPSO, however, fails to produce succinct feature sets for complex optimization problems, thus affecting its overall runtime and optimization performance. Enhanced BPSO (EBPSO) was developed to overcome these shortcomings. Thus, this study will validate unique candidate gene signatures for different underlying biology from EBPSO on transcriptomics cohorts. EBPSO was consistently seen to be as accurate as BPSO with substantially smaller feature signatures and significantly faster runtimes. 100% accuracy was achieved in all but two of the selected data sets. Using clinical transcriptomics cohorts, EBPSO has demonstrated the ability to identify accurate, succinct, and significantly prognostic signatures that are unique from one another. This has been proposed as a promising alternative to overcome the issues regarding traditional single gene signature generation. Interpretation of key genes within the signatures provided biological insights into the associated functions that were well correlated to their cancer type.
Abstract Identifying robust predictive biomarkers to stratify colorectal cancer (CRC) patients based on their response to immune-checkpoint therapy is an area of unmet clinical need. Our evolutionary algorithm Atlas Correlation Explorer (ACE) represents a novel approach for mining The Cancer Genome Atlas (TCGA) data for clinically relevant associations. We deployed ACE to identify candidate predictive biomarkers of response to immune-checkpoint therapy in CRC. We interrogated the colon adenocarcinoma (COAD) gene expression data across nine immune-checkpoints (PDL1, PDCD1, CTLA4, LAG3, TIM3, TIGIT, ICOS, IDO1 and BTLA). IL2RB was identified as the most common gene associated with immune-checkpoint genes in CRC. Using human/murine single-cell RNA-seq data, we demonstrated that IL2RB was expressed predominantly in a subset of T-cells associated with increased immune-checkpoint expression (P < 0.0001). Confirmatory IL2RB immunohistochemistry (IHC) analysis in a large MSI-H colon cancer tissue microarray (TMA; n = 115) revealed sensitive, specific staining of a subset of lymphocytes and a strong association with FOXP3+ lymphocytes (P < 0.0001). IL2RB mRNA positively correlated with three previously-published gene signatures of response to immune-checkpoint therapy (P < 0.0001). Our evolutionary algorithm has identified IL2RB to be extensively linked to immune-checkpoints in CRC; its expression should be investigated for clinical utility as a potential predictive biomarker for CRC patients receiving immune-checkpoint blockade.
Abstract Modern methods of acquiring molecular data have improved rapidly in recent years, making it easier for researchers to collect large volumes of information. However, this has increased the challenge of recognizing interesting patterns within the data. Atlas Correlation Explorer (ACE) is a user-friendly workbench for seeking associations between attributes in The Cancer Genome Atlas (TCGA) database. It allows any combination of clinical and genomic data streams to be searched using an evolutionary algorithm approach. To showcase ACE, we assessed which RNA sequencing transcripts were associated with estrogen receptor (ESR1) in the TCGA breast cancer cohort. The analysis revealed already well-established associations with XBP1 and FOXA1, but also identified a strong association with CT62, a potential immunotherapeutic target with few previous associations with breast cancer. In conclusion, ACE can produce results for very large searches in a short time and will serve as an increasingly useful tool for biomarker discovery in the big data era. Significance: ACE uses an evolutionary algorithm approach to perform large searches for associations between any combinations of data in the TCGA database.
Abstract Longitudinal next-generation sequencing of cancer patient samples has enhanced our understanding of the evolution and progression of various cancers. As a result, and due to our increasing knowledge of heterogeneity, such sampling is becoming increasingly common in research and clinical trial sample collections. Traditionally, the evolutionary analysis of these cohorts involves the use of an aligner followed by subsequent stringent downstream analyses. However, this can lead to large levels of information loss due to the vast mutational landscape that characterizes tumor samples. Here, we propose an alignment-free approach for sequence comparison—a well-established approach in a range of biological applications including typical phylogenetic classification. Such methods could be used to compare information collated in raw sequence files to allow an unsupervised assessment of the evolutionary trajectory of patient genomic profiles. In order to highlight this utility in cancer research we have applied our alignment-free approach using a previously established metric, Jensen–Shannon divergence, and a metric novel to this area, Hellinger distance, to two longitudinal cancer patient cohorts in glioma and clear cell renal cell carcinoma using our software, NUQA. We hypothesize that this approach has the potential to reveal novel information about the heterogeneity and evolutionary trajectory of spatiotemporal tumor samples, potentially revealing early events in tumorigenesis and the origins of metastases and recurrences. Key words: alignment-free, Hellinger distance, exome-seq, evolution, phylogenetics, longitudinal.
OBJECTIVE:High grade serous carcinoma (HGSC) is the most common and most aggressive, subtype of epithelial ovarian cancer. It presents as advanced stage disease with poor prognosis. Recent pathological evidence strongly suggests HGSC arises from the fallopian tube via the precursor lesion; serous tubal intraepithelial carcinoma (STIC). However, further definition of the molecular evolution of HGSC has major implications for both clinical management and research. This study aims to more clearly define the molecular pathogenesis of HGSC.METHODS:Six cases of HGSC were identified at the Northern Ireland Gynaecological Cancer Centre (NIGCC) that each contained ovarian HGSC (HGSC), omental HGSC (OMT), STIC, normal fallopian tube epithelium (FTE) and normal ovarian surface epithelium (OSE). The relevant formalin-fixed paraffin embedded (FFPE) tissue samples were retrieved from the pathology archive via the Northern Ireland Biobank following attaining ethical approval (NIB11:005). Full microarray-based gene expression profiling was performed on the cohort. The resulting data was analysed bioinformatically and the results were validated in a HGSC-specific in-vitro model.RESULTS:The carcinogenesis of HGSC was investigated and showed the molecular profile of HGSC to be more closely related to normal FTE than OSE. STIC lesions also clustered closely with HGSC, indicating a common molecular origin.CONCLUSION:This study provides strong evidence suggesting that extrauterine HGSC arises from the fimbria of the distal fallopian tube. Furthermore, several potential pathways were identified which could be targeted by novel therapies for HGSC. These findings have significant translational relevance for both primary prevention and clinical management of the disease.
Identifying robust predictive biomarkers to enable stratification of colorectal cancer (CRC) patients, based on their response to immune checkpoint therapy, is an area of unmet clinical need. Genetic algorithms represent an exciting branch of artificial intelligence which can be used to extract meaningful associations from ‘big data’ now emerging more frequently in oncological research. We have employed Atlas Correlation Explorer (ACE), a user-friendly workbench that utilises a genetic algorithm to mine data deposited in The Cancer Genome Atlas (TCGA). Our aim was to establish common intersections between gene expression analyses in ACE using nine well established immune checkpoint markers (CD274, PDCD1, CTLA4, LAG3, TIM3, TIGIT, ICOS, IDO1 and BTLA). We observed IL2RB to be the common gene associated with immune checkpoints in both microarray and RNA sequencing data from the TCGA (7/9 gene lists). Assessment of IL2RB indicates that it is highly expressed on CD56+ natural killer cells and is associated with an increased infiltration of cytotoxic lymphocytes and a decreased infiltration of fibroblasts. It is also significantly enriched in the immune consensus molecular subtype group CMS1. We next demonstrated that patients with high IL2RB gene expression have better relapse free survival in the TCGA CRC cohort (n = 322, log-rank p = 0.011) and an all stage CRC validation cohort GSE39582 (n = 519, log-rank p = 0.006). It is also an independent prognostic factor by multivariate analysis (p = 0.01). We next observed strong correlations between IL2RB gene expression in CRC and previously published predictive gene signatures for anti-PD1 therapies in other solid tumours (Pearson correlation, R = 0.88). Finally, we optimised assessment of IL2RB immunohistochemistry in a large CRC cohort (n=661) using a digital pathology approach with the open-source QuPath software. To conclude, we have validated IL2RB as prognostic biomarker and have provided evidence to demonstrate that IL2RB expression could be used for CRC patient stratification in future immunotherapy based clinical trials. Citation Format: Matthew Alderdice, Stephanie Craig, Matt Humphries, Alan Gilmore, Victoria Bingham, Nicole Johnston, Stephen McQuaid, Manuel Salto-Tellez, Mark Lawler, Darragh G. McArt. Artificial intelligence approach identifies IL2RB as a common prognostic and potential predictive biomarker associated with immune checkpoints in colorectal cancer [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2019; 2019 Mar 29-Apr 3; Atlanta, GA. Philadelphia (PA): AACR; Cancer Res 2019;79(13 Suppl):Abstract nr 2787.
Purpose Gene expression profiling can uncover biologic mechanisms underlying disease and is important in drug development. RNA sequencing (RNA-seq) is routinely used to assess gene expression, but costs remain high. Sample multiplexing reduces RNA-seq costs; however, multiplexed samples have lower cDNA sequencing depth, which can hinder accurate differential gene expression detection. The impact of sequencing depth alteration on RNA-seq–based downstream analyses such as gene expression connectivity mapping is not known, where this method is used to identify potential therapeutic compounds for repurposing. Methods In this study, published RNA-seq profiles from patients with brain tumor (glioma) were assembled into two disease progression gene signature contrasts for astrocytoma. Available treatments for glioma have limited effectiveness, rendering this a disease of poor clinical outcome. Gene signatures were subsampled to simulate sequencing alterations and analyzed in connectivity mapping to investigate target compound robustness. Results Data loss to gene signatures led to the loss, gain, and consistent identification of significant connections. The most accurate gene signature contrast with consistent patient gene expression profiles was more resilient to data loss and identified robust target compounds. Target compounds lost included candidate compounds of potential clinical utility in glioma (eg, suramin, dasatinib). Lost connections may have been linked to low-abundance genes in the gene signature that closely characterized the disease phenotype. Consistently identified connections may have been related to highly expressed abundant genes that were ever-present in gene signatures, despite data reductions. Potential noise surrounding findings included false-positive connections that were gained as a result of gene signature modification with data loss. Conclusion Findings highlight the necessity for gene signature accuracy for connectivity mapping, which should improve the clinical utility of future target compound discoveries.
Background BRAF mutation occurs in 8-15% of colon cancers (CC), and is associated with poor prognosis in metastatic disease. Compared to wild-type BRAF ( BRAFWT ) disease, stage II/III CC patients with BRAF mutant ( BRAFMT ) tumors have shorter overall survival after relapse; however, time-to-relapse is not significantly different. The aim of this investigation was to identify, and validate, novel predictors of relapse of stage II/III BRAFMT CC. Patients and methods We used gene expression data from a cohort of 460 patients (GSE39582) to perform a supervised classification analysis based on risk-of-relapse within BRAFMT stage II/III CC, to identify transcriptomic biomarkers associated with prognosis within this genotype. These findings were validated using immunohistochemistry in an independent population-based cohort of Stage II/III CC (n=691), applying Cox proportional hazards analysis to determine associations with survival. Results High gene expression levels of Bcl-xL, a key regulator of apoptosis, were associated with increased risk of relapse, specifically in BRAFMT tumors (HR=8.3, 95% CI 1.7-41.7), but not KRASMT/BRAFWT or KRASWT/BRAFWT tumors. High Bcl-xL protein expression in BRAFMT , untreated, stage II/III CC was confirmed to be associated with an increased risk of death in an independent cohort (HR=12.13, 95% CI 2.49-59.13). Additionally, BRAFMT tumors with high levels of Bcl-xL protein expression appeared to benefit from adjuvant chemotherapy (P for interaction =0.006), indicating the potential predictive value of Bcl-xL expression in this setting. Conclusions These findings provide evidence that Bcl-xL gene and/or protein expression identifies a poor prognostic subgroup of BRAFMT stage II/III CC patients, who may benefit from adjuvant chemotherapy. Key Message Using a combination of computational biology discovery and immunohistochemistry validation in independent patient cohorts, we show that high expression of the apoptosis regulator Bcl-xL is associated with disease relapse specifically within BRAF mutant stage II/III colon cancer. This data could enable tailored disease management to reduce relapse rates in the most aggressive subtype.
Abstract Technology advancements have enhanced our abilities to gain greater insight to the tumor environment. However, such emergent methods have brought new considerations in data storage, access and analysis. Modern large data projects and clinical trial materials could be explored to a greater degree if appropriate infrastructure could be built to support efforts. Such a scaffold would be an integrated and dynamic framework where novel hypotheses could be investigated in modern big data collections. Placing discovery back in the hands of the researcher through a reactive and supportive framework will enhance our understanding of cancer aetiology. The Cancer Integromics Research Application Framework (CIRAFm) was created to emulate current platforms that have a rigid analytical interface. We designed a robust architecture that supports a reactive user-friendly interface blended to several cross-platform coding technologies. It has been developed to accommodate an individualised framework where we can create an ‘app-store' of key software to fit research questions. In efforts to encapsulate and accelerate differing data types it sits astride of two NoSQL based database management systems, minimising data redundancy. Initial requirements have begun to create algorithms to investigate alignment-free applications on next generation sequencing (NGS) data enhancing analysis of spatial and temporal heterogeneity in cancer. Key drivers can be explored by assisting software to build correlative marker associations using Darwinian approaches such as genetic algorithms. Information requirements from external platforms are assisted through a novel domain specific language (DSL) to enable a singular interface. The modularised architecture of the platform is enabled through an Angular framework supported by interactive and dynamic data visualisation software, D3.js. Data analytics can be explored through the ‘apps' created, investigating new markers in large data and enhancing our understanding of tumor heterogeneity. Alignment-free phylogenetics of NGS data harnesses our capabilities to display fully sequence evolution in patient data. This highlights possible sub-types, markers of interest and maps to therapeutic compounds via the DSL to our externally developed drug discovery software, QUADrATiC. Future scientific endeavours further defining new tumor subtypes are paramount in efforts to help uncover treatment strategies. Unburdened by legacy pipelines and data, flexible and robust models of architecture provide an effective and efficient framework for research on our ever increasing search for precision medicine. Citation Format: Darragh G. McArt, Seedevi Senevirathne, Aideen Roddy, Jessica Black, Alan Gilmore, Suneil Jain, Philip Dunne, David Waugh. Integrative analytics: A framework for precision medicine [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2018; 2018 Apr 14-18; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2018;78(13 Suppl):Abstract nr 290.