Cancer research is undergoing a profound transformation driven by the rapid expansion of clinical, genomic, imaging, and real-world data. As Europe prepares for the implementation of the European Health Data Space (EHDS), the ability of health systems to effectively integrate, govern, and translate these diverse datasets will shape the next era of oncology. However, technological capacity alone is insufficient; sustained impact will depend on building trust, strengthening infrastructure, and supporting the people and cultures that enable data-intensive science. Cancer Research UK (CRUK), the nation's largest cancer charity, invests over £400 million annually in research and has launched a national data strategy to accelerate progress. In partnership with CRUK, we are working to develop a more connected and collaborative cancer data science ecosystem,one that brings together researchers across disciplines, identifies shared challenges, and co-designs practical solutions to overcome them. Through the CRUK Data Science Community and its Data Interest Groups, we highlight common obstacles across health system data, data reuse, public involvement, infrastructure, and training. We also present case studies demonstrating how integrated datasets, AI-enabled analytics, international collaboration, and federated approaches are already reshaping cancer research and clinical practice. By fostering a community-led approach to trustworthy, sustainable and FAIR data access, the UK has an opportunity to unlock the full potential of data-driven research and deliver meaningful benefits for people affected by cancer.
Temozolomide (TMZ) represents the cornerstone of therapy for glioblastoma (GBM). However, acquisition of resistance limits its therapeutic potential. The human kinome is an undisputable source of druggable targets, still, current knowledge remains confined to a limited fraction of it, with a multitude of under-investigated proteins yet to be characterized. Here, following a kinome-wide RNAi screen, pantothenate kinase 4 (PANK4) isuncovered as a modulator of TMZ resistance in GBM. Validation of PANK4 across various TMZ-resistant GBM cell models, patient-derived GBM cell lines, tissue samples, as well as in vivo studies, corroborates the potential translational significance of these findings. Moreover, PANK4 expression is induced during TMZ treatment, and its expression is associated with a worse clinical outcome. Furthermore, a Tandem Mass Tag (TMT)-based quantitative proteomic approach, reveals that PANK4 abrogation leads to a significant downregulation of a host of proteins with central roles in cellular detoxification and cellular response to oxidative stress. More specifically, as cells undergo genotoxic stress during TMZ exposure, PANK4 depletion represents a crucial event that can lead to accumulation of intracellular reactive oxygen species (ROS) and subsequent cell death. Collectively, a previously unreported role for PANK4 in mediating therapeutic resistance to TMZ in GBM is unveiled.
Mutually exclusive loss‐of‐function alterations in gene pairs are those that occur together less frequently than may be expected and may denote a synthetically lethal relationship (SSL) between the genes. SSLs can be exploited therapeutically to selectively kill cancer cells. Here, we analysed mutation, copy number variation, and methylation levels in samples from The Cancer Genome Atlas, using the hypergeometric and the Poisson binomial tests to identify mutually exclusive inactivated genes. We focused on gene pairs where one is an inactivated tumour suppressor and the other a gene whose protein product can be inhibited by known drugs. This provided an abundance of potential targeted therapeutics and repositioning opportunities for several cancers. These data are available on the MexDrugs website, https://bioinformaticslab.sussex.ac.uk/mexdrugs .
Applications of key technologies in biomedical research, such as qRT-PCR or LC-MS-based proteomics, are generating large biological (-omics) datasets which are useful for the identification and quantification of biomarkers in any research area of interest. Genome, transcriptome and proteome databases are already available for a number of model organisms including vertebrates and invertebrates. However, there is insufficient information available for protein sequences of certain invertebrates, such as the great pond snail Lymnaea stagnalis, a model organism that has been used highly successfully in elucidating evolutionarily conserved mechanisms of memory function and dysfunction. Here, we used a bioinformatics approach to designing and benchmarking a comprehensive central nervous system (CNS) proteomics database (LymCNS-PDB) for the identification of proteins from the CNS of Lymnaea by LC-MS-based proteomics. LymCNS-PDB was created by using the Trinity TransDecoder bioinformatics tool to translate amino acid sequences from mRNA transcript assemblies obtained from a published Lymnaea transcriptomics database. The blast-style MMSeq2 software was used to match all translated sequences to UniProtKB sequences for molluscan proteins, including those from Lymnaea and other molluscs. LymCNS-PDB contains 9628 identified matched proteins that were benchmarked by performing LC-MS-based proteomics analysis with proteins isolated from the Lymnaea CNS. MS/MS analysis using the LymCNS-PDB database led to the identification of 3810 proteins. Only 982 proteins were identified by using a non-specific molluscan database. LymCNS-PDB provides a valuable tool that will enable us to perform quantitative proteomics analysis of protein interactomes involved in several CNS functions in Lymnaea, including learning and memory and age-related memory decline.
In this paper we explore computational approaches that enable us to identify genes that have become essential in individual cancer cell lines. Using recently published experimental cancer cell line gene essentiality data, human protein-protein interaction (PPI) network data and individual cell-line genomic alteration data we have built a range of machine learning classification models to predict cell line specific acquired essential genes. Genetic alterations found in each individual cell line were modelled by removing protein nodes to reflect loss of function mutations and changing the weights of edges in each PPI to reflect gain of function mutations and gene expression changes. We found that PPI networks can be used to successfully classify human cell line specific acquired essential genes within individual cell lines and between cell lines, even across tissue types with AUC ROC scores of between 0.75 and 0.85. Our novel perturbed PPI network models further improved prediction power compared to the base PPI model and are shown to be more sensitive to genes on which the cell becomes dependent as a result of other changes. These improvements offer opportunities for personalised therapy with each individual’s cancer cell dependencies presenting a potential tailored drug target. The overriding motivation for predicting cancer cell line specific acquired essential genes is to provide a low-cost approach to identifying personalised cancer drug targets without the cost of exhaustive loss of function screening.
Applications of key technologies in bioscientific and biomedical research, such as qRT-PCR or LC-MS based proteomics, are generating large biological data sets (omics data) which are useful for the identification and quantification of biomarkers involved in molecular mechanisms of any research area of interest. Genome, transcriptome and proteome databases are already available for a number of model organisms including vertebrates and invertebrates. However, there is insufficient information available for protein sequences of certain invertebrates, such as the great pond snail Lymnaea stagnalis, a model organism that has been used highly successfully in elucidating evolutionarily conserved mechanisms of learning and memory, ageing and age-related as well as amyloid beta induced memory decline. Here, we present the design and benchmarking of a new proteomics database (LymSt-PDB) for the identification of proteins from the Central Nervous System (CNS) of Lymnaea stagnalis by LC-MS based proteomics.
Cohesin subunits are frequently mutated in cancer, but how they function as tumor suppressors is unknown. Cohesin mediates sister chromatid cohesion, but this is not always perturbed in cancer cells. Here, we identify a previously unknown role for cohesin. We find that cohesin is required to repress transcription at DNA double-strand breaks (DSBs). Notably, cohesin represses transcription at DSBs throughout interphase, indicating that this is distinct from its known role in mediating DNA repair through sister chromatid cohesion. We identified a cancer-associated SA2 mutation that supports sister chromatid cohesion but is unable to repress transcription at DSBs. We further show that failure to repress transcription at DSBs leads to large-scale genome rearrangements. Cancer samples lacking SA2 display mutational patterns consistent with loss of this pathway. These findings uncover a new function for cohesin that provides insights into its frequent loss in cancer.
Using pan-cancer data from The Cancer Genome Atlas (TCGA), we investigated how patterns in copy number alterations in cancer cells vary both by tissue type and as a function of genetic alteration. We find that patterns in both chromosomal ploidy and individual arm copy number are dependent on tumour type. We highlight for example, the significant losses in chromosome arm 3p and the gain of ploidy in 5q in kidney clear cell renal cell carcinoma tissue samples. We find that specific gene mutations are associated with genome-wide copy number changes. Using signatures derived from non-negative factorisation, we also find gene mutations that are associated with particular patterns of ploidy change. Finally, utilising a set of machine learning classifiers, we successfully predicted the presence of mutated genes in a sample using arm-wise copy number patterns as features. This demonstrates that mutations in specific genes are correlated and may lead to specific patterns of ploidy loss and gain across chromosome arms. Using these same classifiers, we highlight which arms are most predictive of commonly mutated genes in kidney renal clear cell carcinoma (KIRC).
The development of improved cancer therapies is frequently cited as an urgent unmet medical need. Here we describe how genetic interactions are being therapeutically exploited to identify novel targeted treatments for cancer. We discuss the current methodologies that use 'omics data to identify genetic interactions, in particular focusing on synthetic sickness lethality (SSL) and synthetic dosage lethality (SDL). We describe the experimental and computational approaches undertaken both in humans and model organisms to identify these interactions. Finally we discuss some of the identified targets with licensed drugs, inhibitors in clinical trials or with compounds under development.
Introduction: The development of improved cancer therapies is frequently cited as an urgent unmet medical need. Recent advances in platform technologies and the increasing availability of biological big data' are providing an unparalleled opportunity to systematically identify the key genes and pathways involved in tumorigenesis. The discoveries made using these new technologies may lead to novel therapeutic interventions.Areas covered: The authors discuss the current approaches that use big data' to identify cancer drivers. These approaches include the analysis of genomic sequencing data, pathway data, multi-platform data, identifying genetic interactions such as synthetic lethality and using cell line data. They review how big data is being used to identify novel drug targets. The authors then provide an overview of the available data repositories and tools being used at the forefront of cancer drug discovery.Expert opinion: Targeted therapies based on the genomic events driving the tumour will eventually inform treatment protocols. However, using a tailored approach to treat all tumour patients may require developing a large repertoire of targeted drugs.
Bioinformatics approaches are becoming ever more essential in translational drug discovery both in academia and within the pharmaceutical industry. Computational exploitation of the increasing volumes of data generated during all phases of drug discovery is enabling key challenges of the process to be addressed. Here, we highlight some of the areas in which bioinformatics resources and methods are being developed to support the drug discovery pipeline. These include the creation of large data warehouses, bioinformatics algorithms to analyse 'big data' that identify novel drug targets and/or biomarkers, programs to assess the tractability of targets, and prediction of repositioning opportunities that use licensed drugs to treat additional indications.