Summary:The rapid expansion of multi-omics data has enabled the development of molecular signatures-coordinated patterns of molecular features that serve as powerful biomarkers for diagnosis, prognosis, and therapeutic decision-making. Despite their potential, many published signatures suffer from limited reproducibility and narrow applicability, partly due to challenges in summarizing complex, multi-feature profiles into a single, statistically sound and biologically meaningful score. Here, we introduce sigscores, an R package that streamlines the computation of summary scores for molecular signatures. Building on the quality control principles of our earlier tool, sigQC, sigscores supports an extensive array of scoring metrics-including measures of central tendency, dispersion, and aggregation. It incorporates a resampling framework to generate empirical null distributions for rigorous significance assessment and provides integrated visualization tools for diagnostic evaluation. Optimized for parallel execution on multi-core systems, sigscores is well-suited for both exploratory research and high-throughput large-scale applications. Availability and implementation:Source code freely available for download on GitHub at https://github.com/alebarberis/sigscores, implemented in R and supported on MacOS and MS Windows.
High-throughput transcriptomics has made gene signatures central to interpreting gene expression data, with applications in diagnosis, prognosis, and prediction. Quantifying signature activity and assessing its robustness remain challenging because scoring methods primarily rely on various assumptions, and no single approach is universally optimal. Here, we present pysigscore, a Python framework for gene set scoring in bulk and single-cell RNA-seq data. pysigscore integrates 18 built-in scoring methods with a fully customisable scorer, allowing users to define and benchmark new scoring functions. It also provides reliability analyses, including p-value estimation and leave-one-out experiments, to assess the significance of scores and gene-level contributions. We validated pysigscore on the CCLE, TCGA, and PBMC datasets, recovering the expected enrichment in liver, hypoxia, inflammatory, and cell-cycle signatures.
Single-cell RNA sequencing (scRNA-seq) has profoundly reshaped our understanding of cellular diversity and functionality; however, accurate cell-type annotation is required for biological interpretation. Current annotation methods, which are predominantly reliant on gene expression alone or manual curation, suffer from subjectivity and a limited biological context. Here, we introduce a novel approach that integrates textual biological knowledge via gene embeddings, derived from fine-tuning Large Language Models (LLMs), with gene counts to enrich the input space for supervised models in automatic cell-type classification. In particular, we trained an XGBoost model and a multi-layer perceptron (MLP) to automatically classify the cell populations. We demonstrate that combining Modern-BERT embeddings and raw counts enhances the performance of MLPs, particularly in complex classification scenarios that involve subtle cell-subtype distinctions. Our results also show that ModernBERT generated better embeddings than smaller LLM architectures, underlining the value of enriched, biologically informed embeddings. By embedding prior knowledge from curated biological databases and literature, our approach enhances the MLP’s ability to distinguish sub-cell populations and biological signals. This work provides a scalable framework for integrating broader biological context into scRNA-seq analyses, offering new opportunities for downstream tasks such as gene regulatory network inference and cross-species annotation.
Myeloid cells play a central role in modulating immune response to cancer and have been linked to clinical outcomes in patients with solid tumors including nonresponse to immunotherapy. As such, there is a need to investigate their functionality in the tumor to identify new mechanisms of immune modulation for therapeutic benefit. We have found trogocytosis, a poorly understood cell-cell interaction mechanism, to be a prominent antibody-enhanced mode of monocyte-tumor cell communication across several aggressive cancer cell lines. The mechanistic processes governing monocyte trogocytosis of tumor cells and their effect on monocyte activation are largely unexplored. In this study, we leveraged functional analyses to interrogate the signaling pathways driving and resulting from monocyte trogocytosis of tumor cells. Both genome-wide and a targeted CRISPR/Cas9 screen identified Rho GTPase signaling as a significant regulator of trogocytosis by monocytes. Concordantly, transcriptomic analysis revealed an upregulation of integrin response and Rho GTPase signaling in trogocytic monocytes. Trogocytic monocytes also displayed higher expression of MHC II and complement related genes, suggesting an activation of antigen presentation machinery and inflammatory response. These results elucidate monocyte trogocytic function and targetable regulators that provide new insights into myeloid biology in disease and actionable avenues to impact antigen presentation and antibody modes of function. Suported by NIH-Oxford/Cambridge Scholars Program; US NIH ZIA BC 011332 and ZIA BC 011855; and NCI Cancer Moonshot. Tumor Immunology: Cellular Responses and Tumor Microevironment (TIME)
Single-cell technologies have significantly advanced our understanding of cellular heterogeneity by allowing the examination of individual cells at high resolution. Traditional single-cell RNA sequencing (scRNA-Seq) methods, which utilise whole cells, capture comprehensive RNA content. In contrast, emerging Multiome technologies, which simultaneously profile multiple omics such as gene expression (GEX) and chromatin accessibility, rely on nuclear RNA, potentially missing key cytoplasmic information. This discrepancy results in substantial technical and biological differences between GEX and scRNA-Seq datasets, making it challenging to integrate the data and perform downstream tasks, such as cell-type classification. To address this challenge, we introduce GENESIS (Gene Expression Normalisation and Enhancement for Single-cell Integrated Sequencing), a novel computational framework designed to transform GEX data from Multiome experiments into enhanced, scRNA-Seq like profiles. Utilising advanced generative models—including Variational Autoencoders, Generative Adversarial Networks, and a tailored VAE_UNet architecture—GENESIS can generate high-quality data by modelling and compensating for the inherent differences between nuclear and cytoplasmic RNA. Our comprehensive evaluations show that GENESIS, particularly through the VAE_UNet model, generates synthetic scRNA-Seq data that closely resembles the resolution and biological accuracy of whole-cell sequencing, thereby improving downstream tasks, especially cell-type classification.
Summary: GeneFEAST, implemented in Python, is a gene-centric functional enrichment analysis summarisation and visualisation tool that can be applied to large functional enrichment analysis (FEA) results arising from upstream FEA pipelines. It produces a systematic, navigable HTML report, making it easy to identify sets of genes putatively driving multiple enrichments and to explore gene-level quantitative data first used to identify input genes. Further, GeneFEAST can compare FEA results from multiple studies, making it possible, for example, to highlight patterns of gene expression amongst genes commonly differentially expressed in two sets of conditions, and giving rise to shared enrichments under those conditions. GeneFEAST offers a novel, effective way to address the complexities of linking up many overlapping FEA results to their underlying genes and data, advancing gene-centric hypotheses, and providing pivotal information for downstream validation experiments. Availability: GeneFEAST is available at https://github.com/avigailtaylor/GeneFEAST Contact: avigail.taylor@well.ox.ac.uk
Tumor hypoxia drives metabolic shifts, cancer progression, and therapeutic resistance. Challenges in quantifying hypoxia have hindered the exploitation of this potential "Achilles' heel." While gene expression signatures have shown promise as surrogate measures of hypoxia, signature usage is heterogeneous and debated. Here, we present a systematic pan-cancer evaluation of 70 hypoxia signatures and 14 summary scores in 104 cell lines and 5,407 tumor samples using 472 million length-matched random gene signatures. Signature and score choice strongly influenced the prediction of hypoxia in vitro and in vivo. In cell lines, the Tardon signature was highly accurate in both bulk and single-cell data (94% accuracy, interquartile mean). In tumors, the Buffa and Ragnum signatures demonstrated superior performance, with Buffa/mean and Ragnum/interquartile mean emerging as the most promising for prospective clinical trials. This work delivers recommendations for experimental hypoxia detection and patient stratification for hypoxia-targeting therapies, alongside a generalizable framework for signature evaluation.
Abstract Triple-negative breast cancer (TNBC) is the subtype of breast cancer that has the worst clinical outcome due to the absence of specific targeting agents. TNBC is highly heterogeneous and within the tumor microenvironment, hypoxia emerges as a pivotal factor contributing to its aggressive biology. Solid tumors experience a fluctuating oxygen supply, leading to the formation of zones with prolonged oxygen deprivation. There have been few studies describing the response to chronic hypoxia [7-14 days] and this project investigates how chronic hypoxia contributes to tumour heterogeneity in a temporal manner. We performed scRNA-seq analysis on MDA-MB-231 cells subjected to a 14-day hypoxia treatment at 1% oxygen. We collected samples from both normoxia and hypoxia conditions on Days 1, 2, 7, and 14. A total of 40,000 cells were sequenced. There were 6 hypoxic cell clusters which evolved over the time course. Interestingly, chronic hypoxia induces various transient subpopulations. Cells in hypoxia Day 7 grew more slowly, before recovering growth at Day 14, marking a crucial adaptation point at Day 7 in chronic hypoxia. Our study suggests that chronic hypoxia cells conserve energy through G1 cell cycle arrest and Myc suppression as an adaptive mechanism to cope with chronic hypoxia. We also observed a higher expression of cancer stem cell markers [CD44, CXCR4 and KLF4] in hypoxia subpopulations relative to normoxia subpopulations, especially in the chronic hypoxia subpopulations. Notably, we identified a distinct CD24+/CD49b-high population, comprising the largest proportion of the hypoxic Day 7 samples which disappears by Day 14. This population is characterised by the upregulation of cystatin family genes- CST 1, 4 and 7 and may be pro-metastatic [upregulation of EMP1, ST3GAL6 and KRT19 genes] compared to earlier and later hypoxia phases. Interestingly, KRT19 is more abundant in peripheral blood of breast cancer patients with node metastasis, although its role in stemness of breast cancer is controversial. The heterogeneity we found here may explain this, with the existence of a transient population. Chronic hypoxia could perhaps have a role in reprogramming cancer stem cell-like cells into a pro-metastatic state. In summary, our study provides insights on how hypoxia induces protective translational reprogramming and drives plasticity in cancer cells. Understanding these processes may facilitate the development of targeted treatments to combat therapy resistance in TNBC. Citation Format: May Sin Ke, Badran Elshenawy, Anjali Arora, Helen Sheldon, Adrian Harris, Francesca Buffa. Chronic hypoxia in TNBC breast cancer induces development of a novel population of CD24+/CD49b-high cells characterized by high expression of CST1 [abstract]. In: Proceedings of the AACR Special Conference in Cancer Research: Advances in Breast Cancer Research; 2023 Oct 19-22; San Diego, California. Philadelphia (PA): AACR; Cancer Res 2024;84(3 Suppl_1):Abstract nr B056.
Background It is uncertain which biological features underpin the response of rectal cancer (RC) to radiotherapy. No biomarker is currently in clinical use to select patients for treatment modifications. Methods We identfied two cohorts of patients (total N = 249) with RC treated with neoadjuvant radiotherapy (45Gy/ 25) plus fl uoropyrimidine. This discovery set included 57 cases with pathological complete response (pCR) to chemoradiotherapy (23%). Pre-treatment cancer biopsies were assessed using transcriptome-wide mRNA expression and targeted DNA sequencing for copy number and driver mutations. Biological candidate and machine learning (ML) approaches were used to identify predictors of pCR to radiotherapy independent of tumour stage. Findings were assessed in 107 cases from an independent validation set (GSE87211). Findings Three gene expression sets showed signi fi cant independent associations with pCR: Fibroblast-TGF P Response Signature (F-TBRS) with radioresistance; and cytotoxic lymphocyte (CL) expression signature and consensus molecular subtype CMS1 with radiosensitivity. These associations were replicated in the validation cohort. In parallel, a gradient boosting machine model comprising the expression of 33 genes generated in the discovery cohort showed high performance in GSE87211 with 90% sensitivity, 86% speci fi city. Biological and ML signatures indicated similar mechanisms underlying radiation response, and showed better AUC and p-values than published transcriptomic signatures of radiation response in RC. Interpretation RCs responding completely to chemoradiotherapy (CRT) have biological characteristics of immune response and absence of immune inhibitory TGF P signalling. These tumours may be identi fi ed with a potential biomarker based on a 33 gene expression signature. This could help select patients likely to respond to treatment with a primary radiotherapy approach as for anal cancer. Conversely, those with predicted radio resistance may be candidates for clinical trials evaluating addition of immune-oncology agents and stromal TGF P signalling inhibition. Funding The Strati fi cation in Colorectal Cancer Consortium (S:CORT) was funded by the Medical Research Council and Cancer Research UK (MR/M016587/1).
Extension of prostate cancer beyond the primary site by local invasion or nodal metastasis is associated with poor prognosis. Despite significant research on tumour evolution in prostate cancer metastasis, the emergence and evolution of cancer clones at this early stage of expansion and spread are poorly understood. We aimed to delineate the routes of evolution and cancer spread within the prostate and to seminal vesicles and lymph nodes, linking these to histological features that are used in diagnostic risk stratification. We performed whole-genome sequencing on 42 prostate cancer samples from the prostate, seminal vesicles and lymph nodes of five treatment-naive patients with locally advanced disease. We spatially mapped the clonal composition of cancer across the prostate and the routes of spread of cancer cells within the prostate and to seminal vesicles and lymph nodes in each individual by analysing a total of > 19,000 copy number corrected single nucleotide variants. In each patient, we identified sample locations corresponding to the earliest part of the malignancy. In patient 10, we mapped the spread of cancer from the apex of the prostate to the seminal vesicles and identified specific genomic changes associated with the transformation of adenocarcinoma to amphicrine morphology during this spread. Furthermore, we show that the lymph node metastases in this patient arose from specific cancer clones found at the base of the prostate and the seminal vesicles. In patient 15, we observed increased mutational burden, altered mutational signatures and histological changes associated with whole genome duplication. In all patients in whom histological heterogeneity was observed (4/5), we found that the distinct morphologies were located on separate branches of their respective evolutionary trees. Our results link histological transformation with specific genomic alterations and phylogenetic branching. These findings have implications for diagnosis and risk stratification, in addition to providing a rationale for further studies to characterise the genetic changes causally linked to morphological transformation. Our study demonstrates the value of integrating multi-region sequencing with histopathological data to understand tumour evolution and identify mechanisms of prostate cancer spread.
AbstractPurpose: While there are several prognostic classifiers, to date, there are no validated predictive models that inform treatment selection for oropharyngeal squamous cell carcinoma (OPSCC). Our aim was to develop clinical and/or biomarker predictive models for patient outcome and treatment escalation for OPSCC. Experimental Design: We retrospectively collated clinical data and samples from a consecutive cohort of OPSCC cases treated with curative intent at ten secondary care centers in United Kingdom and Poland between 1999 and 2012. We constructed tissue microarrays, which were stained and scored for 10 biomarkers. We then undertook multivariable regression of eight clinical parameters and 10 biomarkers on a development cohort of 600 patients. Models were validated on an independent, retrospectively collected, 385-patient cohort. Results: A total of 985 subjects (median follow-up 5.03 years, range: 4.73–5.21 years) were included. The final biomarker classifier, comprising p16 and survivin immunohistochemistry, high-risk human papillomavirus (HPV) DNA in situ hybridization, and tumor-infiltrating lymphocytes, predicted benefit from combined surgery + adjuvant chemo/radiotherapy over primary chemoradiotherapy in the high-risk group [3-year overall survival (OS) 63.1% vs. 41.1%, respectively, HR = 0.32; 95% confidence interval (CI), 0.16–0.65; P = 0.002], but not in the low-risk group (HR = 0.4; 95% CI, 0.14–1.24; P = 0.114). On further adjustment by propensity scores, the adjusted HR in the high-risk group was 0.34, 95% CI = 0.17–0.67, P = 0.002, and in the low-risk group HR was 0.5, 95% CI = 0.1–2.38, P = 0.384. The concordance index was 0.73. Conclusions: We have developed a prognostic classifier, which also appears to demonstrate moderate predictive ability. External validation in a prospective setting is now underway to confirm this and prepare for clinical adoption.
Ubiquitination is a crucial posttranslational modification required for the proper repair of DNA double-strand breaks (DSBs) induced by ionizing radiation (IR). DSBs are mainly repaired through homologous recombination (HR) when template DNA is present and nonhomologous end joining (NHEJ) in its absence. In addition, microhomology-mediated end joining (MMEJ) and single-strand annealing (SSA) provide backup DSBs repair pathways. However, the mechanisms controlling their use remain poorly understood. By using a high-resolution CRISPR screen of the ubiquitin system after IR, we systematically uncover genes required for cell survival and elucidate a critical role of the E3 ubiquitin ligase SCFcyclin F in cell cycle-dependent DSB repair. We show that SCFcyclin F-mediated EXO1 degradation prevents DNA end resection in mitosis, allowing MMEJ to take place. Moreover, we identify a conserved cyclin F recognition motif, distinct from the one used by other cyclins, with broad implications in cyclin specificity for cell cycle control.
Artificial intelligence (AI) techniques are increasingly applied across various domains, favoured by the growing acquisition and public availability of large, complex datasets. Despite this trend, AI publications often suffer from lack of reproducibility and poor generalisation of findings, undermining scientific value and contributing to global research waste. To address these issues and focusing on the learning aspect of the AI field, we present RENOIR (REpeated random sampliNg fOr machIne leaRning), a modular open-source platform for robust and reproducible machine learning (ML) analysis. RENOIR adopts standardised pipelines for model training and testing, introducing elements of novelty, such as the dependence of the performance of the algorithm on the sample size. Additionally, RENOIR offers automated generation of transparent and usable reports, aiming to enhance the quality and reproducibility of AI studies. To demonstrate the versatility of our tool, we applied it to benchmark datasets from health, computer science, and STEM (Science, Technology, Engineering, and Mathematics) domains. Furthermore, we showcase RENOIR’s successful application in recently published studies, where it identified classifiers for SET2D and TP53 mutation status in cancer. Finally, we present a use case where RENOIR was employed to address a significant pharmacological challenge—predicting drug efficacy. RENOIR is freely available at https://github.com/alebarberis/renoir .
Gene Regulatory Networks (GRNs) play a fundamental role in orchestrating the expression of our genes through complex interactions between DNA, RNA, proteins, and other molecules. Accurately reconstructing such networks from gene expression data is a critical yet challenging task in Systems Biology due to their intricate nature and limited data availability. In this work, we introduce a novel Forest-based Evolutionary Algorithm (FP) designed for reconstructing Boolean GRNs from time series data of gene expressions. Unlike traditional methods that struggle with scalability and accurate representation of regulatory interactions, FP utilizes a forest structure where each tree represents the logical relationships between genes, enhancing the model’s capacity to depict complex networks efficiently. Our comprehensive testing indicates that FP rapidly converges towards potential solutions within a limited number of generations, although a higher fitness score does not always equate to a more accurate GRN representation. Implementing mini-batching techniques, inspired by their effectiveness in gradient descent optimization, shows promise in improving computational efficiency without sacrificing performance. A comparative analysis against the main state-of-the-art approaches reveals FP’s tendency towards conservative predictions, emphasizing precision over recall, making it particularly suitable for contexts where the cost of false positives is high. These initial results suggest that FP is a robust and efficient tool for GRN inference.