Accurately predicting chemotherapy response remains a major challenge in precision oncology. Although machine-learning models based on tumour omics data have shown promise, the majority of existing studies are trained and evaluated on pre-clinical cell-line datasets, leaving their clinical applicability insufficiently characterised. In this study, we systematically evaluate a range of transfer-learning strategies for chemotherapy response prediction under realistic clinical constraints using patient data from The Cancer Genome Atlas (TCGA). Rather than proposing a new predictive model, we focus on assessing the effectiveness and limitations of commonly used approaches for transferring pre-clinical knowledge to clinical settings. These include cell-line-validated biomarkers, biologically informed feature representations, direct application of pre-clinical deep-learning models, model fine-tuning, and hybrid strategies that integrate pre-clinical predictions with clinical data. All methods are evaluated within a unified framework using consistent cohort construction, shared performance metrics, and bias-controlled validation procedures. Across multiple drugs and molecular data types, we find that most transfer strategies—including biomarker-based feature selection and direct pre-clinical model transfer—fail to produce robust or consistent improvements in clinical prediction performance. In contrast, conservative approaches based on fine-tuning pre-clinical models or incorporating pre-clinical predictions as features in clinical models yield more stable and reproducible gains. Further improvements are observed when basic pre-treatment clinical variables are integrated. Together, our results demonstrate the practical boundaries of pre-clinical to clinical transfer for drug response prediction and highlight hybrid and fine-tuning strategies as more reliable baselines for future translational modelling efforts.
The drug discovery process often employs phenotypic and target-based virtual screening to identify potential drug candidates. Despite the longstanding dominance of target-based approaches, phenotypic virtual screening is undergoing a resurgence due to its potential being now better understood. In the context of cancer cell lines, a well-established experimental system for phenotypic screens, molecules are tested to identify their whole-cell activity, as summarized by their half-maximal inhibitory concentrations. Machine learning has emerged as a potent tool for computationally guiding such screens, yet important research gaps persist, including generalization and uncertainty quantification. To address this, we leverage a clustering-based validation approach, called Leave Dissimilar Molecules Out (LDMO). This strategy enables a more rigorous assessment of model generalization to structurally novel compounds. This study focuses on applying Conformal Prediction (CP), a model-agnostic framework, to predict the activities of novel molecules on specific cancer cell lines. A total of 4320 independent models were evaluated across 60 cell lines, 5 CP variants, 2 set features, and training-test splits, providing strong and consistent results. From this comprehensive evaluation, we concluded that, regardless of the cell line or model, novel molecules with smaller CP-calculated confidence intervals tend to have smaller predicted errors once measured activities are revealed. It was also possible to anticipate the activities of dissimilar test molecules across 50 or more cell lines. These outcomes demonstrate the robust efficacy that LDMO-based models can achieve in realistic and challenging scenarios, thereby providing valuable insights for enhancing decision-making processes in drug discovery.
Predicting personalized treatment-specific responses in cancer patients requires not only robust experimental models such as PDX, which generate accurate training data on different treatment responses for the same model, but also computational models that can translate this knowledge into decisions for real patients in a timely manner. The translatability of PDX data into patient-specific outcomes is limited by the challenges in the data size requirements for training machine learning models. Previously, ML models have been developed for the two most abundant cancer types, viz. BRCA and CRC, but the same could not be scaled up to other cancer types with a smaller number of PDX samples. Here, we provide an ML framework to train a single pan-cancer, pan-treatment model for predicting treatment outcomes. We show that such models give promising results for all cancer types considered and reproduce the accuracy levels of individually trained cancer types. In the proposed model, all PDX genomic profiles from all cancer types are used as the training data, and instead of partitioning them into cancer types for each model, the cancer type and treatment name are appended as the input features of the training model. Using genomic-only and treatment-only embeddings and combining them with PCA-based dimensionality reduction, our models show promising results and provide a framework for further improvements and real-time use for best treatment selections for cancer patients. ### Competing Interest Statement The authors have declared no competing interest.
Virtual Screening (VS) of large compound libraries using Artificial Intelligence (AI) models is a highly effective approach for early drug discovery. Data splitting is crucial for benchmarking the performance of such AI models. Traditional random data splits often result in structurally similar molecules in both training and test sets, which conflict with the reality of VS libraries that typically contain structurally diverse compounds. To tackle this challenge, scaffold split, which groups molecules by shared core structure, and Butina clustering, which clusters molecules by chemotypes, have long been used. However, we show that these methods still introduce high similarities between training and test sets, leading to overestimated model performance. Our study examined four representative AI models across 60 NCI-60 datasets, each comprising approximately 33,000–54,000 molecules tested on different cancer cell lines. Each dataset was split in four ways: random, scaffold, Butina clustering and the more realistic Uniform Manifold Approximation and Projection (UMAP) clustering. Using Linear Regression, Random Forest, Transformer-CNN, and GEM, we trained a total of 8400 models and evaluated under four splitting methods. These comprehensive results show that UMAP split provides more challenging and realistic benchmarks for model evaluation, followed by Butina splits, then scaffold splits and closely after random splits. Consequently, we recommend using UMAP splits instead of overly optimistic Butina splits and especially scaffold splits for molecular property prediction, including VS. Lastly, we illustrate how misaligned ROC AUC is with VS goals, despite its common use. The code and datasets for reproducibility are available at https://github.com/Rong830/UMAP_split_for_VS and archived in https://zenodo.org/records/14736486 . Scientific contribution This work advances the field by introducing UMAP clustering as a robust splitting method for molecular datasets, improving over traditional methods like Butina clustering and especially scaffold splits. It offers a new evaluation framework to benchmark AI models under more realistic conditions, fostering progress in molecular property prediction. The findings also show how inappropriate the use of ROC AUC for virtual screening (VS) continues to be, despite its popularity, emphasizing the need for context-specific evaluation metrics.
Atezolizumab is a treatment for metastatic urothelial carcinoma (mUC), yet only 23% of mUC patients benefit from it. Worse yet, accurately predicting such responders remains challenging, despite existing biomarkers. Here we employed eight machine learning (ML) algorithms to predict mUC patient response to atezolizumab using tumours’ gene expression profiling and clinical data from two independent cohorts. The CART-OMC model developed on the discovery dataset achieved the highest performance, with a validation set Matthews correlation coefficient (MCC) of 0.437, using the expressions of just 29 ML-selected genes, including CXCL9 and IFNG. Univariate biomarkers like TMB, TNB, and PD-L1 were less predictive with MCCs of 0, 0.316, and 0, respectively. Upon merging these datasets, the best-performing model (LGBM-OMC; MCC of 0.252) also outperformed top modelling approaches such as EaSIeR (MCC ~ 0) and JADBio (MCC of 0.179). We make these promising ML models freely available to predict atezolizumab response in other mUC patients.
Temozolomide is the primary chemotherapeutic agent and first-line treatment for low-grade glioma. Although low-grade gliomas are generally less aggressive than high-grade gliomas, they can eventually progress into high-grade gliomas, making it crucial to maximise the efficacy of initial treatment. We analysed data from 109 patients with low-grade gliomas in The Cancer Genome Atlas to evaluate the predictive performance of 12 machine learning classification algorithms for temozolomide response, using six types of omics data. Cross-validation and bootstrapping bias correction were applied to compare these models with a conventional biomarker-based model using promoter methylation status of O6-methylguanine-DNA methyltransferase. The Matthews Correlation Coefficient (MCC) was used as the primary evaluation metric. The microRNA-based model using the Extreme Gradient Boosting algorithm achieved the best performance (MCC = 0.447), outperforming both the automated machine learning method JADBio (MCC = 0.250) and the biomarker-based model (MCC = 0.331). Incorporating clinical variables, such as patient age and Karnofsky score, further improved predictive power, with the logistic regression model with optimal model complexity achieving the highest MCC (0.483). Feature importance analysis on the best model revealed six predictive microRNAs, including three tumour-related factors (miR-335, let-7f, and miR-7-2) and three potential biomarkers (miR-204, miR-6513, and miR-376). This study systematically demonstrates the potential of large-scale analyses combining machine learning and omics data to predict temozolomide response, offering superior predictive accuracy compared with standard biomarkers. However, validation in independent clinical datasets remains necessary before clinical translation.
An innovative mechanism to inhibit the PD1/PDL1 interaction is PDL1 dimerization induced by small-molecule PDL1 binders. Structure-based virtual screening is a promising approach to discovering such small-molecule PD1/PDL1 inhibitors. Here we investigate which type of generic scoring functions is most suitable to tackle this problem. We consider CNN-Score, an ensemble of convolutional neural networks, as the representative of machine-learning scoring functions. We also evaluate Smina, a commonly used classical scoring function, and IFP, a top structural fingerprint similarity scoring function. These three types of scoring functions were evaluated on two test sets sharing the same set of small-molecule PD1/PDL1 inhibitors, but using different types of inactives: either true inactives (molecules with no in vitro PD1/PDL1 inhibition activity) or assumed inactives (property-matched decoy molecules generated from each active). On both test sets, CNN-Score performed much better than Smina, which in turn strongly outperformed IFP. The fact that the latter was the case, despite precluding any possibility of exploiting decoy bias, demonstrates the predictive value of CNN-Score for PDL1. These results suggest that re-scoring Smina-docked molecules with CNN-Score is a promising structure-based virtual screening method to discover new small-molecule inhibitors of this important therapeutic target.
The discovery of effective therapeutics remains a complex, costly, and time-consuming endeavor, characterized by high failure rates and significant resource investments. A central bottleneck in early-stage drug discovery is identifying suitable hit compounds with moderate affinity for known biological targets. Although advancements occur, current in silico virtual screening methods are subject to limitations, including model overfitting, data bias, and constrained interpretability in their predictive processes. In this study, we present SCORCH2, a machine learning-based framework designed to simultaneously enhance the performance and interpretability of virtual screening by leveraging interaction features. Comparing with its predecessor SCORCH, SCORCH2 exhibits superior predictive accuracy and generalizability across a wide range of biological targets. Importantly, SCORCH2 demonstrates robust hit identification capabilities on previously unseen targets, indicating strong transferability. Furthermore, SCORCH2 obviates the need for meticulous docking pose selection, streamlining the screening process. These advances highlight the potential of SCORCH2 as a valuable tool in accelerating drug discovery campaigns.
The translatability of patient-derived xenograft (PDX)-generated clinical data into patient-specific outcomes for therapeutic guidance is limited by the challenges in generalizability of models across patients, treatments, and cancer types. Previously, machine learning (ML) models have been developed for the two most abundant cancer types, i.e. breast cancer and colorectal cancer, but these are unusable in other cancer types because each treatment/cancer type requires a different model to be trained. Here, we provide an ML framework to train a single pan-cancer, pan-treatment model for predicting treatment outcomes. We show that such models give promising results for all cancer types considered and reproduce the accuracy levels of individually trained cancer types. In the proposed model, all PDX genomic profiles from all cancer types are used as the training data, and instead of partitioning them into cancer types for each model, the cancer type and treatment name are appended as the input features of the training model. Using genomic-only and treatment-only embeddings and combining them with principal component analysis-based dimensionality reduction, our models show promising results and provide a framework for further improvements and real-time use for best treatment selections for cancer patients.
Artificial intelligence is increasingly driving early drug design, offering novel approaches to virtual screening. Phenotypic virtual screening (PVS) aims to predict how cancer cell lines respond to different compounds by focusing on observable characteristics rather than specific molecular targets. Some studies have suggested that deep learning may not be the best approach for PVS. However, these studies are limited by the small number of tested molecules as well as not employing suitable performance metrics and dissimilar-molecules splits better mimicking the challenging chemical diversity of real-world screening libraries. Here we prepared 60 datasets, each containing approximately 30,000 to 50000 molecules tested for their growth inhibitory activities on one of the NCI-60 cancer cell lines. We evaluated the performance of five machine learning algorithms for PVS on these 60 problem instances. To provide a comprehensive evaluation, we employed two model validation types: the random split and the dissimilar-molecules split. The models were primarily evaluated using hit rate, a more suitable metric in VS contexts. The results show that all models are more challenged by test molecules that are substantially different from those in the training data. In both validation types, the D-MPNN algorithm, a graph-based deep neural network, was found to be the most suitable for building predictive models for this PVS problem. ### Competing Interest Statement The authors have declared no competing interest.
Artificial intelligence is increasingly driving early drug design, offering novel approaches to virtual screening. Phenotypic virtual screening (PVS) aims to predict how cancer cell lines respond to different compounds by focusing on observable characteristics rather than specific molecular targets. Some studies have suggested that deep learning may not be the best approach for PVS. However, these studies are limited by the small number of tested molecules as well as not employing suitable performance metrics and dissimilar-molecules splits better mimicking the challenging chemical diversity of real-world screening libraries. Here we prepared 60 datasets, each containing approximately 30 000-50 000 molecules tested for their growth inhibitory activities on one of the NCI-60 cancer cell lines. We conducted multiple performance evaluations of each of the five machine learning algorithms for PVS on these 60 problem instances. To provide even a more comprehensive evaluation, we used two model validation types: the random split and the dissimilar-molecules split. Overall, about 14 440 training runs aczross datasets were carried out per algorithm. The models were primarily evaluated using hit rate, a more suitable metric in VS contexts. The results show that all models are more challenged by test molecules that are substantially different from those in the training data. In both validation types, the D-MPNN algorithm, a graph-based deep neural network, was found to be the most suitable for building predictive models for this PVS problem.
CIC::DUX4 sarcoma (CDS) is a rare but highly aggressive undifferentiated small round cell sarcoma driven by a fusion between the tumor suppressor Capicua (CIC) and DUX4. Currently, there are no effective treatments and efforts to identify and translate better therapies are limited by the scarcity of patient tumor samples and cell lines. To address this limitation, we generated three genetically engineered mouse models of CDS (Ch7CDS, Ai9CDS, and TOPCDS). Remarkably, chimeric mice from all three conditional models developed spontaneous soft tissue tumors and disseminated disease in the absence of Cre-recombinase. The penetrance of spontaneous (Cre-independent) tumor formation was complete irrespective of bi-allelic Cic function and the distance between adjacent loxP sites. Characterization of soft tissue and presumed metastatic tumors showed that they consistently expressed the CIC::DUX4 fusion protein and many downstream markers of the disease credentialing the models as CDS. In addition, tumor-derived cell lines were generated and ChIP-seq was preformed to map fusion-gene specific binding using an N-terminal HA epitope tag. These datasets, along with paired H3K27ac ChIP-sequencing maps, validate CIC::DUX4 as a neomorphic transcriptional activator. Moreover, they are consistent with a model where ETS family transcription factors are cooperative and redundant drivers of the core regulatory circuitry in CDS.
Background: Gemcitabine is a first-line chemotherapy for pancreatic adenocarcinoma (PAAD), but many PAAD patients do not respond to gemcitabine-containing treatments. Being able to predict such nonresponders would hence permit the undelayed administration of more promising treatments while sparing gemcitabine life-threatening side effects for those patients. Unfortunately, the few predictors of PAAD patient response to this drug are weak, none of them exploiting yet the power of machine learning (ML). Methods: Here, we applied ML to predict the response of PAAD patients to gemcitabine from the molecular profiles of their tumors. More concretely, we collected diverse molecular profiles of PAAD patient tumors along with the corresponding clinical data (gemcitabine responses and clinical features) from the Genomic Data Commons resource. From systematically combining 8 tumor profiles with 16 classification algorithms, each of the resulting 128 ML models was evaluated by multiple 10-fold cross-validations. Results: Only 7 of these 128 models were predictive, which underlines the importance of carrying out such a large-scale analysis to avoid missing the most predictive models. These were here random forest using 4 selected mRNAs [0.44 Matthews correlation coefficient (MCC), 0.785 receiver operating characteristic–area under the curve (ROC-AUC)] and XGBoost combining 12 DNA methylation probes (0.32 MCC, 0.697 ROC-AUC). By contrast, the hENT1 marker obtained much worse random-level performance (practically 0 MCC, 0.5 ROC-AUC). Despite not being trained to predict prognosis (overall and progression-free survival), these ML models were also able to anticipate this patient outcome. Conclusions: We release these promising ML models so that they can be evaluated prospectively on other gemcitabine-treated PAAD patients.
Virtual Screening (VS) of vast compound libraries guided by Artificial Intelligence (AI) models is a highly productive approach to early drug discovery. Data splitting is crucial for better benchmarking of such AI models. Traditional random data splits produce similar molecules between training and test sets, conflicting with the reality of VS libraries which mostly contain structurally distinct compounds. Scaffold split, grouping molecules by shared core structure, is widely considered to reflect this real-world scenario. However, here we show that the scaffold split also overestimates VS performance. The reason is that molecules with different chemical scaffolds are often similar, which hence introduces unrealistically high similarities between training molecules and test molecules following a scaffold split. Our study examined three representative AI models on 60 NCI-60 datasets, each with approximately 30,000 to 50,000 molecules tested on a different cancer cell line. Each dataset was split with three methods: scaffold, Butina clustering and the more accurate Uniform Manifold Approximation and Projection (UMAP) clustering. Regardless of the model, model performance is much worse with UMAP splits from the results of the 2100 models trained and evaluated for each algorithm and split. These robust results demonstrate the need for more realistic data splits to tune, compare, and select models for VS. For the same reason, avoiding the scaffold split is also recommended for other molecular property prediction problems. The code to reproduce these results is available at https://github.com/ScaffoldSplitsOverestimateVS
Abstract Treatment at a young age places survivors of childhood cancers at a significantly elevated risk of developing life-threatening conditions such as cardiac dysfunction and secondary neoplasms. To better define the genomic impact of cancer therapy in children, we studied a diverse cohort of childhood tumors that had been heavily treated with multiple chemotherapies. We retrospectively collected detailed therapeutic data (including drug and dosage for each treatment cycle) for three cohorts of aggressive and hard-to-cure pediatric cancers: KiCS (The Hospital for Sick Children Toronto), ZCC (Zero Childhood Cancer, Sydney, Australia), and MSK (Memorial Sloan Kettering, New York, USA). We observed that therapy had a sizeable contribution to the mutation load in the treated samples, some of which could be detected in the form of mutational signatures. Performing a comprehensive signature analysis, we detected 69 mutational signatures, including 49 COSMIC and 20 novel signatures. We observed almost three times more exclusive signatures in treated tumors (14 vs 5 in untreated tumors). Excluding the hypermutator tumors, the difference became even more striking: 15 signatures were exclusively found in treated tumors while there were no signatures exclusive to treatment-naïve tumors. Platinum chemotherapies were the most potent DNA-damaging therapies – leading to the most somatic mutations, in the shortest time. In addition to known COSMIC platinum signatures, we found five novel putative signatures of platinum resistance. Using detailed therapy exposure data extracted from clinical charts, we found that a minimum 91 days and a burden of 0.5 mut/Mb were required to detect evidence of resistance. Nearly three quarters of the tumors passing both thresholds displayed the relevant resistance-associated signatures. Remarkably, over a third (35%) of tumors treated with platinum drugs displayed measurable resistance-associated mutations after just 12 months. This reached the 50% mark at the 33-month timepoint, then plateaued. Similar, albeit slower, trends were detected for other therapies. We also found that therapy signatures could appear at both clonal and subclonal levels as well as among clustered events. This multi-center analysis of heavily treated childhood tumors presents preliminary insights with potential to monitor drug resistance or interfere with its development. Citation Format: Mehdi Layeghifard, Nicholas Light, Erik Bergstrom, Marcos Diaz Gay, Mathepan J. Mahendralingam, Sasha Blay, Scott Davidson, Pedro L. Ballester, Rawan Hammad, Noemi Fuentes Bolanos, Shimaa Nassif, Nirav H. Thacker, Chelsea Mayoh, David Malkin, Elli Papaemmanuil, Mark J. Cowley, Anita Villani, Ludmil B. Alexandrov, Adam Shlien. Chemotherapy is a major mutagen in relapsed childhood cancer [abstract]. In: Proceedings of the AACR Special Conference in Cancer Research: Advances in Pediatric Cancer Research; 2024 Sep 5-8; Toronto, Ontario, Canada. Philadelphia (PA): AACR; Cancer Res 2024;84(17 Suppl):Abstract nr B070.
The response to targeted therapies and immune checkpoint inhibitors for patients suffering from metastatic clear cell renal cell carcinoma (ccRCC) is heterogeneous and currently not predictable in clinic. In this work, a comprehensive integrated study of 700 ccRCCs profiled by DNA methylation and RNA sequencing showed that the hyper-methylated tumors exhibited a worse prognosis, a higher fraction of cycling tumor cells and a lower activity of homeobox transcription factors. To translate the use of DNA methylation information into a clinical setting, we developed a simple model accurately predicting the ccRCC methylation subtypes (AUC-ROCs of 0.91) from two gene expression ratios (IGF2BP3/PCCA, TNNT1/TMEM88). In addition, these methylation subtypes were significantly associated with the therapeutic outcome of patients to anti-PD-1, mTOR inhibitor or tyrosine kinase inhibitor therapies. Overall, our framework for predicting the ccRCC DNA methylation subtypes from targeted gene expression data is easy to translate in clinic and contributes to better personalization of ccRCC therapies. ### Competing Interest Statement The authors have declared no competing interest.
INTRODUCTION:Artificial intelligence (AI) is exhibiting tremendous potential to reduce the massive costs and long timescales of drug discovery. There are however important challenges currently limiting the impact and scope of AI models. AREAS COVERED:In this perspective, the authors discuss a range of data issues (bias, inconsistency, skewness, irrelevance, small size, high dimensionality), how they challenge AI models, and which issue-specific mitigations have been effective. Next, they point out the challenges faced by uncertainty quantification techniques aimed at enhancing and trusting the predictions from these AI models. They also discuss how conceptual errors, unrealistic benchmarks and performance misestimation can confound the evaluation of models and thus their development. Lastly, the authors explain how human bias, whether from AI experts or drug discovery experts, constitutes another challenge that can be alleviated by gaining more prospective experience. EXPERT OPINION:AI models are often developed to excel on retrospective benchmarks unlikely to anticipate their prospective performance. As a result, only a few of these models are ever reported to have prospective value (e.g. by discovering potent and innovative drug leads for a therapeutic target). The authors have discussed what can go wrong in practice with AI for drug discovery. The authors hope that this will help inform the decisions of editors, funders investors, and researchers working in this area.
We examine structural brain characteristics across three diagnostic categories: at risk for serious mental illness; first-presenting episode and recurrent major depressive disorder (MDD). We investigate whether the three diagnostic groups display a stepwise pattern of brain changes in the cortico-limbic regions. Integrated clinical and neuroimaging data from three large Canadian studies were pooled (total n = 622 participants, aged 12-66 years). Four clinical profiles were used in the classification of a clinical staging model: healthy comparison individuals with no history of depression (HC, n = 240), individuals at high risk for serious mental illness due to the presence of subclinical symptoms (SC, n = 80), first-episode depression (FD, n = 82), and participants with recurrent MDD in a current major depressive episode (RD, n = 220). Whole-brain volumetric measurements were extracted with FreeSurfer 7.1 and examined using three different types of analyses. Hippocampal volume decrease and cortico-limbic thinning were the most informative features for the RD vs HC comparisons. FD vs HC revealed that FD participants were characterized by a focal decrease in cortical thickness and global enlargement in amygdala volumes. Greater total amygdala volumes were significantly associated with earlier onset of illness in the FD but not the RD group. We did not confirm the construct validity of a tested clinical staging model, as a differential pattern of brain alterations was identified across the three diagnostic groups that did not parallel a stepwise clinical staging approach. The pathological processes during early stages of the illness may fundamentally differ from those that occur at later stages with clinical progression.
Background: Recent evidence suggests that patients with schizophrenia may show advanced brain ageing, particularly evident after the first year of onset. However, it is unclear if accelerated ageing relents, persists or continues to increase over time. The underlying causal factors are also poorly understood. Disruptions in glutamate function, oxidative stress, and inflammation may all contribute to progressive brain changes in people with schizophrenia. We examine whether brain ageing differs between early and established stages of schizophrenia, correlates with symptom severity and varies with markers of brain function, oxidative status and inflammatory burden. Methods: Two brain-age prediction models assessed 112 participants (34 recent onset psychosis, 36 established schizophrenia, 42 healthy controls). Brain age gap (BAG) was calculated by subtracting chronological age from predicted age. Shapley's additive explanations (SHAP) identified influential structural magnetic resonance imaging (MRI) features driving brain-age prediction. Linear regression models and partial correlations, adjusting for age, explored associations between BAG and neurometabolites, inflammatory markers, medication exposure and clinical scores in the whole sample. Results: The established schizophrenia group showed higher BAG (Mean = 6.21, SD = 7.30) compared to healthy individuals (Mean = -0.01, SD = 9.10), while recent-onset patients (Mean = 4.23, SD = 9.25) did not differ significantly from healthy individuals. The top 10 SHAP features diving the BAG included ventricular enlargement and total grey matter volume, which was similar in psychosis to healthy individuals. In a combined psychosis group (established + recent-onset), higher BAG correlated with more severe symptoms (PANSS total, general, and anxiodepressive subscales). BAG positively associated with Magnetic Resonance Spectroscopy measured glutathione and negatively with N-Acetyl Aspartate. Discussion: Accelerated brain age in schizophrenia may be related to illness severity and poor defence against oxidative stress. The lack of differences in SHAP features between schizophrenia and healthy individuals suggests that the pattern of brain ageing is in keeping with advanced normal ageing. Findings suggest potential treatment targets to improve brain health in schizophrenia, warranting further research. ### Competing Interest Statement LP reports personal fees for serving as chief editor from the Canadian Medical Association Journals, speaker/consultant fee from Janssen Canada and Otsuka Canada, SPMM Course Limited, UK, Canadian Psychiatric Association; book royalties from Oxford University Press; investigator-initiated educational grants from Janssen Canada, Sunovion and Otsuka Canada outside the submitted work. All other authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper. ### Funding Statement AM was supported by funding from the Medical Research Council (MRC) for doctoral training with RU, MK, and PL related to this manuscript (MR/2434208). RU acknowledges funding from MRC (MR/S037675/1) related to this manuscript. This research is also supported by the NIHR Oxford Health Biomedical Research Centre. The views expressed are those of the author(s) and not necessarily those of the NIHR or the Department of Health and Social Care. LP acknowledges research support from Monique H. Bourgeois Chair in Developmental Disorders and Graham Boeckh Foundation (Douglas Research Centre, McGill University) and salary award from the Fonds de recherche du Quebec-Sante ́ (FRQS). ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: The National Research Ethics Service Committe Northwest - Lancaster UK gave ethical approval for this work. Reference 14/NW/0298, Approval date 18/06/2014 I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All data produced in the present study are available upon reasonable request to the authors