Dynamic contrast-enhanced magnetic resonance imaging (DCE-MRI) is central to breast cancer imaging, but gadolinium administration increases scan burden and motivates contrast-reduced alternatives, including synthetic contrast generation. We propose a latent bridge matching (LBM) framework for synthesizing peak-enhanced breast DCE-MRI from pre-contrast images in the MAMA-SYNTH challenge setting. Instead of starting from Gaussian noise as in conventional latent diffusion models (LDMs), the proposed model learns a conditional bridge between paired pre-contrast and peak-enhanced VAE latents. A latent UNet predicts the remaining correction from intermediate bridge states to the peak-enhanced latent, enabling iterative refinement while keeping the trajectory anchored to patient-specific anatomy. We evaluated two LBM conditioning variants on 91 DUKE validation cases. For the tumor-conditioned variant, tumor masks were used as conditioning inputs. Tumor-conditioning improved performance compared with pre-contrast conditioning, reducing MSE from 1.023 to 0.940 and FRD from 7.523 to 4.716, while increasing tumor SSIM from 0.355 to 0.429. The tumor-conditioned LBM also outperformed the evaluated LDM baseline on this validation cohort. These results suggest that latent bridge matching is a promising pre-contrast-anchored formulation for virtual contrast enhancement, while further work is needed to validate generalization and remove dependence on ground-truth tumor masks at inference.
The hallmarks of cancer were introduced by Hanahan and Weinberg as a conceptual organizing framework to distil the complexity of tumours. This concept of cancer hallmarks has become an enduring theme in cancer research. Moreover, an increasing number of therapeutic strategies are being aimed at targeting these hallmarks. However, translating them into the clinic requires technologies to monitor their effectiveness and biomarkers that can stratify patients for the choice of specific therapies. Tumour heterogeneity and the ability of tumour cells to rapidly mutate and develop evasion strategies makes the development of non-invasive imaging capabilities to interrogate these hallmarks as biomarkers and monitor them longitudinally and quantitatively particularly important. This Review presents a holistic discussion of non-invasive diagnostic imaging capabilities related to the hallmarks of cancer; some hallmarks can be assessed with imaging probes that directly target biomolecules, whereas others can be interrogated indirectly by imaging pathophysiological processes. Additionally, visualizing the hallmarks of cancer can be addressed with artificial intelligence-assisted, multiparametric image analysis (for example, radiomics, radiogenomics and deep learning). The approaches discussed have been evaluated in a translational context, and some of them already have a substantial role in clinical practice, for example, to guide treatment strategies, including surgical resections, radiotherapy and molecularly targeted chemo-, immuno- and radiopharmaceutical therapies.
Homologous recombination deficiency (HRD) can lead to genomic instability, increased cancer susceptibility, and enhanced sensitivity to DNA-targeting therapies. Although radiomics has been used for various medical applications, its application in animal studies remains largely unexplored, primarily due to the typically limited availability of preclinical data. In this study, we applied a state-of-the-art foundation model (FM) on preclinical computed tomography (CT) images in mice, aiming to: (i) distinguish HRD status within isogenic xenografts, and (ii) predict differential therapeutic responses of CP-506, a novel hypoxia-activated DNA-crosslinking agent. The dataset comprises micro-CT scans of 307 mice with balanced HRD status, collected both before and after CP-506 or control treatment. The FM demonstrated robust HRD classification performance, achieving an AUC of 0.88 on the test set, which significantly outperformed the handcrafted radiomics and supervised deep learning (sDL). The highest AUC (0.93) was achieved in the consensus subgroup (71%) between sDL and FM. Additionally, HRD-related features predicted DNA damage and growth delay following the treatment. Interpretability analysis indicated the important role of texture heterogeneity in HRD classification. Therefore, these results suggest that FM successfully overcomes the data scarcity in animal studies and enables HRD classification and treatment response prediction from preclinical CT imaging.
Recent advancements in artificial intelligence (AI) and the vast data generated by modern clinical systems have driven the development of AI solutions in medical imaging, encompassing image reconstruction, segmentation, diagnosis, and treatment planning. Despite these successes and potential, many stakeholders worry about the risks and ethical implications of imaging AI, viewing it as complex, opaque, and challenging to understand, use, and trust in critical clinical applications. The FUTURE-AI guideline for trustworthy AI in healthcare was established based on six guiding principles: Fairness, Universality, Traceability, Usability, Robustness, and Explainability. Through international consensus, a set of recommendations was defined, covering the entire lifecycle of medical AI tools, from design, development, and validation to regulation, deployment, and monitoring. In this paper, we describe how these specific recommendations can be instantiated in the domain of medical imaging, providing an overview of current best practices along with guidelines and concrete metrics on how those recommendations could be met, offering a valuable resource to the international medical imaging community.
Accurate MRI-based identification of extramural vascular invasion (EVI) and mesorectal fascia invasion (MFI) is crucial for risk-stratified rectal cancer treatment. However, subjective visual assessment and inter-institutional variability limit diagnostic consistency. This study developed and evaluated a multi-center, foundation model-driven framework that automatically classifies EVI and MFI on axial and sagittal MRI. A total of 331 pre-treatment rectal cancer T2-weighted MRI scans from three European hospitals were retrospectively recruited. A self-supervised frequency domain harmonization strategy was applied to reduce scanner variability. Three classifiers, SeResNet, the universal biomedical pretrained model (UMedPT) with a multilayer perceptron head, and a logistic-regression variant using frozen UMedPT features (UMedPT_LR), were trained (n = 265) and tested (n = 66). Gradient-weighted class activation mapping (Grad-CAM) visualized model predictions. UMedPT_LR achieved the best EVI performance with multiplanar fusion (AUC = 0.82, test set). For MFI, UMedPT trained on axial harmonized images yielded the highest performance (AUC = 0.77). Both tasks outperformed the CHAIMELEON 2024 benchmark (EVI: 0.82 vs 0.74; MFI: 0.77 vs 0.75). Harmonization enhanced MFI classification, and multiplanar fusion further boosted EVI performance. Grad-CAM confirmed biologically plausible attention on peritumoral regions (EVI) and mesorectal fascia margins (MFI). The proposed foundation model-driven framework, leveraging frequency domain harmonization and multiplanar fusion, achieves state-of-the-art performance for automated EVI and MFI classification on MRI, demonstrating strong generalizability across multiple centers. Addressing inter-center inconsistencies in rectal cancer MRI, a multiplanar foundation model with cross-scanner harmonization significantly improves the detection of EVI and MFI, potentially standardizing staging and guiding therapy.
Colorectal liver metastases (CRLM) are a major cause of cancer-related mortality, and reliable detection on CT remains challenging in multi-centre settings. We developed a foundation model-based AI pipeline for patient-level classification and lesion-level detection of CRLM on contrast-enhanced CT, integrating uncertainty quantification and explainability. CT data from the EuCanImage consortium (n=2437) and an external TCIA cohort (n=197) were used. Among several pretrained models, UMedPT achieved the best performance and was fine-tuned with an MLP head for classification and an FCOS-based head for lesion detection. The classification model achieved an AUC of 0.90 and a sensitivity of 0.82 on the combined test set, with a sensitivity of 0.85 on the external cohort. Excluding the most uncertain 20 percent of cases improved AUC to 0.91 and balanced accuracy to 0.86. Decision curve analysis showed clinical benefit for threshold probabilities between 0.30 and 0.40. The detection model identified 69.1 percent of lesions overall, increasing from 30 percent to 98 percent across lesion size quartiles. Grad-CAM highlighted lesion-corresponding regions in high-confidence cases. These results demonstrate that foundation model-based pipelines can support robust and interpretable CRLM detection and classification across heterogeneous CT data.
The novel hypoxia-activated prodrug CP-506 selectively targets the hypoxic, treatment-resistant tumor microenvironment. Given the alkylating effector metabolites of CP-506, we hypothesized that defects in interstrand crosslink (ICL) and double-strand break repair influence treatment efficacy. In vitro and in vivo isogenic cancer models proficient or deficient in the Fanconi anemia (FA), homologous recombination (HR), or non-homologous end joining (NHEJ) pathway were used to assess CP-506-induced cytotoxicity and DNA damage. Viability and clonogenic assays demonstrated enhanced sensitivity to CP-506 in FA- or HR-deficient cells compared to parental cells, which was confirmed by spheroid growth inhibition studies. In vivo, CP-506 caused greater enhancement ratios in FA- and HR-deficient xenografts versus parental controls (p < 0.0001) but not in NHEJ-deficient xenografts (p = 0.18). Mechanistically, CP-506 increased γH2AX expression (1.9- to 9.3-fold) in FA- and HR-deficient cells and xenografts, whereas NHEJ-deficient models showed a 0.5-fold reduction. Alkaline comet assays confirmed CP-506-induced ICLs and DNA strand breaks but did not explain the differential therapeutic responses among isogenic cancer cells. These data indicate that deficiencies within FA or HR, but not NHEJ or nucleotide excision repair (NER), determine CP-506 sensitivity, consistent with a synthetic-lethal interaction. Therefore, tumor hypoxia and DNA repair status are key biomarkers for stratifying patients in CP-506 clinical trials.
Abstract Colorectal liver metastases (CRLM) remain a major cause of cancer-related mortality. Early, reliable detection on CT imaging is essential for curative treatment planning, yet the diagnostic performance of AI models often declines across scanners and institutions, limiting clinical generalizability. In this study, we developed and evaluated a multi-center foundation-model-based AI pipeline for patient-level classification and lesion-level detection of CRLM on contrast-enhanced CT. Using data from the EuCanImage consortium ( n = 2437) and TCIA_CRLM ( n = 197, all CRLM), we benchmarked several pretrained foundation models and identified UMedPT as the optimal encoder. The final model achieved an AUC of 0.89 and a sensitivity of 0.82 on the EuCanImage test set, with a sensitivity of 0.85 on the external TCIA cohort. Excluding the most uncertain 20% of cases improved AUC to 0.90 and balanced accuracy to 0.85. Decision-curve analysis indicated superior net benefit over “treat-all” and “treat-none” strategies for threshold probabilities between 0.35 and 0.75. The lesion-level detector identified 69.1% of lesions overall, increasing substantially with lesion size. Grad-CAM maps showed strong correspondence between attention regions and metastases in high-confidence predictions. These findings demonstrate that foundation-model-based pipelines enable robust, generalizable, and interpretable solutions for CRLM detection and classification in multi-center CT imaging.
PURPOSE:Radiomics allows for the quantification of medical images and facilitates precision medicine. Many radiomic features derived from computed tomography (CT) are sensitive to variations across scanners, reconstruction settings, and acquisition protocols. In this phantom study, eight different CT reconstruction parameters were varied to explore image- and feature-level harmonization approaches to improve tissue classification. METHODS:Varying reconstructions of an anthropomorphic radiopaque phantom containing three lesion categories (metastasis, hemangioma, and benign cyst) and normal liver tissue were used for evaluating two harmonization methods and their combination: (i) generative adversarial networks (GANs) at the image level; (ii) ComBat at the feature level, and (iii) a combination of (i) and (ii). A total of 93 texture and intensity features were extracted from each tissue class before and after image-level harmonization and were also harmonized at the feature level. Reproducibility and stability were assessed via the Concordance Correlation Coefficient (CCC) and pairwise comparisons using paired stability tests. The ability of features to discriminate between tissue classes was assessed by measuring the area under the receiver operating characteristic curve. The global reproducibility and discriminative power were assessed by averaging over the entire dataset and across all tissue types. RESULTS:ComBat improved reproducibility by 31.58% and stability by 5.24%, while GAN increased reproducibility by 8% it reduced stability by 4.33%. Classification analysis revealed that ComBat increased average AUC by 15.19%, whereas GAN decreased AUC by 2.56%. CONCLUSION:While GAN qualitatively enhances image harmonization, ComBat provides superior statistical improvements in feature stability and classification performance, highlighting the importance of robust feature-level harmonization in radiomics.
Text to image latent diffusion models have recently advanced medical image synthesis, but applications to 3D CT generation remain limited. Existing approaches rely on simplified prompts, neglecting the rich semantic detail in full radiology reports, which reduces text image alignment and clinical fidelity. We propose Report2CT, a radiology report conditional latent diffusion framework for synthesizing 3D chest CT volumes directly from free text radiology reports, incorporating both findings and impression sections using multiple text encoder. Report2CT integrates three pretrained medical text encoders (BiomedVLP CXR BERT, MedEmbed, and ClinicalBERT) to capture nuanced clinical context. Radiology reports and voxel spacing information condition a 3D latent diffusion model trained on 20000 CT volumes from the CT RATE dataset. Model performance was evaluated using Frechet Inception Distance (FID) for real synthetic distributional similarity and CLIP based metrics for semantic alignment, with additional qualitative and quantitative comparisons against GenerateCT model. Report2CT generated anatomically consistent CT volumes with excellent visual quality and text image alignment. Multi encoder conditioning improved CLIP scores, indicating stronger preservation of fine grained clinical details in the free text radiology reports. Classifier free guidance further enhanced alignment with only a minor trade off in FID. We ranked first in the VLM3D Challenge at MICCAI 2025 on Text Conditional CT Generation and achieved state of the art performance across all evaluation metrics. By leveraging complete radiology reports and multi encoder text conditioning, Report2CT advances 3D CT synthesis, producing clinically faithful and high quality synthetic data.
Objectives: Accurate differentiation between usual interstitial pneumonia (UIP) and nonspecific interstitial pneumonia (NSIP) is crucial for guiding treatment in interstitial lung diseases (ILDs). This study evaluates the efficacy of clinical, radiomic, and combined models in classifying UIP and NSIP using high-resolution computed tomography (HRCT) scans. Materials and Methods: A retrospective analysis was performed on 105 HRCT scans (UIP = 60, NSIP = 45) from Faisal Hospital and Research Center. Demographic and pulmonary function data formed the clinical model. Radiomic features, extracted using the pyRadiomics package, were refined using recursive feature elimination. A combined model was developed by integrating clinical and radiomic features to assess their complementary diagnostic value. Model performance was assessed via the area under the receiver operating characteristic curve (AUC). SHapley Additive exPlanations (SHAP) analysis, including both global feature importance and individual-level explanations, was used to interpret the model predictions. Results: The clinical model achieved an AUC of 0.62 with a sensitivity of 54% and a specificity of 78%. The radiomic model outperformed it with an AUC of 0.90 with a sensitivity and specificity above 85%. The combined model showed an AUC of 0.86 with a sensitivity of 88% and a specificity of 78%. SHAP analysis identified texture-based features, such as GLCM_Idmn and NGTDM_Contrast, as influential for classification. Conclusions: Radiomic features enhance classification accuracy for UIP and NSIP compared to clinical models. Integrating HCR into clinical workflows may reduce variability and improve diagnostic accuracy in ILD. Future studies should validate findings using larger, multicenter datasets.
Purpose:The relative biological effectiveness (RBE) of tumor control for proton beam therapy (PBT) compared to photon radiotherapy (RT) is typically assumed to be independent of fractionation. To test this, we modeled published PBT outcome results for early-stage non-small cell lung cancer (NSCLC) treatments across a range of fractionation schedules. Materials and Methods:All published and analyzable cohorts were included (399 patients, 413 treated lesions). Two models were used to fit the data: a previously published tumor simulation model that fits photon RT results of NSCLC across all fractionation regimes and the Fowler LQ model with a kick-off time term. The treatment effect of each cohort was referenced to the photon equivalent dose through mechanistic model simulations in a 2 Gy/weekday scenario, with radiobiological parameters determined to simultaneously best-fit all fractionation results. The tumor control RBE of each published treatment schedule, compared to the modeled photon RT effect of the same schedule, was then estimated. Results:For cohorts whose treatments lasted less than three weeks (i.e., 12 fractions or less), the RBE of PBT was in the range of 1.08 to 1.11. However, for fractionated treatments stretching over four weeks or more (20-25 fractions), the relative effectiveness was much lower, with RBEs in the range of 0.82-0.89. This conclusion was unchanged using the simpler Fowler LQ + time model. Conclusions:The proton RBE for hypo-fractionated schedules was 20-30% higher than for conventional schedules. The derived radiobiological parameters of PBT differ significantly from those of photon RT, indicating that PBT is influenced differentially by radiobiological mechanisms which require further investigation.
Purpose: Predictive models for contrast-enhanced mammography often perform better at detecting and classifying enhancing masses than (non-enhancing) microcalcification clusters. We aim to investigate whether incorporating synthetic data with simulated microcalcification clusters during training can enhance model performance. Approach: Microcalcification clusters were simulated in low-energy images of lesion-free breasts from 782 patients, considering local texture features. Enhancement was simulated in the corresponding recombined images. A deep learning (DL) model for lesion detection and classification was trained with varying ratios of synthetic and real (850 patients) data. In addition, a handcrafted radiomics classifier was trained using delineations and class labels from real data, and predictions from both models were ensembled. Validation was performed on internal (212 patients) and external (279 patients) real datasets. Results: The DL model trained exclusively with synthetic data detected over 60% of malignant lesions. Adding synthetic data to smaller real training sets improved detection sensitivity for malignant lesions but decreased precision. Performance plateaued at a detection sensitivity of 0.80. The ensembled DL and radiomics models performed worse than the standalone DL model, decreasing the area under this receiver operating characteristic curve from 0.75 to 0.60 on the external validation set, likely due to falsely detected suspicious regions of interest. Conclusions: Synthetic data can enhance DL model performance, provided model setup and data distribution are optimized. The possibility to detect malignant lesions without real data present in the training set confirms the utility of synthetic data. It can serve as a helpful tool, especially when real data are scarce, and it is most effective when complementing real data.
PURPOSE:Hepatocellular carcinoma (HCC) remains a global health concern, marked by increasing incidence rates and poor outcomes. This study seeks to develop a robust predictive model by integrating radiomics and deep learning features with clinical data to predict 2-year survival in HCC patients treated with stereotactic body radiation therapy (SBRT). METHODS:This study analyzed a cohort of 186 HCC patients who underwent SBRT. Radiomics features were extracted from CT scans, complemented by collection of clinical data. Training and validation of machine learning models were conducted using nested cross-validation techniques. Deep learning models, leveraging various convolutional neural networks (CNNs), were employed to effectively integrate both image and clinical data. Post-hoc explainability techniques were applied to elucidate the contribution of imaging data to predictive outcomes. RESULTS:Handcrafted radiomics features demonstrated moderate predictive performance, with area under the receiver operating characteristic curve (AUC) values ranging from 0.59 to 0.72. Deep learning models, harnessing the fusion of image and clinical data, exhibited improved predictive accuracy, with AUC values ranging from 0.71 to 0.81. Notably, the ensemble model, amalgamating handcrafted radiomics and deep learning features with clinical data, demonstrated the most robust predictive capability, achieving an AUC of 0.86 (95% CI: 0.80-0.93). CONCLUSION:The ensemble model represents a significant advancement, providing a comprehensive tool for predicting survival outcomes in HCC patients undergoing SBRT. The inclusion of interpretability methods such as Grad-CAM enhances transparency and understanding of these complex predictive models.
BackgroundMultiple sclerosis (MS) is an autoimmune disease of the central nervous system, leading to varying degrees of functional impairment. Conventional tools, such as the Expanded Disability Status Scale (EDSS), lack sensitivity to subtle disease worsening. Radiomics provides a quantitative imaging approach to address this limitation. This study applied machine learning (ML) and radiomics features from T2-weighted Fluid-Attenuated Inversion Recovery (FLAIR) magnetic resonance imaging (MRI) to predict disability worsening in MS.MethodsA retrospective analysis was performed on real-world data from 247 PwMS across two centers. Disability worsening was defined as a change in EDSS over two years. FLAIR MRIs underwent preprocessing and super-resolution reconstruction to enhance low-resolution images. White matter lesions (WML) were segmented using the Lesion Segmentation Toolbox (LST), and tissue segmentation was performed using sequence Adaptive Multimodal Segmentation. Radiomics features from WML and normal-appearing white matter (NAWM) were extracted using Pyradiomics, harmonized with Longitudinal ComBat, followed by recursive feature elimination for feature selection. Elastic Net, Balanced Random Forest (BRFC), and Light Gradient-Boosting Machine (LGBM) models were trained and evaluated.ResultsThe LGBM model with harmonized radiomics and clinical features outperformed the clinical-only model, achieving a test area under the precision-recall curve (PR AUC) of 0.20 and a receiver operating characteristic area under the curve (ROC AUC) of 0.64. Key predictive features, among others, included Gray-Level Co-Occurrence Matrix (GLCM) maximum probability (WML) and Gray-Level Dependence Matrix (GLDM) dependence non-uniformity (NAWM). However, short-term longitudinal changes showed limited predictive power (PR AUC = 0.11, ROC AUC = 0.69).ConclusionThese findings highlight the potential of ML-driven radiomics in predicting disability worsening, warranting validation in larger, balanced datasets and exploration of advanced deep learning approaches.
Radiomics is a tool for medical imaging analysis that could have a relevant role in precision oncology by offering precise quantitative support for clinical decision-making. The Radiomics Quality Score (RQS) is a tool developed to assess the rigour of radiomics studies that has now been widely adopted by researchers. Although RQS version 1.0 established a benchmark, an updated framework is required to account for evolving knowledge and ensure optimal evaluation of the quality of radiomics studies through the inclusion of fairness, explainability, rigorous quality control and harmonization. In this Review, we introduce the updated RQS 2.0, which maintains the scientific rigour of its predecessor and addresses these contemporary needs, and therefore could potentially accelerate clinical translation. Moreover, we introduce the radiomics readiness levels, inspired by the technology readiness level framework, which are integrated in RQS 2.0 and reflect nine distinct levels of incremental improvement in radiomics research with the ultimate aim of clinical implementation. We also detail anticipated future directions in radiomics, outlining a strategic vision to advance precision oncology, which is the ultimate aim of RQS 2.0. The Radiomics Quality Score (RQS) was developed to assess the rigour of studies using radiomics, a tool for medical imaging analysis. The RQS has been widely used in the field and now needs an update (RQS 2.0) to address contemporary needs. The authors of this Review introduce RQS 2.0, which integrates radiomics readiness levels to provide a structured framework towards clinical implementation.
BackgroundMammographic imaging is essential for breast cancer detection and diagnosis. In addition to masses, calcifications are of concern and the early detection of breast cancer also heavily relies on the correct interpretation of suspicious microcalcification clusters. Even with advances in imaging and the introduction of novel techniques such as digital breast tomosynthesis and contrast-enhanced mammography, a correct interpretation can still be challenging given the subtle nature and large variety of calcifications.PurposeComputer simulated lesion models can serve to develop, optimize, or improve imaging techniques. In addition to their use in comparative (virtual clinical trial) detection experiments, these models have potential application in training deep learning models and in the understanding and interpretation of breast lesions. Existing simulation methods, however, often lack the capacity to model the diversity occurring in breast lesions or to generate models relevant for a specific case. This study focuses on clusters of microcalcifications and introduces an automated, flexible toolbox designed to generate microcalcification cluster models customized to specific tasks.MethodsThe toolbox allows users to control a large number of simulation parameters related to model characteristics such as lesion size, calcification shape, or number of microcalcifications per cluster. This leads to the capability of creating models that range from regular to complex clusters. Based on the input parameters, which are either tuned manually or pre-set for a specific clinical type, different sets of models can be simulated depending on the use case. Two lesion generation methods are described. The first method generates three-dimensional microcalcification clusters models based on geometrical shapes and transformations. The second method creates two-dimensional (2D) microcalcification cluster models for a specific 2D mammographic image. This novel method employs radiomics analysis to account for local textures, ensuring the simulated microcalcification cluster is appropriately integrated within the existing breast tissue. The toolbox is implemented in the Python language and can be conveniently run through a Jupyter Notebook interface, openly accessible at . Validation studies performed by radiologists assessed the level of malignancy and realism of clusters tuned with specific parameters and inserted in mammographic images.ResultsThe flexibility of the toolbox with multiple simulation methods is illustrated, as well as the compatibility with different simulation frameworks and image types. The automation allows for the straightforward and fast generation of diverse microcalcification cluster models. The generated models are most likely applicable for various tasks as they can be configured in a variety of ways and inserted in different types of mammographic images of multiple acquisition systems. Validation studies confirmed the capacity to simulate realistic clusters and capture clinical properties when tuned with appropriate parameter settings.ConclusionThis simulation toolbox offers a flexible means of simulating microcalcification cluster models with potential use in both technical and clinical research in mammography imaging. The 3D generation methods allow for specifying many characteristics regarding the calcification shape and cluster architecture, and the 2D generation method presents a novel manner to create microcalcification clusters tailored to existing breast textures.
Purpose: This study evaluates the impact of harmonization and multi-region feature integration on survival prediction in non-small cell lung cancer (NSCLC) patients. We assess the prognostic utility of handcrafted radiomics and pretrained deep features from thoracic CT images, integrating them with clinical data using a multicentre dataset. Methods: Survival models were built using handcrafted radiomic and deep features from lung, tumor, mediastinal nodes, coronary arteries, and coronary artery calcium (CAC) scores from 876 patients across five centres. CT features were harmonized using ComBat, reconstruction kernel normalization (RKN), and RKN-ComBat. Models were constructed at the region of interest (ROI) level and through ensemble strategies. Regularized Cox models estimated overall survival, with performance assessed via the concordance index (C-index), 5-year time-dependent area under the curve (t-AUC), and hazard ratios. SHAP values interpreted feature contributions, while consensus analysis categorized predicted survival probabilities at fixed time points. Results: TNM staging showed prognostic value (C-index = 0.67; hazard ratio = 2.70; t-AUC = 0.85). The clinical and tumor texture radiomics model with ComBat yielded high performance (C-index = 0.76; t-AUC = 0.88). FM deep features from 50 voxel cubes also showed predictive value (C-index = 0.76; t-AUC = 0.89). An ensemble model combining tumor, lung, mediastinal node, CAC, and FM features achieved a C-index of 0.71 and t-AUC of 0.79. Consensus analysis identified a high-confidence patient subset, resulting in a model with a 5-year t-AUC of 0.92, sensitivity of 96.8 Conclusion: Harmonization and multi-region feature integration enhance survival prediction in NSCLC patients using CT imaging, supporting individualized risk stratification in multicentre settings.
Background:Making informed decisions about clinical trial participation can be overwhelming for patients due to the complexity of trial information, potential risks and benefits, and the emotional burden of a recent diagnosis. Patient decision aids (PDAs) simplify this process by providing clear information on treatment options, empowering patients to actively participate in shared decision-making with their doctors. While PDAs have shown promise in various health care contexts, their use in clinical trials, particularly in the form of trial-specific patient decision aids (tPDAs), remains underused. Objective:This study aims to address the challenge of patient comprehension of traditional clinical trial materials. We developed a freely accessible, user-friendly tPDA within the context of the ImmunoSABR phase 2 trial. The tPDA aimed to enhance informed decision-making regarding trial participation. The primary endpoint was usability, quantitatively measured by the System Usability Scale (SUS). Secondary endpoints included time spent on the tPDA, patient satisfaction ratings, and participants' self-reported level of understanding of the trial. Methods:We developed the tPDA following the International Patient Decision Aid Standards and validated it through a structured, 3-phase iterative evaluation process. An initial evaluation was performed with 17 computer scientists who had expertise in biomedical applications, ensuring technical robustness. The content and usability were further refined through evaluations involving 10 clinicians and 8 medical students, focusing on clinical accuracy and user-friendliness. Finally, the tool was tested by 6 patients eligible for the ImmunoSABR trial to assess real-world applicability and patient-centered design. Results:Evaluations demonstrated the tPDA's effectiveness in enhancing informed decision-making, directly addressing our primary end point of usability with an overall mean SUS score of 79.4 (SD 15.9), indicative of good usability. Addressing our secondary endpoints, patients completed the tPDA efficiently, with the majority (4/6) finishing in under 30 minutes, and all but 1 within 60 minutes. Qualitative feedback highlighted significant improvements in patients' understanding of the trial details, reinforcing the tPDA's role in facilitating better patient engagement and comprehension. Conclusions:Our study demonstrates the feasibility and potential of tPDAs to enhance patient comprehension and engagement in clinical trials. Integrating tPDAs offers a valuable addition to traditional paper-based and verbal communication methods, promoting informed decision-making and patient-centered care.