Immunohistochemistry (IHC) remains the gold standard for evaluating protein expression in tumor microenvironment analysis. This approach hinders robust correlation analyses between spatial heterogeneity in the tumor microenvironment and clinical outcomes like disease-free survival (DFS). To address these challenges, we developed an automated pipeline for quantitative IHC feature extraction. Our method integrates deep learning-based tumor segmentation with computational detection of invasive margins at varying distances. Deconvolution algorithms quantify diaminobenzidine (DAB) staining intensity across the tumor body and the invasive margin. Spatial heterogeneous DAB density patterns were subsequently analyzed for DFS correlation. Using 104 patient samples (57 training/47 validation) stained for CD3, CD8, CD31, and HIF-1α, we identified 2 prognostic feature categories (CD3/CD8 aggregated positive areas within the 0.25-mm peripheral zone extending outward from the tumor-invasive front and HIF1-α-positive areas within a 0.75-mm peripheral zone extending outward from the tumor-invasive front). Immune-related features demonstrated C-indices of 0.726 (training) and 0.626 (validation), while hypoxia-associated markers showed C-indices of 0.714 and 0.656, respectively. Integration of these features with pTNM staging enhanced DFS stratification compared to pTNM staging alone, improving C-indices from 0.702 to 0.819 (training) and 0.668 to 0.853 (validation). This automated pipeline addresses critical limitations in traditional IHC analysis by enabling: 1) Objective quantification of spatial DAB heterogeneity. 2) Identification of biologically interpretable prognostic features. 3) Enhanced predictive performance over conventional staging systems. Our findings suggest this methodology could standardize IHC-based prognostic assessments and inform personalized treatment strategies. Further validation in multicenter cohorts is warranted to confirm clinical applicability.
Background: Reliable cancer-system data are needed for capacity planning, resource allocation and accountability. We assessed which patient- and department-level data elements professionals consider important for cancer-service planning and how feasible they are to collect across diverse settings.Methods: We conducted a cross-sectional online survey of cancer-system professionals and service leaders. Respondents rated the importance and local feasibility of 27 data elements linked to 14 recurring planning questions on parallel 5-point scales. Value-feasibility gaps were calculated as importance minus feasibility and summarised by domain and World Bank income group.Results: Among 82 valid respondents from 75 institutions in 48 countries, importance ratings were consistently high, while feasibility was lower and more variable. All 14 planning questions showed positive mean gaps. The largest concerned outcomes, toxicity, patient-reported outcomes, patient experience and treatment-related economic burden. In the small low-income subgroup, substantial gaps also affected baseline patient and tumour characteristics, diagnosis and treatment counts, and timelines. Equipment and workforce data were more closely aligned.Conclusions: Cancer-system data are widely valued for planning but are not consistently feasible to collect. The findings support a harmonised minimum indicator set with staged implementation adapted to local data maturity and resources.
The Digital Imaging and Communications in Medicine (DICOM) standard contains rich metadata describing patients, studies, series, and imaging instances, yet archiving systems typically store only a limited subset of these elements. For research purposes, integration of imaging metadata with clinical and experimental data is essential. To address this, we developed LinkedDicom, a lightweight Python package for converting DICOM metadata into Resource Description Framework (RDF) triples and querying using Semantic Web technologies. Building on lessons learned from our earlier Semantic DICOM (SeDI) toolkit, LinkedDicom was designed to be resource-efficient, and flexible to use in various scenarios. We evaluated the package on a publicly available head and neck cancer dataset and measured query performance time in various scenarios (using an SPARQL endpoint and per-patient file storage), demonstrating that on-disk storage with selective in-memory analysis provides a flexible alternative to resource-intensive live databases. Our findings show that LinkedDicom enables efficient, scalable, and FAIR-compliant representation of imaging metadata for research, with practical trade-offs between database-driven and on-demand querying.
Background and purposePredicting overall survival (OS) for inoperable locally advanced non-small cell lung cancer (LA-NSCLC) treated with immune checkpoint inhibitors remains challenging due to heterogeneous clinical response. Furthermore, the application of advanced deep learning is hindered by limited immunotherapy datasets. This study aimed to develop a novel prognostic framework by integrating voxel-level deep radiomics derived from pretreatment imaging with a knowledge transfer strategy to accurately predict OS.Materials and methodsA total of 526 patients were respectively identified. A non-immunotherapy dataset from the RTOG 0617 clinical trial was used to pre-train a Vision-Mamba deep learning model to learn tumor characteristics within manually delineated tumor regions. Voxel-level radiomics feature maps were generated within tumors and integrated with CT images for dual-input co-training. Using the same dual-input, a cross-dataset transfer learning strategy was then used to adapt the pre-trained models to the immunotherapy context by fine-tuning. The model’s performance was evaluated using the concordance index (C-index), time-dependent area under the receiver operating characteristic curve, Kaplan-Meier survival analysis, calibration curves, and decision curve analysis. Additionally, Gradient-weighted Class Activation Mapping (Grad-CAM) was employed to suggest a possible interpretation of the model’s decision logic.ResultsThe proposed model demonstrated robust generalization ability. In the independent immunotherapy testing dataset, the model achieved a C-index of 0.73 (95% CI:0.63-0.82). The time-dependent AUCs for predicting 1-year and 2-year OS were 0.73 and 0.70, respectively. Calibration curves showed good agreement between predicted and observed survival probability. Stratification analysis showed distinct survival differences, with the high-risk group exhibiting significantly poorer OS compared to low-risk group (P<0.001).ConclusionWe developed a voxel-level deep radiomics framework that bridges the data gap in immunotherapy research through fine-tuning on a limited immunotherapy dataset, and subsequent validation on an independent immunotherapy testing dataset, demonstrating robust generalizability.
Background and purpose Radiation-induced xerostomia remains a major toxicity after radiotherapy for head and neck cancer (HNC). Although QUANTEC-based parotid dose constraints are widely applied, increasing evidence suggests that additional salivary and oral structures may contribute to xerostomia risk. This systematic review aimed to summarize contemporary evidence on dose-volume predictors of post-radiotherapy xerostomia in HNC. Materials and methods This PRISMA-based systematic review was registered in PROSPERO (CRD42024519338). PubMed, Embase, Cochrane Library, and Web of Science were searched for studies published between January 2013 and August 2025. Studies evaluating dose-volume predictors of xerostomia in adult HNC patients treated with curative-intent radiotherapy were included. Results Fifty-one studies involving 10,789 patients were included. Most patients received intensity-modulated radiotherapy. Contralateral parotid gland mean dose (Dmean) was the most consistently validated predictor of xerostomia and was supported by moderate-certainty evidence. Multiple studies reported increased risk of moderate-to-severe xerostomia when contralateral parotid Dmean exceeded 24–26 Gy, whereas Dmean below 20 Gy was associated with better salivary preservation. Bilateral parotid Dmean around 26 Gy was also reported. However, support for specific numerical planning thresholds remained limited. Dosimetric parameters of the submandibular glands and oral cavity were also independently associated with xerostomia risk. Emerging evidence suggested potential roles of parotid substructures, including ducts and stem cell-rich regions. Conclusion Parotid gland mean dose remains the most robust predictor of post-radiotherapy xerostomia in HNC. Additional dosimetric information from the submandibular glands, oral cavity, and parotid substructures may further improve xerostomia risk stratification and support future NTCP modeling.
While considerable effort has been devoted to examining how variations in study protocols, acquisition settings, and annotations influence radiomics features, the role of feature selection (FS) methods largely remains unexplored. This study investigates a self-supervised deep sparse autoencoder ensemble (ensembleAE) and a novel Bayesian variant (bayesianAE) for radiomics FS in a small, class-imbalanced framework. Using a cohort of 100 prostate cancer patients under active surveillance, these models were benchmarked against eight classical FS methods. FS stability was assessed both globally and locally in a soft data-perturbation setting. While global stability measured consistency in overall feature ranking, local stability quantified the agreement in selecting top-ranked feature subsets. Among classical methods, wrappers were the least stable, whereas the filter-based Wilcoxon test (WLCX) exhibited high local stability and performance with moderate global stability. Conversely, bayesianAE demonstrated the highest global stability while matching WLCX in local stability and performance. Furthermore, although bayesianAE yielded local stability comparable to that of ensembleAE, it outperformed the latter in terms of global stability with a significantly lower computational burden (95 × faster and consumes 84% less memory). These findings position bayesianAE as a robust alternative to classical FS methods that could support the development of reproducible radiomics signatures.
Accurate estimation of Overall Survival (OS) in Non-Small Cell Lung Cancer (NSCLC) patients provides critical insights for treatment planning. While previous studies have shown that radiomics or Deep Learning (DL) features improved prediction accuracy, this study aimed to evaluate whether a model that integrates clinical, radiomics, DL, and dosimetric features outperforms other models developed with only a subset of these features. We collected pre-treatment lung CT scans and clinical data for 219 NSCLC patients from the Maastro Clinic: 183 for training and 36 for testing. Radiomics features were extracted using the Python radiomics feature extractor, and DL and dose features were obtained using a 3D ResNet model. An ensemble model comprising XGB and NN classifiers was developed using: (1) clinical features only; (2) clinical and radiomics features; (3) clinical and DL features; (4) clinical and dose features, and (5) clinical, radiomics, dose and DL features. The performance metrics were evaluated for the test and K-fold cross-validation data sets. The prediction model utilizing only clinical variables provided an Area Under the Receiver Operating Characteristic Curve (AUC) of 0.71 and a test accuracy of 72.73%. The best performance came from combining clinical, radiomics, dose and DL features (AUC: 0.84, accuracy: 88.64%). World Health Organisation Performance Status emerged as the factor with the highest importance for the combined model. Integrating radiomics, dose and DL features with clinical characteristics improved the prediction of OS after radiotherapy for NSCLC patients. The increased accuracy of our integrated model could enable personalized, risk-based treatment planning, guiding clinicians toward more effective interventions, improved patient outcomes and enhanced quality of life.
This study applied natural language processing to identify common topics in 12,054 Dutch patient-provider messages in inflammatory bowel disease. Using the BERTopic framework, three embedding models were evaluated with topic diversity, Cv coherence, and topic assignment proportion. A lightweight Dutch-specific embedding model (RobBERT) outperformed two larger multilingual models (MPNet Sentence Transformer and QWEN3-embedding-8B). We identified 120 sentence-level topics covering medical and administrative themes. Results showed that 22% of messages did not require specialist attention, highlighting the potential of neural topic models for automated triage in digital IBD care.
PURPOSE:Accurate prediction of symptomatic radiation pneumonitis (RP) is critical for radiation therapy, however, the generalization of deep learning models is hindered by restricted access to multicenter data. Although federated learning (FL) bypasses data sharing restrictions, standard FL algorithms underperform on highly heterogeneous clinical data across institutions. Therefore, this study aims to evaluate the clinical feasibility of a center-specific FL approach. METHODS AND MATERIALS:We evaluated the Federated Cross-Center Adaptive Alternating Model (FCAAM), a tailored framework designed to decouple globally transferable representations from the center-specific adaptations. The framework uses a dynamic weighting mechanism to handle data heterogeneity and uses differential privacy for enhanced security. The proposed FCAAM was evaluated for the prediction of RP using planning computed tomography and dose images on a diverse cohort of 1238 patients from 4 data sets representing real-world temporal and spatial shifts. Its performance was compared with single-center model, centralized model, and standard federated average model. RESULTS:FCAAM demonstrated improved cross-center performance and consistent robustness compared to baseline. It achieved a stable area under the curve across all 4 data set test sets (0.71-0.77), outperforming the single-center models (all area under the curves < 0.70) and federated averaging. FCAAM's performance was comparable to the centralized model and showed a relative improvement in sensitivity to small-sized data sets. Interpretability analysis confirmed that FCAAM learned clinically relevant features, and a web platform demonstrated the practical feasibility of applying FCAAM for multicenter collaboration. CONCLUSIONS:FCAAM provides a privacy-preserving, robust and interpretable solution for multicenter RP prediction. This center-specific strategy shows potential to enhance clinical decision-making and reduce cross-center performance gaps, supporting safer personalized radiation therapy.
While AI models are developed in oncology for predicting different clinical outcomes, the focus is often on accuracy and many fail to adequately communicate the degree of certainty in these predictions. To improve clinical decision-making in oncology, this work introduces the idea of uncertainty quantification (UQ) for AI models using an illustrative example. Our goal is to help radiologists and oncologists better understand prediction reliability by integrating UQ. Our illustrative example is a Radiomics Risk Model (RM) for Thymic Epithelial Tumours, developed to provide a basic understanding of the mechanism to evaluate the degree to which individual patient data matches the training set. The study demonstrates the concept of measuring uncertainty in artificial intelligence (AI) models using a simple example of distance measures within the feature space and example cases where uncertainty is addressed with probable causes. The paper highlights specifically where the clinicians may need more information to improve their confidence in their AI-driven assessments for clinical diagnostics.
Background: Advancements in artificial intelligence (AI) are transforming health care, particularly through AI-driven clinical decision support systems (AI-CDSS) that aid in predicting disease progression and personalizing treatment. Despite their potential, adoption remains limited due to clinician concerns about algorithm misuse, misinterpretation, and lack of transparency. Objective: This qualitative study explores the informational needs and preferences of clinicians to better understand and appropriately use AI-CDSS in decision-making. In parallel, this study explores AI experts' perspectives on what information should be communicated to enable safe and appropriate use of AI-CDSS. Methods: A qualitative description design study was conducted using semistructured interviews with 16 participants (8 clinicians and 8 AI experts). Discussions focused on experiences with AI, informational needs, and feedback on existing reporting standards, including Model Cards, Model Facts, and the Transparent Reporting of a multivariable prediction model for Individual Prognosis Or Diagnosis-Artificial Intelligence (TRIPOD-AI) checklist. The transcripts were analyzed through codebook thematic analysis. Results: Four key themes were identified: (1) clinicians need clear information on training data, its origin, size, and inclusion and exclusion criteria, to judge model applicability; (2) performance metrics must go beyond the area under the curve (AUC) and be clinically relevant to support informed decisions; (3) limitations and warnings about inappropriate use should be specific and clearly communicated to prevent misuse; and (4) information should be presented in layered, customizable formats within existing clinical software, avoiding unnecessary jargon, and allowing optional deeper explanations. While each of the reviewed reporting standards offered strengths, none were considered sufficient alone. Participants recommended a combined and clinician-centered approach to information delivery. Alignment of reporting standards with clinical workflows and decision thresholds was thought to be crucial to bridge the usability gap. Conclusions: To improve AI-CDSS adoption in clinical practice, reporting standards must be designed for better clinician comprehension and usability. Enhancing transparency, particularly regarding training data and performance, can likely help clinicians assess AI-CDSS more effectively. Information should be delivered in an accessible, layered format, fitting clinical workflows. Co-creation with clinicians throughout AI-CDSS development was a cross-cutting theme, highlighting its importance in ensuring tools are not only technically sound but also practically usable. Future research should explore how to structurally report on performance and validation metrics for clinician understanding and assess the impact of information provision on AI-CDSS adoption.
OBJECTIVES:Neoadjuvant chemoimmunotherapy (nCIT) is gradually becoming an important treatment strategy for patients with locally advanced oesophageal squamous cell carcinoma (LA-OSCC). This study aimed to predict the pathological complete response (pCR) of these patients using variational autoencoder (VAE)-based deep learning and radiomics technology. METHODS:A total of 253 LA-ESCC patients who were treated with nCIT and underwent enhanced CT at our hospital between July 2019 and July 2023 were included in the training cohort. VAE-based deep learning and radiomics were utilized to construct deep learning (DL) models and deep learning radiomics (DLR) models. The models were trained and validated via fivefold cross-validation among 253 patients. Forty patients were recruited from our institution between August 2023 and August 2024 as the test cohort. RESULTS:The AUCs of the DL and DLR models were 0.935 (95% CI: 0.786-0.992) and 0.949 (95% CI: 0.910-0.986) in the validation cohort and 0.839 (95% CI: 0.726-0.853) and 0.926 (95% CI: 0.886-0.934) in the test cohort. The performance gap between Precision and Recall of the DLR model was smaller than that of the DL model. The F1 scores of the DL and DLR models were 0.726 (95% confidence interval [CI]: 0.476-0.842) and 0.766 (95% CI: 0.625-0.842) in the validation cohort and 0.727 (95% CI: 0.645-0.811) and 0.836 (95% CI: 0.820-0.850) in the test cohort. CONCLUSIONS:We constructed a DLR model to predict pCR in nCIT-treated LA-ESCC patients, which demonstrated superior performance compared to the DL model. ADVANCES IN KNOWLEDGE:We innovatively used VAE-based deep learning and radiomics to construct the DLR model for predicting pCR of LA-ESCC after nCIT.
Background and purposeRadiation pneumonitis (RP) is one of the major dose-limiting toxicities of thoracic radiotherapy. Although multiple studies have attempted to predict RP, robust multicenter model development is often hindered by privacy regulations and data-transfer constraints, and many existing models are primarily derived from radiotherapy-alone populations, limiting applicability to contemporary regimens that incorporate immunotherapy. Therefore, this study aimed to develop an RP prediction model within a federated learning framework, incorporating sequential transfer learning strategies to enable separate risk assessment for radiotherapy patients with and without immunotherapy.MethodsMulticenter cohorts of lung cancer patients treated with definitive thoracic radiotherapy with or without immunotherapy were retrospectively collected and stratified by immunotherapy exposure. Radiomics features were extracted from whole-lung regions on pretreatment planning CT scans to construct RP prediction models. A federated learning framework was first applied to non-immunotherapy patients to learn common features of radiation pneumonitis without sharing raw data. The pretrained federated model was then sequentially transferred to immunotherapy treatment cohorts, with targeted fine-tuning to adapt to treatment specific RP patterns. Model performance was evaluated through internal validation and independent external validation, with SHAP analysis exploring feature importance differences across treatment settings.ResultsA total of 610 patients were included from five multicenter cohorts. Using patients without immunotherapy for model development, the federated baseline model showed stable discrimination in external validation across non-immunotherapy cohorts (AUC = 0.77). When this baseline model was directly applied to the immunotherapy cohort without adaptation, performance dropped markedly (AUC = 0.43). After fine-tuning on immunotherapy data, the immunotherapy-adapted model achieved improved performance within the immunotherapy cohort (AUC = 0.76) and remained robust in an independent external immunotherapy validation cohort (AUC = 0.75). Feature attribution analysis showed a shift in model coefficients between immunotherapy-treated and non-immunotherapy patients.ConclusionA federated modeling framework with treatment adaptation improves RP risk prediction across heterogeneous treatment settings under multicenter data constraints, particularly in immunotherapy-treated patients.
PURPOSE:Deep learning (DL) techniques may support localizing the epileptogenic zone (EZ) and improve surgical outcomes in drug-resistant epilepsy. This systematic review synthesizes current evidence on DL-assisted EZ localization from neuroimaging acquisitions, aiming to outline methodological trends, limitations, and future directions that bridge the gap between clinical translation and technological advances. METHODS:We systematically searched PubMed, Scopus, and Embase (via Ovid) on April 15, 2025, for studies applying DL to localize the EZ using neuroimaging data. The bias and applicability of studies was assessed using the PROBAST+AI tool. We extracted methodological details, as well as key performance metrics. RESULTS:Thirty-six studies met the eligibility criteria, most focusing on segmenting epileptogenic lesions using structural MRI. Focal cortical dysplasia was the most commonly targeted pathology, with fully convolutional networks being the predominant DL architecture. Approximately two-thirds of the studies showed high risk of bias and clinical applicability concerns, limited by non-representative cohorts and suboptimal evaluation methods. Five studies reported promising EZ detection rate in MRI-negative cases using large multi-center cohorts, yet progress in fine-grained localization tasks, such as lesion segmentation, remained moderate. CONCLUSION:This review highlights methodological limitations hindering the clinical translation of current DL approaches for EZ localization and provides a comprehensive set of recommendations to address them. Future work should prioritize developing standardized, clinically informative evaluation frameworks and explore research avenues aligned with modern DL practices, spanning from uncertainty quantification to large-scale vision foundation models and synthetic data generation.
Healthcare organisations are increasingly pursuing digital transformation (DT) but often struggle to achieve meaningful progress. Traditional IT operating models (ITOMs), defining how IT delivers value through structure, governance and technology, frequently lack the agility and maturity required for DT. Consultancy-led diagnostics and intervention plans also tend to fall short during execution and often overlook structural evaluations. This practitioner paper presents a scientifically grounded, practice-oriented approach to executing DT initiatives, based on a longitudinal case study at a leading Dutch radiotherapy clinic. In this case, the DT initiative focused on structural interventions within the ITOM, including its governance, sourcing, organisational structures and skills. The DT programme emphasised structural changes in day-to-day IT governance and decision-making process within the ITOM, while digital maturity, user satisfaction and IT cost indicators were used as outcome measures embedded in governance cycles to support continuous learning and prioritisation over time. A distinctive feature of this approach is its explicit integration of human factors through co-creation between IT and clinical leaders, multidisciplinary collaboration and adaptive governance. The methodology operationalises psychological and relational dynamics as early execution enablers rather than treating them as downstream effects of technical change. Ownership of digital initiatives gradually shifts from IT departments to care teams, while maintaining CIO accountability for coherence across the ITOM. These findings offer actionable guidance for healthcare leaders in similar contexts on how to integrate governance, ownership, capability development and measurement into day-to-day governance and prioritisation practices, thereby fostering a psychological shift towards engagement and shared responsibility in DT.
[This corrects the article DOI: 10.3389/fimmu.2026.1787518.].
Precision oncology relies on access to high-quality data for increasingly smaller patient subgroups. The international atomCAT consortium investigates the potential of federated learning to support this, using anal cancer as a rare cancer exemplar. Here, we show that federated multivariable Cox models trained across 14 centres (1428 patients) and externally validated in two additional centres (277 patients) achieve consistent calibration and discrimination during leave-one-centre-out and external validation (c-indices 0.68-0.79). Lower T stage, absence of nodal involvement, smaller tumour volume, female sex, younger age, and mitomycin- or cisplatin-based chemotherapy are associated with improved overall survival. Lower T stage, smaller tumour volume, and female sex are associated with improved locoregional control, while absence of nodal involvement and smaller tumour volume are associated with better freedom from distant metastases. These findings demonstrate that federated learning enables robust, privacy-preserving prognostic modelling for rare cancers using real-world data, supporting international collaboration without data sharing.
Tumor-intrinsic biomarkers alone insufficiently predict pathological complete response (pCR) to neoadjuvant immunochemotherapy (NICT) in non-small cell lung cancer (NSCLC). Artificial intelligence (AI)-based three-dimensional CT-derived body composition may provide complementary predictive value. We evaluated its association with pCR following NICT in NSCLC. This multicenter retrospective study of NSCLC patients treated with NICT in China between July 2019 and July 2024. Pre- and post-treatment CT scans were used for automated T1-T12 localization and volumetric body composition segmentation. Metrics included skeletal muscle, intermuscular, visceral, and subcutaneous adipose volume index (SAVI), and their percentage changes between scans. Among 657 patients (mean age, 61.3 years; 87.4 % men), pCR rates were 39.7 % (training), 38.4 % (internal validation), and 34.9 % (external validation). In multivariable analysis, high baseline skeletal muscle volume index (SMVI) was independently associated with pCR (OR = 2.22). During NICT, each 1 % relative increase in SMVI was associated with a 16 % higher likelihood of pCR (OR = 1.16), whereas every 10 % relative increase in SAVI improved pCR probability (OR = 1.56). A machine learning model integrating clinical variables, baseline SMVI, %ΔSMVI, and %ΔSAVI demonstrated significantly better discrimination than models using clinical variables alone (p < 0.05) in all cohorts. The performance was observed in the internal and external validation cohorts, with sensitivities of 62.1 % and 52.8 %, and specificities of 66.7 % and 74.7 %, respectively. AI-based CT-derived body composition quantification, particularly baseline SMVI and dynamic changes in SMVI and SAVI during NICT, are independently associated with pCR in NSCLC. Incorporating these modifiable biomarkers into predictive models improves performance beyond clinical variables alone.