While federated learning (FL) enables collaborative medical image segmentation without centralizing sensitive data, real-world deployment is frequently complicated by cross-site label imperfections such as contour disagreement, missing or additional structures, and confused labels. Federated noisy label learning (FNLL) aims to mitigate these effects, yet remains underused in practice as existing evidence is largely based on synthetic noise, simplified settings, and limited real-world noisy evaluation. We address this gap by introducing a benchmark suite that combines diverse real-world noisy datasets, deployment-relevant client-noise scenarios, and label-noise-targeted evaluation to support systematic FNLL assessment and informed method selection. The suite combines curated real-world noisy medical image segmentation datasets from diverse sources with a comprehensive federated segmentation framework including various client-noise scenarios and noise-targeted evaluation. To demonstrate its capabilities, we compare representative FNLL methods across approaches, including noise-aware aggregation, robust personalization, label correction, and sample selection. In-depth data analysis shows that real-world segmentation label noise occurs both in isolation and in combination of characterized noise types. The benchmark shows that FedSelect performs strongest within noise-targeted evaluation, underlines FedAvg as a competitive baseline, and provides a practically informed decision guide for FNLL method selection based on label-noise type, client-noise scenario, and technical-cost considerations. The presented suite provides a realistic and discriminative basis for FNLL evaluation in medical image segmentation and establishes a reusable foundation for fair benchmarking, dataset-specific label-noise characterization, and future method development under realistic federated settings. Code is available at https://github.com/MIC-DKFZ/FedSegNoiseBench .
PURPOSE:Response assessment after radiotherapy (RT) of gliomas remains challenging due to radiation-induced reactions that mimic tumor growth (pseudoprogression, PsPD). Unlike standard anatomical MRI, Chemical Exchange Saturation Transfer (CEST) MRI offers molecular image contrasts that could differentiate tumor growth from PsPD. Our aim was to quantify the influence of radiation dose on CEST contrasts in both healthy-appearing brain tissues and tumor tissues. METHODS:This prospective study enrolled 33 glioma patients (26 glioblastoma, 7 IDH-mutant glioma). In total, 81 longitudinal CEST MRI scans were performed before RT until seven months thereafter. CEST MRI included asymmetry analysis of the amide proton transfer-weighted (APTw) signal, and a multipool Lorentzian fitting approach was used to quantify the amide signal, ssMT signal and rNOE signal. CEST contrast changes were analyzed in normal-appearing brain tissues and tumor by cumulative histograms and based on Bayesian linear multilevel models. RESULTS:Normal-appearing brain tissues did not exhibit significant changes in any CEST contrasts after high-dose radiation. Conversely, the CEST contrasts changed in the tumor region of glioma patients according to their response to treatment. Particularly the APTw-signal demonstrated moderate opposing trends for tumor growth versus PsPD as early as four weeks after RT (stable disease: -0.08/month [-0.15 to -0.01], progressive disease: +0.18/month [0.07 - 0.29], PsPD: -0.53/month [-0.65 to 0.42]). CONCLUSIONS:Our findings indicate that CEST contrasts reveal tumor-specific molecular changes in gliomas following RT. These early changes support the potential of CEST MRI to differentiate PsPD from tumor growth at later time points.
Introduction:Non-invasive colorectal cancer (CRC) screening offers an important opportunity to increase colonoscopy participation and reduce mortality. This study evaluates the potential of the gut-liver axis to predict colorectal neoplasia using artificial intelligence (AI)-based analysis of the liver in routine CT images as an opportunistic screening approach. Methods:In this retrospective study, data from 1,997 patients were analyzed, including 1,189 without neoplasia and 808 with colorectal neoplasia (423 adenomas, 385 CRC). Radiomic features were extracted from three-dimensional liver segmentations, and the dataset was split into training (n = 1,397) and test (n = 600) cohorts. Five machine learning models were trained using five-fold cross-validation on the 20 most informative features. Results:The best-performing radiomics-based XGBoost model achieved a test AUROC of 0.810 (95% CI: 0.767-0.837), outperforming a clinical-only model (AUROC: 0.457). After threshold optimization, sensitivity reached 74.1% and specificity 72.3% for detecting colorectal neoplasia. Subclassification between CRC and adenoma was less accurate (AUROC: 0.674). Discussion:These findings demonstrate that AI-based liver analysis from routine CT scans can predict colorectal neoplasia, supporting its potential as an accessible adjunct to CRC screening and highlighting the gut-liver axis as a novel biomarker source.
PURPOSE:To increase performance and generalization ability of artificial intelligence prostate cancer detection systems by simulating physiological size changes of the bladder and rectum and, thereby, associated deformations of the prostate and its lesions. MATERIALS AND METHODS:This retrospective study included 1028 bi-parametric MRI examinations of men (age range: 40-90 years) performed between 2014 and 2019, divided into training/test sets (771/257). We integrated an 'anatomy-informed' transformation into the training of nnU-Net, by simulating soft-tissue deformations of the prostate resulting from size changes of the rectum and bladder. The effects of these strategies were evaluated using free-response receiver operating characteristic (FROC) to assess lesion-level performance, along with a variant: weighted alternative FROC (wAFROC), which prioritizes patient-level effects with localization criteria. Change in sensitivity was tested using a clustered McNemar test. Patient-level performance was assessed with standard and localized receiver operating characteristics (ROC/LROC) analysis. RESULTS:On the independent test set, the anatomy-informed model simulating changes of both rectum and bladder significantly increased lesion-level detection of true positive lesions by 18.8% (from 48 to 57, p = 0.01) and demonstrated significantly higher performance in the wAFROC analysis (from 0.597 to 0.639, p < 0.01). Patient-level ROC increased slightly (from 0.779 to 0.782, p = 0.89), while LROC analysis demonstrated increased performance (from 0.471 to 0.546). CONCLUSION:Simulation of rectum and bladder size variations during model training led to significant improvement in lesion detection performance, which may be crucial for diagnostics and therapeutic measures depending on correct lesion localization, e.g. MRI-guided biopsies or focal therapy regimes.
Accurate organ segmentation is crucial for prostate cancer radiotherapy, but cone-beam computer tomography (CBCT) based models are hindered by low image quality and annotation scarcity. Existing approaches rely on deformable registration, which struggles with softtissue deformations, or direct CBCT training, which suffers from domain shifts and low-quality labels. We propose a domain adaptation framework that enables robust prostate segmentation on CBCT using crossmodality supervision from planning CT (pCT). A cycle-consistent generative adversarial network translates pCT into synthetic CBCT, enabling segmentation models to train on high-quality pCT-derived annotations while adapting to CBCT characteristics. Additionally, anatomy-aware augmentation enhances robustness to organ deformations across diverse patient anatomies. Using a multi-center dataset, our approach achieves segmentation accuracy comparable to pCT-trained models. By eliminating the need for manual CBCT annotations, our method enables practical AI-driven segmentation for adaptive radiotherapy.
Abstract Background To assess the predictive value of different chemical exchange saturation transfer (CEST) contrasts, i.e. of the amide proton transfer (APT), relayed nuclear Overhauser effect (rNOE), and semi-solid magnetization transfer (ssMT), as well as of clinical routine perfusion- and diffusion-weighted MRI, in terms of treatment outcome in patients with glioma following surgery at baseline before radiotherapy at 3 T. Materials and methods From September 2018 to December 2022, 78 study participants (median age 62 years, 27/78 female) prospectively underwent CEST, diffusion, and perfusion imaging. CEST contrasts were reconstructed for the APT-weighted magnetization transfer ratio asymmetry (APTwasym), relaxation-compensated CEST metrics (MTRRexAPT, MTRRexNOE, MTRRexMT), and MTconst. Contrast-enhancing and whole tumor volumes were segmented on T2w-FLAIR and T1w images. Associations of mean contrast values with therapy response were tested using ROC analyses, while relationships with progression-free survival (PFS, median 6.04 months) and overall survival (OS, median 11.58 months), as well as added benefit compared to nCBV and ADC maps, were assessed using dichotomized Cox regression models. Results MTRRexAPT, MTRRexNOE, and MTRRexMT were associated with therapy response (AUC = 0.82, 0.81, 0.68; all p ≤ 0.03), PFS (HR = 2.92, 0.37, 3.40; all p ≤ 0.02), and OS (HR = 2.76, 0.63, 8.09; all p ≤ 0.05). MTconst was correlated with OS (HR = 5.52, p < 0.01), while APTwasym was linked to therapy response (AUC = 0.71, p = 0.02). MTRRexMT (χ² = 13.71, p < 0.01) and MTconst (χ² = 5.62, p = 0.018) provided each additional value to nCBV for OS prediction. Conclusion Relaxation-compensated CEST imaging of the APT, rNOE, and ssMT, as well as conventional APTwasym showed ability to predict treatment outcome, whilst ssMT-weighted imaging provided added benefit for OS prediction in patients with diffuse glioma following surgery at baseline before radiotherapy at 3 T.
Medical AI models trained on patient data can memorize individual samples,exposing them to membership inference attacks. Long-tail theory predicts rareexamples require memorization, yet rarity alone is insufficient: in a skin-lesiondataset, the rarest class is less vulnerable than the commoner ones. Under ran-domly initialised (MAE) and medical-image-pretrained (MedSAM) encoders,the rarest class stays most memorized, but ImageNet- and vision–language-pretrained encoders re-rank the tail toward visually distinctive classes. Across fivedatasets, controlled grayscale manipulation gives causal evidence: the same trans-form produces opposite effects on two equally rare classes, and an independentfeature-space atypicality measure predicts memorization across architectures.Memorization predicts membership inference vulnerability (ρ = 0.36–0.77,p < 0.001), and medical-domain pretraining (MedSAM, BiomedCLIP) does notreduce rare-class risk. DP-LoRA at ε = 1 reduces the most vulnerable class’s memorization by 90% versus a LoRA-only baseline, with modest accuracy coston simple tasks. We release medmem, an open-source memorization-auditing tool.
Accurate and efficient 3D segmentation is essential for both clinical and research applications. While foundation models like SAM have revolutionized interactive segmentation, their 2D design and domain shift limitations make them ill-suited for 3D medical images. Current adaptations address some of these challenges but remain limited, either lacking volumetric awareness, offering restricted interactivity, or supporting only a small set of structures and modalities. Usability also remains a challenge, as current tools are rarely integrated into established imaging platforms and often rely on cumbersome web-based interfaces with restricted functionality. We introduce nnInteractive, the first comprehensive 3D interactive open-set segmentation method. It supports diverse prompts-including points, scribbles, boxes, and a novel lasso prompt-while leveraging intuitive 2D interactions to generate full 3D segmentations. Trained on 120+ diverse volumetric 3D datasets (CT, MRI, PET, 3D Microscopy, etc.), nnInteractive sets a new state-of-the-art in accuracy, adaptability, and usability. Crucially, it is the first method integrated into widely used image viewers (e.g., Napari, MITK), ensuring broad accessibility for real-world clinical and research applications. Extensive benchmarking demonstrates that nnInteractive far surpasses existing methods, setting a new standard for AI-driven interactive 3D segmentation. nnInteractive is publicly available: https://github.com/MIC-DKFZ/napari-nninteractive (Napari plugin), https://www.mitk.org/MITK-nnInteractive (MITK integration), https://github.com/MIC-DKFZ/nnInteractive (Python backend).
Non-invasive colorectal cancer (CRC) screening represents a key opportunity to improve colonoscopy participation rates and reduce CRC mortality. This study explores the potential of the gut-liver axis for predicting colorectal neoplasia through liver-derived radiomic features extracted from routine CT images as a novel opportunistic screening approach. In this retrospective study, we analyzed data from 1,997 patients who underwent colonoscopy and abdominal CT. Patients either had no colorectal neoplasia (n=1,189) or colorectal neoplasia (n_total=808; adenomas n=423, CRC n=385). Radiomics features were extracted from 3D liver segmentations using the Radiomics Processing ToolKit (RPTK), which performed feature extraction, filtering, and classification. The dataset was split into training (n=1,397) and test (n=600) cohorts. Five machine learning models were trained with 5-fold cross-validation on the 20 most informative features, and the best model ensemble was selected based on the validation AUROC. The best radiomics-based XGBoost model achieved a test AUROC of 0.810, clearly outperforming the best clinical-only model (test AUROC: 0.457). Subclassification between colorectal cancer and adenoma showed lower accuracy (test AUROC: 0.674). Our findings establish proof-of-concept that liver-derived radiomics from routine abdominal CT can predict colorectal neoplasia. Beyond offering a pragmatic, widely accessible adjunct to CRC screening, this approach highlights the gut-liver axis as a novel biomarker source for opportunistic screening and sparks new mechanistic hypotheses for future translational research.
Computational competitions are the standard for benchmarking medical image analysis algorithms, but they typically use small curated test datasets acquired at a few centers, leaving a gap to the reality of diverse multicentric patient data. To this end, the Federated Tumor Segmentation (FeTS) Challenge represents the paradigm for real-world algorithmic performance evaluation. The FeTS challenge is a competition to benchmark (i) federated learning aggregation algorithms and (ii) state-of-the-art segmentation algorithms, across multiple international sites. Weight aggregation and client selection techniques were compared using a multicentric brain tumor dataset in realistic federated learning simulations, yielding benefits for adaptive weight aggregation, and efficiency gains through client sampling. Quantitative performance evaluation of state-of-the-art segmentation algorithms on data distributed internationally across 32 institutions yielded good generalization on average, albeit the worst-case performance revealed data-specific modes of failure. Similar multi-site setups can help validate the real-world utility of healthcare AI algorithms in the future.
OBJECTIVES:Breast diffusion-weighted imaging (DWI) has shown potential as a standalone imaging technique for certain indications, eg, supplemental screening of women with dense breasts. This study evaluates an artificial intelligence (AI)-powered computer-aided diagnosis (CAD) system for clinical interpretation and workload reduction in breast DWI. MATERIALS AND METHODS:This retrospective IRB-approved study included: n = 824 examinations for model development (2017-2020) and n = 235 for evaluation (01/2021-06/2021). Readings were performed by three readers using either the AI-CAD or manual readings. BI-RADS-like (Breast Imaging Reporting and Data System) classification was based on DWI. Histopathology served as ground truth. The model was nnDetection-based, trained using 5-fold cross-validation and ensembling. Statistical significance was determined using McNemar's test. Inter-rater agreement was calculated using Cohen's kappa. Model performance was calculated using the area under the receiver operating curve (AUC). RESULTS:The AI-augmented approach significantly reduced BI-RADS-like 3 calls in breast DWI by 29% (P =.019) and increased interrater agreement (0.57 ± 0.10 vs 0.49 ± 0.11), while preserving diagnostic accuracy. Two of the three readers detected more malignant lesions (63/69 vs 59/69 and 64/69 vs 62/69) with the AI-CAD. The AI model achieved an AUC of 0.78 (95% CI: [0.72, 0.85]; P <.001), which increased for women at screening age to 0.82 (95% CI: [0.73, 0.90]; P <.001), indicating a potential for workload reduction of 20.9% at 96% sensitivity. DISCUSSION AND CONCLUSION:Breast DWI might benefit from AI support. In our study, AI showed potential for reduction of BI-RADS-like 3 calls and increase of inter-rater agreement. However, given the limited study size, further research is needed.
Federated learning (FL) enables multi-institutional model training on clinical text without sharing raw data; however, gradient inversion methods can reconstruct sensitive information from shared model updates. The extent of such privacy leakage in FL applied to radiology reports, and the role of tokenizer design, remains unclear. To quantify gradient-based reconstruction of radiology report text in an FL setting and to compare privacy risk across three transformer tokenization strategies in a controlled, tokenizer-aware evaluation. Six FL clients trained a GPT-2–style transformer (117M parameters; sequence length 32) on two public radiology corpora comprising 368,751 diagnostic reports, 98,206 discharge summaries, and 1,500 MIMIC-CXR free-text reports. Models were trained using three tokenizers (GPT-2, RadBERT, LLaMA-2) with batch sizes of 64, 128, and 256. A curious-server threat model was assumed, and analytic gradient inversion was applied to recover text. Reconstruction fidelity was measured over five runs using exact sentence accuracy, S-BLEU, and ROUGE-L. Exact sentence reconstruction ranged from 33% to 42% across tokenizers. At batch size 64, accuracy was 42.1% (GPT-2), 42.3% (RadBERT), and 39.4% (LLaMA-2), decreasing to 37.3%, 37.2%, and 34.3% at batch size 256. S-BLEU scores declined with increasing batch size (e.g., GPT-2: 0.44→0.33; RadBERT: 0.48→0.35; LLaMA-2: 0.39→0.30). RadBERT yielded higher reconstruction fidelity and greater recovery of clinical terms, but no tokenizer prevented leakage. Substantial portions of radiology report text can be reconstructed from FL gradients even with larger batch sizes and domain-specific tokenizers. Tokenizer design influences leakage severity and should be incorporated into privacy evaluations for clinical language models. Integrating safeguards such as secure aggregation and differential privacy is necessary to meet HIPAA and GDPR requirements when deploying FL for radiology NLP. Not applicable.
Developing generalizable AI for medical imaging requires both access to large, multi-center datasets and standardized, reproducible tooling within research environments. However, leveraging real-world imaging data in clinical research environments is still hampered by strict regulatory constraints, fragmented software infrastructure, and the challenges inherent in conducting large-cohort multicentre studies. This leads to projects that rely on ad-hoc toolchains that are hard to reproduce, difficult to scale beyond single institutions and poorly suited for collaboration between clinicians and data scientists. We present Kaapana, a comprehensive open-source platform for medical imaging research that is designed to bridge this gap. Rather than building single-use, site-specific tooling, Kaapana provides a modular, extensible framework that unifies data ingestion, cohort curation, processing workflows and result inspection under a common user interface. By bringing the algorithm to the data, it enables institutions to keep control over their sensitive data while still participating in distributed experimentation and model development. By integrating flexible workflow orchestration with user-facing applications for researchers, Kaapana reduces technical overhead, improves reproducibility and enables conducting large-scale, collaborative, multi-centre imaging studies. We describe the core concepts of the platform and illustrate how they can support diverse use cases, from local prototyping to nation-wide research networks. The open-source codebase is available at https://github.com/kaapana/kaapana
Breast cancer detection, and broadly medical object detection, revolves around discovering and rating lesions. One of the most common ways of measuring performance is FROC (Free-response Receiver Operating Characteristic), which calculates sensitivity at predefined thresholds of false positives per case. However, depending on the clinical context, not all lesions might be of equivocal impact on the long-term outcome of a patient. Some lesions missed e.g. in screening might be detected in the subsequent screening round without impacting the clinical prognosis, whilst missing others might significantly detoriate prognosis and treatment pathways. It is therefore desirable to develop and include consideration of clinical prognosis/risk imbalance in the way machine learning models are developed and evaluated. In this work, we propose risk-adjusted FROC (raFROC), an adaptation of FROC that constitutes a first step on reflecting the underlying clinical need more accurately. Experiments on two independent breast magnetic resonance imaging (MRI) datasets with a total of 1535 lesions in 1735 subjects showcase the clinical potential of the proposed metric and its advantages over traditional evaluation methods. Additionally, by utilizing a risk-adjusted adaptation of focal loss (raFocal) we are able to improve the raFROC results and patient-level performance of nnDetection, at no expense of the regular FROC.
Large-scale pre-training holds the promise to advance 3D medical object detection, a crucial component of accurate computer-aided diagnosis. Yet, it remains underexplored compared to segmentation, where pre-training has already demonstrated significant benefits. Existing pre-training approaches for 3D object detection rely on 2D medical data or natural image pre-training, failing to fully leverage 3D volumetric information. In this work, we present the first systematic study of how existing pre-training methods can be integrated into state-of-the-art detection architectures, covering both CNNs and Transformers. Our results show that pre-training consistently improves detection performance across various tasks and datasets. Notably, reconstruction-based self-supervised pre-training outperforms supervised pre-training, while contrastive pre-training provides no clear benefit for 3D medical object detection. Our code is publicly available at: https://github.com/MIC-DKFZ/nnDetection-finetuning.
Background: Federated learning (FL) enables collaborative training of clinical AI models without centralizing patient data, but adoption is limited by privacy concerns, heterogeneous institutional compliance, and resource disparities; standard differential privacy (DP) applies uniform noise to all clients, penalizing well-compliant or under-resourced institutions. Objective: We introduce a compliance-aware FL framework that adapts DP to institutional compliance, letting lower-compliance sites participate without uniformly penalizing others. Methods: A compliance scoring tool aligned with HIPAA, GDPR, NIST, ISO, and HL7/FHIR maps each client score to a per-step Gaussian noise scale for server-side DP-SGD on a small aggregator dataset. The formal (ε,δ) bound applies to the aggregator dataset under a semi-honest aggregator; client-level DP needs secure aggregation (future work). We evaluate five FL strategies on PneumoniaMNIST and BreastMNIST (16 clients, 50 rounds, five seeds); the cumulative aggregator-dataset ε is about 1434 (Breast) and 513 (Pneumonia) at δ=10^-5. Results: Including 12 lower-compliance clients (Experiment 1) versus a compliant-only baseline (Experiment 4) changed BreastMNIST accuracy by +4.5 (FedAvg), +6.8 (FedMedian), +5.2 (FedProx), +1.6 (FedYogi), and -4.1 (FedAdam) percentage points (pooled +2.8 pp; not significant at n=5; up to +17 pp per configuration); compliance-weighted allocation matched uniform server-side DP at equal mean noise (+0.1 pp), carrying no utility penalty, and first-round noise cost 1.3 pp (Breast) and 2.5 pp (Pneumonia, FedAvg). Conclusions: Compliance-weighted server-side DP lets lower-compliance institutions join FL without degrading performance, giving auditable per-site noise control at no utility cost; formal guarantees apply to the aggregator dataset, with client-level DP requiring secure aggregation.
Federated learning (FL) plays a vital role in boosting both accuracy and privacy in the collaborative medical imaging field. The importance of privacy increases with the diverse security standards across nations and corporations, particularly in healthcare and global FL initiatives. Current research on privacy attacks in federated medical imaging focuses on sophisticated gradient inversion attacks that can reconstruct images from FL communications.
Prognosis for thoracic aortic aneurysms is significantly worse for women than men, with a higher mortality rate observed among female patients. The increasing use of magnetic resonance breast imaging (MRI) offers a unique opportunity for simultaneous detection of both breast cancer and thoracic aortic aneurysms. We retrospectively validate a fully-automated artificial neural network (ANN) pipeline on 5057 breast MRI examinations from public (Duke University Hospital/EA1141 trial) and in-house (Erlangen University Hospital) data. The ANN, benchmarked against 3D-ground-truth segmentations, clinical reports, and a multireader panel, demonstrates high technical robustness (dice/clDice 0.88-0.91/0.97-0.99) across different vendors and field strengths. The ANN improves aneurysm detection rates by 3.5-fold compared with routine clinical readings, highlighting its potential to improve early diagnosis and patient outcomes. Notably, a higher odds ratio (OR = 2.29, CI: [0.55,9.61]) for thoracic aortic aneurysms is observed in women with breast cancer or breast cancer history, suggesting potential further benefits from integrated simultaneous assessment for cancer and aortic aneurysms.