Meningiomas are the most common primary intracranial tumors, frequently requiring radiotherapy as a part of management. Effective radiotherapy planning for meningiomas necessitates accurate and consistent segmentation of target volumes on MRI, a process that is complex, labor-intensive, and dependent on expert expertise. The 2024 Brain Tumor Segmentation Challenge Meningioma Radiotherapy (BraTS-MEN-RT) Dataset addresses this problem by providing the largest multi-institutional collection of systematically annotated radiotherapy planning MRIs for meningiomas. Publicly accessible, this dataset comprises 570 radiotherapy planning 3D T1-weighted post-contrast MRIs at native resolutions, with 500 cases featuring expert-annotated gross tumor volumes (GTV). Annotations follow standardized radiotherapy planning protocols and include both intact and postoperative meningioma cases, ensuring wide clinical relevance. Contributions from seven diverse medical centers across the United States and the United Kingdom enhance the dataset’s generalizability. The dataset aims to accelerate the development of automated segmentation methods for radiotherapy planning, improving workflow efficiency, reducing interobserver variability, and ultimately enhancing patient outcomes.
Despite continuous advancements in cancer treatment, brain metastatic disease remains a significant complication of primary cancer and is associated with an unfavorable prognosis. One approach for improving diagnosis, management, and outcomes is to implement algorithms based on artificial intelligence for the automated segmentation of both pre- and post-treatment MRI brain images. Such algorithms rely on volumetric criteria for lesion identification and treatment response assessment, which are still not available in clinical practice. Therefore, it is critical to establish tools for rapid volumetric segmentations methods that can be translated to clinical practice and that are trained on high quality annotated data. The BraTS-METS 2025 Lighthouse Challenge aims to address this critical need by establishing inter-rater and intra-rater variability in dataset annotation by generating high quality annotated datasets from four individual instances of segmentation by neuroradiologists while being recorded on video (two instances doing "from scratch" and two instances after AI pre-segmentation). This high-quality annotated dataset will be used for testing phase in 2025 Lighthouse challenge and will be publicly released at the completion of the challenge. The 2025 Lighthouse challenge will also release the 2023 and 2024 segmented datasets that were annotated using an established pipeline of pre-segmentation, student annotation, two neuroradiologists checking, and one neuroradiologist finalizing the process. It builds upon its previous edition by including post-treatment cases in the dataset. Using these high-quality annotated datasets, the 2025 Lighthouse challenge plans to test benchmark algorithms for automated segmentation of pre-and post-treatment brain metastases (BM), trained on diverse and multi-institutional datasets of MRI images obtained from patients with brain metastases.
Benchmarking competitions are central to the development of artificial intelligence (AI) in medical imaging, defining performance standards and shaping methodological progress. However, it remains unclear whether these benchmarks provide data that are sufficiently representative, accessible, and reusable to support clinically meaningful AI. In this work, we assess fairness along two complementary dimensions: (1) whether challenge datasets are representative of real-world clinical diversity, and (2) whether they are accessible and legally reusable in line with the FAIR principles. To address this question, we conducted a large-scale systematic study of 241 biomedical image analysis challenges comprising 458 tasks across 19 imaging modalities. Our findings show substantial biases in dataset composition, including geographic location, modality-, and problem type-related biases, indicating that current benchmarks do not adequately reflect real-world clinical diversity. Despite their widespread influence, challenge datasets were frequently constrained by restrictive or ambiguous access conditions, inconsistent or non-compliant licensing practices, and incomplete documentation, limiting reproducibility and long-term reuse. Together, these shortcomings expose foundational fairness limitations in our benchmarking ecosystem and highlight a disconnect between leaderboard success and clinical relevance.
Computational competitions are the standard for benchmarking medical image analysis algorithms, but they typically use small curated test datasets acquired at a few centers, leaving a gap to the reality of diverse multicentric patient data. To this end, the Federated Tumor Segmentation (FeTS) Challenge represents the paradigm for real-world algorithmic performance evaluation. The FeTS challenge is a competition to benchmark (i) federated learning aggregation algorithms and (ii) state-of-the-art segmentation algorithms, across multiple international sites. Weight aggregation and client selection techniques were compared using a multicentric brain tumor dataset in realistic federated learning simulations, yielding benefits for adaptive weight aggregation, and efficiency gains through client sampling. Quantitative performance evaluation of state-of-the-art segmentation algorithms on data distributed internationally across 32 institutions yielded good generalization on average, albeit the worst-case performance revealed data-specific modes of failure. Similar multi-site setups can help validate the real-world utility of healthcare AI algorithms in the future.
Decentralized AI systems, such as federated learning, can play a critical role in further unlocking AI asset marketplaces (e.g., healthcare data marketplaces) thanks to increased asset privacy protection. Unlocking this big potential necessitates governance mechanisms that are transparent, scalable, and verifiable. However current governance approaches rely on bespoke, infrastructure-specific policies that hinder asset interoperability and trust among systems. We are proposing a Technical Policy Blueprint that encodes governance requirements as policy-as-code objects and separates asset policy verification from asset policy enforcement. In this architecture the Policy Engine verifies evidence (e.g., identities, signatures, payments, trusted-hardware attestations) and issues capability packages. Asset Guardians (e.g. data guardians, model guardians, computation guardians, etc.) enforce access or execution solely based on these capability packages. This core concept of decoupling policy processing from capabilities enables governance to evolve without reconfiguring AI infrastructure, thus creating an approach that is transparent, auditable, and resilient to change.
Glioblastoma is the most common malignant adult brain tumor with poor prognosis, largely due to its heterogeneous landscape. The AI-RANO and RANO-RGP groups organized the BraTS-Pathology 2024 challenge, providing a publicly-available dataset and a benchmarking environment for AI models to automatically identify six clinically-relevant histopathologic sub-regions in H&E-stained whole slide images (WSIs). We reclassified the TCGA-GBM and TCGA-LGG data according to the WHO 2021 criteria and identified diagnostic WSIs of 188 glioblastoma (IDH-wt, CNS WHO Gr.4) patients. A well-defined annotation protocol was communicated to eight neuropathologists, across eight institutions, who created “ground truth” annotations for Cellular Tumor (CT), Geographic Necrosis, Cortical Infiltration, Pseudopalisading Necrosis (PN), Microvascular Proliferation, and Penetration into White Matter (PWM). Comprehensive WSI curation for elimination of artifactual content and follow up partitioning resulted in approximately 300,000 data samples/patches. Inter-observer variability assessment highlighted difficulties in establishing a “gold standard”. The least annotated sub-regions (PWM and PN) were the most challenging to the participating methods, whereas all methods achieved their highest accuracy in the most annotated sub-region (CT). The highest average performance across sub-regions was obtained by a foundation model-based method. However, this top-ranked solution achieved an F1 score of 0.51 (Accuracy=0.65, Sensitivity=0.56, Specificity=0.93). Our findings suggest that achieving state-of-the-art performance, in automatically identifying clinically-relevant glioblastoma sub-regions, remains an unsolved research problem. Analysis of the ground-truth annotations underscores the critical need for more precisely defined and neuropathologically standardized annotation protocol. Current AI algorithms are facing difficulties in addressing glioblastoma’s heterogeneous histologic landscape in the absence of large, comprehensively, and consistently annotated data. Building upon these findings, and towards developing more accurate AI algorithms, the 2025 iteration of the BraTS-Path challenge is currently on-going by leveraging two million data samples, aiming to enhance our disease understanding and improve diagnostic accuracy, ultimately contributing to improved patient outcomes.
Progress in automated volumetric segmentation of postoperative glioblastoma has been hindered by limited data availability, primarily related to patient privacy concerns. To overcome this, we developed the Federated Tumor Segmentation (FeTS) 2.0 initiative, which leveraged federated learning to enable training of an AI-based glioblastoma segmentation model by 53 global institutions without the need for data sharing. Multi-parametric MRI of 10,208 follow up timepoints from postoperative glioblastoma patients (8,100 training, 2,108 hold-out validation) were pre-processed and expertly annotated across the 53 contributing institutions. The provided annotation protocol comprised necrosis (NE), peritumoral edematous/infiltrated tissue (TIE), enhancing tumor (ET), and resection cavity (RC). Federated learning of an nnUNet model, executed through the MLCommons’ MedPerf orchestrator and Intel’s OpenFL framework, proceeded for over 600 epochs with trained model weights and validation scores being the only data shared between collaborators. The best performing consensus model achieved hold-out validation Dice Similarity Coefficients (DSC) of 0.95, 0.94, 0.89, and 0.77 for TIE, ET, RC, and NE, respectively, indicating new record performance for postoperative glioblastoma segmentation. In comparison, the centralized BraTS 2024 postoperative glioma study was trained on 1,350 follow up timepoints and tested on 397 hold-out follow up timepoints, with class-specific corresponding DSC of 0.89, 0.79, 0.78, and 0.81 respectively. The FeTS 2.0 initiative demonstrates that federated learning enables scaled, geographically diverse, de-centralized collaboration, expediting big data analysis, while obviating privacy concerns, and hence improves AI performance and generalizability. By enabling reliable and reproducible delineation of postoperative regions-of-interest, our AI model paves the way for large-scale retrospective and prospective analyses aimed at defining volumetric thresholds that correlate with clinical outcomes. Future studies will perform technical and clinical validation on out-of-sample clinical trial datasets, generating evidence to establish three-dimensional cutoff values that can be considered in future RANO criteria, thereby standardizing glioblastoma response assessment and supporting adaptive, data-driven therapeutic decision making worldwide.
Clinnova, a collaborative initiative involving France, Germany, Switzerland, and Luxembourg, is dedicated to unlocking the power of precision medicine through data federation, standardization, and interoperability. This European Greater Region initiative seeks to create an interoperable European standard using artificial intelligence (AI) and data science to enhance healthcare outcomes and efficiency. Key components include multidisciplinary research centers, a federated biobanking strategy, a digital health innovation platform, and a federated AI strategy. It targets inflammatory bowel disease, rheumatoid diseases, and multiple sclerosis (MS), emphasizing data quality to develop AI algorithms for personalized treatment and translational research. The IHU Strasbourg (Institute of Minimal-invasive Surgery) has the lead in this initiative to develop the federated learning (FL) proof of concept (POC) that will serve as a foundation for advancing AI in healthcare. At its core, Clinnova-MS aims to enhance MS patient care by using FL to develop more accurate models that detect disease progression, guide interventions, and validate digital biomarkers across multiple sites. This technical report presents insights and key takeaways from the first cross-border federated POC on MS segmentation of MRI images within the Clinnova framework. While our work marks a significant milestone in advancing MS segmentation through cross-border collaboration, it also underscores the importance of addressing technical, logistical, and ethical considerations to realize the full potential of FL in healthcare settings.
Federated learning (FL) in healthcare suffers from non-identically distributed (non-IID) data, impacting model convergence and performance. While existing solutions for the non-IID problem often do not quantify the degree of non-IID nature between clients in the federation, assessing it can improve training experiences and outcomes, particularly in real-world scenarios with unfamiliar datasets. The paper presents a practical non-IID assessment methodology for a medical segmentation problem, highlighting its significance in medical FL. We propose a simple yet effective solution that utilizes distance measurements in the embedding space of medical images and statistical measurements calculated over their metadata. Our method, designed for medical imaging and integrated into federated averaging, improves model generalization by downgrading the contribution from the most distant client, treating it as an outlier. Additionally, it enhances model personalization by introducing distance-based clustering of clients. To the best of our knowledge, this method is the first to use distance-based techniques for providing a practical solution to the non-IID problem within the medical imaging FL domain. Furthermore, we validate our approach on three public FL imaging radiology datasets (FeTS, Prostate, and Fed-KITS2019) to demonstrate its effectiveness across various radiology imaging scenarios.
The 2024 Brain Tumor Segmentation Meningioma Radiotherapy (BraTS-MEN-RT) challenge aimed to advance automated segmentation algorithms using the largest known multi-institutional dataset of 750 radiotherapy planning brain MRIs with expert-annotated target labels for patients with intact or postoperative meningioma that underwent either conventional external beam radiotherapy or stereotactic radiosurgery. Each case included a defaced 3D post-contrast T1-weighted radiotherapy planning MRI in its native acquisition space, accompanied by a single-label "target volume" representing the gross tumor volume (GTV) and any at-risk post-operative site. Target volume annotations adhered to established radiotherapy planning protocols, ensuring consistency across cases and institutions, and were approved by expert neuroradiologists and radiation oncologists. Six participating teams developed, containerized, and evaluated automated segmentation models using this comprehensive dataset. Team rankings were assessed using a modified lesion-wise Dice Similarity Coefficient (DSC) and 95
Pediatric central nervous system tumors are the leading cause of cancer-related deaths in children. The five-year survival rate for high-grade glioma in children is less than 20%. The development of new treatments is dependent upon multi-institutional collaborative clinical trials requiring reproducible and accurate centralized response assessment. We present the results of the BraTS-PEDs 2023 challenge, the first Brain Tumor Segmentation (BraTS) challenge focused on pediatric brain tumors. This challenge utilized data acquired from multiple international consortia dedicated to pediatric neuro-oncology and clinical trials. BraTS-PEDs 2023 aimed to evaluate volumetric segmentation algorithms for pediatric brain gliomas from magnetic resonance imaging using standardized quantitative performance evaluation metrics employed across the BraTS 2023 challenges. The top-performing AI approaches for pediatric tumor analysis included ensembles of nnU-Net and Swin UNETR, Auto3DSeg, or nnU-Net with a self-supervised framework. The BraTS-PEDs 2023 challenge fostered collaboration between clinicians (neuro-oncologists, neuroradiologists) and AI/imaging scientists, promoting faster data sharing and the development of automated volumetric analysis techniques. These advancements could significantly benefit clinical trials and improve the care of children with brain tumors.
While laparoscopic liver resection is less prone to complications and maintains patient outcomes compared to traditional open surgery, its complexity hinders widespread adoption due to challenges in representing the liver's internal structure. Laparoscopic intraoperative ultrasound offers efficient, cost-effective and radiation-free guidance. Our objective is to aid physicians in identifying internal liver structures using laparoscopic intraoperative ultrasound. We propose a patient-specific approach using preoperative 3D ultrasound liver volume to train a deep learning model for real-time identification of portal tree and branch structures. Our personalized AI model, validated on ex vivo swine livers, achieved superior precision (0.95) and recall (0.93) compared to surgeons, laying groundwork for precise vessel identification in ultrasound-based liver resection. Its adaptability and potential clinical impact promise to advance surgical interventions and improve patient care.
Glioblastoma is the most common primary adult brain tumor, with a grim prognosis - median survival of 12-18 months following treatment, and 4 months otherwise. Glioblastoma is widely infiltrative in the cerebral hemispheres and well-defined by heterogeneous molecular and micro-environmental histopathologic profiles, which pose a major obstacle in treatment. Correctly diagnosing these tumors and assessing their heterogeneity is crucial for choosing the precise treatment and potentially enhancing patient survival rates. In the gold-standard histopathology-based approach to tumor diagnosis, detecting various morpho-pathological features of distinct histology throughout digitized tissue sections is crucial. Such "features" include the presence of cellular tumor, geographic necrosis, pseudopalisading necrosis, areas abundant in microvascular proliferation, infiltration into the cortex, wide extension in subcortical white matter, leptomeningeal infiltration, regions dense with macrophages, and the presence of perivascular or scattered lymphocytes. With these features in mind and building upon the main aim of the BraTS Cluster of Challenges https://www.synapse.org/brats2024, the goal of the BraTS-Path challenge is to provide a systematically prepared comprehensive dataset and a benchmarking environment to develop and fairly compare deep-learning models capable of identifying tumor sub-regions of distinct histologic profile. These models aim to further our understanding of the disease and assist in the diagnosis and grading of conditions in a consistent manner.
Objective: Evaluation of prognostic factors is crucial in patients with endometrial cancer for optimal treatment planning and prognosis assessment. This study proposes a deep learning pipeline for tumor and uterus segmentation from magnetic resonance imaging (MRI) images to predict deep myometrial invasion and cervical stroma invasion and thus assist clinicians in pre-operative workups. Methods: Two experts consensually reviewed the MRIs and assessed myometrial invasion and cervical stromal invasion as per the International Federation of Gynecology and Obstetrics staging classification, to compare the diagnostic performance of the model with the radiologic consensus. The deep learning method was trained using sagittal T2-weighted images from 142 patients and tested with a 3-fold stratified test with 36 patients in each fold. Our solution is based on a segmentation module, which employed a 2stage pipeline for efficient uterus in the whole MRI volume and then tumor segmentation in the uterus predicted region of interest. Results: A total of 178 patients were included. For deep myometrial invasion prediction, the model achieved an average balanced test accuracy over 3-folds of 0.702, while experts reached an average accuracy of 0.769. For cervical stroma invasion prediction, our model demonstrated an average balanced accuracy of 0.721 on the 3-fold test set, while experts achieved an average balanced accuracy of 0.859. Additionally, the accuracy rates for uterus and tumor segmentation, measured by the Dice score, were 0.847 and 0.579 respectively. Conclusion: Despite the current challenges posed by variations in data, class imbalance, and the presence of artifacts, our fully automatic approach holds great promise in supporting in pre-operative staging. Moreover, it demonstrated a robust capability to segment key regions of interest, specifically the uterus and tumors, highlighting the positive impact our solution can bring to health care imaging.
The translation of AI-generated brain metastases (BM) segmentation into clinical practice relies heavily on diverse, high-quality annotated medical imaging datasets. The BraTS-METS 2023 challenge has gained momentum for testing and benchmarking algorithms using rigorously annotated internationally compiled real-world datasets. This study presents the results of the segmentation challenge and characterizes the challenging cases that impacted the performance of the winning algorithms. Untreated brain metastases on standard anatomic MRI sequences (T1, T2, FLAIR, T1PG) from eight contributed international datasets were annotated in stepwise method: published UNET algorithms, student, neuroradiologist, final approver neuroradiologist. Segmentations were ranked based on lesion-wise Dice and Hausdorff distance (HD95) scores. False positives (FP) and false negatives (FN) were rigorously penalized, receiving a score of 0 for Dice and a fixed penalty of 374 for HD95. The mean scores for the teams were calculated. Eight datasets comprising 1303 studies were annotated, with 402 studies (3076 lesions) released on Synapse as publicly available datasets to challenge competitors. Additionally, 31 studies (139 lesions) were held out for validation, and 59 studies (218 lesions) were used for testing. Segmentation accuracy was measured as rank across subjects, with the winning team achieving a LesionWise mean score of 7.9. The Dice score for the winning team was 0.65 ± 0.25. Common errors among the leading teams included false negatives for small lesions and misregistration of masks in space. The Dice scores and lesion detection rates of all algorithms diminished with decreasing tumor size, particularly for tumors smaller than 100 mm3. In conclusion, algorithms for BM segmentation require further refinement to balance high sensitivity in lesion detection with the minimization of false positives and negatives. The BraTS-METS 2023 challenge successfully curated well- annotated, diverse datasets and identified common errors, facilitating the translation of BM segmentation across varied environments and providing the tools for future development of personalized volumetric reports to patients undergoing BM treatment.
Validation metrics are key for the reliable tracking of scientific progress and for bridging the current chasm between artificial intelligence (AI) research and its translation into practice. However, increasing evidence shows that particularly in image analysis, metrics are often chosen inadequately in relation to the underlying research problem. This could be attributed to a lack of accessibility of metric-related knowledge: While taking into account the individual strengths, weaknesses, and limitations of validation metrics is a critical prerequisite to making educated choices, the relevant knowledge is currently scattered and poorly accessible to individual researchers. Based on a multi-stage Delphi process conducted by a multidisciplinary expert consortium as well as extensive community feedback, the present work provides the first reliable and comprehensive common point of access to information on pitfalls related to validation metrics in image analysis. Focusing on biomedical image analysis but with the potential of transfer to other fields, the addressed pitfalls generalize across application domains and are categorized according to a newly created, domain-agnostic taxonomy. To facilitate comprehension, illustrations and specific examples accompany each pitfall. As a structured body of information accessible to researchers of all levels of expertise, this work enhances global comprehension of a key topic in image analysis validation.
We describe the design and results from the BraTS 2023 Intracranial Meningioma Segmentation Challenge. The BraTS Meningioma Challenge differed from prior BraTS Glioma challenges in that it focused on meningiomas, which are typically benign extra-axial tumors with diverse radiologic and anatomical presentation and a propensity for multiplicity. Nine participating teams each developed deep-learning automated segmentation models using image data from the largest multi-institutional systematically expert annotated multilabel multi-sequence meningioma MRI dataset to date, which included 1000 training set cases, 141 validation set cases, and 283 hidden test set cases. Each case included T2, FLAIR, T1, and T1Gd brain MRI sequences with associated tumor compartment labels delineating enhancing tumor, non-enhancing tumor, and surrounding non-enhancing FLAIR hyperintensity. Participant automated segmentation models were evaluated and ranked based on a scoring system evaluating lesion-wise metrics including dice similarity coefficient (DSC) and 95% Hausdorff Distance. The top ranked team had a lesion-wise median dice similarity coefficient (DSC) of 0.976, 0.976, and 0.964 for enhancing tumor, tumor core, and whole tumor, respectively and a corresponding average DSC of 0.899, 0.904, and 0.871, respectively. These results serve as state-of-the-art benchmarks for future pre-operative meningioma automated segmentation algorithms. Additionally, we found that 1286 of 1424 cases (90.3%) had at least 1 compartment voxel abutting the edge of the skull-stripped image edge, which requires further investigation into optimal pre-processing face anonymization steps.