This paper presents our submission to the AortaSeg Challenge at MICCAI 2024, which focuses on multiclass segmentation of the aorta into anatomically defined zones and branches. We adopted a data-centric strategy, emphasizing rigorous data preparation and carefully designed pre- and post-processing steps over architectural novelty. Our approach is built upon a classical 3D RUNet-based model, demonstrating that such architectures remain highly competitive in low-data scenarios when paired with robust engineering practices. Specifically, we implemented a two-step, patch-based segmentation pipeline, incorporating targeted data augmentation and class-aware sampling during training. This design aimed to improve performance on small and underrepresented anatomical structures. On the official test dataset, our method achieved an average Dice Similarity Coefficient of 0.755 ± 0.038 and a Normalized Surface Distance of 0.788 ± 0.042, outperforming the baseline in the majority of evaluated regions. These results highlight the effectiveness of prioritizing data quality and processing techniques, and underscore the continued relevance of classical segmentation models in practical, data-constrained medical imaging tasks.
Spatial mapping of lung adenocarcinoma (LUAD) growth patterns across whole slide images (WSIs) requires resolving architectural context at the region level, yet existing methods operate at the individual tile level and produce generic morphological clusters rather than clinically defined pattern maps. We propose a weakly supervised Bag-of-Visual-Words (BoVW) pipeline that learns a visual vocabulary from frozen foundation model embeddings extracted from a small set of annotated regions of interest (ROIs). Pattern prototypes are constructed as mean BoVW histograms of same-label ROIs and used for nearest-prototype classification of sliding-window regions under Jensen–Shannon divergence. The resulting predictions are projected onto the WSI tile grid to produce interpretable spatial pattern maps. We evaluate the method on 87 CPTAC-LUAD patients using three foundation model encoders and multiple vocabulary sizes on two clinically motivated tasks. For tumour/healthy classification, the best configuration achieves a balanced accuracy of 0.974 with H-Optimus-1, approaching the 0.987 obtained by a supervised SVM trained on mean-pooled WSI embeddings. For binary histologic grade classification, the BoVW pipeline achieves higher balanced accuracy than the supervised baseline for all encoders, suggesting that ROI-level pattern decomposition preserves grade-relevant heterogeneity that is attenuated by global mean pooling.
Retrieving histopathological images can assist in the recognition and treatment planning of several diseases. Nevertheless, high-dimensional features can make this process complex and inefficient. These challenges can be addressed by encoding the feature domain into binary codes of different lengths utilizing deep hashing approaches. Still, the vanishing gradient challenge remains a concern in these approaches. According to several studies, quadruplet deep hashing models have exhibited promising performance in retrieving images from multi-category datasets. Furthermore, adding an attention module to a convolutional neural network architecture can increase the efficiency of feature extraction. Thus, we introduce an adaptive quadruplet deep hashing model to retrieve histopathological images. Four designed deep hashing models with matching structures and parameters are utilized to produce hash codes. The resulting codes are trained according to a novel adaptive quadruplet loss function. The adaptive structure is capable of improving retrieval performance. The presented approach also suggests a novel hash layer for the vanishing gradient issue. In addition, a simple yet effective attention module is implemented to enhance feature extraction performance. Our model is evaluated on three publicly available histopathology datasets: Kather, Kimia Path960, and Kimia Path24C. The results indicate that the suggested approach achieves the highest mean average precision (MAP) of approximately 0.9940, 0.9983, and 0.9968 for the respective datasets. Based on experiments performed on the datasets, our model surpasses current hashing techniques.
Hypoxic Ischemic Encephalopathy (HIE) represents a brain dysfunction, affecting approximately 1 to 5 per 1000 full-term neonates. The precise delineation and segmentation of HIE-related lesions in neonatal brain Magnetic Resonance Images (MRI) are pivotal in advancing outcome predictions, identifying patients at high risk, elucidating neurological manifestations, and assessing treatment efficacies. Despite its importance, the development of algorithms for segmenting HIE lesions from MRI volumes has been impeded by data scarcity. Addressing this critical gap, we organized the first BONBID-HIE challenge with diffusion MRI data (Apparent Diffusion Coefficient (ADC) maps) for HIE lesion segmentation, in conjunction with the MICCAI 2023. Totally 14 algorithms were submitted, employing a gamut of cutting-edge automatic machine-learning-based segmentation algorithms. Our comprehensive analysis of HIE lesion segmentation and submitted algorithms facilitates an in-depth evaluation of the current technological zenith, outlines directions for future advancements, and highlights persistent hurdles. To foster ongoing research and benchmarking, the annotated HIE dataset, developed algorithm dockers, and unified evaluation codes are accessible through a dedicated online platform (https://bonbid-hie2023.grand-challenge.org).
Radiological reports are essential for clinical diagnosis. However, their preparation is laborious and time-consuming for radiologists. Artificial intelligence (AI) has the potential to help radiologists reduce this workload. This scoping review aims to map current research on AI-based radiological report generation, highlighting its applications, limitations, and future directions. We systematically retrieved 321 records from three major scientific databases (PubMed, Embase, Web of Science) and included 58 studies after screening. Benchmark datasets, radiological subspecialties, AI models, output formats, and evaluation methods were extracted and analyzed. Nine public benchmark datasets and six radiological subspecialties were identified. Transformer-based architectures have become the dominant approach, with outputs primarily generated in natural language. Dataset diversity is limited, with most studies relying on chest X-ray images from two public datasets. Structured outputs for machine readability remain underexplored, and expert human evaluation is infrequently used as an assessment method. Current AI research in radiological report generation is constrained by limited dataset diversity and evaluation practices. Future studies should expand imaging modalities and subspecialties, adopt unified vision-language Transformer models, generate more structured outputs for programmatic use, and incorporate expert human evaluation for quality assessment.
OBJECTIVE:Information extraction (IE) from clinical texts has advanced rapidly with recent advances in natural language processing, particularly the advent of large language models (LLMs). However, inconsistent and incomplete reporting of methodologies limits reproducibility, comparability, and clinical translation. We aimed to develop a consensus-based reporting guideline tailored to clinical IE studies. MATERIAL AND METHODS:We developed the Clinical Information Extraction Reporting Guideline (CINEX) through a multi-phase process. The initiative was prospectively registered on the EQUATOR Network as a reporting guideline under development, and a detailed Delphi study protocol was published in advance. A scoping review informed an initial set of items, which was refined through a 3-round electronic Delphi study with 20 international experts, followed by a final consensus meeting. Items were iteratively refined based on predefined inclusion criteria and expert feedback. RESULTS:The final CINEX guideline comprises 29 checklist items grouped into 5 domains: information model, architecture, data, annotation, and outcomes. The 3 eDelphi rounds included 20, 15, and 12 experts, respectively. Two items were added after round one. Consensus for inclusion was reached for a 21 and additional 7 items after rounds 2 and 3, respectively. DISCUSSION:CINEX provides a structured framework to improve transparency, reproducibility, and interpretability in clinical IE research. By standardizing reporting of key methodological components, such as data provenance, annotation processes, and evaluation strategies, it facilitates meaningful comparison across studies and supports safer clinical implementation. CONCLUSION:CINEX complements existing AI reporting standards by addressing domain-specific challenges in clinical IE.
Background:Risk of bias (RoB) assessment of randomized clinical trials (RCTs) is vital to answering systematic review questions accurately. Manual RoB assessment for hundreds of RCTs is a cognitively demanding and lengthy process. Automation has the potential to assist reviewers in rapidly identifying text descriptions in RCTs that indicate potential risks of bias. However, no RoB text span annotated corpus could be used to fine-tune or evaluate large language models (LLMs), and there are no established guidelines for annotating the RoB spans in RCTs. Objective:The revised Cochrane RoB 2 test (RoB 2) tool provides comprehensive guidelines for RoB assessment; however, due to the inherent subjectivity of this tool, it cannot be directly used as RoB annotation guidelines. The study aimed to develop precise RoB text span annotation instructions that could address this subjectivity and thus aid the corpus annotation. Methods:We leveraged RoB 2 guidelines to develop visual instructional placards that serve as annotation guidelines for RoB spans and risk judgments. Expert annotators used these visual placards to annotate a dataset named RoBuster, consisting of 41 full-text RCTs from the domains of physiotherapy and rehabilitation. We report interannotator agreement (IAA) between 2 annotators for text span annotations before and after applying visual instructions on a subset (n=9) of RoBuster. We also provide IAA on bias risk judgments using Cohen κ. Moreover, we used a portion of RoBuster (n=10) to evaluate an LLM using a straightforward evaluation framework. This evaluation aimed to gauge the performance of an LLM (here GPT 3.5) in the challenging task of RoB span extraction and demonstrate the utility of this corpus using a straightforward framework. Results:We present a corpus of 41 RCTs with fine-grained text span annotations comprising more than 28,427 tokens belonging to 22 RoB classes. The IAA at the text span level calculated using the F1 measure varies from 0% to 90%, while Cohen κ for risk judgments ranges between -0.235 and 1.0. Using visual instructions for annotation increases the IAA by more than 17 percentage points. LLM (GPT-3.5) shows promising but varied observed agreements with the expert annotation across the different bias questions. Conclusions:Despite having comprehensive bias assessment guidelines and visual instructional placards, RoB annotation remains a complex task. Using visual placards for bias assessment and annotation enhances IAA compared to cases where visual placards are absent; however, text annotation remains challenging for the subjective questions and the questions for which annotation data are unavailable in RCTs. Similarly, while GPT-3.5 demonstrates effectiveness, its accuracy diminishes with more subjective RoB questions and low information availability.
Since its inception in 2003, the various ImageCLEF challenges have provided large and complex datasets targeting a wide array of subjects in medicine, argumentation, reasoning, content recommendation, data generation, and question answering. In its 24th edition at CLEF, ImageCLEF will have five main tasks: (i) a Medical task, which aims to promote the synergy between four medical challenges: Caption, involving concept detection and caption prediction in radiology images, Synthetic Medical Image Generation in the GANs task, Visual Question Answering for improving the diagnosis and classification of real medical gastrointestinal images, and multimodal dermatology response generation and a new MEDIQA-CORE challenge focusing on predicting or correcting tumor type labels and identifying and summarizing major and minor differences between pairs of radiology reports; (ii) the ToPicto task, involving text to pictogram translation and prediction, (iii) the Multimodal Reasoning task on visual, multi-language, interdisciplinary question answering, and a task using multi-spectral remote sensing images: (iv) AI4Agriculture, involving predicting agricultural potential before planing and crop type identification, and (v) the Deepfake detection and generation task. In its last edition, 56 teams finished our challenges, continuing to show the impact in the community.
This study compares baseline (J0) and 24-hour (J1) diffusion magnetic resonance imaging (MRI) for predicting three-month functional outcomes after acute ischemic stroke (AIS). Seventy-four AIS patients with paired apparent diffusion coefficient (ADC) scans and clinical data were analyzed. Three-dimensional ResNet-50 embeddings were fused with structured clinical variables, reduced via principal component analysis (<=12 components), and classified using linear support vector machines with eight-fold stratified group cross-validation. J1 multimodal models achieved the highest predictive performance (AUC = 0.923 +/- 0.085), outperforming J0-based configurations (AUC <= 0.86). Incorporating lesion-volume features further improved model stability and interpretability. These findings demonstrate that early post-treatment diffusion MRI provides superior prognostic value to pre-treatment imaging and that combining MRI, clinical, and lesion-volume features produces a robust and interpretable framework for predicting three-month functional outcomes in AIS patients.
Multi-class segmentation of the aorta in computed tomography angiography (CTA) scans is essential for diagnosing and planning complex endovascular treatments for patients with aortic dissections. However, existing methods reduce aortic segmentation to a binary problem, limiting their ability to measure diameters across different branches and zones. Furthermore, no open-source dataset is currently available to support the development of multi-class aortic segmentation methods. To address this gap, we organized the AortaSeg24 MICCAI Challenge, introducing the first dataset of 100 CTA volumes annotated for 23 clinically relevant aortic branches and zones. This dataset was designed to facilitate both model development and validation. The challenge attracted 121 teams worldwide, with participants leveraging state-of-the-art frameworks such as nnU-Net and exploring novel techniques, including cascaded models, data augmentation strategies, and custom loss functions. We evaluated the submitted algorithms using the Dice Similarity Coefficient (DSC) and Normalized Surface Distance (NSD), highlighting the approaches adopted by the top five performing teams. This paper presents the challenge design, dataset details, evaluation metrics, and an in-depth analysis of the top-performing algorithms. The annotated dataset, evaluation code, and implementations of the leading methods are publicly available to support further research. All resources can be accessed at https://aortaseg24.grand-challenge.org.
Whole slide images (WSIs) in digital histopathology are acquired at discrete magnification levels encoding complementary diagnostic information from global tissue architecture to fine-grained cellular morphology. Yet, deep learning models remain sensitive to scale variation. Existing magnification-invariant methods rely on multi-scale architectures at predefined discrete resolutions, while in clinical deployment the acquisition magnification varies continuously, rarely aligns with a model's fixed training resolution, and intermediate scales are common, so robust coverage otherwise demands a costly ensemble of magnification-specific models. We propose Conditional Layer Normalization (CLN), a lightweight mechanism that generates affine normalization parameters from input pixel size via a small MLP, integrated into standard CNN architectures for both WSI classification and segmentation. Trained on patches sampled continuously across a range of pixel sizes, the model decouples inference from scanner-dependent magnification and generalizes to arbitrary, previously unseen scales at test time. On the PANDA prostate cancer dataset, our approach on average matches or exceeds independently trained single-magnification models and ranks among the top three performers at every evaluated magnification, including those unseen during training. This collapses a five-model ensemble into a single network and reduces training, and inference cost roughly 4-5 times, while leaving the multiply-accumulate count unchanged. The code is available at: https://github.com/aflorkowska/OneModelToMagnifyThemAll.
Electroencephalography (EEG) is establishing itself as an important, low-cost, noninvasive diagnostic tool for the early detection of Parkinson’s Disease (PD). In this context, EEG-based Deep Learning (DL) models have shown promising results due to their ability to discover highly nonlinear patterns within the signal. However, current state-of-the-art DL models suffer from poor generalizability caused by high inter-subject variability. This high variability underscores the need for enhancing model generalizability by developing new architectures better tailored to EEG data. This paper introduces TransformEEG, a hybrid Convolutional-Transformer designed for Parkinson’s disease detection using EEG data. Unlike transformer models based on the EEGNet structure, TransformEEG incorporates a depthwise convolutional tokenizer. This tokenizer is specialized in generating tokens composed of channel-specific features, which enables more effective feature mixing within the self-attention layers of the transformer encoder. To evaluate the proposed model, four public datasets comprising 290 subjects (140 PD patients, 150 healthy controls) were harmonized and aggregated. A 10-outer, 10-inner Nested-Leave-N-Subjects-Out (N-LNSO) cross-validation was performed to provide an unbiased comparison against seven other consolidated EEG deep learning models. TransformEEG achieved the highest balanced accuracy’s median (78.45 %) as well as the lowest interquartile range (6.37 %) across all the N-LNSO partitions. When combined with data augmentation and threshold correction, median accuracy increased to 80.10 %, with an interquartile range of 5.74 %. In conclusion, TransformEEG produces more consistent and less skewed results. It demonstrates a substantial reduction in variability and more reliable PD detection using EEG data compared to the other investigated models.
BACKGROUND:Computed tomography (CT) is widely used in clinical practice due to its ability to provide detailed anatomical information. However, variations in radiation dose can affect image quality, potentially compromising the performance and reliability of artificial intelligence (AI) models applied to these images. PURPOSE:To evaluate the robustness of radiomics-based and deep learning-based models to variations in CT dose levels using a standardized dataset obtained from a 3D-printed anthropomorphic phantom simulating liver tissue with anomalies, as well as in the publicly available dataset CT-ORG with real patient data for organ classification. This study is in an early experimental stage, tested only on retrospective data. METHODS:A total of 1378 image series from 649 scans were acquired across 13 scanners from four manufacturers at five dose levels. Features were extracted from six regions of interest (ROIs), representing four liver tissue types (normal, cyst, hemangioma, metastasis), using four methods: PyRadiomics, a shallow convolutional neural network (CNN), SwinUNETR, and a CT foundation model (CT-FM). Feature stability was assessed using the Intraclass Correlation Coefficient (ICC), while Uniform Manifold Approximation and Projection (UMAP) was employed to evaluate tissue types separability and the influence of scanner variations. Generalizability was tested by training liver tissue classifiers on one dose level and testing on others, alongside a dose classification task (10-fold cross-validation) to determine the sensitivity of each method to dose variations. In addition, we compared the four methods in addressing the task of organ classification (10-fold cross-validation) with the CT-ORG dataset containing 140 CT scans acquired with varying dose levels. RESULTS:Radiomic features showed limited robustness to dose variations, leading to reduced performance in liver tissue classification and the lowest ICC among methods (ICC: 0.8355 ± $\pm$ 0.1705). SwinUNETR and CT-FM exhibited the highest stability (SwinUNETR ICC: 0.9528 ± $\pm$ 0.0272; CT-FM ICC: 0.9347 ± $\pm$ 0.0420), clearly above the Shallow CNN (ICC: 0.8416 ± $\pm$ 0.2018). CT-FM also showed strong generalization across dose levels: its features effectively distinguished between liver tissue types and dose levels simultaneously, without compromising performance in either task. Consistent with these trends in dose sensitivity, CT-FM obtained the highest dose-classification accuracy (0.6517 ± $\pm$ 0.0179), whereas SwinUNETR showed the lowest (0.3796 ± $\pm$ 0.0250). These trends were confirmed in the context of organ classification with real patient data on the CT-ORG dataset, where CT-FM achieved the highest accuracy (0.965). CONCLUSIONS:The study highlights the limited robustness of traditional radiomics and deep models to CT dose variation and underscores the potential of foundation models like CT-FM to enable robust clinical applications by mitigating dose-related variability. This enhanced performance is likely due to the model's pretraining on large and diverse datasets, allowing it to learn robust and generalizable representations across varying acquisition conditions.
Deep learning-based denoising has enabled substantial reductions in radiotracer dose for PET imaging, yet the absence of full-dose ground truth presents challenges in evaluating image fidelity. This study proposes a supervised classification framework to assess the similarity between AI-denoised and standard-dose PET images without relying on ground truth. A total of 872 features, including radiomic descriptors, quantitative metrics, and visual quality assessments, were extracted at patient and lesion levels. Binary classifiers were trained to predict image similarity across six dose levels. Ground truth labels were derived from a structured visual scoring grid co-developed with expert users and validated against objective similarity metrics. Explainability was integrated using global and local SHapley Additive exPlanations (SHAP), calibrated similarity scores, and textual feature descriptions. The most predictive features included anatomical detail, artefact presence and textural radiomics in the lungs (patient level), and SUV skewness, detection score, and texture strength (lesion level). The combined approach enables interpretable evaluation of AI-denoised PET reconstructions in settings where objective references are lacking, supporting transparency and clinical trust.
Cortical lesions (CLs) have emerged as valuable biomarkers in multiple sclerosis (MS), offering high diagnostic specificity and prognostic relevance. However, their routine clinical integration remains limited due to subtle magnetic resonance imaging (MRI) appearance, challenges in expert annotation, and a lack of standardized automated methods. We present a multi-centric comparative study of CL detection and segmentation in MRI. A total of 656 MRI scans, including clinical trial and research data from four institutions, were acquired at 3T and 7T using MP2RAGE and MPRAGE sequences with expert-consensus annotations. We rely on the self-configuring nnU-Net framework, designed for medical imaging segmentation, and propose adaptations tailored to the improved CL detection. We evaluated model generalization through out-of-distribution testing, demonstrating promising lesion detection capabilities with an F1-score of 0.64 and 0.5 in and out of the domain, respectively. We also analyze internal model features and model errors for a better understanding of AI decision-making. Our study examines how data variability, lesion ambiguity, and protocol differences impact model performance, offering future recommendations to address these barriers to clinical adoption. Furthermore, we designed and implemented a medical expert questionnaire for better assessment of clinical value of the model predictions. To reinforce the reproducibility, the implementation and models will be publicly accessible and ready to use at GitHub and Zenodo.
Information extraction (IE) from clinical texts is increasingly important in healthcare, yet reporting practices remain inconsistent. Existing guidelines do not fully address the unique challenges of IE studies. To develop CINEX, a consensus-based reporting guideline for studies on clinical IE. CINEX is developed following established guideline methodology, including a three-round electronic Delphi (eDelphi) study with domain experts and a final in-person consensus meeting. A preliminary set of 28 items was drafted from a scoping review and existing frameworks. The draft guideline includes five key dimensions: information model, architecture, data, annotation, and outcome. This draft guideline will be refined through the eDelphi process starting May 2025. CINEX will provide structured, expert-driven guidance for reporting clinical IE studies, improving transparency, reproducibility, and comparability.
Explainable Artificial Intelligence (XAI) encompasses a broad spectrum of methods that aim to enhance the transparency of deep learning models, with Class Activation Mapping (CAM) methods widely used for visual interpretability. However, systematic evaluations of these methods in veterinary radiography remain scarce. This study presents a comparative analysis of eleven CAM methods, including GradCAM, XGradCAM, ScoreCAM, and EigenCAM, on a dataset of 7362 canine and feline X-ray images. A ResNet18 model was chosen based on the specificity of the dataset and preliminary results where it outperformed other models. Quantitative and qualitative evaluations were performed to determine how well each CAM method produced interpretable heatmaps relevant to clinical decision-making. Among the techniques evaluated, EigenGradCAM achieved the highest mean score and standard deviation (SD) of 2.571 (SD = 1.256), closely followed by EigenCAM at 2.519 (SD = 1.228) and GradCAM++ at 2.512 (SD = 1.277), with methods such as FullGrad and XGradCAM achieving worst scores of 2.000 (SD = 1.300) and 1.858 (SD = 1.198) respectively. Despite variations in saliency visualization, no single method universally improved veterinarians’ diagnostic confidence. While certain CAM methods provide better visual cues for some pathologies, they generally offered limited explainability and didn’t substantially improve veterinarians’ diagnostic confidence.
OBJECTIVE:Artificial intelligence (AI) is increasingly used in radiology, but its environmental implications have not been sufficiently studied, so far. This study aims to synthesize existing literature on the environmental sustainability of AI in radiology and highlights strategies proposed to mitigate its impact. METHODS:A scoping review was conducted following the Joanna Briggs Institute methodology. Searches across MEDLINE, Embase, CINAHL, and Web of Science focused on English and French publications from 2014 to 2024, targeting AI, environmental sustainability, and medical imaging. Eligible studies addressed environmental sustainability of AI in medical imaging. Conference abstracts, non-radiological or non-human studies, and unavailable full texts were excluded. Two independent reviewers assessed titles, abstracts, and full texts, while four reviewers conducted data extraction and analysis. RESULTS:The search identified 3,723 results, of which 13 met inclusion criteria: nine research articles and four reviews. Four themes emerged: energy consumption (n = 10), carbon footprint (n = 6), computational resources (n = 9), and water consumption (n = 2). Reported metrics included CO2-equivalent emissions, training time, power use effectiveness, equivalent distance travelled by car, energy demands, and water consumption. Strategies to enhance sustainability included lightweight model architectures, quantization and pruning, efficient optimizers, and early stopping. Broader recommendations encompassed integrating carbon and energy metrics into AI evaluation, transitioning to cloud computing, and developing an eco-label for radiology AI systems. CONCLUSIONS:Research on sustainable AI in radiology remains scarce but is rapidly growing. This review highlights key metrics and strategies to guide future research and practice toward more transparent, consistent, and environmentally responsible AI development in radiology. ABBREVIATIONS:AI, Artificial intelligence; CNN, Convolutional neural networks; CT, Computed tomography; CPU, Central Processing Unit; DL, Deep learning; FLOP, Floating-point operation; GHG, Greenhouses gas; GPU, Graphics Processing Unit; LCA, Life Cycle Assessment; LLM, Large Language Model; MeSH, Medical Subject Headings; ML, Machine learning; MRI, Magnetic resonance imaging; NLP, Natural language processing; PUE, Power Usage Effectiveness; TPU, Tensor Processing Unit; USA, United States of America; ViT, Vision Transformer; WUE, Water Usage Effectiveness.