
Accurate outer-to-outer wall aortic segmentation is essential for abdominal aortic aneurysm (AAA) assessment and endovascular aneurysm repair (EVAR) planning, yet manual segmentation is time-consuming and impractical for routine use. This study aimed to validate a pre-trained open-source nnU-Net–based model for fully automated outer-to-outer wall segmentation of the abdominal aorta in AAA patients. In this retrospective multicentre observational validation study, 75 consecutive AAA patients (25 per centre) from three Dutch hospitals were included. Manual segmentations of pre-EVAR CTA scans were performed. A second observer independently segmented 30 cases to assess interobserver variability. Manual and automatic segmentations were compared using the Dice similarity coefficient (DSC), Jaccard index, maximum and 95th percentile Hausdorff distances (HD max, HD95), mean surface distance, and absolute and relative volume difference. Agreement was assessed for the full abdominal aorta and the infrarenal neck separately. Automatic segmentation achieved excellent agreement with manual segmentations, with a median DSC of 0.97 (IQR 0.96–0.97) and median HD95 of 5.5 (IQR 3.8–7.6) mm. The mean volume bias was 2.06 ml. Performance was comparable to interobserver variability between manual segmentations (median DSC 0.96 [IQR 0.95–0.97]). For the infrarenal neck, a median DSC of 0.95 (IQR 0.93–0.96) was achieved. Automatic segmentation required less than one minute per scan on a dedicated GPU, compared to 45–60 minutes for manual segmentation. Automated outer-to-outer wall aortic segmentation using a pre-trained nnU-Net–based model achieves excellent agreement with manual segmentations, comparable to interobserver variability, across a consecutive multicentre dataset. This approach may support more reproducible and efficient AAA analysis in clinical and research settings.
Deep learning methods are becoming increasingly important for medical image segmentation. However, previous approaches often performed poorly in spatial-frequency-domain fusion and feature fusion, leading to inaccurate lesion localization, hazy boundaries, and low accuracy. This paper proposes a reliability-guided fusion and dynamic dual-domain mask-modulated Shearlet network for medical image segmentation (RF-DMSNet for short). Its two primary components are the reliability-guided multi-scale feature fusion module (RGMF) and the dynamic dual-domain mask with a coordinate attention-modulated Shearlet operator (DAMS). The former module exhibits reliability in both the channel and spatial dimensions, allowing adaptive weighted fusion to increase feature quality. Meanwhile, the latter uses a dynamic dual-domain mask and coordinate attention to improve lesion boundary and location segmentation performance. The frequency-domain re-enhancement module (FREM) and critical feature guided module (CFGM) are designed using DAMS. These two modules work together to optimize a unified, detailed enhancement module (DEM), enabling progressive refinement of segmentation predictions from coarse to fine. Extensive experimental results demonstrate that RF-DMSNet achieves superior segmentation performance compared to state-of-the-art approaches for medical image segmentation.
Medical image classification has advanced substantially with convolutional neural networks (CNNs), Vision Transformers (ViTs), and hybrid CNN–ViT architectures, yet clinical translation remains limited by dataset dependency, inconsistent evaluation practices, and insufficient external validation. This review provides a clinically grounded comparative synthesis of these model families across diverse medical imaging modalities. A structured literature review was conducted across PubMed, IEEE Xplore, Scopus, and Web of Science for studies published between 2016 and 2025. Following predefined eligibility criteria, 81 studies were included in the qualitative review, of which 74 contributed to a dataset-aware descriptive quantitative aggregation. The quantitative synthesis was therefore restricted to descriptive aggregation; a formal meta-analysis was not performed because of substantial methodological heterogeneity and insufficient reporting of study-level variance information across the included studies. CNN-based models demonstrated the most consistent performance, achieving the highest weighted accuracy (0.934) and weighted recall (0.901). ViT-based models achieved competitive weighted accuracy (0.906) and recall (0.893), particularly for OCT and X-ray imaging, but appeared more sensitive to dataset scale and quality. Hybrid CNN–ViT models achieved a weighted accuracy of 0.814 and weighted recall of 0.698, with the greatest performance variability. Only 10 of the 81 reviewed studies (12.3
Automated Alzheimer’s disease (AD) classification from structural MRI typically employs either feature-engineered machine learning (ML) or end-to-end 3D deep learning (DL). Addressing a lack of rigorous statistical comparison and multi-dataset validation in current literature, this study evaluates classical ML algorithms utilizing volumetric and voxel-based morphometry (VBM) features against 3D CNNs processing raw MRI tensors. Using the ADNI and OASIS datasets, models underwent internal 4-fold cross-validation and zero-shot cross-cohort testing to assess predictive performance, domain shift resilience, and clinical sensitivity. Welch’s t-tests confirmed that 16 of 19 macro-regional brain aspects differed significantly between AD and cognitively normal subjects ( p_FDR < 0.05 ), led by the medial-temporal lobe. While DL architectures achieved marginal numerical superiority in peak accuracy on ADNI (DenseNet121: accuracy 0.92 ± 0.02 , F1 0.85 ± 0.04 , ROC-AUC 0.96 ± 0.02 ), they exhibited higher variance and lacked statistical significance compared to optimized ML baselines (XGBoost-VOL: F1 0.82 ± 0.03 ; p ≥ 0.28 for all 3D CNNs). On the highly imbalanced OASIS dataset, VBM-enhanced Logistic Regression matched top DL models in overall F1-score ( 0.68 ± 0.05 ) while delivering statistically superior AD recall ( 0.67 ± 0.04 ; p = 0.0026 versus baseline), exceeding every 3D CNN. Although both paradigms generalized robustly across cohorts (only a 2–4
Clinical histories are essential for accurate radiologic interpretation but are often lengthy and time-consuming to review amid increasing radiologist workloads. Large language models (LLMs) offer a potential solution by generating concise, clinically focused summaries from existing documentation; however, the optimal presentation format for radiologists remains unclear. This retrospective reader study evaluated the readability, efficiency, and radiologist preference for a two-stage LLM-based summarization pipeline designed to identify the optimal summary format for radiologic use. Ninety imaging studies across multiple modalities and care settings were processed to generate structured timeline summaries (stage 1) and brief narrative summaries (stage 2). Text length was reduced by approximately 86.8
Cerebral aneurysms are complex vascular lesions whose rupture can result in subarachnoid hemorrhage (SAH), a condition associated with substantial morbidity and mortality. Although current clinical risk assessment relies primarily on morphological characteristics such as aneurysm size and location, these factors alone are often insufficient for individualized rupture risk prediction. Advances in artificial intelligence (AI) and machine learning (ML) have enabled the integration of imaging, clinical, morphological, and computational data, creating new opportunities to improve aneurysm detection, risk stratification, and treatment planning. This systematic review evaluates recent applications of AI and ML across the cerebral aneurysm clinical pipeline, focusing on three major domains: (i) automated detection and segmentation from medical imaging, (ii) rupture risk prediction using clinical, morphological, radiomic, and computational fluid dynamics (CFD)-derived hemodynamic features, and (iii) clinical decision support for treatment planning and outcome prediction. The reviewed studies are examined with respect to their methodological approaches, input features, predictive performance, validation strategies, and clinical applicability. Across the literature, multimodal models that integrate heterogeneous data sources generally demonstrate superior predictive performance compared with approaches relying on a single feature category. However, widespread clinical implementation remains limited by retrospective study designs, heterogeneous datasets, inconsistent validation practices, limited external validation, and challenges related to model interpretability and generalizability. Emerging directions, including explainable AI, multimodal learning, and physics-informed ML, offer promising opportunities to improve model robustness and facilitate clinical translation. Overall, the reviewed evidence indicates that AI has considerable potential to support the detection, risk assessment, and management of cerebral aneurysms, while emphasizing the need for standardized datasets, prospective multicenter validation, and interpretable models to enable reliable clinical adoption.
Myocardial scar is a critical pathological substrate, often identified using cardiac magnetic resonance (CMR) as the gold standard. However, CMR accessibility is limited by cost, time, and expertise requirements. This study aimed to develop an artificial intelligence (AI)-powered dual-sequence computed tomography (CT) framework integrating coronary CT angiography (CCTA) and coronary artery calcium (CAC) for detecting myocardial scar, using CMR as the reference. We retrospectively included patients who underwent CCTA, CAC, and CMR. A deep learning nnU-Net model enabled automated cardiac segmentation and rule-based subdivision following the American Heart Association 17-segment model. Radiomic features were extracted from CCTA and CAC, along with volumetric indices, stenosis grading, and clinical/echocardiographic variables. A total of 533 patients were included. CCTA radiomics achieved an AUC of 0.85 at the patient level. Dual-sequence analyses were performed in 257 patients. In this subcohort, the integrated CT model achieved 0.87 and the multimodal model combining CT, clinical, and echocardiographic features achieved an AUC of 0.91. At the segmental level, the combined CCTA-CAC model yielded an AUC of 0.80 among patients with CMR-confirmed myocardial scar. AI-derived dual-sequence CT features demonstrated good discriminatory performance for detecting myocardial scar at the patient level and identifying scar-positive regions at the segmental level. By integrating CT features with clinical and echocardiographic variables, the multimodal model achieved favorable patient-level performance. These findings support the potential role of the proposed CT-based framework in providing complementary scar-related information and identifying high-risk patients who may benefit from further CMR evaluation in appropriate clinical settings.
The published performance of artificial intelligence (AI) models in radiology is typically based on the reporting of sensitivity, specificity, and receiver operating characteristic area under the curve, both in the peer-reviewed literature and for Food and Drug Administration 510(k) submissions. Interestingly, these metrics cannot inform radiologists, the users of these systems, of the rate or quantity of each type of error to anticipate if they implement the candidate AI product(s) into their own clinical practice. Only the positive predictive value (PPV) and negative predictive value (NPV), or rather, their complements, the false discovery rate (1-PPV, FDR), and the false omission rate (1-NPV, FOR) can provide these error rates. Although some published articles and 510(k) submissions include the PPV and NPV of the concerned AI models, many do not, and the ones that do sometimes test on enriched datasets with artificially high prevalence rates of the target condition, thus inflating PPV and deflating NPV that would be found in clinical practice. This manuscript demonstrates how clinical practices can estimate and evaluate an AI's FDR and FOR for their clinical population using Bayes' Theorem. We also propose a risk-based evaluation matrix (RADDE) which allows radiologists to consider the medical, legal, financial, workflow, psychological, and reputational impact of these estimated AI error rates.
Enlarged perivascular spaces (ePVS) are small, sparse MRI-visible markers of cerebral small vessel disease and brain ageing. Their size, low contrast, and severe foreground–background imbalance make automated segmentation challenging. Existing methods are mainly cross-sectional and do not model temporal consistency, limiting their utility for tracking longitudinal change. We developed Long-PVSUNet, a two-timepoint framework for count-oriented ePVS segmentation on paired baseline/follow-up T1-weighted and FLAIR MRI using baseline-guided attention fusion and imbalance-aware training. After screening/QC, longitudinal data were available from UK Biobank (UKB; n = 4568), ADNI (n = 277), and MAS (n = 403). Expert dot annotations were obtained in labelled UKB, ADNI, and MAS subsets (n = 250, 100, 100) for basal ganglia (BG) and centrum semiovale, regions commonly used for clinically relevant ePVS counting. Performance and generalisability were evaluated against cross-sectional and longitudinal baselines using Dice, lesion-level F1-localisation, and cross-cohort, few-shot, ablation, and data-efficiency analyses. Exploratory analyses tested hypertension associations with model-derived ePVS counts and progression in UKB/MAS. Long-PVSUNet achieved strong UKB performance (Dice 0.802 ± 0.010, F1 0.842 ± 0.011) and generalised to ADNI/MAS. It remained data-efficient with fewer labels. Attention fusion significantly outperformed mean and temporal-difference fusion (Dice 0.800 ± 0.004). Long-PVSUNet outperformed cross-sectional U-Net and longitudinal benchmark model, suggesting its utility for longitudinal ePVS segmentation. In clinical analysis, hypertension was associated with faster BG-ePVS progression in UKB (incidence rate ratio (IRR) = 1.06, 95
Post hoc attribution is the standard transparency layer for deep learning in imaging decision support, yet its two dominant families capture conflicting evidence: activation maps are spatially coherent but coarse, while gradient-based attributions are precise but unstable. Gradient-weighted activation maps also show localisation quality that scales with lesion size, so explanation reliability degrades for the smallest lesions. We address both problems with Evidential Dempster–Shafer Fusion (EDSF), which combines an activation-based spatial prior with voxel-level path-integral attributions inside the belief-function formalism. The framework is instantiated on 3D pulmonary nodule analysis with a 3D ResNet-18 trained on LUNA16, reaching 97.2
Large-scale chest X-ray report datasets are widely used to train multimodal models for medical report generation. However, these datasets often contain longitudinal or comparative expressions that implicitly refer to prior examinations. When the corresponding prior images are unavailable, such references create dataset-level inconsistencies and may propagate non-self-contained language patterns to downstream models. To improve report self-containment, this study develops an automated, data quality–oriented preprocessing pipeline that rewrites radiology reports at the sentence level to reduce external references. The pipeline uses an open-source large language model fine-tuned with a parameter-efficient method on a small set of reports manually revised according to predefined editing guidelines that preserve grammatical coherence. The pipeline is applied to the MIMIC-CXR dataset. Content containing external references is quantified using keyword-based statistics and a discriminative detection model. The edited dataset is then used to fine-tune a vision–language model for report generation. External references are substantially less prevalent in the edited reports than in reports from the original MIMIC-CXR dataset and a previously published cleaned dataset. The vision–language model fine-tuned on the edited dataset produces fewer statements that rely on unavailable longitudinal context. Automated report editing can therefore serve as an effective preprocessing step for improving the suitability of chest X-ray report datasets for multimodal model development. The proposed pipeline is open source, computationally efficient, and suitable for fully local deployment in privacy-sensitive medical data environments.
Removing patient-identifying information from medical images is a prerequisite for sharing image data directly, as in public dataset release and open benchmarks, where the images themselves, rather than only model updates must leave the originating institution. However, many methods currently used for de-identification, e.g., cropping or blacking out image regions to eliminate burned-in text, can have negative effects on downstream image analysis tasks because of removal of relevant but non-identifiable information. This work presents an end-to-end deep learning framework for transforming raw clinical image volumes into de-identified, analysis-ready datasets without compromising downstream utility. The methodology developed and tested in this work first detects and redacts regions likely to contain protected health information (PHI), such as burned-in text and metadata, and then uses a generative deep learning model to inpaint the redacted areas with anatomically and imaging-plausible content. The proposed pipeline leverages a lightweight hybrid architecture, combining CRNN-based redaction with a latent-diffusion inpainting restoration module (Stable Diffusion 2). We evaluate the approach using both privacy-oriented metrics, which quantify residual PHI and success of redaction, and image-quality and task-based metrics, which assess the fidelity of restored volumes for representative deep learning applications. The binary mask performance shows strong overall PHI identification (F1 score = 0.891 ± 0.037, recall of 0.912 ± 0.053, and precision of 0.875 ± 0.058, indicating accurate localization of PHI-containing regions with few missed detections or false-positive redactions. Downstream anatomy identification (segmentation) tasks remain markedly similar across after applying varied inpainting strategies (Diffusion with/without context and Telea with Dice ranging from 0.936 to 0.948 on one dataset and 0.955 to 0.959 on another. Our results suggest that the proposed method yields de-identified medical images that are visually coherent, maintaining fidelity for downstream models and clinical tasks, while substantially reducing the risk of patient re-identification. By automating de-identification and image reconstruction within a single workflow and disseminating large-scale medical imaging collections, thereby lowering a key barrier to data sharing and multi-institutional collaboration in medical imaging AI.
Digital subtraction angiography (DSA) roadmap imaging is widely used in endovascular interventions, but conventional mask-based subtraction is often affected by device contamination and motion-related misregistration. Existing mask-free methods commonly rely on unconstrained image synthesis, with limited interpretability and temporal stability. This study developed an interpretable virtual subtraction framework that uses temporal information to improve guidewire visualization and background reconstruction without a physical mask. A temporal-consistency segmentation-inpainting framework was developed to generate device-only subtraction images from consecutive fluoroscopic frames. A motion-adaptive spatiotemporal attention module adjusted temporal feature aggregation according to inter-frame guidewire motion. Historical frames were aligned and combined through confidence-weighted aggregation to guide mean-reverting stochastic differential equation-based background inpainting. An explicit temporal-consistency objective was used to reduce fluctuations between consecutive reconstructed backgrounds. The framework was evaluated on clinical fluoroscopic sequences acquired at three hospitals through comparisons with recent guidewire segmentation and synthetic DSA generation methods, component-wise ablation studies, sensitivity analyses, and quantitative image-quality assessment. The proposed segmentation network achieved a Dice coefficient of 0.917 and a guidewire breakage rate of 3.7
Noninvasive identification of advanced liver fibrosis (ALF), commonly defined as Scheuer stage S3 or higher, is important for the management of patients with liver disease, as it indicates an increased risk of progression to cirrhosis and helps determine the need for closer monitoring and timely intervention. In this retrospective study, we developed and evaluated clinical, deep learning (DL), radiomics, and multimodal fusion models for the noninvasive classification of advanced liver fibrosis using non-contrast magnetic resonance imaging (MRI) combined with clinical data. A total of 366 patients from three clinical centers between May 2018 and June 2024 were included. Radiomics and DL features were extracted from preoperative multiparametric non-contrast MRI, including T1-weighted imaging (T1WI), T2-weighted imaging (T2WI), and diffusion-weighted imaging (DWI). Three single-modality models, namely clinical, radiomic, and DL models, were constructed, and two multimodal fusion models were further developed, including one based on feature-level fusion and the other on decision-level integration. Model performance was evaluated using the area under the receiver operating characteristic curve (AUC), calibration curves, and decision curve analysis (DCA). Among all models, the decision-level fusion model achieved the best overall performance, with an AUC of 0.7891 in the internal validation cohort and 0.7851 in the external validation cohort. In addition, it demonstrated superior clinical net benefit and good calibration compared with the feature-level fusion, clinical, radiomic, and DL models. Overall, the proposed decision-level fusion model showed potential for the noninvasive identification of advanced liver fibrosis by integrating routine MRI and clinical data; however, its external validation performance indicates that further refinement and validation are needed before clinical application.
Research on foundation models is actively progressing. The segment anything model (SAM) and MedSAM are representative foundation models for image segmentation. Recently, low-rank adaptation (LoRA) has been developed, allowing parameter updates without retraining the entire model, thus solving the problem with large data and time required for task-specific fine-tuning. Although many studies have used public databases, few have focused on local data. Moreover, to our knowledge, no studies have fine-tuned MedSAM using LoRA. We aimed to evaluate SAM, MedSAM, and their LoRA-tuned variants (SAM-LoRA and MedSAM-LoRA) using brain magnetic resonance images of gliomas from five centers in Japan and to compare their performance. We used 2D-based fluid-attenuated inversion recovery axial images and conducted parameter optimization based on four-fold cross-validation (189 cases) and external test evaluation (75 cases) using cases collected retrospectively. Dice coefficients, intersection over union (IoU), and the 95
Brain single-photon emission computed tomography (SPECT) imaging using I-123 DaTSCAN is an effective tool for the diagnosis and follow-up of Parkinson disease. Reducing acquisition time decreases the likelihood of patient motion, improves patient comfort, and increases scanner throughput. However, shorter acquisition times often result in degraded image quality with reduced diagnostic value. This study aimed to use deep learning (DL) approaches to enhance the image quality of accelerated SPECT DaTSCAN acquisitions. A total of 85 patients were scanned on a CZT-based GE StarGuide scanner with a standard acquisition time of 15 min. To simulate fast acquisition protocols, raw list-mode data were retrospectively undersampled to represent 20
The diagnosis of temporomandibular joint (TMJ) anterior disc displacement primarily relies on magnetic resonance imaging (MRI). Clinical assessment depends heavily on manual segmentation by radiologists, which is time-consuming and highly subjective. We aimed to develop and validate a novel PE AttU-Net for accurate automatic multi-structure segmentation of the articular disc and condyle on TMJ MRI images. The proposed PE AttU-Net integrates attention mechanisms, boundary enhancement modules, and small-target optimization strategies to achieve joint multi-structure segmentation. A total of 512 TMJ MRI images from 334 patients covering five displacement grades (normal, slight, mild, moderate, severe) were collected and preprocessed uniformly. Model validation, quantitative evaluation, and comparative experiments were performed to verify the model’s performance. Evaluated using the five-grade dataset, the model obtained satisfactory segmentation results for the articular disc. The Dice coefficients were 78.44
Prostate-specific membrane antigen positron emission tomography/computed tomography (PSMA PET/CT) provides rich molecular imaging across the prostate cancer care continuum, yet its full utility in artificial intelligence applications remains underexplored. This study developed a multi-task UNETR–based deep learning framework to simultaneously perform lesion segmentation, disease burden quantification, CHAARTED and LATITUDE treatment stratification and survival prediction from a single PSMA PET/CT image. A retrospective cohort study included 212 prostate cancer cases (478 observations, 2018–2024) from The Brunei Cancer Centre, with 160 anonymized whole-body PSMA PET/CT scans and longitudinal clinical data, following the TRIPOD-AI checklist. Preprocessing included bounding-box cropping, Z-score normalization and Otsu-based lesion masking. A simple convolutional neural network (CNN) and a 2D multi-task UNETR model were trained for lesion segmentation, treatment classification and survival prediction. Performance was evaluated using Dice coefficient, multiclass area under the curve (AUC) and concordance index (C-index). Analysis of 205 matched PSMA PET/CT scans from 115 patients over 6 years showed that the CNN achieved strong treatment classification performance (AUC = 0.91), with highest accuracy for active surveillance (AUC = 0.99) and chemotherapy (AUC = 0.97). The multi-task UNETR produced a lower weighted AUC (0.561) but enabled simultaneous lesion segmentation, metastatic burden quantification and CHAARTED/LATITUDE stratification. Deceased patients showed higher Dice scores (mean = 0.556, SD = 0.065) than survivors (mean = 0.324, SD = 0.165). Multi-task UNETR improved clinical interpretability through segmentation-based tumour quantification and disease burden stratification. Larger multicentre datasets and external validation are needed to enhance generalizability and predictive performance.
To assess radiation burden and contrast medium load, image quality, and the performance of AI-reader and junior radiologists in evaluating coronary stenosis using 80-kVp triple-rule-out (TRO) CT with deep learning image reconstruction (DLIR), 159 patients scheduled for TRO were prospectively recruited and randomized into group A (n = 80; 80-kVp with DLIR) or group B (n = 79; 100-kVp with adaptive statistical iterative reconstruction-V [ASIR-V 60
Assessment of root canal treatment quality is central to improving endodontic outcomes, yet interpretation of periapical radiographs remains susceptible to observer variability, fatigue, and anatomical complexity. To provide a more objective and efficient approach to post-endodontic evaluation, we developed MS-WEFNet, an enhanced YOLOv10n-based object detection framework for automatic assessment of root canal filling quality. The network incorporates a Multi-Scale Edge Information Enhancement module to reinforce fine anatomical boundaries and a Wavelet-guided Contrastive Feature Aggregation module to improve multi-scale structural representation. MS-WEFNet was evaluated on the T2k dataset, which comprised 2564 periapical radiographs containing 6772 annotated teeth. Experienced endodontists assigned the teeth to three clinical categories: 4032 non-target teeth, 1417 qualified teeth, and 1323 teeth requiring warning. The warning category encompassed underfilling, overfilling, and instrument separation. The dataset was partitioned into training, validation, and testing subsets at a ratio of 7:2:1. The proposed model achieved an mAP _50 of 0.902 ± 0.003 and an mAP _50 – _95 of 0.817 ± 0.005, representing improvements of 4.0