
MotivationArtificial intelligence (AI) is reshaping radiology through advances in machine learning, deep learning, and generative AI. As these technologies become embedded in diagnostic workflows, radiology education must adapt to prepare learners for AI-enabled clinical practice. This review aimed to synthesize current evidence, identify educational priorities, and highlight challenges and opportunities for implementation.IntroductionThe growing adoption of AI in medical imaging requires radiology curricula to extend beyond image interpretation and encompass AI literacy, ethical reasoning, critical appraisal, and human–AI collaboration. Understanding how AI is currently incorporated into radiology education is essential for developing effective training frameworks. This review examined educational approaches, learner outcomes, and implementation challenges associated with AI in radiology education.MethodologyA scoping review was conducted and reported in accordance with the Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews (PRISMA-ScR), using the Population–Concept–Context(PCC) framework to define the review question and eligibility criteria. Searches were performed in Scopus, PubMed, and IEEE Xplore for studies published between 2020 and 2025. After screening and eligibility assessment, 29 original studies were included. Educational outcomes were classified into skill development, engagement, and performance domains.Results and discussionSkill development was the most frequently investigated outcome (55.2%), followed by performance (27.6%) and engagement (17.2%). Approximately 86% of studies reported positive or improved educational outcomes. AI-based interventions enhanced learner confidence, AI literacy, diagnostic reasoning, and readiness for clinical implementation. Generative AI tools showed promise for tutoring, assessment, and self-directed learning but raised concerns regarding reliability, hallucinations, bias, and ethical use. Common barriers included limited faculty expertise, insufficient formal training, curricular overcrowding, inadequate infrastructure, and governance challenges.ConclusionAI has significant potential to strengthen radiology education through personalized learning, competency development, and technology-enhanced training. Nevertheless, sustainable implementation requires standardized curricula, faculty development, competency frameworks, and robust ethical oversight. Future research should focus on longitudinal outcomes and evidence-based educational models that prepare radiology professionals for increasingly AI-integrated healthcare systems.
ObjectiveTo evaluate the clinical utility of a wide-detector CT protocol with variable helical pitch (vHP) for one-stop cardiovascular and cerebrovascular CT angiography.MethodsThis prospective randomized study enrolled 61 patients (25 in the Group A and 36 in the Group B) with suspected cardiovascular and cerebrovascular diseases. Group A underwent the integrated one-stop CTA protocol (which incorporates vHP technology), whereas Group B underwent conventional separate scans (coronary CTA and head-neck CTA). Both groups used automated bolus-tracking triggering. Comparisons included effective scanning time, total operating time, contrast medium volume, radiation dose (dose-length product, DLP), and subjective and objective image quality. Normality was assessed using the Shapiro–Wilk test; continuous variables were compared using independent samples t-test or Mann–Whitney U test as appropriate. The false discovery rate method was applied to correct for multiple comparisons, and analysis of covariance was performed to adjust for potential confounders.ResultsThe integrated protocol group (Group A) demonstrated significantly higher contrast-to-noise ratio (CNR) in the coronary arteries and major aortic branches(all FDR-corrected P < 0.001). In the neck region, Group A showed significantly higher SNR in the common carotid arteries (FDR-corrected P < 0.001 for RCCA and P = 0.0044 for LCCA), whereas Group B demonstrated better CNR in the right internal carotid artery (FDR-corrected P < 0.001) and lower image noise in several vessels (FDR-corrected P < 0.05 for RICA SD and LVA SD). Notably, ANCOVA models for SNR exhibited poor fit with adjusted R2 values frequently near zero or negative, warranting cautious interpretation; the primary evidence for protocol-related image quality improvement is therefore based on the robust CNR findings. ANCOVA confirmed that these intergroup differences remained significant after adjustment (all adjusted P < 0.05 for key CNR/SNR comparisons), indicating protocol-related rather than baseline-driven advantages. Subjective image quality was excellent in both groups (median 4, IQR 4–4), with no significant intergroup differences. DLP was higher in Group A (909.00 vs. 694.50 mGy·cm, P < 0.001), but effective scanning time was significantly shorter (8.00 vs. 19.00 s, 57.9% reduction, P < 0.001), and total operating time was markedly reduced (480.00 vs. 1020.00 s, 52.9% reduction, P < 0.001). Contrast medium volume was halved in Group A (50–60 vs. 100–120 mL).ConclusionThe integrated wide-detector CT protocol evaluated in this study (which incorporates variable helical pitch as a key enabling component) is technically feasible and offers workflow advantages over conventional scanning, including shorter scan time, reduced contrast volume, and maintained image quality, with a somewhat higher radiation dose. However, lacking an independent reference standard, its diagnostic accuracy for detecting significant stenosis remains unproven. Further validation with invasive angiography is needed before routine clinical use. However, because the integrated protocol differs from the conventional protocol in multiple parameters simultaneously (detector configuration, rotation speed, pitch, and trigger threshold), the individual contribution of variable helical pitch cannot be isolated from these combined hardware and protocol differences.
IntroductionMolecular subtyping of breast cancer, particularly the discrimination between Luminal A and Basal-like tumours, is critical for guiding treatment decisions and establishing prognosis. Radiomics offers a promising complementary, non-invasive imaging biomarker to support biopsy-based subtyping; however, two methodological gaps limit the validity of existing studies: the absence of controlled 2D-vs.-3D feature comparisons, and the systematic underreporting of patient-level data leakage in multi-slice pipelines.MethodsTo address both limitations, this study benchmarks slice-level (2D) and volumetric (3D) radiomic pipelines for the classification of Luminal A vs. Basal-like tumours on the TCGA-BRCA cohort, evaluating seven classifiers combined with two feature selection strategies under a Stratified Group k-Fold (k = 5) cross-validation scheme that enforces strict patient-level isolation.ResultsThe 3D pipeline achieved a mean AUC of 0.682 vs. 0.631 for the 2D pipeline. In the 3D configuration, the Support Vector Machine with SelectKBest-MI feature selection attained the highest performance (AUC = 0.768 ± 0.066), while the 2D configuration was led by Logistic Regression (AUC = 0.769 ± 0.109), albeit with substantially higher fold-to-fold variability. Six out of seven classifiers yielded higher AUC scores under the 3D configuration, although a paired significance test on fold-level AUC values did not reach statistical significance (p = 0.313).DiscussionThese results establish a reproducible, leakage-free baseline for non-invasive breast cancer subtype classification through radiomics, with 3D radiomics offering improved robustness rather than superior peak performance. Results are derived from a single cohort (TCGA-BRCA) and external validation is required before clinical generalizability can be claimed.
ObjectivesTo develop a contrastive learning model for lung disease classification using discriminative CT imaging embeddings.MethodsA total of 1,187 subjects were included: asthma (n = 315), COPD (n = 355), post-COVID-19 (n = 375), and healthy controls (n = 142). Of these, 1,003 subjects had a single visit with similarly protocoled CT scans acquired at total lung capacity (TLC) and residual volume (RV), and 92 (33 asthma and 59 post-COVID-19) completed a follow-up visit, with two scans per visit. We developed a modified contrastive learning model incorporating an expert-conditioned routing network and adaptive temperature scaling to learn discriminative embeddings. Model performance was evaluated using the area under the receiver operating characteristic curve (AUC). The embeddings were further validated via k-means clustering, and quantitative CT (qCT) metrics were compared across the derived clusters. The embedding space was used to track disease progression or improvement in the follow-up disease subgroup and to evaluate the model's ability to predict qCT metrics, quantified by the coefficient of determination (R2).ResultsThe model achieved a macro-AUC of 89.3% (95% CI: 86.5, 91.8; P < 0.001) in differentiating the four classes. Post-COVID-19 emerged as a distinct class from asthma and COPD in the t-SNE embedding space, and its embeddings across two visits captured disease improvement. Additionally, the learned embeddings showed predictive power for several qCT metrics, particularly the Jacobian (R2=0.61).ConclusionsThe proposed model effectively differentiated these three lung diseases and provided meaningful embeddings for phenotype characterization, longitudinal assessment, and qCT metric prediction.
ObjectivesThis study developed an AI-driven radiomics model fusing pre-radiofrequency ablation (RFA) multi-sequence MRI features and clinical variables to predict early recurrence of hepatocellular carcinoma (HCC) after RFA.MethodsA total of 169 HCC patients who underwent pre-RFA MRI (January 2015–December 2021) were retrospectively enrolled and randomly assigned to a training set (n = 135) and an internal hold-out test set (n = 34) at an 8:2 ratio using stratified sampling (training: 49 recurrent, 86 non-recurrent; test: 12 recurrent, 22 non-recurrent). Radiomics feature selection involved a three-step strategy of variance threshold filtering, SelectKBest, and least absolute shrinkage and selection operator (LASSO) regression. All data preprocessing, feature selection, and model training were performed exclusively within the training set to prevent information leakage. Support vector machine (SVM), logistic regression (LR), and random forest (RF) classifiers were used to construct radiomics-only, clinical-only, and integrated models. Performance was evaluated using ROC curves (AUC as primary metric), calibration curves, Brier scores, Hosmer–Lemeshow test, and decision curve analysis (DCA). The optimal threshold was determined by the Youden index.ResultsAlpha-fetoprotein (AFP), platelet count (PLT), and tumor location were identified as independent clinical predictors (all VIF < 5). Of 11,320 extracted features, 61.5% (n = 6,965) achieved ICC > 0.75, and 16 informative radiomics features were ultimately selected (inter-observer Dice = 0.85 ± 0.06). The radiomics-only models showed robust performance (test AUC: SVM = 0.826, LR = 0.830, RF = 0.826), significantly outperforming clinical-only models (test AUC: SVM = 0.688, LR = 0.706, RF = 0.724; P < 0.05). The integrated models achieved further improvement, with the RF-based integrated model demonstrating superior performance (training AUC = 0.975, test AUC = 0.909; P < 0.0083 after Bonferroni correction). Calibration showed good agreement (Brier score = 0.143, Hosmer–Lemeshow P = 0.367), and DCA demonstrated net clinical benefit across threshold probabilities of 15%–70%. At the optimal threshold (0.42), the RF-based integrated model achieved sensitivity = 83.3%, specificity = 90.9%, PPV = 83.3%, NPV = 90.9%, and accuracy = 88.2%.ConclusionPreoperative MRI-based radiomics features combined with clinical variables effectively predict early post-RFA HCC recurrence. The integrated radiomics-clinical model outperforms both clinical-only and radiomics-only models, with the RF classifier yielding optimal performance, providing a reliable tool for individualized treatment decision-making.
Background and purposeLarge language model (LLM)-assisted radiology research and tool-development work is hampered by a practical form of memory loss we term cross-interface fragmentation: project context built in a web-based conversational interface is not available in a local coding-agent environment, and neither environment retains state across sessions. This forces researchers to manually reconstruct context at every session start and every interface transition. This technical development describes the design and a single-user, proof-of-concept functional demonstration of a Claude-based persistent-memory system intended to reduce this fragmentation for the software-engineering and research-organization layer of radiology AI projects. The system is not a clinical tool and does not process patient data.Materials and methodsWe implemented an end-to-end pipeline integrating a custom web application (Created.app/Supabase), a self-hosted n8n automation server, the Claude API (Anthropic), GitHub version control, and the Obsidian knowledge-management application. A web session is summarized by the Claude API into six structured fields (summary, decisions, pending tasks, artifacts, errors, continuation point) and written as a per-project MEMORY.md file to GitHub; a complementary CLAUDE.md file holds accumulated technical context for the local coding agent. The study design is a single-user, single-operator proof-of-concept engineering validation across three concurrent development projects; no efficacy, usability, or clinical-outcome claim is made.ResultsIn functional testing, the pipeline transferred project context between the web and local-coding environments without manual copy-paste of session content. A representative 558-token session (n = 1 observation) was processed end-to-end with complete six-field extraction; create and edit operations against the GitHub Contents API succeeded across five consecutive runs without overwrite errors. On a fresh launch, the local coding agent read the memory files and reported project identity, prior technical decisions, and pending tasks without user-provided context. Reported timings are single-run observations, not benchmarked measurements.ConclusionA persistent, cross-interface memory layer for LLM-assisted radiology research workflows is technically feasible with currently available components. As a proof of concept it is explicitly limited: single user, Claude-based, text-only, without quantitative extraction-accuracy benchmarking, and dependent on commercial APIs whose cost scales with use (fixed infrastructure plus variable per-session API cost). We discuss a vendor-neutral, open-weights configuration as the recommended direction for cost- and dependency-sensitive settings such as Latin America, and we report that the system has been used in practice to help organize the development of a separate radiology report quality-control tool.
Objectives:CT brain (CTB) scans are frequently performed in older adults, a population at increased risk of osteoporosis and low bone mineral density (BMD), presenting an opportunity for opportunistic screening. This study aims to evaluate an automated deep learning approach for opportunistic screening of low BMD and osteoporosis from routine CTB imaging. Materials and methods:A single-centre retrospective analysis was conducted on 2,014 patients (mean age 69.7 ± 14.9 years; 61% female) who underwent non-contrast CTB and dual-energy x-ray absorptiometry (DEXA) within one year of each other. A convolutional neural network incorporating automatically selected CTB slices cranial to the lateral ventricles, as well as age and sex, was trained to perform two binary classification tasks: low BMD screening (T-score < -1.0) and osteoporosis screening (T-score ≤ -2.5). Model performance was evaluated on a 10% hold-out test set using AUC, balanced accuracy, sensitivity, specificity, positive predictive value, and negative predictive value, with subgroup analysis by sex. Results:22% of patients scanned had normal BMD, 44% had osteopenia, and 34% had osteoporosis. For low BMD screening, the model achieved an AUC of 0.83 (95% CI: 0.76-0.90), with AUCs of 0.88 in females and 0.76 in males. For osteoporosis screening, the model achieved an AUC of 0.78 (95% CI: 0.72-0.85), with AUCs of 0.78 in females and 0.74 in males. Conclusion:Automated analysis of routine CT brain imaging showed good discriminatory performance for opportunistic low BMD screening, particularly among females. With further validation, this approach could support earlier identification of at-risk individuals using imaging already acquired in routine clinical care.
PurposeAutomated monitoring of multiple sclerosis (MS) lesion progression remains challenging in clinical practice. This study evaluates LongiSeg, a longitudinal deep learning architecture, on heterogeneous clinical data to assess its practical value for routine MS radiological monitoring and extends its use for new lesion detection.MethodsWe trained LongiSeg on 470 patients from a diverse clinical cohort acquired across 15 scanner types (both 1.5T and 3T) with mixed 2D/3D T2w FLAIR sequences. We trained and evaluated models for both cross-sectional lesion segmentation and new lesion detection. Performance was compared against a single-timepoint nnU-Net on a cross-sectional test cohort consisting of 67 patients as well as 22 patients presenting with new lesions.ResultsLongiSeg achieved superior cross-sectional segmentation performance compared to single-timepoint nnU-Net on the clinical test set (n = 67, T2w FLAIR input, DSC: 0.684 ± 0.124 vs. DSC: 0.665 ± 0.144). LongiSeg for new lesion detection obtained a low performance on the small clinical test cohort (n = 22, DSC: 0.260, 95%CI: 0.143–0.383).ConclusionLongiSeg demonstrated moderate improvements in cross-sectional MS lesion segmentation. However, the added complexity of processing longitudinal scans may not be justified by these modest gains. For new lesion segmentation, performance was low, especially in cases with few lesions.
BackgroundPericardial lipomas are rare, benign cardiac tumors. Encasement of a coronary artery by such a tumor is extremely uncommon and significantly elevates surgical risk. We report the first case in which multimodality imaging clearly demonstrates a pericardial lipoma enveloping the right coronary artery (RCA) without luminal stenosis, adding a novel anatomical observation to the literature.Case presentationA 32-year-old woman presented with a one-year history of exertional chest tightness and dyspnea that had worsened over three days. Physical examination and laboratory tests were unremarkable except for elevated thyroid-stimulating hormone (13.26 μIU/mL). Echocardiography revealed a 7.8 cm hypoechoic mass above the aortic root with a segment of the RCA coursing within it; color Doppler showed no flow disturbance. Contrast-enhanced cardiac computed tomography (CT) and cardiac magnetic resonance (CMR) confirmed a homogeneous fat-density mass (CT attenuation −85 HU, no enhancement, complete signal suppression on STIR) consistent with a pericardial lipoma. The tumor measured 5.5 cm in maximum transverse diameter and was seen to completely encase the RCA without stenosis. A diagnosis of pericardial lipoma with RCA encasement and concomitant subclinical hypothyroidism was made. After multidisciplinary discussion, levothyroxine was initiated and a staged plan including preoperative coronary angiography followed by surgical resection with RCA protection was formulated. The patient's symptoms resolved at discharge. At one- and three-month follow-ups she remained asymptomatic, but she subsequently declined all further contact and was lost to follow-up; no adverse cardiac events were recorded during the observation period.ConclusionPericardial lipomas can encase a coronary artery while maintaining luminal patency due to the compliance of adipose tissue. When a pericardial mass is detected on echocardiography, the coronary course must be actively traced. Multimodality imaging with CT and cardiac magnetic resonance (CMR) is essential for delineating the tumor-artery relationship and guiding surgical strategy. Coexisting endocrine disorders should be corrected preoperatively. Although the loss of long-term follow-up limits our conclusions, the imaging findings themselves provide an important cautionary lesson.
Portal vein thrombosis (PVT) is a heterogeneous and clinically significant complication of liver cirrhosis, primary hepatobiliary malignancy, and various systemic prothrombotic states. It is associated with advanced portal hypertension, increased surgical complexity during liver transplantation, and higher mortality. The aim of this manuscript is to provide a narrative review of PVT and portal vein recanalization combined with transjugular intrahepatic portosystemic shunt (PVR-TIPS), detailing epidemiology, pathophysiology, and diagnostic classifications while exploring evolving medical and interventional management strategies. A comprehensive narrative review of the existing literature was conducted, incorporating clinical practice guidelines from major societies (AASLD, Baveno VII, EASL, and AGA) and recent prospective and retrospective studies regarding the management of non-malignant PVT in both cirrhotic and non-cirrhotic patients. PVT occurs in approximately 1% of the general population but affects up to 44% of candidates for liver transplantation (LT). Its development is governed by Virchow's triad, with reduced portal flow velocity (<15 cm/s) being the primary independent predictor in cirrhosis. Diagnosis relies on Doppler ultrasound as a first-line screening tool, followed by contrast-enhanced computed tomography or magnetic resonance imaging to define the degree of occlusion, extent, and chronicity using standardized nomenclature. Medical management with anticoagulation is the cornerstone for recent PVT (<6 months), achieving recanalization in 40%–75% of cases. For chronic obliterative PVT or refractory complications (variceal bleeding or refractory ascites), PVR-TIPS has emerged as a safe and highly successful intervention, with technical success rates ranging from 75% to 100% in expert centers. PVR-TIPS facilitates physiologic end-to-end portal anastomosis during LT in up to 96%–100% of cases, significantly improving transplant feasibility and outcomes. The management of PVT has transitioned toward highly individualized strategies. Standardized classification remains essential for guiding therapy. While medical management is effective for recent-onset disease, PVR-TIPS represents a transformative interventional advancement for chronic obliterative PVT, restoring physiologic flow and optimizing candidates for curative liver transplantation.
Primary cardiac tumors are extremely rare, and hemangiomas represent only a small subset, with right atrial involvement being particularly uncommon, posing diagnostic challenges due to non-specific symptoms and overlapping imaging features. We present a 55-year-old male with an incidentally discovered right atrial mass initially suspected to be a myxoma. Multimodal imaging with transthoracic echocardiography (TTE) and cardiac magnetic resonance (CMR) was performed: TTE identified a large mobile mass causing partial tricuspid obstruction, while CMR with quantitative diffusion-weighted imaging/apparent diffusion coefficient (DWI/ADC) provided detailed tissue characterization and anatomical mapping, revealing typical features of a hemangioma. Surgical resection was performed via median sternotomy under cardiopulmonary bypass. The mass, arising from the right atrial posterior wall and interatrial septum, was completely excised, and the atrial defect was repaired with a pericardial patch. Histopathology and immunohistochemistry confirmed a benign capillary hemangioma. Postoperative echocardiography demonstrated reduced right atrial diameter, resolution of obstruction, and mild stable tricuspid regurgitation, with no major complications. Combined TTE and CMR multimodal imaging is crucial for accurate preoperative diagnosis and surgical planning of right atrial hemangiomas. Complete surgical resection achieves favorable outcomes, and long-term multimodal imaging follow-up is recommended to exclude recurrence.
Three barriers significantly hinder the use of deep learning in medical imaging: poor generalization to new clinical domain shifts, label scarcity, and data privacy. A unified framework that learns from unlabeled, decentralized data while optimizing for generalization is desperately needed, even if Federated Learning (FL), Self-Supervised Learning (SSL), and Domain Generalization (DG) provide partial solutions that often operate under contradictory assumptions. We present FedAD: Adaptive Federated Disentanglement, a unique framework that uses two key ideas to handle these problems in a synergistic way. First, a federated semantic disentanglement objective (FedSD) explicitly distinguishes between the domain-invariant semantic characteristics and domain-specific variants using a non-adversarial orthogonality constraint. Second, in order to prevent premature convergence and enhance resilience, an adaptive teacher-student alignment (ATSA) curriculum dynamically modifies the generalization pressure based on the stability of the global model. This dual technique creates a strong feature encoder by forcing the model to learn what it sees as opposed to where it sees it. FedAD outperforms current approaches in terms of generalization to unseen target domains, as demonstrated by its validation on publicly available medical datasets. Our strategy concurrently addresses privacy, label scarcity, and domain change, paving the road for useful, reliable, and fair medical AI.
BackgroundThoracic ossification of the ligamentum flavum (TOLF) is frequently underrecognized in its early stage because radiographic abnormalities on routine chest radiographs are often subtle. We aimed to develop and externally validate a deep learning model for opportunistic screening of TOLF using routine chest radiographs.MethodsThis retrospective multicenter diagnostic study included an internal development cohort from Changzheng Hospital and an independent external validation cohort from South China Hospital. The internal cohort comprised 250 patients with TOLF and 250 control subjects collected between January 2017 and January 2023. The external cohort comprised 150 patients with TOLF and 150 control subjects. TOLF status was established on CT using predefined radiological criteria, whereas frontal and lateral chest radiographs were used only as model inputs. We evaluated multiple backbone architectures, including ResNet101, DenseNet169, Vision Transformer, and Swin Transformer, and additionally explored three dual-view fusion strategies. Model development was performed using 10-fold cross-validation in the internal cohort, and performance was summarized using bootstrap-derived 95% confidence intervals. Human-reader comparison was conducted in the internal cohort.ResultsIn backbone screening within the internal cohort, ResNet101 emerged as the best-performing architecture. After subsequent input-resolution optimization, the final lateral-view ResNet101 model achieved an accuracy of 97.0%, sensitivity of 94.0%, specificity of 100.0%, and an AUC of 0.970. None of the evaluated dual-view fusion strategies outperformed the best single lateral-view model, and the poorer performance of posterior-fusion models was mainly attributable to reduced sensitivity. Compared with experienced spine surgeons and imaging physicians, the internal ResNet101 model showed significantly higher sensitivity and overall accuracy (both p < 0.001). In the external validation cohort, the locked model maintained robust discrimination, with an AUC of 0.954 for frontal radiographs and 0.995 for lateral radiographs. The corresponding accuracy/sensitivity/specificity values were 90.0%/84.7%/95.3% for frontal radiographs and 93.7%/89.3%/98.0% for lateral radiographs.ConclusionA deep learning model based on routine chest radiographs may provide accurate and generalizable screening for TOLF across institutions. The lateral-view model showed the most consistent diagnostic performance, supporting its potential role as an opportunistic screening tool to prompt confirmatory CT evaluation.
BackgroundAutomated chest x-ray reporting could substantially reduce the burden on radiology services worldwide; however, the implementation of current vision–language models (VLMs) in clinical workflows remains limited by factual errors, hallucinations, and inadequate clinical reliability. Bridging this implementation gap requires frameworks that are not only technically sound but also designed for safe integration into real-world healthcare settings.ObjectiveThis study aimed to improve the clinical accuracy and factual consistency of radiology report generation by introducing an agent-inspired, three-stage reasoning framework and evaluating its feasibility as research prototype for potential implementation in resource-constrained and high-throughput clinical environments.MethodsWe propose the Swin-Qwen3 vision–language architecture, which integrates a Swin Transformer visual encoder with a Q-Former-style query-driven cross-attention module aligned with a large language model (Qwen3-0.6B). A three-stage generation strategy—comprising initial report drafting, clinical verification, and structured refinement—was guided by role-specific prompts to simulate distinct clinical reasoning behaviours. Parameter-efficient fine-tuning via low-rank adaptation (LoRA) enabled training within standard GPU constraints. The framework was evaluated on the full IU x-Ray test set (321 Samples) using lexical, semantic, clinical, and factuality metrics and compared with one- and two-stage ablations. Computational feasibility and inference overhead were assessed for offline or batch processing contexts.ResultsThe three-stage framework showed modest but consistent improvements over ablation baselines. CheXpert-F1 reached 0.7038 (compared to 0.7154 for one-stage and 0.7123 for two-stage), Clinical-F1 reached 0.5401 (compared to 0.5478 for one-stage and 0.5449 for two-stage), and RadGraph-F1 scored 0.5467 (compared to 0.5478 for one-stage and 0.5445 for two-stage). The factuality score reached 0.843 (compared to 0.845 for one-stage and 0.845 for two-stage), while the hallucination rate remained at 0.6116 (compared to 0.6109 for one-stage and 0.6109 for two-stage). Clinical recall decreased from 0.8361 (one-stage) and 0.8337 (two-stage) to 0.8044 (three-stage), reflecting a trade-off between sensitivity and precision. However, the clinical verifier demonstrated substantial effectiveness, correcting 85.6% of identified hallucinations. Lexical quality improved: BLEU-4 was 0.0628 (+10.9% vs. one-stage; +13.1% vs. two-stage), ROUGE-2-F was 0.0957 (+10.2% vs. one-stage; +13.1% vs. two-stage), and METEOR was 0.3728 (+4.5% vs. one-stage; +6.3% vs. two-stage). Improvements in METEOR (p < 0.001) and ROUGE-2-F (p < 0.05) were statistically significant. A 14.8 × inference overhead was identified, indicating suitability for offline reporting contexts.ConclusionsThe Swin-Qwen3 three-stage framework demonstrates that explicit clinical verification and iterative refinement can modestly enhance the clinical reliability of VLMs for automated radiology reporting. The parameter-efficient design is scalable across diverse clinical settings, including resource-limited environments, and supports a practical deployment pathway aligned with SDGs 3 (Good Health and Well-being), 9 (Industry, Innovation and Infrastructure), and 17 (Partnerships for the Goals). However, the residual hallucination rate of 61.2% indicates that the current system is a promising research prototype rather than a clinically deployable tool. Future work should prioritize multi-institutional validation on larger datasets (e.g., MIMIC-CXR) with radiologist expert review, regulatory evaluation, and prospective clinical integration studies.
Context and objectivesLung ultrasound (LUS) is a safe low-cost tool that enables diagnosis, monitoring and guidance for interventional procedures at the patient's bedside. However, its expansion is hindered by a lack of training programs and the inherent difficulty of interpreting ultrasound images. In this context, Artificial Intelligence (AI) is emerging as a supportive tool for LUS interpretation, ensuring diagnostic efficacy and mitigating the shortage of experts. This systematic review aims to summarize and analyze recent advances in AI-based tools to support LUS interpretation.MethodsA systematic literature search was conducted across Web of Science, IEEE Xplore, and PubMed databases to identify peer-reviewed original journal articles published between 2015 and November 2025 that employed AI for the identification and localization of lung artifacts, anatomical structures, and pathological findings. Risk of bias was assessed using PROBAST + AI.ResultsTwenty-four studies were included, identifying three main strategies: segmentation (10 studies), object detection (4 studies), and the generation of visual explanations through saliency maps (10 studies). All employed CNN-based architectures. The evaluation metrics used were heterogeneous. The PROBAST + AI assessment showed relevant risk-of-bias concerns, mainly concentrated in the participants and analysis domains.ConclusionsThe development of AI systems to support LUS interpretation shows high potential; however, current studies exhibit significant heterogeneity in their objectives, methodologies, and evaluation metrics. It is necessary to move towards solutions designed for specific clinical environments and to adopt standardized protocols and evaluations that facilitate their implementation in clinical practice.Systematic Review Registrationhttps://www.crd.york.ac.uk/PROSPERO/view/CRD420261322517, PROSPERO CRD420261322517.
IntroductionReconstruction-based methods offer a promising solution for unsupervised anomaly detection in medical imaging tasks. These methods train generative models on healthy data alone and identify anomalies as deviations between an input image and its pseudo-healthy reconstruction. The downfall of these methods is their dependence on two assumptions that often fail in practice: that models cannot reproduce unseen pathologies, yet can faithfully reconstruct healthy tissue. A recent image-conditioned diffusion approach explicitly addresses these issues by training a model to restore synthetic anomalies inserted into healthy images. However, it operates in 2D pixel space, discarding inter-slice context and incurring high computational cost.MethodsWe address both limitations by performing image-conditioned restoration in a 3D latent space using a pretrained VAE and rectified flow, capturing volumetric context whilst drastically reducing computational overhead. To mitigate false positives introduced by VAE compression, we propose using the restoration change which measures the difference between the pseudo-healthy latent restoration and the VAE reconstruction of the original, rather than the standard reconstruction error. We further experiment with applying the synthetic anomaly training task directly in latent space to improve sensitivity to low-contrast anomalies.ResultsWe perform extensive experiments across various medical imaging benchmarks, including brain MRI and the newly released AADD dataset, comparing against reconstruction-based, feature-modelling, attention-based and self-supervised anomaly detection methods. Our image-conditioned rectified flow models establish a new state-of-the-art, with an ensemble of models trained with pixel-space and latent-space anomalies yielding the strongest overall performance.DiscussionThese results demonstrate how incorporating 3D context enables better anomaly detection performance whilst also being ∼10 times faster. The difference in performance between models trained using latent-space and pixel-space anomalies suggests that further broadening of the anomaly imputation process could continue to improve the robustness of these models. Such improvements are certainly necessary, as the AADD benchmark is far from being saturated. Code is available at https://github.com/matt-baugh/img-cond-latent-rflow-model-ad.
Cerebral venous sinus thrombosis (CVST) accounts for less than 1% of cerebrovascular events worldwide and is associated with significant morbidity; in severe or anticoagulation-refractory cases, mechanical thrombectomy is considered, yet no consensus exists on optimal device selection. We report the case of a 32-year-old woman with a history of migraine and oral contraceptive use who presented with sudden-onset headache on awakening, disorientation, and right hemiparesis, followed by focal seizures and National Institutes of Health Stroke Scale (NIHSS) progression from 3 to 7. Magnetic resonance venography confirmed extensive CVST involving the superior sagittal sinus and bilateral transverse and sigmoid sinuses. Mechanical thrombectomy was performed on hospital day 1 via a right jugular approach using a combined aspiration and NeVa stent retriever (5.5 × 37 mm) technique with a 6-second integration time; after six passes, near-complete recanalization of the superior sagittal sinus was achieved with restoration of antegrade venous drainage and no intraprocedural complications. The post-procedural course included neurological improvement (Medical Research Council grade 4/5) and secondary intracranial hypertension (opening pressure 33 cmH₂O), which was managed conservatively. The patient was discharged after 12 days of hospitalization, with a modified Rankin Scale score of 0 at the 34-day follow-up visit. To our knowledge, this represents the first report of off-label NeVa stent retriever use for thrombectomy in the cerebral venous system; the device's higher mesh density may theoretically offer advantages for organized venous thrombi in large-caliber sinuses, expanding therapeutic options where consensus on optimal technique remains elusive.
Lung cancer remains one of the leading causes of cancer-related mortality worldwide, and early detection is essential to improve patient survival. Computed tomography (CT) is currently the reference imaging modality for lung cancer screening. Recent advances in Artificial Intelligence (AI), particularly Deep Learning (DL), have significantly improved automated medical image analysis and diagnostic accuracy. In this study, we propose a Convolutional Neural Network (CNN)-based approach for pulmonary nodule classification. The model was trained using a merged dataset composed of public CT image databases (IQ-OTH/NCCD and SPIE-AAPM) and real-world CT scans collected at Cheikh Zaid Hospital, Morocco. The experimental dataset included 1,103 malignant images, 508 benign images, and 427 normal images. Contrast Limited Adaptive Histogram Equalization (CLAHE) was applied to enhance image contrast prior to training. The proposed CNN achieved an accuracy of 99.84%, precision of 99.97%, sensitivity of 99.84%, and specificity of 99.82% for a three-class classification task. These results demonstrate the robustness and clinical potential of the model in assisting radiologists with the classification of indeterminate pulmonary nodules. This work represents an initial translational step toward real-world clinical implementation of AI-based lung nodule classification at Cheikh Zaid Hospital and highlights the feasibility of integrating AI tools into routine radiology workflows.
Radiology artificial intelligence (AI) is increasingly developed on large external datasets and deployed across institutions, but real-world model performance may vary substantially after implementation. Imaging AI interacts with a local ecosystem shaped by scanner hardware, acquisition protocols, reconstruction methods, technologist practices, disease prevalence, patient demographics, reporting conventions, and clinical workflow. These factors can produce domain shift, degrade calibration, alter false-positive and false-negative patterns, and affect clinical utility. In this article, we argue that radiology AI should move beyond a “plug-and-play” deployment paradigm toward institution-calibrated AI stewardship. We propose an institution-specific radiology AI performance profile, conceptually analogous to a local antibiogram, to summarize how AI tools perform within a specific clinical environment. Unlike prior MLOps and radiology AI governance frameworks, the proposed profile translates lifecycle management into a radiology-specific, locally maintainable artifact that captures technical context, clinical context, performance metrics, reference standards, equity checks, workflow effects, and governance triggers. The framework applies not only to diagnostic decision support, but also to research cohort generation and radiology education, where local reporting language, imaging protocols, and case mix may strongly influence model reliability. We emphasize that institution-specific calibration does not require every hospital to develop AI models de novo. Rather, externally developed, vendor-based, and foundation-model approaches should be paired with local validation, cautious threshold adjustment, calibration when applicable, surveillance for drift and protocol changes, and multidisciplinary governance. Responsible radiology AI deployment should therefore ask not only whether a model is accurate, but whether it remains accurate, relevant, equitable, and useful locally over time.