Automated anatomical vessel labeling of abdominal arteries is important in medical image processing, as the abdominal arterial system plays a vital role in the human body. Such labeling supports disease diagnosis, surgical planning and population-level epidemiological studies. This work focuses on the automatic labeling of 13 major abdominal arteries from contrast-enhanced computed tomography scans. However, achieving accurate and fully automatic labeling remains challenging due to the tortuous and highly branched anatomy of the abdominal arterial system, as well as large inter-individual anatomical variations. To address these challenges, we propose a novel subgraph classification framework that captures the topological variants of the abdominal arterial system. First, vessel segmentation is performed. Next, vessel centerlines are extracted from the segmentation to construct a vascular graph. Subgraph classification is then applied for vascular graph labeling, and finally, the 13 abdominal arterial segments are labeled by mapping centerline labels back onto the segmentation. We evaluate our method on a dataset of 200 contrast-enhanced CT scans, and results show superior performance compared to state-of-the-art methods, with an increase of 1.7% in the Dice score and a reduction of 0.6 mm in the Hausdorff distance.
OBJECTIVE:To assess the utility of AI-driven quantification of abdominal aortic calcification (AAC) extracted from contrast-enhanced CT (CECT) scans using a fully automated AI tool for predicting all-cause mortality and cardiovascular events (CVEs). MATERIALS AND METHODS:In this retrospective cohort study, a fully automated deep learning tool for quantifying AAC from the aortic hiatus to iliac bifurcation was applied to abdominal CT scans in a large, adult population undergoing CT between January 2001 and February 2021. The aortic tool was applied to noncontrast CT (NCCT) and CECT scans. Validated linear adjustments were applied to CECT scans to derive NCCT-equivalent (NCE) measures. Death and cardiovascular events following CT were documented from the electronic health record. ROC curve and time-to-event analyses were performed to generate 5 and 10-year AUCs, HRs comparing the highest and lowest risk quartiles, and Kaplan-Meier survival curves. RESULTS:A total of 123,500 adults (mean age, 51; M:F, 58,442: 65,058) underwent abdominal CT over the study interval (43,455 NCCT; 79,885 CECT). 21,937 patients died, and 16,599 patients had a documented CVE over the study interval (median follow-up: 59 mo). Similar 10-year AUCs were observed for mortality [0.744 (NCCT), 0.732 (CECT), 0.732 (NCE)] and CVE[0.716 (NCCT), 0.718 (CECT), and 0.718 (NCE)] prediction. Comparing the highest and lowest risk quartiles, higher AAC was associated with increased mortality risk: HR of 2.08 (CI: 1.94, 2.24) for NCCT, HR of 1.8 (CI: 1.70, 1.91) for CECT ( P <0.001 for all). Higher AAC was also associated with increased CVE risk: HR of 1.73 (CI: 1.61,1.86) for NCCT, HR of 1.45 (CI: 1.35,1.55) for CECT ( P <0.001 for all). No differences were observed after applying adjustments for IV contrast. CONCLUSIONS:AAC extracted from clinical CECT scans is predictive of mortality and CVD risk, with similar performance to NCCT. AI-driven quantification of AAC from CECT is feasible and can expand the scope and impact of population-level "opportunistic screening."
In previous work, we integrated large language models (LLMs) into the lesion segmentation model based on the ULS23 DeepLesion dataset, using short-form findings from the reports. In this study, we developed a unified 2D lesion analysis framework that integrates LLM-based reasoning, lesion bounding box detection, segmentation, and radiology report generation from the original DeepLesion dataset. In the testing phase, we achieved relatively high lesion bounding box detection accuracy with mAP50 of 70.1
Contrast-enhanced CT (CECT) is routinely used by clinicians to diagnose metabolic diseases such as diabetes. However, a patient’s clinical indication and contrast agent hypersensitivity dictates the specific CT imaging protocol, which can result in missing contrast CT series. For retrospective research studies, it is challenging to re-scan patients with incomplete CT studies. Consequently, this pilot study explores the feasibility of synthesizing missing abdominal CECT series from the available series in a CT study. Given the non-contrast CT and desired CT phase conditioning, we propose to generate an abdominal CECT series using spatial transform control (STControl) in a pre-trained conditional diffusion model. To evaluate the CECT translation quality, a proxy pancreas segmentation task was conducted. Results on the internal Institution-A and external Vin-Dr datasets revealed that STControl improved peak signal-to-noise ratio (PSNR) by 0.42 dB and structural similarity image metric (SSIM) by ∼ 1.28. The proposed technique is generalizable to translate any CT phase into another and may be of particular value when organ morphology rather than focal lesions is critical as in diabetes. Code is available at https://github.com/rsummers11/STControl .
Frontier large language models (LLMs), such as GPT-5, Claude 4.5, Gemini 3, Llama 4 and DeepSeek-R1, represent a transformative class of artificial intelligence tools capable of revolutionizing various aspects of healthcare by generating human-like responses across diverse contexts and adapting to novel tasks following human instructions. Their potential application spans a broad range of medical tasks, such as clinical documentation, matching patients to clinical trials and answering medical questions. Here in this Tutorial, we discuss an actionable set of best practices to help healthcare professionals utilize LLMs more effectively and efficiently. The overall workflow follows sequential phases from formulating the task, choosing the most appropriate LLMs, engineering the prompts, fine-tuning the requests and through to model deployment. We discuss a set of critical considerations in identifying medical tasks that align with the core capabilities of LLMs and selecting models based on the required task, data, performance and model interface. We then review the strategies, such as prompt engineering and fine-tuning, to adapt standard LLMs to specialized medical tasks. We then cover deployment considerations, including regulatory compliance, ethical guidelines and continuous monitoring for fairness and bias. By providing a structured step-by-step methodology, this entry-level tutorial aims to equip healthcare professionals with the tools necessary to effectively integrate LLMs into clinical practice, ensuring that these powerful technologies are applied in a safe, reliable, and impactful manner.
Splenomegaly, or spleen enlargement, arises from a wide range of diseases and can lead to cytopenia, abdominal discomfort, and immune dysfunction. Magnetic resonance imaging (MRI) is commonly used to assess splenomegaly, providing detailed, radiation-free visualization of spleen size and structure. However, existing segmentation methods often show limited accuracy on MRI scans of enlarged spleens, as they are typically trained on cases with normal spleen sizes. This study aims to enhance spleen and neighboring organ segmentation accuracy in MRI scans by incorporating a small set of labeled splenomegaly cases into training. We propose an improved method, MRISegmenter++ (MS++), based on the nnU-Net segmentation framework, and benchmarked its performance against four established multi-organ segmentation methods: MRISegmenter (MS), MRAnnotator (MA), TotalSegmentator MRI (TS), and TotalVibeSegmentator (TV). On both internal and external datasets, MS++ performed comparably to MS, TS, and TV ( p >0.8 ) on scans without splenomegaly, while significantly outperforming them on dedicated splenomegaly datasets ( p < .05 ). These findings suggest that MS++ not only maintains high segmentation accuracy in patients with normal spleen size but also substantially improves performance in cases of splenomegaly. Accurate spleen segmentation may contribute to improved diagnosis and prognosis of splenomegaly in clinical practice.
Intrapancreatic fat deposition is a well-established biomarker for type 2 diabetes mellitus (T2DM), but little is known about the role of pancreatic perivascular adipose tissue (PVAT), fat surrounding blood vessels supplying the pancreas. The objectives of this study were to develop a fully automated pipeline to derive imaging biomarkers of pancreatic PVAT from abdominal CT scans and evaluate their utility for diagnosing T2DM. 1,350 contrast-enhanced CT (CECT) scans from the public PANORAMA dataset were used to train a 3D deep neural network to segment pancreatic anatomy (parenchyma, vasculature, ducts) and extrapancreatic structures. The model was applied to an internal dataset containing 615 CECT scans. Intrapancreatic, extrapancreatic, and PVAT biomarkers were derived from these segmentations. Logistic regression with feature selection identified optimal biomarker subsets for classifying T2DM. All PVAT biomarkers differed significantly between diabetics and non-diabetics. The best multivariate model for T2DM achieved an AUC of 0.93, sensitivity of 0.73, and specificity of 0.91. Exploratory analyses for incident diabetes (developed within the following 4 years) showed similarly high discrimination (AUC 0.93). Automated CT-derived PVAT biomarkers are associated with existing and future T2DM and may enable opportunistic identification of undiagnosed or high-risk individuals undergoing routine abdominal CT.
Type 2 Diabetes Mellitus (T2DM) is a chronic metabolic disease that affects millions of people worldwide. Early detection is crucial as it can alter pancreas function through morphological changes and increased deposition of ectopic fat, eventually leading to organ damage. While studies have shown an association between T2DM and pancreas volume and fat content, the role of increased pancreatic surface lobularity (PSL) in patients with T2DM has not been fully investigated. In this pilot work, we propose a fully automated approach to delineate the pancreas and other abdominal structures, derive CT imaging biomarkers, and opportunistically screen for T2DM. Four deep learning-based models were used to segment the pancreas in an internal dataset of 584 patients (297 males, 437 non-diabetic, age: 45±15 years). PSL was automatically detected and it was higher for diabetic patients (p=0.01) at 4.26 ± 8.32 compared to 3.19 ± 3.62 for non-diabetic patients. The PancAP model achieved the highest Dice score of 0.79 ± 0.17 and lowest ASSD error of 1.94 ± 2.63 mm (p<0.05). For predicting T2DM, a multivariate model trained with CT biomarkers attained 0.90 AUC, 66.7% sensitivity, and 91.9% specificity. Our results suggest that PSL is useful for T2DM screening and could potentially help predict the early onset of T2DM.
There is a growing awareness that body CT scans contain rich cardiometabolic information that can be leveraged for additional patient benefits. However, the clinical implementation of opportunistic CT screening in routine practice has been hindered by valuable yet onerous manual measurements and subjective assessments. Explainable artificial intelligence (AI) algorithms are now poised to change this. The potential impact of opportunistic screening is further enhanced by the large volume of CT scans being obtained. In this "How I Do It" installment, the authors briefly outline some current approaches that can be obtained "on the fly," while focusing more on emerging automated solutions. Detecting unsuspected or presymptomatic conditions, such as osteoporosis, cardiovascular disease, sarcopenia, and hepatic steatosis, could lead to preventive interventions, regardless of the original indication for imaging. Composite models that combine multiple cardiometabolic CT biomarkers can be applied to survival prediction and assessment of biologic aging, frailty, cancer cachexia, metabolic syndrome, and fracture risk, among other factors. For clinical reporting, a range of logistical, actuarial, and ethical issues must be carefully considered. However, if executed properly, we believe that opportunistic CT screening can add substantial value, be cost saving, and provide a new level of personalized precision medicine befitting the dawning AI information era.
Artificial intelligence (AI) can automatically delineate lesions on computed tomography (CT) and generate radiology report content, yet progress is limited by the scarcity of publicly available CT datasets with lesion-level annotations. To bridge this gap, we introduce CT-Bench, a first-of-its-kind benchmark dataset comprising two components: a Lesion Image and Metadata Set containing 20,335 lesions from 7,795 CT studies with bounding boxes, descriptions, and size information, and a multitask visual question answering benchmark with 2,850 QA pairs covering lesion localization, description, size estimation, and attribute categorization. Hard negative examples are included to reflect real-world diagnostic challenges. We evaluate multiple state-of-the-art multimodal models, including vision-language and medical CLIP variants, by comparing their performance to radiologist assessments, demonstrating the value of CT-Bench as a comprehensive benchmark for lesion analysis. Moreover, fine-tuning models on the Lesion Image and Metadata Set yields significant performance gains across both components, underscoring the clinical utility of CT-Bench.
BACKGROUND & AIMS:Metabolic dysfunction-associated steatotic liver disease is a major cause of liver disease that is growing in prevalence. Typically asymptomatic, is often undiagnosed. Imaging studies, including noncontrast computed tomography (CT) scans, can detect hepatic steatosis opportunistically, but the reporting rate is unknown. We hypothesized that incidental finding of steatosis on noncontrast CT is often not reported if not specifically sought in the imaging request. METHODS:This study was a retrospective, cross-sectional, single-center analysis of abdominal noncontrast CT scans performed between 2012 and 2020 for any indication in adult subjects. Images were analyzed using an automated deep learning liver segmentation and attenuation assessment algorithm to obtain a mean volumetric liver attenuation value. Image-based steatosis was defined as mean hepatic attenuation <40 HU and compared with textual radiology reports. Manual review of a random subset of scans and reports was used to verify results. RESULTS:A total of 3646 noncontrast CT scans from 2710 adult patients were analyzed. The mean liver attenuation derived from the deep-learning algorithm was 50.4 ± 11.8 HU. Image-based steatosis was found in 480 (13.1%) scans, with a mean liver attenuation of 29.8 ± 14 HU. Radiologists reported steatosis in only 157 (32.7%) of these low-attenuation scans. Predictors of unreported steatosis included higher average liver attenuation (even if < 40 HU), high variability of fat distribution in the liver and low body mass index. CONCLUSIONS:We found that incidental hepatic steatosis in noncontrast CT is reported in a minority of scans. Incorporating artificial intelligence-based hepatic attenuation measurement in computed tomography scan reading may increase reporting rates.
Transformers have shown strong potential in medical image segmentation, particularly as foundation models. However, comparative studies indicate that nnU-Net, a CNN-based framework, often outperforms transformer-based models. This study investigates why nnU-Net achieves superior performance, focusing on small object segmentation in CT scans and examining the role of transformers in developing foundation models for medical image segmentation. Six transformer-based and two CNN-based methods were trained in the MONAI framework using identical data augmentations and training settings. nnU-Net employed its self-configuration strategy to automatically determine optimal parameters. Segmentation of calcified plaques and mediastinal lymph nodes was evaluated. nnU-Net achieved higher accuracy for both tasks, with a statistically significant improvement over transformers in plaque segmentation $(p<.001$). The self-configuration mechanism of nnUNet played a key role in optimizing CNN training. Within the same framework, MedFormer outperformed SegResNet $(p=.04$), which shares the same ResU-Net architecture as nnU-Net, suggesting that transformers may benefit from enhanced self-configuration strategies. Overall, transformer architectures remain promising candidates for developing foundation models and achieving higher segmentation accuracy through self-configuration in supervised medical image segmentation.
This retrospective study of 118 pediatric rhabdomyosarcoma patients aimed to determine the relationship between skeletal muscle characteristics on diagnostic CT scans and outcome. Univariate analysis revealed that older age and higher muscle bulk were associated with increased risk of events and mortality. While age remained a significant predictor in Cox regression for event-free survival, higher muscle density was associated with improved overall survival. Interestingly, a significant interaction between dichotomized muscle bulk and sex was observed for overall survival, suggesting higher bulk was associated with worse outcomes in males but potentially better outcomes in females. Our findings indirectly emphasize the importance of maintaining good nutrition and physical activity during chemotherapy to preserve muscle and improve outcomes, but further multicenter research is necessary to understand the complex interplay between body composition, sex and outcome in pediatric rhabdomyosarcoma.
Chest X-ray (CXR) interpretation is hindered by the long-tailed distribution of pathologies and the open-world nature of clinical environments. Existing benchmarks often rely on closed-set classes from a single institution, failing to capture the prevalence of rare diseases or the appearance of novel findings. To address this, we present the CXR-LT challenge. The first event, CXR-LT 2023, established a large-scale benchmark for long-tailed multi-label CXR classification and identified key challenges in rare disease recognition. CXR-LT 2024 further expanded the label space and introduced a zero-shot task to study generalization to unseen findings. Building on the success of CXR-LT 2023 and 2024, this third iteration of the benchmark introduces a multi-center dataset comprising over 145,000 images from PadChest and NIH Chest X-ray datasets. Additionally, all development and test sets in CXR-LT 2026 are annotated by radiologists, providing a more reliable and clinically grounded evaluation than report-derived labels. The challenge defines two core tasks this year: (1) Robust Multi-Label Classification on 30 known classes and (2) Open-World Generalization to 6 unseen (out-of-distribution) rare disease classes. This paper summarizes the overview of the CXR-LT 2026 challenge. We describe the data collection and annotation procedures, analyze solution strategies adopted by participating teams, and evaluate head-versus-tail performance, calibration, and cross-center generalization gaps. Our results show that vision-language foundation models improve both in-distribution and zero-shot performance, but detecting rare findings under multi-center shift remains challenging. Our study provides a foundation for developing and evaluating AI systems in realistic long-tailed and open-world clinical conditions.
To report longitudinal intra-patient changes in CT-based body composition using fully automated AI tools in an adult patient sample. This retrospective longitudinal study included 15,616 adult patients (mean age at first CT, 53.0 ± 14.7 years; 7096 male, 8520 female) who underwent at least two abdominal CT examinations at least five years apart (mean study interval, 9.1 ± 3.3 years, range, 5.0-20.4 years) at a single academic institution between January 1, 2000 and February 28, 2021. CT examinations were not restricted based on patient setting, clinical indication, or IV contrast media use. Seven fully automated AI body composition tools quantifying vertebral trabecular attenuation, skeletal muscle area and attenuation, visceral adipose tissue (VAT) area and attenuation, subcutaneous adipose tissue (SAT) area, and VAT/SAT ratio (VSR) were applied to each patient’s first and last available abdominal CT. Change in body composition per year were determined using the first and last available CT scans. T-test and linear regression were used to assess sex and age as predictors of longitudinal body composition change, respectively. Significant differences in sex-specific mean rates of change were observed for all measures (p < 0.05) except muscle attenuation. Certain CT biomarkers showed varying rates of change among younger, middle-age, and older adults. Age significantly predicted body composition measures except VSR in female patients, although effect size was small (R2 values 0.002–0.040). There are significant age and sex-specific differences in longitudinal, intra-patient body composition changes over time.
ARTICLE HIGHLIGHTS:Pancreatic perivascular adipose tissue (PVAT) is a potentially important but understudied fat depot surrounding vessels supplying the pancreas, and its role in type 2 diabetes mellitus (T2DM) remains unclear. We aimed to quantify pancreatic PVAT automatically on routine abdominal computed tomography (CT) scans and assess whether PVAT-derived imaging biomarkers are associated with T2DM and future T2DM risk. PVAT biomarkers were associated with T2DM and were informative predictors in multivariable models. Exploratory analyses showed similar findings for incident T2DM (disease onset within 4 years of follow-up). PVAT may represent a novel CT-based biomarker for early metabolic dysfunction and T2DM risk stratification.
Purpose To develop a deep learning model that automatically delineates the eight liver Couinaud segments and the spleen at CT for future liver remnant (FLR) volumetry. Materials and Methods In this retrospective study (January 2001-October 2025), eight liver Couinaud segments and the spleen were manually labeled on CT scans of patients from institution A and the public Medical Segmentation Decathlon dataset. A three-dimensional nnU-Net segmentation model was trained on this dataset and evaluated on three datasets (one internal and two external). Results The training dataset included 498 patients (442 from the public Medical Segmentation Decathlon dataset and 56 from institution A, mean age ± SD, 55 years ± 7; 38 men), and the testing dataset included 64 patients from institution A (50 had liver fibrosis and eight underwent portal vein embolization; PVE), 197 patients from the publicly available colorectal liver metastases (CRLM) dataset (mean age, 59 years ± 12; 117 men), and 50 patients (25 were healthy and 25 had cirrhosis) from an external site (institution B; mean age, 49 years ± 9; 29 men). For the whole liver in institution A and institution B, Dice scores of 0.98 ± 0.02 (95% CI: 0.97, 0.99) and 0.98 ± 0.03 (95% CI: 0.97, 0.99), and 95% percentile Hausdorff distance (HD) errors of 2.5 mm ± 3.8 (95% CI: 1.6, 3.3) and 3.3 mm ± 6.6 (95% CI: 1.4, 5.2) were obtained, respectively. The pre-PVE FLR% and post-PVE FLR% volume differences (manual vs automated, eight patients) were 0.03 ± 2.4 and -0.39 ± 3.0, respectively. For the FLR in the CRLM dataset, a Dice score of 0.99 ± 0.01 (95% CI: 0.99, 0.993) and an HD error of 0.9 mm ± 1.8 (95% CI: 0.6, 1.1) were achieved. Conclusion The model accurately estimated preoperative FLR volumetry and generalized well to patients with colorectal liver metastases, fibrosis, and cirrhosis and healthy controls. Keywords: CT, Deep Learning, Couinaud, Spleen, Future Liver Remnant, Portal Vein Embolization, Cirrhosis, Fibrosis Supplemental material is available for this article. © RSNA, 2026.
Medical image segmentation is important for quantitative disease diagnosis and treatment but relies on accurate pixel-wise labels, which are costly, time-consuming, and require domain expertise. This work introduces MIST (MIxed supervision, Self, and Transfer learning) to reduce manual labeling in medical image segmentation. A small set of cases was manually annotated ("strong labels"), while the rest used automated, less accurate labels ("weak labels"). Both label types trained a dual-branch network with a shared encoder and two decoders. Self-training iteratively refined weak labels, and transfer learning reduced computational costs by freezing the encoder and fine-tuning the decoders. Applied to segmenting muscle, subcutaneous, and visceral adipose tissue, MIST used only 100 manually labeled slices from 20 CT scans to generate accurate labels for all slices of 102 internal scans, which were then used to train a 3D nnU-Net model. Using MIST to update weak labels significantly improved nnU-Net segmentation accuracy compared to training directly on strong and weak labels. Dice similarity coefficient (DSC) increased for muscle (89.2 ± 4.3% to 93.2 ± 2.1%), subcutaneous (75.1 ± 14.4% to 94.2 ± 2.8%), and visceral adipose tissue (66.6 ± 16.4% to 77.1 ± 19.0% ) on an internal dataset (p<.05). DSC improved for muscle (80.5 ± 6.9% to 86.6 ± 3.9%) and subcutaneous adipose tissue (61.8 ± 12.5% to 82.7 ± 11.1%) on an external dataset (p<.05). MIST reduced the annotation burden by 99%, enabling efficient, accurate pixel-wise labeling for medical image segmentation. Code is available at https://github.com/rsummers11/NIH_CADLab_Body_Composition.
Robust localization of lymph nodes (LNs) in multiparametric MRI (mpMRI) is critical for the assessment of lymphadenopathy. Radiologists routinely measure the size of LN to distinguish benign from malignant nodes, which would require subsequent cancer staging. Sizing is a cumbersome task compounded by the diverse appearances of LNs in mpMRI, which renders their measurement difficult. Furthermore, smaller and potentially metastatic LNs could be missed during a busy clinical day. To alleviate these imaging and workflow problems, we propose a pipeline to universally detect both benign and metastatic nodes in the body for their ensuing measurement. The recently proposed VFNet neural network was employed to identify LN in T2 fat suppressed and diffusion weighted imaging (DWI) sequences acquired by various scanners with a variety of exam protocols. We also use a selective augmentation technique known as Intra-Label LISA (ILL) to diversify the input data samples the model sees during training, such that it improves its robustness during the evaluation phase. We achieved a sensitivity of ∼83% with ILL vs. ∼80% without ILL at 4 FP/vol. Compared with current LN detection approaches evaluated on mpMRI, we show a sensitivity improvement of ∼9% at 4 FP/vol.