Background: Distinguishing invasive adenocarcinoma (IAC) from non-IAC in pure ground-glass nodules (pGGNs) remains a critical clinical challenge. We aimed to develop and multicenter-validate a CT-based deep learning model to differentiate IAC from non-IAC lesions in pGGNs and compare its performance with that of human experts, thereby identifying candidates for definitive surgical management among pGGNs likely to represent IAC. Methods: This retrospective study included 1707 surgically resected pathologically confirmed pGGNs from six medical institutions. We developed Lung-PNetV2, a modular deep-learning framework that integrates cross-scanner normalization, 3D volumetric encoding (via ResNet-18), and the multimodal fusion of imaging, nodular, and clinical features. The model was trained on 847 pGGNs and validated internally (203 pGGNs) and externally (657 pGGNs). Seven clinicians (four radiologists and surgeons) independently evaluated the holdout test set using the NCCN-guided 5-point scoring. Results: Lung-PNetV2 achieved AUCs of 0.892 (training), 0.831 (internal test), and 0.827 (external test), significantly outperforming all the human readers (AUC: 0.681–0.722; P < 0.01). At the clinical decision threshold (0.6 probability for IAC), the model demonstrated balanced performance in the external test set: 81.0% accuracy, 67.7% sensitivity, 84.1% specificity, and 91.8% negative predictive value. The deep learning model surpassed radiologists in sensitivity (67.7% vs 60.9%) and surgeons in specificity (84.1% vs 78.3%) while maintaining superior F1-macro (72.5% vs reader range: 54.1–67.0%; P < 0.05, 4/7 readers). Conclusion: CT-based Lung-PNetV2, a deep learning model, provides a generalizable performance that surpasses that of human experts in stratifying the invasiveness of pGGNs through the effective cross-modal fusion of imaging and clinical data. Its balanced sensitivity and specificity profiles support risk-stratified management, allowing predicted non-IAC pGGNs to avoid unnecessary procedures, whereas high-risk cases receive timely definitive surgery, enabling individualized and precise treatment.
Objectives To establish a 3D V-Net-based segmentation model for adrenal glands on abdominal CT images and validate its performance in multicentre datasets, including chest CT images. Methods CT images of adrenal glands were retrospectively collected for the training of the adrenal segmentation model. Abdominal CT scans with normal and abnormal adrenal glands (N = 5660) were recruited as the model development cohort and were split into training, internal validation, and internal test sets for the development of the segmentation model. Two groups of health screening subjects were included for model validation: 1 from the same institution (N = 6126, validation cohort 1) and 1 from an outside institution (N = 931, validation cohort 2). Their chest CT images were used for model validation. The Dice similarity coefficient (DSC) was used to evaluate the efficacy of the model. Results The DSC of the test set for left and right adrenal segmentation were 0.920 (0.890-0.930) and 0.910 (0.890-0.930), respectively. In the validation cohorts, the DSC were 0.816 (0.744-0.866) for the left adrenal gland and 0.819 (0.743-0.865) for the right adrenal gland in validation cohort 1, and 0.752 (0.666-0.820) for the left adrenal gland and 0.747 (0.673-0.812) for the right adrenal gland in validation cohort 2. Conclusions The 3D V-Net-based adrenal segmentation model achieves considerable segmentation efficacy and demonstrates generalizability from abdominal CT to chest CT, making it suitable for use in CT images with various scanning protocols. Advances in knowledge The study developed a deep learning model using 3D V-Net for the segmentation of adrenal glands on CT images, achieving good performance of normal and abnormal glands in validation cohorts with different scanning protocols and from multiple institutions, demonstrating its potential as a “flagging” system aiding diagnosis.
Background:Preoperative prediction of renal capsule invasion in small renal masses (SRMs) is crucial for treatment planning but challenging on computed tomography (CT). This study developed a deep learning radiomics (DLR) model using CT to noninvasively predict capsule invasion in SRM. Methods:We analyzed 413 SRMs from three centers (July 2017 to September 2024). Data from Centers 1 (the First Affiliated Hospital of Shandong First Medical University) and 2 (the Third Affiliated Hospital of Shenzhen University) (330 patients, 57.27±11.58 years) comprised the training set, and Center 3 (the Union Hospital, Tongji Medical College, Huazhong University of Science and Technology) (83 patients, 57.67±10.76 years) served as the external test set. Radiomics and deep learning features were extracted using PyRadiomics and a pre-trained ResNet50. Feature selection used maximum relevance and minimum redundancy (mRMR) and least absolute shrinkage and selection operator (LASSO). Model performance was evaluated by the area under the curve (AUC), with interpretability assessed via SHapley Additive exPlanations (SHAP) and clinical utility by calibration and decision curves. Results:On the training set, the radiomics (Rad), deep transfer learning (DTL), and DLR models showed AUCs of 0.846, 0.890, and 0.855, respectively. On the external test set, corresponding AUCs were 0.746, 0.715, and 0.734. SHAP analysis revealed greater contribution from deep learning features. All models demonstrated good calibration and clinical utility. Conclusions:The DLR model is feasible for noninvasive prediction of renal capsule invasion in SRM. While not outperforming individual Rad or DTL models, it provides a valuable exploratory tool for preoperative assessment.
INTRODUCTION:The spleen is a key immune organ whose volume and attenuation can change in inflammatory, infectious, and hematologic diseases. In routine practice, splenic enlargement is often assessed using semi-quantitative linear criteria, such as rib-unit-based estimates, which can miss modest but clinically relevant volume changes. While deep learning enables reliable organ segmentation on CT, large-scale descriptions of normal splenic morphometrics in Chinese adults, including age- and sex-related patterns, remain limited. MATERIALS AND METHODS:This study developed a 3D V-Net for automated spleen segmentation, and trained it on 2,856 CT examinations from four public datasets: Abdomen CT-1K, FLARE23, AMOS, and the RSNA 2023 Abdominal Trauma dataset. The model’s performance was assessed using the Dice similarity coefficient (DSC), volume similarity (VS), Hausdorff distance (HD), and average Hausdorff distance. For morphometric analysis and institutional validation, this model was then applied to 1,520 CT examinations in 2023 of individuals with a clinically and radiologically normal spleen, comprising a total of 3,490 CT images across phases (NOC/AP/PVP/DP), from which spleen volume, mean CT value, and three-dimensional diameters were extracted. RESULTS:On the public test set, the model achieved a Dice similarity coefficient (DSC) of 0.982 [interquartile range (IQR): 0.975-0.988], a volumetric similarity (VS) of 0.995 [IQR: 0.991-0.998], and a Hausdorff distance (HD) of 3.047 mm [IQR: 2.578-4.653 mm]. Within the institutional cohort, the median spleen volume was significantly greater in males (213.74 cm3 [IQR: 158.40-284.88]) than in females (163.70 cm3 [IQR: 125.99-217.72]; p < 0.001). Age-related volumetric patterns also differed by sex: males exhibited an inverted U-shaped trajectory, peaking at 273.75 ± 92.96 cm3 in the 28-37 year age group, while females displayed a bimodal pattern, peaking at 225.01 ± 65.53 cm3in the 18-27 year age group. Furthermore, splenic attenuation was significantly higher in females across the arterial, portal venous, and delayed contrast phases (all p < 0.001). DISCUSSION:The 3D V-Net enabled accurate and efficient spleen segmentation on CT, which permitted automated and reproducible measurement of splenic morphometrics. The observed sex- and age-specific patterns in spleen volume and attenuation likely reflect the combined effects of physiology, including hormonal status, as well as metabolic and body-composition factors. Reported differences across populations suggest that splenic reference values may not be directly transferable between ethnic groups, underscoring the need for population-calibrated morphometric data. Limitations of this study include potential residual confounding, incomplete adjustment for anthropometric covariates, and validation within a single institutional center. CONCLUSION:This study establishes a practical deep learning-based pipeline for automated spleen segmentation and standardized morphometric extraction on CT in Chinese adults. The resulting population-specific reference distributions may support more objective interpretation of spleen size and attenuation in routine practice and provide a foundation for future disease-oriented applications. Further multicenter validation is warranted.
4533 Background: The multicenter randomized phase 2/3 FRUSICA-2 trial (NCT05522231) demonstrated that fruquintinib plus sintilimab (F+S) significantly improved progression-free survival (PFS) (22.2 months vs 6.9 months) and objective response rate (ORR) (60.5% vs 24.3%) by blinded independent central review (BIRC) assessment compared to axitinib or everolimus (A/E) in Chinese patients (pts) with advanced renal cell carcinoma (aRCC) who had failed prior tyrosine kinase inhibitor therapy (Ye D, et al; 2025 ESMO). Considering that baseline tumor burden may correlate with efficacy outcome, we present the results of a relevant post-hoc subgroup analysis. Methods: Overall, 234 eligible pts were 1:1 randomized to receive either F+S or A/E. The primary efficacy endpoint was PFS assessed by BIRC per RECIST 1.1; secondary endpoints included investigator-assessed PFS, ORR, disease control rate, duration of response, time to response, and overall survival. This subgroup analysis evaluated BIRC-assessed PFS and ORR across subgroups defined by the number of target lesions (TLs) and metastatic sites (METs) at baseline. Results: At baseline, 79, 79, 32 and 44 pts had 1, 2, 3 and ≥4 TL(s), respectively. Metastatic disease was present in 117 (98.3%) pts in F+S arm and 110 (95.7%) pts in A/E arm, with a higher proportion of pts in F+S arm (88, 73.9%) having ≥3 METs than in A/E arm (67, 58.3%). By the data cut-off date of Feb 17, 2025, median follow-up for PFS was 16.6 months. As summarized in the table, F+S demonstrated superior PFS versus A/E across all subgroups, with unstratified hazard ratios (HRs) ranged from 0.28 to 0.50, as well as consistently longer median PFS. Similarly, improvements in ORR were observed across subgroups with odds ratios (ORs) ranged 2.46~7.22. Notably, in F+ S arm, fewer baseline TLs and METs appeared to correlate with longer median PFS, though such trend was not observed for ORR. Conclusions: Consistent with the primary analysis, F+S showed superior efficacy compared to A/E in terms of PFS and ORR in the second-line treatment of aRCC, regardless of the amount of baseline TLs or METs. Clinical trial information: NCT05522231 . Subgroup(F+S v A/E) 1 TL(38 v 41) 2 TLs(38 v 41) 3 TLs(16 v 16) ≥4 TLs(27 v 17) 1 MET(29 v 43) 2 METs(39 v 28) ≥3 METs(49 v 39) PFS, HR (95% CI) a 0.39 (0.20, 0.76) 0.33 (0.17, 0.66) 0.43 (0.17, 1.10) 0.28 (0.12, 0.62) 0.28 (0.13, 0.60) 0.50 (0.24, 1.04) 0.31 (0.18, 0.545) Median PFS, months b 24.9 vs 8.3 22.2 vs 6.9 15.3 vs 4.2 13.8 vs 6.9 24.9 vs 8.3 22.2 vs 8.3 NE vs 4.2 ORR, OR (95% CI) c 3.95 (1.35, 11.91) 4.65 (1.63, 13.43) 7.22 (1.17, 52.78) 5.53 (1.21, 28.76) 7.18 (2.22, 23.84) 2.46 (0.81, 7.76) 6.12 (2.13, 18.44) ORR, % 52.6 vs 22.0 65.8 vs 29.3 62.5 vs 18.8 63.0 vs 23.5 65.5 vs 20.9 53.8 vs 32.1 61.2 vs 20.5 NE, not estimable. a Based on an unstratified Cox proportional risk model. b Estimated using Kaplan-Meier method. c Exact 95% CI for OR was calculated using Cochran-Mantel-Haenszel method.
Gynecologic and breast cancers (GBC) represent one of the most prevalent malignancies among women globally, and the Dietary Inflammation Index (DII) may influence its associated risk. This research aims to investigate the impact of DII on the risk of GBC in adult female smokers in the United States, utilizing National Health And Nutrition Examination Survey data from 2007 to 2020. A descriptive analysis was initially conducted to evaluate the dataset, followed by a binomial logistic regression model to assess the relationship between DII and GBC. Three models were developed: Model I (unadjusted), Model II (adjusted for age, race, education level, and income ratio), and Model III (further adjusted for confounding variables including body mass index, chronic diseases, alcohol consumption, and physical activity). A multifaceted sensitivity analysis was performed to validate the robustness of the findings. The analysis included 7501 participants, of whom 1291 were smokers. Results indicated that a higher DII was significantly associated with increased GBC risk. In Model I, each unit increase in DII corresponded to a 24% increase in risk (Odds ratios [OR]: 1.24, confidence intervals [CI]: 1.07-1.44, P = .005). Model II showed a 32% increase (OR: 1.32, CI: 1.13-1.55, P = .001), while Model III indicated a 27% increase (OR: 1.27, CI: 1.05-1.53, P = .017). Sensitivity analysis revealed no significant effects of DII on former smokers and nonsmokers (P > .05). Additionally, subgroup analysis based on race, education level, income, body mass index, and DII quartiles did not yield significant results (P > .05). Restricted cubic splines analysis did not identify a nonlinear relationship between DII and GBC (P = .307). The DII is significantly associated with the risk of GBC among female smokers, underscoring the necessity of incorporating dietary factors into cancer prevention and intervention strategies.
Upper urinary tract urothelial carcinoma (UTUC) is an aggressive malignancy treated primarily with radical nephroureterectomy (RNU), yet reliable preoperative prognostic indicators remain limited. Chronic kidney disease (CKD) is common in patients with UTUC and may influence both treatment options and oncological outcomes, but its prognostic value before surgery has not been fully clarified. We retrospectively analyzed 298 UTUC patients who underwent RNU between January 2014 and December 2023. To balance baseline characteristics, 1:1 propensity score matching (PSM) was performed. Cox proportional hazards models and Kaplan–Meier analyses were used to evaluate the association between preoperative CKD and overall survival (OS), intravesical recurrence-free survival (IVRFS), and metastasis-free survival (MFS). Subgroup analyses were conducted in patients with impaired renal function (eGFR < 90 mL/min/1.73 m² or CKD stage ≥ 2). Additional preoperative clinical-stage-adjusted models included clinical T and N stage. The median follow-up was 39 months (IQR, 18.2–69.0), and preoperative CKD was present in 133 patients (44.6
Objective: The aim of this study was to investigate multiomics (MO) integration with stacked-ensemble learning for predicting neoadjuvant chemotherapy (NAC) response and recurrence risk in breast cancer (BC). Impact Statement: This study demonstrates that a stacked-ensemble learning model integrating clinicopathologic and magnetic resonance imaging (MRI)-based intratumoral heterogeneity biomarkers effectively predicts NAC response and postoperative recurrence risk in BC patients. These findings underscore MO and machine learning’s potential to optimize clinical decision-making. Introduction: Selecting BC patients who will benefit from NAC remains challenging. Methods: We retrospectively analyzed 124 BC patients receiving NAC (3 to 8 cycles) prior to mastectomy. Two radiomics signatures—RadSET and RadSITH—were derived from pre-NAC high-resolution dynamic MRI to track entire-tumor and intratumoral heterogeneous characteristics, respectively. These signatures were integrated with clinicopathologic indicators using stacked-ensemble learning algorithms to predict pathological complete response (pCR) and 3-year disease-free survival (DFS). Results: Among the 124 patients, the pCR rate was 26.6%. For pCR prediction, RadSITH and RadSET yielded areas under the curve (AUCs) of 0.798 and 0.770, respectively. The MO-integrated model, combining RadSITH, RadSET, clinical N stage, and molecular subtype, achieved a significantly higher AUC (0.917; 95% confidence interval [CI], 0.860 to 0.958; P < 0.05) than individual models. Postoperative recurrence occurred in 13.6% of patients. The elastic-net Cox model achieved a DFS concordance index of 0.78 (95% CI, 0.72 to 0.83) using pre-NAC variables (MO-predicted pCR, Response Evaluation Criteria in Solid Tumors response, RadSITH), and 0.81 (95% CI, 0.76 to 0.92) with post-NAC variables (pathologic grade, pCR status, pT stage, and pN stage). Conclusion: The MO integration with stacked-ensemble learning effectively predicts NAC response and recurrence risk in BC.
ObjectiveTo systematically summarize study-level factors associated with postoperative hematoma after ultrasound-guided vacuum-assisted breast lesion excision(VAE) using a vacuum-assisted breast biopsy (VABB) system, and to organize the available evidence according to a novel T-P-B framework comprising tumor-related, position-related, and breast- or peri-procedural management-related factors.MethodsA systematic search was conducted in PubMed, China National Knowledge Infrastructure, Wan fang Data, VIP Database, and Elsevier Clinical Key for studies published between January 1995 and October 2025. Observational studies investigating risk factors for hematoma after VAE using a VABB system were included. Study-derived exposure categories and cutoff values were retained. Study quality was assessed using the Newcastle–Ottawa Scale (NOS). Meta-analysis was performed using R (version 4.5.1). Effect sizes were expressed as odds ratios (ORs) with corresponding 95% confidence intervals (CIs) and were pooled using the inverse-variance method. Heterogeneity was assessed using the I² statistic, and a random-effects model was applied. Sensitivity analysis was conducted by sequentially excluding individual studies (leave-one-out analysis). A two-sided p value < 0.05 was considered statistically significant.ResultsA total of 11 retrospective studies involving 3,516 patients who underwent VAE using a VABB system were included, among whom 444 cases of postoperative hematoma were reported. Five study-level factors were identified. Meta-analysis demonstrated that, within the T dimension, tumor numbers (OR = 4.21, 95% CI = 2.59–6.85), numbers of cutting passes(OR=3.87,95%CI=2.16-6.95)were significantly associated with hematoma formation. Within the P dimension, non-moderate tumor depth, including superficial or deep tumor, was associated with postoperative hematoma (OR = 4.39, 95% CI: 1.21–15.92), as was higher vascularity grade (OR = 2.60, 95% CI: 1.42–4.76). Within the B dimension, postoperative compression duration <48 h was associated with hematoma formation (OR = 4.34, 95% CI: 2.53–7.45). Sensitivity analyses showed generally consistent directions of association.ConclusionsThis systematic review and meta-analysis identified five study-level factors associated with postoperative hematoma after VAE using a VABB system: tumor numbers, a higher number of cutting passes, non-moderate tumor depth including superficial or deep tumor, higher vascularity grade, and postoperative compression duration <48 h. The T-P-B framework may provide a clinically interpretable structure for organizing these factors and informing perioperative risk assessment.
To develop a deep learning model for the automated detection and measurement of gallstones on non-contrast CT images. A total of 3,231 CT scans were retrospectively enrolled as an internal cohort, while 753 CT scans from the public AMOS (Abdominal Multi-Organ Segmentation) dataset were employed as an independent external test set. Paired MR imaging served as the reference standard for the internal cohort (including development and hold-out datasets). A three-stage 3D V-Net convolutional neural network was developed for automated gallstone detection. The first stage performed coarse localization of the gallbladder, followed by refined 3D segmentation in the second stage to generate precise anatomical masks. In the final stage, gallstones were detected within the segmented gallbladder volume. Gallstone detection was categorized based on the spatial overlap (Dice similarity coefficient, DSC > 0) between reference labels and predictions. For gallbladder segmentation, the DSCs were 0.992, 0.989, and 0.990 across training, validation, and internal test sets. Gallstone detection achieved mean DSCs of 0.794, 0.742, and 0.759 across the subgroups. For gallstone detection and segmentation, the model achieved overall sensitivities of 97.2
4531 Background: FRUSICA-2 (NCT05522231), a randomized, open-label, controlled phase 2/3 study, demonstrated a significant efficacy of F plus S vs A or E in previously treated advanced RCC (BIRC-assessed mPFS 22.21 vs 6.90 mo; HR 0.373, p<0.0001; ORR 60.5% vs 24.3%, OR 4.622, p<0.0001; D Ye, et al; 2025 ESMO). Given the IMDC classification as the most common prognostic model and PD-L1 expression has demonstrated prognostic value across various cancer types, we conducted this exploratory subgroup analysis to evaluate clinical outcomes by baseline scores of IMDC risk factors and PD-L1 expression. Methods: Patients (pts) with histologically confirmed advanced RCC previously treated with VEGFR-TKI were randomized 1:1 to receive either F (5 mg QD, 2 weeks on/1 week off) plus S (200 mg iv, every 3 weeks), or investigator's choice of A (5 mg twice daily) or E (10 mg once daily). Efficacy was analyzed by IMDC risk factor scores (0, 1, 2, and ≥3) and PD-L1 combined positive score (CPS≥1, <1, or unknown). Data cutoff was February 17, 2025. Results: Among 234 patients in the phase 3 part, IMDC risk scores and PD-L1 CPS distributions are detailed in table. With a median follow-up of 16.56 mo, BIRC-assessed mPFS for F+S vs A/E across IMDC subgroups were: score 0, not estimable(NE) vs 8.31 mo (stratified HR 0.270, p=0.0009); score 1, 24.87 vs 8.25 mo (HR 0.289, p<0.0001); score 2, 13.8 vs 6.9 mo (HR 0.436, p=0.0164); and score ≥3, 9.69 vs 4.21 mo (HR 0.591, p=0.3267). Across PD-L1 expression subgroups, BIRC-assessed mPFS for F+S vs A/E were: CPS ≥1, 22.21 vs 3.22 mo (stratified HR 0.232, p=0.0004); CPS <1, NE vs 8.28 mo (HR 0.350, p<0.0001). An additional 82 patients had unknown PD-L1 CPS status (F+S: n=40; A/E: n=42) and were not included in PD-L1 subgroup analysis. ORR analyses consistently favored F+S over A/E across all subgroups defined by IMDC risk factors and PD-L1 expression, with detailed results presented in table. Conclusions: In the exploratory analyses, these findings suggest that the efficacy benefit of F+S vs A/E is maintained across different IMDC risk scores and PD-L1 CPS in previously treated advanced RCC patients, supporting its broad application. Clinical trial information: NCT05522231 . Efficacy results by IMDC risk scores and PD-L1 expression. F+S v A/E IMDC score 0(33 v 32) IMDC score 1(43 v 41) IMDC score 2(30 v 31) IMDC score ≥3(13 v 11) PD-L1 CPS≥1(23 v 20) PD-L1 CPS<1(56 v 53) mPFS mo NE v 8.31 24.87 v 8.25 13.80 v 6.90 9.69 v 4.21 22.21 v 3.22 NE v 8.28 HR (95%CI) 0.27 (0.12, 0.62) 0.29 (0.15, 0.55) 0.44 (0.22, 0.88) 0.59 (0.20, 1.72) 0.23 (0.10, 0.56) 0.35 (0.20, 0.61) Unstratified Log-rank p 0.0009 < 0.0001 0.0164 0.3267 0.0004 < 0.0001 ORR, % 63.6 v 25.0 62.8 v 26.8 60.0 v 19.4 46.2 v 27.3 78.3 v 10.0 64.3 v 35.8 Odds Ratio (95%CI) 5.25(1.61, 17.71) 4.60(1.66, 12.98) 6.25(1.75, 23.80) 2.29(0.32, 19.10) 32.40(4.67, 338.60) 3.22(1.37, 7.61)
Objective:To systematically review the literature on ductoscopy-related operative failures, establish a classification system for "Complex Ductoscopy Operating Procedure (CDOP)", and propose standardized clinical response strategies. Methods:A comprehensive search of PubMed, Embase, Web of Science, Scopus, and Cochrane Library was conducted for publications from January 1991 to July 2026. Causes of ductoscopy failure were extracted, categorized, and quantitatively summarized. The methodological quality of included studies was evaluated. Based on literature evidence and accumulated clinical experience, a standard CDOP-PFN (Passing, Finding, Non-discharge) classification system and stepwise clinical management framework were developed. Results:From 25 included studies (2,853 procedures), meta-analysis showed a pooled failure rate of 10.0% (95%CI: 6.8%-13.7%; 95% PI: 0.0%-31.4%) with high heterogeneity (I² = 87.1%). Primary causes were ductal perforation (22.5%), ductal stenosis (21.9%), and nipple deformity/retraction (12.1%). Failures were stratified into three major types (PFN): (i) presence of nipple discharge with inability to access the ductal system via ductoscopy (type P, passing); (ii) bloody discharge with failure to identify intraductal lesions (type F, finding); and (iii) absence of spontaneous discharge but ultrasonographic evidence of ductal dilatation and intraductal lesions (type N, non-discharge). For each type, a tiered strategy was formulated, incorporating pre-procedural evaluation, duct entry optimization, image-guided localization, troubleshooting, and surgical approaches when needed. Conclusion:The proposed CDOP-PFN classification and management framework standardizes the recognition and handling of complex ductoscopy operating scenarios. These strategies may reduce dependency on individual operator experience, improve procedural standardization and safety, enhance diagnostic yield for intraductal lesions, and advance minimally invasive breast disease management.
Abstract Objectives To evaluate the performance of a 3D V-Net-based segmentation model of adrenal lesions in characterizing adrenal glands as normal or abnormal. Methods A total of 1086 CT image series with focal adrenal lesions were retrospectively collected, annotated, and used for the training of the adrenal lesion segmentation model. The dice similarity coefficient (DSC) of the test set was used to evaluate the segmentation performance. The other cohort, consisting of 959 patients with pathologically confirmed adrenal lesions (external validation dataset 1), was included for validation of the classification performance of this model. Then, another consecutive cohort of patients with a history of malignancy (N = 479) was used for validation in the screening population (external validation dataset 2). Parameters of sensitivity, accuracy, etc., were used, and the performance of the model was compared to the radiology report in these validation scenes. Results The DSC of the test set of the segmentation model was 0.900 (0.810–0.965) (median (interquartile range)). The model showed sensitivities and accuracies of 99.7%, 98.3% and 87.2%, 62.2% in external validation datasets 1 and 2, respectively. It showed no significant difference comparing to radiology reports in external validation datasets 1 and lesion-containing groups of external validation datasets 2 (p = 1.000 and p > 0.05, respectively). Conclusion The 3D V-Net-based segmentation model of adrenal lesions can be used for the binary classification of adrenal glands. Critical relevance statement A 3D V-Net-based segmentation model of adrenal lesions can be used for the detection of abnormalities of adrenal glands, with a high accuracy in the pre-surgical scene as well as a high sensitivity in the screening scene. Key Points Adrenal lesions may be prone to inter-observer variability in routine diagnostic workflow. The study developed a 3D V-Net-based segmentation model of adrenal lesions with DSC 0.900 in the test set. The model showed high sensitivity and accuracy of abnormalities detection in different scenes. Graphical Abstract
Preoperative detection of muscle-invasive bladder cancer (MIBC) remains a great challenge in practice. We aimed to develop and validate a deep Vesical Imaging Network (ViNet) model for the detection of MIBC using high-resolution T2-weighted MR imaging (hrT2WI) in a multicenter cohort. ViNet was designed using a modified 3D ResNet, in which, the encoder layers were pretrained using a self-supervised foundation model on over 40,000 cross-modal imaging datasets for transfer learning, and the classification modules were weakly supervised by an experiential knowledge-domain mask indicated by a nnUNet segmentation model. Optimal ViNet model was trained in derivation data (cohort 1, n = 312) and validated in multicenter data (cohort 2, n = 79; cohort 3, n = 44; cohort 4, n = 56) across a multi-ablation-test for model selection. In internal validation, ViNet using hrT2WI outperformed all ablation-test models (odds ratio [OR], 7.41 versus 1.85-2.70; all P < 0.05). In external validation, the performance of ViNet using hrT2WI versus ablation-test models was heterogeneous (OR, 1.31-3.89 versus 0.89-9.75; P = 0.03-0.15). In addition, clinical benefit of ViNet was evaluated between six readers using the Vesical Imaging-Reporting and Data System (VI-RADS) versus ViNet-adjusted VI-RADS. As a result, ViNet-adjusted VI-RADS upgraded 62.9 % (17/27) of MIBC missed in VI-RADS score 2, while downgraded 84.1 % (69/84), 62.5 % (35/56) and 67.9 % (19/28) of non-muscle-invasive bladder cancer (NMIBC) overestimated in VI-RADS score 3-5. We concluded that ViNet presents a promising alternative for diagnosing MIBC using hrT2WI instead of conventional multiparametric MRI.
Objective: To develop a deep learning (DL)-based automated segmentation model for normal spleens to enable automated assessment of splenic diameter, volume, and CT (computed tomography) values; and to investigate key physiological factors influencing normal adult splenic volume to ultimately develop a population-specific predictive model for the Chinese adult population. Methods and Materials: To train the 3D U-Net segmentation model, Dataset 1 was randomly split into training (n=3418), validation (n=413), and test (n=443) sets. For internal validation (Dataset 2), we utilized 1,996 thin-slice CT images from upper abdominal scans conducted at our institution (January–April 2024); all scans exhibited no indications of splenic lesions or structural anomalies. The external validation set included 2,856 publicly available thin-slice CT images. Model performance on the test set was evaluated using the Dice coefficient, Hausdorff distance, and volume similarity. Step 2: Develop a predictive model by investigating physiological factors influencing normal adult spleen volume. Dataset 2 and an additional 578 upper abdominal CT scans (Dataset 3) performed at our institution (January–April 2023) showed no evidence of splenic tumors or structural abnormalities. The model developed in Step 1 was used to segment the spleen and calculate its volume. Spleen volume, three-dimensional dimensions, mean CT values, and contrast enhancement patterns were analyzed for scans with adequate segmentation. Physiological parameters influencing normal adult spleen volume were evaluated using portal venous phase images. Results: The model's training and validation revealed a volume similarity of 0.997, a Hausdorff distance of 0.015 [0.013, 0.018] mm, and a Dice similarity coefficient of 0.988 [0.984, 0.989] (median [interquartile range]). The subjects' spleen volumes ranged from 51,086.25 to 644,376.86 mm³, with a median of 177,903.06 mm³. The measurement ranges for x, y, and z are 87.56 ± 11.61 mm (mean ± standard deviation), 92.02 mm [81.98, 104.61], and 91.00 mm [80.00, 103.00], respectively. Age (r = −0.24, p < 0.0001) and gender (male=0, female=1; r = −0.32, p < 0.0001) were negatively correlated with splenic volume (SV), while height (r = 0.35, p < 0.0001), weight (W; r = 0.45, p < 0.0001), body mass index (BMI; r = 0.32, p < 0.0001), and body surface area (BSA; r = 0.46, p < 0.0001) were positively correlated with SV. Using portal venous phase thin-slice CT images, the standard splenic volume (SSV) for the Chinese population was derived using the formula SSV = −202,839.48 + 25.26W + 214,521.59BSA (where SSV = standard splenic volume (mm³), W = weight (kg), and BSA = body surface area (m²). Conclusion: The 3D U-Net architecture enabled the development of an effective automated splenic segmentation model, facilitating automated assessment of quantitative metrics including splenic volume, 3D dimensions, and mean CT values. Among the anthropometric variables studied, splenic volume exhibited the strongest correlations with body weight and body surface area.
Clear cell renal cell carcinoma (ccRCC) exhibits marked clinical heterogeneity, limiting the prognostic accuracy of traditional staging. We developed an unsupervised radiomics-based subtyping system integrating multi-omics data to decode tumor biology and improve risk stratification. Analyzing five cohorts (n = 1700, including surgical cohorts and an advanced ccRCC cohort receiving combined tyrosine kinase inhibitor and immunotherapy [T-I] treatment), we extracted 1834 CT radiomic features, applying consensus clustering to a discovery cohort (n = 748) and validating across centers. Two subtypes emerged with distinct recurrence risks: Cluster 1 and Cluster 2 (adjusted HR = 2.75 for recurrence, 95% CI 1.42-5.33, P = 0.003). Cluster 2's high recurrence risk was validated in three external cohorts (adjusted HRs: 1.76, 4.33, and 3.09; all P < 0.05). Radiogenomic analysis revealed Cluster 2 showed a higher frequency of VHL mutations and KDM5C mutations compared to Cluster 1, a more immunosuppressive microenvironment (reduced CD8+ T cell infiltration, P < 0.01; suppressed interferon signaling pathways, Gene Set Enrichment Analysis P < 0.05), and lower PD-L1 expression. In the T-I treated advanced ccRCC cohort, Cluster 2 patients had shorter overall survival. This first unsupervised radiomic system stratifies ccRCC by recurrence risk, molecular drivers, and treatment efficacy, offering a framework for precision oncology.
Vertebral compression fractures (VCFs) are the most common type of osteoporotic fractures, yet they are often clinically silent and undiagnosed. Chest frontal radiographs (CFRs) are frequently used in clinical practice and a portion of VCFs can be detected through this technology. This study aimed to develop an automatic artificial intelligence (AI) tool using deep learning (DL) model for the opportunistic screening of VCFs from CFRs. The datasets were collected from four medical centers, comprising 19,145 vertebrae (T6-T12) from 2735 patients. Patients from Center 1, 2 and 3 were divided into the training and internal testing datasets in an 8:2 ratio (n = 2361, with 16,527 vertebrae). Patients from Center 4 were used as the external test dataset (n = 374, with 2618 vertebrae). Model performance was assessed using sensitivity, specificity, accuracy and the area under the curve (AUC). A reader study with five clinicians of different experience levels was conducted with and without AI assistance. In the internal testing dataset, the model achieved a sensitivity of 83.0 % and an AUC of 0.930 at the fracture level. In the external testing dataset, the model demonstrated a sensitivity of 78.4 % and an AUC of 0.942 at the fracture level. The model's sensitivity outperformed that of five clinicians with different levels of experience. Notably, AI assistance significantly improved sensitivity at the patient level for both junior clinicians (from 56.1 % without AI to 81.6 % with AI) and senior clinicians (from 65.0 % to 85.6 %). In conclusion, the automatic AI tool significantly increases clinicians' sensitivity in diagnosing fractures on CFRs, showing great potential for the opportunistic screening of VCFs.
Purpose This multicenter study aims to externally validate Lung-PNet, a pre-trained artificial intelligence (AI) model, in differentiating invasive adenocarcinoma (IAC) from non-IAC in pure ground glass nodules (pGGNs); to quantify institutional and technical factors influencing AI performance; and to assess its reliability as a clinical decision-support tool. Methods Chest CT images of resected pGGNs (n = 720; IAC = 143, non-IAC = 577) from seven hospitals were retrospectively analyzed. The dataset was stratified into a local cohort (four hospitals, n = 334) and a national cohort (three hospitals, n = 386). Using the local cohort, receiver operating characteristic (ROC) curve analysis determined a Youden’s index-optimized cutoff threshold for binary classification, which was then validated in the national cohort. Model performance was evaluated using the area under the ROC curve (AUC-ROC). Factors influencing AI diagnostic accuracy were analyzed via random forest and mixed-effects logistic regression. Results Lung-PNet achieved an overall AUC-ROC of 0·800 (95% CI: 0·760–0·839), with comparable performance between the local (0·801, 95% CI: 0·740–0·861) and national (0·832, 95% CI: 0·782–0·881) cohorts (P = 0·438, DeLong test). The negative predictive value (NPV) was 0·930 (95% CI: 0·903–0·952), though hospital-specific NPVs varied (0·700–1·000). Key factors influencing the AI model’s accuracy included pGGN type, lesion volume, pixel spacing, and hospital, with hospital modeled as a random effect and pGGN type interacting with lesion volume and pixel spacing as fixed effects. Conclusion This multicenter validation confirms Lung-PNet’s generalizability for distinguishing IAC from non-IAC in pGGNs. Its high NPV supports its role as a reliable “rule-out” tool to reduce unnecessary interventions for non-IAC lesions. Institutional variability underscores the need for standardized imaging protocols to optimize AI integration into clinical workflows. Future studies should prospectively validate thresholds and assess long-term clinical outcomes.
Objective To investigate the value of mDIXON-Quant imaging–based proton density fat fraction and diffusion-weighted imaging in detecting metabolic syndrome–related renal injury. Methods A total of 24 patients with metabolic syndrome and 21 age-matched healthy volunteers were prospectively enrolled in this study and underwent mDIXON-Quant imaging and diffusion-weighted imaging. The participants were divided into those with metabolic syndrome–normal estimated glomerular filtration rate (eGFR; expressed in mL/min/1.73 m 2 ) (n = 10, eGFR ≥ 90; metabolic syndrome group 1), those with metabolic syndrome–mildly decreased eGFR (n = 14, 60 ≤ eGFR < 90; metabolic syndrome group 2), and controls (n = 21). Observer A measured radiological data, including renal fat fraction, renal apparent diffusion coefficient, liver fat fraction, and perirenal fat fraction. Statistical analyses were performed to compare the imaging parameters among groups, assess correlations with eGFR, determine the diagnostic performance of renal fat fraction using receiver operating characteristic analysis, identify eGFR predictors via multivariate analysis, and evaluate inter- and intra-observer consistency. Results Renal fat fraction was significantly elevated in metabolic syndrome groups compared with that in the control group (control: 3.83% ± 0.52%; metabolic syndrome group 1: 4.61% ± 0.59%; and metabolic syndrome group 2: 5.32% ± 0.47%; p < 0.001). No significant difference was observed in the apparent diffusion coefficient among the three groups ( p = 0.938). Renal fat fraction exhibited an inverse correlation with eGFR ( r = −0.688) across the study cohort. No significant correlation was observed between apparent diffusion coefficient values and eGFR ( r = −0.104, p = 0.495). The area under the curve value was 0.94 (95% confidence interval: 0.87–1.00) for distinguishing controls from patients with metabolic syndrome. A multivariable-adjusted analysis revealed a significant negative association between the renal fat fraction and eGFR (standardized B = −12.59). The interclass correlation coefficient between observers A and B for renal fat fraction measurements was 0.86. Conclusion mDIXON-Quant imaging exhibits greater clinical utility than diffusion-weighted imaging in detecting metabolic syndrome–related renal injury and may serve as a promising imaging tool for monitoring renal lipid deposition and therapeutic efficacy in lipid-mediated renal injury.