Rationale and Objectives While guidelines recommend tests like the FIB-4 index and liver stiffness measurements (LSM) to identify high-risk Metabolic dysfunction-associated steatotic liver disease (MASLD) patients, their actual accuracy in predicting liver-related events (LREs) remains unclear. Here, we systematically evaluate their prognostic performance to better inform clinical decisions. Materials and Methods We systematically searched Cochrane Library, Embase, PubMed and Web of Science (up to January 7, 2026), to identify studies utilizing non-invasive tests (NITs) for predicting incident LREs and reporting their prognostic performance. Subgroup analyses and meta-regression were performed based on dataset characteristics to explore potential sources of heterogeneity. All pooled analyses were conducted using a random-effects model. Results This meta-analysis included a total of 30 studies involving 67,533 patients with MASLD. The results demonstrated that for predicting incident LREs, the pooled C-index for Vibration-Controlled Transient Elastography (VCTE), the FIB-4 index, and Magnetic Resonance Elastography (MRE) were 0.83 (95% CI: 0.79–0.87), 0.82 (95% CI: 0.78–0.85), and 0.78 (95% CI: 0.64–0.91), respectively. In the evaluation of secondary outcomes, the pooled C-index of VCTE for predicting hepatocellular carcinoma (HCC) was 0.78 (95% CI: 0.70–0.86). Regarding mortality risk prediction, the pooled C-index for VCTE and the FIB-4 index were 0.83 (95% CI: 0.76–0.89) and 0.75 (95% CI: 0.66–0.84), respectively. Conclusion The FIB-4 index and LSM measured by VCTE exhibit robust prognostic discrimination for LREs in MASLD. Existing evidence supports the FIB-4 index and VCTE as routine non-invasive prognostic tests. Furthermore, the prognostic performance of MRE requires additional validation in future research.
OBJECTIVES:To develop and test a convolutional neural network model for automated segmentation of complicated cystic renal masses (cCRMs) on MRI. METHODS:This multicenter retrospective study analysed 210 cCRMs between October 2019 and May 2021, divided into training/internal validation (n = 150, Institution 1) and test sets (n = 60, Institutions 2-4). Comparative 3D V-Net and U-Net models were developed across 7 MRI sequences (T2-weighted, diffusion-weighted, apparent diffusion coefficient maps, unenhanced T1-weighted, and enhanced corticomedullary, nephrographic, and excretory phases images). A total of 14 models were developed, and 7 pairwise comparisons were performed between the 3D V-Net and U-Net models. Segmentation performance was evaluated using Dice similarity coefficient (DSC) and Hausdorff distance (HD), with subgroup analysis of small cCRMs (≤40 mm). RESULTS:In the test set, the excretory-phase V-Net (EPV-Net model) showed the highest DSC, and perform better than the corresponding U-Net (EPU-Net model) across all cCRMs (DSC: 0.74 ± 0.05 vs 0.70 ± 0.06, P < .001; HD: 27.41 ± 7.44 mm vs 39.18 ± 11.07 mm, P < .001) and the 35 small cCRMs subgroup (DSC: 0.74 ± 0.05 vs 0.70 ± 0.06, P < .001; HD: 27.48 mm ± 6.32 vs 38.72 ± 10.69 mm, P < .001). CONCLUSIONS:The 3D EPV-Net model demonstrated good segmentation accuracy, even for small lesions, supporting its clinical utility for cCRMs evaluation. ADVANCES IN KNOWLEDGE:This automated approach may streamline workflow compared to manual segmentation in cCRMs assessment.
Automatic lesion classification holds great promise for improving clinical diagnostic workflows, yet its practical application is hindered by domain shifts in multi-center medical images, which cause significant performance degradation. Although domain generalization (DG) with multi-source data offers a solution, a significant challenge remains: effectively suppressing domain-specific style variations to learn subtle, discriminative lesion features, which are often obscured by pronounced domain discrepancies. To address this challenge, we propose a novel framework named Frequency swapping and Soft-mask disentanglement for Domain Generalization (FSDG) for Domain Generalization. Our FSDG is designed to learn discriminative lesion features robust to domain shift through three key synergistic modules: a Fourier Amplitude Swap Module (FASM) that synthesizes domain-disturbed images by swapping high-frequency components to mitigate style variations; a Complementary Soft-mask Disentanglement Module (CSDM) that decomposes these images into distinct lesion-discriminative and domain-specific features; and a Domain Mixup with Consistency Regularization Module (DMCR) that enforces hybrid consistency constraints to ensure robust feature learning. These modules form a cohesive pipeline that can be flexibly integrated with different feature backbones, which we demonstrate on both single-modality and multi-modality tasks to enhance domain generalization for lesion diagnosis. Extensive experiments are conducted on two types of datasets: a publicly available multi-center skin lesion dataset from six hospitals, and a private multi-center Hepatocellular Carcinoma Microvascular Invasion (HCC-MVI) multi-parameters Magnetic Resonance Imaging(mp-MRI) dataset from five hospitals. Results show that the proposed FSDG achieves average accuracy of 81.89% and 79.05% on the single-modality skin lesion and multi-modality HCC-MVI tasks, respectively. On the skin lesion task, this represents an improvement of 11.43% over the baseline and a 0.87% margin over the second-best method. Our approach consistently outperforms several baseline and state-of-the-art methods across various domain combinations. These findings validate the effectiveness and strong generalization capability of FSDG in medical image domain generalization, demonstrating its potential to facilitate the broader clinical application of diagnostic models trained on limited multi-center data.
To refine the diagnostic criteria for rim arterial phase hyperenhancement (Rim APHE) and to evaluate the impact of this modification on the diagnostic performance for primary liver malignancies. This multicenter, retrospective study included patients with pathologically confirmed primary liver malignancies who underwent preoperative magnetic resonance imaging (MRI) before June 2021. Thirteen different measurement methods for Rim APHE were evaluated to determine the optimal criterion based on diagnostic performance. Two radiologists independently reviewed all observations, assigned Liver Imaging Reporting and Data System (LI-RADS) categories according to version 2018. The diagnostic performance of LI-RADS using original versus modified Rim APHE criteria was compared. The study enrolled 272 patients, including 204 with Hepatocellular carcinoma (HCC) (170 men and 34 women; mean age, 57 years ±10) and 68 with non-HCC malignancies (53 men and 15 women; mean age, 56 years ±10). The optimal criterion for Rim APHE was the total thickness of the thickest part of peripheral hyperenhancement to tumor diameter (area under the receiver operating characteristic (ROC) curve [AUC] = 0.852; P < .001). With implementation of the modified Rim APHE criterion, the LR-5 criteria showed significantly superior AUC (0.770 vs. 0.752, p = .007), sensitivity (76.0
Background:Large language models (LLMs) have shown considerable potential for extracting information from free-text radiology reports, enabling efficient data use, large-scale data mining, and a wide range of secondary analyses and clinical applications. This study aimed to evaluate the performance of LLMs in extracting diagnostically relevant information from multicenter free-text liver magnetic resonance imaging (MRI) reports, explore the clinical utility of LLM-generated structured reports, and investigate optimal prompting strategies for multicenter data. Methods:In this retrospective multicenter study, 800 free-text liver MRI reports from four medical centers (Beijing Friendship Hospital, Tianjin Medical University General Hospital, The Second Affiliated Hospital of Xi'an Jiaotong University, Sir Run Run Shaw Hospital) were collected to evaluate the information extraction performance of two LLMs-DeepSeek-V3 and ChatGPT-4o-using radiologist-annotated structured data as the reference standard. Three prompting strategies were applied: zero-shot prompting, global few-shot prompting (using shared examples across centers), and center-specific few-shot prompting (using examples specific to each center), with example counts set to 2-12. Model performance was evaluated using field-level F1 scores, and report-level extraction success was defined as the correct extraction of ≥80% report fields. Additionally, an exploratory clinical evaluation was conducted using 20 reports from one center, in which 10 radiologists and 10 clinicians evaluated the readability and clinical usability of free-text, manually structured, and LLM-generated reports on a 5-point Likert scale. Results:Few-shot prompting significantly outperformed zero-shot prompting for both LLMs, with the largest gains in macro F1 observed when k was increased from 0 to 2 (∆DeepSeek-V3: global 0.106, center-specific 0.127; ∆ ChatGPT-4o: global 0.086, center-specific 0.107). Performance plateaued at k=4 [DeepSeek-V3: global 0.848 (0.838-0.858), center-specific 0.865 (0.856-0.875); ChatGPT-4o: global 0.835 (0.824-0.845), center-specific 0.861 (0.851-0.870)], with adjacent-k gains <0.01. Center-specific prompting consistently outperformed global prompting (∆F1: 0.017-0.024 for DeepSeek-V3; 0.014-0.026 for ChatGPT-4o). In the exploratory clinical evaluation, structured reports received higher scores for clarity and communication than free-text reports (both P<0.001), while LLM-generated reports received scores comparable to those of manually structured reports (both P>0.05). Conclusions:LLMs demonstrated strong performance in extracting diagnostically relevant information from Chinese multicenter liver MRI reports. In the exploratory clinical evaluation, the LLM-generated structured reports showed the potential to improve report clarity and facilitate clinical communication. Global prompting showed good performance across centers, while center-specific prompting further improved accuracy by adapting to local reporting styles.
Renal biopsy has certain limitations for diagnosing membranous nephropathy (MN). The aim is to explore the value of MRI for diagnosing MN. MN patients were divided into two subgroups based on estimated glomerular filtration rate, including the mild group and moderate to severe group. Quantitative T1 mapping and renal blood flow (RBF) of bilateral kidneys were measured, including renal cortical T1 mapping (cT1) value, medullary T1 mapping (mT1) value, cortical RBF value (cRBF), and medullary RBF (mRBF) value. The Student’s t-test, Mann–Whitney U test, chi-square test, and one-way analysis of variance were used. Forty-seven MN patients and 54 matched healthy controls (HC) were prospectively enrolled. The cT1 and mT1 average values of HC were significantly lower than those of both MN subgroups (all p < 0.001) after adjusting for age and sex. Compared with the mild group and HC group, the moderate to severe group had lower cRBF (all p < 0.050) and mRBF average values (p = 0.012 and p < 0.001, respectively). The combination model of the T1 mapping and RBF values for differentiating MN from HC had a higher area under the curve of 0.87 (95
BACKGROUND:Although the clear cell likelihood score (ccLS) v2.0 demonstrates high specificity for clear cell renal cell carcinoma (ccRCC), its performance to characterize general malignancy in small renal masses (SRMs) remains limited. PURPOSE:To develop and validate a modified clear cell likelihood score (m-ccLS) incorporating the pseudocapsule to improve malignancy detection in SRMs while preserving specificity for diagnosing ccRCC. STUDY TYPE:This study was retrospective in type. SUBJECTS:352 patients with pathologically proven SRMs were included: development (n = 235), internal validation (n = 60), and external validation (n = 57). FIELD STRENGTH/SEQUENCE:Imaging was performed at 3.0 and 1.5 T using fast spin-echo T2-weighted imaging, single-shot echo planar diffusion-weighted imaging, 3D spoiled gradient echo (GRE) T1-weighted dynamic contrast-enhanced imaging, and in- and opposed-phase using T1-weighted GRE. ASSESSMENT:14 radiologists blinded to histopathology independently evaluated each SRM using ccLS v2.0 and m-ccLS scores in separate reading sessions; four, five, and five readers interpreted the development, internal, and external cohorts, respectively. STATISTICAL TESTS:Random-effects logistic regression, receiver operating characteristic curve, DeLong test, net reclassification improvement (NRI), integrated discrimination improvement (IDI), and Fleiss Kappa test were used. The statistical significance level was p < 0.05. RESULTS:For malignancy detection, m-ccLS showed a significantly higher area under the curve (AUC) than ccLS v2.0 across the development (0.850 vs. 0.772), internal validation (0.856 vs. 0.779), and external validation (0.803 vs. 0.720) cohorts with improved classification (NRI = 0.270, 0.045, and 0.028) and discrimination (IDI = 0.132, 0.206, and 0.120). For diagnosing ccRCC, m-ccLS and ccLS v2.0 showed similar results (0.908 vs. 0.894, p = 0.250; 0.912 vs. 0.898, p = 0.134; 0.865 vs. 0.838, p = 0.065) in development, internal, and external validation cohorts, respectively. m-ccLS category 3 contained fewer ccRCCs (33.3% vs. 72.5%; 15.8% vs. 47.2%; 7.6% vs. 26.9%) and malignancies (79.2% vs. 88.7%; 71.6% vs. 73.0%; 55.4% vs. 63.9%) than ccLS v2.0 category 3. DATA CONCLUSION:m-ccLS improves malignancy detection in SRMs compared with ccLS v2.0 without impairing diagnostic performance for ccRCC. EVIDENCE LEVEL:4. TECHNICAL EFFICACY:Stage 2.
Due to the scarcity of labeled data, semi-supervised segmentation learning has gained significant attention. However, accurate predictions of hard-to-identify boundaries in medical images remains challenging, especially when learning from unlabeled data. To address this issue, we propose a semi-supervised boundary-aware medical image segmentation method Via Symmetric Boundary-Foreground Collaboration (SBFC). Specifically, SBFC framwork constructs a symmetric dual-task segmentation (SDTS) network containing two symmetric segmentation models, each consisting of a dual-task U-shaped net with one encoder and two task-specific decoders for foreground and boundary segmentation. Using the predicted foreground, boundary probability maps and segmentation and derived boundary labels, a novel compound optimization objective function is proposed. This function integrates Cross-Task Consistency Regularization (CTCR) and Cross-Model Consistency Regularization (CMCR) for unlabeled data with supervised optimization for labeled data. Comprehensive experiments conducted on five public medical image datasets show that our method outperforms state-of-the-art comparative methods in terms of multiple consensus segmentation and boundary evaluation metrics.
Objectives To establish normative T1 and T2 relaxation times for the liver, pancreas, and kidneys at 5.0 T MRI, providing reference data for imaging parameter optimization and quantitative diagnosis. Materials and Methods Standardized T1/T2 phantoms were used for in vitro validation. From January to March 2025, healthy adults underwent 5.0 T abdominal MRI. T1/T2 maps were obtained using Modified Look-Locker Inversion Recovery (MOLLI) and T2-prepared gradient echo (T2 prep) sequences. Values, reproducibility, and correlations with age, sex, and other factors were evaluated. Results Results from the 10 tubes showed excellent agreement with MOLLI-based T1 and T2-prep-based T2 values and reference standards. The final cohort consisted of 41 participants. In healthy volunteers at 5.0 T MRI, mean T1 values were 1131.68±123.03ms(liver), 1164.63±57.81ms(pancreas), 1753.62±80.44ms (kidney cortex), and 2175.06±100.80ms (kidney medulla); mean T2 values were 31.16±3.37ms, 47.48±4.55ms, 59.36±6.35ms, and 41±4.07ms, respectively. T1 and T2 showed excellent reproducibility. Correlation analysis demonstrated a significant linear negative correlation between R2* and liver T1, and the R2* corrected liver T1 was 1156.57±69.25ms. Liver T2 decreased with age, and both measured and corrected liver T1 decreased with waist-to-hip ratio. Measured T1 and T2 of liver and kidney (cortex and medulla) were all lower in males than females, but liver corrected T1 showed no sex difference. Conclusion This study reports normative T1 and T2 relaxation times of the liver, pancreas, renal cortex, and medulla at 5.0 T MRI, confirming the method's feasibility and reproducibility, thus providing reference values for parameter optimization and future disease-related research. Key Points Question Quantitative 5.0 T MRI is diagnostically significant, but normative reference values for abdominal T1/T2 relaxation times are currently lacking.Findings We established normative T1/T2 values for abdominal organs at 5.0 T MRI, providing a reference for protocol optimization and diagnosis.Clinical relevance This study establishes normative T1 and T2 relaxation times for abdominal organs at 5.0 T MRI. These reproducible, quantitative reference values are essential for optimizing imaging protocols and provide a robust baseline for diagnosing abdominal diseases.
Few-shot medical image segmentation aims to delineate previously unseen anatomical structures using only a few labeled samples. Prototype-based methods have emerged as the dominant paradigm due to their strong generalization ability. However, most existing methods follow a query-agnostic prototype generation paradigm, overlooking intra-class variation. To overcome this limitation, we propose the IteRative Self-guided Prototype Enhancement Network (IR-SPENet), a query-centric framework that progressively refines support prototypes to better adapt to diverse query appearances. Specifically, we design a Hybrid Prototype Generation (HPG) module that constructs a global prototype to represent holistic semantics, together with an adaptive number of local prototypes to encode fine-grained details, enabling multi-granularity modeling of support-query discrepancies. Nevertheless, due to large appearance variations, not all local support prototypes are equally informative for a given query. To selectively emphasize relevant prototypes, we introduce a Query-guided Support Prototype Refinement (Q-SPR) module, which leverages optimal transport to re-weight local support prototypes according to query-specific information. By iteratively applying Q-SPR, IR-SPENet progressively enhances prototype quality and robustness against intra-class variation. Extensive experiments on three public medical image segmentation benchmarks demonstrate that IR-SPENet consistently outperforms existing methods, achieving leading performance.
PURPOSE:The purpose of this study was to accurately identify patients with locally advanced rectal cancer (LARC) who are likely to achieve pathologic complete response (pCR) after neoadjuvant therapy (NAT). This study develops a Multimodal Spatiotemporal Attentive Fusion Network (MSTAF-Net) to predict pCR and derives a Multimodal Spatiotemporal Signature (MSTAF-MSS) for disease-free survival (DFS) and explored its association with immune-related transcriptomic features. METHODS:This retrospective multicenter study included 642 patients with LARC. Longitudinal multiparametric magnetic resonance imaging (MRI) acquired before and after NAT and pretreatment hematoxylin and eosin-stained whole-slide images were collected. A dual-stream MSTAF-Net was designed to integrate longitudinal MRI features and pathomics features. Model performance was evaluated using receiver operating characteristic curves and survival analysis. Transcriptomic analyses were conducted to investigate the biologic correlates of the MSTAF-MSS and its association with immune-related transcriptomic features. RESULTS:The multimodal transformer fusion model achieved AUC values of 0.894 (95% CI, 0.847 to 0.942) in the internal validation cohort and 0.865 (95% CI, 0.786 to 0.943) in the external validation cohort, outperforming single-modality model. The MSTAF-MSS enabled effective risk stratification, with low-risk patients showing significantly longer DFS than high-risk patients (log-rank P < .05). Cox regression analyses identified MSTAF-MSS as an independent predictor of DFS in multivariate models (hazard ratio, 0.31 [95% CI, 0.13 to 0.73], P = .007). Distinct patterns of immune cell infiltration were observed between MSTAF-MSS-defined groups. CONCLUSION:The proposed MSTAF-Net integrates pathomics and longitudinal multiparametric MRI to capture spatial heterogeneity and treatment-related temporal dynamics of tumors. It demonstrates robust performance in predicting response to NAT and enables prognostic risk stratification. Furthermore, transcriptomic analysis supports potential biologic relevance of the model in associations with immune-related tumor biology.
Membranous nephropathy (MN) and IgA nephropathy (IgAN) exhibit similar clinical symptoms but differ substantially in treatment strategies and prognoses, highlighting the need for a non-invasive and precise diagnostic method. Multiparametric MRI (mpMRI) has shown great potential for the non-invasive assessment and quantification of renal function. This study aims to explore the potential value of mpMRI in distinguishing between MN and IgAN. Cortical renal blood flow (cRBF), cortical true diffusion coefficient (cD), cortical pseudo-diffusion coefficient (cD*), cortical perfusion fraction (cf), cortical R2* (cR2*), cortical T1 mapping (cT1), cortical mean diffusivity (cMD), and cortical mean kurtosis (cMK) were prospectively measured in 58 patients with MN, 47 patients with IgAN, and 75 healthy controls. One-way analysis of variance was used to compare MRI parameters among the three groups. The correlations between laboratory data and MRI parameters were evaluated using Spearman correlation analysis. Logistic regression analysis was performed to construct diagnostic models. The diagnostic performances of models for differentiating MN and IgAN were assessed using receiver operating characteristic curves with area under the curve (AUC). Compared with the IgAN group, patients with MN were older and had higher cRBF, cf, and cT1 values (all P < 0.050), but cMK value was significantly lower than that in IgAN (P = 0.009). In the MN group, a significant correlation was observed between eGFR and cf. (r = 0.61, P < 0.001), whereas in the IgAN group, eGFR showed a significant correlation with cRBF (r = 0.51, P < 0.001). Age, eGFR, 24 h urinary protein, and cMK demonstrated the ability to distinguish between MN and IgAN in both univariable and multivariate logistic regression analysis and were used to construct diagnostic models. Among the multiparametric models, the combination model incorporating age, eGFR, 24 h urinary protein, and cMK exhibited the highest AUC of 0.92 (95
OBJECTIVE:This study aims to develop a cascaded deep learning (DL) system based on multiparametric MRI to establish an automated pipeline for the segmentation and classification of small renal masses (SRMs). MATERIALS AND METHODS:A retrospective collection of SRM patients with pathologically confirmed from three institutions was conducted. MRI data from Institution 1 were randomly divided into a training set and an internal test set. Data from other institutions served as the external test set. A cascaded DL system was developed, incorporating automated segmentation and benign-malignant classification. Diagnostic performance was evaluated using receiver operating characteristic analysis and compared against three radiologists of varying experience. RESULTS:A total of 965 patients with SRM were included. Institution 1 contributed 888 cases, with 712 used for training and 176 as an internal test set; Institutions 2 and 3 provided 77 cases as an external test set. The optimal classification model using automated segmentation labels achieved AUCs of 0.936 and 0.788 on internal and external test sets, respectively. Performance was comparable to models using manual segmentation (internal: 0.936 vs. 0.944, P = 0.671; external: 0.788 vs. 0.832, P = 0.629). On the external test set, the model performed comparably to the senior radiologist, while it significantly outperformed the senior radiologist on the internal test set. The model significantly outperformed the junior radiologist on both test sets. This finding remained consistent in the subgroup of tumors smaller than 3 cm. CONCLUSION:The cascaded DL system demonstrated robust performance across multiple centers, enabling non-invasive and efficient discrimination of SRM malignancy, showing promise as a clinical support tool.
OBJECTIVES:To evaluate the ability of the maximum standardized uptake value (SUVmax) to predict the lymphovascular space invasion (LVSI) status in endometrial cancer (EC). METHOD:PubMed/MEDLINE, Web of Science, Embase, and the Cochrane Library were systematically searched for all original studies evaluating the diagnostic efficacy of LVSI using PET/CT or PET/MR. Methodological quality was assessed using the Quality Assessment of Diagnostic Accuracy Studies 2 (QUADAS-2). A bivariate random effects model was used to acquire pooled sensitivity, specificity, heterogeneity, and the area under the summary receiver operating characteristic curve (AUROC). Meta-regression and sensitivity analysis were performed to identify sources of heterogeneity. RESULTS:A total of 6 studies (257 patients) were included. Most studies had a low risk of bias, and all studies had minimal applicability concerns. The summary AUROC values, pooled sensitivity and specificity of SUVmax in detecting LVSI in EC were 0.77, 62% and 83%, respectively. One study may have contributed to the unstable results of this study according to the sensitivity analysis. CONCLUSION:Our study showed that SUVmax has moderate accuracy in noninvasively predicting LVSI in EC. More original studies with large samples are needed in the future to evaluate the role of SUVmax in differentiating LVSI. Advances in knowledge: LVSI is closely related to the prognosis of EC, and it can only be obtained by surgical pathology. SUVmax has moderate diagnostic performance in preoperatively predicting LVSI in EC. Future studies with large samples are needed to confirm the clinical value of SUVmax in the preoperative prediction of LVSI.
Multi-modal medical image segmentation plays a pivotal role in the diagnosis, treatment planning, and monitoring of focal liver lesions (FLLs). However, obtaining sufficient multimodal labels is challenging due to the high cost and labor intensity of manual annotation. Additionally, spatial misalignment between modalities acquired at different times or from diverse sources can hinder accurate segmentation. To address these issues, we propose automated MRI focal liver lesion segmentation based on multimodal alignment and interaction, which contains an Unsupervised Cross-Modal Interaction based Registration (UCMIR) module and a Multi-Scale Modality-Contribution-Aware multimodal medical image segmentation network (MSMCA). UCMIR performs multiscale cross-modal feature interaction and registration to align unlabeled modalities with labeled ones, generating a deformation field for medical image registration. The aligned image pairs are then fed into the MSMCA network to obtain the final segmentation result. MSMCA effectively fuses multiscale information from different modalities through coordinate attention, boosting segmentation performance. Experimental results on a focal liver lesions dataset demonstrate that the Dice values of our approach achieve 7.34% and 4.40% improvement compared with Tri-Attention Net in two testing modality groups, respectively.
To develop and validate a machine learning (ML)-based pipeline for automated segmentation and classification of complicated cystic renal masses (cCRMs) on MRI. This multicenter retrospective study enrolled 275 patients (median age, 48 years; 85 females) with pathologically confirmed 275 cCRMs (203 malignant) who underwent renal MRI from January 2013 to December 2023. cCRMs from one institution were used as a training set (n = 215), while those from the other three institutions served as a test set (n = 60). 3D V-Net and random forest algorithms were employed for segmentation and classification, respectively. Segmentation and classification performance was evaluated using the Dice similarity coefficient (DSC) and the area under the curve (AUC), respectively. Two junior and two senior radiologists independently classified cCRMs in the test set into Bosniak categories II–IV based on the Bosniak classification, version 2019. In the test set, the ML pipeline achieved DSC of 0.718 for cCRMs (n = 60) on excretory phase images. Additionally, classification performance of the ML pipeline (AUC = 0.835, 95
OBJECTIVES:This study aims to develop an artificial intelligence (AI)-based automated segmentation method for small renal masses (SRMs) using multi-center, multi-scanner, multi-sequence MRI data. METHODS:MR images from 988 pathologically confirmed SRM patients from three different centers were retrospectively included. Segmentation networks were independently developed for each MRI sequence using deep learning techniques. A GE dataset of 733 patients from Center 1 was used for training and validation. A GE test set, consisting of internal (99 from Center 1) and external test sets (81 from Center 2 and 3), was created for evaluation. Furthermore, a non-GE generalization set, consisting of 75 patients from Center 2 and 3, was used to assess the generalization ability. The method's performance was evaluated in terms of detection rate and segmentation accuracy (Dice similarity coefficient [DSC]). Subgroup analysis and multiple linear regression were used for further exploration. RESULTS:Our method demonstrated promising results in the detection and segmentation of SRMs. All patients in the GE test set were correctly detected in at least one sequence. Our model achieved a median DSC of 0.769-0.855 across five MRI sequences and demonstrated reasonable generalization to non-GE scanners (median DSC range: 0.523-0.785). CONCLUSIONS:The implementation of automated segmentation achieved encouraging outcomes in both correct-detection rates and segmentation accuracy across a diverse cohort spanning multiple centers and scanners, suggesting its potential as a key component of future diagnostic pipelines for SRMs.
This study aimed to develop an interpretable, domain-generalizable deep learning model for microvascular invasion (MVI) assessment in hepatocellular carcinoma (HCC). Utilizing a retrospective dataset of 546 HCC patients from five centers, we developed and validated a clinical-radiological model and deep learning models aimed at MVI prediction. The models were developed on a dataset of 263 cases consisting of data from three centers, internally validated on a set of 66 patients, and externally tested on two independent sets. An adversarial network-based deep learning (AD-DL) model was developed to learn domain-invariant features from multiple centers within the training set. The area under the receiver operating characteristic curve (AUC) was calculated using pathological MVI status. With the best-performed model, early recurrence-free survival (ERFS) stratification was validated on the external test set by the log-rank test, and the differentially expressed genes (DEGs) associated with MVI status were tested on the RNA sequencing analysis of the Cancer Imaging Archive. The AD-DL model demonstrated the highest diagnostic performance and generalizability with an AUC of 0.793 in the internal test set, 0.801 in external test set 1, and 0.773 in external test set 2. The model’s prediction of MVI status also demonstrated a significant correlation with ERFS (p = 0.048). DEGs associated with MVI status were primarily enriched in the metabolic processes and the Wnt signaling pathway, and the epithelial-mesenchymal transition process. The AD-DL model allows preoperative MVI prediction and ERFS stratification in HCC patients, which has a good generalizability and biological interpretability. The adversarial network-based deep learning model predicts MVI status well in HCC patients and demonstrates good generalizability. By integrating bioinformatics analysis of the model’s predictions, it achieves biological interpretability, facilitating its clinical translation.