Objective Recurrent laryngeal nerve lymph nodes (RLNLN) dissection in resectable esophageal squamous cell carcinoma (ESCC) is challenging due to increased post-operative complications and unfavorable outcomes. We aimed to develop and validate a CT-based radiomic model to predict RLNLN metastasis in ESCC patients to optimize treatment strategies. Methods We retrospectively enrolled 645 ESCC patients from four centers (2015-2022) with pathologically confirmed RLNLN status, stratified into a training cohort, an internal validation cohort, and two external validation cohorts to ensure model generalizability. The radiomics model was built from both non-contrast and contrast-enhanced CT images. The CT-based combined model was developed by integrating the radiomic model with clinicopathological factors (including T/N stage, tumor location, and lymph node size) through multivariate logistic regression. The model’s performance was evaluated for discriminative ability, calibration, and clinical usefulness. Results Incorporating ten features, the radiomics model effectively distinguished ESCC patients with RLNLN metastasis from those without, achieving AUCs of 0.819 in training, 0.708 in internal validation, and 0.806 and 0.635 in the two external validation cohorts. The combined model demonstrated strong discriminative ability, achieving AUCs of 0.866 in the training cohort, 0.749 in the internal validation cohort, and 0.710 and 0.694 in the two external validation cohorts. Decision curve analysis validated the clinical utility of the combined model across various threshold probabilities, highlighting its potential to inform treatment decisions in ESCC patients. Conclusion The CT-based combined model demonstrated robust predictive performance for RLNLN metastasis in resectable ESCC across multiple centers, providing clinically actionable support for personalized therapeutic decision-making.
Preoperative discrimination between follicular thyroid carcinoma (FTC) and follicular thyroid adenoma (FTA) remains challenging, as imaging and cytological approaches often show limited efficacy. Even fine-needle aspiration (FNA) biopsy and intraoperative frozen sections frequently fail to provide conclusive results. Thus, follicular thyroid neoplasms (FNs) typically necessitate complete surgical excision for definitive diagnosis, leading to unnecessary thyroidectomies for benign conditions or delayed treatment for malignancies. To address this gap, we developed FTC-Net, a vision-language foundation model, to preoperatively classify FNs using ultrasound images. In a multicenter retrospective study of 2421 patients (6477 images) from 14 institutions, FTC-Net was trained on 1462 patients and validated in two independent cohorts (n = 578 and n = 381). FTC-Net achieved AUCs of 0.836 and 0.841 in external validation, outperforming benchmark deep learning models and established TI-RADS systems. It also substantially reduced both total FNA rates and unnecessary FNA rates compared to ACR TI-RADS and C-TI-RADS. FTC-Net has the potential to serve as a non-invasive and advanced tool for the preoperative diagnosis of FNs, thereby improving clinical decision-making and reducing unnecessary procedures.
BACKGROUND AND PURPOSE:The cerebellum is increasingly recognized as a key contributor to language and cognitive processing, but its dynamic network alterations in poststroke aphasia remain poorly understood. This study investigated dynamic cerebellar networks in patients with poststroke aphasia using resting-state functional MRI. We examined intracerebellar and cerebellar-cortical dynamic functional connectivity and quantified their temporal properties and graph-theoretical topology. MATERIALS AND METHODS:Seventy-seven right-handed patients with poststroke aphasia and 79 healthy controls underwent 3T resting-state functional MRI. Dynamic cerebellar functional networks were constructed using the Seitzman-27 cerebellar atlas. A sliding window approach (30 TR window, 1 TR step) was applied, followed by K-means clustering to identify distinct connectivity states. Graph-theoretical analyses were performed to quantify state-specific network topology. The variability of dynamic functional connectivity between the cerebellar and cortical regions was calculated. Partial correlation analyses were performed to examine the relationships among dynamic network measures, lesion volume, and language and cognitive function. RESULTS:Two cerebellar dynamic functional connectivity states were identified in poststroke aphasia: a predominant segregated state (78.93%) with widespread reductions in connectivity and decreased clustering coefficient (d = -1.29), characteristic path length (d = -0.62), and local efficiency (d = -1.11) but higher global efficiency (d = 1.06) and a less frequent integrated state (21.07%) with enhanced connectivity and a higher clustering coefficient (d = 0.57) and characteristic path length (d = 0.70) and diminished global efficiency (d = -1.25) and small-worldness (d = -0.92) and small-world index (d = -0.89). Poststroke aphasia showed reduced variability of dynamic functional connectivity between the cerebellar and cortical regions involved in language and cognition (Gaussian random field correction, voxel-level P < .001, cluster-level P < .05). Lesion volume negatively correlated with Aphasia Quotient, repetition, memory, executive function, and attention (P < .005). State-specific network metrics and variability measures were associated with language and cognitive performance independent of lesion volume. CONCLUSIONS:Patients with poststroke aphasia exhibited a segregated cerebellar state with reduced intracerebellar connectivity and efficiency and an integrated state with enhanced connectivity and small-world properties, together with reduced variability in cerebellar-cortical connections to language- and cognition-related regions. These state-specific network alterations were linked to distinct behavioral domains independent of lesion volume, highlighting a dissociation between structural constraints and dynamic, lesion-independent plasticity.
Rationale and Objectives This study evaluated the efficacy of a combined artificial intelligence (AI)-assisted and traditional teaching model in enhancing radiology residents' diagnostic competency for pulmonary nodules on chest Combined Teaching (CT). Materials and Methods In this randomized controlled trial, 36 s-year residents were allocated to three groups: Traditional Teaching (TT), AI-Assisted Teaching (AI-AT), and CT. All received standardized theory instruction. TT underwent 2 h of mentored film-reading. AI-AT completed 2 h of AI platform training. CT received 1 h of AI training followed by 1 h of expert-led review of challenging cases. Assessments included theoretical and practical exams, diagnostic process metrics, and a 1-month follow-up survey on confidence, cognitive load, and AI perceptions. Results The CT group significantly outperformed both the TT and AI-AT groups in post-test theoretical (CT: 91.8 ± 3.4 vs. TT: 82.6 ± 3.2, P < 0.001; vs. AI-AT: 86.2 ± 2.9, P = 0.003) and practical scores (CT: 93.5 ± 2.7 vs. TT: 81.9 ± 3.0, P < 0.001; vs. AI-AT: 86.8 ± 3.0, P < 0.001). Analysis of diagnostic process metrics revealed the CT group engaged in more deliberate analysis (longer reading time, P = 0.01) and demonstrated superior completeness in describing key imaging features (P < 0.001). At the 1-month follow-up, the CT group maintained higher diagnostic confidence (P = 0.002) and exhibited a more balanced perception of AI as a collaborative tool. Conclusion This exploratory randomized controlled trial demonstrates that the integration of AI-AT with traditional expert-led sessions is significantly more effective than either method alone in improving radiology residents’ diagnostic performance and clinical reasoning for pulmonary nodules. This combined model fosters deeper cognitive engagement and sustainable learning outcomes, providing a framework for competency-based education in the AI era.
Preoperative distinction between follicular thyroid carcinoma (FTC) and follicular thyroid adenoma (FTA) remains a clinical challenge, largely due to the substantial cytomorphological overlap observed in fine-needle aspiration samples. We hypothesised that deep learning could capture subvisual cytomorphological patterns that reflect underlying biological differences between these entities.Preoperative distinction between follicular thyroid carcinoma (FTC) and follicular thyroid adenoma (FTA) remains a clinical challenge, largely due to the substantial cytomorphological overlap observed in fine-needle aspiration samples. entities. We hypothesised that deep learning could capture subvisual cytomorphological patterns that reflect underlying biological differences between these entities. This retrospective multicentre study enrolled 255 patients with surgically confirmed follicular thyroid neoplasms (FTNs) across 10 institutions, all of whom underwent preoperative liquid-based cytology (LBC) testing. Patients were allocated to a training cohort (n = 127; 48,560 patches), an internal validation cohort (n = 44; 17,226 patches), and an external validation cohort (n = 84; 17,192 patches). We developed a multi‑bag clustering‑ constrained attention multiple instance learning (CLAM_MB) model to classify LBC whole‑slide images. The model leveraged pathologist‑annotated follicular cell clusters to guide instance sampling without requiring instance‑level labels, thereby enabling phenotype pattern discovery under weak supervision. Model performance was benchmarked against single‑bag CLAM (CLAM_SB) and max‑pooling multiple instance learning (MIL) architectures. The CLAM_MB model demonstrated robust classification performance across all cohorts, with no significant performance degradation between validation settings (all P > 0.05). In the internal validation cohort, CLAM_MB achieved an area under the receiver operating characteristic curve (AUC) of 0.817, significantly outperforming CLAM_SB (0.766) and max‑pooling MIL (0.786; both P < 0.001). Similarly, in the external validation cohort, CLAM_MB yielded an AUC of 0.827, compared with 0.796 for CLAM_SB and 0.619 for max‑pooling MIL (both P < 0.001). At optimal cutoffs, the model attained sensitivities of 68.8% and 81.8%, and specificities of 78.4% and 92.9% in the internal and external validation cohorts, respectively. Subgroup analyses further confirmed consistent diagnostic performance across age and sex strata (both P > 0.05). The CLAM_MB model effectively identifies subvisual cytomorphological signatures in LBC samples that distinguish FTC from FTA with high accuracy. These findings support its potential utility as a preoperative decision-support tool for the differential diagnosis of FTNs, warranting further prospective validation. None
Purpose To develop and validate a deep learning model integrating tumor and visceral adipose tissue (VAT) CT scan features with clinical indicators to predict postoperative peritoneal metastasis in serosa-invasive gastric cancer. Materials and Methods This multicenter, retrospective study between April 2008 and January 2018 included patients with pathologically confirmed serosa-invasive gastric cancer. Patients were divided into training, internal test, and independent external test sets. Tumor and VAT regions were segmented at preoperative CT. Deep features were extracted using a ResNet18 network. A fused tumor-VAT deep learning signature (F-DLS) was generated, incorporating clinical variables into a multimodal deep learning radiomics model (MDLR) using a sparse Bayesian extreme learning machine. Model performance was assessed using receiver operating characteristic curve, integrated discrimination improvement, calibration, decision curve analysis, and recurrence-free survival. Results Among 416 patients (mean age, 56.6 years ± 11.6; 66.1% male patients), the F-DLS achieved area under the receiver operating characteristic curve (AUC) values of 0.81 (95% CI: 0.73, 0.88) in the internal test set and 0.79 (95% CI: 0.71, 0.86) in the external test set. Compared with the tumor tissue DLS and VAT-DLS, the F-DLS showed numerically higher AUCs without statistical significance. The MDLR achieved the strongest predictive performance, with AUCs of 0.86 (95% CI: 0.79, 0.92) in the internal test set and 0.86 (95% CI: 0.78, 0.92) in the external test set. The MDLR statistically significantly outperformed clinical and deep learning-only models (integrated discrimination improvement, P < .001), showed good calibration, and provided favorable net benefit on decision curve analysis. High-risk patients identified by the MDLR had significantly shorter recurrence-free survival (log-rank P < .001). Conclusion The MDLR integrating CT scan features and clinical indicators enabled noninvasive prediction of peritoneal metastasis risk in serosa-invasive gastric cancer and may facilitate postoperative risk stratification. Keywords: Gastric Cancer, Peritoneal Metastasis, CT, Visceral Adipose Tissue, Deep Learning Supplemental material is available for this article. © RSNA, 2026.
Here, we design a pH-responsive nanoplatform, PD@MIL, for photodynamic therapy (PDT) against nasopharyngeal carcinoma (NPC). The NH2-MIL-101(Fe) is employed to co-encapsulate pyropheophorbide-a (PPa) and doxycycline (Doxy) with high loading capacities (34.6 % and 39.1 %, respectively), which can be released via a pH-triggered method. Within the acidic tumor microenvironment (TME), PD@MIL generates cytotoxic reactive oxygen species (ROS) under near-infrared (NIR) irradiation to selectively eradicate NPC cells. Crucially, the platform simultaneously normalizes tumor vasculature, reducing metastasis and ameliorating hypoxia. Western blot (WB) and magnetic resonance imaging (MRI) analyses confirm that Doxy drives this vascular normalization by suppressing vascular endothelial growth factor A (VEGFA) expression during PDT, thereby enhancing oxygen perfusion. This synergetic PDT nanoplatform operates at an ultralow injected dose, underscoring the potential of intelligent, TME-responsive nanodelivery systems to overcome key limitations of conventional PDT for clinical translation.
Early detection of nasopharyngeal carcinoma through Epstein-Barr virus serology is hampered by a low positive predictive value. This study aims to develop a hierarchical dynamic model to refine risk stratification among individuals initially identified as medium- or high-risk by Epstein-Barr virus serology. By integrating longitudinal Epstein-Barr virus antibody data with age, sex, and family history, the model is trained using data from the PRO-NPC-001 program. In the validation cohort, the high-risk model using one-year data achieves a positive predictive value of 18.2% (about a fourfold increase over serology screening), with a negative predictive value of 97.7% and an area under the curve of 0.783. With two-year data, the positive predictive values for the high-risk and medium-risk models are 8.8% and 1.1%, respectively, with area under the curve values of 0.859 and 0.687. Compared with serology-only screening, the hierarchical dynamic models reduce the need for follow-up examinations by 74.2% in high-risk individuals, yielding cost savings of up to 65.6%. These findings demonstrate that hierarchical dynamic models significantly enhance current serological screening strategies for nasopharyngeal carcinoma, though further prospective validation is warranted.
Background:Conversion therapies after immune checkpoint inhibitors (ICIs) plus tyrosine-kinase inhibitors (TKIs) provide curative surgery chance and prolong survival for unresectable hepatocellular carcinoma (uHCC). However, only some patients have the opportunity to receive conversion therapies. To this end, we aimed to develop and validate a machine-learning model to identify patients who may have the chance to undergo conversion therapy. Methods:This retrospective cohort study included 443 patients with uHCC who received ICIs and TKIs from four centers. Variables were analyzed using univariate and multivariate logistic regression to identify independent indicators of conversion therapy. The Gradient Boosting Machine (GBM) algorithm was used to develop and validate model, and the Shapley additive explanation algorithm was used to mechanically explain the prediction of the model. Results:Overall, 84 (19%) patients underwent conversion therapy, and their prognosis were significantly longer than those did not (P < 0.05). CA125 level, pre-TKI therapy, pre-antiviral therapy, lymph node metastasis status, and number of intrahepatic lesions were identified as indicators of conversion therapy. The GBM-based combined model outperformed the BCLC classification (P < 0.05), yielding an AUC of 0.76 and 0.74 in the training and external validation cohorts, respectively. Survival analyses indicated that patients who underwent surgery as conversion therapy had a better prognosis than those who underwent ablation therapy (P < 0.05). Conclusion:The GBM-based combined model could identify patients who may benefit from conversion therapy for uHCC treated with ICIs and TKIs. Surgical resection as curative conversion therapy may provide better survival benefits than ablation therapy.
Hypoxia can substantially impact clinical outcomes in patients with head and neck cancer (HNC) by promoting tumor invasion, metastasis, immune escape, and therapy resistance. Given the growing interest in targeting hypoxia for cancer therapy, noninvasive methods are needed to accurately detect hypoxia and evaluate the tumor response to treatment. This review summarizes recent advances in hypoxia-targeted probes and imaging techniques, emphasizing their imaging mechanisms, strengths, and limitations. We focused on the promising clinical applications of hypoxia imaging, especially those currently used in clinics, such as positron emission tomography and magnetic resonance imaging, and highlighted their roles in guiding personalized therapy. Future directions include optimizing imaging probes to improve safety profiles, integrating multimodal imaging, applying machine learning models to analyze multiparametric data, and establishing standardized 3-dimensional in vitro models to better mimic hypoxia heterogeneity. These advancements are expected to considerably improve the management of patients with HNC.
INTRODUCTION:Placenta accreta spectrum (PAS) disorders result from abnormal placental attachment, leading to varying degrees of myometrial invasion. Magnetic resonance imaging (MRI) plays a crucial role in assessing the depth and extent of placental invasion. This study aims to evaluate the correlation between quantified MRI findings and the diagnosis of PAS, as classified according to the FIGO system. MATERIALS AND METHODS:A retrospective analysis was conducted on 556 high-risk PAS patients, defined as those with placenta previa or a history of previous cesarean sections. Ten predefined MRI signs were assessed board certified radiologists. Multivariate logistic regression was used to identify independent predictors of invasive PAS. The positive predictive value (PPV) and negative predictive value (NPV) were calculated to assess the diagnostic performance of signs. RESULTS:Among the 556 cases, 150 (26.98 %) were classified as non-PAS, 180 (32.37 %) as placenta accreta, 158 (28.42 %) as placenta increta, and 68 (12.23 %) as placenta percreta. Four MRI signs were identified as significant predictors of invasive PAS: bladder wall interruption (odd ratio [OR] = 160.17), placental ischemic infarction (OR = 19.91), placental protrusion (OR = 14.66), and myometrial thinning (OR = 14.07). The PPV of these signs ranged from 70 % to 85 %, while the NPV ranged from 65 % to 72 %. Multivariate analysis confirmed these MRI findings as independent predictors of invasive PAS. CONCLUSIONS:This study identified four key MRI signs as reliable predictors of invasive PAS, which can effectively inform clinical decision-making regarding surgical interventions, such as cesarean hysterectomy.
Virtual reality (VR) simulation has become a useful tool for students to develop clinical skills in medical schools. To investigate the effectiveness of VR simulation teaching thyroid ultrasonography skills in medical students. A total of 106 medical students who had finished basic medical courses were recruited and randomized to VR group and control group. All participants received 2 theoretical lessons, followed by a practical class. The control group had traditional PowerPoint (PPT) teaching to learn the thyroid ultrasonography process, while VR group studied the complete process via the VR simulation platform after PPT teaching. All students took a skill examination to evaluate their performance after two weeks. They were scored by 2 examiners who were blinded to group assignment on the global rating scale of ultrasonography, with a total score of 40 points for 8 components. A month later, all participants were required to complete the thyroid ultrasound examination on the VR platform. Ninety-one students completed the program (VR group, n=47; control group, n=44). The score of VR group (median, 32.5; interquartile range [IQR], 29.5-35.0) outperformed the control group (median, 31.0; IQR, 27.5-33.0, p=0.028) on the global rating scale of thyroid ultrasound examination. Meanwhile, students in the VR group (mean score, 51.5; time, 16.2min) demonstrated better performance and less time than the control group (mean score, 44.3; time, 23.7min) when assessed in the VR platform. The VR simulation platform of thyroid ultrasound examination showed promising results in improving medical students’ skills, making it an alternative or additional approach to traditional teaching patterns in medical education. ChiCTR2200064833
The clinical significance, digital attributes, and underlying high‐dimensional information in medical images make them a key area for the artificial intelligence (AI) revolution in health care. Generative AIs (GAIs) provide unprecedented abilities in synthesizing diverse and accurate simulated medical images for AI model training as well as personalized disease management. However, several hurdles must be overcome prior to clinical implementation, such as biases introduced during training in synthesized images and the risk of medical and research falsification. This review outlines the current landscape of medical image synthesis through GAIs, with a specific focus on the variety of medical images to be synthesized, various real‐world issues to be solved, and the evaluation of the quality and utility of the synthesized images. We finally summarize the key challenges, propose potential solutions, and highlight promising directions for future research, with the aim of providing guidance for upcoming research.
BACKGROUND AND PURPOSE:Deep learning can non-invasively depict the radiological phenotype of tumor. We aimed to propose an end-to-end deep learning framework called NPC-SurvAI to perform prognosis assessment using MRI in nasopharyngeal carcinoma (NPC). METHODS:This retrospective study included 2180 NPC patients who underwent baseline MRI. The NPC-SurvAI comprised an AttVNet for image segmentation and a DenseNet-ICAM for prognosis evaluation, including progression-free survival (PFS) and overall survival (OS). The clinical model was built with age, T-stage, N-stage, and EBV DNA. The image and combined models were developed by the NPC-SurvAI framework. The integrated area under the curve (iAUC) and thetime-dependent AUC (tAUC) were leveraged to measure the predictive accuracy. K-means clustering and Kaplan-Meier survival analysis were utilized to stratify patients into subtypes and compare their prognoses. RESULTS:In the validation cohorts, the AttVNet achieved average Dice similarity coefficients of 0.726-0.764 tumor segmentation. The dynamic change curves of the AUCs over time suggested that the combined model outperformed both the clinical and image models in predicting PFS (iAUC: 0.838-0.884 vs 0.788-0.844 vs 0.738-0.798) and OS (iAUC: 0.842-0.894 vs 0.793-0.853 vs 0.754-0.807) at any time point from 1 to 8 years. Specially, the combined model achieved time-AUCs of 0.844-0.930 for 3-year PFS and 0.827-0.896 for 5-year PFS; 0.838-0.978 for 3-year OS and 0.788-0.871 for 5-year OS. Additionally, patients could be stratified into two subtypes with different survivals (all P < 0.05). CONCLUSIONS:NPC-SurvAI has the potential to automatically stratify patients with diverse prognoses, which helps clinicians in optimizing treatment decisions and surveillance.
Background:Conventional diagnostic tools, including ultrasound, fine-needle aspiration cytology, and intraoperative frozen section pathology, may fail to reliably distinguish between benign and malignant follicular-patterned thyroid neoplasms (FNs), leading to unnecessary or inadequate surgical interventions. We aimed to develop and validate a deep learning (DL) system for the preoperative diagnosis of FNs using routine ultrasound images, with the goal of improving diagnostic accuracy and reducing unnecessary procedures. Methods:In this multicenter, retrospective study, we included 3817 patients (2877 [75.4%] female) with a definitive diagnosis of FNs from 11 centers across China. All patients underwent preoperative ultrasound examinations. The dataset comprised 9393 ultrasound images, including thyroid follicular adenoma (n = 1787, 4317 images), follicular carcinoma (n = 446, 1593 images), and follicular variant of papillary thyroid carcinoma (n = 1584, 3483 images) collected between 2012 and 2025. A state-of-the-art OverLoCK (Overview-first-Look-Closely-next ConvNet with Context-Mixing Dynamic Kernels) model was developed on a dataset comprising 2728 patients (6625 images) and validated on an internal cohort (n = 683, 1905 images) and an external cohort (n = 406, 863 images). Model performance was evaluated using the area under the curve (AUC), accuracy, sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), and F1 score. Model calibration was evaluated using calibration curves, while clinical usefulness was assessed through decision curve analysis (DCA). Findings:The OverLoCK model exhibited excellent performance in both the internal and external validation sets. In the internal validation cohort, the OverLoCK model achieved an AUC of 0.937 (95% confidence interval [CI]: 0.919-0.954), with accuracy of 90.9% (95% CI: 87.7-92.0), sensitivity of 93.9% (95% CI: 91.5-95.6), specificity of 84.8% (95% CI: 82.6-86.0), PPV of 92.7% (95% CI: 90.7-93.8), NPV of 87.2% (95% CI: 86.0-91.0), and F1 score of 0.911 (95% CI: 0.887-0.932). In the external validation cohort, the model yielded an AUC of 0.853 (95% CI: 0.832-0.876), accuracy of 82.8% (95% CI: 81.7-84.4), sensitivity of 84.5% (95% CI: 82.5-86.2), specificity of 81.1% (95% CI: 79.2-84.5), PPV of 80.4% (95% CI: 79.0-84.0), NPV of 85.1% (95% CI: 83.2-87.7), and F1 score of 0.839 (95% CI: 0.802-0.877). The DL model demonstrates good agreement between the predicted and actual probabilities of malignancy. DCA confirmed that the model was clinically useful. Interpretation:Our study demonstrates that a DL-based system can provide a noninvasive, accurate, and reliable tool for the preoperative diagnosis of FNs. By improving diagnostic precision, this approach has the potential to optimize clinical decision-making and reduce the burden of overtreatment in patients with FNs. Further prospective studies are warranted to validate these findings in real-world clinical settings. Funding:This work was supported by the National Key Research and Development Program of China (2023YFF1204600), the National Natural Science Foundation of China (82227802 and 82302190), the Clinical Frontier Technology Program of the First Affiliated Hospital of Jinan University (No. JNU1AF-CFTP-2022-a01201), the Science and Technology Projects in Guangzhou (202201020022, 2023A03J1036, 2023A03J1038, 2025A04J7006), the Outstanding Young Talents of Guangdong Special Support Program (Health Commission of Guangdong Province) (0720240213), and the Science and Technology Youth Talent Nurturing Program of Jinan University (21623209).
OBJECTIVE:To propose a deep learning (DL) system for the preoperative diagnosis of follicular-like thyroid neoplasms (FNs) using routine ultrasound images. SUMMARY BACKGROUND DATA:Preoperative diagnosis of malignancy in nodules suspicious for an FN remains challenging. Ultrasound, fine-needle aspiration cytology, and intraoperative frozen section pathology cannot unambiguously distinguish between benign and malignant FNs, leading to unnecessary biopsies and operations in benign nodules. METHODS:This multicenter, retrospective study included 3634 patients who underwent ultrasound and received a definite diagnosis of FN from 11 centers, comprising thyroid follicular adenoma (n=1748), follicular carcinoma (n=299), and follicular variant of papillary thyroid carcinoma (n=1587). Four DL models including Inception-v3, ResNet50, Inception-ResNet-v2, and DenseNet161 were constructed on a training set (n=2587, 6178 images) and were verified on an internal validation set (n=648, 1633 images) and an external validation set (n=399, 847 images). The diagnostic efficacy of the DL models was evaluated against the ACR TI-RADS regarding the area under the curve (AUC), sensitivity, specificity, and unnecessary biopsy rate. RESULTS:When externally validated, the four DL models yielded robust and comparable performance, with AUCs of 82.2%-85.2%, sensitivities of 69.6%-76.0%, and specificities of 84.1%-89.2%, which outperformed the ACR TI-RADS. Compared to ACR TI-RADS, the DL models showed a higher biopsy rate of malignancy (71.6% -79.9% vs 37.7%, P<0.001) and a significantly lower unnecessary FNAB rate (8.5% -12.8% vs 40.7%, P<0.001). CONCLUSION:This study provides a noninvasive DL tool for accurate preoperative diagnosis of FNs, showing better performance than ACR TI-RADS and reducing unnecessary invasive interventions.