Medical ultrasound (US) image segmentation faces significant challenges due to limited labeled data and characteristic imaging artifacts, including speckle noise and low-contrast boundaries. While semi-supervised learning (SSL) approaches have emerged to address data scarcity, existing methods suffer from suboptimal unlabeled data utilization and lack robust feature representation mechanisms. In this article, we propose Switch, a novel SSL framework with two key innovations: 1) a multiscale switch (MSS) strategy that employs hierarchical patch mixing to achieve uniform spatial coverage; and 2) a frequency-domain switch (FDS) with contrastive learning that performs amplitude switching in Fourier space for robust feature representations. Our framework integrates these components within a teacher-student architecture to effectively leverage both labeled and unlabeled data. Comprehensive evaluation across six diverse US datasets (lymph nodes, breast lesions, thyroid nodules, and prostate) demonstrates consistent superiority over state-of-the-art (SOTA) methods. At a 5% labeling ratio, Switch achieves remarkable improvements: 80.04% Dice on LN-INT, 85.52% Dice on DDTI, and 83.48% Dice on Prostate datasets, with our semi-supervised approach even exceeding fully supervised baselines. The method maintains parameter efficiency (1.8 M parameters) while delivering superior performance, validating its effectiveness for resource-constrained medical imaging applications. The source code is publicly available at https://github.com/jinggqu/Switch.
Renal fibrosis is a key pathological feature of chronic kidney disease (CKD), yet its noninvasive evaluation remains challenging. Radiomics provides a quantitative approach for extracting image-based biomarkers from ultrasound, and machine learning techniques may further enhance diagnostic accuracy for fibrosis severity assessment. This study aimed to compare and interpret machine learning models that utilize radiomics features extracted from ultrasound images for the evaluation of renal fibrosis severity in CKD patients. A total of 182 CKD patients (mean age, 40.91 ± 14.55 years; 101 men and 81 women) who underwent renal ultrasound and kidney biopsy were included. Radiomics features were extracted from ultrasound images to generate a radiomics signature. Five machine learning classifiers, including eXtreme Gradient Boosting, logistic regression, support vector machine, K-Nearest Neighbor, and random forest, were developed by combining the radiomics signature with key clinical variables identified through multiple algorithms. Model performance was assessed using receiver operating characteristic and precision-recall curves. Interpretability was achieved through SHapley Additive Explanations (SHAP). The logistic regression model achieved the most favorable diagnostic performance, with an area under the curve of 0.86 (95% confidence interval [CI]: 0.80-0.92) and an F1 score of 0.81 (95% CI: 0.78-0.84) in the primary cohort, and an area under the curve of 0.84 (95% CI: 0.71-0.98) and an F1 score of 0.82 (95% CI: 0.75-0.89) in cross-validation. SHAP analysis identified estimated glomerular filtration rate as the most influential feature, followed by the radiomics signature, age, and renal parenchyma thickness. The logistic regression model combining ultrasound-based radiomics and clinical information demonstrates strong potential for noninvasive renal fibrosis stratification in CKD, with SHAP facilitating transparent model interpretation.
Purpose Machine learning has been extensively applied in nephrology. This study aims to evaluate and compare the effectiveness of various information fusion strategies for differentiating mild from moderate-to-severe renal fibrosis in patients with chronic kidney disease (CKD).Methods This prospective study enrolled CKD patients who underwent renal ultrasound and biopsy at our institution from April 2019 to June 2022. Clinical laboratory indicators and ultrasound parameters were collected and analyzed. Two fusion strategies using machine learning techniques were developed: one at the feature level, which combines clinical and ultrasound features, and the other at the decision level, which integrates decisions from individual modality model outputs. Diagnostic performance was assessed using receiver operating characteristic (ROC) curves and the area under the ROC curve (AUC), along with sensitivity, specificity, and accuracy metrics.Results The multimodality fusion strategy demonstrated enhanced diagnostic performance compared to the single modality models. In the test cohort, the decision-level fusion model yielded optimal diagnostic performance, with an AUC of 0.92 (95% CI: 0.84-0.99), sensitivity of 0.80, specificity of 0.90, and accuracy of 0.84. This was followed by the feature-level fusion model (AUC: 0.86; 95% CI: 0.76-0.95), the clinical model (AUC: 0.82; 95% CI: 0.71-0.92), and the ultrasound model (AUC: 0.80; 95% CI: 0.68-0.93).Conclusion Information fusion strategies enhance diagnostic performance in differentiating mild from moderate-to-severe renal fibrosis in CKD patients. The decision-level fusion approach, which integrates outputs from individual models, proved most effective. These advancements support personalized CKD management by facilitating risk stratification and targeted treatment approaches.
Introduction The influence of ultrasound imaging plane on the diagnostic performance of artificial intelligence-based computer-aided diagnosis systems for thyroid nodules remains unclear. This study evaluated S-Detect in transverse and longitudinal ultrasound planes and compared four plane-based interpretation strategies. Methods This prospective cross-sectional study included 157 patients with 207 surgically confirmed thyroid nodules. Each nodule was independently analyzed by S-Detect in transverse (S-Detect_T) and longitudinal (S-Detect_L) planes. Four diagnostic strategies were evaluated: transverse-plane assessment, longitudinal-plane assessment, a serial strategy classifying a nodule as possibly malignant only when both planes indicated possible malignancy, and a parallel strategy classifying a nodule as possibly malignant when either plane indicated possible malignancy. Diagnostic performance, agreement, and interplane discordance were assessed. Results Of the 207 nodules, 140 were malignant and 67 were benign. S-Detect_L showed numerically higher sensitivity and specificity than S-Detect_T (93.6% vs 92.1% and 62.7% vs 59.7%, respectively), but neither difference was statistically significant. Transverse- and longitudinal-plane classifications agreed in 89.4% of nodules, whereas 10.6% showed interplane discordance. The parallel strategy achieved the highest sensitivity (97.9%) and negative predictive value (92.5%) but the lowest specificity (55.2%). The serial strategy achieved the highest specificity (67.2%) and the lowest false-positive rate (32.8%) but had lower sensitivity (87.9%). Compared with the serial strategy, the parallel strategy identified 14 additional malignant nodules while producing eight additional false-positive classifications. A sensitivity analysis restricted to one nodule per patient showed the same overall pattern. Conclusions Transverse- and longitudinal-plane S-Detect assessments showed similar diagnostic performance, although interplane discordance occurred in a subset of nodules. The parallel strategy favored sensitivity, whereas the serial strategy favored specificity and fewer false-positive classifications. No imaging plane or combination rule was consistently superior across all diagnostic measures.
PURPOSE:Machine learning has been extensively applied in nephrology. This study aims to evaluate and compare the effectiveness of various information fusion strategies for differentiating mild from moderate-to-severe renal fibrosis in patients with chronic kidney disease (CKD). METHODS:This prospective study enrolled CKD patients who underwent renal ultrasound and biopsy at our institution from April 2019 to June 2022. Clinical laboratory indicators and ultrasound parameters were collected and analyzed. Two fusion strategies using machine learning techniques were developed: one at the feature level, which combines clinical and ultrasound features, and the other at the decision level, which integrates decisions from individual modality model outputs. Diagnostic performance was assessed using receiver operating characteristic (ROC) curves and the area under the ROC curve (AUC), along with sensitivity, specificity, and accuracy metrics. RESULTS:The multimodality fusion strategy demonstrated enhanced diagnostic performance compared to the single modality models. In the test cohort, the decision-level fusion model yielded optimal diagnostic performance, with an AUC of 0.92 (95% CI: 0.84-0.99), sensitivity of 0.80, specificity of 0.90, and accuracy of 0.84. This was followed by the feature-level fusion model (AUC: 0.86; 95% CI: 0.76-0.95), the clinical model (AUC: 0.82; 95% CI: 0.71-0.92), and the ultrasound model (AUC: 0.80; 95% CI: 0.68-0.93). CONCLUSION:Information fusion strategies enhance diagnostic performance in differentiating mild from moderate-to-severe renal fibrosis in CKD patients. The decision-level fusion approach, which integrates outputs from individual models, proved most effective. These advancements support personalized CKD management by facilitating risk stratification and targeted treatment approaches.
AIMS:Stroke is a leading cause of death and disability worldwide and occurs primarily due to impaired cerebrovascular health. Aerobic exercise training (AET) has the potential to improve deconditioned cerebrovascular status in post-stroke patients, but its utility remains underexplored. This study assessed the effects of AET on cerebral artery morphology and hemodynamics in post-stroke patients using advanced ultrasonography techniques. METHODS:A randomized controlled trial (RCT) involving post-stroke patients randomly assigned to either a 12-week, high-intensity supervised cycling AET program (3 sessions per week, 30 min per session) or stretching (control) program was conducted. Duplex carotid ultrasound novel applications assessed pre- and post-intervention morphological features (carotid intima-media thickness (CIMT), arterial stiffness, and 3D features) and hemodynamics (resistive index (RI), pulsatility index (PI)). Transcranial color-coded Doppler (TCCD) assessed middle cerebral artery (MCA) hemodynamics. RESULTS:A total of 42 post-stroke patients, mean age (64.4 ± 7.6 years) were enrolled (cycling AET, n = 21 and stretching, n = 21). Mixed design ANOVA revealed significant time*exercise group interactions for CIMT, compliance, and distensibility (all p < 0.05). Cycling AET significantly improved CIMT (between group mean difference (MD) = -0.04 mm, p < 0.043), distensibility (MD = 0.003 (1/kPa)), compliance (MD = 0.11 mm/kPa, all p < 0.001). Modest within-group improvements were observed in internal carotid artery, RI (MD = -0.05, p < 0.001), and PI (MD = -0.2, p < 0.001). No significant changes observed in MCA hemodynamics. CONCLUSION:High-intensity cycling AET improved cerebral arteries' morphology and modestly enhanced extracranial hemodynamics in post-stroke patients, without affecting MCA hemodynamics. Further studies are recommended to explore long-term vascular benefits of AET. TRIAL REGISTRATION:ClinicalTrials.gov identifier: NCT05706168.
BACKGROUND:Multimodal large language models (LLMs) are increasingly being explored for medical image analysis, but their relative performance in thyroid ultrasound remains unclear. OBJECTIVE:This study aimed to compare six publicly available multimodal LLMs for grayscale ultrasound-based classification of thyroid nodules. METHODS:This prospective cross-sectional study included 178 patients with 239 thyroid nodules who underwent preoperative thyroid ultrasound followed by histopathological confirmation. Cropped grayscale ultrasound images of the maximal transverse and longitudinal views were analyzed by six publicly available multimodal LLMs: ChatGPT-5.4, Gemini 3.1 Pro, Claude Opus 4.6, Qwen3.6-Plus, Kimi K2.5, and ERNIE 5.0. All models were evaluated using the same image-input workflow and a standardized prompt, without fine-tuning or task-specific retraining. Agreement was assessed using Cohen's kappa, and diagnostic performance was evaluated using receiver operating characteristic (ROC) analysis. Radiologist benchmarks were included for comparison. RESULTS:All six LLMs significantly distinguished benign from malignant nodules (all P ≤ 0.001). Gemini 3.1 Pro achieved the best overall performance, with a kappa value of 0.580 and an area under the ROC curve (AUC) of 77.1% (95% CI, 71.5%-82.7%). ChatGPT-5.4 and Qwen3.6-Plus each yielded an AUC of 73.5%, and Kimi K2.5 achieved an AUC of 71.3%. Claude Opus 4.6 and ERNIE 5.0 showed lower overall performance, with AUCs of 65.7% and 59.9%, respectively. The senior radiologist achieved higher diagnostic performance than all six LLMs. CONCLUSION:Publicly available multimodal LLMs showed measurable but heterogeneous performance in grayscale ultrasound-based thyroid nodule classification. Gemini 3.1 Pro demonstrated the best overall results, but none of the models matched senior radiologist-level performance.
Objective:Accurate assessment of renal fibrosis is critical for managing chronic kidney disease (CKD). This study aimed to develop a multi-modality ultrasound-based model that integrates radiomics signatures derived from grayscale ultrasound and color Doppler ultrasound, along with shear wave elastography (SWE) measurements, to assess the severity of renal fibrosis in CKD patients. Methods:A total of 125 CKD patients were enrolled and classified into mild and moderate-to-severe fibrosis groups based on renal biopsy. Radiomics features were extracted from grayscale and color Doppler ultrasound images, and key features were identified using machine learning algorithms to construct radiomics signatures for each modality. SWE measurements were used to assess renal stiffness. A multi-modality ultrasound model was constructed using logistic regression, combining the dual-modality radiomics signatures with SWE data. Model performance was evaluated using receiver operating characteristic (ROC) curves, five-fold cross-validation, calibration, and decision curve analysis (DCA). Results:Individual modalities for SWE, grayscale ultrasound radiomics, and color Doppler ultrasound radiomics showed area under the ROC curves (AUCs) of 0.74 (95% CI: 0.65-0.82), 0.79 (95% CI: 0.71-0.87), and 0.70 (95% CI: 0.60-0.79), respectively. The multi-modality model, integrating all three modalities, achieved an AUC of 0.88 (95% CI: 0.82-0.94), sensitivity of 0.82 (95% CI: 0.70-0.91), specificity of 0.84 (95% CI: 0.73-0.92), and accuracy of 0.83 (95% CI: 0.75-0.89). In cross-validation, the model showed robust generalizability (AUC: 0.89, 95% CI: 0.74-1.00). The calibration curve showed excellent agreement between predicted and observed outcomes, and the DCA curve confirmed the clinical utility of the model. A nomogram based on the multi-modality model was developed for individualized risk assessment of moderate-to-severe renal fibrosis. Conclusion:The multi-modality ultrasound model enhances non-invasive renal fibrosis assessment in CKD patients. By combining dual-modality radiomics with SWE measurements, this model offers a promising tool for personalized clinical decision-making and better management of CKD progression.
ObjectivesThis study aimed to develop and compare radiomics signatures derived from different renal regions on ultrasound images to assess fibrosis severity in chronic kidney disease (CKD) patients.MethodsA total of 146 CKD patients who underwent renal ultrasound and biopsy were enrolled. Radiomics features were extracted from the whole kidney, parenchyma, and mid-portion to generate region-specific signatures: radscore_whole, radscore_parenchyma, and radscore_mid-portion. Diagnostic performance in distinguishing mild from moderate-to-severe fibrosis was evaluated using receiver operating characteristic (ROC) curve analysis. Performance improvements were assessed via net reclassification improvement (NRI) and integrated discrimination improvement (IDI).ResultsThe radscore_mid-portion displayed the highest discriminatory accuracy, yielding an area under the ROC curve (AUC) of 0.74 (95% CI: 0.65-0.82), which exceeded both the radscore_whole (AUC = 0.61, 95% CI: 0.51-0.70; P = 0.035) and radscore_parenchyma (AUC = 0.66, 95% CI: 0.56-0.74; P = 0.181). Reclassification analysis confirmed the added diagnostic value of the mid-portion signature, with significant improvements compared with both the whole-kidney (NRI = 36.35%; IDI = 10.57%) and parenchyma signatures (NRI = 42.23%; IDI = 10.42%).ConclusionsRadiomics signatures from different renal regions offer varying diagnostic utility. The mid-portion-based signature demonstrated improved performance and added value in identifying moderate-to-severe renal fibrosis in CKD patients.
Accurate classification of lymphadenopathy is essential for determining the pathological nature of lymph nodes (LNs), which plays a crucial role in treatment selection. The biopsy method is invasive and carries the risk of sampling failure, while the utilization of non-invasive approaches such as ultrasound can minimize the probability of iatrogenic injury and infection. With the advancement of artificial intelligence (AI) and machine learning, the diagnostic efficiency of LNs is further enhanced. This study evaluates the performance of ultrasound-based AI applications in the classification of benign and malignant LNs. The literature research was conducted using the PubMed, EMBASE, and Cochrane Library databases as of June 2024. The quality of the included studies was evaluated using the QUADAS-2 tool. The pooled sensitivity, specificity, and diagnostic odds ratio (DOR) were calculated to assess the diagnostic efficacy of ultrasound-based AI in classifying benign and malignant LNs. Subgroup analyses were also conducted to identify potential sources of heterogeneity. A total of 1,355 studies were identified and reviewed. Among these studies, 19 studies met the inclusion criteria, and 2,354 cases were included in the analysis. The pooled sensitivity, specificity, and DOR of ultrasound-based machine learning in classifying benign and malignant LNs were 0.836 (95
Background/Objectives: Advances in large language models like ChatGPT-4o have extended their use to medical image analysis. Accurate assessment of thyroid nodule ultrasound features using ACR TI-RADS is crucial for clinical practice. This study aims to evaluate ChatGPT-4o’s intra-observer consistency and its agreement with an expert in analyzing these features from ultrasound image assessments based on ACR TI-RADS. Methods: This cross-sectional study used ultrasound images from 100 thyroid nodules collected prospectively between May 2019 and August 2021. Ultrasound images were analyzed by ChatGPT-4o, following ACR TI-RADS guidelines, to assess features of thyroid nodule including composition, echogenicity, shape, margin, and echogenic foci. The analysis was repeated after one week to evaluate intra-observer reliability. The ultrasound images were also analyzed by another ultrasound expert for the evaluation of inter-observer reliability. Agreement was measured using Cohen’s Kappa coefficient, and concordance rates were calculated based on alignment with the expert’s reference classifications. Results: Intra-observer agreement for ChatGPT-4o was moderate for composition (Kappa = 0.449) and echogenic foci (Kappa = 0.404), with substantial agreement for echogenicity (Kappa = 0.795). Agreement was notably low for shape (Kappa = −0.051) and margin (Kappa = 0.154). Inter-observer agreement between ChatGPT-4o and the expert was generally low, with Kappa values ranging from −0.006 to 0.238, the highest being for echogenic foci. Overall concordance rates between ChatGPT-4o and expert evaluations ranged from 46.6% to 48.2%, with the highest for shape (65%) and the lowest for echogenicity (29%). Conclusions: ChatGPT-4o showed favorable consistency in assessing some thyroid nodule features in intra-observer analysis, but notable variability in others. Inter-observer comparisons with expert evaluations revealed generally low agreement across all features, despite acceptable concordance for certain imaging characteristics. While promising for specific ultrasound features, ChatGPT-4o’s consistency and accuracy still vary significantly compared to expert assessments.
This study aims to develop a machine learning model that accurately diagnoses microvascular invasion (MVI) in hepatocellular carcinoma by using radiomic features from MVI-positive regions of interest (ROIs). Unlike previous studies, which do not account for the location and distribution of MVI, this research focuses on correlating preoperative imaging with postoperative pathological MVI. This study involves obtaining ex vivo 3D ultrasound images of 36 hepatic specimens from nine rabbits. These images are fused with whole-slide images to localize MVI regions precisely. The identified MVI regions are segmented into MVI-positive ROIs, with a 1:3 ratio of positive to negative ROIs. Radiomic features are extracted from each ROI, and 30 features highly associated with MVI are selected for model development. The performance of several machine learning models is evaluated using metrics such as sensitivity, specificity, accuracy, the area under the curve (AUC), and F1 score. The GBDT model achieves the best results, with an AUC of 0.91, an F1 score of 0.85, a sensitivity of 0.76, a specificity of 0.92, and an accuracy of 0.86. The high diagnostic accuracy of these models highlights the potential for future clinical application in the precise diagnosis of MVI using radiomic features from MVI-positive ROIs.
Automatic lymph node segmentation is the cornerstone for advances in computer vision tasks for early detection and staging of cancer. Traditional segmentation methods are constrained by manual delineation and variability in operator proficiency, limiting their ability to achieve high accuracy. The introduction of deep learning technologies offers new possibilities for improving the accuracy of lymph node image analysis. This study evaluates the application of deep learning in lymph node segmentation and discusses the methodologies of various deep learning architectures such as convolutional neural networks, encoder-decoder networks, and transformers in analyzing medical imaging data across different modalities. Despite the advancements, it still confronts challenges like the shape diversity of lymph nodes, the scarcity of accurately labeled datasets, and the inadequate development of methods that are robust and generalizable across different imaging modalities. To the best of our knowledge, this is the first study that provides a comprehensive overview of the application of deep learning techniques in lymph node segmentation task. Furthermore, this study also explores potential future research directions, including multimodal fusion techniques, transfer learning, and the use of large-scale pre-trained models to overcome current limitations while enhancing cancer diagnosis and treatment planning strategies.
Background/Objectives Recent advancements in large language models, such as ChatGPT-4o, have created new opportunities for analyzing complex multi-modal data, including medical images. This study aims to assess the potential of ChatGPT-4o in distinguishing between benign and malignant thyroid nodules via multi-modality ultrasound imaging: grayscale ultrasound, color Doppler ultrasound (CDUS), and shear wave elastography (SWE). Materials and Methods Patients who underwent thyroid nodule ultrasound examinations and had confirmed pathological diagnoses were included. ChatGPT-4o analyzed the multi-modality ultrasound data using two approaches: (1.) a dual-modality strategy which employed grayscale ultrasound and CDUS, and (2.) a triple-modality strategy which incorporated grayscale ultrasound, CDUS, and SWE. The diagnostic performance was compared against pathological findings utilizing receiver operating characteristic (ROC) curve analysis, while consistency was evaluated through Cohen’s Kappa analysis. Results A total of 106 thyroid nodules were evaluated; 65.1% were benign and 34.9% malignant. In the dual-modality approach, ChatGPT-4o achieved an area under the ROC curve (AUC) of 66.3%, moderate agreement with pathology results (Kappa = 0.298), a sensitivity of 70.3%, a specificity of 62.3%, and an accuracy of 65.1%. Conversely, the triple-modality approach exhibited higher specificity at 97.1% but lower sensitivity at 18.9%, with an accuracy of 69.8% and a reduced overall agreement (Kappa = 0.194), resulting in an AUC of 58.0%. Conclusions ChatGPT-4o exhibits potential, to some extent, in classifying thyroid nodules using multi-modality ultrasound imaging. However, the dual-modality approach unexpectedly outperforms the triple-modality approach. This indicates that ChatGPT-4o might encounter challenges in integrating and prioritizing different data modalities, particularly when conflicting information is present, which could impact diagnostic effectiveness.
PURPOSE:Recent advances in multimodal large language models (LLMs) have demonstrated promising potential for medical image analysis, yet their diagnostic capability in thyroid ultrasound remains unverified. This study explored the feasibility of ChatGPT-5, the latest multimodal LLM, for thyroid nodule classification and contextualized its diagnostic performance against S-Detect, an FDA-approved commercial computer-aided diagnosis system. METHODS:In this prospective study, 141 patients with 186 nodules who underwent preoperative ultrasound and subsequent surgery were enrolled. For S-Detect, the largest transverse grayscale ultrasound image of each nodule was analyzed with automated contouring for binary classification. For ChatGPT-5, cropped transverse and longitudinal nodule ultrasound images were analyzed using a standardized diagnostic prompt for binary classification. Agreement with histopathology was assessed using Kappa statistics; sensitivity, specificity, accuracy, and area under the receiver operating characteristic curve (AUC) were calculated. RESULTS:Both systems showed statistically significant ability to distinguish benign from malignant nodules (P < 0.05). Agreement with histopathology was fair for ChatGPT-5 (Kappa = 0.224) and moderate for S-Detect (Kappa = 0.579). ChatGPT-5 demonstrated sensitivity 50.8 %, specificity 75.8 %, and accuracy 59.1 %, whereas S-Detect achieved higher sensitivity (91.9 %) and accuracy (82.3 %) but lower specificity (62.9 %). The AUC for S-Detect (77.4 %) was significantly greater than that for ChatGPT-5 (63.3 %, P < 0.001). CONCLUSIONS:ChatGPT-5 demonstrated feasibility for thyroid nodule classification but showed lower diagnostic performance than the licensed, pre-trained S-Detect system and is not yet adequate for medical imaging applications.
Objective: Renal fibrosis is a final common pathological hallmark in the progression of chronic kidney disease (CKD). Non-invasive evaluation of renal fibrosis by mapping renal stiffness obtained by shear wave elastography (SWE) may facilitate the clinical therapeutic regimen for CKD patients. Methods: A cohort of 162 patients diagnosed with CKD, who underwent renal biopsy, was prospectively and consecutively recruited between April 2019 and December 2021. The assessment of renal cortex stiffness was performed using SWE imaging. The patients were classified into different groups based on pathological renal fibrosis (mild group: n = 74; moderate-to-severe group: n = 88). Binary logistic regression model and generalized additive model were conducted to investigate the association of renal elasticity with renal fibrosis. Results: Compared with the mildly impaired group, the moderate-to-severe group showed a significant decline in renal elasticity (P < .001). In the fully adjusted model, each 10 kPa drop in renal elasticity was associated with a 3.5-fold increment in the risk of moderate-to-severe renal fibrosis (fully adjusted odds ratio, 4.54; 95% CI, 2.41-8.57). Particularly, participants in the lowest elasticity group (<= 29.92 kPa) had a 20-fold increased chance of moderate-to-severe renal fibrosis than those in the group with highest elasticity (>= 37.93 kPa). An inverse linear association was observed between renal elasticity increment and moderate-to-severe renal fibrosis risk. Conclusion: There is a negative linear association between increased renal elasticity and moderate-to-severe renal fibrosis risk among CKD patients. Patients with diminished renal stiffness have a higher risk of moderate-to-severe renal fibrosis. Advances in knowledge: CKD patients with reduced renal stiffness have a higher likelihood of moderate-to-severe renal fibrosis.
Debate continues regarding the potential of the ultrasonic renal length to serve as an indicator for evaluating the advancement of renal fibrosis in chronic kidney disease (CKD). This study investigates the independent association between renal length and renal fibrosis in non-diabetic CKD patients and assesses its diagnostic performance. From April 2019 to December 2021, 144 non-diabetic patients diagnosed with CKD who underwent a renal ultrasound examination and kidney biopsy were prospectively enrolled. Patients were categorized into the mild fibrosis group (n = 70) and the moderate-severe group (n = 74) based on the extent of fibrotic involvement. Ultrasonic renal length was measured from pole-to-pole in the coronal plane. A receiver operating characteristic (ROC) curve, multivariable logistic regression analysis, and a generalized additive model were performed. A negative linear correlation was found between renal length and moderate-severe renal fibrosis risk. Each centimeter increase in renal length decreased the odds of moderate-severe fibrosis by 38
Background Non-invasive renal fibrosis assessment is critical for tailoring personalized decision-making and managing follow-up in patients with chronic kidney disease (CKD). We aimed to exploit machine learning algorithms using clinical and elastosonographic features to distinguish moderate-severe fibrosis from mild fibrosis among CKD patients. Methods A total of 162 patients with CKD who underwent shear wave elastography examinations and renal biopsies at our institution were prospectively enrolled. Four classifiers using machine learning algorithms, including eXtreme Gradient Boosting (XGBoost), Support Vector Machine (SVM), Light Gradient Boosting Machine (LightGBM), and K-Nearest Neighbor (KNN), which integrated elastosonographic features and clinical characteristics, were established to differentiate moderate-severe renal fibrosis from mild forms. The area under the receiver operating characteristic curve (AUC) and average precision were employed to compare the performance of constructed models, and the SHapley Additive exPlanations (SHAP) strategy was used to visualize and interpret the model output. Results The XGBoost model outperformed the other developed machine learning models, demonstrating optimal diagnostic performance in both the primary (AUC = 0.97, 95% confidence level (CI) 0.94–0.99; average precision = 0.97, 95% CI 0.97–0.98) and five-fold cross-validation (AUC = 0.85, 95% CI 0.73–0.98; average precision = 0.90, 95% CI 0.86–0.93) datasets. The SHAP approach provided visual interpretation for XGBoost, highlighting the features’ impact on the diagnostic process, wherein the estimated glomerular filtration rate provided the largest contribution to the model output, followed by the elastic modulus, then renal length, renal resistive index, and hypertension. Conclusion This study proposed an XGBoost model for distinguishing moderate-severe renal fibrosis from mild forms in CKD patients, which could be used to assist clinicians in decision-making and follow-up strategies. Moreover, the SHAP algorithm makes it feasible to visualize and interpret the feature processing and diagnostic processes of the model output. Graphical Abstract
Large language models (LLMs) are pivotal in artificial intelligence, demonstrating advanced capabilities in natural language understanding and multimodal interactions, with significant potential in medical applications. This study explores the feasibility and efficacy of LLMs, specifically ChatGPT-4o and Claude 3-Opus, in classifying thyroid nodules using ultrasound images. This study included 112 patients with a total of 116 thyroid nodules, comprising 75 benign and 41 malignant cases. Ultrasound images of these nodules were analyzed using ChatGPT-4o and Claude 3-Opus to diagnose the benign or malignant nature of the nodules. An independent evaluation by a junior radiologist was also conducted. Diagnostic performance was assessed using Cohen’s Kappa and receiver operating characteristic (ROC) curve analysis, referencing pathological diagnoses. ChatGPT-4o demonstrated poor agreement with pathological results (Kappa = 0.116), while Claude 3-Opus showed even lower agreement (Kappa = 0.034). The junior radiologist exhibited moderate agreement (Kappa = 0.450). ChatGPT-4o achieved an area under the ROC curve (AUC) of 57.0
The early and accurate stratification of intracranial cerebral artery stenosis (ICAS) is critical to inform treatment management and enhance the prognostic outcomes in patients with cerebrovascular disease (CVD). Digital subtraction angiography (DSA) is an invasive and expensive procedure but is the gold standard for the diagnosis of ICAS. Over recent years, transcranial color-coded Doppler ultrasound (TCCD) has been suggested to be a useful imaging method for accurately diagnosing ICAS. However, the diagnostic accuracy of TCCD in stratifying ICASs among patients with CVD remains unclear. Therefore, this systematic review and meta-analysis aimed at evaluating the diagnostic accuracy of TCCD in the stratification of intracranial steno-occlusions among CVD patients. A total of six databases-Embase, CINAHL, Medline, PubMed, Google Scholar, and Web of Science (core collection)-were searched for studies that assessed the diagnostic accuracy of TCCD in stratifying ICASs. The meta-analysis was performed using Meta-DiSc 1.4. The Quality Assessment of Diagnostic Accuracy Studies tool version 2 (QUADAS-2) assessed the risk of bias. Eighteen studies met all of the eligibility criteria. TCCD exhibited a high pooled diagnostic accuracy in stratifying intracranial steno-occlusions in patients presenting with CVD when compared to DSA as a reference standard (sensitivity = 90%; specificity = 87%; AUC = 97%). Additionally, the ultrasound parameters peak systolic velocity (PSV) and mean flow velocity (MFV) yielded a comparable diagnostic accuracy of "AUC = 0.96". In conclusion, TCCD could be a noble, safe, and accurate alternative imaging technique to DSA that can provide useful diagnostic information in stratifying intracranial steno-occlusions in patients presenting with CVD. TCCD should be considered in clinical cases where access to DSA is limited.