
Background Monkeypox (Mpox) has emerged as a growing public health concern, requiring effective tools for early patient risk assessment and clinical decision-making. This study aimed to develop an artificial intelligence (AI)-based machine learning framework using a multilayer perceptron (MLP) neural network to analyze a Mpox dataset collected from 11 non-endemic countries, identify the most significant clinical symptoms, investigate the relationships among travel history, hospitalization, isolation, and infection status, and predict the hospitalization and isolation status of Mpox patients. Methods A total of 365 non-severe Mpox cases from 11 non-endemic countries were included in the analysis provided in publicly available World Health Organization (WHO) reports from May 6 to May 29, 2022. Patients were classified into two groups: Hospitalized and isolated patients. Clinical data were converted to binary dummy variables and demographic and travel-history data were added as features. The data were analyzed using cross-tabulation and descriptive visualization for symptom patterns and relationships between clinical and demographic variables. A supervised AI-based MLP neural network was then developed to predict the status of hospitalization and isolation, and the performance of the model was assessed by accuracy, recall, false omission rate (FOR) and F1 score. Results The symptoms most commonly reported in the analysis included genital ulcers, fever, and lesions. The eleven countries included in the study reported varying symptom patterns, with Argentina showing the highest symptom diversity, with up to four symptoms recorded per case.. The overall training accuracy of the proposed MLP neural network for the prediction of hospitalization was 94.7% and the overall testing accuracy was 90.7%. The model showed good predictive performance for each patient status classification task with an accuracy of 94.2% at training phase and 90.1% at test phase for isolation prediction. Conclusion The proposed AI-based multilayer perceptron (MLP) neural network provides an effective approach for predicting the hospitalization and isolation status of Mpox patients using clinical and demographic variables. Symptom distributions varied across countries, indicating regional heterogeneity in disease presentation.
Background Cardiac arrest (CA) remains a critical medical emergency characterized by high mortality rates despite the advancements in resuscitation and intensive care. The heterogeneous and evolving clinical course of patients with CA hinders early and accurate prognostication. We aimed to construct and externally validate a predictive model for 28-day intensive care unit (ICU) mortality in patients with CA by leveraging time-weighted clinical data and traditional and machine-learning approaches including extreme gradient boosting (XGBoost), gradient boosting with component-wise linear models (Glmboost), random forest (RF), generalized linear models with elastic net regularization (Glmnet), and multivariate Cox regression. Methods We retrospectively analyzed adult patients with CA from three large-scale ICU databases: eICU Collaborative Research Database (training cohort, n = 1,077), Medical Information Mart for Intensive Care IV (MIMIC-IV v3.1), and MIMIC-III Clinical Database CareVue subset v1.4 (combined validation cohort, n = 368). Time-weighted averages (TWA) of clinical parameters—including the Glasgow Coma Scale (GCS), vital signs, and arterial blood gas indices—were computed over the first three ICU days. Variable selection was performed using the least absolute shrinkage and selection operator regression, followed by multivariable Cox regression analysis. Five prognostic models—Glmboost, XGBoost, RF, Glmnet, and Cox regression—were developed and evaluated using C-index, time-dependent area under the curve (AUC), Graf score, and decision curve analysis. Model interpretability and time-varying effects were examined via SHapley Additive exPlanations (SHAP) and Piece-wise exponential additive mixed model (PAMM). Results Nine independent predictors of ICU mortality were identified, with TWA Glasgow Coma Scale (TWA-GCS) (hazard ratio: 0.75, 95% CI: 0.71–0.78, P < 0.01) and duration of mechanical ventilation (hazard ratio: 0.72, 95% CI: 0.69–0.75, P < 0.01) showing the strongest associations. In external validation, the Glmboost model showed the lowest Graf score. SHAP analyses underscored TWA-GCS and duration of mechanical ventilation as the most impactful variables. PAMM analysis revealed an increasingly protective temporal effect of higher TWA-GCS values. Subgroup analysis showed 28-day ICU mortality rates of 76%, 57%, and 32% in patients with low (< 4), moderate (4–10), and high (> 10) TWA-GCS scores, respectively (P < 0.01). Conclusion Incorporating time-weighted clinical data significantly improves ICU mortality prediction in patients with CA. Glmboost showed the best calibration/error-related performance based on Graf score while demonstrating comparable discrimination, with TWA-GCS identified as a key prognostic factor.
Background Pancreatic cancer continues to be one of the most challenging malignancies to diagnose and treat, and it is often detected at an advanced stage owing to asymptomatic early phases and limitations in current diagnostic tools, highlighting the need for improved diagnostic methodologies. Methods We aimed to evaluate the potential for integrating information from thermal liquid biopsy (TLB) thermograms with clinical biomarkers to improve the diagnosis and prognosis of pancreatic cancer. Serum samples from 381 Danish patients with pancreatic cancer and 325 patients referred with non-organ specific signs and symptoms of cancer, but without confirmed cancer after clinical and computed tomography evaluation, were analyzed. TLB thermograms provided information on the partial excess heat capacity of the serum as a function of temperature. Three classification models were constructed using machine learning algorithms based on variable selection with penalization, applying cross-validation and resampling techniques: iClin Model (age, Eastern Cooperative Oncology Group Performance Status, carbohydrate antigen 19.9, and C-reactive protein), iTLB Model (discordant pairs of temperatures from thermograms), and iTLB + iClin Model (discordant pairs of temperatures and clinical data). Results The iClin Model achieved high diagnostic performance, with a validation area under the curve (AUC) of 0.95 ± 0.01 for differentiating symptomatic controls from patients with pancreatic ductal adenocarcinoma (PDAC), but showed no significant association with overall survival (OS). The iTLB Model showed limited diagnostic performance, with a validation AUC of 0.59 ± 0.03, but was associated with OS in patients with PDAC. The combined iTLB + iClin Model preserved high diagnostic performance, with a validation AUC of 0.95 ± 0.02, sensitivity of 96.20%, specificity of 91.20%, positive predictive value of 90.48%, and negative predictive value of 96.59%. For early-stage PDAC, the iTLB + iClin Model achieved an AUC of 0.93 (95% CI: 0.89–0.98), compared with 0.96 (95% CI: 0.95–0.98) for stages III–IV. Conclusion This study demonstrates the potential of combining specific clinical biomarker information with advanced techniques such as TLB, to improve the accuracy of pancreatic cancer diagnosis and prognosis.
Critical care medicine is at a pivotal stage of transformation from symptom-based approaches toward systems science. Building on a first-principles understanding of the mechanisms underlying critical illness, we propose the entropic critical illness theory (ECIT). Grounded in the second law of thermodynamics, ECIT posits that critical illness arises when the body’s capacity to regulate entropy production is impaired, leading to systemic disorder. This process is manifested by the concurrent escalation of host response entropy and hemodynamic entropy, resulting in disruptions in blood flow and oxygen delivery, ischemia, and hypoxia at the level of the critical unit, which is defined as the terminal microcirculatory–mitochondrial functional unit, ultimately culminating in multi-organ dysfunction. This study delineates the theoretical foundations of ECIT and outlines an entropy-based critical care framework that may inform future multimodal entropy-informed monitoring, risk stratification, and AI-assisted precision intervention in critical care medicine.
Surgery is a real-time, embodied and safety-critical form of care in which decisions, anatomy, physical action and patient outcomes are tightly coupled. Artificial intelligence (AI) has improved the recognition of instruments, anatomy, workflow, technical performance and multimodal surgical context, but these capabilities have limited clinical value if they remain confined to retrospective benchmarks, narrow procedural settings and isolated perception tasks. The central challenge is to determine whether surgical models can support safer operative care under uncertainty, across procedures, institutions, devices and patient contexts, without obscuring surgical responsibility. This is not a scaling problem alone, but a problem of representation, evidence and accountability. Generalist surgical foundation models may provide a shared basis for clinically useful surgical understanding, but surgery demands a more constrained formulation than general biomedical AI, one that is anchored in operative accountability, real-time decision-making and patient-level consequences. Clinical-grade surgical intelligence should accordingly be defined as a qualification standard, not a capability threshold, for models intended to assist surgeons. First, surgery should be treated as a generalization problem across sensory-action regimes, longitudinal perioperative episodes and heterogeneous clinical environments. Second, model development should progress from machine-readable perception to temporally anchored procedural understanding, evidence-grounded reasoning, operating-room multimodal intelligence and surgeon-supervised action support. Third, clinical-grade claims should be earned through evidence that extends beyond benchmark accuracy to external and cross-regime validity, real-time reliability under intraoperative uncertainty, human-facing utility, patient-level relevance and lifecycle governance. We therefore reframe generalist surgical foundation models not as autonomous substitutes or increasingly capable pretraining assets, but as components of a clinically accountable intelligence layer for safer operative care. The clinical value of these models will depend on whether broader representations, stronger evidence and explicit accountability mature together.
The deep integration of artificial intelligence (AI) in the healthcare sector is revolutionizing medical services, clinical decision-making, and academic research and education in medical institutions. However, its autonomous evolutionary characteristics based on continuous learning and iterative updates pose fundamental challenges to traditional, static regulatory frameworks of medical devices. The “Expert consensus on the application and governance of artificial intelligence in medical institutions (2026),” developed by over 40 leading Chinese medical and scientific research institutions, involved experts from fields such as medicine, hospital management, medical informatics, health policy, law, and medical ethics. Drawing on the latest Chinese and international regulations and best practices, the expert panel systematically compiled a comprehensive governance guide for AI applications across the entire lifecycle of medical institutions, focusing on six thematic pillars: access evaluation, clinical application, patient rights protection, data governance, risk management, and competency enhancement. For each domain, the consensus details critical factors, including tiered access, multidisciplinary review, real-world validation, cross-validation of competing algorithms, delineation of human-machine collaboration liabilities, tiered informed consent, algorithmic traceability and explainability, dynamic risk monitoring, circuit-breaker mechanisms, and progressive AI competency training systems, that medical institutions should consider during AI implementation and management. This consensus aims to establish a compliance baseline and development direction for AI applications in medical institutions that is characterized by safety, efficacy, fairness, and interpretability, thereby ensuring that technological evolution remains within a legally and ethically sound framework. Ultimately, it serves to promote equitable access to high-quality medical resources and achieve substantial improvements in national health standards.
Background: Axial spondyloarthritis (axSpA) can potentially progress to ankylosing spondylitis, and diagnostic delay may lead to irreversible structural damage. Although sacroiliac joint magnetic resonance imaging (MRI) can detect bone marrow edema (BME) for early diagnosis, the complex etiologies and reliance on expert interpretation impede efficient identification in chronic low back pain populations. Therefore, we developed an artificial intelligence (AI) system that integrates MRI analysis and automates report generation to facilitate the detection of active sacroiliitis and BME quantification. Methods: This retrospective study analyzed 691 patients (540 with axSpA and 151 with non-spondyloarthritis) from the Chinese People’s Liberation Army General Hospital (2011–2023). Data were split into training and testing cohorts (4:1 ratio), with five-fold cross-validation applied to the training set. The system comprises four modules: (1) image preprocessing (intensity normalization), (2) coarse-to-fine 3D U-Net-based quadrant segmentation, (3) ResNet18-driven edema recognition (depth/intensity classification), and (4) diagnostic report generation using the QWEN large language model. Results: Among 691 patients (335 active sacroiliitis-positive and 356 negative), the AI system achieved Spondyloarthritis Research Consortium of Canada (SPARCC) scores of (15.12±10.73) vs. (0.88±1.13) in positive versus negative groups, respectively. The quadrant segmentation module achieved DICE similarity coefficients above 0.7 on the training, validation, and testing datasets. The edema inflammation, depth, and intensity classifiers exhibited good performance on the testing dataset, with balanced accuracies of 77.54%, 76.26%, and 80.91%, and area under the curve values of 0.85, 0.87, and 0.92, respectively. The intraclass correlation coefficient between SPARCC scores by our system and those by rheumatologists was 0.84 on the testing dataset. At the patient-level, our system achieved 90.81% sensitivity for diagnosis of active sacroiliitis on the testing dataset. Conclusion: The developed automated system enhances axSpA diagnostic efficiency by automating the identification of active sacroiliitis and enabling quantitative assessment of BME.
Background: Fractal dimension (D-f) quantifies vascular network complexity and spatial filling density, providing a non-invasive biomarker for microvascular pathology assessment. We aimed to establish an artificial intelligence (AI)-powered framework for automated D-f quantification and investigate its clinical implications in a general population of the Beijing Eye Study. Methods: This retrospective study utilized data from the Beijing Eye Study 2011. Fundus images meeting quality criteria were processed with an AI algorithm to segment retinal vasculature and calculate D-f. Multiple linear and logistic regression models were used to assess the associations of retinal vascular D-f with systemic and ocular parameters. Results: AI-based quantification of retinal vascular D-f was successfully performed in 3,298 participants (95.1% of 3,468 participants), demonstrating high model performance: segmentation accuracy (0.9660), sensitivity (0.8879), specificity (0.9743), and intersection-over-union (IoU) (0.7110). The mean D-f was 1.51 +/- 0.09 (median: 1.53; interquartile range: 1.50-1.55). In multiple linear regression analysis, reduced D-f was significantly associated with older age (standardized regression coefficient (s beta) = -2.346; P < 0.001), higher systolic blood pressure (s beta = -0.341; P < 0.001), smaller hip circumference (s beta = 0.518; P = 0.014), lower equivalent diopter (s beta = 3.589; P < 0.001), thinner retinal nerve fiber layer thickness (s beta = 0.390; P = 0.002), smaller arteriovenous ratio (s beta = 524.590; P < 0.001), and thinner subfoveal choroidal thickness (s beta = 0.071; P < 0.001). In the logistic regression model, the risk prevalence of hypertension increased with the decrease in D-f (OR = 0.853, 95% CI: 0.651-0.801). Conclusion: AI-powered D-f analysis may be used as a novel quantitative platform for characterizing retinal microvascular morphological alterations.
Objective Neoadjuvant chemoradiotherapy (nCRT) has become the standard preoperative treatment for patients with locally advanced rectal cancer (LARC). Accurately predicting patient prognosis is helpful for formulating individualized treatment strategies. This study aimed to establish and validate a clinical radiomics model based on magnetic resonance imaging and clinical characteristics to predict the prognosis of patients with LARC who underwent radical surgery after nCRT. Methods This retrospective study included 234 patients with LARC who underwent radical surgery after nCRT at the Affiliated Hospital of Qingdao University between December 2019 and October 2023. The patients were randomly divided into training (164 patients) and testing (70 patients) sets. Imaging and clinical data were collected and analyzed. 1172 radiomic features were extracted from preoperative pelvic magnetic resonance T2-weighted imaging (MRI T2WI). After screening, the radiomic feature score (Rad-Score) of the tumor and mesorectal regions was calculated, and three radiomics models were developed. A clinical model was constructed using the identified clinical predictors. The Rad-Score and clinical predictors were integrated using the random forest algorithm to establish a clinical radiomics model. The accuracy of the clinical radiomic model was evaluated using the area under the receiver operating characteristic (AUC) curve, calibration, and decision curves analysis. Results A total of 8 tumor radiomic features and 12 mesenteric radiomic features that were highly correlated with patient prognosis from pelvic MRI T2WI of patients after nCRT. We then established a hybrid radiomics model using all the radiomic features. The AUCs of the hybrid radiomics model for the training and testing sets are 0.832 (95% CI: 0.762–0.902) and 0.764 (95% CI: 0.646–0.881), respectively. The clinical radiomics model combining 20 radiomic features and 4 clinical predictors achieved an AUC of 0.874 (95% CI: 0.811–0.936) in the training set and 0.813 (95% CI: 0.710–0.916) in the testing set. Conclusion The clinical radiomics model based on MRI developed in this study has a high predictive performance for the prognosis of patients with LARC who underwent radical surgery after nCRT, making helping clinicians to perform nCRT in rectal cancer patients.
Artificial intelligence (AI) in medicine is advancing steadily toward real clinical practice, not only through improved predictive performance but also through more decision-relevant modeling, evaluation, and interaction. The studies highlighted here illustrate three complementary directions. First, contemporary imaging models—exemplified by three-dimensional vision transformer-based analysis of preoperative chest computed tomography (CT)—are being used to infer clinically consequential phenotypes, advancing from image recognition toward decision support in high-stakes settings such as surgical planning. Second, image-to-biomarker pipelines such as automated quantification of retinal vascular fractal dimension demonstrate how medical images can be transformed into reproducible quantitative markers suitable for population-level analysis and risk stratification. Third, large language models (LLMs) are increasingly being evaluated and positioned as clinical communication and interpretation components, most notably structured assessments in telepharmacy and bilingual patient education move beyond fluency to safety, actionability, empathy, and readability, whereas emerging perspectives consider LLMs as interpretive interfaces for complex and temporally evolving health data, including wearable sensing. At the same time, these advances also expose persistent bottlenecks that limit real-world deployment, particularly the continued dominance of single-task and static formulations, fragmented systems driven by task-specific fine-tuning, limited reasoning over disease trajectories and evolving clinical contexts, and evaluation practices that remain insufficiently coupled to clinical workflows and downstream consequences. We argue that the next stage of medical AI should shift from accuracy-centered prediction toward decision-oriented, practice-ready systems that are robust across settings, clinically aligned in evaluation, and deployable at scale in routine care.
Malignant tumors pose a significant global health challenge. While established screening methods exist, they are largely site-specific, limiting comprehensive prevention. Multi-cancer screening, which detects multiple cancers simultaneously, offers substantial advantages including optimized sample utilization, reduced participant costs, enhanced efficiency, and improved resource allocation. Artificial intelligence (AI) has emerged as a transformative technology, revolutionizing healthcare by analyzing complex biomedical data and enhancing diagnostic accuracy. Integrating AI into multi-cancer screening holds immense potential for advancing early cancer detection. This review provides a comprehensive overview of AI-driven multi-cancer screening. We examine foundational technologies and current applications, including biomarker data and medical imaging data analysis, as well as core AI techniques like machine learning, deep learning, natural language processing, explainable AI and essential preprocessing steps. We assess key technical bottlenecks such as data sparsity, model generalizability, false positives and false negatives, and affordability, alongside solutions like transfer learning, federated learning, and Bayesian optimization. Additionally, we highlight clinical validation, regulatory approval, and ethical considerations for multi-cancer screening. Furthermore, we explore future prospects, envisioning enhanced accuracy and expanded coverage, deeper multi-modal data fusion, personalized and dynamic screening, and intelligent decision-support systems with improved accessibility. We also outline targeted recommendations for developing countries conducting AI-driven multi-cancer screening, building on global best practices while adapting to local realities including those in China. This review offers a forward-looking perspective on how AI will evolve multi-cancer screening into a more personalized, dynamic and accessible cornerstone of cancer prevention.
Large language models (LLMs) have demonstrated encouraging performance for medical natural language processing (NLP) tasks, approaching human-equivalent performance in some of the standard benchmarks, positioning them as game-changers in healthcare. However, there remains a persistent gap between high benchmark performance and clinical utility of NLP algorithms due to limitations of existing evaluation paradigms. Existing benchmarks tend to use static, task-specific benchmarks, and as a result, they do not capture the full dimension of complexity, safety, interpretability, and integration in workflow required for the safe deployment in the clinic. The editorial advocates for a shift in evaluation paradigms from narrow score-based metrics to a dynamic, multi-dimensional, and patient-centered system of clinical gatekeeping. The proposed framework integrates a four-phase process, including retrospective benchmarking, pilot testing, multi-center validation, and real-world monitoring, alongside a capability-task-behavior-value progression and a continuous human-in-the-loop feedback mechanism. This comprehensive strategy ensures not only technical robustness but also clinical relevance, ethical accountability, and adaptive improvement, transforming LLMs from experimental tools into reliable clinical partners for safer and more patient-centric healthcare delivery.
Background: Antifreeze proteins (AFPs) are key in combating cold in living organisms and preventing ice morphogenesis. These proteins have applications in cryopreservation, food preservation, and biotechnology. Factors such as accurate prediction of AFP are considered essential for advancing these fields. Methods: In this study, a novel method, StackAFP, was developed using the stacking method and latent semantic analysis as the feature extraction technique for predicting antifreeze proteins. Four machine learning algorithms, random forest, XGboost, CatBoost, and LightGBM (LGBM), were used as the baseline models, and LGBM was employed as the meta-classifier to develop StackAFP. StackAFP was compared with different conventional machine learning methods to ensure the robustness of the proposed method. Results: StackAFP showed potentiality with an accuracy of 0.9997, a Matthews correlation coefficient, and a Kappa value of 0.9944. StackAFP outperformed the entire applied conventional machine learning model. Furthermore, StackAFP also outperformed the existing methods for identifying AFPs. Conclusion: The performance of StackAFP demonstrated its effectiveness, highlighted its potential in bioinformatics, and advanced our knowledge of AFPs.
Dry eye, a common eye disease globally, poses significant challenges to clinical diagnosis and management due to its complex pathogenesis and high incidence rate. The development of artificial intelligence (AI) technology has provided new opportunities for the analysis and auxiliary diagnosis of dry eye imaging. This expert consensus focuses on the classification and annotation methods of dry eye imaging, in line with the application needs of AI technology. It summarizes the scope and tasks of research on the classification and annotation of dry eye imaging and provides detailed standards for the principles and methods of classification and annotation of major imaging modalities, including lipid layer of the tear film, tear meniscus height, tear film breakup time, corneal fluorescein staining, and meibomian gland images. It also clarifies the tools and processes for classification and annotation. The consensus proposes systematic quality control requirements, including annotation consistency assessment, multi-round review, and data cleaning methods. Finally, the consensus summarizes the current challenges and proposes targeted solutions. The launch of this consensus aims to provide high-quality data support for the development of AI in dry eye, enhance the application effects of AI in dry eye diagnosis, disease monitoring, and personalized treatment, and offer scientific references and technical support for clinical and research applications of AI in the field of dry eye.
Objective Large language models (LLMs) show great promise in processing medical texts; The increasing specialization of clinical medicine has led to a greater demand for efficient and accurate referral. This study aims to evaluate the ability of the large language model DeepSeek-R1 to generate transfer notes for gastrointestinal surgery patients. Its performance is compared with that of clinician-provided transfer notes in terms of completeness, accuracy, and clinician preference. Methods A retrospective clinical analysis was conducted on 204 referral patients who underwent gastrointestinal surgery at Qingdao University Affiliated Hospital between January 2022 and June 2025. The LLM was trained using a small-sample study of four cases, and 200 cases were used as the test set. LLM-generated transfer notes were based on a structured template comprising predefined units. A thorough completeness review of the LLM-generated transfer records was conducted by two trained clinicians (κ = 0.719, p < 0.05). We quantitatively assessed the LLM's extraction performance by calculating recall, precision, and F1 scores within the LLM-generated transfer notes. McNemar's test was used to compare the completeness of LLM-generated and clinician-provided transfer notes. Five clinicians conducted blinded, paired preference evaluations. Results DeepSeek demonstrated excellent overall performance in information extraction among the 200 transfer notes generated, achieving high precision (99% [95% CI: 98%, 99%]), recall (97% [95% CI: 96%, 98%]), and an F1 score of 0.98 [95% CI: 0.97, 0.98]. LLM-generated transfer notes were comparable in completeness to clinician-provided notes. Within the “Current Diagnosis” unit, LLM-generated notes were significantly more complete than clinician-provided notes (90% vs. 81.5%; 180 vs. 163; P < 0.05). There were no statistically significant differences across the remaining five assessment units (all P > 0.05). In preference evaluation, clinicians were observed to demonstrate a pronounced preference for referral notes generated by LLMs (39% [78/200] vs. 13% [26/200], respectively; 48% [96/200] rated them as equivalent). Conclusion DeepSeek can generate transfer notes that are accurate and of a quality similar to that of clinician-provided notes. Further evaluation in actual clinical settings is necessary.
Background: Lymph node metastasis (LNM) status serves as a key prognostic marker in patients with locally advanced rectal cancer (LARC). Neoadjuvant chemoradiotherapy (nCRT) can influence lymph node status by inducing regression, with treatment responses exhibiting substantial variability among patients. We aimed to develop and validate a random forest-based radiomics model for non-invasively predicting lymph node regression status following neoadjuvant therapy in patients with LARC using pretreatment magnetic resonance imaging (MRI). Methods: This retrospective study included 285 patients with LARC who were treated with nCRT followed by elective resection at Qingdao University Affiliated Hospital between October 2019 and October 2023. Patients were randomly allocated to training and testing sets in a 7:3 ratio. Baseline characteristics including gender, age, body mass index (BMI), and other 13 clinical variables were compared between the groups, with no significant differences observed (P >0.05). High-resolution T2-weighted MRI scans of the rectal region were collected preoperatively. Regions of interest corresponding to primary tumor lesions were manually delineated on the MRI scans, followed by radiomic feature extraction. Feature selection was performed using the least absolute shrinkage and selection operator (LASSO) regression model. A machine learning random forest predictive model was subsequently constructed, incorporating selected radiomics features and clinical variables. Results: Model performance for predicting lymph node status was assessed using receiver operating characteristic curve analysis, area under the curve (AUC), and calibration curves. Four features were selected from 1,051 radiomic features to construct a radiomics model, achieving an AUC of 0.743 in the test set. Four features were also extracted from 19 clinical parameters to develop a clinical data model, with an AUC of 0.727. Integrating radiomic features and clinical data yielded a combined model with superior performance in the test set (AUC 0.794). Conclusion: The radiomics model derived from pretreatment rectal MRI in patients with LARC demonstrated strong predictive capability for assessing metastatic lymph node responses to nCRT.
Background: The risk of rupture associated with intradural internal carotid artery (ICA) aneurysms warrants considerable attention. We aimed to develop the first machine learning (ML) model that integrates standardized hemodynamic profiling with clinical and morphological data to stratify rupture risk in intradural ICA aneurysms. Methods: We consecutively enrolled 511 intradural ICA aneurysms that underwent DSA examinations at four hospitals from July 2017 to July 2022. Utilizing the electronic medical record system and computational fluid dynamics of AneuFlow software, we extracted 10 clinical baseline characteristics, 13 morphological, and 12 hemodynamic features for the aneurysms. Subsequently, the risk of aneurysm rupture was stratified by random forest (RF), XGBoost (XGB), LightGBM (LGB), and logistic regression (LR) models. Data from three hospitals were used to develop the internal training cohort (n = 331) and internal validation cohort (n = 83), while data from the fourth hospital contributed to the external validation cohort (n = 97). The models’ performance across the three cohorts was evaluated using area under the curve (AUC), sensitivity, specificity, and the Youden index. Additionally, we determined the feature importance ranking of the ML models. Results: The RF model achieved the highest AUC of 0.980 (95% CI: 0.969–0.989) in the internal training cohort. The AUC for the RF, XGB, LGB, and LR models in the internal validation cohort was 0.872 (95% CI: 0.792–0.929), 0.874 (95% CI: 0.794–0.931), 0.852 (95% CI: 0.769–0.914), and 0.827 (95% CI: 0.740–0.894), respectively. In the external validation cohort, the AUC for these models was 0.820 (95% CI: 0.729–0.891), 0.772 (95% CI: 0.675–0.851), 0.782 (95% CI: 0.686–0.859), and 0.782 (95% CI: 0.686–0.859), respectively. Moreover, the RF model achieved the greatest Youden index (0.517) in the external validation cohort, indicating superior discrimination ability. Hemodynamics accounted for 57% of the stratification power, with irregular geometry (nonsphericity index > 0.15) and microvascular inflammation markers (minimum wall shear stress < 0.3 Pa) identified as the key drivers. Conclusion: The ML framework designed for intradural ICA aneurysms demonstrated strong risk stratification capabilities, allowing more timely and personalized clinical diagnosis and treatment.
Background Large language models (LLMs), a revolutionary breakthrough in artificial intelligence, can be leveraged to automatically generate impressions for radiology reports, which usually require time, effort, and training. Our objective was to evaluate the performance of five recent LLMs (GPT-4, GPT-4o mini, Gemini 1.5-Pro, Gemini 1.5-Flash, and Llama 3.1) for impression generation. Methods In this retrospective study, 100 radiology reports were sampled (20 from each of the report-groups 0-400, 400-800, 800-1,200, 1,200-2,000, and 2,000-8,000 based on character count of the Findings section) from the publicly available "BioNLP 2023 report summarization" dataset (collected between 2001-2016, training subset of size 59,320 considered for sampling), sourced from PhysioNet. Then, each of the five LLMs was zero-shot prompted to generate impressions using the findings from the sample. Generated impressions were evaluated: (a) subjectively for coherence, comprehensiveness, conciseness, and medical harmfulness by two radiology fellows and a large reasoning model (LRM) Gemini 2.5-Pro, and (b) objectively using a composite accuracy metric including recall-oriented understudy for gisting evaluation (ROUGE)-1, bilingual evaluation understudy (BLEU) and cosine similarity, against the original human expert-generated impressions. The LLMs were ranked according to the percentage agreement ranking of subjective and composite scores. Statistical tests ( Friedman and post-hoc Nemenyi tests) were used to assess inter-model differences. Results The top-ranked models were Gemini 1.5-Pro, GPT-4, and Gemini 1.5-Flash. Performance varied across models for both human and LRM raters (Friedman test: Human P <1.82 & times;10(-6); LRM P <9.10 & times;10(-40)). Composite accuracy scores were significantly higher for the top three models (0.69, 0.68, and 0.68) versus others (0.65; Nemenyi P <1.11 & times;10(-)& sup1;(6)). The LRM aligned closely with human raters (2.15% complete disagreement) and identified all human-rated inaccurate impressions. Conclusion Gemini 1.5-Pro outperformed GPT-4, in terms of coherence, comprehensiveness, and medical harmfulness, at a lower cost. Human and LRM evaluations were generally consistent, though the LRM was more conservative.
Background: Early-stage high-grade lung invasive adenocarcinoma (IAC) has poor prognosis and is hard to identify using conventional radiological assessment. Current reliance on postoperative histology to identify high-grade subtypes delays risk-adapted surgical planning. Three-dimensional (3D) vision transformers (ViTs) may improve prediction by modeling long-range dependencies in computed tomography (CT) scans. We aimed to develop and validate 3D-ViT and Swin Transformer (SwinT) for preoperative CT-based prediction of early-stage high-grade IAC subtypes (micropapillary/solid), benchmarking against ResNet. Methods: A multicenter cohort of 1028 patients with surgically confirmed early-stage lung adenocarcinoma was divided into training (n = 806), validation (n = 100), and external test (n = 122) sets. 3D-ViT, SwinT, and ResNet models were trained on CT to classify nodules harboring high-grade histologic patterns. A novel decision-aid tool for IAC surgery was provided. Performance was evaluated using area under the curve (AUC), accuracy, sensitivity, specificity, and precision. Attention mapping was performed to interpret 3D-ViT decision-making. Results: The 3D-ViT model achieved AUC values of 0.856 (95% CI: 0.845–0.877) (validation) and 0.806 (95% CI: 0.790–0.816) (testing), compared to 0.854 (95% CI: 0.841–0.872) (validation) and 0.760 (95% CI: 0.743–0.776) (testing) for the ResNet baseline. 3D-ViT showed balanced accuracy, sensitivity, specificity, and precision in the validation set. In external testing, 3D-ViT significantly outperformed ResNet in all metrics with P < 0.01. The SwinT-based AlignSen model from decision-aid tool prioritized sensitivity for high-grade IAC (88.0% validation, 93.4% testing), while maintaining specificity (76.0% validation, 54.1% testing) which significantly outperformed ViT and ResNet-based AlignSen models’ specificity (60.0% and 54% validation, 52.5% and 24.6% testing, respectively). Attention maps highlighted 3D nodule heterogeneity and peripheral irregularities. Conclusion: The 3D-ViT model demonstrated robust accuracy and generalizability in predicting high-grade IAC subtypes using CT images. Integration of preoperative SwinT into clinical workflows may offer a viable alternative to intraoperative pathology subtyping, potentially reducing reliance on frozen sections while optimizing surgical planning.