
Objective Digital chronic disease management (CDM) service packages are an emerging service type of telemedicine platforms. This study aims to examine patients’ willingness to use (WTU), willingness to pay (WTP), and preferences for digital CDM service packages of internet hospitals among hospital outpatients in China’s high-income regions. Methods A cross-sectional study was conducted across the outpatient settings of multiple hospitals in Zhejiang province, which ranked first in per capita disposable income of urban residents among all China’s provinces. Results Among 1262 participants, 63.47% (801/1262) had WTU, 45.25% (571/1262) had WTP. Binary logistic regression analyses showed that younger age, higher education levels, having chronic disease, and being an online health information seeker were significantly associated with greater odds of both WTU and WTP. Higher monthly income (8001-17,000 CNY vs. ≤2000 CNY) was significantly associated with WTU, while higher monthly income (>17,000 CNY vs. ≤2000 CNY) and having basic medical insurance were significantly associated with WTP. Among 571 participants with WTP, more than 80% preferred physician-led packages at >300 CNY/month, 10.51% preferred nurse- and health manager-led packages (100-300 CNY/month), and 7.53% preferred AI-assisted platform and medical staff packages (<100 CNY/month). Ordered logistic regression found that marital status (divorced) was the only significant factor for package preferences, though this should be interpreted with caution due to limited subgroup size. Conclusion Findings replicated and extended the Andersen Behavioral Model and Grossman Model to the context of digital CDM services. A distinct economic threshold separated WTU from WTP for digital CDM service packages: WTU emerged at moderately high income levels, while WTP required high income and basic medical insurance coverage. Financial affordability constituted a substantial barrier. Among willing-to-pay patients, the majority preferred physician-led packages, suggesting that doctors should retain a central role in digital CDM service delivery.
Objective Spirometry, the gold-standard diagnostic tool for chronic obstructive pulmonary disease, is limited in primary care due to shortages of qualified technicians and inconsistent testing quality. Evidence for augmented reality (AR)-assisted spirometry remains limited. This three-stage study aims to evaluate the task performance, operational efficiency, cost-effectiveness, and usability of an AR-assisted spirometry system, AR-SPIRO. Methods The pre-evaluation stage included system development, technical verification, and usability testing. In Phase 1, a single-center randomized crossover agreement validation and usability evaluation study will compare the quality of spirometry between the AR-assisted and senior technician-led modes using weighted Kappa coefficients. In Phase 2, a multi-center, open-label, pragmatic, randomized implementation evaluation will be conducted at 14 primary care institutions. High-risk participants will be randomly assigned to AR-assisted or technician-led guidance. AR-SPIRO integrates a portable spirometer, a mobile terminal for real-time flow data processing, and AR glasses for synchronized visual–auditory feedback. Results The AR-SPIRO system has been developed, and technical verification and preliminary usability testing have been completed. The study protocol has been finalized, and the Phase 1 agreement validation and Phase 2 pragmatic implementation evaluation are planned according to the predefined protocol. Conclusion AR-SPIRO is expected to provide standardized spirometry guidance in primary care. This trial will provide evidence for the usability and implementation value of AR-SPIRO in resource-limited settings.
Background Lung cancer remains the leading cause of cancer-related mortality worldwide. Large language models (LLMs), including ChatGPT, DeepSeek, and Grok, have shown promise in clinical decision support, but differences in training and alignment may lead to variable performance. Current evaluations often rely on aggregate metrics or isolated tasks, which may not capture real-world clinical complexity. Methods We conducted a structured evaluation of three LLMs using nine simulated lung cancer cases across five clinical domains. LLMs’ outputs were anonymized, randomized, and independently scored by five senior lung cancer specialists under a double-blind design using a five-point Likert scale evaluating accuracy, comprehensiveness, relevance, and clinical applicability. Qualitative error analysis was also performed. Results Inter-rater agreement was moderate (Fleiss’ κ = 0.463; ICC (2, k) = 0.675). All LLMs achieved high scores across evaluation dimensions without statistically significant differences (P > 0.05). Given the limited number of simulated cases, these findings should be interpreted cautiously. Descriptive analyses suggested context-dependent performance patterns across clinical domains: Grok tended to show more consistent performance in diagnosis and treatment decision-making, DeepSeek showed comparatively lower descriptive performance in therapeutic decisions but higher applicability in prognosis and rehabilitation, and GPT exhibited relatively stable intermediate performance. No single LLM consistently outperformed others across all clinical scenarios. Conclusion LLMs demonstrate substantial potential in supporting lung cancer clinical workflows, but their performance appears to be context-dependent. The present findings are exploratory and suggest that task-specific evaluation may provide a more clinically informative framework than overall model ranking. Continued validation using larger and more diverse clinical datasets, together with appropriate governance and specialist oversight, remains essential for the safe integration of LLMs into clinical practice.
Background Artificial intelligence (AI) has become a transformative tool in knee osteoarthritis (KOA) research, providing new opportunities for diagnosis, prognosis, disease monitoring, and personalized management. However, the development trajectory, collaboration patterns, and emerging hotspots of this field remain insufficiently mapped. This study aimed to provide a bibliometric and visualized analysis of AI applications in KOA research. Methods Publications on AI applications in KOA from 2004–October 2025 were retrieved from the Web of Science Core Collection. Bibliometric indicators and collaboration networks, intellectual structure, and thematic evolution were analyzed using Bibliometrix, VOSviewer, CiteSpace, and Bibliometric.com . Results A total of 663 publications were included in the final analysis, comprising 582 articles, 2 early access articles, 3 proceedings papers, and 76 reviews. The annual growth rate was 28.03%, indicating rapid expansion of the field. China led in publication output, while the United States ranked first in total citations and average citations. International collaboration exhibited a tripolar structure dominated by North America, Europe, and Asia, although cross-regional integration remained limited. Osteoarthritis and Cartilage was the most productive and influential journal. Thematic evolution revealed a shift from algorithm-centered methodological development, including deep learning and radiomics, toward clinically oriented applications such as predictive modeling, imaging-based assessment, arthroplasty, and decision support. Conclusion AI-related KOA research has grown rapidly, evolving from methodological exploration toward clinical translation. Future studies should strengthen international collaboration, multicenter validation, and integration of AI into real-world KOA management.
Objectives Large language models (LLMs) are increasingly studied for clinical decision support, but high-risk cardiology exposes persistent weaknesses in hallucination control, guideline adherence, and medication-safety reasoning. Heart failure with reduced ejection fraction (HFrEF) is a demanding test case because safe care requires structured guideline-directed therapy, comorbidity-aware monitoring, and reliable risk warnings. Methods We developed a dynamic alignment framework using 1087 retrospective HFrEF cases from Affiliated Zhongshan Hospital of Dalian University. An open-source LLaMA-3.1 backbone was optimized through four sequential stages: continual pre-training for heart-failure domain adaptation, supervised fine-tuning for structured clinical responses, reinforcement policy optimization for safety-oriented alignment, and retrieval-augmented generation for guideline grounding. Models were assessed with dual-track clinical and linguistic metrics. Results LLaMA-3.1 was the strongest supervised baseline, but supervised fine-tuning alone did not fully resolve guideline-adherence limitations. Staged alignment produced a measurable Alignment Tax: the final retrieval-grounded variant improved the Clinical Score from 0.716 to 0.864 and reached a Guideline Score of 0.881, while BLEU-4 decreased from 0.371 to 0.272. The decline in surface overlap coincided with stronger risk safety, stricter structure, and more guideline-directed outputs. Conclusions Dynamic alignment shifted the model from linguistic mimicry toward clinically constrained HFrEF decision support. These findings suggest that staged optimization with policy alignment and retrieval grounding can improve evidence-based recommendations, while conventional language-overlap metrics may underestimate clinically safer generation.
Objective Understanding real-world movement behavior is essential for optimizing rehabilitation for people with lower limb amputation, yet clinical measures reflect performance mainly in controlled settings. Free-living studies are often small, and individuals with transtibial amputation generally demonstrate reduced physical activity and high sedentary time relative to population norms. Clinicians lack robust reference values that reflect daily-life behavior. Therefore, the aim of this study was to characterize 24-hour (h) movement behavior patterns in a larger cohort than typically reported and establish percentile-based reference values to support individualized rehabilitation. Methods A cross-sectional, observational design using objective monitoring of free-living movement behaviors was implemented. Ninety-six adults with unilateral transtibial amputation were recruited from prosthetic clinics across the United States (mean age 58±13.5 years; 24 female). Participants self-reported their age, sex, ethnicity, diabetes status, amputation etiology, years of prosthesis use, and the age of their prosthesis. Height and weight were measured, and amputation-adjusted BMI was calculated. Participants wore activPAL and Fitbit devices for seven days to quantify steps, sedentary time, prolonged sedentary bouts (>1 h), standing time, stepping time, sleep duration, and sit-to-stand transitions. Percentile distributions summarized movement behaviors, and time-use compositions described daily allocation of sedentary, standing, stepping, and sleep time. Results Participants averaged 4,712 steps/day with activPAL and 5,815 steps/day with Fitbit. Median sedentary time was 616 minutes (min) (10.26 h)/day, with mean sleep duration of 500 min (8.33 h)/night. At the extremes, the 5th percentile accumulated 701 steps/day with 386 min (6.43 h) of sedentary behavior, whereas the 95th percentile accumulated 10,691 steps/day and 905 min (15.08 h) of sedentary behavior. Stepping time ranged from 11 to 164 min (2.73 h)/day. Significant differences (p<0.006) were observed by functional mobility classification, obesity status, and diabetes status. Conclusion Free-living movement behavior showed substantial variability and pervasive sedentary time. Percentile-based benchmarks provide clinically useful context for assessing mobility and informing personalized rehabilitation targets.
Objectives In recent years, research interest in digital biomarkers has grown rapidly, driven by their capacity to enable continuous, objective, and personalized health monitoring. However, comprehensive bibliometric analyses of global research output in this field remain limited. This study aims to systematically evaluate the current status, hotspots, and emerging trends of global digital biomarker research using bibliometric analysis. Methods On August 11, 2026, we retrieved publications related to digital biomarkers from the Web of Science Core Collection (WoSCC). This study encompassed articles and reviews published from January 1, 2014 to August 11, 2026. Publication years, journals, authors, institutions, countries/regions, cited references, and keywords were systematically analyzed. VOSviewer was employed to conduct co-authorship, co-occurrence, and co-citation analyses and to construct network visualization maps. Results We evaluated a total of 1056 publications from 96 countries/regions, of which the United States was the main contributor. Harvard University, King’s College London and the Massachusetts General Hospital were the primary research institutions. Among the 7,188 contributing authors, Najafi, Bijan was the most prolific, while Horak, Fay B. was the most frequently cited. Keyword cluster analysis reveals four main research topics: (1) Parkinson’s disease and mobility monitoring in older adults, (2) AI-assisted diagnosis, classification, and prediction, (3) remote mental health assessment and passive physiological monitoring, and (4) cognitive dysfunction and neurodegenerative disease assessment. Conclusion Digital biomarker research shows the characteristics of continuous expansion and increasingly diversified themes. This bibliometric analysis clarifies major research directions and knowledge structures, supporting digital biomarker development and clinical translation.
Objective To evaluate large language models (LLMs) for stroke care in guideline-based question answering (Q&A) and individualized draft rehabilitation-plan generation, with explicit assessment of response quality, readability, and repeated-generation score stability. Methods A two-stage study was conducted. Stage 1: Four LLMs generated best answers to 100 stroke-related questions; the top two models by accuracy advanced to Stage 2. In Stage 2a, ChatGPT-o3 and DeepSeek answered 20 guideline-derived clinical questions, each repeated three times. In Stage 2b, the two models generated individualized rehabilitation plans for 60 de-identified stroke cases, with three independent generations per case. Three senior clinicians evaluated outputs across correctness, completeness, readability, helpfulness, and safety using 5-point Likert scales. Chinese readability was assessed using the LDU-TGP platform, and English readability was assessed using the Flesch-Kincaid Grade Level. Generalized estimating equations were used for Stage 2a and Stage 2b comparisons to account for repeated generations clustered within clinical questions and patient cases, respectively. Results In Stage 1, DeepSeek achieved the highest accuracy (91%), followed by ChatGPT-o3 (90%), Gemini (85%), and ChatGPT-4o (83%). In Stage 2a, no between-model differences across the five clinician-rated domains remained statistically significant after correction for multiple comparisons. ChatGPT-o3 generated clinical Q&A responses with a higher recommended reading age and English Flesch-Kincaid Grade Level, whereas the difference in Chinese Reading Difficulty Score did not remain significant after correction. In Stage 2b, ChatGPT-o3 significantly outperformed DeepSeek across correctness, completeness, readability, helpfulness, and safety in the GEE analysis. Repeated-generation score stability did not differ significantly between models. Conclusions Within this single-center, expert-rated benchmark, ChatGPT-o3 and DeepSeek performed comparably across most guideline-based Q&A dimensions, whereas ChatGPT-o3 achieved higher clinician-rated scores for draft rehabilitation-plan generation. These findings do not establish clinical effectiveness. Prospective patient-centered validation and clinician-supervised review are required before implementation in routine stroke rehabilitation.
Objective This study aims to examine the linkage between incidental information exposure and individuals’ continuous intention to use AI chatbots in healthcare, and whether this process varies by pre-existing attitudes toward human doctors. Method Data were collected through a cross-sectional survey ( N = 779) from June 11 to June 16, 2025. Structural equation modeling and multigroup comparison analysis were conducted to test the hypothesized relationships and group differences. Result Incidental exposure to information about AI chatbots in healthcare was positively linked to both subjective ( β = .378, p < .001) and objective knowledge ( β = .095, p < .05). Both subjective ( β = .222, p < .001) and objective knowledge ( β = .224, p < .001) about AI chatbots in healthcare were positively linked to positive attitudes toward these technologies, which was subsequently associated with continuous intention ( β = .468, p < .001). Multigroup analysis showed that the link between objective knowledge and positive attitudes toward AI chatbots in healthcare was stronger among those with high ( β = .299, p < .001) versus low ( β = .135, p = .014) positive attitudes toward human doctors, as was the link between positive attitudes and continuous intention (high: β = .506; low: β = .271, p s < .001). Conclusion Incidental exposure to media information about AI chatbots in healthcare was positively linked to continuous usage intention through knowledge acquisition and attitude formation. Individuals with divergent levels of pre-existing attitudes toward human doctors performed differently within the mechanism.
Objectives This study aims to systematically evaluate, through a meta-analysis, the effects of Virtual Reality (VR) interventions on depression and anxiety levels among older adults during and after the COVID-19 pandemic, and to explore variations in the effectiveness of VR interventions under different conditions. Method Databases including PubMed, Embase, Web of Science, the Cochrane Library, CINAHL, PsycINFO, CNKI, and Wanfang Data were searched. Studies were screened according to predefined inclusion criteria, and data on study characteristics and intervention outcomes were extracted. The quality of the included studies was assessed using the appropriate risk of bias tools, and meta-analysis was performed using Review Manager (RevMan) version 5.4.1 and Comprehensive Meta-Analysis (CMA) software. Results A total of 12 studies involving 724 older adults aged 60 years and above were included. Compared with control conditions, VR interventions significantly reduced symptoms of depression and anxiety among older adults. Exploratory subgroup analyses indicated that fully immersive VR was associated with a more pronounced pooled effect than semi- and low-immersive VR. Both shorter (<6 weeks) and longer (≥6 weeks) VR interventions were associated with significant improvements in anxiety and depressive symptoms, with no statistically significant difference between the duration subgroups. Conclusion During and after the COVID-19 pandemic, VR interventions may represent a promising complementary non-pharmacological approach for improving depressive and anxiety symptoms among older adults. Exploratory findings suggest that immersion level may contribute to variability in intervention effects; however, these findings should be interpreted cautiously given the limited available evidence. Further high-quality, large-sample randomized controlled trials (RCTs) are needed to clarify the potential influence of intervention characteristics on VR efficacy.
Objective Poor adherence to prescribed treatment for people with type 2 diabetes increases the risk of complications, mortality, and associated healthcare costs. Mobile phone text-messaging interventions such as DiabeText can support diabetes self-management and improve health outcomes. This study aimed to explore patients’ experiences with the DiabeText intervention and the contexts and mechanisms through which behavioral change might occur. Methods A post-hoc process evaluation using a mixed-methods design was conducted within primary care settings in Spain. The analysis combined a questionnaire-based study with 371 participants, aimed at rating the behavior change techniques embedded in the text messages, and a qualitative interview study with 19 participants. Quantitative data were analyzed descriptively, while qualitative data underwent thematic analysis. Results Participants reported predominantly positive experiences with the intervention, highlighting increased support, motivation, and self-care capacity. They valued the reliable and practical content as well as the user-friendly format, which were perceived as facilitating healthier lifestyle changes and supporting medication adherence. However, participants also suggested incorporating greater personalization to enhance acceptability. The mechanisms through which the intervention may foster behavioral change likely relate to the inclusion of behavior change techniques (BCTs) previously associated with positive outcomes—and which participants rated as highly helpful in supporting change. Nevertheless, further research is needed to determine the specific active ingredients that would maximize DiabeText’s overall impact. Conclusion DiabeText was well received by individuals with T2DM to support self-care. Introducing personalized features and messages that utilize BCTs previously associated with better outcomes could be a good starting point for better outcomes.
Background With the increasing volume and complexity of medical data, clustering methods have become valuable tools for discovering hidden patterns in unlabeled datasets. Objective This study aims to compare the performance of different clustering algorithms for analyzing medical datasets, and evaluate the effectiveness of different clustering quality criteria, including internal, stability and external metrics. Methods A descriptive analysis was conducted on five publicly available medical datasets: Indian Liver Patient Disease (ILPD), Breast Cancer, Pima Indians Diabetes, Statlog (Heart) and Polycystic Ovary Syndrome (PCOS). After preprocessing (handling missing values, normalization, and variable encoding), eight clustering algorithms—Hierarchical Clustering (HC), k-means clustering, Fuzzy ANalYsis (FANNY), Self-Organizing Tree Algorithm (SOTA), Divisive Analysis (DIANA), Partitioning Around Medoids (PAM), Clustering Large Applications (CLARA) and AGglomerative NESting (AGNES) were applied. Clustering quality was evaluated using internal measures (Connectivity, Silhouette Width, Dunn Index), stability measures (Average Proportion of Non-overlap (APN), Average Distance (AD), Average Distance between Means (ADM) and Figure of Merit (FOM)) and external measures (Accuracy and Adjusted Rand Index (ARI)). Results HC and AGNES consistently achieved the best internal and stability performance on the ILPD, Pima, and Statlog datasets, while k-means performed best on the Breast Cancer dataset. On the PCOS dataset, HC, AGNES, k-means, and DIANA showed comparable performance. However, external validation revealed that higher internal and stability scores did not necessarily correspond to greater agreement with clinical class labels, and the optimal algorithm varied across datasets. Conclusion No clustering algorithm consistently outperformed the others across all medical datasets. Although hierarchical methods showed strong internal and stability performance, these metrics alone were insufficient to predict agreement with clinical labels. Therefore, external validation should complement internal and stability metrics to provide a more comprehensive evaluation of clustering performance in medical data analysis.
Objective To provide an in-depth understanding of the care processes involving the use of the Long-Term Care Facilities (LTCF) tool in Nova Scotia, Canada. Specifically, this constructivist grounded theory study explored the perspectives of long-term care (LTC) residents, substitute decision-makers (SDMs), and nurses, with the aim of improving residents’ care. Methods A constructivist grounded theory study was conducted in LTC homes in Nova Scotia, Canada, between September 2022 and April 2023. Data were collected through 38 semi-structured interviews with residents, SDMs, nurses, supplemented by 41 field notes and 10 analytical memos. Data were analyzed using constructivist grounded theory methods. Results: The emerging substantive theory was developed primarily from the perspectives of nurses and LTC management, while residents’ and SDMs’ perspectives provided contextual and supportive insights. The participants described InterRAI LTCF care process improvement through relationship-centredness, intentionality, and effective change management. If implemented properly at the point of care, these strategies could facilitate consistent care improvement for LTC residents. Conclusions The findings have important implications for practice, policy, and future research. Care process improvement in LTC settings requires attention to relational, organizational, and technological factors. These implications are particularly relevant for LTC management, point-of-care nurses, and LTCF software designers seeking to enhance the quality and consistency of residents’ care.
Introduction Given the rapid development of digital healthcare interventions, incorporating patients’ perspectives is essential. However, experiences of internet-based exercise programs for chronic whiplash-associated disorders (WAD) remain unexplored. The aim of this study was to explore and describe experiences of an internet-based rehabilitation program combined with four physiotherapy visits for individuals with chronic WAD from a quality-of-care perspective. Methods This study employed a qualitative explorative and descriptive study design. Individual semi-structured interviews were conducted with 13 individuals with chronic WAD who had participated in an internet-based rehabilitation program combined with four physiotherapy visits. Data were analysed using a deductive qualitative content analysis based on Donabedian’s model of quality of care (structure, process, and outcome). Results The analysis resulted in three deductively-developed categories: appraisal of content and technical structure of the internet-supported program (structure); patient actions during the rehabilitation process with or without interaction with the internet-supported program (process); and describing valuable outcomes from the patient’s perspective (outcomes). Each category was supported by 3–5 sub-categories. Conclusions Individuals with chronic WAD participating in a neck exercise program with digital support experienced positive health outcomes. The digital format was appreciated, but a more user-friendly format would have been preferred. Instead, the informants relied more on the physiotherapists’ instructions, and the digital support served as a good complement to the visits at the physiotherapy clinic. Further knowledge is needed regarding how to combine digital and face-to-face visits and how to make digital interventions in healthcare more user-friendly for both patients and healthcare providers.
Background The prognostic value of the Prognostic Nutritional Index (PNI) in critically ill patients with coronary heart disease (CHD) remains unclear. This multicenter retrospective cohort study aimed to investigate the association between PNI and in-hospital mortality in severe CHD patients and develop a machine learning-based predictive model. Methods This study is a multicenter retrospective cohort study, extracting data of adult critically ill patients with CHD who met the inclusion criteria from the MIMIC-IV database and the eICU-CRD database. Multivariate logistic regression and restricted cubic spline (RCS) analyses were used to evaluate the relationship between PNI and in-hospital mortality and 365-day mortality. The Boruta algorithm and LASSO regression were used to screen predictive features, six machine learning models were constructed, and SHapley Additive explanation (SHAP) was used to interpret the importance of model features. Results Low PNI was significantly associated with higher in-hospital mortality (OR=0.95, 95% CI 0.94–0.97). RCS revealed a linear negative relationship between PNI and mortality. The Random Forest model demonstrated superior predictive performance (AUC=0.846), significantly outperforming traditional scoring systems. SHAP analysis identified CRRT, Sepsis, Vasopressor, PNI, and BUN as key predictors. An online clinical calculator was developed based on these findings. Conclusions This study developed a machine learning-based calculator to predict in-hospital mortality in patients with CHD in the ICU ( https://leongxj.shinyapps.io/chd-mortality-prediction/ ). This calculator may guide clinical decision-making and improve the treatment of patients with CHD by identifying those at higher risk of in-hospital death.
Background Telehealth is a system that makes use of information and technology communication to ensure that patients get access to healthcare services regardless of the distance and shortage of healthcare professionals. Telehealth is a strategy that can be applied to complement the management of noncommunicable diseases (NCDs) in Africa. This review therefore aims to: (1) collate the available data about the process and outcomes of telehealth programs implemented in Africa and (2) explore the challenges experienced with the implementation of telehealth in the African context. Methods The scoping review will summarize the information by accessing the following databases; Ebscohost (CINAHL + MEDLINE, PsychINFO, Web of Science and PubMED), Scopus and google scholar. The databases will be searched using specific search terms that are aligned with the aim of the review. Studies will be included if they investigated the telehealth interventions for NCDs in Africa. A screening process that includes screening of the abstracts and full text will identify the studies to be included. A data capture template designed specifically for the study will be used to capture data. The findings will be presented descriptively and thematically. Conclusion This study underscores the increasing opportunity that exists for telehealth interventions to address NCDs, in resource-constrained environments. The intended review could confirm the potential of digital health to increase access to care, promote patient self-management, and engage health systems where resources are lacking.
Objective Magnetic Resonance Imaging (MRI) is essential for understanding post-stroke cognitive impairment (PSCI), yet the global research landscape remains unclear. This study aimed to analyze the trends and hotspots in MRI research on PSCI through bibliometric analysis. Methods Publications were retrieved from the Web of Science Core Collection on June 2, 2026, covering January 1, 2000 to June 2, 2026. Only English-language records classified as “Article” were included. Bibliometric indicators, collaboration networks, and keyword dynamics were analyzed using VOSviewer (version 1.6.20), CiteSpace (version 6.4.R1), and the R package bibliometrix (version 4.5.3). Results A total of 2,022 articles were included. The USA (438 corresponding-author articles) and China (422) were the leading contributors. The USA had the highest total citations (25,505), while average citation impact was higher in the UK and the Netherlands than in China. The University of California System (314 publications), Boston University (295), and Harvard University (272) were the most productive institutions. Stroke published the most articles (171) and showed the strongest network prominence in journal analyses. Keyword co-occurrence revealed five major thematic clusters spanning vascular risk and cerebral small vessel disease markers, cognitive phenotype and assessment, and network-oriented neuroimaging. Burst detection indicated current research frontiers through 2026 centered on “white matter hyperintensities,” “cerebral small vessel disease,” “functional connectivity,” “connectivity,” and “recovery.” Conclusion MRI research on PSCI expanded substantially over the study period, while thematic emphasis shifted from mainly structural and vascular lesion markers toward network-based and prognosis-oriented imaging. Future progress will require stronger international collaboration, more standardized MRI and analytic workflows, and external validation of multimodal prediction models in well-phenotyped cohorts before clinical translation.
Background Digital health technologies improve access to healthcare and offer innovative ways to assist its users, as health systems increasingly integrate telemedicine strategies into medical service routines. Understanding the factors that influence the acceptance and use of telemedicine-related technologies is fundamental. The unified theory of acceptance of technology (UTAUT) is one of the most frequently applied theoretical frameworks for explaining technology adoption and has evolved since its development. Mapping the literature on the application of UTAUT in telemedicine studies is key to understanding the future of research in this area. Objective The main objective of this study is to map the literature on the application of the UTAUT model for telemedicine acceptance and use, in order to assist in developing a framework for future research agendas. Methods A systematic literature review of the application of the UTAUT model to telemedicine acceptance and use was conducted, following the PRISMA protocol. The textual corpus obtained from the collected data was analyzed, with the support of VOSviewer and IRAMUTEQ software, to assist in developing a framework for future research agendas in telemedicine. Results The current literature on the application of the UTAUT model to the acceptance and use of telemedicine was mapped, analyzed, and applied to develop a research agenda framework that provides guidance for understanding and structuring research directions. Conclusion This study maps UTAUT's application in telemedicine and identifies key factors shaping technology acceptance and use. It advances theoretical understanding, offers a framework for future research, and supports the integration of digital health technologies to improve healthcare delivery.
Objectives Large language models (LLMs) are being used to facilitate academic writing. We aimed (1) to assess the quality, integrity, and factual reliability of LLM-assisted academic writing in medicine, and (2) to compare the performance of open-source versus proprietary LLMs, reflecting differences in model transparency. Methods In this prospective, controlled evaluation study, an open-source model (DeepSeek-R1) and proprietary model (ChatGPT-o1) completed ten academic tasks: generating five scientific essays and five evidence-based question-answers covering clinical topics in neuroradiology. Two radiologists rated outputs using 14 Likert and 7 binary criteria across the domains of academic quality, linguistic expression, factual reliability. Paired Wilcoxon/McNemar (Holm) assessed differences; ICC/Cohen’s κ assessed reliability. Results DeepSeek-R1 achieved higher overall mean Likert scores than ChatGPT-o1 (mean ratings ± standard deviation: 3.23 ± 0.44 vs 3.02 ± 0.30, p = 0.021). It significantly outperformed in reasoning depth (p = 0.015), contextual coherence (p = 0.043), subtlety (p = 0.037), and evidence integration (p = 0.027). Although both models cited sources in every output, citation reliability was poor. DeepSeek-R1 cited more real publications (55.1% vs. 33.3%, p=0.053) but showed a higher confirmed fabrication rate (36.7% vs. 2.6%, p<0.001). DeepSeek-R1 respected word limits in 90%, ChatGPT-o1 only in 50%. Neither reviewer reliably noticed AI‐generated text features. Inter-rater reliability was poor for Likert criteria (ICC = 0.466) and substantial for binary items (κ = 0.746). Conclusion DeepSeek-R1 outperformed ChatGPT-o1 in generating academic content for neuroradiology, highlighting the potential of transparent, locally deployable models. However, moderate quality, citation errors and hallucinations indicate that LLMs are not yet sufficient for unsupervised academic writing and demand rigorous human review.
Objectives To develop and internally validate an interpretable random survival forest (RSF) model for short-term mortality in adults with traumatic brain injury (TBI) and intracranial hematoma and to evaluate the association between the systemic immune-inflammation index (SII) and mortality. Methods This retrospective single-center study included 356 patients allocated to training n = 249 and internal validation n = 107 cohorts. All-cause mortality from admission was evaluated at 30, 90, and 180 days. The study included patients admitted between August 1, 2024, and December 31, 2025. Thirty demographic, clinical, imaging, and laboratory variables were considered. Univariate screening and least absolute shrinkage and selection operator regression identified eight predictors. Seven time-to-event models were compared using concordance, time-dependent discrimination, calibration, decision-curve analysis, individualized survival estimates, and interpretability. The RSF model was interpreted using SHAP. SII was calculated as platelet count multiplied by neutrophil count and divided by lymphocyte count and was natural log-transformed. Results During follow-up, 188 patients died; 141, 154, and 169 deaths occurred within 30, 90, and 180 days, respectively. RSF provided the best overall balance of time-dependent discrimination, calibration at 90 and 180 days, clinical net benefit, stable individualized survival estimation, and interpretability. Validation AUCs were 0.902, 0.899, and 0.952 at 30, 90, and 180 days, respectively. Continuous ln-SII was independently associated with mortality after multivariable adjustment HR 1.29 , 95 . This association remained consistent with administrative censoring at 30, 90, 180, and 365 days (HR range 1.30-1.34; all P<0.001), and 1,000 bootstrap resamples supported estimate stability. Conclusion Among the evaluated models, RSF provided the most clinically informative balance of internal predictive performance, dynamic risk estimation, and patient-level explainability. Higher ln-SII was consistently associated with mortality; data-driven nonlinear breakpoints and subgroup findings require independent validation.