OBJECTIVE:We examined whether GPT-4o, a widely used large language model (LLM), could produce age- and education-appropriate versions of complex pediatric traumatic brain injury case descriptions, while preserving clinical accuracy and emotional tone. METHODS:Five cases were adapted into four audience scenarios. Text complexity was assessed via Flesch-Kincaid (FKS), Gunning Fog, and SMOG indices. Clinical human experts rated text fidelity and emotional appropriateness on a 3-point scale. RESULTS:Original texts showed very high complexity (FKS 18.2-20.5), equivalent to 18-20 years of education. Adaptations for parents with high school education were often over-simplified (FKS 4.75-7.1), while versions for 12-year-olds were well-matched (FKS ~5-6). Texts for 8-year-olds had FKS scores of 4.0-6.8 (above grade 2-3 targets) and reduced fidelity (scores 1-2). Emotional tone was consistently rated appropriate across all audiences. CONCLUSION:Clinicians may use LLMs to draft explanations, but must carefully review and tailor them.
Background In post-acute stroke rehabilitation, cognitive assessment is clinically important but frequently incomplete. Complete-case analysis excludes much of the sample, reducing representativeness.Objectives To determine whether incomplete routine post-stroke cognitive assessments can still yield clinically interpretable relationships among cognitive, clinical, and contextual variables, and support group-level stratification using a reduced feasible cognitive set.Methods Retrospective cohort study of adults admitted to post-acute stroke neuro-rehabilitation between 2007 and 2026 with at least one neuropsychological assessment. The baseline cohort included 2654 patients, with 5417 assessments within 6 months available for confirmatory analysis. The 24-subtest baseline battery showed 37.5% to 88.4% missingness. Five representative cognitive measures were selected based on data availability, domain coverage, and within-domain Spearman correlations. Bayesian networks were fitted in the baseline cohort, their stability was assessed by bootstrapping, key dependencies were examined across all available assessments.Results Stable, bootstrap-supported dependency patterns were recovered and broadly replicated in the confirmatory analysis. Replicated associations linked age and time since injury to verbal learning and naming; stroke subtype to verbal learning, timed orientation/performance, and phonemic fluency. Verbal learning occupied a central position, with additional replicated links to naming, timed orientation/performance, and attentional span. Among the contextual variables, economic status was the only one showing a direct link to a cognitive measure (attentional span). Group-level stratification of verbal learning showed moderate performance (R2 = 0.42).Conclusions Incomplete routine cognitive assessments need not rely on complete-case restriction or score imputation; they can still support clinically meaningful interpretation and group-level stratification in post-acute stroke rehabilitation.
PurposeTo identify and characterize latent trajectories of depression using growth mixture modeling (GMM) applied to irregularly timed assessments spanning up to 20 years post-injury, examine baseline rehabilitation variables significantly associated with trajectory membership, and determine predictors of sustained depressive burden.MethodsThis retrospective observational cohort study included adults with traumatic or non-traumatic spinal cord injury admitted for inpatient rehabilitation within 3 months post-injury (2005-2023), with follow-up extending until May 2025. Depressive symptoms were assessed using the Hospital Anxiety and Depression Scale (HADS-D) at admission, discharge, and follow-up, totaling 3,258 assessments (n = 679 patients). GMM identified latent trajectories while accommodating irregularly spaced assessments. Predictors of trajectory membership were analyzed through multivariable regression with quantified model discrimination.ResultsA 3-class GMM solution provided optimal fit (entropy=0.75). Most participants (81.6%) followed a stable low-depression trajectory (Class 1). A borderline depression trajectory (Class 2; 8.4%) remained persistently elevated, while a probable depression trajectory (Class 3; 10.0%) displayed delayed worsening peaking at 5-7 years post-injury. Class 3 included a significantly higher proportion of non-traumatic injuries (63%) and females (44.1%), with most patients (85.3%) showing no depressive symptoms during rehabilitation, but later exhibiting a marked increase in depressive burden. Logistic regression predicting Class 2 achieved good discrimination (AUC=0.81; 95% CI, 0.65-0.97), identifying baseline depressive symptoms, tetraplegia, female sex, and primary level of education as significant predictors.ConclusionsIrregularly sampled follow-ups revealed distinct depression trajectories, including delayed-onset risk. Findings emphasize early rehabilitation-based screening and long-term monitoring to target follow-up and psychological support.
CONTEXT:Effective communication of complex medical information is critical for individuals with spinal cord injury (SCI) and their families, but this need remains largely unmet. Artificial intelligence (AI), including large language models (LLMs) like GPT-4o, may help simplify clinical texts while preserving essential medical information. However, their ability to adapt content across ages, education levels, and languages without compromising accuracy has not been systematically evaluated. FINDINGS:We analyzed short excerpts (∼100 words) from de-identified Spanish-language SCI clinical reports that had historically posed communication challenges. A bilingual co-author produced human English translations, creating parallel English and Spanish corpora (10 texts per language: 4 originals and 6 audience-tailored simplifications). GPT-4o was applied for within-language simplification, not translation. Baseline readability quantifies the complexity of real-world clinical documentation directly. All original English texts had very high complexity by Flesch-Kincaid Grade Level (FKGL >18), whereas simplified texts reached FKGL 7.3 for children and 9.8-10.6 for adolescents and adults. Experts rated child-directed fidelity 2-3/5; adult-directed texts scored 4-5/5 for fidelity and appropriateness.Original Spanish texts ranged from "Difficult" to "Very Difficult" by the Fernández-Huerta (FH) index, showing greater variability than English. Readability classifications were not concordant across languages (3/10 cases, 30%); e.g. one simplified version was "Plain English" by FKGL but "Fairly Difficult" by FH, highlighting language-specific behavior. CONCLUSION/CLINICAL RELEVANCE:GPT-4o can tailor complex SCI clinical excerpts to specific audiences in English and Spanish, but child-directed versions may lose clinically relevant information. Clinician oversight remains essential for safe patient communication.
Background: Transcutaneous spinal cord stimulation (tSCS) is a promising approach to enhance functional recovery after spinal cord injury (SCI). However, evidence on repeated sessions and their effects on gait and lower-limb strength remains limited. This study evaluated the effects of multisegmental tSCS on walking ability, muscle strength, and functional independence in individuals with SCI. Methods: In this randomized controlled trial with a partial crossover design, twelve individuals received tSCS combined with gait rehabilitation, while ten underwent gait rehabilitation alone, for three weeks. Four participants crossed over to the tSCS group after a minimum one-week washout period following the control intervention. We assessed the American Spinal Injury Association Impairment Scale (AIS), Total Motor Score (TMS), Lower Extremity Motor Score (LEMS), Walking Index for Spinal Cord Injury II (WISCI-II), 10- and 6-Meter Walking Tests (10MWT, 6meterWT), Timed Up and Go (TUG) test, maximal voluntary contraction (MVC) of the quadriceps (QM) and tibialis anterior (TA), and the Spinal Cord Independence Measure (SCIM-III); tSCS was applied at three spinal segments during gait rehabilitation over 15 sessions. Results: tSCS significantly improved MVC in both muscles, as well as SCIM-III and TUG, and these improvements were maintained at follow-up, with no significant adverse events reported. Other clinical assessments also showed significant improvement in both groups. Conclusions: tSCS was well tolerated and conferred additional benefits in lower-limb muscle strength, walking ability (as assessed by TUG), and functional independence, supporting its potential as a valuable adjunct to rehabilitation.
OBJECTIVE:Inpatients undergoing stroke rehabilitation experience high malnutrition rates, requiring strict dietary management. However, manual and time-pressured dietary provision can cause errors in diet composition, highlighting the need for innovation. Therefore, we aimed to evaluate whether GPT-4o can accurately identify dietary errors in hospital-based stroke rehabilitation menus, analyze differences in AI vs. expert rationale for decisions, and explore AI's potential role in clinical workflows through a structured collaboration framework. METHODS:A TRIPOD-compliant validation study analyzing 264 hospital-based menus designed for stroke rehabilitation inpatients requiring specialized diets (e.g., dysphagia, diabetes). GPT-4o's dietary compliance classifications were assessed using a structured 0-error, 1-error, and 2+ error framework, with expert dietitians as ground-truth in a rehabilitation hospital nutrition department, where expert dietitians selected menus from existing clinical practices for inpatients on specialized diets. AI-expert agreement, overall accuracy, sensitivity, and specificity in dietary error classification were assessed. AI vs. expert justifications were analyzed thematically to identify differences in decision rationale. Cohen's Kappa (95% CI) measured inter-rater reliability. Overall accuracy, sensitivity, and specificity were calculated using a 3 × 3 confusion matrix, comparing AI classifications (0-error, 1-error, 2+ error) to the expert-labeled ground truth. Thematic analysis categorized AI vs. expert justifications for flagged dietary errors. RESULTS:Out of 264 menus (1,000+ food items), 26 (9.8%) had discrepancies. Among these, 57.7% (15 cases) were PAS-based dysphagia diets, followed by diabetic (19.2%, 5 cases) and allergen-related (15.4%, 4 cases) diets. The remaining two cases involved low-sodium and low-fat diets. Cohen's Kappa: 0.892 (95% CI: 0.845-0.939, p < 0.001). 0-errors: Sensitivity 94.3%, specificity 100%; 1-error: Sensitivity 86.2%, specificity 96.6%; 2+-errors: Sensitivity 97.8%, specificity 92.6%. Thematic analysis revealed GPT-4o followed strict rule-based interpretations, whereas dietitians incorporated patient tolerance and food preparation considerations. CONCLUSION:GPT-4o demonstrated high accuracy but over-flagged violations, supporting its role as a prescreening tool with expert collaboration.
Meaningful participation in occupations or employment and/or the ability to engage in societal roles holds significant implications for one’s wellbeing and is internationally recognised as a fundamental right for all persons, nowadays representing an emerging policy-making goal. We aimed to identify novel classes of individuals with chronic spinal cord injury (SCI) having similar long-term trajectories of community integration and relate them to their demographic and clinical features using a retrospective observational design. Community Integration Questionnaire (CIQ) follow-up assessments, motor Functional Independence Measure (mFIM) categorised as poor, fair or good, and Hospital Anxiety and Depression Scale (HADS) were analysed. Growth mixture models (GMM) were fitted to identify individuals with similar CIQ trajectories, classes’ predictors were identified using multivariate logistic regression. GMM identified three classes of trajectories of community-dwelling adults with SCI (n=238) living in Catalonia, Spain assessed in-person (between 2002 and 2022) up to 19 years post-injury: Class 1 (n=46, 19.3 per cent): male (56.5 per cent), aged 53.5 (16.6) years at injury, mFIM (poor 39.1 per cent, fair 23.9 per cent, good 37.0 per cent), mean total CIQ=9.9 (3.4), depressive (21.7 per cent), tetraplegia (39.1 per cent). Class 2 (n=41, 17.3 per cent): male (56.1 per cent), 57.4 (14.8) years at injury, mFIM (poor 26.8 per cent, fair 12.2 per cent, good 61.0 per cent), CIQ=9.3 (3.8), depressive (7.3 per cent), paraplegia (65.9 per cent). Class 3 (n=151, 63.4 per cent): male (68.9 per cent), 43.6 (15.9) years at injury, mFIM (poor 11.9 per cent, fair 13.9 per cent, good 74.2 per cent), CIQ=17.7 (3.3), depressive (4.0 per cent), paraplegia (74.2 per cent). Admission age, higher education (university), mFIM and HADS depression predict good community integration, AUC: 0.82 (0.73–0.91). Our results suggest possible course of action focusing on specific aspects to promote community integration.
BackgroundLack of information is a critical challenge in occupational health. With over 180 million users, ChatGPT has become a prominent trend, swiftly addressing a wide array of queries, yet it critically needs validation in occupational health.ObjectiveThis study evaluated GPT-3.5 (free version) and GPT-4 (paid version) on their ability to respond to Occupational Risk Prevention formal multiple-choice questions.MethodsA total of 303 questions were assessed, categorized across four levels of complexity-task-specific, national, European, and global-within various Spanish regions.ResultsGPT-3.5 achieved an overall accuracy of 56.8%, while GPT-4 reached 73.9% (p < 0.001). GPT-3.5 showed particularly limited performance on domain-specific content. Both models shared similar error patterns, with incorrect response rates ranging from 18-24% across regions.ConclusionDespite GPT-4's improved performance, both models display notable limitations in occupational health applications. To enhance reliability, four strategies are proposed: formal validation, continuous training, error analysis, and regional adaptation.
BACKGROUND:Digital Twins (DTs) have transitioned from theory to reality, with growing applications in healthcare. Data generated by technologies (e.g. rehabilitation robots), essential for DT implementation, though widely produced in clinical settings, remains untapped in DT stroke rehabilitation, highlighting a gap compared to broader healthcare use. OBJECTIVES:We conducted a scoping review to i) define DT rehabilitation objectives, their input data, generation methods and user involvement; ii) analyze mechanisms underpinning DT models and outputs; iii) map key stakeholders driving innovation; iv) identify desirable properties for DT studies from broader healthcare literature and map them to stroke rehabilitation DT studies. METHODS:Following PRISMA-ScR guidelines, PubMed, Scopus, Web of Science and Google Scholar were searched for studies including only empirical data. Full-text reviews were conducted by three reviewers through repeated calibration. RESULTS:Sixteen studies were included, addressing five rehabilitation objectives: upper-limb (10), gait (3), and engagement, mental health, and general/planning (1 each). Patient sample sizes varied widely, with one retrospective study including 1,216 patients, while 15 studies involved 54 patients in total (median = 1).We identified 16 DTs mechanisms (e.g. variational autoencoders, Hill muscle models) and outcomes (e.g. exoskeleton control, upper-limb exercise delivery, gait torque estimation, impaired hand-mobility quantification). Academic institutions conducted 12 studies, Europe contributed 8 studies across 6 countries. Of 25 desirable properties identified, 8 (e.g. reproducible algorithms) showed high adoption, while 15 (e.g. cost-effectiveness, clinical integration) showed low/very low adoption by included studies. CONCLUSIONS:DTs in stroke rehabilitation show promise, though challenges remain (e.g. patient involvement, scalability).
Background Stroke now represents the condition with the highest need for physical rehabilitation worldwide, with only low or moderate-level evidence testing telerehabilitation compared to in-person care. We compared functional ambulation in subacute patients with stroke following telerehabilitation and matched in-person controls with no biopsychosocial differences at baseline. Methods We conducted a matched case-control study to compare functional ambulation between individuals with stroke following telerehabilitation and in-person rehabilitation, assessed using the Functional Ambulation Categories (FAC) and the Functional Independence Measure™ (FIM). Results The telerehabilitation group (n = 38) achieved significantly higher FAC gains (1.5 (1.3) vs 1.0 (1.0)) than the in-person rehabilitation group, with no differences in ambulation efficiency, in individuals: admitted to rehabilitation within 60 days after stroke onset; aged 49.8 (±11.4) years at admission; 55.3% female sex; moderate stroke severity; 42.1% with ‘good’ motor FIM at baseline; mostly living with sentimental partner (73.7%); with 21.1% holding an university education degree. Conclusions The groups showed no significant differences in ambulation efficiency, though the telerehabilitation group achieved higher FAC gains. Our results suggest that home telerehabilitation can be considered a good alternative to in-person rehabilitation when addressing ambulation in patients with moderate stroke severity and whose home situation mostly includes a cohabiting partner.
Hospital-based retrospective epidemiological research. To describe the epidemiological and demographic characteristics of patients with traumatic spinal cord injury (TSCI) in Catalonia from 1972–2022. Neurological university hospital in Catalonia. All patients diagnosed with TSCI admitted to the hospital from 1972–2022 were retrospectively reviewed. Etiology categories, neurological level of injury, American Spinal Injury Association Impairment Scale (AIS), Functional Independence Measure (FIM), and crude incidence rates were analyzed. A total of 3092 individuals with TSCI met the criteria. The crude annual incidence rate was 0.94/100,000 inhabitants. Mean age rose significantly over the years, from 27.7 (SD = 10.9) during 1972–1981 to 41.8 (SD = 17.1) during 2012–2022. The proportion of females constantly increased during 1982–2022 (18.4−23.1
Objectives: To (1) compare baseline clinical and demographic characteristics of postacute stroke inpatients who were diagnosed with first-time urinary tract infection (UTI) versus inpatients who were not; (2) compare rehabilitation outcomes between both groups; and (3) examine associations between time to UTI event and risk factors. Design: Retrospective observational cohort study. Setting: Institution for inpatient neurologic rehabilitation. Participants: Inpatients (n=1683) admitted within 3 months poststroke to a rehabilitation facility between 2005 and 2023. Interventions: Not applicable. Main Outcome Measures: Functional independence measure (FIM), functional ambulation categories (FACs) at admission. Cox proportional hazard models analyzed the association between UTI event timing and risk factors. Results: Of the (n=1683) included patients, 196 (11.6%) experienced a UTI. In 32.1% of cases, the UTI occurred during the first week after admission to rehabilitation and 47.9% of UTIs occurred during the first 2 weeks. The median (interquartile range) time to UTI was 16 (5-37) days since admission. Most common germs were Escherichia coli (40.5%), Klebsiella pneumoniae (23.7%), and Pseudomonas aeruginosa (6.4%). Patients who acquired a UTI had older age, higher stroke severity, higher proportion of dysphagia, hypertension, neglect, bilateral affectation, atrial fibrillation, hemiplegia, lower levels of functional independence, and lower FAC. We identified no differences in gender, type of stroke (ischemic or hemorrhagic), time to admission, aphasia, diabetes, dyslipidemia, chronic obstructive pulmonary disease, dominant side affected, and educational level between both groups. Patients with UTI presented significantly poorer rehabilitation outcomes including lower discharge FIM and FAC, larger length of stay, lower FIM efficiency, and decreased FIM effectiveness. Multivariable Cox proportional hazards identified hypertension HR=1.60 (1.13-2.27), admission FIM HR=0.98 (0.97-0.99), admission body mass index HR=0.96 (0.93-0.99), and admitted with catheter HR=1.80 (1.22-2.64) as significant predictors of time to first UTI event (Concordance-index=0.754). Conclusions: UTIs identification, characterization, and predictive factors can support postacute stroke mitigation strategies to minimize UTI-related complications and optimize rehabilitation outcomes. Archives of Physical Medicine and Rehabilitation 2025;106:729-37 (c) 2024 by the American Congress of Rehabilitation Medicine.
Sleep quality critically influences recovery in neurological patients, yet its longitudinal monitoring during hospitalization remains limited. Nursing narrative notes offer an underutilized resource to track sleep trajectories objectively across time.To propose and apply a formal pipeline that integrates structured clinical data and unstructured nursing annotations to monitor sleep trajectories during post-acute inpatient neurorehabilitation, relying exclusively on free-to-use software tools and without increasing nursing workload.A total of 17,039 nighttime nursing annotations were extracted and categorized into four sleep quality states. Two expert raters manually labeled a training set of 2,000 annotations (κ = 0.84). A random forest classifier achieved 0.93 sensitivity and 0.94 specificity and was used to classify the remaining notes. Sleep sequences were constructed and clustered using sequence analysis (TraMineR) and hierarchical clustering (AGNES, Ward's method). The obtained clusters (silhouette = 0.40) were compared using non-parametric statistics across clinical, functional, and social variables in a cohort of 303 post-acute consecutive neurorehabilitation inpatients.Four distinct sleep trajectory clusters were identified, each characterized by unique functional and socio-environmental profiles. The first group (n = 102; 33.7%) combined high functional independence, strong social support, stable economy, short hospitalization, and favorable sleep quality. The second group (n = 76; 25.1%) presented moderate functional independence, precarious economic conditions, and the highest proportion of poor sleep quality. The third group (n = 76; 25.1%) exhibited severe functional impairment, long hospitalization, poor housing conditions, but paradoxically the highest proportion of good sleep quality. The fourth group (n = 49; 16.2%) showed profound disability, relatively favorable socio-economic conditions, and predominance of intermediate sleep quality, likely influenced by medication. Distinctive sets of social and functional keywords emerged for each cluster.This pipeline identified clinically meaningful sleep profiles from nursing notes, highlighting functional and social determinants' role in shaping neurorehabilitation sleep trajectories.
PURPOSE:Our study aimed to: i) Assess the readability of textbook explanations using established indexes; ii) Compare these with GPT-4's default explanations, ensuring similar word counts for direct comparisons; iii) Evaluate GPT-4's adaptability by simplifying high-complexity explanations; iv) Determine the reliability of GPT-3.5 and GPT-4 in providing accurate answers. MATERIAL AND METHODS:We utilized a textbook designed for ABPMR certification. Our analysis covered 50 multiple-choice questions, each with a detailed explanation, focusing on non-traumatic spinal cord injury (NTSCI). RESULTS:Our analysis revealed statistically significant differences in readability scores, with the textbook achieving 14.5 (SD = 2.5) compared to GPT-4's 17.3 (SD = 1.9), indicating that GPT-4's explanations are generally more complex (p < 0.001). Using the Flesch Reading Ease Score, 86% of GPT-4's explanations fell into the 'Very difficult' category, significantly higher than the textbook's 58% (p = 0.006). GPT-4 successfully demonstrated adaptability by reducing the mean readability score of the top-nine most complex explanations, maintaining the word count. Regarding reliability, GPT-3.5 and GPT-4 scored 84% and 96% respectively, with GPT-4 outperforming GPT-3.5 (p = 0.046). CONCLUSIONS:Our results confirmed GPT-4's potential in medical education by providing highly accurate yet often complex explanations for NTSCI, which were successfully simplified without losing accuracy.
Addressing a significant gap in post-acute research and recognizing the high prevalence of depression (22.2
Since the publication of "What is the Current and Future Status of Digital Mental Health Interventions?" the exponential growth and widespread adoption of ChatGPT have underscored the importance of reassessing its utility in digital mental health interventions. This review critically examined the potential of ChatGPT, particularly focusing on its application within clinical psychology settings as the technology has continued evolving through 2023 and 2024. Alongside this, our literature review spanned US Medical Licensing Examination (USMLE) validations, assessments of the capacity to interpret human emotions, analyses concerning the identification of depression and its determinants at treatment initiation, and reported our findings. Our review evaluated the capabilities of GPT-3.5 and GPT-4.0 separately in clinical psychology settings, highlighting the potential of conversational AI to overcome traditional barriers such as stigma and accessibility in mental health treatment. Each model displayed different levels of proficiency, indicating a promising yet cautious pathway for integrating AI into mental health practices.
ChatGPT often "hallucinates" or misleads, underscoring the need for formal validation at the professional level for reliable use in nursing education. We evaluated two free chatbots (Google Gemini and GPT-3.5) and a commercial version (GPT-4) on 250 standardized questions from a simulated nursing licensure exam, which closely matches the content and complexity of the actual exam. Gemini achieved 73.2 percent (183/250), GPT-3.5 achieved 72 percent (180/250), and GPT-4 reached a notably higher performance with 92.4 percent (231/250). GPT-4 exhibited its highest error rate (13.3%) in the psychosocial integrity category.
Background: Chat GPT produces factual inaccuracies ("hallucinations"), outputs must be rigorously checked before use in nursing education, given limited validation across diverse cultural contexts. Aims: To evaluate GPT-3.5 and Google GEMINI on publicly available NCLEX-style nursing exam questions in the US and official EU general nursing exam questions for Spanish nationals. Methods: We used publicly available U.S. National Council Licensure Examination for Registered Nurses (NCLEX-RN) and for Practical Nurses (NCLEX-PN) style questions, and official Spanish general nursing exam questions (CONVALIDATE-EU-SPAIN). Results: Accuracy was the same for GPT-3.5 and GEMINI in NCLEX-PN (67.5%, 81/120), in NCLEX-RN was higher for GPT-3.5 (69.2%, 83/120) than for GEMINI (65.8%, 79/120). Regarding CONVALIDATE-EU-SPAIN accuracy was the same for both chatbots (76.7%, 92/120). By language, in English, GPT-3.5 performed slightly better (68.3%, 164/240) than GEMINI (66.7%, 160/240). In Spanish, both chatbots achieved the same accuracy (76.7%, 92/120). We identified specific NCLEX-PN concepts where both chatbots struggled (e.g., pregnancy). Conclusions: In the US, chatbots' accuracy was below 70%, and in Spain, below 80%, highlighting the need to assess them comprehensively across languages. (c) 2025 Organization for Associate Degree Nursing. Published by Elsevier Inc. All rights are reserved, including those for text and data mining, AI training, and similar technologies.
John Kelleher合作论文数School of Computing,
Dublin Institute of Technology,5