BACKGROUND:Patients with Limited English Proficiency (LEP) face poorer surgical outcomes, yet language-access capacity varies by hospital. OBJECTIVE:Compare 7-day readmission after common surgeries among adult patients with LEP at Language Serving Hospitals (LSH) and non-LSH, stratified by Spanish, Common Non-English, Non-Spanish (NENS) and Rare NENS Languages using the New Jersey State Inpatient Database. RESULTS:34,342 adult surgical patients were discharged from LSH with 2.5% readmitted compared to 5.0% of those from non-LSH. Patients from LSH who spoke Spanish, (aOR 0.46, 95% CI 0.40-0.54), Common NENS (aOR 0.49, 95% CI 0.31-0.79), and Rare NENS Languages (aOR 0.50, 95% CI 0.40-0.63) had reduced odds of readmission compared to those from non-LSH. CONCLUSIONS:Surgical patients with LEP discharged from LSH had lower odds of readmission, suggesting LSH may be better equipped than non-LSH. Worse disparities for Common and Rare NENS Languages suggest the need to expand resources beyond Spanish.
OBJECTIVE:This study aimed to evaluate the location-specific and time-sensitive trajectories of pressure injuries (PrIs) stages using real-world electronic health record (EHR) datasets. APPROACH:Using a dataset of 29,475 patients with records of PrIs documented from 2015 to 2023, we developed four PrI patient sub-cohorts with common PrI locations, including coccyx, buttocks, sacrum and heel. We estimated transition intensities between three PrI states: stage 1, stage 2, and a severe stage in each group. Stages and transition paths were derived from domain knowledge provided by clinical experts and The National PrI Advisory Panel (NPIAP) guidelines. RESULTS:The trajectory analysis suggested that stage 2 serves as a "gateway state" in all four locations, meaning that once a PrI reaches stage 2, the likelihood of transiting to severe stages increases significantly. The commonly used Braden Scale and its sub-components are more likely to be associated with transitions from stage 2 to severe stages, suggesting that manual risk assessment tools are suboptimal for predicting early-stage PrI transitions. Further, we observed race-dependent variations across injury location groups. INNOVATION:To our knowledge, this is the first study to introduce multi-state trajectory analysis in PrI research. Our model can investigate PrI status in a dynamic manner, which fills an important gap in the field. CONCLUSION:Our findings underscore the lack of time-sensitive information in existing PrI risk assessment tools, revealing a critical gap in their ability to capture the dynamic nature of PrI progression. Clinical decision support using time sensitive data is needed for delivering personalized, timely, and effective PrI prevention.
Abstract Background Adverse drug events (ADEs) are a critical indicator of patient safety but are often documented only in free-text clinical notes. The potential of recent advances in natural language processing (NLP), particularly generative large language models (LLMs), to identify ADEs remains understudied. This study aimed to compare the performance of multiple LLMs in identifying ADE-Drug relationships in inpatient and ambulatory clinical notes. Methods We used clinical notes from the 2018 National NLP Clinical Challenge (n2c2) ADE dataset (inpatient; n=505) and from outpatient encounters (n=2,555) between October 1, 2018, and December 31, 2019, at a large academic medical center based in New England. Notes were pre-processed into snippets for model input. Evaluated Models included: GPT-4o, GPT-4o-mini, LLAMA 3.3-70B and their instruction fine-tuned variants (including low-rank adapters for LLAMA). Performance was assessed using both strict and relaxed evaluations (precision, recall, and F1) for all models, followed by manual evaluation (exact semantic match, partial match, missing ADE, drug mention only, not a drug, or wrong) of the two best-performing models. Results GPT-4o and GPT-4o-mini were the top-performing models among those evaluated. GPT-4o consistently outperformed GPT-4o-mini in ADE extraction across both datasets, with higher F1-scores (0.524 vs. 0.381) and a more balanced precision-recall profile. Both models captured ADEs effectively in explicit and complex clinical contexts, although limitations included misclassification of pre-existing allergies and occasional conflation of therapeutic indications with adverse effects. GPT-4o achieved higher exact match coverage and fewer errors across clinical notes, indicating more reliable performance in both inpatient and ambulatory settings. Conclusion This work establishes a foundation for integrating LLM methods into real-world drug safety surveillance, with direct implications for improving patient safety.
Artificial intelligence (AI) is increasingly utilized in healthcare, including in language access services, but certain aspects remain understudied. We offer a research agenda to guide the development of evidence on how AI language access services are perceived by patients and how they impact trust and comprehension in clinical encounters, and to inform implementation strategies. We recommend a governance system to mitigate potential harm and capitalize on benefits for patients with a non-English language preference.
This environmental scan aimed to inform specifications of an electronic clinical quality measure for timely diagnostic follow-up after inconclusive/abnormal screening mammograms through a review of current breast cancer screening practice standards. PubMed, UpToDate, and clinical quality measure databases were searched for relevant studies, guidelines, and measures, respectively. Data abstracted included screening eligibility criteria, screening and diagnostic modalities, impact of diagnostic delays, definitions for timely follow-up and performance benchmarks, follow-up rates, factors associated with missed or delayed follow-up, and specifications of informatics tools to measure follow-up. The results were narratively summarized. Twelve guidelines, 23 peer-reviewed articles, and five measures were included. Most guidelines recommended routine screening for average-risk females aged ≥ 40 years using 2D or 3D mammography. Recommended follow-up modalities were additional imaging (diagnostic mammography, breast ultrasound, breast magnetic resonance imaging) and breast biopsy. Five studies and one systematic review reported associations between longer wait times to follow-up and increased tumor sizes and lymph node metastases. Three studies defined timely follow-up as ≤ 60 days after an inconclusive/abnormal mammogram. Five studies and five measures used administrative health/electronic health record data to identify clinically eligible patients, primarily excluding patients requiring individualized care. Diagnostic follow-up rates ranged from 63.6% to 96.8% within 3 months across seven studies. Non-white patients, non-English speakers, and patients with lower educational attainment, lower socioeconomic status, and/or residence in rural/suburban settings had a higher risk of missed or delayed follow-up. A standard definition for timeliness of diagnostic follow-up after inconclusive/abnormal mammographic screening is needed to inform meaningful measurement and targeted interventions that improve clinical outcomes in the breast cancer screening process.
There is an urgent need for scalable strategies for treatment of overweight and obesity that can be implemented in clinical settings. To implement and evaluate an online weight management program in a large, diverse population of patients. Clinical implementation project in primary care and specialty clinics. Eligible patients were ≥ 20 years old, spoke English or Spanish, and had a body mass index (BMI) of ≥ 30 kg/m2 or a BMI of 25–29.9 kg/m2 plus ≥ 1 cardiovascular risk factor or obesity-related condition. Enrolled patients were encouraged to register for a 12-month digital health program called RestoreHealth, which included an online program/app and coaching. We examined recruitment and enrollment in the program, as well as engagement, absolute and percent weight change, and predictors of weight change during the first six months. Subgroup analyses were conducted by use of anti-obesity medications. A total of 5056 patients enrolled between November 2022 and October 2023, and 4511 (89.2
Delays in the completion of two-step colorectal cancer (CRC) screening, consisting of follow-up colonoscopy after positive non-invasive testing, increase the risk of advanced-stage CRC diagnoses and poor health outcomes. This environmental scan aimed to inform the measurement of timely follow-up colonoscopy after positive stool tests. Two main data sources, UpToDate and PubMed, were used to identify society screening guidelines and peer-reviewed literature, respectively. Four websites and two databases were searched for related clinical quality measures. Additional society guidelines and literature were identified by handsearching. All data sources were used to find information on the types of stool tests and follow-up procedures used in two-step CRC screening, target populations for stool tests, follow-up rate benchmarks, and definitions of timely follow-up. Data were extracted and narratively synthesized. Eighteen guidelines, 32 peer-reviewed articles, and eight related clinical quality measures were included. Recent guidelines recommended asymptomatic, average-risk individuals aged 45–75 years undergo routine CRC screening. Modalities for stool tests were the fecal immunochemical test, high-sensitivity guaiac-based fecal occult blood test, and the multi-target stool DNA panel. Colonoscopy was the gold standard follow-up procedure for positive stool tests. Two guidelines proposed benchmarks of ≥ 80
Background:Over the last decade, fentanyl use in the U.S. has experienced a dramatic shift, largely driven by a rise in illicit fentanyl and its role in the opioid overdose crisis. Characterizing individual-level illicit fentanyl exposure poses significant challenges, as such use often occurs outside clinical settings and is not directly captured in clinical data, making it difficult to measure its true scale and impact. Methods:We conducted a retrospective cohort study using Epic Cosmos, a large U.S. electronic health record dataset comprising over 300 million patients across inpatient and outpatient settings nationwide (January 1, 2015, to December 31, 2023). The dataset provides individual-level clinical data, including diagnoses, medication records, and laboratory testing data, enabling longitudinal characterization of fentanyl exposure. To infer sources of fentanyl exposure, we linked urine drug testing (UDT) results to fentanyl medication records using a sequential time window screening method (e.g., UDT positive with pre-fentanyl records or without). Temporal and Cox proportional hazards analyses were used to examine longitudinal patterns and risks of opioid-related harmful outcomes associated with medical-source versus illicit-source fentanyl exposure. Regional variation was explored. Confounding was adjusted using stabilized inverse probability weighting based on demographics, social vulnerability index, and 37 baseline conditions. Subgroup analyses tested effects of underlying clinical burden. Sensitivity analyses evaluated alternative UDT screening windows and follow-up periods to assess sensitivity to exposure and outcome definitions. Findings:Our study cohort included 295,728 patients, consisting of 85,535 (28.9%) in the medical-source cohort (MSC) and 210,193 (71.1%) in the illicit-source cohort (ISC). From 2015 to 2023, the prevalence of nonfatal opioid overdose was consistently higher in the ISC. Overdose prevalence in the ISC increased markedly over time, reaching 18.6%, whereas only a modest increase was observed in the MSC, reaching 4.3%. For opioid dependence and abuse, the ISC had higher prevalence, and steeper year-over-year increases, versus the MSC until 2020. After 2020, prevalence in the ISC declined, particularly for dependence, but remained consistently higher than in the MSC throughout the study period. Region-stratified temporal patterns followed the overall cohort-level patterns. Illicit-source fentanyl initiation was associated with significantly elevated risk of 30-day opioid-related harmful outcomes, with adjusted hazard ratios of 2.99 (95% CI, 2.71-3.29; P < 0.001) for overdose, 1.96 (95% CI, 1.84-2.10; P < 0.001) for abuse, and 3.05 (95% CI, 2.89-3.22; P < 0.001) for dependence, compared with medical-source initiation, after confounding adjustment. Across subgroups stratified by clinical conditions, illicit-source initiation was consistently associated with increased risk of outcomes. Interpretation:Illicit-source fentanyl exposure is associated with a markedly higher risk of opioid-related harmful outcomes than medical-source exposure, providing evidence that illicit fentanyl is a driver of adverse patient outcomes. Whilst acknowledging the limitations of this analysis, these findings underscore the substantial contribution of illicit fentanyl to the ongoing opioid epidemic and highlight the need for deeper investigation of how medical and illicit fentanyl use interact over time. Further research is warranted. Funding:National Institute on Drug Abuse (NIDA) and National Science Foundation (NSF).
In this Viewpoint, we advocate for direct tokenisation of medical data by breaking them into discrete units, such as laboratory results, medications, and vital signs, similar to word tokenisation in language models. This approach enables transformer-based models to learn from the temporal structure of patient health timelines without relying on textual translation, potentially leading to more accurate and personalised care. Enhanced Transformer for Health Outcome Simulation, an example of a model that uses tokenisation, forecasts health timelines and supports clinical decision making using tokenised medical records. We outline a privacy-preserving model-sharing framework, in which models are trained locally and only trained models—not sensitive data—are shared, allowing collaborative development across institutions. We also emphasise that access to large, diverse datasets enhances fairness, generalisability, and equity in health-care generative artificial intelligence. Although challenges such as data complexity and interpretability remain, this Viewpoint underscores that embracing tokenised representations opens a path towards scalable, multimodal, and equitable artificial intelligence in medicine.
Diagnostic errors represent a major cause of patient harm. One way to reduce diagnostic errors is to support learning health systems with standardized event review practices and data integration across event types and health systems. The extent to which this occurs and the barriers to doing so remain poorly characterized. Characterize event review practices and analyze the presence and distribution of key patient safety learning system constructs, diagnostic error analysis capabilities, and implementation barriers across a health system. We developed a survey of participant safety event review practices, associated learning health system activities, diagnostic error analysis capabilities, data integration, and implementation barriers. The survey was electronically distributed to a purposive sample of health system employees who routinely perform safety event reviews at one US academic medical center. Descriptive statistics were reported for survey constructs. Chi-square analysis was used to assess for non-response bias. One hundred and six of 249 possible respondents (42.6
BACKGROUND AND OBJECTIVES:Electronic health records (EHRs) contain valuable information for research and decision-making, but much resides in unstructured notes that are challenging to analyze at scale. We developed SPELL (Snippet-Primed rEgex LLM Pipeline), a scalable natural language processing workflow that combines regular-expression-based snippet retrieval with locally hosted large language model (LLM) inference to extract structured variables from large collections of clinical narratives. METHODS:SPELL uses task-specific regular expressions to retrieve short context windows ("snippets") from unstructured texts and applies task-prompted LLM inference on snippets rather than full documents. All processing occurs within institutional computing environments. Accuracy was evaluated on randomly sampled, clinician-annotated benchmark sets of 50 documents per obstetric task, with separate retrieval-recall audits of 20 regex-negative documents per task. We evaluated accuracy and efficiency across three obstetric information-extraction tasks: numerical value (blood loss volume), date (estimated due date), and diagnosis (hemolysis, elevated liver enzymes, and low platelets [HELLP] syndrome). We quantified computational scalability using elapsed time, out-of-memory events, energy consumed, and GPU telemetry, and audited retrieval recall using clinician-annotated regex-negative notes enriched with relevant structured metadata. Generalizability was assessed on the public MT Samples corpus (5013 notes across 40 specialties) for ventricular tachycardia detection. RESULTS:SPELL processed 31 million clinical notes spanning 1976-2024 from eight hospitals. Snippet-based inference reduced processing time by 71-87% versus full-document LLM inference and by >95% versus manual physician annotation. On the 50-document benchmark sets, snippet-based evaluation achieved 98% exact-match accuracy for blood-loss extraction, 92% exact-match accuracy for estimated-due-date extraction, and 94% accuracy with an F1-score of 0.97 for HELLP syndrome classification. As an exploratory external evaluation on MT Samples, ventricular tachycardia detection achieved 84% accuracy and an F1-score of 0.67. CONCLUSIONS:A hybrid regex-snippet-LLM pipeline can enable accurate and computationally efficient extraction from unstructured EHR narratives.
BACKGROUND:Little is known about large language model (LLM) performance on palliative care (PC)-related knowledge-based tasks. We evaluated two LLMs in answering PC-related test questions and explaining their answer choice rationale. METHODS:LLMs were prompted to answer 25 randomly selected questions from the Fast Facts Quiz and provide their answer choice rationale. Three PC educators ranked and rated LLM-generated answer choice explanations versus the test's answer key explanations. Linear fixed-effect models evaluated reviewer ranking, and ordinal logistic regression evaluated reviewer ratings of quality, suitability, accuracy, relevance, and comprehensiveness. RESULTS:Both LLMs answered 96% of selected questions correctly. Reviewers rated LLM-generated explanations more highly than Fast Facts Quiz explanations. Five themes emerged from reviewer comments: perceived inaccuracies, clarity of writing, educational value, linguistic style, and miscellaneous. CONCLUSIONS:LLMs demonstrated high answer choice accuracy and generated preferable answer explanations when compared to the Fast Facts Quiz answer key.
BACKGROUND:The use of generative large language models (LLMs) with electronic health record (EHR) data is rapidly expanding to support clinical and research tasks. This systematic review characterizes the clinical fields and use cases that have been studied and evaluated to date. METHODS:We followed the Preferred Reporting Items for Systematic Review and Meta-Analyses guidelines to conduct a systematic review of articles from PubMed and Web of Science published between January 1, 2023, and November 9, 2024. Studies were included if they used generative LLMs to analyze real-world EHR data and reported quantitative performance evaluations. Through data extraction, we identified clinical specialties and tasks for each included article, and summarized evaluation methods. RESULTS:Of the 18 735 articles retrieved, 196 met our criteria. Most studies focused on radiology (26.0%), oncology (10.7%), and emergency medicine (6.6%). Regarding clinical tasks, clinical decision support made up the largest proportion of studies (62.2%), while summarizations and patient communications made up the smallest, at 5.6% and 5.1%, respectively. In addition, GPT-4 and GPT-3.5 were the most commonly used generative LLMs, appearing in 60.2% and 57.7% of studies, respectively. Across these studies, we identified 22 unique non-NLP metrics and 35 unique NLP metrics. While NLP metrics offer greater scalability, none demonstrated a strong correlation with gold-standard human evaluations. CONCLUSION:Our findings highlight the need to evaluate generative LLMs on EHR data across a broader range of clinical specialties and tasks, as well as the urgent need for standardized, scalable, and clinically meaningful evaluation frameworks.
Introduction:The eligibility of anti-amyloid disease-modifying therapies (DMTs) and their integration into clinical practice in some institutions requires a specific range of Mini-Mental State Examination (MMSE) scores. Reliance on this pencil-and-paper psychometric instrument imposes operational burdens and risks of perpetuating health disparities, given the test's known educational and cultural biases. This study evaluates the efficacy of the Digital Clock and Recall (DCR™)-a rapid, FDA-listed digital cognitive assessment-to crosswalk to MMSE scores using machine learning, thereby offering a faster, scalable, and equitable mechanism for patient triage. Methods:We conducted a retrospective analysis using data from the multi-site Bio-Hermes-001 (BH) study (NCT04733989, N = 945). Participants were clinically classified as cognitively unimpaired, mild cognitive impairment, or probable Alzheimer's dementia. We trained a Poisson elastic net regression model on 70% of the sample, using age and multimodal digital features derived from the DCR (including drawing kinematics and voice acoustics) to predict MMSE scores. The model was validated using the remaining 30% of Bio-Hermes-001 and an independent external validation cohort from the Apheleia study (NCT05364307, N = 238). Results:The machine learning model predicted MMSE scores with a root-mean-squared error (RMSE) of 2.43 in the BH test set. This error margin falls within the established test-retest reliability range of the manual MMSE itself (∼4.0-4.2 points at short inter-test intervals), providing evidence that the predicted score is of comparable precision to a repeat human administration of the MMSE. External validation in the Apheleia cohort demonstrated robust generalizability (RMSE = 2.62). In the BH held-out test set, the model showed comparable performance across Race (White RMSE = 2.46; Non-White RMSE = 2.25) and Ethnicity (Hispanic RMSE = 2.19; Non-Hispanic RMSE = 2.45), a balanced pattern also observed in the Apheleia-001 external cohort. Exploratory demographic analyses on prediction errors, including Age, Sex, Race, and Ethnicity, yielded significant differences only for Sex and Age in Apheleia, with signed errors becoming progressively more negative (i.e., increasing under-prediction) at older ages for the latter. This scarcity of statistical differences across cohorts suggested that our predictions were fair. Discussion:Machine learning can leverage multimodal features from the DCR to accurately and equitably crosswalk to MMSE scores in support of current guidelines, transforming a time-intensive manual test into a rapid, automated assessment. By deploying this "digital triage" engine, where traditional assessments are still used for DMT eligibility, healthcare systems can streamline the identification of DMT-eligible patients, reduce specialist referral bottlenecks, and ensure that access to life-altering therapies is determined by pathology rather than demography.
Clinical diagnostic evaluation should not only assess whether models can provide correct diagnoses, but also reflect the realities of clinical practice, including progressive disclosure of multimodal information, dynamic updating of diagnostic hypotheses, and continuous refinement of clinical reasoning. However, existing evaluations of multimodal large language models (MLLMs) typically rely on single-turn or isolated tasks, making it difficult to fully capture the complexity of real-world clinical diagnosis. To bridge this gap, we developed ClinMM-Bench, the largest multi-turn multimodal clinical diagnostic evaluation benchmark to date. ClinMM-Bench contains 1,089 challenging real-world clinical cases and 3,760 medical images across eight specialties. We systematically evaluated 15 representative MLLMs using a two-level evaluation framework that assessed both diagnostic accuracy and diagnostic reasoning quality. Results showed that proprietary models achieved the highest overall diagnostic accuracy, but the proportion of completely correct diagnoses remained limited across all models. In terms of diagnostic reasoning quality, current models can identify plausible diagnostic directions but still have considerable limitations in generating reliable diagnostic reasoning. Error analysis further identified five representative failure modes: information synthesis failure, knowledge mapping error, perception error, premature closure, and visual hallucination.
OBJECTIVE:This initiative aimed to identify and classify contributory factors to diagnostic errors in abdominal imaging using an annotation process to identify, quantify, and categorize process-related errors. METHODS:This was a retrospective, quality improvement, institutional review board-approved study conducted at an academic health system on 200 randomly selected adult patients who completed abdominal diagnostic imaging examinations between January 1, 2022, and December 31, 2022, in which a radiologist recommended additional imaging. We used an expert panel to develop a taxonomy based on two previously described taxonomies for process-related errors. We then applied an iterative process to train annotators to apply the taxonomy and assess medical records for diagnostic errors. We assessed the proportion of patients with diagnostic errors, as well as interrater reliability in classifying contributory factors. We also measured percentage agreement in identifying diagnostic errors using an annotation taxonomy versus a gold standard developed by two abdominal radiologists. RESULTS:The final sample included 185 patients. A total of 14.6% (27 of 185; 95% confidence interval 10.2%-20.4%) patients undergoing abdominal imaging experienced at least one diagnostic error. Interannotator agreement with the gold standard increased from 61.1% (33 of 54) at baseline for 5 reviewers to 80% (40 of 50) (χ2P = .04) after four rounds of annotation. "Delay in performing ordered test(s)" is the major contributory factor to diagnostic errors. DISCUSSION:Process-related diagnostic errors occurred in almost one in seven patients undergoing abdominal imaging with recommended additional imaging. After iterative training, an annotation process for identifying patients who have imaging diagnostic errors achieved significant interannotator agreement compared with a gold standard of radiology specialist review.
Medicare beneficiaries frequently visit the emergency department (ED) at the end of life, but little is known about the epidemiology of patients admitted to hospice from the ED. We used 100 percent Medicare fee-for-service claims from the period 2018-20 to describe the frequency of direct ED-to-hospice enrollments, associated patient and hospice agency characteristics, and patient outcomes. In this sample, 4.3 percent of initial enrollments in hospice originated from the ED. ED-to-hospice admissions featured short lengths-of-stay (21.7 percent were two days or less) and high rates of general inpatient level of care at the time of enrollment (23.6 percent). The 10 percent of hospice agencies with the highest proportion of ED-to-hospice enrollments were less often for-profit than agencies ranked below the fiftieth percentile in respect to proportion of ED-to-hospice enrollments. Further research is needed to increase understanding of how much patients benefit from ED-to-hospice transfers when their hospice stays before death are very short, and what drivers lead to these ED-to-hospice transfers.
Stigmatizing language in clinical documentation, which conveys negative stereotypes, attitudes, or judgments toward patients, is a recognized source of documentation bias and is associated with poorer care and adverse health outcomes. Although prior stigma-related research has explored on clinician-written EHR notes, the increasing use of large language model (LLM)-generated documentation in clinical workflows raises new concerns about its potential to produce or amplify bias and affect patient safety. In this study, we conducted a large-scale assessment of stigmatizing language in LLM-generated reasoning text on 35 real-world clinical tasks across 107 LLMs. We applied a psychiatrist-validated, natural language processing (NLP) system to detect stigma terms in LLM reasoning text and quantified stigma rates of LLM-generated reasoning texts across 3,745 model-task pairs. Results showed that stigma rates ranged from 0% to 33.33%, with 84.06% of pairs containing stigma terms. Reasoning models showed higher stigma rates than non-reasoning models (2.35% vs. 1.70%; p < 0.0001), whereas medical models did not show significantly lower stigma rates than general-purpose models (1.80% vs. 2.00%; p = 0.26). Stigma rates of LLM outputs correlated negatively with task accuracy (r = -0.283; p < 0.001) and positively with input clinical-text stigma (r = 0.569; p < 0.001), with 19.76% of LLM-task pairs amplifying stigma in the original input notes. The effectiveness of prompt engineering as a destigmatizing approach varied across models, with stigma rates reduced by up to 91.91% without compromising model performance. This study shows that stigmatizing language generation is common but modifiable in LLM-generated reasoning traces, highlighting the need for direct stigma-related safety evaluation and robust destigmatization strategies before clinical deployment.