Background Medication related morbidity due to inappropriate prescribing, delays in appropriate treatment, and adverse drug events are major contributors to mortality in acutely hospitalized adults. Comprehensive medication management (CMM) is a standard for medication therapy care provided by pharmacists in collaboration with the interprofessional rounding team. Optimization of CMM via appropriate pharmacist staffing practices may reduce mortality. Methods Adults admitted to an intensive care unit (ICU) without restrictions to care for at least 24 hours from 64 centers were enrolled in a prospective, observational trial that collected patient level and healthcare team staffing data from August 2023 to August 2024. The primary exposure was patient-level, pharmacist staffing measured by the pharmacist to patient ratio averaged over each patient s ICU stay. The primary outcome was hospital mortality assessed by multi level logistic regression that considered 15 variables. Results A total of 213 pharmacists provided ICU based CMM care to 28,795 patients at a median (interquartile range) pharmacist to patient ratio of 1:17 (13-23). The hospital mortality rate was 14.7%. Patients without pharmacist CMM for 1 day had an increased risk of mortality of 17.9% (95% Confidence Interval (CI) 1.061 1.311, p=0.002). For every one patient increase in the pharmacist to patient ratio, the odds of mortality increased by 0.8% (Odds Ratio (OR) 1.008, 95% CI 1.003 1.012, p<0.001). Approximately 42 patients need to receive daily CMM at a pharmacist to patient ratio </=1:15 to prevent one additional hospital death. Conclusions CMM delivered by pharmacists in collaboration with the interprofessional team is associated with improved hospital survival, particularly when pharmacists care for 15 or fewer ICU patients. ### Competing Interest Statement Authors with conflicts of interest are listed below. If an author is not listed they reported no conflicts of interest. Marisha Burden- Dr. Burden reports funding from the Agency for Healthcare Research and Quality, the National Institute for Occupational Health and Safety, University of Colorado Innovations digiSPARK award, Med IQ, and the American Medical Association not related to this work. Dr. Burden contributed to the development of GrittyWork, a digital workforce application, and a registered trademark of the University of Colorado not related to this work. Ashley Hawthorne- Speakers Bureau for Vericel Corporation ### Clinical Protocols [https://journals.lww.com/ccejournal/fulltext/2023/09000/optimizing\_pharmacist\_team\_integration\_for_icu.3.aspx][1] ### Funding Statement Funding through the Agency for Healthcare Research and Quality for Drs. Sikora and Smith was provided through R21HS028485 and R01HS029009. Funding through the University of Maryland, Baltimore, Institute for Clinical & Translational Research voucher program was provided to Dr. Heavner. Funding through the ASHP Research and Education Foundation was provided through a research grant to Dr. Heavner. Funding through the Board of Pharmacy Specialties was provided through a research grant to Dr. Smith. Funding through the American College of Clinical Pharmacy Critical Care Practice & Research Network was provided through research grants to Drs. Sikora, Henry, and Murray. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: The protocol was reviewed and approved by the Institutional Review Board at the University of Georgia (PROJECT00007120, April 11, 2023). A waiver of informed consent was granted for ICU patient enrollment, and a partial waiver of consent was granted for pharmacist enrollment. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All data produced in the present study are available upon reasonable request to the authors [1]: https://journals.lww.com/ccejournal/fulltext/2023/09000/optimizing_pharmacist_team_integration_for_icu.3.aspx
Background:Electrolyte replacement is ubiquitous in the acute care setting, but its familiarity cannot belie that even small dosing errors with potassium can cause lethal cardiac arrhythmias. Recently, MedAgentBench offered a benchmark for agentic artificial intelligence (AI) including the ability to correctly dose potassium based on a single rule; however, this does not adequately reflect the clinical complexity or safety concerns of an agent that has been used as the lethal injection. The purpose of this analysis was to a probe leaderboard large language model (LLM) capabilities to follow basic dosing rules to safely replace potassium in a series of clinician-annotated cases. Methods:Using a clinician panel, we developed a series of dosing principles and 20 clinical cases reflective of the complexity of potassium replacement. External clinicians were surveyed to assess practice variability and agreement to clinician panel answers. We tested GPT-5-chat with each case in triplicate, with and without the clinician curated dosing principles, and prompted the model to answer six questions involving potassium goals, dosing, route, lab frequency, concurrent interventions, and the model's perceived level of confidence for the output and complexity of the case. The primary outcome was the rate of appropriate recommendations in comparison to clinician answers. Results:A total of 54 clinicians reviewed the 20 hypokalemia cases and hypokalemia dosing guideline. Clinicians expressed "highly agree" or "somewhat agree" for 66.8% of the cases evaluated when asked if they agree with the guideline-recommended management. When given the potassium dosing guideline, total errors dropped from 165 to 104, and average accuracy improved from 45% to 65% with GPT-5-Chat. GPT-5-Chat conveyed a high level of confidence for 100% of responses, while labeling 80% and 76% of cases as highly complex with and without the criteria, respectively. Potential harm scores were considerable in both groups, however, a notable reduction in severity scores occurred with the dosing guidance document. Recommendations on concurrent interventions and dosing had the highest rate of errors in both groups. Conclusions:Benchmarks must appropriately reflect clinical complexity to be considered valuable for the deployment of agentic artificial intelligence tools in the healthcare domain. GPT-5-Chat assessment on a comprehensive medication management task for potassium replacement showed improvement with dosing guidance, yet unfit benchmarking performance.
Background:The accuracy and safety of generating medication orders by large language models (LLMs) must be demonstrated. Without standardization, performance evaluation is limited to time and resource-intensive clinician grading. This evaluation aimed to develop a standardized medication format that supports automated performance evaluation (MedMatch). Methods:First, a survey of 40 medication prompts was given to clinicians to assess agreement in medication order communication. Second, a clinician panel developed a standardized medication format (MedMatch) for oral and intravenous medications. Third, a clinician-annotated dataset of medication prompts and standardized answers in the MedMatch format was developed for LLM testing. Finally, LLMs were retested with the same dataset, adjusted to exclude route information, to evaluate the appropriate categorization of medication route. Results:The formal medication orders consistently showed low omission rates and high overlap for all entities, compared to the verbal and brief written communication types. Lexical overlap results demonstrated pattern norms amongst clinicians with entities appearing most commonly in positions 1-5 in the order of drug name, dose, unit, route, and frequency. In the second survey, the formal written group performed the highest with 78.3% of prompts considered appropriate as a computer-generated response. LLM accuracy on MedMatch order standardization was highest in oral solid (64.2-72.5%), intravenous intermittent (72.5-84.3%), and intravenous push (62.7-74.5%) categories. LLMs performed the worst at categorizing medication orders accurately into intravenous push (18-61%) and intravenous intermittent (51-100%) routes. Conclusions:Standardized format for computer-based outputs may support automated performance analysis and enhance the clarity of medication communication.
BACKGROUND:The 2019 medication regimen complexity-intensive care unit (MRC-ICU) score is associated with patient outcomes, ICU complications, and critical care pharmacist workload. This score was developed using heuristic component selection and validated in a single-center cohort of 130 ICU patients. We sought to apply data-driven reweighting methodology in a large, multicenter cohort of ICU adults to improve the predictive capabilities of MRC-ICU. METHODS:This was a retrospective, observational cohort study of adults admitted to an ICU between 2015 and 2023 at two academic health systems. Machine learning-based methods, including Principal Component Analysis and Random Forest, were used to create an updated MRC-ICU score optimized to predict three outcomes: hospital mortality, ICU fluid overload (FO) occurrence, and invasive mechanical ventilation (IMV) use. MRC-ICU 2.1 used average mortality, FO, and IMV use; MRC-ICU 2.2 used average mortality and FO and adjusted for prolonged IMV use. Data from one center were used for training and testing, and data from the other for validation. The predictive abilities of MRC-ICU 2.1 and 2.2 for each outcome were compared to MRC-ICU 1.0 and to severity of illness scores (i.e., Acute Physiology and Chronic Health Evaluation [APACHE] II and Sequential Organ Failure Assessment [SOFA]). RESULTS:A total of 19,117 patients across training, testing, and validation datasets were included. MRC-ICU 2.0 scores outperformed MRC-ICU 1.0 for predicting most outcomes, with improvements in Area Under the Receiver Operating Characteristic (AUROC) ranging from +0.03 to +0.08 across datasets. MRC-ICU 2.1 and 2.2 did not consistently outperform APACHE II and SOFA in predicting mortality. The addition of MRC-ICU 2.0 scores to models including APACHE II or SOFA resulted in statistically significant improvements in discrimination in several settings (DeLong p < 0.05), with AUROC increases generally ranging from approximately +0.01 to +0.13 depending on outcome and dataset. CONCLUSIONS:The updated MRC-ICU 2.0 score (MRC-ICU 2.1 and 2.2) demonstrated consistently improved discrimination compared with the original MRC-ICU 1.0 across outcomes and datasets. The performance of MRC-ICU 2.0 (MRC-ICU 2.1 and 2.2) was generally comparable to established severity-of-illness scores (SOFA and APACHE II), although it did not consistently outperform these measures. When incorporated into combined models, MRC-ICU 2.0 provided additional predictive value, indicating that it captures information complementary to traditional severity-of-illness scores. Overall, these findings suggest that MRC-ICU 2.0 represents an improved and clinically interpretable measure of medication regimen complexity that is useful as a complementary predictor.
INTRODUCTION:Prediction algorithms for prolonged mechanical ventilation (PMV) in the intensive care unit (ICU) have rarely incorporated detailed medication data, despite medications being important causal contributors to patient outcomes. The purpose of this study was to develop and validate PMV prediction models to assess the contribution of medication-related variables alongside established physiologic predictors. METHODS:In this retrospective cohort study, models were developed using data from a random sample of 318 adults admitted to ICUs within the University of North Carolina (UNC) health system who received mechanical ventilation for ≥ 24 h from October 2015 to October 2020. Validation was performed in two datasets: a temporally distinct cohort from UNC from June 2021 to June 2023, and a cohort from Oregon Health Sciences University from June 2020 to June 2023. Logistic regression and supervised, classification-based machine learning (ML) models [XGBoost, Random Forest, Support Vector Machine (SVM)] were trained on 30 demographic, clinical, laboratory, and medication-related variables. The primary outcome was area under the receiver operating characteristic (AUROC) of developed prediction models for the occurrence of PMV. RESULTS:The base logistic regression model with medication regimen complexity and severity of illness data added was the best-performing regression model, achieving an AUROC of 0.75. Random Forest and SVM ML models achieved AUROCs of 0.78. Model discrimination decreased modestly in external validation. Explainability analyses of ML models expectedly included severity of illness scores and respiratory indices among the most important features, but also consistently included the medication regimen complexity-intensive care unit (MRC-ICU) score and other medication metrics. Incorporation of medication data yielded modest improvements in overall discrimination and negative predictive value. CONCLUSIONS:Medication-related variables contributed incremental value to PMV prediction. ML methods provided marginal improvements over regression models. These findings highlight the potential value of medication data in prediction modeling for patient outcomes but emphasize the need to contextualize the value of complex models over simpler alternatives.
Heart rate variability (HRV) reflects autonomic nervous system function and has emerged as a potential noninvasive biomarker for early detection of physiologic deterioration in critical illness. HRV-based prediction models show promise; however, translation into routine ICU practice has been limited. A major barrier is the insufficient characterization of medication effects on HRV. Pharmacologic agents commonly used in critical care, including vasopressors, steroids, and antiarrhythmics, can directly or indirectly alter autonomic tone, yet existing studies rarely account for these influences. As a result, medication-induced HRV changes may represent meaningful therapeutic response or misleading confounding noise, complicating interpretation. Current studies do not adequately account for medication exposure when evaluating HRV in critical illness. We outline research priorities focused on quantifying medication effects, integrating medication exposure into predictive modeling, evaluating HRV as a marker of treatment response, and determining the utility of HRV as a treatment target.
Background:The use of large language models (LLMs) is increasing in the medical field; however, LLMs are often subject to "confabulations." Notably, LLMs have vulnerability to adversarial attacks, or fabricated details within prompts, which is concerning given both health misinformation and inadvertent errors in the medical record. This purpose of this study was to determine the effect of adversarial attacks by embedding one fabricated medication into a list of existing medicines. Methods:A total of 250 cases were created, which included 4-6 medications and one fabricated medication (a Pokémon character). Four LLMs (GPT-4o-mini, Gemma-3-27B-IT, Llama-3.3-70B-Instruct, and Qwen3-32B) were tested in triplicate for both dosing information and disease indication with a default prompt, mitigation prompt, and the default prompt with a temperature of 0. If the LLM responded as if the Pokémon were a real medication, it was deemed a confabulation. The primary outcome was the rate of confabulations; exact paired-permutation tests were used to evaluate differences among LLMs and prompting approaches. Results:Confabulation rates for the default and temperature 0 drug dosing prompt ranged from 86-98.8% and 86.9-98.8% across models, respectively, and from 42-95.6% and 41.6-95.5% for the indication prompt. Incorporating the mitigation prompt substantially reduced confabulation rates to 8.3-76.3% (dosing) and 1.7-28.3% (indication). The best-performing model, Llama-3.3-70B-Instruct, demonstrated confabulation rates spanning 1.7-91.9% (p<0.001). Conclusions:LLMs are susceptible to adversarial attacks, especially with medications. Further model improvement is imperative before LLMs are considered safe and reliable for routine use in the medical field.
BACKGROUND:Although numerous research studies have demonstrated the positive impact of clinical pharmacy services, these benefits do not translate into sustained practice changes without support from hospital pharmacy leaders. Factors influencing leadership decisions to expand pharmacy services remain unclear. This study aimed to identify barriers to implementing pharmacy practice model changes and gain insights into potential methods of overcoming these barriers from the hospital pharmacy leader perspective. METHODS:We conducted a cross-sectional survey of hospital pharmacy leaders distributed via email over 3 weeks between September and October 2025. The survey included questions about perceptions related to implementation of practice model changes, resources/evidence used to justify clinical positions, barriers to expanding clinical pharmacy services, and demographics of health care systems they represented. The survey included Likert scales and open-ended questions. The primary outcome was the types of evidence most compelling to justify clinical pharmacist positions. Secondary outcomes included resources currently in use for decision-making and perceived barriers. RESULTS:The survey was completed in full by 84 leaders and highlighted key factors influencing administrative decision-making regarding the expansion of clinical pharmacy services and revealed significant barriers to justifying clinical positions related to knowledge gaps. The types of evidence considered most compelling included data indicating pharmacist impact on budget and impact on patient outcomes, followed by data indicating impact on hospital workload. Resources currently used for decision-making were most frequently benchmarking, productivity, or workload metrics; physicians or other champions; and regulatory or accreditation requirements. CONCLUSION:This survey provided valuable insights into the perspectives of hospital pharmacy leaders on resources and evidence needed to support expanded pharmacy services and justify clinical pharmacist positions. These insights can inform future research by ensuring relevant clinical and administrative metrics are included in outcomes.
The American College of Clinical Pharmacy (ACCP) advocates for board certification of clinical pharmacists in one of the recognized specialties as a fundamental qualification and requirement for providing and supervising trainees in the provision of direct patient care. However, data linking pharmacist board certification to patient outcomes are limited, and many methodological challenges exist to effectively evaluate the impact of board-certified pharmacists (BCPs). The purpose of this ACCP white paper is to critically analyze study design characteristics in order to inform future studies assessing the impact of BCPs on patient outcomes. Value depends on the stakeholder perspective, which informs the study design. This white paper considers the feasibility and applicability of various study methods that may be used to evaluate the impact of BCPs, including overall study design, study group assignment, outcome measures, reporting of clinical practice activities, identification of confounding variables, and statistical analyses, including methods to control for confounders. The 2025 ACCP Research Affairs Committee was surveyed to rate the feasibility and applicability of studying the impact of BCPs on outcomes by stakeholder perspective. Studies from the health care system or provider perspectives were rated as providing the best balance between feasibility and applicability, whereas various study design characteristics (e.g., prospective study designs, use of “length of” or disease-related surrogate markers as primary outcomes, adjusting for specific confounders) were rated highly and are desirable for future studies.
PURPOSE:Large language models (LLMs) are promising artificial intelligence (AI) tools to support clinical decision-making. The ability of LLMs to evaluate medication regimens, identify drug-drug interactions (DDIs), and provide clinical recommendations has undergone limited evaluation. The purpose of this study was to compare the performance of 3 LLMs in recognizing DDIs, determining clinical relevance, and generating management recommendations. METHODS:A total of 15 patient cases with medication regimens were created; each contained a commonly encountered DDI. Two separate study phases were developed: (1) DDI identification and determination of clinical relevance; and (2) DDI identification and generation of a clinical recommendation. The primary outcome was the ability of the LLMs (GPT-4, Gemini 1.5, and Claude 3) to identify the DDI within each medication regimen. Secondary outcomes included the ability of the LLMs to identify the clinical relevance of each DDI and generate a recommendation of high quality relative to ground truth. RESULTS:Claude 3 identified all DDIs, followed by GPT-4 (14/15, 93.3%) and Gemini 1.5 (12/15, 80.0%). All LLMs were significantly more likely than clinical experts to categorize the DDI as clinically relevant (P < 0.01). DDI management recommendations provided by GPT-4 were rated as optimal in 8 of 13 (61.5%) of the cases (P = 0.05 for comparison to ground truth). Two recommendations from GPT-4 and one recommendation from Gemini 1.5 were deemed to result in potential patient harm. CONCLUSION:While LLMs demonstrate promising potential to identify DDIs, application to clinical cases requires ongoing development. Findings from this study may assist in future development and refinement of LLMs for clinical decision-making related to DDIs.
Background:Drug-drug interactions (DDIs) are a significant source of morbidity and adverse drug events (ADEs), particularly in situations of polypharmacy and complex medication regimens. While rules-based software integrated in electronic health records (EHRs) has demonstrated proficiency in identifying DDIs present in medication regimens, large language model (LLM) based identification requires thorough benchmarking and performance evaluation using high-quality datasets for safe use. The purpose of this study was to develop a series of performance benchmarking experiments specifically for LLM performance in identification and management of DDIs using a specifically curated clinician-annotated dataset of clinically-relevant DDIs. Methods:We evaluated three LLMs (GPT-4o-mini, MedGemma-27B, LLaMA3-70B) using a clinician-annotated benchmark dataset of 750 DDI scenarios spanning three levels of diagnostic complexity. Tasks were aligned with flexible judgment formats: (1) a pointwise two-drug classification task, (2) a pairwise three-drug discrimination task, and (3) a listwise 4-6 drug selection task. Standardized zero-shot prompting with task-specific instructions was applied for all models. Performance was assessed using precision, recall, F1 score, and accuracy. Reliability was quantified using self-consistency across repeated runs and confidence-aligned metrics to capture stability in model reasoning. Results:Across the three experiments, model performance varied by task structure and interaction severity. LLaMA3-70B demonstrated the highest recall and F1 score in the pointwise task, whereas GPT-4o-mini achieved superior accuracy and consistency in the pairwise and listwise tasks. MedGemma-27B showed competitive performance in identifying Category D interactions. Self-consistency decreased as task complexity increased, highlighting reduced stability in multi-drug reasoning. No model exhibited uniformly high reliability across all judgment formats. Conclusions:Current LLMs show promising but uneven capabilities in identifying DDIs across clinically relevant task structures. Performance degrades as the reasoning space expands, and stability across repeated queries remains limited. These findings emphasize the need for multi-format evaluation frameworks and reliability-aware assessment when considering LLMs for medication-safety applications.
Background:Large language models (LLMs) have shown the ability to diagnose complex medical cases, but only limited studies have evaluated the performance of LLMs in the development of evidence-based treatment plans. The purpose of this evaluation was to test four LLMs on their ability to develop safe and efficacious treatment plans on complex patients managed in the intensive care unit (ICU). Methods:Eight high-fidelity patient cases focusing on medication management were developed by critical care clinicians including history of present illness, laboratory values, vital signs, home medications, and current medications. Four LLMs [ChatGPT (GPT-3.5), ChatGPT (GPT-4), Claude-2, and Llama-2-70b] were prompted to develop an optimized medication regimen for each case. LLM generated medication regimens were then reviewed by a panel of seven critical care clinicians to assess safety and efficacy, as defined by medication errors identified and appropriate treatment for the clinical conditions. Appropriate treatment was measured by the average rate of clinician agreement to continue each medication in the regimen and compared using analysis of variance (ANOVA). Results:Clinicians identified a median of 4.1-6.9 medication errors per recommended regimen, and life-threatening medication recommendations were present in 16.3%-57.1% of the regimens, depending on LLM. Clinicians continued LLM-recommended medications at a rate of 54.6%-67.3%, with GPT-4 having the highest rate of medication continuation among all LLMs tested (p < 0.001) and the lowest rate of life-threatening medication errors (p < 0.001). Conclusion:Caution is warranted using present LLMs for medication regimens given the number of medication errors that were identified in this pilot study. However, LLMs did demonstrate potential to serve as clinical decision support for the management of complex medication regimens given the need for domain specific prompting and testing.
BackgroundLarge language models (LLMs) have demonstrated impressive performance on medical licensing and diagnosis-related exams. However, comparative evaluations to optimize LLM performance and ability in the domain of comprehensive medication management (CMM) are lacking. The purpose of this evaluation was to test various LLMs performance optimization strategies and performance on critical care pharmacotherapy questions used in the assessment of Doctor of Pharmacy students.MethodsIn a comparative analysis using 219 multiple-choice pharmacotherapy questions, five LLMs (GPT-3.5, GPT-4, Claude 2, Llama2-7b and 2-13b) were evaluated. Each LLM was queried five times to evaluate the primary outcome of accuracy (i.e., correctness). Secondary outcomes included variance, the impact of prompt engineering techniques (e.g., chain-of-thought, CoT) and training of a customized GPT on performance, and comparison to third year doctor of pharmacy students on knowledge recall vs. knowledge application questions. Accuracy and variance were compared with student’s t-test to compare performance under different model settings.ResultsChatGPT-4 exhibited the highest accuracy (71.6%), while Llama2-13b had the lowest variance (0.070). All LLMs performed more accurately on knowledge recall vs. knowledge application questions (e.g., ChatGPT-4: 87% vs. 67%). When applied to ChatGPT-4, few-shot CoT across five runs improved accuracy (77.4% vs. 71.5%) with no effect on variance. Self-consistency and the custom-trained GPT demonstrated similar accuracy to ChatGPT-4 with few-shot CoT. Overall pharmacy student accuracy was 81%, compared to an optimal overall LLM accuracy of 73%. Comparing question types, six of the LLMs demonstrated equivalent or higher accuracy than pharmacy students on knowledge recall questions (e.g., self-consistency vs. students: 93% vs. 84%), but pharmacy students achieved higher accuracy than all LLMs on knowledge application questions (e.g., self-consistency vs. students: 68% vs. 80%).ConclusionChatGPT-4 was the most accurate LLM on critical care pharmacy questions and few-shot CoT improved accuracy the most. Average student accuracy was similar to LLMs overall, and higher on knowledge application questions. These findings support the need for future assessment of customized training for the type of output needed. Reliance on LLMs is only supported with recall-based questions.