Background: Dynamic treatment regimes (DTR) are a sequence of treatment rules prescribing patient management based on evolving patient characteristics and are typical in the Intensive Care Unit (ICU). Determining the optimal time to stop a specific treatment under a DTR is clinically and computationally challenging: premature stopping can adversely affect patient outcomes, while prolonging it beyond clinical necessity can be harmful, expensive, and resource-intensive. This problem can be mathematically formulated as an Optimal Stopping Problem (OSP), where the goal is to identify the optimal time to stop the treatment that maximizes expected benefit or minimizes harm. OSPs are frequently encountered in treatment decisions, and unbiased decision-making requires causal reasoning to (i) correct treatment-related biases that naive prediction models fail to address, and (ii) enable sequential prediction under hypothetical interventions across all clinically permissible stopping options. Aims: In this scoping review, we aimed to identify literature that employ causal methods and facilitate sequential predictions under hypothetical interventions to solve OSPs, with emphasis on ICU applications. We further evaluated how different methodological frameworks address assumptions, estimands, and analytical challenges within the ICU context, emphasizing how experimental settings shape their applicability. Methods: Causal reasoning for OSP under DTRs remains understudied in intensive care. To address this gap, we extended the literature search scope to include (i) causal methods under DTRs for OSPs in general healthcare, and (ii) causal methods under DTRs applied to any ICU task. We applied the two search strategies to systematically query PubMed, arXiv, Web of Science, and Scopus from inception to November 3, 2025, screening results according to PRISMA guidelines for scoping reviews. Both methodological and applied studies were included. Additionally, we proposed a set of desiderata to evaluate the applicability of identified approaches to ICU decision-making and their capacity to address OSP. Results: We identified 52 relevant studies: 33 from general healthcare, and 19 from ICU settings. The studies originated primarily from epidemiology, biostatistics, and computer science. Each study was systematically evaluated and scored using the proposed desiderata to assess their relevance to ICU decision-making and applicability to OSP. Conclusion: Two promising approaches to estimate optimally timed DTRs emerged: (i) target trial emulation with three design patterns - thresholding covariates, sequential trials, and adaptive treatment-length strategies, (ii) algorithmic search with reinforcement learning and sequential decision tree methods. A key contribution of this review is the introduction of desiderata as a structured basis for evaluating methodological suitability to ICU decision-making. More broadly, this work underscores the significance of causal reasoning for the unbiased optimization of sequential treatment durations in intensive care.
Introduction Parathyroid hormone remains the primary biomarker used to classify and monitor CKD-MBD, yet its reliability for risk stratification and treatment guidance is limited. Multidimensional biomarker clustering that integrates indicators of bone turnover, vascular calcification, inflammatory and oxidative stress, may better capture the heterogeneity of clinical phenotypes, improve risk stratification, and inform on new therapeutic approaches in dialysis care. Methods We conducted a computational analysis of a multicentric cohort study including 471 hemodialysis patients. The Partitioning Around Medoids algorithm was used to cluster the cohort by a pathophysiology-based biomarker panel (noxPTH®, iPTH, OPG, sRANKL, ImAnOx®, PerOx®, β-Crosslaps and hsCRP) at baseline. The cluster phenotypes were characterized by anthropometrics, imaging and clinical data. Cluster-specific outcomes and event patterns were assessed after one year. Results Four distinct clusters were identified – Cluster 1 reflected a high bone turnover state with high antioxidative capacity and high bone-specific alkaline phosphatase. Cluster 2 showed low bone turnover with high antioxidative capacity and high sclerostin levels. Cluster 3 exhibited an inflammatory–oxidative phenotype, the highest one-year mortality and cluster membership improved prediction of events beyond clinical covariates. This phenotype was marked by high aortic calcium burden, vascular disease, markers of vitamin D degradation, low vaccine response, and hepatic steatosis. Cluster 4 showed a balanced biomarker profile and the highest rate of reclassification after one year. Conclusions Multidimensional biomarker clustering captures heterogeneity in bone and vascular disease phenotypes of CKD-MBD exhibiting differential mortality risk and event patterns, highlighting inflammation as an important pathway for further studies.
OBJECTIVES:Development and validation of two prediction models for obstetric anal sphincter injury (OASI). DESIGN:Population-based cohort study. SETTING:Nationwide (the Netherlands). POPULATION:Data from the Netherlands Perinatal Registry, describing nulliparous women who delivered a singleton live born infant in cephalic presentation at term from 2016 to 2020, with spontaneous (SVD) or operative vaginal delivery (OVD). METHODS:Based on literature and clinical expertise, a set of potential predictors was defined and derived from the national perinatal registry. A predictive model was constructed, and accessible nomograms provided. Internal and temporal external validation was performed. MAIN OUTCOME MEASURES:OASI rate. RESULTS:The risk of OASI in 171 046 women with SVD was 4.1%. After logistic regression with step-wise backward selection using Akaike Information Criterion (AIC), ten predictors were retained. These were: mediolateral episiotomy (MLE), expected fetal birth weight, duration of the 2nd stage, occipitoposterior presentation, induction of labour, epidural analgesia, Asian ethnicity, maternal age, gestational age and fetal sex. The final model had a moderate discriminative ability (AUC 0.67, 95% CI 0.67-0.68) and excellent calibration (Brier score 0.039). The average risk of OASI in 37 547 women with OVD was 3.5%. Seven predictors were retained in the model: MLE, expected fetal birth weight, duration 2nd stage of labour, occipitoposterior fetal presentation, epidural analgesia, Asian ethnicity and gestational age. The final model had moderate discrimination (AUC 0.68, 95% CI 0.67-0.70) and excellent calibration (Brier score 0.032). CONCLUSIONS:A prediction model for OASI was developed and validated for both nulliparous women with spontaneous vaginal delivery and with operative vaginal delivery. These models can form a basis to identify women with a high risk of OASI.
Clinical large language models (LLMs) are increasingly used for documentation, diagnosis, and decision support, but their opaque reasoning can limit clinician trust, regulatory assessment, and safe deployment. Explainability research has expanded rapidly, yet existing reviews largely address traditional machine learning or general-domain LLMs. We conducted a PRISMA-ScR scoping review to map explainability approaches for decoder-only clinical LLMs with over one billion parameters, searching PubMed, Scopus, Web of Science, ACM Digital Library, and arXiv through early 2026. Among 69 included studies, LLM-native generative and interactive methods dominated (58.0%, n = 40), spanning chain-of-thought rationales, retrieval-augmented evidence citation, and agentic decomposition. Intrinsic by-design methods accounted for 24.6% (n = 17); post-hoc XAI methods accounted for 17.4% (n = 12). General medicine was the most represented clinical domain, and diagnosis was the dominant task. Proprietary models were used in 75.4% of studies, yet every mechanistic analysis relied on open-source models, revealing a transparency asymmetry: most deployed models are the least transparent. Although 59.4% of studies quantitatively evaluated explanations, metrics remain non-standardized and rarely assess faithfulness. Local explanations predominated, and no study prospectively evaluated explanations in live clinical workflows. These findings show that clinical LLM explainability has shifted toward fluent generative rationales, but evidence that such explanations reflect model reasoning remains limited. To support trustworthy deployment, we highlight three regulatory priorities: prioritizing explanations that enable independent verification or logic auditing over plausibility-only rationales; preferring inspectable models where regulatory documentation is required; and prospectively validating explanations in clinical workflows before scaling.
Linear classifiers trained on hidden states of a large language model (LLM), linear probes, can flag factual errors from a single forward pass. Geometrically, that implies that true and false statements separate along a stable direction in hidden state space, i.e., the truth direction. Prior work disagrees on whether this generalises across input shifts, but the disagreement is hard to interpret because cross-dataset probe transfer experiments confound several kinds of input change at once. We isolate three such variables in medical question-answering (QA): writing style (register), domain (medical specialty), and corpus (dataset). We build a benchmark using 500 MedQA entries, each rewritten into four styles (textbook, patient, clinical note, colloquial), annotated with clinical specialty, and grouped with two other exam corpora, MedMCQA and MMLU-medical, for cross-dataset evaluation. Probing four open-weight LLMs (2–8B), we find that the truth direction is largely robust to writing style (mean Δ_register≈ 0.10 AUROC on held-out facts) and to medical specialty (Δ_specialty≈ 0.03), but degrades unevenly across corpora: by 0.12 AUROC on MMLU-medical and by 0.21 on MedMCQA, roughly twice the register gap. The register result replicates with a second generator and carries over to human-written patient questions. The truth direction is therefore largely stable within the medical domain but breaks under some corpus shifts, and question format does not explain the break, which suggests that the signal a linear probe recovers is partly bound to dataset structure rather than to medical knowledge alone.
Large language models (LLMs) are increasingly used to generate summaries from clinical notes. However, their ability to preserve essential diagnostic information remains underexplored, which could lead to serious risks for patient care. This study introduces DistillNote, an evaluation framework for LLM summaries that targets their functional utility by applying the generated summary downstream in a complex clinical prediction task, explicitly quantifying how much prediction signal is retained. We generated over 192,000 LLM summaries from MIMIC-IV clinical notes with increasing compression rates: standard, section-wise, and distilled section-wise. Heart failure diagnosis was chosen as the prediction task, as it requires integrating a wide range of clinical signals. LLMs were fine-tuned on both the original notes and their summaries, and their diagnostic performance was compared using the AUROC metric. We contrasted DistillNote's results with evaluations from LLM-as-judge and clinicians, assessing consistency across different evaluation methods. Summaries generated by LLMs maintained a strong level of heart failure diagnostic signal despite substantial compression. Models trained on the most condensed summaries (about 20 times smaller) achieved an AUROC of 0.92, compared to 0.94 with the original note baseline (97 percent retention). Functional evaluation provided a new lens for medical summary assessment, emphasizing clinical utility as a key dimension of quality. DistillNote introduces a new scalable, task-based method for assessing the functional utility of LLM-generated clinical summaries. Our results detail compression-to-performance tradeoffs from LLM clinical summarization for the first time. The framework is designed to be adaptable to other prediction tasks and clinical domains, aiding data-driven decisions about deploying LLM summarizers in real-world healthcare settings.
Objective: Heart failure (HF) patients present with diverse phenotypes affecting treatment and prognosis. This study evaluates models for phenotyping HF patients based on left ventricular ejection fraction (LVEF) classes, using structured and unstructured data, assessing performance and interpretability. Materials and Methods: The study analyzes all HF hospitalizations at both Amsterdam UMC hospitals (AMC and VUmc) from 2015 to 2023 (33,105 hospitalizations, 16,334 patients). Data from AMC were used for model training, and from VUmc for external validation. The dataset was unlabelled and included tabular clinical measurements and discharge letters. Silver labels for LVEF classes were generated by combining diagnosis codes, echocardiography results, and textual mentions. Gold labels were manually annotated for 300 patients for testing. Multiple Transformer-based (black-box) and Aug-Linear (white-box) models were trained and compared with baselines on structured and unstructured data. To evaluate interpretability, two clinicians annotated 20 discharge letters by highlighting information they considered relevant for LVEF classification. These were compared to SHAP and LIME explanations from black-box models and the inherent explanations of Aug-Linear models. Results: BERT-based and Aug-Linear models, using discharge letters alone, achieved the highest classification results (AUC=0.84 for BERT, 0.81 for Aug-Linear on external validation), outperforming baselines. Aug-Linear explanations aligned more closely with clinicians' explanations than post-hoc explanations on black-box models. Conclusions: Discharge letters emerged as the most informative source for phenotyping HF patients. Aug-Linear models matched black-box performance while providing clinician-aligned interpretability, supporting their use in transparent clinical decision-making.
In this study, we set a benchmark for adverse drug event (ADE) detection in Dutch clinical free text documents using several transformer models, clinical scenarios and fit-for-purpose performance measures. We trained a Bidirectional Long Short-Term Memory (Bi-LSTM) model and four transformer-based Dutch and/or multilingual encoder models (BERTje, RobBERT, MedRoBERTa.nl, and NuNER) for the tasks of named entity recognition (NER) and relation classification (RC) using 102 richly annotated Dutch ICU clinical progress notes. Anonymized free text clinical progress notes of patients admitted to intensive care unit (ICU) of one academic hospital and discharge letters of patients admitted to Internal Medicine wards of two non-academic hospitals were reused. We evaluated our ADE RC models internally using gold standard (two-step task) and predicted entities (end-to-end task). In addition, all models were externally validated on detecting ADEs at the document level. We report both micro- and macro-averaged F1 scores, given the imbalance of ADEs in the datasets. Although differences for the ADE RC task between the models were small, MedRoBERTa.nl was the best performing model with macro-averaged F1 score of 0.63 using gold standard and 0.62 using predicted entities. The MedRoBERTa.nl models also performed the best in our external validation and achieved recall of between 0.67 to 0.74 using predicted entities, meaning between 67 to 74
An artificial intelligence boom is currently ongoing, mainly due to large language models, leading to significant interest in artificial intelligence and subsequently also in machine learning (ML). One area where ML is often applied, prediction modelling, has also long been a focus of conventional statistics. As a result, multiple studies have aimed to prove superiority of one of the two scientific disciplines over the other. However, we argue that ML and conventional statistics should not be competing fields. Instead, both fields are intertwined and complementary to each other. To illustrate this, we discuss some essentials of prediction modelling, elaborate on prediction modelling using techniques from conventional statistics, and explain prediction modelling using common ML techniques such as support vector machines, random forests, and artificial neural networks. We then showcase that conventional statistics and ML are in fact similar in many aspects, including underlying statistical concepts and methods used in model development and validation. Finally, we argue that conventional statistics and ML can and should be seen as a single integrated field. This integration can further improve prediction modelling for both disciplines (e.g. regarding fairness and reporting standards) and will support the ultimate goal: developing the best performing prediction models for the patient and healthcare provider.
Patients admitted to the intensive care unit (ICU) are often treated with multiple high-risk medications. Over- and underprescribing of indicated medications, and inappropriate choice of medications frequently occur in the ICU. This risk has to be minimized. We evaluate the performance of recommendation methods in suggesting appropriate medications and examine whether incorporating clinical patient data beyond the medication list improves recommendations. Using the MIMIC-III dataset, we formulate medication list completion as a recommendation task. Our analysis includes four autoencoder-based approaches and two strong baselines. We used as inputs either only known medications, or medications together with patient data. We showed that medication recommender systems based on autoencoders may successfully recommend medications in the ICU.
Objective To compare various methods for extracting daily dosage information from prescription signatures (sigs) and identify the best performers.Materials and Methods In this study, 5 daily dosage extraction methods were identified. Parsigs, RxSig, Sig2db, a large language model (LLM), and a bidirectional long short-term memory (BiLSTM) model were selected. The methods were analyzed with regard to positive predictive value (PPV), sensitivity, F1-score, cost to compute, and time to finish on a sig dataset in the context of heart failure with reduced ejection fraction.Results The dataset consisted of 29 896 free-text sigs, which were split into training and validation sets of 70% and 30%, respectively. The BiLSTM model scored lowest with an F1-score of 0.71. The LLM GPT-4o and regular expression-based RxSig achieved the highest F1-scores with 0.98 and 0.95, respectively. The LLM outperformed RxSig in sensitivity. RxSig outperformed the LLM in PPV. Additionally, RxSig had a lower run time and no costs compared to a cost of 25 dollars.Discussion In practical usage, it would be preferable for an algorithm to score high on PPV and F1-score, to reduce false positive assertions of daily dosage. Additionally, long running times and high costs are not scalable for larger datasets. Thus, RxSig is likely the most scalable approach. Further research is needed to investigate the generalizability of the findings.Conclusion This study demonstrates that both the LLM and RxSig models excel in daily dose extraction from free-text sigs, with the RxSig model appearing to be the more scalable approach. When medications are ordered, a prescription signature (sig) specifies how the patient should take the medication (eg, "take one tablet twice daily"). Accurately extracting medication dosages from these unstructured sigs is essential for ensuring patients receive safe and effective treatment. This study compared 5 methods for automating this process: RxSig (a rule-based approach), GPT-4o (a large language model), Sig2db, Parsigs, and a bidirectional long short-term memory (BiLSTM) (a machine learning model). The methods were tested on 29 896 sigs for heart failure medications to assess their performance in positive predictive value (PPV), sensitivity, F1-score, speed, and cost.The results showed that GPT-4o and RxSig performed best, with high PPV, sensitivity, and F1-scores. GPT-4o scored the highest on F1-score and sensitivity but incurred higher costs and longer processing times. In contrast, RxSig offered similar scores with no cost and faster processing, making it the most practical and scalable solution for large datasets. The BiLSTM scored the lowest.This study highlights the potential of rule-based approaches, like RxSig, for efficiently extracting daily dosages from unstructured sigs. Future research should explore ways to integrate advanced models like GPT-4o into workflows, particularly for cases where RxSig encounters challenges, while ensuring scalability and cost-effectiveness for broader clinical use.
While conventional comparative effectiveness studies report average treatment effects (ATEs), they may not capture the complex patterns of individual treatment responses. Causal forest, a tree-based machine learning approach that identifies patient-level factors influencing differential treatment responses, facilitates the estimation of conditional average treatment effects (CATEs) based on an individual’s covariates. Our aim was to identify patient characteristics that predict differential responses to adalimumab vs. methotrexate in psoriasis using causal forest analysis, and to characterize heterogeneous treatment effects across clinically relevant subgroups. The analysis compared biologic-naive adult patients who initiated either adalimumab or methotrexate between September 2007 and April 2024 across the UK and Republic of Ireland following the protocol recorded by the prospective registry the British Association of Dermatologists Biologics and Immunomodulators Register (BADBIR). Missing data were handled using one randomly selected dataset from multiple imputation by chained equations. A causal forest model was employed to estimate CATE for achieving Psoriasis Area and Severity Index (PASI) ≤ 2 during the treatment period, incorporating baseline patient characteristics and comorbidities. Variable importance analysis was used to identify key predictors of differential treatment response between adalimumab and methotrexate, with their effects quantified through best linear projection. In total 6810 patients (adalimumab 3974, methotrexate 2836) were included in the analysis. While adalimumab demonstrated superior overall effectiveness [ATE 0.41, 95% confidence interval (CI) 0.38–0.44], CATE analysis revealed substantial variation in individual treatment effects (median 0.42, interquartile range 0.37–0.47). Key predictors of differential treatment effects included male sex (absolute risk difference 0.15, P < 0.001). Effects were greater with increasing baseline PASI (0.007 per unit, P = 0.001), but lower with increasing age (−0.004 per year, P = 0.006) and weight (−0.003 per kg, P = 0.002). Subgroup analysis (CATE with 94% CI) showed higher treatment effects in younger patients (< 30 years: 0.46, 0.41–0.51 vs. > 70 years: 0.36, 0.34–0.39), female patients (0.45, 0.41–0.50 vs. male: 0.38, 0.35–0.41) and those with higher baseline PASI (> 20: 0.45, 0.41–0.49 vs. < 10: 0.41, 0.37–0.46). Increasing weight (< 90 kg: 0.43, 0.38–0.48 vs. > 120 kg: 0.40, 0.36–0.45) and comorbidity burden (no comorbidities: 0.44, 0.40–0.49 vs. ≥ 3 comorbidities: 0.37, 0.33–0.40) were associated with reduced treatment effects. The causal forest analysis reveals significant heterogeneity in responses to adalimumab and methotrexate, underscoring the impact of individual characteristics on treatment effectiveness. Factors such as age, sex, baseline PASI, weight and number of comorbidities are associated with differential treatment responses when prescribing adalimumab or methotrexate. These findings provide a data-driven framework to support personalized treatment decisions in clinical practice.
Background: Improving prediction models to timely detect lung cancer is paramount. Our aim is to develop and validate prediction models for early detection of lung cancer in primary care, based on free-text consultation notes, that exploit the order and context among words and sentences. Methods: Data of all patients enlisted in 49 general practices between 2002 and 2021 were assessed, and we included those older than 30 years with at least one free-text note. We developed two models using a hierarchical architecture that relies on attention and bidirectional long short-term memory networks. One model used only text, while the other combined text with clinical variables. The models were trained on data excluding the five months leading up to the diagnosis, using target replication and a tuning set, and were tested on a separate dataset for discrimination, PPV, and calibration. Results: A total of 250,021 patients were enlisted, with 1507 having a lung cancer diagnosis. Included in the analysis were 183,012 patients, of which 712 had the diagnosis. From the two models, the combined model showed slightly better performance, achieving an AUROC on the test set of 0.91, an AUPRC of 0.05, and a PPV of 0.034 (0.024, 0.043), and showed good calibration. To early detect one cancer patient, 29 high-risk patients would require additional diagnostic testing. Conclusions: Our models showed excellent discrimination by leveraging the word and sentence structure. Including clinical variables in addition to text slightly improved performance. The number needed to treat holds promise for clinical practice. Investigating external validation and model suitability in clinical practice is warranted.
Our objective was to create a gold standard Dutch language annotated corpus of clinical notes with adverse drug event (ADE) mentions, specifically for Intensive Care patients with drug-related acute kidney injury. We used anonymized clinical notes from 102 adult intensive care unit (ICU) patients suspected of acute kidney injury (AKI) and admitted to Amsterdam University Medical Centre, The Netherlands, over a four-year period (November 2015– January 2020). The notes were extracted from the electronic health record (EHR) system and manually reviewed for drug-related causes. Each clinical note contained at least one ADE mention (drug-related AKI). Annotation guidelines were developed over three rounds of annotation based on review of annotations and clarifications during the process. Two clinical expert annotators labelled mentions of drugs and disorders, as well as the relationship between these entities indicating an ADE. The final gold standard corpus was a result of adjudication of the two sets of expert labels. The corpus contains 102 notes with 16,470 labels, consisting of 8,914 Disorder entities, 5,307 Drug entities, 134 Qualitative Concept entities, 1,501 Indication relations, and 614 ADE relations. Annotation reached high agreement for all entities (F1 score 0.7724) with an expected lower agreement for relations (F1 score 0.4327). The Dutch ADE corpus is a real-world data set that can be used to evaluate natural language processing pipelines for ADE detection tasks. Although the corpus was developed for drug-related AKI, 158 additional ADEs were identified. The combination of iterative annotation guideline development and double annotation followed by adjudication produced high quality annotations. Future work will use this gold standard annotated corpus to train and validate NLP models to detect ADEs in Dutch clinical text.
Increased monitoring of health-related data for ICU patients holds great potential for the early prediction of medical outcomes. Research on whether the use of clinical notes and concepts from knowledge bases can improve the performance of prediction models is limited. We investigated the effects of combining clinical variables, clinical notes, and clinical concepts. We focus on the early prediction of Acute Kidney Injury (AKI) in the intensive care unit (ICU). AKI is a sudden reduction in kidney function measured by increased serum creatinine (SCr) or decreased urine output. AKI may occur in up to 30% of ICU stays. We developed three models based on convolutional neural networks using data from the Medical Information Mart for Intensive Care (MIMIC) database. The models used clinical variables, free-text notes, and concepts from the Elsevier H-Graph. Our models achieved good predictive performance (AUROC 0.73-0.90). These models were assessed both when using Scr and urine output as predictors and when omitting them. When Scr and urine output were used as predictors, models that included clinical notes and concepts together with clinical variables performed on par with models that only used clinical variables. When excluding SCr and urine output, predictive performance improved by combining multiple modalities. The models that used only clinical variables were externally validated on the eICU dataset and transported fairly to the new population (AUROC 0.68-0.77). Our in-depth comparison of modalities and text representations may further guide researchers and practitioners in applying multimodal models for predicting AKI and inspire them to investigate multimodality and contextualized embeddings for other tasks. Our models can support clinicians to promptly recognize and treat deteriorating AKI patients and may improve patient outcomes in the ICU.
PURPOSE:The potential of vancomycin to cause acute kidney injury (AKI) in adult intensive care patients is subject to debate due to suboptimal designs of past studies. Therefore, we aimed to estimate the effect of initiating vancomycin versus one of several minimally nephrotoxic alternative antibiotics on the 14-day risk of AKI using the target trial emulation framework. METHODS:A hypothetical trial was emulated using routinely collected data from 15 Dutch intensive care units (ICUs) spanning 2010-2019. We used an active comparator control group with the following alternative antibiotics: clindamycin, linezolid, teicoplanin, meropenem, cefazolin, and daptomycin. AKI was diagnosed according to the KDIGO serum creatinine (SCr) criteria. Cumulative incidence curves were estimated using the Aalen-Johansen method and adjusted for confounding and selection bias through inverse probability of treatment and censoring weighting. Given the time lag of 24-48 h between changes in renal function and SCr, we summarized the estimates by calculating the absolute risks and risk differences at both 2 and 14 days after initiation. RESULTS:We included 1809 ICU admissions. After adjustment, vancomycin was associated with a higher risk of AKI at 14 days of follow-up compared to the alternative antibiotics (0.28 [95% confidence interval (CI) 0.21-0.34] vs. 0.17 [95% CI 0.14-0.20]; risk difference 0.11 [95% CI 0.04-0.19]), but not at 2 days of follow-up (0.10 [95% CI 0.06-0.12] vs. 0.10 [95% CI 0.08-0.11]; risk difference 0.00 [95% CI -0.03-0.03]). CONCLUSIONS:Our findings indicate that vancomycin causes a higher risk of AKI compared to the alternative antibiotics. We recommend clinicians to be compliant with vancomycin-induced AKI prevention strategies, such as therapeutic drug monitoring or the consideration of an alternative antibiotic if possible.