OBJECTIVE:Machine learning applications for longitudinal electronic health records often forecast the risk of events at fixed time points, whereas survival analysis achieves dynamic risk prediction by estimating time-to-event distributions. Here, we propose a novel conditional variational autoencoder-based method, DySurv, which uses a combination of static and longitudinal measurements from electronic health records to estimate the individual risk of death dynamically. MATERIALS AND METHODS:DySurv directly estimates the cumulative risk incidence function without making any parametric assumptions on the underlying stochastic process of the time-to-event. We evaluate DySurv on 6 time-to-event benchmark datasets in healthcare, as well as 2 real-world intensive care unit (ICU) electronic health records (EHR) datasets extracted from the eICU Collaborative Research (eICU) and the Medical Information Mart for Intensive Care database (MIMIC-IV). RESULTS:DySurv outperforms other existing statistical and deep learning approaches to time-to-event analysis across concordance and other metrics. It achieves time-dependent concordance of over 60% in the eICU case. It is also over 12% more accurate and 22% more sensitive than in-use ICU scores like Acute Physiology and Chronic Health Evaluation (APACHE) and Sequential Organ Failure Assessment (SOFA) scores. The predictive capacity of DySurv is consistent and the survival estimates remain disentangled across different datasets. DISCUSSION:Our interdisciplinary framework successfully incorporates deep learning, survival analysis, and intensive care to create a novel method for time-to-event prediction from longitudinal health records. We test our method on several held-out test sets from a variety of healthcare datasets and compare it to existing in-use clinical risk scoring benchmarks. CONCLUSION:While our method leverages non-parametric extensions to deep learning-guided estimations of the survival distribution, further deep learning paradigms could be explored.
Electronic health records (EHRs) capture evolving physiological processes, yet most machine learning models impose static or sequential assumptions that flatten their temporal and relational complexity. We introduce DynaGraph, a dynamic and interpretable graph learning framework that constructs evolving spatio-temporal graphs from multivariate clinical time-series. Unlike previous methods, DynaGraph learns the structure of relationships between different clinical variables over time without predefined graphs, integrates sequential embeddings with contrastive graph augmentation, and incorporates a pseudo-attention mechanism to reveal temporally resolved risk factors. Trained end-to-end with a novel multi-loss objective that combines focal, structural, and contrastive components, DynaGraph addresses two pervasive challenges in real-world clinical modelling: class imbalance and temporal instability. We evaluated DynaGraph on four large-scale EHR datasets totalling 40,856 patients: MIMIC-III (17,279 ICU admissions), eICU (1433 cardiac ICU patients), HiRID-ICU (33,000 patients), and EHRSHOT (2378 primary care patients). DynaGraph consistently outperforms 14 state-of-the-art baselines, achieving 6-8% relative improvements in area under the precision-recall curve (AUPRC) and significant gains in sensitivity (12-22% over leading methods). Beyond predictive performance, DynaGraph offers time-specific interpretability aligned with clinical reasoning, providing gradient-based feature importance scores at 3-hour intervals that identify which physiological relationships drive predictions. This framework explicitly models temporal attribution of risk factors across patient trajectories in a millisecond inference time.
AIM:To explore staff and patient perception of the newly co-developed wearable monitoring system (WMS), including acceptability of use in clinical practice. DESIGN:Pragmatic qualitative descriptive study. METHODS:Semi-structured interviews were conducted with 12 patient participants and eight staff members between June 2023 and August 2024, and were analysed thematically. RESULTS:Three themes were identified, building on previous qualitative work around the use of WMS on hospital wards. The first theme-centralised continuous monitoring enhances care-explores how WMS provides staff with a means to provide safe, efficient care with the ability to see the vital signs away from the patient. Patients reported feeling safer, knowing they were being monitored when staff were not at the bedside. The second theme-human connection at the bedside-considers how both patients and staff emphasised that the system should not replace nurse/patient interactions and face-to-face care, even though it provided patients with a stronger sense of independence. The final theme-system usability and integration into care-focuses on use of the system in clinical practice and implications for the future. CONCLUSION:Wearable monitoring systems have the potential to support nurses to provide safer, more efficient care, whilst providing reassurance to patients. However, centralised monitoring should not replace face-to-face clinical contact, and careful consideration should be given to who would benefit most from the technology. IMPACT:This study extends existing knowledge of the impact of WMS from being a tool to enhance patient safety to an intervention to improve nurse efficiency and patient experience, within the context of a high-demand surgical ward. PATIENT AND PUBLIC CONTRIBUTION:Patients and members of the public were involved in study design and data collection. Their contributions included participating in advisory groups, ensuring the research addressed patient-relevant priorities.
AIMS:Discharge from an intensive care unit (ICU) to the ward has long been recognised as a difficult and dangerous time for patients. This study aims to explore the experiences of patients, family members and staff of ongoing care following discharge from ICU to the ward. DESIGN:An exploratory qualitative interview study, part of a mixed methods research project exploring post-ICU ward care to inform practice changes to improve outcomes for critical care survivors. METHODS:Semi-structured interviews were conducted with 55 purposively sampled patients, family members and multi-professional staff between 2017 and 2018 at three NHS trusts in the UK. Data were analysed thematically. FINDINGS:Three main themes were identified: Being wardable discusses the tension between perceptions of readiness for ICU between ICU and the ward staff; Post-ICU patients as other analyses the characterisation of otherness of post-ICU patients in comparison with other ward patients; and Fear and anxiety considers the impact of this perceived otherness on both patients and staff. CONCLUSION:The term wardable is used by staff to indicate the suitability or otherwise of patient transfer. There is a tension between being deemed wardable from an ICU and ward perspective. This exacerbates the perception of otherness and compounds the fear and anxiety related to post-ICU care experienced by both patients and ward staff. IMPLICATIONS FOR PATIENT CARE:By recognising the tension between being ready for ICU discharge (not requiring organ support) and being wardable (having care needs which can be fulfilled on the ward they are being discharged to), clinicians may better support patients during this transition of care. This has the potential to improve outcomes for post-ICU patients, as well as improve the experience by both patients and the staff and family members caring for them. REPORTING METHOD:The COREQ reporting checklist was used in the reporting of this manuscript. PATIENT AND PUBLIC CONTRIBUTION:Patients and public were involved throughout the REFLECT project, from design through to dissemination. This included advising on the approach to patient participants and supporting dissemination of results via social media. TRIAL REGISTRATION:ISRCTN: 14658054.
BACKGROUND:Nutrition during hospitalisation following critical illness is fundamental to rehabilitation, but provision is often poor. AIM:To analyse the process of delivering nutrition to post-ICU patients on the ward. STUDY DESIGN:This work forms part of a mixed methods study. In three representative UK hospitals, we conducted: a structured judgement review (SJR) of 300 patients who died following discharge from ICU; in-depth reviews of 20 survivors and 20 deaths judged to be 'probably avoidable' in the SJR; and interviews with 55 patients, family members and staff about their experiences of post-ICU ward care. We extracted nutrition provision information from the primary data. Using these data and the Functional Resonance Analysis Method (FRAM), we worked with stakeholders to map the process of delivering enteral feed to patients discharged from ICU to hospital wards. RESULTS:The stakeholder meeting included a dietitian and a medical registrar from two of the three primary data collection sites, two researchers with knowledge of the primary data (with nursing and physiotherapy backgrounds) and a human factors facilitator. The FRAM revealed that providing enteral feeding on the ward is not a linear process, with three clusters of functions delivering distinct steps within the wider process: establishing the need for nasogastric feeding, the nasogastric placement cycle and nasogastric feed delivery. There are multiple points in these processes where failures in multi-professional teamwork result in the absence of the required steps to move through the processes in a timely manner. In particular, the process for confirming nasogastric tube placement risked system-related delays to feed administration, significantly affecting the volume of feed delivered to patients. CONCLUSIONS:The FRAM identified multiple process problems affecting nutritional support that may have led to profound consequences for post-ICU patients, with multi-professional collaboration a key factor for effective delivery of timely enteral nutrition. RELEVANCE TO CLINICAL PRACTICE:Improving collaborative working processes and addressing common nutritional support problems after ICU discharge could improve nutritional delivery and expedite recovery from critical illness. TRIAL REGISTRATION:ISRCTN14658054.
BACKGROUND:Many studies have evaluated the use of wearable monitoring systems to improve patient safety in hospital. Although some have demonstrated effects on intensive care admissions, there remains little evidence of impact on patient outcomes such as mortality, hospital length of stay, and time to antibiotic administration. Very few studies have focused on how wearable monitoring systems are used in clinical practice, including how the rate of manual vital sign measurements (MVSMs) is affected. OBJECTIVE:Our primary aim was to describe the physiological pattern of vital signs in hospitalized patients treated for COVID-19 outside of critical care. We also report an exploratory post hoc analysis of the impact of displaying wearable monitoring system data on the frequency of intermittent MVSMs. METHODS:We conducted a retrospective study during the COVID-19 pandemic following deployment of a wearable monitoring system that continuously displayed heart rate, respiratory rate, and oxygen saturation levels. We included patients treated for COVID-19 in 3 isolation wards in a large UK hospital. Wearable monitoring system data were displayed on a dashboard in the center of each ward. We analyzed the patterns of vital signs in patients monitored using the wearable monitoring system. We compared the time to next observation (led by nursing staff) for routinely collected MVSMs between periods when patients were continuously monitored and those when they were not. In exploratory post hoc analysis, we tested whether the difference varied between stable (early warning score [EWS] above the escalation threshold) and unstable patients. RESULTS:Patients (N=144) had continuous vital signs above the EWS threshold for escalation for 32.7% (2133/6528) of time monitored. The unadjusted median time between MVSMs for continuously monitored periods was 39 minutes (95% CI 29-49; P<.001) longer than for unmonitored periods. When adjusted for EWS category and participant-level clustering, the effect was attenuated but remained significant (14.6 minutes; P<.001). In exploratory post hoc analysis, we found that increases were larger during stable observation periods (51 minutes, 95% CI 39-62; P<.001) than during unstable periods (16 minutes, 95% CI 8-24; P<.001). However, adjusted analyses did not support a significant difference between stable and unstable periods. CONCLUSIONS:Patients in this study were at elevated risk of deterioration, spending a third of monitored time at or above the escalation threshold. We found that, by offering additional vital sign data between manual measurements, the time between routine MVSMs increased, which may reflect changes in nursing task prioritization. Although patient safety outcomes were not directly measured, we found no indication that reducing observation frequency adversely affected patient safety.
Myocardial infarction (MI) remains one of the greatest contributors to mortality, and patients admitted to the intensive care unit (ICU) with myocardial infarction are at higher risk of death. In this study, we use two retrospective cohorts extracted from two US-based ICU databases, eICU and MIMIC-IV, to develop an explainable pseudo-dynamic machine learning framework for mortality prediction in the ICU. The method provides accurate prediction for ICU patients up to 24 hours before the event and provides time-resolved interpretability. We compare standard supervised machine learning algorithms with novel tabular deep learning approaches and find that an integrated XGBoost model in our EHR time-series extraction framework (XMI-ICU) performs best. The framework was evaluated on a held-out test set from eICU and externally validated on the MIMIC-IV cohort using the most important features identified by time-resolved Shapley values. XMI-ICU achieved AUROCs of 92.0 (balanced accuracy of 82.3) for a 6-hour prediction of mortality. We demonstrate that XMI-ICU maintains reliable predictive performance across different prediction horizons (6, 12, 18, and 24 hours) during ICU stay while also achieving successful external validation in a separate patient cohort from MIMIC-IV without any previous training on that dataset. We also evaluated the framework for clinical risk analysis by comparing it to the standard APACHE IV system in active use. We show that our framework successfully leverages time-series physiological measurements from ICU health records by translating them into stacked static prediction problems for mortality in heart attack patients and can offer clinical insight from time-resolved interpretability through the use of Shapley values.
Background:Severe maternal morbidity (SMM) is an important indicator for the improvement of maternity care. Measurement of SMM varies, limiting global comparisons. To promote concordance we studied how SMM has been defined in epidemiological practice. Methods:Comprehensive composite definitions of SMM in pregnancy or up to 6 weeks postnatal that captured both obstetric and non-obstetric processes in high-income settings were identified through a prospectively registered (PROSPERO CRD42023421377) systematic search of PubMed, Embase, and Google Scholar 01/01/1993-31/08/2024. Clinical concepts, diagnostic and procedural codes captured by definitions of SMM were compared and the variation between definitions was described. Findings:The initial search identified 7852 records and 40 studies were included: 28 studies that reported 32 definitions of SMM for use with administrative data, with median incidence of 11.4/1000, and 13 studies that reported 13 definitions for use with the primary medical record, with median SMM incidence of 6.7/1000. The majority of definitions included cardiac, respiratory, and renal dysfunction or failure; haemorrhagic, thrombotic or infective morbidity; and critical interventions. Up to 75% of cases of SMM under some definitions involved transfusion. The main source of variation between definitions was the selection and definition of common obstetric diagnoses. Variation in the sources of additional routine data required to construct a definition also limited comparability. Interpretation:Despite common approaches to defining SMM, there are opportunities to improve comparability. No two definitions for use with administrative data in different settings involved a similar incidence and set of components and involved a similar distribution of components among cases. Harmonization of the purpose, constituent codes, and sources of data would facilitate comparisons between maternity systems. Funding:This work was supported by the Medical Research Council [MR/X006115/1] as well as the National Institute for Health Research [NIHR204430].
Estimating heterogeneous treatment effects (HTEs) of continuous-valued interventions on survival, that is, time-to-event (TTE) outcomes, is crucial in various fields, notably in clinical decision-making and in driving the advancement of next-generation clinical trials. However, while HTE estimation for continuous-valued (i.e., dosage-dependent) interventions and for TTE outcomes have been separately explored, their combined application remains largely overlooked in the machine learning literature. We propose DoseSurv, a varying-coefficient network designed to estimate HTEs for different dosage-dependent and non-dosage treatment options from TTE data. DoseSurv uses radial basis functions to model continuity in dose-response relationships and learns balanced representations to address covariate shifts arising in HTE estimation from observational TTE data. We present experiments across various treatment scenarios on both simulated and real-world data, demonstrating DoseSurv's superior performance over existing baseline models.
Patients discharged from intensive care units (ICU) commonly experience multiple problems in care during the acute hospital period. These can negatively impact their recovery and contribute to poor outcomes such as ICU readmission or in-hospital mortality. Many studies have aimed to address this through non-pharmacological interventions. However there has been no comprehensive synthesis of this literature. To improve care during the post-ICU in-hospital period, it is important to understand existing interventions and highlight priorities to guide future research. We therefore aimed to assess the extent of current literature relating to non-pharmacological interventions delivered to critical care survivors in hospital. We systematically searched five electronic databases (MEDLINE, EMBASE, CINAHL, AMED and CENTRAL) and grey literature to 4th February 2025. Search results were independently screened for eligibility by two reviewers at title, abstract and full text. Reports relating to non-pharmacological interventions delivered to adult patients in hospital following discharge from critical care were included. Study characteristics, intervention delivery and development, and outcomes were extracted using a formal data charting process. Searches yielded 41,242 reports, from which 202 met the inclusion criteria. The most common interventions were critical care outreach/follow-up (CCOT/FU) (n = 93, 46
Purpose of review Perioperative risk scores aim to risk-stratify patients to guide their evaluation and management. Several scores are established in clinical practice, but often do not generalize well to new data and require ongoing updates to improve their reliability. Recent advances in machine learning have the potential to handle multidimensional data and associated interactions, however their clinical utility has yet to be consistently demonstrated. In this review, we introduce key model performance metrics, highlight pitfalls in model development, and examine current perioperative risk scores, their limitations, and future directions in risk modelling. Recent findings Newer perioperative risk scores developed in larger cohorts appear to outperform older tools. Recent updates have further improved their performance. Machine learning techniques show promise in leveraging multidimensional data, but integrating these complex tools into clinical practice requires further validation, and a focus on implementation principles to ensure these tools are trusted and usable. Summary All perioperative risk scores have some limitations, highlighting the need for robust model development and validation. Advancements in machine learning present promising opportunities to enhance this field, particularly through the integration of diverse data sources that may improve predictive performance. Future work should focus on improving model interpretability and incorporating continuous learning mechanisms to increase their clinical utility.
BACKGROUND:It has long been suspected that the vital sign abnormalities that accompany bacterial infection are subtle or absent in older adults. This review summarises the evidence for whether older adults present with different vital sign abnormalities to younger adults when hospitalised with bacterial infection. METHODS:MEDLINE, EMBASE and CINAHL EBSCO were searched from inception to 19 December 2024 for English-language research articles of patients hospitalised with bacterial infection reporting age and admission vital signs. We used meta-regression to assess how vital signs vary with age. Where studies reported vital signs in multiple age groups, we undertook a meta-analysis in younger (<65) and older patients (≥65). Evidence quality was assessed using an adapted Quality Assessment of Diagnostic Accuracy Studies-2 tool. RESULTS:Our search yielded 14 487 studies; 132 were included after screening. Older adults were less likely to be tachycardic (RR 0.82, 0.69 to 0.97, I2 = 86.5%) with a mean difference in heart rate of 5 bpm (-7 to -3 bpm, I2 = 88.3%). Older adults were less likely to be febrile (RR 0.89, 0.83 to 0.95, I2 = 85.9%) with a mean difference in temperature of 0.14°C (-0.26 to -0.02°C, I2 = 94.6%). Most (129/132) studies were at high risk of bias. CONCLUSIONS:Whilst differences in absolute values were small, there was consistency in the finding that older adults were less likely than younger adults to be tachycardic or febrile. As vital signs at presentation may prompt suspicion of infection, influencing investigations and treatment, special consideration for the possibility of infection in older patients with normal vital signs may be warranted.
Large language models show remarkable potential in healthcare but face critical explainability challenges that must be addressed before widespread clinical deployment. Here, we examine technical and regulatory solutions needed to develop trustworthy, transparent large language models for responsible healthcare integration. Large language models show remarkable potential in healthcare but face critical explainability challenges that must be addressed before widespread clinical deployment. Here, Munib Mesinovic, Peter Watkinson and Tingting Zhu examine technical and regulatory solutions needed to develop trustworthy, transparent large language models for responsible healthcare integration.
Learning from longitudinal electronic health records is limited if it does not capture the temporal trajectories of the patient's state in a clinical setting. Graph models allow us to capture the hidden dependencies of the multivariate time-series when the graphs are constructed in a similar dynamic manner. Previous dynamic graph models require a pre-defined and/or static graph structure, which is unknown in most cases, or they only capture the spatial relations between the features. Furthermore in healthcare, the interpretability of the model is an essential requirement to build trust with clinicians. In addition to previously proposed attention mechanisms, there has not been an interpretable dynamic graph framework for data from multivariate electronic health records (EHRs). Here, we propose DynaGraph, an end-to-end interpretable contrastive graph model that learns the dynamics of multivariate time-series EHRs as part of optimisation. We validate our model in four real-world clinical datasets, ranging from primary care to secondary care settings with broad demographics, in challenging settings where tasks are imbalanced and multi-labelled. Compared to state-of-the-art models, DynaGraph achieves significant improvements in balanced accuracy and sensitivity over the nearest complex competitors in time-series or dynamic graph modelling across three ICU and one primary care datasets. Through a pseudo-attention approach to graph construction, our model also indicates the importance of clinical covariates over time, providing means for clinical validation.
Background Early warning scores (EWS) are routinely used in hospitals to assess a patient’s risk of deterioration. EWS are traditionally recorded on paper observation charts but are increasingly recorded digitally. In either case, evidence for the clinical effectiveness of such scores is mixed, and previous studies have not considered whether EWS leads to changes in how deteriorating patients are managed. Objective This study aims to examine whether the introduction of a digital EWS system was associated with more frequent observation of patients with abnormal vital signs, a precursor to earlier clinical intervention. Methods We conducted a 2-armed stepped-wedge study from February 2015 to December 2016, over 4 hospitals in 1 UK hospital trust. In the control arm, vital signs were recorded using paper observation charts. In the intervention arm, a digital EWS system was used. The primary outcome measure was time to next observation (TTNO), defined as the time between a patient’s first elevated EWS (EWS ≥3) and subsequent observations set. Secondary outcomes were time to death in the hospital, length of stay, and time to unplanned intensive care unit admission. Differences between the 2 arms were analyzed using a mixed-effects Cox model. The usability of the system was assessed using the system usability score survey. Results We included 12,802 admissions, 1084 in the paper (control) arm and 11,718 in the digital EWS (intervention) arm. The system usability score was 77.6, indicating good usability. The median TTNO in the control and intervention arms were 128 (IQR 73-218) minutes and 131 (IQR 73-223) minutes, respectively. The corresponding hazard ratio for TTNO was 0.99 (95% CI 0.91-1.07; P=.73). Conclusions We demonstrated strong clinical engagement with the system. We found no difference in any of the predefined patient outcomes, suggesting that the introduction of a highly usable electronic system can be achieved without impacting clinical care. Our findings contrast with previous claims that digital EWS systems are associated with improvement in clinical outcomes. Future research should investigate how digital EWS systems can be integrated with new clinical pathways adjusting staff behaviors to improve patient outcomes.
The data that support the findings of this study are available from the corresponding author upon reasonable request.
Objectives This study was undertaken to identify potential predictors of atrial fibrillation after cardiac surgery (AFACS) through a modified Delphi process and expert consensus. These will supplement predictors identified through a systematic review and cohort study to inform the development of two AFACS prediction models as part of the PARADISE project (NCT05255224). Atrial fibrillation is a common complication after cardiac surgery. It is associated with worse postoperative outcomes. Reliable prediction of AFACS would enable risk stratification and targeted prevention. Systematic identification of candidate predictors is important to improve validity of AFACS prediction tools.Design This study is a Delphi consensus exercise.Setting This study was undertaken through remote participation.Participants The participants are an international multidisciplinary panel of experts selected through national research networks.Interventions This is a two-stage consensus exercise consisting of generating a long list of variables, followed by refinement by voting and retaining variables selected by at least 40% of panel members.Results The panel comprised 15 experts who participated in both stages, comprising cardiac intensive care physicians (n=3), cardiac anaesthetists (n=2), cardiac surgeons (n=1), cardiologists (n=4), cardiac pharmacists (n=1), critical care nurses (n=1), cardiac nurses (n=1) and patient representatives (n=2). Our Delphi process highlighted candidate AFACS predictors, including both patient factors and those related to the surgical intervention. We generated a final list of 72 candidate predictors. The final list comprised 3 demographic, 29 comorbidity, 4 vital sign, 13 intraoperative, 10 postoperative investigation and 13 postoperative intervention predictors.Conclusions A Delphi consensus exercise has the potential to highlight predictors beyond the scope of existing literature. This method proved effective in identifying a range of candidate AFACS predictors. Our findings will inform the development of future AFACS prediction tools as part of the larger PARADISE project.Trial registration number NCT05255224.
In the rapidly changing healthcare landscape, the implementation of offline reinforcement learning (RL) in dynamic treatment regimes (DTRs) presents a mix of unprecedented opportunities and challenges. This position paper offers a critical examination of the current status of offline RL in the context of DTRs. We argue for a reassessment of applying RL in DTRs, citing concerns such as inconsistent and potentially inconclusive evaluation metrics, the absence of naive and supervised learning baselines, and the diverse choice of RL formulation in existing research. Through a case study with more than 17,000 evaluation experiments using a publicly available Sepsis dataset, we demonstrate that the performance of RL algorithms can significantly vary with changes in evaluation metrics and Markov Decision Process (MDP) formulations. Surprisingly, it is observed that in some instances, RL algorithms can be surpassed by random baselines subjected to policy evaluation methods and reward design. This calls for more careful policy evaluation and algorithm development in future DTR works. Additionally, we discussed potential enhancements toward more reliable development of RL-based dynamic treatment regimes and invited further discussion within the community. Code is available at https://github.com/GilesLuo/ReassessDTR.