OBJECTIVE:To determine if hyperinflammatory and hypoinflammatory pediatric acute respiratory distress syndrome (PARDS) subphenotypes defined using serum biomarkers can be determined solely from electronic health record (EHR) data using machine learning. DESIGN:Retrospective, exploratory analysis using data from 2014 to 2022. SETTING:Single-center quaternary care PICU. PATIENTS:Two temporally distinct cohorts of PARDS patients, 2014-2019 and 2019-2022. INTERVENTIONS:None. MEASUREMENTS AND MAIN RESULTS:Patients in the derivation cohort ( n = 333) were assigned to hyperinflammatory or hypoinflammatory subphenotypes using biomarkers and latent class analysis. A machine learning model was trained on 165 EHR-derived variables to identify subphenotypes. The most important variables were selected for inclusion in a parsimonious model. The model was validated in a separate cohort ( n = 114). The EHR-based classifier achieved an area under the receiver operating characteristic curve (AUC) of 0.93 (95% CI, 0.87-0.98), with a sensitivity of 88% and specificity of 83% for determining hyperinflammatory PARDS. The parsimonious model, using only five laboratory values, achieved an AUC of 0.92 (95% CI, 0.86-0.98) with a sensitivity of 76% and specificity of 87% in the validation cohort. CONCLUSIONS:This proof-of-concept study demonstrates that biomarker-based PARDS subphenotypes can be identified using EHR data at 24 hours of PARDS diagnosis. Further validation in larger, multicenter cohorts is needed to confirm the clinical utility of this approach.
To review pediatric artificial intelligence (AI) implementation studies from 2010 to 2021 and analyze reported performance measures.We searched PubMed/Medline, Embase CINHAL, Cochrane Library CENTRAL, IEEE, and Web of Science with controlled vocabulary. Inclusion criteria: AI intervention in a pediatric clinical setting that learns from data (i.e., data-driven, as opposed to rule-based) and takes actions to make patient-specific recommendations; published between 01/2010 and 10/2021; must have agency (AI must provide guidance that affects clinical care, not merely running in the background). We extracted study characteristics, target users, implementation setting, time span, and performance measures.Of 126 articles reviewed as full text, 17 met inclusion criteria. Eight studies (47%) reported both clinical outcomes and process measures, six (35%) reported only process measures and two (12%) reported only clinical outcomes. Five studies (30%) reported no difference in clinical outcomes with AI, four (24%) reported improvement in clinical outcomes compared with controls, two (12%) reported positive effects on clinical outcomes with use of AI but had no formal comparison or controls, and one (6%) reported poor clinical outcomes with AI. Twelve studies (71%) reported improvement in process measures, while two (12%) reported no improvement. Five (30%) studies reported on at least 1 human performance measure.While there are many published pediatric AI models, the number of AI implementations is minimal with no standardized reporting of outcomes, care processes, or human performance measures. More comprehensive evaluations will help elucidate mechanisms of impact.
Objective To assess the influence of an implemented artificial intelligence model predicting pediatric sepsis (defined by IPSO—Improving Pediatric Sepsis Outcomes collaborative) in the emergency department (ED) on human performance measures. Materials and Methods Two ED sites within a large pediatric health system in the Southeastern United States between January 1, 2021 and April 1, 2024. We interviewed ED providers and nurses within 72 hours of caring for a patient identified as potentially having sepsis by the predictive model. Thematic analysis of qualitative data was combined with electronic health record queries to assess measures of human performance, including situation awareness, explainability, human-computer agreement, workload, trust, automation bias, and relationship between staff and patients. Results We interviewed 40 clinicians. Participants found that the sepsis alert improved situation awareness, leading to changes in patient care management, resource allocation, and/or monitoring. Participants reported an average trust in the model-based alert of 3.8/5. Only 28% (555/1977) of sepsis huddles were done without alert firing, suggesting some automation bias. Treatment with antibiotics for IPSO sepsis cases was similar pre- and post-intervention without a huddle (9.3% vs 10.5%), though treatment doubled with huddle intervention (22.7%). NASA Task Load Index increased from 43 to 57 post-intervention. There was no report of adverse relationships with patients post-intervention. Discussion Human performance appeared to be generally positive with improved situation awareness and satisfaction with the alert-driven huddle. However, there was some evidence of automation bias and a slight increase in workload with the intervention Conclusion This study demonstrates the feasibility of evaluating multiple dimensions of human performance using a mixed methods approach for an AI model implemented in clinical practice. Future studies should aim to reduce the measurement burden of human performance metrics associated with AI implementation in acute care settings and assess the correlation between human performance measures and clinical outcomes.
We present UNIPHY+, a unified physiological foundation model (physioFM) framework designed to enable continuous human health and diseases monitoring across care settings using ubiquitously obtainable physiological data. We propose novel strategies for incorporating contextual information during pretraining, fine-tuning, and lightweight model personalization via multi-modal learning, feature fusion-tuning, and knowledge distillation. We advocate testing UNIPHY+ with a broad set of use cases from intensive care to ambulatory monitoring in order to demonstrate that UNIPHY+ can empower generalizable, scalable, and personalized physiological AI to support both clinical decision-making and long-term health monitoring.
Central line-associated bloodstream infections (CLABSIs) are associated with substantial pediatric morbidity and mortality. The capacity to predict which children with central lines are at greatest risk of CLABSI could inform surveillance and prevention efforts. Our team previously published in silico predictive models for CLABSI.To prospectively implement a pediatric CLABSI predictive model and achieve adequate performance in offline validation for implementation in clinical practice.Most performant predictive models were deep learning models requiring substantial pre-processing of many features into 8-hour windows including the current day and up to 56 days prior for the current admission. To replicate this pre-processing, we created a novel infrastructure to (1) organize current-day data for all the relevant features and (2) create a staged historical data store for those same features with application programming interfaces to connect the two. We compared predictive performance of these scores for CLABSI in the next 48 hours with two labels, one based on manual review of positive blood cultures in children with central lines and another based on positive blood culture and receipt of at least 4 days of new IV antibiotics.The area under the receiver-operating characteristic (AUROC) fell from 0.97 from retrospective data to <0.60 despite multiple iterations of troubleshooting. Primary root causes included train/serve skew, feature leakage, and overfitting. Hypothesized secondary drivers were complex model specification, poor data governance, inadequate testing, challenging feature translation between real-time and historical data models, limited monitoring and logging infrastructure for troubleshooting, and suboptimal handoff between the model development and deployment teams.Bridging the gap from predictive model development to clinical deployment requires early and close coordination between data governance, data science, clinical informatics, and implementation engineers. Balancing predictive performance with implementation feasibility can accelerate the adoption of predictive clinical decision support systems.
OBJECTIVES:To develop and externally validate an intubation prediction model for children admitted to a PICU using objective and routinely available data from the electronic medical records (EMRs). DESIGN:Retrospective observational cohort study. SETTING:Two PICUs within the same healthcare system: an academic, quaternary care center (36 beds) and a community, tertiary care center (56 beds). PATIENTS:Children younger than 18 years old admitted to a PICU between 2010 and 2022. INTERVENTIONS:None. MEASUREMENTS AND MAIN RESULTS:Clinical data was extracted from the EMR. PICU stays with at least one mechanical ventilation event (>= 24 hr) occurring within a window of 1-7 days after hospital admission were included in the study. Of 13,208 PICU stays in the derivation PICU cohort, 1,175 (8.90%) had an intubation event. In the validation cohort, there were 1,165 of 17,841 stays (6.53%) with an intubation event. We trained a Categorical Boosting (CatBoost) model using vital signs, laboratory tests, demographic data, medications, organ dysfunction scores, and other patient characteristics to predict the need of intubation and mechanical ventilation using a 24-hour window of data within their hospital stay. We compared the CatBoost model to an extreme gradient boost, random forest, and a logistic regression model. The area under the receiving operating characteristic curve for the derivation cohort and the validation cohort was 0.88 (95% CI, 0.88-0.89) and 0.92 (95% CI, 0.91-0.92), respectively. CONCLUSIONS:We developed and externally validated an interpretable machine learning prediction model that improves on conventional clinical criteria to predict the need for intubation in children hospitalized in a PICU using information readily available in the EMR. Implementation of our model may help clinicians optimize the timing of endotracheal intubation and better allocate respiratory and nursing staff to care for mechanically ventilated children.
To monitor trends for late recognition of deterioration, developing a visual analytics dashboard helps address clinical deterioration rates in pediatric health systems. The dashboard enables ongoing trend analysis and detecting outliers in patient demographics or hospital units through control chart implementation using control limits set at 3 standard errors. The deterioration outcomes are defined by published evidence where EHR documentation can inform the cohort needing ICU interventions within a specific time.
BACKGROUND The molecular signature of pediatric acute respiratory distress syndrome (ARDS) is poorly described, and the degree to which hyperinflammation or specific tissue injury contributes to outcomes is unknown. Therefore, we profiled inflammation and tissue injury dynamics over the first 7 days of ARDS, and associated specific biomarkers with mortality, persistent ARDS, and persistent multiple organ dysfunction syndrome (MODS).METHODS In a single-center prospective cohort of intubated pediatric patients with ARDS, we collected plasma on days 0, 3, and 7. Nineteen biomarkers reflecting inflammation, tissue injury, and damage-associated molecular patterns (DAMPs) were measured. We assessed the relationship between biomarkers and trajectories with mortality, persistent ARDS, or persistent MODS using multivariable mixed effect models.RESULTS In 279 patients (64 [23%] nonsurvivors), hyperinflammatory cytokines, tissue injury markers, and DAMPs were higher in nonsurvivors. Survivors and nonsurvivors showed different biomarker trajectories. IL-1α, soluble tumor necrosis factor receptor 1, angiopoietin 2 (ANG2), and surfactant protein D increased in nonsurvivors, while DAMPs remained persistently elevated. ANG2 and procollagen type III N-terminal peptide were associated with persistent ARDS, whereas multiple cytokines, tissue injury markers, and DAMPs were associated with persistent MODS. Corticosteroid use did not impact the association of biomarker levels or trajectory with mortality.CONCLUSIONS Pediatric ARDS survivors and nonsurvivors had distinct biomarker trajectories, with cytokines, endothelial and alveolar epithelial injury, and DAMPs elevated in nonsurvivors. Mortality markers overlapped with markers associated with persistent MODS, rather than persistent ARDS.FUNDING NIH (K23HL-136688, R01-HL148054).
ABSTRACTBackgroundExposure to patients and clinical diagnoses drives learning in graduate medical education (GME). Measuring practice data, how trainees each experience that exposure, is critical to planned learning processes including assessment of trainee needs. We previously developed and validated an automated system to accurately identify resident provider-patient interactions (rPPIs). In this follow-up study, we employ user-centered design methods to meet two objectives: 1) understand trainees’ planned learning needs; 2) design, build, and assess a usable, useful, and effective tool based on our automated rPPI system to meet these needs.MethodsWe collected data at two institutions new to the American Medical Association’s “Advancing Change” initiative, using a mixed-methods approach with purposive sampling. First, interviews and formative prototype testing yielded qualitative data which we analyzed with several coding cycles. These qualitative methods illuminated the work domain, broke it into learning use cases, and identified design requirements. Two theoretical models—the Systems Engineering Initiative for Patient Safety (SEIPS) and Master-Adaptive Learner (MAL)—structured coding efforts. Feature-prioritization matrix analysis then transformed qualitative analysis outputs into actionable prototype elements that were refined through formative usability methods. Lastly, qualitative data from a summative usability test validated the final prototype with measures of usefulness, usability, and intent to use. Quantitative methods measured time on task and task completion rate.ResultsWe represent GME work domain learnings through process-map-design artifacts which provide target opportunities for intervention. Of the identified decision-making opportunities, trainee-mentor meetings stood out as optimal for delivering reliable practice-area information. We designed a “mid-point” report for the use case of such meetings, integrating features from qualitative analysis and formative prototype testing into iterations of the prototype. A final version showed five essential visualizations. Usability testing resulted in high performance in subjective and objective metrics. Compared to currently available resources, our tool scored 50% higher in terms of Perceived Usability and 60% higher on Perceived Ease of Use.ConclusionsWe describe the multi-site development of a tool providing visualizations of log level electronic health record data, using human-centered design methods. Delivered at an identified point in graduate medical education, the tool is ideal for fostering the development of master adaptive learners. The resulting prototype is validated with high performance on a summative usability test. Additionally, the design, development, and assessment process may be applied to other tools and topics within medical education informatics.
Exposure to patients and clinical diagnoses drives learning in graduate medical education (GME). However, variation exists in the breadth of experiences. Measuring such variation would provide practice data to inform residents’ understanding of the breadth of their patient experiences. We have developed an automated system to identify resident provider-patient interactions (rPPIs) and demonstrated accurate attribution at a single institution. The objective of this study was to understand the landscape of trainee planned learning, and iteratively design a tool to be used for this goal. To achieve these objectives at two institutions new to the AMA “Advancing Change” initiative, we used a mixed-methods approach to develop and evaluate a “mid-point report” of patients encounters. Qualitative outcomes include a guided exploration of usefulness, usability, and intent to use, as well as understanding the resources trainees would use for learning and how our system may deliver these resources. Quantitative outcomes from a summative usability test of the midpoint report will include time on task, task completion rate, and proportion of trainees who perceive the report to be useful to identify gaps in clinical experiences and guide learning.
Introduction: Clinical decision support (CDS) systems are intended to improve adherence to standard practices, enhance awareness, and increase care quality. Excessive passive decision support, such as laboratory result highlighting, can reduce CDS effectiveness. We quantify result highlighting burden across two institutions, and measure change in highlighting burden when applying published data-driven cutoffs to identify abnormal laboratory values. Methods: This multi-site retrospective study includes patients admitted to the pediatric ICU between 9/2016 and 9/2019. We describe the frequency of abnormal laboratory result highlighting by laboratory group, compared across sites. We use a logistic regression model to analyze the relationship between abnormal result highlighting to selected covariates and report odds ratios. We apply modified cutoff values based on mortality odds ratios (MOR) [Pollack et al, Peds Crit Care Med, 2021], indicating “abnormal” results when MOR >2.0. We re-calculate the frequency of abnormal result highlighting and compare to existing cutoffs. Results: We report a total of 19,087 ICU encounters and 12,507 unique patients. Across all laboratory groups, abnormal result highlighting differed significantly by institution (A: 54% vs B: 50%, p < 0.001). Percent of results with passive alerts was greatest in coagulation studies (A: 77%, B: 36%) and least in basic chemistry panels (A: 47%, B: 43%). In a logistic regression model, the only covariate consistently associated with an abnormal result was “collected < 24 hours from ICU admission,” for which odds of an abnormal result decreased [A: 12% decrease; B: 15% decrease]. When MOR-based cutoffs were applied, abnormal result highlight significantly decreased across all laboratory groups by >14%, except for complete blood counts (A: 14% increase, B: 16% increase). Conclusions: Abnormal result highlighting is a pervasive form of passive decision support that is differentially present across PICUs. Modifying laboratory result thresholds using data-driven MOR-based cutoffs reduces this passive CDS burden, which may better allow providers to identify physiologic signals from data noise. This work represents a first step in improving upon passive CDS to improve patient care quality and safety.
Introduction: In an effort to mitigate heterogeneity in Acute Respiratory Distress Syndrome (ARDS), biomarker-based sub-phenotyping has been used. Two distinct ARDS sub-phenotypes (hypoinflammatory and hyperinflammatory) have been identified in multiple adult and pediatric cohorts using a combination of biomarkers and clinical variables, with some evidence of differential response to therapeutic interventions according to sub-phenotype. A limit to this approach is the cost and lack of immediate availability of biomarkers. Machine learning algorithms can predict ARDS sub-phenotypes in adults using exclusively clinical data. We aimed to determine if machine learning can identify biomarker-based ARDS sub-phenotypes in children using readily available clinical data. Methods: This was a retrospective cohort study of 333 patients ≤18 years of age admitted to a quaternary-care PICU between 2014-2019 who met Berlin criteria for ARDS. Patients were assigned hypo- or hyperinflammatory based on biomarker-based latent class analysis (LCA). Electronic Health Record (EHR) data was incorporated into a gradient-boosted machine learning algorithm to develop a classifier model from 157 predictor variables including demographics, vital signs, vasoactive use, lab values, and oxygen requirements. Training and validation were performed using K-fold cross validation. Our primary outcome was the ability of EHR data in the first 24 hours of ARDS to discriminate between hypo- and hyperinflammatory sub-phenotypes. Results: Using EHR data in the first 24 hours of diagnosis, the model accurately classified the cohort with an Area Under the Receiver Operating Characteristic (AUROC) curve of 0.91 (95% CI, 0.88-0.94), a sensitivity of 70% and specificity of 91%. In sensitivity analyses using either 12 hours (AUROC 0.89, 95% CI 0.85-0.93) or 36 hours of clinical data (AUROC 0.92, 95% CI 0.88-0.96), the model continued to perform well. The most significant variables for discrimination in order of importance were median prothrombin time, median INR, mean lactate, minimum WBC, and minimum procalcitonin. Conclusions: Biomarker-based pediatric ARDS sub-phenotypes can be identified using routinely available EHR data in an acute timeframe, improving the ability to leverage sub-phenotypes for trial enrichment.