OBJECTIVE:Beta-blockers are commonly prescribed for chronic cardiovascular diseases. Despite potential benefits in septic shock, beta-blockers are often held at hospital admission for patients with suspected infection and possible sepsis. We compared the effects of chronic beta-blocker continuation vs. discontinuation on 90-day all-cause mortality among patients admitted from the emergency department with suspected infection. DESIGN:Retrospective cohort study using the target trial emulation framework. We used Cox regression to compare 90-day mortality between treatment groups, with inverse probability of treatment weights to account for baseline differences in sex, race, ethnicity, age, body mass index, presence of a "do not resuscitate" order, comorbidities, and acute illness severity. SETTING:A single large, academic, tertiary care emergency department in the Midwest United States. PATIENTS:Patients 18 years or older on beta-blockers prior to admission hospitalized for suspected infection (defined by orders for blood cultures and broad-spectrum antibiotics). Patients with shock, heart rates less than 40 or greater than 120, or who required an IV beta- or calcium channel blocker at a clinician's discretion were excluded. INTERVENTIONS:Continuation of oral beta-blockers within 48 hours of admission vs. no continuation. MEASUREMENTS AND MAIN RESULTS:Of 4635 eligible patients, 1172 (25.3%) received an oral beta-blocker, whereas 3463 (74.7%) did not receive an oral beta-blocker. Beta-blocker continuation was associated with a reduced risk of all-cause mortality within 90 days of hospital admission (hazard ratio 0.77; 95% CI, 0.61-0.98; p = 0.03) and shorter hospital stay (incidence rate ratio 0.39; 95% CI, 0.38-0.41; p < 0.001). There was no significant association between beta-blocker continuation and in-hospital mortality (odds ratio 0.60; 95% CI, 0.30-1.20; p = 0.15). CONCLUSIONS:Continuation of chronic beta-blockers in a broad population of patients admitted with suspected infection was associated with improved clinical outcomes. Our findings support the need for controlled experimental studies evaluating the role of chronic beta-blocker continuation among patients hospitalized with possible sepsis.
Predictive artificial intelligence (AI) models enhance clinical workflows with applications such as prognostication and decision support, yet suffer from postdeployment performance challenges due to dataset shifts. Regulatory guidelines emphasize the need for continuous monitoring, but actionable strategies are lacking. A significant issue is postdeployment assessment of predictive AI models due to confounding medical interventions where effective interventions modify outcomes, introducing bias into performance assessment. This can falsely suggest model decay, leading to unwarranted updates or decommissioning, harming clinical outcomes. Proposed solutions include withholding model outputs, monitoring outcomes as surrogates, or including clinician interventions in models, each with ethical or practical limitations. The lack of effective solutions for this problem can lead to an abundance of models that cannot be later evaluated, tuned, or withdrawn if they become ineffective, leading to patient harm. Advanced causal modeling to assess counterfactual outcomes may offer a reliable validation method. Until effective methods for postdeployment monitoring of predictive models are developed and validated, decisions on model updates should consider the causal pathways and be evidence based, ensuring the sustained utility of AI models in dynamic clinical environments.
Objectives/Goals: The objective of this study is to explore strategies for AI-physician collaboration in diagnosing acute respiratory distress syndrome (ARDS) using chest X-rays. By comparing the diagnostic accuracy of different AI deployment methods, the study aims to identify optimal strategies that leverage both AI and physician expertise to improve outcomes. Methods/Study Population: The study analyzed 414 frontal chest X-rays from 115 patients hospitalized between August 15 and October 2, 2017, at the University of Michigan. Each X-ray was reviewed by six physicians for ARDS presence and diagnostic confidence. We developed a deep learning AI model for detecting ARDS and explored the strengths, weaknesses, and blind spots of both physicians and AI systems to inform optimal system deployment. We then investigated several AI-physician collaboration strategies, including: 1) AI-aided physician: physicians interpret chest X-rays first and defer to the AI model if uncertain, 2) physician-aided AI: the AI model interprets chest X-rays first and defers to a physician if uncertain, and 3) AI model and physician interpreting chest X-rays separately and then averaging their interpretations. Results/Anticipated Results: While the AI model (84.7% accuracy) had higher accuracy than physicians (80.8%), we found evidence that AI and physician expertise are complementary. When physicians lacked confidence in a chest X-ray’s interpretation, the AI model had higher accuracy. Conversely, in cases of AI uncertainty, physicians were more accurate. The AI excelled with easier cases, while physicians were better with difficult cases, defined as those where at least two physicians disagreed with the majority label. Collaboration strategies tested include AI-aided physician (82.4%), physician-aided AI (86.9%), and averaging interpretations (86%). The physician-aided AI approach had the highest accuracy, could off-load the human expert workload on the reading of up to 79% chest X-rays, allowing physicians to focus on challenging cases. Discussion/Significance of Impact: This study shows AI and physicians complement each other in ARDS diagnosis, improving accuracy when combined. A physician-aided AI strategy, where the AI defers to physicians when uncertain, proved most effective. Implementing AI-physician collaborations in clinical settings could enhance ARDS care, especially in low-resource environments.
Background: Brain injury is a major cause of death and disability after cardiac arrest (CA). Quantitative histology enables the assessment of neuronal damage in preclinical models and CA patients. However, manual quantification of labeled neurons is time-consuming and variable, posing a significant challenge to the comprehensive assessment of the brain injury severity and neuroprotective therapy's efficacy. Hypothesis: Machine learning (ML) approach can identify and quantify neuronal damage from brain histological images of swine CA models with comparable accuracy to human raters. Methods: We developed a swine CA model to simulate out-of-hospital CA with 5 or 10 minutes of untreated ventricular fibrillation. Following 24 hours of standardized post-CA care, the animals were euthanized by transcardial perfusion with 4% paraformaldehyde. The brains were post-fixed, cryoprotected, and cryosectioned. Coronal sections (20 μm) containing the caudate putamen were stained with Fluoro-Jade C to label injured neurons. Three blinded human raters quantified Fluoro-Jade C-positive neurons in 136 images from 15 animals. These images were split into training (n=54), validation (n=27), and testing (n=55) sets for ML model development. We compared transfer learning models including VGG16, MobileNetV2, DeepLabv3+, and SegFormer. Model performance was evaluated on individual cells via precision, recall, and F1-score, then by comparing cell counts to the human raters for the best performing model. Results: Human raters showed strong reliability in image-wise counts of Fluoro-Jade C neurons with an average pairwise correlation coefficient of R=0.936. The SegFormer model demonstrated the best performance, with a test-set R=0.989 compared to neurons identified by 2 of 3 human raters, or an average R=0.967 when compared to each rater individually (Figures 1 and 2). On an individual cell level, the model yielded a precision of 0.789, a recall of 0.709, and an F1-score of 0.747 (Table 1). Conclusions: We developed and validated a reliable automated ML approach to quantify neuronal damage after CA in a swine model. Future studies will focus on validating the ML models for other brain regions, other stainings, and application in quantitative histology for CA patients.
Rationale: Machine learning models can identify bilateral airspace disease on chest radiographs consistent with ARDS with high accuracy. It is unknown whether such predictions could serve as a lung imaging biomarker that also quantifies severity of acute lung disease. We sought to determine whether machine learning model predictions of bilateral airspace disease predict clinical outcomes beyond what is captured by traditional clinical indicators of hypoxemia severity in patients receiving invasive mechanical ventilation. Methods: We applied a previously developed deep learning model to chest radiographs performed after intubation in patients receiving invasive mechanical ventilation at Michigan Medicine between 2019 and 2021. The model generates a probability that a chest radiograph has bilateral airspace disease consistent with ARDS. Model probabilities were transformed to a linear ARDS-CXR score using the logit function, with positive scores representing an ARDS probability above 50%. We calculated the AUROC for the ARDS-CXR score and PaO2/FiO2 for 28-day morality for the overall populations and in relevant patient subgroups. We also evaluated the score's association with 28-day mortality and time to extubation after adjusting for PaO2/FiO2 or APACHE-IV. Results: 6,591 patients who received invasive mechanical ventilation were analyzed. Their average ARDS-CXR score was -1.72, corresponding to an ARDS probability of 15%, and 1,089 (17%) patients had a positive score. The ARDS-CXR was strongly associated with 28-day mortality (Figure), and had an AUROC for mortality of 0.68 (95% CI 0.66 – 0.70) compared to 0.63 (95% CI 0.61-0.65) for PaO2/FiO2. ARDS-CXR scores also had higher discrimination of mortality than the PaO2/FiO2 in post-operative and non-operative patients, and in patients with sepsis or heart failure. A one-point increase in ARDS-CXR score was associated with a 28-day mortality odds ratio of 1.35 (95% CI, 1.30-1.40) after adjusting for the patient's concurrent PaO2/FiO2. The ARDS-CXR score was significantly associated with 28-day mortality after adjusting for APACHE-IV score, with an odds ratio of 1.17 (95% CI, 1.12-1.22). The score also predicted time to extubation after adjusting for APACHE-IV, with a one-point increase in score associated with a hazard ratio of 0.91 (95% CI, 0.89–0.92) for successful extubation. Conclusions: A machine-learning model identifying ARDS findings on chest radiographs can also capture acute lung disease severity and independently predicts clinical outcomes. Such machine-learning tools could potentially quantify lung disease severity for clinical care, research studies, or used as part of a future ARDS definition.
Disease risk prediction models play an important role in preventing disease developments in modern healthcare. However, the lack of focus on high-risk patients has hindered the large-scale practical application of these models, especially considering the limitation of medical resources available for following up on patients who are deemed high-risk. In this study, we propose a novel and practical approach that focuses on minimizing the number of false positive observations among high-risk patients by introducing the Highest-k Loss. The solution is to estimate the weights of the highest k scores with a differentiable estimation of the sorting operation and apply the weights to the loss function. We extracted 253,680 survey responses from a public dataset of the U.S. health survey system to define a diabetes prediction task. This study employs nested cross-validation as well as an aggregated model applied to an independent test set to systematically evaluate the proposed method. Compared with traditional binary cross entropy loss and Focal loss, the Highest-k loss improved the precision (positive predictive value) for the highest 1% scores by 0.05 (95% CI: 0.041-0.055), the highest 5% scores by 0.03 (95% CI: 0.024-0.032), and the highest 10% scores by 0.02 (95% CI: 0.016-0.021). The introduced Highest-k loss function addresses the problem of prevailing risk prediction models and offers a practical solution that focuses on patients with the k highest predictive scores who can realistically receive an intervention as opposed to the entire patient population.
Background & Aims Paracentesis is commonly used to manage patient discomfort due to ascites. The relationship between ascites pressure, ascites volume, and patient discomfort has not been elucidated. Methods We prospectively enrolled adult patients with non-malignant ascites undergoing outpatient therapeutic paracenteses from 2021 to 2024 at a tertiary care hospital. Patients completed a validated symptom questionnaire (ASI-7, maximum score 35) before, immediately after, and 1 week after paracentesis. An open-ended manometer was used to measure ascites pressure at the beginning and end of paracentesis. Mixed effect linear regression was performed to evaluate the relationships between patient characteristics, pressure, volume, and symptoms. Results One hundred and fifty paracentesis procedures among 48 unique patients with an average Model for End Stage Liver Disease-Sodium 3.0 of 16.7 were included. An average of 6.5 L was drained, which reduced abdominal pressure from a mean of 13.7 to 6.0 cm H2O (10.1 to 4.4 mmHg, p < 0.001) and mean symptom score from 22.6 to 6.5 (p < 0.001). Regression models identified that symptoms and abdominal pressure linearly correlated above a pressure of 6 cm H2O or ASI-7 score of 16 (p < 0.01). Taller patients required about 670 ml additional drainage per inch above the cohort mean height (5 ' 8 '') to achieve the same symptom relief. Conclusions Pressure measured at the bedside can be used to explore changes in abdominal pressure during paracentesis. Pressure, volume, and patient level factors such as height contribute to patient symptoms but cannot fully explain discomfort associated with ascites and relief after paracentesis.
Purpose Pediatric acute respiratory distress syndrome (PARDS) is underrecognized in the pediatric intensive care unit and the interpretation of chest radiographs is a key step in identification. We sought to test the performance of a machine learning model to detect PARDS in a cohort of children with respiratory failure. Materials and methods A convolutional neural network (CNN) model previously developed to detect ARDS on adult chest radiographs was applied to a cohort of children age 7 days to 18 years, admitted to the PICU, and mechanically ventilated through a tracheostomy, endotracheal tube or full-face non-invasive positive pressure mask between May 2016 and January 2017. Two pediatric critical care physicians and a pediatric radiologist reviewed chest radiographs to evaluate if the chest radiographs were consistent with ARDS (bilateral airspace disease) and PARDS (any airspace disease) and the CNN model was tested against clinicians. Results A total of 328 chest radiographs were evaluated from 66 patients. Clinicians identified 84% (276/328) of the radiographs as potentially consistent with PARDS. Inter-rater reliability between individual clinicians and between the model and clinicians was similar (Cohen’s kappa 0.48 [95% CI 0.37–0.59] and 0.45 [95% CI 0.33–0.57], respectively). The model was better at identifying PARDS (AUC 0.882, F1 0.897) than ARDS (AUC 0.842, F1 0.742) and had equivalent or better performance to individual clinicians. Conclusions An ARDS detection model trained on adults performed well in detecting PARDS in children. Computer-assisted identification of PARDS on chest radiographs could improve the diagnosis of PARDS for enrollment in clinical trials and application of PARDS guidelines through improved diagnosis.
ImportanceBreath analysis has been explored as a noninvasive means to detect COVID-19. However, the impact of emerging variants of SARS-CoV-2, such as Omicron, on the exhaled breath profile and diagnostic accuracy of breath analysis is unknown.ObjectiveTo evaluate the diagnostic accuracies of breath analysis on detecting patients with COVID-19 when the SARS-CoV-2 Delta and Omicron variants were most prevalent.Design, Setting, and ParticipantsThis diagnostic study included a cohort of patients who had positive and negative test results for COVID-19 using reverse transcriptase polymerase chain reaction between April 2021 and May 2022, which covers the period when the Delta variant was overtaken by Omicron as the major variant. Patients were enrolled through intensive care units and the emergency department at the University of Michigan Health System. Patient breath was analyzed with portable gas chromatography.Main Outcomes and MeasuresDifferent sets of VOC biomarkers were identified that distinguished between COVID-19 (SARS-CoV-2 Delta and Omicron variants) and non-COVID-19 illness.ResultsOverall, 205 breath samples from 167 adult patients were analyzed. A total of 77 patients (mean [SD] age, 58.5 [16.1] years; 41 [53.2%] male patients; 13 [16.9%] Black and 59 [76.6%] White patients) had COVID-19, and 91 patients (mean [SD] age, 54.3 [17.1] years; 43 [47.3%] male patients; 11 [12.1%] Black and 76 [83.5%] White patients) had non-COVID-19 illness. Several patients were analyzed over multiple days. Among 94 positive samples, 41 samples were from patients in 2021 infected with the Delta or other variants, and 53 samples were from patients in 2022 infected with the Omicron variant, based on the State of Michigan and US Centers for Disease Control and Prevention surveillance data. Four VOC biomarkers were found to distinguish between COVID-19 (Delta and other 2021 variants) and non-COVID-19 illness with an accuracy of 94.7%. However, accuracy dropped substantially to 82.1% when these biomarkers were applied to the Omicron variant. Four new VOC biomarkers were found to distinguish the Omicron variant and non-COVID-19 illness (accuracy, 90.9%). Breath analysis distinguished Omicron from the earlier variants with an accuracy of 91.5% and COVID-19 (all SARS-CoV-2 variants) vs non-COVID-19 illness with 90.2% accuracy.Conclusions and RelevanceThe findings of this diagnostic study suggest that breath analysis has promise for COVID-19 detection. However, similar to rapid antigen testing, the emergence of new variants poses diagnostic challenges. The results of this study warrant additional evaluation on how to overcome these challenges to use breath analysis to improve the diagnosis and care of patients.
There is a growing gap between studies describing the capabilities of artificial intelligence (AI) diagnostic systems using deep learning versus efforts to investigate how or when to integrate AI systems into a real-world clinical practice to support physicians and improve diagnosis. To address this gap, we investigate four potential strategies for AI model deployment and physician collaboration to determine their potential impact on diagnostic accuracy. As a case study, we examine an AI model trained to identify findings of the acute respiratory distress syndrome (ARDS) on chest X-ray images. While this model outperforms physicians at identifying findings of ARDS, there are several reasons why fully automated ARDS detection may not be optimal nor feasible in practice. Among several collaboration strategies tested, we find that if the AI model first reviews the chest X-ray and defers to a physician if it is uncertain, this strategy achieves a higher diagnostic accuracy (0.869, 95% CI 0.835–0.903) compared to a strategy where a physician reviews a chest X-ray first and defers to an AI model if uncertain (0.824, 95% CI 0.781–0.862), or strategies where the physician reviews the chest X-ray alone (0.808, 95% CI 0.767–0.85) or the AI model reviews the chest X-ray alone (0.847, 95% CI 0.806–0.887). If the AI model reviews a chest X-ray first, this allows the AI system to make decisions for up to 79% of cases, letting physicians focus on the most challenging subsets of chest X-rays.
OBJECTIVES: Implementing a predictive analytic model in a new clinical environment is fraught with challenges. Dataset shifts such as differences in clinical practice, new data acquisition devices, or changes in the electronic health record (EHR) implementation mean that the input data seen by a model can differ significantly from the data it was trained on. Validating models at multiple institutions is therefore critical. Here, using retrospective data, we demonstrate how Predicting Intensive Care Transfers and other UnfoReseen Events (PICTURE), a deterioration index developed at a single academic medical center, generalizes to a second institution with significantly different patient population. DESIGN: PICTURE is a deterioration index designed for the general ward, which uses structured EHR data such as laboratory values and vital signs. SETTING: The general wards of two large hospitals, one an academic medical center and the other a community hospital. SUBJECTS: The model has previously been trained and validated on a cohort of 165,018 general ward encounters from a large academic medical center. Here, we apply this model to 11,083 encounters from a separate community hospital. INTERVENTIONS: None. MEASUREMENTS AND MAIN RESULTS: The hospitals were found to have significant differences in missingness rates (> 5% difference in 9/52 features), deterioration rate (4.5% vs 2.5%), and racial makeup (20% non-White vs 49% non-White). Despite these differences, PICTURE’s performance was consistent (area under the receiver operating characteristic curve [AUROC], 0.870; 95% CI, 0.861–0.878), area under the precision-recall curve (AUPRC, 0.298; 95% CI, 0.275–0.320) at the first hospital; AUROC 0.875 (0.851–0.902), AUPRC 0.339 (0.281–0.398) at the second. AUPRC was standardized to a 2.5% event rate. PICTURE also outperformed both the Epic Deterioration Index and the National Early Warning Score at both institutions. CONCLUSIONS: Important differences were observed between the two institutions, including data availability and demographic makeup. PICTURE was able to identify general ward patients at risk of deterioration at both hospitals with consistent performance (AUROC and AUPRC) and compared favorably to existing metrics.
The QRS complex is the most prominent feature of the electrocardiogram (ECG) that is used as a marker to identify the cardiac cycles. Identification of QRS complex locations enables arrhythmia detection and heart rate variability estimation. Therefore, accurate and consistent localization of the QRS complex is an important component of automated ECG analysis which is necessary for the early detection of cardiovascular diseases. This study evaluates the performance of six popular publicly available QRS complex detection methods on a large dataset of over half a million ECGs in a diverse population of patients. We found that a deep-learning method that won first place in the 2019 Chinese physiological challenge (CPSC-1) outperforms the remaining five QRS complex detection methods with an F1 score of 98.8% and an absolute sdRR error of 5.5 ms. We also examined the stratified performance of the studied methods on various cardiac conditions. All six methods had a lower performance in the detection of QRS complexes in ECG signals of patients with pacemakers, complete atrioventricular block, or indeterminate cardiac axis. We also concluded that, in the presence of different cardiac conditions, CPSC-1 is more robust than Pan-Tompkins which is the most popular model for QRS complex detection. We expect that this study can potentially serve as a guide for researchers on the appropriate QRS detection method for their target applications.Clinical Relevance—This study highlights the overall performance of publicly available QRS detection algorithms in a large dataset of diverse patients. We showed that there are specific cardiac conditions that are associated with the poor performance of QRS detection algorithms and may adversely influence the performance of algorithms that rely on accurate and reliable QRS detection.
There has been a proliferation of machine learning (ML) electrocardiogram (ECG) classification algorithms reaching > 85% accuracy for various cardiac pathologies. Although the accuracy within institutions might be high, models trained at one institution might not be generalizable enough for accurate detection when deployed in other institutions due to differences in type of signal acquisition, sampling frequency, time of acquisition, device noise characteristics and number of leads. In this proof-of-concept study, we leverage the publicly available PTB-XL dataset to investigate the use of time-domain (TD) and frequency-domain (FD) convolutional neural networks (CNN) to detect myocardial infarction (MI), ST/T-wave changes (STTC), atrial fibrillation (AFIB) and sinus arrhythmia (SARRH). To simulate interinstitutional deployment, the TD and FD implementations were also compared on adapted test sets using different sampling frequencies 50 Hz, 100 Hz and 250 Hz, and acquisition times of 5 s and 10s at 100 Hz sampling frequency from the training dataset. When tested on the original sampling frequency and duration, the FD approach showed comparable results to TD for MI (0.92 FD - 0.93 TD AUROC) and STTC (0.94 FD - 0.95 TD AUROC), and better performance for AFIB (0.99 FD - 0.86 TD AUROC) and SARRH (0.91 FD - 0.65 TD AUROC). Although both methods were robust to changes in sampling frequency, changes in acquisition time were detrimental to the TD MI and STTC AUROCs, at 0.72 and 0.58 respectively. Alternatively, the FD approach was able to maintain the same level of performance, and, therefore, showed better potential for interinstitutional deployment.
Department of Internal Medicine and Department of EmergencyMedicine, Medical School, and Department of Epidemiology, School of Public Health, University of Michigan, Ann Arbor, Michigan; Veterans Affairs Center for Clinical Management Research, LTC Charles Kettles Veterans Affairs Medical Center, Ann Arbor, Michigan; and Weil Institute for Critical Care Research and Innovation, Ann Arbor, Michigan
The clinical significance of volatile organic compounds (VOC) in detecting diseases has been established over the past decades. Gas chromatography (GC) devices enable the measurement of these VOCs. Chromatographic peak alignment is one of the important yet challenging steps in analyzing chromatogram signals. Traditional semi-automated alignment algorithms require manual intervention by an operator which is slow, expensive and inconsistent. A pipeline is proposed to train a deep-learning model from artificial chromatograms simulated from a small, annotated dataset, and a postprocessing step based on greedy optimization to align the signals.Clinical Relevance— Breath VOCs have been shown to have a significant diagnostic power for various diseases including asthma, acute respiratory distress syndrome and COVID-19. Automatic analysis of chromatograms can lead to improvements in the diagnosis and management of such diseases.
Despite promising results in disease detection, breath analysis has not reached its full potential. With the rise of portable gas chromatography (GC) devices and the large volume of data, traditional manual GC peak detection methods are impractical. This work proposes a new approach to chromatographic peak detection which is a necessary step in the GC analysis pipeline. A deep-learning model was trained on simulated breath data to localize the chromatogram peaks and validated on manual annotations. The approach was specifically designed to account for the co-eluted peaks that are often overlooked. The results show that the model outperformed threshold- and derivative-based approaches, as well as other proposed models with high sensitivity and precision.