Harmonic activation and transport (HAT) is a stochastic process that rearranges finite subsets of $\mathbb{Z}^d$, one element at a time. Given a finite set $U \subset \mathbb{Z}^d$ with at least two elements, HAT removes $x$ from $U$ according to the harmonic measure of $x$ in $U$, and then adds $y$ according to the probability that simple random walk from $x$, conditioned to hit the remaining set, steps from $y$ when it first does so. In particular, HAT conserves the number of elements in $U$. We study the classification of HAT as recurrent or transient, as the dimension $d$ and number of elements $n$ in the initial set vary. It was recently proved that the stationary distribution of HAT (on sets viewed up to translation) exists when $d = 2$, for every number of elements $n \geq 2$. We prove that HAT exhibits a phase transition in both $d$ and $n$, in the sense that HAT is transient when $d \geq 5$ and $n \geq 4$. Remarkably, transience occurs in only one "way": The set splits into clusters of two or three elements, which then grow steadily, indefinitely separated. We call these clusters dimers and trimers. Underlying this characterization of transience is the fact that, from any set, HAT reaches a set consisting exclusively of dimers and trimers, in a number of steps and with at least a probability which depend on $d$ and $n$ only.
Clinical decision support tools rooted in machine learning and optimization can provide significant value to healthcare providers, including through better management of intensive care units. In particular, it is important that the patient discharge task addresses the nuanced trade-off between decreasing a patient's length of stay (and associated hospitalization costs) and the risk of readmission or even death following the discharge decision. This work introduces an end-to-end general framework for capturing this trade-off to recommend optimal discharge timing decisions given a patient's electronic health records. A data-driven approach is used to derive a parsimonious, discrete state space representation that captures a patient's physiological condition. Based on this model and a given cost function, an infinite-horizon discounted Markov decision process is formulated and solved numerically to compute an optimal discharge policy, whose value is assessed using off-policy evaluation strategies. Extensive numerical experiments are performed to validate the proposed framework using real-life intensive care unit patient data.
Natural collectives, despite comprising individuals who may not know their numerosity, can exhibit behaviors that depend sensitively on it. This paper proves that the collective behavior of number-oblivious individuals can even have a critical numerosity, above and below which it qualitatively differs. We formalize the concept of critical numerosity in terms of a family of zero--one laws and introduce a model of collective motion, called chain activation and transport (CAT), that has one. CAT describes the collective motion of $n \geq 2$ individuals as a Markov chain that rearranges $n$-element subsets of the $d$-dimensional grid, $m < n$ elements at a time. According to the individuals' dynamics, with each step, CAT removes $m$ elements from the set and then progressively adds $m$ elements to the boundary of what remains, in a way that favors the consecutive addition and removal of nearby elements. This paper proves that, if $d \geq 3$, then CAT has a critical numerosity of $n_c = 2m+2$ with respect to the behavior of its diameter. Specifically, if $n < n_c$, then the elements form one "cluster," the diameter of which has an a.s.--finite limit infimum. However, if $n \geq n_c$, then there is an a.s.--finite time at which the set consists of clusters of between $m+1$ and $2m+1$ elements, and forever after which these clusters grow apart, resulting in unchecked diameter growth. The existence of critical numerosities means that collectives can exhibit "phase transitions" that are governed purely by their numerosity and not, for example, their density or the strength of their interactions. This fact challenges prevalent beliefs about collective behavior and suggests new functionality for programmable matter. More broadly, it demonstrates an opportunity to explore the possible behaviors collectives through the study of random processes that rearrange sets.
Abstract For an n-element subset U of $\mathbb {Z}^2$ , select x from U according to harmonic measure from infinity, remove x from U and start a random walk from x. If the walk leaves from y when it first enters the rest of U, add y to it. Iterating this procedure constitutes the process we call harmonic activation and transport (HAT). HAT exhibits a phenomenon we refer to as collapse: Informally, the diameter shrinks to its logarithm over a number of steps which is comparable to this logarithm. Collapse implies the existence of the stationary distribution of HAT, where configurations are viewed up to translation, and the exponential tightness of diameter at stationarity. Additionally, collapse produces a renewal structure with which we establish that the center of mass process, properly rescaled, converges in distribution to two-dimensional Brownian motion. To characterize the phenomenon of collapse, we address fundamental questions about the extremal behavior of harmonic measure and escape probabilities. Among n-element subsets of $\mathbb {Z}^2$ , what is the least positive value of harmonic measure? What is the probability of escape from the set to a distance of, say, d? Concerning the former, examples abound for which the harmonic measure is exponentially small in n. We prove that it can be no smaller than exponential in $n \log n$ . Regarding the latter, the escape probability is at most the reciprocal of $\log d$ , up to a constant factor. We prove it is always at least this much, up to an n-dependent factor.
Many models of one-dimensional local random growth are expected to lie in the Kardar-Parisi-Zhang (KPZ) universality class. For such a model, the interface profile at advanced time may be viewed in scaled coordinates specified via characteristic KPZ scaling exponents of one-third and two-thirds. When the long time limit of this scaled interface is taken, it is expected -- and proved for a few integrable models -- that, up to a parabolic shift, the Airy$_2$ process $\mathcal{A}:\mathbb{R} \to \mathbb{R}$ is obtained. This process may be embedded via the Robinson-Schensted-Knuth correspondence as the uppermost curve in an $\mathbb{N}$-indexed system of random continuous curves, the Airy line ensemble. Among our principal results is the assertion that the Airy$_2$ process enjoys a very strong similarity to Brownian motion $B$ (of rate two) on unit-order intervals; as a consequence, the Radon-Nikodym derivative of the law of $\mathcal{A}$ on say $[-1,1]$, with respect to the law of $B$ on this interval, lies in every $L^p$ space for $p \in (1,\infty)$. Our technique of proof harnesses a probabilistic resampling or {\em Brownian Gibbs} property satisfied by the Airy line ensemble after parabolic shift, and this article develops Brownian Gibbs analysis of this ensemble begun in [CH14] and pursued in [Ham19a]. Our Brownian comparison for scaled interface profiles is an element in the ongoing programme of studying KPZ universality via probabilistic and geometric methods of proof, aided by limited but essential use of integrable inputs. Indeed, the comparison result is a useful tool for studying this universality class. We present and prove several applications, concerning for example the structure of near ground states in Brownian last passage percolation, or Brownian structure in scaled interface profiles that arise from evolution from any element in a very general class of initial data.
Background Electronic health records (EHRs) contain individualized patient data that can be used to develop diagnostic and risk prediction models with artificial intelligence (AI) algorithms. Explicit and implicit sources of bias embedded in EHRs may hinder generalizable model performance and perpetuate bias. Objective This study explores the question of how existing sex and racial disparities in the clinical assessment and management of acute coronary syndrome (ACS) are reflected in patients’ EHRs. We outlined recommendations on how these sources of bias should be examined and mitigated within the framework of development and validation of AI-based diagnosis and risk stratification algorithms. Methods This retrospective study examined several previously unrecognized EHR-embedded biases in a multisite emergency departments (ED) setting. We assessed sex and race differences in EHR data missingness and timeliness of several key ED procedures following ACS suspicion. We additionally conducted a data-driven clustering analysis to detect latent groups of ACS patients. Results 10,043 ACS-associated ED visits were included. We identified sex and race differences in the prevalence of ACS symptoms, data missingness (vitals and labs), and waiting time for troponin order and treatment initiation within 24 hours post-admission. Our cluster analysis discovered four groups of ACS patients corresponding to common symptoms. Additionally, we identified differences in clinical management and patient demographics across these clusters. Conclusions We discovered several sources of bias inherent in the clinical practice of ACS diagnosis and treatment that are also present in EHR data. Our study supports the inclusion of a wider range of symptoms into AI-based diagnosis and risk stratification tools. These models should be also validated in ED admissions without chest pain and in patients who experience a longer waiting time for ED procedures.
Abstract Background Non‐alcoholic fatty liver (NAFL) can progress to the severe subtype non‐alcoholic steatohepatitis (NASH) and/or fibrosis, which are associated with increased morbidity, mortality, and healthcare costs. Current machine learning studies detect NASH; however, this study is unique in predicting the progression of NAFL patients to NASH or fibrosis. Aim To utilize clinical information from NAFL‐diagnosed patients to predict the likelihood of progression to NASH or fibrosis. Methods Data were collected from electronic health records of patients receiving a first‐time NAFL diagnosis. A gradient boosted machine learning algorithm (XGBoost) as well as logistic regression (LR) and multi‐layer perceptron (MLP) models were developed. A five‐fold cross‐validation grid search was utilized for hyperparameter optimization of variables, including maximum tree depth, learning rate, and number of estimators. Predictions of patients likely to progress to NASH or fibrosis within 4 years of initial NAFL diagnosis were made using demographic features, vital signs, and laboratory measurements. Results The XGBoost algorithm achieved area under the receiver operating characteristic (AUROC) values of 0.79 for prediction of progression to NASH and 0.87 for fibrosis on both hold‐out and external validation test sets. The XGBoost algorithm outperformed the LR and MLP models for both NASH and fibrosis prediction on all metrics. Conclusion It is possible to accurately identify newly diagnosed NAFL patients at high risk of progression to NASH or fibrosis. Early identification of these patients may allow for increased clinical monitoring, more aggressive preventative measures to slow the progression of NAFL and fibrosis, and efficient clinical trial enrollment.
Respiratory syncytial virus (RSV) causes millions of infections among children in the US each year and can cause severe disease or death. Infections that are not promptly detected can cause outbreaks that put other hospitalized patients at risk. No tools besides diagnostic testing are available to rapidly and reliably predict RSV infections among hospitalized patients. We conducted a retrospective study from pediatric electronic health record (EHR) data and built a machine learning model to predict whether a patient will test positive to RSV by nucleic acid amplification test during their stay. Our model demonstrated excellent discrimination with an area under the receiver-operating curve of 0.919, a sensitivity of 0.802, and specificity of 0.876. Our model can help clinicians identify patients who may have RSV infections rapidly and cost-effectively. Successfully integrating this model into routine pediatric inpatient care may assist efforts in patient care and infection control.
BACKGROUND The aim of the study was to quantify the relationship between acute kidney injury (AKI) and alcohol use disorder (AUD). METHODS We used a large academic medical center and the MIMIC-III databases to quantify AKI disease and mortality burden as well as AKI disease progression in the AUD and non-AUD subpopulations. We used the MIMIC-III dataset to compare two different methods of encoding AKI: ICD-9 codes, and the Kidney Disease: Improving Global Outcomes scheme (KDIGO) definition. In addition to the AUD subpopulation, we also present analyses for the hepatorenal syndrome (HRS) and alcohol-related cirrhosis subpopulations identified via ICD-9/ICD-10 coding. RESULTS In both the ICD-9 and KDIGO encodings of AKI, the AUD subpopulation had a higher incidence of AKI (ICD-9: 43.3% vs. 37.92% AKI in the non-AUD subpopulations; KDIGO: 48.65% vs. 40.53%) in the MIMIC-III dataset. In the academic dataset, the AUD subpopulation also had a higher incidence of AKI than the non-AUD subpopulation (ICD-9/ICD-10: 12.76% vs. 10.71%). The mortality rate of the subpopulation with both AKI and AUD, HRS, or alcohol-related cirrhosis was consistently higher than that of the subpopulation with only AKI in both datasets, including after adjusting for disease severity using two methods of severity estimation in the MIMIC-III dataset. Disease progression rates were similar for AUD and non-AUD subpopulations. CONCLUSIONS Our work shows that the AUD patient subpopulation had a higher number of AKI patients than the non-AUD subpopulation, and that patients with both AKI and AUD, HRS, or alcohol-related cirrhosis had higher rates of mortality than the non-AUD subpopulation with AKI.
Background: Interventions to better prevent or manage Clostridioides difficile infection (CDI) may significantly reduce morbidity, mortality, and healthcare spending. Methods: We present a retrospective study using electronic health record data from over 700 United States hospitals. A subset of hospitals was used to develop machine learning algorithms (MLAs); the remaining hospitals served as an external test set. Three MLAs were evaluated: gradient-boosted decision trees (XGBoost), Deep Long Short Term Memory neural network, and one-dimensional convolutional neural network. MLA performance was evaluated with area under the receiver operating characteristic curve (AUROC), sensitivity, specificity, diagnostic odds ratios and likelihood ratios. Results: The development dataset contained 13,664,840 inpatient encounters with 80,046 CDI encounters; the external dataset contained 1,149,088 inpatient encounters with 7,107 CDI encounters. The highest AUROCs were achieved for XGB, Deep Long Short Term Memory neural network, and one-dimensional convolutional neural network via abstaining from use of specialized training techniques, resampling in isolation, and resampling and output bias in combination, respectively. XGBoost achieved the highest AUROC. Conclusions: MLAs can predict future CDI in hospitalized patients using just 6 hours of data. In clinical practice, a machine-learning based tool may support prophylactic measures, earlier diagnosis, and more timely implementation of infection control measures. (c) 2021 The Author(s). Published by Elsevier Inc. on behalf of Association for Professionals in Infection Control and Epidemiology, Inc. This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/)
(1) Background: Ventilator-associated pneumonia (VAP) causes high mortality among patients with respiratory disease and imposes major burdens on healthcare infrastructure. Models that use electronic health record data to predict the onset of VAP may spur earlier treatment and improve patient outcomes. We developed and studied the performance of interpretable machine learning (ML) models that predict the onset of VAP from electronic health records (EHRs); (2) Methods: We trained Logistic Regression (LR), full feature Explainable Boosting Machine (fEBM), and eXtreme Gradient Boosting (XGBoost) ML models on data from the MIMIC- III (v1.3) database. Model performance was measured by area under the receiver operating characteristic curves (AUCs). We trained a minimal-feature EBM model (mEBM) with features derived from white blood cell (WBC) counts, duration of ventilation, and Glasgow Coma Scale (GCS). Finally, model robustness was evaluated on randomly sparsified EHR datasets; (3) Results: The fEBM model outperformed the XGBoost and LR models at 24 hours post-intubation. The mEBM model maintained an AUC of 0.893. The fEBM model performance remained robust on sparsified datasets; (4) Conclusions: Our novel interpretable ML algorithm reliably predicts the onset of VAP in intubated patients. Integration of this EBM-based model into clinical practice may enable clinicians to better anticipate and prevent VAP.
Abstract Importance: Despite sex and race disparities in the symptom presentation, diagnosis, and management of acute coronary syndrome (ACS), these differences have not been investigated in the development and validation of machine learning (ML) models using individualized patient information from electronic health records (EHRs) to diagnose ACS. Objective: To evaluate ML-based ACS diagnosis performance across different subpopulations in a multi-site emergency department (ED) setting and determine how bias mitigating techniques influence ML performance. Design, Setting, and Participants: This retrospective observational study included data from 2,334,316 ED patients ( >18 years) from January 2007 to June 2020. Exposure: Logistic regression (LR) and neural network (NN) models were assessed in ED encounters grouped by sex, race, presence or absence of chest pain, EHR data quality, and timeliness of several key ED procedures. Prejudice regularization, reweighting, and within-subpopulation training were evaluated for bias mitigation. Main Outcomes/Measures: Metrics including area under the receiver operating characteristic (AUROC) were used to assess performances. Results: We analyzed 4,268,165 ED visits in which patient demographics by race were 67.40% White, 19.20% Black, 2.40% Asian, and 11.00% Other or Unknown. Patient composition was 54.80% female and 45.20% male. Both models’ AUROCs were significantly higher in White vs. Black patients (LR: z-score = 3.23 and NN: 4.26 for NN; P < 0.0006), in males vs. females (z-score = 3.81 for LR and 4.16 for NN; P < 0.0001) and in no chest pain subpopulation vs. chest pain (z-score = 13.32 for LR and 17.70 for NN; P < 0.0001). Prejudice regularization and reweighting techniques did not reduce biases. Training in race-specific and sex-specific training populations also did not yeild statistically signficant improvements in ML algorithm performance. Chest pain-specific training led to significantly improved AUROC.Conclusion: EHR-derived ML models trained and tested within similar demographic subpopulations and symptom groups may perform better than ML models that are trained in random populations, and provide less biased clinical decision support for ACS diagnosis.
Timely recognition of sepsis in hospital patients increases the likelihood of patient survival. The value of sepsis alert systems in clinical settings is diminished if these alerts are generated after clinically relevant times: after clinician suspicion of sepsis (clinical evaluation or treatment for sepsis) or onset of sepsis (defined by SOFA score). We evaluate and compare two models - standard and early - using traditional time-agnostic methods (area under the curve: AUC; true positive rate: TPR) and a time-dependent approach with redefined metrics that account for the timing of alerts; i.e., the clinical relevancy or earliness (eAUC and eTPR), for three different endpoints. When evaluating the models at the clinically relevant alert-before-sepsis-onset endpoint, the early model outperforms the standard model and achieves eAUC values on the holdout test and external validation sets of (0.892; 0.880) vs (0.832; 0.781), respectively. The early model generated more than 70% of the correct alerts within the clinically optimizable timeframe for preventive measures from 24 hours prior up to sepsis onset at 0.8 TPR. We demonstrate that time-agnostic evaluations of algorithm performance may not accurately represent the usefulness of the model in practice; while an algorithm may achieve high accuracy as measured by high AUC values, this does not describe the potential clinical relevance of the alerts.
BACKGROUND:Short-term fall prediction models that use electronic health records (EHRs) may enable the implementation of dynamic care practices that specifically address changes in individualized fall risk within senior care facilities.OBJECTIVE:The aim of this study is to implement machine learning (ML) algorithms that use EHR data to predict a 3-month fall risk in residents from a variety of senior care facilities providing different levels of care.METHODS:This retrospective study obtained EHR data (2007-2021) from Juniper Communities' proprietary database of 2785 individuals primarily residing in skilled nursing facilities, independent living facilities, and assisted living facilities across the United States. We assessed the performance of 3 ML-based fall prediction models and the Juniper Communities' fall risk assessment. Additional analyses were conducted to examine how changes in the input features, training data sets, and prediction windows affected the performance of these models.RESULTS:The Extreme Gradient Boosting model exhibited the highest performance, with an area under the receiver operating characteristic curve of 0.846 (95% CI 0.794-0.894), specificity of 0.848, diagnostic odds ratio of 13.40, and sensitivity of 0.706, while achieving the best trade-off in balancing true positive and negative rates. The number of active medications was the most significant feature associated with fall risk, followed by a resident's number of active diseases and several variables associated with vital signs, including diastolic blood pressure and changes in weight and respiratory rates. The combination of vital signs with traditional risk factors as input features achieved higher prediction accuracy than using either group of features alone.CONCLUSIONS:This study shows that the Extreme Gradient Boosting technique can use a large number of features from EHR data to make short-term fall predictions with a better performance than that of conventional fall risk assessments and other ML models. The integration of routinely collected EHR data, particularly vital signs, into fall prediction models may generate more accurate fall risk surveillance than models without vital signs. Our data support the use of ML models for dynamic, cost-effective, and automated fall predictions in different types of senior care facilities.
Background: Central line-associated bloodstream infections (CLABSIs) are associated with significant morbidity, mortality, and increased healthcare costs. Despite the high prevalence of CLABSIs in the U.S., there are currently no tools to stratify a patient's risk of developing an infection as the result of central line placement. To this end, we have developed and validated a machine learning algorithm (MLA) that can predict a patient's likelihood of developing CLABSI using only electronic health record data in order to provide clinical decision support. Methods: We created three machine learning models to retrospectively analyze electronic health record data from 27,619 patient encounters. The models were trained and validated using an 80:20 split for the train and test data. Patients designated as having a central line procedure based on International Statistical Classification of Diseases and Related Health Problems 10 codes were included. Results: XGBoost was the highest performing MLA out of the three models, obtaining an AUROC of 0.762 for CLABSI risk prediction at 48 hours after the recorded time for central line placement. Conclusions: Our results demonstrate that MLAs may be effective clinical decision support tools for assessment of CLABSI risk and should be explored further for this purpose. (c) 2021 The Author(s). Published by Elsevier Inc. on behalf of Association for Professionals in Infection Control and Epidemiology, Inc. This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/)
Moisture fluctuations in pavement foundation due to environmental conditions (e.g., heavy precipitation, freeze-thaw cycles, groundwater table variation, etc.) can significantly affect pavements' short- and long-term performance. Moisture variation in the pavement foundation is often monitored as a way to anticipate the structural capacity of pavements. Thus, the ability to monitor moisture variation through a proven nondestructive technology (NDT) such as Ground Penetrating Radar (GPR) would be beneficial for asset management by transportation agencies. This paper summarizes the Minnesota Department of Transportation's (MnDOT) efforts at validating the use of GPR to monitor moisture in the pavement foundation through comparison with Falling-Weight Deflectometer (FWD) parameters that represent the structural condition of the pavement. The results in this paper are based on GPR, FWD and in-place moisture sensor data collected over 17 months on MnROAD instrumented test sections. The pavement test sections considered in this study included unbound aggregate bases (UAB) with virgin and recycled materials (i.e., RCA and RAP) covering a broad range of geotechnical behavior. Previous efforts established a strong correlation between GPR-based moisture measurements and sensor-based moisture measurements. This correlation is further validated by investigating the relationship between GPRbased moisture measurements and FWD-based indices for structural capacity. The observed correlation is reasonable and in good agreement with the expected correlation between volumetric moisture content (VMC) of the base layer and the structural capacity of a pavement. This agreement validates the use of GPR to monitor moisture in pavement foundation. GPR is not a replacement for FWD, but it can be used in asset management efforts in addition to FWD to assess moisture fluctuation in pavement foundation during and after extreme environmental events (e.g., flooding, freeze-thawing) to determine when and where FWD testing may be necessary.
Background: Pulmonary embolism (PE) is a life-threatening condition associated with ~10% of deaths of hospitalized patients. Machine learning algorithms (MLAs) which predict the onset of pulmonary embolism (PE) could enable earlier treatment and improve patient outcomes. However, the extent to which they generalize to broader patient populations impacts their clinical utility. Objective: To conduct the first large-scale external validation of a machine learning-based PE prediction model which uses EHR data from the first three hours of a patient's hospital stay to predict the occurrence of PE within the next 10 days of the inpatient stay. Methods: This retrospective study included approximately two million adult hospital admissions across 44 medical institutions in the US from 2011 to 2017. Demographics, vital signs, and lab tests from adult inpatients at 12 institutions (n = 331,268; 3.3% PE positive) were used for training an XGBoost model. External validation of the model was conducted on patient populations from each of 32 medical institutions (total n = 1,660,715; 3.7% PE positive) without retraining. Model performance was assessed using the area under the receiver operating characteristic curve (AUROC). Backward elimination regression was used to identify correlations between characteristics of the external validation sets and AUROC. Results: The model performed well (AUROC = 0.87) on the 20% hold-out subset of the training set. Despite demographic differences between the 32 external validation populations (percent PE positive: min = 1.54%, max = 6.47%), without retraining, the model had excellent discrimination, with a mean AUROC of 0.88 (min = 0.79, max = 0.93). Fixing sensitivity at 0.80, the model had a mean specificity of 0.85 (min = 0.64, max = 0.93). Backward elimination regression identified a negative association (beta = -0.015, p < 0.001) between the percentage of PE positive encounters and AUROC. Conclusions: A PE prediction model performed remarkably well across 32 different external patient populations without retraining and despite significant differences in demographic characteristics, demonstrating its generalizability and potential as a clinical decision support tool to aid PE detection and improve patient outcomes in a clinical setting.
Introduction: The objective of this study was to assess seven configurations of six convolutional deep neural network architectures for classification of chest X-rays (CXRs) as COVID-19 positive or negative. Methods: The primary dataset consisted of 294 COVID-19 positive and 294 COVID-19 negative CXRs, the latter comprising roughly equally many pneumonia, emphysema, fibrosis, and healthy images. We used six common convolutional neural network architectures, VGG16, DenseNet121, DenseNet201, MobileNet, NasNetMobile and InceptionV3. We studied six models (one for each architecture) which were pre-trained on a vast repository of generic (non-CXR) images, as well as a seventh DenseNet121 model, which was pre-trained on a repository of CXR images. For each model, we replaced the output layers with custom fully connected layers for the task of binary classification of images as COVID-19 positive or negative. Performance metrics were calculated on a holdout test set with CXRs from patients who were not included in the training/validation set. Results: When pre-trained on generic images, the VGG16, DenseNet121, DenseNet201, MobileNet, NasNetMobile, and InceptionV3 architectures respectively produced hold-out test set areas under the receiver operating characteristic (AUROCs) of 0.98, 0.95, 0.97, 0.95, 0.99, and 0.96 for the COVID-19 classification of CXRs. The X-ray pre-trained DenseNet121 model, in comparison, had a test set AUROC of 0.87. Discussion: Common convolutional neural network architectures with parameters pre-trained on generic images yield high-performance and well-calibrated COVID-19 CXR classification.
Purpose : Coronavirus Disease 2019 (COVID-19) continues to be a global threat and remains a significant cause of hospitalizations. Recent clinical guidelines have supported the use of corticosteroids and remdesivir in the treatment of COVID-19. However, uncertainty remains about which patients are most likely to benefit from treatment with either drug; such knowledge is crucial for avoiding preventable side effects, minimizing costs, and effectively allocating resources. This study presents a machine learning system capable of identifying patients for whom treatment with corticosteroids or remdesivir is associated with improved survival time. Methods : Gradient boosted decision tree models to predict treatment benefit were trained and tested on data from patients hospitalized at 10 hospitals in the United States between December 18, 2019 and October 18, 2020. 893 patients were treated with remdesivir, and 1,471 were treated with corticosteroids. Models were evaluated for their ability to identify patients that exhibited longer survival times when treated with corticosteroids or remdesivir. Fine and Gray models for the proportional hazard were evaluated comparing treated and untreated patients in the full COVID-19 population, in patients receiving supplemental oxygen, and in patients identified by the algorithm. Inverse probability of treatment weights were used to adjust for confounding. Models for each treatment were trained and tested separately. Findings : .Adult patients (age ≥ 18) were included in this study, with men comprising slightly more than 50% of the sample. After adjusting for confounding, neither corticosteroids nor remdesivir were associated with increased survival time in the full hospitalized COVID-19 population or in the population receiving supplemental oxygen. However, in the populations identified by the algorithms, both corticosteroids and remdesivir were significantly associated with an increase in survival time, with hazard ratios of 0.56 (p = 0.04) and 0.40 (p = 0.04), respectively Implications : Machine learning methods are capable of identifying hospitalized COVID-19 patients for whom treatment with corticosteroids or remdesivir is associated with an increase in survival time. These methods may help improve patient outcomes and allocate resources during the COVID-19 crisis.
Ektefaie, Yasha1; Mataraso, Samson1; Barnes, Gina1; Lynn-Palevsky, Anna1; Pellegrini, Emily1; Green-Saxena, Abigail1; Hoffman, Jana1; Calvert, Jacob1; Das, Ritankar1 Author Information