Rare variant association analysis, which assesses the aggregate effect of rare damaging variants within a gene, is a powerful strategy for advancing knowledge of human biology. Numerous models have been proposed to identify damaging coding variants, with the most recent ones employing deep learning and large language models (LLMs) to predict the impact of changes in coding sequences. Here, we use newly available proteomics data on 2898 proteins across 46,665 individuals to evaluate and refine LLM predictors of damaging variants. Using one of these refined models, we evaluate the association between rare damaging variants and human phenotypes at 241 positive control gene-trait pairs. Among these gene-trait pairs, our proteomics-guided model outperforms an ensemble of conventional approaches including PolyPhen2, MutationTaster, SIFT, and LRT, as well as newer machine learning approaches for identifying damaging missense variants, such as CADD, ESM-1v, ESM-1b, and AlphaMissense. When attempting to recover known associations by correctly separating damaging singleton missense variants from other singleton variants, our approach recapitulates 36.5% of gene-trait pairs with known associations, exceeding all the alternatives we considered. Furthermore, when we apply our model to 10 example traits from the UK Biobank, we identify 177 gene-trait associations-again exceeding all other approaches. Our results demonstrate that summary statistics from large-scale human proteomics data enable evaluation and refinement of coding variant classification LLMs, improving discovery potential in human genetic studies.
Meta-analysis of gene-based tests using single-variant summary statistics is a powerful strategy for genetic association studies. However, current approaches require sharing the covariance matrix between variants for each study and trait of interest. For large-scale studies with many phenotypes, these matrices can be cumbersome to calculate, store and share. Here, to address this challenge, we present REMETA-an efficient tool for meta-analysis of gene-based tests. REMETA uses a single sparse covariance reference file per study that is rescaled for each phenotype using single-variant summary statistics. We develop new methods for binary traits with case-control imbalance, and to estimate allele frequencies, genotype counts and effect sizes of burden tests. We demonstrate the performance and advantages of our approach through meta-analysis of five traits in 469,376 samples in UK Biobank. The open-source REMETA software will facilitate meta-analysis across large-scale exome sequencing studies from diverse studies that cannot easily be combined.
OBJECTIVES: Implementing a predictive analytic model in a new clinical environment is fraught with challenges. Dataset shifts such as differences in clinical practice, new data acquisition devices, or changes in the electronic health record (EHR) implementation mean that the input data seen by a model can differ significantly from the data it was trained on. Validating models at multiple institutions is therefore critical. Here, using retrospective data, we demonstrate how Predicting Intensive Care Transfers and other UnfoReseen Events (PICTURE), a deterioration index developed at a single academic medical center, generalizes to a second institution with significantly different patient population. DESIGN: PICTURE is a deterioration index designed for the general ward, which uses structured EHR data such as laboratory values and vital signs. SETTING: The general wards of two large hospitals, one an academic medical center and the other a community hospital. SUBJECTS: The model has previously been trained and validated on a cohort of 165,018 general ward encounters from a large academic medical center. Here, we apply this model to 11,083 encounters from a separate community hospital. INTERVENTIONS: None. MEASUREMENTS AND MAIN RESULTS: The hospitals were found to have significant differences in missingness rates (> 5% difference in 9/52 features), deterioration rate (4.5% vs 2.5%), and racial makeup (20% non-White vs 49% non-White). Despite these differences, PICTURE’s performance was consistent (area under the receiver operating characteristic curve [AUROC], 0.870; 95% CI, 0.861–0.878), area under the precision-recall curve (AUPRC, 0.298; 95% CI, 0.275–0.320) at the first hospital; AUROC 0.875 (0.851–0.902), AUPRC 0.339 (0.281–0.398) at the second. AUPRC was standardized to a 2.5% event rate. PICTURE also outperformed both the Epic Deterioration Index and the National Early Warning Score at both institutions. CONCLUSIONS: Important differences were observed between the two institutions, including data availability and demographic makeup. PICTURE was able to identify general ward patients at risk of deterioration at both hospitals with consistent performance (AUROC and AUPRC) and compared favorably to existing metrics.
The Pharma Proteomics Project is a precompetitive biopharmaceutical consortium characterizing the plasma proteomic profiles of 54,219 UK Biobank participants. Here we provide a detailed summary of this initiative, including technical and biological validations, insights into proteomic disease signatures, and prediction modelling for various demographic and health indicators. We present comprehensive protein quantitative trait locus (pQTL) mapping of 2,923 proteins that identifies 14,287 primary genetic associations, of which 81% are previously undescribed, alongside ancestry-specific pQTL mapping in non-European individuals. The study provides an updated characterization of the genetic architecture of the plasma proteome, contextualized with projected pQTL discovery rates as sample sizes and proteomic assay coverages increase over time. We offer extensive insights into trans pQTLs across multiple biological domains, highlight genetic influences on ligand–receptor interactions and pathway perturbations across a diverse collection of cytokines and complement networks, and illustrate long-range epistatic effects of ABO blood group and FUT2 secretor status on proteins with gastrointestinal tissue-enriched expression. We demonstrate the utility of these data for drug discovery by extending the genetic proxied effects of protein targets, such as PCSK9, on additional endpoints, and disentangle specific genes and proteins perturbed at loci associated with COVID-19 susceptibility. This public–private partnership provides the scientific community with an open-access proteomics resource of considerable breadth and depth to help to elucidate the biological mechanisms underlying proteo-genomic discoveries and accelerate the development of biomarkers, predictive models and therapeutics 1 .
Clonal haematopoiesis involves the expansion of certain blood cell lineages and has been associated with ageing and adverse health outcomes1-5. Here we use exome sequence data on 628,388 individuals to identify 40,208 carriers of clonal haematopoiesis of indeterminate potential (CHIP). Using genome-wide and exome-wide association analyses, we identify 24 loci (21 of which are novel) where germline genetic variation influences predisposition to CHIP, including missense variants in the lymphocytic antigen coding gene LY75, which are associated with reduced incidence of CHIP. We also identify novel rare variant associations with clonal haematopoiesis and telomere length. Analysis of 5,041 health traits from the UK Biobank (UKB) found relationships between CHIP and severe COVID-19 outcomes, cardiovascular disease, haematologic traits, malignancy, smoking, obesity, infection and all-cause mortality. Longitudinal and Mendelian randomization analyses revealed that CHIP is associated with solid cancers, including non-melanoma skin cancer and lung cancer, and that CHIP linked to DNMT3A is associated with the subsequent development of myeloid but not lymphoid leukaemias. Additionally, contrary to previous findings from the initial 50,000 UKB exomes6, our results in the full sample do not support a role for IL-6 inhibition in reducing the risk of cardiovascular disease among CHIP carriers. Our findings demonstrate that CHIP represents a complex set of heterogeneous phenotypes with shared and unique germline genetic causes and varied clinical implications.
Understanding of lung-specific metabolic changes to ARDS development is limited due to the paucity of longitudinal studies and that the lung is difficult to repeatedly sample. We obtained a "metabolic lung biopsy" by simultaneous acquisition of trans-pulmonary whole blood (WB) from the pulmonary (PA) and carotid (CA) arteries every 4h and hourly volatile organic compounds (VOCs) in exhaled breath from a swine model that develops both radiographic and histo-morphologic evidence of ARDS. WB was assayed by quantitative proton nuclear magnetic resonance spectroscopy and exhaled breath by gas chromatography. Data were analyzed by a mixed effect model with random intercept. P/F ratio was used as the fixed effect to determine the average slope for each metabolite. 37 WB metabolites and 55 named VOCs were detected from 4 of 5 and 5 of 5 swine, respectively. With declining P/F ratio, there was increased excretion/production (PA<CA) of key WB energy metabolites including (A) ATP (p=0.007), (B) glucose (p=0.03), (C) ADP (p=0.12) and (D) lysine (p=0.14). In exhaled breath, as ARDS develops, (E) 3-methylheptane (p<0.001), which may be a by-product of lipid peroxidation and (F) a-pinene (p<0.001), a terpene with anti-inflammatory activity, increased and decreased, respectively. These findings suggest that a "metabolic lung biopsy" is feasible and could aid in informing lung-specific metabolic processes that participate in ARDS pathogenesis.
Despite the enormous impact on human health, acute respiratory distress syndrome (ARDS) is poorly defined, and its timely diagnosis is difficult, as is tracking the course of the syndrome. The objective of this pilot study was to explore the utility of breath collection and analysis methodologies to detect ARDS through changes in the volatile organic compound (VOC) profiles present in breath. Five male Yorkshire mix swine were studied and ARDS was induced using both direct and indirect lung injury. An automated portable gas chromatography device developed in-house was used for point of care breath analysis and to monitor swine breath hourly, starting from initiation of the experiment until the development of ARDS, which was adjudicated based on the Berlin criteria at the breath sampling points and confirmed by lung biopsy at the end of the experiment. A total of 67 breath samples (chromatograms) were collected and analysed. Through machine learning, principal component analysis and linear discrimination analysis, seven VOC biomarkers were identified that distinguished ARDS. These represent seven of the nine biomarkers found in our breath analysis study of human ARDS, corroborating our findings. We also demonstrated that breath analysis detects changes 1-6 h earlier than the clinical adjudication based on the Berlin criteria. The findings provide proof of concept that breath analysis can be used to identify early changes associated with ARDS pathogenesis in swine. Its clinical application could provide intensive care clinicians with a noninvasive diagnostic tool for early detection and continuous monitoring of ARDS.
Clonal hematopoiesis (CH) refers to the expansion of certain blood cell lineages and has been associated with aging and adverse health outcomes. Here, we use exome sequence data on 628,388 individuals to identify 40,208 carriers of clonal hematopoiesis of indeterminate potential (CHIP). Using genome-wide and exome-wide association analyses, we identify 27 loci (24 novel) where germline genetic variation influences CH/CHIP predisposition, including missense variants in the DNA-repair gene PARP1 and the lymphocytic antigen coding gene LY75 that are associated with reduced incidence of CH/CHIP. Analysis of 5,194 health traits from the UK Biobank (UKB) found relationships between CHIP and severe COVID outcomes, cardiovascular disease, hematologic traits, malignancy, smoking, obesity, infection, and all-cause mortality. Longitudinal analyses revealed that one of the CHIP subtypes, DNMT3A-CHIP, is associated with the subsequent development of myeloid but not lymphoid leukemias, and with solid cancers including prostate and lung. Additionally, contrary to previous findings from the initial 50,000 UKB exomes, our results in the full sample do not support a role for IL-6 inhibition in reducing the risk of cardiovascular disease among CHIP carriers. Our findings demonstrate that CHIP represents a complex set of heterogenous phenotypes with shared and unique germline genetic causes and varied clinical implications.
The UK Biobank Pharma Proteomics Project (UKB-PPP) is a collaboration between the UK Biobank (UKB) and thirteen biopharmaceutical companies characterising the plasma proteomic profiles of 54,306 UKB participants. Here, we describe results from the first phase of UKB-PPP, including protein quantitative trait loci (pQTL) mapping of 1,463 proteins that identifies 10,248 primary genetic associations, of which 85% are newly discovered. We also identify independent secondary associations in 92% of cis and 29% of trans loci, expanding the catalogue of genetic instruments for downstream analyses. The study provides an updated characterisation of the genetic architecture of the plasma proteome, leveraging population-scale proteomics to provide novel, extensive insights into trans pQTLs across multiple biological domains. We highlight genetic influences on ligand-receptor interactions and pathway perturbations across a diverse collection of cytokines and complement proteins, and illustrate long-range epistatic effects of ABO blood group and FUT2 secretor status on proteins with gastrointestinal tissue-enriched expression. We demonstrate the utility of these data for drug target discovery by extending the genetic proxied effect of PCSK9 levels on lipid concentrations, cardio- and cerebro-vascular diseases, and additionally disentangle specific genes and proteins perturbed at COVID-19 susceptibility loci. This public-private partnership provides the scientific community with an open-access proteomics resource of unprecedented breadth and depth to help elucidate biological mechanisms underlying genetic discoveries and accelerate the development of novel biomarkers and therapeutics.
Sepsis‐induced metabolic dysfunction contributes to organ failure and death. L‐carnitine has shown promise for septic shock, but a recent phase II study of patients with vasopressor‐dependent septic shock demonstrated a non‐significant reduction in mortality. We undertook a pharmacometabolomics study of these patients (n = 250) to identify metabolic profiles predictive of a 90‐day mortality benefit from L‐carnitine. The independent predictive value of each pretreatment metabolite concentration, adjusted for L‐carnitine dose, on 90‐day mortality was determined by logistic regression. A grid‐search analysis maximizing the Z‐statistic from a binomial proportion test identified specific metabolite threshold levels that discriminated L‐carnitine responsive patients. Threshold concentrations were further assessed by hazard ratio and Kaplan‐Meier estimate. Accounting for L‐carnitine treatment and dose, 11 1H‐NMR metabolites and 12 acylcarnitines were independent predictors of 90‐day mortality. Based on the grid‐search analysis numerous acylcarnitines and valine were identified as candidate metabolites of drug response. Acetylcarnitine emerged as highly viable for the prediction of an L‐carnitine mortality benefit due to its abundance and biological relevance. Using its most statistically significant threshold concentration, patients with pretreatment acetylcarnitine greater than or equal to 35 µM were less likely to die at 90 days if treated with L‐carnitine (18 g) versus placebo (p = 0.01 by log rank test). Metabolomics also identified independent predictors of 90‐day sepsis mortality. Our proof‐of‐concept approach shows how pharmacometabolomics could be useful for tackling the heterogeneity of sepsis and informing clinical trial design. In addition, metabolomics can help understand mechanisms of sepsis heterogeneity and variable drug response, because sepsis induces alterations in numerous metabolite concentrations.
BACKGROUND COVID-19 has led to an unprecedented strain on health care facilities across the United States. Accurately identifying patients at an increased risk of deterioration may help hospitals manage their resources while improving the quality of patient care. Here, we present the results of an analytical model, Predicting Intensive Care Transfers and Other Unforeseen Events (PICTURE), to identify patients at high risk for imminent intensive care unit transfer, respiratory failure, or death, with the intention to improve the prediction of deterioration due to COVID-19. OBJECTIVE This study aims to validate the PICTURE model’s ability to predict unexpected deterioration in general ward and COVID-19 patients, and to compare its performance with the Epic Deterioration Index (EDI), an existing model that has recently been assessed for use in patients with COVID-19. METHODS The PICTURE model was trained and validated on a cohort of hospitalized non–COVID-19 patients using electronic health record data from 2014 to 2018. It was then applied to two holdout test sets: non–COVID-19 patients from 2019 and patients testing positive for COVID-19 in 2020. PICTURE results were aligned to EDI and NEWS scores for head-to-head comparison via area under the receiver operating characteristic curve (AUROC) and area under the precision-recall curve. We compared the models’ ability to predict an adverse event (defined as intensive care unit transfer, mechanical ventilation use, or death). Shapley values were used to provide explanations for PICTURE predictions. RESULTS In non–COVID-19 general ward patients, PICTURE achieved an AUROC of 0.819 (95% CI 0.805-0.834) per observation, compared to the EDI’s AUROC of 0.763 (95% CI 0.746-0.781; n=21,740; P<.001). In patients testing positive for COVID-19, PICTURE again outperformed the EDI with an AUROC of 0.849 (95% CI 0.820-0.878) compared to the EDI’s AUROC of 0.803 (95% CI 0.772-0.838; n=607; P<.001). The most important variables influencing PICTURE predictions in the COVID-19 cohort were a rapid respiratory rate, a high level of oxygen support, low oxygen saturation, and impaired mental status (Glasgow Coma Scale). CONCLUSIONS The PICTURE model is more accurate in predicting adverse patient outcomes for both general ward patients and COVID-19 positive patients in our cohorts compared to the EDI. The ability to consistently anticipate these events may be especially valuable when considering potential incipient waves of COVID-19 infections. The generalizability of the model will require testing in other health care systems for validation.
A major goal in human genetics is to use natural variation to understand the phenotypic consequences of altering each protein-coding gene in the genome. Here we used exome sequencing 1 to explore protein-altering variants and their consequences in 454,787 participants in the UK Biobank study 2 . We identified 12 million coding variants, including around 1 million loss-of-function and around 1.8 million deleterious missense variants. When these were tested for association with 3,994 health-related traits, we found 564 genes with trait associations at P ≤ 2.18 × 10 −11 . Rare variant associations were enriched in loci from genome-wide association studies (GWAS), but most (91%) were independent of common variant signals. We discovered several risk-increasing associations with traits related to liver disease, eye disease and cancer, among others, as well as risk-lowering associations for hypertension ( SLC9A3R2 ), diabetes ( MAP3K15 , FAM234A ) and asthma ( SLC27A3 ). Six genes were associated with brain imaging phenotypes, including two involved in neural development ( GBE1 , PLD1 ). Of the signals available and powered for replication in an independent cohort, 81% were confirmed; furthermore, association signals were generally consistent across individuals of European, Asian and African ancestry. We illustrate the ability of exome sequencing to identify gene–trait associations, elucidate gene function and pinpoint effector genes that underlie GWAS signals at scale.
Background Acute respiratory distress syndrome (ARDS) is a common, but under-recognised, critical illness syndrome associated with high mortality. An important factor in its under-recognition is the variability in chest radiograph interpretation for ARDS. We sought to train a deep convolutional neural network (CNN) to detect ARDS findings on chest radiographs. Methods CNNs were pretrained on 595 506 radiographs from two centres to identify common chest findings (eg, opacity and effusion), and then trained on 8072 radiographs annotated for ARDS by multiple physicians using various transfer learning approaches. The best performing CNN was tested on chest radiographs in an internal and external cohort, including a subset reviewed by six physicians, including a chest radiologist and physicians trained in intensive care medicine. Chest radiograph data were acquired from four US hospitals. Findings In an internal test set of 1560 chest radiographs from 455 patients with acute hypoxaemic respiratory failure, a CNN could detect ARDS with an area under the receiver operator characteristics curve (AUROC) of 0.92 (95% CI 0.89-0.94). In the subgroup of 413 images reviewed by at least six physicians, its AUROC was 0.93 (95% CI 0.88-0.96), sensitivity 83.0% (95% CI 74.0-91.1), and specificity 88.3% (95% CI 83.1-92.8). Among images with zero of six ARDS annotations (n=155), the median CNN probability was 11%, with six (4%) assigned a probability above 50%. Among images with six of six ARDS annotations (n=27), the median CNN probability was 91%, with two (7%) assigned a probability below 50%. In an external cohort of 958 chest radiographs from 431 patients with sepsis, the AUROC was 0.88 (95% CI 0.85-0.91). When radiographs annotated as equivocal were excluded, the AUROC was 0.93 (0.92-0.95). Interpretation A CNN can be trained to achieve expert physician-level performance in ARDS detection on chest radiographs. Further research is needed to evaluate the use of these algorithms to support real-time identification of ARDS patients to ensure fidelity with evidence-based care or to support ongoing ARDS research. (C) 2021 The Author(s). Published by Elsevier Ltd.
To ensure scientific reproducibility of metabolomics data, alternative statistical methods are needed. A paradigm shift away from the p-value toward an embracement of uncertainty and interval estimation of a metabolite's true effect size may lead to improved study design and greater reproducibility. Multilevel Bayesian models are one approach that offer the added opportunity of incorporating imputed value uncertainty when missing data are present. We designed simulations of metabolomics data to compare multilevel Bayesian models to standard logistic regression with corrections for multiple hypothesis testing. Our simulations altered the sample size and the fraction of significant metabolites truly different between two outcome groups. We then introduced missingness to further assess model performance. Across simulations, the multilevel Bayesian approach more accurately estimated the effect size of metabolites that were significantly different between groups. Bayesian models also had greater power and mitigated the false discovery rate. In the presence of increased missing data, Bayesian models were able to accurately impute the true concentration and incorporating the uncertainty of these estimates improved overall prediction. In summary, our simulations demonstrate that a multilevel Bayesian approach accurately quantifies the estimated effect size of metabolite predictors in regression modeling, particularly in the presence of missing data.
When using tree-based methods to develop predictive analytics and early warning systems for preventive healthcare, it is important to use an appropriate imputation method to prevent learning the missingness pattern. To demonstrate this, we developed a novel simulation that generated synthetic electronic health record data using a variational autoencoder with a custom loss function, which took into account the high missing rate of electronic health data. We showed that when tree-based methods learn missingness patterns (correlated with adverse events) in electronic health record data, this leads to decreased performance if the system is used in a new setting that has different missingness patterns. Performance is worst in this scenario when the missing rate between those with and without an adverse event is the greatest. We found that randomized and Bayesian regression imputation methods mitigate the issue of learning the missingness pattern for tree-based methods. We used this information to build a novel early warning system for predicting patient deterioration in general wards and telemetry units: PICTURE (Predicting Intensive Care Transfers and other UnfoReseen Events). To develop, tune, and test PICTURE, we used labs and vital signs from electronic health records of adult patients over four years (n=133,089 encounters). We analyzed primary outcomes of unplanned intensive care unit transfer, emergency vasoactive medication administration, cardiac arrest, and death. We compared PICTURE with existing early warning systems and logistic regression at multiple levels of granularity. When analyzing PICTURE on the testing set using all observations within a hospital encounter (event rate = 3.4%), PICTURE had an area under the receiver operating characteristic curve (AUROC) of 0.83 and an adjusted (event rate = 4%) area under the precision-recall curve (AUPR) of 0.27, while the next best-tested method-regularized logistic regression-had an AUROC of 0.80 and an adjusted AUPR of 0.22. To ensure system interpretability, we applied a state-of-the-art prediction explainer that provided a ranked list of features contributing most to the prediction. Though it is currently difficult to compare machine learning-based early warning systems, a rudimentary comparison with published scores demonstrated that PICTURE is on par with state-of-the-art machine learning systems. To facilitate more robust comparisons and development of early warning systems in the future, we have released our variational autoencoder's code and weights so researchers can (a) test their models on data similar to our institution and (b) make their own synthetic datasets.
Objective The objective of this review is to discuss the therapeutic use and differential treatment response to Levo-carnitine (l-carnitine) treatment in septic shock, and to demonstrate common lessons learned that are important to the advancement of precision medicine approaches to sepsis. We propose that significant interpatient variability in the metabolic response tol-carnitine and clinical outcomes can be used to elucidate the mechanistic underpinnings that contribute to sepsis heterogeneity. Methods A narrative review was conducted that focused on explaining interpatient variability inl-carnitine treatment response. Relevant biological and patient-level characteristics considered include genetic, metabolic, and morphomic phenotypes; potential drug interactions; and pharmacokinetics (PKs). Main Results Despite promising results in a phase I study, a recent phase II clinical trial ofl-carnitine treatment in septic shock showed a nonsignificant reduction in mortality. However,l-carnitine treatment induces significant interpatient variability inl-carnitine and acylcarnitine concentrations over time. In particular, administration ofl-carnitine induces a broad, dynamic range of serum concentrations and measured peak concentrations are associated with mortality. Applied systems pharmacology may explain variability in drug responsiveness by using patient characteristics to identify pretreatment phenotypes most likely to derive benefit froml-carnitine. Moreover, provocation of sepsis metabolism withl-carnitine offers a unique opportunity to identify metabolic response signatures associated with patient outcomes. These approaches can unmask latent metabolic pathways deranged in the sepsis syndrome and offer insight into the pathophysiology, progression, and heterogeneity of the disease. Conclusions The compiled evidence suggests there are several potential explanations for the variability in carnitine concentrations and clinical response tol-carnitine in septic shock. These serve as important confounders that should be considered in interpretation ofl-carnitine clinical studies and broadly holds lessons for future clinical trial design in sepsis. Consideration of these factors is needed if precision medicine in sepsis is to be achieved.
This article concerns the PhysioNet/Computing in Cardiology Challenge 2020 which focused on building computational methods to identify cardiac abnormalities from 12-lead ECGs. Our team, MCIRCC, utilized a large secondary dataset of 12-lead ECGs obtained from the Section of Electrophysiology at the University of Michigan, called the MUSE dataset, to pre-train multiple residual neural networks that were later re-trained on the challenge dataset. To do so, the diagnosis statements that existed in our dataset were utilized to assign the same labels to our ECGs as the challenge data. After parameter optimization, we selected a subset of top performing models and created an ensemble model that achieved a challenge validation score of 0.616, and full test score of 0.141, placing us 27th out of 41 teams in the official ranking.
Introduction The 2019 coronavirus (COVID-19) has led to unprecedented strain on healthcare facilities across the United States. Accurately identifying patients at an increased risk of deterioration may help hospitals manage their resources while improving the quality of patient care. Here we present the results of an analytical model, PICTURE (Predicting Intensive Care Transfers and other UnfoReseen Events), to identify patients at a high risk for imminent intensive care unit (ICU) transfer, respiratory failure, or death with the intention to improve prediction of deterioration due to COVID-19. We compare PICTURE to the Epic Deterioration Index (EDI), a widespread system which has recently been assessed for use to triage COVID-19 patients. Methods The PICTURE model was trained and validated on a cohort of hospitalized non-COVID-19 patients using electronic health record data from 2014-2018. It was then applied to two hold-out test sets: non-COVID-19 patients from 2019 and patients testing positive for COVID-19 in 2020. PICTURE results were aligned to the EDI for head-to-head comparison via Area Under the Receiver Operator Curve (AUROC) and Area Under the Precision Recall Curve (AUPRC). We compared the models' ability to predict an adverse event (defined as ICU transfer, mechanical ventilation use, or death) at two levels of granularity: (1) maximum score across an encounter with a minimum lead time before the first adverse event and (2) predictions at every observation with instances in the last 24 hours before the adverse event labeled as positive. PICTURE and the EDI were also compared on the encounter level using different lead times extending out to 24 hours. Shapley values were used to provide explanations for PICTURE predictions. Results PICTURE successfully delineated between high- and low-risk patients and consistently outperformed the EDI in both of our cohorts. In non-COVID-19 patients, PICTURE achieved an AUROC (95% CI) of 0.819 (0.805 - 0.834) and AUPRC of 0.109 (0.089 - 0.125) on the observation level, compared to the EDI AUROC of 0.762 (0.746 - 0.780) and AUPRC of 0.077 (0.062 - 0.090). On COVID-19 positive patients, PICTURE achieved an AUROC of 0.828 (0.794 - 0.869) and AUPRC of 0.160 (0.089 - 0.199), while the EDI scored an AUROC of 0.792 (0.754 - 0.835) and AUPRC of 0.131 (0.092 - 0.159). The most important variables influencing PICTURE predictions in the COVID-19 cohort were a rapid respiratory rate, a high level of oxygen support, low oxygen saturation, and impaired mental status (Glasgow coma score). Conclusion The PICTURE model is more accurate in predicting adverse patient outcomes for both general ward patients and COVID-19 positive patients in our cohorts compared to the EDI. The ability to consistently anticipate these events may be especially valuable when considering a potential incipient second wave of COVID-19 infections. PICTURE also has the ability to explain individual predictions to clinicians by ranking the most important features for a prediction. The generalizability of the model will require testing in other health care systems for validation.
Acute respiratory distress syndrome (ARDS) is the most severe form of acute lung injury, responsible for high mortality and long-term morbidity. As a dynamic syndrome with multiple etiologies its timely diagnosis is difficult as is tracking the course of the syndrome. Therefore, there is a significant need for early, rapid detection and diagnosis as well as clinical trajectory monitoring of ARDS. Here we report our work on using human breath to differentiate ARDS and non-ARDS causes of respiratory failure. A fully automated portable 2-dimensional gas chromatography device with high peak capacity, high sensitivity, and rapid analysis capability was designed and made in-house for on-site analysis of patients’ breath. A total of 85 breath samples from 48 ARDS patients and controls were collected. Ninety-seven elution peaks were separated and detected in 13 minutes. An algorithm based on machine learning, principal component analysis (PCA), and linear discriminant analysis (LDA) was developed. As compared to the adjudications done by physicians based on the Berlin criteria, our device and algorithm achieved an overall accuracy of 87.1% with 94.1% positive predictive value and 82.4% negative predictive value. The high overall accuracy and high positive predicative value suggest that the breath analysis method can accurately diagnose ARDS. The ability to continuously and non-invasively monitor exhaled breath for early diagnosis, disease trajectory tracking, and outcome prediction monitoring of ARDS may have a significant impact on changing practice and improving patient outcomes.
Although the slit diaphragm proteins in podocytes are uniquely organized to maintain glomerular filtration assembly and function, little is known about the underlying mechanisms that participate in trafficking these proteins to the correct location for development and homeostasis. Identifying these mechanisms will likely provide novel targets for therapeutic intervention to preserve podocyte function following glomerular injury. Analysis of structural variation in cases of human nephrotic syndrome identified rare heterozygous deletions of EXOC4 in two patients. This suggested that disruption of the highly-conserved eight-protein exocyst trafficking complex could have a role in podocyte dysfunction. Indeed, mRNA profiling of injured podocytes identified significant exocyst down-regulation. To test the hypothesis that the exocyst is centrally involved in podocyte development/function, we generated homozygous podocyte-specific Exoc5 (a central exocyst component that interacts with Exoc4) knockout mice that showed massive proteinuria and died within 4 weeks of birth. Histological and ultrastructural analysis of these mice showed severe glomerular defects with increased fibrosis, proteinaceous casts, effaced podocytes, and loss of the slit diaphragm. Immunofluorescence analysis revealed that Neph1 and Nephrin, major slit diaphragm constituents, were mislocalized and/or lost. mRNA profiling of Exoc5 knockdown podocytes showed that vesicular trafficking was the most affected cellular event. Mapping of signaling pathways and Western blot analysis revealed significant up-regulation of the mitogen-activated protein kinase and transforming growth factor-beta pathways in Exoc5 knockdown podocytes and in the glomeruli of podocyte-specific Exoc5 KO mice. Based on these data, we propose that exocyst-based mechanisms regulate Neph1 and Nephrin signaling and trafficking, and thus podocyte development and function.