BACKGROUND:Near real-time electronic health record (EHR) data offers significant potential for secondary use in research, operations, and clinical care, yet challenges remain in ensuring data quality and stability. While prior studies have assessed retrospective EHR datasets, few have systematically examined the integrity of real-time data for research readiness. METHODS:We developed an automated benchmarking pipeline to evaluate the stability and completeness of real-time EHR data from the Yale New Haven Health clinical data warehouse, transformed into the OMOP common data model. Twenty-nine weekly snapshots of the EHR collected from July to November 2024 and twenty-two daily snapshots collected from April to May 2025 were analyzed. Benchmarks focused on (1) clinical actions such as patient additions, deletions, and merges; (2) changes in demographic variables (date of birth, gender, race, ethnicity); and (3) stability of discharge information (time and status). A synthetic dataset derived from MIMIC-III was used to validate the benchmarking code prior to large-scale analyses. RESULTS:Benchmarking revealed frequent updates due to clinical actions and demographic corrections across consecutive snapshots. Demographic changes were most frequently related to race and ethnicity, highlighting potential workflow and data entry inconsistencies. Discharge time and status values demonstrated instability for several days post-encounter, typically reaching a stable state within 4-7 days. These findings indicate that while near real-time EHR data provide valuable insights, the timing of data stabilization is critical for accurate secondary use. CONCLUSIONS:This study demonstrates the feasibility of automated benchmarking to assess the integrity of real-time EHR data and identify when such data become analysis ready. Our findings highlight key challenges for secondary use of dynamic clinical data and provide an automated framework that can be applied across health systems to support high-quality research, surveillance, and clinical trial readiness.
BACKGROUND:Quantum machine learning (QML) is an emerging field that may offer unique advantages over classical machine learning but has not been extensively studied with real-world healthcare data of practical size. This study evaluates the performance of a recently proposed quantum machine learning algorithm-quantum circuits with data re-uploading (QC-REUP)-in comparison to classical machine learning and other QML methods. Performance was assessed using both benchmarking data sets and real-world laboratory medicine data. METHODS:Four data sets containing between 2 and 30 features were selected for evaluation. Initial baseline comparisons of classification performance (F1 score) were conducted using all algorithms-QC-REUP, 2 QML, and 4 classical machine learning (ML) algorithms-across the four data sets. Configuration parameters were then optimized for the QC-REUP algorithm using a previously published data set of plasma amino acid (PAA) profiles to determine the impact of optimization on classification performance, followed by a final comparison against classical ML algorithms. RESULTS:Baseline comparisons showed that QC-REUP outperformed quantum and linear classical algorithms on lower-dimensional data sets. However, as input dimensionality increased, QML F -score declined. Following optimization, QC-REUP improved relative to baseline, and again performed comparably to linear algorithms, but ultimately demonstrated lower performance than nonlinear classical ML algorithms on the PAA data set. CONCLUSION:This study suggests that data re-uploading algorithms can perform comparably to classical approaches in specific contexts, particularly with low-dimensional data. While optimization can enhance performance, further improvements in quantum hardware and algorithmic development are likely needed before QML can be effectively applied in laboratory medicine and broader biomedical research.
Objectives To evaluate the feasibility for use of electronic health record (EHR) data in conducting adverse event surveillance among women who received mid-urethral slings (MUS) to treat stress urinary incontinence (SUI) in five health systems.Design Retrospective observational study using EHR data from 2010 through 2021. Women with a history of MUS were identified using common data models; a common analytic code was executed at each site. A manual chart review was conducted in a per-site random patient subset to establish a reference standard. Automated text processing (Text Processed Integrated (TPI)) was developed and evaluated at each site to determine the surgical approach and synthetic mesh implantation. Patients were characterized and surgical outcomes were ascertained over 730 subsequent days.Setting Five large tertiary care academic medical centers.Participants Across five health systems, 9,906 eligible patients (mean age 57–60 per site) were identified.Main outcome measures Determination of surgical approach, synthetic mesh implantation, and assessment of the duration of surveillance for mortality and reoperation rates following MUS implantation.Results In the TPI cohort analysis, 3,331 patients were identified. Surgical approach per site was retropubic (42% to 77%), transobturator (6% to 44%), single incision (0% to 24%), and adjustable sling (0% to <4%). Concordance rates for TPI using chart review were 71%–90% at each site for the surgical approach and 28%–85% for synthetic mesh implantation. Patient follow-up observation rates for mortality and reoperation ranged from 22% to 36% at 90 days, 15% to 30% at 365 days, and 8% to 19% at 730 days.Conclusion Using EHR data alone, identification of medical devices and surgical approaches was feasible among women with MUS surgery for SUI, but long-term follow-up ascertainment rates were low. Medical device surveillance using EHR data should be evaluated in the context of the clinical use case, as applicability may vary.
Background:There is a significant delay between symptom onset and diagnosis of childhood asthma, but the impact of this delay on asthma outcomes has not been well understood. Objectives:We sought to study the association of delayed diagnosis of asthma with asthma exacerbations (AEs) in children. Methods:Using the Mayo Clinic birth cohort, we identified children with a diagnosis of asthma from electronic health records. We defined onset date as the date when subjects first met predetermined asthma criteria ascertained by an electronic health records-based natural language processing algorithm. Delay in diagnosis (DD) was defined as first diagnosis >30 days from onset date (vs timely diagnosis [TD] within 30 days). The primary outcome was AE after the index date (for DD: first diagnosis date vs for TD: clinic visit at similar delay from diagnosis as matched DD counterpart). A Cox proportional hazard model was used to test the association between delayed diagnosis status and risk of AE, adjusting for sociodemographics, care quality, and asthma severity. Results:Among 537 matched pairs of DD and TD (median age at index date: 4.1 years), a total of 344 and 253 children in DD and TD, respectively, had ≥1 AE during median follow-up period of 9.3 years. Children in the DD group had a significantly increased risk of AE compared to TD (adjusted hazard ratio: 1.53; 95% CI: 1.28, 1.80; P < .001). Conclusions:DD of asthma in children is associated with an increased risk of AE compared to TD. TD of asthma should be an important priority in childhood asthma management.
BACKGROUND:The COVID-19 pandemic revealed an urgent need for practical screening tests to rule out respiratory virus infection, both for managing outbreaks and for routine screening in high-risk settings. PCR is the gold standard test for respiratory virus diagnosis but requires specialised equipment, uses different assays for each virus, and often excludes emerging viruses. The goal of this study was to evaluate a pan-viral host biomarker to rule out respiratory virus infection. We used CXCL10, a cytokine induced in the nasal mucosa in response to diverse respiratory viruses. METHODS:We compared immunoassay for CXCL10 to respiratory virus PCR panel results in 1088 nasopharyngeal samples from adults and children with an overall viral prevalence of 32.6% by PCR. Using this data, we mathematically modelled the impact of CXCL10 biomarker testing on patient triage and resource savings at different viral prevalences. We also explored clinical features associated with false negatives using automated data extraction from electronic medical records. FINDINGS:CXCL10 accurately predicted virus positivity (A.U.C. 0.87, 95% C.I. 0.85-0.90). Mathematical modelling predicted that CXCL10 screening would enable a significant reduction in PCR testing, especially when viral prevalence is low (e.g. 92% of samples testing negative when viral prevalence is 5%, NPV = 0.975). Outlier analysis identified specific chemotherapeutic drugs and low viral load as features associated with false negatives. INTERPRETATION:These results demonstrate the utility of a nasopharyngeal biomarker to rule out respiratory infection, with potential applications in outbreak management and/or routine screening in high-risk settings. FUNDING:Yale-New Haven Hospital Innovation Fund and NIH.
BACKGROUND:The symptomatic and immune responses to COVID-19 vaccination of people with Long COVID are poorly characterized. METHODS:In this prospective study, we evaluated changes in symptoms and immune responses after COVID-19 vaccination in 16 vaccine-naïve individuals with Long COVID. Surveys were administered before vaccination and at 2, 6, and 12 weeks after receiving the first vaccine dose of the primary series. Simultaneously, SARS-CoV-2-reactive TCR enrichment, SARS-CoV-2-specific antibody responses, antibody responses to other viral and self-antigens, and circulating cytokines were quantified before vaccination and at 6 and 12 weeks after vaccination. RESULTS:At 12 weeks post-vaccination, self-reported improved health is seen in 10 out of 16 participants, 3 have no change, and 3 have worse health although 2 report transient improvement after vaccination. One participant reporting worse health was hospitalized twice with chest pain (after each dose). Symptom outcomes are most associated with plasma biosignatures. Higher baseline sIL-6R is associated with symptom improvement, and stably elevated levels of IFN-β and CNTF are associated with no improvement. Significant elevation in SARS-CoV-2-specific TCRs and spike protein-specific IgG are observed at 6 and 12 weeks after vaccination. No changes in reactivities are observed against herpes viruses and self-antigens. CONCLUSIONS:In this study of 16 people with Long COVID, vaccination is associated with increased SARS-CoV-2 spike protein-specific IgG and T cell expansion in most participants. Specific immune features are associated with symptom change after vaccination and most participants experience improved health or no change following vaccination.
This article describes the Cell Maps for Artificial Intelligence (CM4AI) project and its goals, methods, standards, current datasets, software tools , status, and future directions. CM4AI is the Functional Genomics Data Generation Project in the U.S. National Institute of Health's (NIH) Bridge2AI program. Its overarching mission is to produce ethical, AI-ready datasets of cell architecture, inferred from multimodal data collected for human cell lines, to enable transformative biomedical AI research.
The use of the Sequential Organ Failure Assessment (SOFA) score, originally developed to describe disease morbidity, is commonly used to predict in-hospital mortality. During the COVID-19 pandemic, many protocols for crisis standards of care used the SOFA score to select patients to be deprioritized due to a low likelihood of survival. A prior study found that age outperformed the SOFA score for mortality prediction in patients with COVID-19, but was limited to a small cohort of intensive care unit (ICU) patients and did not address whether their findings were unique to patients with COVID-19. Moreover, it is not known how well these measures perform across races. In this retrospective study, we compare the performance of age and SOFA score in predicting in-hospital mortality across two cohorts: a cohort of 2,648 consecutive adult patients diagnosed with COVID-19 who were admitted to a large academic health system in the northeastern United States over a 4-month period in 2020 and a cohort of 75,601 patients admitted to one of 335 ICUs in the eICU database between 2014 and 2015. We used age and the maximum SOFA score as predictor variables in separate univariate logistic regression models for in-hospital mortality and calculated area under the receiver operator characteristic curves (AU-ROCs) and area under precision-recall curves (AU-PRCs) for each predictor in both cohorts. Among the COVID-19 cohort, age (AU-ROC 0.795, 95% CI 0.762, 0.828) had a significantly better discrimination than SOFA score (AU-ROC 0.679, 95% CI 0.638, 0.721) for mortality prediction. Conversely, age (AU-ROC 0.628 95% CI 0.608, 0.628) underperformed compared to SOFA score (AU-ROC 0.735, 95% CI 0.726, 0.745) in non-COVID-19 ICU patients in the eICU database. There was no difference between Black and White COVID-19 patients in performance of either age or SOFA Score. Our findings bring into question the utility of SOFA score-based resource allocation in COVID-19 crisis standards of care.
Significant variations have been observed in viral copies generated during SARS-CoV-2 infections. However, the factors that impact viral copies and infection dynamics are not fully understood, and may be inherently dependent upon different viral and host factors. Here, we conducted virus whole genome sequencing and measured viral copies using RT-qPCR from 9,902 SARS-CoV-2 infections over a 2-year period to examine the impact of virus genetic variation on changes in viral copies adjusted for host age and vaccination status. Using a genome-wide association study (GWAS) approach, we identified multiple single-nucleotide polymorphisms (SNPs) corresponding to amino acid changes in the SARS-CoV-2 genome associated with variations in viral copies. We further applied a marginal epistasis test to detect interactions among SNPs and identified multiple pairs of substitutions located in the spike gene that have non-linear effects on viral copies. We also analyzed the temporal patterns and found that SNPs associated with increased viral copies were predominantly observed in Delta and Omicron BA.2/BA.4/BA.5/XBB infections, whereas those associated with decreased viral copies were only observed in infections with Omicron BA.1 variants. Our work showcases how GWAS can be a useful tool for probing phenotypes related to SNPs in viral genomes that are worth further exploration. We argue that this approach can be used more broadly across pathogens to characterize emerging variants and monitor therapeutic interventions.
Objectives To introduce quantum computing technologies as a tool for biomedical research and highlight future applications within healthcare, focusing on its capabilities, benefits, and limitations.Target Audience Investigators seeking to explore quantum computing and create quantum-based applications for healthcare and biomedical research.Scope Quantum computing requires specialized hardware, known as quantum processing units, that use quantum bits (qubits) instead of classical bits to perform computations. This article will cover (1) proposed applications where quantum computing offers advantages to classical computing in biomedicine; (2) an introduction to how quantum computers operate, tailored for biomedical researchers; (3) recent progress that has expanded access to quantum computing; and (4) challenges, opportunities, and proposed solutions to integrate quantum computing in biomedical applications.
Electronic Health Records (EHRs) represent a crucial data source for real-world evidence generation. To facilitate biomedical studies using EHRs, standard data models like the OMOP CDM have been developed. Nevertheless, recent advancements in biomedical AI research that leverage EHRs have introduced new challenges, encompassing security considerations, large-scale data retrieval, and computational resource management, including GPUs. This paper introduces Kamino, an innovative architectural solution tailored to support biomedical AI research using EHR data. Kamino offers a user-friendly interface with features designed for efficient team access management in accordance with regulatory requirements. It facilitates direct data retrieval from an OMOP CDM instance and includes a resource allocation system based on Kubernetes orchestration. Here, we demonstrate the practical application and utility of Kamino through a clinical natural language processing task. We firmly believe that such a tool will significantly expedite AI research conducted with EHR data within academic institutions.
Background The digital transformation of medical data enables health systems to leverage real‐world data from electronic health records to gain actionable insights for improving hypertension care. Methods and Results We performed a serial cross‐sectional analysis of outpatients of a large regional health system from 2010 to 2021. Hypertension was defined by systolic blood pressure ≥140 mm Hg, diastolic blood pressure ≥90 mm Hg, or recorded treatment with antihypertension medications. We evaluated 4 methods of using blood pressure measurements in the electronic health record to define hypertension. The primary outcomes were age‐adjusted prevalence rates and age‐adjusted control rates. Hypertension prevalence varied depending on the definition used, ranging from 36.5% to 50.9% initially and increasing over time by ≈5%, regardless of the definition used. Control rates ranged from 61.2% to 71.3% initially, increased during 2018 to 2019, and decreased during 2020 to 2021. The proportion of patients with a hypertension diagnosis ranged from 45.5% to 60.2% initially and improved during the study period. Non‐Hispanic Black patients represented 25% of our regional population and consistently had higher prevalence rates, higher mean systolic and diastolic blood pressure, and lower control rates compared with other racial and ethnic groups. Conclusions In a large regional health system, we leveraged the electronic health record to provide real‐world insights. The findings largely reflected national trends but showed distinctive regional demographics and findings, with prevalence increasing, one‐quarter of the patients not controlled, and marked disparities. This approach could be emulated by regional health systems seeking to improve hypertension care.
Background:Improving hypertension control is a public health priority. However, consistent identification of uncontrolled hypertension using computable definitions in electronic health records (EHR) across health systems remains uncertain. Methods:In this retrospective cohort study, we applied two computable definitions to the EHR data to identify patients with controlled and uncontrolled hypertension and to evaluate differences in characteristics, treatment, and clinical outcomes between these patient populations. We included adult patients (≥ 18 years) with hypertension receiving ambulatory care within Yale-New Haven Health System (YNHHS; a large US health system) and OneFlorida Clinical Research Consortium (OneFlorida; a Clinical Research Network comprised of 16 health systems) between October 2015 and December 2018. We identified patients with controlled and uncontrolled hypertension based on either a single blood pressure (BP) measurement from a randomly selected visit or all BP measurements recorded between hypertension identification and the randomly selected visit). Results:Overall, 253,207 and 182,827 adults at YNHHS and OneFlorida were identified as having hypertension. Of these patients, 83.1% at YNHHS and 76.8% at OneFlorida were identified using ICD-10-CM codes, whereas 16.9% and 23.2%, respectively, were identified using elevated BP measurements (≥ 140/90 mmHg). Uncontrolled hypertension was observed among 32.5% and 43.7% of patients at YNHHS and OneFlorida, respectively. Uncontrolled hypertension was disproportionately higher among Black patients when compared with White patients (38.9% versus 31.5% in YNHHS; p < 0.001; 49.7% versus 41.2% in OneFlorida; p < 0.001). Medication prescription for hypertension management was more common in patients with uncontrolled hypertension when compared with those with controlled hypertension (overall treatment rate: 39.3% versus 37.3% in YNHHS; p = 0.04; 42.2% versus 34.8% in OneFlorida; p < 0.001). Patients with controlled and uncontrolled hypertension had similar rates of short-term (at 3 and 6 months) and long-term (at 12 and 24 months) clinical outcomes. The two computable definitions generated consistent results. Conclusions:Our findings illustrate the potential of leveraging EHR data, employing computable definitions, to conduct effective digital population surveillance in the realm of hypertension management.
Studies during the COVID-19 pandemic showed that children had heightened nasal innate immune responses compared with adults. To evaluate the role of nasal viruses and bacteria in driving these responses, we performed cytokine profiling and comprehensive, symptom-agnostic testing for respiratory viruses and bacterial pathobionts in nasopharyngeal samples from children tested for SARS-CoV-2 in 2021-22 (n = 467). Respiratory viruses and/or pathobionts were highly prevalent (82% of symptomatic and 30% asymptomatic children; 90 and 49% for children <5 years). Virus detection and load correlated with the nasal interferon response biomarker CXCL10, and the previously reported discrepancy between SARS-CoV-2 viral load and nasal interferon response was explained by viral coinfections. Bacterial pathobionts correlated with a distinct proinflammatory response with elevated IL-1β and TNF but not CXCL10. Furthermore, paired samples from healthy 1-year-olds collected 1-2 wk apart revealed frequent respiratory virus acquisition or clearance, with mucosal immunophenotype changing in parallel. These findings reveal that frequent, dynamic host-pathogen interactions drive nasal innate immune activation in children.
Background Symptomatic patients who test negative for common viruses are an important possible source of unrecognised or emerging pathogens, but metagenomic sequencing of all samples is inefficient because of the low likelihood of finding a pathogen in any given sample. We aimed to determine whether nasopharyngeal CXCL10 screening could be used as a strategy to enrich for samples containing undiagnosed viruses. Methods In this pathogen surveillance and detection study, we measured CXCL10 concentrations from nasopharyngeal swabs from patients in the Yale New Haven health-care system, which had been tested at the Yale New Haven Hospital Clinical Virology Laboratory (New Haven, CT, USA). Patients who tested negative for a panel of respiratory viruses using multiplex PCR during Jan 23-29, 2017, or March 3-14, 2020, were included. We performed host and pathogen RNA sequencing (RNA-Seq) and analysis for viral reads on samples with CXCL10 higher than 1 ng/mL or CXCL10 testing and quantitative RT-PCR (RT-qPCR) for SARS-CoV-2. We used RNA-Seq and cytokine profiling to compare the host response to infection in samples that were virus positive (rhinovirus, seasonal coronavirus CoV-NL63, or SARS-CoV-2) and virus negative (controls). Findings During Jan 23-29, 2017, 359 samples were tested for ten viruses on the multiplex PCR respiratory virus panel (RVP). 251 (70%) were RVP negative. 60 (24%) of 251 samples had CXCL10 higher than 150 pg/mL and were identified for further analysis. 28 (47%) of 60 CXCL10-high samples were positive for seasonal coronaviruses. 223 (89%) of 251 samples were PCR negative for 15 viruses and, of these, CXCL10-based screening identified 32 (13%) samples for further analysis. Of these 32 samples, eight (25%) with CXCL10 concentrations higher than 1 ng/mL and sufficient RNA were selected for RNA-Seq. Microbial RNA analysis showed the presence of influenza C virus in one sample and revealed RNA reads from bacterial pathobionts in four (50%) of eight samples. Between March 3 and March 14, 2020, 375 (59%) of 641 samples tested negative for 15 viruses on the RVP. 32 (9%) of 375 samples had CXCL10 concentrations ranging from 100 pg/mL to 1000 pg/mL and four of those were positive for SARS-CoV-2. CXCL10 elevation was statistically significant, and a distinguishing feature was found in 28 (8%) of 375 SARS-CoV-2-negative samples versus all four SARS-CoV-2-positive samples (p=4 center dot 4 x 10(-5)). Transcriptomic signatures showed an interferon response in virus-positive samples and an additional neutrophil-high hyperinflammatory signature in samples with high amounts of bacterial pathobionts. The CXCL10 cutoff for detecting a virus was 166 center dot 5 pg/mL for optimal sensitivity and 1091 center dot 0 pg/mL for specificity using a clinic-ready automated microfluidics-based immunoassay. Interpretation These results confirm CXCL10 as a robust nasopharyngeal biomarker of viral respiratory infection and support host response-based screening followed by metagenomic sequencing of CXCL10-high Copyright (c) 2022 The Author(s). Published by Elsevier Ltd. This is an Open Access article under the CC BY 4.0 license.samples as a practical approach to incorporate clinical samples into pathogen discovery and surveillance efforts.
ObjectiveWe aimed to discover computationally-derived phenotypes of opioid-related patient presentations to the ED via clinical notes and structured electronic health record (EHR) data.MethodsThis was a retrospective study of ED visits from 2013-2020 across ten sites within a regional healthcare network. We derived phenotypes from visits for patients ≥18 years of age with at least one prior or current documentation of an opioid-related diagnosis. Natural language processing was used to extract clinical entities from notes, which were combined with structured data within the EHR to create a set of features. We performed latent dirichlet allocation to identify topics within these features. Groups of patient presentations with similar attributes were identified by cluster analysis.ResultsIn total 82,577 ED visits met inclusion criteria. The 30 topics were discovered ranging from those related to substance use disorder, chronic conditions, mental health, and medical management. Clustering on these topics identified nine unique cohorts with one-year survivals ranging from 84.2-96.8%, rates of one-year ED returns from 9-34%, rates of one-year opioid event 10-17%, rates of medications for opioid use disorder from 17-43%, and a median Carlson comorbidity index of 2-8. Two cohorts of phenotypes were identified related to chronic substance use disorder, or acute overdose.ConclusionsOur results indicate distinct phenotypic clusters with varying patient-oriented outcomes which provide future targets better allocation of resources and therapeutics. This highlights the heterogeneity of the overall population, and the need to develop targeted interventions for each population.
Introduction: Health systems are increasingly leveraging real-world data to improve community health outcomes. Methods: We conducted a longitudinal assessment of health outcomes in patients with severe hypertension using electronic health record data from the Sentara Healthcare System 2010-2021. Severe hypertension was defined as at least two consecutive blood pressure (BP) readings above 160/100 mmHg. We examined follow-up visit rates at 3 and 6 months, and BP control rates (<140/90 mmHg) at 6 and 12 months after the second BP elevation. We also assessed the incidence of cardiovascular disease (CVD) until the end of 2022. Results: The study included 75,657 patients with severe hypertension. Mean age was 61.5 (13.9) years; 55% were female, 57% were White, 37% were Black, and 21.3% had a history of CVD. The median follow-up time was 3.8 years. Among patients with severe hypertension, 72% and 84% had follow-up visits at 3 and 6 months; 38% and 41% achieved BP control at 6 and 12 months. Among patients without prior history of CVD, 22.5% experienced at least one cardiovascular event during the follow-up, including 8.7% with coronary arteriosclerosis, 11.9% with heart failure, 65.2% with cerebral infarction, and 3.5% with myocardial infarction. The median time to cardiovascular events was 793 days. The risk of CVD was significantly higher in Black and older patients, as well as those with blood pressure levels exceeding 180/120 mmHg. Conclusions: This study successfully identified a longitudinal digital cohort of patients with severe hypertension and linked it to health outcomes. The findings have implications for health systems utilizing real-world data in population health research.
Back Cover In article number 2300594, Bai, Hafler, Halene, Fan, and co-workers develop a microfluidic high-plex assay toward highly informative and fully integrated immunological and serological characterization combined in a single platform. The test results allow for rapid evaluation of a patient's immune status at the systems biology level and inform the strategies to help immunocompromised patients to manage COVID-19 and potentially other infectious diseases.
Over the last century, outbreaks and pandemics have occurred with disturbing regularity, necessitating advance preparation and large-scale, coordinated response. Here, we developed a machine learning predictive model of disease severity and length of hospitalization for COVID-19, which can be utilized as a platform for future unknown viral outbreaks. We combined untargeted metabolomics on plasma data obtained from COVID-19 patients (n = 111) during hospitalization and healthy controls (n = 342), clinical and comorbidity data (n = 508) to build this patient triage platform, which consists of three parts: (i) the clinical decision tree, which amongst other biomarkers showed that patients with increased eosinophils have worse disease prognosis and can serve as a new potential biomarker with high accuracy (AUC = 0.974), (ii) the estimation of patient hospitalization length with ± 5 days error (R 2 = 0.9765) and (iii) the prediction of the disease severity and the need of patient transfer to the intensive care unit. We report a significant decrease in serotonin levels in patients who needed positive airway pressure oxygen and/or were intubated. Furthermore, 5-hydroxy tryptophan, allantoin, and glucuronic acid metabolites were increased in COVID-19 patients and collectively they can serve as biomarkers to predict disease progression. The ability to quickly identify which patients will develop life-threatening illness would allow the efficient allocation of medical resources and implementation of the most effective medical interventions. We would advocate that the same approach could be utilized in future viral outbreaks to help hospitals triage patients more effectively and improve patient outcomes while optimizing healthcare resources.