Long COVID affects a substantial proportion of the over 778 million individuals infected with SARS-CoV-2, yet predictive models remain limited in scope. While existing efforts, such as the National COVID Cohort Collaborative (N3C), have leveraged electronic health record (EHR) data for risk prediction, accumulating evidence points to additional contributions from social, behavioral, and genetic factors. Using a diverse cohort of SARS-CoV-2-infected individuals (n>17,200) from the NIH All of Us Research Program, we investigated whether integrating EHR data with survey-based and genomic information improves model performance. Our multi-scale approach outperformed EHR-only models original AUROC 0.736 (95% CI: 0.730, 0.741), achieving an AUROC of 0.748 (0.741,0.755). Among the top predictors, active-duty service status, self-reported fatigue, and chr19:4719431:G:A_A were among the most informative survey and genetic features. These findings highlight the importance of incorporating multi-scale data to improve risk stratification and inform personalized interventions for long COVID.
Electronic health records (EHRs) may be a promising alternative to traditional health surveys for population health surveillance due to their detailed health and patient information, low cost, and minimal respondent burden. However, concerns about the population representativeness of EHRs raise questions about their validity for public health monitoring and tracking social determinants of health. This study addresses these concerns by evaluating the representativeness of EHRs from UNC Health, a large integrated health delivery system in North Carolina, by linking individual-level EHRs (2018-2022; n = 2.12 million unique patients) with individual-level microdata from the nationally representative American Community Survey (ACS, 2018-2022). Specifically, we evaluate how demographic factors (age, sex, race/ethnicity), socioeconomic factors (education, employment, poverty, food stamps, public assistance), and health insurance impact the likelihood that a North Carolina ACS respondent will appear in the UNC Health EHRs. Linear probability models indicate that although UNC Health patients are not fully representative of the state population, selection biases are small and align with known patterns of healthcare utilization (e.g., overrepresentation among females, older adults, and individuals with health insurance). Moderate selection is observed by race/ethnicity and socioeconomic status, with overrepresentation at both the high and low ends of the socioeconomic spectrum. These findings provide cautious reassurance for the use of appropriately weighted EHR data in population health monitoring while demonstrating the value of evaluating and improving the utility of EHRs in public health research through linkages with individual-level nationally representative data.
The utility model offers a framework for ethical stewardship, patient empowerment, and distributed innovation.
Introduction:Healthcare organizations have begun incorporating screening procedures for social determinants of health (SDOH) into care, recognizing the impact these factors can have on health outcomes. We aimed to present methods for evaluating redundancy in the risk information gained across SDOH questions and for evaluating whether demographic biases are present in whether patients were asked SDOH questions and whether they declined to answer them. Methods:SDOH question data were analyzed for 1.8 million UNC Health patients. To evaluate risk information redundancy, response agreement was analyzed for pairs of questions. Demographic biases were evaluated using logistic regression models. Results:Risk information redundancy was identified, particularly across food and financial insecurity questions. Furthermore, female and White patients were more likely to be asked some questions than other groups, and American Indian or Alaska Native and Hispanic or Latino patients were less likely to decline to answer questions. Conclusions:We demonstrated methods healthcare organizations can use to evaluate their SDOH screening procedures. These methods yielded insights for (1) reducing burden in clinical workflows by identifying where redundancy could be eliminated and (2) reducing bias in SDOH data collection through more systematic screening protocols.
Objectives:To assess whether existing nationwide Health Information Networks (HINs) and Health Information Exchanges (HIEs) can be leveraged for electronic health record (EHR) data acquisition for research-specifically, for the NIH's All of Us Research Program. Materials and Methods:The All of Us Center for Linkage and Acquisition of Data (CLAD) collaborated with eHealth Exchange, the nation's largest HIN, to transmit participant-authorized queries to an HIE and a hospital system. Returned FHIR and C-CDA records were mapped to the OMOP Common Data Model and compared with existing All of Us EHR data. Results:Newly obtained records provided complementary information, enhancing completeness and research utility. Discussion:This pilot demonstrates the feasibility of using HINs/HIEs for patient-authorized data exchange for research. Barriers remain-including inconsistent technical capacity to transact authorizations and data quality variations. Conclusion:HINs and HIEs can support secure, patient-authorized EHR exchange for research, offering a pathway to more comprehensive, participant-driven biomedical discovery.
Electronic health record (EHR) data vary substantially in documentation density across patients, independent of disease burden. Existing tools such as the Charlson Comorbidity Index (CCI) and Elixhauser Comorbidity Index measure disease burden but do not capture differences in data volume, leaving a common source of bias unaddressed in EHR-based analyses. To address this gap, we developed the EHR Density Index (EDI), which characterizes the quantity, depth, and breadth of EHR data per patient per year, normalized by utilization patterns, using records from 24,987 adult patients at UNC Health (2018-2024). The EDI combines a utilization cluster assigned via Gaussian Mixture Model with within-cluster residuals quantifying documentation volume across four clinical domains. Four interpretable clusters emerged; while CCI predicted cluster membership, its associations with within-cluster residuals were weak, confirming the EDI captures dimensions of the patient record distinct from disease burden. The EDI is intended as a covariate to address documentation density as a source of confounding in real-world data-driven research.
Cancer registries enable cancer surveillance at the population level. These registries require significant human-time to read through many different parts of the electronic health record, including structured data and lengthy, free-text clinical reports, to abstract values for hundreds of required variables. Large language models (LLMs) offer the possibility to significantly improve this process by supporting and speeding up cancer registry data abstraction. However, it is unclear how well these models perform at real-world cancer registry abstraction involving multiple cancer types and large patient volumes. Here, we evaluate five foundational LLMs for their ability to reliably abstract cancer registry variables. We leverage hospital cancer registry data from a large regional health system as the ground truth and use LLMs to abstract from clinical reports eight registry variables for 5,939 patients with seven different cancer types. We use a zero-shot prompting strategy to compare LLM ability on commonly abstracted cancer variables with different data types. The results show that larger and more advanced models (Claude Sonnet 4.5, GPT-OSS-120b, GPT-OSS-20b) generally outperform smaller models (Gemma 12b, LLaMA 3.1 8b). The best performing models show F1 scores around 0.8 for cancer registry variables with low cardinality (grade, summary stage, laterality), with only slightly lower F1 scores for variables with high cardinality (primary site, regional nodes examined, regional nodes positive). On the more complex task of precise date extraction, all models showed decreased performance on both diagnosis and treatment dates (exact accuracy ~0.55 for the best performing models), which increased to ~0.85 for a tolerance within ±30 days. These results quantify the performance of various models as well as the potential and limitations of LLMs in cancer registry abstraction tasks.
Abstract Purpose Persistent opioid use is one of the most common post-operative complications. Identification of at-risk patients pre-operatively is key to reducing post-operative opioid use. We sought to develop a predictive model for persistent post-operative opioid used and to determine if geographic factors from community databases improve model prediction based solely on electronic health records (EHRs) and claims data. Methods EHR and claims data for 4,116 opioid-naïve surgical patients older than 18 in North Carolina were linked with census tract-level unemployment data from the American Community Survey and Centers for Disease Control and Prevention data on opioid prescriptions and deaths attributed to drug poisoning. Primary outcome was new persistent opioid use and covariates included patient factors from EHR, claims data, and geographic factors. Multivariable logistic regression models of potential risk factors were evaluated. Results 6.0% of patients developed new persistent opioid use. Associated risk factors based on multivariable logistic regressions include age (adjusted odds ratio [AOR] 1.08; 95% confidence interval [CI] 1.00, 1.16), back and neck pain (1.82; 1.39, 2.39), joint disorders (1.58; 1.18, 2.11), mood disorders (1.71; 1.28, 2.28), opioid retail prescription (1.04; 1.00, 1.07) and drug poisoning rates (1.33; 1.09, 1.62). On Monte-Carlo cross-validation, the addition of geographic factors to EHRs and claims may modestly improve prediction performance (area under the curve, AUC) of logistic regression models compared to those based on EHRs and claims data (AUC 0.667 (95% CI 0.619, 0.717) vs AUC 0.653 (0.600, 0.706)). Conclusions Co-morbidities and area-based factors are predictive of new persistent post-operative opioid use. As the addition of geographic-based factors did not significantly improve performance of multivariable logistic regression, larger samples are needed to fully differentiate models.
Recent years have seen an increase in the number and size of integrated health care delivery systems in the USA. The size and sophistication of these systems afford a greater focus on population health, leading to a fundamental question: How do the patients of these systems compare to the underlying regional populations that the systems serve? To demonstrate an approach to answering this question for a large public integrated delivery system, with a particular focus on neighborhood social determinants of health (SDOH). We present a descriptive, graphical comparison of the neighborhood characteristics of UNC Health patients and the overall population of North Carolina (NC). We leveraged electronic health record data from a 5-year period for patients at UNC Health, an integrated health care delivery system focused on serving the NC population. Estimates for the NC population were obtained from the American Community Survey (ACS). Measures included neighborhood SDOH indices for NC census tracts derived from ACS data as well as race and ethnicity. Overall, patients were more concentrated in neighborhoods with the least and greatest disadvantage. However, the density patterns of specific racial and ethnic groups across neighborhood SDOH scores were similar between the patients and NC population. Using a large, public integrated health care delivery system, we illustrate an approach for comparing the demographic and neighborhood characteristics of the patients of such a system and its underlying regional population using freely available data and open-source software. Our findings indicate many similar patterns between the health care system patients and regional population, but overall higher concentrations of patients in neighborhoods with the least and greatest disadvantage.
BACKGROUND:Incidence estimates of post-acute sequelae of severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) infection, also known as long COVID, have varied across studies and changed over time. We estimated long COVID incidence among adult and pediatric populations in 3 nationwide research networks of electronic health records (EHRs) participating in the RECOVER (Researching COVID to Enhance Recovery) Initiative using different classification algorithms (computable phenotypes). METHODS:This EHR-based retrospective cohort study included adult and pediatric patients with documented acute SARS-CoV-2 infection and 2 control groups: contemporary coronavirus disease 2019 (COVID-19)-negative and historical patients (2019). We examined the proportion of individuals identified as having symptoms or conditions consistent with probable long COVID within 30-180 days after COVID-19 infection (incidence proportion). Each network (the National COVID Cohort Collaborative [N3C], National Patient-Centered Clinical Research Network [PCORnet], and PEDSnet) implemented its own long COVID definition. We introduced a harmonized definition for adults in a supplementary analysis. RESULTS:Overall, 4% of children and 10%-26% of adults developed long COVID, depending on computable phenotype used. Excess incidence among SARS-CoV-2 patients was 1.5% in children and ranged from 5% to 6% among adults, representing a lower-bound incidence estimation based on our control groups. Temporal patterns were consistent across networks, with peaks associated with introduction of new viral variants. CONCLUSIONS:Our findings indicate that preventing and mitigating long COVID remains a public health priority. Examining temporal patterns and risk factors for long COVID incidence informs our understanding of etiology and can improve prevention and management.
Background Nirmatrelvir with ritonavir (Paxlovid) is indicated for patients with Coronavirus Disease 2019 (COVID-19) who are at risk for progression to severe disease due to the presence of one or more risk factors. Millions of treatment courses have been prescribed in the United States alone. Paxlovid was highly effective at preventing hospitalization and death in clinical trials. Several studies have found a protective association in real-world data, but they variously used less recent study periods, correlational methods, and small, local cohorts. Their estimates also varied widely. The real-world effectiveness of Paxlovid remains uncertain, and it is unknown whether its effect is homogeneous across demographic strata. This study leverages electronic health record data in the National COVID Cohort Collaborative’s (N3C) repository to investigate disparities in Paxlovid treatment and to emulate a target trial assessing its effectiveness in reducing severe COVID-19 outcomes. Methods and findings This target trial emulation used a cohort of 703,647 patients with COVID-19 seen at 34 clinical sites across the United States between April 1, 2022 and August 28, 2023. Treatment was defined as receipt of a Paxlovid prescription within 5 days of the patient’s COVID-19 index date (positive test or diagnosis). To emulate randomization, we used the clone-censor-weight technique with inverse probability of censoring weights to balance a set of covariates including sex, age, race and ethnicity, comorbidities, community well-being index (CWBI), prior healthcare utilization, month of COVID-19 index, and site of care provision. The primary outcome was hospitalization; death was a secondary outcome. We estimated that Paxlovid reduced the risk of hospitalization by 39% (95% confidence interval (CI) [36%, 41%]; p < 0.001), with an absolute risk reduction of 0.9 percentage points (95% CI [0.9, 1.0]; p < 0.001), and reduced the risk of death by 61% (95% CI [55%, 67%]; p < 0.001), with an absolute risk reduction of 0.2 percentage points (95% CI [0.1, 0.2]; p < 0.001). We also conducted stratified analyses by vaccination status and age group. Absolute risk reduction for hospitalization was similar among patients that were vaccinated and unvaccinate, but was much greater among patients aged 65+ years than among younger patients. We observed disparities in Paxlovid treatment, with lower rates among black and Hispanic or Latino patients, and within socially vulnerable communities. This study’s main limitation is that it estimates causal effects using observational data and could be biased by unmeasured confounding. Conclusions In this study of Paxlovid’s real-world effectiveness, we observed that Paxlovid is effective at preventing hospitalization and death, including among vaccinated patients, and particularly among older patients. This remains true in the era of Severe Acute Respiratory Syndrome Coronavirus 2 (SARS-CoV-2) Omicron subvariants. However, disparities in Paxlovid treatment rates imply that the benefit of Paxlovid’s effectiveness is not equitably distributed.
Background:Preventing and treating post-acute sequelae of SARS-CoV-2 infection (PASC), commonly known as Long COVID, has become a public health priority. Researchers have begun to explore whether Paxlovid treatment in the acute phase of COVID-19 could help prevent the onset of PASC. Methods and Findings:We used electronic health records from the National Clinical Cohort Collaborative (N3C) to define a cohort of 410,026 patients who had COVID-19 since April 1, 2022, and were eligible for Paxlovid treatment due to risk for progression to severe COVID-19. We used the target trial emulation framework to estimate the effect of Paxlovid treatment on PASC incidence. The treatment group was defined as outpatients prescribed Paxlovid within five days of COVID-19 index, and the control group was defined as all patients meeting eligibility criteria not in the treatment group. The follow-up period was 180 days. We estimated overall PASC incidence using a computable phenotype. We also measured incident cognitive, fatigue, and respiratory symptoms in the post-acute period. Paxlovid treatment had a small effect on overall PASC incidence (relative risk [RR] 0.94; 95% CI [0.90, 0.99]; p=0.011). It had a slightly stronger protective effect against cognitive (RR 0.86; 95% CI [0.77, 0.95]; p<0.001) and fatigue (RR 0.92; 95% CI [0.86, 0.97]; p=0.002) symptoms. Conclusions:In this study, Paxlovid had a weaker preventative effect on PASC than in prior observational studies, suggesting that Paxlovid is unlikely to become a definitive solution for preventing PASC. Differing effects by symptom cluster suggest that the etiology of cognitive and fatigue symptoms may be more closely related to viral load than that of respiratory symptoms. Future research should explore potential heterogeneous treatment effects across PASC subphenotypes.
BACKGROUND:In 2021, we used the National COVID Cohort Collaborative (N3C) as part of the National Institutes of Health RECOVER Initiative to develop a machine learning pipeline to identify patients with a high probability of having post-acute sequelae of SARS-CoV-2 infection or long COVID. However, the increased home testing, missing documentation, and reinfections that characterise the pandemic beyond 2022 necessitated the re-engineering of our original model to account for these changes in the COVID-19 research landscape. METHODS:Trained on 72 745 patient records (36 238 with long COVID and 36 507 with no evidence of long COVID), our updated XGBoost model gathered data for each patient in overlapping 100-day periods that progressed through time and issued a probability of long COVID for each 100-day period. We ran the model on patients in N3C (n=5 875 065) who met at least one of the following criteria from Jan 1, 2020, to June 22, 2023: a U07·1 (COVID-19) diagnosis code; a positive SARS-CoV-2 test; a U09·9 (post-acute sequelae of SARS-CoV-2 infection) diagnosis code; a prescription for nirmatrelvir-ritonavir or remdesivir; or an M35·81 (multisystem inflammatory syndrome in children [MIS-C]) diagnosis code. Each patient was given a model score that predicted long COVID status for each 100-day window in which they were aged ≥18 years. If a patient had known acute COVID-19 during any 100-day window (including reinfections), we censored the data from 7 days before the diagnosis or positive test date to 28 days after. We ran the model on controls selected from pre-2020 data to assess the likelihood of false positives. FINDINGS:The updated model had an area under the receiver operating characteristic curve of 0·90. Precision and recall could be adjusted according to a given use case, depending on whether greater sensitivity or specificity was warranted. Using our model, we estimate the overall prevalence of long COVID among the COVID-19 positive cohort within N3C repository to be 10.4%. INTERPRETATION:By eschewing the COVID-19 index date as an anchor point for analysis, we can assess the probability of long COVID among patients who might have tested at home, or with suspected (but untested) cases of COVID-19, or multiple SARS-CoV-2 reinfections. We view this exercise as a model for maintaining and updating any machine learning pipeline used for clinical research and operations. FUNDING:National Institutes of Health RECOVER Initiative.
BACKGROUND:Shared symptoms and biological abnormalities between post-acute sequelae of SARS-CoV-2 infection (PASC) and myalgic encephalomyelitis/chronic fatigue syndrome (ME/CFS) could suggest common pathophysiological bases and would support coordinated treatment efforts. Empirical studies comparing these syndromes are needed to better understand their commonalities and differences. METHODS:We analyzed electronic health record data from 6.5 million adult patients from the National COVID Cohort Collaborative. PASC and ME/CFS diagnostic groups were defined based on recorded diagnoses, and other recorded conditions within the two groups were used to train separate machine learning-driven computable phenotypes (CPs). The most predictive conditions for each CP were examined and compared, and the overlap of patients labeled by each CP was examined. Condition records from the diagnostic groups were also used to statistically derive condition clusters. Rates of subphenotypes based on these clusters were compared between PASC and ME/CFS groups. RESULTS:Approximately half of patients labeled by one CP are also labeled by the other. Dyspnea, fatigue, and cognitive impairment are the most-predictive conditions shared by both CPs, whereas other most-predictive conditions are specific to one CP. Recorded conditions separate into cardiopulmonary, neurological, and comorbidity clusters, with the cardiopulmonary cluster showing partial specificity for the PASC groups. CONCLUSIONS:Data-driven approaches indicate substantial overlap in the condition records associated with PASC and ME/CFS diagnoses. Nevertheless, cardiopulmonary conditions are somewhat more commonly associated with PASC diagnosis, whereas other conditions, such as pain and sleep disturbances, are more associated with ME/CFS diagnosis. These findings suggest that symptom management approaches to these illnesses could overlap.
Post-Acute Sequelae of SARS-CoV-2 infection (PASC), also known as Long-COVID, encompasses a variety of complex and varied outcomes following COVID-19 infection that are still poorly understood. We clustered over 600 million condition diagnoses from 14 million patients available through the National COVID Cohort Collaborative (N3C), generating hundreds of highly detailed clinical phenotypes. Assessing patient clinical trajectories using these clusters allowed us to identify individual conditions and phenotypes strongly increased after acute infection. We found many conditions increased in COVID-19 patients compared to controls, and using a novel method to associate patients with clusters over time, we additionally found phenotypes specific to patient sex, age, wave of infection, and PASC diagnosis status. While many of these results reflect known PASC symptoms, the resolution provided by this unprecedented data scale suggests avenues for improved diagnostics and mechanistic understanding of this multifaceted disease.
Objectives To provide a foundational methodology for differentiating comorbidity patterns in subphenotypes through investigation of a multi-site dementia patient dataset.Materials and Methods Employing the National Clinical Cohort Collaborative Tenant Pilot (N3C Clinical) dataset, our approach integrates machine learning algorithms-logistic regression and eXtreme Gradient Boosting (XGBoost)-with a diagnostic hierarchical model for nuanced classification of dementia subtypes based on comorbidities and gender. The methodology is enhanced by multi-site EHR data, implementing a hybrid sampling strategy combining 65% Synthetic Minority Over-sampling Technique (SMOTE), 35% Random Under-Sampling (RUS), and Tomek Links for class imbalance. The hierarchical model further refines the analysis, allowing for layered understanding of disease patterns.Results The study identified significant comorbidity patterns associated with diagnosis of Alzheimer's, Vascular, and Lewy Body dementia subtypes. The classification models achieved accuracies up to 69% for Alzheimer's/Vascular dementia and highlighted challenges in distinguishing Dementia with Lewy Bodies. The hierarchical model elucidates the complexity of diagnosing Dementia with Lewy Bodies and reveals the potential impact of regional clinical practices on dementia classification.Conclusion Our methodology underscores the importance of leveraging multi-site datasets and tailored sampling techniques for dementia research. This framework holds promise for extending to other disease subtypes, offering a pathway to more nuanced and generalizable insights into dementia and its complex interplay with comorbid conditions.Discussion This study underscores the critical role of multi-site data analyzes in understanding the relationship between comorbidities and disease subtypes. By utilizing diverse healthcare data, we emphasize the need to consider site-specific differences in clinical practices and patient demographics. Despite challenges like class imbalance and variability in EHR data, our findings highlight the essential contribution of multi-site data to developing accurate and generalizable models for disease classification. This study aims to enhance our understanding and classification of dementia subtypes using data from multiple healthcare sites. Dementia includes forms like Alzheimer's, Vascular, and Lewy Body dementia, each with unique health conditions. Researchers analyzed data from 9 US sites using a multi-stage approach with machine learning techniques, specifically logistic regression and eXtreme Gradient Boosting (XGBoost).The methodology involved 3 steps. First, the dataset was refined to focus on well-represented dementia subtypes. Next, advanced techniques balanced the data for fair representation. Finally, machine learning models classified the dementia types based on comorbidities and gender differences, achieving up to 70% accuracy for Alzheimer's and Vascular dementia, but finding Lewy Body dementia more challenging. A hierarchical model was used to address site-specific variations, revealing disparities among sites and improving generalization across populations.This study highlights the complexity of diagnosing dementia subtypes and the limitations of single-site studies, which often suffer from biases. By leveraging data from multiple sites, the research underscores the importance of multi-site dataset analysis for better generalization. This approach enhances understanding of dementia and provides a framework applicable to other diseases.
Objective: Determine the incidence of vestibular disorders in patients with SARS-CoV-2 compared to the control population. Study Design: Retrospective. Setting: Clinical data in the National COVID Cohort Collaborative database (N3C). Methods: Deidentified patient data from the National COVID Cohort Collaborative database (N3C) were queried based on variant peak prevalence (untyped, alpha, delta, omicron 21K, and omicron 23A) from covariants.org to retrospectively analyze the incidence of vestibular disorders in patients with SARS-CoV-2 compared to control population, consisting of patients without documented evidence of COVID infection during the same period. Results: Patients testing positive for COVID-19 were significantly more likely to have a vestibular disorder compared to the control population. Compared to control patients, the odds ratio of vestibular disorders was significantly elevated in patients with untyped (odds ratio [OR], 2.39; confidence intervals [CI], 2.29–2.50; P < 0.001), alpha (OR, 3.63; CI, 3.48–3.78; P < 0.001), delta (OR, 3.03; CI, 2.94–3.12; P < 0.001), omicron 21K variant (OR, 2.97; CI, 2.90–3.04; P < 0.001), and omicron 23A variant (OR, 8.80; CI, 8.35–9.27; P < 0.001). Conclusions: The incidence of vestibular disorders differed between COVID-19 variants and was significantly elevated in COVID-19-positive patients compared to the control population. These findings have implications for patient counseling and further research is needed to discern the long-term effects of these findings.
Background While many patients seem to recover from SARS-CoV-2 infections, many patients report experiencing SARS-CoV-2 symptoms for weeks or months after their acute COVID-19 ends, even developing new symptoms weeks after infection. These long-term effects are called post-acute sequelae of SARS-CoV-2 (PASC) or, more commonly, Long COVID. The overall prevalence of Long COVID is currently unknown, and tools are needed to help identify patients at risk for developing long COVID. Methods A working group of the Rapid Acceleration of Diagnostics-radical (RADx-rad) program, comprised of individuals from various NIH institutes and centers, in collaboration with REsearching COVID to Enhance Recovery (RECOVER) developed and organized the Long COVID Computational Challenge (L3C), a community challenge aimed at incentivizing the broader scientific community to develop interpretable and accurate methods for identifying patients at risk of developing Long COVID. From August 2022 to December 2022, participants developed Long COVID risk prediction algorithms using the National COVID Cohort Collaborative (N3C) data enclave, a harmonized data repository from over 75 healthcare institutions from across the United States (U.S.). Findings Over the course of the challenge, 74 teams designed and built 35 Long COVID prediction models using the N3C data enclave. The top 10 teams all scored above a 0.80 Area Under the Receiver Operator Curve (AUROC) with the highest scoring model achieving a mean AUROC of 0.895. Included in the top submission was a visualization dashboard that built timelines for each patient, updating the risk of a patient developing Long COVID in response to clinical events. Interpretation As a result of L3C, federal reviewers identified multiple machine learning models that can be used to identify patients at risk for developing Long COVID. Many of the teams used approaches in their submissions which can be applied to future clinical prediction questions. Funding Research reported in this RADx® Rad publication was supported by the National Institutes of Health. Timothy Bergquist, Johanna Loomba, and Emily Pfaff were supported by Axle Subcontract: NCATS-STSS-P00438.
Observational studies based on cohorts built from electronic health records (EHR) form the backbone of our current understanding of the risk of new-onset diabetes following COVID. EHR-based research is a powerful tool for medical research but is subject to multiple sources of bias. In this viewpoint, we define key sources of bias that threaten the validity of EHR-based research on this topic (namely misclassification, selection, surveillance, immortal time, and confounding biases), describe their implications, and suggest best practices to avoid them in the context of COVID-diabetes research.
Abstract Background Although the COVID-19 pandemic has persisted for over 3 years, reinfections with SARS-CoV-2 are not well understood. We aim to characterize reinfection, understand development of Long COVID after reinfection, and compare severity of reinfection with initial infection. Methods We use an electronic health record study cohort of over 3 million patients from the National COVID Cohort Collaborative as part of the NIH Researching COVID to Enhance Recovery Initiative. We calculate summary statistics, effect sizes, and Kaplan–Meier curves to better understand COVID-19 reinfections. Results Here we validate previous findings of reinfection incidence (6.9%), the occurrence of most reinfections during the Omicron epoch, and evidence of multiple reinfections. We present findings that the proportion of Long COVID diagnoses is higher following initial infection than reinfection for infections in the same epoch. We report lower albumin levels leading up to reinfection and a statistically significant association of severity between initial infection and reinfection (chi-squared value: 25,697, p-value: <0.0001) with a medium effect size (Cramer’s V: 0.20, DoF = 3). Individuals who experienced severe initial and first reinfection were older in age and at a higher mortality risk than those who had mild initial infection and reinfection. Conclusions In a large patient cohort, we find that the severity of reinfection appears to be associated with the severity of initial infection and that Long COVID diagnoses appear to occur more often following initial infection than reinfection in the same epoch. Future research may build on these findings to better understand COVID-19 reinfections.