Interoperability between data sources, one of the FAIR (Findability, Accessibility, Interoperability, and Reusability) principles for scientific data management, can enable multi-modality research. The purpose of our study was to investigate the potential for interoperability between an imaging resource, the Medical Imaging and Data Resource Center (MIDRC), and a clinical record resource, the National COVID Cohort Collaborative (N3C). The use case was the prediction of COVID-19 severity, defined as evidence for invasive ventilatory support, extracorporeal membrane oxygenation, death, or discharge to hospice in the N3C clinical record. Patient-level matching between MIDRC and N3C was identified using Privacy Preserving Record Linking via an honest broker. We identified positive COVID-19 tests and chest radiograph procedures in N3C and used the interval between them to identify images with matching intervals in MIDRC. Of the 236 patients (306 unique images) meeting initial inclusion criteria in MIDRC, 117 patients (and 139 unique images) remained after date interval matching between repositories and exclusion of patients with multiple potential matches. The Charlson Comorbidity Index (CCI) and the minimum mean arterial pressure (MAP) on the day of the chest radiograph were used as clinical indicators. The AUC in the task of predicting severe COVID-19 was evaluated using the computer-extracted imaging index alone (MIDRC), clinical indicators alone (N3C), and both together. Our model combining imaging and clinical indicators (CCI over 2 and MAP below 70) to predict severe COVID had an AUC of 0.73 (95% CI 0.62-0.84), and the models including imaging or clinical indicators alone were 0.67 (95% CI 0.56-0.79) and 0.69 (95% CI 0.59-0.80), respectively. This study highlights the potential for cross-platform data sharing to facilitate future multi-modality research and broader collaborative studies.
Objective: Determine the incidence of vestibular disorders in patients with SARS-CoV-2 compared to the control population. Study Design: Retrospective. Setting: Clinical data in the National COVID Cohort Collaborative database (N3C). Methods: Deidentified patient data from the National COVID Cohort Collaborative database (N3C) were queried based on variant peak prevalence (untyped, alpha, delta, omicron 21K, and omicron 23A) from covariants.org to retrospectively analyze the incidence of vestibular disorders in patients with SARS-CoV-2 compared to control population, consisting of patients without documented evidence of COVID infection during the same period. Results: Patients testing positive for COVID-19 were significantly more likely to have a vestibular disorder compared to the control population. Compared to control patients, the odds ratio of vestibular disorders was significantly elevated in patients with untyped (odds ratio [OR], 2.39; confidence intervals [CI], 2.29–2.50; P < 0.001), alpha (OR, 3.63; CI, 3.48–3.78; P < 0.001), delta (OR, 3.03; CI, 2.94–3.12; P < 0.001), omicron 21K variant (OR, 2.97; CI, 2.90–3.04; P < 0.001), and omicron 23A variant (OR, 8.80; CI, 8.35–9.27; P < 0.001). Conclusions: The incidence of vestibular disorders differed between COVID-19 variants and was significantly elevated in COVID-19-positive patients compared to the control population. These findings have implications for patient counseling and further research is needed to discern the long-term effects of these findings.
Background: Post-acute sequelae of COVID-19 (PASC) produce significant morbidity, prompting evaluation of interventions that might lower risk. Selective serotonin reuptake inhibitors (SSRIs) potentially could modulate risk of PASC via their central, hypothesized immunomodulatory, and/or antiplatelet properties although clinical trial data are lacking. Materials and Methods: This retrospective study was conducted leveraging real-world clinical data within the National COVID Cohort Collaborative (N3C) to evaluate whether SSRIs with agonist activity at the sigma-1 receptor (S1R) lower the risk of PASC, since agonism at this receptor may serve as a mechanism by which SSRIs attenuate an inflammatory response. Additionally, determine whether the potential benefit could be traced to S1R agonism. Presumed PASC was defined based on a computable PASC phenotype trained on the U09.9 ICD-10 diagnosis code. Results: Of the 17,908 patients identified, 1521 were exposed at baseline to a S1R agonist SSRI, 1803 to a nonS1R agonist SSRI, and 14,584 to neither. Using inverse probability weighting and Poisson regression, relative risk (RR) of PASC was assessed. A 29% reduction in the RR of PASC (0.704 [95% CI, 0.58-0.85]; P = 4 x10-4) was seen among patients who received an S1R agonist SSRI compared to SSRI unexposed patients and a 21% reduction in the RR of PASC was seen among those receiving an SSRI without S1R agonist activity (0.79 [95% CI, 0.67 - 0.93]; P = 0.005). Thus, SSRIs with and without reported agonist activity at the S1R were associated with a significant decrease in the risk of PASC.
Abstract Introduction Research driven by real‐world clinical data is increasingly vital to enabling learning health systems, but integrating such data from across disparate health systems is challenging. As part of the NCATS National COVID Cohort Collaborative (N3C), the N3C Data Enclave was established as a centralized repository of deidentified and harmonized COVID‐19 patient data from institutions across the US. However, making this data most useful for research requires linking it with information such as mortality data, images, and viral variants. The objective of this project was to establish privacy‐preserving record linkage (PPRL) methods to ensure that patient‐level EHR data remains secure and private when governance‐approved linkages with other datasets occur. Methods Separate agreements and approval processes govern N3C data contribution and data access. The Linkage Honest Broker (LHB), an independent neutral party (the Regenstrief Institute), ensures data linkages are robust and secure by adding an extra layer of separation between protected health information and clinical data. The LHB's PPRL methods (including algorithms, processes, and governance) match patient records using “deidentified tokens,” which are hashed combinations of identifier fields that define a match across data repositories without using patients' clear‐text identifiers. Results These methods enable three linkage functions: Deduplication, Linking Multiple Datasets, and Cohort Discovery. To date, two external repositories have been cross‐linked. As of March 1, 2023, 43 sites have signed the LHB Agreement; 35 sites have sent tokens generated for 9 528 998 patients. In this initial cohort, the LHB identified 135 037 matches and 68 596 duplicates. Conclusion This large‐scale linkage study using deidentified datasets of varying characteristics established secure methods for protecting the privacy of N3C patient data when linked for research purposes. This technology has potential for use with registries for other diseases and conditions.
Background Multi-institution electronic health records (EHR) are a rich source of real world data (RWD) for generating real world evidence (RWE) regarding the utilization, benefits and harms of medical interventions. They provide access to clinical data from large pooled patient populations in addition to laboratory measurements unavailable in insurance claims-based data. However, secondary use of these data for research requires specialized knowledge and careful evaluation of data quality and completeness. We discuss data quality assessments undertaken during the conduct of prep-to-research, focusing on the investigation of treatment safety and effectiveness. Methods Using the National COVID Cohort Collaborative (N3C) enclave, we defined a patient population using criteria typical in non-interventional inpatient drug effectiveness studies. We present the challenges encountered when constructing this dataset, beginning with an examination of data quality across data partners. We then discuss the methods and best practices used to operationalize several important study elements: exposure to treatment, baseline health comorbidities, and key outcomes of interest. Results We share our experiences and lessons learned when working with heterogeneous EHR data from over 65 healthcare institutions and 4 common data models. We discuss six key areas of data variability and quality. (1) The specific EHR data elements captured from a site can vary depending on source data model and practice. (2) Data missingness remains a significant issue. (3) Drug exposures can be recorded at different levels and may not contain route of administration or dosage information. (4) Reconstruction of continuous drug exposure intervals may not always be possible. (5) EHR discontinuity is a major concern for capturing history of prior treatment and comorbidities. Lastly, (6) access to EHR data alone limits the potential outcomes which can be used in studies. Conclusions The creation of large scale centralized multi-site EHR databases such as N3C enables a wide range of research aimed at better understanding treatments and health impacts of many conditions including COVID-19. As with all observational research, it is important that research teams engage with appropriate domain experts to understand the data in order to define research questions that are both clinically important and feasible to address using these real world data.
Effective small molecule therapies to combat the SARS-CoV-2 infection are still lacking as the COVID-19 pandemic continues globally. High throughput screening assays are needed for lead discovery and optimization of small molecule SARS-CoV-2 inhibitors. In this work, we have applied viral pseudotyping to establish a cell-based SARS-CoV-2 entry assay. Here, the pseudotyped particles (PP) contain SARS-CoV-2 spike in a membrane enveloping both the murine leukemia virus (MLV) gag-pol polyprotein and luciferase reporter RNA. Upon addition of PP to HEK293-ACE2 cells, the SARS-CoV-2 spike protein binds to the ACE2 receptor on the cell surface, resulting in priming by host proteases to trigger endocytosis of these particles, and membrane fusion between the particle envelope and the cell membrane. The internalized luciferase reporter gene is then expressed in cells, resulting in a luminescent readout as a surrogate for spike-mediated entry into cells. This SARS-CoV-2 PP entry assay can be executed in a biosafety level 2 containment lab for high throughput screening. From a collection of 5,158 approved drugs and drug candidates, our screening efforts identified 7 active compounds that inhibited the SARS-CoV-2-S PP entry. Of these seven, six compounds were active against live replicating SARS-CoV-2 virus in a cytopathic effect assay. Our results demonstrated the utility of this assay in the discovery and development of SARS-CoV-2 entry inhibitors as well as the mechanistic study of anti-SARS-CoV-2 compounds. Additionally, particles pseudotyped with spike proteins from SARS-CoV-2 B.1.1.7 and B.1.351 variants were prepared and used to evaluate the therapeutic effects of viral entry inhibitors.
Asymptomatic SARS-CoV-2 infection and delayed implementation of diagnostics have led to poorly defined viral prevalence rates in the United States and elsewhere. To address this, we analyzed seropositivity in 9089 adults in the United States who had not been diagnosed previously with COVID-19. Individuals with characteristics that reflected the U.S. population ( n = 27,716) were selected by quota sampling from 462,949 volunteers. Enrolled participants ( n = 11,382) provided medical, geographic, demographic, and socioeconomic information and dried blood samples. Survey questions coincident with the Behavioral Risk Factor Surveillance System survey, a large probability-based national survey, were used to adjust for selection bias. Most blood samples (88.7%) were collected between 10 May and 31 July 2020 and were processed using ELISA to measure seropositivity (IgG and IgM antibodies against SARS-CoV-2 spike protein and the spike protein receptor binding domain). The overall weighted undiagnosed seropositivity estimate was 4.6% (95% CI, 2.6 to 6.5%), with race, age, sex, ethnicity, and urban/rural subgroup estimates ranging from 1.1% to 14.2%. The highest seropositivity estimates were in African American participants; younger, female, and Hispanic participants; and residents of urban centers. These data indicate that there were 4.8 undiagnosed SARS-CoV-2 infections for every diagnosed case of COVID-19, and an estimated 16.8 million infections were undiagnosed by mid-July 2020 in the United States.
Abstract Objective Coronavirus disease 2019 (COVID-19) poses societal challenges that require expeditious data and knowledge sharing. Though organizational clinical data are abundant, these are largely inaccessible to outside researchers. Statistical, machine learning, and causal analyses are most successful with large-scale data beyond what is available in any given organization. Here, we introduce the National COVID Cohort Collaborative (N3C), an open science community focused on analyzing patient-level data from many centers. Materials and Methods The Clinical and Translational Science Award Program and scientific community created N3C to overcome technical, regulatory, policy, and governance barriers to sharing and harmonizing individual-level clinical data. We developed solutions to extract, aggregate, and harmonize data across organizations and data models, and created a secure data enclave to enable efficient, transparent, and reproducible collaborative analytics. Results Organized in inclusive workstreams, we created legal agreements and governance for organizations and researchers; data extraction scripts to identify and ingest positive, negative, and possible COVID-19 cases; a data quality assurance and harmonization pipeline to create a single harmonized dataset; population of the secure data enclave with data, machine learning, and statistical analytics tools; dissemination mechanisms; and a synthetic data pilot to democratize data access. Conclusions The N3C has demonstrated that a multisite collaborative learning health network can overcome barriers to rapidly build a scalable infrastructure incorporating multiorganizational clinical data for COVID-19 analytics. We expect this effort to save lives by enabling rapid collaboration among clinicians, researchers, and data scientists to identify treatments and specialized care and thereby reduce the immediate and long-term impacts of COVID-19.
The extent of SARS-CoV-2 infection throughout the United States population is currently unknown. High quality serology is key to avoiding medically costly diagnostic errors, as well as to assuring properly informed public health decisions. Here, we present an optimized ELISA-based serology protocol, from antigen production to data analyses, that helps define thresholds for IgG and IgM seropositivity with high specificities. Validation of this protocol is performed using traditionally collected serum as well as dried blood on mail-in blood sampling kits. Archival (pre-2019) samples are used as negative controls, and convalescent, PCR-diagnosed COVID-19 patient samples serve as positive controls. Using this protocol, minimal cross-reactivity is observed for the spike proteins of MERS, SARS1, OC43 and HKU1 viruses, and no cross reactivity is observed with anti-influenza A H1N1 HAI. Our protocol may thus help provide standardized, population-based data on the extent of SARS-CoV-2 seropositivity, immunity and infection.
Asymptomatic SARS-CoV-2 infection and delayed implementation of diagnostics have led to poorly defined viral prevalence rates. To address this, we analyzed seropositivity in US adults who have not previously been diagnosed with COVID-19. Individuals with characteristics that reflect the US population (n = 11,382) and who had not previously been diagnosed with COVID-19 were selected by quota sampling from 241,424 volunteers (ClinicalTrials.gov NCT04334954). Enrolled participants provided medical, geographic, demographic, and socioeconomic information and 9,028 blood samples. The majority (88.7%) of samples were collected between May 10th and July 31st, 2020. Samples were analyzed via ELISA for anti-Spike and anti-RBD antibodies. Estimation of seroprevalence was performed by using a weighted analysis to reflect the US population. We detected an undiagnosed seropositivity rate of 4.6% (95% CI: 2.6 - 6.5%). There was distinct regional variability, with heightened seropositivity in locations of early outbreaks. Subgroup analysis demonstrated that the highest estimated undiagnosed seropositivity within groups was detected in younger participants (ages 18-45, 5.9%), females (5.5%), Black/African American (14.2%), Hispanic (6.1%), and Urban residents (5.3%), and lower undiagnosed seropositivity in those with chronic diseases. During the first wave of infection over the spring/summer of 2020 an estimate of 4.6% of adults had a prior undiagnosed SARS-CoV-2 infection. These data indicate that there were 4.8 (95% CI: 2.8-6.8) undiagnosed cases for every diagnosed case of COVID-19 during this same time period in the United States, and an estimated 16.8 million undiagnosed cases by mid-July 2020.
Understanding the SARS-CoV-2 virus' pathways of infection, virus-host-protein interactions, and mechanisms of virus-induced cytopathic effects will greatly aid in the discovery and design of new therapeutics to treat COVID-19. Chloroquine and hydroxychloroquine, extensively explored as clinical agents for COVID-19, have multiple cellular effects including alkalizing lysosomes and blocking autophagy as well as exhibiting dose-limiting toxicities in patients. Therefore, we evaluated additional lysosomotropic compounds to identify an alternative lysosome-based drug repurposing opportunity. We found that six of these compounds blocked the cytopathic effect of SARS-CoV-2 in Vero E6 cells with half-maximal effective concentration (EC50) values ranging from 2.0 to 13 μM and selectivity indices (SIs; SI = CC50/EC50) ranging from 1.5- to >10-fold. The compounds (1) blocked lysosome functioning and autophagy, (2) prevented pseudotyped particle entry, (3) increased lysosomal pH, and (4) reduced (ROC-325) viral titers in the EpiAirway 3D tissue model. Consistent with these findings, the siRNA knockdown of ATP6V0D1 blocked the HCoV-NL63 cytopathic effect in LLC-MK2 cells. Moreover, an analysis of SARS-CoV-2 infected Vero E6 cell lysate revealed significant dysregulation of autophagy and lysosomal function, suggesting a contribution of the lysosome to the life cycle of SARS-CoV-2. Our findings suggest the lysosome as a potential host cell target to combat SARS-CoV-2 infections and inhibitors of lysosomal function could become an important component of drug combination therapies aimed at improving treatment and outcomes for COVID-19.