Our study examined the heterogeneity of phenotype algorithms (PA) in literature on Alzheimer's disease (AD), major depressive disorder (MDD), and pulmonary arterial hypertension (PAI), focusing on the impact of PA differences on patient overlap and incidence rate variability across conditions in six observational databases. We reviewed 49 replicated PAs (13 for AD, 23 for MDD, and 13 for PAI) and found significant heterogeneity. These varied PAs identified distinct patient cohorts, resulting in significant incidence rate heterogeneity. Despite some papers reporting primary condition codes and inclusion. comprehensive documentation ensuring reproducibility was often lacking, underscoring the need for more transparent and robust research practices.
Objective: The objective of this study was to evaluate the discoverability of supporting research materials, including supporting documents, individual participant data (IPD), and associated publications, in US federally funded COVID-19 clinical study records in ClinicalTrials.gov (CTG). Methods: Study registration records were evaluated for (1) links to supporting documents, including protocols, informed consent forms, and statistical analysis plans; (2) information on how unaffiliated researchers may access IPD and, when applicable, the linking of the IPD record back to the CTG record; and (3) links to associated publications and, when applicable, the linking of the publication record back to the CTG record. Results: 206 CTG study records were included in the analysis. Few records shared supporting documents, with only 4% of records sharing all 3 document types. 27% of records indicated they intended to share IPD, with 45% of these providing sufficient information to request access to the IPD. Only 1 dataset record was located, which linked back to its corresponding CTG record. The majority of CTG records did not have links to publications (61%), and only 21% linked out to at least 1 results publication. All publication records linked back to their corresponding CTG records. Conclusion: With only 4% of records sharing all supporting document types, 12% sufficient information to access IPD, and 21% results publications, improvements can be made to the discoverability of research materials in federally funded, COVID-19 CTG records. Sharing these materials on CTG can increase their discoverability, therefore increasing the validity, transparency, and reusability of clinical research.
There are many initiatives attempting to harmonize data collection across human clinical studies using common data elements (CDEs). The increased use of CDEs in large prior studies can guide researchers planning new studies. For that purpose, we analyzed the All of Us (AoU) program, an ongoing US study intending to enroll one million participants and serve as a platform for numerous observational analyses. AoU adopted the OMOP Common Data Model to standardize both research (Case Report Form [CRF]) and real-world (imported from Electronic Health Records [EHRs]) data. AoU standardized specific data elements and values by including CDEs from terminologies such as LOINC and SNOMED CT. For this study, we defined all elements from established terminologies as CDEs and all custom concepts created in the Participant Provided Information (PPI) terminology as unique data elements (UDEs). We found 1 033 research elements, 4 592 element-value combinations and 932 distinct values. Most elements were UDEs (869, 84.1%), while most CDEs were from LOINC (103 elements, 10.0%) or SNOMED CT (60, 5.8%). Of the LOINC CDEs, 87 (53.1% of 164 CDEs) originated from previous data collection initiatives, such as PhenX (17 CDEs) and PROMIS (15 CDEs). On a CRF level, The Basics (12 of 21 elements, 57.1%) and Lifestyle (10 of 14, 71.4%) were the only CRFs with multiple CDEs. On a value level, 61.7% of distinct values are from an established terminology. AoU demonstrates the use of the OMOP model for integrating research and routine healthcare data (64 elements in both contexts), which allows for monitoring lifestyle and health changes outside the research setting. The increased inclusion of CDEs in large studies (like AoU) is important in facilitating the use of existing tools and improving the ease of understanding and analyzing the data collected, which is more challenging when using study specific formats.
BackgroundMaintenance drugs are used to treat chronic conditions. Several classes of maintenance drugs have attracted attention because of their potential to affect susceptibility to and severity of COVID-19. MethodsUsing claims data on 20% random sample of Part D Medicare enrollees from April to December 2020, we identified patients diagnosed with COVID-19. Using a nested case-control design, non-COVID-19 controls were identified by 1:5 matching on age, race, sex, dual-eligibility status, and geographical region. We identified usage of angiotensin-converting enzyme inhibitors (ACEI), angiotensin-receptor blockers (ARB), statins, warfarin, direct factor Xa inhibitors, P2Y12 inhibitors, famotidine and hydroxychloroquine based on Medicare prescription claims data. Using extended Cox regression models with time-varying propensity score adjustment we examined the independent effect of each study drug on contracting COVID-19. For severity of COVID-19, we performed extended Cox regressions on all COVID-19 patients, using COVID-19-related hospitalization and all-cause mortality as outcomes. Covariates included gender, age, race, geographic region, low-income indicator, and co-morbidities. To compensate for indication bias related to the use of hydroxychloroquine for the prophylaxis or treatment of COVID-19, we censored patients who only started on hydroxychloroquine in 2020. ResultsUp to December 2020, our sample contained 374,229 Medicare patients over 65 who were diagnosed with COVID-19. Among the COVID-19 patients, 278,912 (74.6%) were on at least one study drug. The three most common study drugs among COVID-19 patients were statins 187,374 (50.1%), ACEI 97,843 (26.2%) and ARB 83,290 (22.3%). For all three outcomes (diagnosis, hospitalization and death), current users of ACEI, ARB, statins, warfarin, direct factor Xa inhibitors and P2Y12 inhibitors were associated with reduced risks, compared to never users. Famotidine did not show consistent significant effects. Hydroxychloroquine did not show significant effects after censoring of recent starters. ConclusionMaintenance use of ACEI, ARB, warfarin, statins, direct factor Xa inhibitors and P2Y12 inhibitors was associated with reduction in risk of acquiring COVID-19 and dying from it.
In response to the COVID-19 pandemic many clinical studies have been initiated leading to the need for efficient ways to track and analyze study results. We expanded our previous project that tracked registered COVID-19 clinical studies to also track result articles generated from these studies. We conducted searches of ClinicalTrials.gov and PubMed to identify articles linked to COVID-19 studies, and developed criteria based on the trial phase, intervention, location, and record recency to develop a prioritized list of result publications. We found 760 articles linked to 419 interventional trials (15.7% of all 2 669 COVID-19 interventional trials as of 15 August 2021), with 418 identified via abstract-link in PubMed and 342 via registry-link in ClinicalTrials.gov. Of the 419 trials publishing at least one article, 123 (29.4%) have multiple linked publications. We used an attention score to develop a prioritized list of all publications linked to COVID-19 trials and identified 58 publications that are result articles from late phase (Phase 3) trials with at least one US site and multiple study record updates. For COVID-19 vaccine trials, we found 69 linked result articles for 40 trials (13.9% of 290 total COVID-19 vaccine trials). Our method allows for the efficient identification of important COVID-19 articles that report results of registered clinical trials and are connected via a structured article-trial link.
Measurement concepts are essential to observational healthcare research; however, a lack of concept harmonization limits the quality of research that can be done on multisite research networks. We developed five methods that used a combination of automated, semi-automated and manual approaches for generating measurement concept sets. We validated our concept sets by calculating their frequencies in cohorts from the Columbia University Irving Medical Center (CUIMC) database. For heart transplant patients, the preoperative frequencies of basic metabolic panel concept sets, which we generated by a semi-automated approach, were greater than 99%. We also made concept sets for lumbar puncture and coagulation panels, by automated and manual methods respectively.
Background: Routinely collected real world data (RWD) have great utility in aiding the novel coronavirus disease (COVID-19) pandemic response [1,2]. Here we present the international Observational Health Data Sciences and Informatics (OHDSI) [3] Characterizing Health Associated Risks, and Your Baseline Disease In SARS-COV-2 (CHARYBDIS) framework for standardisation and analysis of COVID-19 RWD. Methods: We conducted a descriptive cohort study using a federated network of data partners in the United States, Europe (the Netherlands, Spain, the UK, Germany, France and Italy) and Asia (South Korea and China). The study protocol and analytical package were released on 11 th June 2020 and are iteratively updated via GitHub [4]. Findings: We identified three non-mutually exclusive cohorts of 4,537,153 individuals with a clinical COVID-19 diagnosis or positive test, 886,193 hospitalized with COVID-19 , and 113,627 hospitalized with COVID-19 requiring intensive services . All comorbidities, symptoms, medications, and outcomes are described by cohort in aggregate counts, and are available in an interactive website: https://data.ohdsi.org/Covid19CharacterizationCharybdis/. Interpretation: CHARYBDIS findings provide benchmarks that contribute to our understanding of COVID-19 progression, management and evolution over time. This can enable timely assessment of real-world outcomes of preventative and therapeutic options as they are introduced in clinical practice.
Objective: This study was intended to (1) provide clinical trial data-sharing platform designers with insight into users’ experiences when attempting to evaluate and access datasets, (2) spark conversations about improving the transparency and discoverability of clinical trial data, and (3) provide a partial view of the current information-sharing landscape for clinical trials. Methods: We evaluated preview information provided for 10 datasets in each of 7 clinical trial data-sharing platforms between February and April 2019. Specifically, we evaluated the platforms in terms of the extent to which we found (1) preview information about the dataset, (2) trial information on ClinicalTrials.gov and other external websites, and (3) evidence of the existence of trial protocols and data dictionaries. Results: All seven platforms provided data previews. Three platforms provided information on data file format (e.g., CSV, SAS file). Three allowed batch downloads of datasets (i.e., downloading multiple datasets with a single request), whereas four required separate requests for each dataset. All but one platform linked to ClinicalTrials.gov records, but only one platform had ClinicalTrails.gov records that linked back to the platform. Three platforms consistently linked to external websites and primary publications. Four platforms provided evidence of the presence of a protocol, and six platforms provided evidence of the presence of data dictionaries. Conclusions: More work is needed to improve the discoverability, transparency, and utility of information on clinical trial data-sharing platforms. Increasing the amount of dataset preview information available to users could considerably improve the discoverability and utility of clinical trial data.
It is difficult to arrive at an efficient and widely acceptable set of common data elements (CDEs). Trial outcomes, as defined in a clinical trial registry, offer a large set of elements to analyze. However, all clinical trial outcomes is an overwhelming amount of information. One way to reduce this amount of data to a usable volume is to only use a subset of trials. Our method uses a subset of trials by considering trials that support drug approval (pivotal trials) by Food and Drug Administration. We identified a set of pivotal trials from FDA drug approval documents and used primary outcomes data for these trials to identify a set of important CDEs. We identified 76 CDEs out of a set of 172 data elements from 192 pivotal trials for 100 drugs. This set of CDEs, grouped by medical condition, can be considered as containing the most significant data elements.
The objective of this paper is to determine the temporal trend of the association of 66 comorbidities with human immunodeficiency virus (HIV) infection status among Medicare beneficiaries from 2000 through 2016. We harvested patient level encounter claims from a 17-year long 100% sample of Medicare records. We used the chronic conditions warehouse comorbidity flags to determine HIV infection status and presence of comorbidities. We prepared 1 data set per year for analysis. Our 17 study data sets are retrospective annualized patient level case histories where the comorbidity status reflects if the patient has ever met the comorbidity case definition from the start of the study to the analysis year. We implemented one logistic binary regression model per study year to discover the maximum likelihood estimate (MLE) of a comorbidity belonging to our binary classes of HIV+ or HIV- study populations. We report MLE and odds ratios by comorbidity and year. Of the 66 assessed comorbidities, 35 remained associated with HIV- across all model years, 19 remained associated with HIV+ across all model years. Three comorbidities changed association from HIV+ to HIV- and 9 comorbidities changed association from HIV- to HIV+. The prevalence of comorbidities associated with HIV infection changed over time due to clinical, social, and epidemiological reasons. Comorbidity surveillance can provide important insights into the understanding and management of HIV infection and its consequences.
As the SARS-CoV-2 virus (COVID-19) continues to affect people across the globe, there is limited understanding of the long term implications for infected patients1–3. While some of these patients have documented follow-ups on clinical records, or participate in longitudinal surveys, these datasets are usually designed by clinicians, and not granular enough to understand the natural history or patient experiences of ‘long COVID’. In order to get a complete picture, there is a need to use patient generated data to track the long-term impact of COVID-19 on recovered patients in real time. There is a growing need to meticulously characterize these patients’ experiences, from infection to months post-infection, and with highly granular patient generated data rather than clinician narratives. In this work, we present a longitudinal characterization of post-COVID-19 symptoms using social media data from Twitter. Using a combination of machine learning, natural language processing techniques, and clinician reviews, we mined 296,154 tweets to characterize the post-acute infection course of the disease, creating detailed timelines of symptoms and conditions, and analyzing their symptomatology during a period of over 150 days.
OBJECTIVES:To characterize the demographics, comorbidities, symptoms, in-hospital treatments, and health outcomes among children and adolescents diagnosed or hospitalized with coronavirus disease 2019 (COVID-19) and to compare them in secondary analyses with patients diagnosed with previous seasonal influenza in 2017-2018. METHODS:International network cohort using real-world data from European primary care records (France, Germany, and Spain), South Korean claims and US claims, and hospital databases. We included children and adolescents diagnosed and/or hospitalized with COVID-19 at age <18 between January and June 2020. We described baseline demographics, comorbidities, symptoms, 30-day in-hospital treatments, and outcomes including hospitalization, pneumonia, acute respiratory distress syndrome, multisystem inflammatory syndrome in children, and death. RESULTS:A total of 242 158 children and adolescents diagnosed and 9769 hospitalized with COVID-19 and 2 084 180 diagnosed with influenza were studied. Comorbidities including neurodevelopmental disorders, heart disease, and cancer were more common among those hospitalized with versus diagnosed with COVID-19. Dyspnea, bronchiolitis, anosmia, and gastrointestinal symptoms were more common in COVID-19 than influenza. In-hospital prevalent treatments for COVID-19 included repurposed medications (<10%) and adjunctive therapies: systemic corticosteroids (6.8%-7.6%), famotidine (9.0%-28.1%), and antithrombotics such as aspirin (2.0%-21.4%), heparin (2.2%-18.1%), and enoxaparin (2.8%-14.8%). Hospitalization was observed in 0.3% to 1.3% of the cohort diagnosed with COVID-19, with undetectable (n < 5 per database) 30-day fatality. Thirty-day outcomes including pneumonia and hypoxemia were more frequent in COVID-19 than influenza. CONCLUSIONS:Despite negligible fatality, complications including hospitalization, hypoxemia, and pneumonia were more frequent in children and adolescents with COVID-19 than with influenza. Dyspnea, anosmia, and gastrointestinal symptoms could help differentiate diagnoses. A wide range of medications was used for the inpatient management of pediatric COVID-19.
Medicaid is a significant health insurance plan providing healthcare coverage to up to a third of the population of the United Sates. We describe two different formats of Medicaid data within Center for Medicare and Medicaid Services Virtual Research Data Center. We analyze record length, age and enrollment justification among patients for both data formats. As of December 2016, the total size of Medicaid population available from CMS is 92,953,389; 45% of patients are aged 0 to 18, 26.6% are aged 19-35 and 23.2% are aged 36-64. In terms of Medicaid eligibility, 35.6% qualify due to (child) age and 26.8% qualify due to income. We also compare the volume of Medicaid to Medicare for year 2016. We conclude that Medicaid data includes patients with significant record lengths and relatively well documented enrollment justification, which are high value assets for data reuse researchers that are willing to balance known data limitations with careful analysis design and interpretation.
Background With increasing use of real world data in observational health care research, data quality assessment of these data is equally gaining in importance. Electronic health record (EHR) or claims datasets can differ significantly in the spectrum of care covered by the data. Objective In our study, we link provider specialty with diagnoses (encoded in International Classification of Diseases) with a motivation to characterize data completeness. Methods We develop a set of measures that determine diagnostic span of a specialty (how many distinct diagnosis codes are generated by a specialty) and specialty span of a diagnosis (how many specialties diagnose a given condition). We also analyze ranked lists for both measures. As use case, we apply these measures to outpatient Medicare claims data from 2016 (3.5 billion diagnosis-specialty pairs). We analyze 82 distinct specialties present in Medicare claims (using Medicare list of specialties derived from level III Healthcare Provider Taxonomy Codes). Results A typical specialty diagnoses on average 4,046 distinct diagnosis codes. It can range from 33 codes for medical toxicology to 25,475 codes for internal medicine. Specialties with large visit volume tend to have large diagnostic span. Median specialty span of a diagnosis code is 8 specialties with a range from 1 to 82 specialties. In total, 13.5% of all observed diagnoses are generated exclusively by a single specialty. Quantitative cumulative rankings reveal that some diagnosis codes can be dominated by few specialties. Using such diagnoses in cohort or outcome definitions may thus be vulnerable to incomplete specialty coverage of a given dataset. Conclusion We propose specialty fingerprinting as a method to assess data completeness component of data quality. Datasets covering a full spectrum of care can be used to generate reference benchmark data that can quantify relative importance of a specialty in constructing diagnostic history elements of computable phenotype definitions.
Many research sponsors require sharing of data from human clinical trials. We created the CONSIDER statement, a set of recommendations to improve data sharing practices and increase the availability and re-usability of individual participant data from clinical trials. We developed the recommendations by reviewing shared individual participant data and study artifacts from a set of completed studies, as well as study data deposited on ClinicalTrials.gov and on several data sharing platforms. The CONSIDER statement is comprised of seven sections including: format, data sharing, study design, case report forms, data dictionary, data de-identification and choice of data sharing platform. We developed several different forms of CONSIDER which includes a brief form (the checklist), a full form (detailed descriptions and examples), and a scoring methodology. The checklist can be used to evaluate adherence to various progressive data sharing recommendations. We are currently in Phase 2 of collecting feedback on the CONSIDER statement.
HIV medication adherence is a topic of major public health concern in the United States. Adherent patients may be less likely to experience treatment failure, AIDS presentations and extreme medical costs. We evaluate a cohort of highly adherent Medicare beneficiaries to establish if the out of pocket costs of HIV medications are an inherent barrier to adherence. We analyzed a 100% sample of Medicare Part-D prescription medications. The drug and out ofpocket costs for HIV and non-HIV medications of highly adherent cohort were extracted and analyzed. The average gross drug cost per beneficiary was $34,029for HIV medications and $11,439for non-HIV medications. Average out of pocket costs per beneficiary was $454for HIV medications and $129 for non-HIV medications. Out of pocket costs do not reasonably appear to be a barrier to adherence for Part-D beneficiaries.
The premise of Open Science is that research and medical management will progress faster if data and knowledge are openly shared. The value of Open Science is nowhere more important and appreciated than in the rare disease (RD) community. Research into RDs has been limited by insufficient patient data and resources, a paucity of trained disease experts, and lack of therapeutics, leading to long delays in diagnosis and treatment. These issues can be ameliorated by following the principles and practices of sharing that are intrinsic to Open Science. Here, we describe how the RD community has adopted the core pillars of Open Science, adding new initiatives to promote care and research for RD patients and, ultimately, for all of medicine. We also present recommendations that can advance Open Science more globally.