Rigid prespecification can be impractical for noninterventional studies using secondary datasets, where data-driven flexibility is often required. Using target trial emulations comparing immunomodulator treatments for COVID-19, we piloted an adaptive strategy that accommodates warranted mid-course refinements within a prespecified framework. Our preregistered protocol outlined an initial study plan along with predetermined diagnostic thresholds and contingencies. Implementation proceeded through sequential phases, allowing researcher decisions to be guided by prespecified criteria under varying degrees of blinding to results. The adaptive approach led to alterations in the underlying target trial and to the analysis plan used for emulation, strengthening the plausibility of causal assumptions and improving the relevance of findings. During the initial baseline phase, indicated contingencies included sample restrictions, redefining treatments from class-level to product-specific comparisons, a revised propensity score model, and weight truncation. In the subsequent postbaseline phase, diagnostic checks triggered a modified causal contrast, inverse probability of censor weighting to address noncompliance, cause-specific hazard estimation to contextualize competing events, and additional reporting of hazard ratios for progressively truncated follow-up periods. For a secondary study objective, the adaptive framework allowed for some iterative attempts to improve validity while providing a clear stopping point. Similar approaches could lend transparent structure to the process of learning what causal questions the data are equipped to support. Beyond guarding against researcher bias, prespecification of adaptive protocols may promote more robust designs by encouraging investigators to be explicit about their assumptions, strategies for interrogating those assumptions, and specific criteria for determining when and how deviations may be required.
PURPOSE:To characterize select laboratory tests ordered versus reported for patients diagnosed with COVID-19 in administrative healthcare and commercial laboratory data. METHODS:Among patients with an outpatient COVID-19 diagnosis claim in HealthVerity data (01/01/2021-12/31/2022), this study described baseline characteristics and descriptively compared SARS-CoV-2 diagnostic tests and liver function tests from administrative healthcare (insurance claims and hospital billing data) and commercial laboratories, overall and by code type (e.g., CPT, LOINC). Select liver function tests were also described by method-specific and methodless LOINC. RESULTS:Among 214 998 patients with COVID-19, 46.1% had a SARS-CoV-2 molecular diagnostic test recorded within 7 days of diagnosis (in either administrative or laboratory data); 44.5% had a corresponding CPT in medical claims, while only 10.0% had a corresponding LOINC in laboratory data. In contrast, the six most common liver function tests (albumin, aspartate aminotransferase, total protein, alkaline phosphatase, alanine aminotransferase, and total bilirubin) were identified in 55.7%-56.6% of patients via LOINC, but only in 3.2%-4.2% via CPT claims. Of the total count of select liver function tests performed in the laboratory data, 99.7% of aspartate aminotransferase, 96.1% of direct bilirubin, and 82.9% of lactate dehydrogenase were reported by methodless LOINC rather than method-specific LOINC. CONCLUSIONS:Important differences were identified between orders for SARS-CoV-2 diagnostic tests and liver function tests, as well as missingness of LOINC method, highlighting challenges related to completeness of laboratory data in real-world data sources. These challenges underscore a need to improve data quality when considering the utility of laboratory data for research.
PURPOSE:To understand the impact of standardizing administrative healthcare data to the Sentinel common data model for cohort selection and descriptive findings. METHODS:Among patients with an outpatient COVID-19 diagnosis (January 2021-December 2022) in HealthVerity using the data in its native and the standardized format, we descriptively compared cohort attrition and sample size, patient characteristics, and healthcare resource utilization during baseline and incidence of selected conditions after COVID-19 diagnosis. RESULTS:The standardized cohort included fewer patients than the native (164 445 vs. 198 317), but age (median 48 years) and sex (70% female) were the same. The distribution of race was similar; however, the standardized cohort mapped patients with "Other" race to the "Unknown/Missing" race category, which created differences among those categories. Distributions were similar, albeit slightly lower for comorbidities (differences < 1%), and lower for SARS-CoV-2 diagnostic tests (59% vs. 70%). Medical encounter counts were also lower, with substantial differences that were attenuated after limiting encounter counts to one event per day (e.g., mean count of 6.0 vs. 27.7 specialty care visits reduced to 2.9 vs. 3.5). Incidence rates were lower, with the greatest difference for hepatotoxicity (29.6 vs. 37.1 per 1000 person-years). CONCLUSIONS:The data standardization refines the data (e.g., removes duplicate claims and variables or variable categories), which may reduce outliers and errors but yield lower distributions and counts of certain variables than observed in native format data. Therefore, it is critical to understand how standardization impacts the data and subsequently its fitness for use.
In response to the COVID-19 pandemic, a collaborative public–private partnership was launched to harness evidence from rapidly accruing real-world data (RWD) in various healthcare settings, with the goal of characterizing and understanding COVID-19 in near real-time, by applying rigorous epidemiological methods and defining research best practices. Projects were conducted in 4 phases: Research Planning and Prioritization, Protocol Development, Protocol Implementation, and Results Dissemination. During these projects, areas were identified with a current or future need to enhance existing best practices. This report provides a summary of our research processes, including application of new and existing practices, along with key learnings related to the challenges of conducting research when the clinical landscape is rapidly evolving as was the case during the first year of the COVID-19 pandemic. Such processes and learnings may be helpful to the broader research community when using RWD to understand or address future public health priorities.
Artificial intelligence (AI) has rapidly evolved from experimental applications in pharmacovigilance (PV) to being considered for routine use. This review critically examines AI’s potential to revolutionize drug safety monitoring, focusing on practical implementation challenges such as ensuring AI’s consistent and transparent performance, reducing multiple sources of bias, and addressing interpretability issues. It emphasizes the transition from experimental use to a routine, scalable capability within PV. It examines AI’s evidence base in specific applications, its ability to enhance actionable insights, and how organizations can safeguard against unintended consequences in multi-AI system environments. These considerations are vital as AI moves from theory to practice in PV.
Background:There is a dearth of drug utilization studies for coronavirus disease 2019 (COVID-19) treatments in 2021 and beyond after the introduction of vaccines and updated guidelines; such studies are needed to contextualize ongoing COVID-19 treatment effectiveness studies during these time periods. This study describes utilization patterns for corticosteroids, interleukin-6 (IL-6) inhibitors, Janus kinase inhibitors, and remdesivir among hospitalized adults with COVID-19, over the entire hospitalization, and within hospitalization periods categorized by respiratory support requirements. Methods:This descriptive cohort study included United States adults hospitalized with COVID-19 admitted from 1 January 2021 through 1 February 2022; data included HealthVerity claims and hospital chargemaster. The number and distribution of patients were reported for the first 3 drug regimen lines initiated. Results:The cohort included 51 066 patients; the most common initial drug regimens were corticosteroids (23.4%), corticosteroids plus remdesivir (25.1%), and remdesivir (4.4%). IL-6 inhibitors and Janus kinase inhibitors were included in later drug regimens and were more commonly administered with both corticosteroids and remdesivir than with corticosteroids alone. IL-6 inhibitors were more commonly administered than Janus kinase inhibitors when patients received high-flow oxygen or ventilation. Conclusions:These findings provide important context for comparative studies of COVID-19 treatments with study periods extending into 2021 and later. While prescribing generally aligned with National Institutes of Health COVID-19 treatment guidelines during this period, these findings suggest that prescribing preference, potential confounding by indication, and confounding by prior/concomitant use of other therapeutics should be considered in the design and interpretation of comparative studies.
BackgroundAs diagnostic tests for COVID-19 were broadly deployed under Emergency Use Authorization, there emerged a need to understand the real-world utilization and performance of serological testing across the United States. MethodsSix health systems contributed electronic health records and/or claims data, jointly developed a master protocol, and used it to execute the analysis in parallel. We used descriptive statistics to examine demographic, clinical, and geographic characteristics of serology testing among patients with RNA positive for SARS-CoV-2. ResultsAcross datasets, we observed 930,669 individuals with positive RNA for SARS-CoV-2. Of these, 35,806 (4%) were serotested within 90 days; 15% of which occurred <14 days from the RNA positive test. The proportion of people with a history of cardiovascular disease, obesity, chronic lung, or kidney disease; or presenting with shortness of breath or pneumonia appeared higher among those serotested compared to those who were not. Even in a population of people with active infection, race/ethnicity data were largely missing (>30%) in some datasets-limiting our ability to examine differences in serological testing by race. In datasets where race/ethnicity information was available, we observed a greater distribution of White individuals among those serotested; however, the time between RNA and serology tests appeared shorter in Black compared to White individuals. Test manufacturer data was available in half of the datasets contributing to the analysis. ConclusionOur results inform the underlying context of serotesting during the first year of the COVID-19 pandemic and differences observed between claims and EHR data sources-a critical first step to understanding the real-world accuracy of serological tests. Incomplete reporting of race/ethnicity data and a limited ability to link test manufacturer data, lab results, and clinical data challenge the ability to assess the real-world performance of SARS-CoV-2 tests in different contexts and the overall U.S. response to current and future disease pandemics.
PURPOSE:The U.S. Food and Drug Administration's Sentinel System is a national medical product safety surveillance system consisting of a large multisite distributed database of administrative claims supplemented by electronic health-care record data. The program seeks to improve data capture of race and ethnicity for pharmacoepidemiology studies.METHODS:We conducted a narrative literature review of published research on data augmentation and imputation methods to improve race and ethnicity capture in U.S. health-care systems databases. We focused on methods with limited (five-digit ZIP codes only) or full patient identifiers available to link to external sources of self-reported data. We organized the literature by themes: (1) variation in data capture of self-reported data, (2) data augmentation from external sources of self-reported data, and (3) imputation methods, including Bayesian analysis and multiple regression.RESULTS:Researchers reduced data missingness with high validity for Asian, Black, White, and Pacific Islander racial groups and Hispanic ethnicity. Native American and multiracial groups were difficult to validate due to relatively small sample sizes.CONCLUSIONS:Limitations on accessible self-reported data for validation will dictate methods to improve race and ethnicity data capture. We recommend methods leveraging multiple sources that account for variations in geography, age, and sex.
The terms real-world data (RWD) and real-world evidence (RWE) are often used inconsistently or interchangeably, including in submissions to the US Food and Drug Administration (FDA) involving RWE to evaluate the effectiveness of drugs and biologic products. Although a misconception sometimes exists that only non-interventional studies utilize RWD to generate RWE, the spectrum of study designs involves various combinations of data sources and design architectures which determine whether RWE is generated or not. Attempts by various stakeholders to identify the role of RWE in regulatory decision-making often do not focus on the contribution of RWD in specifically evaluating drug-outcome associations. Prior examples of FDA approvals demonstrate how RWD can be utilized to generate RWE as part of a marketing application for regulatory decision-making. In accordance with the 21st Century Cures Act of 2016 (Cures Act),1 the Food and Drug Administration (FDA or the Agency) launched a program for evaluating the use of real-world evidence (RWE) to support regulatory decision-making. The Agency's 2018 Framework for FDA's Real-World Evidence Program defines real-world data (RWD) as "data relating to patient health status and/or the delivery of health care routinely collected from a variety of sources" and RWE as "clinical evidence about the usage and potential benefits or risks of a medical product derived from analysis of RWD."2 Since passage of the Cures Act, the Center for Drug Evaluation and Research (CDER), in cooperation with the Center for Biologics Evaluation and Research (CBER) and the Oncology Center of Excellence, published a series of guidance documents related to RWD and RWE.3 (Although beyond the scope of this commentary, FDA's Center for Devices and Radiological Health (CDRH) published a guidance4 describing expectations for the use of RWD and RWE for regulatory decision-making for medical devices and also published examples5 of "RWE-based" CDRH approvals.) Despite the focus of the Cures Act on promoting research using RWD, types of data and study designs have not fundamentally changed since passage of the Cures Act. For example, the randomized controlled trial (RCT) is the archetype of study designs for assessing the safety and efficacy of a medical product. Other study design types can be used, however, including when randomization has feasibility challenges or ethical concerns. Although methodological challenges exist with nonrandomized research, such studies can contribute to drug development and support regulatory decision-making regarding effectiveness when adequately meeting evidentiary standards.6, 7 For example, data sources with well-characterized covariates and clinical endpoints are now more available for exploration using existing study design approaches and statistical methods in lieu of randomization. Nonrandomized studies also offer opportunities to study diverse patient populations and better understand long-term outcomes among patients receiving medical products. Although FDA has been using what is now called RWE for years to assess the safety of medical products, since passage of the Cures Act, we have observed increased submissions involving RWE to evaluate effectiveness of drugs and biological products that analyze data collected during routine clinical care. At the same time, and given varying operational definitions of RWD and RWE in the stakeholder community, these terms have often been used inconsistently and sometimes interchangeably. As a result, confusion can arise when similar data sources and study designs are characterized differently in different settings.6 This commentary addresses RWD/RWE terminology for drugs and biological products, specifically related to studies of effectiveness submitted to FDA and whether they are classified as RWE by the agency. We have encountered a misconception8 that only non-interventional (observational) research7 utilizes RWD to generate RWE—in other words, a dichotomy of randomized controlled trials versus real-world evidence is said to exist.6 In reality, the spectrum of study design involves various combinations of data sources and design architectures; see Table 1. For example, externally controlled trials that utilize RWD in the comparator arm generate RWE, despite the treatment arm generating data according to a study protocol in a clinical trial environment. As another example, a randomized trial generates RWE if the primary outcome is based on an assessment of RWD (often referred to as a point-of-care trial). Conversely, although RWD can be utilized to identify potential participants or trial sites in a traditional RCT (along with the ability to promote diversity of study populations), such data are not generating RWE to evaluate the drug-outcome association. In another scenario, a non-interventional study can use RWD to generate RWE even if a prespecified protocol exists to collect additional data, as in a patient registry, as long as the treatment is administered as part of routine clinical care (i.e., per a clinician's judgment).7 Additional considerations arise when characterizing data from various RWD sources. A specific consideration is how to characterize data generated from digital health technologies (DHTs), such as software applications and sensors. If a DHT is used in a clinical trial according to protocol-driven procedures, the data are not considered RWD. In contrast, if data are obtained from personal use of DHTs outside of research setting, the data are considered RWD. When such data are determined to be reliable and relevant, they can be used to generate RWE for regulatory purposes. Another consideration is how to characterize summary-level aggregated data from the medical literature. Although stakeholders sometimes classify non-patient-level data as RWD generating RWE, such information from the literature is not central to the goal of using RWE for regulatory approvals. For example, literature citations that report on disease prevalence or drug utilization are included in most regulatory submissions and may not be viewed as RWE for regulatory purposes. Two FDA approvals help illustrate how RWD can be utilized to generate RWE as part of a marketing application for regulatory decision making. In 2021, the FDA approved Prograf® (tacrolimus) in combination with other immunosuppressant drugs to prevent organ rejection in adult and pediatric patients receiving lung transplantation.9 The evidence in support of approval of this new indication included a non-interventional study using RWD from a US-based registry, compared to historical controls. The FDA considered the RWD fit-for-use and the non-interventional study using these data to be the adequate and well-controlled clinical investigation necessary for establishing substantial evidence of effectiveness for approval. RCTs of Prograf® in other (liver, kidney, and heart) transplant settings provided confirmatory evidence. Another scenario, wherein the RWE played a lesser role, was the approval in 2019 of Ibrance® (palbociclib) for male patients with metastatic breast cancer (MBC) in combination with letrozole, an aromatase inhibitor. The sponsor submitted a supplemental new drug application including data from prior RCTs (PALOMA-1, PALOMA-2, and PALOMA-3) including only women, along with descriptive analyses of electronic health record and medical claims data for men with MBC.10, 11 Substantial evidence of effectiveness relied on the previous RCT data in women, based on the knowledge that the natural history of the disease, response to therapy, and safety would be expected to be similar in men and women. The data from electronic health records on men provided information on safety, indicating the safety profile for the use of palbociclib in combination with hormonal therapies in men was consistent with the known adverse event profile. More generally, external attempts to identify the role of RWD and RWE in regulatory decision-making can lead to different characterizations than those made by the FDA.12-16 For example, an article14 examined the role of RWE in FDA-approved new drug and biologics license applications from 2019 to 2021 and identified studies as RWE when "used to support the application's therapeutic context (e.g., prevalence and incidence of a disease)." Another article15 tracking submissions to the European Medicines Agency from 2018 to 2019 considered the use of RWD "to assess the representativeness of the control arm [of a traditional randomized trial]" as constituting RWE for label expansion and marketing authorization. By not focusing directly on the evaluation of drug-outcome associations, these examples indicate how different applications of real-world terminology can create inconsistency when tracking approvals based on RWE across regulatory agencies. As part of the FDA RWE Program, published guidance17 includes recommendations for sponsors to accurately describe data sources and design attributes in their submission cover letter to the Agency. In that guidance,17 FDA recommends providing specific information describing the regulatory purpose, types of study design, and RWD source; see Table 2. More recently, FDA announced an Advancing RWE Program18 to fulfill a commitment under the Prescription Drug User Fee Act (PDUFA) VII for fiscal years 2023 through 2027. The program includes a new mechanism for identifying approaches to generate RWE that meet regulatory requirements in support of labeling for effectiveness; it also includes a commitment to publicly report on RWE submissions to CDER and CBER starting in 2024. Use of consistent terminology can help FDA classify and quantify RWE and promote better understanding of reports by external entities regarding the use of RWE for regulatory purposes. The FDA Real-World Evidence Program seeks to address current challenges in using RWE for regulatory decision-making, along a spectrum from supportive to pivotal contributions. Inconsistent use of the terms RWD and RWE complicates efforts among regulators to track such data and evidence, causing potential confusion during communications among regulatory agencies, sponsors, and other stakeholders. Although the distinction between studies that generate RWE and those that do not may at times seem confusing to some in the stakeholder community, careful consideration of data and design elements can help sponsors and regulators better describe and characterize RWE. There is no funding information to report. The authors declare no conflict of interest.
Unique challenges pertain when studying children, although many research principles are the same as those when studying adult populations. This truism extends to the use of real-world data (RWD). RWD are particularly relevant to pediatrics because they may potentially provide an additional source of data to inform pediatric labeling and practice patterns when clinical trials have not been or cannot be conducted. The purpose of this commentary is to provide a brief overview of the unique issues in using RWD to study the effectiveness or safety of medical therapies in children.
Purpose Algorithms for classification of inpatient COVID-19 severity are necessary for confounding control in studies using real-world data. Methods Using Healthverity chargemaster and claims data, we selected patients hospitalized with COVID-19 between April 2020 and February 2021, and classified them by severity at admission using an algorithm we developed based on respiratory support requirements (supplemental oxygen or non-invasive ventilation, O2/NIV, invasive mechanical ventilation, IMV, or NEITHER). To evaluate the utility of the algorithm, patients were followed from admission until death, discharge, or a 28-day maximum to report mortality risks and rates overall and by stratified by severity. Trends for heterogeneity in mortality risk and rate across severity classifications were evaluated using Cochran-Armitage and Logrank trend tests, respectively. Results Among 118 117 patients, the algorithm categorized patients in increasing severity as NEITHER (36.7%), O2/NIV (54.3%), and IMV (9.0%). Associated mortality risk (and 95% CI) was 11.8% (11.6-12.0%) overall and increased with severity [3.4% (3.2-3.5%), 11.5% (11.3-11.8%), 47.3% (46.3-48.2%); p < 0.001]. Mortality rate per 1000 person-days (and 95% CI) was 15.1 (14.9-15.4) overall and increased with severity [5.7 (5.4-6.0), 14.5 (14.2-14.9), 32.7 (31.8-33.6); p < 0.001]. Conclusion As expected, we observed a positive association between the algorithm-defined severity on admission and 28-day mortality risk and rate. Although performance remains to be validated, this provides some assurance that this algorithm may be used for confounding control or stratification in treatment effect studies.
Objective To describe differences by race and ethnicity in treatment patterns among hospitalized COVID-19 patients in the US from March-August 2020. Methods Among patients in de-identified Optum electronic health record data hospitalized with COVID-19 (March-August 2020), we estimated odds ratios of receiving COVID-19 treatments of interest (azithromycin, dexamethasone, hydroxychloroquine, remdesivir, and other steroids) at hospital admission, by race and ethnicity, after adjusting for key covariates of interest. Results After adjusting for key covariates, Black/African American patients were less likely to receive dexamethasone (adj. OR [95% CI]: 0.83 [0.71, 0.96]) and more likely to receive other steroids corticosteroids (adj. OR [95% CI]: 2.13 [1.90, 2.39]), relative to White patients. Hispanic/Latino patients were less likely to receive dexamethasone than Not Hispanic/Latino patients (adj. OR [95% CI]: 0.69 [0.58, 0.82]). Conclusions Our findings suggest that COVID-19 treatments patients received in Optum varied by race and ethnicity after adjustment for other possible explanatory factors. In the face of rapidly evolving treatment landscapes, policies are needed to ensure equitable access to novel and repurposed therapeutics to avoid disparities in care by race and ethnicity.