Time-to-event outcomes are widely used in clinical and epidemiological research. For instance, studies of medical product safety often involve comparative analyses of rare time-to-event outcomes. The effects of misclassified outcomes and error in survival times for time-to-event data have not been widely investigated. In this Monte Carlo simulation study, we compared the relative bias of absolute and relative measures of effect under varying degrees of outcome misclassification, outcome incidences, direction of error in survival times, and the time point of inference. Relative measures of effect were susceptible to considerable downward bias, which was larger when the outcome incidence and specificity were lower, error in survival times led to earlier times, time point of inference was earlier, and the estimation excluded samples for which an estimate could not be obtained. For absolute measures of effect, the pattern of bias was much simpler, greater downward bias was primarily a function of the degree of sensitivity. The results suggest when the outcome incidence is rare, specificity and sensitivity are high, absolute measures of effect may be preferable to relative measures of effect.
PURPOSE:In the United States, the Food and Drug Administration has provided guidance on the contexts in which real-world evidence (RWE) can be used to support regulatory decisions for medical devices. METHODS:In this paper, we review different sources of real-world data (RWD) for medical devices used in surgical procedures and how to optimize the design and analysis features of RWD studies to maximize the strength of RWE. The design and analysis of a study are critical to ensuring that the RWE generated from the RWD can ultimately be used to support a regulatory decision or fulfill a commitment for a medical device. We consider several topics relevant to medical device RWE in our review: misclassification and measurement error, causal inference, and incomplete follow-up. RESULTS:Some key challenges include indication and outcome misclassification in non‑registry data, tenuous assumptions of single‑arm designs, device‑ and surgeon‑level confounding in comparative studies, and potentially informative loss to follow‑up from censoring (e.g., insurance termination). CONCLUSION:The review is intended as a guide to researchers and regulators interested in generating or evaluating medical device RWE for regulatory purposes.
When a drug or medical device is suspected of having a safety problem, observational studies are often utilized with an active comparator cohort design and covariate balancing using the propensity score. Each covariate balancing method is an estimator for a particular estimand, with each estimand characterizing the target population of interest differently as it relates to the treatment effect. In this article, I argue that characterizing the average treatment effect in the treated population (ATT), has a distinct inferential advantage over estimands characterizing the treatment effect in either the comparator population, entire population, and the overlap population, and as such the ATT should be prioritized. Regulatory guidance offers little direction with respect to estimand selection in observational studies, and a review of recent pharmacoepidemiology studies suggests that the ATT is infrequently used. Guidance is offered with respect to selecting among ATT estimators and identifying contexts where alternative estimands might be more informative. An empirical example is used to illustrate the implementation of the described methods. The implications of adopting the recommendations set forth in this article are considered.
Background:The use of surgical staplers for lung resection is widespread; however, determining the optimal surgical stapler remains a challenge. This study compared clinical and economic outcomes of the ECHELON™ 3000 Stapler (ECH3000) against its previous generation ECHELON™+ Stapler (ECH+) among patients undergoing lung resection. Methods:This was a retrospective, multi-institutional, comparative cohort study among patients undergoing lung resection using ECH3000 or ECH+ between April 1, 2022 and September 30, 2023 in the Premier Healthcare Database (PHD). The primary outcome was 30-day prolonged air leak (PAL). Secondary outcomes included 30-day bleeding-related complications and all-cause inpatient readmission, and index admission total hospital costs, length of stay (LOS), discharge status, and mortality. Covariate balancing propensity score (CBPS) weighting was used to balance baseline characteristics between groups. The difference in the cumulative incidence (binary outcomes) and mean (continuous outcomes) of study outcomes was calculated using covariate-balanced data. For the primary outcome, the non-inferiority in the difference in the cumulative incidence of 30-day PAL between groups was evaluated using a prespecified non-inferiority margin of 10%. Results:A total of 910 (ECH3000: 277; ECH+: 633) patients met the study eligibility criteria. The cumulative incidences of 30-day PAL were comparable between ECH3000 vs. ECH+ {13.7% vs. 13.0%; difference =0.7% [two-sided 95% confidence interval (CI): -4.6%, 5.9%]}. Since the upper bound of the 95% CI for the risk difference in 30-day PAL between groups (5.9%) was significantly lower than the prespecified non-inferiority margin of 10% (P value<0.001), the primary outcome was met. A lower cumulative incidence of 30-day bleeding-related complications was observed among patients using ECH3000 vs. ECH+ [9.4% vs. 16.4%; difference =-7.0% (95% CI: -12.2%, -1.8%)]. There were no significant differences between groups for other secondary outcomes. Conclusions:Among patients undergoing lung resection, ECH3000 had a comparable cumulative incidence of 30-day PAL and a lower cumulative incidence of bleeding-related complications compared to ECH+.
Observational studies are frequently used in clinical research to estimate the effects of treatments or exposures on outcomes. To reduce the effects of confounding when estimating treatment effects, covariate balancing methods are frequently implemented. This study evaluated, using extensive Monte Carlo simulation, several methods of covariate balancing, and two methods for propensity score estimation, for estimating the average treatment effect on the treated using a hazard ratio from a Cox proportional hazards model. With respect to minimizing bias and maximizing accuracy (as measured by the mean square error) of the treatment effect, the average treatment effect on the treated weighting, fine stratification, and optimal full matching with a conventional logistic regression model for the propensity score performed best across all simulated conditions. Other methods performed well in specific circumstances, such as pair matching when sample sizes were large (n = 5000) and the proportion treated was < 0.25. Statistical power was generally higher for weighting methods than matching methods, and Type I error rates were at or below the nominal level for balancing methods with unbiased treatment effect estimates. There was also a decreasing effective sample size with an increasing number of strata, therefore for stratification-based weighting methods, it may be important to consider fewer strata. Generally, we recommend methods that performed well in our simulations, although the identification of methods that performed well is necessarily limited by the specific features of our simulation. The methods are illustrated using a real-world example comparing beta blockers and angiotensin-converting enzyme inhibitors among hypertensive patients at risk for incident stroke.
Abstract Background Autoimmune disorders have primary manifestations such as joint pain and bowel inflammation but can also have secondary manifestations such as non-infectious uveitis (NIU). A regulatory health authority raised concerns after receiving spontaneous reports for NIU following exposure to Remicade®, a biologic therapy with multiple indications for which alternative therapies are available. In assessment of this clinical question, we applied validity diagnostics to support observational data causal inferences. Methods We assessed the risk of NIU among patients exposed to Remicade® compared to alternative biologics. Five databases, four study populations, and four analysis methodologies were used to estimate 80 potential treatment effects, with 20 pre-specified as primary. The study populations included inflammatory bowel conditions Crohn’s disease or ulcerative colitis (IBD), ankylosing spondylitis (AS), psoriatic conditions plaque psoriasis or psoriatic arthritis (PsO/PsA), and rheumatoid arthritis (RA). We conducted four analysis strategies intended to address limitations of causal estimation using observational data and applied four diagnostics with pre-specified quantitative rules to evaluate threats to validity from observed and unobserved confounding. We also qualitatively assessed post-propensity score matching representativeness, and bias susceptibility from outcome misclassification. We fit Cox proportional-hazards models, conditioned on propensity score-matched sets, to estimate the on-treatment risk of NIU among Remicade® initiators versus alternatives. Estimates from analyses that passed four validity tests were assessed. Results Of the 80 total analyses and the 20 analyses pre-specified as primary, 24% and 20% passed diagnostics, respectively. Among patients with IBD, we observed no evidence of increased risk for NIU relative to other similarly indicated biologics (pooled hazard ratio [HR] 0.75, 95% confidence interval [CI] 0.38–1.40). For patients with RA, we observed no increased risk relative to similarly indicated biologics, although results were imprecise (HR: 1.23, 95% CI 0.14–10.47). Conclusions We applied validity diagnostics on a heterogenous, observational setting to answer a specific research question. The results indicated that safety effect estimates from many analyses would be inappropriate to interpret as causal, given the data available and methods employed. Validity diagnostics should always be used to determine if the design and analysis are of sufficient quality to support causal inferences. The clinical implications of our findings on IBD suggests that, if an increased risk exists, it is unlikely to be greater than 40% given the 1.40 upper bound of the pooled HR confidence interval.
The ThermoCool STSF catheter is used for ablation of ischemic ventricular tachycardia (VT) in routine clinical practice, although outcomes have not been studied and the catheter does not have Food and Drug Administration (FDA) approval for this indication. We used real-world health system data to evaluate its safety and effectiveness for this indication. Among patients undergoing ischemic VT ablation with the ThermoCool STSF catheter pooled across two health systems (Mercy Health and Mayo Clinic), the primary safety composite outcome of death, thromboembolic events, and procedural complications within 7 days was compared to a performance goal of 15%, which is twice the expected proportion of the primary composite safety outcome based on prior studies. The exploratory effectiveness outcome of rehospitalization for VT or heart failure or repeat VT ablation at up to 1 year was averaged across health systems among patients treated with the ThermoCool STSF vs. ST catheters. Seventy total patients received ablation for ischemic VT using the ThermoCool STSF catheter. The primary safety composite outcome occurred in 3/70 (4.3%; 90% CI, 1.2–10.7%) patients, meeting the pre-specified performance goal, p = 0.0045. At 1 year, the effectiveness outcome risk difference (STSF-ST) at Mercy was − 0.4% (90% CI: − 25.2%, 24.3%) and at Mayo Clinic was 12.6% (90% CI: − 13.0%, 38.4%); the average risk difference across both institutions was 5.8% (90% CI: − 12.0, 23.7). The ThermoCool STSF catheter was safe and appeared effective for ischemic VT ablation, supporting continued use of the catheter and informing possible FDA label expansion. Health system data hold promise for real-world safety and effectiveness evaluation of cardiovascular devices.
Observational studies are increasingly being used in medicine to estimate the effects of treatments or exposures on outcomes. To minimize the potential for confounding when estimating treatment effects, propensity score methods are frequently implemented. Often outcomes are the time to event. While it is common to report the treatment effect as a relative effect, such as the hazard ratio, reporting the effect using an absolute measure of effect is also important. One commonly used absolute measure of effect is the risk difference or difference in probability of the occurrence of an event within a specified duration of follow-up between a treatment and comparison group. We first describe methods for point and variance estimation of the risk difference when using weighting or matching based on the propensity score when outcomes are time-to-event. Next, we conducted Monte Carlo simulations to compare the relative performance of these methods with respect to bias of the point estimate, accuracy of variance estimates, and coverage of estimated confidence intervals. The results of the simulation generally support the use of weighting methods (untrimmed ATT weights and IPTW) or caliper matching when the prevalence of treatment is low for point estimation. For standard error estimation the simulation results support the use of weighted robust standard errors, bootstrap methods, or matching with a naïve standard error (i.e., Greenwood method). The methods considered in the article are illustrated using a real-world example in which we estimate the effect of discharge prescribing of statins on patients hospitalized for acute myocardial infarction.
Left atrium (LA) geometric changes with advancing chronicity and severity of atrial fibrillation (AF). During catheter ablation for AF, electroanatomical mapping (e.g. with CARTO) of the LA can be used to measure the LA volume and atrial width to heigh ratio to guide ablation procedures.
Importance:The ThermoCool SmartTouch catheter (ablation catheter with contact force and 6-hole irrigation [CF-I6]) is approved by the US Food and Drug Administration (FDA) for paroxysmal atrial fibrillation (AF) ablation and used in routine clinical practice for persistent AF ablation, although clinical outcomes for this indication are unknown. There is a need to understand whether data from routine clinical practice can be used to conduct regulatory-grade evaluations and support label expansions. Objective:To use health system data to compare the safety and effectiveness of the CF-I6 catheter for persistent AF ablation with the ThermoCool SmartTouch SurroundFlow catheter (ablation catheter with contact force and 56-hole irrigation [CF-I56]), which is approved by the FDA for this indication. Design, Setting, and Participants:This retrospective, comparative-effectiveness cohort study included patients undergoing catheter ablation for persistent AF at Mercy Health or Mayo Clinic from January 1, 2014, to April 30, 2021, with up to a 1-year follow-up using electronic health record data. Exposures:Use of the CF-I6 or CF-I56 catheter. Main Outcomes and Measures:The primary safety outcome was a composite of death, thromboembolic events, and procedural complications within 7 to 90 days. The exploratory effectiveness outcome was a composite of AF-related hospitalization events after a 90-day blanking period. Propensity score weighting was used to balance baseline covariates. Risk differences were estimated between catheter groups and averaged across the 2 health care systems, testing for noninferiority of the CF-I6 vs the CF-I56 catheter with respect to the safety outcome using 2-sided 90% CIs. Results:Overall, 1450 patients (1034 [71.3%] male; 1397 [96.3%] White) underwent catheter ablation for persistent AF, including 949 at Mercy Health (186 CF-I6 and 763 CF-I56; mean [SD] age, 64.9 [9.2] years) and 501 at Mayo Clinic (337 CF-I6 and 164 CF-I56; mean [SD] age, 63.7 [9.5] years). A total of 798 (55.0%) had been treated with class I or III antiarrhythmic drugs before ablation. The safety outcome (CF-I6 - CF-I56) was similar at both Mercy Health (1.3%; 90% CI, -2.1% to 4.6%) and Mayo Clinic (-3.8%; 90% CI, -11.4% to 3.7%); the mean difference was noninferior, with a mean of 0.5% (90% CI, -2.6% to 3.5%; P < .001). The effectiveness was similar at 12 months between the 2 catheter groups (mean risk difference, -1.8%; 90% CI, -7.3% to 3.7%). Conclusions and Relevance:In this cohort study, the CF-I6 catheter met the prespecified noninferiority safety criterion for persistent AF ablation compared with the CF-I56 catheter, and effectiveness was similar. This study demonstrates the ability of electronic health care system data to enable safety and effectiveness evaluations of medical devices.
Greater impedance drop during radiofrequency catheter ablation (RFA) reflects greater catheter contact with atrial tissue, which contributes to durable lesion formation. The role of impedance drop in influencing ablation success is not fully understood.
Traditional approaches to hypothesis testing in comparative post-approval safety and effectiveness studies of medical products are often inadequate because of a limited scope of possible inferences (e.g., superiority or inferiority). Often there is interest in simultaneously testing for superiority, equivalence, inferiority, non-inferiority, and non-superiority, which can be achieved using a partition testing framework. Partition testing only requires selection of an equivalence margin and calculation of a two-sided Wald confidence interval. In addition to permitting a broader range of inferences, the strengths of the approach include: mitigating publication bias, avoiding use of a clinically irrelevant nil hypothesis, and more transparent and impartial appraisal of the clinical importance of a study's findings by pre-specifying an equivalence margin. However, a challenge in implementing the approach can be the process for identifying an equivalence margin. The methodology is illustrated using a published study of the safety of Ondansetron for the off-label treatment of nausea and vomiting during pregnancy. Applying the method to the study results would have led to a conclusion that women exposed to Ondansetron in comparison to those that are not, are equivalent with respect to risk of cardiac malformations and oral clefts. These conclusions are more in line with the magnitude of the observed effects than the conclusions resulting from a traditional inferiority/superiority testing conducted by the study authors.
Prospectively designed studies that make use of real-world data (RWD) are increasingly being used for studies with regulatory implications, including both premarket and post-market studies, aided b...
Purpose Linear surgical staplers reduce rates of surgical adverse events (bleeding, leaks, infections) compared to manual sutures thereby reducing patient risks, surgeon workflow disruption, and healthcare costs. However, further improvements are needed. Ethicon Gripping Surface Technology (GST) reloads, tested and approved by regulatory authorities in combination with powered staplers, may reduce surgical risks through improved tissue grip. While manual staplers are used in some regions due to affordability, clinical data on GST reloads used with manual staplers are unavailable. This study compared surgical adverse event rates of manual staplers with GST vs standard reloads. These data may be used for label changes in China and Latin America. Patients and Methods Patients undergoing general or thoracic surgery between October 1, 2015 and August 31, 2021 using ECHELON FLEX™ manual staplers with GST or standard reloads were identified from the Premier Healthcare Database. GST reloads were compared to standard reloads for non-inferiority in bleeding and anastomotic leak for general surgery. Secondary outcomes included sepsis for general surgery, and bleeding and prolonged air leak for thoracic surgery. Covariate balancing was performed using stable balancing weights. Results The general and thoracic surgery cohorts contained 4571 (GST: 2780; standard: 1791) and 814 (GST: 514; standard: 300) patients, respectively. GST reloads were non-inferior to standard reloads for bleeding and anastomotic leak (adjusted cumulative incidence ratio: 1.02 [90% CI: 0.71, 1.45] and 1.03 [90% CI: 0.72, 1.46], respectively) for general surgery. Compared with standard reloads, GST reloads had a similar incidence of sepsis (2.2% vs 2.1%) for general surgery and lower incidences of bleeding (9.5% vs 16.0%) and prolonged air leak (12.6% vs 14.0%,) for thoracic surgery. Conclusion GST reloads, compared to standard reloads, used with ECHELON FLEX™ manual staplers had comparable perioperative bleeding and anastomotic leak for general surgery, and lower incidences of safety events for thoracic surgery.
In an observational study, to investigate the treatment effect, one common strategy is to match the control subjects to the treated subjects. The outcomes between the two groups are then compared after the TC (treatment-control) match. However, when the outcome is rare, detection of an outcome difference can be challenging. An alternative approach is to compare the treatment or exposure discrepancy after matching subjects with the outcome (cases) to subjects without the outcome (referents). Throughout the article, we follow the tradition to call this the matched "case-control" approach instead of the matched "case-referent" approach. We reserve "control" to mean not taking the treatment, and the abbreviation TC and CC (case-control) when possible confusion may arise. We derive conditions when the matched CC approach has more power for testing the treatment effect and examine its empirical performance in simulations and in our data example. We also show that the CC approach gives better match quality in our study of the effect of long vs. short stay in the hospital after joint surgery.