Study objective: Randomized controlled trials govern evidence-based clinical practice, and it is therefore critical that their results be robust. We aim to investigate the fragility of randomized controlled trials in emergency medicine by determining how often significance would be nullified with small changes in outcomes using the fragility index. Methods: We conducted a methodological systematic review of randomized controlled trials in emergency medicine published in the top 10 general medicine journals and the top 10 emergency medicine journals. Inclusion criteria required that trials be emergency medicine studies structured with a 2-arm or 2-by-2 factorial design and report at least 1 statistically significant dichotomous outcome. Results: A total of 180 trials met inclusion criteria. The median fragility index across all trials in emergency medicine was 4 (interquartile range [IQR] 2 to 10) and the median sample size was 140 (IQR 69.5 to 286). For trials from general medicine journals (n = 32), the median fragility index was 9 (IQR 4 to 16.5) and the median sample size was 415.5 (IQR 219.5 to 901); for trials from emergency medicine journals (n = 148), the median fragility index was 4 (IQR 1 to 9) and the median sample size was 119 (IQR 60 to 227.25). One third of all trials (62/180) had a loss to follow-up that was greater than or equal to the fragility index. There was a modest correlation between fragility index and total number of events (r = 0.36; 95% confidence interval [CI] 0.23 to 0.48) and a weak correlation between fragility index and total sample size (r = 0.26; 95% CI 0.12 to 0.39). There was no correlation between fragility index and either P value (r = -0.14; 95% CI -0.28 to -0.006) or Science Citation Index (r = 0.07; 95% CI -0.08 to 0.22). Conclusion: The statistical significance of the results of randomized controlled trials in emergency medicine was often contingent on a small number of events. Until frequentist interpretation of clinical trials is replaced with Bayesian analysis, the fragility index may have utility as a tool to aid clinicians in assessing the robustness of randomized controlled trials in emergency medicine when considered in conjunction with the fragility quotient and other reported metrics.
STUDY OBJECTIVE:We aim to investigate spin in emergency medicine abstracts, using a sample of randomized controlled trials from high-impact-factor journals with statistically nonsignificant primary endpoints.METHODS:This study investigated spin in abstracts of emergency medicine randomized controlled trials from emergency medicine literature, with studies from 2013 to 2017 from the top 5 emergency medicine journals and general medical journals. Investigators screened records for inclusion and extracted data for spin. We considered evidence of spin if trial authors focused on statistically significant results, interpreted statistically nonsignificant results as equivalent or noninferior, used favorable rhetoric in the interpretation of nonsignificant results, or claimed benefit of an intervention despite statistically nonsignificant results.RESULTS:Of 772 abstracts screened, 114 randomized controlled trials reported statistically nonsignificant primary endpoints. Spin was found in 50 of 114 abstracts (44.3%). Industry-funded trials were more likely to have evidence of spin in the abstract (unadjusted odds ratio 3.4; 95% confidence interval 1.1 to 11.9). In the abstracts' results, evidence of spin was most often due to authors' emphasizing a statistically significant subgroup analysis (n=9). In the abstracts' conclusions, spin was most often due to authors' claiming they accomplished an objective that was not a prespecified endpoint (n=14).CONCLUSION:Spin was prevalent in the selected randomized controlled trial, emergency medicine abstracts. Authors most commonly incorporated spin into their reports by focusing on statistically significant results for secondary outcomes or subgroup analyses when the primary outcome was statistically nonsignificant. Spin was more common in studies that had some component of industry funding.
We commend the work by Brown et al1Brown J. Lane A. Cooper C. et al.The results of randomized controlled trials in emergency medicine are frequently fragile.Ann Emerg Med. 2019; 73: 565-576Google Scholar that, using the fragility index and fragility quotient, assesses the fragility of randomized controlled trials in the emergency medicine literature. Although they provide a good overview of the limitations of the 2 metrics, we would like to further question the utility of fragility measures. In their critique of fragility measures, Carter et al2Carter R.E. McKie P.M. Storlie C.B. The fragility index: a P-value in sheep’s clothing?.Eur Heart J. 2017; 38: 346-348Google Scholar suggested that fragility index is a “P-value in sheep’s clothing.” In their article, they demonstrated the strong relationship between P values, fragility index, and sample size. The results of their study were based on 60,000 simulated clinical trials using a combination of 10 sample sizes (100 to 1,000), with relative risks of 1.0, 0.67, and 0.33 for intervention versus control.2Carter R.E. McKie P.M. Storlie C.B. The fragility index: a P-value in sheep’s clothing?.Eur Heart J. 2017; 38: 346-348Google Scholar They showed that P values decrease as sample size increases when the effect size is nonzero; the inverse relationship is observed when fragility index and sample size are compared. Stated another way, if the effect size is held constant and the sample size increases, the fragility index likewise increases.2Carter R.E. McKie P.M. Storlie C.B. The fragility index: a P-value in sheep’s clothing?.Eur Heart J. 2017; 38: 346-348Google Scholar In their study, Brown et al noted that the overall sample size for included trials was a median of 140 (interquartile range 69.5 to 286), with the overall fragility index for primary outcomes of the 74 included studies equal to 5 (interquartile range 2 to 11.75) and a fragility quotient of 0.039 (interquartile range 0.015 to 0.081). The small median sample size in randomized controlled trials in this study could be one reason for the low fragility index, as mentioned above. Furthermore, as Carter et al2Carter R.E. McKie P.M. Storlie C.B. The fragility index: a P-value in sheep’s clothing?.Eur Heart J. 2017; 38: 346-348Google Scholar observed, randomized controlled trials “are designed to balance the sample size with expected efficacy. In doing so, the [fragility index] is also minimized and results will necessarily hinge on fewer events. This is unavoidable, particularly in the context of clinical equipoise and finite resources.” Although certainly fragility quotient (fragility index normalized to study size) addresses some of these concerns, relative measures of fragility have not proven to be a reliable indicator of study quality, as the authors suggested (for example, the Collaborative Study Group captopril study, which established the use of angiotensin-converting enzyme inhibitors for the prevention of worsening diabetic nephropathy).3Lewis E.J. Hunsicker L.G. Bain R.P. et al.The effect of angiotensin-converting-enzyme inhibition on diabetic nephropathy. Collaborative Study Group.N Engl J Med. 1993; 329: 1456-1462Google Scholar The primary endpoint of Collaborative Study Group was a doubling of the baseline serum creatinine concentration. Using the methods of Brown et al, we calculated the fragility index of Collaborative Study Group’s primary endpoint to be 4, with a fragility quotient of 0.009, or 1 event per 100 patients. Or consider the West of Scotland Coronary Prevention trial, the first major study to demonstrate the efficacy of statins in the primary prevention of nonfatal myocardial infarction or death from coronary artery disease in men with hypercholesterolemia.4Shepherd J. et al.Prevention of coronary heart disease with pravastatin in men with hypercholesterolemia.N Engl J Med. 1995; 333: 1301-1308Google Scholar The fragility index of this study was 34, with a fragility quotient of 0.005. The fragility index and fragility quotient of these studies suggest these trials are quite fragile, yet have proved clinically meaningful in daily practice. Despite the more recent use of the fragility index and fragility quotient in the literature, we recommend caution in relation to their use until research establishes their utility in clinical context. Until then, their clinical utility for randomized controlled trials remains uninterpretable. The Results of Randomized Controlled Trials in Emergency Medicine Are Frequently FragileAnnals of Emergency MedicineVol. 73Issue 6PreviewRandomized controlled trials govern evidence-based clinical practice, and it is therefore critical that their results be robust. We aim to investigate the fragility of randomized controlled trials in emergency medicine by determining how often significance would be nullified with small changes in outcomes using the fragility index. Full-Text PDF In reply:Annals of Emergency MedicineVol. 73Issue 6PreviewWe thank Niforatos et al for their contribution to the discussion in regard to the limitations of the fragility index and fragility quotient. In their letter, they expand on some of the limitations to the application of the fragility index and fragility quotient discussed in our original work and ultimately question their utility. Although they highlight some important considerations, we believe that the measures have clinical utility. Here, we explore further some key points about their application and provide insight into how clinicians might use these tools. Full-Text PDF
In budding yeast meiosis, homologous chromosomes become linked by chiasmata and then move back and forth on the spindle until they are bioriented, with the kineto-chores of the partners attached to microtubules from opposite spindle poles. Certain mutations in the conserved kinase, Mps1, result in catastrophic meiotic segregation errors but mild mitotic defects. We tested whether Dam1, a known substrate of Mps1, was necessary for its critical meiotic role. We found that kinetochore-microtubule attachments are established even when Dam1 is not phosphorylated by Mps1, but that Mps1 phosphorylation of Dam1 sustains those connections. But the meiotic defects when Dam1 is not phosphorylated are not nearly as catastrophic as when Mps1 is inactivated. The results demonstrate that one meiotic role of Mps1 is to stabilize connections that have been established between kinetochores and microtubles by phosphorylating Dam1.