Discrepancies may occur between the results of large randomized trials and the pooled results of several small trials in meta-analyses (14). Previous studies have suggested that discrepancies may be due to publication bias, that is, the fact that small trials are more likely to be published if they show a statistically significant intervention effect (58). Previous empirical studies of the association between methodologic quality and intervention effects have had inconsistent conclusions (912). In theory, adequate randomization requires adequate generation of the allocation sequence and adequate allocation concealment. The assumption is partly supported by studies from Schulz and colleagues (10) and Moher and associates (11, 12), who found that trials with inadequate allocation concealment exaggerate intervention effects significantly compared with trials reporting adequate allocation concealment. However, Emerson and coworkers (9) found no association between reported allocation concealment and intervention effects. Furthermore, none of the studies (912) found a significant association between generation of allocation sequence and intervention effects, although Schulz and colleagues found a nonsignificant trend (10). Schulz and colleagues (10), who analyzed trials in pregnancy and childbirth, found that trials without double blinding exaggerate intervention effects significantly compared with double-blind trials. However, Emerson and coworkers (9) and Moher and associates (11), who included trials from several therapeutic areas, found no association between double blinding and intervention effects. Methodologic quality can be assessed by using separate components, as in the studies discussed above, or by using one of several quality scales (13). One popular scale, developed by Jadad and colleagues in 1996 (14), has thus far received 233 registered citations (Institute for Scientific Information, Philadelphia, Pennsylvania). The scale includes assessment of the reported generation of allocation sequence, double blinding, and follow-up. Moher and associates (11, 12) found that trials with a low score on this scale exaggerate intervention effects significantly compared with trials that have high quality scores. However, the use of this and other quality scales has been disputed by Jni and coworkers (15), who showed that several quality scales produce inconsistent conclusions. We studied the potential association between reported methodologic quality and intervention effects to assess whether methodologic quality may explain discrepancies between the results of large and small randomized trials in meta-analyses. Methods Identification and Selection of Meta-Analyses and Trials According to suggestions in other studies (24), we arbitrarily defined trials with 1000 or more participants as large. We searched the Cochrane Library, MEDLINE on PubMed (using meta-analysis, review, randomi*ed, and controlled clinical trial as free text search words), and reference lists of relevant articles to identify potentially eligible meta-analyses that included at least one large trial. We identified 23 eligible meta-analyses. Nine were excluded because they included trials that were also included in larger eligible meta-analyses (n = 5), lacked references to the primary trials (n = 3), or excluded low-quality trials (n = 1). Accordingly, 14 meta-analyses (1626) were included. Three of the included Cochrane Reviews included two meta-analyses each (22, 25, 26). The meta-analyses included 248 trials, of which 58 were excluded because they were unpublished (n = 33), were quasi-randomized (n = 15), or were published as abstracts (n = 8). We were unable to translate 2 Spanish-language articles. The remaining 190 trials, published as English-language (n = 188) or German-language (n = 2) full-length articles, were included. Our analyses included 23 large and 167 small randomized trials and a total of 136 164 participants. On the basis of the study by Schulz and colleagues (10), who analyzed 250 controlled trials with 62 091 participants, we estimated that our sample size would be large enough to show significant differences between intervention benefits in high-quality and low-quality trials. Assessment of Methodologic Quality Methodologic quality was defined as the confidence that the trial's design, conduct, analysis, and presentation minimized or avoided biases in the trial's intervention comparisons (12). The reported methodologic quality was assessed in an unmasked manner by using four separate components and a composite quality scale. The four components were generation of the allocation sequence (adequate [computer-generated random numbers or similar] or inadequate [not described]), allocation concealment (adequate [central independent unit, sealed envelopes, or similar] or inadequate [not described or open table of random numbers or similar]), double blinding (adequate [identical placebo tablets or similar] or inadequate [not performed or tablets versus injections or similar]), and follow-up (adequate [number and reasons for dropouts and withdrawals described] or inadequate [number or reasons for dropouts and withdrawals not described]). The five-point quality scale included generation of the allocation sequence (2 points, computer-generated random numbers or similar; 1 point, not described; 0 points, quasi-randomized trial [which we excluded]), double blinding (2 points, identical placebo tablets or similar; 1 point, not described; 0 points, no blinding or inadequate method, such as tablets versus injections or similar), and follow-up (1 point, number and reasons for dropouts and withdrawals described; 0 points, number or reasons for dropouts and withdrawals not described). The quality score was ranked as low ( 2 points) or high ( 3 points), as suggested elsewhere (11). Two reviewers assessed the effect of masking and the interobserver reliability of the quality assessments. The reported methodologic quality of 30 trials was assessed with and without masking the names of the authors and journal, the year of publication, acknowledgments, institutional affiliations, and funding. The unmasked assessments were performed first. The masked assessments were performed 3 months later, ensuring that the assessors did not remember the trials. The difference between masked and unmasked quality assessments was not significant (mean [SE] quality score, 3.70 0.19 vs. 3.63 0.21; P > 0.2). Interobserver reliability was assessed by using 30 randomized trials randomly selected from the Cochrane Hepato-Biliary Group Controlled Trials Register and was found to be high (intraclass correlation coefficient, 0.96 [95% CI, 0.92 to 0.98]) (27). After assessing the quality of 100 trials, we reassessed the methodologic quality of the trials from the Cochrane Hepato-Biliary Group Controlled Trials Register and found that the testretest reliability was high (0.98 [CI, 0.97 to 0.99]) (27). Data Extraction Data were extracted independently by two reviewers. First, the primary binary outcome measure described by the largest number of trials in each meta-analysis was identified. We then extracted the number of events in the intervention and control groups and the number of participants randomly assigned to the intervention and control groups. All disagreements were due to inaccurate data extraction and were resolved through further consulting of the original articles and meta-analyses. Consensus was achieved before analyses were done in all cases. Statistical Analysis Analyses were performed by using SAS for Windows, version 6.12 (SAS Institute, Inc., Cary, North Carolina) or SPSS for Windows, version 10.0 (SPSS, Inc., Chicago, Illinois). Differences between the masked and unmasked quality assessments were estimated by using the KruskalWallis test. The number of participants and the year of publication in trials with adequate versus inadequate generation of allocation sequence, allocation concealment, double blinding, and follow-up and high versus low quality scores were compared by using the Pearson chi-square test. To estimate evidence of publication bias and other biases, we used linear regression to analyze funnel-plot asymmetry (7). The standard normal deviate, defined as the log odds ratio divided by its standard error, was regressed against the precision (the inverse of the standard error). If funnel-plot asymmetry is present, the regression line will not run through the origin, and the intercept will provide a measure of asymmetry. Intervention effects were estimated by using the number of events and participants in the treatment group and the number of events and participants in the control group (Appendix). Accordingly, two observations were needed per trial, one for the intervention group and one for the control group. When necessary, positive outcomes were re-expressed as unwanted end points, for example, mortality instead of survival. Discrepancies between intervention effects in large and small trials were estimated by the ratio of odds ratios (10), which is the summary odds ratio of large trials divided by the summary odds ratio of small trials. In this modeling convention, a ratio of odds ratios less than 1.0 indicates that a group of trials (for example, small trials with inadequate allocation concealment) exaggerates the intervention effect compared with the referent group (for example, large trials). Variance and confidence intervals were increased to adjust for overdispersion (Appendix). The Pearson chi-square test was used to estimate the potential overlap between quality components. Results Characteristics of Included Trials We were able to locate and retrieve all 190 included trials and to reproduce the results reported in the meta-analyses (1626) (Table 1). Of the 190 trials, 81 (43%) reported adequate generation of the allocation sequence, 68 (36%) reported adequate allocation concealment, 103 (54%) reported adequate double blinding, and 1
更多