Assessing heterogeneity between studies is a critical step in determining whether studies can be combined and whether the synthesized results are reliable. The I-2 statistic has been a popular measure for quantifying heterogeneity, but its usage has been challenged from various perspectives in recent years. In particular, it should not be considered an absolute measure of heterogeneity, and it could be subject to large uncertainties. As such, when using I-2 to interpret the extent of heterogeneity, it is essential to account for its interval estimate. Various point and interval estimators exist for I-2. This article summarizes these estimators. In addition, we performed a simulation study under different scenarios to investigate preferable point and interval estimates of I-2. We found that the Sidik-Jonkman method gave precise point estimates for I-2 when the between-study variance was large, while in other cases, the DerSimonian-Laird method was suggested to estimate I-2. When the effect measure was the mean difference or the standardized mean difference, the Q-profile method, the Biggerstaff-Jackson method, or the Jackson method was suggested to calculate the interval estimate for I-2 due to reasonable interval length and more reliable coverage probabilities than various alternatives. For the same reason, the Kulinskaya-Dollinger method was recommended to calculate the interval estimate for I-2 when the effect measure was the log odds ratio.
Trial sequential analysis (TSA) is an increasingly used tool in systematic reviews to monitor synthesized evidence. However, the current practice of TSAs often overlooks the order of same-year studies, which are typically ordered alphabetically based on the last names of the studies’ authors by default in the widely used TSA software application. This practice is inappropriate and contrary to the TSA’s definition. This issue is particularly concerning in systematic reviews on time-sensitive topics, such as COVID-19, where reviews include many studies within a short period. In this article, we use a case study to illustrate the impact of the order of same-year studies on TSA conclusions. It shows dramatically different patterns of evidence accumulation when same-year studies are ordered alphabetically vs. in their actual temporal order. This article offers suggestions for authors to pay attention to study ordering in future TSAs.
Biomedical studies, such as clinical trials, often require the comparison of measurements from two correlated tests in which each unit of observation is associated with a binary outcome of interest via relative risk. The associated confidence interval is crucial because it provides an appreciation of the spectrum of possible values, allowing for a more robust interpretation of relative risk. Of the available confidence interval methods for relative risk, the asymptotic score interval is the most widely recommended for practical use. We propose a modified score interval for relative risk and we also extend an existing nonparametric U-statistic-based confidence interval to relative risk. In addition, we theoretically prove that the original asymptotic score interval is equivalent to the constrained maximum likelihood-based interval proposed by Nam and Blackwelder. Two clinically relevant oncology trials are used to demonstrate the real-world performance of our methods. The finite sample properties of the new approaches, the current standard of practice, and other alternatives are studied via extensive simulation studies. We show that, as the strength of correlation increases, when the sample size is not too large the new score-based intervals outperform the existing intervals in terms of coverage probability. Moreover, our results indicate that the new nonparametric interval provides the coverage that most consistently meets or exceeds the nominal coverage probability.
Systematic reviews and meta-analyses are principal tools to synthesize evidence from multiple independent sources in many research fields. The assessment of heterogeneity among collected studies is a critical step when performing a meta-analysis, given its influence on model selection and conclusions about treatment effects. A common-effect (CE) model is conventionally used when the studies are deemed homogeneous, while a random-effects (RE) model is used for heterogeneous studies. However, both models have limitations. For example, the CE model produces excessively conservative confidence intervals with low coverage probabilities when the collected studies have heterogeneous treatment effects. The RE model, on the other hand, assigns higher weights to small studies compared to the CE model. In the presence of small-study effects or publication bias, the over-weighted small studies from a RE model can lead to substantially biased overall treatment effect estimates. In addition, outlying studies may exaggerate between-study heterogeneity. This article introduces penalization methods as a compromise between the CE and RE models. The proposed methods are motivated by the penalized likelihood approach, which is widely used in the current literature to control model complexity and reduce variances of parameter estimates. We compare the existing and proposed methods with simulated data and several case studies to illustrate the benefits of the penalization methods.