
Abstract There are various statistical models with different features and assumptions in meta-analysis. Popular models include the fixed-effect (or common effect), random-effects, and mixed-effects models. Apart from these models, several alternative models, such as the multiplicative error model (also known as the unrestricted weighted least squares model), the hybrid of additive and multiplicative error models, and the location-scale model, have been proposed in the literature. Understanding these models can be challenging for researchers without a solid mathematical background. Implementing or modifying these models is even more challenging for researchers unless they have advanced statistical and programming knowledge. This tutorial elucidates a structural equation modeling (SEM) framework to understand these models. Several R packages are introduced to facilitate the specification and generation of graphical models for these meta-analytic models. Researchers may fit these models with the full information maximum likelihood estimation method. This holds significant potential across two domains. First, it can be used as an educational tool in teaching and learning meta-analytic models. Second, the framework supports the development and evaluation of novel meta-analytical models, which have not yet been implemented in current meta-analysis software, thereby fostering advancements in meta-analytic methods. This tutorial also demonstrates, using two real datasets, how to extend it to address research questions involving complex meta-analytic data. Limitations and future directions to extend the SEM-based meta-analysis are discussed.
Abstract Meta-analysis is a widely used statistical tool for estimating the diagnostic accuracy of tests across multiple studies. Existing methods and available R packages primarily focus on a single diagnostic test, typically under the assumption that all studies include a gold standard. Greater efficiency can be achieved by modeling multiple diagnostic tests together and drawing on studies with or without a gold standard reference test across diverse designs. To address this challenge, recent work has extended both the Bayesian hierarchical model and the Bayesian hierarchical summary receiver operating characteristic model to the framework of network meta-analysis of diagnostic tests, enabling simultaneous comparison of multiple tests when some data are missing. Despite the importance of these methods, their computational complexity has limited their broad application. This article introduces NMADTA , an R package that implements these models with user-friendly functions. The package allows researchers to evaluate the accuracy of multiple diagnostic tests simultaneously and provides comprehensive graphical displays of the results.
Individual patient data meta-analyses (IPDMAs) provide powerful tools for synthesizing evidence across studies, yet methods for addressing unmeasured confounding in observational IPDMAs with survival outcomes are rarely implemented. Instrumental variable (IV) approaches offer causal inference capabilities but face practical challenges in hierarchical data structures, particularly the lack of standard diagnostics for instrument strength in nonlinear mixed-effects models. We adapt and evaluate a frequentist mixed-effects two-stage residual inclusion (2SRI) framework for survival IPDMAs, extending traditional IV methods to accommodate study-level and temporal clustering while handling time-to-event outcomes through Cox proportional hazards models. Because classical F-statistics are unavailable for logistic mixed-effects first-stage models, we propose the Wald chi 2 $\chi <^>2$ chi squared statistic as a practical instrument-strength diagnostic and empirically characterize its relationship to estimator performance. Through a comprehensive simulation study with 48 scenarios-varying unmeasured confounding (weak to very strong), instrument-treatment association strength (0.3-1.0), and cross-study IV allocation patterns-we evaluated 2SRI against naive mixed-effects Cox models using bias, coverage, variance, and mean squared error. The design was anchored to realistic IPDMA structure (10 studies, N approximate to 4 , 357 $N \approx 4,357$ upper N almost equals 4 comma 357 ) from pooled Ebola data, with 1,000 replications per scenario. Results show that under weak confounding, naive models dominate on all metrics. With moderate-to-strong confounding and realized Wald chi 2 $\chi <^>2$ chi squared exceeding 150-200, mixed-effects 2SRI substantially reduces bias and achieves near-nominal coverage, though with inflated variance. We provide empirical guideposts linking realized first-stage strength to expected performance, enabling analysts to judge when 2SRI will outperform conventional approaches in hierarchical survival IPDMAs. All simulations assume a common treatment effect across studies. Performance under heterogeneous effects remains to be established.
Flexible meta-regression is an important facet of biomedical research. These methods employ basis expansions to model the relationship between exposure and outcome. Examples include splines, fractional polynomial, and local regression models. The extent to which flexible meta-regression has been appropriately employed and reported in the past has not been studied. This study aims to provide an overview of flexible meta-regression for researchers as well as summarize the practical use and reporting of these methods in the epidemiological literature. A systematic literature search of EMBASE, MEDLINE, and the Cochrane Library identified systematic reviews published between 2011 and 2021 which employed flexible meta-regression methods. Data on study characteristics were extracted and validated by two independent researchers. Given the lack of guidance on reporting standards for flexible meta-regression, we proposed a novel 9-item checklist. This was used to assess reporting quality in included studies. A total of $N=346$ eligible studies were identified and $N=86$ were randomly sampled for full data extraction and analysis. The number of reviews using flexible regression methods increased over time. Spline models were the most frequently employed class of methods. There were no apparent trends in the use of the three methods over country of affiliation, number of authors, or clinical context. In over two thirds of reviews sampled, methods were found to have not been reported consistently. Further work is needed to improve the reporting of these methods to ensure research transparency and reproducibility.
This study investigated information retrieval of preprint records in the context of evidence synthesis work and compared 12 sources used to discover preprints. Identification of grey literature is often required or recommended in evidence synthesis guidance, and preprints are categorized as grey literature. The purpose of this work is to inform how and where to search for preprints to maximize coverage (through exploration of preprint server across aggregators and databases) while balancing search efficiency. Authors selected aggregators and databases hosting two or more preprint servers and then tested search functionality and extracted characteristics and features. Authors analyzed and compared the selected sources, tabulated the number of essential features, and created comparison tables reflecting database and aggregator features. The study protocol was registered in Open Science Framework registries. Preprint aggregators and databases differ in their content coverage, and their ability to design a comprehensive and reproducible search strategy. Limitations such as character or word limits for queries, limited advanced search operators, and missing export functionality affect the usability of aggregators for evidence synthesis searches. Ongoing updates to search interfaces and functionality and differing approaches to versioning make it challenging to study discovery of preprints across sources. The recommendations and scenarios in this article will assist searchers engaged in evidence synthesis to make informed decisions about where to search for preprints.
A key task in conducting systematic reviews is deduplicating the results from database searching. Deduplication using reference management software can be time-consuming and prone to error, while automated tools can be expensive and lack transparency. To support review teams, we evaluated eight deduplication tools: (1) The Automated Systematic Search Deduplicator (ASySD); (2) Covidence; (3) Deduklick; (4) EPPI-Reviewer; (5) PICO Portal; (6) Rayyan; (7) The Systematic Review Accelerator (SRA) Deduplicator: Focused; (8) The SRA Deduplicator: Relaxed. Five randomly selected Cochrane reviews had their searches rerun to create five gold standard sets. We compared the gold standard sets to the outputs of the eight deduplication tools and evaluated the results for: (1) unique records removed; (2) duplicate records retained; (3) time taken to deduplicate. Summed across all five reviews, the unique records removed in error ranged from 2 to 22. The three best tools were: (1) Rayyan; (2) Covidence; (3) SRA Deduplicator: Focused. The duplicate records retained in error ranged from 34 to 280, the three best tools were: (1) ASySD; (2) Rayyan; (3) EPPI-Reviewer. The time taken to deduplicate ranged from one minute to 20 hours and 34 minutes, the three fastest tools were: (1) SRA Deduplicator: Relaxed; (2) Deduklick; (3) Covidence. No tool performed so poorly that we don't recommend using it. But, as all the tools had strengths and weaknesses, some are expensive while others require large amounts of manual checking time, we recommend review teams compare the tools across all three outcomes and choose the tool that best suits their needs.
Scoping reviews are increasingly used for evidence synthesis, yet a comprehensive overview of benchmarks for their conduct is lacking. This study aimed to extract and analyze key metrics of preregistered scoping reviews to aid researchers in planning, resource allocation, methodological rigor, and innovation. We examined 2,038 scoping review protocol registrations on the Open Science Framework between April and August 2024, identifying 891 corresponding publications. Extracted variables included review type (e.g., rapid scoping review), information about the number of studies in the screening process, methodological practices, and funding. Using Stata, we calculated process metrics such as review duration, yield rate, and team composition. Qualitative analysis was applied to further metrics, including aims and eligibility criteria frameworks. Among the 891 publications, 91.47% were published as scoping reviews. On average, reviews involved 6.11 authors and took 80.41 weeks from registration to publication, with rapid scoping reviews completed in 55.63 weeks. Most publications (80.13%) originated from high-income countries, where funding was more common. An average of 5.19 databases were searched, identifying a mean of 6,920.66 potential inclusions, with 92.64 studies ultimately included (yield rate: 5.68%). Endnote was the most frequently employed screening tool. Language restrictions were reported in 37.60% of cases, and 75.42% did not conduct or report critical appraisal. This study provides valuable insights into common practices within scoping review processes, offering researchers a foundation to allocate resources more efficiently and enhance research quality through methodological benchmarks. It underscores the need for greater transparency and improved methodological rigor.
Knowledge synthesis involves bringing together findings from individual research studies to answer a specific question or explore a particular topic, often at a global level. People with lived or living experience of health conditions or of accessing healthcare services (PWLE), such as patients and their families or friends, can provide valuable perspectives throughout this process. PWLE involvement can help shape the design and conduct of a knowledge synthesis to better reflect real-world concerns through a process called patient engagement. Patient engagement is the meaningful and active involvement of PWLE throughout the research process. While previous research shows that PWLE can contribute meaningfully when appropriate supports are in place, there is limited practical guidance on how to enact patient engagement in knowledge syntheses. This tutorial paper aims to guide researchers and PWLE who would like to engage in knowledge syntheses together. Building on existing research, our team (comprised of PWLE, researchers, and an academic librarian) shares lessons we have learned from working together throughout the stages of a knowledge synthesis. We present considerations for successful patient engagement when research teams are: (1) planning to engage, (2) recruiting PWLE, (3) establishing roles and rapport, (4) supporting capacity development, (5) conducting the knowledge synthesis, and (6) mobilizing findings. We also identify 12 common barriers to patient engagement and offer strategies to address them, alongside recommendations for integrating PWLE across every stage of the knowledge synthesis process.
Under the network meta-analysis (NMA) framework for aggregate data, there are limited possibilities for evidence synthesis across multiple treatment doses or timepoints. Model-based network meta-analysis (MBNMA) has been recommended as a framework for either evidence synthesis across multiple dose levels or across multiple timepoints to circumvent the limitations of the NMA. A joint dose-response and time-course MBNMA (DT-MBNMA) is proposed that combines the strengths of both DT-MBNMA. This framework allows for combining data at multiple timepoints from studies in early clinical development with a broad range of doses to late-stage clinical studies with a limited range of doses. The method respects randomization and allows for assessment of consistency and hence satisfies the requirements from reimbursement agencies. The method was validated in a simulation study, showing that the drug effect parameters and therefore indirect treatment effects could be recovered without bias, while the precision of the treatment effects was dependent on the simulated network. Compared with a standard NMA, the methodology increased the statistical efficiency of the indirect treatment comparison (ITC). The use of the method was further illustrated on a dataset consisting of seven randomized clinical trials (RCTs) (26 treatment arms) in the treatment of obesity with Glucagon-like peptide-1 receptor agonists (GLP-1 RAs), with a broad range of doses and follow-up times. The method integrated phase 2 and 3 data seamlessly into the meta-analysis and provided greater precision on the treatment effects compared to NMA. Finally, the statistical framework may be used to support clinical decision-making, providing a framework for ITC during drug development.
Data extraction in systematic reviews, maps, and meta-analyses is time-consuming and prone to human error or subjective judgment. Large Language Models offer the potential for saving time, yet their performance has been evaluated in a limited range of platforms, disciplines, and review types. We assessed the performance of the Elicit platform across diverse data extraction tasks using journal articles from seven systematic reviews in life and environmental sciences. Human-extracted data served as the gold standard. For each review, we used eight articles for prompt development and another eight for testing. Initial prompts were iteratively refined to exceed 87% accuracy or up to five rounds. We then tested extraction accuracy, reproducibility across user accounts, and the effect of Elicit's high-accuracy mode. Of 90 considered prompts, 70 exceeded the 87% accuracy when compared to gold standard, but tended to be lower when tested on a new set of articles. Repeating data extractions with different Elicit user accounts resulted in 90% agreement on extracted values, though supporting quotes and reasoning matched in only 46% and 30% of cases, respectively. In high-accuracy mode, value matches dropped to 77%, with just 10% quote matches and 0% reasoning matches. Extraction accuracy did not differ by data types. Elicit also helped identify eight (<1%) errors in the gold standard data. Our results show that Elicit can complement, but not replace, human data extractors. Elicit may be best used for sanity checks and to evaluate the clarity of data extraction protocols. Prompts must be fine-tuned and independently validated.
In meta-analyses, effect size measures with bounded or non-normal sampling distributions are commonly analyzed on a transformed scale to justify normality assumptions. While point estimates and confidence intervals (CIs) are routinely back-transformed to the original scale for interpretation, this practice is nontrivial in random-effects models. In particular, standard inverse back-transformations yield estimates of the median rather than the mean effect size due to Jensen's inequality. Integral back-transformations provide a principled solution for recovering the mean on the original scale, but their use entails practical issues. We study integral back-transformations for several effect size measures, including correlation coefficients, proportions, odds and risk ratios, and Cronbach's alpha. We derive general formulations for integral back-transformations and corresponding CIs that are applicable across different transformation functions and provide a software implementation. Although required to obtain correct mean estimates, these approaches must be used with caution, as they are sensitive to heterogeneity estimation and can be unstable for unbounded transformations. Certain asymmetric transformations can also lead to inconsistent inference results. We illustrate these issues using analytical considerations and re-analyses of several meta-analyses. Importantly, we stress that the choice of back-transformation depends on the analyst's goals. The standard inverse back-transformation remains well suited for descriptive purposes and is often preferable in practice, provided that it is correctly interpreted as a back-transformed median effect size. We recommend using the integral back-transformation only when a mean estimate is explicitly required, and restricting its use for CIs to cases where they are demanded, but not for hypothesis testing.
Categorical moderators are often found in meta-analysis and examined using meta-regression models. When multiple effect sizes are present within studies, several methods can be used for meta-regression: multivariate models, three-level models, correlated-effects models with robust variance estimation (RVE), three-level models with RVE, and correlated-effects models with RVE and cluster wild bootstrapping (CWB). This study aimed to compare the performance of these methods through a simulation study. Cohen's d values were generated under a multivariate model, incorporating a binary variable that could represent either study-level or effect size-level characteristics. When the moderator referred to an effect size-level characteristic, its effect was allowed to vary across studies. Factors manipulated in the simulation included number of studies, number of outcomes per study, and the distribution of effect sizes across the categories of the moderator variable, ranging from balanced to highly unbalanced. The methods were applied and compared in terms of bias, Type I error, and power. The results showed that all methods exhibited lower power to detect effects when the moderator variable referred to study-level characteristics and the effect size distribution was very unbalanced. Methods based on RVE (correlated-effects with RVE or with RVE and CWB, and three-level models with RVE) effectively controlled Type I error rates but tended to be overconservative. In contrast, three-level models achieved higher power but at the cost of inflated Type I error. The best balance between Type I error control and power was observed when using a combination of three-level models and RVE.
To explore the perceptions of, and barriers to, grey literature searching among medical researchers and journal editors. A cross-sectional survey of authors of systematic reviews and the editors of the journals in which the reviews were published. Systematic reviews indexed in MEDLINE, spanning a 4-week period in 2019. We excluded protocols. We asked whether the reviewers were performing a grey literature search. If they were, we asked about their approach to the grey literature search and relevant guidance. If they were not, we asked about their rationale for this. We elucidated understandings of grey literature from all reviewers. The survey to journal editors asked about their perceptions towards grey literature. A consecutive sample of 1,229 systematic reviews was included. A total of 155 authors responded, and a total of 46 journal editors responded. The majority (57%) of reviewers reported performing a grey literature search. However, there was no consensus on types of grey literature items or sources among the reviewers. The most frequent barrier to grey literature searches was concern about the quality/detail of the grey literature items. Editors expressed negative perceptions towards grey literature, rooted in suspicion of its quality. While many reviewers reported performing a grey literature search, there was a diverse understanding of the term grey literature. Concerns about the quality of grey literature exist among reviewers and editors alike. There is a marked discrepancy between the best-practice guidelines and the gatekeepers of journals.
To examine the extent to which information sources other than journal articles are sought for systematic reviews. Cross-sectional study of published systematic reviews. We examined all published systematic reviews included in MEDLINE in a 4-week period in 2019. Both systematic reviews and protocols of reviews were eligible for inclusion. (1) Number and types of information sources sought in systematic reviews; (2) proportion of reviews that explicitly searched for study reports other than journal articles; (3) proportion of reviews that searched resources containing study reports other than journal articles. A total of 1,262 systematic reviews fulfilled the eligibility criteria. The median number of information resources searched for all systematic reviews was 4. Of the 1,262 reviews, study reports other than journal articles were sought in 40% (n = 502) of systematic reviews (97% (n = 64) of Cochrane reviews and 37% (n = 438) of non-Cochrane reviews). Trial registers were searched in 88% of Cochrane reviews and 21% of non-Cochrane reviews. In 99.3% (n = 1,253) of all the systematic reviews, the searches performed had the potential to identify study reports other than journal articles. Between a third and a half of systematic reviews search for study reports other than journal articles. Systematic review searches often search resources that include study reports other than journal articles, whether or not the reviewers explicitly sought them.
Systematic reviews (SRs) are critical for evidence-based research but are time-consuming and labor-intensive. The rapid expansion of academic publications further challenges the performance and applicability of existing screening and classification methods. While large language models (LLMs) present new opportunities for automation, limited research has examined whether they can achieve classification performance comparable to human reviewers in large-scale, multi-class settings. With the goal of improving classification performance, we proposed an LLM-based framework that leverages full-text key-insight extraction to enhance literature classification. We constructed a manually curated dataset of 900 articles from 17 published SRs to quantitatively evaluate the classification capabilities of LLMs. The results provided empirical evidence of LLMs' potential in supporting large-scale SRs and introduced a practical pathway for improving efficiency and reliability in evidence synthesis. Empirical results showed that key-insight-based classification (KBC) significantly outperforms abstract-based classification (ABC). We implemented a confidence-weighted voting (CWV) mechanism using multiple LLMs to improve robustness. The CWV method achieved the highest macro F1-score of 0.796, substantially exceeding KBC (0.732), ABC (0.676), and unsupervised K-means clustering (0.446). By employing zero-shot LLMs, our approach demonstrated the potential for enhanced adaptability across diverse domains and classification tasks without requiring fine-tuning, demonstrating that a carefully designed pipeline can enable LLMs to achieve classification performance comparable to human reviewers.
Effective data management is essential for tasks involving decisions based on data, including knowledge synthesis and literature reviews. Despite this, how to carry out data management in literature reviews effectively remains unclear. With the increasing volume of research papers and the expansion of computational techniques for processing data (e.g., machine learning or large language models), it becomes imperative to consider data management as a crucial element for the advancement of literature review practices and tools. Presently, there are shortcomings related to (1) handling the growth of research to be synthesized, (2) addressing data quality issues when applying computational techniques or facilitating the verification of content produced by generative artificial intelligence, (3) enabling efficient reuse of datasets and innovative recombination of tools, and (4) facilitating transparent collaboration across heterogeneous review teams. To address these shortcomings, we develop the C5-DM Framework with conceptual principles to address data management challenges across five areas relevant to literature reviews: data conceptualization, collection, curation, control, and consumption. Methodological guidance for researchers with respect to these five areas is necessary to reduce errors, save time on repetitive tasks, and allow review teams to develop insightful syntheses.
Meta-analysis is a cornerstone of evidence synthesis, yet challenges arise when studies report heterogeneous summary statistics, such as means and standard deviations (SDs) versus medians, interquartile ranges (IQRs), or other percentiles. Excluding studies that report only medians and IQRs can introduce bias and reduce precision, particularly when outcomes are skewed, which is common in clinical research. Although several methods exist to estimate means and SDs from alternative summaries, many rely on strong normality assumptions, exhibit computational burden, or fail to adequately account for the precision of reported quantiles (e.g., extreme values versus medians). To address these limitations, we propose two flexible weighted estimators for estimating the mean and SD from reported quantiles. The methods leverage inverse-variance and inverse-variance-covariance weighting, respectively, to enhance both accuracy and precision. Additionally, our methods are flexible enough to accommodate any set of reported quantiles and various underlying distributions, and they can be readily implemented using standard statistical software. Simulation studies demonstrate that the weighted estimators provide nearly unbiased estimates of the mean and SD with high precision in most cases, especially for large sample sizes. In a real-world meta-analysis, the estimates obtained using the proposed estimators closely aligned with those derived from true sample statistics. These approaches are particularly valuable for skewed outcomes and offer a practical and user-friendly solution for researchers seeking to integrate heterogeneous data while improving accuracy and precision.
Interest in large language models (LLMs) as a tool for meta-analyses and systematic reviews (MA/SRs) is growing. We prospectively developed 515 unique prompts by predefined screening-related categories and tested with open-access LLMs (Llama, Mistral) against four gold-standard MA/SRs from different medical fields published after the LLMs' training cut-offs, using a Python-based pipeline. Heterogeneity between prompts was quantified, and hypothetical workload/cost reduction with top-performing prompts calculated. Across 12,360 pipeline runs, LLMs versus MA/SRs reached average recall/sensitivity = 83.6 ± 17.0%, precision = 18.5 ± 15.6%, specificity = 36.6 ± 23.7% F1-score = 27.6 ± 17.2%, and accuracy = 61.1 ± 11.0%. F1-scores were significantly higher when prompts focused on methods (0.78 ± 0.40%), explicitly mentioned MA/SR screening (0.81 ± 0.37%), included the comparison MA/SR's title (5.64 ± 0.37%) or selection criteria (8.05 ± 0.68%), and with more LLM parameters (70b = 4.48 ± 0.31%, 123b = 7.77 ± 0.31%), but lower when screening abstracts instead of titles (-3.67 ± 0.28%). In LLM-base preselection, top-performing F1-score prompts (recall/sensitivity = 72.2%, specificity = 66.1%, precision = 28.6%) would reduce screening demands by 34.5%-37.5%, saving 8.4-8.8 weeks of work and 17,592-18,552. Recall/sensitivity increased with less MA/SR information contrasting F1-score results, which highlights a recall/sensitivity-precision/specificity trade-off. F1-score increased with detailed MA/SR information, while recall/sensitivity increased with shorter, zeroshot prompts. We provide the first prospectively assessed prompt engineering framework for early-stage LLM-based paper screening across medical fields. The publicly available Python pipeline and full prompt list used here support further development of LLM-based evidence synthesis.
Dissemination bias can occur when qualitative research is published selectively, potentially reducing the confidence in qualitative evidence. This retrospective cohort study aims to quantify the extent of non-dissemination of qualitative health research by following 1,123 conference abstracts. The proportion of non-dissemination, the time to publication, as well as associations between author or study characteristics and full publication were examined. For 22.8% of these studies, no full publication could be identified within at least 6 and up to 8 years after their presentation. For those that were published, median time to publication was 11 months (95% CI 10 to 12). Studies from authors affiliated with institutions in Australia were more likely to be published than those from North America (OR 4.47; 95% CI 1.58 to 18.74). Oral presentations were more likely to be published than poster presentations (OR 3.40; 95% CI 1.57 to 8.20). Studies that used two qualitative data collection methods were more likely to be published than studies that used one qualitative method only (OR 1.53; 95% CI 1.01 to 2.38). Conference abstracts that reported no funding were less likely to be published than those which reported funding (OR 0.71; 95% CI 0.51 to 0.99). Publicly funded research was more likely to be published than privately funded research (OR 2.24; 95% CI 1.16 to 4.28). Given the considerable proportion of unpublished health-related qualitative studies, there is a reason to believe that dissemination bias may impact negatively on qualitative evidence synthesis. This can, in turn, impair decision-making that uses qualitative evidence.
It is difficult to understand the safety profile of drugs based on a single clinical trial since clinical trials are often designed to prove efficacies, and sample size is not powered for safety assessment. Thus, meta-analysis would be a valuable tool to infer the safety profiles utilizing multiple studies. Individual clinical trials usually report the incidence proportions of adverse events (AEs) observed in the study. The follow-up duration may be study-specific, and furthermore different between the treatment groups within a single study. It often occurs in oncology clinical trials and if this is the case, it is hard to interpret the aggregated relative risk of AEs and compare the risk of AEs between the treatment groups with the standard meta-analysis techniques. The progression-free survival or the overall survival is often used as the primary endpoint in oncology clinical trials and the Kaplan-Meier estimates of the survival functions for the primary endpoint are often demonstrated graphically, which give us information of the follow-up duration of the AEs. We propose novel meta-analysis methods for AEs that address differences in follow-up durations by efficiently utilizing the Kaplan-Meier estimates of the primary endpoint. We adapt our approach using both simulated data and real data from a meta-analysis of bevacizumab. Simulation studies demonstrate that the proposed methods perform well when follow-up time differs between trials and groups.