Incomplete and inaccurate reporting of research can prevent research from being used. Reporting guidelines aim to circumvent this by providing recommendations that support researchers in writing complete and accurate accounts of their work. While reporting guidelines have many benefits, challenges exist for both developers and users. In this article, we outline initiatives implemented to address challenges within the PRISMA (Preferred Reporting Items for Systematic reviews and Meta-Analyses) family of guidelines and highlight opportunities for future work.
Defining outcomes completely before conducting clinical trials helps to mitigate reporting biases; however, there is limited guidance to help investigators define outcomes completely. We aimed to develop a structured approach for defining trial outcomes completely and consistently. We reviewed literature, developed preliminary rules for defining outcomes, and refined them iteratively. We randomly selected randomized controlled trials (RCTs) on ClinicalTrials.gov that registered before their start dates and posted results by January 4, 2024. The 225 included RCTs evaluated 3,424 outcomes. Two raters independently applied preliminary rules to define each outcome. When raters encountered outcomes they could not define, we refined the rules. We continued this process until no further changes were needed. We discussed and finalized our approach in a consensus meeting. We define an “outcome” as a value for each participant that will be used in analysis to generate study results. A complete outcome definition includes six elements: outcome domain, specific measurement, specific metric, cutoff, variable type, and timepoint. We developed rules for naming specific measurements for both subjective and objective outcomes. We expanded on prior work by developing more comprehensive categories for specific metrics. We introduced "cutoff" as a distinct element with three subelements. To clarify the boundary between outcome definitions and statistical methods, we replaced a previously described element, "method of aggregation," with "variable type," which refers to whether the value for each individual is continuous or categorical. Trialists and sponsors could use this approach alongside other guidelines to define outcomes in trial registrations, protocols, and result reports.
ObjectivesThe interrupted time series (ITS) design is commonly used to investigate the impact of an intervention or exposure in public health. There are many statistical methods that can be used to analyse ITS data and to meta-analyse their results. We undertook two empirical studies to investigate: (i) how effect estimates (and associated statistics) compared when six statistical methods were applied to 190 real-world datasets; and (ii) how meta-analysis effect estimates (and associated statistics) compared when the combinations of two ITS analysis methods and five meta-analysis methods were applied to 17 real-world meta-analyses including 283 ITS datasets. Here we present a curated repository of a subset of ITS datasets from these studies.Data descriptionThe repository includes 430 ITS datasets curated from the two empirical studies. The datasets are diverse in the populations, interruptions and outcomes examined, and are methodologically diverse in the outcome types, aggregation time intervals, number of timepoints and segments. Most of the datasets are from public health. For each dataset, we provide the outcome value at each timepoint and the segment (indicating different interruptions), along with characteristics of the dataset. This repository may be of value for future research of ITS studies, and as a source of examples of ITS for use in teaching.
Network meta-analysis allows the synthesis of relative effects from several treatments. Two broad approaches are available to synthesize the data: arm-synthesis and contrast-synthesis, with several models that can be fitted within each. Limited evaluations comparing these approaches are available. We re-analyzed 118 networks of interventions with binary outcomes using three contrast-synthesis models (CSM; one fitted in a frequentist framework and two in a Bayesian framework) and two arm-synthesis models (ASM; both fitted in a Bayesian framework). We compared the estimated log odds ratios, their standard errors, ranking measures and the between-trial heterogeneity using the different models and investigated if differences in the results were modified by network characteristics. In general, we observed good agreement with respect to the odds ratios, their standard errors and the ranking metrics between the two Bayesian CSMs. However, differences were observed when comparing the frequentist CSM and the ASMs to each other and to the Bayesian CSMs. The network characteristics that we investigated, which represented the connectedness of the networks and rareness of events, were associated with the differences observed between models, but no single factor was associated with the differences across all of the metrics. In conclusion, we found that different models used to synthesize evidence in a network meta-analysis (NMA) can yield different estimates of odds ratios and standard errors that can impact the final ranking of the treatment options compared.
Floods of unprecedented intensity and frequency have been observed. However, evidence regarding the impacts of floods on hospitalization remains limited. Here we collected daily hospitalization counts during 2000-2019 from 747 communities in Australia, Brazil, Canada, Chile, New Zealand, Taiwan, Thailand and Vietnam. For each community, flooded days were defined as days from the start dates to the end dates of flood events. Lag-response associations between flooded day and daily hospitalization risks were estimated for each community using a quasi-Poisson regression model with a distributed lag nonlinear function. The community-specific estimates were then pooled using a random-effects meta-analysis. Based on the pooled estimates, attributable fractions of hospitalizations due to floods were calculated. We found that hospitalization risks increased and persisted for up to 210 days after flood exposure, with the overall relative risks being 1.26 (95% confidence interval 1.15-1.38) for all causes, 1.35 (1.21-1.50) for cardiovascular diseases, 1.30 (1.13-1.49) for respiratory diseases, 1.26 (1.10-1.44) for infectious diseases, 1.30 (1.17-1.45) for digestive diseases, 1.11 (0.98-1.25) for mental disorders, 1.61 (1.39-1.86) for diabetes, 1.35 (1.21-1.50) for injury, 1.34 (1.21-1.48) for cancer, 1.34 (1.20-1.50) for nervous system disorders and 1.40 (1.22-1.60) for renal diseases. The associations were modified by climate types, flood severity, age, population density and socioeconomic status. Flood exposure contributed to hospitalizations by up to 0.27% from all causes. This study revealed that flood exposure was associated with increased all-cause and ten cause-specific hospitalization risks within up to 210 days after exposure.
Objective: The objective of this scoping review is to develop a list of items for potential inclusion in the Preferred Reporting Items for Systematic reviews and Meta-Analyses (PRISMA) reporting guidelines for network meta-analysis (NMA), scoping reviews (ScRs), and rapid reviews (RRs). Introduction: The PRISMA extensions for NMA and ScRs were published in 2015 and 2018. However, since then, their methodologies and innovations, including automation, have evolved. There is no reporting guideline for RRs. In 2020, an updated PRISMA statement was published, reflecting advances in the conduct and reporting of systematic reviews. These advances are not yet incorporated into these PRISMA extensions. We will update our previous methods for scoping reviews to inform the update of PRISMA-NMA and PRISMA-ScR as well as the development of the PRISMA-RR reporting guidelines. Inclusion criteria: This review will include any study design evaluating the completeness of reporting, offering reporting guidance, or assessing methods relevant to NMA, ScRs, or RRs. Editorial guidelines and tutorials that describe items related to reporting completeness will also be eligible. Methods: We will follow the JBI guidance for scoping reviews. For each PRISMA extension, we will i) search multiple electronic databases from inception to present, ii) search for unpublished studies, and iii) scan the reference lists of included studies. There will be no language limitations. Screening and data extraction will be conducted by 2 researchers independently. A third researcher will resolve discrepancies. We will conduct frequency analyses of the identified items. The final list of items will be considered for potential inclusion in the relevant PRISMA reporting guidelines. Review registration: NMA protocol (OSF: osf.io/7bkwy); ScR protocol (OSF: osf.io/7bkwy); RR protocol (OSF: osf.io/3jcpe); EQUATOR registration link: https://www.equator-network.org/library/reporting-guidelines-under-development/reporting-guidelines-under-development-for-systematic-reviews/
Background and Objective This scoping review is the first step in the process of updating the 2015 Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) extension for network meta-analysis (NMA). It builds up on a 2014 scoping review by our team and aims to enhance the usability, completeness, and transparency of NMA reporting for diverse audiences, including patients and the public. The updated extension will align with PRISMA 2020 and will address gaps in reporting, such as documenting methods for assessing NMA homogeneity and transitivity, defining intervention nodes in a network of studies, and considering advances in statistical modeling with NMA. Methods We registered the study protocol with the Open Science Framework and published it in Joanna Briggs Institute (JBI) Evidence Synthesis. We searched multiple databases and gray literature sources, and screened studies in duplicate. Data extraction was also conducted in duplicate using a standardized form, focusing on study characteristics, authors' reporting recommendations, and proposed additions to PRISMA-NMA and/or PRISMA 2020. Results Sixty-one studies met eligibility criteria, including 23 guidance documents and 38 overviews of reviews assessing the completeness or quality of NMA reporting. We identified 37 additional reporting items relevant to NMAs, which will inform the next stage of the PRISMA-NMA update (a Delphi consensus process). Conclusion Our findings support the urgent need to update the PRISMA-NMA guideline. Addressing persistent reporting gaps and incorporating recent methodological developments is critical to improving the transparency, reproducibility, and trustworthiness of NMAs.
This article presents the CONSORT (consolidated standards of reporting trials) extension for cluster randomised crossover trials. A cluster randomised crossover trial involves randomisation of groups of individuals (known as clusters) to different sequences of interventions over time. The design has gained popularity in settings where cluster randomisation is required because it can largely overcome the loss in power due to clustering in parallel cluster trials. However, the design has many methodological complexities, requiring tailored reporting guidance. The guideline was developed using a survey and in-person consensus meeting, informed by a systematic review examining the quality of reporting in cluster randomised crossover trials and relevant CONSORT statements for individual, crossover, cluster, and stepped wedge designs. This article also provides recommended reporting items, along with explanations and examples.
BACKGROUND:Exposure to floods might increase the risks of adverse birth outcomes. However, the current evidence is scarce, inconsistent, and has knowledge gaps. This study aims to estimate the associations of flood exposure before and during pregnancy with adverse birth outcomes and to identify susceptible exposure windows and effect modifiers. METHODS:In this cohort study, we obtained all the birth records occurring in Greater Sydney, Australia, from Jan 1, 2001, to Dec 31, 2020, from the New South Wales Midwives Data Collection and in the Brisbane metropolitan region, Australia, from Jan 1, 1995, to Dec 31, 2014, from the Queensland Health Perinatal Data Collection. For each birth, residential address and historical flood information from the Dartmouth Flood Observatory were used to estimate the numbers of days with floods during five exposure windows (Pre-1 was defined as 13-24 weeks before the last menstrual period [LMP], Pre-2 was 0-12 weeks before the LMP, trimester 1 [Tri-1] was 0-12 weeks after the LMP, trimester 2 [Tri-2] was 13-28 weeks after the LMP, and trimester 3 [Tri-3] was ≥29 weeks after the LMP). We estimated the hazard ratios (HRs) of adverse birth outcomes (preterm births, stillbirths, term low birthweight [TLBW], and small for gestational age [SGA]) associated with flood exposures in the five exposure windows using Cox proportional hazards regression models. FINDINGS:1 338 314 birth records were included in our analyses, which included 91 851 (6·9%) preterm births, 9831 (0·7%) stillbirths, 25 567 (1·9%) TLBW, and 108 658 (8·1%) SGA. Flood exposure in Pre-1 was associated with increased risks of TLBW (HR 1·06 [95% CI 1·01-1·12]) and SGA (1·04 [1·01-1·06]); flood exposure during Tri-1 was associated with increased risks of preterm births (1·03 [1·002-1·05]), stillbirth (1·11 [1·03-1·20]), and SGA (1·03 [1·01-1·06]). In contrast, flood exposures during Pre-2 and Tri-3 were associated with reduced risks. INTERPRETATION:Exposures to floods in Pre-1 and Tri-1 are both associated with increased risks of adverse birth outcomes, and the risks increase with a higher exposure. Upon planning for conception and prenatal care, individuals and health practitioners should raise awareness of the increased risks of adverse birth outcomes after experiencing floods. FUNDING:The Australian Research Council and the Australian National Health and Medical Research Council.
OBJECTIVES:In research evaluating statistical analysis methods, a common aim is to compare point estimates and CIs calculated from different analyses. This can be challenging when the outcomes (and their scale ranges) differ across datasets. We therefore developed a graphical method, the "Banksia plot", to facilitate pairwise comparisons of different statistical analysis methods by plotting and comparing point estimates and CIs from each analysis method, both within and across datasets. STUDY DESIGN AND SETTING:The plot is constructed in three stages. Stage 1: To compare the results of two statistical analysis methods, for each dataset, the point estimate from the reference analysis method is centered on zero, and its confidence limits are scaled to range from -0.5 to 0.5. The same centering and scale adjustment values are then applied to the corresponding comparator analysis point estimate and confidence limits. Stage 2: A Banksia plot is constructed by plotting the centered and scaled point estimates from the comparator method for each dataset on a rectangle centered at zero, ranging from -0.5 to 0.5, which represents the reference method results. Stage 3: Optionally, a matrix of Banksia plots is graphed, showing all pairwise comparisons from multiple analysis methods. We illustrate the Banksia plot using two examples. RESULTS:Illustration of the Banksia plot demonstrates how the plot makes it immediately apparent whether there are differences in point estimates and CIs when using different analysis methods (example 1) or different data extractors (example 2). Furthermore, we demonstrate how different bases for ordering the CIs can be used to highlight particular differences (ie, in point estimates or CI widths). CONCLUSION:The Banksia plot provides a visual summary of pairwise comparisons of different analysis methods, allowing patterns and trends in the point estimates and CIs to be easily identified.
Background The Interrupted Time Series (ITS) is a robust design for evaluating public health and policy interventions or exposures when randomisation may be infeasible. Several statistical methods are available for the analysis and meta-analysis of ITS studies. We sought to empirically compare available methods when applied to real-world ITS data. Methods We sourced ITS data from published meta-analyses to create an online data repository. Each dataset was re-analysed using two ITS estimation methods. The level- and slope-change effect estimates (and standard errors) were calculated and combined using fixed-effect and four random-effects meta-analysis methods. We examined differences in meta-analytic level- and slope-change estimates, their 95% confidence intervals, p-values, and estimates of heterogeneity across the statistical methods. Results Of 40 eligible meta-analyses, data from 17 meta-analyses including 282 ITS studies were obtained (predominantly investigating the effects of public health interruptions (88%)) and analysed. We found that on average , the meta-analytic effect estimates, their standard errors and between-study variances were not sensitive to meta-analysis method choice, irrespective of the ITS analysis method. However, across ITS analysis methods, for any given meta-analysis, there could be small to moderate differences in meta-analytic effect estimates, and important differences in the meta-analytic standard errors. Furthermore, the confidence interval widths and p -values for the meta-analytic effect estimates varied depending on the choice of confidence interval method and ITS analysis method. Conclusions Our empirical study showed that meta-analysis effect estimates, their standard errors, confidence interval widths and p -values can be affected by statistical method choice. These differences may importantly impact interpretations and conclusions of a meta-analysis and suggest that the statistical methods are not interchangeable in practice.
Background Systematic reviews that aim to synthesize evidence on the effects of interventions targeted at populations often include interrupted time-series (ITS) studies. However, the suppression of ITS studies or results within these studies (known as reporting bias) has the potential to bias conclusions drawn in such systematic reviews, with potential consequences for healthcare decision-making. Therefore, we aim to determine whether there is evidence of reporting bias among ITS studies. Methods We will conduct a search for published protocols of ITS studies and reports of their results in PubMed, MEDLINE, and Embase up to December 31, 2022. We contact the authors of the ITS studies to seek information about their study, including submission status, data for unpublished results, and reasons for non-publication or non-reporting of certain outcomes. We will examine if there is evidence of publication bias by examining whether time-to-publication is influenced by the statistical significance of the study’s results for the primary research question using Cox proportional hazards regression. We will examine whether there is evidence of discrepancies in outcomes by comparing those specified in the protocols with those in the reports of results, and we will examine whether the statistical significance of an outcome’s result is associated with how completely that result is reported using multivariable logistic regression. Finally, we will examine discrepancies between protocols and reports of results in the methods by examining the data collection processes, model characteristics, and statistical analysis methods. Discrepancies will be summarized using descriptive statistics. Discussion These findings will inform systematic reviewers and policymakers about the extent of reporting biases and may inform the development of mechanisms to reduce such biases.
In this report we present the Consolidated Standards of Reporting Trials (CONSORT) extension for the cluster randomised crossover (CRXO) trial. A CRXO trial involves randomisation of groups of individuals (‘clusters’) to different sequences of treatments over time. The design has gained popularity in settings where cluster randomisation is required as it can largely overcome the loss in power due to clustering in parallel cluster trials. However, the design has many methodological complexities, requiring tailored reporting guidance. The guideline was developed using a survey and in-person consensus meeting, informed by a systematic review examining the quality of reporting in CRXO trials and relevant CONSORT statements for individual, crossover, cluster and stepped-wedge designs. We provide recommended reporting items, along with explanations and examples.
Objective: To generate a bank of items describing application and interpretation errors that can arise in pairwise meta-analyses in systematic reviews of interventions.Study design and setting: Medline, Embase and Scopus were searched to identify studies describing types of errors in meta-analyses. Descriptions of errors and supporting quotes were extracted by multiple authors. Errors were reviewed at team meetings to determine if they should be excluded, reworded, or combined with other errors, and were categorised into broad categories of errors and subcategories within.Results: 50 articles met our inclusion criteria, leading to the identification of 139 errors. We identified 25 errors covering data extraction/manipulation, 74 covering statistical analyses, and 40 covering interpretation. Many of the statistical analysis errors related to the meta-analysis model (e.g. using a two-stage strategy to determine whether to select a fixed or random-effects model) and statistical heterogeneity (e.g. not undertaking an assessment for statistical heterogeneity).Conclusions: We generated a comprehensive bank of possible errors that can arise in the application and interpretation of meta-analyses in systematic reviews of interventions. This item bank of errors provides the foundation for developing a checklist to help peer reviewers detect statistical errors.
Background Interrupted time-series (ITS) studies are commonly used to examine the effects of interventions targeted at populations. Suppression of ITS studies or results within these studies, known as reporting bias, has the potential to bias the evidence-base on a particular topic, with potential consequences for healthcare decision-making. Therefore, we aim to determine whether there is evidence of reporting bias among ITS studies. Methods We will conduct a search for published protocols of ITS studies and reports of their results in PubMed, MEDLINE, and Embase up to December 31, 2022. We contact the authors of the ITS studies to seek information about their study, including submission status, data for unpublished results, and reasons for non-publication or non-reporting of certain outcomes. We will examine if there is evidence of publication bias by examining whether time-to-publication is influenced by the statistical significance of the study’s results for the primary research question using Cox proportional hazards regression. We will examine whether there is evidence of discrepancies in outcomes by comparing those specified in the protocols with those in the reports of results, and we will examine whether the statistical significance of an outcome’s result is associated with how completely that result is reported using multivariable logistic regression. Finally, we will examine discrepancies between protocols and reports of results in the methods by examining the data collection processes, model characteristics, and statistical analysis methods. Discrepancies will be summarized using descriptive statistics. Discussion These findings will inform systematic reviewers and policymakers about the extent of reporting biases and may inform the development of mechanisms to reduce such biases.
We aimed to explore, in a sample of systematic reviews (SRs) with meta-analyses of the association between food/diet and health-related outcomes, whether systematic reviewers selectively included study effect estimates in meta-analyses when multiple effect estimates were available. We randomly selected SRs of food/diet and health-related outcomes published between January 2018 and June 2019. We selected the first presented meta-analysis in each review (index meta-analysis), and extracted from study reports all study effect estimates that were eligible for inclusion in the meta-analysis. We calculated the Potential Bias Index (PBI) to quantify and test for evidence of selective inclusion. The PBI ranges from 0 to 1; values above or below 0.5 suggest selective inclusion of effect estimates more or less favourable to the intervention, respectively. We also compared the index meta-analytic estimate to the median of a randomly constructed distribution of meta-analytic estimates (i.e., the estimate expected when there is no selective inclusion). Thirty-nine SRs with 312 studies were included. The estimated PBI was 0.49 (95% CI 0.42-0.55), suggesting that the selection of study effect estimates from those reported was consistent with a process of random selection. In addition, the index meta-analytic effect estimates were similar, on average, to what we would expect to see in meta-analyses generated when there was no selective inclusion. Despite this, we recommend that systematic reviewers report the methods used to select effect estimates to include in meta-analyses, which can help readers understand the risk of selective inclusion bias in the SRs.
Meta-analysis is a statistical method used to combine results from multiple studies, providing a quantitative summary of their findings. One of the fundamental decisions in conducting a meta-analysis is choosing an appropriate model to estimate the overall effect size and its confidence interval. In this article, we focus on the common-effect (also referred to as the fixed-effect) model, and in a companion article, the random-effects model. These models are the two prevailing meta-analysis models employed in the literature. In this article we outline the key assumption underlying the common-effect model, describe different common-effect methods (i.e., inverse variance, Peto, and Mantel-Haenszel), and highlight characteristics of the meta-analysis that should be considered when selecting a method. Furthermore, we demonstrate the application of these methods to a dataset. Understanding the common-effect model is important for knowing when to use the model and how to interpret the overall effect size and its confidence interval.