Abstract Background Hierarchical composite outcomes, analyzed using the win ratio, are increasingly used in randomized clinical trials (RCTs). However, methods for covariate adjustment in this context are underdeveloped, despite evidence that adjusting for prognostic variables can increase statistical power. Objectives Introducing a new covariate adjustment method for hierarchical outcomes using ordinal logistic regression, comparing it with existing approaches, and assessing whether adjustment improves power in randomized trials with hierarchical outcomes. Methods We developed an ordinal regression-based method for covariate adjustment of the win ratio and compared it with three alternatives: probability index models, inverse probability weighting, and a randomization-based estimator. Methods were applied to the EMPEROR-Preserved rial and tested through extensive simulations involving two common hierarchical outcome structures: time-to-event composites, and composites combining time-to-event with quantitative measures. Simulations assessed impacts on estimates, standard errors, and power across prognostic and non-prognostic settings. Results In RCT data and simulations, covariate adjustment consistently increased power when adjusting for prognostic baseline variables. Gains were comparable to or greater than those in conventional Cox models, with no power loss for non-prognostic covariates. Our ordinal approach performed similarly to existing methods while providing interpretable covariate effect estimates. Adjusting for baseline values of quantitative components yielded power gains according to the baseline-to-follow-up correlation. Conclusions Covariate adjustment for prognostic variables meaningfully improves efficiency in win ratio analyses for hierarchical outcomes. Our ordinal method is easily implemented and facilitates covariate effect interpretation. We recommend the broader adoption of covariate adjustment and our ordinal method in randomized trials using hierarchical outcomes.
The hazard ratio, typically estimated using Cox's famous proportional hazards model, is the most common effect measure used to describe the association or effect of a covariate on a time-to-event outcome. In recent years the hazard ratio has been argued by some to lack a causal interpretation, even in randomised trials, and even if the proportional hazards assumption holds. This is concerning, not least due to the ubiquity of hazard ratios in analyses of time-to-event data. We review these criticisms, describe how we think hazard ratios should be interpreted, and argue that they retain a valid causal interpretation. Nevertheless, alternative measures may be preferable to describe effects of exposures or treatments on time-to-event outcomes.
Estimation of hypothetical estimands in clinical trials typically does not make use of data that may be collected after the intercurrent event (ICE). Some recent papers have shown that such data can be used for estimation of hypothetical estimands, and that statistical efficiency and power can be increased compared to using estimators that only use data before the ICE. In this paper we critically examine the efficiency and bias of estimators that do and do not exploit data collected after ICEs, in a simplified setting. We find that certain imputation G-formula and G-estimators are equivalent in important special cases. We moreover find that efficiency can only be improved by assuming certain covariate effects are common between patients who do and do not experience ICEs. Even when such an assumption holds, we find that in a simple setting gains in efficiency will typically be modest. We derive analytical results for the bias of the estimators when some of their modeling assumptions are violated and evaluate these in simulations. We conclude with a discussion of the relative merits of estimators that do and do not make use of post-ICE data.
Flexible machine-learning (ML) models to generate imputations within the Multiple Imputation (MI) framework has recently gained traction, particularly in non-randomised observational settings. For randomised controlled trials (RCTs), it is unclear whether ML approaches to MI, in combination with Rubin’s Rules, provide valid inference in terms bias, confidence interval coverage, mean squared error (MSE), Type I error and power of the treatment effect estimator. We conducted two simulation studies in RCT settings that have an incomplete continuous outcome but fully observed covariates and treatment assignment. We compared Complete Cases, standard MI (MI-norm), MI with predictive mean matching (MI-PMM) and ML-based approaches to MI, including classification and regression trees (MI-CART), Random Forests (MI-RF) and SuperLearner when outcomes are missing completely at random or missing at random conditional on treatment/covariate. The first simulation explored a cross-sectional outcome with non-linear covariate-outcome relationships in the presence/absence of covariate-treatment interactions. The second simulation explored skewed repeated measures, motivated by a trial with digital outcomes. For the cross-sectional simulation without interaction, we found that Complete Cases yielded valid inference; MI-norm performed similarly, except when there is a non-linear covariate-outcome relationship and missingness depends on the covariate. ML approaches led to smaller MSE in specific non-linear settings, but provided unreliable inference for others. MI-PMM, which is the default setting in the R package mice, led to unreliable inference under some settings. In the presence of complex treatment-covariate interactions, performing MI separately by arm, either with MI-norm, MI-RF or MI-CART, provided inference that had comparable or better properties compared to Complete Cases when the analysis model omits the interaction term. In the repeated measures setting, ML approaches to MI and MI-PMM led to bias and under- or over-coverage, particularly when missingness depended on treatment. Based on the simulation findings, Complete Cases and MI-norm are more appropriate than ML approaches to MI for making inference, especially for late phase RCTs where Type I error control is crucial. While ML approaches may provide gains in MSE for complex covariate-outcome relationships, results should be interpreted with the caveat that Rubin’s Rules are not guaranteed to be valid when used with ML imputation methods, and can lead to bias in the estimated effect and/or its standard error.
When treatment policy estimands are of interest, clinical trials often attempt to collect patient data after intercurrent events (ICEs), although such data are often limited. Retrieved dropout imputation methods, which use pre-ICE and available post-ICE data to impute missing post-ICE outcomes, are commonly applied but often yield treatment effect estimates with large standard errors (SEs) and may encounter convergence issues when post-ICE data are sparse. Reference-based imputation methods are also used, but they rely on strong assumptions about post-ICE outcomes, which can lead to biased estimates if these assumptions are incorrect. To address these limitations, we previously proposed the reference-based Bayesian causal model (BCM), which incorporates a prior on the maintained effect parameter to reflect uncertainty in reference-based assumptions for missing post-ICE data. Our earlier work assumed no post-ICE data were observed. Here, we extend the BCM to incorporate available post-ICE outcomes, providing an approach that mitigates limitations of both retrieved-dropout and standard reference-based methods. We propose both a fully Bayesian model and an imputation-based approach. A simulation study was conducted to evaluate the frequentist properties of the proposed methods in settings with partially observed post-ICE data and to compare performance with existing approaches. Retrieved-dropout methods produced higher estimated SEs than the BCM, particularly when post-ICE data were sparse. Under the BCM, treatment effect SEs increased as post-ICE data became more limited for both modelling approaches. Importantly, this increase can be controlled through the prior variance of the maintained effect parameter, with more informative priors stabilising estimation when post-ICE data are scarce.
G-formula is a popular approach for estimating the effects of time-varying treatments or exposures from longitudinal data. G-formula is typically implemented using Monte-Carlo simulation, with non-parametric bootstrapping used for inference. In longitudinal data settings missing data are a common issue, which are often handled using multiple imputation, but it is unclear how G-formula and multiple imputation should be combined. We show how G-formula can be implemented using Bayesian multiple imputation methods for synthetic data, and that by doing so, we can impute missing data and simulate the counterfactuals of interest within a single coherent approach. We describe how this can be achieved using standard multiple imputation software and explore its performance using a simulation study and an application from cystic fibrosis.
BACKGROUND:Infections may increase the risk of age-related diseases such as dementia. Accelerated immunological ageing, measurable by telomere length (TL), may be a potential mechanism. However, the relationship between different infections and TL or telomere attrition remains unclear. This systematic review synthesises existing evidence on whether infections contribute to TL or telomere attrition and highlights research gaps to inform future studies. OBJECTIVE:To summarise the literature on associations between infections and telomere length or attrition. METHODS:We conducted comprehensive searches across six databases (MEDLINE, EMBASE, Web of Science, Scopus, Global Health, Cochrane Library) from inception to 22 May 2025, using concepts of infections, TL, and study type. Two researchers independently screened studies, extracted data, and assessed risk of bias (ROB) using the ROBINS-E tool. Meta-analysis was unfeasible due to heterogeneity, so a narrative synthesis was conducted. Studies were grouped by infection type, telomere measurement assay, cell type, and statistical approach. A GRADE assessment was performed to evaluate evidence quality. RESULTS:Our searches identified 10,349 studies, of which 73 met eligibility criteria. Most (59) were cross-sectional and most were published after 2000, with the earliest from 1996. Most studies were from the USA (17). HIV was the most frequently studied infection (35 studies), with 79% (excluding overlapping samples) reporting an association between HIV and reduced TL or increased telomere attrition. Findings for other infections, including herpesviruses and Human Papillomavirus were more variable. Variation in infection type, measurement assay, cell type, and statistical approach made cross-study comparisons challenging. Most studies had a high ROB, mainly due to unmeasured confounding. The GRADE assessment rated evidence quality as very low. CONCLUSIONS:Our review highlights a potential link between HIV and TL and telomere attrition. More robust longitudinal studies with standardised measurements and better confounder control are needed, particularly for non-HIV infections. PROSPERO (ID:CRD42023444854).
To precisely define the treatment effect of interest in a clinical trial, the ICH E9 estimand addendum describes that relevant so-called intercurrent events should be identified and strategies specified to deal with them. Handling intercurrent events with different strategies leads to different estimands. In this paper, we focus on estimands that involve addressing one intercurrent event with the treatment policy strategy and another with the hypothetical strategy. We define these estimands using potential outcomes and causal diagrams, considering the possible causal relationships between the two intercurrent events and other variables. We show that there are different causal estimand definitions and assumptions one could adopt, each having different implications for estimation, which is demonstrated in a simulation study. The different considerations are illustrated conceptually using a diabetes trial as an example.
The creation of the ICH E9 (R1) estimands framework has led to more precise specification of the treatment effects of interest in the design and statistical analysis of clinical trials. However, it is unclear how the new framework relates to causal inference, as both approaches appear to define what is being estimated and have a quantity labeled an estimand. Using illustrative examples, we show that both approaches can be used to define a population-based summary of an effect on an outcome for a specified population and highlight the similarities and differences between these approaches. We demonstrate that the ICH E9 (R1) estimand framework offers a descriptive, structured approach that is more accessible to non-mathematicians, facilitating clearer communication of trial objectives and results. We then contrast this with the causal inference framework, which provides a mathematically precise definition of an estimand and allows the explicit articulation of assumptions through tools such as causal graphs. Despite these differences, the two paradigms should be viewed as complementary rather than competing. The combined use of both approaches enhances the ability to communicate what is being estimated. We encourage those familiar with one framework to appreciate the concepts of the other to strengthen the robustness and clarity of clinical trial design, analysis, and interpretation.
The recently published ICH E9 addendum on estimands in clinical trials provides a framework for precisely defining the treatment effect that is to be estimated, but says little about estimation methods. Here we report analyses of a clinical trial in type 2 diabetes, targeting the effects of randomised treatment, handling rescue treatment and discontinuation of randomised treatment using the so-called hypothetical strategy. We show how this can be estimated using mixed models for repeated measures, multiple imputation, inverse probability of treatment weighting, G-formula and G-estimation. We describe their assumptions and practical details of their implementation using packages in R. We report the results of these analyses, broadly finding similar estimates and standard errors across the estimators. We discuss various considerations relevant when choosing an estimation approach, including computational time, how to handle missing data, whether to include post intercurrent event data in the analysis, whether and how to adjust for additional time-varying confounders, and whether and how to model different types of ICE separately.
Measurement error and misclassification can cause bias or loss of power in epidemiological studies. Software performing quantitative bias analysis (QBA) to assess the sensitivity of results to mismeasurement are available. However, QBA is still not commonly used in practice, partly due to a lack of knowledge of these software implementations. The features and particular use cases of these tools have not been systematically evaluated. We reviewed and summarised the latest available software tools for QBA in relation to mismeasured variables in health research. We searched the electronic database Web of Science for studies published between 1^st January 2014 and 1^st May 2024 (inclusive). We included epidemiological studies that described the use of software tools for QBA in relation to mismeasurement. We also searched for tools catalogued on the CRAN archive, in Stata manuals, and via Stata’s net command, available from within Stata or from the IDEAS/RePEc database. Tools were included if they were purpose-built, had documentation, and were applicable to epidemiological research. Data on the tools’ features and use cases were then extracted from the full article texts and software documentation. 17 publicly available software tools for QBA were identified, accessible via R, Stata, and online web tools. The tools cover various types of analysis, including regression, contingency tables, mediation analysis, longitudinal analysis, survival analysis and instrumental variable analysis. However, there is a lack of software tools performing QBA for misclassification of categorical variables and measurement error outside of the classical model. Additionally, the existing tools often require specialist knowledge. Despite the availability of several software tools, there are still gaps in the existing collection of tools that need to be addressed to enable wider usage of QBA in epidemiological studies. Efforts should be made to create new tools to assess multiple mismeasurement scenarios simultaneously, and also to increase the clarity of documentation for existing tools, and provide tutorials and examples for their usage. By doing so, the uptake of QBA techniques in epidemiology can be improved, leading to more accurate and reliable research findings.
In various missing data problems, values are not entirely missing, but are coarsened. For coarsened observations, instead of observing the true value, a subset of values - strictly smaller than the full sample space of the variable - is observed to which the true value belongs. In our motivating example for patients with endometrial carcinoma, the degree of lymphovascular space invasion (LVSI) can be either absent, focally present, or substantially present. For a subset of individuals, however, LVSI is reported as being present, which includes both non-absent options. In the analysis of such a dataset, difficulties arise when coarsened observations are to be used in an imputation procedure. To our knowledge, no clear-cut method has been described in the literature on how to handle an observed subset of values, and treating them as entirely missing could lead to biased estimates. Therefore, in this paper, we evaluated the best strategy to deal with coarsened and missing data in multiple imputation. We tested a number of plausible ad hoc approaches, possibly already in use by statisticians. Additionally, we propose a principled approach to this problem, consisting of an adaptation of the SMC-FCS algorithm (SMC-FCS CoCo $$ {}_{\mathrm{CoCo}} $$ : Coarsening compatible), that ensures that imputed values adhere to the coarsening information. These methods were compared in a simulation study. This comparison shows that methods that prevent imputations of incompatible values, like the SMC-FCS CoCo $$ {}_{\mathrm{CoCo}} $$ method, perform consistently better in terms of a lower bias and RMSE, and achieve better coverage than methods that ignore coarsening or handle it in a more naïve way. The analysis of the motivating example shows that the way the coarsening information is handled can matter substantially, leading to different conclusions across methods. Overall, our proposed SMC-FCS CoCo $$ {}_{\mathrm{CoCo}} $$ method outperforms other methods in handling coarsened data, requires limited additional computation cost and is easily extendable to other scenarios.
The Fine-Gray model for the subdistribution hazard is commonly used for estimating associations between covariates and competing risks outcomes. When there are missing values in the covariates included in a given model, researchers may wish to multiply impute them. Assuming interest lies in estimating the risk of only one of the competing events, this paper develops a substantive-model-compatible multiple imputation approach that exploits the parallels between the Fine-Gray model and the standard (single-event) Cox model. In the presence of right-censoring, this involves first imputing the potential censoring times for those failing from competing events, and thereafter imputing the missing covariates by leveraging methodology previously developed for the Cox model in the setting without competing risks. In a simulation study, we compared the proposed approach to alternative methods, such as imputing compatibly with cause-specific Cox models. The proposed method performed well (in terms of estimation of both subdistribution log hazard ratios and cumulative incidences) when data were generated assuming proportional subdistribution hazards, and performed satisfactorily when this assumption was not satisfied. The gain in efficiency compared to a complete-case analysis was demonstrated in both the simulation study and in an applied data example on competing outcomes following an allogeneic stem cell transplantation. For individual-specific cumulative incidence estimation, assuming proportionality on the correct scale at the analysis phase appears to be more important than correctly specifying the imputation procedure used to impute the missing covariates.
OBJECTIVE To compare the effectiveness of three commonly prescribed oral antidiabetic drugs added to metformin for people with type 2 diabetes mellitus requiring second line treatment in routine clinical practice. DESIGN Cohort study emulating a comparative effectiveness trial (target trial). SETTING Linked primary care, hospital, and death data in England, 2015-21. PARTICIPANTS 75 739 adults with type 2 diabetes mellitus who initiated second line oral antidiabetic treatment with a sulfonylurea, DPP-4 inhibitor, or SGLT-2 inhibitor added to metformin. MAIN OUTCOME MEASURES Primary outcome was absolute change in glycated haemoglobin A 1c (HbA 1 c ) between baseline and one year follow-up. Secondary outcomes were change in body mass index (BMI), systolic blood pressure, and estimated glomerular filtration rate (eGFR) at one year and two years, change in HbA 1c at two years, and time to >= 40% decline in eGFR, major adverse kidney event, hospital admission for heart failure, major adverse cardiovascular event (MACE), and all cause mortality. Instrumental variable analysis was used to reduce the risk of confounding due to unobserved baseline measures. RESULTS 75 739 people initiated second line oral antidiabetic treatment with sulfonylureas (n=25 693, 33.9%), DPP4 inhibitors (n=34 464 ,45.5%), or SGLT-2 inhibitors (n=15 582, 20.6%). SGLT-2 inhibitors were more effective than DPP-4 inhibitors or sulfonylureas in reducing mean HbA 1c values between baseline and one year. After the instrumental variable analysis, the mean differences in HbA 1c change between baseline and one year were -2.5 mmol/mol (95% confidence interval (CI) -3.7 to -1.3) for SGLT-2 inhibitors versus sulfonylureas and -3.2 mmol/mol (-4.6 to -1.8) for SGLT-2 inhibitors versus DPP-4 inhibitors. SGLT-2 inhibitors were more effective than sulfonylureas or DPP-4 inhibitors in reducing BMI and systolic blood pressure. For some secondary endpoints, evidence for SGLT-2 inhibitors being more effective was lacking- the hazard ratio for MACE, for example, was 0.99 (95% CI 0.61 to 1.62) versus sulfonylureas and 0.91 (0.51 to 1.63) versus DPP-4 inhibitors. SGLT-2 inhibitors had reduced hazards of hospital admission for heart failure compared with DPP-4 inhibitors (0.32, 0.12 to 0.90) and sulfonylureas (0.46, 0.20 to 1.05). The hazard ratio for a >= 40% decline in eGFR indicated a protective effect versus sulfonylureas (0.42, 0.22 to 0.82), with high uncertainty in the estimated hazard ratio versus DPP-4 inhibitors (0.64, 0.29 to 1.43). CONCLUSIONS This emulation study of a target trial found that SGLT-2 inhibitors were more effective than sulfonylureas or DPP-4 inhibitors in lowering mean HbA 1 c , BMI, and systolic blood pressure and in reducing the hazards of hospital admission for heart failure ( v DPP-4 inhibitors) and kidney disease progression ( v sulfonylureas), with no evidence of differences in other clinical endpoints.
When using multiple imputation (MI) for missing data, maintaining compatibility between the imputation model and substantive analysis is important for avoiding bias. For example, some causal inference methods incorporate an outcome model with exposure-confounder interactions that must be reflected in the imputation model. Two approaches for compatible imputation with multivariable missingness have been proposed: Substantive-Model-Compatible Fully Conditional Specification (SMCFCS) and a stacked-imputation-based approach (SMC-stack). If the imputation model is correctly specified, both approaches are guaranteed to be unbiased under the "missing at random" assumption. However, this assumption is violated when the outcome causes its own missingness, which is common in practice. In such settings, sensitivity analyses are needed to assess the impact of alternative assumptions on results. An appealing solution for sensitivity analysis is delta-adjustment using MI, specifically "not-at-random" (NAR)FCS. However, the issue of imputation model compatibility has not been considered in sensitivity analysis, with a naive implementation of NARFCS being susceptible to bias. To address this gap, we propose two approaches for compatible sensitivity analysis when the outcome causes its own missingness. The proposed approaches, NAR-SMCFCS and NAR-SMC-stack, extend SMCFCS and SMC-stack, respectively, with delta-adjustment for the outcome. We evaluate these approaches using a simulation study that is motivated by a case study, to which the methods were also applied. The simulation results confirmed that a naive implementation of NARFCS produced bias in effect estimates, while NAR-SMCFCS and NAR-SMC-stack were approximately unbiased. The proposed compatible approaches provide promising avenues for conducting sensitivity analysis to missingness assumptions in causal inference.
We appreciate Cro et al.'s efforts to bring wider attention to the debate surrounding variance estimation for reference-based imputation methods. However, we believe that the way this debate is presented as "multiple imputation" versus "conditional mean imputation" can be misleading. Both of these imputation methods rely on identical assumptions and provide essentially identical treatment effect estimates. While conditional mean imputation naturally focuses on the frequentist repeated sampling variance, we show here that it can be easily adapted to target a variance with similar properties to Rubin's variance. Therefore, conditional mean imputation combined with jackknife resampling remains a valid and effective deterministic method for handling missing data under missing-at-random or reference-based assumptions regardless of the user's preference for variance estimation. We also reappraise the frequentist variance by arguing that it correctly reflects the strong assumptions of reference-based imputation. In contrast, we are not aware of any frequentist or Bayesian framework under which Rubin's variance provides correct inference.
Background: Mismeasurement (measurement error or misclassification) can cause bias or loss of power. However, sensitivity analyses (e.g. using quantitative bias analysis, QBA) are rarely used. Methods: We reviewed software tools for QBA for mismeasurement in health research identified by searching Web of Science, the CRAN archive, and the IDEAS/RePEc software components database. Tools were included if they were purpose-built, had documentation and were applicable to epidemiological research. Results: 16 freely available software tools for QBA were identified, accessible via R and online web tools. The tools handle various types of mismeasurement, including classical measurement error and binary misclassification. Only one software tool handles misclassification of categorical variables, and few tackle non-classical measurement error. Conclusions: Efforts should be made to create tools that can assess multiple mismeasurement scenarios simultaneously, to increase the clarity of documentation for existing tools, and provide tutorials for their usage. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement CJCW is supported by the Engineering and Physical Sciences Research Council (EPSRC) (grant EP/S023569/1). RAH is supported by a Sir Henry Dale Fellowship that is jointly funded by the Wellcome Trust and the Royal Soci- ety (grant 215408/Z/19/Z). KMT works in the MRC Integrative Epidemiology Unit, which is supported by the University of Bristol and the Medical Research Council (grant MC UU 00032/2). JWB is supported by the UK Medical Research Council (grant MR/T023953/1). ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes This study did not involve any underlying data. Computing code is available as supplemental digital content.
Nonproportional hazards (NPH) have been observed in confirmatory clinical trials with time to event outcomes. Under NPH, the hazard ratio does not stay constant over time and the log rank test is no longer the most powerful test. The weighted log rank test (WLRT) has been introduced to deal with the presence of nonproportionality. We focus our attention on the WLRT and the complementary Cox model based on time varying treatment effect proposed by Lin and León. We investigate whether the proposed weighted hazard ratio (WHR) approach is unbiased in scenarios where the WLRT statistic is the most powerful test. In the diminishing treatment effect scenario where the WLRT statistic would be most optimal, the time varying treatment effect estimated by the Cox model estimates the treatment effect very close to the true one. However, when the true hazard ratio is large the proposed model overestimates the treatment effect and the treatment profile over time. In the delayed treatment scenario, the estimated treatment effect profile over time is typically close to the true profile. For both scenarios, we have demonstrated analytically that the hazard ratio functions are approximately equal under small treatment effects. When the assumed rate of how quickly the treatment effect profile is diminishing or delaying differs in the analysis from that in the true data generating mechanism, the estimated hazard ratio profile from the WHR approach is biased. Since in practice the true HR time profile may differ from that assumed in the WHR analysis, it may be preferable to use alternative approaches for effect estimation.
Introduction Telomeres are a measure of cellular ageing with potential links to diseases such as cardiovascular diseases and cancer. Studies have shown that some infections may be associated with telomere shortening, but whether an association exists across all types and severities of infections and in which populations is unclear. Therefore we aim to collate available evidence to enable comparison and to inform future research in this field.Methods and analysis We will search for studies involving telomere length and infection in various databases including MEDLINE (Ovid interface), EMBASE (Ovid interface), Web of Science, Scopus, Global Health and the Cochrane Library. For grey literature, the British Library of electronic theses databases (ETHOS) will be explored. We will not limit by study type, geographical location, infection type or method of outcome measurement. Two researchers will independently carry out study selection, data extraction and risk of bias assessment using the ROB2 and ROBINS-E tools. The overall quality of the studies will be determined using the Grading of Recommendations Assessment, Development and Evaluation criteria. We will also evaluate study heterogeneity with respect to study design, exposure and outcome measurement and if there is sufficient homogeneity, a meta-analysis will be conducted. Otherwise, we will provide a narrative synthesis with results grouped by exposure category and study design.Ethics and dissemination The present study does not require ethical approval. Results will be disseminated via publishing in a peer-reviewed journal and conference presentations.PROSPERO registration number CRD42023444854.
The main aim of many epidemiological studies is to estimate the causal effect of an exposure on an outcome. When data is obtained for such studies, there is potential for some of the exposure, confounders, mediators, effect modifiers, or outcomes to be measured with error. Where we have categorical variables, we refer to this measurement error as misclassification. If measurement error and misclassification are not appropriately accounted for, erroneous study conclusions may be reached. Quantitative bias analysis (QBA) can be applied to studies that have not accounted for measurement error and be used to quantify the potential impact of measurement error, or how much measurement error would be needed to result in changes to the study conclusions. Currently, QBA methods are not implemented as a standard practise, in some part due to a lack of awareness about accessible software for the purpose. With this review, we aim to identify the available software that implements a QBA for studies with measurement error or misclassification.