
Population attributable risk percentage (PAR%) estimates the proportion of disease outcomes in a population that can be attributed to a risk factor which was developed using univariable models. However, clinical and epidemiological research is often multifactorial, particularly in infectious disease settings, where several factors may impact the outcome of interest which requires methods that can account for their correlation structure. We developed a generalisable methodology to estimate both full and partial PAR% using multivariable logistic and Cox regression models. The approach incorporated delta-method variance estimation and Fisher’s Z transformation to derive the 95 % confidence intervals after accounting for their correlation structure. Methodology was implemented by developing a SAS macro applicable to both cross-sectional and longitudinal designs. This methodology was applied to nationally representative data from South Africa’s COVID-19 Vaccine Surveys (CVACS-1 and CVACS-2) conducted during the 2021–2022 vaccination rollout to estimate the population-level impacts of various beliefs and attitudes on vaccine uptake. After adjusting for demographic and socioeconomic covariates, we examined the population-level impact of vaccine-related beliefs and behaviours on two outcomes: receipt of at least one COVID-19 vaccine dose and intention to receive a booster. The belief “I intend to get vaccinated” had the highest partial PAR% for uptake (62 %), followed by belief in vaccine effectiveness (43 %) and safety (34 %). The combined impact of all factors, i.e. full PAR%s, were 96 % (95 % CI: 89 %, 99 %) and 74 % (95 % CI: 67 %, 80 %) for on developing trust and intention to receive booster, respectively. These estimates highlight the significant contribution of key behavioural factors and the need for model-based estimation, as simple addition of partial PAR%s would result in overestimation. This study presented a methodological framework for deriving full and partial PAR% in multifactorial infectious disease settings. By combining advanced statistical techniques with real-world data and developing a SAS macro, we provided an epidemiological tool using advanced statistical methodologies for evaluating modifiable risk factors and informing targeted interventions in public health.
Objectives The approach of using HIV recency assay to estimate the counterfactual incidence rate is being used as the primary efficacy method in a few ongoing large-scale HIV pre-exposure prophylaxis (PrEP) trials, and the current available approach for the inference is based on the Wald method that leverages the asymptotic distribution of the estimators. One issue with the Wald test is that it does not work well when the number of HIV infections are small in the active arm, and it fails to work when there are zero HIV infections. As future long-acting PrEP products are becoming more efficacious, it is very likely that a small or zero number of infections will be observed in HIV prevention trials, especially for subgroup analyses or interim analyses, hence there is a pressing need to develop inference methods that work under such scenarios. Methods It is well known that when the sample size is small to moderate, likelihood ratio tests are more reliable than Wald tests in terms of actual error probabilities coming close to matching nominal levels. In this manuscript we derive the likelihood ratio test and the likelihood-based confidence intervals for HIV prevention trials based on recency assays. Results Compared with the Wald test, the proposed method works when there are zero infections. Additionally, unlike the Wald test, the p-value from the likelihood ratio test is an increasing function with respect to the number of infections, which is a desirable property as otherwise it will cause confusions. Conclusions For HIV PrEP trials based on recency assay, the likelihood-based p-value and confidence interval can be preferable to the Wald based inference methods when the number of HIV infections is expected to be small.
Objectives:Vigorous discussions are ongoing about future efficacy trial designs of candidate human immunodeficiency virus (HIV) prevention interventions. The study design challenges of HIV prevention interventions are considerable given rapid evolution of the prevention landscape and evidence of multiple modalities of highly effective products; future trials will likely be 'active-controlled', i.e., not include a placebo arm. Thus, novel design approaches are needed to accurately assess new interventions against these highly effective active controls. Methods:To discuss active control design challenges and identify solutions, an initial virtual workshop series was hosted and supported by the International AIDS Enterprise (October 2020-March 2021). Subsequent symposia discussions continue to advance these efforts. As the non-inferiority design is an important conceptual reference design for guiding active control trials, we adopt several of its principles in our proposed design approaches. Results:We discuss six potential study design approaches for formally evaluating absolute prevention efficacy given data from an active-controlled HIV prevention trial including using data from: 1) a registrational cohort, 2) recency assays, 3) an external trial placebo arm, 4) a biomarker of HIV incidence/exposure, 5) an anti-retroviral drug concentration as a mediator of prevention efficacy, and 6) immune biomarkers as a mediator of prevention efficacy. Conclusions:Our understanding of these proposed novel approaches to future trial designs remains incomplete and there are many future statistical research needs. Yet, each of these approaches, within the context of an active-controlled trial, have the potential to yield reliable evidence of efficacy for future biomedical interventions.
Randomization inference is a powerful tool in early phase vaccine trials when estimating the causal effect of a regimen against a placebo or another regimen. Randomization-based inference often focuses on testing either Fisher's sharp null hypothesis of no treatment effect for any participant or Neyman's weak null hypothesis of no sample average treatment effect. Many recent efforts have explored conducting exact randomization-based inference for other summaries of the treatment effect profile, for instance, quantiles of the treatment effect distribution function. In this article, we systematically review methods that conduct exact, randomization-based inference for quantiles of individual treatment effects (ITEs) and extend some results to a special case where naïve participants are expected not to exhibit responses to highly specific endpoints. These methods are suitable for completely randomized trials, stratified completely randomized trials, and a matched study comparing two non-randomized arms from possibly different trials. We evaluate the usefulness of these methods using synthetic data in simulation studies. Finally, we apply these methods to HIV Vaccine Trials Network Study 086 (HVTN 086) and HVTN 205 and showcase a wide range of application scenarios of the methods. R code that replicates all analyses in this article can be found in first author's GitHub page at https://github.com/Zhe-Chen-1999/ITE-Inference.
Abstract Objectives An exceptional effort by the scientific community has led to the development of multiple vaccines against COVID-19. Efficacy estimates for these vaccines have been widely communicated to the general public, but are nonetheless challenging to compare because they are based on phase 3 trials that differ in study design, definition of vaccine efficacy and the handling of cases arising shortly after vaccination. We investigate the impact of these choices on vaccine efficacy estimates, both theoretically and by re-analyzing the Janssen and Pfizer COVID-19 trial data under a uniform protocol. We moreover study the causal interpretation that can be assigned to per-protocol analyses typically performed in vaccine trials. Finally, we propose alternative estimands to measure the intrinsic vaccine efficacy in settings with delayed immune response. Methods The data of the Janssen COVID-19 trials were recreated, based on the published Kaplan-Meier curves. An estimator for the alternative causal estimand was developed using a Structural Distribution Model. Results In the data analyses, we observed rather large differences between intention-to-treat and per-protocol effect estimates. In contrast, the causal estimand and the different estimators used for per-protocol effects lead approximately to the same estimates. Conclusions In these COVID-10 vaccine trials, per-protocol effects can be interpreted as the number of cases that can be avoided by vaccination, if the vaccine would immediately induce an immune response. However, it is unclear whether this interpretation also holds in other settings.
Abstract Objectives Characterizing features of the viral rebound trajectories and identifying host, virological, and immunological factors that are predictive of the viral rebound trajectories are central to HIV cure research. We investigate if key features of HIV viral decay and CD4 trajectories during antiretroviral therapy (ART) are associated with characteristics of HIV viral rebound following ART interruption. Methods Nonlinear mixed effect (NLME) models are used to model viral load trajectories before and following ART interruption, incorporating left censoring due to lower detection limits of viral load assays. A stochastic approximation EM (SAEM) algorithm is used for parameter estimation and inference. To circumvent the computational intensity associated with maximizing the joint likelihood, we propose an easy-to-implement three-step method. Results We evaluate the performance of the proposed method through simulation studies and apply it to data from the Zurich Primary HIV Infection Study. We find that some key features of viral load during ART (e.g., viral decay rate) are significantly associated with important characteristics of viral rebound following ART interruption (e.g., viral set point). Conclusions The proposed three-step method works well. We have shown that key features of viral decay during ART may be associated with important features of viral rebound following ART interruption.
Abstract Objectives The causal impact method (CIM) was recently introduced for evaluation of binary interventions using observational time-series data. The CIM is appealing for practical use as it can adjust for temporal trends and account for the potential of unobserved confounding. However, the method was initially developed for applications involving large datasets and hence its potential in small epidemiological studies is still unclear. Further, the effects that measurement error can have on the performance of the CIM have not been studied yet. The objective of this work is to investigate both of these open problems. Methods Motivated by an existing dataset of HCV surveillance in the UK, we perform simulation experiments to investigate the effect of several characteristics of the data on the performance of the CIM. Further, we quantify the effects of measurement error on the performance of the CIM and extend the method to deal with this problem. Results We identify multiple characteristics of the data that affect the ability of the CIM to detect an intervention effect including the length of time-series, the variability of the outcome and the degree of correlation between the outcome of the treated unit and the outcomes of controls. We show that measurement error can introduce biases in the estimated intervention effects and heavily reduce the power of the CIM. Using an extended CIM, some of these adverse effects can be mitigated. Conclusions The CIM can provide satisfactory power in public health interventions. The method may provide misleading results in the presence of measurement error.
Abstract Objectives The use of correlates of protection (CoPs) in vaccination trials offers significant advantages as useful clinical endpoint substitutes. Vaccines with very high vaccine efficacy (VE) are documented in the literature (95% or above). Callegaro, A., and F. Tibaldi. 2019. “Assessing Correlates of Protection in Vaccine Trials: Statistical Solutions in the Context of High Vaccine Efficacy.” BMC Medical Research Methodology 19: 47 showed that the rare infections observed in the vaccinated groups of these trials poses challenges when applying conventionally-used statistical methods for CoP assessment such as the Prentice criteria and meta-analysis. The objective of this work is to investigate the impact of this problem on another statistical method for the assessment of CoPs called Principal stratification. Methods We perform simulation experiments to investigate the effect of high vaccine efficacy on the performance of the Principal Stratification approach. Results Similarly to the Prentice framework, simulation results show that the power of the Principal Stratification approach decreases when the VE grows. Conclusions It can be challenging to validate principal surrogates (and statistical surrogates) for vaccines with very high vaccine efficacy.
Abstract Objectives The averted infections ratio (AIR) is a novel measure for quantifying the preservation-of-effect in active-control non-inferiority clinical trials with a time-to-event outcome. In the main formulation, the AIR requires an estimate of the counterfactual placebo incidence rate. We describe two approaches for calculating confidence limits for the AIR given a point estimate of this parameter, a closed-form solution based on a Taylor series expansion (delta method) and an iterative method based on the profile-likelihood. Methods For each approach, exact coverage probabilities for the lower and upper confidence limits were computed over a grid of values of (1) the true value of the AIR (2) the expected number of counterfactual events (3) the effectiveness of the active-control treatment. Results Focussing on the lower confidence limit, which determines whether non-inferiority can be declared, the coverage achieved by the delta method is either less than or greater than the nominal coverage, depending on the true value of the AIR. In contrast, the coverage achieved by the profile-likelihood method is consistently accurate. Conclusions The profile-likelihood method is preferred because of better coverage properties, but the simpler delta method is valid when the experimental treatment is no less effective than the control treatment. A complementary Bayesian approach, which can be applied when the counterfactual incidence rate can be represented as a prior distribution, is also outlined.
Abstract Objectives: Our objective is to propose a robust approach to model daily new cases and daily new deaths due to covid-19 infection in Turkey. Methods: We consider the generalized linear model (GLM) approach for the autoregressive process (AR) with log link for modelling. We study the data between March 11, 2020 that is the date first confirmed case occurred and October 20, 2020. After a month of the first outbreak in Turkey, the first official curfew has been imposed during the weekend. Since then there have been curfews each weekend till June 1st. Hence, we include intervention effects as well as some outlying data points in the model where necessary. We use the data between March 11 and September 15 to build the models, and test the performance on the data from September 16 till October 20. We also study the consistency of the model statistics. Results: Estimated models fit data quite well. Results reveal that after the first curfew daily new Covid-19 cases decrease 18.5%. As expected, effect of the curfew gets more significant once a month is past, and daily new cases cut down 24.9%. Our approach also gives a robust estimate for the effective reproduction number that is approximately 2 meaning as of October 20, 2020 there is still a risk for an infected person to cause 2 secondary infections despite all the interventions, preventions, and rules. Conclusion: The GLM approach for AR process with log link produces consistent and robust estimates for the daily new cases and daily new deaths for the data covering almost the first year of the pandemic in Turkey. The proposed approach can also be used to model the cases in other countries.
Abstract Infectious disease transmission between individuals in a heterogeneous population is often best modelled through a contact network. This contact network can be spatial in nature, with connections between individuals closer in space being more likely. However, contact network data are often unobserved. Here, we consider the fit of an individual level model containing a spatially-based contact network that is either entirely, or partially, unobserved within a Bayesian framework, using data augmented Markov chain Monte Carlo (MCMC). We also incorporate the uncertainty about event history in the disease data. We also examine the performance of the data augmented MCMC analysis in the presence or absence of contact network observational models based upon either knowledge about the degree distribution or the total number of connections in the network. We find that the latter tend to provide better estimates of the model parameters and the underlying contact network.
Abstract Objectives The past decade has seen tremendous progress in the development of biomedical agents that are effective as pre-exposure prophylaxis (PrEP) for HIV prevention. To expand the choice of products and delivery methods, new medications and delivery methods are under development. Future trials of non-inferiority, given the high efficacy of ARV-based PrEP products as they become current or future standard of care, would require a large number of participants and long follow-up time that may not be feasible. This motivates the construction of a counterfactual estimate that approximates incidence for a randomized concurrent control group receiving no PrEP. Methods We propose an approach that is to enroll a cohort of prospective PrEP users and aug-ment screening for HIV with laboratory markers of duration of HIV infection to indicate recent infections. We discuss the assumptions under which these data would yield an estimate of the counterfactual HIV incidence and develop sample size and power calculations for comparisons to incidence observed on an investigational PrEP agent. Results We consider two hypothetical trials for men who have sex with men (MSM) and transgender women (TGW) from different regions and young women in sub-Saharan Africa. The calculated sample sizes are reasonable and yield desirable power in simulation studies. Conclusions Future one-arm trials with counterfactual placebo incidence based on a recency assay can be conducted with reasonable total screening sample sizes and adequate power to determine treatment efficacy.
Abstract Background Human immunodeficiency virus (HIV) viral failure occurs when antiretroviral therapy fails to suppress and sustain a person’s viral load count below 1,000 copies of viral ribonucleic acid per milliliter. For those newly diagnosed with HIV and living in a setting where healthcare resources are limited, such as a low- and middle-income country, the World Health Organization recommends viral load monitoring six months after initiation of antiretroviral treatment and yearly thereafter. Deviations from this schedule are made in cases where viral failure occurs or at the discretion of the clinician. Failure to detect viral failure in a timely fashion can lead to delayed administration of essential interventions. Clinical prediction models based on information available in the patient medical record are increasingly being developed and deployed for decision support in clinical medicine and public health. This raises the possibility that prediction models can be used to detect potential for viral failure in advance of viral measurements, particularly when those measurements occur infrequently. Objective Our goal is to use electronic health record data from a large HIV care program in Kenya to characterize and compare the predictive accuracy of several statistical machine learning methods for predicting viral failure at the first and second measurements following initiation of antiretroviral therapy. Predictive accuracy is measured in terms of sensitivity, specificity and area under the receiver-operator characteristic curve. Methods We trained and cross-validated 10 statistical machine learning models and algorithms on data from over 10,000 patients in the Academic Model Providing Access to Healthcare care program in western Kenya. These included parametric, non-parametric, ensemble, and Bayesian methods. The input variables included 50 items from the clinical record, hand picked in consultation with clinician experts. Predictive accuracy measures were calculated using 10-fold cross validation. Results Viral load failure rate is about 20% in this patient cohort at both the first and second measurements. Ensemble techniques generally outperformed other methods. For predicting viral failure at the first follow up measure, specificity was over 90% for these methods, but sensitivity was typically in the 50–60% range. Predictive accuracy was greater for the second follow up measure, with sensitivities over 80%. Super Learner, gradient boosting and Bayesian additive regression trees consistently outperformed other methods. For a viral failure rate of 20%, the positive predictive value for the top-performing methods is between 75 and 85%, while the negative predictive value is over 95%. Conclusion Evidence from this study suggests that machine learning techniques have potential to identify patients at risk for viral failure prior to their scheduled measurements. Ultimately, prognostic virologic assessment can help guide the administration of earlier targeted intervention such as enhanced drug resistance monitoring, rigorous adherence counseling, or appropriate next-line therapy switching. External validation studies should be used to confirm the results found here.
Abstract Great efforts are devoted to end the HIV epidemic as it continues to have profound public health consequences in the United States and throughout the world, and new interventions and strategies are continuously needed. The use of HIV sequence data to infer transmission networks holds much promise to direct public heath interventions where they are most needed. As these new methods are being implemented, evaluating their benefits is essential. In this paper, we recognize challenges associated with such evaluation, and make the case that overcoming these challenges is key to the use of HIV sequence data in routine public health actions to disrupt HIV transmission networks.
Abstract Objectives: Using the MTN-020/ASPIRE HIV prevention trial as a motivating example, our objective is to construct a joint model for the HIV exposure process through vaginal intercourse and the time to HIV infection in a population of sexually active women. By modeling participants’ HIV infection in terms of exposures, rather than time exposed, our aim is to obtain a valid estimate of the per-act efficacy of a preventive intervention.Methods: Within the context of HIV prevention trials, in which the frequency of sex acts is self-reported periodically by the participants, we model the exposure process of the trial participants with a non-homogeneous Poisson process. This approach allows for variability in the rate of sexual contacts between participants as well as variability in the rate of sexual contacts over time. The time to HIV infection for each participant is modeled as the time to the exposure that results in HIV infection, based on the modeled sexual contact rate. We propose an empirical Bayes approach for estimation. Results: We report the results of a simulation study where we evaluate the performance of our proposed approach and compare it to the traditional approach of estimating the overall reduction in HIV incidence using a Proportional Hazards Cox model. The proposed approach is also illustrated with data from the MTN-020/ASPIRE trial. Conclusions: The proposed joint modeling, along with the proposed empirical Bayes estimation approach, can provide valid estimation of the per-exposure efficacy of a preventive intervention.
Abstract West Nile virus (WNV) outbreaks raise the concern of WNV infection in donated blood and blood products destined for transfusion. We describe methods we developed to estimate time-dependent risk of WNV infection in donated blood, including improvements not previously detailed. The methods are then extended for use in estimation of the risk of WNV infection in donated cadaveric tissues by introducing stratification and stratum-specific weighting to address novel aspects of this application. Data from the WNV outbreak in Colorado in 2003 are used to estimate risk for donated cardiac tissue.
Abstract Objectives Exploratory studies that aim to evaluate novel therapeutic strategies in human cohorts often involve the collection of hundreds of variables measured over time on a small sample of individuals. Stringent error control for testing hypotheses in this setting renders it difficult to identify statistically signification associations. The objective of this study is to demonstrate how leveraging prior information about the biological relationships among variables can increase power for novel discovery. Methods We apply the class level association score statistic for longitudinal data (CLASS-LD) as an analysis strategy that complements single variable tests. An example is presented that aims to evaluate the relationships among 14 T-cell and monocyte activation variables measured with CD4 T-cell count over three time points after antiretroviral therapy (n=62). Results CLASS-LD using three classes with emphasis on T-cell activation with either classical vs. intermediate/inflammatory monocyte subsets detected associations in two of three classes, while single variable testing detected only one out of the 14 variables considered. Conclusions Application of a class-level testing strategy provides an alternative to single immune variables by defining hypotheses based on a collection of variables that share a known underlying biological relationship. Broader use of class-level analysis is expected to increase the available information that can be derived from limited sample clinical studies.
Abstract Objectives Estimation of the cascade of HIV care is essential for evaluating care and treatment programs, informing policy makers and assessing targets such as 90-90-90. A challenge to estimating the cascade based on electronic health record concerns patients “churning” in and out of care. Correctly estimating this dynamic phenomenon in resource-limited settings, such as those found in sub-Saharan Africa, is challenging because of the significant death under-reporting. An approach to partially recover information on the unobserved deaths is a double-sampling design, where a small subset of individuals with a missed clinic visit is intensively outreached in the community to actively ascertain their vital status. This approach has been adopted in several programs within the East Africa regional IeDEA consortium, the context of our motivating study. The objective of this paper is to propose a semiparametric method for the analysis of competing risks data with incomplete outcome ascertainment. Methods Based on data from double-sampling designs, we propose a semiparametric inverse probability weighted estimator of key outcomes during a gap in care, which are crucial pieces of the care cascade puzzle. Results Simulation studies suggest that the proposed estimators provide valid estimates in settings with incomplete outcome ascertainment under a set of realistic assumptions. These studies also illustrate that a naïve complete-case analysis can provide seriously biased estimates. The methodology is applied to electronic health record data from the East Africa IeDEA Consortium to estimate death and return to care during a gap in care. Conclusions The proposed methodology provides a robust approach for valid inferences about return to care and death during a gap in care, in settings with death under-reporting. Ultimately, the resulting estimates will have significant consequences on program construction, resource allocation, policy and decision making at the highest levels.
Objectives: Observational data derived from patient electronic health records (EHR) data are increasingly used for human immunodeficiency virus/acquired immunodeficiency syndrome (HIV/AIDS) research. There are challenges to using these data, in particular with regards to data quality; some are recognized, some unrecognized, and some recognized but ignored. There are great opportunities for the statistical community to improve inference by incorporating validation subsampling into analyses of EHR data.Methods: Methods to address measurement error, misclassification, and missing data are relevant, as are sampling designs such as two-phase sampling. However, many of the existing statistical methods for measurement error, for example, only address relatively simple settings, whereas the errors seen in these datasets span multiple variables (both predictors and outcomes), are correlated, and even affect who is included in the study.Results/Conclusion: We will discuss some preliminary methods in this area with a particular focus on time-to-event outcomes and outline areas of future research.
Abstract Objective To compare empirical and mechanistic modeling approaches for describing HIV-1 RNA viral load trajectories after antiretroviral treatment interruption and for identifying factors that predict features of viral rebound process. Methods We apply and compare two modeling approaches in analysis of data from 346 participants in six AIDS Clinical Trial Group studies. From each separate analysis, we identify predictors for viral set points and delay in rebound. Our empirical model postulates a parametric functional form whose parameters represent different features of the viral rebound process, such as rate of rise and viral load set point. The viral dynamics model augments standard HIV dynamics models–a class of mathematical models based on differential equations describing biological mechanisms–by including reactivation of latently infected cells and adaptive immune response. We use Monolix, which makes use of a Stochastic Approximation of the Expectation–Maximization algorithm, to fit non-linear mixed effects models incorporating observations that were below the assay limit of quantification. Results Among the 346 participants, the median age at treatment interruption was 42. Ninety-three percent of participants were male and sixty-five percent, white non-Hispanic. Both models provided a reasonable fit to the data and can accommodate atypical viral load trajectories. The median set points obtained from two approaches were similar: 4.44 log10 copies/mL from the empirical model and 4.59 log10 copies/mL from the viral dynamics model. Both models revealed that higher nadir CD4 cell counts and ART initiation during acute/recent phase were associated with lower viral set points and identified receiving a non-nucleoside reverse transcriptase inhibitor (NNRTI)-based pre-ATI regimen as a predictor for a delay in rebound. Conclusion Although based on different sets of assumptions, both models lead to similar conclusions regarding features of viral rebound process.