Surrogate endpoints are often used in place of expensive, delayed, or rare clinical endpoints in clinical trials. However, regulatory authorities require thorough evaluation to accept these surrogate endpoints as reliable substitutes. One evaluation approach is the information-theoretic causal inference framework, which quantifies surrogacy using the individual causal association (ICA). Like most causal inference methods, this approach relies on models that are only partially identifiable. For continuous outcomes, a normal model is often used. In this study, we explored the effects of model misspecification across various scenarios. We first considered true data-generating mechanisms based on multivariate t $$ t $$ and log-normal distributions. We then used D-vine copulas with Gaussian, Clayton, Gumbel, and Frank families to vary the unidentifiable copulas involving counterfactual pairs while preserving the observable bivariate margins, and considered intuitive restrictions on nonidentified correlations, including positivity and conditional independence. In all settings, the identifiability issue was addressed through sensitivity analysis. Finally, we illustrate the proposed sensitivity analyses using clinical-trial data from schizophrenia studies, evaluating the ICA under several modeling assumptions. The results show that, in most scenarios considered, the impact of model misspecification is small; however, certain departures from the assumed model can materially affect the surrogacy assessment.
In progressive diseases such as Alzheimer's, treatments that slow progression should start early to preserve higher levels of functioning for a longer period. In corresponding clinical trials, treatment effects are usually expressed as mean differences on a clinical scale at fixed time points. Early in the disease course, however, these mean differences may appear small but may nonetheless correspond to an important slowing of disease progression. This complicates the appreciation of the relevance of observed treatment effects. We introduce a class of target parameters that quantify treatment effects on the time scale in longitudinal studies; for instance, in terms of time saved or percentage slowing of progression. We focus on data from randomized trials where the target parameters are identified under regularity assumptions. These target parameters remain well defined if treatment was not randomized, but additional untestable assumptions are required for identification. We propose general two-step estimators. In the first step, the data can be analyzed with standard methods for longitudinal data and standard software can thus be used. In the second step, summary statistics from the first step are used for inferences about the target parameters. The second step has been implemented in the TCT R package. We study the asymptotic properties and efficiency of these two-step estimators, and evaluate them in an extensive simulation study. These estimators are used in a phase 2/3 clinical trial for Alzheimer's disease, leading to important additional insights into the treatment effect.
In clinical trials, surrogate endpoints, that are more cost-effective, occur earlier, or are more frequently measured, are sometimes used to replace costly, late, or rare true endpoints. Regulatory authorities typically require thorough evaluation and validation to accept these surrogate endpoints as reliable substitutes. To this end, the meta-analytic framework is considered a very viable approach to validate surrogates at both trial and individual levels. However, this framework requires data from multiple trials or centers, posing challenges when data sharing is not feasible. In this article, we propose a federated data analysis approach that allows organizations to maintain control over their datasets while still enabling surrogate validation through meta-analytic techniques. In this approach, there is no longer a need for raw data sharing. Instead, independent analyses are conducted at each organization. Thereafter, the results of these independent analyses are aggregated at a central analysis hub and the metrics for surrogate evaluation are extracted. We apply this approach to simulated and real clinical data, demonstrating how this federated approach can overcome data-sharing constraints and validate surrogate endpoints in decentralized settings.
Surrogate endpoints provide several significant advantages in evaluating new drugs, including reduced follow-up times, smaller required sample sizes, and lower overall costs. The Individual Causal Association (ICA), first introduced by Alonso et al. within the causal inference framework, assesses the validity of a surrogate based on the mutual information between the individual causal treatment effects on both the surrogate and the true endpoint. Mutual information is a specific instance from the broader family of information-theoretic measures known as the R & eacute;nyi divergence family. In this article, we extend the ICA by proposing a new family of surrogacy metrics derived from R & eacute;nyi divergence. We evaluate the performance of this extended family through theoretical analysis and an extensive simulation study. Finally, we provide an illustrative application using clinical trial data from schizophrenia studies.
In empirical studies, multiple outcomes are often measured repeatedly over time, and interest frequently lies in studying the association between these longitudinal outcomes and a time-to-event outcome. Therefore, shared-parameter joint models for longitudinal and time-to-event outcomes have been developed. However, while such joint models in theory also allow for multiple longitudinal outcomes, they are often restricted to a limited number of outcomes due to computational complexity when fitting the models. To address this problem, we propose a new joint model, which is based on correlated instead of shared random effects, and for which a pairwise-modelling strategy can be used. In this approach, the longitudinal outcomes are modelled with (generalized) linear mixed models and the survival outcome with a Weibull proportional hazards frailty model. Instead of fitting the full joint model, this approach involves fitting all possible bivariate models, and inference is based on pseudo-likelihood theory. The main advantage of our approach is that there is no restriction on the number of longitudinally measured outcomes that are jointly modelled with the time-to-event outcome.
AbstractIn progressive diseases, like Alzheimer’s disease, treatments that slow progression should start early in the disease course to longer maintain higher levels of functioning. In corresponding clinical trials, the treatment effect is usually expressed in terms of mean differences on a clinical scale. Early in the disease course, however, treatment effects expressed on a clinical scale are often small but may nonetheless correspond to an important slowing of disease progression. This complicates the appreciation of the relevance of observed treatment effects. For example, it may be difficult to determine whether a 2-point improvement on a clinical scale is relevant for clinical practice. In this paper, we propose the meta Time-Component Tests (meta TCT). This new approach leads to estimators of treatment effects on the time scale, in terms of time saved or percentage slowing of progression, that are easy to interpret. This approach is based on estimates obtained from an arbitrary model for longitudinal data and is, therefore, very flexible. Asymptotic properties of the Meta TCT estimators are derived and evaluated in an extensive simulation study. Meta TCT is then applied to a phase 2/3 clinical trial for Alzheimer’s disease, which was first analyzed with a mixed model. In this trial, meta TCT leads to important additional insights into the treatment effect. We believe that meta TCT will facilitate the estimation of interpretable treatment effects in clinical trials for progressive diseases, and that this, in turn, will fine-tune the evaluation of the clinical relevance of new treatments.
Putative surrogate endpoints must undergo a rigorous statistical evaluation before they can be used in clinical trials. Numerous frameworks have been introduced for this purpose. In this study, we extend the scope of the information-theoretic causal-inference approach to encompass scenarios where both outcomes are time-to-event endpoints, using the flexibility provided by D-vine copulas. We evaluate the quality of the putative surrogate using the individual causal association (ICA)—a measure based on the mutual information between the individual causal treatment effects. However, in spite of its appealing mathematical properties, the ICA may be ill defined for composite endpoints. Therefore, we also propose an alternative rank-based metric for assessing the ICA. Due to the fundamental problem of causal inference, the joint distribution of all potential outcomes is only partially identifiable and, consequently, the ICA cannot be estimated without strong unverifiable assumptions. This is addressed by a formal sensitivity analysis that is summarized by the so-called intervals of ignorance and uncertainty. The frequentist properties of these intervals are discussed in detail. Finally, the proposed methods are illustrated with an analysis of pooled data from two advanced colorectal cancer trials. The newly developed techniques have been implemented in the R package Surrogate.
One of the key tools to understand and reduce the spread of the SARS-CoV-2 virus is testing. The total number of tests, the number of positive tests, the number of negative tests, and the positivity rate are interconnected indicators and vary with time. To better understand the relationship between these indicators, against the background of an evolving pandemic, the association between the number of positive tests and the number of negative tests is studied using a joint modeling approach. All countries in the European Union, Switzerland, the United Kingdom, and Norway are included in the analysis. We propose a joint penalized spline model in which the penalized spline is reparameterized as a linear mixed model. The model allows for flexible trajectories by smoothing the country-specific deviations from the overall penalized spline and accounts for heteroscedasticity by allowing the autocorrelation parameters and residual variances to vary among countries. The association between the number of positive tests and the number of negative tests is derived from the joint distribution for the random intercepts and slopes. The correlation between the random intercepts and the correlation between the random slopes were both positive. This suggests that, when countries increase their testing capacity, both the number of positive tests and negative tests will increase. A significant correlation was found between the random intercepts, but the correlation between the random slopes was not significant due to a wide credible interval.
The selection of the primary endpoint in a clinical trial plays a critical role in determining the trial’s success. Ideally, the primary endpoint is the clinically most relevant outcome, also termed the true endpoint. However, practical considerations, like extended follow-up, may complicate this choice, prompting the proposal to replace the true endpoint with so-called surrogate endpoints. Evaluating the validity of these surrogate endpoints is crucial, and a popular evaluation framework is based on the proportion of treatment effect explained (PTE). While methodological advancements in this area have focused primarily on estimation methods, interpretation remains a challenge hindering the practical use of the PTE. We review various ways to interpret the PTE. These interpretations—two causal and one non-causal—reveal connections between the PTE principal surrogacy, causal mediation analysis, and the prediction of trial-level treatment effects. A common limitation across these interpretations is the reliance on unverifiable assumptions. As such, we argue that the PTE is only meaningful when researchers are willing to make very strong assumptions. These challenges are also illustrated in an analysis of three hypothetical vaccine trials.
In a causal inference framework, a new metric has been proposed to quantify surrogacy for a continuous putative surrogate and a binary true endpoint, based on information theory. The proposed metric, termed the individual causal association (ICA), was quantified using a joint causal inference model for the corresponding potential outcomes. Due to the non-identifiability inherent in this type of models, a sensitivity analysis was introduced to study the behavior of the ICA as a function of the non-identifiable parameters characterizing the aforementioned model. In this scenario, to reduce uncertainty, several plausible yet untestable assumptions like monotonicity, independence, conditional independence or homogeneous variance-covariance, are often incorporated into the analysis. We assess the robustness of the methodology regarding these simplifying assumptions via simulation. The practical implications of the findings are demonstrated in the analysis of a randomized clinical trial evaluating an inactivated quadrivalent influenza vaccine.
Surrogate endpoints are often used in place of expensive, delayed, or rare true endpoints in clinical trials. However, regulatory authorities require thorough evaluation to accept these surrogate endpoints as reliable substitutes. One evaluation approach is the information-theoretic causal inference framework, which quantifies surrogacy using the individual causal association (ICA). Like most causal inference methods, this approach relies on models that are only partially identifiable. For continuous outcomes, a normal model is often used. Based on theoretical elements and a Monte Carlo procedure we studied the impact of model misspecification across two scenarios: 1) the true model is based on a multivariate t-distribution, and 2) the true model is based on a multivariate log-normal distribution. In the first scenario, the misspecification has a negligible impact on the results, while in the second, it has a significant impact when the misspecification is detectable using the observed data. Finally, we analyzed two data sets using the normal model and several D-vine copula models that were indistinguishable from the normal model based on the data at hand. We observed that the results may vary when different models are used.
Within the causal association paradigm, a method is proposed to assess the validity of a continuous outcome as a surrogate for a binary true endpoint. The methodology is based on a previously introduced information-theoretic definition of surrogacy and has two main steps. In the first step, a new model is proposed to describe the joint distribution of the potential outcomes associated with the putative surrogate and the true endpoint of interest. The identifiability issues inherent to this type of models are handled via sensitivity analysis. In the second step, a metric of surrogacy new to this setting, the so-called individual causal association is presented. The methodology is studied in detail using theoretical considerations, some simulations, and data from a randomized clinical trial evaluating an inactivated quadrivalent influenza vaccine. A user-friendly R package Surrogate is provided to carry out the evaluation exercise.
May 15, 2015 Type Package Title Evaluation of Surrogate Endpoints in Clinical Trials Version 0.1-6 Date 2015-05-14 Author Wim Van der Elst, Ariel Alonso & Geert Molenberghs Maintainer Wim Van der Elst Description In a clinical trial, it frequently occurs that the most credible outcome to evaluate the effectiveness of a new therapy (the true endpoint) is difficult to measure. In such a situation, it can be an effective strategy to replace the true endpoint by a biomarker that is easier to measure and that allows for a prediction of the treatment effect on the true endpoint (a surrogate endpoint). The package 'Surrogate' allows for an evaluation of the appropriateness of a candidate surrogate endpoint based on the meta-analytic, informationtheoretic, and causal-inference frameworks. Part of this software has been developed using funding provided from the European Union's 7th Framework Programme for research, technological development and demonstration under Grant Agreement no 602552. Depends MASS, nlme, msm, lme4 Imports rgl, lattice, latticeExtra License GPL (>= 2) BugReports Wim Van der Elst Repository CRAN NeedsCompilation no Date/Publication 2015-05-15 00:12:46
The identification of good surrogate endpoints is a challenging endeavor. This may, at least partially, be attributable to the fact that most researchers have focused on the identification of a single surrogate endpoint. It is thus implicitly assumed that the treatment effect on the true endpoint (T) can be accurately predicted based on the treatment effect on one surrogate endpoint (S) only. Given the complex nature of many diseases and the different therapeutic pathways in which a treatment can impact T, this assumption may be too optimistic. For example, in oncology, the effect of a treatment often depends on both the treatment's efficacy and its toxicity. In the present article, the meta-analytic framework of? is extended to the setting where multiple S are considered. To cope with potential model convergence issues that often arise in a meta-analytic framework, several simplified model fitting strategies are proposed. Further, simulation studies are conducted to evaluate the properties of the estimated surrogacy metrics, and the new methodology is applied on a case study in schizophrenia. An online Appendix that details how the analyses can be conducted in practice (using the R package Surrogate) is also provided.
To monitor the COVID-19 epidemic in Cuba, data on several epidemiological indicators have been collected on a daily basis for each municipality. Studying the spatio-temporal dynamics in these indicators, and how they behave similarly, can help us better understand how COVID-19 spread across Cuba. Therefore, spatio-temporal models can be used to analyze these indicators. Univariate spatio-temporal models have been thoroughly studied, but when interest lies in studying the association between multiple outcomes, a joint model that allows for association between the spatial and temporal patterns is necessary. The purpose of our study was to develop a multivariate spatio-temporal model to study the association between the weekly number of COVID-19 deaths and the weekly number of imported COVID-19 cases in Cuba during 2021. To allow for correlation between the spatial patterns, a multivariate conditional autoregressive prior (MCAR) was used. Correlation between the temporal patterns was taken into account by using two approaches; either a multivariate random walk prior was used or a multivariate conditional autoregressive prior (MCAR) was used. All models were fitted within a Bayesian framework.
Multivariate surrogate endpoints can improve the efficiency of the drug development process, but their evaluation raises many challenges. Recently, the so-called individual causal association (ICA) has been introduced for validation purposes in the causal-inference paradigm. The ICA is a function of a partially identifiable correlation matrix (R) and, hence, it cannot be estimated without making untestable assumptions. This issue has been addressed via a simulation-based analysis. Essentially, the ICA is assessed across a set of values for the non-identifiable entries in R that lead to a valid correlation matrix and this has been implemented using a fast algorithm based on partial correlations (PC). Using theoretical arguments and simulations, it is shown that, in spite of its computational efficiency, the PC algorithm may lead to the spurious effect that adding non-informative surrogates, i.e., surrogates that convey no information on the treatment effect on the true endpoint, seemingly reduces the ICA range. To address this, a modified PC algorithm (MPC) is proposed. Based on simulations, it is shown that the MPC algorithm removes this nuisance effect and increases computational efficiency.
The meta-analytic approach has become the gold-standard methodology for the evaluation of surrogate endpoints and several implementations are currently available in SAS and R. The methodology is based on hierarchical models that are numerically demanding and, when the amount of data is limited, maximum likelihood algorithms may not converge or may converge to an ill-conditioned maximum such as a boundary solution. This may produce misleading conclusions and have negative implications for the evaluation of new drugs. In the present work, we explore the use of two distinct functions in R (lme and lmer) and the MIXED procedure in SAS to assess the validity of putative surrogate endpoints in the meta-analytic framework, via simulations and the analysis of a real case study. We describe some problems found with the lmer function in R that led to a poorer performance as compared with the lme function and MIXED procedure.
In this study, a new cost-benefit economic model for hemodialyzer reuse has been developed considering all of the direct costs (dialyzer price, disinfection fluid price, reverse osmosis water cost, personnel, and miscellaneous) as well as the number of disinfections applied to the hemodialyzer. The maximum number of disinfections/reuses for different models of hemodialyzer was estimated using statistical analysis based on the information obtained from a total of 60 adult patients on maintenance hemodialysis for approximately 4 years; from a hospital in Santiago de Cuba province, Cuba. An equal number of treatments (100) was evaluated for each hemodialyzer including 2,800 total reuses. The total cost savings for reuse/disinfection using the new economic approach is compared with the single-use modality. Obtained results by applying the proposed model indicated that the correlation between the economic advantages of the reuse/disinfection process in the total cost of the hemodialysis treatment significantly depends on the type of hemodialyzer used for the treatment, the disinfection price, the virgin hemodialyzer price, the disposal cost and its price reduction because of the number of reuses and the maximum possible reuses that could be applied to the hemodialyzer.
A surrogate endpoint is a biomarker, intended for substituting a clinical endpoint. Surrogate endpoints can play a role in the earlier detection of safety signals that could point to toxicity problems with new drugs. This chapter gives a perspective on data from a single trial. It presents the meta-analytic evaluation framework in the context of normally distributed outcomes. M. Buyse and G. Molenberghs proposed quantity for the validation of a surrogate endpoint: the relative effect, which is the ratio of the effects of treatment upon the true and the surrogate endpoint. A variety of surrogate marker evaluation strategies have been proposed, cast within a meta-analytic framework. The meta-analytic approach was formulated originally for two continuous, normally distributed outcomes, and extended in the meantime to various outcome types, ranging from continuous, binary, ordinal, time-to-event, and longitudinally measured outcomes. The chapter discusses the settings of binary endpoints, failure-time endpoints, the combination of an ordinal and a survival endpoint, and longitudinal endpoints.
In the meta-analytic surrogate evaluation framework, the trial-level coefficient of determination Rtrial2 quantifies the strength of the association between the expected causal treatment effects on the surrogate (S) and the true (T) endpoints. Burzykowski and Buyse supplemented this metric of surrogacy with the surrogate threshold effect (STE), which is defined as the minimum value of the causal treatment effect on S for which the predicted causal treatment effect on T exceeds zero. The STE supplements Rtrial2 with a more direct clinically interpretable metric of surrogacy. Alonso et al. proposed to evaluate surrogacy based on the strength of the association between the individual (rather than expected) causal treatment effects on S and T. In the current paper, the individual-level surrogate threshold effect (ISTE) is introduced in the setting where S and T are normally distributed variables. ISTE is defined as the minimum value of the individual causal treatment effect on S for which the lower limit of the prediction interval around the individual causal treatment effect on T exceeds zero. The newly proposed methodology is applied in a case study, and it is illustrated that ISTE has an appealing clinical interpretation. The R package surrogate implements the methodology and a web appendix (supporting information) that details how the analyses can be conducted in practice is provided.