Obtaining valid real-world evidence about intervention effects from observational cohorts or administrative health records data is challenging. Visits to healthcare providers tend to occur more often during periods of increased disease activity and symptom exacerbation, or upon disease progression. Treatments likewise tend to change when it is apparent that disease activity has increased or a meaningful progression has occurred. This creates a dual problem in which patient visits are disease-related and treatments changes are driven by disease condition and clinical presentation. Disease-related visits and treatment by indication can produce a biased impression of the disease process in the target population and of the effects of treatment. We discuss how these challenges can be addressed through the use of joint models for the disease, marker, and treatment processes, as well as the observation (visit) process. Using illustrative multistate models, we demonstrate the biases that can arise from various types of analyses and show how estimators from fitting such joint models to persons with psoriatic arthritis can be used to gain scientific insights and address common questions about treatment effects.
[This corrects the article DOI: 10.1371/journal.pgen.1011192.].
Cohort studies of disease processes deal with events and other outcomes that may occur in individuals following disease onset. The particular goals are often the evaluation of interventions and estimation of the effects of risk factors that may affect the disease course. Models and methods of event history analysis and longitudinal data analysis provide tools for understanding disease processes, but there are numerous challenges in practice. These are related to the complexity of the disease processes and to the difficulty of recruiting representative individuals and acquiring detailed longitudinal data on their disease course. Our objectives here are to describe some of these challenges and to review methods of addressing them. We emphasize the appeal of multistate models as a framework for understanding both disease processes and the processes governing recruitment of individuals for cohort studies and the collection of data. The use of other observational data sources in order to enhance model fitting and analysis is discussed.
In genome-wide association studies (GWAS), it is often desirable to test for interactions, such as gene-environment (G x E) or gene-gene (G x G) interactions, between single-nucleotide polymorphisms (SNPs, G's) and environmental variables (E's). However, directly accounting for interaction is often infeasible, because the interacting variable is latent or the computational burden is too large. For quantitative traits (Y) that are approximately normally distributed, it has been shown that indirect testing on GxE can be done by testing for heteroskedasticity of Y between genotypes. However, when traits are binary, the existing methodology based on testing the heteroskedasticity of the trait across genotypes cannot be generalized. In this paper, we propose an approach to indirectly test interaction effects for binary traits and subsequently propose a joint test that accounts for the main and interaction effects of each SNP during GWAS. The final method is straightforward to implement in practice-it simply involves adding a non-additive (i.e., dominance) term to standard GWAS additive models for binary traits and testing its significance. We illustrate the statistical features including type-I-error control and power of the proposed method through extensive numerical studies. Applying our method to the UK Biobank dataset, we showcase the practical utility of the proposed method, revealing SNPs and genes with strong potential for latent interaction effects.
In complex diseases, individuals are often at risk of several types of possibly semi-competing events and may experience recurrent symptomatic episodes. This complex disease course makes it challenging to define target estimands for clinical trials. While composite endpoints are routinely adopted, recent innovations involving the win ratio and other methods based on ranking the disease course have received considerable attention. We emphasize the usefulness of multistate models for addressing challenges arising in complex diseases, along with the simplicity and interpretability that come from defining utilities to synthesize evidence of treatment effects on different aspects of the disease process. Robust variance estimation based on the infinitesimal jackknife means that such methods can be used as the basis of primary analyses of clinical trials. We illustrate the use of utilities for the assessment of bleeding outcomes in a trial of cancer patients with thrombocytopenia.
In life history analysis of data from cohort studies, it is important to address the process by which participants are identified and selected. Many health studies select or enrol individuals based on whether they have experienced certain health related events, for example, disease diagnosis or some complication from disease. Standard methods of analysis rely on assumptions concerning the independence of selection and a person's prospective life history process, given their prior history. Violations of such assumptions are common, however, and can bias estimation of process features. This has implications for the internal and external validity of cohort studies, and for the transportabilty of results to a population. In this paper, we study failure time analysis by proposing a joint model for the cohort selection process and the failure process of interest. This allows us to address both independence assumptions and the transportability of study results. It is shown that transportability cannot be guaranteed in the absence of auxiliary information on the population. Conditions that produce dependent selection and types of auxiliary data are discussed and illustrated in numerical studies. The proposed framework is applied to a study of the risk of psoriatic arthritis in persons with psoriasis.
Clinical trials with random assignment of treatment provide evidence about causal effects of an experimental treatment compared to standard care. However, when disease processes involve multiple types of possibly semi-competing events, specification of target estimands and causal inferences can be challenging. Intercurrent events such as study withdrawal, the introduction of rescue medication, and death further complicate matters. There has been much discussion about these issues in recent years, but guidance remains ambiguous. Some recommended approaches are formulated in terms of hypothetical settings that have little bearing in the real world. We discuss issues in formulating estimands, beginning with intercurrent events in the context of a linear model and then move on to more complex disease history processes amenable to multistate modeling. We elucidate the meaning of estimands implicit in some recommended approaches for dealing with intercurrent events and highlight the disconnect between estimands formulated in terms of potential outcomes and the real world.
To advance scientific understanding of disease processes and related intervention effects, study results should be free from bias and replicable. More broadly, investigators seek results that are transportable, that is, applicable to a perceived study population as well as in other environments and populations. We review fundamental statistical issues that arise in the analysis of observational data from disease cohorts and other sources and discuss how these issues affect the transportability and replicability of research results. Much of the literature focuses on estimating average exposure or intervention effects at the population level, but we argue for more nuanced analyses of conditional effects that reflect the complexity of disease processes.
The HostSeq initiative recruited 10,059 Canadians infected with SARS-CoV-2 between March 2020 and March 2023, obtained clinical information on their disease experience and whole genome sequenced (WGS) their DNA. We analyzed the WGS data for genetic contributors to severe COVID-19 (considering 3,499 hospitalized cases and 4,975 non-hospitalized after quality control). We investigated the evidence for replication of loci reported by the International Host Genetics Initiative (HGI); analyzed the X chromosome; conducted rare variant gene-based analysis and polygenic risk score testing. Population stratification was adjusted for using meta-analysis across ancestry groups. We replicated two loci identified by the HGI for COVID-19 severity: the LZTFL1/SLC6A20 locus on chromosome 3 and the FOXP4 locus on chromosome 6 (the latter with a variant significant at P < 5E-8). We found novel significant associations with MRAS and WDR89 in gene-based analyses, and constructed a polygenic risk score that explained 1.01% of the variance in severe COVID-19. This study provides independent evidence confirming the robustness of previously identified COVID-19 severity loci by the HGI and identifies novel genes for further investigation.
In genome-wide association studies (GWAS), it is desirable to test for interactions ( GxE ) between single-nucleotide polymorphisms (SNPs, G ’s) and environmental variables ( E ’s). However, directly accounting for interaction is often infeasible, because E is latent. For quantitative traits ( Y ) that are approximately normally distributed, it has been shown that indirect testing on GxE can be done by testing for heteroskedasticity of Y between genotypes. However, when traits are binary, the existing methodology based on testing the heteroskedasticity of the trait across genotypes cannot be generalized. In this paper, we propose an approach to indirectly test GxE for binary traits based on the non-additive effect G , and subsequently propose a joint test that accounts for the main and interaction effects of each SNP during GWAS. We illustrate the statistical features including type-I-error control and power of the proposed method through extensive numerical studies. Applying our method to the UK Biobank dataset, we showcase the practical utility of the proposed method, revealing SNPs and genes with strong potential for latent interaction effects. ### Competing Interest Statement The authors have declared no competing interest.
Regression analyses based on transformations of cumulative incidence functions are often adopted when modeling and testing for treatment effects in clinical trial settings involving competing and semi-competing risks. Common frameworks include the Fine-Gray model and models based on direct binomial regression. Using large sample theory we derive the limiting values of treatment effect estimators based on such models when the data are generated according to multiplicative intensity-based models, and show that the estimand is sensitive to several process features. The rejection rates of hypothesis tests based on cumulative incidence function regression models are also examined for null hypotheses of different types, based on which a robustness property is established. In such settings supportive secondary analyses of treatment effects are essential to ensure a full understanding of the nature of treatment effects. An application to a palliative study of individuals with breast cancer metastatic to bone is provided for illustration.
Intensity-based multistate models provide a useful framework for characterizing disease processes, the introduction of interventions, loss to followup, and other complications arising in the conduct of randomized trials studying complex life history processes. Within this framework we discuss the issues involved in the specification of estimands and show the limiting values of common estimators of marginal process features based on cumulative incidence function regression models. When intercurrent events arise we stress the need to carefully define the target estimand and the importance of avoiding targets of inference that are not interpretable in the real world. This has implications for analyses, but also the design of clinical trials where protocols may help in the interpretation of estimands based on marginal features.
Two-phase study designs are ideal for focused sub-studies based on large prospective cohorts when the outcome of interest is an event that is rare in the full cohort, and additional covariates are expensive or difficult to measure. Researchers often wish to examine large numbers of covariates for association with outcomes of interest. In the context of cancer, hundreds to millions of genetic markers may be considered, along with environmental exposures. A computationally efficient variable selection method is proposed for two-phase failure time studies with stratified sampling under the Cox proportional hazards model. The penalized estimator is obtained from a penalized (weighted) Cox log partial likelihood using a pathwise cyclical coordinate descent algorithm which is scalable for high dimensional datasets where the number of features is much larger than the sample size (p≫n). A detailed simulation study to examine the performance of the proposed methodology is described. The variable selection and estimation procedure is then used to obtain a model for predicting acute myeloid leukaemia using somatic stem cell mutation profiles derived from blood samples, based on a two-phase sample from the European Prospective Investigation into Cancer and Nutrition (EPIC) study.
Two-phase outcome dependent sampling (ODS) is widely used in many fields, especially when certain covariates are expensive and/or difficult to measure. For two-phase ODS, the conditional maximum likelihood (CML) method is very attractive because it can handle zero Phase 2 selection probabilities and avoids modeling the covariate distribution. However, most existing CML-based methods use only the Phase 2 sample and thus may be less efficient than other methods. We propose a general empirical likelihood method that uses CML augmented with additional information in the whole Phase 1 sample to improve estimation efficiency. The proposed method maintains the ability to handle zero selection probabilities and avoids modeling the covariate distribution, but can lead to substantial efficiency gains over CML in the inexpensive covariates, or in the influential covariate when a surrogate is available, because of an effective use of the Phase 1 data. Simulations and a real data illustration using NHANES data are presented.
Studies of chronic disease often involve modeling the relationship between marker processes and disease onset or progression. The Cox regression model is perhaps the most common and convenient approach to analysis in this setting. In most cohort studies, however, biospecimens and biomarker values are only measured intermittently (e.g. at clinic visits) so Cox models often treat biomarker values as fixed at their most recently observed values, until they are updated at the next visit. We consider the implications of this convention on the limiting values of regression coefficient estimators when the marker values themselves impact the intensity for clinic visits. A joint multistate model is described for the marker-failure-visit process which can be fitted to mitigate this bias and an expectation-maximization algorithm is developed. An application to data from a registry of patients with psoriatic arthritis is given for illustration.
Disease progression is often monitored by intermittent follow-up "visits" in longitudinal cohort studies, resulting in interval-censored failure time outcomes. Furthermore, the timing and frequency of visits is often found related to a person's history of disease-related variables in practice. This article develops a semiparametric estimation approach using weighted binomial regression and a kernel smoother to analyze interval-censored failure time data. Visit times are allowed to be subject-specific and outcome-dependent. We consider a collection of widely used semiparametric regression models, including additive hazards and linear transformation models. For additive hazards models, the nonparametric component has a closed-form estimator and the estimators of regression coefficients are shown to be asymptotically multivariate normal with sandwich-type covariance matrices. Simulations are conducted to examine the finite sample performance of the proposed estimators. A data set from the Toronto Psoriatic Arthritis (PsA) Cohort Study is used to illustrate the proposed methodology.
Life history analysis has evolved in the last 50 years as a methodology for analyzing processes associated with human health, education, employment, and other areas. The complexity of many processes, the difficulty of obtaining complete and accurate data, and the increased use of observational data from registries and administrative sources have posed many recent challenges. We review the evolution of life history analysis, discuss some recent work, and consider three areas currently receiving much attention. A theme we stress is the use of expanded models that include selection and observation processes for studies in addition to the life history process of interest. Examples from health research are presented.
Canadian Journal of StatisticsVolume 49, Issue 4 p. 990-991 Editorial Covid-19-related content in The Canadian Journal of Statistics Jerry Lawless, Jerry Lawless Guest Editor University of Waterloo, CanadaSearch for more papers by this author Jerry Lawless, Jerry Lawless Guest Editor University of Waterloo, CanadaSearch for more papers by this author First published: 25 November 2021 https://doi.org/10.1002/cjs.11672Read the full textAboutPDF ToolsRequest permissionExport citationAdd to favoritesTrack citation ShareShare Give accessShare full text accessShare full-text accessPlease review our Terms and Conditions of Use and check box below to share full-text version of article.I have read and accept the Wiley Online Library Terms and Conditions of UseShareable LinkUse the link below to share a full-text version of this article with your friends and colleagues. Learn more.Copy URL Share a linkShare onEmailFacebookTwitterLinkedInRedditWechat No abstract is available for this article. Volume49, Issue4December 2021Pages 990-991 RelatedInformation