
Competing risks, referring to multiple mutually exclusive cause-specific events in survival data, are commonly seen in medical and engineering studies. In the presence of covariates, the Cox proportional hazards model is the most common approach to modeling covariate effects as the relative risk component for each cause-specific hazard function. However, the linear form of covariate effects in the existing models can be too restrictive to satisfy in practice. In this work, we propose a Cox regression model for competing risk survival data where the relative risk of each cause-specific hazard takes a flexible nonparametric form. The effects of multiple covariates are modeled under a smoothing spline ANOVA framework allowing for nonparametric effect components among covariates. For each cause-specific hazard, we develop an estimation procedure based on the optimization of penalized partial likelihood. We show that our spline estimators of the log relative risk functions can achieved the optimal nonparametric rates of convergence. Through simulation studies, we demonstrate the numerical performance of the proposed method. We then apply our approach to a multiple myeloma dataset, where the effects of gene expressions on cancer and non-cancer related death are investigated. The nonlinear trend revealed in our analysis offers a strong support to the need of a nonparametric modeling approach like ours.
Early identification of Parkinson’s disease is critical for timely intervention. Disruptions in sleep architecture and changes in physical activity patterns have been reported years before motor symptom onset, yet prodromal behavioral changes spanning sleep and daytime activity patterns are subtle and difficult to detect with conventional clinical tools. Wearable sensors provide a scalable means of monitoring these behaviors in natural settings, but extracting meaningful, interpretable features from high-frequency, unlabeled time series remains a major challenge. We present an end-to-end framework for assessing Parkinson’s disease risk from wrist-worn accelerometry via automated feature extraction and survival modeling of time to diagnosis. Behavioral states are derived from unlabeled data using pretrained models including random forests with hidden Markov models for physical activity classification and a self-supervised learning based sleep staging model. We jointly model both sleep stage and physical activity sequences using hierarchical nonstationary Markov chains stratified by time of day and temporal resolution, characterizing individual-level behavioral rhythms and transition dynamics. Gradient-boosted Cox proportional hazards models are then used to estimate Parkinson’s disease risk from these transition-based features. Applied to accelerometer data from the UK Biobank, our approach outperforms baselines that exclude dynamic modeling or rely on traditional functional data analysis, while providing interpretable predictors from long sequences of wearable sensor data. This demonstrates the potential of integrating pretrained AI models, temporal sequence modeling, and survival analysis to detect early behavioral signatures of neurodegeneration.
Given a double-truncated sample of lifespans, we test the hypothesis of a parametric distribution family for the lifespan. Demography typically certifies the life expectancy to be nonstationary in time. We model the resulting dependence between a lifespan and the birthday of an individual with a copula. Our main example is the Farlie-Gumbel-Morgenstern copula. The asymptotic null distribution of the test is based on Donsker-class arguments and the functional delta method for empirical processes. One assumption for the test is consistency in the estimation of the parameter for the model under the null hypothesis. Requirements for consistency are fulfilled by our set of assumptions. The test is carried out in two stages, the computation of the test statistics and the simulation of the critical value. The statistic can be reached after finitely many computation steps, whereas the critical value can only be an approximation. With the exponential distribution as an example for the lifespan distribution, and for the application to 55,000 German double-truncated enterprise lifespans, the constructed Kolmogorov-Smirnov test rejects clearly an age-homogeneous closure hazard.
The threshold regression model, first introduced by Lee and Whitmore in 2006 by combining the concepts of the first hitting time model and linear regression, has been widely applied in the medical context. The model features enhanced capabilities in incorporating information from auxiliary variables to model the target variable. These advances are evident not only in the medical context, but also in finance. Therefore, in this study, we extend the concept of the threshold regression model to financial data focusing on stock market prices. We construct a threshold regression model based on geometric Brownian motion, the conventional underlying process for the first hitting time models in financial data. We demonstrate advantages of the proposed model via applications to the stock exchange of Thailand (SET) compared with three reference models including the first hitting time model based on Brownian motion, the first hitting time model based on geometric Brownian motion, and the threshold regression model based on Brownian motion. Moreover, we enhance the quality of the proposed model by adapting a mixture cure fraction to the model. The results indicate that the proposed threshold regression based on geometric Brownian motion with a cure fraction outperforms other competing models.
Analysis of left-truncated and interval-censored survival data is challenging, particularly when the failure time and observation process are dependent. Existing methods, including Sun et al. (2023), model dependency via copulas, which may be restrictive in the presence of complex interval-censoring mechanisms. To address this gap, we propose the first shared frailty model specifically designed for left-truncated, interval-censored data, capturing heterogeneity and dependency between the failure time and observation processes. A sieve maximum likelihood approach is developed, using I-splines and M-splines to approximate the unknown baseline hazard and examination intensity functions. The asymptotic properties of the estimators are established, and an extensive simulation study demonstrates that the method provides consistent, efficient, and robust parameter estimates under a variety of scenarios. The approach is illustrated through a real data application to the AIDS cohort study, highlighting its ability to account for left truncation, interval censoring, and dependent observation process.
In clinical trials with prioritized composite outcomes, win ratios are commonly employed to evaluate the efficacy of investigational interventions. Adjusting such win ratios with respect to covariates enhances the accuracy and precision of treatment effect estimates, by controlling for baseline clinical characteristics, demographics, etc. Effects of the covariates on composite outcomes are often of complex and non-linear nature, reflecting the complexity of underlying biological mechanisms. Parametric approaches that assume an additive effect of covariates on the log-hazard often fail to capture such complexities, resulting in unreliable inferences and reduced predictive accuracy. In this article, we introduce a flexible win fraction regression framework based on B -splines, that is capable of assessing the extent and nature of the functional effects of covariates on the outcome adaptively. By leveraging the moment condition model framework equipped with the generalized method of moments (GMM) technique, we develop an efficient computational algorithm to carry out inference based on the proposed model. Using the asymptotic distribution of the estimated spline coefficients, we develop a large-sample test to assess the significance of the functional covariates. The favorable operating characteristics of the proposed methodology, compared to the current state-of-the-art, are assessed through extensive simulations. Finally, the practical utility of our proposal is demonstrated through the analysis of composite time-to-event datasets arising from cardiovascular and breast cancer clinical trials.
Assessing causal treatment effect on a time-to-event outcome and identifying important risk factors that contribute to the outcome of interest are crucial in many scientific studies. Although existing instrumental variable (IV) methods can address the endogenous treatment selection and yield an unbiased causal treatment effect estimate in the presence of censoring, the corresponding variable selection technique has not been investigated. In this paper, we propose a variable selection method for a wide class of causal semiparametric transformation models with all-or-nothing treatment compliance and right-censored data. Specifically, the minimum information criterion is embedded in the optimization step of the proposed expectation-maximization algorithm, rendering sparse estimators of the complier causal treatment effect and other regression parameters. The asymptotic properties of our method are established, including consistency and oracle property. Extensive simulation studies are conducted to evaluate the finite sample performance of the proposed method. An application to a colorectal cancer screening dataset is provided.
Analyses of recurrent hypoglycemia are critical for effective treatment management in diabetic patients. Typically, within-subject dependency in such analyses is captured through subject-level frailty. Recent research has modeled recurrent hypoglycemia using the first hitting times of a reflected Brownian motion. A close examination of this approach reveals that it does not adequately account for varying frailties among individuals, which indicate notable heterogeneity. To address this gap, we propose a finite mixture model of the first hitting time distribution of the reflected Brownian motion. This model allows for component-specific regression coefficients and frailty parameters, providing nuanced insights into how risk factors differently affect patient subgroups. We employ a Bayesian framework for inference, utilizing Markov chain Monte Carlo for estimation. Model selection is conducted using the Deviance Information Criterion and the Logarithm of the Pseudo-Marginal Likelihood. The effectiveness of these criteria is assessed through simulation studies. Application to recurrent hypoglycemia modeling revealed two subgroups with different risk profiles, as reflected in their volatilities. Bayesian model comparison criteria favor the model with component specific regression coefficients for volatilities. The subgroup with lower volatility exhibits a larger variance and, hence, a greater level of heterogeneity.
k-Sample versions of the Kolmogorov–Smirnov and Cramér–von Mises tests are proposed for data subject to left truncation and right censoring. Their asymptotic behaviour is studied and a bootstrap resampling plan is proposed to approximate the null distribution of the new tests. The performance of the tests with finite sample sizes is investigated in a simulation study, where the classical log-rank test is also considered. An illustration with a real dataset regarding unemployment times is discussed.
Clinical trials and cohort studies often aim to assess treatment effects or exposure associations in relation to the risk of one or more diseases, with death of the study participant as a competing risk. If the diseases under study are major health concerns, it may not be appropriate to assume that death acts as an independent source of right-censoring. When this occurs, a summary of treatment or exposure influences should consider disease incidence and death jointly. Here we consider some modeling approaches to doing so, starting with type-specific (cause-specific) hazard functions. We also model marginal hazard rates for disease-free survival and death, along with their dual outcome hazard functions, with emphasis on Cox models for each hazard function. Furthermore, a simple hazard ratio summary statistic is proposed for covariate effects on disease incidence and death jointly. Analyses of data from the Women's Health Initiative hormone therapy trials provide illustration.
Analyzing survival data from the National Cancer Institute’s Surveillance, Epidemiology, and End Results (SEER) Program offers invaluable insight to guide cancer management. In these data the event times are recorded in units of months. In such datasets, it can be important to consider the selection of covariates, distinguish between time-varying and time-independent effects, and capture important interaction terms. However, variable selection is difficult to combine with the differentiation between time-varying and time-independent effects. Moreover, accurately selecting interaction terms is more challenging because of the hierarchy restriction that an interaction term should only enter a model if one or both associated main effects are also included. To resolve these issues, we develop a coordinate ascent-based gradient boosting procedure for discrete failure time models. This approach achieves variable selection in high-dimensional settings, differentiates between time-varying and time-independent variables, and selects important interaction terms under hierarchy restrictions. Traditional boosting stopping criteria are ad hoc, while preferable information-criteria-based rules necessitate defining degrees of freedom, a challenge in survival models. Our approach provides well-defined degrees of freedom and offers information-criteria-based stopping rules. Simulation studies show the proposed method achieves good selection performance. We apply the proposed method to SEER melanoma cancer data and select important risk factors, distinguish between time-varying and time-dependent effects and detect important two-way interactions.
The timing of an intervention such as surgery may happen sometime after treatment inception (e.g., diagnosis date) due to being put on a wait-list for treatment or being scheduled for treatment at a later time. If interest lies in the change in risk upon receiving the intervention, a Cox model with a time-dependent exposure may be employed. However, if there exists unmeasured confounding, maximum partial likelihood estimators of the hazard ratio are biased. An instrumental variable is a cause of the exposure of interest but not of the outcome except through the exposure. We propose a new framework involving potential outcomes for each possible waiting time, and derive an estimating equation for the hazard ratio of a treatment with a time-dependent start using an instrumental variable without making any assumptions about the form of involvement of unmeasured confounders. We conducted simulations to evaluate bias of our hazard ratio estimator. We illustrate the approach using two examples: First, we analyzed a procedural registry for which instrumental variable based estimators may be the only unbiased approach. Second, we demonstrate causal hazard ratio estimation for a randomized trial with time-dependent treatment initiation.
In healthcare, personalized treatment strategies are vital for improving patient outcomes, especially under right-censored survival data. We propose Dynamic Deep Buckley-James Q-Learning, a novel counterfactual Q-learning algorithm that integrates deep learning with the Buckley-James method to simultaneously address censoring and nonlinear modeling challenges. By explicitly capturing complex, nonlinear interactions between covariates and treatment effects, the algorithm robustly estimates optimal dynamic treatment regimes. Leveraging a counterfactual framework, we define and estimate potential survival outcomes under hypothetical treatment sequences, enabling unbiased Q-function estimation even in the presence of time-dependent covariates and right censoring. The algorithm maximizes the expected imputed survival reward under these counterfactual scenarios. Simulation studies and real-world data analysis demonstrate its superior performance in predictive accuracy and treatment decision-making, offering a powerful framework for individualized care in complex clinical settings.
Two-stage randomized trials, or the more general sequential multiple assignment randomized trials (SMART), have been increasingly used in studying adaptive treatment strategies for treating chronic diseases or conditions where treatments need to be adjusted during the treatment course. Methods for analyzing two-stage randomized trials with continuous, longitudinal or right-censored survival endpoints have been developed. However, no method exists for data analysis for two-stage randomized trials with grouped survival endpoints, which are frequently encountered in practice. In this article, we propose methods for analyzing grouped survival data from two-stage randomized trials. We first extend the methods in Prentice and Gloeckler (1978) to allow for patient missing pre-specified visits, and use an efficient score test to make inferences on treatment effects with grouped data in traditional randomized trials. Based on this, we propose a weighted efficient score test to compare two adaptive treatment strategies with grouped survival data in two-stage randomized trials with missing visits. Asymptotic properties are formulated, and simulation studies are conducted to evaluate the performances of the proposed methods. We also apply the weighted score test in analyzing data from the sequenced treatment alternatives to relieve depression (STAR*D) trial.
Estimating the causal effect of a continuous treatment on survival data, particularly in cases where there is a cured fraction from observational studies, is a significant issue. However, this topic is not well addressed in the existing literature. Current methods either rely on strong parametric assumptions or struggle to effectively control for confounding variables. In this study, we propose a novel nonparametric estimation method that utilizes a weighted generalized Kaplan-Meier survival estimator. This method aims to estimate the average effects of a continuous treatment on both the probability of being cured and the restricted mean survival time. Notably, our approach does not require any parametric assumptions about the effects, and it can efficiently control for multiple confounding variables. A simulation study demonstrates that our proposed method outperforms existing approaches, particularly when the average effects are complex or when confounding is strong. We apply this method to data from a study of chlamydia patients to evaluate the average effects of years of schooling on the probability of being immune to reinfection, as well as on the restricted mean survival time.
In disease prevention research, researchers often need to assess a prevention strategy that targets key disease-associated risk factors to reduce a population’s disease burden. In this article, the fraction of the total disease burden associated with the risk factors targeted by the prevention strategy is calculated by time-varying attributable risk functions (ARF) when the disease outcome is censored time-to-event. We study some generic ARFs and develop nonparametric and semiparametric model-based procedures to estimate, compare, and predict ARFs. In addition to numerical simulation studies, we demonstrate the use of ARFs for a human immunodeficiency virus (HIV) behavior intervention trial in prevention of HIV transmissions among men who have sex with men (MSM).
In the past few decades, artificial neural networks (ANNs) have exhibited their superior capabilities for capturing nonlinear data patterns in supervised learning. Inspired by such desirable advantages, we herein explore the integration of ANNs with partially linear Cox model in the context of left truncation. This sampling scheme tends to enroll individuals with slower disease progression and thus leads to a biased sample of survival times in the target population. The proposed model comprises both parametric covariate effects and a nuisance function of uninterested covariates that is approximated with ANNs, offering a balance between interpretability and flexibility. We consider the conditional maximum likelihood estimation and derive a profile likelihood that is free of the baseline hazard function. An iterative algorithm embedded with stochastic gradient descent is proposed to minimize the negative profile log-likelihood, leading to the estimators of the regression parameters and ANNs simultaneously. Extensive simulation studies and an application demonstrate that the proposed method outperforms the traditional approaches regarding the estimation accuracy and predictive capability.
The original case-cohort design obtains detailed covariate information on a random sample of subjects from the cohort (subcohort) and on the subjects who developed the event of interest (cases). Recently, there was some work on case-cohort estimation of pure risk, i.e., the hypothetical probability that the event occurs, assuming it is the only risk. But competing events can preclude the occurrence of the event of interest, and the pure risk thus overestimates the probability of experiencing the event of interest (absolute risk). Under the cause-specific hazard Cox model, methods for case-cohort inference have been published for relative hazards and cumulative baseline hazards; we have not seen treatments of absolute risk, however. In this work we focus on absolute risk inference under the cause-specific hazard Cox model when using a sample of subjects from the cohort. We propose an influence-based variance estimation formula and consider two sampling designs: (1) a case-cohort with exhaustive sampling of subjects who developed the event of interest or a competing event; and (2) an event-stratified sample of the cohort that only includes fractions of these subjects. Our proposed variance estimate properly accounts for the sampling features and allows appropriate analysis of the sampled data. We illustrate our method and designs in simulation and on the Prostate, Lung, Colorectal and Ovarian Cancer Screening Trial. These analyses also suggest that the "robust" variance originally proposed by Barlow (Biometrics, 50:1064-1072, 1994) may be too large for the absolute risk when using a cohort subsampling design.
Interval-censored competing risks data frequently arise in medical and clinical studies among others and furthermore, the cause of failure may be missing in some situations. In this paper, we consider regression analysis of such data under the framework of an additive subdistribution hazard model and propose a two-step sieve and weighted maximum likelihood estimation procedure. The method explicitly imposes constraints on the cumulative incidence functions to ensure valid survival function estimation and adopts an augmented inverse probability weighting strategy to address the issue of missing event types. Also in the proposed approach, Bernstein polynomials are employed to approximate unknown functions and the proposed estimators are shown to be consistent and asymptotically normal. An extensive simulation study is conducted and indicates that the proposed method works well in practical situations. Finally the proposed approach is applied to the real data from a breast cancer study.
To examine the causal effects of time-varying treatments on survival, structural nested cumulative survival time models (SNCSTMs) are flexible and theoretically promising semiparametric models characterized by causally interpretable parameters. One concern is the prerequisite for uniformly scheduled data collection and complete data for time-varying confounders. For example, in pharmacoepidemiological studies using medical information databases, laboratory test results can be missing due to unscheduled hospital visits or non-compliance with health checkups. Furthermore, missing mechanisms data may be non-ignorable and non-monotone, invalidating the typical missing-data methods that assume ignorable or monotone missing mechanisms. We propose a novel g-estimation method for SNCSTMs with non-ignorable, non-monotonic missing data for time-varying confounders. We augment the g-estimation functions using missing probability and imputation models, incorporating a user-defined selection function, which allows sensitivity analyses to evaluate the departure of missing data from ignorable mechanisms. Using a proper selection function, our estimator is doubly robust in the sense that it is consistent if either model for missing probability or imputation of missing data is correct at each time point and if either model for propensity score or conditional expectation of counterfactual counting processes is correct. Moreover, applying frequentist-type multiple imputation yields a closed-form solution for calculating the estimator, even if time-varying confounders are missing. A simulation study evaluated our proposed method's finite sample performance and the estimator's double robustness. We also conducted sensitivity analyses in a pharmacoepidemiological study using a Japanese medical claims database, assessing the risk of hypoglycemia in sulfonylurea-treated patients with incomplete hemoglobin A1c values.