Clustered competing-risks data often arise in clinical studies, such as multi-center clinical trials, where the occurrence of an event within a cluster hinders the observation of other types of events. The correlation resulting from clustering can be modeled using random effects. These competing-risks data have usually been analyzed using hazard-based models, rather than survival times themselves. Hao et al. proposed a cause-specific joint accelerated failure time (AFT) random-effect modeling approach for analyzing the clustered competing-risks data, which is easy to interpret. In this article, we propose a variable selection method for fixed effects using a penalized h-likelihood (HL) procedure in the joint AFT competing-risk model. Simulation studies were conducted to evaluate the performance of the proposed variable selection procedure, which concluded that the penalized methods of SCAD and HL are more appropriate than that of LASSO. The proposed method is illustrated with two real clinical datasets.
There is a growing interest in subject-specific predictions using neural networks, as large-scale biomedical data often exhibit dependency due to high-cardinality categorical features, which have been largely overlooked by traditional neural network frameworks. This article proposes a novel hierarchical likelihood learning framework that captures both nonlinear overall effects and subject-specific effects by incorporating gamma random effects into Poisson neural networks. The global maximizer of the proposed objective function yields maximum likelihood estimators for fixed parameters and best unbiased predictors for random effects. The proposed framework provides a robust end-to-end algorithm for clustered biomedical count data, in the sense that the corresponding estimating equations remain unbiased even when the random-effects distribution is misspecified. To enhance learning efficiency, we introduce an adjustment procedure for the random effects and variance component. Extensive simulation studies and real data analyses demonstrate the practical effectiveness of the proposed method for clustered biomedical count data. The proposed method achieves competitive predictive performance in terms of mean squared Pearson error and mean deviance across various random-effects distributions and real-world datasets.
This research integrates deep learning, copula functions, and survival analysis to effectively handle highly correlated and right-censored multivariate survival data. It introduces copula-based activation functions (Clayton, Gumbel, and their combinations) to model the nonlinear dependencies inherent in such data. Through simulation studies and analysis of real breast cancer data, our proposed CNN-LSTM with copula-based activation functions for multivariate multi-types of survival responses enhances prediction accuracy by explicitly addressing right-censored data and capturing complex patterns. The model's performance is evaluated using Shewhart control charts, focusing on the average run length (ARL).
This study introduces a novel approach to modeling competing risks in survival analysis by integrating learnable Copula functions (Clayton, Frank, and Gaussian) with deep learning architectures, including Convolutional Neural Networks (CNN), Long Short-Term Memory (LSTM) networks, and a hybrid CNN-LSTM model. Here, we are interested in classifying competing risks outcomes. The proposed method captures complex dependencies within the data. Our approach demonstrates improved predictive performance in survival data modeling by effectively capturing intricate dependency structures and event relationships. We validate the proposed models using both simulated data and real-world clinical data. This research highlights the potential of integrating Copula-based dependency structures into deep learning models for survival analysis with competing risks. The results emphasize how Copula-based neural networks can enhance prediction accuracy and handle competing risks in survival analysis.
We introduce a multivariate functional principal component analysis (MFPCA) residual control chart for multivariate functional data. Our method utilizes the vine copula technique and is applied to high-frequency financial data. We employ functional eigenfunctions to uncover hidden dependence structures and explain variations in sparse multivariate longitudinal data through MFPCA. With these functional eigenfunctions, we create a vine copula-based residual control chart for sparse multivariate longitudinal data. To handle sparse multivariate longitudinal data in this context, we employ predictive mean matching imputation. As part of real-world applications, we conduct analysis on high-frequency time series data for five technology stocks listed on the Nasdaq exchange, as well as high-frequency air quality data obtained from a significantly polluted area within an Italian city.
Dependence among clustered survival data (i.e. multivariate correlated time-to-event data) can be described through copula survival hazard models or frailty models. Here, the risk prediction is based on the linear predictor which can be a strong assumption. The copula survival models have been relatively less studied as compared to frailty models. The deep neural networks (DNN) is a powerful structured neural network consisting of three layers (input, hidden and output layers) for constructing (or modelling) the functional relationship (mainly nonlinear) between input and output variables. In this paper, we propose a DNN approach for the copula models with clustered survival data. For simplicity, we use Weibull marginal hazard functions under the Clayton copula function. The proposed DNN copula model is trained by using a negative copula likelihood as a loss function. The predictive performance of the proposed method is evaluated by comparing it with existing methods in terms of C-index and integrated Brier score using simulation study and a real data set.
Deep learning (DL) includes various architectures, such as deep neural networks (DNNs) and convolutional neural networks (CNNs). DL is very powerful and flexible for non-tabular (non-structured) data (e.g. image, text). However, in tabular data, standard DNNs often do not outperform traditional machine learning (ML) methods such as tree-based models (e.g. random forest, XGBoost). CNNs carry out dimensionality reduction for non-tabular (especially image) data, but may be useful in tabular data too. In this paper, we present a unified framework of one-dimensional CNN (1D-CNN)-based approaches for various types of tabular data, which provides an end-to-end learning framework. We also propose two novel 1D-CNN-based models, i.e. a negative binomial CNN (NB-CNN) model for over-dispersed count data and a Cox-based CNN Self-Attention model for high-dimensional survival data. The predictive performance of the proposed method is evaluated by comparing it with existing ML/DL methods using four types of real tabular data, i.e. a binary response data with high dimensional features, over-dispersed count data, high-dimension survival data, and time-series data with substantial variability. The experimental results show that the proposed methods overall outperform existing ML/DL models. In particular, the NB-CNN achieves lower root mean squared error (RMSE) and higher coefficient of determination (R2) on over-dispersed count data than tree-based methods. Similarly, the Cox-based CNN Self-Attention model yields higher C-index values for high-dimensional survival tasks relative to state-of-the-art approaches.
There is a growing interest in subject-specific predictions using deep neural networks (DNNs) because real-world data often exhibit correlations, which has been typically overlooked in traditional DNN frameworks. In this paper, we propose a novel hierarchical likelihood learning framework for introducing gamma random effects into the Poisson DNN, so as to improve the prediction performance by capturing both nonlinear effects of input variables and subject-specific cluster effects. The proposed method simultaneously yields maximum likelihood estimators for fixed parameters and best unbiased predictors for random effects by optimizing a single objective function. This approach enables a fast end-to-end algorithm for handling clustered count data, which often involve high-cardinality categorical features. Furthermore, state-of-the-art network architectures can be easily implemented into the proposed h-likelihood framework. As an example, we introduce multi-head attention layer and a sparsemax function, which allows feature selection in high-dimensional settings. To enhance practical performance and learning efficiency, we present an adjustment procedure for prediction of random parameters and a method-of-moments estimator for pretraining of variance component. Various experiential studies and real data analyses confirm the advantages of our proposed methods.
We consider a parametric modelling approach for survival data where covariates are allowed to enter the model through multiple distributional parameters (i.e., scale and shape). This is in contrast with the standard convention of having a single covariate-dependent parameter, typically the scale. Taking what is referred to as a multi-parameter regression (MPR) approach to modelling has been shown to produce flexible and robust models with relatively low model complexity cost. However, it is very common to have clustered data arising from survival analysis studies, and this is something that is under developed in the MPR context. The purpose of this article is to extend MPR models to handle multivariate survival data by introducing random effects in both the scale and the shape regression components. We consider a variety of possible dependence structures for these random effects (independent, shared and correlated), and estimation proceeds using a h-likelihood approach. The performance of our estimation procedure is investigated by a way of an extensive simulation study, and the merits of our modelling approach are illustrated through applications to two real data examples, a lung cancer dataset and a bladder cancer dataset.
With the increasing popularity of big data analysis, research on zero-inflated count data with the copula method has garnered significant attention because zero-inflated count data do not follow a normal distribution and have high correlation among variables. Within the domain of quality control, there has been limited emphasis on multivariate statistical process control (SPC) techniques that specifically address the challenge of multicollinearity within regression models for multivariate zero-inflated count responses. In this paper, we explain a computational challenge in handling big data with parametric generalized linear models, such as the zero-inflated Poisson model. This challenge motivates us to introduce a copula-based deep learning and neural network model involving multivariate zero-inflated count response variables with highly correlated explanatory variables. This approach builds upon existing deep learning and neural network models designed for univariate count responses, expanding them to encompass multivariate count response models through the integration of a copula regression framework. To evaluate the performance of our proposed methodology, we conduct a comparative analysis of accuracies using the zero-inflated Poisson model, univariate deep learning and neural network models, multivariate count response deep learning and neural network models, and our proposed copula deep learning and neural network models. This assessment involves employing metrics such as root mean square error (RMSE), weighted mean absolute percentage error (WMAPE), and mean absolute deviation (MAD). We include both copula-based asymmetrical zero-inflated simulated data and real-world data. We also propose a temporal dependence control chart for assessing the temporal dependence between bivariate zero-inflated count response variables. Our proposed copula deep learning and neural network temporal control charts can check time-varying dependence and outliers by leveraging the Shewhart statistical process control chart and t-copula ARMA-GARCH dynamic control correlation.
Accelerated failure time (AFT) models with random effects, a useful alternative to frailty models, have been widely used for analyzing clustered (or correlated) time-to-event data. In the AFT model, the distribution of the unobserved random effect is conventionally assumed to be parametric, often modeled as a normal distribution. Although it has been known that a misspecfied random-effect distribution has little effect on regression parameter estimates, in some cases, the impact caused by such misspecification is not negligible. Particularly when our focus extends to quantities associated with random effects, the problem could become worse. In this paper, we propose a semi-parametric maximum likelihood approach in which the random-effect distribution under the AFT models is left unspecified. We provide a feasible algorithm to estimate the random-effect distribution as well as model parameters. Through comprehensive simulation studies, our results demonstrate the effectiveness of this proposed method across a range of random-effect distribution types (discrete or continuous) and under conditions of heavy censoring. The efficacy of the approach is further illustrated through simulation studies and real-world data examples.
The deep neural network (DNN) model can be viewed as a highly non-linear and semi-parametric generalization of statistical regression models such as the generalized linear model (GLM). The fitting (i.e. learning) of DNN models using training data is usually implemented by minimizing a squared loss function, which is equivalent to a Gaussian likelihood. We present how to understand and fit the DNN models via the GLM framework, with simulated and real data analyses, which are useful for examining the behavior of the GLM-based DNN. Furthermore, we extend the GLM-based DNN to Cox’s proportional hazards models with censored survival data, including an numerical study. The Appendix provides Python codes for a TensorFlow-Keras implementation so that a GLM-based DNN becomes directly accessible to interested readers, including the codes of Cox-based DNN.
The proportional hazards model and accelerated failure time model are two important classes for analyzing survival data. However, the two models cannot capture time-scaled effects such as crossing hazard or the gradual effect of treatment over time. For the purpose of this detection, accelerated hazards model has been studied. In this paper, we propose the use of log-logistic distribution for the baseline hazard function of the accelerated hazards model which allows for time-dependent group effect. In particular, the log-logistic distribution gives various forms (e.g. non-monotone) of hazard function and also an easy implementation for likelihood-based model fitting due to an explicit form of hazard and survival functions. The likelihood-based estimation procedures of the accelerated hazards models are derived. Our method is demonstrated with two practical data sets, via the procedures of systematic model checking.
Recently, deep learning has become a pervasive tool in prediction problems for structured and/or unstructured big data in various areas including science and engineering. In particular, deep neural network models (i.e. a basic core model of deep learning) can be viewed as an extension of statistical models by going through the incorporation of hidden layers. In this paper, we study the relationship between both models in terms of model structures and model learning. For this purpose, we also compare the predictive performances of both models, with two practical examples.
Competing risks data arise when occurrence of an event hinders observation of other types of events, and they are encountered in various research areas including biomedical research. These data have been usually analyzed using the hazard-based models, not survival times themselves. In this paper, we propose a joint accelerated failure time (AFT) modeling approach to model clustered competing risks data. Times to competing events are assumed to be log-linear with normal errors and correlated through a scaled random effect that follows a zero-mean normal distribution. Inference on the model parameters is based on the h-likelihood. Performance of the proposed method is evaluated through extensive simulation studies. The simulation results show that the estimated regression parameters are robust against the violation of the assumed parametric distributions. The proposed method is illustrated with three real competing risks data sets.
In a semi-competing risks model in which a terminal event censors a non-terminal event but not vice versa, the conventional method can predict clinical outcomes by maximizing likelihood estimation. However, this method can produce unreliable or biased estimators when the number of events in the datasets is small. Specifically, parameter estimates may converge to infinity, or their standard errors can be very large. Moreover, terminal and non-terminal event times may be correlated, which can account for the frailty term. Here, we adapt the penalized likelihood with Firth’s correction method for gamma frailty models with semi-competing risks data to reduce the bias caused by rare events. The proposed method is evaluated in terms of relative bias, mean squared error, standard error, and standard deviation compared to the conventional methods through simulation studies. The results of the proposed method are stable and robust even when data contain only a few events with the misspecification of the baseline hazard function. We also illustrate a real example with a multi-centre, patient-based cohort study to identify risk factors for chronic kidney disease progression or adverse clinical outcomes. This study will provide a better understanding of semi-competing risk data in which the number of specific diseases or events of interest is rare.
Jonghun Park合作论文数Seoul National University, Seoul, South Korea2