This paper centers on estimating the turning point (or mode) of the hazard function in survival studies, particularly in health-related research. The hazard function represents the instantaneous risk of an event, like death, and it often follows a unimodal pattern in these contexts. Identifying the turning point is important because it reveals when the risk is highest, which helps in developing more targeted treatment plans and determining the best times for interventions. Additionally, considering that some patients may be cured, estimating cure rates provides a fuller picture of survival outcomes. To tackle these issues, the paper introduces a cure rate regression model based on the Dagum distribution. A key feature of this model is that it reparameterizes the Dagum distribution to include the mode of the hazard function, which is linked to covariates through a logarithmic function. For estimating the proportion of cured patients, a logistic regression model is used to capture the effects of different covariates. The parameters of the model are estimated using the maximum likelihood method, and its performance is tested through extensive Monte Carlo simulations. The paper also illustrates the practical benefits of the model by applying it to a dataset on COVID-19 in maternal populations.
Survival models with cure fractions, known as long-term survival models, are widely used in epidemiology to account for both immune and susceptible patients regarding a failure event. In such studies, it is also necessary to estimate unobservable heterogeneity caused by unmeasured prognostic factors. Moreover, the hazard function may exhibit a non-monotonic shape, specifically, an unimodal hazard function. In this article, we propose a long-term survival model based on a defective version of the Dagum distribution, incorporating a power variance function frailty term to account for unobservable heterogeneity. This model accommodates survival data with cure fractions and non-monotonic hazard functions. The distribution is reparameterized in terms of the cure fraction, with covariates linked via a logit link, allowing for direct interpretation of covariate effects on the cure fraction-an uncommon feature in defective approaches. We present maximum likelihood estimation for model parameters, assess performance through Monte Carlo simulations, and illustrate the model's applicability using two health-related datasets: severe COVID-19 in pregnant and postpartum women and patients with malignant skin neoplasms.
This paper presents a parametric quantile regression model for survival data that incorporates a cure fraction, addressing limitations of traditional survival models related to clinical interpretability and their limited capacity to account for cured individuals. The proposed model is built upon the exponentiated Weibull distribution and employs a logarithmic link between survival quantiles and covariates, offering a flexible and interpretable framework. Parameter estimation is conducted via an expectation-maximization algorithm within the maximum likelihood framework. Monte Carlo simulation studies assess the model's performance under varying censoring levels and sample sizes, confirming its ability to yield stable estimates and reliable inference across diverse scenarios. An application to gastric cancer data demonstrates the model's effectiveness in capturing heterogeneity and uncovering clinically relevant patterns. This methodology advances survival analysis by integrating cure fraction modeling with quantile-based inference in a coherent and robust manner.
In this article, we particularly address the problem of assessing the impact of clinical stage and age on the specific survival times of men with breast cancer when cure is a possibility, where there is also the interest of explaining this impact on different quantiles of the survival times. To this end, we developed a quantile regression model for survival data in the presence of long-term survivors based on the generalized distribution of Gompertz in a defective version, which is conveniently reparametrized in terms of the q-th quantile and then linked to covariates via a logarithm link function. This proposal allows us to obtain how each variable affects the survival times in different quantiles. In addition, we are able to study the effects of covariates on the cure rate as well. We consider Markov Chain Monte Carlo (MCMC) methods to develop a Bayesian analysis in the proposed model and we evaluate its performance through a Monte Carlo simulation study. Finally, we illustrate the advantages of our model in a data set about male breast cancer from Brazil.
Survival models incorporating cure fractions, commonly known as cure fraction models or long-term survival models, are widely employed in epidemiological studies to account for both immune and susceptible patients in relation to the failure event of interest under investigation. In such studies, there is also a need to estimate the unobservable heterogeneity caused by prognostic factors that cannot be observed. Moreover, the hazard function may exhibit a non-monotonic form, specifically, an unimodal hazard function. In this article, we propose a long-term survival model based on the defective version of the Dagum distribution, with a power variance function (PVF) frailty term introduced in the hazard function to control for unobservable heterogeneity in patient populations, which is useful for accommodating survival data in the presence of a cure fraction and with a non-monotone hazard function. The distribution is conveniently reparameterized in terms of the cure fraction, and then associated with the covariates via a logit link function, enabling direct interpretation of the covariate effects on the cure fraction, which is not usual in the defective approach. It is also proven a result that generates defective models induced by PVF frailty distribution. We discuss maximum likelihood estimation for model parameters and evaluate its performance through Monte Carlo simulation studies. Finally, the practicality and benefits of our model are demonstrated through two health-related datasets, focusing on severe cases of COVID-19 in pregnant and postpartum women and on patients with malignant skin neoplasms.
A flexible parametric mixture cure model, called bi-lognormal cure rate model or simply BLN model, is defined and studied. The BLN model can be effectively used to analyze survival dataset in the presence of long-term survivors, especially when the dataset presents the underlying phenomenon of latent competing risks or when there is evidence that a bimodal hazard function is appropriated to described it, which are advantages over other cure rate models found in the literature. We discuss the maximum likelihood estimation for the model parameters considering interval-censored data through the differential evolution algorithm that is a nature-inspired computing metaheuristic used for global optimization of functions defined in multidimensional spaces. This approach is also used because the likelihood function of the model is multimodal and the direct application of gradient methods in this case is not ideal, since such methods are local search methods with a high chance of getting stuck at a local maximum when the starting point is chosen outside the basin of attraction of a global maximum. In addition, a simulation study was implemented to compare the performance of differential evolution algorithm with the performance of the Newton-Raphson algorithm in terms of bias, root mean square error, and the coverage probability of the asymptotic confidence intervals for the parameters. Finally, an application of the BLN model to real data is presented to illustrate that it can provide a better fit than other mixture cure rate models.
It has come to our attention that the results described in Remarks 1 and 2 do not hold, since the joint distribution given by in Eq. (2) (Shmueli et al., 2005) is not marginally compatible, that is, if the Bernoulli variables Z(i) (i = 1,..., m) have a joint distribution given by Eq. (2), then the joint distribution of (Z(1),..., Z(m-1)) is not of the same form, with m replaced by (m - 1), so that Sigma(m)(i=1) Z(i) does not have its distribution in the same form as Sigma(m)(i=1) Z(i) (see Kadane, 2016). The authors thank Dr. Christian Wei ss for alerting us about this point with respect to Remarks 1 and 2. (c) 2014 Elsevier B.V. All rights reserved.
It has come to our attention that the results described in Remarks 1 and 2 do not hold, since the joint distribution given by in Eq. (2) (Shmueli et al., 2005) is not marginally compatible, that is, if the Bernoulli variables Zi (i=1,…,m) have a joint distribution given by Eq. (2), then the joint distribution of (Z1,…,Zm−1) is not of the same form, with m replaced by (m−1), so that ∑i=1m−1Zi does not have its distribution in the same form as ∑i=1mZi (see Kadane, 2016). The authors thank Dr. Christian Weiß for alerting us about this point with respect to Remarks 1 and 2.
The hazard function plays an important role in cancer patient survival studies, as it quantifies the instantaneous risk of death of a patient at any given time. Often in cancer clinical trials, unimodal hazard functions are observed, and it is of interest to detect (estimate) the turning point (mode) of hazard function, as this may be an important measure in patient treatment strategies with cancer. Moreover, when patient cure is a possibility, estimating cure rates at different stages of cancer, in addition to their proportions, may provide a better summary of the effects of stages on survival rates. Therefore, the main objective of this paper is to consider the problem of estimating the mode of hazard function of patients at different stages of cervical cancer in the presence of long-term survivors. To this end, a mixture cure rate model is proposed using the log-logistic distribution. The model is conveniently parameterized through the mode of the hazard function, in which cancer stages can affect both the cured fraction and the mode. In addition, we discuss aspects of model inference through the maximum likelihood estimation method. A Monte Carlo simulation study assesses the coverage probability of asymptotic confidence intervals.
The log-linear Poisson model, characterized by linear variance function and a logarithmic relation between means and covariates, embedded in the exponential family regression framework provided by generalized linear models (GLM) is still the standard approach for analyzing count data responses with regression models. In practice, however, count data are often overdispersed and, thus, not conducive to Poisson regression. Therefore, the main goal of this article is to introduce a log-linear model based on the P[Formula: see text]lya–Aeppli (PA) distribution, which is an extension of the Poisson distribution by including a dispersion parameter ρ, to address the problem of overdispersion. Maximum likelihood (ML) estimation procedure is discussed as well as a test for determining the need for a PA regression over a standard Poisson regression. In addition, a simple EM-type algorithm for iteratively computing ML estimates is presented. In order to study departures from the error assumption as well as the presence of outliers, we perform residual analysis based on the standardized Pearson residuals. Furthermore, for different parameter settings and sample sizes, various simulations are performed. Finally, we also illustrated the new method on three real datasets, two of them are from biological researches and the other is from a violence study.
In this paper, we introduce a new non-negative integer-valued autoregressive time series model based on a new thinning operator, so called generalized zero-modified geometric (GZMG) thinning operator. The first part of the paper is devoted to the distribution, GZMG distribution, which is obtained as the convolution of the zero-modified geometric (ZMG) distributed random variables. Some properties of this distribution are derived. Then, we construct a thinning operator based on the counting processes with ZMG distribution. Finally, an INAR(1) time series model is introduced and its properties including estimation issues are derived and discussed. A small Monte Carlo experiment is conducted to evaluate the performance of maximum likelihood estimators in finite samples. At the end of the paper, we consider an empirical illustration of the introduced INAR(1) model.
In this paper we develop a regression model for survival data in the presence of long-term survivors based on the generalized Gompertz distribution introduced by El-Gohary et al. [The generalized Gompertz distribution. Appl Math Model. 2013;37:13-24] in a defective version. This model includes as special case the Gompertz cure rate model proposed by Gieser et al. [Modelling cure rates using the Gompertz model with covariate information. Stat Med. 1998;17:831-839]. Next, an expectation maximization algorithm is then developed for determining the maximum likelihood estimates (MLEs) of the parameters of the model. In addition, we discuss the construction of confidence intervals for the parameters using the asymptotic distributions of the MLEs and the parametric bootstrap method, and assess their performance through a Monte Carlo simulation study. Finally, the proposed methodology was applied to a database on uterine cervical cancer.
SummaryIn this paper we propose a new stationary first‐order non‐negative integer valued autoregressive process with geometric marginals based on a generalised version of the negative binomial thinning operator. In this manner we obtain another process that we refer to as a generalised stationary integer‐valued autoregressive process of the first order with geometric marginals. This new process will enable one to tackle the problem of overdispersion inherent in the analysis of integer‐valued time series data, and contains the new geometric process as a particular case. In addition various properties of the new process, such as conditional distribution, autocorrelation structure and innovation structure, are derived. We discuss conditional maximum likelihood estimation of the model parameters. We evaluate the performance of the conditional maximum likelihood estimators by a Monte Carlo study. The proposed process is fitted to time series of number of weekly sales (economics) and weekly number of syphilis cases (medicine) illustrating its capabilities in challenging cases of highly overdispersed count data.
In this paper, we propose a new stationary first-order non-negative integer valued autoregressive [INAR(1)] process with geometric marginals based on a modified version of the binomial thinning operator. This new process will enable one to tackle the problem of overdispersion inherent in the analysis of integer-valued time series data that may arise due to the presence of some correlation between underlying events, heterogeneity of the population, excess to zeros, among others. In addition, it includes as special cases the geometric INAR(1) [GINAR(1)] (Alzaid and Al-Osh, 1988) and new geometric [NGINAR(1)] (Ristic et al., 2009) processes, making it be very useful in discriminating between nested models. The innovation structure of the new process is very simple. The main properties of the process are derived, such as conditional distribution, autocorrelation structure, innovation structure and jumps. The method of conditional maximum likelihood is used for estimating the process parameters. Some numerical results of the estimators are presented with a brief discussion. In order to illustrate the potential for practice of our process we apply it to a real data set. (C) 2016 Elsevier B.V. All rights reserved.
Shmueli et al. (2005) introduced the COM-Poisson-binomial distribution, but they did not study the mathematical properties of this family of distributions. In this paper, we discuss some properties and an asymptotic approximation of it by the COM-Poisson distribution. Moreover, three datasets are also considered. (C) 2014 Elsevier BM. All rights reserved.
In this article, we present a simple generalization of the Bernoulli trials model to a Markov chain with an additional parameter that measures dependence. We then formulate a Markov correlated Poisson process which, due to its flexibility, has great potential for analyzing many practical processes including those for long-term survival analysis.
In this paper, we introduce the complementary exponential power series distributions, with failure rate increasing, which is complementary to the exponential power series model proposed by Chahkandi and Ganjali [Comput. Statist. Data Anal. 53 (2009) 4433-4440]. The new class of distribution arises on latent complementary risks scenarios, where the lifetime associated with a particular risk is not observable, rather we observe only the maximum lifetime value among all risks. This new class contains several distributions as a particular case. The properties of the proposed distribution class are discussed, such as quantiles, moments and order statistics. Estimation is carried out via maximum likelihood. Simulation results on maximum likelihood estimation are presented. A real application illustrates the usefulness of the new distribution class.