In this article we consider an extension of the penalized splines approach in the context of censored semiparametric modelling using Kaplan-Meier weights to take into account the effect of censorship. We proposed an estimation method and develop statistical inferences in the model. Using various simulation studies we show that the performance of the method is quite satisfactory. A real data set is used to illustrate that the proposed method is comparable to parametric approaches when assuming a probability distribution of the response variable and/or the functional form. However, our proposal does not need these assumptions since it avoids model specification problems.
The combination of P-splines and Kaplan–Meier weights provide a flexible approach to nonparametric modelling in the context of censored data. To apply this methodology, it is necessary to choose the smoothing parameter and the number and location of the knots. In this paper, we propose a new criterion for choosing the smoothing parameter adapted to the case of uncensored data. In addition, alternatives to the methods used in the literature on uncensored data are proposed for choosing the location and number of knots. Using a simulation study we analyse the effectiveness of the various alternatives proposed in situations with differences in the information available and show that their performance is quite satisfactory. A real dataset from Mayo Clinic Primary Biliary Cirrhosis data is also used to illustrate the methodology proposed. Finally, we offer some guidelines to help the user choose the parameters in the practical application of the methodology.
In this paper, we consider the problem of nonparametric curve fitting in the specific context of censored data. We propose an extension of the penalized splines approach using Kaplan-Meier weights to take into account the effect of censorship and generalized cross-validation techniques to choose the smoothing parameter adapted to the case of censored samples. Using various simulation studies, we analyze the effectiveness of the censored penalized splines method proposed and show that the performance is quite satisfactory. We have extended this proposal to a generalized additive models (GAM) framework introducing a correction of the censorship effect, thus enabling more complex models to be estimated immediately. A real dataset from Stanford Heart Transplant data is also used to illustrate the methodology proposed, which is shown to be a good alternative when the probability distribution for the response variable and the functional form are not known in censored regression models.
The Local Whittle estimator is one of the most popular techniques for estimating the memory parameter in long memory series due to its simple implementation and nice asymptotic properties under mild conditions. However, its empirical performance depends heavily on the bandwidth, that is the band of frequencies used in the estimation. Different choices may lead to different conclusions about, for example, the stationarity of the series or its mean reversion. Optimal bandwidth selection is thus of crucial importance for accurate estimation of the memory parameter, but few strategies for assuring this have been proposed to date, and their results in applied contexts are poor. A new strategy based on minimising a bootstrap approximation of the mean square error is proposed here and its performance is shown to be convincing in an extensive Monte Carlo analysis and in applications to real series.
The asymptotic properties of the Local Whittle estimator of the memory parameter d have been widely analysed and its consistency and asymptotic distribution have been obtained for values of d∈(−1/2,1] in a wide range of situations. However, the asymptotic distribution may be a poor approximation of the exact one in several cases, e.g. with small sample sizes or even with larger samples when d>0.75. In other situations the asymptotic distribution is unknown, as for example in a noninvertible context or in some nonlinear transformations of long memory processes, where only consistency is obtained. For all these cases a bootstrap strategy based on resampling a (perhaps locally) standardised periodogram is proposed. A Monte Carlo analysis shows that this strategy leads to a good approximation of the exact distribution of the Local Whittle estimator in those situations where the asymptotic distribution is not reliable.
Stute (1993, Consistent estimation under random censorship when covariables are present. Journal of Multivariate Analysis 45, 89-103) proposed a new method to estimate regression models with a censored response variable using least squares and showed the consistency and asymptotic normality for his estimator. This article proposes a new bootstrap-based methodology that improves the performance of the asymptotic interval estimation for the small sample size case. Therefore, we compare the behavior of Stute's asymptotic confidence interval with that of several confidence intervals that are based on resampling bootstrap techniques. In order to build these confidence intervals, we propose a new bootstrap resampling method that has been adapted for the case of censored regression models. We use simulations to study the improvement the performance of the proposed bootstrap-based confidence intervals show when compared to the asymptotic proposal. Simulation results indicate that, for the new proposals, coverage percentages are closer to the nominal values and, in addition, intervals are narrower.
This paper proposes a bias reduction of the coefficients' estimator for linear regression models when observations are randomly censored and the error distribution is unknown. The proposed bias correction is applied to the weighted least squares estimator proposed by Stute [28] [W. Stute: Consistent estimation under random censorship when covariables are present. J. Multivariate Anal. 45 (1993), 89-103.], and it is based on model-based bootstrap resampling techniques that also allow us to work with censored data. Our bias-corrected estimator proposal is evaluated and its behavior assessed in simulation studies concluding that both the bias and the mean square error are reduced with the new proposal.
We propose the study of the relevant factors for the survival of Russian commercial banks during the transition period. The accelerated development of the Russian commercial banking industry after the banking reform caused high rates of entry followed by a period of high rates of exit. As a consequence, many banks had to exit the market without refunding their deposits. Therefore, both for the banks and for the banks' depositors, it is of interest to identify the relevant factors that motivate the exit or the closing of the bank. We propose a different methodology based on penalized weighted least squares that represents a very general, flexible and innovative approach for this type of analysis. That is, the proposed methodology does not require the assumptions of a probability distribution for the variable under study or of any covariate parametric functional form for the covariate whose effect on the survival variable we wish to address. Copyright © 2010 John Wiley & Sons, Ltd.
In this work we study the effect of several covariates X on a censored response variable T with unknown probability distribution. In this context, most of the studies in the literature can be located in two possible general classes of regression models: models that study the effect the covariates have on the hazard function; and models that study the effect the covariates have on the censored response variable. Proposals in this paper are in the second class of models and, more specifically, on least squares based model approach. Thus, using the bootstrap estimate of the bias, we try to improve the estimation of the regression parameters by reducing their bias, for small sample sizes. Simulation results presented in the paper show that, for reasonable sample sizes and censoring levels, the bias is always smaller for the new proposals. Keywords—censored response variable, regression, bias.
The choice of the bandwidth in the local log-periodogram regression is of crucial importance for estimation of the memory parameter of a long memory time series. Different choices may give rise to completely different estimates, which may lead to contradictory conclusions, for example about the stationarity of the series. We propose here a data-driven bandwidth selection strategy that is based on minimizing a bootstrap approximation of the mean-squared error (MSE). Its behaviour is compared with other existing techniques for optimal bandwidth selection in a MSE sense, revealing its better performance in a wider class of models. The empirical applicability of the proposed strategy is shown with two examples: the widely analysed in a long memory context Nile river annual minimum levels and the input gas rate series of Box and Jenkins.
The log periodogram regression is widely used in empirical applications because of its simplicity to estimate the memory parameter, d, its good asymp- totic properties and its robustness to misspeciflcation of the short term behavior of the series. However, the asymptotic distribution is a poor approximation of the (unknown) flnite sample distribution if the sam- ple size is small. Here the flnite sample performance of difierent nonparametric residual bootstrap proce- dures is analyzed when applied to construct confl- dence intervals. In particular, in addition to the basic residual bootstrap the local bootstrap that might ad- equately replicate the structure that may arise in the errors of the regression is considered when the series shows weak dependence in addition to the long mem- ory component. Bias correcting bootstrap to adjust the bias caused by that structure is also considered.
—The log periodogram regression is widely used in em- pirical applications because of its simplicity, since only a least squares regression is required to estimate the memory parameter, d , its good asymptotic properties and its robustness to misspecification of the short term behavior of the series. However, the asymptotic distribution is a poor approximation of the (unknown) finite sample distribution if the sample size is small. Here the finite sample performance of dif- ferent nonparametric residual bootstrap procedures is analyzed when applied to construct confidence intervals. In particular, in addition to the basic residual bootstrap, the local and block bootstrap that might adequately replicate the structure that may arise in the errors of the regression are considered when the series shows weak dependence in addition to the long memory component. Bias correcting bootstrap to adjust the bias caused by that structure is also considered. Finally, the performance of the bootstrap in log periodogram regression based confidence intervals is assessed in different type of models and how its performance changes as sample size increases.
The choice of the bandwidth in the local log-periodogram regression is of crucial importance for estimation of the memory parameter of a long memory time series. Different choices may give rise to completely different estimates, which may lead to contradictory conclusions, for example about the stationarity of the series. We propose here a data driven bandwidth selection strategy that is based on minimizing a bootstrap approximation of the mean squared error and compare its performance with other existing techniques for optimal bandwidth selection in a mean squared error sense, revealing its better performance in a wider class of models. The empirical applicability of the proposed strategy is shown with two examples: the widely analyzed in a long memory context Nile river annual minimum levels and the input gas rate series of Box and Jenkins.
The problem of lifetime data in which censored observations are present is considered. In addition, it introduces the different characteristics that censored data have, together with the different scenarios that would lead to the application of the appropriate statistical approaches. The analysis of these scenarios will be mainly centered on the knowledge of the distribution for the survival time and the functional relationship between the survival time and the different covariates available in heterogeneous populations. The proposals are applied to a real data set where the survival time of AIDS-diagnosed patients in the Basque Country (Spain) is studied.
Abstract. The relevance of statistical time to event analysis in the social sciences has proved to be of great importance in the last few years, especially in applications related to labor-market analysis, employment and/or unemployment issues, duration of strikes, and survival of new firms, and in financial applications related to the time a company spends in a given status, for example, bankruptcy. We review some of the techniques that have proved to be adequate for analyzing this type of data and the conditions they require for their proper use. In addition, we extend these techniques in order to be able to analyze specific and more complex situations by using a more general and flexible model. All of these techniques and their extensions are illustrated with an example that studies the duration of firms under bankruptcy in the United States.