In cancer clinical trials, health-related quality of life (HRQoL) is an important endpoint, providing information about patients' well-being and daily functioning. However, missing data due to premature dropout can lead to biased estimates, especially when dropouts are informative. This paper introduces the extJMIRT approach, a novel tool that efficiently analyzes multiple longitudinal ordinal categorical data while addressing informative dropout. Within a joint modeling framework, this approach connects a latent variable, derived from HRQoL data, to cause-specific hazards of dropout. Unlike traditional joint models, which treat longitudinal data as a covariate in the survival submodel, our approach prioritizes the longitudinal data and incorporates the log baseline dropout risks as covariates in the latent process. This leads to a more accurate analysis of longitudinal data, accounting for potential effects of dropout risks. Through extensive simulation studies, we demonstrate that extJMIRT provides robust and unbiased parameter estimates and highlight the importance of accounting for informative dropout. We also apply this methodology to HRQoL data from patients with progressive glioblastoma, showcasing its practical utility.
Health-related quality-of-life (HRQoL) outcomes are increasingly incorporated into oncology research to complement traditional survival endpoints by capturing patients' well-being over time. These outcomes are typically collected through multidimensional questionnaires yielding longitudinal ordinal data, and are often subject to dropout due to disease progression or death. In this context, joint models provide a well-established framework to account for the dependence between longitudinal HRQoL trajectories and time-to-event outcomes, but fully joint estimation rapidly becomes computationally prohibitive when multiple latent dimensions and random effects are involved. We propose a novel slope-corrected two-stage (SC2S) approach for the joint analysis of multivariate ordinal HRQoL data and survival outcomes within a multidimensional latent trait framework. The proposed approach propagates longitudinal information to the survival model through informative priors on the random effects, while additionally re-estimating longitudinal slope parameters. This strategy substantially reduces bias in both longitudinal and survival submodels while preserving much of the computational efficiency of two-stage procedures. Through simulation studies and an application to HRQoL data from patients with progressive glioblastoma, we show that the proposed method closely approximates fully joint Bayesian estimation while requiring notably less computation time.
Extended cure survival models enable to separate covariates that affect the long-term probability of an event (or long-term survival) from those only affecting the dynamics of events (or short-term survival). We propose to generalize the bounded cumulative hazard model to handle exogenous covariates frequently changing values over time and jointly impacting long- and short-term survival. The selection of the penalty parameters tuning the smoothness of additive terms is a challenge in that framework. A fast algorithm based on Laplace approximations in Bayesian P-spline models is proposed. The methodology is motivated by fertility studies where women’s characteristics such as the employment status and the income (to cite a few) can vary in a non-trivial and frequent way during the individual follow-up. The method is furthermore illustrated by drawing on register data from the German Pension Fund which enabled us to study how women’s time-varying earnings relate to first birth transitions.
We consider the class of Erlang mixtures for the task of density estimation on the positive real line when the only available information is given as local moments, a histogram with potentially higher order moments in some bins. By construction, the obtained moment problem is ill-posed and requires regularization. Several penalties can be used for such a task, such as a lasso penalty for sparsity of the representation, but we focus here on a simplified roughness penalty from the P-splines literature. We show that the corresponding hyperparameter can be selected without cross-validation through the computation of the so-called effective dimension of the estimator, which makes the estimator practical and adapted to these summarized information settings. The flexibility of the local moments representations allows interesting additions such as the enforcement of Value-at-Risk and Tail Value-at-Risk constraints on the resulting estimator, making the procedure suitable for the estimation of heavy-tailed densities.
In medical studies, repeated measurements of biomarkers and time-to-event data are often collected during the follow-up period. To assess the association between these two outcomes, joint models are frequently considered. The most common approach uses a linear mixed model for the longitudinal part and a proportional hazard model for the survival part. The latter assumes a linear relationship between the survival covariates and the log hazard. In this work, we propose an extension allowing the inclusion of nonlinear covariate effects in the survival model using Bayesian penalized B-splines. Our model is valid for non-Gaussian longitudinal responses since we use a generalized linear mixed model for the longitudinal process. A simulation study shows that our method gives good statistical performance and highlights the importance of taking into account the possible nonlinear effects of certain survival covariates. Data from patients with a first progression of glioblastoma are analysed to illustrate the method.
Extended cure survival models enable to separate covariates that affect the probability of an event (or `long-term' survival) from those only affecting the event timing (or `short-term' survival). We propose to generalize the bounded cumulative hazard model to handle additive terms for time-varying (exogenous) covariates jointly impacting long- and short-term survival. The selection of the penalty parameters is a challenge in that framework. A fast algorithm based on Laplace approximations in Bayesian P-spline models is proposed. The methodology is motivated by fertility studies where women's characteristics such as the employment status and the income (to cite a few) can vary in a non-trivial and frequent way during the individual follow-up. The method is furthermore illustrated by drawing on register data from the German Pension Fund which enabled us to study how women's time-varying earnings relate to first birth transitions.
Laplace P-splines (LPS) combine the P-splines smoother and the Laplace approximation in a unifying framework for fast and flexible inference under the Bayesian paradigm. The Gaussian Markov random field prior assumed for penalized parameters and the Bernstein-von Mises theorem typically ensure a razor-sharp accuracy of the Laplace approximation to the posterior distribution of these quantities. This accuracy can be seriously compromised for some unpenalized parameters, especially when the information synthesized by the prior and the likelihood is sparse. Therefore, we propose a refined version of the LPS methodology by splitting the parameter space in two subsets. The first set involves parameters for which the joint posterior distribution is approached from a non-Gaussian perspective with an approximation scheme tailored to capture asymmetric patterns, while the posterior distribution for the penalized parameters in the complementary set undergoes the LPS treatment with Laplace approximations. As such, the dichotomization of the parameter space provides the necessary structure for a separate treatment of model parameters, yielding improved estimation accuracy as compared to a setting where posterior quantities are uniformly handled with Laplace. In addition, the proposed enriched version of LPS remains entirely sampling-free, so that it operates at a computing speed that is far from reach to any existing Markov chain Monte Carlo approach. The methodology is illustrated on the additive proportional odds model with an application on ordinal survey data.
Data on a continuous variable are often summarized by means of histograms or displayed in tabular format: the range of data is partitioned into consecutive interval classes and the number of observations falling within each class is provided to the analyst. Computations can then be carried in a nonparametric way by assuming a uniform distribution of the variable within each partitioning class, by concentrating all the observed values in the center, or by spreading them to the extremities. Smoothing methods can also be applied to estimate the underlying density or a parametric model can be fitted to these grouped data. For insurance loss data, some additional information is often provided about the observed values contained in each class, typically class-specific sample moments such as the mean, the variance or even the skewness and the kurtosis. The question is then how to include this additional information in the estimation procedure. The present paper proposes a method for performing density and quantile estimation based on such augmented information with an illustration on car insurance data.
Building on a thick strand of the literature on the determinants of higher-order births, this study uses a gender and class perspective to analyse second birth progression rates in Germany. Using data from the German Socio-Economic Panel from 1990 to 2020, individuals are classified based on their occupation into: upper service, lower service, skilled manual/higher-grade routine nonmanual, and semi-/unskilled manual/lower-grade routine nonmanual classes. Results highlight the "economic advantage" of men and women in service classes who experience strongly elevated second birth rates. Finally, we demonstrate that upward career mobility post-first birth is associated with higher second birth rates, particularly among men.
Aerosol delivery equipment, suitable for the treatment of bovine respiratory dysfunctions and including two parallelly positioned jet nebulizers, was studied in depth in order to determine the optimal working conditions in the field. Indeed, some factors might reasonably alter the performance of this equipment. Among those factors, the influence of the parallel position of the jet nebulizers (in order to accomodate the breathing requirements of the cattle and to achieve a quick treatment), of the long feed pipe delivering compressed air (in order to keep the animal away from the compressor unit), and finally of the ambient temperature were studied, this equipment being essentially used during the winter season. It has been shown that this equipment is able to accomodate the breathing needs of bovine weighing up to 225 kg if a pressure of 600 kPa is developed upstream to the nebulizers. That pressure provides the highest proportion of particles with a diameter less than 5 µm. The rate of atomization is significantly reduced when working at ambient air temperatures (272.25 K < T < 274.65 K) close to those encountered in winter. This is especially true when pressure upstream to the nebulizers does not reach 500 kPa. Finally, the residual volume is significantly decreased when the feed pipe for compressed air is immersed in hot water and the rate of atomization significantly increased when the face mask including the nebulizers is maintained so that the nebulizers are in a vertical position or at an angle not less than 60° with respect to the ground.
Penalized B-splines are routinely used in additive models to describe smooth changes in a response with quantitative covariates. It is typically done through the conditional mean in the exponential family using generalized additive models with an indirect impact on other conditional moments. Another common strategy consists in focussing on several low-order conditional moments, leaving the complete conditional distribution unspecified. Alternatively, a multi-parameter distribution could be assumed for the response with several of its parameters jointly regressed on covariates using additive expressions. Our work can be connected to the latter proposal for a right- or interval-censored continuous response with a highly flexible and smooth nonparametric density. We focus on location-scale models with additive terms in the conditional mean and standard deviation. Starting from recent results in the Bayesian framework, we propose a quickly converging algorithm to select penalty parameters from their marginal posteriors. It relies on Laplace approximations to the conditional posterior of the spline parameters. Simulations suggest that the so-obtained estimators own excellent frequentist properties and increase efficiency as compared to approaches with a working Gaussian hypothesis. We illustrate the methodology with the analysis of imprecisely measured income data.
Generalized additive models (GAMs) are a well-established statistical tool for modeling complex nonlinear relationships between covariates and a response assumed to have a conditional distribution in the exponential family. To make inference in this model class, a fast and flexible approach is considered based on Bayesian P-splines and the Laplace approximation. The proposed Laplace-P-spline model contributes to the development of a new methodology to explore the posterior penalty space by considering a deterministic grid-based strategy or a Markov chain sampler, depending on the number of smooth additive terms in the predictor. The approach has the merit of relying on a simple Gaussian approximation to the conditional posterior of latent variables with closed form analytical expressions available for the gradient and Hessian of the approximate posterior penalty vector. This enables to construct accurate posterior pointwise and credible set estimators for (functions of) regression and spline parameters at a relatively low computational budget even for a large number of smooth additive components. The performance of the Laplace-P-spline model is confirmed through different simulation scenarios and the method is illustrated on two real datasets.
Cure survival models are used when we desire to acknowledge explicitly that an unknown proportion of the population studied will never experience the event of interest. An extension of the promotion time cure model enabling the inclusion of time-varying covariates as regressors when modelling (simultaneously) the probability and the timing of the monitored event is presented. Our proposal enables us to handle non-monotone population hazard functions without a specific parametric assumption on the baseline hazard. This extension is motivated by and illustrated on data from the German Socio-Economic Panel by studying the transition to second and third births in West Germany.
This article introduces double additive models to describe the effect of continuous covariates in cure survival models, thereby relaxing the traditional linearity assumption in the two regression parts. This class of models extends the classical event history models when an unknown proportion of the population under study will never experience the event of interest. They are used on data from the German Socio-Economic Panel (GSOEP) to examine how age at first birth relates to the timing and quantum of fertility for given education levels of the respondents. It is shown that the conditional probability of having further children decreases with the mother's age at first birth. While the effect of age at first birth in the third birth's probability model is fairly linear, this is not the case for the second child with an accelerating decline detected for women that had their first kid beyond age 30.
Cure survival models are used when we desire to acknowledge explicitly that an unknown proportion of the population studied will never experience the event of interest. An extension of the promotion time cure model enabling the inclusion of time-varying covariates as regressors when modelling (simultaneously) the probability and the timing of the monitored event is presented. Our proposal enables us to handle non-monotone population hazard functions without a specific parametric assumption on the baseline hazard. This extension is motivated by and illustrated on data from the German Socio-Economic Panel by studying the transition to second and third births in West Germany.
The promotion time cure model is a survival model acknowledging that an unidentified proportion of subjects will never experience the event of interest whatever the duration of the follow-up. We focus our interest on the challenges raised by the strong posterior correlation between some of the regression parameters when the same covariates influence long- and short-term survival. Then, the regression parameters of shared covariates are strongly correlated with, in addition, identification issues when the maximum follow-up duration is insufficiently long to identify the cured fraction. We investigate how, despite this, plausible values for these parameters can be obtained in a computationally efficient way. The theoretical properties of our strategy will be investigated by simulation and illustrated on clinical data. Practical recommendations will also be made for the analysis of survival data known to include an unidentified cured fraction.
Bayesian methods for flexible time-to-event models usually rely on the theory of Markov chain Monte Carlo (MCMC) to sample from posterior distributions and perform statistical inference. These techniques are often plagued by several potential issues such as high posterior correlation between parameters, slow chain convergence and foremost a strong computational cost. A novel methodology is proposed to overcome the inconvenient facets intrinsic to MCMC sampling with the major advantage that posterior distributions of latent variables can rapidly be approximated with a high level of accuracy. This can be achieved by exploiting the synergy between Laplace’s method for posterior approximations and P-splines, a flexible tool for nonparametric modeling. The methodology is developed in the class of cure survival models, a useful extension of standard time-to-event models where it is assumed that an unknown proportion of unidentified (cured) units will never experience the monitored event. An attractive feature of this new approach is that point estimators and credible intervals can be straightforwardly constructed even for complex functionals of latent model variables. The properties of the proposed methodology are evaluated using simulations and illustrated on two real datasets. The fast computational speed and accurate results suggest that the combination of P-splines and Laplace approximations can be considered as a serious competitor of MCMC to make inference in semi-parametric models, as illustrated on survival models with a cure fraction.