A common feature of much survival data is censoring due to incompletely observed lifetimes. Survival analysis methods have been designed to take account of this and provide appropriate relevant summaries, such as the Kaplan-Meier plot and the median is easily read off this plot. However, a single summary is not really a relevant quantity for communication to an individual patient, as it conveys no notion of variability and uncertainty. The aim of this paper is to consider censored data as a form of missing data and impute them using Bayesian methods. We introduce two novel parametric and non-parametric Bayesian approaches for imputing right censored observations to be used as a complement to formal inferential methods and to allow more interpretable displays to be made for physicians and patients.
The modelling of count data in real world scenarios often requires models that address over-dispersion. Within the generalized linear modeling framework, various over-dispersion models are available as extensions of the basic Poisson model. Graphical assessment of goodnessof-fit commonly involves half-normal plots, for which a simulated envelope can be added to aid interpretation. The simulated envelope is such that, under a well-fitted model, the majority of points should fall within its bounds. Nonetheless, closely related models tend to produce very similar graphs. Here, we propose an objective statistic based on half-normal plots with a simulated envelope to aid goodness-of-fit assessment.
Traditional methods of model diagnostics may include a plethora of graphical techniques based on residual analysis, as well as formal tests (e.g. Shapiro-Wilk test for normality and Bartlett test for homogeneity of variance). In this paper we derive a new distance metric based on the half-normal plot with a simulation envelope, a graphical model evaluation method, and investigate its properties through simulation studies. The proposed metric can help to assess the fit of a given model, and also act as a model selection criterion by being comparable across models, whether based or not on a true likelihood. More specifically, it quantitatively encompasses the model evaluation principles and removes the subjective bias when closely related models are involved. We validate the technique by means of an extensive simulation study carried out using count data, and illustrate with two case studies in ecology and fisheries research.
This paper revisits the common problem of analysing counts recorded over time through the modelling of the underlying rate, motivated by the analysis of a cancer treatment related study. The baseline Poisson model is simply implemented though the inclusion of an offset for the different exposure/recording times and the underlying Poisson process gives other nice well-known properties. We consider how this approach can be extended to models for overdispersed data. The use of a simple offset with a negative binomial model is common practice and we consider the appropriateness of this and the resulting implications. We discuss how these ideas extend to more general mixed Poisson models, including ZIP, to handle zero-inflation, and ZINB for zero-inflation and overdispersion. The simple offset approach does not extend to other extended count models such as the COM-Poisson and general weighted Poisson distributions.
We consider discrete mortality data for groups of individuals observed over time. The fitting of cumulative mortality curves as a function of time involves the longitudinal modelling of the multinomial response. Typically such data exhibit overdispersion, that is greater variation than predicted by the multinomial distribution. To model the extra-multinomial variation (overdispersion) we consider a Dirichlet-multinomial model, a random intercept model and a random intercept and slope model. We construct asymptotic and robust covariance matrix estimators for the regression parameter standard errors. Applying this model to a specific insect bioassay of the fungus Beauveria bassiana, we note some simple relationships in the results and explore why these are simply a consequence of the data structure. Fitted models are used to make inferences on the effectiveness and consistency of different isolates of the fungus to provide recommendations for its use as a biological control in the field.
In this chapter, we discuss the analysis of data that typically arise from entomological studies using generalized linear models. We focus on techniques that can be used to assess model goodness of fit, which is an important step in statistical modelling to ensure the reliability of the inferences made. Specifically, we demonstrate the utility of half-normal plots with a simulated envelope as a complementary tool for assessing model assumptions. We illustrate the concepts with two examples, one involving count responses and another involving continuous responses.
A common feature of much survival data is censoring due to incompletely observed lifetimes. Survival analysis methods and models have been designed to take account of this and provide appropriate relevant summaries, such as the Kaplan–Meier plot and the commonly quoted median survival time of the group under consideration. However, a single summary is not really a relevant quantity for communication to an individual patient, as it conveys no notion of variability and uncertainty, and the Kaplan–Meier plot can be difficult for the patient to understand and also is often mis-interpreted, even by some physicians. This paper considers an alternative approach of treating the censored data as a form of missing, incomplete data and proposes an imputation scheme to construct a completed dataset. This allows the use of standard descriptive statistics and graphical displays to convey both typical outcomes and the associated variability. We propose a Bayesian approach to impute any censored observations, making use of other information in the dataset, and provide a completed dataset. This can then be used for standard displays, summaries, and even, in theory, analysis and model fitting. We particularly focus on the data visualisation advantages of the completed data, allowing displays such as density plots, boxplots, etc, to complement the usual Kaplan–Meier display of the original dataset. We study the performance of this approach through a simulation study and consider its application to two clinical examples.
We consider the analysis of count data in which the observed frequency of zero counts is unusually large, typically with respect to the Poisson distribution. We focus on two alternative modelling approaches: over-dispersion (OD) models and zero-inflation (ZI) models, both of which can be seen as generalisations of the Poisson distribution; we refer to these as implicit and explicit ZI models, respectively. Although sometimes seen as competing approaches, they can be complementary; OD is a consequence of ZI modelling, and ZI is a by-product of OD modelling. The central objective in such analyses is often concerned with inference on the effect of covariates on the mean, in light of the apparent excess of zeros in the counts. Typically, the modelling of the excess zeros per se is a secondary objective, and there are choices to be made between, and within, the OD and ZI approaches. The contribution of this paper is primarily conceptual. We contrast, descriptively, the impact on zeros of the two approaches. We further offer a novel descriptive characterisation of alternative ZI models, including the classic hurdle and mixture models, by providing a unifying theoretical framework for their comparison. This in turn leads to a novel and technically simpler ZI model. We develop the underlying theory for univariate counts and touch on its implication for multivariate count data.
A virtual interview with Murray Aitkin by Brian Francis and John Hinde, two of the original members of the Centre for Applied Statistics that Murray created at Lancaster University. The talk ranges over Murray's reflections of a career in statistical modelling and the many different collaborations across the world that have been such a significant part of it.
Survival models have been extensively used to analyse time-until-event data. There is a range of extended models that incorporate different aspects, such as overdispersion/frailty, mixtures, and flexible response functions through semi-parametric models. In this work, we show how a useful tool to assess goodness-of-fit, the half-normal plot of residuals with a simulated envelope, implemented in the hnp package in R, can be used on a location-scale modelling context. We fitted a range of survival models to time-until-event data, where the event was an insect predator attacking a larva in a biological control experiment. We started with the Weibull model and then fitted the exponentiated-Weibull location-scale model with regressors both for the location and scale parameters. We performed variable selection for each model and, by producing half-normal plots with simulated envelopes for the deviance residuals of the model fits, we found that the exponentiated-Weibull fitted the data better. We then included a random effect in the exponentiated-Weibull model to accommodate correlated observations. Finally, we discuss possible implications of the results found in the case study.
The orange variety "x11", which is a spontaneous mutant of the sweet orange, has a short juvenile period with early flowering. The data used in this paper are from a randomized design experiment that aimed to assess the plants' flowering characteristics when grafted onto two different varieties of lemon rootstock. The plants were pruned in each of the four seasons, and on each pruning occasion, the number of branches on each plant was counted and classified into four mutually exclusive flowering categories. The data presented large variability and many zeros. The statistical analysis included the use of generalized linear mixed models with a Bayesian approach. The results showed that flowering is not equal over the seasons, i.e., there are significant differences in the classification of the branches across the four seasons and the two varieties, with interactions between seasonal and branch effects.
Background and Objective Observational studies and experiments in medicine, pharmacology and agronomy are often concerned with assessing whether different methods/raters produce similar values over the time when measuring a quantitative variable. This article aims to describe the statistical package lcc, for are, that can be used to estimate the extent of agreement between two (or more) methods over the time, and illustrate the developed methodology using three real examples. Methods The longitudinal concordance correlation, longitudinal Pearson correlation, and longitudinal accuracy functions can be estimated based on fixed effects and variance components of the mixed-effects regression model. Inference is made through bootstrap confidence intervals and diagnostic can be done via plots, and statistical tests. Results The main features of the package are estimation and inference about the extent of agreement using numerical and graphical summaries. Moreover, our approach accommodates both balanced and unbalanced experimental designs or observational studies, and allows for different within-group error structures, while allowing for the inclusion of covariates in the linear predictor to control systematic variations in the response. All examples show that our methodology is flexible and can be applied to many different data types. Conclusions The lcc package, available on the CRAN repository, proved to be a useful tool to describe the agreement between two or more methods over time, allowing the detection of changes in the extent of agreement. The inclusion of different structures for the variance-covariance matrices of random effects and residuals makes the package flexible for working with different types of databases.
In the analysis of count data often the equidispersion assumption is not suitable, hence the Poisson regression model is inappropriate. As a generalization of the Poisson distribution, the COM-Poisson distribution can deal with under-, equi- and overdispersed count data. It is a member of the exponential family of distributions and has well known special cases. In spite of the nice properties of the COM-Poisson distribution, its location parameter does not correspond to the expectation, which complicates the interpretation of regression models. In this paper, we propose a straightforward reparametrization of the COM-Poisson distribution based on an approximation to the expectation of this distribution. The main advantage of our new parametrization is the straightforward interpretation of the regression coefficients in terms of the expectation, as usual in the context of generalized linear models. Furthermore, the estimation and inference for the new COM-Poisson regression model can be done based on the likelihood paradigm. We carried out simulation studies to verify the finite sample properties of the maximum likelihood estimators. The results from our simulation study show that the maximum likelihood estimators are unbiased and consistent for both regression and dispersion parameters. We observed that the empirical correlation between the regression and dispersion parameter estimators is close to zero, which suggests that these parameters are orthogonal. We illustrate the application of the proposed model through the analysis of three data sets with over-, under- and equidispersed count data. The study of distribution properties through a consideration of dispersion, zero-inflated and heavy tail indexes, together with the results of data analysis show the flexibility over standard approaches.
Transition models are an important framework that can be used to model longitudinal categorical data. They are particularly useful when the primary interest is in prediction. The available methods for this class of models are suitable for the cases in which responses are recorded individually over time. However, in many areas, it is common for categorical data to be recorded as groups, that is, different categories with a number of individuals in each. As motivation we consider a study in insect movement and another in pig behaviou. The first study was developed to understand the movement patterns of female adults of Diaphorina citri, a pest of citrus plantations. The second study investigated how hogs behaved under the influence of environmental enrichment. In both studies, the number of individuals in different response categories was observed over time. We propose a new framework for considering the time dependence in the linear predictor of a generalized logit transition model using a quantitative response, corresponding to the number of individuals in each category. We use maximum likelihood estimation and present the results of the fitted models under stationarity and non‐stationarity assumptions, and use recently proposed tests to assess non‐stationarity. We evaluated the performance of the proposed model using simulation studies under different scenarios, and concluded that our modeling framework represents a flexible alternative to analyze grouped longitudinal categorical data.
We propose a flexible class of regression models for continuous bounded data based on second-moment assumptions. The mean structure is modelled by means of a link function and a linear predictor, while the mean and variance relationship has the form [Formula: see text], where [Formula: see text], [Formula: see text] and [Formula: see text] are the mean, dispersion and power parameters respectively. The models are fitted by using an estimating function approach where the quasi-score and Pearson estimating functions are employed for the estimation of the regression and dispersion parameters respectively. The flexible quasi-beta regression model can automatically adapt to the underlying bounded data distribution by the estimation of the power parameter. Furthermore, the model can easily handle data with exact zeroes and ones in a unified way and has the Bernoulli mean and variance relationship as a limiting case. The computational implementation of the proposed model is fast, relying on a simple Newton scoring algorithm. Simulation studies, using datasets generated from simplex and beta regression models show that the estimating function estimators are unbiased and consistent for the regression coefficients. We illustrate the flexibility of the quasi-beta regression model to deal with bounded data with two examples. We provide an R implementation and the datasets as supplementary materials.