Since their introduction in 1972, generalized linear models (GLMs) have proven useful in the generalization of classical normal models. Presenting methods for fitting GLMs with random effects to data, Generalized Linear Models with Random Effects: Unified Analysis via H-likelihood explores a wide range of applications, including combining informati
A modelling approach has been useful for the analysis of data from robust designs for quality improvement. Recently, Robinson et al. (J. Qual. Technol. 2006; 38:65-38) proposed the use of generalized linear mixed models (GLMMs) and they used the marginal quasi-likelihood (MQL) method of Breslow and Clayton (J. Am. Statist, Ass. 1983; 88:9-25). Hierarchical generalized linear models (HGLMs) extend GLMMs by allowing structured dispersions and conjugate distributions of arbitrary GLM families for random effects. In this paper we use two examples to illustrate how these additional features in HGLMs can be used for the analysis of data from quality-improvement experiments. We also show that the hierarchical likelihood (HL, or h-likelihood) estimators have better statistical properties than the MQL estimators. Copyright (C) 2010 John Wiley & Sons, Ltd.
SummaryIn a regression model, the joint distribution for each finite sample of units is determined by a function px(y) depending only on the list of covariate values x=(x(u1),…,x(un)) on the sampled units. No random sampling of units is involved. In biological work, random sampling is frequently unavoidable, in which case the joint distribution p(y,x) depends on the sampling scheme. Regression models can be used for the study of dependence provided that the conditional distribution p(y|x) for random samples agrees with px(y) as determined by the regression model for a fixed sample having a non-random configuration x. The paper develops a model that avoids the concept of a fixed population of units, thereby forcing the sampling plan to be incorporated into the sampling distribution. For a quota sample having a predetermined covariate configuration x, the sampling distribution agrees with the standard logistic regression model with correlated components. For most natural sampling plans such as sequential or simple random sampling, the conditional distribution p(y|x) is not the same as the regression distribution unless px(y) has independent components. In this sense, most natural sampling schemes involving binary random-effects models are biased. The implications of this formulation for subject-specific and population-averaged procedures are explored.
Two papers have recently been published in this journal which purport to deal with the mixed-models controversy ( Lencina et al. , 2005 ; Lencina & Singer, 2006 , to be referred to as L1 and L2). In my view they do not represent the current state of thinking on this subject. My own work begins with Nelder (1977) and continues with papers in 1982, and particularly with Nelder (1994, 1995) . Further papers are Nelder (1997) , whose title includes the phrase ‘The great mixed-model muddle’, and Nelder (1998) . In Nelder (1994) I describe what I regard as three false steps that have generated confusion and show how a consistent treatment may be developed. The two most important ideas are (1) marginality relations between terms in a factorial model, and (2) why constraints must not be put on parameters because they are required on estimates.
In recent issues of this journal it has been asserted in two papers that the use of h-likelihood is wrong, in the sense of giving unsatisfactory estimates of some parameters for binary data (Kuk and Cheng, 1999; Waddington and Thompson, 2004) or theoretically unsound (Kuk and Cheng, 1999). We wish to refute both these assertions.
International Statistical ReviewVolume 76, Issue 1 p. 134-135 What is the Mixed-Models Controversy? John A. Nelder, John A. Nelder Imperial College, London, UKSearch for more papers by this author John A. Nelder, John A. Nelder Imperial College, London, UKSearch for more papers by this author First published: 21 November 2007 https://doi.org/10.1111/j.1751-5823.2007.00022_1.xCitations: 6Read the full textAboutPDF ToolsRequest permissionExport citationAdd to favoritesTrack citation ShareShare Give accessShare full text accessShare full-text accessPlease review our Terms and Conditions of Use and check box below to share full-text version of article.I have read and accept the Wiley Online Library Terms and Conditions of UseShareable LinkUse the link below to share a full-text version of this article with your friends and colleagues. Learn more.Copy URL Share a linkShare onEmailFacebookTwitterLinkedInRedditWechat Abstract Two papers have recently been published in this journal which purport to deal with the mixed-models controversy (Lencina et al., 2005; Lencina & Singer, 2006, to be referred to as L1 and L2). In my view they do not represent the current state of thinking on this subject. My own work begins with Nelder (1977) and continues with papers in 1982, and particularly with Nelder (1994, 1995). Further papers are Nelder (1997), whose title includes the phrase 'The great mixed-model muddle', and Nelder (1998). In Nelder (1994) I describe what I regard as three false steps that have generated confusion and show how a consistent treatment may be developed. The two most important ideas are (1) marginality relations between terms in a factorial model, and (2) why constraints must not be put on parameters because they are required on estimates. References Lencina, V.B. & Singer, J.M. (2006). Measure for measure: exact F tests and the mixed models controversy. Int. Stat. Rev., 74, 391–402. Lencina, V.B., Singer, J.M. & Stanek, E.J. (2005). Much ado about nothing: the mixed models controversy revisited. Int. Stat. Rev., 73, 9–20. Nelder, J.A. (1977). A reformulation of linear models. J. Roy. Stat. Soc. A, 140, 48–77. Nelder, J.A. (1982). Linear models and non-orthogonal data. Utilitas Math., 21B, 141–152. Nelder. J.A. (1994). The Statistics of linear models: back to basics. Stat. Comput., 4, 221–234. Nelder, J.A. (1995). Rejoinder to comments on 'The statistics of linear models: back to basics'. Stat. Comput., 5, 109–111. Nelder, J.A. (1997). The great mixed-model muddle is alive and flourishing—alas! Food Quality and Preference, 9, 157–159. Nelder, J.A. (1998). The selection of terms in response-surface models—How strong is the weak-heredity principle? Amer. Stat., 52, 315–318. Citing Literature Volume76, Issue1April 2008Pages 134-135 ReferencesRelatedInformation
When there are two alternative random-effect models leading to the same marginal model, inferences from one model can be used for the other model. We illustrate how a likelihood method for fitting models with independent random effects can be applied to seemingly very different models with correlated random effects. We also discuss some merits of using these alternative models.
Summary A general class of statistical models for a univariate response variable is presented which we call the generalized additive model for location, scale and shape (GAMLSS). The model assumes independent observations of the response variable y given the parameters, the explanatory variables and the values of the random effects. The distribution for the response variable in the GAMLSS can be selected from a very general family of distributions including highly skew or kurtotic continuous and discrete distributions. The systematic part of the model is expanded to allow modelling not only of the mean (or location) but also of the other parameters of the distribution of y, as parametric and/or additive nonparametric (smooth) functions of explanatory variables and/or random-effects terms. Maximum (penalized) likelihood estimation is used to fit the (non)parametric models. A Newton–Raphson or Fisher scoring algorithm is used to maximize the (penalized) likelihood. The additive terms in the model are fitted by using a backfitting algorithm. Censored data are easily incorporated into the framework. Five data sets from different fields of application are analysed to emphasize the generality of the GAMLSS class of models.
There has existed controversy about the use of marginal and conditional models, particularly in the analysis of data from longitudinal studies. We show that alleged differences in the behavior of parameters in so-called marginal and conditional models are based on a failure to compare like with like. In particular, these seemingly apparent differences are meaningless because they are mainly caused by preimposed unidentifiable constraints on the random effects in models. We discuss the advantages of conditional models over marginal models. We regard the conditional model as fundamental, from which marginal predictions can be made.
Restricted likelihood was originally introduced as the criterion for the estimation of dispersion components in normal mixed linear models. Lee & Nelder (2001a) showed that it can be extended to a much wider class of models via double extended quasi-likelihood. We give a detailed description of the new method and show that it gives an efficient estimation procedure for dispersion components.
A single data transformation may fail to satisfy all the required properties necessary for an analysis. With generalized linear models (GLMs), the identification of the mean-variance relationship and the choice of the scale on which the effects are to be measured can be done separately, overcoming the shortcomings of the data-transformation approach. GLMs also provide an extension of the response surface approach. In this paper, we set out the current status of the GLM approach to the analysis of data from quality-improvement experiments and discuss its merits.
A search for a good parsimonious model is often required in data analysis. However, unfortunately we may end up with a falsely parsimonious model. Misspecification of the variance structure causes a loss of efficiency in regression estimation and this can lead to large standard-error estimates, producing possibly false parsimony. With generalized linear models (GLMs) we can keep the link function fixed while changing the variance function, thus allowing us to recognize false parsimony caused by such increased standard errors. With data transformation, any change of transformation automatically changes the scale for additivity, making false parsimony hard to recognize.
In multi‐centre clinical trials, heterogeneities in individual hospital treatment effects can be modelled as random effects. Estimates of the individual hospital treatment effects and estimate of the mean treatment effect, allowing for the presence of overall hospital differences, are required, together with some measure of their uncertainty. Systematic inferences from the hierarchical‐likelihood are now possible, using hierarchical generalized linear models. We show how to construct profile likelihoods for the treatment effects of individual hospitals. Copyright © 2002 John Wiley & Sons, Ltd.
Hierarchical generalised linear models are developed as a synthesis of generalised linear models, mixed linear models and structured dispersions. We generalise the restricted maximum likelihood method for the estimation of dispersion to the wider class and show how the joint fitting of models for mean and dispersion can be expressed by two interconnected generalised linear models. The method allows models with (i) any combination of a generalised linear model distribution for the response with any conjugate distribution for the random effects, (ii) structured dispersion components, (iii) different link and variance functions for the fixed and random effects, and (iv) the use of quasilikelihoods in place of likelihoods for either or both of the mean and dispersion models. Inferences can be made by applying standard procedures, in particular those for model checking, to components of either generalised linear model. We also show by numerical studies that the new method gives an efficient estimation procedure for substantial class of models of practical importance. Likelihood-type inference is extended to this wide class of models in a unified way.
We introduce a model class that includes many types of correlation structures for non-Gaussian models. We then show how to check the underlying model assumptions to discriminate between different correlation patterns and demonstrate how to select suitable models. Strawberry data are used to discuss the choice between fixed-and random-effect models for the fertility effect in agricultural experiments. Prostatecancer data are used to demonstrate the method applied to the analysis of longitudinal studies and Scottish lipcancer data to illustrate an application to spatial statistics.
For contingency tables with one factor as a response, log-linear models can be used provided that a minimal model, which constrains the predicted sample sizes of the response factor to equal the actual sizes, is fitted first. Only models that contain this minimal term make inferential sense. For response factors with two levels the results from fitting such log-linear models are identical with those from the corresponding logistic regression. The correspondence is illustrated with two examples from recent literature, both of which have been previously misanalysed.
SUMMARY For non-normal data assumed to have distributions, such as the Poisson distribution, which have an a priori dispersion parameter, there are two ways of modelling overdispersion: by a quasi-likelihood approach or with a random-effect model. The two approaches yield different variance functions for the response, which may be distinguishable if adequate data are available. The epilepsy data of Thall and Vail and the fabric data of Bissell are used to exemplify the ideas.