A fast way is proposed based on the multiple outputation idea (Hoffman et al., ; Follmann et al., ) to calculate the precision of parameter estimates for high‐dimensional multivariate joint models using a pairwise approach (Fieuws and Verbeke, ; Fieuws et al., ). Simulation results as well as data analysis shows possibly more than 2500 times faster computations using the proposed method. In our real data illustration, the time gain is more than 330 times.
In this paper, we study the psychometric properties of the ten-item version of the Big Five Inventory as a subset of the original BFI in a Flemish sample. The data come from the Divorce in Flanders study and consist of a full sample of 7533 individuals from 4460 families. Factor analysis shows the presence of the Big Five factor structure with very high primary loadings for most items. However, one of the Agreeableness items loads exclusively on the Extraversion factor and within-factor correlations are also low. Despite this, the BFI-10 correlates well with the BFI. Therefore, while more research is needed before validity and reliability of the Dutch-language version of the test can be concluded, it is clear that the BFI-10 may prove very effective in the assessment of the Big Five factors in the Flemish and Dutch cultures when assessment with longer questionnaires is not feasible.
For item response theory (IRT) models, which belong to the class of generalized linear or non-linear mixed models, reliability at the scale of observed scores (i.e., manifest correlation) is more difficult to calculate than latent correlation based reliability, but usually of greater scientific interest. This is not least because it cannot be calculated explicitly when the logit link is used in conjunction with normal random effects. As such, approximations such as Fisher's information coefficient, Cronbach's α, or the latent correlation are calculated, allegedly because it is easy to do so. Cronbach's α has well-known and serious drawbacks, Fisher's information is not meaningful under certain circumstances, and there is an important but often overlooked difference between latent and manifest correlations. Here, manifest correlation refers to correlation between observed scores, while latent correlation refers to correlation between scores at the latent (e.g., logit or probit) scale. Thus, using one in place of the other can lead to erroneous conclusions. Taylor series based reliability measures, which are based on manifest correlation functions, are derived and a careful comparison of reliability measures based on latent correlations, Fisher's information, and exact reliability is carried out. The latent correlations are virtually always considerably higher than their manifest counterparts, Fisher's information measure shows no coherent behaviour (it is even negative in some cases), while the newly introduced Taylor series based approximations reflect the exact reliability very closely. Comparisons among the various types of correlations, for various IRT models, are made using algebraic expressions, Monte Carlo simulations, and data analysis. Given the light computational burden and the performance of Taylor series based reliability measures, their use is recommended.
Inference in mixed models is often based on the marginal distribution obtained from integrating out random effects over a pre-specified, often parametric, distribution. In this paper, we present the so-called gradient function as a simple graphical exploratory diagnostic tool to assess whether the assumed random-effects distribution produces an adequate fit to the data, in terms of marginal likelihood. The method does not require any calculations in addition to the computations needed to fit the model, and can be applied to a wide range of mixed models (linear, generalized linear, non-linear), with univariate as well as multivariate random effects, as long as the distribution for the outcomes conditional on the random effects is correctly specified. In case of model misspecification, the gradient function gives an important, albeit informal, indication on how the model can be improved in terms of random-effects distribution. The diagnostic value of the gradient function is extensively illustrated using some simulated examples, as well as in the analysis of a real longitudinal study with binary outcome values.
In statistical practice, incomplete measurement sequences are the rule rather than the exception. Fortunately, in a large variety of settings, the stochastic mechanism governing the incompleteness can be ignored without hampering inferences about the measurement process. While ignorability only requires the relatively general missing at random assumption for likelihood and Bayesian inferences, this result cannot be invoked when non-likelihood methods are used. A direct consequence of this is that a popular non-likelihood-based method, such as generalized estimating equations, needs to be adapted towards a weighted version or doubly-robust version when a missing at random process operates. So far, no such modification has been devised for pseudo-likelihood based strategies. We propose a suite of corrections to the standard form of pseudo-likelihood to ensure its validity under missingness at random. Our corrections follow both single and double robustness ideas, and is relatively simple to apply. When missingness is in the form of dropout in longitudinal data or incomplete clusters, such a structure can be exploited toward further corrections. The proposed method is applied to data from a clinical trial in onychomycosis and a developmental toxicity study.
Whereas marginal models, random-effects models, and conditional models are routinely considered to be the three main modeling families for continuous and discrete repeated measures with linear and generalized linear mean structures, respectively, it is less common to consider nonlinear models, let alone frame them within the above taxonomy. In the latter situation, indeed, when considered at all, the focus is often exclusively on random-effects models. In this article, we consider all three families, exemplify their great flexibility and relative ease of use, and apply them to a simple but illustrative set of data on tree circumference growth of orange trees. This article has supplementary material online.
For over two decades, following the pioneering work of Rubin (1976) and Little (1976), there has been a growing literature on incomplete data, with a lot of emphasis on longitudinal data. Following the original work of Rubin and Little, there has evolved a general view that “likelihood methods” that ignore the missing value mechanism are valid under an MAR process, where likelihood is interpreted in a frequentist sense. The availability of flexible standard software for incomplete data, such as PROC MIXED, and the advantages quoted in Section 17.3 contribute to this point of view. This statement needs careful qualification however. Kenward and Molenberghs (1998) provided an exposition of the precise sense in which frequentist methods of inference are justified under MAR processes.