A previous Monte Carlo study examined the relative powers of several simple and more complex procedures for testing the significance of difference in mean rates of change in a controlled, longitudinal, treatment evaluation study. Results revealed that the relative powers depended on the correlation structure of the simulated repeated measurements. Tests on dropout-weighted linear slope coefficients fitted to all of the available measurements for each participant were found to provide superior power in the presence of compound symmetry (CS), but tests of significance applied to simple baseline-to-endpoint difference scores provided superior power in the presence of a strongly autoregressive (AR) correlation structure. Type I error rates appeared in an acceptable range for both of those analyses. Insofar as the previous study considered only two widely disparate correlation structures, the present work was undertaken to examine where along a continuum of correlation structures lying between strongly AR and CS the power balance shifts from favoring the simple endpoint difference-score analysis to favoring a regression analysis that utilizes all of the available repeated measurements for each participant. With power calculated from the relative frequencies of rejecting Ho at different levels of autoregression, the results indicate superior power for the simple endpoint analysis across more than half the distance from strongly AR to CS. To examine replicability of the simulation results using real data from a previously published study, sampling with replacement from a double-blind controlled study examining the treatment of depression was used to create a Monte Carlo data set from which power could be calculated from relative frequencies of rejecting Ho.
This article concerns methodology for testing the significance of differences in mean rates of change in controlled repeated measurements designs with limited sample sizes, autoregressive error structures, nonlinear patterns of underlying true mean change, dropout rates exceeding 50%, plus other missing data. Each of these is problematic for ordinary repeated measures analysis of variance, and a complex generalized linear mixed model formulation popularly advocated for the ability to deal with autoregressive error structures and missing data is shown to perform poorly in such circumstances. Monte Carlo simulation methods confirm that simple two-stage analyses of dropout-weighted linear slope coefficients provide conservative Type 1 error protection, although adequate power requires the presence of large treatment effects in studies with the limited sample sizes and high proportions of missing data. No other analysis has been documented to provide both conservative Type 1 error protection and competitive power under similarly taxing conditions.
Last-observation-carried-forward (LOCF) originated as a way of replacing missing measurements for use in classical repeated measurements analysis of variance (ANOVA) where complete data are required. As mathematical statisticians developed newer methods in the attempt to deal effectively with missing data, LOCF fell largely by the wayside. This paper reports use of simulation methods to evaluate validity of tests of significance for differences in mean rates of change applied to LOCF-completed data and compares results with those from alternative procedures that do not require replacement of the missing values for dropouts. Across several different data structures and dropout mechanisms, tests of significance for difference in means of slope coefficients fitted to LOCF data maintained appropriate Type 1 error protection. In no case where the Type 1 error rates of more complex generalized linear mixed model (GLMM) tests appeared appropriately conservative did the power of those tests exceed power of the simpler test on LOCF-completed data.
Abstract. Differences in mean rates of change are of primary interest in many controlled treatment evaluation studies. Generalized linear mixed model (GLMM) procedures are widely conceived to be the preferred method of analysis for repeated measurement designs when there are missing data due to dropouts, but systematic dependence of the dropout probabilities on antecedent or concurrent factors poses a problem for testing the significance of differences in mean rates of change across time in such designs. Controlling for the dependence of dropout probabilities on baseline values poses a special problem because a theoretically correct GLMM random-effects model does not permit including the same baseline score as both covariate and dependent variable. Monte Carlo methods are used herein to evaluate the actual Type 1 error rates and power resulting from two commonly-illustrated GLMM random-effects model formulations for testing the GROUPS × TIMES linear interaction effect in group-randomized repeated measurem...
This article is about a simple two-stage analysis that utilizes slope coefficients as the dependent variable for testing the significance of difference in mean rates of change in repeated measurement designs with missing data. The ANCOVA test on the doubly weighted slope coefficients provides power comparable to that of more complex maximum likelihood procedures when data are missing completely at random, requires fewer assumptions and is more generally applicable under realistic nonrandom dropout conditions, and most importantly can be readily understood and explained by those who actually do most controlled clinical research.
Recent contributions to the statistical literature have provided elegant model‐based solutions to the problem of estimating sample sizes for testing the significance of differences in mean rates of change across repeated measures in controlled longitudinal studies with differentially correlated error and missing data due to dropouts. However, the mathematical complexity and model specificity of these solutions make them generally inaccessible to most applied researchers who actually design and undertake treatment evaluation research in psychiatry. In contrast, this article relies on a simple two‐stage analysis in which dropout‐weighted slope coefficients fitted to the available repeated measurements for each subject separately serve as the dependent variable for a familiar ANCOVA test of significance for differences in mean rates of change. This article is about how a sample of size that is estimated or calculated to provide desired power for testing that hypothesis without considering dropouts can be adjusted appropriately to take dropouts into account. Empirical results support the conclusion that, whatever reasonable level of power would be provided by a given sample size in the absence of dropouts, essentially the same power can be realized in the presence of dropouts simply by adding to the original dropout‐free sample size the number of subjects who would be expected to drop from a sample of that original size under conditions of the proposed study. Copyright © 2006 John Wiley & Sons, Ltd.
OBJECTIVE:The authors examined clinical differences between divalproex sodium and generic immediate-release valproic acid.METHOD:This 6-year prospective, quasi-experimental clinical trial compared the effectiveness and tolerability of divalproex and valproic acid. The dependent variables were length of hospital stay, rehospitalization rate, and adverse drug reactions in 9,260 psychiatric admissions.RESULTS:Inpatients who initially received divalproex sodium had a 32.7% longer hospital stay and 3.8% higher readmission rate than did patients who initially received valproic acid. Initial treatment with divalproex prolonged length of stay by 30.3% in patients treated with divalproex and valproic acid during different admissions. After other variables were controlled by multiway analysis of variance, the hospital stay of patients who continued the initial medication was 15.2% longer (2.0 days) for divalproex than valproic acid. Switching medications was more common for valproic acid, partly because of study design. Medication intolerance occurred in approximately 6.4% more patients taking valproic acid than divalproex. However, switching from valproic acid to divalproex did not significantly prolong length of stay, over that for continuous divalproex, or increase the rehospitalization rate.CONCLUSIONS:Lower peak valproate concentrations with divalproex sodium may have enhanced tolerability but may also explain the lower effectiveness. Extended-release divalproex could lower effectiveness further and require higher doses. Thus, inpatients are better served by beginning with generic valproic acid and by changing to delayed-release divalproex only if intolerance occurs. This would save up to one-third of inpatient costs and two-thirds of a billion dollars yearly in medication costs.
Autocorrelated error and missing data due to dropouts have fostered interest in the flexible general linear mixed model (GLMM) procedures for analysis of data from controlled clinical trials. The user of these adaptable statistical tools must, however, choose among alternative structural models to represent the correlated repeated measurements. The fit of the error structure model specification is important for validity of tests for differences in patterns of treatment effects across time, particularly when maximum likelihood procedures are relied upon. Results can be affected significantly by the error specification that is selected, so a principled basis for selecting the specification is important. As no theoretical grounds are usually available to guide this decision, empirical criteria have been developed that focus on model fit. The current report proposes alternative empirical criteria that focus on bootstrap estimates of actual type I error and power of tests for treatment effects. Results for model selection before and after the blind is broken are compared. Goodness‐of‐fit statistics also compare favourably for models fitted to the blinded or unblinded data, although the correspondence to actual type I error and power depends on the particular fit statistic that is considered. Copyright © 2004 Whurr Publishers Ltd.
A split-sample replication criterion originally proposed by J. E. Overall and K. N. Magee (1992) as a stopping rule for hierarchical cluster analysis is applied to multiple data sets generated by sampling with replacement from an original simulated primary data set. An investigation of the validity of this bootstrap procedure was undertaken using different combinations of the true number of latent populations, degrees of overlap, and sample sizes. The bootstrap procedure enhanced the accuracy of identifying the true number of latent populations under virtually all conditions. Increasing the size of the resampled data sets relative to the size of the primary data set further increased accuracy. A computer program to implement the bootstrap stopping rule is made available via a referenced Web site.
Generalized linear model analyses of repeated measurements typically rely on simplifying mathematical models of the error covariance structure for testing the significance of differences in patterns of change across time. The robustness of the tests of significance depends, not only on the degree of agreement between the specified mathematical model and the actual population data structure, but also on the precision and robustness of the computational criteria for fitting the specified covariance structure to the data. Generalized estimating equation (GEE) solutions utilizing the robust empirical sandwich estimator for modeling of the error structure were compared with general linear mixed model (GLMM) solutions that utilized the commonly employed restricted maximum likelihood (REML) procedure. Under the conditions considered, the GEE and GLMM procedures were identical in assuming that the data are normally distributed and that the variance-covariance structure of the data is the one specified by the user.The question addressed in this article concerns relative sensitivity of tests of significance for treatment effects to varying degrees of misspecification of the error covariance structure model when fitted by the alternative procedures. Simulated data that were subjected to monte carlo evaluation of actual Type I error and power of tests of the equal slopes hypothesis conformed to assumptions of ordinary linear model ANOVA for repeated measures except for autoregressive covariance structures and missing data due to dropouts. The actual within-groups correlation structures of the simulated repeated measurements ranged from AR(1) to compound symmetry in graded steps, whereas the GEE and GLMM formulations restricted the respective error structure models to be either AR(I), compound symmetry (CS), or unstructured (UN). The GEE-based tests utilizing empirical sandwich estimator criteria were documented to be relatively insensitive to misspecification of the covariance structure models, whereas GLMM tests which relied on restricted maximum likelihood (REML) were highly sensitive to relatively modest misspecification of the error correlation structure even though normality, variance homogeneity, and linearity were not an issue in the simulated data. Goodness-of-fit statistics were of little utility in identifying cases in which relatively minor misspecification of the GLMM error structure model resulted in inadequate alpha protection for tests of the equal slopes hypothesis. Both GEE and GLMM formulations that relied on unstructured (UN) error model specification produced nonconservative results regardless of the actual correlation structure of the repeated measurements. A random coefficients model produced robust tests with competitive power across all conditions examined.
This paper examines the implications of the correlational structure of repeated measurements for three indices of change that can be used to evaluate treatment effects in longitudinal studies with scheduled assessment times and fixed total duration. The generalized least squares (GLS) regression of repeated measurements on time, which is usually reserved for complex mixed model solutions, takes the correlational structure of the repeated measurements into account, whereas simple gain scores and ordinary least squares (OLS) regression calculations do not. Nevertheless, the GLS solution is equivalent to OLS under conditions of compound symmetry and is equivalent to the analysis of simple gain scores in the presence of an autoregressive (order 1) correlational structure. The understanding of these relationships is important with regard to the frequently heard criticisms of the simpler definitions of treatment response in repeated measurement designs.
This paper examines the implications of the correlational structure of repeated measurements for three indices of change that can be used to evaluate treatment effects in longitudinal studies with scheduled assessment times and fixed total duration. The generalized least squares (GLS) regression of repeated measurements on time, which is usually reserved for complex mixed model solutions, takes the correlational structure of the repeated measurements into account, whereas simple gain scores and ordinary least squares (OLS) regression calculations do not. Nevertheless, the GLS solution is equivalent to OLS under conditions of compound symmetry and is equivalent to the analysis of simple gain scores in the presence of an autoregressive (order 1) correlational structure. The understanding of these relationships is important with regard to the frequently heard criticisms of the simpler definitions of treatment response in repeated measurement designs.
Controlled clinical trials in neuropsychopharmacology, as in numerous other clinical research domains, tend to employ a conventional parallel-groups design with repeated measurements. The hypothesis of primary interest in the relatively short-term, double-blind trials, concerns the difference between patterns or magnitudes of change from baseline. A simple two-stage approach to the analysis of such data involves calculation of an index or coefficient of change in stage 1 and testing the significance of difference between group means on the derived measure of change in stage 2. This article has the aim of introducing formulas and a computer program for sample size and/or power calculations for such two-stage analyses involving each of three definitions of change, with or without baseline scores entered as a covariate, in the presence of homogeneous or heterogeneous (autoregressive) patterns of correlation among the repeated measurements. Empirical adjustments of sample size for the projected dropout rates are also provided in the computer program.
A project that originated with the aim of documenting the implications of dropouts for tests of significance based on general linear mixed model procedures resulted in recognition of problems in the use of SAS PROC.MIXED for this purpose. In responding to suggestions and criticisms, we have further analyzed simulated clinical trial data with realistic autoregressive structure, using alternative error model formulations, different approaches to the use of covariates to model dropout patterns, and different ways to include the critical time variable in the mixed model. Results emphasize the sensitivity of the PROC.MIXED tests of significance for GROUP and TIME × GROUP equal slopes hypothesis to less than optimal modeling of the error covariance structure. Even with the authoritatively recommended best available modeling of the error structure, model formulations that made use of the REPEATED statement did not maintain conservative test sizes when covariates were required to model dropout data patterns. Random coefficients models that employed the RANDOM statement did permit appropriate covariate controls, but the tests of significance for treatment effects were lacking in power. After examining a variety of alternative PROC.MIXED model formulations, it is concluded that none provided both Type I error protection and power comparable to that of simple two-stage analysis of covariance (ANCOVA) procedures for confirming the presence of true treatment effects in controlled clinical trials. Other issues examined in this article concern treating baseline scores as both covariate and initial repeated measurement to which a linear means model is fitted, failure to take advantage of the regression of repeated measurements on time in modeling time as an unordered categorical variable, and fitting linear regression models to nonlinear response patterns.
The work reported in this article was undertaken to evaluate the utility of SAS PROC.MIXED for testing hypotheses concerning GROUP and TIME x GROUP effects in repeated measurements designs with drop-outs. If dropouts are not completely at random, covariate control over informative individual differences on which dropout data patterns depend is widely recognized to be important. However, the inclusion of baseline scores and time-in-study as between-subject covariates in an otherwise well formulated SAS PROC.MIXED model resulted in inadequate control over type I error in simulated data with or without drop-outs present. The inadequate model formulations and resulting deviant test sizes are presented here as a warning for others who might be guided by the same information sources to employ similar model specifications when analyzing data from actual clinical trials. It is important that the complete model specification be provided in detail when reporting applications of the general linear mixed-model procedure. A single random-coefficients model produced appropriate test sizes, but it provided inferior power when informative covariates were added in the attempt to adjust for dropouts. As an alternative, the incorporation of covariate controls in simpler two-stage endpoint or random regression analyses is documented to be effective in dealing with dropouts under specifiable conditions.
The power of univariate and multivariate tests of significance is compared in relation to linear and nonlinear patterns of treatment effects in a repeated measurement design. Bonferroni correction was used to control the experiment-wise error rate in combining results from univariate tests of significance accomplished separately on average level, linear, quadratic, and cubic trend components. Multivariate tests on these same components of the overall treatment effect, as well as a multivariate test for between-groups difference on the original repeated measurements, were also evaluated for power against the same representative patterns of treatment effects. Results emphasize the advantage of parsimony that is achieved by transforming multiple repeated measurements into a reduced set of mean ngful composite variables representing average levels and rates of change. The Bonferroni correction applied to the separate univariate tests provided experiment-wise protection against Type I error, produced slightly greater experiment-wise power than a multivariate test applied to the same components of the data patterns, and provided substantially greater power than a multivariate test on the complete set of original repeated measurements. The separate univariate tests provide interpretive advantage regarding locus of the treatment effects.
A two-stage mixed model analysis of repeated measurement calculates participant-specific regression slopes relating change in available measurements to associated assessment times, and then the difference between mean regression slopes in two or more treatment groups is tested for significance against the within-groups variability of the participant-specific regression slopes. It is not necessary that all participants have the same schedule or number of repeated measurements. However, when dropouts are included in an "intent to treat" analysis, the shortened treatment exposures for the dropouts substantially increase variability and reduce power of tests for differences in rates of change. Previous work has suggested that normalizing the time scale to unit length for all participants prior to fitting the individual regression equations materially reduces the power attenuation produced by dropouts. This article reports a more detailed evaluation of the enhanced robustness against dropouts that is achieved by rescaling the time dimension. The robust analysis is recognized to be equivalent to weighting ordinary least squares regression on the original time scale by the duration of treatment for each participant. Slope coefficients calculated across a shortened time span for dropouts are less stable, so they are given less weight in defining the (linear) treatment effects.
Two equations for calculating sample sizes that are required for power in testing differences in rates of change in repeated measurement designs have been presented by different authors. One equation provides support for the conclusion that increased frequency of measurements across a treatment period of fixed duration enhances power of the tests. The other equation supports the counterintuitive conclusion that increased frequency of measurements actually tends to decrease power in the presence of realistic serial dependencies in the data. Monte Carlo methods confirm that the equation providing support for the latter conclusion is accurate, whereas the alternative equation tends to underestimate sample sizes required for power in testing differences in slopes of regression lines fitted to changes in the repeated measurements across time when symmetry is absent from the covariance structure.