
Intensive longitudinal data (ILD) are commonly used to study within-person processes, and psychometric methods for computing within-person reliability with ILD have recently appeared in the methodological literature. The dense, frequent nature of data collection in ILD often leads to copious amounts of missing data-recent reviews find that 30%-40% missing data rates are typical. However, given the nascent state of psychometrics for ILD, the intersection of psychometrics and missing ILD has yet to be explored despite the pervasiveness of missing ILD in empirical studies. In addition, given that ILD are frequently used to study sensitive topics such as mental health and substance use, there is a heightened risk of missing not at random (MNAR) data. That is, the reason data are missing is associated with what the value would have been (e.g., depression responses are missing when a person is momentarily too depressed to respond). The goal of this article is to (a) clarify how missing values affect estimates of within-person reliability with ILD and (b) to extend one recently proposed method (the measurement error autoregressive model) with a Diggle-Kenward selection process to evaluate sensitivity to a potential MNAR mechanism. Simulations in the article find that the accuracy of within-person reliability estimates from standard approaches can deteriorate when missing data are present, but the proposed model can better recover population within-person reliability-even with a large amount of MNAR data-under the conditions and missingness mechanisms studied. Limitations and extensions to other methods for computing within-person reliability are also discussed. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
Missing data are pervasive in psychological and educational assessments. Naive remedies, including listwise deletion and item-mean imputation, often degrade research validity and misinform subsequent decisions. Recent advances in artificial neural networks have demonstrated their efficacy in prediction-related tasks by using observed features to infer unknown values. Building on this potential, we propose the columnwise neural imputation (COLNI) algorithm to impute missing ordinal responses in psychometric data. Simulation studies demonstrated that, when benchmarked against conventional methods, COLNI more accurately recovered item means, inter-item correlations, and person and item parameters under the multidimensional graded response model. We further evaluated COLNI using data from the Short Dark Triad test, confirming its effectiveness in a multidimensional empirical setting. We conclude with implementation guidelines and avenues for refining and extending this artificial neural network-based imputation approach in future research. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
There is an increasing call for individualized treatment rules, which leverage individual patient characteristics to recommend treatments or interventions, tailoring recommendations based on their covariates. This is particularly of interest for the care of conditions such as depression, for which many treatment options are available with similar average effectiveness but with large heterogeneity in individual responses. In parallel, there has been a growing interest in machine learning methods for causal inference and variable selection. We compared several strategies for variable adjustment in a dynamic marginal structural modeling approach to estimating an optimal individualized treatment rule and investigated the performance of outcome adaptive lasso, group lasso and doubly robust estimation, double-index propensity score, the high-dimensional balancing propensity score, and the causal ball lasso as variable selection methods for the propensity score. Our results demonstrate that these all provided similar unbiased estimates. However, methods differed in their ability to exclude extraneous variables and in computational burden. We found statistical efficiency is gained when variable selection approaches were for the propensity score were used and by including variables in the outcome model. We applied all methods to determine the optimal treatment rule, treating with either selective serotonin reuptake inhibitors or serotonin and norepinephrine reuptake inhibitors, for unipolar depression in individuals aged 13 years and older. This analysis, which used electronic health records from 74,058 Kaiser Permanente Washington patients with a new antidepressant dispensing between 2008 and 2018, suggested tailoring treatment based on baseline symptom severity did not impact symptom severity 6 months later. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
Procedural knowledge space theory (PKST) integrates knowledge space theory and problem space theory to model human problem-solving. This study extends PKST to competitive, two-player games and presents the first empirical application of homomorphisms to derive abstract representations of complex problem spaces. Using tic-tac-toe and its numerical variant pick15 as test cases, the study demonstrates how PKST can model complex competitive scenarios effectively. An empirical study involving 148 participants validated these PKST extensions by analyzing their performance on pick15 tasks. Problem and knowledge spaces were derived at multiple levels of abstraction and tested using the Markov solution process model. Results indicated that abstract representations provided a better fit for participants with prior experience in tic-tac-toe, suggesting that expertise promotes the use of generalizable strategies. This work provides a rigorous proof of concept for applying homomorphisms to test theoretical predictions on problem-solving tasks. By linking abstract representations, like tactics, to problem-solving performance, the proposed methodology offers a concrete approach to evaluating cognitive theories of problem-solving and decision-making. These findings highlight the potential of PKST to model broader cognitive phenomena in interactive contexts. The study emphasizes the practical utility and flexibility of PKST for formalizing and empirically testing hypotheses, extending its relevance to diverse domains of cognitive science. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
Post hoc comparisons following a significant analysis of variance (ANOVA) are typically conducted using pairwise methods such as Tukey's honestly significant difference or Holm-Bonferroni corrections. Although these procedures control family-wise error rates, they can produce nontransitive patterns of equivalence that do not correspond to any coherent partition of the factor levels. We propose a sequential fusion procedure that takes a modeling perspective: Starting from the saturated ANOVA model, the algorithm iteratively fuses the pair of groups with the most similar means, testing each fusion via a likelihood-ratio test using the ANOVA error term. A calibrated sequence of critical values ensures family-wise error rate control at the nominal level (e.g., 5%). Through extensive Monte Carlo simulations across balanced designs with 3-50 groups and 5-50 observations per group, we derive a simple approximation for the calibration factor as a function of the number of groups and sample size. Comparisons with Tukey's honestly significant difference, Holm-Bonferroni, Benjamini-Hochberg, and Scheffé procedures demonstrate that the fusion cascade achieves comparable or superior power while guaranteeing interpretational coherence: The output is always a single, transitive partition of groups. We illustrate the method's practical utility by reanalyzing data from Moreno and Mayer (2000) on multimedia learning, demonstrating how factorial designs are naturally handled by treating them as one-way ANOVA on cell means, with interaction structure revealed directly as a partition. Simulation studies confirm that the fusion cascade consistently identifies exact partitions under well-separated conditions, whereas pairwise methods produce nontransitive patterns in the majority of cases. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
Analyzing zero-inflated count data in single-case experimental designs presents a significant analytical challenge. Although traditional zero-inflated generalized linear mixed models are available, they estimate a conditional treatment effect given an individual coming from the count process, which often mismatches the applied researcher's interest in the overall effect of an intervention for each subject. These conditional models also face interpretational and estimation challenges within the small-sample context of single-case experimental designs. This study introduces and evaluates a Bayesian marginalized zero-inflated Poisson (mZIP) model with random effects. This framework reparameterizes the model to estimate the marginal intervention effect, aligning the statistical estimand with the typical research question. A Monte Carlo simulation study was conducted to evaluate the mZIP model's performance. We compared its performance with other methods that also target the marginal treatment effect: the log response ratio (LRR), a Poisson generalized linear mixed-effects model (GLMM), and a negative binomial GLMM. Simulation results indicate that the Bayesian mZIP model consistently recovers unbiased estimates of the marginal treatment effect and provides reliable statistical inference. The LRR, Poisson GLMM, and negative binomial GLMM produced biased estimates for the marginal effect under conditions with smallest sample size and highest zero-inflation rate. They also suffered from low coverage rate, making their inferential statistics invalid. The LRR indicated lower statistical power than the mZIP model. We apply the mZIP model to two empirical data sets to illustrate its application and interpretation. Finally, we discuss the distinction between conditional and marginal estimands, as well as the limitations and future directions. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
When planning a study that uses multiple measurement trials, both the number of participants and the number of measurement trials are fundamental design elements. In multiple measurement trial studies, each participant is repeatedly exposed to a particular treatment condition, and the dependent variable is the average of multiple measurement trials. Increasing the number of measurement trials will increase the reliability of the dependent variable, which will reduce the sample size needed to obtain a confidence interval with desired precision or a hypothesis test with desired power. In studies where the dependent variable is an average of multiple measurement trials, the investigator may want to compare the required sample size for different numbers of measurement trials. Closed-form sample size formulas are derived to determine the sample size needed for desired confidence interval precision or statistical power, given a specified number of measurement trials. Sample size formulas are derived for the two-group design and the paired-samples design and then extended to the case of general linear contrasts in between-subjects and within-subjects designs. R functions for each sample size formula are given in the online supplemental materials. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
The ability of a student can be conceptualized as either a continuously varying entity (e.g., conventional analysis using dichotomous item response theory [IRT] models; Lord & Novick, 1968) or a bundle of latent classes (e.g., cognitive diagnostic models [CDMs]; Rupp et al., 2010; von Davier & Lee, 2019). This article builds on recent efforts to focus on predictive differences between measurement models in an attempt to examine the degree to which such approaches, which utilize quite distinctive notions regarding the nature of ability, produce different predictions of response behavior in the real world. We first present two simulation studies in which data are generated from variants of CDMs that differ in sample size, attribute hierarchical structures, and attribute estimation methods. We illustrate that, as we would expect given that they were used to generate data, CDM-based predictions uniformly outperform those of IRT models. We then compare the performance of CDM- and IRT-based approaches across 11 empirical data sets previously analyzed using CDMs. Our findings indicate that overfitting is a pervasive issue across CDM-based predictions, particularly with the generalized deterministic inputs, noisy "and" gate model. Furthermore, in no case does the CDM show superior performance when using the maximum a posteriori estimator, and only six out of 11 data sets show improved model fit for CDMs over the two-parameter logistic model when using the marginal mastery probabilities estimator. Researchers and practitioners may need to balance the diagnostic appeal of CDMs with the fact that their complexity can come at the cost of predictive accuracy. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
Model evaluation is crucial in psychological methodology, particularly in the context of Bayesian latent variable models. Bayesian evaluation methods require integrating out latent variables to compute marginal likelihoods for information criteria and integrating both latent variables and model parameters to compute fully marginal likelihoods for Bayes factors. These processes can be computationally demanding, as likelihoods in many latent variable models are intractable. Moreover, the role of latent variables in model evaluation has been insufficiently addressed in applied research, likely due to a lack of practical guidance and user-friendly software. This tutorial fills these gaps by (a) offering step-by-step instructions for approximating marginal likelihoods using an efficient numerical method (i.e., adaptive Gauss-Hermite quadrature), (b) developing an R package, bleval, designed specifically for computing information criteria and fully marginal likelihoods in Bayesian latent variable models, and (c) demonstrating the application of this package to various empirical scenarios, including structural equation models, item response theory models, and multilevel models. The empirical examples also investigate practical issues such as the sensitivity of model evaluation results to the number of quadrature nodes, the performance of the numerical quadrature method in high-dimensional latent variable models, and the handling of latent class variables in Bayesian mixture models. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
The growing use of intensive longitudinal studies has heightened demand for advanced modeling approaches to better understand temporal dynamics of individual change. Although widely used multilevel vector autoregressive models can capture both intraindividual dynamics and interindividual differences, current implementations are restricted to two-level analyses. This can be problematic, as three-level nested structures are common in practice-for instance, individuals may be further clustered within higher level units, or weekday assessments are nested within weeks that are further nested within individuals. Existing two-level frameworks cannot adequately model such hierarchies, potentially leading to incorrect estimation and inferences about intraindividual dynamics. To address this issue, we introduce a three-level vector autoregressive modeling framework and provide corresponding Bayesian estimation algorithms. We evaluate the estimation performance of three key model specifications-unconditional univariate, conditional univariate (with Level 2 and Level 3 covariates), and unconditional bivariate-through a set of simulation studies. By varying sample sizes at all three levels and key population parameters, these simulations offer initial evidence on how sample sizes across levels influence model performance. An empirical example using ratings of children's daily emotional states during their first kindergarten month illustrates the three-level vector autoregressive framework's utility. We conclude by summarizing key findings, offering practical guidance for application, and pointing out limitations and future directions. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
In multilevel analysis, Level-1 dichotomous predictors (e.g., treatment vs. control, minority vs. majority; D) are commonly used to examine the existence of average within-cluster mean differences. To address this research question, we examine the best practices for coding D, considering within-cluster group compositions and between-cluster heterogeneity in them, as well as possible between-group heterogeneity in residual variances. We compare traditional dummy coding and unweighted effect coding with two proposed methods: overall and within-cluster weighted effect coding. Each is combined with cluster-mean centering to clearly define the within-cluster mean difference. Analytical results show that the intraclass correlation of D (ICC-D) substantially influences differences across coding methods regarding test results for the average within-cluster effect (group mean difference). When ICC-D is zero, all strategies yield identical test results (i.e., the significance tests based on the point estimates and standard error estimates) for the average within-cluster effect. However, as the ICC-D increases, test statistics (e.g., the t statistic) decrease for all coding strategies. An empirical study and simulations further confirm that ICC-D strongly influences results. Based on the simulations, we recommend dummy coding, unweighted effect coding, and overall weighted effect coding, which generally provide comparable or higher power than within-cluster weighted effect coding. The power contrast is most pronounced when both ICC-D and Level-1 residual variance heterogeneity increase, with the difference in power reaching up to 40%. Furthermore, cluster-mean centering not only disaggregates within- and between-cluster effects but also achieves generally reasonable Type I error rates even when homogeneity assumptions are not met. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
Hierarchical models for ordinal responses, in which responses are modeled successively by partitioning groups of categories into finer subgroups are proposed. These partitions reflect conceptually meaningful distinctions among categories. Such models are particularly well suited for Likert items, which typically differentiate between disagreement, agreement, and, in some cases, a neutral category. The hierarchical framework offers a parsimonious representation of predictor effects and often provides a better fit than traditional ordinal models. It also enables the investigation of dispersion effects, that is, systematic tendencies of respondents to prefer either extreme categories or middle categories, independently of the substantive content. In addition to specialized fitting tools for ordinal models, we provide a more general procedure that can be used to fit any hierarchically structured model. The practical use of these methods is demonstrated through illustrative examples. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
Although previous research has described that intervention effects vary across replication studies, less effort has been devoted to identifying causes of this effect heterogeneity with regard to differences in study implementations. However, knowing in which way study characteristics (such as population, measurement instrument, setting, or treatment implementation) impact the study results may not only help to better infer the impact of research practices but also provide evidence for theory building. Causal effects can be easily identified if all study characteristics but the one under investigation are kept constant across two studies. This is, however, not always possible in practice and unintended differences between the studies to be compared may confound the relationship of the study characteristic of interest and the treatment effect. In this article, we present a statistical approach for identifying effects of study characteristics on study-specific treatment effects from randomized experiments in cases in which unintended differences in study implementation across studies cannot be prevented. We present formal definitions of the causal effects of interest, identification assumptions, and derive respective causal estimands. The assumptions can more likely be fulfilled in prospective replication studies or many-lab studies, where researchers have more control over design and measurement of covariates in both studies. We also provide ways to test the assumptions and illustrate consequences of not meeting the assumptions. The approach is illustrated using an empirical example on the imagined intergroup contact effect in social psychology. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
In both single-level and multilevel regression analyses, researchers commonly report R-squared to quantify the proportion of outcome variance that is explained by a model or its component parts. It is well-established that the classic estimator of R-squared for single-level regression (via ordinary least squares estimation) is biased upward, and thus an adjusted version of R-squared is commonly reported. In the multilevel modeling literature, however, there is inherently more complexity in defining R-squared, and the possibility of upward bias in estimation has received very little attention, and thus researchers almost never report adjusted versions of measures. The purposes of this article are to (a) provide a pedagogical overview of the logic and computation of adjusted R-squared in single-level regression and show how this same logic can be extended to multilevel models; (b) evaluate the degree to which R-squared measures developed specifically for multilevel contexts are susceptible to upward bias, both analytically via expectation algebra and empirically via simulation; and (c) provide and evaluate a set of adjustments that correct for this bias. We ultimately show that the proposed adjustments do, in fact, yield less bias than the unadjusted versions of measures and discuss factors affecting the discrepancy between adjusted and unadjusted versions, as well as the differential impact such factors have across different measures. We also discuss software with which researchers can compute adjusted measures and illustrate their computation with an empirical example. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
Time series (or intensive longitudinal) data are now widely used in psychological science. These data allow researchers to study the dynamics of human functioning at an unprecedented level of granularity. Capturing these dynamics requires appropriate statistical models. One powerful class of models for such data is the hidden Markov model (HMM), which captures processes that switch between a number of latent states over time, each associated with a different probability distribution. Unlike the currently predominant linear models, such as the vector autoregressive model, HMMs can capture more complex behaviors. They can characterize multiple equilibrium states (for instance, manic and depressed states as seen in bipolar disorder) and quantify how the system transitions between them. HMMs can also fit common empirical data patterns such as highly skewed or multimodal distributions. Despite their effectiveness in modeling within-person dynamics, HMMs are not widely used. We see three reasons: (1) limited familiarity among researchers, (2) only recent availability of software for multilevel HMMs, and (3) the relative complexity of estimating, analyzing, and reporting HMMs. Here, we address these issues by providing a gentle introduction to HMMs for researchers working with time series data in psychology, including a fully reproducible tutorial that walks the reader through each of the steps of analyzing psychological time series data using the R-package mHMMbayes. We hope our tutorial helps researchers working with time series data with adding (multilevel) HMMs into their methodological toolbox. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
Standard two-level SEM models require access to raw data. In this study we propose a method for fitting two-level SEMs using meta-analytic structural equation modelling (MASEM) on sum-mary statistics. Although our focus is on the meta-analytic case of individuals nested in studies, the approach could be applied to any two-level structure. An illustration using PISA data showed that fitting a two-level model on individual data versus fitting a two-level model using summary statistics produces essentially identical results. One advantage of the MASEM approach is that it is straightforward to model heterogeneity in the within-cluster covariances. Using simulated da-ta, we demonstrated that the MASEM method performed well in such heterogeneous conditions. In contrast, two-level SEM analyses resulted in inflated Type I errors based on the test statistic and a significant underestimation of the parameters' standard errors. Through an empirical ex-ample, we show how two-level SEM through the MASEM method can be employed to evaluate measurement invariance across all studies and how the model can be expanded to incorporate study-level variables that explain some of the variation in parameters across studies. Implications and limitations of the proposed method are discussed, and directions for future research are pro-vided.
Ordinary least squares (OLS) estimation, which is frequently applied in psychology, assumes constant variance of errors across predictor levels. This assumption is known as homoscedasticity, whereas its violation is referred to as heteroscedasticity. In categorical predictors, heteroscedasticity can be quantified by calculating the ratio of variances across groups. For continuous predictors, diagnostic residual plots are often used to assess whether the assumption had been met, but there is currently no measure that can quantify the amount of heteroscedasticity in an interpretable way. In this study, we have developed and evaluated a measure that constructs a quantile locally weighted scatterplot smoothing interval (QLI) around the residuals and estimates the linear, quadratic, cubic, and quartic change in the width of this interval as a function of the predictor or the fitted values. Furthermore, we evaluated simple linear models under different patterns of heteroscedasticity in a simulation to provide benchmark values of QLI estimates associated with inadequate control over false-positive results, loss of power, and loss of coverage probability of confidence intervals. The QLI method provided consistent estimates of trends for models with 60 or more cases, and this was true across variance patterns. We discuss QLI-generated estimates in relation to performance of OLS linear models. Finally, we present an example of how to apply the QLI method to quantify heteroscedasticity and how to interpret the estimates it provides, focusing on the implications for the OLS analysis. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
We introduce a novel approach to analyzing responses of individuals to items at two or more time points. Existing longitudinal item response models do not capture interactions among individuals and items that evolve over time. We construct time-varying interaction maps with a view to capturing and visualizing time-varying interactions among individuals and items in a low-dimensional Euclidean space. A time-varying interaction map provides a window into the strengths and weaknesses of an individual on specific items, in addition to tracking changes in the underlying trait of the individual. We provide a data-driven Bayesian approach to determining whether time-varying interaction maps have added value, along with Bayesian methods for learning time-varying interaction maps from observed responses. We present multiple simulation and empirical studies to showcase the merits of time-varying interaction maps, including applications to behavior ratings of children and problem-solving assessments of students. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
Multivariate analysis of variance (MANOVA) has a long history of use in psychological science, is a staple of many advanced statistics textbooks and classes, and remains widely used in diverse areas of psychology. However, the way in which it is typically explained and taught relies on multivariate concepts, terms, and procedures that may be highly nonintuitive for many students, teachers, and applied researchers. This may limit many individuals' ability to learn, teach, interpret, use, and communicate MANOVA effectively and comfortably. Fortunately, MANOVA can be understood via a largely univariate perspective that is likely intuitive and accessible for many who are interested in MANOVA. Although other sources allude to this alternative perspective, those allusions are minimal and are not comprehensive, systematic, or illustrated in accessible ways. Moreover, no existing sources illustrate how, or even whether, this perspective applies to factorial MANOVA. The current tutorial explains and illustrates a largely univariate framework for understanding both one-way and factorial MANOVA. From this perspective, MANOVA begins with combinations of dependent variables, followed by a univariate analysis of variance (ANOVA) conducted on each combination, and by aggregation of effect sizes obtained from those ANOVAs to obtain multivariate effect sizes and multivariate inferential statistics. Thus, all key MANOVA results can be interpreted in relation to familiar results emerging from univariate ANOVA. This alternative perspective may help many students, teachers, and researchers gain a more intuitive understanding of how MANOVA works, what its results mean, and when/how to use it effectively. (PsycInfo Database Record (c) 2026 APA, all rights reserved).