A growing body of methodological literature suggests that simulation studies are affected by issues similar to those underlying the replication crisis in empirical research, including inadequate reporting practices and the exploitation of researcher degrees of freedom (RDF). Open Science practices originally developed for empirical research have therefore been proposed as potential remedies. In this paper, we propose a template for transparently documenting the entire life-cycle of a simulation study. It is structured in three parts: (1) a Statement of Intent that can be used to specify and preregister hypotheses and core design choices early in the process; (2) a Piloting Documentation that records all explored simulation settings, including any choices not selected for the final study; and (3) a Final Documentation that supports honest reporting and reproducibility of the published results. Together, these components aim to improve the planning, documentation, and reporting of simulation studies. We motivate the template by discussing parallels and differences between empirical and simulation research. We argue that while exploitable RDF in empirical studies arise primarily during data analysis, simulation study RDF appear primarily during the design and piloting stages. Documenting all explored conditions therefore helps to make decisions in simulation research more transparent.
Model-based recursive partitioning (MOB; Zeileis et al., 2008) is a flexible framework for detecting subgroups of persons showing different effects in a wide range of parametric models. It provides a versatile tool for detecting and explaining heterogeneity in, for example, intervention studies. In this tutorial article, we introduce the general MOB framework. In two specific case studies, we illustrate how MOB-based methods can be used to detect and explain heterogeneity in two widely used frameworks in educational studies: (a) The generalized linear mixed model (GLMM) and (b) item response theory (IRT). In the first case study, we show how GLMM trees (Fokkema et al., 2018) can be used to detect subgroups with different parameters in mixed-effects models. We apply GLMM trees to longitudinal data from a study on the effects of the Head Start pre-school program to identify subgroups of families where children show comparatively larger or smaller gains in performance. In a second case study, we show how Rasch trees (Strobl et al., 2015) can be used to detect subgroups with different item parameters in IRT models (i.e. differential item functioning [DIF]). DIF should be investigated before using test results for group comparisons. We show how a recently developed stopping criterion (Henninger et al., 2023) can be used to guide subgroup detection based on DIF effect sizes.
Partial credit trees (PCtree) from the model-based recursive partitioning framework combine the partial credit measurement model for polytomous items with decision trees from machine learning. This method allows researchers to investigate measurement invariance by detecting differential item and differential step functioning (DIF/DSF) in a data driven way. In this manuscript, we extend PCtrees by an effect size measure for DIF/DSF in polytomous items, the partial gamma coefficient from psychometrics. We evaluate this extension of PCtrees in a series of simulation studies. Our results show that the partial gamma coefficient supports researchers in evaluating whether splits in the tree are meaningful, identifying DIF and DSF items, and can stop the tree from growing in case of negligible effect sizes. Furthermore, we assess and implement a correction for item-wise testing that is particularly crucial in longer tests. Finally, we illustrate the extension of PCtrees using data from the LISS panel to showcase its enhanced interpretability.
Machine learning models have recently become popular in the social sciences, psychology, education, or economics. However, many machine learning models lack interpretable parameters that researchers are used to from parametric models, such as linear or logistic regression. To gain insights into how the machine learning model has made its predictions, different interpretation techniques have been proposed. In this article, we review two local interpretation techniques that are widely used in machine learning: Local Interpretable Model-Agnostic Explanations (LIME) and Shapley values. LIME aims at explaining machine learning predictions in the close neighborhood of a specific person. Shapley values can be understood as a measure of predictor relevance or contribution of predictor variables for specific persons. Using two illustrative, simulated examples, we explain the idea behind LIME and Shapley values, demonstrate their characteristics, and discuss challenges that might arise in their application and interpretation. For LIME, we demonstrate how the choice of the size of the neighborhood may impact conclusions. For Shapley values, we show how they can be interpreted individually for a specific person and jointly across persons, and we compare the results to global interpretation techniques. The aim of this article is to support researchers to safely use these interpretation techniques themselves, but also to critically evaluate interpretations when they encounter the interpretation techniques in research articles.
Random forests are a nonparametric machine learning method, which is currently gaining popularity in the behavioral sciences. Despite random forests' potential advantages over more conventional statistical methods, a remaining question is how reliably informative predictor variables can be identified by means of random forests. The present study aims at giving a comprehensible introduction to the topic of variable selection with random forests and providing an overview of the currently proposed selection methods. Using simulation studies, the variable selection methods are examined regarding their statistical properties, and comparisons between their performances and the performance of a conventional linear model are drawn. Advantages and disadvantages of the examined methods are discussed, and practical recommendations for the use of random forests for variable selection are given.
Tree-based methods are being both successfully applied and critically discussed in corpus linguistics. In this article, we would like to contribute a few aspects to this discussion from a methodological point of view. These aspects include the interpretation of interaction effects in single trees and random forests, as well as more general aspects like stability and overfitting. In particular, we have conducted a simulation study to investigate an approach suggested by Gries for computing the importance of interactions in random forests more systematically than the previous literature. The evidence of this simulation study shows that, even when interaction predictors are explicitly added, the permutation variable importance is not suited for distinguishing between main effects and interaction effects or between interaction effects of different orders. We also discuss the use of partial dependence (PD) and individual conditional expectation (ICE) plots for illustrating the functional form and potential interaction effects, and other means of interpretable machine learning.
In recent years, machine learning methods have become increasingly popular prediction methods in psychology. At the same time, psychological researchers are typically not only interested in making predictions about the dependent variable, but also in learning which predictor variables are relevant, how they influence the dependent variable, and which predictors interact with each other. However, most machine learning methods are not directly interpretable. Interpretation techniques that support researchers in describing how the machine learning technique came to its prediction may be a means to this end. We present a variety of interpretation techniques and illustrate the opportunities they provide for interpreting the results of two widely used black box machine learning methods that serve as our examples: random forests and neural networks. At the same time, we illustrate potential pitfalls and risks of misinterpretation that may occur in certain data settings. We show in which way correlated predictors impact interpretations with regard to the relevance or shape of predictor effects and in which situations interaction effects may or may not be detected. We use simulated didactic examples throughout the article, as well as an empirical data set for illustrating an approach to objectify the interpretation of visualizations. We conclude that, when critically reflected, interpretable machine learning techniques may provide useful tools when describing complex psychological relationships.
This simulation study investigated to what extent departures from construct similarity as well as differences in the difficulty and targeting of scales impact the score transformation when scales are equated by means of concurrent calibration using the partial credit model with a common person design. Practical implications of the simulation results are discussed with a focus on scale equating in health-related research settings. The study simulated data for two scales, varying the number of items and the sample sizes. The factor correlation between scales was used to operationalize construct similarity. Targeting of the scales was operationalized through increasing departure from equal difficulty and by varying the dispersion of the item and person parameters in each scale. The results show that low similarity between scales goes along with lower transformation precision. In cases with equal levels of similarity, precision improves in settings where the range of the item parameters is encompassing the person parameters range. With decreasing similarity, score transformation precision benefits more from good targeting. Difficulty shifts up to two logits somewhat increased the estimation bias but without affecting the transformation precision. The observed robustness against difficulty shifts supports the advantage of applying a true-score equating methods over identity equating, which was used as a naive baseline method for comparison. Finally, larger sample size did not improve the transformation precision in this study, longer scales improved only marginally the quality of the equating. The insights from the simulation study are used in a real-data example.
A family of score-based tests has been proposed in recent years for assessing the invariance of model parameters in several models of item response theory (IRT). These tests were originally developed in a maximum likelihood framework. This study discusses analogous tests for Bayesian maximum-a-posteriori estimates and multiple-group IRT models. We propose two families of statistical tests, which are based on an approximation using a pooled variance method, or on a simulation approach based on asymptotic results. The resulting tests were evaluated by a simulation study, which investigated their sensitivity against differential item functioning with respect to a categorical or continuous person covariate in the two- and three-parametric logistic models. Whereas the method based on pooled variance was found to be useful in practice with maximum likelihood as well as maximum-a-posteriori estimates, the simulation-based approach was found to require large sample sizes to lead to satisfactory results.
A family of score-based tests has been proposed in recent years for assessing the invariance of model parameters in several models of item response theory (IRT). These tests were originally developed in a maximum likelihood framework. This study discusses analogous tests for Bayesian maximum-a-posteriori estimates and multiple-group IRT models. We propose two families of statistical tests, which are based on an approximation using a pooled variance method, or on a simulation approach based on asymptotic results. The resulting tests were evaluated by a simulation study, which investigated their sensitivity against differential item functioning with respect to a categorical or continuous person covariate in the two- and three-parametric logistic models. Whereas the method based on pooled variance was found to be useful in practice with maximum likelihood as well as maximum-a-posteriori estimates, the simulation-based approach was found to require large sample sizes to lead to satisfactory results.