In modern psychological science, many researchers use psychophysical tasks to study individual differences in perception, attention, aging, and personality. Psychophysics, however, traditionally uses designs with a great many trials and few individuals. These designs are inappropriate for individual-difference studies. Here, we develop a random-intercept psychophysics that jointly models variation across trials and individuals within a Bayesian hierarchical framework. We show that the model is ideal for measuring thresholds in small-trial designs. Because the model jointly accounts for variation across trials and individuals, it provides an assessment of correlation across tasks without the pernicious attenuation caused by trial noise. The resulting correlations are accompanied by measures of uncertainty that reflect both the number of trials per individual and the number of individuals. Because the framework is Bayesian, it is flexible, and we leverage this flexibility in two ways: First, we place factor models on the thresholds themselves, demonstrating how the structure of individual differences across a battery of tasks may be assessed. Second, we develop a custom-tailored psychophysical model for assessing whether stimulation is subliminal or superliminal. The threshold divides at chance-performance form above-chance performance, and the approach serves as a principled approach for assessing truly subliminal priming.
It is popular to study individual differences in cognition with experimental tasks, and the main goal in such approaches is to analyze the pattern of correlations across a battery of tasks and measures. One difficulty is that experimental tasks are often low in reliability as effects are small relative to trial-by-trial variability. Consequently, it remains difficult to accurately estimate correlations. One approach that seems attractive is hierarchical modeling where trial-by-trial variability and variability across conditions, tasks, and individuals are modeled separately. Here we show that hierarchical models may reduce the error in estimating correlations up to 46%, but only if substantive constraint is imposed. The approach here is Bayesian, and we develop novel Bayesian hierarchical factor models for experiments where trials are nested in conditions, tasks, and individuals. The prior on covariances across tasks can either be unconstrained, in which there is little error reduction, or constrained, in which there is substantial error reduction. The constraints are: 1. There is a low-dimension factor structure underlying the covariation across tasks; and 2. All loadings are nonnegative leading to a positive manifold on correlations. We argue that both of these assumptions are reasonable in cognitive domains, and that with them, researchers may profitably use hierarchical models to estimate correlations across tasks in low-reliability settings.
Although factor models have been pivotal in understanding the relations among variables, they are difficult to apply to data from experimental designs. In experiments, data are from trials which are nested in conditions, tasks, and individuals. Data at the trial level is quite noisy, so much so that even averages on critical contrasts are subject to excessive measurement error. In this paper, we develop hierarchical factor models where the first level are linear models that account for trial noise and are parameterized with separate, latent individual-by-task scores. The second level consists of factor models on latent individual-by-task scores. Because the model is hierarchical, Bayesian analysis is convenient and conceptually straightforward. We provide development with good computational properties. The advatages of the model are that (a) correlations across tasks may be both disattenuated and localized fairly precisely; (b) uncertainty from both variability in individuals and across trials may be accurately assessed; and (c) factor structures may be recovered even in high-measurement-error contexts. Applications to visual illusions and to cognitive control are included.
People tend to judge repeated information as more likely true compared with new information. A key explanation for this phenomenon, called the illusory truth effect, is that repeated information can be processed more fluently, causing it to appear more familiar and trustworthy. To consider the function of time in investigating its underlying cognitive and affective mechanisms, our design comprised two retention intervals. Seventy-five participants rated the truth of new and repeated statements 10 min, as well as 1 week after first exposure while spontaneous facial expressions were assessed via electromyography. Our data demonstrate that repetition results not only in an increased probability of judging information as true (illusory truth effect) but also in specific facial reactions indicating increased positive affect, reduced mental effort, and increased familiarity (i.e., relaxations of musculus corrugator supercilii and frontalis) during the evaluation of information. The results moreover highlight the relevance of time: both the repetition-induced truth effect as well as EMG activities, indicating increased positive affect and reduced mental effort, decrease significantly after a longer interval.
Understanding how people covary in performance across experimental tasks is a critical component of psychometrics and individual-difference psychology. A key goal, therefore, is the accurate and precise measurement of correlation coefficients in real-world settings. The difficulty is that real-world settings in experiments contain multiple sources of variation such as those from trials, conditions, individuals, and tasks; moreover, not appropriately modeling these sources leads to asymptotically attenuated estimates with nonsense confidence intervals. In these contexts, Bayesian hierarchical models are essential. The problem addressed here is how the choice of prior affects posterior correlation distributions for correlations in real-world settings. We compare through simulation the performance on Inverse Wishart and LKJ priors across a range of settings, and find both priors do well with reasonable settings. The advantage of the Inverse Wishart is computational speed (especially in large designs); the advantage of the LKJ is greater robustness to variation in prior settings. Our recommendation to use LKJ rather than the Inverse Wishart as a default unless speed is prioritized and some scaling information about the data is known a priori.
A central element of statistical inference is good model specification where researchers specify models that capture differing theoretical position. We argue that methods of inference forcing researchers to use models that may not be appropriate for their research question are not as desirable as methods with no such constraints. We ask how posterior-predictive model assessment methods such as WAIC and LOO-CV perform when theoretical positions correspond to different space restrictions on a common parameter space. One of the main theoretical relations is nesting — where the parameter space of one model is a subset of that for another. A good example is a general model that admits any set of preferences; a nested model is one that admits only preferences that obey transitivity. We find that posterior-predictive methods fail in these cases: More constrained models are not favored even when data are compatible with the constraint. Researchers who use posterior predictive methods are forced to partition the parameter space into non-overlapping subspaces, even if these subspaces have no theoretical interpretation. Fortunately, Bayes factor model comparison accommodates overlapping models without such difficulties. We argue given that posterior predictive approaches force certain specifications that may not be ideal for scientific questions, they are less desirable in many contexts.
Although individual-difference studies have been invaluable in several domains of psychology, there has been less success in cognitive domains using experimental tasks. The problem is often called one of reliability: Individual differences in cognitive tasks, especially cognitive-control tasks, seem too unreliable. In this article, we use the language of hierarchical models to define a novel reliability measure-a signal-to-noise ratio-that reflects the nature of tasks alone without recourse to sample sizes. Signal-to-noise reliability may be used to plan appropriately powered studies as well as understand the cause of low correlations across tasks should they occur. Although signal-to-noise reliability is motivated by hierarchical models, it may be estimated from a simple calculation using straightforward summary statistics.
Stimulus duration has the opposite effect in masking and fusion tasks: longer durationsenhance performance in masking tasks but impair it in fusion tasks.Several visual theoriesexplain these and related phenomena with recourse to small temporal window wherestimuli are integrated or superimposed. Accordingly, individuals with a long temporalwindow should exibit good performance in the fusion task and poor performance in themasking task. Therefore, performance in these two tasks should result in a negativecorrelation. We tested this negative correlation and found decisive evidence to the contrary,a positive correlation (N = 21, BF = 256). People who perform well on the fusion taskalso perform well on the masking task. Hence, individual variation in a temporal windowdoes not drive individual differences in vision. Instead, we suspect the positive correlationreflects a common ability to read out from and to refresh iconic storage.
Are people who are susceptible to one illusion susceptible to others? Previous research has shown small correlations, but might small values reflect attenuation from measurement error? Data from 149 participants on 2 variants of 5 illusions were collected using an adjustment paradigm. The resulting data are of notable high quality inasmuch as there is relatively little within-subject variability and relatively much between-subject variability ($\gamma^2 \approx 1.14$, reliability $\approx .93$). Because the data are of such high quality, correlations may be estimated to high precision. In line with previous research, these cross-illusion correlations are low in value, about .22$\pm .07$. A Bayesian hierarchical analysis reveals that there is almost no attenuation from measurement error in these values. Though correlations are low, latent variable analysis reveals that the pattern among these correlations yields a single, common latent factor. This factor loads on every illusion and accounts for about 25\% of the variance in each; thus it is a {\em susceptibility to illusions} factor. We provide a novel set of statistical and graphical analyses focused on understanding the uncertainty in effects, correlations, and latent variables.
One of the most popular and influential measures of cognitive control is the antisaccade task where participants identify briefly presented and subsequently masked targets at a peripheral location. Prior to mask presentation, there is a cue at an opposing location that must be inhibited. Performance reflects the ability to inhibit We assess the degree to which antisaccade accuracy measure reflects an independent inbhibition component vs. a more generic processing-speed component by comparing performance to that in a prosaccade accuracy condition. Target durations in both tasks are adjusted with an adaptive 2-down/1-up staircase to produce constant accuracies. Individual target durations in the antisaccade condition are highly correlated to those in the prosaccade condition (r = .83). With a series of hierarchical models, we show there is but a single source of systemic variation across the conditions. We conclude that individual variation in inhibition in the antisaccade task is due to individual variation in processing speed.
The most prominent goal when conducting a meta-analysis is to estimate the true effect size across a set of studies. This approach is problematic whenever the analyzed studies have qualitatively different results, i.e. some studies show an effect in the predicted direction while others show no effect and still others show an effect in the opposite direction. In case of such qualitative differences, the average effect may be a product of different mechanisms, and therefore uninterpretable. The first question in any meta-analysis should therefore be whether all studies show an effect in the same, expected direction. To tackle this question a model with ordinal constraints is proposed where the ordinal constraint holds each study in the set. This "every study" model is compared to a set of alternative models, such as an unconstrained model that predicts effects in both directions. If the ordinal constraints hold, one underlying mechanism may suffice to explain the results from all studies, and this result could be supported by reduced between-study heterogeneity. A major implication is then that average effects become interpretable. We illustrate the model comparison approach using Carbajal et al.'s (2021) meta-analysis on the familiar-word-recognition effect, show how predictor analyses can be incorporated in the approach, and provide R-code for interested researchers. As common in meta-analysis, only surface statistics (such as effect size and sample size) are provided from each study, and the modeling approach can be adapted to suit these conditions.
Individual difference exploration of cognitive domains is predicated on being able to ascertain how well performance on tasks covary. Yet, establishing correlations among common inhibition tasks such as Stroop or flanker tasks has proven quite difficult. It remains unclear whether this difficulty occurs because there truly is a lack of correlation or whether analytic techniques to localize correlations perform poorly real-world contexts because of excessive measurement error from trial noise. In this paper, we explore how well correlations may localized in large data sets with many people, tasks, and replicate trials. Using hierarchical models to separate trial noise from true individual variability, we show that trial noise in 24 extant tasks is about 8 times greater than individual variability. This degree of trial noise results in massive attenuation in correlations and instability in Spearman corrections. We then develop hierarchical models that account for variation across trials, variation across individuals, and covariation across individuals and tasks. These hierarchical models also perform poorly in localizing correlations. The advantage of these models is not in estimation efficiency, but in providing a sense of uncertainty so that researchers are less likely to misinterpret variability in their data. We discuss possible improvements to study designs to help localize correlations.
The authors assessed a battery of number skills in a sample of over 500 preschoolers, including both monolingual and bilingual/multilingual learners from households at a range of socio-economic levels. Receptive vocabulary was measured in English for all children, and also in Spanish for those who spoke it. The first goal of the study was to describe entailment relations among numeracy skills by analyzing patterns of co-occurrence. Findings indicated that transitive and intransitive counting skills are jointly present when children show understanding of cardinality and that cardinality and knowledge of written number symbols are jointly present when children successfully use number lines. The study’s second goal was to describe relations between symbolic numeracy and language context (i.e., monolingual vs. bilingual contexts), separating these from well-documented socio-economic influences such as household income and parental education: Language context had only a modest effect on numeracy, with no differences detectable on most tasks. However, a difference did appear on the scaffolded number-line task, where bilingual learners performed slightly better than monolinguals. The third goal of the study was to find out whether symbolic number knowledge for one subset of children (Spanish/English bilingual learners from low-income households) differed when tested in their home language (Spanish) vs. their language of preschool instruction (English): Findings indicated that children performed as well or better in English than in Spanish for all measures, even when their receptive vocabulary scores in Spanish were higher than in English.
In his 1956 APA Presidential Address, Lee Cronbach called for a merging of differential and experimental psychology. One main component was the use true experiments from experimental psychology to study individual differences. True experiments had control conditions, and used statistical contrasts to define effects that were free of nuisance variation. Cronbach reasoned that such experimentally defined effects would have greater construct validity than raw performance scores. Cronbach's merger has been repeatedly attempted, but the results have been lackluster, especially for measuring individual differences in cognitive control. Here we show through simulation that the merger is difficult because experimentally-defined contrasts are too noise-prone to be useful at the individual level. As a consequence, it is difficult to uncover even simple latent structures such as clusters or factors. To explore the merger, we provide a new measure of task goodness that is invariant across experiments with different numbers of people and, most importantly, trials. This new measure is a signal-to-noise ratio; how variable people are relative to how variable noise is in repeated trials. We survey the literature to establish typical SNR levels, which are then used in simulations. Latent cluster or factor structures were only recoverable in the largest of experiments comprising hundreds of people each observing hundreds of trials in each condition for each task, and only for the simplest structures. These disappointing results serve as a warning. Although Cronbach's merger is a great idea in theory, it faces substantial hurdles in practice.
van Doorn et al. (2021) outlined various questions that arise when conducting Bayesian model comparison for mixed effects models. Seven response articles offered their own perspective on the preferred setup for mixed model comparison, on the most appropriate specification of prior distributions, and on the desirability of default recommendations. This article presents a round-table discussion that aims to clarify outstanding issues, explore common ground, and outline practical considerations for any researcher wishing to conduct a Bayesian mixed effects model comparison.
ANOVA—the workhorse of experimental psychology—seems well understood in that behavioral sciences have agreed-upon contrasts and reporting conventions. Yet, we argue this consensus hides considerable flaws in common ANOVA procedures, and these flaws become especially salient in the within-subject and mixed-model cases. The main thesis is that these flaws are in model specification. The specifications underlying common use are deficient from a substantive perspective, that is, they do not match reality in behavioral experiments. The problem, in particular, is that specifications rely on coincidental rather than robust statements about reality. We provide specifications that avoid making arguments based on coincidences, and note these Bayes factor model comparisons among these specifications are already convenient in the BayesFactor package. Finally, we argue that model specification necessarily and critically reflects substantive concerns, and, consequently, is ultimately the responsibility of substantive researchers. Source code for this project is at github/PerceptionAndCognitionLab/stat_aov2 .
Doctoral students were randomly assigned to a five-week (30-h) faculty-led writing workshop intervention, either preceded by a five-week (waiting list) control phase or followed by a five-week maintenance phase. In the workshop, students wrote together, received instruction in genres of academic writing (literature reviews, scientific articles, funding proposals, and presentations), and exchanged feedback on drafts. As a result of the workshop students enjoyed writing more, found writing easier, and gained confidence in themselves as academic writers. They felt able to write productively in shorter blocks of time, and they engaged in more short-term, medium-term, and long-term planning of their research. The intervention also caused participants to pause more frequently for reflection or positive thinking and to generate more new writing. Effects were maintained in a peer-led writing maintenance group for at least five weeks after the intervention ended. This is the first randomized controlled trial of a doctoral-level writing intervention to date and has the potential to support doctoral training in academic and scientific writing across the Social Sciences, Education, and the Humanities.
Ordinal-scale items—say items that assess agreement with a proposition on an ordinal rating scale from strongly disagree to strongly agree—are exceedingly popular in psychological research. A common research question concerns the comparison of response distributions on ordinal-scale items across conditions. In this context, there is often a lingering question of whether metric-level descriptions of the results and parametric tests are appropriate. We consider a different problem, perhaps one that supersedes the parametric-vs-nonparametric issue: When is it appropriate to reduce the comparison of two (ordinal) distributions to the comparison of simple summary statistics (e.g., measures of location)? In this paper, we provide a Bayesian modeling approach to help researchers perform meaningful comparisons of two response distributions and draw appropriate inferences from ordinal-scale items. We develop four statistical models that represent possible relationships between two distributions: an unconstrained model representing a complex, non-ordinal relationship, a nonparametric stochastic-dominance model, a parametric shift model, and a null model representing equivalence in distribution. We show how these models can be compared in light of data with Bayes factors and illustrate their usefulness with two real-world examples. We also provide a freely available web applet for researchers who wish to adopt the approach.
An increasingly popular approach to statistical inference is to focus on the estimation of effect size. Yet this approach is implicitly based on the assumption that there is an effect while ignoring the null hypothesis that the effect is absent. We demonstrate how this common null-hypothesis neglect may result in effect size estimates that are overly optimistic. As an alternative to the current approach, a spike-and-slab model explicitly incorporates the plausibility of the null hypothesis into the estimation process. We illustrate the implications of this approach and provide an empirical example.
Mixed-effects models are becoming common in psychological science. Although they have many desirable features, there is still untapped potential that has not yet been fully realized. It is customary to view homogeneous variance as an assumption to satisfy. We argue to move beyond that perspective, and to view modeling within-person variance (``noise'') as an opportunity to gain a richer understanding of psychological processes. This can provide important insights into behavioral (in)stability. The technique to do so is based on the mixed-effects location scale model. The formulation can simultaneously estimate mixed-effects sub-models to both the mean (location) and within-person variance (scale) for clustered data common to psychology. We develop a framework that goes beyond assessing the sub-models in isolation of one another, and allows for testing structural relations between the mean and within-person variance with the Bayes factor. We first present a motivating example, which makes clear how the model can characterize mean--variance relations. We then apply the method to reaction times gathered from two cognitive inference tasks. We find there are more individual differences in the within-person variance than the mean structure, as well as a complex web of structural mean--variance relations in the random effects. This stands in contrast to the dominant view of within-person variance--i.e., measurement ``error'' or ``noise.'' The results also point towards paradoxical within-person, as opposed to between-person, effects. That is, in both tasks, several people had \emph{slower} and \emph{less} variable incongruent responses. This contradicts the typical pattern, wherein \emph{larger} means are expected to be \emph{more} variable. We conclude with future directions. These span from methodological to theoretical inquires that can be answered with the presented methodology.