The Same-Different task presents two stimuli in close succession and participants must indicate whether they are completely identical or if there are any attributes that differ. While the task is simple, its results have proven difficult to explain. Notably, response times are characterized by a fast-same effect whereby Same responses are faster than Different responses even though identical stimuli should be exhaustively processed to be accurate. Herein, we examine a little more than a quarter million response times (N = 255,744) obtained from 327 participants who participated in one of 14 variants of the task involving minor changes in the stimuli or their durations. We performed distribution fitting and analyzed estimated parameters stemming from the ex-Gaussian, lognormal, and Weibull distributions to infer the cognitive processing characteristics underlying this task. The results exclude serial processing of the stimuli and do not support dual-route processing. The fast-same effect appears only through a shift of the entire response time distributions, a feature impossible to detect solely with mean response time analyses. An attention-modulated process driven by entropy may be the most adequate model of the fast-same effect. (PsycInfo Database Record (c) 2023 APA, all rights reserved).
The same-different task is a classic paradigm that requires participants to judge whether two successively presented stimuli are the same or different. While this task is simple, with results that have been replicated many times, response times (RTs) and accuracy for both same and different decisions remain difficult to model. The biggest obstacle in modeling the task lies within its effect referred to as the fast-same phenomenon whereby participants are much faster at responding "same" than "different," while most standard cognitive models predict the opposite. In this study, we investigated whether this effect is the result of identity priming activated by the first stimulus. We ran four variants of the same-different task in which identity priming is intended to be attenuated or cancelled in half of the trials. Results for all four variants show that a complete visual match between both stimuli is necessary to observe a fast-same effect and that hampering this relation attenuates same RTs while different RTs remained relatively unchanged. (PsycInfo Database Record (c) 2022 APA, all rights reserved).
Mean comparisons can be an abstract concept for students when they first encounter it in the classroom. Students must grasp not only the notions underlying the comparison but also the concepts surrounding inference. Students must also have a definite understanding of the relationship between sample and population. Traditionally, we have taught this concept using a variety of tools such as paint-by-number recipes or decision trees, and students typically put into practice their learnings by analysing datasets that are limited in size and novelty. Nevertheless, the actual concept may not be grasped if the analysis is only ever run as a routine. This statistical vignette aims to help students understand mean comparison and differences between groups by creating intuitions that are based in observing the overlap between group distributions. The activities presented within this vignette also intend to showcase the limitations of using p-value alone and to leverage effect size when addressing the importance of the difference between means. Using random number generators facilitate this process.
Plotting the data of an experiment allows researchers to illustrate the main results of a study, show effect sizes, compare conditions, and guide interpretations. To achieve all this, it is necessary to show point estimates of the results and their precision using error bars. Often, and potentially unbeknownst to them, researchers use a type of error bars—the confidence intervals—that convey limited information. For instance, confidence intervals do not allow comparing results (a) between groups, (b) between repeated measures, (c) when participants are sampled in clusters, and (d) when the population size is finite. The use of such stand-alone error bars can lead to discrepancies between the plot’s display and the conclusions derived from statistical tests. To overcome this problem, we propose to generalize the precision of the results (the confidence intervals) by adjusting them so that they take into account the experimental design and the sampling methodology. Unfortunately, most software dedicated to statistical analyses do not offer options to adjust error bars. As a solution, we developed an open-access, open-source library for R— superb —that allows users to create summary plots with easily adjusted error bars. Keywords error bars , confidence interval , plots , results , statistics , standard error , within-subjects designs , sampling method , research methods , open materials
Plotting the data of an experiment allows researchers to illustrate the main results of a study, show effect sizes, compare conditions, and guide interpretations. To achieve all this, it is necessary to show point estimates of the results and their precision using error bars. Often, and potentially unbeknownst to them, researchers use a type of error bars—the confidence intervals—that convey limited information. For instance, confidence intervals do not allow comparing results (a) between groups, (b) between repeated measures, (c) when participants are sampled in clusters, and (d) when the population size is finite. The use of such stand-alone error bars can lead to discrepancies between the plot’s display and the conclusions derived from statistical tests. To overcome this problem, we propose to generalize the precision of the results (the confidence intervals) by adjusting them so that they take into account the experimental design and the sampling methodology. Unfortunately, most software dedicated to statistical analyses do not offer options to adjust error bars. As a solution, we developed an open-access, open-source library for R—superb—that allows users to create summary plots with easily adjusted error bars.
GRD is a popular tool to genenrate random data on the fly. It is most useful in statistic classes where the students can generate with a single short syntax, or using a graphical interface, random data that differs on every run but yet can implement effect sizes, outliers, etc. With the new versions of SPSS (version 27 and above) which is now using a new version of Python, it was necessary to upgrade the extension. Here, GRD 2.1 is presented which works with SPSS 27 and above.
The diffusion model is useful for analyzing data from decision making experiments as it gives information about a dataset that regular statistical tests cannot, including: the rate of processing, the encoding and motor response times, and decision thresholds. The EZ diffusion model is a restricted version of the diffusion model with some parameter variability set to zero, allowing for quicker analyses. Here we describe the EZ diffusion model-including how it was derived mathematically- the measurement units of the parameters, and how it can be generalized to starting points other than the mid-point. We also show how its parameters can be estimated using computer software (the model is available with many software programs such as R and Excel, to which we add SPSS and a Mathematica code). Finally, an EZ analysis was run on one dataset obtained from a "Same"-"Different" experiment.
Statistics courses prove to be a common difficulty to many social sciences students. To address this problem, we developed a tool in the R programming language (R Core Team, 2018) that can be used to easily and quickly generate data to be analysed. The Generator of Random Data or GRD, allows the user to build datasets according to any design type, both within and between, with or without statistical effects and/or correlations. By default, GRD creates normally distributed data, but any type of distribution defined in R can be specified. With GRD, it is possible to generate samples of any size and see the benefit of larger sample sizes on the precision of statistical measures. The students of statistics can acquire better skills in analyzing custom-made datasets, skipping long data-acquisition processes. They can also experience first-hand concepts such as statistical power and type-I and type-II errors. Each sample being different, students can appreciate randomness at the tips of their fingers.
Since its inception, Systems Factorial Technology (SFT; Townsend & Nozawa, 1995) has been used alongside many research paradigms to detect the characteristics underlying a cognitive process. Here, we show how thresholds variability in a coactive architecture can result in an ambiguous diagnosis even when all SFT assumptions are met. We implemented two independent race models: the well-known Linear Ballistic Accumulator (LBA; Brown & Heathcote, 2008) and a discrete accumulator model with varying thresholds (DAVT), a suitable model for demonstration purposes. When threshold variability increases in both models, all architectures other than coactive can be correctly identified by SFT. The coactive SIC curve is affected by the magnitude of the variability and converges towards a parallel self-terminating SIC curve. To avoid possible misdiagnoses, we show the importance of exhausting the entire SFT toolbox, including the capacity curve. We also present the SIC centerline, which can be used to discriminate architectures when threshold variability is suspected.
Discrimination decisions are at the forefront of human cognition. For this reason, many different types of models aim to predict how they are made. In this research, we compared the discrimination capabilities of a Recurrent Associative Memory (RAM) with the predictions of an accumulator model to show that, although the discrimination processes of both model classes differs, both make similar predictions regarding trends in the results. We did this by measuring the performances of a RAM within the context of a discrimination task using different stimuli (i.e., letters and randomly generated stimuli) and fitting the obtained results with an accumulator model possessing a coactive architecture. The experimental conditions varied with regard to the correlation between the tested stimuli, the amount of redundancy of the stimuli used in a trial, and the number of total stimuli presented to the network in the learning phase. Results showed that high inter-stimulus correlation led to slower recall speed, and that low redundancy also resulted in slower recall speed. Results also indicated that an increased number of exemplars contained in the network's memory increased recall speed for the letter stimuli but randomly generated stimuli received no apparent benefits. Ultimately, exploring neural networks and accumulator models jointly provides a broader and deeper understanding of the cognitive processes behind discrimination decisions.
Les statistiques sont une matière notoirement difficile à enseigner aux étudiants des sciences humaines. L’anxiété statistique, une forme d’anxiété bien documentée chez ces étudiants, est présente dès le début du cours et explique donc une partie des difficultés rencontrées par ces étudiants. Cependant, nous croyons que l’anxiété statistique est la conséquence plutôt que la source de ces difficultés. Pour enclencher une discussion qui, à terme, pourrait bonifier la façon d’enseigner les statistiques dans les sciences humaines, nous examinons ici trois causes qui pourraient expliquer cet anxiété. D’une part, les humains ont très peu d’intuition du hasard (comme le montrent les études sur les jeux d’argent dans lesquels des conceptions erronées sont dévoilées), et donc, les notions de distribution et d’échantillonnage restent des concepts opaques pour plusieurs. De plus, certains concepts statistiques reposent sur des raisonnements « méta-statistiques » dans lesquels il faut concevoir des statistiques sur des statistiques. Finalement, la notion de prise de décision dans un contexte où l’information est partielle et incertaine est souvent mal comprise. Dans ce texte, nous précisons ces trois difficultés et suggérons des recommandations pour en amoindrir les conséquences. Cependant, les pistes présentées ici nécessitent d’être validées par des études formelles; ce texte se veut avant tout un catalyseur de discussions visant l’amélioration de l’état de nos classes de méthodes quantitatives.
The study of mental processes is at the forefront of research in cognitive psychology. However, the ability to identify the architectures responsible for specific behaviors is often quite difficult. To alleviate this difficulty, recent progress in mathematical psychology has brought forth Systems Factorial Technology (SFT; Townsend and Nozawa, 1995). Encompassing a series of analyses, SFT can diagnose and discriminate between five types of information processing architectures that possibly underlie a mental process. Despite the fact that SFT has led to new discoveries in cognitive psychology, the methodology itself remains far from intuitive to newcomers. This article therefore seeks to provide readers with a simple tutorial and a rudimentary introduction to SFT. This tutorial aims to encourage newcomers to read more about SFT and also to add it to their repertoire of analyses.
A group distribution is a synthesis of a set of individual distributions. To be adequate, a method for creating group distributions should not introduce characteristics that are not present in the individual distributions and preserve those that are present. A method occasionally used is quantile averaging (sometimes called vincentizations), applied generally to response time distributions. However, it is shown here using quantile-quantile plots on empirical response times that this method is inadequate. As shown by Thomas and Ross (1980, Journal of Mathematical Psychology), to solve this problem, quantile averaging can be generalised using an appropriate nonlinear transformation of the data. Here we argue that the correct transformation is the log transform of response times to which the base response time has been removed. Equivalently, the geometric mean of the quantiles can be used. We first propose 4 estimates of the base response times. We next examine empirical data in a same-different task, in a redundant-attribute target detection task and in a visual search task. The results show that this approach is appropriate to construct group distributions. It can be used to aggregate distributions over multiple participants, over multiple sessions of training for a given participant, or both. (PsycINFO Database Record
Statistical analyses have grown immensely since the inception of computational methods. However, many quantitative methods classes teach sampling and sub-sampling at a very abstract level despite the fact that, with the faster computers of today, these notions could be demonstrated live to the students. For this reason, we have created a simple extension module for SPSS that can sub-sample and Bootstrap data, GSD (Generator of Sub-sampled Data). In this paper, we describe and show how to use the GSDmodule as well as provide short descriptions of both the subsampling and Bootstrap methods. In addition, as this article aims to inspire instructors to introduce these concepts in their statistics classes of all levels, we provide three short exercises that are ready for curriculum implementation.
The GRD extension command for SPSS (Harding & Cousineau, 2014) has been used in a variety of applications since its inception. Ranging from a teaching tool to demonstrate statistical analyses, to an inferential tool used to find critical values instead of looking into a z-table, GRD has been very well received. However, some users have requested other data generation components that would make GRD a more complete extension command: the possibility to add contaminants to the generated dataset as well as the ability to generate correlated variables. Another component we added is a graphical user interface (or GUI) that makes GRD accessible through the drop-down menus in the SPSS Data Editor window. This GUI allows users to generate a simple dataset by entering parameters in dedicated fields rather than writing out the full script. Finally, we devised a small series of exercises to help users get acquainted with the new sub-commands and GUI.
The Pearson skew is a measure of asymmetry of a distribution, based on the difference between the mean and the median of a distribution. Here we show how to calculate the Pearson skew, estimate its standard error and the confidence interval. The derivation is based on a population following a normal distribution. Simulations explored the validity of this expression when the normality assumption is met in comparison to when the normality assumption is not met. The standard error of the Pearson skew revealed very robust in case of non-normal populations, compared to the Fisher Skew as presented in Harding, Tremblay & Cousineau (2014).
Characteristics of a population are often unknown. To estimate such characteristics, random sampling must be used. Sampling is the process by which a subgroup of a population is examined in order to infer the values of the population's true characteristics. Estimates based on samples are approximations of the population's true value; therefore, it is often useful to know the reliability of such estimates. Standard errors are measures of reliability of a given sample's descriptive statistics with respect to the population's true values. This article reviews some widely used descriptive statistics as well as their standard error estimators and their confidence intervals. The statistics discussed are: the arithmetic mean, the median, the geometric mean, the harmonic mean, the variance, the standard deviation, the median absolute deviation, the quantile, the interquartile range, the skewness, as well as the kurtosis. Evaluations using Monte-Carlo simulations show that standard errors estimators, assuming a normally distributed population, are almost always reliable. In addition, as expected, smaller sample sizes lead to less reliable results. The only exception is the estimate of the confidence interval for kurtosis, which shows evidence of unreliability. We therefore propose an alternative measure of confidence interval based on the lognormal distribution. This review provides easy to find information about many descriptive statistics which can be used, for example, to plot error bars or confidence intervals.
System Factorial Technology is a recent methodology for the analysis of information processing architectures. SFT can discriminate between three processing architectures, namely serial, parallel and coactive processing. In addition, it can discriminate between two stopping rules, self-terminating and exhaustive. Although the previously stated architectures fit to many psychological skills as performed by human beings (i.e. recognition task, categorization, visual search, etc.), the analysis of processing architectures that lie outside of the five original choices remain unclear. An example of such architecture is the recall process as performed by iterative systems. Results indicate that an iterative recall neural network is mistakenly detected by SFT as being a serial exhaustive architecture. This research shows a limit of SFT as an analytic tool but could lead to advancements in cognitive modeling by improving the strategies used for the analysis of underlying information processing architectures.