Equivalence testing, also called negligible effect significance testing (NEST), is appropriate when a researcher would like to find evidence of a negligible association. However, since equivalence testing/NEST procedures are newer and considerably less popular than traditional difference-based null hypothesis significance testing, it is useful to give a gentle introduction to these methods. Accordingly, this tutorial paper aims to provide an overview of NEST/equivalence testing procedures by describing the nature of the procedures, explaining when they should be used, defining what considerations should go into their application (including selecting a minimally meaningful effect size), and outlining how they may be conducted and interpreted. The tutorial paper also includes examples and code in open-source software to illustrate how these procedures may be applied to real data.
This study examined the feasibility of online group therapy for trauma survivors. Community-based participants (N = 178; 64% women; mean age = 43) self-referred for an eight-week psychoeducational, skills-based trauma therapy group, which was offered via secure video conferencing. Participants were diverse in terms of socio-economic status and ethnic identity. Participants completed the PCL-5 at registration (T0), pre-treatment (T1), and post-treatment (T2). T0-T1 reflected a waitlist period; T1-T2 represented the treatment phase. Mixed effects models showed significant overall PCL-5 reductions from T0-T2, with clinically meaningful change during treatment (T1-T2). PTSD symptom subtypes as measured by the PCL-5 (reexperiencing, avoidance, negative cognition/mood, arousal) also improved. Findings support telemedicine as a feasible option for fostering stabilization skills and reducing barriers to trauma treatment.
Bayesian statistics has gained substantial popularity in the social sciences, particularly in psychology. Despite its growing prominence in the psychological literature, many researchers remain unacquainted with Bayesian methods and their advantages. This tutorial addresses the needs of curious applied psychology researchers and introduces Bayesian analysis as an accessible and powerful tool. We begin by comparing Bayesian and frequentist approaches, redefining fundamental terms from both perspectives with practical illustrations. Our exploration of Bayesian statistics includes Bayes's Theorem, likelihood, prior and posterior distributions, various prior types, and Markov-Chain Monte Carlo (MCMC) methods, supplemented by graphical aids for clarity. To bridge theory and practice, we employ a psychological research example with real, open data. We analyse the data using both frequentist and Bayesian approaches, providing R code and comprehensive supporting information, and emphasising best practices for interpretation and reporting. We discuss and demonstrate how to interpret parameter estimates and credible intervals, among other essential topics. Throughout, we maintain an accessible and user-friendly language, focusing on practical implications, intuitive examples, and actionable recommendations.
We predicted that low statistics grades, anxiety sensitivity, and trait perfectionism would be associated with increased statistics anxiety and worsened statistics attitudes. We also expected a grades by personality interaction, consistent with the vulnerability-stress model. Participants included 423 students currently taking a statistics class. We used a two-wave longitudinal design using self-reported online surveys at the beginning of term and after final grades were released. Grades were self-reported letter grades in statistics classes. Grades predicted increased statistics anxiety and worsened attitudes. Anxiety sensitivity predicted increased statistics anxiety. Self-critical perfectionism positively predicted statistics anxiety, but not attitudes. Rigid perfectionism was not significantly associated with either outcome. No interaction effects were statistically significant, failing to support the vulnerability-stress model.
Whenever researchers test multiple hypotheses, the risk that one or more of the hypotheses might be falsely supported (i.e., Type I errors) increases with the number of hypotheses evaluated (i.e., the multiplicity problem). Because most studies evaluate multiple hypotheses, the risk that at least one hypothesis is a Type I error appears substantial. However, it is necessary to evaluate how consistently multiplicity control (MC) is applied and the rationale behind its application, to understand the merit of MC in psychological research. We conducted a systematic review of MC practices in 250 articles from 10 high-impact psychology journals. There was a median of 76 hypotheses tested per article; however, only 7.5% of all hypotheses were protected by MC. The most popular type of MC was familywise error control, via the Bonferroni method. The results highlight the wide range of situations wherein multiplicity occurs, the inconsistency with which MC is applied, and the lack of rationale for MC decisions. We hope that the results will create an active discussion on the merit of MC within the field of psychology. Lorsque les chercheurs testent plusieurs hypoth & egrave;ses, le risque qu'une ou plusieurs d'entre elles soient faussement confirm & eacute;es (c'est-& agrave;-dire les erreurs de type I) augmente avec le nombre d'hypoth & egrave;ses & eacute;valu & eacute;es (c'est-& agrave;-dire le probl & egrave;me de la multiplicit & eacute;). Comme la plupart des & eacute;tudes & eacute;valuent plusieurs hypoth & egrave;ses, le risque qu'au moins l'une d'entre elles soit une erreur de type I semble important. Toutefois, il est n & eacute;cessaire d'& eacute;valuer la coh & eacute;rence du contr & ocirc;le de la multiplicit & eacute; (CM) et la logique qui sous-tend son application, afin de comprendre l'int & eacute;r & ecirc;t du contr & ocirc;le de la multiplicit & eacute; dans la recherche psychologique. Nous avons proc & eacute;d & eacute; & agrave; un examen syst & eacute;matique des pratiques de CM dans 250 articles provenant de 10 revues de psychologie & agrave; fort impact. En moyenne, 76 hypoth & egrave;ses ont & eacute;t & eacute; test & eacute;es par article ; cependant, seules 7,5 % de toutes les hypoth & egrave;ses & eacute;taient prot & eacute;g & eacute;es par le CM. Le type de CM le plus r & eacute;pandu & eacute;tait le contr & ocirc;le d'erreur global, via la m & eacute;thode de Bonferroni. Les r & eacute;sultats mettent en & eacute;vidence le large & eacute;ventail de situations dans lesquelles la multiplicit & eacute; se produit, le manque de coh & eacute;rence dans l'application de la CM et l'absence de justification des d & eacute;cisions en mati & egrave;re de CM. Nous esp & eacute;rons que les r & eacute;sultats susciteront une discussion active sur les m & eacute;rites du CM dans le domaine de la psychologie.
Statistics play an important role in psychology, but statistics modules are notoriously unpopular amongst psychology students. We examined attitudes toward statistics and attitudes toward the statistical software package R in both undergraduate and postgraduate students across the duration of a statistics module. Participants’ responses were analysed using both quantitative and qualitative techniques. Results demonstrated that, on average, students in introductory level modules held neutral (not negative) attitudes, but students at higher study levels held somewhat positive attitudes towards R and statistics. While not all students enjoyed learning R, our findings demonstrate that many students enjoyed statistics or R, most students found statistics/R valuable and they generally reported feeling competent using software by the end of their module. These results challenge the argument that R is not suitable for undergraduate psychology students. Consequently, benefits, challenges, and implications of teaching R to psychology students are discussed.
Linear models are particularly vulnerable to influential observations which disproportionately affect the model's parameter estimates. Multiple statistics and numerous cut-off values have been proposed to detect highly influential observations including Cook’s Distance (CD), Standardized Difference of Fits (DFFITS) and Standardized Difference of Beta (DFBETAS). This paper reports on a Monte Carlo simulation study that assesses the effectiveness of these methods and recommended cut-off values under various conditions, including different sample sizes, numbers of predictors, strengths of variable associations, and non-sequential versus sequential analysis approaches within a multiple linear regression framework. The findings suggest that the proportion of observations identified as highly influential varies significantly based on the chosen diagnostic method and the thresholds used for detection. Consequently, researchers should consider the implications of their methodological choices and the thresholds they apply when identifying influential data points.
It has been suggested that equivalence testing (otherwise known as negligible effect testing) should be used to evaluate model fit within structural equation modelling (SEM). In this study, we propose novel variations of equivalence tests based on the popular root mean squared error of approximation and comparative fit index fit indices. Using Monte Carlo simulations, we compare the performance of these novel tests to other existing equivalence testing-based fit indices in SEM, as well as to other methods commonly used to evaluate model fit. Results indicate that equivalence tests in SEM have good Type I error control and display considerable power for detecting well-fitting models in medium to large sample sizes. At small sample sizes, relative to traditional fit indices, equivalence tests limit the chance of supporting a poorly fitting model. We also present an illustrative example to demonstrate how equivalence tests may be incorporated in model fit reporting. Equivalence tests in SEM also have unique interpretational advantages compared to other methods of model fit evaluation. We recommend that equivalence tests be utilized in conjunction with descriptive fit indices to provide more evidence when evaluating model fit.
A popular measure of model fit in structural equation modeling (SEM) is the standardized root mean squared residual (SRMR) fit index. Equivalence testing has been used to evaluate model fit in structural equation modeling (SEM) but has yet to be applied to SRMR. Accordingly, the present study proposed equivalence-testing based fit tests for the SRMR (ESRMR). Several variations of ESRMR were introduced, incorporating different equivalence bounds and methods of computing confidence intervals. A Monte Carlo simulation study compared these novel tests with traditional methods for evaluating model fit. The results demonstrated that certain ESRMR tests based on an analytic computation of the confidence interval correctly reject poor-fitting models and are well-powered for detecting good-fitting models. We also present an illustrative example with real data to demonstrate how ESRMR may be incorporated into model fit evaluation and reporting. Our recommendation is that ESRMR tests be presented in addition to descriptive fit indices for model fit reporting in SEM.
Self-critical perfectionism and anxiety sensitivity are potential vulnerability factors for increased distress following performance failure. We hypothesized that participants who fail a statistics quiz will have lower state self-esteem, lower positive affect, and greater negative affect at post-test than those who get a good grade, after controlling for pre-test scores and the effect of experimental condition would become larger as self-critical perfectionism and anxiety sensitivity increase. Exploratory analyses examined rigid perfectionism and a newly introduced construct (statistics anxiety sensitivity) as moderators. We tested this vulnerability-stress model in 329 post-secondary students using a two-group, pre-post, between-subjects design. Students completed an easy or hard statistics test and were assessed on pre- and post-test state self-esteem (social & performance) and state affect (anxiety, dysphoria, hostility, & positive affect). Across outcomes, main effects of experimental condition predicted between 7-33% of the variance, with the largest effects for performance self-esteem. Personality by condition interactions predicted 0.1-2% of the variance; 16 of 24 interactions were statistically significant in the expected direction (i.e., the effect of experimental condition was larger for participants high in measured personality traits). Findings suggest personality traits are vulnerability factors for decreased self-esteem and increased negative affect following failure in a statistics assessment.
This large, international dataset contains survey responses from N = 12,570 students from 100 universities in 35 countries, collected in 21 languages. We measured anxieties (statistics, mathematics, test, trait, social interaction, performance, creativity, intolerance of uncertainty, and fear of negative evaluation), self-efficacy, persistence, and the cognitive reflection test, and collected demographics, previous mathematics grades, self-reported and official statistics grades, and statistics module details. Data reuse potential is broad, including testing links between anxieties and statistics/mathematics education factors, and examining instruments’ psychometric properties across different languages and contexts. Data and metadata are stored on the Open Science Framework website (https://osf.io/mhg94/).
Reporting and interpreting effect sizes (ESs) has been recommended by all major bodies within the field of psychology. In this systematic review, we investigated the reporting of ESs in six social-personality psychology journals from 2018, given that this area has been at the center of psychology’s replication crisis. Our results highlight that although ES reporting is near perfect (even for follow-up tests), interpreting the magnitude of ESs, including confidence intervals for ESs, and interpreting the precision of the confidence intervals needs development. We also highlight widespread confusion regarding the interpretations of the magnitude of ESs within the context of the research.
Equivalence testing (ET) is a framework to determine if an effect is small enough to be considered meaningless, wherein meaningless is expressed as an equivalence interval (EI). Although traditional effect sizes (ESs) are important accompaniments to ET, these measures exclude information about the EI. Incorporating the EI is valuable for quantifying how far the effect is from the EI bounds. An ES measure we propose is the proportional distance (PD) from an observed effect to the smallest effect that would render it meaningful. We conducted two Monte Carlo simulations to evaluate the PD when applied to (1) mean differences and (2) correlations. The coverage rate and bias of the PD were excellent within the investigated conditions. We also applied the PD to two recent psychological studies. These applied examples revealed the beneficial properties of the PD, namely its ability to supply information above and beyond other statistical tests and ESs.
The treatment-control pre-post-follow-up (TCPPF) design is a popular means to demonstrate that a treatment group is superior to a control group over time.The TCPPF design can be analyzed using traditional methods (e. g., between-within ANOVA) or with modern multilevel (also known as mixed or hierarchical) modeling.In spite of TCPPF's widespread popularity, there is sparse and confusing guidance for applied researchers on how to analyze data from TCPPF designs using SPSS, one of the most popular software packages for data analysis.We present an introductory tutorial on methods for analyzing TCPPF data.Advantages, disadvantages, and cautions related to applying these approaches are discussed.
Behavioral science researchers are often interested in whether there is negligible interaction among continuous predictors of an outcome variable. For example, a researcher might be interested in demonstrating that the effect of perfectionism on depression is very consistent across age. In this case, the researcher is interested in assessing whether the interaction between the predictors is too small to be meaningful. Unfortunately, most researchers address the above research question using a traditional association-based null hypothesis test (e.g. regression) where their goal is to fail to reject the null hypothesis of no interaction. Common problems with traditional tests are their sensitivity to sample size and their opposite (and hence inappropriate) hypothesis setup for finding a negligible interaction effect. In this study, we investigated a method for testing for negligible interaction between continuous predictors using unstandardized and standardized regression-based models and equivalence testing. A Monte Carlo study provides evidence for the effectiveness of the equivalence-based test relative to traditional approaches.
BACKGROUND:Many youth with neurodevelopmental disorders (NDDs) experience mental health problems such as anxiety, depression or anger, and these are often associated with impairments of cognition and emotion regulation. The mechanisms that may be linking cognitive difficulties, emotion regulation and mental health are not known.AIMS:The current study examined whether adaptive and maladaptive (dysregulated) emotion regulation mediated the link between different cognitive control processes (working memory, inhibition and shifting) and internalizing/externalizing symptoms in children with NDDs.METHODS:Participants included 48 children (8-13 years of age) with one or more diagnoses of autism, attention deficit hyperactivity disorder, cerebral palsy and learning disability, who were enrolled in a larger study of cognitive behaviour therapy targeting emotion regulation. Multiple mediation analyses were implemented using the PROCESS macro. The mediation effects of adaptive and maladaptive emotion regulation were examined on the relationships between (1) working memory and internalizing/externalizing symptoms, (2) inhibition and internalizing/externalizing symptoms and (3) shifting and internalizing/externalizing symptoms. All data were collected prior to intervention, at baseline.RESULTS:Shifting, inhibitory control and working memory predicted increased emotion dysregulation, which functioned as a full mediator to both internalizing and externalizing problems in children with NDDs.CONCLUSIONS:In the presence of emotionally triggering situations, children with greater cognitive challenges experience greater maladaptive emotion regulation, which results in both internalizing and externalizing problems. For youth with NDDs, therapeutic plans that include strengthening of working memory, inhibition and shifting abilities in addition to emotion regulation skills training may be helpful in alleviating externalizing and internalizing behaviour.
Suppressor variables increase the predictive power of one or more predictors by suppressing irrelevant variance. Although theoretically and statistically useful, no research has addressed the frequency or interpretation of statistical suppression (SS) in the psychological literature. In two studies, we explored the nature and interpretation of SS. In the first study, we reviewed regression analyses to determine the frequency with which SS occurs in psychological articles published in 2017. Approximately one-third of articles showed evidence of SS, although researchers almost never acknowledged or attempted to interpret the SS. In the second study, we reviewed articles containing the keyword "suppression" to assess the interpretations provided by researchers that identified SS. Results indicate that most researchers do not attempt to classify or interpret SS. Therefore, although SS is common in psychology, scarcely any attempts are made to identify, classify, and/or interpret it.
In this paper we endeavour to provide a largely non-technical description of the issues surrounding unbalanced factorial ANOVA and review the arguments made for and against the use of Type I, Type II and Type III sums of squares. Though the issue of which is the 'best' approach has been debated in the literature for decades, to date confusion remains around how the procedures differ and which is most appropriate. We ultimately recommend use of the Type II sums of squares for analysis of main effects because when no interaction is present it tests meaningful hypotheses and is the most statistically powerful alternative.