Low power, often the result of an insufficient number of participants, contributes to replication failures (e.g. Button et al., 2013). It is usually recommended that studies have at least an 80% chance of identifying an effect. However, following Cohen’s (1962) work, many reviews have reported that studies are under-powered. We calculated the power to detect small, medium and large effects of a test from the first study in all relevant Psychonomic Society journal articles published in 2017. Overall, mean power to detect medium sized effects was only 65%, and two-thirds of papers failed to achieve 80% power. Nearly half of papers in Psychonomic Bulletin & Review and Memory & Cognition achieved the 80% criteria, but other journals were much lower. The mean power to detect small effects was 22%, indicating that they would almost always be missed. Larger samples are generally needed to avoiding missing real effects.
The Psychonomic Society (PS) adopted New Statistical Guidelines for Journals of the Psychonomic Society in November 2012. To evaluate changes in statistical reporting within and outside PS journals, we examined all empirical papers published in PS journals and in the Experimental Psychology Society journal, The Quarterly Journal of Experimental Psychology (QJEP), in 2013 and 2015, to describe these populations before and after effects of the Guidelines. Comparisons of the 2013 and 2015 PS papers reveal differences associated with the Guidelines, and QJEP provides a baseline of papers to reflect changes in reporting that are not directly influenced by the Guidelines. A priori power analyses increased from 5% to 11% in PS papers, but not in QJEP papers (2%). The reporting of effect sizes in PS papers increased from 61% to 70%, similar to the increase for QJEP from 58% to 71%. Only 18% of papers reported confidence intervals (CIs) for means; only two PS papers in 2015 reported CIs for effect sizes. Although variability statistics are important to understanding data, and to further analysis, they were only reported as numbers in just over half of the PS journal papers. Almost all PS and QJEP papers relied exclusively on null hypothesis significance testing to guide interpretation of the data. Changes associated with the Guidelines are in the desired direction with respect to reporting effect sizes and power analyses but are not yet reflected in researchers’ practices in describing their data, addressing data assumptions, and thinking beyond the p value when interpreting their data.
Previous research has shown that little benefit is achieved through spaced study and recall of text passages after the first recall attempt, an effect that we term the failure-of-further-learning. We hypothesized that the effect occurs because a situation model of the text's gist is formed when the text is first comprehended and is consolidated when recalled; it dominates later recall after verbatim memories of more recent study episodes have been lost. Experiments 1 and 2 attempted to circumvent the effect by varying the activities of participants and requiring interactive exploration. In both experiments, recall after four, weekly sessions showed little benefit beyond performance on the first recall. Experiment 3 interfered with the formation of an immediate situation model by introducing passages that were hard to comprehend without a title. Performance improved substantially across four sessions when titles were not supplied, but the standard effect was replicated when titles were given. Experiment 4 made verbatim memories available by incorporating all re-presentations and tests into one session; as predicted, recall improved over successive tests.
Previous research has shown that little benefit is achieved through spaced study and recall of text passages after the first recall attempt, an effect that we term the failure-of-further-learning. We hypothesized that the effect occurs because a situation model of the text's gist is formed when the text is first comprehended and is consolidated when recalled; it dominates later recall after verbatim memories of more recent study episodes have been lost. Experiments 1 and 2 attempted to circumvent the effect by varying the activities of participants and requiring interactive exploration. In both experiments, recall after four, weekly sessions showed little benefit beyond performance on the first recall. Experiment 3 interfered with the formation of an immediate situation model by introducing passages that were hard to comprehend without a title. Performance improved substantially across four sessions when titles were not supplied, but the standard effect was replicated when titles were given. Experiment 4 made verbatim memories available by incorporating all re-presentations and tests into one session; as predicted, recall improved over successive tests.
Reports an error in "Neuroanatomical Correlates of Behavioral Rating Versus Performance Measures of Working Memory in Typically Developing Children and Adolescents" by Nazlie Faridi, Sherif Karama, Miguel Burgaleta, Matthew T. White, Alan C. Evans, Vladimir Fonov, D. Louis Collins and Deborah P. Waber (Neuropsychology, Advanced Online Publication, Jul 7, 2014, np). In the original version of the article, Matthew T. White's name was missing his middle initial. All versions of this article have been corrected. (The following abstract of the original article appeared in record 2014-27816-001.)The frequent lack of correspondence between performance and observational measures of executive functioning, including working memory, has raised questions about the validity of the observational measures. This study was conducted to investigate sources of this discrepancy through correlation of volumetric and cortical thickness (CT) neuroimaging values with performance and questionnaire measures of working memory (WM).Using longitudinal data from the NIH MRI Study of Normal Brain Development (Volumes, N= 347, 54.3% female; CT, N= 350, 54.6% female; age range: 6 to 16.9 years), scores on the Behavioral Rating Inventory of Executive Function (BRIEF) WM, Emotional Control (EC) and Inhibition (INH) scales; Wechsler Scale of Intelligence for Children-III Digit Span; and Cambridge Neuropsychological Test Battery Spatial Working Memory (CANTAB SWM) were correlated with each other and with morphometric measurements using mixed effects linear regression models.BRIEF WM was correlated with CANTAB SWM (p < .001). With whole brain correction, BRIEF WM and EC were both correlated with CT of the posterior parahippocampal gyrus (PHG), EC on the right side only. Performance measures of WM were unrelated to lobar volumes or CT, but were associated with volumes of hippocampus and amygdala.The known role of PHG in contextual learning suggests that the BRIEF WM assesses contextualized learning/memory, potentially explaining its loose correspondence to the decontextualized performance measures. Observational measures can be useful and valid functional metrics, complementing performance measures. Labels used to characterize scales should be interpreted with caution, however. (PsycINFO Database Record (c) 2015 APA, all rights reserved).
In four experiments, we extended the study of part-set cuing to expository texts and pictorial scenes. In Experiment 1, recall of expository text was tested with and without part-set cues in the same order as the original text; cues strongly impaired recall. Experiment 2 repeated Experiment 1 but used cues in random order and found significant but reduced impairment with cuing. Experiments 3 and 4 examined the part-set cuing of objects presented in a scene or matrix and found virtually no effect of cuing. More objects were recalled from the scene than from the matrix, indicating that the scene's organization aided memory, but the cues did not assist recall. These results extend the domains in which part-set cues have either impaired or failed to improve recall. Implications for education and eyewitness accounts are briefly considered.
Past research has reported a consistent but small relationship (e.g. r=.23) between conscientiousness and university academic performance. However, in almost all cases the nature of the academic work has not been divided into the major elements of coursework and examination performance. We examined the relationships between conscientiousness and procrastination and the coursework and examination performance of psychology students in their second and third year modules. Both conscientiousness (r=.45) and procrastination (r=−.39) were significant predictors of overall coursework marks and significantly predicted coursework marks for all but one of the individual modules. Correlations with examination marks were smaller and less consistent. Regression analysis showed that conscientiousness was the more dominant predictor than procrastination. These results extend the literature relating conscientiousness to academic performance, demonstrating that the relationship is stronger with coursework than with exams.
In this research we explored the use of short-answer questions to improve learning from chapter-like texts (3395 words). Experiment 1 investigated the influence of pre-questions on recall from a text passage when tested a week later; two question sets were counterbalanced within the experimental group. Participants with pre-questions scored higher both overall (d = 3.6, 95%CI [2.4, 4.8]) and on novel questions (d = 2.0 [1.6, 2.4]). In Experiment 2, questions were made available immediately after studying the text either alongside the text, open-book, or closed-book with the opportunity to check answers, or not at all with additional study time. Learning was tested after a week. Although the immediate test scores were substantially higher for open- than closed-book tests, week-delayed performance on the same items was much worse for open-book tests and was moderately improved for closed-book tests. For seen questions, closed-book tests led to better delayed recall than did open-book tests, d = 0.7 [0.02, 1.5]. For novel questions, observed differences were small; ds = .2 [-0.6, 0.9] for both comparisons.
This short paper reviews the reasons why effect sizes are worthy of reporting and consideration when interpreting results. It briefly describes the use of the more common effect sizes and provides examples where the use of effect sizes adds meaning to the research results
Effect sizes are omitted from many research articles and are rarely discussed. To help researchers evaluate effect sizes we collected values for the more commonly reported effect size measures (partial eta squared and d) from papers reporting memory research published in 2010. Cohen's small, medium, and large generic guideline values for d mapped neatly onto the observed distributions, but his values for partial eta squared were considerably lower than those observed in current memory research. We recommend interpreting effect sizes in the context of either domain-specific guideline values agreed for an area of research or the distribution of effect size estimates from published research in the domain. We provide cumulative frequency tables for both partial eta squared and d enabling authors to report and consider not only the absolute size of observed effects but also the percentage of reported effects that are larger or smaller than those observed.