This study applied multidimensional scaling (MDS) to the Cognitive Assessment System-Second Edition (CAS2) to investigate the structural validity of PASS theory (Planning, Attention, Simultaneous, Successive) across two age groups (5-7, 8-18 years) in the normative sample (N = 1342). MDS is a statistical technique that creates a visual map showing similarities and differences between objects in a visual array of the distances between them. Accordingly, MDS offered a spatial representation of subtest relationships, revealing near-perfect model fit for both groups. Results indicated some clustering of Planning and Attention measures, partially supporting prior findings that these constructs may be psychometrically fused. Expressive Attention displayed atypical spatial placement from ages 8 to 18, suggesting factorial complexity beyond those linkages. A radex-like structure emerged, with cognitively complex tasks centrally located and more differentiated configurations observed in older participants, consistent with the developmental mutualism hypothesis. Findings highlight persistent challenges in cleanly separating PASS processes, with implications for CAS2 interpretation and intervention design. MDS proved a valuable complement to factor analysis, offering nuanced insights into test dimensionality and the ongoing debate over the validity of PASS theory.
The Wechsler Adult Intelligence Scale-Fifth Edition (WAIS-5) latent factor structure was assessed using complementary hierarchical exploratory factor analyses (EFA) with the Schmid and Leiman procedure and confirmatory factor analyses (CFA) using the standardization sample (N = 2,020) correlation matrix and descriptive statistics of the 20 primary and secondary WAIS-5 subtests. The WAIS-5 Technical and Interpretive Manual did not include EFA, CFA with fewer than five first-order (group) factors, CFA with rival bifactor models, or model-based reliability and dimensionality estimates; thus, the present independent structural validity assessment corrects this evidential lacuna to help guide ethical and evidence-based interpretation. EFA results did not support five latent factors with separate Visual Spatial and Fluid Reasoning factors. Instead, a four-factor model with Visual Spatial and Fluid Reasoning factors merged into the former Perceptual Reasoning factor and measurement dominated by a general intelligence (g) factor-similar to the WAIS-IV structure-was supported. CFA results indicated that a bifactor model with four group factors provided the best fit, consistent with the EFA findings. Overall, the EFA and CFA results did not support the purported WAIS-5 structure and instead replicated findings from independent assessments of the WISC-V with standardization and clinical samples, that indicated primary, if not exclusive, interpretation of the FSIQ as an estimate of psychometric g.
The present study examined the posited structure of the Wechsler Intelligence Scale for Children-Fifth Edition (WISC-V) ancillary index scores with normative sample participants aged 6-16 years (N = 2200) using a series of confirmatory factor analyzes (CFA) with maximum likelihood estimation. CFA results supported the retention of auditory working memory (AWM) but not quantitative reasoning (QR) as narrow dimensions in an extended WISC-V measurement model. Additional results from models explicating the structures for each of the posited ancillary composite-level indexes (nonverbal [NVI], general ability [GAI], cognitive processing [CPI]) provided support, in part, that these indexes represent global dimensions with differing degrees of generality. Though some of these scores may be used in the manner intended by the test publisher (e.g., comparing and contrasting performance on different composites, specific learning disability identification), provisional limitations for using the ancillary indexes as a focal point of clinical decision-making are discussed.
Although specific learning disabilities (SLD) represent the largest category for which school-age children receive special education services, the science and practice of SLD identification continues to evade consensus. Our goal is to bring together trainers and researchers with different perspectives on SLD identification to help spur a move toward a potential consensus, discuss agreements and disagreements on SLD identification in the field including amongst ourselves, and work toward productive discussion that may help move the field forward. We review essential conceptual questions that require greater scrutiny and thought to build a stronger understanding of SLD. We then discuss current assessment and identification practices, focusing on the not-so-controversial and the controversial issues in the field. Finally, we conclude with questions and considerations that challenge many of the established assumptions and systems currently in place. The aim of this article is to support constructive discussion on the topic of SLD that may have profound effects on the perennial issues the field continues to face.
The use of Bayesian structural equation modeling (BSEM) provided additional insight into the WISC-V theoretical structure beyond that offered by traditional factor analytic approaches (e.g., exploratory factor analysis and maximum likelihood confirmatory factor analysis) through the specification of all cross loadings and correlated residual terms. The results indicated that a five-factor higher-order model with a correlated residual between the Visual-Spatial and Fluid Reasoning group factors provided a superior fit to the four bifactor model that has been preferred in prior research. There were no other statistically significant correlated residual terms or cross loadings in the measurement model. The results further suggest that the WISC-V ten subtest primary battery readily attains simple structure and its index level scores may be interpreted as suggested in the WISC-V's scoring and interpretive manual. Moreover, BSEM may help to advance IQ theory by providing contemporary intelligence researchers with a novel tool to explore complex interrelationships among cognitive abilities-relationships that traditional structural equation modeling methods may overlook. It can also help attenuate the replication crises in school psychology within the area of cognitive assessment structural validity research through systematic evaluation of complex structural relationships obviating the need for CFA based post hoc specification searches which can be prone to confirmation bias and capitalization on chance.
This study investigated the incremental validity of the successive-level approach to intelligence test interpretation, specifically within the context of school psychology. The successive-level approach assumes that unique information is captured at various levels of intelligence tests (e.g., full scale, index-level, subtest-level). However, previous research has indicated that lower-level scores often fail to explain significant or meaningful variance in achievement outcomes beyond what is accounted for by the global score, suggesting potential redundancy in interpreting lower-level scores. Using the Woodcock-Johnson IV (WJ IV) standardization sample, this study examined the relationship between cognitive ability scores (GIA and CHC clusters) and achievement outcomes. The results indicated that, although the CHC subscale scores contributed some unique variance in achievement, their incremental validity was limited and frequently overshadowed by the GIA composite. These findings align with previous research, suggesting that clinicians may unwittingly commit a duplication fallacy when relying on successive-level score interpretation or related guidance in clinical practice.
Given the interdisciplinary influences on school psychology along with its requirement to comply with federal and state law in the United States, scientific progress in the area of cognitive assessment and specific learning disabilities (SLD) identification has experienced slow, if not stagnant, progress. Extrapolation of research from one discipline to that of assessment is common in school psychology where test authors and creators of interpretive and diagnostic systems make theoretical and empirical justification for their claims with correlational research and factor analysis. Although these methodologies may appear to support an underlying theory or interpretive approach, they can produce divergent results depending upon sample size and methodological choice. Consequently, greater replication and reproduction is required. Federal and state law in the United States may perpetuate low value practices among practitioners who view them as acceptable since they are legal. School psychology does not have regulatory agencies to oversee practices. All of these influences impinge on scientific progress in cognitive assessment and SLD identification. Fortunately, Canada is not beholden to omnibus special education law so its academic institutions and agencies (e.g., school districts) may be better poised to engender scientific progress in cognitive assessment and SLD identification.
The present study examined the structure of the NEPSY-II within the norming sample using exploratory factor analysis. For the 3–4-year-old group, our results were conceptually uninterpretable. As a result, a unidimensional model was retained by default as a remedy to local fit issues. For the 7–12-year-old group, our analysis supported some aspects of the NEPSY-II conceptual domains in the form of a six-factor model that yielded the best fit to the data. While variance partitioning results indicate that the majority of NEPSY-II subtests at ages 7–12 contain adequate specificity to be interpreted in isolation, caution is suggested for interpreting the Social Perception subtests; in particular, given the inability to locate that latent dimension in either of the analyses conducted. Implications for the clinical interpretation of the instrument moving forward are discussed.
This study aimed to evaluate the tenability of the proposed scoring/interpretive structure for the Woodcock-Johnson IV Test of Cognitive Abilities (WJ IV COG) Standard Battery configuration of subtests using confirmatory factor analysis (CFA) at school age. Results indicated that a three-factor hierarchical model, consistent with the CHC theory (Crystallized Ability, Fluid Reasoning, Short-Term Memory/Working Memory), provided the best fit to the WJ IV COG normative data. Whereas the preferred CHC interpretive structure was largely replicated, indices of interpretive relevance indicated that, among the Stratum II/III attributes that were located, only the omnibus general intelligence dimension should be interpreted with confidence. Nevertheless, several subtests contained adequate specificity to be interpreted in isolation apart from broad abilities. Implications for clinical interpretation are discussed.
One important aspect of construct validity is structural validity. Structural validity refers to the degree to which scores of a psychological test are a reflection of the dimensionality of the construct being measured. A factor analysis, which assumes that unobserved latent variables are responsible for the covariation among observed test scores, has traditionally been employed to provide structural validity evidence. Factor analytic studies have variously suggested either four or five dimensions for the WISC–V and it is unlikely that any new factor analytic study will resolve this dimensional dilemma. Unlike a factor analysis, an exploratory graph analysis (EGA) does not assume a common latent cause of covariances between test scores. Rather, an EGA identifies dimensions by locating strongly connected sets of scores that form coherent sub-networks within the overall network. Accordingly, the present study employed a bootstrap EGA technique to investigate the structure of the 10 WISC–V primary subtests using a large clinical sample (N = 7149) with a mean age of 10.7 years and a standard deviation of 2.8 years. The resulting structure was composed of four sub-networks that paralleled the first-order factor structure reported in many studies where the fluid reasoning and visual–spatial dimensions merged into a single dimension. These results suggest that discrepant construct and scoring structures exist for the WISC–V that potentially raise serious concerns about the test interpretations of psychologists who employ the test structure preferred by the publisher.
Developed more than 2 decades ago, the MEZURE (Assessment Technologies, 1995-2020; https://www.mezure.com/) has received increased attention as a result of the COVID-19 pandemic. It is the first individualized test of cognitive ability created to use an online (local or remote) assessment modality. The MEZURE claims to be aligned both with the extended Gf-Gc theory and the Cattell-Horn-Carroll model of abilities. Whereas the test publisher claims it used exploratory factor analysis to investigate the instrument's factor structure, only the subtest factor loadings on the Gf-Gc factors were furnished. No other structural validity information was provided, suggesting that users of the instrument should interpret the scores produced by the MEZURE with caution. Accordingly, the present study used exploratory and confirmatory factor analysis to more fully investigate the structural validity of the MEZURE. The results revealed that the MEZURE contains a combined perceptual reasoning (i.e., [Gf/Gv]/working memory [Gwm]) group factor, a verbal ability group factor, and a relatively weak general factor that is dominated by perceptual reasoning. The finding of a paltry general factor that is weakly loaded by verbal subtests is inconsistent with the broader research on traditional cognitive ability assessment and could be related to the online administration format of the test. Future research is required to better understand this finding. (PsycInfo Database Record (c) 2023 APA, all rights reserved).
Although the field of school psychology has made progress toward the use of tests and assessment practices with empirical support over the past 20 years, many school psychology practitioners still engage in what can be described as low-value value assessment practices that lack compelling scientific support potentially taking time and resources away from practices that have a demonstrated evidence-base. Why do school psychologists engage in questionable assessment and interpretive practices despite decades of discrediting scientific evidence? This article critically examines several plausible explanations for the perpetuation of low-value practices in school psychology assessment. It also underscores the importance of critical thinking when evaluating assessment and interpretation practices, and discusses practical recommendations to assist in advancing evidence-based assessment in school psychology training and practice as the field progresses well-into the 21st century.Impact StatementMany school psychologists engage in assessment practices that lack compelling scientific support potentially taking time, resources, and energy away from more effective practices. This article critically reviews reasons why these questionable assessment practices persist long after discrediting scientific evidence has been aptly presented. Recommendations are offered to promote the use of evidence-based practices and discourage the use of assessment methods lacking compelling empirical support in training and clinical practice.
School psychology contributes to the science of human behavior and utilizes this science to inform an evidence-based practice. The usefulness of this science is dependent on scientists making good faith efforts to minimize bias in their research. Nonetheless, implicit biases can still influence scientists’ decisions and, hence, the outcomes of their investigations. One source of such bias comes from conflicts of interest (COIs). In this article, we discuss COIs within the context of science, with a particular focus on financial COIs. In addition, we discuss how financial COIs can arise in school psychology as well as some ways the COIs may influence psychological science. We conclude by discussing how financial COIs are typically handled and some suggestions for handling them in the future.
This article addresses the use of hype in the promotion of clinical assessment practices and instrumentation. Particular focus is given to the role of school psychologists in evaluating the evidence associated with clinical assessment claims, the types of evidence necessary to support such claims, and the need to maintain a degree of “healthy self-doubt” about one’s own beliefs and preferred practices. Included is a discussion of topics that may facilitate developing and refining scientific thinking skills related to clinical assessment across common coursework, and how this framework fits with both the scientist-practitioner and clinical science perspectives for training.
Patterns of strengths and weaknesses represent relatively novel methods for identifying specific learning disabilities (SLD) with proponents asserting that the incorporation of multiple sources of assessment data and professional judgment play a key role in their utility. In this study, we examined if the sequential presentation of assessment data impacted school psychologists' ratings as to whether or not hypothetical students depicted in special education evaluation vignettes should be identified with SLD. Results showed that when participants viewed vignettes that were indicative of SLD (i.e., SLD positive), SLD likelihood ratings increased with the additional presentation of assessment data sources over time. However, when participants viewed vignettes that were indicative of a student not having SLD (i.e., SLD negative), SLD likelihood ratings were relatively consistent over time. Moreover, participants demonstrated relatively high levels of confidence in their SLD identification decisions, and in SLD negative vignettes, confidence increased after the fourth assessment data source was presented. Implications for SLD identification are discussed.
This study investigated the stability of Wechsler Intelligence Scale for Children-Fifth Edition (WISC-V) scores for 225 children and adolescents from an outpatient neuropsychological clinic across, on average, a 2.6 year test-retest interval. WISC-V mean scores were relatively constant but subtest stability score coefficients were all below 0.80 (M = 0.66) and only the Verbal Comprehension Index (VCI), Visual Spatial Index (VSI), and omnibus Full Scale IQ (FSIQ) stability coefficients exceeded 0.80. Neither intraindividual subtest difference scores nor intraindividual composite difference scores were stable across time (M = 0.26 and 0.36, respectively). Rare and unusual subtest and composite score differences as well as subtest and index scatter at initial testing were unlikely to be repeated at retest (kappa = 0.03 to 0.49). It was concluded that VCI, VSI, and FSIQ scores might be sufficiently stable to support normative comparisons but that none of the intraindividual (i.e. idiographic, ipsative, or person-relative) measures were stable enough for confident clinical decision making.