
Curriculum-based measurement (CBM) tools for early reading remain largely unvalidated in Latin American contexts, limiting the availability of evidence-based screening instruments for intervention decision-making. This study examined the technical adequacy and screening utility of PRELEC ( Predictores de Lectura para la Identificacion Temprana ), a CBM battery designed to support early reading intervention decisions for first-grade Chilean children. A total of 738 students from low–, middle–, and high–socioeconomic status (SES) backgrounds completed five PRELEC tasks across three waves during first grade, with a standardized reading criterion administered 1 year later. Analyses provided evidence based on internal structure, relations to other variables, and classification accuracy, consistent with current test validation standards. Results supported a stable unidimensional structure across waves, with full measurement invariance across sex and partial invariance across SES and time. Latent mean analyses revealed no sex differences but persistent, large SES disparities from school entry. Classification accuracy ranged from acceptable to outstanding (area under the curve [AUC] = .72–.91), with composite scores consistently outperforming individual indicators. Findings support PRELEC as a valid and equitable screening tool for identifying students at risk for reading disabilities, with direct implications for intervention placement decisions within tiered instructional frameworks and the Chilean regulatory context.
Although associated with desirable student and teacher outcomes, many educators lack sufficient training in evidence-based, effective classroom management practices. The Direct Behavior Rating-Classroom Management (DBR-CM) is a low-inference, flexible, and feasible observational assessment tool developed to support professional development in classroom management. Validation efforts have yielded strong psychometric evidence in support of DBR-CM use thus far; however, early validation work was conducted with participants who received only brief, familiarization training. Given the positive influence of training on the reliability and accuracy of observational assessment, this study evaluated the impact of brief, comprehensive training on the accuracy of DBR-CM ratings. Results indicated that training significantly improved DBR-CM rating accuracy. The impact of training on the perceived social validity of the DBR-CM was also examined. Training predicted higher levels of understanding and acceptability, but not other dimensions of social validity (e.g., feasibility, system climate, system support). Findings support the efficacy of comprehensive training, which included modeling, guided practice, performance feedback, and calibration, in improving DBR-CM rating accuracy.
Curriculum-based measurement (CBM) can be used for screening and progress monitoring academic skill development, including written expression skills (WE-CBM). While research on CBM use with autistic students is growing, questions linger regarding its use with autistic students. One question is whether autistic students' WE-CBM scores are impacted by handwriting legibility, especially since fine motor challenges are common in the population. This study explored the potential relationship between handwriting legibility and WE-CBM scores among elementary-age autistic children. Thirty-three WE-CBM samples were collected from 10 children as part of a prior study and then scored using the Handwriting Legibility Scale (HLS). Handwriting legibility in this sample scored, on average, in the "fair" range, but within-student performance varied widely. Within-student correlations between handwriting legibility and writing performance ranged from r = -.54 to .94, with handwriting legibility accounting for between 3 and 88% of the variance in writing outcomes. The considerable variability in how these skills covaried across participants underscores the need for continued focus on writing measurement in ways that inform and support individual development and measurement of writing skills in this population.
Bullying is a pervasive problem in U.S. educational settings, yet research has predominantly focused on traditional K-12 schools, with limited attention to restrictive educational environments such as juvenile justice facilities and residential treatment centers. These settings present unique challenges, such as highly structured conditions that restrict youth autonomy in movement, social interactions, and peer selection, potentially intensifying bullying behaviors and undermining rehabilitative goals. Despite the critical need for accurate assessment in these high-risk contexts, few validated measures are available for this population. This study evaluated the psychometric properties of a modified version of the Illinois Bully/Victim/Fight Scale in restrictive educational settings. We adopt exploratory structural equation modeling (ESEM), which combines the advantages of exploratory factor analysis (EFA) and confirmatory factor analysis (CFA), allowing for a more accurate representation of construct overlap. Results demonstrated that the four-factor structure (bullying, victimization, fighting, and anger) provided an adequate fit with ESEM, whereas traditional CFA yielded a less satisfactory fit. The ESEM revealed theoretically meaningful cross-loadings, with several items loading on multiple dimensions. Internal consistency was acceptable (omega = .74-.93). Findings support cautious use of this modified scale in restrictive education settings, with ESEM recommended to accommodate construct overlap.
Behavior needs in the classroom are a persistent problem for teachers. Teachers often need additional training to determine behavioral function, so they can effectively promote student behavioral success. Trial-based functional analysis (TBFA) is a brief method for determining the function of a behavior, which aligns better with typical classroom routines. There is evidence that individual and small group trainings help teachers acquire TBFA implementation skills. However, large group training paradigms and outcomes have limited research. This study assessed the level of performance on outcomes from two separate, successive large group trainings on TBFA. Training 1 included 11 undergraduate and graduate students from a regional university. Training 2 included 18 in-service special education teachers who attended a conference workshop. Post-test assessments indicated generally accurate knowledge of TBFA procedures across trainings (range = 83%-100%). Descriptive comparisons between participants in Training 1 and Training 2 suggested similar levels of TBFA knowledge. However, adjustments were made based on the setting and participants that impacted outcomes. While large group training for TBFA is promising, more research is needed to improve delivery and apply it within schools.
Spelling is a foundational literacy skill that supports both word reading and written expression. For students with or at risk of learning disabilities, difficulties in spelling often constrain the fluency and complexity of writing, making effective interventions essential. Yet the conclusions drawn about intervention efficacy depend heavily on how outcomes are measured. This review synthesizes outcome measurement practices across 59 spelling intervention studies conducted over the past five decades. All outcome measures (n = 233) were coded by type (researcher-developed vs. norm-referenced) and by linguistic level (sublexical, lexical, sentence, discourse) using the Interactive Dynamic Literacy (IDL) framework. Descriptive analyses revealed that nearly four out of five outcomes were lexical, most often researcher-developed lexical-level spelling probes, with comparatively few outcomes at the sentence or discourse levels. Standardized assessments were similarly concentrated at the word level, with the Wide Range Achievement Test-Spelling subtest and Test of Written Spelling most commonly used. Finally, the pairing of proximal and standardized outcomes was inconsistent, particularly among group designs. Taken together, findings highlight a measurement bottleneck: spelling interventions are evaluated primarily through lexical-level accuracy, offering limited insight into whether gains transfer to the higher-level writing processes for students with or at risk for LD.
The Ages & Stages Questionnaires: Social-Emotional, Second Edition (ASQ:SE-2) is a caregiver-completed screening tool for children ages 1 to 72 months. The measure has been adapted for use in Taiwan. This study examined the psychometric properties of the Traditional Chinese version of the 60-month interval (ASQ:SE-2-TC) using item response theory (IRT) and gender differences using differential item functioning analyses. A sample of 702 children, ages 54 months 0 days to 72 months 30 days, was collected to proportionally reflect the population distribution across regions in Taiwan. Results indicated that (a) item fit statistics ranged from 0.84 to 1.34 (M = 1.01, SD = 0.13), (b) item difficulty ranged from -0.41 to 3.75 (M = 2.00, SD = 0.74), (c) reliability (EAP/PV) was 0.80, and (d) all items demonstrated negligible differential item functioning by gender. One misfitting item was further examined using item characteristic curves to evaluate its performance across the latent trait continuum of social-emotional competence. Findings provide psychometric support for the use of the ASQ:SE-2-TC with Taiwanese children. The promising evidence suggests the importance of continued validation efforts to further establish its utility in developmental screening.
Bilingual education programs, including immersion language programs, are rapidly increasing in U.S. schools. Although multi-tiered systems of support frameworks have been modified to include bilingual approaches, facets of data-based decision making such as expected reading fluency growth have yet to be fully explored in these contexts. This brief report describes oral reading fluency growth rates across languages for English-speaking German immersion learners across grades 1 (n = 143), 2 (n = 145), 3 (n = 202), and 4 (n = 201). These growth rates were compared to publisher-provided oral reading fluency growth rates for monolingual ELs. For English-speaking German immersion learners, oral reading fluency growth rates in English were initially substantially slower than oral reading fluency growth rates for monolingual ELs. However, English oral reading fluency growth rates increased such that immersion learners achieved similar levels of oral reading fluency in English toward the end of fourth grade. Oral reading fluency growth rates in German peaked in second grade, after which they slowed down considerably. Results suggest that for bilingual learners in diverse language learning environments, including additive ones such as immersion, and potentially subtractive ones such as transitional bilingual programs, local reading fluency growth rate norms are required to accurately identify students who are making expected reading growth.
The Integrated MTSS Fidelity Rubric (IMFR) is a 14-item measure of elementary school implementation of integrated multi-tiered systems of support (I-MTSS). With I-MTSS, schools strategically use assessment to guide intervention that addresses the integrated nature of students' academic, social, emotional, and behavioral needs. The goal for this study was to develop and validate the IMFR as a measurement tool that is useful and relevant for school practitioners and researchers who deliver and study intervention for students in elementary schools. The IMFR was validated through 3 years of iterative administration, psychometric testing, and refinement, using a nationwide sample that ranged from 65 to 87 elementary schools across 13-20 districts in a given administration year. Analyses examined content validity, substantive validity, structural validity, and generalizability. In addition, the study examined usability through focus groups with participating school teams. Overall, we conclude that the IMFR is a reliable and valid measure of I-MTSS implementation and is useful to school and district practitioners.
Strengths-based assessment (SBA) focuses on traits and resources that foster resilience, as opposed to the symptoms and impairments that are the focus of traditional deficit-based assessment. In schools, SBA can provide an effective framework for social, emotional, behavioral, and academic intervention. The Social Emotional Health Survey (SEHS) system is a set of SBAs that assess the synergistic effects of multiple psychological strengths (i.e., covitality), though more research is needed. This study examined the measurement invariance of the Social Emotional Health Survey-Primary (SEHS-P) for elementary school-age children, across four subscales (Gratitude, Zest, Optimism, and Persistence), as well as the higher order construct of covitality, using multiple-group categorical confirmatory factor analysis. We analyzed responses from 1,030 students across 16 elementary schools in two southern states, comparing results by state, grade (fourth, fifth), gender, and race (Black, White). The results suggest configural, weak, strong, and latent means invariance across all comparisons. Covitality was confirmed as a higher order construct, with configural and weak invariance holding across all comparisons. These results suggest the SEHS-P is a reliable tool for assessing psychological strengths and covitality in diverse elementary student populations, supporting its generalizability and value in promoting student well-being and informing school-based interventions.
Technical quality of language sample analysis (LSA) metrics using sentence-level writing curriculum-based measures was examined with 73 emergent bilinguals (English learners) in Grades 1-3. Alternate-form reliability, criterion-related validity between LSA metrics with writing curriculum-based measure metrics, predictive validity between fall LSA scores and winter scores on a standardized English proficiency measure, discrimination among grades, and sensitivity to growth were evaluated. The LSA metric mean length of T-Unit in words showed technical quality using the mean of two forms in the fall for Grades 2 and 3, while a number of different words maintained technical quality in Grades 2 and 3 across seasons using individual and the mean of two forms. Discrimination among grades and sensitivity to growth evidence were weaker.
Universal screeners of academic skills in schools are intended to predict the probability of academic risk in an efficient and economical manner. Recent methods of calculating post-test risk probabilities (Klingbeil et al., 2019; 2021) have been demonstrated to be simple and efficient to calculate, improving data-based decision-making practices in schools. However, these methods do not leverage the full advantages of Bayesian statistical inference, thereby limiting the quantification of uncertainty in the calculation of posterior probabilities of risk. This could produce overly deterministic data-based decisions. Bayesian ordinal regression models (BORMs) are a fully Bayesian extension of existing posterior probability calculations, and they offer multiple potential advantages for enhancing universal screening practices in schools. Through simulations and an applied example using real screening data, we elucidate some of the issues around BORMs in screening, including potential strengths (e.g., multilevel modeling) and barriers to practice (difficulty of interpretation/implementation). We discuss how BORMs can further advance both research and practice of data-based decision making in universal screening in schools.
Students with mathematics difficulty (MD) often struggle with both computation and word-problem solving, which are foundational skills emphasized in national standards such as the Common Core State Standards. As most past-error analysis research has primarily focused on a single mathematical topic, little is known about whether students with MD demonstrate error patterns consistently across both computation and word problems. The present study examined error types and consistency that Grade 4 and 5 students with MD made on 22 computation problems and 10 word problems. Using a researcher-developed coding protocol adapted from prior literature, we identified that the most common errors were miscalculation, regrouping subtract smaller integer, and wrong operation in computation; and wrong schema, miscalculation, copy, and regrouping in word problems. The majority of students did not demonstrate overlapping errors.
This study examined students' growth in algebra readiness progress monitoring measures for middle school students with math learning difficulties within the context of Project STAIR, a federally funded initiative supporting teachers' use of data-based individualization. Participating teachers received professional development and coaching focused on evidence-based instructional strategies and data use to enhance students' algebra readiness. Using multilevel modeling, we examined students' growth over time on the three measures-Quantity Discrimination, Number Properties, and Proportional Reasoning-which represent foundational skills critical for success in algebra. A total of 82 students who completed weekly progress monitoring measures over a period of 12 to 15 weeks were included in the analysis. Among the three measures, only Proportional Reasoning showed a significant increase in mean scores over time, indicating its sensitivity to student growth. These findings underscore the need for more targeted instructional supports and refined progress monitoring measures to meet the diverse needs of students with math learning difficulties. Practical implications highlight the ongoing challenges of implementing data-based individualization with fidelity and the critical role of sustained coaching and data-informed instructional decision-making in improving outcomes for students at risk in mathematics.
Childhood stress affects physical and mental health, making its proper assessment crucial. While several stress questionnaires for youth are available, their psychometric quality is questionable. Our aim was to develop a brief, age-appropriate questionnaire to measure current stress. Two-hundred thirty German children (6-17 years) completed the Stress Questionnaire for Children (SQC) and measures of stress symptoms, anxiety, depression, and quality of life via an online survey. The SQC consists of 17 items assessing current stress in school, social life, and leisure. An exploratory factor analysis indicated a three-factor solution (school stress, time stress, social stress) with a good model fit. The reliability of the total score (alpha = .90) and the subscales is good. Convergent validity was confirmed through correlations with stress-related symptoms, anxiety, depression, and quality of life. Total, school, and time stress increased with age, and girls endorsed higher stress than boys. Ratings on perceived difficulty, comprehensibility, and age-appropriateness indicated good acceptance. The SQC is a reliable and valid tool for assessing current stress in children and adolescents. It is child-friendly and suitable for a wide age range, making it suitable for research, clinical, and school settings. Further studies should confirm its factorial structure and its applicability across different cultural contexts.
Ongoing professional development is a critical component of high-quality early childhood education systems. To guide the content of such professional development, teacher and classroom quality assessments are often used. These assessments generally address universal or tier 1 instruction but omit information to guide teachers' practices to support children with disabilities. In addition, these assessments can be particularly onerous to deliver given that they require direct observation by a trained rater. As a step toward supporting the professional development of teachers serving children with disabilities, we evaluated a revised version of a newly developed resource-sensitive assessment called the Brief Preschool Progress Monitoring Measure. The assessment functioned as an online, test-based measure, to be completed by a teacher. The assessment provided information about teachers' abilities with collecting and using progress monitoring data to individualize instruction for children needing interventions and supports beyond those typically provided at tier 1 of a support system. Using Rasch analysis, findings revealed strong unidimensionality and item reliability, though limitations exist in detecting extreme ability levels. The revised assessment demonstrates potential as a tool for supporting targeted professional development initiatives and program evaluation in early childhood education but should not be incorporated into teacher accountability systems.
This study investigated the preliminary psychometric properties of the Pathway to Independence Inventory (P2I), a transition assessment tool designed to meet the needs of transition-age students with disabilities (SWD) with support needs in the areas of adaptive skills, executive functions, and social skills. Analyses examined the item-total correlations (ITC), factor structure, internal consistency, and concurrent validity evidence of the instrument and group differences based on gender identity, age, geographic location, and disability status. Results of the ITC and exploratory factor analysis (EFA) supported dimension reduction of items that did not discriminate well, resulting in separate versions of the instrument for students and informants. Results also indicated evidence of internal consistency, limited concurrent validity evidence for the informant version, and differences in mean scores based on gender identity and age for the informant report version and on disability status for both versions. Implications for both research and practice are discussed.
School climate plays a critical role in student well-being, academic success, and behavioral outcomes, yet the length of existing school climate surveys limits their feasibility for frequent administration within intervention frameworks such as Multi-Tiered Systems of Support (MTSS) and Positive Behavioral Interventions and Supports (PBIS). The purpose of this study was to evaluate the psychometric properties of a shortened school climate survey adapted from the validated U.S. Department of Education School Climate Surveys (EDSCLS)-Student Survey. A sample of 596 secondary students responded to 19 strategically selected items from the original survey. Using exploratory and confirmatory factor analysis, two items were removed due to low factor loadings, resulting in a refined 17-item measure with strong psychometric properties across three factors: embracing school diversity, school belonging, and student-teacher relationships. Findings suggest that this shortened survey maintains measurement integrity while enhancing practical usability within a data-driven intervention framework. Implications for integrating school climate assessment into MTSS and PBIS models to inform timely and effective school-wide interventions are discussed.
A reveal of how and why Diagnostique changed to Assessment for Effective Intervention.
Social and Emotional Learning (SEL) promotes positive mental health, strong relationships, and success in school and life. Identifying SEL skills and competencies relies heavily on self-report scales, but few of these scales have been developed and validated in Brazil, a country that requires all schools to implement SEL. We assessed 12,887 students (50% male) across five grade levels in three Brazilian states using a brief self-report measure that is based on the Collaborative for Academic Social and Emotional Learning's (CASEL) SEL framework. We conducted a Confirmatory Factor Analysis (CFA) of the measure, identified risk for below-average SEL using latent scores <= 1 SD below the mean, and evaluated the relationships between students' sociodemographic characteristics and SEL delay. Results of the CFA indicated acceptable fit, chi(2)(221) = 17,183.888, p < .001, comparative fit index (CFI) = .922, Tucker-Lewis index (TLI) = .911, root mean square error of approximation (RMSEA) = .077 (90% confidence interval [CI] = [.076, .078]), and standardized root mean square residual (SRMR) = .066 for the CASEL five-factor model including self-awareness, self-management, social awareness, relationship skills, and responsible decision-making. Results of the risk analyses indicated that race, grade level, and household size were associated with SEL risk status. Implications of these findings for future research and practice efforts are discussed.