
Improving early-grade reading in linguistically diverse, resource-constrained settings remains challenging. This study evaluated a nationwide bilingual reading intervention in Ghana that provided reading materials, teacher guides, sequenced training, and ongoing coaching. Using a randomized trial with 3,559 second- and third-grade students in 198 schools, we examined program impacts and how student, teacher, and school languages moderate these effects. The intervention improved reading outcomes across letter-sound knowledge, decoding, fluency, and comprehension, yielding main effect sizes of 0.09 to 0.43 standard deviations for Ghanaian languages and 0.12 to 0.37 standard deviations for English. These gains were amplified when students and teachers shared the same first language for classroom oral exchanges. Effect sizes for students in concordant student-teacher pairs were 0.17 to 0.42 standard deviations higher than those for students in discordant pairs. These results show that large-scale reading improvements are achievable when classroom instruction is strengthened, and language-of-instruction policies account for teacher-student language concordance, not just school- or district-level assignments.
Randomized controlled trials (RCTs) produce impact evidence with high internal validity but uncertain external validity. External validity bias has been defined as the expected difference between the average impact in the sample and the average impact in the population. This study estimated the external validity bias from several site selection methods by simulating hypothetical RCTs of the Head Start program in which some selected sites decline to participate. Three main findings emerged from the analysis. First, purposive site selection consistently produced biased impact estimates that varied in magnitude based on the outcome examined, the factors used to select sites, and the factors that influenced site decisions to participate. Second, simple random site selection yielded less external validity bias than purposive site selection under most tested conditions. Third, stratified random site selection yielded virtually no external validity bias, but the results likely overstate the method's performance when data on impact moderators are unavailable before sites are selected. These findings offer lessons on how to select sites in future RCTs to minimize external validity bias.
The study presents evidence from a multi-arm cluster randomized evaluation examining the impact and cost-effectiveness of three teacher professional development (TPD) components implemented within a large-scale early grade reading intervention in Pakistan. A sample of 200 schools, 185 teachers and 3068 and Grade 1 students was randomly assigned to receive a full TPD package -including face-to-face training (FtF), teacher inquiry groups (TIGS), and school support visits (SSVs)- or one of three reduced variants omitting one component. Teacher instructional practices were measured using a Teacher Classroom Observation tool, student literacy using the Early Grade Reading Assessment (EGRA), and cost using an ingredients-based approach. All TPD components improved instructional practices; however, only SSVs generated statistically significant gains in reading outcomes (0.23 standard deviations in overall literacy and 0.30 in reading comprehension). Although SSVs were the most expensive component (US$6 per child) they were also the most cost-effective. FtF (US$3 per child) and TIGs (US$5 per child) improved instruction but did not yield detectable literacy gains. Findings provide policy-relevant evidence on cost effective TPD in conflict affected settings.
Value-added measures (VAMs) are used to hold teachers accountable for their teaching effectiveness, yet their interpretation and communication are challenged by both statistical uncertainty and vulnerability to systematic bias. To complement conventional measures of uncertainty (e.g., standard errors or confidence intervals), we apply and extend the Robustness of Inference to Replacement (RIR) framework to quantify how sensitive VAM-based evaluations are to departures from key modeling assumptions, including violations of the Stable Unit Treatment Value Assumption such as peer effects and heterogeneity. RIR captures robustness to systematic bias rather than sampling variability alone, resulting in statements such as "One would have to replace __ % of the students in the classroom with average students to increase the teacher's VAM above the threshold used to categorize performance." Ultimately, RIR is intended primarily as a tool for research analysts and methodologists who evaluate and communicate VAM results, offering a counterfactual framing that aligns with how educators reason about classroom composition. Using data from Project STAR (Student-Teacher Achievement Ratio), we demonstrate how RIR can enrich interpretation of VAMs and support more transparent discussions.
Research on teacher preparation programs (TPPs) continues to debate the extent to which program quality meaningfully shapes teacher effectiveness. While early evidence from high-income contexts documented nontrivial differences across programs, more recent studies suggest that most variation in teacher effectiveness occurs within rather than between programs. Using national administrative data from Chile and a three-level value-added model linking students, teachers, and preparation programs, this study estimates the share of variance in student achievement attributable to TPPs. Focusing on novice mathematics teachers linked to the only cohort with available lagged test scores (students assessed in 2015 and 2017), we find that TPPs account for approximately 1% of the total variance in student outcomes, while substantially more variation occurs within programs across teachers. Rather than interpreting this limited differentiation as evidence of uniformly strong preparation, we discuss how these patterns are consistent with institutional convergence and bounded instructional learning in a highly regulated system. The study contributes new empirical evidence from Latin America and offers a theoretically grounded interpretation of how accountability regimes may shape the distribution-rather than the overall level-of teacher effectiveness across preparation programs.
In recent years, scholars across the social and educational sciences have increasingly advanced a subtle but consequential falsehood-that specific instances of research fall into two categories, "causal" research, which uses a finite set of empirical strategies to clearly establish cause and effect relationships, and all other, non-causal research. We term this false logic "the causal dichotomy." In this article we trace the origins of the causal dichotomy, review why causal robustness lies on a continuum, provide evidence of its prevalence from a novel coding of education journals, and discuss its potential consequences. Fortunately, the causal dichotomy can be easily avoided, even as researchers continue to maintain an emphasis on inferential rigor that benefits the field.
There is considerable variability in the literacy assessments taken in Kindergarten through second grade, across schools and between multilingual learners and other students, and within students over time. This makes it difficult to study changes in students' acquisition of ELA skills in these formative years, or to evaluate policies and practices meant to support literacy development. Here we examine several popular early grade assessments-the MAP, ACCESS, DIBELS, TRC English & Spanish versions, and apply a novel approach to combining information to develop latent scores of students' literacy development. We find each assessment provides information that is predictive of students' development toward third grade literacy outcomes (ELA grades and state assessment scores), with different strengths and weaknesses, and considerable overlap among them. We further provide evidence of strong predictive validity for the combined scale, even in post-COVID-19 years, suggesting that we could leverage existing assessment information to produce metrics for studying school, district, and state policies and practices around literacy development.
Building on prior research that found more restrictive educational placements for secondary than elementary students, we explored grade-level placement variations of the population of 171,216 Utah public school students receiving special education services between kindergarten and 11th grade (2016-2022). Of these students, 9,037 (5%) were eligible for the state alternative assessment based on significant cognitive disability (SCD). Linear regression predicting mean placement level showed that students without SCD in later grades tend to have more restrictive placements due to a different distribution of disability categories-specifically, more students with Specific Learning Disability and Other Health Impairments and fewer students with Speech Language Impairment. Multilevel linear regression showed SCD had the strongest negative effect on placement, followed by disability categories often associated with complex support needs (multiple disabilities, intellectual disability, autism). Prior research and federal reports indicating increasing inclusion rates of students with disabilities across years only apply to students without SCD. Highly segregated placements for students with SCD remain largely unchanged across grades and years. Our findings highlight a substantial, longstanding gap in improving inclusive opportunities for students with SCD.
Education researchers are increasingly interested in understanding "what works" for whom and under what conditions, often through analyses of subgroup findings. Yet such findings are frequently reported without strong empirical or theoretical justification, raising concerns about their interpretability. Drawing on the What Works Clearinghouse (WWC) study database, this study uses meta-analysis to examine systematic patterns in subgroup and compositional effects across educational evaluations and investigates empirical explanations for those patterns. Within studies, we find that treatment effects are modestly more favorable for economically disadvantaged students and female students and somewhat less favorable for high-performing students. Across studies, evaluations with higher proportions of economically disadvantaged students or students with disabilities have larger average treatment effects, but these compositional effects do not persist when accounting for variation in intervention characteristics. Despite these systematic patterns, we observe considerable treatment effect heterogeneity and wide prediction intervals, highlighting the need to interpret average treatment effects and subgroup findings with caution. We discuss implications for researchers, funders of research, and policymakers working to improve the design, evaluation, and equitable impact of educational interventions.
Randomized control trials (RCTs) are the most common type of experimental research design and are widely regarded as the gold standard for investigating causal effects. A key reason RCTs enable valid estimation of intervention effects is that random assignment to treatment conditions theoretically minimizes or eliminates confounding by balancing outcome-related covariates (e.g., pretest scores and demographic characteristics) across groups. This assumption generally holds in large-sample studies, where the law of large numbers ensures covariate balance-an asymptotic property of randomization. However, in small-sample studies, chance imbalances may occur, potentially biasing estimates of the intervention's effect. The present study highlights this often-overlooked issue and synthesizes approaches from educational research and other behavioral science fields into a structured implementation guide. This guide aims to help researchers systematically address the practical challenge of confounder imbalance in small-sample RCTs. To illustrate its application, we include a real-world example from special education research, where small-sample studies are common due to the need to pilot novel interventions before larger efficacy trials, limited target populations (e.g., students with disabilities), and the high operational costs of individualized interventions.
The Workforce Innovation Opportunity Act of 2014 prioritizes supporting disabled youth and young adults by requiring Vocational Rehabilitation (VR) agencies and schools collaboratively deliver pre-employment transition services (Pre-ETS). The Pre-ETS are comprised of five key areas that impact employment outcomes: (a) counseling on postsecondary education opportunities, (b) instruction in self-advocacy, (c) job exploration counseling, (d) workplace readiness training, and (e) work-based learning experiences. This study examined the relationship between Pre-ETS and VR services on postsecondary education participation among youth with disabilities (ages 14-24). Additionally, we established a process to track exit and reentry specifically as a precursor to design longitudinal studies using the Rehabilitation Services Administration (RSA-911), a national administrative dataset that includes over 15,000 youth and young adults with disabilities who receive VR services. Our findings show that the combination of certain Pre-ETS (counseling on postsecondary education opportunities) and receipt of individualized VR services during postsecondary enrollment is particularly important for disabled youth. Implications for research and practice that enhance collaborative structures among professionals are discussed.
Staff turnover is a critical issue for practitioners and policymakers alike. While extant research indicates teacher turnover has a detrimental impact on student outcomes generally, students with disabilities may be disproportionately impacted, as they receive services from both general educators and from special educators and paraeducators, who have higher turnover rates. Yet, prior research has not examined effects of turnover on students with versus without disabilities. Using administrative data from Washington State, we provide the first causal estimates of special and general education teacher and paraeducator turnover on outcomes for students with and without disabilities. We find general education teacher turnover negatively impacts test scores for both groups and there is an outsized impact of special education teacher turnover on test scores for students with disabilities. Estimates of paraeducator turnover, however, are small and not statistically significant. Our findings suggest the stability of the general education teacher workforce is important for all students, and the stability of the special education teacher workforce should be a key priority for improving outcomes for students with disabilities.
The Individuals with Disabilities Education Act (IDEA) mandates that special education teachers be fully licensed, either by obtaining state certification or passing a licensure examination. However, despite this federal mandate, many states have historically issued emergency licenses to fill special education teacher positions due to persistent shortages in the workforce. We use longitudinal state data from Indiana spanning years 2012 to 2021 to document the proportion of special education teachers working under emergency licensure, the students they serve, and the instructional settings in which they work. We find that emergency-licensed teachers make up a growing proportion of the special education teacher workforce, and that they serve more students with autism and with intellectual disabilities in more restrictive settings than their non-emergency licensed counterparts.
As more students with disabilities are educated in inclusive classrooms, dual licensure policies have emerged as a promising strategy to improve the quality of special education teaching. However, there are concerns that these policies might worsen the chronic shortage of special education teachers by introducing additional requirements for teaching candidates. This study seeks to explore these concerns by examining the association between state dual licensure policies and the number of prepared special education teachers. We analyzed Title II data in conjunction with policy enactment data collected from various sources in three states. Our findings indicate that, overall, the policy was not significantly associated with the number of prepared teachers. We discuss these results in the context of their implications for future research and policy.
Early intervention (EI) and early childhood special education (ECSE) services for children with disabilities have expanded substantially across the U.S. over the past few decades, necessitating efforts to recruit and retain a qualified workforce to meet their needs. Despite widespread reports of staffing challenges in this sector, few contemporary studies provide large-scale evidence on this workforce. Using administrative data for all EI/ECSE employees in Oregon from 2008 to 2023, we provide longitudinal descriptive evidence on their composition, distribution, and stability. We show that the workforce has increased significantly, is growing more racially/ethnically diverse, and is more highly educated but less experienced than the state's K-12 workforce. Turnover remained fairly constant during this period, with the exception of paraprofessionals and non-licensed staff whose retention steadily declined to historic lows. Finally, we show that staff are distributed somewhat inequitably throughout the state, with areas serving more low-income students having the highest child-staff ratios and fewer highly-educated teachers/interventionists. Together these analyses contribute the first longitudinal portrait of an EI/ECSE workforce, providing key insights into their staffing dynamics at scale.
This article reports empirical evidence to support the design of evaluations that estimate the impacts of programs that provide postsecondary credentials and/or job training on earnings. Statistical power analyses are strengthened by having accurate empirical estimates of the standard deviation of earnings, share of earnings variance explained by covariates (R2), and, for some designs, the intra-class correlation (ICC). We compute and report values of these inputs for a large sample of control group members from three large studies of such programs. Using our estimated properties of quarterly earnings, we calculate the minimum sample size needed to detect an earnings impact of reasonable magnitude. This calculation demonstrates that many recently published experimental program evaluations have samples sufficient to detect only very large earnings impacts.
Ample research investigates returns to teacher preparation and other instructional inputs for the general student population, yet evidence is lacking for students with disabilities (SWDs). This study uses North Carolina data to estimate achievement returns to teacher preparation by classroom type and level of classroom support for SWDs. The findings show that SWDs perform better when placed in integrated classrooms and when these classrooms have co-teachers. Regardless of classroom type, SWDs benefit from more experienced teachers, but only gain from special education certified teachers in certain classroom configurations. These results indicate that education leaders can optimize resource allocation by reducing reliance on separate classrooms for SWDs, investing in an experienced teacher workforce supported by co-teachers, and aligning certification requirements and training more closely with the diverse needs of SWDs.
This paper uses a propensity score weighting approach to explore how the impact of dual enrollment varies by locale (rural and urban/suburban settings); we then apply an established methodological framework to systematically explore possible reasons for variation in impact. Our findings suggest that students in rural settings benefit more from dual enrollment than students in urban settings. A large part of the explanation appears to be the counterfactual experience of students, particularly the fact that urban students have higher levels of access to other advanced courses. We also see increased support services for dual enrollment in rural settings as well as differential impacts by locale for certain subgroups.
Universal, school-based social-emotional learning (SEL) interventions hold the promise to promote children's academic learning and social and emotional skills necessary to cope with adversity in the context of education in conflict and crisis. This study evaluated the sequential implementation of two sets of skill-targeted SEL activities-first, testing the impact of 11-week implementation of Mindfulness activities alone, then the combined effect of Mindfulness and Brain Games implemented sequentially over the course of 22 weeks-with Nigerian refugee and Nigerien local second to fourth graders (N = 1,795) attending a remedial education programming in Diffa, Niger, via a cluster randomized controlled trial. Positive impacts of Mindfulness activities were observed on a select number of children's SEL outcomes, with some additional benefits for refugees and girls, demonstrating the potential for low-cost SEL approaches to support vulnerable children at scale in crisis-affected contexts. These findings provide actionable insights for policymakers, program developers and implementers in conflict-affected, low-income countries, such as the potential for targeted interventions to improve SEL outcomes for marginalized children and the need to invest in contextually-responsive and sustained program models to maximize long-term impact.
Recently, researchers have advocated for adopting a strengths-based approach to investigate the academic and socio-emotional development of historically minoritized and linguistically diverse students. However, there is little guidance on how to apply this approach in practice. Here, we outline the theoretical framework and key principles underlying a strengths-based approach. We illustrate how to apply this approach in various stages of the research process, from program design and evaluation to result interpretation and dissemination, using three examples from recent Randomized Controlled Trials (RCTs). These examples have successfully enhanced educational outcomes for kindergarten, elementary, and middle school students of Latino and African American backgrounds. We offer five recommendations for education researchers: 1) maintain collaboration with the community and other stakeholders at every stage in the research process; 2) evaluate the program using both standardized and culturally relevant measures; 3) consider the challenges of customizing for one group vs. generalizing across different groups; 4) test strengths-based approaches against other intervention approaches; and 5) secure supplemental funds for this work. The five recommendations do not need to be adopted all at the same time.