The third National Charter School Study (NCSS III) aimed to test whether charter school were effective and to highlight outcomes on academic progress. The authors reported that typical charter school students outperformed similar students in non-charter public schools by 6 days in mathematics and 16 days in reading. This "days of learning" metric used to claim relatively higher performance in charter schools than in comparable public schools. This logic of this metric is critiqued in this paper, and an alternative method of reporting outcomes is proposed.
Camilli (2024) proposed a methodology using natural language processing (NLP) to map the relationship of a set of content standards to item specifications. This study provided evidence that NLP can be used to improve the mapping process. As part of this investigation, the nominal classifications of standards and items specifications were used to examine construct equivalence. In the current paper, we determine the strength of empirical support for the semantic distinctiveness of these classifications, which are known as "domains" for Common Core standards, and "strands" for National Assessment of Educational Progress (NAEP) item specifications. This is accomplished by separate k-means clustering for standards and specifications of their corresponding embedding vectors. We then briefly illustrate an application of these findings.
Natural language processing (NLP) is rapidly developing for applications in educational assessment. In this paper, I describe an NLP-based procedure that can be used to support subject matter experts in establishing a crosswalk between item specifications and content standards. This paper extends recent work by proposing and demonstrating the use of multivariate similarity based on embedding vectors for sentences or texts. In particular, a hybrid regression procedure is demonstrated for establishing the match of each content standard to multiple item specifications. The procedure is used to evaluate the match of the Common Core State Standards (CCSS) for mathematics at grade 4 to the corresponding item specifications for the 2026 National Assessment of Educational Progress (NAEP).
Test fairness, which had a late arrival on the assessment scene, has been closely linked to the earlier concept of test validity. However, among cognitive, social, and behavioral science professionals and the general public, the concept of test validity has been continuously debated since the 1950s. As a result, test fairness does not have a clear or generally agreed-upon fit within the topic of validity. The Standards for Educational and Psychological Testing (American Educational Research Association, American Psychological Association, & National Council on Measurement in Education, 2014) provides a useful starting point for examining fairness and its relationship to validity. However, the relationship of fairness and justice in assessment is a topic that extends well beyond the Standards. The goal of this chapter is to argue for a pluralistic and interdisciplinary approach to fairness and justice in order to promote more useful conversations about the social context of assessment. The authors illustrate how communication about fairness and justice can be sharpened with legal and philosophical illustrations.
Using results from surveys conducted in 2003 and 2018, we examined the perceived importance attributed to a set of specific tasks taught by faculty in required law school courses, with each questionnaire item in the survey corresponding to a task that describes a particular competency, such as critical reading. In this study, the results from two surveys completed by instructors of required law school courses are reported for a set of survey tasks representing competencies that regularly appear as topics in the legal education literature. We asked survey respondents to indicate the importance of each set of tasks for success in such courses. While one goal was to identify the competencies perceived as important by the faculty in required courses, another goal was to identify which competencies are becoming more or less important over time.
After 25 years with small to moderate gains in performance in mathematics, scores on the National Assessment of Educational Progress (NAEP) main assessment declined between 2013 and 2015 in Grades 4 and 8. Previous research has suggested the decline may be linked to the implementation of the Common Core state standards. In this article, the decline in the NAEP composite score is shown to be driven primarily by losses in the content strands of Geometry and of Data Analysis, Statistics, and Probability. A gain in fractions achievement is also evident in an item-level examination of the NAEP results, but not in reported NAEP scores. These effects are discussed with respect to the CCSS, the rationale for evaluating national progress, and a potential redesign of the NAEP assessment.
This study compared candidates' scores based on the normalised model and the two-parameter item response theory (2PL IRT) model using simulated multi-form exam data. Candidates' calculated scores, rankings, qualification status and score ties from the two models were compared with their true values. The results suggest that the 2PL IRT model outperformed the normalised model when the candidate ability distributions varied across forms. It was found that candidate scores based on the 2PL model were more closely related to the true scores. The qualification status of candidates belonging to the top 10% group were more accurately classified by the 2PL model than the normalised model when group abilities differed.
Free online Law School Admission Test (LSAT) preparation resources from the Law School Admission Council and Khan Academy have been widely utilized by LSAT and LSAT-Flex test takers. In September 2020, nearly 70,000 individuals engaged with Khan Academy’s Official LSAT® Prep platform. The purpose of this study was to examine the potential effects of engagement on actual LSAT performance. Our analyses showed that a higher level of engagement (measured in terms of practice time and number of practice exams taken) was associated with higher performance on the LSAT. These results held not only for the overall population but also across multiple demographic subgroups. The results also showed that the performance of test takers with lower initial practice exam scores was associated with slightly higher LSAT score gains per practice minute, indicating that these students benefitted at least as much as students who scored higher initially. Because this was a quasi-experimental controlled study, the possibility of alternative influences on LSAT performance cannot be ruled out. However, we believe that engagement with the Khan Academy platform is currently the best explanation for the LSAT score increases observed in this study.
The Law School Admission Council (LSAC) has a long-standing commitment to diversity, equity, and inclusion in legal education and in the legal profession. In line with its mission to promote quality, access, and equity in legal education, LSAC is providing a report, Understanding and Interpreting Law School Enrollment Data: A Focus on Race and Ethnicity, to help law schools, admission professionals, and other legal education stakeholders understand how we are measuring who is in the pipeline. The purpose of the report is to inform conversations about diversity, equity, and inclusion in law school and recruitment efforts. The report outlines the history of the Office of Management and Budget (OMB) data reporting standards, how these differ from LSAC data collection and reporting practices, and the social and cultural implications of different race and ethnicity data collection and reporting methods. The report includes examples of how the different methods affect conclusions that can be drawn from analyses of subgroup trends over time.
National profiles in mathematics achievement were obtained from a multidimensional analysis of item response data from the Trends in Mathematics and Science Study (TIMSS) for 2011 and 2015 by means of IRT factor analysis. Empirical subscores were then obtained at Grade 4 Grade 8. These subscores were less correlated than TIMSS-reported subscores, and revealed previously unknown information for some countries. The empirical subscores also resulted in ranks that differed from TIMSS-reported unidimensional ranks. In general, the analysis demonstrates that educational systems are not usefully described by a single mathematics score, or in terms of traditional subscores.
In large-scale international assessment programs, results for mathematics proficiency are typically reported for jurisdictions such as provinces or countries. An overall score is provided along with subscores based on content subdomains defined in the test specifications. In this paper, an alternative method for obtaining empirical subscores is described, where the empirical subscores are based on an exploratory item response theory (IRT) factor solution. This alternative scoring is intended to augment rather than to replace traditional scoring procedures. The IRT scoring method is applied to the mathematics achievement data from the Trends in International Mathematics and Science Study (TIMSS). A brief overview of the method is given, and additional material is given for validation of the empirical subscores. The ultimate goal of scoring is to provide diagnostic feedback in the form of naturally occurring item clustering. This provides useful information in addition to traditional subscores based on test specifications. As shown by Camilli and Dossey (2019), the achievement ranks of countries may change depending on which empirical subscore of mathematics is considered. Traditional subscores are highly correlated and tend to provide similar rank orders. •The methods takes advantage of the TIMSS sampling design, specifically pairs of jackknife zones, to aggregate categorical to higher-order sampling units for IRT factor analysis.•Once factor scores are estimated for sampling units and interpreted, they are aggregated to the jurisdiction level (countries, states, provinces) using sampling weights. The procedure for obtaining standard errors of jurisdictional level scores combines cross-sampling-unit variance and Monte Carlo sampling variation.•Full technical details of the IRT factoring procedures are given in Camilli and Fox (2015). Fox (2010) provides additional background for Bayesian item response modeling techniques. The estimation algorithm is based on stochastic approximation expectation-maximization (SAEM).
An assumption is often made that STEM shortages can be remedied by either increasing the number of STEM graduates or enlargingSTEMlabor supply through immigration. Yet in some STEM fields, there are classic signs of adequate supply or even oversupply. The issue is further complicated by nonlinear career dynamics and rapidly evolving international pressures. The goals of this special issue are to summarize the research, and more importantly, to go beyond thecurrent debate toidentify critical policy issues in preparing individuals for STEM careers that are personally satisfying and meeting the needs of industry and the public sector. A common theme is that broader skill sets will be required that span STEM and non-STEM fields. However, political and other expedient considerations have continued to shape workforce policies.
A stochastic approximation EM algorithm (SAEM) is described for exploratory factor analysis of dichotomous or ordinal variables. The factor structure is obtained from sufficient statistics that are updated during iterations with the Robbins‐Monro procedure. Two large‐scale simulations are reported that compare accuracy and CPU time of the proposed SAEM algorithm to the Metropolis‐Hasting Robbins‐Monro procedure and to a generalized least squares analysis of the polychoric correlation matrix. A smaller‐scale application to real data is also reported, including a method for obtaining standard errors of rotated factor loadings. A simulation study based on the real data analysis is conducted to study bias and error estimates. The SAEM factor algorithm requires minimal lines of code, no derivatives, and no large‐matrix inversion. It is programmed entirely in R.
The goal of the Law School Admission Council (LSAC) 2018 Skills Analysis Study is to identify the skills that law school faculty consider important for success in required law school courses. If certain tasks are required of all or most law school required courses, the skills involved in those tasks can be inferred to be essential to success in law school. This report provides evidence for assessing the validity of the current Law School Admission Test (LSAT), which will guide the development of new item types, item formats, and test specifications for future versions of the LSAT, including digital versions.
A summary is not available for this content so a preview has been provided. Please use the Get access link above for information on how to access this content.
Comparisons of US student achievement to other countries, conducted since the 1960s, have received extensive media coverage in the USA. Policy studies have cited these comparisons as evidence that the quality of US educational performance in mathematics and science requires federal intervention. Policy makers have been particularly concerned that the US economy would be impacted by inadequate student achievement levels especially in technology-related industries. This paper explores evidence about whether policy makers react to the citations of international student achievement rankings by changing funding levels for educational research and whether the studies have motivated the nation’s educators to support a trend toward common educational standards.
The trend in mathematics achievement from preschool to kindergarten is studied with a longitudinal growth item response theory model. The three measurement occasions included the spring of preschool and the spring and fall of kindergarten. The growth trend was nonlinear, with a steep drop between spring of preschool and fall of kindergarten. The modeling results provide validation for the argument that a classroom assessment in mathematics can be used to assess developmental skill levels that are consistent with a theory of early mathematics acquisition. The statistical model employed enables an effective illustration of overall gains and individual variability. Implications of the summer loss are discussed as well as model limitations.
This article focuses on the topic of how item response theory (IRT) scoring models reflect the intended content allocation in a set of test specifications or test blueprint. Although either an adaptive or linear assessment can be built to reflect a set of design specifications, the method of scoring is also a critical step. Standard IRT models employ a set of optimal scoring weights, and these weights depend on item parameters in the two-parameter logistic (2PL) and three-parameter logistic (3PL) models. The current article is an investigation of whether the scoring models reflect an intended set of weights defined as the proportion of item falling into each cell of the test blueprint. The 3PL model is of special interest because the optimal scoring weights depend on ability. Thus, the concern arises that for examinees of low ability, the intended weights are implicitly altered.
In item response theory (IRT), the scaling constant D = 1.7 is used to scale a discrimination coefficient a estimated with the logistic model to the normal metric. Empirical verification is provided that Savalei's [1] proposed a scaling constant of D = 1.749 based on Kullback-Leibler divergence appears to give the best empirical approximation. However, the understanding of this issue as one of the accuracy of the approximation is incorrect for two reasons. First, scaling does not affect the fit of the logistic model to the data. Second, the best scaling constant to the normal metric varies with item difficulty, and the constant D = 1.749 is best thought of as the average of scaling transformations across items. The reason why the traditional scaling with D = 1.7 is used is simply because it preserves historical interpretation of the metric of item discrimination parameters.
An aggregation strategy is proposed to potentially address practical limitation related to computing resources for two-level multidimensional item response theory (MIRT) models with large data sets. The aggregate model is derived by integration of the normal ogive model, and an adaptation of the stochastic approximation expectation maximization algorithm is used for estimation. This methodology is used to conduct an exploratory factor analysis of the 2007 mathematics data from Trends in International Mathematics and Science Study (TIMSS) fourth grade to illustrate potential uses. A comparison to flexMIRT and two brief simulations indicate the aggregate model provides accurate estimates of Level 2 parameters despite loss of information ensuing from key assumption.