
Abstract This paper traces more than a decade of collaboration between Bob Mislevy and Cisco, documenting his pivotal role in transforming educational assessment during the emerging digital revolution (2000–2014). The Cisco Networking Academy Program, serving millions of students globally, became the first large‐scale implementation of Evidence Centered Design (ECD) and a proving ground for his ideas. Five innovations are described: (a) application of ECD to a global assessment ecosystem; (b) NetPASS, a pioneering system for automated performance assessment using Bayesian inference networks; (c) integration of assessment into the Packet Tracer simulation platform; and (d) extensions to game‐based assessment through projects like SimCityEDU, and (e) broad impacts of the ECD logic on software development in general. The paper illustrates how Mislevy contributed to a paradigm shift from technology‐enhanced assessment to assessment‐enhanced technology and concludes with reflections on the intellectual and personal qualities that distinguished his contributions, including his integration of epistemological rigor with practical problem‐solving and his concern for the social impacts of assessment.
Abstract The analysis of eye‐tracking data has been shown to be useful in supporting inferences about the underlying processes of test takers responding to multiple‐choice questions (MCQs; Yaneva et al.), especially when framed as one piece of evidence within an overall validity argument. The same type of eye‐tracking data can also be collected with other item formats such as short‐answer questions (SAQs). In the present study, we leveraged eye‐tracking data to evaluate the extent to which test takers engage in different processes when responding to MCQs and SAQs. We randomly assigned medical students matched test items presented in either MCQ or SAQ format and collected eye‐tracking data as they responded to each item. We then compared various eye‐tracking measures (sequencing order, location, and duration of visual fixations) across the two item formats. The results provided little evidence of differences in the response process for the two item formats.
Abstract This research evaluates a Hybrid Angoff Method for standard setting procedures that utilize polytomous items. The method combines an Option Angoff with a Mean Angoff method to reduce the complexity of implementing judgments on polytomous items. The results were encouraging both with support for utilizing the method and panelist perceptions. The hybrid method provides a potential option for programs using polytomously scored items on their exams when implementing a standard setting panel.
In this essay, we recount Robert Mislevy's influence on the three of us and several others in the context of the National Research Council's Committee on the Foundations of Assessment, which produced the report Knowing What Student's Know: The Science and Design of Education Assessment (aka KWSK; NRC). We start by providing a context for the Committee's work, describe critical elements of the KWSK report, including their impact on the field of educational assessment over the last 25 years, and emphasize ways in which Bob influenced their articulation and impact.
As part of this EM:IP special issue honoring Robert (Bob) J. Mislevy, this paper highlights his critical role in shaping the emerging field of computational psychometrics, which integrates computational methods from data science and artificial intelligence (AI) with core psychometric principles to address the challenges posed by complex, high-volume multimodal data from interactive digital assessments. We discuss how Bob's vision, mentorship, and intellectual leadership helped to establish the conceptual foundations of this field, and how his foresight continues to guide the ongoing transformation of assessment in the era of AI.
Machine learning (ML) and generative artificial intelligence (AI) are rapidly transforming the field of educational measurement. This module focuses on illustrating the process of (1) automated machine learning (AutoML) using AutoGluon via an application of detecting aberrant test behavior and (2) AI-based item generation using Amazon Bedrock. To support these demonstrations, two tools are used: AutoGluon, an open-source automated machine learning (AutoML) system, and Amazon Bedrock, a fully managed AWS service for accessing foundation models from leading AI providers. By the end of this module, participants will (1) understand the key concepts and fundamentals underlying these two applications and (2) be able to programmatically train a classification ML model via AutoML using the provided data, as well as conduct AI-based item generation via the LLMs.
Traditional alignment studies that are narrowly focused on content standards and assessment items and are conducted postadministration often lack the timeliness and utility needed to improve assessment systems. This paper introduces a flexible alignment framework centered on assessment-system coherence. Our framework broadens alignment evaluation to include all interconnected elements, from domain definition to score use. We advocate for proactive alignment evaluation by defining expectations in advance and gathering both procedural and empirical evidence throughout the assessment life cycle. By addressing three guiding questions regarding system elements that need to be aligned, the purpose of the alignment evidence, and the types of evidence to use, developers can create customized evaluation plans to enhance alignment during development. The paper provides example alignment questions and an example application of the framework. We discuss the framework's potential to enhance the coherence of assessment systems in the service of meaningful policy and practice.
Bob Mislevy's conception of evidence-centered design (ECD) is one of the major contributions to modern validity theory. Building on prior work in psychometrics, philosophy, legal reasoning, and statistics, ECD is a generative framework developed to support the principled design or redesign and scaling of both existing and emerging assessment practices. This paper provides an overview of the context that Mislevy saw as an opportunity to reimagine assessment design, some of the history of ECD development, including several seminal efforts. I examine some of the impacts of ECD and consider how ECD might evolve in the future.
Bob Mislevy made several important contributions to large-scale educational survey assessments (LSAs). The contributions include implementing marginal maximum likelihood estimation of item response theory model parameters, using latent variable regression to condition proficiency estimates on background characteristics, and using multiple imputation (plausible values) to report assessment results. All these pieces were needed to analyze LSA data, report the results, and produce the data products needed for secondary users. In addition, Mislevy provided analyses of potential bias, giving users confidence that this novel approach to producing results yielded reliable information. Although procedures have since evolved, the approach and combination of techniques introduced in the 1984 NAEP reading assessment remain the cornerstones of LSA analysis and reporting in NAEP and internationally, some 40 years later.
Educational measurement has long been characterized by a productive tension between technical sophistication and the conceptual frameworks used to interpret what measurement outcomes mean. While the field has produced major methodological advances, comparatively fewer contributions have reshaped how psychometric evidence is connected to theories of cognition, epistemology, and use. This article examines Robert J. Mislevy's Sociocognitive Foundations of Educational Measurement as a landmark effort to address this imbalance. We argue that Mislevy's framework advances the field along three dimensions: reframing constructs through sociocognitive theories of learning and practice; grounding assessment in an epistemology of evidentiary reasoning under uncertainty, aligned with Bayesian inference; and articulating a principled pluralism across methodological traditions. Central to this contribution is a disciplined stance toward inference that treats epistemic restraint as rigor. We discuss the framework's relevance for contemporary challenges, including fairness, socially responsible use, and artificial intelligence in assessment.
This article describes what educational measurement can contribute to the theory and practice of classroom formative assessment. Three contributions are highlighted. The first contribution, assessment as evidentiary reasoning, serves as the conceptual frame within which the other two contributions are elaborated. Those contributions are (1) uncertainty and strategies for its reduction and (2) bias with a focus on approaches to facilitating fairness. The foundational concept running through all three contributions is that of inference as to what students know and can do at a given time, the basis for taking next instructional steps. Such inferences may be strengthened when response consistency is observed across problem variations generated by the teacher or by the teacher in partnership with Large Language Models (LLMs).
This study evaluates four clustering methodologies-Hierarchical Agglomerative Clustering (HAC), K-Prototypes, KAMILA, and Spectral Clustering-for detecting organized test collusion using mixed-type data. A two-phase design employed simulated data across 12 conditions and real certification exam data from a documented security breach. Results showed clustering efficacy increased substantially with higher exact response match rates and larger collusion group proportions. HAC emerged as the most robust method in simulations, while all four methods demonstrated exceptional convergence with empirically identified collusion groups in real data. These findings advance test security practice by providing empirically validated guidelines for selecting clustering methodologies and developing strategic implementation approaches, establishing a methodological foundation that equips testing organizations with enhanced collusion detection capabilities for evidence-based enforcement decisions.
The relationship between classroom assessment and educational measurement has been under discussion for some time. This article uses the TISM framework (Theory, Instrumentation, Scales and units, and Modeling) to clarify which aspects of classroom assessment are educational measurement (e.g., a grade on a performance assessment keyed to a learning standard) and which are not (e.g., extended elaborated feedback on that same assessment). We conclude that classroom assessment which produces ordinal or interval-level quantitative scores-by whatever name they are called, including scores, grades, and performance levels-is educational measurement because it implicates theory, instrumentation, scales and units, and modeling of error. On this basis, we claim that work in classroom assessment and educational measurement can and should be mutually informative.
In large-scale assessments, constructed response items are often scored using hybrid scoring systems, which combine human and automated scores. In this study, we augment automated scoring with confidence modeling to strategically route difficult-to-score responses for human review. We utilize hybrid performance curves to visualize the impact of routing on performance. Additionally, we propose several hybrid scoring policies for selecting optimal routing thresholds given practical constraints. Our findings reveal that hybrid scoring systems can achieve an overall performance that exceeds that of human- and automated-only systems. Moreover, the superior performance of the hybrid system is less expensive than a human-only system. These findings highlight the complementarity of human raters and automated scoring engines. Although current standards focus on the performance of human raters and automated scoring engines in isolation, we recommend that practitioners also report on the performance of the hybrid scoring system as a whole.
Large language models (LLMs) have been widely explored for automated scoring in educational assessment to facilitate learning and instruction. However, empirical evidence regarding which LLMs produce the most reliable scores and induce the least rater effects remains limited. This study compared 10 LLMs (ChatGPT 3.5, ChatGPT 4, ChatGPT 4o, OpenAI o1, Claude 3.5 Sonnet, Gemini 1.5, Gemini 1.5 Pro, Gemini 2.0, DeepSeek V3, and DeepSeek R1) with human expert raters in scoring two types of writing tasks. Their performance was evaluated in terms of score accuracy, intra-rater consistency, and rater effects estimated using the Many-Facet Rasch model. Although the results generally supported the use of ChatGPT 4o, Gemini 1.5 Pro, and Claude 3.5 Sonnet with high scoring accuracy, better intra-rater consistency, and less rater effects, the study is not intended to support substantive comparisons or rankings of LLMs or to identify a single "best" model, given the small sample size.
Diagnostic classification models (DCMs) assess students' mastery of cognitive attributes to provide personalized ability profiles. Retrofitting DCMs to large-scale mathematics assessments usually relies on inferred Q-matrices, which can reduce accuracy and diagnostic value. This study evaluated whether constructing items from cognitive models-yielding Q-matrices directly-and incorporating hierarchical relationships among attributes improve diagnostic outcomes. Responses from 5,336 third-grade students to a Luxembourgish image-based, large-scale standardized mathematics exam were analyzed using multiple DCMs and their hierarchical extensions. Items were constructed based on a Q-matrix, derived from the curriculum and cognitive models. The hierarchical A-CDM outperformed other models, classifying students into 60 latent classes with acceptable attribute- and test-level accuracy and more interpretable results than the G-DINA model. Using cognitive model-based item generation and Q-matrices as well as specifying attribute hierarchies enhance the accuracy and interpretability of DCM-based diagnostics in large-scale assessments, complementing traditional psychometric approaches by discerning meaningful within-score differences.
Unlike standardized testing applications of validity, teachers need a simple and efficient way to reflect on the accuracy of the claims based on student performance, then consider whether the uses of those claims are appropriate. A two-phase reasoning process of validation, consisting of a proficiency claim /argument and a use/argument, is presented as a way for teachers to understand and apply the central tenets of validation to their classroom assessments. Since classroom assessment is contextualized with multiple purposes, each teacher is obligated to use validation for their situation. The accuracy of teachers' conclusions about the proficiency claims, and uses, will depend on their skill in gathering supportive evidence and considering alternative explanations. Examples of the proposed classroom assessment validation process are presented.
To assess the interrater reliability of human ratings of constructed responses (CR), or the accuracy of scores given by automated scoring engines, concordance metrics quantify agreement between measures. This article examines the quadratic weighted kappa (QWK) in these contexts and highlights its practical limitations compared to other metrics. Both empirical and simulation study results reveal how different factors including the shape of the marginal distributions and score scale length may impact the estimates and how we can adjust for these properties of the contingency table. The results highlight the QWK's sensitivities and suggest that additional caution should be taken before decisions about whether to keep a CR item on a test form are made. If using QWK without the proper interpretive supports, such decisions may be misinformed. Consequently, we make suggestions for best practices to promote responsible evaluation of agreement in the context of CR scoring in educational testing.
The rapid advancement of large language models (LLMs) has enabled the generation of coherent essays, making AI-assisted writing increasingly common in educational and professional settings. Using large-scale empirical data, we examine and benchmark the characteristics and quality of essays generated by popular LLMs and discuss their implications for two key components of writing assessments: automated scoring and academic integrity. Our findings highlight limitations in existing automated scoring systems when applied to essays generated or heavily influenced by AI, and identify areas for improvement, including the development of new features to capture deeper thinking and recalibrating feature weights. Despite growing concerns that the increasing variety of LLMs may undermine the feasibility of detecting AI-generated essays, our results show that detectors trained on essays generated from one model can often identify texts from others with high accuracy, suggesting that effective detection could remain manageable in practice.
Value-added models (VAMs) are both common and controversial in education policy and accountability research. While the sensitivity of VAM results to model specification and covariate selection is well documented, the extent to which test scoring methods (e.g., mean scores vs. item response theory based scores) may affect Value-added (VA) estimates is less studied. We examine the sensitivity of VA estimates to the scoring method using empirical item response data from 18 education datasets. We find that VA estimates can be sensitive to the choice of scoring method, holding constant students and items. While the various test scores are highly correlated, on average, using different scoring approaches leads to variation in VA percentile ranks of over 20 points, and more than 50% of teachers or schools are classified in multiple quartiles of the VA distribution. Dispersion in VA ranks is reduced with more complete item response data. Our findings suggest that consideration of both measurement error and model uncertainty are important for the appropriate interpretation of VAMs.