The usual interpretation of the person and task variables in between-persons measurement models such as item response theory (IRT) is as attributes of persons and tasks, respectively. They can be viewed instead as ensemble descriptors of patterns of interactions among persons and situations that arise from sociocognitive complex adaptive system (CASs). This view offers insights for interpreting and using between-persons measurement models and connecting with sociocognitive research. In this article, we use data generated from an agent-based model to illustrate relations between "social" and "cognitive" features of a simple underlying CAS and the variables of an IRT model fit to resulting data. We note how the ideas connect to explanatory item response modeling and briefly comment on implications for score interpretations and uses in practice.
Sijtsma, Ellis, and Borsboom (Psychometrika, 89:84-117, 2024. https://doi.org/10.1007/s11336-024-09964-7 ) provide a thoughtful treatment in Psychometrika of the value and properties of sum scores and classical test theory at a depth at which few practicing psychometricians are familiar. In this note, I offer comments on their article from the perspective of evidentiary reasoning.
Rapid advances in psychology and technology open opportunities and present challenges beyond familiar forms of educational assessment and measurement. Viewing assessment through the perspectives of complex adaptive sociocognitive systems and argumentation helps us extend the concepts and methods of educational measurement to new forms of assessment, such as those involving interaction in simulation environments and automated evaluation of performances. I summarize key ideas for doing so and point to the roles of measurement models and their relation to sociocognitive systems and assessment arguments. A game-based learning assessment SimCityEDU: Pollution Challenge! is used to illustrate ideas.
Evidence-centered design (ECD) is a comprehensive framework that makes explicit the structure of assessment arguments, the elements and processes through which they are instantiated, and the interrelationships among them. This article provides an overview of ECD. It highlights the idea of layers in this process, structures and representations within each layer, and terms and concepts that can be used to guide the design of educational assessment. Examples from large-scale assessment and game-based assessment illustrate how ECD helps assessment designers articulate the design process.
An overarching mission of the educational assessment community today is strengthening the connection between assessment and learning. To support this effort, researchers draw variously on developments across technology, analytic methods, assessment design frameworks, research in learning domains, and cognitive, social, and situated psychology. The study lays out the connection among three such developments, namely learning progressions, evidence-centered assessment design (ECD), and dynamic Bayesian modeling for measuring students' advancement along learning progression in a substantive domain. Their conjunction can be applied in both formative and summative assessment uses. In addition, this study conducted an application study in domain of beginning computer network engineering for illustrating the ideas with data drawn from the Cisco Networking Academy's online assessment system.
• Background: This study advances a sociocognitive approach to modeling complex communication tasks. Using an integrative perspective of linguistic, cultural, and substantive (LCS) patterns, we provide a framework for understanding the nature and acquisition of people’s adaptive capabilities in social/cognitive complex adaptive systems. We also illustrate the application of the framework to learning and assessment. As we will show, understanding the connection between measurement models and users’ needs is important to increase assessments’ educative usefulness.
Virtual performance-based assessments (VPBAs) are environments for test takers to interact with systems, sometimes including other persons or agents, in order to provide evidence about their knowledge, skills, or other attributes. Examples include tasks based on interactive simulations, games, branching scenarios, and collaboration among students communicating through digital chats. They may be used for summative purposes, as in certification examinations, or for other purposes, as in intelligent tutoring systems and exploratory learning environments. They afford opportunities to obtain direct evidence about capabilities that inherently involve interaction, such as inquiry and collaboration. Our focus here is digital, usually with regard to the environment but always with regard to the form of data. Digital data capture makes it possible to acquire rich details about students' actions and the evolving situations in which they occur. The challenges they pose to psychometrics lie in designing VPBAs to optimally evoke the targeted capabilities, providing students with affordances that evidence that cognition, capturing the relevant aspects of the performances, identifying meaningful patterns in performances that constitute evidence about the targeted capabilities, and providing an inferential framework for synthesizing the evidence and characterizing its properties. This chapter provides an introduction to VPBAs and psychometric considerations in VPBA design and analysis.
In this chapter we articulate what is computational psychometrics, why we need a volume focused on it, and how this book contributes to the expansion of psychometric toolbox to include methodologies from machine learning and data science in order to address the complexities of big data collected from virtual learning and assessment systems. We also discuss here the structure of the edited volume, how each chapter contributes to enhancing the psychometrics science and our recommendations for further readings.
Digitally based learning and assessment systems generate large volumes of complex process data. The next generation psychometricians need to acquire new data science skills to meet the data challenge. In this chapter, we summarize data science skills and identify the subset that psychometricians need to prioritize. We introduce an evidence identification centered data design (EICDD) process during the task design, as an important way to address the data challenges from digitally based assessments. We describe some specific data techniques to parse and process complex process data with example codes in Python programming language. We also outline the general methodological strategies when dealing with process data from digitally based assessments.
This chapter surveys the history of Bayesian inference in educational measurement and testing. The mid-century revival of Bayesian inference led the measurement community to look backward and forward: backward to recast earlier developments in Bayesian terms, and forward to advance the field of measurement by way of using Bayesian approaches. In the 1970s, researchers began investigating Bayesian approaches to tackle problems in criterion-referenced testing. This work was closely related to the research on prediction reviewed. Bayesian efforts addressed this principal problem of classifying students with respect to the criterion, as well as problems of determining test length, setting cut scores, evaluating reliability, and marrying probabilistic classifications with utilities in a decision-theoretic frame. Concurrently, researchers such as Embretson were drawing on developments in cognitive psychology to define constructs, construct items, and build measurement models to analyze resulting data. In the 1980s, work in artificial intelligence began to develop in ways that would soon influence educational measurement in several ways.
The validity of the human scores on which the automated score is based must also be established. Wolfe shows us that much more is going on in human scoring than meets the eye and how to investigate the many issues of evidentiary reasoning involved, how to articulate alternative explanations that arise, and how to identify sources of backing that can be marshaled. 'Scoring' is a term inherited from familiar processes of comparing multiple-choice responses with keys and eliciting evaluative ratings from human raters. The warrants posit that the scores characterize targeted qualities of a performance and that the raters are providing those scores on the basis of those qualities and with sufficient accuracy. The correspondence between the assessment design/interpretation elements and corresponding elements of the score-use situations can render scores more or less informative for various subsequent uses, quite beyond the construct labels assigned to the scores of performances and the assessment as a whole.
This chapter provides an overview of changes in the higher education environment that inform admissions and placement decisions. Various factors that must be considered when reconceptualizing current admission and placement practices are discussed. An expanded assessment framework based on two models – the multilevel design model and the complementarity model – are described. These models aim to better support diverse students' learning by improving the connection between assessments and instruction once students are admitted to higher education institutions. Finally, the contributions of technological advancements, measurement of noncognitive skills, and innovations in task design are described.
In his 2019 NCME Career Contributions Award address, Dr. Shelby Haberman uses examples of three kinds to illustrate how his training in theoretical statistics influenced his contributions to educational measurement. I bracket my comments on his address, and his contributions more generally, by considering two questions: Why might any theoretical statisticians receive this award? Why aren't all recipients theoretical statisticians? Through the course of the discussion emerges the answer to a third, easier, question: Why Shelby Haberman?