Among Broadfoot's many contributions is the four 'C's framework, which includes the insight that one of the social functions of assessment is to impose particular educational content through assessment design. A second insight is that 'An individual's performance on a particular task, cannot be divorced from their human characteristics - from culture, from personality, from motivation - even from the external conditions prevailing at the time the performance is measured'. In this piece, I elaborate on these insights. I note the connection of content with method, and that the interaction of content, method, human characteristics, and external conditions inevitably lead to assessments that may perpetuate inequity. A potential approach to achieving greater equity may be to personalise assessment to the characteristics of the individual and the environments from which they come.
Personalized assessments are of increasing interest because of their potential to lead to more equitable decisions about the examinees. However, one obstacle to the widespread use of personalized assessments is the lack of a measurement toolkit that can be used to analyze data from these assessments. This article takes one step toward building such a toolkit by proposing a validation framework for personalized assessments. The framework is built on the threats-to-validity approach. We demonstrate applications of the suggested framework using the AP 3D Art and Design Portfolio examination and a more restrictive culturally relevant assessment as examples.
Using the method of narrative review, this paper considers the impact of structural inequity in US society and its implications for educational assessment. Focusing on African Americans, some of the many past and present examples of structural inequity and their effects are delineated. Described next are how these effects can be connected to opportunity to learn (OTL), very broadly conceived, and how the persistence of so-called achievement gaps might be seen from that OTL perspective. Based on the conception derived from the review, a graphical representation is given positing how structural inequity, through OTL, works to constrain achievement and life chances cumulatively over time. Understanding OTL from this broad-ranging, cumulative perspective suggests ways in which the design and use of K-12 and higher education assessment might be rethought.
Over our field's 100-year-plus history, standardization has been a central assumption in test theory and practice. The concept's justification turns on leveling the playing field by presenting all examinees with putatively equivalent experiences. Until relatively recently, our field has accepted that justification almost without question. In this article, I present a case for standardization's antithesis, personalization. Interestingly, personalized assessment has important precedents within the measurement community. As intriguing are some of the divergent ways in which personalization might be realized in practice. Those ways, however, suggest a host of serious issues. Despite those issues, both moral obligation and survival imperative counsel persistence in trying to personalize assessment.
This entry focuses on automated scoring of questions calling for answers that cannot be graded using exact-matching techniques and that are used operationally in assessment programs and learning products. The entry describes the benefits of automated scoring, the contexts in which it has been used, and approaches to validation. Operational uses are briefly discussed for the automated scoring of essay writing, spoken responses, short text responses, STEM responses, and response processes. A discussion of emerging trends closes the entry.
“Toward a Theory of Socioculturally Responsive Assessment” assembled design principles from multiple literatures and wove them into a working definition and a network of empirically testable propositions. The intention was to offer a coherent theoretical framework within which to understand why and how particular assessment designs might work, what actions testing programs should consider, how they might move forward with those actions, and how to evaluate the impact. Dr. Solano Flores offers many comments on these ideas, with which I mostly agree. In this response, I detail those agreements, as well as some points of departure. I close with some implications for revising the Standards.
In the United States, opposition to traditional standardized tests is widespread, particularly obvious in the admissions context but also evident in elementary and secondary education. This opposition is fueled in significant part by the perception that tests perpetuate social injustice through their content, design, and use. To survive, as well as contribute positively, the measurement field must rethink assessment, including how to make it more socioculturally responsive. This paper offers a rationale for that rethinking and then employs provisional design principles drawn from various literatures to formulate a working definition and the beginnings of a theory. In the closing section, a path toward implementation is suggested.
Grouping individuals according to a set of measured characteristics, or profiling, is frequently used in describing, understanding, and acting on a phenomenon. The advent of computer-based assessment offers new possibilities for profiling writing because aspects can be captured that were not heretofore observable. We explored whether writing processes could be profiled of over 30,000 adults taking a high-school equivalency examination. Process features were extracted from keystroke logs, aggregated into composite indicators, and used with essay score to assign individuals to profiles. Analyses included computing the percentages of individuals that could be classified, using MANOVA to examine differences among profiles on external variables, and examining if profiles could be distinguished from one another based on patterns derived from cluster analysis. Results showed that about 30% of examinees could be classified into profiles that were largely distinct. These results contribute toward a foundation for using such profiles in describing how individuals compose and in how their writing might be improved.
This commentary focuses on one of the positive impacts of COVID-19, which was to tie societal inequity to testing in a manner that could motivate the reimagining of our field. That reimagining needs to account for our nation's dramatically changing demographics so that assessment generally, and standardized testing specifically, better fit the needs of a multicultural society.
本研究基于一个中学同等学力测验,考察了教育高危群体在写作过程上的性 别差异。研究涉及了来自美国23个州的合计三万多考生,每一考生均参与了 该语言测验的12副本考卷中的一个。研究借助键盘记录中抽取出的特征推断 背后的写作过程,并将之整合为7个过程指标。研究结果发现女性被试的作文 得分和语言测验总分均领先于男性,但领先程度很微弱。更重要的是,当控 制了语言测验总分、年龄和作文题目后,全部7个过程指标均显示出显著的性 别差异,其中,最突出的指标是流畅性和编辑性的不同方面。当前研究结果 在许多重要方面与先前一些对在校生和成人的研究结果相一致,也与在线和 纸笔写作任务的研究结果相吻合。关于对使用字符类语言(如汉语) 进行写作 的个体如何开展类似研究,文章结尾给出了一些建议。
This study examined differences in the composition processes used by educationally at-risk males and females who wrote essays as part of a high-school equivalency examination. Over 30,000 individuals were assessed, each taking one of 12 forms of the examination’s language arts writing subtest in 23 US states. Writing processes were inferred using features extracted from keystroke logs and aggregated into seven composite indicators. Results showed that females earned higher essay and total language arts writing composite scores than did males, but only by trivial amounts. More pertinent was that, after controlling for language arts writing composite score, age, and essay prompt, all seven process indicators showed nontrivial, statistically significant differences, the most notable being for indicators related to fluency and different aspects of editing. The study’s findings are consistent in important ways with those from other investigations of school-age students and adults, and with results from both online and paper-based writing tasks. Implications are offered for conducting similar research for individuals composing in character-based languages like Chinese.
This study investigates the effects of a scenario-based assessment design on students' writing processes. An experimental data set consisting of four design conditions was used in which the number of scenarios (one or two) and the placement of the essay task with respect to the lead-in tasks (first vs. last) were varied. Students' writing processes on the essay task were recorded using keystroke logs. Each keystroke action was classified into one of four writing states: planning, text production, local edit, or jump edit, and a semi-Markov model was fit to the data. Results showed that the single-scenario and essay-last design encouraged fewer but longer editing states compared to the alternative designs. Additionally, this task ordering appeared to have enabled more fluent and efficient text production when paired with a single scenario. These results seem explainable from cognitive writing theory, particularly with respect to working memory load. Limitations and future directions for research are also discussed.
We evaluate how higher- vs. lower-scoring middle-school students differ in their composition processes when writing persuasive essays from source materials. We examined differences on four individual process features-time taken before beginning to write, typing speed, total time spent, and number of words started. Next, we examined differences for four aggregated process measures: fluency, local editing, macro editing, and interstitial pausing (suspending text entry at locations associated with planning). Results showed that higher vs. lower scoring students were most consistently differentiated by total time, number of words started, and fluency. These differences persisted across two persuasive subgenres and two proficiency criteria, essay score and English language arts total-test score. The study's findings give a more complete picture of how the processes employed by more- and less-successful students differ, which contributes to cognitive writing theory and may have eventual implications for education policy and instructional practice.
Writing from source text is critical for developing college-and-career readiness because it is required in advanced academic environments and many vocations. Scenario-based assessment (SBA) represents one approach to measuring this ability. In such assessment, the scenario presents an issue that the student is to read and write about. Before writing, lead-in exercises are presented to encourage the examinee to engage with the source materials and to model the process used in a classroom writing project. This study experimentally manipulated a middle-school assessment design to understand if (1) the lead-in/essay structure increased scores erroneously with a concomitant decrease in test technical quality, and (2) the presence of a single unifying scenario affected scores or score meaning. In general, the SBA design did not appear to artificially increase total-test or essay scores. As importantly, it functioned as well as, sometimes better than, the alternative designs in terms of the measurement characteristics examined.
This study compared gender groups on the processes used in writing essays in an online assessment. Middle-school students from four grades responded to essays in two persuasive subgenres, argumentation and policy recommendation. Writing processes were inferred from four indicators extracted from students' keystroke logs. In comparison to males, on average females not only obtained higher essay scores but differed from males in their writing processes. Females entered text more fluently, engaged in more macro and local editing, and showed less need to pause at locations associated with planning (e.g., between bursts of text, at sentence boundaries). That these differences were detected after controlling for essay scores suggests that they cannot be attributed solely to disparities in group writing skill.
This paper presents a theoretical and empirical case for the value of scenario-based assessment (SBA) in the measurement of students' written argumentation skills. First, we frame the problem in terms of creating a reasonably efficient method of evaluating written argumentation skills, including for students at relatively low levels of competency. We next present a proposed solution in the form of an SBA and lay out the design for such an assessment. We then describe the results of prior research done within our group using this design. Fourth, we present the results of two new analyses of prior data that extend our previous results. These analyses concern whether the test items behave in ways consistent with the learning progressions underlying the design, how items measuring reading and writing component skills relate to essay performance, how measures of transcription fluency and proficiency in oral and academic language relate to writing skill, and whether the scenario-based design affects the fluency and vocabulary used in an essay. Results suggest that students can be differentiated by learning progression level, with variance in writing scores accounted for by a combination of performance on earlier tasks in the scenario and automated linguistic features measuring general literacy skills. The SBA structure, with preliminary tasks leading up to the final written performance, appears to result in more fluent (and also more efficient) writing behavior, compared to students' performances when they write an essay in isolation.
This chapter discusses the application of measurement principles to the theory and practice of classroom formative assessment. It describes assessment generally—and formative assessment particularly—from the perspective of evidentiary reasoning. The chapter explains how principles from educational measurement and the practice of formative assessment might be brought together. The development and application of evidentiary reasoning to assessment comes primarily from the work of Mislevy and colleagues on Evidence Centered Design. That model weights evidence according to its value, generating a score, a qualitative characterization, or both. That measurement model also provides an estimate of the uncertainty associated with that score or characterization. When formative assessment is embedded in educational software, formal measurement models can be employed because student responses can be automatically captured, the responses used in characterizing underlying proficiencies, and those characterizations acted upon to adjust instruction.
本文根据作者于2018年4月在纽约召开的(美国)全国教育测量学会(NCME)年会上的主席演讲稿修改而成.作者首先介绍了未来教育测量发展变化的11个可能特征、每个特征之所以重要的原因,以及应该如何看待这些变化.随后概述了未来教育测量领域不太可能发生变化的几个方面.最后对今后十年的教育发展进行了展望,并就这些发展对教育测量工作者可能产生的影响进行了讨论.
We used an unobtrusive approach, keystroke logging, to examine students’ cognitive states during essay writing. Based on data contained in the logs, we classified writing process data into three states: text production, long pause, and editing. We used semi-Markov processes to model the sequences of writing states and compared the state transition time and probability for demographic subgroups that were matched on writing proficiency. Results suggested that the subgroups employed different processes in essay writing.