
A frequently encountered security issue in writing tests is nonauthentic text submission: Test takers submit texts that are not their own but rather are copies of texts prepared by someone else. In this report, we propose AutoESD, a human-in-the-loop and automated system to detect nonauthentic texts for a large-scale writing tests, and report its performance on an operational data set. The AutoESD system utilizes multiple automated text similarity measures to identify suspect texts and provides an analytics-enhanced web application to help human experts review the identified texts. To evaluate the performance of AutoESD, we obtained its similarity measures on TOEFL iBT ® test writing responses collected from multiple remote administrations and examined their distributions. The results were highly encouraging in that the distributional characteristics of AutoESD similarity measures were effective in identifying suspect texts and the measures could be computed quickly without affecting the operational score turnaround timeline.
This paper presents a multidimensional model of variation in writing quality, register, and genre in student essays, trained and tested via confirmatory factor analysis of 1.37 million essay submissions to ETS' digital writing service, Criterion®. The model was also validated with several other corpora, which indicated that it provides a reasonable fit for essay data from 4th grade to college. It includes an analysis of the test‐retest reliability of each trait, longitudinal trends by trait, both within the school year and from 4th to 12th grades, and analysis of genre differences by trait, using prompts from the Criterion topic library aligned with the major modes of writing (exposition, argumentation, narrative, description, process, comparison and contrast, and cause and effect). It demonstrates that many of the traits are about as reliable as overall e‐rater® scores, that the trait model can be used to build models somewhat more closely aligned with human scores than standard e‐rater models, and that there are large, significant trait differences by genre, consistent with genre differences in trait patterns described in the larger literature. Some of the traits demonstrated clear trends between successive revisions. Students using Criterion appear to have consistently improved grammar, usage, and spelling after getting Criterion feedback and to have marginally improved essay organization. Many of the traits also demonstrated clear grade level trends. These features indicate that the trait model could be used to support more detailed scoring and reporting for writing assessments and learning tools.
Originating in adult education, the approach of task-based language teaching (TBLT) has been promoted in young language learner (YLL) education. However, its application often encounters challenges due to varying interpretations of what constitutes a “task.” Previous research has repeatedly highlighted gaps in teachers' understanding of tasks, often reducing them to mere exercises rather than opportunities for genuine communication. A potential issue could be that some of the criteria of a task as defined in the literature that focuses on adult second/foreign language (L2) learners do not necessarily apply or may need to be modified in YLL education. For example, tasks have traditionally been defined as having “authenticity,” but this may vary, as YLLs are often engaged in play and driven by imagination. Additionally, for children, school represents their “real world,” so their concept of an “authentic” task may differ from that of adult L2 learners, who may be attending classes to improve workplace skills. In this study, we aimed to explore the concept of task in the context of teaching an additional language to YLLs in primary education. Utilizing a Delphi method, 16 well-known experts who work at the intersection of applied linguistics, TBLT, and YLLs participated in three rounds of data collection via email. After providing written definitions of a task and its characteristics in the YLL classroom in Round 1, the experts rated each other's definitions on a 4-point Likert scale and provided comments on the definitions in two subsequent rounds. Additionally, we conducted follow-up interviews with a subsample of the participants ( n = 6) relative to a particular task characteristic: “authenticity.” Using both quantitative and qualitative analyses, we identified key aspects from the data, including task characteristics, learner considerations, and implementation details. Findings showed a distinction between “activity” and “task,” with the latter being understood as featuring certain characteristics. Accordingly, a task in the YLL classroom has a goal orientation, an orientation to meaning rather than linguistic form, a need for YLLs to use their L2 repertoire, a type of information gap, and a real-life connection. While largely congruent with the concept of task in the L2 adult literature, the experts particularly highlighted a learner-oriented approach to tasks that stresses cognitive, social-emotional, and affective development of YLLs. In particular, experts highlighted the significance of imagination as part of children's authentic world. Thus an “authentic” task for adults may reference a “real-world” domain, whereas an authentic task for YLLs may reference an imaginary one. We discuss the findings and emphasize that the concept of task in YLL education should be broadened to include aspects of imaginary worlds and make-believe.
The multistage testing (MST) design has been gaining attention and popularity in educational assessments. For testing programs that have small test‐taker samples, it is challenging to calibrate new items to replenish the item pool. In the current research, we used the item pools from an operational MST program to illustrate how research studies can be built upon literature and program‐specific data to help to fill the gaps between research and practice and to make sound psychometric decisions to address the small‐sample issues. The studies included choice of item calibration methods, data collection designs to increase sample sizes, and item response theory models in producing the score conversion tables. Our results showed that, with small samples, the fixed parameter calibration (FIPC) method performed consistently the best for calibrating new items, compared to the traditional separate‐calibration with scaling method and a new approach of a calibration method based on the minimum discriminant information adjustment. In addition, the concurrent FIPC calibration with data from multiple administrations also improved parameter estimation of new items. However, because of the program‐specific settings, a simpler model may not improve current practice when the sample size was small and when the initial item pools were well‐calibrated using a two‐parameter logistic model with a large field trial data.
Existing research reveals a robust relationship between self-reported print exposure and long-term literacy development, yet few studies have demonstrated how reading skills change as children read a book in the short term. In this study, 50 children (mean age 9.7 years, SD = .8) took turns with a prerecorded narrator reading aloud a popular children's novel, producing 6,092 oral reading responses over 1,093 book passages. Each oral reading response was evaluated by a speech engine that calculated words-correct-per-minute (WCPM). Mixed effect models revealed that text level differences, between-individual differences, and within-individual variations explained 13%, 56% and 32% of variance in WCPM, respectively. On average, children started reading the book at about 93 WCPM, and they improved by 2.26 WCPM for every 10,000 words of book reading. Random effects showed that the standard deviation of the growth rate was 1.85 WCPM, suggesting substantial individual difference in growth rate. Implications for reading instruction and assessment were discussed.
Assessment refers to a broad array of approaches for measuring or evaluating a person's (or group of persons') skills, behaviors, dispositions, or other attributes. Assessments range from standardized tests used in admissions, employee selection, licensure examinations, and domestic and international large-scale assessments of cognitive and behavioral skills to formative K–12 classroom curricular assessments. The various types of assessments are used for a wide variety of purposes, but they also have many common elements, such as standards for their reliability, validity, and fairness—even classroom assessments have standards. We believe the future of assessment will involve a shift in emphasis on what skills will be measured, innovations in how we go about measuring them, the use of advanced technologies for test operations, and an expansion in the value and kinds of information that test takers will receive from taking the assessment. In this paper, we argue and provide evidence for our belief that the future of assessment contains challenges but is promising. The challenges include risks associated with security and exposure of personal data, test score bias, and inappropriate test uses, all of which may be exacerbated by the growing infiltration of artificial intelligence (AI) into our lives. The promise is increasing opportunities for testing to help individuals achieve their education and career goals and contribute to well-being and overall quality of life. To help achieve this promise we focus on the evidence-based science of measurement in education and workplace learning, a theme throughout this paper.
Collaborative learning environments that support students' problem solving have been shown to promote better decision-making, greater academic achievement, and more reasonable argumentation about controversial issues. In this research, we developed a technology-based critical discussion platform to support middle school students' argumentation, with a focus on evidence-based reasoning and perspective taking. A feasibility study was conducted to examine the patterns of group interaction and individual students' contributions to the critical discussion and their perceptions of the critical discussion activity. We found that more students used text-based communications than audio, but students who used audio collaborated with each other more frequently. In addition, student engagement in argumentative discourse varied greatly across groups as well as individuals. At the end of the discussion, most groups provided a solution that integrated both sides of the controversial issue. Survey and interview results suggest an overall positive experience with this technology-supported critical discussion activity. Using the insights from our research, we develop a conceptual dialogue analysis framework that identifies relevant skills under the argumentation and collaboration dimensions. In this report, we discuss our design considerations, feasibility study results, and implications of engaging students in computer-supported collaborative argumentation.
Studies of test score comparability have been conducted at different stages in the history of testing to ensure that test results carry the same meaning regardless of test conditions. The expansion of at-home testing via remote proctoring sparked another round of interest. This study uses data from three licensure tests to assess potential mode effects associated with the dual option of on-site testing at test centers and at-home testing via remote proctoring. We generated propensity score weights to balance the two self-selected groups in order to detect the mode effect on the test outcomes. We also assessed the potential impact of omitted variables on the estimated mode effect. Results of the study indicate that the demographic compositions of the test takers are similar before and after the introduction of the RP option. Examinees under the two testing modes differ slightly on certain background variables. Once the group differences are adjusted by propensity score weighting, the estimated mode effects are small and nonsystematic across test titles overall. We note some variations across subgroups based on gender and race.
In a study of 370 postsecondary students in electronics, engineering, and other science classes, we investigated collaborative problem-solving (CPS) skills that best predict performance at individual levels in an online electronics environment. The results showed that while monitoring was a consistent predictor across levels, other skills such as executing, sharing information, planning, and maintaining communication each predicted individual performance at one or more levels of the task. The availability of background data on students' content classes and associated content knowledge to analyze the model results can help identify possible cues for instructors across domains to help students improve specific CPS skills to achieve high performance in activities conducted in collaborative learning environments.
Over the last 20 years, many methods have been proposed to use process data (e.g., response time) to detect changes in engagement during the test-taking process. However, many of these methods were developed and evaluated in highly similar testing contexts: 30 or more single-select multiple-choice items presented in a linear, fixed sequence in which an item must be answered before progressing to the next item. However, this testing context becomes less and less representative of testing contexts in general as the affordances of technology are leveraged to provide more diverse and innovative testing experiences. The 2019 National Assessment of Educational Progress (NAEP) mathematics administration for grades 8 and 12 testing context represents an example use case that differed significantly from assessments that were typically used in previous research on test-taking engagement (e.g., number of items, item format, navigation). Thus, we leveraged this use case to re-evaluate the utility of an existing engagement detection method: normative threshold method. We decomposed the normative threshold method to evaluate its alignment with this use case and then evaluated 25 variations of this threshold-setting method with previously established evaluation criteria. Our findings revealed that this critical analysis of the threshold-setting method's alignment with the NAEP testing context could be used to identify the most appropriate variation of this method for this use case. We discuss the broader implications for engagement detection as testing contexts continue to evolve.
This paper reports partial results from a larger study of how three different groups of stakeholders—university admissions officers, faculty in graduate programs involved in admissions decisions, and Intensive English Program (IEP) faculty—interpret and use TOEFL iBT® scores in making admissions decisions or preparing students to meet minimum test score requirements. Our overall goal was to gain a better understanding of the perceived role of English language proficiency in admissions decisions and the internal and external factors that inform decisions about acceptable ways to demonstrate proficiency and minimal standards. To that end, we designed surveys for each stakeholder group that contained questions for all groups and questions specific to each group. This report focuses on the questions that were common to all three groups across two areas: (1) understandings of and participation in institutional policy making around English language proficiency tests and (2) knowledge of and attitudes toward the TOEFL iBT test itself. Our results suggested that, as predicted, university admissions staff were the most aware of and involved in policy making but frequently consulted with ESL experts such as IEP faculty when setting policies. This stakeholder group was also the most knowledgeable about the TOEFL iBT test. Faculty in graduate programs varied in their understanding of and involvement in policy making and reported the least familiarity with the test. However, they reported that more information about many aspects of the test would help them make better admissions decisions. The results of the study add to the growing literature on language assessment literacy among various stakeholder groups, especially in terms of identifying aspects of assessment literacy that are important to different groups of stakeholders.
The goal of this paper is to find better ways to estimate the internal consistency reliability of scores on tests with a specific type of design that are often encountered in practice: tests with constructed‐response items clustered into sections that are not parallel or tau‐equivalent, and one of the sections has only one item. To estimate the reliability of scores on this kind of test, we propose a two‐step approach (denoted as CA_STR) that first estimates the reliability of scores on the section with a single item using the correction for attenuation method and then estimates the reliability of scores on the whole test using the stratified coefficient alpha. We compared the CA_STR method with three other reliability estimation approaches under various conditions using both real and simulated data. We found that overall, the CA_STR method performed the best and it was easy to implement.
The principle of fairness in testing traditionally involves an assertion about the absence of bias, or that measurement should be impartial (i.e., not provide an unfair advantage or disadvantage), across groups of test takers. In more general‐purposes language testing, a test taker's background knowledge is not typically considered relevant to the measurement of language proficiency; consequently, if there are systematic differences in background knowledge between groups of test takers this background knowledge should not provide an unfair advantage or disadvantage. As a general‐purposes assessment of English for everyday life and the international workplace, the TOEIC® Listening and Reading test is designed to assess the listening and reading comprehension skills of second language (L2) users of English. In this study, we investigated whether a group of test takers with more workplace experience (full‐time employees) have an unfair advantage over test takers with less workplace experience (full‐time students). We conducted DIF analysis using nine forms of the test (1,800 items) and flagged 18 items (1.0%) for statistical differential functioning. An expert panel reviewed the items and concluded that none of the items could be clearly identified as biased in favor of employed (or student) test takers. Follow‐up analyses using score equity assessment found that test scores do not unfairly advantage fulltime employed (versus student) test takers. Finally, we performed a content review using two expert panels that led to examples of how workplace‐oriented content is incorporated into test items without disadvantaging full‐time students (versus full‐time employees). The results of these analyses provide support for claims about the impartiality (or fairness) of TOEIC Listening and Reading test scores for postsecondary test takers and add to current research on the role of background knowledge and fairness for more general‐purposes language assessments.
The TOEFL Junior ® tests are designed to evaluate young language students' English reading, listening, speaking, and writing skills in an English-medium secondary instructional context. This paper articulates a validity argument constructed to support the use and interpretation of the TOEFL Junior test scores for the purpose of placement, progress monitoring, and evaluation of a test taker's English skills. The validity argument is built within an argument-based approach to validation and consists of six validity inferences that provide a coherent narrative about the measurement quality and intended uses of the TOEFL Junior test scores. Each validity inference is underpinned by specific assumptions and corresponding evidential support. The claims and supporting evidence presented in the validity argument demonstrate how the TOEFL Junior research program takes a rigorous approach to supporting the uses of the tests. The compilation of validity evidence serves as a resource for score users and stakeholders, guiding them to make informed decisions regarding the use and interpretation of TOEFL Junior test scores within their educational contexts.
At a time when institutions of higher education are exploring alternatives to traditional admissionstesting, institutions are also seeking to better support students and prepare them for academicsuccess. Under such an engaged model, one may seek to measure not just the accumulatedknowledge and skills that students would bring to a new academic program, but also their ability togrow and learn through the academic program. To help prepare students for law school before theymatriculate, the JD-Next is a fully-online, non-credit, 7-10 week course to train potential juris doctor(JD) students in case reading and analysis skills. This study builds upon the work presented forprevious JD-Next cohorts by introducing new scoring and reliability estimation methodologies basedon a recent redesign of the assessment for the 2021 cohort, as well as presenting updated validity andfairness findings, using first-year grades, rather than merely first-semester grades as in prior cohorts.Results support the claim that the JD-Next exam is reliable and valid for predicting law schoolsuccess, providing a statistically significant increase in predictive power over baseline modelsincluding entrance exam scores and grade-point average. In terms of fairness across racial and ethnicgroups, smaller score disparities are found with JD-Next than with traditional admissionsassessments, and the assessment is shown to be equally predictive for students from underrepresentedminority groups and first-generation students. These findings, in conjunction with those fromprevious research, support the use of the JD-Next exam for both preparation and admissions offuture law school students.
In a targeted double‐scoring procedure for performance assessments that are used for licensure and certification purposes, a subset of responses receives an independent second rating if the first rating falls into a preidentified critical score range (CSR) where an additional rating would lead to considerably more reliable pass‐fail decisions. This study evaluates the CSRs using two approaches—one based on imputation of missing scores and the other based on statistical decision theory—using data from the Performance Assessment for School Leaders (PASL). Results from the evaluation indicate that the currently used CSRs are effective.
Culturally responsive personalized learning (CRPL) emphasizes the importance of aligning personalized learning approaches with previous research on culturally responsive practices to consider social, cultural, and linguistic contexts for learning. In the present discussion, we briefly summarize two bodies of literature considered in defining and developing a framework for CRPL: technology‐enabled personalized learning and culturally relevant, responsive, and sustaining pedagogy. We then provide a definition and framework consisting of six key principles of CRPL, along with a brief discussion of theories and empirical evidence to support these principles. These six principles include agency, dynamic adaptation, connection to lived experiences, consideration of social movements, opportunities for collaboration, and shared power. These principles fall into three domains: fostering flexible student‐centered learning experiences, leveraging relevant content and practices, and supporting meaningful interactions within a community. Finally, we conclude with some implications of this framework for researchers, policymakers, and practitioners working to ensure that all students receive high‐quality learning opportunities that are both personalized and culturally responsive.
Recent criticisms of large‐scale summative assessments have claimed that the assessments are biased against historically excluded groups because of the assessments' lack of cultural representation. Accompanying these criticisms is a call for more culturally responsive assessments—assessments that take into account the background characteristics of the students; their beliefs, values, and ethics; their lived experiences; and everything that affects how they learn and behave and communicate. In this paper, we present provisional principles, based on a review of research, that we deem necessary for fostering cultural responsiveness in assessment. We believe the application of these principles can address the criticisms of current assessments.
This report presents results from a survey of 64 elementary mathematics and reading language arts teacher educators providing feedback on a new type of short performance task. The performance tasks each present a brief teaching scenario and then require a short performance as if teaching actual students. Teacher educators participating in the study first reviewed six performance tasks, followed by a more in‐depth review of two of the tasks. After reviewing the tasks, teacher educators completed an online survey providing input on the value of the tasks and on potential uses to support teacher preparation. The survey responses were positive with the majority of teacher educators supporting a variety of different uses of the performance tasks to support teacher preparation. The report concludes by proposing a larger theory for how the performance tasks can be used as both formative assessment tools to support teacher learning and summative assessments to guide decisions about candidates' readiness for the classroom.
Though a substantial amount of research exists on imputing missing scores in educational assessments, there is little research on cases where responses or scores to an item are missing for all test takers. In this paper, we tackled the problem of imputing missing scores for tests for which the responses to an item are missing for all test takers. We considered three missing-data imputation methods—the median method, the item response theory (IRT) method, and the two-way method—for imputing scores. We compared the performance of these three imputation methods with respect to their accuracy in estimating scaled scores and test reliability for the aforementioned problem. Real data were used in the comparison. All three methods performed well in imputing scaled scores with negligible imputation error: The IRT method and the median method provided slightly more accurate scaled scores. The two-way method provided the most accurate reliability estimates. Recommendations for practice are provided.