We propose IRAISE (Impactful and Responsible AI Systems for Education), a one-day workshop at the Festival of Learning 2026 (in Seoul, South Korea). Building on two successful AAAI workshops (2024 and 2025) on responsible innovation in AI for education, IRAISE is organized around three pillars that address the most pressing challenges facing AI in education today. From static to continuous: the field must move beyond one-time evaluations toward continuous improvement cycles that keep pace with AI’s rapid evolution while respecting education’s slower, evidence-based rhythms. From tech-first to learning theory-driven: AI tools must be grounded in learning science, cognitive science, and psychometric theory—not merely technically sophisticated. From labs to classrooms: translating research into deployable products requires multi-stakeholder co-design with teachers, students, and policymakers. Through a keynote, invited talks, posters, roundtables, and a panel, IRAISE will convene researchers, practitioners, policymakers, and industry leaders to catalyze best practices at the intersection of responsible AI innovation and real-world educational impact.
Remote proctors make decisions about whether test takers have violated testing rules and, as a result, whether to certify test takers' scores. These decisions rely on both AI signals and human evaluation of test-taking behaviors. Given that fairness is a key component of test validity evidence, it is critical that proctors' decisions are unbiased with respect to proctor and test-taker background characteristics (e.g., gender, age, and nationality). In this study, we empirically evaluate whether proctor or test-taker background characteristics affect whether a test taker is flagged for rule violations. Results suggest that proctor and test-taker nationality may influence proctoring decisions, whereas gender and age do not. The direction of the influence generally reflects an "in-group, out-group" bias: proctors are less likely to identify rule violations among test takers with similar nationalities as proctors (in-group favoring) and more likely to identify rule violations among test takers of different nationalities (out-group disfavoring). Results also suggest that decisions based on AI signals may be less prone to in-group/out-group bias than decisions based on human evaluation only, although more research is needed to support this finding.
The core principles of validity, reliability, and fairness have been the foundation of ethical assessment practices, as discussed in classical validation theories (e.g., Chapelle et al., 2008; Kane, 1992, 2013) and the American Educational Research Association (AERA), the American Psychological Association (APA), and the National Council on Measurement in Education Standards (NCME) (AERA & APA & NCME, 2014). The Standards have provided best practices for AI use in high-stakes testing, particularly in the automated scoring of written and spoken responses. Responsible AI (RAI) is essential across industry domains, including educational assessments. With recent advancements in generative AI, new policies and guidance on applying RAI principles in assessment have emerged. Expanding on Chapelle et al.'s (2008) work, this paper introduces a unified assessment framework that integrates traditional validation theory with both assessment-specific and domain-agnostic RAI principles. This framework supports responsible AI use, aligns with ethical principles to uphold human values and oversight, and promotes broader social responsibility in AI-driven assessments.
Practice tests for high-stakes assessment are intended to build test familiarity, and reduce construct-irrelevant variance which can interfere with valid score interpretation. Generative AI-driven, automated item generation (AIG) scales the creation of large item banks and multiple practice tests, enabling repeated practice opportunities. We conducted a large-scale observational study (N = 25,969) using the Duolingo English Test (DET) – a digital, high-stakes, computer-adaptive English language proficiency test to examine how increased access to repeated test practice relates to official DETscores, test-taker affect (e.g., confidence), and score-sharing for university admissions. To our knowledge, this is the first large-scale study exploring the use of AIG-enabled practice tests in high-stakes language assessment. Results showed that taking 1-3 practice tests was associated with better performance (scores), positive affect (e.g., confidence) toward the official DET, and increased likelihood of sharing scores for university admissions for those who also expressed positive affect. Taking more than 3 practice tests was related to lower performance, potentially reflecting washback – i.e., using the practice test for purposes other than test familiarity, such as language learning or developing test-taking strategies. Findings can inform best practices regarding AI-supported test readiness. Study findings also raise new questions about test-taker preparation behaviors and relationships to test-taker performance, affect, and behaviorial outcomes.
The use of digital technologies in higher education is continually increasing, leading to changes in language use and presumably altering the language skills needed for academic studies. However, scores from high-stakes English language proficiency (ELP) tests used in postsecondary admissions only ensure the prerequisite level of traditional English skills (reading, writing, listening, and speaking). Such tests generally do not directly assess technology-mediated language skills (e.g., using online dictionaries and communicating via text message) that likely facilitate successful degree completion for international students. We present results from a needs analysis survey to re-evaluate the English-medium postsecondary linguistic landscape (i.e., update the target language use domain description), to inform ELP admissions tests. We specifically investigate international student (n = 379) and disciplinary instructor (n = 427) perceptions of the importance and frequency of technology-mediated language skills. Results show that student and instructor responses differ on certain technology-mediated skills, such as typing on a smartphone, underscoring the need to consider diverse perspectives in domain analysis research. Findings may inform how digital ELP admissions tests are developed, and how English for academic purposes curricula are designed, in order to better align test/classroom tasks with the academic language skills postsecondary students need.
Artificial intelligence (AI) creates opportunities for assessments, such as efficiencies for item generation and scoring of spoken and written responses. At the same time, it poses risks (such as bias in AI-generated item content). Responsible AI (RAI) practices aim to mitigate risks associated with AI. This chapter addresses the critical role of RAI practices in achieving test quality (appropriateness of test score inferences), and test equity (fairness to all test takers). To illustrate, the chapter presents a case study using the Duolingo English Test (DET), an AI-powered, high-stakes English language assessment. The chapter discusses the DET RAI standards, their development and their relationship to domain-agnostic RAI principles. Further, it provides examples of specific RAI practices, showing how these practices meaningfully address the ethical principles of validity and reliability, fairness, privacy and security, and transparency and accountability standards to ensure test equity and quality.
Validity, reliability, and fairness are core ethical principles embedded in classical argument-based assessment validation theory. These principles are also central to the Standards for Educational and Psychological Testing (2014) which recommended best practices for early applications of artificial intelligence (AI) in high-stakes assessments for automated scoring of written and spoken responses. Responsible AI (RAI) principles and practices set forth by the AI ethics community are critical to ensure the ethical use of AI across various industry domains. Advances in generative AI have led to new policies as well as guidance about the implementation of RAI principles for assessments using AI. Building on Chapelle's foundational validity argument work to address the application of assessment validation theory for technology-based assessment, we propose a unified assessment framework that considers classical test validation theory and assessment-specific and domain-agnostic RAI principles and practice. The framework addresses responsible AI use for assessment that supports validity arguments, alignment with AI ethics to maintain human values and oversight, and broader social responsibility associated with AI use.
The use of digital technologies in higher education is continually increasing, leading to changes in language use and presumably altering the language skills needed for academic studies. However, scores from high-stakes English language proficiency (ELP) tests used in postsecondary admissions only ensure the prerequisite level of traditional English skills (reading, writing, listening, and speaking). Such tests generally do not directly assess technology-mediated language skills (e.g., using online dictionaries and communicating via text message) that likely facilitate successful degree completion for international students. We present results from a needs analysis survey to re-evaluate the English-medium postsecondary linguistic landscape (i.e., update the target language use domain description), to inform ELP admissions tests. We specifically investigate international student (n = 379) and disciplinary instructor (n = 427) perceptions of the importance and frequency of technology-mediated language skills. Results show that student and instructor responses differ on certain technology-mediated skills, such as typing on a smartphone, underscoring the need to consider diverse perspectives in domain analysis research. Findings may inform how digital ELP admissions tests are developed, and how English for academic purposes curricula are designed, in order to better align test/classroom tasks with the academic language skills postsecondary students need.
Large language models (LLMs) have been a catalyst for the increased use of AI for automatic item generation on high-stakes assessments. Standard human review processes applied to human-generated content are also important for AI-generated content because AI-generated content can reflect human biases. However, human reviewers have implicit biases and gaps in cultural knowledge which may emerge where the test population is diverse. Quantitative analyses of item responses via differential item functioning (DIF) can help to identify these unknown biases. In this paper, we present DIF results based on item responses from a high-stakes English language assessment (Duolingo English Test - DET). We find that human- and AI-generated content, both of which were reviewed for fairness and bias by humans, show similar amounts of DIF overall but varying amounts by certain test-taker groups. This finding suggests that humans are unable to identify all biases beforehand, regardless of how item content is generated. To mitigate this problem, we recommend that assessment developers employ human reviewers which represent the diversity of the test-taking population. This practice may lead to more equitable use of AI in high-stakes educational assessment.
Essay scoring is a critical task used to evaluate second-language (L2) writing proficiency on high-stakes language assessments. While automated scoring approaches are mature and have been around for decades, human scoring is still considered the gold standard, despite its high costs and well-known issues such as human rater fatigue and bias. The recent introduction of large language models (LLMs) brings new opportunities for automated scoring. In this paper, we evaluate how well GPT-3.5 and GPT-4 can rate short essay responses written by L2 English learners on a high-stakes language assessment, computing inter-rater agreement with human ratings. Results show that when calibration examples are provided, GPT-4 can perform almost as well as modern Automatic Writing Evaluation (AWE) methods, but agreement with human ratings can vary depending on the test-taker's first language (L1).
Writing is a critical competency in 4-year college curricula. Yet, it is acknowledged that many students lack the writing skills required in college. College retention remains a national concern. However, to our knowledge, there is a research gap around the relationship between writing skills and retention in college. This chapter examines the potential for an innovative application of automated writing evaluation (AWE) to fill the gap in what we know about college writing skills and retention. Using AWE beyond its traditional function for automated scoring and feedback, the chapter applies AWE to assess multiple facets of students' writing, using writing samples from their coursework and a standardized writing test. It then tests whether these AWE features predict retention. A sociocognitive writing achievement framework defined the evaluated components of writing. AWE was applied to 997 coursework writing samples from 476 students from six universities. A survival analysis was conducted to examine the implications of student writing for retention. Study findings show relationships between some AWE features and student retention. Findings from this research have implications for real-time analytics which can serve as a contributing support mechanism for college retention.
The popularization of large language models (LLMs) such as OpenAI's GPT-3 and GPT-4 have led to numerous innovations in the field of AI in education. With respect to automated writing evaluation (AWE), LLMs have reduced challenges associated with assessing writing quality characteristics that are difficult to identify automatically, such as discourse coherence. In addition, LLMs can provide rationales for their evaluations (ratings) which increases score interpretability and transparency. This paper investigates one approach to producing ratings by training GPT-4 to assess discourse coherence in a manner consistent with expert human raters. The findings of the study suggest that GPT-4 has strong potential to produce discourse coherence ratings that are comparable to human ratings, accompanied by clear rationales. Furthermore, the GPT-4 ratings outperform traditional NLP coherence metrics with respect to agreement with human ratings. These results have implications for advancing AWE technology for learning and assessment.
• Background: This exploratory writing analytics study uses argumentative writing samples from two performance contexts—standardized writing assessments and university English course writing assignments—to compare (1) linguistic features in argumentative writing and (2) relationships between linguistic characteristics and academic performance outcomes. Writing data from this study come from 180 students enrolled at five four-year universities in the United States.
• Background: Researchers interested in quantitative measures of student “success” in writing cannot control completely for contextual factors which are local and site-based (i.e., in context of a specific instructor’s writing classroom at a specific institution). (In)ability to control for curriculum in studies of student writing achievement complicates interpretation of features measured in student writing. This article demonstrates how identifying and analyzing features of writing curriculum can provide dimensions of local context not captured in analysis of student-generated texts alone. Using a dataset of 48 curricular texts collected from 21 instructors teaching in five disciplines across six four-year public universities in the United States, this article: 1) presents a set of curriculum scoring rubrics developed through qualitative analysis, 2) describes a protocol for training raters to use the rubrics to score curricular texts to achieve rater agreement and generate quantitative data, and 3) explores how this framework
In the past few years, our lives have changed due to the COVID-19 pandemic; many of these changes resulted in pivoting our activities to a virtual environment, forcing many of us out of traditional face-to-face activities into digital environments. Digital-first learning and assessment systems (LAS) are delivered online, anytime, and anywhere at scale, contributing to greater access and more equitable educational opportunities. These systems focus on the learner or test-taker experience while adhering to the psychometric, pedagogical, and validity standards for high-stakes learning and assessment systems. Digital-first LAS leverage human-in-the-loop artificial intelligence to enable personalized experience, feedback, and adaptation; automated content generation; and automated scoring of text, speech, and video. Digital-first LAS are a product of an ecosystem of integrated theoretical learning and assessment frameworks that align theory and application of design and measurement practices with technology and data management, while being end-to-end digital. To illustrate, we present two examples—a digital-first learning tool with an embedded assessment, the Holistic Educational Resources and Assessment (HERA) Science, and a digital-first assessment, the Duolingo English Test.