Undergraduate psychology research methods courses offer a foundational opportunity for students to learn scientific writing, yet many students find these courses challenging and are not uniformly successful in acquiring writing skills. Generative artificial intelligence (GenAI) holds promise as a pedagogical support for emerging writing skills. Thus, the current study examined whether providing undergraduate students with in-class training on how to use GenAI for scientific writing improves their grades on scientific writing assignments and self-efficacy when writing. One section of a psychology research methods course was provided with weekly in-class GenAI training (GenAI condition), and the other section was not given any training (Control condition). An examination of students’ grades on two research papers (Papers 1 and 2) revealed no significant differences between conditions, but students in the Control condition significantly improved their grades across time, whereas students in the GenAI condition remained stable. Students’ writing self-efficacy did not vary by condition or time. Students’ perceptions of GenAI improved from pre- to post-course but did not differ by condition. Overall, in-class GenAI training did not improve writing performance beyond standard instruction, but we highlight future directions for potentially more effective GenAI training.
Recent assessments of students’ reading comprehension in the United States have found that many 3rd, 4th, and 5th grade readers are not developing the skills necessary for academic success. iSTART-Early is an intelligent tutoring system that provides instruction on reading comprehension strategies and offers practice with skill application by answering review questions, designed to improve literacy outcomes. Across two studies, we measured students’ engagement, enjoyment, subjective difficulty, and objective difficulty of iSTART-Early lesson videos and review questions, by iteratively adapting and testing aspects of the system design based on findings. Study 1 revealed that the lessons were grade-appropriate for 3rd through 5th grade students regarding difficulty and enjoyment levels, with lessons being too simple for older children and too challenging for younger children. However, engagement was low across some lessons. In Study 2, modifications to iSTART-Early’s design were made based on Study 1’s findings, which were found to increase engagement while also maintaining appropriate enjoyment and difficulty levels for the target audience. Overall, these findings suggest that the design of lesson videos and review questions in iSTART-Early is grade-appropriate and has strong potential to enhance students’ reading comprehension skills.
As generative AI expands the technical frontiers of prediction, measurement, and design, a growing tension has emerged between algorithmic fluency and institutional trust. This conceptual article offers a narrative synthesis of recent work in learning analytics, educational data science, human–AI interaction, and AI governance to propose stewardship as a necessary fourth paradigm of educational data science. Stewardship represents the professional, epistemic, and institutional work of governing judgment in an environment where analytic systems are increasingly generative and persuasive. Rather than treating stewardship as a general ethics checklist, the article positions it as the governance of epistemic and pedagogical authority: who determines what counts as evidence, interpretation, and educational action when AI systems help produce those judgments. The synthesis suggests that while GenAI can support bounded analytic tasks, evidence for systemic educational transformation remains limited and uneven. The field’s primary challenge is therefore not technical performance alone, but the governance of interpretation, validation, delegation, and action. By centering provenance, uncertainty, accountable oversight, learner agency, and institutional learning, stewardship provides an actionable framework for anchoring analytic innovation in responsible educational improvement.
Knowledge Tracing (KT) is widely used for personalization in Intelligent Tutoring Systems (ITSs). However, it is typically implemented as a hidden component that autonomously makes learning decisions, leaving students with little control over their own learning. Generative AI (GenAI) creates an opportunity to rethink the relationship between KT and student agency through natural language interaction, enabling learners to contest KT-driven recommendations and take an active role in learning decisions. Yet, few have examined how a GenAI-based KT could be configured and evaluated to support ITSs’ adaptivity and student agency. In this paper, we first present a multi-agent ITS for AI-assisted French learning (AIFL), where a KT agent drives adaptivity, and learners can contest topic recommendations, ask follow-up questions, and track progress via a learning analytics dashboard (LAD). We then compared the KT agent’s performance against Bayesian Knowledge Tracing (BKT) with simulated students under experimental conditions. Our results showed that with appropriate multi-agent design and prompt engineering, a GenAI-based KT could achieve adaptivity comparable to an established framework, and maintain adaptivity when learners exercise agency in topic selection.
Generating multiple text sequences and refining them through feedback is essential for improving the quality of outputs in many NLP tasks. While Large Language Models can leverage iterative feedback during inference, smaller models often lack this capability due to limited capacity and the absence of suitable training paradigms. In this paper, we propose a novel Feedback-Aware Inference approach that enables iterative sequence generation with integration of feedback signals. Our method allows models to generate multiple sequences, incorporate feedback from previous iterations, and refine outputs accordingly. This approach dynamically adjusts to different quality metrics, making it adaptable to various contexts and objectives. We evaluate our approach on two distinct tasks: Answer Selection for Question Generation and Keyword Generation, arguing for its generalizability and effectiveness. Results show that our method outperforms strong baselines, maintaining high performance across iterations and achieving superior results even with smaller, open-source models.
Extensive research has demonstrated that testing enhances long-term memory by strengthening retrieval processes for the tested memory traces. Such memory models explain why learners benefit from repeated testing but do not account for how learners construct understanding across multiple ideas. In contrast, theories of discourse comprehension emphasize the construction of inferences, whereby learners actively connect ideas into a coherent mental representation. These two perspectives—memory-based retrieval versus inference-based integration—make distinct predictions about what learners retain and understand. The present study tested these predictions by comparing self-testing, self-explanation, and reading strategies on both fact memory and inference performance. In Experiment 1 (n = 254), despite superior retention of individual facts, self-testing produced lower inference performance than self-explanation. Experiment 2 (n = 68), a preregistered replication with a minimum 24-hour delay, confirmed and extended these findings: repeated retrieval improved recall of facts but did not facilitate the assembly of coherent knowledge structures. A mixed-methods analysis revealed that generating inferences during learning specifically predicted inference performance, whereas noting facts predicted fact recall, providing converging evidence that these strategies produce qualitatively different mental representations. Furthermore, participants generated fewer inferences when self-explaining science texts compared with narratives, highlighting domain differences in the ease of knowledge integration. Together, these results demonstrate that while self-testing strengthens memory for discrete information, it fails to foster the integrative processing necessary for comprehension. Effective learning requires more than strengthening retrieval processes—it depends on the active generation of inferences that bind ideas into coherent mental models.
Educational assessment systems have primarily relied on episodic forms of assessment, including examinations, assignments, grades, and credentials. These approaches provide efficient and scalable summaries of achievement and yet capture only part of the developmental process through which learners build competence. Moreover, learning increasingly unfolds across digital platforms, workplaces, collaborative networks, and AI-mediated environments, generating rich evidence of learner development that remains fragmented across systems and contexts. Advances in artificial intelligence, learning analytics, multimodal analytics, learner modeling, and semantic interoperability make it increasingly feasible to connect, integrate, and interpret this evidence across contexts and over time. This paper introduces the AI-Mediated Continuous Assessment Infrastructure (AIM-CAI), a sociotechnical framework supporting longitudinal, probabilistic interpretation of distributed evidence of learning. Within AIM-CAI, continuous assessment refers to the ongoing accumulation and dynamic interpretation of evidence generated through learning activities. The framework integrates distributed evidence systems, evidence serialization mechanisms, AI-mediated semantic translation, probabilistic learner models, dynamic competency profiles, and federated governance architectures to support context-sensitive interpretations of learner development while maintaining human judgment, privacy, accountability, and learner agency. The authors examine implications for assessment, credentialing, lifelong learning, institutional roles, interoperability, and governance and outline a research agenda addressing key psychometric, ethical, and governance challenges, including validity, fairness, surveillance, semantic instability, and ownership of learning evidence.
Rendering understanding as a general causal account is problematic. This leads us toward accusations of partial understanding, or illusory understanding, characterized by understanding that lacks depth. The discourse comprehension literature offers an alternative: rendering understanding as the construction of a situation model, a model that integrates knowledge, purpose, and context. On this account, the question is never whether understanding is complete, but whether it is adequate for what it aims to achieve.
Effective source integration is a complex but essential skill in academic writing; yet it remains difficult to teach and evaluate. Assessing source integration has important implications for formative feedback and instructional practice, but existing approaches face limitations, particularly in automated systems. The purpose of this study was to develop and validate an automated measure of source integration by extracting linguistic features from student essays and modeling them as a latent construct with confirmatory factor analysis. Predictive validity with human scores and generalizability across datasets using measurement invariance was tested and examined. The source integration construct included linguistic features related to citation, quotation, plagiarism, and semantic overlap with the source text. Results indicated strong alignment with human ratings (beta = .81, R2 = .65) and evidence of structural consistency across a new dataset with novel prompts and sources. Predictive utility analyses showed that the latent construct improved machine learning models and enhanced agreement with human ratings when paired with BERT embeddings. GPT-5.2 produced interpretable justifications but lower scoring reliability. These findings suggest that a source integration construct grounded in linguistic features can complement modern AI methods, providing a foundation for formative feedback on source-based writing that addresses issues of fairness and interpretability.
This work explores user needs for educational games and gamification that incorporates Generative Artificial Intelligence (GenAI). As GenAI is increasingly incorporated in educational settings, we must consider both the wide-spanning literature on gamification and games that have been shown to benefit learning, and characterize the needs and desires of relevant stakeholders in developing educational games that incorporate GenAI generally, and specifically for higher education. A mixed-methods questionnaire inquired 345 undergraduate students about their perceptions, use patterns, needs, and desires related to GenAI, educational and non-educational games, and text-based games. GenAI tools are widely used for educational purposes already, but mostly as a supplementary source. Despite the wide use, participants expressed being concerned with accuracy, transparency, and quality. Participants also expressed a desire for an educational game/tool to have scaffolded interactions and to help with learning material in math, science, and language arts. Taken together the findings provide a road map and specific recommendations for developing an educational game incorporating GenAI. The roadmap includes instructional design (i.e., the gamified tools’ content and type(s) of instruction and interaction) through information regarding preferred platforms, game genres, gamified properties (e.g., characters, challenges), and lastly, clear information about concerns students have related to trust and equity that will need to be addressed in an educational game incorporating GenAI.
Recent advances in large language models (LLMs) have made automated multiple-choice question (MCQ) generation increasingly feasible; however, reliably producing items that satisfy controlled cognitive demands remains a challenge. To address this gap, we introduce ReQUESTA, a hybrid, multi-agent framework for generating cognitively diverse MCQs that systematically target text-based, inferential, and main idea comprehension. ReQUESTA decomposes MCQ authoring into specialized subtasks and coordinates LLM-powered agents with rule-based components to support planning, controlled generation, iterative evaluation, and post-processing. We evaluated the framework in a large-scale reading comprehension study using academic expository texts, comparing ReQUESTA-generated MCQs with those produced by a single-pass GPT-5 zero-shot baseline. Psychometric analyses of learner responses assessed item difficulty and discrimination, while expert raters evaluated question quality across multiple dimensions, including topic relevance and distractor quality. Results showed that ReQUESTA-generated items were consistently more challenging, more discriminative, and more strongly aligned with overall reading comprehension performance. Expert evaluations further indicated stronger alignment with central concepts and superior distractor linguistic consistency and semantic plausibility, particularly for inferential questions. These findings demonstrate that hybrid, agentic orchestration can systematically improve the reliability and controllability of LLM-based generation, highlighting workflow design as a key lever for structured artifact generation beyond single-pass prompting.
Learning engineering applies data and learning science principles to better understand outcomes and support improvement research. One important approach is A/B testing-common in large software companies and also represented academically at conferences like the Annual Conference on Digital Experimentation (CODE), and the International Consortium for Innovation and Collaboration in Learning Engineering (IEEE ICICLE). Several systems supporting A/B testing in educational applications have arisen recently, including UpGrade, E-TRIALS, and Terracotta. A/B testing can help improve educational platforms, yet there are challenging issues unique to conducting such work in these contexts. In response, a number of digital learning platforms have opened their systems to learning-improvement research by instructors and/or third-party researchers, with specific supports necessary for education-specific research designs. This workshop will explore how A/B testing is conducted in educational contexts, how digital learning platforms are accelerating education research, and how empirical approaches can be used to drive powerful gains in student learning. It will also discuss opportunities for funding to conduct platform-enabled learning engineering.
The purpose of this feasibility study was to examine the potential impact of reading digital interactive e-books on essential skills that support reading comprehension with third-fifth grade students. Students read two e-Books that taught word learning and comprehension monitoring strategies in the service of learning difficult vocabulary and targeted science concepts about hurricanes. We investigated whether specific comprehension strategies including word learning and strategies that supported general reading comprehension, summarization, and question generation, show promise of effectiveness in building vocabulary knowledge and comprehension skills in the e-Books. Students were assigned to read one of three versions of each of the e-Books, each version implemented one strategy. The books employed a choose-your-adventure format with embedded comprehension questions that provided students with immediate feedback on their responses. Paired samples t-tests were run to examine pre-to-post differences in learning the targeted vocabulary and science concepts taught in both e-Books. For both e-Books, students demonstrated significant gains in word learning and on the targeted hurricane concepts. Additionally, Hierarchical Linear Modeling (HLM) revealed that no one strategy was more associated with larger gains than the other. Performance on the embedded questions in the books was also associated with greater posttest outcomes for both e-Books. This work discusses important considerations for implementation and future development of e-books that can enhance student engagement and improve reading comprehension.
The purpose of this study was to engage high school science teachers as co-design partners in refining and extending instructional frameworks to support multiple-document reading and writing in science classrooms. Using a participatory mixed-methods design, the project adapted the InSPECT framework for secondary science, developed professional development (PD) materials to introduce the framework, and explored the role of generative artificial intelligence (AI) in lesson planning. Five virtual focus group sessions guided the co-design of PD activities, followed by a pilot implementation in one biology classroom. Data included focus group and interview transcripts, surveys, and student work artifacts. Analyses examined teachers’ perceptions of PD features, framework usability, and student engagement. Teachers valued PD that was practical, relevant, and feasible within classroom constraints and described the frameworks as clear, stepwise structures that supported lesson design and literacy integration. Student work showed that paraphrasing was an accessible entry point, while bridging, elaboration, and source evaluation required additional modeling. Teachers viewed generative AI as a promising planning aid but expressed concerns about accuracy and ethics. Findings informed revisions emphasizing discipline-specific exemplars, scaffolds for higher-order strategies, and AI-literacy modules, illustrating how participatory design can yield feasible, teacher-centered PD.