Games are recognized as valuable tools for learning and interest development. However, the association between behavioral player archetypes and these important outcomes is still underexplored. This study explores the relationship between learning, interest development, and player archetypes (Roamers, Explorers, and Scientists) identified within the WHIMC project, a game-based learning environment where students engage with scientifically accurate hypothetical astronomy scenarios in Minecraft. Grounded in human-computer interaction and player typologies frameworks, we analyze data from 57 participants across four summer camps using Ordered Network Analysis (ONA) and k-means clustering to identify player archetypes emerging from student actions. We then examine how these archetypes relate to learning outcomes and motivational factors. Statistical tests reveal significant differences in in-game actions across archetypes and correlations between player behaviors and learning outcomes. These findings contribute to the design of serious educational games by increasing understanding of how to optimize experiences and enhance science engagement for learners with differing playing styles.
The emergence of generative AI has accelerated the development of conversational tutoring systems that interact with students through natural language dialogue. Unlike prior intelligent tutoring systems (ITS), which largely function as adaptive and interactive problem sets with feedback and hints, conversational tutors hold the potential to simulate high-quality human tutoring by engaging with students' thoughts, questions, and misconceptions in real time. While some previous ITS, such as AutoTutor, could respond conversationally, they were expensive to author and lacked a full range of conversational ability. Generative AI has changed the capacity of ITS to engage conversationally. However, realizing the full potential of conversational tutors requires careful consideration of what research on human tutoring and ITS has already established, while also unpacking what new research will be needed. This paper synthesizes tenets of successful human tutoring, lessons learned from legacy ITS, and emerging work on conversational AI tutors. We use a keep, change, center, study framework for guiding the design of conversational tutoring. We argue that systems should keep proven methods from prior ITS, such as knowledge tracing and affect detection; change how tutoring is delivered by leveraging generative AI for dynamic content generation and dialogic scaffolding; and center opportunities for meaning-making, student agency, and granular diagnosis of reasoning. Finally, we identify areas requiring further study, including efficacy testing, student experience, and integration with human instruction. By synthesizing insights from human tutoring, legacy ITS, and emerging generative AI technologies, this paper outlines a research agenda for developing conversational tutors that are scalable, pedagogically effective, and responsive to the social and motivational dimensions of learning.
ABSTRACT Background Video data provides rich opportunities to examine student behaviour in game‐based learning environments, capturing not only observable actions but also subtle indicators of cognitive engagement. However, traditional video analysis is labor‐intensive, and current applications of AI to this task are limited by models trained on general‐purpose datasets, which often fail to capture the pedagogical meaning of student actions in authentic educational contexts. Objectives This study explores how multimodal large language models (MLLMs), specifically LLaVA‐Video‐7B‐Qwen2, can support qualitative video analysis of student exploration behaviours in a Minecraft‐based STEM learning environment. Methods We conducted an exploratory case study using screen recordings from a Minecraft‐based STEM learning environment. We tested multiple prompt strategies for guiding MLLM‐generated video descriptions and found that role‐assignment prompting performed most effectively. We evaluated model outputs using a mixed‐method framework that included quantitative scoring, GPT‐based judgement, and researcher validation. Results and Conclusions Our findings show that MLLMs can reliably identify surface‐level behaviours, such as navigation patterns and object interactions, but struggle to infer the intent or goals behind student actions, leading to a significant rate of over‐interpretation (26.5% when explaining student strategies). The model's outputs are sensitive to prompt phrasing, underscoring the importance of prompt engineering. While current MLLMs show promise for streamlining parts of the video analysis workflow, their use in educational contexts requires structured oversight and careful interpretation to ensure reliability and relevance.
Large language models (LLMs) are rapidly transforming knowledge work by improving the quality and efficiency of tasks such as writing, coding, and data analysis. However, their growing use in education has exposed a learning-performance paradox: while they can enhance short-term task performance, they may also undermine genuine learning, including cognitive growth, knowledge transfer, and metacognitive development. This paper addresses the question of how artificial intelligence should be designed and used to support learning rather than merely improve immediate outputs. We introduce the concept of AI learning companions, defined as adaptive, pedagogically informed, LLM-powered agents designed for integration into learning environments. We propose a framework for their design built on three interrelated foundations: a pedagogical foundation focused on how students learn with AI, an adaptive foundation focused on how AI learns about students, and a responsible design foundation ensuring systems remain transparent, accountable, inclusive, and secure. The framework is illustrated through five case studies spanning diverse educational contexts, levels, and tool designs, revealing both the promise and current limitations of existing tools. We conclude that there is a necessary shift away from LLMs designed for task-oriented performance, and beyond simply prompting them to act as tutors, toward deliberately developed AI learning companions that are pedagogically sound, adapt to their learners, and foster durable understanding, metacognitive growth, and learner agency.
The rise of generative artificial intelligence (GenAI) and accelerated globalization have necessitated a fundamental recalibration of higher education to prioritize domain-agnostic, 21st-century professional competencies. While institutional commitment to these skills is high, their systematic integration into the curriculum and evaluation remains fragmented, highlighting a critical gap between traditional academic success metrics and demonstrated workforce readiness. This special issue presents five complementary studies that investigate how the intersection of learning analytics (LA) and GenAI can bridge the gap between institutional rhetoric and demonstrated professional readiness. The contributions collectively advance a research agenda across four dimensions: 1) benchmarking large language models (LLMs) for curricular-competency alignment using reasoning-based prompting, 2) the iterative design of Socratic-style GenAI chatbots to scaffold self-regulated learning, 3) the application of psychometric modelling and Latent Profile Analysis to quantify 21st-century professional competencies, and 4) institutional governance and adoption of curriculum analytics. Collectively, these studies advocate for an epistemological shift toward processsensitive assessments that move beyond static, episodic indicators toward dynamic, longitudinal representations of learner capability. We conclude by outlining the sociotechnical infrastructure, including robust governance and interdisciplinary collaboration, required to responsibly transition these AI-driven innovations from research prototypes to sustainable enterprise infrastructure, ensuring that analytics serve the evolving needs of students, educators, and professional bodies.
This study investigates how situational interest (SI) influences student interview responses during game-based learning. Using real-time interviews conducted within the What-If Hypothetical Implementations in Minecraft (WHIMC) environment, we analyzed differences in discourse between high- and low-SI students. Interview transcripts were coded using a structured codebook through a hybrid approach combining humans and GPT-4o. Epistemic Networks of student reflections revealed that high-SI students were more likely to offer Brief and Enthusiastic responses across all question types. These students used excited language, reacted aloud to game events, and expressed interest in specific gameplay elements. Low-SI students provided more Explanatory and Neutral responses. They often paused to describe their plans, explain in-game decisions, or reason through moments of uncertainty. Ordered Networks of interviewer statements revealed that interviewers’ strategies remained consistent for low- and high-SI groups, ruling out interviewer behavior as the cause of differences in student responses. These results shed light on the ways student interest levels may impact interview responses during game-based learning experiences.
As student learning transitions to being increasingly 24/7, online courses struggle to provide support for learners on the same schedule. Human TAs are bound by time constraints and are often available only during limited working hours. This can result in a wait time for students seeking answers to their coursework questions. This paper introduces a novel approach to address this need by developing a Virtual Teaching Assistant (TA) that leverages OpenAI’s text embeddings to format and search for data and GPT to adapt the style and content of its responses to align with the typical discourse found in a discussion forum. Our virtual TA, JeepyTA, offers round-the-clock assistance to students, much more rapidly addressing their academic queries. Although still limited in what it can respond to, JeepyTA provides students with responses to their logistic, conceptual, and programming questions, tailored to specific courses. In this paper, we outline the development process, discuss the results, and outline our future plans for a more generalized and versatile Virtual TA catering to a broader range of courses and their differing learning support needs.
There has been considerable research on confusion and frustration that has treated them as two unitary constructs, distinct from each other. In this article, we argue that there is instead a constellation of different types of confusion and frustration, with different antecedents, manifestations, and impacts, and that the commonalities between many types of confusion and frustration justify thinking of them as part of the same constellation of affect, distinct from other prominent affective categories. We discuss how these types of affect have been considered historically and in key models. We then discuss unusual manifestations of each form of affect that have been documented in the literature, and what light they shed on the broader constructs. We conclude with a discussion of a new theoretical framing that treats confusion and frustration as a confrustion constellation, and the opportunities and open questions that this perspective presents.
Programming syntax is highly complex, which has made it consistently difficult to automate feedback for compiler errors in educational settings. Large language models (LLMs) show promise for addressing this issue at scale by providing personalized feedback tailored to specific code submissions, but the effectiveness of this feedback remains uncertain. This study evaluated the impact of GPT-4o at generating real-time feedback for compiler errors during a randomized controlled trial. A total of 248 CS1 students participated, submitting 22,674 pieces of code to an automated programming assessment platform. Students in the Experimental group received LLM feedback, while the Control group did not. Results showed that students who received LLM feedback rated it highly for usefulness, and submitted fewer non -compiling code attempts. These students also had significantly improved performance in terms of resolving errors in consecutive attempts compared to the Control group. Affective surveys revealed that the LLM feedback group self-reported higher focus and lower levels of "confrustion" (a combination of confusion and frustration) after encountering compiler errors. When LLM feedback was temporarily disabled, students in the Experimental group solved programming problems more quickly and demonstrated significant improvement in resolving errors across attempts. However, no significant differences were observed between groups in terms of final scores on a simulated exam. These findings suggest that LLM-generated feedback can improve students' coding experience, engagement, and problem-solving efficiency in the initial phases of computer science education, though it may not lead to better final performance.
This study explores the potential of the large language model GPT-4 as an automated tool for qualitative data analysis by educational researchers, exploring which techniques are most successful for different types of constructs. Specifically, we assess three different prompt engineering strategiesA- Zero-shot, Few-shot, and Fewshot with contextual informationA- as well as the use of embeddings. We do so in the context of qualitatively coding three distinct educational datasets: Algebra I semi-personalized tutoring session transcripts, student observations in a game-based learning environment, and debugging behaviours in an introductory programming course. We evaluated the performance of each approach based on its inter-rater agreement with human coders and explored how different methods vary inAeffectiveness depending on a construct's degree of clarity, concreteness, objectivity, granularity, and specificity. Our findings suggest that while GPT-4 can code aAbroad range of constructs, no single method consistently outperforms the others, and the selection of a particular method should be tailored to the specific properties of the construct and context being analyzed. We also found that GPT-4 has the most difficulty with the same constructs than human coders find more difficult to reach inter-rater reliability on.
Data Driven Classroom Interviews (DDCIs) are an interviewing technique that is facilitated by recent technological developments in the learning analytics community. DDCIs are short, targeted interviews that allow researchers to contextualize students' interactions with a digital learning environment (e.g., intelligent tutoring systems or educational games) while minimizing the amount of time that the researcher interrupts that learning experience, and focusing researcher time on the events they most want to focus on DDCIs are facilitated by a research tool called the Quick Red Fox (QRF)–an open-source server-client Android app that optimizes researcher time by directing interviewers to users that have just displayed an interesting behavior (previously defined by the research team). QRF integrates with existing student modeling technologies (e.g., behavior-sensing, affect-sensing, detection of self-regulated learning) to alert researchers to key moments in a learner's experience. This manual documents the tech while providing training on the processes involved in developing triggers and interview techniques; it also suggests methods of analyses.
This study explores the ability of GPT-4 working together with humans to generate a codebook to analyze scientific observations from middle school learners in the What-if Hypothetical Implementations in Minecraft (WHIMC) project. It compares this Hybrid codebook to one fully developed by Humans using a variety of techniques to evaluate how the codes developed by each approach relate to one another and to external measures of student interest. Results show that the Hybrid GPT-Human codes consist of broader categories that align more consistently with the external interest metrics, whereas the Human codes offer finer-grained insights into specific student behaviors. However, the complementary insights offered by each suggest that combining both approaches could improve our understanding of student engagement and inform more effective strategies in educational game design and intervention.
Prior work has developed a range of automated measures ("detectors") of student self-regulation and engagement from student log data. These measures have been successfully used to make discoveries about student learning. Here, we extend this line of research to an underexplored aspect of self-regulation: students' decisions about when to start and stop working on learning software during classwork. In the first of two analyses, we build on prior work on session-level measures (e.g., delayed start, early stop) to evaluate their reliability and predictive validity. We compute these measures from year-long log data from Cognitive Tutor for students in grades 8-12 (N = 222). Our findings show that these measures exhibit moderate to high month-to-month reliability (G > .75), comparable to or exceeding gaming-the-system behavior. Additionally, they enhance the prediction of final math scores beyond prior knowledge and gaming-the-system behaviors. The improvement in learning outcome predictions beyond time-on-task suggests they capture a broader motivational state tied to overall learning. The second analysis demonstrates the cross-system generalizability of these measures in i-Ready, where they predict state test scores for grade 7 students (N = 818). By leveraging log data, we introduce system-general naturally embedded measures that complement motivational surveys without extra instrumentation or disruption of instruction time. Our findings demonstrate the potential of session-level logs to mine valid and generalizable measures with broad applications in the predictive modeling of learning outcomes and analysis of learner self-regulation.
This study investigates how students explored a Minecraft-based science learning environment by analyzing their in-game movement trajectories. We use GPT-4o to identify recurring trajectory patterns from gameplay visualizations and to automatically label trajectory images, with some constructs labeled and reviewed by humans. These patterns are examined in relation to students’ self-reported survey measures. Epistemic Network Analysis is used to compare how different movement behaviors co-occurred across different learner profiles. The findings showed that students with high or improving outcomes engaged in more flexible exploration. They often wandered, changed directions, and alternated between looking around the environment before closely examining specific objects. In contrast, students with low or declining outcomes tended to concentrate on specific areas and frequently backtracked to previously visited locations. These findings highlight the importance of how, not just how much, students explore in open-ended environments as part of the learning process.
Understanding the sequence of player decisions in open-ended educational games provides insight into how those decisions influence player persistence or readiness for later challenges. This study uses Ordered Network Analysis to examine how players move between jobs of varying difficulty in the educational game Wake. To scaffold players, Wake breaks down multi-phase scientific investigations into smaller “jobs”. Each job is manually coded based on its difficulty level in Experimentation, Modeling, or Argumentation. We use these difficulty ratings, along with whether the player completed or quit it, as codes to model player progression. We see that players who completed a job with a high quit rate on their first attempt more often followed paths with gradually increasing difficulty prior to accepting that job. In contrast, other players who quit the same job with a high quit rate on their first attempt were more likely to have failed prior jobs requiring basic skills in Experimentation or Modeling when moving from jobs that did not involve such components. They also tended to remain within jobs without such components across multiple transitions, which may reflect lower preparedness or content knowledge compared to those who completed the later difficult job. Findings also show that players who completed the difficult job on their second attempt spent the time between attempts completing jobs with lower difficulty, which may have helped strengthen foundational skills relevant to the target job or restore confidence. These findings point to opportunities for progression-aware intervention design based on how successful and unsuccessful players move through different types of jobs.