Educational content is typically designed for a broad audience, often failing to address the specific needs, contexts, and backgrounds of individual learners. While certain educational campaigns (e.g., public health outreach) develop multiple targeted versions of learning content, historically it has been infeasible to do at scale. However, Large Language Models (LLMs) offer the potential to adapt content based on detailed descriptions of learner characteristics, such as demographic information, situational context, resources availability, and risk factors. To explore techniques to generate such content, we developed Generative AI for Micro-Tailored Adaptation (GAIMA), a multi-agent LLM framework designed to personalize educational and training documents. GAIMA employs a feedback-driven pipeline architecture where a content modification agent generates personalized adaptations and a feedback moderator agent evaluates quality, safety, and educational value. This iterative process refines content through multiple cycles until it meets detailed standards for personalization depth and learner appropriateness. We evaluate the system using a composite framework that includes ROUGE-L and BERTScore, style transfer ratio, expertise recall, NLI-based faithfulness, and separate LLM-as-judge scores for relevance and grounding. We present a comparative analysis against a zero-shot LLM baseline, quantifying the value of iterative feedback. Our results demonstrate that GAIMA achieves 23.1% improvement over zero-shot baselines while generating more personalized, context-aware content that maintains educational integrity and safety standards.
As AI tools become common across jobs and industries, it is critical to broaden education about AI beyond teaching computer scientists how to build AI systems. To expand AI education, we are researching AI for AI learning: a personalized and adaptive learning system that integrates dialog-based tutoring and gamified programming activities. To study this problem, we adapted and expanded an existing smartphone adaptive coach to develop the Game-if-AI system. Using a design-based research approach, Game-if-AI was iteratively tested and improved across four semesters of optional use in a course designed for technician-level understanding of AI: mastering programming skills to apply AI libraries and established models. In this study, we measured the interests and needs of these technical learners, based on both survey data and on how they engaged with topics in the system. Based on this data, new topics were added and the system was refined. In this paper, we report students' usability ratings for system components and student preferences based on completion rates of AI topics available each semester. Students rated the adaptive system positively overall (93% rated as a "good idea"), but more complex learning activities (tutoring dialogs, programming) were rated lower than traditional ones (e.g., multiple choice, reading). Students were most likely to master topics highly aligned to the course materials, as well as self-directed learning toward easier high-interest topics (e.g., LLM Prompting).
Research to improve Automated Short Answer Grading has recently focused on Large Language Models (LLMs) with prompt engineering and no- or few-shot prompting to achieve best results. This is in contrast to the fine-tuning approach, which has historically required large-scale compute clusters inaccessible to most users. New closed-model approaches such as OpenAI's fine-tuning service promise results with as few as 100 examples, while methods using open weights such as quantized low-rank adaptive (QLORA) can be used to fine-tune models on consumer GPUs. We evaluate both of these fine-tuning methods, measuring their interaction with few-shot prompting for automated short answer grading (ASAG) with structured (JSON) outputs. Our results show that finetuning with small amounts of data has limited utility for Llama open-weight models, but that fine-tuning methods can outperform few-shot baseline instruction-tuned LLMs for OpenAI's closed models. While our evaluation set is limited, we find some evidence that the observed benefits of finetuning may be impacted by the domain subject matter. Lastly, we observed dramatic improvement with the LLama 3.1 8B-Instruct open-weight model by seeding the initial training examples with a significant amount of cheaply generated synthetic training data.
Assessing teams and providing feedback on scenario-based training typically requires human observers or scenario-specific metrics crafted by experts, due to the complexity of general-purpose automated tools to assess team performance. Machine learning can help infer team performance patterns, but labeled data for a specific training scenario is often sparse. To address this issue, the Semi-Supervised Learning for Assessing Team Simulations (SLATS) project investigated the feasibility of semi-supervised learning and transfer learning which leverages training data from related scenarios to classify performance on a target scenario with the same metrics but a different terrain context. To this approach, we analyzed performance of teams in the first-person shooter Team Fortress 2 (TF2). TF2 teams for the “Capture Point” mode were classified into archetypes based on the performance of the team and the performance of individual members of the team across the corpus: novice, weak link, team of experts, and expert team. To investigate the feasibility of transfer learning, we isolated matches from two of the most frequent maps/terrains. Results found that leveraging data from the source map always improved classification F1-scores compared to relying solely upon target (test) map training data. The greatest benefits were observed when target data was limited (0 to 42 target examples). While further research is required to explore the effectiveness of transfer learning across training scenarios that are more dissimilar (e.g., different simulations, rather than just different maps), these results offer a promising direction to help bootstrap team assessments on new training scenarios by leveraging data from earlier, comparable scenarios. However, efficiently calculating reusable metrics for model features based on low-level scenario events and logs remains a challenge that requires further research.
Despite the critical role of teachers in the educational process, few advanced learning technologies have been developed to support teacher-instruction or professional development. This lack of support is particularly acute for middle school math teachers, where only 37% felt well prepared to scaffold instruction to address the needs of diverse students in a national sample. To address this gap, the Advancing Middle School Teachers’ Understanding of Proportional Reasoning project is researching techniques to apply pedagogical virtual agents and dialog-based tutoring to enhance teachers' content knowledge and pedagogical content knowledge. This paper describes the design of a conversational, agent-based intelligent tutoring system to support teachers' professional development. Pedagogical strategies are presented that leverage a virtual human facilitator to tutor pedagogical content knowledge (how to teach proportions to students), as opposed to content knowledge (understanding proportions). The roles for different virtual facilitator capabilities are presented, including embedding actions into virtual agent dialog, open-response versus choice-based tutoring, ungraded pop-up sub-activities (e.g. whiteboard, calculator, note-taking). Usability feedback for a small cohort of instructors pursuing graduate studies was collected. In this feedback, teachers rated the system ease of use and perceived usefulness moderately well, but also reported confusion about what to expect from the system in terms of flow between lessons and support by the facilitator.
Reinforcement Learning (RL) has been applied successfully to Intelligent Tutoring Systems (ITSs) in a limited set of well-defined domains such as mathematics and physics. This work is unique in using a large state space and for applying RL to tutoring interpersonal skills. Interpersonal skills are increasingly recognized as critical to both social and economic development. In particular, this work enhances an ITS designed to teach basic counseling skills that can be applied to challenging issues such as sexual harassment and workplace conflict. An initial data collection was used to train RL policies for the ITS, and an evaluation with human participants compared a hand-crafted ITS which had been used for years with students (control) versus the new ITS guided by RL policies. The RL condition differed from the control condition most notably in the strikingly large quantity of guidance it provided to learners. Both systems were effective and there was an overall significant increase from pre- to post-test scores. Although learning gains did not differ significantly between conditions, learners had a significantly higher self-rating of confidence in the RL condition. Confidence and learning gains were both part of the reward function used to train the RL policies, and it could be the case that there was the most room for improvement in confidence, an important learner emotion. Thus, RL was successful in improving an ITS for teaching interpersonal skills without the need to prune the state space (as previously done).
Facial expression trackers output measures for facial action units (AUs), and are increasingly being used in learning technologies. In this paper, we compile patterns of AUs seen in related work as well as use factor analysis to search for categories implicit in our corpus. Although there was some overlap between the factors in our data and previous work, we also identified factors seen in the broader literature but not previously reported in the context of learning environments. In a correlational analysis, we found evidence for relationships between factors and self-reported traits such as academic effort, study habits, and interest in the subject. In addition, we saw differences in average levels of factors between a video watching activity, and a decision making activity. However, in this analysis, we were not able to isolate any facial expressions having a significant positive or negative relationship with either learning gain, or performance once question difficulty and related factors were also considered. Given the overall low levels of facial affect in the corpus, further research will explore different populations and learning tasks to test the possible hypothesis that learners may have been in a pattern of “Over-Flow” in which they were engaged with the system, but not deeply thinking about the content or their errors.
Scenario-based tutoring systems influence affective states due to two distinct mechanisms during learning: (1) reactions to performance feedback and (2) responses to the scenario context or events. To explore the role of affect and engagement, a scenario-based ITS was instrumented to support unobtrusive facial affect detection. Results from a sample of university students showed relatively few traditional academic affective states such as confusion or frustration, even at decision points and after poor performance (e.g., incorrect responses). This may show evidence of “over-flow,” with a high level of engagement and interest but insufficient confusion/disequilibrium for optimal learning.
Scenario-based training systems pose an especially difficult challenge for an intelligent tutoring system (ITS). In addition to the basic problems of deciding when to intervene and what guidance to provide, the ITS must decide whether to give guidance directly (e.g., a hint message), indirectly through positive/negative results in the scenario, or to delay guidance until a post-scenario review session. There are a number of factors that an adaptive ITS should consider and we use self-report survey instruments to investigate the relationship between traits, learning strategies, expectations, learner behaviors derived from log files, post-use perceptions of the system, and pre-test and post-test results. We use the ELITE Lite Counseling training system as a testbed for our experiments. This system uses virtual role players to allow learners to practice leadership counseling skills, and is in use at the United States Military Academy (USMA). This paper analyzes two data sets. We collected data from local university students, a nonmilitary population of roughly the same age as USMA Cadets using the system. For these local participants, we could administer surveys and pre-tests and post-tests, and collect log files recording clicks made while using ELITE Lite. The second data set comes from USMA itself but is limited to log files. In both populations, the ITS’s hints are effective at boosting scenario performance, and for the university students, the overall experience promoted learning, and survey results suggest that higher levels of organization in study habits may lead to greater learning with ELITE Lite. For the USMA Cadets, ELITE Lite is part of their Military Leadership course rather than an experiment, which could explain why we found higher scenario performance on average than the non-military population, and more use of the post-scenario review feature.
We describe the Situated Pedagogical Authoring (SitPed) system that seeks to allow non-technical authors to create ITS content for soft-skills training, such as counseling skills. SitPed is built on the assertion that authoring tools should use the learner’s perspective to the greatest extent possible. SitPed provides tools for creating tasks lists, authoring assessment knowledge, and creating tutor messages. We present preliminary findings of a two-phase study comparing authoring in SitPed to an ablated version of the same system and a spreadsheet-based control. Findings suggest modest advantages for SitPed in terms of the quality of the authored content and student learning.
Serious games are generally designed with two goals in mind: promoting learning and creating compelling and engaging experiences (sometimes termed a sense of presence). Presence itself is believed to promote learning, but serious games often attempt to further increase pedagogical value. One way to do so is to use an intelligent tutoring system (ITS) to provide feedback during gameplay. Some researchers have expressed concern that, because feedback from an ITS is often extrinsic (i.e., it operates outside of the primary game mechanic), attending to it disrupts players’ sense of presence. As a result, learning may be unintentionally hindered by an ITS. However, the most beneficial conditions of instruction are often counterintuitive; in this paper, we challenge the assumption that feedback during learning hinders sense of presence. Across three experiments, we examined how an ITS that provided extrinsic feedback during a serious game affected presence. Across different modalities and conditions, we found that feedback and other ITS features do not always affect presence. Our results suggest that it is possible to provide extrinsic feedback in a serious game without detracting from the immersive power of the game itself.
We describe Coach Mike, an animated pedagogical agent for informal computer science education, and report findings from two experiments that provide initial evidence for the efficacy of the system. In the first study, we found that Coach Mike’s presence led to 20% longer holding times, increased acceptance of programming challenges, and reduced misuse of the exhibit, but had limited cumulative impact on attitudes, awareness, and knowledge beyond what the host exhibit already achieved. In the second study, we compared two different versions of Coach Mike and found that the use of enthusiasm and self-regulatory feedback led to greater self-efficacy for programming.
In the context of practicing intercultural communication skills, we investigated the role of fidelity in a game-based, virtual learning environment as well as the role of feedback delivered by an intelligent tutoring system. In 2 experiments, we compared variations on the game interface, use of the tutoring system, and the form of the feedback. Our findings suggest that for learning basic intercultural communicative skills, a 3-dimensional (3-D) interface with animation and sound produced equivalent learning to a more static 2-D interface. However, learners took significantly longer to analyze and respond to the actions of animated virtual humans, suggesting a deeper engagement. We found large gains in learning across conditions. There was no differential effect with the tutor engaged, but it was found to have a positive impact on learner success in a transfer task. This difference was most pronounced when the feedback was delivered in a more general form versus a concrete style.
In this paper, we describe Coach Mike, a virtual staff member at the Boston Museum of Science that seeks to help visitors at Robot Park, an interactive exhibit for computer programming. By tracking visitor interactions and through the use of animation, gestures, and synthesized speech, Coach Mike provides several forms of support that seek to improve the experiences of museum visitors. These include orientation tactics, exploration support, and problem solving guidance. Additional tactics use encouragement and humor to entice visitors to stay more deeply engaged. Preliminary analysis of interaction logs suggest that visitors can follow Coach Mike's guidance and may be less prone to immediate disengagement, but further study is needed.
We investigate the role of presence in a serious game for intercultural communication and negotiation skills by comparing two interfaces: a 3D version with animated virtual humans and sound against a 2D version using text-only interactions with static images and no sound. Both versions provide identical communicative action choices and are driven by the same underlying simulation engine. In a study, the 3D interface led to a significantly greater self-reported sense of presence, but produced significant, but equivalent learning on immediate posttests for declarative and conceptual knowledge related to intercultural communication. Log data reveals that 3D learners needed fewer interactions with the system than those in the 2D environment, suggesting they benefited equally with less practice and may have treated the experience as more authentic.
The role of explicit feedback in learning has been studied from a variety of perspectives and in many contexts. In this paper, we examine the impact of the specificity of feedback delivered by an intelligent tutoring system in a game-based environment for cultural learning. We compared two versions: one that provided only “bottom-out” hints and feedback versus one that provided only conceptual messages. We measured during-training performance, in-game transfer, and long-term retention. Consistent with our hypotheses, specific feedback utterances produced inferior learning on the in-game transfer task when compared to conceptual utterances. No differences were found on a web-based post-test. We discuss possible explanations for these findings, particularly as they relate to the learning of loosely defined skills and serious games.
Truly generic, reusable intelligent tutoring software frameworks remain elusive. As part of our effort to develop ITSs for simulations, a software framework with minimal dependencies on domain specifics has emerged. Herein, we describe this framework, its functionality, components, configurability, and use of natural language generation.