The broad adoption of Generative AI (GenAI) is impacting Computer Science education, and recent studies found its benefits and potential concerns when students use it for programming learning. However, most existing explorations focus on GenAI tools that primarily support text-to-text interaction. With recent developments, GenAI applications have begun supporting multiple modes of communication, known as multimodality. In this work, we explored how undergraduate programming novices choose and work with multimodal GenAI tools, and their criteria for choices. We selected a commercially available multimodal GenAI platform for interaction, as it supports multiple input and output modalities, including text, audio, image upload, and real-time screen-sharing. Through 16 think-aloud sessions that combined participant observation with follow-up semi-structured interviews, we investigated student modality choices for GenAI tools when completing programming problems and the underlying criteria for modality selections. With multimodal communication emerging as the future of AI in education, this work aims to spark continued exploration on understanding student interaction with multimodal GenAI in the context of CS education.
Educational technology (EdTech) solutions have shown promise for disseminating educational opportunities to last-mile learners, particularly in the Global South. Low-infrastructure EdTech as a digital learning resource is especially critical to understand in remote contexts where educational opportunities and resources are limited. Our work investigated insights from 81 learners who engaged with a remote course that provided engineering education through radios and mobile phones in rural Uganda. Findings revealed that the course facilitated goal-driven and practical motivations in safe, adaptable environments. Our work goes beyond the idea that low-infrastructure EdTech can easily facilitate learning, highlighting diverse learner experiences navigating radio and phone use and presenting novel findings on community skepticism towards the course. Our research extends the EdTech and HCI literature by bringing light to the underrepresented voices of last-mile learners, sharing their insights on interacting with low-infrastructure EdTech, and how these insights can guide the design of contextually aligned EdTech.
K–12 educators increasingly use large language models (LLMs) to draft lesson plans and learning activities, but outputs often lack the pedagogical structure needed for high-quality instruction. While sophisticated prompting can help, it requires substantial teacher time and expertise. We therefore shift pedagogical expertise from prompts into system architecture by embedding the Knowledge–Learning–Instruction (KLI) framework in two multi-agent systems (MAS), and compare them with a single-agent baseline that simulates typical teacher prompting. We evaluated the three systems on secondary math and science activity generation in a teacher study using an adapted Quality Matters K–12 rubric and open-ended feedback. Overall rubric score differences were not statistically significant across systems. However, MAS-CMD received the highest mean total score and received more favorable qualitative feedback, particularly around engagement, collaboration, and differentiation. More broadly, this study positions theory-guided AI instructional design as a promising direction for further investigation.
This quasi-experimental study integrates a large language model (LLM) with expert qualitative analysis to examine how instructional design variations in computer-supported collaborative learning (CSCL) shape collaboration in data-centric programming. We collected 73 team transcripts from two contrasting CSCL designs deployed across five course offerings: a closed-ended variant with prescribed solution paths and auto-graded milestones, and an open-ended variant supporting exploratory tasks with multiple valid paths. LLM annotation revealed statistically significant differences in knowledge co-construction patterns: the open-ended design yielded a higher proportion of utterances focused on developing a shared understanding of problems and solutions. Guided by these quantitative results, human experts conducted qualitative coding that confirmed and enriched these findings, showing how open-ended tasks fostered elaborative solution negotiation while closed-ended structures promoted non-elaborative exchanges. Our contributions are: (1) instructional insights for data science education, demonstrating how open-ended CSCL designs better support collaborative sense-making essential for real-world data science projects; and (2) a documented workflow for human-LLM co-evaluation, providing the methodological detail necessary for others to replicate our process and apply it to future studies.
Generative artificial intelligence (GenAI) is rapidly entering K-12 classrooms worldwide, initiating urgent debates about its potential to either reduce or exacerbate educational inequalities. Drawing on interviews with 30 K-12 teachers across the United States, South Africa, and Taiwan, this study examines how teachers navigate this GenAI tension around educational equalities. We found teachers actively framed GenAI education as an equality-oriented practice: they used it to alleviate pre-existing inequalities while simultaneously working to prevent new inequalities from emerging. Despite these efforts, teachers confronted persistent systemic barriers, i.e., unequal infrastructure, insufficient professional training, and restrictive social norms, that individual initiative alone could not overcome. Teachers thus articulated normative visions for more inclusive GenAI education. By centering teachers' practices, constraints, and future envisions, this study contributes a global account of how GenAI education is being integrated into K-12 contexts and highlights what is required to make its adoption genuinely equal.
As Artificial Intelligence (AI) becomes increasingly integrated into daily life, there is a growing need to equip the next generation with the ability to apply, interact with, evaluate, and collaborate with AI systems responsibly. Prior research highlights the urgent demand from K-12 educators to teach students the ethical and effective use of AI for learning. To address this need, we designed a Large-Language Model (LLM)-based module to teach prompting literacy. This includes scenario-based deliberate practice activities with direct interaction with intelligent LLM agents, aiming to foster secondary school students' responsible engagement with AI chatbots. We conducted two iterations of classroom deployment in 11 authentic secondary education classrooms, and evaluated 1) AI-based auto-grader's capability; 2) students' prompting performance and confidence changes towards using AI for learning; and 3) the quality of learning and assessment materials. Results indicated that the AI-based auto-grader could grade student-written prompts with satisfactory quality. In addition, the instructional materials supported students in improving their prompting skills through practice and led to positive shifts in their perceptions of using AI for learning. Furthermore, data from Study 1 informed assessment revisions in Study 2. Analyses of item difficulty and discrimination in Study 2 showed that True/False and open-ended questions could measure prompting literacy more effectively than multiple-choice questions for our target learners. These promising outcomes highlight the potential for broader deployment and highlight the need for broader studies to assess learning effectiveness and assessment design.
Learning engineering applies data and learning science principles to better understand outcomes and support improvement research. One important approach is A/B testing-common in large software companies and also represented academically at conferences like the Annual Conference on Digital Experimentation (CODE), and the International Consortium for Innovation and Collaboration in Learning Engineering (IEEE ICICLE). Several systems supporting A/B testing in educational applications have arisen recently, including UpGrade, E-TRIALS, and Terracotta. A/B testing can help improve educational platforms, yet there are challenging issues unique to conducting such work in these contexts. In response, a number of digital learning platforms have opened their systems to learning-improvement research by instructors and/or third-party researchers, with specific supports necessary for education-specific research designs. This workshop will explore how A/B testing is conducted in educational contexts, how digital learning platforms are accelerating education research, and how empirical approaches can be used to drive powerful gains in student learning. It will also discuss opportunities for funding to conduct platform-enabled learning engineering.
Strengthening our informal reasoning skills may be our best defense against the misinformation and disinformation permeating this new media landscape. However, designing theoretically-grounded instructional systems that mimic the unique constraints of real-world everyday reasoning is challenging. In a series of two experiments, we explore the viability of distracted reasoning as a paradigm for difficulty factors assessment in the informal reasoning space. Participants in both experiments saw performance improvements on both the detection and classification of informal reasoning errors. Importantly, the level of distraction affected the impact of theorized difficulty factors (such as argument plausibility and political alignment). These results suggest that a distracted reasoning paradigm may offer a reasonable analogue for the unique constraints of everyday informal reasoning.
Instructional designers are increasingly integrating generative AI (genAI) into their workflows to assist with the creation of educational content. However, particularly for novices, it remains unclear whether this integration genuinely elevates the quality of the resulting content. We conducted an A/B field experiment within a 14-week graduate educational technology course (n = 27) to examine this potential impact on content quality. Students created eight microlessons across four modules, alternating between receiving genAI assistance (GPT-4 via ChatGPT) and working without genAI. The quality of the microlessons was assessed via a 5-criteria rubric. Results showed that genAI-assisted microlessons received significantly higher quality scores than non-genAI microlessons for half of the assignments and never scored lower on average. Our findings suggest that thoughtful integration of genAI can enhance the quality of materials created by novices.
Culturally Relevant Pedagogy (CRP) is vital in K-12 education, yet teachers struggle to implement CRP into practice due to time, training, and resource gaps. This study explores how Large Language Models (LLMs) can address these barriers by introducing CulturAIEd, an LLM tool that assists teachers in adapting AI literacy curricula to students' cultural contexts. Through an exploratory pilot with four K-12 teachers, we examined CulturAIEd's impact on CRP integration. Results showed CulturAIEd enhanced teachers' confidence in identifying opportunities for cultural responsiveness in learning activities and making culturally responsive modifications to existing activities. They valued CulturAIEd's streamlined integration of student demographic information, immediate actionable feedback, which could result in high implementation efficiency. This exploration of teacher-AI collaboration highlights how LLM can help teachers include CRP components into their instructional practices efficiently, especially in global priorities for future-ready education, such as AI literacy.
Expanding access to education in rural African communities remains difficult, largely due to limited internet connectivity. Mobile learning courses delivered via radio and offline mobile phones offer a promising, scalable solution. However, it is challenging to track student engagement in these environments due to the absence of tools that monitor students' interactions with the radio. In this study, we investigate the potential of ''Prize Codes'' -- codes read aloud during broadcasts that students enter via text message -- to serve as a real-time measure of student engagement with mobile-learning broadcasts. Using data from a 2024 implementation of Yiya AirScience, a mobile engineering course in Uganda, we evaluate the validity of Prize Codes as an engagement metric. Specifically, we test whether Prize Code measures (1) demonstrate reliability, with students who enter correct codes in one lesson being more likely to do so in subsequent lessons; (2) demonstrate convergent validity with existing measures of engagement; and (3) demonstrate predictive validity, predicting learning outcomes in the course. Our findings suggest that Prize Codes are a reliable and valid measure of engagement. Prize-Code accuracy demonstrates strong internal consistency (alpha = .97) and moderate test-retest reliability (ICC = .44). The measure aligns closely with synchronous participation (87% agreement, Cohen's kappa = .50), indicating it captures similar engagement patterns. Importantly, students who consistently enter correct Prize Codes perform significantly better on assessments, with Prize Code engagement predicting final exam scores above and beyond other engagement metrics. After establishing the measure's validity, we use it to (1) characterize patterns of engagement with Yiya broadcasts, (2) investigate early engagement with the broadcasts as a predictor of course persistence, and (3) replicate findings about the benefits of learning by doing. This study suggests that Prize Codes can be a feasible, scalable approach for tracking real-time engagement in resource-limited mobile learning settings at scale.
Adaptive learning technologies observe learner performance, infer mastery, and dynamically tailor instruction. However, this approach can fall short when encountering a learning phenomenon that we term "deceptive overgeneralization," where learners perform correct actions based on incomplete understanding. This phenomenon "deceives" adaptive systems into prematurely stopping necessary practice, leaving overgeneralization unaddressed. To address this, we propose a theoretical model of deceptive overgeneralization, grounded in Adaptive Control of Thought-Rational (ACT-R) and the Knowledge-Learning-Instruction (KLI) framework. This work contributes both theoretically and practically to enhancing the precision of adaptive learning, enabling more accurate mastery assessments and improved learning outcomes.
With the proliferation of large language model (LLM) applications since 2022, their use in education has sparked both excitement and concern. Recent studies consistently highlight students' (mis)use of LLMs can hinder learning outcomes. This work aims to teach students how to effectively prompt LLMs to improve their learning. We first proposed pedagogical prompting, a theoretically-grounded new concept to elicit learning-oriented responses from LLMs. To move from concept design to a proof-of-concept learning intervention in real educational settings, we selected early undergraduate CS education (CS1/CS2) as the example context. We began with a formative survey study with instructors (N=36) teaching early-stage undergraduate-level CS courses to inform the instructional design based on classroom needs. Based on their insights, we designed and developed a learning intervention through an interactive system with scenario-based instruction to train pedagogical prompting skills. Finally, we evaluated its instructional effectiveness through a user study with CS novice students (N=22) using pre/post-tests. Through mixed methods analyses, our results indicate significant improvements in learners' LLM-based pedagogical help-seeking skills, along with positive attitudes toward the system and increased willingness to use pedagogical prompts in the future. Our contributions include (1) a theoretical framework of pedagogical prompting; (2) empirical insights into current instructor attitudes toward pedagogical prompting; and (3) a learning intervention design with an interactive learning tool and scenario-based instruction leading to promising results on teaching LLM-based help-seeking. Our approach is scalable for broader implementation in classrooms and has the potential to be integrated into tools like ChatGPT as an on-boarding experience to encourage learning-oriented use of generative AI.
GPT has become nearly synonymous with large language models (LLMs), an increasingly popular term in AIED proceedings. A simple keyword-based search reveals that 61% of the 76 long and short papers presented at AIED 2024 describe novel solutions using LLMs to address some of the long-standing challenges in education, and 43% specifically mention GPT. Although LLMs pioneered by GPT create exciting opportunities to strengthen the impact of AI on education, we argue that the field's predominant focus on GPT and other resource-intensive LLMs (with more than 10B parameters) risks neglecting the potential impact that small language models (SLMs) can make in providing resource-constrained institutions with equitable and affordable access to high-quality AI tools. Supported by positive results on knowledge component (KC) discovery, a critical challenge in AIED, we demonstrate that SLMs such as Phi-2 can produce an effective solution without elaborate prompting strategies. Hence, we call for more attention to developing SLM-based AIED approaches.
We propose the third annual workshop on Learnersourcing: Student-generated Content @ Scale, reimagined for an era where AI is rapidly transforming the educational landscape. This full- day workshop is designed to explore the vast potential of learnersourcing, which combines the efforts of humans, AI, and other data sources to create and assess educational materials. invites instructors, researchers, learning engineers, and professionals from diverse backgrounds to explore the innovative intersection of human insight and AI-sourced content in learnersourcing. By drawing on principles from education, crowdsourcing, learning analytics, data mining, and natural language processing, we aim to foster an environment where all participants can learn from and contribute to the conversation. As AI tools become increasingly integral to content creation, our program will delve into pedagogically sound learnersourcing practices that effectively integrate these technologies. Participants will examine the broader implications of AI on learnersourcing, from ensuring content quality to fostering authentic student engagement, and will be provided with practical guidelines for incorporating AI in ways that enrich learning. Through bringing together these different perspectives, attendees will leave equipped with both a conceptual framework and practical strategies to engage with learnersourcing.
The Doer Effect states that completing more active learning activities, like practice questions, is more strongly related to positive learning outcomes than passive learning activities, like reading, watching, or listening to course materials. Although broad, most evidence has emerged from practice with tutoring systems in Western, Industrialized, Rich, Educated, and Democratic (WEIRD) populations in North America and Europe. Does the Doer Effect generalize beyond WEIRD populations, where learners may practice in remote locales through different technologies? Through learning analytics, we provide evidence from N = 234 Ugandan students answering multiple-choice questions via phones and listening to lectures via community radio. Our findings support the hypothesis that active learning is more associated with learning outcomes than passive learning. We find this relationship is weaker for learners with higher prior educational attainment. Our findings motivate further study of the Doer Effect in diverse populations. We offer considerations for future research in designing and evaluating contextually relevant active and passive learning opportunities including leveraging familiar technology, increasing the number of practice opportunities, and aligning multiple data sources.