Generative AI poses a significant threat to data integrity on crowdsourcing platforms like Prolific, which behavioral scientists widely rely on for data collection. Large language models (LLMs) allow users to generate fluent and relevant responses to open-ended questions, which can mask inattention and compromise experimental validity. To empirically estimate the prevalence of this behavior, we analyzed keystroke data from three studies (N = 928) on Prolific between May and July 2025. Using an embedded JavaScript tool, we flagged participants who pasted text or whose keystroke count was anomalously low compared to their response length. For each flagged participant, we manually compared detected keystrokes to their final response to determine if the text could have been typed. This confirmed that, despite deterrence measures, approximately 9% of participants submitted responses consistent with AI assistance or other forms of outsourced responding. These participants outperformed non-cheaters (by up to 1.5 SDs), were over twice as likely to share geolocations with other participants (suggesting possible proxy use), and exhibited lower internal consistency on questionnaire scales. Simulated power analyses indicate that this level of undetected cheating can diminish observed effect sizes by 10% and inflate required sample sizes by up to 30%. These findings highlight the urgent need for new detection methods like keystroke logging, which offers verifiable evidence of cheating that is difficult to obtain from manual review of LLM-generated text alone. As AI continues to evolve, maintaining data quality in crowdsourced research will require active monitoring, methodological adaptation, and communication between researchers and platforms.
Persistence after failure is critical for learning—but when students make mistakes in intelligent tutoring systems, they often choose not to try again. How can digital platforms encourage students to persist at these moments? We conducted a randomized controlled trial in an intelligent tutoring system for math and science, involving 164,532 students (Grades 8-12) who completed 17 million practice problems. We tested two scalable interventions: a brief persuasive prompt encouraging students to try again, and a visual default nudge that highlighted the retry option. Both interventions increased persistence after failure, and when combined, their effects were additive—suggesting they operate through distinct psychological mechanisms. The nudge had a much larger immediate effect, but the prompt showed proportionally greater spillover to untreated problems. These findings advance theories of persuasive design, demonstrating that implicit, interface-level nudges and explicit motivational prompts can be combined to avoid redundancy while amplifying impact.
Given the cumulative nature of computer science, success in introductory computing (CS1) courses requires students to not only learn the material but also develop effective self-regulated learning (SRL) habits. While theories of SRL emphasize planning, performance, and self-reflection as essential phases of effective learning, there is limited evidence on how to help learners put these phases into practice. In this context, Mastery-Based Tests (MBT), which allow students to retake assessments after receiving feedback, have shown promise for improving learning outcomes. However, prior work in computer science is largely observational and does not directly test MBT's impact on SRL behaviors. This paper presents a pilot study (N = 6) exploring this relationship in CS1. Using a between-subjects design, we observed that learners who first completed an MBT achieved higher post-test scores, demonstrated higher metacognitive accuracy, and self-reported more productive SRL behaviors. These patterns suggest that MBTs warrant further investigation as a viable scaffold for fostering self-regulation in CS1.
The benefits of learning in one's mother tongue are well documented, yet colonial languages dominate education, marginalizing local languages and limiting access for learners who rely on their mother tongue for understanding. With the rapid growth of educational technology, there is potential to integrate multilingual instruction supporting both colonial and local languages. This study is part of a larger quasi-experiment conducted in Uganda, where learners could choose to learn in English, Leb-Lango (a local language), or in Hybrid mode (a combination of both) in a remote EdTech course. We examined how learners who chose the Hybrid option navigated English and Leb-Lango. While many Hybrid learners did not consistently use both languages, those who did persisted longer in the course. Learners also shared how they managed language complexities. We provide the first empirical evidence of learner agency in bilingual remote EdTech instruction and offer insights for designing inclusive multilingual learning solutions.
Most learners worldwide are multilingual, yet implementing multilingual education remains challenging in practice. EdTech offers an opportunity to bridge this gap and expand access for linguistically diverse learners. We conducted a quasi-experiment in Uganda with 2,931 participants enrolled in a non-formal radio- and mobile-based engineering course, where learners self-selected instruction in Leb Lango (a local language), English, or a Hybrid option combining both languages. The Leb Lango version of the course was used disproportionately by learners from rural areas, those with less formal education, and those with lower prior knowledge, broadening participation among disadvantaged learners. Moreover, the availability of Leb Lango instruction was associated with higher active participation, even among learners who registered for English instruction. Although Leb Lango learners began with lower performance, they demonstrated faster learning gains and achieved comparable final examination outcomes to English and Hybrid learners. These results suggest that providing local language options to learners is an effective way to make EdTech more accessible.
As AI tutors become increasingly capable of delivering rich, personalized feedback at scale, a key challenge remains: novice learners often struggle to process detailed explanations on their own. Structured reflection, grounded in decades of self-explanation research, is a theoretically compelling solution. By helping learners parse feedback and prompting them to actively interpret it, reflection activities are designed to reduce cognitive overload and deepen understanding. But does adding reflection to already rich AI-generated feedback actually help, or does it simply add friction? We tested this in a randomized experiment comparing Python practice with AI-generated, personalized feedback to "reflective practice," which paired identical feedback with structured self-explanation prompts. Contrary to our predictions, reflection never improved performance on any measure. Instead, it proved to be a temporal bottleneck: it doubled time spent on feedback and reduced practice volume by 40%, without making each learning opportunity more effective. Learners who cycled through more practice-and-feedback iterations outperformed reflective learners at the end of the session and maintained a small, nonsignificant advantage on transfer. Notably, reflection did not provide the scaffolding benefit we predicted for novices—and when individual differences did emerge, they favored higher-volume practice for more knowledgeable learners. Both practice conditions also substantially outperformed a high-quality video baseline (d = 0.66-0.93), replicating benefits of active practice with AI feedback over passive instruction. These findings suggest that when AI feedback is already elaborated and personalized, self-explanation activities may be a redundant time sink. As AI-generated feedback reaches learners at scale, these findings underscore the necessity of empirically validating pedagogical scaffolds—even those with strong theoretical support—before deploying them broadly.
Existing accounts of testing effects assume a dedicated study trial is necessary. Students must encode material before practice can strengthen memory. We tested this assumption by removing the study trial entirely. Across two experiments using prequestions and postquestions formats, in addition to clear effects of adding practice to a passage, a Practice with Feedback Only condition, in which participants answered questions about never studied material and received feedback, matched the accuracy of full instruction conditions that included a study passage. The study trial added time but not learning. These findings challenge the foundational assumption that a study trial is a necessary precondition for practice-based learning. We propose that the functional unit of learning is the generation of a response and its precise resolution through targeted feedback, not retrieval from a prior study event.
Generative artificial intelligence (AI) poses a significant threat to data integrity on crowdsourcing platforms, such as Prolific, which behavioral scientists widely rely on for data collection. Large language models (LLMs) allow users to generate fluent and relevant responses to open-ended questions, which can mask inattention and compromise experimental validity. To empirically estimate the prevalence of this behavior, we analyzed keystroke data from three studies ( N = 928) on Prolific between May and July 2025. Using an embedded JavaScript tool, we flagged participants who pasted text or whose keystroke count was anomalously low compared with their response length. For each flagged participant, we manually compared detected keystrokes with their final response to determine if the text could have been typed. This confirmed that despite deterrence measures, approximately 9% of participants submitted responses consistent with AI assistance or other forms of outsourced responding. These participants outperformed noncheaters (by up to 1.5 SD ), were more than twice as likely to share geolocations with other participants (suggesting possible proxy use), and exhibited lower internal consistency on questionnaire scales. Simulated power analyses indicate that this level of undetected cheating can diminish observed effect sizes by 10% and inflate required sample sizes by up to 30%. These findings highlight the urgent need for new detection methods, such as keystroke logging, which offers verifiable evidence of cheating that is difficult to obtain from manual review of LLM-generated text alone. As AI continues to evolve, maintaining data quality in crowdsourced research will require active monitoring, methodological adaptation, and communication between researchers and platforms.
The growing use of large language models by students poses challenges to the validity of open-ended responses in online learning and educational research. Keystroke tracking has shown promise for surfacing behavioral signals associated with outsourced responding, but existing implementations often require researchers to parse raw JSON logs with custom scripts, limiting adoption in educational settings. We present two open-source tools intended to support scalable screening and human review: (1) a lightweight JavaScript Software Development Kit (SDK) that integrates into web-based surveys and learning platforms with minimal setup, capturing keypress, paste, and copy events per question, and (2) a browser-based visualization dashboard for heuristic flagging, interactive keystroke timeline inspection, and one-click data export, requiring no programming expertise. Both tools process data entirely on the client-side during analysis, supporting privacy-preserving workflows. We report preliminary deployment evidence from an online calculus study, where the tools surfaced suspicious response-production patterns for manual inspection and enabled exploratory sensitivity analyses of downstream findings. Together, these tools lower the barrier to using behavioral process data to review response authenticity in open-ended learning activities, classroom assessments, and educational studies. The tools are available at https://keystroke-viz.theoaklab.org and GitHub https://github.com/the-oak-lab/aied26-keystroke-viz .
What conditions are necessary for students to learn from practice and feedback, without the need for upfront lecture? Across two experiments (N = 597), we examined how practice with feedback can support memory, generalization, metacognition, and motivation. Participants were randomly assigned to one of three instructional formats: a traditional lecture, practice with correct-answer feedback, or practice with explanatory feedback (predefined or adaptive and AI generated). In both studies, the lecture condition introduced linear regression through definitions and a worked example, while the practice conditions used matched problem sets with feedback that either (a) provided only correct answers or (b) explained why answers were correct. Study 1 used multiple-choice questions; Study 2 used open-ended questions with personalized explanatory feedback generated in real time by GPT-4o. For memory, both types of feedback outperformed lecture, suggesting that attempting a response and receiving feedback—even without explanations—enhances encoding. For generalization, however, feedback needed to include explanations, and learners needed sufficient prior knowledge to benefit. Study 2 also showed that practice—regardless of feedback type—improved metacognitive calibration compared to lecture, helping learners more accurately assess their understanding. While lecture produced greater situational interest for less-confident learners in Study 1, this pattern reversed in Study 2, where personalized, AI-generated feedback elicited higher interest for this group. Together, these findings clarify when and for whom practice with feedback can replace lecture-based instruction, and they highlight the potential of generative AI to scale personalized, explanatory feedback.
According to expectancy-value theories of motivation, individuals choose to pursue tasks that they expect to succeed at and find personally valuable. Historically, researchers have often suggested that these two factors interact to motivate behavior. However, expectancy × value interactions are rarely observed in empirical research and, when detected, they are often small in magnitude. Does this mean they can safely be ignored in models of motivation? In this paper we conduct two power analyses with simulated data to argue that expectancy × value interactions are likely far more important than a straightforward interpretation of effect sizes would suggest, and that downplaying them risks oversimplifying theory and recommendations for intervention. Specifically, Study 1 demonstrates that a realistic combination of three constraints (measurement error, skew, and correlation) can negatively bias expectancy × value interaction estimates by more than 50%. Study 2 shows that these interactions can create meaningful variability in motivation interventions and may contribute to a better understanding of treatment heterogeneity.
Expanding access to education in rural African communities remains difficult, largely due to limited internet connectivity. Mobile learning courses delivered via radio and offline mobile phones offer a promising, scalable solution. However, it is challenging to track student engagement in these environments due to the absence of tools that monitor students' interactions with the radio. In this study, we investigate the potential of ''Prize Codes'' -- codes read aloud during broadcasts that students enter via text message -- to serve as a real-time measure of student engagement with mobile-learning broadcasts. Using data from a 2024 implementation of Yiya AirScience, a mobile engineering course in Uganda, we evaluate the validity of Prize Codes as an engagement metric. Specifically, we test whether Prize Code measures (1) demonstrate reliability, with students who enter correct codes in one lesson being more likely to do so in subsequent lessons; (2) demonstrate convergent validity with existing measures of engagement; and (3) demonstrate predictive validity, predicting learning outcomes in the course. Our findings suggest that Prize Codes are a reliable and valid measure of engagement. Prize-Code accuracy demonstrates strong internal consistency (alpha = .97) and moderate test-retest reliability (ICC = .44). The measure aligns closely with synchronous participation (87% agreement, Cohen's kappa = .50), indicating it captures similar engagement patterns. Importantly, students who consistently enter correct Prize Codes perform significantly better on assessments, with Prize Code engagement predicting final exam scores above and beyond other engagement metrics. After establishing the measure's validity, we use it to (1) characterize patterns of engagement with Yiya broadcasts, (2) investigate early engagement with the broadcasts as a predictor of course persistence, and (3) replicate findings about the benefits of learning by doing. This study suggests that Prize Codes can be a feasible, scalable approach for tracking real-time engagement in resource-limited mobile learning settings at scale.
Decades of research show that tests, beyond assessing student knowledge, are powerful tools to promote learning, though they can also cause stress and disengagement. To utilize tests to encourage and motivate students, we implemented a mastery-based testing system in a large-enrollment general chemistry course (N = 234), allowing students to take three versions of each unit test, studying in between to increase their mastery of the content. Students took advantage of the mastery grading system when they struggled with unit tests, averaging six total repeated attempts. This level of repeated testing was associated with a 60% increase in students' use of study resources over the duration of the course, and a five-point overall increase in final exam scores (11 points for first-generation college students). This research suggests that mastery-based testing systems can leverage the benefits of test-enhanced learning while also providing motivation and structure to support students’ self-regulation.
Many college students drop science, technology, engineering, and math (STEM) majors after struggling in gateway courses, in part because these courses place large demands on students' time. In three online experiments with two different lessons (measures of central tendency and multiple regression), we identified a promising approach to increase the efficiency of STEM instruction. When we removed lectures and taught participants exclusively with practice and feedback, they learned at least 15% faster. However, our research also showed that this instructional strategy has the potential to undermine interest in course content for less confident students, who may be discouraged when challenged to solve problems without upfront instruction and learn from their mistakes. If researchers and educators can develop engaging and efficacy-building activities that replace lectures, STEM courses could become better learning environments.
When deciding how to study, students often choose suboptimal strategies. Our experiment investigated how preferences for instructional methods—video, practice, or both combined—affect learning outcomes. By collecting preferences both before and after instruction, we tested whether learners update their decisions once they have experiential data. We randomly assigned 130 participants to receive their preferred method (honoring initial choice) or a different method (dishonoring choice). Honoring preferences did not significantly influence recall or efficiency, so control over instructional methods may be less important than the methods themselves. Contrary to previous research showing preferences for lectures, most participants initially preferred approaches involving practice (35% practice-only; 50% combined). After instruction, preferences shifted further towards also including lectures as part of instruction (74% combined). Low post-test self-efficacy predicted changing to combined instruction, indicating learners with low confidence may overvalue comprehensive approaches, even though practice alone was equally effective at promoting recall and reduced instructional time by 66%. By measuring pre-learning and post-learning preferences, along with motivation, we were able to show for the first time that learners with lower confidence were more likely to change their preference to the most time-consuming option, revealing a miscalibration in post-learning judgments not seen before learning. These findings suggest that students rely on evidence from recent experiences, including both the instruction they received and their self-efficacy after the instruction, to make decisions about future study strategies.
The Doer Effect states that completing more active learning activities, like practice questions, is more strongly related to positive learning outcomes than passive learning activities, like reading, watching, or listening to course materials. Although broad, most evidence has emerged from practice with tutoring systems in Western, Industrialized, Rich, Educated, and Democratic (WEIRD) populations in North America and Europe. Does the Doer Effect generalize beyond WEIRD populations, where learners may practice in remote locales through different technologies? Through learning analytics, we provide evidence from N = 234 Ugandan students answering multiple-choice questions via phones and listening to lectures via community radio. Our findings support the hypothesis that active learning is more associated with learning outcomes than passive learning. We find this relationship is weaker for learners with higher prior educational attainment. Our findings motivate further study of the Doer Effect in diverse populations. We offer considerations for future research in designing and evaluating contextually relevant active and passive learning opportunities including leveraging familiar technology, increasing the number of practice opportunities, and aligning multiple data sources.
Many college students drop STEM majors after struggling in gateway courses, in part because these courses place large demands on students’ time. In three online experiments with two different lessons (measures of central tendency and multiple regression), we identified a promising approach to increase the efficiency of STEM instruction. When we removed lectures and taught participants exclusively with practice and feedback, they learned at least 15% faster. However, our research also showed that this instructional strategy has the potential to undermine interest in course content for less-confident students, who may be discouraged when challenged to solve problems without upfront instruction and learn from their mistakes. If researchers and educators can develop engaging and efficacy-building activities that replace lectures, STEM courses could become better learning environments.