In supervised machine learning (SML) research, large training datasets are essential for valid results. However, obtaining primary data in learning analytics (LA) is challenging. Data augmentation can address this by expanding and diversifying data, though its use in LA remains underexplored. This paper systematically compares data augmentation techniques and their impact on prediction performance in a typical LA task: prediction of academic outcomes. Augmentation is demonstrated on four SML models, which we successfully replicated from a previous LAK study based on AUC values. Among 21 augmentation techniques, SMOTE-ENN sampling performed the best, improving the average AUC by 0.01 and approximately halving the training time compared to the baseline models. In addition, we compared 99 combinations of chaining 21 techniques, and found minor, although statistically significant, improvements across models when adding noise to SMOTE-ENN (+0.014). Notably, some augmentation techniques significantly lowered predictive performance or increased performance fluctuation related to random chance. This paper’s contribution is twofold. Primarily, our empirical findings show that sampling techniques provide the most statistically reliable performance improvements for LA applications of SML, and are computationally more efficient than deep generation methods with complex hyperparameter settings. Second, the LA community may benefit from validating a recent study through independent replication.
There has been considerable research on confusion and frustration that has treated them as two unitary constructs, distinct from each other. In this article, we argue that there is instead a constellation of different types of confusion and frustration, with different antecedents, manifestations, and impacts, and that the commonalities between many types of confusion and frustration justify thinking of them as part of the same constellation of affect, distinct from other prominent affective categories. We discuss how these types of affect have been considered historically and in key models. We then discuss unusual manifestations of each form of affect that have been documented in the literature, and what light they shed on the broader constructs. We conclude with a discussion of a new theoretical framing that treats confusion and frustration as a confrustion constellation, and the opportunities and open questions that this perspective presents.
The promise of game-based learning relies on the assumption that games provide more engaging learning experiences than conventional instructional methods. The effects of playing a mobile health literacy game were compared with reading digital text material on situational interest, epistemic emotions, and satisfaction. A total of 251 Finnish high- and middle-school students were assigned one of two conditions: the mobile game or digital text condition. The mobile game condition played the Antidote COVID-19 game, while the digital text condition read the same health-related content. Situational interest increased significantly more in the mobile game condition than in the digital text condition. The game also induced more surprise, enjoyment, and confusion, and less boredom than the text material. Lastly, the mobile game condition showed higher satisfaction with the learning material. These findings suggest that game-based learning is more interesting and emotionally engaging than reading digital text, evoking stronger interest and epistemic emotions relevant to learning.
Productively engaging in SRL is challenging for learners since it involves coordinating multiple motivational, affective, cognitive, and metacognitive processes. Researchers have investigated methods to adaptively scaffold learners' productive engagement using SRL processes automatically captured by SRL detectors. However, most previous studies relied solely on the frequency of SRL processes to drive adaptive scaffolds (e.g., feedback, hints), possibly missing the sequential characteristics inherent to self-regulation, a crucial dimension of productive SRL. To address this gap, this study analysed the impact of sequential transitions between multiple SRL processes on learners' performance on a reading-writing task with a hypermedia environment called Flora. A sample of 66 secondary-school learners completed the task and trace data were collected. Grounded in the COPES model of SRL, a rule-based SRL detector was employed to capture SRL processes from collected trace data. We employed a method combining logistic regression with ordered network analysis (ONA) to analyse the transitions between the detected SRL processes. This exploratory study revealed several influential transitions to learners' performance in different temporal learning blocks of self-regulation. The implications suggest the potential of using COPES SRL process transitions to drive adaptive scaffolds to facilitate engagement in productive SRL, benefiting performance outcomes in hypermedia environments.
A student's ability to accurately evaluate the quality of their work holds significant implications for their self-regulated learning and problem-solving proficiency in introductory programming. A widespread cognitive bias that frequently impedes accurate self- assessment is overconfidence, which often stems from a misjudgment of contextual and task-related cues, including students' judgment of their peers' competencies. Little research has explored the role of overconfidence on novice programmers' ability to accurately monitor their own work in comparison to their peers' work and its impact on performance in introductory programming courses. The present study examined whether novice programmers exhibited a common cognitive bias called the "hard-easy effect", where students believe their work is better than their peers on easier tasks (overplace) but worse than their peers on harder tasks (underplace). Results showed a reversal of the hard-easy effect, where novices tended to overplace themselves on harder tasks, yet underplace themselves on easier ones. Remarkably, underplacers performed better on an exam compared to overplacers. These findings advance our understanding of relationships between the hard-easy effect, monitoring accuracy across multiple tasks, and grades within introductory programming. Implications of this study can be used to guide instructional decision making and design to improve novices' metacognitive awareness and performance in introductory programming courses.
In collaborative learning, emotions serve crucial functions that regulate task progress and social relationships among the team members. Although the research on the role of emotions in collaborative learning has been emerging, there is limited knowledge on how emotions emerge at different levels (individuals, teams) and its impact on collaborative learning, particularly when teams fail to complete a collaborative task. Therefore, the aim of this study is to explore how facial expressions of emotions unfold at the individual and the collective team level in a digital game-based learning setting in which collaborative teams eventually quit the collaborative task. A total of 30 high-school students' emotions were captured using facial expressions software and high-resolution video cameras from beginning to the quitting of the digital game on Biology (Min = 43 min; Max = 60 min). Before the collaborative learning task, participants were randomly assigned to work in a triad involving one of two experimental conditions: 1) a within-team competition (n = 5 teams) and a no within-team competition (n = 5 teams). Utilizing FaceReader, participants' momentary facial expressions of basic (i.e., happy, sad, anger, surprise, fear, and disgust) and neutral emotions were classified, and recurrence quantification analysis was utilized to study the recurrence of facial expressions within and between the members across the team conditions. Our findings indicate significant differences in the recurrence of facial expressions between the experimental conditions when they eventually quit the task. The current findings provide novel insights on the interplay between emotions and competitive dynamics in collaborative learning settings.
Game-based learning (GBL) environments are designed to foster emotional experiences conducive to learning; yet, there are mixed findings regarding their effectiveness. The inconsistent results may stem from challenges in measuring and modeling emotions as multi-dimensional constructs during GBL. Traditional approaches often use one data channel and conventional statistics to study emotions, which limit our understanding of the multicomponential interactions that underlie emotional states during GBL. In this study, we merged non-linear dynamical systems (NLDS) theory with the component process model of emotion to examine interactions and synchrony among two emotion signals during GBL, facial expressions and heart rate variability (HRV), and assessed its relation to knowledge and learning gain. Data were collected from 58 participants (n = 58) at a university in Central Finland while they learned about pathology with a tower defense game called Antidote COVID-19. Results showed a significant improvement in knowledge after GBL. A NLDS technique called cross-wavelet transformation showed there were varying degrees of synchrony between facial expressions and HRV. Neutral expressions showed the highest degree of synchrony with HRV, followed closely by happiness and anger with HRV. However, the synchrony between facial expressions and HRV did not affect knowledge and learning gain. This research contributes to the field by studying emotions as multidimensional systems during GLB.
When students procrastinate on programming assignments, it can hinder the quality of their code and negatively impact their grades. In contrast, when students actively delay working on assignments to prepare to code (e.g., reading or seeking help), it can be an effective self-regulated learning (SRL) strategy beneficial to programming performance. However, distinguishing active delay from procrastination is methodologically challenging. To address this, we tracked what students did when they behaviorally delayed starting an assignment. Most students prepared to code by using multiple course resources across programming assignments. We found that many students delayed starting to code by seeking help in the Q&A platform, and this was beneficial to the quality of their code. Also, some pre-coding activities were related to behavioral delay in starting to code, but benefitted students' grades, and thus may indicate active delay, but not all pre-coding activities were beneficial. By considering pre-coding activities, we gain a comprehensive view of students' approach to coding in CS education.
Programming courses can be challenging for first year university students, especially for those without prior coding experience. Students initially struggle with code syntax, but as more advanced topics are introduced across a semester, the difficulty in learning to program shifts to learning computational thinking (e.g., debugging strategies). This study examined the relationships between students' rate of programming errors and their grades on two exams. Using an online integrated development environment, data were collected from 280 students in a Java programming course. The course had two parts. The first focused on introductory procedural programming and culminated with exam 1, while the second part covered more complex topics and object-oriented programming and ended with exam 2. To measure students' programming abilities, 51095 code snapshots were collected from students while they completed assignments that were autograded based on unit tests. Compiler and runtime errors were extracted from the snapshots, and three measures -- Error Count, Error Quotient and Repeated Error Density -- were explored to identify the best measure explaining variability in exam grades. Models utilizing Error Quotient outperformed the models using the other two measures, in terms of the explained variability in grades and Bayesian Information Criterion. Compiler errors were significant predictors of exam 1 grades but not exam 2 grades; only runtime errors significantly predicted exam 2 grades. The findings indicate that leveraging Error Quotient with multiple error types (compiler and runtime) may be a better measure of students' introductory programming abilities, though still not explaining most of the observed variability.
Building adaptive game-based learning (GBL) interventions (e.g., immediate feedback) has been a recent effort to maximize learning effectiveness. State-of-the-art algorithms often overlook the concurrent emotions and motivation of individuals and its impact on GBL interventions. This pilot study utilized a 3 (feedback type: results, elaborative, attribution) x 2 (graph type: misleading, non-misleading) within-subjects design with MediaWatch, a GBL environment built to improve critical graph literacy. At an Austrian university, 41 students' concurrent emotions and motivation were measured using validated surveys immediately after different types of feedback on tasks during GBL. Results showed a significant improvement in graph literacy after GBL. Different types of feedback and task performance influenced concurrent emotions and interest, but individual differences accounted for the largest variability explained in emotions and interest. The findings suggest that within-subject variability is crucial for understanding concurrent emotions and motivation to feedback types and task performance during GBL.
Numerous studies aim to enhance learning in digital environments through emotionally-sensitive interventions. The D'Mello and Graesser (2012) model of affect dynamics hypothesizes that when a learner encounters confusion, the degree to which it is prolonged (and transitions into frustration) or resolved, significantly affects their learning outcomes in digital environments. However, studies yield inconclusive results regarding relations between confusion, frustration, and learning. More research is needed to explore how confusion and frustration manifest during learning and its relation to outcomes. We go beyond past work looking at the rate, duration, and transitions of confusion and frustration by treating these affective states as non-linear dynamical systems consisting of expressive and behavioral components. We examined the frequency and recurrence of facial expressions associated with basic emotions (as automatically labeled by AffDex, a standard tool for analyzing emotions with video data) during confused and frustrated states (as automatically labeled with BROMP-based detectors applied to students' interaction data). We compare these co-occurring patterns to learning outcomes (pre-tests, post-tests, and learning gains) within a digital learning environment, Betty ' s Brain. Results showed that the frequency and recurrence rate of basic emotions expressed during confusion and frustration are complex and remain incompletely understood. Specifically, we show that confusion and frustration have different relationships with learning outcomes, depending on which basic emotion expressions they co-occur with. Implications of this study open avenues for better understanding these emotions as complex and non-linear dynamical systems, in the long-term enabling personalized feedback and emotional support within digital learning environments that enhance learning outcomes.
Self-regulated learning (SRL) is important for computer science education. Yet, students often do not have SRL skills to benefit their learning. In this study, we examined 187 (n=187) students' SRL behaviors while they built programs with an automated feedback tool. Anchored in Winne and Hadwin's (1998) COPES model of SRL, our results showed that novices used more operators to debug compiler errors, while more experienced programmers used more operators to debug non-compiler errors. Finally, a random forest classifier showed that prior knowledge was the most important COPES feature predicting learning gain, followed closely by the student's perceived programming ability, use of evaluations with the automated feedback tool, and operators used to debug non-compiler errors on failed programs.
The study investigates the impact of cognitive biases on middle-school students' affective experiences while learning about math in a game-based learning environment (GBLE). The study focused on students' confrustion, an affect construct that unifies the several manifestations of confusion and frustration. We studied confrustion in the context of students self-explaining erroneous examples, where they had to find and fix common errors in given math problems and self- explain their problem-solving processes either with or without scaffolding. Text replays were utilized to examine student interactions during game-based learning and identify behaviors that emerged in response to cognitive biases and affect and its impact on learning and performance outcomes. The results revealed that students who demonstrated more pseudo-confidence in their self-explanations had higher self-reported self-efficacy, but were more likely to submit incongruent responses, exhibit confrustion, make errors, and take longer to finish the game. Overall, the findings show that students were vulnerable to cognitive biases and did not always respond in ways that accurately reflected their approach to solving math problems. The insights into how students approach and learn from math games inform the design and implementation of GBLEs by addressing cognitive biases.
Computing education researchers often study the impact of online help-seeking behaviors that occur across multiple online resources in isolation. Such separation fails to capture the interconnected nature of online help-seeking behaviors that occur across multiple online resources and its affect on course grades. This is particularly important for programming education, which arguably has more online resources to seek help from other people (e.g., computer-mediated conversations) than other majors. Using data from an introductory programming course (CS1) at a large US university, we found that students (n=301) sought help in multiple computer-mediated conversations, both Q&A forum and online office hours (OHQ), differently. Results showed the more prior knowledge about programming students had, the more they sought help in the Q&A compared to students with less prior knowledge. In general, higher-performing students sought help online in the Q&A more than the lower-performing groups on all the homework assignments, but not for the OHQ. By better understanding how students seek help online across multiple modalities of computer-mediated conversations and the relationship between help-seeking and grades, we can re-design online resources that best support all students in introductory programming courses at scale.
Accumulating evidence indicates that game-based learning is emotionally engaging. However, little is known about the nature of emotions in game-based learning. We extended previous game-based learning research by examining epistemic emotions and their relations to motivational constructs. One-hundred-thirty-one (n=131) 15–18-year-old students played the Antidote COVID-19 game for 25 minutes. Data were collected on their epistemic emotions, flow experience, situational interest, and satisfaction that were measured after the game-playing session. Learners reported significantly higher intensity levels of positive epistemic emotions (excitement, surprise, and curiosity) than negative ones (boredom, anxiety, frustration, and confusion). The co-occurrence network analyses provided new insights into the relationships between motivational and emotional states, where high-intensity flow experience, situational interest, and satisfaction co-occurred the most often with positive epistemic emotions. Results also revealed that a high-intensity flow can be experienced without high levels of situational interest in the topic. That is, gameplay can engage learners even though the learning topic does not interest them. This highlights the importance of intrinsically integrating the learning content with core game mechanics, ensuring the processing of the learning content. The study demonstrated that epistemic emotions, flow experience, satisfaction, and situational interest reveal different qualities of game-based learning. The results suggest that at least flow, situational interest, and epistemic emotions should be measured to understand different dimensions of engagement in game-based learning. Overall, the study advances prior research by clarifying relationships between epistemic emotions and motivational constructs.
This study examined 57 learners' emotions (i.e., joy, anger, confusion, frustration) as they engaged with scientific content while learning about microbiology with Crystal Island, a game-based learning environment (GBLE). Measures of learners' prior knowledge, in-game text comprehension, facial expressions of emotion, and posttest reading comprehension were collected to examine the relationship between emotions and single- and multiple-text comprehension. Analyses found that both discrete and non-discrete emotions were expressed during reading and answering in-game assessments of single-text comprehension. Learners expressed greater joy during reading and greater expressions of anger, confusion, and frustration during in-game assessments. Further results found that learners who expressed a high number of different emotions throughout reading and completing in-game assessments tended to have lower in-game comprehension scores whereas a higher number of different expressed emotions while completing in-game assessments was associated with greater posttest comprehension. Finally, while increased prior knowledge was associated with higher single- and multiple-text comprehension, there was no interaction between prior knowledge and emotions on multiple-text comprehension. Overall, this study found that (1) learners often express more than one emotion during GBLE activities, (2) emotions expressed while learning with a GBLE shift across different activities, and (3) emotions are related to demonstrated comprehension, but the type of activity influences this relationship. Results from this study provide implications for how emotions can be examined as learners engage in GBLE activities as well as the design of GBLEs to support learners' emotions accounting for different activity demands to increase comprehension of single and multiple texts.
Undergraduate students (N = 82) learned about microbiology with Crystal Island, a game-based learning environment (GBLE), which required participants to interact with instructional materials (i.e., books and research articles, non-player character [NPC] dialogue, posters) spread throughout the game. Participants were randomly assigned to one of two conditions: full agency, where they had complete control over their actions, and partial agency, where they were required to complete an ordered play-through of Crystal Island. As participants learned with Crystal Island, log-file and eye-tracking time series data were collected to pinpoint instances when participants interacted with instructional materials. Hierarchical linear growth models indicated relationships between eye gaze dwell time and (1) the type of representation a learner gathered information from (i.e., large sections of text, poster, or dialogue); (2) the ability of the learner to distinguish relevant from irrelevant information; (3) learning gains; and (4) agency. Auto-recurrence quantification analysis (aRQA) revealed the degree to which repetitive sequences of interactions with instructional material were random or predictable. Through hierarchical modeling, analyses suggested that greater dwell times and learning gains were associated with more predictable sequences of interaction with instructional materials. Results from hierarchical clustering found that participants with restricted agency and more recurrent action sequences had greater learning gains. Implications are provided for how learning unfolds over learners' time in game using a non-linear dynamical systems analysis and the extent to which it can be supported within GBLEs to design advanced learning technologies to scaffold self-regulation during game play.