Generating hints for incorrect code is a cognitively demanding task that fosters learning and metacognitive development. This study investigates three designs for personalized, scalable, and reflective hint-writing activities within a data science course: (i) writing a hint independently, (ii) writing a hint with on-demand AI assistance, and (iii) deferred AI assistance, in which students first write a hint independently and then revise it with the help of an AI-generated one. We examine how AI support can scaffold the learning process without diminishing students' productive cognitive effort. Through a randomized controlled experiment with graduate-level students (N=97), we found that deferring AI assistance leads to the highest-quality hints. Further, this design helps students identify a wide range of mistakes they otherwise struggle to identify without any AI assistance. Students valued these activities as opportunities to practice debugging and critically engage with AI outputs–skills that are now critical for learners to acquire as programming becomes increasingly automated and the use of AI for learning grows. Our findings also highlight key considerations for designing student-AI collaborative learning experiences to sustain student engagement, maintain appropriate cognitive load, and mitigate negative effects of AI, such as introducing redundancies and extraneous information into student work.
This paper explores the space of optimizing feedback mechanisms in complex domains such as data science, by combining two prevailing approaches: Artificial Intelligence (AI) and learnersourcing. Towards addressing the challenges posed by each approach, this work compares traditional learnersourcing with an AI-supported approach. We report on the results of a randomized controlled experiment conducted with 72 Master's level students in a data visualization course, comparing two conditions: students writing hints independently versus revising hints generated by GPT-4. The study aimed to evaluate the quality of learnersourced hints, examine the impact of student performance on hint quality, gauge learner preference for writing hints with versus without AI support, and explore the potential of the student-AI collaborative exercise in fostering critical thinking about LLMs. Based on our findings, we provide insights for designing learnersourcing activities leveraging AI support and optimizing students' learning as they interact with LLMs.
In this study, we designed a tool to investigate the relationship between students' ability to render accurate judgements of learning (JOLs) with decision-making behavior when annotating their own work and comparing it with perceived AI-generated annotations. Our findings suggest that students rarely adjust their JOLs after seeing the AI annotations, indicative of a strong self-confirmation bias. Trust in the AI tool was associated with a decreased likelihood of changing initial judgments, in part due to the similarity of AI annotations with their own. The process of using the tool to self-annotate was found to enhance performance on a post-test. Emphasizing clear learning objectives and being transparent with the limitations of AI functions may improve the effectiveness of such tools as a way to provide quick feedback and mitigate hesitancy towards wider-scale adoption.
While learning analytics researchers have been diligently integrating trace log data into their studies, learners’ achievement goals are still predominantly measured by self-reported surveys. This study investigated the properties of trace data and survey data as representations of achievement goals. Through the lens of goal complex theory, we generated achievement goal clusters using latent variable mixture modeling applied to each kind of data. Findings show significant misalignment between these two data sources. Self-reported goals stated before learning do not translate into goal-relevant behaviors tracked using trace data collected during learning activities. While learners generally articulate an orientation towards mastery learning in self-report surveys, behavioral trace data showed a higher incidence of less engaged learning activities. These findings call into question the utility of survey-based measures when up-to-date achievement goal data are needed. Our results advance methodological and theoretical understandings of achievement goals in the modern age of learning analytics.
Learning analytics defines itself with a focus on data from learners and learning environments, with corresponding goals of understanding and optimizing student learning. In this regard, learning analytics research, ideally, should be characterized by studies that make use of data from learners engaged in education systems, should measure student learning, and should make efforts to intervene and improve these learning environments. However, a common concern among members of the learning analytics research community is that these standards are not being met. In two analysis waves, we review a large and comprehensive sample of research articles from the proceedings of the three most recent Learning Analytics and Knowledge conferences, the premier conference venue for learning analytics research, and from articles published during the same time in the Journal of Learning Analytics (over the years of 2020, 2021, and 2022). We find that 37.4% of articles do not analyze data from learners in an education system, 71.1% do not include any measure of learning, and 89.0% of articles do not attempt to intervene in the learning environment. We contrast these findings with the stated definition of learning analytics and infer, like others before us, that scholarship in learning analytics research presently lacks clear direction toward its stated goals. We invite critical discussion of these findings from the learning analytics community, through open peer commentary.
Learning analytics defines itself with a focus on learner data, with corresponding goals of understanding and optimizing student learning. In this regard, learning analytics research, ideally, should be characterized by studies that make use of data from learners engaged in education systems, should measure student learning, and should make efforts to intervene and improve these learning environments. However, a common concern among members of the learning analytics research community is that these standards are not being met. In the current study, we review a large and comprehensive sample of research articles from the proceedings of two recent Learning Analytics and Knowledge conferences, the premier conference venue for learning analytics research. We find that 36.3% of articles do not analyze data from learners in a formal education system, 70.5% do not include any measure of learning, and 91.4% of articles do not attempt to intervene in the learning environment. We contrast these findings with the stated definition of learning analytics, and infer, like others, that scholarship in learning analytics research presently lacks clear direction toward its stated goals.
Considering that the central theme of learning analytics is data, there are uniquely strong benefits to sharing data, materials, and analysis code in this field. These open research practices facilitate transparency in research, and transparency can improve the quality of learning analytics methods, improve the reliability of findings, and improve the scope of impact of new developments. The Society for Learning Analytics research recently adopted a motion to advance on these goals. To inform this effort, we conducted a systematic review of articles published in the past two proceedings of the Learning Analytics and Knowledge conference, measuring the frequency of openly-shared data, materials, and analysis code. We find that 7.5% of articles made their data available, 13.7% made materials available, and 5.5% made analysis code available. We discuss these findings in the context of possible barriers to open practices, and suggest that the principal barrier to improved transparency is researcher’s reluctance to share, rather than privacy or legal constraints.
Question generation as a form of learnersourcing is both a metacognitive learning activity for students that encourages the development of higher-order thinking skills and a method for producing question banks and assessments. To better understand the motivations for learners who engage in learnersourcing and its impacts on student learning, we conducted an experiment that measured the effects of Multiple Choice Question (MCQ) generation in an introductory data science MOOC. We compared two approaches to question generation: (i) as a required activity, and (ii) as an optional activity. In both cases, the learnersourcing activity was part of the student summative evaluation. We found that learners value creating questions more, and create higher quality questions when they choose to do so compared to when it is required. At the same time there is a significant reduction in instructor evaluation workload in large-scale courses when learners engage by choice due to self-selection. Thus, we propose choice-based learnersourcing as a new form of scalable personalized learning design for MOOCs in particular. In addition, we contribute an exploration of the factors that influence learner choice to create (or not create) an MCQ, which can help contextualize the propensity of learners to engage in such learnersourcing activities.
Use of university students’ educational data for learning analytics has spurred a debate about whether and how to provide students with agency regarding data collection and use. A concern is that students opting out of learning analytics may skew predictive models, in particular if certain student populations disproportionately opt out and biases are unintentionally introduced into predictive models. We investigated university students’ propensity to consent to learning analytics through an email prompt, and collected respondents’ perceived benefits and privacy concerns regarding learning analytics in a subsequent online survey. In particular, we studied whether and why students’ consent propensity differs among student subpopulations bysending our email prompt to a sample of 4,000 students at our institution stratified by ethnicity and gender. 272 students interacted with the email, of which 119 also completed the survey. We identified that institutional trust, concerns with the amount of data collection versus perceived benefits, and comfort with instructors’ data use for learning engagement were key determinants in students’ decision to participate in learning analytics. We find that students identifying ethnically as Black were significantly less likely to respond and self-reported lower levels of institutional trust. Female students reported concerns with data collection but were also more comfortable with use of their data by instructors for learning engagement purposes. Students’ comments corroborate these findings and suggest that agency alone is insufficient; institutional leaders and instructors also play a large role in alleviating the issue of bias.
Privacy concerns may lead people to opt-in or opt-out of having their educational data collected. These decisions may impact the performance of educational predictive models. To understand this, we conducted a survey to determine the propensity of students to withhold or grant access to their data for the purposes of training predictive models. We simulated the effects of opt-out on the accuracy of educational predictive models by dropping a random sample of data over a range of increments, and then contextualize our findings using the survey results. We find that grade predictive models are fairly robust and that kappa scores do not decrease unless there is signiicant opt-out, but when there is, the deteriorating performance disproportionately affects certain subpopulations.
In this paper, we evaluate the complete undergraduate co-enrollment network over a decade of education at a large American public university. We provide descriptive and exploratory analyses of the network, demonstrating that the co-enrollment networks evaluated follow power-law degree distributions similar to many other large-scale networks; that they reveal strong performance-based assortativity; and that network-based features can improve GPA-based student performance predictors. We model the university-wide undergraduate co-enrollment network as an undirected graph, and implement multiple network-augmented approaches to student grade prediction, including an adaption of the structural modelling approach from (Getoor, 2005, Lu, 2003}. We compare the performance of this predictor to traditional methods used for grade prediction in undergraduate university courses, and demonstrate that a multi-view ensembling approach outperforms both prior ``flat'' and network-based models for grade prediction across several classification metrics. These findings demonstrate the usefulness of combining diverse approaches in models of student success, and demonstrate specific network-based modelling strategies that are likely to be most effective for grade prediction.