
Professional learning environments offer a unique opportunity for exploring professional competencies due to the use of workplace data, which captures important aspects of professional performance. This data is different from that used in formal learning analytics contexts, as it is not collected by systems designed to understand learning. Instead, workplace data represents the digital fingerprint professionals leave behind when interacting with technologies to do their jobs. Harnessing this data to support workplace learning is challenging for a range of reasons, including being able to identify meaningful metrics to identify individual performance and scaffold the use of this data to support professional learning interventions on knowledge and performance. Generative AI (GenAI) has great potential to enhance professional learning analytics by addressing some of the challenges inherent in using workplace data. These challenges include needing to process large amounts of unstructured data to understand individual performance and transform this data into interfaces and interventions that can support learning. In our paper, we modify Clow’s learning analytics cycle to inform a modified framework describing the intersection of learning analytics with professional learning. Subsequently, we illustrate the potential power of GenAI for supporting professional learning across the framework through three case studies in health professions education.
Online discussion forums are a key data source in STEM education for analyzing student interaction, knowledge construction, and social learning. From a learning analytics perspective, understanding the temporal dynamics of instructional behaviours, for example, the differences between instructors and peers, can offer valuable insights into how instructional support unfolds in online discussions. However, few studies have examined these interactions with attention to their sequential and temporal organization. This study adopts a learning analytics approach to investigate the behavioural dynamics of instructional support in a large-scale online math discussion forum. We analyzed 83,569 posts from Math Nation, a widely used online learning platform supporting K–12 math education in the United States. Instructional behaviours were automatically coded based on a theoretically grounded coding scheme using transformer-based language models, and their temporal patterns were modelled using multilevel vector autoregression. Network visualizations were employed to illustrate the dynamic relationships among instructional strategies. Our findings show that peers posted more frequently than instructors and were more likely to engage in acknowledgement and feedback behaviours. When examining behaviour proportions and temporal transitions within discussion threads, peer interactions were associated with more diverse behavioural sequences, whereas instructor interactions exhibited more focused and directive patterns. In mixed-participant threads, instructors often responded following peer contributions, which was associated with a reduced likelihood of subsequent direct peer interventions. Rather than making claims about instructional effectiveness or learning outcomes, this work contributes to the field of learning analytics by demonstrating how natural language processing (NLP)-based behaviour modelling can be integrated with temporal interaction analysis to characterize role-based instructional support dynamics in asynchronous, discussion-based STEM learning environments.
Assessment is foundational to learning analytics, especially in evaluating instructional interventions and guiding improvement in online learning environments. With the growing use of large language models (LLMs) to score open-ended responses, questions arise about the reliability of these model-generated scores, particularly in short pre-post formats where learners are expected to improve. This study introduces a novel method for estimating test reliability that adjusts for learning gains using a Rasch-based split-half approach. We validated this approach through simulation under realistic conditions of missing data and score change, showing tangible improvements in reliability estimation compared to baseline methods. Applying this method to a dataset of 985 tutors completing 12 online lessons, we find that GPT-4-based scoring achieves satisfactory reliability, with open-ended responses (0.733) outperforming multiple-choice items (0.652). Both item types jointly yielded the highest reliability (0.774). Hence, as few as 14 open-ended items (across an average of 3-4 completed lessons) were sufficient to surpass common reliability thresholds of 0.7 or higher. Principal component analysis revealed a skill structure with a strong primary dimension shared across almost all lessons and interpretable subdimensions—socio-emotional, cognitive, and fairness-related tutoring skills—supporting a bifactor-like model. These findings demonstrate that GPT-4 and similar LLMs can be effectively used for formative assessment of complex instructional skills in online and personalized learning contexts, provided their reliability is empirically verified. This study contributes an open-source, learning-aware framework for scalable and reliable AI-supported assessment in learning analytics contexts.
This study expands the concept of curricular complexity as defined in the Curricular Analytics framework (Heileman et al., 2018) by providing qualitative evidence of the suitability of three new metrics, concerning timing of course offerings, extended time-to-degree, and credit loss, that more adequately address curricular challenges encountered by transfer students. Curricular Analytics is a method for analyzing a curriculum that enables practitioners and researchers to quantify and systematically analyze the impacts of course sequencing in a plan of study on student outcomes. However, the original conceptualization falls short of capturing the substantive challenges faced by transfer students who enter an undergraduate program at various points in the curricular sequence. This study was guided by the following research question: “How do three new measures of curricular complexity (i.e., inflexibility factor, transfer delay factor, and credit loss) align with transfer professionals’ perceptions of curricular barriers for transfer students?” Using a grounded theory approach, we conducted seven focus groups with 38 transfer professionals across the United States. We presented these transfer experts with each new measure and prompted them to reflect on its validity based on their experiences supporting transfer students. We found transfer professionals resonated strongly with all three new metrics, suggesting strong initial construct and content validity.
The growing emphasis on 21st-century competencies in postsecondary education, intensified by the transformative impact of generative artificial intelligence (GenAI) on the economy and society, underscores the urgent need to evaluate how they are embedded in curricula and how effectively academic programs align with evolving workforce and societal demands. Curricular analytics, particularly recent advancements powered by GenAI, offer a promising data-driven approach to this challenge. However, the analysis of 21st-century competencies requires pedagogical reasoning beyond surface-level information retrieval, and the capabilities of large language models (LLMs) in this context remain underexplored. In this study, we extend prior research on curricular analytics of 21st-century competencies across a broader range of curriculum documents, competency frameworks, and models. Using 7,600 manually annotated curriculum-competency alignment scores (38 competencies and 200 courses across five curriculum document types), we evaluate the informativeness of different curriculum document sources, benchmark the performance of general-purpose LLMs on mapping curricula to competencies, and analyze error patterns. We further introduce a reasoning-based prompting strategy, curricular chain-of-thought (CoT), to strengthen LLMs’ pedagogical reasoning. Our results show that detailed instructional activity descriptions are the most informative type of curriculum document for competency analytics. Open-weight LLMs achieve accuracy comparable to proprietary models on coarse-grained tasks, demonstrating their scalability and cost-effectiveness for institutional use. However, no model reaches human-level precision in fine-grained pedagogical reasoning. Our proposed curricular CoT yields modest improvements by reducing bias in instructional keyword inference and improving the detection of nuanced pedagogical evidence in long text. Together, these findings highlight the untapped potential of institutional curriculum documents and provide an empirical foundation for advancing AI-driven curricular analytics.
The growing emphasis on competency-based education (CBE) has heightened the need for clearly defined metrics and robust assessment frameworks to evaluate 21st-century competencies. Curriculum analytics (CA) provides a promising avenue for assessing learning outcomes (LOs) and informing continuous improvement in higher education. However, challenges persist in differentiating academic performance from actual LO development and in translating assessment data into meaningful program-level actions. This study examines how CA tools support the direct assessment of LOs and contribute to continuous improvement processes in higher education. Using a twocase study design, we analyzed CA implementation in two universities through interviews, cognitive walkthroughs, and institutional document analysis. Data triangulation identified 18 themes, nine of which reached full consensus among the three researchers. Findings indicate that CA tools effectively support the assessment of LOs aligned with 21st-century competencies by generating actionable insights that guide faculty toward more authentic and reflective teaching practices. The study contributes to the LA field by providing empirical evidence of how CA tools can bridge assessment and pedagogical improvement, offering both theoretical and practical implications for researchers and practitioners.
Computational thinking (CT) is a vital skill set for pre-service teachers who will need to foster computational literacy in K-12 classrooms, yet the factors influencing their CT skills remain less understood than those for K-12 students or in-service teachers. This study leverages multimodal data to investigate how pre-service teachers (n=128) differ in CT skills, the predictive role of metacognitive strategies and prior coding experience, and variations in online behaviours. Using latent profile analysis, we identified three profiles based on digital literacy, problem-solving, and coding comfort (Novice, Developing, and Proficient), revealing heterogeneity in CT, and supporting non-linear skill acquisition. Linear discriminant analysis revealed that metacognitive strategies and prior coding experience significantly predict profile membership, validating the interplay of technical and cognitive factors in the development of CT skills. Behavioural data from an interactive problem-solving task showed that, compared to Novices and Developing learners, Proficient learners were more task efficient and perceived fewer challenges during task completion. Implications for designing a learning analytics dashboard to visualize profiles and behavioural metrics to support adaptive, equitable, and personalized teacher training are discussed, thereby enhancing pre-service teachers' readiness to integrate CT into K-12 education.
Online learning platforms have expanded access to education but also raise concerns about biased content, particularly in text-based learning materials such as textbooks, lesson plans, and course excerpts. Such biases can perpetuate discrimination, can harm student outcomes, and can often be difficult to detect, as identification typically relies on time-consuming human review. Learning analytics (LA) can enhance this process by supporting human reviewers through automated detection, offering a scalable solution while retaining human judgment for nuanced evaluations. Accordingly, this LA study explores two research questions: RQ1: Which features might support the identification of ethnic bias in text-based online learning materials? and RQ2: Which classification approaches might be suitable for identifying ethnic bias in text-based online learning materials? First, we identified features signalling potential ethnic bias (presence or absence) in textual content using a dataset (N = 345) labelled by 193 students from diverse ethnic backgrounds. Then, we evaluated multiple machine learning (ML) models for their effectiveness in bias classification. The results suggest significant correlations between perceived bias and content from social sciences. Additionally, through bootstrap analysis, support vector machines and random forest classifiers showed consistent performance in bias identification (with F1-scores of 0.71 and 0.70 on the test set, respectively). In contrast, the naive Bayes (NB) model demonstrated the highest precision (0.75 on the test set). We discuss these findings and their implications for LA, emphasizing the importance of quality and inclusive educational tools. As an initial step toward automated bias classification in education, this study provides a foundation for spotting ethnic bias in learning content, supporting fairer technologies for more inclusive learning environments. Notes for Research and Practice center dot There are statistically significant correlations between ethnic bias and social sciences content. center dot Random forest (RF) and stacking (STK) classifier models were more reliable for ethnic bias classification. center dot The naive Bayes (NB) model is recommended in scenarios that prioritize precision in bias detection. center dot An extended labelled dataset is made available for promoting fairer artificial intelligence (AI) applications. center dot A baseline approach for more inclusive learning analytics (LA) in online environments is provided.
Student workload analysis has the potential to play a crucial role in providing both actionable insights to inform course design and curricular adjustments that promote student learning and well-being. While numerous studies have emphasized the need for analyzing workload beyond single-value metrics, such as credit hours, the interpretation and practical application of these metrics for educational interventions remains unclear. In this study, we explore the interplay between time-on-task measurements with student-perceived learning and difficulty. We move beyond average indicators of time-on-task by proposing and examining various metrics related to the dynamics of workload over time. Across 14 engineering courses taught at Pontificia Universidad Cat olica de Chile, we analyze three different sources of data: (1) self-reported time-on-task and perceived difficulty obtained through a weekly timesheet survey, (2) interactions with the learning management system (LMS), and (3) perceived learning attainment obtained from the course evaluation survey. Our results show that LMS-based and self-reported time-on-task were highly correlated. Also, workload dynamics metrics, such as the presence of workload peaks, were highly correlated with perceived learning and perceived difficulty. As such, this study provides evidence in support of considering workload dynamics, rather than average measures of time-on-task, to predict variables related to student learning. The metrics proposed by this framework could be used to implement practical tools for educators and administrators willing to optimize course design and improve learning attainment.
This study explores how students across Grades 8 to 12 engage with mathematical functions in creative, visual ways through function art-an innovative STEAM-based educational approach. Grounded in the Trends in International Mathematics and Science Study (TIMSS) framework and employing a Design-Based Research methodology, the project involved 400 students from the Philippines who created digital artworks using GeoGebra. To uncover learner profiles, a person-centred clustering method-hierarchical clustering on principal components- was applied to variables representing the number and types of functions used. The results revealed three distinct student profiles: Repetitivists (high function quantity, low diversity), Simplists (low quantity and diversity), and Multifunctionists (high diversity, low quantity). Further analysis showed meaningful associations between cluster membership, grade level, and function strategies. Qualitative evaluation using TIMSS cognitive domains-Knowing, Applying, and Reasoning-highlighted that students' use of mathematical strategies and precision in transformations varied widely, often independently of the quantity or diversity of functions used. These findings suggest that function art, when analyzed through learning analytics, provides a rich lens for understanding students' mathematical thinking and offers valuable insights for tailoring interdisciplinary instruction in STEAM education.
The rise of generative artificial intelligence (GenAI) and accelerated globalization have necessitated a fundamental recalibration of higher education to prioritize domain-agnostic, 21st-century professional competencies. While institutional commitment to these skills is high, their systematic integration into the curriculum and evaluation remains fragmented, highlighting a critical gap between traditional academic success metrics and demonstrated workforce readiness. This special issue presents five complementary studies that investigate how the intersection of learning analytics (LA) and GenAI can bridge the gap between institutional rhetoric and demonstrated professional readiness. The contributions collectively advance a research agenda across four dimensions: 1) benchmarking large language models (LLMs) for curricular-competency alignment using reasoning-based prompting, 2) the iterative design of Socratic-style GenAI chatbots to scaffold self-regulated learning, 3) the application of psychometric modelling and Latent Profile Analysis to quantify 21st-century professional competencies, and 4) institutional governance and adoption of curriculum analytics. Collectively, these studies advocate for an epistemological shift toward processsensitive assessments that move beyond static, episodic indicators toward dynamic, longitudinal representations of learner capability. We conclude by outlining the sociotechnical infrastructure, including robust governance and interdisciplinary collaboration, required to responsibly transition these AI-driven innovations from research prototypes to sustainable enterprise infrastructure, ensuring that analytics serve the evolving needs of students, educators, and professional bodies.
It is widely recognized that higher education (HE) graduates require a broad range of professional skills and abilities to succeed in their future careers. However, despite this acknowledgement, assessment practices in HE remain focused on content-based knowledge. This narrow emphasis limits the capacity to effectively and holistically evaluate a student's professional competency and readiness for employment. This issue is particularly acute for HE degrees that require graduates to demonstrate attainment of externally regulated professional standards. While the curricula are mapped to professional standards for accreditation purposes, demonstrating a student's attainment of these standards is not straightforward and has mostly been done through self-reported surveys. This study offers a novel curriculum analytics method for mapping assessment grades to the attainment of professional standards across a Teacher Education program. Specifically, we present an approach that uses psychometric modelling and learning analytics to identify distinct patterns in learners' acquisition of professional standards. This method does not alter current assessment practices in HE. Instead, the approach offers a scalable, automated means to infer a learner's attainment of documented professional standards, complementing current measures of academic success, such as GPA. The study underscores the advantages of complementing the current HE assessment practises with an outlined curriculum analytics approach, providing a holistic representation of a student's learning progress.
Generative AI (GAI) is increasingly integrated into education, particularly through chatbots that support students without direct human intervention. While these tools show promise as personalized learning companions, concerns persist about their potential to foster overreliance, limit creativity, and hinder the development of critical thinking. These risks highlight the need to strengthen students' metacognitive skills and promote structured self-reflection. This paper presents the design and iterative development of a GAI-based chatbot aimed at scaffolding self-regulated learning. Through three design cycles involving 276 students in 10 courses, the chatbot evolved from a static assistant into a dynamic, course-integrated tool capable of supporting personalized, Socratic-style dialogue. Thematic analysis of diverse qualitative data sources revealed that students seek scaffolded support, such as human tutoring, and require explicit guidance to engage meaningfully with AI. Findings emphasize the importance of dialogic competence, personalization, and educator involvement in shaping effective AI-mediated reflection. This study underscores the need for pedagogically grounded AI tools that position chatbots as collaborative agents, complementing rather than replacing the roles of teachers and learners. It advocates for reflective teaching practices that clearly define the responsibilities of students, educators, and AI systems to ensure that GAI enhances deep learning and independent thought.
This study aims to develop a set of heuristics tailored for evaluating learning analytics in simulation-based professional learning, focusing on the following research questions: (1) What heuristics are appropriate for evaluating learning analytics in simulation-based professional learning contexts? (2) How can theoretical frameworks and empirical findings be combined in the development of such heuristics? (3) How can expert evaluation inform their refinement and applicability? The study combines a top-down approach, drawing on a theoretical framework for learning experience design, with a bottom-up analysis of empirical findings from prior studies in the context of a design project. An initial set of heuristics was iteratively reviewed and refined in collaboration with experts in user and learning experience design. The outcome is a detailed heuristic framework that supports the evaluation of learning analytics in simulation-based settings and accounts for the technological, pedagogical, and social dimensions of professional learning.
This study analyzed data from prototyping sessions to support learners' engagement in collaborative problem-solving through design-mode thinking. Knowledge creation (KC) is essential for addressing 21st-century complex problems, yet learners struggle to share incomplete ideas due to unfamiliarity with design-mode thinking. Although prototyping is known to be effective for sharing and improving incomplete ideas in the business sector, few studies analyze design activities during prototyping from a KC perspective across different design disciplines. This study examined three teams (engineering, product design, service design) during 30-minute prototyping sessions. Using a novel combination of techniques-temporal socio-semantic network analysis (tSSNA) and ordered network analysis (ONA)-we analyzed teams' shared epistemic agency, segmenting activities into three phases based on idea improvement patterns. Engineers focused on generative collaborative actions, product designers emphasized creating shared understanding, and service designers concentrated on alleviating lack of knowledge. Statistical tests revealed significant differences between teams and phases. The findings suggest three design principles for KC practice: attending to disciplinary differences in interdisciplinary teams, providing timely educator support for concept creation, and using short task durations to encourage sharing incomplete ideas. Furthermore, this work demonstrates the potential of combining tSSNA and ONA for analyzing collaborative KC processes.
Learning analytics has the potential to enhance education through data-informed decision-making, but persistent challenges around generalizability and scalability continue to limit its real-world impact. In this paper, we introduce the concept of a modelable world: a learning ecosystem purposefully designed to support the development of predictive models that generalize across diverse contexts. We outline three core design principles of modelability: (1) valid and interpretable measurements, (2) scalable and stable implementation, and (3) a collaborative research-practice-technology ecosystem. We then illustrate how these principles can be operationalized in the real world through a case study of CourseKata, a platform offering a fully instrumented online textbook adopted across a wide range of institutions and disciplines. Using CourseKata data, we developed early prediction models of students' final course grades using behavioral measures and tested the model generalizability across institutions (something rarely done in the modeling literature). Results show that a system designed with modelability in mind can produce predictive models that generalize effectively across diverse educational contexts.
Documenting and understanding self-regulated learning (SRL) processes can inform the design of learning activities and scaffolds to enhance student success in STEM. Think-aloud protocols (TAPs)—prompting students to verbalize thoughts during task performance—reveal real-time, ecologically sensitive verbalizations of students’ SRL processes such as planning, monitoring, and strategy use. However, coding TAP data requires substantial resources. We investigated the capabilities of Large Language Models (LLMs) in automating the coding of SRL processes in TAPs across undergraduate STEM courses. To examine how task features, prompt engineering strategies, and SRL codes are linked to LLMs’ coding accuracy, we used a factorial design comparing different LLMs (GPT-4o, Claude 3.5 Sonnet, and Gemini 1.5 Pro) across six SRL TAP codes (representing distinct cognitive and metacognitive processes), six prompt conditions (varying few-shots and context levels), and two STEM tasks (mathematics and biology). Analysis of 600 student verbalizations revealed that mathematics tasks yielded significantly higher classification accuracy compared to biology tasks, with GPT-4o and Claude 3.5 Sonnet outperforming Gemini 1.5 Pro. Few-shot prompting showed code-specific effects, with descriptively higher accuracy for monitoring negative judgments of learning but significantly decreased accuracy for subgoal setting. The addition of contextual information showed minimal impact across tasks. Cognitive process TAP codes (e.g., mathematical problem-solving) demonstrated the most consistent cross-task classification accuracy. In contrast, metacognitive monitoring (e.g., judgment of learning) showed substantial task-dependent variations. These findings highlighted both the promise and limitations of LLMs for scaling SRL research. They suggest that theoretical alignment in prompt engineering is essential for effective automated coding of dynamic regulatory processes.
The growing emphasis on 21st-century competencies in postsecondary education, intensified by the transformative impact of generative artificial intelligence (GenAI) on the economy and society, underscores the urgent need to evaluate how they are embedded in curricula and how effectively academic programs align with evolving workforce and societal demands. Curricular analytics, particularly recent advancements powered by GenAI, offer a promising data-driven approach to this challenge. However, the analysis of 21st-century competencies requires pedagogical reasoning beyond surface-level information retrieval, and the capabilities of large language models (LLMs) in this context remain underexplored. In this study, we extend prior research on curricular analytics of 21st-century competencies across a broader range of curriculum documents, competency frameworks, and models. Using 7,600 manually annotated curriculum-competency alignment scores (38 competencies and 200 courses across five curriculum document types), we evaluate the informativeness of different curriculum document sources, benchmark the performance of general-purpose LLMs on mapping curricula to competencies, and analyze error patterns. We further introduce a reasoning-based prompting strategy, curricular chain-of-thought (CoT), to strengthen LLMs' pedagogical reasoning. Our results show that detailed instructional activity descriptions are the most informative type of curriculum document for competency analytics. Open-weight LLMs achieve accuracy comparable to proprietary models on coarse-grained tasks, demonstrating their scalability and cost-effectiveness for institutional use. However, no model reaches human-level precision in fine-grained pedagogical reasoning. Our proposed curricular CoT yields modest improvements by reducing bias in instructional keyword inference and improving the detection of nuanced pedagogical evidence in long text. Together, these findings highlight the untapped potential of institutional curriculum documents and provide an empirical foundation for advancing AI-driven curricular analytics.
The aim of learning analytics (LA) is to turn educational data into insights, decisions, and actions to improve learning and teaching. The reasoning of the provided insights, decisions, and actions is often not transparent to the end-user, and this can lead to trust and acceptance issues when interventions, feedback, and recommendations fail. In this paper, we shed light on achieving transparent LA by following a transparency through exploration approach. To this end, we present the design, implementation, and evaluation details of the Indicator Editor, which aims to support self-service LA (SSLA) by empowering end-users to take control of the indicator implementation process. We systematically designed and implemented the Indicator Editor through an iterative human-centred design (HCD) approach. Further, we conducted a qualitative user study (n = 15) to investigate the impact of following an SSLA approach on users' perceptions of and interactions with the Indicator Editor. Our study showed qualitative evidence that supporting user interaction and providing user control in the indicator implementation process can have positive effects on different crucial aspects of LA, namely transparency, trust, satisfaction, and acceptance.
Effective learning design (LD) grounded in sound pedagogy is a critical driver of student success. Therefore, it is important to explore how LD of online learning environments influences student ability to manage their own learning. This understanding can inform the development of online programs that prioritize student-driven learning. Research increasingly shows that students can monitor and regulate their own motivation, significantly impacting their academic achievement. Based on the concept of motivational regulation (MR) as a context-dependent process, this systematic review aims to identify existing learning analytics research that examines the link between MR and LD as a key contextual factor. The findings reveal that: 1) there is a lack of consistency in how motivation is measured, operationalized, and applied within the learning analytics literature; 2) self-reporting through surveys remains the most common approach for measuring and operationalizing MR; 3) LD is primarily operationalized at the session and learning activity level, with descriptions focusing on pedagogical principles and strategies; and 4) most studies address the relation between motivational constructs and persistence or academic achievements by looking at variables such as performance or/and outcomes rather than the processes that influence motivational changes and the relationship between MR and LD.