Effective communication is crucial for agile software development teams, significantly influencing collaboration, decision-making quality, and productivity. Despite growing awareness around gender equity, challenges remain in ensuring equal participation during team interactions. This study analyzes communication in agile teams from a gender perspective, employing a Multimodal Learning Analytics approach to identify interaction patterns. Advanced natural language processing and epistemic network analysis were utilized to explore differences in communication behaviors between male and female team members during effort estimation activities. Findings revealed that women concentrated their contributions in specific categories—clarification and understanding, estimation and consensus, and confirmation and acceptance—suggesting structured, reflective decision-making strategies. Men exhibited a more evenly distributed communication style across various categories. Additionally, dependencies and risks were central to women’s epistemic networks, highlighting gender-based differences in how they managed uncertainty. Psychological safety significantly shaped these communication patterns; lower perceived safety among women resulted in more structured interactions, while equitable participation contexts facilitated balanced communication. These results highlight the importance of fostering psychological safety and inclusive communication in agile teams. Future research should examine intersectional aspects, such as cultural background or experience level, to deepen understanding of gender dynamics and enhance agile practices.
Introduction Programming courses continue to grapple with several pedagogical challenges, such as class size, variety of student prior experiences, and the necessity of timely, individualized feedback. Although Intelligent Tutoring Systems (ITS) have been targeted to solve them for some time, the rise of Generative Artificial Intelligence (GenAI) instruments prospective for transformation. But advanced systems such as these need to be designed based on educational stakeholders' expectations. In this paper we introduce a modular architecture of a GenAI-driven ITS that caters to the needs of both students and teachers for introductory programming. To validate and refine this framework we conducted a mixed-methods survey of programming teachers. The results indicate a high consensus regarding the desired features: teachers place valued emphasis in the ability of an ITS to auto grade questions, track learners’ progress, and provide deep analytics of most common mistakes and engagement patterns. For students, the focus is on immediate, explainer feedback and guided hints. Specifically, concerns from instructors were identified as being incredibly substantial regarding GenAI in terms of student overuse and the possible “hallucinations” of AI-generated misconceptions. These results can be directly related to our proposed approach, highlighting the need of a separate teacher's interface for monitoring, and a strong pedagogical module, that enables to steer the GenAI to generate hints and not solutions. This teacher driven approach guarantees not only that the artifact is technologically advanced, but also pedagogically fit and meets an authentic educational need.
Predictive Learning Analytics (PLA) systems are increasingly deployed in public education to support dropout prevention, raising concerns about algorithmic bias in contexts of structural inequality. This paper presents a large-scale case study of a PLA system in Brazilian public secondary education, evaluating predictive performance, probabilistic calibration, and fairness across protected groups (race/ethnicity, sex, and socioeconomic vulnerability). Random Forest models were trained at two temporal moments per school year across three grade levels (six models total), using administrative records from over 100,000 student enrollments. Calibration analysis stratified by protected groups revealed that bias manifests primarily through differences in probabilistic reliability rather than aggregate metrics. Post-processing via Platt Scaling reduced these deviations without introducing new disparities across groups. These findings highlight that fairness in PLA is a dynamic, temporally sensitive property, and that calibration-aware evaluation is essential when risk scores function as decision-support signals in institutional governance.
As the demand for data-driven strategies in education grows, learning analytics dashboards have proven to be essential tools for enhancing transparency, monitoring, and decision-making at various administrative levels. This study examines the acceptance of a prototype control panel developed to provide relevant information to municipal education departments. Utilizing the Technology Acceptance Model (TAM), the study evaluates dashboard experience satisfaction, perceived ease of use, and behavioral intention to use the implemented dashboards. Data were collected through a survey that included secretaries of education, technical agents, administrative directors, pedagogical coordinators, school principals, and secretaries, after testing and analyzing a functional prototype. The results indicate a high level of satisfaction with the dashboards during phase 3 of the experiment, with an overall dashboard experience satisfaction score of 5.13 on a 7-point scale and a perceived ease of use score of 4.89. Additionally, the behavioral intention to use the dashboards was significant, with a score of 4.89. These findings indicate that participants viewed the dashboards positively as support tools for municipal educational management. The study discusses the practical implications of these results.
School absenteeism undermines learning and increases the risk of dropout, especially among low-income students. This study uses clustering techniques to analyze attendance data from the Sistema Gestão Presente (SGP), covering 1.4 million high school records from four Brazilian states. We combined absenteeism profiles with indicators of physical and pedagogical infrastructure, including transportation, internet access, accessibility, and inclusion resources. Applying K-Means separately to urban and rural schools revealed five profiles in each context. Results indicate that stronger infrastructure correlates with lower absenteeism, but highlight persistent weaknesses in pedagogical inclusion, particularly in support for students with disabilities and in the provision of adapted materials.
Artificial Intelligence in Education (AIED) has advanced significantly over the past four decades, generating powerful tools for personalization, assessment, learning analytics, and predictive modeling. However, the dominant design assumptions underlying many AIED systems, stable connectivity, continuous electricity supply, and access to cloud-based infrastructure, reflect conditions typical of high-income contexts and risk reinforcing existing global inequalities. In large parts of the Global South, structural constraints such as limited internet access, unreliable energy provision, and scarce digital resources restrict the feasibility of mainstream AIED solutions, producing what can be described as a double digital exclusion: first from basic connectivity, and subsequently from AI-driven educational innovation. In response to this challenge, this paper advances the concept of Artificial Intelligence in Education Unplugged (AIED Unplugged) as a distinct sociotechnical paradigm. We argue that it should not be understood merely as a technical adaptation, such as deploying Edge AI or lightweight models, but as a broader pedagogical, institutional, and political reorientation of AI design toward equity, contextual adequacy, and sustainability. Through a conceptual and analytical examination, we clarify the core principles of AIED Unplugged, differentiate it from related approaches, discuss implementation requirements and constraints, and explore its implications for digital sovereignty and public policy in the Global South.
This paper investigates the development of predictive models to identify students at academic risk in higher education, even in scenarios with incomplete historical data. Using real data provided by Monterrey Institute of Technology and Higher Education, we simulated four distinct scenarios with varying levels of information availability, reflecting real-world situations in the educational context. The methodology adopted included exploratory analysis for feature selection and engineering, sampling techniques for class balancing, and the application of several machine learning classifiers with default settings. The database used contains academic and sociodemographic information of 1.796 unique students. Eighty model combinations were evaluated, using the AUC-ROC metric as the main performance indicator. The four scenarios were designed to progressively incorporate different categories of student information, ranging from demographics and admissions data to engagement and academic performance metrics, and allow us to analyze the impact of information completeness on model performance. The results indicate that even with limited data, it is possible to achieve competitive predictive performance (AUC ≈ 0.87), whereas scenarios with complete academic records achieve AUCs of up to 0.96. Variables such as current academic performance, admissions test scores, study load, and demographic attributes were consistently influential. These findings reinforce the possibility of early prediction of academic risk and provide a flexible framework for adapting predictive strategies based on available data.
In Brazil, undergraduate curricula are required to include at least 10
Early identification of students at academic risk is a central challenge for large and decentralized educational systems. In countries such as Brazil, pronounced regional disparities raise concerns not only about predictive performance, but also about whether machine learning models generalize equitably across territories. This study examines the role of regionalization in academic risk prediction by comparing national and state-level models trained under a unified experimental pipeline using longitudinal administrative data from over six million upper secondary student enrollments across all Brazilian states. Multiple supervised learning algorithms are evaluated, with Random Forest selected for detailed analysis due to its robust overall performance. Territorial fairness is assessed through an operationalization of Equal Opportunity and Equalized Odds based on state-level true and false positive rates. Results show that while national and state-level models achieve similar aggregate performance, substantial disparities persist in Recall across states. State-level models improve local risk detection in a small subset of states, often at the cost of increased false positives. These findings indicate that regional specialization is not uniformly beneficial and should be understood as a context-dependent trade-off between improved local detection and governance complexity. By separating territorial fairness auditing from performance-based model comparison, this study provides an evidence-based framework for reasoning about regionalization in large-scale educational risk prediction.
Recent advances in Natural Language Processing (NLP) have significantly enhanced text analysis and generation capabilities, introducing new tools that may help reduce student dropout rates. Concurrently, the growing adoption of Learning Management Systems (LMS) has contributed to an increase in educational data volume, creating opportunities for intervention through learning analytics applications. In this context, this study conducts a systematic review of research published over the past decade to investigate how NLP has been used to identify or mitigate the risk of student dropout in LMS such as Moodle and Canvas LMS. A search was performed across the Scopus and Web of Science databases, yielding a total of 142 results. After applying inclusion and exclusion criteria, 19 studies employing diverse approaches were selected and analyzed. Although relevant applications have been proposed, the results suggest that few studies provide data demonstrating the efficacy of such interventions, underscoring the need for further evidence to guide the implementation of NLP in these environments.
Early Warning Systems (EWS) have evolved substantially over recent decades, demonstrating strong potential to reduce school dropout and grade repetition. While early implementations relied on rule-based detection of academic risk, recent advances in artificial intelligence and machine learning enable risk prediction by analyzing historical data. Despite the growing accuracy of predictive models, the primary challenge has shifted from risk detection to ensuring that predictions translate into timely, equitable, and effective interventions. This gap is particularly salient in Low- and Middle-Income Countries (LMICs), where institutional constraints often hinder the operationalization of predictive insights. This paper introduces the Systemic Early Warning Systems Framework (SEWS), a systemic architecture designed to bridge the prediction-to-action gap in educational early warning systems. SEWS focuses on protecting students’ trajectories by integrating technological, institutional, and human-centered strategies. The framework comprises six core elements: institutional capacity-building assessment, predictive analytics, structured prioritization mechanisms, psychometric diagnostics of disengagement, stakeholder communication strategies, and continuous monitoring. The framework conceptualizes it as one component within a coordinated response to educational risk. Emphasizing systemic readiness and implementation, this work contributes to the ongoing shift from model-centric research toward actionable educational transformation. The SEWS framework provides a structured pathway for integrating EWS into real-world educational settings, supporting more responsive interventions and promoting student persistence.
O objetivo deste estudo é compreender como as Instituições de Educação Superior (IES), vinculadas ao Sistema Acafe de Santa Catarina(SC), abordam a regulamentação do uso da Inteligência Artificial (IA), com o intuito de preencher uma lacuna de pesquisa, tendo em vista o potencial impacto da IA no ecossistema educacional catarinense. Este estudo se caracteriza como uma pesquisa de natureza aplicada, de abordagem exploratória. Os dados foram coletados por meio de um questionário on-line (survey), enviado às 14 IES, obtendo-se o retorno de 09 delas. Os resultados obtidos revelaram que nenhuma das IES pesquisadas possuía regulamentação de uso da IA formalizada no período da pesquisa e, embora a implementação de políticas estivesse em estágio embrionário, já havia um esforço das IES pesquisadas para compreender essa tecnologia e equilibrar seus usos e riscos. Concluiu-se que diretrizes básicas, mesmo que provisórias, podem contribuir para uma integração responsável da IA no ambiente acadêmico e mitigar os desafios éticos. Este estudo contribui para a literatura ao apresentar o estágio de maturidade da regulamentação de IA em SC. Elas têm o desafio e a oportunidade de liderar iniciativas pioneiras na regulamentação e no uso ético da IA tanto em SC quanto no Brasil.
This workshop proposes a discussion of the AIED Unplugged framework as a lens for designing, implementing, and evaluating AI-in-Education solutions that operate offline-first, require low digital skill, support shared devices, and prioritize teacher mediation, with a focus on contexts of constrained infrastructure, especially in the Global South. Building on the diagnosis of educational inequalities exacerbated by gaps in connectivity, capacity building, and resources, the workshop aims to: (i) synthesize offline-first, low-skill, shared-device practices for equitable learning; (ii) examine evidence from real cases (e.g., mathematics and writing assessment via low-cost capture and offline analysis); (iii) connect leapfrogging and innovation to policy roadmaps; and (iv) co-create a prioritized research and policy agenda for underserved contexts. The program combines a keynote, contributed talks, a policy panel, and two thematic breakout blocks to produce practical artifacts, a public report, and comparable case vignettes. The main contribution of the I Workshop on AIED Unplugged is to consolidate a pragmatic pathway to “leapfrog” toward equity with educational AI in constrained environments.
Student dropout remains a critical challenge in global education, representing a significant loss of human capital. Although data-driven approaches like predictive modeling have achieved high accuracy in identifying at-risk students, a “trust gap” persists among educators. Prediction alone is insufficient; stakeholders require interpretable insights to guide effective intervention. This survey explores the paradigm shift from passive “black-box” detection to “actionable prescription” through the lens of Counterfactual Explanations (CFE). We analyze 23 primary studies published between 2021 and 2026, categorizing them using a novel Intervention Pipeline taxonomy that bridges four levels: Identification (predictive models), Interpretation (XAI), Visual Analytics (interactive visualization and human-in-the-loop analysis), and Actionable Recommendations (counterfactuals). Adopting an analytical framework, we systematically interrogate the literature regarding variable risk definitions, data source diversity, prediction horizons, and feature actionability. Our analysis reveals a critical disconnect: while many models prioritize accuracy using immutable features, few focus on the actionable variables necessary for generating feasible pedagogical recommendations. We conclude that while CFE offers a promising path for personalized support, future research must address challenges regarding causal validity, explanation stability, and scalability (individual vs. institutional) to successfully bridge the gap between algorithmic output and practical pedagogical action.
Effort estimation in software engineering has been evolving from prescriptive, process-centered approaches to social, collaborative approaches under the umbrella of agile methods since the beginning of the 21th century. To foster the fair participation of all team members, collaboration dynamics such as planning poker have been adopted. Thus, encouraging such collaboration settings in the classroom is important in software engineering education. However, gender stereotypes and roles could hinder the equitable participation of women in such processes, affecting not only technical aspects of the process and product but also core agile principles such as team members' motivation, team reflection capabilities, and process sustainability. In this study, we aim to illuminate the question of how women and men behave during a planning poker session in terms of key collaboration indicators. We examined the voice recordings of seven groups in planning poker role-playing teaching sessions; the interventions of each participant were coded in terms of the collaboration constructs: contribution, assimilation, self-regulating, team coordination, cultivation of environment, and integration. We applied Epistemic Network Analysis to study such indicators' frequency and epistemic connections. The results show statistically significant differences in the epistemic networks of women and men. While their collaboration indicators seem similar, women tend to be less expressive and have fewer connections than men regarding self-regulation, assimilation, and integration. This could be explained by the asymmetric conformation of groups, which may hamper the psychological safety of women, and suggest the need for equitably composed teams in design activities.
This paper presents a systematic review of the application of Sequential Pattern Mining (SPM) techniques in Virtual Learning Environments (VLEs) to improve learning analytics. The authors used the PRISMA-ScR methodology to explore studies that analyze log data from VLEs, particularly focusing on how SPM can identify student behaviors and predict academic performance. Out of 22 initially selected studies, 9 met the criteria and were fully analyzed. These studies primarily applied SPM in high school and undergraduate settings using algorithms like GSP, PrefixSpan, and Apriori. The review highlights challenges such as data granularity, the lack of real-time implementations, and the effectiveness of SPM in predicting student success and risk. The findings emphasize the need for better data preprocessing and integration of SPM into educational systems for early identification of at-risk students. Future work suggests enhancing SPM algorithms with real-time data analysis to support timely interventions in students' learning processes.
The widespread adoption of virtual and remote laboratories has transformed practical education by offering scalable, safe, and flexible environments for developing technical skills. Despite these benefits, challenges remain in providing timely, meaningful feedback that supports self-regulated learning and improves student outcomes. In this study, we analyze anonymized interaction logs from a virtual laboratory system focused on the Industrial Electrical Control Training Lab, used by 839 students enrolled in Electrical, Electromechanical, Mechanical, or Industrial Automation vocational programs offered by the National Service of Industrial Apprenticeship of Santa Catarina (SENAI/SC, Brazil). We introduce a novel methodology that combines sequential pattern mining (SPM) with automatic performance-based segmentation to analyze student behavior in virtual laboratory environments. Using interaction logs from students containing time-stamped cable connections and interface interactions, we segment learners into performance (High, Intermediate, and Low) based on the z-scores of accuracy and speed, using two strategies: (i) a K-means clustering approach, and (ii) a rule-based decision function. Building on this segmentation, we propose an algorithm to identify frequent action sequences and recurrent error patterns specific to each group. Chi-square tests confirmed significant associations (p < 0.05) between distinct error types and performance levels, and one-way ANOVA validated that both segmentation methods produce statistically distinct clusters. Our findings indicate that students in the high-performance group are likely to refer to schematics beforehand, whereas those in the lower performance group frequently begin with montage connections. These results can be integrated into LA dashboards to support real-time feedback for students and instructors, ultimately enhancing teaching strategies and learning in practical disciplines.
The widespread adoption of distance education (DE) has increased the use of Virtual Learning Environments (VLEs), resulting in large educational data repositories. This study explores applying the PrefixSpan algorithm for sequential pattern mining (SPM) in educational data to identify patterns in student interactions that predict academic success or failure. Using Python and the SPMF framework, we analyzed log data from students at the Federal University of Santa Catarina across three courses-Programming, Data Structures, and Algorithms-over four years. Our findings show that successful students engage more frequently within the VLE, demonstrating higher interaction levels. The study also highlights challenges in mining educational data, such as significant memory requirements and difficulties in identifying patterns in smaller datasets or early in the semester. Integrating SPM techniques in VLEs can offer insights for early intervention to support atrisk students. Future research should focus on enhancing real-time applications and improving result visualization for better usability by educators and learners.
In modern organizations, agile methods have become key strategies for project development, promoting collaboration and adaptability within teams. These approaches optimize communication and cooperation, enabling effective responses to evolving environmental demands. However, collaboration relies on effective communication, coordinated action, and cooperative participation elements that are difficult to evaluate without automated methodological support. Meanwhile, the emergence of transformer-based natural language processing (NLP) models has enabled the identification of semantic features for analyzing communication and collaboration. This study presents a verbal intervention classification system based on natural language processing (NLP) techniques, utilizing multimodal learning analytics to transcribe audio into text and characterize interactions. The system, built upon DistilBERT and trained with manually annotated examples, identifies and categorizes interventions into five classes: question, answer, feedback, suggestion, and comment. The model has been implemented in a functional platform that visualizes results through an interactive interface, allowing facilitators of collaborative activities to analyze the dynamics of interactions and participant contributions, thereby supporting the continuous improvement of such activities.
Studying collaborative dynamics in agile development teams requires multi- modal data that captures verbal and non-verbal communication. However, few experimental datasets provide this level of depth in real or simulated teamwork contexts. This article presents a multimodal dataset with experimental data collected during controlled sessions involving simulated agile development teams, each composed of four computer science students. A total of 19 groups (76 different participants) were organized, each participating in two collaborative activities: one without a coordination technique and another using the Planning Poker method. Three of these teams were designated as control groups. The resulting dataset includes audio recordings of verbal interactions and non- verbal behaviour data, such as body posture, facial expressions, visual attention, and gestures, captured using MediaPipe, YOLOv8, and DeepSort. It also contains time-aligned automatic transcriptions generated with WhisperX, attention logs, mimicry labels, and surveys on perceived equity in interactions. This re- source aims to provide a comprehensive view of collaborative behaviour in agile contexts, supporting both qualitative analysis of interactions and the development of predictive models of group performance. The dataset explores how shared visual attention and behavioural synchrony influence team effectiveness and decision-making through this multimodal approach. This work contributes a unique dataset valuable to researchers across multiple fields of study.