Scaling writing interventions in resource-constrained educational systems is challenging due to the heavy workload of manual essay assessment. This study investigates the feasibility of using AI-based text analysis to monitor primary education argumentative writing under the AIED Unplugged paradigm. Analyzing 261 handwritten essays from Brazilian public school elementary students, we compared AI-powered and manual transcription methods. The AI transcription achieved a high mean F1-score of 0.918, performing slightly better in higher grades. Critically, while structural metrics like sentence length showed differences due to AI segmentation errors, lexical and discourse-level features remained remarkably robust across both methods. These findings suggest AI transcription is a reliable, scalable solution for supporting intensive writing pedagogical programs in low-resource environments.
Most AIED systems rely on cloud infrastructure and stable internet, limiting applicability in resource-constrained contexts. This paper proposes an offline-first architectural blueprint that enables on-device LLM integration under strict computational and connectivity constraints. Grounded in Design Science Research, the architecture combines pedagogical control, local inference, and persistent storage. We instantiate this design in a mobile educational application aligned with the Brazilian National Common Core Curriculum, demonstrating the feasibility of fully offline deployment on commodity smartphones. The main contribution lies in providing transferable architectural insights that balance technical feasibility, pedagogical quality, and responsible AI principles.
Large language models (LLMs) increasingly generate formative feedback for students, yet little is known about how teachers revise this feedback before it reaches learners. Teachers’ revisions shape what students receive, making revision practices central to evaluating AI classroom tools. We analyze a dataset of 1,349 instances of AI-generated feedback and corresponding teacher-edited explanations from 117 teachers. We examine (i) textual characteristics associated with teacher revisions, (ii) whether revision decisions can be predicted from the AI feedback text, and (iii) how revisions change the pedagogical type of feedback delivered. First, we find that teachers accept AI feedback without modification in about 80 AUC=0.75 ). Third, qualitative coding shows that when revisions occur, teachers often simplify AI-generated feedback, shifting it away from high-information explanations toward more concise, corrective forms. Together, these findings characterize how teachers engage with AI-generated feedback in practice and highlight opportunities to design feedback systems that better align with teacher priorities while reducing unnecessary editing effort.
This study explores the integration of Augmented Intelligence (AuI) in Intelligent Tutoring Systems (ITS) to address challenges in Artificial Intelligence in Education (AIED), including teacher involvement, AI reliability, and resource accessibility. We present MathAIde, an ITS that uses computer vision and AI to correct mathematics exercises from student work photos and provide feedback. The system was designed through a collaborative process involving brainstorming with teachers, high-fidelity prototyping, A/B testing, and a real-world case study. Findings emphasize the importance of a teacher-centered, user-driven approach, where AI suggests remediation alternatives while teachers retain decision-making. Results highlight efficiency, usability, and adoption potential in classroom contexts, particularly in resource-limited environments. The study contributes practical insights into designing ITSs that balance user needs and technological feasibility, while advancing AIED research by demonstrating the effectiveness of a mixed-methods, user-centered approach to implementing AuI in educational technologies.
Lightweight Large Language Models (LLMs) enable educational content generation in resource-constrained environments, bringing AIED benefits to underserved populations. Yet, questions remain about output quality and pedagogical appropriateness when using lightweight models on offline, low-cost devices. This study investigates mathematical question generation through dual evaluation: human expert assessment and automated model-based evaluation. We generated 160 arithmetic questions aligned with Brazilian National Common Curricular Base (BNCC) standards across six competencies for elementary grades using Gemma 2 2b quantized model on Samsung Galaxy A34 smartphone. Results revealed small but statistically significant correlation between evaluation methods ( r = 0.31 , p < .001 ) and 75 r = 0.32 ), while division questions showed only 30 r = 0.47 ). EF02MA05 competency (basic facts) demonstrated highest consistency (mean = 4.97, 100
Feedback is a key factor in improving the writing skills of students, but providing it at a large scale remains a significant challenge. Large Language Models (LLMs) emerge as a promising solution; however, the pedagogical quality of the automatically generated feedback requires further investigation. This study examined how effective teachers perceived feedback texts generated by an LLM when the prompts were designed to reflect different aspects of feedback theory. To this end, 450 feedback texts were generated for 30 essays and evaluated by five teachers with varying feedback literacy profiles. The results showed that teachers' preferences for feedback models were strongly shaped by their profiles, indicating that there was no single universally "best" feedback model. From a learning analytics perspective, these findings demonstrate how combining theory-informed prompt design with evaluations from diverse teacher profiles can generate actionable insights for the development of scalable, AI-powered feedback systems. The study contributes to learning by showing that the effectiveness of LLM-based feedback depends on personalization not only for students but also for teachers, highlighting new directions for designing adaptive, equitable, and context-sensitive feedback analytics.
This paper describes how two AI-generated tools, personalized study guides and individualized feedback, were found to contribute to students’ sense of belonging in distance learning at a higher education institution in Brazil. Developed through the integration of semantic search (FAISS) and Large Language Models (LLMs), particularly GPT-4, the interventions were implemented within the educational system of the Brazilian largest university network, where most students come from low-income backgrounds. These learners often face academic isolation and must reconcile studies with work and family demands. Thus, the study guide aimed to help students focus on key course content and prepare effectively for exams, while the feedback provided personalized reflections on learning outcomes, highlighting mastered competencies and areas for improvement. Both tools included emotionally engaging language and were intentionally structured to enhance student motivation and recognition. Quantitative results demonstrated high levels of satisfaction, and qualitative responses revealed notable pedagogical, affective, and relational impacts. Students interpreted the materials not only as academic support but also as expressions of institutional care and encouragement. Many reported feeling more motivated, seen, and connected to their academic journey. By showing how AI tools can foster both academic performance and emotional engagement, this study highlights their transformative potential to scale personalization in distance education.
AIED-Unplugged (AIED-U) systems seek to expand access to educational AI in low-infrastructure contexts, but the Small Language Models (SLMs) they often rely on remain susceptible to hallucinations and pedagogical inconsistencies, making teacher validation essential. This study investigates how a Participatory Design process supports the redesign of validation workflows in an AIED-U system for generating mathematics exercise lists. Two mathematics teachers from the Brazilian public school system participated in three remote workshops. The results revealed seven interaction breakdowns in the original workflow, mainly related to limited operational control, interface noise, and AI-generated inconsistencies. Participants prioritized structural improvements over cosmetic adjustments, especially mechanisms for preview, revision, and direct intervention in generated outputs. The redesigned workflow reduced observed breakdowns and strengthened teachers’ sense of control and pedagogical agency, suggesting that transparency and interaction control are central design requirements for meaningful human–AI collaboration in the broader AIED-U context.
Access to AI-assisted educational technologies remains uneven in contexts with limited computational resources and digital skills, highlighting the need for infrastructure-aware and ethically grounded AIED approaches. The AIED Unplugged framework addresses this challenge by advocating offline-first systems that can operate under constrained conditions while remaining adaptable as educational infrastructure improves. In this work, an offline-first mobile application that instantiates the Unplugged principles is presented and used as a testbed to evaluate three instruction-tuned large language models across heterogeneous mobile devices operating fully offline. The results show that lightweight on-device models provide more stable and broadly deployable behavior, whereas larger models may improve benchmark performance but fail to operate reliably on lower-end hardware. These findings are interpreted as system-level evidence for pedagogical elasticity in Unplugged AIED: different local model-device configurations sustain different levels of feasible support, motivating adaptive forms of assistance rather than a uniform pedagogical role across contexts.
BackgroundA key skill for self-regulated learners is the ability to critically interpret and act on feedback-key components of feedback literacy. Yet, the connection between feedback literacy and self-regulated learning (SRL) remains underexplored, particularly in terms of how different levels of feedback literacy influence SRL processes in authentic learning contexts.ObjectivesThis study aimed to investigate the interplay between feedback literacy and SRL processes among secondary school students while working on a multi-source writing task using an online learning analytics (LA) platform.MethodsThe study involved 99 secondary school students from multiple nations (Brazil, UAE, India, and Australia) engaged in multi-source writing tasks. Students received personalised scaffolding feedback designed to enhance their SRL processes. Using K-medoids clustering, students were grouped based on their self-reported feedback literacy levels. Ordered Network Analysis (ONA) was employed to visualise and analyse trace data from the learning analytics platform, revealing SRL strategies across different feedback literacy profiles.Results and ConclusionsAnalysis of data from the 99 participants revealed two distinct feedback literacy groups with different SRL patterns. Proactive Feedback Engagers (n = 55) initially showed less effective SRL strategies but demonstrated significant improvement after receiving scaffolding interventions, adopting more balanced and adaptive regulation processes. In contrast, Moderate Feedback Engagers (n = 48) began with a more strategic approach but showed less adaptability as the task progressed, diverging from suggested SRL processes. These findings imply the need for more adaptive scaffolding approaches based on students' feedback literacy levels, highlighting the importance of tailored support in developing learning strategies based on different levels of feedback literacy.
Short open-ended questions represent a central resource in formative and summative assessments both face-to-face and online settings, ranging from elementary to higher education. However, grading these questions remains challenging for instructors, raising attention to the field of Automatic Short Answer Grading (ASAG). While ASAG has yielded valuable contributions to learning analytics, it often faces generalizability issues. Accordingly, the rapid advancement in Large Language Models (LLMs) has motivated their adoption to empower ASAG systems. Despite that, previous research has not investigated whether LLMs are fair graders in the context of ASAG. Therefore, this paper presents an empirical analysis aimed to understand LLMs' fairness in ASAG by using human grades as a baseline, comparing them to GPT-4's answers, and investigating whether the LLM's grades are equivalent in grading answers from varied groups of humans. Our results demonstrated that, while GPT-4 tended to be more lenient in its grading, it maintained consistent evaluation standards when assessing responses from different groups of students. GPT-4 remained consistent for questions of different subjects and levels of Bloom's taxonomy and for people with different demographics. These findings suggest GPT-4 is a fair grader, supporting its potential to empower educators and developers in using and designing ASAG systems. Nevertheless, we recommend further research to investigate these findings and understand how to optimize GPT-4's grades.
A Correção Automática de Respostas Curtas (ASAG) busca reduzir o esforço humano em avaliações educacionais de larga escala, mas ainda há poucas investigações em português brasileiro. Este estudo compara três grandes modelos de linguagem (GPT-4o-mini, Sabiazinho-3 e Gemini 2.0-Flash) e analisa o impacto de sete elementos de engenharia de prompt no desempenho dos modelos. Com base em um conjunto de dados em português, avaliamos todas as combinações possíveis desses elementos. A combinação de exemplos few-shot com rubrica explícita foi a mais eficaz; o raciocínio passo a passo beneficiou especialmente o GPT-4o-mini. Sabiazinho-3 teve maior concordância com humanos, Gemini 2.0-Flash obteve menor erro médio absoluto, mas com mais alucinações, e o GPT-4o-mini gerou saídas numéricas mais limpas.
Detecting cognitive presence in online learning environments helps understand and enhance educational interactions in online discussions. However, the lack of annotated data and class imbalance challenges building accurate machine-learning models for this task. This study explores the impact of oversampling techniques using GPT-4 and BERT for the automatic classification of cognitive presence in online discussion, aiming to address data scarcity and improve model generalization. We found that oversampling enhances classification accuracy and Cohen's Kappa, reaching up to 0.59 and 0.40, respectively, for a small sample of the original dataset. These results highlight the potential of recent oversampling methods in improving the detection of cognitive presence, informing future research and practical applications in online education.
Peer feedback (PF) is essential for improving student learning outcomes, particularly in Computer-Supported Collaborative Learning (CSCL) settings. When using digital tools for PF practices, student data (e.g., PF text entries) is generated automatically. Analyzing these large datasets can enhance our understanding of how students learn and help improve their learning. However, manually processing these large datasets is time-intensive, highlighting the need for automation. This study investigates the use of six machine learning models to classify PF messages from 231 students in a large university course. The models include Multi-Layer Perceptron (MLP), Decision Tree, BERT, RoBERTa, DistilBERT, and ChatGPT4o. The models were evaluated based on Cohen’s accuracy and F1-score. Preprocessing involved removing stop words, and the impact of this on model performance was assessed. Results showed that only the Decision Tree model improved with stop-word removal, while performance decreased in the other models. RoBERTa consistently outperformed the others across all metrics. Explainable AI was used to understand RoBERTa’s decisions by identifying the most predictive words. This study contributes to the automatic classification of peer feedback which is crucial for scaling learning analytics efforts aiming to provide better in-time support to students in CSCL settings.
Question-answering systems facilitate adaptive learning and respond to student queries, making education more responsive. Despite that, challenges such as natural language understanding and context management complicate their widespread adoption, where Large Language Models (LLMs) offer a promising solution. However, existing research is predominantly focused on English, proprietary models, and often limited to a single question type, subject, or skill, leaving a gap in understanding LLMs' performance in languages like Brazilian Portuguese and across questions of various characteristics. This study investigates how LLMs could be integrated in an educational question-answering system efficiently to answer different question types (multiple-choice, cloze, open-ended), subjects (mathematics and Portuguese language), and skills (summation/subtraction, multiplication, interpretation, and grammar), evaluating answers by GPT-4 - the main LLM at the time of writing - and Sabia - the open-source Brazilian Portuguese LLM - based on grades assigned by two experienced teachers. Overall, both LLMs demonstrated strong overall performance, with mean scores close to 9.8 out of 10. However, specific challenges emerged, with distinct strengths and weaknesses observed for each model, such as GPT-4's error in a multiple-choice subtraction question and Sabia's misinterpretation of a cloze question.
While providing feedback is a critical task in education, assessing open-ended responses and crafting personalized feedback for each student is time-consuming and challenging. Large Language Models (LLM) can be used to support this process by leveraging their natural language understanding capabilities to evaluate open-ended answers and generate automated feedback. However, in real-world learning environments, empirical gaps remain regarding: I) How do students perceive LLM-generated feedback compared to teacher-provided feedback? and II) What impact does LLM assistance have on instructor feedback practices? Thus, this paper presents a controlled experiment with 60 secondary school students in Brazil, comparing traditional teacher feedback with LLM-assisted feedback using the Tutoria platform. Results showed no significant difference in students’ perceptions of feedback quality between the two approaches, with 85
Alunos de escolas públicas brasileiras precisam de auxílio em matemática. Nesse sentido, a literatura científica tem explorado maneiras de, por meio da tecnologia, auxiliar professores e alunos no processo de ensino-aprendizagem. No entanto, as pesquisas geralmente visam propor soluções sem antes investigar as necessidades de quem está na linha de frente: os professores. Assim, este artigo tem como objetivo investigar aspectos relevantes do cotidiano em sala de aula sob a ótica de professores de matemática de escolas públicas brasileiras. Utilizamos as duas primeiras fases do Design Thinking: empatizar e definir. Na primeira fase, realizamos um grupo focal com quatro professores de quatro escolas públicas de ensino fundamental, que analisamos qualitativamente utilizando a Teoria Fundamentada (Grounded Theory - GT). Com base nos resultados da GT, na segunda fase do Design Thinking, propusemos a persona e o mapa de empatia. Nossos resultados fornecem implicações para designers e desenvolvedores de soluções tecnológicas educacionais a partir da aplicação prática da pesquisa em experiência do usuário. Além disso, esperamos que nosso estudo seja um ponto de partida para novos pesquisadores da área explorarem as necessidades dos professores de escolas públicas brasileiras.
A correção manual de provas de múltipla escolha consome tempo e atrasa o feedback aos estudantes. Scanners são uma solução, mas inacessíveis para muitas instituições. Este artigo propõe uma alternativa baseada em câmeras de smartphones, preenchendo uma lacuna sobre o custo-benefício de soluções móveis. Desenvolvemos uma ferramenta com o modelo You Only Look Once (YOLO) para reconhecer respostas em folhas de gabarito. Após testes e refinamentos, nossa solução atingiu 97% de acurácia com um tempo de inferência de 25ms, otimizando o processo de correção e acelerando o retorno avaliativo.
Escreva Mais é um aplicativo móvel inovador multiplataforma (compatível com IOS e Android) que usa Inteligência Artificial para auxiliar professores na correção de redações. é voltada para contextos com baixa infraestrutura digital, onde ela transcreve textos manuscritos a partir de imagens, atribui pontuações automáticas e gera feedbacks para os alunos. O sistema utiliza o Gemini 2.0 Flash para a análise textual e foi desenvolvido com backend em Django. O objetivo é demonstrar uma aplicação de Aprendizado Aprimorado por Tecnologia (TEL) que otimiza o fluxo de trabalho docente e facilita a avaliação em diversas realidades educacionais.