
The generation and identification of paraphrases are essential challenges in natural language processing (NLP), involving considerable consequences for education, plagiarism detection, and content analysis. Despite significant advancements in sentence-level paraphrasing, paragraph-level paraphrasing is still inadequately investigated, especially across various fields. This research is the first to examine the effectiveness of SALAC algorithms and Transformer-based models in generating coherent and semantically accurate paraphrases across various domains, utilising the ALECS dataset, which includes text samples from five different educational fields, such as Economics and Anthropology. The methodology uses SALAC algorithms for sentence reordering, according to coherence scores obtained from the ALBERT model's sentence Order Prediction (SOP). Real-world human evaluation is performed on the paraphrased paragraphs, measuring semantic similarity and coherence via a Likert scale of five pre-defined points. Evaluators involve academic students and researchers, ensuring the reliability of outcomes. Research findings show that SALAC algorithms successfully maintain semantic integrity and coherence across several domains, highlighting their generalisability. Additionally, the study examines how domain-specific readability impacts annotation reliability and paraphrasing performance. This research contributes to the development of robust, domain-independent paraphrasing techniques at the paragraph-level for educational applications and plagiarism detection, thereby advancing NLP solutions across multiple domains.
The problem of updating the logic of ITS operation when introducing innovative functions and adapting to the requirements of a particular educational institution (localization) is considered. Arguments in favor of multi-stage decision-making based on the model of P.K. Anokhin are given. The architectural solution of such ITS is proposed and the stages of generalization of processed data from the digital educational footprint are described. The possibilities of using the mechanism of recommender systems and fuzzy logic for reconfiguring the logic of decision-making are discussed. By the example of using the mechanism of competency development assessment and course individualization, the process of changing the logic to fit the new functionality of the system is shown. The results of implementation in the ITS are demonstrated for an experimental ITS used in the educational process of the master's degree program at the Siberian Federal University.
Question-answer generation (QAG) focuses on creating answers to a set of questions based on a pre-specified context. Some applications for QAG are included in the field of information retrieval and data augmentation for QA models. However, one of the most popular uses of automatic answer retrieval presently is notably in the field of education. This study focuses on harnessing advanced artificial intelligence (AI), in particular, large language models (LLMs) for creating a Question-Answering system by generating answers for a set of predefined questions in healthcare. This sub-system will feed into and be integrated into a larger framework for a project with a goal of building a bilingual (English and Spanish) intelligent tutoring system (ITS) geared towards educating survivorship skills to low-literacy Hispanic breast cancer survivors. Several Hugging Face’s Models and evaluation metrics were used for generating and evaluating answers based on previously-generated health-related questions. Synthetic data generation and Code-switching functionality was also incorporated to determine how well the models perform when accounting for both English and Spanish. GPT-generated English answers achieved higher accuracy than Spanish answers (96.38
The evaluation of neural network models, particularly for classification tasks, is a critical phase in the development pipeline, ensuring models perform reliably across diverse datasets and real-world scenarios. A robust evaluation framework relies on multiple performance assessment techniques, each addressing different aspects of model accuracy, generalization, and robustness. This paper presents a study on the development and evaluation of a deep learning-based classification system for UML diagrams, focusing on measuring accuracy through various techniques. The research leverages deep learning techniques to classify UML diagrams into predefined categories, such as class, sequence, state, other UML and non-UML diagrams. The system was evaluated using cross-validation with a k-folding approach, to confirm the model does have a reliable performance estimation and minimizing overfitting. Automatic evaluation and self-training methods were explored to enhance the model’s robustness and adaptability to new datasets. The experiments demonstrated an average accuracy of 82.8
This paper explores a neuromorphic approach to semantic knowledge representation using spiking neural networks (SNNs) for relational inference within knowledge graphs (KGs). While traditional KG models often rely on dense embeddings that lack interpretability, SNNs offer a biologically inspired alternative by encoding relationships through discrete, event-driven spikes. We evaluate three architectures-LIF-Base, RecurrentSNN, and LiquidSNN-on a dataset of 54 computer science-related KG triplets. Performance is assessed using relationship classification accuracy, temporal and spatial stability, and spike variability. Results show that relationship types produce distinct spike activation patterns, with LiquidSNN achieving the highest accuracy and spatial coherence. These findings support the potential of SNNs for structured, interpretable KG reasoning, with future applications in adaptive learning systems and explainable AI.
Achieving deeper understanding in online learning requires adaptive support that evolves with learner needs. This paper traces the iterative design of Ivy, a generative AI coach that combines structured knowledge representations with large language models to deliver pedagogically aligned responses. Across three versions: pre-Ivy, Ivy 1.0, and Ivy 2.0, we analyzed learner interactions and preferences. Expert users consistently favored Ivy's structured, example-driven responses, while novices often preferred the more conversational tone of a ChatGPT-powered assistant. These insights shaped Ivy's refinement and suggest its potential to support novices in developing conceptual understanding and progressing toward expertise.
The use of generative AI tools in medical education, such as Copilot and Gemini, has shown great promise for improving student evaluation and learning experiences. This study assesses the effectiveness of these tools in producing diagnostic microbiological evaluations, with a focus on the quality and discrimination level of their questions. The results show significant variations between the two instruments, especially in terms of question scores and concentration regions. Copilot-generated evaluations produced better total scores and featured a variety of theoretical and practical questions, whereas Gemini's assessments were more visually focused and had a broader range of score variances between questions. The analysis of discrimination levels in both evaluations revealed that good, fair, and poor discrimination questions appeared in both instruments. Questions with significant discrimination related to practical skills and critical thinking, whereas basic knowledge questions frequently had weaker discrimination. Despite the promising results, limitations were noted, including the necessity for thorough editing to assure the accuracy of AI-generated questions and the clarity of phrasing to avoid student confusion. This study demonstrates the potential for generative AI to advance medical education by increasing students' practical and rational skills. However, it emphasizes the significance of human monitoring in validating and refining AI-generated material. With appropriate integration, generative AI techniques can provide a viable path for novel and successful instructional practices in diagnostic microbiology and beyond.
The article presents a method and tool that use deep learning to enable the preparation of educational graphic materials adapted to the needs of visually impaired people, focusing on supporting the teacher's work in adapting exercises to an alternative audio-tactile form. The challenges educators face, the proposed alternative in the form of an artificial intelligence tool, and the research carried out are presented. The research group comprised 9 teachers with extensive experience working with visually impaired students, and the study spanned one semester in a high school setting. The proposed evaluation method encompassed several factors, including a comparison of the effectiveness of the obtained artificial intelligence models for image analysis, a comparison of the time required for exercise preparation in contrast to previously used solutions, the usability scale of the proposed system measured by subjective indicators, and the reduction of the workload of the tutor by using a standard task-load index. The findings suggest that the proposed method and the teacher support tool attain high-efficiency scores in the quality of the classifier and markedly reduce the workload necessary to adapt educational material to the needs of blind individuals. Further development of the tool toward generative models can contribute to the creation of a comprehensive and automated adaptation tool, thereby reducing exercise preparation time to the requisite minimum.
In this paper, we describe the enhancement of an intelligent tutoring system for teaching details of program-element scope in the C++ programming language. We expanded the range of learned content, introduced hints with informative descriptions, and provided detailed feedback on errors during premature completion of a learning problem. In addition to scope and visibility, the new tutor can also teach the concept of object lifetime and data flow direction. We conducted an experiment to evaluate the effects of the tutor improvements on the learning process. The results showed that while the overall level of learning gains increased, the gains in some of the topics remained the same and for the most complex topic, they decreased because the students were distracted by the new material. It shows that when expanding intelligent tutors, the teachers should also consider expanding the time spent with the tutor. It was also shown that the students who actively used hints learned significantly more than the students who ignored that feature.
The rapid advancement of interactive systems and Extended Reality (XR) technologies has the potential to transform healthcare and well-being. To this end, this paper introduces a fully customizable VR application as a proof of concept designed to support diverse therapeutic needs. A practical memory rehabilitation scenario and an engaging upper limb recovery exercise was integrated, demonstrating the system's versatility for both cognitive and motor therapy. By improving accessibility and engagement in rehabilitation, our VR application empowers healthcare professionals, paving the way for more effective, personalized therapy solutions.
In this study, we propose an improved selective cross-shaped window self-attention mechanism, incorporated into a depthwise separable CNN, that guides self-attention to attend only to the most relevant regions associated with Alzheimer's disease and the relevant long dependencies and relationships between these affected regions. This new method was enhanced by multi-modal Conditional Progressive GAN. Our proposed method not only enhances the feature extraction efficiency but also maintains computational effectiveness. The results underscore the key role of advanced generative models in developing advanced deep learning methods for early Alzheimer's detection. We evaluated our proposed method on the Alzheimer's Disease Neuroimaging Initiative (ADNI) dataset, achieving a high accuracy of 99%.
Existing Virtual World (VW) based curriculum oriented educational systems use conventional non-player characters (NPCs) to interact with users (players), represented as avatars, to guide and help them to accomplish learning activities. Also, few of them use some kind of gamification and keep data for user interactions and activities. In this paper, we present the design and implementation of an educational system based on VW technology, which employs gamification features, two types of NPCs, a conventional and an LLM based, and a Database that stores, apart from educational information, information about interactions of users with the NPCs. Evaluation of the system shows that both types of NPCs are useful for different reasons. LLM based ones offer more interesting dialogues, and conventional ones more targeted help.
In high-stress aviation environments, maintaining optimal pilot performance is critical for flight safety. This study investigates the relationship between cognitive workload, stress, and facial thermal responses during simulated Airbus A320 takeoff procedures. Utilizing a multimodal approach, facial thermography, heart rate monitoring, and EEG-derived workload measures were recorded from ten participants-five with and five without prior simulation experience-while they navigated 12 distinct flight scenarios comprising both standard and engine failure conditions. The experimental design featured two rounds of simulation tasks interleaved with rest periods to capture both cooling (during task performance and stress) and warming (rest and recovery) phases. Results revealed that elevated cognitive workload induced significant cooling across the nose, forehead, and cheeks, with the nasal region exhibiting the most rapid and pronounced temperature decline. These thermal changes were synchronized with increases in heart rate and subjective workload ratings, thereby validating the use of facial thermography as an objective, real-time indicator of stress. Additionally, preliminary analyses suggest that individual differences, such as prior simulation experience, moderate the magnitude of these physiological responses. The findings underscore the potential of integrating facial thermal imaging into cockpit monitoring systems as an early warning mechanism for cognitive overload. Future work should focus on refining critical temperature thresholds and exploring real-world applicability to further enhance pilot training and operational safety.
This paper presents a cognitive assistance approach using ontologies to enhance pilot training and decision-making. A reference ontology, combining a domain ontology representing the aircraft's external environment and a task ontology describing flight procedures, is leveraged. Based on this, a synthetic pilot, developed using the ACT-R cognitive architecture, provides personalized cognitive assistance to pilots, helping them in their training and decision-making processes. The synthetic pilot adapts to individual learning preferences and progress, offering tailored feedback and guidance. The aim is to increase pilots' skills and enhance aviation safety through this personalized ontological cognitive assistance.
Intelligent tutoring systems (ITSs) have long held the gold standard for learning in digital learning environments. However, ITSs have historically required substantial authoring effort, limiting scaling. Computer assisted instruction has continued to be more widely available to students, and generative AI now enables richer student-computer interactions. One key feature of ITSs is the delivery of personalized feedback on student responses. In this paper, we discuss the nature of successful personalized feedback, the development of personalized feedback using generative AI as part of an automatic question generation system, and initial analysis of this feedback performance using data from students in a traditional university course.
This paper presents an interdisciplinary assessment method using items grouped according to the disciplines targeted to identify student competencies. For each subject, a weight is used depending on the requirements of the assessor. The model we present uses different databases, one for each subject used for assessment, and the selection of items for assessment is done using genetic algorithms. The model implementation is done using the Python language. Based on the results, the genetic algorithm efficiently optimizes the generation of multidisciplinary tests and its performance can be improved by dynamically adapting parameters and diversifying solutions to avoid stagnation.
This paper explores how a Low-Code/No-Code (LCNC) platform can be used by non-technical users, such as educators, to design and deploy a Sequential Agent-Based Generative AI System to facilitate instructional design. The system deploys an LLM-based sequential workflow of AI agents to support educators in the first three stages of the ADDIE instructional design model: Analysis, Design, and Development. It follows a co-design, Human-In-The-Loop (HITL) approach, where AI agents guide instructional designers on needs analysis, content validation and generation, while allowing user intervention. The system also explores the role of self-checking agents for fact-checking, bias detection, and instructional quality review, based on specific prompts. However, the identified potential remains theoretical, requiring empirical validation through user testing to assess usability, effectiveness, and adoption by non-IT users.
The rapid rise of Artificial Intelligence (AI) tools prompts educators to revisit how learning theories inform instruction and assessment. While AI supports personalization, adaptive feedback, and automation, its alignment with established learning theories remains unclear. This study examines how AI tools engage cognitive, metacognitive, and affective dimensions using an open dataset of AI-powered educational tools. Using topic modeling techniques and cluster analysis, we identify thematic categories and their alignment with Bloom's Taxonomy, the Two-Level Model of Metacognition, and Control-Value Theory. Findings show a strong emphasis on lower-order processes (e.g., Remembering and Understanding), with fewer designed for higher-order thinking. Some tools enable self-monitoring and reflection, but few support strategic learning. In the affective domain, AI enhance motivation via personalization but lacks emotional adaptation. These results highlight the need for theory-informed AI design to promote deeper, self-regulated, and emotionally responsive learning.
Language models (LMs) are a new technology attracting research interest in many fields, including educational technology, for their ability to generate media, including text, from a prompt. For teachers creating new questions is a time-consuming process yet essential for developing and confirming a learner’s understanding of the material. This research extends existing research on LM-based question generation by investigating what challenges exist when developing a question generator to support both teachers and learners. In addition, to support personalization, the generation of questions at three difficulty levels is investigated. Finally, different settings are systematically evaluated for creating the best possible questions. Using training data from a course, two small LMs, Gemma 7B and Gemma-7B-it (instruction pre-tuned), were trained under 84 experimental settings. 7 prompts styles (10 prompts each), plus 9 human-generated prompts, for each course unit were used for each experimental setting. It was found that considerable manual effort to prepare the training data and extract the generated questions from the LM’s response was required. Depending on the settings, a question with the proper difficulty was generated between 61.5
The present study investigates the ability of Large Language Models (LLMs) in generating and evaluating formal proofs within propositional logic. Specifically, we examine whether an LLM can accurately construct formal proofs to determine the validity of logical arguments and if other independent LLMs can reliably assess the correctness of such proofs. That is, when a LLM plays the problem solver role, the other involved LLMs play the role of evaluator. The evaluation comprises 12 diverse propositional logic proof problems, classified into distinct characteristics. Experimental scenarios were designed such that one model, exemplified by DeepSeek, performed the solver role by generating formal proofs, while three other models, represented by Qwen, GPT, and Gemini, independently evaluated the validity of these proofs. Our findings for each different configuration are described in this article, revealing positive results in favor of the LLMs used.