Evaluating students’ natural language explanations of source code remains challenging in computer science education, especially for intelligent tutoring systems needing real-time feedback. While Large Language Models (LLMs) can assess such explanations via prompting, these methods require extensive prompt engineering and yield inconsistent results. This paper explores instruction fine-tuning of open-source LLMs for automated evaluation of students’ line-by-line Java code explanations. Four models (CodeGemma 7B, Mistral 7B, CodeLlama 7B, Llama 3.1 8B) were fine-tuned on 2,415 training examples and evaluated on 302 test instances. All models achieved strong correlations with human judgments (Pearson r = 0.652–0.681, p < 0.0001), with CodeGemma 7B performing best (r = 0.681). Statistical testing revealed no significant differences between models (all pairwise p > 0.10), indicating that instruction fine-tuning enables both code-specialized and general-purpose models to achieve comparable performance. All instruction fine-tuned models substantially outperformed GPT-4 and GPT-5.4 prompting baselines (r = 0.330–0.376) by +0.276 to +0.351, demonstrating that task-specific fine-tuning is far more effective than few-shot prompting for educational assessment while being computationally efficient, supporting scalable and accurate automated assessment for code comprehension tutoring systems.
Code comprehension theories postulate that programmers need to identify the logical steps of a computer program. This work examines the ability of large language models (LLMs) to identify and explain logical steps and their corresponding blocks of code in well-structured programming tasks. To evaluate the LLMs’ performance, we compare the LLM-identified logical steps with those identified by human experts. We assess the match between human and LLM annotations on this task using automated similarity analysis under multiple alignment strategies. Results show that the highest similarity between LLM-generated and human expert descriptions of the logical steps reaches up to 64.4%.
Code comprehension is a critical skill for computer science students who spend a substantial portion of their time engaged in reading and understanding code. While prior research has explored students’ use of Large Language Models (LLMs) for tasks such as code generation or bug fixing, there is very limited understanding of how effectively these students can prompt LLMs to get help for code comprehension activities. In this paper, we present a novel study exploring how intro-toprogramming students, i.e., novices to programming, freely prompt LLMs for code explanations. The goal was to understand how well LLMs can support students’ code comprehension activities with no training on advanced LLM prompting techniques. Our analysis reveals that while students’ prompts vary significantly, the quality of the LLM-generated code explanations for typical intro-to-programming code examples was considerably accurate, and complete. Students primarily use three types of prompts: whole-program explanation, specific logic explanation, and conceptual explanation while interacting with LLM. We also observed that access to LLM assistance is associated with a statistically significant increase in students’ confidence and improvements in code comprehension tasks.
This paper investigates various approaches using Large Language Models (LLMs) to identify gaps and misconceptions in students' self-explanations of specific instructional material, in our case explanations of code examples. This research is a part of our larger effort to automate the assessment of students' freely generated responses, focusing specifically on their self-explanations of code examples during activities related to code comprehension. In this work, we experiment with zero-shot prompting, Supervised Fine-Tuning (SFT), and preference alignment of LLMs to identify gaps in students' self-explanation. With simple prompting, GPT-4 consistently outperformed LLaMA3 and Mistral in identifying gaps and misconceptions, as confirmed by human evaluations. Additionally, our results suggest that fine-tuned large language models are more effective at identifying gaps in students' explanations compared to zero-shot and few-shot prompting techniques. Furthermore, our findings show that the preference optimization approach using Odds Ratio Preference Optimization (ORPO) outperforms SFT in identifying gaps and misconceptions in students' code explanations.
Assessing student responses is a critical task in adaptive educational systems. More specifically, automatically evaluating students' self-explanations contributes to understanding their knowledge state which is needed for personalized instruction, the crux of adaptive educational systems. To facilitate the development of Artificial Intelligence (AI) and Machine Learning models for automated assessment of learners' self-explanations, annotated datasets are essential. In response to this need, we developed the SelfCode2.0 corpus, which consists of 3,019 pairs of student and expert explanations of Java code snippets, each annotated with semantic similarity, correctness, and completeness scores provided by experts. Alongside the dataset, we also provide performance results obtained with several baseline models based on TF-IDF and Sentence-BERT vectorial representations. This work aims to enhance the effectiveness of automated assessment tools in programming education and contribute to a better understanding and supporting student learning of programming.
Assessing students' responses, especially natural language responses, is a major challenge in education. In general, in education contexts, automatically evaluating what learners do or say is important as it enables personalized instruction, e.g., based on what the learner knows tailored tasks and feedback are given to the learner. Recently, deep learning techniques led to state-of-the-art methods in NLP such as transformer-based methods which resulted in significant performance improvements for many NLP tasks such as text classification and question answering. However, there is not much work exploring such methods for assessing students' free answers, particularly in the context of code comprehension, which brings additional challenges as the student explanations include code references as well. This paper explores the potential of applying automated assessments methods using transformers to code comprehension. We fine-tuned pre-trained transformer models, including BERT, RoBERTa, CodeBERT, and SciBERT, to see how well they can automatically judge students' responses to code comprehension tasks. Our results demonstrate that these models can significantly enhance the accuracy and reliability of automated assessments, offering insights into how the latest NLP techniques can be leveraged in computer science education to support personalized learning experiences.
Understanding how students use math strategies is important to help us build tools and techniques that improve cognitive flexibility in students, i.e., select strategies that are appropriate and efficient for a problem. In this work, we focus on instructional content within MATHia, where students choose between strategies that were previously taught to them independently. Some problems favor one strategy over the other, giving us the opportunity to understand how/if students learn to pay attention to problem characteristics that suggest one strategy over the other. Using data from over 600 schools, we show that students find it hard to adapt their strategies to suit a problem. Further, we learn a BERT model to learn embeddings for strategies, develop a prediction task to distinguish between successful and unsuccessful strategies, and analyze its results to reveal deeper insights into student strategies.
AI models have shown a remarkable ability to perform representation learning using large-scale data. In particular, the emergence of Large Language Models (LLMs) attests to the capability of AI models to learn complex hidden structures in a bottom-up manner without requiring a lot of human expertise. In this paper, we leverage these models to learn Math learning strategies at scale. Specifically, we use student interaction data from the MATHia Intelligent Tutoring System to learn strategies based on sequences of actions performed by students. To do this, we develop an AI model based on BERT (Bidirectional Encoder Representations From Transformers) that has two main components. First, we pre-train BERT using an approach known as Masked Language Modeling to learn embeddings for strategies. The embeddings represent strategies in a vector form while preserving their semantics. Next, we fine-tune the model to predict if students are likely to apply a correct strategy to solve a novel problem. We demonstrate using a large dataset collected from 655 schools that our approach where we pre-train to learn strategies from a sample of schools can be fine-tuned with a small number of examples to make accurate predictions over student data collected from other schools.
The proventriculus has an important adjustment function in broiler chickens, potentially impacting nutrient availability and performance. Given their contribution to protecting the gastric lining and facilitating digestion, the understanding of the mucus secretion cells within this organ is essential. Therefore, this study aims to investigate, through histochemical methods, the types of mucins secreted at the proventriculus level in 10 day old Ross 308 broiler chickens. Fragments of the proventriculus (five chickens) were histologically processed by paraffin embedding and the slides were stained using three techniques: periodic acid–Schiff (PAS) reaction, alcian blue (AB) staining pH 2.5, and combined PAS-AB staining. The color of the mucus present in the cells of the gastric mucosa was qualitatively assessed using the Photoshop Color Picker using the red, green, blue (RGB) color model and the numerical values for the RGB spectrum were analyzed from a statistical point of view. The synthesized mucus in the proventricular mucosa is predominantly PAS-positive in the superficial half of the mucosa and in the deep half, the synthesized mucins are both neutral and acidic with a predominance of acidic mucins. In the submucosa, only a few cells lining the central lumen of the glandular lobules are moderately PAS- and AB-positive.
Assessing self-explanations of code is a critical task in educational Natural Language Processing (NLP), essential for providing automated feedback in programming education. While recent studies demonstrate that LLMs exhibit analogical reasoning capabilities, their application to educational code explanation assessment remains underexplored. To address this gap, we propose a novel three-stage approach: (1) prompting the LLM to extract Open Information Extraction (OIE) units from student and expert explanations, (2) prompting the LLM to construct semantic graphs from these units, and (3) employing LLM-based analogical reasoning to assess explanation similarity. We evaluate our approach on the Self-code corpus using Pearson and Spearman correlations. Our method achieves correlations of 0.8 and 0.76, respectively, significantly outperforming supervised models like BERT (0.74/0.75) and unsupervised approaches like text-embedding-ada-002 (0.67/0.64). These results demonstrate the effectiveness of structured semantic representation combined with analogical reasoning for educational assessment tasks.
The birds, in contrast to mammals, have two cranial cava veins. The two cranial cava veins and caudal cava vein are the largest veins in chickens. In general, veins have a wall consisting of 3 tunics: intima, media, and adventitia. This study aimed to describe the microscopical structure of the left and right cranial cava veins in 10-day-old chicken broiler. Fragments from the left and right cranial cava veins were collected during the necropsy and were histologically processed by paraffin inclusion and stained with the Verhoeff-trichrome method. The left and right cranial vena cava have relatively thin walls compared to the lumen. The intima is formed by an endothelium, the media is formed by circular smooth muscle cells and the adventitia is formed by dense non-oriented connective tissue. A particular aspect is the fact that smooth muscle cells with a predominantly longitudinal orientation are present in the structure of the adventitia. Both in the left and right cranial veins, the proportion between the amounts of muscle tissue in the media relative to that in the adventitia is not identical on the entire circumference of the vessel. Moreover, in some areas of the venous wall, the quantity of muscle cells present in the tunica adventitia is higher than in tunica media.
The feeding behaviour of fish is drastically influenced by food availability and dietary preferences. In omnivorous fish, chemoreception plays an important role in feeding and morphological adaptations may be observed in several regions of the body. However, it is not clearly known if the fish need to feel the taste before swallowing the food. The present study aims to describe the receptors for taste perception and their distribution at the level of the gill rackers, underlining their importance in the food sorting behaviour. Paired gills were harvested from Carpathian gudgeon fish Gobio carpathicus Vladykov, 1925 and immersed in 10% buffered formalin. The samples were processed according to the current paraffin embedding technique and stained with Goldner’s trichrome method. The obtained results suggest that the Carpathian gudgeon presents, up to a point, the common gill morphology. However, on the pharyngeal face of the gills, more exactly on the gill rackers, are present several structures with an onion-like shape, disposed through the surface of the epithelium. Those elements consist of sensorial cells, sustained by sustentacular and basal cells, forming a taste bud. Due to their disposition on the inner surface of the gills, those structures may act like a sorter, enhancing the rackers sieve activity. In conclusion, the histological findings suggest that the Carpathian gudgeon, a common omnivorous fish, may use taste reception at the level of the gill rackers before swallowing the food.
Assessing student's answers and in particular natural language answers is a crucial challenge in the field of education. Advances in machine learning, including transformer-based models such as Large Language Models(LLMs), have led to significant progress in various natural language tasks. Nevertheless, amidst the growing trend of evaluating LLMs across diverse tasks, evaluating LLMs in the realm of automated answer assesment has not received much attention. To address this gap, we explore the potential of using LLMs for automated assessment of student's short and open-ended answer. Particularly, we use LLMs to compare students' explanations with expert explanations in the context of line-by-line explanations of computer programs. For comparison purposes, we assess both Large Language Models (LLMs) and encoder-based Semantic Textual Similarity (STS) models in the context of assessing the correctness of students' explanation of computer code. Our findings indicate that LLMs, when prompted in few-shot and chain-of-thought setting perform comparable to fine-tuned encoder-based models in evaluating students' short answers in programming domain.
Worked examples, which present an explained code for solving typical programming problems are among the most popular types of learning content in programming classes. Most approaches and tools for presenting these examples to students are based on line-by-line explanations of the example code. However, instructors rarely have time to provide explanations for many examples typically used in a programming class. In this paper, we assess the feasibility of using LLMs to generate code explanations for passive and active example exploration systems. To achieve this goal, we compare the code explanations generated by chatGPT with the explanations generated by both experts and students.
This research paper explores to what extent large language models (LLMs) can generate line-by-line explanations of code examples used in intro-to-programming courses such as CS1 (Computer Science) and CS2. While it is known that LLMs can generate code explanations, a systematic analysis of those explanations and their appropriateness for instructional and learning purposes is needed, which is the goal of this paper. Specifically, the paper explores how different types of prompts impact the nature and quality of line-by-line explanations relative to human expert explanations. We report a quantitative and qualitative analysis that compares AI-generated explanations with explanations produced by human experts. Furthermore, we investigate to what degree LLM can generate explanations for learners of various levels of mastery.
Adapting to a student's problem solving strategy can lead to improved engagement and motivation. In this work, we develop an AI-based approach to analyze math learning strategies at scale. Specifically, we use a state-of-the-art AI model, namely, BERT to learn structure within strategies observed in large datasets. In particular, we consider the MATHia ITS and define strategies as sequences of steps that a student follows in solving the problem. We apply BERT pre-training to learn semantic representations of strategies from a workspace in MATHia that allows for different strategies. Further, we fine-tune these embeddings to train them on downstream tasks such as identifying a strategy and understanding drift in strategy. Our preliminary results are encouraging and demonstrate that BERT can uncover hidden structure in strategies and therefore is a promising direction to analyze large-scale math learning data.
BACKGROUND:Rabbits are herbivores with a distinctive digestive strategy that differs significantly from other caecal fermenters (e.g., horses, guinea pigs) and ruminants. In view of this, the current study aimed to highlight distinctive histological and morphometric features of the caecal mucosa in adult rabbits that accentuate its major role in digestion. The caecal and jejunal samples were harvested from five 1-year-old domestic rabbits and processed by regular paraffin-embedding histological technique followed by Goldner's trichrome staining. A comprehensive morphological and morphometrical analysis of the jejunal mucosa vs. caecal mucosa was performed. RESULTS:Microscopically, as in the case of the jejunal mucosa, the caecal mucosa presents long and often branched finger-like villi covered by a simple columnar epithelium mostly made of enterocytes with a prominent microvillous brush border. Besides, the caecal villi include a lacteal along with the villous muscle. Statistically, except for villus length, all the parameters assessed in the caecal mucosa, including villus width, villus count, thickness of the brush border and enterocyte/goblet cells ratio, revealed a high grade of similarities with the jejunal villi. CONCLUSIONS:According to the obtained results, the caecal mucosa in adult domestic rabbits includes unique features, namely caecal villi, structures infrequently presented in the large intestine of other adult mammals. Those structures once more emphasize the major role of the caecum not only in fermentation but also subliminally in local absorption. To our knowledge, this is the first reliable microanatomical and morphometric report of caecal villi in adult domestic rabbits.
The Automated Essay Scoring (AES) task is an important NLP research problem given its significance for the education ecosystem. Recently, researchers started to apply a hybrid approach to this task. This hybrid approach incorporates into a deep learning model expert features that assess a particular dimension of the essay. Motivated by these successes, we propose to automatically assess essays using a hybrid approach that relies on external discourse knowledge. Our proposed model consists of using transformer-based embeddings to generate semantic representations of essays. Then, we incorporate several discourse features into these representations. Finally, we apply a linear classifier to generate the final score. To evaluate the effectiveness of this approach, we have conducted extensive experiments using the Automated Student Assessment Prize dataset (ASAP). The performance of the proposed model has been evaluated using the Quadratic Weighted Kappa (QWK) metric. The experimental results demonstrate the effectiveness of this approach in comparison with several existing solutions in literature.