Automated scoring models are increasingly used to assign rubric-based quality ratings to complex language performances, including classroom transcripts, yet they typically provide little insight into why a particular score is produced. We propose a general framework for sentence-level interpretability of rubric-based scoring that combines model-agnostic Shapley-value attributions with rationales generated by large language models (LLMs). Instantiated on the Quality of Feedback dimension of the CLASS framework using the NCTE corpus, the framework enables systematic comparison of fine-tuned pretrained language models (PLMs) and prompted LLMs on both scoring performance and explanation faithfulness. Across 6k annotated transcript segments, fine-tuned PLMs outperform LLMs in prediction accuracy but exhibit label compression toward mid-scale scores. Deletion-based tests show that SHAP identifies sentences that reliably drive model predictions, producing typically larger and more coherent prediction shifts than LLM-generated rationales. Cross-model analyses further reveal that SHAP attributions transfer robustly across architectures, whereas LLM rationales exert limited and inconsistent influence. Overall, the findings demonstrate that SHAP provides more faithful and transferable explanations for rubric-based scoring, and that the proposed framework offers a principled basis for evaluating both scoring models and their explanations in high-stakes educational settings and other rubric-based language assessment tasks.
Background: Teaching quality is commonly assessed through classroom observation. However, observer ratings of teaching quality frequently exhibit limited psychometric quality. Beyond evaluating the ratings themselves, exploring the observation process and the design of the video environment may offer valuable insights into conditions that enhance rating accuracy. Aims: This study explored the classroom observation process of preservice teachers using eye-tracking technology and examined how their gaze behavior relates to the accuracy of their teaching-quality ratings. We also investigated the impact of different video environments by comparing traditional classroom videos presented on computer screens with immersive 360-degree classroom videos presented on virtual-reality headsets. Sample: N = 75 preservice teachers participated in a controlled lab study. Method: Each participant observed two randomly assigned mathematics classroom videos-one in a traditional screen-based and one in an immersive video environment. Eye trackers recorded gaze behavior during critical classroom events. Participants rated the quality of the observed teaching after the video observations. Results: Overall, observers adjusted their gaze according to the focus of critical events (teacher-or student-focused). In immersive 360-degree videos, they showed a stronger visual focus on the teacher. Both visual focus of attention and cognitive arousal during critical events, indicated by a larger pupil diameter, predicted the accuracy of teaching-quality ratings, with gaze being a stronger predictor of rating accuracy in immersive compared to screen-based videos. Conclusions: The findings indicate that gaze behavior offers actionable insights into classroom observation processes and informs the design of observation environments in both research and teacher education.
Artificial intelligence is increasingly used in hiring, raising concerns about how applicants perceive these systems. While prior work on algorithmic fairness has emphasized technical bias mitigation, little is known about how avatar identity cues influence applicants' justice attributions in an interview context. We conducted a crowdsourcing study with 215 participants who completed an interview with photorealistic AI avatars varied in phenotypic traits (race and sex), followed by a standardized rejection. Using self-reports, sentiment analysis, and eye tracking, we measured perceptions of trust, fairness, and bias. Results show that racial mismatch heightened perceptions of ethnic bias, while partial match (sharing only one identity) reduced fairness judgments compared to both full and no match. This work extends the Computers-Are-Social-Actors paradigm by demonstrating that avatar appearances shape justicerelated evaluations of AI. We contribute to HCI by revealing how identity cues influence fairness attributions and offer actionable insights for designing equitable AI interview systems.
Virtual reality (VR) offers promising opportunities for procedural learning, particularly in preserving intangible cultural heritage. Advances in generative artificial intelligence (Gen-AI) further enrich these experiences by enabling adaptive learning pathways. However, evaluating such adaptive systems using traditional temporal metrics remains challenging due to the inherent variability in Gen-AI response times. To address this, our study employs multimodal behavioural metrics, including visual attention, physical exploratory behaviour, and verbal interaction, to assess user engagement in an adaptive VR environment. In a controlled experiment with (n = 54) participants, we compared three levels of adaptivity (high, moderate, and non-adaptive baseline) within a Neapolitan pizza-making VR experience. Results show that moderate adaptivity optimally enhances user engagement, significantly reducing unnecessary exploratory behaviour and increasing focused visual attention on the AI avatar. Our findings suggest that a balanced level of adaptive AI provides the most effective user support, offering practical design recommendations for future adaptive educational technologies.
BACKGROUND:Existing conceptions of teaching quality assume that classroom interactions serve as the foundation for effective teaching. The resulting data necessitates analytical approaches capable of extracting the semantics of these interactions. AIM:This study investigates whether and to what extent lesson semantics provide insights into teaching quality (i.e., cognitive engagement, encouragement and warmth, multiple approaches, and the nature of discourse). To achieve this, GPT-4 was applied as a tool for analysing lesson transcripts. SAMPLE:The study is based on data from the TALIS Video study, which included N = 50 teachers delivering two consecutive mathematics lessons in 9th grade. Teaching quality was annotated by trained observers across multiple dimensions. METHOD:The analysis involved embedding segmented lesson transcripts to examine their semantic characteristics and associations with human annotations of teaching quality. Additionally, we applied content-informed prompting to evaluate the interpretability of semantic characteristics for the considered dimensions. RESULTS:GPT-4 identified five distinct semantic representations of transcripts, varying at both the teacher and lesson levels. These representations were related to teaching quality, accounting for up to 20% of variance in teaching quality annotations. Content-informed prompting aligned lesson segments more closely with semantic representations, supporting their interpretability. CONCLUSION:The findings suggest that lesson semantics serve as indicators of teaching quality, offering a promising approach to understanding effective classroom learning.
This article provides a step-by-step guideline for measuring and analyzing visual attention in 3D virtual reality (VR) environments based on eye-tracking data. We propose a solution to the challenges of obtaining relevant eye-tracking information in a dynamic 3D virtual environment and calculating interpretable indicators of learning and social behavior. With a method called ''gaze-ray casting,'' we simulated 3D-gaze movements to obtain information about the gazed objects. This information was used to create graphical models of visual attention, establishing attention networks. These networks represented participants' gaze transitions between different entities in the VR environment over time. Measures of centrality, distribution, and interconnectedness of the networks were calculated to describe the network structure. The measures, derived from graph theory, allowed for statistical inference testing and the interpretation of participants' visual attention in 3D VR environments. Our method provides useful insights when analyzing students’ learning in a VR classroom, as reported in a corresponding evaluation article with N = 274 participants. • Guidelines on implementing gaze-ray casting in VR using the Unreal Engine and the HTC VIVE Pro Eye. • Creating gaze-based attention networks and analyzing their network structure. • Implementation tutorials and the Open Source software code are provided via OSF: https://osf.io/pxjrc/?view_only=1b6da45eb93e4f9eb7a138697b941198.
Background: Studies have shown that the effectiveness of virtual reality (VR) can be improved by using learning prompts (e.g., summary prompts, pretraining, and priming). However, less is known about how such learning prompts affect students' processing of the learning content within the virtual learning environment.Aims: This study aimed to shed light on students' processing of learning content in VR by analyzing their eye -movements during a virtual biology lesson after they had been primed to perceive VR as useful for learning or for daily life.Sample: Participants were 171 German students in the 10th grade from ten high schools.Methods: Participants were randomly assigned to a learning-usefulness condition (i.e., priming video highlighting VR's usefulness for learning), a daily-life-usefulness condition (i.e., priming video highlighting VR's usefulness in daily life), or a control condition (i.e., no priming video). We integrated the priming interventions at the beginning of the virtual lesson and used VR headsets with integrated eye tracking to monitor students' eye movements during the lesson. Results: When compared with the control condition, the learning-usefulness intervention positively affected students' fixation duration and saccade amplitude, which in turn were positively related to students' learning achievement. Regarding students' learning achievement, the daily-life-usefulness condition did not differ from the other two conditions. There was no difference between the three conditions in terms of interest in the virtual lesson.Conclusions: These results indicate that a brief priming intervention about VR's usefulness for learning can induce appropriate cognitive processing, which can positively affect students' learning achievement.
This paper explores gaze entropy as a metric for detecting classroom discourse events in a virtual reality (VR) classroom. Using data from a laboratory experiment with N = 240 secondary school students, we distinguished between events of teacher-centered classroom discourse (question, hand raising, answer) and teacher explanation by analyzing their transition and stationary gaze entropy. Employing multi-level regression models, both entropy measures effectively discriminated between the two events and distinguished different levels of classroom participation as indicated by the degree of hand-raising by virtual students. Furthermore, using both measures in a logistic regression model, the potential of gaze entropy could be demonstrated by predicting the two events with 67% accuracy. By analyzing transition and stationary entropy, the study attempts to uncover different gaze patterns associated with learning events in a virtual classroom. The results contribute to the research and development of VR scenarios that help to simulate effective learning environments.
Mental rotation is the ability to rotate mental representations of objects in space. Shepard and Metzler’s shape-matching tasks, frequently used to test mental rotation, involve presenting pictorial representations of 3D objects. This stimulus material has raised questions regarding the ecological validity of the test for mental rotation with actual visual 3D objects. To systematically investigate differences in mental rotation with pictorial and visual stimuli, we compared data of $$N=54$$ N = 54 university students from a virtual reality experiment. Comparing both conditions within subjects, we found higher accuracy and faster reaction times for 3D visual figures. We expected eye tracking to reveal differences in participants’ stimulus processing and mental rotation strategies induced by the visual differences. We statistically compared fixations (locations), saccades (directions), pupil changes, and head movements. Supplementary Shapley values of a Gradient Boosting Decision Tree algorithm were analyzed, which correctly classified the two conditions using eye and head movements. The results indicated that with visual 3D figures, the encoding of spatial information was less demanding, and participants may have used egocentric transformations and perspective changes. Moreover, participants showed eye movements associated with more holistic processing for visual 3D figures and more piecemeal processing for pictorial 2D figures.
Higher-achieving peers have repeatedly been found to negatively impact students’ evaluations of their own academic abilities (i.e., Big-Fish-Little-Pond Effect). Building on social comparison theory, this pattern is assumed to result from students comparing themselves to their classmates; however, based on existing research designs, it remains unclear how exactly students make use of social comparison information in the classroom. To determine the extent to which students ( N = 353 sixth graders) actively attend and respond to social comparison information in the form of peers’ achievement-related behaviour, we used eye-tracking data from an immersive virtual reality (IVR) classroom. IVR classrooms offer unprecedented opportunities for psychological classroom research as they allow to integrate authentic classroom scenarios with maximum experimental control. In the present study, we experimentally varied virtual classmates’ achievement-related behaviour (i.e., their hand-raising in response to the teacher’s questions) during instruction, and students’ eye and gaze data showed that they actively processed this social comparison information. Students who attended more to social comparison information (as indicated by more frequent and longer gaze durations at peer learners) had less favourable self-evaluations. We discuss implications for the future use of IVR environments to study behaviours in the classroom and beyond.
Currently, VR technology is increasingly being used in applications to enable immersive yet controlled research settings. One such area of research is expertise assessment, where novel technological approaches to collecting process data, specifically eye tracking, in combination with explainable models, can provide insights into assessing and training novices, as well as fostering expertise development. We present a machine learning approach to predict teacher expertise by leveraging data from an off-the-shelf VR device collected in a VirATec study. By fusing eye-tracking and controller-tracking data, teachers' recognition and handling of disruptive events in the classroom are taken into account or considered. Three classification models were compared, including SVM, Random Forest, and LightGBM, with Random Forest achieving the best ROC-AUC score of 0.768 in predicting teacher expertise. The SHAP approach to model interpretation revealed informative features (e.g., fixations on identified disruptive students) for distinguishing teacher expertise. Our study serves as a pioneering effort in assessing teacher expertise using eye tracking within an interactive virtual setting, paving the way for future research and advancements in the field.
Pupil diameter is a reliable indicator of mental effort, but it must be baseline corrected to account for its idiosyncratic nature. Established methods for measuring baselines cannot be applied in virtual reality (VR) experiments. To reliably measure a pupil diameter baseline in VR, we propose a short testing environment of visual arithmetic tasks. In an experiment with 66 university students, we analyzed external reliability and internal validity criteria for pupil diameter measures during counting and summation tasks. During the counting task, we found a high retest reliability between stimulus intervals. Acceptable retest reliability was found for task repetition at a second measuring time. Analyzing internal validity, we found that pupil diameter increased with task difficulty comparing both tasks. Further, a linear effect was found between the pupil diameter amplitude and luminance levels. Our findings highlight the potential of counting tasks as a pupil diameter baseline for VR experiments.
Immersive virtual reality (IVR) provides great potential to experimentally investigate effects of peers on student learning in class and to strategically deploy virtual peer learners to improve learning. The present study examined how three social-related classroom configurations (i.e., students' position in the classroom, visuali-zation style of virtual avatars, and virtual classmates' performance-related behavior) affect students' visual attention toward information presented in the IVR classroom using a large-scale eye-tracking data set of N = 274 sixth graders. ANOVA results showed that the IVR configurations were systematically associated with differences in learners' visual attention on classmates or the instructional content and their overall gaze distribution in the IVR classroom (Cohen's d ranging from 0.28 to 2.04 for different IVR configurations and gaze features). Gaze-based attention on classmates was negatively related to students' interest in the IVR lesson (d = 0.28); specif-ically, the more boys were among the observed peers, the lower students' situational self-concept (d = 0.24). In turn, gaze-based attention on the instructional content was positively related to students' performance after the IVR lesson (d = 0.26). Implications for the future use of IVR classrooms in educational research and practice are discussed.
Recent developments in computer graphics and hardware technology enable easy access to virtual reality headsets along with integrated eye trackers, leading to mass usage of such devices. The immersive experience provided by virtual reality and the possibility to control environmental factors in virtual setups may soon help to create realistic digital alternatives to conventional classrooms. The importance of such settings has become especially evident during the COVID-19 pandemic, forcing many schools and universities to provide the digital teaching. Researchers foresee that such transformations will continue in the future with virtual worlds becoming an integral part of education. Until now, however, students' behaviors in immersive virtual environments have not been investigated in depth. In this work, we study students' attention by exploiting object-of-interests using eye tracking in different classroom manipulations. More specifically, we varied sitting positions of students, visualization styles of virtual avatars, and hand-raising percentages of peer-learners. Our empirical evidence shows that such manipulations play an important role in students' attention towards virtual peer-learners, instructors, and lecture material. This research may contribute to understanding of how visual attention relates to social dynamics in the virtual classroom, including significant considerations for the design of virtual learning spaces.