While many digital distractions can be managed, real-world interruptions, such as phone calls, notifications, and office noise, are harder to control and can harm productivity, well-being, and learning. Mixed reality systems like Augmented Reality (AR) are often described as immersive-a property which might protect users from such disruptions. We tested this assumption by comparing a head-mounted AR interface that overlays digital annotations on physical objects with a traditional flat screen during vocabulary learning under common office distractions. In a user study (n = 32), AR users reported feeling less distracted and recalled less task-irrelevant information, but their learning performance did not improve. Instead, distraction-related performance decline was greater in AR. Physiological and self-report measures showed no reduction in effort or workload, and participants with higher auditory distractibility did not benefit. Overall, AR annotation alone may not sufficiently shield learners from real-world distractions, motivating new design approaches.
The “keyword method” is a mnemonic technique often used in vocabulary learning, in which a target word is linked to a familiar, phonetically similar keyword through a vivid mental image. For example, the Japanese word for “tree” is “ki”, which sounds like “key” (keyword), so a learner might imagine “a tree with key-shaped leaves” (association). Research in non-contextualised settings on screen has shown that externalising personalised associations as images enhanced recall while Augmented Reality (AR) further strengthened retention by anchoring learning to the real-world context. However, prior work on externalised associations was conducted in linear, non-contextualised workflow without opportunities to revisit studied content, whereas existing AR studies have focused only on predefined (non-personalised) keyword-associations. It therefore remains unclear how predefined and personalised approaches compare in contextualised AR learning, where learners can interact with and revisit previously encountered content. To explore this, we developed ARText, an AR system that annotates real-world objects with target words, keywords, and text-to-image-generated visual representation of association. Participants used the system in both the personalised condition (keyword-associations and visual representations created by users) and the predefined condition (all designed by experts). Our findings show that personalising keyword-associations and visual representations reduced engagement with the learning content, evidenced by fewer revisits and shorter viewing times. Participants also preferred predefined keyword associations, showed better immediate and delayed recall, and achieved higher learning efficiency. This results suggest that, for novice learners in short AR learning sessions, reducing cognitive load and sustaining engagement through predefined keyword associations may be more important than personalisation alone. We discuss possible reasons for these outcomes and their implications for designing future AR-based vocabulary learning systems.
Radar-based gesture recognition has emerged as a promising approach for unobtrusive interaction. Unlike camera-based systems, radar sensors can detect gestures through opaque materials, enabling seamless embedding into various everyday objects. However, it remains unclear how to train models efficiently for robust gesture recognition through diverse materials. To investigate this, we collected a dataset of 17,520 gesture recordings performed through 73 everyday materials. By comparing several material-sampling and data-augmentation strategies, we found that a small carefully selected representative subset of training materials was sufficient to match the performance of a classifier trained on the full material dataset. Our results showed that the models trained on 14 quota-sampled materials achieved accuracies of 95.8% and 91.2%, comparable to training on all 73 materials (96.8% and 91.6%) and significantly better than training without material data (66.8% and 65.8%). Among the evaluated sampling approaches, Quota sampling also provided the best overall trade-off between performance and practicality. In contrast, classifiers trained on augmented data performed worse than those trained on actual material-specific data. Taken together, these findings indicate that, for the tested sensor, gesture set, and material collection, carefully selected real-material data offer a practical route to reducing material-specific data collection in radar-based gesture recognition while preserving generalisation. Code, models, and data are available in the public repository, with additional details provided in the supplementary materials: https://gitlab.com/hicuplab/seeing-through .
Augmented reality (AR) magic-lens (ML) displays, such as handheld devices, offer a convenient and accessible way to enrich our environment using virtual imagery. Several display technologies, including conventional monocular, less common stereoscopic, and varifocal displays, are currently being used. Vergence and accommodation effects on depth perception, as well as vergence-accommodation conflict, have been studied, where users interact only with the content on the display. However, little research exists on how vergence and accommodation influence user performance and cognitive-task load when users interact with the content on a display and its surroundings in a short timeframe. Examples of this are validating augmented instructions before making an incision and performing general hand-eye coordinated tasks such as grasping augmented objects. To improve interactions with future AR displays in such scenarios, we must improve our understanding of this influence. To this end, we conducted two fundamental visual-acuity user studies with 28 and 27 participants, while investigating eye vergence and accommodation distances on four ML displays. Our findings show that minimizing the accommodation difference between the display and its surroundings is crucial when the gaze between the display and its surroundings shifts rapidly. Minimizing the difference in vergence is more important when viewing the display and its surroundings as a single context without shifting the gaze. Interestingly, the vergence-accommodation conflict did not significantly affect the cognitive-task load nor play a pivotal role in the accuracy of interactions with AR ML content and its physical surroundings.
Heavy goods vehicles (HGVs) have a significant impact on road and bridge infrastructure, with overloaded vehicles accelerating structural deterioration and increasing safety risks. Bridge weigh‐in‐motion (B‐WIM) systems estimate gross vehicle weight (GVW) using strain measurements, but inaccuracies in axle configuration recognition can reduce reliability. This study presents a low‐cost computer vision (CV) extension for existing B‐WIM installations that verifies strain‐inferred axle configurations using traffic camera images and flags GVW estimates as reliable or unreliable. Experiments on a data set of over 30,000 HGV records show that by combining convolutional neural networks with strain‐based heuristics, GVW reliability can improve from 96.7% to 99.89%, effectively excluding nearly all erroneous measurements. The approach operates without interrupting ongoing B‐WIM operations and can be applied retrospectively to historical data. Limitations include the inability to detect raised axles (RAs), which the method excludes as unreliable. This method provides a practical, high‐precision enhancement for structural health monitoring of bridges.
Although deformable objects are not typically designed for digital interaction, they offer a largely unexplored potential-any such object could be repurposed as a medium for controlling digital content. While existing approaches embed sensors into deformable objects to enable interaction, this limits scalability and practicality of such systems. An alternative is to perform gesture recognition on deformable objects using a wrist-worn radar sensor. However, when analysing reflected radar signals it is difficult to separate reflections originating from the continues deformations of the object shape and those from the user's hand and fingers. Additionally, the continuous shape changes of deformable objects introduce changes in radar cross-section, affecting signal variability. Furthermore, user ergonomics-such as variations in hand size, finger dexterity, and strength-are likely to influence the degree of object deformation during interaction. In this paper, we explore whether radar sensing can be used for robust gesture detection on deformable objects, focusing on how well does a system generalize to previously unseen users and what can we do to improve such generalisability. In pursuit of this goal, we record a dataset of 4.3k labelled gestures with Google Soli millimeter-wave radar sensor on a plush toy and demonstrates robust classification performance, achieving accuracy of up to 90% on a five-gesture set. Furthermore, we investigate model generalizability and show that transfer learning improves recognition for previously unseen users, yielding performance gains of up to 20%. These findings highlight the potential of radar-based sensing for spontaneous and practical interaction with deformable objects.
The 'keyword method' is an effective technique for learning vocabulary of a foreign language. It involves creating a memorable visual link between what a word means and what its pronunciation in a foreign language sounds like in the learner's native language. However, these memorable visual links remain implicit in the people's mind and are not easy to remember for a large set of words. To enhance the memorisation and recall of the vocabulary, we developed an application that combines the keyword method with text-to-image generators to externalise the memorable visual links into visuals. These visuals represent additional stimuli during the memorisation process. To explore the effectiveness of this approach we first run a pilot study to investigate how difficult it is to externalise the descriptions of mental visualisations of memorable links, by asking participants to write them down. We used these descriptions as prompts for text-to-image generator (DALL-E2) to convert them into images and asked participants to select their favourites. Next, we compared different text-to-image generators (DALL-E2, Midjourney, Stable and Latent Diffusion) to evaluate the perceived quality of the generated images by each. Despite heterogeneous results, participants mostly preferred images generated by DALL-E2, which was used also for the final study. In this study, we investigated whether providing such images enhances the retention of vocabulary being learned, compared to the keyword method only. Our results indicate that people did not encounter difficulties describing their visualisations of memorable links and that providing corresponding images significantly improves memory retention.
We examined eye and head movements to gain insights into skill development in clinical settings. A total of 24 practitioners participated in simulated baby delivery training sessions. We calculated key metrics, including pupillary response rate, fixation duration, or angular velocity. Our findings indicate that eye and head tracking can effectively differentiate between trained and untrained practitioners, particularly during labor tasks. For example, head-related features achieved an F1 score of 0.85 and AUC of 0.86, whereas pupil-related features achieved F1 score of 0.77 and AUC of 0.85. The results lay the groundwork for computational models that support implicit skill assessment and training in clinical settings by using commodity eye-tracking glasses as a complementary device to more traditional evaluation methods such as subjective scores.
Cross-reality (XR) systems facilitate interaction between devices with differing levels of virtual content. By engaging with a variety of such devices, XR systems offer the flexibility to choose the most suitable modality for specific task or context. This capability enables rich applications in training and education, including vocabulary learning. Vocabulary acquisition is a vital part of language learning, employing techniques such as words rehearsing, flashcards, labelling environments with post-it notes, and mnemonic strategies such as the keyword method. Traditional mnemonics typically rely on visual stimuli or mental visualisations. Recent research highlights that AR can enhance vocabulary learning by combining real objects with augmented stimuli such as in labelling environments. Additionally,advancements in generative AI now enable high-quality, synthetically generated images from text descriptions, facilitating externalisation of personalised visual stimuli of mental visualisations. However, creating interfaces for effective real-world augmentation remains challenging, particularly given the limited text input capabilities of Head-Mounted Displays (HMDs). This work presents an XR system that combines smartphones and HMDs by leveraging Augmented Reality (AR) for contextually relevant information and a smartphone for efficient text input. The system enables users to visually annotate objects with personalised images of keyword associations generated with DALL-E 2. To evaluate the system, we conducted a user study with 16 university graduate students, assessing both usability and overall user experience.
Improvisation is an important skill in music instrument learning, but remains a less-taught topic in traditional piano education. To improvise effectively, learners must develop musical vocabulary, creative confidence, and comfort in performance. These demands make piano improvisation a complex teaching challenge where technology interventions may offer support. Prior short-term studies on augmented piano roll visualisations have shown promise for teaching sight-reading and motor coordination in novice students. However, how such approaches can support advanced learners in acquisition of improvisational skills remains under-explored. To address this gap, we present ImproVisAR, an interactive piano training system that teaches improvisation through augmented piano roll visualisations. Concepts and tools derived from a co-design process with improvisation experts are integrated as structured learning modes. We validated the system through a four-day controlled study ( n=6 ) comparing an AR-based condition with a traditional sheet music condition following a mixed-methods approach to data analysis. We collected and analysed subjective ratings of cognitive load, creativity support, user-experience, expert evaluation of performances, interaction logs, and qualitative insights collected from daily post-study interviews. Our findings show that participants experienced reduced cognitive load over time, sustained engagement across sessions, and AR participants showed higher expert-rated scores, particularly in rhythm, flow, musicality and overall musical impression. Participants also reported greater immersion, freedom to create musical content and motivation to continue playing. We discuss these findings in relation to user experience and creativity support, and offer design recommendations for AR systems that aim to teach complex, expressive skills such as musical improvisation.
Eye gaze is considered an important indicator for understanding and predicting user behaviour, as well as directing their attention across various domains including advertisement design, human-computer interaction and film viewing. In this paper, we present a novel method to enhance the analysis of user behaviour and attention by (i) augmenting video streams with automatically annotating and labelling areas of interest (AOIs), and (ii) integrating AOIs with collected eye gaze and fixation data. The tool provides key features such as time to first fixation, dwell time, and frequency of AOI revisits. By incorporating the YOLOv8 object tracking algorithm, the tool supports over 600 different object classes, providing a comprehensive set for a variety of video streams. This tool will be made available as open-source software, thereby contributing to broader research and development efforts in the field.
Robots have the potential to enhance teaching of advanced computer science topics, making abstract concepts more tangible and interactive. In this paper, we present Timmy-a GoPiGo robot augmented with projections to demonstrate shortest path algorithms in an interactive learning environment. We integrated a JavaScript-based application that is projected around the robot, which allows users to construct graphs and visualise three different shortest path algorithms with colour-coded edges and vertices. Animated graph exploration and traversal are augmented by robot movements. To evaluate Timmy, we conducted two user studies. An initial study (n=10) to explore the feasibility of this type of teaching where participants were just observing both robot-synced and the on-screen-only visualisations. And a pilot study (n=6) where participants actively interacted with the system, constructed graphs and selected desired algorithms. In both studies we investigated the preferences towards the system and not the teaching outcome. Initial findings suggest that robots offer an engaging tool for teaching advanced algorithmic concepts, but highlight the need for further methodological refinements and larger-scale studies to fully evaluate their effectiveness.