
Due to the increasing availability of Augmented Reality (AR) devices, it is now possible to 3D sketch over physical objects. However, users still encounter challenges that affect their stroke accuracy and drawing experience when sketching in AR. This paper presents the design and implementation of the Surface Extension Guides for Augmented Reality (SEGAR) system, a pen-based user interface that allows users to extend the physical surfaces to mid-air while sketching. SEGAR provides users with visual and haptic feedback to help them seamlessly transition between sketching on physical surfaces and mid-air. We ran a pilot study, where four participants sketched a cube in mid-air using SEGAR. Our results show that participants found the system useful.
Extended-reality (XR) training with generative AI could give athletes instant coaching, but current systems may produce feedback that diverges from expert practice. This work aims to bridge this in sabre fencing by giving an AI agent a set of tools for analyzing athletes' performances. We designed these tools based on a formative study involving specialized literature and interviews with three elite coaches. We evaluated the impact of pose analysis tools in AI-generated feedback through an expert evaluation with a two-alternative forced-choice task, comparing a tool-augmented vision-language model with a baseline without tools. Surprisingly, the tool-augmented agent was preferred in 40% of cases (p = 0.15), indicating no significant advantage at this sample size. Coaches valued its richer kinematic insights but noted limitations in capturing intentional tactical deviations. Our study contributes to design considerations that can inform the next generation of AI fencing coaches in XR that are aligned with human coaching practices.
Agility ladder training is commonly used in fitness and therapy to improve coordination in lower-limb exercises by training precise foot placement. Previous research adapted this training approach to immersive environments and emphasized both the potential and the need for enhanced visual feedback in virtual reality (VR) to improve training outcomes. However, effective visual guidance for such tasks in VR remains underexplored, as it is unclear how spatial user interfaces (UIs) should be designed to support accurate foot placement. To address this gap, we conducted a within-subject study with 40 participants, investigating the effects of two visualization techniques: a Foot-Aligned UI (exocentric, attached to the feet) and a Head-Aligned UI (egocentric, floating in view); combined with color-coded performance feedback. Foot positioning accuracy, rotational control, success rate, and perceived workload were measured during VR-based agility ladder tasks. Results show that the Foot-Aligned UI significantly improved foot placement success rates without increasing cognitive load, compared to the Head-Aligned UI and No UI conditions, which was supported by qualitative feedback. In contrast, color-coded step feedback was perceived as helpful but showed no measurable performance benefit. Based on these findings, we derive design recommendations for spatial UIs to support lower-limb motor training in VR.
We present Flick-in, a Japanese text entry method for indirect touch on a smartwatch. Indirect touch is performed without looking at the input surface, which makes it difficult to touch down accurately at the correct location. This difficulty limits the usability of conventional Japanese text entry methods, which require visual confirmation of the input surface. In contrast, few Japanese text entry methods have been proposed specifically for indirect touch. In Flick-in, users first select a vowel using a bezel-initiated swipe, followed by selecting a consonant using a touch-up gesture while observing the output surface. This makes Japanese text entry with indirect touch feasible. We conducted two studies to evaluate the performance of Flick-in using an external display and a mixed reality (MR) environment as output surfaces. The results showed a text entry speed of 29.5 CPM with a total error rate (TER) of 10.0% on the external display and 36.5 CPM with a TER of 6.14% in the MR environment.
Despite its centrality to the overall text entry experience, the task of editing text in virtual reality has received limited attention in the literature. In this paper, we focus on the task of editing text in virtual reality and explore the opportunities afforded by Large Language Models (LLMs) in supporting this task. We specifically investigate how an LLM can be employed to mediate free-form speech-based text editing commands given by users. This flexible approach is inspired by the concept of shared control and allows users to potentially focus less on the specific edit operations required, and more on the desired outcome. However, such flexibility in the user commands can introduce potential ambiguity regarding the desired target of the edit. We therefore study three different interaction conditions representing promising alternative methods for expressing the intended location of an edit: (i) directly via voice commands; (ii) by highlighting words in the text field with a touchbased interaction; and (iii) implicitly via gaze fixations. We find that participants develop effective strategies for collaborating with the LLM-based agent to make edits but also appreciate the more explicit interaction based on touch for expressing edit context.
Augmented reality displays are evolving into devices that can seamlessly integrate virtual content into users' physical environments in a context-sensitive manner, promising to support users more fluidly in their tasks. However, dynamically adapting applications to the user's environment presents challenges that require a real-time understanding of the user's current context. While recent computer vision approaches for context detection have made progress, they often lack a semantic understanding of the interplay between the user, their surroundings, and the system state. In this paper, we investigate how context-aware application recommendations affects user experience and task performance in AR. To this end, we evaluate a real-time system that uses a large multimodal model integrated with a Microsoft HoloLens 2. The system processes application metadata and front-facing camera images to recommend contextually relevant apps. In a user study, we compared three interaction modes for selecting apps: a fully automatic mode that proactively switches to the suggested app, a recommendation-based mode that highlights a suggested app without launching it, and a manual mode requiring user selection. Our study showed that automatic context-aware app switching significantly improved task efficiency and reduced cognitive load compared to the other modes, without diminishing users' sense of control. Participants preferred the automated mode, which enabled smoother workflows and enhanced usability.
Narrow field of view (FOV) of augmented reality head-mounted displays (AR-HMDs) can challenge user interaction, as AR content is often positioned off-screen, necessitating frequent and potentially cumbersome head movements for interactions. To address this limitation, we propose extending the interactable workspace into peripheral areas beyond the narrow FOV. Although the visual FOV is limited, AR-HMDs track hands across a wider area, which current strategies for guiding user interactions to off-screen objects underutilize. We compare three methods that utilize the full area for facilitating interaction with off-screen objects, thereby extending the AR workspace. Our evaluation focuses on the methods' feasibility in terms of accuracy, speed, and user preference. Results from two studies suggest that these methods successfully aid users by extending the interactable workspace by enabling off-screen interactions, mitigating some limitations of narrow FOV displays.
This paper investigates how designers envision using virtual reality as a spatial interaction tool for collecting user insights during product development. We developed KitchXR, a virtual kitchen prototype, to elicit expectations from 17 professionals in the home appliance industry, including UX/UI designers, industrial designers, and consumer insight specialists via an expert evaluation study. Through thematic analysis of semi-structured interviews, we identify key challenges in current UX research workflows, particularly limited access to user data and constraints on the timing of insight collection. Our findings highlight 3 opportunity areas for VR-based user insight tools: (1) enabling complex and context-rich testing scenarios, (2) integrating bio-sensory and behavioral data, and (3) supporting the revisitation of early design decisions. By centering designers as users of spatial interaction tools, this study contributes to emerging discourse at the intersection of extended reality, user research, and collaborative design workflows.
In Handheld augmented reality (AR), selecting targets behind the user often requires considerable body or device rotation, increasing selection time and physical effort. In this paper, we present using a smartphone's front camera to enable rear-target selection with minimal body or device movement. This approach supports interaction even when turning around is difficult and can be easily integrated into existing AR applications via a simple camera-switching function. The results of our two user studies show that our method enabled faster and more accurate selection of rear targets, although using both the rear and front cameras introduced confusion due to differences in their operability, leading to increased operation time. Despite this, the fact that most participants chose to use the front camera highlights its potential for supporting 360-degree interaction in Handheld AR.
Smart textiles embedded with capacitive touch sensors offer significant potential for intuitive gesture-based interaction, yet recognizing complex gestures on resource-constrained wearable devices remains challenging. This paper presents a minimalist neural network architecture specifically optimized for knitted capacitive touch interfaces. Our approach efficiently recognizes single and multi-touch gestures including taps, swipes, and pinches with accuracy exceeding 90% on training data and 80% on testing data. To assess real-world usability, we conducted a comparative user study evaluating participant performance with our knitted interface against a conventional trackpad in a gesture-controlled gaming scenario. Results demonstrated comparable overall performance between both interfaces, with participants achieving similar game scores despite the novelty of the textile interface. Statistical analysis revealed rapid user adaptation to the textile interface, with performance stabilizing after initial trials while the standard trackpad showed continuous improvement throughout testing. Quantitative metrics were supplemented by qualitative feedback highlighting the comfort and tactile appeal of the textile interface. This work advances the practical deployment of smart textile gesture recognition systems by addressing both technical performance requirements and usability considerations for real-world applications.
This paper examines how virtual reality (VR) can be used to create customizable optokinetic drums for motion sickness (MS) training. In a study with 11 participants, we tested how stripe thickness and rotation speed affect gaze patterns and MS symptoms. Results showed that thicker rotating stripes replicated the gaze behaviour of traditional horizontal drums while reducing discomfort. The findings provide design guidelines for developing more effective and comfortable VR-based optokinetic experiences.
Material sustainability is a topic of growing interest in the field of Human Computer Interaction (HCI). In this interactive artwork series Unraveling, we provide an example of how we can use regular activities in our spatial environment to deconstruct materials for reuse. For this demonstration, we designed machine-knit panels for deconstruction using digital fabrication, and then created a corresponding unraveling machine that responds to sounds in the vicinity to unravel the knits. During the conference, attendees talking (and any other ambient noises) will cause a knit panel to unravel. With this interactive demonstration, we aim to encourage discussion on how computing can support sustainable computational fabrication where we can continually create and unmake for reuse.
This paper presents a wrist-based spatial gesture interface enabling simultaneous control of multiple avatar dance parameters in real time. Using quaternion data from an IMU sensor worn on the back of the hand, the system maps wrist roll (qz) to turning direction and upper-body rotation, and wrist pitch (qx) to hip inclination and animation speed. This design leverages the wrist's high degrees of freedom compared to other hand joints, allowing expressive multidimensional control with minimal physical effort. A three-part user study evaluated (1) trajectory reproduction while dancing, (2) free control to fast-tempo music, and (3) free control to slow-tempo music. Results show that participants naturally exploited the mapping to synchronize movement parameters with music tempo, producing distinct control patterns for different BPM contexts. The findings highlight how spatial wrist gestures can serve as an efficient and expressive control modality for embodied interaction in XR environments, offering both precision for guided tasks and freedom for creative performance.
I-IFA is a user-friendly avatar motion authoring framework that uses a single IMU[3, 5] and a preconfigured dataset to provide finegrained control of distal segments in the kinematic chain. Unlike keyframe tools and full mocap-which are costly and hard to personalize our method supports real time, intention driven manipulation to boost immersion and interaction diversity in XR. Targeting complex dances, I-IFA employs a minimal hip-kneeankle hierarchy that preserves naturalness while enabling user intent edits. The goal is real time control under ultralight sensing, allowing immersive interaction with virtual humans. By improving personalization, responsiveness, and efficiency, I-IFA offers a practical, accessible UI for intention-aware motion control, bridging interaction design and immersive avatar manipulation in XR.
Low-vision (LV) individuals often face challenges with visual search and object recognition due to partial loss of visual function. While rehabilitation centers help clients develop compensatory strategies using residual vision, interpreting clients' visual behavior, especially in spatially rich contexts or when using head movements for eccentric viewing, can be difficult. This paper explores the use of co-located shared augmented reality (S-AR) with projected gaze cues (eye gaze, head gaze, eye trace, and field of view) to support LV therapist (LVT)-guided training. Built on the Microsoft HoloLens 2, our system projects these cues in real-time within an S-AR environment. Grounded in a field observation and discussions with two certified LVTs, we designed an AR-based visual search task inspired by current rehabilitation practices. Through an exploratory study with nine LV clients and LVTs, we assessed the potential of gaze projections to support guided training in the future. Our findings highlight the system's potential to enhance LVTs' understanding of client strategies, engage clients in training tasks, and facilitate more personalized guidance, while pointing out areas for further refinement in cue interpretability and system adaptability.
Augmented Reality (AR) tools have the capacity to enhance humanrobot collaboration, such as complex navigation tasks, by providing more immersive environmental cues and enhanced operational control information. While AR has recently demonstrated promise with this concept, controlling more complicated, versatile quadrupedal robots remains underexplored despite their potential applications in dangerous environments, laser scanning, and addressing monotonous tasks. This study explores the impact of AR interfaces on controlling a Boston Dynamics Spot quadrupedal robot during complex navigation tasks. We conducted an empirical study with 33 non-expert participants and two experts performed three navigational routes comparing AR interfaces and controls to more traditional tablet-based approaches. Our findings indicate that while task duration increased, the AR system performed comparably to a conventional tablet approach in terms of usability and trust, and offered more desirable increased robot movement smoothness. This study provides a foundation for leveraging AR interfaces in advancing robot-assisted navigation to enable a broader range of users to effectively operate sophisticated quadrupedal robotic systems and accelerate the adoption and trust of these systems in operation.
Many mixed reality applications rely on computer vision to perform operations like object tracking, identification, and registration. While platform-specific solutions like OpenCV for Unity and ARKit exist, there is currently no cross-platform software library for performing these tasks across most consumer mixed reality hardware. Here, we present an open-source, cross-platform Unity library that facilitates common computer vision workflows within mixed reality applications. The package is written primarily in C++ using OpenCV as a backend, with APIs in C and C# to maximize compatibility. The library is modular, and new modules can be easily created to extend its functionality. The source code is available on GitHub at https://github.com/surreality-lab/SurrealityCV under the MIT license.
Spaces shared by Mixed Reality (MR) users and bystanders (nonMR users) pose unique challenges because bystanders cannot see the virtual interfaces that the MR user is interacting with. As a result, they may unknowingly occlude these interfaces, leading to physical-virtual conflicts that disrupt the MR user's experience. To address this issue, we explore how projecting shadows of virtual interfaces onto the floor can enhance bystander awareness. We conducted a user study (N = 20) in which a bystander performed an independent task in the same space as an MR user. The bystander experienced one of three visualization conditions - none (no projection), dynamic (shadows appear only upon collision),and always-on (shadows of all interfaces are continuously visible). Our findings show that always-on significantly mitigated physicalvirtual conflicts and was most preferred by participants. In contrast, the dynamic condition was found to be distracting, while none required extra communication and led to more interference. Our results highlight the value of persistent, easily perceived cues for bridging the perceptual divide between MR users and bystanders in shared physical spaces, ultimately promoting more seamless integration of MR systems into everyday environments.
We present a real-time visual feedback system that overlays dynamic circular indicators of sEMG activity on target and compensatory muscle locations. The indicators change in angle and color according to muscle activation, aiming to promote target muscle engagement and suppress compensatory muscle overuse. A user study with nine participants (beginners, intermediates, advanced) performing lateral raises measured the signal-to-noise ratio between target and compensatory muscles. Beginners showed the largest improvement and retention, intermediates had smaller gains with some reliance on feedback, and advanced participants showed minimal change due to a ceiling effect. These results provide preliminary evidence of the system's effectiveness.
Prolonged sedentary behavior is a health concern in screen-based work environments. Mixed Reality (MR) offers opportunities to subtly influence posture and physical activity through virtual screens. This study explores using slow, imperceptible screen motion in MR to prompt movement, such as sit-to-stand transitions and lateral displacement, without affecting task performance. We evaluated three conditions: a static baseline, a vertically moving screen for sit-to-stand, and a horizontally shifting screen for lateral movement while standing. Participants performed Fitts' Law and Stroop tasks in each condition, with measurements on physical displacement, performance, workload, and usability.