Robots are moving beyond industrial settings into creative, educational, and public environments where interaction is open-ended and improvisational. Yet much of human-AI-robot interaction remains framed around performance and efficiency, positioning humans as supervisors rather than collaborators. We propose a re-framing of AI interaction with robots as scaffolding: infrastructure that enables humans to shape robotic behaviour over time while remaining meaningfully in control. Through scenarios from creative practice, learning-by-teaching, and embodied interaction, we illustrate how humans can act as executive directors, defining intent and steering revisions, while AI mediates between human expression and robotic execution. We outline design and evaluation implications that foreground creativity, agency, and flow. Finally, we discuss open challenges in social, scalable, and mission-critical contexts. We invite the community to rethink interacting with Robots and AI not as autonomy, but as sustained support for human creativity.
While many digital distractions can be managed, real-world interruptions, such as phone calls, notifications, and office noise, are harder to control and can harm productivity, well-being, and learning. Mixed reality systems like Augmented Reality (AR) are often described as immersive-a property which might protect users from such disruptions. We tested this assumption by comparing a head-mounted AR interface that overlays digital annotations on physical objects with a traditional flat screen during vocabulary learning under common office distractions. In a user study (n = 32), AR users reported feeling less distracted and recalled less task-irrelevant information, but their learning performance did not improve. Instead, distraction-related performance decline was greater in AR. Physiological and self-report measures showed no reduction in effort or workload, and participants with higher auditory distractibility did not benefit. Overall, AR annotation alone may not sufficiently shield learners from real-world distractions, motivating new design approaches.
The “keyword method” is a mnemonic technique often used in vocabulary learning, in which a target word is linked to a familiar, phonetically similar keyword through a vivid mental image. For example, the Japanese word for “tree” is “ki”, which sounds like “key” (keyword), so a learner might imagine “a tree with key-shaped leaves” (association). Research in non-contextualised settings on screen has shown that externalising personalised associations as images enhanced recall while Augmented Reality (AR) further strengthened retention by anchoring learning to the real-world context. However, prior work on externalised associations was conducted in linear, non-contextualised workflow without opportunities to revisit studied content, whereas existing AR studies have focused only on predefined (non-personalised) keyword-associations. It therefore remains unclear how predefined and personalised approaches compare in contextualised AR learning, where learners can interact with and revisit previously encountered content. To explore this, we developed ARText, an AR system that annotates real-world objects with target words, keywords, and text-to-image-generated visual representation of association. Participants used the system in both the personalised condition (keyword-associations and visual representations created by users) and the predefined condition (all designed by experts). Our findings show that personalising keyword-associations and visual representations reduced engagement with the learning content, evidenced by fewer revisits and shorter viewing times. Participants also preferred predefined keyword associations, showed better immediate and delayed recall, and achieved higher learning efficiency. This results suggest that, for novice learners in short AR learning sessions, reducing cognitive load and sustaining engagement through predefined keyword associations may be more important than personalisation alone. We discuss possible reasons for these outcomes and their implications for designing future AR-based vocabulary learning systems.
In everyday life, physical effort is often minimized and convenience is prioritized, making it difficult for many people to sustain light exercise and stretching despite well-known long-term benefits. This challenge often arises not from objective movement limitations, but from whether an action feels doable in the moment and, therefore worth continuing. This position paper argues that subtle VR hand redirection (HR) can be reframed as a form of cross-sensory support for sustained practice by targeting perceived doability: a moment-to-moment cognitive appraisal that an action is within one's capability while requiring manageable effort. We propose that conservative HR, applied within known perceptual limits, can create repeated micro-success experiences (e.g., reaching a virtual goal earlier with similar physical movement). These micro-successes may increase continuation intention and early re-engagement without relying on overt pressure or intensive coaching. At the same time, such support raises questions about autonomy and authenticity. We therefore articulate two research questions: (RQ1) how HR shifts perceived doability to support sustained practice and positive behavior change; and (RQ2) when HR functions as acceptable support versus becoming counterproductive by undermining authenticity, agency, trust, or fostering dependence. We present an initial sit-and-reach VR prototype, outline a research plan, and identify key design tensions to spark community discussions on autonomy-preserving cross-sensory futures in HCI.
The 'keyword method' is an effective technique for learning vocabulary of a foreign language. It involves creating a memorable visual link between what a word means and what its pronunciation in a foreign language sounds like in the learner's native language. However, these memorable visual links remain implicit in the people's mind and are not easy to remember for a large set of words. To enhance the memorisation and recall of the vocabulary, we developed an application that combines the keyword method with text-to-image generators to externalise the memorable visual links into visuals. These visuals represent additional stimuli during the memorisation process. To explore the effectiveness of this approach we first run a pilot study to investigate how difficult it is to externalise the descriptions of mental visualisations of memorable links, by asking participants to write them down. We used these descriptions as prompts for text-to-image generator (DALL-E2) to convert them into images and asked participants to select their favourites. Next, we compared different text-to-image generators (DALL-E2, Midjourney, Stable and Latent Diffusion) to evaluate the perceived quality of the generated images by each. Despite heterogeneous results, participants mostly preferred images generated by DALL-E2, which was used also for the final study. In this study, we investigated whether providing such images enhances the retention of vocabulary being learned, compared to the keyword method only. Our results indicate that people did not encounter difficulties describing their visualisations of memorable links and that providing corresponding images significantly improves memory retention.
Cross-reality (XR) systems facilitate interaction between devices with differing levels of virtual content. By engaging with a variety of such devices, XR systems offer the flexibility to choose the most suitable modality for specific task or context. This capability enables rich applications in training and education, including vocabulary learning. Vocabulary acquisition is a vital part of language learning, employing techniques such as words rehearsing, flashcards, labelling environments with post-it notes, and mnemonic strategies such as the keyword method. Traditional mnemonics typically rely on visual stimuli or mental visualisations. Recent research highlights that AR can enhance vocabulary learning by combining real objects with augmented stimuli such as in labelling environments. Additionally,advancements in generative AI now enable high-quality, synthetically generated images from text descriptions, facilitating externalisation of personalised visual stimuli of mental visualisations. However, creating interfaces for effective real-world augmentation remains challenging, particularly given the limited text input capabilities of Head-Mounted Displays (HMDs). This work presents an XR system that combines smartphones and HMDs by leveraging Augmented Reality (AR) for contextually relevant information and a smartphone for efficient text input. The system enables users to visually annotate objects with personalised images of keyword associations generated with DALL-E 2. To evaluate the system, we conducted a user study with 16 university graduate students, assessing both usability and overall user experience.
Improvisation is an important skill in music instrument learning, but remains a less-taught topic in traditional piano education. To improvise effectively, learners must develop musical vocabulary, creative confidence, and comfort in performance. These demands make piano improvisation a complex teaching challenge where technology interventions may offer support. Prior short-term studies on augmented piano roll visualisations have shown promise for teaching sight-reading and motor coordination in novice students. However, how such approaches can support advanced learners in acquisition of improvisational skills remains under-explored. To address this gap, we present ImproVisAR, an interactive piano training system that teaches improvisation through augmented piano roll visualisations. Concepts and tools derived from a co-design process with improvisation experts are integrated as structured learning modes. We validated the system through a four-day controlled study ( n=6 ) comparing an AR-based condition with a traditional sheet music condition following a mixed-methods approach to data analysis. We collected and analysed subjective ratings of cognitive load, creativity support, user-experience, expert evaluation of performances, interaction logs, and qualitative insights collected from daily post-study interviews. Our findings show that participants experienced reduced cognitive load over time, sustained engagement across sessions, and AR participants showed higher expert-rated scores, particularly in rhythm, flow, musicality and overall musical impression. Participants also reported greater immersion, freedom to create musical content and motivation to continue playing. We discuss these findings in relation to user experience and creativity support, and offer design recommendations for AR systems that aim to teach complex, expressive skills such as musical improvisation.
While maps provide upfront content, this might not always be the most effective way for users to remember information. With the proliferation of interactive displays for tourists and visitors in public spaces, we can create a more playful user experience with maps than just exploring them. Adding interactions with the map could also help users retain more information as they use them. In this paper, we investigated whether completing a jigsaw puzzle of a map supports users in retaining more information about a specific map. The results of a between-subject study with a sample of n=28 indicate that additional interaction helped improve mean scores of textual and spatial recall but not visual recall. However, the results are not statistically significant, and the topic is subject to further investigation. Our findings contribute to discussions on using interactive touchscreen displays in similar learning scenarios involving memory retention.
In this paper, we describe various technologies that are being used in virtual garment fitting and simulation. There, we have focused about the usage of anthropometry in clothing industry and avatar generation of virtual garment fitting. Most commonly used technologies for avatar generation in virtual environment have been discussed in this paper such as generic body model and laser scanning. Moreover, this paper includes the real-time tracking technologies used in virtual garment fitting like markers and depth cameras in various related researches as well as how the virtual cloth generation and simulation carried out in the related researches. Apart from these, virtual clothing methods such as geometrical, physical and hybrid based models were also discussed in this paper. As ease allowance has a major impact on virtual cloth fitting, it is also considered in this paper related to similar researches. Within this paper, all the above mentioned areas were described thoroughly while stating the existing gap of the virtual garment fitting in online marketplaces.
While virtual reality (VR) has been explored in the field of architecture, its implications on people who experience their future office space in such a way has not been extensively studied. In this explorative study, we are interested in how VR and other representation methods support users in projecting themselves into their future office space and how this might influence their willingness to relocate. In order to compare VR with other representations, we used (i) standard paper based floor plans and renders of the future building (as used by architects to present their creations to stakeholders), (ii) a highly-detailed virtual environment of the same building experienced on a computer monitor (desktop condition), and (iii) the same environment experienced on a head mounted display (VR condition). Participants were randomly assigned to conditions and were instructed to freely explore their representation method for up to 15 min without any restrictions or tasks given. The results show, that compared to other representation methods, VR significantly differed for the sense of presence, user experience and engagement, and that these measures are correlated for this condition only. In virtual environments, users were observed looking at the views through the windows, spent time on terraces between trees, explored the surroundings, and even “took a walk” to work. Nevertheless, the results show that representation method influences the exploration of the future building as users in VR spent significantly more time exploring the environment, and provided more positive comments about the building compared to users in either desktop or paper conditions. We show that VR representation used in our explorative study increased users’ capability to imagine future scenarios involving their future office spaces, better supported them in projecting themselves into these spaces, and positively affected their attitude towards relocating.
Computer programming is a demanding task requiring users to understand a syntax of a programming language, logic flows, and complex abstract concepts. Adult users commonly start learning how to program in a text-based programming environment. However, such tools are not optimal for children as they require understanding of a high level of abstraction. Instead, visual programming languages were developed to hide the syntax and error messages, but are still capable to teach concepts such as parallelism and event handling. Despite, these languages still require children to handle virtual block-like elements on a computer screen. In this research, we examine the feasibility of teaching young children programming using physical blocks coupled with Augmented Reality (AR). To this end, we conducted a between subject design study with 8 participants. We compared a traditional 2D desktop application for visual programming to a 3D setting with tangible physical programming blocks coupled with a Head Mounted Device displaying programming results as AR content. The preliminary results show that the immersive 3D learning environment scored better in mental effort, task completion time, performance and subjective user satisfaction. The results should be validated with a larger sample size.
Learning vocabulary in a primary or secondary language is enhanced when we encounter words in context. This context can be afforded by the place or activity we are engaged with. Existing learning environments include formal learning, mnemonics, flashcards, use of a dictionary or thesaurus, all leading to practice with new words in context. In this work, we propose an enhancement to the language learning process by providing the user with words and learning tools in context, with VocabulARy. VocabulARy visually annotates objects in AR, in the user's surroundings, with the corresponding English (first language) and Japanese (second language) words to enhance the language learning process. In addition to the written and audio description of each word, we also present the user with a keyword and its visualisation to enhance memory retention. We evaluate our prototype by comparing it to an alternate AR system that does not show an additional visualisation of the keyword, and, also, we compare it to two non-AR systems on a tablet, one with and one without visualising the keyword. Our results indicate that AR outperforms the tablet system regarding immediate recall, mental effort and task completion time. Additionally, the visualisation approach scored significantly higher than showing only the written keyword with respect to immediate and delayed recall and learning efficiency, mental effort and task-completion time.
A critical component of user studies is gaining access to a representative sample of the population researches intend to investigate. Nevertheless, the vast majority of human-computer interaction (HCI)studies, including augmented reality (AR) studies, rely on convenience sampling. The outcomes of these studies are often based on results obtained from university students aged between 19 and 26 years. In order to investigate how the results from one of our studies are affected by convenience sampling, we replicated the AR-supported language learning study called VocabulARy with 24 teenagers, aged between 14 and 19 years. The results verified most of the outcomes from the original study. In addition, it also revealed that teenagers found learning significantly less mentally demanding compared to young adults, and completed the study in a significantly shorter time. All this at no cost to learning outcomes.
Textual sources provide limited information to their readers which could be underwhelming and may reduce engagement. Augmentation approaches have been introduced to present information more engagingly and have shown the potential in supporting information retention. In this research, we inquire further into this opportunity through the use of interactive touchscreen visual elements such as puzzle pieces. We present Retzzles, where users get to solve puzzles applied in a tourist use-case. To evaluate this, we will do a within-subject study with participants n = 30 to determine whether such elements promote engagement which thereby supports information retention. Our preliminary findings shed light on some perspectives on the use of touchscreen displays for engagement but are subject to further investigation. We contribute to more discussions on the use of interactive screens in other similar learning scenarios.
Virtual Reality (VR) provides new possibilities for modern knowledge work. However, the potential advantages of virtual work environments can only be used if it is feasible to work in them for an extended period of time. Until now, there are limited studies of long-term effects when working in VR. This paper addresses the need for understanding such long-term effects. Specifically, we report on a comparative study $i$, in which participants were working in VR for an entire week—for five days, eight hours each day—as well as in a baseline physical desktop environment. This study aims to quantify the effects of exchanging a desktop-based work environment with a VR-based environment. Hence, during this study, we do not present the participants with the best possible VR system but rather a setup delivering a comparable experience to working in the physical desktop environment. The study reveals that, as expected, VR results in significantly worse ratings across most measures. Among other results, we found concerning levels of simulator sickness, below average usability ratings and two participants dropped out on the first day using VR, due to migraine, nausea and anxiety. Nevertheless, there is some indication that participants gradually overcame negative first impressions and initial discomfort. Overall, this study helps lay the groundwork for subsequent research, by clearly highlighting current shortcomings and identifying opportunities for improving the experience of working in VR.
Experiential learning (ExL) is the process of learning through experience or more specifically “learning through reflection on doing”. In this paper, we propose a simulation of these experiences, in Augmented Reality (AR), addressing the problem of language learning. Such systems provide an excellent setting to support “adaptive guidance”, in a digital form, within a real environment. Adaptive guidance allows the instructions and learning content to be customised for the individual learner, thus creating a unique learning experience. We developed an adaptive guidance AR system for language learning, we call Arigatō (Augmented Reality Instructional Guidance & Tailored Omniverse), which offers immediate assistance, resources specific to the learner's needs, manipulation of these resources, and relevant feedback. Considering guidance, we employ this prototype to investigate the effect of the amount of guidance (fixed vs. adaptive-amount) and the type of guidance (fixed vs. adaptive-associations) on the engagement and consequently the learning outcomes of language learning in an AR environment. The results for the amount of guidance show that compared to the adaptive-amount, the fixed-amount of guidance group scored better in the immediate and delayed (after 7 days) recall tests. However, this group also invested a significantly higher mental effort to complete the task. The results for the type of guidance show that the adaptive-associations group outperforms the fixed-associations group in the immediate, delayed (after 7 days) recall tests, and learning efficiency. The adaptive-associations group also showed significantly lower mental effort and spent less time to complete the task.
Experiential learning is the process of learning through experience or more specifically “learning through reflection on doing”. These experiences can be simulated in Extended Reality (XR) systems– interactive environments generated by computer technology combining real and virtual worlds. Such systems have shown to be excellent tools for providing guidance to teach different skills and activities. In this research, we introduce a comprehensive design space grounded in (i) XR technologies, and (ii) theoretical underpinnings of experiential learning and guidance. We use this design space to describe key design decisions in prior work and elucidate existing gaps. Employing the design space, we intend to build an adaptive, personalised instructional guidance system in XR, which offers various levels of guidance that adapt to users’ learning progress. We utilise the Kolb’s experiential learning cycle considering all four stages and transitions within it. With XR technologies, we offer the environment that can provide guidance to support learners to move spontaneously from one stage to another in the experiential learning cycle and further progress on the learning spiral. It is the latter that we explore in this research by adopting potential learning scenarios such as language learning and spatial navigation. We attempt to investigate how different factors of instructional guidance, such as the amount (minimal guidance, full guidance, and adaptive guidance) and the type (from direct instructions to associative models based examples), would affect the experiential learning process and learning outcomes of different learning scenarios in XR environments.
Our experiences of artworks in galleries and museums are generally passive. While this is a widely adopted practice for preservation purposes, it hinders the engagement potential with younger visitors. However, we could adopt novel technologies to change this, and augmented reality (AR) is one of the most promising because of its’ capacity to increase the engagement and add value to the experience. In this paper, we present the development of a mobile augmented reality application designed based on "Google Tango" technology. The playful experience provides by the AR application is divided into three steps, which allow users to colour a contour of a 3D object on a paper, scan it with the AR application and embark on a treasure hunt to find a 3D object (statue) somewhere in the gallery on which the coloured contour would wrap. The paper demonstrates that such a concept is feasible on a modern camera phone with Google Tango technology creating a playful experience in the gallery.