
Recent advancements in eXtended Reality (XR) technologies have opened new opportunities for integrating virtual and physical environments, enabling natural immersive user experiences. This paper addresses the challenge of developing socially acceptable XR agents that can engage in complex, meaningful interactions across various social and public settings. Specifically, we propose a modular architecture for enhancing user-XR agent interaction, integrated into the framework of the Horizon Europe project “Socially-acceptable Extended Reality Models And Systems (SERMAS)”. The architecture is designed to ensure adaptability, scalability, and the integration of diverse communication modalities, including verbal and non-verbal cues. Its main components work together to enable seamless user detection and communication. In particular, the Detection module plays a key role in managing the agent’s awareness of the surrounding environment. Its functions include identifying user intentions, monitoring group dynamics, and inferring emotional states, thereby increasing responsiveness and enabling more contextually grounded behavior. In contrast, the Communication module employs Large Language Models for personalized verbal responses, and recognizes gestures and body language for enriched non-verbal communication. Together, these approaches guarantee contextually relevant, personalized, and emotionally intelligent interactions. The architecture supports scalability and the flexible integration or replacement of modules, fostering socially acceptable user-XR agent interactions. Its real-world applicability is demonstrated across three deployment scenarios, including a digital receptionist integrated with a physical robot, an XR-based security training system for journalists practicing safety-critical procedures, and a digital assistant supporting customers in a post-office environment. The functionality of the proposed system is presented through the implementation of the XR Agent interacting with users.
Large Language Models (LLMs) are increasingly integrated into research practices due to their ability to simplify knowledge-intensive tasks. Although qualitative research is often seen as difficult to automate, qualitative scholars have begun exploring LLMs for analytical support. Even though existing work highlights concerns about depth, transparency, and responsibility, little is known about how senior qualitative researchers perceive these tools. In this article, we present an exploratory study that investigates the attitudes of 19 senior researchers in psychology, sociology, and anthropology toward the actual or potential use of LLMs in qualitative analysis. Findings show that participants view qualitative analysis as a subjective endeavor and are reluctant to delegate it to LLMs. This resistance stems primarily from anxieties about AI, including fears of losing control and authorship over their work and of being replaced as researchers. Moreover, we show that senior researchers tend to devalue LLM technology when they feel their scientific identity is threatened.
Role-play is an effective method for training conflict resolution skills. However, traditional programs relying on theoretical models, such as the Thomas-Kilmann Conflict Mode Instrument (TKI), often employ questionnaire-based exercises or human-led role-plays, which can limit ecological validity and scalability. To address these limitations, we developed the AI-based Immersive Role-Play Experience Platform (AIRPEP), which integrates large language models (LLMs) and immersive virtual reality (IVR) to simulate realistic training scenarios. AIRPEP features virtual characters, natural interactions, immersive environments, multimodal data collection, and LLM-driven assistants for role-play and personalized feedback. We also proposed a Thematic Content Creation Workflow to create AI-driven scenarios on AIRPEP and applied it to create the Student Conflict Simulator VR, designed for Korean university students. This tool incorporates TKI’s five conflict styles (Collaborating, Competing, Avoiding, Compromising, Accommodating) and the Six Steps of the No-Lose Method. To evaluate the system’s feasibility and user experience, we conducted a pretest-posttest user study with an experimental group (n=32) and a control group (n=23) comprising Korean university students. The experimental group demonstrated a significant improvement in self-expressive communication skills and significantly decreased scores for the Collaborating, Competing, and Avoiding conflict resolution styles. In contrast, the control group’s results indicated a significant decrease in communication skills, with significantly decreased Avoiding conflict resolution style. While the satisfaction and usability ratings of the experimental group were high, the absence of avatars negatively affected their sense of self-presence. Finally, we proposed ten design principles for developing AI-driven role-play training IVR applications. Overall, the initial feasibility and user experience evaluation results indicate that the AIRPEP platform can be used to develop immersive, engaging, and repeatable role-play training interventions, but its long-term effects remain to be evaluated.
In many endoscopic surgical procedures, the surgical team must identify and remove pathological tissue while avoiding critical structures such as arteries and nerves. Augmented reality (AR) offers potential support by overlaying visual information about the location of pathology and critical structures directly onto the operative field, enhancing spatial awareness and surgical navigation. However, limited research has evaluated how best to design and present AR overlays in ways that align with surgical workflow and perception. This study investigates surgeons’ preferences across three key AR overlay dimensions: Design (how anatomy is visualised: outlines, heatmaps, masks, or centroids), Trigger (how and when overlays are activated: always visible, activated by the user, or triggered by instrument position), and Placement (where the overlay appears: above or below the surgical instrument). We take endoscopic pituitary adenoma surgery as a high-risk exemplar. Using a web-based prototype, 38 neurosurgeons ranked options and provided qualitative feedback. Surgeons preferred outline designs for clarity, with a trend towards user-activated triggers for control of information flow and distraction minimisation, and below-instrument placement for better spatial awareness. Preferences were consistent across experience levels and emphasised the importance of balancing visual saliency with cognitive load, to facilitate surgical navigation without distraction or disruption. These findings inform AR interface design, but require evaluation for impact on surgical performance and safety in further physical simulation and clinical studies.
Sustaining children’s engagement and response quality in self-reporting diaries is a longstanding challenge, as repetitive tasks often reduce compliance and motivation. While conversational agents powered by large language models (LLMs) offer a promising solution, most designs for children emphasize semantic adaptation (e.g., phrasing, wording) and leave underexplored the sound-pattern form of language (e.g., rhyme/rhythm) that may shape children’s enjoyment. We conducted a seven-session field study with 32 children (aged 7–13) using an LLM-powered voice-based chatbot that supported two conversational styles: a prose baseline and a phonologically enriched (rhyming) style. Children selected their preferred style before each session, allowing analysis of behavioral preferences, compliance, response quality, and perceived enjoyment. Results show that younger children strongly preferred the rhyming style, whereas older children exhibited more varied preferences, and rhyming preference attenuated over sessions. Session-level rhyming choice was associated with higher completion, fewer skipped days before returning, and higher response quality. Mediation analyses were consistent with enjoyment as a candidate pathway linking style choice to response quality. We contribute a design account of conversational style as semantic scaffolding plus phonological enrichment, with implications for dosage-controlled and adaptive style management in agents for children.
Welcoming behavior is a fundamental and indispensable component of human social interaction, necessary for establishing positive initial encounters and effectively conveying interaction intent. This study aims to gain a deep understanding and quantify such behaviors, thereby laying the foundation for designing more effective, contextually appropriate, and emotionally resonant welcoming behaviors for smart systems. We conducted a content analysis of welcoming behaviors observed in film scenarios featuring human and anthropomorphic characters (e.g., animated figures), identifying 13 terms that summarize various welcoming contexts. Subsequently, utilizing semantic differential scales and exploratory factor analysis (EFA), we identified five underlying dimensions of welcoming behavior: support and acceptance, energetic, ritualistic, resonance, and intimacy. These findings provide a foundation for a framework that enables Human-Computer Interaction (HCI) researchers and practitioners to apply these dimensions of welcome, derived from human and human-like behaviors to the design of smart systems that foster welcoming behaviors.
Web forms are a critical gateway to online services, yet they remain challenging for blind and visually impaired (BVI) users. Despite accessibility standards, screen reader workflows often prove frustrating and error-prone. This paper investigates how conversational agents (CAs) powered by Large Language Models (LLMs) can enable more accessible, voice-based interaction with online forms. Through a human-centered process involving guidelines analysis, stakeholder consultations, prototyping and a study with 18 BVI participants, we explored how LLM-driven CAs can support data entry, navigation, and input verification. Findings show that participants valued the system’s intuitiveness and reliability but highlighted tensions around verbosity, pacing, and control. From these insights, we derive design patterns that translate abstract voice interaction guidelines into concrete strategies for conversational form filling, contributing to more inclusive digital services for BVI users and beyond.
As technologies increasingly converge with users’ psychophysiology and mental states, perception and mind-altering are emerging as plausibly designable components of user experience. This gives rise to emergent phenomena of digitally induced altered states of consciousness (DIAL) that find application not only in therapeutic and mental health contexts but also in communication and leisure. Modern consciousness-altering technologies primarily exist as early prototypes, and their empirical research tends to constrain the context of use to instrumental functionality. Fictional narrative analysis proposes to push these limitations by considering concepts of technology in imaginary contexts, thereby examining the experiences, social rituals, and ethical issues associated with their use far beyond traditional technical scenarios. This article employs a reflexive thematic analysis of popular science fiction narratives to expand our understanding of how ASCs are integrated into HCI, considering broader sociocultural and socio-technical contexts. As a result, we identify key aspects of interaction (purpose and context) and combine their subcategories into a framework that facilitates rethinking traditional HCI, revealing potentials that remain inaccessible within clear consciousness. This framework was used to construct five themes, highlighting that the benefits of mitigating cognitive load and deepening technology-mediated communication with DIAL might entail sacrificing identity and autonomy and increasing the pressures of the modern, overcompetitive, hyper-individualistic lifestyle. Thus, the article contributes to the HCI agenda by expanding existing paradigms of interaction with mind-altering technologies and by demonstrating the use of science fiction as a powerful avenue for exploring the design space and critical inquiries about these types of interactions.
Research on psychoactive substance use during video game play has traditionally adopted a medical lens, focusing on the addictive potential of licit and illicit drugs and their pathological impacts on players. This framing highlights negative outcomes of the co-occurrence between excessive gaming and substance use, leaving the motivations, meanings, and perceptions that players themselves attach to substance use during recreational play under-examined. In this article, we explore a first-person perspective on substance-involved play through a qualitative study with 24 participants (11 Italians and 13 Americans) who reported using psychoactive substances during gameplay. Semi-structured interviews were analyzed using reflexive thematic analysis, which led us to develop three themes: first, adult players describe a “lost immersion”, an effortless absorption they remember from youth and no longer attain spontaneously, which leads them to deploy a repertoire of strategies, including substance use, to re-enter it; second, participants engage in “situated modulation”: they use substances as precise, context-sensitive levers for shaping experience, emotion, and cognition, aligning substance choice with game genre, in-game task, and the stakes of the session; third, participants perform legitimacy work around fairness, treating the same substance as benign in casual play and suspect in competitive play, and practicing situational rather than categorical self-restraint. Substance use also produced unintended effects that sometimes diminished, rather than enhanced, players’ agency over play and over their internal states. By foregrounding first-person accounts, this article contributes to HCI scholarship on digital games, altered states, and player agency in ways that complement the medical lens.
Child helplines require well-trained counsellors to support children in need. Such training typically involves role-playing, which is effective but costly and difficult to organise at scale. A promising addition, therefore, is to offer a simulation-based training, where trainees interact with a conversational agent that mimics a child contacting the helpline. However, interacting with a simulation alone may not be sufficient, as augmenting it with feedback and reflection can provide more guided learning. The learning effects of sequentially adding these elements—simulation, feedback, and reflection—to training systems are not yet fully understood. In this paper, we extend Lilobot, a conversational BDI-based virtual child, to investigate the effects of these elements on learning outcomes. In a randomised controlled online trial (N = 346), participants were randomly assigned to one of four conditions: no intervention, simulation only, simulation with feedback, or simulation with feedback and reflection. Participants interacted with the interventions in five sessions over fourteen days. Compared with the no-intervention condition, simulation training improved participants’ task knowledge and reflective-writing capability. Adding feedback led to further gains in knowledge and conversational outcomes (performance). Also, adding reflection on top of this further increased reflective-writing capability, though it dampened conversational outcomes improvements. These findings clarify how simulation, feedback, and reflection each contribute to learning in early-stage counsellor training.
Large Language Models (LLMs) have become increasingly integrated into critical activities of daily life. This raises concerns about equitable access and utilization across diverse demographics. This study investigates the adoption and usage of LLMs in two waves of a survey with a quota-based U.S. sample (, ). The share of respondents who reported utilizing an LLM rose from 42 % in 2023 to 82 % in 2025. We found a gender gap in early LLM adoption (more male than female users) with complex interaction patterns regarding age and technology education as a mediator. Over time, the gender gap for usage disappeared, but male users were still more frequent users and rated their perceived competence as higher. Common usage scenarios show an increase in information search and editorial use as opposed to a decrease in entertainment purposes and experimental use. Adoption barriers include a lack of knowledge and need, privacy, and ethical concerns. These results underscore the importance of providing education in artificial intelligence in our technology-driven society to promote equitable access to and benefits from LLMs.
Human-automation interaction is vital to ensure safety in partially and conditionally automated vehicles. While previous research mostly focused on the design of takeover requests (TORs) for the automated driving system (ADS), silent failure scenarios without TORs are more common and even more safety-critical. Improving the transparency of the ADS may support drivers in these scenarios. However, little is known about the effectiveness of different types of ADS information on drivers’ performance in handling silent ADS failures. In this study, a driving simulator experiment with 48 participants was conducted to explore the impact of environmental perception (EP, i.e., showing a bounding box around detected items in the environment) and planned maneuver (PM, i.e., illustrating the planned trajectory of the ego vehicle with projection on the road) information on driver takeover performance when facing silent automation failures. Drivers encountered two types of hazards (invisible hazards, which were not visible but could be predicted based on the environment or traffic setup, and visible hazards that were directly visible) in two lighting conditions (i.e., nighttime and daytime). The findings revealed that neither EP nor PM information increased the likelihood of successful takeovers, but both EP and PM, as well as their combination, facilitated earlier takeover actions, which may be attributed to drivers’ earlier detection of hazards with the support of these types of information. The nighttime conditions posed challenges for drivers in responding to silent failures, especially when there were invisible hazards. Finally, EP information could mitigate the negative effects of salient failures on drivers’ trust in ADS. The results from our study can provide insights for human-machine interface designs in vehicles with ADS to facilitate safer driving-task transitions.
As AI systems increasingly support high-stakes decisions, understanding how users evaluate and adopt AI advice is critical. This study theorizes and tests how two types of user knowledge gaps – Prior Unknowns (PUK) and Emergent Unknowns (EUK) – undermine users’ confidence when adopting AI-generated advice. PUK captures self-acknowledged gaps in task-relevant knowledge needed to interpret the decision situation (e.g., not knowing decision-relevant medical or diagnostic concepts). It exists regardless of whether AI is involved. On the other hand, EUK captures interaction-surfaced gaps that arise when users encounter information or prompts they cannot confidently interpret or assimilate during AI use. These are gaps made apparent by the AI system during the interaction. Drawing on Cognitive Fit Theory (CFT), we propose two novel forms of cognitive alignment: prescriptive fit (aligning users’ need for decisiveness with the directive strength of AI advice) and logical fit (aligning users’ need for causal understanding with the depth of explanation). A randomized field experiment in hospitals shows that both PUK and EUK reduce user confidence in AI advice adoption, but require distinct design remedies. Reducing AI advice’s presentation ambiguity (UAMB) increases confidence for both groups, while enhancing the advice’s reasoning transparency (EXP) further increases confidence only under PUK. These findings challenge the assumption that more explanation is always better and provide design guidance for tailoring XAI to users’ cognitive readiness across high-stakes contexts.
The growing integration of Artificial Intelligence (AI), and particularly Machine Learning (ML), into educational technology raises enduring questions about when and why educators trust these systems enough to rely on them. Yet most acceptance models tend to overlook the calibration of trust, often conflating system design quality with actual usage. In this study, we extend current research models by distinguishing perceived trustworthiness (system-attributed properties) from trust (user-anchored reliance), and by integrating human factors including knowledge and skill, cognitive biases, and personality. Drawing on insights from technology acceptance and cognitive psychology, we model how these factors interact with critical design cues (usability, usefulness, social influence, hedonic motivation, and facilitating conditions) to shape adoption behavior. Survey data from 429 educators were analyzed using partial least squares structural equation modelling (PLS-SEM) with confirmatory composite analysis (CCA). The model demonstrated strong psychometric validity and explained substantial variance in trust (R² = .52) and adoption (R² = .25). Knowledge and skills most strongly predicted perceived trustworthiness (β ≈ .67), while cognitive biases significantly weakened the trustworthiness–trust pathway. These findings show that trust is neither automatic nor purely attitudinal, but emerges from the alignment between user competence and transparent system design. We conclude by outlining bias-aware and interpretable design principles to support robust and accountable human–AI collaboration in educational contexts.
This work presents the User Privacy Communication (UPC) Catalogue, a structured collection of research-based guidelines designed to bridge the gap between privacy research and privacy-aware interaction design in personal data-driven systems. The catalogue was derived from a systematic mapping study of 127 user-involving studies, in which common problems, proposed solutions, and underlying rationales were qualitatively analysed and synthesised into guidelines, which were then mapped to a unified set of privacy attributes and classified within design spaces for privacy notices and privacy choices proposed in prior research. In addition to describing the catalogue’s conceptual and structural design, this paper reports an empirical evaluation conducted in two studies involving a total of 92 participants. Participants used the catalogue to analyse diverse digital platforms, identify privacy communication issues, select relevant guidelines, and propose guideline-informed improvements for personal data-driven interfaces. This evaluation enabled examination of guideline effectiveness in terms of alignment among identified problems, selected guidelines, and proposed solutions, as well as participants’ perceptions of guideline relevance, ease of use, and contribution to understanding of data protection principles. The results indicate that the UPC Catalogue supports the formulation of concrete improvement proposals grounded in prior user-centred research, and serves as a practical resource to foster critical reflection and more informed system design decisions regarding personal data-driven interactions.
Dark commercial patterns and deceptive user interface (UI) designs (or DPs) trick consumers into actions that benefit the shareholders. The legal and ethical implications of DPs are shaped by the sociocultural context. Special types of DPs exist in Japan, but the impact of these DPs on consumer attitudes and behaviour remains underexplored. We report on the first comparative mixed methods user study with Japanese consumers (N=84), half of whom (n=40) experienced a range of DPs—including the Japanese varieties—in a simulated e-commerce website. We discovered that the Japanese DPs were among the least noticeable and caused the highest simulated financial harm, with Untranslation perceived as highly disruptive and Alphabet Soup highly deceptive. Comparative analyses with a group that experienced the DP-free version of the online store (N=44) revealed a sharp negative difference in positive emotions and acceptance. Qualitative analyses surfaced cultural norms in consumer–business relationships, notably unease, ambivalence, and endured disloyalty (不誠実 or fuseijitsu). No evidence of sampling biases was found for the participants involved in a publicly broadcast programme on the study (n=10), indicating true deception and unacceptability. Our findings suggest that while reactions toward and ability to recognize a given DP may vary across individuals, the mere presence of DPs tends to have negative effects on most Japanese consumers.
Emerging generative AI (GenAI) tools, which are known to encompass cultural biases, are anticipated to significantly impact creative work, especially as they become incorporated into professional software suites such as Adobe Creative Cloud, Canva, Figma, and Miro. Professionals and pundits predict positive and negative outcomes, ranging from substantial productivity gains to the risk of reduced employment opportunities for creative practitioners, concerns around intellectual property, and fears that a lack of diversity in training data will result in diminished variety of creative outputs. This uncertainty makes it vital to examine how generative AI is used and perceived at the cutting edge of adoption across professional contexts in both the Global North and South. We conducted semi-structured interviews with 20 professional User Experience (UX), User Interface (UI) and graphic designers from across cultural contexts, with most from Global South backgrounds. Participants were at varying stages of GenAI adoption; some were exploring and others were integrating GenAI tools into their professional workflows. Using reflexive thematic analysis, we analysed participants’ experiences to understand how socio-economic and cultural situatedness shaped their engagement with GenAI tools. Our findings revealed a complex, asymmetric landscape of AI adoption within professional design workflows. We contribute insights into asymmetries within extended creative cognition, uncovering a paradox of access wherein AI provides essential scaffolding for participants’ work while simultaneously imposing a standardisation tax that constrains situated cultural expression. We propose design directions for GenAI tools that may promote cultural inclusivity and ethically responsible Human-AI collaboration.
Directing a person’s attention by guiding their gaze to an object of interest is a promising application of extended reality (XR) to support people’s everyday lives. XR gaze guidance is not only useful for assisting visual search, but also for influencing a person’s viewing order of objects, which affects memory and comprehension. For example, following an expert’s eye movements can enhance one’s proficiency in identifying relevant information. Existing gaze guidance techniques typically guide users to one target at a time. However, many real-world tasks require guidance to multiple targets, where the viewing sequence is crucial for comprehension or problem-solving. This paper addresses this gap by exploring how to guide users to multiple targets in a sequential order by applying guidance cues several times consecutively (i.e., chaining them). This paper presents a user study (N = 35) investigating a continuous gaze guidance technique that chains individual cues based on subtle flickering (i.e., subtle gaze direction). Participants perform visual foraging–a controlled search task with multiple target and distractor objects. The results suggest that subtle chaining is effective when it supports users’ search strategies, but less effective in guiding users to targets that require a higher degree of cognitive processing. We discuss how the results can improve the design of gaze guidance techniques for multiple targets in XR.
Remote supervisory-based operations, like satellite operations, are becoming increasingly complex, creating a need for improved training. Challenges exist in training operators to have the necessary skills and understanding to perform complex satellite operations, such as on-orbit servicing. Virtual reality (VR), which creates a sense of immersion, has been demonstrated as a learning tool across many fields, however, it has not been investigated for remote supervision environments. This paper investigates the effects of immersion and 3D visualizations as a way to improve training for a satellite rendezvous mission. The transfer of skills after training in three different displays, but then performing the task in a traditional display is compared through human subject testing (n = 45) on measures of situation awareness (SA), workload, usability, performance, and subjective utility. The three displays include an immersive VR display, a computer screen with 3D visualizations, and a representative traditional display with only 2D elements. It was found that training in VR improved level 2 SA (p = 0.006) and increased usability (p = 0.005) during operational tasks. VR training also had higher perceived subjective utility, particularly in aspects such as understanding collision likelihood (p = 0.022) and orbital motion (p = 0.041). Training in 3D screen visualizations led to improved performance in more difficult scenarios (p = 0.034) and subjective understanding of event awareness (p = 0.024). No differences were found between workload and levels 1 and 3 SA. This research demonstrates the benefits of using VR as a training tool for complex supervisory operations and may lead to improved operational outcomes.
Crowdsourcing depends on engaging qualified participants, referred to as solvers, and gamification can facilitate their engagement. As task complexity increases, solver motivation decreases, making it critical to design gamification systems that sustain motivation in complex tasks requiring skilled solvers. Existing gamification designs typically select game elements based on popularity rather than motivational theory, limiting their effectiveness in complex tasks where solvers face cognitive demands and evolving goals. Motivational affordance theory (MAT) provides a systematic framework mapping system features to user motivations through affordances, yet MAT remains underexplored in complex crowdsourcing gamification. To address this gap, this study integrates MAT principles into gamification design to systematically support solver motivation in complex crowdsourcing tasks. We employ a two-stage research design. In Study 1, we compared four crowdsourcing task types (processing, rating, solving, creating) to identify which tasks solvers perceive as complex and how they expect to benefit from gamification. Based on these insights, we developed a Motivational Affordance Perspective (MAP) design incorporating motivational affordances to guide gamification system design. In Study 2, we implemented a mobile application featuring three gamification designs: baseline, popular, and MAP, and evaluated their effects over a three-week field deployment with 68 registered participants. The analysis sample comprised 44 participants who met the predefined minimum participation criterion across all three systems. The study revealed the influence of task familiarity and topical interest on psychological and behavioral outcomes, indicating that gamification systems should consider both task complexity and individual solver factors. Based on our findings, we provide implications for designing gamification systems for complex crowdsourcing tasks.