We make sense of the world around us through our bodies; however, this somatic dimension of meaning-making is often overlooked in the development of AI systems. This workshop (re-)positions the body as central to the design of human-AI interactions by critically exploring the frictions and possibilities that emerge when attempting to incorporate our somatic dimension into the design of predominantly disembodied AI systems. By using self-knowledge as a point of departure to explore the potential of soma-aligned AI as a research territory, our workshop hosts (1) participant-driven discussion on tensions and opportunities between AI design and soma-centric approaches and (2) practical exercises where we together experiment with designing forms of such interactions that are more embodied, sensuous and poetic. We aim to extend these activities beyond the workshop, both by establishing a long-term community of design researchers and through a public Poetry Jam event.
Foreign language speaking anxiety (FLSA) poses a major challenge for English-language learners, suppressing confidence and triggering a cycle of avoidance that hinders language acquisition. To address this, we explored the use of LLM-based embodied conversational agents (ECA) in social virtual reality (VR), which provide personalized support and multimodal interaction in a contextualized environment. We developed three English-language learning scenarios in social VR and conducted a five-day mixed-methods study where participants (N=20) engaged in daily 30-minute role-play practice with an LLM-based ECA to evaluate the efficacy of the system. Quantitative results showed a significant reduction in self-reported FLAS after 3 days, along with subtle gains in speaking proficiency measures. Qualitatively, learners perceived increased confidence, attributing it to the LLM-based ECA’s non-judgmental stance, linguistic scaffolding, affective encouragement, and adaptive feedback. Our findings suggest the potential of LLM-based ECAs in social VR for language learning and offer considerations for future agent design.
Generative AI (GenAI) is increasingly transforming human-centered design (HCD) by enabling the generation and simulation of AI-driven personas that can role‑play different stakeholders, ideate solutions, and perform usability tests. With their increased efficiency and accessibility, these AI personas offer advantages such as scalability, rapid ideation, coverage of diverse and edge cases, and reduced reliance on costly human subject studies at early stages of design. However, they also pose ethical, representational, and methodological challenges. Although evidence on the use, benefits, and limitations of AI personas is growing, the field still lacks shared validity standards and reporting norms to guide the use of AI personas and their synthetic feedback, and decisions about when to involve actual human participants. This half-day workshop will examine the roles, opportunities, and responsibilities associated with integrating GenAI simulated personas into HCI design practice. Through the presentation of emerging evidence, collaborative discussions, and hands-on activities, we aim to explore responsible methods and develop practical guidelines for the responsible use of AI personas in human-centered design and research.
Users feel frustrated when they do not know when to speak with LLM-based agents. Technical delays disrupt the natural rhythm of conversation (turn-taking), yet there is little understanding of how these specific delays impact the back-and-forth flow of interaction. To address this, we analyzed human-agent conversations in social VR to measure timing differences. We used conversation analysis techniques to track specific timing metrics, such as how long it takes to respond (response latencies) and how agents handle interruptions (repair attempts). We found that agents are significantly slower to respond with a median of 4.1 seconds compared to a human’s 1.2 seconds. We identified a “conversational timing drift”, noting that agents struggle with start-up latency, i.e., taking too long to start speaking, and wind-down latency, i.e., failing to stop speaking quickly when a user interrupts them. This is the first study to empirically quantify human-agent conversational latencies within VR. We offer design suggestions to help future agents manage conversational timing better, ultimately improving natural conversation and user experience.
As Large Language Model (LLM)-based AI agents emerge in social Virtual Reality (VR), understanding the roles they might play is critical. We conducted interviews in VRChat with experienced users and developers $(\mathrm{N}=6)$ and, using thematic analysis, identified five primary roles: information sharing, personal companionship, moderation, content generation, and self-development and learning. While participants recognized the potential for increased accessibility and support, they expressed concerns regarding overmoderation, the replacement of human creativity, and the psychological impact of AI co-dependency. These preliminary findings highlight the need to balance AI utility with the preservation of authentic human connection in social VR.
Speech-language pathology (SLP) training programs face persistent challenges in providing sufficient opportunities for clinical placements or simulations that allow students to practice clinical communication with patients with communication disorders. We designed and implemented V.O.I.C.E., an AI-powered virtual patient system that enables realistic, open-ended clinical simulation for SLP students. V.O.I.C.E. integrates (1) a large language model (LLM) role-play agent designed to produce responses consistent with symptoms of post-stroke expressive aphasia (Broca’s aphasia), (2) a multimodal emotional expression pipeline that renders context-appropriate emotions of the virtual patient through coordinated speech prosody, facial expressions, and gestures, and (3) an LLM-based, rubric-guided debriefing module that generates structured formative feedback across core clinical communication competencies. We conducted a pilot mixed-methods, quasi-experimental evaluation with 11 graduate-level SLP students, where they completed three simulation sessions using the V.O.I.C.E system with progressive challenges. Results show significant gains in students' overall self-efficacy for interacting with Broca's aphasia patients in clinical settings. The self-efficacy gains were moderated by students’ perceptions of multimodal realism of the simulations. Qualitative findings highlight the system’s value as a low-stakes “bridge” for transition to clinical practice, which supports learning through iterative practice, reflection, and strategy refinement. This work provides a framework for AI-based simulation training applications that emphasizes interpersonal communication.
We present a case study of Persona-L, a system that leverages large language models (LLMs) and retrieval-augmented generation (RAG) to model personas of people with Down syndrome. Existing approaches to persona creation can often lead to oversimplified or stereotypical profiles of people with Down Syndrome. To that end, we built stereotype detection capabilities into Persona-L. Through interviews with caregivers and healthcare professionals (N=10), we examine how Down Syndrome stereotypes could manifest in both, content and delivery of LLMs, and interface design. Our findings show the challenges in stereotypes definition, and reveal the potential stereotype emergence from the training data, interface design, and the tone of LLM output. This highlights the need for participatory methods that capture the heterogeneity of lived experiences of people with Down Syndrome.
In this paper, we present Speak Ease: an augmentative and alternative communication (AAC) system to support users' expressivity by integrating multimodal input, including text, voice, and contextual cues (conversational partner and emotional tone), with large language models (LLMs). Speak Ease combines automatic speech recognition (ASR), context-aware LLM-based outputs, and personalized text-to-speech technologies to enable more personalized, natural-sounding, and expressive communication. Through an exploratory feasibility study and focus group evaluation with speech and language pathologists (SLPs), we assessed Speak Ease's potential to enable expressivity in AAC. The findings highlight the priorities and needs of AAC users and the system's ability to enhance user expressivity by supporting more personalized and contextually relevant communication. This work provides insights into the use of multimodal inputs and LLM-driven features to improve AAC systems and support expressivity.
We present Persona-L, a novel approach for creating personas using Large Language Models (LLMs) and an ability-based framework, specifically designed to improve the representation of users with complex needs. Traditional methods of persona creation often fall short of accurately depicting the dynamic and diverse nature of complex needs, resulting in oversimplified or stereotypical profiles. Persona-L enables users to create and interact with personas through a chat interface. Persona-L was evaluated through interviews with UX designers (N=6), where we examined its effectiveness in reflecting the complexities of lived experiences of people with complex needs. We report our findings that indicate the potential of Persona-L to increase empathy and understanding of complex needs while also revealing the need for transparency of data used in persona creation, the role of the language and tone, and the need to provide a more balanced presentation of abilities with constraints.
Many people struggle with learning a new language, with traditional tools falling short in providing contextualized learning tailored to each learner's needs. The recent development of large language models (LLMs) and embodied conversational agents (ECAs) in social virtual reality (VR) provide new opportunities to practice language learning in a contextualized and naturalistic way that takes into account the learner's language level and needs. To explore this opportunity, we developed ELLMA-T, an ECA that leverages an LLM (GPT-4) and situated learning framework for supporting learning English language in social VR (VRChat). Drawing on qualitative interviews (N=12), we reveal the potential of ELLMA-T to generate realistic, believable and context-specific role plays for agent-learner interaction in VR, and LLM's capability to provide initial language assessment and continuous feedback to learners. We provide five design implications for the future development of LLM-based language agents in social VR.
We present the Requirement Elicitation Tool that leverages Large Language Model (LLM) (gpt-4o-mini) to enable simulated real-world interactions of requirements gathering from three synthetic personas. We demonstrate the use case of Computer Science (CS) students in Database Management Systems leveraging the tool to build a conceptual model and Entity-Relationship (ER) diagrams. Our preliminary findings show the potential of this tool to engage students in discovery process without providing predefined solutions and set the directions for future work.
Recent advancements in augmented reality and virtual reality have significantly enhanced workflows for drawing 3D objects. Despite these technological strides, existing AR tools often lack the necessary precision and struggle to maintain quality when scaled, posing challenges for larger-scale drawing tasks. This paper introduces a novel AR tool that uniquely integrates bitmap drawing and vectorization techniques. This integration allows engineers to perform rapid, real-time drawings directly on 3D models, with the capability to vectorize the data for scalable accuracy and editable points, ensuring no loss in fidelity when modifying or resizing the drawings. We conducted user studies involving professional engineers, designers, and contractors to evaluate the tool's integration into existing workflows, its usability, and its impact on project outcomes. The results demonstrate that our enhancements significantly improve the efficiency of drawing processes. Specifically, the ability to perform quick, editable, and scalable drawings directly on 3D models not only enhances productivity but also ensures adaptability across various project sizes and complexities.
In this paper, we present Safe Guard, an LLM-agent for the detection of hate speech in voice-based interactions in social VR (VRChat). Our system leverages Open AI GPT and audio feature extraction for real-time voice interactions. We contribute a system design and evaluation of the system that demonstrates the capability of our approach in detecting hate speech, and reducing false positives compared to currently available approaches. Our results indicate the potential of LLM-based agents in creating safer virtual environments and set the groundwork for further advancements in LLM-driven moderation approaches.
Recent progress in large language model (LLM) technology has significantly enhanced the interaction experience between humans and voice assistants (VAs). This project aims to explore a user's continuous interaction with LLM-based VA (LLM-VA) during a complex task. We recruited 12 participants to interact with an LLM-VA during a cooking task, selected for its complexity and the requirement for continuous interaction. We observed that users show both verbal and nonverbal behaviors, though they know that the LLM-VA can not capture those nonverbal signals. Despite the prevalence of nonverbal behavior in human-human communication, there is no established analytical methodology or framework for exploring it in human-VA interactions. After analyzing 3 hours and 39 minutes of video recordings, we developed an analytical framework with three dimensions: 1) behavior characteristics, including both verbal and nonverbal behaviors, 2) interaction stages–exploration, conflict, and integration–that illustrate the progression of user interactions, and 3) stage transition throughout the task. This analytical framework identifies key verbal and nonverbal behaviors that provide a foundation for future research and practical applications in optimizing human and LLM-VA interactions.
In this paper, we introduce the design and evaluation of an LLM-based AI agent for human-agent interaction in Virtual Reality (VR). Our AI agent system leverages GPT-4, a Large Language Model (LLM) to simulate human behavior. Our LLM-based agent, deployed in VRChat as a Non-playable Character (NPC), exhibits the ability to respond to a player by providing context-relevant responses followed by appropriate facial expressions and body gestures. Our preliminary evaluation yielded the most optimal parameters for generating the most plausible responses. With our system, we lay the groundwork for future development of LLM-based NPCs in VR.
Synthetic personae and data powered by artificial intelligence (AI) are emerging in many HCI areas, including education and training, gaming, and piloting research studies. Recently, Large Language Models (LLMs) have shown promise for synthetic AI personae, experimenting with human and social simulacra and producing synthetic data. This presents challenges and opportunities for extending HCI research via LLMs and AI. In this proposed workshop, we engage HCI researchers interested in working with LLMs, synthetic personae, and synthetic data through speculative design and producing visions, desiderata, and requirements for future HCI research engaging with synthetic personae/data. The outcomes of this workshop may be disseminated to the HCI community through scientific publications or special issues to facilitate continued discussion and advance knowledge on a timely HCI topic.
An emerging class of MobileHCI applications demand high-fidelity 3D rendering of everything from accurate geo-spatial information to real-time dynamic data sources. This requires a rethinking of interfaces for everything from visualizing crowd-sourced mapping applications to highly realistic, real-time simulation environments. Whether these applications are running on sophisticated head-mounted mobile devices or commodity phones, these generated environments need to take advantage of spatial cues, and need to deliver uninterrupted immersive user experiences in extended reality (XR). This workshop will focus on this emerging class of applications and HCI challenges in mobile applications involving rich interactive XR assets. Topics include but are not limited to:
Third wave HCI initiated a slow transformation in the methods of UX research: from widely used quantitative approaches to more recently employed qualitative techniques. Articulating the nuances, complexity, and diversity of a user's experience beyond surface descriptions remains a challenge within design. One qualitative method — micro-phenomenology — has been used in HCI/Design research since 2001. Yet, no systematic understanding of micro-phenomenology has been presented, particularly from the perspective of HCI/Design researchers who actively use it in design contexts. We interviewed 5 HCI/Design experts who utilize micro-phenomenology and present their experiences with the method. We illustrate how this method has been applied by the selected experts through developing a practice, and present conditions under which the descriptions of the experience unfold, and the values that this method can provide to HCI/Design field. Our contribution highlights the value of micro-phenomenology in articulating the experience of designers and participants, developing vocabulary for multi-sensory experiences, and unfolding embodied tacit knowledge.