Virtual meetings often exclude individuals who rely on text-based communication, such as autistic adults and those with social anxiety. This paper introduces a prototype that converts typed text into emotive avatars using LLM technology, which convey emotional tone through modulated vocal and facial expressions. We reflect on design choices for using LLMs in accessible meetings and discuss insights from our semi-structured interviews with 18 autistic adults and adults with social anxiety. Our qualitative analysis revealed the following key insights: 1) Participants found the avatars helpful in alleviating challenges like masking and exhaustion, with some noting that the avatars enhanced their communication, increasing participation and confidence; 2) While they valued the avatars' affective capabilities, including both vocal and facial animations, they were sensitive to inaccuracies in vocal expression; and 3) Participants desired more personalized control over the avatars' affect to balance societal expectations with authentic expression.
We introduce the In Real Life (IRL) Ditto, an AI-driven embodied agent designed to represent remote colleagues in shared office spaces, creating opportunities for real-time exchanges even in their absence. IRL Ditto offers a unique hybrid experience by allowing in-person colleagues to encounter a digital version of their remote teammates, initiating greetings, updates, or small talk as they might in person. Our research question examines: How can the IRL Ditto influence interactions and relationships among colleagues in a shared office space? Through a four-day study, we assessed IRL Ditto's ability to strengthen social ties by simulating presence and enabling meaningful interactions across different levels of social familiarity. We find that enhancing social relationships depended deeply on the foundation of the relationship participants had with the source of the IRL Ditto. This study provides insights into the role of embodied agents in enriching workplace dynamics for distributed teams.
Since the Covid-19 pandemic, video calling (VC) has become a staple means of daily communication. Beyond socializing, VC in the United States (U.S.) now supports remote work, healthcare and education. The sudden ubiquity of VC could have presented both advantages and challenges for chronically ill people. However, our understanding of chronically ill people’s experiences with VC remains limited. To address this gap, we conducted the largest online survey study (N=55) on chronically ill people’s VC experiences in the U.S.—investigating their routines, facilitators and barriers. Our quantitative and qualitative findings established that chronically ill people heavily depend on VC to cope with everyday life. At the same time, VC can also detrimentally exacerbate cognitive (e.g., brain fog), emotional (e.g., self-consciousness) and physical challenges (e.g., migraines) for chronically ill people. In response, we offer actionable design opportunities to improve the accessibility and experience of VC for chronically ill people.
Hybridge is an experimental system for exploring the design of remote inclusion for hybrid meetings. In-room users see remote participants on individual displays positioned around a table, and remotes see video feeds from the room integrated into a digital twin of the meeting room. Remotes can choose where to appear in and view the meeting room from. We designed two digital interfaces for remote attendees, one using a 2D canvas, and the other using a 3D digital twin of the room as the medium of interaction. To decide which interface to use for future evaluation, we conducted a within-subjects comparison of 24 groups completing survival tasks. We found that 3D outperformed 2D in the participants’ perceived sense of awareness, sense of agency, and physical presence. The majority of participants also subjectively preferred 3D over 2D. We discuss design recommendations based on usage patterns and participant comments, and plans for further research.
Hybrid meetings limit inclusion for remote participants. The Hybridge experimental system provides different interfaces for remote and room endpoints, focusing on improving inclusion via shared spatiality and remote agency. In-room participants see remotes on displays around a table, and remotes see video integrated into a digital twin. Remotes can choose where to appear and from where they view the room. We tested Hybridge in a within-subjects study of group survival tasks. An in-person condition was followed by a counterbalanced order of hybrid traditional videoconferencing ("Gallery") and Hybridge. We found that co-presence and agency differences between in-room and remotes were alleviated in Hybridge but remained in Gallery. Physical presence for remotes was higher in Hybridge than Gallery. Conversation flow was better in Hybridge than Gallery, but ease of awareness was not different. We argue that asymmetry should be embraced when designing hybrid meeting systems, with inclusivity achieved by tailoring features for the needs of different endpoints.
Virtual environments (VEs) afford similar interactions to those in physical environments: individuals can navigate and manipulate objects. Yet, a prerequisite for these interactions is being able to view the environment. Despite the existence of numerous scene-viewing techniques (i.e., interaction techniques that facilitate the visual perception of virtual scenes), there is no guidance to help designers choose which techniques to implement. We propose a scene taxonomy based on the visual structure and task within a VE by drawing on literature from cognitive psychology and computer vision, as well as virtual reality (VR) applications. We demonstrate how the taxonomy can be used by applying it to an accessibility problem, namely limited head mobility. We used the taxonomy to classify existing scene-viewing techniques and generate three new techniques that do not require head movement. In our evaluation of the techniques with 16 participants, we discovered that participants identified trade-offs in design considerations such as accessibility, realism, and spatial awareness, that would influence whether they would use the new techniques. Our results demonstrate the potential of the scene taxonomy to help designers reason about the relationships between VR interactions, tasks, and environments.
Imagine being able to send a personalized embodied agent to meetings you are unable to attend. This paper explores the idea of a Ditto—an agent that visually resembles a person, sounds like them, possesses knowledge about them, and can represent them in meetings. This paper reports on results from two empirical investigations: 1) focus group sessions with six groups (n=24) and 2) a Wizard of Oz (WOz) study with 10 groups (n=39) recruited from within a large technology company. Results from the focus group sessions provide insights on what contexts are appropriate for Dittos, and issues around social acceptability and representation risk. The focus group results also provide feedback on visual design characteristics for Dittos. In the WOz study, teams participated in meetings with two different embodied agents: a Ditto and a Delegate (an agent which did not resemble the absent person). Insights from this research demonstrate the impact these embodied agents can have in meetings and highlight that Dittos in particular show promise in evoking feelings of presence and trust, as well as informing decision making. These results also highlight issues related to relationship dynamics such as maintaining social etiquette, managing one's professional reputation, and upholding accountability. Overall, our investigation provides early evidence that Dittos could be beneficial to represent users when they are unable to be present but also outlines many factors that need to be carefully considered to successfully realize this vision.
Presenters often screen-share slides over remote meetings as visual aids. Screen reader users however often do not have adequate ways to access slide content during presentations. To inform the design of more accessible interfaces, we ran a formative design workshop to elicit what screen reader users value during live screen-shared presentations, which are Prioritize Key Information, Reduce Cognitive Load, Provide Exploration Independence, Minimize Control Effort, and Encode Spatial Information. Based on these values, we created two prototypes that users experienced in a follow-up design probe. These interactions provide users control over what, when, and how slide elements are read out and introduce spatial audio separation between the presenter's voice and screen reader audio. Participants found that a laptop-based prototype designed for the first four values significantly improved the accessibility of screen-shared content over existing tools, even with a presenter who was perceived to present materials accessibly. The importance of concurrent exploration is highlighted by participants' frequent and diverse use of access to consume information complementary to the presenter, align focus, and understand information using alternative means. A phone-based prototype that encoded spatial information and greater exploration independence at the cost of increased cognitive load received mixed feedback, illustrating the balance between values that interaction design needs to strike. We conclude with insights and considerations for how the identified values and corresponding interactions can be used to improve the accessibility of screen-shared presentations.
With the shift to hybrid meetings in work spaces, there is an increasing need to create a more inclusive hybrid meeting experience where people meeting together in a room interact with those joining remotely. This paper describes a design exploration, implementation, and evaluation of Perspectives, a novel hybrid meeting system that aimed to create an inclusive and equitable space for hybrid meetings. Perspectives digitally composites everyone into a virtual room so that each person has a unique but spatially consistent viewpoint into the meeting. The user study compared Perspectives with three commercially available UX designs for hybrid meetings: Gallery, Together Mode, and Front Row. Results from this study revealed key benefits of Perspectives, including supporting natural interactions, creating a strong sense of co-presence, and reducing cognitive load. Results from the study also helped iterate on the design principles of Perspectives, which offer important insights on supporting hybrid meetings.
Virtual Reality (VR) applications often require users to perform actions with two hands when performing tasks and interacting with objects in virtual environments. Although bimanual interactions in VR can resemble real-world interactions -- thus increasing realism and improving immersion -- they can also pose significant accessibility challenges to people with limited mobility, such as for people who have full use of only one hand. An opportunity exists to create accessible techniques that take advantage of users' abilities, but designers currently lack structured tools to consider alternative approaches. To begin filling this gap, we propose Two-in-One, a design space that facilitates the creation of accessible methods for bimanual interactions in VR from unimanual input. Our design space comprises two dimensions, bimanual interactions and computer assistance, and we provide a detailed examination of issues to consider when creating new unimanual input techniques that map to bimanual interactions in VR. We used our design space to create three interaction techniques that we subsequently implemented for a subset of bimanual interactions and received user feedback through a video elicitation study with 17 people with limited mobility. Our findings explore complex tradeoffs associated with autonomy and agency and highlight the need for additional settings and methods to make VR accessible to people with limited mobility.
People with limited mobility often use multiple devices when interacting with computing systems, but little is known about the impact these multi-modal configurations have on daily computing use. A deeper understanding of the practices, preferences, obstacles, and workarounds associated with accessible multi-modal input can uncover opportunities to create more accessible computer applications and hardware. We explored how people with limited mobility use multi-modality through a three-part investigation grounded in the context of video games. First, we surveyed 43 people to learn about their preferred devices and configurations. Next, we conducted semi-structured interviews with 14 participants to understand their experiences and challenges with using, configuring, and discovering input setups. Lastly, we performed a systematic review of 74 YouTube videos to illustrate and categorize input setups and adaptations in-situ. We conclude with a discussion on how our findings can inform future accessibility research for current and emerging computing technologies.
Interactions in VR for Users with Limited Movement Two-In-One Momona Yamagami* Microsoft Research, Redmond, Washington, USA, my13@uw.edu Sasa Junuzovic Microsoft Research, Redmond, Washington, USA, Sasa.Junuzovic@microsoft.com Mar Gonzalez-Franco Microsoft Research, Redmond, Washington, USA, margon@microsoft.com Eyal Ofek Microsoft Research, Redmond, Washington, USA, eyalofek@microsoft.com Edward Cutrell Microsoft Research, Redmond, Washington, USA, cutrell@microsoft.com John Porter Microsoft Research, Redmond, Washington, USA, John.Porter@microsoft.com Andrew Wilson Microsoft Research, Redmond, Washington, USA, awilson@microsoft.com Martez Mott Microsoft Research, Redmond, Washington, USA, Martez.Mott@microsoft.com
Virtual reality (VR) leverages sight, hearing, and touch senses to convey virtual experiences. For d/Deaf and hard of hearing (DHH) people, however, information conveyed through sound may not be accessible. While prior work has explored making every day sounds accessible to DHH users, the context of VR is, as yet, unexplored. In this paper, we provide a first comprehensive investigation of sound accessibility in VR. Our primary contributions include a design space for developing visual and haptic substitutes of VR sounds to support DHH users and prototypes illustrating several points within the design space. We also characterize sound accessibility in commonly used VR apps and discuss findings from early evaluations of our prototypes with 11 DHH users and 4 VR developers.
Virtual reality (VR) leverages human sight, hearing and touch senses to convey virtual experiences. For d/Deaf and hard of hearing (DHH) people, information conveyed through sound may not be accessible. To help with future design of accessible VR sound representations for DHH users, this paper contributes a consistent language and structure for representing sounds in VR. Using two studies, we report on the design and evaluation of a novel taxonomy for VR sounds. Study 1 included interviews with 10 VR sound designers to develop our taxonomy along two dimensions: sound source and intent. To evaluate this taxonomy, we conducted another study (Study 2) where eight HCI researchers used our taxonomy to document sounds in 33 VR apps. We found that our taxonomy was able to successfully categorize nearly all sounds (265/267) in these apps. We also uncovered additional insights for designing accessible visual and haptic-based sound substitutes for DHH users.
We propose Nearmi, a framework that enables designers to create customizable and accessible point-of-interest (POI) techniques in virtual reality (VR) for people with limited mobility. Designers can use Nearmi by creating and combining instances of its four components—representation, display, selection, and transition. These components enable users to gain awareness of POIs in virtual environments, and automatically re-orient the virtual camera toward a selected POI. We conducted a video elicitation study where 17 participants with limited mobility provided feedback on different Nearmi implementations. Although participants generally weighed the same design considerations when discussing their preferences, their choices reflected tradeoffs in accessibility, realism, spatial awareness, comfort, and familiarity with the interaction. Our findings highlight the need for accessible and customizable VR interaction techniques, as well as design considerations for building and evaluating these techniques.
Through an iterative design process using Wizard of Oz (WOz) prototypes, we designed a video calling application for people with Autism Spectrum Disorder. Our Video Calling for Autism prototype provided an Expressiveness Mirror that gave feedback to autistic people on how their facial expressions might be interpreted by their neurotypical conversation partners. This feedback was in the form of emojis representing six emotions and a bar indicating the amount of overall expressiveness demonstrated by the user. However, when we built a working prototype and conducted a user study with autistic participants, their negative feedback caused us to reconsider how our design process led to a prototype that they did not find useful. We reflect on the design challenges around developing AI technology for an autistic user population, how Wizard of Oz prototypes can be overly optimistic in representing AI-driven prototypes, how autistic research participants can respond differently to user experience prototypes of varying fidelity, and how designing for people with diverse abilities needs to include that population in the development process.
P Dewan合作论文数UNC Department of Computer Sciences7