Virtual meetings often exclude individuals who rely on text-based communication, such as autistic adults and those with social anxiety. This paper introduces a prototype that converts typed text into emotive avatars using LLM technology, which convey emotional tone through modulated vocal and facial expressions. We reflect on design choices for using LLMs in accessible meetings and discuss insights from our semi-structured interviews with 18 autistic adults and adults with social anxiety. Our qualitative analysis revealed the following key insights: 1) Participants found the avatars helpful in alleviating challenges like masking and exhaustion, with some noting that the avatars enhanced their communication, increasing participation and confidence; 2) While they valued the avatars' affective capabilities, including both vocal and facial animations, they were sensitive to inaccuracies in vocal expression; and 3) Participants desired more personalized control over the avatars' affect to balance societal expectations with authentic expression.
We introduce ChartA11y, an app developed to enable accessible 2-D visualizations on smartphones for blind users through a participatory and iterative design process involving 13 sessions with two blind partners. We also present a design journey for making accessible touch experiences that go beyond simple auditory feedback, incorporating multimodal interactions and multisensory data representations. Together, ChartA11y aimed at providing direct chart accessing and comprehensive chart understanding by applying a two-mode setting: a semantic navigation framework mode and a direct touch mapping mode. By re-designing traditional touch-to-audio interactions, ChartA11y also extends to accessible scatter plots, addressing the under-explored challenges posed by their non-linear data distribution. Our main contributions encompass the detailed participatory design process and the resulting system, ChartA11y, offering a novel approach for blind users to access visualizations on their smartphones.
The opportunity for artificial intelligence, or AI, to enable accessibility is rapidly growing, but widely impactful applications can be challenging to build given the diversity of user need within and across disability communities. Teachable AI systems give users with disabilities a way to leverage the power of AI to personalize applications for their own specific needs, as long as the effort of providing examples is balanced with the benefit of the personalization received. As an example, this paper presents the design and evaluation of Find My Things, an end-to-end application that can be taught by people who are blind or low vision to find their personal things. Through synthesis of the design process, this paper offers design considerations for the teaching loop that is so critical to realizing the power of teachable AI for accessibility.
Imagine being able to send a personalized embodied agent to meetings you are unable to attend. This paper explores the idea of a Ditto—an agent that visually resembles a person, sounds like them, possesses knowledge about them, and can represent them in meetings. This paper reports on results from two empirical investigations: 1) focus group sessions with six groups (n=24) and 2) a Wizard of Oz (WOz) study with 10 groups (n=39) recruited from within a large technology company. Results from the focus group sessions provide insights on what contexts are appropriate for Dittos, and issues around social acceptability and representation risk. The focus group results also provide feedback on visual design characteristics for Dittos. In the WOz study, teams participated in meetings with two different embodied agents: a Ditto and a Delegate (an agent which did not resemble the absent person). Insights from this research demonstrate the impact these embodied agents can have in meetings and highlight that Dittos in particular show promise in evoking feelings of presence and trust, as well as informing decision making. These results also highlight issues related to relationship dynamics such as maintaining social etiquette, managing one's professional reputation, and upholding accountability. Overall, our investigation provides early evidence that Dittos could be beneficial to represent users when they are unable to be present but also outlines many factors that need to be carefully considered to successfully realize this vision.
Three-dimensional virtual environments are currently inaccessible to people who are blind, as current screen-reading solutions for 2D content are not fully extensible to achieve the needed embodied spatial presence. Forefronting perceptual agency as key to any access approach for users who are blind, we offer Scene Weaving as an interactional metaphor that allows users to choose how and when they perceive the environment and the people in it. We illustrate how this metaphor can be implemented in an example prototype system. In this interactivity, users can control how and when they perceive a virtual museum environment and people within it through a range of interaction mechanisms.
Only a small percentage of blind and low-vision people use traditional mobility aids such as a cane or a guide dog. Various assistive technologies have been proposed to address the limitations of traditional mobility aids. These devices often give either the user or the device majority of the control. In this work, we explore how varying levels of control affect the users' sense of agency, trust in the device, confidence, and successful navigation. We present Glide, a novel mobility aid with two modes for control: Glide-directed and User-directed. We employ Glide in a study (N=9) in which blind or low-vision participants used both modes to navigate through an indoor environment. Overall, participants found that Glide was easy to use and learn. Most participants trusted Glide despite its current limitations, and their confidence and performance increased as they continued to use Glide. Users' control mode preferences varied in different situations; no single mode "won" in all situations.
SSVEP-based BCIs are amongst the most promising BCIs in terms of speed and accuracy. However, despite significant effort from the community in order to make them more practical and user friendly, they remain particularly annoying to use. In this paper, we investigate the effect of the size and contrast of the SSVEP visual stimulations on both of the classification accuracy and the annoyance of the interface, with the global aim to find a trade-off between performance and user-friendliness. We conducted a user study on twelve (12) participants in order to evaluate the joint effect of different stimulation sizes and contrasts on the SSVEP classification accuracy in a Virtual Reality context. The results of this experiment suggest that the size of the stimulation has a significant impact on both of the classification accuracy, below a certain threshold, and on the perceived annoyance. No effect of the contrast was however found neither on the classification accuracy nor on the perceived annoyance, suggesting that it is still possible to accurately operate SSVEP-based BCIs using lower contrast stimulation.
Even though screen readers are a core accessibility tool for blind and low vision individuals (BLVIs), most visualizations are incompatible with screen readers. To improve accessible visualization experiences, we partnered with 10 BLV screen reader users (SRUs) in an iterative co-design study to design and develop accessible visualization experiences that afford SRUs the autonomy to interactively read and understand visualizations and their underlying data. During the five-month study, we explored accessible visualization prototypes with our design partners for three one-hour sessions. Our results provide feedback on the synthesized design concepts we explored, why (or why not) they aid comprehension and exploration for SRUs, and how differing design concepts can fit into cohesive accessible visualization experiences. We contribute both Chart Reader, a web-based accessibility engine resulting from our design iterations, and our distilled study findings—organized by design dimensions—in the creation of comprehensive accessible visualization experiences.
Profile pictures can convey rich social signals that are often inaccessible to blind and low vision screen reader users. Although there have been efforts to understand screen reader users’ preferences for alternative (alt) text descriptions when encountering images online, profile pictures evoke distinct information needs. We conducted semi-structured interviews with 16 screen reader users to understand their preferences for various styles of profile picture image descriptions in different social contexts. We also interviewed seven sighted individuals to explore their thoughts on authoring alt text for profile pictures. Our findings suggest that detailed image descriptions and user narrated alt text can provide screen reader users enjoyable and informative experiences when exploring profile pictures. We also identified mismatches between how sighted individuals would author alt text with what screen reader users prefer to know about profile pictures. We discuss the implications of our findings for social applications that support profile pictures.
Brain-computer interfaces (BCIs) employ various paradigms which afford intuitive, augmented control for users to navigate digital technologies. In this study we explore the application of these BCI concepts to predictive text systems: commonplace interactive and assistive tools with variable usage contexts and user behaviors. We conducted an experiment to analyze user neurophysiological responses under these different usage scenarios and evaluate the feasibility of a closed-loop, adaptive BCI for use with such technologies. We recorded electroencephalogram (EEG) and eye tracking (ET) data from participants while they completed a self-paced typing task in a simulated predictive text environment. Participants completed the task with different degrees of reliance on the predictive text system (completely dependent, completely independent, or their choice) and encountered both correct and incorrect text generations. Data suggest that erroneous text generations may evoke neurophysiological responses that can be measured with both EEG and pupillometry. Moreover, these responses appear to change according to users’ reliance on the predictive text system. Results show promise for use in a passive, hybrid, BCI with a closed-loop, adaptive framework, and support a neurophysiological approach to the challenge of real-time human feedback on system performance.
Notifications and alerts can both convey critical data needed to make decisions as we go about our day, and at the same time be completely ignored as a result of the overuse of these channels. Auditory notifications can be especially intrusive given their demand for synchronous attention. However, during some activities, such as navigation guidance via GPS, immediate information cues are necessary. In this case audio alerts and instructions are typically implemented as voice commands, but processing language requires significant cognitive resources and attention. In this study we report findings about the feasibility of a novel ambient pedestrian guidance system inspired by neuroscience. Results suggest that it is possible to communicate precise spatial actions without the intrusion of verbal turn-by-turn cues. We also share learnings from our use of low-fidelity prototyping to evaluate the novel audio-AR interaction.
Brain-Computer Interface (BCI) technology may provide individuals with motor impairments or even the general population a new way to interact with the world around them. However, current BCI systems using electroencephalography (EEG) can be unreliable and produce large variations in performance. Most studies seek to improve performance by focusing on signal processing and classification techniques. However, it may also be beneficial to investigate different control strategies. For this reason, the main objective of this pilot study was to investigate the use of visual imagery, a control paradigm that has not been much tested for EEG BCI applications. Visual imagery may provide a more intuitive control strategy with a greater number of available classes than other popular imagery-based methods such as motor imagery. Using this paradigm, we have demonstrated above chance binary classification accuracy (59.9%, p < 0.05) during offline decoding of face and scene visual imagery. Furthermore, the participant in this study achieved significantly above chance performance during a three-class, closed-loop BCI interaction (47.2%, p = 0.05). The initial results of this pilot study demonstrate the feasibility of using visual imagery as an alternative EEG BCI control paradigm.
Virtual Reality (VR) applications often require users to perform actions with two hands when performing tasks and interacting with objects in virtual environments. Although bimanual interactions in VR can resemble real-world interactions -- thus increasing realism and improving immersion -- they can also pose significant accessibility challenges to people with limited mobility, such as for people who have full use of only one hand. An opportunity exists to create accessible techniques that take advantage of users' abilities, but designers currently lack structured tools to consider alternative approaches. To begin filling this gap, we propose Two-in-One, a design space that facilitates the creation of accessible methods for bimanual interactions in VR from unimanual input. Our design space comprises two dimensions, bimanual interactions and computer assistance, and we provide a detailed examination of issues to consider when creating new unimanual input techniques that map to bimanual interactions in VR. We used our design space to create three interaction techniques that we subsequently implemented for a subset of bimanual interactions and received user feedback through a video elicitation study with 17 people with limited mobility. Our findings explore complex tradeoffs associated with autonomy and agency and highlight the need for additional settings and methods to make VR accessible to people with limited mobility.
Brain-computer interfaces (BCIs) using Electroencephalography (EEG) have drawn attention to providing alternative control pathways for users with motor disabilities or even the general public in real-world environments due to their robustness, relatively low cost, and high portability. However, EEG still suffers from large variability between subjects or between sessions of an individual subject. To obtain optimal performance, a BCI usually requires a user to go through a calibration process to fine-tune the model. This calibration process is usually long and could hinder the practicality of a BCI. In this study, we propose a closed-loop framework that monitors the user EEG responses to the action of a BCI. If an Error-related Potential (ErrP) is detected in the response, it is indicated that the BCI is making a wrong prediction. By using the information from this ErrP detector, we can include online testing trials into the training pool and further fine-tune the model over the time the BCI is used. Results suggest that the proposed framework can reach better results with a few additional trials when compared to the model pre-trained from some existing data. Also, the performance of the proposed model can gradually converge to a fully calibrated model, which suggests that the conventional calibration process could be replaced by online training.
Approximately 15% of the world's population has a disability and 80% live in low resource-settings, often in situations of severe social isolation. Technology is often inaccessible or inappropriately designed, hence unable to fully respond to the needs of people with disabilities living in low resource settings. Also lack of awareness of technology contributes to limited access. This workshop will be a call to arms for researchers in HCI to engage with people with disabilities in low resourced settings to understand their needs and design technology that is both accessible and culturally appropriate. We will achieve this through sharing of research experiences, and exploration of challenges encountered when planning HCI4D studies featuring participants with disabilities. Thanks to the contributions of all attendees, we will build a roadmap to support researchers aiming to leverage post-colonial and participatory approaches for the development of accessible and empowering technology with truly global ambitions.
Object recognition has made great advances in the last decade, but predominately still relies on many high-quality training examples per object category. In contrast, learning new objects from only a few examples could enable many impactful applications from robotics to user personalization. Most few-shot learning research, however, has been driven by benchmark datasets that lack the high variation that these applications will face when deployed in the real-world. To close this gap, we present the ORBIT dataset and benchmark, grounded in the real-world application of teachable object recognizers for people who are blind/low-vision. The dataset contains 3,822 videos of 486 objects recorded by people who are blind/low-vision on their mobile phones. The benchmark reflects a realistic, highly challenging recognition problem, providing a rich playground to drive research in robustness to few-shot, high-variation conditions. We set the benchmark's first state-of-the-art and show there is massive scope for further innovation, holding the potential to impact a broad range of real-world vision applications including tools for the blind/low-vision community. We release the dataset at https://doi.org/10.25383/city.14294597 and benchmark code at https://github.com/microsoft/ORBIT-Dataset.
Artificial Intelligence (AI) for accessibility is a rapidly growing area, requiring datasets that are inclusive of the disabled users that assistive technology aims to serve. We offer insights from a multi-disciplinary project that constructed a dataset for teachable object recognition with people who are blind or low vision. Teachable object recognition enables users to teach a model objects that are of interest to them, e.g., their white cane or own sunglasses, by providing example images or videos of objects. In this paper, we make the following contributions: 1) a disability-first procedure to support blind and low vision data collectors to produce good quality data, using video rather than images; 2) a validation and evolution of this procedure through a series of data collection phases and 3) a set of questions to orient researchers involved in creating datasets toward reflecting on the needs of their participant community.
Alternative (alt) text provides access to descriptions of digital images for people who use screen readers. While prior work studied screen reader users’ (SRUs’) preferences about alt text and automatic alt text (i.e., alt text generated by artificial intelligence), little work examined the alt text author’s experience composing or editing these descriptions. We built two types of prototype interfaces for two tasks: authoring alt text and providing feedback on automatic alt text. Through combined interview-usability testing sessions with alt text authors and interviews with SRUs, we tested the effectiveness of our prototypes in the context of Microsoft PowerPoint. Our results suggest that authoring interfaces that support authors in choosing what to include in their descriptions result in higher quality alt text. The feedback interfaces highlighted considerable differences in the perceptions of authors and SRUs regarding “high-quality” alt text. Finally, authors crafted significantly lower quality alt text when starting from the automatic alt text compared to starting from a blank box. We discuss the implications of these results on applications that support alt text.
driven communication by business, government, and science. Furthermore, the use and need for visualizations is not just confined to data experts: Data visualizations are becoming ubiquitous in textbooks, presentations, and reports, as well as in popular media, both online and in print. The design of these visualizations, however, is premised on implicit assumptions about the reader's sensory, cognitive, and motor abilities. People without these abilities are ultimately disenfranchised, and access to the benefits of data visualization and to the underlying information is limited. Data visualizations, such as statistical charts, diagrams, and maps, are an effective means to represent, analyze, and explore data as well as identify and communicate insights. They take advantage of the human visual system’s high bandwidth, parallel processing, and ability to quickly recognize patterns. For instance, a table of numbers may be hard to understand, while those same numbers shown in a graphic form (such as a line chart) will immediately reveal a steadily increasing trend. For these reasons, interactive data visualization is central to both exploratory data analysis and dataD Insights → Lack of accessible access to data visualizations is a significant equity issue. → It's not only visual impairments that can restrict access but also other kinds of disabilities including cognitive and learning disabilities, and motor disabilities. → Overcoming this challenge requires visualization practitioners, visualization and accessibility researchers, and the relevant disability communities to work together.
Despite the large body of work in accessibility concerning the design of novel navigation technologies, little is known about commonly available technologies that people with visual impairments currently use for navigation. We address this gap with a qualitative study consisting of interviews with 23 people with visual impairments, ten of whom also participated in a follow-up diary study. We develop the idea of complementarity first introduced by Williams et al. [53] and find that in addition to using apps to complement mobility aids, technologies and apps complemented each other and filled in for the gaps inherent in one another. Furthermore, the complementarity between apps and other apps/aids was primarily the result of the differences in information and modalities in which this information is communicated by apps, technology and mobility aids. We propose design recommendations to enhance this complementarity and guide the development of improved navigation experiences for people with visual impairments.
Christian Holz合作论文数Department of Computer Science, Eidgenössische Technische Hochschule Zürich;Sensing, Interaction & Perception Lab, Eidgenössische Technische Hochschule Zürich4