Interactive digital maps have revolutionized how people travel and learn about the world; however, they rely on pre-existing structured data in GIS databases (e.g., road networks, POI indices), limiting their ability to address geo-visual questions related to what the world looks like. We introduce our vision for Geo-Visual Agents–multimodal AI agents capable of understanding and responding to nuanced visual-spatial inquiries about the world by analyzing large-scale repositories of geospatial images, including streetscapes (e.g., Google Street View), place-based photos (e.g., TripAdvisor, Yelp), and aerial imagery (e.g., satellite photos) combined with traditional GIS data sources. We define our vision, describe sensing and interaction approaches, provide three exemplars, and enumerate key challenges and opportunities for future work.
The rapid emergence of generative AI has changed the way that technology is designed, constructed, maintained, and evaluated. Decisions made when creating AI-powered systems may impact some users disproportionately, such as people with disabilities. In this paper, we report on an interview study with 25 AI practitioners across multiple roles (engineering, research, UX, and responsible AI) about how their work processes and artifacts may impact end users with disabilities. We found that practitioners experienced friction when triaging problems at the intersection of responsible AI and accessibility practices, navigated contradictions between accessibility and responsible AI guidelines, identified gaps in data about users with disabilities, and gathered support for addressing the needs of disabled stakeholders by leveraging informal volunteer and community groups within their company. Based on these findings, we offer suggestions for new resources and process changes to better support people with disabilities as end users of AI.
This paper reports on disability representation in images output from text-to-image (T2I) generative AI systems. Through eight focus groups with 25 people with disabilities, we found that models repeatedly presented reductive archetypes for different disabilities. Often these representations reflected broader societal stereotypes and biases, which our participants were concerned to see reproduced through T2I. Our participants discussed further challenges with using these models including the current reliance on prompt engineering to reach satisfactorily diverse results. Finally, they offered suggestions for how to improve disability representation with solutions like showing multiple, heterogeneous images for a single prompt and including the prompt with images generated. Our discussion reflects on tensions and tradeoffs we found among the diverse perspectives shared to inform future research on representation-oriented generative AI system evaluation metrics and development processes.
This study examined how to design tools that build independence with Blind or Visually Impaired (BVI) children and their families. Beyond core academics, BVI children require instruction on independent living skills, with their curriculum necessitating parent-school cooperation to support continued education at home. However, most technology for BVI children focus on academics, spatial orientation, and physical mobility. In this work, we aim to design a tool for independence that aligns with existing familial structures and activities. Through interviews and diary studies with fve families, we explored development practices parents used with their BVI children, parent-teacher relationships, and how a prompting and refection tool supported family goals. This study highlights home routines and independence skills that beneft from customized prompting, how activity prompts can encourage parents to scale back their assistance and propel independence, and how refection builds optimism and empowers parents in the learning process.
Cooking is an essential activity that enhances quality of life by enabling individuals to prepare their own meals. However, cooking often requires multitasking between cooking tasks and following instructions, which can be challenging to cooks with vision impairments if recipes or other instructions are inaccessible. To explore the practices and challenges of recipe access while cooking, we conducted semi-structured interviews with 20 people with vision impairments who have cooking experience and four cooking instructors at a vision rehabilitation center. We also asked participants to edit and give feedback on existing recipes. We revealed unique practices and challenges to accessing recipe information at different cooking stages, such as the heavy burden of hand-washing to interact with recipe readers. We also presented the preferred information representation and structure of recipes. We then highlighted design features of technological supports that could facilitate the development of more accessible kitchen technologies for recipe access. Our work contributes nuanced insights and design guidelines to enhance recipe accessibility for people with vision impairments.
Generative AI (GAI) is proliferating, and among its many applications are to support creative work (e.g., generating text, images, music) and to enhance accessibility (e.g., captions of images and audio). As GAI evolves, creatives must consider how (or how not) to incorporate these tools into their practices. In this paper, we present interviews at the intersection of these applications. We learned from 10 creatives with disabilities who intentionally use and do not use GAI in and around their creative work. Their mediums ranged from audio engineering to leatherwork, and they collectively experienced a variety of disabilities, from sensory to motor to invisible disabilities. We share cross-cutting themes of their access hacks, how creative practice and access work become entangled, and their perspectives on how GAI should and should not fit into their workflows. In turn, we offer qualities of accessible creativity with responsible AI that can inform future research.
AI-generated images are proliferating as a new visual medium. However, state-of-the-art image generation models do not output alternative (alt) text with their images, rendering them largely inaccessible to screen reader users (SRUs). Moreover, less is known about what information would be most desirable to SRUs in this new medium. To address this, we invited AI image creators and SRUs to evaluate alt text prepared from various sources and write their own alt text for AI images. Our mixed-methods analysis makes three contributions. First, we highlight creators’ perspectives on alt text, as creators are well-positioned to write descriptions of their images. Second, we illustrate SRUs’ alt text needs particular to the emerging medium of AI images. Finally, we discuss the promises and pitfalls of utilizing text prompts written as input for AI models in alt text generation, and areas where broader digital accessibility guidelines could expand to account for AI images.
AbstractAccelerating text input in augmentative and alternative communication (AAC) is a long-standing area of research with bearings on the quality of life in individuals with profound motor impairments. Recent advances in large language models (LLMs) pose opportunities for re-thinking strategies for enhanced text entry in AAC. In this paper, we present SpeakFaster, consisting of an LLM-powered user interface for text entry in a highly-abbreviated form, saving 57% more motor actions than traditional predictive keyboards in offline simulation. A pilot study on a mobile device with 19 non-AAC participants demonstrated motor savings in line with simulation and relatively small changes in typing speed. Lab and field testing on two eye-gaze AAC users with amyotrophic lateral sclerosis demonstrated text-entry rates 29–60% above baselines, due to significant saving of expensive keystrokes based on LLM predictions. These findings form a foundation for further exploration of LLM-assisted text entry in AAC and other user interfaces.
Individuals with vision impairments employ a variety of strategies for object identification, such as pans or soy sauce, in the culinary process. In addition, they often rely on contextual details about objects, such as location, orientation, and current status, to autonomously execute cooking activities. To understand how people with vision impairments collect and use the contextual information of objects while cooking, we conducted a contextual inquiry study with 12 participants in their own kitchens. This research aims to analyze object interaction dynamics in culinary practices to enhance assistive vision technologies for visually impaired cooks. We outline eight different types of contextual information and the strategies that blind cooks currently use to access the information while preparing meals. Further, we discuss preferences for communicating contextual information about kitchen objects as well as considerations for the deployment of AI-powered assistive technologies.
Technology companies continue to invest in efforts to incorporate responsibility in their Artificial Intelligence (AI) advancements, while efforts to audit and regulate AI systems expand. This shift towards Responsible AI (RAI) in the tech industry necessitates new practices and adaptations to roles—undertaken by a variety of practitioners in more or less formal positions, many of whom focus on the user-centered aspects of AI. To better understand practices at the intersection of user experience (UX) and RAI, we conducted an interview study with industrial UX practitioners and RAI subject matter experts, both of whom are actively involved in addressing RAI concerns throughout the early design and development of new AI-based prototypes, demos, and products, at a large technology company. Many of the specific practices and their associated challenges have yet to be surfaced in the literature, and distilling them offers a critical view into how practitioners' roles are adapting to meet present-day RAI challenges. We present and discuss three emerging practices in which RAI is being enacted and reified in UX practitioners' everyday work. We conclude by arguing that the emerging practices, goals, and types of expertise that surfaced in our study point to an evolution in praxis, with associated challenges that suggest important areas for further research in HCI.
Making engages young people with the material world and reflection-in-action, creating promising science learning contexts.Emphasizing relational and social dimensions of making, we conducted a week-long workshop for middle schoolers who are current and aspiring pet companions.Supporting participants' inquiry into pets' senses and related behaviors, we asked them to work on maker projects meant to improve their pets' lives.Following a qualitative analysis of participants' positioning in relation to their pets, we present case studies of two female participants' positioning.We find that through the process of making, the two participants demonstrated an increased awareness of pets' biology and related behavior and their personal interests in pet care, while also differing in what aspects of human-pet relations they focused on.We conclude that through making, especially in contexts with a robust relational draw, youth become attentive to complex and otherwise difficult-to-notice transactions central to taking care of pets.
Text-to-image generation models have grown in popularity due to their ability to produce high-quality images from a text prompt. One use for this technology is to enable the creation of more accessible art creation software. In this paper, we document the development of an alternative user interface that reduces the typing effort needed to enter image prompts by providing suggestions from a large language model, developed through iterative design and testing within the project team. The results of this testing demonstrate how generative text models can support the accessibility of text-to-image models, enabling users with a range of abilities to create visual art.
People with cognitive disabilities may experience challenges in consistently performing daily activities because they skip steps, struggle to track progress, or lack the motivation to complete them. These challenges are often along a range; people need assistive devices customized to their specific needs. However, existing assistive technologies, like prompting systems, lack the capabilities to customize support for diverse needs. With the advent of smart home devices, there are opportunities to design prompting systems that support diverse accessibility and motivational needs, thereby supporting the regular practice of daily activities. To understand design factors for such devices, we interviewed adults with cognitive disabilities, parents, and caregivers. Our participants described their needs for future prompting systems, including structuring tasks, supporting motivation, and introducing community support. This paper presents insights and design suggestions for context-aware assistive technologies that could help people with cognitive disabilities regularly perform everyday activities.
Based on the widely recognized situated nature of identity and youth as social producers and products, this qualitative case study reports findings from a week-long informal pet-sciences workshop for middle schoolers who have existing relationships with pets or a strong interest in future pet companionship.Mindful of the structure-agency dialectic, we analyze youth's wayfaring and trajectories of identification as they learn about their pets at the workshop, accounting for how youth see themselves and their pets and are seen by others.In contrast to a commonly assumed analytic directionality seeing people as moving towards or away from STEM, we find that there were different ways for youth to meaningfully engage themselves in learning about their pets at the workshop.We conclude that attention to fluidity in youth's identifications can inform us, the adults in the community, of the need to affirm the many possible trajectories that youth may follow.
Large language models (LLMs) trained on real-world data can inadvertently reflect harmful societal biases, particularly toward historically marginalized communities. While previouswork has primarily focused on harms related to age and race, emerging research has shown that biases toward disabled communities exist. This study extends prior work exploring the existence of harms by identifying categories of LLM-perpetuated harms toward the disability community. We conducted 19 focus groups, during which 56 participants with disabilities probed a dialog model about disability and discussed and annotated its responses. Participants rarely characterized model outputs as blatantly offensive or toxic. Instead, participants used nuanced language to detail how the dialog model mirrored subtle yet harmful stereotypes they encountered in their lives and dominant media, e.g., inspiration porn and able-bodied saviors. Participants often implicated training data as a cause for these stereotypes and recommended training the model on diverse identities from disability-positive resources. Our discussion further explores representative data strategies to mitigate harm related to different communities through annotation co-design with ML researchers and developers.
BackgroundNatureculture (Fuentes, 2010; Haraway, 2003) constructs offer a powerful framework for science education to explore learners' interactions with and understanding of the natural world. Technologies such as Augmented Reality (AR) designed to reveal pets' sensory worlds and companionship with pets can facilitate learners' harmonious relationships with significant others in naturecultures.MethodsAt a two-week virtual summer camp, we engaged teens in inquiring into dogs' and cats' senses using selective color filters, investigations, experience design projects, and understanding how the umwelt (von Uexkull, 2001) of pets impacts their lives with humans. We qualitatively analyzed participants' talk, extensive notes, and projects completed at the workshop.FindingsWe found that teens engaged in the science and engineering practices of planning and carrying out investigations, constructing explanations and designing solutions, and questioning while investigating specific aspects of their pets' lives. Further, we found that teens checking and taking pets' perspectives while caring for them shaped their productive engagement in these practices. The relationship between pets and humans facilitated an ecological and relational approach to science learning.ContributionOur findings suggest that relational practices of caring and perspective-taking coexist with scientific practices and enrich scientific inquiry.
Users of augmentative and alternative communication (AAC) devices sometimes find it difficult to communicate in real time with others due to the time it takes to compose messages. AI technologies such as large language models (LLMs) provide an opportunity to support AAC users by improving the quality and variety of text suggestions. However, these technologies may fundamentally change how users interact with AAC devices as users transition from typing their own phrases to prompting and selecting AI-generated phrases. We conducted a study in which 12 AAC users tested live suggestions from a language model across three usage scenarios: extending short replies, answering biographical questions, and requesting assistance. Our study participants believed that AI-generated phrases could save time, physical and cognitive effort when communicating, but felt it was important that these phrases reflect their own communication style and preferences. This work identifies opportunities and challenges for future AI-enhanced AAC devices.
Accessibility solutions often focus on the experiences of people with more severe disabilities, such as those who are unable to perform certain tasks unassisted. However, disability exists on a spectrum, and people with moderate disabilities may be overlooked when studying or sampling for differences. As a result, these individuals and their needs are excluded from relevant research. In this study, we interviewed 12 adults with mild-to-moderate dexterity impairments about their experiences using smartphones and other mobile devices. While our participants frequently experienced accessibility challenges, they struggled to know where to find help, in part because of discomfort with traditional labels of disability and accessibility. We found four key themes: (1) There were large gaps in available and usable accessibility tools for this population, (2) Users were unlikely to seek out accessibility features due to complex disability identity, (3) Contextual concerns impacted mobile device use, and (4) Users relied on self-created adaptations and modifications to improve usability. We suggest that individuals with mild-to-moderate dexterity challenges are a unique cohort that would benefit from further consideration from the accessibility community and accessibility features that support their needs.
Accelerating communication for users with severe motor and speech impairments, in particular for eye-gaze-based augmentative and alternative communication (AAC) device users, is a longstanding area of research. However, observation of such users' communication over extended durations has been limited. This case study presents the real-world experience of developing and field-testing a tool for observing and curating the gaze typing-based communication of an eye-gaze AAC user with amyotrophic lateral sclerosis (ALS). With the intent to observe and develop technology to accelerate eye-gaze typed communication, we designed a tool and a protocol called the SpeakFaster Observer to measure everyday conversational text entry by the gaze-typing user, as well as several consenting conversation partners of the AAC user. We detail the design of the Observer software and data curation protocol, along with considerations for privacy protection. The deployment of the data protocol from November 2021 to April 2022 yielded a rich dataset of gaze-based AAC text entry from everyday life, consisting of 130+ hours of gaze keystrokes and 5,000+ curated speech utterances from the AAC user and the conversation partners. We present the key statistics of the data, including the speed (8.1 ± 3.9 words per minute) and keystroke saving rate (-0.14 ± 0.83) of gaze typing, patterns of utterance repetition and reuse, and the temporal dynamics of conversation turn-taking in gaze-based communication. We share our findings and also open source our data collection tools to further research in this domain.
Finding ways to accelerate text input for individuals with profound motor impairments has been a long-standing area of research. Closing the speed gap for augmentative and alternative communication (AAC) devices such as eye-tracking keyboards is important for improving the quality of life for such individuals. Recent advances in neural networks of natural language pose new opportunities for re-thinking strategies and user interfaces for enhanced text-entry for AAC users. In this paper, we present SpeakFaster, consisting of large language models (LLMs) and a co-designed user interface for text entry in a highly-abbreviated form, allowing saving 57% more motor actions than traditional predictive keyboards in offline simulation. A pilot study with 19 non-AAC participants typing on a mobile device by hand demonstrated gains in motor savings in line with the offline simulation, while introducing relatively small effects on overall typing speed. Lab and field testing on two eye-gaze typing users with amyotrophic lateral sclerosis (ALS) demonstrated text-entry rates 29-60% faster than traditional baselines, due to significant saving of expensive keystrokes achieved through phrase and word predictions from context-aware LLMs. These findings provide a strong foundation for further exploration of substantially-accelerated text communication for motor-impaired users and demonstrate a direction for applying LLMs to text-based user interfaces.