Describe an animal without using the verb look. Can you effectively provide an alternative method for interpreting complex microscopy images while preserving the length scale? The world is filled with features too small for our eyes to see: the setae on a gecko's feet, the cuticles covering a rat's whisker, or the fuzziness of a bat's wing. Furthermore, these structures are non-homogeneous, often shifting from stiff to soft. We provide a workflow for producing low-data, low-cost, and open-source lithograph files, allowing tactile accessibility in microscopy images. The lithographs made with this workflow can be printed on a 350 USD 3D printer using 3D files under 100 Mb, for a total cost per print of 0.75 USD. This work seeks to leverage advanced 3D printing to create tactile graphics and art that make science more accessible and enable tactile exploration of biological structures. This framework in this text is aligned with a GitHub repository that will be constantly updated, allowing tactile media to be created as 3D printing and lithography become more streamlined in the years to come.
We propose a novel task, hierarchical instance tracking, which entails tracking all instances of predefined categories of objects and parts, while maintaining their hierarchical relationships. We introduce the first benchmark dataset supporting this task, consisting of 2,765 unique entities that are tracked in 552 videos and belong to 40 categories (across objects and parts). Evaluation of seven variants of four models tailored to our novel task reveals the new dataset is challenging. Our dataset is available at https://vizwiz.org/tasks-and-datasets/hierarchical-instance-tracking/
Describe an animal without using the verb look. The world is filled with features too small for our eyes to see: the setae on a gecko's feet, the cuticles covering a rat's whisker, or the fuzziness of a bat's wing. Can you effectively provide an alternative method for interpreting complex microscopy images while preserving the length scale? In this work, we provide a fully open-source lithograph workflow, allowing rapid creation of three-dimensional (3D) printable lithograph files for tactile accessibility for generalized images. The lithographs made in this workflow utilize a customizable API, allowing users to manually input various output variables (resolution, relief height, and blur) while outputting various file types (STL, STP, 3MF, and GLB). This workflow is validated on several different commercial 3D printers. This work seeks to leverage 3D printing to create tactile graphics and art, making science more accessible and enabling novel tactile exploration of biological structures. This workflow is paired with an open-source GitHub repository for Tactile Accessible Microscopy Printing (TAMP-OS; https://github.com/nagova/TAMP-OS). The TAMP-OS workflow includes a step-by-step guide, allowing other users to branch and augment the code for their customizable use cases while maintaining an open-source codebase allowing users to adapt all parts of the workflow for their needs.
Individuals who are blind or have low vision (BLV) are at a heightened risk of sharing private information if they share photographs they have taken. To facilitate developing technologies that can help them preserve privacy, we introduce BIV-Priv-Seg, the first localization dataset originating from people with visual impairments that shows private content. It contains 1,028 images with segmentation annotations for 16 private object categories. We first characterize BIV-Priv-Seg and then evaluate modern models' performance for locating private content in the dataset. We find modern models struggle most with locating private objects that are not salient, small, and lack text as well as recognizing when private content is absent from an image. We facilitate future extensions by sharing our new dataset with the evaluation server at https://vizwiz.org/tasks-and-datasets/object-localization/
In this paper, we share a case study of using critical disability studies to design for chronic illness. Specifically, we draw from the Political/Relational model of disability to explore designs for rest. Through a 6-week design collaboration, we brought together different design perspectives and worked through tensions between utility-oriented product design approaches and critical disability approaches for access. Our process yielded three design strategies: 1) moving from designing a product to designing a provocation; 2) using mapping as a process of building collective understanding of the built and social environment; and 3) re-imagining institutional practices around access as a starting point for design. We end by unpacking tensions in our design process and sharing some reflections on how to critically design for access in HCI.
3D printing in principle enables Blind and Low-Vision users to create tactile reference materials and objects, but inaccessible software creates a major barrier for these users. Many graphical user interface (GUI) elements in desktop 3D printing software are undetectable by standard accessibility APIs, such as UI Automation, used by screen readers to expose interfaces. We compare the hierarchies of API-detected elements with those obtained from the software source code, emphasizing detectability as a key metric in desktop accessibility. We assess the screen reader-based operability of core tasks in the 3D printing workflow, such as model positioning and slicing, and identify interface design and framework implementation patterns linked to accessibility failures. This evaluation is conducted on three open-source 3D printing software Ultimaker Cura, PrusaSlicer, and Bambu Studio - using the NVDA screen reader. These findings provide framework-specific insights to inform retrofit and redesign decisions, contributing to accessible interaction paradigms in fabrication.
This paper explores the design space of tactile interfaces made of paper electronics for eyes-free scenarios. We conducted a user study to examine the usability and preferences of paper-based tactile interfaces, focusing on on/off buttons, sliders, and pressure sensors, and identified challenges in eyes-free usage. To address these challenges, we led a participatory workshop with eight design students, resulting in a set of design guidelines. Our recommendations cover: (1) highlighting a starting point; (2) reducing hand travel; (3) allowing more tolerance; (4) assisting multi-touch; (5) using simple patterns; (6) optimizing size and gaps; (7) assisting precise touch; and (8) balancing comfort and use. These guidelines are intended for a broad range of designers, including novices interested in paper-based tactile interfaces. To demonstrate how the guidelines can be applied, we present three proof-of-concept application examples that employ them.
In this paper, we contribute three design manifestos that start from our queer, crip experiences to resist dominant designs and practices of productivity. Through our manifestos, we explore tensions in glitching three technologies of productivity (Mendeley, Figma, and ChatGPT) by reorienting their intended uses and design scripts. By sharing our perspectives and design processes, we invite new ways of relating to technologies of productivity, offer design provocations for queering and cripping technologies in HCI, and call for building intersectional coalitions that contribute towards a slow, non-linear resistance.
Metaphors enrich language by allowing us to express complex ideas through familiar concepts, enhancing both understanding and creativity in communication. Crossmodal metaphors are metaphors where one sensory modality is understood in terms of another (e.g, a sharp smell). Crossmodality is an integral part of how we make sense of and create meaning about the world. However, there is a lack of research on how children generate crossmodal metaphors and the interpretation of such metaphors. We present Sense-O-Nary, a game we designed to explore how children react when asked to create crossmodal metaphors in a novel environment. Children are presented with one sensory input and then asked to describe it using a different sense, for another team to guess what the original sensory input is. We engaged children (n=65, aged 8-10) to play this crossmodal metaphor generation game. We qualitatively analysed children’s exchange of crossmodal metaphors to define a set of crossmodal association strategies and then use this to categorise the metaphors they created. We discuss how engaging with crossmodal metaphors can enhance children’s linguistic development and how our findings can inform the design of interactions that involve multiple senses.
Collaborative ideation plays a vital role in driving creativity and innovation across various professional and educational contexts. This study investigates the experiences of disabled individuals within the collaborative ideation process, specifically examining their utilization of digital whiteboarding tools. Through interviews with 19 professionals and academics with disabilities, alongside a thematic analysis of online forum posts for two popular digital whiteboarding platforms (Miro and Figma), we delve into the access barriers encountered by disabled individuals and the strategies they employ to create access in collaborative ideation. Our findings illuminate the multifaceted nature of access barriers, encompassing issues such as inaccessible visual features, technology-induced discomfort, unstructured nature of freeform content, and complex communication setups. Furthermore, we uncover the intricate dynamics involved in negotiating diverse access needs and conflicts within teams involving people with different disabilities. Through this analysis, we highlight tensions around proficiency with inaccessible technologies stemming from ableist standards of professional success and discuss the implications of our findings for the design of accessible collaborative ideation systems.
Blind individuals commonly share photos in everyday life. Despite substantial interest from the blind community in being able to independently obfuscate private information in photos, existing tools are designed without their inputs. In this study, we prototyped a preliminary screen reader-accessible obfuscation interface to probe for feedback and design insights. We implemented a version of the prototype through off-the-shelf AI models (e.g., SAM, BLIP2, ChatGPT) and a Wizard-of-Oz version that provides human-authored guidance. Through a user study with 12 blind participants who obfuscated diverse private photos using the prototype, we uncovered how they understood and approached visual private content manipulation, how they reacted to frictions such as inaccuracy with existing AI models and cognitive load, and how they envisioned such tools to be better designed to support their needs (e.g., guidelines for describing visual obfuscation effects, co-creative interaction design that respects blind users’ agency).
While audio description (AD) is the standard approach for making videos accessible to blind and low vision (BLV) people, existing AD guidelines do not consider BLV users' varied preferences across viewing scenarios. These scenarios range from how-to videos on YouTube, where users seek to learn new skills, to historical dramas on Netflix, where a user's goal is entertainment. Additionally, the increase in video watching on mobile devices provides an opportunity to integrate nonverbal output modalities (e.g., audio cues, tactile elements, and visual enhancements). Through a formative survey and 15 semi-structured interviews, we identified BLV people's video accessibility preferences across diverse scenarios. For example, participants valued action and equipment details for how-to videos, tactile graphics for learning scenarios, and 3D models for fantastical content. We define a six-dimensional video accessibility design space to guide future innovation and discuss how to move from "one-size-fits-all" paradigms to scenario-specific approaches.
The relentless pace of video production exacerbates the digital accessibility gap that individuals who are blind or low vision (BLV) face on a daily basis, resulting in disproportionate exclusion from community opportunities and risk management. Whereas previous automated audio description (AD) systems provide single-tool approaches for delivering minimum viable description (MVD) or delivering on-demand visual question answering (VQA), we present a tandem AI-based AD tool that combines MVD and on-demand VQA. A user study with 26 BLV individuals explored how the tandem system may be used under the conditions of delivering MVD and/or on-demand VQA with AI-only or human-in-the-loop support. When each tool was used in isolation, AI-only conditions scored significantly lower in both user enjoyment and comprehension. When used in tandem, AI-only conditions matched outcomes delivered with human-in-the-loop, which suggests that AI-only AD tools may be most effective when both types of tools are used in tandem. A multimodal analysis of interactions with the tandem system revealed areas for system improvement in terms of the timing of AD delivery and accurate content delivery. We discuss how the use of both types of tools in a tandem system can mitigate some of the digital frictions that have plagued efforts in machine learning and automated tools for accessibility.
We present the design and creation of a disability-first dataset, “BIV-Priv,” which contains 728 images and 728 videos of 14 private categories captured by 26 blind participants to support downstream development of artificial intelligence (AI) models. While best practices in dataset creation typically attempt to eliminate private content, some applications require such content for model development. We describe our approach in creating this dataset with private content in an ethical way, including using props rather than participants’ own private objects and balancing multi-disciplinary perspectives (e.g., accessibility, privacy, computer vision) to meet the tangible metrics (e.g., diversity, category, amount of content) to support AI innovations. We observed challenges that our participants encountered during the data collection, including accessibility issues (e.g., understanding foreground vs. background object placement) and issues due to the sensitive nature of the content (e.g., discomfort in capturing some props such as condoms around family members).
Many people who are blind take and post photos to share about their lives and connect with others. Yet, current technology does not provide blind people with accessible ways to handle when private information is unintentionally captured in their images. To explore the technology design in supporting them with this task, we developed a design probe for blind people-ImageAlly- that employs a human-AI hybrid approach to detect and redact private image content. ImageAlly notifies users when potential private information is detected in their images, using computer vision, and enables them to transfer those images to trusted sighted allies to edit the private content. In an exploratory study with pairs of blind participants and their sighted allies, we found that blind people felt empowered by ImageAlly to prevent privacy leakage in sharing images on social media. They also found other benefits from using ImageAlly, such as potentially improving their relationship with allies and giving allies the awareness of the accessibility challenges they face.
Visual assistance technologies provide people who are blind with access to information about their visual surroundings by digitally connecting them to remote humans or artificial intelligence systems that describe visual content such as objects, people, scenes, and text observed in their live image/video feeds. Prior work has revealed that users have concerns about how such technologies handle private visual content captured in their image/video feeds. Yet, it remains unclear how users want technologies to manage such private content. To fill this gap, we interviewed 16 totally blind individuals to learn about their expectations for visual privacy when using visual assistance technologies. Our findings reveal three overarching user-centered expectations associated with visual privacy-preservation in this domain, as well as the broader ethical challenges involved with developing AI-based privacy-preserving visual assistance technologies.
Visual arts play an important role in cultural life and provide access to social heritage and self-enrichment, but most visual arts are inaccessible to blind people. Researchers have explored different ways to enhance blind people's access to visual arts (e.g., audio descriptions, tactile graphics). However, how blind people adopt these methods remains unknown. We conducted semi-structured interviews with 15 blind visual arts patrons to understand how they engage with visual artwork and the factors that influence their adoption of visual arts access methods. We further examined interview insights in a follow-up survey (N=220). We present: 1) current practices and challenges of accessing visual artwork in-person and online (e.g., Zoom tour), 2) motivation and cognition of perceiving visual arts (e.g., imagination), and 3) implications for designing visual arts access methods. Overall, our findings provide a roadmap for technology-based support for blind people's visual arts experiences.
The privacy dimensions of accessibility technologies are often understudied and overlooked. Very little prior research has investigated the privacy concerns of disabled people, and much less has studied the barriers of privacy-preserving techniques. In order to address this gap and bridge between two separate communities (accessibility and privacy), our one-day workshop explores how researchers might design and build technologies that are both accessible and privacy-preserving.
Auditory interfaces increasingly support access to website content, through recent advances in voice interaction. Typically, however, these interfaces provide only limited audio styling, collapsing rich visual design into a static audio output style with a single synthesized voice. To explore the potential for more aesthetic and intuitive sound design for websites, we prompted 14 professional sound designers to create auditory website mockups and interviewed them about their designs and rationale. Our findings reveal their prioritized design considerations (aesthetics and emotion, user engagement, audio clarity, information dynamics, and interactivity), specific sound design ideas to support each consideration (e.g., replacing spoken labels with short, memorable audio expressions), and challenges with applying sound design practices to auditory websites. These findings provide promising direction for how to support designers in creating richer auditory website experiences.