For people with noise sensitivity, everyday soundscapes can be overwhelming. Existing noise-management tools, such as active noise cancellation, typically apply broad suppression across the entire environment, easing discomfort at the expense of situational awareness and participation. We present Sona, a system that reframes acoustic accessibility as selective, user-steerable soundscape mediation, attenuating triggering sounds while preserving other ambient sounds users may want to hear. Sona combines real-time acoustic context sensing to reduce interaction effort during moments of overload, steerable suppression that supports continuous negotiation between comfort and environmental awareness, and end-user personalization that extends filtering beyond fixed sound taxonomies using short recordings from users’ own environments. We describe Sona’s prototype implementation and derive design implications for supporting customizable, context-aware soundscape mediation for people with noise sensitivity.
As virtual 3D environments become more prevalent, equitable access is essential for blind and low-vision (BLV) users, who face challenges with spatial awareness, navigation, and interaction. Prior work has explored supplementing visual information with auditory or haptic modalities, but these methods are static and offer limited support for dynamic, in-context adaptation. Recent advances in generative AI allow users to query and modify 3D scenes via natural language, introducing a paradigm that offers greater flexibility and control for accessibility. We present RAVEN, a system that enables BLV users to issue queries and modification prompts to improve the runtime accessibility of 3D virtual scenes. We evaluated RAVEN with eight BLV people and six Unity developers, generating empirical insights into how conversational programming can support personalized accessibility in 3D environments. Our work highlights both the promise of natural language interaction—intuitive, flexible, and empowering—and the challenges of ensuring reliability, transparency, and trust in generative AI–driven accessibility systems.
For people with noise sensitivity, everyday soundscapes can be overwhelming. Existing tools such as active noise cancellation reduce discomfort by suppressing the entire acoustic environment, often at the cost of awareness of surrounding people and events. We present Sona, an interactive mobile system for real-time soundscape mediation that selectively attenuates bothersome sounds while preserving desired audio. Sona is built on a target-conditioned neural pipeline that supports simultaneous attenuation of multiple overlapping sound sources, overcoming the single-target limitation of prior systems. It runs in real time on-device and supports user-extensible sound classes through in-situ audio examples, without retraining. Sona is informed by a formative study with 68 noise-sensitive individuals. Through technical benchmarking and an in-situ study with 10 participants, we show that Sona achieves low-latency, multi-target attenuation suitable for live listening, and enables meaningful reductions in bothersome sounds while maintaining awareness of surroundings. These results point toward a new class of personal AI systems that support comfort and social participation by mediating real-world acoustic environments.
Background Speech recognition technology is widely used by individuals who are Deaf/deaf and hard-of-hearing (DHH) in everyday communication, but its clinical applications remain underexplored. Communication barriers in health care can compromise safety, understanding, and autonomy for individuals who are DHH. Objective This study aimed to evaluate a real-time speech recognition system (SRS) tailored for clinical settings, examining its usability, perceived effectiveness, and transcription accuracy among users who are DHH. Methods We conducted a pilot study with 10 adults who are DHH participating in mock outpatient encounters using a custom SRS powered by Google’s speech-to-text application programming interface. We used a convergent parallel mixed-methods design, collecting quantitative usability ratings and qualitative interview data during the same study session. These datasets were subsequently merged and jointly interpreted. Participants completed postscenario surveys and structured exit interviews assessing distraction, trust, ease of use, satisfaction, and emotional response. Caption accuracy was benchmarked against professional communication access real-time translation transcripts using word error rate (WER). Because WER assigns equal weight to all tokens, it does not differentiate between routine transcription errors and those involving safety-critical clinical terms (eg, medications or diagnoses). Therefore, WER may underestimate the potential impact of certain errors in medical contexts. Results Across 29 clinical scenario simulations, 86% (25/29) of participants found captions nondistracting, 90% (26/29) reported them easy to follow and trustworthy, and 76% (22/29) were satisfied with the experience. Participants described the SRS as intuitive, emotionally grounding, and preferable to lip reading in masked settings. WER ranged from 12.7% to 22.8%, consistent with benchmarks for automated SRSs. Interviews revealed themes of increased confidence in following clinical conversations and staying engaged despite masked communication. Participants reported less anxiety about missing critical medical information and expressed a strong interest in expanding the tool to real-world settings, especially for older adults or those with cognitive impairments. Conclusions Our findings support the potential of real-time captioning to enhance accessibility and reduce the cognitive and mental burden of communication for individuals who are DHH in clinical care. Participants described the SRS as both functionally effective and personally empowering. While accuracy for complex medical terminology remains a limitation, participants consistently expressed trust in the system and a desire for its integration into clinical care. Future research should explore real-world implementation, domain-specific optimization, and the development of user-centered evaluation metrics that extend beyond transcription fidelity to include trust, autonomy, and communication equity.
Communication Access Realtime Translation (CART) is a widely used captioning technology among deaf and hard of hearing (DHH) individuals, valued for its high accuracy and ability to convey speaker cues and contextual sounds in real time. However, CART performance can degrade in challenging conditions such as background noise, technical jargon, or rapid speech-reducing caption quality and impacting comprehension. We introduce CARTGPT, a real-time captioning system that enhances CART transcripts by leveraging large language models (LLMs) and automatic speech recognition (ASR) input to detect and correct transcription errors. To inform the design of CARTGPT, we conducted a formative study with 10 professional CART captioners to identify common sources of error and their perspectives on using AI for caption correction. We evaluated CARTGPT on a 39.7-hour speech dataset spanning medical, technical, and conversational domains, observing a 5.6% improvement in word accuracy over standard CART and 17.3% over a state-of-the-art ASR model. In a user study with 16 DHH participants, CARTGPT captions were rated as significantly more comprehensible, particularly in technical scenarios, while maintaining real-time responsiveness. These findings demonstrate the potential of LLM-assisted captioning to improve accessibility and comprehension for DHH users in real-world settings.
Digital media is increasingly audio-rich, yet much of its sound content remains inaccessible for deaf and hard-of-hearing (DHH) individuals. While prior work has focused on captioning and sound recognition, little research has explored how sound itself can be transformed to better align with hearing needs, preferences, and contexts for people with partial hearing. In this paper, we present findings from a formative study with 24 DHH participants that examines their experiences and unmet needs around digital media audio. Participants emphasized the importance of features such as selective speaker amplification, contextual sound control, semantic summarization, and adaptive personalization. Based on these insights, we introduce the ReMediaTion Framework, a layered model that articulates user goals, audio transformation dimensions, interaction strategies, contextual modulation, and expressive engagement. Our work provides a foundation for designing future audio accessibility systems that go beyond substitution, empowering DHH users to reshape how they experience and interpret sound.
With the popularity of virtual 3D applications, from video games to educational content and virtual reality scenarios, the accessibility of 3D scene information is vital to ensure inclusive and equitable experiences for all. Previous work include information substitutions like audio description and captions, as well as personalized modifications, but they could only provide predefined accommodations. In this work, we propose SceneGenA11y, a system that responds to the user's natural language prompts to improve accessibility of a 3D virtual scene in runtime. The system primes LLM agents with accessibility-related knowledge, allowing users to explore the scene and perform verifiable modifications to improve accessibility. We conducted a preliminary evaluation of our system with three blind and low-vision people and three deaf and hard-of-hearing people. The results show that our system is intuitive to use and can successfully improve accessibility. We discussed usage patterns of the system, potential improvements, and integration into apps. We ended with highlighting plans for future work.
Sound recognition enhances safety, social interaction, and situational awareness for deaf and hard of hearing (DHH) individuals. However, existing sound recognition technologies primarily classify sounds into predefined categories (e.g., door opening, speech), which fail to capture the full complexity of real-world auditory scenes (e.g., temporal variations, sound transitions, overlapping sound layers). In this work, we introduce SoundNarratives, a realtime system that generates rich, contextual auditory scene descriptions tailored to DHH users. We began with conducting a formative study with 10 DHH participants to identify nine key auditory scene parameters (e.g., sound class, loudness, emotion, semantic description), and used these insights to guide prompt engineering with a state-of-the-art audio language model. A user study with 10 DHH participants demonstrated a significant preference for SoundNarratives over a baseline model, along with a potential for improved confidence and situational awareness.
Augmented reality (AR) has shown promise for supporting Deaf and hard-of-hearing (DHH) individuals by captioning speech and visualizing environmental sounds, yet existing systems do not allow users to create personalized sound visualizations. We present SonoCraftAR, a proof-of-concept prototype that empowers DHH users to author custom sound-reactive AR interfaces using typed natural language input. SonoCraftAR integrates real-time audio signal processing with a multi-agent LLM pipeline that procedurally generates animated 2D interfaces via a vector graphics library. The system extracts the dominant frequency of incoming audio and maps it to visual properties such as size and color, making the visualizations respond dynamically to sound. This early exploration demonstrates the feasibility of open-ended sound-reactive AR interface authoring and discusses future opportunities for personalized, AI-assisted tools to improve sound accessibility.
Current ASR systems struggle to reliably recognize the speech of Deaf and Hard of Hearing (DHH) individuals, particularly in realtime communication. Existing personalization methods typically require extensive pre-recorded data and place the burden entirely on DHH users. We present EvolveCaptions, a live ASR adaptation system that supports collaborative, in-the-moment personalization. Hearing participants correct ASR errors during conversation, and the system generates short, phonetically relevant phrases for the DHH speaker to record. These recordings are then used to iteratively fine-tune the ASR model. In a preliminary evaluation, our system reduced word error rate from 0.53 to 0.27 over four adaptation rounds with minimal user effort. This work introduces a low-effort, socially collaborative method for adapting ASR to diverse DHH voices in real-world settings.
Automatic Speech Recognition (ASR) systems often fail to accurately transcribe speech from Deaf and Hard of Hearing (DHH) individuals, especially during real-time conversations. Existing personalization approaches typically require extensive pre-recorded data and place the burden of adaptation on the DHH speaker. We present EvolveCaptions, a real-time, collaborative ASR adaptation system that supports in-situ personalization with minimal effort. Hearing participants correct ASR errors during live conversations. Based on these corrections, the system generates short, phonetically targeted prompts for the DHH speaker to record, which are then used to fine-tune the ASR model. In a study with 12 DHH and six hearing participants, EvolveCaptions reduced Word Error Rate (WER) across all DHH users within one hour of use, using only five minutes of recording time on average. Participants described the system as intuitive, low-effort, and well-integrated into communication. These findings demonstrate the promise of collaborative, real-time ASR adaptation for more equitable communication.
We demonstrate CapTune, an interactive system that enables personalized transformation of non-speech captions for deaf and hard of hearing (DHH) viewers while preserving creators' narrative intents. Unlike traditional one-size-fits-all captioning approaches, CapTune allows creators to define transformation boundaries and viewers to customize captions across four dimensions: level of detail, expressiveness, sound representation method, and genre alignment. Our demonstration showcases: (1) the Creator Tool for defining acceptable caption transformation ranges, and (2) the Viewer Client enabling real-time caption personalization during video playback. Demo attendees can experience both interfaces hands-on, exploring how generative AI can support more inclusive and engaging cinematic experiences for DHH audiences.
Previous VR sound accessibility work substituted sounds with visual or haptic output to increase VR accessibility for deaf and hard of hearing (DHH) people. However, deafness occurs on a spectrum, and many DHH people (e.g., those with partial hearing) can also benefit greatly from having more control over the audio instead of substituting it with another modality. In this paper, we explore the possibilities of modifying sounds in VR to support DHH people. To understand the best modification features for this goal, we designed and implemented 18 VR sound modification tools spanning four categories, including prioritizing sounds, modifying sound parameters, providing spatial assistance, and adding additional sounds. We evaluated our tools in five diverse VR scenarios with 10 DHH people, finding that our tool can improve DHH users’ VR experience, but could be further improved by providing more customization options and decreasing distraction. We then compiled a Unity toolkit from select tools and conducted a preliminary evaluation with six Unity VR developers. Findings show that our toolkit is easy to use and debug but could be enhanced through modularization and better documentation. We close by discussing further implications of sound modification in VR.
Current sound recognition systems for deaf and hard of hearing (DHH) people identify sound sources or discrete events. However, these systems do not distinguish similar sounding events (e.g., a patient monitor beep vs. a microwave beep). In this paper, we introduce HACS, a novel futuristic approach to designing human-AI sound awareness systems. HACS assigns AI models to identify sounds based on their characteristics (e.g., a beep) and prompts DHH users to use this information and their contextual knowledge (e.g., “I am in a kitchen”) to recognize sound events (e.g., a microwave). As a first step for implementing HACS, we articulated a sound taxonomy that classifies sounds based on sound characteristics using insights from a multi-phased research process with people of mixed hearing abilities. We then performed a qualitative (with 9 DHH people) and a quantitative (with a sound recognition model) evaluation. Findings demonstrate the initial promise of HACS for designing accurate and reliable human-AI systems.
Noise sensitivity is a frequently reported characteristic in many autistic individuals. While strategies like sound isolation (e.g., noise-canceling headphones) and avoidance behaviors (e.g., leaving a crowded room) can help, they can reduce situational awareness and limit social engagement. In this paper, we examine an alternate approach to managing noise sensitivity: introducing ambient background sounds to reduce the perception of disruptive noises, i.e., sound masking. Through two studies (with ten and nine autistic individuals respectively), we investigated the autistic individuals’ preferred sound masks (e.g., white noise, brown noise, calming water sounds) for different contexts (e.g., traffic, speech) and elicited reactions for a future interactive tool to deliver effective sound masks. Our findings have implications not just for the accessibility community, but also for designers and researchers working on sound augmentation technology.
Communication Access Realtime Translation (CART) is a commonly used real-time captioning technology used by deaf and hard of hearing (DHH) people, due to its accuracy, reliability, and ability to provide a holistic view of the conversational environment (e.g., by displaying speaker names). However, in many real-world situations (e.g., noisy environments, long meetings), the CART captioning accuracy can considerably decline, thereby affecting the comprehension of DHH people. In this work-in-progress paper, we introduce CARTGPT, a system to assist CART captioners in improving their transcription accuracy. CARTGPT takes in errored CART captions and inaccurate automatic speech recognition (ASR) captions as input and uses a large language model to generate corrected captions in real-time. We quantified performance on a noisy speech dataset, showing that our system outperforms both CART (+5.6% accuracy) and a state-of-the-art ASR model (+17.3%). A preliminary evaluation with three DHH users further demonstrates the promise of our approach.
To increase VR sound accessibility for deaf and hard of hearing users, previous work has substituted sounds with visual or haptic feedback. However, many DHH people (e.g., those with partial hearing) can also benefit from modifying audio (e.g., changing volume based on priorities) instead of fully substituting it with another modality. In this demo paper, we present a toolkit that allows modifying sounds in VR to support DHH people. We designed and implemented 18 VR sound modification tools spanning four categories, including prioritizing sounds, modifying sound parameters, providing spatial assistance, and adding additional sounds. We present five demo scenarios with tools incorporated, covering common VR use cases.
Previous VR sound accessibility work have substituted sounds with visual or haptic output to increase VR accessibility for deaf and hard of hearing (DHH) people. However, deafness occurs on a spectrum, and many DHH people (e.g., those with partial hearing) can also benefit from manipulating audio (e.g., increasing volume at specific frequencies) instead of substituting it with another modality. In this demo paper, we present a toolkit that allows modifying sounds in VR to support DHH people. We designed and implemented 18 VR sound modification tools spanning four categories, including prioritizing sounds, modifying sound parameters, providing spatial assistance, and adding additional sounds. Evaluation of our tools with 10 DHH users across five diverse VR scenarios reveal that our toolkit can improve DHH users’ VR experience but could be further improved by providing more customization options and decreasing cognitive load. We then compiled a Unity toolkit and conducted a preliminary evaluation with six Unity VR developers. Preliminary insights show that our toolkit is easy to use but could be enhanced through modularization.
In this panel, we follow up on conversations that have been happening around hybrid environments, conducted in open sessions by the ACM Special Interest Group on Computer-Human Interaction (SIGCHI) Executive Committee (EC), at SIGs at CHI 2022 and CHI 2023 and on twitter, facebook, medium and similar media. The COVID-19 pandemic led to a shift to virtual conferences. As we go back to in-person events, it is important to reflect on what we have learned about these configurations, the types of events we desire, and how hybrid intersects with SIGCHI values such as increased accessibility, sustainability and inclusion. With this panel, we expect to engage the CSCW community in the discussion of what lies beyond the current in-person format, the possibilities created by the hybrid and what other innovations might further these values.