Interest in learning American Sign Language (ASL) is growing across higher education institutions in North America, as reflected in rising enrollments. Yet this growth is constrained by limited program availability and few opportunities to practice outside the classroom. AI-based technologies show promise for supporting ASL learning, but educators – who bring essential pedagogical, linguistic, and cultural expertise – have been largely absent from conversations on the design of these tools, with prior work focusing primarily on learners. To address this, we conducted formative interviews with eleven Deaf and one hearing ASL instructor, followed by two focus groups with six Deaf educators, to examine how AI tools could support ASL education. Findings revealed priorities for technology design and considerations for integration into existing pedagogical practices, with attention to curricular, linguistic, and access factors. We offer insights for designing and researching technologies aimed at (1) providing adaptive, structured feedback on signing performance and (2) supporting immersive conversational practice with virtual signing partners.
Hypertension disparities persist despite effective lifestyle and pharmacological therapies, underscoring the need for strategies that deliver evidence-based care more equitably. This commentary examines the potential for patient-facing large language model-based systems to reduce hypertension disparities and proposes an implementation science approach to their development, integration, and evaluation. These systems could expand access to personalized health education, behavioral support, care navigation, and clinical escalation between traditional care encounters. However, biased training data, inaccurate or overly agreeable responses, unequal digital access and literacy, and mistrust of automated systems could reinforce existing inequities. Technical safeguards are necessary but insufficient. Equitable implementation should begin with meaningful community engagement and human-centered co-design, incorporate trusted knowledge sources and human oversight, complement team-based care, and iteratively address local barriers and facilitators. Evaluation should assess both clinical effectiveness and implementation outcomes, including reach, adoption, acceptability, feasibility, cost, sustainability, and equitable distribution of benefits. Patient-facing large language model systems should augment rather than replace established care models. Their value will depend not only on technical capabilities, but on whether they safely, sustainably, and equitably extend evidence-based hypertension care in real-world settings.
This paper explores a multimodal approach for translating emotional cues present in speech, designed with Deaf and Hard-of-Hearing (dhh) individuals in mind. Prior work has focused on visual cues applied to captions, successfully conveying whether a speaker’s words have a negative or positive tone (valence), but with mixed results regarding the intensity (arousal) of these emotions. We propose a novel method using haptic feedback to communicate a speaker’s arousal levels through vibrations on a wrist-worn device. In a formative study with 16 dhh participants, we tested six haptic patterns and found that participants preferred single per-word vibrations at 75 Hz to encode arousal. In a follow-up study with 27 dhh participants, this pattern was paired with visual cues, and narrative engagement with audio-visual content was measured. Results indicate that combining haptics with visuals significantly increased engagement compared to a conventional captioning baseline and a visuals-only affective captioning style.
dDeaf and Hard of Hearing (DHH) people often face significant barriers in medical settings, leading to miscommunication and reduced access to care. While American Sign Language (ASL) interpretation is essential for effective communication with DHH signers, it is frequently unavailable in emergency contexts. Emergency Medical Responders (EMRs)-frontline responders trained to deliver basic emergency care-often struggle to obtain accurate medical histories, particularly from DHH people with limited English literacy. To address this, we designed an AI-based ASL learning tool tailored for EMRs, featuring medical vocabulary modules and AI-powered vocabulary testing support. We present a preliminary evaluation of the tool with five EMRs and publicly release a working prototype with this paper. Insights from the study inform new features and vocabulary expansion.
Recruiting participants from disability communities for accessibility research presents unique challenges that require careful consideration of ethical practices, intersectional representation, methodological rigor, and community sustainability. As accessibility research continues to grow and evolve, researchers face tensions between meaningfully including participants with disabilities and addressing emerging concerns around recruited participants not adequately representing the diversity of the community, overburdening certain participants, participant verification, and fair compensation practices. This workshop will bring together members of the ASSETS community to examine current recruiting practices and document insights into ethical, rigorous, and inclusive participant recruitment in disability research. Through facilitated discussions, we will explore three main themes: (1) methods and models, (2) eligibility criteria and participant verification, and (3) ethical and sustainability considerations. The workshop aims to share current practices, identify key challenges, and develop preliminary guidelines to support accessibility researchers in more sustainable participant recruitment.
People with disabilities represent linguistically diverse communities. For example, among Deaf and Hard of Hearing (DHH) people, many of whom use sign language as their primary language, there is significant variation in written language literacy, highlighting that some might benefit from reading comprehension support tools. Prior research has demonstrated the benefits of lexical and syntactic approaches to Automatic Text Simplification for DHH readers and explored design considerations. Building on this work, we present a fully automatic, GPT-based text comprehension tool that provides in-situ reading support. The tool, released with this demo paper, is easily customizable and adaptable to support a range of disability communities and literacy levels. We present usage scenarios to spark conversations around broader applicability, personalization needs, and future studies comparing in-situ reading support to chatbot-style GPT interfaces.
Searching for unfamiliar American Sign Language (ASL) signs is challenging for learners because, unlike spoken languages, they cannot type a text-based query to look up an unfamiliar sign. Advances in isolated sign recognition have enabled the creation of video-based dictionaries, allowing users to submit a video and receive a list of the closest matching signs. Previous HCI research using Wizard-of-Oz prototypes has explored interface designs for ASL dictionaries. Building on these studies, we incorporate their design recommendations and leverage state-of-the-art sign-recognition technology to develop an automated video-based dictionary. We also present findings from an observational study with twelve novice ASL learners who used this dictionary during video-comprehension and question-answering tasks. Our results address human-AI interaction challenges not covered in previous WoZ research, including recording and resubmitting signs, unpredictable outputs, system latency, and privacy concerns. These insights offer guidance for designing and deploying video-based ASL dictionary systems.
Progress in machine understanding of sign languages has been slow and hampered by limited data. In this paper, we present FSboard, an American Sign Language finger-spelling dataset situated in a mobile text entry use case, collected from 147 paid and consenting Deaf signers using Pixel 4A selfie cameras in a variety of environments. Finger-spelling recognition is an incomplete solution that comprises only a small part of sign language translation, but it could provide some immediate benefit to Deaf/Hard of Hearing signers while more broadly capable technology develops. At >3 million characters in length and >250 hours in duration, FSboard is the largest fingerspelling recognition dataset to date by a factor of >10x. As a simple baseline, we finetune 30 Hz MediaPipe Holistic landmark inputs into ByT5-Small and achieve 11.1% Character Error Rate (CER) on a test set with unique phrases and signers. This quality degrades gracefully when decreasing frame rate and excluding face/body landmarks-plausible optimizations to help with on-device performance-but falls short of human performance measured at 2.2% CER.
Reading is a vital skill for social, educational, and professional development, yet various disabilities can impact a person's ability to read and develop literacy skills. HCI and accessibility researchers have explored a wide range of technologies to support reading for people with disabilities. To understand trends in this space, we analyzed 101 publications from Association for Computing Machinery (ACM) venues (2000-2024), coding for target user communities, research methods, technologies, types of support, and contributions. Most research focused on people with dyslexia, followed by people who are Blind or Low Vision, Deaf or Hard of Hearing, or who have intellectual and cognitive disabilities. The majority of studies involved artifact development and short-term lab-based evaluations, with common technologies including visual augmentations, text modifications, and simplification-primarily aimed at improving readability, comprehension, and reading speed. However, participatory approaches and longitudinal evaluations were rarely employed, and the body of work has disproportionately focused on web-based digital reading. Following the initial coding, we conducted community-specific analyses of individual publications to identify patterns and limitations. Based on these analyses, we offer a set of open research questions and community-specific directions to guide future
Closed-captioning is an essential part of viewing audio-visual content for many people, including those who are D/deaf and Hard-of-Hearing. Traditional closed-captioning systems generally consist of a single track of timed text that offers limited options for personalization. Research into extending the capabilities of captioning, such as affective, poetic, and customizable captions has shown a desire among a subset of users for these features, but only in specific contexts. However, due to the difficulty in creating custom stimuli videos utilizing the custom captioning system, comparisons between systems and longitudinal studies have not been pursued. This demo paper introduces Rich Captions, a structured system that allows for a single closed-caption file to be tagged with additional information that can then be flexibly leveraged to render different customizable, creative, and poetic captions from the same file. Additionally, we introduce the Rich Caption Editor 1, a free, open-source software system designed to author, edit, and render rich captions. The system design was informed by a formative design workshop with closed-captioning researchers and advocates. The current design allows researchers to generate reproducible stimuli for closed-captioning studies. Once the design space and user preferences are better understood, the rich captioning framework could be refined to serve a general audience.
As they develop comprehension skills, American Sign Language (ASL) learners often view challenging ASL videos, which may contain unfamiliar signs. Current dictionary tools require students to isolate a single sign they do not understand and input a search query, by selecting linguistic properties or by performing the sign into a webcam. Students may struggle with extracting and re-creating an unfamiliar sign, and they must leave the video-watching task to use an external dictionary tool. We investigate a technology that enables users, in the moment, i.e., while they are viewing a video, to select a span of one or more signs that they do not understand, to view dictionary results. We interviewed 14 American Sign Language (ASL) learners about their challenges in understanding ASL video and workarounds for unfamiliar vocabulary. We then conducted a comparative study and an in-depth analysis with 15 ASL learners to investigate the benefits of using video sub-spans for searching, and their interactions with a Wizard-of-Oz prototype during a video-comprehension task. Our findings revealed benefits of our tool in terms of quality of video translation produced and perceived workload to produce translations. Our in-depth analysis also revealed benefits of an integrated search tool and use of span-selection to constrain video play. These findings inform future designers of such systems, computer vision researchers working on the underlying sign matching technologies, and sign language educators.
Social stigma negatively impacts the well-being of neurodivergent individuals. Specifically for autistic people, the social isolation and pressure to conform to normative ways of being can take a tremendous toll; such as a thwarted sense of belonging to the point of higher rates of suicide. Yet, few technologies are directly targeting the problem of stigma. Socio-technical systems have tremendous potential to shift public perception of traditionally marginalized populations. However, these systems are not consistently designed to reflect the values and needs of neurodivergent individuals. This work explores the use of design sprints to envision CT-a spectrum of technology that could reduce social stigma by increasing the public's awareness, accommodations, acceptance, advocacy, and appreciation. This work reports on two design sprints across 25 HCI community members with varying lived experiences with neurodiversity and knowledge of design practice. The resulting design concepts were discussed in the groups and then analyzed to reflect on how they might combat stigma. Results reveal designs that support the freedom to be oneself via (1) safe spaces (2) public understanding, and (3) authentic expression of strengths, challenges, and needs.
Public sector leverages artificial intelligence (AI) to enhance the efficiency, transparency, and accountability of civic operations and public services. This includes initiatives such as predictive waste management, facial recognition for identification, and advanced tools in the criminal justice system. While public-sector AI can improve efficiency and accountability, it also has the potential to perpetuate biases, infringe on privacy, and marginalize vulnerable groups. Responsible AI (RAI) research aims to address these concerns by focusing on fairness and equity through participatory AI. We invite researchers, community members, and public sector workers to collaborate on designing, developing, and deploying RAI systems that enhance public sector accountability and transparency. Key topics include raising awareness of AI's impact on the public sector, improving access to AI auditing tools, building public engagement capacity, fostering early community involvement to align AI innovations with public needs, and promoting accessible and inclusive participation in AI development. The workshop will feature two keynotes, two short paper sessions, and three discussion-oriented activities. Our goal is to create a platform for exchanging ideas and developing strategies to design community-engaged RAI systems while mitigating the potential harms of AI and maximizing its benefits in the public sector.
PopSign is a smartphone-based bubble-shooter game that helps hearing parents of deaf infants learn sign language. To help parents practice their ability to sign, PopSign is integrating sign language recognition as part of its gameplay. For training the recognizer, we introduce the PopSign ASL v1.0 dataset that collects examples of 250 isolated American Sign Language (ASL) signs using Pixel 4A smartphone selfie cameras in a variety of environments. It is the largest publicly available, isolated sign dataset by number of examples and is the first dataset to focus on one-handed, smartphone signs. We collected over 210,000 examples at 1944x2592 resolution made by 47 consenting Deaf adult signers for whom American Sign Language is their primary language. We manually reviewed 217,866 of these examples, of which 175,022 (approximately 700 per sign) were the sign intended for the educational game. 39,304 examples were recognizable as a sign but were not the desired variant or were a different sign. We provide a training set of 31 signers, a validation set of eight signers, and a test set of eight signers. A baseline LSTM model for the 250-sign vocabulary achieves 82.1% accuracy (81.9% class-weighted F1 score) on the validation set and 84.2% (83.9% class-weighted F1 score) on the test set. Gameplay suggests that accuracy will be sufficient for creating educational games involving sign language recognition.
Caption text conveys salient auditory information to deaf or hard-of-hearing (DHH) viewers. However, the emotional information within the speech is not captured. We developed three emotive captioning schemas that map the output of audio-based emotion detection models to expressive caption text that can convey underlying emotions. The three schemas used typographic changes to the text, color changes, or both. Next, we designed a Unity framework to implement these schemas and used it to generate stimuli videos. In an experimental evaluation with 28 DHH viewers, we compared DHH viewers’ ability to understand emotions and their subjective judgments across the three captioning schemas. We found no significant difference in participants’ ability to understand the emotion based on the captions or their subjective preference ratings. Open-ended feedback revealed factors contributing to individual differences in preferences among the participants and challenges with automatically generated emotive captions that motivate future work.
Advancements in AI will soon enable tools for providing automatic feedback to American Sign Language (ASL) learners on some aspects of their signing, but there is a need to understand their preferences for submitting videos and receiving feedback. Ten participants in our study were asked to record a few sentences in ASL using software we designed, and we provided manually curated feedback on one sentence in a manner that simulates the output of a future automatic feedback system. Participants responded to interview questions and a questionnaire eliciting their impressions of the prototype. Our initial findings provide guidance to future designers of automatic feedback systems for ASL learners.
Television captions blocking visual information causes dissatisfaction among Deaf and Hard of Hearing (DHH) viewers, yet existing caption evaluation metrics do not consider occlusion. To create such a metric, DHH participants in a recent study imagined how bad it would be if captions blocked various on-screen text or visual content. To gather more ecologically valid data for creating an improved metric, we asked 24 DHH participants to give subjective judgments of caption quality after actually watching videos, and a regression analysis revealed which on-screen contents’ occlusion related to users’ judgments. For several video genres, a metric based on our new dataset out-performed the prior state-of-the-art metric for predicting the severity of captions occluding content during videos, which had been based on that prior study. We contribute empirical findings for improving DHH viewers’ experience, guiding the placement of captions to minimize occlusions, and automated evaluation of captioning quality in television broadcasts.
Deaf and hard of hearing individuals regularly rely on captioning while watching live TV. Live TV captioning is evaluated by regulatory agencies using various caption evaluation metrics. However, caption evaluation metrics are often not informed by preferences of DHH users or how meaningful the captions are. There is a need to construct caption evaluation metrics that take the relative importance of words in a transcript into account. We conducted correlation analysis between two types of word embeddings and human-annotated labeled word-importance scores in existing corpus. We found that normalized contextualized word embeddings generated using BERT correlated better with manually annotated importance scores than word2vec-based word embeddings. We make available a pairing of word embeddings and their human-annotated importance scores. We also provide proof-of-concept utility by training word importance models, achieving an F1-score of 0.57 in the 6-class word importance classification task.
Searching for the meaning of an unfamiliar sign-language word in a dictionary is difficult for learners, but emerging sign-recognition technology will soon enable users to search by submitting a video of themselves performing the word they recall. However, sign-recognition technology is imperfect, and users may need to search through a long list of possible results when seeking a desired result. To speed this search, we present a hybrid-search approach, in which users begin with a video-based query and then filter the search results by linguistic properties, e.g., handshape. We interviewed 32 ASL learners about their preferences for the content and appearance of the search-results page and filtering criteria. A between-subjects experiment with 20 ASL learners revealed that our hybrid search system outperformed a video-based search system along multiple satisfaction and performance metrics. Our findings provide guidance for designers of video-based sign-language dictionary search systems, with implications for other search scenarios.
While the availability of captioned television programming has increased, the quality of this captioning is not always acceptable to Deaf and Hard of Hearing (DHH) viewers, especially for live or unscripted content broadcast from local television stations. Although some current caption metrics focus on textual accuracy (comparing caption text with an accurate transcription of what was spoken), other properties may affect DHH viewers’ judgments of caption quality. In fact, U.S. regulatory guidance on caption quality standards includes issues relating to how the placement of captions may occlude other video content. To this end, we conducted an empirical study with 29 DHH participants to investigate the effect on user’s judgements of caption quality or their enjoyment of the video, when captions overlap with an onscreen speaker’s eyes or mouth, or when captions overlap with onscreen text. We observed significantly more negative user-response scores in the case of such overlap. Understanding the relationship between these occlusion features and DHH viewers’ judgments of the quality of captioned video will inform future work towards the creation caption evaluation metrics, to help ensure the accessibility of captioned television or video.