Captions rarely convey emotional nuances in speech, leaving Deaf and Hard-of-Hearing (dhh) viewers without access to tonal and affective information. We present a two-part mixed-methods study on how haptic feedback can communicate vocal emotion without adding visual load. In Part 1, we replicated an arousal-driven captioning approach using speech-emotion-recognition to modulate typographic weight and vibration intensity. Participants showed divergent mental models and often mapped “more vibration” to loudness rather than emotional arousal, underscoring the construct’s conceptual fuzziness. In Part 2, we evaluated five acoustic-to-haptic mappings that bypass affective inference and translate pitch, rhythm, and waveform cues into vibration patterns. No single pattern dominated, but participants associated options such as pulse or sawtooth with high-arousal emotions, and pitch-normalized signals with calmer states. We derive design guidelines emphasizing contrastive, acoustically grounded mappings and user control for integrating emotional haptics into short-form, captioned media.
Interest in learning American Sign Language (ASL) is growing across higher education institutions in North America, as reflected in rising enrollments. Yet this growth is constrained by limited program availability and few opportunities to practice outside the classroom. AI-based technologies show promise for supporting ASL learning, but educators – who bring essential pedagogical, linguistic, and cultural expertise – have been largely absent from conversations on the design of these tools, with prior work focusing primarily on learners. To address this, we conducted formative interviews with eleven Deaf and one hearing ASL instructor, followed by two focus groups with six Deaf educators, to examine how AI tools could support ASL education. Findings revealed priorities for technology design and considerations for integration into existing pedagogical practices, with attention to curricular, linguistic, and access factors. We offer insights for designing and researching technologies aimed at (1) providing adaptive, structured feedback on signing performance and (2) supporting immersive conversational practice with virtual signing partners.
This paper explores a multimodal approach for translating emotional cues present in speech, designed with Deaf and Hard-of-Hearing (dhh) individuals in mind. Prior work has focused on visual cues applied to captions, successfully conveying whether a speaker’s words have a negative or positive tone (valence), but with mixed results regarding the intensity (arousal) of these emotions. We propose a novel method using haptic feedback to communicate a speaker’s arousal levels through vibrations on a wrist-worn device. In a formative study with 16 dhh participants, we tested six haptic patterns and found that participants preferred single per-word vibrations at 75 Hz to encode arousal. In a follow-up study with 27 dhh participants, this pattern was paired with visual cues, and narrative engagement with audio-visual content was measured. Results indicate that combining haptics with visuals significantly increased engagement compared to a conventional captioning baseline and a visuals-only affective captioning style.
Affective and prosodic captions convey not only what a speaker says, but also how they say it-louder words may appear thicker, quieter ones thinner; angry in red, calm in blue. These captions can improve access, satisfaction, and engagement for d/Deaf and Hard-of-Hearing (dhh) users. While prior work has explored their design space, it has focused largely on dhh participants in North America, limiting generalizability beyond English and Latin-based scripts. To uncover the role of culture and language, we ran an exploratory study with 49 dhh participants from North America and South Korea using CuCap, a tool that allowed them to personalize which speech features were displayed, and how. While emotion visualization was a universally favored choice, confirming prior findings, prosody preferences varied across cultures, reflecting linguistic and hearing factors. These findings point to the need for flexible captioning systems that account for cultural, linguistic, and individual differences.
Use of automatic speech recognition (ASR) technologies is growing as an accessibility tool for the Deaf, however there has not been much research investigating communication practices among hearing and DHH conversational partners and whether hearing participants can be encouraged to exhibit more beneficial communication behaviors through redesigning ASR systems. To that end, we implemented a Wizard-of-Oz study to investigate three designs (Icons, Pop-Up, and Overlay) for a notification system to influence hearing speakers while using ASR-supported videoconferencing. These designs could notify speakers to slow down or talk louder in real-time to aid DHH participants in understanding them, particularly in impromptu work meetings without an interpreter. We recruited 15 deaf participants and 11 hearing participants in a preliminary study to individually show them various designs. Based on our preliminary results, we selected top designs to show to 7 DHH and 7 hearing participants in a follow-up group experimental study, allowing them to experience notifications in virtual collaborative meetings. Overall, participants rated certain notification designs highly, and analysis of open-ended feedback revealed nuances about conversation flow, responsible communication, and showed optimism about including notification systems in future ASR systems for real-world contexts, including in the workplace.
Affective captions employ visual typographic modulations to convey a speaker’s emotions, improving speech accessibility for Deaf and Hard-of-Hearing (dhh) individuals. However, the most effective visual modulations for expressing emotions remain uncertain. Bridging this gap, we ran three studies with 39 dhh participants, exploring the design space of affective captions, which include parameters like text color, boldness, size, and so on. Study 1 assessed preferences for nine of these styles, each conveying either valence or arousal separately. Study 2 combined Study 1’s top-performing styles and measured preferences for captions depicting both valence and arousal simultaneously. Participants outlined readability, minimal distraction, intuitiveness, and emotional clarity as key factors behind their choices. In Study 3, these factors and an emotion-recognition task were used to compare how Study 2’s winning styles performed versus a non-styled baseline. Based on our findings, we present the two best-performing styles as design recommendations for applications employing affective captions.
Research has observed benefits from providing lexical and syntactic approaches to Automatic Text Simplification (ATS) to Deaf and Hard-of-hearing (DHH) readers. However, little research has explored DHH readers' design preferences and interactions with these approaches. This work first explores the design space of ATS systems with DHH readers, identifying potential design configurations for evaluation. Open-ended discussion of participants' design preferences reveal values informing those preferences, including maintaining reading fluency and effciency, and control over the tool. Using popular design choices from our formative study, we evaluated a prototype that provides various simplification types to explore DHH readers' interactions with the system. We observed potential conflicts between participants' values and design preferences, such as the prototype's impact on participants' reading speed and participants' perceived need to reread simplifications suggested by the tool. However, participants found the tool useful, showing a nuanced preference towards world-level lexical simplifications using pop-ups. Our findings highlight the importance of the tool's design on users' reading experiences, and provide implications for the design and evaluation of ATS prototypes with target readers.
As they develop comprehension skills, American Sign Language (ASL) learners often view challenging ASL videos, which may contain unfamiliar signs. Current dictionary tools require students to isolate a single sign they do not understand and input a search query, by selecting linguistic properties or by performing the sign into a webcam. Students may struggle with extracting and re-creating an unfamiliar sign, and they must leave the video-watching task to use an external dictionary tool. We investigate a technology that enables users, in the moment, i.e., while they are viewing a video, to select a span of one or more signs that they do not understand, to view dictionary results. We interviewed 14 American Sign Language (ASL) learners about their challenges in understanding ASL video and workarounds for unfamiliar vocabulary. We then conducted a comparative study and an in-depth analysis with 15 ASL learners to investigate the benefits of using video sub-spans for searching, and their interactions with a Wizard-of-Oz prototype during a video-comprehension task. Our findings revealed benefits of our tool in terms of quality of video translation produced and perceived workload to produce translations. Our in-depth analysis also revealed benefits of an integrated search tool and use of span-selection to constrain video play. These findings inform future designers of such systems, computer vision researchers working on the underlying sign matching technologies, and sign language educators.
In this paper, we propose a machine learning-based multi-stream framework to recognize American Sign Language (ASL) manual signs and nonmanual gestures (face and head movements) in real time from RGB-D videos. Our approach is based on 3D Convolutional Neural Networks (3D CNNs) by fusing the multi-modal features including hand gestures, facial expressions, and body poses from multiple channels (RGB, Depth, Motion, and Skeleton joints). To learn the overall temporal dynamics in a video, a proxy video is generated by selecting a subset of frames for each video which are then used to train the proposed 3D CNN model. We collected a new ASL dataset, ASL-100-RGBD, which contains 42 RGB-D videos captured by a Microsoft Kinect V2 camera. Each video consists of 100 ASL manual signs, along with RGB channel, Depth maps, Skeleton joints, Face features, and HD face. The dataset is fully annotated for each semantic region (i.e. the time duration of each sign that the human signer performs). Our proposed method achieves 92.88% accuracy for recognizing 100 ASL sign glosses in our newly collected ASL-100-RGBD dataset. The effectiveness of our framework for recognizing hand gestures from RGB-D videos is further demonstrated on a large-scale dataset, ChaLearn IsoGD, achieving the state-of-the-art results.
Captions blocking visual information in live television news leads to dissatisfaction among Deaf and Hard of Hearing (DHH) viewers, who cannot see important information on the screen. Prior work has proposed generic guidelines for caption placement but not specifically for live television news, and important genre of television with dense placement of onscreen information regions, e.g., current news topic, scrolling news, etc. To understand DHH viewers' gaze behavior while watching television news, both spatially and temporally, we conducted an eye-tracking study with 19 DHH participants. Participants' gaze behavior varied over time as measured by their proportional fixation time on information regions on the screen. An analysis of gaze behavior coupled with open-ended feedback revealed four thematic categories of information regions. Our work motivates considering the time dimension when considering caption placement, to avoid blocking information regions, as their importance varies over time.
Speech is expressive in ways that caption text does not capture, with emotion or emphasis information not conveyed. We interviewed eight Deaf and Hard-of-Hearing (dhh) individuals to understand if and how captions' inexpressiveness impacts them in online meetings with hearing peers. Automatically captioned speech, we found, lacks affective depth, lending it a hard-to-parse ambiguity and general dullness. Interviewees regularly feel excluded, which some understand is an inherent quality of these types of meetings rather than a consequence of current caption text design. Next, we developed three novel captioning models that depicted, beyond words, features from prosody, emotions, and a mix of both. In an empirical study, 16 dhh participants compared these models with conventional captions. The emotion-based model outperformed traditional captions in depicting emotions and emphasis, with only a moderate loss in legibility, suggesting its potential as a more inclusive design for captions.
Despite the recent improvements in automatic speech recognition (ASR) systems, their accuracy is imperfect in live conversational settings. Classifying the importance of each word in a caption transcription can enable evaluation metrics that best reflect Deaf and Hard of Hearing (DHH) readers' judgment of the caption quality. Prior work has proposed using word embeddings, e.g., word2vec or BERT embeddings, to model word importance in conversational transcripts. Recent work also disseminated a human-annotated word importance dataset. We conducted a word-token level analysis on this dataset and explored Part-of-Speech (POS) distribution. We then augmented the dataset with POS tags and reduced the class imbalance by generating 5% additional text using masking. Finally, we investigated how various supervised models learn the importance of words. The best performing model trained on our augmented dataset performed better than prior models. Our findings can inform the design of a metric for measuring live caption quality from DHH users' perspectives.
Live TV news and interviews often include multiple individuals speaking, with rapid turn-taking, which makes it difficult for viewers who are Deaf and Hard of Hearing (DHH) to follow who is speaking when reading captions. Prior research has proposed several methods of indicating who is speaking. While recent studies have observed various preferences among DHH viewers for speaker-identification methods for videos with different numbers of speakers onscreen, there has not yet been a study that has systematically explored whether there is a formal relationship between the number of people onscreen and the preferences among DHH viewers for how to indicate the speaker in captions. We conducted an empirical study followed by a semi-structured interview with 17 DHH participants to record their preferences among various speaker-identifier types for videos that vary in the number of speakers onscreen. We observed an interaction effect between DHH viewers’ preference for speaker identification and the number of speakers in a video. An analysis of open-ended feedback from participants revealed several factors that influenced their preferences. Our findings guide broadcasters and captioners in selecting speaker-identification methods for captioned videos.
We are releasing a dataset containing videos of both fluent and non-fluent signers using American Sign Language (ASL), which were collected using a Kinect v2 sensor. This dataset was collected as a part of a project to develop and evaluate computer vision algorithms to support new technologies for automatic detection of ASL fluency attributes. A total of 45 fluent and non-fluent participants were asked to perform signing homework assignments that are similar to the assignments used in introductory or intermediate level ASL courses. The data is annotated to identify several aspects of signing including grammatical features and non-manual markers. Sign language recognition is currently very data-driven and this dataset can support the design of recognition technologies, especially technologies that can benefit ASL learners. This dataset might also be interesting to ASL education researchers who want to contrast fluent and non-fluent signing.
Advancements in AI will soon enable tools for providing automatic feedback to American Sign Language (ASL) learners on some aspects of their signing, but there is a need to understand their preferences for submitting videos and receiving feedback. Ten participants in our study were asked to record a few sentences in ASL using software we designed, and we provided manually curated feedback on one sentence in a manner that simulates the output of a future automatic feedback system. Participants responded to interview questions and a questionnaire eliciting their impressions of the prototype. Our initial findings provide guidance to future designers of automatic feedback systems for ASL learners.
Television captions blocking visual information causes dissatisfaction among Deaf and Hard of Hearing (DHH) viewers, yet existing caption evaluation metrics do not consider occlusion. To create such a metric, DHH participants in a recent study imagined how bad it would be if captions blocked various on-screen text or visual content. To gather more ecologically valid data for creating an improved metric, we asked 24 DHH participants to give subjective judgments of caption quality after actually watching videos, and a regression analysis revealed which on-screen contents’ occlusion related to users’ judgments. For several video genres, a metric based on our new dataset out-performed the prior state-of-the-art metric for predicting the severity of captions occluding content during videos, which had been based on that prior study. We contribute empirical findings for improving DHH viewers’ experience, guiding the placement of captions to minimize occlusions, and automated evaluation of captioning quality in television broadcasts.
Research has explored the use of automatic text simplification (ATS), which consists of techniques to make text simpler to read, to provide reading assistance to Deaf and Hard-of-hearing (DHH) adults with various literacy levels. Prior work in this area has identified interest in and benefits from ATS-based reading assistance tools. However, no prior work on ATS has gathered judgements from DHH adults as to what constitutes complex text. Thus, following approaches in prior NLP work, this paper contributes new word-complexity judgements from 11 DHH adults on a dataset of 15,000 English words that had been previously annotated by L2 speakers, which we also augmented to include automatic annotations of linguistic characteristics of the words. Additionally, we conduct a supplementary analysis of the interaction effect between the linguistic characteristics of the words and the groups of annotators. This analysis highlights the importance of collecting judgements from DHH adults for training ATS systems, as it revealed statistically significant interaction effects for nearly all of the linguistic characteristics of the words.
Deaf signers who wish to communicate in their native language frequently share videos on the Web. However, videos cannot preserve privacy—as is often desirable for discussion of sensitive topics—since both hands and face convey critical linguistic information and therefore cannot be obscured without degrading communication. Deaf signers have expressed interest in video anonymization that would preserve linguistic content. However, attempts to develop such technology have thus far shown limited success. We are developing a new method for such anonymization, with input from ASL signers. We modify a motion-based image animation model to generate high-resolution videos with the signer identity changed, but with the preservation of linguistically significant motions and facial expressions. An asymmetric encoder-decoder structured image generator is used to generate the high-resolution target frame from the low-resolution source frame based on the optical flow and confidence map. We explicitly guide the model to attain a clear generation of hands and faces by using bounding boxes to improve the loss computation. FID and KID scores are used for the evaluation of the realism of the generated frames. This technology shows great potential for practical applications to benefit deaf signers.
Automatic Text Simplification (ATS) software aims at automatically rewrite complex text to make it simpler to read. Prior research has explored the use of ATS as a reading assistance technology, identifying benefits from providing these technologies to different groups of users, including Deaf and Hard-of-hearing (DHH) adults. However, little work has investigated the interests and requirements of specific groups of potential users of this technology. Considering prior work establishing that computing professionals often need to read about new technologies in order to stay current in their profession, in this study, we investigated the reading experiences and interests of DHH individuals with work experience in the computing industry in ATS-based reading assistance tools, as well as their perspective on the social accessibility of those tools. Through a survey and two sets of interviews, we found that these users read relatively often, especially in support of their work, and were interested in tools to assist them with complicated texts; but misperceptions arising from public use of these tools may conflict with participants’ desired image in a professional context. This empirical contribution motivates further research into ATS-based reading assistance tools for these users, prioritizing which reading activities users are most interested in seeing the application of this technology, and highlighting design considerations for creating ATS tools for DHH adults, including considerations for social accessibility.
PURPOSE:Phoneme categorization (PC) for voice onset time and second formant transition was studied in adult cochlear implant (CI) users with early-onset deafness and hearing controls.METHOD:Identification and discrimination tasks were administered to 30 participants implanted before 4 years of age, 21 participants implanted after 7 years of age, and 21 hearing individuals.RESULTS:Distinctive identification and discrimination functions confirmed PC within all groups. Compared to hearing participants, the CI groups generally displayed longer/higher category boundaries, shallower identification function slopes, reduced identification consistency, and reduced discrimination performance. A principal component analysis revealed that identification consistency, discrimination accuracy, and identification function slope, but not boundary location, loaded on a single factor, reflecting general PC performance. Earlier implantation was associated with better PC performance within the early CI group, but not the late CI group. Within the early CI group, earlier implantation age but not PC performance was associated with better speech recognition. Conversely, within the late CI group, better PC performance but not earlier implantation age was associated with better speech recognition.CONCLUSIONS:Results suggest that implantation timing within the sensitive period before 4 years of age partly determines the level of PC performance. They also suggest that early implantation may promote development of higher level processes that can compensate for relatively poor PC performance, as can occur in challenging listening conditions.