We are releasing a dataset containing videos of both fluent and non-fluent signers using American Sign Language (ASL), which were collected using a Kinect v2 sensor. This dataset was collected as a part of a project to develop and evaluate computer vision algorithms to support new technologies for automatic detection of ASL fluency attributes. A total of 45 fluent and non-fluent participants were asked to perform signing homework assignments that are similar to the assignments used in introductory or intermediate level ASL courses. The data is annotated to identify several aspects of signing including grammatical features and non-manual markers. Sign language recognition is currently very data-driven and this dataset can support the design of recognition technologies, especially technologies that can benefit ASL learners. This dataset might also be interesting to ASL education researchers who want to contrast fluent and non-fluent signing.
Online video is an important information source, yet its pace of growth, including user-submitted content, is so rapid that automatic captioning technologies are needed to make content accessible for people who are Deaf or Hard-of-Hearing (DHH). To support future creation of a research dataset of online videos, we must prioritize which genres of online video content DHH users believe are of greatest importance to be accurately captioned. Our first contribution is to validate that the Best-Worst Scaling (BWS) methodology is able to accurately gather judgments on this topic by conducting an in-person study with 25 DHH users, using a card-sorting methodology to rank the importance for various YouTube genres of online video to be accurately captioned. Our second contribution is to identify video genres of highest captioning importance via an online survey with 151 DHH individuals, and those participants highly ranked: News and Politics, Education, and Technology and Science.
We discuss issues of Artificial Intelligence (AI) fairness for people with disabilities, with examples drawn from our research on HCI for AI-based systems for people who are Deaf or Hard of Hearing (DHH). In particular, we discuss the need for inclusion of data from people with disabilities in training sets, the lack of interpretability of AI systems, ethical responsibilities of access technology researchers and companies, the need for appropriate evaluation metrics for AI-based access technologies (to determine if they are ready to be deployed and if they can be trusted by users), and the ways in which AI systems influence human behavior and influence the set of abilities needed by users to successfully interact with computing systems.
Many Deaf and Hard-of-Hearing (DHH) individuals rely on sign language interpreting to communicate with hearing peers. If on-site interpreting is not available, DHH individuals may use remote interpreting over a smartphone video-call. However, this solution requires the DHH individual to give up either 1) the use of one signing hand by holding the smartphone or 2) their ability to multitask and move around by propping the smartphone up in a fixed location. We explore this problem within the context of the workplace, and present a prototype hands-free device using augmented reality glasses with a hat-mounted fisheye camera and mic/speaker. To explore the validity of our design, we conducted 1) a video interpretability experiment, and 2) a user study with 18 participants (9 DHH, 9 hearing) in a workplace environment. Our results suggest that a hands-free device can support accurate interpretation while enhancing personal interactions.
We have collected a new dataset consisting of color and depth videos of fluent American Sign Language (ASL) signers performing sequences of 100 ASL signs from a Kinect v2 sensor. This directed dataset had originally been collected as part of an ongoing collaborative project, to aid in the development of a sign-recognition system for identifying occurrences of these 100 signs in video. The set of words consist of vocabulary items that would commonly be learned in a first-year ASL course offered at a university, although the specific set of signs selected for inclusion in the dataset had been motivated by project-related factors. Given increasing interest among sign-recognition and other computer-vision researchers in red-green-blue-depth (RBGD) video, we release this dataset for use by the research community. In addition to the RGB video files, we share depth and HD face data as well as additional features of face, hands, and body produced through post-processing of this data.
As the accuracy of Automatic Speech Recognition (ASR) nears human-level quality, it might become feasible as an accessibility tool for people who are Deaf and Hard of Hearing (DHH) to transcribe spoken language to text. We conducted a study using in-person laboratory methodologies, to investigate requirements and preferences for new ASR-based captioning services when used in a small group meeting context. The open-ended comments reveal an interesting dynamic between: caption readability (visibility of text) and occlusion (captions blocking the video contents). Our 105 DHH participants provided valuable feedback on a variety of caption-appearance parameters (strongly preferring familiar styles such as closed captions), and in this paper we start a discussion on how ASR captioning could be visually styled to improve text readability for DHH viewers.
Developing successful sign language recognition, generation, and translation systems requires expertise in a wide range of fields, including computer vision, computer graphics, natural language processing, human-computer interaction, linguistics, and Deaf culture. Despite the need for deep interdisciplinary knowledge, existing research occurs in separate disciplinary silos, and tackles separate portions of the sign language processing pipeline. This leads to three key questions: 1) What does an interdisciplinary view of the current landscape reveal? 2) What are the biggest challenges facing the field? and 3) What are the calls to action for people working in the field? To help answer these questions, we brought together a diverse group of experts for a two-day workshop. This paper presents the results of that interdisciplinary workshop, providing key background that is often overlooked by computer scientists, a review of the state-of-the-art, a set of pressing challenges, and a call to action for the research community.
To promote greater inclusion of people who are Deaf and Hard of Hearing (DHH) in studies conducted by Human-Computer Interaction (HCI) researchers or professionals, we have undertaken a project to formally translate several standardized usability questionnaires from English to ASL. Many deaf adults in the U.S. have lower levels of English reading literacy, but there are currently no standardized usability questionnaires available in American Sign Language (ASL) for these users. A critical concern in conducting such a translation is to ensure that the meaning of the original question items has been preserved during translation, as well as other key psychometric properties of the instrument, including internal reliability, criterion validity, and construct validity. After identifying best-practices for such a translation and evaluation project, a bilingual team of domain experts (including native ASL signers who are members of the Deaf community) translated the System Usability Scale (SUS) and Net Promoter Score (NPS) instruments into ASL and then conducted back-translation evaluations to assess the faithfulness of the translation. The new ASL instruments were employed in usability tests with DHH participants, to assemble a dataset of response scores, in support of the psychometric validation. We are disseminating these translated instruments, as well as collected response values from DHH participants, to encourage greater participation in HCI studies among DHH users.
As Automatic Speech Recognition (ASR) improves in accuracy, it may become useful for transcribing spoken text in real-time for Deaf and Hard-of-Hearing (DHH) individuals. To quantify users' comprehension and opinion of automatic captions, which inevitably contain some errors, we must identify appropriate methodologies for evaluation studies with DHH users, including quantitative measurement instruments suitable to the various literacy levels among the DHH population. A literature review guided our selection of several probes (e.g. multiple-choice comprehension-question accuracy or response time, scalar-questions about user estimation of ASR errors or their impact, users' numerical estimation of accuracy), which we evaluated in a lab study with DHH users, wherein their literacy levels and the actual accuracy of each caption stimulus were factors. For some probes, participants with lower literacy had more positive subjective responses overall, and, for participants with particular literacy score ranges, some probes were insufficiently sensitive to distinguish between caption accuracy levels.
To enable more websites to provide content in the form of sign language, we investigate software to partially automate the synthesis of animations of American Sign Language (ASL), based on a human-authored message specification. We automatically select: where prosodic pauses should be inserted (based on the syntax or other features), the time-duration of these pauses, and the variations of the speed at which individual words are performed (e.g. slower at the end of phrases). Based on an analysis of a corpus of multi-sentence ASL recordings with motion-capture data, we trained machine-learning models, which were evaluated in a cross-validation study. The best model out-performed a prior state-of-the-art ASL timing model. In a study with native ASL signers evaluating animations generated from either our new model or from a simple baseline (uniform speed and no pauses), participants indicated a preference for speed and pausing in ASL animations from our model.
Recent advances in Automatic Speech Recognition (ASR) have made this technology a potential solution for transcribing audio input in real-time for people who are Deaf or Hard of Hearing (DHH). However, ASR is imperfect; users must cope with errors in the output. While some prior research has studied ASR-generated transcriptions to provide captions for DHH people, there has not been a systematic study of how to best present captions that may include errors from ASR software nor how to make use of the ASR system's word-level confidence. We conducted two studies, with 21 and 107 DHH participants, to compare various methods of visually presenting the ASR output with certainty values. Participants answered subjective preference questions and provided feedback on how ASR captioning could be used with confidence display markup. Users preferred captioning styles with which they were already most familiar (that did not display confidence information), and they were concerned about the accuracy of ASR systems. While they expressed interest in systems that display word confidence during captions, they were concerned that text appearance changes may be distracting. The findings of this study should be useful for researchers and companies developing automated captioning systems for DHH users.
To compare methods of displaying speech-recognition confidence of automatic captions, we analyzed eye-tracking and response data from deaf or hard of hearing participants viewing videos.
In usability studies, designers and researchers frequently use subjective questions to evaluate participants' impression of the usability of some product. The System Usability Scale (SUS) is a popular standardized questionnaire consisting of ten English statements about the usability of a product, to which participants indicate their agreement on a five-point scale. Many deaf adults in the U.S. have lower levels of English reading literacy, but there are currently no standardized questionnaires similar to SUS for Deaf and Hard-of-Hearing (DHH) users who are fluent in American Sign Language (ASL). To facilitate the inclusion of such users in studies, we created an ASL translation of SUS following accepted methods of survey translation: using a bilingual team including native ASL signers who are members of the Deaf community, along with back-translation evaluation to determine whether the meaning of the original was preserved. To validate whether key psychometric properties were preserved during translation, we deployed the ASL instrument in a study with 30 DHH participants. By comparing the results to users? responses to another measurement instrument, along with scores from 10 additional DHH participants responding to the original English SUS, we verified the criterion validity and internal reliability of the new "ASL-SUS." We are disseminating the translated instrument to promote the inclusion of DHH users in HCI research studies or in usability testing of consumer products.
As the accuracy and latency of Automatic Speech Recognition (ASR) technology improves over time, it may become a viable method for transcribing audio input in real-time for specific situations. Such technology can provide access to spoken language for people who are Deaf or Hard of Hearing (DHH). However, ASR is imperfect and will remain in that state for a while, thus there is a need for users to cope with errors in the output. My research focuses on how to best present captions that make use of the ASR system's word-level confidence. This summary will describe the proposed solution, current state of study, and the planned contribution to the field of HCI and accessibility for DHH individuals.
32nd Annual International Technology and Persons with Disabilities Conference Scientific/Research Proceedings, San Diego, 2017
Generating sentences from a library of signs implemented through a sparse set of key frames derived from the segmental structure of a phonetic model of ASL has the advantage of flexibility and efficiency, but lacks the lifelike detail of motion capture. These difficulties are compounded when faced with real-time generation and display. This paper describes a technique for automatically adding realism without the expense of manually animating the requisite detail. The new technique layers transparently over and modifies the primary motions dictated by the segmental model and does so with very little computational cost, enabling real-time production and display. The paper also discusses avatar optimizations that can lower the rendering overhead in real-time displays.
A new extension to ELAN offers expanded n-gram analysis tools including improved search capabilities and an extensive library of statistical measures of association for n-grams. This paper presents an overview of the new tools and a case study in American Sign Language synthesis that exploits these capabilities for computing more natural timing in generated sentences. The new extension provides a time-saving convenience for language researchers using ELAN.
American Sign Language, the preferred language of the Deaf community in the USA is its own language; complete with a rich collection of grammatical features. The DePaul University team has been working on an automatic English/ASL translator implemented as a 3D avatar in order to facilitate better communication between hearing and deaf people. Animations suffered from various timing inconsistencies and were awkward in appearance. This project attempts to address them by using corpus analysis to discern subtle features of ASL in order to improve the coarticulation model in the avatar.