PopSignAI is an educational smartphone app that uses sign language recognition to help players learn American Sign Language. It focuses on vocabulary from the MacArthur-Bates Communicative Development Inventories (MB-CDI), which are the first concepts one teaches a child in any language. For training the recognizer, we extend the PopSign ASL v1.0 dataset from 250 isolated signs to the full 562 concepts from the MB-CDI. The dataset focuses on one-handed signing as is often performed when communicating using a smartphone. We collected over 372,000 examples at 1944x2592 resolution made by 76 consenting Deaf adult signers for whom American Sign Language is their primary language. We manually reviewed 372,147 of these examples, of which 314,190 (approximately 559 per sign) were the intended signs for the educational game. 49,512 examples were recognizable as signs but were not the desired variant or were a different sign. We provide a training set of 51 signers, a validation set of 11 signers, and a test set of 14 signers. A baseline LSTM model for the 562-sign vocabulary achieves $87.4 \%$ accuracy and a 97.0% top-5 accuracy after resolving homonyms (82.2% accuracy and 82.2% class-weighted F1 score before).
Presbyopia, a common age-related vision condition affecting most people as they age, often remains inadequately understood by those unaffected. To help bridge the gap between abstract accessibility knowledge and a more grounded appreciation of perceptual challenges, this study presents OpticalAging, an optical see-through simulation approach. Unlike VR-based methods, OpticalAging uses dynamically controlled tunable lenses to simulate the first-person visual perspective of presbyopia's distance-dependent blur during real-world interaction, aiming to enhance awareness. While acknowledging critiques regarding simulation's limitations in fully capturing lived experience, we position this tool as a complement to user-centered methods. Our user study (N = 19, 18-35 years old) provides validation: quantitative measurements show statistically significant changes in near points across three age modes (40s, 50s, 60s), while qualitative results suggest increases in reported understanding and empathy among participants. The integration of our tool into a design task showcases its potential applicability within age-inclusive design workflows when used critically alongside direct user engagement.
To help novice signers learn American Sign Language, we develop PopSignAI, a proof-of-concept smartphone-based bubble-shooter game that facilitates real-time interaction through isolated sign language recognition. In a 20-person user study, we demonstrate that encouraging novice signers to practice generating sign in PopSignAI is more efficient for teaching ASL skills than a version of PopSign focused on receptive signing ability. We use over 200,000 examples of 250 signs from 47 signers to train and test a user-independent LSTM recognizer that achieves 82.9% accuracy on an independent test set. For the purposes of the game, the recognizer averages 99.6% accuracy with a 7ms inference time using a 2.5MB model. Ablation studies suggest that as few as eight signers are need for training in order for adequate recognition accuracy for PopSignAI’s gameplay. To encourage future sign language recognition games, we release the PopSignAI recognition pipeline and software. We identify hearing parents of deaf children as important potential users of sign games and conduct interviews with eight of these parents, investigating their motivation and challenges in learning sign.
The NASA-Task Load Index (NASA-TLX) is a widely used multidimensional scale for measuring subjective workload in Human-Computer Interaction (HCI) research. Despite widespread adoption, diverse reporting methods hinder cross-study comparisons and standardization. We systematically reviewed 522 CHI papers (2006–2024) to identify usage variations and common implementation errors, and collected subscale values from 683 tasks across 185 papers for meta-analysis. Correlation, regression, and factor analyses revealed stronger inter-subscale correlations than the original development study, different beta weights, and weakened discriminability. Factor analysis identified a two-factor structure, with Performance showing notably low communality as a distinct dimension. We propose a Short TLX (STLX) that retains Mental Demand , Physical Demand , and Performance as a condensed instrument, given that many studies already omit subscales. We provide standardized implementation guidelines and contribute an interactive research database and a data collection toolkit to facilitate consistent use of NASA-TLX across the HCI community.
The deployment of exoskeletons offers a means to avert musculoskeletal disorders in the construction industry; however, adequate training, especially in equipping students to recognize ergonomic hazards and utilize these devices for construction activities, is still insufficient. Prior research has suggested virtual reality environments as a viable solution to enhance student learning, owing to their immersive and interactive features. Hence, this study investigated a virtual reality learning environment (ViRLE) for facilitating the identification of ergonomic risks and exoskeleton education in construction. The study assessed the effectiveness of the learning environment by examining the cognitive demands imposed on students during interactions with ViRLE. Twelve (12) participants were recruited to engage with ViRLE, which comprised three scenes: (1) construction site walkthrough, (2) ergonomic risks identification, and (3) exoskeleton training. Data was collected using the NASA Task Load Index alongside an eye tracker. Both descriptive and inferential statistics were employed to analyze the data. Subjectively, the findings indicate that the cognitive demand imposed on students across learning Scenes was generally moderate, with overall workload increasing slightly between Scenes. Objectively, mean fixation duration and fixation rate provided converging evidence. Specifically, the mean fixation duration, which reflects the depth of processing at each point of gaze, remained consistent across all Scenes. This suggests that students were able to maintain a consistent level of cognitive processing across different scenes. The study enhances construction education by demonstrating the moderate cognitive demand ViRLE imposes on construction students, which may guide educators in the development of learning environments.
A significant health and safety concern in the construction industry is musculoskeletal disorders (MSDs), typically arising from the physically strenuous nature of tasks and inadequate ergonomic practices. Exoskeletons have been proposed as a potential intervention to mitigate MSDs. To effectively utilize exoskeletons and adopt proper postures during construction tasks, there is a pressing need to educate the future workforce on the operation of exoskeletons and the awareness of ergonomic hazards. Nevertheless, there is a scarcity of research addressing this necessity. To address this need, this study examines the effectiveness of a virtual learning environment (ViRLE) designed to enhance students’ comprehension and practical understanding of ergonomic risks and the implementation of exoskeletons. This study employed a pre-post survey design to evaluate improvement in experiences, knowledge, and skills among 12 construction engineering and management students. Participants initially completed a pre-survey that assessed their confidence in evaluating ergonomic risks, as well as their familiarity and knowledge of the various exoskeleton support types. Subsequently, they engaged with the ViRLE designed to train students on exoskeleton implementation and ergonomic risk identification. A post-survey was conducted to assess changes in their confidence, familiarity, and knowledge levels. The results demonstrated a significant enhancement in students’ confidence regarding ergonomic risks, as well as their knowledge and familiarity with exoskeletons. The findings indicate that ViRLE can serve as an effective educational tool for construction education, especially in safety training. By investigating educational platforms for workforce development, this study contributes to the advancement of robotics in the construction industry.
Directional cues are crucial for environmental interaction. Conventional methods rely on symbolic visual or auditory reminders that require semantic interpretation, a process that proves challenging in demanding dual-tasking scenarios. We introduce a novel alternative for conveying directional cues on wearable displays: directly triggering motion perception using monocularly presented peripheral stimuli. This approach is designed for low visual interference, with the goal of reducing the need for gaze-switching and the complex cognitive processing associated with symbols. User studies demonstrate our method's potential to robustly convey directional cues. Compared to a conventional arrow-based technique in a demanding dual-task scenario, our motion-based approach resulted in significantly more accurate interpretation of these directional cues (p=.008) and showed a trend towards reduced errors on the concurrent primary task (p=.066).
Head-worn displays for everyday wear in the form of regular eyeglasses are technically feasible with recent advances in waveguide technology. One major design decision is determining where in the user's visual field to position the display. Centering the display in the principal point of gaze (PPOG) allows the user to switch attentional focus between the virtual and real images quickly, and best performance often occurs when the display is centered in PPOG or is centered vertically below PPOG. However, these positions are often undesirable in that they are considered interruptive or are associated with negative social perceptions by users. Offsetting the virtual image may be preferred when tasks involve driving, walking, or social interaction. This paper consolidates findings from recent studies on monocular optical see-through HWDs (OST-HWDs), focusing on potential for interruption, comfort, performance, and social perception. For text-based tasks, which serve as a proxy for many monocular OST-HWD tasks, we recommend a 15 horizontal field of view (FOV) with the virtual image in the right lens vertically centered but offset to +8.7 to +23.7 toward the ear. Glanceable content can be offset up to +30 for short interactions.
Approximately 95% of deaf infants are born to hearing parents who are unfamiliar with American Sign Language (ASL). Addressing this need, we are developing ASL smartphone games that utilize sign language recognition (SLR) models to help teach sign production. However, many smartphone users are not accustomed to such computer vision-based game interactions, making their on-boarding process challenging. Through playtests and user interviews, we evaluate the intuitiveness and efcacy of various interaction mechanics for SLR models in games. For beginners, isolated mechanics-where sign recognition is manually triggered-prove efective but require structured pacing constraints to prevent misclassifcation due to delayed or transitional inputs. As players become more experienced, they prefer integrated mechanics, where recognition is tied to gameplay actions, streamlining interaction.
Learning to see the world like an expert is a critical step in mastering complex skills, yet transferring these implicit visual strategies remains a significant challenge. We present Pro's Eyes, a wearable see-through display system designed to bridge this expert-novice gap by synchronizing a novice's observational patterns with an expert's. The system implicitly guides a user's attention by creating a clear viewing aperture aligned with an expert's gaze path while dimming the visual periphery. To validate our approach, we conducted a formal user study (N=17) focusing on an art appreciation task. The results provide strong empirical evidence of the system's efficacy: Pro's Eyes significantly improved gaze synchronization between novices and a pre-recorded expert's path (p <.05) compared to free observation. Subjectively, participants reported that the guidance helped them identify details they would have otherwise missed. Our work's primary contribution is an empirically-validated wearable system that demonstrates the potential of implicit gaze guidance for transferring expert observational skills.
Progress in machine understanding of sign languages has been slow and hampered by limited data. In this paper, we present FSboard, an American Sign Language finger-spelling dataset situated in a mobile text entry use case, collected from 147 paid and consenting Deaf signers using Pixel 4A selfie cameras in a variety of environments. Finger-spelling recognition is an incomplete solution that comprises only a small part of sign language translation, but it could provide some immediate benefit to Deaf/Hard of Hearing signers while more broadly capable technology develops. At >3 million characters in length and >250 hours in duration, FSboard is the largest fingerspelling recognition dataset to date by a factor of >10x. As a simple baseline, we finetune 30 Hz MediaPipe Holistic landmark inputs into ByT5-Small and achieve 11.1% Character Error Rate (CER) on a test set with unique phrases and signers. This quality degrades gracefully when decreasing frame rate and excluding face/body landmarks-plausible optimizations to help with on-device performance-but falls short of human performance measured at 2.2% CER.
AI-Augmented Reasoning systems are cognitive assistants that support human reasoning by providing AI-based feedback that can help users improve their critical reasoning skills. Made possible with new techniques like argumentation mining, fact-checking, crowdsourcing, attention nudging, and large language models, AI augmented reasoning systems can provide real-time feedback on logical reasoning, help users identify and avoid flawed arguments and misinformation, suggest counter-arguments, provide evidence-based explanations, and foster deeper reflection. The goal of this workshop is to bring together researchers from AI, HCI, cognitive and social science to discuss recent advances in AI-augmented reasoning, to identify open problems in this area, and to cultivate an emerging community on this important topic.
Electronic systems for cognitive enrichment in zoos open avenues for monitoring to better understand how and when the animals are engaged. We present a pilot study of an automated elephant trunk detection system utilizing computer vision techniques such as background subtraction and region-based motion detection on a Google Pixel smartphone for real-time monitoring of enrichment use by four African elephants (Loxodonta africana) at Zoo Atlanta. Preliminary results show promise in using camera-based methods for trunk detection; however, more testing and refining of the algorithm is required to obtain an accurate and reliable method that could replace physical sensing methods.
Therapy that incorporates equine movement is a valuable tool for strength and rehabilitation training. While several commercial products exist to monitor the saddle pressure and fit on a horse, therapists utilizing hippotherapy currently do not have an objective means for assessing the progress of their clients' control of weight distribution while engaging in therapy. In this work, we present a device for monitoring the leaning patterns and weight distribution of clients participating in therapy using equine movement. Our prototype uses commercial off-the-shelf parts including barometers that are enclosed in custom-made inflated bladders placed inside a bareback pad to detect changes in air pressure as the bladders are deformed by the weight of the patient. Currently, we are able to show that a client's lean could be detected with an accuracy range of 91.4-98.7%, precision of 48.2%, and an average recall of 42.9% using a simple thresholding algorithm based on rolling sums of pressure changes. Ongoing work seeks to explore methods of designing bladders and processing data that may improve these metrics.