Socially assistive robots (SARs) have shown great potential for supplementing well-being support. However, prior studies have found that existing dialogue pipelines for SARs remain limited in real-time latency, back-channeling, and personalized speech dialogue. Toward addressing these limitations, we propose using integrated end-to-end speech-language models (SLMs) with SARs. This work 1) evaluated the usability of an SLM-enabled SAR dialogue system through a small user study, and 2) identified remaining limitations through study user feedback to inform future improvements. We conducted a small within-participant user study with university students (N = 11) whose results showed that participants perceived an SLM-enabled SAR system as capable of providing empathetic feedback, natural turn-taking, back-channeling, and adaptive responses. We also found that participants reported the robot's nonverbal behaviors as lacking variability and synchronization with conversation, and the SLM's verbal feedback as generic and repetitive. These findings highlighted the need for real-time robot movement synchronized with conversation, improved prompting or fine-tuning to generate outputs better aligned with mental health practices, and more expressive, adaptive vocal generation.
Missing significant amounts of school during K-12 education is known to put students’ cognitive and social development at risk. Alternatives such as home instruction and online learning are common, but lack sufficient interaction with peers and teachers in the classroom. Mobile remote presence systems, or telepresence robots, are promising for homebound students because they provide embodiment and mobility in addition to the real-time participation offered by video conferencing technologies. Research is needed, however, to identify what actual benefits and challenges would be experienced by homebound students using telepresence robots in the K-12 classroom context. We present findings from four multi-week deployments with homebound K-12 students attending classes via telepresence robots. The homebound students’ experiences were documented in a total of 15 interviews and analyzed qualitatively as case studies. The homebound student participants and their deployment contexts differed from one another along multiple dimensions, and while some benefits of mobile remote attendance were enjoyed by all participants, each participant also experienced unique benefits. Challenges with hearing, seeing, and moving the robot around the classroom warranted improvements to the design of the telepresence system. Other challenges suggested priorities for managing a classroom deployment, such as ensuring that the remote student is included in classroom activities, accountable to the teacher, and treated with respect by classmates. Based on insights from the study, we make recommendations for real-world deployment procedures and design of future studies in similar contexts.
Both haptic signals and simple, non-anthropomorphic robots can convey complex emotions and enhance remote communication. In this study, we integrated a zoomorphic socially expressive Blossom robot and a haptic sleeve to create a novel multimodal telepresence platform for remote social interaction. Through a within-subject user study with 16 participants, we explored the individual and combined effects of socially expressive robots and mediated social touch on affective communication and social presence during a semi-collaborative LEGO assembly task. Across all participants, the robot and wearable device significantly impacted how participants perceived expressions of gratitude, calming, attention-grabbing, and sadness, evaluated through self-reported valence and arousal. The robot and wearable device in our setting did not show a significant effect on social presence. The observations from this exploratory study can inform the design of multimodal telepresence systems and interactions using non-anthropomorphic robots and mediated touch.
This work presents a novel approach to fabricating soft capacitive tactile sensors using a surface crochet technique to embed conductive thread within crocheted textile substrates. The sensors are mechanically compliant, low-cost, removable, and can be incorporated into a wide range of semi-open-mesh textile substrates, including crocheted, knitted, and loosely woven fabrics. To examine the influence of textile structure and fiber material on sensing performance, we fabricated sensors from acrylic, bamboo, and faux fur yarns, and evaluated their binary touch detection accuracy across four force levels and their signal-to-noise ratio over 30 trials per material. A user study with 15 participants revealed that integrating the sensors significantly affected the perceived tactile qualities of each textile substrate. Finally, we evaluated the sensors in a potential real-world use case: enabling touch-based interactions with a soft, zoomorphic socially assistive robot. Quantitative and qualitative findings highlight trade-offs between sensor performance, perceived tactile qualities, and affective impressions of the robot, informing design considerations for integrating textile-based tactile sensing in soft robotic systems.
Artificial intelligence (AI) has demonstrated significant potential in supporting pediatric speech therapy. However, no systematic review has examined the quality and expanding landscape of commercially available AI-enabled speech therapy mobile apps for children. We conducted a systematic review and analysis of 21 identified commercially available AI -enabled apps designed for speech-sound practice with children. Using 15 evaluation criteria consolidated and extended from prior HCI design guidelines for human-AI and child-AI interaction, our content and risk analysis revealed that existing apps have not yet leveraged the full potential of existing AI tools, and often fail to rigorously adhere to the best recommended practice. Incorporating feedback from the speech therapy community, we categorized our findings into four areas of ethical concern. To address these concerns, we propose actionable recommendations for app designers, families, and speech therapists to guide more ethical development and safe use of AI-enabled applications for pediatric speech therapy.
Hair care robots can help address labor shortages in elderly care while enabling those with limited mobility to maintain their hair-related identity. We present MOE-Hair, a soft robot system that performs three hair-care tasks: head patting, finger combing, and hair grasping. The system features a tendon-driven soft robot end-effector (MOE) with a wrist-mounted RGBD camera, leveraging both mechanical compliance for safety and visual force sensing through deformation. In testing with a force-sensorized mannequin head, MOE achieved comparable hair-grasping effectiveness while applying significantly less force than rigid grippers. Our novel force estimation method combines visual deformation data and tendon tensions from actuators to infer applied forces, reducing sensing errors by up to 60.1% and 20.3% compared to actuator current load-only and depth image-only baselines, respectively. A user study with 12 participants demonstrated statistically significant preferences for MOE-Hair over a baseline system in terms of comfort, effectiveness, and appropriate force application. These results demonstrate the unique advantages of soft robots in contact-rich hair-care tasks, while highlighting the importance of precise force control despite the inherent compliance of the system. Videos, data, and code are available at moehair.github.io.
People have a variety of preferences for how robots behave. To understand and reason about these preferences, robots aim to learn a reward function that describes how aligned robot behaviors are with a user's preferences. Good representations of a robot's behavior can significantly reduce the time and effort required for a user to teach the robot their preferences. Specifying these representations-what ''features'' of the robot's behavior matter to users-remains a difficult problem; Features learned from raw data lack semantic meaning and features learned from user data require users to engage in tedious labeling processes. Our key insight is that users tasked with customizing a robot are intrinsically motivated to produce labels through exploratory search; they explore behaviors that they find interesting and ignore behaviors that are irrelevant. To harness this novel data source of exploratory actions, we propose contrastive learning from exploratory actions (CLEA) to learn trajectory features that are aligned with features that users care about. We learned CLEA features from exploratory actions users performed in an open- ended signal design activity (N=25) with a Kuri robot, and evaluated CLEA features through a second user study with a different set of users (N=42). CLEA features outperformed self- supervised features when eliciting user preferences over four metrics: completeness, simplicity, minimality, and explainability.
Perceptions of gender have a significant impact on human-human interaction, and gender has wide-reaching social implications for robots intended to interact with humans. This work explored two flexible modalities for communicating gender in robots–voice and appearance–and we studied their individual and combined influences on a robot’s perceived gender. We evaluated the perception of a robot’s gender through three online studies. First, we conducted a voice design study (n = 65) on the gender perception of robot voices by varying speaker identity and pitch. Second, we conducted a clothing design study (n = 93) on the gender perception of robot clothing designed for two different tasks. Finally, building on the results of the first two studies, we completed a large integrative video study (n = 273) involving two human-robot interaction tasks. We found that voice and clothing can be used to reliably establish a robot’s perceived gender, and that combining these two modalities can have different effects on the robot’s perceived gender. Taken together, these results inform the design of robot voices and clothing as individual and interacting components in the perceptions of robot gender.
Support groups allow individuals with similar challenges to come together, share their experiences, and receive support. The psychological benefits of participating in a support group heavily rely on peer support and group connectedness. Hence, a high degree of dyadic alliance between group participants is crucial to successful support groups. This paper investigates the verbal and nonverbal behaviors associated with peer-to-peer dyadic alliance in support groups. Multimodal behavioral data were collected from 96 participants of 18 support group sessions, moderated by a virtual embodied conversational agent (ECA). Statistical analysis revealed that select facial expressions and head gestures are significant predictors of dyadic alliance. We develop a multimodal machine learning model to quantify alliance between dyads (i.e., participant pairs in sessions), achieving a weighted F1 score of 0.752 and a balanced accuracy of 0.590 under session-independent cross-validation. This work represents, to the best of our knowledge, the first computational study on dyadic alliance and its behavioral markers. The findings offer potential opportunities for improving support-group facilitator training and automated facilitation systems, with the aim of expanding access to care.
In peer mediation--an approach to conflict resolution used in many K-12 schools in the United States--students help other students to resolve conflicts. For schools without peer mediation programs, socially assistive robots (SARs) may be able to provide an accessible option to practice peer mediation. We investigate how elementary school students react to a peer mediator role-play activity through an exploratory study with SARs. We conducted a small single-session between-subjects study with 12 participants. The study had two conditions, one with two robots acting as disputants, and the other without the robots and just the tablet. We found that a majority of students had positive feedback on the activity, with many students saying the peer mediation practice helped them feel better about themselves. Some said that the activity taught them how to help friends during conflict, indicating that the use of SARs for peer mediation practice is promising. We observed that participants had varying reading levels that impacted their ability to read and dictate the turns in the role-play script, an important consideration for future study design. Additionally, we found that some participants were more expressive while reading the script and throughout the activity. Although we did not find statistical differences in pre-/post-session self-perception and quiz performance between the robot and tablet conditions, we found strong correlations (p
Automatic speech recognition (ASR) models often experience performance degradation due to data domain shifts introduced at test time, a challenge that is further amplified for child speakers. Test-time adaptation (TTA) methods have shown great potential in bridging this domain gap. However, the use of TTA to adapt ASR models to the individual differences in each child's speech has not yet been systematically studied. In this work, we investigate the effectiveness of two widely used TTA methods-SUTA, SGEM-in adapting off-the-shelf ASR models and their fine-tuned versions for child speech recognition, with the goal of enabling continuous, unsupervised adaptation at test time. Our findings show that TTA significantly improves the performance of both off-the-shelf and fine-tuned ASR models, both on average and across individual child speakers, compared to unadapted baselines. However, while TTA helps adapt to individual variability, it may still be limited with non-linguistic child speech.
Large language models (LLMs) have been shown to exhibit a wide range of capabilities, such as writing robot code from language commands – enabling non-experts to direct robot behaviors, modify them based on feedback, or compose them to perform new tasks. However, these capabilities (driven by in-context learning) are limited to short-term interactions, where users' feedback remains relevant for only as long as it fits within the context size of the LLM, and can be forgotten over longer interactions. In this work, we investigate fine-tuning the robot code-writing LLMs, to remember their in-context interactions and improve their teachability i.e., how efficiently they adapt to human inputs (measured by average number of corrections before the user considers the task successful). Our key observation is that when human-robot interactions are formulated as a partially observable Markov decision process (in which human language inputs are observations, and robot code outputs are actions), then training an LLM to complete previous interactions can be viewed as training a transition dynamics model – that can be combined with classic robotics techniques such as model predictive control (MPC) to discover shorter paths to success. This gives rise to Language Model Predictive Control (LMPC), a framework that fine-tunes PaLM 2 to improve its teachability on 78 tasks across 5 robot embodiments – improving non-expert teaching success rates of unseen tasks by 26.9 number of human corrections from 2.4 to 1.9. Experiments show that LMPC also produces strong meta-learners, improving the success rate of in-context learning new tasks on unseen robot embodiments and APIs by 31.5 code, and demos at: https://robot-teaching.github.io/.
Understanding and respecting personal space preferences is essential for socially assistive robots (SARs) designed for older adult users. This work introduces and evaluates a novel personalized context-aware method for modeling users' proxemics preferences during human-robot interactions (HRIs). Using an interactive augmented reality (AR) interface, we collected a set of user-preferred distances from the robot and employed an active transfer learning (ATL) approach to fine-tune a specialized deep learning model. We evaluated this approach through two user studies: 1) a convenience population study (n = 24) to validate the efficacy of the ATL approach, and 2) a user study involving older adults (n = 15) to assess the system's usability. We compared the data collected with the AR interface and with the physical robot to examine the relationship between proxemics preferences for a virtual robot versus a physically embodied robot. We found that fine-tuning significantly improved model performance: on average, the error in testing decreased by 26.97% after fine-tuning. The system was well-received by older adult participants, who provided valuable feedback and suggestions for future work.
Hair-care robots have the potential to alleviate labor shortages in elderly care and enable those with limited mobility to express their identities through hair styling. In this work, we highlight two advantages that soft robotic manipulators have in hair-care applications: safety through mechanical compliance and sensing through observing deformation. To demonstrate these advantages, we introduce a soft robotic end-effector which we call Multi-finger Omnidirectional End-effector (MOE) for hair-care applications. We validate that in hair-grasping tasks, MOE exerts 74.1% less force on the head while being able to grasp a similar amount of hair compared to rigid grippers. We further demonstrate that we can reliably estimate the mesh shape of MOE during interaction with a head and that we can infer useful information about the head such as its occluded shape. The results suggest that soft robots are uniquely advantaged in hair-care tasks.
Assistive robots interact with humans and must adapt to different users' preferences to be effective. An easy and effective technique to learn non-expert users' preferences is through rankings of robot behaviors, for example, robot movement trajectories or gestures. Existing techniques focus on generating trajectories for users to rank that maximize the outcome of the preference learning process. However, the generated trajectories do not appear to reflect the user's preference over repeated interactions. In this work, we design an algorithm to generate trajectories for users to rank that we call Covariance Matrix Adaptation Evolution Strategies with Information Gain (CMA-ES-IG). CMA-ES-IG prioritizes the user's experience of the preference learning process. We show that users find our algorithm more intuitive and easier to use than previous approaches across both physical and social robot tasks. This project's code is hosted at github.com/interaction-lab/CMA-ES-IG
Robots that cooperate with humans must be effective at communicating with them. However, people have varied preferences for communication based on many contextual factors, such as culture, environment, and past experience. To communicate effectively, robots must take those factors into consideration. In this work, we present the Robot Signal Design (RoSiD) tool to empower people to easily self-specify communicative preferences for collaborative robots. We show through a participatory design study that the RoSiD tool enables users to create signals that align with their communicative preferences, and we illuminate how this tool can be further improved.
Mothers of infants have specific demands in fostering emotional bonds with their children, characterized by dynamics that are different from adult-adult interactions, notably requiring heightened maternal emotional regulation. In this study, we analyzed maternal emotional state by modeling maternal emotion regulation reflected in smiles. The dataset comprises N=94 videos of approximately 3 +/- 1-minutes, capturing free play interactions between 6 and 12-month-old infants and their mothers. Corresponding demographic details of self-reported maternal mental health provide variables for determining mothers' relations to emotions measured during free play. In this work, we employ diverse methodological approaches to explore the temporal evolution of maternal smiles. Our findings reveal a correlation between the temporal dynamics of mothers' smiles and their emotional state. Furthermore, we identify specific smile features that correlate with maternal emotional state, thereby enabling informed inferences with existing literature on general smile analysis. This study offers insights into emotional labor, defined as the management of one's own emotions for the benefit of others, and emotion regulation entailed in mother-infant interactions.
Cognitive behavioral therapy (CBT) is a widely used therapeutic method for guiding individuals toward restructuring their thinking patterns as a means of addressing anxiety, depression, and other challenges. We developed a large language model (LLM)-powered prompt-engineered socially assistive robot (SAR) that guides participants through interactive CBT at-home exercises. We evaluated the performance of the SAR through a 15-day study with 38 university students randomly assigned to interact daily with the robot or a chatbot (using the same LLM), or complete traditional CBT worksheets throughout the duration of the study. We measured weekly therapeutic outcomes, changes in pre-/post-session anxiety measures, and adherence to completing CBT exercises. We found that self-reported measures of general psychological distress significantly decreased over the study period in the robot and worksheet conditions but not the chatbot condition. Furthermore, the SAR enabled significant single-session improvements for more sessions than the other two conditions combined. Our findings suggest that SAR-guided LLM-powered CBT may be as effective as traditional worksheet methods in supporting therapeutic progress from the beginning to the end of the study and superior in decreasing user anxiety immediately after completing the CBT exercise.
A common denominator for most therapy treatments for children who suffer from an anxiety disorder is daily practice routines to learn techniques needed to overcome anxiety. However, applying those techniques while experiencing anxiety can be highly challenging. This paper presents the design, implementation, and pilot study of a tactile hand-held pocket robot “AffectaPocket”, designed to work alongside therapy as a focus object to facilitate coping during an anxiety attack. The robot does not require daily practice to be used, has a small form factor, and has been designed for children 7 to 12 years old. The pocket robot works by sensing when it is being held and attempts to shift the child's focus by presenting them with a simple three-note rhythm-matching game. We conducted a pilot study of the pocket robot involving four children aged 7 to 10 years, and then a main study with 18 children aged 6 to 8 years; neither study involved children with anxiety. Both studies aimed to assess the reliability of the robot's sensor configuration, its design, and the effectiveness of the user tutorial. The results indicate that the morphology and sensor setup performed adequately and the tutorial process enabled the children to use the robot with little practice. This work demonstrates that the presented pocket robot could represent a step toward developing low-cost accessible technologies to help children suffering from anxiety disorders.
Adriana Tapus合作论文数Human-Robot Interaction (HRI) conference 200918