
Large Language Models (LLMs) are considered state of the art for many tasks in robotics and AI. At the same time, there is increasing evidence of their critical limitations such as generating arbitrary responses in new situations, inability to support rapid incremental updates based on limited examples, and opacity. Toward addressing these limitations, our architecture leverages the complementary strengths of LLMs and knowledge-based reasoning. Specifically, the architecture enables an AI agent assisting a human to use an LLM to provide generic abstract predictions of upcoming tasks. The agent also reasons with domain-specific knowledge, recent history of interactions with the human, and semantic databases to: (a) provide contextual prompts to the LLM; and (b) compute a plan of concrete actions that jointly implements the current task and prepares for the anticipated task, replanning as needed. Furthermore, the agent solicits and uses high-level human feedback based on need and availability to incrementally revise the domain-specific knowledge and interactions with the LLM. We ground and evaluate our architecture’s abilities in the realistic VirtualHome simulation environment, demonstrating a substantial performance improvement compared with just using an LLM or an LLM and logical reasoner. Project website: https://brianej.github.io/igfmrdskaa.github.io/
Robot navigation plays a critical role in how people perceive, accept, and collaborate with robots in shared environments. This study presents PRoMo (Preference for Robot Motion Questionnaire), a user-centered tool designed to capture human expectations of robot navigation behavior, independent of specific robot forms or tasks. The questionnaire consolidates 28 empirically grounded behaviors into five thematic categories: safety, predictability, proximity, speed and path selection, and responsiveness. Responses from 142 participants reveal strong preferences for navigation strategies that respect personal space, avoid blind spots, and signal awareness through subtle motion cues. Open-ended responses highlight additional concerns, including robot noise and emotional comfort, suggesting that movement is experienced not only as spatial but also as sensory and expressive. Importantly, subjective familiarity with robots showed stronger correlations with behavior preferences than objective experience. These findings provide a generalizable framework for designing socially appropriate robot navigation strategies in human-centered environments. The questionnaire also serves as a practical evaluation tool to guide the development and testing of real-world robot navigation systems.
The integration of robots into creative domains presents new opportunities for artistic expression. This work introduces a robotic painting system designed to facilitate intuitive, non-programmatic interaction through two distinct modes of engagement. In the first mode, a human user and a robot co-create an abstract painting by taking turns making marks, with the robot responding based on image analysis and predefined artistic rules. The second mode allows the robot to autonomously generate an outline portrait of the user based on image segmentation and facial feature extraction. The system employs ROS2 for task orchestration, animatronic eyes to enhance user engagement, and ChatGPT for conversational feedback.
An ageing population and the need of providing adequate care have led to developing robots to relieve healthcare workers and to assist individuals in their own homes. However, the successful integration of robots in such settings relies on more than just ensuring physical safety associated with physical risks (e.g., collisions): it also requires the user's perceived safety - the users perceiving the robot as not doing any harm. This paper explores the potential influence of a robot gripper's visual and tactile properties, such as materials and texture, on the users' perceived safety and comfort of human-robot interaction. An initial survey was distributed to 53 participants, exploring five (n=5) robot gripper designs focusing on the robots' gripper shape. One design shape was thereafter selected to be constructed as a cover to be placed over the parallel grippers of the TIAGo robot, by using 1) wood filament and 2) plastic. The covers were then tested in an experimental setting with 11 participants. The covers were attached to the TIAGo mobile manipulator robot and participants interacted with both of the designed gripper covers within a controlled laboratory environment. A questionnaire was distributed to all 11 experiment participants, at different stages of the interactions. The findings indicate that the material of the gripper influenced participants' sense of comfort, familiarity, and perceived capabilities of the robot. The study suggests that perceived safety in human-robot interaction (HRI) is shaped not only by physical factors but also by how materials are personally and contextually interpreted. To better support safe and comfortable interactions, further research is needed to understand how material choices shape users' perceived safety.
Most countries are ageing rapidly, creating significant challenges in providing adequate care to elderly population. Older adults’ care centers face several difficulties in ensuring support able to address their diverse needs, due to several factors including increasing caregiver shortages, high variability in cognitive and health conditions of elderly, and challenges in delivering personalized interventions to them. Additionally, maintaining elderly engagement during cognitive training is often problematic due to the repetitive and impersonal nature of involved tasks. To address these limitations, we carried out a study investigating the use of a humanoid robot to deliver interactive, personalised serious games based on older adults' personal memories, to enhance relevance and engagement for them. The approach has been evaluated in a trial conducted in a center for older adults, involving users with varying cognitive abilities. Results indicated that such personalised games were well received by them, with a positive impact on their experience.
A robot’s physical design impacts user acceptance, engagement, and trust while also influencing social and functional expectations about the robot’s capabilities. Robot design, and industrial design in general, can also be a driver of differentiation in a competitive market. In this paper, we leverage generative AI to incrementally explore design spaces of robot appearances. With the goal of overcoming training data bias of text-to-image models that favor stereotypical morphologies and designs, we propose a set of generation methods that enable a designer to guide the exploration process through text-, style-, and structure-based specifications from user or client feedback. Further, using Low-Rank Adaptation for model fine-tuning, the method allows to define the aesthetic direction by an image collection that conveys a particular style or theme ("mood boards"). The experiments demonstrate that our extensions retain image quality in terms of statistical and structural features and allow for both diversity and specificity in the design process. In a case study, we apply this method in the user-centered design process and discuss its opportunities and limitations.
In an interdisciplinary and evolving research field like human-robot interaction, clear and precise results reporting is essential for study comparability and replicability. To address the lack of a standard for such reporting and, at the same time, provide guidance for novices in the field, we have developed a web-based reporting form to capture human-robot interaction studies, serving as a model for how conferences could adopt it into the submission pipeline. In this work, we present a formative evaluation of this form regarding its level of detail, format and clarity, and the perceived benefits for authors, reviewers, and the community as a whole. We report the expert review of nine researchers who highlight the substantial value of this tool. In addition, these experts also provide suggestions for improvements to its form and the addition of details surrounding qualitative reporting.
Isolated sign language recognition is a challenging task involving the learning of complex relationships between spatial and temporal features. Due to the high complexity and relatively small datasets available, state-of-the-art methods often adopt language modeling and convolutional neural network based multimodal designs, achieving high accuracy at the cost of significant architectural complexity. Conceptually simpler, transformers have gained widespread adoption in related computer vision tasks, outperforming 3D convolutional network competitors. However, due to a lack of training data, video transformers struggle with sign language recognition and have not demonstrated competitive accuracy compared to 3D convolutional neural network designs. We introduce DeepSign, a family of vision transformer based sign language recognition models with superior performance to 3D convolutional neural network designs. Through careful model ablation we select the UniFormerV2 and VideoMAE V2 architectures and perform mixture of dataset pretraining. Our strongest model DeepSign UniFormerV2-L achieves state-of-the-art on the WLASL100 and MSASL100 benchmarks, producing 92.64% and 94% top-1 accuracies respectively. Armed with VideoMAE V2’s powerful pretrained backbone, DeepSign ViT base offers greater efficiency for a small accuracy tradeoff. We hope DeepSign will help advance future sign language research by providing strong foundational models to kickstart experiments.
Pomegranate harvesting remains a challenging task due to the fruit's tough stem, dense canopy, and sensitivity to mechanical damage. Traditional harvesting robots rely on vision-based stem localization, which increases computational complexity and reduces robustness in unstructured orchard environments. This paper presents a dual-shear ring end-effector designed to eliminate the need for precise stem detection, utilizing a self-locking shear mechanism that allows the stem to naturally align between the cutting blades. The system integrates a vision-assisted robotic manipulator for fruit detection and a torque regulation mechanism for optimized cutting force application. Experimental validation demonstrates a success rate of over 90% for stems up to 8 mm in diameter and robust performance even under partial and full occlusion conditions. The results confirm that the proposed system achieves efficient, adaptable, and damage-free harvesting, providing a viable solution for autonomous pomegranate harvesting.
As AI-driven robots become more integrated into daily life, understanding user perceptions is crucial for improving their design and interaction. This study investigates the impact of interruptibility and response unpredictability on user engagement with Keirzo, an AI-powered musical robot. Using the Robotic Social Attributes Scale (RoSAS), participants engaged with Keirzo under two conditions: one allowing interruptions and one requiring them to wait for complete responses. Findings suggest that while the ability to interrupt offered a greater sense of control, it did not significantly increase engagement. Participants generally rated Keirzo as more competent when its responses were structured and coherent, whereas repetitive or unpredictable replies reduced perceived intelligence. Perceptions of personality were mixed; some found the robot engaging and expressive, while others viewed it as mechanical or detached. These results highlight the importance of balancing control, coherence, and expressiveness in AI-driven musical interactions. As the findings are exploratory, future work should involve more adaptive systems and larger sample sizes to further examine these dynamics in creative HRI contexts.
Human-interactive robot learning allows a robot to learn tasks more effectively with the help of humans in the role of teacher. While there is a large body of work on algorithms that leverage human input for better robot learning, there has been little attention to understanding how humans teach robots. In this paper, we provide preliminary results on how users strategize the use of demonstrations and evaluative feedback under a budget, and how these choices are influenced by demographic variables such as gender. We implemented a learning algorithm that allows a simulated robot arm to learn three reaching tasks with the help of a human. We collected interaction data for a total of 58 participants, which shows that participants demonstrate a tendency to provide evaluative feedback earlier in their interactions compared to demonstrations, and that gender may have an influence on teaching strategy. This preliminary analysis lays the foundation for future research aimed at developing tuneable computational models of different human teachers.
The use of social robots in education is increasingly being explored as a way to enhance learner engagement and improve learning outcomes. However, most research to date has focused on one-to-one tutoring in high-resource settings, leaving open questions about how social robots perform in group learning contexts—especially in low-resource environments. This study is one of the first to investigate human-robot interaction (HRI) in a low-resource African context, specifically in Tanzanian primary schools. We examined how a social robot tutor can support group-based mathematics learning, comparing the effects of adaptive versus non-adaptive tutoring strategies. Through an experimental, mixed-methods research design, we evaluated pupils’ learning outcomes, engagement, and classroom interactions. Our findings show that social robot tutoring has a significant positive impact on learning outcomes, with adaptive tutoring leading to slightly higher knowledge gains than non-adaptive tutoring. Qualitative observations further reveal that the presence of the robot fostered motivation, engagement, and collaborative classroom dynamics. This work demonstrates the potential of social robots to support group learning in under-resourced educational settings and highlights the importance of extending HRI research beyond well-resourced contexts.
In collaborative assembly, humans and robots must coordinate their actions efficiently to achieve a common goal. The quality of this coordination is often measured by fluency, which refers to interactions being perceived as mutually engaging and highly synchronized. While fluency has been studied in various task assignment models, from sequential turn-taking to reciprocal effort, most research focuses on predictable interactions. However, in-the-wild applications of human-robot collaboration also include workflow interruptions, such as when a robot performs an incorrect task. We investigated which interaction modality best facilitates both communicating errors to the robot and fluent workflow recovery. In our study (n=29), participants engaged in a collaborative assembly task with deliberate robot failures. They then had to communicate the failure to the robot, and we examined whether the interaction modality influenced perceived fluency and the time it took to complete the task after signalling the error. The three interaction modalities tested included a graphical user interface, haptic interactions, and detection of implicit cues via a camera-based system. Our results indicate that implicit communication led to the fastest task recovery, the highest perceived fluency, and the strongest sense of user-robot bonding.
Robots interacting with humans pose data privacy risks, potentially leading to uncontrollable leaks of personal information. Robots interact with older adults in public places, homes, care facilities, and hospitals. Older adults may have concerns about privacy issues related to robots and would benefit from developing their robot literacy skills regarding data privacy. Research on older adults’ robot literacy in relation to data privacy is scarce. We conducted a qualitative and explorative human-centered design study with care home residents (N=9) to explore their perceptions of and interest in robot data privacy literacy. In the study, they interacted with an early prototype of a robot-assisted learning application implemented on the social robot QTrobot. Participants were concerned about the "superpowers" and data storage of robots. Based on our findings and existing literature, we redesigned the prototype into "Myth Buster," a robot-assisted learning application aimed at enhancing older adults’ data privacy literacy regarding robots. Our work contributes to the understanding of older adults’ data privacy literacy, which is currently under-researched in Human-Robot Interaction. We also present design-relevant insights for developing robot-assisted data privacy learning applications to enhance robot literacy of older adults.
Unexpected robot encounters in public spaces can cause discomfort for third parties, yet the acceptability of accompanying robots is not yet fully understood. Visual cues strongly influence robot impressions; however, the interaction between explicit visual indicators and individual robot resistance in determining acceptability remains unexplored. We investigated how visual relationship indicators—connection visibility (the physical linkage between the robot and handler) and control visibility (the evident authority of the handler)—influence acceptance based on individuals’ levels of robot resistance. In our experiment, 23 participants encountered a mobile robot under three operation methods: Autonomous (no visual indicators), Joystick (only control visibility), and Leash (both connection and control visibility), with participants divided into high-resistance (n=12) and low-resistance (n=11) groups based on their NARS scores. Results indicate that Leash had the highest acceptability, with high-resistance participants showing significant differences across methods and benefiting from explicit visual indicators, unlike low-resistance participants who were largely unaffected. These findings offer important design implications for accompanying robots in public spaces, suggesting that employing visually explicit relationship indicators is an effective strategy for enhancing acceptability, particularly among individuals with robot resistance.
Various strategies have been explored for robots to prevent navigation blocks. However, such blocks may still happen and robots then need to resolve them. Blocks may happen either on the robot’s path or target location and may be caused either by human or object obstacles. In this paper we explore the design of robot behaviors to resolve such navigation blocks by asking humans for help. These behaviors combine different communication modalities in steps with increasing urgency. We focus on robots with limited sensing capabilities and present findings from an in-person experiment evaluating these behaviors. Our findings illustrate that humans can be more easily engaged to solve blocks caused by themselves rather than by third party objects. They also highlight the complexity of having robots with limited sensing capabilities successfully enact sequential interactions with their bystanders.
This study investigates whether personalized voice cloning can improve a robot’s likability compared to a design-congruent voice and a distinctly dissimilar voice. Participants interacted with a gender-ambiguous android robot in three different voice conditions. We compared: (1) a personalized voice clone based on the participant’s voice, (2) a design-congruent voice matching the robot’s appearance, and (3) a dissimilar voice, which differs from both the participant’s and the robot’s features.The cloned and design-congruent voices significantly increased likability compared to the dissimilar voice, while anthropomorphism and familiarity showed no significant differences across conditions. Most participants did not immediately recognize their cloned voice until informed that one of the voices was a clone. However, most of the participants were successful when asked to pick out their cloned voice from those used. We assume that voice personalization through similarity to the user improves likability even before the user is aware of this similarity.Our results show that personalized voice cloning is a simple alternative to other methods for the design of robotic voices. It significantly increases robot likability while requiring minimal user effort.
Assistive robots will be more effective if they can accurately reason about the intentions and beliefs of the user (i.e., have Theory of Mind (ToM)). ToM benchmarks allow us to examine how well an artificial agent (e.g., robot) is able to do ToM reasoning in a given scenario. However, there is a need for ToM benchmarks that are more representative of the challenges faced in assistive robotics. Existing benchmarks from AI and HRI make simplifying assumptions, such as simply defined goals, plans that are indicative of goals, and no user errors. To address the challenges from relaxing these assumptions, we propose the Theory of Mind of Children Assembling Tangrams (ToMCAT) dataset. The data is derived from videos of children building tangram puzzles while being assisted by a social robot. As a baseline benchmark, we evaluated two approaches for how well they can recognize which puzzle this child is building based on a single observation. Analogical reasoning correctly recognized the puzzle more than 75% of the time and had perfect accuracy for puzzle states that were close to complete. However, an out-of-the-box commercial LLM correctly recognized the puzzle only 60% of the time and was accurate on less than 80% of the completed puzzles. Our results suggest that the ToMCAT dataset offers challenges for recognizing the intended puzzle of a child. Furthermore, the dataset provides opportunities to examine additional ToM reasoning capabilities. Overall, the ToMCAT dataset provides a useful benchmark to facilitate the advancement of ToM reasoning for assistive robotics.
In this paper, we study the problem of gesture recognition as a method for divers to communicate with an underwater robot. Gesture is a common method of communication between divers, and yet autonomous underwater vehicles have very limited capacity to understand gesture given lighting and visibility constraints (e.g., from water turbidity and diver depth). Traditional deep learning methods are limited in this domain because of a lack of sufficient training data. We show that it is not enough to learn a gesture in a laboratory setting, because the appearance changes dramatically underwater. We show how hyperdimensional computing can solve this problem by permitting hypervectors to serve as abstract representations of gestures, yielding rapid adaptation to new environments and new gestures. We experimentally verify this approach using a novel dataset of 6 diving relevant gestures. We show that we can accurately adapt to a gesture learned in a laboratory setting to work with a gesture observed underwater. Our approach compares favorably to a ResNet-18, which performs well in laboratory conditions (91.9% accuracy), but performs poorly underwater (53.9% accuracy). Our proposed approach is capable of rapid adaptation, resulting in an accuracy of 83.8% on underwater gestures with just one additional example from each class added to the support set. Finally, we also show the ability to adapt to new gestures not present in our original training set. We use hypervectors to learn new gestures from the Sign Language MNIST dataset, providing a high level of accuracy with a limited amount of training data.
The perception of danger in HRI settings has become increasingly important as interactions between robots and humans become more commonplace. Previously, a perceived danger scale was developed and validated. Here, we shortened this scale to create the Perceived Danger-Short Form (PD-SF) scale. Experiment 1 used pre-existing data and standard procedures to shorten the scale from 12 items to 4. Experiment 2 validated the short form in a new experiment where participants observed images of robots holding kitchen items of varying levels of danger in close proximity to a human. PD-SF was able to capture differences across the kitchen items. Results from both experiments indicate that PD-SF is a reliable and psychometrically valid measure of perceived danger in HRI contexts.