For effective human-robot collaboration, a robot must align its actions with human goals, even as they change mid-task. Prior approaches often assume fixed goals, reducing goal prediction to a one-time inference. However, in real-world scenarios, humans frequently shift goals, making it challenging for robots to adapt without explicit communication. We propose a method for detecting goal changes by tracking multiple candidate action sequences and verifying their plausibility against a policy bank. Upon detecting a change, the robot refines its belief in relevant past actions and constructs Receding Horizon Planning (RHP) trees to actively select actions that assist the human while encouraging Differentiating Actions to reveal their updated goal. We evaluate our approach in a collaborative cooking environment with up to 30 unique recipes and compare it to three comparable human goal prediction algorithms. Our method outperforms all baselines, quickly converging to the correct goal after a switch, reducing task completion time, and improving collaboration efficiency.
IntroductionWhile shared activities foster connection between people living with dementia (PLWD) and their care partners, emotional distress and daily caregiving responsibilities often make them difficult to initiate. This paper investigates the adaptation of a socially assistive robot, Ommie, to guide shared deep breathing and singing activities for these pairs.MethodsWe refined the robot’s behaviors through two interaction design sessions with people living with dementia and care partners, mediated by an occupational therapist. In a subsequent study with 17 pairs, participants engaged in deep breathing and singing activities with the robot as well as in-session semi-structured interviews, and we conducted post-hoc video analysis to explore their interactional dynamics.ResultsParticipants reported the interactions as easy to follow, calming, and familiar. Post-hoc video analysis revealed patterns of intimacy and synchrony, including frequent physical touch, mutual gaze, and rhythmic coordination. We also observed instances of personal memory recall and a playful atmosphere, in which pairs often used humor as a coping mechanism after deviations from the robot’s instructions.DiscussionFrom our observations, we discuss three design opportunity spaces: the robot as the focus for synchronization, as an instrument of joint play, and as a source of familiarity versus variety.
To collaborate with humans, robots must infer goals that are often ambiguous, difficult to articulate, or not drawn from a fixed set. Prior approaches restrict inference to a predefined goal set, rely only on observed actions, or depend exclusively on explicit instructions, making them brittle in real-world interactions. We present BALI (Bidirectional ActionLanguage Inference) for goal prediction, a method that integrates natural language preferences with observed human actions in a receding-horizon planning tree. BALI combines language and action cues from the human, asks clarifying questions only when the expected information gain from the answer outweighs the cost of interruption, and selects supportive actions that align with inferred goals. We evaluate the approach in collaborative cooking tasks, where goals may be novel to the robot and unbounded. Compared to baselines, BALI yields more stable goal predictions and significantly fewer mistakes.CCS Concepts center dot Computing methodologies -> Reasoning about belief and knowledge.
Human-Robot Interaction (HRI) continues to rely on commercial social robot platforms to support academic research. Yet again and again, these systems prove short-lived, inaccessible, or misaligned with research needs. We argue that this is not an industry problem the goals, needs, and constraints of industry are inherently distinct. Instead, this is a fundamental structural problem in HRI research, and one that must be solved from within. In short, HRI researchers must build their own products. In this paper, we trace the recent problems of industry-supplied robots and frame a new type of HRI research artifact in response: Deployable Research Products (DRPs), which bridge the gap between lab prototypes and commercial products. Drawing on mental models from business and innovation theory, we outline the mindset shifts that HRI must embody to move towards DRPs. We conclude with three emerging examples of this alternative path in the HRI community. These projects differ in scope and approach but share a common thread: to ensure the longevity of our science, we cannot outsource what we value most.
The fields of human-robot interaction (HRI) and robotics at large have developed around a stable set of assumptions about what robots are and how they should behave. These assumptions arise from the constitutive traits of robots, which together shape social expectations. Over time, these expectations have hardened into tacit rules that quietly govern research and design: robots should always engage, help, be productive, remain polite, never lie, never err, and never model harm. While these prevailing norms have merit, they also constrain the field's imagination of the interactions robots can meaningfully support. We propose rule-breaking as a generative design strategy and illustrate how deliberate violations-robots that interrupt, refuse, mislead, or err-can produce interactions that are more ethical, effective, and socially intelligent. In doing so, we argue for a more reflexive and imaginative HRI that learns as much from breaking the rules as from following them.
Robots are often designed to help, but help is not always helpful. In everyday situations, it is a socially delicate act: the right offer of help at the wrong moment can be intrusive, unnecessary, or even undermining. In this article, we challenge the prevailing assumption that robots should always offer help, prompting an essential discussion of how robots can discern when to offer help. We introduce a theoretical framework that enables robots to assess the appropriateness of offering help by considering factors such as the relative skill levels of the robot and human user, as well as the social value and cost of assistance. To validate this framework, we conducted a large-scale online study in which participants rated the appropriateness of robot assistance across diverse task scenarios. Their responses supported our core predictions and highlighted additional contextual factors. Building on these results, we discuss potential extensions of the simplified model for real-world settings, including uncertainty management, perception of ability, autonomy preferences, and social presence. We present these directions as opportunities for future research.
We propose a unified strategy for fast goal inference in human–robot interaction. The core idea is to drive the human toward Critical Decision Points (CDPs)–states where competing human strategies prescribe different next actions and thus maximally reveal the goal. We formalise CDPs using a goal-conditioned policy divergence measure and incorporate them into a Receding-Horizon Planner that explores future action sequences while optimizing a cost function balancing task progress and information gain. We evaluate this approach in both a collaborative, fully observable cooking task and a competitive, partially observable hide-and-seek game, each in simulation and on real robots. In both scenarios, our method infers human goals more accurately and earlier than baseline strategies.
Transformer adaptation is typically distributed across model depth, even when the intended change is narrow. We investigate how adaptation site shapes what a model learns, how well that learning generalizes, and how selectively it is applied. We introduce a controlled benchmark spanning five objectives (lexical binding, factual association, behavioral policy learning, causal mapping, and procedural reasoning) and define each objective's "adaptation geometry" as its profile of acquisition, transfer, and boundedness under full-stack and early-, middle-, or late-layer LoRA. The objectives exhibit distinct geometries. Lexical binding favors early-layer adaptation for acquisition and boundedness but requires broader updates for transfer; factual association favors later layers among localized adapters; behavioral learning separates late-layer action acquisition from middle-layer policy gating; and causal and procedural transfer benefit most from middle- or full-stack adaptation. These patterns largely persist under parameter-matched controls, and most corresponding directional contrasts replicate across five model families. These findings establish adaptation site as a key design variable for controlling what models learn, generalize, and leave unchanged.
Instruction-based suppression is widely used to prevent language models from generating prohibited content, yet it remains unclear whether suppression reduces internal representation or merely suppresses expression. We investigate this question through representational probing, attention analysis, and behavioral semantic leakage experiments across multiple transformer models. We find that prohibited concepts remain highly recoverable from hidden representations under suppression, continue to influence attention routing, and measurably shape downstream generations despite successful lexical avoidance. These effects persist across pooling strategies, indirect semantic controls, and multiple model families. Our results expose a fundamental gap between behavioral and representational alignment.
A growing body of work on robot-mediated therapy suggests that robots can elicit engagement and novel social behaviors among users with autism. In this narrative review, we trace the full historical arc of this field to date, from its inception in 2001 to 2024, covering 304 studies that present a robot for autism support. Early work largely consisted of short, highly structured sessions conducted in controlled laboratory or clinical environments. More recent research has shifted toward longer-term, real-world deployments in which robots operate with greater autonomy and engage users over multiple days or weeks. However, evidence for lasting and generalized benefits is still limited. The literature also remains focused primarily on children, with comparatively little research involving adults or individuals across a wider range of support needs. Despite rapid growth over these past two decades, the research remains fragmented across disciplinary boundaries, with robotics and clinical communities advancing similar goals without a cohesive, shared research framework. To address this, we offer a translational roadmap that details what interventions work, for whom, under which conditions, and why. We highlight current methodological gaps, outline key design considerations, and propose priorities for future research.
Foundation models are increasingly deployed in socially sensitive domains such as education, mental health, and caregiving, where failures are often cumulative and context-dependent. Existing guardrail approaches – ranging from training-time alignment to prompting, decoding constraints, and post-hoc moderation – primarily provide empirical risk reduction rather than enforceable behavioral guarantees, and largely treat safety as a property of individual outputs rather than interaction trajectories. We reframe guardrails as a problem of runtime behavioral control over interaction trajectories, drawing on robotics to introduce formal constructs for constraint enforcement in uncertain, closed-loop systems. We instantiate these ideas in the Grounded Observer framework and apply it across three real-world deployments: small talk, in-home autism therapy, and behavioral de-escalation in schools. Across settings, the framework enables runtime interventions that mitigate drift into undesirable interaction regimes while adapting to diverse social contexts. We discuss extensions to the framework and propose research directions toward stronger guarantees.
As robot deployments become more commonplace, people are likely to take on the role of supervising robots (i.e., correcting their mistakes) rather than directly teaching them. Prior works on Learning from Corrections (LfC) have relied on three key assumptions to interpret human feedback: (1) people correct the robot only when there is significant task objective divergence; (2) people can accurately predict if a correction is necessary; and (3) people trade off precision and physical effort when giving corrections. In this work, we study how two key factors (robot competency and motion legibility) affect how people provide correction feedback and their implications on these existing assumptions. We conduct a user study (N=60) under an LfC setting where participants supervise and correct a robot performing pick-and-place tasks. We find that people are more sensitive to suboptimal behavior by a highly competent robot compared to an incompetent robot when the motions are legible (p=0.0015) and predictable (p=0.0055). In addition, people also tend to withhold necessary corrections (p < 0.0001) when supervising an incompetent robot and are more prone to offering unnecessary ones (p = 0.0171) when supervising a highly competent robot. We also find that physical effort positively correlates with correction precision, providing empirical evidence to support this common assumption. We also find that this correlation is significantly weaker for an incompetent robot with legible motions than an incompetent robot with predictable motions (p = 0.0075). Our findings offer insights for accounting for competency and legibility when designing robot interaction behaviors and learning task objectives from corrections.
In the past two decades, the field of social robotics has undergone significant growth, witnessing a surge in long-term human-robot interaction (HRI) studies. This review paper provides an in-depth analysis of 120 long-term HRI studies conducted between 2003 and 2023, spanning 7 major domains including education, entertainment, and physical and mental health. We define "long-term" as studies deploying social robots with the same users for more than three sessions across 3 consecutive days, aiming to employ a comprehensive approach and identify trends in this dynamic field. Our analysis explores various aspects of these studies, from participant demographics to the characteristics of the HRI and engagement measures. The findings reveal promising trends, such as diverse age group representation, a strong focus on real-world contexts, and autonomous robot operation. We also identify gaps, notably the limited representation of studies involving teenagers and those studying workplace settings. By presenting this overview, we aim to empower the HRI community to address challenges, refine methodologies, and foster innovation in the domain of long-term HRI.
Mistakes, failures, and transgressions committed by a robot are inevitable as robots become more involved in our society. When a wrong behavior occurs, it is important to understand what factors might affect how the robot is perceived by people. In this paper, we investigated how the type of transgressor (human or robot) and type of backstory depicting the transgressor's mental capabilities (default, physio-emotional, socio-emotional, or cognitive) shaped participants' perceptions of the transgressor's morality. We performed an online, between-subjects study in which participants (N=720) were first introduced to the transgressor and its backstory, and then viewed a video of a real-life robot or human pushing down a human. Although participants attributed similarly high intent to both the robot and the human, the human was generally perceived to have higher morality than the robot. However, the backstory that was told about the transgressors' capabilities affected their perceived morality. We found that robots with emotional backstories (i.e., physio-emotional or socio-emotional) had higher perceived moral knowledge, emotional knowledge, and desire than other robots. We also found that humans with cognitive backstories were perceived with less emotional and moral knowledge than other humans. Our findings have consequences for robot ethics and robot design for HRI.
This article discusses the design, development, and evaluation of Ommie , a novel socially assistive robot that supports deep breathing practices for the purposes of anxiety reduction. Research has shown that practicing deep breathing (breathing while extending one’s inhales, holds, and exhales) has a strong capacity to calm the autonomic nervous system and reduce anxiety. The robot’s primary function is to guide users through a series of deep breaths by way of haptic interactions and audio cues. We utilized a user-centered design approach and present our design methodology in addition to core decisions across robot morphology, tactility, and interactivity. As reported in prior work, the final robot prototype was tested with a two-cohort usability study (n = 43) at a local university wellness center, including participants with anxiety and those with varying levels of experience with deep breathing. Interacting with Ommie resulted in a significant reduction in STAI-6 anxiety measures across all participants, who also found the robot intuitive, approachable, and engaging. Participants also reported feelings of focus and companionship when using the robot, often elicited by the haptic interaction. This article describes how our design process and design goals contributed to these results showing Ommie’s capacity for supporting those with anxiety. Our work also serves as an example of how researchers can design robots for behavioral practices for mental health.
In this work, we introduce and formalize the Zero-Knowledge Task Planning (ZKTP) problem, i.e., formulating a sequence of actions to achieve some goal without task-specific knowledge. Additionally, we present a first investigation and approach for ZKTP that leverages a large language model (LLM) to decompose natural language instructions into subtasks and generate behavior trees (BTs) for execution. If errors arise during task execution, the approach also uses an LLM to adjust the BTs on-the-fly in a refinement loop. Experimental validation in the AI2-THOR simulator demonstrate our approach’s effectiveness in improving overall task performance compared to alternative approaches that leverage task-specific knowledge. Our work demonstrates the potential of LLMs to effectively address several aspects of the ZKTP problem, providing a robust framework for automated behavior generation with no task-specific setup.
Many schools have built de-escalation and sensory rooms to support students who experience heightened emotional states, sensory overload, or difficulty self-regulating in traditional classroom settings. Yet, effective implementation remains challenging due to diverse student needs and resource constraints. Hence, we developed RESET (Robot-Enhanced Social-Emotional Therapy), a robot for facilitating students’ self-regulation in their school’s existing de-escalation space. We present our co-design process, iterative development, and final system components. Following a fully autonomous, month-long deployment in an elementary school, we assessed the robot’s usability and impacts. Results indicate RESET integrated well into the school environment, promoting more efficient deescalation, smoother transitions back to classroom learning, and lasting impacts beyond its deployment period.
Background: Allergy procedures, such as oral food and drug challenges, are the diagnostic criterion standards, but may cause increased anxiety in patients and caregivers. Interventions that may mitigate allergy-related anxiety are now being explored to ease the burden these procedures pose. Objective: The purpose of this feasibility study was to explore the use of robots to reduce the psychosocial effects of allergy procedures. Methods: A robot designed to support and guide a user through deep breathing practices was made available to 10 patients and their caregivers during outpatient oral challenges. Aspects of the interactions between the patients and caregivers and the robot were recorded. Results: The interest and the number of interactions with the robot varied among the individuals, but most patients interacted with the robot several times. Most caregivers reported that the robot had a positive impact on their child’s experience and would like the robot to be present at future procedures. Conclusion: This feasibility study explored the integration of a robot in the setting of allergy procedures. In future work, validated measures of allergy-related anxiety and comparison with other forms of distraction should be used to further assess the robot’s clinical impact. This study lends support to the idea that robotics may serve as a valuable and complementary tool to improve patient and caregiver experiences during allergy procedures.
Adriana Tapus合作论文数Human-Robot Interaction (HRI) conference 20094