
To recognize continuous hand gestures from real-time video streams quickly and enable natural human-robot interaction, this paper proposes continuous hand gesture-based human-robot interaction for indoor mobile robot based on Multi-scale Local and Global skeleton Spatiotemporal feature extraction Network (MLGSNet). Firstly, we propose spatiotemporal feature extraction method of hand skeletal joint flow based on uniform sampling to address the challenge of extracting spatiotemporal features from videos. Subsequently, we propose multi-scale local and global spatiotemporal graph convolution network for hand gesture recognition to effectively capture both local and global dynamic spatiotemporal features during gesture execution, which is characterized by combining multi-scale attention, graph and temporal convolution. Furthermore, we design hand gesture activation based on mean filtering of dual confidence to accurately activate continuous gesture streams and recognize gesture command units under uneven gesture stream distribution. Experimental results on two public datasets demonstrate that the MLGSNet achieves state-of-the-art performance in both isolated and continuous gesture recognition. Finally, case study on the robot platform shows an overall human-robot interaction performance of 93.05
Socially aware robots must adhere to human social norms, such as respecting social groups by avoiding robot intrusions. Consequently, identifying individuals and social groups constitutes a critical perception task for social navigation. However, existing laser scanner-based perception methods face challenges such as partial occlusion and high latency. In this study, an improved laser scanner-based neural radiance field (NeRF) is proposed to identify social groups with two-dimensional (2D) laser scanner inputs. This NeRF method utilizes a distance field to reconstruct partially occluded laser points and a density field to filter out distant noise. To further minimize identification errors in dynamic and crowded environments, a novel social interaction field is introduced, which assesses social features including motion direction consistency and density invariance. The improved NeRF method has been validated in both public datasets and real-world experiments. Experimental results show that the proposed method achieves equal error rates (EERs) of 0.213, 0.170, and 0.195 in laboratory, corridor, and outdoor environments, respectively. Furthermore, the method achieves occlusion recovery rates (ORR) of 84.2
The effectiveness of older adult care is one of the main challenges in healthcare as the world’s population ages. In this work, we present the application of a rehabilitation platform based on social assistive robotics, in which a humanoid robot acts as a co-therapist in sessions prescribed by clinical professionals. The platform, known as Inrobics Rehab, was originally developed as a research project focused on rehabilitation for pediatric patients with motor impairments. We propose extending this system to support physical and cognitive stimulation in older adults. A five-week evaluation was conducted with 10 residents of a nursing home, in which users participated in one-on-one therapy sessions with the robot twice a week. The study primarily assessed usability, user experience, social acceptance, and social impact using qualitative and quantitative methods. Results show that users found the platform easy to understand and use and maintained a high engagement during all the sessions of the study. The robot was accepted by older adult users, who expressed a desire to continue using the robot in their rehabilitation process. These findings highlight the feasibility of integrating this tool into the daily lives of the nursing home. Clinical outcomes were beyond the scope of this study and require further research, along with the system’s accessibility and adaptation to individual user capabilities.
Social robots have developed rapidly in educational settings, where they are used as tutors, teaching assistants, and learning partners. However, earlier quantitative syntheses have often pooled embodied robots with chatbots, virtual agents, and broader AI-based learning systems, making it difficult to determine the specific contribution of physical embodiment. To address this construct-validity issue, the present systematic review and meta-analysis narrowed the inclusion criteria to controlled empirical studies involving physically present social robots. Twenty-nine studies involving 1,543 participants and 37 effect sizes were included. Hedges’ g was calculated for learning achievement, knowledge retention, and learning motivation. The results showed a significant positive effect on learning achievement, whereas the effects on knowledge retention and motivation were positive but not statistically significant. Considerable heterogeneity, significant funnel-plot asymmetry, and limited evidence for retention and motivation indicate that the findings should be interpreted cautiously. The study provides more focused evidence on embodied social robots in education and identifies priorities for future research.
Social robots are increasingly present in homes, schools, and care settings, but in regulations and international standards they are typically classified as mobile service robots (MSRs). What qualifies as an MSR varies across jurisdictions and standards, creating definitional ambiguity with practical safety consequences: for example, Paro, the robotic seal, is classified as a medical device in the U.S. but treated as non-medical in the EU, complicating which safety requirements apply. This problem is amplified by ongoing changes in standards, including ISO 13482’s shift from ‘personal care robots’ (2014) to broader ‘service robots’ in the 2024 draft. To support clearer hazard analysis, testing, and oversight, we systematically review MSR definitions in research and standards and map resulting misalignments. Using PRISMA, we synthesize recent literature and compare it with ISO 13482:2014, the revised ISO/DIS 13482:2024, and ISO 8373:2021 on robotics vocabulary. Our results show that vulnerable groups, especially children, older adults, and pregnant women, remain insufficiently protected; emotional and cognitive risks are rarely specified; and standards continue to emphasize physical hazards. Many robots described in the literature do not clearly align with ISO MSR criteria, creating uncertainty for hazard analysis and conformity assessment. Based on these findings, we propose a refined MSR definition incorporating diverse embodiments, services, and social interaction.
This paper presents a protocol to assess interactions between children and social robots in real-world educational environments, with the objective of evaluating the impact of social robotics as an educational tool to support teachers in their daily classroom activities. Although previous studies have proposed methods and variables for assessing child–robot interaction, most have been conducted in laboratory settings and based on 1:1 interaction paradigms. This work introduces a multi-construct protocol and experimental procedure designed for natural classroom environments. Four constructs - Intentional Social Acceptance, Perceived Enjoyment, Social Presence and Anxiety- were used to evaluate the interaction. The protocol was validated with 93 first-grade students in a public primary school in Madrid, Spain. High levels of acceptance and enjoyment, more than 90
This study aims to explore the integration of Large Language Models (LLMs) in Socially Assistive Robots (SARs) to enhance personalized interaction and companionship for older adults living independently. It evaluates user acceptance, perceived quality of interaction, and conversational capabilities of an LLM-powered social robot in real-world domestic environments. The SHARA robot (with GPT-4o-mini based conversational system and personalized memory) was deployed for one week in the homes of 20 adults aged 55–82. Participants freely interacted with the robot, completed interaction diaries, and evaluated the system using the Almere model and Godspeed questionnaire. Robot logs and qualitative feedback were also analyzed. Participants reported high acceptance and enjoyment. Almere model scores exceeding 4.0 in most dimensions, particularly Perceived Sociability (M = 4.588), Perceived Enjoyment (M = 4.510) and Attitude Towards Technology (M = 4.550). Anxiety remained very low (M = 1.34). In the Godspeed test, the LLM-powered robot was perceived as likeable (M = 4.75) and intelligent (M = 4.56). Users reported developing emotional bonds with the robot, despite being aware of its artificial nature. Diaries and interaction logs revealed frequent, diverse, and engaging conversations, with an average of 14.3 interactions per conversation and 4.5 conversations per day. LLM-powered SARs can foster emotionally supportive and engaging interactions with older adults. Positive perceptions of sociability and usefulness were strongly correlated, indicating the value of conversational competence in assistive contexts. Future improvements should focus on local LLM integration, proactive behaviors, and extended functionality to enhance long-term adoption and impact.
Human-Robot Interaction (HRI) presents significant challenges in accurately assessing situations, adapting robotic behavior to human intentions, ensuring explainability, pertinence, and acceptability, and effectively managing uncertainty. Traditional model-based approaches provide reliability but struggle with human unpredictability, often approximating human behavior through specific models that fail to encompass all possible scenarios. Conversely, statistical approaches, such as Large Language Models (LLMs), offer greater adaptability but lack deterministic guarantees. This paper introduces a hybrid architecture that integrates structured methodologies with the flexibility of LLMs to enhance robotic coaching in dynamic environments. The proposed architecture leverages deterministic modules for critical constraints and safety guarantees while employing LLMs for contextual understanding, natural language interaction, and adaptation to unpredictable human behaviors. Through a healthcare robotic coach scenario implementation and the conduction of a user study, we aim to demonstrate how this balanced approach enables effective monitoring of task execution, dynamic adaptation to human states, and seamless verbal interaction while maintaining system reliability. By bridging deterministic and statistics-based techniques, the proposed architecture aims to advance HRI toward safer, more transparent, flexible, and adaptive interactions.
The present work investigated humans’ responses to robots’ verbal emotional expressions. To this end, the interplay of agent characteristics and signals of communicative intent on emotional word processing was examined. In two behavioral studies (N1 = 36, N2 = 72), participants watched short videos featuring different agents (humans, androids, and humanoid robots) alongside the auditory presentation of single words of varying valence (negative, neutral, and positive). In the communicative condition, the agent signaled communicative intent by looking at the participant while speaking; in the non-communicative condition, the audio played while the agent remained passive with eyes and mouth closed. After each video, participants rated the valence and arousal of the presented word. The results of both studies revealed a consistent pattern, with differences in emotional word processing based on agent type: For human agents, signals of communicative intent enhanced the emotional impact (arousal dimension) of emotional words. In contrast, for android and humanoid robot agents, the corresponding effects were smaller and statistically non-significant, indicating a relatively lower perceived relevance of direct emotional speech from robots. For all agents, the emotional meaning (valence dimension) of words was robustly conveyed without being affected by the communication manipulation. Further analyses identified the dimension of perceived animacy as a basis for the observed arousal effects. In line with theories that emphasize contextual influences on language processing, our results highlight the role of mind attribution and perceived relevance for emotional language processing in human–robot interaction.
Healthcare robots at home are increasingly essential for promoting the independence of older adults, yet their widespread acceptance is hindered by a lack of clarity regarding optimal design features, particularly among users with varying levels of knowledge and attitudes towards this emerging technology. To address this, this study applies the Kano model to classify and prioritize healthcare robot features based on their impact on user satisfaction and design decisions, factoring in older adults diverse robot-related knowledge and attitudes towards robots. Following a thorough literature review that highlighted 27 distinct robot features, we conducted a survey with 253 community-dwelling older adult participants and identified essential features such as ‘Medication Management’ and ‘Managing Illness and Monitoring Health’ as one-dimensional features, whereas ‘Animal-like Appearance’ was negatively received. The Kano model classifications including must-be, one-dimensional, attractive, indifferent, and reverse, offer direct guidance for design priorities by identifying which features are most likely to enhance satisfaction when included and cause dissatisfaction when absent or poorly implemented. The analysis also showed that user preferences vary significantly with their knowledge and perception of robots. These insights emphasize the need to tailor healthcare robots to the initial expectations of community-dwelling seniors, prioritizing functional features that support daily independence over therapeutic care.
Expressing emotional motions is one of the key social capabilities for the humanoid robots. However, most humanoid robots are mainly used to complete specific tasks, lacking the ability to express emotional motions. In this paper, the Pleasure-Arousal-Dominance (PAD) model is introduced for emotional motion design of the humanoid robots, in order to enrich its emotional expression. A detailed rule that converts PAD parameters to 24 motion parameters of the humanoid robot is proposed. Based on this rule, 16 emotional motions can be rapidly implemented in a humanoid robot via a dual-arm exoskeleton. To evaluate the effectiveness of the proposed conversion rule, an evaluation is conducted. Participants undertake the experiment in an online or on-site format, and observe the humanoid robot’s emotional motions from two distinct viewpoints. The results show that the recognition rate of emotional motions is up to 71% , and the average score for the human-robot interaction impression is 3.58 out of 5. Additionally, the velocity and frequency, compared with other motion parameters, have greater influence on the effectiveness of the humanoid robot expressing emotional motions. This work can contribute to guiding the emotional motion design of the humanoid robots in social interaction.
Despite increasing interest in and widespread use of robots in education, there is limited research on the effect of voice gender and communication style on affective outcomes. Based on a Q A interaction model, a within-subject experiment was conducted with a 2 × 2 design, manipulating voice gender (male vs. female) and communication style (functional vs. relational). The study combined subjective evaluations with eye-tracking and fNIRS measurements. We confirmed the effects of voice gender and communication style on subjective perceptions, reported the influence of voice gender on users’ visual attention, and demonstrated the interaction between voice gender and communication style on users’ cerebral activity. The results suggest that the female voice and relational communication style are preferable for educational robots. This study enhances the understanding of users’ visual attention and cerebral activity during interactions with educational robots, providing valuable insights for designing robot voices in educational contexts.
Global aging has intensified staffing shortages in nursing homes worldwide, while Socially Assistive Robots (SARs) often underperform in real-world care settings relative to laboratory conditions. Yet elderly residents may still report positive emotional benefits even when technical performance is limited. This seemingly contradictory pattern, observed in our fieldwork, is described here as the tech-emotion paradox. To explain this pattern, we propose a Robot-Caregiver-Resident (RCR) framework grounded in Social Support Theory that distinguishes three mechanisms: Remedial Intervention, Emotional Support, and Motivational Guidance. We examine these mechanisms using an institution-linked serial-parallel mediation model. Using a mixed-methods design, we first conducted qualitative fieldwork in a representative nursing home in eastern China and then used Structural Equation Modeling (SEM) with survey data from seven nursing homes equipped with SARs (Caregiver N = 251, Resident N = 238) to examine associations among robot performance, caregiver practices, and residents’ emotional outcomes. The model showed a good fit (CFI = 0.96). The results indicated no significant direct association between robot performance and residents’ emotional outcomes, but significant indirect associations through Emotional Support and Motivational Guidance, whereas Remedial Intervention functioned more as damage control than as a pathway to positive emotional outcomes. Together, these findings provide institution-linked quantitative support for the RCR framework and show that positive resident emotional outcomes in nursing home SAR use are more closely linked to caregiver-mediated socio-emotional pathways than to technical performance alone.
Commensality, the social act of dining collectively, serves as a fundamental mechanism for fostering social bonds and intimacy. However, a significant number of older adults experience diminished commensality due to a scarcity of companionship in aging populations. Past studies have suggested that robotic companionship may be one solution, especially since commensality robots that can track their interlocutors’ emotions can provide higher-quality companionship. However, there has been a lack of research on how to accurately measure users’ genuine emotional state during interaction with a robot to justify the robot’s design. This study employed two techniques, electroencephalography (EEG) and a self-assessment manikin (SAM), to measure reflective emotions during co-eating interactions between older adults and either robots or stuffed animals, including high- and low-social-presence robot groups, as well as within-group comparisons of robot and stuffed-animal interactions. Results showed that the SAM arousal effect was significant only between the robot and stuffed animal groups. The values of other SAM and EEG measures were consistent between and within groups, with no significant differences in SAM valence or EEG. The measures demonstrated convergent validity for valence, consistent across and within conditions, despite a weak individual-level correlation. Divergent validity was observed for arousal, aligning with expectations since frontal alpha asymmetry indexes valence rather than arousal. EEG and SAM thus show convergent validity for valence and appropriate divergence for arousal, though further studies with greater statistical power are needed to confirm the strength of valence convergence. The findings of this study support the feasibility of testing users’ emotions in real-world human-robot interactions, underscoring the significance of this research.
With the integration of large language models (LLMs) in humanoid robots, expectations are rising for face-to-face dialogue and personalised services. However, as these models are primarily trained on text, their responses may not be well suited for voice interactions, where response brevity and contextual relevance are particularly important. This study investigates how three response styles (i.e. generic, concise contingent, and lengthy contingent) affect user satisfaction. A between-subjects experiment (N = 96) was conducted, with contingent responses generated by an LLM and generic responses selected from a predefined phrase library. Results indicated that compared to generic responses, concise contingent responses produced higher satisfaction through an increased sense of feeling heard. Furthermore, feeling heard enhanced satisfaction through perceived intelligence and perceived connection, indicating a cognitive–affective dual pathway. However, when the contingent responses shifted from concise to lengthy, participants experienced greater annoyance, which in turn reduced satisfaction. This indirect effect of lengthy (vs. concise) contingent responses on satisfaction via annoyance was moderated by user sociability: it was significant among high-sociability participants but not among low-sociability participants. Practically, these findings suggest that service robots should tailor response length to individual traits and achieve a balance between informativeness and brevity to maintain user satisfaction.
Internet-based applications of Cognitive Behavioral Therapy (iCBT) are promising to alleviate mood problems and depression. However, versions without personal support (i.e., unguided) still fall short in establishing meaningful therapeutic relationships, despite recent versions also embedding virtual avatars, potentially resulting in low levels of adherence. The current paper systematically examined whether a physically present embodied social robot, automatically operating (i.e., no WOz), may yield stronger therapeutic alliance and increase adherence compared to screen-based therapeutic interventions through experimental research designs. In two separate experiments, participants were randomly assigned to either a robot intervention or screen-based control condition with or without an avatar, addressing mild mood issues in three subsequent intervention sessions. Overall, results showed the social robot intervention to increase therapeutic alliance, adherence, and satisfaction compared to the same screen-based intervention, with alliance in a mediating role, but did not significantly improve participants’ mood. Future research is warranted with larger samples suffering mental health issues. The studies highlight the potential of social robots to act as guidance in unguided iCBT interventions and enhancing its effectiveness.
‘Anthropomorphism’, broadly understood as associating humanness with nonhuman entities, has been widely studied in the context of human interactions with social robots and computers. However, despite its centrality in human-machine interaction studies, the idea remains confusing and requires further clarification. This paper provides a critical conceptual review of how anthropomorphism is essentially characterized or conceptualized by scholars for robots and artificial intelligence (AI) systems, which will be increasingly incorporated into social robotics. We identify key dimensions of anthropomorphism that have not been sufficiently unpacked, namely: ‘object-related’ anthropomorphism; three types of ‘subject-related’ anthropomorphism; alleged attributive processes linking human-like features to machines; and normative views about whether such attributions constitute epistemic errors. We then discuss vagueness, ambiguity, and disagreement surrounding characterizations of anthropomorphism which may generate confusion and impede clear communication. In a constructive provocation, we ask if the term itself might be too compromised to remain useful, before offering recommendations for more precise and clearer terminology. By analyzing fundamental characterizations of anthropomorphism, this paper offers insights to scholars in robotics and AI, Human-Computer Interaction, cognitive science, and social sciences, aiming to support more coherent discussions and effective investigations.
The use of nursing robots to participate in nursing work is an effective means to solve the contradiction between the growing elderly population and the serious shortage of nursing staff, and robot assisted transfer is an important task in nursing work. In this study, a motion control method of manipulator is proposed, with the aim of helping robots achieve a comfortable and safe process of lifting and hugging people. Based on the characteristics of the human musculoskeletal system, a dynamic model of the human body is first established to analyze the balance of joint torque and contact force between human joints and manipulators, and propose a human tactile force model. Based on the analysis of human posture and joint torque, the constraint conditions for the human joint angle and joint torque of the tactile force model are provided. By solving the human tactile force model, the hug posture parameters of the manipulator can be obtained. Based on these hug parameters, the trajectory of the robotic arm is planned, and the end effector trajectory of the manipulator is optimized for end effector error. Then a motion control method for manipulators based on the optimal human tactile force is proposed, combining the kinematics and dynamics of the manipulator. Finally, a manipulator hug experiment was conducted, and the effectiveness of the proposed manipulator motion control method was verified through a combination of subjective evaluation and qualitative analysis.
As artificial intelligence (AI) systems increasingly operate in socially meaningful settings, it becomes important to understand how design choices such as physical embodiment shape user judgements. This mixed-methods study examines how people evaluate an embodied robot (Pepper) and a disembodied voice assistant (Alexa) in a competitive rock-paper-scissors task (N = 71). Rather than treating trust as a single unitary construct, we operationalise several task-relevant trust-related judgements (e.g., perceived fair play, comfort competing, reliance on strategic choices) alongside engagement and opponent-realism indicators. Across within-participant comparisons, Pepper was associated with higher enjoyment and stronger opponent realism, and was judged as less likely to be ‘cheating’. Alexa, however, was rated as more predictable and was preferred by some participants for reliance in a more complex game scenario. Exploratory analyses further suggest that embodiment-related cues (e.g., perceived influence of Pepper’s physical actions) are associated with individual differences in these judgements. Qualitative themes contextualise these findings, indicating that visible actions can be interpreted as transparency and ‘fairness’, while familiarity with voice assistants may support competence-oriented trust. Overall, the results underline that embodiment can shift which dimensions of trust become salient in competitive interaction, with implications for designing AI that is both engaging and appropriately trusted. Exploratory within-subject mediation analysis (difference-score formulation with bootstrap confidence intervals) further indicates that perceived realism/‘real-opponent’ judgements partially mediate Pepper’s advantage on trust and enjoyment, while a sensitivity analysis suggests that implausibly large order-based novelty effects would be required to overturn the main Pepper-Alexa differences.
Operators of teleoperated social robots often face challenges in adjusting their voice due to limited feedback about the remote situation. This makes it hard to determine a suitable speaking volume compared to face-to-face conversations. Such difficulties can lead to unclear communication and operator insecurity about their vocal delivery. A primary difficulty is estimating the appropriate voice volume based on the physical distance between the robot and the interlocutor. To tackle this, we introduce a straightforward interface improvement: a voice volume gauge that shows the difference between the operator’s current speaking volume and an ideal volume estimated by the system. We developed the gauge and conducted a controlled user study comparing scenarios with and without the gauge. The findings indicate that (1) the gauge was simple to use and reduced operator anxiety, (2) operator vocalizations became acoustically closer to those observed in face-to-face settings, and (3) interlocutors perceived the vocalizations as more natural when the gauge was used. This research offers a practical design to help operators adjust their voice in teleoperated robot systems, specifically by reducing the difference in sensory information between remote and in-person communication settings.