This paper describes the development of algorithms that decide when to move, where to move, and how to look for people in a home environment. We introduce a design framework as a tool to guide the development of a social robot to proactively be with people for companionship and assistance in the home. Through a series of experiments ranging from simulations to longitudinal A/B studies, we demonstrate how to utilize the design framework to help guide the evaluation and selection of solutions. We deployed our autonomous robot in a long-term in-situ study and found our proposed approach to be more capable of being co-present with its household members compared to a baseline approach. Conducted in an industry setting, our research approach departs from typical academic practices as the motivations are inherently different. We share our perspective on the differences of industry research when developing a social robot as a commercial product.
This paper defines a dual computational framework to nonverbal communication for human-robot interactions. We use a Bayesian Theory of Mind approach to model dyadic storytelling interactions where the storyteller and the listener have distinct roles. The role of storytellers is to influence and infer the attentive state of listeners using speaker cues, and we computationally model this as a POMDP planning problem. The role of listeners is to convey attentiveness by influencing perceptions through listener responses, which we computational model as a DBN with a myopic policy. Through a comparison of state estimators trained on human-human interaction data, we validate our storyteller model by demonstrating how it outperforms current approaches to attention recognition. Then through a human-subjects experiment where children told stories to robots, we demonstrate that a social robot using our listener model more effectively communicates attention compared to alternative approaches based on signaling.
Understanding social-emotional behaviors in storytelling interactions plays a critical role in the development of interactive educational technologies for children. A challenge when designing for such interactions using technology like social robots, virtual agents, and tablets is understanding the social-emotional behaviors pertinent to storytelling especially when emulating a natural peer-to-peer relation between the child and the technology. We present P2PSTORY, a dataset of young children (5-6 years old) engaging in natural peer-to-peer storytelling interactions with fellow classmates. The dataset consists of rich social behaviors of children without adult supervision, with each participant demonstrating being a storyteller and a listener. The dataset contains 58 video recorded sessions along with a diverse set of behavioral annotations as well as developmental and demographic profiles of each child participant. We describe the main characteristics of the dataset in addition to findings that reveal perceptual differences between adults and children when evaluating the attentiveness of listeners.
Current state-of-the-art approaches to emotion recognition primarily focus on modeling the nonverbal expressions of the sole individual without reference to contextual elements like the co-presence of the partner. In this paper, we demonstrate that the accurate inference of listeners' social-emotional state of attention depends on also accounting for the nonverbal behaviors of their storytelling partner, namely their speaker cues. To gain a deeper understanding of the role of speaker cues in attention inference, we conduct investigations into real world interactions of children storytelling with their peers. Through in-depth analysis of human-human interaction data, we first identify nonverbal speaker cues (i.e., backchannel-inviting cues) and listener responses (i.e., backchannel feedback) to later demonstrate how speaker cues can modify the interpretation of attention-related backchannels as well as serve as a means to regulate the responsiveness of listeners. We then discuss the design implications of our findings toward our primary goal of developing attention recognition models for storytelling robots. Social robots can use speaker cues to form more accurate inferences about the attentive state of their human partner.
This paper investigates how a robot that can produce contingent listener response, i.e., backchannel, can deeply engage children as a storyteller. We propose a backchannel opportunity prediction (BOP) model trained from a dataset of children's dyad storytelling and listening activities. Using this dataset, we gain better understanding of what speaker cues children can decode to find backchannel timing, and what type of nonverbal behaviors they produce to indicate engagement status as a listener. Applying our BOP model, we conducted two studies, within- and between-subjects, using our social robot platform, Tega. Behavioral and self-reported analyses from the two studies consistently suggest that children are more engaged with a contingent backchanneling robot listener. Children perceived the contingent robot as more attentive and more interested in their story compared to a non-contingent robot. We find that children significantly gaze more at the contingent robot while storytelling and speak more with higher energy to a contingent robot.
In this video, we provide an overview of the analyses, design, and evaluation of a backchannel opportunity prediction (BOP) model for a social robot listener.
While there has been a growing body of work in child-robot interaction, we still have very little knowledge regarding young children's speaking and listening dynamics and how a robot companion should decode these behaviors and encode its own in a way children can understand. In developing a backchannel prediction model based on observed nonverbal behaviors of 4-6 year-old children, we investigate the effects of an attentive listening robot on a child's storytelling. We provide an extensive analysis of young children's nonverbal behavior with respect to how they encode and decode listener responses and speaker cues. Through a collected video corpus of peer-to-peer storytelling interactions, we identify attention-related listener behaviors as well as speaker cues that prompt opportunities for listener backchannels. Based on our findings, we developed a backchannel opportunity prediction (BOP) model that detects four main speaker cue events based on prosodic features in a child's speech. This rule-based model is capable of accurately predicting backchanneling opportunities in our corpora. We further evaluate this model in a human-subjects experiment where children told stories to an audience of two robots, each with a different backchanneling strategy. We find that our BOP model produces contingent backchannel responses that conveys an increased perception of an attentive listener, and children prefer telling stories to the BOP model robot.
Tega is a new expressive “squash and stretch”, Android-based social robot platform, designed to enable long-term interactions with children.
We deployed an autonomous social robotic learning companion in three preschool classrooms at an American public school for two months. Before and after this deployment, we asked the teachers and teaching assistants who worked in the classrooms about their views on the use of social robots in preschool education. We found that teachers' expectations about the experience of having a robot in their classrooms often did not match up with their actual experience. These teachers generally expected the robot to be disruptive, but found that it was not, and furthermore, had numerous positive ideas about the robot's potential as a new educational tool for their classrooms. Based on these interviews, we provide a summary of lessons we learned about running child-robot interaction studies in preschools. We share some advice for future researchers who may wish to engage teachers and schools in the course of their own human-robot interaction work. Understanding the teachers, the classroom environment, and the constraints involved is especially important for microgenetic and longitudinal studies, which require more of the school's time-as well as more of the researchers' time-and is a greater opportunity investment for everyone involved.
Socially assistive robots have been used successfully to attract childrens attention, stimulate sustainable interactions and improve communication skills in young children. Many recent research studies propose various approaches for increasing the interaction effectiveness of such robots. Our research objective is to develop an autonomous social robot learning companion that, through contingent backchannel feedback, can successfully foster the development of early language skills of preschoolers over long-term interaction in an educational storytelling context. In this study we used Tega and we investigated the effect of the robot nonverbal behaviors during storytelling activities of young children. The findings of this study contribute to the effective use of interactive robots in education. Our results suggest we can build models that capture individual differences in backchannel style, captivate children’s attention and enhance the whole storytelling and learning experience. In the near future we will possibly identify individual traits from observations of backchannel behavior and create autonomous and personalized activities so as to increase the children’s curriculum in a social and engaging venue.
Though substantial research has been dedicated towards using technology to improve education, no current methods are as effective as one-on-one tutoring. A critical, though relatively understudied, aspect of effective tutoring is modulating the student's affective state throughout the tutoring session in order to maximize long-term learning gains. We developed an integrated experimental paradigm in which children play a second-language learning game on a tablet, in collaboration with a fully autonomous social robotic learning companion. As part of the system, we measured children's valence and engagement via an automatic facial expression analysis system. These signals were combined into a reward signal that fed into the robot's affective reinforcement learning algorithm. Over several sessions, the robot played the game and personalized its motivational strategies (using verbal and non-verbal actions) to each student. We evaluated this system with 34 children in preschool classrooms for a duration of two months. We saw that (1) children learned new words from the repeated tutoring sessions, (2) the affective policy personalized to students over the duration of the study, and (3) students who interacted with a robot that personalized its affective feedback strategy showed a significant increase in valence, as compared to students who interacted with a non-personalizing robot. This integrated system of tablet-based educational content, affective sensing, affective policy learning, and an autonomous social robot holds great promise for a more comprehensive approach to personalized tutoring.
We created a socially assistive robotic learning companion to support English-speaking children’s acquisition of a new language (Spanish). In a two-month microgenetic study, 34 preschool children will play an interactive game with a fully autonomous robot and the robot’s virtual sidekick, a Toucan shown on a tablet screen. Two aspects of the interaction were personalized to each child: (1) the content of the game (i.e., which words were presented), and (2) the robot’s affective responses to the child’s emotional state and performance. We will evaluate whether personalization leads to greater engagement and learning.
This paper describes an extended (6-session) interaction between an ethnically and geographically diverse group of 26 first-grade children and the DragonBot robot in the context of learning about healthy food choices. We find that children demonstrate a high level of enjoyment when interacting with the robot, and a statistically significant increase in engagement with the system over the duration of the interaction. We also find evidence of relationship-building between the child and robot, and encouraging trends towards child learning. These results are promising for the use of socially assistive robotic technologies for long-term one-on-one educational interventions for younger children.
By designing socially intelligent robots that can more effectively communicate and interact with us, we can increase their capacity to function as collaborative partners. Our research goal is to develop robots capable of engaging in nonverbal communication, which has been argued to be at the core of social intelligence. We take a human-centric approach that closely aligns with how people are theorized to model nonverbal communication. We propose a unified computational approach to interactively learn the meaning of nonverbal behaviors for inference and production. More specifically, we use a partially observable Markov decision process model to infer an interactant’s mental state based on their observed nonverbal behaviors as well as produce appropriate nonverbal behaviors for a robot to communicate its internal state. By interactively learning the connection between nonverbal behaviors and the mental states producing them, an agent can more readily generalize to new people and situations.
We present a computational model capable of predicting—above human accuracy—the degree of trust a person has toward their novel partner by observing the trust-related nonverbal cues expressed in their social interaction. We summarize our prior work, in which we identify nonverbal cues that signal untrustworthy behavior and also demonstrate the human mind's readiness to interpret those cues to assess the trustworthiness of a social robot. We demonstrate that domain knowledge gained from our prior work using human-subjects experiments, when incorporated into the feature engineering process, permits a computational model to outperform both human predictions and a baseline model built in naiveté of this domain knowledge. We then present the construction of hidden Markov models to investigate temporal relationships among the trust-related nonverbal cues. By interpreting the resulting learned structure, we observe that models built to emulate different levels of trust exhibit different sequences of nonverbal cues. From this observation, we derived sequence-based temporal features that further improve the accuracy of our computational model. Our multi-step research process presented in this paper combines the strength of experimental manipulation and machine learning to not only design a computational trust model but also to further our understanding of the dynamics of interpersonal trust.
People are increasingly working with robots in teams and recent research has focused on how human-robot teams function, but little attention has yet been paid to the role of social signaling behavior in human-robot teams. In a controlled experiment, we examined the role of backchanneling and task complexity on team functioning and perceptions of the robots' engagement and competence. Based on results from 73 participants interacting with autonomous humanoid robots as part of a human-robot team (one participant, one confederate, and three robots), we found that when robots used backchanneling team functioning improved and the robots were seen as more engaged. Ironically, the robots using backchanneling were perceived as less competent than those that did not. Our results suggest that backchanneling plays an important role in human-robot teams and that the design and implementation of robots for human-robot teams may be more effective if backchanneling capability is provided.
We describe research towards creating a computational model for recognizing interpersonal trust in social interactions. We found that four negative gestural cues— leaning-backward, face-touching, hand-touching, and crossing-arms—are together predictive of lower levels of trust. Three positive gestural cues—leaning- forward, having arms-in-lap, and open-arms—are predictive of higher levels of trust. We train a probabilistic graphical model using natural social interaction data, a “Trust Hidden Markov Model” that incorporates the occurrence of these seven important gestures throughout the social interaction. This Trust HMM predicts with 69.44% accuracy whether an individual is willing to behave cooperatively or uncooperatively with their novel partner; in comparison, a gesture-ignorant model achieves 63.89% accuracy. We attempt to automate this recognition process by detecting those trust-related behaviors through 3D motion capture technology and gesture recognition algorithms. We aim to eventually create a hierarchical system—with low-level gesture recognition for high-level trust recognition—that is capable of predicting whether an individual finds another to be a trustworthy or untrustworthy partner through their non- verbal expressions.
Because trusting strangers can entail high risk, an ability to infer a potential partner’s trustworthiness would be highly advantageous. To date, however, little evidence indicates that humans are able to accurately assess the cooperative intentions of novel partners by using nonverbal signals. In two studies involving human-human and human-robot interactions, we found that accuracy in judging the trustworthiness of novel partners is heightened through exposure to nonverbal cues and identified a specific set of cues that are predictive of economic behavior. Employing the precision offered by robotics technology to model and control humanlike movements, we demonstrated not only that experimental manipulation of the identified cues directly affects perceptions of trustworthiness and subsequent exchange behavior, but also that the human mind will utilize such cues to ascribe social intentions to technological entities.
Being able to establish successful human-robot teamwork in ambiguous, real life situations requires the ability for robots to dynamically engage in collaborative interactions whenever the need for collaboration arises. Understanding what distinguishes successful from less successful human-robot interaction patterns is therefore crucial for the future development of social robotic applications. This paper explores the social and task oriented dynamics occurring during collaboration with an autonomous robot and demonstrates that a crowd-sourced model can drive not only task completion but also social interactions successfully. We propose that bids and subtle cues signal interaction attempts which may play an important role in facilitating the establishment and sustenance of collaborative human robot interactions. We demonstrate the applicability of bids focused interaction perspective in three exploratory cases by modifying and applying a coding scheme that had originally been developed for the analysis of bids in marital interactions. Implications for further research and design of robots capable of teamwork are discussed.
Introduction Seymour Papert in Mindstorms highlights the powerful potential technology, especially computers, has in education. He envisions a future where the computer's role goes beyond just " computer-aided instruction, " and instead it becomes an instrument that helps people think and learn creatively. Rather than for computers to " program the child, " the child should take on an active role and " program the computer " [1]. We present a technology called the Huggable, an embodied computer, that can encourage children to read aloud (see Figure 5). The Huggable is an autonomous robotic teddy bear capable of fun and dynamic interactions with its reader. As a child reads a story out-loud, the Huggable follows along and listens and occasionally looks and nods at the reader and at the pictures in the book. The Huggable can also understand the content of the story and emotionally reacts to the different events and scenes in the book. Through these rich and meaningful interactions, we hope to keep a child engaged as he/she reads the book aloud. By Reading with Robots, we hope to make reading a fun and positive experience for all kids. And by associating reading as a fun interactive activity, children will genuinely want to comeback and read again to the Huggable. And by practicing their oral fluency, children can improve their reading comprehension and other literacy skills [7]. Robots are already having a positive impact in classroom settings. Robots like LEGO Mindstorms are helping children effectively learn about science, technology, and mathematics [16]. They even encourage children to work together in teams, helping them develop social skills in the process. Robots also foster creative imagination and innovation through the making of robot designs and computer algorithms. Classrooms in Japan have introduced an interactive robot tutor called Robovie that speaks English to the Japanese children. In just 2 weeks, researchers found that the classroom's English skills had significantly improved [15]. Robots have a power potential in education. And as we step into this new frontier of robots and education, we consider principles from new learning sciences, principles of human-robot-interaction, designs of previous similar works, and the latest advancements in research technology to drive and shape our design of Reading with Robots, a fun activity where children can read aloud to robots that listen and keep the reader engaged throughout the book.