One of the earliest forms of interaction between mothers and infants is smiling games. While the temporal dynamics of these games have been extensively studied, they are still not well understood. Why do mothers and infants time their smiles the way they do? To answer this question we applied methods from control theory, an approach frequently used in robotics, to analyze and synthesize goal-oriented behavior. The results of our analysis show that by the time infants reach 4 months of age both mothers and infants time their smiles in a purposeful, goal-oriented manner. In our study, mothers consistently attempted to maximize the time spent in mutual smiling, while infants tried to maximize mother-only smile time. To validate this finding, we ported the smile timing strategy used by infants to a sophisticated child-like robot that automatically perceived and produced smiles while interacting with adults. As predicted, this strategy proved successful at maximizing adult-only smile time. The results indicate that by 4 months of age infants interact with their mothers in a goal-oriented manner, utilizing a sophisticated understanding of timing in social interactions. Our work suggests that control theory is a promising technique for both analyzing complex interactive behavior and providing new insights into the development of social communication.
Our aim is to demonstrate the usefulness of photoplethysmography (PPG) for analyzing heart rate variability (HRV) using a standard 5-min test at rest with paced breathing, comparing the results with real RR intervals and testing supine and sitting positions. Simultaneous recordings of R-R intervals were conducted with a Polar system and a non-contact PPG, based on facial video recording on 20 individuals. Data analysis and editing were performed with individually designated software for each instrument. Agreement on HRV parameters was assessed with concordance correlations, effect size from ANOVA and Bland and Altman plots. For supine position, differences between video and Polar systems showed a small effect size in most HRV parameters. For sitting position, these differences showed a moderate effect size in most HRV parameters. A new procedure, based on the pixels that contained more heart beat information, is proposed for improving the signal-to-noise ratio in the PPG video signal. Results were acceptable in both positions but better in the supine position. Our approach could be relevant for applications that require monitoring of stress or cardio-respiratory health, such as effort/recuperation states in sports.
L'invention concerne un appareil, des procedes et des articles fabriques pour mettre en œuvre des pipelines d'externalisation ouverte qui generent des exemples de formation pour des classificateurs d'expression d'apprentissage machine. Des fournisseurs d'externalisation ouverte generent activement des images avec des expressions, selon des signaux ou des objectifs. Les signaux ou objectifs peuvent etre destines a imiter une expression ou a apparaitre d'une certaine maniere, ou a « casser » un element de reconnaissance d'expression existant. Les images sont recoltees et classees par les memes fournisseurs d'externalisation ouverte ou des fournisseurs d'externalisation ouverte differents, et les images qui satisfont un premier critere de qualite sont ensuite examinees minutieusement par un ou par plusieurs experts. Les images examinees minutieusement sont ensuite utilisees comme exemples positifs ou negatifs dans la formation de classificateurs d'expression d'apprentissage machine.
Student engagement is a key concept in contemporary education, where it is valued as a goal in its own right. In this paper we explore approaches for automatic recognition of engagement from students' facial expressions. We studied whether human observers can reliably judge engagement from the face; analyzed the signals observers use to make these judgments; and automated the process using machine learning. We found that human observers reliably agree when discriminating low versus high degrees of engagement (Cohen's κ = 0.96). When fine discrimination is required (four distinct levels) the reliability decreases, but is still quite high ( κ = 0.56). Furthermore, we found that engagement labels of 10-second video clips can be reliably predicted from the average labels of their constituent frames (Pearson r=0.85), suggesting that static expressions contain the bulk of the information used by observers. We used machine learning to develop automatic engagement detectors and found that for binary classification (e.g., high engagement versus low engagement), automated engagement detectors perform with comparable accuracy to humans. Finally, we show that both human and automatic engagement judgments correlate with task performance. In our experiment, student post-test performance was predicted with comparable accuracy from engagement labels ( r=0.47) as from pre-test scores ( r=0.44).
To deal with the question of what a sociable robot is, we describe how an educational robot is encountered by children, teachers and designers in a preschool. We consider the importance of the robot’s body by focusing on how its movements are contingently embedded in interactional situations. We point out that the effects of agency that these movements generate are inseparable from their grounding in locally coordinated, multimodal actions and interactions.
Biologically inspired humanoid robots present new challenges for system identification and control due to the presence of many degrees of freedom, highly compliant actuators, and non-traditional force transmission mechanisms. In this thesis, we address these challenges using machine learning approaches. The key idea is to replace classical laborious manual model calibration and motion programming with statistical inference and learning from multi-modal sensory data. To this end, we develop several new parametric models and their parameter identification algorithms enabling new sensor/ actuator configurations beyond the scope of previous approaches. In addition, we also develop a semi-parametric model to learn from experiences not predicted by the parametric model. Using similar approaches grounded in machine learning, we also develop methods to allow humanoid robots to learn to make facial expressions, kick a ball, and to reach for objects while collaborating with people. We collected a unique dataset that describes development of infant reaching behavior while interacting with an adult caregiver. We compared the observed development of social reaching in human infants with the machine learning based development behavior in a complex humanoid robot.
Crowdsourcing services such as the Amazon Mechanical Turk [1] are increasingly being used to annotate large datasets for machine learning and data mining applications. The crowdsourced data labels must then be somehow combined to form a final judgment on the “true” label value of each data instance. Recently developed algorithms for combining multiple opinions [13, 12, 17, 4, 9, 15] have shown promising results, but they are lacking in several ways: (1) Each labeler is typically treated independently. In practice, labelers may share much in common, and their labeling proficiency may vary as a function of easily-queried attributes such as age and geographical location. Such commonality both among data labelers and among data instances could be better exploited. (2) Interaction effects between labelers and data and the reliability of a given label are ignored. For instance, some labelers may be good at labeling particular kinds of data but not others. In this paper, we present a probabilistic model for combining crowdsourced opinions that discovers and exploits both commonality and interaction effects automatically using a latent product-of-factors approach. We evaluate our proposed method on a facial expression labeling task and on a geography knowledge test. Empirical results show increased accuracy of the label estimation compared to competing methods. In addition, the model revealed interesting and useful trends relating the labeler and data features to the probability of correct labeling.
In this chapter we define the problem space and describe the core components of automatic facial expression recognition systems. In particular, we discuss the most prominent methods of face segmentation, face registration, feature extraction, classification, and temporal integration. We then present several practical applications of expression recognition technology. Finally, we comment on the core future challenges to the field, including generalization to multiple poses and ethnicities, collection of better training data, evaluation infrastructure, and commercialization.
To interact with objects effectively, a robot can use model-based or model-free control approaches. The superior performance typical of model-based control comes at the cost of developing or learning an accurate model of the system to be controlled. In this paper, we suggest an approach that generates models for novel objects based on visual features of those objects. These models can then be used for anticipatory control. We demonstrate this approach by replicating an infant experiment on a pneumatic humanoid robot. Infants seem to use visual information to estimate the mass of rods, and when they are presented a rod with an unexpected length-to-mass relationship, infants produce a large overcompensating arm movement when compared to an object with an expected mass. Our replication shows that the visual model-based control approach qualitatively replicates the behavior observed in the infant experiment, whereas a popular model-free approach, PID control, does not.
Handling Digital Brains proves that ethnography of the laboratory is still capable of making a significant contribution in the field of social studies on science and technology. The reviewed work presents details of interactions between researchers, as well as between researchers and their material equipment, which are key to explaining the methods of solving research problems when analyzing brain scans generated during fMRI experiments. Significantly, the reconstructed multimodal embodied practices shed light not only on the process of scientific cognition, but also on a broader spectrum of human cognitive activities. The book constitutes a challenge of a kind to neurocognitive sciences. As the author shows, cognitive neuroscientists utilizing fMRI declare that they study the embodied mind; yet, in practice, they reduce the body to the brain, and cognition - to purely internal processes. Such a model of cognition, ( tacitly) assumed by experimental neurocognitive scientists, turns out to be insufficient when used reflexively in order to explain the way neuroscientists themselves solve problems.
Badacze reprezentujący robotykę społeczną projektują swoje roboty tak, aby funkcjonowały one jako społeczni agenci w interakcji z ludźmi oraz z innymi robotami. Jakkolwiek nie przeczymy, że fizyczne cechy robota oraz jego oprogramowanie są istotne dla osiągniecia tego celu, pragniemy zwrócić uwagę na znaczenie organizacji przestrzennej oraz procesów koordynacji interakcji robota z ludźmi. Interakcje te badaliśmy, prowadząc obserwacje w [„rozszerzonym”] laboratorium robotyki społecznej. W tekście dokonujemy multimodalnej analizy interakcyjnej dwóch momentów praktyki projektantów robotów społecznych. Opisujemy kluczową rolę samych robotyków oraz grupy małych dzieci nieposługujących się jeszcze językiem, które zaangażowano w proces projektowania robota. Twierdzimy tu, że społeczny charakter projektowanej maszyny jest w istotny sposób powiązany z subtelnością ludzkich zachowań w laboratorium. To ludzkie zaangażowanie w proces tworzenia społecznego sprawstwa robota nie jest kwestią woli indywidualnych osób. Raczej jest tak, że dopasowania maszyn i ludzi wymaga dynamika sytuacyjna, w której osadzony jest robot.
Faced with the challenge of learning about its environment, an important first step for a robot is deciding how to direct its sensors to extract meaningful information. Computational models of visual salience have been developed to predict where humans tend to look in visual scenes, and thus they may provide useful heuristics for orienting robotic cameras. However, as we show here, current models of visual salience exhibit some important problems when applied to active robotic cameras. Here, we describe a new model of visual salience, named EMI, specifically designed to work on robotic cameras. The intuition behind this model is that, regardless of the task at hand, it is critical for robot cameras to keep track of motion, be it caused by camera movement or by world movement. Thus, it is reasonable for cameras to focus on image regions that are expected to provide high information about future motion, and to do so in a way that blurs the image as little as possible. We show that EMI predicts human fixations at least as well as current models of visual salience. In addition, we show that EMI overcomes the limitations that current visual salience models have when applied to robotic, active cameras.
The RUBI project [2,3,4,5,6,7,8,9] started in 2004 with the goal of designing sociable robots for early childhood education. The project relies on an immersive, iterative approach to robot development and scientific progress. From the beginning of the project scientists and engineers immersed themselves at the Early Childhood Education Center at UC San Diego, designed robot prototypes to interact and teach toddlers in close collaboration with teachers, toddlers, and parents. As part of this process we learned many lessons on how to develop hardware and software components, that are safe, reliable, and that can withstand the rigors of daily interactions with toddlers. Here we present our last robot platforms: RUBI-5. The robot components are off-the-shelf and the 3D-CAD files are being made available for free to the research community to ease replication and adaptation by other groups. One of the goals of the new platform is to greatly accelerate the design, data-gathering, and statistical analysis of early education experiments.
Heart Rate Variability (HRV) is an indicator of health status in the general population and of adaptation to stress in athletes. In this paper we compare the performance of two systems to measure HRV: (1) A commercial system based on recording the physiological cardiac signal with (2) A computer vision system that uses a standard video images of the face to estimate RR from changes in skin color of the face. We show that the computer vision system performs surprisingly well. It estimates individual RR intervals in a non-invasive manner and with error levels comparable to those achieved by the physiological based system.
From a Bayesian point of view, learning is simply the process of making inferences about the world based on incoming data. The efficiency of this learning is determined by the quality of the information provided by the sensors. Thus, a critical part of learning is the existence of a sensory-motor system designed to maximize the information required to achieve goals. Here we show that a wide range of primate eye movement phenomena can be elegantly explained from the point of view of infomax control. The proposed approach describes the velocity profiles of saccadic eye movements as well as previously existing models. In addition, the infomax approach explains phenomena that are beyond the scope of previous models: non-saccadic eye movements, and the difference in end point and velocity profiles observed in saccade-to-target and reach-to-target tasks. The results suggest that the occulomotor control system evolved to be a very efficient real time learning machine.
Gwen Littlewort合作论文数Machine Perception Laboratory;;Institute for Neural Computation24
Garrison W. Cottrell合作论文数Computer Science & Engineering Department, University of California, San Diego4