Intrinsic motivation is a common method to facilitate exploration in reinforcement learning agents. Curiosity is thereby supposed to aid the learning of a primary goal. However, indulging in curiosity may also stand in conflict with more urgent or essential objectives such as self-sustenance. This paper addresses the problem of balancing curiosity, and correctly prioritising other needs in a reinforcement learning context. We demonstrate the use of the multi-objective reinforcement learning framework C-MORE to integrate curiosity, and compare results to a standard linear reinforcement learning integration. Results clearly demonstrate that curiosity can be modelled with the priority-objective reinforcement learning paradigm. In particular, C-MORE is found to explore robustly while maintaining self-sustenance objectives, whereas the linear approach is found to over-explore and take unnecessary risks. The findings demonstrate a significant weakness of the common linear integration method for intrinsic motivation, and the need to acknowledge the potential conflicts between curiosity and other objectives in a multi-objective framework.
Intelligent agents often have to cope with situations in which their various needs must be prioritised. Efforts have been made, in the fields of cognitive robotics and machine learning, to model need prioritization. Examples of existing frameworks include normative decision theory, the subsumption architecture and reinforcement learning. Reinforcement learning algorithms oriented towards active goal prioritization include the options framework from hierarchical reinforcement learning and the ranking approach as well as the MORE framework from multi-objective reinforcement learning. Previous approaches can be configured to make an agent function optimally in individual environments, but cannot effectively model dynamic and efficient goal selection behaviour in a generalisable framework. Here, we propose an altered version of the MORE framework that includes a threshold constant in order to guide the agent towards making economic decisions in a broad range of priority-objective reinforcement learning’ (PORL) scenarios. The results of our experiments indicate that pre-existing frameworks such as the standard linear scalarization, the ranking approach and the options framework are unable to induce opportunistic objective optimisation in a diverse set of environments. In particular, they display strong dependency on the exact choice of reward values at design time. However, the modified MORE framework appears to deliver adequate performance in all cases tested. From the results of this study, we conclude that employing MORE along with integrated thresholds, can effectively simulate opportunistic objective prioritization in a wide variety of contexts.
We find ourselves at a unique point of time in history. Following over two millennia of debate amongst some of the greatest minds that ever existed about the nature of morality, the philosophy of ethics and the attributes of moral agency, and after all that time still not having reached consensus, we are coming to a point where artificial intelligence (AI) technology is enabling the creation of machines that will possess a convincing degree of moral competence. The existence of these machines will undoubtedly have an impact on this age old debate, but we believe that they will have a greater impact on society at large, as AI technology deepens its integration into the social fabric of our world. The purpose of this special issue on Computing Morality is to bring together different perspectives on this technology and its impact on society. The special issue contains four very different and inspiring contributions.
An early integration of tactile sensing into motor coordination is the norm in animals, but still a challenge for robots. Tactile exploration through touches on the body gives rise to first body models and bootstraps further development such as reaching competence. Reaching to one’s own body requires connections of the tactile and motor space only. Still, the problems of high dimensionality and motor redundancy persist. Through an embodied computational model for the learning of self-touch on a simulated humanoid robot with artificial sensitive skin, we demonstrate that this task can be achieved 1) effectively and 2) efficiently at scale by employing the computational frameworks for the learning of internal models for reaching: intrinsic motivation and goal babbling. We relate our results to infant studies on spontaneous body exploration as well as reaching to vibrotactile targets on the body. We analyze the reaching configurations of one infant followed weekly between 4 and 18 months of age and derive further requirements for the computational model: accounting for 3) continuous rather than sporadic touch and 4) consistent redundancy resolution. Results show the general success of the learning models in the touch domain, but also point out limitations in achieving fully continuous touch.
The mechanisms of infant development are far from understood. Learning about one's own body is likely a foundation for subsequent development. Here we look specifically at the problem of how spontaneous touches to the body in early infancy may give rise to first body models and bootstrap further development such as reaching competence. Unlike visually elicited reaching, reaching to own body requires connections of the tactile and motor space only, bypassing vision. Still, the problems of high dimensionality and redundancy of the motor system persist. In this work, we present an embodied computational model on a simulated humanoid robot with artificial sensitive skin on large areas of its body. The robot should autonomously develop the capacity to reach for every tactile sensor on its body. To do this efficiently, we employ the computational framework of intrinsic motivations and variants of goal babbling-as opposed to motor babbling-that prove to make the exploration process faster and alleviate the ill-posedness of learning inverse kinematics. Based on our results, we discuss the next steps in relation to infant studies: what information will be necessary to further ground this computational model in behavioral data.
Robots that cohabitate in social spaces must abide by the same behavioural cues humans follow, including interpersonal distancing. Proxemics investigates the appropriate distances and the impact of factors affecting it, such as gender and age. This paper investigates people's attitudes towards a robot that can learn Proxemics rules by gauging direct individual feedback from a person, and utilizing it in a reinforcement learning framework. Previous learning attempts have relied on larger robots, for which physical safety is a primary concern. In contrast, our study uses a handheld sized robot that allows us to focus on the impact of distance on engageability in dialogue. General consensus between interviewees was a feeling of ease and safety during interactions, as well as disparity regarding the invasion of personal space, which was influenced by cultural background.
Both biological and artificial agents need to coordinate their behavior to suit various needs at the same time. Reconciling conflicts of different needs and contradictory interests such as self-preservation and curiosity is the central difficulty arising in the design and modelling of need and value systems. Current models of multi-objective reinforcement learning do either not provide satisfactory power to describe such conflicts, or lack the power to actually resolve them. This paper aims to promote a clear understanding of these limitations, and to overcome them with a theory-driven approach rather than ad hoc solutions. The first contribution of this paper is the development of an example that demonstrates previous approaches' limitations concisely. The second contribution is a new, non-linear objective function design, MORE, that addresses these and leads to a practical algorithm. Experiments show that standard RL methods fail to grasp the nature of the problem and ad-hoc solutions struggle to describe consistent preferences. MORE consistently learns a highly satisfactory solution that balances contradictory needs based on a consistent notion of optimality.
Robotic tele-operation systems have vast potential in areas ranging from surgical robotics and underwater exploration to disposing of toxic, explosive and nuclear materials. While visual camera feeds for the human operator are typically available and well studied, tactile sensory information is often vital for successful and efficient manipulation. Previous studies have largely focused on execution time alone as measure of success of feedback methods on individual tasks. The present study complements this by a comparative analysis of vibration and visual feedback of tactile information across a range of manipulation tasks. Results show a significant reduction in perceived workload with the implementation of vibration feedback and an improvement of error rates for visual feedback. Contrary to expectation, we did not find a reduction in task completion time. The negative finding on completion time challenges the belief that the mere existence of task-relevant feedback aids efficient task completion. The reduced workload, however, clearly points out potential for enhancing performance on more difficult and prolonged tasks with highly skilled operators.
Exploration is one of the fundamental problems in mobile robotics. Efforts to address this problem made over the past two decades divide into two approaches: reactive approaches, that make only instantaneous decisions, and map-based approaches involving e.g. grid, metric, or topological representations. Comparative studies have so far largely focused on comparing different map-based algorithms, while no common framework to compare them to purely reactive approaches currently exists. This paper aims at creating a framework to simulate, evaluate, and compare exploratory algorithms as different as reactive and map-based approaches. Preliminary results are demonstrated for two reactive algorithms, random walk and wall follower, and one map based approach, pheromone potential field, have been implemented. Measurements of navigation success, time to success, as well as computational and memory usage reveal a dominance of simple wall-following over the map-based potential field approach, and a distinct load/efficacy trade off for random walks. These preliminary results challenge the common assumptions that maps are needed for successful and efficient exploration and navigation.
AI and robot ethics have recently gained a lot of attention because adaptive machines are increasingly involved in ethically sensitive scenarios and cause incidents of public outcry. Much of the debate has been focused on achieving highest moral standards in handling ethical dilemmas on which not even humans can agree, which indicates that the wrong questions are being asked. We suggest to address this ethics debate strictly through the lens of what behavior seems socially acceptable, rather than idealistically ethical. Learning such behavior puts the debate into the very heart of developmental robotics. This paper poses a roadmap of computational and experimental questions to address the development of socially acceptable machines. We emphasize the need for social reward mechanisms and learning architectures that integrate these while reaching beyond limitations of plain reinforcement-learning agents. We suggest to use the metaphor of “needs” to bridge rewards and higher level abstractions such as goals for both communication and action generation in a social context. We then suggest a series of experimental questions and possible platforms and paradigms to guide future research in the area.
This paper presents a study of the movements of a humanoid head-and-neck robot called Eddie. Eddie has a musculo-skeletal structure similar to that found in human necks enabling it to perform head movements that are comparable with human head movements. This study compares the movements of Eddie with those of a more conventional robotic neck structure and with those of a human head. Results show that Eddie's movements are perceived as significantly more natural and by trend more lifelike than the conventional head's. No differences were found with respect to the impression of human-likeness, consciousness, and elegance.
Future personal robots might possess the capability to autonomously generate novel goals that exceed their initial programming as well as their past experience. We discuss the ethical challenges involved in such a scenario, ranging from the construction of ethics into such machines to the standard of ethics we could actually demand from such machines. We argue that we might have to accept those machines committing human-like ethical failures if they should ever reach human-level autonomy and intentionality. We base our discussion on recent ideas that novel goals could be originated from agents’ value system that express a subjective goodness of world or internal states. Novel goals could then be generated by extrapolating what future states would be good to achieve. Ethics could be built into such systems not just by simple utilitarian measures but also by constructing a value for the expected social acceptance of a the agent’s conduct.
The bionic handling assistant is one of the largest soft continuum robots and very special in being a pneumatically operated platform that is able to bend, stretch, and grasp in all directions. It nevertheless shares many challenges with smaller continuum and other soft robots such as parallel actuation, complex movement dynamics, slow pneumatic actuation, non-stationary behavior, and a lack of analytic models. To master the control of this challenging robot, we argue for a tight integration of standard analytic tools, simulation, control, and state-of-the-art machine learning into an overall architecture that can serve as blueprint for control design also beyond the BHA. To this aim, we show how to integrate specific modes of operation and different levels of control in a synergistic manner, which is enabled by using modern paradigms of software architecture and middleware. We thereby achieve an architecture with unique overall control abilities for a soft continuum robot that allow for flexible experimentation toward compliant user-interaction, grasping, and online learning of internal models.
Goals are concepts used in many different areas of robotics, artificial intelligence, psychology, neuroscience, and also philosophy. Despite the wide usage, there is no common definition of a "goal". Rather, the term is used in substantially different ways even within disciplines. This paper discusses these notions and potentially unified views on goals, and points out how different perspectives on the same term lead to different arguments and can cause communication difficulties in the interdisciplinary community. We discuss how far goal terminologies can be generally considered as desired end states of action and point out the pivotal aspect of their explicit representation. As a major point we discuss the relation of such goals with reward and value systems from various perspectives.
Reinforcement learning is a paradigm that is both very general and widely applied for interacting agents. Despite tremendous progress on both model-based and model-free algorithms, reinforcement learning does however still requires a substantial amount of manual task design. One of the major burdens for a truly autonomous operation of RL agents is the design of task-appropriate features [Kober and Peters, 2012] of state and action. These features need to be comprehensive for RL to perform effectively, yet compact in terms of dimension to perform efficiently.
Goals are abstractions of high-dimensional world states that express intelligent agents’ intentions underlying their actions. Goals are considered to organize the behavior of both humans and robots. For instance in robot planning as well as motor control goals describe the desired outcome of future actions. Goals are a fundamental concept also in neuroscience and psychology, e.g. in formulations of internal models [1], motivation psychology [2], or teleological action understanding [3]. We recently argued [4] that the achievement semantics of goals point out an immediate need for an evaluation of the own action’s effect (see Fig. 1). In hand-eye coordination this evaluation, or rather its learning, is often referred to as self-detection [5] or body schema [6]: the hand needs to be localized e.g. from vision data. Goals are only useful when this “ground-truth” position is available. Indeed, there very relation allows for versatile motor control as well as selfsupervised motor learning. Due to the vital relation between both, we argue to learn them within a joint framework. Yet, how could an agent learn such goal systems in which goals, body-schema, and their relation are identified?
Current approaches to artificial attention are largely limited to the visual domain. Only some consider audition as a source of information at the same time. Yet, attention is not necessarily limited to a single modality or a mere agglomeration of several modalities in human perception. Cross-modal attention, and its manipulation by cross-modal cues, seems to play a vital role in asymmetric interactions such as a parent tutoring a child. We discuss previous efforts [23] to reflect such perceptual processes with an artificial attention system that considers signal-level synchrony between vision and audition to guide visual attention. Results show that the system is receptive to infant directed cues from parents.
Goals express agents' intentions and allow them to organize their behavior based on low-dimensional abstractions of high-dimensional world states. How can agents develop such goals autonomously? This paper proposes a detailed conceptual and computational account to this longstanding problem. We argue to consider goals as high-level abstractions of lower-level intention mechanisms such as rewards and values, and point out that goals need to be considered alongside with a detection of the own actions' effects. We propose Latent Goal Analysis as a computational learning formulation thereof, and show constructively that any reward or value function can by explained by goals and such self-detection as latent mechanisms. We first show that learned goals provide a highly effective dimensionality reduction in a practical reinforcement learning problem. Then, we investigate a developmental scenario in which entirely task-unspecific rewards induced by visual saliency lead to self and goal representations that constitute goal-directed reaching.
Bionic soft robots offer exciting perspectives for more flexible and safe physical interaction with the world and humans. Unfortunately, their hardware design often prevents analytical modeling, which in turn is a prerequisite to apply classical automatic control approaches. On the other hand, also modeling by means of learning is hardly feasible due to many degrees of freedom, high-dimensional state spaces and the softness properties like e.g. mechanical elasticity, which cause limited repeatability and complex dynamics. Nevertheless, the realization of basic control modes is important to leverage the potential of soft robots for applications. We therefore propose a hybrid approach combining classical and learning elements for the realization of an interactive control mode for an elastic bionic robot. It superimposes a low-gain feedback control with a feed-forward control based on a learned simplified model of the inverse dynamics which considers only equilibria of the robot's dynamics. We demonstrate on the Bionic Handling Assistant how a respective inverse equilibrium model can be learned and effectively exploited for quick and agile control. In a second step, the control scheme is extended to an active compliant control mode. It implements a kind of gravitation compensation to allow for kinesthetic teaching of the robot based on the implicit knowledge of gravitational and mechanical forces that are encoded in the learned equilibrium model. We finally discuss that this control scheme may be implemented also on other soft robots to provide the avenue towards their applications in general manipulation tasks.
Goals are abstractions that express agents' intention and allow them to organize their behavior appropriately. How can agents develop such goals autonomously? This paper proposes a conceptual and computational account to this longstanding problem. We argue to consider goals as abstractions of lower-level intention mechanisms such as rewards and values, and point out that goals need to be considered alongside with a detection of the own actions' effects. Then, both goals and self-detection can be learned from generic rewards. We show experimentally that task-unspecific rewards induced by visual saliency lead to self and goal representations that constitute goal-directed reaching.