Building artificial intelligence (AI) that aligns with human values is an unsolved problem. Here we developed a human-in-the-loop research pipeline called Democratic AI, in which reinforcement learning is used to design a social mechanism that humans prefer by majority. A large group of humans played an online investment game that involved deciding whether to keep a monetary endowment or to share it with others for collective benefit. Shared revenue was returned to players under two different redistribution mechanisms, one designed by the AI and the other by humans. The AI discovered a mechanism that redressed initial wealth imbalance, sanctioned free riders and successfully won the majority vote. By optimizing for human preferences, Democratic AI offers a proof of concept for value-aligned policy innovation.
Artificial learning agents are mediating a larger and larger number of interactions among humans, firms, and organizations, and the intersection between mechanism design and machine learning has been heavily investigated in recent years. However, mechanism design methods often make strong assumptions on how participants behave (e.g. rationality), on the kind of knowledge designers have access to a priori (e.g. access to strong baseline mechanisms), or on what the goal of the mechanism should be (e.g. total welfare). Here we introduce HCMD-zero, a general purpose method to construct mechanisms making none of these three assumptions. HCMD-zero learns to mediate interactions among participants and adjusts the mechanism parameters to make itself more likely to be preferred by participants. It does so by remaining engaged in an electoral contest with copies of itself, thereby accessing direct feedback from participants. We test our method on a stylized resource allocation game that highlights the tension between productivity, equality and the temptation to free ride. HCMD-zero produces a mechanism that is preferred by human participants over a strong baseline, it does so automatically, without requiring prior knowledge, and using human behavioral trajectories sparingly and effectively. Our analysis shows HCMD-zero consistently makes the mechanism policy more and more likely to be preferred by human participants over the course of training, and that it results in a mechanism with an interpretable and intuitive policy.
'Intuitive physics' enables our pragmatic engagement with the physical world and forms a key component of 'common sense' aspects of thought. Current artificial intelligence systems pale in their understanding of intuitive physics, in comparison to even very young children. Here we address this gap between humans and machines by drawing on the field of developmental psychology. First, we introduce and open-source a machine-learning dataset designed to evaluate conceptual understanding of intuitive physics, adopting the violation-of-expectation (VoE) paradigm from developmental psychology. Second, we build a deep-learning system that learns intuitive physics directly from visual data, inspired by studies of visual cognition in children. We demonstrate that our model can learn a diverse set of physical concepts, which depends critically on object-level representations, consistent with findings from developmental psychology. We consider the implications of these results both for AI and for research on human cognition.
In order to build agents with a rich understanding of their environment, one key objective is to endow them with a grasp of intuitive physics; an ability to reason about three-dimensional objects, their dynamic interactions, and responses to forces. While some work on this problem has taken the approach of building in components such as ready-made physics engines, other research aims to extract general physical concepts directly from sensory data. In the latter case, one challenge that arises is evaluating the learning system. Research on intuitive physics knowledge in children has long employed a violation of expectations (VOE) method to assess children's mastery of specific physical concepts. We take the novel step of applying this method to artificial learning systems. In addition to introducing the VOE technique, we describe a set of probe datasets inspired by classic test stimuli from developmental psychology. We test a baseline deep learning system on this battery, as well as on a physics learning dataset ("IntPhys") recently posed by another research group. Our results show how the VOE technique may provide a useful tool for tracking physics knowledge in future research.
Motor adaptation displays a structure-learning effect: adaptation to a new perturbation occurs more quickly when the subject has prior exposure to perturbations with related structure. Although this `learning-to-learn' effect is well documented, its underlying computational mechanisms are poorly understood. We present a new model of motor structure learning, approaching it from the point of view of deep reinforcement learning. Previous work outside of motor control has shown how recurrent neural networks can account for learning-to-learn effects. We leverage this insight to address motor learning, by importing it into the setting of model-based reinforcement learning. We apply the resulting processing architecture to empirical findings from a landmark study of structure learning in target-directed reaching (Braun et al., 2009), and discuss its implications for a wider range of learning-to-learn phenomena.
The interpretation of other agents as intentional actors equipped with mental states has been connected to the attribution of rationality to their behavior. But a workable definition of “rationality” is difficult to formulate in complex situations, where standard normative definitions are difficult to apply. In this study, we explore a notion of rationality based on the idea of evolutionary fitness. We ask whether agents that are more adapted to their environment are, consequently, perceived as more rational and intentional. We created a 2-D virtual environment populated with autonomous virtual agents, each of which behaves according to a built-in program equipped with simulated perception, memory, and decision making. We then introduced a process of simulated evolution that pressured the agents’ programs toward behavior more adapted to the simulated environment. We showed these agents to human subjects in 2 experiments, in which we respectively asked them to judge their intelligence and to dynamically estimate their “mental states.” The results confirm that subjects construed evolved agents as more intelligent, and judged evolved agents’ mental states more accurately, relative to nonevolved agents. These results corroborate a view that the interpretation of agent behavior is connected to a concept of rationality based on the apparent fit between an agent’s actions and its environment.
The application of ideas from computational reinforcement learning has recently enabled dramatic advances in behavioral and neuroscientific research. For the most part, these advances have involved insights concerning the algorithms underlying learning and decision making. In the present article, we call attention to the equally important but relatively neglected question of how problems in learning and decision making are internally represented. To articulate the significance of representation for reinforcement learning we draw on the concept of efficient coding, as developed in perception research. The resulting perspective exposes a range of novel goals for behavioral and neuroscientific research, highlighting in particular the need for research into the statistical structure of naturalistic tasks.
Recent work has reawakened interest in goal-directed or ‘model-based’ choice, where decisions are based on prospective evaluation of potential action outcomes. Concurrently, there has been growing attention to the role of hierarchy in decision-making and action control. We focus here on the intersection between these two areas of interest, considering the topic of hierarchical model-based control. To characterize this form of action control, we draw on the computational framework of hierarchical reinforcement learning, using this to interpret recent empirical findings. The resulting picture reveals how hierarchical model-based mechanisms might play a special and pivotal role in human decision-making, dramatically extending the scope and complexity of human behaviour.
Inferring the mental states of other agents, including their goals and intentions, is a central problem in cognition. A critical aspect of this problem is that one cannot observe mental states directly, but must infer them from observable actions. To study the computational mechanisms underlying this inference, we created a two-dimensional virtual environment populated by autonomous agents with independent cognitive architectures. These agents navigate the environment, collecting “food” and interacting with one another. The agents’ behavior is modulated by a small number of distinct goal states: attacking, exploring, fleeing, and gathering food. We studied subjects’ ability to detect and classify the agents’ continually changing goal states on the basis of their motions and interactions. Although the programmed ground truth goal state is not directly observable, subjects’ responses showed both high validity (correlation with this ground truth) and high reliability (correlation with one another). We present a Bayesian model of the inference of goal states, and find that it accounts for subjects’ responses better than alternative models. Although the model is fit to the actual programmed states of the agents, and not to subjects’ responses, its output actually conforms better to subjects’ responses than to the ground truth goal state of the agents.
We focus on effective sample-based planning in the face of underactuation, high-dimensionality, drift, discrete system changes, and stochasticity. These are hallmark challenges for important problems, such as humanoid locomotion. In order to ensure broad applicability, we assume domain expertise is minimal and limited to a generative model. In order to make the method responsive, computational costs that scale linearly with the amount of samples taken from the generative model are required. We bring to bear a concrete method that satisfies all these requirements; it is a receding-horizon open-loop planner that employs cross-entropy optimization for policy construction. In simulation, we empirically demonstrate near-optimal decisions in a small domain and effective locomotion in several challenging humanoid control tasks.
Cross-entropy optimization (CE) has proven to be a powerful tool for search in control environments. In the basic scheme, a distribution over proposed solutions is repeatedly adapted by evaluating a sample of solutions and refocusing the distribution on a percentage of those with the highest scores. We show that, in the kind of noisy evaluation environments that are common in decision-making domains, this percentage-based refocusing does not optimize the expected utility of solutions, but instead a quantile metric. We provide a variant of CE (Proportional CE) that effectively optimizes the expected value. We show using variants of established noisy environments that Proportional CE can be used in place of CE and can improve solution quality.
Recently, rollout-based planning and search methods have emerged as an alternative to traditional tree-search methods. The fundamental operation in rollout-based tree search is the generation of trajectories in the search tree from root to leaf. Game-playing programs based on Monte-Carlo rollouts methods such as UCT have proven remarkably effective at using information from trajectories to make state-of-the-art decisions at the root. In this paper, we show that trajectories can be used to prune more aggressively than classical alpha-beta search. We modify a rollout-based method, FSSS, to allow for use in game-tree search and show it outprunes alpha-beta both empirically and formally.
Recent research leverages results from the continuous-armed bandit literature to create a reinforcement-learning algorithm for continuous state and action spaces. Initially proposed in a theoretical setting, we provide the first examination of the empirical properties of the algorithm. Through experimentation, we demonstrate the effectiveness of this planning method when coupled with exploration and model learning and show that, in addition to its formal guarantees, the approach is very competitive with other continuous-action reinforcement learners.
In some decision-making environments, successful solutions are common. If the evaluation of candidate solutions is noisy, however, the challenge is knowing when a “good enough” answer has been found. We formalize this problem as an infinite-armed bandit and provide upper and lower bounds on the number of evaluations or “pulls” needed to identify a solution whose evaluation exceeds a given threshold r0. We present several algorithms and use them to identify reliable strategies for solving screens from the video games Infinite Mario and Pitfall! We show order of magnitude improvements in sample complexity over a natural approach that pulls each arm until a good estimate of its success probability is known.
In this paper, we present a new algorithm that integrates recent advances in solving continuous bandit problems with sample-based rollout methods for planning in Markov Decision Processes (MDPs). Our algorithm, Hierarchical Optimistic Optimization applied to Trees (HOOT) addresses planning in continuous-action MDPs. Empirical results are given that show that the performance of our algorithm meets or exceeds that of a similar discrete action planner by eliminating the problem of manual discretization of the action space.
Perception of intentions and mental states in autonomous virtual agents Peter C. Pantelis, Steven Cholewiak, Paul Ringstad, Kevin Sanik, Ari Weinstein, Chia-Chien Wu, Jacob Feldman (petercp@eden.rutgers.edu, jacob@ruccs.rutgers.edu) Departments of Psychology, Center for Cognitive Science, Rutgers University-New Brunswick 152 Frelinghuysen Road, Piscataway, NJ 08854 USA Abstract 1995; Johnson, 2000). But the adult capacity to understand animate motion in terms of intelligent behavior has been re- searched less. Computational approaches to the problem of intention estimation are still scarce (Baker, Tenenbaum, & Saxe, 2006; Feldman & Tremoulet, 2008), in part because of the difficulty in specifying the problem in computational terms. Almost without exception, video stimuli used in past ex- periments in this area have consisted of hand-crafted an- imations with motions chosen subjectively by the experi- menters in order to achieve particular psychological impres- sions (Porter & Susman, 2000). This makes it difficult to in- vestigate the way subjects estimate intentionality, because the object of the estimation procedure—the actual mental state of the agent under observation—does not actually exist. Our proposed solution to this problem is to indeed endow our stimuli agents with “minds,” which our subjects then attempt to “read.” Comprehension of goal-directed, intentional motion is an im- portant but understudied visual function. To study it, we created a two-dimensional virtual environment populated by independently-programmed autonomous virtual agents, which navigate the environment, collecting food and competing with one another. Their behavior is modulated by a small number of distinct “mental states”: exploring, gathering food, attacking, and fleeing. In two experiments, we studied subjects’ ability to detect and classify the agents’ continually changing men- tal states on the basis of their motions and interactions. Our analyses compared subjects’ classifications to the ground truth state occupied by the observed agent’s autonomous program. Although the true mental state is inherently hidden and must be inferred, subjects showed both high validity (correlation with ground truth) and high reliability (correlation with one an- other). The data provide intriguing evidence about the factors that influence estimates of mental state—a key step towards a true “psychophysics of intention.” Keywords: animate motion perception; theory of mind; inten- tionality; action understanding; goal inference. Introduction Comprehension of the goals and intentions of other intelligent agents is an essential aspect of cognition. Motion is an espe- cially important cue to intention, as vividly illustrated by the famous short film by Heider and Simmel (1944). The “cast” of this film consists only of two triangles and a circle, but the motions of these simple geometrical figures are universally interpreted in terms of dramatic narrative. Indeed, it is practi- cally impossible to understand many naturally occurring mo- tions without comprehending the intentions that helped cause them: a person running is interpreted as trying to get some- where; a hand lifting a Coke can is automatically understood as a person intending to raise the can, not simply as two ob- jects moving upwards together (Mann, Jepson, & Siskind, 1997). Much of the motion in a natural environment—and certainly some of the most behaviorally important motion— is caused by other agents, and is impossible to understand except in terms of how and why they might have caused it. Human subjects readily attribute mentality and goal- directedness to moving objects as a function of properties of their motion (Tremoulet & Feldman, 2000), and in par- ticular on how that motion seems to relate to the motion of other agents and objects in the environment (Blythe, Todd, & Miller, 1999; Dittrich & Lea, 1994; Gao, McCarthy, & Scholl, 2010; Pantelis & Feldman, 2010; Tremoulet & Feld- man, 2006; Zacks, Kumar, Abrams, & Mehta, 2009). The broad problem of attributing mentality to others has received a great deal of attention in the philosophical literature (of- ten under the term mindreading), and has been most widely studied in infants and children (Gelman, Durgin, & Kaufman, A virtual environment of autonomous agents We developed a two-dimensional interactive virtual envi- ronment populated with autonomous virtual agents (Fig. 1). These agents (referred to as Independent Mobile Personali- ties, or IMPs), are simple but cognitively independent virtual robots, equipped with perception, planning, decision making, and goals. They move about in the virtual environment, in- teracting with each other, making intelligent though unpre- dictable decisions and taking steps to achieve simple goals. The IMPs are endowed with potentially distinct personalities and cognitive faculties, including variations in intelligence, memory, aggression, and strategy. The result is a complex, dynamic microcosm in which goal-directed behavior, and the perception thereof, can be studied in a controlled way. The inspiration is taken from the substantial literature on artificial life (Shao & Terzopoulos, 2007; Yaeger, 1994) in which inter- actions among virtual creatures have been extensively mod- eled. But unlike previous environments, our agents are cog- nitively complete, meaning that their behavior is entirely de- termined by autonomous decisions based on input they have received via their own senses, and are presented visually to subjects so that we may study how their intentions are inter- preted by observers. Our focus is on what can be understood from the IMPs’ motion alone; to this end, we depict the IMPs as triangles, so that they have clearly identifiable main axes and front ends, but otherwise minimal shape. Because we have direct access to the “actual” intentions and mental states of the agents—represented by a simple state variable in the autonomous program—we can compare this “ground truth”
This article presents a short history of the Kosovo mining-metallurgy industry within the course of the last 800 years. Kosovo, as a relatively small piece of land in the Balkans, contains that region's highest concentration of mineral wealth (silver, lead, zinc, tin, coal). This paper explains the social and political impacts of metal production in medieval Serbia, the Ottoman Empire, and Yugoslavia. The recent clashes between Serbs and Albanians are viewed in the light of the "geopolitics of minerals."