People assume different and important roles within social networks. Some roles have received extensive study: that of influencers who are well-connected, and that of brokers who bridge unconnected parts of the network. However, very little work has explored another potentially important role, that of creating opportunities for people to interact and facilitating conversation between them. These individuals bring people together and act as social catalysts. In this paper, we test for the presence of social catalysts on the online social network Facebook. We first identify posts that have spurred conversations between the poster's friends and summarize the characteristics of such posts. We then aggregate the number of catalyzed comments at the poster level, as a measure of the individual's "catalystness." The top 1% of such individuals account for 31% of catalyzed interactions, although their network characteristics do not differ markedly from others who post as frequently and have a similar number of friends. By collecting survey data, we also validate the behavioral measure of catalystness: a person is more likely to be nominated as a social catalyst by their friends if their posts prompt discussions between other people more frequently. The measure, along with other conversation-related features, is one of the most predictive of a person being nominated as a catalyst. Although influencers and brokers may have gotten more attention for their network positions, our findings provide converging evidence that another important role exists and is recognized in online social networks.
Seventy-three children between 6 and 7 years of age were presented with a problem having ambiguous subgoal ordering. Performance in this task showed reliable fingerprints: (a) a non-monotonic dependence of performance as a function of the distance between the beginning and the end-states of the problem, (b) very high levels of performance when the first move was correct, and (c) states in which accuracy of the first move was significantly below chance. These features are consistent with a non-Markov planning agent, with an inherently inertial decision process, and that uses heuristics and partial problem knowledge to plan its actions. We applied a statistical framework to fit and test the quality of a proposed planning model (Monte Carlo Tree Search). Our framework allows us to parse out independent contributions to problem-solving based on the construction of the value function and on general mechanisms of the search process in the tree of solutions. We show that the latter are correlated with children's performance on an independent measure of planning, while the former is highly domain specific.
Human behavior has long been recognized to display hierarchical structure: actions fit together into subtasks, which cohere into extended goal-directed activities. Arranging actions hierarchically has well established benefits, allowing behaviors to be represented efficiently by the brain, and allowing solutions to new tasks to be discovered easily. However, these payoffs depend on the particular way in which actions are organized into a hierarchy, the specific way in which tasks are carved up into subtasks. We provide a mathematical account for what makes some hierarchies better than others, an account that allows an optimal hierarchy to be identified for any set of tasks. We then present results from four behavioral experiments, suggesting that human learners spontaneously discover optimal action hierarchies.
Studies suggest that dopaminergic neurons report a unitary, global reward prediction error signal. However, learning in complex real-life tasks, in particular tasks that show hierarchical structure, requires multiple prediction errors that may coincide in time. We used functional neuroimaging to measure prediction error signals in humans performing such a hierarchical task involving simultaneous, uncorrelated prediction errors. Analysis of signals in a priori anatomical regions of interest in the ventral striatum and the ventral tegmental area indeed evidenced two simultaneous, but separable, prediction error signals corresponding to the two levels of hierarchy in the task. This result suggests that suitably designed tasks may reveal a more intricate pattern of firing in dopaminergic neurons. Moreover, the need for downstream separation of these signals implies possible limitations on the number of different task levels that we can learn about simultaneously.
This work was supported by AFOSR FA9550-07-1-0075 and ONR N00014-07-1-0937. SJG was supported by a Graduate Research Fellowship from the NSF.
The cultural evolution of introspective thought has been recognized to undergo a drastic change during the middle of the first millennium BC. This period, known as the “Axial Age,” saw the birth of religions and philosophies still alive in modern culture, as well as the transition from orality to literacy—which led to the hypothesis of a link between introspection and literacy. Here we set out to examine the evolution of introspection in the Axial Age, studying the cultural record of the Greco-Roman and Judeo-Christian literary traditions. Using a statistical measure of semantic similarity, we identify a single “arrow of time” in the Old and New Testaments of the Bible, and a more complex non-monotonic dynamics in the Greco-Roman tradition reflecting the rise and fall of the respective societies. A comparable analysis of the twentieth century cultural record shows a steady increase in the incidence of introspective topics, punctuated by abrupt declines during and preceding the First and Second World Wars. Our results show that (a) it is possible to devise a consistent metric to quantify the history of a high-level concept such as introspection, cementing the path for a new quantitative philology and (b) to the extent that it is captured in the cultural record, the increased ability of human thought for self-reflection that the Axial Age brought about is still heavily determined by societal contingencies beyond the orality-literacy nexus.
The 10th European Workshop on Reinforcement Learning (EWRL), held June 30–July 1, 2012 in Edinburgh, Scotland, served as a forum to discuss the current state-of-the-art and future research directions in the continuously growing field of reinforcement learning (RL). We made EWRL an exciting and well-received event for the international RL community. Therefore, we appreciate that we could attract a wide spectrum of researchers by co-locating EWRL 2012 with the International Conference on Machine Learning. We were very excited about our outstanding invited speakers Martin Riedmiller (University of Freiburg), Drew Bagnell (Carnegie Mellon University), Shie Mannor (Technion), and Richard Sutton (University of Alberta), who gave great overviews of current research and diverse application areas of reinforcement learning. EWRL 2012 had 63 submissions from which 43 (68%) were accepted for presentation at the workshop. These post-proceedings contain 12 (19%) selected revised papers submitted to EWRL 2012. We thank our program committee for a fantastic job: Abdeslam Boularias, Adam White, Alborz Geramifard, Alessandro Lazaric, Amir-massoud Farahmand, Andre Damotta Salles Barreto, Andrew McHutchon, Bert Kappen, Bradley Knox, Byron Boots, Carlos Diuk Wasser, Christian Daniel, Christian Igel, Damien Ernst, David Silver, Doina Precup, Dvijotham Krishnamurthy, Emma Brunskill, Evangelos Theodorou, Fernand Fernandez, Francisco Melo, Gerhard Neumann, Hado van Hasselt, Jan Peters, Jens Kober, Jose Antonio Martin H., Jun Morimoto, Katharina Mülling, Kristian Kersting, Manuel Lopes, Marco Wiering, Martijn …
This paper presents a new algorithm for online linear regression whose efficiency guarantees satisfy the requirements of the KWIK (Knows What It Knows) framework. The algorithm improves on the complexity bounds of the current state-of-the-art procedure in this setting. We explore several applications of this algorithm for learning compact reinforcement-learning representations. We show that KWIK linear regression can be used to learn the reward function of a factored MDP and the probabilities of action outcomes in Stochastic STRIPS and Object Oriented MDPs, none of which have been proven to be efficiently learnable in the RL setting before. We also combine KWIK linear regression with other KWIK learners to learn larger portions of these models, including experiments on learning factored MDP transition and reward functions together.
This paper develops a generalized apprenticeship learning protocol for reinforcement-learning agents with access to a teacher who provides policy traces (transition and reward observations). We characterize sufficient conditions of the underlying models for efficient apprenticeship learning and link this criteria to two established learnability classes (KWIK and Mistake Bound). We then construct efficient apprenticeship-learning algorithms in a number of domains, including two types of relational MDPs. We instantiate our approach in a software agent and a robot agent that learn effectively from a human teacher.
The evolution of literary styles in the western tradition has been the subject of extended research that arguably has spanned centuries. In particular, previous work has conjectured the existence of a gradual yet persistent increase of the degree of self-awareness or introspection, i.e. that capacity to expound on one's own thought processes and behaviors, reflected in the chronology of the classical literary texts. This type of question has been traditionally addressed by qualitative studies in philology and literary theory. In this paper, we describe preliminary results based on the application of computational linguistics techniques to quantitatively analyze this hypothesis. We evaluate the appearance of introspection in texts by searching words related to it, and focus on simple studies on the Bible. This preliminary results are highly positive, indicating that it is indeed possible to statistically discriminate between texts based on a semantic core centered around introspection, chronologically and culturally belonging to different phases. In our opinion, the rigurous extension of our analysis can provide not only a stricter statistical measure of the evolution of introspection, but also means to investigate subtle differences in aesthetic styles and cognitive structures across cultures, authors and literary forms.
Reinforcement learning (RL) deals with the problem of an agent that has to learn how to behave to maximize its utility by its interactions with an environment (Sutton & Barto, 1998; Kaelbling, Littman & Moore, 1996). Reinforcement learning problems are usually formalized as Markov Decision Processes (MDP), which consist of a finite set of states and a finite number of possible actions that the agent can perform. At any given point in time, the agent is in a certain state and picks an action. It can then observe the new state this action leads to, and receives a reward signal. The goal of the agent is to maximize its long-term reward. In this standard formalization, no particular structure or relationship between states is assumed. However, learning in environments with extremely large state spaces is infeasible without some form of generalization. Exploiting the underlying structure of a problem can effect generalization and has long been recognized as an important aspect in representing sequential decision tasks (Boutilier et al., 1999). Hierarchical Reinforcement Learning is the subfield of RL that deals with the discovery and/or exploitation of this underlying structure. Two main ideas come into play in hierarchical RL. The first one is to break a task into a hierarchy of smaller subtasks, each of which can be learned faster and easier than the whole problem. Subtasks can also be performed multiple times in the course of achieving the larger task, reusing accumulated knowledge and skills. The second idea is to use state abstraction within subtasks: not every task needs to be concerned with every aspect of the state space, so some states can actually be abstracted away and treated as the same for the purpose of the given subtask.
The purpose of this paper is three-fold. First, we formalize and study a problem of learning probabilistic concepts in the recently proposed KWIK framework. We give details of an algorithm, known as the Adaptive k-Meteorologists Algorithm, analyze its sample-complexity upper bound, and give a matching lower bound. Second, this algorithm is used to create a new reinforcement-learning algorithm for factored-state problems that enjoys significant improvement over the previous state-of-the-art algorithm. Finally, we apply the Adaptive k-Meteorologists Algorithm to remove a limiting assumption in an existing reinforcement-learning algorithm. The effectiveness of our approaches is demonstrated empirically in a couple benchmark domains as well as a robotics navigation problem.
Rich representations in reinforcement learning have been studied for the purpose of enabling generalization and making learning feasible in large state spaces. We introduce Object-Oriented MDPs (OO-MDPs), a representation based on objects and their interactions, which is a natural way of modeling environments and offers important generalization opportunities. We introduce a learning algorithm for deterministic OO-MDPs and prove a polynomial bound on its sample complexity. We illustrate the performance gains of our representation and algorithm in the well-known Taxi domain, plus a real-life videogame.
We present an adaptive end-host anomaly detector where a supervised classifier trained as a traffic predictor is used to control a time-varying detection threshold. Using real enterprise traffic traces for both training and testing, we show that our detector outperforms a fixed-threshold detector. This comparison is robust to the choice of off-the-shelf classifier and to a variety of performance criteria, i.e., the predictor's error rate, the reduction in the "threshold gap," and the ability to detect incremental worm traffic that is added to real life traces. Our adaptive-threshold detector is intended as a part of a distributed worm detection system. This distributed system infers system-wide threats from end-host detections, thereby avoiding the sensing and resource limitations of conventional centralized systems. The system places a constraint on this end-host detector to appear consistent over time and host variability
We consider the problem of reinforcement learning in factored-state MDPs in the setting in which learning is conducted in one long trial with no resets allowed. We show how to extend existing efficient algorithms that learn the conditional probability tables of dynamic Bayesian networks (DBNs) given their structure to the case in which DBN structure is not known in advance. Our method learns the DBN structures as part of the reinforcement-learning process and provably provides an efficient learning algorithm when combined with factored Rmax.
Factored representations, model-based learning, and hierarchies are well-studied techniques for improving the learning efficiency of reinforcement-learning algorithms in large-scale state spaces. We bring these three ideas together in a new algorithm. Our algorithm tackles two open problems from the reinforcement-learning literature, and provides a solution to those problems in deterministic domains. First, it shows how models can improve learning speed in the hierarchy-based MaxQ framework without disrupting opportunities for state abstraction. Second, we show how hierarchies can augment existing factored exploration algorithms to achieve not only low sample complexity for learning, but provably efficient planning as well. We illustrate the resulting performance gains in example domains. We prove polynomial bounds on the computational effort needed to attain near optimal performance within the hierarchy.