Predicting human decisions under risk and uncertainty remains a fundamental challenge across disciplines. Existing models often struggle even in highly stylized tasks like choice between lotteries. Here we introduce BEAST gradient boosting (BEAST-GB), a hybrid model integrating behavioural theory (BEAST) with machine learning. We first present CPC18, a competition for predicting risky choice, in which BEAST-GB won. Then, using two large datasets, we demonstrate that BEAST-GB predicts more accurately than neural networks trained on extensive data and dozens of existing behavioural models. BEAST-GB also generalizes robustly across unseen experimental contexts, surpassing direct empirical generalization, and helps to refine and improve the behavioural theory itself. Our analyses highlight the potential of anchoring predictions on behavioural theory even in data-rich settings and even when the theory alone falters. Our results underscore how integrating machine learning with theoretical frameworks, especially those-like BEAST-designed for prediction, can improve our ability to predict and understand human behaviour.
We study human performance in two classical NP-hard optimization problems: Set Cover and Maximum Coverage. We suggest that Set Cover and Max Coverage are related to means selection problems that arise in human problem-solving and in pursuing multiple goals: The relationship between goals and means is expressed as a bipartite graph where edges between means and goals indicate which means can be used to achieve which goals. While these problems are believed to be computationally intractable in general, they become more tractable when the structure of the network resembles a tree. Thus, our main prediction is that people should perform better with goal systems that are more tree-like. We report three behavioral experiments which confirm this prediction. Our results suggest that combinatorial parameters that are instrumental to algorithm design can also be useful for understanding when and why people struggle to choose between multiple means to achieve multiple goals.
Predicting and understanding how people make decisions has been a long-standing goal in many fields, with quantitative models of human decision-making informing research in both the social sciences and engineering. We show how progress toward this goal can be accelerated by using large datasets to power machine-learning algorithms that are constrained to produce interpretable psychological theories. Conducting the largest experiment on risky choice to date and analyzing the results using gradient-based optimization of differentiable decision theories implemented through artificial neural networks, we were able to recapitulate historical discoveries, establish that there is room to improve on existing theories, and discover a new, more accurate model of human decision-making in a form that preserves the insights from centuries of research.
The explosion of data generated during human interactions online presents an opportunity for psychologists to evaluate cognitive models outside the confines of the laboratory. Moreover, the size of these online data sets can allow researchers to construct far richer models than would be feasible with smaller in-lab behavioral data. In the current article, we illustrate this potential by evaluating 3 popular psychological models of generalization on 2 web-scale online data sets typically used to build automated recommendation systems. We show that each psychological model can be efficiently implemented at scale and in certain cases can capture trends in human judgments that standard recommendation systems from machine learning miss. We use these results to illustrate the opportunity Internet-scale data sets offer to psychologists and to underscore the importance of using insights from cognitive modeling to supplement the standard predictive-analytic approach taken by many existing machine learning approaches. (PsycInfo Database Record (c) 2020 APA, all rights reserved).
nbgrader is a flexible tool for creating and grading assignments in the Jupyter Notebook (Kluyver et al., 2016). nbgrader allows instructors to create a single, master copy of an assignment, including tests and canonical solutions. From the master copy, a student version is generated without the solutions, thus obviating the need to maintain two separate versions. nbgrader also automatically grades submitted assignments by executing the notebooks and storing the results of the tests in a database. After auto-grading, instructors can manually grade free responses and provide partial credit using the formgrader Jupyter Notebook extension. Finally, instructors can use nbgrader to leave personalized feedback for each student’s submission, including comments as well as detailed error information.
Human decision-making underlies all economic behavior. For the past four decades, human decision-making under uncertainty has continued to be explained by theoretical models based on prospect theory, a framework that was awarded the Nobel Prize in Economic Sciences. However, theoretical models of this kind have developed slowly, and robust, high-precision predictive models of human decisions remain a challenge. While machine learning is a natural candidate for solving these problems, it is currently unclear to what extent it can improve predictions obtained by current theories. We argue that this is mainly due to data scarcity, since noisy human behavior requires massive sample sizes to be accurately captured by off-the-shelf machine learning methods. To solve this problem, what is needed are machine learning models with appropriate inductive biases for capturing human behavior, and larger datasets. We offer two contributions towards this end: first, we construct "cognitive model priors" by pretraining neural networks with synthetic data generated by cognitive models (i.e., theoretical models developed by cognitive psychologists). We find that fine-tuning these networks on small datasets of real human decisions results in unprecedented state-of-the-art improvements on two benchmark datasets. Second, we present the first large-scale dataset for human decision-making, containing over 240,000 human judgments across over 13,000 decision problems. This dataset reveals the circumstances where cognitive model priors are useful, and provides a new standard for benchmarking prediction of human decisions under uncertainty.
People impose structure onto other agents’ sequential problem-solving behavior. That is, they interpret actions in terms of a likely program that the observed agent was executing to solve a problem. But what prior expectations do people have about these programs? For example, in both cognitive science and computer science, shortest description length has been proposed as a general principle for inducing a program. However, there may be other criteria that bias how people reconstruct others’ solutions: That they are symmetric, balanced, or organize child and parent processes in particular ways. Here, we report preliminary experiments and models that investigate peoples’ priors on others’ problem-solving programs. We first present a novel experimental paradigm in which participants were given examples of how a problem was solved and needed to reconstruct the program that generated the solution. Then, we discuss the application of our model of human program priors to these data. We find that shortest description length inadequately explains how people reconstruct others’ problem solving programs.
The explosion of data generated during human interactions online presents an opportunity for cognitive scientists to evaluate their models on popular real-world tasks outside the confines of the laboratory. We demonstrate this approach by evaluating two cognitive models of generalization against two machine learning approaches to recommendation on an online dataset of over 100K human playlist selections. Across two experiments we demonstrate that a model from cognitive science can both be efficiently implemented at scale and can capture generalization trends in human recommendation judgments which neither machine learning model is capable of replicating. We use these results to illustrate the opportunity internet-scale datasets offer to cognitive scientists, as well as to underscore the importance of using insights from cognitive modeling to supplement the standard predictive-analytic approach taken by many existing machine learning approaches.
The importance of hierarchically structured representations for tractable planning has long been acknowledged. However, the questions of how people discover such abstractions and how to define a set of optimal abstractions remain open. This problem has been explored in cognitive science in the problem solving literature and in computer science in hierarchical reinforcement learning. Here, we emphasize an algorithmic perspective on learning hierarchical representations in which the objective is to efficiently encode the structure of the problem, or, equivalently, to learn an algorithm with minimal length. We introduce a novel problem-solving paradigm that links problem solving and program induction under the Markov Decision Process (MDP) framework. Using this task, we target the question of whether humans discover hierarchical solutions by maximizing efficiency in number of actions they generate or by minimizing the complexity of the resulting representation and find evidence for the primacy of representational efficiency.
Evolutionary theory describes the dynamics of population change in settings affected by reproduction, selection, mutation, and drift. In the context of human cognition, evolutionary theory is most often invoked to explain the origins of capacities such as language, metacognition, and spatial reasoning, framing them as functional adaptations to an ancestral environment. However, evolutionary theory is useful for understanding the mind in a second way: as a mathematical framework for describing evolving populations of thoughts, ideas, and memories within a single mind. In fact, deep correspondences exist between the mathematics of evolution and of learning, with perhaps the deepest being an equivalence between certain evolutionary dynamics and Bayesian inference. This equivalence permits reinterpretation of evolutionary processes as algorithms for Bayesian inference and has relevance for understanding diverse cognitive capacities, including memory and creativity.
When people encode a representation of a scene, they do not necessarily represent the exact locations and orientations of the constituent elements. Instead, people rely on preexisting inductive biases to simplify their encoding of new scene configurations. We investigated people’s inductive biases in their memory for configurations of simple 2D shapes (such as circles, triangles, etc.) using a serial reproduction paradigm (Bartlett, 1932). This paradigm establishes an iterative process in which information is transmitted through a chain of people (like the ”telephone” game). In our experiment, we asked people to memorize configurations of simple shapes (which were either generated at random or by other participants) and then asked them to reproduce those configurations. In analyzing the final generation of reproductions, we found that people have strong preferences for the scale of individual shapes, as well as the alignment, distance, overlap, and relative rotation between pairs of shapes.
Most psychological theories attribute people’s failure to achieve their goals exclusively to insufficient motivation or lack of skill. Here, we offer a complementary explanation that emphasizes the inherent complexity of the computational problems that arise from the structure of people’s goal systems. Concretely, we hypothesize that people’s capacity to achieve their goals can be predicted from combinatorial parameters of the structure of the network connecting their goals to the means available to pursue them. To test this hypothesis, we expressed the relationship between goals and means as a bipartite graph where edges between means and goals indicate which means can be used to achieve which goals. This allowed us to map two computational challenges that arise in goal achievement onto two classic NP-hard problems: Set Cover and Maximum Coverage. The connection between goal pursuit and NP-hard problems led us to predict that people should perform better with goal systems that are tree-like. Three behavioral experiments confirmed this prediction. Our results imply that network parameters that are instrumental to algorithm design could also be useful for understanding when and why people struggle in their goal pursuits.
Empirical Evidence for Markov Chain Monte Carlo in Memory Search David D. Bourgin (ddbourgin@berkeley.edu) Joshua T. Abbott (joshua.abbott@berkeley.edu) Thomas L. Griffiths (tom griffiths@berkeley.edu) Department of Psychology University of California Berkeley Berkeley, CA 94720 USA Abstract Kevin A. Smith (k2smith@ucsd.edu) Edward Vul (evul@ucsd.edu) Department of Psychology University of California San Diego La Jolla, CA 92093 USA mote Associates Test (RAT; S. A. Mednick, 1962), a classic multiply-constrained search task. Building on recent work using random walks on semantic networks to model semantic fluency (Griffiths, Steyvers, & Firl, 2007) and singly-constrained search in memory (Abbott, Austerweil, & Griffiths, 2012), we consider a range of MCMC-inspired models for search on the RAT. We find that human response patterns are well-matched by a model closely related to the Metropolis-Hastings (M-H) algorithm, provid- ing empirical support for MCMC as a candidate for modeling creative ideation. The plan for the paper is as follows. We first review pre- vious work on human response patterns on the Remote Asso- ciates Test and discuss existing theories of search on this task. We then introduce Markov chain Monte Carlo methods with an emphasis on their application to creative search. We end by formally introducing the model used in the current study, and offer an analysis of its behavior in relation to human data collected by Smith, Huber, and Vul (2013). Previous theoretical work has proposed the use of Markov chain Monte Carlo as a model of exploratory search in mem- ory. In the current study we introduce such a model and eval- uate it on a semantic network against human performance on the Remote Associates Test (RAT), a commonly used creativ- ity metric. We find that a family of search models closely re- sembling the Metropolis-Hastings algorithm is capable of re- producing many of the response patterns evident when human participants are asked to report their intermediate guesses on a RAT problem. In particular we find that when run our model produces the same response clustering patterns, local depen- dencies, undirected search trajectories, and low associative hi- erarchies witnessed in human responses. Keywords: Creativity; Remote Associates Test; Information retrieval; Semantic networks; Markov chain Monte Carlo. Introduction Exploratory search in memory is a major component of cre- ative problem solving. Often, the demands of a task impose constraints on the solution space: when searching for a word to rhyme with A that also means B, a would-be poet is un- likely to include items that mean T or that rhyme with Q (as- suming that Q does not rhyme with A and T is not a synonym of B) in her search. Similarly, an inventor would be remiss to consider inventions that do not satisfy certain requirements (e.g., novelty, usefulness) if she plans to pursue a patent. The current study proposes a formal model for the process by which people perform this type of exploratory search under multiple constraints. Locating relevant pieces of information in memory re- quires a strategy for quickly traversing the space of poten- tial solutions. Markov chain Monte Carlo (MCMC) methods have been a particularly useful in this regard, offering effi- cient statistical techniques for exploring spaces that would otherwise prove computationally intractable to traverse fully (Brooks, 1998). MCMC methods have been employed to model phenomena as diverse as theory change (Ullman, Goodman, & Tenenbaum, 2012), perceptual multistability (Gershman, Vul, & Tenenbaum, 2012), and conceptual flu- idity (Gabora, 2000), making them a popular tool within sta- tistical and computational cognitive modeling. MCMC meth- ods have also been proposed within the creativity literature to model search behavior in memory (e.g., Martindale, 1995; Paulus, Levine, Brown, Minai, & Doboli, 2010), although to date there exists little empirical evidence on which to evaluate these proposals. In the current study we aim to fill this gap by evaluating a formal model of MCMC search on the Re- Background Stage models of problem solving propose that individuals first search through memory to identify potential answers to a problem and then test those candidate answers against the constraints of the problem to determine acceptability (Gruenewald & Lockhead, 1980; Raaijmakers & Shiffrin, 1981). In the current paper it is assumed that these search and test processes constitute qualitatively distinct cognitive tasks. Thus in reviewing recent work on the RAT and mem- ory search we emphasize accounts that focus not only on the identification of correct answers but also on the path by which these answers are selected. The Remote Associates Test In the Remote Associates Test participants are shown a set of three cue words (e.g., ‘surprise,’ ‘line,’ ‘birthday’) and in- structed to generate a fourth word to relate them (in this case, ‘party’). Although originally introduced to test for individual differences creative ability (S. A. Mednick, 1962), early work established that performance on the RAT correlated with IQ (M. T. Mednick & Andrews, 1967) and originality during brainstorming (Forbach & Evans, 1981). More recently the RAT has also been used to measure the effects of manipu- lations related to creative ability, including intuition and in- cubation (Bowers, Regehr, Balthazard, & Parker, 1990; Vul & Pashler, 2007), the role of affect during problem solving