Damage to the ventromedial prefrontal cortex (vmPFC) in humans disrupts planning abilities in naturalistic settings. However, it is unknown which components of planning are affected in these patients, including selecting the relevant information, simulating future states, or evaluating between these states. To address this question, we leveraged computational paradigms to investigate the role of vmPFC in planning, using the board game task "Four-in-a-Row" (18 lesion patients, 9 female; 30 healthy control participants, 16 female) and the simpler "Two-Step" task measuring model-based reasoning (49 lesion patients, 27 female; 20 healthy control participants, 13 female). Damage to vmPFC disrupted performance in Four-in-a-Row compared with both control lesion patients and healthy age-matched controls. We leveraged a computational framework to assess different component processes of planning in Four-in-a-Row and found that impairments following vmPFC damage included shallower planning depth and a tendency to overlook game-relevant features. In the "Two-Step" task, which involves binary choices across a short future horizon, we found little evidence of planning in all groups and no behavioral differences between groups. Complex yet computationally tractable tasks such as "Four-in-a-Row" offer novel opportunities for characterizing neuropsychological planning impairments, which in vmPFC patients we find are associated with oversights and reduced planning depth.
Human planning tends to be efficient, focusing on a relatively small number of options when considering future paths. Recent proposals have suggested that this efficiency reflects intelligent deployment of the limited resources available for planning. A prediction of this and related proposals is that when individuals spend time thinking should depend on the benefits and costs of additional computation. We tested this hypothesis by measuring how much time humans spent thinking before acting in over 12 million online chess games. Players spent more time thinking in board positions where additional computation was more beneficial. This relationship was greater in stronger players, and was strengthened by considering only the information available to the player at the time of choice. A simple model based on measuring the actual cost of spending time thinking in online chess was able to capture qualitative features of this relationship. These results provide evidence that the amount of time humans spend thinking is appropriately sensitive to the value of computation.
Board, card, or video games have been played by virtually every individual in the world population, with both children and adults participating. Games are popular because they are intuitive and fun. These distinctive qualities of games also make them ideal as a platform for studying the mind. By being intuitive, games provide a unique vantage point for understanding the inductive biases that support behavior in more complex, ecological settings than traditional lab experiments. By being fun, games allow researchers to study new questions in cognition such as the meaning of "play'' and intrinsic motivation, while also supporting more extensive and diverse data collection by attracting many more participants. We describe both the advantages and drawbacks of using games relative to standard lab-based experiments and lay out a set of recommendations on how to gain the most from using games to study cognition. We hope this article will lead to a wider use of games as experimental paradigms, elevating the ecological validity, scale, and robustness of research on the mind.
Acting intelligently in complex environments poses a challenging learning problem: faced with many different situations and possible actions, how do people learn which action to take in each situation? While traditional laboratory-based experiments have been used to study specific learning mechanisms, these experiments often employ relatively simple tasks conducted over a short period of time. Thus, it is unclear to what extent these mechanisms are used in the significantly more complex and temporally extended environments people encounter in their everyday lives. To understand the processes by which people learn policies to guide their decisions, we investigate the opening strategies of novice online chess players over their first months of play. We use a large online data set consisting of 2,499,783 games, providing us with the necessary scale to explore learning mechanisms in a complex setting. In particular, we focus on two types of learning: reinforcement learning, or learning from rewards given repeated experiences, and social learning, or learning from the actions of others. We show that players’ choices are modulated by both game outcomes and observing their opponents’ actions, and that they exhibit important hallmarks of adaptive decision-making such as exploration and expertise. Our results provide evidence that people use sophisticated learning algorithms in naturalistic strategic behavior.
A hallmark of human intelligence is the ability to plan multiple steps into the future(1,2). Despite decades of research(3-5), it is still debated whether skilled decision-makers plan more steps ahead than novices(6-8). Traditionally, the study of expertise in planning has used board games such as chess, but the complexity of these games poses a barrier to quantitative estimates of planning depth. Conversely, common planning tasks in cognitive science often have a lower complexity(9,10) and impose a ceiling for the depth to which any player can plan. Here we investigate expertise in a complex board game that offers ample opportunity for skilled players to plan deeply. We use model fitting methods to show that human behaviour can be captured using a computational cognitive model based on heuristic search. To validate this model, we predict human choices, response times and eye movements. We also perform a Turing test and a reconstruction experiment. Using the model, we find robust evidence for increased planning depth with expertise in both laboratory and large-scale mobile data. Experts memorize and reconstruct board features more accurately. Using complex tasks combined with precise behavioural modelling might expand our understanding of human planning and help to bridge the gap with progress in artificial intelligence.
How will superhuman artificial intelligence (AI) affect human decision-making? And what will be the mechanisms behind this effect? We address these questions in a domain where AI already exceeds human performance, analyzing more than 5.8 million move decisions made by professional Go players over the past 71 y (1950 to 2021). To address the first question, we use a superhuman AI program to estimate the quality of human decisions across time, generating 58 billion counterfactual game patterns and comparing the win rates of actual human decisions with those of counterfactual AI decisions. We find that humans began to make significantly better decisions following the advent of superhuman AI. We then examine human players' strategies across time and find that novel decisions (i.e., previously unobserved moves) occurred more frequently and became associated with higher decision quality after the advent of superhuman AI. Our findings suggest that the development of superhuman AI programs may have prompted human players to break away from traditional strategies and induced them to explore novel moves, which in turn may have improved their decision-making.
This study aimed to distinguish between the developmental trajectories of different cognitive component processes underlying planning decisions. Participants (8-25 year olds) completed a planning task called Four-in-a-row. By fitting a computational model, we distinguished between three cognitive component processes of planning: planning depth, heuristic quality, and attentional oversights. All three contributed to better playing strength, but they differed in their developmental trajectories. Specifically, from early to mid-adolescence, heuristic quality rapidly improved and contributed to better playing strength. From mid to late-adolescence, planning depth increased and supported better playing strength. Fewer attentional oversights were associated with better playing strength and this relation did not show age differences. Together, these results suggest an order in which the use of cognitive component processes of planning develop, starting by first refining the heuristic strategies, then gradually increasing the number possible future actions, states, and outcomes considered towards young adulthood. The findings move the field of cognitive development towards a more complete account of the development of planning and its component processes.
The nature of eye movements during visual search has been widely studied in psychology and neuroscience. Virtual reality (VR) paradigms provide an opportunity to test whether computational models of search can predict naturalistic search behavior. However, existing ideal observer models are constrained by strong assumptions about the structure of the world, rendering them impractical for modeling the complexity of environments that can be studied in VR. To address these limitations, we frame naturalistic visual search as a problem of allocating limited cognitive resources, formalized as a meta-level Markov decision process (meta-MDP) over a representation of the environment encoded by a deep neural network. We train reinforcement learning agents to solve the meta-MDP, showing that the agents’ optimal policy converges to a classic ideal observer model of search developed for simplified environments. We compare the learned policy with human gaze data from a visual search experiment conducted in VR, finding a qualitative and quantitative correspondence between model predictions and human behavior. Our results suggest that gaze behavior in naturalistic visual search is consistent with rational allocation of limited cognitive resources.### Competing Interest StatementThe authors have declared no competing interest.
SignificanceMany bad decisions and their devastating consequences could be avoided if people used optimal decision strategies. Here, we introduce a principled computational approach to improving human decision making. The basic idea is to give people feedback on how they reach their decisions. We develop a method that leverages artificial intelligence to generate this feedback in such a way that people quickly discover the best possible decision strategies. Our empirical findings suggest that a principled computational approach leads to improvements in decision-making competence that transfer to more difficult decisions in more complex environments. In the long run, this line of work might lead to apps that teach people clever strategies for decision making, reasoning, goal setting, planning, and goal achievement.
Human decision-making is plagued by many systematic errors. Many of these errors can be avoided by providing decision aids that guide decision-makers to attend to the important information and integrate it according to a rational decision strategy. Designing such decision aids used to be a tedious manual process. Advances in cognitive science might make it possible to automate this process in the future. We recently introduced machine learning methods for discovering optimal strategies for human decision-making automatically and an automatic method for explaining those strategies to people. Decision aids constructed by this method were able to improve human decision-making. However, following the descriptions generated by this method is very tedious. We hypothesized that this problem can be overcome by conveying the automatically discovered decision strategy as a series of natural language instructions for how to reach a decision. Experiment 1 showed that people do indeed understand such procedural instructions more easily than the decision aids generated by our previous method. Encouraged by this finding, we developed an algorithm for translating the output of our previous method into procedural instructions. We applied the improved method to automatically generate decision aids for a naturalistic planning task (i.e., planning a road trip) and a naturalistic decision task (i.e., choosing a mortgage). Experiment 2 showed that these automatically generated decision-aids significantly improved people's performance in planning a road trip and choosing a mortgage. These findings suggest that AI-powered boosting might have potential for improving human decision-making in the real world.
Making good decisions requires thinking ahead, but the huge number of actions and outcomes one could consider makes exhaustive planning infeasible for computationally constrained agents, such as humans. How people are nevertheless able to solve novel problems when their actions have long-reaching consequences is thus a long-standing question in cognitive science. To address this question, we propose a model of resource-constrained planning that allows us to derive optimal planning strategies. We find that previously proposed heuristics such as best-first search are near-optimal under some circumstances, but not others. In a mouse-tracking paradigm, we show that people adapt their planning strategies accordingly, planning in a manner that is broadly consistent with the optimal model but not with any single heuristic model. We also find systematic deviations from the optimal model that might result from additional cognitive constraints that are yet to be uncovered.
Many human abilities rely on cognitive algorithms discovered by previous generations. Cultural accumulation of innovative algorithms is hard to explain because complex concepts are difficult to pass on. We found that selective social learning preserved rare discoveries of exceptional algorithms in a large experimental simulation of cultural evolution. Participants ( N = 3450) faced a difficult sequential decision problem (sorting an unknown sequence of numbers) and transmitted solutions across 12 generations in 20 populations. Several known sorting algorithms were discovered. Complex algorithms persisted when participants could choose who to learn from but frequently became extinct in populations lacking this selection process, converging on highly transmissible lower-performance algorithms. These results provide experimental evidence for hypothesized links between sociality and cognitive function in humans.
Author(s): Becker, Frederic; Skirzynski, Julian Mateusz; van Opheusden, Bas; Lieder, Falk | Abstract: People often fall victim to decision-making biases, e.g. short-sightedness, that lead to unfavorable outcomes in their lives. It is possible to overcome these biases by teaching people better decision-making strategies. Finding effective interventions is an open problem, with a key challenge being the lack of transfer to the real world. Here, we tested a new approach to improving human decision-making that leverages Artificial Intelligence to discover procedural descriptions of effective planning strategies. Our benchmark problem regarded improving far-sightedness. We found our intervention elicits transfer to a similar task in a different domain, but its effects in more naturalistic financial decisions were not statistically significant. Even though the tested intervention is on par with conventional approaches, which also struggle in far-transfer, further improvements are required to help people make better decisions in real life. We conclude that future work should focus on training decision-making in more naturalistic scenarios.
In recent years, artificial intelligence has made great progress in improving machine performance in tasks that require planning many steps ahead. By comparison, cognitive science has lagged behind in understanding human planning in complex tasks. One question of long-standing interest in this domain is whether skilled decision-makers plan further into the future than novices. Traditionally, the study of expertise in planning has focused on board games like chess, but the complexity of these games poses a barrier to detailed behavioral modeling. Conversely, common planning tasks in cognitive science are often lower-complexity and impose a ceiling for the depth to which any player can plan. Here, we investigate expertise in a complex board game that offers ample opportunity for skilled players to plan deeply. Despite this complexity, we show that human behavior can be captured using a computational cognitive model based on heuristic search. To validate this model, we predict human choices, response times, eye movements and perform a Turing test. Using the model, we find robust evidence for increased planning depth with expertise in both laboratory and large-scale mobile data. Our results highlight the promise of investigating human planning in complex tasks with precise behavioral modeling.
Visual search is a ubiquitous human behavior and canonical example of selectively sampling sensory information to attain a goal. Previous research has studied optimality in visual search with artificial laboratory tasks (Najemnik and Geisler, 2005; Yang et al. 2016). To understand how people search in naturalistic environments, we conducted a study of visual search in virtual reality. Participants (N=21) viewed scenes generated with the Unity game engine through a head-mounted display equipped with an eye-tracker. On each of 300 trials, participants were shown a target object and teleported into a virtual cluttered room where they searched for the item from a fixed viewpoint. They had 8 seconds to identify the target object among 60-100 distractors. Participants had a 76% success rate of finding the target with a median response time on successful trials of 2.89s (IQR: 1.99-4.44s). To understand what features drive people’s search, we annotated gaze samples with semantic scene information such as the identity, shape, color, and texture of the object at the center of gaze. Concretely, we used the object asset (3D mesh and texture) to compute low-dimensional shape and color representations of each object. We found that people’s gaze is primarily directed to task-relevant objects (i.e. targets or distractors), and that the distractors that people look at are close in representational space to the target. Furthermore, this distance decreased over time, suggesting that representational similarity guides eye movements. We discuss these results in the context of a meta-level Markov Decision Process model (Callaway et al. 2018), which frames visual search as optimal information sampling under computational constraints.
The fate of scientific hypotheses often relies on the ability of a computational model to explain the data, quantified in modern statistical approaches by the likelihood function. The log-likelihood is the key element for parameter estimation and model evaluation. However, the log-likelihood of complex models in fields such as computational biology and neuroscience is often intractable to compute analytically or numerically. In those cases, researchers can often only estimate the log-likelihood by comparing observed data with synthetic observations generated by model simulations. Standard techniques to approximate the likelihood via simulation either use summary statistics of the data or are at risk of producing substantial biases in the estimate. Here, we explore another method, inverse binomial sampling (IBS), which can estimate the log-likelihood of an entire data set efficiently and without bias. For each observation, IBS draws samples from the simulator model until one matches the observation. The log-likelihood estimate is then a function of the number of samples drawn. The variance of this estimator is uniformly bounded, achieves the minimum variance for an unbiased estimator, and we can compute calibrated estimates of the variance. We provide theoretical arguments in favor of IBS and an empirical assessment of the method for maximum-likelihood estimation with simulation-based models. As case studies, we take three model-fitting problems of increasing complexity from computational and cognitive neuroscience. In all problems, IBS generally produces lower error in the estimated parameters and maximum log-likelihood values than alternative sampling methods with the same average number of samples. Our results demonstrate the potential of IBS as a practical, robust, and easy to implement method for log-likelihood evaluation when exact techniques are not available.
Effective use of limited computational resources is a hallmark of intelligent systems. Here we explore how people make use of limited perceptual resources in a naturalistic visual search task. We hypothesize that people optimally trade off performance against the cost of sampling information. To formalize this hypothesis, we frame the problem of attention allocation in visual search as a meta-level Markov decision process (meta-MDP) and show that a classic Bayesian model of visual search can be interpreted as a heuristic policy for this meta-MDP. We test the heuristic policy against gaze data from 21 human participants in a virtual reality (VR) visual search study, finding that human gaze trajectories share qualitative structure with trajectories simulated from the model.
An ideal Mixed Reality (MR) system would only present virtual information (e.g., a label) when it is useful to the person. However, deciding when a label is useful is challenging: it depends on a variety of factors, including the current task, previous knowledge, context, etc. In this paper, we propose a Reinforcement Learning (RL) method to learn when to show or hide an object's label given eye movement data. We demonstrate the capabilities of this approach by showing that an intelligent agent can learn cooperative policies that better support users in a visual search task than manually designed heuristics. Furthermore, we show the applicability of our approach to more realistic environments and use cases (e.g., grocery shopping). By posing MR object labeling as a model-free RL problem, we can learn policies implicitly by observing users' behavior without requiring a visual search model or data annotation.
What algorithms do people use to make decisions with future consequences in complex environments? In order to investigate the cognitive processes underlying sequential planning, we collected large-scale behavioral data in a challenging variant of tic-tac-toe. This task is at an intermediate level of complexity, providing rich behavior for which modeling is still tractable. We argue that a data set of this nature is necessary for distinguishing theoretical frameworks for integration between prospective and retrospective decision-making, and show preliminary evidence for the existence of both systems in our task. We outline a computational model based on an intuitive value function and decision tree search to demonstrate that people engage in prospective planning. We then explain discrepancies between the model’s predictions and observed data in early game choices, finding behavioral patterns consistent with retrospective learning.