Large language models (LLMs) are increasingly used as surrogates for human participants, but it remains unclear which models best capture human behavior and why. To address this, we introduce Psych-201, a novel dataset that enables us to measure behavioral alignment at scale. We find that post-training – the stage that turns base models into useful assistants – consistently reduces alignment with human behavior across model families, sizes, and objectives. Moreover, this misalignment widens in newer model generations even as base models continue to improve. Finally, we find that persona-induction – a popular technique for eliciting human-like behavior by conditioning models on participant-specific information – does not improve predictions at the level of individuals. Taken together, our results suggest that the very processes that are currently employed to turn LLMs into useful assistants also make them less accurate models of human behavior.
Two hallmarks of biological computation are its flexibility and efficiency. These features are often attributed to cognitive control processes that balance external utility against computational cost. However, how the brain could implement such adaptive control remains unknown. Here, we provide one possible answer by combining the computational theory of rational meta-reasoning with a meta-learning algorithm recently proposed as a model of prefrontal cortex. This yields a recurrent neural network model that learns to select computations. In simple choice tasks, the model approximates the algorithms and representations of optimal symbolic models and reproduces neural dynamics observed in macaque orbitofrontal cortex. In multi-step planning tasks, the model replicates key behavioral signatures of human planning strategies and captures human neural dynamics associated with step-by-step mental simulation. Our framework unifies meta-reasoning and meta-learning by showing that learning to reason can be understood as learning to learn from information generated by one's own cognitive operations, providing a mechanistic account of how adaptive control of thought can be implemented in neural systems.
Despite decades of work, we still lack a robust, task-general theory of human behavior even in the simplest domains. In this paper we tackle the generality problem head-on, by aiming to develop a unified model for all tasks embedded in a task-space. In particular we consider the space of binary sequence prediction tasks where the observations are generated by the space parameterized by hidden Markov models (HMM). As the space of tasks is large, experimental exploration of the entire space is infeasible. To solve this problem we propose the adversarial construction approach, which helps identify tasks that are most likely to elicit a qualitatively novel behavior. Our results suggest that adversarial construction significantly outperforms random sampling of environments and therefore could be used as a proxy for optimal experimental design in high-dimensional task spaces.
If one important function of curiosity is to foster learning, what does curiosity direct agents to learn? The present research investigates what kinds of situations spark curiosity. Prior work has proposed many candidate triggers of curiosity, but they have rarely been disentangled in a single study. Using a Bayesian model applied to a trial-and-error learning task, we investigated the correspondence between optimal triggers of curiosity (those that maximize learning), heuristic triggers of curiosity (surprise and uncertainty), and participants’ reported curiosity. In Studies 1-2 (N = 848), adults’ curiosity was most sensitive to a “local” optimal trigger: how much would be learned about the immediate target of curiosity. Curiosity was less sensitive to heuristic triggers or a “global” optimal trigger: how much would be learned about broader learning goals. Study 3 (N = 310) showed that curiosity’s sensitivity to local learning was stronger in adults than in 5- to 9-year-olds. These studies suggest that curiosity is heightened by opportunities for learning about immediate targets of curiosity, but not always broader learning goals.
The nature of eye movements during visual search has been widely studied in cognitive science. Virtual reality (VR) paradigms are an opportunity to test whether computational models of search can predict naturalistic search behavior. However, existing ideal observer models are constrained by strong assumptions about the structure of the world, rendering them impractical for modeling the complexity of environments which can be studied in VR. To address these limitations, we modeled immersive visual search as a reinforcement learning problem, in which sequential decisions are made over a multidimensional representation of the environment learned by a convolutional neural network. In our formulation, RL agents learned a policy over latent states-effectively solving what is known as a meta-Markov decision process (meta-MDP), where each decision concerns how to allocate attention to information in the environment. Training deep-RL agents on the meta-MDP showed that learned (i.e., optimal) search policies converge to a classic ideal-observer model of search developed for simple (1D) stimuli. We compared the learned resource-rational policy with human gaze data from a visual-search experiment conducted in VR and found qualitative and quantitative alignment between model predictions and human behavior. However, both the model's simulated performance and its correspondence with human behavior depended strongly on the representational features available to the policy. These results suggest that naturalistic visual search behavior can partially be explained by resource-rational allocation of limited cognitive resources, and the choice of representation influences the degree of alignment between model and human behavior.
In complex environments, the space of possible plans is vast. Generating a good plan therefore requires judicious selection of which parts of the plan space to mentally explore. Drawing on past studies of human exploration, we propose that mental exploration might invoke similar mechanisms. In particular, we test the hypothesis that mental exploration during planning is uncertainty-driven, such that people will exhibit a tendency to explore parts of the plan space that have high epistemic uncertainty. We developed a route-planning task, displayed as a binary tree, where participants were instructed to collect as many treats (rewards) as possible by traversing the tree. By separating the planning and execution phases, we encouraged participants to externalize their planning process. We manipulated uncertainty by varying the number of potential future states available from each current state. Across two studies, the data suggest that people preferred to explore options with more successor states after controlling for value differences, supporting the uncertainty-driven planning hypothesis. We also found that uncertainty played a larger role during the planning phase than during the execution phase, consistent with the hypothesis that the uncertainty effect primarily reflects a property of human planning algorithms rather than an intrinsic preference for uncertainty.
Many real-world problems are defined by complex systems of interlocking constraints. How people are able to solve these problems with such limited working memory capacity remains poorly understood.We propose a formal model of human problem-solving that uses metacognitive knowledge of its own memory limits and imperfect reasoning to guide subproblem choice. We compare our model to human gameplay in two experiments using a variant of the classic game Minesweeper. In Experiment 1, we find that participants' accuracy was influenced both by the order of subproblems and their ability to externalize intermediate results, indicative of a memory bottleneck in reasoning. In Experiment 2, we used a mouse-tracking paradigm to assess participants' subproblem choice and time allocation. The model captures key patterns of subproblem ordering, error, and time allocation. Our results point toward memory limits and strategies for navigating those limits --- including the careful choice of subproblems and memory-offloading --- as central elements of human problem-solving.
When making decisions, we often have more information about some options than others. Previous work has shown that people are more likely to choose options that they look at more and those that they are more confident in. But should one always prefer options one knows more about? Intuition suggests not. Rather, how additional information impacts our preferences should depend critically on how valuable we expect the options to be. Here, we formalize this intuition in a Bayesian sequential sampling model where attention and confidence influence the precision of momentary evidence. Our model makes a key prediction: attention and confidence both increase choice probability for better-than-average options, and both decrease choice probability for worse-than-average options. We confirm this prediction in two experiments, in which we independently manipulate value and attention. Our results offer a novel perspective on prior work on the role of attention and confidence in decision-making, showing that people rely on contextual knowledge and uncertainty estimates to adaptively learn about their options and make better decisions.
Human learners often seek information rationally: they prefer more informative over less informative data, and ask questions that optimally discriminate between competing hypotheses. But this picture holds only in simple, controlled settings: in more complex tasks, the computations needed for optimal information seeking become intractable and human inquiry empirically is often far from optimal. How do people effectively navigate the information landscapes the real world presents? We propose a model of information-seeking under representational constraints, which simplifies complex data used to make inferences and seeks information rationally with respect to its costs. Six behavioral experiments using two popular search games support this account: we find that people reason imprecisely and make queries in ways that are objectively uninformative under standard models, but are actually efficient given their limited representational capacity. Participants also increasingly prefer simpler, less informative queries as they gain experience in a task, suggesting rapid adaptation to their cognitive limits. Our findings challenge theories of human information-seeking based on an idealized notion of rationality and instead support a more nuanced and realistic view: people are limited in what information they can represent but have information-seeking strategies that are efficient given those limits.
If curiosity is the engine of learning, what does it direct agents to learn? The present research investigates what kinds of situations spark curiosity. Using a Bayesian computational model applied to a modified multi-armed bandit task, we investigated the correspondence between optimal triggers of curiosity (those that maximize learning potential), heuristic triggers of curiosity (surprise and uncertainty), and participants’ reported curiosity. In Studies 1-2 (N = 848), we found that adults’ curiosity was most sensitive to “local” learning potential (the extent to which information would shed light on the immediate target of curiosity). Curiosity was less sensitive to “global” learning potential (the extent to which information would contribute to broader learning goals), uncertainty, or surprise. Study 3 (N = 310) showed that curiosity’s sensitivity to local learning potential strengthened between ages 5 to 9 years and adulthood. Together, these studies suggest that curiosity tracks opportunities for learning, especially those that support learning about immediate targets of curiosity rather than broader learning goals.
Most of us have experienced moments when we could not recall some piece of information, but felt that it was just out of reach. Research in metamemory has established that such judgments are often accurate; but what adaptive purpose do they serve? Here, we present an optimal model of how metacognitive monitoring (feeling of knowing) could dynamically inform metacognitive control of memory (the direction of retrieval efforts). In two experiments, we find that, consistent with the optimal model, people report having a stronger memory for targets they are likely to recall, and direct their search efforts accordingly, cutting off search when it is unlikely to succeed and prioritizing search for stronger memories. Our results suggest that metamemory is indeed adaptive, and motivate the development of process-level theories that account for the dynamic interplay between monitoring and control.
Human behavior is often assumed to be hierarchically structured, made up of abstract actions that can be decomposed into concrete actions. However, behavior is typically measured as a sequence of actions, which makes it difficult to infer its hierarchical structure. In this paper, we explore how people form hierarchically structured plans, using an experimental paradigm with observable hierarchical representations: participants create programs that produce sequences of actions in a language with explicit hierarchical structure. This task lets us test two well-established principles of human behavior: utility maximization (i.e. using fewer actions) and minimum description length (MDL; i.e. having a shorter program). We find that humans are sensitive to both metrics, but that both accounts fail to predict a qualitative feature of human-created programs, namely that people prefer programs with reuse over and above the predictions of MDL. We formalize this preference for reuse by extending the MDL account into a generative model over programs, modeling hierarchy choice as the induction of a grammar over actions. Our account can explain the preference for reuse and provides better predictions of human behavior, going beyond simple accounts of compressibility to highlight a principle that guides hierarchical planning.
Establishing a unified theory of cognition has been a major goal of psychology. While there have been previous attempts to instantiate such theories by building computational models, we currently do not have one model that captures the human mind in its entirety. Here we introduce Centaur, a computational model that can predict and simulate human behavior in any experiment expressible in natural language. We derived Centaur by finetuning a state-of-the-art language model on a novel, large-scale data set called Psych-101. Psych-101 reaches an unprecedented scale, covering trial-by-trial data from over 60,000 participants performing over 10,000,000 choices in 160 experiments. Centaur not only captures the behavior of held-out participants better than existing cognitive models, but also generalizes to new cover stories, structural task modifications, and entirely new domains. Furthermore, we find that the model’s internal representations become more aligned with human neural activity after finetuning. Taken together, Centaur is the first real candidate for a unified model of human cognition. We anticipate that it will have a disruptive impact on the cognitive sciences, challenging the existing paradigm for developing computational models.
Perfectly rational decision-making is almost always out of reach for people because their computational resources are limited. Instead, people may rely on computationally frugal heuristics that usually yield good outcomes. Although previous research has identified many such heuristics, discovering good heuristics and predicting when they will be used remains challenging. Here, we present a theoretical framework that allows us to use methods from machine learning to automatically derive the best heuristic to use in any given situation by considering how to make the best use of limited cognitive resources. To demonstrate the generalizability and accuracy of our method, we compare the heuristics it discovers against those used by people across a wide range of multi-attribute risky choice environments in a behavioral experiment that is an order of magnitude larger than any previous experiments of its type. Our method rediscovered known heuristics, identifying them as rational strategies for specific environments, and discovered novel heuristics that had been previously overlooked. Our results show that people adapt their decision strategies to the structure of the environment and generally make good use of their limited cognitive resources, although their strategy choices do not always fully exploit the structure of the environment.
Inferring an individual’s preferences from their observable behavior is a key step in the development of assistive decision-making technology. Although machine learning models such as neural networks could in principle be deployed toward this inference, a large amount of data is required to train such models. Here, we present an approach in which a cognitive model generates simulated data to augment limited human data. Using these data, we train a neural network to invert the model, making it possible to infer preferences from behavior. We show how this approach can be used to infer the value that people assign to food items from their eye movements when choosing between those items. We demonstrate first that neural networks can infer the latent preferences used by the model to generate simulated fixations, and second that simulated data can be beneficial in pretraining a network for predicting human-reported preferences from real fixations. Compared to inferring preferences from choice alone, this approach confers a slight improvement in predicting preferences and also allows prediction to take place prior to the choice being made. Overall, our results suggest that using a combination of neural networks and model-simulated training data is a promising approach for developing technology that infers human preferences.
People's judgments and decisions often deviate from classical notions of rationality, incurring costs both to themselves and to society. One way to reduce the costs of poor decisions is to redesign the decision problems people face to encourage better choices. While often subtle, these \emph{nudges} can have dramatic effects on behavior and are increasingly popular in public policy, healthcare, and marketing. Although nudges are often designed with psychological theories in mind, they are typically not formalized in computational terms and their effects can be hard to predict. As a result, designing nudges can be difficult and time-consuming. To address this challenge, we propose a computational framework for understanding and predicting the effects of nudges. Our framework builds on recent work modeling human decision-making as adaptive use of limited cognitive resources, an approach called \emph{resource-rational analysis}. Concretely, nudges change the optimal sequence of cognitive operations an agent should execute, which in turn influences the agent's behavior. We first show that our framework can account for known effects of nudges based on default options, suggested alternatives, and information highlighting. In each case, we validate the model's predictions in an experimental process-tracing paradigm. We then show how the framework can be used to automatically construct optimal nudges, and demonstrate that these nudges improve people's decisions more than intuitive heuristic approaches. Overall, our results show that resource-rational analysis is a promising framework for formally characterizing and constructing nudges.