Large language models (LLMs) are increasingly deployed in teams, yet existing coordination approaches often occupy two extremes. Highly structured methods rely on fixed roles, pipelines, or task decompositions assigned a priori. In contrast, fully unstructured teams enable adaptability and exploration but suffer from inefficiencies such as error propagation, inter-agent conflicts, and wasted resources (measured in time, tokens, or file operations). We introduce Language Agent Teams for Task Evolution (LATTE), a framework for coordinating LLM teams inspired by distributed systems, where processors must operate under partial observability and communication constraints. In LATTE, a team of agents collaboratively construct and maintain a shared, evolving coordination graph which encodes sub-task dependencies, individual agent assignment, and the current state of sub-task progress. This protocol maintains consistency while empowering agents to dynamically allocate work, adapt coordination, and discover new tasks. Across multiple collaborative tasks and a variety of base models, we demonstrate how LATTE reduces token usage, wall-clock time, communication, and coordination failures (e.g. file conflicts and redundant outputs) while matching or exceeding the accuracy of standard designs including MetaGPT, decentralized teams, top-down Leader-Worker hierarchies, and static decompositions.
Reliable human-machine discrimination is becoming increasingly important as large language models and autonomous agents are deployed in online settings. Existing approaches evaluate whether a system can produce behavior or responses indistinguishable from those of a human, following the emphasis on outputs as a criterion for intelligence proposed by Alan Turing. Cognitive science offers an alternative perspective: evaluating the process by which behavior is produced. To test whether cognitive processes can reliably distinguish humans from machines, we introduce CogCAPTCHA30, a battery of 30 cognitive tasks designed to elicit diagnostic process-level features even when task performance is matched. Across the battery, process-level features provide stronger discriminative signal than performance metrics alone, reliably distinguishing humans from agents even under output matching (mean process-feature classifier AUC = 0.88). To evaluate agentic process differences, we compare off-the-shelf frontier agents (Claude Sonnet 4.5, GPT-5, Gemini 2.5 Pro), Centaur (a language model fine-tuned on 10.7M human decisions), and two task-specific fine-tuning approaches applied to Qwen2.5-1.5B-Instruct: action-level supervised fine-tuning (A-SFT) and process-level fine-tuning (P-SFT), which directly optimizes process features. Broad fine-tuning on human decisions improves human-like task processes relative to off-the-shelf agents, while task-specific process-level supervision further improves behavioral mimicry. However, this advantage diminishes under cross-task transfer when supervised process targets do not naturally generalize across tasks. Explicit process-level supervision can improve human behavioral mimicry, but only if appropriate task-specific process representations are available, highlighting process specification as a bottleneck for achieving human-like cognitive processes in machines.
Humans flexibly adapt their reasoning strategies to the requirements of a given problem. Large language models (LLMs) have performed well on many cognitive tasks, however, it is unclear whether this accuracy is a result of pattern matching from training data or flexible reasoning. Here, we introduce a novel paradigm to test this question: the riddle riddle paradigm. Riddle riddles are word problems written to mimic popular riddles, but altered so their answers only require literal interpretations. Identifying correct answers requires looking past the structure of each question and flexibly apply different reasoning strategies based on the content. If LLMs respond to surface features, such as form, a riddle-like structure should cause models to use an inventive reasoning strategy even when a literal interpretation suffices. Alternatively, if LLMs reason based on content, they should flexibly switch strategies when appropriate. Across two experiments with nine state-of-the-art LLMs and 100 human participants, we show humans and LLMs fail on this paradigm in opposite directions. LLMs were far more accurate on genuine riddles than on riddle riddles (84.9
A central goal of cognitive modeling is to develop models that not only predict human behavior but also provide insight into the underlying cognitive mechanisms. While neural network models trained on large-scale behavioral data often achieve strong predictive performance, they typically fall short in offering interpretable explanations of the cognitive processes they capture. In this work, we explore the potential of pretrained large language models (LLMs) to serve as dual-purpose cognitive models--capable of both accurate prediction and interpretable explanation in natural language. Specifically, we employ reinforcement learning with outcome-based rewards to guide LLMs toward generating explicit reasoning traces for explaining human risky choices. Our findings demonstrate that this approach produces high-quality explanations alongside strong quantitative predictions of human decisions.
Some of the strongest evidence that human minds should be thought of in terms of symbolic systems has been the way they combine ideas, produce novelty, and learn quickly. We argue that modern neural networks—and the artificial intelligence systems built upon them—exhibit similar abilities. This potentially undermines the argument that the cognitive processes and representations used by human minds are symbolic. We consider possible interpretations of these results—that modern neural networks implement symbolic systems, or that they approximate them subsymbolically—and the theoretical consequences of these two possibilities for explanations of human cognition at different levels of analysis. This consideration leads us to offer a new agenda for research on the symbolic basis of the mind.
Magic tricks provide a uniquely powerful window onto the boundaries of what people consider possible, and the ways in which those boundaries can be challenged. Although there exists an effectively infinite variety of magic tricks, relatively little is known about how different magical effects relate to one another at a psychological level. Previous attempts to classify magic have largely relied on the judgments of a small number of expert magicians, rather than capturing people's intuitive understanding of magical phenomena. Here, we introduce a new multi-dimensional taxonomy of magic that treats the observable effect (e.g., an object vanishing) and the attributed cause (e.g., a magic wand) as independent dimensions. This framework allows a wide range of magic tricks to be systematically described and captures the vast majority of known magical effects. We evaluated the psychological validity of this taxonomy by asking participants to judge the perceived similarity between magic tricks drawn from each category. These similarity judgments were used to construct a dendrogram, revealing a hierarchical structure that reflects how people cognitively organize magical effects. The results corroborate several key distinctions traditionally made by magicians, such as the separation between mental and physical effects, while also revealing more subtle patterns in the perception of magic that have not previously been documented. Together, these findings provide a psychologically grounded framework for understanding how people categorize and experience magical impossibilities.
Even competitive activities require cooperation to agree on the rules. How do people navigate this tension, and what conditions help them reach agreement? We study rule negotiation in a collaborative game design task (N = 184) where dyads propose and vote on what game to play, crossing two factors: communication (chat available vs.~vote-only) and role uncertainty (whether players know which role they will occupy). We find communication increased agreement rates and role uncertainty had only a minimal impact on agreement. When players knew their role, they proposed games that subtly favored themselves, balancing self-interest against the need for acceptance. Chat content was largely coordinative rather than argumentative: players declared preferences and deferred rather than playing hardball. Our findings highlight how deliberation helps people coordinate on rules, even when they know which side they're on.
Today's large language models (LLMs) are trained to align with user preferences through methods such as reinforcement learning. Yet models are beginning to be deployed not merely to satisfy users, but also to generate revenue for the companies that created them through advertisements. This creates the potential for LLMs to face conflicts of interest, where the most beneficial response to a user may not be aligned with the company's incentives. For instance, a sponsored product may be more expensive but otherwise equal to another; in this case, what does (and should) the LLM recommend to the user? In this paper, we provide a framework for categorizing the ways in which conflicting incentives might lead LLMs to change the way they interact with users, inspired by literature from linguistics and advertising regulation. We then present a suite of evaluations to examine how current models handle these tradeoffs. We find that a majority of LLMs forsake user welfare for company incentives in a multitude of conflict of interest situations, including recommending a sponsored product almost twice as expensive (Grok 4.1 Fast, 83
Creativity is the ability to come up with novel ideas, a capacity crucial for human development and flourishing. Are large language models (LLMs) creative in the same way humans are, and can the same interventions increase creativity in both? We study a promising but largely untested intervention for creativity: forcing creators to draw an analogy from a random, remote source domain (”cross-domain mapping”). Human participants and LLMs generated novel designs for ten daily products (e.g., backpack, TV) under two prompts: (i) cross-domain mapping, which required drawing inspiration from a randomly assigned source (e.g., octopus, cactus, GPS), and (ii) user need, which required proposing innovations targeting unmet user needs. We show that humans reliably benefit from randomly assigned cross-domain mappings, while LLMs, on average, generate more original ideas than humans and do not show a statistically significant effect of cross-domain mappings. However, in both systems, the impact of cross-domain mapping increases when the inspiration source becomes more semantically distant from the target, and above a distance threshold, cross-domain mapping benefits the most capable LLMs too. Humans and LLMs differed in how they used the source concept: Humans tended to transfer surface features of the source, whereas LLMs transferred structural and functional properties. Our results highlight the role of remote association in creative ideation and systematic differences in how humans and LLMs respond to the same intervention for creativity.
Post-training Large Language Models (LLMs) for reasoning typically focuses on deductive tasks such as mathematics and coding where correctness is verifiable. Yet, many real-world reasoning problems are inductive: agents must infer uncertain beliefs from sparse, ambiguous observations. There are challenges to using standard fine-tuning methods for inductive reasoning, including difficulties in curating large-scale, high-quality labeled datasets and in handling targets that are inherently distributional. In this work, we introduce a novel approach, called Program-based Posterior Training (PPT), to address these limitations: we use an LLM to generate diverse open-world scenarios as probabilistic programs, run probabilistic inference to produce distributional target responses to queries, and then fine-tune on these probabilistic soft labels. Using this approach, we fine-tune LLMs on 10,000 programmatically generated scenarios and evaluate on held-out motifs, human-labeled judgments, and external benchmarks. Overall, PPT substantially improves estimation accuracy on held-out inductive tasks, increases alignment with human judgments, and transfers to external benchmarks for estimation and calibration. Additionally, the gains in raw calibration are not subsumed by post-hoc temperature scaling, showing that the models have more deeply internalized uncertainty compared to output rescaling. Together, these results suggest that probabilistic-program-mediated fine-tuning is a promising approach for post-training LLMs to reliably perform approximate inductive inference.
Psychologists have long been fascinated with understanding the nature of Aha! moments, moments when we transition from not knowing to suddenly realizing the solution to a problem. In this work, we present a theoretical framework that explains why we experience Aha! moments. Our theory posits that during problem-solving, in addition to solving the problem, people also maintain a metacognitive model of their ability to solve the problem as well as a prediction about the time it would take them to solve that problem. Aha! moments arise when we experience a positive error in this metacognitive prediction, i.e. when we solve a problem much faster than we expected, signaling to ourselves that we are more competent than we had realized. We posit that this metacognitive error is analogous to a positive reward prediction error thereby explaining why we feel so good after an Aha! moment. We provide support to our theory across two large-scale pre-registered experiments on problem solving, demonstrating a link between metacognitive prediction errors and Aha! moments. These results highlight the importance of metacognitive prediction errors and deepen our understanding of human metareasoning.
Generative AI systems like foundation models (FMs) must align well with human values to ensure their behavior is helpful and trustworthy. While Reinforcement Learning from Human Feedback (RLHF) has shown promise for optimizing model performance using human judgments, existing RLHF pipelines predominantly rely on immediate feedback, which can fail to accurately reflect the downstream impact of an interaction on users' utility. We demonstrate that feedback based on evaluators' foresight estimates of downstream consequences systematically induces Goodhart's Law dynamics, incentivizing misaligned behaviors like sycophancy and deception and ultimately degrading user outcomes. To alleviate this, we propose decoupling evaluation from prediction by refocusing RLHF on hindsight feedback. Our theoretical analysis reveals that conditioning evaluator feedback on downstream observations mitigates misalignment and improves expected human utility, even when these observations are simulated by the AI system itself. To leverage this insight in a practical alignment algorithm, we introduce Reinforcement Learning from Hindsight Simulation (RLHS), which first simulates plausible consequences and then elicits feedback to assess what behaviors were genuinely beneficial in hindsight. We apply RLHS to two widely-employed online and offline preference optimization methods – Proximal Policy Optimization (PPO) and Direct Preference Optimization (DPO) – and show empirically that misalignment is significantly reduced with both methods. Through an online human user study, we show that RLHS consistently outperforms RLHF in helping users achieve their goals and earns higher satisfaction ratings, despite being trained solely with simulated hindsight feedback. These results underscore the importance of focusing on long-term consequences, even simulated ones, to mitigate misalignment in RLHF.
AI systems based on artificial neural networks are being developed with aspirations of pushing the boundary of human mathematical knowledge. A key question for these systems is how much they can reach beyond their training data. Mathematical discovery requires a strong form of out of distribution generalization; the ability to hypothesize genuinely new - and potentially logically more powerful - mathematical structures. It has been hypothesized that language abilities support such generalizations in human cognition. In this work, we use simple arithmetic as a case study for examining how modern AI models could expand their mathematical horizons, evaluating whether these models can independently discover the concept of "zero". We show that We show that (1) language models of a GPT-2 size are unable to perform this generalization at test time regardless of language pretraining, but (2) models can improve substantially after training on tens or hundreds of examples of zero. Additionally, we find that language pretraining reduces the number of required examples by approximately $50\%$, showing that language abilities can scaffold mathematical discovery in neural models.
Determining the degree of similarity between different stimuli is central for intelligent behavior. Psychologists have long debated the nature of similarity, considering whether it is continuous and space-like or discrete and feature-like. Here, we argue that the idea that stimuli are similar to the extent that they are likely to have been generated by the same process unifies these spatial and featural accounts as special cases. We formulate this notion of generative similarity in terms of Bayesian inference over hierarchical generative processes and show how it can illuminate key ideas from the literature including the universal law of generalization, feature contrast models, exemplar and prototype theories, and metric violations in semantic organization. Moreover, just as previous work on similarity drew on connections to multi-dimensional scaling and additive clustering, generative similarity provides a new way to construct representations of stimuli by drawing on contrastive learning, a modern machine learning procedure for learning representations by pushing similar stimuli together and dissimilar stimuli apart. We show how this approach can provide a way to translate hierarchical generative processes into spatial representations that support human-like abstraction and generalization, even in complex domains that are parameterized by probabilistic programs.
Games have long been a microcosm for studying planning and reasoning in both natural and artificial intelligence, often focusing on expert-level or even super-human play1-6. But real life also pushes human intelligence along a different frontier, requiring people to flexibly navigate decision-making problems that they have never thought about before. Here we use novice gameplay to study how people reason about new problem settings. Through a series of large-scale behavioural studies with over 1,000 participants and 121 two-player strategic board games (almost all novel to our participants), we show that people are systematic and adaptively rational in how they play a game for the first time or evaluate a game (for example, how fair or how fun it is likely to be) before they have played it even once. We explain these capacities via a computational cognitive model that we call the 'Intuitive Gamer': a model based on mechanisms of fast and flat (depth-limited) goal-directed probabilistic simulation. Our work offers insights into how people rapidly evaluate, act and make suggestions when encountering novel problems, and could inform the design of more flexible and human-like artificial intelligence systems that can determine not just how to solve new tasks but also whether a task is worth thinking about at all.
A fundamental property of memory is its decline over time, a process widely observed to follow a power law. Two primary accounts have been proposed to explain this phenomenon. A rational account considers the power law of forgetting arising from optimal adaptation to environmental statistics, but does not specify how human memory carries out such adaptation. A mechanistic account considers the power law of forgetting arising from a weighted combination of exponential trace decays, constrained by the immutable structure of memory, but it remains unclear why these weights are combined in this particular way. To reconcile these two accounts, we propose that the power law of forgetting emerges from both environmental demands and architectural constraints. Our theoretical analysis and simulations demonstrate that the memory system's architectural constraints can support estimating how quickly information decays in the environment, aligning its own rate of forgetting to match that of the environment. This adaptation would ensure that information is forgotten quickly when it becomes irrelevant, and retained longer when it remains useful. This provides a cognitively plausible model that explains past empirical results where human memory is observed to adapt to changing environmental statistics on a short timescale. Together, these results offer a unified, mechanistic, and rational understanding of what gives rise to the shape of memory forgetting.
Teaching is a foundational social behaviour that can result from mentally effortful reasoning or cognitively frugal heuristics. When do people use these strategies to teach? Here we investigated this using behavioural experiments and computational modelling with adult participants recruited via Prolific. Experiment 1 (N = 100) revealed robust individual differences: some participants taught by reasoning about a learner's knowledge, consistent with an optimal Bayesian pedagogy model, while others relied on simple heuristics that do not require mentalizing. In two preregistered follow-up experiments, we found that people persist in using heuristics even when they are no longer effective (experiment 2, N = 253, P < 0.001, rank-biserial r = 0.287, 95% confidence interval 0.149-0.419) but that this tendency is pre-empted when inference about a learner's knowledge is scaffolded using an auxiliary task (experiment 3, N = 759, P < 0.001, partial η p 2 = 0 . 107 , 95% confidence interval 0.068-0.148). These results demonstrate sophisticated arbitration between planning and heuristics during teaching and elucidate the more general mechanisms involved in adapting mental effort during social interactions.
Four-term word analogies (A:B::C:D) are classically modeled geometrically as parallelograms: adding the vector B-A+C produces D. Recent work suggests that this model poorly captures how humans produce analogies, with simple local-similarity heuristics often providing a better account (Peterson et al., 2020). But does the parallelogram model fail because it is a bad model of analogical relations, or because people are not very good at generating relation-preserving analogies? We compared human and large language model (LLM) analogy completions on the set of problems from Peterson et al. (2020). We find that LLM-generated analogies are reliably judged as better than human-generated ones, and are also more consistent with parallelograms in a distributional embedding space. Crucially, we show that the improvement over human analogies is driven by greater parallelogram alignment and reduced reliance on accessible words rather than enhanced sensitivity to local similarity. Finally, fine-tuning GloVe to better satisfy the parallelogram constraint makes the model's top-ranked candidates more likely to be the completions humans and LLMs actually produced, and improves its prediction of human ratings. Overall, these results provide support for the parallelogram model of word analogies.
Rigorously evaluating machine intelligence against the broad spectrum of human general intelligence has become increasingly important and challenging in this era of rapid technological advance. Conventional AI benchmarks typically assess only narrow capabilities in a limited range of human activity. Most are also static, quickly saturating as developers explicitly or implicitly optimize for them. We propose that a more promising way to evaluate human-like general intelligence in AI systems is through a particularly strong form of general game playing: studying how and how well they play and learn to play all conceivable human games, in comparison to human players with the same level of experience, time, or other resources. We define a "human game" to be a game designed by humans for humans, and argue for the evaluative suitability of this space of all such games people can imagine and enjoy – the "Multiverse of Human Games". Taking a first step towards this vision, we introduce the AI GameStore, a scalable and open-ended platform that uses LLMs with humans-in-the-loop to synthesize new representative human games, by automatically sourcing and adapting standardized and containerized variants of game environments from popular human digital gaming platforms. As a proof of concept, we generated 100 such games based on the top charts of Apple App Store and Steam, and evaluated seven frontier vision-language models (VLMs) on short episodes of play. The best models achieved less than 10% of the human average score on the majority of the games, and especially struggled with games that challenge world-model learning, memory and planning. We conclude with a set of next steps for building out the AI GameStore as a practical way to measure and drive progress toward human-like general intelligence in machines.
A significant challenge for Bayesian models of cognition is understanding how the abstract assumptions behind these models connect to psychological mechanisms. We examine the consequences of considering one of the psychological processes involved in inductive inference: generating hypotheses. We analyze the predictions of a simple model that separates the processes of generating hypotheses and evaluating those hypotheses. This analysis shows that the ease of generating a hypothesis and its plausibility are confounded if we simply analyze people’s behavior in terms of Bayesian inference without taking the underlying mechanisms into account. We then establish through three experiments that the processes of generation and evaluation are separable, and that influencing the ease of hypothesis generation has different consequences from simply manipulating the plausibility of a hypothesis. These experiments produce results that could not be accounted for by a Bayesian model that assumes people assign a fixed prior probability to hypotheses.