Strategic decision-making is a crucial component of human interaction. Here we conduct a large-scale study of strategic decision-making in the context of initial play in two-player matrix games, analysing over 90,000 human decisions across more than 2,400 procedurally generated games that span a much wider space than previous datasets. We show that a deep neural network trained on this dataset predicts human choices with greater accuracy than leading theories of strategic behaviour, revealing systematic variation unexplained by existing models. By modifying this network, we develop an interpretable behavioural model that uncovers key insights: individuals' abilities to respond optimally and reason about others' actions are highly context dependent, influenced by the complexity of the game matrices. Our findings illustrate the potential of machine learning as a tool for generating new theoretical insights into complex human behaviours.
Establishing a unified theory of cognition has been an important goal in psychology1,2. A first step towards such a theory is to create a computational model that can predict human behaviour in a wide range of settings. Here we introduce Centaur, a computational model that can predict and simulate human behaviour in any experiment expressible in natural language. We derived Centaur by fine-tuning a state-of-the-art language model on a large-scale dataset called Psych-101. Psych-101 has an unprecedented scale, covering trial-by-trial data from more than 60,000 participants performing in excess of 10,000,000 choices in 160 experiments. Centaur not only captures the behaviour of held-out participants better than existing cognitive models, but it also generalizes to previously unseen cover stories, structural task modifications and entirely new domains. Furthermore, the model's internal representations become more aligned with human neural activity after fine-tuning. Taken together, our results demonstrate that it is possible to discover computational models that capture human behaviour across a wide range of domains. We believe that such models provide tremendous potential for guiding the development of cognitive theories, and we present a case study to demonstrate this.
In order for AI systems to communicate effectively with people, they must understand how we make decisions. However, people's decisions are not always rational, so the implicit internal models of human decision-making in Large Language Models (LLMs) must account for this. Previous empirical evidence seems to suggest that these implicit models are accurate --- LLMs offer believable proxies of human behavior, acting how we expect humans would in everyday interactions. However, by comparing LLM behavior and predictions to a large dataset of human decisions, we find that this is actually not the case: when both simulating and predicting people's choices, a suite of cutting-edge LLMs (GPT-4o \& 4-Turbo, Llama-3-8B \& 70B, Claude 3 Opus) assume that people are more rational than we really are. Specifically, these models deviate from human behavior and align more closely with a classic model of rational choice --- expected value theory. Interestingly, people also tend to assume that other people are rational when interpreting their behavior. As a consequence, when we compare the inferences that LLMs and people draw from the decisions of others using another psychological dataset, we find that these inferences are highly correlated. Thus, the implicit decision-making models of LLMs appear to be aligned with the human expectation that other people will act rationally, rather than with how people actually act.
Predicting human decisions under risk and uncertainty remains a fundamental challenge across disciplines. Existing models often struggle even in highly stylized tasks like choice between lotteries. Here we introduce BEAST gradient boosting (BEAST-GB), a hybrid model integrating behavioural theory (BEAST) with machine learning. We first present CPC18, a competition for predicting risky choice, in which BEAST-GB won. Then, using two large datasets, we demonstrate that BEAST-GB predicts more accurately than neural networks trained on extensive data and dozens of existing behavioural models. BEAST-GB also generalizes robustly across unseen experimental contexts, surpassing direct empirical generalization, and helps to refine and improve the behavioural theory itself. Our analyses highlight the potential of anchoring predictions on behavioural theory even in data-rich settings and even when the theory alone falters. Our results underscore how integrating machine learning with theoretical frameworks, especially those-like BEAST-designed for prediction, can improve our ability to predict and understand human behaviour.
Establishing a unified theory of cognition has been a major goal of psychology. While there have been previous attempts to instantiate such theories by building computational models, we currently do not have one model that captures the human mind in its entirety. Here we introduce Centaur, a computational model that can predict and simulate human behavior in any experiment expressible in natural language. We derived Centaur by finetuning a state-of-the-art language model on a novel, large-scale data set called Psych-101. Psych-101 reaches an unprecedented scale, covering trial-by-trial data from over 60,000 participants performing over 10,000,000 choices in 160 experiments. Centaur not only captures the behavior of held-out participants better than existing cognitive models, but also generalizes to new cover stories, structural task modifications, and entirely new domains. Furthermore, we find that the model’s internal representations become more aligned with human neural activity after finetuning. Taken together, Centaur is the first real candidate for a unified model of human cognition. We anticipate that it will have a disruptive impact on the cognitive sciences, challenging the existing paradigm for developing computational models.
Shepard's universal law of generalization is a remarkable hypothesis about how intelligent organisms should perceive similarity. In its broadest form, the universal law states that the level of perceived similarity between a pair of stimuli should decay as a concave function of their distance when embedded in an appropriate psychological space. While extensively studied, evidence in support of the universal law has relied on low-dimensional stimuli and small stimulus sets that are very different from their real-world counterparts. This is largely because pairwise comparisons -- as required for similarity judgments -- scale quadratically in the number of stimuli. We provide direct evidence for the universal law in a naturalistic high-dimensional regime by analyzing an existing dataset of 214,200 human similarity judgments and a newly collected dataset of 390,819 human generalization judgments (N=2406 US participants) across three sets of natural images.
The rapid development of machine learning has led to new opportunities for applying these methods to the study of human decision making. We highlight some of these opportunities and discuss some of the issues that arise when using machine learning to model the decisions people make. We first elaborate on the relationship between predicting decisions and explaining them, leveraging findings from computational learning theory to argue that, in some cases, the conversion of predictive models to interpretable ones with comparable accuracy is an intractable problem. We then identify an important bottleneck in using machine learning to study human cognition-data scarcity-and highlight active learning and optimal experimental design as a way to move forward. Finally, we touch on additional topics such as machine learning methods for combining multiple predictors arising from known theories and specific machine learning architectures that could prove useful for the study of judgment and decision making. In doing so, we point out connections to behavioral economics, computer science, cognitive science, and psychology.
Background: Cementless metaphyseal filling stems rely on fixation in the medial-to-lateral and anterior-to-posterior (AP) planes. The purpose of this preclinical study was to develop Insignia, a new metaphyseal filling system to match the anatomy of the proximal femur, and then compare it to clinically successful stems in multiple simulations. Methods: In this preclinical study, the geometry of the proximal femur in the AP plane among 1321 healthy subjects was evaluated using computed tomography. This data was then used to design insignia. Preclinical studies were performed to compare the broaching effort required to prepare a canal using this system, assess the reliability of seating heights for the stem, and compare in vitro micromotion testing of the stem under simulated stair climb activity. Results: The proximal femur decreased approximately 50% in the AP plane spanning 20 mm above the lesser trochanter to 30 mm below the lesser trochanter. Additional bench top testing was performed, and the new stem system was found to demonstrate significantly reduced broaching effort (average 6 vs 29 hits, P-value = .000), reliable seating heights on stem placement, and 70% less proximal micromotion on 10,000-cyclic testing (P < .05) compared to another clinically successful metaphyseal filling stem. Conclusions: The AP dimension of the proximal femur decreases nearly 50% throughout its length. Metaphyseal filling stems that match the AP anatomy of the proximal femur may require fewer hits during broaching, yield reproducible seating heights, and reduce micromotion on cyclic testing.
Delayed gratification is an important focus of research, given its potential relationship to forms of behavior, such as savings, susceptibility to addiction, and pro-social behaviors. The COVID-19 pandemic may be one of the most consequential recent examples of this phenomenon, with people's willingness to delay gratification affecting their willingness to socially distance themselves. COVID-19 also provides a naturalistic context by which to evaluate the ecological validity of delayed gratification. This article outlines four large-scale online experiments (total N = 12, 906) where we ask participants to perform Money Earlier or Later (MEL) decisions (e.g., $5 today vs. $10 tomorrow) and to also report stress measures and pandemic mitigation behaviors. We found that stress increases impulsivity and that less stressed and more patient individuals socially distanced more throughout the pandemic. These results help resolve longstanding theoretical debates in the MEL literature as well as provide policymakers with scientific evidence that can help inform response strategies in the future. (PsycInfo Database Record (c) 2023 APA, all rights reserved).
Deep neural networks are increasingly being used in cognitive modeling as a means of deriving representations for complex stimuli such as images. While the predictive power of these networks is high, it is often not clear whether they also offer useful explanations of the task at hand. Convolutional neural network representations have been shown to be predictive of human similarity judgments for images after appropriate adaptation. However, these high-dimensional representations are difficult to interpret. Here we present a method for reducing these representations to a low-dimensional space which is still predictive of similarity judgments. We show that these low-dimensional representations also provide insightful explanations of factors underlying human similarity judgments.
Supervised learning typically focuses on learning transferable representations from training examples annotated by humans. While rich annotations (like soft labels) carry more information than sparse annotations (like hard labels), they are also more expensive to collect. For example, while hard labels only provide information about the closest class an object belongs to (e.g., "this is a dog"), soft labels provide information about the object's relationship with multiple classes (e.g., "this is most likely a dog, but it could also be a wolf or a coyote"). We use information theory to compare how a number of commonly-used supervision signals contribute to representation-learning performance, as well as how their capacity is affected by factors such as the number of labels, classes, dimensions, and noise. Our framework provides theoretical justification for using hard labels in the big-data regime, but richer supervision signals for few-shot learning and out-of-distribution generalization. We validate these results empirically in a series of experiments with over 1 million crowdsourced image annotations and conduct a cost-benefit analysis to establish a tradeoff curve that enables users to optimize the cost of supervising representation learning on their own datasets.
Modern large-scale physics experiments create datasets with sizes and streaming rates that can exceed those from industry leaders such as Google Cloud and Netflix. Fully processing these datasets requires both sufficient compute power and efficient workflows. Recent advances in Machine Learning (ML) and Artificial Intelligence (AI) can either improve or replace existing domain-specific algorithms to increase workflow efficiency. Not only can these algorithms improve the physics performance of current algorithms, but they can often be executed more quickly, especially when run on coprocessors such as GPUs or FPGAs. In the winter of 2023, MIT hosted the Accelerating Physics with ML at MIT workshop, which brought together researchers from gravitational-wave physics, multi-messenger astrophysics, and particle physics to discuss and share current efforts to integrate ML tools into their workflows. The following white paper highlights examples of algorithms and computing frameworks discussed during this workshop and summarizes the expected computing needs for the immediate future of the involved fields.
Technician is an essential role in a medical procedure. A technician team is often faced with a high workload variability due to uncertain procedure case number and case duration. Traditional manual technician schedule with a fixed shift may not be able to address the variability, leading to technician shortage, extra cost on technician overtime and technician dissatisfaction. This paper is motivated by a technician scheduling and staffing problem in the Heart Rhythm Services. Historical data of medical procedures are evaluated and show high uncertainty of medical procedures, resulting in the daily technician demand distribution that has a peak in the morning and a long right tail. Hospital may employ excessive number of technicians but still experience technician shortage. An integer programming model is proposed to schedule technicians, and a technician scheduling tool with a user-friendly interface is developed. The tool is shown to be able to improve technician coverage and reduce the use of overtime. Besides, we also present the capability of the tool in supporting technician staffing and determining technician team configuration.
The diversity of human faces and the contexts in which they appear gives rise to an expansive stimulus space over which people infer psychological traits (e.g., trustworthiness or alertness) and other attributes (e.g., age or adiposity). Machine learning methods, in particular deep neural networks, provide expressive feature representations of face stimuli, but the correspondence between these representations and various human attribute inferences is difficult to determine because the former are highdimensional vectors produced via black-box optimization algorithms. Here we combine deep generative image models with over 1 million judgments to model inferences of more than 30 attributes over a comprehensive latent face space. The predictive accuracy of our model approaches human interrater reliability, which simulations suggest would not have been possible with fewer faces, fewer judgments, or lower-dimensional feature representations. Our model can be used to predict and manipulate inferences with respect to arbitrary face photographs or to generate synthetic photorealistic face stimuli that evoke impressions tuned along the modeled attributes.
Author(s): Dubey, Rachit; Peterson, Joshua | Abstract: The climate crisis is one of the most alarming issues of our time. Our planet is deteriorating at an unprecedented scale and accelerating rate, putting human societies and countless biological species in grave danger. The root cause of this problem is human behavior and thus it could prove crucial to examine the psychology behind the human behaviors that drive unsustainable living and impede enactment of climate policy. Unfortunately, despite the importance of psychological research in responding to the climate crisis, the field has had very little influence on the climate policy process as well as in mobilizing action on climate change, handicapping progress towards a sustainable future. This workshop aims to bring together scientists working in the broad area of climate change and sustainability, along with cognitive scientists, to engage in the development of ideas related to using cognitive science research in understanding, reducing, and responding to the climate crisis.
The remarkable successes of convolutional neural networks (CNNs) in modern computer vision are by now well known, and they are increasingly being explored as computational models of the human visual system. In this paper, we ask whether CNNs might also provide a basis for modeling higher-level cognition, focusing on the core phenomena of similarity and categorization. The most important advance comes from the ability of CNNs to learn high-dimensional representations of complex naturalistic images, substantially extending the scope of traditional cognitive models that were previously only evaluated with simple artificial stimuli. In all cases, the most successful combinations arise when CNN representations are used with cognitive models that have the capacity to transform them to better fit human behavior. One consequence of these insights is a toolkit for the integration of cognitively motivated constraints back into CNN training paradigms in computer vision and machine learning, and we review cases where this leads to improved performance. A second consequence is a roadmap for how CNNs and cognitive models can be more fully integrated in the future, allowing for flexible end-to-end algorithms that can learn representations from data while still retaining the structured behavior characteristic of human cognition.
Predicting and understanding how people make decisions has been a long-standing goal in many fields, with quantitative models of human decision-making informing research in both the social sciences and engineering. We show how progress toward this goal can be accelerated by using large datasets to power machine-learning algorithms that are constrained to produce interpretable psychological theories. Conducting the largest experiment on risky choice to date and analyzing the results using gradient-based optimization of differentiable decision theories implemented through artificial neural networks, we were able to recapitulate historical discoveries, establish that there is room to improve on existing theories, and discover a new, more accurate model of human decision-making in a form that preserves the insights from centuries of research.
The diversity in appearance of human faces and their naturalistic viewing conditions give rise to an expansive stimulus space over which humans perceive numerous psychological traits (e.g., perceived trustworthiness). Current scientific models characterize only few of these traits, and over only a tiny fraction of possible faces. Here we show that generative image models from machine learning combined with over 1 million human judgments can capture more than 30 traits over a near-infinite set of face stimuli. This makes it possible to then seamlessly infer and manipulate the psychological traits of arbitrary face photograph inputs and generate infinite synthetic photorealistic face stimuli along those dimensions. The predictive accuracy of our model approaches human inter-rater reliability, which our simulations suggest would not have been possible with previous datasets having fewer faces, fewer trait ratings, or using low-dimensional feature representations.
Whilst knowledge regarding the pathophysiology of congenital heart disease (CHDs) has advanced greatly in recent years, the underlying developmental processes affecting the cardiac outflow tract (OFT) such as bicuspid aortic valve, tetralogy of Fallot and transposition of the great arteries remain poorly understood. Common among CHDs affecting the OFT, is a large variation in disease phenotypes. Even though the different cell lineages contributing to OFT development have been studied for many decades, it remains challenging to relate cell lineage dynamics to the morphologic variation observed in OFT pathologies. We postulate that the variation observed in cellular contribution in these congenital heart diseases might be related to underlying cell lineage dynamics of which little is known. We believe this gap in knowledge is mainly the result of technical limitations in experimental methods used for cell lineage analysis. The aim of this review is to provide an overview of historical fate mapping and cell tracing techniques used to study OFT development and introduce emerging technologies which provide new opportunities that will aid our understanding of the cellular dynamics underlying OFT pathology.