Do Large Language Models (LLMs) possess a Theory of Mind (ToM)? Research into this question has focused on evaluating LLMs against benchmarks and found success across a range of social tasks. However, these evaluations do not test for the actual representations posited by ToM: namely, a causal model of mental states and behavior. Here, we use a cognitively-grounded definition of ToM to develop and test a new evaluation framework. Specifically, our approach probes whether LLMs have a coherent, domain-general, and consistent model of how mental states cause behavior -- regardless of whether that model matches a human-like ToM. We find that even though LLMs succeed in approximating human judgments in a simple ToM paradigm, they fail at a logically equivalent task and exhibit low consistency between their action predictions and corresponding mental state inferences. As such, these findings suggest that the social proficiency exhibited by LLMs is not the result of an domain-general or consistent ToM.
People’s actions often leave physical traces (e.g., footprints in the mud, a dirty mug) that adults can use to reconstruct what another person was doing, what their goals were, and what they knew. Here we test whether children can make these inferences in cases that require going beyond a direct mapping from traces to actions, focusing on reasoning about a conspicuous absence of evidence, and relying on an agent’s preferences to interpret physical traces that are otherwise ambiguous. In Experiment 1, five- and six-year-olds inferred that an agent had lifted a rock-filled bowl or a rice-filled bowl when they saw the corresponding physical trace (i.e., rocks or rice spilled on the table). Critically, when no traces were left behind, children also inferred that the agent had lifted the rock-filled bowl, suggesting they understood that this action was less likely to produce observable evidence. Experiments 2 and 3 then show that children can integrate the agent’s preferences when the physical evidence alone does not reveal what happened. Our results suggest that, by age five, children begin to draw social inferences from indirect evidence by integrating physical and social reasoning.
Everyday social interactions require us to infer the cognitive processes happening in other minds—their reasoning, distractions, and recall. Despite the importance of these inferences, computational models of social reasoning focus exclusively on attributions of mental states like knowledge and preferences, leaving the mechanisms underlying cognitive process inference largely uncharacterized. Here we introduce Bayesian Inverse Reasoning (BIR), a new computational framework for social cognition that represents other minds as computational systems and performs Bayesian inference over abstract representations of computation to reconstruct the cognitive processes unfolding in other minds. Across six experiments, we show that this framework quantitatively captures how people make graded inferences about what others are thinking about (Experiment 1), their reasoning speed (Experiment 2), whether they engaged in recall (Experiment 3), and whether they were distracted (Experiment 4). These inferences were fine-grained and easily evoked. We further used the BIR framework to probe the algorithmic nature of this capacity, revealing that people use first-person thinking to estimate the computational demands of different problems (Experiment 5), yet integrate these self-derived estimates into a causal model of other minds in a way that produces unbiased inferences about others (Experiment 6). Together, these findings reveal a previously unknown cognitive architecture underlying human social cognition: one in which people represent minds not merely as containers of beliefs and desires, but as computational engines engaged in dynamic, flexible information processing.
When talking about the world in front of us, humans are remarkably efficient communicators. Our referential expressions help listeners find what we’re talking about by strategically adding adjectives as needed. But most conversations are about things that are not physically in front of us. Does this efficiency break down when referents are in the mind (rather than on the table), or are we also able to efficiently help a listener retrieve an item from memory? Across three experiments, we asked participants to describe images to help a listener recall each image. In Experiment 1 (total n = 600), participants showed efficiency, spontaneously incorporating expectations about memorability by providing relatively longer descriptions for images that people expect to be less memorable. Participants did not need access to or knowledge of their listener’s prior experience to adjust efficiently (Experiment 2, n = 300), but when shared experience directly modulated expected efficiency participants used that context accordingly (Experiment 3, n = 300). Interestingly, people’s descriptions were more aligned with subjective estimates of memorability, rather than objective, empirically-derived metrics. Together, this work provides new evidence that speakers spontaneously guide listeners’ mental processes to effectively facilitate recall.
An unwritten expectation in our everyday social interactions is that intimate personal information about someone-"insider knowledge"-is usually confined within close relationships. For example, it would be odd, or even unsettling, if a stranger knew about your favorite movie. Such expectations about who knows what about whom constitute a cornerstone of complex social behavior, but much remains unknown about their cognitive underpinnings and developmental origins. Drawing on parental report (Study 1) as well as a novel experimental approach using controlled but naturalistic videochat conversations (Study 2 & 3), we find that 4- to 5-y-old children have an abstract, theory-like understanding of how social connections give rise to interpersonal knowledge. Self-report, facial expressions, and memory errors provide converging evidence that children were surprised when someone possessed insider knowledge that is misaligned with their relationships, such as a stranger knowing their favorite food (Study 2a) or their own parent knowing a stranger's favorite movie (Study 2b). Children also generated coherent ad-hoc explanations about how someone might have acquired that knowledge, appealing to either first-hand observations or second-hand sources (Study 3). These findings demonstrate an early-emerging understanding of how individual minds are shaped in the context of their social networks, supporting a precocious ability to detect and explain anomalies in what people know about each other in real-time conversations. The current work also opens possibilities for leveraging open-ended online interactions to study social cognition without compromising experimental control.
Where someone looks, and for how long, is a remarkably rich signal about their mind. It reveals what they find interesting, what they like, and what is new to them. These inferences are possible because as adults, we understand that people don’t just encode whatever is visible to them. This is instead modulated by what they choose to attend to, and for how long. Across four experiments (N = 220), we asked whether children understand these nuances. We found that unlike adults, children showed only partial expectations about what draws attention, expecting longer looking for larger sets by age six, but not for more complex or more varied objects (Exp. 1). However, children as young as five understood that an agent who looked longer formed a more accurate representation of what they saw (Exp. 2), and that they are more likely to look longer at things they find desirable (Exp. 3). Finally, six-year-olds, but not five- year-olds, inferred that agents who already know about a toy are less likely to look for an extended duration (Exp. 4). Together, these results suggest that an intuitive theory of attention is partly in place by age five and continues to develop past age six.
Human social life unfolds within richly structured networks of overlapping relationships, including friendships, hierarchies, and collaborations. Yet the observable interactions that reveal these networks are often sparse and noisy, making it unclear how people could infer the latent structure of their social environments from such limited evidence. We propose that humans integrate domain-general statistical learning with domain-specific models of social structures to rapidly construct causal representations that support explanation, prediction, and planning. Across three behavioral experiments, we show that participants can infer underlying social structures (Experiment 1), predict social behavior (Experiment 2), and reason about the spread of social influence (Experiment 3), based on brief, abstract videos of social interactions. These judgments were closely captured by a computational model grounded in our account and could not be explained by simpler cue-based accounts. Statistical learning and causal reasoning operate in concert to support rapid, flexible understanding of social structures.
Everyday life requires that we navigate dozens of interactions with strangers quickly and effectively. Here we hypothesized that representations of roles (e.g., cashier, mechanic, doctor) enable people to build rapid expectations about what others will do, what they know, and who else serves the same function. To test this, we used a self-paced reading paradigm in which we timed how long it took participants to read short passages about social interactions. Across three studies (N=300 participants per study), we show that role representations support real-time expectations about how other people might act (Study 1), the knowledge they might possess (Study 2), and whether it is appropriate to generalize across agents (Study 3). Moreover, people reported more surprise when the events deviated from role expectations and were more likely to misreport what happened in a way that conformed to role expectations. Our results suggest that roles are a powerful route for social understanding that has been previously understudied in social cognition. Using a self-paced reading paradigm, we show that, from just the mention of a role, people build rapid expectations about how other people will act, what they know, and whether they can be interchangeable with others in the same role.
Vision-Language models (VLMs) can now produce fluent, socially appropriate dialogue, leading to interest in whether they have Theory of Mind (ToM). While recent work suggests that VLMs still lack coherent mental-state reasoning, this work has focused on classical propositional belief representations. Here we test a complementary, communication-relevant form of ToM: attention-based social micro-processes that support referential communication in the here and now (i.e., selecting efficient descriptions to identify a particular object for someone else). In face-to-face communication, people strategically add redundant color adjectives to facilitate the listener’s visual search, and omit them when they provide no benefit. We evaluate whether VLMs show the same strategy in a referential communication paradigm where the usefulness of redundant color words varies based on the set size and color distribution of objects. We find that VLMs can produce successful referential expressions but lack the attention-guiding strategies that make human communication so efficient. This suggests that VLMs lack the more implicit representations of attention people use in everyday communication.
When talking about the world in front of us, humans are remarkably efficient communicators. Our referential expressions help listeners efficiently find what we’re talking about by strategically adding color or material words as needed. But most conversations are about things that are not physically in front of us. In these cases, do we also use language to efficiently help a listener retrieve an item from memory? Across two experiments, we asked participants to describe images to help a listener recall each image. In Experiment 1 (total n = 600), participants spontaneously incorporated expectations about memorability, providing relatively more description for images that people expect to be less memorable. In Experiment 2 (n = 300), we replicated this pattern even when participants had no access to or knowledge of their listener’s prior experience. Interestingly, people’s descriptions were more aligned with subjective estimates of memorability, rather than objective, empirically-derived metrics. Together, this work provides new evidence that speakers spontaneously guide listeners’ mental processes to effectively facilitate their memory recall.
Moffett points to humans’ use of physical markers to signal group identity as crucial to human society. We characterize the developmental and cognitive bases of this capacity, arguing that it is part of an early-emerging, intuitive socio-physical interface which allows the inanimate world to encode rich social meaning about individuals’ identities, and the values of the society as a whole.
Understanding the relationship between seeing and knowing is fundamental to social cognition. While research demonstrates that even infants grasp basic aspects of this relationship, prior work often treats perceptual access and knowledge as equivalent (e.g., "if you see it, you know it"). In reality, their connection is richer: more complex objects require longer to encode, and agents’ looking patterns often reveal how well they have encoded something and how much they want it. Across three experiments, we investigated whether children understand these nuances. In Experiment 1, we found that by age six, children expect more objects to require longer looking times. In Experiment 2, children inferred that agents who looked longer were more likely to form accurate representations of what they observed. In Experiment 3, children reasoned that agents who looked longer at an object were more likely to want it. Together, these findings suggest that by age six, children develop an intuitive theory of attention, enabling them to make sophisticated inferences about others' mental states based on looking behaviors.
Students often receive encouragement but do not always find it motivating. Whose encouragement motivates students and what cognitive mechanisms underlie this process? We propose that students' responses to positive feedback (e.g., encouragement) hinge on mental state representations, specifically what the speaker knows. Across three studies, we find that U.S. adolescents (n = 581-759 11- to 19-year-olds per study, preregistered; > 80% racial/ethnic minorities; > 36% low income) report being more motivated by, more confident in, and more likely to seek out encouragement from hypothetical and real-world speakers (e.g., parents, teachers, peers) who are knowledgeable about both their abilities (e.g., students' math skills) and the task at hand (e.g., math). To make feedback most effective, our findings suggest that students should seek and receive encouragement from those who know them and their activities well. (PsycInfo Database Record (c) 2025 APA, all rights reserved).
Human success in navigating the social world is typically attributed to our capacity to represent other minds-a mentalistic stance. We argue that humans are endowed with a second equally powerful intuitive theory: an institutional stance. In contrast to the mentalistic stance, which helps us predict and explain unconstrained behavior via unobservable mental states, the institutional stance interprets social interactions in terms of role-based structures that constrain and regulate behavior via rule-like behavioral expectations. We argue that this stance is supported by a generative grammar that builds structured models of social collectives, enabling people to rapidly infer, track, and manipulate the social world. The institutional stance emerges early in development and its precursors can be traced across social species, but its full-fledged generative capacity is uniquely human. Once in place, the ability to reason about institutional structures takes on a causal role, allowing people to create and modify social structures, supporting new forms of institutional life. Human social cognition is best understood as an interplay between a system for representing the unconstrained behavior of individuals in terms of minds and a system for representing the constrained behavior of social collectives in terms of institutional structures composed of interlocking sets of roles.
When determining what others know, we intuitively consider not only whether they succeed but also their probability of success in the absence of knowledge (e.g., random guessing). Across three experiments ( n = 240 North American 4–6-year-olds, data collected between 2020–2023) we find that 4-year-olds understand that tasks with a lower probability of chance success are harder. However, it is not until age 6 that children use this understanding to gauge (Experiment 1) and infer (Experiments 2–3) what others know. These results suggest that, although basic probabilistic reasoning and representations of knowledge are well in place by age 4, children do not integrate the two to make mental-state inferences until much later, pointing to an area of important developmental change in Theory of Mind.
Beneficial societal outcomes cannot be guaranteed by aligning individual AI systems with the intentions of their operators or users. Even an AI system that is perfectly aligned to the intentions of its operating organization can lead to bad outcomes if the goals of that organization are misaligned with those of other institutions and individuals. For this reason, we need full-stack alignment, the concurrent alignment of AI systems and the institutions that shape them with what people value. This can be done without imposing a particular vision of individual or collective flourishing. We argue that current approaches for representing values, such as utility functions, preference orderings, or unstructured text, struggle to address these and other issues effectively. They struggle to distinguish values from other signals, to support principled normative reasoning, and to model collective goods. We propose thick models of value will be needed. These structure the way values and norms are represented, enabling systems to distinguish enduring values from fleeting preferences, to model the social embedding of individual choices, and to reason normatively, applying values in new domains. We demonstrate this approach in five areas: AI value stewardship, normatively competent agents, win-win negotiation systems, meaning-preserving economic mechanisms, and democratic regulatory institutions.
When determining what others know, we intuitively consider not only whether they succeed but also their probability of success in the absence of knowledge (e.g., random guessing). Across three experiments (n = 240 North American 4-6-year-olds, data collected between 2020-2023) we find that 4-year-olds understand that tasks with a lower probability of chance success are harder. However, it is not until age 6 that children use this understanding to gauge (Experiment 1) and infer (Experiments 2-3) what others know. These results suggest that, although basic probabilistic reasoning and representations of knowledge are well in place by age 4, children do not integrate the two to make mental-state inferences until much later, pointing to an area of important developmental change in Theory of Mind.
People have a remarkable ability to infer the hidden causes of things. From physical evidence, such as muddy foot prints on the floor, we can figure out what happened and who did it. Here, we investigate another source of evidence: social evaluations. Social evaluations, such as praise or blame, are commonplace in everyday conversations. While such evaluations don't fully reveal what happened, they provide valuable clues. Across three experiments, we present situations where a person was praised or blamed, and participants' task is to use that information to figure out what happened. In Experiment 1, we find that people draw systematic inferences from social evaluations about situational factors, a person's actions, capabilities, and social roles. In Experiments 2 and 3 we develop computational models that generate praise and blame judgments by considering what causal role a person's action played, and what action they should have taken. Inverting these generative models of praise and blame via Bayesian inference yields accurate predictions about what inferences participants draw based on social evaluations.
Social biases are prevalent in everyday social interactions, but they are often expressed in subtle ways that can make them difficult to detect. Yet, intuitively, people can often recognize when they are the subject of a bias, even in the absence of any overt behavior. How do we do this? While much research has focused on the negative consequences of being the subject of a bias, less is known about the cognitive mechanisms that allow people to explicitly detect biases in the first place. In this paper, we propose an account of bias detection which is grounded on mental state representations. We propose that people infer biases by detecting a gap between expected unbiased behavior and observed real-world behavior, which in turns reveals the hidden biases influencing other people's beliefs. We present a formal computational model of this account and, across three preregistered studies (n=720 total), we show that this model captures participants' inferences about an observer's prior beliefs (Experiment 1), general social biases (Experiment 2), and specific real-world biases (Experiments 3a--3c). Moreover, our model captures key patterns of variance in participant responses which simpler alternative models fail to capture. These findings highlight the role of Theory of Mind in social bias detection, and broaden our understanding of the human capacity to detect and reason about implicit prejudices.