The first problem of modeling the world is not just estimating the right parameters or causal structure, but deciding what should be represented at all. We frame this problem as continual model construction: an agent maintains an environment-specific model M of an inaccessible world W and curates a persistent library L of reusable representational elements across environments. We propose Representational Empowerment (RepEmp) to score candidate elements by how much they expand the agent's future capacity to model and plan, complementing the classic definition of empowerment, but redefined as control over internal representations instead of external states. We realize the framework as a hierarchical Curator-Actor architecture and test it across three experiments. In a closed-vocabulary causal-learning task, human participants construct causal models at varying abstraction granularities to maximize goal reachability rather than fidelity to the world, a signature better predicted by RepEmp than by information-gain alternatives. Matched simulations reveal that RepEmp-guided construction contributes more than exploration to sufficient structure recovery and cross-task transfer. Finally, in an open-vocabulary planning domain, an LLM-augmented Curator builds more compact symbolic libraries, which also generalize better than baselines. Ablating RepEmp eliminates these benefits. Together, these results identify RepEmp as a key principle for continual model construction: deciding what to build, retain, and reuse under bounded resources.
A long-standing finding in the causal learning literature is that adults struggle to identify conjunctive causal rules, where an effect requires the simultaneous presence of multiple causes, while performing better in disjunctive settings. However, most demonstrations of this “conjunctive handicap” rely on passive observation paradigms with limited evidence, where learners have no control over evidence generation. This paper asks whether this bias persists when adults are granted agency through active exploration. Using a modified “blicket detector” task, adult participants freely intervened to identify causal objects under conjunctive or disjunctive rule structures. We show that active exploration substantially improves adults' conjunctive causal reasoning, although conjunctive rules still require more tests to infer than disjunctive rules. We further compare human performance to a range of large language models in the same setting. While some state-of-the-art models approach human-level performance on hypothesis inference accuracy, they often exhibit less efficient exploration strategies and similar conjunctive-disjunctive performance gaps.
Learning about the causal structure of the world is a fundamental problem for human cognition. Causal models and especially causal learning have proved to be difficult for Large Models using standard techniques of deep learning. In contrast, cognitive scientists have applied advances in our formal understanding of causation in computer science, particularly within the Causal Bayes Net formalism, to understand human causal learning. In the very different tradition of reinforcement learning, researchers have described an intrinsic reward signal called “empowerment” which maximizes mutual information between actions and their outcomes. “Empowerment” may be an important bridge between classical Bayesian causal learning and reinforcement learning and may help to characterize causal learning in humans and enable it in machines. If an agent learns an accurate causal world model they will necessarily increase their empowerment, and increasing empowerment will lead to a more accurate causal world model. Empowerment may also explain distinctive empirical features of children’s causal learning, as well as providing a more tractable computational account of how that learning is possible. In an empirical study, we systematically test how children and adults use cues to empowerment to infer causal relations, design effective causal interventions and appropriately generalize to new contexts.
Here we explore whether adults and children differ in the way they search through a semantic network, and the role those differences may play in hypothesis generation. Participants generated sequential guesses about a novel causal relation. A week later, they completed a similarity spacing task and we measured the average “distance” between guesses. We find that adults show greater dependencies between sequential guesses than preschoolers, and generate a less diverse set of options. These findings may support the idea that development can be viewed as analogous to simulated annealing strategies in machine learning that start “hot” (in early childhood), generating wider and more variable searches, and eventually cool (in adulthood) to generate narrower searches.
This paper investigates visual analogical reasoning in large multimodal models (LMMs) compared to human adults and children. A "visual analogy" is an abstract rule inferred from one image and applied to another. While benchmarks exist for testing visual reasoning in LMMs, they require advanced skills and omit basic visual analogies that even young children can make. Inspired by developmental psychology, we propose a new benchmark of 4,300 visual transformations of everyday objects to test LMMs on visual analogical reasoning and compare them to children (ages three to five) and to adults. We structure the evaluation into three stages: identifying what changed (e.g., color, number, etc.), how it changed (e.g., added one object), and applying the rule to new scenarios. Our findings show that while GPT-o1, GPT-4V, LLaVA-1.5, and MANTIS identify the "what" effectively, they struggle with quantifying the "how" and extrapolating this rule to new objects. In contrast, children and adults exhibit much stronger analogical reasoning at all three stages. Additionally, the strongest tested model, GPT-o1, performs better in tasks involving simple surface-level visual attributes like color and size, correlating with quicker human adult response times. Conversely, more complex tasks such as number, rotation, and reflection, which necessitate extensive cognitive processing and understanding of extrinsic spatial properties in the physical world, present more significant challenges. Altogether, these findings highlight the limitations of training models on data that primarily consists of 2D images and text.
Unlike other primates, young children have been shown to exhibit seemingly irrational overimitation—faithfully copying unnecessary steps from a demonstration. We tested 3- to 5-year-old children (N = 39) and capuchin monkeys (N = 21) on a causal reasoning task in which a sequence of two actions was demonstrated, followed by reward production. We manipulated (1) causal plausibility and (2) the degree and nature of demonstrator intentionality, to explore the hypothesis that children—but not capuchins—integrate information about demonstrator intent and causal relations to infer which actions are necessary. We compared both species’ behavior to Bayesian computational models with the same varying social and physical expectations. Our results suggest that both species can learn from causal demonstrations, and that their copying behavior is affected by both the demonstration’s causal plausibility and the demonstrator’s communicative cues, but that children may be unique in interpreting communicative cues as having pedagogical intent.
Implications draw on the history of transformative information systems from the past.
Early childhood researchers frequently use learning materials and assessments involving pictures, across different cultures and contexts. However, there is variation in when and how children across cultures and contexts begin to understand and learn from pictures. While children growing up in high-income contexts often have more experience with picture books and other kinds of two-dimensional visual symbols, children growing up in low-income, rural contexts in low- and middle-income countries often have less experience with pictures and other kinds of visual symbols. The current research leverages variation in picture experience within a geographical region to investigate whether previous picture experience is related to toddlers' (1) performance on a picture-based word learning task, and (2) referential understanding, controlling for maternal education, number of toys, caregiver talk, and caregiver play. One hundred and twenty-eight toddlers in urban and rural western Kenya (n = 64 per area), who had varying amounts of picture experience, participated in a picture-based word learning task. Preregistered analyses with the entire sample showed no relation between picture experience and performance on a picture-based word learning task, or between picture experience and referential understanding. However, exploratory analyses found a positive association between picture experience and performance on the picture-based word learning task in the urban sample, but not the rural sample. We found no association between toddlers' referential understanding and picture experience, in either sample. We discuss how these results may inform the efficacy of learning materials and the validity of assessments used with children from diverse global backgrounds.
Curiosity is adaptive, enhances learning, and reduces uncertainty. Social curiosity is defined as the motivation to gain information about the actions, relationships, and psychology of others. Little is known about the developmental and evolutionary roots of social curiosity. Here, across three comparative studies, we investigate whether chimpanzees (n = 27) and young children (4-6 years old, n = 94) show particular interest in social interactions among third parties. Chimpanzees and children preferred to watch videos of social interactions compared with videos of a single conspecific (Experiment 1) and young children and male chimpanzees even paid a material cost to gain social information (Experiment 2). Finally, our results show that boys become more curious about negative social interactions whereas girls become more curious about positive social interactions as they develop, while chimpanzees demonstrated no preference for negative versus positive social interactions (Experiment 3). Taken together, these findings suggest that social curiosity emerges early in human ontogeny and is shared with one of our two closest living relatives, the chimpanzees.
Play is important in many cultures and species, but the basic motivations behind play remain unclear. In two preregistered experiments, we examined what 5- to 10-year-old children (n = 124) think makes play rewarding under internally and externally motivated contexts using a novel game design task. We specifically compared children's choices about how to best configure a novel tossing game when either playing for fun or playing to win. We found that for "win-relevant" variables, children chose easier settings when playing to win than when playing for fun. By contrast, for "win-irrelevant" variables, children generally preferred similar settings across conditions. Children also judged "win-relevant" variables as more important to winning than "win-irrelevant" variables and judged both as irrelevant to having fun. These results suggest that playing to win and playing for fun are distinct motivational contexts to which children can appropriately adapt their decisions during play. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
Almost all of human infants' experience and learning takes place in the context of caregiving relationships. This essay considers how infants understand the care they receive. We begin by outlining plausible features of an "intuitive theory" of care. In this intuitive theory, caregiving has both a distinctive foundational structure and distinctive features that differentiate it from other social relationships. We then review methods and findings from research on infants' understanding of people and social relationships. We propose that even before infants can use language, they may understand caregiving as an abstract intuitive theory with some features in common with how adults think about caregiving. In particular, infants understand care relationships as intimate, altruistic, and asymmetric. We review work that starts to shed light on this proposal, including the findings that infants distinguish between intimate relationships and merely positive ones and that they have asymmetric expectations of responses to distress in intimate relationships between large and small individuals. The proposal that infants can make these inferences has societal and political implications for how we structure caregiving in early life.
Many childhood assessments rely on picture stimuli, but children in diverse early environments possess varying amounts of experience with pictures. Two preregistered experiments, conducted in 2022-2023, investigated whether picture assessments are valid across diverse contexts. Low-to-middle-income children (n = 192, 2-7 years, 85 females, all Black) in their first month of formal schooling in Mombasa County, Kenya, an early environment with relatively few pictures, performed more accurately on an object vocabulary task than a picture vocabulary task (β = 0.07, p < .001; Experiment 1). Middle-to-high-income children (n = 96, 2-3 years, 52 females, predominantly White and Asian) in the San Francisco Bay Area, an early environment with relatively more pictures, performed similarly on object and picture vocabulary tasks (β = 0.02, p = .60; Experiment 2). Consequently, these results tentatively suggest that assessments involving pictures may underestimate children's capacities in some contexts. To accurately measure developing capacities in children from diverse backgrounds, it is critical that assessment tools are appropriately adapted to environmental contexts. (PsycInfo Database Record (c) 2025 APA, all rights reserved).
What drives exploration? Understanding intrinsic motivation is a long-standing challenge in both cognitive science and artificial intelligence; numerous objectives have been proposed and used to train agents, yet there remains a gap between human and agent exploration. We directly compare adults, children, and AI agents in a complex open-ended environment, Crafter, and study how common intrinsic objectives: Entropy, Information Gain, and Empowerment, relate to their behavior. We find that only Entropy and Empowerment are consistently positively correlated with human exploration progress, indicating that these objectives may better inform intrinsic reward design for agents. Furthermore, across agents and humans we observe that Entropy initially increases rapidly, then plateaus, while Empowerment increases continuously, suggesting that state diversity may provide more signal in early exploration, while advanced exploration should prioritize control. Finally, we find preliminary evidence that private speech utterances, and particularly goal verbalizations, may aid exploration in children.