In some contexts, abstract stimulus representations can effectively promote the pursuit of reward, whereas in others, more detailed representations are needed to guide choice. Here, using a novel reinforcement-learning task, we asked how children, adolescents, and adults flexibly adjust the specificity of the representations used for learning based on experienced reward statistics, as well as how the specificity of these learning representations influences subsequent memory. Across two experiments (total N = 224), we found that children, adolescents, and adults flexibly up- and down-weighted more detailed versus broader stimulus representations, depending on the reward structure of the environment. The representations used for learning shaped mnemonic specificity; when participants placed greater weight on granular representations during value-guided learning, they demonstrated enhanced subsequent memory for stimulus details, whereas when they placed greater weight on broader, categorical representations, they showed enhanced memory only for categorical information. Moreover, the relation between learning and memory strengthened with age; relative to adults, children exhibited reduced coupling between the specificity of the representations used for value-based choice and the specificity of their subsequent memories. Our work demonstrates that from early in life, reward shapes the granularity with which the world is partitioned, which, increasingly with age, influences how experiences are remembered.
By building a mental model of how the world works and using it to forecast the outcomes of different actions, a learner can make flexible choices in changing environments. However, while children and adolescents readily acquire structured knowledge about their environments, relative to adults, they tend to demonstrate weaker signatures of leveraging this knowledge to plan actions. One explanation for these developmental differences is that using a mental model to prospectively simulate potential choices and their outcomes is computationally costly, taxing cognitive control and working memory mechanisms that continue to develop into adulthood. Here, we ask whether children might effectively leverage structured knowledge to make flexible choices by relying on two alternative strategies that do not require costly mental simulation at choice time. First, through offline replanning, models can be queried before the time of choice to generate possible scenarios and update the values of potential actions. Second, an abstracted predictive model, known as a Successor Representation, can be built and harnessed to enable simplified computation of long-run reward values of candidate actions, without requiring iterative simulation of multiple time steps. To assess whether children, adolescents, and adults aged 7 - 23 years similarly harness these learning strategies, we ran three experiments. In Experiments 1 and 2, we used a reward revaluation task in which we manipulated the opportunity for offline replanning during rest, and found that children flexibly updated their behavior by leveraging structured knowledge in an adult-like manner. Surprisingly, across age, rest did not mediate flexible replanning, raising the possibility that participants may have behaved adaptively by harnessing predictive representations online. In Experiment 3, we directly tested whether children use predictive representations. Here, we observed early-emerging use of the SR, providing a mechanistic account of how children use structured knowledge to guide choice without detailed model-based simulation.
Across development, people tend to demonstrate a preference for contexts in which they have the opportunity to make choices. However, it is not clear how children, adolescents, and adults learn to calibrate this preference based on the costs and benefits of agentic choice. Here, in both a primary, in-person reinforcement-learning experiment (N = 92; age range 10 - 25 years) and a preregistered, online replication study (N = 150; age range 8 - 25 years), we found that participants overvalued agentic choice, but also calibrated their agency decisions to the reward structure of the environment, increasingly selecting agentic choice when choice had greater instrumental value. Regression analyses and computational modeling of participant choices revealed that participants’ bias toward agentic choice — reflecting the intrinsic value of choice — remained consistent across age, while sensitivity to the instrumental value of agentic choice increased from childhood to early adulthood.
To understand human learning and progress, it is crucial to understand curiosity. But how consistent is curiosity’s conception and assessment across scientific research disciplines? We present the results of a large collaborative project assessing the correspondence between curiosity measures in personality psychology and cognitive science. A total of 820 participants completed 15 personality trait measures and 9 cognitive tasks that tested multiple aspects of information demand. We show that shared variance across the cognitive tasks was captured by a dimension reflecting directed (uncertainty-driven) versus random (stochasticity-driven) exploration and individual differences along this axis were significantly and consistently predicted by personality traits. However, the personality metrics that best predicted information demand were not the central curiosity traits of openness to experience, deprivation sensitivity, and joyous exploration, but instead included more peripheral curiosity traits (need for cognition, thrill seeking, and stress tolerance) and measures not traditionally associated with curiosity (extraversion and behavioral inhibition). The results suggest that the umbrella term “curiosity” reflects a constellation of cognitive and emotional processes, only some of which are shared between personality measures and cognitive tasks. The results reflect the distinct methods that are used in these fields, indicating a need for caution in comparing results across fields and for future interdisciplinary collaborations to strengthen our emerging understanding of curiosity.
From early in life, we encounter both controllable environments, in which our actions can causally influence the reward outcomes we experience, and uncontrollable environments, in which they cannot. Environmental controllability is theoretically proposed to exert an organizing influence on our behavior. In controllable contexts, we can learn to proactively identify and select instrumental actions that bring about desired outcomes. In uncontrollable environments, we can instead rely on simpler Pavlovian learning processes that enable hard-wired reflexive reactions to anticipated, motivationally salient events. Recent empirical work has demonstrated that adults exhibit flexible controllability-dependent arbitration between instrumental and Pavlovian learning systems. In this study, we examined how the flexible arbitration between learning systems based on the degree of environmental controllability changes from childhood to adulthood. Based on prior empirical findings from a cross-species literature, we hypothesized that adolescents’ action selection may be particularly sensitive to environmental controllability. Ninety participants, aged 8-27, performed a probabilistic learning task that enables estimation of the degree of Pavlovian bias on instrumental learning. Participants completed the task in both controllable and uncontrollable conditions. We fit participants’ data with a reinforcement learning model in which controllability inferences adaptively modulate the dominance of Pavlovian versus instrumental control. Relative to children and adults, adolescents exhibited greater flexibility in calibrating the expression of Pavlovian learning biases to the degree of environmental controllability. These findings suggest that adolescence may be a period of heightened sensitivity to environmental reward statistics that organize motivated behavior across the lifespan.
Binz et al. argue that meta-learned models are essential tools for understanding adult cognition. Here, we propose that these models are particularly useful for testing hypotheses about why learning processes change across development. By leveraging their ability to discover optimal algorithms and account for capacity limitations, researchers can use these models to test competing theories of developmental change in learning.
Determining how environments shape how people learn is central to understanding individual differences in goal-directed behaviour. Studies of the effects of early-life adversity on reward learning have revealed that the environments that infants and children experience exert lasting influences on reward-guided behaviour. However, the varied findings from this research are difficult to reconcile under a unified computational account. Studies of adaptive reinforcement learning have demonstrated that learning algorithms and parameters dynamically adapt to support reward-guided behaviour in varied contexts, but this body of research has largely focused on learning that proceeds within the short timeframes of experimental tasks. In this Perspective, we argue that, to understand how the structure of experienced environments shapes reward learning across development, computational accounts of the effects of environmental statistics on reinforcement learning need to be extended to encompass learning across multiple nested timescales of experience. To this end, we consider the development of reward learning through the lens of meta-learning models, in particular meta-reinforcement learning. This computational formalization can inspire new hypotheses and methods for empirical research to understand how features of experienced environments give rise to individual differences in learning and adaptive behaviour across development. Environments shape reward learning, which can result in individual differences in behaviour. In this Perspective, Nussenbaum and Hartley consider the development of reward learning through the lens of meta-learning models, in particular meta-reinforcement learning.
A central goal of computational psychiatry is to identify systematic relationships between transdiagnostic dimensions of psychiatric symptomatology and the latent learning and decision-making computations that inform individuals' thoughts, feelings, and choices. Most psychiatric disorders emerge prior to adulthood, yet little work has extended these computational approaches to study the development of psychopathology. Here, we lay out a roadmap for future studies implementing this approach by developing empirically and theoretically informed hypotheses about how developmental changes in model-based control of action and Pavlovian learning processes may modulate vulnerability to anxiety and addiction. We highlight how insights from studies leveraging computational approaches to characterize the normative developmental trajectories of clinically relevant learning and decision-making processes may suggest promising avenues for future developmental computational psychiatry research.
Cross-species evidence suggests that the ability to exert control over a stressor is a key dimension of stress exposure that may sensitize frontostriatal-amygdala circuitry to promote more adaptive responses to subsequent stressors. The present study examined neural correlates of stressor controllability in young adults. Participants (N = 56; Mage = 23.74, range = 18-30 years) completed either the controllable or uncontrollable stress condition of the first of two novel stressor controllability tasks during functional magnetic resonance imaging (fMRI) acquisition. Participants in the uncontrollable stress condition were yoked to age- and sex-matched participants in the controllable stress condition. All participants were subsequently exposed to uncontrollable stress in the second task, which is the focus of fMRI analyses reported here. A whole-brain searchlight classification analysis revealed that patterns of activity in the right dorsal anterior insula (dAI) during subsequent exposure to uncontrollable stress could be used to classify participants' initial exposure to either controllable or uncontrollable stress with a peak of 73% accuracy. Previous experience of exerting control over a stressor may change the computations performed within the right dAI during subsequent stress exposure, shedding further light on the neural underpinnings of stressor controllability.
Across the lifespan, individuals frequently choose between exploiting known rewarding options or exploring unknown alternatives. A large body of work has suggested that children may explore more than adults. However, because novelty and reward uncertainty are often correlated, it is unclear how they differentially influence decision-making across development. Here, children, adolescents, and adults (ages 8–27 years, N = 122) completed an adapted version of a recently developed value-guided decision-making task that decouples novelty and uncertainty. In line with prior studies, we found that exploration decreased with increasing age. Critically, participants of all ages demonstrated a similar bias to select choice options with greater novelty, whereas aversion to reward uncertainty increased into adulthood. Computational modeling of participant choices revealed that whereas adolescents and adults demonstrated attenuated uncertainty aversion for more novel choice options, children’s choices were not influenced by reward uncertainty.
Fear conditioning is a widely used laboratory model to investigate learning, memory, and psychopathology across species. The quantification of learning in this paradigm is heterogeneous in humans and psychometric properties of different quantification methods can be difficult to establish. To overcome this obstacle, calibration is a standard metrological procedure in which well-defined values of a latent variable are generated in an established experimental paradigm. These intended values then serve as validity criterion to rank methods. Here, we develop a calibration protocol for human fear conditioning. Based on a literature review, series of workshops, and survey of N = 96 experts, we propose a calibration experiment and settings for 25 design variables to calibrate the measurement of fear conditioning. Design variables were chosen to be as theory-free as possible and allow wide applicability in different experimental contexts. Besides establishing a specific calibration procedure, the general calibration process we outline may serve as a blueprint for calibration efforts in other subfields of behavioral neuroscience that need measurement refinement.
Optimal integration of positive and negative outcomes during learning varies depending on an environment’s reward statistics. The present study investigated the extent to which children, adolescents, and adults (N = 142 8 - 25 year-olds, 55% female, 42% White, 31% Asian, 17% mixed race, and 8% Black) adapt their weighting of better-than-expected and worse-than-expected outcomes when learning from reinforcement. Participants made a series of choices across two contexts: one in which weighting positive outcomes more heavily than negative outcomes led to better performance, and one in which the reverse was true. Reinforcement learning modeling revealed that across age, participants shifted their valence biases in accordance with the structure of the environment. Exploratory analyses revealed increases in context-dependent flexibility with age.
Previously rewarding experiences can influence choices in new situations. Past work has demonstrated that existing reward associations can either help or hinder future behaviors and that there is substantial individual variability in the transfer of value across contexts. Developmental changes in reward sensitivity may also modulate the impact of prior reward associations on later goal-directed behavior. The current study aimed to characterize how reward associations formed in the past affected learning in the present from childhood to adulthood. Participants completed a reinforcement learning paradigm using high- and low-reward stimuli from a task completed 24 h earlier, as well as novel stimuli, as choice options. We found that prior high-reward associations impeded learning across all ages. We then assessed how individual differences in the prioritization of high- versus low-reward associations in memory impacted new learning. Greater high-reward memory prioritization was associated with worse learning performance for previously high-reward relative to low-reward stimuli across age. Adolescents also showed impeded early learning regardless of individual differences in high-reward memory prioritization. Detrimental effects of previous reward on choice behavior did not persist beyond learning. These findings indicate that prior reward associations proactively interfere with future learning from childhood to adulthood and that individual differences in reward-related memory prioritization influence new learning across age.
The model‐free algorithms of “reinforcement learning” (RL) have gained clout across disciplines, but so too have model‐based alternatives. The present study emphasizes other dimensions of this model space in consideration of associative or discriminative generalization across states and actions. This “generalized reinforcement learning” (GRL) model, a frugal extension of RL, parsimoniously retains the single reward‐prediction error (RPE), but the scope of learning goes beyond the experienced state and action. Instead, the generalized RPE is efficiently relayed for bidirectional counterfactual updating of value estimates for other representations. Aided by structural information but as an implicit rather than explicit cognitive map, GRL provided the most precise account of human behavior and individual differences in a reversal‐learning task with hierarchical structure that encouraged inverse generalization across both states and actions. Reflecting inference that could be true, false (i.e., overgeneralization), or absent (i.e., undergeneralization), state generalization distinguished those who learned well more so than action generalization. With high‐resolution high‐field fMRI targeting the dopaminergic midbrain, the GRL model's RPE signals (alongside value and decision signals) were localized within not only the striatum but also the substantia nigra and the ventral tegmental area, including specific effects of generalization that also extend to the hippocampus. Factoring in generalization as a multidimensional process in value‐based learning, these findings shed light on complexities that, while challenging classic RL, can still be resolved within the bounds of its core computations.
Cross-species research suggests that exploratory behaviors increase during adolescence and relate to the social, affective, and risky behaviors characteristic of this developmental stage. However, how these typical adolescent behaviors manifest and relate in real-world settings remains unclear. Using geolocation tracking to quantify exploration—variability in daily movement patterns—over a 3-month period in 58 adolescents and adults (ages 13–27) in New York City, we investigated whether daily exploration varied with age and whether exploration related to social connectivity, risk taking, and momentary positive affect. In our cross-sectional sample, we found an association between daily exploration and age, with individuals near the transition to legal adulthood exhibiting the highest exploration levels. Days of higher exploration were associated with greater positive affect irrespective of age. Higher mean exploration was associated with greater social connectivity in all participants but was linked to higher risk taking selectively among adolescents. Our results highlight the interplay of exploration and socioemotional behaviors across development and suggest that societal norms may modulate their expression in naturalistic contexts.
As individuals learn through trial and error, some are more influenced by good outcomes, while others weight bad outcomes more heavily. Such valence biases may also influence memory for past experiences. Here, we examined whether valence asymmetries in reinforcement learning change across adolescence, and whether individual learning asymmetries bias the content of subsequent memory. Participants ages 8–27 learned the values of ‘point machines,’ after which their memory for trial-unique images presented with choice outcomes was assessed. Relative to children and adults, adolescents overweighted worse-than-expected outcomes during learning. Individuals’ valence biases modulated incidental memory, such that those who prioritized worse- (or better-) than-expected outcomes during learning were also more likely to remember images paired with these outcomes, an effect reproduced in an independent dataset. Collectively, these results highlight age-related changes in the computation of subjective value and demonstrate that a valence-asymmetric valuation process influences how information is prioritized in episodic memory.
How does human cognition adapt to idiosyncratic features of our real-world experiences across our lifetimes? The dynamic interaction between individuals and their natural environments is rarely the focus of study within cognitive science, but I argue that a more ecological approach will be critical for advancing developmental science and revealing the adaptive nature of cognition.
This study aimed to distinguish between the developmental trajectories of different cognitive component processes underlying planning decisions. Participants (8-25 year olds) completed a planning task called Four-in-a-row. By fitting a computational model, we distinguished between three cognitive component processes of planning: planning depth, heuristic quality, and attentional oversights. All three contributed to better playing strength, but they differed in their developmental trajectories. Specifically, from early to mid-adolescence, heuristic quality rapidly improved and contributed to better playing strength. From mid to late-adolescence, planning depth increased and supported better playing strength. Fewer attentional oversights were associated with better playing strength and this relation did not show age differences. Together, these results suggest an order in which the use of cognitive component processes of planning develop, starting by first refining the heuristic strategies, then gradually increasing the number possible future actions, states, and outcomes considered towards young adulthood. The findings move the field of cognitive development towards a more complete account of the development of planning and its component processes.
Reward motivation enhances memory through interactions between mesolimbic, hippocampal, and cortical systems, both during and after encoding. Developmental changes in these distributed neural circuits may lead to age-related differences in reward-motivated memory and the underlying neural mechanisms. Converging evidence from cross-species studies suggests that subcortical dopamine signaling is increased during adolescence, which may lead to stronger memory representations of rewarding, relative to mundane, events and changes in the contributions of underlying subcortical and cortical brain mechanisms across age. Here, we used fMRI to examine how reward motivation influences the "online" encoding and "offline" postencoding brain mechanisms that support long-term associative memory from childhood to adulthood in human participants of both sexes. We found that reward motivation led to both age-invariant enhancements and nonlinear age-related differences in associative memory after 24 h. Furthermore, reward-related memory benefits were linked to age-varying neural mechanisms. During encoding, interactions between the prefrontal cortex (PFC) and ventral tegmental area (VTA) were associated with better high-reward memory to a greater degree with increasing age. Preencoding to postencoding changes in functional connectivity between the anterior hippocampus and VTA were also associated with better high-reward memory, but more so at younger ages. Our findings suggest that there may be developmental differences in the contributions of offline subcortical and online cortical brain mechanisms supporting reward-motivated memory.SIGNIFICANCE STATEMENT A substantial body of research has examined the neural mechanisms through which reward influences memory formation in adults. However, despite extensive evidence that both reward processing and associative memory undergo dynamic change across development, few studies have examined age-related changes in these processes. We found both age-invariant and nonlinear age-related differences in reward-motivated memory. Moreover, our findings point to developmental differences in the processes through which reward modulates the prioritization of information in long-term memory, with greater early reliance on offline subcortical consolidation mechanisms and increased contribution of systems-level online encoding circuitry with increasing age. These results highlight dynamic developmental changes in the cognitive and neural mechanisms through which motivationally salient information is prioritized in memory from childhood to adulthood.
Accurate assessment of environmental controllability enables individuals to adaptively adjust their behavior — exploiting rewards when desirable outcomes are contingent upon their actions and minimizing costly deliberation when their actions are inconsequential. However, it remains unclear how estimation of environmental controllability changes from childhood to adulthood. Ninety participants (ages 8-25) completed a task that covertly alternated between controllable and uncontrollable conditions, requiring them to explore different actions to discover the current degree of environmental controllability. We found that while children were able to distinguish controllable and uncontrollable conditions, accuracy of controllability assessments improved with age. Computational modeling revealed that whereas younger participants’ controllability assessments relied on evidence gleaned through random exploration, older participants more effectively recruited their task structure knowledge to make highly informative interventions. Age-related improvements in working memory mediated this qualitative shift toward increased use of an inferential strategy. Collectively, these findings reveal an age-related shift in the cognitive processes engaged to assess environmental controllability. Improved detection of environmental controllability may foster increasingly adaptive behavior over development by revealing when actions can be leveraged for one’s benefit.