Adaptive decision-making in dynamic environments requires flexible adjustment of learning speed to balance stability and flexibility. When outcomes are highly stochastic, learners must avoid over-interpreting noise and update more slowly, whereas in volatile environments where contingencies change frequently, learning should accelerate to rapidly incorporate new evidence. Theories propose that internal estimates of uncertainty tune learning rates through neuromodulation-dependent mechanisms. Here, we investigated how noradrenergic inputs from the locus coeruleus (LC) to the orbitofrontal cortex (OFC) support adaptive learning under uncertainty. We show that, in a probabilistic reversal learning task performed across different levels of stochasticity, rats exhibited behavior best explained by an adaptive reinforcement-learning model in which learning rates dynamically adjusted according to model-estimated stochasticity and volatility, outperforming standard fixed-rate models. Noradrenaline release in the OFC closely tracked trial-by-trial, model-derived volatility estimates around contingency changes. Disrupting LC→OFC noradrenergic inputs reproduced the model-predicted impairment in adaptive learning-rate adjustment associated with model-estimated volatility. Together, these findings identify OFC noradrenergic signaling as a key circuit mechanism supporting learning-rate adjustment in response to internal estimates of volatility during adaptive decision-making under uncertainty.
Adaptive behavior often relies on learning associations between stimuli that were never directly paired with reinforcement. Sensory preconditioning provides a powerful paradigm to investigate such indirect learning: when neutral stimuli A and X are paired during preconditioning, and X is subsequently paired with an unconditioned stimulus (US) during conditioning, stimulus A elicits conditioned response at test despite never being directly paired with the US. While the hippocampus is known to be critical for this process, its involvement in unimodal sensory preconditioning remains unclear. Here, we investigated the role of the hippocampus (dorsal and ventral) during different phases of gustatory preconditioning using chemogenetic approach in rats. We demonstrated that disruption of the hippocampus during either preconditioning or test selectively impaired the mediated aversion to A while leaving the direct conditioned aversion to X intact. These findings provide the first evidence that the hippocampus is necessary for unimodal gustatory preconditioning, extending previous demonstrations of hippocampal involvement in polymodal protocols. Our work therefore clarifies the critical role of the hippocampus in forming and retrieving purely gustatory associations to guide behavior, highlighting its function as a central hub for memory integration.
In a dynamic environment, organisms must continuously update learned action-outcome associations. Central to this flexibility is the prefrontal cortex, whose computations are finely tuned by neuromodulatory inputs. Yet, the temporal dynamics and circuit specificity of this regulation remain unclear. Here, we investigate the contribution of orbitofrontal noradrenaline (OFC-NA) to flexible updating in rats performing an instrumental reversal learning task. Using fiber photometry, we observe transient increases in OFC-NA release following reward deliveries on reversal day, and we find that the magnitude of these responses predicts the speed of behavioral adaptation. Chemogenetic and optogenetic manipulations of NA projections from the locus coeruleus (LC) to the OFC show that perturbing these signals delays reversal learning in a graded, mode-dependent manner, with chemogenetic inhibition having the strongest impact. Together, our findings establish OFC-NA as a temporally precise neuromodulatory mechanism, gating flexible adaptation to changing environmental contingencies.
Adaptive decision-making in dynamic environments requires flexible adjustment of learning speed to balance stability and flexibility. When outcomes are highly stochastic, learners must avoid over interpreting noise and update more slowly, whereas in volatile environments where contingencies change frequently, learning should accelerate to rapidly incorporate new evidence. Theories propose that internal estimates of uncertainty tune learning rates through neuromodulatory-dependent mechanisms. Here, we investigated how noradrenergic inputs from the locus coeruleus (LC) to the orbitofrontal cortex (OFC) support adaptive learning under uncertainty. We show that rats performing a probabilistic reversal learning task exhibited behavior that was best explained by an adaptive reinforcement-learning model in which learning rates dynamically adjust according to estimated environmental volatility and stochasticity, outperforming standard fixed-rate models. Noradrenaline release in the OFC closely tracked trial-by-trial, model-derived volatility estimates around contingency changes. Disrupting LC→OFC noradrenergic inputs reproduced the model-predicted deficit in volatility-dependent adjustments of learning rate. Together, these findings identify OFC noradrenergic signaling as a key circuit mechanism for volatility-dependent modulation of learning rates during adaptive decision-making. ### Competing Interest Statement The authors have declared no competing interest.
Obesity is associated with neurocognitive dysfunction, including memory deficits. This is particularly worrisome when obesity occurs during adolescence, a maturational period for brain structures critical for cognition. In rodent models, we recently reported that memory impairments induced by obesogenic high-fat diet (HFD) intake during the periadolescent period can be reversed by chemogenetic manipulation of the ventral hippocampus (vHPC). Here, we used an intersectional viral approach in HFD-fed male mice to chemogenetically inactivate specific vHPC efferent pathways to nucleus accumbens (NAc) or medial prefrontal cortex (mPFC) during memory tasks. We first demonstrated that HFD enhanced activation of both pathways after training and that our chemogenetic approach was effective in normalizing this activation. Inactivation of the vHPC–NAc pathway rescued HFD-induced deficits in recognition but not location memory. Conversely, inactivation of the vHPC–mPFC pathway restored location but not recognition memory impairments produced by HFD. Either pathway manipulation did not affect exploration or anxiety-like behaviour. These findings suggest that HFD intake throughout adolescence impairs different types of memory through overactivation of specific hippocampal efferent pathways and that targeting these overactive pathways has therapeutic potential.
In uncertain environments in which resources fluctuate continuously, animals must permanently decide whether to stabilise learning and exploit what they currently believe to be their best option, or instead explore potential alternatives and learn fast from new observations. While such a trade-off has been extensively studied in pretrained animals facing non-stationary decision-making tasks, it is yet unknown how they progressively tune it while learning the task structure during pretraining. Here, we compared the ability of different computational models to account for long-term changes in the behaviour of 24 rats while they learned to choose a rewarded lever in a three-armed bandit task across 24 days of pretraining. We found that the day-by-day evolution of rat performance and win-shift tendency revealed a progressive stabilisation of the way they regulated reinforcement learning parameters. We successfully captured these behavioural adaptations using a meta-learning model in which either the learning rate or the inverse temperature was controlled by the average reward rate.
A dynamic environment, such as the one we inhabit, requires organisms to continuously update their knowledge of the setting. While the prefrontal cortex is recognized for its pivotal role in regulating such adaptive behavior, the specific contributions of each prefrontal area remain elusive. In the current work, we investigated the direct involvement of two major prefrontal subregions, the medial prefrontal cortex (mPFC, A32D+A32V) and the orbitofrontal cortex (OFC, VO+LO), in updating Pavlovian stimulus-outcome (S-O) associations following contingency degradation in male rats. Specifically, animals had to learn that a particular cue, previously fully predicting the delivery of a specific reward, was no longer a reliable predictor. First, we found that chemogenetic inhibition of mPFC, but not of OFC, neurons altered the rats’ ability to adaptively respond to degraded and non-degraded cues. Next, given the growing evidence pointing at noradrenaline (NA) as a main neuromodulator of adaptive behavior, we decided to investigate the possible involvement of NA projections to the two subregions in this higher-order cognitive process. Employing a pair of novel retrograde vectors, we traced NA projections from the locus coeruleus (LC) to both structures and observed an equivalent yet relatively segregated amount of inputs. Then, we showed that chemogenetic inhibition of NA projections to the mPFC, but not to the OFC, also impaired the rats’ ability to adaptively respond to the degradation procedure. Altogether, our findings provide important evidence of functional parcellation within the prefrontal cortex and point at mPFC-NA as key for updating Pavlovian S-O associations.Significant statementThe ability to update stimulus-outcome (S-O) associations is a key adaptive behavior, essential for surviving and thriving in an ever-changing environment. The prefrontal cortex is well-known for playing a key role in this process. The discrete contribution of each prefrontal subregion and of different neurotransmitters, however, remains unclear. In the current study, we show that inhibiting medial prefrontal (mPFC), but not orbitofrontal cortex (OFC), neurons impairs rats’ ability to update S-O associations following contingency degradation. Moreover, we demonstrate that discrete noradrenergic projections to the two subregions exist and that inhibiting the ones projecting to the mPFC, but not to the OFC, once again impairs the animals’ behavior, thereby implying a substantial contribution of noradrenaline in orchestrating this higher-order cognitive process.
The article “In vitro neurons learn and exhibit sentience when embodied in a simulated game-world” by Kagan et al. 1 Kagan B.J. Kitchen A.C. Tran N.T. Habibollahi F. Khajehnejad M. Parker B.J. Bhat A. Rollo B. Razi A. Friston K.J. In vitro neurons learn and exhibit sentience when embodied in a simulated game-world. Neuron. 2022; 110: 3952-3969.e8https://doi.org/10.1016/j.neuron.2022.09.001 Abstract Full Text Full Text PDF PubMed Scopus (17) Google Scholar triggered a wave of positive mainstream and scientific media coverage as well as a widespread negative reaction from the scientific community. Here, we discuss why this negative reaction is legitimate and must be taken seriously. We raise concerns about the key claim of the article: that it demonstrates that “a single layer of in vitro cortical neurons can self-organize activity to display intelligent and sentient behavior when embodied in a simulated game-world.” Scientific communication and the semantics of sentienceKagan et al.NeuronMarch 01, 2023In BriefThe use of language to describe specific phenomena has always been, and will likely remain, a contentious aspect of scientific discourse. Effective scientific communication must be considered in the context of the field in which the given signifiers are used. While the response by Balci et al. may be described as polemic and contains reasoning subject to equivocation and other fallacies, we do appreciate the core concerns. Here, we address these concerns and suggest some constructive pathways that might help improve scientific communication in the future. Full-Text PDF In vitro neurons learn and exhibit sentience when embodied in a simulated game-worldKagan et al.NeuronOctober 12, 2022In BriefThe DishBrain system is the first real-time synthetic biological intelligence platform that demonstrates that biological neurons can adjust firing activity in a way that suggests the ability to learn to perform goal-oriented tasks when provided with simple electrophysiological sensory input and feedback while embodied in a game-world. Full-Text PDF Open AccessConceptual conundrums for neuroscienceRommelfanger et al.NeuronMarch 01, 2023In BriefIn their recently published work, Kagan and colleagues argue that cultures of human and mouse neurons can exhibit goal-directed activity adaptations and are thus “sentient.” In their commentary to this piece, Balci and colleagues1 offer a critical review of the work. In addition to discussing the value of the research and the soundness of the methodology, they question the appropriateness of using language such as “sentience” in the context of the study. Research in isolated neural networks provides crucial knowledge, Balci and colleagues note, but the terms that Kagan et al. Full-Text PDF
In uncertain environments in which resources fluctuate continuously, animals must permanently decide whether to exploit what they currently believe to be their best option, or instead explore potential alternatives in case better opportunities are in fact available. While such a trade-off has been extensively studied in pretrained animals facing non-stationary decision-making tasks, it is yet unknown how they progressively tune it while progressively learning the task structure during pretraining. Here, we compared the ability of different computational models to account for long-term changes in the behaviour of 24 rats while they learned to choose a rewarded lever in a three-armed bandit task across 24 days of pretraining. We found that the day-by-day evolution of rat performance and win-shift tendency revealed a progressive stabilization of the way they regulated the exploration-exploitation trade-off. We successfully captured these behavioural adaptations using a meta-learning model in which the exploration-exploitation trade-off is controlled by the animal’s average reward rate.
In a constantly changing environment, organisms must track the current relationship between actions and their specific consequences and use this information to guide decision-making. Such goal-directed behaviour relies on circuits involving cortical and subcortical structures. Notably, a functional heterogeneity exists within the medial prefrontal, insular, and orbitofrontal cortices (OFC) in rodents. The role of the latter in goal-directed behaviour has been debated, but recent data indicate that the ventral and lateral subregions of the OFC are needed to integrate changes in the relationships between actions and their outcomes. Neuromodulatory agents are also crucial components of prefrontal functions and behavioural flexibility might depend upon the noradrenergic modulation of the prefrontal cortex. Therefore, we assessed whether noradrenergic innervation of the OFC plays a role in updating action-outcome relationships in male rats. We used an identity-based reversal task and found that depletion or chemogenetic silencing of noradrenergic inputs within the OFC rendered rats unable to associate new outcomes with previously acquired actions. Silencing of noradrenergic inputs in the prelimbic cortex or depletion of dopaminergic inputs in the OFC did not reproduce this deficit. Together, our results suggest that noradrenergic projections to the OFC are required to update goal-directed actions.
The ability to engage into flexible behaviors is crucial in dynamic environments. We recently showed that in addition to the well described role of the orbitofrontal cortex (OFC), its thalamic input from the submedius thalamic nucleus (Sub) also contributes to adaptive responding during Pavlovian degradation. In the present study, we examined the role of the mediodorsal thalamus (MD) which is the other main thalamic input to the OFC. To this end, we assessed the effect of both pre- and post-training MD lesions in rats performing a Pavlovian contingency degradation task. Pre-training lesions mildly impeded the establishment of stimulus-outcome associations during the initial training of Pavlovian conditioning without interfering with Pavlovian degradation training when the sensory feedback provided by the outcome rewards were available to animals. However, we found that both pre- and post-training MD lesions produced a selective impairment during a test conducted under extinction conditions, during which only current mental representation could guide behavior. Altogether, these data suggest a role for the MD in the successful encoding and representation of Pavlovian associations.
Choosing between different course of behavioral response is an essential process to survive in a complex environment. Numerous studies have demonstrated that basic processes of action control may be investigated using instrumental conditioning, as instrumental response may be dissociated in goal-directed action or habitual response depending both on different, but interacting, neuronal circuits. The dopamine system is a central element in the coordination between actions and habits. In this chapter, we describe in details the different behavioral procedures used to investigate actions and habits in rodent models, including instrumental learning, outcome devaluation, and contingency degradation. We also discuss how these procedures can be combined with other techniques to specifically investigate the role of the dopamine system in these different processes.
SUMMARYIn a constantly changing environment, organisms must track the current relationship between actions and their specific consequences and use this information to guide decision-making. Such goal-directed behavior relies on circuits involving cortical and subcortical structures. Notably, a functional heterogeneity exists within the medial prefrontal, insular, and orbitofrontal cortices (OFC) in rodents. The role of the latter in goal-directed behavior has been debated, but recent data indicate that the ventral and lateral subregions of the OFC are needed to integrate changes in the relationships between actions and their outcomes. Neuromodulatory agents are also crucial components of prefrontal functions and behavioral flexibility might depend upon the noradrenergic modulation of prefrontal cortex. Therefore, we assessed whether noradrenergic innervation of the OFC plays a role in updating action-outcome relationships. We used an identity-based reversal task and found that depletion or chemogenetic silencing of noradrenergic inputs within the OFC rendered rats unable to associate new outcomes with previously acquired actions. Silencing of noradrenergic inputs in the medial prefrontal cortex or depletion of dopaminergic inputs in the OFC did not reproduce this deficit. Together, our results indicate that noradrenergic projections to the OFC are required to update goal-directed actions.GRAPHICAL ABSTRACTHIGHLIGHTSRats learn initial action-outcome associations in an instrumental taskNoradrenergic depletion in the OFC prevents the encoding and expression of these associations following reversal learningDopaminergic depletion in the OFC does not result in behavioral deficitsLC:OFC noradrenergic projections are required to update action-outcome associationsIN BRIEFCerpa et al. investigate whether noradrenergic projections from the locus coeruleus (LC) to the orbitofrontal cortex are involved in updating previously established goal-directed actions following environmental change. They find that these LC projections are required to both encode and express reversed action-outcome associations in rats.
The prefrontal cortex is considered to be at the core of goal-directed behaviors. Notably, the medial prefrontal cortex (mPFC) is known to play an important role in learning action-outcome (A-O) associations, as well as in detecting changes in this contingency. Previous studies have also highlighted a specific engagement of the dopaminergic pathway innervating the mPFC in adapting to changes in action causality. While previous research on goal-directed actions has primarily focused on the mPFC region, recent findings have revealed a distinct and specific role of the ventral and lateral orbitofrontal cortex (vlOFC). Indeed, vlOFC is not necessary to learn about A-O associations but appears specifically involved when outcome identity is unexpectedly changed. Unlike the mPFC, the vlOFC does not receive a strong dopaminergic innervation. However, it receives a dense noradrenergic innervation which might indicate a crucial role for this neuromodulator. In addition, several lines of evidence highlight a role for noradrenaline in adapting to changes in the environment. We, therefore, propose that the vlOFC's function in action control might be under the strong influence of the noradrenergic system. In the present article, we review anatomical and functional evidence consistent with this proposal and suggest a direction for future studies that aim to shed light on the orbitofrontal mechanisms for flexible action control. Specifically, we suggest that dopaminergic modulation in the mPFC and noradrenergic modulation in the vlOFC may underlie distinct processes related to updating one's actions. (PsycInfo Database Record (c) 2021 APA, all rights reserved).
In addition to numerous metabolic comorbidities, obesity is associated with several adverse neurobiological outcomes, especially learning and memory alterations. Obesity prevalence is rising dramatically in youth and is persisting in adulthood. This is especially worrying since adolescence is a crucial period for the maturation of certain brain regions playing a central role in memory processes such as the hippocampus and the amygdala. We previously showed that periadolescent, but not adult, exposure to obesogenic high-fat diet (HFD) had opposite effects on hippocampus- and amygdala-dependent memory, impairing the former and enhancing the latter. However, the causal role of these two brain regions in periadolescent HFD-induced memory alterations remains unclear. Here, we first showed that periadolescent HFD induced long-term, but not short-term, object recognition memory deficits, specifically when rats were exposed to a novel context. Using chemogenetic approaches to inhibit targeted brain regions, we then demonstrated that recognition memory deficits are dependent on the activity of the ventral hippocampus, but not the basolateral amygdala. On the contrary, the HFD- induced enhancement of conditioned odor aversion specifically requires amygdala activity. Taken together, these findings suggest that HFD consumption throughout adolescence impairs long-term object recognition memory through alterations of ventral hippocampal activity during memory acquisition. Moreover, these results further highlight the bidirectional effects of adolescent HFD on hippocampal and amygdala functions.
The prefrontal cortex is considered to be at the core of goal-directed behaviours. Notably, the medial prefrontal cortex (mPFC) is known to play an important role in learning action-outcome associations, as well as in detecting changes in this contingency. Previous studies have also highlighted a specific engagement of the dopaminergic pathway innervating the mPFC in adapting to changes in action causality. While previous research on goal-directed actions has primarily focused on the mPFC region, recent findings have revealed a distinct and specific role of the ventral and lateral orbitofrontal cortex (vlOFC). Indeed, vlOFC is not necessary to learn about action-outcome associations but appears specifically involved when outcome identity is unexpectedly changed. Unlike the mPFC, the vlOFC does not receive a strong dopaminergic innervation. However, it receives a dense noradrenergic innervation which might indicate a crucial role for this neuromodulator. In addition, several lines of evidence highlight a role for noradrenaline in adapting to changes in the environment. We therefore propose that the vlOFC’s function in action control might be under the strong influence of the noradrenergic system. In the present paper, we review anatomical and functional evidence consistent with this proposal and suggest a direction for future studies that aims to shed light on the orbitofrontal mechanisms for flexible action control. Specifically, we suggest that dopaminergic modulation in the mPFC and noradrenergic modulation in the vlOFC may underlie distinct processes related to updating one’s actions.
Techniques that allow the manipulation of specific neural circuits have greatly increased in the past few years. DREADDs (Designer receptors exclusively activated by designer drugs) provide an elegant way to manipulate individual brain structures and/or neural circuits, including neuromodulatory pathways. Considerable efforts have been made to increase cell-type specificity of DREADD expression while decreasing possible limitations due to multiple viral vectors injections. In line with this, a retrograde canine adenovirus type 2 (CAV-2) vector carrying a Cre-dependent DREADD cassette has been recently developed. In combination with Cre-driver transgenic animals, the vector allows one to target neuromodulatory pathways with cell-type specificity. In the present study, we specifically targeted catecholaminergic pathways by injecting the vector in knock-in rat line containing Cre recombinase cassette under the control of the tyrosine hydroxylase promoter. We assessed the efficacy of infection of the nigrostriatal pathway and the catecholaminergic pathways ascending to the orbitofrontal cortex (OFC) and found cell-type-specific DREADD expression.
In a volatile environment where rewards are uncertain, successful performance requires a delicate balance between exploitation of the best option and exploration of alternative choices. It has theoretically been proposed that dopamine contributes to the control of this exploration-exploitation trade-off, specifically that the higher the level of tonic dopamine, the more exploitation is favored. We demonstrate here that there is a formal relationship between the rescaling of dopamine positive reward prediction errors and the exploration-exploitation trade-off in simple non-stationary multi-armed bandit tasks. We further show in rats performing such a task that systemically antagonizing dopamine receptors greatly increases the number of random choices without affecting learning capacities. Simulations and comparison of a set of different computational models (an extended Q-learning model, a directed exploration model, and a meta-learning model) fitted on each individual confirm that, independently of the model, decreasing dopaminergic activity does not affect learning rate but is equivalent to an increase in random exploration rate. This study shows that dopamine could adapt the exploration-exploitation trade-off in decision-making when facing changing environmental contingencies.
Article Figures and data Abstract Introduction Results Discussion Materials and methods Data availability References Decision letter Author response Article and author information Metrics Abstract Cues in the environment can elicit complex emotional states, and thereby maladaptive behavior, as a function of their ascribed value. Here we capture individual variation in the propensity to attribute motivational value to reward-cues using the sign-tracker/goal-tracker animal model. Goal-trackers attribute predictive value to reward-cues, and sign-trackers attribute both predictive and incentive value. Using chemogenetics and microdialysis, we show that, in sign-trackers, stimulation of the neuronal pathway from the prelimbic cortex (PrL) to the paraventricular nucleus of the thalamus (PVT) decreases the incentive value of a reward-cue. In contrast, in goal-trackers, inhibition of the PrL-PVT pathway increases both the incentive value and dopamine levels in the nucleus accumbens shell. The PrL-PVT pathway, therefore, exerts top-down control over the dopamine-dependent process of incentive salience attribution. These results highlight PrL-PVT pathway as a potential target for treating psychopathologies associated with the attribution of excessive incentive value to reward-cues, including addiction. https://doi.org/10.7554/eLife.49041.001 Introduction Learning to associate environmental stimuli with the availability of valuable resources, such as food, is critical for survival. Such stimulus-reward associations rely on Pavlovian conditioning, a learning process during which a once neutral stimulus becomes a conditioned stimulus (CS), as it reliably predicts the delivery of an unconditioned stimulus (US) (e.g. food). The CS, then, attains predictive value and comes to elicit a conditioned response. Yet, we know from both preclinical and clinical studies that CSs can also acquire incentive value and elicit complex emotional and motivational states (Robinson and Berridge, 2008; Robinson and Flagel, 2009; Tibboel et al., 2015; Pool et al., 2016). When a CS is attributed with incentive salience and transformed into an incentive stimulus, it becomes attractive and desirable in its own right. That is, the CS becomes a 'motivational magnet', now capable of capturing attention and eliciting approach behavior (Berridge, 2009a). However, individuals vary in their propensity to attribute incentive value to reward cues, and only for some individuals do such cues acquire inordinate control over behavior and the ability to elicit maladaptive tendencies (Flagel et al., 2007; Flagel et al., 2009) that are characteristic of psychopathology. Indeed, several psychiatric disorders have been associated with the excessive attribution of motivational significance to environmental cues, including substance use disorder (Robinson and Berridge, 1993; Berridge and Robinson, 2016; Kwako et al., 2017; MacNiven et al., 2018), eating disorders (Berridge et al., 2009b; Robinson, 2014), gambling disorder (Limbrick-Oldfield et al., 2017), post-traumatic stress disorder (PTSD) (Coffey et al., 2010), and bipolar disorder (Mason et al., 2012; Whitton et al., 2015). In recent years, an animal model has been established that enables us to parse the neurobiological mechanisms that may bias the way an individual responds to reward-cues. While the preferential use of predictive vs. incentive learning strategies may be adaptive under the right conditions; an extreme bias for the selective use of a single strategy may contribute to increased risk for psychopathology. When rats are exposed to a Pavlovian conditioned approach (PavCA) paradigm in which the presentation of a lever-CS is immediately followed by the response-independent delivery of food-US, distinct conditioned responses emerge. Some rats, goal-trackers (GTs), approach the location of impending food delivery upon the lever-CS presentation, while others, sign-trackers (STs), approach and interact with the lever-CS itself. For both GTs and STs the lever-CS acquires predictive value and elicits a conditioned response, but for STs, the CS also acquires incentive value. This animal model, therefore, provides a unique platform to investigate the neurobiological determinants of individual differences in the propensity to attribute incentive salience to reward-cues. Previous studies suggest that sign-tracking behavior results from enhanced activity in subcortical brain circuits known to mediate motivated behaviors, including the striatal dopamine system, the amygdala, and the hypothalamus (Flagel et al., 2011a; Flagel et al., 2011b; Saunders and Robinson, 2012; Yager et al., 2015; Singer et al., 2016; Haight et al., 2017). In addition, relative to GTs, STs appear to have deficits in top-down cognitive control originating in the prefrontal cortex (Paolone et al., 2013). Thus, we hypothesize that sign-tracking behavior arises from an imbalance between top-down cognitive control and bottom-up motivational processes. One brain region that is ideally situated to act as a fulcrum between cortical, limbic and homeostatic circuits is the paraventricular nucleus of the thalamus (PVT). The PVT receives cortical afferents from the medial PFC, including the infralimbic (IL) and prelimbic (PrL) cortices, and subcortical afferents from the hypothalamus, amygdala, and several brainstem regions involved in visceral functions and homeostatic regulation (Hsu and Price, 2007; Li and Kirouac, 2012; Kirouac, 2015). The PVT sends projections to various brain regions that have been associated with reward-learning and motivated behaviors, including the PrL and IL cortices, nucleus accumbens (NAc) core (NAcC) and shell (NAcS), lateral bed nucleus of the stria terminalis and central amygdala (Hsu and Price, 2007; Li and Kirouac, 2008; Hsu and Price, 2009). Recent findings surrounding the functional role of these PVT circuits (Do-Monte et al., 2015; Haight et al., 2017; Millan et al., 2017; Giannotti et al., 2018) have garnered recognition of the PVT as the 'thalamic gateway' (Millan et al., 2017) for appetitive motivation; acting to integrate cognitive, emotional, motivational and viscerosensitive information, and, in turn, guide behavioral responses (Kirouac, 2015; Millan et al., 2017). Consistent with this view, the PVT has been implicated in the propensity to attribute incentive motivational value to reward-cues (Haight and Flagel, 2014; Haight et al., 2015; Kuhn et al., 2018). The functional connectivity of the PVT in response to cue-induced neuronal activity differentiates STs from GTs (Flagel et al., 2011a; Haight and Flagel, 2014; Haight et al., 2017). In STs, cue-induced activity in the PVT is correlated with activity in subcortical areas, including the NAc; whereas in GTs, cue-induced activity in the PVT is correlated with activity in the PrL (Flagel et al., 2011a; Haight and Flagel, 2014). Further investigation of PVT-associated circuity in STs and GTs revealed that these phenotypes exhibit the same degree of cue-induced neural activity in PrL neurons that project to the PVT; but STs show greater cue-induced activity in subcortical afferents to the PVT, including the hypothalamus, and efferents from the PVT, including those to the NAc (Haight et al., 2017). These data suggest that the predictive value of the reward-cue is encoded in the PrL-PVT circuit. In STs, however, cognitive information about the predictive value of the reward-cue presumably competes with overriding subcortical motivational circuits, thus rendering them more prone to attribute incentive value. From this, we hypothesized that stimulating the PrL-PVT circuit (i.e. enhancing top-down control) in STs would reduce the tendency to attribute incentive value to a food-cue, by counteracting their inherent bias towards bottom-up motivational mechanisms. In contrast, we hypothesized that inhibiting the PrL-PVT circuit (i.e. attenuating top-down control) in GTs would increase the tendency to attribute incentive value to a food-cue, by weakening the top-down cognitive component of the system and permitting bottom-up motivational mechanisms to act. Because the PVT sends dense projections to the NAc (Berendse and Groenewegen, 1990; Li and Kirouac, 2008; Kirouac, 2015), and can affect local dopamine (DA) release (Jones et al., 1989; Pinto et al., 2003; Parsons et al., 2007; Choi et al., 2012; Perez and Lodge, 2018), which is critical for incentive learning (Berridge and Robinson, 1998; Flagel et al., 2011b; Saunders and Robinson, 2012), we also hypothesized that manipulations of PrL-PVT activity would affect extracellular DA levels in the NAcS, where PVT connections are most dense (Li and Kirouac, 2008). Specifically, we predicted that stimulation of the PrL-PVT circuit in STs would decrease DA, whereas inhibition of the PrL-PVT circuit in GTs would increase DA in the NAcS. To test these hypotheses, we used a dual-vector approach (Soudais et al., 2001; Boender et al., 2014; Kerstetter et al., 2016) to express either the stimulatory Gq- or inhibitory Gi/o- DREADD (Designer Receptors Exclusively Activated by Designer Drugs) in neurons of the PrL that project to the PVT, and examined how bidirectional manipulations of this pathway affect the attribution of incentive salience to a food-cue (Experiment 1; Figure 1). In addition, we used in-vivo microdialysis to assess extracellular levels of DA in the NAcS following manipulations of the PrL-PVT pathway (Experiment 2, Figure 5a–f). Figure 1 with 1 supplement see all Download asset Open asset Experiment 1 methods. (a) Timeline of the experimental procedures. Rats were trained in a Pavlovian Conditioned Approach (PavCA) paradigm for five consecutive days (Acquisition, Sessions 1–5) and phenotyped as sign- (STs) or goal-trackers (GTs). Following acquisition, STs and GTs underwent DREADD surgeries for delivering Gi- or Gq- DREADDs in neurons of the prelimbic cortex (PrL) projecting to the paraventricular nucleus of the thalamus (PVT). Incubation time for DREADD expression was 3–5 weeks. After incubation, rats were re-screened for sign- and goal-tracking behavior (Re-screening, Sessions 6–10). All rats received an i.p. VEH injection 25 min before session 10 to habituate them to the injection procedure. CNO (3 mg/kg) or VEH were then administered i.p. every day during the Test phase (Sessions 11–16), 25 min before the start of each session. 24 hr after the last session of PavCA, rats received an additional injection of CNO or VEH 25 min before being exposed to a Conditioned Teinforcement Test (CRT, Session 17). (b) Schematic of the dual-vector strategy used for expressing Gi- or Gq- DREADDs in the PrL-PVT pathway. (c,d) Photomicrographs representing mCherry expression in pyramidal neurons of the PrL projecting to the PVT at (c) 4x magnification and (d) 40x magnification. (e,f) Photomicrographs of mCherry expression representing terminal fibers in the anterior PVT coming from the PrL at (e) 10x magnification and at (f) 40x magnification. (g,h) Photomicrographs of mCherry expression representing terminal fibers in the posterior PVT coming from the PrL at (g) 10x magnification and (h) 40x magnification. https://doi.org/10.7554/eLife.49041.002 Results Experiment 1 Acquisition of pavlovian conditioned approach behaviors The average PavCA index from sessions 4–5 was used to classify rats as STs (PavCA index ≥ +0.30) or GTs (PavCA index ≤−0.30) (Figure 2b). As explained in the Methods below, the intermediate population of rats (PavCA index between −0.30 and +0.30) were excluded from this study. Linear mixed-effects models revealed a significant effect of phenotype, session and a significant phenotype x session interaction for all measures of sign- and goal-tracking behavior. Across the 5 sessions of PavCA training, STs had a greater number of lever contacts (F4,184.225=57.778, p<0.001), a greater probability to contact the lever (F4,303.698=68.278, p<0.001), and a lower latency to approach the lever (F4,336.578=63.676, p<0.001) (Figure 2c–e). These significant differences were apparent during sessions 1–5 of PavCA training (post-hoc analyses, p<0.001). In contrast, across the 5 sessions of PavCA training, GTs showed a greater number of magazine entries (F4,243.952=57.436, p<0.001), a greater probability to enter the magazine (F4,225.359=76.968, p<0.001), and a lower latency to enter the magazine (F4,348.976=71.788, p<0.001) (Figure 2f–h), and these significant differences were apparent during PavCA sessions 2–5 (post-hoc analyses, p<0.05). Figure 2 Download asset Open asset Acquisition of sign- and goal-tracking behaviors during 5 sessions of Pavlovian Conditioned Approach (PavCA) training (i.e. prior to surgery or CNO administration). (a) Schematic representing the PavCA task. Rats were presented with an illuminated lever (conditioned stimulus, CS) for 8 s followed by the delivery of a food pellet (unconditioned stimulus, US) immediately upon lever-CS retraction. Each PavCA session consisted of 25 lever-food pairings. (b) PavCA index scores (composite index of Pavlovian conditioned approach behavior) for individual rats across 5 sessions of Pavlovian conditioning. PavCA index from session 4–5 were averaged to determine the behavioral phenotype. Rats with a PavCA score <−0.3 were classified as goal-trackers (GTs, n = 59), rats with a PavCA score >+0.3 were classified as sign-trackers (STs, n = 56). (c–e) Acquisition of lever-directed behaviors (sign-tracking) during PavCA training. Mean ± SEM for (c) number of lever contacts, (d) probability to contact the lever, and (e) latency to contact the lever. (f–h) Acquisition of magazine-directed behaviors (goal-tracking) during PavCA training. Mean ± SEM for (f) number of magazine entries, (g) probability to enter the magazine, and (h) latency to enter the magazine. (i) Data are expressed as individual data points with mean ± SEM plotted on violin plots for PavCA index. Rats with similar PavCA scores were assigned to receive different G-protein coupled receptor (GPCR; Gi, Gq or no-DREADD) and different treatment (CNO, VEH). Baseline differences in PavCA index between experimental groups were assessed by using a 3-way ANOVA with phenotype (GT, ST), GPCR (Gi, Gq and no-DREADD) and treatment (CNO, VEH) as independent variables and PavCA index as the dependent variable. A significant effect of phenotype was found (p<0.001), but no significant differences between experimental groups and no significant interactions. Sample sizes: GT-Gi = 32, GT-Gq = 12, ST-Gi = 14, ST-Gq = 25, GT-no DREADD = 15, ST-no DREADD = 17. https://doi.org/10.7554/eLife.49041.004 Figure 2—source data 1 Acquisition of lever- and magazine-directed behaviors during 5 sessions of PavCA training (Figure 2c–h). https://doi.org/10.7554/eLife.49041.005 Download elife-49041-fig2-data1-v2.xlsx Figure 2—source data 2 Average PavCA index scores during session 4–5 of PavCA training (Figure 2b,i). https://doi.org/10.7554/eLife.49041.006 Download elife-49041-fig2-data2-v2.xlsx Experimental groups (i.e. G-protein coupled receptor (GPCR) and treatment groups) were counterbalanced based on the average PavCA index from sessions 4–5 (Figure 2i). A three-way ANOVA (phenotype x GPCR x treatment) showed a significant effect of phenotype (F1,114=2685.054, p<0.001) on PavCA index (Figure 2i), but no significant effects of GPCR or treatment groups, and no significant interactions (see also Supplementary file 1 and 2). PavCA rescreening vs. PavCA test PavCA index is presented as the primary dependent variable in the main text, but analyses for other dependent variables indicative of Pavlovian conditioned approach behavior including, contacts, probability and latency directed towards either the lever-CS or food magazine are included in Supplementary file 3 and 4. PavCA index during each daily session of rescreening is presented in Figure 3—figure supplement 1. Stimulation of the PrL-PVT pathway attenuates the incentive value of the food cue in STs Stimulation of the PrL-PVT pathway in STs significantly decreased the PavCA index (Figure 3c), which, in this case, is reflective of both a decrease in lever-directed behaviors (Supplementary file 3) and an increase in goal-directed behaviors (Supplementary file 4). There was a significant effect of treatment (F1,23 = 33.251, p<0.001), session (F1,23 = 10.799, p = 0.03), and a significant treatment x session interaction (F1,23 = 14.051, p = 0.01, 1-β = 1) for the PavCA index. Post-hoc analyses revealed a significant difference between VEH- and CNO-treated STs during both rescreening (p=0.005, Cohen's d = 1.27) and test (p<0.001, Cohen's d = 2.74). The significant difference between treatment groups during rescreening, prior to actual treatment, is due, in part, to the fact that counterbalancing was disrupted once animals were eliminated because of inaccurate DREADD expression. Importantly, however, only the CNO-treated rats exhibited a change in behavior during the test sessions relative to rescreening (p<0.001, Cohen's d = 1.18). For GTs (Figure 3d), there was not a significant effect of treatment (F1,10 = 0.169, p = 0.690), session (F1,10 = 0.511, p = 0.491) nor a significant treatment x session interaction (F1,10 = 0.351, p = 0.567). The modest sample size (n = 6) may have contributed to the lack of effects in the GT-Gq rats, as suggested by a post-hoc power analysis (1-β = 0.27). Nonetheless, taken together, these results suggest that 'turning on' the top-down PrL-PVT circuit appears to selectively attenuate the incentive value of the cue in STs. Figure 3 with 2 supplements see all Download asset Open asset Chemogenetic stimulation of the PrL-PVT pathway decreases sign-tracking behavior in sign-trackers, while chemogenetic inhibition of the PrL-PVT pathway increases sign-tracking behavior in goal-trackers. (a,b) Drawing representative of sign-tracking (a) and goal-tracking (b) conditioned responses during PavCA training. (c–h) Data are expressed as individual data points with mean ± SEM plotted on violin plots for PavCA index during Rescreening (Res., average of PCA Sessions 6–10) and Test (average of PCA Sessions 11–16) periods. (c,d) Chemogenetic stimulation (Gq) of the PrL-PVT circuit decreases the PavCA index in (c) sign-trackers, but has no effect in (d) goal-trackers. There was a significant treatment x session interaction in sign-trackers (p<0.01). Pairwise comparisons showed that CNO decreased the PavCA index in STs (*p<0.001, CNO vs. VEH; #p<0.001, Test vs. Res.). (e,f) Chemogenetic inhibition of the PrL-PVT circuit has no effect on (e) sign-trackers, but increases the PavCA index in (f) goal-trackers. There was a significant treatment x session interaction in goal-trackers (*p<0.012 CNO vs. VEH; #p<0.002 Test vs. Res.). (g,h) Effects of CNO administration on the PavCA index in non-DREADD expressing (g) STs and GTs (h). When DREADD receptors were not expressed in the brain, CNO had no effect on the PavCA index in either sign-trackers nor goal-trackers. Sample sizes: GT-Gi = 32, GT-Gq = 12, ST-Gi = 14, ST-Gq = 25, GT-no DREADD = 15, ST-no DREADD = 17. https://doi.org/10.7554/eLife.49041.007 Figure 3—source data 1 Behavioral data during the rescreening and test sessions for ST-Gq (Figure 3c). https://doi.org/10.7554/eLife.49041.010 Download elife-49041-fig3-data1-v2.xlsx Figure 3—source data 2 Behavioral data during the rescreening and test sessions for GT-Gq (Figure 3d). https://doi.org/10.7554/eLife.49041.011 Download elife-49041-fig3-data2-v2.xlsx Figure 3—source data 3 Behavioral data during the rescreening and test sessions for ST-Gi (Figure 3e). https://doi.org/10.7554/eLife.49041.012 Download elife-49041-fig3-data3-v2.xlsx Figure 3—source data 4 Behavioral data during the rescreening and test sessions for GT-Gi (Figure 3f). https://doi.org/10.7554/eLife.49041.013 Download elife-49041-fig3-data4-v2.xlsx Figure 3—source data 5 Behavioral data during the rescreening and test sessions for ST-No DREADD controls (Figure 3g). https://doi.org/10.7554/eLife.49041.014 Download elife-49041-fig3-data5-v2.xlsx Figure 3—source data 6 Behavioral data during the rescreening and test sessions for GT-No DREADD controls (Figure 3h). https://doi.org/10.7554/eLife.49041.015 Download elife-49041-fig3-data6-v2.xlsx Inhibition of the PrL-PVT pathway increases the incentive value of the food cue in GTs Inhibition of the PrL-PVT pathway in GTs significantly increased the PavCA index (Figure 3f). This effect appears to be driven primarily by a change in the 'response bias' score (F1,30=4.136 p=0.051, Cohens d = 1.04, 1-β=0.99; data not shown), which is a measure of: [(total lever-CS contacts − total food magazine entries) / (total lever-CS contacts + total food magazine entries)] (Meyer et al., 2012). Other specific metrics of lever- or magazine-directed behaviors were not significantly different between treatment groups (Supplementary file 3 and 4). For PavCA index, however, there was a significant effect of treatment (F1,30=5.975, p=0.021), session (F1,30=7.106, p=0.012,), and a significant treatment x session interaction (F1,30=4.403, p=0.044, 1-β=0.98). Post-hoc analyses revealed that inhibiting the PrL-PVT pathway increased the PavCA index relative to both the rescreening session (p=0.002, Cohen's d = 0.85) and to VEH controls during the test session (p=0.012, Cohen's d = 0.96). For STs (Figure 3e), there was not a significant effect of treatment (F1,12=0.876, p=0.368), session (F1,12=4.359, p=0.059), nor a significant treatment x session interaction (F1,12=1.479, p=0.247, 1-β=0.85), suggesting that 'turning off' the PrL-PVT pathway permits the attribution of incentive motivational value to a reward-cue selectively in GTs. CNO administration in the absence of DREADD receptors does not affect the Pavlovian conditioned approach response in STs or GTs Administration of CNO in the absence of DREADD had no effect on behavior during PavCA in either STs or GTs (Figure 3g and h, respectively). Behavior during the intertrial interval was not affected by manipulation of the PrL-PVT pathway or by administration of CNO in the absence of DREADD receptors To assess the effects of manipulating the PrL-PVT pathway on general locomotor activity and motivated behavior, we examined head entries into the food magazine during the intertrial interval (ITI), when the lever-CS was not present (Figure 3—figure supplement 2). Consistent with prior findings, head entries into the food magazine during the ITI tended to decrease with training; thus responses during the 'test' sessions were generally less than those during the 'rescreening' sessions (see statistics in Figure 3—figure supplement 2 legend). Importantly, however, there were no significant effects of treatment and no significant interactions for this metric for either phenotype or any of the experimental groups (i.e. Gq, Gi, No-DREADD controls). It should also be noted, that all rats continued to consume all of their food pellets during the ITI, regardless of treatment. Thus, the effects described above following manipulation of the PrL-PVT circuit appear to be specific to Pavlovian conditioned approach behavior and not reflective of a change in general locomotor activity or motivated behavior. Conditioned reinforcement test A conditioned reinforcement test (CRT) was conducted to assess the reinforcing properties of the lever-CS (Robinson and Flagel, 2009). During this test, responses into a port designated 'active' results in the brief presentation of the lever-CS; whereas those in the 'inactive' port have no consequence. If a rat responds more into the active port relative to the inactive port, the lever-CS is considered to have reinforcing properties (Robinson and Flagel, 2009). Moreover, if the rat approaches and interacts with the lever-CS during its brief presentation, it is considered to have incentive properties (Robinson and Flagel, 2009; Hughson et al., 2019). Here we use the incentive value index, a composite metric ((pokes in active port + lever-CS contacts) – (pokes in inactive port)) as a primary measure of the conditioned reinforcing properties of the lever-CS (Figure 4; see also Hughson et al., 2019) and additionally report nosepoke responding and lever-CS contacts in Figure 4—figure supplement 1. Figure 4 with 1 supplement see all Download asset Open asset Chemogenetic stimulation of the PrL-PVT pathway decreases the conditioned reinforcing properties of a reward-paired cue in sign-trackers. (a) Schematic representing the conditioned reinforcement test (CRT). Data are expressed as individual data points with mean ± SEM plotted on violin plots for Incentive value index ((active nosepokes + lever presses) – (inactive nosepokes)). (b,c) Relative to VEH controls, administration of CNO (3 mg/kp, i.p.) significantly decreased the incentive value of the reward cue for (b) ST-Gq (*p<0.05), but not (c) GT-Gq rats. (d,e) CNO-induced inhibition of the PrL-PVT circuit did not affect the conditioned reinforcing properties of the reward cue in (d) ST-Gi or (e) GT-Gi rats. (f,g) CNO administration did not affect the conditioned reinforcing properties of the reward cue in non-DREADD expressing (f) STs or (g) GTs. Sample sizes: GT-Gi = 32, GT-Gq = 12, ST-Gi = 14, ST-Gq = 25, GT-no DREADD = 15, ST-no DREADD = 17. https://doi.org/10.7554/eLife.49041.016 Figure 4—source data 1 Lever presses (Figure 4—figure supplement 1h), nosepokes (Figure 4—figure supplement 1a) and incentive value index (Figure 4b) during conditioned reinforcement for ST-Gq. https://doi.org/10.7554/eLife.49041.018 Download elife-49041-fig4-data1-v2.xlsx Figure 4—source data 2 Lever presses (Figure 4—figure supplement 1h), nosepokes (Figure 4—figure supplement 1b) and incentive value index (Figure 4c) during conditioned reinforcement for GT-Gq. https://doi.org/10.7554/eLife.49041.019 Download elife-49041-fig4-data2-v2.xlsx Figure 4—source data 3 Lever presses (Figure 4—figure supplement 1i), nosepokes (Figure 4—figure supplement 1c) and incentive value index (Figure 4d) during conditioned reinforcement for ST-Gi. https://doi.org/10.7554/eLife.49041.020 Download elife-49041-fig4-data3-v2.xlsx Figure 4—source data 4 Lever presses (Figure 4—figure supplement 1j), nosepokes (Figure 4—figure supplement 1d) and incentive value index (Figure 4e) during conditioned reinforcement for GT-Gi. https://doi.org/10.7554/eLife.49041.021 Download elife-49041-fig4-data4-v2.xlsx Figure 4—source data 5 Lever presses (Figure 4—figure supplement 1k), nosepokes (Figure 4—figure supplement 1e) and incentive value index (Figure 4f) during conditioned reinforcement for ST-No DREADD controls. https://doi.org/10.7554/eLife.49041.022 Download elife-49041-fig4-data5-v2.xlsx Figure 4—source data 6 Lever presses (Figure 4—figure supplement 1l), nosepokes (Figure 4—figure supplement 1f) and incentive value index (Figure 4g) during conditioned reinforcement for GT-No DREADD controls. https://doi.org/10.7554/eLife.49041.023 Download elife-49041-fig4-data6-v2.xlsx Stimulation of the PrL-PVT pathway attenuates the conditioned reinforcing properties of a reward-cue in STs Stimulation of the PrL-PVT pathway in STs significantly attenuated the incentive value index during the CRT (t25 = −3.574, p=0.002, Cohen's d = 1.48, 1-β=0.92, Figure 4b). The same manipulation had no effect on the incentive value index in GTs (Figure 4c). In agreement, there was a significant effect of treatment (F1,36 = 19.021, p<0.001), port (F1,36 = 30.501, p<0.001) and a significant treatment x port interaction (F1,36 = 7.024, p = 0.012) for nosepoke responding for STs (Figure 4—figure supplement 1a). Post-hoc analysis revealed that, relative to VEH controls, stimulation of the PrL-PVT in STs decreased the number of nosepokes into the active port (p<0.001, Cohen's d = 1.55). Furthermore, while VEH-treated controls responded more in the active port relative to the inactive port (p<0.001, Cohen's d = 2.421), this discrimination between ports was abolished following CNO treatment (p=0.061). Similarly, stimulation of the PrL-PVT pathway reduced the number of lever contacts in STs (t18.142 = -3.615, p<0.05, Cohen's d = 1.47, Figure 4—figure supplement 1g). For GTs, there was not a significant effect of treatment, port, nor a significant treatment x port interaction for nosepoke responding (Figure 4—figure supplement 1b), nor a significant effect of treatment for lever-CS interactions (Figure 4—figure supplement 1h). These data are consistent with those reported above for PavCA behavior, as 'turning on' this top-down cortico-thalamic control attenuates the incentive value of a food-cue selectively in STs. Inhibition of the PrL-PVT pathway does not affect the conditioned reinforcing properties of a reward-cue in either STs or GTs Inhibition of the PrL-PVT pathway did not affect the incent