There are often sudden changes in the state of environment. For a decision maker, accurate prediction and detection of change points are crucial for optimizing performance. Still unclear, however, is whether rodents are simply reactive to reinforcements, or if they can be proactive to estimate future change points during value-based decision making. In this study, we characterize head-fixed mice performing a two-armed bandit task with probabilistic reward reversals. Choice behavior deviates from classic reinforcement learning, but instead suggests a strategy involving belief updating, consistent with the anticipation of change points to exploit the task structure. Excitotoxic lesion and optogenetic inactivation implicate the anterior cingulate and premotor regions of medial frontal cortex. Specifically, over-estimation of hazard rate arises from imbalance across frontal hemispheres during the time window before the choice is made. Collectively, the results demonstrate that mice can capitalize on their knowledge of task regularities, and this estimation of future changes in the environment may be a main computational function of the rodent dorsal medial frontal cortex.
Selective serotonin transport (SERT) inhibitors such as fluoxetine are the most commonly prescribed treatments for depression. Although efficacious for many symptoms of depression, motivational impairments such as psychomotor retardation, anergia, fatigue and amotivation are relatively resistant to treatment with SERT inhibitors, and these drugs have been reported to exacerbate motivational deficits in some people. In order to study motivational dysfunctions in animal models, procedures have been developed to measure effort-related decision making, which offer animals a choice between high effort actions leading to highly valued reinforcers, or low effort/low reward options. In the present studies, male and female rats were tested on two different tests of effort-based choice: a fixed ratio 5 (FR5)/chow feeding choice procedure and a running wheel (RW)/chow feeding choice task. The baseline pattern of choice differed across tasks for males and females, with males pressing the lever more than females on the operant task, and females running more than males on the RW task. Administration of the SERT inhibitor and antidepressant fluoxetine suppressed the higher effort activity on each task (lever pressing and wheel running) in both males and females. The serotonin receptor mediating the suppressive effects of fluoxetine is uncertain, because serotonin antagonists with different patterns of receptor selectivity failed to reverse the effects of fluoxetine. Nevertheless, these studies uncovered important sex differences, and demonstrated that the suppressive effects of fluoxetine on high effort activities are not limited to tasks involving food reinforced behavior or appetite suppressive effects. It is possible that this line of research will contribute to an understanding of the neurochemical factors regulating selection of voluntary physical activity vs. sedentary behaviors, which could be relevant for understanding the role of physical activity in psychiatric disorders.
Learning from experience is essential to the optimization of behavior. In particular, we learn from past choices and outcomes to infer the predicted values of the actions to be taken. Then based on the values, we may select an informed choice. However, despite the many neural correlates identified, we still do not have a clear picture for how values are computed and translated into informed behavior. Here, we trained head-fixed mice to perform a two-armed bandit task. Animals based their decisions on past choices and reinforcements, consistent with having an internal representation of action values. To determine the causal contributions of the medial prefrontal cortex, we tested the animals before and after an excitotoxic lesion of the medial secondary motor cortex (M2). We found that unilateral M2 lesion led to side-specific effects on the animal’s ability to learn from past choices. To quantify the decision-making process, we fitted the animal’s choice behavior with Q-learning models to extract learning parameters such as learning rate, forgetting rate, and inverse temperature. Altogether, the results provide insights into the causal involvement of mouse mM2 in value-based decision making.