When receiving a reward after a sequence of multiple events, how do we determine which event caused the reward? This problem, known as temporal credit assignment, can be difficult for humans to solve given the temporal uncertainty in the environment. Research to date has attempted to isolate dimensions of delay and reward during decision-making, but algorithmic solutions to temporal learning problems and the effect of uncertainty on learning remain underexplored. To further our understanding, we adapted a reward learning task that creates a temporal credit assignment problem by combining sequentially delayed rewards, intervening events, and varying uncertainty via the amount of information presented during feedback. Using computational modeling, two learning strategies were developed: an eligibility trace, whereby previously selected actions are updated as a function of the temporal sequence, and a tabular update, whereby only systematically related past actions (rather than unrelated intervening events) are updated. We hypothesized that reduced information uncertainty would correlate with increased use of the tabular strategy, given the model's capacity to incorporate additional feedback information. Both models effectively learned the task, and predicted choices made by participants (N = 142) as well as specific behavioral signatures of credit assignment. Consistent with our hypothesis, the tabular model outperformed the eligibility model under low information uncertainty, as evidenced by more accurate predictions of participants' behavior and an increase in tabular weight. These findings provide new insights into the mechanisms implemented by humans to solve temporal credit assignment and adapt their strategy in varying environments.
Reinforcement learning is used to align language models with human preference signals after first pre-training the model to predict the next token of text within a large corpus using likelihood maximization. Before being deployed in a specific domain, models are often further fine-tuned on task specific data. Since human preferences are often unavailable for the last step, it is performed using likelihood maximization as that is the typical default method. However, reinforcement learning has other advantages besides facilitating alignment to a human derived reward function. For one, whereas likelihood maximization is a form of imitation learning in which the model is trained on what to do under ideal conditions, reinforcement learning is not limited to demonstrating actions just for optimally reached states and trains a model what to do under a range of scenarios as it explores the policy space. In addition, it also trains a model what not to do, suppressing competitive but poor actions. This work develops a framework for last-mile fine-tuning using reinforcement learning and tests whether it garners performance gains. The experiments center on abstractive summarization, but the framework is general and broadly applicable. Use of the procedure produced significantly better results than likelihood maximization when comparing raw predictions. For the specific data tested, the gap could be bridged by employing post-processing of the maximum likelihood outputs. Nonetheless, the framework offers a new avenue for model optimization in situations where post-processing may be less straightforward or effective, and it can be extended to include more complex classes of undesirable outputs to penalize and train against, such as hallucinations.
The anterior cingulate cortex (ACC) has been implicated across multiple highly specialized cognitive functions-including task engagement, motivation, error detection, attention allocation, value processing, and action selection. Here, we ask if ACC lesions disrupt task performance and firing in dorsomedial striatum (DMS) during the performance of a reward-guided decision-making task that engages many of these cognitive functions. We found that ACC lesions impacted several facets of task performance-including decreasing the initiation and completion of trials, slowing reaction times, and resulting in suboptimal and inaccurate action selection. Reductions in movement times towards the end of behavioral sessions further suggested attenuations in motivation, which paralleled reductions in directional action selection signals in the DMS that were observed later in recording sessions. Surprisingly, however, beyond altered action signals late in sessions-neural correlates in the DMS were largely unaffected, even though behavior was disrupted at multiple levels. We conclude that ACC lesions result in overall deficits in task engagement that impact multiple facets of task performance during our reward-guided decision-making task, which-beyond impacting motivated action signals-arise from dysregulated attentional signals in the ACC and are mediated via downstream targets other than DMS.
Our prior research has identi fi ed neural correlates of cognitive control in the anterior cingulate cortex (ACC), leading us to hypothesize that the ACC is necessary for increasing attention as rats fl exibly learn new contingencies during a complex reward -guided decision -making task. Here, we tested this hypothesis by using optogenetics to transiently inhibit the ACC, while rats of either sex performed the same two -choice task. ACC inhibition had a profound impact on behavior that extended beyond de fi cits in attention during learning when expected outcomes were uncertain. We found that ACC inactivation slowed and reduced the number of trials rats initiated and impaired both their accuracy and their ability to complete sessions. Furthermore, drift - diffusion model analysis suggested that free -choice performance and evidence accumulation (i.e., reduced drift rates) were degraded during initial learning - leading to weaker associations that were more easily overridden in later trial blocks (i.e., stronger bias). Together, these results suggest that in addition to attention -related functions, the ACC contributes to the ability to initiate trials and generally stay on task.
Background: Functional connectivity has garnered interest as a potential biomarker of psychiatric disorders including borderline personality disorder (BPD). However, small sample sizes and lack of within-study replications have led to divergent findings with no clear spatial foci. Aims: Evaluate discriminative performance and generalizability of functional connectivity markers for BPD. Method: Whole-brain fMRI resting state functional connectivity in matched subsamples of 116 BPD and 72 control individuals defined by three grouping strategies. We predicted BPD status using classifiers with repeated cross-validation based on multiscale functional connectivity within and between regions of interest (ROIs) covering the whole brain-global ROI-based network, seed-based ROI-connectivity, functional consistency, and voxel-to-voxel connectivity-and evaluated the generalizability of the classification in the left-out portion of nonmatched data. Results: Full-brain connectivity allowed classification (-70 %) of BPD patients vs. controls in matched inner cross-validation. The classification remained significant when applied to unmatched out-of-sample data (-61-70 %). Highest seed-based accuracies were in a similar range to global accuracies (-70-75 %), but spatially more specific. The most discriminative seed regions included midline, temporal and somatomotor regions. Univariate connectivity values were not predictive of BPD after multiple comparison corrections, but weak local effects coincided with the most discriminative seed-ROIs. Highest accuracies were achieved with a full clinical interview while self-report results remained at chance level. Limitations: The accuracies vary considerably between random sub-samples of the population, global signal and covariates limiting the practical applicability. Conclusions: Spatially distributed functional connectivity patterns are moderately predictive of BPD despite heterogeneity of the patient population.
Treatments for depression were developed largely through trial and error, with limited knowledge of necessary processes to target for symptom improvement. Learning and valuation are possible treatment targets in depression. Reinforcement learning (RL) may help specify and target learning dysfunctions in depression.
Previous studies have demonstrated that the rate of evidence integration during perceptual decision making, a specific computationally defined parameter, is negatively correlated with both subclinical symptoms of OCD measured on a continuum and categorically diagnosed patient status. However, the neural mechanisms underlying this deficit are unknown. Separate work has shown that both gamma and beta-band power are related to evidence integration, and differences in beta-band power in particular have been hypothesized to hinder flexible behavioral control. We sought to unify these two disparate literatures, one on OCD-related information processing differences constrained by behavioral data alone, and the other on the neural correlates of evidence integration. Using computational modeling and scalp EEG, we tested (N = 67) the relationships between subclinical symptom scores, drift rate, and gamma/beta-band activity during perceptual decision making. We replicated both prior work showing deficits in evidence integration as a function of OCD symptoms, and work showing a relationship between evidence integration and gamma and beta-band power. As predicted, the slope of beta-band power was correlated with OCD symptoms. However, the relationships between OCD symptoms and drift rate and the slopes of gamma and beta-band power and drift rate remained unchanged when simultaneously accounting for all variables, speaking against the hypothesis that differences in band-band power explain drift rate deficits.
Computational models of decision making have identified a relationship between obsessive-compulsive symptoms (OCS), both in the general population and in patients, and impairments in perceptual evidence accumulation. Some studies have interpreted these deficits to reflect global disease traits which give rise to clusters of OCS. Such assumptions are not uncommon, even if implicit, in computational psychiatry more broadly. However, it is well established that state- and trait-symptom scores are often correlated (e.g., state and trait anxiety), and the extent to which perceptual deficits are actually explained by state-based symptoms is unclear. State-based symptoms may give rise to information processing differences in a number of ways, including the mechanistically less interesting possibility of tying up working memory and attentional resources for off-task processing. In a general population sample (N = 150), we investigated the extent to which previously identified impairments in perceptual evidence accumulation were related to trait vs stated-based OCS. In addition, we tested whether differences in working memory capacity moderated state-based impairments, such that impairments were worse in individuals with lower working memory capacity. We replicated previous work demonstrating a negative relationship between the rate of evidence accumulation and trait-based OCS when state-based symptoms were unaccounted for. When state-based effects were included in the model, they captured a significant degree of impairment while trait-based effects were attenuated, although they did not disappear completely. We did not find evidence that working memory capacity moderated the state-based effects. Our work suggests that investigating the relationship between information processing and state-based symptoms may be important more generally in computational psychiatry beyond this specific context.
A large literature has accumulated suggesting that human and animal decision making is driven by at least two systems, and that important functions of these systems can be captured by reinforcement learning algorithms. The "model-free" system caches and uses stimulus-value or stimulus-response associations, and the "model-based" system implements more flexible planning using a model of the world. However, it is not clear how the two systems interact during deliberation and how a single decision emerges from this process, especially when they disagree. Most previous work has assumed that while the systems operate in parallel, they do so independently, and they combine linearly to influence decisions. Using an integrated reinforcement learning/drift-diffusion model, we tested the hypothesis that the two systems interact in a non-linear fashion similar to other situations with cognitive conflict. We differentiated two forms of conflict: action conflict, a binary state representing whether the systems disagreed on the best action, and value conflict, a continuous measure of the extent to which the two systems disagreed on the difference in value between the available options. We found that decisions with greater value conflict were characterized by reduced model-based control and increased caution both with and without action conflict. Action conflict itself (the binary state) acted in the opposite direction, although its effects were less prominent. We also found that between-system conflict was highly correlated with within-system conflict, and although it is less clear a priori why the latter might influence the strength of each system above its standard linear contribution, we could not rule it out. Our work highlights the importance of non-linear conflict effects, and provides new constraints for more detailed process models of decision making. It also presents new avenues to explore with relation to disorders of compulsivity, where an imbalance between systems has been implicated.
Deficits in primary recognition memory and confidence have previously been tested as potential contributors to excessive checking behavior in obsessive-compulsive disorder. Studies have tested both recognition for actions and, hypothesizing that recognition may be disrupted more generally across content domains, verbal recognition memory. However, studies of verbal recognition memory have yielded mixed results. We revisited this work with the benefit of hindsight, running two new experiments with larger samples, the manipulation of recognition difficulty, and a computational model-based approach to data analysis. In both datasets, we found that discriminability, defined as the difference in drift rate for old versus new stimuli in the drift-diffusion model, was reduced as a function of subclinical OCD symptoms in the general population. Paralleling work on drift rate deficits in perceptual decision making in OCD, these reductions were larger for easier recognition decisions. We also asked participants about their confidence in each recognition decision and parcellated confidence into bias, or the difference in overall confidence, and sensitivity, which represents the ability to appropriately map confidence to objective accuracy. We found no consistent evidence of a relationship between OCD symptoms and either quantity.
BackgroundFunctional connectivity measures have garnered interest as possible biomarkers of psychiatric disorders including borderline personality disorder (BPD). However, small sample sizes and lack of within-study replications have led to divergent findings with no clear spatial foci. Therefore, we adopted an exploratory full-brain approach in the current study to evaluate which combinations of regions are most consistently predictive of BPD diagnosis.MethodsWe studied fMRI resting state functional connectivity in matched subsamples of 116 BPD and 72 control individuals defined by three grouping strategies: 1) referral diagnosis, 2) clinical diagnostic interview excluding patients no longer filling diagnostic criteria or controls scoring above threshold in a screening questionnaire and 3) self-reported symptom severity. We predicted BPD status using classifiers with repeated cross-validation based on multiscale functional connectivity within and between regions of interest (ROIs) covering the whole brain— global ROI-based network, seed-based ROI-connectivity, functional consistency and voxel-to-voxel connectivity within and between ROIs. Finally, we evaluated the generalizability of the classification in the left-out portion of non-matched data.ResultsFull-brain connectivity allowed successful classification (~70%) of BPD patients vs. control individuals in matched inner cross-validation. The classification remained significant when applied to unmatched out-of-sample data, but accuracies were lower (~61–70%) than in fully matched samples. The over-estimation of inner cross-validation accuracy was exacerbated by univariate regression of nuisance variables, particularly in smaller samples. Highest seed-based accuracies were in a similar range to global accuracies (~70–75%), but spatially more specific. In the seed-based classification, the regions implicated most often included midline, temporal and somatomotor regions. Highest accuracies were achieved with the clinical interview followed by referral diagnosis group definition. Self-report results remained at chance level. The accuracies were affected by an interaction of medications and global signal and univariate nuisance regression. Pairwise correlations, local consistencies and fine-scale connectivity matrices were not significantly predictive of BPD after multiple comparison corrections, but weak local effects coincided with the most discriminative ROIs in the classification. ConclusionsOur multivariate results indicate that complex global functional connectivity differences are moderately predictive of BPD despite heterogeneity of the patient population. However, univariate nuisance regression applied to full cross-validation dataset can cause inflation of accuracies compared with left-out test data.
Key Points Question Are depression symptoms associated with features of reinforcement learning, and if so, is treatment-related symptom change associated with learning changes? Findings In this mixed cross-sectional–cohort study including 101 participants, participants with and without depression completed a probabilistic learning task during functional magnetic resonance imaging; participants with depression were reassessed after cognitive behavioral therapy (CBT). Computational model–based analyses of behavioral choices and neural data identified associations of learning with symptoms during reward learning and loss learning, respectively; symptom improvement following CBT was associated with normalization of learning parameters. Meaning Mapping reinforcement learning processes to symptoms of depression reveals mechanistic features of these symptoms and points to possible learning-based therapeutic processes and targets.
Obsessive-compulsive disorder is an illness characterized by intrusive thoughts and a heightened focus on thought experiences (i.e., cognitive self-consciousness; CSC). Individuals with obsessive-compulsive symptoms demonstrate broad deficits in executive functioning, including impairments in basic decision processes. Yet, it is unclear whether individual differences in such impairments are greater accounted for at the trait-level or state-level of OCD symptomatology and the extent to which CSC contributes.
Real-life decisions are often repeated. Whether considering taking a job in a new city, or doing something mundane like checking if the stove is off, decisions are frequently revisited even if no new information is available. This mode of behavior takes a particularly pathological form in obsessive-compulsive disorder (OCD), which is marked by individuals' redeliberating previously resolved decisions. Surprisingly, little is known about how information is transferred across decision episodes in such circumstances, and whether and how such transfer varies in OCD. In two experiments, data from a repeated decision-making task and computational modeling revealed that both implicit and explicit memories of previous decisions affected subsequent decisions by biasing the rate of evidence integration. Further, we replicated previous work demonstrating impairments in baseline decision-making as a function of self-reported OCD symptoms, and found that information transfer effects specifically due to implicit memory were reduced, offering computational insight into checking behavior.
Reward-based decision making is thought to be driven by at least two different types of decision systems: a simple stimulus-response cache-based system which embodies the common-sense notion of "habit," for which model-free reinforcement learning serves as a computational substrate, and a more deliberate, prospective, model-based planning system. Previous work has shown that loss aversion, a well-studied measure of how much more on average individuals weigh losses relative to gains during decision making, is reduced when participants take all possible decisions and outcomes into account including future ones, relative to when they myopically focus on the current decision. Model-based control offers a putative mechanism for implementing such foresight. Using a well-powered data set (N = 117) in which participants completed two different tasks designed to measure each of the two quantities of interest, and four models of choice data for these tasks, we found consistent evidence of a relationship between loss aversion and model-based control but in the direction opposite to that expected based on previous work: loss aversion had a positive relationship with model-based control. We did not find evidence for a relationship between either decision system and risk aversion, a related aspect of subjective utility.
This chapter reviews recent work on computational models describing the temporal dynamics of reward-based goal-directed decision-making, from a conceptual rather than technical perspective. Models of temporal dynamics make predictions about the time course and duration of deliberation in addition to the choices that people make. We begin by reviewing the simple one-step choice decision situation and the idea of evidence integration, a core feature of dynamic models, within the context of the drift-diffusion model. We then present more recent extensions of the evidence integration perspective to decisions that span more than one time step.
The laboratory study of how humans and other animals trade-off value and time has a long and storied history, and is the subject of a vast literature. However, despite a long history of study, there is no agreed upon mechanistic explanation of how intertemporal choice preferences arise. Several theorists have recently proposed model-based reinforcement learning as a candidate framework. This framework describes a suite of algorithms by which a model of the environment, in the form of a state transition function and reward function, can be converted on-line into a decision. The state transition function allows the model-based system to make decisions based on projected future states, while the reward function assigns value to each state, together capturing the necessary components for successful intertemporal choice. Empirical work has also pointed to a possible relationship between increased prospection and reduced discounting. In the current paper, we look for direct evidence of a relationship between temporal discounting and model-based control in a large new data set (n = 168). However, testing the relationship under several different modeling formulations revealed no indication that the two quantities are related.
SEE CORRESPONDING ARTICLE ON PAGE 781 Selective Inhibition of Amygdala Neuronal Ensembles Encoding Nicotine-Associated Memories Inhibits Nicotine Preference and RelapseBiological PsychiatryVol. 82Issue 11PreviewNicotine craving and relapse often occurs after reactivation of nicotine reward memories. We recently developed a memory retrieval–reconsolidation interference procedure in which reactivating nicotine reward memories by acute exposure to nicotine (the unconditioned stimulus [UCS]) and then pharmacologically interfering with memory reconsolidation decreased relapse to nicotine seeking in rats and nicotine craving in smokers. Here, we investigated underlying mechanisms. Full-Text PDF
Research on the dynamics of reward-based, goal-directed decision making has largely focused on simple choice, where participants decide among a set of unitary, mutually exclusive options. Recent work suggests that the deliberation process underlying simple choice can be understood in terms of evidence integration: Noisy evidence in favor of each option accrues over time, until the evidence in favor of one option is significantly greater than the rest. However, real-life decisions often involve not one, but several steps of action, requiring a consideration of cumulative rewards and a sensitivity to recursive decision structure. We present results from two experiments that leveraged techniques previously applied to simple choice to shed light on the deliberation process underlying multistep choice. We interpret the results from these experiments in terms of a new computational model, which extends the evidence accumulation perspective to multiple steps of action.