Abstract Action selection is traditionally thought to be governed by dual top-down control systems: a fast, affect-dependent controller and a slow, experience-driven controller. Emerging evidence indicates that bottom-up cardiac phase signals also contribute to action selection by facilitating (i.e., go) or suppressing (i.e., nogo) motor responses. Despite growing interest in cardiac-brain interactions, no prior study has examined whether cardiac phase effects on feedback-driven action decisions are modulated by appetitive and aversive motivational states. We conducted a secondary analysis of a go/nogo functional magnetic resonance imaging (fMRI) dataset ( n = 29), retrospectively determining the onset of decision cues relative to physiologically defined cardiac phases. We tested competing behavioural and computational hypotheses proposing that cardiac phase modulated action selection via motivational action biases, or through a general, affect-independent prepotent urge to act. Additionally, we performed an exploratory fMRI analysis to investigate neural activity underlying action decisions as a function of cardiac phase. Converging behavioural and computational evidence showed that systole was associated with a higher likelihood of go responding, independent of motivational action biases. Preliminary fMRI results further suggested that systolic facilitation of action was accompanied by diminished neural activity supporting action inhibition. Together, these findings extend existing evidence that cardiac phase impacts action decisions.
People often make decisions in contexts where rewards and punishments co-occur, yet most human research still examines reward and punishment learning as independent processes. Here, across three studies, we address this gap by demonstrating that punishments amplify reward learning and its neural correlates in healthy human participants. In Study 1 (N = 102, 69 females and 33 males), participants performed a probabilistic learning task involving monetary rewards and punishments presented in either intermixed or separated contexts. In intermixed contexts, punishments enhanced reward learning accompanied by changes in computational parameters, including higher learning rates from reward prediction errors. In Study 2 (N = 26, 18 females and 8 males), fMRI revealed that punishments amplified reward prediction errors signals in the caudate. Study 3, an fMRI meta-analysis, confirmed that striatal reward responses are consistently stronger when punishments are present. Across studies, we found no reciprocal enhancement of punishment learning by rewards. Together, these findings demonstrate that punishments sharpen reward learning through striatal modulation and underscore the extent to which reward learning is influenced by its broader outcome context.
The brain must detect and evaluate rewards amidst multiple stimuli to generate adaptive behavior. Physically salient stimuli draw greater attention, potentially influencing their subsequent evaluation. An influential framework proposes that rewards are processed through a two-component dopaminergic response: an early, value-agnostic salience signal, followed by a partially overlapping but temporally lagged signal reflecting stimulus evaluation. Yet, evidence for this framework in humans is lacking due to spatiotemporal limitations of neuroimaging methods. Using bespoke simultaneous 7T functional magnetic resonance imaging (fMRI)-electroencephalogram (EEG) developments, we decoupled reward-anticipatory midbrain signals with distinct spatiotemporal profiles: an early signal in posterior substantia nigra (SN) consistent with physical salience, followed by a lagged signal in anterior SN and ventral tegmental area, likely reflecting evaluative processes such as value and/or motivational salience. We also demonstrate that the early physical salience signal enhances this later evaluative response, offering the first evidence of attention-guided reward processing in the human midbrain.
Computerised Cognitive Behavioural therapy (CBT) is an effective psychological intervention for mild to moderate depression. While CBT aims to correct maladaptive cognitive biases and ensuing disadvantageous decision-making, our current understanding of decision-making signatures linked to CBT response remains limited. Preliminary behavioural evidence has shown that the process of evidence accumulation (EA), indexing the efficiency of decision dynamics, is impaired in depression. However, little is known about the role of EA in the context of CBT for depression. In this study we recruited 37 (18 females) unmedicated depressed subjects. Participants attended two task-based functional resonance imaging sessions before and two months after completing an online self-help CBT-based intervention. We fitted a hybrid reinforcement learning drift diffusion model to the probabilistic reversal learning task data and investigated accumulator-like brain activity as a function of response to computerised CBT. We found that at baseline, compared to nonresponders, responders exhibited weaker left prefrontal and parieto-occipital EA neural signatures, which subsequently increased in proportion to the sustained symptomatic improvement observed following computerised CBT. We thus provide preliminary evidence that attenuated EA neural signatures in the left prefrontal and parieto-occipital cortical areas are associated with response to computerised CBT in depression. Crucially, the observed increase of accumulator-like brain activity following computerised CBT warrants further replication in future experimental work probing neurocomputational mechanisms of change in CBT.
Personalised music-based interventions offer a powerful means of supporting motor rehabilitation by dynamically tailoring auditory stimuli to provide external timekeeping cues, modulate affective states, and stabilise gait patterns. Generalisable Brain-Computer Interfaces (BCIs) thus hold promise for adapting these interventions across individuals. However, inter-subject variability in EEG signals, further compounded by movement-induced artefacts and motor planning differences, hinders the generalisability of BCIs and results in lengthy calibration processes. We propose Individual Tangent Space Alignment (ITSA), a novel pre-alignment strategy incorporating subject-specific recentering, distribution matching, and supervised rotational alignment to enhance cross-subject generalisation. Our hybrid architecture fuses Regularised Common Spatial Patterns (RCSP) with Riemannian geometry in parallel and sequential configurations, improving class separability while maintaining the geometric structure of covariance matrices for robust statistical computation. Using leave-one-subject-out cross-validation, `ITSA' demonstrates significant performance improvements across subjects and conditions. The parallel fusion approach shows the greatest enhancement over its sequential counterpart, with robust performance maintained across varying data conditions and electrode configurations. The code will be made publicly available at the time of publication.
Learning to reinforce rewarding actions and avoid repeated mistakes is crucial for survival in dynamic environments. Yet, it remains unclear how distinct neural signals coordinate to implement reward-based decision-making and behavioural adjustment. We obtained simultaneous electroencephalography (EEG) and pupillometry during a probabilistic reversal learning task. Leveraging single-trial EEG, we first replicate the presence of two feedback-locked neural representations; an early signal previously linked to alertness and switching behaviours following negative feedback and a late signal associated with value updating and reward learning. Using single-trial pupillometry, we then show that differences in feedback-evoked pupil responses between positive and negative feedback are driven primarily by negative feedback encoding. Jointly examining these EEG and pupillometry signatures, we show that following negative feedback, increased trial-by-trial coupling between the pupil response and the early, but not the late, EEG signal is linked to increased uncertainty and exploration tendency as well as reduced accuracy and evidence accumulation on the next trial. Consistent with previous research implicating the locus-coeruleus-noradrenaline system in uncertainty signalling and network resets, we propose that when internal estimates of contextual uncertainty are high following negative feedback, an early signal, likely regulated by locus coeruleus activity, implements a network reset in reward learning structures of a later learning signal. This interruption may simultaneously increase the neural gain related to the processing of novel information and decrease the influence of existing representations in reward learning structures, in turn improving performance by creating new, more accurate internal representations of the external world. Significance Statement The current study jointly examines EEG and pupillometry signatures associated with reversal-learning during reward-based learning. It suggests that when internal estimates of contextual uncertainty are high following negative feedback, an early neural signal, likely regulated by locus coeruleus activity, implements a network reset in reward learning structures of a later learning signal. This interruption may simultaneously increase the neural gain related to the processing of novel information and decrease the influence of existing representations in reward learning structures, in turn improving performance by creating new, more accurate internal representations of the external world. ### Competing Interest Statement The authors have declared no competing interest. European Research Council, https://ror.org/0472cxd90, DyNeRfusion; 865003 Economic and Social Research Council, https://ror.org/03n0ht308, ES/L012995/1 UKRI FLF, MR/Y034368/1 BBSRC, BB/Y001494/1 ARIA grant, SCNI-PR01-P15 University of Glasgow MVLS doctoral training programme
Music has a powerful effect in entraining brain networks that influence both affective states and motor control. The use of Rhythmic Auditory Stimulation (RAS) has shown promising results in regularising and stabilising gait control in patients with neurological problems while alleviating associated depressive symptoms. Brain-computer interfaces (BCIs) can play a pivotal role in shaping music stimulus during these interventions. However, this requires robust detection of mental states during gait adaptation. In this work we investigate the use of Regularised Common Spatial Patterns (RCSP) and Riemannian geometry to detect gait states based on Electroencephalogram (EEG) signals. RCSP are particularly effective on small and noisy datasets while reducing overfitting. Riemannian geometry has proven powerful in analysing the covariance structure of EEG signals that reflect functional brain connectivity. Using a publicly available dataset, we extensively evaluate our methods using two dataset splits. We demonstrated statistically significant results in the dataset split 'individual subjects' with the combination of Regularised Common Spatial Patterns (RCSP) and Riemannian geometry.
Older adults (OAs) are typically slower and/or less accurate in forming perceptual choices relative to younger adults. Despite perceptual deficits, OAs gain from integrating information across senses, yielding multisensory benefits. However, the cognitive processes underlying these seemingly discrepant ageing effects remain unclear. To address this knowledge gap, 212 participants (18-90 years old) performed an online object categorisation paradigm, whereby age-related differences in Reaction Times (RTs) and choice accuracy between audiovisual (AV), visual (V), and auditory (A) conditions could be assessed. Whereas OAs were slower and less accurate across sensory conditions, they exhibited greater RT decreases between AV and V conditions, showing a larger multisensory benefit towards decisional speed. Hierarchical Drift Diffusion Modelling (HDDM) was fitted to participants' behaviour to probe age-related impacts on the latent multisensory decision formation processes. For OAs, HDDM demonstrated slower evidence accumulation rates across sensory conditions coupled with increased response caution for AV trials of higher difficulty. Notably, for trials of lower difficulty we found multisensory benefits in evidence accumulation that increased with age, but not for trials of higher difficulty, in which increased response caution was instead evident. Together, our findings reconcile age-related impacts on multisensory decision-making, indicating greater multisensory evidence accumulation benefits with age underlying enhanced decisional speed.
Asymmetry in choice patterns across rewarding and punishing contexts has long been observed in behavioural economics. Within existing theories of reinforcement learning, the mechanistic account of these behavioural differences is still debated. We propose that motivational salience-the degree of bottom-up attention attracted by a stimulus with relation to motivational goals-offers a potential mechanism to modulate stimulus value updating and decision policy. In a probabilistic reversal learning task, we identified post-feedback signals from EEG and pupillometry that captured differential activity with respect to rewarding and punishing contexts. We show that the degree of between-context distinction in these signals predicts interindividual asymmetries in decision accuracy. Finally, we contextualise these effects in relation to the neural pathways that are currently centred in theories of reward and punishment learning, demonstrating how the motivational salience network could plausibly fit into a range of existing frameworks.
Metacognitive evaluations of confidence provide an estimate of decision accuracy that could guide learning in the absence of explicit feedback. We examine how humans might learn from this implicit feedback in direct comparison with that of explicit feedback, using simultaneous EEG-fMRI. Participants performed a motion direction discrimination task where stimulus difficulty was increased to maintain performance, with intermixed explicit- and no-feedback trials. We isolate single-trial estimates of post-decision confidence using EEG decoding, and find these neural signatures re-emerge at the time of feedback together with separable signatures of explicit feedback. We identified these signatures of implicit versus explicit feedback along a dorsal-ventral gradient in the striatum, a finding uniquely enabled by an EEG-fMRI fusion. These two signals appear to integrate into an aggregate representation in the external globus pallidus, which could broadcast updates to improve cortical decision processing via the thalamus and insular cortex, irrespective of the source of feedback. Confidence could act as an implicit learning signal when explicit feedback is unavailable. The authors show confidence can also provide a distinct value signal in the presence of explicit feedback, both of which are integrated to drive perceptual learning via basal ganglia circuits.
The prior probability of an upcoming stimulus has been shown to influence the formation of perceptual decisions. Computationally, these effects have typically been attributed to changes in the starting point (i.e., baseline) of evidence accumulation in sequential sampling models. More recently, it has also been proposed that prior probability might additionally lead to changes in the rate of evidence accumulation. Here, we introduce a neurally-informed behavioural modelling approach to understand whether prior probability influences the starting point, the rate of evidence accumulation or both. To this end, we employ a well-established visual object categorisation task for which two neural components underpinning participants' choices have been characterised using single-trial analysis of the electroencephalogram. These components are reliable measures of trial-by-trial variability in the quality of the relevant decision evidence, which we use to constrain the estimation of a hierarchical drift diffusion model of perceptual choice. We find that, unlike previous computational accounts, constraining the model with the endogenous variability in the relevant decision evidence results in prior probability effects being explained primarily by changes in the rate of evidence accumulation rather than changes in the starting point or a combination of both. Ultimately, our neurally-informed modelling approach helps disambiguate the mechanistic effect of prior probability on perceptual decision formation, suggesting that prior probability biases primarily the interpretation of sensory evidence towards the most likely stimulus.
BackgroundAltered neural haemodynamic activity during decision making and learning has been linked to the effects of inflammation on mood and motivated behaviours. So far, it has been reported that blunted mesolimbic dopamine reward signals are associated with inflammation-induced anhedonia and apathy. Nonetheless, it is still unclear whether inflammation impacts neural activity underpinning decision dynamics. The process of decision making involves integration of noisy evidence from the environment until a critical threshold of evidence is reached. There is growing empirical evidence that such process, which is usually referred to as bounded accumulation of decision evidence, is affected in the context of mental illness.MethodsIn a randomised, placebo-controlled, crossover study, 19 healthy male participants were allocated to placebo and typhoid vaccination. Three to four hours post-injection, participants performed a probabilistic reversal-learning task during functional magnetic resonance imaging. To capture the hidden neurocognitive operations underpinning decision-making, we devised a hybrid sequential sampling and reinforcement learning computational model. We conducted whole brain analyses informed by the modelling results to investigate the effects of inflammation on the efficiency of decision dynamics and reward learning.ResultsWe found that during the decision phase of the task, typhoid vaccination attenuated neural signatures of bounded evidence accumulation in the dorsomedial prefrontal cortex, only for decisions requiring short integration time. Consistent with prior work, we showed that, in the outcome phase, mild acute inflammation blunted the reward prediction error in the bilateral ventral striatum and amygdala.ConclusionsOur study extends current insights into the effects of inflammation on the neural mechanisms of decision making and shows that exogenous inflammation alters neural activity indexing efficiency of evidence integration, as a function of choice discriminability. Moreover, we replicate previous findings that inflammation blunts striatal reward prediction error signals.
When considering whether to purchase consumer products, people consider both the items' attractiveness and their brand labels. Brands may affect the decision process through various mechanisms. For example, brand labels may provide direct support for their paired products, or they may indirectly affect choice outcomes by changing the way that people evaluate and compare their options. To examine these possibilities, we combined computational modeling with an eye-tracking experiment in which subjects made clothing choices with brand labels either present or absent. Subjects' choices were consistent with both the attractiveness of the clothing items and, to a smaller extent, the appeal of the brands. In line with the direct support mechanism, subjects who spent more time looking at the brands were more likely to choose the options with the preferred brands. When a clothing item was more attractive, subjects were more likely to look longer at the associated brand label, but not vice versa. In line with indirect mechanisms, in the presence of brand labels subjects exerted more caution and showed marginally less attentional bias in their choices. This research sheds light on the interplay between gaze and choice in decisions involving brand information, indicating that brands have both direct and indirect influences on choice.
Signatures of confidence emerge during decision-making, implying confidence may be of functional importance to decision processes themselves. We formulate an extension of sequential sampling models of decision-making in which confidence is used online to actively moderate the quality and quantity of evidence accumulated for decisions. The benefit of this model is that it can respond to dynamic changes in sensory evidence quality. We highlight this feature by designing a dynamic sensory environment where evidence quality can be smoothly adapted within the timeframe of a single decision. Our model with confidence control offers a superior description of human behaviour in this environment, compared to sequential sampling models without confidence control. Using multivariate decoding of electroencephalography (EEG), we uncover EEG correlates of the model's latent processes, and show stronger EEG-derived confidence control is associated with faster, more accurate decisions. These results support a neurobiologically plausible framework featuring confidence as an active control mechanism for improving behavioural efficiency. Feelings of confidence may have a functional role in decision-making. Here, the authors develop a neurobiologically plausible framework featuring confidence as an active control mechanism for improving behavioural efficiency in dynamic environments.
Cognitive behavioral therapy (CBT) is an effective intervention for depression. At present there is no clinically viable predictor of CBT response. Here we combined computational modelling and multivariate classification of neuroimaging data to develop mechanistically interpretable neurocomputational predictors of CBT response in depression.
Motivational (i.e., Pavlovian) values interfere with instrumental responding and can lead to suboptimal decision-making. In humans, task-based neuroimaging studies have only recently started illuminating the functional neuroanatomy of Pavlovian biasing of instrumental control. To provide a mechanistic understanding of the neural dynamics underlying the Pavlovian and instrumental valuation systems, analysis of neuroimaging data has been informed by computational modeling of conditioned behavior. Nonetheless, because of collinearities in Pavlovian and instrumental predictions, previous research failed to tease out hemodynamic activity that is parametrically and dynamically modulated by coexistent Pavlovian and instrumental value expectations. Moreover, neural correlates of Pavlovian to instrumental transfer effects have so far only been identified in extinction (i.e., in the absence of learning). In this study, we devised a modified version of the orthogonalized go/no-go paradigm, which introduced Pavlovian-only catch trials to better disambiguate trial-by-trial Pavlovian and instrumental predictions in both sexes. We found that hemodynamic activity in the ventromedial pFC covaried uniquely with the model-derived Pavlovian value expectations. Notably, modulation of neural activity encoding for instrumental predictions in the supplementary motor cortex was linked to successful action selection in conflict conditions. Furthermore, hemodynamic activity in regions pertaining to the limbic system and medial pFC was correlated with synergistic Pavlovian and instrumental predictions and improved conditioned behavior during congruent trials. Altogether, our results provide new insights into the functional neuroanatomy of decision-making and corroborate the validity of our variant of the orthogonalized go/no-go task as a behavioral assay of the Pavlovian and instrumental valuation systems.
Learning to seek rewards and avoid punishments, based on positive and negative choice outcomes, is essential for human survival. Yet, the neural underpinnings of outcome valence in the human brainstem and the extent to which they differ in reward and punishment learning contexts remain largely elusive. Here, using simultaneously acquired electroencephalography and functional magnetic resonance imaging data, we show that during reward learning the substantia nigra (SN)/ventral tegmental area (VTA) and locus coeruleus are initially activated following negative outcomes, while the VTA subsequently re-engages exhibiting greater responses for positive than negative outcomes, consistent with an early arousal/avoidance response and a later value-updating process, respectively. During punishment learning, we show that distinct raphe nucleus and SN subregions are activated only by negative outcomes with a sustained post-outcome activity across time, supporting the involvement of these brainstem subregions in avoidance behavior. Finally, we demonstrate that the coupling of these brainstem structures with other subcortical and cortical areas helps to shape participants' serial choice behavior in each context.
Music therapy has emerged recently as a successful intervention that improves patient outcomes in a large range of neurological and mood disorders without adverse effects. Brain networks are entrained to music in ways that can be explained both via top-down and bottom-up processes. In particular, the direct interaction of auditory with the motor and the reward system via a predictive framework explains the efficacy of music-based interventions in motor rehabilitation. In this article, we provide a brief overview of current theories of music perception and processing. Subsequently, we summarize the evidence of music-based interventions primarily in motor, emotional, and cardiovascular regulation. We highlight opportunities to improve the quality of life and reduce the stress beyond the clinic environment and in healthy individuals. This relatively unexplored area requires an understanding of how we can personalize and automate music selection processes to fit individual needs and tasks via feedback loops mediated by measurements of neurophysiological responses.
Sensorimotor decision-making is believed to involve a process of accumulating sensory evidence over time. While current theories posit a single accumulation process prior to planning an overt motor response, here, we propose an active role of motor processes in decision formation via a secondary leaky motor accumulation stage. The motor leak adapts the "memory" with which this secondary accumulator reintegrates the primary accumulated sensory evidence, thus adjusting the temporal smoothing in the motor evidence and, correspondingly, the lag between the primary and motor accumulators. We compare this framework against different single accumulator variants using formal model comparison, fitting choice, and response times in a task where human observers made categorical decisions about a noisy sequence of images, under different speed-accuracy trade-off instructions. We show that, rather than boundary adjustments (controlling the amount of evidence accumulated for decision commitment), adjustment of the leak in the secondary motor accumulator provides the better description of behavior across conditions. Importantly, we derive neural correlates of these 2 integration processes from electroencephalography data recorded during the same task and show that these neural correlates adhere to the neural response profiles predicted by the model. This framework thus provides a neurobiologically plausible description of sensorimotor decision-making that captures emerging evidence of the active role of motor processes in choice behavior.
To date, social and nonsocial decisions have been studied largely in isolation. Consequently, the extent to which social and nonsocial forms of decision uncertainty are integrated using shared neurocomputational resources remains elusive. Here, we address this question using simultaneous electroencephalography (EEG)-functional magnetic resonance imaging (fMRI) in healthy human participants (young adults of both sexes) and a task in which decision evidence in social and nonsocial contexts varies along comparable scales. First, we identify time-resolved build-up of activity in the EEG, akin to a process of evidence accumulation (EA), across both contexts. We then use the endogenous trial-by-trial variability in the slopes of these accumulating signals to construct parametric fMRI predictors. We show that a region of the posterior-medial frontal cortex (pMFC) uniquely explains trial-wise variability in the process of evidence accumulation in both social and nonsocial contexts. We further demonstrate a task-dependent coupling between the pMFC and regions of the human valuation system in dorso-medial and ventro-medial prefrontal cortex across both contexts. Finally, we report domain-specific representations in regions known to encode the early decision evidence for each context. These results are suggestive of a domain-general decision-making architecture, whereupon domain-specific information is likely converted into a “common currency” in medial prefrontal cortex and accumulated for the decision in the pMFC.SIGNIFICANCE STATEMENTLittle work has directly compared social-versus-nonsocial decisions to investigate whether they share common neurocomputational origins. Here, using combined electroencephalography (EEG)-functional magnetic resonance imaging (fMRI) and computational modeling, we offer a detailed spatiotemporal account of the neural underpinnings of social and nonsocial decisions. Specifically, we identify a comparable mechanism of temporal evidence integration driving both decisions and localize this integration process in posterior-medial frontal cortex (pMFC). We further demonstrate task-dependent coupling between the pMFC and regions of the human valuation system across both contexts. Finally, we report domain-specific representations in regions encoding the early, domain-specific, decision evidence. These results suggest a domain-general decision-making architecture, whereupon domain-specific information is converted into a common representation in the valuation system and integrated for the decision in the pMFC.