An important dimension of cognitive control is the adaptive regulation of the balance between exploitation (pursuing known sources of reward) and exploration (seeking new ones) in response to changes in task utility. Recent studies have suggested that the locus coeruleus-norepinephrine system may play an important role in this function and that pupil diameter can be used to index locus coeruleus activity. On the basis of this, we reasoned that pupil diameter may correlate closely with control state and associated changes in behavior. Specifically, we predicted that increases in baseline pupil diameter would be associated with decreases in task utility and disengagement from the task (exploration), whereas reduced baseline diameter (but increases in task-evoked dilations) would be associated with task engagement (exploitation). Findings in three experiments were consistent with these predictions, suggesting that pupillometry may be useful as an index of both control state and, indirectly, locus coeruleus function.
Previous theoretical work has shown that a single-layer neural network can implement the optimal decision process for simple, two-alternative forced-choice (2AFC) tasks. However, it is likely that the mammalian brain comprises multilayer networks, raising the question of whether and how optimal performance can be approximated in such an architecture. Here, we present theoretical work suggesting that the noradrenergic nucleus locus coeruleus (LC) may help optimize 2AFC decision making in the brain. This is based on the observations that neurons of the LC selectively fire following the presentation of salient stimuli in decision tasks and that the corresponding release of norepinephrine can transiently increase the responsivity, or gain, of cortical processing units. We describe computational simulations that investigate the role of such gain changes in optimizing performance of 2AFC decision making. In the tasks we model, no prior cueing or knowledge of stimulus onset time is assumed.Performance is assessed in terms of the rate of correct responses over time (the reward rate). We first present the results of a single-layer model that accumulates (integrates) sensory input and implements the decision process as a threshold crossing. Gain transients, representing the modulatory effect of the LC, are driven by separate threshold crossings in this layer. We optimize over all free parameters to determine the maximum reward rate achievable by this model and compare it to the maximum reward rate when gain is held fixed. We find that the dynamic gain mechanism yields no improvement in reward for this single-layer model.We then examine a two-layer model, in which competing sensory accumulators in the first layer (capable of implementing the task relevant decision) pass activity to response accumulators in a second layer. Again, we compare a version in which threshold crossing in the first (decision) layer elicits an LC response (and a concomitant increase in gain) with a fixed-gain version of the model. Here, we find that gain transients modeling the LC phasic response yield an improvement in reward rate of 12 to 24. Furthermore, we show that the timing characteristics of these gain transients agree with observations concerning LC firing patterns reported in recent experimental studies. This provides converging evidence for the hypothesis that the LC optimizes processes underlying 2AFC decision making in multilayer networks.
We propose a model by which dopamine (DA) and norepinepherine (NE) combine to alternate behavior between relatively exploratory and exploitative modes. The model is developed for a target detection task for which there is extant single neuron recording data available from locus coeruleus (LC) NE neurons. An exploration-exploitation trade-off is elicited by regularly switching which of the two stimuli are rewarded. DA functions within the model to change synaptic weights according to a reinforcement learning algorithm. Exploration is mediated by the state of LC firing, with higher tonic and lower phasic activity producing greater response variability. The opposite state of LC function, with lower baseline firing rate and greater phasic responses, favors exploitative behavior. Changes in LC firing mode result from combined measures of response conflict and reward rate, where response conflict is monitored using models of anterior cingulate cortex (ACC). Increased long-term response conflict and decreased reward rate, which occurs following reward contingency switch, favors the higher tonic state of LC function and NE release. This increases exploration, and facilitates discovery of the new target.
We review simple connectionist and firing rate models for mutually inhibiting pools of neurons that discriminate between pairs of stimuli. Both are two-dimensional nonlinear stochastic ordinary differential equations, and although they differ in how inputs and stimuli enter, we show that they are equivalent under state variable and parameter coordinate changes. A key parameter is gain: the maximum slope of the sigmoidal activation function. We develop piecewise-linear and purely linear models, and one-dimensional reductions to Ornstein–Uhlenbeck processes that can be viewed as linear filters, and show that reaction time and error rate statistics are well approximated by these simpler models. We then pose and solve the optimal gain problem for the Ornstein–Uhlenbeck processes, finding explicit gain schedules that minimize error rates for time-varying stimuli. We relate these to time courses of norepinephrine release in cortical areas, and argue that transient firing rate changes in the brainstem nucleus locus coeruleus may be responsible for approximate gain optimization.
We show that adaptive gain changes, by hypothesis mediated by the locus coeruleus (LC), can help optimize performance on simulated sensory discrimination tasks, even when no knowledge of stimulus timing is assumed. The metric of performance used here is the rate of correct responses (or ‘reward rate’) achieved by the simulated decision network. The primary model that we study has two layers: the first integrates sensory input directly and the second accumulates this filtered input (as well as noise from other brain areas) and translates it into motor responses via threshold crossing events. Gain transients occur after a physiologically-motivated delay following threshold crossings in the first layer − in this sense, gain schedules adapt according to accumulated sensory information. We adopt a linearization and reduction of the two layer model that allows a clear understanding of parameter effects and which is used to obtain a simpler set of optimization problems without sacrificing the generality of their solutions. By comparing optimal model reward rates in the presence of simulated LC-mediated gain changes with the (separately) optimized reward rates in the absence of such gain changes, the extent to which these gain changes contribute to enhanced task performance is determined, all for a ‘standard parameter set’ derived from fits to experiments. The results indicate that a significant improvement in reward (12−24%) is attributable to the LC-mediated gain mechanism. Additionally, the statistical variations in the optimal model gain transients from trial-to-trial agree with trends reported in recent experimental studies involving direct recordings from the LC. This provides converging evidence for the hypothesis that the LC plays a part in optimizing the dynamics of simple decision tasks.
An understanding of attention is arguably one of the most important goals of the cognitive sciences and yet also has proven to be one of the most elusive. Most attention researchers will agree that a major problem has been agreeing on a definition of the term and the scope of the phenomena to which it applies. There are no doubt as many explanations for this state of affairs as there are those who consider themselves “attention researchers.” However, most will probably agree that, in large measure, this is because attention is not a unitary phenomenon—at least not in the sense that it reflects the operation of a single mechanism, or a single function of one or a set of mechanisms. Rather, attention is the emergent property of the cognitive system that allows it to successfully process some sources of information to the exclusion of others, in the service of achieving some goals to the exclusion of others. This begs an important question: If attention is so varied a phenomenon, how can we make progress in understanding it? There are two simple answers to this question: Be precise about the specific (aspects of the) phenomena to be studied, and be precise about the mechanisms thought to explain them. In this chapter, we address a particular type of attentional phenomenon—that associated with cognitive control. Furthermore, we focus on an account that addresses not only the functional characteristics of this form of attention but also how it is implemented in neural machinery. This neurally oriented approach is attractive not only because it is intrinsically interesting to understand how the mechanisms of the brain give rise to the processes of the mind but more specifically because this exercise has proven useful in generating insights into how controlled attention operates at the systems level. By assuming that information is
We propose a model by which dopamine (DA) and norepinepherine (NE) combine to alternate behavior between relatively exploratory and exploitative modes. The model is developed for a target detection task for which there is extant single neuron recording data available from locus coeruleus (LC) NE neurons. An exploration-exploitation trade-off is elicited by regularly switching which of the two stimuli are rewarded. DA functions within the model to change synaptic weights according to a reinforcement learning algorithm. Exploration is mediated by the state of LC firing, with higher tonic and lower phasic activity producing greater response variability. The opposite state of LC function, with lower baseline firing rate and greater phasic responses, favors exploitative behavior. Changes in LC firing mode result from combined measures of response conflict and reward rate, where response conflict is monitored using models of anterior cingulate cortex (ACC). Increased long-term response conflict and decreased reward rate, which occurs following reward contingency switch, favors the higher tonic state of LC function and NE release. This increases exploration, and facilitates discovery of the new target.
We propose a model by which dopamine (DA) and norepinephrine (NE) combine to alternate behavior between relatively exploratory and exploitative modes. The model is developed for a target detection task for which there is extant single neuron recording data available from locus coeruleus (LC) NE neurons. An exploration-exploitation trade-off is elicited by regularly switching which of the two stimuli are rewarded. DA functions within the model to change synaptic weights according to a reinforcement learning algorithm. Exploration is mediated by the state of LC firing, with higher tonic and lower phasic activity producing greater response variability. The opposite state of LC function, with lower baseline firing rate and greater phasic responses, favors exploitative behavior. Changes in LC firing mode result from combined measures of response conflict and reward rate, where response conflict is monitored using models of anterior cingulate cortex (ACC). Increased long-term response conflict and decreased reward rate, which occurs following reward contingency switch, favors the higher tonic state of LC function and NE release. This increases exploration, and facilitates discovery of the new target.