More than four decades ago, Gibbon and Balsam (1981) showed that the acquisition of Pavlovian conditioning in pigeons is directly related to the informativeness of the conditioning stimulus (CS) about the unconditioned stimulus (US), where informativeness is defined as the ratio of the US-US interval (C) to the CS-US interval (T). However, the evidence for this relationship in other species has been equivocal. Here, we describe an experiment that measured the acquisition of appetitive Pavlovian conditioning in 14 groups of rats trained with different C/T ratios (ranging from 1.5 to 300) to establish how learning is related to informativeness. We show that the number of trials required for rats to start responding to the CS is determined by the C/T ratio, and the specific scalar relationship between the rate of learning and informativeness is similar to that previously obtained with pigeons. We also found that the response rate after extended conditioning is strongly related to T, with the terminal CS response rate being a scalar function of the CS reinforcement rate (1 /T). Moreover, this same scalar relationship extended to the rats’ response rates during the inter-trial interval, which was directly proportional to the overall rate of reinforcement in the context (1 /C). The findings establish that animals encode rates of reinforcement, and that conditioning is directly related to how much information the CS provides about the US. The consistency of these observations across species, captured by a simple regression function, suggests a universal model of conditioning.
Contemporary theories guiding the search for neural mechanisms of learning and memory assume that associative learning results from the temporal pairing of cues and reinforcers resulting in coincident activation of associated neurons, strengthening their synaptic connection. While enduring, this framework has limitations: Temporal pairing-based models of learning do not fit with many experimental observations and cannot be used to make quantitative predictions about behavior. Here, we present behavioral data that support an alternative, information-theoretic conception: The amount of information that cues provide about the timing of reward delivery predicts behavior. Furthermore, this approach accounts for the rate and depth of both inhibitory and excitatory learning across paradigms and species. We also show that dopamine release in the ventral striatum reflects cue-predicted changes in reinforcement rates consistent with subjects understanding temporal relationships between task events. Our results reshape the conceptual and biological framework for understanding associative learning.
Reinforcement learning inspires much theorizing in neuroscience, cognitive science, machine learning, and AI. A central question concerns the conditions that produce the perception of a contingency between an action and reinforcement-the assignment- of- credit problem. Contemporary models of associative and reinforcement learning do not leverage the temporal metrics (measured intervals). Our information- theoretic approach formalizes contingency by time- scale invariant temporal mutual information. It predicts that learning may proceed rapidly even with extremely long action-reinforcer delays. We show that rats can learn an action after a single reinforcement, even with a 16- min delay between the action and reinforcement (15- fold longer than any delay previously shown to support such learning). By leveraging metric temporal information, our solution obviates the need for windows of associability, exponentially decaying eligibility traces, microstimuli, or distributions over Bayesian belief states. Its three equations have no free parameters; they predict one- shot learning without iterative simulation.
We develop a mathematical approach to formally proving that certain neural computations and representations exist based on patterns observed in an organism's behaviour. To illustrate, we provide a simple set of conditions under which an ant's ability to determine how far it is from its nest would logically imply neural structures isomorphic to the natural numbers ℕ . We generalise these results to arbitrary behaviours and representations and show what mathematical characterisation of neural computation and representation is simplest while being maximally predictive of behaviour. We develop this framework in detail using a path integration example, where an organism's ability to search for its nest in the correct location implies representational structures isomorphic to two-dimensional coordinates under addition. We also study a system for processing a n b n strings common in comparative work. Our approach provides an objective way to determine what theory of a physical system is best, addressing a fundamental challenge in neuroscientific inference. These results motivate considering which neurobiological structures have the requisite formal structure and are otherwise physically plausible given relevant physical considerations such as generalisability, information density, thermodynamic stability and energetic cost.
In Pavlovian conditioning, the strength of a conditioned response is a function of the probability of reinforcement. However, manipulations of probability are often confounded with changes in the rate of reinforcement. Two between-group experiments in mice evaluated the effect of the probability of reinforcement, while controlling the rate of reinforcement, on appetitive conditioning and extinction. Experiment 1 equated the reinforcement rate by manipulating the number of reinforcements received in each reinforced trial in a critical group (one vs. two consecutive rewards). The results of this experiment showed that probability influenced the rate of responses in acquisition, even when controlling the reinforcement rate. Experiment 2 further assessed the role of probability on behavior while controlling the rate of reinforcement during the conditioned stimulus (CS) using a split-trial design, in which the total CS time was held constant but presented in different numbers of discrete trials (e.g., 50% reinforcement with two 12 s CS's vs. 100% reinforcement with a 24 s average CS duration). This experiment confirmed that probability influenced response rates, and both the probability and rate of reinforcement affected the proportion of trials with responses. Together, these results suggest that the probability of reinforcement, while having little effect on the speed at which responses emerged, affects responding even when the rate of reinforcement is held constant. The results challenge formal learning theories to account for the effects of both the probability and rate of reinforcement. (PsycInfo Database Record (c) 2024 APA, all rights reserved).
Honeybees (Apis mellifera carnica) communicate the direction and distance to a food source by means of a waggle dance. We ask whether bees recruited by the dance use it only as a flying instruction, with the technical form of a polar vector, or also translate it into a location vector that enables them to set courses directed toward the food source from arbitrary locations within their familiar territory. The flights of recruits captured on exiting the hive and released at distant sites were tracked by radar. The recruits performed first a straight flight in approximately the compass direction indicated by the dance. However, this "vector" portion of their flights and the ensuing tortuous "search" portion were strongly and differentially affected by the release site. Searches were biased toward the true location of the food and away from the location specified by translating the origin for the danced polar vector to the release site. We conclude that by following the dance recruits get two messages, a polar flying instruction (bearing and range from the hive) and a location vector that enables them to approach the source from anywhere in their familiar territory. The dance communication is much richer than thought so far.
Temporal information-processing is critical for adaptive behavior and goal-directed action. It is thus crucial to understand how the temporal distance between behaviorally relevant events is encoded to guide behavior. However, research on temporal representations has yielded mixed findings as to whether organisms utilize relative versus absolute judgments of time intervals. To address this fundamental question about the timing mechanism, we tested mice in a duration discrimination procedure in which they learned to correctly categorize tones of different durations as short or long. After being trained on a pair of target intervals, the mice were transferred to conditions in which cue durations and corresponding response locations were systematically manipulated so that either the relative or absolute mapping remained constant. The findings indicate that transfer occurred most readily when relative relationships of durations and response locations were preserved. In contrast, when subjects had to re-map these relative relations, even when positive transfer initially occurred based on absolute mappings, their temporal discrimination performance was impaired, and they required extensive training to re-establish temporal control. These results demonstrate that mice can represent experienced durations both as having a certain magnitude (absolute representation) and as being shorter or longer of the two durations (an ordinal relation to other cue durations), with relational control having a more enduring influence in temporal discriminations. (PsycInfo Database Record (c) 2023 APA, all rights reserved).
Interval timing refers to the ability to perceive and remember intervals in the seconds to minutes range. Our contemporary understanding of interval timing is derived from relatively small-scale, isolated studies that investigate a limited range of intervals with a small sample size, usually based on a single task. Consequently, the conclusions drawn from individual studies are not readily generalizable to other tasks, conditions, and task parameters. The current paper presents a live database that presents raw data from interval timing studies (currently composed of 68 datasets from eight different tasks incorporating various interval and temporal order judgments) with an online graphical user interface to easily select, compile, and download the data organized in a standard format. The Timing Database aims to promote and cultivate key and novel analyses of our timing ability by making published and future datasets accessible as open-source resources for the entire research community. In the current paper, we showcase the use of the database by testing various core ideas based on data compiled across studies (i.e., temporal accuracy, scalar property, location of the point of subjective equality, malleability of timing precision). The Timing Database will serve as the repository for interval timing studies through the submission of new datasets.
Bayesian parameter estimation and Shannon's theory of information provide tools for analysing and understanding data from behavioural and neurobiological experiments on interval timing-and from experiments on Pavlovian and operant conditioning, because timing plays a fundamental role in associative learning. In this tutorial, we explain basic concepts behind these tools and show how to apply them to estimating, on a trial-by-trial, reinforcement-by-reinforcement and response-by -response basis, important parameters of timing behaviour and of the neurobiological manifesta-tions of timing in the brain. These tools enable quantification of relevant variables in the trade-off between acting as an ideal observer should act and acting as an ideal agent should act, which is also known as the trade-off between exploration (information gathering) and exploitation (information utilization) in reinforcement learning. They enable comparing the strength of the evidence for a measurable association to the strength of the behavioural evidence that the association has been perceived. A GitHub site and an OSF site give public access to well-documented Matlab and Python code and to raw data to which these tools have been applied.
The engram encoding the interval between the conditional stimulus (CS) and the unconditional stimulus (US) in eyeblink conditioning resides within a small population of cerebellar Purkinje cells. CSs activate this engram to produce a pause in the spontaneous firing rate of the cell, which times the CS-conditional blink. We developed a Bayesian algorithm that finds pause onsets and offsets in the records from individual CS-alone trials. We find that the pause consists of a single unusually long interspike interval. Its onset and offset latencies and their trial-to-trial variability are proportional to the CS-US interval. The coefficient of variation (CoV = σ/μ) are comparable to the CoVs for the conditional eye blink. The average trial-to-trial correlation between the onset latencies and the offset latencies is close to 0, implying that the onsets and offsets are mediated by two stochastically independent readings of the engram. The onset of the pause is step-like; there is no decline in firing rate between the onset of the CS and the onset of the pause. A single presynaptic spike volley suffices to trigger the reading of the engram; and the pause parameters are unaffected by subsequent volleys. The Fano factors for trial-to-trial variations in the distribution of interspike intervals within the intertrial intervals indicate pronounced non-stationarity in the endogenous spontaneous spiking rate, on which the CS-triggered firing pause supervenes. These properties of the spontaneous firing and of the engram read out may prove useful in finding the cell-intrinsic, molecular-level structure that encodes the CS-US interval.
Optimal behavior requires interpreting environmental cues that indicate when to perform actions. Dopamine is important for learning about reward-predicting events, but its role in adapting to inhibitory cues is unclear. Here we show that when mice can earn rewards in the absence but not presence of an auditory cue, dopamine level in the ventral striatum accurately reflects reward availability in real-time over a sustained period (80 s). In addition, unpredictable transitions between different states of reward availability are accompanied by rapid (~1-2 s) dopamine transients that deflect negatively at the onset and positively at the offset of the cue. This Dopamine encoding of reward availability and transitions between reward availability states is not dependent on reward or activity evoked dopamine release, appears before mice learn the task and is sensitive to motivational state. Our findings are consistent across different techniques including electrochemical recordings and fiber photometry with genetically encoded optical sensors for calcium and dopamine.
The question of whether single cells can learn led to much debate in the early 20th century. The view prevailed that they were capable of non-associative learning but not of associative learning, such as Pavlovian conditioning. Experiments indicating the contrary were considered either non-reproducible or subject to more acceptable interpretations. Recent developments suggest that the time is right to reconsider this consensus. We exhume the experiments of Beatrice Gelber on Pavlovian conditioning in the ciliate Paramecium aurelia, and suggest that criticisms of her findings can now be reinterpreted. Gelber was a remarkable scientist whose absence from the historical record testifies to the prevailing orthodoxy that single cells cannot learn. Her work, and more recent studies, suggest that such learning may be evolutionarily more widespread and fundamental to life than previously thought and we discuss the implications for different aspects of biology.
Rescorla's first theoretical and experimental papers on the truly random control (random, independent presentations of CSs and USs) showed that associative learning was driven by contingency, that is, by the information that events at one time provide about events located elsewhere in time.This discovery has revolutionary neurobiological and philosophical implications.The problem was that Rescorla was unable to derive a function that mapped conditional probabilities into contingencies.Rescorla and Wagner (1972) proposed a hugely influential model for explaining Rescorla's results, but their model ignored his earlier insights about time, temporal order, information and contingency in conditioning.Their paper pioneered an empirically indefensible treatment of time that has continued in associative theorizing down to the present day.A key to a more defensible approach to the cue competition problem (aka the temporal assignment of credit problem) in Pavlovian and instrumental conditioning is to measure the information that cues and responses provide about the wait for reinforcement and the information that reinforcement provides about the recency of a response.
Numbers are symbols manipulated in accord with the axioms of arithmetic. They sometimes represent discrete and continuous quantities (e.g., numerosities, durations, rates, distances, directions, and probabilities), but they are often simply names. Brains, including insect brains, represent the rational numbers with a fixed-point data type, consisting of a significand and an exponent, thereby conveying both magnitude and precision.
Honeybees ( Apis mellifera carnica ) communicate the rhumb line to a food source (its direction and distance from the hive) by means of a waggle dance. We ask whether bees recruited by the dance use it only as a flying instruction or also translate it into a location vector in a map-like memory, so that information about spatial relations of environmental cues informs their attempts to find the source. The flights of recruits captured on exiting the hive and released at distant sites were tracked by radar. The recruits performed first a straight flight in the direction of the rhumb line. However, the vector portions of their flights and the ensuing tortuous search portions were strongly and differentially affected by the release site. Searches were biased toward the true location of the food and away from the location specified by translating the rhumb line origin to the release site. We conclude that by following the dance a recruit gets two messages, a polar flying instruction (the rhumb line) and its conversion to Cartesian map coordinates.
Information theory provides a quantitative conceptual framework for understanding the flow of information from the world into and through brains. It focuses our attention on the sets of possible messages a brain's anatomy and physiology enable it to receive. The meanings of the messages arise from the inferences licensed by the brain's processing of them. Different meanings arise at different levels because different representations of the input license different inferences.
We measured rate of acquisition, trials to extinction, cumulative responses in extinction, and the spontaneous recovery of anticipatory hopper poking in a Pavlovian protocol with mouse subjects. We varied by factors of 4 number of sessions, trials per session, intersession interval, and span of training (number of days over which training extended). We find that different variables affect each measure: Rate of acquisition [1/(trials to acquisition)] is faster when there are fewer trials per session. Terminal rate of responding is faster when there are more total training trials. Trials to extinction and amount of responding during extinction are unaffected by these variables. The number of training trials has no effect on recovery in a 4‐trial probe session 21 days after extinction. However, recovery is greater when the span of training is greater, regardless of how many sessions there are within that span. Our results and those of others suggest that the numbers and durations and spacings of longer‐duration “episodes” in a conditioning protocol (sessions and the spans in days of training and extinction) are important variables and that different variables affect different aspects of subjects' behavior. We discuss the theoretical and clinical implications of these and related findings and conclusions—for theories of conditioning and for neuroscience.
Karl Lashley began the search for the engram nearly seventy years ago. In the time since, much has been learned but divisions remain. In the contemporary neurobiology of learning and memory, two profoundly different conceptions contend: the associative/connectionist (A/C) conception and the computational/representational (C/R) conception. Both theories ground themselves in the belief that the mind is emergent from the properties and processes of a material brain. Where these theories differ is in their description of what the neurobiological substrate of memory is and where it resides in the brain. The A/C theory of memory emphasizes the need to distinguish memory cognition from the memory engram and postulates that memory cognition is an emergent property of patterned neural activity routed through engram circuits. In this model, learning re-organizes synapse association strengths to guide future neural activity. Importantly, the version of the A/C theory advocated for here contends that synaptic change is not symbolic and, despite normally being necessary, is not sufficient for memory cognition. Instead, synaptic change provides the capacity and a blueprint for reinstating symbolic patterns of neural activity. Unlike the A/C theory, which posits that memory emerges at the circuit level, the C/R conception suggests that memory manifests at the level of intracellular molecular structures. In C/R theory, these intracellular structures are information-conveying and have properties compatible with the view that brain computation utilizes a read/write memory, functionally similar to that in a computer. New research has energized both sides and highlighted the need for new discussion. Both theories, the key questions each theory has yet to resolve and several potential paths forward are presented here.