Cerebellar-like networks, in which input activity patterns are separated by projection to a much higher-dimensional space before classification, are a recurring neurobiological motif, present in the cerebellum, dentate gyrus, insect olfactory system, and electrosensory system of the electric fish. Their relatively well-understood design presents a promising test-case for probing principles of biological learning. The circuits' expansive projections have long been modelled as random, enabling effective general purpose pattern separation. However, electron-microscopy studies have discovered interesting hints of structure in both the fly mushroom body and mouse cerebellum. Recent numerical work suggested that this non-random connectivity enables the circuit to prioritise learning of some, presumably natural, tasks over others. Here, rather than numerical results, we present a robust mathematical link between the observed connectivity patterns and the cerebellar circuit's learning ability. In particular, we extend a simplified kernel regression model of the system and use recent machine learning theory results to relate connectivity to learning. We find that the reported structure in the projection weights shapes the network's inductive bias in intuitive ways: functions are easier to learn if they depend on inputs that are oversampled, or on collections of neurons that tend to connect to the same hidden layer neurons. Our approach is analytically tractable and pleasingly simple, and we hope it continues to serve as a model for understanding the functional implications of other processing motifs in cerebellar-like networks.
The efficient coding hypothesis presents a compelling success story for theoretical and systems neuroscience. It marshals a unifying idea, that neural codes can be understood as efficient encodings of natural stimuli, to explain phenomena from across sensory systems, sometimes with exquisite precision. However, similar normative assaults on cognitive representations, such as those in prefrontal or entorhinal cortex, have been less comprehensively successful. We argue that this reflects efficient coding’s focus on encoded variables, overlooking the computations that neural circuits implement. Here, instead, we develop an efficient computing theory that studies optimal implementations of recurrent cognitive computations. We apply this framework to the prefrontal cortex, in particular, to a rich vein of neural recordings: structured working memory tasks, such as recalling a sequence. Despite using near-identical tasks, the literature has reported two distinct coding schemes: one contextual, with neurons active only in a particular sequence, the other compositional, with neurons tuned to a single sequence element. Just as efficient coding links stimuli statistics to optimal representation, our theory relates task structure and statistics to optimal representation. In so doing, we find that compositional and contextual codes can be understood as two extremes of a spectrum of optimal representations, with the correlations amongst sequence elements determining a task’s position on this spectrum. Our theory highlights previously underappreciated discrepancies between measured representations, explaining them via subtle task differences; and allows us to infer the algorithm underlying otherwise ambiguous neural data. In sum, we demonstrate an efficient computing approach that makes normative statements about cognitive representations, and serves as a tool for understanding a swathe of neural data.
Sleep is crucial for consolidating all forms of memory and a core mechanism underlying this process is offline replay. Current models propose that replay originates in the hippocampus and triggers reactivation across cortical and subcortical networks. However, conflicting evidence about the role of the hippocampus in offline consolidation of nondeclarative memories raises the question of whether hippocampal replay drives their consolidation. Here we show that replay occurs in the dorsal striatum during offline consolidation of a procedural memory in mice, independently of the hippocampus, and that its content predicts subsequent performance improvements. Neural sequences linked to salient behavioral events were prioritized for replay, with positive and negative behavioral outcomes having opposing effects on individual replay events. All features of replay persisted despite complete bilateral hippocampal lesions. These findings demonstrate that procedural replay occurs independently of the hippocampus, indicating that replay-driven memory consolidation can operate through parallel, independent mechanisms.
How much of the brain's learned algorithms depend on the fact it is a brain? We argue: a lot, but surprisingly few details matter. We point to simple biological details – e.g. nonnegative firing and energetic/space budgets in connectionist architectures – which, when mixed with the requirements of solving a task, produce models that predict brain responses down to single-neuron tuning. We understand this as details constraining the set of plausible algorithms, and their implementations, such that only `brain-like' algorithms are learned. In particular, each biological detail breaks a symmetry in connectionist models (scale, rotation, permutation) leading to interpretable single-neuron responses that are meaningfully characteristic of particular algorithms. This view helps us not only understand the brain's choice of algorithm but also infer algorithm from measured neural responses. Further, this perspective aligns computational neuroscience with mechanistic interpretability in AI, suggesting a more unified approach to studying the mechanisms of intelligence, both natural and artificial.
Why do neurons encode information the way they do? Normative answers to this question model neural activity as the solution to an optimisation problem; for example, the celebrated efficient coding hypothesis frames neural activity as the optimal encoding of information under efficiency constraints. Successful normative theories have varied dramatically in complexity, from simple linear models (Atick & Redlich, 1990), to complex deep neural networks (Lindsay, 2021). What complex models gain in flexibility, they lose in tractability and often understandability. Here, we split the difference by constructing a set of tractable but flexible normative representational theories. Instead of optimising the neural activities directly, following (Sengupta et al. 2018), we instead optimise the representational similarity, a matrix formed from the dot products of each pair of neural responses. Using this, we show that a large family of interesting optimisation problems are convex. This includes problems corresponding to linear and some non-linear neural networks, and problems from the literature not previously recognised as convex such as modified versions of semi-nonnegative matrix factorisation or nonnegative sparse coding. We put these findings to work in two ways. First, we extend previous results on modularity and mixed selectivity in neural activity; in so doing we provide the first necessary and sufficient identifiability result for a form of semi-nonnegative matrix factorisations. Second, we seek to understand the meaningfulness of single neural tuning curves as compared to neural representations. In particular we derive an identifiability result stating that, for an optimal representational similarity matrix, if neural tunings are `different enough' then they are uniquely linked to the optimal representational similarity, partially justifying the use of single neuron tuning analysis in neuroscience. In sum, we identify an interesting space of convex problems, and use that to derive neural coding results.
For 20 years the beautiful structure in the grid cell code has presented an attractive puzzle: what computation do these representations subserve, and why does it manifest so curiously in neurons. The first question quickly attracted an answer: grid cells subserve path-integration, the ability to keep track of one's position as you move about the world. Subsequent work has only solidified this link: bottom-up mechanistic models that perform path-integration match the measured neural responses, while experimental perturbations that selectively disrupt grid cell activity impair performance on path-integration dependent tasks. A more controversial area of work has been top-down normative modelling: why has the brain chosen to compute like this? Floods of ink have been spilt attempting to build a precise link between the population's objective and the measured implementation. The holy grail is a normative link with broad predictive power which generalises to other neural systems. We review this literature and argue that, despite some controversies, the literature largely agrees that grid cells can be explained as a (1) biologically plausible (2) high fidelity, non-linearly decodable code for position that (3) subserves path-integration. As a rare area of neuroscience with mature theoretical and experimental work, this story holds lessons for normative theories of neural computations, and on the risks and rewards of integrating task-optimised neural networks into such theorising.
Why do biological and artificial neurons sometimes modularise, each encoding a single meaningful variable, and sometimes entangle their representation of many variables? In this work, we develop a theory of when biologically inspired networks---those that are nonnegative and energy efficient---modularise their representation of source variables (sources). We derive necessary and sufficient conditions on a sample of sources that determine whether the neurons in an optimal biologically-inspired linear autoencoder modularise. Our theory applies to any dataset, extending far beyond the case of statistical independence studied in previous work. Rather we show that sources modularise if their support is ``sufficiently spread''. From this theory, we extract and validate predictions in a variety of empirical studies on how data distribution affects modularisation in nonlinear feedforward and recurrent neural networks trained on supervised and unsupervised tasks. Furthermore, we apply these ideas to neuroscience data, showing that range independence can be used to understand the mixing or modularising of spatial and reward information in entorhinal recordings in seemingly conflicting experiments. Further, we use these results to suggest alternate origins of mixed-selectivity, beyond the predominant theory of flexible nonlinear classification. In sum, our theory prescribes precise conditions on when neural activities modularise, providing tools for inducing and elucidating modular representations in brains and machines.
Sleep is critical for consolidating all forms of memory[1][1]-[3][2], from episodic experience to the development of motor skills[4][3]-[6][4]. A core feature of the consolidation process is offline replay of neuronal firing patterns that occur during experience[7][5],[8][6]. This replay is thought to originate in the hippocampus and trigger the reactivation of ensembles of cortical and subcortical neurons1,3,9-18. However, non-declarative memories do not require the hippocampus for learning or for sleep-dependent consolidation[19][7]-[26][8] meaning what drives their consolidation is unknown. Here we show, using an unsupervised method, that replay occurs in the dorsal striatum of mice during offline consolidation of a non-declarative, procedural, memory and that this replay is generated independently of the hippocampus. Replay occurred at both real-world and time-compressed speeds and was also prioritised both at the level of the individual neurons and the type of neural sequence. Complete bilateral lesions of the hippocampus had no effect on any feature of this replay. Our results demonstrate that procedural replay during consolidation of a non-declarative memory is independent of the hippocampus. These results support the view that replay drives active consolidation of all types of memory during sleep but challenges the idea that the hippocampus is the source of this replay. ### Competing Interest Statement The authors have declared no competing interest. [1]: #ref-1 [2]: #ref-3 [3]: #ref-4 [4]: #ref-6 [5]: #ref-7 [6]: #ref-8 [7]: #ref-19 [8]: #ref-26
Remembering events is crucial to intelligent behavior. Flexible memory retrieval requires a cognitive map and is supported by two key brain systems: hippocampal episodic memory (EM) and prefrontal working memory (WM). Although an understanding of EM is emerging, little is understood of WM beyond simple memory retrieval. We develop a mathematical theory relating the algorithms and representations of EM and WM by unveiling a duality between storing memories in synapses versus neural activity. This results in a formalism of prefrontal WM as structured, controllable neural subspaces (activity slots) representing dynamic cognitive maps without synaptic plasticity. Using neural networks, we elucidate differences, similarities, and trade-offs between the hippocampal and prefrontal algorithms. Lastly, we show that prefrontal representations in tasks from list learning to cue-dependent recall are unified as controllable activity slots. Our results unify frontal and temporal representations of memory and offer a new understanding for dynamic prefrontal representations of WM.
ABSTRACT To flexibly adapt to new situations, our brains must understand the regularities in the world, but also in our own patterns of behaviour. A wealth of findings is beginning to reveal the algorithms we use to map the outside world 1–6 . In contrast, the biological algorithms that map the complex structured behaviours we compose to reach our goals remain enigmatic. Here we reveal a neuronal implementation of an algorithm for mapping abstract behavioural structure and transferring it to new scenarios. We trained mice on many tasks which shared a common structure organising a sequence of goals, but differed in the specific goal locations. Animals discovered the underlying task structure, enabling zero-shot inferences on the first trial of new tasks. The activity of most neurons in the medial Frontal cortex tiled progress-to-goal, akin to how place cells map physical space. These “goal-progress cells” generalised, stretching and compressing their tiling to accommodate different goal distances. In contrast, progress along the overall sequence of goals was not encoded explicitly. Instead a subset of goal-progress cells was further tuned such that individual neurons fired with a fixed task-lag from a particular behavioural step. Together these cells implemented an algorithm that instantaneously encoded the entire sequence of future behavioural steps, and whose dynamics automatically retrieved the appropriate action at each step. These dynamics mirrored the abstract task structure both on-task and during offline sleep. Our findings suggest that goal-progress cells in the medial frontal cortex may be elemental building blocks of schemata that can be sculpted to represent complex behavioural structures.
Each olfactory cortical hemisphere receives ipsilateral odor information directly from the olfactory bulb and contralateral information indirectly from the other cortical hemisphere. Since neural projections to the olfactory cortex (OC) are disordered and nontopographic, spatial information cannot be used to align projections from the two sides like in the visual cortex. Therefore, how bilateral information is integrated in individual cortical neurons is unknown. We have found, in mice, that the odor responses of individual neurons to selective stimulation of each of the two nostrils are significantly correlated, such that odor identity decoding optimized with information arriving from one nostril transfers very well to the other side. Nevertheless, these aligned responses are asymmetric enough to allow decoding of stimulus laterality. Computational analysis shows that such matched odor tuning is incompatible with purely random connections but is explained readily by Hebbian plasticity structuring bilateral connectivity. Our data reveal that despite the distributed and fragmented sensory representation in the OC, odor information across the two hemispheres is highly coordinated.
Neurons in the brain are often finely tuned for specific task variables. Moreover, such disentangled representations are highly sought after in machine learning. Here we mathematically prove that simple biological constraints on neurons, namely nonnegativity and energy efficiency in both activity and weights, promote such sought after disentangled representations by enforcing neurons to become selective for single factors of task variation. We demonstrate these constraints lead to disentanglement in a variety of tasks and architectures, including variational autoencoders. We also use this theory to explain why the brain partitions its cells into distinct cell types such as grid and object-vector cells, and also explain when the brain instead entangles representations in response to entangled task factors. Overall, this work provides a mathematical understanding of why single neurons in the brain often represent single human-interpretable factors, and steps towards an understanding task structure shapes the structure of brain representation.
ABSTRACT Remembering events in the past is crucial to intelligent behaviour. Flexible memory retrieval, beyond simple recall, requires a cognitive map, or model of how sensations, actions, and latent environmental or task states are all related to one another. Two key brain systems are implicated in this process: the hippocampal episodic memory (EM) system and the prefrontal working memory (WM) system. While an understanding of the hippocampal system, from computation to algorithm and representation, is emerging, less is understood about how the prefrontal WM system can give rise to flexible computations beyond simple memory retrieval, and even less is understood about how the two systems relate to each other. Here we develop a mathematical theory relating the algorithms and representations of EM and WM by unveiling a duality between storing memories in synapses versus neural activity. In doing so, we develop a formal theory of the algorithms and representations of prefrontal WM in terms of structured, and controllable, neural subspaces (termed activity slots) that together can represent a dynamic cognitive map without any need for synaptic plasticity. By building models using this formalism, we elucidate the differences, similarities, and trade-offs between the hippocampal and prefrontal algorithms. Lastly, we show that several prefrontal representations in tasks ranging from list learning to cue dependent recall are unified as controllable activity slots. Our results unify frontal and temporal representations of memory, and offer a new basis for understanding dynamic prefrontal representations of WM.
To afford flexible behaviour, the brain must build internal representations that mirror the structure of variables in the external world. For example, 2D space obeys rules: the same set of actions combine in the same way everywhere (step north, then south, and you won't have moved, wherever you start). We suggest the brain must represent this consistent meaning of actions across space, as it allows you to find new short-cuts and navigate in unfamiliar settings. We term this representation an `actionable representation'. We formulate actionable representations using group and representation theory, and show that, when combined with biological and functional constraints - non-negative firing, bounded neural activity, and precise coding - multiple modules of hexagonal grid cells are the optimal representation of 2D space. We support this claim with intuition, analytic justification, and simulations. Our analytic results normatively explain a set of surprising grid cell phenomena, and make testable predictions for future experiments. Lastly, we highlight the generality of our approach beyond just understanding 2D space. Our work characterises a new principle for understanding and designing flexible internal representations: they should be actionable, allowing animals and machines to predict the consequences of their actions, rather than just encode.
Training data is always finite, making it unclear how to generalise to unseen situations. But, animals do generalise, wielding Occam's razor to select a parsimonious explanation of their observations. How they do this is called their inductive bias, and it is implicitly built into the operation of animals' neural circuits. This relationship between an observed circuit and its inductive bias is a useful explanatory window for neuroscience, allowing design choices to be understood normatively. However, it is generally very difficult to map circuit structure to inductive bias. Here, we present a neural network tool to bridge this gap. The tool meta-learns the inductive bias by learning functions that a neural circuit finds easy to generalise, since easy-to-generalise functions are exactly those the circuit chooses to explain incomplete data. In systems with analytically known inductive bias, i.e. linear and kernel regression, our tool recovers it. Generally, we show it can flexibly extract inductive biases from supervised learners, including spiking neural networks, and show how it could be applied to real animals. Finally, we use our tool to interpret recent connectomic data illustrating its intended use: understanding the role of circuit features through the resulting inductive bias.
Animals receive noisy and incomplete information, from which we must learn how to react in novel situations. A fundamental problem is that training data is always finite, making it unclear how to generalise to unseen data. But, animals do react appropriately to unseen data, wielding Occam's razor to select a parsimonious explanation of the observations. How they do this is called their inductive bias, and it is implicitly built into the operation of animals' neural circuits. This relationship between an observed circuit and its inductive bias is a useful explanatory window for neuroscience, allowing design choices to be understood normatively. However, it is generally very difficult to map circuit structure to inductive bias. In this work we present a neural network tool to bridge this gap. The tool allows us to meta-learn the inductive bias of neural circuits by learning functions that a neural circuit finds easy to generalise, since easy-to-generalise functions are exactly those the circuit chooses to explain incomplete data. We show that in systems where the inductive bias is known analytically, i.e. linear and kernel regression, our tool recovers it. Then, we show it is able to flexibly extract inductive biases from differentiable circuits, including spiking neural networks. This illustrates the intended use case of our tool: understanding the role of otherwise opaque pieces of neural functionality, such as non-linearities, learning rules, or connectomic data, through the inductive bias they induce.
In disentangled representation learning, a model is asked to tease apart a dataset's underlying sources of variation and represent them independently of one another. Since the model is provided with no ground truth information about these sources, inductive biases take a paramount role in enabling disentanglement. In this work, we construct an inductive bias towards encoding to and decoding from an organized latent space. Concretely, we do this by (i) quantizing the latent space into discrete code vectors with a separate learnable scalar codebook per dimension and (ii) applying strong model regularization via an unusually high weight decay. Intuitively, the latent space design forces the encoder to combinatorially construct codes from a small number of distinct scalar values, which in turn enables the decoder to assign a consistent meaning to each value. Regularization then serves to drive the model towards this parsimonious strategy. We demonstrate the broad applicability of this approach by adding it to both basic data-reconstructing (vanilla autoencoder) and latent-reconstructing (InfoGAN) generative models. For reliable evaluation, we also propose InfoMEC, a new set of metrics for disentanglement that is cohesively grounded in information theory and fixes well-established shortcomings in previous metrics. Together with regularization, latent quantization dramatically improves the modularity and explicitness of learned representations on a representative suite of benchmark datasets. In particular, our quantized-latent autoencoder (QLAE) consistently outperforms strong methods from prior work in these key disentanglement properties without compromising data reconstruction.
In this paper we explore a neural control architecture that is both biologically plausible, and capable of fully autonomous learning. It consists of feedback controllers that learn to achieve a desired state by selecting the errors that should drive them. This selection happens through a family of differential Hebbian learning rules that, through interaction with the environment, can learn to control systems where the error responds monotonically to the control signal. We next show that in a more general case, neural reinforcement learning can be coupled with a feedback controller to reduce errors that arise non-monotonically from the control signal. The use of feedback control can reduce the complexity of the reinforcement learning problem, because only a desired value must be learned, with the controller handling the details of how it is reached. This makes the function to be learned simpler, potentially allowing learning of more complex actions. We use simple examples to illustrate our approach, and discuss how it could be extended to hierarchical architectures.
Twisted van der Waals heterostructures have recently emerged as a tunable platform for studying correlated electrons. However, these materials require laborious and expensive effort for both theoretical and experimental exploration. Here we numerically simulate twistronic behavior in acoustic metamaterials composed of interconnected air cavities in two stacked steel plates. Our classical analog of twisted bilayer graphene perfectly replicates the band structures of its quantum counterpart, including mode localization at a magic angle of 1.12(circle). By tuning the thickness of the interlayer membrane, we reach a regime of strong interlayer tunneling where the acoustic magic angle appears as high as 6.01(circle), equivalent to applying 130 GPa to twisted bilayer graphene. In this regime, the localized modes are over five times closer together than at 1.12(circle), increasing the strength of any emergent non-linear acoustic couplings.