Scaling data and artificial neural networks has transformed AI, driving breakthroughs in language and vision. Whether similar principles apply to modeling brain activity remains unclear. Here we leveraged a dataset of 3.3 million neurons from the visual cortex of 78 mice across 323 sessions, totaling more than 150 billion neural tokens recorded during natural movies, images and parametric stimuli, and behavior. We train multi-modal, multi-task transformer models (1M–300M parameters) that support three regimes flexibly at test time: neural prediction (predicting neuronal responses from sensory input and behavior), behavioral decoding (predicting behavior from neural activity), neural forecasting (predicting future activity from current neural dynamics), or any combination of the three. We find that performance scales reliably with more data, but gains from increasing model size saturate -- suggesting that current brain models are limited by data rather than compute. This inverts the standard AI scaling story: in language and computer vision, massive datasets make parameter scaling the primary driver of progress, whereas in brain modeling -- even in the mouse visual cortex, a relatively simple and low-resolution system -- models remain data-limited despite vast recordings. These findings highlight the need for richer stimuli, tasks, and larger-scale recordings to build brain foundation models. The observation of systematic scaling raises the possibility of phase transitions in neural modeling, where larger and richer datasets might unlock qualitatively new capabilities, paralleling the emergent properties seen in large language models.
Detection and tracking of animals is an important first step for automated behavioral studies using videos. Animal tracking is currently done mostly using deep learning frameworks based on keypoints, which show remarkable results in lab settings with fixed cameras, backgrounds, and lighting. However, multi-animal tracking in the wild presents several challenges such as high variability in background and lighting conditions, complex motion, and occlusion. We propose PriMAT, an approach for tracking nonhuman primates in the wild. PriMAT learns to detect and track primates and other objects of interest from labeled videos or single images using bounding boxes instead of keypoints. Using bounding boxes significantly facilitates data annotation and robustness. Our one-stage model is conceptually simple but highly flexible, and we add a classification branch that allows us to train individual identification. To evaluate the performance of our approach, we applied it in two case studies with Assamese macaques (Macaca assamensis) and redfronted lemurs (Eulemur rufifrons) in the wild. Additionally, we show transfer to other settings and species, particularly, Barbary macaques (Macaca sylvanus), Guinea baboons (Papio papio), chimpanzees (Pan troglodytes), and gorillas (Gorilla spp.). We show that with only a few hundred frames labeled with bounding boxes, we can achieve robust tracking results. Combining these results with the classification branch for the lemur videos, the lemur identification model shows an accuracy of 84% in predicting identities. Our approach presents a promising solution for accurately tracking and identifying animals in the wild, offering researchers a tool to study animal behavior in their natural habitats. Our code, models, training images, and evaluation video sequences are publicly available at https://github.com/ecker-lab/PriMAT-tracking, facilitating their use for animal behavior analyses and future research in this field.
Sensory systems support generalization by representing features that persist under input variation; however, identifying the neuronal basis of these invariances remains difficult due to high-dimensional and nonlinear neural computations. Here we leverage the inception loop paradigm, iterating between large-scale recordings, predictive models and in silico experiments with in vivo verification, to characterize neuronal invariances in mouse primary visual cortex (V1). We synthesize varied exciting inputs (VEIs), dissimilar images that drive target neurons. These VEIs revealed a new bipartite invariance: one subfield encodes a shift-tolerant high-frequency texture and the other encodes a fixed low-frequency pattern. This division aligns with object boundaries defined by spatial frequency differences in highly activating images, suggesting a contribution to segmentation. Analysis of the MICrONS dataset revealed a hierarchy of excitatory neurons in mouse V1 layers 2/3: postsynaptic neurons exhibited greater invariance than their presynaptic inputs, while neurons with lower invariance formed more connections. Together, these results provide insights and scalable methodology for mapping neuronal invariances.
Non-human primates are our closest living relatives, and analyzing their behavior is central to research in cognition, evolution, and conservation. Computer vision could greatly aid this research, but existing methods often rely on human-centric pretrained models and focus on single datasets, which limits generalization. We address this limitation by shifting from a model-centric to a data-centric approach and introduce PriVi, a large-scale primate-centric video pretraining dataset. PriVi contains 424 hours of curated video, combining 174 hours from behavioral research across 11 settings with 250 hours of diverse web-sourced footage, assembled through a scalable data curation pipeline. We continue pretraining V-JEPA on PriVi to learn primate-specific representations and evaluate it using a lightweight frozen classifier. Across four benchmark datasets — ChimpACT, PanAf500, BaboonLand, and ChimpBehave — our approach consistently outperforms prior work, including fully finetuned baselines, and scales favorably with fewer labels. These results demonstrate that primate-centric pretraining substantially improves data efficiency and generalization, making it a promising approach for low-label applications. Code, models, and the majority of the dataset will be made available.
The retina encodes a broad range of stimuli, adapting its computations to features like brightness, contrast, and motion. However, it is unclear whether it also adapts when switching between natural scenes and white noise (WN). To address this, we analyzed the neural activity of male marmoset retinal ganglion cells (RGCs) in response to WN and naturalistic movies. We trained linear-nonlinear models on both stimuli, evaluated their performance, and compared their receptive fields across stimulus domains. We found that models with spatial filters trained on one stimulus ensemble were less accurate when predicting neural activity on the other compared to models trained directly on the target stimulus. This suggests that spatial processing adapts to stimulus statistics. Different RGC types exhibited distinct changes: The OFF midget cells' receptive fields became enlarged under natural movies (NMs), resulting in a lower cutoff frequency. Parasol cells and large OFF cells did not significantly change their receptive field sizes. All cell types exhibited stronger surrounds under NMs, resembling the whitening filters predicted by efficient coding for stimulus decorrelation, prompting us to test whether these changes were related to the different spectral content of the two stimulus types. Quantifying the effects of the filters' enhanced surrounds on the stimulus power spectrum showed a significant contribution toward whitening only in ON parasol cells, where a whitening effect emerged regardless of the training stimulus. These results suggest that while RGCs adapt to the differences between WN and NM stimuli, efficient coding can only partially account for this adaptation.
Vision is context dependent, with neuronal responses shaped not only by local features but also by surrounding visual input. While classical studies, using grating stimuli, show that iso-oriented surrounds suppress responses more than orthogonal surrounds, the role of contextual modulation under natural stimulus conditions remains less clear. Using recordings from mouse primary visual cortex (V1), we trained convolutional neural network models to predict neuronal responses to natural images and synthesized surround stimuli that selectively suppressed or facilitated responses to optimal center inputs. In vivo experiments confirmed these predictions. Facilitatory surrounds resembled naturalistic continuations of the optimal center stimulus, consistent with natural image statistics, whereas suppressive surrounds deviated from these predictions. Applying the same approach to macaque V1 revealed similar principles across species. Both models accurately predicted responses to classical grating stimuli. We formalize these results in a normative Bayesian model, showing that neuronal activity for preferred center features reflects posterior beliefs about likely center-surround configurations.
Retinal ganglion cells, the output neurons of the vertebrate retina, often display nonlinear summation of visual signals over their receptive fields. This creates sensitivity to spatial contrast, letting the cells respond to spatially structured visual stimuli even when no net change in overall illumination of the receptive field occurs. Yet, computational models of ganglion cell responses are often based on linear receptive fields, and typical nonlinear extensions, which separate receptive fields into nonlinearly combined subunits, are often cumbersome to fit to experimental data. Previous work has suggested to model spatial-contrast sensitivity in responses to flashed images by combining signals from the mean and variance of light intensity inside the receptive field. Here, we extend and adjust this spatial contrast model for application to spatiotemporal stimulation and explore its performance on spiking responses that we recorded from ganglion cells of marmosets under artificial and naturalistic movies. We show how the model can be fitted to experimental data and that it outperforms common models with linear spatial integration to different degrees for different types of ganglion cells. Finally, we use the model framework to infer the cells' spatial scale of nonlinear spatial integration. Our work shows that the spatial contrast model can capture aspects of nonlinear spatial integration in the primate retina with only few free parameters. The model can be used to assess the cells' functional properties under natural stimulation and provides a simple-to-obtain benchmark for comparison with more detailed nonlinear encoding models.
Computer vision methods are increasingly used for the automated analysis of large volumes of video data collected through camera traps, drones, or direct observations of animals in the wild. While recent advances have focused primarily on detecting individual actions, much less work has addressed the detection and annotation of interactions – a crucial aspect for understanding social and individualized animal behavior. Existing open-source annotation tools support either behavioral labeling without localization of individuals, or localization without the capacity to capture interactions. To bridge this gap, we present SiLVi, an open-source labeling software that integrates both functionalities. SiLVi enables researchers to annotate behaviors and interactions directly within video data, generating structured outputs suitable for training and validating computer vision models. By linking behavioral ecology with computer vision, SiLVi facilitates the development of automated approaches for fine-grained behavioral analyses. Although developed primarily in the context of animal behavior, SiLVi could be useful more broadly to annotate human interactions in other videos that require extracting dynamic scene graphs. The software, along with documentation and download instructions, is available at: https://silvi.eckerlab.org.
The complexity of neural circuits makes it challenging to decipher the brain's algorithms of intelligence. Recent breakthroughs in deep learning have produced models that accurately simulate brain activity, enhancing our understanding of the brain's computational objectives and neural coding. However, it is difficult for such models to generalize beyond their training distribution, limiting their utility. The emergence of foundation models1 trained on vast datasets has introduced a new artificial intelligence paradigm with remarkable generalization capabilities. Here we collected large amounts of neural activity from visual cortices of multiple mice and trained a foundation model to accurately predict neuronal responses to arbitrary natural videos. This model generalized to new mice with minimal training and successfully predicted responses across various new stimulus domains, such as coherent motion and noise patterns. Beyond neural response prediction, the model also accurately predicted anatomical cell types, dendritic features and neuronal connectivity within the MICrONS functional connectomics dataset2. Our work is a crucial step towards building foundation models of the brain. As neuroscience accumulates larger, multimodal datasets, foundation models will reveal statistical regularities, enable rapid adaptation to new tasks and accelerate research.
Sensory neurons are traditionally viewed as feature detectors that respond with an increase in firing rate to preferred stimuli while remaining unresponsive to others. Here, we identify a dual-feature encoding strategy in macaque visual cortex, wherein many neurons in areas V1 and V4 are selectively tuned to two distinct visual features-one that enhances and one that suppresses activity-around an elevated baseline firing rate. By combining neuronal recordings with functional digital twin models-deep learning-based predictive models of biological neurons-we were able to systematically identify each neuron's preferred and non-preferred features. These feature pairs served as anchors for a continuous, low-dimensional axis in natural image similarity space, along which neuronal activity varied approximately linearly. Within a single visual area, visual features that strongly or weakly activated individual neurons also had a high probability of modulating the activity of other neurons, suggesting a shared feature selectivity across the population that structures stimulus encoding. We show that this encoding strategy is conserved across species, present in both primary and lateral visual areas of mouse cortex. Dual-feature selectivity is consistent with recent anatomical evidence for feature-specific inhibitory connectivity, complementing the feature-detector principle through circuit mechanisms in which selective excitation and inhibition may together enhance the representational capacity of the neuronal population.
Computer vision for animal behavior offers promising tools to aid research in ecology, cognition, and to support conservation efforts. Video camera traps allow for large-scale data collection, but high labeling costs remain a bottleneck to creating large-scale datasets. We thus need data-efficient learning approaches. In this work, we show that we can utilize self-supervised learning to considerably improve action recognition on primate behavior. On two datasets of great ape behavior (PanAf and ChimpACT), we outperform published state-of-the-art action recognition models by 6.1
Advances in computer vision as well as increasingly widespread video-based behavioral monitoring have great potential for transforming how we study animal cognition and behavior. However, there is still a fairly large gap between the exciting prospects and what can actually be achieved in practice today, especially in videos from the wild. With this perspective paper, we want to contribute towards closing this gap, by guiding behavioral scientists in what can be expected from current methods and steering computer vision researchers towards problems that are relevant to advance research in animal behavior. We start with a survey of the state-of-the-art methods for computer vision problems that are directly relevant to the video-based study of animal behavior, including object detection, multi-individual tracking, (inter)action recognition and individual identification. We then review methods for effort-efficient learning, which is one of the biggest challenges from a practical perspective. Finally, we close with an outlook into the future of the emerging field of computer vision for animal behavior, where we argue that the field should move fast beyond the common frame-by-frame processing and treat video as a first-class citizen.
Wide-field amacrine cells (ACs) play a unique role in retinal processing by integrating visual information across a large spatial area. Their inhibitory influence has been implicated in multiple retinal functions such as differential motion detection and the suppression of retinal activity during eye movements. However, a coherent understanding of their general function is lacking due to difficulties in directly recording from these cells and identifying effective visual stimuli to activate them. In this study, we used convolutional neural networks (CNNs) to investigate widefield inhibition mediated by wide-field ACs in the marmoset retina. We trained CNNs to mimic the function of the retina by predicting retinal ganglion cell (RGC) population responses to naturalistic movie stimuli and optimising the most exciting inputs (MEIs) to visualise RGCs’ receptive field (RF) structures. We then optimized suppressive surrounds beyond classical RGC RF boundaries, intended to capture the inhibitory effect of wide-field ACs on RGC activity. These optimized surrounds reduced MEI-elicited activity by 10% to 30%, demonstrating that CNNs not only mimic retinal responses but can also reveal hidden computational aspects of wide-field inhibition. However, suppression strength and generalization varied across architectures and datasets, indicating potential model-specific effects, high-lighting the importance of cautious interpretation. Overall, our approach illustrates how interpretability methods applied to artificial neural networks can offer new hypotheses regarding biological retinal computation, paving the way for targeted experimental validation. The code is available here .
A child’s first step towards social interaction begins with attention to desired social partners, typically by focusing on the partner’s face. While laboratory studies using static images of people’s faces suggest that children look more at adult than child faces, such paradigms fail to capture the dynamic, reciprocal nature of real-world social interactions. In the current study, using mobile headmounted eyetrackers, we examined children’s attention to adults’ and children’s faces as they navigated a crowded room with a naturally changing number of adults and children. Our results reveal that children exhibit a preference for adult over child faces, with a bias towards female faces, and surprisingly, they attend more to unfamiliar than familiar faces. Together, these results reveal the flexible context-sensitive strategies that children use to navigate their daily interactions, and document how social attention unfolds in dynamic social settings.
Rotary Positional Embeddings (RoPE) have demonstrated exceptional performance as a positional encoding method, consistently outperforming their baselines. While recent work has sought to extend RoPE to higher-dimensional inputs, many such extensions are non-commutative, thereby forfeiting RoPE's shift-equivariance property. Spherical RoPE is one such non-commutative variant, motivated by the idea of rotating embedding vectors on spheres rather than circles. However, spherical rotations are inherently non-commutative, making the choice of rotation sequence ambiguous. In this work, we explore a quaternion-based approach – Quaternion Rotary Embeddings (QuatRo) – in place of Euler angles, leveraging quaternions' ability to represent 3D rotations to parameterize the axes of rotation. We show Mixed RoPE and Spherical RoPE to be special cases of QuatRo. Further, we propose a generalization of QuatRo to Clifford Algebraic Rotary Embeddings (CARE) using geometric algebra. Viewing quaternions as the even subalgebra of Cl(3,0,0), we extend the notion of rotary embeddings from quaternions to Clifford rotors acting on multivectors. This formulation enables two key generalizations: (1) extending rotary embeddings to arbitrary dimensions, and (2) encoding positional information in multivectors of multiple grades, not just vectors. We present preliminary experiments comparing spherical, quaternion, and Clifford-based rotary embeddings.
Hierarchical clustering is an effective, interpretable method for analyzing structure in data. It reveals insights at multiple scales without requiring a predefined number of clusters and captures nested patterns and subtle relationships, which are often missed by flat clustering approaches. However, existing hierarchical clustering methods struggle with high-dimensional data, especially when there are no clear density gaps between modes. In this work, we introduce t-NEB, a probabilistically grounded hierarchical clustering method, which yields state-of-the-art clustering performance on naturalistic high-dimensional data. t-NEB consists of three steps: (1) density estimation via overclustering; (2) finding maximum density paths between clusters; (3) creating a hierarchical structure via bottom-up cluster merging. t-NEB uses a probabilistic parametric density model for both overclustering and cluster merging, which yields both high clustering performance and a meaningful hierarchy, making it a valuable tool for exploratory data analysis. Code is available at https://github.com/ecker-lab/tneb clustering.
Deep neural networks trained to predict neural activity from visual input and behaviour have shown great potential to serve as digital twins of the visual cortex. Per-neuron embeddings derived from these models could potentially be used to map the functional landscape or identify cell types. However, state-of-the-art predictive models of mouse V1 do not generate functional embeddings that exhibit clear clustering patterns which would correspond to cell types. This raises the question whether the lack of clustered structure is due to limitations of current models or a true feature of the functional organization of mouse V1. In this work, we introduce DECEMber -- Deep Embedding Clustering via Expectation Maximization-based refinement -- an explicit inductive bias into predictive models that enhances clustering by adding an auxiliary $t$-distribution-inspired loss function that enforces structured organization among per-neuron embeddings. We jointly optimize both neuronal feature embeddings and clustering parameters, updating cluster centers and scale matrices using the EM-algorithm. We demonstrate that these modifications improve cluster consistency while preserving high predictive performance and surpassing standard clustering methods in terms of stability. Moreover, DECEMber generalizes well across species (mice, primates) and visual areas (retina, V1, V4). The code is available at https://github.com/Nisone2000/DECEMber, https://github.com/ecker-lab/cnn-training.
In modern deep neural networks, the learning dynamics of individual neurons are often obscure, as the networks are trained via global optimization. Conversely, biological systems build on self-organized, local learning, achieving robustness and efficiency with limited global information. Here, we show how self-organization between individual artificial neurons can be achieved by designing abstract bio-inspired local learning goals. These goals are parameterized using a recent extension of information theory, Partial Information Decomposition (PID), which decomposes the information that a set of information sources holds about an outcome into unique, redundant and synergistic contributions. Our framework enables neurons to locally shape the integration of information from various input classes, i.e., feedforward, feedback, and lateral, by selecting which of the three inputs should contribute uniquely, redundantly or synergistically to the output. This selection is expressed as a weighted sum of PID terms, which, for a given problem, can be directly derived from intuitive reasoning or via numerical optimization, offering a window into understanding task-relevant local information processing. Achieving neuron-level interpretability while enabling strong performance using local learning, our work advances a principled information-theoretic foundation for local learning strategies.
Transformers rely on positional encoding to compensate for the inherent permutation invariance of self-attention. Traditional approaches use absolute sinusoidal embeddings or learned positional vectors, while more recent methods emphasize relative encodings to better capture translation equivariances. In this work, we propose RollPE, a novel positional encoding mechanism based on traveling waves, implemented by applying a circular roll operation to the query and key tensors in self-attention. This operation induces a relative shift in phase across positions, allowing the model to compute attention as a function of positional differences rather than absolute indices. We show this simple method significantly outperforms traditional absolute positional embeddings and is comparable to RoPE. We derive a continuous case of RollPE which implicitly imposes a topographic structure on the query and key space. We further derive a mathematical equivalence of RollPE to a particular configuration of RoPE. Viewing RollPE through the lens of traveling waves may allow us to simplify RoPE and relate it to processes of information flow in the brain.
Deciphering the brain’s structure-function relationship is key to understanding the neuronal mechanisms underlying perception and cognition. The cortical column, a vertical organization of neurons with similar functions, is a classic example of primate neocortex structure-function organization. While columns have been identified in primary sensory areas using parametric stimuli, their prevalence across higher-level cortex is debated, particularly regarding complex tuning in natural image space. However, a key hurdle in identifying columns is characterizing the complex, nonlinear tuning of neurons to high-dimensional sensory inputs. Building on prior findings of topological organization for features like color and orientation, we investigate functional clustering in macaque visual area V4 in non-parametric natural image space, using large-scale recordings and deep learning–based analysis. We combined linear probe recordings with deep learning methods to systematically characterize the tuning of >1,200 V4 neurons using in silico synthesis of most exciting images (MEIs), followed by in vivo verification. Single V4 neurons exhibited MEIs containing complex features, including textures and shapes, and even high-level attributes with eye-like appearance. Neurons recorded on the same silicon probe, inserted orthogonal to the cortical surface, often exhibited similarities in their spatial feature selectivity, suggesting a degree of functional organization along the cortical depth. We quantified MEI similarity using human psychophysics and distances in a contrastive learning-derived embedding space. Moreover, the selectivity of the V4 neuronal population showed evidence of clustering into functional groups of shared feature selectivity. These functional groups showed parallels with the feature maps of units in artificial vision systems, suggesting potential shared encoding strategies. These results demonstrate the feasibility and scalability of deep learning–based functional characterization of neuronal selectivity in naturalistic visual contexts, offering a framework for quantitatively mapping cortical organization across multiple levels of the visual hierarchy.
M Bethge合作论文数Computational Vision & Neuroscience Group4