Detecting a moving object against a moving background is fundamental to survival. The retina is known to perform this computation, extracting object motion from global motion before signals reach the brain. Under the prevailing hierarchical view of vision, downstream areas should inherit and elaborate this retinal computation rather than rebuild it. Using multielectrode recordings of matched stimuli in mouse retina and cortex along with computational modeling, we show instead that post-retinal processing recomputes rather than inherits object motion sensitivity. Retinal and cortical object-motion sensitivities have distinct tuning: the retina prefers fine jitter, whereas cortex prefers coarse drift, revealing a division of labor in which the retina detects moving objects, whereas the cortex conveys information about object pattern. Vision synthesizes object motion through parallel specialization, not hierarchical refinement.
Abstract Background Mice make substantial eye movements during head-fixed visual stimulation, and uncorrected gaze shifts corrupt receptive field measurements and confound stimulus-response relationships. Corneal-reflection video oculography in rodents has provided the methodological foundation for calibrated angular gaze tracking since Stahl (2004) but the calibration procedures used by existing methods — physical camera rotation, motorized stages, behavioral tasks, or precisely co-aligned dual cameras — have limited their adoption in many mouse neuroscience laboratories. Most studies instead use uncalibrated pupil tracking, deep learning pose estimation that returns pixel coordinates without angular calibration, or learned shifter networks that lack independent validation. New method We present an open-source corneal-reflection eye tracking system for head-fixed mice with two methodological contributions. First, a geometric model recovers gaze in calibrated angular units from the pixel displacements of the pupil and corneal reflections, using the known 3D positions of multiple fiducial LEDs as the source of angular scale. The model requires no estimate of R p , the per-animal eye-geometry parameter that earlier corneal-reflection methods determine through physical calibration. Second, a self-calibration procedure exploits the redundancy of multiple stationary fiducial LEDs: each LED produces an independent gaze estimate from the same geometric model, and a single residual calibration parameter is determined by minimizing the disagreement between per-LED estimates. This replaces the physical camera-rotation calibrations of earlier video oculography (Sakatani and Isa, 2004, 2007; van Alphen et al., 2013; Kretschmer et al., 2017), the motorized stages of Zoccolan et al. (2010), and the precision dual-camera alignment of Payne and Raymond (2017) with a software operation that requires no moving parts, no behavioral task, and no per-animal procedure. The system provides three interactive GUI stages: (1) pupil and LED detection via Difference-of-Gaussians filtering, (2) 3D geometry definition, and (3) gaze angle computation with blink detection, fiducial correction, and manual curation. Results Validation against a rotary-encoder-controlled artificial eye demonstrated mean absolute errors below 1° across all four fiducial LEDs over the ±20° working range of mouse eye movements, with Pearson correlations exceeding 0.998 between our method’s estimation and encoder ground truth. The self-calibration reduced inter-LED disagreement by a factor of 4–6 in mouse recordings. Gaze-corrected stimulus reconstruction applied to Neuropixels recordings from mouse V1 produced qualitatively sharper receptive field estimates with improved signal-to-noise ratios. Comparison with existing methods Our method is the first multi-LED, single-camera, fully software-calibrated corneal-reflection eye tracker for mice and includes an integrated open-source pipeline for detection, calibration, blink handling, and artifact correction. The multi-LED redundancy doubles as an internal consistency check — if two LEDs disagree on gaze direction, the calibration is wrong — providing a guarantee that learned approaches relying on neural-data-derived correction cannot offer. Conclusions Our method makes calibrated corneal-reflection eye tracking accessible to non-specialist mouse laboratories using consumer-grade hardware (∼ $2,000–2,700 USD), eliminates the per-animal calibration procedures of earlier methods, and is validated by two independent ground truths at both the absolute angular (artificial eye) and functional (V1 receptive fields) levels. Highlights Open-source corneal-reflection eye tracking for head-fixed mice using a single camera and multiple stationary fiducial LEDs. Geometric gaze model derives angular scale from LED positions, eliminating per-animal eye-geometry calibration. Self-calibration via multi-LED redundancy replaces physical camera rotation, motorized stages, and dual-camera precision alignment. Validated to sub-degree accuracy against a rotary-encoder ground truth across the ± 20° range of mouse eye movements. Gaze correction produces sharper V1 receptive field estimates in Neuropixels recordings.
Sensory systems discriminate between stimuli to direct behavioral choices, a process governed by two distinct properties - neural sensitivity to specific stimuli, and neural variability that importantly includes correlations between neurons. Two questions that have received extensive investigation and debate are whether visual systems are optimized for natural scenes, and whether correlated neural variability contributes to this optimization. However, the lack of sufficient computational models has made these questions inaccessible in the context of the normal function of the visual system, which is to discriminate between natural stimuli. Here we take a direct approach to analyze discriminability under natural scenes for a population of salamander retinal ganglion cells using a model of the retinal neural code that captures both sensitivity and variability. Using methods of information geometry and generative machine learning, we analyzed the manifolds of natural stimuli and neural responses, finding that discriminability in the ganglion cell population adapts to enhance information transmission about natural scenes, in particular about localized motion. Contrary to previous proposals, correlated noise reduces information transmission and arises simply as a natural consequence of the shared circuitry that generates changing spatiotemporal visual sensitivity. These results address a long-standing debate as to the role of retinal correlations in the encoding of natural stimuli and reveal how the highly nonlinear receptive fields of the retina adapt dynamically to increase information transmission under natural scenes by performing the important ethological function of local motion discrimination.
Understanding the circuit mechanisms of the visual code for natural scenes is a central goal of sensory neuro-science. We show that a three-layer network model predicts retinal natural scene responses with an accuracy nearing experimental limits. The model's internal structure is interpretable, as interneurons recorded separately and not modeled directly are highly correlated with model interneurons. Models fitted only to natural scenes reproduce a diverse set of phenomena related to motion encoding, adaptation, and predictive coding, establishing their ethological relevance to natural visual computation. A new approach decomposes the computations of model ganglion cells into the contributions of model interneurons, allowing automatic generation of new hypotheses for how interneurons with different spatiotemporal responses are combined to generate retinal computations, including predictive phenomena currently lacking an explanation. Our results demonstrate a unified and general approach to study the circuit mechanisms of ethological retinal computations under natural visual scenes.
The ability for the brain to discriminate among visual stimuli is constrained by their retinal representations. Previous studies of visual discriminability have been limited to either low-dimensional artificial stimuli or pure theoretical considerations without a realistic encoding model. Here we propose a novel framework for understanding stimulus discriminability achieved by retinal representations of naturalistic stimuli with the method of information geometry. To model the joint probability distribution of neural responses conditioned on the stimulus, we created a stochastic encoding model of a population of salamander retinal ganglion cells based on a three-layer convolutional neural network model. This model not only accurately captured the mean response to natural scenes but also a variety of second-order statistics. With the model and the proposed theory, we computed the Fisher information metric over stimuli to study the most discriminable stimulus directions. We found that the most discriminable stimulus varied substantially across stimuli, allowing an examination of the relationship between the most discriminable stimulus and the current stimulus. By examining responses generated by the most discriminable stimuli we further found that the most discriminative response mode is often aligned with the most stochastic mode. This finding carries the important implication that under natural scenes, retinal noise correlations are information-limiting rather than increasing information transmission as has been previously speculated. We additionally observed that sensitivity saturates less in the population than for single cells and that as a function of firing rate, Fisher information varies less than sensitivity. We conclude that under natural scenes, population coding benefits from complementary coding and helps to equalize the information carried by different firing rates, which may facilitate decoding of the stimulus under principles of information maximization.
Cortical function relies on the balanced activation of excitatory and inhibitory neurons. However, little is known about the organization and dynamics of shaft excitatory synapses onto cortical inhibitory interneurons, which cannot be easily identified morphologically. Here, we fluorescently visualize the excitatory postsynaptic marker PSD-95 at endogenous levels as a proxy for excitatory synapses onto layer 2/3 pyramidal neurons and parvalbumin-positive (PV+) inhibitory interneurons in the mouse barrel cortex. Longitudinal in vivo imaging reveals that, while synaptic weights in both neuronal types are log-normally distributed, synapses onto PV+ neurons are less heterogeneous and more stable. Markov-model analyses suggest that the synaptic weight distribution is set intrinsically by ongoing cell type-specific dynamics, and substantial changes are due to accumulated gradual changes. Synaptic weight dynamics are multiplicative, i.e., changes scale with weights, though PV+ synapses also exhibit an additive component. These results reveal that cell type-specific processes govern cortical synaptic strengths and dynamics.
Neuromodulation imposes powerful control over brain function, and cAMP-dependent protein kinase (PKA) is a central downstream mediator of multiple neuromodulators. Although genetically encoded PKA sensors have been developed, single-cell imaging of PKA activity in living mice has not been established. Here, we used two-photon fluorescence lifetime imaging microscopy (2pFLIM) to visualize genetically encoded PKA sensors in response to the neuromodulators norepinephrine and dopamine. We screened available PKA sensors for 2pFLIM and further developed a variant (named tAKAR alpha) with increased sensitivity and a broadened dynamic range. This sensor allowed detection of PKA activation by norepinephrine at physiologically relevant concentrations and kinetics, and by optogenetically released dopamine. In vivo longitudinal 2pFLIM imaging of tAKARa tracked bidirectional PKA activities in individual neurons in awake mice and revealed neuromodulatory PKA events that were associated with wakefulness, pharmacological manipulation, and locomotion. This new sensor combined with 2pFLIM will enable interrogation of neuromodulation-induced PKA signaling in awake animals.
Understanding how the visual system encodes natural scenes is a fundamental goal of sensory neuroscience. We show here that a three-layer network model predicts the retinal response to natural scenes with an accuracy nearing the fundamental limits of predictability. The model’s internal structure is interpretable, in that model units are highly correlated with interneurons recorded separately and not used to fit the model. We further show the ethological relevance to natural visual processing of a diverse set of phenomena of complex motion encoding, adaptation and predictive coding. Our analysis uncovers a fast timescale of visual processing that is inaccessible directly from experimental data, showing unexpectedly that ganglion cells signal in distinct modes by rapidly (< 0.1 s) switching their selectivity for direction of motion, orientation, location and the sign of intensity. A new approach that decomposes ganglion cell responses into the contribution of interneurons reveals how the latent effects of parallel retinal circuits generate the response to any possible stimulus. These results reveal extremely flexible and rapid dynamics of the retinal code for natural visual stimuli, explaining the need for a large set of interneuron pathways to generate the dynamic neural code for natural scenes.
The normal function of the retina is to convey information about natural visual images. It is this visual environment that has driven evolution, and that is clinically relevant. Yet nearly all of our understanding of the neural computations, biological function, and circuit mechanisms of the retina comes in the context of artificially structured stimuli such as flashing spots, moving bars and white noise. It is fundamentally unclear how these artificial stimuli are related to circuit processes engaged under natural stimuli. A key barrier is the lack of methods for analyzing retinal responses to natural images. We addressed both these issues by applying convolutional neural network models (CNNs) to capture retinal responses to natural scenes. We find that CNN models predict natural scene responses with high accuracy, achieving performance close to the fundamental limits of predictability set by intrinsic cellular variability. Furthermore, individual internal units of the model are highly correlated with actual retinal interneuron responses that were recorded separately and never presented to the model during training. Finally, we find that models fit only to natural scenes, but not white noise, reproduce a range of phenomena previously described using distinct artificial stimuli, including frequency doubling, latency encoding, motion anticipation, fast contrast adaptation, synchronized responses to motion reversal and object motion sensitivity. Further examination of the model revealed extremely rapid context dependence of retinal feature sensitivity under natural scenes using an analysis not feasible from direct examination of retinal responses. Overall, these results show that nonlinear retinal processes engaged by artificial stimuli are also engaged in and relevant to natural visual processing, and that CNN models form a powerful and unifying tool to study how sensory circuitry produces computations in a natural context.
Stoichiometric labeling of endogenous synaptic proteins for high-contrast live-cell imaging in brain tissue remains challenging. Here, we describe a conditional mouse genetic strategy termed endogenous labeling via exon duplication (ENABLED), which can be used to fluorescently label endogenous proteins with near ideal properties in all neurons, a sparse subset of neurons, or specific neuronal subtypes. We used this method to label the postsynaptic density protein PSD-95 with mVenus without overexpression side effects. We demonstrated that mVenus-tagged PSD-95 is functionally equivalent to wild-type PSD-95 and that PSD-95 is present in nearly all dendritic spines in CA1 neurons. Within spines, while PSD-95 exhibited low mobility under basal conditions, its levels could be regulated by chronic changes in neuronal activity. Notably, labeled PSD-95 also allowed us to visualize and unambiguously examine otherwise-unidentifiable excitatory shaft synapses in aspiny neurons, such as parvalbumin-positive interneurons and dopaminergic neurons. Our results demonstrate that the ENABLED strategy provides a valuable new approach to study the dynamics of endogenous synaptic proteins in vivo.