A long-standing observation about primary visual cortex (V1) is that the stimulus selectivity of neurons can be well explained with a cascade of linear computations followed by a nonlinear rectification stage. This framework remains highly influential in systems neuroscience and has also inspired recent efforts in artificial intelligence. The success of these models include describing the disparity-selectivity of binocular neurons in V1. Some aspects of real neuronal disparity responses are hard to explain with simple linear-nonlinear models, notably the attenuated response of real cells to “anticorrelated” stimuli which violate natural binocular image statistics. General linear-nonlinear models can account for this attenuation, but no one has yet tested whether they quantitatively match the response of real neurons. Here, we exhaustively test this framework using recently developed optimisation techniques. We show that many cells are very poorly characterised by even general linear-nonlinear models. Strikingly, the models can account for neuronal responses to unnatural anticorrelated stimuli as well as to most natural, correlated stimuli. However, the models fail to capture the particularly strong response to binocularly correlated stimuli at the preferred disparity of the cell. Thus, V1 neurons perform an amplification of responses to correlated stimuli which cannot be accounted for by a linear-nonlinear cascade. The implication is that even simple stimulus selectivity in V1 requires more complex computations than previously envisaged.Significance statement A long-standing question in sensory systems neuroscience is whether the computations performed by neurons in primary visual cortex can be described by repeated elements of linear-nonlinear units (a linear filtering/pooling stage followed by a subsequent output nonlinearity, such as a squaring). This question goes back to the Nobel-prize winning work by Hubel & Wiesel who argued that orientation selectivity in V1 can qualitatively be explained in this way. In this paper, we show that V1 neurons have an amplification of their response to stimuli which are contrast matched in the two eyes, and that the recovered models cannot describe this property. We argue that this likely represents more sophisticated computations than can be compactly described by the linear-nonlinear cascade framework.
Stereopsis is the ability to estimate distance based on the different views seen in the two eyes [1-5]. It is an important model perceptual system in neuroscience and a major area of machine vision. Mammalian, avian, and almost all machine stereo algorithms look for similarities between the luminance-defined images in the two eyes, using a series of computations to produce a map showing how depth varies across the scene [3, 4, 6-14]. Stereopsis has also evolved in at least one invertebrate, the praying mantis [15-17]. Mantis stereopsis is presumed to be simpler than vertebrates' [15, 18], but little is currently known about the underlying computations. Here, we show that mantis stereopsis uses a fundamentally different computational algorithm from vertebrate stereopsis-rather than comparing luminance in the two eyes' images directly, mantis stereopsis looks for regions of the images where luminance is changing. Thus, while there is no evidence that mantis stereopsis works at all with static images, it successfully reveals the distance to a moving target even in complex visual scenes with targets that are perfectly camouflaged against the background in terms of texture. Strikingly, these insects outperform human observers at judging stereoscopic distance when the pattern of luminance in the two eyes does not match. Insect stereopsis has thus evolved to be computationally efficient while being robust to poor image resolution and to discrepancies in the pattern of luminance between the two eyes.
Reducing contrast has little effect on psychophysical stereoacuity, except at very low contrast. Differences in contrast between the eyes are more disruptive. The effect of contrast on disparity selectivity in cortical neurons has been investigated only in the cat with grating stimuli. Here we report the effects of stimulus contrast on disparity selectivity in 38 disparity-tuned neurons recorded from V1 of the awake fixating primate. The stimulus was a dynamic 1-dimensional noise pattern ("barcode"), to which disparity was applied. The stimulus (duration 750ms) was presented at either 20% (L) or 100% (H) contrast in each eye, and all four combinations (HH,LL,LH,HL) were used. We used the high contrast condition (HH) as a reference, and plotted responses to correlated disparity in the other three conditions relative to this. The slope of a type II regression was then used to quantify relative response strength in the other three conditions. Reducing contrast in either or both eyes reduced the strength of disparity selectivity (median ratio 0.83 for LL vs HH, 0.59 for LH and HL vs HH, both significantly different from 1, p=0.03 and < 0.001, sign test). Reducing contrast in one eye only does have a slightly greater effect on disparity tuning compared compared to reducing it in both (median ratio 1.27 for LL vs LH and HL; significantly different from 1, p=0.02), which is qualitatively in agreement with psychophysics. However, the 40% reduction in signal strength for LH and HL relative to HH is larger than the psychophysical effects reported for interocular contrasts in the same range. One explanation for this difference could be that those neurons most affected by interocular contrast differences are given less weight in stereoacuity tasks. Meeting abstract presented at VSS 2018
Stereo vision relies on disparity-selective cells in primary visual cortex. The binocular energy model (BEM) has been successful in capturing a range of properties of these neurons. While the BEM's ability to capture mean rates have received substantial attention, little effort has been directed to understanding response variability. We have previously shown that the BEM predicts that the spike rate variance should increase with the square of the mean rate. Importantly, this variance is stimulus driven, reflecting the effect of different dot patterns while disparity is fixed. Recording from V1 neurons in the macaque, we used a two-pass method to separate stochastic variability ("internal variance") from stimulus induced variability ("external variance"). We found that V1 neurons show much less external variance than the BEM. This failure is partly due to the BEM's highly constrained structure. We fit more general linear-nonlinear (LN) models, with the number of subunits as a free parameter, to neuronal data. This general architecture is able to capture both the mean and variability of real cells much better. In particular, by incorporating multiple orthogonal excitatory subunits, the new model is able to achieve lower variability for a given mean than is possible in the BEM. However, problems remain. Notably, while the BEM produced too much variability, the new model produces too little. We show that while the new model captures the "internal" variability due to the spike generation process, it underestimates the "external" variability produced by different random noise patterns with a given correlation and disparity. The new model also still shows the characteristic relationship between the Fano Factor (Variance/Mean) and the mean which we reported previously in the BEM, and which is absent in real cells. Thus, this substantial generalization of the BEM is still not an accurate model of real V1 neurons Meeting abstract presented at VSS 2017
The first step in binocular stereopsis is to match features on the left retina with the correct features on the right retina, discarding 'false' matches. The physiological processing of these signals starts in the primary visual cortex, where the binocular energy model has been a powerful framework for understanding the underlying computation. For this reason, it is often used when thinking about how binocular matching might be performed beyond striate cortex. But this step depends critically on the accuracy of the model, and real V1 neurons show several properties that suggest they may be less sensitive to false matches than the energy model predicts. Several recent studies provide empirical support for an extended version of the energy model, in which the same principles are used, but the responses of single neurons are described as the sum of several subunits, each of which follows the principles of the energy model. These studies have significantly improved our understanding of the role played by striate cortex in the stereo correspondence problem.This article is part of the themed issue 'Vision in our three-dimensional world'.
A recent study provides compelling evidence that binocular vision uses two separate channels; one channel adds the images from the two eyes, and the other subtracts them.
Recent technological advances now allow for the collection of vast data sets detailing the intricate neural connectivity patterns of various organisms. Oh et al. (2014) recently published the most complete description of the mouse mesoscale connectome acquired to date. Here we give an in-depth characterization of this connectome and propose a generative network model which utilizes two elemental organizational principles: proximal attachment ‒ outgoing connections are more likely to attach to nearby nodes than to distant ones, and source growth ‒ nodes with many outgoing connections are likely to form new outgoing connections. We show that this model captures essential principles governing network organization at the mesoscale level in the mouse brain and is consistent with biologically plausible developmental processes.
The binocular energy model (BEM) has proven to be a very successful description of disparity selective neurons in area V1. However, most tests of the model compare mean responses to different stimuli - little attention has been given to the distribution of responses predicted by the BEM. To explore the distribution of responses, we computed the responses of BEM units to dynamic 1D noise stimuli at a range of disparities, with either positive or negative binocular correlation. Each frame was presented for 30 ms. The distribution of the model's responses to these stimuli was highly kurtotic, similar to an exponential distribution. As for exponential distributions, the variance grows with the square of the mean, meaning that the Fano factor (variance/mean) in the model is proportional to the mean. This is in stark contrast to cells in V1, where the Fano factor depends only weakly on the mean. This means that disparity-related changes in the mean response of linear-nonlinear models are largely caused by a small number of frames which elicit very large responses, rather than, for example, reflecting a change in the mean of a Gaussian distribution. We show that the dependence of Fano factor on disparity holds even for linear-nonlinear cascade models of binocular cells that have been fit to real cells using recently developed optimization techniques. Importantly the neuronal data (to which the models were fit) do not show the same dependence of Fano factor on disparity. We show that it is possible to greatly reduce the Fano factor variation in the model by introducing a form of monocular gain control. We propose that incorporating monocular gain control is critical for adequately modeling the trial-to-trial dynamics of disparity-selective cells in V1. Meeting abstract presented at VSS 2016
Human stereopsis can operate in dense "cyclopean" images containing no monocular objects. This is believed to depend on the computation of binocular correlation by neurons in primary visual cortex (V1). The observation that humans perceive depth in half-matched random-dot stereograms, although these stimuli have no net correlation, has led to the proposition that human depth perception in these stimuli depends on a distinct "matching" computation possibly performed in extrastriate cortex. However, recording from disparity-selective neurons in V1 of fixating monkeys, we found that they are in fact able to signal disparity in half-matched stimuli. We present a simple model that explains these results. This reinstates the view that disparity-selective neurons in V1 provide the initial substrate for perception in dense cyclopean stimuli, and strongly suggests that separate correlation and matching computations are not necessary to explain existing data on mixed correlation stereograms. SIGNIFICANCE STATEMENT The initial step in stereoscopic 3D vision is generally thought to be a correlation-based computation that takes place in striate cortex. Recent research has argued that there must be an additional matching computation involved in extracting stereoscopic depth in random-dot stereograms. This is based on the observation that humans can perceive depth in stimuli with a mean binocular correlation of zero (where a correlation-based mechanism should not signal depth). We show that correlation-based cells in striate cortex do in fact signal depth here because they convert fluctuations in the correlation level into a mean change in the firing rate. Our results reinstate the view that these cells provide a sufficient substrate for the perception of stereoscopic depth.
In order to extract retinal disparity from a visual scene, the brain must match corresponding points in the left and right retinae. This computationally demanding task is known as the stereo correspondence problem. The initial stage of the solution to the correspondence problem is generally thought to consist of a correlation-based computation. However, recent work by Doi et al suggests that human observers can see depth in a class of stimuli where the mean binocular correlation is 0 (half-matched random dot stereograms). Half-matched random dot stereograms are made up of an equal number of correlated and anticorrelated dots, and the binocular energy model-a well-known model of V1 binocular complex cell-sfails to signal disparity here. This has led to the proposition that a second, match-based computation must be extracting disparity in these stimuli. Here we show that a straightforward modification to the binocular energy model-adding a point output nonlinearity-is by itself sufficient to produce cells that are disparity-tuned to half-matched random dot stereograms. We then show that a simple decision model using this single mechanism can reproduce psychometric functions generated by human observers, including reduced performance to large disparities and rapidly updating dot patterns. The model makes predictions about how performance should change with dot size in half-matched stereograms and temporal alternation in correlation, which we test in human observers. We conclude that a single correlation-based computation, based directly on already-known properties of V1 neurons, can account for the literature on mixed correlation random dot stereograms.
Depth perception from binocular disparity depends upon correctly matching image features seen by the left and right eyes. A local cross-correlation between left and right images, similar to the operation of the binocular energy model, is a good candidate mechanism. Recently, Doi et al. (2011, 2013, 2014) showed that human observers can detect depth in “half-matched” stereograms containing equal numbers of correlated and anti-correlated dots. These stimuli have a binocular correlation of 0 for all disparities, leading Doi et al. to argue that a correlation computation cannot explain human performance. However, these stimuli do contain local fluctuations in correlation. We explore whether it is possible account for these responses with the binocular energy model, and have begun testing the predictions in disparity-selective neurons from V1. In half-matched stereograms the standard binocular energy model responds equally to all disparities. Simply adding an expansive nonlinearity on the output of a model complex cell renders it disparity selective. At the preferred disparity, local fluctuations in correlation produce a greater variance in the output of an energy model, compared to a non-preferred disparity. When passed through the nonlinearity, this greater variance results in a greater mean. The strength of disparity selectivity decreases with dot density (low density produces greater fluctuations). We examined this in disparity selective neurons in monkey V1. In neurons showing attenuated responses to anticorrelated stimuli, we find robust disparity selectivity at low dot densities. At high dot densities disparity selectivity is much weaker, but remains significant. A single computation – the binocular energy model followed by an output nonlinearity, can explain depth perception in both correlated and “half-matched” random dot stereograms. A key property of this simple model, that disparity selectivity depends on dot density (only in half-matched stereograms), is true in appropriately selected disparity selective V1 neurons. Meeting abstract presented at VSS 2015