There has been a long-standing debate about the mechanisms underlying the perception of stereoscopic depth and the computation of the relative disparities that it relies on. Relative disparities between visual objects could be computed in two ways: (a) using the difference in the object's absolute disparities (Hypothesis 1) or (b) using relative disparities based on the differences in the monocular separations between objects (Hypothesis 2). To differentiate between these hypotheses, we measured stereoscopic discrimination thresholds for lines with different absolute and relative disparities. Participants were asked to judge the depth of two lines presented at the same distance from the fixation plane (absolute disparity) or the depth between two lines presented at different distances (relative disparity). We used a single stimulus method involving a unique memory component for both conditions, and no extraneous references were available. We also measured vergence noise using Nonius lines. Stereo thresholds were substantially worse for absolute disparities than for relative disparities, and the difference could not be explained by vergence noise. We attribute this difference to an absence of conscious readout of absolute disparities, termed the absolute disparity anomaly. We further show that the pattern of correlations between vergence noise and absolute and relative disparity acuities can be explained jointly by the existence of the absolute disparity anomaly and by the assumption that relative disparity information is computed from absolute disparities (Hypothesis 1).
In the causal inference model of multi-sensory cue integration (Körding et al., 2007) the integration of sensory signals from multiple modalities is determined by the discrepancy between sensory signals. Crossmodal integration is the strongest when signal discrepancies are small. When signal discrepancies are large, the perceptual system assumes that the signals are caused by different sources. Consequently, the influence of one modality on another is flexibly adjusted in a manner contingent on the degree of crossmodal signal discrepancy. Little, however, is known about how aging affects this adaptive crossmodal integration. Methods: 10 older adults (age>60) and 10 young adults participated in a motion direction judgment task. In the vision condition, subjects reported the motion direction of random dot motion. In the haptic condition, subjects reported the motion direction of visually occluded right hand motion that was controlled by a robot arm. In the multimodal condition, haptic motion and visual motion were presented synchronously. In 40% of multimodal trials, both hand and dot direction were identical, whereas in the remaining trials their motion differed by 7, 15, 30, or 50 deg. Participants were always asked to report the direction of visual motion. Results: For young adults the weight of hand motion on their visual estimates significantly decreased from 0.45 to 0.16 (F(3,27)=10.28, p=.0001) as the discrepancy between visual motion and hand motion direction increased from 7 deg to 50 deg. This result is consistent with the prediction of the causal inference model. For older adults,, there was no significant change in the weight of hand motion (F(3,27)=0.93, p=.43). The interaction between group and discrepancy was significant (F(3,54)=3.54, p=.021). Conclusion: These results show that older adults do not adaptively adjust their cue combination strategy in a manner that depends on the discrepanc Meeting abstract presented at VSS 2016
A large body of research has established that, under relatively simple task conditions, human observers integrate uncertain sensory information with learned prior knowledge in an approximately Bayes-optimal manner. However, in many natural tasks, observers must perform this sensory-plus-prior integration when the underlying generative model of the environment consists of multiple causes. Here we ask if the Bayes-optimal integration seen with simple tasks also applies to such natural tasks when the generative model is more complex, or whether observers rely instead on a less efficient set of heuristics that approximate ideal performance. Participants localized a "hidden" target whose position on a touch screen was sampled from a location-contingent bimodal generative model with different variances around each mode. Over repeated exposure to this task, participants learned the a priori locations of the target (i.e., the bimodal generative model), and integrated this learned knowledge with uncertain sensory information on a trial-by-trial basis in a manner consistent with the predictions of Bayes-optimal behavior. In particular, participants rapidly learned the locations of the two modes of the generative model, but the relative variances of the modes were learned much more slowly. Taken together, our results suggest that human performance in a more complex localization task, which requires the integration of sensory information with learned knowledge of a bimodal generative model, is consistent with the predictions of Bayes-optimal behavior, but involves a much longer time-course than in simpler tasks.
Stereopsis is the rich impression of three-dimensionality, based on binocular disparity—the differences between the two retinal images of the same world. However, a substantial proportion of the population is stereo-deficient, and relies mostly on monocular cues to judge the relative depth or distance of objects in the environment. Here we trained adults who were stereo blind or stereo-deficient owing to strabismus and/or amblyopia in a natural visuomotor task—a ‘bug squashing’ game—in a virtual reality environment. The subjects' task was to squash a virtual dichoptic bug on a slanted surface, by hitting it with a physical cylinder they held in their hand. The perceived surface slant was determined by monocular texture and stereoscopic cues, with these cues being either consistent or in conflict, allowing us to track the relative weighting of monocular versus stereoscopic cues as training in the task progressed. Following training most participants showed greater reliance on stereoscopic cues, reduced suppression and improved stereoacuity. Importantly, the training-induced changes in relative stereo weights were significant predictors of the improvements in stereoacuity. We conclude that some adults deprived of normal binocular vision and insensitive to the disparity information can, with appropriate experience, recover access to more reliable stereoscopic information. This article is part of the themed issue ‘Vision in our three-dimensional world’.
Amblyopia is a neuro-developmental disorder of the visual cortex that arises from abnormal visual experience early in life. Amblyopia is clinically important because it is a major cause of vision loss in infants and young children. Amblyopia is also of basic interest because it reflects the neural impairment that occurs when normal visual development is disrupted. Amblyopia provides an ideal model for understanding when and how brain plasticity may be harnessed for recovery of function. Over the past two decades there has been a rekindling of interest in developing more effective methods for treating amblyopia, and for extending the treatment beyond the critical period, as exemplified by new clinical trials and new basic research studies. The focus of this review is on stereopsis and its potential for recovery. Impaired stereoscopic depth perception is the most common deficit associated with amblyopia under ordinary (binocular) viewing conditions (Webber & Wood, 2005). Our review of the extant literature suggests that this impairment may have a substantial impact on visuomotor tasks, difficulties in playing sports in children and locomoting safely in older adults. Furthermore, impaired stereopsis may also limit career options for amblyopes. Finally, stereopsis is more impacted in strabismic than in anisometropic amblyopia. Our review of the various approaches to treating amblyopia (patching, perceptual learning, videogames) suggests that there are several promising new approaches to recovering stereopsis in both anisometropic and strabismic amblyopes. However, recovery of stereoacuity may require more active treatment in strabismic than in anisometropic amblyopia. Individuals with strabismic amblyopia have a very low probability of improvement with monocular training; however they fare better with dichoptic training than with monocular training, and even better with direct stereo training.
Despite growing evidence for perceptual interactions between motion and position, no unifying framework exists to account for these two key features of our visual experience. We show that percepts of both object position and motion derive from a common object-tracking system-a system that optimally integrates sensory signals with a realistic model of motion dynamics, effectively inferring their generative causes. The object-tracking model provides an excellent fit to both position and motion judgments in simple stimuli. With no changes in model parameters, the same model also accounts for subjects' novel illusory percepts in more complex moving stimuli. The resulting framework is characterized by a strong bidirectional coupling between position and motion estimates and provides a rational, unifying account of a number of motion and position phenomena that are currently thought to arise from independent mechanisms. This includes motion-induced shifts in perceived position, perceptual slow-speed biases, slowing of motions shown in visual periphery, and the well-known curveball illusion. These results reveal that motion perception cannot be isolated from position signals. Even in the simplest displays with no changes in object position, our perception is driven by the output of an object-tracking system that rationally infers different generative causes of motion signals. Taken together, we show that object tracking plays a fundamental role in perception of visual motion and position.
While we know that humans are extremely sensitive to optic flow information about direction of heading, we do not know how they integrate information across the visual field. We adapted the standard cue perturbation paradigm to investigate how young adult observers integrate optic flow information from different regions of the visual field to judge direction of heading. First, subjects judged direction of heading when viewing a three-dimensional field of random dots simulating linear translation through the world. We independently perturbed the flow in one visual field quadrant to indicate a different direction of heading relative to the other three quadrants. We then used subjects' judgments of direction of heading to estimate the relative influence of flow information in each quadrant on perception. Human subjects behaved similarly to the ideal observer in terms of integrating motion information across the visual field with one exception: Subjects overweighted information in the upper half of the visual field. The upper-field bias was robust under several different stimulus conditions, suggesting that it may represent a physiological adaptation to the uneven distribution of task-relevant motion information in our visual world.
A growing body of scientific evidence suggests that visual working memory and statistical learning are intrinsically linked. Although visual working memory is severely resource limited, in many cases, it makes efficient use of its available resources by adapting to statistical regularities in the visual environment. However, experimental evidence also suggests that there are clear limits and biases in statistical learning. This raises the intriguing possibility that performance limitations observed in visual working memory tasks can to some degree be explained in terms of limits and biases in statistical-learning ability, rather than limits in memory capacity.
Self-generated body movements have reliable visual consequences. This predictive association between vision and action likely underlies modulatory effects of action on visual processing. However, it is unknown whether actions can have generative effects on visual perception. We asked whether, in total darkness, self-generated body movements are sufficient to evoke normally concomitant visual perceptions. Using a deceptive experimental design, we discovered that waving one’s own hand in front of one’s covered eyes can cause visual sensations of motion. Conjecturing that these visual sensations arise from multisensory connectivity, we showed that grapheme-color synesthetes experience substantially stronger kinesthesis-induced visual sensations than nonsynesthetes do. Finally, we found that the perceived vividness of kinesthesis-induced visual sensations predicted participants’ ability to smoothly track self-generated hand movements with their eyes in darkness, which indicates that these sensations function like typical retinally driven visual sensations. Evidently, even in the complete absence of external visual input, the brain predicts visual consequences of actions.
We have previously shown that an optimal object tracking model accounts for the well-known illusory motion-induced position shift (in which the position of a stationary envelope that contains a moving pattern is shifted in the direction of pattern motion). The model works by optimally using noisy sensory signals to estimate both the motion of the object containing a pattern and the motion of the pattern within the object (VSS 2013). The model makes the novel prediction that when the pattern motion within an envelope differs from the motion of the envelope one's percept of object motion should conflict with temporal changes in one's percept of object position. METHODS: To measure both perceived position and perceived motion of an object, we adapted the well-known 'curveball illusion' stimulus. In this illusion, an object moves downward, while the pattern within the object moves horizontally. In peripheral vision (11°), the object appears to move obliquely. We measured subjects' perceived object position and motion direction for different stimulus durations (20-400 ms). In position blocks, subjects reported the final perceived object position. In motion blocks, subjects reported the final object motion direction. RESULTS: Subjects' horizontal object position biases initially increased with stimulus duration, but asymptoted after 200 ms. (suggesting a constant estimate of horizontal position after 200 msec.) In contrast, their perceived object motion trajectory was oblique at all durations, indicating a non-zero percept of the object's horizontal component of motion. While seemingly irrational, this conflict is consistent with the prediction of the optimal tacking model. Meeting abstract presented at VSS 2014
Dieter, Kevin C., Hu, Bo, Knill, David C., Blake, Randolph, & Tadin, Duje. (2013). Kinesthesis Can Make an Invisible Hand Visible. Psychological Science, 25(1), 66-75. (Original DOI: 10.1177/0956797613497968)
There has been a long-standing debate about the mechanisms underlying the human perception of stereoscopic relative depth. Relative depth between visual objects could be recovered in two different ways. The first one is using the difference of their absolute disparities, which are the differences of the monocular distances between each object and fixation point. A second more direct route consists of using relative disparities, i.e. differences in monocular distances between objects. Studies have claimed the existence of an independent relative disparity system from better performances in two-alternative forced choice discriminations between simultaneously presented stimuli compared to two-interval forced choice (2IFC) between successively presented stimuli, designed to isolate absolute disparities. However, memory noise and vergence noise can substantially reduce performance in the 2IFC task. Further, no previous study has controlled for visual references, leaving open the possible use of relative disparities in 2IFC tasks. We measured depth performance from absolute and relative horizontal disparities with a single stimulus method, involving a single memory component for both conditions. No fixation point was presented and the screen-border shape was in binocular rivalry. Participants were asked to judge the distance to the screen of two lines at the same depth (absolute condition) or the distance between two lines at different depths (relative condition). Vergence noise was also measured using nonius lines. If relative disparity is computed from absolute disparities, one can predict performance in the relative condition from performances in the absolute condition and from the vergence noise. Performance in the relative condition was significantly better than predicted performance, and better than absolute disparity performance. Interestingly, dress-makers displayed significantly better stereoscopic and vergence performance compared to a group of control participants. We conclude that either the relative disparity system is independent from the absolute disparity system, or that absolute disparities cannot be accessed directly. Meeting abstract presented at VSS 2014
To estimate the location of a visual target, an ideal observer must combine what is seen (corrupted by sensory noise) with what is known (from prior experience). Bayesian inference provides a principled method for optimally accomplishing this. Here we provide evidence from two experiments that observers combine sensory information with complex prior knowledge in a Bayes-optimal manner, as they estimate the locations of targets on a touch screen. On each trial, the x-y location of a target was drawn from one of two underlying distributions (the "priors") centered on mean positions on the left and right of the display, and with different variances. The observer, however, only saw a cluster of dots normally distributed around that location (the "likelihood"). Across 1200 trials, the variance of the dot cluster was manipulated to provide three levels of reliability for the likelihood. Feedback on observers’ accuracy in each trial was provided post-touch using dots representing the touched location and the true target location on the screen. In Experiment 1, consistent with the Bayesian model, observers not only relied less on the likelihoods (and more on the priors) as the cluster of dots increased in variance, but they also assigned greater weight to the more reliable prior. In Experiment 2, we obtained a direct estimate of observers’ priors by additionally having them localize the target in the absence of any sensory information. We found that within a few hundred trials, observers reliably learned the true means of both prior distributions, but it took them much longer to learn the relative reliabilities of the two prior distributions. In sum, human observers optimized their performance in a novel spatial localization task by learning the relevant environmental statistics (the two different distributions of target locations) and optimally integrating these statistics with sensory information. Meeting abstract presented at VSS 2013
Introduction: When planning movements to intercept a moving target, humans use adaptive statistical models of speed distributions and temporal correlations to estimate target velocity. The current experiment tested whether and how subjects learn category-contingent statistical models of object speed when presented with targets with mixed distributions. Methods: Subjects viewed target objects that moved briefly (200 – 300 msec.) in a virtual display before disappearing behind a variable-length occluder. Their task was to hit an "impact-zone" drawn on the occluder when the object passed behind it. In a first ("hard") experiment, two types of targets were randomly interleaved from trial-to-trial - red-square targets had speeds drawn from a distribution with a small variance while green-circular targets had speeds drawn from a distribution with the same mean but with a large variance (or vice-versa for half of the subjects). In a second ("easy") experiment, the means of the distributions also differed significantly. Results: In both experiments, subjects showed a significantly larger bias to the mean for the low-variance category than the high, had a marginally significant greater bias to the previous stimulus when it was from the same category than when it was from a different category and adapted their timing behavior to the feedback on the previous trial significantly more when the previous stimulus was from the same category. Despite these differences, their behavior on a trial was still significantly affected by both the speed of the target and the hitting error on the previous trial even when it was from a different category. Conclusions: Subjects partially learn separate priors for different categories, but retain temporal biasing affects from stimuli across categories. This suggests a model in which observers use categorization cues probabilistically, rather than as absolute cues to category membership for purposes of imposing prior models on speed estimates. Meeting abstract presented at VSS 2013
Numerous perceptual demonstrations show that motion influences the spatial coding of object position. For example, the perceived position of a static object is shifted in the direction of motion contained within the object. We postulate that this motion induced position shift (MIPS) results from a process of statistical inference in which position and motion estimates are derived by integrating noisy sensory inputs with the prediction of a forward model that reflects natural dynamics. The model predicts a broad range of known MIPS characteristics, including MIPS’ dependency on stimulus speed and position uncertainty and the asymptotic increase in MIPS with increasing stimulus duration. The model also predicts a novel visual illusion. To confirm this prediction, we presented translational motion (low-pass filtered white noise moving at 7.8°/s) within a stationary Gaussian envelope (s = 0.5°). Crucially, the direction of translation changed at a constant and relatively slow rate (0.72 Hz). Subjects perceive this stimulus as moving along a circular path in a way that its perceived motion conspicuously lags behind the direction of the motion within the object. To quantify this illusion, subjects adjusted the radius and phase of a circularly moving comparison disk to match the perceived motion path of the test stimulus. We found that the matching radius of motion (i.e., perceived illusory rotation) gradually increased from 0.6° to 1.0° as the eccentricity increased from 7.9° to 22.6°. Notably, the phase of the perceived object motion lagged behind the motion in the test stimulus by almost a quarter of the cycle, increasing from 73° to 82° as eccentricity increased. These results are consistent with the behavior of a Kalman filter that integrates sensory signals over time to estimate the evolving position and motion of objects. The model provides a unifying account of perceptual interactions between motion and position signals. Meeting abstract presented at VSS 2013
Because of uncertainty and noise, the brain should use accurate internal models of the statistics of objects in scenes to interpret sensory signals. Moreover, the brain should adapt its internal models to the statistics within local stimulus contexts. Consider the problem of hitting a baseball. The impoverished nature of the visual information available makes it imperative that batters use knowledge of the temporal statistics and history of previous pitches to accurately estimate pitch speed. Using a laboratory analog of hitting a baseball, we tested the hypothesis that the brain uses adaptive internal models of the statistics of object speeds to plan hand movements to intercept moving objects. We fit Bayesian observer models to subjects' performance to estimate the statistical environments in which subjects' performance would be ideal and compared the estimated statistics with the true statistics of stimuli in an experiment. A first experiment showed that subjects accurately estimated and used the variance of object speeds in a stimulus set to time hitting behavior but also showed serial biases that are suboptimal for stimuli that were uncorrelated over time. A second experiment showed that the strength of the serial biases depended on the temporal correlations within a stimulus set, even when the biases were estimated from uncorrelated stimulus pairs subsampled from the larger set. Taken together, the results show that subjects adapted their internal models of the variance and covariance of object speeds within a stimulus set to plan interceptive movements but retained a bias to positive correlations.
Derivation of the optimal memory channelIn this section we derive the ideal observer model of visual working memory using results from rate-distortion theory.To begin our analysis, we assume that there is an information source in the world, labeled x, which can be characterized by a Gaussian distribution with mean µ w and variance σ 2 w .We further assume that the process of sensory encoding can be approximated by adding corrupting zero-mean Gaussian noise to the true signal.This results in a distribution over aerent sensory signals which is again Gaussian, with mean µ w and variance σ 2 w + σ 2 s , where σ 2 s is the variance of the corrupting sensory noise.For many visual features, psychophysical performance is more accurately modeled by assuming that sensory noise is additive with respect to the logarithm of the physical stimulus value (this is just a restatement of the Weber-Fechner law).To apply our modeling framework to such cases, one can design an experiment in which the information source is log-normal distributed (or in other words, so that the logarithm of the feature value has a Gaussian distribution).In this case, the sensory signal x s is assumed to represent the log-transform of the physical feature, so that p(x s ) still follows a Gaussian distribution with mean and variance as given above.The task is to specify a memory channel distribution, p(x m | x s ), which relates the output of VWM to its input, and show that this information channel is optimal for a fixed information rate.The memory channel describes the combined eect of two processes, a memory encoding mechanism, and a memory decoder (this can be thought of as analogous to memory encoding and retrieval in VWM).For the time being, we assume a particular (noisy) memory encoding mechanism, and pair it with a decoder that is optimal with respect to the encoder.We subsequently show that our choice for the encoding mechanism does not impact the level of performance of the resulting model, as it is shown to achieve the theoretical rate-distortion bound.That is to say, although we assume a particular parametric model of memory encoding (namely, corrupting Gaussian noise), no alternative model can achieve lower estimation error for a given capacity.If x s is the sensory signal that is input to VSTM, then we define its memory encoding, x e as a Gaussian distribution with mean x s and variance σ 2 e .The output of VWM represents an attempt to reconstruct the incoming sensory signal in face of the degradation incurred as part of the encoding process.(Note that since the af-
When reaching for objects, humans make saccades to fixate the object at or near the time the hand begins to move. In order to address whether the CNS relies on a common representation of target positions to plan both saccades and hand movements, we quantified the contributions of visual short-term memory (VSTM) to hand and eye movements executed during the same coordinated actions. Subjects performed a sequential movement task in which they picked up one of two objects on the right side of a virtual display (the "weapon"), moved it to the left side of the display (to a "reloading station") and then moved it back to the right side to hit the other object (the target). On some trials, the target was perturbed by 1° of visual angle while subjects moved the weapon to the reloading station. Although subjects did not notice the change, the original position of the target, encoded in VSTM, influenced the motor plans for both the hand and the eye back to the target. Memory influenced motor plans for distant targets more than for near targets, indicating that sensorimotor planning is sensitive to the reliability of available information; however, memory had a larger influence on hand movements than on eye movements. This suggests that spatial planning for coordinated saccades and hand movements are dissociated at the level of processing at which online visual information is integrated with information in short-term memory.
Motivation: It has long been reported that perceptual estimation of stimulus is biased toward the immediately preceding stimulus (Holland & Lockhead, 1968), and we have recently shown that the bias is modulated by the temporal correlation of stimuli history (Kwon and Knill, VSS 2011). Here, we asked whether the temporal dependency in perceptual estimation can be modulated by the observers’ prior knowledge on how the stimuli velocities are generated. Method: We used a motion extrapolation task in which a target moved and disappeared behind an occluder and subjects had to hit the target when it was supposed to be in the designated hitting zone. In the first condition, subjects actively generated the velocity of target (active condition) by hitting the virtual target with a pusher. The target launched when the pusher contacted target at the speed of the pusher. The same group of subjects participated in the second condition. In the second condition, the stimuli velocity was passively presented, but the sequence of stimuli velocities was what the same subjects generated in the first condition (active-passive condition). In the third condition, a new group of subjects participated. The sequence of stimuli velocities was the same as the other conditions (passive condition). Results: The bias toward the immediately preceding velocity was close to zero in the active and the active-passive conditions, whereas the bias was significantly different from zero in the passive condition. On the contrary, the bias toward the mean velocity was comparable in all three conditions. Conclusions: Temporal dependency in estimation of velocity disappears when subjects actively generated stimuli. Surprisingly, the same effect was observed when the stimuli were presented passively, if subjects know that the sequence of stimuli generated by themselves. Perceptual estimation of stimulus velocity is modulated by subjects’ prior knowledge on how the stimuli are generated. Meeting abstract presented at VSS 2012
Limits in visual working memory (VWM) strongly constrain human performance across many tasks. However, the nature of these limits is not well understood. In this article we develop an ideal observer analysis of human VWM by deriving the expected behavior of an optimally performing but limited-capacity memory system. This analysis is framed around rate-distortion theory, a branch of information theory that provides optimal bounds on the accuracy of information transmission subject to a fixed information capacity. The result of the ideal observer analysis is a theoretical framework that provides a task-independent and quantitative definition of visual memory capacity and yields novel predictions regarding human performance. These predictions are subsequently evaluated and confirmed in 2 empirical studies. Further, the framework is general enough to allow the specification and testing of alternative models of visual memory (e.g., how capacity is distributed across multiple items). We demonstrate that a simple model developed on the basis of the ideal observer analysis-one that allows variability in the number of stored memory representations but does not assume the presence of a fixed item limit-provides an excellent account of the empirical data and further offers a principled reinterpretation of existing models of VWM.