
When we look at a scene, how do we consciously see surfaces infused with lightness and color at the correct depths? Random-Dot Stereograms (RDS) probe how binocular disparity between the two eyes can generate such conscious surface percepts. Dense RDS do so despite the fact that they include multiple false binocular matches. Sparse stereograms do so even across large contrast-free regions with no binocular matches. Stereograms that define occluding and occluded surfaces lead to surface percepts wherein partially occluded textured surfaces are completed behind occluding textured surfaces at a spatial scale much larger than that of the texture elements themselves. Earlier models suggest how the brain detects binocular disparity, but not how RDS generate conscious percepts of 3D surfaces. This article proposes a neural network model that predicts how the layered circuits of visual cortex generate these 3D surface percepts using interactions between visual boundary and surface representations that obey complementary computational rules. The model clarifies how interactions between layers 4, 3B and 2/3A in V1 and V2 contribute to stereopsis, and proposes how 3D perceptual grouping laws in V2 interact with 3D surface filling-in operations in V1, V2 and V4 to generate 3D surface percepts in which figures are separated from their backgrounds. The model explanations of 3D surface percepts raised by various RDS are demonstrated by computer simulations. The model hereby unifies the explanation of data about stereopsis and data about 3D figure-ground separation and completion of partially occluded object surfaces. It shows how these model mechanisms convert the complementary rules for boundary and surface formation into consistent visual percepts of 3D surfaces.
Recovering shape in three dimensions has obvious importance for visual perception. Hence one principal goal for stereopsis should be to recover good estimates of 3D shape. But this is impossible if disparity processing is hardwired, because at different fixation distances a fixed angular disparity will correspond to quite different distance increments. An experiment confirms previous evidence that the disparity computation is not hardwired. Specifically, as fixation distance changes, the perceived relation between depth and disparity changes. The changes are consistent with a remapping that partially preserves the constancy of 3D shape over a wide range of fixation distances.
We investigate the relation between the physical world and its mental representation in the 'cognitive map', and test if this representation is image-like and complies with the laws of Euclidean geometry. We have developed a new experimental technique using 'impossible' virtual environments (VE) to directly influence the representational development. Subjects explore a number of VEs -- some 'normal', others with severe violations of Euclidean metrics or planar topology. We check if these manipulated properties cause problems in navigation performance. A consistent VE should be easily represented mentally in a map-like fashion, while a VE with severe violations should prove difficult. Surprisingly, we found no substantial influence of the impossible VEs on navigation performance, and forced-choice tests showed little evidence that subjects were aware of manipulations. This suggests that the representation does not resemble a two-dimensional image-like map. Alternatives to consider are sensorimotor and graph-like representations.
We deal with the analysis of eye movements made on natural movies in free-viewing conditions. Saccades are detected and used to label two classes of movie patches as attended and non-attended. Machine learning techniques are then used to determine how well the two classes can be separated, i.e., how predictable saccade targets are. Although very simple saliency measures are used and then averaged to obtain just one average value per scale, the two classes can be separated with an ROC score of around 0.7, which is higher than previously reported results. Moreover, predictability is analysed for different representations to obtain indirect evidence for the likelihood of a particular representation. It is shown that the predictability correlates with the local intrinsic dimension in a movie.
The unresolved questions relating to binocular processing of motion include: Is the perceived speed of the motion in depth (MID) of an approaching object inversely proportional to the time to collision?; What visual information supports judgements of the direction of MID?; What is the relation between binocular and monocular processing in the perception of MID? We review whether the perception of stereomotion in depth of a monocularly visible object is caused entirely by a rate of change of disparity, and conclude that the difference between the horizontal velocities of the object's left and right retinal images makes at most only a small contribution to speed discrimination, but conclusions may be different for detection, perceived speed and directional discrimination. We review laboratory evidence on the relative importance of binocular and monocular information for interceptive action and collision avoidance and conclude that, in addition to the effect of considerable intersubject variability, the relative importance depends on the physical size of the approaching object, its distance and, if nonspherical, its direction of motion and whether it is rotating. We compare attempts to find whether the human visual system contains a mechanism specialized for the speed of cyclopean motion within a frontoparallel plane, and find the question ill-posed.
This article describes the relationship between Art, as painting or sculpture, and a new theory of perceptual meaning, which builds on and now further develops the Gestalt principles. A key new idea in the theory is that higher-order groupings principles exist which, like the spatial grouping articulated by the principle of Prägnanz, helps to associate and combine stimuli, but which, unlike the Gestalt laws, can explain combinations of dissimilar as well as similar forms of visual information in a lawful manner. Similarities and dissimilarities are put together again by virtue of another and more global grouping factor that overcomes the dissimilarities of the components: it is some kind of meaning principle that perceptually solves the differences among whole and elements at a higher level, making them appear strongly linked just by virtue of the differences. In this way, similarities and dissimilarities complement and do not exclude each other. Such higher-order principles of grouping-by-meaning are articulated and illustrated using Art, from prehistoric to modern.
Metacontrast masking is by no means a unitary phenomenon, as is evidenced in recent studies showing differences between masking of surface- and contour properties of target stimuli (Breitmeyer et al., 2006; Ishikawa et al., 2006). Optima of masking appear earlier for contour processing and feature-specific operations compared to the variety of brightness processing that shows up in area filling-in phenomena. The present study explored whether this rule of processing - contours first and area filling-in afterwards - will be sustained if target and mask are, respectively, a central and a peripheral part of a coherent or incoherent meaningful visual object. Observers were presented with gray-level targets (images of the central part of a visual object) that were masked by a following, spatially surrounding mask, which was a complementary part of that object. Consistently with earlier findings, it appeared that salient visibility of contours which belonged to the internal spatial area of the target part of the object was established earlier and the whole-surface brightness quality (i.e. gray level) later in the course of target microgenesis. The unexpected facilitative effect of within-object coherence on target visibility which appeared at longer stimuli onset asynchrony (SOA) between target and mask parts of the object and only with large target and mask supports either some bias effects or lateral facilitatory interaction between iso-oriented parts of target-mask configuration having long time constants. The absence of the effects of coherence and inversion of target-plus-mask composite with small stimuli does not support the reentrant, top-down accounts of object processing in the context of metacontrast interactions.
While repetition of a feature (position) unrelated to a response is acknowledged to be facilitatory, there is disagreement on whether priming for response-defining feature or spatial position is facilitatory or inhibitory. To address this question, we used simple feature targets to analyze the interactions between facilitatory and inhibitory mechanisms associated to the repetition of features and position, for responses given either to the feature or to the position. We were able to reproduce the general facilitatory effect when a feature was repeated, and the inhibitory effect when it was changed, although these feature priming effects were always in interaction with repetition effects of spatial position. The most interesting finding, however, was that repetition of spatial position showed facilitation when non-response-defining, and inhibition when coincident with the response (response-defining); that is, repetition effects of spatial position are strictly dependent on the object of the motor response (a feature vs the position itself), whereas repetition priming for features is not, suggesting the involvement of a different mechanism and different neural substrate in the two cases. These effects interact, resulting in an ecologically plausible heuristic of visual discrimination that facilitates recently viewed features appearing in recently visited positions, but inhibits recently visited positions containing features recently associated with a distractor.
Two experiments examined whether filling-in occurred at the blind spot when a line segment was presented on only one side of the blind spot. We used static and dynamic stimuli: a static test line segment and a pair of probe line segments were presented in Experiment 1 and a moving test line segment was presented in Experiment 2. We compared the probability that the proximal end was perceived to be on the blind spot side when the test line segment came into contact with the blind spot (blind spot condition) with that when the test line segment was outside the blind spot (control condition). The results of the two experiments showed that the proximal end was perceived to be more on the blind spot side in the blind spot condition than in the control condition. Notably, when a dynamic stimulus was presented below the blind spot, the mean amount of filling-in reached 2.84 degrees. Therefore, filling-in occurred at the blind spot even when a line segment was presented on only one side of the blind spot.
Contrast sensitivity for a Gabor target can be increased by a factor of two when identical patches are separated by about three wavelengths (lambda) and positioned collinearly (Polat and Sagi, 1993, 1994a, 1994b). The facilitation effect was found for a wide range of spatial frequencies but was tested with well-experienced observers. Since practice modifies the range of lateral interactions, in this study naive observers were tested in order to document the initial stage of collinear facilitation. Surprisingly, we found that facilitation is maximal for the high spatial frequencies and minimal for the low spatial frequencies. We also found that when experienced observers were tested, facilitation at the low spatial frequencies was evident, suggesting that the initially reduced facilitation was due to inefficient lateral interactions. We suggest that the absence of facilitation for low spatial frequencies is due to the slow propagation velocity of the remote input, resulting in a mismatch between the flanker's input and the target's integration time.
Several unresolved issues in stereopsis are discussed that are related to the visual processing of dynamic disparity information. These unresolved issues include: (1) how well does the visual system compute temporal change in disparity and what kind of computation is used; (2) what is the neurophysiological basis of such processing in humans; and (3) how is the information gleaned from such processing used for guiding human action. The resolution of these issues will likely involve an adoption of a philosophical perspective in which dynamic disparity is viewed as playing an important role in the control of behavior and in which stereopsis is studied within the context of control-systems analysis.
One of the major difficulties in graph classification is the lack of mathematical structure in the space of graphs. The use of kernel machines allows us to overcome this fundamental limitation in an elegant manner by addressing the pattern recognition problem in an implicitly existing feature vector space instead of the original space of graphs. In this paper we propose three novel error-tolerant graph kernels -- a diffusion kernel, a convolution kernel, and a random walk kernel. The kernels are closely related to one of the most flexible graph matching methods, graph edit distance. Consequently, our kernels are applicable to virtually any kind of graph. They also show a high degree of robustness against various types of distortion. In an experimental evaluation involving the classification of line drawings, images, diatoms, fingerprints, and molecules, we demonstrate the superior performance of the proposed kernels in conjunction with support vector machines over a standard nearest-neighbor reference method and several other graph kernels including a standard random walk kernel.
Repeating the same target's features or spatial position, as well as repeating the same context (e.g. distractor sets) in visual search leads to a decrease of reaction times. This modulation can occur on a trial by trial basis (the previous trial primes the following one), but can also occur across multiple trials (i.e. performance in the current trial can benefit from features, position or context seen several trials earlier), and includes inhibition of different features, position or contexts besides facilitation of the same ones. Here we asked whether a similar implicit memory mechanism exists for the size of the attentional focus. By manipulating the size of the attentional focus with the repetition of search arrays with the same vs. different size, we found both facilitation for the same array size and inhibition for a different array size, as well as a progressive improvement in performance with increasing the number of repetition of search arrays with the same size. These results show that implicit memory for the size of the attentional focus can guide visual search even in the absence of feature or position priming, or distractor's contextual effects.
It has been shown that isometric matching problems can be solved exactly in polynomial time, by means of a Junction Tree with small maximal clique size. Recently, an iterative algorithm was presented which converges to the same solution an order of magnitude faster. Here, we build on both of these ideas to produce an algorithm with the same asymptotic running time as the iterative solution, but which requires only a single iteration of belief propagation. Thus our algorithm is much faster in practice, while maintaining similar error rates.
We consider the problem of spatial-temporal modeling of interactive image interpretation. The interactive process is composed of a sequential prediction step and a change detection step. Combining the two steps leads to a semi-automatic predictor that can be applied to a time-series, yields good predictions, and requests new human input when a change point is detected. The model can effectively capture changes of image features and gradually adapts to them. We propose an online framework that naturally addresses these problems in a unified manner. Our empirical study with a synthetic data set and a road tracking dataset demonstrate the efficiency of the proposed approach.
It has been suggested that the deleterious effect of contrast reversal on visual recognition is unique to faces, not objects.Here we show from priming, supervised category learning, and generalization that there is no such thing as general invariance of recognition of non-face objects against contrast reversal and, likewise, changes in direction of illumination.However, when recognition varies with rendering conditions, invariance may be restored, and effects of continuous learning may be reduced, by providing prior object knowledge from active sensation.Our findings suggest that the degree of contrast invariance achieved reflects functional characteristics of object representations learned in a task-dependent fashion.
In this paper we define the content of information in an image and show how it can be computed by taking into account different levels of resolution, in the framework of information theory and the thermodynamics of irreversible transformations. The results thus obtained will eventually be exploited to derive a mechanism for active exploration of visual space suitable to perform a dynamic coupling between the agent and its environment.
A striking effect of selective attention on perception of first- and second-order motion has been termed 'attention-induced motion blindness' or AMB (Sahraie et al., 2001). The AMB paradigm is based on a rapid serial visual presentation (RSVP) task and causes a severe transient impairment of the detection of coherent motion in a random dot kinematogram (RDK). The effect crucially depends on irrelevant motion intervals (distractors) prior to the motion target. To account for this phenomenon, both psychophysical and electrophysiological studies point to the existence of a post-perceptual gate operated by attentional mechanisms that limits access to the encoded motion signals by higher cortical areas. Here, we report in a first experiment that the presentation of motion distractors reduces motion sensitivity (operationalised as motion coherence threshold) which is in line with the assumption of a temporal carry-over effect of distractor inhibition. In a second experiment, we show that the rate of recovery of AMB is independent of target salience. The results of the third experiment provide evidence against the assumption that AMB is due to a shift or expansion of the 'attentional spotlight'.
Attention modifies our visual experience by selecting certain aspects of a scene for further processing. It is therefore important to understand factors that govern the deployment of selective attention over the visual field. Both location and feature-specific mechanisms of attention have been identified and their modulatory effects can interact at a neural level (Treue and Martinez-Trujillo, 1999). The effects of spatial parameters on feature-based attentional modulation were examined for the feature dimensions of orientation, motion and color using three divided-attention tasks. Subjects performed concurrent discriminations of two briefly presented targets (Gabor patches) to the left and right of a central fixation point at eccentricities of +/-2.5 degrees , 5 degrees , 10 degrees and 15 degrees in the horizontal plane. Gabors were size-scaled to maintain consistent single-task performance across eccentricities. For all feature dimensions, the data show a linear increase in the attentional effects with target separation. In a control experiment, Gabors were presented on an isoeccentric viewing arc at 10 degrees and 15 degrees at the closest spatial separation (+/-2.5 degrees ) of the main experiment. Under these conditions, the effects of feature-based attentional effects were largely eliminated. Our results are consistent with the hypothesis that feature-based attention prioritizes the processing of attended features. Feature-based attentional mechanisms may have helped direct the attentional focus to the appropriate target locations at greater separations, whereas similar assistance may not have been necessary at closer target spacings. The results of the present study specify conditions under which dual-task performance benefits from sharing similar target features and may therefore help elucidate the processes by which feature-based attention operates.
In three experiments, the rate of acquisition of information from a visual display was measured, using Sperling's method of backward masking (Sperling, 1963). Experiment I (which included a control for letter redundancy) showed that the rate of acquisition from an array of letters or words is determined not by the number of visual features or letters to be processed, but by the number of names into which they are to be encoded. In a second, control experiment, no effect was found on acquisition rate of restricting the letter ensemble-size. The third experiment showed that the time needed to identify words of the same frequency of occurrence was identical for words of three or six letters and of one or two syllables. The results are contrasted with those obtained in RT and comparison tasks. They are interpreted as evidence that (1) the backward masking paradigm provides a direct measure of word identification latency; (2) visually presented words are processed as wholes: that is, prior to word identification there is no intermediate stage of representation - not subject to backward masking - in units (e.g., syllables, spelling patterns) smaller than a complete word; (3) words are initially represented in an abstract lexical code, which does not reflect either visual or motor attributes of the word name.