Two recent versions of a single channel model of motion perception have had impressive success in explaining direction discrimination by human observers for spatially filtered noise images in two-flash apparent motion. It has been argued that the dramatic breakdown in motion perception which occurs when one image in the two-flash sequence is low-pass filtered can be explained only by a single channel model. We show that neither version of the single channel model which has been proposed can explain performance for noise images chosen to provide comparable stimulation in the spatial channels known to subserve human vision. A multi-channel model of motion perception has little difficulty in explaining these results.
The combination of visual motion information over visual space (spatial summation) and stimulus duration (temporal integration) was investigated using a random-pixel array (spatiotemporally broad-band) apparent motion stimulus designed to isolate specific populations of visual motion detectors. The results indicate that, in agreement with results from spatiotemporally narrow-band stimuli, spatial summation follows the form of linear probabilistic summation rather than non-linear probabilistic summation. Linear probabilistic summation holds for a wide range of stimulus parameters and when changing either motion stimulus height or width. Linear probabilistic summation breaks down when the motion display region approaches a height and/or width that is related to the spatial displacement size, not the speed, of the random-pixel array. This height and width (termed the critical height and width, or critical dimension), increases with spatial displacement size and can be interpreted as a measure of the basic dimensions of the selected motion detector population's receptive field. The critical height is smaller than the critical width, a result that is consistent with a motion detector receptive field that is elongated in the direction of motion. Perhaps most importantly, the mechanisms of temporal integration and spatial summation can work independently under a wide range of conditions. Finally, the results provide evidence for a short-term inhibitory phenomenon from the edges of the useful display area that affects the visibility of the motion.
A bi-local detector array model was assumed to describe the functional performance of monocular motion perception. Distributions of model parameters were measured in human vision at several positions in the visual field. The stimulus paradigm was designed to measure directional motion perception thresholds for individual combinations of spatial displacement and temporal delay in random dot apparent motion stimuli. The resulting data support previous results on perceivable spatial displacement limits in human vision but also indicate that both minimum and maximum perceivable spatial displacement thresholds in human observers have a similar dependence on temporal delay. This dependence changes with eccentricity in the visual field in a qualitatively similar manner but by quantitatively different factors. A description of possible biological properties of the bi-local detector population is presented that may explain how detection of spatio-temporal pattern displacements can be performed by a single system. Such a system also predicts that minimum and maximum perceivable spatial displacement thresholds should scale with visual field eccentricity in a manner consistent with our results.
The algorithms used in computing the intensity axis of symmetry (IAS) for 2D and 3D medical images are described. The basic 2D algorithms, along with the algorithms needed to incorporate scale space, are described. A brief discussion of the extensions needed to work with 3D images is given. The basic approach is to treat the image as a deformable intensity surface which is contracted onto the IAS. The primitive regions of the segmentation are identified by the branches in the resulting tree-like structure. A hierarchy is produced by following the simplification of the branching through scale space. For comparison papers see ibid., p.94-101 and ibid., p.108-14
Multiscale geometric image structure analysis is used to produce a hierarchical labeling of image regions. The regions provide a language for fast, interactive object definition. The approach allows human analysts to quickly inject semantics into the image representation, enhancing rather than trying to replace the human operator's capabilities.
A means is described of analyzing two- and three-dimensional images into a directed acyclic graph of visually sensible, coherent regions and of using this DAG as the basis for interactive object definition. The image analysis is in terms of the geometry of the intensity surface via a multiscale approach with a focus on symmetry properties about ridges. The image analysis method, a system for interactive object definition, and results of their use on two-dimensional images are reported.
The authors describe an image analysis that provides a quasi-hierarchy of image regions in terms of which humans can quickly build object regions by an interactive approach. This hierarchy is generated by the ridge structure of the intensity surface corresponding to the image. These ridges, in turn, are defined by an intensity axis of symmetry (IAS), which forms a branching structure in which the branches correspond to image regions and the parent/child relationships indicate ridge/subridge structure. The parent/child relationships are computed by following the IAS structure through changes in spatial scale, with scale change achieved by a diffusion in which conduction may be related to edge strength