Optic flow patterns have been identified as the primary cues in extracting 3-D shape features (Jain & Zaidi, PNAS 2011), deformations (Jain & Zaidi, JOV 2011) and material properties (Doerschner et al., Current Biology 2011) from motion signals. These patterns can be parsed into combinations of motion divergence and shear, which in turn have been linked to 3-D shape features and deformations (Koenderink and Van Doorn, 1975), and which can selectively activate MT/MST cells. We measured human performance on identification of nonrigid shapes, classification of deformations, and detection of shear and divergence motion patterns. We compared performance on foveal versus 4 degrees peripheral stimuli with a cortical magnification factor of 2.09. In Experiment 1, observers performed an 8AFC shape identification task on point-light ellipsoidal 3-D shapes with three Gaussian features (indentations or projections), and we estimated identification thresholds as a function of indentation/projection height. Performance was similar for rigid and nonrigid shapes, but was better at fovea than at periphery. In Experiment 2, observers performed a 3AFC deformation classification task on horizontal point-light cylinders that were either rigid or flexed nonrigidly along depth or in the image plane. Observers were consistently better at identifying cylinders that flexed in the image-plane than those that flexed in depth. Surprisingly, their performance was better in the periphery than at the fovea for both nonrigidities. In Experiment 3, observers' performance was similar in the fovea and periphery for both shear and divergence patterns, indicating that the magnification factor was successful in equating sensitivity for the elementary patterns, but not for shape or deformation identification. Sensitivities to combinations of motion patterns cannot thus be predicted from sensitivities to elementary motion patterns alone. These results suggest that estimating 3-D shapes and deformations may involve heuristics that employ non-linear functions or derivatives of the elementary motion patterns. Meeting abstract presented at VSS 2014
Color spaces are invaluable for specifying colors. However, color-matching spaces only predict which spectral distributions will match, while most perceptual color spaces describe the discrimination of small color differences. Yet, color perception involves more than matching and discrimination. For example, similarities between colors (Zaidi & Bostic, 2008) and between color changes (Zaidi, 1998) guide material identification across illuminants. Unfortunately, uniform color spaces based on multi-dimensional scaling of similarity ratings rely on untenable Euclidean assumptions (Wuerger, et al, 1995). We investigated the geometrical structure underlying relative similarity judgments. In a metric space, distance would represent similarity magnitude, but even in a weaker Affine space, ratios of distances along a line would provide measures of relative similarity; parallelism would define similarity between color changes. We tested whether Affine geometry holds for a mid-point setting task. We chose quadrilaterals in the MacLeod-Boynton (1979) equiluminant color plane. Observers viewed three colored patches. Two test patches were vertices of one color quadrilateral edge. Observers were instructed to consider the color change between the test patches in terms of “reddish-greenish“ and “bluish-yellowish” components and adjusted the hue and saturation of the third patch to the combined midpoint on these two dimensions. After finding the four edge midpoints, observers set the midpoints between the two pairs of opposing mid-points. The two final mid-points for each quadrilateral coincided, satisfying Varignon's Theorem (an Affine test). These settings also held for different adaptation conditions. A perceptual color space based on relative similarities across large color differences might have Affine structure.
Color spaces (e.g., CIE, Macleod-Boynton, and CIELUV) are invaluable for specifying colors. However, CIE and M-B space only predict which spectral distributions will match, while CIELUV deals with the discrimination of small color differences. There is much more to color perception than matching and discrimination. For example, similarities between colors (Zaidi & Bostic, 2008) and between color changes (Zaidi, 1998) can be used to identify materials across illuminants. Uniform color spaces based on multi-dimensional scaling of similarity ratings do exist, but these rely on Euclidean assumptions shown to be untenable (Wuerger, et al, 1995). We investigated the geometrical structure underlying relative similarity judgments. In a metric space, the distance between chromaticities would represent their magnitude of similarity, but even in a weaker affine space, ratios of distances between colors on a line would provide measures of relative similarity, and parallelism would define similarity between color changes. We tested whether affine geometry holds for a mid-point setting task. We chose two large quadrilaterals in the M-B equiluminant color plane. On each trial, observers viewed three colored patches, two of which were the endpoint colors forming one side of the quadrilateral. Observers were instructed to consider the color change between the test patches in terms of "reddish-greenish" and "bluish-yellowish" components and to set the color of the middle patch, by adjusting its hue and saturation, to the combined midpoint of the change on the two dimensions. After finding the mid-point for the four sides, observers set the midpoints between the two pairs of facing mid-points. For four observers, the two final mid-points for each quadrilateral coincided, thus satisfying Varignon’s Theorem and passing the affine test. A perceptual color space based on relative similarities across large color differences thus has an affine structure. Meeting abstract presented at VSS 2013
Frequent pitfalls of relying solely on visual appearances are theories that confuse the products of perception with the processes of perception. Being blatantly reductionist and seeking cell-level explanations helps to conceive of underlying mechanisms and avoid this pitfall. Sometimes the best way to uncover a neural substrate is to find physically distinct stimuli that appear identical, while ignoring absolute appearance. The prime example was Maxwell's use of color metamers to critically test for trichromacy and estimate the spectral sensitivities of three classes of receptors. Sometimes it is better to link neural substrates to particular variations in appearance. The prime example was Mach's inference of the spatial gradation of lateral inhibition between neurons, from what are now called Mach-bands. In both cases, a theory based on neural properties was tested by its perceptual predictions, and both strategies continue to be useful. I will first demonstrate a new method of uncovering the neural locus of color afterimages. The method relies on linking metamers created by opposite adaptations to shifts in the zero-crossings of retinal ganglion cell responses. I will then use variations in appearance to show how 3-D shape is inferred from orientation flows, relative distance from spatial-frequency gradients, and material qualities from relative energy in spatial-frequency bands. These results elucidate the advantages of the parallel extraction of orientations and spatial frequencies by striate cortex neurons, and suggest models of extra-striate neural processes. Phenomenology is thus made useful by playing with identities and variations, and considering theories that go below the surface. Meeting abstract presented at VSS 2013
Jain & Zaidi (2011) in a shape-from-motion paradigm showed that human observers are as good at making categorical judgments (fat vs thin) for nonrigid shapes as they are for rigid shapes based solely on motion cues. They showed for the first time that an explicit rigidity assumption was not required for extracting 3D shape from motion cues, at least for simple categorical judgments. In the current study we examined whether a more objective task such as shape identification would reveal an advantage for rigid objects, thus lending support to the rigidity assumption hypothesis. Stimulus consisted of point-light versions of ellipsoids that contained three bumps or dimples at three locations on the surface, thus leading to eight possible shapes. The ellipsoids were either rigid or deformed smoothly in the depth plane or the image plane. The stimuli were presented either at the fovea or at 4 deg. eccentricity after adjusting for the cortical magnification factor. Observers (N=4) performed an 8AFC shape-identification task and we measured their performance as a function of bump/dimple height/depth. Observers’ performance was identical for rigid and the two nonrigid ellipsoids both at the fovea and under peripheral viewing, thus providing further evidence against an explicit rigidity assumption. Observers’ performance in the periphery was comparable to their performance in the fovea after adjusting for cortical magnification factor. Our results suggest that extraction of 3D shape from motion cues in the periphery is limited by mechanisms that extract optic flow and lie lower in the visual system hierarchy than by higher level mechanisms that process the optic flow to extract 3D shapes. We propose multi-scale first-order optic flow analyses (div, def and curl) to extract both 3D shapes and deformations. Meeting abstract presented at VSS 2013
We have previously introduced frequency-band based analyses as an approach to infer material properties like roughness, thickness, and volume from images of fabrics (Giesel & Zaidi, Frequency based perception of material properties, Current Biology, under review). Here, we address how these material properties and their frequency-band representations are influenced by viewing parameters such as the distance or angle, and the illumination direction. To determine whether the viewing distance has an influence on the perception of material properties, we varied it by using two CRTs. The reference monitor was placed at a distance of 66cm, while the comparison monitor was placed at 33, 66, or 132cm. The original stimulus was displayed on the comparison monitor, while either the original or one of four manipulated versions of the original image was shown on the reference monitor. The observers' task was to indicate which of the images of a fabric displayed on the two monitors appeared as having more volume, as being thicker, or rougher, respectively. The results showed that over the tested range of viewing distances the perception of material properties remained largely constant indicating that the visual inference of material properties is more likely to be based on estimated material spatial frequency than on retinal frequency. We complemented the psychophysical results by image analyses using images from the KTH-TIPS database containing images of materials photographed at different distances, slants, and illumination directions. We show that the changes in material appearance are closely reflected in the frequency-band signatures. Finally, we present results from an experiment in which observers ranked printed versions of fabric images according to their volume, thickness, and roughness. We show that observers' ratings were closely correlated with amplitudes in the three frequency bands we identified previously. Meeting abstract presented at VSS 2013
The time course of visual responses is thought to play a major role in visual processing. For example, X and Y thalamic cells in the cat (M and P cells in the primate) have different temporal properties and are presumed to serve different functions. In contrast to X and Y visual pathways, ON and OFF pathways were originally thought to differ only in contrast polarity. However, using multi-unit recordings from cortical neurons in layer 4 we found that response latency to dark is shorter than light stimuli. Besides latency difference, we also found that ON and OFF responses are biphasic in nature and that the rebounds are stronger in ON than OFF responses. To evaluate the perceptual consequence of a latency difference we presented two square targets as dark/light pairs on either sides of a central fixation spot. 3 observers were instructed to report the location of the target that appeared first and the proportion of correct responses were calculated for different inter-target onset delays. Observers showed consistent temporal advantage for dark targets when presented on a uniform noise background. To measure the effect of rebound on perception, we used a two interval forced choice paradigm to present two successive spatially overlapping targets of like polarity in a randomly chosen interval. 3 observers were asked to report the interval in which they saw a flicker. The inter-target interval to perceive a flicker at 75% threshold performance was significantly lower for dark targets than light targets consistent with the physiological finding. We thus demonstrate using psychophysics, the functional correlates of latency and rebound differences observed in the neural responses to increments and decrements. Meeting abstract presented at VSS 2013
OBJECTIVE: Investigate role of top-down influences on recovering 3D shape from motion information, using objects varying in familiarity to test familiarity’s role on the tendency to perceive concave surfaces as convex. BACKGROUND: We reported (Papathomas et al, VSS 2011) that rotating hollow masks are perceived as convex faces rotating in the opposite direction, even in conditions where shape-from-motion signals have previously generated concave 3D percepts for artificial stimuli. We now test directly whether these results were dominated by object familiarity. METHODS: Experiment 1 used hollow, realistically painted, physical stimuli rotating on a turntable: (1) facial mask, (2) watermelon. Experiment 2 used four computer-generated concave stimuli: (1) Realistic human mask, using FaceGenTM; (2) ellipsoid rendered as watermelon; (3) ellipsoid with random-dot texture; (4) ellipsoid shown by longitude and latitude gridlines. In both experiments, the center (C) of the turntable was at a fixed distance from the observer. For artificial stimuli, motion parallax signals dominate the percept (Zaidi et al, 2011). We manipulated parallax by using 6 different rotational radii (distance between C and stimulus centroid). The illusion-strength was estimated by the time reported in the illusion divided by the total time that the concave side faced the observer. RESULTS: Experiment 1: The illusion was obtained for significantly longer intervals for the face than the watermelon; illusion-strength did not vary significantly with rotational radius. Experiment 2: Illusion-strength, averaged across rotational radii, was significantly higher for the human mask (44%) and watermelon (47%) than for the random-textured (35%) or gridline (28%) ellipsoids. CONCLUSIONS: The experiments provide evidence for a top-down bias to perceive familiar objects as convex that is greater than the bias for less familiar objects. Real objects are predominantly convex, so familiarity significantly influences the recovery of 3D structure and shape from bottom-up data-driven motion signals. Meeting abstract presented at VSS 2012
Rapid and reliable identification of material properties is important for successful interactions with the environment. Real materials exhibit characteristic configurations of low-level features and it is possible that features corresponding to particular properties are present generically and are stored as knowledge by human observers. We demonstrated at VSS2011, that some common properties perceivable from images of fabrics can be altered by increasing or decreasing the relative energy in specific spatial-frequency bands of the amplitude spectra. This result suggests that simple neural "detectors" for material properties could just combine the outputs of sets of V1 frequency-selective neurons. Can such detectors be revealed through selective adaptation? If observers adapt to a specific frequency-band, thus shifting the balance of sensitivities to other parts of the spectrum, does that shift their judgment of the associated material property? Using frequency-bands identified by our image analyses, we answer these questions for the fabric properties of volume, roughness, and thickness. For test stimuli, we used images of fabrics, along with two versions of each image with increased relative energy in the specific frequency band (constant total energy), and two with decreased relative energy. Baseline psychometric functions for each property were measured by comparing the original image side-by-side with its manipulated or unaltered version. During adaptation, bandpass-filtered dynamic white noise patches were presented on the location of the original image, and complementary notch-filtered dynamic noise patches on the location of the comparison images. The post-adaptation psychometric curve was measured by interleaving adaptation and test presentations. As predicted for each adaptation frequency-band, observers judged fabrics as having systematically less surface volume, softer texture, or thinner weave/knit than the original percept. The results demonstrate that observers directly use broad-band spatial frequency information in perceiving material properties, and suggest that V1 neurons transmit spatial-frequency signals in parallel for material perception. Meeting abstract presented at VSS 2012
Many objects deform when moving and are often partially occluded, requiring observers to integrate disparate local motion signals into coherent object motion. To test whether stereo information can help do this task, we measured the efficiency of extracting purely stereo-driven object motion by comparing tolerance of deformation noise by human observers to an optimal observer. We also compared the efficacy of global shape defined by stereo-disparities to the efficacy of local motion signals by presenting them in opposition. Stimuli consisted of 3-D shapes defined by disk-shaped random-dot stereograms uniformly arranged around a circle, varying randomly in stereoscopic depth. The disparities were oscillated to simulate clockwise or counterclockwise rotation of the 3-D shape. Observers performed a 2AFC direction discrimination task. The mean shape on a trial was constructed by assigning random depths to the disks (Shape Amplitude). The shape was dynamically deformed on every frame by independent random depth perturbations of each disk (Jitter Amplitude). Observers’ percent-correct discrimination declined monotonically as a function of Jitter Amplitude, but improved with Shape Amplitude. Observers were roughly 80% as efficient as the optimal shape-matching Bayesian decoder. Next, on each frame, we rotated the shape by 80% of the inter-disk angle, resulting in 2-D local motions opposite to the global object motion direction. Observers’ performances were largely unaffected at small dynamic distortions, but favored local motion signals at large dynamic distortions. Our original stimuli were devoid of monocular motion information, hence observers had to extract stereo-defined 3-D shapes in order to perform the task. These stimuli simulate disparity signals from both 3-D solid shapes and transverse-waves viewed through fixed multiple windows. Thus, our results provide strong evidence for the general role of 3-D shape inferences on object-motion perception. The cue-conflict stimuli reveal a trade-off between local motion and global depth cues based on cue-reliability. Meeting abstract presented at VSS 2012
Material perception may be as important as object perception for successful interactions with the environment. Possibly the most important facet of visual material identification is to infer properties from the available information that can provide predictions for the appearance of the material in other states, for effects of the material on other sensory modalities, for the calibration of motor actions, and most importantly for potential uses, i.e. affordances. We present an investigation of the inference of useful properties of fabrics from exclusively visual information. We used 261 color images of flat fabrics, presented on a monitor as similar sized squares surrounded by black. We asked observers to rate each image separately on four property dimensions anchored by the poles: soft – rough, stiff – flexible, warm – cool, and water-repellent – water-absorbent. To guide observers in inferring each property, we presented them with a question aimed at potential uses of the material. For data analysis, we used images that were rated consistently across repetitions and observers. We show that between 30% and 50% of the images rated by the observers as belonging to one property class were assigned consistently across observers. Based on verbal descriptions from the sources where we obtained the images, it seems that visual judgments of affordances are close to the ground truth, but to make more definite statements we will run physical and tactile tests on fabrics. Attempts at modeling fabrics have shown that since stitches/knits can be resolved visually, the surface cannot be assumed to be flat, so 2-D image-statistics and texture-mapping schemes are insufficient. As an alternative, we are using results that affordance groups like soft, flexible, and water-absorbent, and rough, stiff, and water-repellent, contain a large proportion of common images, to identify the critical perceptual qualities that underlie inferences of multiple material qualities.
At VSS 2010, we presented a new psychophysical method for measuring color afterimages. Colors of two halves of a bipartite disk were modulated sinusoidally (1/16 Hz to 2 Hz) from mid-gray to opposite ends of a color axis and back, e.g. grey > red > grey on one half and grey > green > grey on the other. The two halves appeared identical initially, increased in difference, then decreased to no difference, then increased again in opposite phase, so that when the physical modulation returned to grey, negative afterimages were perceived: the half modulated through red appeared green and vice versa. The physical contrast between the two halves, when they appeared identical, provided a Class A measurement of the after-image magnitude. Here, we present an early neural substrate for the afterimages by measuring the responses of retinal ganglion cells (RGC) to similar stimuli. Parafoveal RGCs were shown uniform circular patches modulating towards each of the poles of the preferred axis of the cell at 1/32 Hz and 1/16 Hz. The responses of Parvocellular and Koniocellular RGCs vigorously tracked modulations in their preferred direction, but decreased to base-rate 1–2 sec before physical modulations returned to mid-gray, dipped below base-rate and then recovered. Cell responses to modulation in the non-preferred direction, tracked the sinusoidal dip, but the response recovered faster than the stimulus, firing was significantly above base-rate when the stimulus reached grey, and the excitation persisted for a short time. Together, the excitation and inhibition of RGCs tuned to opposite directions along a color axis provide an early neural explanation for the afterimages: cells responding vigorously at mid-grey propagate an after-image signal to subsequent stages. RGC responses to both modulation frequencies were well described by a cone-opponent subtractive adaptation with a slow time constant of 5–10 seconds. Slow neural adaptation of the RGC population thus accounted for the after-image psychophysics.
Coloration and pattern are thought to play roles in crypsis for evading predators, and display for attracting mates or warning enemies. Colors of animals and their natural surroundings form distributions elongated between spectral reflectances and illumination spectra. We investigate whether color distributions contribute to camouflage and display independent of spatial information.We obtained 30 images of animals distinctly visible on the natural surround, and 29 images of animals camouflaged against their background. Colors were extracted from each pixel in the animal's image, and each pixel in the surrounding annulus of roughly the same width as the animal. The surround colors were randomly placed in the pixels of an 8 deg square, and the animal colors randomly in a 1.5 × 6 deg rectangle. Random masks were made as averages of all the images. Duration thresholds for discrimination of each animal color distribution from its background distribution, were measured by presenting rectangles randomly as horizontal or vertical in the center of the square, followed by a mask, in a 2AFC method of constant stimuli for durations from 8.5 to 100 msec. Besides the original distributions, the experiment also used isoluminant and achromatic versions of the random displays. Thresholds were determined from psychometric curves fitted to each image for each condition and for each observer. Comparing isoluminant thresholds to thresholds for the original distributions, reveals whether an animal distribution can be reliably detected on the basis of chromatic information alone. Comparing achromatic thresholds to thresholds for the original distributions reveals whether the addition of chromatic information enhances detection of an animal distribution. In general, both tests separated the images into the same groups, animals that could be detected using chromatic signals, and animals that could not. We will show that the usefulness of chromatic signals is related to the angular difference between color distributions.
perceivers to parse retinal images object motions, object shapes, and shape deformations. We have discovered interesting details about human perception of deforming shapes from motion cues, by using movies of rigid and flexing point-light cylinders rotating simultaneously around the depth and vertical axes in perspective (Jain & Zaidi, 2011): • Observers can discern cross-sectional shapes of flexing and rigid cylinders equally well, suggesting no advantage for structure-from-motion models using rigidity assumptions. • When cylinder rotations lead to asymmetric velocity profiles, symmetric cylinders appear asymmetric, highlighting the primacy of velocity patterns in shape perception. • Inexperienced observers are generally incapable of using motion cues to detect inflation/deflation of cylinders, but this handicap can be overcome with practice equally well for rigid and flexing objects. In this study we explore the ability of observers to classify dynamic deformations from motion cues. Classifying Dynamic 3-D Shape Deformations from Motion Cues
OBJECTIVE: In the rotating hollow face illusion (HFI), viewers perceive a hollow face as a convex face rotating in the opposite direction. Our objective: compare the strength of data-driven motion-perspective signals (faster-is-closer) to schema-driven influences (faces-are-convex) in the HFI. BACKGROUND: Meng and Zaidi (VSS 2008) obtained evidence for strong motion-perspective depth effects: two half-cycles of a sinusoidal corrugation (one concave, one convex) rotating about the zero crossing both appear convex; while rotating about their apex, both appear concave. HFI studies generally use masks rotating about an axis through the center of gravity. This generates retinal velocities that are larger for the nose than the cheeks and eyes. These relative velocities signal that the nose is closer, enhancing the illusion. We varied the relative velocities by placing the rotation axis at various distances from the nose. METHODS: We used: (1) physical masks painted realistically on both sides (convex and concave) rotating on a turntable; (2) computer-generated painted virtual masks (using FaceGen), rotating as in (1), allowing a greater range of axis-to-nose distances. We varied the relative velocities of the features, including conditions where the nose had a smaller velocity than the cheeks and eyes. We estimated the illusion's predominance as the time spent in the illusory percept divided by the total time that the mask had its concave side facing the observer. RESULTS: In both experiments, the concave face was seen as convex for significant intervals, even when the nose had lower velocity than the cheeks and eyes, which should reinforce the concave percept based on motion perspective. There was relatively little variation of the illusion predominance as the axis position was varied. CONCLUSIONS: The results provide evidence that, in the HFI, the top-down prior of convex faces dominates the velocity-driven 3D shape-from-motion signals resulting from object rotation.
After fixating on a colored pattern, observers see a similar pattern in complementary colors when the stimulus is removed [1-6]. Afterimages were important in disproving the theory that visual rays emanate from the eye, in demonstrating interocular interactions, and in revealing the independence of binocular vision from eye movements. Afterimages also prove invaluable in exploring selective attention, filling in, and consciousness. Proposed physiological mechanisms for color afterimages range from bleaching of cone photopigments to cortical adaptation [4-9], but direct neural measurements have not been reported. We introduce a time-varying method for evoking afterimages, which provides precise measurements of adaptation and a direct link between visual percepts and neural responses [10]. We then use in vivo electrophysiological recordings to show that all three classes of primate retinal ganglion cells exhibit subtractive adaptation to prolonged stimuli, with much slower time constants than those expected of photoreceptors. At the cessation of the stimulus, ganglion cells generate rebound responses that can provide afterimage signals for later neurons. Our results indicate that afterimage signals are generated in the retina but may be modified like other retinal signals by cortical processes, so that evidence presented for cortical generation of color afterimages is explainable by spatiotemporal factors that modify all signals.
Jeremy S. De Bonet合作论文数Skyward Mobile, LLC, Woburn, Ma3
Thomas V Papathomas合作论文数Department of Biomedical Engineering, School of Engineering, Rutgers University;Laboratory of Vision Research, Rutgers Center for Cognitive Science, School of Arts and Sciences, Rutgers University2