Head tracking is a key technical component for AR and VR applications that use head‐mounted displays. Many different head‐tracking systems are currently in use, but one called “inside‐out” tracking seems to have the edge for consumer displays.
There have been three hypotheses about the utility of slit pupils: 1) Larger adjustments in area with simple musculature; 2) better image quality for contours perpendicular to the pupil’s long axis; 3) preserves chromatic-aberration correction in some lenses when pupil is constricted. These hypotheses do not explain why slits are always vertical or horizontal relative to the upright head, nor why they are vertical in terrestrial predators (e.g., domestic cats) and horizontal in terrestrial grazers (horses). Humans use blur to estimate depth in front of and behind fixation where depth from disparity is imprecise (Held et al., 2011). We simulated retinal images with various pupils to determine depth-of-field blur for different kinds of natural scenes. This leads to a new hypothesis concerning slit pupils. With slit pupils, depth of field is astigmatic: shorter for contours orthogonal to the pupil’s long axis; longer for perpendicular contours. Thus, depth from blur is more precise for orthogonal contours. The ground is a common environmental feature for terrestrial predators and grazers. With the head upright, the ground is foreshortened vertically in the retinal image, increasing the prevalence of horizontal contours. Vertical slits of terrestrial predators align the orientation of the shorter depth of field with horizontal contours allowing these animals to make finer depth discriminations along the ground, an advantage in their niche. Eyes of terrestrial grazers are laterally positioned in the head, so when the head pitches downward to graze, the pupils are roughly vertical relative to the ground. Again this aligns the orientation of the shorter depth of field with horizontal contours along the ground, which is advantageous (at least while grazing). We hypothesize that the orientation of slit pupils is an adaptation that provides some animals advantageous depth discrimination relative to common contours in their environment. Meeting abstract presented at VSS 2012
We present a system for producing 3D animations using physical objects (i.e., puppets) as input. Puppeteers can load 3D models of familiar rigid objects, including toys, into our system and use them as puppets for an animation. During a performance, the puppeteer physically manipulates these puppets in front of a Kinect depth sensor. Our system uses a combination of image-feature matching and 3D shape matching to identify and track the physical puppets. It then renders the corresponding 3D models into a virtual set. Our system operates in real time so that the puppeteer can immediately see the resulting animation and make adjustments on the fly. It also provides 6D virtual camera \\rev{and lighting} controls, which the puppeteer can adjust before, during, or after a performance. Finally our system supports layered animations to help puppeteers produce animations in which several characters move at the same time. We demonstrate the accessibility of our system with a variety of animations created by puppeteers with no prior animation experience.
Estimating depth from binocular disparity is extremely precise, and the cue does not depend on statistical regularities in the environment. Thus, disparity is commonly regarded as the best visual cue for determining 3D layout. But depth from disparity is only precise near where one is looking; it is quite imprecise elsewhere. Away from fixation, vision resorts to using other depth cues-e.g., linear perspective, familiar size, aerial perspective. But those cues depend on statistical regularities in the environment and are therefore not always reliable. Depth from defocus blur relies on fewer assumptions and has the same geometric constraints as disparity but different physiological constraints. Blur could in principle fill in the parts of visual space where disparity is imprecise. We tested this possibility with a depth-discrimination experiment. Disparity was more precise near fixation and blur was indeed more precise away from fixation. When both cues were available, observers relied on the more informative one. Blur appears to play an important, previously unrecognized role in depth perception. Our findings lead to a new hypothesis about the evolution of slit-shaped pupils and have implications for the design and implementation of stereo 3D displays.
Properly constructed stereoscopic images are aligned vertically on the display screen, so on-screen binocular disparities are strictly horizontal. If the viewer's inter-ocular axis is also horizontal, he/she makes horizontal vergence eye movements to fuse the stereoscopic image. However, if the viewer's head is rolled to the side, the on-screen disparities now have horizontal and vertical components at the eyes. Thus, the viewer must make horizontal and vertical vergence movements to binocularly fuse the two images. Vertical vergence movements occur naturally, but they are usually quite small. Much larger movements are required when viewing stereoscopic images with the head rotated to the side. We asked whether the vertical vergence eye movements required to fuse stereoscopic images when the head is rolled cause visual discomfort. We also asked whether the ability to see stereoscopic depth is compromised with head roll. To answer these questions, we conducted behavioral experiments in which we simulated head roll by rotating the stereo display clockwise or counter-clockwise while the viewer's head remained upright relative to gravity. While viewing the stimulus, subjects performed a psychophysical task. Visual discomfort increased significantly with the amount of stimulus roll and with the magnitude of on-screen horizontal disparity. The ability to perceive stereoscopic depth also declined with increasing roll and on-screen disparity. The magnitude of both effects was proportional to the magnitude of the induced vertical disparity. We conclude that head roll is a significant cause of viewer discomfort and that it also adversely affects the perception of depth from stereoscopic displays.
Perceiving three-dimensional video imagery appropriately in a display requires matching parameters throughout the imaging pathway, such as inter-aperture distance at the stereoscopic camera side with parallax shifting at the display side. In addition, many tradeoffs and compromises are often made at different points in the imaging pathway, leading to common perceptual distortions. Some of these may be simple two-dimensional image distortions such as display surface noise, while others are three-dimensional distortions, such as global geometric scene distortions and localized depth errors around edges. There is an increasing use of various forms of signal processing to modify the images, either for compensation of distortions due to system limitations, display constraints, formatting and compression for efficient transmission, or making depth range adjustments dependent on the display viewing conditions. Perceptual issues are critical to the design of the entire imaging pathway and this paper will highlight some of those due to stereoscopic signal processing.
Estimating the three-dimensional (3D) structure of the environment is challenging because the third dimension – depth – is not directly available in the retinal images. This “inverse optics problem” has for years been a core area in vision science (Berkeley, 1709; Helmholtz, 1867). The traditional approach to studying depth perception defines “cues” – identifiable sources of depth information – that could in principle provide useful information. This approach can be summarized with a depth-cue taxonomy, a categorization of potential cues and the sort of depth information they provide (Palmer, 1999). The categorization is usually based on a geometric analysis of the relationship between scene properties and the retinal images they produce.
Stereoscopic displays can potentially improve many aspects of medicine. However, weighing the advantages and disadvantages of such displays remains difficult, and more insight is needed to evaluate whether stereoscopic displays are worth adopting. In this article, we begin with a review of monocular and binocular depth cues. We then apply this knowledge to examine how stereoscopic displays can potentially benefit diagnostic imaging, medical training, and surgery. It is apparent that the binocular depth information afforded by stereo displays 1) aid the detection of diagnostically relevant shapes, orientations, and positions of anatomical features, especially when monocular cues are absent or unreliable; 2) help novice surgeons orient themselves in the surgical landscape and perform complicated tasks; and 3) improve the three-dimensional anatomical understanding of students with low visual-spatial skills. The drawbacks of stereo displays are also discussed, including extra eyewear, potential three-dimensional misperceptions, and the hurdle of overcoming familiarity with existing techniques. Finally, we list suggested guidelines for the optimal use of stereo displays. We provide a concise guide for medical practitioners who want to assess the potential benefits of stereo displays before adopting them.
Stereoscopic displays afford more accurate 3D percepts than conventional displays due to the added depth cue of disparity. However, 3D shape and scene layout are often misperceived when viewing stereoscopic displays. For example, viewing from the wrong distance alters an object’s perceived size and shape. It is crucial to understand the causes of such misperceptions so one can determine the best approaches for minimizing them. We develop the mathematics of an existing geometric model for calculating misperceptions, and then describe common viewing situations in which the model fails to make a prediction. We show how the visual system’s interpretation of vertical disparities can supplement the existing model and help predict the percepts associated with improper viewing of stereoscopic displays. We also discuss blur as a previously under-appreciated depth cue present in both stereo and non-stereo images. We present a probabilistic model that explains how the pattern of blur in an image together with relative depth cues indicates the apparent scale of the image’s contents. To examine the correspondence between the model/algorithm and actual viewer experience, we conducted an experiment with human viewers and compared their estimates of absolute distance to the model’s predictions. We did this for images with geometrically correct blur due to defocus and for images with commonly used approximations to the correct blur. The agreement between the experimental data and model predictions was excellent. A semi-automated algorithm is included, which helps one apply the correct pattern of blur to change the apparent size of a scene. The algorithm and model allow one to manipulate blur precisely and to achieve the desired perceived scale efficiently. Finally, we discuss the utility of stereoscopic displays for medical imaging. The technology’s greatest benefits arise in applications in which monocular cues are uninformative. Its specific strengths and shortcomings are presented in terms of diagnostics, education, surgical planning, minimally invasive surgery, and telesurgery. General guidelines and common errors are also listed to help potential users avoid unwanted misperceptions and visual fatigue.
Disparity is generally considered the most precise cue to depth, while blur is considered a coarse, qualitative cue. Depth from disparity and depth from blur have similar underlying geometries: one is based on triangulation between images collected by different eyes and the other is based on triangulation between images collected through different parts of the pupil. Thus, from a geometric standpoint, they provide complementary distance information. Physiologically, the two cues have very different sensitivities. Disparity thresholds, expressed as just-noticeable differences in depth, are low near fixation, but increase rapidly away from fixation. In contrast, blur thresholds are relatively large and do not vary significantly with position relative to fixation. Thus, one might expect disparity to determine depth discrimination near fixation and blur to determine discrimination away from fixation. We tested this expectation in a psychophysical experiment. Observers were presented a reference and test stimulus on each trial. The two stimuli were either both in front of fixation or both behind fixation. After each trial, observers indicated which stimulus appeared more distant. We used a novel volumetric display (Love et al., 2009) to present stimuli that contained 1) only disparity information (Gaussian dot viewed binocularly), 2) only blur information (a disk with 1/f noise viewed monocularly), or 3) both disparity and blur information (1/f disk viewed binocularly). As expected, thresholds were lower in the disparity-only condition than in the blur-only condition and the two-cue thresholds were similar to the disparity-only thresholds when the reference and test were near fixation. The situation reversed, however, behind fixation where blur thresholds were lower than disparity thresholds and two-cue thresholds were similar to blur-only thresholds. Thus, disparity and blur are complementary sources of information with disparity providing the best depth information near fixation and blur providing the best information away from fixation.
We present a probabilistic model of how viewers may use defocus blur in conjunction with other pictorial cues to estimate the absolute distances to objects in a scene. Our model explains how the pattern of blur in an image together with relative depth cues indicates the apparent scale of the image's contents. From the model, we develop a semiautomated algorithm that applies blur to a sharply rendered image and thereby changes the apparent distance and scale of the scene's contents. To examine the correspondence between the model/algorithm and actual viewer experience, we conducted an experiment with human viewers and compared their estimates of absolute distance to the model's predictions. We did this for images with geometrically correct blur due to defocus and for images with commonly used approximations to the correct blur. The agreement between the experimental data and model predictions was excellent. The model predicts that some approximations should work well and that others should not. Human viewers responded to the various types of blur in much the way the model predicts. The model and algorithm allow one to manipulate blur precisely and to achieve the desired perceived scale efficiently.
Focus cues—blur and accommodation—are regarded as very weak depth cues. The research that led to that assessment has, however, been greatly limited by an inability to present focus cues in a fashion consistent with natural viewing. We recently developed a novel volumetric display that allows the presentation of near-correct focus cues along with standard depth cues like binocular disparity. We used the display to re-examine the usefulness of blur and accommodation as cues to depth. We created stereograms with anisotropic textures that created disparities specifying a disk in front of a background. Probability-based theories of cue combination predict that as the disparity signal becomes less reliable, the percept should be more heavily determined by the focus cues. We varied the reliability of the disparity signal by changing the dominant orientation of the texture. Reliable (texture vertically oriented) and unreliable (texture horizontally oriented) disparity signals were presented in three conditions: i) focus cues specified zero depth, as they do in conventional 3d displays; ii) focus cues and disparity specified the same depth, as they do in natural viewing; iii) focus cues specified more depth than disparity. We presented stimuli in two intervals and observers reported the interval with more apparent depth. As expected, when focus cues specified zero depth, subjects saw less depth when the disparity signal was unreliable than when it was reliable; when focus cues specified more depth than disparity, the effect was reversed: subjects saw more depth with unreliable disparity. Based on pupil-diameter fluctuation, accommodative fluctuation, and optical aberrations, we computed the theoretically expected depth-from-focus likelihood functions. They are reasonably similar to the likelihood functions estimated from the data. These results show that focus cues provide a metric depth signal that is combined in a statistically reasonable fashion with disparity.
As stereoscopic displays become more commonplace, it is more important than ever for those displays to create a faithful impression of the 3-D structure of the object or scene being portrayed. This article reviews current research on the ability of a viewer to perceive the 3-D layout specified by a stereo display.
3d shape and scene layout are often misperceived when viewing stereoscopic displays. For example, viewing from the wrong distance alters an object's perceived size and shape. It is crucial to understand the causes of such misperceptions so one can determine the best approaches for minimizing them. The standard model of misperception is geometric. The retinal images are calculated by projecting from the stereo images to the viewer's eyes. Rays are back-projected from corresponding retinal-image points into space and the ray intersections are determined. The intersections yield the coordinates of the predicted percept. We develop the mathematics of this model. In many cases its predictions are close to what viewers perceive. There are three important cases, however, in which the model fails: 1) when the viewer's head is rotated about a vertical axis relative to the stereo display (yaw rotation); 2) when the head is rotated about a forward axis (roll rotation); 3) when there is a mismatch between the camera convergence and the way in which the stereo images are displayed. In these cases, most rays from corresponding retinal-image points do not intersect, so the standard model cannot provide an estimate for the 3d percept. Nonetheless, viewers in these situations have coherent 3d percepts, so the visual system must use another method to estimate 3d structure. We show that the non-intersecting rays generate vertical disparities in the retinal images that do not arise otherwise. Findings in vision science show that such disparities are crucial signals in the visual system's interpretation of stereo images. We show that a model that incorporates vertical disparities predicts the percepts associated with improper viewing of stereoscopic displays. Improving the model of misperceptions will aid the design and presentation of 3d displays.
Stereoscopic displays are becoming more common in fields as diverse as medical imaging and oil exploration. However, to ensure the usefulness of such displays, it is important to minimize any visual misperceptions experienced by their potential users. We present a novel graphical user interface designed for the exploration of misperceptions of stereoscopic images. The software includes controls for stimulus, image acquisition, and viewing parameters. As these settings are adjusted, the user can see their effects on the predicted stereoscopic percept in real time. The software employs two models for stereoscopic distortions to generate these predictions; one based purely on geometry and found throughout the stereocinema literature, and the other based on higher perceptual processes within the visual system. We demonstrate the software's utility in discovering various types of distortions that arise from improper image acquisition and viewing conditions. We believe this functionality would be useful for engineers who wish to optimize 3D displays for specific viewing situations. We also demonstrate how the software can be used to explore the differences between the two models of stereoscopic misperceptions, and suggest a way to test each model's accuracy against psychophysical data from human observers. CR Categories and Subject Descriptors: I.3.7 (Computer
Scott J. Daly合作论文数Dolby Laboratories2