
We describe a system for obtaining a "generic" parts-based 3D object representation. We use range image data as the input, obtaining a 3D object representation based on 12 geon-like 3D part primitives as the output. The 3D parts-based representation consists of parts detected in the image and their identities. Unlike previous work, we do not make simplifying assumptions such as the availability of perfect line drawings, perfect segmentation, or manual segmentation.We propose a novel method of specifying "generic" 3D parts, i.e., by means of surface adjacency graphs (SAGs). Using the SAGs, we derive an extremely compact multi-view representation of the part primitives, consisting of a total of only 74 views for all 12 primitives. Based on the multi-view representation of parts, we present a method of performing part segmentation from range images, given a good surface segmentation. This method for partsegmentation is more general than common approaches based on Hoffman and Richards′ "principle of transversality." We present two approaches for identifying the parts as one of the 12 3D part primitives. The first approach applies statistical pattern classification methods using parameters estimated by superquadric fitting. Five features derived from the estimated superquadric parameters are used to distinguish between the 12 part primitives. Classification error rates are estimated for k-nearest-neighbor and binary tree classifiers, for real as well as for synthetic range images. The second approach for part identification draws inferences from the distribution of angles between surface normals and the principal axis of a part.We show that intensity data can be used to recover from some misclassifications yielded by the purely range-based methods of part identification. A simple test is applied to check the concavity or convexity of the part silhouette in the intensity image. This serves as a reliable test of whether the part axis is straight orcurved.Results of part segmentation and identification are presented for real range images of several multi-part objects. Our system successfully performs part segmentation and identifies the parts.
An approach to labeling the components of human faces from range images is proposed. The components of interest are those humans usually find significant for recognition. To cope with the nonrigidity of faces, a qualitative approach is used. The preprocessing stage employs a multi-stage diffusion process to identify convexity and concavity points. These points are grouped into components and qualitative reasoning about possible interpretations of the components is performed. Consistency of hypothesized interpretations is carried out using context-based reasoning. Experimental results on real images of several faces are provided.<>
Recently, the assumed goal of computer vision, reconstructing a representation of the scene, has been critcized as unproductive and impractical. Critics have suggested that the reconstructive approach should be supplanted by a new purposive approach that emphasizes functionality and task driven perception at the cost of general vision. In response to these arguments, we claim that the recovery paradigm central to the reconstructive approach is viable, and, moreover, provides a promising framework for understanding and modeling general purpose vision in humans and machines. An examination of the goals of vision from an evolutionary perspective and a case study involving the recovery of optic flow support this hypothesis. In particular, while we acknowledge that there are instances where the purposive approach may be appropriate, these are insufficient for implementing the wide range of visual tasks exhibited by humans (the kind of flexible vision system presumed to be an end-goal of artificial intelligence). Furthermore, there are instances, such as recent work on the estimation of optic flow, where the recovery paradigm may yield useful and robust results. Thus, contrary to certain claims, the purposive approach does not obviate the need for recovery and reconstruction of flexible representations of the world.
Several models of statistical estimation of motion from visual input are derived and analyzed theoretically and experimentally. We study a wide variety of models, ones that use least squares and ones that use maximum likelihood, with several different assumptions (dependent and independent noise, isotropic and non-isotropic noise), spherical and planar image surfaces, and different preprocessing (one based on correspondence and one based on disparity). We do all this analysis using only a few fundamental concepts from statistical estimation, so the relative merits and shortcomings of all the methods become evident. The experimental results provide a quantitative measure of these merits.
This paper presents a bibliography of nearly 1300 references related to computer vision and image analysis, arranged by subject matter. The topics covered include computational techniques; feature detection and segmentation; image analysis; two-dimensional shape; pattern; color and texture; matching and stereo; three-dimensional recovery and analysis; three-dimensional shape; and motion. A few references are also given on related topics, such as geometry, graphics, coding and processing, sensors and optical processing, visual perception, neural nets, pattern recognition, and artificial intelligence, as well as on applications.
This paper describes a physically inspired method for the recovery of the surface of 3D solid objects from sparse data. The method is based on a model of closed elastic thin surface under the action of radial springs which can be considered as the analogous, in spherical coordinates, to the well-known thin plate model. The model is a representation for whole-body surfaces which has the degrees of freedom for representing fine details. We formulate the surface recovery problem as the problem of minimizing a non-quadratic energy functional. In the hypothesis of small deformations, this functional is approximated with a quadratic one which is then discretized with the finite element method. We provide steepest-descent-like algorithms both for the case of small deformations and for that of large ones. Then we introduce a representation of our model in terms of its free deformation modes. This representation is extremely concise and is therefore suited for shape analysis and recognition tasks. Finally, we report on the results of experiments with synthetic and real data which show the performance of the method
Scale-space representation is a topic of active research in computer vision. The focus of the research so far has been on coarse-to-fine focusing methods, image reconstruction, and computational aspects. However, not much work has been done on the signal detection problem, i.e., detecting the presence or absence of signal models from noisy image scans using the scale-space. In this paper we propose four 1-D signal detection algorithms for separating pulse signals in an image scan from the background in the scale-space domain. These algorithms do not need any thresholding to detect the zero-crossings (zc′s) at any of the scales. The different algorithms are applicable to image scans with different noise and clutter characteristics. A simple algorithm works best for scans having low noise and clutter. When noise and clutter increase sufficiently, a more sophisticated algorithm must be used. The 1-D algorithms for pulse and edge detection can be used to detect 2-D closed objects in cluttered and noisy backgrounds. This is done by scanning the image row-wise (and column-wise) and working on the individual scans. Using this method, the algorithms are demonstrated on several real life images. Another objective of this paper is to conduct comparative analysis of (i) a single-scale system vs a multiscale system and (ii) white noise vs clutter. This is done by conducting an experimental statistical analysis on single-scale and multiscale systems corrupted by white noise or clutter. Performance indices such as probability of detection, probability of false alarms, and delocalization errors are computed. The results indicate that (i) the multiscale approach is better than the single-scale approach and (ii) the degradation in performance is greater with clutter than with white noise.
This paper mathematically analyzes and proposes new solutions for the problem of estimating the camera 3D location and orientation (pose determination) from a matched set of 3D model and 2D image landmark features. Least-squares techniques for line tokens, which minimize both rotation and translation simultaneously, are developed and shown to be far superior to the earlier techniques which solved for rotation first and then translation. However, least-squares techniques fail catastrophically when outliers (or gross errors) are present in the match data. Outliers arise frequently due to incorrect correspondences or gross errors in the 3D model. Robust techniques for pose determination are developed to handle data contaminated by fewer than 50.0% outliers. Finally, the sensitivity of pose determination to incorrect estimates of camera parameters is analyzed. It is shown that for small field of view systems, offsets in the image center do not significantly affect the location of the camera in a world coordinate system. Errors in the focal length significantly affect only the component of translation along the optical axis in the pose computation.
In Tarr and Black′s paper it is stated that computer vision research should be based on reconstruction, as it offers the most promising framework for achieving insight into human visual cognition. It is further stated that it is in agreement with evolution. The competing school, the purposive, is considered too specific and relevant mainly for construction of robotic related systems with a limited functionality. In this paper it is argued that the two schools should not be viewed as competing, but rather as complementary. The reconstruction approach is used for research in vision functionalities, which may be combined into operational systems through a purposive analysis from a global point of view. Such a combined approach to vision is necessary for addressing critical issues such as continuous operation and achievement of specific visual tasks, while maintaining the generality needed to obtain insight into visual cognition.
In this paper we pose the question of how well segmentation of images done by humans are well suited to be used as ground truth in training and evaluating image segmentation methods. In general we show that humans prefer image segmentation done by other humans, even though there are cases, where subjects reject segmentation done by other humans. In fact we found cases in which humans gave a higher marks to image segmentation by the algorithms then of those done by humans. We also show that humans do not like segmentation methods that tend to produce regions having (approximately) same sizes, pixel-wise. They preferred segmentation with variation in region size.
This paper addresses the issue of optimal motion and structure estimation from monocular image sequences of a rigid scene. The new method has the following characteristics: (1) the dimension of the search space in the nonlinear optimization is drastically reduced by exploiting the relationship between structure and motion parameters; (2) the degree of reliability of the observations and estimates is effectively taken into account; (3) the proposed formulation allows arbitrary interframe motion; (4) the information about the structure of the scene, acquired from previous images, is systematically integrated into the new estimations; (5) the integration of multiple views using this method gives a large 2.5D visual map, much larger than that covered by any single view. It is shown also that the scale factor associated with any two consecutive images in a monocular sequence is determined by the scale factor of the first two images. Our simulation results and experiments with long image sequences of real world scenes indicate that the optimization method developed in this paper not only greatly reduces the computational complexity but also substantially improves the motion and structure estimates over those produced by the linear algorithms.
A robust algorithm for edge detection is presented. The algorithm detects both roof- and step-type edges. A pixel is declared as an edge pixel if there is a consensus between different processes that try to determine if the pixel lies on a discontinuity. Robust methods are used to estimate local fits to windows in the pixel's neighborhood and accumulate votes from each fit. The use of robust estimators allows the transformation of any window possibly containing a discontinuity to a binary window containing a step edge in the location of the discontinuity. Conventional methods are used to detect this step edge.< >
Edge and line detection are important processes in many computer vision systems. This paper examines errors in each of these processes and how they are related, i.e. how do errors in edge detection propagate to cause errors in line detection. Much use is made of simulated images which produce ground truth information, allows many trials to be performed, and allows analysis of variance. The analysis is restricted to polyhedral models which is appropriate for line finding algorithms. Results are presented that show how edge error affects line error, which error dominates for one variable, and how the errors vary for one parameter namely the scale or sigma for the edge detector. A hypothesis addressed is whether knowing the performance of edge and line detectors independently with respect to ground truth allows the overall performance of the two processes on real data to be predicted.
In robot navigation a model of the environment needs to be reconstructed for various applications, including path planning, obstacle avoidance, and determining where the robot is located. Traditionally, the model was acquired using two images (two-frame structure from motion) but the acquired models were unreliable and inaccurate. Recently, research has shifted to using several frames (multiframe structure from motion) instead of just two frames. However, almost none of the reported multiframe algorithms have produced accurate and stable reconstructions for general robot motion. The main reason seems to be that the primary source of error in the reconstruction-the error in the underlying motion-has been mostly ignored. Intuitively, if a reconstruction of the scene is made up of points, this motion error affects each reconstructed point in a systematic way. For example, if the translation of the robot is erroneous in a certain direction, all the reconstructed points would be shifted along the same direction. The contributions of this paper include mathematically isolating the effect of the motion error (as correlations in the structure error) and showing theoretically that these correlations can improve existing multiframe structure from motion techniques. Finally it is shown that new experimental results and previously reported work confirm the theoretical predictions.
In this paper, an algorithm for grouping edges belonging to straight lines is presented. The algorithm uses as input data a labeled set of edge points represented by a list of coordinate-label pairs. The output is a graph whose nodes are rectilinear segments linked by relational properties. Collinearity, convergence, and parallelism can be easily taken into account. The main novelty of the method lies in extending the use of the Hough transform to a symbolic domain (i.e., labeled edges); it is shown that edge labeling can be used to partition the Hough space and to isolate contributions coming from different image areas. Moreover, it is demonstrated that a simple focusing mechanism can be applied (in order to speed up the matching with 3D models) by using relational properties provided by the output graph. In order to confirm the algorithm′s performances, results on synthetic images containing randomly generated textures of straight lines are presented. Finally, a complex road image is considered to point out the advantages of using the proposed representation and the attention-focusing mechanism to solve real-world problems.