Humans looking around in the world can, seemingly without effort, segment out and distinguish different objects in the world. The corresponding capability has largely eluded the efforts of researchers in computer vision. Figure-ground segmention in general needs both context and task to be well-defined, i.e. may not be addressed using information in the visual scene alone. However, 3D cues play a special role: they indicate physical chunks that in turn can be ascribed visually observable 3D properties, such as position, location and motion, and object intrinsic properties such as shape, color and maybe surface and material characteristics. In the paper we will discuss segmentation of the scene into figure and ground and more generally into layers. Cues from stereo and motion will be used together with monocular cues from e.g. colour and texture. The goal is to acquire appearance models of the objects that can be used for subsequent processing, such as recognition. We will consider both moving and static objects, in the latter case assuming that 3D cues are available from either binocular stereo or observer motion. Integrating multiple cues is a key aspect of our approach and two techniques for this will be compared. One is a probabilistic approach where the likelihood of observing the data given a model of each layer is computed followed by a classification of each pixel using Bayes' rule. A second scheme is a voting method, the key difference being that each cue makes an independent decision regarding membership before these decisions are combined using a weighted sum. The advantage of voting in data fusion is that measurements drawn from very different spaces can easily be combined. With probabilistic methods more care must be taken in designing the model of each so that the different cues combine in the desired manner. Experiments on everyday scenes will show the performance of our methods and the type of object appearance models that can be acquire
Classifying materials from their appearance is challenging. Impressive results have been obtained under varying illumination and pose conditions. Still, the effect of scale variations and the possibility to generalise across different material samples are still largely unexplored. This paper (A preliminary version of this work was presented in Hayman et al. [E. Hayman, B. Caputo, M.J. Fritz, J.-O. Eklundh, On the significance of real world conditions for material classification, in: Proceedings of the ECCV, Lecture Notes in Computer Science, vol. 4, Springer, Prague, 2004, pp. 253–266].) addresses these issues, proposing a pure learning approach based on support vector machines. We study the effect of scale variations first on the artificially scaled CUReT database, showing how performance depends on the amount of scale information available during training. Since the CUReT database contains little scale variation and only one sample per material, we introduce a new database containing 10 CUReT materials at different distances, pose and illumination. This database provides scale variations, while allowing to evaluate generalisation capabilities: does training on the CUReT database enable recognition of another piece of sandpaper? Our results demonstrate that this is not yet possible, and that material classification is far from being solved in scenarios of practical interest.
This paper introduces a fast texture descriptor, the LU-transform. Itis inspired by previous methods, the SVD-transform and Eigen-transform, whichyield measures of image roughness by considering th ...
This document provides a brief Users’ Guide to the KTH-TIPS image database (KTH is the abbreviation of our university, and TIPS stands for Textures under varying Illumination, Pose and Scale). The guide describes which materials are contained in the database (Section 2), how images were acquired (Section 3) and subsequently cropped to remove the background (Section 4), and we also discuss some non-ideal artifacts, like poor focus, in some pictures (Section 5). This document concludes by outlining how we intend to extend the database in the future (Section 6).
This paper introduces a novel texture descriptor, the Eigen-transform. The transform provides a measure of roughness by considering the eigenvalues of a matrix which is formed very simply by inserting the greyvalues of a square patch around a pixel directly into a matrix of the same size. The eigenvalue of largest magnitude turns out to give a smoothed version of the original image, but the eigenvalues of smaller magnitude encode high frequency information characteristic of natural textures. A major advantage of the Eigen-transform is that it does not fire on straight, or locally straight, brightness edges, instead it reacts almost entirely to the texture itself. This is in contrast to many other descriptors such as Gabor filters or the standard deviation of greyvalues of the patch. These properties make it remarkably well suited to practical applications. Our experiments focus on two main areas. The first is in bottom-up visual attention where textured objects pop out from the background using the Eigen-transform. The second is unsupervised texture segmentation with particular emphasis on real-world, cluttered indoor environments. We compare results with other state-of-the-art methods and find that the Eigen-transform is highly competitive, despite its simplicity and low dimensionality.
Although a considerable amount of work has been published on material classification, relatively little of it studies situations with considerable variation within each class. Many experiments use the exact same sample, or different patches from the same image, for training and test sets. Thus, such studies are vulnerable to effectively recognising one particular sample of a material as opposed to the material category. In contrast, this paper places firm emphasis on the capability to generalise to previously unseen instances of materials. We adopt an appearance-based strategy, and conduct experiments on a new database which contains several samples of each of eleven material categories, imaged under a variety of pose, illumination and scale conditions. Together, these sources of intra-class variation provide a stern challenge indeed for recognition. Somewhat surprisingly, the difference in performance between various state-of-the-art texture descriptors proves rather small in this task. On the other hand, we clearly demonstrate that very significant gains can be achieved via different SVM-based classification techniques. Selecting appropriate kernel parameters proves crucial. This motivates a novel recognition scheme based on a decision tree. Each node contains an SVM to split one class from all others with a kernel parameter optimal for that particular node. Hence, each decision is made using a different, optimal, class-specific metric. Experiments show the superiority of this approach over several state-of-the-art classifiers
A powerful and popular approach for estimating radial lens distortion parameters is to use the fact that lines which are straight in the scene should be imaged as straight lines under the pinhole camera model. This paper revisits this problem using the division model to parameterise the lens distortion. This turns out to have significant advantages over the more conventional parameterisation, especially for a single parameter model. In particular, we demonstrate that the locus of distorted points from a straight line is a circular arc. This allows distortion estimation to be reformulated as circle-fitting for which many algorithms are available. We compare a number of suboptimal methods offering closed-form solutions with an optimal, iterative technique which minimises a cost function on the actual image plane as opposed to existing techniques which suffer from a bias due to the fact that they optimise a geometric cost function on the undistorted image plane.
Classifying materials from their appearance is a challenging problem, especially if illumination and pose conditions are permitted to change: highlights and shadows caused by 3D structure can radically alter a sample’s visual texture. Despite these difficulties, researchers have demonstrated impressive results on the CUReT database which contains many images of 61 materials under different conditions. A first contribution of this paper is to further advance the state-of-the-art by applying Support Vector Machines to this problem. To our knowledge, we record the best results to date on the CUReT database. In our work we additionally investigate the effect of scale since robustness to viewing distance and zoom settings is crucial in many real-world situations. Indeed, a material’s appearance can vary considerably as fine-level detail becomes visible or disappears as the camera moves towards or away from the subject. We handle scale-variations using a pure-learning approach, incorporating samples imaged at different distances into the training set. An empirical investigation is conducted to show how the classification accuracy decreases as less scale information is made available during training. Since the CUReT database contains little scale variation, we introduce a new database which images ten CUReT materials at different distances, while also maintaining some change in pose and illumination. The first aim of the database is thus to provide scale variations, but a second and equally important objective is to attempt to recognise different samples of the CUReT materials. For instance, does training on the CUReT database enable recognition of another piece of sandpaper? The results clearly demonstrate that it is not possible to do so with any acceptable degree of accuracy. Thus we conclude that impressive results even on a well-designed database such as CUReT, does not imply that material classification is close to being a solved problem under real-world conditions.
Statistical background modelling and subtraction has proved to be a popular and effective class of algorithms for segmenting independently moving foreground objects out from a static background, without requiring any a priori information of the properties of foreground objects. We present two contributions on this topic, aimed towards robotics where an active head is mounted on a mobile vehicle. In periods when the vehicle's wheels are not driven, camera translation is virtually zero, and background subtraction techniques are applicable. This is also highly relevant to surveillance and video conferencing. The first part presents an efficient probabilistic framework for when the camera pans and tilts. A unified approach is developed for handling various sources of error, including motion blur, subpixel camera motion, mixed pixels at object boundaries, and also uncertainty in background stabilisation caused by noise, unmodelled radial distortion and small translations of the camera. The second contribution regards a Bayesian approach to specifically incorporate uncertainty concerning whether the background has yet been uncovered by moving foreground objects. This is an important requirement during initialisation of a system. We cannot assume that a background model is available in advance since that would involve storing models for each possible position, in every room, of the robot's operating environment.. Instead the background mode must be generated online, very possibly in the presence of moving objects.
Reconstructing the scene from image sequences captured by moving cameras with varying intrinsic parameters is one of the major achievements of computer vision research in recent years. However, there remain gaps in the knowledge of what is reliably recoverable when the camera motion is constrained to move in particular ways. This paper considers the special case of multiple cameras whose optic centres are fixed in space, but which are allowed to rotate and zoom freely, an arrangement seen widely in practical applications. The analysis is restricted to two such cameras, although the methods are readily extended to more than two. As a starting point an initial self-calibration of each camera is obtained independently. The first contribution of this paper is to provide an analysis of near-ambiguities which commonly arise in the self-calibration of rotating cameras. Secondly we demonstrate how their effects may be mitigated by exploiting the epipolar geometry. Results on simulated and real data are presented to demonstrate how a number of self-calibration methods perform, including a final bundle-adjustment of all motion and structure parameters.
This paper presents algorithms for tracking unknown objects in the presence of zoom. Since prior models are unavailable, point and line matches in affine views are used to characterize the structure and to transfer a fixation point into new images in a sequence. Because any affine projection matrix is permitted, the intrinsic camera parameters such as focal length may change freely. Also, since the techniques do not require long feature tracks, a further desirable property is insensitivity to partial occlusion caused, for instance, by part of the object falling off the image plane while zooming in. If only point matches are available, a previous method based on factorization is applied. When also incorporating lines, the affine trifocal and quadrifocal tensors are used for tracking in monocular and stereo systems respectively. Methods for computing the tensors, minimizing algebraic error, are developed. In comparison with their projective counterparts, the affine tensors offer significant advantages in terms of computation time and convenience of parameterization, and the relations between the different tensors are shown to be much simpler. Successful tracking is demonstrated on several real image sequences.
Algorithms for self-calibrating cameras whose changes in calibration parameters are confined to rotation and zooming are useful since many real-world imaging situations do not permit translations-consider, for instance, cameras mounted on tripods and desk or wall-mounted active heads. In practice, however, the assumption of pure rotation is often violated because the optic center of the camera and the rotation center do not completely coincide. This work determines how such misalignments affect the estimation of the camera focal length. Expressions for the errors in focal length and recovered rotations are derived and results are confirmed with experiments on synthetic data. We show that the approximation of pure rotation is indeed sufficient in many cases, especially since other sources of error, such as noise and particularly radial distortion, tend to be more detrimental.
This paper describes techniques for fusing the output of multiple cues to robustly and accurately segment foreground objects from the background in image sequences. Two different methods for cue integration are presented and tested. The first is a probabilistic approach which at each pixel computes the likelihood of observations over all cues before assigning pixels to foreground or background layers using Bayes Rule. The second method allows each cue to make a decision independent of the other cues before fusing their outputs with a weighted sum. A further important contribution of ourwork concerns demonstrating how models for some cues can be learnt and subsequently adapted online. In particular, regions of coherent motion are used to train distributions for colour and for a simple texture descriptor. An additional aspect of our framework is in providing mechanisms for suppressing cues when they are believed to be unreliable, for instance during training or when they disagree with the general consensus. Results on extended video sequences are presented.
This paper describes techniques for fusing the output of multiple cues to robustly and accurately segment foreground objects from the background in image sequences. Two different methods for cue integration are presented and tested. The first is a probabilistic approach which at each pixel computes the likelihood of observations over all cues before assigning pixels to foreground or back- ground layers using Bayes Rule. The second method allows each cue to make a decision independent of the other cues before fusing their outputs with a weighted sum. A further important contribution of our work concerns demonstrating how models for some cues can be learnt and subsequently adapted online. In particular, regions of coherent motion are used to train distributions for colour and for a simple texture descriptor. An additional aspect of our framework is in providing mechanisms for suppressing cues when they are believed to be unreliable, for instance during training or when they disagree with the general consensus. Results on extended video sequences are presented.
Algorithms for self-calibrating cameras whose changes in calibration parameters are confined to rotation and zooming are useful since many real-world imaging situations do not permit translations — consider for instance cameras mounted on tripods and deskor wall-mounted active heads. In practice, however, the assumption of pure rotation is often violated because the optic centre of the camera and the rotation centre do not completely coincide. This work determines how such misalignments affect the estimation of the camera focal length. Expressions for the errors in focal length and recovered rotations are derived, and results are confirmed with experiments on synthetic data. We show that the approximation of pure rotation is indeed sufficient in many cases, especially since other sources of error such as noise and particularly radial distortion tend to be more detrimental.
In this paper we will show how rotation matrices obtained from image measurements, for instance via self-calibration of rotating cameras, may be used to self-align an active stereo head such that the optic axes of the cameras are parallel, horizontal and perpendicular to the elevation axis. Whereas general rotation matrices have three degrees of freedom, with robot heads we may generate motions with fewer degrees of freedom, and use prior knowledge of the kinematic chain to constrain the rotation matrices to have a particular form. We demonstrate how these constraints provide sufficient alignment information, and we provide fast, accurate and stable algorithms to achieve this task. Both a linear method, which minimizes an algebraic error, and two nonlinear methods which minimize meaningful errors, are presented, and we show in theory and experiments that they give virtually identical results. Results from both simulated data and from real images obtained from an active stereo head are presented
Zoom lenses appear to fit very naturally into the framework of active vision — controlling a zoom lens allows an adjustment of the image, enabling either an analysis of a wide scene or a close look at a region or object of particular interest. However, their integration into vision systems is not without difficulty since zoom interacts insidiously with both low and high level processes. This thesis concerns developing and analyzing algorithms that function in spite of zoom; algorithms for visual tracking, camera calibration and Euclidean reconstruction. An approach grounded in visual geometry is adopted, motivated by the notion that the geometric descriptions of point (corner) and line (straight edge) features are zoom-invariant. For tracking, point and line matches in affine views are used to transfer a fixation point into new images in a sequence. Since any affine projection matrix is permitted, the intrinsic camera parameters such as focal length may change freely. If only point matches are available, a previous method based on factorization is applied. When also incorporating lines, the affine trifocal and quadrifocal tensors are used for tracking with zoom in monocular and stereo systems respectively. Methods for the computation of the tensors, minimizing algebraic error, are developed. It is shown that the computation of the affine tensors offer significant advantages in comparison with the projective counterparts in terms of speed and convenience of parameterization. Although dealing with uncalibrated cameras is appealing, calibration information is necessary to set the gains in a visuo-control loop and for extracting Euclidean measurements from the scene. Further contributions of the thesis concern analyzing the reliability of the self-calibration of rotating and zooming cameras under practical situations, motivated by the observation that this configuration of cameras is common for instance in surveillance applications and in sports broadcasts. In particular, degenerate motions are identified and characterized, an analysis of near-ambiguities that may arise when the perspective effects are small is presented, and the effects of violations of the assumption that the motion of the camera is a pure rotation about its optic centre are characterized. Finally, work on self-calibration and Euclidean reconstruction from two or more such cameras is presented. An independent self-calibration of each rotating camera provides a starting point, but this process may be unable to recover all parameters accurately due to couplings between the intrinsic and extrinsic parameters. It is shown that the epipolar geometry may be utilized to alleviate the problem in a collaborative algorithm.
This paper considers the problem of self-calibration of a camera from an image sequence in the case where the camera's internal parameters (most notably focal length) may change. The problem of camera self-calibration from a sequence of images has proven to be a difficult one in practice, due to the need ultimately to resort to non-linear methods, which have often proven to be unreliable. In a stratified approach to self-calibration, a projective reconstruction is obtained first and this is successively refined first to an affine and then to a Euclidean (or metric) reconstruction. It has been observed that the difficult step is to obtain the affine reconstruction, or equivalently to locate the plane at infinity in the projective coordinate frame. The problem is inherently non-linear and requires iterative methods that risk not finding the optimal solution. The present paper overcomes this difficulty by imposing chirality constraints to limit the search for the plane at infinity to a 3-dimensional cubic region of parameter space. It is then possible to carry out a dense search over this cube in reasonable time. For each hypothesised placement of the plane at infinity, the calibration problem is reduced to one of calibration of a nontranslating camera, for which fast non-iterative algorithms exist. A cost function based on the result of the trial calibration is used to determine the best placement of the plane at infinity. Because of the simplicity of each trial, speeds of over 10,000 trials per second are achieved on a 256 MHz processor. It is shown that this dense search allows one to avoid areas of local minima effectively and find global minima of the cost function
This paper describes methods for tracking using point and line features in affine views to provide fundamental invariance to changes of focal length. It first demonstrates how an earlier method of transfer-based tracking using spatio-temporal matching of point features in a stereo active head is indeed zoom-invariant. In order to also make use of lines the paper then illustrates how the affine tri- and quadrifocal tensors may be applied to tracking with zoom in monocular and stereo systems respectively. The usefulness of the tensors is evident from their ability to transfer a fixation point in two uncalibrated images into novel views. Whereas the trifocal tensor is already familiar and in common use for matching and reconstruction we believe this to be the first practical application of a quadrifocal tensor. We develop expressions for affine tri- and quadrifocal tensors, and using novel affine specializations of existing projective algorithms we show how computation of the tensors is faster simpler and more stable. Experiments on real images are presented
A linear self-calibration method is given for computing the calibration of a stationary but rotating camera. The internal parameters of the camera are allowed to vary from image to image, allowing for zooming (change of focal length) and possible variation of the principal point of the camera. In order for calibration to be possible some constraints must be placed on the calibration of each image. The method works under the minimal assumption of zero-skew (rectangular pixels), or the more restrictive but reasonable conditions of square pixels, known pixel aspect ratio, and known principal point. Being linear the algorithm is extremely rapid, and avoids the convergence problems characteristic of iterative algorithms