Infrared images can often provide information missing in visible images in a night driving system, but infrared images lack key features needed to operate a vehicle. Image fusion techniques can be used to combine the relevant information from both the visible and infrared images. The need for high frame rates in an automotive application motivates the investigation into computationally simple methods of combining the visible and infrared images. In this paper we present a computationally simple image fusion technique based on the discrete Haar wavelet transform and apply the technique to combine three images from cameras operating in different wavelength bands.
In a digital image network for traffic monitoring a large number of cameras are connected to control centers through a hierarchical network. Compressed image data and recognition results are transmitted over the network. With conventional approaches, each control center receives compressed image data along with preliminary recognition results from low level control centers or surveillance cameras. Each center needs to decompress image data for further recognition processing, and if necessary the center sends the compressed image data and recognition results to the upper-level control center. In order to increase the cost-efficiency of the digital image network, we propose eliminating the decompression required at each center by developing a recognition method which works in the compressed domain. The main stream of conventional image compression methods such as discrete cosine transform is based on spatial frequency which makes it difficult to carry out recognition processes in the compressed domain. In contrast, we will compress the image data by using attributes which are relevant both for compression and recognition. Examples of the common attributes are binary edge locations and the color information surrounding the edge. This and other information is retained in the compression domain to enable recognition without decompression
A major obstacle in the application of stereo vision to intelligent transportation system is high computational cost. In this paper, a PC based three-camera stereo vision system constructed with off-the-shelf components is described. The system serves as a tool for developing and testing robust algorithms which approach real-time performance. We present an edge based, subpixel stereo algorithm which is adapted to permit accurate distance measurements to objects in the field of view using a compact camera assembly. Once computed, the 3D scene information may be directly applied to a number of in-vehicle applications, such as adaptive cruise control, obstacle detection, and lane tracking. Moreover, since the largest computational costs is incurred in generating the 3D scene information, multiple applications that leverage this information can be implemented in a single system with minimal cost. On-road applications, such as vehicle counting and incident detection, are also possible. Preliminary in-vehicle road trial results are presented.
The original optical flow algorithm [1] dealt with a flow field that could vary from place to place in the image, as would typically occur when a camera is moved through a three-dimensional environment—or if objects moved in front of a fixed camera. A related, but simpler problem, is that of recovering the motion of an image, all parts of which move with the same velocity (section 4.3 in [2]). The solution of this problem enables “optical mice,” as well as motion compensation in hand-held video cameras—critical to their operation. We provide here background on moving images, a method for the estimation of image velocity using least squares applied to brightness derivatives, and an analysis of the sensitivity of the estimated image velocity to noise in brightness measurement.
An algorithm is described for recovering the six degrees of freedom of motion of a vehicle from a sequence of range images of a static environment taken by a range camera rigidly attached to the vehicle. The technique utilizes a least-squares minimization of the difference between the measured rate of change of elevation at a point and the rate predicted by the so-called elevation rate constraint equation. It is assumed that most of the surface is smooth enough so that local tangent planes can be constructed, and that the motion between frames is smaller than the size of most features in the range image. This method does not depend on the determination of correspondences between isolated high-level features in the range images. The algorithm has been successfully applied to data obtained from the range imager on the Autonomous Land Vehicle (ALV). Other sensors on the ALV provide an initial approximation to the motion between frames. It was found that the outputs of the vehicle sensors themselves are not suitable for accurate motion recovery because of errors in dead reckoning resulting from such problems as wheel slippage. The sensor measurements are used only to approximately register range data. The algorithm described here then recovers the difference between the true motion and that estimated from the sensor outputs.
Relative orientation is the recovery of the position and orientation of one imaging system relative to another from correspondences among five or more ray pairs. It is one of four core problems in photogrammetry and is of central importance in binocular stereo as well as in long-range motion vision. While five ray correspondences are sufficient to yield a finite number of solutions, more than five correspondences are used in practice to ensure an accurate solution with least-squares methods. Most iterative schemes for minimizing the sum of the squares of weighted errors require a good guess as a starting value. The author has previously published a method that results in the best solution without requiring an initial guess [ J. Opt. Soc. Am. A4, 629 ( 1987)] An even simpler method is presented here that utilizes the representation of rotations by unit quaternions.
We address the problem of recovering the motion of a monocular observer relative to a rigid scene. We do not make any assumptions about the shapes of the surfaces in the scene, nor do we use estimates of the optical flow or point correspondences. Instead, we exploit the spatial gradient and the time rate of change of brightness over the whole image and explicitly impose the constraint that the surface of an object in the scene must be in front of the camera for it to be imaged.
Finding the relationship between two coordinate systems by using pairs of measurements of the coordinates of a number of points in both systems is a classic photogrammetric task. The solution has applications in stereophotogrammetry and in robotics. We present here a closed-form solution to the least-squares problem for three or more points. Currently, various empirical, graphical, and numerical iterative methods are in use. Derivation of a closed-form solution can be simplified by using unit quaternions to represent rotation, as was shown in an earlier paper [ J. Opt. Soc. Am. A4, 629 ( 1987)]. Since orthonormal matrices are used more widely to represent rotation, we now present a solution in which 3 × 3 matrices are used. Our method requires the computation of the square root of a symmetric matrix. We compare the new result with that obtained by an alternative method in which orthonormality is not directly enforced. In this other method a best-fit linear transformation is found, and then the nearest orthonormal matrix is chosen for the rotation. We note that the best translational offset is the difference between the centroid of the coordinates in one system and the rotated and scaled centroid of the coordinates in the other system. The best scale is equal to the ratio of the root-mean-square deviations of the coordinates in the two systems from their respective centroids. These exact results are to be preferred to approximate methods based on measurements of a few selected points.
Finding the relationship between two coordinate systems using pairs of measurements of the coordinates of a number of points in both systems is a classic photogrammetric task. It finds applications in stereophotogrammetry and in robotics. I present here a closed-form solution to the least-squares problem for three or more points. Currently various empirical, graphical, and numerical iterative methods are in use. Derivation of the solution is simplified by use of unit quaternions to represent rotation. I emphasize a symmetry property that a solution to this problem ought to possess. The best translational offset is the difference between the centroid of the coordinates in one system and the rotated and scaled centroid of the coordinates in the other system. The best scale is equal to the ratio of the root-mean-square deviations of the coordinates in the two systems from their respective centroids. These exact results are to be preferred to approximate methods based on measurements of a few selected points. The unit quaternion representing the best rotation is the eigenvector associated with the most positive eigenvalue of a symmetric 4 × 4 matrix. The elements of this matrix are combinations of sums of products of corresponding coordinates of the points.
In this correspondence, we show how to recover the motion of an observer relative to a planar surface from image brightness derivatives. We do not compute the optical flow as an intermediate step, only the spatial and temporal brightness gradients (at a minimum of eight points). We first present two iterative schemes for solving nine nonlinear equations in terms of the motion and surface parameters that are derived from a least-squares fomulation. An initial pass over the relevant image region is used to accumulate a number of moments of the image brightness derivatives. All of the quantities used in the iteration are efficiently computed from these totals without the need to refer back to the image. We then show that either of two possible solutions can be obtained in closed form. We first solve a linear matrix equation for the elements of a 3 × 3 matrix. The eigenvalue decomposition of the symmetric part of the matrix is then used to compute the motion parameters and the plane orientation. A new compact notation allows us to show easily that there are at most two planar solutions.
This paper describes a system which locates and grasps parts from a pile. The system uses photometric stereo and binocu lar stereo as vision input tools. Photometric stereo is used to make surface orientation measurements. With this informa tion the camera field is segmented into isolated regions of a continuous smooth surface. One of these regions is then selected as the target region. The attitude of the physical ob ject associated with the target region is determined by histo graming surface orientations over that region and comparing them with stored histograms obtained from prototypical objects. Range information, not available from photometric stereo, is obtained by the PRISM binocular stereo system. A collision-free grasp configuration is computed and executed using the attitude and range data.
We develop a systematic approach to the discovery of parallel iterative schemes for solving the shape-from-shading problem on a grid. A standard procedure for finding such schemes is outlined, and subsequently used to derive several new ones. The shape-from-shading problem is known to be mathematically equivalent to a nonlinear first-order partial differential equation in surface elevation. To avoid the problems inherent in methods used to solve such equations, we follow previous work in reformulating the problem as one of finding a surface orientation field that minimizes the integral of the brightness error. The calculus of variations is then employed to derive the appropriate Euler equations on which iterative schemes can be based. The problem of minimizing the integral of the brightness error term is ill posed, since it has an infinite number of solutions in terms of surface orientation fields. A previous method used a regularization technique to overcome this difficulty. An extra term was added to the integral to obtain an approximation to a solution that was as smooth as possible. We point out here that surface orientation has to obey an integrability constraint if it is to correspond to an underlying smooth surface. Regularization methods do not guarantee that the surface orientation recovered satisfies this constraint. Consequently, we attempt to develop a method that enforces integrability, but fail to find a convergent iterative scheme based on the resulting Euler equations. We show, however, that such a scheme can be derived if, instead of strictly enforcing the constraint, a penalty term derived from the constraint is adopted. This new scheme, while it can be expressed simply and elegantly using the surface gradient, unfortunately cannot deal with constraints imposed by occluding boundaries. These constraints are crucial if ambiguities in the solution of the shape-from-shading problem are to be avoided. Differrent schemes result if one uses different parameters to describe surface orientation. We derive two new schemes, using unit surface normals, that facilitate the incorporation of the occluding boundary information. These schemes, while more complex, have several advantages over previous ones.
This is a primer on extended Gaussian images. Extended Gaussian images are useful for representing the shapes of surfaces. They can be computed easily from: 1. needle maps obtained using photometric stereo; or 2. depth maps generated by ranging devices or binocular stereo. Importantly, they can also be determined simply from geometric models of the objects. Extended Gaussian images can be of use in at least two of the tasks facing a machine vision system: 1. recognition, and 2. determining the attitude in space of an object. Here, the extended Gaussian image is defined and some of its properties discussed. An elaboration for nonconvex objects is presented and several examples are shown.
The problem of producing a colored image from a colored original is analyzed. Conditions are determined for the production of an image in which the colors cannot be distinguished form those in the original by a human observer. If the final image is produced by superposition of controlled amounts of colored lights, only a simple linear transform need be applied to the outputs of the image sensors to produce the control inputs required for the image generators. In systems which depend instead on control of the concentration or the fractional area covered by colored dyes, a more difficult computation is called for. This calculation may for practical purposes be expressed in table lookup form. The conditions for exact reproduction of colored images should prove useful in the design and analysis of image processing systems whose final output is intended for human viewing. Judging by the design of some existing systems, these rules are not generally known or adhered to. Modern computational techniques make it practical to tackle this problem now. Adherence to design constraints developed here is of particular importance where colors are to be judged when the original is not directly accessible to the observer as, for example, when it is on another planet.
In this paper, we show how to recover the motion of an observer relative to a planar surface directly from image brightness derivatives. We do not con~pute the optical flow as an intermediate step. We derive a set of nine non-linear equations using a least-squares formulation. A simple iterative scheme allows us to find either of two possible solutions of these equations. An initial pass over the relevmt image region is used to accumulate a number of moments of the image brightness derivatives. All of the quantities used in the iteration can be efficiently computed from these totals, without the need to refer back to the image. A new, compact notation allows us to show easily that there are a t most two planar solutions.
It is possible to obtain useful maps of surface albedo from remotely sensed images by eliminating effects due to topography and the atmosphere, even when the atmospheric state is not known. A simple phenomenological model of earth radiance that depends on six empirically determined parameters is developed given certain simplifying assumptions. The model incorporates path radiance and illumination from sun and sky and their dependencies on surface altitude and orientation. It takes explicit account of surface shape, represented by a digital terrain model, and is therefore especially suited for use in mountainous terrain. A number of ways of determining the model parameters are discussed, including the use of shadows to obtain path radiance and to estimate local albedo and sky irradiance. The emphasis is on extracting as much information from the image as possible, given a digital terrain model of the imaged area and a minimum of site-specific atmospheric data. The albedo image, introduced as a representation of surface reflectance, provides a useful tool to evaluate the simple imaging model. Criteria for the subjective evaluation of albedo images are established and illustrated for Landsat multispectral data of a mountainous region of Switzerland. The method exposes some of the limitations found in computing reflectance information using only the image-forming equation.
Shaded overlays for maps give the user an immediate appreciation for the surface topography since they appeal to an important visual depth cue. A brief review of the history of manual methods is followed by a discussion of a number of methods that have been proposed for the automatic generation of shaded overlays. These techniques are compared using the reflectance map as a common representation for the dependence of tone or gray level on the orientation of surface elements.
Optical flow cannot be computed locally, since only one independent measurement is available from the image sequence at a point, while the flow velocity has two components. A second constraint is needed. A method for finding the optical flow pattern is presented which assumes that the apparent velocity of the brightness pattern varies smoothly almost everywhere in the image. An iterative implementation is shown which successfully computes the optical flow for a number of synthetic image sequences. The algorithm is robust in that it can handle image sequences that are quantized rather coarsely in space and time. It is also insensitive to quantization of brightness levels and additive noise. Examples are included where the assumption of smoothness is violated at singular points or along lines in the image.
In a previous paper a technique was developed for finding reconstruction algorithms for arbitrary ray-sampling schemes. The resulting algorithms use a general linear operator, the kernel of which depends on the details of the scanning geometry. Here this method is applied to the problem of reconstructing density distributions from arbitrary fan-beam data. The general fan-beam method is then specialized to a number of scanning geometries of practical importance. Included are two cases where the kernel of the general linear operator can be factored and rewritten as a function of the difference of coordinates only and the superposition integral consequently simplifies into a convolution integral. Algorithms for these special cases of the fan-beam problem have been developed previously by others. In the general case, however, Fourier transforms and convolutions do not apply, and linear space-variant operators must be used. As a demonstration, details of a fan-beam method for data obtained with uniform ray-sampling density are developed.