An overview of the current status and research directions of the vehicle exemplar is given. We discuss the function of the component modules of the vision system, and the need to integrate the operation of diverse knowledge sources in visual recognition of vehicles. The concept of a reasoning strategy is introduced and illustrated. The identification of effective reasoning strategies appropriate to a wide range of images is a prerequisite to specifying the control mechanisms needed to guide the automatic recognition of vehicles. We have developed an interactive environment allowing the informed user to investigate the performance of visual cues, and the subsequent use of processing modules to form coherent reasoning strategies. The environment forms a support tool, used to develop a range of strategies which cover the application domain. It also represents a practical framework within which to develop a context-dependent control framework for autonomous vehicle recognition in natural scenes. This paper presents an overview of part of the work of the Alvey MMI-007 Consortium, and discusses issues concerned with the management of the various processing modules developed for the "vehicle exemplar". The design of a computer vision system to locate and recognise vehicles in natural daylight scenes has proved extremely challenging. The human observer in carrying out this task is able to draw upon knowledge of many different aspects of vehicles and their expected circumstances, including both the geometrical structure and disposition of vehicles and schematic expectations of general scenes. Machine vision needs a similar ability to exploit diverse sources of knowledge to contribute, as opportunity allows, to the understanding of the sensed image. In the course of the project, the consortium has explored a variety of methods for expressing and exploiting knowledge about different facets of the overall problem of seeing vehicles. We have created a set of "modules", each capable of reasoning about some limited visual or contextual problem, in the specific application of vehicle recognition. Associated papers give details of methods for using knowledge concerned with • The geometrical form of the vehicles, and their expected relationships with other objects in the scene 1 • The identification of the context of vehicles according to expectations about different types of scene 2 • Viewpoint independent perceptual groups which may be associated (ambiguously) with labelled parts of vehicles 3 • View-and model-specific patterns, of particular significance for the initial cueing of vehicles and the viewpoint 4 • The …
Experiments are reported on the use of an Assumption-based Truth Maintenance System (ATMS) [6] to establish a match between a 3-d model and a single 2-d image. We show that the ATMS improves the efficiency of the search for maximal combinations of consistently labelled features. A memory cost is incurred, associated with the recording system of the ATMS; this can be reduced by simple heuristics. Empirical evidence is presented quantifying the costs and benefits of the method.
This paper concerns model-based vision for road traffic scene analysis. A traffic vision system has recently been developed at The University of Reading. The three main modules of the system are Movement Detection, Vehicle Localisation and Discrimination, and Vehicle Tracking. This paper outlines our work on the localisation and discrimination module. Emphasis is on recovering 3D poses of road vehicles in given image regions. Two classes of algorithms are described, one based on symbolic image features (line segments), and the other simply on image intensity gradients. A priori knowledge about traffic scenes and vehicles is exploited to improve the performance and efficiency of the algorithms. The algorithms are tested extensively with routine outdoor traffic images, and examples are included to demonstrate their principles.
A driver controls a car by turning the steering wheel or by pressing on the accelerator or the brake. These actions are modelled by Gaussian processes, leading to a stochastic model for the motion of the car. The stochastic model is the basis of a new filter for tracking and predicting the motion of the car, using measurements obtained by fitting a rigid 3D model to a monocular sequence of video images. Experiments show that the filter easily outperforms traditional filters.
This paper concerns the interactive construction of geometric models of objects from image sequences. We show that when the objects are constrained to move on the ground plane, a simple direct SFM algorithm is possible, which is vastly superior to conventional methods. The proposed algorithm is non-iterative, and in general requires a minimum of three points in two frames. Experimental comparisons with other methods are presented in the paper. It is shown to be greatly superior to general linear SFM algorithms not only in computational cost but also in accuracy and noise robustness. It provides a practical method for modelling moving objects from monocular monochromatic image sequences.
A new algorithm is described for refining the pose of a model of a rigid object, to conform more accurately to the image structure. Elemental 3D forces are considered to act on the model. These are derived from directional derivatives of the image local to the projected model features. The convergence properties of the algorithm is investigated and compared to a previous technique. Its use in a video sequence of a cluttered outdoor traffic scene is also illustrated and assessed.
This paper describes and demonstrates a view-independent relational model (VIRM) in a vision system designed for recognising a known 3D object from single monochromatic images. The aim is to derive a model of an object able to effect recognition without invoking pose information. The system inspects a CAD model of the object from a number of different viewpoints to identify relatively view-independent relationships among component parts of the object. These relations are represented in the form of a hypergraph. The VIRM can be searched using a best-first technique to obtain hypotheses of vehicle poses which match image features.
This paper concerns the pose determination and recognition of vehicles in traffic scenes, which under normal conditions stand on the ground-plane. Novel linear and closed-form algorithms are described for pose determination from an arbitrary number of known line matches. A form of the generalised Hough transform is used in conjuction with explicit probability-based voting models to find consistent matches. The algorithms are fast and robust. They cope well with complex outdoor scenes.
Objects are often constrained to lie on a known plane. This paperconcerns the pose determination and recognition of vehicles in trafficscenes, which under normal conditions stand on the ground-plane. Theground-plane constraint reduces the problem of localisation and recognitionfrom 6 dof to 3 dof.The ground-plane constraint significantly reduces the pose redundancy of2D image and 3D model line matches. A form of the generalised Houghtransform is used in conjuction with explicit probability-based votingmodels to find consistent matches and to identify the approximate poses. Thealgorithms are applied to images of several outdoor traffic scenes andsuccessful results are obtained. The work reported in this paper illustratesthe efficiency and robustness of context-based vision in a practicalapplication of computer vision.Multiple cameras may be used to overcome the limitations of a singlecamera. Data fusion in the proposed algorithms is shown to be simple andstraightforward.
This paper presents recent developments to a vision-based traffic surveillance system which relies extensively on the use of geometrical and scene context. Firstly, a highly parametrised 3-D model is reported, able to adopt the shape of a wide variety of different classes of vehicle (e.g. cars, vans, buses etc.), and its subsequent specialisation to a generic car class which accounts for commonly encountered types of car (including saloon, hatchback and estate cars). Sample data collected from video images, by means of an interactive tool, have been subjected to principal component analysis (PCA) to define a deformable model having 6 degrees of freedom.Secondly, a new pose refinement technique using ''active'' models is described, able to recover both the pose of a rigid object, and the structure of a deformable model; an assessment of its performance is examined in comparison with previously reported ''passive'' model-based techniques in the context of traffic surveillance. The new method is more stable, and requires fewer iterations, especially when the number of free parameters increases, but shows somewhat poorer convergence.Typical applications for this work include robot surveillance and navigation tasks.
A novel algorithm is presented for determining the orientation of road vehicles in traffic scenes using video images. The algorithm requires no specific 3-D vehicle models and only uses local image gradient values. It may easily be implemented in real-time. Experimental results with a variety of vehicles in routine traffic scenes are included to demonstrate the effectiveness of the algorithm.
This paper reports the current state of work to simplify our previous model-based methods for visual tracking of vehicles for use in a real-time system intended to provide continuous monitoring and classification of traffic from a fixed camera on a busy multi-lane motorway. The main constraints of the system design were: (i) all low level processing is to be carried out by low-cost auxiliary hardware; (ii) all 3-D reasoning is to be carried out automatically off-line, at set-up time. The system developed uses three main stages: (i) pose and model hypothesis using 1-D templates, (ii) hypothesis tracking, and (iii) hypothesis verification, using 2-D templates. Stages (i) and (iii) have radically different computing performance and computational costs, and need to be carefully balanced for efficiency. Together, they provide an effective way to locate, track and classify vehicles.
This paper concerns the recovery of pose and scale of vehicles in traffic scenes which, under normal conditions, are constrained to be in contact with the ground-plane. Several closed-form algorithms are described for pose and scale recovery using known 2D-to-3D line matches. The algorithms directly exploit the ground-plane constraint and are applicable to an arbitrary number of line matches. The algorithms are tested extensively with both synthetic and real outdoor traffic images. They are found to be robust and perform satisfactorily with real images.
In this paper a semi-automatic approach towards efficient and flexible image database population is described. An interactive process aided by computer vision and image processing techniques is employed. Computer aided object selection, using region growing based on colour and texture, allows new labels to be added quickly and efficiently. Extracted features are used to propagate the label through the database. As the images are very unlikely to be acquired at the same viewpoint, the features should ideally be viewpoint invariant-an area which has been mostly overlooked despite its importance. This paper focuses on rotation invariance.
The motion of a car is described using a stochastic model in which the driving processes are the steering angle and the tangential acceleration. The model incorporates exactly the kinematic constraint that the wheels do not slip sideways. Two filters based on this model have been implemented, namely the standard EKF, and a new filter (the CUF) in which the expectation and the covariance of the system state are propagated accurately. Experiments show that i) the CUF is better than the EKF at predicting future positions of the car; and ii) the filter outputs can be used to control the measurement process, leading to improved ability to recover from errors in predictive tracking.
This paper reports novel algorithms for the efficient localisation and recognition of vehicles in traffic scenes, which eliminate the need for explicit symbolic feature extraction and matching. The algorithms make use of two a priori sources of knowledge about the scene and the objects: (i) the ground-plane constraint, and (ii) the fact that road vehicles are strongly rectilineal: The algorithms are demonstrated and tested using routine outdoor traffic images. Success with a variety of vehicles demonstrates the efficiency and robustness of context-based computer vision in road traffic scenes. The limitations of the algorithms are also addressed in the paper.
This workshop paper reports recent developments to a vision system for traffic interpretation which relies extensively on the use of geometrical and scene context. Firstly, a new approach to pose refinement is reported, based on forces derived from prominent image derivatives found close to an initial hypothesis. Secondly, a parameterised vehicle model is reported, able to represent different vehicle classes. This general vehicle model has been fitted to sample data, and subjected to a Principal Component Analysis to create a deformable model of common car types having 6 parameters. We show that the new pose recovery technique is also able to operate on the PCA model, to allow the structure of an initial vehicle hypothesis to be adapted to fit the prevailing context. We report initial experiments with the model, which demonstrate significant improvements to pose recovery.
This paper concerns the determination of orientation of road vehicles from monocular intensity images. A novel algorithm is presented which exploits known physical and geometric knowledge about traffic scenes to allow fast and model-independent determination of vehicle orientations. The algorithm eliminates the need for symbolic image feature extraction and image-to-model matching, and the computational cost is substantially reduced. In fact, since the algorithm only requires local gradient data, object orientation can be determined from the input video data on-the-fly, and the overall algorithm can easily be implemented in real-time. The algorithm is tested with both indoor and outdoor data. Successful results are obtained for a variety of vehicles in routine traffic scenes.
This paper reports the development of a highly parameterised 3-D model able to adopt the shapes of a wide variety of different classes of vehicles (cars, vans, buses, etc), and its subsequent specialisation to a generic car class which accounts for most commonly encountered types of car (includng saloon, hatchback and estate cars). An interactive tool has been developed to obtain sample data for vehicles from video images. A PCA description of the manually sampled data provides a deformable model in which a single instance is described as a 6 parameter vector. Both the pose and the structure of a car can be recovered by fitting the PCA model to an image. The recovered description is sufficiently accurate to discriminate between vehicle sub-classes.