The Holy Grail of autonomous ground robotics has been to make ground vehicles that behave like humans. Over the years, as a community, we have realized the difficulty of this task, and we have back pedaled from the initial Holy Grail and have constrained and narrowed the domains of operation in order to get robotic systems fielded. This has lead to phrases such as "operation in structured environments" and "open-and-rolling terrain" in the context of autonomous robot navigation. Unfortunately, constraining the problem in this way has only put off the inevitable, i.e., solving the myriad of difficult robotics problems that we identified as long ago as the 1980's on the Autonomous Land Vehicle Project and in most cases are still facing today. These "Tall Poles" have included but are not limited to navigation through complex terrain geometry, navigation through thick vegetation, the detection of geometry-less obstacles such as negative obstacles and thin obstacles, the ability to deal with diverse and dynamic environmental conditions, the ability to function in dynamic and cluttered environments alongside other humans, and any combination of the above. This paper is an overview of the progress we have made at Autonomous Systems over the last three years in trying to knock down some of the tall poles remaining in the field of autonomous ground robotics.
A standard method for handling Bayesian models is to use Markov chain Monte Carlo methods to draw samples from the posterior. We demonstrate this method on two core problems in computer vision—structure from motion and colour constancy. These examples illustrate a samplers producing useful representations for very large problems. We demonstrate that the sampled representations are trustworthy, using consistency checks in the experimental design. The sampling solution to structure from motion is strictly better than the factorisation approach, because: it reports uncertainty on structure and position measurements in a direct way; it can identify tracking errors; and its estimates of covariance in marginal point position are reliable. Our colour constancy solution is strictly better than competing approaches, because: it reports uncertainty on surface colour and illuminant measurements in a direct way; it incorporates all available constraints on surface reflectance and on illumination in a direct way; and it integrates a spatial model of reflectance and illumination distribution with a rendering model in a natural way. One advantage of a sampled representation is that it can be resampled to take into account other information. We demonstrate the effect of knowing that, in our colour constancy example, a surface viewed in two different images is in fact the same object. We conclude with a general discussion of the strengths and weaknesses of the sampling paradigm as a tool for computer vision.
It is now well-established that k nearest-neighbour classifiers offer a quick and reliable method of data classification. In this paper we extend the basic definition of the standard k nearest-neighbour algorithm to include the ability to resolve conflicts when the highest number of nearest neighbours are found for more than one training class (model-1). We also propose model-2 of nearest-neighbour algorithm that is based on finding the nearest average distance rather than nearest maximum number of neighbours. These new models are explored using image understanding data. The models are evaluated on pattern recognition accuracy for correctly recognising image texture data of five natural classes: grass, trees, sky, river reflecting sky and river reflecting trees. On noise contaminated test data, the new nearest neighbour models show very promising results for further studies. We evaluate their performance with increasing values of neighbours (k) and discuss their future in scene analysis research.
In this paper we compare four classification techniques for classifying texture data of various natural objects found in forward-looking infrared (FLIR) images. The techniques compared include linear discriminant analysis, mean classifier and two different models of k-nearest neighbour methods. Hermite functions are used for texture feature extraction from segmented regions of interest in natural scenes taken as a video sequence. A total of 2680 samples for a total of twelve different classes are used for object recognition. The results on correctly identifying twelve natural objects in scenes are compared across the four classifiers on both unnormalised and normalised data. On unnormalised data, the average best recognition rate obtained using a ten fold cross-validation is 96.5%, and on unnormalised data it is 86.1% with a single nearest neighbour technique
Directional features extracted from Gabor responses are used as primitives for perceptual grouping. In previous work, we extracted Gabor features in 8 directions and then applied two self-organising maps, thus classifying each pixel in the image within a neuron-map, each corner of which represents one of four main directions. In this work we group pixels with similar directional features to detect salient structures within an image. Results obtained from application to forward-looking infrared (FLIR) images are very promising
In this paper we apply artificial neural networks for classifying texture data of various natural objects found in FLIR images. Hermite functions are used for texture feature extraction from segmented regions of interest in natural scenes taken as a video sequence. A total of 2680 samples for a total of twelve different classes are used for object recognition. The results on correctly identifying twelve natural objects in scenes are compared across ten folds of the cross-validation study. Neural networks are found to be extremely effective in robust classification of our data giving an average recognition rate of 91.8%.
The detection of image segmented objects in video sequences is constrained by the a priori information available with a classifier. An object recognizer labels image regions based on texture and shape information about objects for which historical data is available. The introduction of a new object would culminate in its misclassification as the closest possible object known to the recognizer. Neural networks can be used to develop a strategy to automatically recognize new objects in image scenes that can be separated from other data for manual labeling. In this paper, one such strategy is presented for natural scene analysis of FLIR images. Appropriate threshold tests for classification are developed for separating known from unknown information. The results show that very high success rates can be obtained using neural networks for the labeling of new objects in scene analysis.
1 © British Crown Copyright 1999/DERA; Published with the permission of the controller of Britannic Majesty's Stationary Office ABSTRACT In this paper we apply artificial neural networks for classifying texture data of various natural objects found in FLIR images. Hermite functions are used for texture feature extraction from segmented regions of interest in natural scenes taken as a video sequence. A total of 2680 samples for a total of twelve different classes are used for object recognition. The results on correctly identifying twelve natural objects in scenes are compared across ten folds of the cross-validation study. Neural networks are found to be extremely effective in robust classification of our data giving an average recognition rate of 91.8%.
Nearest Neighbour algorithms for pattern recognition have been widely studied. It is now well-established that they offer a quick and reliable method of data classification. In this paper we further develop the basic definition of the standard k-nearest neighbour algorithm to include the ability to resolve conflicts when the highest number of nearest neighbours are found for more than one training class (kNN model). We also propose aNN model of nearest neighbour algorithm that is based on finding the nearest average distance rather than nearest maximum number of neighbours. These new models are explored using image understanding data. The models are evaluated on pattern recognition accuracy for correctly recognising image texture data of five natural classes: grass, trees, sky, river reflecting sky and river reflecting trees. On noise contaminated test data, the new nearest neighbour models show very promising results for further studies when compared with neural networks. 1 © British Crown Copyright 1999/DERA Published with the permission of the controller of Britannic Majesty's Stationary Office; S. Singh, J.F. Haddon and M. Markou. Nearest Neighbour Strategies for Image Understanding, Proc. Workshop on Advanced Concepts for Intelligent Vision, Systems (ACIVS'99), Baden-Baden, (2-7 August, 1999).
Digital library applications require very general object recognition techniques. We describe an object recognition strategy that operates by grouping together image primitives in increasingly distinctive collections. Once a sufficiently large group has been found, we declare that an object is present. We demonstrate this method on applications such as finding unclothed people in general images and finding horses in general images. Finding clothed people is difficult, because the variation in colour and texture on the surface of clothing means that it is hard to find regions of clothing in the image. We show that our strategy can be used to find clothing by marking the distinctive shading patterns associated with folds in clothing, and then grouping these patterns.
Diffuse interreflections cause effects that make current theories of shape from shading unsatisfactory. We show that distant radiating surfaces produce radiosity effects at low spatial frequencies. This means that, if a shading pattern has a small region of support, unseen surfaces in the environment can only produce effects that vary slowly over the support region. It is therefore relatively easy to construct matching processes for such patterns that are robust to interreflections. We call regions with these patterns “shading primitives”. Folds and grooves on surfaces provide two examples of shading primitives; the shading pattern is relatively independent of surface shape at a fold or a groove, and the pattern is localised. We show that the pattern of shading can be predicted accurately by a simple model, and derive a matching process from this model. Both groove and fold matchers are shown to work well on images of real scenes
Perceptual organisation can be defined as the ability to impose structural organisation on sensory data, so as to group sensory primitives arising from a common underlying cause. Our organisational philosophy is hierarchical, with complex organisations being formed from simpler ones. In this paper, directional features extracted from Gabor responses are used as the primitives for perceptual grouping. In previous work, we extracted Gabor features in 8 directions and then applied two SOMs, thus classifying each pixel in the image within a 8x10 neuronmap, each corner of which represents one of four main directions, (horizontal, vertical, left diagonal and right diagonal). In the present work we group pixels with similar directional features, thereby detecting salient structures within an image. These detected-structures will be used as tokens from which to create the next level of abstraction in the hierarchy of the system. This approach is an alternative to the use of sets of edges as primary features: the directional features that Gabor filters provide are a potentially richer source of information. Preliminary results obtained from application to forward-looking infrared (FLIR) images are very promising. At present only four main directions are utilised, i.e. vertical, horizontal, right diagonal and left diagonal: the technique may be readily extended to the eight utilised in previous work. The next stage will be to group the tokens by the application of additional Gestalt-laws in order to detect objects.
This paper present an algorithm based on co-occurrence matrices for the automatic analysis of images. This includes the simultaneous segmentation of images into key regions and the detection of the main boundaries. The texture is described using discrete Hermite functions to decompose cooccurrence matrices of the segmented regions and are subsequently labelled using neural network based classifiers. Local consistency of interpretation of the analysis is imposed using relaxation labelling techniques. The algorithm presented is generally applicable to a wide range of imaging domains and applications with only minor variations on the basic theme, for example, the incorporation of a priori knowledge in an intelligent manner to ensure that parameters are adapted to the particular applications. Experimental results ave presented for 2 scenarios: scene analysis of a sequence of infrared images taken from a low flying aircraft and the analysis of snow profiles for the assessment of snowpack stability.
This paper discusses ego motion recovery from planar scenes with a camera view parallel to the direction of movement. Reasons for the failure of two published algorithms are considered, and a novel method a successful recovery is presented.
This paper present a relaxation labelling technique for enforcing local consistency of region classification, both spatially within an image and temporally between images in a sequence of images. The technique is region based as opposed to more traditional relaxation labelling techniques which are pixel based and uses region neighbours spatially within an image and temporally from previous images in the sequence. The technique is demonstrated on a sequence of 300 segmented infrared images in which the initial region labelling has been performed using texture features and neural network classifiers. The relaxation labelling is extremely successful at enforcing local consistency and is able to label regions which were too small for an initial classification
Texture can be interpreted as a measure of the 'edginess' about a pixel and can thus be described by edge co-occurrence matrices. The matrix can be decomposed using 2-dimensional orthogonal Hermite functions, the coefficients of which provide a low order feature vector which is characteristic of the texture. The Hermite coefficients for 240 hand-segmented regions of grass, trees, sky and river from 60 forward looking infrared (FLIR) images have been used to train and validate 2 neural networks, which have subsequently been used to label FLIR images segmented using co-occurrence techniques [1].
A cooccurrence space is defined by utilising the combinations of pixel strengths defined by a Canny edge operator. A region and boundary segmentation derived from this space is first edge thinned by non-maximal suppression and then hysteresis is used as a post-processing step to improve the edges. The distributions in cooccurrence space define the thresholds employed in the hysteresis post-processing.