Kohonen's self-organizing map is extended by a technique allowing the neurons in the feature map to compete in a selective manner. This is accomplished by introducing gated neurons prior to the winner-take-all layer. These gated neurons are activated by a cosine function with time-varying frequency. This results in a spatio-temporal signature at the output for each input pattern over a predetermined interval. This pattern is found to be unique in its characteristics and leads to very high degree of recognition results. The simulations performed on a standard texture recognition problem indicate excellent performance
The complex retinal neural layer in the human visual system is considered to perform certain early visual processing. The authors examine how an array of such complex neurons and their associated neural circuitry could be used for static input-output mapping without losing stability. The advantage of using the proposed complex retinal type neural model is that it has a spatio-temporal structure whose transient and steady-state responses provide adequate information for use in image recognition systems. In order to utilize these networks' full capability, it is necessary to adapt the network parameters by computing the gradients. A suitable method to obtain gradients for parameter adjustment is given. A constraint satisfaction approach is developed with observations concerning stability criteria for these networks.<>
In this paper, Kohonen's self-organizing feature map is modified by a novel technique of allowing the neurons in the feature map to compete in a selective manner. The selective competition is achieved by grating the N-dimensional feature space using a spatial frequency and setting a criterion for the neurons to compete based on the region in which the input pattern resides. The spatial grating and selective competition are achieved by introducing a gated neuronal architecture in the feature map. As the selection criterion changes with time, it generates a time sequence of winning node indexes providing more input information and potentially allowing higher classification performance. These time sequences are then used to predict the class label of the input pattern more accurately. Three possible class label prediction algorithms are formulated based on evidential reasoning method and Bayes conditional probability theorem. These are tested on real world 8-class texture and a synthetic 12-class 3D object recognition problems. The classification performance is then compared with the results obtained by using a standard statistical linear discriminant analysis
A novel approach of temporal feature extraction over a predefined window at every pixel of range images is presented. Range images are first smoothed by an edge preserving nonlinear filter and then decomposed into a time sequence of image frames by a gated neuronal array. The image pixels are selected via neuronal gates controlled by a spatial grating function whose parameters are the 3D distance to the center pixel of the window and the time dependent spatial grating frequency. As the frequency is varied from a maximum value down to zero over a preset discrete time period, statistical information on the set of pixels selected is gathered. This temporal information is considered to represent the surface characteristics over the window. In this paper, as a first step it is demonstrated by using realistic range data that the entropy information of the sequence of images at every pixel can be effectively used to identify the edges
In this paper it is shown that the spatio-temporal signature generated for any input pattern on a topologically ordered feature map using a gated neuronal architecture is invariant over a neighbourhood of the input pattern provided the input patterns lie in the interior of the decision space and the regions of competition created by n-dimensional spatial grating function at any given spatial frequency are open. The spatio-temporal signature in a Gated Neuronal Architecture uniquely represents a collection of disjoint regions in the feature space. For pattern classification the labeling of the set of disjoint regions represented by the spatio-temporal signature is obtained by using Bayes conditional probabilities. Simulation results indicate improved performance.< >
We elucidate the conditions for a simple relationship between a geometrical transformation of an image and the associated integral transform.
Human observers are generally capable of recognizing patterns invariant to their orientation, position and size within an image. Though techniques are available for similar performance in computer visual systems, most suffer from lack of uniqueness or computational complexity. In this paper we introduce a new adaptive approach to invariant pattern recognition which overcomes both these problems. This technique is based upon the intrinsic invariance properties of the pattern and the recognition criterion. Our simulations demonstrate that the number of templates required to gain efficient pattern recognition is considerably lower than previously thought.
Matched filters are often used to detect image objects. However, for an image scene consisting of many different patterns corrupted by noise, the direct use of matched filters is time consuming and the performance of the matched filter deteriorates if the noise is not stationary and white. We develop a hierarchical approach for the detection of multiple objects which is divided into three steps: (a) prefiltering; (b) pattern recognition; and (c) object detection. This approach reduces computation time by more than 50% and increases classification efficiency compared with the direct matched filtering approach.
In this paper we consider a technique for pattern classification based upon the development of prototypes which capture the distinguishing features (“disjunctive prototypes”) of each pattern class and, via cross-correlation with incoming test images, enable efficient pattern classification. We evaluate such a classification procedure with prototypes based on the images per se (direct code), Gabor scheme (multiple fixed filter representation) and an edge (scale space-based) coding scheme. Our analyses, and comparisons with human pattern classification performance, indicate that the edge-only disjunctive prototypes provide the most discriminating classification performance and are the more representative of human behaviour.
The paper proposes a modified observed image model which takes local statistical information into account. Using this model and local statistical measurements, a two-dimensional iterative image restoration algorithm is developed. Experimental results show that the visual qualities of the processed image are improved considerably over those obtained by conventional processing, particularly with respect to sharpness of restored images. It is envisaged that this approach could also be used to improve the performance of two-dimensional Kalman-type recursive restoration algorithms in image processing.
Matched filtering methods are often used to detect objects in images. However, for an image scene consisting of many different patterns which are corrupted by noise, direct use of such matched filters would be very time-consuming, and the performance of the matched filter is suboptimal if the noise is nonwhite. In these cases more complex filtering processes are required which are based on the statistical characteristics of the images involved.
In this paper we develop an adaptive matched filter for detection of signals embedded in nonwhite noise. This filter is based on minimization of the difference between signal autocorrelation and the signal-to-image cross-correlation function filtered accordingly. Our results show this filter lies somewhere between the optimal filter and worst-case detection performance—the latter being evaluated by standard receiver operating characteristic methods.
Although sufficient evidence exists for the perceptual decomposition of spatial information via essentially uncorrected detector arrays with specific signals, their general form and applicability to superthreshold spatial vision is still questionable. In this paper we propose a nonlinear adaptive matched filter model for spatial coding which not only estimates the underlying detector characteristics to represent the processes of image alignment and identification, but also produces quantitative bounds on the efficiency of the decision processes based on same. Of particular importance to predicting results or visual texture discrimination, alignment of images, detection of signals embedded in scenes, and edge extraction are (1) the apparent nonlinear thresholding activity of such detectors, (2) the importance of the gradient or edge response, and (3) the use of cross correlation or matching as a comparison procedure. All processes are illustrated and compared with human psychophysical data. This formulation for superthreshold identification phenomena reformulates image amplitude and phase registration in terms of underlying (derivative) matched filter responses.
Uttal (1969, 1971, 1973) has conducted a series of experiments investigating the role of dot configurations in the detection of visual forms. The results indicated that regular dot spacings and linear forms are optimal for shape detection. Uttal (1973) contends that these data support a linear autocorrelation theory of pattern detection. This paper emphasises the difference between pattern detection and recognition, and proposes a model for this latter process. The model is based on a concept of interpolation whereby the visual system constructs a shape representation by encoding rates of change of spatial parameters with respect to distance.
can be seen for higher order maps, in particular for the projection type C. Higher order SOMs' classiication performance is lower than the standard SOM under no projection pursuit on account of correlated inputs i.e. higher the order of correlation lower the recognition performance. 6 Conclusions In this paper, we have proposed a feature projection approach to enhance the classiication performance of the SOMs and demonstrated on a variety of SOMs that the performance can be improved by choosing an appropriate high dimensional manifold for feature projection. Tests on real world texture classiication indicate consistent performance for at least two projection types A and C, while projection type B is found to be good for SO-and TO-SOMs. Further tests to integrate these maps for improving prediction accuracies are under progress. Performance evaluation of spatio-temporal feature maps with gated neuronal architecture. Figure 3: Texture images used in the classiication experiment 4.2 SOM setup SOMs of size 10x10 are chosen for testing the projection pursuit hypothesis. These SOMs are trained on the input feature space spanned by the 1024 vectors. The initial conditions for all the maps were neighborhood size 5, learning rate 0.05 and 50 epochs. These were modiied in subsequent batch mode training as The following table 1 indicates the number of patterns correctly classiied for all three types of SOMs with and without projection pursuits. Percentage classiication is computed over 1024 input pattern set. In the above Table 1, FO-SOM, SO-SOM and TO-SOM indicate rst, second, and third order maps. While the improvements to classiication performance is marginal in rst order, satisfactory from 1 to input dimension m (extended from original dimension n). The learning rate function () has values in the interval 0,1] and is a monotonically decreasing function of the discrete time index t. The winning node is based on the neuron having a least distance to the input pattern. The distance metric is given by: