This paper presents a real-time classification algorithm for 2D object contours using a multi-resolution tree model which is implemented in a modular VLSI architecture. The hardware implementation takes advantage of pipelining, parallelism, and the speed of VLSI technology to perform real-time object classification. Using the multi-resolution tree model, the classification algorithm is invariant under 2D similarity transformations and recognizes the visible portions of occluded objects. The VLSI classification system is implemented in 0.8 /spl mu/m CMOS and is capable of performing 34000 matchings per second.
An advanced approach to man-machine interaction is proposed, in which computer vision techniques are used for interpreting user actions. The key idea of the approach is the combined use of head motions for visual navigation and eye pupil positions for context switching within the graphical human-computer interface. This allows a partial decoupling of the visual models used for tracing eye features, with beneficial effects on both computational speed and adaptation to user characteristics. The applications range from navigation and selection in virtual reality and multimedia systems, to aids for the disabled and the monitoring of typical user actions in front of advanced terminals. The feasibility of the approach is tested and discussed in the case of a virtual reality application, the virtual museum
We present a fast parallel implementation of linear feature extraction on IBM SP-2. We first analyze the machine features and the problem characteristics to understand the overheads in parallel solutions to the problem. Based on these, we propose an asynchronous algorithm which enhances processor utilization and overlaps communication with computation by maintaining algorithmic threads in each processing node. Our implementation shows that, given a 512/spl times/512 image, the linear feature extraction task can be performed in 0.065 seconds on a SP-2 having 64 processing nodes. A serial implementation takes 3.45 seconds on a single processing node of SP-2. A previous implementation on CM-5 takes 0.1 second on a partition of 512 processing nodes. Experimental results on various sizes of images using 4, 8, 16, 32, and 64 processing nodes are also reported.
In nature, the visual detection of motion appears to be used in a variety of tasks, ranging from collision avoidance to posture maintenance. Many insects seem to rely primarily on information provided by an array of elementary movement detectors in order to navigate. Moreover, experimental evidence suggests that motion information is interpreted at an early stage of the insect visual system, and may be closely linked to motor control. A motion detector, whose design is based on some of the characteristics of the insect visual system, has been implemented on a single VLSI chip. This paper shows the manner in which motion information, provided by the chip in real-time, may be utilised by the control system of an autonomous vehicle in low-level perceptual tasks.
High speed image and video processing is a key technology in multimedia applications, and, therefore, currently, many hardware accelerators to speed up such processing are developed and used. However, for the next generation advanced multimedia applications in the next generation, such as high quality virtual reality, bidirectional visual interface, etc., the hardware accelerators can not deal with tasks involved in these applications. This is because these tasks consist of complex and irregular computation structures, and therefore, it is not easy to implement simple hardware accelerators for these tasks. Considering the above situation, for the next generation multimedia applications, we have been developing a MIMD based multimedia processor, KUMP/D (Kyushu University Multimedia Processor on Datarol-II). KUMP/D is a flexible parallel processor, based on fine grain parallel processing, which is which indispensable in complex and irregular computation, and is also equipped with a specialized I/O network. This I/O network has enough throughput for real time video I/O, providing a mechanism, which supports the synchronization of process executions and real time video frames.
This paper discusses two components of a Robot Eye intended as an active vision system to be mounted on a mobile robot. The first component is a foveated vision sensor which is based on an overlapping receptive field model for data reduction. We present the adapted scan-line algorithm used to compute so-called retinal images and a description of the implementation of the system on a network of DSP's. The second component computes salient points in the foveated image and is motivated by the biological processes which guide primate gaze fixation. The model of attention and its real-time implementation are described. Experimental results obtained with these algorithms are also presented.
The image represents an information structure of an extreme complexity. We present a method for symbolic and structured description with several levels of abstraction. The application area concerns the construction and interpretation knowledge on 3D objects in a machine perception. The knowledge representation and scenes interpretation tasks based on 2D image used "Perceived Aspects Tables" and "Quality Tables" proposed by our vision system. The necessary knowledge to this task are formalised. Then, we resolve the main interpretation problem, the control and the identification of objects, thanks to the approach called: "Prediction-Checking of Hypotheses". The representation, which is suggested, is based on the "Frames" Model. With this end in view, an organisation system of perception and scenes interpretation tasks in order to maintain a coherent representation of a structured and evolutive environment is designed. The released concepts and the proposed system are realised in a LE-LISP Object Oriented environment based on the SHIRKA knowledge representation system.
Although SIMD arrays have been built since the 1960's, they have undergone few empirical studies. The underlying problems-which have included the lack of a unified architectural framework and the computational intractability of simulating large PE arrays-are addressed through the use of trace compilation, a novel approach to trace driven simulation. The results indicate the benefits of adding another level to current SIMD array memory designs. Also, surprising results were obtained about performance effects of varying cache associativity and block size. Together, they indicate that while SIMD array programs have sufficient locality to make PE caches worthwhile, the type of locality may differ fundamentally from that of serial machine and multiprocessor programs. Other results demonstrate the limitations of increasing the datapath width and inter PE communication bandwidth without corresponding improvements in other processor features.
From a robot vision point of view, none of the individual visual operators presented features much originality. However they all make up a complete perceptual task, allowing a miniature vehicle to follow a specific textured target. Furthermore, most of them are performed inside a programmable artificial retina that fits, together with its controller, within a few cm/sup 3/ and make the vehicle a miniature autonomous one. After a presentation of the scientific motivations, we describe the whole process and its implementation and we take the opportunity through the elementary steps to illustrate the generic issues of such retina based vision systems.
The exploitation of analog VLSI techniques combined with computer vision knowledge offers spectacular possibilities. Limitations of current VLSI technologies do not allow to create sensors with extremely complex pixel architecture, but the coupling of external CMOS analog processing units is a great solution for rapid low level segmentation processes. This paper presents a novel sensing approach where photo-transduction, multiresolution feature extraction, scale-space integration, and edge tracking combined with sub-pixel interpolation are performed on a mixed-signal (digital-analog) VLSI architecture. The paper also discusses how we implement the curvature primal sketch into the system for higher level scene representation. The main sensory part of this integrated image acquisition system is a CMOS sensor called Multiport Access photo-Receptor (MAR). VLSI also provides means to integrate analog computing, digital controller, and DSP co-processor modules which define a powerful sensory chip set for focal plane image processing. A current version of the MAR sensor which implements 256/spl times/256 pixels includes 16 analog spatial filters which simultaneously compute multiresolution edge maps. This novel smart image sensor approach with associated low level segmentation capability presents good opportunities for real time automated process for the particular case of unstructured environment.
An automated inspection method is sketched and its parallel implementation on a MIMD multiprocessor is discussed. The method is based on a segmentation of gray level views of the test object and on the extraction and measurement of meaningful blobs from the segmented images. Segmentation is performed processing the histogram of pixel intensities; blob extraction is based on the application of the distance transformation to interest regions. The method is able to detect shape defects affecting a type of moulded plastic drippers; typical defects revealed are: incompleteness, excess of moulded material along the joint lines, incompleteness of a labyrinth-like mask moulded on the dripper surface. The algorithm consists of several steps that can be easily implemented on a MIMD architecture using both a shared or a distributed memory approach. The implementation on the hierarchical shared memory ViP multiprocessor is discussed; drippers can be efficiently analyzed on this kind of system, configured with 3 clusters of 4 processors each, at the production rate of two per second, making possible single piece quality certification.
Today, in the digitized satellite image domain, the need for high-dimension images is increasing considerably. To transmit or to store such images (more than 6000/spl times/6000 pixels), we need to reduce their data volume, and so we have to use image compression techniques. In most cases, these operations have to be processed in real time. The large amount of computations required by classical image compression algorithms prohibits the use of common sequential processors. To solve this problem, CEA (in collaboration with CNES) has tried to define the best-suited architecture for image compression. In order to achieve this aim, we developed and evaluated a new parallel image compression algorithm for general-purpose parallel computers using data-parallelism. This paper presents this new parallel image compression algorithm. We present implementation results on several parallel computers. We also examine load balancing and data mapping problems. We end by defining optimal characteristics of the parallel machine for real-time image compression.
Low level image processing needs to apply simple operations to large sets of data. This processing has to be done quickly. Systolic processors enable this to be done. This study describes the design of a reconfigurable systolic processor for the iconic processing of images in real time (video rate). The computers of the processor (basic cells, delay elements, and interconnection network) are presented paying special attention to the basic cell. The basic operations used in the iconic processing of images are studied. The cell is designed so that it can carry out these operations. In the results section, the configurations of the processor for some preprocessing algorithms (convolution, median filtering, narrow edge extraction) are presented and the typical performance of the processor and the techniques used (segmentation and parallelism) are also shown.
Discusses the intention to provide robots with a computer system which gives them the capability to build an abstract and condensed description of their environment from low level data provided by an artificial vision system. This intention rapidly faced a complexity barrier: on one hand, the importance of the volume of information conveyed by an image they manipulate; on the other hand, the access to pertinent information to validate decision making. To reach this goal, research is oriented towards extracting a set of computing tools learning in two ways: descriptive process of the objects to manipulate, and constructive process of the pertinent information for each operation to apply to an object.
The goal of the Digital Elevation Model is to generate an accurate three-dimensional scene using a stereo vision technique. In the stereo matching process two techniques are utilized, an area-based and feature-based to generate a disparity map. In our application we use an area-based approach coupled with the prediction validation techniques. The computation of the Digital Elevation Model (DEM) is based on correlation, also called matching, to determine the pixel correspondence in a pair of stereo spatial images. It is a fundamental step in digital mapping. The French "Institut Geographique National" (IGN) has developed a system to provide DEM. The kernel of this system is based on an incremental correlation method, which is the bottleneck in the map production because of its expenditure of computing time. In the same way the CEA-LETI has developed in collaboration with the IRIT laboratory (Toulouse University), a SIMD calculator SYMPATI2 dedicated to image processing, and integrated in the OPENVISION real-time system. The IGN DEM of SPOT images (6000/spl times/6000) takes 20 hours using a Sparc 10 workstation. In order to reduce this computation time we studied the parallelization and the implementation the IGN algorithm on OPENVISION.
The paper describes the architecture of a SIMD massively parallel computer, based on a model called associative mesh. Associative mesh is a reconfigurable interconnection whose basic primitives combine efficiently communications and computations, thanks to an asynchronous scheme. We present the physical implementation of an associative network, with 8-connected 2D mesh topology specifically designed for applications in the computer vision domain. Some results of simulations and evaluations are shown.
The flow of visual input reaching the eye consists of huge amounts of time-varying information. It is crucial for both biological vision and automated systems to perceive and comprehend such a constantly changing environment within a relatively short processing time. To cope with such a computational challenge, one should locate and analyze only the information relevant to the current task by quickly focusing on selected areas of the scene as needed. Attention makes perception computationally tractable and helps with tasks such as object recognition. Attention permeates the whole stream of visual computation, it is both hierarchical and modular, and it involves representations, processing and strategies. Attentional mechanisms are intimately related to adaptation processes, and high-level attention corresponds to competitive, functional and learned behavioral programs. Attention consists of both data- and model-driven processes and their relationships, and it covers several levels such as sensory, reactive and behavioral processes. An example of how attention can be implemented considers time-varying imagery and it shows how functional linked pyramids and zoom lens operations lead to the generation of visual saccades. Both the time-varying imagery and the corresponding recognition memory are organized as pyramids and uniform indexing and classification interfaces using an attention pyramid are established. This paper concludes with a discussion on promising venues for future research that are most likely to enhance our understanding of attentional mechanisms.
The paper gives a brief overview of the hardware design of RTA/1 (with 1024 PEs), a parallel machine designed based on the Recursive Torus Architecture (T. Matsuyama et al., 1993). We developed a small scale prototype machine with 16 PEs, RTA/0, to evaluate its performance. Then, we propose a scheme of data level parallel processing on RTA/1 and demonstrate its utilities by implementing complex parallel processes for bottom up object recognition on RTA/0
Few problems in computer vision have been investigated more vigorously than stereo. Nevertheless, the main obstacle on the way to their practical application is the excessively long computation time needed to match stereo images. This paper presents parallel algorithms for edge-based stereo that are suitable for depth computation. Edge-based stereo techniques produce only sparse depth maps; thus we present, in addition, an efficient parallel algorithm for dense stereo matching that can be employed in scene reconstruction. Both approaches are implemented on several different computers to measure their performance. We compared single-processor and multiple-processor implementations to evaluate the profit of parallel realizations. We show that both approaches are very suitable for parallel implementations and that the computing time can be considerably reduced with parallel implementations. Furthermore, we present the results that are obtained when employing the different approaches to stereo images.
A cooperation between the Microcomputer Laboratory and the Artificial Vision Laboratory of the Informatics and Systemics Department of the University of Pavia led to a design project for a control system for active vision, named PAVIA (Project for an Active VIsion Apparatus). The system allows one to control the movements and the independent positioning of two cameras. A workstation network processes the data received from the cameras and consequentially sends commands to drive four step motors to implement independent 4-axis movements. Furthermore, it is possible to control the shot parameters.
Ranganathan合作论文数Department of Computer Science and Engineering;University of South Florida1