
In this paper, a novel approach to the impulsive noise removal in color images is presented. The proposed technique employs the switching scheme based on the impulse detection mechanism using the so-called peer group concept. Compared to the vector median filter and other commonly used multichannel filters, the proposed technique consistently yields very good results in suppressing both the random and fixed-valued impulsive noise. The main advantage of the proposed noise detection framework is its enormous computational speed, which enables efficient filtering of color images in real-time applications.
This paper presents a description of the ''Pairs Of Lines'' object recognition algorithm used in the MIMAS Computer Vision toolkit. This toolkit was developed at Sheffield Hallam University and the ''Pairs of Lines'' method was used in a recent European Union funded project. The algorithm was developed to enable a micro-robot system (Amavasai et al., InstMC Journal of Measurement and Control) to recognize geometric planar objects in real-time, in a noisy environment. The method involves using straight line segments which are extracted from both the known object models and from the visual scene that the objects are to be located in. Pairs of these straight lines are then compared. If there is a geometric match between the two pairs an estimate of the possible position, orientation and scale of the model in the scene is made. The estimates are collated, as all possible pairs of lines are compared. The process yields the position, orientation and scale of the known models in the scene. The algorithm has been optimized for speed. This paper describes the method in detail and presents experimental results which indicate that the technique exhibits robustness to camera noise and partial occlusion and produces recognition in times under 1s on a desktop PC. Recognition times are shown to be from 2 to 16 times faster than with the well-studied pairwise geometric histograms method. Recognition rates of up to 80% were achieved with scenes having signal to noise ratios of 2.5.
In today's fast-paced information-driven society, the need for accurate, timely, and cost-effective data collection is very critical. Optical mark reader (OMR) systems can be used to achieve these aspects. This paper describes the development of a low-cost and high-speed OMR system prototype for marking multiple-choice questions. The novelty of this approach is the implementation of the complete system into a single low-cost Field Programmable Gate Array (FPGA) to achieve the high processing speed. Effective mark detection and verification algorithms have been developed and implemented to achieve real-time performance at low computational cost. The OMR is capable of processing a high-resolution CCD linear sensor with 3456pixels at 5000frame/s at the effective maximum clock rate of the sensor of 20MHz (4x5MHz). The performance of the prototype system is tested for different marker colours and marking methods. At the end of the paper the proposed OMR system is compared with commercially available systems and the pro and cons are discussed.
MPEG-4 introduces the concept of video object to support content-based functionalities. Video object segmentation is a crucial step for object-based coding and manipulation. In this paper, a robust semi- automatic video object segmentation scheme is proposed. To efficiently and accurately define the initial object contour, modified intelligent scissors is proposed on the basis of original intelligent scissors. It can improve about 6–8 times the processing speed with only slight sacrifice of accuracy, which just meets the requirements of initial object extraction for semi-automatic approach. To avoid errors accumulating and propagating during object tracking, an adaptive frame skipping scheme is proposed to decompose video sequence into video clips. For rigid and non-rigid video objects, two different image segmentation algorithms are utilized, and then region-based backward projection technique is adopted to interpolate the video object plane (VOPs) of other frames within every video clip. The proposed approach can cope with occlusion/disocclusion problem to most extent. Experimental results demonstrate the effectiveness and robustness of the method.
In the textile industry, scoured wool contains different types of foreign materials (contaminants) that need to be separated out before it goes into further processing, so that the textile machines are protected from damage and the quality of the final woollen products is ensured. This paper presents an automated visual inspection (AVI) system for detecting and sorting contaminants from wool in real time. The techniques were first developed in the lab and subsequently applied to a large-scale factory system. The combinative use of image processing algorithms in RGB and HSV colour spaces can segment 96% of contaminant types (minimum size around 4 cm long and 5 mm in diameter) in real-time on the lab test rig. One of the most important aspects of the system is to use the non-linear colour space transformation and merge the threshold algorithm in HSV colour space into the image processing algorithms in RGB colour space to enhance the contaminant identification in real time. The real-time capability of the system is also analysed in detail. The experimental results demonstrate that the factory AVI system could identify and remove the contaminants at a camera speed of around 800 lines/s and the conveyor speed of 20 m/min in real time.
Recently, Shen et al. [IEEE Transactions on Image Processing 2003;12:283–95] presented an efficient adaptive vector quantization (AVQ) algorithm and their proposed AVQ algorithm has a better peak signal-to-noise ratio (PSNR) than that of the previous benchmark AVQ algorithm. This paper presents an improved AVQ algorithm based on the proposed hybrid codebook data structure which consists of three codebooks—the locality codebook, the static codebook, and the history codebook. Due to easy maintenance advantage, the proposed AVQ algorithm leads to a considerable computation-saving effect while preserving the similar PSNR performance as in the previous AVQ algorithm by Shen et al. [IEEE Transactions on Image Processing 2003;12:283–95]. Experimental results show that the proposed AVQ algorithm over the previous AVQ algorithm has about 75% encoding time improvement ratio while both algorithms have the similar PSNR performance.
A new adaptive vector filter is proposed for impulse noise suppression and its relationship with the recent impulse reduction filters is investigated. The new filter detects outliers presented in the image through a novel neighborhood evaluation process, which significantly improves the accuracy of noise detection and detail preservation. The computational complexity of the new filter is very competitive. Its two parameters can be configured efficiently using online/offline optimization processes. Extensive simulations indicated that the new filter outperforms other prior-art methods in suppressing impulse noise in natural color images.
Computer-assisted vision plays an important role in our society, in various fields such as personal and goods safety, industrial production, telecommunications, robotics, etc. However, technical developments are still rare and slowed down by various factors linked to sensor cost, lack of system flexibility, difficulty of rapidly developing complex and robust applications, and lack of interaction among these systems themselves, or with their environment. This paper describes our proposal for a smart camera with real-time video processing capabilities. A CMOS sensor, processor and, reconfigurable unit associated in the same chip will allow scalability, flexibility, and high performance.
For robust real-time transmission of scalable image and video data over packet-loss networks, a commonly used approach is FEC-based multiple description coding, which protects a scalable bitstream with a fixed number of packets of equal length. In this paper, by considering the problem of applying this approach of multicasting a source to a collection of clients with heterogeneous bandwidths, we propose a novel technique that changes packet length but fixes packet number. In one scenario where different clients access the server via separate links, we study the sensitivity of an optimal solution to the change of packet length, and propose a local search procedure which refines the already computed optimal solutions for other bandwidths. Compared to the scheme that computes an optimal solution individually for each bandwidth, this procedure can achieve comparable performance, however with quite lower time complexity, thus can be used in real-time applications. In another scenario where many clients share a bottleneck link, we present an embedded packetization framework for layered multiple description coding, in which even simple methods with low complexities can achieve good performance. We also propose a local search algorithm to optimize the weighted average performance in case of two layers, and a fast heuristic algorithm which can achieve very good performance tradeoff among all clients in case of more than two layers.
The purpose of this paper is to investigate a real-time system to detect context-independent events in video shots. We test the system in video surveillance environments with a fixed camera. We assume that objects have been segmented (not necessarily perfectly) and reason with their low-level features, such as shape, and mid-level features, such as trajectory, to infer events related to moving objects. Our goal is to detect generic events, i.e., events that are independent of the context of where or how they occur. Events are detected based on a formal definition of these and on approximate but efficient world models. This is done by continually monitoring changes and behavior of features of video objects. When certain conditions are met, events are detected. We classify events into four types: primitive, action, interaction, and composite. Our system includes three interacting video processing layers: enhancement to estimate and reduce additive noise, analysis to segment and track video objects, and interpretation to detect context-independent events. The contributions in this paper are the interpretation of spatio-temporal object features to detect context-independent events in real time, the adaptation to noise, and the correction and compensation of low-level processing errors at higher layers where more information is available. The effectiveness and real-time response of our system are demonstrated by extensive experimentation on indoor and outdoor video shots in the presence of multi-object occlusion, different noise levels, and coding artifacts.
In this paper, a novel real-time 3D and color sensor for the mid-distance range (0.1-3m) based on color-encoded structured light is presented. The sensor is integrated using low-cost of-the-shelf components and allows the combination of 2D and 3D image processing algorithms, since it provides a 2D color image of the scene in addition to the range data. Its design is focused on enabling the system to operate reliably in real-world scenarios, i.e. in uncontrolled environments and with arbitrary scenes. To that end, novel approaches for encoding and recognizing the projected light are used, which make the system practically independent of intrinsic object colors and minimize the influence of the ambient light conditions. The system was designed to assist and complement a face authentication system integrating both 2D and 3D images. Depth information is used for robust face detection, localization and 3D pose estimation. To cope with illumination and pose variations, 3D information is used for the normalization of the input images. The performance and robustness of the proposed system is tested on a face database recorded in conditions similar to those encountered in real-world applications.
This paper presents a new progressive refinement algorithm for full spectral rendering. This algorithm adopts wavelet transformation to efficient represent full spectral data. To our knowledge, this is the first approach to employing such a transformation for progressive, full spectral rendering, where the radiance calculation through multiplications of two spectral functions is computed under a wavelet basis. We implemented the proposed technique for Monte Carlo direct lighting, and divide the rendering process into 9 stages (i=1–9), each of which employs the first leading 2i coefficients to produce progressive results. In the fourth progressive stage, our algorithm renders a spectral image that is 95% similar to the final non-progressive approach but only requires less than 70% of execution time. The quality of the rendered image is visually plausible being indistinguishable to those rendered by the non-progressive method. Our algorithm demonstrates features of fast convergence and high image fidelity. It is graceful, efficient, progressive, and flexible for full spectral rendering.
The objective of this research was to develop a ground-based real-time remote sensing system for detecting diseases in arable crops under field conditions and in an early stage of disease development, before it can visibly be detected. This was achieved through sensor fusion of hyper-spectral reflection information between 450 and 900 nm and fluorescence imaging. The work reported here used yellow rust (Puccinia striiformis) disease of winter wheat as a model system for testing the featured technologies. Hyper-spectral reflection images of healthy and infected plants were taken with an imaging spectrograph under field circumstances and ambient lighting conditions. Multi-spectral fluorescence images were taken simultaneously on the same plants using UV-blue excitation. Through comparison of the 550 and 690 nm fluorescence images, it was possible to detect disease presence. The fraction of pixels in one image, recognized as diseased, was set as the final fluorescence disease variable called the lesion index (LI). A spectral reflection method, based on only three wavebands, was developed that could discriminate disease from healthy with an overall error of about 11.3%. The method based on fluorescence was less accurate with an overall discrimination error of about 16.5%. However, fusing the measurements from the two approaches together allowed overall disease from healthy discrimination of 94.5% by using QDA. Data fusion was also performed using a Self-Organizing Map (SOM) neural network which decreased the overall classification error to 1%. The possible implementation of the SOM-based disease classifier for rapid retraining in the field is discussed. Further, the real-time aspects of the acquisition and processing of spectral and fluorescence images are discussed. With the proposed adaptations the multi-sensor fusion disease detection system can be applied in the real-time detection of plant disease in the field.
This paper presents a simple, fast coding technique for lossless compression of mosaic video data. The design of a video codec needs to strike a balance between the compression performance and the codec throughput. Aiming to make the encoding throughput high enough for real-time lossless video compression, we propose a hybrid scheme of inter and intraframe coding. Interframe predictive coding is invoked only when the motion between adjacent frames is modest and a simple motion compensation operation can significantly improve the compression performance. Otherwise, still frame compression is performed to keep the complexity low. Experimental results show that the proposed scheme achieves higher lossless video compression ratio than existing methods such as JPEG-LS and JPEG-2000.
SmartSpectra is a smart multispectral system for industrial, environmental, and commercial applications where the use of spectral information beyond the visible range is needed. The SmartSpectra system provides six spectral bands in the range 400–1000 nm. The bands are configurable in terms of central wavelength and bandwidth by using electronic tunable filters. SmartSpectra consists of a multispectral sensor and the software that controls the system and simplifies the acquisition process. A first prototype called Autonomous Tunable Filter System is already available. This paper describes the SmartSpectra system, demonstrates its performance in the estimation of chlorophyll in plant leaves, and discusses its implications in real-time applications.
We present a real-time algorithm for foreground-background segmentation.Sample background values at each pixel are quantized into codebooks which represent a compressed form of background model for a long image sequence.This allows us to capture structural background variation due to periodic-like motion over a long period of time under limited memory.The codebook representation is efficient in memory and speed compared with other background modeling techniques.Our method can handle scenes containing moving backgrounds or illumination variations, and it achieves robust detection for different types of videos.We compared our method with other multimode modeling techniques.In addition to the basic algorithm, two features improving the algorithm are presented-layered modeling/detection and adaptive codebook updating.For performance evaluation, we have applied perturbation detection rate analysis to four background subtraction algorithms and two videos of different types of scenes.
The 3-D visual tracking of human limbs is fundamental to a wide array of computer vision applications including gesture recognition, interactive entertainment, biomechanical analysis, vehicle driver monitoring, and electronic surveillance. The problem of limb tracking is complicated by issues of occlusion, depth ambiguities, rotational ambiguities, and high levels of noise caused by loose fitting clothing. We attempt to solve the 3-D limb tracking problem using only monocular imagery (a single 2-D video source) in largely unconstrained environments. The approach presented is a movement towards full real-time operating capabilities. The described system presents a complete visual tracking system which incorporates target detection, target model acquisition/initialization, and target tracking components into a single, cohesive, probabilistic framework. The presence of a target is detected, using visual cues alone, by recognition of an individual performing a simple pre-defined initialization cue. The physical dimensions of the limb are then learned probabilistically until a statistically stable model estimate has been found. The appearance of the limb is learned in a joint spatial-chromatic domain which incorporates normalized color data with spatial constraints in order to model complex target appearances. The target tracking is performed within a Monte Carlo particle filtering framework which is capable of maintaining multiple state-space hypotheses and propagating ambiguity until less ambiguous data is observed. Multiple image cues are combined within this framework in a principled Bayesian manner. The target detection and model acquisition components are able to perform at near real-time frame rates and are shown to accurately recognize the presence of a target and initialize a target model specific to that user. The target tracking component has demonstrated exceptional resilience to occlusion and temporary target disappearance and contains a natural mechanism for the trade-off between accuracy and speed. At this point, the target tracking component performs at sub real-time frame rates, although several methods to increase the effective operating speed are proposed.
A novel progressive image transmission scheme based on the quadtree segmentation technique is introduced in this paper. A 3-level quadtree is used in the quadtree segmentation technique to partition the original image into blocks of different sizes. Image blocks of different sizes are encoded by their block mean values. The relatively addressing technique is employed to cut down the storage cost of block mean values.In the proposed scheme, the number of image hierarchies can be adaptively selected according to the specific applications. By exploiting inter-pixel correlation and differently sized blocks for segmentation, the proposed scheme provides good image qualities at low bit rates and consumes very little computational cost in both image encoding and decoding procedures. It is quite suitable for real-time progressive image transmission.
Intelligent surveillance has become an important research issue due to the high cost and low efficiency of human supervisors, and machine intelligence is required to provide a solution for automated event detection. In this paper we describe a real-time system that has been used for detecting tailgating, an example of complex interactions and activities within a vehicle parking scenario, using an adaptive background learning algorithm and intelligence to overcome the problems of object masking, separation and occlusion. We also show how a generalized framework may be developed for the detection of other complex events.