
The concept of colour and and multispectral image recognition connects all the topics we are considering. Colour (multispectral) image processing is investigated using an algebraic approach based on triplet numbers. In the algebraic approach, each image element is considered not as a 3D vector, but as a triplet number. The main goal of the paper is to show that triplet algebra can be used to solve colour image processing problems in a natural and effective manner. We propose novel methods for wavelet transforms and spline implementation in colour space.
We present an adapted nonlinear multiresolution decomposition (with some derivatives) of still images that permits perfect reconstruction. A hierarchical pyramidal decomposition with maximal decimation is used. For one level of decomposition, the input image I is partitioned into two subimages I/sub 1/ and I/sub 2/ obtained by downsampling. The subimage I/sub 1/ is unchanged. The subimage I/sub 2/ is replaced by the rounded output I/sub h/ of a 2D FIR filter whose coefficients have first been adapted to I and which has the entries I/sub 1/ and I/sub 2/. A similar processing is then applied one time to each of the subimages I/sub 1/ and I/sub 2/. The nonlinearity introduced by the rounding permits us to perfectly inverse the process, even when inverse filters (which are all ARMA) are not BIBO stable. Applications are given in lossless image coding, with possibility of embedded zerotree coding.
Detection and tracking of the lip contour is an important issue in lipreading. While there are solutions for lip tracking once a good contour initialization in the first frame is available, the problem of finding such a good initialization is not yet solved automatically, but done manually. Solutions based on edge detection and tracking have failed when applied to real world mouth images. In this paper, we propose a solution to lip contour detection that minimizes user interaction by requiring a minimal number of points to be marked manually on the mouth image. The proposed approach is based on edge detection using gradient masks and edge following. The method is based on the examination of gradient direction patterns in the lip area, and makes use of the local direction constancy along the lip contours, as opposed to the other regions of the mouth image that are characterized by random edge directions.
We discuss parallel algorithms for wavelet-based image and video coding. After reviewing fundamentals of the parallel discrete wavelet transform, we cover the parallelization of two state-of-the-art compression schemes: a C++ 3D SPIHT video codec and a Java JPEG-2000 implementation.
A novel approach to face recognition based on a multi-pose image sequence is presented in this paper. In this approach, faces are represented by their pattern vectors (projections to eigenfaces) in eigenspace. Instead of recognising a face from a single view, a sequence of images showing face movement (from left to the right profile) is used for recognition. Pattern vectors corresponding to multiple poses build a trajectory in eigenspace where each trajectory belongs to one face sequence (profile to profile). In the training phase, sequences of poses construct prototype trajectories, and in recognition phase, an unknown face trajectory is compared with prototypes. New matching models are presented and analysed as well as the influence of some parameters on the recognition ratio.
Because of their abundance and simplicity, planes are used in several computer vision tasks. Their simplicity results in that, under perspective projection, the transformation between a world plane and its corresponding image plane is projective linear, or a homography. These relations also hold between perspective views of a plane in different images. This paper proposes an algorithm that detects planar homographies in uncalibrated image pairs. It then demonstrates how this plane identification method can be used as a first step in an image analysis process, when point matching between images is unreliable. The detection is performed using a RANSAC scheme based on the linear computation of the homography matrix elements using four points. Results are shown on real image pairs.
A classical computer does not allow to calculate a discrete cosine transform on N points in less than linear time. This trivial lower bound is no longer valid for a computer that takes advantage of quantum mechanical superposition, entanglement, and interference principles. In fact, we show that it is possible to realize the discrete cosine transforms and the discrete sine transforms of size NxN and types I,II,III, and IV with as little as O(log^2 N) operations on a quantum computer, whereas the known fast algorithms on a classical computer need O(N log N) operations.
The fractional Fourier transform (FRFT) is one-parametric generalization of the classical Fourier transform. FRFT was introduced in the 1980s and found a lot of applications in signal processing. The time and spectral domains are both special cases of the fractional Fourier domain. They correspond to the Oth and 1st functional Fourier domains, respectively. We introduce the classical and quantum fractional Walsh transforms and develop corresponding fast algorithms.
Support vector machine is a special kind of learning machine, proposed by Vapnik. The learning capability of support vector machines depends on the Vapnik-Chervonenkis (VC) dimension of the kernel function used. In this paper, we construct a new kernel function for support vector machine, which is based on Walsh functions. We prove some theoretical results related to the VC-dimension of the support vector machines which are built in the space of the Walsh functions. First experimental results for face detection are reported.
SIREV (Sector Imaging Radar for Enhanced Vision) is an innovative radar system which can supply high-quality images of a sector in front of an aircraft under almost all weather conditions. In contrast to synthetic aperture radar (SAR) systems, no relative motion between sensor and targets is required in SIREV. The high frame repetition frequency in SIREV allows us to detect even very rapid changes in the imaged scenes. Besides a map of the Earth's surface, the complex-valued radar images can be further processed to supply additional information about the topography and objects in the field of view. A fast processing algorithm has been developed for SIREV which allows a very accurate, phase-preserving and efficient image formation. Raw data were obtained by a first demonstration flight using a helicopter, as a platform and the data processing results show good agreement with the theory. Motion errors of the platform could be extracted from the range compressed raw data avoiding the need of an inertial navigation system. After the image formation, coherent and incoherent image averaging processes have been applied to improve the image quality. The evaluation of the computational effort shows that a real-time hardware realization can be carried out using off-the-shelf digital components.
Third-order low-pass filters are analyzed. Two different configurations of general impedance converter (GIC) based filters are compared to two Sallen and Key (SAK) based structures. For GIC-based third-order filters, realization procedures are given. The transfer functions and the component values for the third-order filters with Butterworth response are presented. Sensitivity analysis is done, and the results of Monte Carlo runs, as well as Schoeffler sensitivities are shown. The best results are obtained with the GIC-based filter which uses four capacitances for the realization
In this article we present a panoramic depth imaging system. The system is mosaic-based which means that we use a single rotating camera and assemble the captured images in a mosaic. Due to a setoff of the camera's optical center from the rotational center of the system we are able to capture the motion parallax effect which enables the stereo reconstruction. The camera is rotating on a circular path with the step defined by an angle, equivalent to one column of the captured image. The equation for depth estimation can be easily extracted from system geometry. To find the corresponding points on a stereo pair of panoramic images the epipolar geometry needs to be determined. It can be shown that the epipolar geometry is very simple if we are doing the reconstruction based on a symmetric pair of stereo panoramic images. We get a symmetric pair of stereo panoramic images when we take symmetric columns on the left and on the right side from the captured image center column. Epipolar lines of the symmetrical pair of panoramic images are image rows. We focused mainly on the system analysis. Results of the stereo reconstruction procedure and quality evaluation of generated depth images are quite promissing. The system performs well in the reconstruction of small indoor spaces. Our finall goal is to develop a system for automatic navigation of a mobile robot in a room.
Systems of finite order with minimum time-bandwidth products are considered. The time and frequency response spreads are defined by moments of the higher order in both domains. Minimizing the product of the moments, causal systems with the largest energy concentration in the time and frequency domain are obtained. The optimization is carried out for transfer functions up to the eighth order with two and four complex zeros. The moment orders are chosen from two to eight. The influence of the zeros and the moment order on system properties are considered. The complete data suitable for filter design are given.
In this paper, we propose a technique for 3D motion estimation of the left ventricle from an image sequence of a beating human heart. Accurate motion estimation of the cardiac wall has been shown to be very important in studying coronary diseases. The proposed technique requires initial 3D segmentation of the left ventricle area obtained for each time frame during the cardiac cycle. Characteristic points at its surface are detected, and matched in two consecutive frames by matching shape properties. Optical flow is computed from the sequence of images using the gradient-based Horn-Schunck method, additionally constrained with motion estimates for characteristic surface points. Our work demonstrates the application of the Horn-Schunck optical flow algorithm for 3D cardiac motion estimation, and proposes to improve the accuracy of estimation by introducing constraints obtained by a shape-based matching method.
This paper presents a method for automatic segmentation and labeling of computerised tomography (CT) head images of stroke lesions. The method is composed of three steps. The first step is automatic determination of head symmetry axis, with the possibility of manual improvement of the result if necessary. Symmetry axis calculation is based on moments. In the second step, the seeded region-growing (SRG) algorithm is used to segment the input image into a number of regions having uniform brightness. Features of these regions, such as brightness, area, neighborhood and relative position to the symmetry axis are used to create facts for a rule-based expert system. Based on created facts and pre-defined rules as input, the rule-based expert system is used in the third step to label regions as background, skull, gray/white matter, CSF and stroke. Experimental results have been conducted and have demonstrated the feasibility and accuracy of the proposed method.
The paper presents a number of image processing and pattern recognition applications using coordinate logic filters which execute coordinate logic operations among the pixels of the image. These filters are very efficient in various 1D, 2D or higher-dimensional digital signal processing applications, such as noise removal, magnification, opening, closing, skeletonization, coding, edge detection, feature extraction, and fractal modeling. The key issue in the coordinate logic analysis of images is the method of fast successive filtering and managing of the residues. The desired processing is achieved by executing only direct logic operations among the pixels of the given image. Coordinate logic filters can be easily and quickly implemented using logic circuits or cellular automata; this is their primary advantage.
An object-based wavelet-based video codec has been designed and implemented. The codec takes a colour video sequence of arbitrary size as input and performs intra-frame object-based compression on the sequence. Video sequences are segmented using an optic flow-based algorithm, before being transformed with a symmetrical wavelet transform, prior to Lloyd-Max quantisation and entropy encoding. A video codec quality analyser (CQA) is used to assess the subjective quality of the coded stream with respect to the uncoded one in order to provide a single quality measure that correlates with a subjective assessment of the data. Experiments using standard test sequences show that a higher compression ratio is achieved than obtainable using a similar non-object-based codec, at the same quality level.
It is observed that by designing a filter with marginally stricter specifications than the desired one without any increase in order, i.e., length of the filter, a multiplierless implementation of recursive filters is feasible utilizing some class of low sensitivity structure. These implementations are not associated with an increase in the order of the filter that involves more shift registers, data paths, control circuits, etc., and, hence, an increase in complexity, i.e. indirect overheads. The approach appears to be especially suitable for filters with high sensitivity. In low sensitivity structures the modified coefficients can be realized with multipliers of shorter wordlength, i.e., in fewer bits. When these are implemented in minimum numbers of signed powers of two (MNSPT) form, we have a multiplierless implementation.
The development of an airborne remote sensing system at the Faculty of Transport and Traffic Engineering started in 1998, as support to postgraduate study. The system is intended for the airborne acquisition of data on urban traffic dynamic characteristics, for the acquisition of data on the bottleneck phenomenon in road traffic, for determining the status of vegetation and forests, and for sensing of mine polluted areas. The aircraft Cessna 172R is used as the platform for the gimbal on which more, mutually different sensors can be placed. The characteristics of the Sharp View Cam VL-H420S as the sensor are considered here. The static spatial resolution of the camcorder is determined by means of test bars. Further, the effects of platform disturbance and vibration on the aerial images are considered