Segmenting an image into meaningful regions is an important step in many computer vision applications such as facial recognition, target tracking and medical image analysis. Because image segmentation is an ill-posed problem, parameters are needed to constrain the solution to one that is suitable for a given application. For a user, setting parameter values is often unintuitive. We present a method for automating segmentation parameter selection using an efficient search method to optimize a segmentation objective function. Efficiency is improved by utilizing prior knowledge about the relationship between a segmentation parameter and the objective function terms. An adaptive sampling of the search space is created which focuses on areas that are more likely to contain a minimum. When compared to parameter optimization approaches based on genetic algorithm, Tabu search, and multi-locus hill climbing the proposed method was able to achieve equivalent optimization results with an average of 25% fewer objective function evaluations.
A fundamental limitation of hyperspectral imaging is the inter-band misalignment correlated with subject motion during data acquisition. One way of resolving this problem is to assess the alignment quality of hyperspectral image cubes derived from the state-of-the-art alignment methods. In this paper, we present an automatic selection framework for the optimal alignment method to improve the performance of face recognition. Specifically, we develop two qualitative prediction models based on: 1) a principal curvature map for evaluating the similarity index between sequential target bands and a reference band in the hyperspectral image cube as a full-reference metric; and 2) the cumulative probability of target colors in the HSV color space for evaluating the alignment index of a single sRGB image rendered using all of the bands of the hyperspectral image cube as a no-reference metric. We verify the efficacy of the proposed metrics on a new large-scale database, demonstrating a higher prediction accuracy in determining improved alignment compared to two full-reference and five no-reference image quality metrics. We also validate the ability of the proposed framework to improve hyperspectral face recognition.
Image blind deconvolution is well known as a challenging, ill-posed problem due to the uncertainty of the blur kernel and the noise condition. Based on our observations, blind deconvolution algorithms tend to generate disconnected and noisy blur kernels, which would yield a serious ringing effect in the restored image if the input image is noisy. Therefore, there is still room for further improvement, especially for noisy images captured under poor illumination conditions. In this paper, we propose a robust blind deconvolution algorithm by adopting a penalty-weighted anisotropic diffusion prior. On one hand, the anisotropic diffusion prior effectively eliminates the discontinuity in the blur kernel caused by the noisy input image during the process of kernel estimation. On the other hand, the weighted penalizer reduces the speckle noise of the blur kernel, thus improving the quality of the restored image. The effectiveness of the proposed algorithm is verified by both synthetic and real images with defocused or motion blur. (C) 2016 Elsevier Inc. All rights reserved.
Spectral imaging typically generates a large amount of high-dimensional data that are acquired in different sub-bands for each spatial location of interest. The high dimensionality of spectral data imposes limitations on numerical analysis. As such, there is an emerging demand for robust data compression techniques with loss of less relevant information to manage real spectral data. In this paper, we describe a reduced-order data modeling technique based on local proper orthogonal decomposition (POD) in order to compute low-dimensional models by projecting high-dimensional clusters onto subspaces spanned by local reduced-order bases. We refer to the proposed method as the local-based approach because POD finds locally optimal solutions on each group split by k-means clustering. Experimental results are reported on three public domain databases and an in-house database. Comparisons with three leading spectral recovery techniques, three decomposition techniques used for hyperspectral imaging, and two baseline techniques show that the proposed method leads to promising improvement on spectral and colorimetric accuracy corresponding to the reconstructed spectral reflectance.
Most existing camera placement algorithms focus on coverage and/or visibility analysis, which ensures that the object of interest is visible in the camera's field of view (FOV). However, visibility is inadequate for continuous and automated tracking. In such applications, a sufficient overlap between FOVs should be secured so that camera handoff can be executed successfully and automatically before the object of interest becomes untraceable or unidentifiable. In addition, most of the existing algorithms employ discrete solution space, which suffers from limited solution accuracy and high computational complexity due to the high dimension of sampled solution space. In this paper, we propose to perform the optimisation of the camera parameters in a continuous solution space. In addition, to incorporate the optimisation of coverage and sufficient overlapped FOVs, a weighted sum approach is utilised to translate a multiple objective optimisation problem into a single objective optimisation process in our previous work. Significantly improved handoff success rates are illustrated via experiments using typical office floor plans in comparison with Erdem and Sclaroff's method. Improved accuracy, enhanced robustness, completeness of the solution set, and reduced computational complexity are accomplished in comparison with our previous algorithms.
We study face recognition in unconstrained illumination conditions. A twofold contribution is proposed: First, the robustness of four state-of-the-art algorithms, namely Multi-block Local Binary Pattern (MBLBP), Histogram of Gabor Phase Patterns (HGPP), Local Gabor Binary Pattern Histogram Sequence (LGBPHS) and Patterns of Oriented Edge Magnitudes (POEM-WPCA) against high illumination variation is studied. Second, we propose to enhance the performance of the four mentioned algorithms, which has been drastically decreased upon the day lighted face images provided by IRIS-M3 face database. For this purpose, we use visible narrow band subspectral images selected from the mentioned database. We formulate best spectral bands selection as a pursuit optimization problem wherein the vector of weights determining the importance of each visible spectral band is supposed to be sparse, and hence can be determined by minimizing its L1-norm. Several fusing approaches are then applied on selected best spectral bands using multi-scale and multi-orientation Gabor wavelets. The results highlight further the still challenging problem of face recognition in conditions with high illumination variation, as well as the effectiveness of our subspectral images based approach with its two components; bands selection and bands fusion, to increase the accuracy of the studied algorithms by at least 14 % upon the proposed database.
Bi-modal image processing can be defined as a series of steps taken to enhance a target image with a guidance image. This is done by using exploitable information derived from acquiring two images of the same scene with different image modalities. However, while the potential benefit of bi-modal image processing may be significant, there is an inherent risk; if noise or defects in the guidance image are allowed to transfer to the target image, the target image could become corrupted rather than improved. In this chapter, we present a new method to enhance a noisy depth map from its color information via the joint bilateral filter (JBF) based on common distance transform (CDT). This method is composed of two main steps: CDT map generation and CDT-based JBF. In the first step, a CDT map is generated that represents the degree of pixel-modal similarity between a depth pixel and its corresponding color pixel. Then, based on the CDT map, JBF is carried out in order to enhance depth information with the aid of color information. Experimental results show that CDT-based JBF outperforms other conventional methods objectively and subjectively in terms of noise reduction, as well as inherent visual artifacts suppression.
Surveillance and inspection have an important role in security and industry applications and are often carried out with line-scan cameras. The advantages of line-scan cameras include hyper-resolution (larger than 50 Megapixels), continuous image generation, and low cost, to mention a few. However, due to the physical separation of line CCD sensors for the red (R), green (G), and blue (B) color channels, the color images acquired by multi-line CCD cameras intrinsically exhibit a color misalignment defect, such that the edges of objects in the scene are separated by a certain number of pixels in the R, G, B color planes in the scan direction. This defect, if not corrected properly, can severely degrade the quality of multi-line CCD images and hence impairs the functionality of the cameras. Current techniques for correcting such color misalignments are typically not fully automated, which is undesirable in applications such as inspection and surveillance that depend on fast unmanned responses. This chapter introduces an algorithm to automatically correct the color misalignments in multi-line CCD images for rotational scans as well as for translational scans. Results are presented for two different configurations of multi-line CCD imaging systems: (a) a close-range multi-line CCD imaging system for inspection applications and (b) a long-range imaging system for surveillance applications. Experimental results show that the two imaging systems are able to acquire hyper-resolution images and the color misalignment correction algorithm can automatically and accurately correct those images for their respective applications.
Disparity maps, occlusions, and occlusion filling results on the Middlebury College test images Map, Venus, and Tsukuba (from top to bottom). Occlusions are filled using SLS linear interpolation model.
The relationship between the bilateral kernel function and the recently proposed locally adaptive regression kernel is examined. Despite the difference in implementation, both locally adaptive approaches are designed to prevent averaging across edges while smoothing an image. Their similarity suggests that they can reasonably be linked although both filtering approaches have grown to become well-established theories in their fields. First, the locally adaptive regression kernel is analysed theoretically. Then, the connection between the methods is explored by applying the spectral distance measure to the bilateral kernel. Finally, a direct relation is established between the bilateral kernel and the locally adaptive regression kernel.
In this paper, we investigate face recognition in unconstrained illumination conditions. A twofold contribution is proposed: First, three state of the art algorithms, namely Multiblock Local Binary Pattern (MBLBP), Histogram of Gabor Phase Patterns (HGPP) and Local Gabor Binary Pattern Histogram Sequence (LGBPHS) are challenged against the IRIS-M3 multispectral face data base to evaluate their robustness against high illumination variation. Second, we propose to enhance the Performance of the three mentioned algorithms, which has been drastically decreased because of the non-monotonic illumination variation that distinguishes the IRIS-M3 face database. Instead of the usual braod band images, we use narrow band sub spectral images selected from the visible spectrum. Selection of best spectral bands is formulated as a pursuit optimization problem wherein the vector of weights determining the importance of each visible spectral band is supposed to be sparse, and hence can be determined by minimizing its L1-norm. The results highlight further the still challenging problem of face recognition in conditions with high illumination variation, as well as the effectiveness of our sub spectral images based approach to increase the accuracy of the studied algorithms by at least 14% upon the proposed database.
Due to increasing security concerns, a complete security system should consist of two major components, a computer-based face-recognition system and a real-time automated video surveillance system. A computerbased face-recognition system can be used in gate access control for identity authentication. In recent studies, multispectral imaging and fusion of multispectral narrow-band images in the visible spectrum have been employed and proven to enhance the recognition performance over conventional broad-band images, especially when the illumination changes. Thus, we present an automated method that specifies the optimal spectral ranges under the given illumination. Experimental results verify the consistent performance of our algorithm via the observation that an identical set of spectral band images is selected under all tested conditions. Our discovery can be practically used for a new customized sensor design associated with given illuminations for an improved face recognition performance over conventional broad-band images. In addition, once a person is authorized to enter a restricted area, we still need to continuously monitor his/her activities for the sake of security. Because pantilt-zoom (PTZ) cameras are capable of covering a panoramic area and maintaining high resolution imagery for real-time behavior understanding, researches in automated surveillance systems with multiple PTZ cameras have become increasingly important. Most existing algorithms require the prior knowledge of intrinsic parameters of the PTZ camera to infer the relative positioning and orientation among multiple PTZ cameras. To overcome this limitation, we propose a novel mapping algorithm that derives the relative positioning and orientation between two PTZ cameras based on a unified polynomial model. This reduces the dependence on the knowledge of intrinsic parameters of PTZ camera and relative positions. Experimental results demonstrate that our proposed algorithm presents substantially reduced computational complexity and improved flexibility at the cost of slightly decreased pixel accuracy as compared to Chen and Wang’s method [18].
In this paper, we present a new method for a locally adaptive region detector called Bilateral kernel-based Region Detector (BIRD). This work is to detect stable regions from images by consecutively computing a multiscale decomposition based on the bilateral kernel. The BIRD regards a region as covariant if it exhibits predictability in its photometric distance over spatial distance. Distinctiveness and robustness across scales are achieved by selecting the extremely stable regions through sequential scales. Our method is simple and easy to implement. Experimental results show that our method outperforms competing affine region detection methods in efficiency on region detection.
This chapter contains sections titled: Introduction Matching Stereo Images Conclusions Acknowledgments References Further Reading
In this paper, we propose a novel outdoor scene image segmentation algorithm based on background recognition and perceptual organization. We recognize the background objects such as the sky, the ground, and vegetation based on the color and texture information. For the structurally challenging objects, which usually consist of multiple constituent parts, we developed a perceptual organization model that can capture the nonaccidental structural relationships among the constituent parts of the structured objects and, hence, group them together accordingly without depending on a priori knowledge of the specific objects. Our experimental results show that our proposed method outperformed two state-of-the-art image segmentation approaches on two challenging outdoor databases (Gould data set and Berkeley segmentation data set) and achieved accurate segmentation quality on various outdoor natural scene environments.
Most existing performance evaluation methods concentrate on defining various metrics over a wide range of conditions and generating standard benchmarking video sequences to examine the effectiveness of a video tracking system. It is a common practice to incorporate a robustness margin or factor into the system/algorithm design. However, these methods, deterministic approaches, often lead to overdesign, thus increasing costs, or underdesign, causing frequent system failures. In order to overcome the aforementioned limitations, we propose an alternative framework to analyze the physics of the failure process via the concept of reliability. In comparison with existing approaches where system performance is evaluated based on a given benchmarking sequence, the advantage of our proposed framework lies in that a unified and statistical index is used to evaluate the performance of an automated video surveillance system independent of input sequences. Meanwhile, based on our proposed framework, the uncertainty problem of a failure process caused by the system’s complexity, imprecise measurements of the relevant physical constants and variables, and the indeterminate nature of future events can be addressed accordingly.
In this paper, we present a method to calibrate and enhance depth information captured by an infrared (IR)-based time-of-flight video-plus-depth camera called “Kinect camera”. For depth data calibration, we use color and IR images of a chessboard with on-off halogen light sources to calculate camera parameters of the video and IR sensors in the Kinect camera. For depth data enhancement, we introduce weighted joint bilateral filtering based on distance transform of the color and depth images. Experimental results show that our method calibrates the video and depth sensors successfully and reduces the noise in captured depth images efficiently.