Edge computing is becoming very popular. Many researchers are using edge devices for computing artificial intelligence, especially in real-time applications of computer vision by the use of Convolutional Neural Networks (CNNs). Edge devices can encounter non-ideal environments that impact the effectiveness of a CNN model's predictions. We propose a tool flow methodology for implementing a real-time intelligent image enhancement technique for deploying and validating CNNs on edge devices. In this work, we implement this proposed methodology in a modular and user-friendly way. We show that this methodology is capable of real-time intelligent image enhancement on various CNN models.
A real-time non-contact vital sign detection system is developed by utilizing neural network-based detection, multi-object tracking, and direction of arrival (DoA) techniques. The DoA produces a spatial-based image, which is fed into the detector. The detector is a convolutional neural network (CNN), which produces a list potential subject locations. These locations are propagated and associated via a tracking method called BYTE. All of these methods allow the system to accurately localize and track subjects as well as improve the robustness of vital sign estimation for stationary, multi-subject scenarios. We demonstrate that this real-time system produces low error rates of less than 1 and 3 BPM for breathing and heart rate estimations respectively in both single and multi-subject scenarios. All this is done while maintaining an average of 14 FPS on a portable Jetson Xavier NX.
Mapping vegetation species is critical to facilitate related quantitative assessment, and mapping invasive plants is important to enhance monitoring and management activities. Integrating high-resolution multispectral remote-sensing (RS) images and lidar (light detection and ranging) point clouds can provide robust features for vegetation mapping. However, using multiple sources of high-resolution RS data for vegetation mapping on a large spatial scale can be both computationally and sampling intensive. Here, we designed a two-step classification workflow to potentially decrease computational cost and sampling effort and to increase classification accuracy by integrating multispectral and lidar data in order to derive spectral, textural, and structural features for mapping target vegetation species. We used this workflow to classify kudzu, an aggressive invasive vine, in the entire Knox County (1362 km2) of Tennessee (U.S.). Object-based image analysis was conducted in the workflow. The first-step classification used 320 kudzu samples and extensive, coarsely labeled samples (based on national land cover) to generate an overprediction map of kudzu using random forest (RF). For the second step, 350 samples were randomly extracted from the overpredicted kudzu and labeled manually for the final prediction using RF and support vector machine (SVM). Computationally intensive features were only used for the second-step classification. SVM had constantly better accuracy than RF, and the producer’s accuracy, user’s accuracy, and Kappa for the SVM model on kudzu were 0.94, 0.96, and 0.90, respectively. SVM predicted 1010 kudzu patches covering 1.29 km2 in Knox County. We found the sample size of kudzu used for algorithm training impacted the accuracy and number of kudzu predicted. The proposed workflow could also improve sampling efficiency and specificity. Our workflow had much higher accuracy than the traditional method conducted in this research, and could be easily implemented to map kudzu in other regions as well as map other vegetation species.
Sparse representation has attracted extensive attention and performed well on image super-resolution (SR) in the last decade. However, many current image SR methods face the contradiction of detail recovery and artifact suppression. We propose a multi-resolution dictionary learning (MRDL) model to solve this contradiction, and give a fast single image SR method based on the MRDL model. To obtain the MRDL model, we first extract multi-scale patches by using our proposed adaptive patch partition method (APPM). The APPM divides images into patches of different sizes according to their detail richness. Then, the multi-resolution dictionary pairs, which contain structural primitives of various resolutions, can be trained from these multi-scale patches. Owing to the MRDL strategy, our SR algorithm not only recovers details well, with less jag and noise, but also significantly improves the computational efficiency. Experimental results validate that our algorithm performs better than other SR methods in evaluation metrics and visual perception.
The human visual system is not very sensitive to the absolute luminance of an image, but rather responds to local luminance changes, i.e., the gradient of an image. In this paper, we propose a constrained optimization approach for image gradient enhancement. The gradient strength of the enhanced image can be controlled directly using a target gradient strength parameter in the cost function. To suppress artifacts and ensure that contrast improves, a novel constraint is included in the optimization. Due to the number of variables in optimization-based image enhancement techniques being equal to the number of gray scales, we quantize the image using a $k$ -means clustering-based histogram mergence (KCHM) method before enhancement. KCHM can significantly reduce the number of image gray scales while effectively preserving the subjective quality. This is useful considering the reduction of variables is good for solving optimization and reducing computation cost. Experimental results demonstrate that the proposed method can significantly improve the subjective image quality by enhancing both the contrast and the image gradient.
In this paper, we present a novel denoising algorithm based on the Rodin-Osher-Fatemi (ROF) model. The goal is to ensure maximum noise removal while preserving image details. To achieve this goal, we developed a new edge detector based on the structure tensor, Non-Local Mean filtering and fuzzy complement. This edge detector is incorporated in the objective function of the ROF model to introduce more control over the amount of regularization allowing more denoising in smooth regions and less denoising when processing edge regions. Experiments on synthetic images demonstrate the efficiency of the edge detector. Furthermore, denoising experiments and comparison with other algorithms show that the proposed method presents good performance in terms of Peak Signal-to-Noise Ratio and Structure Similarity Index.
Outdoor images captured in bad-weather conditions usually have poor intensity contrast and color saturation since the light arriving at the camera is severely scattered or attenuated. The task of improving image quality in poor conditions remains a challenge. Existing methods of image quality improvement are usually effective for a small group of images but often fail to produce satisfactory results for a broader variety of images. In this paper, we propose an image enhancement method, which makes it applicable to enhance outdoor images by using content-adaptive contrast improvement as well as contrast-dependent saturation adjustment. The main contribution of this work is twofold: (1) we propose the content-adaptive histogram equalization based on the human visual system to improve the intensity contrast; and (2) we introduce a simple yet effective prior for adjusting the color saturation depending on the intensity contrast. The proposed method is tested with different kinds of images, compared with eight state-of-the-art methods: four enhancement methods and four haze removal methods. Experimental results show the proposed method can more effectively improve the visibility and preserve the naturalness of the images, as opposed to the compared methods.
This chapter presents an adaptive regularized image interpolation algorithm, which is developed in a general framework of data fusion, to enlarge noisy-blurred, low-resolution (LR) image sequences. Initially, the assumption is made that each LR image frame is obtained by subsampling the corresponding original high-resolution (HR) image frame. Then the mathematical model of the subsampling process is obtained. Given a sequence of LR image frames and the mathematical model of subsampling, the general regularized image interpolation estimates HR image frames by minimizing the residual between the given LR image frame and the subsampled estimated solution with appropriate smoothness constraints. The proposed algorithm adopts spatial adaptivity which can preserve the high-frequency components along the edge orientation in a restored HR image frame. This multiframe image interpolation algorithm is composed of two levels of data fusion. At the first level, an LR image is obtained and used as an input of the adaptive regularized image interpolation. At the second level, the spatially adaptive, fusion-based regularized interpolation is implemented by using steerable orientation analysis. In order to apply the regularization approach to the interpolation procedure, an observation model of the LR video formation system is first presented. Based on the observation model, an interpolated image can be obtained, where the residual between the original HR and the interpolated images is minimized under a priori constraints. In addition, directional high-frequency components are preserved in the noise-smoothing process by combining spatially adaptive constraints. By experimentation, interpolated images using the conventional algorithms are compared with the proposed adaptive fusion-based algorithm. Experimental results show that the proposed algorithm has the advantage of preserving directional high-frequency components and suppressing undesirable artifacts such as noise.
In most practical image processing and computer vision problems involving optimization, the solutions do not span the entire range of real values. This resulting bounding property is often imposed upon the solution process in the form of constraints placed upon system parameters. Optimization problems with constraints are generally referred to as constrained optimization problems.The basic idea of solving most constrained optimization problems is to determine the solution that satisfies the stated performance index, subject to the given constraints. In this chapter, proper methods of finding the solutions to constrained optimization problems, according to the characteristics of various types of constraints, are discussed.If the solution and its local perturbation do not violate the given constraints, we consider the problem as being unconstrained. Therefore, the important issue of constrained optimization is analyzing the local behavior of the solution on or near the boundary of constraints.This chapter includes three sections. The first section describes the constrained optimization problem and defines the terminology commonly used with this problem, both in the literature and in this chapter. The second and the third sections describe optimality conditions and the corresponding optimization methods with linear and nonlinear constraints, respectively.
Many applications require the use of 3D graphics to create models of real environments. These models are usually built from range or depth images. In the scene modeling process, the use of additional 2D digital sensorial information leads to multimodal scene representation, where an image acquired by a 2D sensor is used as a texture map for a geometric model of a scene. In this chapter we present, as an example of optimization, a photo-realistic scene reconstruction procedure using laser range data and color photographs.The reconstruction system involves the creation of triangle meshes from range images as a scene surface representation, but the main emphasis is made on the registration of laser range and photographic images. Major 3D data acquisition techniques are discussed in Appendix C, and a real range data is acquired by using a light amplitude detection and ranging ( LADAR) range scanner.The proposed multimodal image registration approach uses random distributions of pixels to measure the amount of dependence between two images and estimates the relative pose of one imaging system to the other. The similarity metric used in the proposed automatic registration algorithm is based on the chi(2) measure of dependence, which is presented as an alternative to the standard mutual information criterion. These two criteria belong to the class of information theoretic similarity measures, which quantify the dependence in terms of the information provided by one image about the other. For the maximization of the similarity measure, a robust optimization scheme is needed. To achieve both accurate and robust results, genetic algorithms are investigated in the heuristic manner.
In solving optimization problems, proper constraints play an important role in both making the problem well posed and in making the solution approach the desired and appropriate results. Once constraints are imposed on the solution of a problem, it becomes a constrained optimization problem. Since there are a variety of methods available for solving unconstrained optimization problems, rather than the constrained ones, we usually replace the constrained problem by an unconstrained counterpart by using regularization techniques. In many image processing and computer vision applications, optimization problems are used for estimating the complete data. This assumes that the incomplete data, and the transformation that relates the complete and the incomplete data, are given. If the transformation occurs in a space-invariant manner, the optimization process can be performed in the discrete Fourier transform domain, as discussed in the previous chapter. Although frequency domain implementation is extremely efficient for solving space-variant optimization problems, its application is limited because the space-invariance assumption does not hold for many problems. The use of iterative type methods for solving general optimization problems is widely accepted. They include (a) direct search methods, (b) derivative-based methods, (c) conjugate-gradient methods, and (d) quasi-Newton methods. Iterative type methods in the above list have been developed for solving rather general-purpose numerical optimization problems. For this reason, they are not very efficient in solving more stringent optimization problems. This is especially true for image processing and computer vision applications, where special constraints, such as non-negativity and smoothness, are widely used. In this chapter we introduce the regularized iterative method for solving optimization problems in image processing and computer vision applications. Particular emphasis is given to incorporation of constraints and discussion of convergence issues.
In many image processing problems, we need to estimate the original complete data sets, generally from incomplete and, most often, from degraded observations. One simplified example is to estimate the original pixel intensity value, which has been attenuated in the imaging system, without any correlation with neighboring pixels. If we know the nonzero attenuation factor for the imaging system, we can easily estimate the original value by multiplying by this attenuation factor. Figure 1.1 shows the corresponding attenuation and the restoration processes.
This book presents practical optimization techniques used in image processing and computer vision problems. Ill-posed problems are introduced and used as examples to show how each type of problem is r
Regularization methods play an important role in solving linear equations of the form $$ y=Hx, $$ with prior knowledge about the solution. The corresponding regularization results in minimization of $$ f(x)={\left\Vert y-Hx\kern0.1em \right\Vert}^2+\lambda {\left\Vert \kern0.1em Cx\kern0.1em \right\Vert}^2. $$
Most practical optimization problems arise with constraints on the solutions. Nevertheless, unconstrained optimization techniques serve as a major tool in finding solutions for both unconstrained and constrained optimization problems. In this chapter we present techniques for solving the unconstrained optimization problem. In Chap. 5 we will see how the unconstrained solution methods aid in finding the solutions to constrained optimization problems. In order to find the solution of a given unconstrained optimization problem, a general procedure includes an initial solution estimate ( guess) followed by iterative updates of the solution directed toward minimizing the objective function. If we consider the solution of an N-dimensional objective function as an N-vector, iterative updates can be performed by adding a vector to the current solution vector. We call the vector augmented to the current solution vector the step vector or simply the step. Various unconstrained optimization methods can be classified by the method of determining step vectors. We place direct search methods, which are the simplest, in group one. In this group, the step vectors are randomly determined. The second group is known as derivative-based methods. In this case, the step vectors are determined based on the derivative of the objective function. Gradient methods and Newton's methods fall into this category. In practice, more sophisticated methods, such as conjugate gradient and quasi-Newton methods, have shown comparable performance with greatly reduced computational cost and memory space. A brief discussion of the conjugate gradient algorithm will also be included in this chapter.
Segmenting an image into meaningful regions is an important step in many computer vision applications such as facial recognition, target tracking and medical image analysis. Because image segmentation is an ill-posed problem, parameters are needed to constrain the solution to one that is suitable for a given application. For a user, setting parameter values is often unintuitive. We present a method for automating segmentation parameter selection using an efficient search method to optimize a segmentation objective function. Efficiency is improved by utilizing prior knowledge about the relationship between a segmentation parameter and the objective function terms. An adaptive sampling of the search space is created which focuses on areas that are more likely to contain a minimum. When compared to parameter optimization approaches based on genetic algorithm, Tabu search, and multi-locus hill climbing the proposed method was able to achieve equivalent optimization results with an average of 25% fewer objective function evaluations.
A fundamental limitation of hyperspectral imaging is the inter-band misalignment correlated with subject motion during data acquisition. One way of resolving this problem is to assess the alignment quality of hyperspectral image cubes derived from the state-of-the-art alignment methods. In this paper, we present an automatic selection framework for the optimal alignment method to improve the performance of face recognition. Specifically, we develop two qualitative prediction models based on: 1) a principal curvature map for evaluating the similarity index between sequential target bands and a reference band in the hyperspectral image cube as a full-reference metric; and 2) the cumulative probability of target colors in the HSV color space for evaluating the alignment index of a single sRGB image rendered using all of the bands of the hyperspectral image cube as a no-reference metric. We verify the efficacy of the proposed metrics on a new large-scale database, demonstrating a higher prediction accuracy in determining improved alignment compared to two full-reference and five no-reference image quality metrics. We also validate the ability of the proposed framework to improve hyperspectral face recognition.