Calibration of the internal and external parameters of a stereo vision camera is a well-known research problem in the computer vision. Usually, to get accurate 3D results the camera should be manually calibrate accurately as well. This paper proposes a robust approach to Auto Calibration stereo camera Without intervention of the user. There are several methods and techniques of calibration that have been proven, in this work we exploiting the geometric constraint, namely, the epipolar geometry. We specifically focuses to use 7 techniques for Features Extraction (SURF, BRISK, FAST, FREAK, MinEigen, MSERF, SIFT), however tries to establish the correspondences between points extracted in stereo images with Various Matching Techniques (SSD, SAD, Hamming). Then we exploits the Fundamental Matrix to estimate the epipolar Line by choosing the perfect Eight-point algorithms (Norm8Point, LMedS, RANSAC, MSAC, LTS). A large number of experiments have been carried out, and very good results have been obtained by Comparison & choice the perfect technique in every stage.
We present a new multi-modal technique for assisting visually-impaired people in recognizing objects in public indoor environment. Unlike common methods which aim to solve the problem of multi-class object recognition in a traditional single-label strategy, a comprehensive approach is developed here allowing samples to take more than one label at a time. We jointly use appearance and depth cues, specifically RGBD images, to overcome issues of traditional vision systems using a new complex-valued representation. Inspired by complex-valued neural networks (CVNNs) and multi-label learning techniques, we propose two methods in order to associate each input RGBD image to a set of labels corresponding to the object categories recognized at once. The first one, ML-CVNN, is formalized as a ranking strategy where we make use of a fully complex-valued RBF network and extend it to be able to solve multi-label problems using an adaptive clustering method. The second method, L-CVNNs, deals with problem transformation strategy where instead of using a single network to formalize the classification problem as a ranking solution for the whole label set, we propose to construct one CVNN for each label where the predicted labels will be later aggregated to construct the resulting multi-label vector. Extensive experiments have been carried on two newly collected multi-labeled RGBD datasets prove the efficiency of the proposed techniques.
Computation of stereoscopic depth and disparity map extraction are dynamic research topics. A large variety of algorithms has been developed, among which we cite feature matching, moment extraction, and image representation using descriptors to determine a disparity map. This paper proposes a new method for stereo matching based on Fourier descriptors. The robustness of these descriptors under photometric and geometric transformations provides a better representation of a template or a local region in the image. In our work, we specifically use generalized Fourier descriptors to compute a robust cost function. Then, a box filter is applied for cost aggregation to enforce a smoothness constraint between neighboring pixels. Optimization and disparity calculation are done using dynamic programming, with a cost based on similarity between generalized Fourier descriptors using Euclidean distance. This local cost function is used to optimize correspondences. Our stereo matching algorithm is evaluated using the Middlebury stereo benchmark; our approach has been implemented on parallel high-performance graphics hardware using CUDA to accelerate our algorithm, giving a real-time implementation.
In this article, we deal with the problem of understanding human-to-human interactions as a fundamental component of social events analysis. Inspired by the recent success of multi-modal visual data in many recognition tasks, we propose a novel approach to model dyadic interaction by means of features extracted from synchronized 3D skeleton coordinates, depth, and Red Green Blue (RGB) sequences. From skeleton data, we extract new view-invariant proxemic features, named Unified Proxemic Descriptor (uProD), which is able to incorporate intrinsic and extrinsic distances between two interacting subjects. A novel key frame selection method is introduced to identify salient instants of the interaction sequence based on the joints' energy. From Red Green Blue Depth (RGBD) videos, more holistic CNN features are extracted by applying an adaptive pre-trained Convolutional Neural Networks (CNNs) on optical flow frames. For better understanding the dynamics of interactions, we expand the boundaries of dyadic interactions analysis by proposing a fundamentally new modeling for non-treated problem aiming to discern the active from the passive interactor. Extensive experiments have been carried out on four multi-modal and multi-view interactions datasets. The experimental results demonstrate the superiority of our proposed techniques against the state-of-the-art approaches.
Dense depth map extraction is a dynamic research field in a computer vision that tries to recover three-dimensional information from a stereo image pair. A large variety of algorithms has been developed. The local methods based on block matching that are prevalent due to the linear computational complexity and easy implementation. This local cost is used on global methods as graph cut and dynamic programming in order to reduce sensitivity to local to occlusion and uniform texture. This paper proposes a new method for matching images based on a two-stage of block matching as local cost function and dynamic programming as energy optimization approach. In our work introduce the two stage of the zero-mean sum of absolute differences (ZSAD) combined with dynamic programming: the smoothness and ordering constraints are used to optimize correspondences. Stereo matching accuracy and runtime are the fundamental metrics to evaluate the stereo matching methods. The real-time has become a reality through the complexity reduction of the calculation and the use of parallel high-performance graphics hardware. In this paper we evaluate the developed method on using Middlebury stereo benchmark and, we propose a GPU CUDA implementation in order to accelerate our algorithm and reach the real time.
Object recognition methods usually tend to focus on single cues coming from traditional vision based systems but ignore to incorporate multi-modal data. With the advent of depth RGB-D sensors which provide synchronized multi-modal data with good quality, new opportunities have been emerged. In this paper, we make use of RGB and depth images to propose a new object recognition approach. Using a pixel-wise scheme, we propose a novel method to describe RGB-D images with a complex-valued representation. By means of neural network, we introduce a new CVNN (Complex-Valued Neural Network) with RBF neurons. Different from many RGB-D features, the proposed approach is able to jointly use RGB and depth data within a unified end-to-end learning framework. Category and instance object recognition tasks are evaluated through experiments carried out on a large scale RGB-D object dataset. Results show that our method can efficiently recognize objects in RGB-D images and outperforms state-of-the-art approaches.
In this paper a robust and simple scheme is presented for three dimensional (3D) shape reconstruction of real object. A novel composite pattern technique is proposed for projecting the light pattern on the object of interest. The proposed scheme reduces the number of patterns by combining the primary color coded channels into one composite format. Our approach uses both spatial and temporal intensity variation for calibration and construction phase. Gamma calibration is considered with the propose scheme. High quality depth map is obtained from the linear light reflected by the shape of object without complex calculations. Experimental results demonstrated that proposed technique is fast and exhibit high level of precision. In addition hardware cost is minimized as compare to current calibration procedures used in structured light scanning system. Our scheme requires a digital camera, flashlight and mask of pattern only.
This paper addresses the problem of foreground and background segmentation. Multi-modal data specifically RGBD data has gain many tasks in computer vision recently. However, techniques for background subtraction based only on single-modality still report state-of-the-art results in many benchmarks. Succeeding the fusion of depth and color data for this task requires a robust formalization allowing at the same time higher precision and faster processing. To this end, we propose to make use of kernel density estimation technique to adapt multi-modal data. To speed up kernel density estimation, we explore the fast Gauss transform which allows the summation of a mixture of M kernel at N evaluation points in O(M+N) time as opposed to O(MN) time for a direct evaluation. Extensive experiments have been carried out on four publicly available RGBD foreground/background datasets. Results demonstrate that our proposal outperforms state-of-the-art methods for almost all of the sequences acquired in challenging indoor and outdoor contexts with a fast and non-parametric operation. (C) 2017 Elsevier B.V. All rights reserved.
This paper contributes to 3D facial synthesis by presenting a novel method for parameterization using Landmark Point detection. The approach presented aims at improving facial recognition even in varying facial expressions, and missing data in 3D facial models. As such, the prime objective was to develop an automatically embedded process that can detect any frontal face in 3D face recognition systems, with face segmentation and surface curvature information. Using the hybrid interpolation method, experiments on facial landmarks were performed on 4950 images from Face Recognition Grand Challenge database (FRGC). Distinctive facial landmarks from the nose–tips, Limits mouth and two eye corners formed the statistical inputs for Iterative Closest Point (ICP) in the Point Distribution Model (PDM). Performance or landmark localization is reported by using percentage deviation from the mean 3D profile. Localization results and estimated data on landmark locations demonstrate that the method confirms its effectiveness for proposed application.
This paper presents our methodology for Landmark Point detection to improve 3D face recognition in a presence of variant facial expression. The objective was to develop an automatic process for distinguishing and segmenting to be embedded in a 3D face recognition system using only 3D Point Distribution Model (PDM) as input. The approach used hydride method to extract this features from the surface curvature information. Landmark Localization is done on the segmented face via finding the change that decreases the deviation of the model from the mean profile. Face registering is achieved using previous anthropometric information and the localized landmarks. The results confirm that the method used is accurate and robust for the proposed application.
This paper presents a method of background subtraction that uses multimodal information, specifically depth and appearance cues, to robustly separate the foreground in dynamic indoor scenes. To this end, RGB-Depth data from a Microsoft Kinect sensor are exploited. We propose an extension of one from the most effective technique for background modeling in real time: Kernel Density Estimation with Fast Gauss Transform technique. Experimental results show that our proposed deals well with gradual/sudden illumination changes, shadows and dynamic backgrounds.
The aim of this paper is to present the use of Fourier descriptors for color images recognition. We are interested in the generalized Clifford Fourier descriptors (GCFD).We Present these family of descriptors, evaluate its performance by its application to color images using a pure software implementation. This implementation does not meet the real time constraint where the need for hardware implementation in order to speed up the recognition system. So we seek a better partitioning of our system in two-part HW / SH. The design and implementation of GCFD on XILINX FPGA by exploiting FFT IP are detailed in two different architectures.
This paper presents an image processing system based on smart camera platform, whose two principle elements are a Pan-Tilt-Zoom (PTZ) camera and a Field Programmable Gate Array (FPGA). The latter is used to control the various sensor parameter configurations and, where desired, to receive and process the images captured by the PTZ Camera. With the advent of today's highly integrated Field Programmable Gate Array (FPGA) it is possible to have a software programmable processor and hardware computing resources on the same chip. Apart from having sufficient logic blocks on which the hardware is implemented these chips also have an embedded processor with system software to implement the application software around it. In this paper, the Spartan3A DSP based Xilinx VSK platform is used for developing the proposed extensible hardware-software video streaming and processing modules. In order to develop the required hardware and software in an integrated fashion, Xilinx Embedded Development Kit (EDK) design tool has been used. A number of Xilinx provided IPs are customized to realize the hardware modules in the FPGA fabric.
This paper presents an image processing system based on smart camera platform, whose two principle elements are a Wide-VGA CMOS Sensor and a Field Programmable Gate Array (FPGA). The latter is used to control the various sensor parameter configurations and, where desired, to receive and process the images captured by the CMOS sensor. With the advent of today's highly integrated Field Programmable Gate Array (FPGA) it is possible to have a software programmable processor and hardware computing resources on the same chip. Apart from having sufficient logic blocks on which the hardware is implemented these chips also have an embedded processor with system software to implement the application software around it. In this paper, the Spartan-3A DSP based Xilinx VSK platform is used for developing the proposed extensible hardware-software video streaming and processing modules. In order to develop the required hardware and software in an integrated fashion, Xilinx Embedded Development Kit (EDK) design tool has been used. A number of Xilinx provided IPs are customized to realize the hardware modules in the FPGA fabric,
The present work aims to exploit the new generation of 3D vision systems for detecting people. We present a challenge process dedicated to test the feasibility of detection over disparity maps by exploiting techniques used with monocular cues, specifically HOG/SVM. Disparity maps are extracted by a developed stereoscopic vision system using two passive sensors with an algorithm stack well adopted to real time constraint with lower processing speeds. This detection module can improve systems’perception ability in complex scenes under shadows, gradual/sudden illumination changes and animated texture. Another key point is to estimate their exact locations to predict intrusions in monitored areas. Results indicate a clear advantage of the proposed method to enhance the rate of performance up to 99.6%.
A methodology for implementing real-time DSP applications on a field programmable gate arrays (FPGA) using Xilinx System Generator (XSG) for Matlab is presented in this paper.It presents architecture for Edge Detection using Sobel Filter for image processing using Xilinx System Generator. The design was implemented targeting a Spartan3A DSP 3400 device (XC3SD3400A-4FGG676C) then a Virtex 5 (xc5vlx50-1ff676). The Edge Detection method has been verified successfully with no visually perceptual errors in the resulted images.
This paper presents a study on human detection using the multi-scale covariance descriptor (MSCOV) proposed in a previous work [1] in which we showed the performance of this descriptor for human re-identification. In this work, we evaluate its performance in human detection. We propose a fast tree based method for multi-scale features covariance computation. This method considerably speed up the image scan process for fast object detection. Furthermore, we experimentally evaluate the human detection performance using region covariance descriptor (COV), multi-scale covariance descriptor (MSCOV) and histogram of oriented gradients (HOG). In term of classifier, we consider the popular Support Vector Machines (SVM). The experiments are performed on both benchmarking datasets INRIA and MIT CBCL. Experiments on both datasets show the high detection performance of the MSCOV based detector.
Image Processing algorithms implemented in hardware have emerged as the most viable solution for improving the performance of image processing systems. The introduction of reconfigurable devices and high level hardware programming languages has further accelerated the design of image processing in FPGA. This paper briefly presents the design of Sobel edge detector system on FPGA. The design is developed in System Generator and integrated as a dedicated hardware peripheral to the Microblaze 32 bit soft RISC processor with the EDK embedded system. The input comes from a live video acquired from a CMOS camera and the detected edges are displayed on a DVI display screen.
Object matching is the process of determining the presence and the location of a reference object inside a scene image. Matching accuracy requires robust image description and efficient similarity measures. In this paper, we present a tree based object matching approach using a descriptor proposed in a previous work [1]. Visual objects are described by a collection of multi-scale covariance matrices structured in a tree form. Tree matching is then performed to match visual objects. With this approach, matching accuracy considerably increases compared to traditional image matching techniques. The proposed matching approach is evaluated on CAVIAR dataset. Overall, our approach is an important contribution to a complete system for object reidentification and tracking over different camera views.
In wireless camera networks, the communication load between cameras is a major concern for visual tracking. To save the bandwidth, traditional applications transfer the spatial coordinates under the precondition of camera calibration, which is computationally unreasonable for large and mobile camera networks. In this chapter, we exploit the use of distinctive and fast to compute local features to represent the non-rigid targets. Transmission of feature descriptors between cameras is done without any calibration. Combining the haar-like patterns and relative color information, our local features succeed to re-identify and relocate the target among the distributed cameras. Furthermore, efficient interest point detection and matching scheme are proposed for the visual tracking under real-time constraints.