The U.S. Defense Advanced Research Projects Agency's (DARPA) Neovision2 program aims to develop artificial vision systems based on the design principles employed by mammalian vision systems. Three such algorithms are briefly described in this paper. These neuromorphic-vision systems' performance in detecting objects in video was measured using a set of annotated clips. This paper describes the results of these evaluations including the data domains, metrics, methodologies, performance over a range of operating points and a comparison with computer vision based baseline algorithms.
Applications of a simple spatiotemporal characterization of human gait in the surveillance domain are presented. The approach is based on decomposing a video sequence into x-t slices, which generate periodic patterns referred to as double helical signatures (DHSs). The features of DHS are given as follows: 1) they naturally encode the appearance and kinematics of human motion and reveal geometric symmetries and 2) they are effective and efficient for recovering gait parameters and detecting simple events. We present an iterative local curve embedding algorithm to extract the DHS from video sequences. Two applications are then considered. First, the DHS is used for simultaneous segmentation and labeling of body parts in cluttered scenes. Experimental results showed that the algorithm is robust to size, viewing angles, camera motion, and severe occlusion. Then, the DHS is used to classify load-carrying conditions. By examining various symmetries in DHS, activities such as carrying, holding, and walking with objects that are attached to legs are detected. Our approach possesses several advantages: a compact representation that can be computed in real time is used; furthermore, it does not depend on silhouettes or landmark tracking, which are sensitive to errors in background subtraction stage.
Soft biometrics, as a prescreening filter, contribute to a much smaller candidate pool and allow the overall query to perform better and faster. In this paper, we focus on the efficiency and effectiveness of several soft biometrics for surveillance applications. We propose a temporal signature in x-t slices. Such a signature has explicitly embedded body articulation and enables direct mensuration. The algorithms determine characteristics for gender, body size, height, cadence, and stride of the subject using a novel gait analysis tool. We have evaluated algorithm performance under various poses, ranges, and illuminations. Preliminary experiments have shown promising results.
We describe an approach to characterize the signatures generated by walking humans in spatio-temporal domain. To describe the computational model for this periodic pattern, we take the mathematical theory of Geometry Group Theory, which is widely used in crystallographic structure research. Both empirical and theoretical analysis prove that spatio-temporal helical patterns generated by legs belong to the Frieze Groups because they can be characterized by a repetitive motif along the direction of walking. The theory is applied to an automatic detection-and-tracking system capable of counting heads and handling occlusion by recognizing such patterns. Experimental results for videos acquired from both static and moving ground sensors are presented. Our algorithm demonstrates robustness to nonrigid human deformation as well as background clutter.
In target tracking, fusing multi-modal sensor data under a power-performance trade-off is becoming increasingly important. Proper fusion of multiple modalities can help in achieving better tracking performance while decreasing the total power consumption. In this paper, we present a framework for tracking a target given joint acoustic and video observations from a co-located acoustic array and a video camera. We demonstrate on field data that tracking of the direction-of-arrival of a target improves significantly when the video information is incorporated at time instants when the acoustic signal-to-noise ratio is low.
Visual surveillance systems have gained a lot of interest in the last few years. In this paper, we present a visual surveillance system that is based on the integration of motion detection and visual tracking to achieve better performance. Motion detection is achieved using an algorithm that combines temporal variance with background modeling methods. The tracking algorithm combines motion and appearance information into an appearance model and uses a particle filter framework for tracking the object in subsequent frames. The systems was tested on a large ground-truthed data set containing hundreds of color and FLIR image sequences. A performance evaluation for the system was performed and the average evaluation results are reported in this paper.
We describe algorithms for detecting pedestrians in videos acquired by infrared (and color) sensors. Two approaches are proposed based on gait. The first employs computationally efficient periodicity measurements. Unlike other methods, it estimates a periodic motion frequency using two cascading hypothesis testing steps to filter out non-cyclic pixels so that it works well for both radial and lateral walking directions. The extraction of the period is efficient and robust with respect to sensor noise and cluttered background. In order to integrate shape and motion, we convert the cyclic pattern into a binary sequence by Maximal Principal Gait Angle (MPGA) fitting in the second method. It does not require alignment and continuously estimates the period using a Phase-locked Loop. Both methods are evaluated by experimental results that measure performance as a function of size, movement direction, frame rate and sequence length.
Video based human motion analysis has been actively studied over the past decades. We propose novel approaches that are able to analyze human motion under such challenges and apply them to surveillance and security applications. Part I analyses the cyclic property of human motion and presents algorithms to classify humans in videos by their gait patterns. Two approaches are proposed. The first employs the computationally efficient periodogram, to characterize periodicity. In order to integrate shape and motion, we convert the cyclic pattern into a binary sequence using the angle between two legs when the toe-to-toe distance is maximized during walking. Part II further extends the previous approaches to analyze the symmetry in articulation within a stride. A feature that has been shown in our work to be a particularly strong indicator of the presence of pedestrians is the X-junction generated by bipedal swing of body limbs. The proposed algorithm extracts the patterns in spatio-temporal surfaces. In Part III, we present a compact characterization of human gait and activities. Our approach is based on decomposing an image sequence into x-t slices, which generate twisted patterns defined as the Double Helical Signature (DHS). It is shown that the patterns sufficiently characterize human gait and a class of activities. The features of DHS are: (1) it naturally codes appearance and kinematic parameters of human motion; (2) it reveals an inherent geometric symmetry (Frieze Group); and (3) it is effective and efficient for recovering gait and activity parameters. Finally, we use the DHS to classify activities such as carrying a backpack, briefcase, etc. The advantage of using DHS is that we only need a small portion of 3D data to recognize various symmetries.
In this paper, the problem of simultaneous motion estimation of multiple independently moving objects is addressed. A novel Bayesian approach is designed for solving this problem using the sequential importance sampling (SIS) method. In the proposed algorithm, a balancing step is added into the SIS procedure to preserve samples of low weights so that all objects have enough samples to propagate empirical motion distributions. By using the proposed algorithm, the relative motions of all moving objects with respect to camera can be simultaneously estimated . This algorithm has been tested on both synthetic and real image sequences. Improved results have been achieved.
This paper describes an efficient pedestrian detection system for videos acquired from moving platforms. Given a detected and tracked object as a sequence of images within a bounding box, we describe the periodic signature of its motion pattern using a twin-pendulum model. Then a principle gait angle is extracted in every frame providing gait phase information. By estimating the periodicity from the phase data using a digital phase locked loop (dPLL), we quantify the cyclic pattern of the object, which helps us to continuously classify it as a pedestrian. Past approaches have used shape detectors applied to a single image or classifiers based on human body pixel oscillations, but ours is the first to integrate a global cyclic motion model and periodicity analysis. Novel contributions of this paper include: i) development of a compact shape representation of cyclic motion as a signature for a pedestrian, ii) estimation of gait period via a feedback loop module, and iii) implementation of a fast online pedestrian classification system which operates on videos acquired from moving platforms.
Surface reconstruction is a critical step in three-dimensional image processing and understanding. In this letter, a discontinuity-embedded deformable model has been developed to model surfaces with discontinuities. Governed by the Lagrange motion equation, a finite-element representation of the model can dynamically fit the data in both continuous and discontinuous components, reaching its equilibrium in response to induced forces. Experimental results on synthetic and range images demonstrate a significant improvement in preserving depth discontinuities over conventional approaches.
This paper describes a periodic motion based pedestrian segmentation algorithm for videos acquired from moving platforms. Given a sequence of bounding boxes containing the detected and tracked walking human, the goal is to analyze the low D structure by considering every object sample as a point in the high D manifold space and use the learned structure for segmentation. In this work, we introduce a novel bottom-up learning approach. We represent the human stride as a cascade of models with increasing parameter numbers. These parameters describe the dynamics of pedestrians from coarse to fine. By applying the learned manifold structure, we can predict the location of body parts, especially legs, with high accuracy at every frame. The segmentation in consecutive images is done by EM clustering. With the accuracy for prediction using the twin-pendulum model, EM is more likely to converge to global maximums. Experimental results for real videos are presented The algorithm has demonstrated a reliable performance for videos acquired from moving platforms.
This paper describes a periodicity motion detection based object classification algorithm for infrared videos. Given a detected and tracked object, the goal is to analyze the periodic signature of its motion pattern. We propose an efficient and robust solution, which is related to the frequency estimation in speech recognition. Periodic reference functions are correlated with the video signal. Experimental results for both infrared and visible videos acquired by ground-based as well as airborne moving sensors are presented.
Abstract : We present an approach for vehicle classification in IR video sequences by integrating detection, tracking and recognition. The method has two steps. First, the moving target is automatically detected using a detection algorithm. Next, we perform simultaneous tracking and recognition using an appearance-model based particle filter. The tracking result is evaluated at each frame. Low confidence in tracking performance initiates a new cycle of detection, tracking and classification. We demonstrate the robustness of the proposed method using outdoor IR video sequences.
We present a robust algorithm for tracking moving objects from a moving platform. Robustness is achieved by incorporating temporal differencing and shape detection in an appearance-based object tracking algorithm. In addition, the incorporation of these two methods also improves accuracy and computational efficiency of detection. Some experimental results using airborne-video tracking are given to illustrate the effectiveness of this method.
In this paper, we propose an algorithm for robust 3D motion estimation of wide baseline cameras from noisy feature correspondences. The posterior probability density function of the camera motion parameters is represented by weighted samples. The algorithm employs a hierarchy coarse-to-fine strategy. First, a coarse prior distribution of camera motion parameters is estimated using the random sample consensus scheme (RANSAC). Based on this estimate, a refined posterior distribution of camera motion parameters can then be obtained through importance sampling. Experimental results using both synthetic and real image sequences indicate the efficacy of the proposed algorithm.
Yang Ran合作论文数University of Maryland, College Park10
Thomas M. Strat合作论文数Defense Advanced Research Projects Agency (DARPA) Information Systems Office2