Humans use context and scene knowledge to easily localize moving objects in conditions of complex illumination changes, scene clutter, and occlusions. In this paper, we present a method to leverage human knowledge in the form of annotated video libraries in a novel search and retrieval-based setting to track objects in unseen video sequences. For every video sequence, a document that represents motion information is generated. Documents of the unseen video are queried against the library at multiple scales to find videos with similar motion characteristics. This provides us with coarse localization of objects in the unseen video. We further adapt these retrieved object locations to the new video using an efficient warping scheme. The proposed method is validated on in-the-wild video surveillance data sets where we outperform state-of-the-art appearance-based trackers. We also introduce a new challenging data set with complex object appearance changes.
Tracking and re-identification in wide-area camera networks is a challenging problem due to non-overlapping visual fields, varying imaging conditions, and appearance changes. We consider the problem of person re-identification and tracking, and propose a novel clothing context-aware color extraction method that is robust to such changes. Annotated samples are used to learn color drift patterns in a non-parametric manner using the random forest distance (RFD) function. The color drift patterns are automatically transferred to associate objects across different views using a unified graph matching framework . A hypergraph representation is used to link related objects for search and re-identification. A diverse hypergraph ranking technique is proposed for person-focused network summarization . The proposed algorithm is validated on a wide-area camera network consisting of ten cameras on bike paths. Also, the proposed algorithm is compared with the state of the art person re-identification algorithms on the VIPeR dataset .
Camera networks provide opportunities for practical video surveillance and monitoring, but tracking people across the network presents many computational and modeling hurdles that researchers have yet to surmount.
In a wide-area camera network, cameras are often placed such that their views do not overlap. Collaborative tasks such as tracking and activity analysis still require discovering the network topology including the extrinsic calibration of the cameras. This work addresses the problem of calibrating a fixed camera in a wide-area camera network in a global coordinate system so that the results can be shared across calibrations. We achieve this by using commonly available mobile devices such as smartphones. At least one mobile device takes images that overlap with a fixed camera's view and records the GPS position and 3D orientation of the device when an image is captured. These sensor measurements (including the image, GPS position, and device orientation) are fused in order to calibrate the fixed camera. This article derives a novel maximum likelihood estimation formulation for finding the most probable location and orientation of a fixed camera. This formulation is solved in a distributed manner using a consensus algorithm. We evaluate the efficacy of the proposed methodology with several simulated and real-world datasets.
Coastal marine ecosystems are highly productive and diverse, but biodiversity of underwater habitats is poorly described due to logistical and financial limitations of diving and submersible operations. Imagery is a promising way to address this challenge, but the complexity of diverse organisms thwarts simple automated analysis. We consider the problem of automated annotation of complex communities of sessile marine invertebrates and macroalgae in order to automate percent coverage estimation. We propose an efficient fusion technique amongst diverse classifiers based on the idea of "dropout" in machine learning. We use dropout technique to weight each classifier implicitly and for each specie we optimize the region of interest (ROI) for highest accuracy. The preliminary results are promising and show 20% increase in average accuracy (over 30 species) when compared with the best base performance of Random Forest classifiers. The data set along with human "ground truth" annotations are available to the public.
For the instance search task, we are given a set of query images with the corresponding textual meta-data and objects masks to retrieve video shots containing query objects from FLICKR video database. We extract meaningful regions in the key-frames using Maximally Stable Extremal Regions (MSER) and use SIFT descriptors for representation. We use standard Bag of visual Word (BoW) model to represent database images. Additionally, we crawled training images for each query topic using the textual meta-data from Google and FLICKR images databases to train a discriminative classifier using Support Vector Machines (SVM). We use a discriminative model to rerank candidate images obtained by initial BoW search. The experimental results demonstrates the efficacy of the overall system. Finally, we highlight the need for domain adaptation when the source and target domains are completely different.
This paper addresses the novel and challenging problem of aligning camera views that are unsynchronized by low and/or variable frame rates using object trajectories. Unlike existing trajectory-based alignment methods, our method does not require frame-to-frame synchronization. Instead, we propose using the intersections of corresponding object trajectories to match views. To find these intersections, we introduce a novel trajectory matching algorithm based on matching Spatio-Temporal Context Graphs (STCGs). These graphs represent the distances between trajectories in time and space within a view, and are matched to an STCG from another view to find the corresponding trajectories. To the best of our knowledge, this is one of the first attempts to align views that are unsynchronized with variable frame rates. The results on simulated and real-world datasets show trajectory intersections are a viable feature for camera alignment, and that the trajectory matching method performs well in real-world scenarios.
This paper addresses the problem of context-aware object search and retrieval in a wide area distributed camera network. With the proliferation of smart cameras in urban networks, it is a challenge to process this big data in an efficient manner. A novel graph based model is proposed to represent relationships, and for search and retrieval tasks. This representation exploits the fact that objects occurring in close spatial-temporal proximity are not completely independent and serve as context for each other. Additional information such as appearance and scene context can also be encoded into the graph model to improve the overall accuracy. A manifold ranking strategy is used to order the items based on similarity with an emphasis on diversity. Extensive experimental results on a ten camera network are presented.
This paper presents a novel and computationally efficient multi-object tracking-by-detection algorithm with interacting particle filters. The proposed online tracking methodology could be scaled to hundreds of objects and could be completely parallelized. For every object, we have a set of two particle filters, i.e. local and global. The local particle filter models the local motion of the object. The global particle filter models the interaction with the other objects and scene. These particle filters are integrated into a unified Interacting Markov Chain Monte Carlo (IMCMC) framework. The local particle filter improves its performance by interacting with the global particle filter while they both are run in parallel. We indicate the manner in which we bring in object interaction and domain specific information into account by using global filters without further increase in complexity. Most importantly, the complexity of the proposed methodology varies linearly in the number of objects. We validated the proposed algorithms on two completely different domains 1) Pedestrian Tracking in urban scenarios 2) Biological cell tracking (Melanosomes). The proposed algorithm is found to yield favorable results compared to the existing algorithms.
Wide-area wireless camera networks are being increasingly deployed in many urban scenarios. The large amount of data generated from these cameras pose significant information processing challenges. In this work, we focus on representation, search and retrieval of moving objects in the scene, with emphasis on local camera node video analysis. We develop a graph model that captures the relationships among objects without the need to identify global trajectories. Specifically, two types of edges are defined in the graph: object edges linking the same object across the whole network and context edges linking different objects within a spatial-temporal proximity. We propose a manifold ranking method with a greedy diversification step to order the relevant items based on similarity as well as diversity within the database. Detailed experimental results using video data from a 10-camera network covering bike paths are presented.
This paper proposes a distributed multi-camera tracking algorithm with interacting particle filters. A robust multi-view appearance model is obtained by sharing training samples between views. Motivated by incremental learning and [1], we create an intermediate data representation between two camera views with generative subspaces as points on a Grassmann manifold, and sample along the geodesic between training data from two views to uncover the meaningful description due to viewpoint changes. Finally, a Boosted appearance model is trained using the projected training samples on to these generative subspaces. For each object, a set of two particle filters i.e., local and global is used. The local particle filter models the object motion in the image plane. The global particle filter models the object motion in the ground plane. These particle filters are integrated into a unified Interacting Markov Chain Monte Carlo (IMCMC) framework. We show the manner in which we induce priors on scene specific information into the global particle filter to improve tracking accuracy. The proposed algorithm is validated with extensive experimentation in challenging camera network data, and compares favorably with state of the art object trackers.
We detect seven activities defined by TRECVID SED task such as CellToEar, Embrace, ObjectPut, PeopleMeet, PeopleSplitUp, PersonRuns, and Pointing. We employ two different strategies to detect these activities based on their characteristics. Activities like CellToEar, Embrace, ObjectPut, and Pointing are the results of articulated motion of human parts. Therefore, we employ local spatio-temporal interest point (STIP) feature based bag of words strategy for these activities. Visual vocabularies are constructed from the STIP features and each activity is described by the histograms of visual words. We also construct activity probability map for each camera-activity pair that reflects the spatial distribution of an activity in a camera. We train a discriminative SVM classifier using Gaussian kernel for each camera-activity pair. During evaluation we employ sliding window based technique. We slide spatio-temporal cuboids in both spatial and temporal direction to find a likely activity. The cuboid is also described by the histograms of visual words and final decision is made using the SVM classifier and the activity probability map. For the activities like PeopleMeet, PeopleSplitUp, and PersonRuns, the characteristics of trajectories of persons of interest in the activities are discriminative. For instance, trajectories of PeopleMeet converge along time while those of PeopleSplitUp diverge along time. Therefore, we use track-based string of feature graph (SFG) to recognize these activities. Results of our experimental runs on the evaluation videos are comparable with other participants. Our performances in all the activities are among the top five teams.
For the instance search task, we are given a set of query images with the corresponding textual meta-data and objects masks to retrieve video shots containing query objects from FLICKR video database. We extract meaningful regions in the key-frames using Maximally Stable Extremal Regions (MSER) and use SIFT descriptors for representation. We use standard Bag of visual Word (BoW) model to represent database images. Additionally, we crawled training images for each query topic using the textual meta-data from Google and FLICKR images databases to train a discriminative classifier using Support Vector Machines (SVM). We use a discriminative model to rerank candidate images obtained by initial BoW search. The experimental results demonstrates the efficacy of the overall system. Finally, we highlight the need for domain adaptation when the source and target domains are completely different.
This paper proposes a distributed algorithm for object tracking in a camera sensor network. At each camera node, an efficient online multiple instance learning algorithm is used to model object's appearance. This is integrated with particle filter for camera's image plane tracking. To improve the tracking accuracy, each camera node shares its particle states with others and fuses multi-camera information locally. In particular, particle weights are updated according to the fused information. Then, appearance model is updated with the re-weighted particles. The effectiveness of the proposed algorithm is demonstrated on human tracking in challenging environments.
This paper addresses the problem of object tracking by learning a discriminative classifier to separate the object from its background. The online-learned classifier is used to adaptively model object's appearance and its background. To solve the typical problem of erroneous training examples generated during tracking, an online multiple instance learning (MIL) algorithm is used by allowing false positive examples. In addition, particle filter is applied to make best use of the learned classifier and help to generate a better representative set of training examples for the online MIL learning. The effectiveness of the proposed algorithm is demonstrated in some challenging environments for human tracking.