Current state-of-the-art methods for object tracking perform adaptive trackingby-detection, meaning that a detector predicts the position of an object and adapts its parameters to the object’s appearance at the same time. While suitable for cases when the object does not disappear from the scene, these methods tend to fail on occlusions. In this work, we build on a novel approach called Tracking-Learning-Detection (TLD) that overcomes this problem. In methods based on TLD, a detector is trained with examples found on the trajectory of a tracker that itself does not depend on the object detector. By decoupling object tracking and object detection we achieve high robustness and outperform existing adaptive tracking-by-detection methods. We show that by using simple features for object detection and by employing a cascaded approach a considerable reduction of computing time is achieved. We evaluate our approach both on existing standard single-camera datasets as well as on newly recorded sequences in multi-camera scenarios.
Data acquisition by multidomain data acquisition provides means for environment perception usable for detecting unusual and possibly dangerous situations. When being automated, this approach can simplify surveillance tasks required in, for example, airports or other security sensitive infrastructures. This paper describes a novel architecture for surveillance networks based on combining multimodal sensor information. Compared to previous methodologies using only video information, the proposed approach also uses audio data thus increasing its ability to obtain valuable information about the sensed environment. A hierarchical processing architecture for observation and surveillance systems is proposed, which recognizes a set of predefined behaviors and learns about normal behaviors. Deviations from “normality” are reported in a way understandable even for staff without special training. The processing architecture, including the physical sensor nodes, is called smart embedded network of sensing entities (SENSE).
Deployment of existing vision approaches in camera networks for applications such as human tracking show a large gap between user expectation and current results. Calibrated cameras could push these approaches closer to applicability, as physical constraints greatly complement the ill-posed acquisition process. Calibrated cameras promise also new applications as spatial relationships among cameras and the environment capture additional information. However, a convenient calibration is still a challenge on its own. This paper presents a novel calibration framework for large networks including non-overlapping cameras. The framework purely relies on visual information coming from walking people. Since non-overlapping scenarios make point correspondences impossible, time constancy of a person’s motion introduces the missing complementary information. The framework obtains calibrated cameras starting from single camera calibration thereby bringing the problem to a reduced form suitable for multi-view calibration. It extends the standard bundle adjustment by a smoothness constraint to avoid the ill-posed problem arising from missing point correspondences. The stratified optimization suppresses the danger to get stuck in local minima. Experiments with synthetic and real data validate the approach.
The paper presents a novel approach for single object tracking across non-overlapping camera views, which searches for the optimal association of single view trajectories. We map the tracking problem to a tree structure and introduce a branch and bound approach to efficiently explore the search space. We use an optimization criterion based solely on the geometric cues coming from the calibration of the network. The cost function is defined as to enforce consistency of geometrical and kinematic properties over the whole trajectory path. We show how the information content of the geometric properties of the network brings a substantial contribution to solve the association problem. Experiments in a set-up of four cameras using both synthetic and real trajectories validate the advantages of the approach both in terms of performance and information content of the geometric cue.
This paper presents a stratified auto-calibration framework for typical large surveillance set-ups including non-overlapping cameras. The framework avoids the need of any calibration target and purely relies on visual information coming from walking people. Since in non-overlapping scenarios there are no point correspondences across the cameras the standard techniques cannot be employed. We show how to obtain a fully calibrated camera network starting from single camera calibration and bringing the problem to a reduced form suitable for multi-view calibration. We extend the standard bundle adjustment by a smoothness constraint to avoid the ill-posed problem arising from missing point correspondences. The proposed framework optimizes the objective function in a stratified manner thus suppressing the problem of local minima. Experiments with synthetic and real data validate the approach.
The tracking of storm centres in radar data is of particular importance for short term weather prediction and specifically thunderstorm prediction. This paper presents a method to track storm centres in terrestrial radar images. Mean shift segmentation is used to outline storm centres and mean shift tracking to locate the storm in the consecutive images. Re sults demonstrate the ability of the method to deal with deformable objects such as storm centres. Moreover the method is able to handle the splitting and merging of the convective cells.
Branislav Mičušík合作论文数Department of Computer Science
George Mason University4