This paper presents key techniques for real-time, multi-person tracking and association on doorbell surveillance cameras at the edge. The challenges for this task are: significant person size changes during tracking caused by person approaching or departing from the doorbell camera, person occlusions due to limited camera field and occluding objects in the camera view, and the requirement for a lightweight algorithm that can run in real time on the doorbell camera at the edge. To address these challenges, we propose a multi-person tracker that uses a detect-track-associate strategy to achieve good performance in speed and accuracy. The person detector only runs at every n-th frame, and between person detection frames a low-cost point-based tracker is used to track the subjects. To maintain subject tracking accuracy, at each person detection frame, a person association algorithm is used to associate persons detected in the current frame to the current and recently tracked subjects and identify any new subjects. To improve the performance of the point-based tracker, human-shaped masks are used to filter out background points. Further, to address the challenge of drastic target scale change during the tracking we introduced an adaptive image resizing strategy to dynamically adjust the tracker input image size to allow the point-based tracker to operate at the optimal image resolution given a fixed number of feature points. For fast and accurate person association, we introduced the Sped-Up LOMO, a fast version of the popular local maximal occurrence (LOMO) person descriptor. The experimental results on doorbell surveillance videos illustrate the efficacy of the proposed person tracking and association framework.
Videos from an Unmanned Aerial Vehicle (UAV) platform have become more available in the surveillance community. As the number of video samples are increasing, automated analysis of these videos have become important. In this paper, we describe the task of vehicle movement event classification in videos taken from UAV. Given object track information from UAV video sequences, we processed the noisy trajectories to make reliable so that we can classify each trajectory as movement event using trajectory curvature and speed. We also derived an evaluation methodology framework and tested our method with a public UAV dataset.
In this paper, we describe how to detect abnormal human activities taking place in an outdoor surveillance environment. Human tracks are provided in real time by the baseline video surveillance system. Given trajectory information, the event analysis module will attempt to determine whether or not a suspicious activity is currently being observed. However, due to real-time processing constrains, there might be false alarms generated by video image noise or non-human objects. It requires further intensive examination to filter out false event detections which can be processed in an off-line fashion. We propose a hierarchical abnormal event detection system that takes care of real time and semi-real time as multi-tasking. In low level task, a trajectory-based method processes trajectory data and detects abnormal events in real time. In high level task, an intensive video analysis algorithm checks whether the detected abnormal event is triggered by actual humans or not.
The task of understanding video content has seen great interest from computer vision community with the increase in camera based surveillance at grocery stores, airports, train stations, etc. What makes up a scene (objects) and what happens in the scene (actions) are two important dimensions of video understanding. In this work, we aim to identify both actions and objects in the video, however, we focus only on the objects with which human interacts. We use videos which may have multiple actions taking place during possibly overlapping intervals. Our system can recognize actions having high intra-class variance performed in complex environments using objects of different types, sizes and shapes. We produce structured descriptions for the videos as output. The descriptions identify the subject, the object, the verb and the interval of each activity recognized.
The SESAME team submitted runs as a full participant in the MED13 evaluation, and submitted video, motion, and audio features; high-level semantic concepts for visual objects, scenes, persons, and actions; automatic speech recognition (ASR); and video optical character recognition (OCR). The individual types of features and concepts produced a total of eight event classifiers. We combined the event detection results for these classifiers using arithmetic mean and log-likelihood ratio fusion methods, and developed and applied a method for selecting the detection threshold. The SESAME system generated event recountings by selecting intervals based on the semantic concepts, and on concepts recognized by ASR and OCR. Our major findings are: Our strategy of first selecting the most informative interval for a video, and then determining the most appropriate event-related semantic concepts within that interval to display for multimedia event recounting (MER), produced the best ObsTextScore in the evaluation. (The ObsTextScore measures the judges’ responses to the question “How well does the text of this observation describe the snippet(s)?”.)
Effective perimeter protection mechanisms for industrial sites and critical infrastructure must contend with a large variety of potential threats as well as with the fact that normal site activity can be both complex and diverse. This paper documents the development of a system level approach capable of functioning under such challenging conditions. A multi-view tracking system is used to provide real-time site wide trajectories of all observed individuals. A Radar-based system is also used for tracking if and when camera coverage of various regions is not available. Track information is then analyzed with respect to articulated motion analysis, complex event analysis and normalcy analysis. In addition, object recognition is used to classify left behind objects using high resolution PTZ imagery. A real-time integrated version of this comprehensive approach to perimeter protection was deployed using a single standard off-the-shelf desktop computer.
In this paper, we describe Object Pixel Mixture Classifiers (OPMCs) which classify an object not only apart from background but also from other objects based on Gaussian Mixture Model (GMM) classification. The proposed OPMC is different from general GMM based classifiers in the respect that novel pairwise threshold is applied for final classification. Pairwise thresholds are different thresholds depending on predicted mixture component index combination by a positive and a negative GMMs. We train the pairwise threshold using discriminative model so that generative GMM can take advantage from it. We demonstrate that OPMCs are robust to noise in train data and can keep tracking objects after missing tracks even with occlusion. Also, we show that OPMCs can generate meaningful blob of object, and can separate the region of objects from merged blobs.
An effective 3D method incorporating user assistance for modeling complex buildings is described. This method utilizes the connectivity and similar structure information among unit blocks in a multi-component building structure, to enable the user to incrementally construct models of many types of buildings. The system attempts to minimize the complexity and the number of user interactions needed to assist an existing automatic system in this task. Several examples are presented that demonstrate significant improvement and efficiency, compared with other approaches and with purely manual systems.
Video surveillance applications such as smart room and security system are prevailing nowadays. Camera calibration information (e.g. camera position, orientation, and focal length) is very useful for various surveillance systems because it can provide scene knowledge and limit search space for object detection or tracking. In this paper, we describe a camera calibration tool that does not require any calibration object or specific geometric objects by using vanishing points. In urban environment, vanishing points are easily obtainable since there exist many parallel lines such as street lines, light poles, buildings, etc in either outdoor or indoor scene images. Experimental results from various surveillance cameras are presented.
Detection of events in the surveillance video selected for TRECVID 2008 is an extremely difficult task. It was not possible for us to evaluate on a variety of events and the entire length of the dataset. We participated in an exploratory task and selected a single “people meeting” event. This event was selected due to its frequency, importance and difficulty. It is actually a collection of similar tasks as there can be several styles of meeting: for example two people coming towards each other, one person waiting for others, a person joining an existing group etc. Difficulty of the task can vary based on the density of the crowd in the scene at the moment and the extent to which the participants are occluded. It is not too useful to just combine the results of all these conditions and simply average; we selected video segments where participants are partly occluded and the crowd density is medium. We first detect and track pedestrians using an existing tracking method; in fact, tracking in this complex environment is the strongest challenge for the task. Then, we detect meeting events based on analysis of the trajectories. We present results which include the precision and recall rates and some graphical results.
A speed performance improved vehicle tracking system on a given set of evaluation videos of a street surveillance system is presented. We implement multi-threading technique to meet the requirement of real-time performance which demanded in the practical surveillance systems. Through multi-threading technique, we can accomplish near real-time performance. An analysis of results is also presented.
The production of geospatial information from overhead imagery is generally a labor-intensive process. Analysts must accurately delineate and extract important features, such as buildings, roads, and landcover from the imagery. Automated feature extraction (AFE) tools offer the prospect of reducing analyst's workload. This paper presents a new tool, called iMVS, for extracting buildings and discusses user testing conducted by the National Geospatial-Intelligence Agency (NGA). Using a semi-automated approach, iMVS processes two or more images to form a set of hypothesized 3-D buildings. When the user clicks on one of the building vertices, the system determines which hypothesis is the best fit and extracts the building. A set of powerful editing tools support rapid clean-up of the extraction, including extraction of complex buildings. User testing of iMVS provides an assessment of the benefits and identifies areas for system improvement.
The rapid increase in the availability of geospatial data has motivated the effort to seamlessly integrate this information into an information-rich and realistic 3D environment. However, heterogeneous data sources with varying degrees of consistency and accuracy pose a challenge to such efforts. We describe the geospatial decision making (GeoDec) system, which accurately integrates satellite imagery, three-dimensional models, textures and video streams, road data, maps, point data and temporal data. The system also includes a glove-based user interface
Accurate 3D building models of a city are useful for a variety of applications such as 2D and 3D GIS, fly-through rendering, and simulation for mission planning. Each application may require different aspects of the model. Different levels of information are computed from different data sources. For example, information such as the roof boundary or building height can be obtained by using aerial images or LIDAR (LIght Detection And Ranging) data. Facade information, such as its texture or detailed 3D structure, can be computed from multiple ground view images. In this thesis, acquisition processes for the different levels of building model from various data sources are integrated by deriving a multi-level representation of 3D building model. The concept of Level Of Detail (LOD) from virtual reality literature is exploited to represent the different levels of knowledge for the 3D building model. The multi-level representations of the 3D building model are defined as followings: Level 1 (Initial 3D building) model, Level 2 (Building facade texture) model, and Level 3 (Detailed facade structure) model. The reconstruction process of more complex level models is aided by simpler level models. The knowledge that building roof is parallel to the ground and its wall is perpendicular to the ground, is used to obtain Level 1 model (simpler model). 3D information of Level I model such as its 3D vertices, boundary lines, and facade faces can be used to estimate the pose of uncalibrated ground view images to acquire Level 2 model (facade texture). The calibrated ground view information from Level 2 and 3D model information from Level 1 are used to reduce user interactions in creating Level 3 model, which capture detailed facade structures such as entry ways, recesses, or window structures. Complex and large area site modeling tasks can be done more efficiently by integrating multiple data sources. The final result is a rich multi-level 3D model useful for a variety of applications depending on their needs.
Details of the building facades are needed for high quality fly-through visualization or simulation applications. Windows form a key structure in the detailed facade reconstruction. In this paper, given calibrated facade texture (i.e. the rectified texture), we extract and reconstruct the 3D window structure of the building. We automatically extract windows (rectangles in the rectified image) using a profile projection method, which exploits the regularity of the vertical and horizontal window placement. We classify the extracted windows using 2D dimensions and image texture information. The depth of the extracted windows is automatically computed using window classification information and image line features. A single ground view image is enough to compute 3D depths of the facade windows in our approach.
Modeling and visualization of city scenes is important for many applications including entertainment and urban mission planning. Models covering wide areas can be efficiently constructed from aerial images. However only roof details are visible from aerial views and ground views are needed to provide details of the building facades for high quality fly-through visualization or simulation applications. Different data sources provide different levels of necessary detail knowledge. We need a method that integrates the various levels of data. We propose a hierarchical representation of 3D building models for urban areas that integrates different data sources including aerial and ground view images. Each data source gives us different details and each level of the model has its own application as well. Through the hierarchical representation of 3D building models, large area site modeling can be done efficiently and cost-effectively. This proposal suggests efficient approaches for acquiring each level model and demonstrates some results of each level including the integration results.
3D models of urban sites with geometry and facadetextures are needed for many planning and visualizationapplications. Approximate 3D wireframe model can bederived from aerial images but detailed textures must beobtained from ground level images. Integrating such viewswith the 3D models is difficult as only small parts of buildingsmay be visible in a single view. We describe a methodthat uses two or three vanishing points, and three 3D to2D line correspondences to estimate the rotational andtranslational parameters of the ground level cameras. Thevalid set of multiple combinations of 3D to 2D line pairs ischosen by a hypotheses generation and evaluation Someexperimental results are presented.
Visualization of city scenes is important for many applications including entertainment and urban mission planning. Models covering wide areas can be efficiently constructed from aerial images. However, only roof details are visible from aerial views; ground views are needed to provide details of the building facades for high quality 'fly-through' visualization or simulation applications. We present an automatic method of integrating facade textures from ground view images into 3D building models for urban site modeling. We first segment the input image into building facade regions using a hybrid feature extraction method, which combines global feature extraction with Hough transform on an adaptively tessellated Gaussian Sphere and local region segmentation. We estimate the external camera parameters by using the corner points of the extracted facade regions to integrate the facade textures into the 3D building models. We validate our approach with a set of experiments on some urban sites.
Arnold Smeulders合作论文数Intelligent Systems Lab Amsterdam, Informatics Institute, Faculty of Science, University of Amsterdam3
Cyrus Shahabi合作论文数Department of Computer Science, Viterbi School of Engineering, University of Southern California1