Time‐of‐flight (TOF) sensors provide real‐time depth information at high frame‐rates. One issue with TOF sensors is the usual high level of noise (i.e . the depth measure's repeatability within a static setting). However, until now, TOF sensors’ noise has not been well studied. The authors show that the commonly agreed hypothesis that noise depends only on the amplitude information is not valid in practice. They empirically establish that the noise follows a signal‐dependent Gaussian distribution and varies according to pixel position, depth and integration time. They thus consider all these factors to model noise in two new noise models. Both models are evaluated, compared and used in the two following applications: depth noise removal by depth filtering and uncertainty (repeatability) estimation in three‐dimensional measurement.
Time-Of-Flight (TOF) sensors provide real time depth information at high frame-rates. One issue with TOF sensors is the usually high level of noise (i.e. the depth measure’s repeatability within a static setting). However, until now, TOF sensors’ noise has not been well studied. We show that the commonly agreed hypothesis that noise depends only on the amplitude information is not valid in practice. We empirically establish that the noise follows a signal dependent Gaussian distribution and varies according to pixel position, depth and Integration Time (IT ). We thus consider all these factors to model noise in two new noise models. Both models are evaluated, compared and used in the two following applications: depth noise removal by depth filtering and uncertainty (repeatability) estimation in 3D measurement.
Automatic human action annotation is a challenging problem, which overlaps with many computer vision fields such as video-surveillance, human-computer interaction or video mining. In this work, we offer a skeleton based algorithm to classify segmented human-action sequences. Our contribution is twofold. First, we offer and evaluate different trajectory descriptors on skeleton datasets. Six short term trajectory features based on position, speed or acceleration are first introduced. The last descriptor is the most original since it extends the well-known bag-of-words approach to the bag-of-gestures ones for 3D position of articulations. All these descriptors are evaluated on two public databases with state-of-the art machine learning algorithms. The second contribution is to measure the influence of missing data on algorithms based on skeleton. Indeed skeleton extraction algorithms commonly fail on real sequences, with side or back views and very complex postures. Thus on these real data, we offer to compare recognition methods based on image and those based on skeleton with many missing data.
Time-of-Flight (TOF) cameras are active real time depth sensors. One issue of TOF sensors is measurement noise. In this paper, we present a method for providing the uncertainty associated to 3D TOF measurements based on noise modelling. Measurement uncertainty is the combination of pixel detection error and sensor noise. First, a detailed noise characterization is presented. Then, a continuous model which gives the noise's standard deviation for each depth-pixel is proposed. Finally, a closed-form approximation of 3D uncertainty from 2D pixel detection error is presented. An applicative example is provided that shows the use of our 3D uncertainty modelling on real data.
Cet article adresse le probleme de suivi automatique de pietons au travers de reseaux de cameras a champs de vue disjoints. Le suivi dans l'image est traite de maniere locale par un algorithme de Suivi-par-Detections et re-identification. Avec du filtrage particulaire a etat mixte, nous introduisons la notion d'identite globale dans un algorithme de suivi multi-pistes pour caracteriser les personnes au niveau du reseau et pallier aux discontinuites d'observations. Nous venons renforcer la decision de re-identification en proposant un schema decisionnel haut niveau integrant les hypotheses de chaque traqueur confrontees a la topologie du reseau. La composante suivi multi-personnes et re-identification est d'abord testee en contexte monocamera. Nous evaluons ensuite notre approche complete sur un reseau de 3 cameras a champs de vue disjoints et un ensemble de 7 personnes. La seule connaissance a priori requise est la carte topologique du reseau.
This article tackles the problem of automatic multi-pedestrian tracking in non-overlapping fields of view camera networks, using monocular, uncalibrated cameras. Tracking is locally addressed by a Tracking-by-Detection and reidentification algorithm. We propose here to introduce the concept of global identity into a multi-target tracking algorithm, qualifying people at the network level, to allow us to rebound observation dis-continuities. We embed that identity into the tracking loop thanks to the mixed-state particle filter framework, thus including it in the search space. Doing so, each tracker maintains a mutli-modality on the identity in the network of its target. We increase the decision strength introducing a high level decision scheme which integrates all the trackers hypothesis over all the cameras of the network with previous reidentification results and the topology of the network. The tracking and reidentification module is first tested with a single camera. We then evaluate the whole framework on a 3 non-overlapping fields of view network with 7 identities. The only a priori knowledge assumed is a topological map of the network.
L’analyse video pour la video-surveillance necessite d’avoir une bonne resolution pour pouvoir analyser les flux video avec un maximum de robustesse. Dans le contexte de la detection d’objets stationnaires dans les grandes zones, telles que les parkings, le compromis entre la largeur du champ d’observation et la bonne resolution est difficile avec un nombre limite de cameras. Nous allons utiliser une paire de cameras a focale variable de type Pan-Tilt-Zoom (PTZ). Les cameras parcourent un ensemble de positions (pan, tilt, zoom) predefinies afin de couvrir l’ensemble de la scene a une resolution adaptee. Chacune de ces positions peut etre vue comme une camera stationnaire a tres faible taux de rafraichissement. Dans un premier temps notre approche considere les positions des PTZ comme des cameras independantes. Une soustraction de fond robuste aux changements de luminosite reposant sur une grille de descripteurs SURF est effectuee pour separer le fond du premier plan. La detection des objets stationnaires est effectuee par re-identification des descripteurs a un modele du premier plan. Dans un deuxieme temps afin de filtrer certaines fausses alarmes et pouvoir localiser les objets en 3D une phase de mise en correspondance des silhouettes entre les deux cameras et effectuee. Les silhouettes des objets stationnaires sont placees dans un repere commun aux deux cameras en coordonnees rectifiees. Afin de pouvoir gerer les erreurs de segmentation, des groupes de silhouettes s’expliquant mutuellement et provenant des deux cameras sont alors formes. Chacun de ces groupes (le plus souvent constitue d’une silhouette de chaque camera, mais parfois plus) correspond a un objet stationnaire. La triangulation des points frontiere haut et bas permet ensuite d’acceder a sa localisation 3D et a sa taille.
In this article we propose a novel approach for the detection and localisation of stationary objects using a pair of Pan-Tilt-Zoom (PTZ) cameras monitoring a wide scene. Our contribution is twofold. First we propose a stationary object detection and segmentation technique. It relies on the re-identification of foreground descriptors followed by a segmentation of these regions into objects, using Markov Random Fields. Our method allows the foreground to be dated and, under some conditions, to segment the different objects composing a single foreground blob. The second contribution concerns the matching of object silhouettes detected in each camera. This correspondence stage is only based on geometric constraints. Finally we tested our system on sequences which highlight its robustness to occlusions, even in the case of non planar scenes whose geometry is unknown.
We present a pedestrian tracking system that uses re-identification to monitor nonoverlapping cameras. As tracking, re-identification is an assignment problem, the difficulties being to generate an accurate representation and to prune unlikely pairings. The assignments are realised in two stages. First, a Markovian multi-target trackingby-detection framework which includes identification in the search space is run in the cameras. This generates tracks in the cameras and a first assignment between them thanks to the local identification. This solution is then optimized globally by a network supervisor benefiting from coarse topology knowledge over a sliding window with MCMC sampling. The tracking results obtained on a large ground-truthed dataset demonstrate the effectiveness of the approach.
Time-of-Flight (TOF) cameras measure, in real-time, the distance between the camera and objects in the scene. This opens new perspectives in different application fields: 3D reconstruction, Augmented Reality, video-surveillance, etc. However, like any sensor, TOF cameras have limitations related to their technology. One of them is distance distortion. In this paper, we present a new depth calibration method (estimation of distance distortion) for TOF cameras. Our approach has several advantages. First, it is based on a non-parametric model, contrary to most of the other methods. Second, it models under the same formalism the distortion variation according to the distance and the pixel position in the image. This improves calibration accuracy even at the image boundaries which are typically more distorted than the image center. A comparison with two state of the art parametric methods is presented.
This article tackles the problem of automatic multi-pedestrian tracking in non overlapping fields of view camera networks, using monocular, uncalibrated cameras. Tracking is locally addressed by a Tracking-by-Detection and reidentification algorithm. We propose here to introduce the concept of global identity into a multi-target tracking algorithm, qualifying people at the network level, to allow us to rebound observation discontinuities. We embed that identity into the tracking loop thanks to the mixed-state particle filter framework, thus including it in the search space. Doing so, each tracker maintains a mutli-modality on the identity in the network of its target. We increase the decision strength introducing a high level decision scheme which integrates all the trackers hypothesis over all the cameras of the network with previous reidentification results and the topology of the network. The tracking and reidentification module is first tested with a single camera. We then evaluate the whole framework on a 3 non-overlapping fields of views network with 7 identities. The only a priori knowledge assumed is a topological map of the network.
Depth cameras open new possibilities in fields such as 3D reconstruction, Augmented Reality and video-surveillance since they provide depth information at high frame-rates. However, like any sensor, they have limitations related to their technology. One of them is depth distortion. In this paper, we present a method to estimate depth correction for depth cameras. The proposed method is based on two steps. The first one is a nonplanarity correction that needs depth measurement of different plane views. The second one is an affinity correction that,contrary to state of the art approaches, requires a very small set of ground truth measurements. Thus, it is more easy to use compared to other methods and does not need a large set of accurate ground truth that is extremely difficult to obtain in practice. Experiments on both simulated and real data show that the proposed approach improve also the depth accuracy compare to state of the art methods.
Background subtraction is often one of the first tasks involved in video surveillance applications. Classical methods only use temporal modelling of the background pixels. Using pixel blocks with fixed size allows robust detection but these approaches lead to a loss of precision. We propose in this paper a model of the scene which combines a temporal and local model with a spatial model. This whole representation of the scene both models fixed elements (background) and mobile ones. This allows improving detection accuracy by transforming the detection problem in a two classes classification problem.
The goal of the MobileMii platform is to create a platform for world-class ICT research in ambient intelligence for the comfort and security in the place of life. Activities will build on research topics of the CEA-LIST and the Institut Mines-Telecom. Research will focus on major prospects of ambient intelligence at home and at work such as telecare, the activity monitoring (or surveillance) at the place of life, or smart assistance for work tasks. Specifically, it is to create a platform localized in buildings of NANO-INNOV and DIGITEO on the Plateau de Saclay in the South of Paris. This platform features hundred square meters including development zones, a showroom, and a realistic space for metrology equipment designed in laboratories. MobilMii also addresses a user-driven research whose objective is a short-term return on research investment. MobilMii will demonstrate the industrial relevance, performance of developments, with the idea to influence developments to meet specific needs, to develop resulting technologies developed, thus reducing time to market. This platform will promote exchanges with institutions and the regional community of national and international ambient intelligence community (scientific, industrial users). It will give opportunities for demonstrations in a realistic context and provide support equipment for future research projects. A major challenge of the platform is to develop innovative devices and services. The innovative nature will be generated thanks to the multidisciplinary technical teams, from CEA and Institut Mines-Telecom, acting together on a same place. This will be an opportunity to simultaneously develop services based on several technologies (communications, sensors, HMIs, reasoning, ...)
Pan Tilt Zoom cameras have the ability to cover wide areas with an adapted resolution. Since the logical downside of high resolution is a limited field of view, a guard tour can be used to monitor a large scene of interest. However, this greatly increases the duration between frames associated to a specific location. This constraint makes most background algorithms ineffective. In this article we propose a background subtraction algorithm suitable to cameras with very low frame rate. Its main interest consists in the resulting robustness to sudden illumination changes. The background model which describes a wide scene of interest consisting of a collection of images can thus be successfully maintained. This algorithm is compared with the state of the art and a discussion regarding its properties follows.
Cet article presente une nouvelle approche pour le suivi de personnes par reseau de cameras a champs de vue disjoints. Le probleme du suivi dans l'image est traite par des filtres a particules distribues utilisant un modele de couleurs hierarchiques. La nouveaute de notre approche reside dans l'insertion d'une base de personnes deja rencontrees dans le reseau, dans le formalisme du filtre a particule. Ce faisant, les filtres ne realisent plus seulement une estimation de position dans l'image mais aussi etablissent une identite potentielle pour les cibles, relativement a la base de personnes. Ainsi nous envisageons la re-identification en ligne de personnes pour introduire de la continuite et pouvoir suivre les cibles dans un reseau a champs de vue disjoints. Aucune calibration n'est requise. Nous evaluons notre approche sur un reseau de 5 cameras a champs de vue disjoints et un ensemble de 16 personnes.