We introduce an active learning framework for content-based image retrieval for video surveillance that can be trained ad-hoc for a single camera in a matter of minutes. This technique allows searching for both, known and unknown objects, given a region of interest. The process does not require prior labelled data and treats image retrieval as a binary classification task, in which frames can be similar or different from a query image. The technique is compatible with any pre-trained deep feature extractor. In addition, we propose a novel label propagation algorithm that benefits from (1) visual similarity of image pairs and (2) the semantic representation of the feature vectors from a pre-trained deep feature extractor. This approach allows to reduce the amount of labels needed, while avoiding the propagation of errors. Our experiments with three use-cases from a nuclear facility show the validity of the proposed method, which achieves high precision and recall while requiring minimal amounts of labelled data.
In this paper, we face the point-cloud segmentation problem for spinning laser sensors from a deep-learning (DL) perspective. Since the sensors natively provide their measurements in a 2D grid, we directly use state-of-the-art models designed for visual information for the segmentation task and then exploit the range information to ensure 3D accuracy. This allows us to effectively address the main challenges of applying DL techniques to point clouds, i.e., lack of structure and increased dimensionality. To the best of our knowledge, this is the first work that faces the 3D segmentation problem from a 2D perspective without explicitly re-projecting 3D point clouds. Moreover, our approach exploits multiple channels available in modern sensors, i.e., range, reflectivity, and ambient illumination. We also introduce a novel data-mining pipeline that enables the annotation of 3D scans without human intervention. Together with this paper, we present a new public dataset with all the data collected for training and evaluating our approach, where point clouds preserve their native sensor structure and where every single measurement contains range, reflectivity, and ambient information, together with its associated 3D point. As experimental results show, our approach achieves state-of-the-art results both in terms of performance and inference time. Additionally, we provide a novel ablation test that analyses the individual and combined contributions of the different channels provided by modern laser sensors.
This paper addresses the challenges of rendering massive indoor point clouds in Virtual Reality. In these kind of visualizations the point of view is never static, imposing the need of a one-shot (i.e. non-iterative) rendering strategy, in contrast with progressive refinement approaches that assume that the camera position does not change between most consecutive frames. Our approach benefits from the static nature of indoor environments to pre-compute a visibility map that enables us to boost real-time rendering performance. The key idea behind our visibility map is to exploit the cluttered topology of buildings in order to effectively cull the regions of the space that are occluded by structural elements such as walls. This does not only improve performance but also the visual quality of the final render, allowing us to display in full detail the space and preventing the user to see the contiguous spaces through the walls. Additionally, we introduce a novel hierarchical data structure that enables us to display the point cloud with a continuous level of detail with a minimal impact on performance. Experimental results show that our approach outperforms state-of-the-art techniques in complex indoor environments and achieves comparable results in outdoor ones, proving the generality of our method.
In this paper we introduce a novel public dataset for developing and benchmarking indoor localization systems. We have selected and 3D mapped a set of representative indoor environments including a large office building, a conference room, a workshop, an exhibition area and a restaurant. Our acquisition pipeline is based on a portable LiDAR SLAM backpack to map the buildings and to accurately track the pose of the user as it moves freely inside them. We introduce the calibration procedures that enable us to acquire and geo-reference live data coming from different independent sensors rigidly attached to the backpack. This has allowed us to collect long sequences of spherical and stereo images, together with all the sensor readings coming from a consumer smartphone and locate them inside the map with centimetre accuracy. The dataset addresses many of the limitations of existing indoor localization datasets regarding the scale and diversity of the mapped buildings; the number of acquired sequences under varying conditions; the accuracy of the ground-truth trajectory; the availability of a detailed 3D model and the availability of different sensor types. It enables the benchmarking of existing and the development of new indoor localization approaches, in particular for deep learning based systems that require large amounts of labeled training data.
This paper presents a new technique to solve the Indoor Visual Place Recognition problem from the Deep Learning perspective. It consists on an image retrieval approach supported by a novel image similarity metric. Our work uses a 3D laser sensor mounted on a backpack with a calibrated spherical camera i) to generate the data for training the deep neural network and ii) to build a database of geo-referenced images for an environment. The data collection stage is fully automatic and requires no user intervention for labelling. Thanks to the 3D laser measurements and the spherical panoramas, we can efficiently survey large indoor areas in a very short time. The underlying 3D data associated to the map allows us to define the similarity between two training images as the geometric overlap between the observed pixels. We exploit this similarity metric to effectively train a CNN that maps images into compact embeddings. The goal of the training is to ensure that the L2 distance between the embeddings associated to two images is small when they are observing the same place and large when they are observing different places. After the training, similarities between a query image and the geo-referenced images in the database are efficiently retrieved by performing a nearest neighbour search in the embeddings space.
We present a robust Global Matching technique focused on 3D mapping applications using laser range-finders. Our approach works under the assumption that places can be recognized by analyzing the projection of the observed points along the gravity direction. Relative poses between pairs of 3D point clouds are estimated by aligning their 2D projective representations and benefiting from the corresponding dimensional reduction. We present the complete processing pipeline for two different applications that use the global matcher as a core component: First, the global matcher is used for the registration of static scan sets where no a-priori information of the relative poses is available. It is combined with an effective procedure for validating the matches that exploits the implicit empty space information associated to single acquisitions. In the second use case, the global matcher is used for the loop detection required for 3D SLAM applications. We use an Extended Kalman Filter to obtain a belief of the map poses, which allows to validate matches and to execute hierarchical overlap tests, which reduce the number of potential matches to be evaluated. Additionally, the global matcher is combined with a fast local technique. In both use cases, the global reconstruction problem is modeled as a sparse graph, where scan poses (nodes) are connected through matches (edges). The graph structure allows formulating a sparse global optimization problem that optimizes scan poses, considering simultaneously all accepted matches. Our approach is being used in production systems and has been successfully evaluated on several real and publicly available datasets.
Detailed ground truth measurements and error visualization for each team, as well as the 3D point cloud of the evaluation area related to the Microsoft Indoor Localization Competition 2018. Additional details can be found here: https://www.microsoft.com/en-us/research/event/microsoft-indoor-localization-competition-ipsn-2018/
We present a robust Global Registration technique focused on environment survey applications using laser range-finders. Our approach works under the assumption that places can be recognized by analyzing the projection of the observed points along the gravity direction. Candidate 3D matches are estimated by aligning the 2D projective representations of the acquired scans, and benefiting from the corresponding dimensional reduction. Each single candidate match is then validated exploiting the implicit empty space information associated to scans. The global reconstruction problem is modeled as a directed graph, where scan poses (nodes) are connected through matches (edges). This is exploited to compute local matches (instead of global ones) between pairs of scans that are in the same reference frame. As a consequence, both performance and recall ratio increase w.r.t. using only global matches. Additionally, the graph structure allows formulating a sparse global optimization problem that optimizes scan poses, considering simultaneously all accepted matches. Our approach is being used in production systems and has been successfully evaluated on several real datasets.
•We face the ego-motion estimation and localization in known large environments.•A portable 3D sensor is used to solve the place recognition and tracking problems.•An efficient search space reduction technique is proposed.•Global localization is addressed using a robust place recognizer.•A tracking algorithm is introduced to update the sensor pose as it moves.
Precise 3D mapping and 6DOF trajectory estimation using exteroceptive sensors are key problems in many fields. Real-time moving laser sensors gained popularity due to their precise depth measurements, high frame rate and large field of view. We propose an optimization framework for Simultaneous Localization And Mapping that properly models the acquisition process in a scanning-while-moving scenario. Each measurement is correctly reprojected in the map reference frame by considering a continuous time trajectory which is defined as the linear interpolation of a discrete set of control poses in SE3. The trajectory estimation is performed using the sensor readings only, i.e., no external motion measurement units are used. An efficient data structure that makes use of a hybrid sparse voxelized representation for large map management allows to perform global optimization over trajectories, resetting the accumulated drift when loops are detected. We experimentally show that such framework improves localization and mapping w.r.t. solutions that compensate the distortion effects without including them in the optimization step. Moreover, we show that the proposed map structure provides linear or constant time operations w.r.t. the map size in order to perform real time SLAM and it can handle very large maps.
We present a portable system for tracking and mapping using a real time 3D sensor. In particular our current implementation is composed by a spinning laser sensor mounted on a portable backpack that sends positioning and acquisition data to an hand-held tablet. The system can be directly employed as a real time odometer system if no information on the environment are available or may take advantage of the known map. In the latter case an initial automatic off-line stage builds a 3D map of an unknown environment (either using high definition laser scanner acquisitions or our system alone). A second off-line stage extracts useful information from the generated map that are used in the subsequent on-line tracking stage. Finally, real-time pose tracking is performed using a robust ICP implementation that efficiently selects potentially descriptive points, removes outliers and that fuses a local odometer to allow the user to navigate through non-mapped areas. Our system provides accurate real time positioning information in large indoor environments and, optionally, overlays the current real time 3D acquisition on the original map assisting the user in accurately identifying regions of the environment that undergone changes.
Ferns ensembles offer an accurate and efficient multiclass non-linear classification, commonly at the expense of consuming a large amount of memory. We introduce a two-fold contribution that produces large reductions in their memory consumption. First, an efficient L0 regularised cost optimisation finds a sparse representation of the posterior probabilities in the ensemble by discarding elements with zero contribution to valid responses in the training samples. As a by-product this can produce a prediction accuracy gain that, if required, can be traded for further reductions in memory size and prediction time. Secondly, posterior probabilities are quantised and stored in a memory-friendly sparse data structure. We reported a minimum of 75% memory reduction for different types of classification problems using generative and discriminative ferns ensembles, without increasing prediction time or classification error. For image patch recognition our proposal produced a 90% memory reduction, and improved in several percentage points the prediction accuracy.
We describe a method to identify ambiguous poses during tracking and localization based on depth sensors. In particular we distinguish between ambiguities related to specific acquisitions (track ambiguities) that hinder a good registration of the current pose with previous acquisitions and ambiguities related to repetitive elements observed in a particular (known) environment visited (map ambiguities). We propose a measure of both types of ambiguities to scale tracking and localization problems to large environments and to obtain more accurate results. We also propose a two level classifier that firstly labels an input observation as ambiguous or not, and then provides a prediction of candidate poses from the subset of unambiguous ones. We show that by identifying such poses, real time SLAM systems can reduce processing time in the real-time relocalization step. Furthermore, it permits the generation of more compact, highly-discriminative relocalization classifiers. We combine these proposals and use them as a proof of concept. Our preliminary results on real datasets justify the integration in SLAM or Ego-Motion pipelines of such concepts.
We describe a method for non-invasive, accurate and efficient 3D reconstruction of occluded scenes, from a minimal number of X-ray and range scan image acquisitions. The residuals of generalised epipolar constraints (GEC) are incorporated in a highly efficient bundle adjustment minimization, to obtain maximum likelihood estimations of the X-ray image calibration parameters from correspondences between scene points, image points and apparent contours of scene objects. Furthermore, we propose a multimodal template adequate for accurate joint calibration of X-ray and range scan images. It offers crucial advantages for security applications, such as minimal scene occlusion and an agile data acquisition. Finally we describe a shape-from-silhouettes method based on the state of the art, able to reconstruct scene objects with general 3D shapes. We combine these proposals in a full system for 3D reconstruction of occluded scenes, and use it to demonstrate the practical and computational advantages of the methods herein described, with respect to previous proposals, using both synthetic and real data experiments.
Closed receptacle contents can be reconstructed without invasive approaches by coupling X-ray and 3D data acquisitions. We present a solution to the reconstruction problem that exploits 3D data acquisitions of the environment to initialize the X-ray projection calibration and to align the final reconstruction to a global frame. In particular we show that the problem can be solved, without employing calibration patterns, by precisely locating the X-ray emitters and receiver plates w.r.t. a global reference frame. We then exploit matching contours extracted from the X-ray images to calibrate the projection system and to perform the reconstruction of parametric objects. We provide experimental results to measure the reconstruction precision of the overall system using both synthetic data and real experiments.
This paper presents a new robotic device whose main objective is to increase the flexibility of current systems used to carry out inspection and manipulation of spent nuclear fuel inside dry storage unit cells. Instead of using a rigid kinematic chain, the proposed device is based on the idea of a capsule that can be deployed inside a cell with the appropriate tool by means of a conventional crane and cable. The device is equipped with extendable actuators that immobilize the whole system inside the cell by pushing against the walls. This provides a stable platform for those tools that have to execute precision tasks. The system also comprises an industrial tool changer that allows the deployment of the required tool for each specific task. A first prototype of the device and a tool to retrieve spent nuclear fuel were built, and the results of the initial experiments are reported at the end of this paper.