Lifelong mapping presents unique challenges to household robots which operate in the same environment over long durations. One is the growth of redundant information in the map as it evolves over time, which can easily overwhelm the limited computation resources of a household robot. Another is the possibility of mapping errors. An error in robot pose estimate, which if not corrected fast enough, will result in incorrect occupancy and semantic representation, rendering the map unusable. Finally, for a lifelong mapping system where the map is updated continuously, avoiding these errors altogether is infeasible. In this paper, we present a comprehensive overview of novel strategies for eliminating redundant information from the map and preventing and correcting mapping errors. We also present a detailed evaluation of these novel strategies on 10,000 robots running in indoor environments across different geographic locations of the world to demonstrate map stability and accuracy over time.
A Graph SLAM system is only as good as the edges in its pose graph. Critical mistakes in the generation of these edges can instantly render a map inconsistent, misleading, and ultimately unusable. For a lifelong mapping system, where the map is updated continuously, avoiding these errors altogether is infeasible. Instead, we propose a system for detection of and recovery from severe errors in edge generation. Our system remedies both edges created by view observations and edges created by an odometry motion model. For observation edges, we pair a novel method for monitoring ambiguous views with an intelligent graph-merging algorithm capable of rejecting a relocalization in progress. For motion edges, we propose a qualitative geometric approach for detecting structural aberrations characteristic of odometry failures. We conclude with an analysis of our results based on an empirical study of thousands of robot runs.
Occupancy mapping enables a mobile robot to make intelligent planning decisions to accomplish its tasks. Adaptive local maps is an algorithm which represents the occupancy information as a set of overlapping local maps anchored to poses in the robot's trajectory. At any time, a global occupancy map can be rendered from the local maps to be used for path planning. The advantage of this approach is that the occupancy information stays consistent despite the changes in the pose estimates resulting from loop closures and localization updates. The disadvantage, however, is that the number of local maps grows over time. For long robot runs, or for multiple runs in the same space, this growth will result in redundant occupancy information, which will in turn increase the time it takes to render the global map, as well as the memory footprint of the system. In this paper, we propose a novel approach for the maintenance of an adaptive local maps system, which intelligently prunes redundant local maps, ensuring the robustness and stability required for lifelong mapping.
The time complexity of making observations and loop closures in a graph-based visual SLAM system is a function of the number of views stored [1], [2]. Clever algorithms, such as approximate nearest neighbor search, can make this function sub-linear. Despite this, over time the number of views can still grow to a point at which the speed and/or accuracy of the system becomes unacceptable, especially in computation- and memory-constrained SLAM systems. However, not all views are created equal. Some views are rarely observed, because they have been created in an unusual lighting condition, or from low quality images, or in a location whose appearance has changed. These views can be removed to improve the overall performance of a SLAM system. In this paper, we propose a method for pruning views in a visual SLAM system to maintain its speed and accuracy for long term use.
The stage-by-stage modeling of the automobile body by the multiparameter equations with alphabetic parameters for geometrical characteristics and a technique of the surfaces equations of construction with continuous curvature function with the help of R-functions is considered in this work. The work of new high-speed system of the equations of geometrical objects surfaces visualization in 3D is illustrated.
Visual search experiments in static displays have long established that size, color, and orientation are elementary features whose attributes are processed in parallel and available to guide the deployment of attention. Using a gaze-tracked flicker paradigm for change blindness and stimuli rendered identically in space and separately in the 3 feature dimensions, we investigate whether and how these features distinguish themselves in the active deployment of attention during prolonged visual search. We find out that visual search does not show any attentional modulation in orientation, whereas it engages spatial attention in color with shorter saccades between the same color, and it engages featural attention in size with shorter fixation from previewing the same size as well as tuning into a particular size. Thus, in terms of dynamic attribute processing over time, size, color, and orientation are highly distinctive: Between successive fixations, only orientation is truly pre-attentive without any form of priming, whereas size and color deploy attention in the featural and spatial domains respectively.
We propose a novel object detection approach that combines the discriminative power of object category classifiers with a simple pixel level focus of attention mechanism. The pixel-level foreground/background detectors evolve to classify each pixel as either being part of an object of interest or noise. Unlike background subtraction algorithms, the decision of what is foreground is influenced by object level knowledge rather than it being an outlier to a background distribution. The approach outperforms many background subtraction techniques in challenging scenarios. Combined with the proposed focus of attention mechanism, a robust object classifier(capable of classifying known objects or rejecting noise) runs in real-time while processing 1920x1080 videos on an off-the-shelf DSP.
We extend the application of block-floating point arrays to multi-operand algebraic expressions consisting of additions and multiplications. The proposed method enables automatic runtime calculation of binary shifts of array elements. The shifts are computed for all elementary operations in an expression using a dataflow graph. The method attempts to preserve accuracy across the entire expression while avoiding overflow. A variety of common computer vision and image processing operations can be efficiently realized in fixed-point processors using the proposed technique. It eliminates the need to hand-craft block-floating point implementations for each new operation or processor. The result is a reduction in development time and the likelihood of errors.
Camera system (100) with: an image capture device (102) having a field of view that generates image data representing a plurality of images of the field of view; an object detection module (204) coupled to the image capture device (102) and receiving the image data, the object detection module (204) operable to detect objects appearing in one or more of the plurality of images; an object tracking module (206) coupled to the object detection module (204) operable to time associate instances of a first object detected in a first group of the plurality of images, the first object having a first signature, the features of the first object derived from the images of the first group; and a match classifier (218) coupled to the object tracking module (206) for comparing object cases, wherein the match classifier (218) is operable to resolve object cases by analyzing data from the first signature of the first object and a second signature of a first second object ...
The goal of lossy image compression ought to be reducing entropy while preserving the perceptual quality of the image. Using gaze-tracked change detection experiments, we discover that human vision attends to one scale at a time. This evidence suggests that saliency should be treated on a per-scale basis, rather than aggregated into a single 2D map over all the scales. We develop a compression algorithm which adaptively reduces the entropy of the image according to its saliency map within each scale, using the Laplacian pyramid as both the multiscale decomposition and the saliency measure of the image. We finally return to psychophysics to evaluate our results. Surprisingly, images compressed using our method are sometimes judged to be better than the originals.
Many classification techniques expect class instances to be represented as feature vectors, i.e. points in a feature space. In computer vision classification problems, it is often possible to generate an informative feature vector representation of an image, for example using global texture or shape descriptors. However, in other cases, it may be beneficial to treat images as variable size unordered sets or bags of features, in which each feature represents a localized salient image structure or patch. These local features do not require a segmentation, and can be useful for object recognition in the presence of occlusion and clutter. The local features are often used to find point correspondences between images to be later used for 3D reconstruction, object recognition, detection, or image retrieval. However, there are many cases when exact correspondences are difficult or even impossible to compute. Furthermore, point correspondences may not be necessary, unless one is interested in recovering the 3D shape of an object. If the correspondences are not computed, then this representation indeed constitutes an unordered set of local features. In this dissertation we present methods for object class recognition using bags of features without relying on point correspondences. We also show that using bags of features and more traditional feature vector representation of images together can improve classification accuracy. We then propose and evaluate several methods of combining the two representations. The proposed techniques are applied to a challenging marine science domain.
Edward M. Riseman合作论文数Manning College of Information & Computer Sciences, University of Massachusetts Amherst3
Mario E. Munich合作论文数Evolution Robotics2