
Object detection in large-scale real-world scenes requires efficient multi-class detection approaches. Random forests have been shown to handle large training datasets and many classes for object detection efficiently. The most prominent example is the commercial application of random forests for gaming [37]. In this paper, we describe the general framework of random forests for multi-class object detection in images and give an overview of recent developments and implementation details that are relevant for practitioners.
Conventional rigid structure from motion (SFM) addresses the problem of recovering the camera parameters (motion) and the 3D locations (structure) of scene points, given observed 2D image feature points. In this chapter, we propose a new formulation called Semantic Structure From Motion (SSFM). In addition to the geometrical constraints provided by SFM, SSFM takes advantage of both semantic and geometrical properties associated with objects in a scene. These properties allow to jointly estimate the structure of the scene, the camera parameters as well as the 3D locations, poses, and categories of objects in a scene. We cast this problem as a max-likelihood problem where geometry (cameras, points, objects) and semantic information (object classes) are simultaneously estimated. The key intuition is that, in addition to image features, the measurements of objects across views provide additional geometrical constraints that relate cameras and scene parameters. These constraints make the geometry estimation process more robust and, in turn, make object detection more accurate. Our framework has the unique ability to: i) estimate camera poses only from object detections, ii) enhance camera pose estimation, compared to feature-point-based SFM algorithms, iii) improve object detections given multiple uncalibrated images, compared to independently detecting objects in single images. Extensive quantitative results on three datasets – LiDAR cars, street-view pedestrians, and Kinect office desktop – verify our theoretical claims.
The pyramid transform compresses images while preserving global features such as edges and segments. The pyramid transform is efficiently used in optical flow computation starting from planar images captured by pinhole camera systems, since the propagation of features from coarse sampling to fine sampling allows the computation of both large displacements in low-resolution images sampled by a coarse grid and small displacements in high-resolution images sampled by a fine grid. The image pyramid transform involves the resizing of an image by downsampling after convolution with the Gaussian kernel. Since the convolution with the Gaussian kernel for smoothing is derived as the solution of a linear diffusion equation, the pyramid transform is performed by applying a downsampling operation to the solution of the linear diffusion equation.
Multiple people tracking consists in detecting the subjects at each frame and matching these detections to obtain full trajectories. In semi-crowded environments, pedestrians often occlude each other, making tracking a challenging task. Tracking methods mostly work with the assumption that each pedestrian moves independently unaware of the objects or the other pedestrians around it. In the real world though, it is clear that when walking in a crowd, pedestrians try to avoid collisions, keep a close distance to a group of friends or avoid static obstacles in the scene. In this paper, we present an approach which includes the interaction between pedestrians in two ways: first, including social and grouping behavior as a physical model within the tracking system, and second, using a global optimization scheme which takes into account all trajectories and all frames to solve the data association problem . Results are presented on three challenging publicly available datasets, showing our method outperforms state-of-the-art tracking systems. We also make a thorough analysis of the effect of the parameters of the proposed tracker as well as its robustness against noise, outliers and missing data.
The US National Academy of Engineering recently identified restoring and improving urban infrastructure as one of the grand challenges of engineering. Part of this challenge stems from the lack of viable methods to map/label existing infrastructure. For computer vision, this challenge becomes “How can we automate the process of extracting geometric, object oriented models of infrastructure from visual data?” Object recognition and reconstruction methods have been successfully devised and/or adapted to answer this question for small or linear objects (e.g. columns). However, many infrastructure objects are large and/or planar without significant and distinctive features, such as walls, floor slabs, and bridge decks. How can we recognize and reconstruct them in a 3D model? In this paper, strategies for infrastructure object recognition and reconstruction are presented, to set the stage for posing the question above and discuss future research in featureless, large/planar object recognition and modeling.
We present a feature-based surveillance pipeline which, in contrast to traditional image-based methods, allows to learn a detailed description of the observed background as well as of foreground objects. The pipeline consists of motion segmentation of feature trajectories and subsequent tracking-by-recognition with updates. Furthermore, 3D object representations are learned in order to extract the 3D object pose of a later object recognition. Finally, we show how such sufficiently reliable information is inputted into a reasoning system comparing actual and nominal condition of an airport apron. By this, automatic situation assessment becomes possible in a manageable and reliable way.
The accuracy of stereo algorithms or optical flow methods is commonly assessed by comparing the results against the Middlebury database. However, equivalent data for automotive or robotics applications rarely exist as they are difficult to obtain. As our main contribution, we introduce an evaluation framework tailored for stereo-based driver assistance able to deliver excellent performance measures while circumventing manual label effort. Within this framework one can combine several ways of ground-truthing, different comparison metrics, and use large image databases.
Evaluation of stereo-analysis algorithms is usually done by analysing the performance of stereo matchers on data sets with available ground truth. The trade-off between precise results, obtained with this sort of evaluation, and the limited amount (in both, quantity and diversity) of data sets, needs to be considered if the algorithms are required to analyse real-world environments. This chapter discusses a technique to objectively evaluate the performance of stereo-analysis algorithms using real-world image sequences. The lack of ground truth is tackled by incorporating an extra camera into a multi-view stereo camera system. The relatively simple hardware set-up of the proposed technique can easily be reproduced for specific applications.
Utilization of camera systems for surveillance tasks (e. g. traffic monitoring) has become a standard procedure and has been in use for over 20 years. However, most of the cameras are operated locally and data analyzed manually. Locally means here a limited field of view and that the image sequences are processed independently from other cameras. For the enlargement of the observation area and to avoid occlusions and non-accessible areas multiple camera systems with overlapping and non-overlapping cameras are used. The joint processing of image sequences of a multi-camera system is a scientific and technical challenge. The processing is divided traditionally into camera calibration, object detection, tracking and interpretation. The fusion of information from different cameras is carried out in the world coordinate system. To reduce the network load, a distributed processing concept can be implemented. Object detection and tracking are fundamental image processing tasks for scene evaluation. Situation assessments are based mainly on characteristic local movement patterns (e.g. directions and speed), from which trajectories are derived. It is possible to recognize atypical movement patterns of each detected object by comparing local properties of the trajectories. Interaction of different objects can also be predicted with an additional classification algorithm. This presentation discusses trajectory based recognition algorithms for atypical event detection in multi object scenes to obtain area based types of information (e.g. maps of speed patterns, trajectory curvatures or erratic movements) and shows that two-dimensional areal data analysis of moving objects with multiple cameras offers new possibilities for situational analysis.
When capturing images underwater, image formation is affected in two major ways. First, the light rays traveling underwater are absorbed and scattered depending on their wavelength, creating effects on the image colors. Secondly, the glass interface between air and water refracts the ray entering the camera housing because of a different index of refraction of water, hence the ray is also affected in a geometrical way. This paper examines different camera models and their capabilities to deal with geometrical effects caused by refraction. Using imprecise camera models leads to systematic errors when computing 3D reconstructions or otherwise exploiting geometrical properties of images. In the literature, many authors have published work on underwater imaging by using the perspective pinhole camera model (single viewpoint model - SVP) with a different effective focal length and distortion to compensate for the error induced by refraction at the camera housing. On the other hand, methods were proposed, where refraction is modeled explicitly or where generic, non-single-view-point camera models are used. In addition to discussing all three model categories, an accuracy analysis of using the perspective model on underwater images is given and shows that the perspective model leads to systematic errors that compromise measurement accuracy.
Several new algorithms for camera-based fall detection have been proposed in the literature recently, with the aim to monitor older people at home so nurses or family members can be warned in case of a fall incident. However, these algorithms are evaluated almost exclusively on data captured in controlled environments, under optimal conditions (simple scenes, perfect illumination and setup of cameras), and with falls simulated by actors. In contrast, we collected a dataset based on real life data, recorded at the place of residence of four older persons over several months. We showed that this poses a significantly harder challenge than the datasets used earlier. The image quality is typically low. Falls are rare and vary a lot both in speed and nature. We investigated the variation in environment parameters and context during the fall incidents. We found that various complicating factors, such as moving furniture or the use of walking aids, are very common yet almost unaddressed in the literature. Under such circumstances and given the large variability of the data in combination with the limited number of examples available to train the system, we posit that simple yet robust methods incorporating, where available, domain knowledge (e.g. the fact that the background is static or that a fall usually involves a downward motion) seem to be most promising. Based on these observations, we propose a new fall detection system. It is based on background subtraction and simple measures extracted from the dominant foreground object such as aspect ratio, fall angle and head speed. We discuss the results obtained, with special emphasis on particular difficulties encountered under real world circumstances.
We present a novel approach to segment and classify objects in images into two classes. A binary conditional random field (CRF) framework is augmented with an unsupervised clustering step learning contextual relations of objects, the so-called implicit scene context (ISC). Several experiments with simulated data, images from benchmark data sets, and aerial images of an urban area show improved results compared to a standard CRF.
Robust surface reconstruction from sample points is a challenging problem, especially for real-world input data. We present a new hierarchical surface reconstruction based on volumetric graph-cuts that incorporates significant improvements over existing methods. One key aspect of our method is, that we exploit the footprint information which is inherent to each sample point and describes the underlying surface region represented by that sample. We interpret each sample as a vote for a region in space where the size of the region depends on the footprint size. In our method, sample points with large footprints do not destroy the fine detail captured by sample points with small footprints. The footprints also steer the inhomogeneous volumetric resolution used locally in order to capture fine detail even in large-scale scenes. Similar to other methods our algorithm initially creates a crust around the unknown surface. We propose a crust computation capable of handling data from objects that were only partially sampled, a common case for data generated by multi-view stereo algorithms. Finally, we show the effectiveness of our method on challenging outdoor data sets with samples spanning orders of magnitude in scale.
Human motion capturing (HMC) from multiview image sequences is an extremely difficult problem due to depth and orientation ambiguities and the high dimensionality of the state space. In this paper, we introduce a novel hybrid HMC system that combines video input with sparse inertial sensor input. Employing an annealing particle-based optimization scheme, our idea is to use orientation cues derived from the inertial input to sample particles from the manifold of valid poses. Then, visual cues derived from the video input are used to weight these particles and to iteratively derive the final pose. As our main contribution, we propose an efficient sampling procedure where the particles are derived analytically using inverse kinematics on the orientation cues. Additionally, we introduce a novel sensor noise model to account for uncertainties based on the von Mises-Fisher distribution. Doing so, orientation constraints are naturally fulfilled and the number of needed particles can be kept very small. More generally, our method can be used to sample poses that fulfill arbitrary orientation or positional kinematic constraints. In the experiments, we show that our system can track even highly dynamic motions in an outdoor environment with changing illumination, background clutter, and shadows.
Recent developments in Structure-from-Motion approaches allow the reconstructions of large parts of urban scenes. The available models can in turn be used for accurate image-based localization via pose estimation from 2D-to-3D correspondences. In this paper, we analyze a recently proposed localization method that achieves state-of-the-art localization performance using a visual vocabulary quantization for efficient 2D-to-3D correspondence search. We show that using only a subset of the original models allows the method to achieve a similar localization performance. While this gain can come at additional computational cost depending on the dataset, the reduced model requires significantly less memory, allowing the method to handle even larger datasets. We study how the size of the subset, as well as the quantization, affect both the search for matches and the time needed by RANSAC for pose estimation.
Literally thousands of articles on optical flow algorithms have been published in the past thirty years. Only a small subset of the suggested algorithms have been analyzed with respect to their performance. These evaluations were based on black-box tests, mainly yielding information on the average accuracy on test-sequences with ground truth. No theoretically sound justification exists on why this approach meaningfully and/or exhaustively describes the properties of optical flow algorithms. In practice, design choices are often made based on unmotivated criteria or by trial and error. This article is a position paper questioning current methods in performance analysis. Without empirical results, we discuss more rigorous and theoretically sound approaches which could enable scientists and engineers alike to make sufficiently motivated design choices for a given motion estimation task.
We introduce an (equi-)affine invariant geometric structure by which surfaces that go through squeeze and shear transformations can still be properly analyzed. The definition of an affine invariant metric enables us to evaluate a new form of geodesic distances and to construct an invariant Laplacian from which local and global diffusion geometry is constructed. Applications of the proposed framework demonstrate its power in generalizing and enriching the existing set of tools for shape analysis.
We present a generalized subgraph preconditioning (GSP) technique to solve large-scale bundle adjustment problems efficiently. In contrast with previous work which uses either direct or iterative methods as the linear solver, GSP combines their advantages and is significantly faster on large datasets. Similar to [11], the main idea is to identify a sub-problem (subgraph) that can be solved efficiently by sparse factorization methods and use it to build a preconditioner for the conjugate gradient method. The difference is that GSP is more general and leads to much more effective preconditioners. We design a greedy algorithm to build subgraphs which have bounded maximum clique size in the factorization phase, and also result in smaller condition numbers than standard preconditioning techniques. When applying the proposed method to the "bal" datasets [1], GSP displays promising performance.
This paper describes an approach for Structure from Motion (SfM) for wide baselines image sets and its combination with the dense Semiglobal Matching (SGM) 3D reconstruction approach. Our approach for SfM relies on given information concerning image overlap, but can deal with large baselines and produces highly precise camera parameters and 3D points. At the core of our contribution is robust least squares adjustment with full exploitation of the covariance information from affine point matching to bundle adjustment. Reweighting for robust adjustment is based on covariance information for each individual residual. We use points detected based on Differences of Gaussians including scale and orientation information as well as a variant of the five point algorithm. A strategy similar to the Expectation Maximization (EM) algorithm is employed to extend partial solutions. The key characteristics of the approach is reliability obtained by aiming at a high precision in every step. The capabilities of our approach are demonstrated by presenting results for sets consisting of images from the ground and from small Unmanned Aircraft Systems (UASs).
In 3-source photometric stereo, a Lambertian surface is illuminated from 3 known independent light-source directions, and photographed to give 3 images. The task of recovering the surface reduces to solving systems of linear equations for the gradients of a bivariate function u whose graph is the visible part of the surface [9], [16], [17], [24]. In the present paper we consider the same task, but with slightly more realistic assumptions: the photographic images are contaminated by Gaussian noise, and light-source directions may not be known. This leads to a non-quadratic optimization problem with many independent variables, compared to the quadratic problems resulting from addition of noise to the gradient of u and solved by linear methods in [6], [10], [20], [21], [22], [25]. The distinction is illustrated in Example below. Perhaps the most natural way to solve our problem is by global Gradient Descent, and we compare this with the 2-dimensional Leap-Frog Algorithm [23]. For this we review some mathematical results of [23] and describe an implementation in sufficient detail to permit code to be wrtten. Then we give examles comparing the behavior of Leap-Frog with GradientDescent, and explore an extension of Leap-Frog (not covered in [23]) to estimate light source directions when these are not given, as well as the reflecting surface.