
This paper addresses the problem of detecting and tracking a large number of individuals in aerial image sequences that have been taken from high altitude. We propose a method which can handle the numerous challenges that are associated with this task and demonstrate its quality on several test sequences. Moreover this paper contains several contributions to improve object detection and tracking in other domains, too. We show how to build an effective object detector in a flexible way which incorporates the shadow of an object and enhanced features for shape and color. Furthermore the performance of the detector is boosted by an improved way to collect background samples for the classifier training. At last we describe a tracking-by-detection method that can handle frequent misses and a very large number of similar objects.
This paper reports an investigation of the effect of digitization on the measurement accuracy of the center location of a circle by a centroid method. Although general expressions representing the measurement accuracy of the center location of a circle by the centroid method are unable to be obtained analytically, we have succeeded in obtaining the variances V of measurement errors for 39 quantization bits n ranging from one to infinity by numerical integration. We have succeeded in obtaining the effective approximation formulae of V as a function of the diameter d of the circle for any n as well. The results show that V would oscillate on an approximate one-pixel cycle in d for any n and decrease as n increases. The differences of V among the different n would be negligible when n ≥ 6 . Some behaviors of V with an increase in n are demonstrated.
We propose a context-based classification method for point clouds acquired by full waveform airborne laser scanners. As these devices provide a higher point density and additional information like echo width or type of return, an accurate distinction of several object classes is possible. However, especially in dense urban areas correct labelling is a challenging task. Therefore, we incorporate context knowledge by using Conditional Random Fields. Typical object structures are learned in a training step and improve the results of the point-based classification process. We validate our approach with two real-world datasets and by a comparison to Support Vector Machines and Markov Random Fields.
With the availability of high-resolution commercial satellite images, automated analysis and object extraction became even a more important topic in remote sensing. As shadows cover a significant portion of an image, they play an important role on automated analysis. While they degrade performance of applications such as image registration, shadow is an important cue for information such as man-made structures. In this article, a shadow detection algorithm that makes use of near-infrared information in combination with RGB bands is introduced. The algorithm is applied on an application for automated building detection.
Most recent stereo algorithms are designed to perform well on close range stereo datasets with relatively small baselines and good radiometric conditions. In this paper, different matching costs on the Semi-Global Matching algorithm are evaluated and compared using aerial image sequences and satellite images with ground truth. The influence of various cost functions on the stereo matching performance using datasets with different baseline lengths and natural radiometric changes is evaluated. A novel matching cost merging Mutual Information and Census is introduced and shows the highest robustness and accuracy. Our study indicates that using an adaptively weighted combination of Mutual Information and Census as matching cost can improve the peformance of stereo matching for airborne image sequences and satellite images.
In this paper, we present a new approach for event detection of pedestrian interaction in crowded and cluttered scenes. Existing work is focused on the detection of an abnormal event in general or on the detection of specific simple events incorporating only up to two trajectories. In our approach, event detection in large groups of pedestrians is performed by exploiting motion interaction between pairs of pedestrians in a graph-based framework. Event detection is done by analyzing the temporal behaviour of the motion interaction with Hidden Markov Models (HMM). In addition, temporarily unsteady edges in the graph can be compensated by a HMM buffer which internally continues the HMM analysis even if the representing pedestrians depart from each other awhile. Experimental results show the capability of our graph-based approach for event detection by means of an image sequence in which pedestrians approach a soccer stadium.
Multi-camera systems offer some advantages over classical systems like stereo or monocular camera systems. A multi-camera system with a non-overlapping field of view, able to cover a wide area, might prove superior e.g. in a mapping scenario where less time is needed to cover the entire area. Approaches to determine the parameters of the mutual orientation from common motions exist for more than 30 years. Most work presented in the past neglected or ignored the influence different motion characteristics have on the parameter estimation process. However, for critical motions a subset of the parameters of the mutual orientation can not be determined or only very inaccurate. In this paper we present a strategy and assessment scheme to allow a successful estimation of as many parameters as possible even for critical motions. Furthermore, the proposed approach is validated by experiments.
Classification of remote sensing image and range data is normally done in 2D space, because anyhow most sensors capture the surface of the earth from a close-to vertical direction and thus vertical structures, e.g. at building façades are not visible anyways. However, when the objects of interest are photographed from off-nadir directions, like in oblique airborne images, the question on how to efficiently classify those scenes arises. In this paper a study on classification in 3D object space is presented: image features from individual oblique airborne images, and 3D geometric features derived from matching in those images are projected onto voxels. Those are segmented and classified. The study area is Port-Au-Prince (Haiti), where images have been acquired after the earthquakes in January 2010. Results show that through the combination of image evidence as realized by the projection into object space the classification becomes more accurate compared to single image classification.
This paper presents an automatic method for gable roof detection in terrestrial images. The purpose of this study is to refine the roofs of a 3D city model automatically derived from aerial images. The input images consist of geo-referenced terrestrial images acquired by a mobile mapping system (MMS). The raw images have been rectified and merged into seamless façade texture images (one texture per façade). Firstly, each image is pre-processed in order to remove small structures and to smooth homogeneous areas. Secondly, line segments are extracted and analysed to define the lateral edges of the roof. Finally, the analysis of the lowest part of the roof leads to the classification of the roof as gable or non-gable. The method was tested on more than 150 images and shows promising results.
Most approaches use corresponding points to determining an object’s orientation from stereo-images, but this is not always possible. Imaging modalities that do not produce correspondences for different viewing angles, as in X-ray imaging, require other procedures. Our method works on contours in images that do not need to be equivalent in length or contain corresponding points. It is able to determine corresponding contours and resamples those, creating new sets of corresponding points for registration. Two sets of in-plane transformations from a stereo-system are used to determine spatial orientation. The approach was tested with three ground truth datasets and sub-pixel accuracy was achieved. The approach is originally designed for X-ray based patient alignment, but it is versatile and can also be employed in other close range photogrammetry applications.
We propose a surface segmentation method based on Fast Marching Farthest Point Sampling designed for noisy, visually reconstructed point clouds or laser range data. Adjusting the distance metric between neighboring vertices we obtain robust, edge-preserving segmentations based on local curvature. We formulate a cost function given a segmentation in terms of a description length to be minimized. An incremental-decremental segmentation procedure approximates a global optimum of the cost function and prevents from under- as well as strong over-segmentation. We demonstrate the proposed method on various synthetic and real-world data sets.
The combination of LiDAR data with panoramic images could be of great benefit to many geo-related applications and processes such as measuring and map making. Although it is possible to record both LiDAR points and panoramic images at the same time, there are economic and practical advantages to separating the acquisition of both types of data. However, when LiDAR and image data is recorded separately, poor GPS reception in many urban areas will make registration between the data sets necessary. In this paper, we describe a method to register a sequence of panoramic images to a LiDAR point cloud using a non-rigid version of ICP that incorporates a bundle adjustment framework. The registration is then refined by integrating image-to-reflectance data SIFT correspondences into the bundle adjustment. We demonstrate the validity of this registration method by a comparison against ground truth data.
Stereo vision based mobile mapping systems enable the efficient capturing of directly georeferenced stereo pairs. With today's camera and storage technologies imagery can be captured at high data rates resulting in dense stereo sequences. The overlap within stereo pairs and stereo sequences can be exploited to improve the accuracy and reliability of point measurements. This paper aims at robust and accurate monoscopic 3d measurements in future vision-based mobile mapping services. Key element is an adapted Least Squares Matching approach yielding point matching accuracies at the subpixel level. Initial positions for the matching process along the stereo sequence, are obtained by projecting the matched point position within the reference stereo pair to object space and by reprojecting it to the adjacent pairs. Once homologue image positions have been derived, final 3D point coordinates are estimated. Investigations with real-world data show, that points can successfully and reliably be matched over extended stereo sequences.
The main goal of the federal funded project 'LiDAR based biomass assessment' is the nationwide investigation of the biomass potential coming from wood cuttings of non-forest trees. In this context, first and last pulse airborne laserscanning (F+L) data serve as preferred database. First of all, mandatory field calibrations are performed for pre-defined grove types. For this purpose, selected reference groves are captured by full-waveform airborne laserscanning (FWF) and terrestrial laserscanning (TLS) data in different foliage conditions. The paper is reporting about two methods for the biomass assessment of non-forest trees. The first method covers the determination of volume-to-biomass conversion factors which relate the reference above-ground biomass (AGB) estimated from allometric functions with the laserscanning derived vegetation volume. The second method is focused on a 3D Normalized Cut segmentation adopted for non-forest trees and the follow-on biomass calculation based on segmentation-derived tree features.
In recent years, the classification task of building facade images receives a great deal of attention in the photogrammetry community. In this paper, we present an approach for regionwise classification using an efficient randomized decision forest classifier and local features. A conditional random field is then introduced to enforce spatial consistency between neighboring regions. Experimental results are provided to illustrate the performance of the proposed methods using image from eTRIMS database, where our focus is the object classes building, car, door, pavement, road, sky, vegetation, and window.
This paper presents a method to improve the robustness of automated correspondences while also increasing the total amount of measured points and improving the point distribution. This is achieved by incorporating a tiling technique into existing automated interest point extraction and matching algorithms. The technique allows memory intensive interest point extractors like SIFT to use large images beyond 10 megapixels while also making it possible to approximately compensate for perspective differences and thus get matches in places where normal techniques usually do not get any, few, or false ones. The experiments in this paper show an increased amount as well as a more homogeneous distribution of matches compared to standard procedures.
Most of the image registration/matching methods are applicable to images acquired by either identical or similar sensors from various positions. Simpler techniques assume some object space relationship between sensor reference points, such as near parallel image planes, certain overlap and comparable radiometric characteristics. More robust methods allow for larger variations in image orientation and texture, such as the Scale-Invariant Feature Transformation (SIFT), a highly robust technique widely used in computer vision. The use of SIFT, however, is quite limited in mapping so far, mainly, because most of the imagery are acquired from airborne/spaceborne platforms, and, consequently, the image orientation is better known, presenting a less general case for matching. The motivation for this study is to look at the feasibility of a particular case of matching between different image domains. In this investigation, the co-registration of satellite imagery and LiDAR intensity data is addressed.
Statistical background modeling is a standard technique for the detection of moving objects in a static scene. Nevertheless, the stateof-the-art approaches have several lacks for short sequences or quasistationary scenes. Quasi-static means that the ego-motion of the sensor is compensated by image processing. Our focus of attention goes back to the modeling of the pixel process, as it was introduced by Stauffer and Grimson. For quasi-stationary scenes the assignment of a pixel to an origin is uncertain. This assignment is an independent random process that contributes to the gray value. Since the typical update schemes are biased we introduce a novel update scheme based on the join mean and join variance of two independent distributions. The presented method can be seen as an update for the initial guess for more sophisticated algorithms that optimize the spatial distribution.
In this study, we propose a new line matching and reconstruction methodology for aerial image triplets that are acquired within a single strip. The newly developed stereo reconstruction approach gives us better line predictions in the third image which in turn helps to improve the performance of the matching. The redundancy information generated in each stereo match gives us ability to reduce the number of false matches while preserving high levels of matching completeness. The approach is tested over four test patches and produced highly promising line matching and reconstruction results.
The rapid generation of aerial mosaics is an important task for change detection, e.g. in the context of disaster management or surveillance. Unmanned aerial vehicles equipped with a single camera offer the possibility to solve this task with moderate efforts. Unfortunately, the accumulation of tracking errors leads to a drift in the alignment of images which has to be compensated by loop closing for instance. We propose a novel approach for constructing large, consistent and undistorted mosaics by aligning video images of planar scenes. The approach allows the simultaneous closing of multiple loops possibly resulting from the camera path in a batch process. The choice of the adjustment model leads to statistical rigorous solutions while the used minimal representations for the involved homographies and the exploitation of the natural image order enable very efficient computations. The approach will be empirically evaluated with the help of synthetic data and its feasibility will be demonstrated with real data sets.