In the last years, Microsoft has driven innovation in the aerial photogrammetry community. Besides the market leading camera technology, UltraMap has grown to an outstanding photogrammetric workflow system which enables users to effectively work with large digital aerial image blocks in a highly automated way. Best example is the project-based color balancing approach which automatically balances images to a homogeneous block. UltraMap V3 continues innovation, and offers a revolution in terms of ortho processing. A fully automated dense matching module strives for high precision digital surface models (DSMs) which are calculated either on CPUs or on GPUs using a distributed processing framework. By applying constrained filtering algorithms, a digital terrain model can be derived which in turn can be used for fully automated traditional ortho texturing. By having the knowledge about the underlying geometry, seamlines can be generated automatically by applying cost functions in order to minimize visual disturbing artifacts. By exploiting the generated DSM information, a DSMOrtho is created using the balanced input images. Again, seamlines are detected automatically resulting in an automatically balanced ortho mosaic. Interactive block-based radiometric adjustments lead to a high quality ortho product based on UltraCam imagery. UltraMap v3 is the first fully integrated and interactive solution for supporting UltraCam images at best in order to deliver DSM and ortho imagery.
This paper describes in detail the dense matcher developed since years by Vexcel Imaging in Graz for Microsoft’s Bing Maps project. This dense matcher was exclusively developed for and used by Microsoft for the production of the 3D city models of Virtual Earth. It will now be made available to the public with the UltraMap software release mid-2012. That represents a revolutionary step in digital photogrammetry. The dense matcher generates digital surface models (DSM) and digital terrain models (DTM) automatically out of a set of overlapping UltraCam images. The models have an outstanding point density of several hundred points per square meter and sub-pixel accuracy and are generated automatically. The dense matcher consists of two steps. The first step rectifies overlapping image areas to speed up the dense image matching process. This rectification step ensures a very efficient processing and detects occluded areas by applying a back-matching step. In this dense image matching process a cost function consisting of a matching score as well as a smoothness term is minimized. In the second step the resulting range image patches are fused into a DSM by optimizing a global cost function. The whole process is optimized for multi-core CPUs and optionally uses GPUs if available. UltraMap 3.0 features also an additional step which is presented in this paper, a complete automated true-ortho and ortho workflow. For this, the UltraCam images are combined with the DSM or DTM in an automated rectification step and that results in high quality true-ortho or ortho images as a result of a highly automated workflow. The paper presents the new workflow and first results.
Starting 2007 a dense matching algorithm and an automated ortho/true-ortho workflow has been developed by Vexcel Imaging GmbH that has exclusively been used for the automated 3D city model production of Microsoft’s Virtual Earth project and is now also in use for the production of the current BING maps platform. This famous automated workflow has now been disclosed and has been implemented into UltraMap 3.0. This makes the workflow now commercially available for all UltraCam users world-wide. The dense-matching algorithm generates very dense surface models from overlapping aerial images by multi-ray photogrammetry, superior to airborne Lidar collection. Based on the results of the dense matcher, an automated workflow generates ortho images and true-ortho images automatically within UltraMap 3.0. The paper shows the new UltraMap 3.0 workflow, the results of the automated dense matching such as point clouds and DSM and the results of the automated ortho and true-ortho workflow.
http://www.cgv.tugraz.at/cityfit ABSTRACT: Many approaches for automatic 3D city reconstruction exist, but they are still missing an important feature: detailed facades. The goal of the CityFit project is to reconstruct the facades of 80% of the buildings in the city of Graz fully automatically. The challenge is to establish a complete workflow, ranging from acquisition of images and LIDAR data over 2D/3D feature detection and recognition to the generation of lean polygonal facade models. The desired detail level is to represent all significant facade elements larger than 50 cm by explicit polygonal geometry. All geometry shall also carry descriptive annotations (semantic enrichment). This paper presents an outline of the workflow, important design decisions, and the current state of the project. First results were obtained by case studies of facade analysis followed by manual reconstruction. This gave important hints how to structure grammars for automatic reconstruction.
Accurate and realistic building models of urban environments are increasingly important for applications, like virtual tourism or city planning. Initiatives like Virtual Earth or Google Earth are aiming at offering virtual models of all major cities world wide. The prohibitively high costs of manual generation of such models explain the need for an automatic workflow.This paper proposes an algorithm for fully automatic building reconstruction from aerial images. Sparse line features delineating height discontinuities and dense depth data providing the roof surface are combined in an innovative manner with a global optimization algorithm based on Graph Cuts. The fusion process exploits the advantages of both information sources and thus yields superior reconstruction results compared to the indiviual sources. The nature of the algorithm also allows to elegantly generate image driven levels of detail of the geometry.The algorithm is applied to a number of real world data sets encompassing thousands of buildings. The results are analyzed in detail and extensively evaluated using ground truth data.
This paper describes a highly automatic work flow for urban model generation from digital aerial imagery. The images used in our approach feature one high resolution panchromatic channel and four lower resolution channels (red, green, blue and near infrared). In the first step an initial land use classification is performed where a support vector machine is applied to each image. Subsequently the Aerial Triangulation algorithm integrates area- and feature-based points of interest to determine the position and orientation of all images. Afterwards, a digital surface model is generated by a dense image matching procedure. Based on the AT and dense image matching result a true ortho image is computed. Different layers for building blocks, streets, vegetation or water regions are generated from a refined land use classification. The building block layer is used in a subsequent step to extract polygonized buildings.
In this work we propose a scanline optimization procedure for computational stereo using a linear smoothness cost model performed by programmable graphics hardware. The main idea for an efficient implementation of this dynamic programming approach is a recursive scheme to calculate the min-convolution in a manner suitable for the parallel stream computation model of graphics processing units. Since many image similarity functions can be efficiently calculated by modern graphics hardware, it is reasonable to address the final disparity extraction by graphics processors as well. Our timing results indicate that the proposed approach is beneficial for larger image resolutions and disparity ranges in particular.
We present a method to compare computer generated images of a real scene to the corresponding photograph. The approach is based on a numerical comparison of luminance values. Ground-truth data was obtained by mapping the pixel values of the photograph to luminance values via point measurements obtained from an accurate luminance measurement device. We present results of this comparison performed on a very complex indoor scene and conclude with a critical discussion of the problems encountered with our approach and the task of image comparison in general.
This paper proposes a fast 3D reconstruction approach for efficiently generating watertight 3D models from multiple short baseline views. Our method is based on the combination of a GPU-based plane-sweep approach, to compute individual dense depth maps and a subsequent robust volumetric depth map integration technique. Basically, the dense depth map values are transformed to a volumetric grid, which are further embedded in a graph structure. The edge weights of the graph are derived from the dense depth map values and if available, from sparse 3D information. The final optimized surface is obtained as a min-cut/max-flow solution of the weighted graph.We demonstrate the robustness and accuracy of our proposed approach on several real world data sets.
We present a high performance reconstruction approach, which generates true 3D models from multiple views with known camera parameters. The complete pipeline from depth map generation over depth image integration to the final 3D model visualization is performed on programmable graphics processing units (GPUs). The proposed pipeline is suitable for long image sequences and uses a plane- sweep depth estimation procedure optionally employing robust image similarity functions to generate a set of depth images. The subsequent volumetric fusion step combines these depth maps into an impicit surface representation of the final model, which can be directly displayed using GPU- based ray casting methods. Depending on the number of input views and the desired resolution of the final model the computing times range from several seconds to a few minutes. The quality of the obtained models is illustrated with real-world datasets.
This paper introduces a feature based method for the fast generation of sparse 3D point clouds from multiple images with known pose. We extract sub-pixel edge elements (2D position plus associated orientation) and use a space sweeping scheme to compute the accurate 3D location of these edge features. Our approach relies mainly on the geometric properties of the extracted primitives and incorporates a robust uncertainty estimation to detect outliers. Epipolar constraints between views are used to narrow down the search space of potential candidates in the images. In order to improve the efficiency of spatial queries, for detecting edgels lying close to an epipolar line, we utilise a pairwise stereo rectification scheme. The detection and verification of tentative hypotheses is carried out in 3D-space, thus allowing to perform an event driven search mode. An uncertainty measure that models the location inaccuracies in the feature extraction process and errors introduced in the camera pose estimation stage allows to assign a likelihood value to each hypothesis. Optionally an image-based similarity measure can be used to verify the 3D hypotheses and to identify false positives. We perform experiments on a synthetic data set and on several real datasets. The results indicate that the proposed method yields accurate measurements on depth discontinuities and thus represents a complementary technique to standard dense matching approaches.
In this paper we describe the 3D modeling and interactive visualization of 3D city models generated from multispectral digital aerial images. The proposed 3d modeling approach integrates spectral classification to obtain a semantic description of the reconstructed scene. The workflow consists of several consecutive steps, namely an initial land use classification (LUC), the aerial triangulation (AT), a dense matching process, a refined LUC using the DSM to fuse redundant information from all input images and a texture extraction step.