Supplying realistically textured 3D city models at ground level promises to be useful for pre-visualizing upcoming traffic situations in car navigation systems. Because this pre-visualization can be rendered from the expected future viewpoints of the driver, the required maneuver will be more easily understandable. 3D city models can be reconstructed from the imagery recorded by surveying vehicles. The vastness of image material gathered by these vehicles, however, puts extreme demands on vision algorithms to ensure their practical usability. Algorithms need to be as fast as possible and should result in compact, memory efficient 3D city models for future ease of distribution and visualization. For the considered application, these are not contradictory demands. Simplified geometry assumptions can speed up vision algorithms while automatically guaranteeing compact geometry models. In this paper, we present a novel city modeling framework which builds upon this philosophy to create 3D content at high speed. Objects in the environment, such as cars and pedestrians, may however disturb the reconstruction, as they violate the simplified geometry assumptions, leading to visually unpleasant artifacts and degrading the visual realism of the resulting 3D city model. Unfortunately, such objects are prevalent in urban scenes. We therefore extend the reconstruction framework by integrating it with an object recognition module that automatically detects cars in the input video streams and localizes them in 3D. The two components of our system are tightly integrated and benefit from each other’s continuous input. 3D reconstruction delivers geometric scene context, which greatly helps improve detection precision. The detected car locations, on the other hand, are used to instantiate virtual placeholder models which augment the visual realism of the reconstructed city model.
This paper deals with a step towards a 3D reconstruction system for city modeling from omnidirectional video sequences using structure from motion together with stereo constraints. We concentrate on two issues. First, we show how the tracking and reconstruction paradigm were adapted to use omnidirectional images taken by lenses with 180 degrees field of view. This concerns mainly camera calibration transforming the pixel locations into rays and solving the minimal problem for 3D-to-2D matches using RANSAC. Secondly, we compare the results of the reconstruction using additional stereo constraints to the results when these constraints are not used and show that they are needed to make the reconstruction stable. Performance of the system is demonstrated on a sequence of 870 images acquired while driving in a city.
3D city modeling using computer vision is very challenging. A typical city contains objects which are a nightmare for some vision algorithms, while other algorithms have been designed to identify exactly these parts but, in their turn, suffer from other weaknesses which limit their application. For instance, moving cars with metallic surfaces can degrade the results of a 3D city reconstruction algorithm which is primarily based on the assumption of a static scene with diffuse reflection properties. On the other hand, a specialized object recognition algorithm could be able to detect cars, but also yields too many false positives without the availability of additional scene knowledge. In this paper, the design of a cognitive loop which intertwines both aforementioned algorithms is demonstrated for 3D city modeling, proving that the whole can be much more than the simple sum of its parts. A cognitive loop is the mutual transfer of higher knowledge between algorithms, which enables the combination of algorithms to overcome the weaknesses of any single algorithm. We demonstrate the promise of this approach on a real-world city modeling task using video data recorded by a survey vehicle. Our results show that the cognitive combination of algorithms delivers convincing city models which improve upon the degree of realism that is possible from a purely reconstruction-based approach.
In this paper, we present a system that integrates fully automatic scene geometry estimation, 2D object detection, 3D localization, trajectory estimation, and tracking for dynamic scene interpretation from a moving vehicle. Our sole input are two video streams from a calibrated stereo rig on top of a car. From these streams, we estimate structure-from-motion (SfM) and scene geometry in real-time. In parallel, we perform multi-view/multi-category object recognition to detect cars and pedestrians in both camera images. Using the SfM self-localization, 2D object detections are converted to 3D observations, which are accumulated in a world coordinate frame. A subsequent tracking module analyzes the resulting 3D observations to find physically plausible spacetime trajectories. Finally, a global optimization criterion takes object-object interactions into account to arrive at accurate 3D localization and trajectory estimates for both cars and pedestrians. We demonstrate the performance of our integrated system on challenging real-world data showing car passages through crowded city areas.
Conference on optical 3-D measurement techniques, Date: 2007/07/09 - 2007/07/12, Location: Zürich, Switzerland. 2007-01,. 8th conference on optical 3-D measurement techniques. 3D city modeling integrating recognition and reconstruction. Author: Cornelis, Nico. Leibe, Bastian ; Cornelis, Kurt ; Van Gool, Luc. Keywords: PSI_VISICS …
This paper presents a practical system for vision-based traffic scene analysis from a moving vehicle based on a cognitive feedback loop which integrates real-time geometry estimation with appearance-based object detection. We demonstrate how those two components can benefit from each other's continuous input and how the transferred knowledge can be used to improve scene analysis. Thus, scene interpretation is not left as a matter of logical reasoning, but is instead addressed by the repeated interaction and consistency checks between different levels and modes of visual processing. As our results show, the proposed tight integration significantly increases recognition performance, as well as overall system robustness. In addition, it enables the construction of novel capabilities such as the accurate 3D estimation of object locations and orientations and their temporal integration in a world coordinate frame. The system is evaluated on a challenging real-world car detection task in an urban scenario.
In many sports and surveillance scenarios the action is dynamic and takes place on a planar surface, while being recorded by two or more zoom-pan-tilt cameras. Although their position is fixed, these cameras can typically rotate and zoom independently from each other. When rotation and zoom of each camera are known, one could reconstruct the dynamic event in 3D and generate different views of the action. Sensors exist which report zoom and orientation changes of pan-tilt units. In absence of such sensors, however, we prove that the varying zoom and rotation of two pan-tilt units can be extracted solely from the planar homography which exists between both cameras.
Nowadays, GPS-based car navigation systems mainly use speech and aerial views of simplified road maps to guide drivers to their destination. However, drivers often experience difficulties in linking the simple 2D aerial map with the visual impression that they get from the real environment, which is inherently ground-level based. Therefore, supplying realistically textured 3D city models at ground-level proves very useful for pre-visualizing an upcoming traffic situation. Because this pre-visualization can be rendered from the expected future viewpoints of the driver, the latter will more easily understand the required maneuver. 3D city models can be reconstructed from the imagery recorded by surveying vehicles. The vastness of image material gathered by these vehicles, however, puts extreme demands on vision algorithms to ensure their practical usability. Algorithms need to be as fast as possible and should result in compact, memory efficient 3D city models for future ease of distribution and visualization. For the considered application, these are not contradictory demands. Simplified geometry assumptions can speed up vision algorithms while automatically guaranteeing compact geometry models. We present a novel city modeling framework which builds upon this philosophy to create 3D content at high speed which could allow for pre-visualization of any conceivable traffic situation by car navigation modules.
The paper presents a method for multi-dimensional registration of two video streams. The sequences are captured by two hand-held cameras moving independently with respect to each other, both observing one object rigidly moving apart from the background. The method is based on uncalibrated Structure-from-Motion (SfM) to extract 3D models for the foreground object and the background, as well as for their relative motion. It fixes the relative scales between the scene parts within and between the videos. It also provides the registration between all partial 3D models, and the temporal synchronization between the videos. The crux is that not a single point on the foreground or background needs to be in common between both video streams. Extensions to more than two cameras and multiple foreground objects are possible.
The 3D reconstruction of scenes containing independently moving objects from uncalibrated monocular sequences still poses serious challenges. Even if the background and the moving objects are rigid, each reconstruction is only known up to a certain scale, which results in a one-parameter family of possible, relative trajectories per moving object with respect to the background. In order to determine a realistic solution from this family of possible trajectories, this paper proposes to exploit the increased linear coupling between camera and object translations that tends to appear at false scales. An independence criterion is formulated in the sense of true object and camera motions being minimally correlated. The increased coupling at false scales can also lead to the destruction of special properties such as planarity, periodicity, etc. of the true object motion. This provides us with a second, 'non-accidentalness' criterion for the selection of the correct motion among the one-parameter family.
In sequential Structure from Motion algorithms for extended image or video sequences, error build up caused by drift poses a problem as feature tracks that normally represent a single scene point will have distinct 3D reconstructions. For the final bundle adjustment to remove this drift, it must be told about these 3D-3D correspondences through a change in the cost function. However, as a bundle adjustment is a nonlinear optimization technique, the drift needs to be removed from the supplied initial solution to allow for convergence of the bundle adjustment to the real global optimum. Before drift can be removed, it has to be detected. This is accomplished through understanding of the long term behavior of drift which leaves 3D reconstructions from short sequences intact. Drift detection boils down to identifying reconstructions of the same scene part that only differ up to a projective transformation. After detection, the drift can be removed from future processed images and an Adapted Bundle Adjustment using correspondences supplied by the drift detection can remove the drift from previous images. Several experiments on real video sequences demonstrate the merit of drift detection and removal.
Relative scale ambiguity is one of the most important problems we encounter while reconstructing dynamic scenes from monocular image sequences. As an example, consider a scene where a car is moving on the road and a moving camera is recording it. By just looking at the images, we can not decide whether it is a real car on the road or a small toy car hovering in front of the camera. In this paper, we first analyze this problem and then present our solutions which are based on the non-accidentalness principle.
Reconstructing 3D scenes with independently moving objects from uncalibrated monocular image sequences still poses serious challenges. One important problem is to find the relative scales between these different reconstructed objects. The perspective reconstruction of a single object can only be known up to a certain scale which results in a one-parameter family of relative scales and trajectories per moving object. The paper formulates this ambiguity and proposes solutions in the vein of "non-accidentalness". Two instantiations of its use are analyzed: planar motion and the 'heading constraint'.
In this paper a complete system to build visual models from camera images is presented. The system can deal with uncalibrated image sequences acquired with a hand-held camera. Based on tracked or matched features the relations between multiple views are computed. From this both the structure of the scene and the motion of the camera are retrieved. The ambiguity on the reconstruction is restricted from projective to metric through self-calibration. A flexible multi-view stereo matching scheme is used to obtain a dense estimation of the surface geometry. From the computed data different types of visual models are constructed. Besides the traditional geometry- and image-based approaches, a combined approach with view-dependent geometry and texture is presented. As an application fusion of real and virtual scenes is also shown.
Until recently, archaeologists have had limited 3D recording options because of the complexity and expense of the necessary recording equipment. We outline a system that helps archaeologists acquire 3D models without using equipment more complex or delicate than a standard digital camera.
Over the years archaeologists have been swift to embrace new advances in technology that allow them to more comprehensively document the results of their work. Today it is commonplace to find information technologies, in the form MS Office-type tools with some CAD and GIS, deployed for primary data capture, analysis, presentation and publication. While these computing technologies can be used effectively to record and interpret archaeological sites, the radical developments in 3D recording, reconstruction and visualisation tools have had relatively limited impact upon the archaeological community. This is unfortunate as these new technologies have the potential to (a) enable the archaeologists to record their unrepeatable experiments to unprecedented levels of accuracy, (b) enable the archaeologists to reconstruct artefacts such as pottery from sherds, textures and sites from different eras (c) visualise the wealth of excavated information in dynamic new ways away from the archaeological site during post-excavation analysis, (d) make this wealth of detail available to the scholarly community as part of the publication process and secure its digital longevity through its deposition in a trusted digital library/archive and (e) communicate the excitement and importance of their archaeological site and its finds to an interested non-academic audience. This paper describes the overall concept of the EU funded project, 3D Measurement and Virtual Reconstruction of Ancient Lost Worlds of Europe (3D MURALE), that has developed and created a set of low-cost multimedia tools for recording, reconstructing, encoding, and visualising archaeological artefacts and site.
In this paper we present an automated processing pipeline that, from a sequence of images, reconstructs a 3D model. The approach is particularly flexible as it can deal with a hand-held camera without the need for an a priori calibration or explicit knowledge about the recorded scene. In a fist stage features are extracted and tracked throughout the sequence. Using robust statistics and multiple view relations the 3D structure of the observed features and the camera motion and calibration are computed. In a second stage stereo matching is used to obtain a detailed estimate of the geometry of the observed scene. The presented approach integrates state-of-the-art algorithms developed in computer vision, computer graphics and photogrammetry. Due to its flexibility during image acqusition, this approach is particularly well suited for application in the field of archaeology and architectural conservation.