Accurate registration between LiDAR (Light Detection and Ranging) point clouds and semantic 3D city models is a fundamental topic in urban digital twinning and a prerequisite for downstream tasks, such as digital construction, change detection, and model refinement. However, achieving accurate LiDAR-to-Model registration at the individual building level remains challenging, particularly due to the generalization uncertainty in semantic 3D city models at the Level of Detail 2 (LoD2). This paper addresses this gap by proposing L2M-Reg, a plane-based fine registration method that explicitly accounts for model uncertainty. L2M-Reg consists of three key steps: establishing reliable plane correspondence, building a pseudo-plane-constrained Gauss-Helmert model, and adaptively estimating vertical translation. Overall, extensive experiments on five real-world datasets demonstrate that L2M-Reg is both more accurate and computationally efficient than current leading ICP-based and plane-based methods. Therefore, L2M-Reg provides a novel building-level solution regarding LiDAR-to-Model registration when model uncertainty is present. The datasets and code for L2M-Reg can be found: https://github.com/Ziyang-Geodesy/L2M-Reg.
Although semantic 3D city models are internationally available and becoming increasingly detailed, the incorporation of material information remains largely untapped. However, a structured representation of materials and their physical properties could substantially broaden the application spectrum and analytical capabilities for urban digital twins. At the same time, the growing number of repeated mobile laser scans of cities and their street spaces yields a wealth of observations influenced by the material characteristics of the corresponding surfaces. To leverage this information, we propose radiometric fingerprints of object surfaces by grouping LiDAR observations reflected from the same semantic object under varying distances, incidence angles, environmental conditions, sensors, and scanning campaigns. Our study demonstrates how 312.4million individual beams acquired across four campaigns using five LiDAR sensors on the Audi Autonomous Driving Dataset (A2D2) vehicle can be automatically associated with 6368 individual objects of the semantic 3D city model. The model comprises a comprehensive and semantic representation of four inner-city streets at Level of Detail (LOD) 3 with centimeter-level accuracy. It is based on the CityGML 3.0 standard and enables fine-grained sub-differentiation of objects. The extracted radiometric fingerprints for object surfaces reveal recurring intra-class patterns that indicate class-dominant materials. The semantic model, the method implementations, and the developed geodatabase solution 3DSensorDB are released under: https://github.com/tum-gis/sensordb
Multi-scale Digital Twins (DTs) of the built environment provide valuable spatial and semantic insights into city assets, supporting urban management, sustainability, and mobility. However, aligning semantic building models with large-scale digital city models remains a significant challenge due to discrepancies in geospatial alignment, spatial resolution, and heterogeneity of data sources. This paper proposes a novel pipeline for the automatic alignment of semantic indoor building models with digital city models. The pipeline comprises two main steps: first, multi-level semantic segmentation of using Artificial Intelligence (AI) models, and second, point cloud registration using the Particle Swarm Optimization (PSO) algorithm, which estimates the transformation parameters required to align the corresponding models. The results of testing the proposed pipeline on real-world data from the Technical University of Munich (TUM) demonstrate the effectiveness of the proposed method in aligning semantic digital models across multiple scales.
Urban Digital Twins (UDTs) have become essential for managing cities and integrating complex, heterogeneous data from diverse sources. Creating UDTs involves challenges at multiple process stages, including acquiring accurate 3D source data, reconstructing high-fidelity 3D models, maintaining models' updates, and ensuring seamless interoperability to downstream tasks. Current datasets are usually limited to one part of the processing chain, hampering comprehensive Urban Digital Twin (UDT)s validation. To address these challenges, we introduce the first comprehensive multimodal Urban Digital Twin benchmark dataset: TUM2TWIN. This dataset includes georeferenced, semantically aligned 3D models and networks along with various terrestrial, mobile, aerial, and satellite observations boasting 32 data subsets over roughly 100,000 m2 and currently 767 GB of data. By ensuring georeferenced indoor-outdoor acquisition, high accuracy, and multimodal data integration, the benchmark supports robust analysis of sensors and the development of advanced reconstruction methods. Additionally, we explore downstream tasks demonstrating the potential of TUM2TWIN, including novel view synthesis of NeRF and Gaussian Splatting, solar potential analysis, point cloud semantic segmentation, and LoD3 building reconstruction. We are convinced this contribution lays a foundation for overcoming current limitations in UDT creation, fostering new research directions and practical solutions for smarter, data-driven urban environments. The project is available under: https://tum2t.win.
3D semantic scene understanding remains a long-standing challenge in the 3D computer vision community. One of the key issues pertains to limited real-world annotated data to facilitate generalizable models. The common practice to tackle this issue is to simulate new data. Although synthetic datasets offer scalability and perfect labels, their designer-crafted scenes fail to capture real-world complexity and sensor noise, resulting in a synthetic-to-real domain gap. Moreover, no benchmark provides synchronized real and simulated point clouds for segmentation-oriented domain shift analysis. We introduce TrueCity, the first urban semantic segmentation benchmark with cm-accurate annotated real-world point clouds, semantic 3D city models, and annotated simulated point clouds representing the same city. TrueCity proposes segmentation classes aligned with international 3D city modeling standards, enabling consistent evaluation of synthetic-to-real gap. Our extensive experiments on common baselines quantify domain shift and highlight strategies for exploiting synthetic data to enhance real-world 3D scene understanding. We are convinced that the TrueCity dataset will foster further development of sim-to-real gap quantification and enable generalizable data-driven models. The data, code, and 3D models are available online: https://tum-gis.github.io/TrueCity/
Semantic 3D city models are worldwide easy-accessible, providing accurate, object-oriented, and semantic-rich 3D priors. To date, their potential to mitigate the noise impact on radar object detection remains under-explored. In this paper, we first introduce a unique dataset, RadarCity, comprising 54K synchronized radar-image pairs and semantic 3D city models. Moreover, we propose a novel neural network, RADLER, leveraging the effectiveness of contrastive self-supervised learning (SSL) and semantic 3D city models to enhance radar object detection of pedestrians, cyclists, and cars. Specifically, we first obtain the robust radar features via a SSL network in the radar-image pretext task. We then use a simple yet effective feature fusion strategy to incorporate semantic-depth features from semantic 3D city models. Having prior 3D information as guidance, RADLER obtains more fine-grained details to enhance radar object detection. We extensively evaluate RADLER on the collected RadarCity dataset and demonstrate average improvements of 5.46 in mean avarage precision (mAP) and 3.51 previous radar object detection methods. We believe this work will foster further research on semantic-guided and map-supported radar object detection. Our project page is publicly available athttps://gpp-communication.github.io/RADLER .
It is of significant importance to ensure the secure and efficient movement of both pedestrians and vehicles in order to develop an accessible urban environment. This research aims to generate a CityGML model from point clouds, considering the use of curbsides by both pedestrians and vehicles. Mobile Laser Scanning (MLS) and Handheld Mobile Laser Scanning (HMLS) point clouds from three different cities were employed for the automatic classification of curbside areas, including parking for various users, parking by time, parking entrances, garbage bin spaces, and terrace areas. Subsequently, the point clouds were processed for modeling in accordance with the international OGC standard, CityGML version 3.0. The findings demonstrate an overall accuracy of 0.86 in correctly classifying curbside elements in comparison to ground truth data. Furthermore, the point-to-point analysis indicated an F1-score exceeding 0.8 across categories and an IoU mean of 0.8, which serves to underscore the effectiveness of the method. This approach directly generates a semantic 3D streetspace model of the curbside in CityGML from point clouds, ensuring standardized and interoperable access to the data.
Safety concerns remain a barrier to the widespread adoption of cycling. Assessing cycling safety facilitates planning safer cycling routes, which helps to boost cycling confidence. However, existing research primarily concentrates on assessing cycling safety at regional or urban levels, with few studies assessing safety at the road segment level, often without considering detailed lane information. This article leverages the complementary strengths of OSM's rich semantic information on roads and CityGML with lane-level geometry to facilitate cycling safety assessment. Precisely, an informed map matching using Kernel Density Estimation (KDE) for bidirectional attribute transfer, cycling safety scores calculation, and CityGML enrichment with cycling safety are introduced in detail. OpenDRIVE data from the Test Track for Autonomous and Connected Driving (TAVF) in Hamburg, Germany, was converted to a CityGML 3.0-compliant structure using the r:tr & aring;n tool and used together with the corresponding OSM data for experimental analysis. The experimental results show that integrating OSM and CityGML is conducive to improving cycling safety assessment at the road segment level. The assessment results are further embedded into bicycle-related semantics within CityGML 3.0 for the subsequent 3D representation of cycling safety, paving the way for safest path navigation and enhanced perception of cycling safety in 3D environments.
Large language models (LLMs), such as OpenAI's Generative Pre-trained Transformer (GPT), commonly known as ChatGPT has witnessed a very rapid evolution which has opened the door for new possibilities across various industries and academic fields. These advanced technologies are transforming how we view and interact with data, how we communicate and solve complex problems. In this paper, we present a framework that employs LLMs to interact with an Urban Digital Twin (UDT) of a district. The framework utilizes the semantic richness of CityGML for representing 3D city models and the SensorThings API for managing dynamic sensor data, allowing users to query and visualize geospatial and dynamic information intuitively. Through experiments with different types of queries from stakeholders, varying from city planners, to utility providers, and citizens, we found that LLMs can effectively translate natural language queries into complex geospatial and temporal operations, narrowing the gap between non-expert users and complex urban datasets into a fine margin. The results shed light on the potential of LLMs to support decision-making in smart city applications.
High-detail semantic 3D building models are frequently utilized in robotics, geoinformatics, and computer vision. One key aspect of creating such models is employing 2D conflict maps that detect openings' locations in building facades. Yet, in reality, these maps are often incomplete due to obstacles encountered during laser scanning. To address this challenge, we introduce FacaDiffy, a novel method for inpainting unseen facade parts by completing conflict maps with a personalized Stable Diffusion model. Specifically, we first propose a deterministic ray analysis approach to derive 2D conflict maps from existing 3D building models and corresponding laser scanning point clouds. Furthermore, we facilitate the inpainting of unseen facade objects into these 2D conflict maps by leveraging the potential of personalizing a Stable Diffusion model. To complement the scarcity of real-world training data, we also develop a scalable pipeline to produce synthetic conflict maps using random city model generators and annotated facade images. Extensive experiments demonstrate that FacaDiffy achieves state-of-the-art performance in conflict map completion compared to various inpainting baselines and increases the detection rate by 22% when applying the completed conflict maps for high-definition 3D semantic building reconstruction. The code is be publicly available in the corresponding GitHub repository: https://github.com/ThomasFroech/InpaintingofUnseenFacadeObjects
Owing to the typical long-tail data distribution issues, simulating domain-gap-free synthetic data is crucial in robotics, photogrammetry, and computer vision research. The fundamental challenge pertains to credibly measuring the difference between real and simulated data. Such a measure is vital for safety-critical applications, such as automated driving, where out-of-domain samples may impact a car's perception and cause fatal accidents. Previous work has commonly focused on simulating data on one scene and analyzing performance on a different, real-world scene, hampering the disjoint analysis of domain gap coming from networks' deficiencies, class definitions, and object representation. In this paper, we propose a novel approach to measuring the domain gap between the real world sensor observations and simulated data representing the same location, enabling comprehensive domain gap analysis. To measure such a domain gap, we introduce a novel metric DoGSS-PCL and evaluation assessing the geometric and semantic quality of the simulated point cloud. Our experiments corroborate that the introduced approach can be used to measure the domain gap. The tests also reveal that synthetic semantic point clouds may be used for training deep neural networks, maintaining the performance at the 50/50 real-to-synthetic ratio. We strongly believe that this work will facilitate research on credible data simulation and allow for at-scale deployment in automated driving testing and digital twinning.
Thermal point clouds integrate thermal radiation and laser point clouds effectively. However, the semantic information for the interpretation of building thermal point clouds can hardly be precisely inferred. Transferring the semantics encapsulated in 3D building models at Level of Detail (LoD)3 has a potential to fill this gap. In this work, we propose a workflow enriching thermal point clouds with the geo-position and semantics of LoD3 building models, which utilizes features of both modalities: model point clouds are generated from LoD3 models, and thermal point clouds are co-registered by coarse-to-fine registration. The proposed method can automatically co-register the point clouds from different sources and enrich the thermal point cloud in facade-detailed semantics. The enriched thermal point cloud supports thermal analysis and can facilitate the development of currently scarce deep learning models operating directly on thermal point clouds.
This study focuses on the integration of System Dynamics (SD), Artificial Intelligence (AI) technology, and 3D Urban Information Modeling (CityGML) in the field of urban planning. It aims to optimize urban layout to cope with rapid urban growth, emphasizing resource allocation through a combination of greedy strategies and AI methods. The core contribution of this paper is to provide spatial allocation of feedback from simulation techniques i.e. SD using the application of genetic algorithms (GA) based on a greedy strategy. The methodology was implemented using a fictitious case study in Berlin, where new students’ dorms are required according to the growth in the number of students. The SD tool is used to simulate the growth over 5 years and determine the number of new dorms required. A genetic algorithm is used to optimize the planning objectives of allocating 16 new dorms in 186 possible locations. The algorithm effectively handles the feedback from SD and applies complex multi-objective optimization problems, addressing the challenge of accommodating a growing student population. It proposes strategic dormitory allocations that balance factors such as cost efficiency, campus accessibility, and proximity to public transportation systems. We use CityGML data to simulate and predict future urban transformation, providing dynamic and realistic representations for urban planning decisions and updating the original city models.
Urban Digital Twins (UDTs) have emerged as essential tools for managing city operations, forming the basis of smart city solutions. They offer a digital representation of the physical urban environment, which supports various city applications such as monitoring mobility, air quality, and modelling simulations. To accurately represent the physical world, UDTs need to be updated continuously to reflect the changes in the urban environment on time. The Internet of Things (IoT) enables real-time data collection to capture these changes. Combined with 3D city models, IoT allows the interactive visualisation of patterns and trends in UDTs. In this study, we conduct investigations on the requirements for the web visualisation of semantic 3D city models enriched with time-dependent properties from IoT and simulation data. We explore the 3D models and IoT data integration requirements, 4D web visualisation design considerations, and the technical implementation requirements for rendering dynamic properties for UDTs applications. The paper also presents a workflow and a web viewer prototype for the 4D visualisation of integrated 3D models and dynamic data.
Structured semantic 3D city models are pivotal in creating urban 3D digital twins. The wide adoption of such models has been primarily enabled by robust, model-based, and automatic 3D reconstruction methods. However, these methods impose requirements on the reconstruction, mainly restricting the solution space to several model types and relying on accurate 2D footprints. Recent research shows that deep-learning-based methods promise highly generic solution space and are footprint-free. Yet, the current training and test datasets are limited, hindering the methods’ development. In this work, we analyze the ubiquity of already existing, open 3D city model datasets and their potential to serve as a large-scale training and test set for 3D reconstruction, where 27 potential dataset collections have been identified. Our review shows that more than 215 million building models are readily available. We firmly believe that this review will facilitate further research on robust automatic 3D city model reconstruction and serve as a reference for benchmarking 3D city models.
The growing demand for sustainable mobility has led to an increased focus on the development and improvement of bicycle infrastructure, especially within cities. However, evaluating the quality of existing or planned bicycle paths is a complex task mostly done manually. This paper presents a novel approach for automatically evaluating the service quality of bicycle paths using parameters derived from semantic 3D city and streetspace models compliant with the international OGC standard CityGML version 3.0. These models contain detailed 3D information with lane-level accuracy, including precise outlines of individual surfaces. This allows for accurate and high-resolution evaluations of changing bicycle path widths and slopes, as well as information on adjacent surfaces and local disturbances such as bus stops. Additionally, estimated, measured or simulated bicycle traffic volumes are considered. Based on these parameters a method for calculating the Bicycle Levels of Service (BLOS) described in a national technical regulation is adapted and implemented for a microscopic analysis. Results of this analysis are then transferred back to the original semantic 3D city objects, allowing for the attributive description of BLOS values for bicycle paths. In addition, results are visually represented by coloring corresponding bicycle path segments according to evaluation results and integrating the colored objects within a web-based Cesium visualization of a semantic 3D city model.
Urban Digital Twins have received significant attention in recent years due to their economic and research importance. Although many definitions exist, the general consensus agrees on a continuous two-way data flow between a physical entity and its virtual counterpart in a digital twin. In the context of smart cities and semantic 3D city models, however, no major breakthrough in realizing such complex change detection and analysis systems has yet been achieved. While several methods for change detection in semantic 3D city models have been proposed, the analysis of found changes, especially the identification of patterns among a large number of changes, has not been given as much attention. Without a proper handling of patterns, it is difficult to provide useful interpretation of changes with respect to stakeholders. Therefore, this research proposes a framework to define, detect and decipher complex semantic change patterns in semantic 3D city models. The approach provides a central rule network to describe aggregation relations between changes as well as methods to identify and capture detected change patterns directly in the graph representation of a city model.
The AI4TWINNING project aims at the automated generation of a system of inter-related digital twins of the built environment spanning multiple resolution scales providing rich semantics and coherent geometry. To this end, an interdisciplinary group of researchers develops a multi-scale, multi-sensor, multi-method approach combining terrestrial, airborne, and spaceborne acquisition, different sensor types (visible, thermal, LiDAR, Radar) and different processing methods integrating top-down and bottom-up AI approaches. The key concept of the project lies in intelligently fusing the data from different sources by AI-based methods, thus closing information gaps and increasing completeness, accuracy and reliance of the resulting digital twins. To facilitate the process and improve the results, the project makes extensive use of informed machine learning by exploiting explicit knowledge on the design and construction of built facilities. The final goal of the project is not to create a single monolithic digital twin, but instead a system of interlinked twins across different scales, providing the opportunity to seamlessly blend city, district and building models while keeping them up-to-date and consistent. As testbed and demonstration scenario serves a urban zone around the city campus of TUM, for which large data sets from various sensors are available.