
Abstract. Air pollution is one of the most serious environmental problems in large cities, particularly in developing countries, where it has escalated into a major crisis that endangers public health. To better understand and manage this issue, Computational Fluid Dynamics (CFD) provides an effective tool for analyzing airflow and pollutant dispersion in complex urban settings. This study aims to examine how urban morphology affects ventilation efficiency and pollutant dispersion using three-dimensional (3D) modeling integrated with CFD simulation. The case study focuses on Mashhad, Iran’s second-largest metropolis, within a 300-meter radius of the Avini Air Quality Monitoring Station. First, a detailed 3D model of the study area was created using geospatial building data, including height and geometry. Then, the dispersion of PM₂.₅ and PM₁₀ particles was simulated in the CFD environment based on actual air quality data and prevailing wind conditions. The Residence Time Index (RTI) was calculated to represent the ability of the built environment to retain pollutants. Model validation was performed through comparison with results from a reference study, confirming good agreement and reliability. The findings indicate that areas with higher building density and lower wind velocity tend to trap more pollutants, leading to increased particle accumulation. Variations in building height and differences between PM₁₀ and PM₂.₅ behavior also influence dispersion patterns. Overall, integrating 3D urban modeling with CFD analysis provides a robust approach to understanding pollutant dynamics in urban areas and supports sustainable planning strategies aimed at improving air quality.
Antarctic ice shelves regulate ice sheet mass balance through their "buttressing effect", with major implications for global sea level rise. This study focuses on the Nansen Ice Shelf in Victoria Land, East Antarctica, which exhibits complex topography and sensitivity to environmental changes. Previous research has primarily centered on its significant collapse event in 2016; however, systematic evolutionary patterns over longer timescales remain unclear. This study integrates multi-source remote sensing observations from 1948 to 2025 to systematically reconstruct changes in the Nansen Ice Shelf's geometric characteristics (crevasse width, area) and dynamic parameters (ice flow velocity). Findings reveal distinct activity differences between the northern and southern regions of the ice shelf, closely linked to their respective boundary conditions and structural features.
To plan nature restoration of fluvial corridors on a national level an inventory of existing man-made levees is mandatory. We suggest an automatic method for a river-wise extraction of levees from a high resolution terrain model based on profiles perpendicular to the river axis. This includes a method to cover corridors with non overlapping profiles with a given maximum distance. Levee detection is based on a mathematical formulation of the protective function of levees, i.e. negative elevation differences in outward direction. In an evaluation of 150 km river length distributed over nine different rivers in Austria the method detected 98% of manually extracted levees, and 68% of their length.
Prompted by the rapid advancements in software and hardware, 3D building data for numerous different applications is nowadays often captured via mobile or kinematic laser scanning. However, in contrast to other laser scanning methods, there exist only a few approaches tailored for the planning of a kinematic laser scan survey, and none of them provides an optimality guarantee. Therefore, we propose a novel approach based on Mixed Integer Linear Programming (MILP) to find the optimal trajectory for such a survey. To obtain a high-quality point cloud, we account for scanner-related constraints that influence the quality of the resulting point cloud. Moreover, we enable the introduction of tie points to mitigate the effects of uncertainties in the position estimation that are propagated in the acquired data. In our problem formulation, we aim to find the best tour in a properly weighted graph. For this, we propose two different weight settings to either enable a purely length-based optimization or to increase the redundancy in the measurements by incorporating a Visibility Ratio Factor (VRF) into the objective function.To prove the applicability of our approach for offline panning, we apply our formulation to three different scenarios. In this context, the VRF-based weighting enables a significant speed-up of the solving process while resulting in only slightly prolonged routes. This approach paves the way for applying exact algorithms with an optimality guarantee in the planning process for efficient kinematic laser scanning surveys.
In this paper, a method exploiting aerial LiDAR point clouds to build realistic building meshes suitable for electromagnetic simulation is proposed. One of the main challenges lies in reconstructing regularized building meshes with low polygonal density. Optimization-based methods, commonly used for building reconstruction from point clouds, are highly data-driven, making the quality of results dependent on the quality of input data. Aerial LiDAR scans can be incomplete or sparse, for instance due to occlusion. A novel LoD2 buildings reconstruction method based on deep learning is proposed, assuming that deep learning methods are more robust to incomplete or sparse data than optimization-based methods. A parametric building model is introduced, based on the Weighted Straight Skeleton algorithm, which generates realistic roofs from a building footprint and an associated set of slopes, and subsequently extrudes the roof to the specified building height. This parametric approach guarantees that a given set of parameters (height, footprint and slopes) produces a regularized building mesh with low polygonal density. A multimodal model, named Point2WSS, was trained to recover the variable number of building’s continuous parameters from its corresponding point cloud. This approach enables the generation of realistic building meshes suitable for electromagnetic simulation, if the predicted parameters accurately approximate real-world values. All the code and datasets used in this paper are available at : https://github.com/KWIKERRR/point2wss.
The rapid evolution of three-dimensional (3-D) geospatial science has redefined the standards of national mapping and cadastral agencies (NMCAs). Traditionally bodies of authoritative 2-D topographic products, these organisations now face the challenge of producing, maintaining, and disseminating national-scale 3-D geospatial datasets that support applications ranging from climate adaptation and urban planning to disaster response and digital twins. This paper presents a comparative study of five NMCAs, comprising IGN (France), BKG (Germany), Kadaster (The Netherlands), GSI (Japan) and USGS (United States of America). By examining agency structure, economic models, and 3-D data collection programmes, this paper identifies converging trends in AI integration, national surveys, along with divergences in funding and implementation. The analysis highlights insights and potential lessons for organisations at early stages of national 3-D dataset implementation.
LiDAR-based perception models’ performance can degrade sharply when applied to data from sensors different to those they were trained on. LiDAR super-resolution aims to enhance sparse point clouds from low-cost sensors. This can help to bridge the sensor domain gap to higher resolution LiDAR. Prior work has primarily focused on reconstruction quality metrics for super-resolution with limited evaluation of downstream perception tasks. We address this gap by conducting a systematic analysis of how super-resolution quality impacts 3D object detection performance. We evaluate detection capability through zero-shot transfer experiments on the KITTI object dataset. Four representative detectors (SECOND, PointPillars, PV-RCNN, PointRCNN) trained on high-resolution data are directly applied to super-resolved low-resolution data without fine-tuning. Results reveal a critical insight: reconstruction improvements yield vastly different detection gains across architectures. PointPillars shows minimal improvement until reaching high reconstruction quality, then performance improves significantly. In contrast, PV-RCNN exhibits steady gains throughout. The highest-quality reconstruction closes up to 86% of the performance gap and enables detection in safety-critical scenarios, including distant vehicles and small pedestrians, where lower-quality methods fail entirely. This work establishes that LiDAR super-resolution effectiveness depends on both reconstruction quality and detector architecture.
The rapid acquisition of high-precision parametric railway alignment is a fundamental prerequisite for intelligent railway construction and maintenance. Traditional measurement techniques and alignment fitting methods heavily rely on manual operations, often resulting in inefficiency, high costs, and insufficient accuracy control. To address these challenges, this study proposes an automated method for extracting and optimizing railway alignment from UAV LiDAR point clouds. Initially, the track centerline is extracted by leveraging the geometric smoothness of the railway and the structural characteristics of the track. A multi-constraint energy model integrating distance, orientation, and curvature is constructed to fit the geometric parameters of alignment elements, thereby providing high-quality initial values for subsequent alignment engineering parameter optimization. Finally, a global optimization strategy based on the simulated annealing algorithm is applied to jointly refine the engineering parameters of the standardized alignment composition, ensuring strict compliance with railway design specification. Experimental results demonstrate that the proposed method can efficiently and robustly extract high-precision alignment parameters with well-defined engineering semantics from complex railway point clouds, thereby providing reliable technical support for intelligent construction and full lifecycle management of railway systems.
High dynamic range (HDR) variations in satellite optical imagery arise from extreme differences in surface reflectance and illumination conditions. Conventional satellite NeRF frameworks are typically trained on tone-mapped or radiometrically enhanced images, where nonlinear preprocessing alters the physical relationship between measured pixel values and true scene radiance. This leads to biased photometric optimization and loss of geometric fidelity, especially under strong illumination contrasts. To address these limitations, we propose an HDR-consistent learning framework that integrates RawNeRF-style radiance supervision with shadow regularization. The method trains directly on raw satellite imagery using a logarithmic, tone mapping–aware loss that preserves linear radiance and stabilizes optimization under high dynamic range conditions. In parallel, a soft shadow regularization constrains network-predicted shadows using geometric cues derived from solar ray casting, promoting physically consistent irradiance decomposition. Experiments on four AOIs from the DFC2019 dataset demonstrate that HDR-aware radiance learning significantly improves DSM accuracy by maintaining linear radiometric consistency. The proposed shadow regularization also improves geometric consistency in structure-dominated urban scenes, although its effect is limited in vegetation-dominant areas where shadow cues are less informative. Although performance gains are smaller in vegetation-dominant areas, the results confirm that combining HDR radiance learning with geometric shadow regularization yields more radiometrically consistent and geometrically accurate 3D reconstruction from satellite imagery.
Wearable sensors are essential for gait analysis outside of traditional laboratory environments. However, selection of the right sensor technology involves several trade-offs. Inertial Measurement Units (IMUs) offer high temporal resolution which are ideal for detecting gait events but they suffer from drift. Ultra-Wideband (UWB) provides stable spatial data, but are less precise for detecting event timing. This paper presents a comparative study of three distinct foot-mounted sensor methodologies for heel strike detection and cadence estimation: (1) IMU-Only approach, (2) UWB-Only approach, and (3) a multi-sensor IMU+UWB fusion approach. Each method is evaluated against a camera-based ground truth system using data from four subjects. Results show the IMU-Only method is inconsistent, with moderate event precision (Avg. F1: 0.798), temporal accuracy (Avg. MAE: 47.99 ms), and subject-dependent cadence accuracy (Avg. Acc: 89.59%). The UWB-Only method provides robust event detection (Avg. F1: 0.811) with similar temporal error (Avg. MAE: 49.0 ms) but is exceptionally accurate for cadence estimation (Avg. Acc: 96.94%). The IMU+UWB fusion approach achieves the highest event precision (Avg. Precision: 0.835) and the best temporal accuracy (Avg. MAE: 46.51 ms), while also maintaining robust cadence accuracy (Avg. Acc: 95.62%). In conclusion, while the UWB-Only method is ideal for cadence-only applications, the IMU+UWB fusion approach provides the best overall balance of high event precision, superior temporal accuracy, and reliable cadence estimation.
Monitoring tree health is essential for detecting early signs of stress, defoliation, and potential mortality, supporting effective forest management, ecosystem conservation, and early warning systems. Advances in deep learning have enabled automated analysis of trees in remote sensing imagery through object detection methods that leverage both spectral and spatial information. However, assessing tree defoliation remains challenging, as subtle differences between defoliation levels make accurate classification difficult. To address this, we propose the hybrid ResNet-Swin Transformer, an object detection architecture built on a Faster R-CNN framework, incorporating a fused ResNet and Swin Transformer backbone with attention-based feature fusion. This design captures rich, multiscale representations by combining convolutional and transformer-based features and progressively refines them through channel-wise attention blocks for robust detection and classification. The architecture was evaluated on a very high-resolution aerial dataset from Switzerland, partially annotated with five classes: Conifer (healthy), Conifer (defoliated), Broadleaf (healthy), Broadleaf (defoliated) and Dead. Comparative experiments with state-of-the-art object detection and classification methods demonstrate that the proposed approach achieves higher accuracy and robustness, highlighting its potential for precise and reliable automated tree health monitoring.
Information on the precision of TLS observables is limited. While the range measurement precision can be modeled with respect to the intensity measurement nowadays, the precision of the angular observations still relies on the claims of the manufacturer. This contribution proposes a method to determine the vertical angular variance of a TLS using profile measurements. Supported by a simulation, which serves as proof-of concept, the methodology is laid out. In the end, measurements with a Z+F IMAGER ® 5016A are evaluated. A dependency of the angular standard deviation on the rotational speed of the beam deflection unit is observed. The estimation precision of the angular standard deviation is high with consistent values for differing ranges. The estimated angular standard deviations are much lower than the claims of the manufacturer starting with roughly 2” for the slowest rotating settings, up to 4” for the fastest. All this can be achieved by scanning a reflectivity target with at least two adjacent fields of different homogeneous reflectivity. This needs to be aligned to the scanner to reduce and eliminate as many contributing error sources as possible. The target itself provides the fields and the transitions needed to perform the in-situ estimation of the angular precision.
To meet the growing demand for 3D digital map applications and to better understand the multi-level spatial structure of cities, some cities have implemented citywide 3D digital map programs. In 3D digital map production, vehicle-mounted mobile surveying is a key component. Drawing with a practical project, this paper proposes a technical scheme for road data acquisition and processing based on the SSW VMMS (Vehicle-mounted Mobile Mapping System). Through integrated processing steps, including combined navigation solution, point cloud correction, image coordinate calculation, image deblurring, point cloud coloring, point cloud denoising, and Orbit GT data preparation, the rapid production of colored point cloud data with georeferenced coordinates, 360° panoramic image data, and individual image data is achieved. A technical scheme suitable for 3D digital map production along urban roads was developed and validated. The results produced by this scheme have passed inspection and acceptance, and were released to the public free of charge as the first batch of visualized 3D map data on the Common Spatial Data Infrastructure Portal (portal.csdi.gov.hk), receiving widespread attention and positive recognition from various sectors of society. This scheme not only promotes the broader application of the SSW VMMS but also provides effective reference for similar urban vehicle-mounted mobile mapping projects.
Retrieving information from point clouds for analysis and visualization has gained ever-increasing interest. A growing niche in this regard is ray queries, commonly used for image synthesis. Ray tracing is widely used in computer graphics, with a multitude of solutions based on bounding volume hierarchies. However, these solutions are rarely straightforward to integrate with raw point cloud data and geospatial analytical workflows. To overcome this, we present a novel approach to ray tracing in raw point clouds that builds upon and extends existing geospatial indices. The solution is exemplified by a fast octree implementation that supports versatile query semantics, such as neighborhood queries with constraints on k and radius for both points and rays, while offering configurable data organization schemes, including layered, fixed, and adaptive depth. The evaluation demonstrates satisfactory speed and capabilities for many scientific use cases, while simultaneously exhibiting low implementation costs, high flexibility, and simplicity in integrating ray tracing into analytical point cloud workflows.
Accurate building height estimation plays a crucial role in large-scale 3D urban reconstruction. However, conventional stereo matching approaches often suffer from mismatches around building edges, leading to unreliable height retrieval in dense urban areas. To address this issue, this paper presents a novel method for building height estimation based on contour vector registration integrated with the vertical line locus technique. The proposed framework first automatically matches building contour vectors extracted from stereo high-resolution satellite images. Then, for each paired contour, a range of candidate heights is searched using a rational function model to project the reference contour from the image space to object space and then reproject it onto the conjugate image. The elevation that maximizes the overlap ratio between projected and paired contours is identified as the optimal roof elevation. Building height is subsequently derived by subtracting the ground elevation from the estimated roof elevation. Experiments conducted on SuperView-1 (SV-1) satellite stereo images over Jiuyuan District, Baotou, Inner Mongolia, China, demonstrate the effectiveness of the proposed method. The resulting building height estimates achieve a root mean square error of 0.84 m compared to manual measurements, showing strong agreement (r = 0.9993). The proposed contour-based stereo registration approach provides a robust and efficient solution for building height extraction from high-resolution satellite data, supporting precise urban 3D modeling and large-scale spatial analysis.
Speech is a highly complex and multidimensional process, requiring precise coordination of muscular actions within the vocal tract. Disruptions or delays in speech motor control often lead to speech impairments. Recent advancements in markerless facial tracking technology enable the collection of objective measurements to assess these impairments. To obtain such photogrammetric measurements, a multi-camera network is employed, making accurate camera calibration essential. This paper examines the constraints applied during the calibration process. Two adjustment strategies were evaluated. The first, Independent Adjustment (IDP), performs self-calibration for each camera without introducing constraints. The second, Combined Adjustment (CMB), incorporates object space constraints by ensuring that object point locations observed from all cameras remain consistent. Given the cameras’ narrow fields of view, both IDP and CMB were tested with additional constraints related to the principal point offset. Each adjustment was executed under two conditions: fixing the principal point offset to zero or estimating it as part of the calibration. Results indicate that the choice of adjustment significantly affects the interior orientation parameters (IOPs). IDP with the principal point offset fixed to zero produced the most accurate outcomes. However, variations in IOPs had no meaningful impact on object space coordinates. These findings suggest that the simplest approach—IDP with the principal point offset fixed to zero—offers reliable calibration for multi-camera systems used in speech assessment. This streamlined method can be adopted in future applications to enhance efficiency without compromising accuracy.
The modeling of buildings suffers from a dichotomy between generic and specific representations: the lack of domain knowledge in flexible models that can represent many shapes, and the restricted geometry of pre-specified parametric building primitives. To fill this gap, we propose using general boundary representations enriched with automatically recognized and enforced geometric constraints derived from human-made regularities. The proposed reasoning process relies on the statistics of the planar point groups extracted from airborne-captured point clouds. Hence, a chosen significance level is the only process parameter. To enforce the creation of sound solids, we apply manifold constraints for the generation of the boundary representations. The feasibility and usability of the approach are demonstrated by evaluating an airborne-captured laser scan containing approximately 7,600 buildings over an area of 50 km2 featuring both inner-city and rural landscapes.
Semantic 3D indoor mapping often depends on supervised learning and large annotated datasets, limiting scalability across diverse environments. This work introduces a category-specific prompt strategy for semantic 3D mapping using RGB-D cameras, integrating RGB-D SLAM with the Segment Anything Model 2 (SAM2) to enable annotation-efficient reconstruction. Keyframes and trajectories extracted from SLAM provide spatial references, while SAM2 performs zero-shot segmentation guided by a Category- Wise Prompt Segmentation Strategy (CPSS), which segments structural and functional elements (e.g., floors, doors, staircases) by category to reduce prompt interference and manual effort. The segmented keyframes are then fused with depth and pose data to produce instance-level semantic point clouds. Experiments on custom RGB-D sequences and selected ScanNet scenes demonstrate centimeter-scale geometric consistency and strong semantic consistency, with mIoU values up to 0.89 on the custom dataset and 0.98 on ScanNet. The resulting semantic point clouds are clean, structured, and require minimal post-processing, showing that the proposed strategy provides an efficient and scalable solution for semantic 3D indoor mapping without retraining or environment-specific supervision.
Mobile robots are widely used in unmanned surveying, warehouse logistics, and emergency response. However, achieving safe, reliable, and efficient autonomous navigation in unknown environments remains challenging, where accurate environment representation and feasible trajectory planning are crucial. This paper presents an autonomous navigation method integrating lightweight LiDAR mapping with real-time local planning for ground robots. At the perception level, an incremental single-frame point cloud update is used to accumulate and project locally traversable space, producing a lightweight obstacle map that preserves geometric accuracy while reducing planning computation. At the planning level, A* is employed to generate reference control points, and uniform B-spline curves are used to optimize the trajectory while enforcing kinematic feasibility and smoothness. At the control level, nonlinear model predictive control (NMPC) ensures accurate trajectory tracking by producing control commands that satisfy velocity and acceleration constraints. The framework also supports low-cost evaluation in simulation. Experiments in simulated forests, simulated indoor corridors, and real-world gardens and hallways show average navigation speeds of 2.24 m/s, 0.76 m/s, 0.43 m/s, and 0.38 m/s, respectively. Results demonstrate that the proposed method generates smooth, feasible, and safe trajectories and completes autonomous navigation and mapping tasks across diverse environments.
LoD-2 building models are more informative and practically more useful than LoD-1 representations because they capture the roof structure that defines the essential three-dimensional form of a building. They are important for applications such as urban planning, environmental simulation, and digital heritage. Although recent roof shape extraction methods can derive vectorised 2D roof structures from very-high-resolution imagery, transforming these image-based representations into fully textured 3D buildings remains challenging. In this paper, we present a semi-automated LoD-2 reconstruction pipeline that integrates HEAT-derived roof geometry with airborne LiDAR, satellite and Google Street View imagery. The 2D outputs are reprojected into map coordinates, fused with LiDAR through a two-stage roof reconstruction strategy to derive roof shapes and combined with an adaptive, LiDAR-based ground base initialisation to create a complete 3D wireframe. Roofs are textured using VHR orthophotos while the walls are textured via a process of Street View panorama selection, geometric filtering, Mask2Former segmentation, and homography rectification. Across a large-scale evaluation on 1000 buildings, the proposed two-stage reconstruction strategy improves geometric agreement with the LiDAR reference data achieving a roof-surface RMSE of 0.445 m. The wall texturing process produces convincing facades when suitable panoramas are available. While minor challenges such as sensitivities to LiDAR outliers, incomplete roof geometry, and facade occlusions persist, this pipeline effectively bridges 2D roof parsing and textured LoD-2 model generation, providing a robust and scalable foundation for advancing toward fully automated workflows.