
Circular marker detection plays an important role in high-speed videogrammetry applications, in which their accurate and reliable detection and localization is essential for precise motion measurement and deformation analysis. However, achieving high global detection accuracy with existing ellipse detection methods remains challenging in the presence of noise, background clutter or illumination variation, which often lead to missed detections and false positives. To address these challenges, this paper proposes a robust unsupervised ellipse detection method for circular markers, designed to minimize these errors and enhance detection accuracy. The proposed method comprises four components: initial ellipse set extraction, ellipse candidate generation, target ellipse detection, and target ellipse refinement. To improve detection performance, we introduce a reliable candidate generation approach coupled with a robust detection and refinement strategy. Experimental evaluation on multiple real-world data sets demonstrates that the proposed method outperforms reference methods, in particular in detection accuracy, while maintaining comparative computational efficiency.
With the growing importance of underwater vision tasks in fields such as marine geological research, coral reef monitoring, and aquatic population assessment, underwater image retrieval has become a challenging and urgent problem to address. However, underwater images suffer from color distortion, blur, low contrast, and high visual similarity between different objects. These issues complicate feature extraction, ultimately affecting underwater image retrieval accuracy. To mitigate the aforementioned challenges, this study refines underwater image retrieval with deep orthogonal convolutional aggregation of local and global features and constructs an underwater image retrieval data set (the data set is available at: https://github.com/jojokaka/UIRD.git) dedicated to this task. The method begins with an underwater image enhancement stage to mitigate underwater image degradation and then uses an orthogonal fusion module, which adopts a convolutional global average pooling approach to replace the traditional global average pooling mechanism, effectively integrating local and global features to capture richer local feature information and enhance the model’s representational capacity. Comprehensive experimental evaluations verify that the proposed approach achieves superior performance, outperforming multiple state-of-the-art baseline models by a significant margin in both retrieval accuracy and robustness.
Remote sensing change detection is essential for land use planning, urban monitoring, and disaster response. However, conventional deep learning methods often suffer from blurred building and vegetation edges, shadow misclassification, and low efficiency due to excessive complexity. To address these issues, we propose ELFFNet, a lightweight and efficient network that achieves high detection accuracy with reduced computational cost. ELFFNet integrates Asymmetric Dilated Convolution Modules for deep semantic extraction, Multi Efficient Attention Modules between encoder and decoder to retain fine spatial details, and a simplified Pyramid Pooling Module at the deepest stage for low-cost global context aggregation. Experiments on the GVLM-CD and WHU-CD data sets show superior performance, with Kappa coefficients of 82.86% and 94.48%, F1 scores of 91.72% and 95.24%, and mean intersection over union values of 84.96% and 94.73%. Importantly, ELFFNet requires a parameter size of 7.31 MB and 22.6 giga floating-point operations [GFLOPs])—far fewer than Transformer-based models—while delivering strong edge detection for small targets and complex disaster areas. This balance of accuracy and efficiency makes ELFFNet highly suitable for practical remote sensing applications.
Cross-view localization and cross-view synthesis are commonly organized as separate retrieval and generation problems, although both integrate overhead imagery and ground-level visual evidence for navigation, mapping, digital twins, and immersive scene generation. This survey adopts a unified view in which both tasks are treated as instances of cross-view transfer. Their objectives differ, yet both are shaped by coupled viewpoint, scale, temporal, visibility, and sensor-domain gaps between near-nadir and ground-level observations. The survey formulates the intermediate transfer operation as the view-transformation primitive T, a cross-view transformation stage before task-specific retrieval or generation modules. In localization, T supports descriptor construction and pose-aware matching. In synthesis, it converts source observations, pose cues, layouts, geographic context, or scene supports into conditioning representations for generation. The resulting taxonomy organizes localization by the interaction between transformation and backbone and synthesis by the coupling among generation engine, conditioning representation, geographic objective, and rendering mechanism. The review further examines data sets, objectives, metrics, and protocols, linking empirical conclusions to geographic coverage, sampling density, pose metadata, sequence continuity, and metric design. Across both tasks, the analysis identifies trade-offs among geometric explicitness, scalability, data demand, generalization, photorealism, and geographic fidelity and isolates common failure modes caused by incomplete geometry, temporal drift, scale mismatch, and weak static-scene evidence. The resulting analysis motivates deployment-scale retrieval, continuous pose estimation, temporally aware supervision, multi-view consistent synthesis, geographic identity preservation, and unified discriminative-generative representations.
Modern digital elevation models (DEMs) are generated with increasingly small cell sizes and high accuracy, yet no closed-form relationship links these variables. Previous work relied on empirical formulations validated experimentally. This paper derives an analytical relationship for the traditional elevation mean square error (MSE) accuracy metric and highlights limitations of empirical approaches. Common practice wrongly dismisses the effect of the choice of interpolation method on DEM accuracy. Accuracy reports may even fail to mention which interpolation method was used. The analysis was limited to two local interpolation methods: nearest neighbor and bilinear interpolation. In those cases we show that the model accuracy is bounded by three terms: one depending only on the dataset elevation MSE; mixed monomials depending on both the dataset elevation MSE and powers of the cell size h; and a term independent of the dataset elevation MSE and proportional to a power of h. For a given interpolation method, the exponents of h are constant, whereas the coefficients can be estimated from local terrain characteristics at the control points. For bilinear interpolation, the exponents are twice those for nearest-neighbor interpolation. The availability of this analytical relationship makes it possible to determine the cell size or instrument accuracy required to achieve a specified target DEM accuracy. Because the result depends strongly on the interpolation method, the recommended interpolation method should be routinely and explicitly mentioned in the DEM producer’s accuracy report
Temperature retrieval from unmanned aerial vehicle (UAV) thermal imagery commonly relies on proprietary manufacturer radiometric calibration software. However, the correction procedures used by these workflows are generally undisclosed, and some user-defined parameters, particularly altitude, are restricted to ranges that do not reflect real-world conditions, risking the introduction of bias and higher radiometric uncertainty. Accordingly, this study evaluates the accuracy of proprietary radiometric calibration by examining how its performance changes with flight altitude, while also proposing a practical method to correct residual bias after proprietary processing. Thermal imagery is collected with a DJI Matrice 300 RTK equipped with a Zenmuse H20T sensor at 15, 30, 45, and 60 m. Two temperature-controlled plates (TCPs) spanning cold and hot regimes are used to validate three DJI Thermal SDK workflows: (1) default settings; (2) custom settings using measured emissivity, flight altitude, humidity, and reflected apparent temperature (Treflected); and (3) custom settings with per-image tuning of Treflected. Residual bias is subsequently corrected using empirical line calibration (ELC) derived from TCP observations. Across all proprietary workflows, temperature error increases with altitude, indicating progressive gain and offset bias not fully corrected by manufacturer processing. Default settings produce the largest error, with mean absolute error increasing from 4.39°C at 15 m to 11.69°C at 60 m. Custom settings reduce error moderately (3.86°C to 9.43°C), while Treflected tuning improves close-range performance but does not eliminate altitude-dependent degradation (0.99°C to 8.57°C). Application of ELC substantially reduces residual bias across workflows, with best performance achieved by combining Treflected tuning with ELC, resulting in sub-1°C accuracy across altitudes. These results demonstrate that proprietary UAV thermal image processing alone is insufficient to ensure robust temperature retrieval but that a simple empirically calibrated postprocessing workflow using TCPs can mitigate altitude-dependent bias and enable high-accuracy temperature estimation in operational remote sensing applications
In geophysical exploration, the integrated interpretation of multi-physical parameter data is essential for precisely identifying mineral prospecting targets. Composite isosurfaces constitute a critical framework for the visualization and quantitative analysis of such complex data sets. Nevertheless, conventional marching cubes (MC) algorithms encounter significant limitations—such as topological ambiguities, challenges in multi-parameter data fusion, and inadequate adaptability to diverse clipping scenarios—when dealing with combinations of multiple physical property parameters. To address these issues, this paper introduces an enhanced composite marching cubes (CMC) algorithm. The proposed CMC algorithm systematically redefines the 15 isosurface triangulation configurations of the standard MC algorithm and establishes new strategies for combining grid nodes under multi-physical parameter conditions. These improvements not only resolve topological ambiguities and improve computational efficiency but also facilitate the precise generation of composite isosurfaces representing multiple properties. Furthermore, by incorporating blank grid labeling and intersection point screening mechanisms, the algorithm enables arbitrary and repeatable region clipping, thereby accommodating complex terrains, voids, and areas of high uncertainty. Application of this methodology in the Panzhihua region has demonstrated its efficacy in characterizing the coupling of multi-physical parameters and in accurately delineating areas with high mineralization potential
True digital orthophoto maps (TDOMs) serve as foundational geospatial products for applications in land surveying, urban planning, and emergency management. Conventional TDOM generation relies on differential correction, often resulting in cartographic artifacts such as geometric discontinuities, radiometric inconsistencies, and linear feature misalignments. Although recent methods leverage 3D Gaussian splatting to bypass differential correction, their computational demands hinder real-world deployment. To address these limitations, we propose FastPro-Gaussian, a novel framework enabling rapid high-quality TDOM generation on consumer-grade GPUs. Our contributions are threefold: block-based processing ensuring scalability for large-scale scenes; progressive densification stabilizing model optimization and reducing training iterations; and spherical-to-ellipsoidal Gaussian transformation, accelerating early-stage optimization of structural features (e.g., building edges). Experiments demonstrate that FastPro-Gaussian surpasses commercial solutions (ContextCapture, Metashape, and Pix4Dmapper) in rendering quality for building facades, edges, and roads. Compared to state-of-the-art methods, it achieves comparable TDOM fidelity with over two-fold acceleration in training time (notably more than two times faster than Tortho-GS). These gains in efficiency and efficacy confirm its strong potential for practical deployment in geospatial production pipelines.
Individual trees are essential components of forest ecosystems, and accurate tree-level segmentation provides a crucial foundation for forest ecosystem modeling and biodiversity assessment. We propose a novel point cloud-based individual tree segmentation method guided by morphological prior constraints and single-tree canopy radiative effective extent (SCREE), a metric used to define the canopy boundary for each tree, computed from high-dimensional features output by a network. First, the original point cloud is divided into a central point set and a buffer point set based on morphological prior constraints. On this basis, different strategies are applied for trunk extraction from the central and buffer points: trunks in the central point set are extracted through adaptive density-based filtering, while trunks in the buffer point set are extracted using a semantic segmentation network. Finally, canopy points are assigned to individual trunks to obtain complete tree structures. This assignment is guided by the proposed SCREE, inferred using a diameter at breast height–height generative model to generate a constrained space. Validated on six forest scenes from three public data sets, the proposed method achieves an average instance-level accuracy of 84.71% and an average point-level accuracy of 81.85%, demonstrating strong robustness across diverse forest environments. Moreover, compared with existing methods, our method exhibits higher stability when handling variations in samples and the presence of small trees across different types of scanners.
In the context of large-scale hydropower infrastructure, the accurate extraction of tunnel rock mass structural information is vital for informed engineering design. However, this task remains inherently difficult due to the complexity of subsurface geological conditions and the inherent limitations of conventional survey techniques. This study proposes an integrated methodology that combines high-resolution 3D reconstruction with the Geo-AINet ensemble learning framework to enable automated identification of structural planes. The process involves generating detailed tunnel models, extracting multi-dimensional semantic features, performing initial segmentation, and subsequently applying cluster analysis to refine the structural interpretations. Experimental validation confirms the method’s high level of precision: the extracted structural orientations exhibit average angular deviations of less than 3 degrees, with a maximum error margin not exceeding 5 degrees relative to manual measurements. These findings demonstrate the method’s capacity to meet stringent engineering standards while enhancing both the efficiency and safety of geological data acquisition in tunneling projects.
Groundwater overexploitation has triggered widespread land subsidence across the North China Plain, posing significant risks to urban infrastructure. Using time-series interferometric synthetic aperture radar data from 2017 to 2025, this study investigated the spatiotemporal evolution of vertical land motion in Xingtai, a region influenced by the interplay of anthropogenic interventions, tectonic structures, and extreme climate events. Our analysis reveals a stark spatial contrast: while eastern areas exhibit persistent subsidence exceeding −50 mm/year due to industrial groundwater demand, the western sector and Longyao County have exhibited localized uplift (10 to 20 mm/year) since 2017, affirming the efficacy of mitigation policies such as the South-to-North Water Diversion Project. Time-series analysis identifies two distinct uplift mechanisms: a gradual, policy-driven aquifer recovery confined by local faults, and an abrupt, transient pulse triggered by the July 2021 extreme precipitation event. Furthermore, deformation discontinuities across the Longyao ground fissures are attributed to differential subsidence rather than deep-seated tectonic creep. These findings highlight the critical coupling between long-term water management and acute climatic disturbances in shaping surface deformation, emphasizing the need for integrated strategies for sustainable groundwater governance and hazard mitigation in water-stressed regions.
Airborne laser scanning point clouds constitute a primary data source for city-scale 3D building modeling. However, the automated reconstruction of geometrically accurate and topologically consistent roof wireframes persists as a formidable challenge, hindered by data noise, uneven point density, and the topological complexity of urban roofs. Existing data-driven methods exhibit high sensitivity to local fitting errors, frequently leading to topological inconsistencies, whereas model-driven approaches are constrained by predefined primitive libraries, limiting their generalization to complex composite roofs. To address these limitations, this paper proposes a novel framework for roof wireframe reconstruction that fuses topological data analysis (TDA) with hybrid geometric-topological constraints. A core innovation of this framework is the application of persistent homology to extract globally stable topological skeletons from 2.5D height scalar fields, serving as a robust topological prior independent of local geometric noise. Specifically, the method first generates regularized 3D eave lines constrained by 2D building footprints. Subsequently, a “geometry–topology” dual-domain cross-validation mechanism is used to validate geometric hypotheses derived from planar adjacencies against TDA-extracted topological critical points (ridges and valleys), thereby effectively suppressing spurious structure lines. Finally, a global optimization model governed by the minimum description length principle enforces implicit regularization to rectify local geometric distortions and guarantee topological closure. Experiments on the Trondheim dataset demonstrate that the proposed method achieves a root mean square error of 0.443 m, while enhancing the length-weighted completeness and correctness of structure lines to 96.23% and 97.41%, respectively. These results validate the efficacy of integrating height-field–based topological skeletons for reconstructing complex roof wireframes, offering a scalable and robust solution for automated large-scale urban 3D modeling.
Urban green infrastructure plays a critical role in enhancing ecological resilience and reducing infrastructure vulnerability in metropolitan settings. However, achieving scalable, high-resolution monitoring of urban green infrastructure remains a persistent challenge due to visual occlusion, structural complexity, and the cost or inaccessibility of conventional three-dimensional remote sensing technologies. This study introduces a novel, low-cost, and reproducible framework for near real-time, object-level structural assessment and geo-location of roadside vegetation and infrastructure using commonly available but underused dashboard camera (i.e., dashcam) video data. A pipeline was developed that combines monocular depth estimation, supervised calibration of a monocular depth proxy, and geometric triangulation to generate accurate spatial and structural data from continuous street-level video streams acquired from vehicle-mounted dashcams. Depth outputs were treated as a relative proxy and calibrated via a gradient-boosted regression model to metric camera-to-object distance, particularly for distant objects. The depth correction model achieved strong predictive performance (R 2 = 0.92, mean absolute error = 0.31 on transformed scale), significantly reducing bias beyond 15 m. Further, object locations were estimated using global positioning system–based triangulation, whereas object heights were calculated using pinhole camera geometry. This method was evaluated under varying conditions of camera placement and vehicle speed. The configuration involving interior-mounted cameras and low-speed travel yielded the highest accuracy, with mean geo-location error of 2.83 m (interquartile range = 2.64 m) and mean absolute error in height estimation of 2.09 m for trees and 0.88 m for poles. This approach complements conventional overhead remote sensing methods, such as lidar and stereo imaging by enabling low-cost, frequent object-level monitoring of vegetation risks and infrastructure exposure for utility and urban planners, whereas broader transferability across locations, camera models, and environmental conditions require further multi-site evaluation.
Photogrammetric mapping missions using robots or unmanned aerial vehicles (UAVs) often encounter environments where sparse strong-texture structures are found alongside extensive weak-texture surfaces, such as indoor corridors with frames and aerial views of shorelines. Establishing accurate correspondences under such circumstances remains a major technical challenge. Traditional methods can only match a few or no correspondences. Deep learning algorithms can yield stable matches in weak-texture regions. However, they rely on large-scale annotated data and computational resources, limiting their real-time application and generalization capability. This article proposes an anchor point–assisted image-matching method, using Euclidean distances and angular relationships between corner points in the weak-texture area and anchor points in the strong-texture area to establish structured features, enhancing the distinctiveness of the corner points in the weak-texture area. Meanwhile, epipolar geometry constraints are applied to restrict the candidate corner points search to a one-dimensional range, thereby improving the accuracy, efficiency, and reliability of the corner point matching. Comparative experiments on indoor, outdoor, and UAV data sets demonstrate that the proposed method achieves a greater number of correct matches than scale-invariant feature transform (SIFT), oriented FAST and rotated BRIEF (ORB), KAZE, AKAZE, grid-based motion statistics (GMS), SuperGlue, and LightGlue with a conventional CPU configuration. The presented method, SuperGlue, and LightGlue achieve a success rate (SR) of 100% on three data sets and secure sufficient correspondences in weak-texture regions. In contrast, the SR obtained by traditional methods ranges from 70% to 80%. In addition, our method runs in under 3 seconds, whereas the deep learning methods still require over 10 seconds even on a lightweight CPU. The effects of anchor point numbers, dual threshold settings, and epipolar search range on matching performance are analyzed to identify the optimal parameter configuration. The proposed method provides potential for robot and UAV photogrammetric mapping applications in sparsely textured environments based on resource-constrained platforms.
Railways serve as critical national infrastructures, and anomalies in key facilities can pose serious threats to transportation safety. Due to complex spatial structures and large-scale variations in railway point clouds, existing semantic segmentation methods struggle with multi-scale feature representation and semantic modeling. To address this issue, we propose a graph convolution–based point cloud semantic segmentation method. The network uses a multi-level feature aggregation framework in which dilated residual blocks facilitate adaptive feature fusion, and hierarchical downsampling progressively expands the receptive field to capture global structural information. Furthermore, a coordinate-guided graph convolution network is introduced to enhance local structural perception and encode topological relationships by constructing graphs from point coordinates and propagating features to model geometric relations and long-range dependencies. Together, these components achieve a unified multi-scale semantic representation for railway point clouds. With the WHU-Railway3D data set, our method achieved mean intersection over union scores of 72.40% and overall accuracy of 91.10% in the urban scenarios and mean intersection over union scores of 77.89% and overall accuracy of 95.50% in the rural scenarios, which shows state-of-the-art performance among current methods.
Satellite-derived bathymetry is an important method for obtaining shallow sea water depth. To address the problem that the traditional band log-ratio model exhibits significant discrepancies in bathymetric inversion accuracy in both shallow sea (<2 m) and deep sea (>12 m), where water depth is overestimated in shallow areas and underestimated in deep areas, this paper proposes a bathymetric inversion method from active–passive satellite remote sensing data based on a residual correction model. First, the proposed method uses the denoised Ice, Cloud, and Land Elevation Satellite-2 (ICESat-2) ATL03 data set as precise depth control points to construct and implement the traditional band log-ratio model, generating initial bathymetric inversion results. Second, the initial inversion result is compared with the ICESat-2 depth control points, and the residuals between the initial water depth estimates and the true water depths in the local region are calculated. Then, based on the relationship between the spectral characteristics of remote sensing images and water depth residuals, a random forest model is constructed to fit the residual distribution across the entire study area. Finally, using the residual distribution derived from this fitting process, the initial bathymetric inversion results for the entire region are corrected, thereby improving the accuracy and precision of the water depth data. Experimental results demonstrate that the proposed method achieves a bathymetric inversion accuracy with the root mean square error better than 1.57 m and the mean absolute error better than 1.15 m for islands with varying seabed topographies, providing high-precision shallow sea bathymetric inversion results.
Spectral unmixing is a technique to predict the proportions of land cover classes within mixed pixels. The proportions can serve as inputs for land cover classification at finer spatial resolutions (i.e., by subpixel mapping [SPM]). To achieve high-quality finer spatial resolution land cover mapping, during the spectral unmixing process, it is necessary to accurately identify land cover boundaries, which are typically represented by mixed pixels. The recently proposed unsupervised object-based subpixel mapping model (UO-SPM) incorporates object-level information into spectral unmixing. However, its performance is limited by uncertainties in unsupervised clustering and the automatic threshold segmentation. In this paper, we proposed a supervised strategy for object-based spectral unmixing (SOSU). SOSU segments observed images (i.e., Landsat images in this paper) into objects and uses supervised classifier to produce object-level classification maps. Mixed (i.e., located at object boundaries) and pure pixels (i.e., within object) are subsequently identified using an erosion algorithm applied to the classification maps. The identified mixed pixels are unmixed using the extracted pure pixels as endmembers. The SOSU method not only inherits the advantage of UO-SPM but also exploits supervised and object-based spatial information to achieve more reliable segmentation and unmixing. The effectiveness of SOSU was validated across seven different experimental regions. By introducing supervised information, SOSU increases the correlation coefficient by 0.0347 and reduces the root mean square error and mean absolute error by 0.0352 and 0.0200, respectively, compared with UO (i.e., the spectral unmixing component of UO-SPM).
Meeting strict latency budgets for high-performance approximate nearest neighbor (ANN) search remains a critical challenge in remote sensing edge devices, including microsatellites and UAVs, because of stringent limitations in both primary (RAM) and secondary (disk) storage. To address this challenge, we propose Edge-ANN, an efficient ANN framework specifically engineered for storage-efficient edge retrieval. Instead of explicitly storing high-dimensional partitioning hyperplanes, Edge-ANN uses pairs of in-data-set points, termed anchors, together with a scalar offset to define partitions implicitly. To ensure that these implicit partitions remain balanced and effective, we introduce a binary anchor optimization algorithm. Rigorous experiments on four multi-source and multi-modal data sets, Million-AID, High-resolution Urban Complex, GlobalUrbanNet, and BigEarthNet, demonstrated that under simulated edge environments with dual storage constraints, Edge-ANN achieves a 30% to 40% reduction in secondary storage compared with an unconstrained Annoy baseline, at the cost of only a 3% to 9% reduction in Recall@10. Furthermore, within constrained memory budgets, Edge-ANN delivers superior retrieval performance to other mainstream ANN methods under the adopted near-real-time latency budget. Collectively, these results establish Edge-ANN as a practical solution for enabling large-scale, high-performance remote sensing feature retrieval on edge devices with exceptionally constrained storage. The codes of Edge-ANN are available at https://github.com/huaijiao666/Edge-ANN
Monitoring large-scale surface deformation using satellite radar data is crucial for advancing regional and continental-scale geodynamic observations. However, challenges arise when unifying deformation data from different tracks of the same satellite, primarily due to inconsistencies in reference baselines and variations in radar viewing angles. This study proposes an interferometric synthetic aperture radar (InSAR) deformation datum connection method with a fixed line-of-sight (LOS) direction. The method combines Bayesian inference with a Markov random field model and integrates InSAR and global navigation satellite system deformation measurements to unify deformation datums across multi-track SAR interferometric results using only a single orbit direction (ascending-only or descending-only), without requiring both ascending and descending data sets. In the simulation experiments, the root mean square error (RMSE) of the LOS displacement rate difference in the overlapping regions of adjacent-track SAR images decreased by 99%. Applying the developed methodology for InSAR observations of the 2023 Mw 7.8 Kahramanmaraş earthquake in Türkiye, the RMSE of displacement differences in the overlapping regions of adjacent tracks was reduced from 98 to 34 mm, demonstrating the effectiveness of the proposed unified datum approach.