
ABSTRACT Accurate measurement of individual‐tree structural parameters is critical for forest inventory, carbon stock assessment, and urban ecosystem monitoring; however, low‐cost image‐based approaches often fail to capture under‐canopy structures due to canopy occlusion and limited trunk visibility. This study proposes a fully photogrammetric and cost‐effective workflow based on a single‐sensor unmanned aerial vehicle (UAV) using a hybrid flight geometry to improve trunk observability and automate individual tree parameter extraction. A double‐grid acquisition plan combining nadir (−90°) and oblique (−45°) imagery was applied over a semi‐dense forest plot including 120 reference trees. SfM–MVS processing of 918 images produced a high‐density point cloud (162.5 million points) with 1.54 cm/pixel ground sampling distance. Individual trees were delineated using a hybrid strategy integrating DBSCAN‐based coarse clustering, RANSAC‐validated trunk geometry, and trunk‐seeded 3D splitting. Tree height, diameter at breast height (DBH), and crown area were derived from segmented 3D tree models and validated against field measurements and differential GNSS observations. The proposed method successfully detected 108 of 120 reference trees (Precision = 1.00, Recall = 0.90, F1‐score = 0.95) and demonstrated high agreement with in situ measurements, achieving RMSE values of 0.79 m for tree height ( R 2 = 0.975), 3.02 cm for DBH ( R 2 = 0.931), and 3.51 m 2 for crown area. In addition, crown boundary modeling showed that α‐shape reconstruction provided a more realistic representation (IoU = 0.78) than Convex Hull methods by reducing systematic overestimation.
3D reconstruction has matured into a robust technology. However, small, flexible objects such as conifer seedlings remain challenging due to their fine-scale structures, and susceptibility to movement. This study investigates and evaluates methods for reconstructing spruce (Picea abies) and pine (Pinus sylvestris) seedlings, with the aim of establishing a workflow capable of capturing geometry and texture for applications in machine learning and virtual testing environments. Two acquisition approaches were tested: photogrammetry using a RGB camera and a 3D scanner, both mounted on a robotic arm. While the scanner produced incomplete results, the photogrammetry approach successfully generated point clouds (pcl) with color information. Three different photogrammetry software were tested before relying on Agisoft Metashape and Meshroom for image processing and dense pcl generation, followed by pcl filtering in CloudCompare and meshing in Blender. Six seedlings were reconstructed to textured meshes and quantitatively evaluated using the metrics precision, recall, F1-score, mask intersection-over-union (IoU), and boundary IoU. Results showed an average mask IoU of 75.7% and F1-score of 86.1%. Pine seedlings yielded higher recall and F1-scores, whereas spruce reconstructions demonstrated higher precision. The proposed semi-automated workflow demonstrates the feasibility of reconstructing small and slender structured flexible objects, specifically conifer seedlings.
Land cover classification (LCC) through synthetic aperture radar (SAR) and optical data provides comprehensive scene information for more accurate results. However, existing methods have limitations in two aspects: (1) insufficient generalization of foundation vision models in the remote sensing domain, and (2) modal imbalance and ineffective feature fusion in multi-modal architectures. To address these, we propose APS-Net with three key innovations: Adapter fine-tuning strategy, Primary modality regularization and Adaptive sparse attention. Specifically, we present a parameter-efficient adapter fine-tuning strategy to integrate a vision foundation model with the LCC task, demonstrating its ability to effectively extract features from both SAR and optical remote sensing images. We introduce a knowledge distillation-based architecture that leverages a pretrained optical segmentation model to facilitate multi-modal fusion through employing the primary modality as external regularization. Additionally, we propose an adaptive sparse attention adjustment module that filters out redundant information and suppresses noisy interactions, enabling the effective extraction of relevant SAR and optical features. These components collectively enhance cross-modal synergy modeling and alleviate modal imbalance and unreliable features. Comprehensive experiments on the WHU-OPT-SAR, DDHR and DFC2020 datasets demonstrate optimal performance with APS-Net outperforming suboptimal method by 1.6% OA and 0.9% Kappa on average.
To address the technical challenges of monitoring wind turbine blades under low-light nighttime conditions, this study proposes a long-exposure photogrammetry method based on stereo vision for monitoring wind turbine blades operating under low-light conditions. This method uses frequency-controllable LED markers. By considering parameters such as blade motion speed and distance, an appropriate frequency range is determined. Long-exposure imaging is employed to capture the motion trajectory, and morphological operations, marker extraction, and computational solutions are applied to obtain the blade's trajectory and kinematic parameters. Two experiments were conducted to validate the method. First, the method was applied to measure a rotating blade in the laboratory. Compared with vibration analysis equipment, the errors in the first and second natural frequencies were -1.66% and 0.45%, respectively, demonstrating the method's ability to monitor blade motion trajectories and vibration characteristics. Second, under various motion conditions of the robotic arm, the proposed method yielded a trajectory error of 0.15% and a velocity error of -0.45% compared to the theoretical values. The results showed high consistency with the preset parameters of the robotic arm. This study provides a viable vision-based solution for structural health monitoring of wind turbine blades, showing significant potential for operational monitoring in low-light environments.
The current challenges in remote sensing image segmentation are twofold: balancing global and local information and reducing redundant information in shallow features of the encoder-decoder architectures. To solve the problems, we proposed a semantic segmentation network for remote sensing images, called the Multi-Scale Sliding Refined Network (MSSRNet), which includes a Multi-Scale Sliding Transmit Attention (MSSTA) module and a Shallow Feature Refinement (SFR) module. MSSTA enhances both local detail extraction and long-range dependency modeling by combining sliding window attention (SWA) with standard self-attention (SA) while integrating multi-scale features. SFR effectively reduces redundant information in shallow features and enhances feature representation by reconstructing spatial information and enhancing channel information. The effectiveness of MSSRNet was validated on three publicly available datasets, which are ISPRS Vaihingen, ISPRS Potsdam, and LoveDA. Experimental results demonstrate that MSSRNet achieves state-of-the-art mIoU scores of 85.26%, 87.93%, and 55.97%, compared to recent high-performance models like MMT, MSSRNet improves mIoU by 1.1% on the Vaihingen dataset while reducing parameter counts by approximately 34.6%, proving its superior balance between accuracy and computational efficiency.
Lines are prevalent geometric features in urban environments and serve as an important complement to point features, playing an irreplaceable role in specific scenarios and tasks. Line feature matching is thus a crucial and fundamental task in photogrammetry and computer vision. However, matching line features remains challenging, particularly in scenes with semantic ambiguity caused by dense fragmentation and similar structures. This study embeds a novel line transformation block into a graph neural network, leading to a novel network, that is, Line-Transformer Graph Neural Network (LTGNN), which optimizes the global topology while generating geometrically enhanced semantic descriptors for robust line feature matching. Furthermore, the proposed LTGNN leverages the complementary nature between point and line features to integrate local and global information. Firstly, we construct a unified graph structure called an LSD-Wireframe using the retrieved point and line features. Secondly, we implement position encoding for LSD-Wireframe nodes, which is subsequently passed into the LTGNN for enhancing information propagation and facilitating the learning of spatial and visual relationships among nodes. Finally, we employ a geometry-constrained hybrid loss function to supervise the training process and optimize the network toward producing accurate and consistent line correspondences. We evaluated the proposed LTGNN on three different datasets and compared it to six benchmarking methods. Experimental results demonstrate that the proposed LTGNN outperforms others in feature matching, homography estimation, and camera localization applications.
Point cloud registration within agricultural environments presents unique challenges due to cluttered scenes, repetitive plant structures, partial overlaps and temporal variability. Traditional methods relying on hand-crafted features often struggle with the complexity and irregularity of such datasets, leading to limited performance. In this study, a registration methodology is proposed that utilizes the PointNet++ architecture, designed to automatically learn features from point cloud data and overcome these limitations. The methodology involves three main steps: (1) sampling and canonicalization to preprocess local patches of point clouds, (2) designing a deep learning architecture to extract robust and distinctive features, and (3) registering point clouds using learned descriptors. Experimental validation was conducted using the ETH dataset and two agricultural datasets collected from test sites with cabbage and maize crops. The proposed method achieved a root mean square error (RMSE) of 4.21 cm in a cabbage field and RMSE values of 5.48 cm, 3.45 cm, and 2.06 cm at the seedling, jointing, and flowering growth stages of maize crops, respectively. Additionally, the method outperformed classical and recent techniques in terms of feature matching recall (FMR), demonstrating its reliability, accuracy, and potential for industrial applications. These results underscore the effectiveness of the proposed deep learning-based approach in handling the dynamic and nonrigid nature of agricultural environments.
Wildfires are becoming increasingly frequent and severe in Mediterranean regions, posing growing threats to both natural ecosystems and high-value agricultural systems such as vineyards. Following the August 2025 wildfire near Patras (Achaia, Western Greece), this study develops a rapid workflow for post-fire damage assessment in viticultural landscapes using ultra-high-resolution UAV RGB imagery (approximate to 1.5 cm GSD). A hybrid interpretable and explainable artificial intelligence (AI) framework was designed to compare traditional and deep learning methods for burn severity mapping. Here, explainability refers to methodological transparency and role separation between interpretable RGB-based analysis and deep learning segmentation, rather than post hoc inspection of neural network internals. Results demonstrate that the interpretable PCA-k-means workflow captures vineyard burn severity patterns consistent with U-Net semantic segmentation, while requiring minimal data and computational cost. The classical approach integrated vegetation indices derived from RGB bands (ExG, NGRDI, GRVI, VARI) with principal component analysis (PCA) and k-means clustering, providing interpretable severity zonation from minimal data. In parallel, a U-Net semantic segmentation model was trained to produce pixel-level delineations of burned and surviving canopy, serving as a high-fidelity benchmark. PCA-k-means effectively captured overall burn gradients (PC1 explaining 50.1% of variance), while the U-Net achieved high spatial accuracy (mIoU = 0.91) and improved boundary precision. The combination of both approaches enabled cross-validation and calibration of severity thresholds, revealing their complementarity for operational post-fire analysis. This hybrid AI framework demonstrates that low-cost UAV RGB imagery, when coupled with interpretable and deep models, can deliver fast, reproducible, and transferable assessments of wildfire impacts in agricultural systems, supporting early recovery decisions and resilience planning under climate-driven fire regimes.
Ground surfaces and retaining walls are necessary to monitor to mitigate against the risk of sediment disasters. Currently, the methods employed for monitoring of retaining walls rely heavily on visual inspection of cracks. This is time-consuming and subjective as the results vary depending on the inspector. Hence, it is necessary to establish a more efficient and versatile monitoring method. This paper proposes a surface deformation monitoring method using Structure-from-Motion Photogrammetry, in which a Digital Single-lens Reflex camera is attached to a Real Time Kinetic Global Navigation Satellite System (GNSS) to acquire images with precise position coordinates. In this study, point clouds were generated using Pix4DMapper and Agisoft Metashape and analysed in Cloud Compare. The two point clouds acquired at the same location at separate times, were aligned using the Iterative Closest Point algorithm. Change detection methods of the two point clouds were performed to visualise the surface deformation. The results showed that displacements can be seen from 5, 10, 15 and 20 mm test boards with errors of 0.28 to1.68 mm for Metashape and 0.29 mm to 2.84 mm for Pix4D at the retaining wall.
Monitoring vessel activity within Exclusive Economic Zones (EEZs) is essential for maritime security, environmental protection, and sustainable resource management. This study presents a novel framework that combines high-resolution aerial imagery, photogrammetric geolocation techniques, and deep learning-based object detection to detect and monitor vessels near maritime boundaries. Using the SeaDronesSee dataset, vessels are automatically detected with YOLOv8 and georeferenced through image metadata, enabling accurate transformation from image to geographic coordinates. Spatial queries with EEZ boundary datasets are then applied to identify potential violations. Experimental evaluation demonstrates a detection accuracy of 98.3%, with robust performance across varied vessel types and imaging conditions. The framework is further supported by a lightweight security layer to ensure reliable transmission of boundary violation alerts. This integration of photogrammetric image analysis, automated object detection, and geospatial boundary validation provides an efficient and scalable approach to maritime monitoring, contributing to the advancement of remote sensing and photogrammetric applications in marine surveillance.
Accurate and reliable localization is a key requirement for modern automated vehicles. In densely built-up urban environments such as city canyons, where satellite-based positioning methods are affected by signal interference or blockage, vehicle-mounted lidar can harvest the rich geometric features found on human-made structures and increase the availability of a localization solution. Lidar-based odometry and simultaneous localization and mapping (SLAM) approaches do, however, require loop closing or they suffer from accumulated error. A promising approach to address such problems is to employ voxelized High Definition (HD) maps, which provide a compact, effective, and low-maintenance solution. Despite the advantages of voxelized HD maps, their potential for localization through registration with lidar scans has not been explored. In this work, we evaluate the performance of various registration methods for lidar localization with voxelized HD maps. The results show that the registration methods that rely on matching local surface geometry fail when applied to voxelized point clouds. However, methods that match individual points or superpoints, namely GeoTransformer and two variants of the iterative closest point (ICP), achieve an acceptable performance with voxelized data and potentially support lane-level accuracy when provided with initialization within 10 m from the ground truth location.
For decades, Perspective-n-Point (PnP) algorithms have been widely used for camera pose estimation in incremental Structure-from-Motion (SfM) systems. The simplest example of PnP problems involves registering a new camera to two already-registered cameras in the 3D model, which forms a camera triplet and requires at least four 2D-3D correspondences (i.e., four 3-view points) to determine the pose of the new camera. However, this requirement is not always satisfied in challenging urban scenarios where only 2-view points are available due to insufficient overlap between views, leading to incomplete and fractured models. This work revisits the minimal solver that uses as low as six pure 2D correspondences for pose estimation and provides a comprehensive assessment across various scenes. We modify the solver to support both calibrated and uncalibrated cameras and incorporate Bundle Adjustment (BA) for camera triplet refinement. We name the refined solver Triplet-2-view-points (T2vP), as it leverages 2-view points to estimate the pose of the new camera within the triplet. Extensive experiments on diverse datasets demonstrate substantial improvements in model completeness when T2vP is integrated into existing SfM systems. The results highlight the effectiveness of T2vP in reconstructing challenging urban environments with weak 3-view overlap.
Current machine learning-based landslide susceptibility assessment heavily relies on supervised classification, which necessitates both landslide and non-landslide samples. However, the selection of non-landslide samples (negative samples) suffers from significant epistemic uncertainty and a lack of standardized criteria, introducing bias that compromises model reliability. To bridge this gap, this study proposes a novel framework using One-Class Classification (OCC), which eliminates the dependency on unreliable negative samples by training exclusively on landslide occurrences. We utilize a historical landslide dataset from Luding County, China, prior to the earthquake on September 5, 2022 as the training data, and post-earthquake landslide data as the testing data. We model the data using three one-class classifiers: One-Class Support Vector Machines (OCSVM), Isolation Forest (IForest), and One-Class K-nearest neighbors (OCKNN). Then we compare the results with supervised learning classification algorithms based on Support Vector Machines (SVM), Random Forest, and KNN. The results show that OCSVM has a higher recall rate (0.865) than SVM (0.639) for high susceptibility areas. IForest has a higher recall rate (0.903) than RandomForest (0.884). OCKNN performs the best with a recall rate of 0.968, surpassing KNN classification (0.923). Furthermore, we employ SHAP to interpret the OCKNN model, identifying elevation as the most influential factor in landslide susceptibility, followed by TRI and slope. This enhances the interpretability of the model and provides insights into the driving factors of landslides. The results demonstrate that the proposed one-class classification effectively addresses the issue of negative sample quality in traditional supervised learning. This study provides a new approach for landslide susceptibility assessment in data-scarce regions.
The integration of hyperspectral imagery and LiDAR data offers promising potential for multimodal feature learning in urban land cover classification. However, effectively extracting and fusing these heterogeneous data sources to fully leverage their complementary strengths remains a challenging problem. To address this issue, we propose the Cascade Encoder-Decoder Fusion Network (CEDFNet), a novel multimodal framework designed for high-precision urban land cover classification. CEDFNet employs two parallel Cascade Encoder-Decoder Networks (CEDNets) as its backbone, where the cascaded architecture enables progressive multi-scale feature integration and improves the discrimination of land cover patterns across different spatial resolutions. In addition, the model incorporates two specialized modules: the Complementary Feature Focusing Module (CFFM), which enhances cross-modal complementarity and produces high-quality fused representations, and the Dense Attention Branch (DAB), which adaptively captures both low-level and high-level attentive cues to further strengthen feature expressiveness. Experimental evaluations on the Houston 2018 and MUUFL Gulfport datasets demonstrate that CEDFNet consistently outperforms state-of-the-art baseline models, confirming its effectiveness and robustness in complex urban environments with diverse land cover distributions.
Band Selection (BS) is an essential method in the classification of hyperspectral images, as it effectively decreases spectral redundancy in hyperspectral remote sensing data, lowers computational expenses, and identifies the best band subsets that offer improved discriminative power from numerous spectral dimensions. Evolutionary algorithms (EAs), known for their strong search capabilities, have been effectively utilized as BS techniques in hyperspectral image analysis. However, many current EA-based BS methods encounter two significant issues: (1) a tendency to become trapped in local optima and experience premature convergence, and (2) a high sensitivity to the choice of initialization methods and hyperparameter settings, resulting in variable performance stability. To overcome these challenges, this study introduces a multi-strategy enhanced salp swarm optimization method (MSSA) for optimal spectral band selection, referred to as MSSA-BS. Initially, we improve the population initialization process to achieve a more even distribution of individuals within the initial population across the search space, which enhances diversity and reduces sensitivity to initialization. Furthermore, the algorithm incorporates a L & eacute;vy flight strategy and optimization of inertial weights to enhance search dynamics. This combined approach narrows the search area, speeds up evolutionary development, and aids in escaping local optima, thus boosting optimization effectiveness. Additionally, a chain follower mechanism is implemented to update the positions of the least effective individuals, further enhancing the algorithm's exploration capabilities. Together, these advancements systematically tackle the identified challenges. To assess the performance of MSSA-BS, extensive experiments are carried out on three standard hyperspectral image (HSI) datasets. The findings indicate that MSSA-BS achieves higher classification accuracy compared to various leading BS methods when used in conjunction with a support vector machine (SVM) classifier.
Although accurate classification of large-format, high-resolution remote sensing image is essential for land cover mapping, balancing computational efficiency with classification performance remains challenging. Traditional methods often incur high computational costs and achieve limited accuracy when applied to large-format data. This study addressed the core challenge in the classification of large-format and high-resolution remote sensing images through a graph neural network classification method that combines a parallel fine segmentation strategy with graph neighborhood relationship optimization. It improved the computational efficiency bottleneck and classification accuracy. Our methodology comprises three key components. First, an adaptive compactness parameter-based tiling method using simple linear iterative clustering (SLIC) generates uniform image patches through radiometric resolution downsampling and parallel allocation via Spark. Second, we propose a SLIC algorithm considering ground features (SLIC-GF), which employs the Otsu method and ratio vegetation index (RVI) to distinguish vegetation/non-vegetation pixels before fine segmentation. Finally, image objects are structured into graphs for classification via graph neural networks, with neighborhood-based correction of misclassified nodes. Experimental results from three high-resolution datasets show that our parallel segmentation strategy improves average computational efficiency threefold while reducing standard deviation (SD) and value range (R) by 41.5% and 51.9%, respectively. Compared to original SLIC, SLIC-GF improves achievable segmentation accuracy (ASA) by 3.53%, reduces under-segmentation error (UE) by 5.6%, and increases boundary recall (BR) by 2.8%. Furthermore, the graph attention network (GAT) with neighborhood optimization significantly enhances classification performance, yielding average improvements of 0.0165 in Kappa coefficient and 1.08% in overall accuracy (OA).
In multimodal remote sensing image (MRSI) matching, nonlinear radiometric distortion (NRD), scale/geometric inconsistencies, and illumination changes often cause false or missed correspondences. We propose a method that couples an improved self-similarity index map (SSIM) with absolute phase-orientation features. First, a feature-weighted aggregation jointly captures similarity and edge cues. We then fuse odd- and even-symmetric Log-Gabor filters to derive phase-congruency-based orientation and scale cues, and combine them with Sobel gradients to form a scale-adaptive absolute phase-congruency orientation gradient. Finally, we construct a rank-order self-similarity map (SRSIM) to strengthen rotational invariance. We evaluate the method on representative MRSI datasets with translation, scale, rotation, and illumination differences, and compare against five mainstream algorithms. The results show superior robustness under radiometric distortion, contrast variation, orientation reversal, and abrupt phase-extrema changes. Quantitatively, the average number of matched points (NCM) increases by more than 40%, the average matching success rate by 38%, and the average correct matching rate by 12.23%-31.56%, while the average root-mean-square error (RMSE) drops to 2.12 pixels. Overall, the approach markedly improves the accuracy and robustness of automatic multimodal remote sensing image matching.
The rapid development of deep learning techniques has revolutionized various remote sensing applications, which is especially true for change detection (CD) areas. Consequently, the past few years have seen a surge of deep learning change detection (DLCD) techniques with unparalleled improvements in precision, efficiency, and automation. Despite their huge success, these methods often follow a data-driven routine, where massive labeled data is required to guarantee network parameter learning. However, it is costly and labor-intensive to obtain sufficient labeled data for the CD task, especially pixel-level annotations. In this context, label-efficient DLCD (LE-DLCD) techniques have garnered increasing attention, which are capable of training CD networks with incomplete labels, inexact labels, or even no explicit labels. In this review, we conducted a comprehensive survey of state-of-the-art label-efficient DLCD methods, which are categorized into six schemes of semi-supervised CD , weakly supervised CD , self-supervised CD , active learning CD , few-shot CD , and unsupervised CD . Subsequently, each scheme is further categorized into finer subcategories for in-depth summarization and analysis. Next, we make systematic quantitative comparisons of typical LE–DLCD methods to provide valuable guidance for real-world scenarios. Finally, the challenges and future directions of LE-DLCD are presented in detail, which aims to shed light and inspiration on this area for the CD community. To facilitate transparency, we have shared the selected LE-DLCD methods via https://github.com/daifeng2016/Awesome-Label-efficient-Deep-Learning-Change-Detection-Methods .
Accurate and efficient 3D reconstruction of small-scale objects remains challenging due to intricate geometries, limited imaging volumes, and sensitivity to acquisition conditions. This study presents a quantitative comparison between two close-range photogrammetric acquisition methods: a conventional manual tripod setup and a custom-built, automated turntable platform controlled by an Arduino microcontroller. Four geometrically distinct objects were reconstructed using both approaches and analyzed through a unified Structure-from-Motion (SfM) workflow. Dimensional accuracy was assessed using reference measurements obtained with a digital vernier caliper (+/- 0.01 mm precision), while geometric fidelity was evaluated through Cloud-to-Cloud (C2C) surface deviation analysis. Results consistently favored the automated system. For instance, the used house object achieved a Root Mean Square Error (RMSE) of 0.18 cm with the turntable system versus 0.70 cm manually. The used jug, with complex occlusions, exhibited a C2C mean deviation of 0.411 cm in the manual method. The used jug showed a maximum deviation of 1.1 cm, while the used ceramic swan yielded the lowest mean error of 0.006 cm. In terms of efficiency, the automated platform reduced acquisition time by nearly 50%, improved repeatability, and minimized operator input. These findings underscore the potential of low-cost, semi-automated acquisition systems for improving the accuracy, reliability, and scalability of photogrammetric measurement workflows. The proposed system is especially well suited for technical education, low-budget laboratory environments, and object-scale documentation scenarios requiring consistent measurement standards.
The aerial three-linear push-broom camera, with its simple and rational structure, enables low-cost and rapid acquisition of high-resolution aerial images over large areas. It plays an irreplaceable role in agricultural surveys, forestry investigations, and large-scale topographic mapping. Particularly in China, the ongoing national natural resource surveys in regions such as the northwest and northeast have provided unprecedented opportunities for the application of aerial three-linear push-broom cameras. However, we have identified that stereo measurements based on Level 1 images suffer from issues such as accuracy loss and low efficiency in object points back-projection. To address these challenges, this study introduces a real-time epipolar fitting algorithm for local scenes and a fast back-projection method guided by a 3D spatial grid. The epipolar fitting algorithm dynamically adjusts image orientation based on flight direction (kappa), effectively eliminating edge vertical parallax and improving stereo viewing quality. The fast back-projection algorithm constructs a 3D grid to store the coordinates of corresponding image points, enabling efficient retrieval of image coordinates for any spatial point through localized search. Validation experiments conducted on eight datasets demonstrate that the proposed methods significantly enhance epipolar image quality and stereo mapping efficiency, reducing vertical parallax by an average of 80% and increasing back-projection speed by at least 2.6 times relative to conventional approaches.