
Detailed characterization of forest structure is critical for applications ranging from production forestry and carbon monitoring to fire risk analysis or habitat and biodiversity assessments. Ground-based LiDAR technologies, such as terrestrial (TLS) and mobile (MLS) laser scanning, enable precise 3D mapping of trees, understorey, and ground, offering rich structural detail. Yet, extracting meaning from these dense, unstructured point clouds remains a persistent challenge that requires reliable and scalable segmentation methods.Earlier segmentation approaches have struggled to generalize across diverse forest types and acquisition conditions. Recent advances in deep learning (DL), particularly architectures designed for 3D data, offer new opportunities for robust and scalable forest structure modelling using ground-based point clouds. Nevertheless, their implementation in forestry remains limited, partly due to the difficulty of obtaining large, accurately labelled point clouds for training and evaluation.In this study, we benchmark the performance of six different implementations of four state-of-the-art DL architectures (PointNeXt, SuperPoint Transformer, Point Transformer V3 and OA-CNN) on the public SegmentedForests dataset, which contains over 850 million labelled points from 14 plots representing a range of coniferous and broadleaf forests scanned with both TLS and MLS. We assess segmentation performance across four ecologically meaningful classes: ground, understorey, stems, and canopy. The best-performing model, Point Transformer V3, achieved an overall mean Intersection over Union (mIoU) of 81.9%, confirming the strong potential of transformer-based architectures for forest point cloud segmentation.The benchmark analysis considers how segmentation accuracy varies across forest type, sensor, understorey density and forest maturity. We observe clear differences in segmentation difficulty across conditions, with coniferous plots being consistently easier to segment than broadleaf plots, and TLS point clouds yielding higher accuracies than MLS acquisitions. Furthermore, a dense understorey was found to significantly improve class-wise saliency despite increased terrain occlusion. Cross-testing confirmed high model stability when generalized to entirely unseen scenes. Together, these results provide the first comprehensive benchmark of modern DL models for forest segmentation from ground-based LiDAR point clouds and highlight key factors shaping their performance across forest structural diversity.
The Ramsar site “Sistema Delta Estuarino del Río Magdalena Ciénaga Grande de Santa Marta” (SDERM-CGSM), located in the Colombian Caribbean, is a wetland of international importance due to its ecological and environmental value. However, the available thematic cartography to support territorial decision-making is limited, and alternative approaches for land cover monitoring and assessment are required. For this reason, a multi-source remote sensing methodology is presented, integrating PlanetScope (PS), Sentinel-1 (S1), and Sentinel-2 (S2) satellite data for land use and land cover change (LULCC) analysis over the 2016–2024 period in the SDERM-CGSM. Data processing was conducted using Google Colab (GC), ArcGIS Pro, and Google Earth Engine (GEE). Synthetic SWIR 1 (∼1610 nm) and RedEdge 1 (∼740 nm) bands were estimated for each PS image using multiple linear regressions (MLR) based on S2 data, achieving R2 values between 0.85 and 0.96. The classification workflow was performed through variable selection using a Genetic Algorithm (GA) and the application of a Random Forest (RF) model with a three-stage temporal cross-validation training approach, achieving an overall accuracy (OA) of 92.99% and a Kappa index (K) of 92.36%. Finally, the LULCC diagnosis was carried out through the design and implementation of the Directional Spatial Similarity Index (DSSI), which highlights potential anomalous changes in the classified time series by considering the proportional spatial relationship between the evaluated land covers. It was found that approximately 8% of the Ramsar site may correspond to areas exhibiting unusual changes. In this way, a methodological workflow is presented that enables multitemporal analysis using multi-source data and provides foundational inputs for understanding land cover trajectories in a specific site, offering an initial tool for the development of territorial planning instruments.
Collecting spatially consistent multi-modal remote sensing (MMRS) images remains challenging due to different sensors vary in the imaging principles and acquisition times. This hinders the development of data-driven MMRS technologies, which rely on large-scale training samples. This paper proposes MMDiff, the first text-driven diffusion framework explicitly designed for jointly generating structurally consistent optical (OPT), synthetic aperture radar (SAR), and infrared (IR) remote sensing images from a single text prompt via cross-modality spatial feature transfer. MMDiff first trains the OPT branch with paired optical image–text data to capture rich semantic content, and then trains the SAR/IR branches with simple modality-specific text templates to learn the corresponding style attributes, without relying on complex linguistic descriptions. Specifically, we introduce a LoRA-based modality translation adaptation mechanism to translate the style attributes of optical spatial representations to SAR and IR style attributes while preserving the underlying semantic content. The translated representations are then transferred into the SAR and IR generation branches through the proposed spatial feature transfer mechanism, enabling rich spatial details in the generated SAR/IR images while maintaining cross-modal spatial consistency. Extensive experiments demonstrate that MMDiff achieves superior image quality in terms of modality similarity and semantic consistency compared to the state-of-the-art methods. Furthermore, MMDiff benefits downstream data-driven MMRS applications, e.g., multi-modal image fusion and object classification. Code is available at: https://xinr-tang.github.io/MMDiff-homepage/.
Persistent scatterer (PS) selection is a fundamental step in interferometric synthetic aperture radar (InSAR) precise earth observation, where the quality and density of selected PSs directly influence the monitoring performance. Although deep learning has advanced rapidly, its application to PS selection remains limited; the amplitude and phase information are still underexplored in both temporal and spatial domains. To conclude, a PS selection method based on temporal-spatial vision transformers (PSSformer) was proposed. The backbone consists of dual input branches and a fusion network. In the phase input branch, a spatiotemporal baseline position encoding (STB-PE) is introduced to model interferogram relationships, and a phase-consistent self-attention (PCSA) is proposed to guide the model’s optimization based on explicit physical modeling. The fusion network integrates multi-scale features via a multi-head cross-attention (MCA) and an amplitude-phase fusion (APF) module. A dataset based on multi-sensor amplitude images and interferometric phase was constructed for training. On the in-distribution (ID) test set, PSSformer outperforms two baseline methods, achieving recall improvements of 13.54 % and 7.48 %. In addition, an experiment using Sentinel-1 data in two deformation regions, validated by global navigation satellite system (GNSS) stations, demonstrates that PSSformer significantly improves the number of reliable PSs (49.28 %) and enhances monitoring accuracy, while reducing the PS selection time to only 1.26 % of StaMPS.
As a core component of simultaneous localization and mapping (SLAM), LiDAR-based place recognition (LPR) is essential for mitigating cumulative drift. However, existing methods struggle with stability and generalization across diverse environments. Thus, this paper proposes a novel descriptor, the dual stable triangle descriptor (DSTD), which integrates vertex-based triangles (V-STD) and voxel centroid-based triangles (C-STD). This design enables reliable descriptor generation in both structured and moderately feature-degraded environments. For loop candidate retrieval, we develop a dual-branch hash voting strategy that adaptively switches between two descriptors based on environmental characteristics. Loop closures are further validated using multiple metrics, rigid transformation, scene coverage, and spatial overlap, thus mitigating the bias introduced by fixed overlap thresholds. We introduce a sliding-window optimization to adaptively re-rank loop candidates based on reliable matches within and at the boundaries of the window, while exploiting temporal continuity to recover missed loops and correct false detections. We integrate all modules into a framework and carry out extensive experiments on the MCD, KITTI, and MulRan datasets. The results demonstrate that DSTD achieves the best average performance across datasets. Compared with the strongest baseline, DSTD improves the average Top-1 recall, achieving a maximum gain of 6.7% on a sequence (tuhh_day_04) from the MCD dataset. In addition, it improves the average F1 score by 2.9%, 6.2%, and 4.2% under different evaluation settings, achieving a maximum gain of 8.9% and over 10% for long-range loop closures. The source code will be available at https://github.com/ShiPC-AI/DSTD.
Annual maps of water-related land cover types, including surface water bodies, wetlands, and paddy rice, are essential for biodiversity conservation, climate regulation, food security, and infectious disease risk assessment. However, the existing large-scale data products have moderate accuracy and incomplete representation of these land cover types, and are not up to date. Moreover, those mapping approaches that rely on knowledge-based methods often suffer from insufficient good-quality observations within key phenological windows. The machine learning models are constrained by the lack of good-quality and temporally consistent training data. In this study, we proposed a two-stage mapping framework that explicitly combines knowledge-based training data generation and deep learning algorithms to leverage their complementary strengths. We first applied the established knowledge-based rules (classifiers) to produce maps of individual land cover types, which were geo-processed and filtered to generate training data. The resultant training data were used to train a deep neural network with multi-source time-series observations for classification of multiple land cover types. Using this framework, we produced a 30 m water-related land cover map for Northeast Asia in 2021, named OU-WRLC, achieving an overall accuracy of 93.8 ± 0.2%. Northeast Asia contained a total of 1,239,234 km2 of water-related land cover types, of which 17% was yearlong surface water bodies, 76% wetlands (15% seasonal open-canopy marshes and 61% yearlong closed-canopy marshes), and 7% paddy rice. The map showed strong consistency with existing products for surface water bodies and paddy rice, while providing improved spatial details. It captured extensive yearlong closed-canopy marshes that are less consistently represented in the existing datasets. Overall, this study provides a comprehensive mapping approach, supporting regional-scale environmental monitoring and land management.
In recent years, unified generative models have achieved unprecedented success in photorealistic synthesis. However, since satellite imagery serves as strict physical measurements, the geospatial domain requires capturing the complex multi-modal symbiosis across diverse sensors to achieve rigorous structural and spectral fidelity rather than mere visual realism. Currently, constrained by multi-modal data scarcity and inherent modality barriers, existing studies often fail to reconstruct underlying physical interactions and remain restricted to superficial perceptual synthesis. To address these challenges, we present the MMG-5 dataset and the unified SMS-Diff model. MMG-5 is a globally sampled, strictly aligned benchmark integrating five geospatial modalities (LULC, RGB, SWIR, SAR, DEM) to bridge the multi-modal data gap. Building on MMG-5, SMS-Diff transforms geospatial image translation into true multi-modal physics-aware reconstruction. Specifically, SMS-Diff incorporates a Spectral Manifold Injector (SMI) and a Latent Physics Resonance Loss (LPRL) to explicitly decouple spectral laws from spatial geometry, successfully bridging global frequency alignment and local structural precision. Extensive experiments demonstrate that SMS-Diff achieves consistent improvements over strong baselines, achieving an SSIM of 0.315 and LPIPS of 0.307. Beyond generic computer vision metrics, we integrate classic geospatial physical indicators with downstream task utility as the dual benchmark for validating geospatial consistency. Crucially, these task-specific evaluations reveal a significant “Denoising Effect Tendency” in synthesized features, yielding partial accuracy inversions over raw sensor data. The model shows zero-shot transfer ability for change detection to some extent, while also scaling to reconstruction across 10 diverse modalities and multiple downstream tasks via low-resource fine-tuning. In summary, SMS-Diff drives a paradigm shift toward internally consistent physics-aware reconstruction, proving the feasibility of generative models to simulate the real physical world.
Accurate and up-to-date building polygon maps are essential for urban planning, disaster response, and large-scale geospatial analysis, yet creating them automatically across diverse regions remains challenging. To address this challenge, we present the P3 dataset, a large-scale multimodal benchmark for building vectorization, constructed from aerial LiDAR point clouds, aerial images, and vectorized 2D building outlines, collected across three continents. The dataset contains over 10 billion LiDAR points with decimeter-level accuracy and RGB images at a ground sampling distance of 25 centimeters. While many existing datasets primarily focus on the image modality, P3 offers a complementary perspective by also incorporating dense 3D information. We demonstrate that LiDAR point clouds serve as a robust modality for predicting building polygons, both in hybrid and end-to-end learning frameworks. Moreover, fusing aerial LiDAR and imagery further improves accuracy and geometric quality of predicted polygons. The P3 dataset is publicly available, along with code and pretrained weights of three state-of-the-art models for building polygon prediction at https://github.com/raphaelsulzer/PixelsPointsPolygons.
A consistent trend in polarimetric synthetic aperture radar (PolSAR) remote sensing is to reduce the dependence on dense annotations while maintaining reliable target monitoring capability in complex scenes. Sparse weakly semi-supervised learning addresses this demand by exploiting limited weak instance annotations, but its application to PolSAR ship detection remains constrained by insufficient discriminative priors, incomplete target evidence, and the loss of valuable supervisory signals. To address these issues, this paper proposes a sparse weakly semi-supervised tri-evidence recovery (SWS-TER) network for oriented ship detection. First, an annotation-free contrastive prior construction is introduced to construct region-level reliability priors by combining polarimetric superpixel region cues with momentum contrastive encoding. Second, a sparse-weak evidence completion student is designed to alleviate the over-reliance on strong local scatterers through scale adaptive context compensation and scattering keypoint graph. Third, an uncertainty-guided supervision recovery teacher is developed to transform the hard rejection of uncertain pseudo-labels into a soft semantic recovery process, preventing potentially informative samples from being discarded. Extensive experiments on the GaoFen-3 and RADARSAT − 2 images show the effectiveness of the proposed solution. SWS-TER achieves 75.5 % AP, which is 4.8 % higher than the second-best SPWOOD. As the annotation ratio further decreases, SWS-TER still achieves an AP of 60.7 % under the extremely sparse setting with less than 1 % annotations, demonstrating a more pronounced performance advantage. The G-R dataset is publicly available at https://drive.google.com/file/d/1-jJ_NxLoizAElMlTl6Aaf4Sn2zxAW-PI/view?usp=drive_link, and the code is publicly available at https://github.com/YucongHe462/SWS-TER.