Thick cloud cover severely occludes surface information in optical remote sensing images, limiting the reliability of subsequent applications. Recent advances in multitemporal cloud removal methods have made notable progress. However, features from existing methods are often contaminated by cloud interference and temporal variation noise, masking the true surface information. Directly fusing such features limits the use of useful information and introduces artifacts. Furthermore, discriminators typically rely on single-scale constraints, which cannot adequately supervise multiscale feature recovery, resulting in poor structure and texture restoration. To address these challenges, this article proposes adaptive feature purification and a dual-branch hierarchical discriminator for multitemporal multiscale cloud removal (FDH-CR), which reduces artifacts from channel conflicts and provides more comprehensive multiscale supervision, thereby reconstructing missing surface information in scenes affected by clouds using multitemporal data. To tackle mixed channel information, a feature purification and enhancement block is introduced to utilize an adaptive weighting mechanism to enhance information-rich channels while suppressing redundant and conflicting features, thereby strengthening useful feature representations and reducing artifacts. In addition, to address the limited supervisory capability of single-scale discriminators, a dual-branch hierarchical discriminator (DBHD) offers multiscale supervision, helping generated images preserve structural and textural realism. A feature matching loss based on DBHD is innovatively adopted in the cloud removal framework, which aligns intermediate features from the discriminator to enhance semantic consistency and stabilize adversarial training. Experiments on two public remote sensing datasets demonstrate that FDH-CR achieves superior performance both quantitatively and visually, validating its effectiveness and potential.
Bare soil is an important indicator for evaluating urbanisation and is essential for research on sustainable development. Traditional bare soil indices have limited applicability across different geographic regions and soil types and are sensitive to interference from impervious surfaces. Consequently, few indices are widely adopted for bare soil identification in both urban and rural environments. In this study, a new index, called the Enhanced Normalized Difference Bare Soil Index (ENDBSI), is developed by integrating the spectral characteristics of various land cover types across multiple study areas. The results show that the proposed index enables accurate bare soil extraction in both urban and rural environments. By validating the ENDBSI across 14 study areas in China, each dominated by a different soil type and consisting of urban and rural areas, the index is shown to consistently outperform existing bare soil indices for bare soil identification, namely, the Bare Soil Index (BSI), Normalized Difference Bare Soil Index (NDBSI), and Product Index for Dark Soil (PIDS). Moreover, it achieves a greater spectral contrast between bare soil and impervious surfaces, improving the accuracy of bare soil mapping. The ENDBSI demonstrates excellent performance in bare soil identification across all 14 study areas, with an average spectral discrimination index between bare soil and non-bare soil exceeding 2.41 and an average identification accuracy of 94.33%. The results highlight the index’s strong applicability and transferability across diverse environments. These findings suggest that the ENDBSI is suitable for the large-scale remote sensing monitoring of bare soil and has the potential to significantly improve bare soil observation and mapping across China.
High-precision matching between images and point clouds underpins 3D reconstruction, robotic localization and navigation, and multimodal perception. However, a pronounced modality gap exists between 2D images and 3D point clouds, and existing methods often fail to fully exploit spatial depth geometry within a unified modality, resulting in limited spatial perception and constrained extraction of cross-modal features with spatial geometric consistency. The problem is especially acute in weak-texture and highly similar regions, where obtaining well-aligned cross-modal descriptors is difficult, thereby impeding matching accuracy. To address these issues, we propose 2D3D-SPMatch, a novel cross-modal geometric consistency based on spatial perception for detector-free image and point cloud matching networks. For the first time, spatially aware depth information is used to explicitly model spatial geometric consistency, enabling robust image and point cloud matching. The key idea is to generate an image depth estimated map with a monocular depth estimation model while projecting the 3D point cloud to obtain a projected depth map and use these depth maps to encode cross-modal spatial geometric information and design an image and point cloud spatial geometry attention mechanism that strengthens the geometric consistency and discriminability of cross-modal features. Concretely, we first feed the image depth estimated map and the point-cloud projected depth map into structurally matched Spatial Geometric Encoders (SGE) to extract geometry features with structural consistency, thereby improving spatial perception and promoting consistent alignment of cross-modal features. Next, to mitigate the modality gap and strengthen cross-modal spatial geometric consistency, we design a Spatial Geometry Attention fusion (GeoAttn) module. GeoAttn fuses the geometry features derived from the image and point-cloud depth maps into their respective feature-extraction networks, guiding the model to focus on cross-modal geometrically consistent cues, effectively enhancing geometric expressiveness in weak-texture and highly similar regions while reducing the impact of modality differences. Our approach substantially improves the accuracy and robustness of cross-modal matching and provides a reliable foundation for precise downstream registration. Finally, extensive experiments on 7Scenes and RGB-D Scenes V2 showed registration recall rates of 81.6% and 62.8%, respectively, validating the method’s superior ability to reinforce cross-modal spatial geometric consistency and alleviate modality discrepancies in complex scenes, while achieving the best and overall leading performance under a unified evaluation protocol when compared with the latest representative methods. The source code will be made publicly available https://github.com/RSDPLab/2D3D-SPMatch.
We propose self-supervised rotation-invariant descriptors based on mixed rotation-equivariant convolutional neural networks (CNNs) (MRDes) for heterogeneous remote sensing image matching between visible (VIS) and near-infrared (NIR) images. Existing methods struggle with large rotation variations for VIS-NIR image matching, particularly when georeferenced information is unavailable or inaccurate, limiting their practical applicability. To address this, MRDes employs rotation-equivariant CNNs to extract equivariant features and construct robust rotation-invariant descriptors. Specifically, we design a mixed learning strategy that integrates explicit and implicit equivariance to optimize feature representations, while a contrastive loss enhances their discriminability by refining the distances between positive and negative samples. Experimental results on benchmark datasets demonstrate that MRDes significantly outperforms state-of-the-art methods, achieving a 70.9% improvement in matching success rate over XoFTR and exhibiting strong generalization to unseen UAV data.
Remote sensing retrieval of lake chlorophyll-a (Chl-a) concentration is essential for eutrophication assessment and water quality management, yet conventional models suffer severe degradation in large-scale, long-term monitoring due to domain shifts arising from inconsistent atmospheric correction (regional aerosol variability) and spatiotemporal heterogeneity in aquatic optical properties. This study pioneers Test-Time Training (TTT) for robust lake Chl-a retrieval. The TTT dual-task architecture—a primary regression head paired with a self-supervised spectral reconstruction head—enables real-time adaptation to test-domain distributions without retraining, effectively compensating for residual atmospheric errors and regional apparent optical properties (AOP) distortions. Based on the proposed algorithm, this study comprehensively analyzed the spatiotemporal evolution of Chl-a in 2,621 lakes (≥1 km2) in China from 2018 to 2025. The results showed the proposed model achieved superior performance in Chl-a retrieval (R2 = 0.79, RMSE = 9.99 μg/L), outperforming conventional machine learning approaches by approximately 28 %-36 % and underscoring its robustness in optically complex inland waters. National Chl-a distribution exhibits a clear east-high, west-low pattern, with mean concentrations of 31.29 μg/L in the Eastern Plain versus 13.83 μg/L on the Tibetan Plateau. Regional analysis indicated that most regions exhibited a downward trend in Chl-a concentrations, suggesting an improvement in water quality. Bloom risk assessment, based on over 170,000 observations, indicates that the Eastern Plain lakes face “extremely high risk” (44.7 % bloom occurrence), requiring targeted management. The proposed lake Chl-a remote sensing retrieval algorithm provides a new paradigm for fine-grained global water quality monitoring and spatiotemporal analysis, holding landmark significance for global water environment surveillance and proactive management. The code is available at https://github.com/KustAIRS/SR-TTA.
Intra-day high-temporal-resolution mountainous land surface temperature (MLST) is a key surface parameter for studying climate change at both global and local scales over mountainous areas. However, the current high-resolution MLST retrieved from satellite data is the instantaneous value at the time of the satellite transit, and there is no model suitable for MLST diurnal scale expansion. Thus, a four-parameter physical diurnal MLST scaling extension model was developed by considering topographic effect and environmental radiation contributions, and integrating the energy balance equation with the heat conduction equation. The four parameters include a factor (& micro;) linking ambient temperature to environmental forcing intensity, two parameters (eta(0) and eta(1)) related to meteorological and climatic conditions, and the thermal inertia parameter (P). Validation results indicate that the proposed diurnal scale extension model achieves higher prediction accuracy at mountainous in situ observations than two existing physical models, with accuracy improvements of up to 0.59 K, particularly at sunrise. More importantly, compared to the semi-empirical model, the diurnal MLST scale expansion model proposed demonstrates greater flexibility in predicting MLST at any time of day. In addition, the developed diurnal scaling model theoretically requires only four observational inputs for its implementation. However, as the number of input observations decreases, prediction accuracy declines, with losses ranging from 0.14 K to a maximum of 1.82 K. Therefore, to ensure stable model performance, using more than four observational inputs is recommended.
The SAR pixel offset tracking (POT) technique, which exploits synthetic aperture radar (SAR) intensity images, is an effective tool for measuring large-gradient displacements by overcoming phase decorrelation and phase unwrapping errors that limit interferometric synthetic aperture radar (InSAR) measurements. However, the accuracy of POT depends on image resolution, and its performance degrades significantly when using medium-resolution SAR data such as Sentinel-1, often leading to mismatches, noisy displacement fields, and reduced reliability in capturing localized landslide deformation. To address this challenge, we propose an autofocusing-based adaptive time-series POT (AF-TSPOT), which integrates autofocusing-based temporal stacking into a time-series POT framework for landslide deformation retrieval from medium-resolution SAR image sequences. By refocusing temporally stacked SAR intensity images at the pixel level and integrating displacement estimates from multiple temporal baselines, AF-TSPOT suppresses matching noise and reconstructs a more stable displacement time series. Application to the pre-failure deformation of the 2018 Baige landslide along the Jinsha River demonstrates that AF-TSPOT retrieves deformation patterns comparable to those derived from high-resolution ALOS-2 data, reducing the root mean square error (RMSE) in stable areas by 54.16
Deep learning approaches that jointly learn feature extraction have achieved remarkable progress in image matching. However, current methods often treat central and neighboring pixels uniformly and use static feature selection strategies that fail to account for environmental variations. This results in limited robustness of descriptors and keypoints, thereby affecting matching accuracy. To address these limitations, we propose a robust joint optimization network for feature detection and description in optical and SAR image matching. A center-weighted module (CWM) is designed to enhance local feature representation by emphasizing the hierarchical relationship between central and surrounding features. Furthermore, a multiscale gated aggregation (MSGA) module is introduced to suppress redundant responses and improve keypoint discriminability through a gating mechanism. To address the inconsistency of score maps across heterogeneous modalities, we design a position-constrained repeatability loss to guide the network in learning stable and consistent keypoint correspondences. Experimental results across various scenarios demonstrate that the proposed method outperforms state-of-the-art techniques in terms of both matching accuracy and the number of correct matches, highlighting its robustness and effectiveness.
Existing generative adversarial network-based methods for SAR-to-optical image translation typically utilize a single-branch generator architecture, which makes it challenging to simultaneously ensure the authenticity of local details and the consistency of global information, often resulting in generated images suffering from the loss of land cover information and reduced edge clarity. To address this limitation, this paper proposes a SAR-to-optical image translation method utilizing a dual-generator framework based fusion optimization (A Dual-Generators Framework Based Fusion Optimization cGAN, DGFO-cGAN). The method aims to achieve dual fidelity at both global and local levels through information complementarity between the two branches of the network. The method consists of a Detail-Preserving subnetwork (DP-GAN) focused on preserving fine structure details of images and a Holistic-Balanced subnetwork (HB-GAN) dedicated to ensuring global information consistency. By performing pixel-level weighted fusion and optimization on the outputs of the two subnetworks, ensuring that the images generated by DGFO-cGAN possess high fidelity in both local texture structural details and global information. The experimental results demonstrate that, on the selected three publicly available datasets and a our proprietary dataset, the proposed method outperforms 11 current mainstream image transformation methods in terms of four evaluation metrics: PSNR, SSIM, LPIPS, and RMSE. It is capable of generating optical images with clearer structures and richer details.
The technology of optical and Synthetic Aperture Radar (SAR) remote sensing images change detection (CD) has vast potential applications in natural resources monitoring and disaster investigation. However, due to their inconsistent imaging mechanisms, challenges such as poor extraction of local features and inaccurate identification of change regions arise. A selective kernel convolution with a global attention mechanism for multitask CD in optical and SAR images is proposed. A novel SKGAM module has been designed, which integrates selective kernel (SK) convolution and global attention mechanism (GAM). This module allows the network to flexibly and efficiently capture key local details and global features when processing complex data, thereby enhancing its representational ability. The triple attention mechanism (TAM) is introduced to capture interactions between channels and spatial dimensions, enhance the feature representation of the change regions, and suppress the feature representation of the unchanged regions. Furthermore, an end-to-end multitask network architecture combining image translation and CD is constructed to address problems such as high data demand, low training efficiency, and intermediate intervention, with facilitation provided by the effective coordination mechanism between the dual tasks. Experiments are conducted on four heterogeneous remote sensing datasets, with M-UNet, DTCDN, TDSCCNet, and MTCDN selected for comparison. The Overall Accuracy (OA) of the proposed method is 99.26
Hyperspectral image (HSI) captured by uncrewed aerial vehicles (UAVs) is distinguished by superior spatial resolution and intricate spectral detail, with widespread applications in precise environmental monitoring, agricultural management, and biotic stress detection. The classification of UAV-acquired HSIs constitutes the cornerstone of these sophisticated monitoring endeavors. However, the ultrahigh spatial resolution inherent to these images engenders considerable spatial heterogeneity and spectral variability. Existing deep-learning classification methods overlook multilevel semantic features and global contextual information of images, leading to significant issues of salt-and-pepper noise and recurrent misclassifications. To address these challenges, this article proposes a novel classification framework, multihierarchical semantics Mamba (MHS-Mamba), tailored specifically for UAV-based HSI classification. The framework integrates two innovative components: the multihierarchical semantics spectral-spatial network, which adeptly extracts spatial, spectral, and semantic features from UAV-derived HSIs, and the linear spectral Mamba module, which effectively models and amalgamates short- and long-range spectral dependencies. The experimental results on agricultural and coastal urban scene datasets demonstrate that the proposed method achieves the overall accuracies of 96.58%, 96.98%, and 95.47%, outperforming state-of-the-art approaches by 1.99%, 3.22%, and 3.82%, respectively. These results highlight the robustness and generalizability of the proposed method, which holds strong potential for practical applications in precision agriculture and fine-scale land-cover monitoring.
The multimodal remote sensing image matching is crucial for many applications. However, nonlinear intensity distortion (NID) significantly impairs matching performance, especially when dealing with scale and rotation variations. To address this challenge, we propose a global-to-local invariant feature transformation (GLIFT) method for multimodal remote sensing image matching. The method consists of three key components: feature detection, global search, and local search. First, we introduce a fast dominant orientation assignment approach, which ensures rotational invariance while reducing computational costs. Next, we design a 3-D descriptor structure that effectively integrates both local region information and keypoint self-information, enhancing the robustness and discriminability of the descriptor. To overcome the limitations of image pyramids in handling scale variations in multimodal remote sensing images, we propose a local multiregion description strategy that adapts well to scale changes. In addition, we construct a pixel-based descriptor vector and present a local search strategy to identify optimal matching point pairs within local regions, which effectively improves the matching accuracy. Finally, we validate the matching performance of GLIFT by conducting experiments and comparing it with eight state-of-the-art multimodal matching algorithms on various datasets. Extensive results demonstrate that our method effectively addresses the challenges posed by rotation and scale variations in multimodal remote sensing images. It significantly enhances the number of correct matches (NCMs), matching accuracy, and matching precision. Our code is available at: https://github.com/wdzsc/GLIFT
Soil organic carbon density (SOCD) in swamp wetlands is a critical indicator for assessing global carbon stocks. In plateau wetlands, challenges such as dense vegetation cover and fragmented land distribution complicate SOCD research. The availability of high-resolution optical and radar satellite data introduces new possibilities for precise carbon stock predictions. This study proposes a framework that combines multisource remote sensing data with the sparrow search algorithm random forest (SSA-RF) algorithm to predict SOCD in plateau swamp wetlands. It also compares the effectiveness of laboratory spectroscopy and multisource remote sensing in monitoring SOCD. We integrated 24 features from Sentinel-1 (S1), Sentinel-2 (S2), topographic, and climatic data, along with spectral data ranging from 550 to 1400 nm, to construct the SSA-RF model and map the SOCD distribution of swamp wetlands in Dianchi Basin. Additionally, we estimated the total soil organic carbon (SOC) stock in these wetlands. The results indicate that the multisource remote sensing SSA-RF model (S1+ S2+ topographic + climatic SSA-RF) achieved an R-2 of 0.76, a root-mean-square error (RMSE) of 1.14, a mean absolute error (MAE) of 0.65, and a residual predictive deviation (RPD) of 1.98. Compared to the spectral model, this model improved the R-2 by 0.24 and the RPD by 0.5. Relative to the S1+ S2 SSA-RF model, the R-2 is increased by 0.15, and the RMSE is decreased by 0.46. The total SOC stock of the swamp wetlands in the Dianchi Basin was estimated to be 1.32x 10(5) t. This study provides a new framework for predicting carbon stocks in plateau wetlands, offering a reference for global wetland carbon sink assessments.
The ideal goal of generalizable image matching is to achieve stable and efficient performance in unseen domains. However, many existing learning-based optical-SAR image matching methods, despite demonstrating effectiveness in specific scenarios, often exhibit limited generalization and face challenges in adapting to practical applications. Repeatedly training or fine-tuning matching models to address domain differences not only lacks elegance but also incurs additional computational overhead and data production costs. In recent years, foundation models have shown significant potential for enhancing generalization. However, the disparity in visual domains between natural and remote sensing images poses challenges for their direct application. Consequently, effectively leveraging foundation models to improve the generalization of optical-SAR image matching remains a critical challenge. To address these challenges, we propose PromptMID, a novel approach that constructs modality invariant descriptors using text prompts based on land use classification as priors information for optical and SAR image matching. PromptMID consists of several key stages. Firstly, we finetune the diffusion model (DM) using we collected optical images, SAR images, and text prompts data to obtain the PromptDM model. Secondly, we construct modality-invariant descriptors by integrating multi-scale latent diffusion features extracted from the fine-tuned PromptDM model with multi-scale features derived from pre-trained visual foundation models (VFMs). To efficiently fuse local-global and texture-semantic features of varying granularities, we design a feature aggregation module (FAM) that ensures comprehensive feature representation. Finally, the discriminative power of the descriptors is enhanced through contrastive learning loss functions, aiming to improve the robustness and generalization of matching. Extensive experiments conducted on optical-SAR image datasets from five diverse regions demonstrate that PromptMID outperforms state-of-the-art matching methods, achieving superior performance in both seen and unseen domains while exhibiting strong cross-domain generalization capabilities. The source code will be made publicly available https://github.com/HanNieWHU/PromptMID.
Land Surface Temperature (LST) is a parameter retrieved through the thermal infrared band of remote sensing satellites, and it is a crucial parameter in various climate and environmental models. Compared to other multispectral bands, the thermal infrared bands have lower spatial resolution, which limits their practical applications. Taking the Heihe River Basin in China as a case study, this research focuses on LST data retrieved from the SDGSAT-1 using the three-channel split-window algorithm. In this paper, we propose a novel approach, the Information-Guided Diffusion Model (IGDM), and apply it to downscale the SDGSAT-1 LST image. The results indicate that the downscaling accuracy of the SDGSAT-1 LST image using the proposed IGDM model outperforms that of Linear, Enhanced Deep Super-Resolution Network (EDSR), Super-Resolution Convolutional Neural Network (SRCNN), Discrete Cosine Transform and Local Spatial Attention (DCTLSA), and Denoising Diffusion Probabilistic Models (DDPM). Specifically, the RMSE of IGDM is reduced by 55.16%, 51.29%, 48.39%, 52.88%, and 17.18%. By incorporating auxiliary information, particularly when using NDVI and NDWI as auxiliary inputs, the performance of the IGDM model is significantly improved. Compared to DDPM, the RMSE of IGDM decreased from 0.666 to 0.574, MAE dropped from 0.517 to 0.376, and PSNR increased from 38.55 to 40.27. Overall, the results highlight the effectiveness of the auxiliary information-guided SDGSAT-1 LST downscaling diffusion model in generating high-resolution remote sensing LST data. Additionally, the study reveals the spatial feature impact of different auxiliary information in LST downscaling and the variations in features across different regions and temperature ranges.
Unmanned aerial vehicle (UAV) hyperspectral images are endowed with abundant spectral information and spatial texture details, which are crucial for the precise classification and monitoring of terrestrial features. Despite the imagery offers high spatial and spectral resolution along with exceptional mobility, UAV-borne hyperspectral images exhibit intricate intraclass spectral variability and high spatial heterogeneity among features, which consequently poses a significant obstacle to the accurate classification. Simultaneously, the high spectral resolution of UAV-borne hyperspectral images, constrained by sensor payload, often results in a relatively low signal-to-noise ratio, posing challenges for traditional deep-learning models in effectively extracting features for fine classification tasks. To address the obstacles presented, this article innovatively proposes a multifeature fusion attention residual hybrid model for UAV-borne hyperspectral imagery classification. The proposed network initially employs an innovative sequential mechanism that leverages multiscale filters and attention residuals, and followed by the fusion of a multicascade 2-D-3-D parallel structure to achieve a semantic-spectral-spatial multifeature fusion, aimed at enhancing the fine classification of UAV-borne hyperspectral imagery. Experimental results using HySpex data highlight the superiority of the proposed approach, achieving an overall accuracy of 80.40%, which represents a significant average accuracy improvement of 5.22% over existing advanced classification methods. Furthermore, the proposed network demonstrates robust generalization and stability across UAV-borne hyperspectral imagery captured by diverse sensors with different spatial and spectral resolutions.
Recently, cross-scene hyperspectral image classification(HSIC) via domain adaptation is drawing increasing attention. However, most existing methods either directly align the source domain and target domain without fully mining of SD information, or perform the domain adaptation from semantic and structure aspects with simply characterization method which is sensitive to noise, resulting in the negative transfer and performance decline. To address these issues, in this paper, we propose a novel Dynamic Semantic-Geometric Guidance and Structure Transfer (DSGG-ST) network for cross-scene hyperspectral image classification task. The main aspects of DSGG-ST are twofold. On the one hand, the dynamic semantic-geometric guidance (DSGG) module is designed which consists of the semantic guidance component and geometric guidance component. The proposed DSGG module can align source and target domains under the dynamical guidance of the domain-invariance learning from the semantic and geometric perspectives. On the other hand, the graph attention learning-matching (GALM) module is developed for effectively transferring the structure information between the source domain and target domain. In this module, the graph attention network is adopted to encode the underlying complex structures, and the SeedGNN is exploited for efficient graph matching and alignment. Extensive experiments on three commonly used cross-scene HSI datasets demonstrate that the proposed DSGG-ST obtains a new SOTA performance on cross-scene HSIC, verifying the effectiveness of the proposed DSGG-ST.