
Accurate crop segmentation from multitemporal Sentinel-2 satellite image time series is important for agricultural monitoring and land-cover mapping. Although recent factorized spatial–temporal Transformers alleviate the quadratic cost of full self-attention, they still face challenges in preserving irregular parcel boundaries and handling noisy or invalid observations in long temporal sequences. To address these issues, we propose the sparse temporal–spatial vision Transformer (STSViT), a Transformer framework tailored for multitemporal crop segmentation. STSViT combines time-aware token encoding for temporally irregular observations, classwise temporal aggregation for category-specific feature modeling, and a structured sparse attention mechanism that integrates local sliding-window interactions with global anchor tokens for efficient spatial refinement. Experiments on the MTLCC and PASTIS benchmarks show that STSViT achieves consistent improvements over strong CNN and Transformer-based baselines, while achieving better runtime efficiency. These results demonstrate the effectiveness of STSViT for high-resolution crop mapping from multitemporal Sentinel-2 imagery.
Pansharpening aims to obtain multispectral images with high spatial resolution (HRMS) by fusing high-spatial-resolution panchromatic (PAN) images with low-spatial-resolution multispectral (LRMS) images. Most deep learning-based methods directly extract all band features of LRMS images and fuse them with PAN image features to reconstruct HRMS, which leads to inaccurate feature reconstruction for each band. To address this issue, we propose a progressive group-feature transmission fusion network (PGTFusion) with a multi-scale and multi-branch structure for pansharpening to achieve the extraction and fusion of different group features at different scales. At the minimum scale layer, considering the correlation between adjacent band features in LRMS images, a group-feature extraction transmission block (GETB) is designed to extract and fuse features from adjacent bands to obtain the group features, which were then transmitted to adjacent group for feature supplementation. At the intermediate scale layer, a group-feature fusion transmission block (GFTB) is constructed to achieve the fusion of group features and PAN images features, and the fused features are transferred to adjacent group to enhance the group features. At the maximum scale layer, a spatial information block (SIB) is utilized to achieve precise fusion of group features and PAN images features. Numerous experiments conducted on simulated and real datasets demonstrate that our method outperforms some state-of-the-art methods in both subjective and objective evaluation. The source code will be released on GitHub.
In understory terrain study, high-precision photon point cloud classification is essential for comprehensive analysis of forest ecosystems. However, ICESat-2 Advanced Topographic Laser Altimeter System (ATLAS) photon data contains substantial noise photons and inherently lacks terrain-adaptive properties, leading to frequent misclassification of noise photons as ground signals and thus severely reducing terrain extraction accuracy. To overcome this challenge, we propose a Density-Based Clustering and Filtering (DBCF) framework that integrates density-stratified adaptive denoising (DSAD) with geometry-driven cloth-surface classification, enabling robust classification of ICESat-2 photon point clouds into ground, canopy, and noise categories without using airborne terrain products as classification references. Across four contrasting forest regions—South Carolina, Maryland, Massachusetts, and Georgia—the proposed framework consistently outperformed NASA's official algorithm and conventional DBSCAN. With airborne G-LiHT DTM/DSM reference data, DBCF achieved an average recall of 0.66, precision of 0.72, and $F_{1}$-score of 0.69; with manual reference, it achieved an average recall of 0.79, precision of 0.78, and $F_{1}$-score of 0.78. Notably, the extracted ground photon data points achieved an average R$^{2}$ of approximately 0.99 and an average RMSE of 2.21 m, with RMSE values ranging from 1.92 to 2.40 m in understory terrain estimation, substantially outperforming forest and buildings removed copernicus digital elevation model (FABDEM) across all study areas. Furthermore, slope-stratified evaluation showed that the $F_{1}$-score decreased from 0.61 on flat terrain to 0.57 on moderate slopes and 0.47 on steep slopes, indicating that increasing terrain slope remains an important constraint on ground-photon classification. These results demonstrate the framework's adaptability across diverse forest and terrain conditions, offering a reliable basis for high-precision forest parameter retrieval and ecological assessment.
In recent years, the rapid and accurate extraction of photovoltaic (PV) footprints from Remote Sensing Images (RSIs) has become a prominent research focus, particularly with Deep Learning (DL). However, as RSI resolution improves and the demand for finer extraction increases, current DL-based algorithms face substantial challenges. High-precision, fully supervised PV segmentation relies heavily on massive, high-quality object-level annotated data, which is highly time-consuming and labor-intensive to obtain. In contrast, contour-level annotations are simpler and more accessible, making them ideal for larger-scale model training. Therefore, we propose a novel two-stage weakly supervised PV segmentation framework, termed PV Progressive SAM (PVPSAM), based on the Segment Anything Model (SAM). In the first stage, we fine-tune a SAM-based Contour Decoder (C-Dec) using abundant contour-level annotated samples to learn the PV segmentation task. In the second stage, we use the segmentation outputs as dense prompts to iteratively fine-tune a Refine Decoder (R-Dec) with a small amount of object-level data, progressively achieving high-precision results. PVPSAM achieves state-of-the-art performance on our newly proposed multi-scene dataset, GF-PVCOWS. Additionally, we verify its value for optimizing traditional datasets: through our developed PVRefineTool, original contour-level annotations can be converted into refined object-level data. Our work provides a new paradigm for large-scale RSI-based PV extraction.
Global navigation satellite system (GNSS) radio occultation (RO) water vapor pressure retrieval commonly adopts one-dimensional variational assimilation. This method relies on numerical weather prediction (NWP) background fields and can be limited by degraded background quality, missing background data, or the need for near-real-time processing. Data-driven retrieval methods provide a possible complementary approach that avoids the use of event-specific NWP background profiles during inference. However, existing studies mostly focus on conventional machine learning models, and the modeling of the vertical structure of refractivity and environmental contextual information remains insufficient. To this end, this paper proposes a Decoupled Cross-Attention Network (DCA-Net), which directly retrieves water vapor pressure profiles using refractivity profiles and a limited number of environmental context variables as inputs. DCA-Net first decomposes the refractivity profile into a trend component representing large-scale variations and a disturbance component representing local detailed variations through a high-low frequency decomposition module. Subsequently, differentiated convolutional encoding branches are used to extract the vertical structural features of the two components, and an environment-guided cross-attention mechanism is employed to model the adaptive interaction among trend features, disturbance features, and environmental context features, thereby reconstructing the water vapor pressure profile. Experiments based on COSMIC-2, ERA5, and IGRA RAOB show that DCA-Net achieves competitive retrieval performance among multiple data-driven methods, providing a feasible complementary scheme for GNSS RO water vapor pressure profile retrieval when real-time high-quality background fields are unavailable or degraded.
Long-term deformation monitoring and high precision reservoir parameter inversion are essential for mitigating geohazards and ensuring production safety in oilfield regions. This study utilizes Sentinel-1A imagery from 2017 to 2025 and employs the small baseline subset interferometric synthetic aperture radar (SBAS-InSAR) technology to reveal the spatiotemporal evolution of land subsidence in the Shuguang Oilfield. The results identify a maximum cumulative deformation of 120 cm and a peak subsidence rate of -12 cm/yr, which are rigorously validated by high-precision ground monitoring data. By integrating unmanned aerial vehicle digital orthophoto maps (UAV-DOM), the spatial damage effects of differential subsidence on critical infrastructure are accurately quantified. To overcome the geometric blind spots and non-uniqueness inherent in single line of sight InSAR inversions for anisotropic reservoirs, a “Model-Ground” joint inversion framework is proposed by incorporating high-precision ground observations. This framework utilizes three-dimensional surface displacements as hard constraints in Bayesian inference, significantly narrowing the prior search space for reservoir parameters. Quantitative evaluation demonstrates that this joint constraint reduces the residual variance of the Sill model by approximately 65% and achieves an uncertainty reduction rate (URR) of up to 59% for deep dynamic parameters. The established multi-source synergistic monitoring and high-fidelity inversion paradigm effectively decouples the bias between subsurface stress fields and geometric parameters, providing robust scientific support for geohazard early warning and the sustainable development of complex oil and gas reservoirs with highly compressible overburdens.
The early detection and classification of oil spills in the ocean is crucial to mitigate potential harm to the environment and economy. Optical hyperspectral remote sensing is an effective method for monitoring and tracking these events, primarily because of the high spectral precision and the possibility of multiscale spatial coverage. Despite the great potential of these sensors, the detection of ocean oil spills through optical images, especially in coastal environments, still requires human interpretation, delaying response efforts. Accurately measuring oil spill thickness is crucial for understanding the incident and its impacts. However, it remains a challenge due to limited studies, scarce data, and unreliable measurement techniques. A promising approach to automating hyperspectral imaging processing is utilizing Deep Learning models. This research evaluates the effectiveness of 3D Convolutional Neural Networks in detecting and classifying oil spill thickness through hyperspectral imaging in the visible/near-infrared and short-wave infrared regions. The data used to train the model were acquired through a controlled field experiment in which oil spills were simulated under conditions similar to marine coastal environments. Savitzky- Golay filters and continuous wavelet signal decomposition were applied in data preprocessing before the application of a 3D Convolutional Neural Network. The combination of continuous wavelet analysis and 3D Convolutional Neural Networks proved to be a promising approach to identify and estimate oil spill thickness on hyperspectral data, achieving an overall accuracy of 97.85%, a Kappa coefficient of 0.96, and a weighted F1 score of 0.98. This provides a scalable and accurate approach for early response to oil spill monitoring in coastal environments.
Hyperspectral image (HSI) lifelong classification aims to continuously recognize newly arriving classes while preserving previously acquired knowledge without full model retraining. Although exemplar replay is effective, retaining historical examples is often impractical in real-world applications due to privacy concerns and limited storage resources. Existing exemplar-free methods often suffer from representation drift in feature space and boundary confusion between previous and new classes. To address this issue, we propose an exemplar-free lifelong framework named DC-DMA for HSI classification based on Dual-Constrained Feature Extraction (DCFE) and Dynamic Margin Adaptation (DMA). Specifically, DCFE is designed to stabilize spectral-spatial representation learning. It imposes an embedding consistency constraint to preserve low-level spectral-spatial patterns. DCFE also restricts deep parameter updates via the Deep Orthogonal Gradient Constraint (DOGC) to protect historical manifolds. Furthermore, to mitigate decision boundary ambiguity, we introduce a classifier with DMA which maintains online class prototypes to estimate the semantic similarity between previous and new classes and accordingly adjusts class-specific cosine margins, thereby refining decision boundaries. Experiments on Indian Pines, Houston2013, Houston2018, and Hyrank-Loukia show that DC-DMA achieves competitive early-task performance and superior final-task performance against representative continual learning and hyperspectral incremental learning baselines. Our code is available at https://github.com/jsjcn/Jstars.
Wetland mapping is a fundamental prerequisite for ecological conservation and management, yet large-scale applications are severely hindered by the difficulty of acquiring high-quality and low-noise training samples. Manual interpretation is costly, inefficient, and prone to uncertainty, which substantially limits the generalization capability of intelligent mapping models. To address these challenges, this study proposes a highly automated and reliable framework for training sample generation and optimization in large-scale wetland mapping. Ecological index thresholding combined with Mahalanobis distance-based outlier detection is then employed to ensure initial sample reliability. Furthermore, a novel Sample Selection and Relabeling strategy based on a local Gaussian-Weighted Reverse Knearest Neighbor algorithm is introduced to robustly identify and correct noisy and confused samples by jointly leveraging high-dimensional spectral features and local spatial neighborhood information. Cross-regional generalization experiments conducted across seven ecologically heterogeneous regions demonstrate that the proposed framework yields remarkable performance improvements against initial labels produced by random forest, with the overall accuracy enhanced by 3%–6% and the Kappa coefficient elevated by 4%–10% across all test regions. The proposed approach provides an efficient, robust, and highly generalizable solution for intelligent wetland mapping across diverse regional ecological settings.
Solid waste detection in remote sensing images remains challenging because waste regions often exhibit heterogeneous appearances, irregular boundaries, substantial scale variations, weak visual saliency, and high similarity to surrounding land-cover types. To address these issues, this article proposes CDMNet, a solid waste detector integrating hierarchical frequency modeling, dynamic cross-scale fusion, and magnitude-aware refinement. In the backbone, resolution-specific complex spectral responses are learned for different channels and frequency positions to strengthen global contextual representation and attenuate redundant background responses. In the neck, a two-stage fusion strategy centered on the intermediate feature level promotes repeated interaction between spatial details and semantic information across scales. Magnitude-aware attention is further introduced at the deepest semantic stage to stabilize semantic responses associated with weakly salient and boundary-ambiguous waste regions. Experiments on two solid waste datasets show that CDMNet achieves mAP@50:95 of 56.4% and 37.4%. With 12.0 million parameters, the proposed model achieves a favorable balance between detection accuracy and model complexity. These results demonstrate the effectiveness of CDMNet for solid waste detection in complex remote sensing scenes.
Classical polarimetric synthetic aperture radar (PolSAR) ship detectors can be unified into a quadratic form of polarimetric scattering vectors, offering clear physical interpretability. To further exploit spatial information, the neighborhood polarimetric covariance matrix (NPCM) was introduced to characterize inter-pixel polarimetric coherences. However, NPCM-based approaches fundamentally rely on single-look complex data and suffer from high computational complexity, limiting their applicability in large-scale multi-look scenarios. Conversely, convolutional neural networks (CNNs) have become a dominant paradigm but are often treated as “black boxes” operating in Euclidean space. A fundamental theoretical gap remains: for the pre-activation, single-layer linear convolutional representation analyzed in this paper, the corresponding function class is a strict subset of the NPCM function class. This result shows that such linear convolutional aggregation cannot represent the off-diagonal inter-pixel correlation terms captured by NPCM, especially when phase information is absent. To bridge this gap, a Manifold-Aware Neighborhood Feature Learning (MA-NFL) framework is proposed. Instead of treating polarimetric matrices as flat vectors in Euclidean space, the Riemannian geometric structure of polarimetric covariance matrices is explicitly embedded into the network representation. By constructing a neighborhood quadratic embedding, non-linear inter-pixel polarimetric interactions are effectively learned from incoherent multi-look data, acting as a proxy for the missing phase information. Experiments on both simulated and measured datasets demonstrate that the proposed method outperforms standard linear convolutional representations.
Adaptive subtraction is a critical step in seismic data processing that attenuates multiple reflections by adaptively subtracting modeled multiples from the original seismic records. While U-Net-based approaches outperform traditional linear regression (LR) methods in multiple reflections suppression, they only capture time-domain features and disregard frequency-domain characteristics. Moreover, they ignore the morphological similarity (MS), defined by the dip and strike consistency of seismic events, between modeled and estimated multiples. Therefore, the U-Net-based methods cannot optimize the multiple reflections attenuation while preserving primary reflections and result in residual multiple reflections. To address these limitations, we introduce the Multi-Wavelet Convolutional Neural Network (MWCNN) and integrate a Gabor kernel-based MS constraint into the loss function of the proposed GMS-MWCNN approach. Unlike U-Net, MWCNN substitutes pooling operation with Discrete Wavelet Transform (DWT) and up-sampling convolution operation with Inverse Wavelet Transform (IWT), enabling simultaneous time-frequency feature extraction while preserving data details via DWT's superior time-frequency localization. Moreover, the Gabor kernels extract dip information to enforce MS consistency between modeled and estimated multiple reflections. By minimizing the loss function, the MS value is improved, enabling more effective attenuation of seismic multiple reflections. Synthetic data experiments show that GMS-MWCNN outperforms U-Net across multiple quantitative metrics, including but not limited to SNR and RMSE, and field data tests confirm its notable superiority in eliminating residual multiples.
Ship detection in synthetic aperture radar (SAR) imagery remains challenging in complex maritime scenes. Weak and locally discontinuous ship-related intensity responses may be progressively degraded during multiscale feature transformation. Meanwhile, sea clutter, coastlines, harbor facilities, and other strong background structures may introduce confusing responses during cross-scale feature fusion. Existing detectors commonly inherit fixed downsampling, interpolation-based upsampling, and direct feature concatenation from general object detection frameworks. These generic operations may inadequately preserve contour-sensitive ship structures and may propagate clutter-contaminated features. To address these issues, this paper proposes SAR-RAFNet, a SAR imaging-aware resampling and fusion network for ship detection in complex maritime scenes. First, an Edge-Guided Adaptive Downsampling module (EGADS) uses neighborhood context and Sobel-derived directional edge priors to guide adaptive four-sub-cell aggregation. Contour residual compensation further preserves weak, elongated, and contour-sensitive ship-related intensity structures during resolution reduction. Second, the established CARAFE operator is directly adopted in the top-down pathway to improve the spatial coherence of semantic feature reconstruction. Finally, a Scattering-aware Feature Aggregation Module (SFAM) explicitly models pairwise discrepancy and consistency between aligned feature streams. It performs branch-wise spatial-channel competitive selection with complementary interaction residual compensation, thereby enhancing mutually supported ship-related responses and suppressing branch-specific clutter. Experiments on SSDD, HRSID, and LS-SSDD-v1.0 demonstrate that SAR-RAFNet improves $\text{mAP}_{50:95}$ over the YOLOv26 baseline by 1.71, 2.13, and 2.17 percentage points, respectively. Ablation studies and visual analyses further verify the complementary contributions of EGADS and SFAM, particularly for small, weak-response, densely distributed, and nearshore ships. These results demonstrate that redesigning multiscale resampling and cross-scale interaction according to SAR intensity-domain characteristics can consistently improve ship detection across diverse maritime scenes.
Deep learning (DL)-based synthetic aperture radar (SAR) target recognition depends on measured data that are costly to collect, and synthetic images produced by electromagnetic simulation offer an economical substitute whose direct use is limited by the synthetic-to-measured (S2M) domain gap. Owing to the coherent imaging mechanism of SAR, this gap manifests not only in spatial appearance but also in the distribution of image energy across spatial-frequency bands and orientations, a structure that latent representations learned solely for reconstruction encode only implicitly and therefore cannot align selectively. To reduce this deep-feature gap, we propose LFCS2M, a latent diffusion framework for S2M SAR image translation in a frequency-coupled latent space. A frequency domain feature refinement module (FDFRM) imposes a learnable two-dimensional spectral mask on the latent representation, allowing bandwise and orientationwise feature adjustment without assuming a predefined spectral discrepancy. A bidirectional diffusion bridge guided by Measured Information Guided Cross Attention (MIGCA) further maps synthetic latent features toward the measured distribution. Qualitative and quantitative experiments show that LFCS2M generates high-fidelity SAR images that closely resemble measured data, supporting DL-based SAR target recognition by narrowing the domain gap and reducing data acquisition cost.
Small moving target detection aims to exploit spatio-temporal information to discriminate targets from background clutter. However, the extremely small target size, weak signal intensity, and highly dynamic backgrounds make accurate target extraction particularly challenging. Under highly dynamic backgrounds and weak target conditions, conventional spatio-temporal features may be insufficiently aligned with target information. Although some existing methods incorporate temporal cues at relatively early stages, many architectures still rely on frame-wise spatial encoding, separate spatial-temporal processing, or stage-wise temporal fusion. Consequently, fine-grained temporal variations associated with extremely dim moving targets may be attenuated before sufficiently coupled spatio-temporal representations are established, making them difficult to fully recover in subsequent feature propagation. To address these challenges, we propose a Refined Spatio-Temporal Pipeline Representation Network (RPRNet) for infrared dim and small moving target detection. Specifically, we first design a spatio-temporal pipeline representation block (STRB), which is built upon spatio-temporal deformable 3D convolutions to capture complex spatio-temporal pipeline features of targets. We then introduce We then introduce infrared-aware aware geometric sampling modulation (IR-GSM), which enhances the continuity and consistency of spatio-temporal pipeline representations by adaptively adjusting high-resolution sampling points for small targets during upsampling. Finally, we propose structure-progressive guided loss (SPG Loss), which incorporates a progressive edge expansion guidance mechanism to alleviate class imbalance and provide more stable supervision for dim targets, thereby preserving the integrity of spatio-temporal pipeline features. Experimental results on the NUDT-MIRSDT and TSIRMT datasets demonstrate that our network outperforms state-of-the-art (SOTA) methods.
Remote sensing image interpretation is shifting from closed-set visual recognition toward open-ended, language-driven understanding. Remote sensing vision-language models (RS-VLMs) connect Earth observation imagery with natural language, enabling more flexible recognition, retrieval, generation, and reasoning beyond predefined categories. This review first introduces the foundations of visual representation and cross-modal alignment, and then organizes RS-VLMs into contrastive vision-language pretraining, visual instruction tuning, and diffusion-based multimodal generation. It systematically examines representative model architectures and their capabilities across zero-shot scene classification, image captioning, image-text retrieval, visual question answering, open-vocabulary detection and segmentation, and visual grounding. It further discusses pretraining, instruction-tuning, and reasoning-oriented datasets, and analyzes how RS-VLMs can serve as perceptual and semantic-interface modules in remote sensing agents. Finally, the review identifies future directions in physically grounded reasoning, vertical-domain and lightweight deployment, and unified representation across optical, synthetic aperture radar (SAR), hyperspectral, and multitemporal data.
Building change detection in remote sensing plays an important role in urban planning, land-use monitoring, and rapid post-disaster damage assessment. However, remote sensing images are inherently affected by illumination variations and seasonal changes, which introduce complex pseudo-change noise. In addition, traditional hierarchical network architectures struggle to effectively balance multi-scale structural boundaries and global texture consistency. To address these challenges, this paper proposes a Memory-Sustained Structure-Texture Representation Network (MSST-Net). The network consists of three key components. Hierarchical Group-Grid Coordination Module (HGGCM) is designed within the hierarchical backbone to suppress shallow feature noise, promote cross-scale collaboration among hierarchical features, and align spatial discrepancies. Bitemporal Memory Augmented Module (BMAM) is introduced, which integrates learnable memory matrices into the cross-attention mechanism to establish stable semantic correspondence between bi-temporal images. Structure-Texture Fusion Module (STFM) is further designed based on a structure-texture disentanglement strategy to separately refine the macroscopic structural information and fine-grained texture details of buildings, thereby accomplishing the change detection task. MSST-Net achieves F1 of 92.08%, 94.84%, and 89.39% on three public datasets LEVIR-CD, WHU CD, and GZ-CD, respectively, demonstrating the effectiveness of the proposed network and the rationality of its architectural design.
Semantic segmentation in remote sensing is essential for environmental perception and Earth observation. However, developing segmentation networks relies on large-scale datasets with pixel-wise annotations. When integrating optical (RGB) and thermal data, models face data scarcity and in formation asymmetry between sensors, yet traditional multi modal networks typically fuse features through symmetric spatial matching. To overcome these limitations, this paper proposes the Foundation model-driven Multimodal Asymmetric Gated Network (FMAGNet). Utilizing a teacher-student architecture, this framework mitigates label scarcity by distilling semantic priors from a frozen Segment Anything Model (SAM) encoder, preserving pre-trained knowledge while reducing computational burden. To resolve information asymmetry, we introduce an Asymmetric Feature Mimicry (AFM) module, coupling shallow optical features with deep thermal semantics to bridge the semantic gap before feature fusion. Furthermore, we integrate a Strip Pooling Module (SPM) into the student encoder to establish anisotropic context modeling and correct geometric mismatches. An Uncertainty-Aware Gated Fusion (UAGF) mechanism is designed to filter modal noise caused by environmental degradation. We conducted our experiment on the Caltech Aerial RGB Thermal (CART) dataset and verify the model's generalization capability on the ISPRS Vaihingen dataset. FMAGNet achieves an mIoU of 74.9% on CART dataset and a OA of 91.93% on Vaihingen dataset, confirming the effectiveness of the proposed architecture.
In recent years, CNN-based methods have been extensively applied to hyperspectral image (HSI) and light detection and ranging (LiDAR) joint classification. However, most existing methods rely on fixed sliding windows for local feature extraction, making them susceptible to heterogeneous pixel interference, particularly around object boundaries and detail-rich regions. Moreover, their performance often degrades due to the limited availability of labeled samples. To address these limitations, this paper proposes a joint classification network via superpixel-guided prototypical feature (SPFNet). Specifically, a collaborative superpixel segmentation strategy is first employed by integrating the spectral-spatial characteristics of HSI with the geometric information of LiDAR, enabling the extraction of homogeneous regions. Based on these regions, an adaptive patch restructuring strategy is introduced to generate regular-shape patches. Subsequently, these regular-shape patches are fed into a dual-branch multi-scale convolutional module for hierarchical feature extraction and cross-modal fusion. Finally, class prototypes are constructed in the embedding space, and the Mahalanobis distance is adopted as the similarity metric for classification. Experimental results on three benchmark datasets demonstrate the propose SPFNet achieves competitive classification performance, particularly in complex scenes and few-shot scenarios.