
Remote sensing scene classification is a fundamental task in scene understanding. Several existing methods generate local images containing discriminative regions to assist in scene classification. However, the generation of such local images often neglects the surrounding scene context. In this letter, we propose a local region magnification strategy (LRMS) to enlarge discriminative regions while preserving scene context, thereby producing context-aware local images. Subsequently, a dual-branch distillation framework (DBDF) is developed to integrate features from both original and local images for scene classification. In addition, an online distillation strategy is adopted, where the ensemble logits from the internal components of DBDF are used to generate supervisory signals that guide the optimization of the overall framework. Experimental results on four benchmark datasets demonstrate the effectiveness of the proposed DBDF. Our code is available at: https://github.com/plkfans/DBDF.
Object detection in aerial imagery has garnered significant attention due to its crucial role in applications such as urban planning, environmental monitoring, and disaster response. However, class imbalance remains a persistent challenge, as minority categories are represented by fewer instances than dominant classes in most remote-sensing datasets. To address this issue, we propose a novel context-aware patch-based data augmentation method that mitigates class imbalance by selectively choosing donor images and semantically compatible object instances. The proposed method features two key innovations. First, it preserves contextual coherence by leveraging class-specific donor collections and enforcing geometric and semantic constraints (e.g., boundary, shape, and class). Second, it introduces class-balancing threshold computed from dataset statistics to dynamically regulate augmentation copy-paste probability for minority categories. We evaluate the approach on DOTA-v1.0 demonstrating a significant improvement in mAP@0.5, and mAP@0.5-0.95 for minority categories while preserving competitive performance for majority classes.
The array-shaped microwave calibration target (MCT) is widely applied in the radiometer payloads, with the advantages of compact size and Lambertian thermal radiation. However, it has been noticed that the vertical temperature gradient at unit tips introduce notable brightness temperature (BT) bias, obstructing the achievement to high-accuracy calibration. Meanwhile, the direct testing on the temperature of the coating layer remains as a difficult task, especially when the MCT is mounted on the space-borne payloads, so the BT of the MCT is hard to be accurately determined. In this work, the authors explore the possibility to estimate the brightness temperature based on the widely applied PRT (platinum resistance thermometer) in the unit kernel. It is shown in the thermal measurement of a uniform background that, the PRT tested temperature at different height of metal kernel can be distinguishable, and the temperature difference variation agrees with the trend of target-ambient temperature difference. While the PRT has been widely applied to test the aperture temperature homogeneity of the array-shaped MCT, the results in this work indicate that the PRT has the potential to practically detect the vertical temperature gradient in units which is directly related to the BT bias. Further, the possible manners and main obstruction factors in projecting the PRT tested temperature difference to the BT bias estimation, are discussed in this work.
Two-dimensional phase unwrapping is a key step in interferometric synthetic aperture radar and sonar processing. Existing two-step deep learning methods usually rely on nondifferentiable back-end physical solvers, which interrupt end-to-end optimization and introduce a mismatch between training and inference. This letter proposes CorrectionNet, an end-to-end unwrapping framework based on edge-level flow classification. The wrapped phase is reformulated as horizontal and vertical phase-jump prediction on image boundaries, and a continuous expected relaxation strategy is introduced to replace hard discrete decoding. Combined with a parameter-free discrete cosine transform Poisson solver, the proposed framework establishes a fully differentiable path from jump estimation to phase reconstruction. On the simulated InSAR-DLPU test set, CorrectionNet achieves an RMSE of 0.9861 rad and an SSIM of 0.7871, outperforming RPNet, FPUNet, PhaseNet 2.0, and SNAPHU in both reconstruction accuracy and structural fidelity. Zero-shot tests on 512×512 and 1024×1024 repeat-pass Sentinel-1 interferograms processed by LiCSAR further indicate scale transferability and re-wrapping consistency under low-coherence and decorrelated conditions. For 256×256 inputs, the average inference time is about 13 milliseconds per image, indicating a favorable tradeoff between accuracy and efficiency.
Semantic segmentation of high-resolution remote sensing imagery is computationally demanding, particularly for multimodal fusion of RGB and normalized digital surface model (nDSM) data. Existing multimodal networks improve segmentation accuracy but often introduce substantial computational overhead. This letter presents LiEAF-Net, a lightweight elevation-aware fusion network with only 6.50M parameters. The proposed elevation-aware double-attention selective kernel (EADASK) module enables effective cross-modal feature interaction through SE-based calibration, modality-specific multi-scale feature extraction, and dual spatial–channel attention. Experiments on the ISPRS Vaihingen and Potsdam datasets achieve 83.59% and 86.45% mIoU, respectively, with 4.8–35.7× fewer parameters than state-of-the-art multimodal methods, demonstrating the efficiency of LiEAF-Net.
The Ice, Cloud, and Land Elevation Satellite-2 (ICESat-2) photon-counting lidar provides new opportunities for the large-scale forest aboveground biomass (AGB) estimation. Currently, most AGB estimation models rely on the relative height (RH) metrics that lack of the canopy stem and crown information, limiting the AGB estimation accuracy. This study incorporates a Biomass Index (BI) derived from the gap probability and an allometric formulation to represent the vertical canopy structure. A composite AGB regression model is proposed by integrating the linear RH metrics and power-law BI metric. Using ICESat-2 ATL08 data and the reference AGB maps in the northern Rocky Mountains, the relationships of the AGB with the BI and RH metrics are investigated. Results show that BI exhibits a stronger correlation with the AGB than the RH100, RH75, and RH50. The performance of the proposed nonlinear AGB model is evaluated by comparing with the RH-based linear regression and random forest models based on the coefficient of determination (R2) and the root mean square error (RMSE). The proposed AGB model can achieve the best performance with the R2 of 0.78 and the RMSE of 34.14 Mg/ha. The results demonstrate that the BI can provide a quantitative metric for representing the vertical canopy structure and may be a crucial factor in the estimation of the large-scale AGB using ICESat-2 photon data.
Accurate estimation of the quality factor (Q) using seismic waves remains a challenge due to spectral notching induced by thin-bed tuning and noise. These notches act as outliers in conventional logarithmic spectral ratio (LSR) methods, leading to biased attenuation estimates. We propose a robust LSR inversion framework informed by physical priors and residual statistics. By integrating dynamic Cauchy weights within an iteratively reweighted least squares (IRLS) scheme, we mitigate the impact of interference-related outliers. Simultaneously, a second-order Tikhonov constraint encourages the reconstructed spectral ratio to adhere to the Futterman model. Synthetic tests with known Q values demonstrate improved accuracy and robustness, while the field-data example illustrates applicability and stability on real seismic data.
Least-squares migration (LSM) improves seismic imaging resolution by compensating for Hessian-induced blurring, but the explicit construction and inversion of the full Hessian remain computationally prohibitive. Image-domain least-squares migration (ID-LSM) commonly uses point-spread functions (PSFs) as local representations of the Hessian. However, conventional PSF deconvolution is sensitive to spatial nonstationarity and ill-posedness, and often suffers from a tradeoff between resolution enhancement and artifact or noise amplification. To alleviate this problem, we propose an ID-LSM method based on a deep feature deconvolution network (DFDN). The migrated image is first mapped into a deep feature space, where explicit PSF deconvolution is performed in the feature domain, and the compensated features are then fused through a reconstruction network to recover the final image. In addition, a physics-consistency constraint derived from the PSF degradation model is introduced to improve both stability and physical plausibility. Numerical experiments demonstrate that, compared with reverse time migration (RTM) and standard ID-LSM, the proposed method produces higher-quality imaging results.
Referring remote sensing image segmentation (RRSIS) commonly treats language as a fixed query that only modulates visual features. This open-loop design is brittle in aerial scenes containing repeated objects, weak appearance cues, and relational expressions: visual evidence cannot revise which words and relations should dominate the query. We present MMF-Net, a closed-loop fusion architecture that couples stage-wise reciprocal token refinement with decoder-wide context propagation. Its Vision–Language Mutual Feedback Fusion (VLMFF) module updates visual and linguistic tokens through symmetric cross-attention and gated residual fusion, while Multimodal Context-Aware Fusion (MCAF) consolidates the co-refined state and broadcasts it across the visual hierarchy. Under an identical Swin-B+BERT backbone, decoder, training schedule, and data split, MMF-Net improves a strong internal baseline by 2.87/2.51/1.48 points at Pr@0.5/0.7/0.9 and by 0.68 oIoU and 0.58 mIoU. A factorial ablation further exposes positive VLMFF–MCAF interaction effects of 5.54 points at Pr@0.7 and 2.40 mIoU. MMF-Net reaches 64.86% mIoU and 78.14% oIoU on RRSIS-D with 115.2M parameters and 68.7 GFLOPs.
This study compared machine‑learning, geostatistical, and hybrid methods for predicting soil organic matter, cation exchange capacity, and nitrogen using remotely sensed, topographic, and soil survey covariates to support targeted soil management at the field scale. Using nested cross-validation, we found broadly similar predictive performance across methods, with XGBoost showing a slight advantage over kriging with external drift by reducing pooled standardized RMSE from 0.603 to 0.593, corresponding to a 1.7% relative reduction. Incorporating XGBoost into a regression‑kriging framework provided only marginal additional improvement, reducing pooled standardized RMSE to 0.583. These results suggest that hybridization does not necessarily provide substantial gains when the ML model and covariates already capture most predictable spatial variation. However, model diagnostics showed that geostatistical correction was most beneficial when residual spatial autocorrelation remained after the ML step. Sentinel‑derived surface and vegetation indices were the most consistent predictors, while topography played a secondary role, and soil survey variables contributed minimally. Overall, combining ML and geostatistical tools offers a more robust framework than relying on either alone, because their relative advantages vary with site and soil property.
Weakly supervised remote sensing semantic segmentation aims to achieve pixel-level prediction with limited annotation costs. Recently, CLIP-based methods have shown promising potential by leveraging vision-language alignment for semantic localization; however, they often suffer from ambiguous semantic representations and incomplete pseudo-labels in complex remote sensing scenes. To address these issues, we propose a spatial cue-guided framework for weakly supervised remote sensing semantic segmentation. Specifically, a confidence-aware prototype alignment module is designed to extract reliable semantic cues from high-confidence regions, enhance feature discrimination through contrastive learning, and improve object completeness by exploiting structural information. An adaptive pseudo-label completion strategy is developed to progressively improve pseudo-label coverage while reducing noise propagation. In addition, a structure-aware heterogeneous multi-head attention decoder is introduced to effectively fuse global semantic context, local spatial details, and cue information for refined prediction. Experimental results on two benchmark remote sensing datasets demonstrate the superiority of the proposed method.
Moving infrared small target detection remains challenging because weak targets are easily obscured by complex dynamic backgrounds. Existing methods mainly aggregate temporal features and represent directional evolution discretely, limiting the faithful characterization of continuous motion directions. To address this issue, we propose an Angle-Aware Temporal Motion Representation Network (ATMNet) for moving infrared small target detection. The proposed method introduces a continuous angle-aware motion representation paradigm by jointly learning motion saliency and directional motion dynamics from infrared image sequences. Specifically, a Temporal Motion Representation (TMR) module is designed to construct interpretable motion representations, capturing both motion occurrence and continuous directional evolution. Furthermore, a Motion-Conditioned Encoding-Decoding Strategy is introduced to inject motion representations into feature learning through feature modulation, enabling motion-guided representation learning. Extensive experiments on benchmark datasets demonstrate that the proposed ATMNet achieves superior performance compared with existing state-of-the-art methods, particularly in complex dynamic scenarios. The code is available at: https://github.com/Y-xiaoyang/ATMNet.
Three-dimensional (3D) point cloud change detection is an essential task in urban monitoring. Recent raw-point frameworks preserve native 3D geometric information more effectively than projection-based or voxelization-based pipelines. However, robust change detection in bi-temporal urban scenes remains challenging because irregular sampling, density variation, and temporal misalignment can weaken cross-temporal feature correspondence. To address this problem, we propose Adaptive Kernel Point Convolution X (AKPConvX), a point convolution operator that dynamically constructs geometry-conditioned kernels from a learnable weight bank and incorporates a dual adaptive mechanism to jointly model spatial structure and semantic context. AKPConvX replaces rigid kernel point parameterization with dynamic depthwise kernel assembly, allowing convolution weights to adapt to local point distributions during inference. Based on this operator, we develop a Siamese encoder-decoder network with Inverted Bottleneck blocks for effective multi-scale feature extraction. We further introduce a Position-Adaptive Difference (PAD) fusion module that reweights reference features according to spatial alignment confidence, thereby reducing false detections caused by temporal misregistration. Experiments are conducted on the two point-level sub-datasets of Urb3DCD-V2 defined by the Siamese KPConv benchmark: Urb3DCD-V2-1, corresponding to low-density LiDAR acquisitions, and Urb3DCD-V2-2, corresponding to the Multi-Sensor (MS) configuration. On Urb3DCD-V2-1, the proposed method achieves a mean Intersection over Union for change classes (mIoUch) of 80.24% and an overall accuracy of 96.26%, outperforming the Siamese KPConv baseline by 5.31% and 2.43%, respectively. On Urb3DCD-V2-2, AKPConvX also consistently improves over Siamese KPConv across mIoUch, mIoU, mAcc, and Acc.
Detecting dim infrared small targets in cluttered video sequences remains difficult because their weak spatial signatures can be easily confused with dynamic background responses. Although multiframe methods exploit temporal information, directly aggregating frame-wise features may introduce background fluctuations into target representations, whereas sophisticated alignment or global interaction mechanisms can incur considerable complexity. This paper proposes Motion-Outlier Calibration (MOCA), a lightweight plug-in module that converts multiframe features into an enhanced key-frame representation for a compatible single-frame infrared small target detector. MOCA constructs a multi-statistic background representation from temporal mean, local low-frequency context, and temporal variation to derive background-decoupled residual cues and a stable base correction. It further aggregates signed, average, and maximum interframe contrasts to describe target-related temporal anomalies, and uses local high-frequency priors to selectively inject complementary temporal corrections into the key-frame feature. Experiments on the IRDST and NUDT-MIRSDT datasets, including challenging low-SCR scenes, demonstrate that MOCA achieves competitive detection accuracy and false-alarm suppression with a compact plug-in design. The source code is available at https://github.com/AmazingJ-123/MIRSTD-MOCA.
High-resolution remote sensing image semantic segmentation is vital for ecological assessment, natural resource surveying, and other geospatial applications. Compared with CNN- and Transformer-based methods, superpixel-based graph neural networks (GNNs) have gained attention for modeling topological relationships among irregular geographic objects. However, many existing methods construct graph nodes from shallow features, limiting semantic discriminability and global contextual guidance. Moreover, most methods model topology only at a single or parallel multi-scale level, restricting the capture of geographic organizational relationships and the preservation of structural consistency across spatial granularities. To address these issues, this letter proposes a Hierarchical Topological Reasoning Network (HiTNet) for high-resolution remote sensing image segmentation. HiTNet employs a Hierarchical Local–Global Perception Encoder (HLGPE) to efficiently extract local details and global semantics. A Problem-aware Complementary Fusion Module (PCFM) adaptively integrates these features through problem-aware guidance, providing sufficiently rich representations for subsequent graph node construction. Then, the Hierarchical Topological Reasoning Module (HTRM) organizes superpixel-guided object nodes into an object–community–block hierarchical topology graph and performs region-level reasoning through edge-aware graph convolution and bidirectional gated cross-level propagation, explicitly modeling geographic organizational relationships across multiple granularities. Experiments on Vaihingen, Potsdam, and LoveDA datasets demonstrate that HiTNet consistently outperforms state-of-the-art methods.
Existing remote sensing image segmentation methods rely on predefined category labels, which are often insufficient to capture the complex spatial semantics inherent in geospatial concepts such as flood inundation zones, landslide bodies, and industrial complexes. This paper presents ConceptSeg, a text-guided multimodal segmentation framework that enables zero-shot segmentation of linguistically described geospatial objects in remote sensing imagery. ConceptSeg employs a decoupled architecture comprising a multi-scale image encoder with 2D Rotary Position Embedding (2D-RoPE), a unified multimodal prompt encoder with adaptive gated fusion, and a compact two-way cross-attention decoder, bridging natural-language concept descriptions with pixel-level predictions through Region of Interest (ROI)-guided feature extraction and contrastive alignment. With only ≈35M trainable parameters atop frozen encoders, it stays markedly lighter than recent multimodal large-language-model baselines. Experiments on multiple benchmarks show up to +0.09 mIoU improvement over prior cascaded reasoning baselines out-of-domain, a smaller margin over concurrent open-vocabulary methods that is significant only with an external box prior, and consistent gIoU/cIoU gains in-domain, supporting concept-driven segmentation for complex geospatial reasoning.
GeoHam is a geometric-regularisation framework inspired by Hamiltonian mechanics for remote-sensing change detection (CD). It replaces conventional feature subtraction with multi-scale Störmer-Verlet symplectic integration applied to learned feature dynamics. Unlike physics-informed neural networks that enforce PDE residuals from physical laws, GeoHam exploits the mathematical structure of Hamiltonian mechanics as an inductive bias on bitemporal feature trajectories: the symplectic 2-form is preserved at every leapfrog step, acting as a geometric regulariser that keeps distinct temporal states distinguishable. Three Feature Pyramid Network (FPN) levels (64×64, 32×32, 16×16) share a single integrator projected to a common 64-dimensional space; a raw absolute-difference skip at 128×128 preserves boundary detail. Integration steps per scale are selected by Bayesian hyperparameter optimisation. GeoHam achieves F1 of 91.20%, 95.89%, and 79.56% on LEVIR-CD, CDD, and CLCD respectively, with only 7.53M parameters and 4.72 GFLOPs, 5.3×fewer GFLOPs than StarCD-Net without any transformer, state-space model, or cross-attention block.
Remote sensing (RS) semantic change detection (SCD) confronts key challenges of temporal feature misalignment, fragmented narrow-elongated target prediction, and inconsistent semantic-change representations due to temporal variations and intricate geographic target properties. To address these intertwined challenges, this letter proposes STA-Net, a multi-task framework integrating three synergistic core modules: Cross-Temporal Interaction (CTI), Axial-Semantic Transformer (AST), and Spatio-Temporal Fusion (STF). The CTI module aligns cross-temporal features via channel subset swapping to mitigate misalignment, the AST module preserves narrow-elongated target structural integrity through axial attention for fine-grained semantic representation, and the STF module fuses spatio-temporal cues via 3D convolution to resolve semantic-change inconsistency, forming a closed-loop “alignment-representation-fusion” collaborative framework that addresses each bottleneck. On the SECOND and HRSCD datasets, STA-Net achieves state-of-the-art fine-grained SCD performance and superior robustness across multiple metrics (e.g., FSCD, Weighted Score). With a multi-task architecture for high-resolution RS scenes, it delivers high-precision detection with a favorable trade-off between computational efficiency and detection performance, demonstrating strong effectiveness for RS-SCD tasks. The source code is available at: https://github.com/WangXin81/STA-Net/.
Global particulate matter (PM) forecasting is critical for air quality management, yet regional variability in emission sources and atmospheric processes poses challenges for unified modeling approaches. We present a lightweight nested-domain deep learning framework using a Residual U-Net (ResUNet) architecture for short-range forecasting of PM1, PM2.5, and PM10. Our approach trains separate U-Net models with a shared encoder and one decoder per forecast lead time for each of 10 different spatial regions and 3 PM species, using overlapping 256×256 input grids to predict 192×192 forecast regions with explicit spatial context. Using CAMS analysis data spanning 2021–2024, we train independent U-Net models for each region/PM species combination, each with ≈535 K parameters and a total of ≈16 million parameters across the full 30-model ensemble. Evaluated against the 1.3 billion-parameter Aurora foundation model, our framework achieves substantially lower CRMSE, RMSE, and SEEPS scores at the 6-hour forecast horizon for all three PM species, while remaining competitive at longer lead times (12–24 h). These results demonstrate that lightweight, regionally-specialized models offer a viable alternative to large-scale foundation models for PM forecasting, providing several orders of magnitude reductions in parameter count while achieving comparable or superior short-range forecast skill.
High-resolution aerial and satellite imagery is vital for applications such as urban planning and environmental monitoring, yet its high acquisition cost limits accessibility. We present Boundary Attentioned Super-Resolution GAN (BASR-GAN), a deep learning framework that enhances low-cost, low-resolution RGB aerial imagery using Boundary Edge Loss and Total Variation Loss to preserve spatial structures and object boundaries. Beyond standard visual reconstruction metrics such as PSNR and SSIM, we conduct a comprehensive evaluation of the impact of super-resolution on downstream GeoAnalytics applications, specifically aerial object detection.We perform a large-scale benchmark of multiple super-resolution methods across four diverse datasets (SpaceNet2, DOTA-v2, VEDAI, HRSC) and three generations of lightweight object detection models (YOLOv8n, YOLOv11n, and YOLOv26n), analyzing detection performance for each dataset-detector configuration. The results show that BASR-GAN improves reconstruction quality and substantially reduces the object detection performance gap between super-resolved low-cost imagery and native high-resolution data. By limiting this gap to an average of approximately 14% across the evaluated dataset–detector configurations, compared to larger drops observed with traditional interpolation-based methods, BASR-GAN provides a practical and cost-effective solution for automated geospatial monitoring.