Hyperspectral video data provide complementary contextual cues across spectral, spatial, and temporal dimensions for modeling object dynamics under challenging conditions. Many existing hyperspectral video object tracking (HVOT) approaches organize spatial-spectral and temporal modeling in successive stages, leaving room for closer interaction among video-level contextual cues. To address this, we propose HucrTrack, a unified contextual reasoning framework for HVOT trained by parameter-efficient fine-tuning (PEFT). HucrTrack forms synchronized hyperspectral and false-color representations from each hyperspectral cube and enhances spatial-spectral features through a weight-shared dual-representation backbone with unified contextual cue modeling. To effectively leverage contextual dynamics, we design a unified contextual reasoning module (UCRM) composed of three key components: memory dynamics unit (MDU), contextual injection unit (CIU), and selective retrieval unit (SRU). Specifically, MDU maintains a frame-wise dynamic memory via Mamba's hidden states; CIU hierarchically integrates this memory into the spectral-spatial backbone features; and SRU selectively retrieves relevant contextual information to reinforce the tracking representation. In contrast to representative stepwise designs, HucrTrack enables concurrent, unified reasoning over all three dimensions within one recurrent process. Extensive experiments on ten benchmarks demonstrate that HucrTrack compares favorably with existing trackers in both robustness and generalization.
High-resolution remote sensing imagery presents unique challenges for efficient visual understanding, including dense object distributions, severe scale variations, strong background redundancy, and complex spatial structures. Existing deep models often rely on deep architectures or computationally intensive global modeling strategies, limiting their deployment on resource-constrained platforms. In this paper, we propose an efficient and lightweight backbone network, termed CFMNet (Cooperative Feature Modeling Network), for high-resolution remote sensing image understanding. CFMNet decomposes feature representations into heterogeneous yet complementary subspaces and models them cooperatively within a unified framework. Specifically, it coordinates channel semantics, structure-aware spatial dependencies, local detail enhancement, and global contextual consistency to improve representation efficiency while suppressing redundant computation. Extensive experiments demonstrate the effectiveness and generality of CFMNet. It achieves 96.13%, 95.50%, and 98.10% Top-1 accuracy on NWPU-RESISC45, AID, and UCM, respectively, 79.82% mAP on DOTA-v1.0, 73.02% mAP on DOTA-v1.5, and 90.82% mAP on HRSC2016, as well as 83.8% mIoU on Vaihingen and 53.8% mIoU on LoveDA, while maintaining low parameter count and computational complexity. A scaling-based Pareto analysis on DOTA-v1.0 and LoveDA further shows that CFMNet variants form a favorable efficiency–accuracy frontier compared with representative lightweight backbones. These results indicate that cooperative modeling of heterogeneous features provides an effective and efficient solution for high-resolution remote sensing image understanding. The code will be released at https://github.com/BEIBEIPRINCESS/CFMNet.
Unmanned Aerial Vehicle (UAV) multispectral video object tracking is critical for real-world applications. While multispectral imaging offers complementary spectral cues beyond the visible range, tracking in aerial scenarios remains challenging due to data scarcity, suboptimal spectral-spatial modulation, and discrete sequential temporal modeling. To this end, we propose CASS, a context-aware memory framework with spectral-spatial modulation, which integrates spectral, spatial, and temporal cues for UAV multispectral tracking. CASS introduces two lightweight modules: (i) the efficient spectral-spatial modulation (ESSM) module, which modulates spatial representations through spectral-guided fusion, and (ii) the context state space reasoning (CSSR) module, which leverages evolving state space memory to retain long-term temporal cues and mitigate error propagation during cross-frame reasoning. By integrating these components in a parameter-efficient fine-tuning fashion, CASS achieves both efficient modulation and context-aware tracking. Evaluations on UAV multispectral benchmark (MUST) and ground-based hyperspectral benchmarks (NIR, RedNIR, VIS, MSSOT, MSVT) demonstrate CASS’s superior performance for both general hyperspectral tracking and specific UAV perception.
We propose self-supervised rotation-invariant descriptors based on mixed rotation-equivariant convolutional neural networks (CNNs) (MRDes) for heterogeneous remote sensing image matching between visible (VIS) and near-infrared (NIR) images. Existing methods struggle with large rotation variations for VIS-NIR image matching, particularly when georeferenced information is unavailable or inaccurate, limiting their practical applicability. To address this, MRDes employs rotation-equivariant CNNs to extract equivariant features and construct robust rotation-invariant descriptors. Specifically, we design a mixed learning strategy that integrates explicit and implicit equivariance to optimize feature representations, while a contrastive loss enhances their discriminability by refining the distances between positive and negative samples. Experimental results on benchmark datasets demonstrate that MRDes significantly outperforms state-of-the-art methods, achieving a 70.9% improvement in matching success rate over XoFTR and exhibiting strong generalization to unseen UAV data.
Estimating motion and geometry from dynamic scenes is an open and underconstrained problem in computer vision. Constrained by the limited availability of co-labeled data, current end-to-end methods consider motion segmentation and geometry estimation as two independent tasks, with motion segmentation unilaterally facilitating geometric alignment in downstream applications. However, geometry estimation also offers valuable information for enhancing motion segmentation. This paper proposes a mutually beneficial framework for zero-shot motion segmentation and consistent geometric reconstruction. Specifically, 3D spatio-temporal priors are extracted from a geometry-first 4D reconstruction model to guide generalizable motion segmentation. A Dual-Dimension Multi-path Information Fusion (D2MIF) module is designed to fuse complementary 3D and 2D information at multiple scales through a recursive refinement mechanism, thereby improving zero-shot segmentation in complex dynamic scenes with background distraction and object articulations. Subsequently, the refined motion segmentation mask is utilized to more accurately separate the dynamic foreground from the static background during alignment, thus improving the geometric consistency of the 4D reconstruction. Experimental results validate the proposed framework’s mutual benefits and efficiency in downstream 4D reconstruction tasks; the motion segmentation model exhibits competitiveness and generalizability. The project’s webpage is available at https://mube4d.github.io/.
Remote sensing images acquired by unmanned aerial vehicles (UAVs) and satellites are often degraded by adverse weather, illumination variation, and imaging artifacts, which may co-occur and jointly induce global distribution shifts and local structural corruption. Although All-in-One image restoration offers an appealing unified alternative to task-specific pipelines, existing methods still suffer from weak or implicit degradation cues and parameter redundancy caused by full-rank multi-expert designs with overlapping restoration behaviors. We propose CoRE-UIR (Common and Residual Experts for Universal Image Restoration), a prior-guided global–local framework centered on the Common-and-Residual Expert Block (CoRE). CoRE explicitly decomposes restoration capacity into a common dense expert for degradation-invariant restoration and low-rank residual experts for degradation-specific compensation, enabling adaptive specialization without redundant expert replication. Built on this design, Degradation Prior Embedding (DPE) adapts frozen CLIP features into an explicit restoration-oriented prior, while Global Feature Modulation (GFM) aligns global feature statistics before local residual compensation. We also construct MDVD-108K (Multi-Degradation VisDrone), a large-scale UAV restoration dataset covering both single and compound degradations, together with a real-world test set. Extensive experiments on multiple datasets show that CoRE-UIR improves the overall average PSNR by 1.05 dB while running 11.83× faster and reducing peak memory by 85.3% relative to the strongest baseline, BaryIR, thereby maintaining a favorable quality-efficiency trade-off.Evaluations on downstream tasks and unseen degradation also validate the generalizability of CoRE-UIR. The code and dataset will be released at https://github.com/zzaiyan/CoRE-UIR.
Remote sensing imagery is essential for global environmental monitoring, but frequent cloud cover severely limits the utility of optical images. Fusing cloud-prone optical images with cloud-penetrating Synthetic Aperture Radar (SAR) data offers a path to all-weather Earth observation. However, this task faces a dual challenge: the escalating computational cost of state-of-the-art methods and the inherent ill-posedness of the reconstruction under information loss, which complicates the learning process. To tackle this, we propose ECRformer (Efficient Cloud Removal Transformer). ECRformer pairs an efficient architecture with a principled learning paradigm to address both challenges through: (1) a suite of efficient attention mechanisms, including Cross-Covariance Attention (XCA) for computationally-aware multimodal feature fusion and Multi-Dilation Window Attention (MDWA) for capturing multi-scale spatial context with linear complexity; and (2) the Semantic-Decoupled Feature Learning (SDFL) paradigm, a novel training strategy that decomposes the ill-posed reconstruction task into two well-defined sub-problems: structure recovery and texture rendering. By applying asymmetric supervision (structural loss on the encoder, texture loss on the decoder), SDFL provides a more principled learning process. These improvements enhance reconstruction quality, training stability, and reliability, culminating in new state-of-the-art (SOTA) performance on both the SEN12MS-CR and LuojiaSET-OSFCR large-scale optical-SAR cloud removal datasets. Notably, ECRformer surpasses previous SOTA methods by 1.23/0.90 dB in PSNR, while requiring only 28.9% of the parameters and 24.5% of the FLOPs, providing a powerful, efficient, and reliable solution for multimodal cloud removal. The code is available at https://github.com/zzaiyan/ECRformer.
Existing cloud detection methods often rely on deep neural networks, leading to excessive computational overhead. To address this, we propose a shallow convolutional neural network (CNN)-Transformer hybrid architecture that limits the maximum downsampling rate to 8x. This design preserves local details while effectively capturing global context through a lightweight Transformer branch. To enhance adaptability across diverse cloud scenes, we introduce two novel statistics-driven modules: statistics-adaptive convolution (SAC) and statistical mixing augmentation (SMA). SAC dynamically generates convolutional kernels based on input feature statistics, enabling adaptive feature extraction for varying cloud patterns. SMA improves model generalization by interpolating channel-wise statistics across training samples, increasing feature diversity. Experiments on four datasets show that the proposed method achieves state-of-the-art performance with 732 K parameters and 1G multiply-accumulate operations (MACs). Our code will be available at https://weix-liu.github.io/ for further research.
Accurate crop mapping plays a critical role in optimizing agricultural monitoring and ensuring food security. Although data-driven deep learning methods have demonstrated success in crop mapping with satellite image time series (SITS) data, their promising performances heavily depend on labeled training samples. Nevertheless, the difficulty of annotating crop types often results in labeled data scarcity, leading to a decline in the model's performance. Self-supervised learning (SSL) is a novel technique for crop mapping with limited labels. However, the existing SSL methods applied to SITS data typically explore masking solely on temporal dimension, which cannot guarantee strong spatial representation and therefore hinders the accurate prediction of complex crop fields. Furthermore, these methods sequentially extract spatial and temporal information without fully integrating information across different dimensions. In this study, we propose a spatiotemporal masking strategy for pre-training a SpatioTemporal Collaborative Learning Network (STCLN) to extract informative spatial and temporal representations from SITS data. Additionally, we design a SpatioTemporal Attention (STA) module in STCLN that integrates representations from spatial and temporal dimensions. The experimental results on two crop type mapping benchmarks encompassing various crop types demonstrate the outperformance of our proposed method. STCLN_wp outperforms the previous state-of-the-art (SOTA) methods with 6.49% higher mIoU on PASTIS dataset and 4.04% higher mIoU on MTLCC dataset. The ablation experiments on pre-training, masking strategies, and the STA module validate the effectiveness of our methodological design. Additionally, experiments conducted under varying sizes of the training set highlight the superior generalization ability of our method for crop type mapping in label-scarce situations. The code of our method is available at https://github.co m/XiaoleiQinn/STCLN.
Despite the in-depth understanding of the synthetic aperture-radar (SAR) speckle and its characteristics, despeckling remains an open issue far from being solved. Deep-learning methods with supervised training have made great progress. However, reliable reference images are inconveniently accessible or even non-existent. In this paper, we propose an end-to-end self-supervised method named Speckle2Self for SAR image despeckling, which learns mapping from noisy input to clean output using only the input noisy image itself for training. We formulate the image despeckling as a masked pixel-estimation problem, where a set of masks is carefully designed. The masked pixel values are predicted by the queries of complementary masks indicating the positions of masked pixels through an attention mechanism. Transformer architecture is employed as the network backbone. In addition, a novel loss function is also derived based on the statistics of SAR images, and meanwhile, image downsampling is used to provide guarantees on the white noise assumption involved in our Speckle2Self. We compare the proposed Speckle2Self with reference methods on both synthetic and real images. Experimental results demonstrate that the proposed Speckle2Self achieves comparable despeckling performance with supervised methods, suppressing noise while maintaining structural details. Even compared with self-supervised methods, the proposed Speckle2Self still has significant advantages in SAR image-despeckling metrics.
Models of dense prediction based on traditional Artificial Neural Networks (ANNs) require a lot of energy, especially for image restoration tasks. Currently, neural networks based on the SNN (Spiking Neural Network) framework are beginning to make their mark in the field of image restoration, especially as they typically use less than 10% of the energy of ANNs with the same architecture. However, training an SNN is much more expensive than training an ANN, due to the use of the heuristic gradient descent strategy. In other words, the process of SNN's potential membrane signal changing from sparse to dense is very slow, which affects the convergence of the whole model.To tackle this problem, we propose a novel distillation technique, called asymmetric framework (ANN-SNN) distillation, in which the teacher is an ANN and the student is an SNN. Specifically, we leverage the intermediate features (feature maps) learned by the ANN as hints to guide the training process of the SNN. This approach not only accelerates the convergence of the SNN but also improves its final performance, effectively bridging the gap between the efficiency of the SNN and the superior learning capabilities of ANN. Extensive experimental results show that our designed SNN-based image restoration model, which has only 1/300 the number of parameters of the teacher network and 1/50 the energy consumption of the teacher network, is as good as the teacher network in some denoising tasks.
Change detection (CD) enables the identification of alterations between images of the same area captured at different times. However, existing CD methods still struggle to mitigate pseudo-changes resulting from domain information differences in multi-temporal images and instances of detail errors caused by the loss and contamination of detail features during the upsampling process in the network. To better alleviate these problems, we propose a bi-temporal Gaussian distribution feature-dependent (BGFD) network. Specifically, we first introduce the Gaussian noise domain disturbance (GNDD) module, which approximates the distribution using image statistical features to characterize domain information and samples noise to perturb the network for learning redundant domain information, mitigating domain information differences. Additionally, within the feature dependency facilitation (FDF) module, we integrate a novel mutual information difference loss (${L_{MI}}$LMI) and more sophisticated attention mechanisms to enhance the capabilities of the network, ensuring the acquisition of essential domain information. Subsequently, we have designed a novel detail feature compensation (DFC) module, which compensates for detail feature loss and contamination introduced during the upsampling process from the perspectives of enhancing local features and refining global features. The BGFD has effectively reduced pseudo-changes and enhanced the detection capability of detail information. It has also achieved state-of-the-art performance on four publicly available datasets - DSIFN-CD, SYSU-CD, LEVIR-CD, and S2Looking, surpassing baseline models by +8.58%, +1.28%, +0.31%, and +3.76% respectively, in terms of the F1-Score metric.
The multimodal remote sensing image matching is crucial for many applications. However, nonlinear intensity distortion (NID) significantly impairs matching performance, especially when dealing with scale and rotation variations. To address this challenge, we propose a global-to-local invariant feature transformation (GLIFT) method for multimodal remote sensing image matching. The method consists of three key components: feature detection, global search, and local search. First, we introduce a fast dominant orientation assignment approach, which ensures rotational invariance while reducing computational costs. Next, we design a 3-D descriptor structure that effectively integrates both local region information and keypoint self-information, enhancing the robustness and discriminability of the descriptor. To overcome the limitations of image pyramids in handling scale variations in multimodal remote sensing images, we propose a local multiregion description strategy that adapts well to scale changes. In addition, we construct a pixel-based descriptor vector and present a local search strategy to identify optimal matching point pairs within local regions, which effectively improves the matching accuracy. Finally, we validate the matching performance of GLIFT by conducting experiments and comparing it with eight state-of-the-art multimodal matching algorithms on various datasets. Extensive results demonstrate that our method effectively addresses the challenges posed by rotation and scale variations in multimodal remote sensing images. It significantly enhances the number of correct matches (NCMs), matching accuracy, and matching precision. Our code is available at: https://github.com/wdzsc/GLIFT
Satellite image time series (SITS) data provides continuous observations over time, allowing for the tracking of vegetation changes and growth patterns throughout the seasons and years. Numerous deep learning (DL) approaches using SITS for crop classification have emerged recently, with the latest approaches adopting Transformer for SITS classification. However, the quadratic complexity of self-attention in Transformer poses challenges for classifying long time series. While the cutting-edge Mamba architecture has demonstrated strength in various domains, including remote sensing image interpretation, its capacity to learn temporal representations in SITS data remains unexplored. In this paper, we proposed a Satellite Image Time Series Mamba (SITSMamba) method for crop classification based on remote sensing time series data. The proposed SITSMamba contains a spatial encoder based on Convolutional Neural Networks (CNN) and a Mamba-based temporal encoder. When evaluated on the MTLCC dataset, our SITSMamba achieves an OA of 0.9100, outperforming the previous state-of-the-art (SOTA) methods. The code will be released at https://github.com/XiaoleiQinn/SITSMamba.
In the realm of image super-resolution, learning-based methods have made significant progress. However, limited computational resources still restrict their application. This prompts us to develop an efficient method for achieving effective image super-resolution. In this letter, we propose a novel adaptive feature selection modulation network (AFSMNet) tailored for efficient image super-resolution. Specifically, we design feature modulation blocks, which include the adaptive feature selection modulation (AFSM) module and the self-gating feed-forward network (SFN). The AFSM module dynamically computes the importance of each feature channel. For channels with differing levels of importance, we employ distinct processing strategies, thereby concentrating the computational resources of the network on the more critical features as much as possible. This approach facilitates the maintenance of a low computational cost without compromising performance. The SFN restricts the flow of irrelevant feature information within the network through a simple gating mechanism. In this way, our method achieves efficient and effective image super-resolution. Extensive experiment results show that the proposed method achieves a better trade-off between reconstruction performance and computational efficiency compared to the current state-of-the-art lightweight super-resolution methods.
Satellite image time series (SITS) provide continuous observations of the Earth's surface, making them essential for applications such as environmental management and disaster assessment. However, existing spatiotemporal foundation models rely on plain vision transformers, which encode entire temporal sequences without explicitly capturing multiscale spatiotemporal relationships between land objects. This limitation hinders their effectiveness in downstream tasks. To overcome this challenge, we propose TiMo, a novel hierarchical vision transformer foundation model tailored for SITS analysis. At its core, we introduce a spatiotemporal gyroscope attention mechanism that dynamically captures evolving multiscale patterns across both time and space. For pre-training, we curate MillionST, a large-scale dataset of one million images from 100,000 geographic locations, each captured across 10 temporal phases over five years, encompassing diverse geospatial changes and seasonal variations. Leveraging this dataset, we adapt masked image modeling to pre-train TiMo, enabling it to effectively learn and encode generalizable spatiotemporal representations.Extensive experiments across multiple spatiotemporal tasks-including deforestation monitoring, land cover segmentation, crop type classification, and flood detection-demonstrate TiMo's superiority over state-of-the-art methods. Code, model, and dataset will be released at https://github.com/MiliLab/TiMo.
Restoring images afflicted by complex real-world degradations remains challenging, as conventional methods often fail to adapt to the unique mixture and severity of artifacts present. This stems from a reliance on indirect cues which poorly capture the true perceptual quality deficit. To address this fundamental limitation, we introduce AdaQual-Diff, a diffusion-based framework that integrates perceptual quality assessment directly into the generative restoration process. Our approach establishes a mathematical relationship between regional quality scores from DeQAScore and optimal guidance complexity, implemented through an Adaptive Quality Prompting mechanism. This mechanism systematically modulates prompt structure according to measured degradation severity: regions with lower perceptual quality receive computationally intensive, structurally complex prompts with precise restoration directives, while higher quality regions receive minimal prompts focused on preservation rather than intervention. The technical core of our method lies in the dynamic allocation of computational resources proportional to degradation severity, creating a spatially-varying guidance field that directs the diffusion process with mathematical precision. By combining this quality-guided approach with content-specific conditioning, our framework achieves fine-grained control over regional restoration intensity without requiring additional parameters or inference iterations. Experimental results demonstrate that AdaQual-Diff achieves visually superior restorations across diverse synthetic and real-world datasets.
Self-supervised monocular depth estimation from oblique UAV videos is crucial for enabling autonomous navigation and large-scale mapping. However, existing self-supervised monocular depth estimation methods face key challenges in UAV oblique video scenarios: depth discontinuity from geometric distortion under complex viewing angles, and spatial ambiguity in weakly textured regions. These challenges highlight the need for models that combine global reasoning with geometric awareness. Accordingly, we propose RMTDepth, a self-supervised monocular depth estimation framework for UAV imagery. RMTDepth integrates an enhanced Retentive Vision Transformer (RMT) backbone, introducing explicit spatial priors via a Manhattan distance-driven spatial decay matrix for efficient long-range geometric modeling, and embeds a neural window fully-connected CRF (NeW CRFs) module in the decoder to refine depth edges by optimizing pairwise relationships within local windows. To mitigate noise in COLMAP-generated depth for real-world UAV datasets, we constructed a high-fidelity UE4/AirSim simulation environment, which generated a large-scale precise depth dataset (UAV SIM Dataset) to validate robustness. Comprehensive experiments against seven state-of-the-art methods across UAVID Germany, UAVID China, and UAV SIM datasets demonstrate that our model achieves SOTA performance in most scenarios.
Synthetic aperture radar (SAR) images are inherently affected by speckle noise due to their imaging principles, significantly impacting downstream research. Denoising methods based on deep learning have garnered attention and are gradually maturing, yet they face certain challenges. Denoising single SAR images lacks temporal information, often resulting in structural fitting that introduces artifacts. The existing denoising techniques do not adequately account for the differences between SAR and optical images, despite their distinct structural characteristics. In addition, mainstream deep learning approaches heavily rely on data-driven methods, with limited consideration for statistical properties. In response to these challenges, we propose a novel network for denoising SAR images fusing multitemporal interaction with coherence prior, termed MTCSAR. The proposed network leverages multitemporal interaction (MTI) to gather information from different moments in various regions. The redundancy across these temporal dimensions helps mitigate artifacts and edge blurring. To address SAR image characteristics, we design a dual-branch network combining global and local information to handle large-scale regions and fine structures. In addition, we incorporate prior coherence information from SAR images into the network, utilizing statistical properties to enhance transparency during training. Experimental results on simulated and real datasets demonstrate that injecting MTIs and coherence improves our method's qualitative and quantitative performance, surpassing current state-of-the-art algorithms. This validates the effectiveness of the proposed MTCSAR for multitemporal SAR image denoising.