Hyperspectral image super-resolution (HSISR) is dedicated to reconstructing a high-resolution hyperspectral image (HrHSI) from a low-resolution hyperspectral image (LrHSI) in conjunction with a high-resolution multispectral image (HrMSI). Although recent unsupervised approaches have mitigated the reliance on paired HrHSI supervision, many of them still hinge on pre-defined degradation priors or predominantly leverage local spectral correspondences, thereby constraining their robustness amidst cross-resolution distribution shifts and intricate scene configurations. To tackle these challenges, we present in this paper an unsupervised HSISR approach that integrates spectral mapping learning with overall consistency optimization, SMO-Net for short. Specifically, a degradation information estimation module is employed to adaptively ascertain the spatial and spectral degradations from the observed LrHSI and HrMSI. Subsequently, a multiscale spectral learning network establishes a transferable multispectral-to-hyperspectral mapping within the low-resolution domain, facilitating the generation of an initial HrHSI estimate in the high-resolution domain. Ultimately, an overall consistency optimization module refines this initial estimate through multiscale feature aggregation and cross-resolution interaction. Experiments on three simulated datasets show competitive spectral and reconstruction performance, and a real ZY-1 02D scene further verifies its utility for downstream classification.
In recent years, hyperspectral and multispectral image (HSI-MSI) fusion has attracted considerable attention as an effective technique for producing images with both high spectral and spatial resolution. However, most existing methods focus on overlapping regions where both HSI and MSI observations are available. In practice, narrow sensor swaths in operational satellites produce incomplete hyperspectral coverage, leaving non-overlapping regions without direct spectral observations. Abundance-based spectral unmixing offers a principled solution, where shared endmembers propagate spectral knowledge across regions. Yet accurate endmember learning faces two obstacles: idealized point spread function (PSF) assumptions that mischaracterize realistic sensor blur, and uniform feature extraction across varying observational conditions. To this end, we propose a novel end-to-end framework addressing both issues through complementary mixture-of-experts mechanisms. Specifically, the PSF Mixture-of-Experts (PSF-MoE) module employs a three-stage progressive strategy to capture physically-grounded degradation patterns, providing accurate priors for endmember learning. Additionally, the Region-Aware Mixture-of-Experts (RAMoE) mechanism routes features based on observational regimes, enabling specialized processing across overlapping, non-overlapping, and transition regions. This joint design enables accurate spectral reconstruction across the entire spatial domain, including regions lacking direct hyperspectral observations. Experiments on established benchmarks demonstrate state-of-the-art performance, while validation on operational ZY1-02D satellite imagery against unmanned aerial vehicle (UAV) ground truth confirms practical effectiveness. Our implementation is publicly accessible at https://github.com/Welcome-to-LISA/RAMoE.
As spaceborne hyperspectral imaging technology advances, satellite hyperspectral images (HSIs)-with contiguous narrowband coverage and high-spectral resolution-have been widely applied to resource management and environmental monitoring. However, due to sensor and platform constraints, hyperspectral data typically exhibit a low-temporal resolution, which limits their use in long-term dynamic monitoring. To this end, this letter proposes an interactive multiscale temporal-spectral fusion network (IMFNet) that fuses the high-temporal-frequency information of multispectral images (MSIs) with the high-resolution spectral information of HSIs to reconstruct the HSI at the target time. In particular, features are first extracted from the reference HSIs and multitemporal MSIs, and the spectral attention block (SAB) is introduced to enhance spectral representation and mitigate spectral distortion. The cross-feature interaction module (CFIM) is designed to promote cross-modal feature interaction via mutual feature guidance. Finally, the adaptive multiscale fusion module (AMFM) captures local details and global structural information using different receptive fields for adaptive aggregation. Experimental results show that IMFNet outperforms competing methods on real datasets, indicating its effectiveness and practical potential in hyperspectral time-series reconstruction.
Hyperspectral imaging can capture abundant spectral information and reveal the spectral absorption properties of surface materials. Nevertheless, the tradeoff in spatial resolution reduces its capacity to represent surface object textures and structures; hyperspectral super-resolution (SR) technology is a viable solution to this problem. Yet, mainstream supervised methods depend on low-scale training data and specific data distributions, restricting their generalization capability and practicality in real scenarios. Although unsupervised methods remove the reliance on training data, they still face suboptimal reconstruction quality due to the absence of reference images and inaccuracies in degradation process estimation. Furthermore, bridging the performance gap between simulated datasets and real-world applications remains challenging. To this end, we propose a semi-supervised network that effectively couples unsupervised learning and supervised pretraining in multistage architecture, MCS-Net for short. The network consists of three key components: degradation information estimation (DIE), supervised fusion pretraining (SFP) at low resolution, and unsupervised image generation (UIG) at full resolution. The MCS-Net first estimates deep degradation information from input image pairs using DIE. It then applies supervised learning in SFP to construct a pretrained fusion function and its parameters from the input low-resolution data pairs and their fused outputs. Finally, the pretrained parameters from the previous stage are used to initialize the fusion network of UIG, which is then fine-tuned under the guidance of degradation parameters estimated by DIE, enabling the network to process full-resolution images effectively. Ablation experiments validated the effectiveness of each component. Moreover, the proposed MCS-Net outperforms the existing state of the art (SOTA) methods across six evaluation metrics in the simulation experiments, and the experimental results on real satellite data further validate the outstanding image fusion performance of MCS-Net and its potential for practical applications.
Depending on a large-scale paired dataset of low-resolution hyperspectral image (LrHSI), high-resolution multispectral image (HrMSI), and corresponding high-resolution hyperspectral image (HrHSI), the supervised paradigm has achieved impressive performance in the hyperspectral image super-resolution (HISR). However, the intrinsic data-intensive manner hinders its further application in real scenarios. Fortunately, deep image prior (DIP) allows us to achieve unsupervised super-resolution (SR) by solely utilizing degraded observations. However, its potential to accurately model complicated hyperspectral priors is still not fully exploited due to the following two factors: 1) existing methods tend to reconstruct the unknown HrHSI directly from a randomly generated noise, leaving it hard to leverage the scene-relevant information for prior learning and 2) the vanilla architecture is handcrafted for the generator network, which shows limitations in feature representation and thus fails to characterize the complicated image properties. To unleash the potential of DIP for the HISR task, we propose an enhanced DIP network, called EDIP-Net, by addressing the aforementioned impediments. Specifically, EDIP-Net is built with a two-stage four-component scheme, with a zero-shot learning (ZSL) stage for input image establishment and a deep image generation (DIG) stage for prior learning. First, we exploit the cross-scale spectral relationship inside the observations and thus design a degradation learning network to generate paired training samples from the observations themselves. As such, two image-coarse estimations are derived in a ZSL manner by learning an interactive spectral learning network. By replacing random noise with two estimations, we design a double U-shape architecture for the generator network to capture their hyperspectral prior, each independently generating one HrHSI candidate. Under this premise, we further propose a degradation-aware decision fusion strategy to integrate the optimal results in a pixel-to-pixel manner. Extensive experiments demonstrate our superiority in achieving high-quality SR performance. The code will be available at https://github.com/JiaxinLiCAS.
Due to the limitations of hyperspectral satellite imaging systems, hyperspectral images (HSIs) have a low spatial resolution and spatial missing regions, which hinders their full potential for land cover classification. Through collaboration with the auxiliary high-resolution multispectral image (MSI), the high-resolution HSI (HHSI) can be obtained. Although image fusion deep learning methods have emerged as the predominant approach for enhancing the spatial resolution of HSI, few studies have solved the problem of spatial missing. Moreover, the reliance on the spectral response function and the point spread function limits the applicability and performance of these methods. This study proposes an HSI cross-region reconstruction framework, which can not only fuse the HSI and MSI in overlapping regions but also be used to reconstruct the missing information of HSI in non-overlapping regions. The core of this framework is an unsupervised deep end-to-end fusion network with adaptive attention based on the conception of endmember and abundance. Specifically, a multi-layer attention mechanism with an interaction module is designed to enhance the extraction of endmember and abundance features. In addition, the attached module addresses the dependency on the prior information about degradation models. To verify the practicability of the proposed framework, we conduct evaluations using both simulation datasets and real on-board data for experiments involving both overlapping and non-overlapping regions. The experimental results demonstrate that the proposed method is effective and the reconstructed HHSI is reliable.
Spatiotemporal fusion (STF) technology is an effective means to address the challenge of balancing temporal and spatial resolutions for single satellite sensors. Remote sensing imagery exhibits substantial scale variations in different ground objects. However, existing CNN-based models employ a fixed receptive field during feature extraction, leading to a lack of dynamic adjustment capability when capturing information at different scales, which limits the fusion accuracy. To overcome this limitation, our study proposes the Multi-Kernel Adaptive Network (MKAN), specifically designed for STF tasks. The network incorporates our proposed Dynamic Visual Processing Module (DVPM) to achieve adaptive perception and intelligent fusion of multi-scale geophysical features. Furthermore, to enhance the geometric consistency of DVPM during the process of cross-scale feature fusion of geophysical features, we design the Multi-Dimensional Perceptual Attention Module (MDPA), which integrates channel and spatial domain information enhancement techniques for precise feature identification and optimized pixel-level resource allocation. To comprehensively evaluate the performance advantages of MKAN, we conducted thorough experimental assessments on three datasets and benchmarked it against five spatiotemporal fusion algorithms. The experimental results demonstrate that MKAN excels in multiple evaluation metrics. Compared to optimal comparison algorithms, MKAN increases structural similarity by an average of 1.4% and reduces global dimensionless error by an average of 4.4%, fully demonstrating the rationality and superiority of its design. In addition, this chapter explores the effectiveness of the MKAN network structure and assesses the contributions of the DVPM and MDPA modules to the overall algorithm performance through meticulous ablation experiments conducted from various perspectives.
Deep learning has become increasingly popular in hyperspectral image (HSI) and light detection and ranging (LiDAR) data classification, thanks to its powerful feature learning and representation capabilities. However, HSI often contains substantial redundant information, which can hinder efficient data utilization. Furthermore, the significant disparity in information content between HSI and LiDAR data poses a major challenge in representing and aligning semantic information across these two modalities. To address these challenges, we propose a fusion network structure guided by feature reconstruction embedding (FRE). This approach employs feature decomposition to reconstruct HSI features and incorporates weight embedding to seamlessly integrate the reconstructed information into classification features. Furthermore, we introduce a cross-modal attention fusion module designed to merge extracted HSI and LiDAR features. This module fully exploits the complementary nature of these two type of feature, facilitating effective information exchange and semantic alignment across multimodal data. We evaluated our method on three widely used HSI and LiDAR datasets: Houston 2013, Augsburg, and MUUFL. Experimental results demonstrate that our proposed FRGFNet significantly outperforms traditional probabilistic methods and state-of-the-art deep learning networks, showcasing its effectiveness in multisource data fusion.
The fusion of hyperspectral images (HSIs) and multispectral images (MSIs) is crucial for overcoming the limitations of low spatial resolution in HSI. Currently, supervised learning methods tend to yield satisfactory integration results when applied to data distributions similar to those of the training set; however, they often exhibit insufficient generalization when confronted with real-world application scenarios. In contrast, unsupervised methods exhibit good generalization capabilities; however, they typically require careful tuning of hyperparameters to achieve satisfactory results, primarily due to the lack of sufficiently clear training objectives. To fully leverage the advantages of both supervised and unsupervised learning, this letter proposes an unsupervised pretraining framework (UPFW) guided fusion approach, which effectively enhances the performance of HSI-MSI fusion by introducing low-resolution supervised pretraining and full-resolution unsupervised adaptive strategy. Specifically, in the first stage, the model adapts to the learning spatial and spectral degradation parameter; in the second stage, we propose an adaptive fusion network (ADFNet) and conduct supervised learning on low-resolution scale to obtain a pretrained fusion network model with a clear objective-oriented; in the third stage, we utilize the pretrained model for full-resolution unsupervised fusion, thereby enhancing the model's generalization capabilities and applicability. Experimental results show that compared to traditional methods and other deep learning approaches, the proposed method achieves significant advantages in spectral fidelity and spatial detail recovery across multiple public datasets.
Hyperspectral images (HSI) are renowned for their high spectral resolution and extensive wavelength coverage, but suffer from limited spatial and temporal resolution due to imaging sensor constraints. However, hyperspectral images with low spatial resolution and low temporal resolution are difficult to be applied to subsequent more advanced tasks such as object detection, classification, and anomaly detection. Physical constraints make it impossible for a single satellite sensor to acquire images that simultaneously have high resolution in time, space, and spectrum, so image fusion is the most efficient choice to achieve this goal. Spatial-temporal-spectral fusion (STSF) has the purpose of synthesizing different information which has the advantage in temporal, spatial, and spectral aspect respectively from multisource satellite data to reconstruct HSI with high spatial resolution and high temporal resolution. In order to address the problems that the linear relationship in current spectral reconstruction is difficult to accurately map the complex relationships among space, time and spectrum, and that deep convolutional neural networks are prone to overfitting, as well as to enhance the model's ability to extract spectral features, this paper designs an unsupervised STSF method. The proposed method has three stages: Stage 1 (Spatial-Spectral Downsampling) analyzes the spatial-spectral degradation of time 1 observed images; Stage 2 (Spectral Upsampling) develops a spectral upsampling network (with the shared spatial-spectral downsampling network) to upsample multispectral data to hyperspectral data; Stage 3 uses the trained network to upsample multispectral images of time 2 for high-spatial-resolution HSI. In order to verify the proposed method, it is compared with other state-of-the-art methods on simulated and real datasets. This proves the method's advantages in richer spatial-spectral details and more accurate reconstruction.
Fusing low-resolution hyperspectral images (HSI) with high-resolution multispectral images (MSI) has become a promising technique for generating high-resolution HSI, effectively addressing the low spatial resolution limitations of hyperspectral data. Deep learning-based fusion methods have emerged as the dominant approach in recent research. However, almost unsupervised deep fusion methods rely on degradation processes that may disrupt and lose the dominant spectral-spatial information in the original images. Moreover, they only construct spectral inverse mapping while lacking effective spectral-spatial interactions, resulting in insufficient detail preservation, thereby affecting fusion performance. In this study, we propose a deep unsupervised pretraining fusion framework of spectral-spatial collaborative constraint. Specifically, in the first stage, both spatial and spectral inverse modules are adaptively pretrained from the original HSI and MSI, providing initialize parameters for further optimization; in the second stage, the model learns spatial and spectral degradation modules to support the construction of inverse modules for the subsequent stage; in the third stage, guided by pretrained parameters, we further optimize the inverse mapping and effectively extract spectral-spatial information at a different resolution, thereby enhancing the fusion performance and applicability. Experimental results on simulated and real satellite data verify the superiority of our proposed method in recovering spatial details and preserving spectral fidelity.
Hyperspectral images (HSI) are renowned for their high spectral resolution and extensive wavelength coverage, but they often suffer from limited spatial and temporal resolution due to imaging sensor constraints. This may make it difficult for hyperspectral images to play a role in the acquisition of fine surface information and the observation of continuity of time scales. Spatial-temporal-spectral fusion (STSF) has the purpose of synthesizing different information which from multisource satellite data which has the advantage in temporal, spatial, and spectral aspect respectively to reconstruct HSI with high spatial resolution and high temporal resolution. However most existing spatial-temporal-spectral fusion methods are restricted to the assumption of linear relationship of space temporal and spectrum. Furthermore, the spatial-temporal-spectral fusion methods mostly cannot directly process the current hyperspectral data which has a lower temporal resolution based on Landsat and MODIS. This paper proposed an unsupervised STSF model for hyperspectral images using a global shared convolutional neural network (UGSCNN). The proposed method has two models: 1) Spatial-spectral down model: combine spectral linear theory with deep learning to down sample the space and the spectrum; 2) Spectral up model: utilize the shared global information to up sample the spectrum. To verify the proposed method, we compared it with others using simulated and real on-board data, proving its effectiveness and practical fusion results.
Due to its powerful feature extraction and representation capabilities, deep learning has been successfully applied in the field of hyperspectral image classification. In patchbased hyperspectral image classification, where the central pixel represents the true category, extracting spectral features similar to the central pixel spectrum is crucial. To this end, we develop a new framework, called the inverse Mahalanobis attention network (IMAN), to address the spectral similarity feature extraction problem. The proposed framework develops a self-attention mechanism module based on the Mahalanobis distance to better learn the correlation between feature vectors of pixels, thereby efficiently suppressing noise generated by different land cover categories within patches. The feature extraction capability is further enhanced by integrating a dual-stream network structure that separates spatial and spectral information in hyperspectral images. Experiments conducted on on real hyperspectral datasets demonstrate the effectiveness and superiority of the proposed method compared to several state-of-the-art hyperspectral image classification methods.
By fusing a low-resolution hyperspectral image (LrMSI) with an auxiliary high-resolution multispectral image (HrMSI), hyperspectral image super-resolution (HISR) can generate a high-resolution hyperspectral image (HrHSI) economically. Despite the promising performance achieved by deep learning (DL), there are still two challenges remaining to be solved. First, most DL-based methods heavily rely on large-scale training triplets, which reduces them to limited generalization and poor practicability in real-world scenarios. Second, existing methods pursue higher performance by designing complex structures from off-the-shelf components while ignoring inherent information from the degradation model, hence leading to insufficient integration of domain knowledge and lower interpretability. To address those drawbacks, we propose a model-informed multi-stage unsupervised network, M2U-Net for short, by leveraging both deep image prior (DIP) and degradation model information. Generally, M2U-Net is built with a three-stage scheme, i.e., degradation information learning (DIL), initialized image establishment (IIE), and deep image generation (DIG) stages. The first stage is to exploit the deep information of the degradation model via a tiny network whose parameters and outputs will serve as guidance for the following two stages. Instead of feeding uninformed noise as input for stage three, IIE stage aims to establish an initialized input with expressive HrHSI-relevant information by resorting to a spectral mapping learning network, thus facilitating the extraction of prior information and further magnifying the potential of DIP for high-quality reconstruction. Last, we propose a dual U-shape network as a powerful regularizer to capture image statistics, in which two U-Nets are coupled together by cross-attention guidance (CAG) module to separately achieve spatial feature extraction and final image generation. The CAG module can incorporate abundant spatial information into the reconstruction process and hence guide the network toward a more plausible generation. Extensive experiments demonstrate the effectiveness of our proposed M2U-Net in terms of quantitative evaluation and visual quality. The code will be available at https://github.com/JiaxinLiCAS.
The detection of objects with larger aspect ratios (OLARs) is a challenging problem in a special application scenario, such as remote-sensing object recognition and scene text detection. However, current object detectors perform poorly in OLAR feature extraction because they are incapable of adaptively responding to object shapes, which leads to severe misalignment between impure feature representations and region proposals. In this letter, we aim to solve this problem by proposing our shape-sensitive convolution network (SSC-Net). SSC-Net is carefully embedded with a feature enhancement module (SSC module) specifically suitable for OLAR. This module can use fewer sampling points to achieve more intelligent feature sampling area transformation, thus achieving the goal of enhancing OLAR feature representation. Extensive experiments on benchmark datasets that are rich in OLARs have proved the superiority of our method. Besides, we further verified the plug-and-play performance of the SSC module, and the experimental results show that it can significantly improve the detection performance of the detector for OLAR.
Spatiotemporal fusion (STF) technology effectively addresses the difficulty in obtaining high spatial and temporal resolution images due to compromises in satellite design. Deep learning algorithms are extensively applied in this field, but their performance is significantly impacted by hardware factors, such as the vast differences in sensor resolutions. To address this issue, we have fully considered the scale differences among data from various sources and proposed a two-stage coarse-to-fine STF approach, named SIFnet. SIFnet comprises two stages; the first stage concentrates on learning the scale differences between the data, while the second stage integrates feature maps that represent different spatial resolutions to impose constraints on the network's learning. This approach enables the preservation of detailed information by reusing feature maps at different scales during the learning process. As a result, the mapping relationship with the real image is closer. In this letter, three datasets containing different ground feature characteristics are used for STF experiments. The results demonstrate that SIFnet can effectively utilize the concept of scale transformation to improve feature extraction and enhance the accuracy of reconstructed images. Compared with various advanced STF algorithms, SIFnet achieves an average improvement of 2% in structural similarity, proving the effectiveness of this method.
Hyperspectral image super-resolution (HI-SR) can improve hyperspectral images' spatial resolution to capture more spatial details from the observed scenario. Recently, fusing multispectral and hyperspectral images to implement HI-SR has become a hot topic in remote sensing. Moreover, the development of deep learning has further promoted the advancement of HI-SR over the past few years, bringing many specific HI-SR networks. However, for most HI-SR networks, it is challenging to obtain global features from different images, which limits the reliability of results. We propose a new HI-SR method that adopts a two-stream self-attention network (TSSA-Net) to address the above issues. The proposed deep network consists of two coupled encoders for acquiring the abundance and endmembers from multispectral and hyperspectral images, respectively. Each encoder has a self-attention mechanism to acquire global features and use a convolutional layer to aggregate local features. Meanwhile, cross-stream self-attention is designed for information exchange among different streams, enhancing the robustness of TSSA-Net. Several experiments are conducted to verify the effectiveness and competitiveness of TSSA-Net.
The adequate and finer spectral information in hyperspectral images (HSIs) are benefit for various downstream applications like smart agriculture and environmental monitoring. In HSI classification, dual-stream convolutional networks have gained much attention and have been widely used. In patch-based hyperspectral classification tasks, however, merely using center-labeled patches could lead to an increased unlabeled noise in the data. Moreover, in the application of dual-stream network structures, heterogeneity existed in both the data and feature semantic levels to capture more representative features. To tackle these challenges, we have devised a framework called cross-semantic heterogeneous modeling network (CreatingNet), which aligns more closely with the design principles of dual-stream networks by adjusting the input size. This framework introduces a distance metric attention mechanism (DMAM) based on spectral and spatial distances to strengthen the influence of the center pixel on the entire patch. Additionally, we present a fusion module named CrossViT, which combines features with diverse structures and characteristics, leveraging their complementarity. The proposed multiscale heterogeneous fusion module allows for more effective integration of spatial and spectral features in the images. Extensive experiments on four well-known HSI datasets (Indian Pines, Pavia University, Salinas, and Houston 2013) demonstrate the superior classification performance of the proposed CreatingNet to several state-of-the-art methods. The effectiveness of the proposed model is further validated through ablation studies.