Reconstructing the high-quality hyperspectral image (HSI) from compressed measurement is a core task in spectral compressive imaging. With the development of deep learning, the deep unfolding networks (DUNs) could learn the deep prior for reconstructing the latent HSI. However, the spectral compressive imaging is a heavily-degraded process, the current HSI reconstruction methods still face the challenges in restoring fine-grained and realistic image details from the compressed measurement. As a generative model, the diffusion model has high potential in generating natural and realistic image content. In this study, we jointly leverage the diffusion prior and deep prior, and propose a diffusion and deep priors regularized network for HSI reconstruction (D&D-Net). Specifically, based on the spectral compressive imaging, we first propose a HSI reconstruction model regularized by diffusion prior and deep prior. By optimizing the reconstruction model and unrolling the iterative solutions into a deep network, we build D&D-Net. In the network, the diffusion and deep priors are represented by diffusion and deep proximal operators respectively. We propose a multi-task diffusion model as the diffusion proximal operator, in which both the unknown coded aperture mask and image contents are generated. Furthermore, we also design a mask-guided multi-scale state-space model as the deep proximal operator, which efficiently exploits the long-range spectral-spatial dependency for HSI reconstruction under the guidance of generated mask. In the experiments, the proposed method could achieve competitive performance on both simulated and real measurements, which demonstrates the effectiveness of joint learning diffusion and deep priors.
Hyperspectral image (HSI) is applicable in many fields due to the ability in discriminating different materials. Collecting HSI usually requires expensive hardware and long period. Reconstructing HSI from RGB image, also called spectral super-resolution (SSR), is an affordable and feasible way for HSI acquisition. Despite the SSR results achieved by existing deep unfolding networks (DUNs), they still face challenges in: 1) recovering the fine-grained and realistic details; 2) suppressing the spectral distortion. Diffusion model has advantages in generating diverse and realistic contents, while its fidelity is limited due to the inherent randomness. In this study, to reconstruct a faithful and realistic HSI, we integrate the diffusion model in DUN, and propose a degradation-aware unrolling diffusion model for SSR (deDiff-SSR). The generative diffusion prior is jointly leveraged with the spectral degradation and deep prior learning. Specifically, we first pre-train a channel attention enhanced denoising diffusion probabilistic model (DDPM), the spectral correlation is exploited for learning the diffusion prior of HSI. To aware the degradation, by optimizing a diffusion and deep priors regularized HSI SSR model, we propose a degradation-aware diffusion sampling method, the spectral degradation is learned to refine each diffusion sampling step. Via unrolling the degradation-aware diffusion sampling steps, we build the deDiff-SSR network. It contains diffusion and deep proximal operators to represent the diffusion and deep priors, respectively. We implement the diffusion proximal operator with one sampling step of the pre-trained DDPM. Moreover, we design a state-space Transformer as the deep proximal operator, the spectral-spatial long-range relationship of HSI can be efficiently captured. The experiments on several indoor and remote sensing datasets demonstrate the effectiveness of deDiff-SSR.
Hyperspectral unmixing (HU) is a fundamental task in the analysis and interpretation of hyperspectral images. A large number of HU methods have been developed rapidly, and several review papers have summarized these developments. More recently, spectral variability (SV) has attracted increasing attention in HU, leading to the emergence of various SV-aware hyperspectral unmixing (SVHU) methods. While existing reviews cover general HU, dedicated surveys on SVHU remain scarce. Moreover, most relevant reviews focus on a specific methodological branch and fail to address the latest advances. To address this gap, we present a comprehensive review of SVHU methods. To the best of our knowledge, this is the first work to propose a unified SVHU framework that covers both single-temporal and multi-temporal settings. Existing SVHU methods are systematically categorized into three groups: model-driven, data-driven, and model–data-driven approaches, each of which is discussed in detail. Subsequently, we describe the datasets and evaluation metrics, and compare several representative state-of-the-art methods on multiple commonly used public datasets. In addition, we summarize related downstream applications, including hyperspectral super-resolution, classification, and change detection. Finally, we discuss the main challenges and outline several promising directions for future research.
Diffusion sampling-based Plug-and-Play (PnP) methods produce images with high perceptual quality but often suffer from reduced data fidelity, primarily due to the noise introduced during reverse diffusion. To address this trade-off, we propose Noise Frequency-Controlled Diffusion Sampling (NFCDS), a spectral modulation mechanism for reverse diffusion noise. We show that the fidelity-perception conflict can be fundamentally understood through noise frequency: low-frequency components induce blur and degrade fidelity, while high-frequency components drive detail generation. Based on this insight, we design a Fourier-domain filter that progressively suppresses low-frequency noise and preserves high-frequency content. This controlled refinement injects a data-consistency prior directly into sampling, enabling fast convergence to results that are both high-fidelity and perceptually convincing–without additional training. As a PnP module, NFCDS seamlessly integrates into existing diffusion-based restoration frameworks and improves the fidelity-perception balance across diverse zero-shot tasks.
Image Poisson noise is usually captured by photonlimited imaging systems. It is widely present in fields such as astronomical observations and medical imaging. Unlike Gaussian noise, Poisson noise has signal dependence and multiplicative characteristics, so its removal is a challenging task. In this paper, we propose a Poisson image denoising method via generalized Gaussian scale mixture (GGSM) modeling in the wavelet domain. Firstly, we perform the wavelet transform on the image to obtain wavelet high-frequency coefficients and low-frequency coefficients. Then, the GGSM framework is used to model the sparse prior of the high-frequency coefficients and simultaneously estimate the hidden scale parameters. For the low-frequency coefficients, the Tikhonov regularization is adopted to enforce smoothness. Subsequently, the denoising is formulated as a maximum a posteriori (MAP) estimation problem, which is efficiently solved by the alternating direction method of multipliers (ADMM). Experimental results show that the proposed method can effectively suppress Poisson noise and exhibit good competitiveness.
Spectral variability is one of the significant challenges in hyperspectral image unmixing. To address this issue, most unmixing methods consider spectral variability through linear mapping to model the imaging environment. However, in real-world imaging scenarios, factors such as illumination, atmospheric effects, and terrain can introduce nonlinear variation of the endmembers. This paper proposes a novel dual-stream unmixing network to address linear and nonlinear spectral variability, with each stream being obtained through variational autoencoders. To capture variability, the spectrum is projected into two latent subspaces, each following a Gaussian distribution with different parameters, representing the estimated posterior conditional probability distributions. Subsequently, endmembers and their variabilities, encompassing both linear and nonlinear variations, are derived by decoding the latent features. Additionally, the hyperspectral image is reconstructed by multiplying endmembers with abundances, facilitating network learning. Experiments show that the proposed method achieves superior unmixing performance and significantly enhances the accuracy of endmember extraction.
Hyperspectral and multispectral images fusion is a crucial technique for generating image with both high spatial and spectral resolution. Most existing fusion methods rely on the ideal assumption of spectral consistency between the hyperspectral and multispectral images by omitting the spectral variability with the imaging conditions. To address this issue, a novel fusion method is proposed that explicitly accounts for spectral variation. The hyperspectral and multispectral images are represented via their Block Term Decomposition (BTD), where spectral variability is introduced into the decomposed factors matrices. Additionally, Laplacian regularization is employed to preserve the spatial structure and promote local smoothness in the fused image. This approach enables the model to jointly capture spectral variation and spatial consistency. Experimental results on several benchmark datasets demonstrate that the proposed method outperforms existing techniques in terms of both spectral fidelity and spatial detail preservation.
Diffusion models have shown impressive performance in imaging inverse problems but typically require retraining for specific degradations, which limits their adaptability. To address this issue, we propose a zero-shot framework capable of solving diverse tasks without any model retraining. Our method introduces a likelihood-guided noise refinement strategy that approximates the likelihood score in closed form to refine the predicted noise, aligning the restoration with the generative mechanism of diffusion models. Combined with DDIM sampling for efficient inference, our approach achieves competitive performance on image super-resolution and denoising while maintaining high computational efficiency.
Deep learning has emerged as a prevalent approach for hyperspectral unmixing. However, most existing unmixing methods employ a single network, resulting in moderate estimation errors and less meaningful endmembers and abundances. To address this imitation, this paper proposes a novel double autoencoders-based unmixing method, consisting of an endmember extraction network and an abundance estimation network. In the endmember network, to improve the spectral discrimination, a logarithm spectral angle distance (SAD), integrated with anomaly-guided weight, is developed as the loss function. Specifically, the logarithm function is used to boost the reliability of a pixel based on its high SAD similarity to other pixels. Moreover, the anomaly-guided weight mitigates the influence of outliers. As for the abundance network, a spectral convolutional autoencoder combined with the channel attention module is employed to exploit the spectral features. Additionally, the decoder weight is shared between the two networks to reduce computational complexity. Extensive comparative experiments with state-of-the-art unmixing methods demonstrate that the proposed method achieves superior performance in both endmember extraction and abundance estimation.
Zero-shot image super-resolution methods have attracted considerable attention due to their high potential for applications with few additional training data. However, these methods often encounter challenges such as inconsistent results and blurred details. To mitigate these issues, we propose WGID: Wavelet-Guided Iterative Detail Enhancement Diffusion Models for single-image super-resolution. In this method, diffusion iteration is guided by wavelet transform to enhance details across different scales while maintaining the similarity between the reference and diffusion-generated images. In addition, the reference images are dynamically updated to provide a suitable guidance during the diffusion process. Meanwhile, the diffusion model progressively refines details, suppresses noise and preserves the natural appearance of the generated image. By integrating these two techniques, the proposed method produces reconstructed super-resolution image with enhanced visual quality, clear details, and more realistic results, accompanied by improved assessment metrics.
Blind spot network (BSN) has gained increasing attention with its state-of-the-art performance in self-supervised image denoising. However, most existing BSN models are based on an unrealistic assumption of noise independence and use isotropic mask convolutions, which can lead to the loss of structural details in the denoised image. To address these limitations, we consider the spatially correlated noise and introduce directional adaptive downsampling and mask convolutions to the wavelet domain, resulting in a novel self-supervised denoising method called wavelet-adaptive BSN (WA-BSN). Specifically, we design the direction-adaptive pixel-shuffle downsamplings (PDs) and apply them to the wavelet decomposition subbands, where the spatial-correlated noise is eliminated and the inherent structure is well preserved in the wavelet domain. Then, based on the geometric direction of the wavelet subimages, we propose four shape-adaptive mask convolutions of a smaller size for each wavelet subband in WA-BSN. This enables adaptive pixel prediction within a structural neighborhood for each subband with reduced training time. Finally, total variation (TV) is added to the loss function to further preserve the edges. The results on public real-world datasets demonstrate that our method significantly outperforms existing self-supervised denoising methods and achieves great efficiency.
Diffusion models have achieved remarkable success in imaging inverse problems owing to their powerful generative capabilities. However, existing approaches typically rely on models trained for specific degradation types, limiting their generalizability to various degradation scenarios. To address this limitation, we propose a zero-shot framework capable of handling various imaging inverse problems without model retraining. We introduce a likelihood-guided noise refinement mechanism that derives a closed-form approximation of the likelihood score, simplifying score estimation and avoiding expensive gradient computations. This estimated score is subsequently utilized to refine the model-predicted noise, thereby better aligning the restoration process with the generative framework of diffusion models. In addition, we integrate the Denoising Diffusion Implicit Models (DDIM) sampling strategy to further improve inference efficiency. The proposed mechanism can be applied to both optimization-based and sampling-based schemes, providing an effective and flexible zero-shot solution for imaging inverse problems. Extensive experiments demonstrate that our method achieves superior performance across multiple inverse problems, particularly in compressive sensing, delivering high-quality reconstructions even at an extremely low sampling rate (5
Blind spot network (BSN) is an effective method for self-supervised image denoising, but real-word noises violate its basic assumption of pixel-wise independent noise. Furthermore, the existing CNN-based BSNs use a uniform convolutional kernel to predict the masked pixels across different patches, which limit the denoising performance. Based on this, we propose a novel blind spot network called dynamic adaptive blind spot network (DA-BSN) that can be self-adjustable according to different spatial geometric structures of the image. Specifically, we design a dynamic adaptive blind convolution (DAB-Conv) block that can accurately predict clean pixels without introducing noise from irrelevant pixels by learning the relationship between masked pixels and surrounding pixels of different image structure. The results on public real-world datasets demonstrate that our method significantly outperforms existing self-supervised denoising methods and achieves great efficiency.
Over the past few decades, researchers have proposed various hyperspectral unmixing (HU) methods. Among these methods, deep learning (DL) has emerged as a promising approach for HU, providing new opportunities for advancement. However, accurately quantifying the presence of spectral variability factors within a mixture remains a challenging task. Therefore, numerous literatures have concerned the HU with spectral variability, in which the variation spectra are generated through the network. However, there is a lack of the connection between the network and spectral variability, so they fail to provide physically meaningful interpretability of spectral variability. To this end, we use physics-driven model to represent spectral variability and introduce it to the two-stream autoencoder unmixing network, resulting in the improved endmember and abundance estimations. Specifically, the endmember extraction network learn spectral variability parameters associated the dispersion model to generate the variations of spectra, which enhancing physical interpretability of endmember variability. In addition, the abundance estimation autoencoder network, tied to the endmember extraction network by shared weights, estimates abundances using the reconstructed hyperspectral image. Compared with the state-of-the-art HU approaches on three real hyperspectral image datasets, our method outperforms these techniques with improved unmixing accuracy, especially on endmember estimation.
Spiking Neural Network (SNN) has been recognized as the third generation of neural networks. Conventionally, a SNN can be converted from a pre-trained Artificial Neural Network (ANN) with less computation and memory than training from scratch. But, these converted SNNs are vulnerable to adversarial attacks. Numerical experiments demonstrate that the SNN trained by optimizing the loss function will be more adversarial robust, but the theoretical analysis for the mechanism of robustness is lacking. In this paper, we provide a theoretical explanation by analyzing the expected risk function. Starting by modeling the stochastic process introduced by the Poisson encoder, we prove that there is a positive semidefinite regularizer. Perhaps surprisingly, this regularizer can make the gradients of the output with respect to input closer to zero, thus resulting in inherent robustness against adversarial attacks. Extensive experiments on the CIFAR10 and CIFAR100 datasets support our point of view. For example, we find that the sum of squares of the gradients of the converted SNNs is 13∼160 times that of the trained SNNs. And, the smaller the sum of the squares of the gradients, the smaller the degradation of accuracy under adversarial attack.
Low-light image enhancement is a necessary preprocessing step for target detection and recognition in low-light environment and urgently needed in night vision monitoring, medical imaging, remote sensing imaging and other fields. With the rapid development of machine learning research, machine learning based low-light image enhancement has attracted extensive attention and achieved good results. However, most of the existing machine learning based low-light image enhancement methods rely on the "bright-dark" paired datasets. On the one hand, the construction of the paired dataset has a high cost, which is not conducive to the promotion and practical application. On the other hand, in practical problems, we can usually get partly-paired images with similar background, and there are a lot of shared information between these images. Taking full advantage of this shared information is also conducive to further improve the efficiency of learning methods. This paper focuses on the low-light image enhancement model based on Retinex theory. By mining and modeling the shared prior between partly-paired images of the same scene, and coupling with the existing machine learning methods based on paired dataset training, a Retinex model for partly-paired low-light image enhancement method with learned prior is proposed. Experiments demonstrate that the proposed method can recover more details and richer colors in visual effects, and can improve the numerical results by up to 20%.
For unmixing (UN) of the sequence of hyperspectral images (SHS), spectral variability is an important factor to be considered. However, most existing UN methods tend to model the endmember and its variability in the spatial domain rather than the transform domain. In fact, the intrinsic and invariant features of the spectral curve can be effectively represented by wavelet transform. Therefore, this article proposes to perform SHS UN in the wavelet domain by combining the Bayesian method. First, the assumption of abundance being invariability in both the spatial and wavelet domains is made, and then, the formulation of UN in the wavelet domain using the perturbed linear mixing model (PLMM) is presented. Second, based on the Bayesian framework, the likelihood and prior are both given, in which the parameter priors are divided into two parts: low- and high-frequency wavelet coefficients. Moreover, by considering the sparsity of the high-frequency wavelet coefficients of endmembers, a noninformative prior with zero mean is designed. Meanwhile, for the coefficients of endmember variability, Gaussian distributions are utilized to represent the steady fluctuation along the temporal dimension. Finally, using the maximum a posteriori (MAP) rule, a hierarchical spectral variability UN model in the wavelet domain is built and solved by the Markov chain Monte Carlo (MCMC) sampling algorithm. Numerical experiments show that the proposed method generates more accurate estimates for endmembers and their variation.
The main changellage of multitemporal hyperspectral image (MTHS) unmixing is modeling the variation of the spectra. In this paper, by considering the spectral variability as an additional term of the endmember, a hierarchical bayesian MTHS unmixing model is established via multi-scaless wavelet analysis, in which the prior of the endmember and its variability is the core concern. The low frequency wavelet coefficients of the endmember and its variability are both designed according to their non-informative propotites. Furthermore, the interscale correlationship of high frequency coefficients of the endmember and its variability is fully exploited, resulting in the scale-dependent prior. In addition, the joint posterior distribution is inferred and solved by a Markov chain Monte Carlo (MCMC) sampling. Experimental results show that the propose unmixing method can accurately capture the variation of the spectra and estimate of the abundance.
Hyperspectral image (HSI) super-resolution aims at improving the spatial resolution of HSI by fusing a high spatial resolution multispectral image (MSI). To preserve local submanifold structures in HSI super-resolution, a novel superpixel graph-based super-resolution method is proposed. Firstly, the MSI is segmented into superpixel blocks to form two-directional feature tensors, then two graphs are created using spectral–spatial distance between the unfolded feature tensors. Secondly, two graph Laplacian terms involving underlying BTD factors of high-resolution HSI are developed, which ensures the inheritance of the spatial geometric structures. Finally, by incorporating graph Laplacian priors with the coupled BTD degradation model, a HSI super-resolution model is established. Experimental results demonstrate that the proposed method achieves better fused results compared with other advanced super-resolution methods, especially on the improvement of the spatial structure.
In this paper, we propose a low rank regularized hyperspectral image (HSI) super-resolution method based on transform domain tensor. The spectral dimension subspace is learned from low spatial resolution hyperspectral image (LR-HSI) by singular value decomposition. Then representation coefficients of the subspace are grouped as several low rank tensors that are constrained in a data adaptive tensor transform domain, resulting in a low rank regularized super-resolution model. Numerical experiments show that compared with other methods, the proposed super-resolution method shows outstanding performance on enhancing the spatial structures.