
The paper presents an empirical study concerning filters size adjustment in 3D Convolutional Neural Networks (CNN) for hyperspectral data classification. Based on Heisenberg's uncertainty principle, constraints on the depth-dependent relationships between spatial and spectral filter sizes have been defined. The benefits of CNNs satisfying these constraints have been evaluated for raw and transformed input data, in terms of optimization of both classification accuracy and training time. Experimental results show that the filter setting based on Heisenberg principle provides a significant reduction in training time and high classification accuracy. In particular, it can be a viable and feasible alternative to depth reduction in the case of transformed input data.
Multimodal remote sensing land cover classification is a significant challenging task. Current methods predominantly rely on deep semantic segmentation models; however, the lack of sufficient representative data and class imbalance among samples obstacles the application of these models on a large-scale area. To overcome these challenges, this paper proposes a semantic segmentation framework based on diffusion model-based data augmentation, leveraging existing classification maps to address data scarcity. The framework comprises four components: pseudo-label generation, multispectral image translation, segmentation model training, and post-processing. Experimental results on the Yangtze River Economic Belt dataset (MMSeg-YREB) from the 2024 IEEE WHISPERS Data Fusion Competition demonstrate the effectiveness of the proposed approach. Additionally, our method achieved second place in the 2024 IEEE WHISPERS competition.
Hyperspectral imaging in combination with microscopy can increase material discrimination possibilities with respect to regular microscopy imaging. We explore this discrimination potential to assess exposure to particle contamination. We focus on discriminating health relevant particles such as silica, in the respirable size fraction. To do so, multispectral imaging in combination with transmission microscopy is used for particle material identification. We use two compact snapshots near-infrared (NIR) cameras providing 16–25 spectral bands in the 600–960 nm range. The multispectral microscopy system has been tested for discrimination of silica among fourteen different particle materials. The analysis performed shows potential to accurately discriminate the 14 tested particle materials. In addition, the band relevance analysis shows that only a few specific bands are needed to provide accurate discrimination. The multispectral method presented could therefore enable a faster exposure assessment than traditional techniques for occupational exposure estimation.
Hyperspectral data, with its rich material information, offers significant advantages over RGB data, leading to an increased focus on hyperspectral object tracking. However, the limited availability of hyperspectral tracking datasets presents a major challenge, hindering the development of robust deep learning-based tracking algorithms. To address this issue, we propose a novel method called Low-Rank Adaptation-based Hyperspectral Tracking (LA-HST). Initially, we propose a spectral enhancement module with a low intrinsic dimension to enhance spectral details in images. Additionally, we introduce an adapted transformer encoder that effectively extracts hyperspectral features with rich spectral information. During training, we freeze the parameters of the pre-trained network and only train the parameters of the proposed modules, which account for just 0.5% of the total parameters. This approach allows the model to fully leverage the tracking capabilities of the pre-trained network while significantly enhancing training efficiency and improving generalization. Extensive experiments on the VIS, NIR, and RedNIR videos from the HOTC2024 dataset validate the effectiveness of LAHST, demonstrating its ability to deliver robust and efficient tracking performance across diverse spectral bands.
This paper analyses the challenge of using the same autoencoder model for lossy compression on different hyperspectral sensors and its impact on the complex application of camouflaged target detection. We propose an adaptive 1D convolutional autoencoder architecture for lossy hyperspectral data compression with the property of portability to unknown spectral signatures of different sensors. In our experiment, a model is pre-trained on PRISMA satellite data from around the world and fine-tuned with a small amount of target HySpex sensor data. The comparison model with the same architecture is trained from scratch with a larger amount of target sensor data. Compression rates of 4, 8, and 16 are implemented and evaluated. The evaluation discusses the reconstruction accuracy measured by the SAM, PSNR, and SSIM metrics. In addition, the reconstruction error is evaluated using the camouflaged target detection application. We show that fine-tuning the pretrained model with a small amount of data results in higher reconstruction accuracy than building the model from scratch with a larger amount of target data. This observation also applies to the target detection application.
In hyperspectral super-resolution (HSR), multivariate continuous functions can approximate the nonlinear spectrum mapping. This process is typically managed using multilayer perceptrons (MLPs), which operate as black boxes. However, these opaque operations make it difficult to interpret whether the super-resolution results align with the underlying physical mechanisms of imaging. Building on the Kolmogorov-Arnold representation theorem, which states that any complex nonlinear function can be constructed through the simple superposition of multiple linear functions, an interpretable super-resolution KAN network, termed HyperKAN, is proposed. This model leverages a univariate function, parametrized as a spline, to directly model the continuous spectral sequence without additional nonlinearities. HyperKAN is effectively constrained by the spectral function, allowing for intuitive visualization, supervision, and control over each node within the neural network. We conducted experiments on CAVE and Chikusei datasets to evaluate the effectiveness of the proposed model achieves state-of-the-art (SOTA) performance compared to competitive methods.
We combine 2 m resolution airborne (HySpex) and 30 m resolution satellite (EnMAP) hyperspectral data to address the challenge of mixed pixels in satellite imagery. Endmembers are manually selected from HySpex data, and Non-negative Least Squares (NNLS) spectral unmixing is applied to generate high-resolution spectral abundance maps. These maps are then resampled to match EnMAP's spatial resolution and used to predict an endmember library from the EnMAP scene. This predicted library is then used for unmixing the EnMAP data over a broader area. When compared to spectral abundance maps generated from direct endmember selection from EnMAP alone, the unmixing results using the predicted library closely align with the high-resolution output, despite some land cover changes over time. In contrast, the spectral abundance maps from low-resolution endmembers lack detail. We discuss the implications of our approach for improved spatial and temporal mapping.
Uncontrolled potato diseases can cause significant yield loss. UAV-based hyperspectral imaging offers a promising method to comprehensively inspect and identify diseased plants across entire fields. This study explored how dimensionality reduction of UAV hyperspectral imagery can enable disease detection with deep learning. Data was collected with the Headwall Nano line-scan sensor, which captures 270 bands over a 400 to 1000nm spectral range. The data was converted into three-band imagery and fed into the YOLOv5s model, which successfully detected the plants infected with blackleg and Potato Virus Y (PVY). The pre-trained model achieved an average mAP@.50 of 0.85 and an average AP@.50 of 0.73 for blackleg detection, as well as an average mAP@.50 of 0.82 and an average AP@.50 of 0.69 for PVY detection, each calculated over ten independent experiments. The results demonstrated the potential of using UAV-based hyperspectral imagery with deep learning techniques for precision agriculture.
Ensuring food safety necessitates rapid and accurate detection of foodborne pathogens. This study presents a novel framework integrating hyperspectral imaging (HSI) with advanced deep learning (DL) techniques to identify pathogenic Escherichia coli on spinach leaves. Fresh spinach samples were inoculated with Enterotoxigenic E. coli (ETEC), while deionized water served as the control. Hyperspectral images in the 400–1000 nm range were captured and preprocessed to normalize spectral data. Custom DL architectures were developed and applied directly to the raw hyperspectral images, exploiting the full spectral-spatial information. The models' performance was rigorously evaluated using comprehensive metrics. Results demonstrate that combining HSI with DL significantly enhances the detection of pathogenic E. coli, offering high accuracy. This innovative approach provides a rapid, non-destructive method for identifying foodborne pathogens, advancing food safety monitoring in processing and distribution chains.
Hyperspectral image (HSI) super-resolution has seen significant advancements with the emergence of diffusion models, demonstrating impressive results in prior works. However, existing approaches often take up cascaded architectures for feature extraction, often suffering from high computational costs, gradient instability, and training difficulties. This paper introduces a novel Efficient Spectral Attention Block (ESAB) that addresses these challenges by leveraging atrous convolutional blocks and channel attention within a multiscale feature extraction strategy. ESAB enables effective spectral feature extraction while maintaining computational efficiency, leading to improved performance. Furthermore, a unique loss function is proposed to enhance spectral fidelity by comparing the spectral slopes of the ground truth and predicted image pixels. Experimental results demonstrate that the proposed approach achieves great performance in super-resolving HSIs, achieving high spectral fidelity with reduced model complexity.
Hyperspectral pansharpening is aimed to fuse low resolution hyperspectral (HS) and high resolution panchromatic (PAN) images. Classic methods, such as the generalized Laplacian (GLP) pyramid algorithm, have been proven effective to extract spatial details from PAN images. This work is focused on the joint use of GLP and robust regression schemes for the injection phase, which is aimed to insert spatial details into the low-resolution HS image. More in detail, the statistical analysis performed on two HS/PAN datasets (acquired by PRISMA and EO-l platforms, respectively) suggests that the Kolmogorov-Smirnov (KS) distance can be effectively used to perform a band-wise decision between the classic linear regression and the robust one. Finally, numerical tests are conducted, even performing a comparison with a benchmark consisting of some widely used methods.
Remote sensing images are indispensable tools in critical domains such as environmental monitoring, urban planning, and national security. However, the integrity of these data sources can be compromised due to forgery, leading to mis-information and potentially harmful consequences. This paper presents a novel approach for detecting forgery in remote sensing images. Through extensive experimentation on benchmark datasets and the RSICD (Remote Sensing Image Capturing Dataset) [1], augmented with synthetic data created using forgery techniques such as copy-move, resizing, rotation and splicing, we demonstrate the effectiveness and scalability of our approach. With the variations in U-Net model and with split and merge operations our method is able to locate the forged regions of input images efficiently.
Hyperspectral object tracking using snapshot mosaic cameras is emerging as it provides enhanced spectral information alongside spatial data, contributing to a more comprehensive understanding of material properties. Using transformers, which have consistently outperformed convolutional neural networks (CNNs) in learning better feature representations, would be expected to be effective for Hyperspectral object tracking. However, training large transformers necessitates extensive datasets and prolonged training periods. This is particularly critical for complex tasks like object tracking, and the scarcity of large datasets in the hyperspectral domain acts as a bottleneck in achieving the full potential of powerful transformer models. This paper proposes an effective methodology that adapts large pretrained transformer-based foundation models for hyperspectral object tracking. We propose an adaptive, learnable spatial-spectral token fusion module that can be extended to any transformer-based backbone for learning inherent spatial-spectral features in hyperspectral data. Furthermore, our model incorporates a cross-modality training pipeline that facilitates effective learning across hyperspectral datasets collected with different sensor modalities. This enables the extraction of complementary knowledge from additional modalities, whether or not they are present during testing. Our proposed model also achieves superior performance with minimal training iterations.
Estimating nitrogen (N) content of crops is essential for obtaining critical parameters such as nitrogen use efficiency (NUE). However, most of the methods developed for N content estimation are destructive and are time- and labor-intensive. Here, we show the results of a non-invasive method developed based on unmanned aerial vehicle (UAV) multispectral imaging of crops to estimate N content at different growth stages. To do so, multispectral drone images of canola and wheat were collected at three different stages of crop growth in an experimental field trial at AAFC Lethbridge, AB, Canada. Leaf tissue samples were also collected concurrently for seven different treatments of nitrogen applications, each of which was replicated four times. Several machine learning (ML) models were trained and evaluated for estimating plant N-uptake. The results show that multispectral imaging can estimate N content in canola with an RMSE of 0.38-0.59 and R2 of 0.77-0.92, while these numbers are 0.33-0.52 and 0.71-0.89 for wheat. We also show that N-content estimations based on multispectral imagery significantly benefit from incorporating ancillary data, such as treatments and image acquisition date into ML models, reducing RMSE by 5-10%. These results show the potential of UAV-based multispectral imaging in acquiring nitrogen-related parameters such as plant N-uptake and NUE measurements.
Spectral unmixing of high spatial resolution hyperspectral images is challenging due to the need to accurately identify and extract endmembers from complex mixtures of materials within each pixel. To tackle these challenges, we propose a Sparse Coding Inspired Generative Adversarial Network (SC-GAN) that leverages sparse coding principles within a GAN framework to enhance the unmixing process. Traditional spectral unmixing methods often struggle with mixed pixel complexity, high data dimensionality, noise susceptibility, and the requirement for extensive labeled training data. SC-GAN addresses these issues by generating sparse abundance representations using a learned over-complete dictionary of endmember signatures, which improves interpretability and reduces computational complexity. The discriminator guides the generator to produce realistic abundance maps, enhancing noise robustness and reducing dependency on labeled data. Experimental results demonstrate that SC-GAN improves un-mixing accuracy and provides a robust solution for handling complex spectral mixtures in high-resolution hyperspectral imagery.
Domain shift refers to the overall distribution differences between data used for model development and post-deployment. If not addressed, it can typically lead to performance degradation in operational settings. It is especially emphasized in the context of remote sensing, where scenes commonly capture large areas with significant geographical, temporal, and sensor variations. Domain generalization is a type of transfer learning for addressing this issue, often through feature alignment, that on the contrary of domain adaptation, assumes access to neither target domain labels nor to target domain data. In this study, we explore the progressive alignment of an image's spectral bands, instead of handling them collectively and concurrently. Experiments have been conducted with Sentinel-2 multi-spectral images, with six European countries denoting the domains, using various contemporary domain generalization techniques, and it is shown that a gradual alignment of spectral bands leads to consistent performance improvements.
Hyperspectral images (HSIs) provide rich spectral information, but acquiring high-resolution data is costly and challenging, making spectral super-resolution essential. Inspired by the near-linear efficiency of state space models (SSM) in long-sequence tasks, we propose a spectral reconstruction method based on dual cross-scanning and cross-attention mechanisms (SR-CSCA). This novel network integrates Transformer and Mamba architectures, efficiently extracting multi-resolution contextual information and optimizing computational strategies to enhance spectral reconstruction performance. SR-CSCA method captures spectral and spatial details and dependencies by employing an optimized joint spatial-spectral dimensional scanning strategy. Furthermore, the cross-attention mechanism enhances feature fusion and information extraction. The efficacy of our approach is demonstrated through comparisons with multiple benchmark methods for spectral reconstruction. Experimental results indicate that SR-CSCA achieves state-of-the-art performance in spectral reconstruction while maintaining a low computational cost.
Diverse engineering disciplines rely on highly detailed, up-to-date thematic maps for daily decision-making. Over the last decade, researchers have approached the land cover classification using supervised deep learning, requiring many labels per category. Labeling is costly, error-prone, and challenging to scale for the ever-growing remote sensing data. Self-supervised learning emerged to learn feature representations on unlabeled datasets, facilitating, for instance, the resolution of few-shot downstream tasks using prior acquired knowledge through transfer learning. Since highly detailed maps often rely on hyperspectral and LiDAR data, it is necessary to quantify the potential of recent self-supervised learning techniques to learn multimodal representations that facilitate accurate few-shot hyperspectral-LiDAR classifications. The current work occupies that gap and compares the representation learning ability of four modern self-supervised learning strategies. It first implements modality-specific encoders for individually handling hyperspectral and rasterized LiDAR data. It then couples each regarded method's architecture on top of the encoders, building pseudo-Siamese networks whose objectives are specific to each learning strategy. It then implements a multi-level feature fusion to combine learned features at different depth levels. Ultimately, it performs non-parametric classifications using the k-nearest neighbors and the support vector machine classifiers to assign categories to joint features at test time. Experiments show that the SimSiam-based method learned the most discriminative features across the studied datasets, achieving consistent classifications at four labeling levels.
We present a novel active learning method for hyperspectral images based on representation learning in Wasserstein space. We perform regularized Wasserstein dictionary learning in the space of hyperspectral pixels, then leverage the learned barycentric coefficients to embed the high-dimensional spectra into a low-dimensional space. Sampling in the low-dimensional space leads to high-quality labels that propagate accurately to the remaining pixels in the data. Our method achieves a high level of accuracy with very few training labels, suggesting its utility for hyperspectral image classification in the active labeling setting.
To evaluate the remaining lifetime of high voltage metallic towers, continuous inspections are required. These inspections are done today by experts climbing the towers, which is not only risky and time-consuming, but also costly and limited (not all parts of the towers can be reached). Hence, it is very important to develop an alternative inspection method to properly assess the repair these towers needs and extend their lifetimes. Imaging from drones is considered as a safe and fast tool for monitoring the environment and assets. However, using RGB imaging, it is not possible to distinguish between various types of corrosions, some of which can lead to severe material loss and thus compromised structural integrity. Moreover, RGB based image analysis is prone to false positives. Although hyperspectral imaging solves some of those problems, varying illumination conditions outdoor and the complex geometry of the towers bring extra challenges, which are addressed in the methodology developed in this paper. We designed a drone payload integrating both a LiDAR scanner and hyperspectral sensors. Leveraging reference spectra captured during a manual initialization step and extra features such as high dynamic range (HDR) for hyperspectral image acquisition resulted in a faster, less complex, and more reliable corrosion inspection.