In the present article, an improved Knowledge Distillation (KD) framework has been proposed for efficient compression of deep convolutional neural networks for land-use image classification task. Motivated by the need to achieve competitive classification accuracy while reducing computational complexity, a teacher-student learning paradigm is adopted in which a VGG16 network transfers knowledge to a lightweight MobileNetV2 model. The proposed framework integrates hard supervision from ground truth labels with a soft supervision strategy that combines Kullback-Leibler divergence and Cosine Similarity losses. Experiments conducted on three land-use datasets show that the proposed KD-based method yields improved performance, and achieves an accuracy of 99.04%, outperforming both baseline student training and single-loss distillation approaches, while retaining substantial model compression.
This article presents DMFNet, a dual-backbone multiscale feature fusion framework with residual feature propagation and spatial attention for remote sensing scene classification. Existing approaches often face challenges in effectively capturing multiscale feature interactions and learning robust feature representations from complex aerial scenes with high intra-class variability and inter-class similarity. To address these limitations, the proposed framework employs two pretrained backbone networks to extract diverse hierarchical feature representations. A multiscale feature fusion mechanism with residual feature propagation is introduced to enhance feature interaction across multiple resolution levels. In addition, a spatial attention module is introduced to emphasize informative spatial regions in multi-object scenes. Further, a two-stage training strategy consisting of backbone freezing followed by selective fine-tuning is adopted to ensure stable optimization and improved generalization. Experiments conducted on the benchmark AID dataset demonstrate that the DMFNet achieves an average accuracy of 97.46% ± 0.14%. Ablative analysis further show the importance of various components in unison.
Aerial images play a vital role in urban planning and environmental preservation, as they consist of various structures, representing different types of buildings, forests, mountains, and unoccupied lands. Due to its heterogeneous nature, developing robust models for scene classification remains a challenge. In this study, we conduct a literature review of various machine learning methods for aerial image classification. Our survey covers a range of approaches from handcrafted features (e.g., SIFT, LBP) to traditional CNNs (e.g., VGG, GoogLeNet), and advanced deep hybrid networks. In this connection, we have also designed Aerial-Y-Net, a spatial attention-enhanced CNN with multi-scale feature fusion mechanism, which acts as an attention-based model and helps us to better understand the complexities of aerial images. Evaluated on the AID dataset, our model achieves 91.72% accuracy, outperforming several baseline architectures.
Detecting surface landmines and unexploded ordnances (UXOs) using deep learning has shown promise in humanitarian demining. However, deterministic neural networks can be vulnerable to noisy conditions and adversarial attacks, leading to missed detection or misclassification. This study introduces the idea of uncertainty quantification through Monte Carlo (MC) Dropout, integrated into a fine-tuned ResNet-50 architecture for surface landmine and UXO classification, which was tested on a simulated dataset. Integrating the MC Dropout approach helps quantify epistemic uncertainty, providing an additional metric for prediction reliability, which could be helpful to make more informed decisions in demining operations. Experimental results on clean, adversarially perturbed, and noisy test images demonstrate the model's ability to flag unreliable predictions under challenging conditions. This proof-of-concept study highlights the need for uncertainty quantification in demining, raises awareness about the vulnerability of existing neural networks in demining to adversarial threats, and emphasizes the importance of developing more robust and reliable models for practical applications.
Hyperspectral image (HSI) classification presents a unique challenge due to its inherent high-dimensional 3D voxel structure, where each spatial location is associated with hundreds of contiguous spectral channels. While standalone vision models and language models can be optimized with relative ease on natural image or text tasks, their cross-modal alignment in the hyperspectral domain remains an underexplored problem. In this article, we attempt to align a Vision-Language Model (VLM) for hyperspectral scene understanding, exploiting a CLIP-style contrastive training framework. Our framework maps voxel-level embeddings from a vision transformer onto the latent space of a frozen large embedding model (LEM), where a trainable linear probe aligns vision features with the model's textual token representations. The two modalities are aligned via contrastive loss restricted to a curated set of hard (closest wrong classes) and semi-hard (random distractors) negatives, along with positive pairs, improving both efficiency and discriminability. To enhance alignment, descriptive prompts that encode class semantics are introduced and act as structured anchors for HSI embeddings. It is seen that the proposed method updates only 0.07% of the total parameters, yet yields state-of-the-art performance. For example, on Indian Pines (IP) the model produces better results over unimodal and multimodal baselines by +0.92 OA and $+1.60 \kappa$, while on Pavia University (PU) it provides gains of +0.69 OA and $+0.90 \kappa$. Moreover, this is achieved with the set of parameters, nearly $\mathbf{5 0} \times$ smaller than DCTN and $\mathbf{9 0} \times$ smaller than SS-TMNet.
This work proposes a disentangled feature-based Generative Adversarial Network (GAN) for SAR-to-optical image translation. The generator consists of two encoders and a decoder. During training, one of the encoders extract shared structural features from SAR and optical images, while the other captures domain-specific style information from optical image. These structure and style features are concatenated and fed into the decoder to reconstruct the optical image. The generated and real optical images are compared by a discriminator, which provides a score for the generated image being real or fake. This score is used in the loss calculation. The network is optimized using both contrastive loss, to enforce alignment of content features across domains, and adversarial loss, to achieve realistic optical outputs. After training, the style encoder generates optical style representations that are reduced with PCA to form a representative style vector. During translation, the SAR image is encoded to obtain structural features, which are combined with the fixed style vector and decoded into the optical image. The framework enables unpaired training and style transfer. Evaluation using PSNR, SSIM, and FID confirmed high visual quality and structural consistency in the generated outputs.
This article proposes a novel deep neural network-based method to enhance image classification, focusing on lung cancer detection. Integrating a self-attention mechanism into a pre-trained VGG16 network, the approach combines the robustness of pre-trained convolutional neural networks (CNNs) with the interpretability of self-attention. By using the transformer model's scaled dot-product attention mechanism, attention scores are calculated to effectively focus on critical areas within input images. This increased attention and global relationship enable more precise extraction and representation of features, and addresses nuanced patterns in lung cancer images. Experimental results on the IQ-OTH/NCCD lung cancer dataset show that the proposed model demonstrates high performance in accurately classifying lung cancer, with average precision, recall, and F1-score of 98.25%, 97.96%, and 98.10%, respectively, and an average accuracy of 97.96% (best accuracy 98.64%). The proposed model has also demonstrated an average accuracy of 99.36% (best accuracy 99.54%) on 5-fold cross-validation and is seen to be outperforming 13 other state-of-the-art approaches on similar dataset with only 76,292 parameters. Statistical analysis, including an unpaired t-test and Welch's t-test, confirms the superiority of the proposed method over VGG16, with a two-tailed p value of 0.0319 and 0.0278, respectively. These findings assert the effectiveness of the lightweight model for real-world applications in medical image analysis. This technique allows to combine fine-grained detail extraction with global attention, enabling precise localization and accurate assessment, thus enhancing sensitivity to subtle variations crucial for medical diagnosis.
In this article, we propose a novel approach for plant hierarchical taxonomy classification by posing the problem as an open class problem. It is observed that existing methods for medicinal plant classification often fail to perform hierarchical classification and accurately identifying unknown species, limiting their effectiveness in comprehensive plant taxonomy classification. Thus we address the problem of unknown species classification by assigning it best hierarchical labels. We propose a novel method, which integrates DenseNet121, Multi-Scale Self-Attention (MSSA) and cascaded classifiers for hierarchical classification. The approach systematically categorizes medicinal plants at multiple taxonomic levels, from phylum to species, ensuring detailed and precise classification. Using multi scale space attention, the model captures both local and global contextual information from the images, improving the distinction between similar species and the identification of new ones. It uses attention scores to focus on important features across multiple scales. The proposed method provides a solution for hierarchical classification, showcasing superior performance in identifying both known and unknown species. The model was tested on two state-of-art datasets with and without background artifacts and so that it can be deployed to tackle real word application. We used unknown species for testing our model. For unknown species the model achieved an average accuracy of 83.36
Hyperspectral images provide rich spatial-spectral information, enabling detailed analysis of remotely sensed images. This article presents a novel ultra-lightweight deep neural net model for hyperspectral image (HSI) classification by introducing a patch-based attention mechanism along with a Restricted Isometric Regularization (RIR). The patch attention mechanism is embedded onto a 3D convolutional neural network (CNN) and operates on fixed spatial neighborhoods to dynamically weigh the local spatial-spectral features, thereby enhancing the model’s ability to capture intricate patterns across spectral bands. Due to the presence of a large number of bands in HSI images, preservation of intrinsic structure is a challenge. The present article formulates a novel regularizer, RIR, which facilitates stable feature projection in high-dimensional spaces and effectively addresses the curse of dimensionality while preserving the overall structure of the image. The proposed method achieves 99.79% accuracy on Pavia University dataset, setting a new benchmark in HSI classification. It also performs excellently on other benchmark datasets, e.g. 98.63% OA and 98.44% Kappa on Indian Pines, 99.94% OA and 99.93% Kappa on Salinas, and 99.54% OA and 99.50% Kappa on Botswana. In addition, ablation studies highlight the importance of each mechanism. With only 33,337 parameters, our framework significantly outperforms state-of-the-art methods by a considerable margin.
The present article introduces the Geometric Harmonic Ensemble (GHE) strategy for weather image classification. For varying perspective feature extraction, properties of transfer learned VGG16, Inception V3, and ResNet50 models have been exploited. GHE combines the geometric and harmonic means to balance the models’ consensus and divergence and is seen to achieve a superior classification accuracy of 92.82%, outperforming baseline ensembles on WEAPD dataset. Statistical analysis reveals significant differences favoring GHE across Precision, Recall, F1-Score, and Accuracy, with t-tests confirming the findings at a 0.05 significance level. GHE is also able to consistently outperform other competing methods, including state-of-the-art techniques.
Vision transformers often struggle with sensitivity to spectral-spatial perturbations and inefficiencies in label-scarce regimes for hyperspectral image analysis. To address these, we introduce contextually perturbed diffusion-guided active learning (CPDGAL), integrating diffusion-guided feature refinement (DGFR) and contextualized masking (CM). The DGFR initially injects structured per-turbations into patch embeddings and then reconstructs clean patches via a diffusion-based denoising mechanism. Through this, DGFR refines the learned features while im-plicitly calibrating the aleatoric uncertainty. The CM mechanism applies attention-guided probabilistic masking and enforces context-aware reconstruction to improve generalization. We also introduce a sparse supervision scheme for label-scarce scenarios that selects uncertain samples using confidence-aware ranking, prioritizing challenging data for efficient retraining through active learning (AL). Experiments on benchmark datasets validate the effectiveness of CPDGAL, achieving 97.34% overall accuracy (OA) on Indian Pines, 99.87% on Salinas, and 98.94% on Botswana with a lightweight architecture (0.09M parameters, 5.17 MFLOPs) and outperforms sixteen CNN/transformer-based SOTA methods. Our framework also generalizes better than the vision transformer in extreme low-label settings.
As data requirements continue to grow, efficient learning increasingly depends on the curation and distillation of high-value data rather than brute-force scaling of model sizes. In the case of a hyperspectral image (HSI), the challenge is amplified by the high-dimensional 3D voxel structure, where each spatial location is associated with hundreds of contiguous spectral channels. While vision and language models have been optimized effectively for natural image or text tasks, their cross-modal alignment in the hyperspectral domain remains an open and underexplored problem. In this article, we make an attempt to optimize a Vision-Language Model (VLM) for hyperspectral scene understanding by exploiting a CLIP-style contrastive training framework. Our framework maps voxel-level embeddings from a vision backbone onto the latent space of a frozen large embedding model (LEM), where a trainable probe aligns vision features with the model's textual token representations. The two modalities are aligned via a contrastive loss restricted to a curated set of hard (closest wrong classes) and semi-hard (random distractors) negatives, along with positive pairs. To further enhance alignment, descriptive prompts that encode class semantics are introduced and act as structured anchors for the HSI embeddings. It is seen that the proposed method updates only 0.07 percent of the total parameters, yet yields state-of-the-art performance. For example, on Indian Pines (IP) the model produces better results over unimodal and multimodal baselines by +0.92 Overall Accuracy (OA) and +1.60 Kappa (κ), while on Pavia University (PU) data it provides gains of +0.69 OA and +0.90 κ. Moreover, this is achieved with the set of parameters, nearly 50× smaller than DCTN and 90× smaller than SS-TMNet.
Traditional filter-based methods for unsupervised feature selection often use symmetric dependence measures to detect and eliminate redundant features. Although effective for capturing monotonic relationships, such measures overlook the directionality inherent in many non-linear dependencies. We introduce an asymmetric dependence index, AsymNL, which assesses the relative predictive ability of features, offering a directional dependence measure. Building on AsymNL, we develop a new unsupervised feature selection method that prioritizes the retention of features with strong predictive potential while discarding redundant features. We evaluated the proposed method on four datasets: two hyperspectral remote sensing datasets (Indian Pines and Botswana), an atmospheric radar remote sensing dataset (Ionosphere) and an acoustic signal processing dataset (Sonar). Our method consistently surpasses state-of-the-art unsupervised feature selection techniques, demonstrating its effectiveness.
Hyperspectral images offer rich spectral information but suffer from high dimensionality and redundancy, increasing the computational burden for classification. This paper presents an unsupervised band selection method that leverages Kernel Fuzzy C-Means clustering, guided by a Normalized Mutual Information based dissimilarity measure, to group spectrally similar bands. To ensure robustness against clustering variability, the process is repeated over multiple simulations, and representative medoid bands are consistently selected from each cluster. Finally, these selected bands are ranked based on their selection frequency and average membership strength. Experiments on the Indian Pines and Kennedy Space Center datasets demonstrate the superior performance of the proposed method over several state-of-the-art algorithms.
This article introduces a novel approach to bolster the robustness of Deep Neural Network (DNN) models against adversarial attacks named “Targeted Adversarial Resilience Learning (TARL)”. The initial evaluation of a baseline DNN model reveals a significant accuracy decline when subjected to adversarial examples generated through techniques like FGSM, PGD, Carlini Wagner, and DeepFool attacks. To address this vulnerability, the article proposes an active learning framework, wherein the model iteratively identifies and learns from the most uncertain and misclassified instances. The key components of this approach include uncertainty estimation score in predicting the class of the input sample, selecting challenging samples based on this uncertainty score, labeling these challenging examples and augmenting them into the training set, and thereafter retraining the model with the expanded training set. The iterative active learning process, governed by parameters such as the number of iterations and batch size, demonstrates the potential to systematically enhance the resilience of DNN against adversarial threats. The proposed methodology has been investigated on several popular datasets such as the SARS-CoV-2 CT scan, MNIST, CIFAR-10, and Caltech-101, and demonstrated to be effective. Experiments illustrate that the learning framework improves the adversarial accuracies from 17.4
This article presents a novel approach for training convolutional neural networks (CNNs) using an active learning framework. The core of this methodology lies in iteratively enhancing the training set by incrementally augmenting high-divergent unlabeled samples as labeled data. Initially, a small subset of labeled samples is selected from the dataset for training, while the remaining samples are treated as unlabeled. Three CNN models (with ResNet-18 backbone), each with identical architecture but independent parameter initialization, have been trained separately on this labeled subset of data. While predicting the class of the unlabeled data, diverse data samples are selected from the probability vectors generated as output from each identical model and their ensemble model. Here, the diversity of samples is quantified by the Kullback–Leibler (KL) divergence measure computed using each model’s predictive distribution and the fused predictive distribution of their ensemble version. The fused prediction is derived by averaging the individual models’ probability vectors. At each iteration, samples exhibiting the higher average KL divergence are identified as diverse and are subsequently augmented into the training set. This process involves multiple iterations to ensure the gradual expansion of the training set with high-diversity samples, thereby improving the model’s performance. The effectiveness of this approach is empirically evaluated by assessing the model on the hold-out test set. Experiments have been carried out on CIFAR-10 and CIFAR-100 datasets. Results demonstrate that the proposed approach significantly enhances the model’s generalization capabilities compared to traditional training methods on a fixed dataset. The proposed approach yielded an average accuracy of 90.65
This paper introduces a novel benchmark dataset of Visible and Near-Infrared (VNIR) hyperspectral imagery acquired via an unmanned aerial vehicle (UAV) platform for landmine and unexploded ordnance (UXO) detection research. The dataset was collected over a controlled test field seeded with 143 realistic surrogate landmine and UXO targets, including surface, partially buried, and fully buried configurations. Data acquisition was performed using a Headwall Nano-Hyperspec sensor mounted on a multi-sensor drone platform, flown at an altitude of approximately 20.6 m, capturing 270 contiguous spectral bands spanning 398-1002 nm. Radiometric calibration, orthorectification, and mosaicking were performed followed by reflectance retrieval using a two-point Empirical Line Method (ELM), with reference spectra acquired using an SVC spectroradiometer. Cross-validation against six reference objects yielded RMSE values below 1.0 and SAM values between 1 and 6 degrees in the 400-900 nm range, demonstrating high spectral fidelity. The dataset is released alongside raw radiance cubes, GCP/AeroPoint data, and reference spectra to support reproducible research. This contribution fills a critical gap in open-access UAV-based hyperspectral data for landmine detection and offers a multi-sensor benchmark when combined with previously published drone-based electromagnetic induction (EMI) data from the same test field.
High-dimensional datasets often pose a significant challenge for learning algorithms due to the curse of dimensionality, leading to increased computational complexity, overfitting, and poor generalization. An effective way to address this issue is feature selection, which aims to identify and retain only the most informative features. In problems where label information is not available, unsupervised feature selection becomes crucial. To perform unsupervised feature selection, we can use normalized mutual information to identify features that carry similar information, allowing us to reduce dimensionality without significant loss of relevant information. Here, we present an efficient method for the estimation of normalized mutual information using maximal separation.