
The detection of anomalies within video surveillance becomes a key issue in terms of public safety due to the rapid development of surveillance camera deployment. Hence, the urgent need arises for automatic anomaly detection techniques for identifying unusual behaviours such as accidents, aggressive situations, and fires. Simultaneously, the increasing complexity of video analysis and diversity of behavioural characteristics place severe limitations on the effectiveness of the existing methods of anomaly detection, which usually suffer from low performance due to poor spatial-temporal feature representation. To resolve these problems, an innovative deep learning (DL)-based approach is presented. Initially, the input frames are preprocessed using standard methods like image resizing and normalisation to ensure uniformity and improved learning efficiency. For spatial - temporal feature extraction, a 3D Convolutional Neural Network (3DCNN) is utilised to capture both spatial and short-term motion information. Subsequently, a Bidirectional Long Short-Term Memory (BiLSTM) network integrated with an additive attention mechanism (AMM) is utilised to obtain long-range temporal dependencies and emphasise significant features across the sequence. Finally, a novel GhostBottleneck-ConvNeXt (GBConvNeXt) architecture is introduced for efficient and accurate classification of anomalous events. The proposed approach outperforms existing methods in accuracy, recall, precision, specificity, sensitivity, F-measure, FPR, and FNR.
Rice, a staple food for more than half of the global population, faces significant yield and quality losses due to various diseases and environmental stresses. This paper presents a new Simplicial Finite-Element-Informed Neural Network based on the Musical Chairs Optimisation Algorithm (SFEINN-MCOA) to detect rice leaf disease precisely and efficiently. The first step involves the acquisition of the RGB image from the Rice Leaf Disease dataset. The Iterative Robust Peak-Aware Guided Filter (IRPAGF) is used to remove noise and enhance contrast, thus improving the image. The Graph-Based Soft-Balanced Fuzzy Clustering (GSBFC) method is used to separate diseased regions and then analyse them. The Self-Distillated Masked Autoencoder (SMA) is used to perform feature extraction and capture important attributes of leaves. The SFEINN classifies data under the category of healthy and diseased leaves, and the MCOA optimises the model parameters to achieve maximum accuracy and minimum error. Experimental findings reveal that the SFEINN-MCOA model has an accuracy of 99.9% and an F1-score of 98.9%, which is better and stronger. This smart system offers a secure, automatic, and effective system for early disease identification of rice and helps farmers to enhance crop health and yield sustainability.
Gurugram, a rapidly urbanising satellite city in India's National Capital Region, exemplifies the environmental challenges associated with urban expansion . This study quantifies urbanisation in Gurugram district from 2021 to 2024 and assesses its environmental impacts, focusing on land surface temperature (LST), gas emissions, night-time light intensity, and air quality. Utilising Sentinel-2 imagery, ten binary land use land cover (LULC) maps (built/non-built) were generated using the Google Earth Engine platform, achieving a classification accuracy of 98.20% for 2024 using a Random Forest classifier with spectral bands, 13 optimised indices, texture, and topographic features. The built-up area increased by 18.6% (217.59 km2 to 258.13 km2) between 2021 and 2024, primarily in Cyber City, Sector-51, and along the Delhi-Gurugram Expressway. Environmental analyses integrating Sentinel-5P, VIIRS, Landsat-8, and ground station data revealed a 3.12 degrees C rise in mean land surface temperature (35.05 degrees C to 38.17 degrees C), a 53.8% surge in night-time light intensity (9.15 to 14.07 nW/cm2 /sr), and rising satellite-derived gas emissions including CO (+3.69%) and NO2 (+5.97%). Ground station measurements showed a more pronounced NO2 increase of 57.98% (11.39 to 17.99 & micro;g/m3), while PM2.5 declined by 6.19% (111.17 to 104.28 & micro;g/m3), suggesting possible effects of dust control measures. These findings highlight intensified urban heat island effects and persistent air quality challenges, underscoring the urgent need for integrated sustainable urban planning in Gurugram and comparable rapidly developing cities in the Global South.
Rice leaf diseases significantly reduce crop yield and quality, making early and accurate diagnosis essential for sustainable agriculture. Although deep learning (DL) methods have improved disease identification, precise multiclass classification remains challenging because of image noise, complex textures and varying environmental conditions. To overcome these limitations, this study proposes a Bayesian Multi-Scale Edge-Guided Quantised Network optimised by the Tactical Unit Algorithm (BMEGQN-TUA) for effective rice leaf disease detection. Initially, Fast Gradient Domain Guided Image Filtering (FGDGIF) is applied to remove noise from Rice Leafs and Rice Leaf Diseases datasets. Then, a Maximum-Entropy Regularised Decision Transformer-based U-Net (MERDTB-U-Net) performs accurate segmentation of colour, region and edge information. Important colour, texture, shape and edge features are selected using the Black-Winged Kite Algorithm (BWKA). The proposed BMEGQN integrates Multi-Scale Edge-Guided Attention Network (MSEGAN) with Bayesian Asymmetric Quantised Neural Network (BAQNN) to classify six disease categories: Brown Spot, Healthy, Hispa, Leaf Blast, Bacterial Leaf Blight and Leaf Smut. Finally, the Tactical Unit Algorithm optimises network parameters for improved convergence and classification efficiency. Experimental results demonstrate superior performance with 99.8% accuracy, 99.5% precision, 99.7% recall, 99.8% specificity, 99.8% F1-score and 0.98 MCC, confirming the robustness and reliability of the proposed framework.
High-intensity mining activities in mining areas cause severe disturbances to native vegetation. Native vegetation plays an irreplaceable role in maintaining biodiversity, controlling soil erosion and regulating regional microclimates, and its restoration status is directly related to the environmentally sustainable development of ecologically sensitive areas. At present, accurate assessment of native vegetation in complex mining scenarios still faces challenges. This study constructs an intelligent evaluation framework integrating multi-source remote sensing and deep learning, and proposes the Multi-source Remote Sensing Fusion Deep Attention Convolutional Network (MSR-DACN). Based on optical remote sensing images and synthetic aperture radar (SAR) data, this network achieves accurate identification and coverage estimation for ground objects with complex backgrounds and multi-scales. To improve the ability to acquire multi-scale ground object features and efficiently fuse multi-source information from optical and SAR images, a deep attention mechanism is added to MSR-DACN. The results show that the suggested model's overall accuracy of 0.881 is 5.51% greater than that of the best multi-context attention network. In the meantime, the F1-score rises by 7.41% and the recall rate rises by 8.39%. The coverage estimation results based on this model show that from 2015 to 2024, the high-coverage areas in the Qilian Mountains mining area expanded significantly, reflecting the positive promotion effect of the restoration and governance measures implemented in recent years on ecological reconstruction. The interpretable analysis shows that the fusion feature has the highest contribution value, the optical feature has a strong advantage in vegetation spectral expression, and the SAR feature plays an auxiliary role in surface structure and scattering information supplement. This study improves the intelligent analysis capability of remote sensing data in complex ecological scenarios and can provide strong technical support for mining area ecological restoration monitoring and policy formulation.
In response to the problems of insufficient multi-scale feature extraction, difficulty in preserving edge details, and insufficient road connectivity in remote sensing image road segmentation, a multi-scale feature refinement network is proposed. The network adopts a two-stage design architecture. In the first stage, a multi-scale feature enhancement module is combined with a window transformer encoder, a global information enhancement module, and an attention enhancement module to effectively express multi-scale road targets. Finally, the complete segmentation result is obtained through weighted fusion of spatial and channel alignment. The results showed that the average intersection to union ratio index reached 92.1%, 89.7%, 88.3%, and 88.9% on the Massachusetts Road Dataset, DeepGlobe Road Extraction Dataset, Space Net-3 Dataset, and Urban Landscape Dataset, respectively. The boundary score index reached 0.91, 0.89, 0.88, and 0.88, respectively. The median road connectivity index reached 0.89, the number of road fractures decreased to 3.2 per kilometre, and the road accessibility rate reached 92.6%. Therefore, the multi-scale feature refinement network effectively solves the key technical problems in remote sensing image road segmentation through two-stage progressive optimisation, providing an effective technical solution for road information extraction in intelligent transportation systems.
It is considered impossible to obtain an elevation model from a single optical image; however, if the human brain has the ability to understand the topography of an area by looking at an image, this understanding can be created in machines. In this situation, 3D shapes have to be recovered from 2D images. This ability can be achieved through the combination of Geographic Information System (GIS), Remote Sensing (RS) and Artificial Intelligence (AI). AI can introduce human brain analysis into the machines. In this article, AI is involved with satellite imagery to extract a 3D model from a single image. A new method is developed based on fuzzy logic. Then, some satellite indexes are used to extract geographical objects, then elevation is extracted based on the implementation of fuzzy logic. Finally, the model is tested on Landsat 8 image over the Yazd area, in Iran. This paper provides a detailed description of the model, as well as evidence to show that the elevation model can be extracted from a single image with a complex surface. The obtained results are way better than the primary expectation. Comparing the obtained results and actual observation shows a correlation of 77% between them.
The global population increase, coupled with the effects of climate change, poses an imminent threat to food security. Changes in agricultural practices are necessary to meet the growing food demand while minimising negative impacts on water and soil. To make informed decisions, policymakers need timely, accurate, and efficient crop maps. Spectral similarities between crops are a major limiting factor in crop map accuracy. Crops that have similar spectral responses are difficult to distinguish using traditional multispectral image analysis. The aim of this study is to develop a system that is capable of accurately identifying crop maps regardless of the similarities that might exist between classes. To achieve this, we propose a hierarchical approach that uses NDVI time-series and ground truth data combined with machine learning and deep learning models. For each leaf of the hierarchical tree, a separate model is trained. A case study of the proposed system in the Gharb region of Morocco showed that hierarchical crop mapping reached an F1-measure of 0.90, outperforming commonly used flat classifiers by 12-17%.
Synthetic Aperture Radar (SAR) images are inherently affected by speckle noise, which degrades image quality and complicates subsequent analysis. In this study, a novel metric-driven pixel-level fusion framework is proposed that combines multiple despeckling algorithms to generate a single high-quality despeckled SAR image. For each despeckled result, quality maps are computed using both full-reference metrics, such as Structural Similarity Index (SSIM), Peak Signal-to-Noise Ratio (PSNR) and Universal Quality Index (UQI), as well as no-reference metrics, including entropy, Edge Preservation Index (EPI) and Mean of Ratio (MoR). To address the absence of clean ground-truth images in real SAR applications, the revised framework also incorporates pseudo-reference image generation strategies inspired by recent encoder-decoder and optical-image-assisted SAR restoration approaches. At each pixel location, the algorithm selects the pixel value corresponding to the optimal local quality score, thereby constructing a fused image with enhanced structural preservation and speckle suppression. Experimental results obtained from both simulated and real SAR datasets demonstrate that the proposed method consistently outperforms individual despeckling algorithms in terms of perceptual and statistical quality measures. In addition, the integration of recent state-of-the-art despeckling methods further improves the robustness and flexibility of the proposed fusion framework for practical remote sensing applications.
The effective management of biotic stresses contributes to the agricultural resilience and sustainability. In this respect, the proposed system introduces a hybrid fusion of textural features with deep features derived from a Vision Transformer to enhance the characterisation of the disease severity. Texture features are extracted from the grey level co-occurrence matrix, which is computed on leaf images having undergone a prior Gabor filtering. The classification stage is achieved by machine learning techniques including a multi-layer perceptron and a support vector machine. Additionally, we propose an algorithm to assign labels of severity levels for unlabelled datasets. This algorithm utilises a U-Net architecture to extract masks of diseased regions, where the ratio with the total leaf area allows the severity estimation. Experiments were conducted on datasets including images of pear leaves and wheat spikes labelled into multiple severity levels as well as unlabelled tomato leaves affected by late blight. The proposed system outperforms various deep learning models by more than 3% on the overall classification accuracy and seems to be well-suited for multi-species disease severity classification tasks. Also, it offers an accuracy of 94.2%, 91.8% and 91.11% on the tomato, pear, and wheat spike datasets, exceeding all state-of-the-art results.
The rapid growth of the low-altitude economy has made Unmanned Aerial Vehicle (UAV) scene perception critical for applications like environmental monitoring, but fog degrades image quality-requiring simultaneous dehazing (to restore details) and semantic segmentation (for scene understanding). These tasks conflict: dehazing prioritises high-frequency details, while segmentation relies on low-frequency semantics, and traditional joint frameworks lack dedicated feature control, under-performing in foggy scenes. Conventional methods struggle with frequency issues or fog generalisation, and UAV-specific hazy segmentation datasets are scarce. Thus, we propose Dehaze-Seg, a frequency-aware multi-task framework for foggy UAV perception including: 1) the Local Fusion Module (LFM) uses Haar Wavelet to balance detail restoration and semantics; 2) the Interactive Fusion Module (IFM) corrects mid-level feature distortion; 3) the Semantic Enhance Module (SEM) uses DCT to disentangle fog and deep semantics. We also build two UAV-view hazy segmentation datasets called HazyUAVid, HazyUDD. Experiments show Dehaze-Seg outperforms state-of-the-art methods on dehazing (PSNR, SSIM) and segmentation (mIoU, F-1-score) metrics, with strong generalisation to ground-view foggy scenes. Codes and datasets are available in blue at https://github. com/aresdrw/dehaze-seg.
Satellite image fusion is a vital role in enhancing spatial and spectral information in remote sensing applications for accurate Land Use and Land Cover (LULC) classification. However, existing fusion models often fail to preserve structural details and spectral consistency that leads to low classification accuracy. To address these challenges, a novel SIFMAF-NET is proposed for high-quality satellite image fusion and classification. Adaptive Guided Gaussian Filter (AGGF) is utilised to effectively suppress noise and preserve edges in both multispectral (MS) and panchromatic (PAN) images. The proposed CaMo-Net integrates Capsule Network and MobileNet to extract deep hierarchical and orientation-invariant spatial features with improved computational efficiency. Multi-Attention Fusion Network (MAF-Net) utilises dual-attention mechanisms, namely Channel Attention (CA) and Spatial Attention (SA) to fuse spatial-rich and spectral-rich features while minimising redundancy. Fully connected Layer (FCL) is used to classify LULC classes such as vegetation, water, cropland, rangeland, bare ground and built-up areas. The SIFMAF-NET attains an overall accuracy of 98.30%, F1 score of 96.81%, PSNR of 31.75 dB, SSIM of 0.963 and MSE of 0.0031 that ensures spatial and spectral reliability. The SIFMAF-NET improves the total accuracy by 6.30%, 4.25%, 1.13%, 0.45% and 0.75% compared to TextFusion, RS-IRSAN, Fusion+, STFMamba and SwinMFF, respectively.
This study addresses the critical task of automatically identifying oceanic eddies, essential features for marine energy and chemical distribution, using sea surface temperature data from the Atlantic Ocean. It introduces the Deep Eddy Network, a sophisticated deep-learning framework based on an encoder-decoder architecture. The network performs pixel-wise classification, generating an output map where each pixel is labeled as '0' (non-eddy), '1' (anticyclonic eddy), or '2' (cyclonic eddy). Key innovations include a dedicated morphological module that injects shapebased information into the input data. The architecture is designed for high efficiency, employing advanced techniques in its core components. The encoder block utilizes dilated convolutions combined with activation functions, batch normalization, and an attention mechanism. Similarly, the decoder block integrates activation functions with 2D transpose convolution, batch normalization, and attention. Developed using Python and Keras, the final model demonstrates a superior balance between computational performance and segmentation accuracy. This makes the proposed Deep Eddy Network a practical and powerful tool, particularly suitable for deployment in real-time oceanographic monitoring and analysis applications, advancing our ability to understand these dynamic oceanic phenomena.
Hyperspectral image classification faces multiple challenges in handling complex spectral and spatial features. Existing methods typically focus on either local or global feature extraction, often neglecting the synergy and the underutilisation of frequency domain features. Additionally, these methods show reduced accuracy and robustness with limited training samples. To address these issues, this paper proposes the graph attention and frequency aware siamese network, GATSiamNet, a dual branch network for few-shot hyperspectral image classification. The method extracts global features from a superpixel-based graph structure, while the other branch, based on a convolutional neural network, captures local details by integrating spectral-spatial and frequency domain features. This architecture not only captures multi-dimensional spectral-spatial and frequency domain but also enhances classification performance through an effective feature fusion strategy. Experiments on four benchmark hyperspectral datasets show that the proposed method outperforms state-of-the-art approaches, particularly under limited training samples, demonstrating strong generalisation. With five labelled samples per class, the overall accuracy on the Pavia University dataset is 91.04%, and 82.67% on the University of Houston dataset. Furthermore, ablation studies validate the critical role of frequency domain information, along with spectral and spatial feature fusion, in improving classification performance, proving the effectiveness of GATSiamNet in capturing multi-dimensional features.
This study evaluates the geospatial accuracy of 3D Gaussian Splatting for rapid 3D reconstruction from UAV video integrated with GNSS metadata synchronisation. The reconstruction quality was assessed in terms of relative geometric consistency and absolute geospatial accuracy, using a reference model generated through close-range photogrammetry (CRP) as a baseline. The experiment used 54 extracted video frames and nine geospatial checkpoints to evaluate both geometric and geospatial accuracy. Experimental results show that GS achieves millimetre-level structural consistency, with an ICP RMS of 0.0048 m, which is comparable to the photogrammetric reconstruction accuracy reported in similar UAV mapping studies. In terms of absolute positioning, the GS model demonstrates moderate horizontal accuracy (CE90 = 5.26 m) and strong vertical accuracy (LE90 = 0.56 m) when reconstructed directly from video frames enriched with GNSS metadata. Comparative analysis indicates that photogrammetric workflows provide superior absolute geospatial control when Ground Control Points (GCPs) are available. GS enables faster reconstruction and stable geometric representation, supporting rapid mapping applications. However, horizontal positional drift remains a key limitation, mainly caused by temporal misalignment between video frames and GNSS metadata rather than reconstruction instability. With improved synchronisation or external geospatial constraints, GS shows strong potential for UAV-based rapid 3D mapping workflows.
Although building extraction is vital for urban planning, contemporary deep learning frameworks frequently encounter a trade-off between model complexity and extraction accuracy. To solve this limitation, a Lite Fusion UNet (LF-UNet) is proposed, featuring an encoder - decoder architecture that leverages a hybrid encoder. This encoder combines shallow depthwise separable convolutions with deep, window-based self-attention mechanisms. Furthermore, the architecture integrates a Multi-scale Feature Adaptive Fusion (MFAF) module to bolster multi-scale spatial awareness. In addition, a 'Lite Fusion' strategy that stacks dual-layer convolutions with coordinate attention is introduced to rectify deficiencies in fine-grained detail recovery. Experimental results indicate that LF-UNet achieves superior performance, with F1-scores of 92.56% on the WHU Aerial Imagery dataset and 84.63% on the GF-7 Building dataset, outperforming existing lightweight models. By effectively balancing efficiency and segmentation accuracy, LF-UNet offers a robust solution for large-scale, real-time building extraction. The code and models of our LF-UNet model are available at: https://github.com/troYeBlueQ/LF-UNet.