
ABSTRACT Convolutional Neural Networks (CNNs) and Transformers have achieved remarkable success in hyperspectral image (HSI) classification. However, CNNs are limited in modelling long‐range spatial dependencies, while the quadratic computational complexity of Transformer restricts their application to high‐dimensional HSI data. Recently, state‐space models (SSMs), particularly Mamba, have shown great potential due to their linear computational complexity and powerful sequence modelling capabilities. However, the original Mamba model is designed for one‐dimensional sequence modelling and struggles to effectively handle the inherent three‐dimensional spatial‐spectral characteristics of HSI. Furthermore, it lacks dedicated considerations for the deep fusion of spectral and spatial information and the integration of global context and local details. To address this issue, we propose a novel Multi‐Scale Spatial‐Spectral Mamba Network (MSS‐MambaNet) for efficient and accurate HSI classification. Specifically, MSS‐MambaNet consists of four key components: First, the Multi‐scale Local Spectral Feature Extraction Module (MSLSFE) captures spatial‐spectral features at different scales through a parallel multi‐branch architecture. Second, the Spatial Mamba Encoder employs an innovative dimensionality reorganization strategy to model global spatial context while preserving spatial structure. Third, the Spectral Mamba Encoder incorporates a bidirectional split‐scan mechanism to comprehensively capture the complex dependencies between spectral bands. Finally, the Global Spatial‐Spectral Feature Fusion Module (GSSFF) deeply integrates the complementary information of spatial and spectral features through multi‐scale dilated convolutions and a detail‐enhanced attention mechanism. MSS‐MambaNet achieves overall accuracies of 97.19%, 94.25%, and 97.38% on the Pavia University, Indian Pines, and Salinas datasets, respectively, outperforming state‐of‐the‐art methods in both classification accuracy and computational efficiency. The code will be available at https://github.com/ZhangSaida/Mss‐MambaNet .
ABSTRACT Drone‐based thermal imaging is increasingly used for wildlife monitoring, but tracking individual animals in aerial footage remains challenging in forested habitats due to low texture, occlusions and dynamic viewpoints. Traditional image‐space tracking methods often fail because motion and appearance cues degrade in thermal data. We present a geo‐referenced tracking framework that projects detections into world coordinates using drone localisation, camera pose and elevation models. Building on Deep OC‐Sort, we introduce (i) a geo‐native tracker operating in metric space and (ii) a hybrid tracker combining pixel‐space tracking with geo‐referenced recovery. Both use physically grounded motion models to improve identity preservation under occlusion. Evaluated on a large airborne thermal dataset (225 videos), our methods achieve the fewest ID switches, reducing identity errors by 17% compared to the best and 92% compared to the worst tested state‐of‐the‐art algorithm—in both cases without custom embeddings, which further improve results. Additionally, geo‐referenced outputs enable ecological analyses such as spatial mapping and movement estimation, demonstrating the advantages of global over image‐based tracking.
Automated plant traits recognition from herbarium images is essential for plant sciences, yet remains challenging because background elements (e.g., textual labels, mounting artifacts, and color charts) can introduce shortcut learning, leading models to rely on spurious non-plant cues rather than plant morphology. This bias degrades both generalization and interpretability. In this paper, we introduce AT-ViT, a dual-branch Vision Transformer that jointly encodes raw herbarium scans and their segmented-derived counterparts via a multi-scale, multi-view cross-attention fusion scheme. AT-ViT further incorporates a mask-guided patch weighting mechanism that amplifies plant-relevant regions and attenuates background-driven features. By learning from the original scans while being guided by segmentation masks through the mask-guided patch reweighting mechanism, the model is encouraged to focus on plant organs and learn plant-centric representations more effectively. Across multiple trait classification tasks (e.g., leaf base shape, thorns), AT-ViT delivers consistent accuracy gains, improves attention localization on plant regions, and exhibits increased robustness under synthetic background perturbations. Specifically, AT-ViT substantially improves spatial attention grounding, boosting plant-region alignment (Avg IoU_p: +15.66 to +18.03 pp) while reducing background overlap (Avg IoU_b: -27.92 to -31.02 pp) relative to CrossViT, and remains markedly more robust to background perturbations, outperforming ResNet101 by up to +32.32 accuracy points and CrossViT by up to +5.07 points under background-noise conditions.
The rapid advancement of large-scale pre-trained models in computer vision and natural language processing has necessitated the development of parameter-efficient fine-tuning (PEFT) techniques. Low-Rank Adaptation (LoRA) has emerged as a leading approach, enabling efficient adaptation by decomposing weight updates into low-rank matrices. However, standard LoRA suffers from inefficiencies due to uniform rank allocation, redundant parameter updates, and random initialisation. This paper introduces Sparse Truncated Low-Rank Adaptation (ST-LoRA), a novel PEFT method that enhances LoRA through three key innovations: (1) SVD-based initialisation, leveraging singular value decomposition of pre-trained weights to provide informed starting points for adaptation matrices. (2) Dynamic rank assignment, allocating layer-specific ranks based on singular value distributions to optimise parameter efficiency. (3) Learnable sparse masks, selectively updating only the most critical connections to reduce computational overhead. We evaluate ST-LoRA across vision (ViT-L/16 on ImageNet-1K, CIFAR-100, and fine-grained datasets) and language (LLaMA-2 7B on GLUE and instruction-following tasks) domains. Our comprehensive experiments include rigorous ablation studies, cross-validation, and statistical significance testing. Our results demonstrate that ST-LoRA consistently outperforms existing PEFT methods while significantly reducing computational costs. On ImageNet-1K, ST-LoRA achieves 84.0% top-1 accuracy ( 0.12% std), surpassing LoRA (r = 16) by 0.4% while using only 26.5% of its parameters and reducing training time by 57%. For language tasks, ST-LoRA approaches full fine-tuning performance with a fraction of the computational cost, reducing training time by up to 45% versus standard LoRA and 87% versus full fine-tuning.
Multimodal object detection with RGB and thermal infrared imagery is essential for reliable perception in complex environments, but existing methods often remain limited by insufficient cross-modal feature fusion. To address this issue, we propose PFDet, a progressive cross-domain feature fusion framework for robust multimodal detection. PFDet progressively aligns semantic cues, interacts cross-modal features, and refines multi-scale representations. Specifically, a Semantic Alignment Guidance (SAG) module establishes a unified semantic reference at the input level to guide subsequent fusion. Unified Cross-Modal Fusion Modules (UCMF) are then deployed across the backbone to enable fine-grained bidirectional feature interaction, while a Context Guide Fusion Module (CGFM) performs context-aware multi-scale refinement in the neck network. Experimental results on the M3FD and FLIR-Aligned datasets show that PFDet achieves state-of-the-art performance, reaching 86.7% mAP@0.5 on M3FD and 86.9% mAP@0.5 on FLIR-Aligned, while maintaining real-time inference speed. Ablation studies further validate the effectiveness of the proposed progressive fusion strategy. This work provides a robust and scalable paradigm for multimodal object detection, with the potential to be extended to other heterogeneous sensing modalities, contributing to reliable perception in complex real-world environments.
Open-vocabulary semantic segmentation aims to recognise arbitrary categories beyond a fixed label set, yet existing vision-language methods often struggle with stable pixel-level alignment and suffer feature drift under domain shifts. We propose OSS-CAEA, a framework that freezes the CLIP image and text encoders, Depth Anything V2 and a visual foundation model (VFM), introducing a few learnable adaptation modules. The image passes through three parallel branches: CLIP extracts semantic representations, Depth Anything V2 provides geometric cues and the VFM captures spatial structures, whereas text is encoded by CLIP and projected to obtain shared textual priors. The collaborative attention module (CAM) fuses multiscale features from Depth Anything V2 and the VFM under GeoText Prompt (GTP) guidance, enhancing geometry-semantic consistency and alleviating unstable responses caused by single-layer proxies and appearance variations. CLIP features and CAM outputs are fused, reshaped and fed into a coarse segmentation head to produce supervised predictions that provide spatial and category priors for refinement. Guided by these priors, the pixel-semantic alignment head (PSAH) narrows the gap between pixel semantics and language descriptions, reducing visual-language discrepancy and improving robustness for unseen categories. Experiments on multiple datasets show OSS-CAEA consistently outperforms existing methods; ablations validate each component.
ABSTRACT License plate recognition (LPR)—also widely referred to as automatic license plate recognition (ALPR)—is a critical component of intelligent transportation systems. Despite substantial accuracy improvements driven by deep learning, reliable LPR deployment in real‐world environments remains challenging due to adverse imaging conditions, domain shifts and practical system constraints. This survey focuses on hybrid deep learning frameworks integrating convolutional and attention‐based models. Distinct from existing surveys, this review emphasises the joint design and evaluation of detection–recognition pipelines analysing how architectural coupling affects robustness and deployability. A unified overview of the LPR pipeline is presented, covering detection, recognition, datasets and deployment considerations. Existing methods are categorised into CNN‐based, transformer‐based and hybrid approaches. For detection, the evolution from convolutional detectors to CNN–transformer hybrids is analysed, highlighting trade‐offs among accuracy, robustness and real‐time performance. For recognition, sequence modelling paradigms from CTC‐based methods to attention‐driven Seq2Seq and transformer architectures are reviewed. In addition, public benchmarks and domain generalisation strategies are examined revealing persistent limitations such as regional bias. A consolidated benchmark comparison across representative architectures is provided to facilitate quantitative assessment. Finally, a holistic evaluation perspective is advocated, and future research directions towards robust and deployable LPR systems are outlined.
ABSTRACT As a core task in image understanding, panoptic segmentation integrates instance‐level segmentation for countable objects, referred to as things, and semantic‐level segmentation for uncountable regions, referred to as stuff, thereby enabling comprehensive pixel‐level parsing of image scenes. The original Panoptic FPN relies on unidirectional top–down feature propagation: its high‐resolution shallow features lack sufficient semantic information, and no adaptive filtering is applied to backbone features. These limitations restrict the segmentation accuracy for small objects and those with overlapping features. To address this issue, this study proposes improvements to the Panoptic FPN panoptic segmentation model, with the goal of enhancing segmentation performance and accuracy. Specifically, the traditional feature pyramid network in the original model is replaced with a path aggregation feature pyramid network, hereafter referred to as PAFPN, which transforms the information flow from unidirectional to bidirectional propagation. Additionally, a channel‐spatial attention mechanism is incorporated after PAFPN completes feature fusion on the aggregated pyramid features, so as to obtain higher‐quality features. Experimental results on the COCO2017 dataset show that the improved Panoptic FPN model achieves a 0.556% increase in Panoptic quality, denoted as PQ, compared with the original model—this verifies the effectiveness of the proposed improvements. Compared with mainstream panoptic segmentation models, the improved model also achieves favorable performance.
ABSTRACT Cross‐domain few‐shot learning (CD‐FSL) addresses the challenge of few‐shot classification under significant distribution shifts between source and novel target domains. The core difficulty lies in bridging the domain gap. Existing methods primarily mitigate this issue from the spatial perspective, overlooking the role of frequency information. Empirical studies reveal that augmenting samples through frequency‐space operations can alleviate domain discrepancies. However, current frequency‐based augmentation methods typically perform a holistic replacement of high‐frequency components, which oversimplifies the process and fails to adequately model complex frequency shortcuts (i.e., the tendency of models to prioritise learning the simplest and most class‐discriminative frequency patterns rather than semantically meaningful features). Inspired by gradient‐based adversarial learning, we propose Early‐Stage Feature Frequency Attack (ESFFA). Our method perturbs low‐frequency components along gradient directions and randomly masks high‐frequency components in shallow‐layer feature maps. This joint operation in the feature space, compared to directly replacing components in raw pixels, more effectively disrupts the model's reliance on frequency shortcuts. This approach compels the model to adapt to dynamically changing frequency characteristics, thereby enhancing cross‐domain generalisation. Experiments on eight target datasets validate its effectiveness.
ABSTRACT With the linear complexity and long‐sequence global modelling capability, Mamba becomes a competitor to Transformer architectures in point clouds analysis. However, the designs of traditional 1D convolutions and reordering strategies, do not match the inherently unordered nature of point clouds, which constrain the performance enhancement. In this work, we rethink the ordering and convolution strategy of the PointMamba, and present a novel architecture named PointMamba++ to more effectively aggregate local structural features and achieve a superior accuracy and computation trade‐off. Specifically, we design a point‐edge convolution to aggregate neighbourhood features of point cloud tokens, which replaces 1D convolution layers in traditional Mamba modules and does not perform convolution by sequence but according to geometric relationships. Furthermore, considering that forcibly ordering point clouds is not conducive to learning local geometric features and easily leads to unstable sequence dependencies, we design a sequence‐independent BiMamba module, which adopts two reverse and parallel scanning paths, to reduce the dependency on sequential scanning of Mamba while enhancing point cloud representation abilities. Extensive experiments show that PointMamba++ surpasses typical convolution‐based and Transformer‐based architectures, and achieves state‐of‐the‐art performance on multiple tasks including shape classification, part segmentation, and semantic segmentation.
ABSTRACT This survey provides a focused review of knowledge distillation (KD) techniques in object detection, a key area in computer vision. We categorise existing approaches into three primary types—feature‐based, relation‐based, and response‐based—each defined by the stage of knowledge transfer and the nature of information distilled. Beyond summarising these core paradigms, we examine their adaptations for emerging challenges such as continual learning, semi‐ and weakly‐supervised detection, and 3D object detection. We further conduct a systematic evaluation of representative methods on the COCO dataset, offering an in‐depth analysis of their strengths, limitations, and suitability across scenarios. A distinctive contribution of this work is its cross‐cutting synthesis of shared principles behind diverse KD strategies, revealing how these can be generalised and extended to new domains. We aim to provide researchers and practitioners with both a consolidated conceptual framework and actionable insights for advancing KD in object detection.
ABSTRACT Prevailing image editing methods heavily rely on user‐provided bounding boxes or pixel‐level masks to ensure visual consistency in nonedited regions. Although some attention‐based approaches eliminate the need for manually annotated input, they often unexpectedly alter nontarget areas due to semantic leakage between objects. Our goal is to address this semantic inconsistency challenge with minimal user input by leveraging the mutual exclusion of scene graph nodes, thereby enhancing both editability and background preservation without additional training costs. To address the challenge of semantic inconsistency, we propose a Scene graph‐based ImaGe editing method with Mutually exclusive Attention manipulation, namely SIGMA, to leverage the inherent semantic mutual exclusion properties between scene graph nodes for attention map distribution manipulation. Specifically, we propose a semantic decoupling module to disentangle the desired and nontarget editing areas. We also introduce a semantic injection module to facilitate both foreground editing and background preservation. We validated the effectiveness of SIGMA on the widely used image editing dataset PIE‐Bench. The experimental results demonstrate that SIGMA significantly outperforms existing approaches without any additional training cost.
ABSTRACT Stroke is a leading cause of long‐term disability, with 80% of survivors experiencing acute upper‐limb impairment. Although vision‐based and wearable sensor technologies have the potential to improve rehabilitation, a thorough analysis of their comparative advantages, technical limitations and clinical readiness is still lacking. This systematic review provides a methodologically rigorous analysis of the peer‐reviewed literature from 2005 to 2025, synthesising and critically evaluating vision‐based and wearable sensor technologies for post‐stroke hand rehabilitation. Following PRISMA guidelines, we searched PubMed, Scopus and Web of Science. We analysed 132 included studies to identify a trend towards deep learning‐based computer vision and hybrid wearable systems. However, quantitative synthesis exposed critical gaps: technical benchmarks (e.g., latency and computational cost) were reported in fewer than 5% of studies, and the median sample size was only 17 participants. Methodological quality was low to moderate, with only 12% of studies being randomised controlled trials. We present a new taxonomy classifying systems by sensing modality and maturity, which reveals a lab‐to‐clinic gap. Although innovation is rapid, a lack of standardised benchmarking hinders clinical translation. We propose a decision‐making framework to guide future research and implementation.
Normalising flows are a flexible class of generative models that provide exact likelihoods and are often trained through maximum likelihood estimation. Recent work suggests that discrete‐step flow models trained in this way can assign undesirably high likelihood to out‐of‐distribution image data, bringing their reliability for applications where likelihoods are important (e.g., outlier detection) into question. Continuous‐time normalising flows trained with the conditional flow matching objective (CFM models) also provide unreliable likelihoods, and we investigate whether training them on various feature representations can lead to more reliable likelihoods. We consider features from a pretrained classifier, features from a pretrained perceptual autoencoder and features from an autoencoder trained from scratch with a simple pixel‐based reconstruction loss, and compare their effects on CFM model likelihoods on various in‐ and out‐of‐distribution sets. Autoencoder‐based features are of particular interest as the presence of a decoder preserves the ability to generate images. We find that training CFM models on feature representations can lead to improvements in likelihood reliability, but only for certain datasets and certain parameterisations of the feature space, at a cost in sample quality. Further investigation suggests possible links between likelihood reliability and geometric characteristics of the data and the feature space, and opens avenues for future work.
ABSTRACT Infrared imaging technology holds a prominent position across various applications, particularly with the burgeoning significance of infrared object detection technology. Although prior research has explored physical attacks on infrared object detectors, the practical implementation of these techniques remains intricate. For example, certain methodologies involve the deployment of specialised equipment such as bulb boards or infrared QR suits to execute attacks, necessitating costly optimisation and cumbersome deployment processes. Alternatively, other strategies employ irregular aerogel as physical perturbations for infrared attacks, albeit with associated optimisation expenses and perceptibility issues. In this study, we introduce a novel infrared physical attack termed ‘adversarial infrared geometry’ (AdvIG), which enables efficient black‐box query attacks by modelling diverse geometric shapes such as lines, triangles and ellipses and optimising their physical parameters using particle swarm optimisation (PSO). We conduct extensive experiments to assess the effectiveness, stealthiness and robustness of AdvIG. In digital attack experiments, line, triangle and ellipse patterns achieve success rates of 93.1%, 86.8% and 100.0%, respectively, with average query times of 71.7, 113.1 and 2.6, respectively, thereby confirming the efficiency of AdvIG. Physical attack experiments are performed to evaluate the success rate of AdvIG at various distances. On average, the success rates for lines, triangles and ellipses are 61.1%, 61.2% and 96.2%, respectively. We further conduct comprehensive experiments to analyse AdvIG, including ablation experiments, transfer attack experiments and assessments of adversarial defence mechanisms. Given the superior performance of our method as a simple and efficient black‐box adversarial attack in both digital and physical environments, we advocate broader attention to AdvIG.
Due to the importance of the parathyroid glands (PG) for health, detecting and preserving them during endoscopic thyroid surgery is vital. However, existing parathyroid detection methods face issues from colour variations, target deformation, blurriness, and lighting effects in complex surgical environments. Therefore, they fail to extract high‐quality features and perform poorly in our detection tasks. The essential reasons for these issues are two‐fold: (A) ignoring the spatial relation among targets and (B) failing to identify the shapes, colours, and positions of targets, mainly due to insufficient utilisation of depth information , especially under lighting variations and occlusions, which are unavoidably prone to false or missed detection. To better discover and exploit the inter‐target spatial relation (solving A), inspired by the power of graph neural networks on dependency modelling, we explore an effective graph‐based framework specifically designed for PG detection. Specifically, we propose a novel double‐layer graph attention network (i.e., DL‐GAT for short), which explicitly facilitates local augmentation through the identification of key visual features (e.g., texture and shape) and global interactions. Besides, it also has merits in robustly combating image blur and better differentiating PG targets and background parts, thus improving detection precision. On the other hand, to solve B and better incorporate depth knowledge , we further propose a depth augmentation component, which can adaptively capture the intrinsic geometrical features of targets based on depth information and thus significantly improve light intensity robustness and naturally enhance generalisability. Moreover, because we lack a thyroid endoscopy surgery benchmark to evaluate and compare the performance of models for this task, we meticulously established a novel data set from 838 real surgeries performed (via the fully laparoscopic thoracic‐breast approach) at the Fujian Medical University Union Hospital. Extensive experiments show that our framework achieves superior PG detection accuracy compared to its current state‐of‐the‐art counterparts while maintaining real‐time efficiency.
Pedestrian Attribute Recognition is a foundational computer vision task that provides essential support for downstream applications, including person retrieval in video surveillance and intelligent retail analytics. However, existing research is frequently constrained by the “one-model-per-dataset" paradigm and struggles to handle significant discrepancies across domains in terms of modalities, attribute definitions, and environmental scenarios. To address these challenges, we propose UniPAR, a unified Transformer-based framework for PAR. By incorporating a unified data scheduling strategy and a dynamic classification head, UniPAR enables a single model to simultaneously process diverse datasets from heterogeneous modalities, including RGB images, video sequences, and event streams. We also introduce an innovative phased fusion encoder that explicitly aligns visual features with textual attribute queries through a late deep fusion strategy. Experimental results on the widely used benchmark datasets, including MSP60K, DukeMTMC, and EventPAR, demonstrate that UniPAR achieves performance comparable to specialized SOTA methods. Furthermore, multi-dataset joint training significantly enhances the model's cross-domain generalization and recognition robustness in extreme environments characterized by low light and motion blur. The source code of this paper will be released on https://github.com/Event-AHU/OpenPAR
Convolutional sparse coding (CSC) using learnt convolutional dictionaries has recently emerged as an effective technique for emphasising discriminative structures in signal and image processing applications. In this paper, we propose a multilayer model for convolutional sparse networks (CSNs), based on hierarchical convolutional sparse coding and dictionary learning, as a competitive alternative to conventional deep convolutional neural networks (CNNs). In the proposed CSN architecture, each layer learns a convolutional dictionary from the feature maps of the preceding layer (if available), and then uses it to extract sparse representations. This hierarchical process is repeated to obtain high-level feature maps in the final layer, suitable for pattern recognition and classification tasks. One key advantage of the CSN framework is its reduced sensitivity to training set size and its significantly lower computational complexity compared to CNNs. Experimental results on image classification tasks show that the proposed model achieves up to 7% higher accuracy than CNNs when trained with only 150 samples, while reducing computational cost by at least 50% under similar conditions.
With urbanisation accelerating, predicting the heat release rate (HRR) of building fires using visual data has emerged as a pivotal research focus in the field of fire rescue. However, existing approaches face challenges, such as limited training data and complex models, which lead to suboptimal performance and slow inference speeds. To address these issues and adapt to the rapid morphological changes of smoke in dynamic fire environments, we propose a lightweight neural network prediction model based on adaptive pooling with channel information interaction (APCI). This model can achieve high precision while maintaining faster inference speed. Our approach employs simplified dense connections to propagate shallow smoke features, thereby effectively capturing the relationship between smoke textures and multiscale features to accommodate the variations of smoke morphologies. To mitigate the loss of smoke features caused by spatial misalignment and ventilation disturbances during downsampling, we introduce an adaptive weighted pooling mechanism that fully leverages the detailed information contained in the invoked smoke. Additionally, an enhanced channel shuffle operation in channel information interaction ensures effective cross-level transfer to detail-aware information exchange during sudden escalations in fire intensity in the hybrid feature fusion framework. Experiments on the smoke-heat release rate dataset we created demonstrate that the proposed method can achieve a coefficient of determination R 2 $\left({R}^{2}\right)$ of 0.937, a root mean square error (RMSE) of 23.0 kW, a mean absolute error (MAE) of 17.4 kW and with an inference time of 4.13 ms per image.
This article is protected by copyright. All rights reserved. This article is protected by copyright. All rights reserved.