
Cloth-Changing person Re-Identification (CC-ReID) aims to address the challenge of person retrieval when clothing changes cause significant variations in appearance. This task is particularly challenging due to the large intra-class appearance differences introduced by clothing variations. Although CC-ReID has made considerable progress, most existing methods still mainly focus on decoupling identity features from clothing information. This operation ignores semantic relationships and fine-grained identity details, which are crucial for accurate pedestrian recognition. To address these limitations, we propose the Semantic-aware Relation and Identity Enhancement Network (SRIE-Net), which effectively integrates semantic-aware relationships and fine-grained identity features to enhance pedestrian recognition under clothing changes. The SRIE-Net consists of three key components: Firstly, a Clothing Attention Weakening (CAW) stream employs an orthogonal complement projection to extract clothing features and penalize the model’s attention to clothing variations, thereby mitigating the influence of clothing changes; Then, a Pose-Guided Learning (PGL) stream aggregates features based on semantic relationships and adapts to diverse posture variations by leveraging pose-guided information; Finally, a Fine-grained Identity Enhancement (FIE) stream learns discriminative fine-grained identity features independent of clothing and fuses them with global features to enhance identity representation. Most importantly, these three streams are jointly optimized in an end-to-end unified framework, ensuring that the model learns complementary features from each stream. Extensive experiments on three public CC-ReID datasets (PRCC, LTCC, and NKUP) demonstrate the superiority of SRIE-Net. It outperforms existing methods on multiple datasets, especially in handling the challenges posed by clothing and posture variations. The code can be available from https://github.com/VPMS-Open-Source/SRIE-Net.
Precision livestock farming depends on accurate, non-invasive, and reliable identification of individual cattle to maintain traceability, biosecurity, and overall animal welfare. Currently, ear tags and microchips are the most common cattle identification methods deployed in real farm; however, recent progress in computer vision has highlighted the potential of biometric-based identification techniques. Deep convolutional neural network (CNN) and Vision transformers (ViTs), the trending computer vision models, have demonstrated a strong performance in several biometric applications, but their use in animal identification remains limited. ViTs entail a significant computational cost, while CNNs do not capture long range dependencies. In this paper, we propose a lightweight hybrid CNN–ViT method that strengthens local feature extraction through CNN and provide more informative and refined representation to the transformers for reliable identification of cows using muzzle feature. The main contributions of this work include the creation of a 140-subject cattle video dataset, muzzle localization from video frames, an enhanced self-supervised frame selection strategy for selecting informative frames, a lightweight and hybrid CNN–ViT classifier for capturing local and global biometric features, and attention enhanced CNN and adaptive patch projection enhanced ViT models for robust feature representation. Experimental results demonstrate an identification accuracy of 98.57%, a compact model size of 7.89 (MB), and an inference time of 4.78 ms per image. These results indicate that the proposed method achieves an effective balance between accuracy and efficiency, making it suitable for real-world cattle identification under practical farm conditions.
Gait recognition (GR) has emerged as a crucial non-contact, long-distance biometric technology, driving advancements in intelligent surveillance, public security, and identity verification. Most existing GR methods rely on 3D Convolution Neural Networks (CNNs) to model spatio-temporal dynamics, yet they suffer from high computational costs and feature misalignment issues. Furthermore, conventional approaches typically employ simple max temporal pooling to capture global sequence features, which fails to model global and multi-scale temporal features and thus discards critical discriminative information. To alleviate these issues, we propose a lightweight 2D convolutional backbone network consisting of three novel modules: (1) a Multi- Frame Temporal Shift Module (MFTSM); (2) an Efficient Multiscale Temporal Pooling (EMTP) module; and (3) a Query-based Temporal Pooling (QTP) module. The MFTSM captures local continuous micro-motion features between consecutive frames while expanding the temporal receptive field without increasing parameter count. The EMTP module employs cascaded depthwise separable 3D convolutions along the temporal dimension to progressively enlarge the temporal receptive field, and aggregates hierarchical features to generate multi-scale temporal representations. The QTP module converts variable-length gait sequences into compact and fixed-dimensional temporal representations using learnable queries, while remaining fully compatible with part-based gait modeling. Extensive experiments on four benchmark datasets demonstrate that our method achieves recognition accuracy comparable or exceeding state-of-the-art approaches while preserving a lightweight architecture.
Fairness evaluation in face recognition relies on a growing number of metrics, each designed to capture a particular form of disparity and often validated under limited or synthetic conditions. As a result, practitioners still lack clear guidance on which metric to use in a given context. We address this issue through a unified framework that evaluates nineteen fairness metrics on four public face datasets. Three complementary sources of bias are considered: algorithmic variation through four training loss functions, data-related variation through image quality degradation and dataset resampling, and operational variation through decision-threshold changes. All experiments are conducted on real biometric data under a common evaluation protocol. For the controlled data-related and operational variations, metric validity is assessed through monotonicity and stability across demographic groups. This analysis reveals how the mathematical structure of each metric determines the type of disparity to which it is most responsive. It also provides a common basis for comparing score-based and error-rate-based metrics under the same experimental conditions. The results show that metric suitability depends strongly on the source and manifestation of bias, and that no single metric is appropriate for all conditions. Based on these findings, we propose a practical guideline linking bias conditions to the metric structures best suited to detect them, supporting fairness audits and the future development of biometric standards.
This paper investigates whether there are gender-specific differences in iris images that cause variations in the accuracy of presentation attack detection (PAD). We assemble a dataset of iris images with gender metadata, containing both bona fide images and images of eyes wearing textured contact lenses. This dataset enables a comprehensive experimental analysis of four key research questions relating to the existence of gender-related effects, their generalizability, and their mitigation through training and testing datasets balancing. The analysis employs six different PAD algorithms, two dataset balancing strategies, and three classes of performance metrics to assess PAD performance. Our large-scale and diverse experiments show that although there are statistically significant variations in iris PAD performance for specific combinations of dataset, model, and training regime, such variations are not rooted in population-wide male/female differences. Rather, observed accuracy differences across gender are found to be related to non-gender-based variations between groups that are only accidentally correlated with the bona fide/presentation attack class label.
Electrocardiogram (ECG) biometric recognition has emerged as a promising technique for identity authentication due to its inherent anti-spoofing capabilities and the accessibility of wearable ECG devices. However, existing methods often face the challenge of significant performance degradation in long-term cross-session applications. The ECG non-stationarity and the semantic differences among intra-beat components are the major challenges to the stability of identity representation. To address these issues, we propose PGDC, a physiology-guided divide-and-conquer framework for learning heartbeat representation for cross-time drift-resilient ECG biometrics. Specifically, we design a decomposition strategy that divides a single heartbeat into depolarization (P-QRS), ventricular repolarization (ST-T), and global views. We extract independent features for each component using dedicated Transformer encoders. Subsequently, an adaptive gating fusion dynamically aggregates complementary multi-view features. Moreover, we propose a decoupled training-registration pipeline to leverage multi-session, non-target data, which reduces dependence on limited enrollment samples without fine-tuning during deployment. Extensive experiments conducted on five public and self-collected datasets show that PGDC significantly outperforms state-of-the-art baseline approaches in both identification and verification tasks. PGDC delivers impressive cross-session robustness with a 10-second single-session enrollment. This framework presents a promising solution for achieving long-term, cross-session robust ECG biometrics under minimal enrollment, making it well-suited for open-set, wearable applications.
Accurate identification of individuals in surveillance footage is essential for public safety, crime prevention, law enforcement efficiency, and forensic investigations. While eyewitness accounts are valuable, relying solely on their semantic descriptions to recognize suspects in CCTV footage poses significant challenges. This paper therefore considers the integration of soft biometrics, which are attributes, semantically described by eyewitnesses to enhance face profile recognition in traditional biometric systems even under surveillance conditions and more importantly to demonstrate that face profile biometrics is a standalone biometric modality. By leveraging both computer vision techniques and eyewitnesses’ comparative labels, we develop a hybrid system capable of subject identification and retrieval using both images and eyewitness testimonies. Experiments conducted using XM2VTSDB dataset demonstrate the efficacy of this approach, achieving a recognition accuracy of 98% and EER of 0.86% thereby to validate face profiles as a standalone biometric modality. Additionally, a fusion technique of our soft and traditional biometric systems is also presented in this study. These findings highlight the value of integrating comparative soft biometrics with traditional methods, to offer a powerful framework for semantic-based biometric systems in real-world forensic and surveillance scenarios.
As one of the most popular biometric recognition methods in identity identification, palmprint and palm vein patterns have recently attracted much attention due to their uniqueness and stability. Most existing deep learning-based palmprint and palm vein fusion recognition methods are based on convolutional neural network (CNN) model. However, these methods struggle with preserving global features and the recognition performance will significantly decrease when dealing with incomplete images caused by injury, lighting, or hand posture. In addition, the feature maps output from the network are usually fused through element addition or cascading, without considering different weights of palmprint and palm vein modalities. To alleviate this problem, we present a novel two-stage multimodal dynamic fusion framework for palmprint and palm vein fusion recognition, called TransCNN-MLDFN. Specifically, we design a multi-level dynamic fusion network, which includes a global feature extraction block based on Transformer network and a local feature extraction block based on CNN. Following this, we propose a two-stage dynamic weighted fusion strategy combining shallow and deep features by setting the weights of palmprint and vein images as learnable parameters to fully integrate the feature information of two modalities. Finally, we design a new loss function to improve the accuracy and generalization ability of the model. Numerous experiments on four widely used and public available palmprint and palm vein datasets show that our proposed method outperforms other methods in terms of accuracy and equal error rate on both complete and incomplete image recognition scenarios.
Micro-expression (ME), reflecting individual genuine emotions, can be widely applied in psychological treatments, suspect interrogation. However, ME usually occurs in localized facial regions and is characterized by the brief and low-intensity nature, posing significant challenges to robust micro-expression recognition (MER). To address this issue, this paper proposes a local optical flow feature enhancement capsule network, named LOFECAP, to learn fine-grain and discriminative representation for MER. The LOFECAP model integrates cross-attention mechanism with Transformer to enhance the learning capability on local subtle features. Specifically, it can concentrate on holistic contextual dependence associated with local features and incorporate them with the corresponding regional blocks from input images. Based on this foundation, the aggregated local features are fed into the capsule network for local relation exploration. Additionally, we acquire the optical flow map between the onset and apex frames and then implement the pre-trained model to capture intricate motion features. To evaluate the effectiveness of our MER model, the extensive experiments are conducted on three publicly available databases, namely SMIC, SAMM, and CASME II, with unweighted F1 score (UF1) and unweighted average recall (UAR) as the evaluation metrics. The experimental results show that the LOFECAP model achieves the superior performance to other state-of-the art MER methods, yielding the UF1 and UAR scores of 0.8332 and 0.8624, respectively.
ECG and PPG biometrics are emerging as robust alternatives to traditional authentication methods, particularly in scenarios requiring enhanced security, continuous authentication, and strong resilience against spoofing attacks. This paper investigates and validates the feasibility of launching adversarial attacks against cardiovascular biometric authentication systems by using diffusion models to generate ECG and PPG data that are difficult for authentication algorithms to distinguish from genuine signals. First, we developed a method using a diffusion model to synthesize fake biometric signals for user impersonation, achieving a high attack success rate against completely black-box authentication models. Secondly, we directly extracted users’ rPPG signals from videos and assessed the potential for exploiting video data to compromise users’ ECG-based authentication systems. Our experimental results show that across all algorithms, we achieved an average of 45% attack success rate within five attempts, thus validating the real threat posed by synthetic signals. This finding highlights that even without direct access to the cardiovascular signals used for authentication, an attacker can still launch highly successful attacks by synthesizing signals to deceive the verification system. Our experiments highlight potential vulnerabilities in current biometric systems and underscore the need to develop more secure and attack-resistant authentication technologies.
Morphing attacks remain a critical threat to face recognition systems, particularly in document-style acquisition settings where such manipulations have the most significant operational impact. Most existing detection methods primarily focus on texture-level or pixel-wise inconsistencies, often neglecting geometric cues that more directly reflect morph-induced deformations. This work investigates the use of UV texture maps derived from 3D facial reconstruction as a canonical, geometry-aware representation for morphing attack detection. We hypothesize that UV parametrization can enhance morph detectability by exposing distortions on the reconstructed facial surface and by amplifying mapping inconsistencies propagated through the reconstruction process, including those indirectly influenced by artifacts outside the facial region. Extensive experiments conducted on document-like datasets and across multiple deep learning architectures show that integrating UV-based representations with original images consistently improves the separability between bona fide and morphed samples. Evaluations under challenging cross-dataset and cross-manipulation conditions further confirm the robustness of the proposed approach, which outperforms models trained solely on raw images and achieves superior performance compared to representative state-of-the-art baselines. Overall, these findings indicate that UV texture maps can act as a geometry-induced canonicalization layer for morphing attack detection, largely independent of the underlying classifier or morphing strategy, and represent a promising direction for enhancing the reliability of MAD systems in document-oriented applications.
Robust detection methods are essential to address the growing threat of audio deepfakes and synthetic media, which are increasingly appearing in complex, real-world conditions. Existing research has primarily focused on clean, single-language audio, overlooking critical challenges like bilingualism, background noise, and various forms of audio codec compression. This paper introduces the BhashaBluff dataset, with approximately 2,500 hours of audio data and over 3.88 million samples generated using 21 distinct methods. Its key innovation lies in its extensive variability across three critical dimensions: (i) bilingualism, which encompasses Hindi and English deepfakes as well as code-mixed speech; (ii) noise, with samples corrupted by four distinct environmental noise types; and (iii) compression, incorporating seven neural compression techniques to simulate real-world artifacts. Benchmark evaluations using state-of-the-art models show substantial performance degradation when faced with these challenging conditions, revealing the limited generalizability of current methods and highlighting the urgent need for more adaptive and robust algorithms. BhashaBluff is publicly available1 and serves as a crucial benchmark for developing next-generation detection systems that are effective in diverse, noisy, and compressed audio environments.
Photoplethysmography (PPG) biometrics have gained significant interest in recent years. However, PPG signals are dynamic and are influenced by physiological and environmental factors, making generalization across their distribution challenging. We propose an invariant representation learning (IRL) framework for PPG biometric recognition to alleviate distribution shift, abbreviated as IRL-PPG. To enhance PPG data diversity, we generate augmented samples via the phase-spectrum perturbation and employ a deep cascade framework comprising a disentanglement 1D vision Transformer (DViT) and a dual feature reconstruction module (DFRM). Within the DViT, a locality-aware feed-forward network and a progressive disentanglement module are employed to separate PPG signals into biometric-invariant and distribution-sensitive features. In the DFRM, to enhance the DViT’s generalization under distribution shifts, we reconstruct the disentangled representations to preserve feature integrity. Specifically, invariant features are utilized for PPG biometrics, while the specific features are isolated to minimize distribution-shift effects on biometric models. Extensive experiments demonstrate that IRL-PPG outperforms state-of-the-art methods across four public PPG databases. The code used in these experiments is available at https://github.com/oniopluto/IRL-PPG.
Person re-identification (Re-ID) across visible and infrared modalities is crucial for 24-hour surveillance systems, but existing datasets primarily focus on ground-level perspectives. While ground-based IR systems offer nighttime capabilities, they suffer from occlusions, limited coverage, and vulnerability to obstructions—problems that aerial perspectives uniquely solve. To address these limitations, we introduce AG-VPReID.VIR, the first aerial-ground cross-modality video-based person Re-ID dataset. This dataset captures 1,837 identities across 4,861 tracklets (124,855 frames) using both UAV-mounted and fixed CCTV cameras in RGB and infrared modalities. AG-VPReID.VIR presents unique challenges including cross-viewpoint variations, modality discrepancies, and temporal dynamics. We propose TCC-VPReID, a novel three-stream architecture designed to address the joint challenges of cross-platform and cross-modality person Re-ID. We enhance this with ODE-Fusion, a neural ODE framework for fusing paired RGB-IR captures from the same platform, facilitating effective cross-modal feature fusion through continuous state evolution. Our approach bridges the domain gaps through style-robust feature learning, memory-based cross-view adaptation, intermediary-guided temporal modeling, and ODE-based fusion. Experiments show that AG-VPReID.VIR presents distinctive challenges compared to existing datasets, with our enhanced TCC-VPReID framework achieving significant performance gains across multiple evaluation protocols. Dataset and code will be made available.
Micro-expressions (MEs) are transient in nature and reveal enough visual cues to recognize genuine emotions. Neural Architecture Search (NAS) has recently gained significant attention in the field of micro expression recognition (MER). However, existing NAS approaches in MER rely on general-purpose strategies, often reducing network depth to compensate for limited training data. This restriction can limit the discovery of optimal architectures, especially since dataset sizes in MER vary significantly. Moreover, focusing solely on operation selection and connectivity is inadequate, flexibility in depth and structure is essential to adapt to different data characteristics. While shallow networks are often considered suitable for small datasets, this assumption falls short for MER. Despite the scarcity of training samples, MEs are inherently complex, fleeting, and nuanced, requiring models with high representational power for accurate recognition. To address this trade-off, we propose a novel Dynamic Neural Architecture Search framework for MER (DNAS-MER), designed to adaptively balance model depth and architectural design based on the data. The proposed DNAS hierarchically optimizes the network depth, resolution paths, and cell-level operations through a scalable network search space and a multi-scale context-aware (MSCA) cell search space. A novel depth-aware loss function enables the model to automatically adapt depth based on data needs, while MSCA cell search, embedded with adaptive feature fusion operations, enhances ME-specific feature representation to better capture the complexity of MEs. We evaluate DNAS-MER on a composite dataset MEGC2019, CASME-II, and SMIC. Experimental results demonstrate that DNAS-MER consistently outperforms existing state-of-the-art methods for both video and image-based MER tasks.
Iris recognition technology is widely used in identity verification and other fields due to its high accuracy and strong security. However, with the popularization of the technology, the privacy protection of iris data has become increasingly prominent. Existing systems are often difficult to protect personal data privacy while guaranteeing recognition accuracy, especially when facing data leakage and attacks, iris features cannot be modified once leaked, which is extremely risky. Therefore, proposing an efficient and accurate protection scheme that can ensure the security of iris data has become an important issue that needs to be solved urgently. In this paper, we propose an iris template protection framework based on positive-negative mapping and the negative selective algorithm, called the IrisPN-NS. The framework effectively protects iris data by constructing a complex system of inequality equations and combining it with the negative selection algorithm, thus increasing the difficulty for attackers to crack. Based on this framework, we propose two iris template protection methods, IrisPNM-NS and IrisPNS-NS. We conducted experiments on the CASIA-IrisV3-Interval, CASIA-IrisV4-Lamp, and IIT Delhi Iris datasets. The experimental results show that the scheme achieves high performance and effectively resists the attack of the greedy algorithm, while satisfying the requirements of irreversibility, revocability, and unlinkability.
This work focuses on proposing a presentation attack detection (PAD) system for ID cards based on a meta-learning approach, such as Few-Shot Learning (FSL), to determine whether an image is bona fide, printed, or displayed on a screen with only a few samples (50). This approach involves a commercial PAD system trained in one country, such as Chile, that extends its capabilities to other countries with fewer available images, such as Nicaragua, Honduras, and El Salvador. We demonstrate that Prototypical Networks generalise effectively to new bona fide and attack with an average EER of 3,10% and BPCER20 of 2,80%. Our experiments validate FSL as a comprehensive solution for diverse presentation attack modalities in mobile production environments where acquiring large datasets is impractical1.
Biometric authentication based on dynamic hand gestures has attracted increasing attention due to its integration of both physiological and behavioral traits. Among various approaches, point cloud–based dynamic gesture recognition has emerged as a particularly promising direction, owing to its ability to preserve rich 3D geometric information. However, existing methods still struggle to effectively model point-level motion features while capturing overall temporal dynamics, which limits their ability to capture fine-grained motion information in point cloud sequences. To address these challenges, we first propose a point cloud–based dynamic gesture recognition framework, termed Motion-Temporal Network (MoTeNet). MoTeNet explicitly models cross-frame point-wise motion, enabling enhanced fine-grained dynamic perception and effective extraction of point-level motion features, thereby achieving strong recognition performance. Nevertheless, this explicit motion modeling paradigm relies on multi-stage feature extraction and fusion, and requires repeated neighborhood searches across adjacent frames, resulting in high computational complexity. To further improve efficiency, we propose the Implicit Spatio-Temporal Network (ISTNet), which adopts a novel motion modeling strategy. Specifically, ISTNet constructs a cross-frame neighborhood structure during spatial feature learning, allowing motion information to be seamlessly integrated into spatial representations through local feature aggregation. This design enables implicit motion encoding, effectively avoiding the additional computational overhead introduced by explicit motion modeling. In addition, a Saliency-Aware Motion Encoder (SAME) is introduced to capture complementary global motion representations, further enhancing the modeling of holistic dynamic patterns. Extensive experiments on three public benchmarks, including SHREC’17, DHG, and NVGesture, demonstrate that the proposed methods achieve state-of-the-art performance in terms of both recognition accuracy and computational efficiency, with ablation studies further validating the effectiveness of the proposed framework.
Vision foundation models have demonstrated strong transferability across diverse visual recognition tasks and are increasingly considered for biometric applications. Their suitability for iris Presentation Attack Detection (PAD), particularly under realistic open-set operating conditions, remains insufficiently examined. This work presents a systematic failure analysis of general-purpose vision foundation models for open-set iris PAD using periocular imagery. Five representative foundation models are evaluated under three open-set protocols that explicitly separate different sources of distribution shift: unseen Presentation Attack Instruments (PAIs), unseen datasets captured with different sensors and cross-spectral transfer from near-infrared (NIR) to visible spectrum (VIS) imagery. Both frozen feature representations and parameter-efficient task adaptation using Low-Rank Adaptation (LoRA) are assessed within a unified experimental framework. The results indicate that foundation models can transfer across datasets with similar sensing characteristics, but fail to generalise reliably to unseen attack instruments and degrade sharply under cross-spectral evaluation. While LoRA improves performance in certain cross-dataset settings, it frequently amplifies failure under attack-level and spectral shifts. Additional validation experiments using segmented iris inputs, full backbone fine-tuning, joint cross-dataset and cross-PAI shifts, and reverse VIS to NIR transfer further confirm that these failures are not simply artefacts of periocular input, weak adaptation, or one-directional spectral evaluation. These findings show that strong closed-set or cross-dataset performance should not be treated as evidence of robust open-set security, and highlight the need for PAD representations that maintain sensitivity to presentation artefacts while remaining stable under realistic deployment variation.
As a highly secure biometric modality, finger vein recognition holds significant value in personal identification applications. However, some recognition modules have a large number of parameters and a habit of dependency on large-scale training images. To deal with above problems, this paper proposes a dual convolutional fusion Attention Based parameter-efficient Contrastive learning Network (ABCNet) for finger vein recognition, which achieves a low computational cost and a high recognition accuracy by the compact network architecture and supervised contrastive learning. In detail, we design a parameter-efficient backbone network, which effectively reduces parameters by optimizing residual layers and bottleneck blocks. Additionally, a dual convolution fusion attention (DCFAtt) is proposed to enhance feature representation, in which the attention weight is computed by a depthwise separable convolution path and a pointwise convolution path. Moreover, the supervised contrastive learning strategy makes use of class labels to build positive and negative image pairs for training our network, which reduces the dependency on large-scale training images. The effectiveness of our ABCNet is proved on SDU, HKPU and USM datasets with the recognition accuracies of 99.84%, 99.15%, 99.80% respectively. Compared with the state-of-the-art models, our ABCNet has fewer parameters and higher recognition accuracy.