
Multi-face tracking is a challenging problem, as low-quality face images (e.g., partially occluded faces, lateral faces, motion blur, etc.) are prevalent. A multi-face tracker called FaceQSORT mitigates this problem by combining two different types of features (i.e., biometric face features and generic appearance features) extracted from the same image (face) patch. In particular, the appearance features provide a general visual representation of the image (face) patch and compensate for limited biometric information. To extract these appearance features, a generic image classifier is utilized. In this work, the selection of the appearance feature model (image classifier) and the biometric feature model (face recognition model) is evaluated empirically. For this purpose, the following questions are investigated: (i) How crucial is the choice of appearance model? (ii) How does the choice of appearance model affect the inter-person ID switches? To answer these questions, a comprehensive experimental evaluation is conducted that includes multiple tracking parameter settings and different combinations of biometric and appearance model pairs.
This paper investigates the impact of multi-sample fusion at enrolment on the performance and security of biometric cryptosystems, specifically focusing on the fuzzy vault scheme. Biometric template protection is critical for safeguarding sensitive biometric data, and we propose enhancing robustness by averaging multiple biometric samples from the same instance of a modality during the enrolment phase. This approach aims to reduce intra-class variability and stabilize feature representations, thereby improving recognition accuracy without significantly impacting usability. We evaluate a fuzzy vault system based on deep learning-based fingerprint embeddings, using a comprehensive fingerprint dataset (MCYT330) with varying sample qualities. Our experimental results demonstrate that fusing multiple enrolment samples significantly shifts mated score distributions, leading to notable improvements in recognition rates and security in both unprotected and fuzzy vault systems. This work confirms the practical benefits of multi-sample fusion in developing more reliable and secure privacy-preserving biometric systems.
Iris segmentation is an important step in determining downstream iris recognition performance. Most iris segmentation systems fail to reliably segment the iris images at extreme off-angles that go beyond 30 degrees. Obtaining a manual ground truth for different extreme angles is both a costly and a time-consuming undertaking. In this paper, our aim is to automate this iris segmentation process for different extreme gaze angles, while minimizing the iris recognition performance loss. To address this, we propose an Iris segmentation method using a novel distance-based loss function and a residual-based ellipse fitting method to overcome these obstacles. The novel method shows an overall best performance of 0.9809 Area Under Curve (AUC) and an equal error rate (EER) of 0.0628, reducing the gap to the performance using ideal manual segmentation to 1.88% in terms of AUC and 0.056% in terms of EER.
This paper presents a pilot study on the feasibility of using a selected commercial electronic fingerprint biometric capture subsystem for post mortem biometric recognition. The study investigates the effectiveness of contact and contactless devices in capturing fingerprints from deceased individuals within 48 hours after death. The environmental conditions, legal and medical implications, and post mortem changes that affect skin integrity are examined. Fingerprint quality and interoperability were evaluated using the NFIQ 2 and IDkit tools. The results show that while traditional dactyloscopic cards offer lower image quality for algorithmic comparison, some modern digital devices provide promising results for post mortem fingerprint acquisition and comparison.
Tattoo recognition is a critical task in forensics; however, existing approaches often face limitations under real-world conditions such as occlusion, distortion, and low image quality. This paper presents a novel and robust tattoo retrieval framework that combines the global semantic representation capabilities of OmniGlue with the geometric precision of LightGlue. Notably, the proposed framework is entirely based on foundation models and operates in a training-free manner, enabling rapid deployment without the need for dataset-specific fine-tuning. Extensive experiments conducted on the challenging WebTattoo dataset demonstrate the effectiveness of the approach. In closed-set evaluation, the framework achieves a rank-1 identification rate of 90.91%, significantly outperforming standalone models as well as the state-of-the-art TattTRN. In open-set identification, the method attains an Equal Error Rate of 11.56%, outper-forming equally the state of the art, i.e., TattTRN (20.75%). These results highlight the advantages of integrating complementary foundation models through mechanisms, establishing a new benchmark for tattoo recognition in both controlled and unconstrained environments.
Face recognition systems are primarily tailored to adult faces, and numerous studies have shown that their performance significantly declines when applied to children. The development of child-specific recognition models is further constrained by the scarcity of publicly available datasets and the ethical and privacy concerns surrounding the collection of children's biometric data. To overcome these limitations, we propose GANDiffKids, a novel framework for synthesizing realistic child face images with diverse intra-class variations. Our approach combines a generative adversarial network (GAN) with a diffusion model - a strategy that has proven effective for adult face synthesis in prior work. We demonstrate that GANDiffKids can generate a large-scale dataset of synthetic child faces across a range of image qualities. The resulting dataset is benchmarked using a state-of-the-art face recognition model, establishing its potential as a valuable resource for advancing child-specific face recognition research.
As biometric recognition systems become increasingly prevalent in security-critical applications, ensuring their robustness against attacks is paramount to maintaining trust and reliability. This is particularly important for systems relying on highly accurate biometric traits such as the iris, for which malicious users may attempt to gain unauthorized access by resorting to artifacts or synthetic samples, i.e., through presentation attacks. This paper explores the capabilities of two state-of-the-art attention-based architectures, namely vision transformers (ViTs) and shifted windows (Swin) transformers, for iris presentation attack detection (PAD). Furthermore, a detailed analysis is performed on a relatively underexplored aspect, specifically the effects of image compression (JPEG and JPEG AI) on data-driven PAD. The obtained results testify to the effectiveness of transformer-based PAD frameworks, whose behaviors can be interpreted by exploring visual maps under degraded image quality. The trade-off between compression efficiency and PAD accuracy is also analyzed.
We investigate the capabilities of deep learning methods for the quality assessment of fingerprint samples. Starting from general considerations for the implementation of a deep learning-based fingerprint quality assessment, we modify and fine-tune six convolutional neural networks and one vision transformer for the specific task. For this, we generate a synthetic training database and propose a labelling based on the normalized average sample comparison score. The obtained results in terms of error vs. discard characteristic curves show that all deep learning methods outperform the NFIQ 2.3 baseline. Most notably, the best-performing VGG16 model results in a superior predictive performance than NFIQ 2.3 on all tested datasets. A detailed interpretation of the obtained results paves the way for more specific and tailored solutions. We suggest to consider the proposed algorithms as a starting point for a development of a deep learning-based fingerprint quality assessment methods and make them publicly available.
With the growing prevalence of biometric recognition technology, the demand for diverse large-scale test datasets is significant, and data collection costs are increasing. Consequently, synthetic images attract attention as a potential solution. However, some synthetic images deviate from target representativeness; even if visually close, subtle alterations can negatively impact biometric recognition accuracy. To address the challenge, this paper introduces a novel Applicability Evaluation framework for synthetic images and biometric engines based on the image quality and score similarity. Synthetic Image Quality Assessment (SIQA) evaluates how closely synthetic images approximate target representativeness. Score Similarity Assessment (SSA) compares biometric engine scores and verifies identity similarity that is crucial for performance testing. Experimental results demonstrate that this framework effectively identified combinations of synthetic images and biometric engines that maintain recognition accuracy while ensuring the reliability of biometric performance testing. It can improve the trustworthiness and efficiency of biometric performance testing.
In the era of digital communication, creating avatars that are both expressive and privacy-preserving is increasingly important, particularly in biometric-sensitive domains such as teletherapy, virtual education, and secure online interactions. We present SynthFace, a framework that generates realistic talking avatars which accurately convey a user’s expressions while ensuring unlinkable and non-reversible identity replacement. Unlike conventional methods that depend on stochastic GAN sampling or real facial images, SynthFace employs a deterministic, key-conditioned synthetic identity generator. This generator cryptographically combines a secret key with a reference image to produce a unique and reproducible facial representation, enhancing entropy while maintaining no direct link to any real individual. Facial expressions and head pose are extracted from a source video using DECA-based 3D Morphable Models (3DMMs) and fused at the coefficient level with the synthetic identity, allowing precise control over identity and expression. For animation, a mel-spectrogram-driven module synthesizes temporally coherent lip and head movements for speech-driven scenarios. We evaluate SynthFace in both visual quality and biometric verification contexts, using ArcFace cosine similarity, SSIM, GradSim, and Dlib landmark distances for identity and expression analysis, along with PSNR and optical flow for temporal smoothness. We also report verification-oriented metrics such as similarity threshold analysis and false acceptance/rejection behaviour, aligning with biometric performance evaluation protocols. Experiments on VoxCeleb2, CREMA-D, and a custom dataset show that SynthFace achieves low mean similarity to source identities, high expression fidelity, and smooth animation, achieving a balance of privacy preservation and expression realism that is competitive with existing anonymization approaches.
With the rising importance of biometric authentication in numerous security applications it is critical to implement Presentation Attack Detection (PAD) techniques to protect against counterfeit biometrics. We evaluate a PAD-approach that uses 3D Time-of-Flight (ToF) cameras to perform heart rate measurements based on remote photoplethysmography (rPPG) for PAD. Compared to Red-Green-Blue (RGB)-camera based approaches, ToF technology provides advantages, including integration of 3D depth information, insensitivity to electronic displays and independence from ambient light and robustness against its influence. We introduce a Spectral pulse Peak prominence Ratio (SPR) and motion based metric and evaluate their applicability to PAD and present first promising results.
This work proposes a Presentation of Attack Detection (PAD) on ID card systems, focusing on improving the detection of manual composite attacks. These kinds of attacks are used in a remote verification system to impersonate the identity of one subject in order to obtain benefits. For this purpose, relevant text can be modified, copied, and pasted. Faces and other areas can be manually changed because of the alignment of the originally issued ID card to a fake ID card. This work proposed an “online” methodology to create composite attack images during the training process and shows that using automatic composite attacks may reduce the error rate in the PAD system. To demonstrate it, three experiments were set up that compared the performance of models trained with and without the aforementioned technique on a private dataset and a cross-evaluation dataset.
As biometric authentication becomes more common, protecting biometric data is becoming increasingly important. One widely used protection method is encryption. However, not all encryption methods are suitable for biometric data. On the one hand, the encryption solution can lead to worse performance or on the other hand, not fulfill all required security measures. This paper proposes a solution for iris template protection that results in hash-encrypted data using Maximum Entropy Binary (MEB) codes and Convolutional Neural Networks (CNN). The method is compared to a baseline approach to demonstrate competitive recognition performance. In addition, the privacy and security properties of the template protection are evaluated.
This paper presents a post-quantum secure multiparty computation (MPC) scheme using homomorphic encryption (PQ-MHE) applied for fingerprint template protection. The proposed scheme leverages the CKKS cryptosystem for homomorphic operations and additive secret sharing (ASS) for splitting, storing, and comparing biometric templates across multiple servers. This architecture ensures that no single entity ever has access to the unencrypted templates. We benchmark the system using two-, three-, and four-server configurations. The scheme achieves a comparison time of 1.24 seconds and incurs 35 MB of communication, which can be reduced by approximately 98% through caching and reusing the CKKS cryptographic context on the server side. The inherent approximate arithmetic of CKKS introduces a minor degradation, resulting in a 0.5 percentage point decrease in the true acceptance rate (TAR) compared to unencrypted fingerprint samples. The system is secure under a semi-honest adversary model with non-colluding servers and complies with the biometric information protection standards outlined in ISO/IEC 24745.
This study investigates the use of SHAP (SHapley Additive exPlanations) values as an explainable artificial intelligence (xAI) technique applied on a facial attribute classification task. We analyse the consistency of SHAP value distributions across diverse classifier architectures that share the same feature extractor, revealing that key features driving attribute classification remain stable regardless of classifier architecture. Our findings highlight the challenges in interpreting SHAP values at the individual sample level, as their reliability depends on the model’s ability to learn distinct class-specific features; models exploiting inter-class correlations yield less representative SHAP explanations. Furthermore, pixel-level SHAP analysis reveals that superior classification accuracy does not necessarily equate to meaningful semantic understanding; notably, despite FaceNet exhibiting lower performance than CLIP, it demonstrated a more nuanced grasp of the underlying class attributes. Finally, we address the computational scalability of SHAP, demonstrating that KernelExplainer becomes infeasible for high-dimensional pixel data, whereas DeepExplainer and GradientExplainer offer more practical alternatives with trade-offs. Our results suggest that SHAP is most effective for small to medium feature sets, providing interpretable and computationally manageable explanations.
Gait recognition is a biometric technology that identifies individuals based on their walking posture and motion. It has significant potential for applications such as criminal investigations and security systems. In silhouette-based gait recognition, silhouette normalization is a key component that enhances recognition accuracy. However, in scenarios where partial occlusions frequently occur, normalization can sometimes negatively affect recognition performance due to inappropriate adjustments. To address this issue, we propose a normalization refinement network based on spatial transformer networks, designed to mitigate the adverse effects of improper normalization. To validate the effectiveness of our method, we apply it to multiple gait recognition models and evaluate its performance under various occlusion scenarios using the OU-MVLP dataset. Experimental results demonstrate that our approach improves recognition accuracy across different occlusion conditions and multiple recognition models.
We study the complementarity of different CNNs for periocular verification at different distances on the UBIPr database. We train three architectures of increasing complexity (SqueezeNet, MobileNetv2, and ResNet50) on a large set of eye crops from VGGFace2. We analyse performance with cosine and chi2 metrics, compare different network initialisations, and apply score-level fusion via logistic regression. In addition, we use LIME heatmaps and Jensen-Shannon divergence to compare attention patterns of the CNNs. While ResNet50 consistently performs best individually, the fusion provides substantial gains, especially when combining all three networks. Heatmaps show that networks usually focus on distinct regions of a given image, which explains their complementarity. Our method significantly outperforms previous works on UBIPr, achieving a new state-of-the-art.
Age verification is increasingly critical for regulatory compliance, user trust, and the protection of minors online. Historically, solutions have struggled with poor accuracy, intrusiveness, and significant security risks. More recently, concerns have shifted toward privacy, surveillance, fairness, and the need for transparent, trustworthy systems. In this paper, we propose Biometric Bound Credentials (BBCreds) as a privacy-preserving approach that cryptographically binds age credentials to an individual's biometric features without storing biometric templates. This ensures only the legitimate, physically present user can access age-restricted services, prevents credential sharing, and addresses both legacy and emerging challenges in age verification. enhances privacy.
Despite growing attention to deepfake speech detection, the aspects of bias and fairness remain underexplored in the speech domain. To address this gap, we introduce the Speaker Characteristics Deepfake (SCDF) dataset: a novel, richly annotated resource enabling systematic evaluation of demographic biases in deepfake speech detection. SCDF contains over 237,000 utterances in a balanced representation of both male and female speakers spanning five languages and a wide age range. We evaluate several state-of-the-art detectors and show that speaker characteristics significantly influence detection performance, revealing disparities across sex, language, age, and synthesizer type. These findings highlight the need for bias-aware development and provide a foundation for building non-discriminatory deepfake detection systems aligned with ethical and regulatory standards.
Fingerprints are widely recognized as one of the most unique and reliable characteristics of human identity. Most modern fingerprint authentication systems rely on contact-based fingerprints, which require the use of fingerprint scanners or fingerprint sensors for capturing fingerprints during the authentication process. Various types of fingerprint sensors, such as optical, capacitive, and ultrasonic sensors, employ distinct techniques to gather and analyze fingerprint data. This dependency on specific hardware or sensors creates a barrier or challenge for the broader adoption of fingerprint based biometric systems. This limitation hinders the widespread adoption of fingerprint authentication in various applications and scenarios. Border control, healthcare systems, educational institutions, financial transactions, and airport security face challenges when fingerprint sensors are not universally available. To mitigate the dependence on additional hardware, the use of contactless fingerprints has emerged as an alternative. Developing precise fingerprint segmentation methods, accurate fingerprint extraction tools, and reliable fingerprint matchers are crucial for the successful implementation of a robust contactless fingerprint authentication system. This paper focuses on the development of a deep learning-based segmentation tool for contactless fingerprint localization and segmentation. Our system leverages deep learning techniques to achieve high segmentation accuracy and reliable extraction of fingerprints from contactless fingerprint images. In our evaluation, our segmentation method demonstrated an average mean absolute error (MAE) of 30 pixels, an error in angle prediction (EAP) of 5.92 degrees, and a labeling accuracy of 97.46%. These results demonstrate the effectiveness of our novel contactless fingerprint segmentation and extraction tools.