Existing palm presentation attack detection (PAD) datasets are often limited by static imagery, restricted acquisition conditions, or insufficient multimodal video data, hindering systematic evaluation across environments, modalities, and attack types. We present GBU-Palm, a large-scale multimodal video dataset and benchmark containing 21,326 videos from 105 subjects and 210 palms across six acquisition environments, including bona fide, Print, and Replay presentations, with 6,310 synchronized RGB-NIR samples. We construct leakage-controlled protocols that separate palm identity and attack lineage and benchmark four representative video architectures under environment-matched and held-out-environment settings. Results reveal substantial architecture-dependent degradation under environmental shift and show that RGB-NIR fusion does not consistently outperform RGB-only input. We further analyze model behavior through true accept (TA), true reject (TR), false accept (FA), and false reject (FR) decomposition, spectral masking, temporal-order intervention, and frozen-backbone NIR probing, revealing distinct failure patterns and evidence utilization across architectures. GBU-Palm provides a unified and challenging benchmark for developing and evaluating robust multimodal palm PAD methods under cross-environment conditions.
The primary obstacle in realizing the full potential of finger crease biometrics is the accurate identification of deformed knuckle patterns, often resulting from completely contactless imaging. Current methods struggle significantly with this task, yet accurate matching is crucial for applications ranging from forensic investigations, such as child abuse cases, to surveillance and mobile security. To address this challenge, our study introduces the largest publicly available dataset of deformed knuckle patterns, comprising 805,768 images from 351 subjects. We also propose a novel framework to accurately match knuckle patterns, even under severe pose deformations, by recovering interpretable knuckle crease keypoint feature templates. These templates can dynamically uncover graph structure and feature similarity among the matched correspondences. Our experiments, using the most challenging protocols, illustrate significantly outperforming results for matching such knuckle images. For the first time, we present and evaluate a theoretical model to estimate the uniqueness of 2D finger knuckle patterns, providing a more interpretable and accurate measure of distinctiveness, which is invaluable for forensic examiners in prosecuting suspects.
Contactless 3D fingerprint identification systems have emerged to provide more accurate and hygienic alternatives to contact-based conventional systems that acquire hundreds of millions of fingerprints everyday. However, the intricate process of acquiring 3D fingerprints presents a significant challenge, acting as a key barrier to fully unlocking the potential of 3D fingerprint biometrics. This paper introduces a novel framework to directly recover corresponding 3D minutiae template from a single contactless 2D fingerprint image. Billions of contact-based fingerprints have been acquired and employed everyday for e-governance and other applications. Seamless adoption of contactless 3D fingerprint technologies also requires advanced capabilities to accurately match 3D fingerprints with respective 2D fingerprint templates, which is currently missing in existing literature. We therefore introduce novel capabilities to accurately align minutiae templates in 3D spaces and enable compensation for the unknown perspective transformation. This capability significantly enhances the ability to accurately match 3D to 3D and 3D to 2D fingerprint templates. Furthermore, we introduce a new approach to synthesizing realistic contactless fingerprint images, resulting in the generation of a large synthetic database complete with corresponding 3D ground truths of minutiae points. Finally, we provide a detailed theoretical analysis of formulation for the uniqueness of recovered 3D minutiae templates, providing a theoretical justification for the superiority of such 3D minutiae templates over their 2D counterparts.
We comment on a recently published TPAMI paper presenting an iris recognition algorithm. While the approach is intriguing, we have identified several inconsistencies and errors in this paper. Additionally, their comparison with the state-of-the art methods lacks fairness. We take this opportunity to clarify and underline these errors, aiming to assist fellow researchers like us who are interested in advancing biometrics research.
This paper addresses two key challenges in detecting diffusion-model generated videos: generalization to unseen but sophisticated video synthesizers and threats from the multimodal manipulations where the deepfakes combine the text prompts and visual synthesis to create realistic forgeries. We propose a novel multimodal framework built on a transformer-based architecture, which was originally designed for image forgery detection. Our framework extends this architecture by integrating two complementary components: (i) spatio-temporal feature extraction and (ii) a text-guided enrichment module which uses a frozen Vision Language Models (VLMs) text encoder when prompts are available and a small set of learnable default embeddings when no prompt is provided. Trained on our self-curated dataset comprising KlingAI and StableDiffusion Samples, we present cross-dataset performance from unseen but hyper-realistic fake video generators comprising Sora, Luma, Pika, and Runway. Our model can achieve high accuracy and outperformance results, which demonstrate cross-model generalization for detecting hyper-realistic AI-generated videos.
Deep neural networks have successfully applied for fingerprints and palmprints recognition. However, the performances of most fingerprint and palmprint recognition methods are often degraded in young children, due to the small physical size and poor image quality. Thus, it is still challenging for reliable fingerprint and palmprint recognition of young children. In this study, we propose a hybrid graph transformer network to gradually learn both local and global features of minutiae points, which are integrated to enhance the matching accuracy of fingerprints and palmprints in young children. First, a multi-scale transformer is designed to extract the features of minutiae points detected in fingerprints and palmprints. Then, a graph attention network (GAN) is proposed to capture the topological structure of minutiae set for more powerful representation and precise minutiae matching. Furthermore, to address the class imbalance problem between the positive and negative matched minutiae pairs, a filtering-and-refining strategy is developed for coarse-to-fine minutiae matching of fingerprints and palmprints. The extensive experiments and comparisons are performed on two fingerprint datasets and one palmprint dataset of young children, showing the promising performance of the proposed method.
In the quest to advance solar energy conversion, understanding the intricacies of charge transfer (CT) at interfaces has become crucial for photovoltaic applications. Efficient CT is essential for minimizing energy losses due to recombination and enhancing the efficiency of the devices. Herein, we delve into the significance of the CT process in the improvement of current conduction across the interface by using several spectroscopic and microscopic measurements. A photoinduced CT mechanism in CsPbBr3 perovskite quantum dots (PQDs) is explored with a band-aligned hole acceptor, anthracene. All of the spectroscopic measurements reveal a prominent CT from PQDs to anthracene through hole transfer, as evinced from the energy level diagram and density of states calculations. The electrical measurements across an electrode-semiconductor-electrode junction using conductive atomic force microscopy reflect an increase in the conductance in PQD after the introduction of anthracene. The inclusion of anthracene in the PQD changes the nature of the current-voltage curve from nonlinear to Ohmic one, revealing only direct tunneling and both direct and Fowler-Nordheim tunneling for PQD with and without anthracene, respectively. These findings can contribute to the development of efficient and cost-effective optoelectronic devices with a careful selection of simple molecules as charge transport or some additional layers.
The uniqueness of many biometric modalities can be attributed to the fine-grained and randomly textured patterns that are revealed in respective normalized images. This paper introduces a new approach to accurately match such biometric images using micropattern exemplars. The images acquired using different sensors, or spectrum, from the same biometric surface often reveal localized fine-grained micropatterns that maintain high similarity in such differently sensed images. Therefore, selecting appropriate micropattern exemplars that can encode similar micropatterns in both images is expected to enhance cross-sensor matching capabilities. We introduce specialized masks that are designed to efficiently encode such microstructural patterns for more accurate biometric identification. Our experiments specifically focus on the cross-modal palm patterns and introduce micropattern exemplars to match such biometric features more accurately. The experimental results presented on multiple public and fine-grained biometrics databases validate the effectiveness of the proposed approach in characterizing micro patterns that are retained in the images from different sensors. These results are highly encouraging and validate the effectiveness of the proposed approach for cross-sensor and cross-spectral biometrics identification. Our micropattern exemplars-based approach is quite generalized, and we present reproducible experiments on other biometric databases, from iris and knuckle, to underline its potential for a range of other biometric modalities.
Accurate biometric identification under real environments is one of the most critical and challenging tasks to meet growing demand for higher security. This paper proposes a new framework to efficiently and accurately match periocular images that are automatically acquired under less-constrained environments. Our framework, referred to as semantics-assisted convolutional neural networks (SCNNs) in this paper, incorporates explicit semantic information to automatically recover comprehensive periocular features. This strategy enables superior matching accuracy with the usage of relatively smaller number of training samples, which is often an issue with several biometrics. Our reproducible experimental results on four different publicly available databases suggest that the SCNN-based periocular recognition approach can achieve outperforming results, both in achievable accuracy and matching time, for less-constrained periocular matching. Additional experimental results presented in this paper also indicate that the effectiveness of proposed SCNN architecture is not only limited to periocular recognition but it can also be useful for generalized image classification. Without increasing the volume of training data, the SCNN is able to automatically extract more discriminative features from the input data than a single CNN, therefore can consistently improve the recognition performance. The experimental results presented in this paper validate such an approach to enable faster and more accurate periocular recognition under less constrained environments.
The remarkable progress in neural-network-driven visual data generation, especially with neural rendering techniques like Neural Radiance Fields and 3D Gaussian splatting, offers a powerful alternative to GANs and diffusion models. These methods can produce high-fidelity images and lifelike avatars, highlighting the need for robust detection methods. In response, an unsupervised training technique is proposed that enables the model to extract comprehensive features from the Fourier spectrum magnitude, thereby overcoming the challenges of reconstructing the spectrum due to its centrosymmetric properties. By leveraging the spectral domain and dynamically combining it with spatial domain information, we create a robust multimodal detector that demonstrates superior generalization capabilities in identifying challenging synthetic images generated by the latest image synthesis techniques. To address the absence of a 3D neural rendering-based fake image database, we develop a comprehensive database that includes images generated by diverse neural rendering techniques, providing a robust foundation for evaluating and advancing detection methods.
Contactless fingerprint identification has emerged as an reliable and user friendly alternative for the personal identification in a range of e-business and law-enforcement applications. It is however quite known from the literature that the contactless fingerprint images deliver remarkably low matching accuracies as compared with those obtained from the contact-based fingerprint sensors. This paper develops a new approach to significantly improve contactless fingerprint matching capabilities available today. We systematically analyze the extent of complimentary ridge-valley information and introduce new approaches to achieve significantly higher matching accuracy over state-of-art fingerprint matchers commonly employed today. We also investigate least explored options for the fingerprint color-space conversions, which can play a key-role for more accurate contactless fingerprint matching. This paper presents experimental results from different publicly available contactless fingerprint databases using NBIS, MCC and COTS matchers. Our consistently outperforming results validate the effectiveness of the proposed approach for more accurate contactless fingerprint identification.
Ming-Hsuan Yang合作论文数Vision and Learning Lab, University of California, Merced;Google DeepMind4