Evaluating Open-Source Large Vision-Language Models for Face Presentation Attack Detection | AMiner
Evaluating Open-Source Large Vision-Language Models for Face Presentation Attack Detection
Camilo Andrés Linares Cáceres,Marta Gomez-Barrero
2026 14th International Workshop on Biometrics and Forensics (IWBF)(2026)
Biometrics & Machine Learning Lab (BioML) - RI CODE
被引用0|浏览1
摘要
Biometric authentication systems, which use individual biological and behavioral characteristics for recognition, have become increasingly popular and are central to the security of modern infrastructures. However, these systems remain vulnerable to presentation attacks. The perpetrator can use manipulated biometric artifacts to deceive the sensor. Recent progress in Artificial Intelligence (AI), more exactly in multi-modal large language models (MLLMs), including GPT-5 and Google's Gemini, has created new ways to improve presentation attack detection (PAD). In this work, we investigate the baseline performance of MLLMs, more specifically, large vision-language models (LVLMs), for detecting face presentation attacks without fine-tuning. Our focus is on zero-shot and few-shot learning scenarios. We discuss how these models, trained on large-scale datasets and capable of processing text and images, can offer new strategies to improve the security of biometric systems. We discuss several main challenges, including dataset diversity and prompt sensitivity, and suggest directions for future work, with a focus on fine-tuning and task-specific adaptation. Through an analysis of the possible advantages and limitations of MLLMs for PAD, this paper intends to inform researchers and professionals looking to improve the security and robustness of biometric authentication technologies.