Guangdong Key Lab. of Intelligent Info. Proc. and Shenzhen Key Lab. of Media Security
被引用0|浏览0
摘要
Display-recapture attacks pose a critical threat to the integrity and authenticity of digital document images, particularly by concealing tampering traces through rephotographing displayed content. Existing document presentation attack detection (DPAD) methods often struggle to distinguish forensic artifacts (e.g., chromatic distortions and moiré patterns) from the natural textures inherent in documents with complex backgrounds. To address this texture confusion challenge, we propose a dual-stream LC&DF framework that integrates Local Chromaticity (LC) features with a Masked Attention mechanism and a Discriminative Frequency (DF) branch enhanced via a Frequency-domain Moiré-Aware Adapter (FMAAda). This architecture jointly models local chromatic distortions and global frequency cues to robustly isolate recapture-induced artifacts from genuine document content. Extensive evaluations demonstrate the superiority of our method. Under the cross-dataset protocol, LC&DF achieves state-of-the-art performance. In challenging in-the-wild scenarios evaluated on the ROD_M&F, SRDID162, DLC2021, and KID34K benchmarks, our method consistently outperforms existing baselines, achieving an AUC of 0.9105 and reducing the Equal Error Rate by up to 15.72 percentage points on challenging document benchmarks. Furthermore, we conduct a zero-shot evaluation of Multimodal Large Language Models (MLLMs), revealing that they lack the sensitivity to subtle forensic artifacts. Visualizations of challenging samples further confirm that our claims on distinguishing the forensic artifacts from the document textures. The source code of this work will be available upon acceptance.