ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)(2026)
University of Science and Technology of China
被引用0|浏览6
摘要
The proliferation of photorealistic, AI-generated images threatens public trust, creating an urgent need for detectors that can reliably distinguish them from real content. However, existing detectors, typically trained on limited known generators, struggle to generalize to unseen ones. The rapid evolution of diffusion models has intensified this problem, creating a "diffusion fog" that limits practical applicability of existing detectors. To pierce this fog, we posit that the inherent patterns in the denoising process of all diffusion models are key to generalizable detection. To isolate these patterns, we create minimal fake training images by applying a single-step denoising operation to slightly perturb real images, forcing the detector to learn the subtle artifacts of denoising. We then blend these minimal fakes with their real counterparts, erasing superficial cues from the perturbation operation to prevent detector from learning shortcuts. Considering that the fake images closely resemble real ones, we propose a feature separation loss to enhance detector’s discrimination capacity. For rigorous evaluation, we constructed a new benchmark, DiffuGen, comprising 65K synthetic images from 13 modern diffusion models and 5K real images. Empirical results demonstrate that our detector achieves significantly improved generalization, reaching an average accuracy of 99% on DiffuGen.
更多
查看译文
关键词
AI-generated image detection,diffusion models,minimal fake training samples,deepfake