While specialized detectors for AI-Generated Images (AIGI) achieve near-perfect accuracy on curated benchmarks, they suffer from a dramatic performance collapse in realistic, in-the-wild scenarios. In this work, we demonstrate that simplicity prevails over complex architectural designs. A simple linear classifier trained on the frozen features of modern Vision Foundation Models , including Perception Encoder, MetaCLIP 2, and DINOv3, establishes a new state-of-the-art. Through a comprehensive evaluation spanning traditional benchmarks, unseen generators, and challenging in-the-wild distributions, we show that this baseline not only matches specialized detectors on standard benchmarks but also decisively outperforms them on in-the-wild datasets, boosting accuracy by striking margins of over 30%. We posit that this superior capability is an emergent property driven by the massive scale of pre-training data containing synthetic content. We trace the source of this capability to two distinct manifestations of data exposure: Vision-Language Models internalize an explicit semantic concept of forgery, while Self-Supervised Learning models implicitly acquire discriminative forensic features from the pretraining data. However, we also reveal persistent limitations: these models suffer from performance degradation under recapture and transmission, remain blind to VAE reconstruction and localized editing. We conclude by advocating for a paradigm shift in AI forensics, moving from overfitting on static benchmarks to harnessing the evolving world knowledge of foundation models for real-world reliability.
Consumer health electronic systems increasingly rely on the seamless integration of sensing, communication, computing, and control to deliver intelligent, adaptive, and context-aware services. However, under imbalanced environments and heterogeneous network conditions, these integrated systems often exhibit performance disparities that affect reliability and user trust. In this paper, we propose a novel framework for fair and robust integration of sensing, communication, computing, and control in consumer health electronics. Our approach adopts a three-fold strategy: 1) feature disentanglement to separate environment-dependent signals from task-relevant features, reducing bias in sensing and decision-making; 2) a fairness-aware optimization scheme that balances resource allocation across diverse devices and network conditions, ensuring equitable performance; and 3) loss landscape flattening through sharpness-aware optimization to enhance robustness against distribution shifts in communication channels and computing platforms. Extensive experiments on representative IoT and edge-computing benchmarks demonstrate that our method significantly improves both fairness and stability across heterogeneous consumer health scenarios, providing a promising direction for equitable and reliable smart device ecosystems.
Generative models now produce imperceptible, fine-grained manipulated faces, posing significant privacy risks. However, existing AI-generated face datasets generally lack focus on samples with fine-grained regional manipulations. Furthermore, no researchers have yet studied the real impact of splice attacks, which occur between real and manipulated samples, on detectors. We refer to these as detector-evasive samples. Based on this, we introduce the DiffFace-Edit dataset, which has the following advantages: 1) It contains over two million AI-generated fake images. 2) It features edits across eight facial regions (e.g., eyes, nose) and includes a richer variety of editing combinations, such as single-region and multi-region edits. Additionally, we specifically analyze the impact of detector-evasive samples on detection models. We conduct a comprehensive analysis of the dataset and propose a cross-domain evaluation that combines IMDL methods. Dataset will be available at https://github.com/ywh1093/DiffFace-Edit.
Multi-label Chest X-ray (CXR) classification faces significant challenges from the inherently imperfect nature of clinical data, particularly the complex interplay of co-occurring pathologies, training data with a long-tailed distribution, and high visual similarity between distinct diseases. To address these challenges, we propose a novel framework that synergizes medical prior knowledge with prototype-driven contrastive learning, enabling disentangled and discriminative per-pathology representation learning. In particular, our approach integrates a co-occurrence modulated Label Graph Attention (LGA) module, which leverages semantic prior knowledge from a pre-trained large language model (LLM) and statistical co-occurrence patterns from training data to model inter-pathology relationships. Subsequently, a Label-Aware Decoupling (LAD) decoder is proposed to isolate pathology-specific visual features and mitigate feature suppression by dominant classes. Furthermore, we introduce an Adaptive Proto type Contrastive Learning (APCL) mechanism to enhance the discriminability of visually similar pathologies. Extensive experiments on the NIH ChestX-ray14 and CheXpert datasets demonstrate the framework's superiority, achieving state-of-the-art mean AUCs of 0.834 and 0.840, respectively. Furthermore, cross-dataset evaluations on the external MIMIC-CXR dataset validate the framework's exceptional zero-shot and few-shot generalization capabilities, highlighting its strong robustness and potential for real world clinical deployment. The implementation is available at https://github.com/ZengXHYX/Learning-from-Prototypes
Face inpainting with diffusion models has recently achieved impressive visual quality, yet preserving identity fidelity under significant occlusion and conflicting text guidance remains a major challenge. To address this issue, we present Reference Semantic Inpainting for Face (ReSem-Face), a cascaded diffusion framework that introduces an explicit identity-conditioned semantic prior for multi-reference face inpainting. Our approach distills representative identity features from multiple references to reconstruct missing semantic regions, which then guide the diffusion process through a multi-stream conditioning architecture. This design provides strong semantic constraints when pixels are absent and stabilizes identity reconstruction while remaining compatible with prompt-driven edits. Experiments on CelebAHQ-IDI-5 and VGGFace2 demonstrate that ReSem-Face yields more reliable identity-preserving completion under severe semantic masks and improves text-controlled editing quality compared with representative baselines.
The rapid deployment of large language models (LLMs) in edge-intelligent Internet of Things (IoT) systems has enabled autonomous content generation and intelligent human-device interaction. However, the widespread use of AI-generated text also introduces new forensic challenges, including misinformation dissemination, automated social engineering, and malicious content injection. Effective post-incident investigation requires not only identifying machine-generated content but also extracting informative forensic cues that can support subsequent analysis and event reconstruction. To address this challenge, we propose ChatZoom, an AI-driven forensic framework for analyzing and identifying AI-generated textual evidence in edge-enabled IoT environments. Instead of relying solely on surface-level linguistic cues, ChatZoom leverages linguistic representation models to uncover latent generation patterns embedded in machine-generated content and characterizes these patterns as auxiliary forensic cues for AI-generated text identification. The extracted evidence is subsequently utilized to distinguish AI-generated and human-authored texts while facilitating interpretable forensic analysis. Extensive experiments on publicly available benchmarks demonstrate that ChatZoom achieves highly reliable forensic identification performance, reaching an average precision of 99.61% on in-domain datasets. Furthermore, the proposed framework exhibits strong generalization capability across heterogeneous data distributions, achieving an average precision of 82.09% under cross-domain settings. These results indicate that ChatZoom can serve as an effective component for AI-generated text identification and subsequent forensic analysis. Our code and data can be found in https://github.com/redscarf-liu/lcd.
With the rapid advancement of video generation models such as Veo and Wan, the visual quality of synthetic content has reached a level where macro-level semantic errors and temporal inconsistencies are no longer prominent. However, this does not imply that the distinction between real and cutting-edge high-fidelity fake is untraceable. We argue that AI-generated videos are essentially products of a manifold-fitting process rather than a physical recording. Consequently, the pixel composition logic of consecutive adjacent frames residual in AI videos exhibits a structured and homogenous characteristic. We term this phenomenon `Manifold Projection Fluctuations' (MPF). Driven by this insight, we propose a hierarchical dual-path framework that operates as a sequential filtering process. The first, the Static Manifold Deviation Branch, leverages the refined perceptual boundaries of Large-Scale Vision Foundation Models (VFMs) to capture residual spatial anomalies or physical violations that deviate from the natural real-world manifold (off-manifold). For the remaining high-fidelity videos that successfully reside on-manifold and evade spatial detection, we introduce the Micro-Temporal Fluctuation Branch as a secondary, fine-grained filter. By analyzing the structured MPF that persists even in visually perfect sequences, our framework ensures that forgeries are exposed regardless of whether they manifest as global real-world manifold deviations or subtle computational fingerprints.
Fairness is a core element in the trustworthy deployment of deepfake detection models, especially in the field of digital identity security. Biases in detection models toward different demographic groups, such as gender and race, may lead to systemic misjudgments, exacerbating the digital divide and social inequities. However, current fairness-enhanced detectors often improve fairness at the cost of detection accuracy. To address this challenge, we propose a dual-mechanism collaborative optimization framework. Our proposed method innovatively integrates structural fairness decoupling and global distribution alignment: decoupling channels sensitive to demographic groups at the model architectural level, and subsequently reducing the distance between the overall sample distribution and the distributions corresponding to each demographic group at the feature level. Experimental results demonstrate that, compared with other methods, our framework improves both inter-group and intra-group fairness while maintaining overall detection accuracy across domains.
Image generation algorithms are increasingly integral to diverse aspects of human society, driven by their practical applications. However, insufficient oversight in artificial intelligence-generated content (AIGC) can facilitate the spread of malicious content and increase the risk of unauthorized use. Among the diverse range of image generation models, the Latent Diffusion Model (LDM) is currently the most widely used, dominating the majority of the Text-to-Image model market. Currently, most attribution methods for LDMs rely on directly embedding watermarks into the generated images or their intermediate noise, a practice that compromises both the quality and the robustness of the generated content. To address these limitations, we introduce TraceMark-LDM, a novel algorithm that integrates watermarking to attribute generated images while guaranteeing non-destructive performance. Unlike current methods, TraceMark-LDM leverages watermarks as guidance to rearrange random variables sampled from a Gaussian distribution. To mitigate potential deviations caused by inversion errors, the small-magnitude elements are grouped and strategically rearranged. Additionally, we fine-tune the LDM encoder to enhance the robustness of the watermark. Experimental results show that images synthesized using TraceMark-LDM exhibit superior quality and attribution accuracy compared to state-of-the-art (SOTA) techniques. Notably, TraceMark-LDM demonstrates exceptional robustness against various common attack methods, consistently outperforming SOTA methods. Our code is available at https://github.com/luowhDevSpace/TraceMark.
DeepFake, an AI-driven face-swapping technique, has been weaponized to spread disinformation. In response, researchers have developed forensic detectors to identify such manipulations. To circumvent these defenses, a growing body of work now focuses on generating adversarial samples—carefully perturbed forgeries designed to deceive detection tools. However, most existing adversarial generation methods sacrifice image quality to achieve undetectability, introducing perceptible artifacts that ironically make them more detectable under human scrutiny. To address this limitation, we propose a novel spectral fusion approach to multimodally synthesize forgery traces from authentic facial images. Unlike traditional noise injection methods, our technique integrates diffusion-based noise during image preprocessing, embedding perturbations in the forward process of a diffusion model. This approach not only deceives forensic detectors more effectively but also preserves high visual fidelity. Through extensive experiments, our method achieves state-of-the-art DeepFake anti-forensic performance while preserving high visual fidelity, ensuring that the adversarial samples remain indistinguishable from real images.
With the emergence of adversarial steganography, existing specialized steganalysis models suffer a significant performance decline in detecting non-homologous adversarial steganographic methods (i.e., trained on traditional-based method and tested on adversarial-based method), resulting in insufficient robustness in complex network environments. To address this issue, we propose a two-stage robust steganalysis framework with multi-feature enhancement against adversarial steganography. The framework integrates edge-aware attention with multi-dimensional statistical features to enhance robustness against adversarial steganography. In the first stage, we design a covariance pooling based convolutional neural network and integrate an edge-aware attention mechanism to improve the feature representation of subtle steganographic traces, enabling fast detection for most samples. In the second stage, samples with uncertain confidence scores from the first stage are further analyzed by extracting block-wise entropy features, global entropy features, and SRM co-occurrence features, followed by dimensionality reduction via principal component analysis (PCA) and classification using a random forest. The final decision is made through the collaborative fusion of the two stages. Experimental results demonstrate that the proposed method achieves excellent detection performance (2.57% average improvement over the existing best method) with strong robustness for adversarial steganography, and its generalization capability is further validated in cross-dataset scenarios. Furthermore, comprehensive ablation studies validate the efficacy of the network architecture.
With the rapid advancement of generative AI, synthetic images have become increasingly realistic, raising serious concerns regarding their potential misuse. Several prior methods detect AI-generated images by exploiting low-level fake artifacts left by generators. However, these fake artifacts are fragile and easily suppressed by post-processing (e.g., JPEG compression), leading detectors to misclassify fake images as real. To address this robustness limitation, we propose ArtGate, a novel AI-generated image detector that integrates a CLIP-ViT backbone with a frequency-domain artifact branch, and carefully design a confidence-aware gating mechanism to modulate the artifact branch. Specifically, the artifact branch employs wavelet analysis to extract artifact features from the high-frequency subbands of the image, while a confidence-aware gating mechanism selectively activates this branch only when generation-related fake artifacts are detected. After modulation by the gating mechanism, the artifact features are subsequently injected into the CLIP image encoder and adaptively fused with semantic features to enhance the detection performance. Extensive experiments demonstrate that ArtGate outperforms state-of-the-art methods in both generalization and robustness. Under random JPEG compression, our method achieves improvements of +5.88% in accuracy and +9.19% in F1-score on the AIGCDetectBenchmark.
Current AIGC detectors often achieve near-perfect accuracy on images produced by the same generator used for training but struggle to generalize to outputs from unseen generators. We trace this failure in part to latent prior bias: detectors learn shortcuts tied to patterns stemming from the initial noise vector rather than learning robust generative artifacts. To address this, we propose On-Manifold Adversarial Training (OMAT): by optimizing the initial latent noise of diffusion models under fixed conditioning, we generate on-manifold adversarial examples that remain on the generator's output manifold-unlike pixel-space attacks, which introduce off-manifold perturbations that the generator itself cannot reproduce and that can obscure the true discriminative artifacts. To test against state-of-the-art generative models, we introduce GenImage++, a test-only benchmark of outputs from advanced generators (Flux.1, SD3) with extended prompts and diverse styles. We apply our adversarial-training paradigm to ResNet50 and CLIP baselines and evaluate across existing AIGC forensic benchmarks and recent challenge datasets. Extensive experiments show that adversarially trained detectors significantly improve cross-generator performance without any network redesign. Our findings on latent-prior bias offer valuable insights for future dataset construction and detector evaluation, guiding the development of more robust and generalizable AIGC forensic methodologies.
Faces synthesized by diffusion models (DMs) with high-quality and controllable attributes pose a significant challenge for Deepfake detection. Most state-of-the-art detectors only yield a binary decision, incapable of forgery localization, attribution of forgery methods, and providing analysis on the cause of forgeries. In this work, we integrate Multimodal Large Language Models (MLLMs) within DM-based face forensics, and propose a fine-grained analysis triad framework called VLForgery, that can 1) predict falsified facial images; 2) locate the falsified face regions subjected to partial synthesis; and 3) attribute the synthesis with specific generators. To achieve the above goals, we introduce VLF (Visual Language Forensics), a novel and diverse synthesis face dataset designed to facilitate rich interactions between Visual and Language modalities in MLLMs. Additionally, we propose an extrinsic knowledge-guided description method, termed EkCot, which leverages knowledge from the image generation pipeline to enable MLLMs to quickly capture image content. Furthermore, we introduce a low-level vision comparison pipeline designed to identify differential features between real and fake that MLLMs can inherently understand. These features are then incorporated into EkCot, enhancing its ability to analyze forgeries in a structured manner, following the sequence of detection, localization, and attribution. Extensive experiments demonstrate that VLForgery outperforms other state-of-the-art forensic approaches in detection accuracy, with additional potential for falsified region localization and attribution analysis.
While specialized detectors for AI-generated images excel on curated benchmarks, they fail catastrophically in real-world scenarios, as evidenced by their critically high false-negative rates on `in-the-wild' benchmarks. Instead of crafting another specialized `knife' for this problem, we bring a `gun' to the fight: a simple linear classifier on a modern Vision Foundation Model (VFM). Trained on identical data, this baseline decisively `outguns' bespoke detectors, boosting in-the-wild accuracy by a striking margin of over 20%. Our analysis pinpoints the source of the VFM's `firepower': First, by probing text-image similarities, we find that recent VLMs (e.g., Perception Encoder, Meta CLIP2) have learned to align synthetic images with forgery-related concepts (e.g., `AI-generated'), unlike previous versions. Second, we speculate that this is due to data exposure, as both this alignment and overall accuracy plummet on a novel dataset scraped after the VFM's pre-training cut-off date, ensuring it was unseen during pre-training. Our findings yield two critical conclusions: 1) For the real-world `gunfight' of AI-generated image detection, the raw `firepower' of an updated VFM is far more effective than the `craftsmanship' of a static detector. 2) True generalization evaluation requires test data to be independent of the model's entire training history, including pre-training.
The integration of Cognitive Internet of Things (IoT) sensors with autonomous aerial vehicles (AAVs) has transformed industrial sectors, such as monitoring, logistics, and infrastructure inspection. However, the advancement of visual synthesis technologies like generative adversarial networks and diffusion models has introduced significant risks by enabling the creation of highly realistic AI-manipulated content, making the detection of falsified imagery increasingly challenging. Existing detection methods, largely based on convolutional neural networks (CNNs), focus primarily on global image features and often overlook crucial relational connections, limiting their robustness and generalization. To overcome these limitations, we propose a novel dual-stream architecture that integrates global feature extraction with relational feature learning. By combining the CLIP model with a graph-based topology, our approach identifies hard-to-detect samples and processes them through a graph convolutional network (GCN) to capture both structural and relational information. Extensive evaluations validate the robustness and generalization ability of our method across various generative models and real-world perturbations. This approach offers a scalable and reliable solution to ensure data integrity in industrial IoT systems, helping to preserve societal trust in AI-driven applications.
With the significant success of Generative Adversarial Networks and Diffusion Models in visual synthesis, the risk of disinformation has also increased. Despite many researchers dedicated to solving this issue, they still rely on limited visual features when dealing with complex generated images. In this work, we propose vision-text interactive hybrid granularity proxy representation learning, which overcomes the limitations of previous methods that solely rely on visual information by incorporating cross-modal textual information. First, we design a cross-modal interactive reconstruction framework that uses a reconstruction mechanism to understand and learn information from different modalities, thereby preliminarily establishing semantic relationships between image-text pairs. Based on this, the hybrid granularity proxy representation learning method is introduced, which utilizes fine-grained proxy points within the same modality and coarse-grained proxy points across modalities to reduce the feature space distance between the two modalities. This method not only helps mitigate the interference of label noise on representation learning but also enhances the discrimination of the feature representations through text-guided cross-modal indirect alignment. We have carried out extensive experimental verification on multiple datasets, and the experimental results show that our method shows significant performance improvement in accuracy and generalization.
The high-quality, realistic images generated by generative models pose significant challenges for exposing them. So far, data-driven deep neural networks have been justified as the most efficient forensics tools for the challenges. However, they may be over-fitted to certain semantics, resulting in considerable inconsistency in detection performance across different contents of generated samples. It could be regarded as an issue of detection fairness. In this paper, we propose a novel framework named Fairadapter to tackle the issue. In comparison with existing state-of-the-art methods, our model achieves improved fairness performance. Our project is vailable at https://github.com/AppleDogDog/FairnessDetection
DeepFake, an AI technology that can automatically synthesize facial forgeries, has recently attracted worldwide attention. While DeepFakes can be entertaining, they can also be used to spread falsified information or be weaponized as cognition warfare. Forensic researchers have been dedicated to designing defensive algorithms to combat such disinformation. However, attacking technologies have been developed to make DeepFake products more aggressive. For example, by launching anti-forensics and adversarial attacks, DeepFakes can be disguised as authentic media to evade forensic detectors. However, such manipulations often sacrifice image quality for satisfactory undetectability. To address this issue, we propose a method to generate a novel adversarial sharpening mask for launching black-box anti-forensics attacks. Unlike many existing methods, our approach injects perturbations that allow DeepFakes to achieve high anti-forensics performance while maintaining pleasant sharpening visual effects. Experimental evaluations demonstrate that our method successfully disrupts state-of-the-art DeepFake detectors. Moreover, compared to images processed by existing DeepFake anti-forensics methods, our method’s quality of anti-forensics DeepFakes rendered is significantly improved. Our code is available at https://github.com/fb-reps/HQ-AF_GAN.
Pradeep Atrey合作论文数The University of Winnipeg;Department of Applied Computer Science3