AbstractIn recent years there have been astonishing advances in AI-based synthetic media generation. Thanks to deep learning methods it is now possible to generate visual data with a high level of realism. This is especially true for human faces. Advanced deep learning tools allow one to easily change some specific attributes of a real face or even create brand new identities. Although this opens up a large number of new opportunities, just think of the entertainment industry, it also undermines the trustworthiness of media content and supports the spread of fake identities over the internet. In this context, there is a fundamental need to develop robust and automatic tools capable of distinguishing synthetic faces from real ones. The scientific community is making a huge research effort in this field, proposing several interesting approaches. However, a universal detector is yet to come. Fundamentally, the research in this field is like a cat and mouse game, with new detectors that are designed to deal with powerful synthetic face generators, while the latter keep improving to produce more and more realistic images. In this chapter we will present the most effective techniques proposed in the literature for the detection of synthetic faces. We will analyze their rationale, present real-world application scenarios , and compare different approaches in terms of accuracy and generalization ability.
This work evaluates the use of synthetic data to train deep 6DoF pose estimation models that use a monocular RGB camera as input. We have compared different training strategies combining real and synthetic data (with domain randomization) to investigate how to better handle real-world challenges. We show that it is possible to obtain accurate models using less real data and suggest how to utilize this strategy. In this work, we have captured and made available two datasets: one real and one synthetic, totaling over 110,000 annotated frames. These datasets are organized according to the different cameras used and the challenges present in the sequences, all featuring textureless 3D printed objects. We also show that synthetic data can help models generalize, handling challenges such as fast motion, occlusion, illumination changes, color variation, scale changes, and unexpected geometry. Finally, we evaluated 70 different models to understand how a model trained for one camera sensor performs when used with a different sensor. To this end, we also suggest how to handle this challenge better by using synthetic simulations to supplement training.
Deep neural networks provide unprecedented performance in all image classification problems, taking advantage of huge amounts of data available for training. Recent studies, however, have shown their vulnerability to adversarial attacks, spawning an intense research effort in this field. With the aim of building better systems, new countermeasures and stronger attacks are proposed by the day. On the attacker's side, there is growing interest for the realistic black-box scenario, in which the user has no access to the neural network parameters. The problem is to design efficient attacks which mislead the neural network without compromising image quality. In this work, we propose to perform the black-box attack along a low-distortion path, so as to improve both the attack efficiency and the perceptual quality of the adversarial image. Numerical experiments on real-world systems prove the effectiveness of the proposed approach, both in benchmark classification tasks and in key applications in biometrics and forensics.
The advent of deep learning has brought a significant improvement in the quality of generated media. However, with the increased level of photorealism, synthetic media are becoming hardly distinguishable from real ones, raising serious concerns about the spread of fake or manipulated information over the Internet. In this context, it is important to develop automated tools to reliably and timely detect synthetic media. In this work, we analyze the state-of-the-art methods for the detection of synthetic images, highlighting the key ingredients of the most successful approaches, and comparing their performance over existing generative architectures. We will devote special attention to realistic and challenging scenarios, like media uploaded on social networks or generated by new and unseen architectures, analyzing the impact of suitable augmentation and training strategies on the detectors' generalization ability.
PRNU-based image processing is a key asset in digital multimedia forensics. It allows for reliable device identification and effective detection and localization of image forgeries, in very general conditions. However, performance impairs significantly in challenging conditions involving low quality and quantity of data. These include working on compressed and cropped images, or estimating the camera PRNU pattern based on only a few images. To boost the performance of PRNU-based analyses in such conditions we propose to leverage the image noiseprint, a recently proposed camera-model fingerprint that has proved effective for several forensic tasks. Numerical experiments on datasets widely used for source identification prove that the proposed method ensures a significant performance improvement in a wide range of challenging situations.
Due to limited computational and memory resources, current deep learning models accept only rather small images in input, calling for preliminary image resizing.This is not a problem for high-level vision problems, where discriminative features are barely affected by resizing.On the contrary, in image forensics, resizing tends to destroy precious high-frequency details, impacting heavily on performance.One can avoid resizing by means of patch-wise processing, at the cost of renouncing whole-image analysis.In this work, we propose a CNN-based image forgery detection framework which makes decisions based on fullresolution information gathered from the whole image.Thanks to gradient checkpointing, the framework is trainable end-to-end with limited memory resources and weak (image-level) supervision, allowing for the joint optimization of all parameters.Experiments on widespread image forensics datasets prove the good performance of the proposed approach, which largely outperforms all baselines and all reference methods.
In the last few years, generative adversarial networks (GAN) have shown tremendous potential for a number of applications in computer vision and related fields. With the current pace of progress, it is a sure bet they will soon be able to generate high-quality images and videos, virtually indistinguishable from real ones. Unfortunately, realistic GAN-generated images pose serious threats to security, to begin with a possible flood of fake multimedia, and multimedia forensic countermeasures are in urgent need. In this work, we show that each GAN leaves its specific fingerprint in the images it generates, just like real-world cameras mark acquired images with traces of their photo-response non-uniformity pattern. Source identification experiments with several popular GANs show such fingerprints to represent a precious asset for forensic analyses.
Current developments in computer vision and deep learning allow to automatically generate hyper-realistic images, hardly distinguishable from real ones. In particular, human face generation achieved a stunning level of realism, opening new opportunities for the creative industry but, at the same time, new scary scenarios where such content can be maliciously misused. Therefore, it is essential to develop innovative methodologies to automatically tell apart real from computer generated multimedia, possibly able to follow the evolution and continuous improvement of data in terms of quality and realism. In the last few years, several deep learning-based solutions have been proposed for this problem, mostly based on Convolutional Neural Networks (CNNs). Although results are good in controlled conditions, it is not clear how such proposals can adapt to real-world scenarios, where learning needs to continuously evolve as new types of generated data appear. In this work, we tackle this problem by proposing an approach based on incremental learning for the detection and classification of GAN-generated images. Experiments on a dataset comprising images generated by several GAN-based architectures show that the proposed method is able to correctly perform discrimination when new GANs are presented to the network.
The aim of this paper is to propose an algorithm based on convolutional neural networks (CNN) for iris sensor model identification. This task is important in forensics applications as well as to face the problem of sensor interoperability in large scale systems. When different sensor models are involved in a recognition system, in fact, the overall performance can strongly decrease. A possible solution consists in first identifying the sensor model and then mapping the features extracted from the image from one sensor to the other. To keep low both complexity and memory requirements we propose a simple network architecture and the use of transfer learning to speed-up the training phase and tackle the problem of limited training set availability. Experiments are carried out on several public iris databases. First, we show that the proposed solution outperforms the state-of-the art approaches used for the model identification task. Then, we test the performance of a biometric recognition system and show that improving the sensor model identification step can benefit the iris sensor interoperability. (C) 2017 Elsevier B.V. All rights reserved.
Camera model identification is a fundamental task for many investigative activities, and is drawing great attention in the research community. In this context, convolutional neural networks (CNN) are expected to provide a significant performance gain over the current state of the art, as already happened for a wide range of image processing applications. However, recent studies enlightened the vulnerability of CNNs to adversarial attacks, casting shadows on their reliability for critical applications. In this paper, we investigate the robustness to adversarial attacks of CNN-based methods for camera model identification. Several networks and attack methods are considered, both when the attacker has complete knowledge of the network and when only the training set is available. In addition, the analysis concerns both original and JPEG compressed images, to simulate a social network environment. The experiments, carried out on a publicly available dataset with images coming from 29 different camera models, shed some light on the suitability of CNN-based approaches for this task.
With the ubiquitous diffusion of social networks, images are becoming a dominant and powerful communication channel. Not surprisingly, they are also increasingly subject to manipulations aimed at distorting information and spreading fake news. In recent years, the scientific community has devoted major efforts to contrast this menace, and many image forgery detectors have been proposed. Currently, due to the success of deep learning in many multimedia processing tasks, there is high interest towards CNN-based detectors, and early results are already very promising. Recent studies in computer vision, however, have shown CNNs to be highly vulnerable to adversarial attacks, small perturbations of the input data which drive the network towards erroneous classification. In this paper we analyze the vulnerability of CNN-based image forensics methods to adversarial attacks, considering several detectors and several types of attack, and testing performance on a wide range of common manipulations, both easily and hardly detectable.
The diffusion of fake images and videos on social networks is a fast growing problem. Commercial media editing tools allow anyone to remove, add, or clone people and objects, to generate fake images. Many techniques have been proposed to detect such conventional fakes, but new attacks emerge by the day. Image-to-image translation, based on generative adversarial networks (GANs), appears as one of the most dangerous, as it allows one to modify context and semantics of images in a very realistic way. In this paper, we study the performance of several image forgery detectors against image-to-image translation, both in ideal conditions, and in the presence of compression, routinely performed upon uploading on social networks. The study, carried out on a dataset of 36302 images, shows that detection accuracies up to 95% can be achieved by both conventional and deep learning detectors, but only the latter keep providing a high accuracy, up to 89%, on compressed data.
The Photo Response Non-Uniformity (PRNU) noise can be regarded as a camera fingerprint and used, accordingly, for source identification, device attribution and forgery localization. To accomplish these tasks, the camera PRNU is typically assumed to be known in advance or reliably estimated. However, there is a growing interest for methods that can work in a real-word scenario, where these hypotheses do not hold anymore. In this paper we analyze a PRNU-based framework for forgery localization in a blind scenario. The framework comprises four main steps: PRNU-based blind image clustering, parameter estimation, device attribution, and forgery localization. Each of these steps impacts on the final outcome of the analysis. The aim of this paper is to assess the overall performance of the proposed framework and how it depends on the individual steps.
Camera model identification has great relevance for many forensic applications, and is receiving growing attention in the literature. Virtually all techniques rely on the traces left in the image by the long sequence of in-camera processes which are specific of each model. They differ in the prior assumptions, if any, and in how such evidence is gathered in expressive features. In this work we study a class of blind features, based on the analysis of the image residuals of all color bands. They are extracted locally, based on co-occurrence matrices of selected neighbors, and then used to train a classifier. A number of experiments are carried out on the well-known Dresden Image Database. Besides the full-knowledge case, where all models of interest are known in advance, other scenarios with more limited knowledge and partially corrupted images are also investigated. Experimental results show these features to provide a state-of-the-art performance.
We propose a new algorithm for blind camera identification, based on the Photo-Response Non-Uniformity (PRNU) noise estimated by image residuals. Successful identification relies on the correct clustering of residuals coming from the same camera. We adopt a two-step strategy. First, residuals are efficiently grouped by correlation clustering, setting parameters so as to over-partition the data points and avoid any wrong associations. Then, basic clusters are progressively merged by an ad hoc refinement algorithm. Experiments on the Dresden database prove the effectiveness of the proposed method.
With the powerful image editing tools available today, it is very easy to create forgeries without leaving visible traces. Boundaries between host image and forgery can be concealed, illumination changed, and so on, in a naive form of counter-forensics. For this reason, most modern techniques for forgery detection rely on the statistical distribution of micro-patterns, enhanced through high-level filtering, and summarized in some image descriptor used for the final classification. In this work we propose a strategy to modify the forged image at the level of micro-patterns to fool a state-of-the-art forgery detector. Then, we investigate on the effectiveness of the proposed strategy as a function of the level of knowledge on the forgery detection algorithm. Experiments show this approach to be quite effective especially if a good prior knowledge on the detector is available.
Camera model identification is of interest for many applications. In-camera processes, specific of each model, leave traces that can be captured by features designed ad hoc , and used for reliable classification. In this work we investigate on the use of blind features based on the analysis of image residuals. In particular, features are extracted locally based on co-occurrence matrices of selected neighbors and then used to train an SVM classifier. Experiments on the well-known Dresden database show this approach to provide state-of-the-art performances.
Digital camera identification is a very active research area, with important applications in the forensics field. Several approaches have been proposed in recent years for this task. One of the most promising is based on the estimation of the sensor noise pattern, used as a sort of camera fingerprint. However, a clever attacker can estimate a camera fingerprint and use it maliciously: this calls for new countermeasures, and so on, in a typical two-party game. In this paper we consider the triangle test, a well-know countermeasure against fake fingerprint attacks, and propose a new algorithm for improving the attacker's success rate. Numerical experiments show that, in typical scenarios, the proposed algorithm improves significantly the attacker performance.