Counterfeiting affects diverse industries, including pharmaceuticals, electronics, and food, posing serious health and economic risks. Printable unclonable codes, such as Copy Detection Patterns (CDPs), are widely used as an anti-counterfeiting measure and are applied to products and packaging. However, the increasing availability of high-resolution printing and scanning devices, along with advances in generative deep learning, undermines traditional authentication systems, which often fail to distinguish high-quality counterfeits from genuine prints. In this work, we propose a diffusion-based authentication framework that jointly leverages the original binary template, the printed CDP, and a representation of printer identity that captures relevant semantic information. Formulating authentication as multi-class printer classification over printer signatures lets our model capture fine-grained, device-specific features via spatial and textual conditioning. We extend ControlNet by repurposing the denoising process for class-conditioned noise prediction, enabling effective printer classification. On the Indigo 1 x 1 Base dataset, our method outperforms traditional similarity metrics and prior deep learning approaches. Results show the framework generalises to counterfeit types unseen during training.
The wide availability of modern tools that enable the creation of high-quality facial deepfake videos raises significant concerns about the trustworthiness of online content, particularly in an era of ubiquitous social media. As deepfake detection becomes more challenging, the development of facial deepfake detection methods that can reliably detect tampered videos becomes more important. In this paper, we introduce HFVideoSwin, a novel architecture for deepfake video detection that uses spatio-temporal attention through Video Swin Transformer backbones as well as a discrete wavelet transform pipeline for the extraction of meaningful forgery clues through high-frequency features. We extensively evaluate our method on a wide range of deepfake video datasets, focusing on the challenging generalization scenario. To ensure fair and reliable results, the rigorous training and testing protocol of DeepfakeBench, a popular deepfake detection benchmark, is used. Experimental evaluation results against state-of-the-art image- and video-based deepfake detection architectures demonstrate the efficiency and competitive performance of our method.
Reconstruction-based generative models offer a natural framework for unsupervised out-of-distribution (OOD) detection, but multi-class normality modelling requires a single detector to capture multiple in-distribution manifolds and produce comparable anomaly scores across classes. We study this problem in copy detection pattern (CDP) authentication, where authentic and counterfeit samples are visually similar but differ in subtle printing-and-digitisation (P&D) signatures. We propose a diffusion based multi-class normality framework in which a single class-conditional ControlNet is trained exclusively on authentic CDPs from multiple P&D classes and detects counterfeits through reconstruction error under authentic-class conditioning. We further introduce dual template masking, which hides complementary regions of the input template and scores only withheld pixels, reducing reliance on visible binary structure. On the Indigo 1 x 1 Base dataset, the proposed method outperforms traditional and adapted generative baselines under multi-class authentic-versus-counterfeit evaluation, without using counterfeit samples for training or threshold calibration.
Objects present in paintings help art history specialists interpret and decode artworks. The analysis of large, digitized artistic collections became feasible thanks to modern object detection approaches. Nevertheless, the use of object detection models typically requires fine-tuning for specific tasks. Therefore, art history specialists remain constrained by the categories of objects in existing labeled artistic datasets when using artificial intelligence methods. This limitation can be overcome by using recent models that combine two modalities: vision and text. Vision-language models have made open-vocabulary detection (OVD) possible, allowing detection without restrictions on the applied categories, in contrast to fixed-vocabulary detection. Recent literature lacks a comprehensive review focusing on OVD in artistic images. In this paper we analyze state-of-the-art models for OVD, analyze their transferability to cultural heritage categories and systematically evaluate them on artistic datasets commonly used in literature. The DEArt and IconArt datasets, which are annotated with cultural heritage-specific categories contain paintings from the 11th to the 20th century. While the Watercolor2K dataset, annotated with common object categories consists of watercolor paintings. Based on our analysis, the OWLv2 model achieved the best performance in both object detection and grounding task scenarios on these datasets. Additionally, we discuss existing challenges of open-vocabulary segmentation in artistic images and future tasks.
While studying objects presented in paintings, art history specialists identify their significance, symbolic meaning and historical context. Analyzing big artistic collections can be very time-consuming for the specialists. The search could be relieved by using modern object detectors. However, object detectors have poor performance on artistic images. This problem could be solved by fine-tuning them on specialized annotated datasets. In this paper, we explore the possibilities of using open-vocabulary foundation models for dataset annotation in a semi-automated manner. We propose an approach for artistic dataset annotation for object detection task based on a small set of images annotated on image-level and using Vision Transformer for Open-World Localization (OWL-ViT2) model, the YOLO object detector and an approximate nearest neighbour oh yeah (ANNOY) algorithm. We extend the existing DEArt dataset by 97.2% and introduce the way of adding new classes without exhaustive annotation. With the extended version of the dataset, we achieve 12.2% increase of mAP0.5 metric on average on the test data compared to the model trained on the original dataset.
Museums and art galleries use modern technologies to engage visitors and attract not only specialists but also the general public, by providing recommendations for interactive studying and observing their collections. These recommendations can be created on the base of artwork similarity defined by using fine-tuned deep neural networks. In this paper, we explore the possibilities of fine-tuning foundation models using the Low-Rank-Adaptation (LoRA) fine-tuning technique for the classification task and perform a similarity search based on features extracted with fine-tuned models. Using LoRA technique allows to use and switch easily between several fine-tuned models with only one frozen backbone, this is useful in the case of using large models on mobile devices with limited storage space. During fine-tuning, we examined the influence of hyperparameters on two DINOv2 models’ performance and found their reasonable combination according to a number of trainable parameters and performance, achieved on the relatively small artistic dataset. We achieved state-of-the-art accuracy for genre classification on the WikiArt dataset with the proposed approach.
The ease of use and wide availability of high-quality deepfake creation tools raises significant concerns about the reliability and trustworthiness of online content, and makes the task of detecting facial tampering more complicated. As such, the development of effective deepfake detection methods is of utmost importance. In recent years, the facial deepfake detection task took a leap thanks to the development of deep learning-based methods as well as the availability of large datasets of high-quality deepfake videos. Despite the aforementioned methods achieving excellent results when tasked with detecting deepfakes generated using methods seen during training, the cross-manipulation, or generalization, task—where a trained model is exposed to unseen manipulation techniques—is a major challenge which is attracting the attention of the research community. In this paper, we introduce WaveConViT, a novel spatio-temporal architecture for deepfake detection based on Vision Transformers and a two-dimensional discrete wavelet transform. Additionally, we introduce and evaluate a temporal sampling strategy based on frame skipping. We extensively test and benchmark this architecture in the challenging cross-manipulation scenario on the FaceForensics++, Celeb-DF, and DeeperForensics-1.0 datasets, comparing it to a selection of modern, representative Vision Transformer (ViT) and convolutional neural network (CNN) architectures and demonstrating the value of high-frequency features as well as our frame skipping strategy for deepfake detection.
A number of bridges have collapsed around the world over the past years, with detrimental consequences on safety and traffic. To a large extend, such failures can be prevented by regular bridge inspections and maintenance, tasks that fall in the general category of structural health monitoring (SHM). Those procedures are time and labor consuming, which partly accounts for their neglect. Computer vision and artificial intelligence (AI) methods have the potential to ease this burden, by fully or partially automating bridge monitoring. A critical step in this automation is the identification of a bridge’s structural components. In this work, we propose an extensible synthetic dataset for structural component semantic segmentation of portal frame bridges (PFBridge). We first create a 3 dimensional (3D) generic mesh representing the bridge geometry, while respecting a set of rules. The definition of new, or the extension of the existing rules can adjust the dataset to specific needs. We then add textures and other realistic elements to the model, and create an automatically annotated synthetic dataset. The synthetic dataset is used in order to train a deep semantic segmentation model to identify bridge components on bridge images. The amount of available real images is not sufficient to entirely train such a model, but is used to refined the model trained on the synthetic data. We evaluate the contribution of the dataset to semantic segmentation by training several segmentation models on almost 2,000 synthetic images and then finetuning with 88 real images. The results show an increase of 28
The high visual quality of modern deepfakes raises significant concerns about the trustworthiness of digital media and makes facial tampering detection more challenging. Although current deep learning-based deepfake detectors achieve excellent results when tested on deepfake images or image sequences generated using known methods, generalization—where a trained model is tasked with detecting deepfakes created with previously unseen manipulation techniques—is still a major challenge. In this paper, we investigate the impact of training spatial and spatio-temporal deep learning network architectures in the image noise residual domain using spatial rich model (SRM) filters on generalization performance. To this end, we conduct a series of tests on the manipulation methods of the FaceForensics++, DeeperForensics-1.0 and Celeb-DF datasets, demonstrating the value of image noise residuals and temporal feature exploitation in tackling the generalization task.
The number of medicine counterfeits increases each year due to the accessibility of printing devices and the weak protection of medicine blister foils.The medicine blisters are often produced using the rotogravure printing process.In this paper, we address the problem of rotogravure press identification and printed support identification using similarity metric learning.Both identification problems are difficult as the impact of printing press or of printing support are minimal, moreover the classical techniques (for example, the use of Pearson correlation) cannot identify the rotogravure press or the printing support used for the packaging production.We show that the similarity metric learning can easily identify the press used and the printing support used.Additionally, we explore the possibility to use the proposed approach for packaging authentication.
Modern content-based image retrieval systems demonstrate rather good performance in identifying visually similar artworks. However, this task becomes more challenging when art history specialists aim to refine the list of similar artworks based on their criteria, thus we need to train the model to reproduce this refinement. In this paper, we propose an approach for improving the list of similar paintings according to specific simulated criteria. By this approach, we retrieve paintings similar to a request image using ResNet50 model and ANNOY algorithm. Then, we simulate re-ranking based on the two criteria, and use the re-ranked lists for training LambdaMART model. Finally, we demonstrate that the trained model reproduces the re-ranking for the query painting by the specific criteria. We plan to use the proposed approach for reproducing re-rankings made by art history specialists, when this data will be collected.
Due to development and broad availability of high-quality printing and scanning devices, the number of counterfeited products and documents is dramatically increasing. Therefore, different security elements have been suggested to prevent this socioeconomic plague. One of the most promising and cheap solutions is the use of Copy Detection Pattern (CDP), a maximum entropy image, generated using a secret key. This pattern takes full advantage of information loss principle during printing-and-digitization process to detect copies. Such an unpredictable pattern is highly sensitive to distortions occurring inevitably during production (printing), verification (digitization) and reproduction (duplication) processes. Initially, the detection of counterfeited CDP was devoted to evaluating the level of information loss using Pearson correlation. However, the security of CDP based authentication system was shown to be vulnerable to estimation attacks based on neural network that can infer a CDP after scanning. In this paper, we study how to increase the performance of a detector using a similarity metric learning approach.
The medicine falsification is an important problem nowadays which represents a real danger for human lives. Therefore, it is important to find a cheap and efficient solution that can be used for all kinds of medicines. In this paper, we explore the domain of pharmaceutical packaging printed on blister foils using rotogravure process. It was shown that the chemical etching has a stochastic nature that can be spotted by correlation or classical machine learning methods. We propose an authentication system that uses a novel regular test pattern for authentication of blister foils. Thanks to this regular test pattern we can identify the cylinder used for printing and the position of the regular test pattern engraved on the authentic cylinder, as well as easily reject the fake patterns printed using counterfeiter cylinder and rotogravure press. The proposed identification/authentication system cannot be easily attacked as it is not possible to imitate the signature of chemical etching process. Additionally, we will shortly discussed the possibility to enlarge the proposed system to another types of engraving processes and formulate some future paths for this work
Nowadays, with the use of photo-editing software being mainstream, document integrity verification has become crucial. As we have seen during the pandemic, most administrative documents are printed and then scanned before being transmitted, making these documents noisy. Indeed, a printed and scanned document undergoes geometric transformations, as well as the addition of black spots, not to mention a decrease in color intensity. The relevant features of an original document, which will be matched against a query document, are stored to be used as a template. We propose a 2-step method that compares a template with a query document to ensure that the query document has not been tampered with. Our method first reverts geometric transformations the document underwent, and then extracts the crossing numbers in that image. A Euclidean distance based matching method is applied to the two sets of crossing numbers, and abnormally distant point groups are flagged as potentially modified. A second step in our method is then applied to analyze the statistical properties of these distance values, to ensure that the document has not been altered. Our results when we apply our method to a database containing administrative documents and tampered versions of these documents - all of which underwent a print and scan process - show the validity of our considerations.
Ce chapitre dresse un panorama des approches en authentification de documents matérialisés avant de traiter de la protection par anticopie. Il passe en revue les différents types de dispositifs anticopies de l'état de l'art en distinguant les formes-tests soumis à l’impression et les codes sensibles à la copie. Les procédés d’authentification associés, améliorés/améliorables par des apprentissages profonds, sont conjointement décrits.
Nowadays, the security of both digital and hard-copy documents has become a real issue. As a solution, numerous integrity check approaches have been designed. The challenge lies in finding features which are robust to print-and-scan process. In this paper, we propose a new method of printed-and-scanned character matching based on the adaptation of biometrical features. After the binarization and the skeletonization of a character, feature points are extracted by computing crossing numbers. The feature point set can then be smoothed to make it more suitable for template matching. From various experimental results, we have shown that an accuracy of more than 95% is achieved for print-and-scan resolutions of 300 dpi and 600 dpi. We have also highlighted the feasibility of the proposed method in case of double print-and-scan operation. The comparison with a state-of-the-art method shows that the generalization of proposed matching method is possible while using different fonts.
Copy Detection Patterns (CDP) have received significant attention from academia and industry as a practical mean of detecting counterfeits. Their security level against sophisticated attacks has been studied theoretically and practically in different research papers, but for reasons that will be explained below, the results are not fully conclusive. In addition, the publicly available CDP datasets are not practically usable to evaluate the performance of authentication algorithms. In short, the apparently simple question: “are copy detection patterns secure against copy?”, remains unanswered as of today. The primary contribution of this paper is to present a publicly available dataset of CDPs including multiple types of copies and attacks, allowing to systematically compare the performance level of CDPs against different attacks proposed in the prior art. The specific case in which a CDP is the same for an entire batch of prints, which is of practical importance as it covers applications with widely used industrial printers such as offset, flexo and rotogravure, is also studied. A second contribution is to highlight the role played by the CDP detector and its different processing steps. Indeed, depending on the specific processing involved, the detection performance can widely outperform the CDP bit error rate which has been used as a reference metrics in the prior art.
The number of medicine counterfeits increases significantly. This problem affects not only expensive medicines, but also some low cost ones. In this paper, we study the characteristics of medicine packages printed using rotogravure printing on blister foils and propose an authentication system that identifies the equipment used for printing medicine foils. The rotogravure printing process uses an engraved cylinder and a rotogravure press. Each of these elements has its own signature that can be used for process identification and for packaging authentication. Using constructed database, we show that the signature of engraved cylinder impacts more on printed patterns in comparison with the signature of rotogravure press. The experiments done show that we can identify the cylinder used for the printing using a classical machine learning methods from a small number of training samples.
Olivier Strauss合作论文数Universite Montpellier II10
Serge Miguet合作论文数des Universit??s8
Vasile-Marian Scuturici合作论文数1