Among the severe threats to Face Recognition Systems (FRS), morphs attacks remain challenging and difficult to mitigate. This difficulty stems from the high realism in morphed images, which can deceive both human operators and standard FRS. In this paper, we propose to detect morphed face images as anomalies. Specifically, instead of training a binary classifier with bona fide and morphed images, our method only requires learning the data distribution of those samples that are deemed to be normal; i.e. a set of bona fide images. Once trained, morphs are detected as outliers or abnormal samples. Our method relies on the analysis of the frequency power spectrum to generate a descriptive latent space, which is then reduced via autoencoders into compact and descriptive embeddings. These embeddings are used to generate probability distributions of the observed data. Our method manages to detect morphs generated by several common techniques without prior knowledge of their generation process and without relying on any training morphed images. It achieves state-of-the-art performance on the FRLL dataset and very competitive performance on the SMDD dataset and the Morphed Flick Faces High Quality (MFFHQ), which is a new dataset of morphs generated from the FFHQ dataset. Our evaluation results confirm the strength and practicality of the anomaly detection framework in the compressed frequency domain for the morph attack detection task.
Face detection technology is critical in video surveillance applications. Unfortunately, many of the existing approaches fail to detect tiny faces, especially in videos acquired in uncontrolled public environments. In this work, we address this challenge by proposing a novel detection framework that relies on a new multiscale detector that does not require the input to be re-scaled. Our framework leverages existing detectors to detect relatively large faces along with a new detector specifically designed for tiny faces. The core part of our framework is the use of an integral score to detect tiny faces by densely scanning several support regions at different scales. Although this strategy may pose high computational complexity in high-resolution videos, the integral operations used by our framework significantly reduce the computational cost, making it feasible for high-resolution videos and real-time applications. Experiments on the WIDER FACE and MEVA datasets show a significant improvement in performance, particularly for tiny faces depicted in surveillance videos acquired in uncontrolled environments.
Face Recognition Systems (FRS) are critical and essential components for user authentication via biometrics. To name a few, baking, e-Commerce, and border control are entities propelling their progress. These are of immense importance due to their economic and social relevance. FRS widespread usage leads to security vulnerabilities that need to be identified and mitigated. This paper provides a comprehensive review of potential attacks on recently discovered vulnerabilities from 2017–2024. Our work is significant regarding FRS development because their impact in terms of security. The novelty is a systematic review to properly categorize threat vectors and their severity towards FRS over the past eight years. We categorize, summarize, and analyze the threat vectors towards FRS to this end. We also elaborate on the threat taxonomy for existing Architecture Reference Architecture (ARA) to identify threats on user-based authentication FRS. Our findings show the most persistent attack vectors, usage trends, severity, functionality, and level of sophistication required to perform them. We present a comprehensive description of each to create more resilient and trustable systems for this fast-growing technology. This paper can be used by researchers and practitioners interested in the state-of-the-art FRS attack vectors to develop more secure systems.
Face image synthesis is gaining more attention in computer security due to concerns about its potential negative impacts, including those related to fake biometrics. Hence, building models that can detect the synthesized face images is an important challenge to tackle. In this paper, we propose a fusion-based strategy to detect face image synthesis while providing resiliency to several attacks. The proposed strategy uses a late fusion of the outputs computed by several undisclosed models by relying on random polynomial coefficients and exponents to conceal a new feature space. Unlike existing concealing solutions, our strategy requires no quantization, which helps to preserve the feature space. Our experiments reveal that our strategy achieves state-of-the-art performance while providing protection against poisoning, perturbation, backdoor, and reverse model attacks.
Face image synthesis detection is considerably gaining attention because of the potential negative impact on society that this type of synthetic data brings. In this paper, we propose a data-agnostic solution to detect the face image synthesis process. Specifically, our solution is based on an anomaly detection framework that requires only real data to learn the inference process. It is therefore data-agnostic in the sense that it requires no synthetic face images. The solution uses the posterior probability with respect to the reference data to determine if new samples are synthetic or not. Our evaluation results using different synthesizers show that our solution is very competitive against the state-of-the-art, which requires synthetic data for training.
Face image synthesis has shown remarkable progress in recent years. However, the effect that the demographics of the data used to train synthesizers has on the generation of new face images remains an open question. This paper investigates the effects of the training set demographics in the face image synthesis task. To this end, we propose a strategy that allows synthesizing face images for specific groups of people with a high visual quality. The strategy uses an unsupervised learning approach to discover groups of people in the training set based on Bayesian inference via a probabilistic mixture model. If labels are available to define the groups, our strategy can also exploit such information in lieu of unsupervised learning. Once the groups are defined, our strategy trains a Generative Adversarial Network on each group to generate new face images with specific characteristics. Our results show remarkable performance in terms of image quality compared to several state-of-the-art baselines. More importantly, our strategy allows synthesizing face images with reduced demographic biases.
Although in recent years there has been a trend to incorporate data-driven models to automate the task of video surveillance, there remains a shortage of learning-based solutions for the detection of specific events that do not require examples of such events during training. Current research largely focuses on abnormal event detection, a learning-based approach that allows for the detection of events as anomalies. Unfortunately, such an approach does not provide any details about the types of events detected as the training data only includes normal events. This paper proposes an approach to detect specific events in surveillance videos under an anomaly detection framework by constraining the input space of the detection model, and hence, allowing to determine the type of event detected as an anomaly. Specifically, the proposed approach exploits the benefits of variational Bayesian inference to build probabilistic models for the detection of specific events as anomalies. Such a novel strategy leverages the benefits of the learning mechanism used by abnormal event detection frameworks to detect specific events. The proposed approach achieves a very competitive performance for the detection of several events within the context of video surveillance in public transport stations. It outperforms the state-of-the-art for the detection of pieces of abandoned luggage.
This paper presents a strategy to synthesize face images based on human traits. Specifically, the strategy allows synthesizing face images with similar age, gender, and ethnicity, after discovering groups of people with similar facial features. Our synthesizer is based on unsupervised learning and is capable to generate realistic faces. Our experiments reveal that grouping the training samples according to their similarity can lead to more realistic face images while having semantic control over the synthesis. The proposed strategy achieves competitive performance compared to the state-of-the-art and outperforms the baseline in terms of the Frechet Inception Distance.
HighlightsSynthesizing Faces with Demographic AttributesRoberto Leyva,Gregory Epiphaniou,Carsten Maple,Victor SanchezAbstract Face image synthesis has shown remarkable progress in recent years. However, the effect that the training set demographics has on the generation of new face images remains an open question, especially when there is an interest in generating face images with a wide range of facial characteristics. This question is quite important when severe imbalances are present in the training dataset, e.g., more samples for specific groups of people. This paper investigates the effects of the training set demographics in the face image synthesis task. To this end, we propose a strategy that allows synthesizing face images for specific groups of people with a high visual quality.The strategy first uses an unsupervised learning approach to discover groups of people in the training set based on Bayesian inference via a probabilistic mixture model. If labels are available to define the groups, our strategy can also exploit such information in lieu of unsupervised learning. Once the groups are defined, our strategy trains aa GAN synthesizer on each group to generate new face images with specific characteristics. Our results show remarkable performance in terms of image quality outperforming the state-of-the-art baselines. Finally, our strategy requires significant less training time and training samples, making it a step forward for face image generation with demographic attributes.
Video memorability is a cornerstone in social media platform analysis, as a highly memorable video is more likely to be noticed and shared. This paper proposes a new framework to fuse multi-modal information to predict the likelihood of remembering a video. The proposed framework relies on late fusion of text, visual and motion features. Specifically, two neural networks extract features from the captions describing the video’ s content; two ResNet models extract visual features from specific frames, and two 3DResNet models, combined with Fisher Vectors, extract features from the video’ s motion information. The extracted features are used to compute several memorability scores via Bayesian Ridge regression, which are then fused based on a greedy search of the optimal fusion parameters. Experiments demonstrate the superiority of the proposed framework on the MediaEval2019 dataset, outperforming the state-of-the-art.
•A solution for learning identify-related motion traits using user-agnostic models.•The proposed manifold-based solution outperformed previous, user-specific ones.•The proposed solution presented robustness in cross-dataset and open-set scenarios.•The proposed solution employs a small amount of target user data for personalization.•An extended dataset with motion data from 115 users is made available.
© 2019 for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0). This paper describes a multimodal feature fusion approach for predicting the short and long term video memorability where the goal to design a system that automatically predicts scores reflecting the probability of a video being remembered. The approach performs early fusion of text, image, and video features. Text features are extracted using a Convolutional Neural Network (CNN), an FBResNet152 pre-trained on ImageNet is used to extract image features and video features are extracted using 3DResNet152 pre-trained on Kinetics 400. We use Fisher Vectors to obtain a single vector associated with each video that overcomes the need for using a non-fixed global vector representation for handling temporal information. The fusion approach demonstrates good predictive performance and regression superiority in terms of correlation over standard features.
In this paper, we propose a compact and low-complexity binary feature descriptor for video analytics. Our binary descriptor encodes the motion information of a spatio-temporal support region into a low-dimensional binary string. The descriptor is based on a binning strategy and a construction that binarizes separately the horizontal and vertical motion components of the spatio-temporal support region. We pair our descriptor with a novel Fisher Vector (FV) scheme for binary data to project a set of binary features into a fixed length vector in order to evaluate the similarity between feature sets. We test the effectiveness of our binary feature descriptor with FVs for action recognition, which is one of the most challenging tasks in computer vision, as well as gait recognition and animal behavior clustering. Several experiments on the KTH, UCF50, UCF101, CASIA-B, and TIGdog datasets show that the proposed binary feature descriptor outperforms the state-of-the-art feature descriptors in terms of computational time and memory and storage requirements. When paired with FVs, the proposed feature descriptor attains a very competitive performance, outperforming several state-of-the-art feature descriptors and some methods based on convolutional neural networks.
This paper addresses the problem of activity recognition and people identification using accelerometer signals acquired by personal devices. Specifically, we propose a framework based on a Deep Neural Network that employs an efficient dense trajectory encoding to compute features. These Accelerometer Dense Trajectory (ADT) features, which are similar to those used for action recognition in the spatio-temporal domain of video data, densely map the accelerometer signals into three-dimensional normalised positions. To deal with the unordered nature and dimensional variation of trajectories associated with the classes, the proposed framework employs Fisher Vectors as a high order representation of the extracted features. We evaluate the proposed ADT features and framework on the Sphere2016 Challenge and WISDM datasets for activity recognition. For people identification, we employ the RecodGait dataset. For these two significantly different classification tasks, the performance evaluation results confirm the high descriptiveness of the proposed ADT features and the effectiveness of the proposed framework.
This paper describes a multimodal feature fusion approach for predicting the short and long term video memorability where the goal to design a system that automatically predicts scores reflecting the probability of a video being remembered. The approach performs early fusion of text, image, and video features. Text features are extracted using a Convolutional Neural Network (CNN), an FBResNet152 pre-trained on ImageNet is used to extract image features and video features are extracted using 3DResNet152 pre-trained on Kinetics 400. We use Fisher Vectors to obtain a single vector associated with each video that overcomes the need for using a non-fixed global vector representation for handling temporal information. The fusion approach demonstrates good predictive performance and regression superiority in terms of correlation over standard features.
Millions of surveillance cameras are currently installed in public places around the world, making it necessary to intelligently analyse the acquired data to detect the occurrence of abnormal events. A vast number of methods to detect such events have been recently proposed; unfortunately, there is a lack of methods capable of detecting these events as frames are acquired, also known as online processing. In this paper, we present an online framework for video anomaly detection that employs binary features to encode motion information, and low-complexity probabilistic models for detection. Evaluation results on the popular UCSD dataset and on a recently introduced real-event video surveillance dataset show that our framework outperforms non-online and online methods.
This paper proposes a view-invariant gait recognition algorithm, which builds a unique view invariant model taking advantage of the dimensionality reduction provided by the Direct Linear Discriminant Analysis (DLDA). Proposed scheme is able to reduce the under-sampling problem (USP) that appears usually when the number of training samples is much smaller than the dimension of the feature space. Proposed approach uses the Gait Energy Images (GEIs) and DLDA to create a view invariant model that is able to determine with high accuracy the identity of the person under analysis independently of incoming angles. Evaluation results show that the proposed scheme provides a recognition performance quite independent of the view angles and higher accuracy compared with other previously proposed gait recognition methods, in terms of computational complexity and recognition accuracy.
Nowadays, big imaging data are very common in many fields of study. As a result, detecting small objects in very large images is challenging and computationally demanding. Taking advantage of the intrinsic cumulative properties of the Fisher Score, we propose the Integral Fisher Score (IFS) for lowcomplexity and accurate object detection in big imaging data. The IFS, which is a multi-dimensional extension of the Integral Image, allows computing the Fisher Vector associated with a spatial region using only four operations. This considerably reduces the computational cost of searching for a small query object on a very large target image. Evaluations for the detection of small object on high-resolution HUB telescope and digital pathology images show that IFS attains a high accuracy with short processing times.