In this paper, we address the anomaly detection problem in the context of heterogeneous normal observations and propose an approach that accounts for this heterogeneity. Although prediction-based methods are common to learn normality, the vast majority of previous work predicts a single outcome, which is generally not sufficient to account for the multiplicity of possible normal observations. To address this issue, we introduce a new masked multi-prediction (MMP) approach that produces multiple likely normal outcomes, and show both theoretically and experimentally that it improves normality learning and leads to a better anomaly detection performance. In addition, we observed that normality can be characterized from multiple aspects, depending on the types of anomalies to be detected. Therefore, we propose an adaptation (MMP-AMS) of our approach to cover multiple aspects of normality such as appearance, motion, semantics and location. Since we model each aspect separately, our approach has the advantage of being interpretable and modular, as we can select only a subset of normality aspects. The experiments conducted on several benchmarks show the effectiveness of the proposed approach.
The ongoing research on human action recognition models is achieving very promising results, and the existing models reach very high performances. However, they still suffer from one major challenge: their performance decreases on viewpoints not seen in the training step of the model. In this paper, we introduce a new approach based on virtual viewpoint augmentation in the feature space to increase the robustness of the action recognition models to different camera viewpoints. This approach was evaluated on two action recognition datasets: DAHLIA and Toyota SmartHome. Our model shows promising results, with a significant performance increase on both datasets for viewpoints not seen during the training step.
This paper addresses video anomaly detection problem for videosurveillance. Due to the inherent rarity and heterogeneity of abnormal events, the problem is viewed as a normality modeling strategy, in which our model learns object-centric normal patterns without seeing anomalous samples during training. The main contributions consist in coupling pre-trained object-level action features prototypes with a cosine distance-based anomaly estimation function, therefore extending previous methods by introducing additional constraints to the mainstream reconstruction-based strategy. Our framework leverages both appearance and motion information to learn object-level behavior and captures prototypical patterns within a memory module. Experiments on several well-known datasets demonstrate the effectiveness of our method as it outperforms current state-of-the-art on most relevant spatio-temporal evaluation metrics.
Abnormal event detection in videos is a challenging problem, partly due to the multiplicity of abnormal patterns and the lack of their corresponding annotations. In this paper, we propose new constrained pretext tasks to learn object level normality patterns. Our approach consists in learning a mapping between down-scaled visual queries and their corresponding normal appearance and motion characteristics at the original resolution. The proposed tasks are more challenging than reconstruction and future frame prediction tasks which are widely used in the literature, since our model learns to jointly predict spatial and temporal features rather than reconstructing them. We believe that more constrained pretext tasks induce a better learning of normality patterns. Experiments on several benchmark datasets demonstrate the effectiveness of our approach to localize and track anomalies as it outperforms or reaches the current state-of-the-art on spatio-temporal evaluation metrics.
We present a new framework for Patch Distribution Modeling, PaDiM, to concurrently detect and localize anomalies in images in a one-class learning setting. PaDiM makes use of a pretrained convolutional neural network (CNN) for patch embedding, and of multivariate Gaussian distributions to get a probabilistic representation of the normal class. It also exploits correlations between the different semantic levels of CNN to better localize anomalies. PaDiM outperforms current state-of-the-art approaches for both anomaly detection and localization on the MVTec AD and STC datasets. To match real-world visual industrial inspection, we extend the evaluation protocol to assess performance of anomaly localization algorithms on non-aligned dataset. The state-of-the-art performance and low complexity of PaDiM make it a good candidate for many industrial applications.
Spatial Augmented Reality applications generally use projector-camera systems to control the visual projection appearance by comparing the initial projected and the acquired images. To obtain an accurate geometric compensation, a non-intrusive feature-point matching approach can be exploited which must handle complex photometric distortions due to the spectral devices responses, complex illumination and the mixing of the projected image with the projection surface.This paper first discusses the invariance properties of existing color descriptors in that application for non-intrusive geometric compensation. Their performance is evaluated using the framework of Setkov et al. (2013) extended by adding the several new test cases: modeled synthetic projections, real-world projections under various illuminants on one and two planar surfaces. Our experimental results show two main conclusions: (1) classical color vision models are hardly suitable to model the distortions in a projector-camera system, and (2) the LHE-based descriptor (Local Histogram Equalization) is the most reliable to compensate real-projections. (C) 2016 Elsevier Inc. All rights reserved.
Projector-camera systems, used in Spatial Augmented Reality, automatically adapt the video projections to the scene objects according to the visualization conditions. This paper introduces a novel non-invasive (without Structured Light) method based on a combination of traditional Feature Matching (FM) and more computationally effective Optical Flow (OF). It requires only one projected and one acquired image at a time in the most difficult case when both projected content and geometric transformations change every frame. It detects scene changes when OF fails and thus should be replaced by FM. In the experiments, we show that the method yields a more precise and less shaky compensation for different types of projected videos, and is up to 2.8 times faster than previous FM-based works.
Projector-camera systems are designed to improve the projection quality by comparing original images with their captured projections, which is usually complicated due to high photometric and geometric variations. Many research works address this problem using their own test data which makes it extremely difficult to compare different proposals. This paper has two main contributions. Firstly, we introduce a new database of acquired image projections (DAcImPro) that, covering photometric and geometric conditions and providing data for ground-truth computation, can serve to evaluate different algorithms in projector-camera systems. Secondly, a new object recognition scenario from acquired projections is presented, which could be of a great interest in such domains, as home video projections and public presentations. We show that the task is more challenging than the classical recognition problem and thus requires additional pre-processing, such as color compensation or projection area selection.
The success of matching algorithms relies on the definition of features which are both invariant against the geometric distortions to be considered, and distinctive enough to avoid ambiguities. This paper addresses the problem of color feature points matching under photometric and geometric changes. Considering the popular SURF descriptor, it analyzes its state-of-the-art color versions, and proposes a new extension by using local histogram equalization (LHE). While most existing descriptors stem from color conversions and apply to standard lighting variations acquired by the same device, the proposed feature is device-independent and could fit to very generic changes. The experimental results show that the proposed color descriptors outperform the existing ones under some types of distortions, and are more precise and invariant to different color variations. The paper considers Projector-based Augmented Reality (PAR) as an application field, where one of the evaluation criteria is homography accuracy between real and estimated distorted images. The results show that the proposed method gives the most stable results over all the other techniques and therefore they justify its use for robust color feature matching and its application to geometric correction.
IVORA (Image et Vision par Ordinateur pour la Réalité Augmentée) : Invariance colorimétrique et correspondances pour la définition d'un système projecteur/caméra La Réalité Augmentée Spatiale (SAR) vise à superposer spatialement l'information virtuelle sur des objets physiques. Au cours des dernières décennies ce domaine a connu une grande expansion et est utilisé dans divers domaines, tels que la médecine, le prototypage, le divertissement etc. Cependant, pour obtenir des projections de bonne qualité, on doit résoudre plusieurs problèmes, dont les plus importants sont la gamme de couleurs réduite du projecteur, la lumière ambiante, la couleur du fond, et la configuration arbitraire de la surface de projection dans la scène. Ces facteurs entraînent des distorsions dans les images qui requièrent des processus de compensation complémentaires.Les projections intelligentes (smart projections) sont au cœur des applications de SAR. Composées d'un dispositif de projection et d'un dispositif d'acquisition, elles contrôlent l'aspect de la projection et effectuent des corrections à la volée pour compenser les distorsions. Bien que les méthodes actives de Lumière Structurée aient été utilisées classiquement pour résoudre ces problèmes de compensation géométrique, cette thèse propose une nouvelle approche non intrusive pour la compensation géométrique de plusieurs surfaces planes et pour la reconnaissance des objets en SAR s'appuyant uniquement sur la capture du contenu projeté.Premièrement, cette thèse étude l'usage de l'invariance couleur pour améliorer la qualité de la mise en correspondance entre primitives dans une configuration d'acquisition des images vidéoprojetées. Nous comparons la performance de la plupart des méthodes de l'état de l'art avec celle du descripteur proposé basé sur l'égalisation d'histogramme. Deuxièmement, pour mieux traiter les conditions standard des systèmes projecteur-caméra, deux ensembles de données de captures de projections réelles, ont été spécialement préparés à des fins expérimentales. La performance de tous les algorithmes considérés est analysée de façon approfondie et des propositions de recommandations sont faites sur le choix des algorithmes les mieux adaptés en fonction des conditions expérimentales (paramètres image, disposition spatiale, couleur du fond...). Troisièmement, nous considérons le problème d'ajustement multi-surface pour compenser des distorsions d'homographie dans les images acquises. Une combinaison de mise en correspondance entre les primitives et de Flux Optique est proposée afin d'obtenir une compensation géométrique plus rapide. Quatrièmement, une nouvelle application en reconnaissance d'objet à partir de captures d'images vidéo-projetées est mise en œuvre. Finalement, une implémentation GPU temps réel des algorithmes considérés ouvre des pistes pour la compensation géométrique non intrusive en SAR basée sur la mise en correspondances entre primitives.
Christian Jacquemin合作论文数3