The exponential growth of multimedia content highlights the importance of specific methods to automatically extract the most representative key images, called key-frame - which we will refer to hereafter as tag-images - from videos. Unlike video surveillance, where the objects searched for are clearly defined, cinema offers a wide variety of visual and narrative styles, making the selection of a tag-image per shot particularly complex. In this paper, we propose a new method based on object detection, specially adapted to this context. The proposed method introduces an innovative strategy based on statistical analysis. This analysis relies on histograms of occurrences of detected object classes, enabling a more semantically consistent selection. This method is enriched with spatial, temporal, and sharpness weightings to optimize the relevance of a selected tag-image per shot. A qualitative evaluation conducted with film analysis experts shows that our method outperforms state-of-the-art approaches in terms of direct preference and robustness across various types of shots.
In this paper, we propose a new approach for detecting and removing impulse noise. The method uses the new framework of cloud filtering for detecting noise locations. This new framework use efficient mathematical tools to filter with sets of filters rather than with a single one. Once a (very) noisy pixel is detected, its illumination is estimated by an extension of the median filtering applied on the neighborhood defined by the cloud. Experiments on various images demonstrate the capacity of the algorithm to identify noisy pixels (especially at low noise rate) while well preserving image edges.
Aims. The treatment of astronomical image time series has won increasing attention in recent years. Indeed, numerous surveys following up on transient objects are in progress or under construction, such as the Vera Rubin Observatory Legacy Survey for Space and Time (LSST), which is poised to produce huge amounts of these time series. The associated scientific topics are extensive, ranging from the study of objects in our galaxy to the observation of the most distant supernovae for measuring the expansion of the universe. With such a large amount of data available, the need for robust automatic tools to detect and classify celestial objects is growing steadily.Methods. This study is based on the assumption that astronomical images contain more information than light curves. In this paper, we propose a novel approach based on deep learning for classifying different types of space objects directly using images. We named our approach ConvEntion, which stands for CONVolutional attENTION. It is based on convolutions and transformers, which are new approaches for the treatment of astronomical image time series. Our solution integrates spatio-temporal features and can be applied to various types of image datasets with any number of bands.Results. In this work, we solved various problems the datasets tend to suffer from and we present new results for classifications using astronomical image time series with an increase in accuracy of 13%, compared to state-of-the-art approaches that use image time series, and a 12% increase, compared to approaches that use light curves.
In this paper, we study the performance invariance of convolutional neural networks when confronted with variable image sizes in the context of a more "wild steganalysis". First, we propose two algorithms and definitions for a fine experimental protocol with datasets owning "similar difficulty" and "similar security". The "smart crop 2" algorithm allows the introduction of the Nearly Nested Image Datasets (NNID) that ensure "a similar difficulty" between various datasets, and a dichotomous research algorithm allows a "similar security". Second, we show that invariance does not exist in state-of-the-art architectures. We also exhibit a difference in behavior depending on whether we test on images larger or smaller than the training images. Finally, based on the experiments, we propose to use the dilated convolution which leads to an improvement of a state-of-the-art architecture.
Monocular cameras and multibeam imaging sonars are common sensors of Unmanned Underwater Vehicles (UUV). In this paper, we propose a new method for calibrating a hybrid sonar–vision system. This method is based on motion comparisons between both images and allows us to compute the transformation matrix between the camera and the sonar and to estimate the camera’s focal length. The main advantage of our method lies in performing the calibration without any specific calibration pattern, while most other existing methods use physical targets. In this paper, we also propose a new sonar–vision dataset and use it to prove the validity of our calibration method.
Autonomous vehicles are able to sense their environment and operate without human involvement. In particular, they have to be able to locate themselves. In order to achieve this task, geolocalisation files such as RINEX files are used. Due to the large amount of data, these files are very large and there is a big challenge in reducing their size for real-time navigation. However, this problem has not really been explored for RINEX files. In this paper, we propose an efficient method to losslessly compress RINEX files based on quadratic polynomial interpolation for autonomous vehicles navigation. Indeed, due to the orbits of the satellites around the earth, satellite signals can be modeled by parabolic curves. After interpolation of this model, prediction errors are computed as the differences between each original value of the satellite signal and its associated predicted value. Finally, these prediction errors are compressed using entropy coding. Experimental results show that our proposed method allows us to outperform the compression rate obtained by a previous state-of-the-art approach.
For many years, the image databases used in steganalysis have been relatively small, i.e. about ten thousand images. This limits the diversity of images and thus prevents large-scale analysis of steganalysis algorithms. In this paper, we describe a large JPEG database composed of 2 million colour and grey-scale images. This database, named LSSD for Large Scale Steganalysis Database, was obtained thanks to the intensive use of “controlled” development procedures. LSSD has been made publicly available, and we aspire it could be used by the steganalysis community for large-scale experiments. We introduce the pipeline used for building various image database versions. We detail the general methodology that can be used to redevelop the entire database and increase even more the diversity. We also discuss computational cost and storage cost in order to develop images.
Since the emergence of deep learning and its adoption in steganalysis fields, most of the reference articles kept using small to medium size CNN, and learn them on relatively small databases. Therefore, benchmarks and comparisons between different deep learning-based steganalysis algorithms, more precisely CNNs, are thus made on small to medium databases. This is performed without knowing: 1. if the ranking, with a criterion such as accuracy, is always the same when the database is larger, 2. if the efficiency of CNNs will collapse or not if the training database is a multiple of magnitude larger, 3. the minimum size required for a database or a CNN, in order to obtain a better result than a random guesser. In this paper, after a solid discussion related to the observed behaviour of CNNs as a function of their sizes and the database size, we confirm that the error's power-law also stands in steganalysis, and this in a border case, i.e. with a medium-size network, on a big, constrained and very diverse database.
Image steganography aims to securely embed secret information into cover images. Until now, adaptive embedding algorithms such as S-UNIWARD or Mi-POD, were among the most secure and most often used methods for image steganography. With the arrival of deep learning and more specifically, Generative Adversarial Networks (GAN), new steganography techniques have appeared. Among them is the 3 -player game approach, where three networks compete against each other. In this paper, we propose three different architectures based on the 3-player game. The first architecture is proposed as a rigorous alternative to two recent publications. The second takes into account stego noise power. Finally, our third architecture enriches the second one with a better interaction between embedding and extracting networks. Our method achieves better results compared to existing works Hayes and Danezis (2017), Zhu et al. (2018), and paves the way for future research on this topic.
After 2015, CNN-based steganalysis approaches have started replacing the two-step machine-learning-based steganalysis approaches (feature extraction and classification), mainly due to the fact that they offer better performance. In many instances, the performance of these networks depend on the size of the learning database. Until a certain point, the larger the database, the better the results. However, working with a large database with controlled acquisition conditions is usually rare or unrealistic in an operational context. An easy and efficient approach is thus to augment the database, in order to increase its size, and therefore to improve the efficiency of the steganalysis process. In this article, we propose a new way to enrich a database in order to improve the CNN-based steganalysis performance. We have named our technique "pixels-off". This approach is efficient, generic, and is usable in conjunction with other data-enrichment approaches. Additionally, it can be used to build an informed database that we have named "Side-Channel-Aware databases" (SCA-databases).
Cosmologists are facing the problem of the analysis of a huge quantity of data when observing the sky. The methods used in cosmology are, for the most of them, relying on astrophysical models, and thus, for the classification, they usually use a machine learning approach in two-steps, which consists in, first, extracting features, and second, using a classifier. In this paper, we are specifically studying the supernovae phenomenon and especially the binary classification I.a supernovae versus not-I.a supernovae. We present two Convolutional Neural Networks (CNNs) defeating the current state-of-the-art. The first one is adapted to time series and thus to the treatment of supernovae light-curves. The second one is based on a Siamese CNN and is suited to the nature of data, i.e. their sparsity and their weak quantity (small learning database).
For about 10 years, detecting the presence of a secret message hidden in an image was performed with an Ensemble Classifier trained with Rich features. In recent years, studies such as Xu et al. have indicated that well-designed convolutional Neural Networks (CNN) can achieve comparable performance to the two-step machine learning approaches. In this paper, we propose a CNN that outperforms the state-ofthe-art in terms of error probability. The proposition is in the continuity of what has been recently proposed and it is a clever fusion of important bricks used in various papers. Among the essential parts of the CNN, one can cite the use of a pre-processing filterbank and a Truncation activation function, five convolutional layers with a Batch Normalization associated with a Scale Layer, as well as the use of a sufficiently sized fully connected section. An augmented database has also been used to improve the training of the CNN. Our CNN was experimentally evaluated against S-UNIWARD and WOW embedding algorithms and its performances were compared with those of three other methods: an Ensemble Classifier plus a Rich Model, and two other CNN steganalyzers.
Deep learning and convolutional neural networks (CNN) have been intensively used in many image processing topics during last years. As far as steganalysis is concerned, the use of CNN allows reaching the state-of-the-art results. The performances of such networks often rely on the size of their learning database. An obvious preliminary assumption could be considering that "the bigger a database is, the better the results are". However, it appears that cautions have to be taken when increasing the database size if one desire to improve the classification accuracy i.e. enhance the steganalysis efficiency. To our knowledge, no study has been performed on the enrichment impact of a learning database on the steganalysis performance. What kind of images can be added to the initial learning set? What are the sensitive criteria: the camera models used for acquiring the images, the treatments applied to the images, the cameras proportions in the database, etc? This article continues the work carried out in a previous paper, and explores the ways to improve the performances of CNN. It aims at studying the effects of "base augmentation" on the performance of steganalysis using a CNN. We present the results of this study using various experimental protocols and various databases to define the good practices in base augmentation for steganalysis.
In order to preserve cultural heritage, this paper proposes to develop a digital paintings visualization system. Our proposed system mainly consists of extracting regions of interest (ROI) from a digital painting to characterize them. These close-ups are then animated on the basis of the painting characteristics and the artist's or designer's aim. In order to obtain interesting results from short video clips, we developed a visual saliency map-based method by using the well-known Itti's saliency analysis which already proved its efficiency on paintings. The experimental results show the efficiency of our approach and an evaluation based on a Mean Opinion Score validates the proposed method. Eighteen volunters watched for a total of 216 views at 72 video clips build up from 6 paintings, and we found a significant difference between our saliency-based generated videos and random-based generated videos.
Over the last 15 years, several applications have been developed for digital cultural heritage in the image processing and particularly in the area of digital painting. In order to help preserve cultural heritage, this chapter proposes several applications for digital paintings such as restoration, authentication, style analysis and visualization. For the visualization of digital paintings we present specific methods to visualize digital paintings based on visual saliency and in particular we propose an automatic digital painting visualization method based on visual saliency. The proposed system consists of extracting regions of interest (ROI) from a digital painting to characterize them. These close-ups are then animated on the basis of the paintings characteristics and the artist's or designer's aim. In order to obtain interesting results from short video clips, we developed a visual saliency mapbased method. The experimental results show the efficiency of our approach and an evaluation based on a Mean Opinion Score validates our proposed method.
The most effective superresolution methods proposed in the literature require precise knowledge of the so-called point spread function of the imager, while in practice its accurate estimation is nearly impossible. This paper presents a new superresolution method, whose main feature is its ability to account for the scant knowledge of the imager point spread function. This ability is based on representing this imprecise knowledge via a non-additive neighborhood function. The superresolution reconstruction algorithm transfers this imprecise knowledge to output by producing an imprecise (interval-valued) high-resolution image. We propose some experiments illustrating the robustness of the proposed method with respect to the imager point spread function. These experiments also highlight its high performance compared with very competitive earlier approaches. Finally, we show that the imprecision of the high-resolution interval-valued reconstructed image is a reconstruction error marker.
Pendant environ 10 ans, l'approche classique pour detec-ter la presence d'un message secret insere dans une image etait d'utiliser un ensemble de classifieurs alimentes par des vecteurs de caracteristiques issues des images a traiter. Ces dernieres annees, des etudes telles que Xu et al. ont in-dique que des reseaux de neurones convolutionnels (CNN) bien concus peuvent atteindre des performances compa-rables aux approches classiques d'apprentissage automa-tique. Dans cet article, nous proposons un CNN qui depasse les performances de l'etat de l'art en terme de proba-bilite d'erreur de classification. La proposition s'inscrit dans la continuite de ce qui a ete propose recemment et consiste en une fusion intelligente de briques importantes proposees dans divers articles. Parmi les elements essen-tiels du CNN propose, on peut citer l'utilisation : d'un ensemble de filtres pour le pretraitement de l'image d'en-tree, de la troncature comme fonction d'activation, d'au moins cinq couches convolutionnelles avec une normalisa-tion par lot (Batch normalization layer) et une couche de mise a l'echelle (Scale layer) ainsi qu'une couche entiere-ment connectee correctement dimensionnee.
In this paper, we present a new regularization paradigm for inverse based regularized image reconstruction techniques. These methods usually attempt to minimize a cost function expressed as the sum of a data-fitting term and a regularization term. The trade-off between both terms is determined by a weighting parameter that has to be set by the user since this trade-off is data dependent. In the approach we present here, we first concentrate on finding a set of eligible candidates for the data fitting term minimization and then select the most appropriate candidate according to the regularization criterion. The main advantage of this method is that it does not require any weighting parameter, and guarantees that no over-regularization can occur. We illustrate this method with a super-resolution reconstruction technique to show its efficiency compared to other competitive methods. Comparisons are carried out with simulated and real data.
Source camera identification methods aim at identifying the camera used to capture an image. In this paper we developed a method for digital camera model identification by extracting three sets of features in a machine learning scheme. These features are the co-occurrences matrix, some features related to CFA interpolation arrangement, and conditional probability statistics. These features give high order statistics which supplement and enhance the identification rate. The method is implemented with 14 camera models from Dresden database with multi class SVM classifier. A comparison is performed between our method and a camera fingerprint correlation-based method which only depends on PRNU extraction. The experiments prove the strength of our proposition since it achieves higher accuracy than the correlation-based method.
Olivier Strauss合作论文数Universite Montpellier II32