Résumé. La détection du plagiat passe le plus souvent par la phase de recherche de similitudes la plus naïve, la détection de « copier/coller ». Dans cet article, nous proposons une méthode alternative à l’approche standard de comparaison mot à mot. Le principe étant d’effectuer une intersection des deux textes à comparer, récupérant ainsi un tableau des mots qu’ils ont en commun et de ne conserver que les séquences maximales des mots se suivant dans l’un des textes et existant également dans l’autre. Nous montrons que cette méthode est plus rapide et moins coûteuse en ressources que les méthodes de parcours de textes habituellement utilisées. L’objectif étant de détecter les passages identiques entre deux textes plus rapidement que les méthodes de comparaison mot à mot, tout en étant plus efficace que les méthodes n-grammes.
Résumé. La détection de plagiat extrinsèque devient vite inefficace lorsque l’on n’a pas accès aux documents potentiellement sources du plagiat ou lorsque l’on se confronte à un espace aussi vaste que le Web, ce qui est souvent le cas dans les logiciels anti-plagiat actuels. Dès lors la détection intrinsèque devient nettement plus efficace. Dans cet article, nous traitons justement de la détection automatique d’auteurs qui permet de savoir si un passage d’un texte n’appartient pas au même auteur que le reste du texte et donc en théorie de repérer les passages plagiés d’un document. Nous expliquons notre contribution aux procédures déjà existantes et évaluons les limites de notre approche. L’objectif est de permettre la détection et le regroupement de passages d’un document par auteur.
Comparison of two documents in the plagiarism detection context is often reduced to a word to word comparison, a research of copy and paste. In this article, a naive approach to compare two documents with the aim of automatically detecting whether a copied sentence from one text to the other, or paraphrases and reformulations, is presented. This is achieved by looking for the existence of meaningful words and their potential substitution words. We compare three algorithms using this approach and retain only the most efficient one to evaluate it with existing methods. The goal is to enable detection of similarities between two texts using only keywords. The proposed approach can detect non paraphrastic reformulations, which are impossible to detect with the conventional alignment approach.
La recherche de documents similaires est un processus qui consiste a trouver les documents presentant des similitudes, comme la copie ou la reformu- lation, sur des bases documentaires ou sur internet. Elle est utilisee notamment pour proteger la propriete intellectuelle de productions issues de l'enseignement, de la recherche ou de l'industrie. Dans cet article, nous definissons une approche automatique pour permettant d'extraire des mots-cles d'un document en effec- tuant un bouclage sur une succession de decoupage de plus en plus petit. Cette approche permet d'obtenir des mots-cles impossibles a obtenir par une approche globale notamment quand la thematique, le style ou le contenu d'un document varient dans le document. L'objectif est de permettre la detection des documents presentant des similitudes en utilisant uniquement des mots-cles.
The identification and authentication of individuals by their palmprints is a recent approach in the family of biometric modalities, very interesting in appearance and non-contact non-intrusive. It was recently studied in topic research over the last decade. The proposed approaches are mostly based on methods of classification and learning. However, the complexity of the calculations leads to an inappropriate application in real time. In our work, we investigated the use of basic primitives pictures more precisely the Space Interest Points (SIO) in a realtime process of identification and authentication of palmprint. This process is based on construction and matching of graph. By setting few constraints and working with matching methods and matching specific, experimental results suggest a robust real-time solution as good as the best methods with an error rate authentication below 1% for a population of 20 individuals.
Video media are now emerging with the democratization of broadband, capture tools and content delivery services. Despite the current performance of the 'traditional' search engines, the video search engines are challenged by the semantic difference between requests made verbatim and the data as video. Thus, most engines are not based on content but on the tags associated with documents. The approach presented in this paper is based on extraction of high-level semantics information and offers a visualization of the results based on the content, representativeness, temporal location and redundancy, which ensures the users a high level of relevance of their research results.
Among all the features which can be extracted from videos, we propose to use Space-Time Interest Points (STIPs). STIPs are particularly interesting because they are simple and robust low-level features providing an efficient characterization of moving objects within videos. In this paper, after defining STIPs and after giving some of their properties, we will use STIPs to detect moving objects and to characterize specific changes in the movements of these objects. Proposed results are obtained from two very different types of videos, namely athletic videos and animation movies.
Parmi toutes les caracteristiques qui peuvent etre extraites de videos, les points d'interet spatiotemporels (STIP) sont particulierement interessants car ce sont des caracteristiques de bas niveau simples et robustes qui permettent une bonne caracterisation des objets en mouvement. Dans cet article, nous definissons les STIP et analysons leurs proprietes. Puis, les STIP sont utilises pour detecter des objets en mouvement et pour caracteriser les changements specifiques dans les mouvements de ces objets. Les performances sont etudiees sur des types tres differents de videos : des sequences d'athletisme et des sequences de films d'animation.
In the video indexing framework, we have developed an assistance system for the user to define a new concept as semantic index according to the features automatically extracted from the video. Because the manual indexing is a long and tedious task, we propose to focus the attention of the user on pre selected prototypes that a priori correspond to the concept. The proposed system is decomposed in three steps. In the first one, some basic spatio-temporal blocks are extracted from the video, a particular block is associated to a particular property of one feature. In the second step, a Question/Answer system allows the user to define links between basic blocks in order to define concept block models. And finally, some concept blocks are extracted and proposed as prototypes of the concepts. In this paper, we present the two first steps, particularly the block structure, illustrated by an example of video indexing that corresponds to the concept running in athletic videos.
This papers tests the relevance of interest points to predict eye movements of subjects when viewing video sequences freely. Moreover the papers compares the eye positions of subjects with interest maps obtained using two classical interest point detectors: one spatial and one space-time. We fund that in function of the video sequence, and more especially in function of the motion inside the sequence, the spatial or the space-time interest point detector is more or less relevant to predict eye movements.