Cynthia Basileu* — Sofian Ben Amor ** — Marc Bui ** — Michel Lamure* * Universite Claude Bernard Lyon1, Laboratoire ERIC, Universite Lyon 1 Batiment Odontologie, 11 rue Guillaume Paradin, 69008 Lyon FRANCE cbasileu@yahoo.fr, michel.lamure@univ-lyon1.fr ** Ecole Pratique des Hautes Etudes, Laboratoire LaISC (EPHE-Sorbonne) 41, rue Gay Lussac 75005 Paris {sofiane.benamor, marc.bui}@ephe.sorbonne.fr
Latent semantic analysis is a computation method to demonstrate a major component of language learning and use. Thus, in this sense, it is a theory of meaning, such that it applies to and offers an explanation of phenomena of meaning in words and passages of words. This enables LSA to hold a strong position in the automated document classification, document analysis, etc. Though the experiments show that LSA can reach a very high accuracy in document classification, it also depends on the various factors such as quality and amount of training documents, characteristics of representative vector and composition of the to be classified documents, etc. On the other hand, pretopology is showing its strength in the fields of data classification and modeling. Besides, some applications, which are to strengthen the pretopology with visualization in the domain of classification, have shown promising results. In this paper two document classification algorithms based on pretopology and LSA are proposed, which are suitable for different situations, and their results with deft07 contest data are discussed. This work also shows future possibility of visualization integration, which could help human intervention in the classification process. RESUME. L’Analyse de la Semantique Latente (LSA) est une methode de calcul qui permet de rendre compte de l’apprentissage du langage et de son utilisation. Dans ce sens, LSA est une theorie de la signification des mots et groupes de mots (paragraphes, passages, textes) et de leur emploi. Cette propriete permet a LSA d’occuper une position enviable dans la classification automatique de documents, l’analyse de documents, etc. Bien que de nombreuses experiences indiquent que LSA peut atteindre une grande precision dans la classification de documents, ses Studia Informatica Universalis. performances sont tributaires de facteurs tels que la qualite et la quantite de documents utilises pour l’entrainement, les caracteristiques des vecteurs representatifs et la composition des documents a classer. De son cote, la pretopologie a montre son efficacite dans les domaines de la classification des donnees et de la modelisation. De plus, certaines applications ont renforce la pretopologie en ajoutant la visualisation au domaine de la classification et ont donne des resultats prometteurs. Dans cet article, nous proposons deux algorithmes de classification des documents bases sur LSA et la pretopologie, algorithmes qui sont adaptes a des situations differentes et dont nous discutons les resultats obtenus quand ils sont appliques aux donnees du defi DEFT07. Ce travail dessine egalement les possibilites futures d’integration de la visualisation, integration qui pourra contribuer a l’intervention humaine dans les processus de
Pollution in metropolitan cities has become a serious problem, resulting in poor living conditions and serious health problems. Pollution being qualified as a complex system, we propose a multi-agent approach to model and simulate it, so that we could study, analyze and predict it better. As in the early stage of the project, we have some successful experiments and attempts to integrate the mathematical theory of Pretopology in the modeling and simulation levels. In addition, these interesting results shades some light on our future direction.