Formal Concept Analysis (FCA) is an approach for conceptual classification building and rule discovery from a binary table describing a set of objects by a set of attributes. Extensions have been proposed to deal with non-binary and more complex data, such as Relational Concept Analysis (RCA) for multi-relational data. RCA aims to highlight groups of objects characterized by their relationships with other groups of objects. The richer and more complex nature of the underlying data allows RCA to produce richer results than FCA, at the expense of higher computational and interpretive complexity. The most commonly used conceptual classification structure in FCA is the concept lattice. However, in many applications, concept lattice substructures, such as AOC-posets, are preferred over the full lattice, either to mitigate combinatorial blow-up or to focus on the most informative parts of the structure. Indeed, in an AOC-poset, only concepts introducing an object or an attribute are represented, which makes AOC-posets smaller and easier to compute and use than concept lattices. Although RCA was originally defined on concept lattices, it can also be instantiated on AOC-posets. RCA is iterative and its convergence is guaranteed in the lattice-based setting, but this guarantee is lost when using AOC-posets. In this paper, we investigate this loss of convergence in detail. We show why convergence is no longer guaranteed in the general case, identify conditions under which it can still be ensured, and discuss how a dataset can be transformed to recover convergence. We also propose a convergent variant of the process, which preserves the AOC-poset structure: relational attributes, once created, are never removed, which guarantees convergence at the price of attributes that may refer to concepts absent from the final structures.
Relational Concept Analysis (RCA) and Graph-FCA (GCA) are two extensions of Formal Concept Analysis (FCA) to multi-relational data. Whereas similar when restricted to some specific inputs and parameters, RCA and GCA processes are different and can be applied to a variety of multi-relational data. Analysts should thus be guided to choose between the two methods, their parameters, and the data model, according to the queries to solve. In this paper, we focus on the processing of n-ary relationships, how they can be modeled, and the impact of the chosen model and the chosen method on query answering.
Relational Concept Analysis (RCA) and Graph-FCA (GCA) are two extensions of Formal Concept Analysis (FCA) introduced in order to allow concept analysis on multi-relational data. The two methods have different properties and parameters, but when restricting to binary relationships, existential quantifier and unary concepts, their outputs look similar. On this basis, a theoretical comparison of the two methods is conducted, showing that each RCA concept corresponds to a GCA concept. Furthermore, to allow the comparison of concept intensions, a transformation of RCA results into relational patterns is performed. These results give a sound basis to help interpreting RCA results and to combine the two approaches for data exploration.
Relational Concept Analysis (RCA) and Graph-FCA (GCA) have been defined as Formal Concept Analysis (FCA) extensions for processing relational data and knowledge graphs respectively. Nevertheless, while their purposes and results seem similar, the data modelling and the definition of concepts are different. In this paper, we compare these two approaches on a common basis, considering only unary and binary relations for GCA and the existential quantifier for RCA. We focus on examples showing the similarities and dissimilarities between both methods, and highlighting how cycles are processed differently by RCA and GCA.
La directive cadre européenne sur l’eau (2000) fixe l’atteinte du bon état écologique dans toutes les masses d’eau, à court et moyen termes. Ce bon état est établi par des indices biologiques basés sur les êtres vivants aquatiques. Les conditions hydrologiques sont l’une des caractéristiques physiques importantes de ces écosystèmes, d’autant plus dans le contexte actuel de changement climatique qui pourrait provoquer des phénomènes extrêmes plus marqués. Si l’influence de l’hydrologie est connue sur les êtres vivants des rivières, la recherche d’indicateurs spécifiques est limitée. Notre objectif est de proposer des métriques caractérisant les conditions de basses eaux et de hautes eaux, ainsi que la variabilité hydrologique. Nous avons ainsi retenu finalement six métriques caractérisant le régime hydrologique d’une station hydrométrique donnée et permettant de préciser la situation hydrologique à la fois dans l’année qui précède un prélèvement biologique et dans les jours immédiatement avant son échantillonnage, afin de vérifier si celui-ci a été réalisé en conditions hydrologiques stables. Un travail statistique, réalisé sur des données de 2007 à 2013, a permis également de les discrétiser en trois classes d’intensité : faible, moyen et fort. Ces métriques sont robustes, indépendantes des conditions géographiques et géologiques, mais semblent bien correspondre au contexte climatique. Nous nous sommes heurtés à la non-concordance des réseaux publics français, hors outre-mer, de suivis de la qualité des eaux (Réseau de contrôle de surveillance) et des débits (BD Hydro) : ces métriques sont calculables pour seulement 57 % des stations biologiques. Ce travail est une étape préliminaire à notre objectif de confronter les métriques retenues aux résultats des indices biologiques en utilisant une méthode de fouille de données, que nous avons élaborée et déjà testée sur des altérations physico-chimiques.
Exploring old pharmacopoeias is a promising way to find active ingredients that can be useful to design new drugs. Nevertheless, studying these texts is a laborious task for biologists. Therefore, an interdisciplinary project was undertaken: texts have been annotated to extract relevant information and represent it within a graph database. Formal Concept Analysis (FCA) and Relational Concept Analysis (RCA) have then been used to explore this database, in order to answer questions regarding remedies and their ingredients. This paper presents the data and some results obtained with FCA and RCA. It highlights the suitability of these approaches to explore these data and answer the needs of biologists.
Many applications of Formal Concept Analysis (FCA) and its diverse extensions have been carried out in recent years. Among these extensions, Relational Concept Analysis (RCA) is one approach for addressing knowledge discovery in multi-relational datasets. Applying RCA requires stating a question of interest and encoding the dataset into the input RCA data model, i.e. an Entity-Relationship model with only Boolean attributes in the entity description and unidirectional binary relationships. From the various concrete RCA applications, recurring encoding patterns can be observed, that we aim to capitalize taking software engineering design patterns as a source of inspiration. This capitalization work intends to rationalize and facilitate encoding in future RCA applications. In this paper, we describe an approach for defining such design patterns, and we present two design patterns: “Separate/Gather Views” and “Level Relations”.
We have implemented a specific data mining process to explore the relationship between biological indices and physico-chemical pressures in rivers. Data were collected in the framework of the French National monitoring network set up to assess the ecological status of rivers under the European Water Framework Directive (WFD). Chemical parameters and biological indices were collected regularly from 1.781 locations in metropolitan France from 2007 to 2013. The sequential pattern mining process generates closed partially ordered patterns representing a succession of physico-chemical events that precede a given biological index in a given status, validated using a subset of data. This paper focuses on the patterns and their occurrence. We showed that biological statuses depend on these temporal successions of alterations and not only on the last alterations. The physico-chemical statuses of water bodies usually appeared to be higher than their biological statuses, suggesting synergism between toxicants and/or an additive impact of other stressors related to hydromorphology or hydrology. Patterns found in the highest biological status for the biological indices based on macroinvertebrates, diatoms, macrophytes or fish, were characterised by the constancy of a high physico-chemical status over time. By contrast, before indices based on macroinvertebrates and macrophytes, two types of patterns were observed for bad biological status: (1) a chronic multi-pressure pattern, in which pressure categories such as nitrates, pesticides and other organic hydrocarbons, in moderate, poor or bad status, repeated themselves several times over time, or (2) a single occurrence of a degraded pressure category, such as one moderate nitrogen, excluding nitrate, or one poor oxidizable organic matter, among other pressure categories in good status. Extracting such patterns is a promising solution both to disentangle the effects of the different stressors on water quality, and to identify the key temporal sequences among them in a context of multi-stress conditions, which is a challenge currently facing the WFD.
Most of available data are inherently relational, with e.g. temporal, spatial, causal or social relations. Besides, many datasets involve complex and voluminous data. Therefore, the exploration of relational data is a major challenge for Formal Concept Analysis (FCA). Relational Concept Analysis (RCA) is specifically designed to investigate the relational structure of a dataset in the FCA paradigm. In this chapter, we examine how RCA can take over the issues raised by complex data. Using two datasets, one about the quality monitoring of waterbodies in France, the other about the use of pesticidal and antimicrobial plants in Africa, we study the limitations of different FCA algorithms, and their current implementations to explore these datasets with RCA. We also show how pattern extraction combined with the presentation of data in hierarchical structures is appropriate for the analysis of temporal datasets by the domain expert. Finally, we discuss about the possible directions to investigate.
. This paper is a short feedback on a collaborative research work by computer scentists and hydroecologists. We have applied Relational Concept Analysis on complex data about running water characteristics (physical, biological and chemical parameters), to answer various questions. Two approaches are presented and discussed: the first one extracts patterns from temporal data, the second one extracts rules from a multi-relational dataset
. This paper presents the collaborative work led in an interdisciplinary project on studying remedies in Arabic medieval pharmacopeia. The goal is to find new molecules in plants or combinations of active principles to substitute them for antibiotics which encounter some limits. Formal Concept Analysis is used to discover co-occurrences of ingredients. We describe the difficulties inherent to these data and the results of preliminary analyses.
Cet article s'interesse a l'exploration de jeux de donnees multi-relationnelles, et aux differentes manieres de les analyser en utilisant l'analyse relationnelle de concepts (ARC), une variante de l'analyse formelle de concepts. L'ARC utilise plusieurs quantifieurs d'echelonnage qui rendent le processus d'analyse finement reglable, permettant une grande flexibilite dans l'exploration et dans ses resultats. En contrepartie, l'analyste peut etre submerge par l'ensemble des choix qu'il doit faire au cours de l'analyse. Pour traiter ce probleme, nous proposons trois sur-couches qui aident l'analyste a anticiper et controler les resultats de ses choix. Notre proposition est appliquee a un jeu de donnees sur la qualite des eaux de rivieres.
Relational Concept Analysis (RCA) is one variant of Formal Concept Analysis for multi-relational dataset exploration. The tool RCAExplore is an implementation of the RCA process where several choices can be made before each iteration: the structure to be used (con-cept lattice, AOC-Poset, Iceberg lattice), the scaling quantifier (exist, forall, contains, percentage-quantifiers, etc.), and the considered formal and relational contexts. RCAExplore was developed during Fresqueau ANR 11 MONU 14 project, in order to explore relational hydroecolog-ical data, and has also been used on several datasets in different other domains. The source code and a standalone jar application are available online. An integration of the tool is ongoing within a platform for data exploration and knowledge representation.
Ouzerdine, Amirouche Braud, Agnès Dolques, Xavier Huchard, Marianne Le Ber, FlorenceIn this paper, we focus on the exploration of multi-relational datasets, and the various ways they can be analyzed using Relational Concept Analysis (RCA), an extension of Formal Concept Analysis (FCA). RCA uses several scaling operators that make the process highly tunable, allowing a high flexibility in the exploration and in the results. In return, the multiplicity of choices that can be made when performing an analysis task potentially overwhelms the expert. We thus propose three overlays for helping users control and foresee the results of their choices. Our proposition is exemplified on a dataset about the hydro-ecological state of watercourses.
Relational Concept Analysis (RCA) has been designed to classify sets of objects described by attributes and relations between these objects. This is achieved by iterating on Formal Concept Analysis (FCA). It can be used to discover knowledge patterns and implication rules in multi-relational datasets. The classification output by RCA is a family of lattices whose graphical representation facilitates the analysis by an expert. However, RCA comes with specific complexity issues. It iterates on the building of interconnected concept lattices, so that each concept in a lattice might be the cause of generating other concepts in other lattices. In complex analyses, it relies on the successive choice of scaling operators which affects the size and the understandability of the results. These operators are based on a set of quantifiers which are studied in this paper: we indeed focus on the comparison of scaling quantifiers and highlight a generality relation between them. Our theoretical proposition is complemented by an experimental evaluation of the exploration space size, based on a real dataset upon watercourses. This work is intended for data analysts, to provide them with an overview on the different strategies offered by RCA.
This paper presents a prototype of case-based reasoning, built for the agricultural domain. Its aim is to forecast the allocation of a new energy crop, the miscanthus. Interviews were conducted with french farmers in order to know how they make their decisions. Based on interview analysis, a case base and a rule base have been formalized, together with similarity and adaptation knowledge. Furthermore we have introduced variations in the reasoning modules, for allowing different uses. Tests have been conducted. Results showed that the model can be used in different ways, according to the aim of the user, and e.g. the economic conditions for miscanthus allocation.
In this paper, we consider data analysis methods for knowledge extraction from large water data-sets. More specifically, we try to connect physico-chemical parameters and the characteristics of taxons living in sample sites. Among these data analysis methods, we consider formal concept analysis (FCA), which is a recognized tool for classification and rule discovery on object-attribute data. Relational concept analysis (RCA) relies on FCA and deals with sets of object-attribute data provided with relations. RCA produces more informative results but at the expense of an increase in complexity. Besides, in numerous applications of FCA, the partially ordered set of concepts introducing attributes or objects (AOC poset, for Attribute-Object-Concept poset) is used rather than the concept lattice in order to reduce combinatorial problems. AOC posets are much smaller and easier to compute than concept lattices and still contain the information needed to rebuild the initial data. This paper introduces a variant of the RCA process based on AOC posets rather than concept lattices. This approach is compared with RCA based on iceberg lattices. Experiments are performed with various scaling operators, and a specific operator is introduced to deal with noisy data. We show that using AOC poset on water data-sets provides a reasonable concept number and allows us to extract meaningful implication rules (association rules whose confidence is 1), whose semantics depends on the chosen scaling operator.
This paper presents an approach for mining temporal data, based on Relational Concept Analysis (RCA), that has been developed for a real world application. Our data are sequential samples of biological and physico-chemical parameters taken from watercourses. Our aim is to reveal meaningful relations between the two types of parameters. To this end, we propose a comprehensive temporal data mining process starting by using RCA on an ad hoc temporal data model. The results of RCA are converted into closed partially ordered patterns to provide experts with a synthetic representation of the information contained in the lattice family. Patterns can also be filtered with various measures, exploiting the notion of temporal objects. The process is assessed through some quantitative statistics and qualitative interpretations resulting from experiments carried out on hydroecological datasets.
Cet article presente une methode d'exploration de donnees temporelles, fondee sur l'analyse relationnelle de concepts (ARC) et appliquee a des donnees sequentielles construites a partir d'echantillons physico-chimiques et biologiques preleves dans des cours d'eau. Notre but est de mettre au jour des sous-sequences pertinentes et hierarchisees, associant les deux types de parametres. Pour faciliter la lecture, ces sous-sequences sont representees sous la forme de motifs partiellement ordonnes (po-motifs). Le processus de fouille de donnees se decompose en plusieurs etapes : construction d'un modele temporel ad hoc et mise en oeuvre de l'ARC ; extraction des sous-sequences synthetisees sous la forme de po-motifs ; selection des po-motifs interessants grâce a une mesure exploitant la distribution des extensions de concepts. Le processus a ete teste sur un jeu de donnees reelles et evalue quantitativement et qualitativement.
Petko Valtchev合作论文数Departement d'informatique, University of Quebec at Montreal1
Houari Sahraoui合作论文数department of computer science and operations research (GEODES, software engineering group) of University of Montreal1