Contrast set mining aims at finding differences between different groups. This paper shows that a contrast set mining task can be transformed to a subgroup discovery task whose goal is to find descriptions of groups of individuals with unusual distributional characteristics with respect to the given property of interest. The proposed approach to contrast set mining through subgroup discovery was successfully applied to the analysis of records of patients with brain stroke (confirmed by a positive CT test), in contrast with patients with other neurological symptoms and disorders (having normal CT test results). Detection of coexisting risk factors, as well as description of characteristic patient subpopulations are important outcomes of the analysis.
The goal of exploratory pattern mining is to find patterns that exhibit yet unknown relationships in data and to provide insightful representations of detected relationships, This paper explores contrast set mining and an approach to improving its explanatory potential by using the so called supporting factors that provide additional descriptions of the detected patterns. The proposed methodology is described in a medical data analysis problem of distinguishing between similar diseases in the analysis of patients suffering from brain ischaemia.
The task addressed and the method proposed in this paper aim at improved understanding of differences between similar diseases. In particular we address the problem of distinguishing between thrombolic brain stroke and embolic brain stroke as an application of our approach of contrast set mining through subgroup discovery. We describe methodological lessons learned in the analysis of brain ischaemia data and a practical implementation of the approach within an open source data mining toolbox.
The problem of traceability of genetically modified organisms (GMOs) addresses the detection, identification and quantification of GMOs in food, feed and seed samples. Due to a large number of GMOs in the market, a system for reliable and affordable traceability of GMOs, optimizing the price of testing for a given sample, has to be established. We have defined the input to the future decision support system for assay selection to be of a tabular form, with rows corresponding to GMOs, and columns corresponding to assays they react to. In this paper we present a prototype decision support system for GMO detection and identification.
The paper describes a data mining and visualization experiment performed on a real-world problem and results achieved in the experiment. A subgroup discovery algorithm was used in order to find useful patterns in production data. Workshop and manufacturing work system utilization visualization plots were created in order to gain some useful information. The purpose of this work is to examine the usefulness of data mining and visualization of the shop floor data to support decision making in a manufacturing company.
Closed sets are being successfully applied in the context of compacted data representation for association rule learning. However, their use is mainly descriptive. This paper shows that, when considering labeled data, closed sets can be adapted for prediction and discrimination purposes by conveniently contrasting covering properties on positive and negative examples. We formally justify that these sets characterize the space of relevant combinations of features for discriminating the target class. In practice, identifying relevant/irrelevant combinations of features through closed sets is useful in many applications. Here we apply it to compacting emerging patterns and essential rules and to learn descriptions for subgroup discovery.
This paper presents the state of the art of subgroup visualization methods. Visualization methods are evaluated by different criteria. A novel subgroup visualization method is proposed and its implementation as a part of an interactive interface for subgroup discovery is presented.
This paper applies a recently introduced methodology of closed itemset mining for class labeled data to potato microarray data. The study shows the discovered rules that best distinguish between virus resistant and virus sensitive transgenic potato lines. The discovered rules are interpretable and meaningful to domain experts.