
The acknowledgement section of this paper originally referred to grant DEC-2013/09/B/ST6/01568. The reference to this grant has been removed from the acknowledgement section at the request of one of the authors.
When removing some attributes, the partition induced by a smaller set of attributes will be coarser and the decision regions may be changed. In this paper, we analyze the decision region changes when removing attributes and propose a new type of attribute reducts from the point of view of vector based three-way approximations of a partition. We also present a reduct construction method by using a discernibility matrix.
An interval set is a family of sets restricted by a upper bound and lower bound. Interval-set algebras are concrete models of granular computing. The triarchic theory of granular computing focuses on a multilevel and multi-view granular structure. This paper discusses granular structures of interval sets under inclusion relations between two interval sets from a measurement-theoretic perspective and set-theoretic perspective, respectively. From a measurement-theoretic perspective, this paper discusses preferences on two objects represented by interval sets under inclusion relations on interval sets. From a set-theoretic perspective, this paper uses different inclusion relations and operations on interval sets to construct multilevel and multi-view granular structures.
Principal curves are nonlinear generalizations of principal components analysis. They are smooth self-consistent curves that pass through the middle of the distribution. By analysis of existed principal curves, we learn that a soft k-segments algorithm for principal curves exhibits good performance in such situations in which the data sets are concentrated around a highly curved or self-intersecting curves. Extraction of features are critical to improve the recognition rate of off-line handwritten characters. Therefore, we attempt to use the algorithm to extract structural features of off-line handwritten characters. Experiment results show that the algorithm is not only feasible for extraction of structural features of characters, but also exhibits good performance. The proposed method can provide a new approach to the research for extraction of structural features of characters.
This book constitutes the refereed conference proceedings of the 15th International Conference on Rough Sets, Fuzzy Sets, Data Mining and Granular Computing, RSFDGrC 2015, held in Tianjin, China in November 2015 as one of the co-located conference of the 2015 Joint Rough Set Symposium, JRS 2015. The 44 papers were carefully reviewed and selected from 97 submissions. The papers in this volume cover topics such as rough sets: the experts speak; generalized rough sets; rough sets and graphs; rough and fuzzy hybridization; granular computing; data mining and machine learning; three-way decisions; IJCRS 2015 data challenge.
With the view point of granular computing, the notion of a granule may be interpreted as one of the numerous small particles forming a larger unit. There are different granules at different levels of scale in data sets having hierarchical structures. Human beings often observe objects or deal with data hierarchically structured at different levels of granulations. And in real-world applications, there may exist multiple types of data in interval information systems. Therefore, the concept of multi-scale interval information systems is first introduced in this paper. The lower and upper approximations in multi-scale interval information systems are then defined, and the accuracy and the roughness are also explored. Monotonic properties of these rough set approximations with different levels of granulations are analyzed with illustrative examples.
We summarize our observations on utilizing generalized decision functions to define dependencies between attributes in decision systems. We refer to well-known criteria for attribute selection and less-known results linking generalized decisions with the notions of multivalued dependency and conditional independence. We formulate the problem of finding the simplest ensembles of subsets of attributes which allow to retrieve original decision values of considered objects by intersecting the sets of possible decisions induced by particular attributes.
In many fields including medical research, e-business and road transportation, data may vary over time, i.e., new objects and new attributes are added. In this paper, we present a method for dynamically updating approximations based on rough fuzzy sets under the variation of objects and attributes simultaneously in fuzzy decision systems. Firstly, a matrix-based approach is proposed to construct the rough fuzzy approximations on the basis of relation matrix. Then the method for incrementally computing approximations is presented, which involves the partition of the relation matrix and partly changes its element values based the prior matrices’ information. Finally, an illustrative example is employed to validate the effectiveness of the proposed method.
Multi-granularity thinking, computation and problem solving are effective approaches for human being to deal with complex and difficult problems. Deep learning, as a successful example model of multi-granularity computation, has made significant progress in the fields of face recognition, image automatic labeling, speech recognition, and so on. Its idea can be generalized as a model of solving problems by joint computing on multi-granular information/knowledge representation (MGrIKR) in the perspective of granular computing (GrC). This paper introduces our research on constructing MGrIKR from original datasets and its application in big data processing. Firstly, we have a survey about the study of the multi-granular computing (MGrC), including the four major theoretical models (rough sets, fuzzy sets, quotient space, and cloud model) for MGrC. Then we introduce the five representative methods for constructing MGrIKR based on rough sets, computing with words(CW), fuzzy quotient space based on information entropy, adaptive Gaussian cloud transformation (A-GCT), and multi-granularity clustering based on density peaks, respectively. At last we present an MGrC based big data processing framework, in which MGrIKR is built and taken as the input of other machine learning and data mining algorithms.
Hazard monitoring systems play a key role in ensuring people's safety. The problem of detecting dangerous levels of methane concentration in a coal mine was a subject of IJCRS' 15 Data Challenge competition. The challenge was to predict, from multivariate time series data collected by sensors, if methane concentration reaches a dangerous level in the near future. In this paper we present our solution to this problem based on the ensemble of Deep Neural Networks. In particular, we focus on Recurrent Neural Networks with Long Short-Term Memory (LSTM) cells.
The generalization of Pawlak rough set model always attracts the attentions of the researchers in the rough set society. In this paper, we propose a new subsystem-based definition of generalized rough set model and disclose the corresponding properties. We also discuss the interrelationships between our definition and the existing ones, the outputs show that our definition is effective and reasonable.
Appropriate measures are important for evaluating the performance of a classifier. In existing studies, many performance measures designed for two-way decisions based classification are applied to three-way decisions based classification directly, which may result in an incomprehensive evaluation. However, there is a lack of systematically research on the performance measures for three-way decisions based classification. This paper introduces some numerical measures and graphical measures for three-way decisions based binary classification.
Dominance-based rough set approach (DRSA) has been adopted in solving various multiple criteria classification problems with positive outcomes; its advantage in exploring imprecise and vague patterns is especially useful concerning the complexity of certain financial problems in business environment. Although DRSA may directly process the raw figures of data for classifications, the obtained decision rules (i.e., knowledge) would not be close to how domain experts comprehend those knowledge—composed of granules of concepts—without appropriate or suitable discretization of the attributes in practice. As a result, this study proposes a hybrid approach, composes of DRSA and a multiple attributes decision method, to search for suitable approximation spaces of attributes for gaining applicable knowledge for decision makers (DMs). To illustrate the proposed idea, a case of life insurance industry in Taiwan is analyzed with certain initial experiments. The result not only improves the classification accuracy of the DRSA model, but also contributes to the understanding of financial patterns in the life insurance industry.
We summarize the data mining competition associated with IJCRS’15 conference – IJCRS’15 Data Challenge: Mining Data from Coal Mines, organized at Knowledge Pit web platform. The topic of this competition was related to the problem of active safety monitoring in underground corridors. In particular, the task was to design an efficient method of predicting dangerous concentrations of methane in longwalls of a Polish coal mine. We describe the scope and motivation for the competition. We also report the course of the contest and briefly discuss a few of the most interesting solutions submitted by participants. Finally, we reveal our plans for the future research within this important subject.
We investigate properties of approximation operators being closure and topological closure in a framework of sixteen pairs of dual approximation operators, for the study of covering based rough sets. We extended previous results about approximation operators related with closure operators.
We describe our submission to the IJCRS'15 Data Mining Competition, where the objective is to predict methane outbreaks from multiple sensor readings. Our solution exploits a selective naive Bayes classifier, with optimal preprocessing, variable selection and model averaging, together with an automatic variable construction method that builds many variables from time series records. One challenging part of the challenge is that the input variables are not independent and identically distributed (i.i.d.) between the train and test datasets, since the train data and test data rely on different time periods. We suggest a methodology to alleviate this problem, that enabled to get a final score of 0.9439 (team marcb), second among the 50 challenge competitors.
Coal mining operation continuously balances the trade-off between the mining productivity and the risk of hazards like methane explosion. Dangerous methane concentration is normally a result of increased cutter loader workload and leads to a costly operation shutdown until the increased concentrations abate.We propose a simple yet very robust methane warning prediction model that can forecast imminent high methane concentrations at least 3 minutes in advance, thereby giving enough notice to slow the mining operation, prevent methane warning and avoid costly shutdowns.Our model is in fact an instance of the generic prediction framework able to rapidly compose a predictor of any future events upon the aligned time series big data. The model uses fast greedy backward-forward search applied subsequently upon the design choices of the machine learning model from the data granularity, feature selection, filtering and transformation up to the selection of the predictor, its configuration and complexity.We have applied such framework to the methane concentration warning prediction in real coal mines as a part of the IJCRS' 2015 data mining competition and scored 3rd place with the performance just under 85%. Our top model emerged as a result of the rapid filtering through the large amount of sensors time series and eventually used only the latest 1 minute of aggregated data from just few sensors and the logistic regression predictor. Many other model setups harnessing multiple linear regression, decision trees, naive Bayes or support vector machine predictors on slightly altered feature sets returned nearly equally good performance.
FRNN (Fuzzy Rough Nearest Neighbor) algorithm has exhibited good performance in classifying data with inadequate features. However, FRNN does not perform well on imbalanced data. To overcome this problem, this paper introduces a combination method. An improved SMOTE method is adopted to balance data and FRNN is applied as the classification method. Experiments show that the combination method can obtain a better result rather than classical FRNN algorithm.