微阵列数据广泛而成功地应用于生物医学的癌症分类研究.一个典型的微阵列数据集包含大量(通常成千上万,甚至数十万)的基因、相对少量(往往不足一百)的样本.在这成千上万的基因中,仅仅一少部分基因对癌症分类有贡献.因而,对于癌症分类来说,最重要的一个问题就是识别出对癌症分类最有贡献的基因.这一识别过程称为基因选择.基因选择在统计模式识别、机器学习和数据挖掘领域已得到广泛研究.介绍基因选择问题所涉及到的相关背景知识和基本概念;全面地回顾统计学、机器学习和数据挖掘领域对基因选择问题的解决方法;通过实验展示了几种典型算法在微阵列数据上的性能;指出当前存在的问题和未来的研究方向.
Gene selection is one of the important and frequently used techniques for microarray data classification. In this paper, we introduce a new metric to measure gene-class relevance and gene-gene redundancy. The new metric is based on Grey Relational Analysis (GRA), called Grey Relational Grade (GRG), and never used in gene selection before. Based on the GRG, we develop a new gene selection method, which uses GRG to group similar genes to clusters, and then select informative genes from each cluster to avoid redundancy. Experiments on public data sets demonstrate the effectiveness of the proposed method.
Accurately estimating language model is important to improve the performance of information retrieval. The key problems include solving synonymy and polysemy problem, and smoothing the seen term or not seen term in a document. In this paper, we propose a new method for topic language model. First, concept-based clustering is performed using improved fuzzy c-means. The clustering result is considered as the topics of document collections. The probability of a document generating the topics is estimated by the similarity between the document and each cluster. Then, the probability of the topics generating words is estimated using Expectation Maximization algorithm. At last, we integrate the above algorithms into aspect model to form our topic language model. This new language model accurately describes the distribution probability of the words in different topics and the probability of a document generating a topic. Moreover, it can solve synonymy and polysemy problems. The new method is evaluated on TREC 2004/05 Genomics Track collections. Experiments show that the retrieval performance is greatly improved by the new method compared with the simple language model.
Selecting a small number of discriminative genes from thousands of genes in microarray data is very important for accurate classification of diseases or phenotypes. In this paper, we provide more elaborate and complete definitions of feature relevance and develop a novel feature selection method, which is based on relevance analysis and discernibility matrix to select small enough genes and improve the classification accuracy. The extensive experimental study using microarray data shows the proposed approach is very effective in selecting genes and improving classification accuracy.
Feature (gene) selection is a frequently used preprocessing technology for successful cancer classification task in microarray gene expression data analysis. Widely used gene selection approaches are mainly focused on the filter methods. Filter methods are usually considered to be very effective and efficient for high-dimensional data. This paper reviews the existing filter methods, and shows the performance of the representative algorithms on microarray data by extensive experimental study. Surprisingly, the experimental results show that filter methods are not very effective on microarray data. We analyze the cause of the result and provide the basic ideas for potential solutions.
Rough set theory has been widely and successfully used in data mining, especially in classification field. But most existing rough set based classification approaches require computing optimal attribute reduction, which is usually intractable and many problems related to it have been shown to be NP-hard. Although approximate algorithms exist, they also tend to be computationally expensive. This paper presents a novel rough set method for classification, which does not require computing attribute reduction. It stepwise investigates condition attributes and outputs the classification rules induced by them, which is just like the strategy of "on the fly". The theoretical analysis and the empirical study show that the proposed method is effective and efficient. Index Terms—rough set, attribute reduction, data mining, classification
粗糙集理论作为一种处理不完备信息的有力工具,已广泛应用于人工智能的许多领域,特别是数据挖掘和知识发现领域。文章将基于粗糙集理论的数据挖掘技术应用于液体火箭发动机故障诊断的知识发现,在智能诊断的知识自动获取方面取得较好的结果。
By analyzing the characteristics of the liquid rocket engine and the rough set method,the application of rough set theory to data mining for the hotfiring test data of liquid rocket engine is presented.By investigating 4 groups of hot-firing test data of a certain engine,24 original attributions are reduced to 2 attributions,213 records are reduced to 5 or 6 records.All possible critical decisive combinations for fault detection and diagnosis of liquid rocket engine are obtained,fault detection and diagnosis for liquid rocket engine are implemented.
In this article we describe a method for selecting informative genes from microarray data. The method is based on clustering, namely, it first find similar genes, group them and then select informative genes from these groups to avoid redundancy. A new gene similarity measure based on grey relational analysis (GRA), called grey relational grade (GRG), is used in clustering. Experiments on three public data sets demonstrate the effectiveness of our method
Gene selection is a common task in microarray data classification. The most commonly used gene selection approaches are based on gene ranking, in which each gene is evaluated individually and assigned a discriminative score reflecting its correlation with the class according to certain criteria, genes are then ranked by their scores and top ranked ones are selected. Various discriminative scores have been proposed, including t-test, S2N,RelifF, Symmetrical Uncertainty andχ2-statistic. Among these methods, some require abundant data and require the data follow certain distribution, some require discrete data value. In this work, we propose a gene ranking method based on Grey Relational Analysis (GRA) in grey system theory, which requires less data, does not rely on data distribution and is more applicable to numerical data value. We experimentally compare our GRA method with several traditional methods, including Symmetrical Uncertainty, χ2-statistic and ReliefF. The results show that the performance of our method is comparable with other methods, especially it is much faster than other methods.
粒度计算涵盖了所有在处理问题过程中使用粒度的理论、方法、技术和工具。本文首先简要地介绍了粒度计算的基本思想、基本问题以及它的三个主要模型(模糊集、粗糙集和商空间),然后综述了粒度计算在数据挖掘中的应用。
For traditional methods of classification, the results are very explicit. When these methods are applied to fuzzy objects, the explicit results do not accord with the fact. This paper studies the basic concepts and methods in fuzzy set , and offers a classification model based on fuzzy integrated estimation which is applied to corporation credit evaluation.
AbstrcatBy analyzing the working characteristics of the liquid rocket engine, strategies of applying data mining from the point of view of the data warehouse for FDD of LRE are proposed. The data mining methods possibly applicable to different topics in the fault detection and diagnosis of LRE are compared. Primary research results show that such data mining methods as clustering, classification, association, time-series analysis and outlier analysis are feasible in the FDD of LRE.
引言数据仓库是信息技术领域发展迅速的一门新技术。从20世纪90年代初数据仓库概念的提出到现在短短的十几年时