In the big data era, the large volume and high dimensionality of data challenge many data mining algorithms. Since many classification algorithms are sensitive to data distribution, removing redundant or noisy features and instances can greatly improve their performance. Feature selection and instance selection are two major data reduction techniques that are inherently interconnected. This article proposes an auxiliary-optimization-assisted constrained multiobjective optimization method that concurrently tackles feature and instance selection, with two key constraints: forcing the obtained subsets to have better classification performance than that of using all features and instances under the given classification algorithm, and bounding the worst-class error. A simple but effective initialization method is designed to sample initial solutions relatively uniformly across different regions of the objective space, to provide a diverse and high-quality set of initial feature and instance subsets. An auxiliary-optimization-based search approach is proposed to fully utilize useful infeasible solutions, which can further reduce the number of selected features and instances without compromising classification performance. The proposed method is compared with a number of promising methods on 20 real-world classification datasets, and the experimental results show that it is generally better than those methods. Additionally, the proposed method offers a significant advantage whereby the majority of its solutions exhibit superior or comparable classification performance compared to using all original features and instances, while selecting no more than 30% of the features and 50% of the training instances across most datasets.
更多
查看译文
关键词
Filtering,Filters,Circuits and systems,Internet,Internet of Things,Communication systems,Product development,Electronic mail,Instant messaging,Electronic messaging,Classification,evolutionary multiobjective learning,feature selection,instance selection