Accurate particle identification is an ongoing task in the European organization for nuclear research, known as CERN where the challenge remains that targeted particles/events represent tiny minorities in front of the overwhelming presence of common particles such as protons. This paper presents a directed undersampling using an active learning method named DUAL to handle the high imbalance problem present in the particle identification dataset. The proposed approach was used to reduce the training set size while maintaining classifiers' performance. Compared against various imbalance learning approaches, the experimental results show that using DUAL as a data reduction technique with a random forest classifier enhances classification performance in terms of Macro- $$F_1$$ score and decreases the training time needed to train the models, which is very relevant while dealing with large-scale datasets. Despite being experimented only with particle identification dataset, we believe that DUAL could be adopted as a generic method for multi-class imbalanced classification problems with big data scale difficulties.
更多
查看译文
关键词
Active learning,Directed undersampling,Imbalance learning,Multi-class classification,Particle identification