基于间隔理论的过采样集成算法
2019-08-01张宗堂陈喆戴卫国
张宗堂 陈喆 戴卫国
摘 要:针对传统集成算法不适用于不平衡数据分类的问题,提出基于间隔理论的AdaBoost算法(MOSBoost)。首先通过预训练得到原始样本的间隔; 然后依据间隔排序对少类样本进行启发式复制,从而形成新的平衡样本集; 最后将平衡样本集输入AdaBoost算法进行训练以得到最终集成分类器。在UCI数据集上进行测试实验,利用Fmeasure和Gmean两个准则对MOSBoost、AdaBoost、随机过采样AdaBoost(ROSBoost)和随机降采样AdaBoost(RDSBoost)四种算法进行评价。实验结果表明,MOSBoost算法分类性能优于其他三种算法,其中,相对于AdaBoost算法,MOSBoost算法在Fmeasure和Gmean准则下分别提升了8.4%和6.2%。
关键词:不平衡数据;间隔理论;过采样方法;集成分类器;机器学习
中图分类号:TP181
文献标志码:A
Abstract: In order to solve the problem that traditional ensemble algorithms are not suitable for imbalanced data classification, Over Sampling AdaBoost based on Margin theory (MOSBoost) was proposed. Firstly, the margins of original samples were obtained by pretraining. Then, the minority class samples were heuristic duplicated by margin sorting thus forming a new balanced sample set. Finally, the finall ensemble classifier was obtained by the trained AdaBoost with the balanced sample set as the input. In the experiment on UCI dataset, Fmeasure and Gmean were used to evaluate MOSBoost, AdaBoost, Random OverSampling AdaBoost (ROSBoost) and Random UnderSampling AdaBoost (RDSBoost). The experimental results show that MOSBoost is superior to other three algorithm. Compared with AdaBoost, MOSBoost improves 8.4% and 6.2% respctively under Fmeasure and Gmean criteria.
英文关键词Key words: imbalanced data; margin theory; over sampling method; ensemble classifier; machine learning
0 引言
近些年,不平衡数据分类问题成为了机器学习的热点问题,它广泛存在于现实生产生活中,例如邮件过滤[1]、图像分类[2]、软件缺陷预测[3]、医疗诊断[4]、基因数据分析[5]等。对于二分类问题,不平衡数据中多类的样本数量远大于少类。传统的分类方法以总体分类精度为目标,忽视了类别不平衡性,从而导致少类样本分类准确率降低,然而少类样本往往具有较高的价值,这使得错分代价较大。
针对不平衡数据的处理方法大致分为算法层面和数据层面: 算法层面指构造新的算法或对原有算法进行改造以偏向少类; 数据层面主要是利用重采樣方法获得平衡样本集,再结合现有分类器进行分类。重采样方法,包括欠采样法和过采样法,形式上比较简练,且不影响分类器设计,因此得到了广泛的研究。……
