基于局部密度的高效聚类算法研究
2016-01-19刘立波卜旭松
周 杰,刘立波,卜旭松
(宁夏大学 数学计算机学院,宁夏 银川 750021)
基于局部密度的高效聚类算法研究
周杰,刘立波,卜旭松
(宁夏大学 数学计算机学院,宁夏 银川 750021)
摘要:针对基于密度聚类算法对参数值过于敏感、参数值难以设置的不足,提出一种基于局部密度的高效聚类算法,以有效解决基于密度聚类算法对密度层级不同数据集聚类准确率低的问题。该算法将局部密度概念和k-means算法相结合,在算法迭代过程中动态计算局部密度参数和簇核心距。对于分布在簇核心距之内的数据,采用基于划分的方法直接聚类;对于分布在簇核心距之外的数据采用基于局部密度方法进行聚类。实验结果表明,提出的算法在聚类精度和计算效率两方面均具有较好的性能。
关键词:k-means;局部密度;簇核心距;聚类
中图分类号:TP391
文献标志码:码:A
文章编号:号:2095-4824(2015)06-0021-05
收稿日期:2015-09-11
作者简介:周杰(1990-),女,宁夏银川人,宁夏大学数学计算机学院硕士研究生。
Abstract:In view of the limitation of density-based clustering algorithm that is sensitive to parameter settings and difficult to set the parameters, this paper proposes an efficient local density-based clustering algorithm for the problem that the density-based clustering algorithm is not accurate to the data set with different density-levels. The proposed algorithm combines the local density concepts and K-means algorithms and dynamically calculates the local-density parameters and core-distance of the clusters in the process of the iteration. The data inside the core-distance is clustered into the clusters directly by the partition-based method and the data outside the core distance of the clusters is clustered by the local density method. Experimental results indicate that the proposed method exhibits better performance in both efficiency and accuracy.
聚类[1-2]作为一种有效的数据挖掘工具已经成功应用于许多领域,如网络安全[3]、生物学、商务智能和Web搜索等多个方面。聚类是把一个数据集划分成多个组或簇的过程,使得簇内高度相似,而簇间相异。若待分类数据集中不含噪声数据,采用基于划分的聚类算法具有较好的聚类效果,但当数据集中存在较多噪声数据时,基于划分的聚类算法聚类质量较低[4-5]。尽管基于密度的聚类算法能通过设置全局阈值方式消除噪声数据对聚类准确率的影响[6-8],但全局密度阈值不能很好地刻画数据集的内在聚类结构,当数据集中簇的密度层级相差较大时,全局密度阈值选择非常困难,使得聚类质量难以保证。……
