大数据中数据挖掘模型的模糊改进聚类算法
2020-08-04李小红常振云
李小红 常振云



摘 要: 在大数据的数据挖掘模型中,普遍采用模糊聚类算法进行数据分析。常用的模糊C均值聚类算法即FCM聚类算法,具有较多明显缺点,如抗噪性偏低、收敛速度慢、聚类数目无法自动确定等。常用的增量式模糊聚类方法通常在原有的以一个中心点为集群代表的基础上,改为选取多中心点进行增量式聚类算法的分析。但是,通过这样的算法进行数据分析也存在一定的问题,主要表现在其中心点选择是固定的,灵活性很差。基于以上原因,文中将对原有基础算法做出改进,主要对大数据中数据挖掘模型的增量型模糊聚类算法做出分析,经实践验证,改进后算法切实可行,普适性较强。
关键词: 增量型模糊聚类; 大数据; 数据挖掘模型; 聚类算法; 余弦相似度; 隶属度矩阵
中图分类号: TN911.1?34 文献标识码: A 文章编号: 1004?373X(2020)03?0177?06
Improved fuzzy clustering algorithm for data mining model in big data
LI Xiaohong, CHANG Zhenyun
(School of Information Science and Engineering, Tianshi College, Tianjin 301700, China)
Abstract: The fuzzy clustering algorithm is widely used in data mining model of big data for data analysis. The commonly used fuzzy C?means clustering algorithm, also known as FCM clustering algorithm, has obvious disadvantages, for instance, the noise immunity is poor, the convergence speed is slow, and the number of clusters cannot be determined automatically. In the commonly used incremental fuzzy clustering algorithm, multi?center points are selected for incremental clustering algorithm analysis instead of taking one center point as the cluster representative as before. However, there are still certain problems in the algorithm in the process of data analysis, mainly because the selection of the center point is fixed, resulting the poor flexibility. In view of the above, the existing basic algorithm will be improved, and the incremental fuzzy clustering algorithm for data mining model in big data will be mainly analyzed. The practice shows that the improved algorithm is feasible and universal.
Keywords: incremental fuzzy clustering; big data; data mining model; clustering algorithm; cosine similarity; membership matrix
0 引 言
社会在不断发展和进步,信息时代的到来既让人们享受到了应有的便利,同时也遭受大量信息侵袭的困扰[1]。因此,如何在繁多的数据之中快捷、高效、精确地选取有用信息,成为当下亟需解决的问题。在大数据时代,建立数据挖掘模型就是在庞杂的数据当中对信息进行挖掘,以达到更为高效的信息筛选与获取的目的。在数据挖掘模型之中,聚类算法作为一种常用算法,主要是将数据进行多个集群划分,通过多个不同集群的相似度对比,进而进行数据的选择。本文研究的主要是大数据中数据挖掘模型的模糊改进聚类算法,也即增量型模糊聚类算法,该算法主要依据最小权重阈值展开。……
