APP下载

大数据中数据挖掘模型的模糊改进聚类算法

2020-08-04李小红常振云

现代电子技术 2020年3期
关键词:大数据

李小红 常振云

摘  要: 在大数据的数据挖掘模型中,普遍采用模糊聚类算法进行数据分析。常用的模糊C均值聚类算法即FCM聚类算法,具有较多明显缺点,如抗噪性偏低、收敛速度慢、聚类数目无法自动确定等。常用的增量式模糊聚类方法通常在原有的以一个中心点为集群代表的基础上,改为选取多中心点进行增量式聚类算法的分析。但是,通过这样的算法进行数据分析也存在一定的问题,主要表现在其中心点选择是固定的,灵活性很差。基于以上原因,文中将对原有基础算法做出改进,主要对大数据中数据挖掘模型的增量型模糊聚类算法做出分析,经实践验证,改进后算法切实可行,普适性较强。

关键词: 增量型模糊聚类; 大数据; 数据挖掘模型; 聚类算法; 余弦相似度; 隶属度矩阵

中图分类号: TN911.1?34                       文献标识码: A                          文章编号: 1004?373X(2020)03?0177?06

Improved fuzzy clustering algorithm for data mining model in big data

LI Xiaohong, CHANG Zhenyun

(School of Information Science and Engineering, Tianshi College, Tianjin 301700, China)

Abstract: The fuzzy clustering algorithm is widely used in data mining model of big data for data analysis. The commonly used fuzzy C?means clustering algorithm, also known as FCM clustering algorithm, has obvious disadvantages, for instance, the noise immunity is poor, the convergence speed is slow, and the number of clusters cannot be determined automatically. In the commonly used incremental fuzzy clustering algorithm, multi?center points are selected for incremental clustering algorithm analysis instead of taking one center point as the cluster representative as before. However, there are still certain problems in the algorithm in the process of data analysis, mainly because the selection of the center point is fixed, resulting the poor flexibility. In view of the above, the existing basic algorithm will be improved, and the incremental fuzzy clustering algorithm for data mining model in big data will be mainly analyzed. The practice shows that the improved algorithm is feasible and universal.

Keywords: incremental fuzzy clustering; big data; data mining model; clustering algorithm; cosine similarity; membership matrix

0  引  言

社会在不断发展和进步,信息时代的到来既让人们享受到了应有的便利,同时也遭受大量信息侵袭的困扰[1]。因此,如何在繁多的数据之中快捷、高效、精确地选取有用信息,成为当下亟需解决的问题。在大数据时代,建立数据挖掘模型就是在庞杂的数据当中对信息进行挖掘,以达到更为高效的信息筛选与获取的目的。在数据挖掘模型之中,聚类算法作为一种常用算法,主要是将数据进行多个集群划分,通过多个不同集群的相似度对比,进而进行数据的选择。本文研究的主要是大数据中数据挖掘模型的模糊改进聚类算法,也即增量型模糊聚类算法,该算法主要依据最小权重阈值展开。……

登录APP查看全文

猜你喜欢

大数据
基于在线教育的大数据研究
“互联网+”农产品物流业的大数据策略研究
基于大数据的小微电商授信评估研究
大数据时代新闻的新变化探究
浅谈大数据在出版业的应用
“互联网+”对传统图书出版的影响和推动作用
大数据环境下基于移动客户端的传统媒体转型思路
基于大数据背景下的智慧城市建设研究
数据+舆情:南方报业创新转型提高服务能力的探索