融合贝叶斯深度学习的计算机大数据频繁项挖掘算法
2020-01-21刘兴建原振文
刘兴建 原振文



摘要:随着数据每天呈指数级增长,频繁项集挖掘的效率和可伸缩性问题变得更加严重。因此,提出融合贝叶斯深度学习的计算机大数据频繁项挖掘算法(Sequential growth),并在MapReduce框架上实现。为了测试算法的性能,在具有大型数据集的MapReduce框架上进行了不同方面的实验。结果表明,Sequential growth算法具有良好的效率和可扩展性,尤其在处理大数据和长项目集时。
关键词:频繁模式挖掘;频繁项集挖掘;MapReduce;贝叶斯深度学习;可扩展性
中图分类号:TP301文献标志码:A
文章编号:2095-5383(2020)04-0038-05
Frequent Item Mining Algorithm of Computer Big Data
based on Bayesian Deep Learning
LIU Xingjian1, YUAN Zhenwen2
(1.Department of Computer Application Technology, Guangdong Business and Technology University, Zhaoqing 526040, China;2. Artillery Academy, National University of Defense Technology, Changsha 410111, China)
Abstract:
Frequent itemsets mining (FIM) is an important research topic because it is widely used in the real world to find frequent itemsets and mine human behavior patterns. As data grows exponentially every day, the efficiency and scalability issues of frequent itemset mining have become more serious. Therefore, a frequent item mining algorithm (Sequential growth) of computer big data that integrates Bayesian deep learning was porposed in this paper, and it was implemented on the MapReduce framework. In order to test the performance of the algorithm, experiments in different aspects were carried out on the MapReduce framework with large data sets. The results show that the Sequential growth algorithm has good efficiency and scalability, especially when processing large data and long itemsets.
Keywords:
frequent pattern mining; Frequent Itemset Mining(FIM); MapReduce; Bayesian deep learning; scalability
频繁模式挖掘(frequent pattern mining)是数据挖掘的最重要技术之一,它在各种领域中都有广泛应用,例如市场篮分析、网页点击分析、患者路径分析、DNA序列发现以及最近的无缺陷率改进等[1-4]。频繁项集挖掘(Frequent Itemset Mining,FIM)问题首次出现在1995年,这是基于1994年的工作进行的扩展研究工作,Gupta等[5]介绍了2种基于Apriori的算法,分别称为AprioriAll和AprioriSome,FIM问题逐渐引起了人们的广泛关注,并成为数据挖掘的重要研究课题之一。Apriori算法也成为FIM最重要的概念之一,并以此为基础拓展了许多算法[6-7]。FIM的另一种方法称为Eclat。与Apriori使用自下而上的广度优先搜索策略不同,Eclat使用具有交集的广度优先方法来挖掘频繁项集[8]。 Eclat将事务数据库转换为其垂直格式。 每个项目集与其封面(也称为tid-list)一起存储。垂直数据库D′定义为:D′=ij,Cij=tidij∈X,tid,X∈D,其中Cij是ij的最新列表。在垂直数据库中,项目集Y的支持可以通过与任意2个子集的tid-list相交来计算。它可被表示为supportY=∩tj∈YCij。
基于Apriori的算法以世代扫描方式挖掘频繁项集。将串行Apriori算法应用于MapReduce框架的最便捷方法是MapReduce作业的多次迭代,以生成候选序列并扫描其在数据库中的包含度。……
