MapReduce编程模型中key值二次分类算法
2018-05-02刘帅
刘帅

摘 要: MapReduce编程模型是分布式计算中最常用的编程模型,其主要目的是将单个巨大计算任务分割成多个小计算任务,并分别交由不同的计算机去处理。MapReduce将任务分成map阶段和reduce阶段,每个阶段都是用key/value键值对作为输入和输出。针对MapReduce中Map数量少,Reduce数量多的情况,文章将Map阶段任务中的Key值进行二次划分,提出一种MapReduce编程模型中Key二次分类的方法。实验,证明该方法能够在原有基础上提高数据处理效率。
关键词: MapReduce; Key/value; 二次分类
中图分类号:TP301 文献标志码:A 文章编号:1006-8228(2018)03-58-02
Two times classification algorithm of Key value in MapReduce programming model
Liu Shuai
(Department of Computer Application, Xinzhou Vocational and Technical College, Xinzhou, Shanxi 034000, China)
Abstract: MapReduce programming model is the most commonly used programming model in the distributed computing. It divides a single huge computing task into multiple small computing tasks, which are processed by different computers respectively. MapReduce divides the task into the Map phase and the Reduce phase, each of which is used as input and output with the key/value key value pair. In view of the fact that the number of Map in MapReduce is small and the number of Reduce is large, the Key value of Map phase task is divided in two times, and a method of two times classification of Key value in MapReduce programming model is proposed. Experiments show that the method can improve the efficiency of data processing on the original basis.
Key words: MapReduce; Key/Value; two times classification
0 引言
MapReduce是由google公司提出的一種并行计算框架[1-2],其“分而治之”的思想被广泛用于Hadoop与Spark分布式计算平台。MapReduce处理任务过程中,将任务分为Map阶段与Reduce阶段,Map阶段负责将数据按一定规则整理成
针对MapReduce编程模型,目前大多数的研究是在MapReduce编程思想结合现有分布式平台的使用。对于MapReduce自身思想的改进并不很多[5-6]。本文通过分析MapReduce编程模型中Map阶段与Reduce阶段任务的处理过程,发现会存在Key值数量少但Value值数量多的情况,针对此类任务处理的负载不均衡现象,给出一种MapReduce编程模型中key值二次分类的算法,通过增加Key值的标记,使Value值能够更加均匀的分布在集群不同处理节点上,提高了集群节点的利用率与数据处理效率。
1 MapReduce模型
1.1 基本原理
MapReduce编程模型[7-8]是一个处理和生成超大数据集的算法模型的相关实现。此模型首先创建一个Map函数处理一个基于
