基于隐马尔科夫模型的spark作业异常分析
2018-07-28王欣周云才
王欣 周云才
摘要:随着大数据技术的不断发展,数据分析越来越受到人们的关注,Spark 作为大规模数据处理的快速通用的计算引擎,由于它的高速性而被各大商家应用于实际生产过程中。本文通过隐马尔科夫模型(HMM),选择在实际生产过程中,在进行海量的数据分析过程中出现的异常进行分析,以实际任务执行时的:内存溢出、垃圾回收异常、序列化异常为指标,根据实际出现异常时的提示,来确定HMM状态空间、确定相应的观测值、计算相关的参数,进而构建针对于Spark作业工作过程中的出现异常时的隐马尔科夫模型,用来揭示引发异常的类型,来对实际生产过程中出现此类问题时提供可靠的类型诊断。
关键词:spark;隐马尔科夫模型;内存溢出;异常;内存管理
中图分类号:TP31 文献标志码:A 文章编号:1009-3044(2018)11-0198-03
Spark Operation Anomaly Analysis Based on Hidden Markov Model
WANG Xin,ZHOU Yun-cai
(Yangtze University,Jingzhou 434023,China)
Abstract: With the continuous development of big data technology, data analysis has attracted more and more attention. Spark, a fast and universal computing engine for large-scale data processing, has been used by major merchants in the actual production process due to its high speed. This paper uses hidden Hidden Markov Model (HMM) to select the analysis of abnormalities that occur in the process of mass data analysis in the actual production process. When actual tasks are executed, memory overflow, garbage collection anomalies, and serialization anomalies are Indicators, according to the actual occurrence of abnormal prompts, to determine the HMM state space, determine the corresponding observations, calculate the relevant parameters, and then build a Hidden Markov model for exceptions in the Spark job process, to reveal The type of exception that is thrown to provide a reliable type diagnosis when such problems occur in the actual production process.
Key words:spark; hidden markov model; memory overflow; exceptions; memory management
1 概述
Spark是UC Berkeley计算机教授Ion Stoica 在2009年发起的,随后被Apache软件基金会接管的类似于Hadoop MapReduce的通用并行计算框架,是当前大数据领域最活跃的开源项目之一[1]。Spark是基于MapReduce计算框架实现的分布式计算,拥有Hadoop MapReduce所具有的优点;但不同于MapReduce。
在现实生活中隐马尔科夫模型广泛应用于图像处理、语音识别、模式识别、信息处理预测以及股票风险预测等领域,同时在机器异常状态预测、风险预警、风投决策等生产环境中都有广泛的应用。在大型的电商公司,对于日常的点击、收藏、购买等行为的日志分析尤为重要,而在针对于海量数据的分析过程中spark作为基于内存的一种大数据处理解决方案,越来越受到各大电商公司的关注[2]。……
