基于改进PageRank的调用链异常节点定位研究
2021-05-16陈乐纪炎明肖忠良
陈乐 纪炎明 肖忠良









摘 要:微服务、云原生是当前信息系统发展的主流方向,给信息系统带来高可用的同时也让IT系统变得前所未有的复杂,这对IT运维工作带来了巨大的挑战。机器学习正是当前时代下应对复杂系统和海量信息的可选措施。文章将讨论基于PageRank的算法在接口服务调用链上定位异常节点,并且经过测试,可在错综复杂的调用关系上实现快速准确的异常定位。
关键词:IT运维;机器学习;PageRank;调用链分析;异常节点定位
中图分类号:TP18 文献标识码:A文章编号:2096-4706(2021)22-0059-04
Abstract: Microservices and cloud native are the mainstream development directions of current information systems. While bringing high availability to information systems, they also make IT systems more complex than ever before, which bring huge challenges to IT operation and maintenance. Machine learning is an optional measure to deal with complex systems and massive amounts of information in the current era. This paper will discuss the algorithm based on PageRank to locate the exception node on the interface service call chain, and after testing, it can realize fast and accurate exception location on the complex call relationship.
Keywords: IT operation and maintenance; machine learning; PageRank; call chain analysis; abnormal node location
0 引 言
當前的IT系统大多以采用微服务+云原生的架构[1]。微服务架构是将一个复杂的应用拆解成多个独立自治的服务,服务之间以松耦合的形式交互。如此部署应用的优势很明显,即业务逻辑清晰、部署简单、可拓展、高可用等等。以中国移动某省CRM为例,2018年完成了分布式5层云化IT系统,设计的系统极为庞大[2]。其中某订单系统的规模如图1所示。
微服务的劣势也随着IT系统规模的扩大而愈发明显,各个组件的调用关系错综复杂,对运维的压力也与日俱增。为了掌握系统的实时状态,大多数IT厂商会建设集中化监控系统,在IT系统中部署采集agent,实时采集系统各个层级,比如应用、中间件、主机等的告警信息。这从原来缺少信息的极端,走到信息过载的另一个极端。同时,高度复杂的IT系统架构意味着,一旦某个局部组件发生异常,故障信息就极易在短时间内扩散,触发大量告警。大量的告警信息中存在巨大的冗余,会淹没掉真正有用的信息。……
