基于决策树的钓鱼网页的识别方法
2017-12-13魏盛娜盛超
魏盛娜 盛超
摘要:现如今许多不法分子利用钓鱼网站盗取用户的个人信息,窃取用户的财产,对用户造成巨大损失。因此该文通过使用决策树学习算法,提取其中的关键词,分析并建立钓鱼网站特征模型,对未知网站进行判别。CART是一种决策树算法,但CART决策树的多数表决法会屏蔽小类数据类型的影响,因此该文根据这点对CART决策树进行改进,引入代价函数,不断地利用迭代和最小均方误差调整特征的权重增加惩罚。实验结果表明,改进后的决策树在对未知网站进行分析,成功地降低了负样本的错误率,提升了识别率。
关键词:决策树;URL识别;最小均方误差;代价函数
中图分类号:TP391 文献标识码:A 文章编号:1009-3044(2017)33-0079-02
Abstract: Now many criminals use phishing sites to steal the user's personal information, steal the user's property, causing huge losses to the user. Therefore, this paper uses the decision tree learning algorithm to extract the keywords, analyze and establish the phishing website feature model, and judge the unknown website. CART is a decision tree algorithm, but the majority voting method of CART decision tree will shield the influence of small class data type. Therefore, this paper improves the CART decision tree according to this point, introduces the cost function, and makes use of iteration and minimum mean square error Adjust the weight of the feature to increase the penalty. The experimental results show that the improved decision tree has successfully reduced the error rate of negative samples and improved the recognition rate in the analysis of unknown websites.
Key words: decision tree; URL identification; least-mean-square; cost function
1 背景
钓鱼网站通常是指伪装成合法网站,窃取用户提交的账号、密码等私密信息的网站。目前已出现10余种反钓鱼工具,本文选用决策树方法对钓鱼URL特征进行识别,国内外学者也提出了很多决策树的相关改进算法:
ID3算法是1986年由Quinlan提出的,是基于信息增益的选择[1] 。J.Ma[2]等人分析可疑URL 的词汇和主机属性采用词袋模型表示特征, 获得了成千上万的特征,运用特征匹配加上ID3算法檢测钓鱼网站。但ID3算法也存在缺陷,因为包含较多属性值的特征所含的信息增益一般会越高,所以ID3优先会选择有较多属性值的特征,从而构建的决策树往往不是最优的,只可以用于处理离散数据,不能用于处理连续数据。
C4.5算法是Quinlan本人对ID3算法的改进[3],引入了信息增益比(GainRatio)作为选择的准则。……
