早期糖尿病风险预测模型的比较研究
2021-07-11王成武晏峻峰
王成武 晏峻峰



摘 要:糖尿病是一种比较常见的慢性疾病,并且存在较长的无症状阶段。本文主要介绍了机器学习中的5种分类算法,分别是朴素贝叶斯、支持向量机、逻辑回归、决策树和集成分类器Random Forest,并在Weka数据挖掘平台上,对糖尿病数据进行挖掘分析,根据混淆矩阵、Kappa系数、ROC曲线、均方根误差以及相对绝对误差这几个性能指标对分类器效果进行分析,找到最适合糖尿病疾病预测的算法,为当今医疗行业其他疾病数据的挖掘分析提供思路。
关键词: 糖尿病;机器学习;集成分类器;数据挖掘;Weka
文章编号: 2095-2163(2021)01-0064-05 中图分类号:TP391 文献标志码:A
【Abstract】Diabetes is a relatively common chronic disease, and there is a long asymptomatic stage. This article mainly introduces five classification algorithms in machine learning, which are Naive Bayes, Support Vector Machine, Logistic Regression, Decision Tree, and Random Forest, an integrated classifier. On the Weka data mining platform, the diabetes data is mined and analyzed. The effect of the classifier is analyzed according to the confusion matrix, Kappa coefficient, ROC curve, root mean square error and relative absolute error, and the most suitable algorithm for diabetic disease prediction is achieved, which could provide ideas for the current medical industry data mining.
【Key words】diabetes; machine learning; integrated classifier; data mining; Weka
0 引 言
糖尿病是一种终身疾病,可引发心脏病、血管疾病等并发症[1],不仅影响了患者的生活质量,也会带来相应的经济负担,所以进行早期糖尿病风险预测具有十分重要的意义。
作为重要的数据挖掘技术,机器学习等人工智能技术,在糖尿病预测与治療上应用得很多。例如,Purushottam等人[2]分别用C4.5算法和Partial Tree算法自动提取糖尿病预测规则来预测患者的糖尿病风险。Santhanam等人 [3]用遗传算法对糖尿病数据集进行维数约简并利用支持向量机进行了糖尿病的预测。胡玮[4]基于改进邻域粗糙集和随机森林算法进行了糖尿病的预测研究。黄艳群等人[5]利用患者相似性建立了个性化糖尿病预测模型。
本文将机器学习技术应用在早期糖尿病风险预测数据集上,构建多种分类模型,通过各种性能评价指标对模型进行分析,选择最优分类模型,该模型可通过评估症状来检查用户患糖尿病的风险。……
