APP下载

基于K-means算法的企业信用无监督分类研究

2021-09-14施天虎韦诗玥

电脑知识与技术 2021年22期
关键词:分类

施天虎 韦诗玥

摘要:企业信用分类的应用,能够为商业银行降低信贷业务的风险,随着市场竞争的不断加剧,机器学習和大数据的应用,越来越多的计量方法不断革新,并广泛运用到信用分析领域。本文设计了一个基于K-means算法的企业信用无监督分类方法,通过对企业信息进行大数据分析,提取企业信用相关的内容,再使用K-means算法对企业数据进行聚类,对目标企业根据其聚类所在簇来评估信用等级,以此对企业的信用进行分类。

关键词:企业信用;信贷风险;K-means算法;分类;特征选择

Abstract: The application of corporate credit classification can reduce the risk of credit business for commercial banks. With the continuous intensification of market competition, the application of machine learning and big data, more and more measurement methods continue to innovate and are widely used in the field of credit analysis. This paper designs an unsupervised classification system for corporate credit based on the K-means algorithm. Through big data analysis of corporate information, the content related to corporate credit is extracted, and then the K-means algorithm is used to cluster the companies, and the target companies are based on their The clusters where the clusters are located are used to evaluate the credit rating and thus classify the credit of the enterprise.

Key words: Corporate credit; Credit Risk; K-means algorithm; classification; Feature selection

1引言

金融行业积累了大量的企业脱敏数据信息,企业的有效划分及标识在企业信用评估、企业风险监测中具有重要作用并受到各大平台的重点关注[1]。金融场景中企业作为信贷主体的数据覆盖互联网、政府、线上应用等来源的方方面面,数据量大,来源广泛、涉及企业的维度丰富[2]。企业信用分类的应用,为商业银行降低企业信贷业务风险,创新风险管理理念,探索出一条行之有效的解决办法[3]。随着大数据、人工智能的发展和市场竞争日益加剧,大量基于机器学习的信用评估分类方法提出并广泛应用于企业信用分析[4]。本文将企业脱敏数据信息进行特征选择,提取信用分类相关的内容,再使用K-means算法对数据进行聚类,按聚类簇划分信用等级。

2 关键技术

2.1 K-means算法

2.2 特征选择

特征选择是重要的数据预处理方法,在数据中选出重要特征可以降低数据维度、去除多余的变量,提高算法的精度和效率。……

登录APP查看全文

猜你喜欢

分类
分类算一算
垃圾分类的困惑你有吗
星星的分类
我给资源分分类
垃圾分类,你准备好了吗
按需分类
教你一招:数的分类
说说分类那些事
给塑料分分类吧