文本分类中支持向量机研究
2019-10-21何焱
何焱
摘 要:随着我国现代科技的快速发展,文本分类逐渐在信息化技术与数字化技术领域得到重视。利用计算处理系统处理文本信息,能够有效提升文本分类的质量与效率,提升数据信息的利用率,从而促进信息化技术的普及。而支持向量机是处理文本内容,加强文本分类速度,并通过文档建模、中文分词、分类器评估等形式,构建出的行之有效的统计语言模型,它可以推动文本分类工作的发展。本文结合国内外研究现状,探析文本分类内涵及支持向量机原理,提出基于支持向量机的文本分类算法。
关键词:文本分类;支持向量机;统计语言模型
中图分类号:TP391.1文献标识码:A文章编号:1003-5168(2019)29-0008-03
Research on Support Vector Machine in Text Categorization
HE Yan
(Zunyi Medical and Pharmaceutical College,Zunyi Guizhou 563002)
Abstract: With the rapid development of modern science and technology in China, text classification has gradually gained attention in the field of information technology and digital technology. The use of the computing processing system to process text information can effectively improve the quality and efficiency of text classification, improve the utilization of data information, and promote the popularization of information technology. The support vector machine is a statistical language model that is effective in processing text content, enhancing text classification speed, and constructing it through document modeling, Chinese word segmentation, and classifier evaluation, which can promote the development of text classification work. Based on the research status at home and abroad, this paper analyzed the text classification connotation and the principle of support vector machine, and proposed a text classification algorithm based on support vector machine.
Keywords: text classification;support vector machine;statistical language model
大数据时代,数据信息技术逐渐成为推动我国社会经济快速发展的重要途径,同时也是加速城市智能化、现代化发展的关键手段。随着云计算、物联网等技术的快速发展,数字信息技术得到我国社会各领域的广泛重视。然而,如何提升现代信息的利用效率,凸显数字信息的时代价值呢?人们需要从文本分类手段出发,整合现有的文本信息,使其成为大数据技术及云计算技术的重要组成部分。
1 国内外研究现状
20世纪中叶,文本分类得到了迅速的发展,并利用知识工程理论实现了人为定制分类体系的建构目标。而在21世纪初,相關专家和学者开始尝试利用机器学习的形式实现对文本的分类。这种不需要人为干预的文本分类方法得到快速的发展,并逐渐成为文本分类的主要研究内容[1-3]。……
