基于语义的Web招聘信息抽取关键技术的研究
2019-10-21张晓孪王西锋
张晓孪 王西锋

摘 要: 随着互联网技术的应用,大量求职者期望能从招聘网站中快速、精准获取有用信息,因此分析并抽取这些网站中的招聘信息具有实际应用的价值。针对Web信息抽取技术在招聘信息系统中的应用,提出了一种基于语义的Web招聘信息抽取的方法,首先是构建主题蜘蛛程序抓取网页,然后对预处理过的网页中的命名实体进行识别。经测试采用本文提出的方法进行信息抽取是可行的,命名实体识别的准确率和召回率能达到71%以上。
关键词: 语义; Web招聘信息抽取; 蜘蛛程序; 命名实体识别
中图分类号: TP391
文献标志码: A
文章编号:1007-757X(2019)06-0069-02
Abstract: With the application of the Internet technology, a large number of job seekers expect to obtain useful information quickly and accurately from the recruitment Website. That the recruitment information extraction provides for the majority of job seekers correct employment information is of great importance. Aiming at the application of Web information extraction technology in recruitment information system, this paper proposes a Web recruitment information extraction method based on semantic. The first is to build a topic spider program to crawl the Web page, and then to identify named entity from pre-processed Web pages. After testing, it is feasible to use the method proposed in this paper to extract the information, and the accuracy and recall rate of named entity recognition are all above 71%.
Key words: Semantic; Web recruiting information extraction; Spider program; Named entity recognition
0 引言
隨着互联网技术的应用与普及,越来越多的企业与公司通过网站发布相关招聘信息,这种招聘方式显现出信息量大、信息增长速度快和信息处理难度大等弊端,解决这些问题的关键就是从网页中抽取出人们感兴趣的信息。面对这些海量招聘信息,大量求职者期望能从这些网站中快速、精准的获取有用信息,对他们求职提供参考,因此招聘信息抽取为广大求职者提供正确的就业信息有着非常重要的意义,具有实际应用的价值。
虽然国内外学者已对网络招聘系统做了大量研究,但是却很少涉及对网络招聘信息的抽取、挖掘和分析。本文针对Web信息抽取技术在招聘信息系统中的应用,提出了一种基于语义的Web招聘信息抽取的方法,其目标是将分散在海量Web页面中的动态变化的招聘信息抽取出来,以结构化、语义清晰的形式提供给求职者,帮助求职者正确了解当前的就业趋势,尽快找到称心满意的工作,并进一步提高网络信息中数据的利用率。……
