面向网站群的主题爬虫研究
2020-09-02徐昊沈江明
徐昊 沈江明



摘 要:聚焦爬虫(Focused Crawler)又称为主题爬虫,是从网络上获取特定主题数据的有效工具。为了避免传统聚焦爬虫预训练主题相关性分类器的繁复工作,提出一种自举聚焦爬虫(Bootstrapping Focused Crawler),用于从特定网站群中收集主题数据。自举聚焦爬虫省略了预先训练分类器的步骤,转而采用一些样本页面以相似度排序的方式替代分类器功能。在实验中,自举聚焦爬虫以牺牲一定准确率为代价,取得了0.62的召回率以及0.45的F1值,表现优于传统聚焦爬虫(召回率0.16、F1值0.25)。对于网站群主题数据采集任务,采用相似度排序替代主题分类器,不仅可以减轻分类器训练负担,还可以达到更好的效果。
关键词:爬虫技术;信息检索;自举聚焦爬虫
DOI:10. 11907/rjdk. 201564 开放科学(资源服务)标识码(OSID):
中图分类号:TP393文献标识码:A 文章编号:1672-7800(2020)008-0109-04
Abstract: Focused crawler (also known as theme crawler) is an effective tool to get data in any specific domain from Web. However, conventional focused crawlers need a classifier to filter out the irrelevant webpages, and to get such a classifier is usually labor-intensive. In this paper, we propose a Bootstrapping Focused Crawler (BFC) for collecting information from a group of websites in the same category. Instead of pre-training a tailored classifier, BFC adopts a ranking module to do the classification. In the experiments, the recall and F1-score of BFC is significantly better than conventional focused crawler, from which we could draw the conclusion that our approach is more effective for the crawling tasks within a group of similar websites.
Key Words: Web crawler; information retrieval; bootstrapping focused crawler
0 引言
从Web上收集特定主题数据的技术可分为两类:①基于搜索的发现技术[1-3],主要依靠搜索引擎查找网页;②基于爬行的发现技术[4-6],主要利用Web链接结构从已下载的网页中提取新链接,从而发现更多潜在的目标网页。前者适用于存在一些关键字可区分主题数据和其它数据的情况,后者灵活性更强,代表技术就是聚焦爬虫。
与普通爬虫相比,聚焦爬虫有明确的目标指向性,在爬取网页过程中能够丢弃不相关页面,并始终跟踪可能导向“相关”页面的超链接,因而能更有效地收集特定主题的数据。聚焦爬虫框架与一般爬虫基本相同,也即是说,它从几个种子链接(Seed URL)开始,下载相关页面并提取其中包含的超链接,然后跟踪这些超链接以获取更多页面。……
