基于BeautifulSoup+requests和selenium爬虫网页自动化处理的实现和性能对比
2021-02-28李晨昊





摘 要:网络爬虫是一种按照一定的规则,自动地抓取网页信息的程序或者脚本,因此编写特定的网络爬虫可以用来对网页进行自动化处理,从而达到提升工作效率的目的。文章针对同一个任务清单系统,分别使用BeautifulSoup + requests和selenium两种不同的爬虫方法实现了网页自动化处理功能。并且通过对两种方法的实现原理和运行结果进行分析,对两种爬虫方法进行对比。
关键词:爬虫;网页自动化;BeautifulSoup+requests;selenium
中图分类号:TP391 文献标识码:A文章编号:2096-4706(2021)16-0010-04
Implementation and Performance Comparison of Crawler Web Page Automatic Processing Based on BeautifulSoup + requests and selenium
LI Chenhao
(Wuhan Branch of China Mobile Hubei Co., Ltd., Wuhan 430000, China)
Abstract: Web crawler is a program or script that automatically grabs web page information according to certain rules. Therefore, a specific web crawler can be written to process web pages automatically, which provides efficiency improvement. The paper uses two different crawler methods: BeautifulSoup + requests and selenium to implement webpage automatic processing function for the same task list system. By analyzing the implementation principle and operation results of the two methods, the two crawler methods are compared.
Keywords: crawler; webpage automation; BeautifulSoup+requests; selenium
0 引 言
网络爬虫是一种按照一定的规则,自动地抓取网页信息的程序或者脚本。它的基本工作方式是模拟人工的操作去访问网站,并且在网站查找数据或者发送数据。因此爬虫不但能够用来快速获取网页信息,而且能够对网页进行自动化处理。
在实际工作中,碰到了一类任务清单系统,需要对系统里的任务进行处理。处理这些任务的操作大同小异,平均处理一条任务大约需要30秒到1分钟。如果全程依靠人工来完成,不但耗费时间长,而且还存在着人为误差。因此,为了加快任务处理速度、提高任务处理准确率、提升工作效率,编写了python爬虫脚本来对这些任务进行自动化处理。
最初版本的爬虫是基于BeautifulSoup + requests库的方法设计实现的,并且达到了预期的效果。后期由于新系统使用了远程访问模式,网址发生了变更,使得无法用原有方法进行自动化处理。因此重新編写了一个新的基于selenium实现的爬虫。两种爬虫方式虽然最终的都能达到相同的运行效果,但是实现和性能上存在着不小的差异。……
