基于Python的网上购物数据爬取
2021-02-28高雅婷刘雅举
高雅婷 刘雅举















摘 要:随着网上购物的盛行,淘宝、京东、拼多多等互联网商业巨头也展开了激烈的竞争。收集商品、评论及销量数据以及对各种商品及用户的消费场景进行分析成了必不可少的环节。然而传统的人工收集并整理数据显然效率不足以满足当下各大公司以及其他相关产业对这些数据的需要。近年来Python爬虫技术的逐渐成熟,给网购数据收集并整理带来了极大的便利。
关键词:网购;Python;pip;爬虫技术
中图分类号:TP311.1 文献标识码:A文章编号:2096-4706(2021)16-0026-06
Online Shopping Data Crawling Based on Python
GAO Yating, LIU Yaju
(Hebei Agricultural University, Cangzhou 061000, China)
Abstract: With the popularity of online shopping, Taobao, Jingdong, Pinduoduo and other internet business giants also launches a fierce competition. Collecting product, review, and sales data, as well as analyzing the consumption scenarios of various products and users, has become an essential links. However, the traditional manual collection and sorting of data is obviously not efficient enough to meet the needs of companies and other related industries. In recent years, the gradual maturity of Python crawler technology has brought great convenience to the collection and sorting of online shopping data.
Keywords: online shopping; Python; pip; crawler technology
0 引 言
我们会发现我们常用的购物软件如:淘宝、京东等总是能给我们推荐符合我们兴趣的商品。这就是因为他们收集了消费者的POI(兴趣点)[1],通过调用数据库来分析到每个消费者的消费偏好,进而给消费之提供感兴趣的商品。各大网购公司,以及相关产业都需要消费者的消费数据如某商品的销量以及评论信息、各种商品及用户的消费场景等数据来进行数据分析。本文将以基于Python的網上购物数据-淘宝商品销量数据爬取为例进行分析。
1 Pyhton爬虫技术概述
爬虫的概念就是一段自动抓取互联网信息的程序,从互联网上抓取对于我们有价值的信息[2],其根本原理是递归算法。爬虫技术原本是应用于搜索引擎的,随着程序员前辈们的改进,目前已经成为一项非常通用且实用的数据抓取技术了。能开发爬虫技术的语言有很多种,但是由于Python语言的简单高效,人们更习惯用Python来开发爬虫技术,所以Python爬虫技术已经成为当下主流的爬虫技术。……
