APP下载

基于深度强化学习的动态共享单车重置问题研究

2021-06-15张建同何钰林

上海管理科学 2021年2期

张建同 何钰林

摘 要:共享单车在为城市出行带来便利的同时,也面临着资源分布不平衡问题。针对单车分布动态变化环境下的共享单车重置问题,提出基于强化学习的实时调度策略结构。构建了面向强化学习的共享单车重置问题模型,利用深度确定性策略梯度算法(DDPG)進行求解,以获得实时调度策略。基于实际单车分布数据,构建了调度过程中的环境交互模拟器。最后,利用强化学习在模拟器中进行大规模数据实验,结果表明算法得到的调度策略能提高系统表现,并且效果好于已有方法。

关键词:共享单车重置问题;深度强化学习;摩拜单车

中图分类号:F 57

文献标志码:A

文章编号:1005-9679(2021)02-0081-06

Abstract:While bikes sharing bring convenience to urban travel, they also face the problem of unbalanced distribution of shared bike resources. A real-time scheduling strategy structure based on reinforcement learning was proposed to solve the repositioning problem of shared bikes under dynamic change of bicycle distribution. In this paper, a model of the bike repositioning problem for reinforcement learning is built, which is solved by deep deterministic strategy gradient (DDPG) to obtain real-time scheduling strategy. Based on the actual distribution data of shared bikes, an environmental interaction simulator is constructed for the scheduling process. A large-scale data experiment using reinforcement learning is carried out in the simulator. The experiment results show that the reposi tioning strategy obtained by the algorithm can significantly improve the performance of the system, and the algorithm performance is better than other existing methods.

Key words:bike repositioning problem; deep reinforcement learning; Mobike

共享单车作为一种便捷、环保的出行方式,近年来在国内大部分城市都已经普及,有效地解决了城市公共交通的“最后一公里”问题。但庞大的共享单车系统在运营管理上也面临诸多问题,其中一个主要问题就是共享单车在时空上分布不平衡,导致有些地方单车短缺,无法满足用户需求,而同时在某些地方单车数量过多,不仅浪费资源,同时给城市管理带来了许多麻烦。

针对共享单车分布不平衡现象,许多学者围绕共享单车重置问题(Bike Repositioning Problem,BRP)展开了研究。从调度主体的角度,共享单车重置问题可以分为基于用户的重置问题(User-Based BRP)和基于运营商的重置问题(Operator-Based BRP)。基于用户的重置通过引导用户用车和还车行为实现系统单车再平衡,一般通过动态定价或者对于在指定站点用车与还车行为给予奖励的方式实现。……

登录APP查看全文