随机方差缩减的递归动量策略梯度方法
辛垚辉,李永强,冯宇,胡磊
Stochastic variance-reduced recursive momentum policy gradient method
Yaohui XIN,Yongqiang LI,Yu FENG,Lei HU
表 6
5种算法在不同任务上的运行时间
Tab.6
Running time of five algorithms on different tasks
算法
t
Acrobot
Cartpole
Hopper
Walker
GPOMDP
3′06″
13′10″
33′36″
36′05″
STORM-PG
2′16″
14′22″
38′50″
38′10″
PAGE-PG
2′21″
13′12″
35′18″
35′10″
SHARP
2′19″
14′00″
38′16″
38′52″
SVRRM-PG
2′26″
15′40″
41′24″
42′37″