随机方差缩减的递归动量策略梯度方法
辛垚辉,李永强,冯宇,胡磊
Stochastic variance-reduced recursive momentum policy gradient method
Yaohui XIN,Yongqiang LI,Yu FENG,Lei HU
表 5
不同批量下的平均回报
Tab.5
Average returns under different batches
$ Q $
-
$ B $
R
0
$ \sim $
3000
3000
$ \sim $
10000
20-5
155.87
190.79
50-5
155.87
191.28
50-10
137.84
191.63
100-5
157.91
193.57
100-10
139.50
192.69
100-20
108.91
193.71
150-5
157.01
189.44
150-10
140.52
191.80
150-20
110.50
191.77