随机方差缩减的递归动量策略梯度方法
|
|
辛垚辉,李永强,冯宇,胡磊
|
Stochastic variance-reduced recursive momentum policy gradient method
|
|
Yaohui XIN,Yongqiang LI,Yu FENG,Lei HU
|
|
| 表 1 部分算法对寻找$ \epsilon $-稳定解的复杂度实例 |
| Tab.1 Examples of complexity of partial algorithms for finding $ \epsilon - $stable solutions |
|
| 方法 | VR技术 | 样本复杂度 | 小样本复杂度 | | GPOMDP[7] | — | $ \mathcal{O}({\epsilon }^{-4}) $ | — | | SVRPG[18] | SVRG[13] | $ \mathcal{O}({\epsilon }^{-10/3}) $ | $ \mathcal{O}({\epsilon }^{-4/3}) $ | | STORM-PG[20] | SARAH[14] | $ \mathcal{O}({\epsilon }^{-3}) $ | $ \mathcal{O}(1) $ | | PAGE-PG[21] | PAGE[17] | $ \mathcal{O}({\epsilon }^{-3}) $ | $ \mathcal{O}(1) $ | | SVRRM-PG | SVRRM[24] | $ \mathcal{O}({\epsilon }^{-3}) $ | $ \mathcal{O}(1) $ |
|
|
|