|
|
|
| Stochastic variance-reduced recursive momentum policy gradient method |
Yaohui XIN( ),Yongqiang LI*( ),Yu FENG,Lei HU |
| School of Information Engineering, Zhejiang University of Technology, Hangzhou 310000, China |
|
|
|
Abstract In reinforcement learning, policy gradient methods are often challenged by excessive variance in gradient estimation, which leads to poor convergence. A stochastic variance-reduced recursive momentum policy gradient method (SVRRM-PG) was proposed. This method combines importance sampling technique with a reference-point switching mechanism to achieve large-batch gradient estimation without resampling, effectively reducing variance and improving training stability, while exhibiting provably advanced sample complexity currently. To improve the general applicability, the SVRRM-PG framework was integrated with the proximal policy optimization (PPO) algorithm, which led to the proposed SVR-PPO method for verifying the efficacy of SVRRM-PG estimator in enhancing the performance of the established reinforcement learning algorithms. A lot of benchmark discrete and continuous control tasks were conducted to compare the performance. The results demonstrate that the SVRRM-PG estimator not only accelerates the convergence significantly and improves the training stability, while achieves superior overall performance in multiple benchmark evaluations, confirming its effectiveness and robustness in complex reinforcement learning environments.
|
|
Received: 06 June 2025
Published: 28 July 2026
|
|
|
| Fund: 国家自然科学基金资助项目(U2341216). |
|
Corresponding Authors:
Yongqiang LI
E-mail: a13353745863@163.com;yqli@zjut.edu.cn
|
随机方差缩减的递归动量策略梯度方法
强化学习中的策略梯度方法在梯度估计过程中通常存在方差过大的问题, 导致收敛性较差,为此提出随机方差缩减的递归动量策略梯度方法(SVRRM-PG). 通过结合重要性采样技术与参考点切换机制, 不需要使用重采样即可实现大批量梯度估计, 有效减小方差并提升训练稳定性,拥有可证明的、当前最优的样本复杂度. 为提升算法通用性, 将SVRRM-PG与近端策略优化(PPO)算法进行深度结合, 提出SVR-PPO算法, 以验证SVRRM-PG估计器在增强主流强化学习算法性能方面的有效性. 在多个经典离散与连续控制任务中进行性能对比,结果表明,SVRRM-PG估计器不仅能够有效提升收敛速度和稳定性,而且在多项基准测试中表现出优于现有方法的综合性能,验证了其在复杂强化学习场景下的有效性.
关键词:
强化学习,
策略梯度,
随机方差,
重要性采样,
样本复杂度
|
|
| [1] |
SHALEV-SHWARTZ S, SHAMMAH S, SHASHUA A. Safe, multi-agent, reinforcement learning for autonomous driving [EB/OL]. (2016-10-11). https://arxiv.org/pdf/1610.03295.
|
|
|
| [2] |
DEISENROTH M P, NEUMANN G, PETERS J, et al A survey on policy search for robotics[J]. Foundations and Trends in Robotics, 2013, 2 (1/2): 1- 142
doi: 10.1561/9781601987037
|
|
|
| [3] |
SILVER D, SCHRITTWIESER J, SIMONYAN K, et al Mastering the game of go without human knowledge[J]. Nature, 2017, 550 (7676): 354- 359
doi: 10.1038/nature24270
|
|
|
| [4] |
WANG W Y, LI J W, HE X D. Deep reinforcement learning for NLP [C]// Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics: Tutorial Abstracts. Melbourne: ACL, 2018: 19–21.
|
|
|
| [5] |
SUTTON R S, MCALLESTER D, SINGH S, et al. Policy gradient methods for reinforcement learning with function approximation [C]// Advances in Neural Information Processing Systems. Vancouver: MIT Press, 2000, 12: 1057–1063
|
|
|
| [6] |
WILLIAMS R J Simple statistical gradient-following algorithms for connectionist reinforcement learning[J]. Machine Learning, 1992, 8: 229- 256
doi: 10.1023/A:1022672621406
|
|
|
| [7] |
BAXTER J, BARTLETT P L Infinite-horizon policy-gradient estimation[J]. Journal of Artificial Intelligence Research, 2001, 15: 319- 350
doi: 10.1613/jair.806
|
|
|
| [8] |
SCHULMAN J, MORITZ P, LEVINE S, et al. High-dimensional continuous control using generalized advantage estimation [EB/OL]. [2015-06-18]. https://arxiv.org/pdf/1506.02438.
|
|
|
| [9] |
SCHULMAN J, WOLSKI F, DHARIWAL P, et al. Proximal policy optimization algorithms [EB/OL]. (2017−07−20)[2017−08−28]. https://arxiv.org/pdf/1707.06347.
|
|
|
| [10] |
郭振华, 闫瑞栋, 邱志勇, 等 基于随机采样的方差缩减优化算法[J]. 计算机科学与探索, 2025, 19 (3): 667- 681 GUO Zhenhua, YAN Ruidong, QIU Zhiyong, et al Variance reduction optimization algorithm based on random sampling[J]. Journal of Frontiers of Computer Science and Technology, 2025, 19 (3): 667- 681
|
|
|
| [11] |
HUANG Feihu, GAO Shangqian, HUANG Heng. Bregman gradient policy optimization [EB/OL]. (2021−06−23)[2021−03−16]. https://arxiv.org/pdf/2106.12112.
|
|
|
| [12] |
王爽 基于随机镜像下降对称交替方向乘子法的非凸优化问题研究[J]. 应用数学进展, 2025, 14 (3): 176- 191 WANG Shuang Research on non-convex optimization problems based on stochastic mirror descent symmetric alternating direction multiplier method[J]. Advances in Applied Mathematics, 2025, 14 (3): 176- 191
doi: 10.12677/aam.2025.143104
|
|
|
| [13] |
JOHNSON R, ZHANG Tong. Accelerating stochastic gradient descent using predictive variance reduction [C]// Advances in Neural Information Processing Systems. Lake Tahoe: Curran Associates, 2013, 26: 315-323
|
|
|
| [14] |
NGUYEN L M, LIU J, SCHEINBERG K, et al. SARAH: a novel method for machine learning problems using stochastic recursive gradient [C]// International Conference on Machine Learning. Sydney: PMLR, 2017: 2613–2621.
|
|
|
| [15] |
FANG C, LI C J, LIN Z C, et al. Spider: near-optimal non-convex optimization via stochastic path-integrated differential estimator [C]// Advances in Neural Information Processing Systems. Montreal: Curran Associates, 2018, 31: 689-699
|
|
|
| [16] |
CUTKOSKY A, ORABONA F. Momentum-based variance reduction in non-convex SGD [C]// Advances in Neural Information Processing Systems. Vancouver: Curran Associates, 2019, 32: 15210–15219
|
|
|
| [17] |
LI Z Z, BAO H Y, ZHANG X L, et al. PAGE: a simple and optimal probabilistic gradient estimator for nonconvex optimization [C]// International Conference on Machine Learning. Virtual Event: PMLR, 2021: 6286–6295.
|
|
|
| [18] |
PAPINI M, BINAGHI D, CANONACO G, et al. Stochastic variance-reduced policy gradient [C]// International Conference on Machine Learning. Stockholm: PMLR, 2018: 4026–4035.
|
|
|
| [19] |
XU P, GAO F, GU Q Q. Sample efficient policy gradient methods with recursive variance reduction [EB/OL]. (2019–09–18) [2021–08–01]. https://arxiv.org/pdf/1909.08610.
|
|
|
| [20] |
YUAN H Z, LIAN X R, LIU J, et al. Stochastic recursive momentum for policy gradient methods [EB/OL]. (2020–03–09). https://arxiv.org/pdf/2003.04302.
|
|
|
| [21] |
GARGIANI M, ZANELLI A, MARTINELLI A, et al. PAGE-PG: a simple and loopless variance-reduced policy gradient method with probabilistic gradient estimation [C]// International Conference on Machine Learning. Baltimore: PMLR, 2022: 7223–7240.
|
|
|
| [22] |
SALEHKALEYBAR S, KHORASANI S, KIYAVASH N, et al. Momentum-based policy gradient with second-order information [EB/OL]. (2022–05–17) [2023–09–26]. https://arxiv.org/pdf/2205.08253.
|
|
|
| [23] |
胡磊, 李永强, 冯宇, 等 海森辅助的概率策略梯度方法[J]. 模式识别与人工智能, 2025, 38 (2): 177- 191 HU Lei, LI Yongqiang, FENG Yu, et al Hessian-aided probability strategy gradient method[J]. Pattern Recognition and Artificial Intelligence, 2025, 38 (2): 177- 191
doi: 10.16451/j.cnki.issn1003-6059.202502006
|
|
|
| [24] |
LIAO S C, LIU Y, HAN C Y, et al Momentum-based variance-reduced stochastic Bregman proximal gradient methods for nonconvex nonsmooth optimization[J]. Expert Systems with Applications, 2025, 266: 125960
doi: 10.1016/j.eswa.2024.125960
|
|
|
| [25] |
KOVALEV D, HORVÁTH S, RICHTÁRIK P. Don’t jump through hoops and remove those loops: SVRG and Katyusha are better without the outer loop [C]// Algorithmic Learning Theory. San Diego: PMLR, 2020: 451–467.
|
|
|
| [26] |
FURMSTON T, BARBER D. A unifying perspective of parametric policy search methods for Markov decision processes [C]// Advances in Neural Information Processing Systems. Lake Tahoe: Curran Associates, 2012, 25: 2726−2734
|
|
|
|
Viewed |
|
|
|
Full text
|
|
|
|
|
Abstract
|
|
|
|
|
Cited |
|
|
|
|
| |
Shared |
|
|
|
|
| |
Discussed |
|
|
|
|