Please wait a minute...
Journal of ZheJiang University (Engineering Science)  2026, Vol. 60 Issue (10): 2247-2258    DOI: 10.3785/j.issn.1008-973X.2026.10.017
    
Stochastic variance-reduced recursive momentum policy gradient method
Yaohui XIN(),Yongqiang LI*(),Yu FENG,Lei HU
School of Information Engineering, Zhejiang University of Technology, Hangzhou 310000, China
Download: HTML     PDF(1237KB) HTML
Export: BibTeX | EndNote (RIS)      

Abstract  

In reinforcement learning, policy gradient methods are often challenged by excessive variance in gradient estimation, which leads to poor convergence. A stochastic variance-reduced recursive momentum policy gradient method (SVRRM-PG) was proposed. This method combines importance sampling technique with a reference-point switching mechanism to achieve large-batch gradient estimation without resampling, effectively reducing variance and improving training stability, while exhibiting provably advanced sample complexity currently. To improve the general applicability, the SVRRM-PG framework was integrated with the proximal policy optimization (PPO) algorithm, which led to the proposed SVR-PPO method for verifying the efficacy of SVRRM-PG estimator in enhancing the performance of the established reinforcement learning algorithms. A lot of benchmark discrete and continuous control tasks were conducted to compare the performance. The results demonstrate that the SVRRM-PG estimator not only accelerates the convergence significantly and improves the training stability, while achieves superior overall performance in multiple benchmark evaluations, confirming its effectiveness and robustness in complex reinforcement learning environments.



Key wordsreinforcement learning      policy gradient      stochastic variance      importance sampling      sample complexity     
Received: 06 June 2025      Published: 28 July 2026
CLC:  TP 18  
Fund:  国家自然科学基金资助项目(U2341216).
Corresponding Authors: Yongqiang LI     E-mail: a13353745863@163.com;yqli@zjut.edu.cn
Cite this article:

Yaohui XIN,Yongqiang LI,Yu FENG,Lei HU. Stochastic variance-reduced recursive momentum policy gradient method. Journal of ZheJiang University (Engineering Science), 2026, 60(10): 2247-2258.

URL:

https://www.zjujournals.com/eng/10.3785/j.issn.1008-973X.2026.10.017     OR     https://www.zjujournals.com/eng/Y2026/V60/I10/2247


随机方差缩减的递归动量策略梯度方法

强化学习中的策略梯度方法在梯度估计过程中通常存在方差过大的问题, 导致收敛性较差,为此提出随机方差缩减的递归动量策略梯度方法(SVRRM-PG). 通过结合重要性采样技术与参考点切换机制, 不需要使用重采样即可实现大批量梯度估计, 有效减小方差并提升训练稳定性,拥有可证明的、当前最优的样本复杂度. 为提升算法通用性, 将SVRRM-PG与近端策略优化(PPO)算法进行深度结合, 提出SVR-PPO算法, 以验证SVRRM-PG估计器在增强主流强化学习算法性能方面的有效性. 在多个经典离散与连续控制任务中进行性能对比,结果表明,SVRRM-PG估计器不仅能够有效提升收敛速度和稳定性,而且在多项基准测试中表现出优于现有方法的综合性能,验证了其在复杂强化学习场景下的有效性.


关键词: 强化学习,  策略梯度,  随机方差,  重要性采样,  样本复杂度 
方法VR技术样本复杂度小样本复杂度
GPOMDP[7]$ \mathcal{O}({\epsilon }^{-4}) $
SVRPG[18]SVRG[13]$ \mathcal{O}({\epsilon }^{-10/3}) $$ \mathcal{O}({\epsilon }^{-4/3}) $
STORM-PG[20]SARAH[14]$ \mathcal{O}({\epsilon }^{-3}) $$ \mathcal{O}(1) $
PAGE-PG[21]PAGE[17]$ \mathcal{O}({\epsilon }^{-3}) $$ \mathcal{O}(1) $
SVRRM-PGSVRRM[24]$ \mathcal{O}({\epsilon }^{-3}) $$ \mathcal{O}(1) $
Tab.1 Examples of complexity of partial algorithms for finding $ \epsilon - $stable solutions
任务描述动作空间状态空间最大步长
Cartpole平衡车2维离散4维500
Acrobot两关节连杆3维离散6维500
Hopper跳跃机器人3维连续11维1000
Walker行走机器人6维连续17维1000
Tab.2 Key parameters of the experimental task scenarios
算法$ \gamma $$ \eta $$ \alpha $$ p_{t} $$ {p}_{0} $$ S_{1} $$ B $$ Q $
GPOMDP0.990.01
STORM-PG0.990.010.9205
PAGE-PG0.990.010.3205
SHARP0.990.010.9205
SVRRM-PG0.990.010.90.3205100
Tab.3 Benchmark experimental parameter settings
Fig.1 Reward curves of five algorithms on discrete tasks
Fig.2 Reward curves of 5 algorithms on continuous tasks
算法AcrobotCartpoleHopperWalker
0~30%30%~100%0~30%30%~100%0~30%30%~100%0~30%30%~100%
GPOMDP6383.36762.243726.2583.8073250.8637910.807108.043110.71
STORM-PG5577.74296.932025.4467.3580832.2520085.554739.302500.16
PAGE-PG5178.15247.353472.70104.6057135.1325513.845098.692666.30
SHARP4634.38308.802542.07152.4145235.574477.664100.442607.13
SVRRM-PG3531.65163.941817.9953.9038120.922744.553205.001836.84
Tab.4 Variance of five experimental algorithms at different stages
$ Q $-$ B $R
0$ \sim $30003000$ \sim $10000
20-5155.87190.79
50-5155.87191.28
50-10137.84191.63
100-5157.91193.57
100-10139.50192.69
100-20108.91193.71
150-5157.01189.44
150-10140.52191.80
150-20110.50191.77
Tab.5 Average returns under different batches
Fig.3 Average returns of SVRRM-PG on Cartpole under different batches
算法t
AcrobotCartpoleHopperWalker
GPOMDP3′06″13′10″33′36″36′05″
STORM-PG2′16″14′22″38′50″38′10″
PAGE-PG2′21″13′12″35′18″35′10″
SHARP2′19″14′00″38′16″38′52″
SVRRM-PG2′26″15′40″41′24″42′37″
Tab.6 Running time of five algorithms on different tasks
Fig.4 Average returns of SVR-PPO and PPO on Walker tasks
[1]   SHALEV-SHWARTZ S, SHAMMAH S, SHASHUA A. Safe, multi-agent, reinforcement learning for autonomous driving [EB/OL]. (2016-10-11). https://arxiv.org/pdf/1610.03295.
[2]   DEISENROTH M P, NEUMANN G, PETERS J, et al A survey on policy search for robotics[J]. Foundations and Trends in Robotics, 2013, 2 (1/2): 1- 142
doi: 10.1561/9781601987037
[3]   SILVER D, SCHRITTWIESER J, SIMONYAN K, et al Mastering the game of go without human knowledge[J]. Nature, 2017, 550 (7676): 354- 359
doi: 10.1038/nature24270
[4]   WANG W Y, LI J W, HE X D. Deep reinforcement learning for NLP [C]// Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics: Tutorial Abstracts. Melbourne: ACL, 2018: 19–21.
[5]   SUTTON R S, MCALLESTER D, SINGH S, et al. Policy gradient methods for reinforcement learning with function approximation [C]// Advances in Neural Information Processing Systems. Vancouver: MIT Press, 2000, 12: 1057–1063
[6]   WILLIAMS R J Simple statistical gradient-following algorithms for connectionist reinforcement learning[J]. Machine Learning, 1992, 8: 229- 256
doi: 10.1023/A:1022672621406
[7]   BAXTER J, BARTLETT P L Infinite-horizon policy-gradient estimation[J]. Journal of Artificial Intelligence Research, 2001, 15: 319- 350
doi: 10.1613/jair.806
[8]   SCHULMAN J, MORITZ P, LEVINE S, et al. High-dimensional continuous control using generalized advantage estimation [EB/OL]. [2015-06-18]. https://arxiv.org/pdf/1506.02438.
[9]   SCHULMAN J, WOLSKI F, DHARIWAL P, et al. Proximal policy optimization algorithms [EB/OL]. (2017−07−20)[2017−08−28]. https://arxiv.org/pdf/1707.06347.
[10]   郭振华, 闫瑞栋, 邱志勇, 等 基于随机采样的方差缩减优化算法[J]. 计算机科学与探索, 2025, 19 (3): 667- 681
GUO Zhenhua, YAN Ruidong, QIU Zhiyong, et al Variance reduction optimization algorithm based on random sampling[J]. Journal of Frontiers of Computer Science and Technology, 2025, 19 (3): 667- 681
[11]   HUANG Feihu, GAO Shangqian, HUANG Heng. Bregman gradient policy optimization [EB/OL]. (2021−06−23)[2021−03−16]. https://arxiv.org/pdf/2106.12112.
[12]   王爽 基于随机镜像下降对称交替方向乘子法的非凸优化问题研究[J]. 应用数学进展, 2025, 14 (3): 176- 191
WANG Shuang Research on non-convex optimization problems based on stochastic mirror descent symmetric alternating direction multiplier method[J]. Advances in Applied Mathematics, 2025, 14 (3): 176- 191
doi: 10.12677/aam.2025.143104
[13]   JOHNSON R, ZHANG Tong. Accelerating stochastic gradient descent using predictive variance reduction [C]// Advances in Neural Information Processing Systems. Lake Tahoe: Curran Associates, 2013, 26: 315-323
[14]   NGUYEN L M, LIU J, SCHEINBERG K, et al. SARAH: a novel method for machine learning problems using stochastic recursive gradient [C]// International Conference on Machine Learning. Sydney: PMLR, 2017: 2613–2621.
[15]   FANG C, LI C J, LIN Z C, et al. Spider: near-optimal non-convex optimization via stochastic path-integrated differential estimator [C]// Advances in Neural Information Processing Systems. Montreal: Curran Associates, 2018, 31: 689-699
[16]   CUTKOSKY A, ORABONA F. Momentum-based variance reduction in non-convex SGD [C]// Advances in Neural Information Processing Systems. Vancouver: Curran Associates, 2019, 32: 15210–15219
[17]   LI Z Z, BAO H Y, ZHANG X L, et al. PAGE: a simple and optimal probabilistic gradient estimator for nonconvex optimization [C]// International Conference on Machine Learning. Virtual Event: PMLR, 2021: 6286–6295.
[18]   PAPINI M, BINAGHI D, CANONACO G, et al. Stochastic variance-reduced policy gradient [C]// International Conference on Machine Learning. Stockholm: PMLR, 2018: 4026–4035.
[19]   XU P, GAO F, GU Q Q. Sample efficient policy gradient methods with recursive variance reduction [EB/OL]. (2019–09–18) [2021–08–01]. https://arxiv.org/pdf/1909.08610.
[20]   YUAN H Z, LIAN X R, LIU J, et al. Stochastic recursive momentum for policy gradient methods [EB/OL]. (2020–03–09). https://arxiv.org/pdf/2003.04302.
[21]   GARGIANI M, ZANELLI A, MARTINELLI A, et al. PAGE-PG: a simple and loopless variance-reduced policy gradient method with probabilistic gradient estimation [C]// International Conference on Machine Learning. Baltimore: PMLR, 2022: 7223–7240.
[22]   SALEHKALEYBAR S, KHORASANI S, KIYAVASH N, et al. Momentum-based policy gradient with second-order information [EB/OL]. (2022–05–17) [2023–09–26]. https://arxiv.org/pdf/2205.08253.
[23]   胡磊, 李永强, 冯宇, 等 海森辅助的概率策略梯度方法[J]. 模式识别与人工智能, 2025, 38 (2): 177- 191
HU Lei, LI Yongqiang, FENG Yu, et al Hessian-aided probability strategy gradient method[J]. Pattern Recognition and Artificial Intelligence, 2025, 38 (2): 177- 191
doi: 10.16451/j.cnki.issn1003-6059.202502006
[24]   LIAO S C, LIU Y, HAN C Y, et al Momentum-based variance-reduced stochastic Bregman proximal gradient methods for nonconvex nonsmooth optimization[J]. Expert Systems with Applications, 2025, 266: 125960
doi: 10.1016/j.eswa.2024.125960
[25]   KOVALEV D, HORVÁTH S, RICHTÁRIK P. Don’t jump through hoops and remove those loops: SVRG and Katyusha are better without the outer loop [C]// Algorithmic Learning Theory. San Diego: PMLR, 2020: 451–467.
[26]   FURMSTON T, BARBER D. A unifying perspective of parametric policy search methods for Markov decision processes [C]// Advances in Neural Information Processing Systems. Lake Tahoe: Curran Associates, 2012, 25: 2726−2734
[1] Yuanfeng LIAN,Shuyu FAN,Benzhe ZHANG,Sen WANG. Bio-inspired gait control integrating geometric constraint feature extraction and multi-reward cooperative optimization[J]. Journal of ZheJiang University (Engineering Science), 2026, 60(9): 1862-1871.
[2] Li SONG,Yuanyu WAN,Mingli SONG. Graph convolutional reinforcement learning via hybrid dilated attention mechanism[J]. Journal of ZheJiang University (Engineering Science), 2026, 60(9): 1980-1990.
[3] Haijun XING,Yiqin SHI,Qiwei WANG,Xiao YANG,Shijie ZHUANG. Review on distribution system restoration based on multiple flexible resources[J]. Journal of ZheJiang University (Engineering Science), 2026, 60(9): 2059-2076.
[4] Ningbo CAO,Qichao WAN,Liying ZHAO,Zimeng LI,Baolin HUANG. Collaborative control of mixed traffic intersections integrating multi-agent reinforcement learning and maximum pressure control[J]. Journal of ZheJiang University (Engineering Science), 2026, 60(8): 1819-1831.
[5] Lang CHEN,Zengli LIU,Xuanzhi ZHAO. Decision-making for multi-USV adversarial encirclement based on adaptive curriculum reinforcement learning[J]. Journal of ZheJiang University (Engineering Science), 2026, 60(7): 1369-1380.
[6] Yong CHEN,Xizhi DU,Yiwei JIANG,Wenchao YI,Zhi PEI,Zuzhen JI. Deep reinforcement learning models and algorithms for single-machine scheduling considering deteriorated maintenance[J]. Journal of ZheJiang University (Engineering Science), 2026, 60(7): 1528-1538.
[7] Qiao LUO,Jun CHEN,Jie GENG,Chaoyang TANG,Chunyun FU. Review of energy management strategies for hybrid electric vehicles based on driving information[J]. Journal of ZheJiang University (Engineering Science), 2026, 60(7): 1539-1556.
[8] Yiming SHANG,Changping DU,Rui YANG,Tianrui FANG,Ze’an DU,Yao ZHENG. Safe hierarchical reinforcement learning framework for dynamic UAV navigation[J]. Journal of ZheJiang University (Engineering Science), 2026, 60(6): 1240-1250.
[9] Yiwei ZHANG,Xin CUI,Qinghui ZHAO,Yan CHEN. Collaborative content caching optimization in UAV-assisted internet of vehicle based on NOMA[J]. Journal of ZheJiang University (Engineering Science), 2026, 60(6): 1289-1298.
[10] Yang WANG,Hongchao LIU,Chi TIAN,Bing WU,Di ZHANG. Multi-ship collision avoidance via route exchange mechanism: strategy learning and game-theoretic decision making[J]. Journal of ZheJiang University (Engineering Science), 2026, 60(5): 964-976.
[11] Qingqing YANG,Runpeng TANG,Yi PENG. Joint waveform and phase shift design in integrated sensing and communication systems[J]. Journal of ZheJiang University (Engineering Science), 2026, 60(4): 906-914.
[12] Xingxing LI,Fan YANG,Jie HUANG,Xianzhi LAI,Fenghang YAO,Jieliang CAI,Ni ZHANG. Federated reinforcement learning for resource management in vehicular networks with differentiated interferences[J]. Journal of ZheJiang University (Engineering Science), 2026, 60(10): 2207-2214.
[13] Hongwei GAO,Bingxu SHANG,Xinkang ZHANG,Hongfeng WANG,Wei HE,Xiaofei PEI. Decision-making and planning of intelligent vehicle based on reachable set and reinforcement learning[J]. Journal of ZheJiang University (Engineering Science), 2025, 59(9): 1996-2004.
[14] Jiale LIU,Yali XUE,Shan CUI,Jun HONG. TD3 mapless navigation algorithm guided by dynamic window approach[J]. Journal of ZheJiang University (Engineering Science), 2025, 59(8): 1671-1679.
[15] Kun HAO,Xuan MENG,Xiaofang ZHAO,Zhisheng LI. 3D underwater AUV path planning method integrating adaptive potential field method and deep reinforcement learning[J]. Journal of ZheJiang University (Engineering Science), 2025, 59(7): 1451-1461.