Please wait a minute...
Journal of ZheJiang University (Engineering Science)  2026, Vol. 60 Issue (9): 1862-1871    DOI: 10.3785/j.issn.1008-973X.2026.09.003
    
Bio-inspired gait control integrating geometric constraint feature extraction and multi-reward cooperative optimization
Yuanfeng LIAN1,2(),Shuyu FAN1,Benzhe ZHANG1,Sen WANG1
1. College of Artificial Intelligence, China University of Petroleum, Beijing 102249, China
2. Beijing Key Laboratory of Petroleum Data Mining, China University of Petroleum, Beijing 102249, China
Download: HTML     PDF(4099KB) HTML
Export: BibTeX | EndNote (RIS)      

Abstract  

A new bio-inspired gait control method was proposed to address the challenges of insufficient locomotion stability and environmental adaptability of quadruped robots in complex terrains. A bio-inspired gait encoder module based on dynamic time warping (DTW) and convolutional processing was constructed to extract spatiotemporally correlated gait features through inverse kinematics solution and joint velocity differentiation calculation. A geometric constraint mechanism driven by differential kinematics was incorporated, and the coupled mapping between task-space trajectories and joint torques was embedded into the proximal policy optimization (PPO) network to enhance the joint motion coordination and the sampling efficiency. A hierarchical reward function combining energy efficiency factors and the Lyapunov stability was designed, the and multi-objective collaborative control of bio-inspired gaits was achieved via exponential energy efficiency rewards and stability rewards. The experimental results demonstrated that the proposed method outperformed the comparative algorithms in terms of indicators such as average collision value and adaptive loss function in the simulation environment, and effectively mitigated the joint ground contact phenomenon. This method significantly enhanced the motion stability and the environmental adaptability of the quadruped robot in the deployment test of the real robot in complex terrains.



Key wordsgait feature extraction      geometric constraint      hierarchical reward mechanism      deep reinforcement learning      proximal policy optimization     
Received: 14 August 2025      Published: 20 July 2026
CLC:  TP 393  
Cite this article:

Yuanfeng LIAN,Shuyu FAN,Benzhe ZHANG,Sen WANG. Bio-inspired gait control integrating geometric constraint feature extraction and multi-reward cooperative optimization. Journal of ZheJiang University (Engineering Science), 2026, 60(9): 1862-1871.

URL:

https://www.zjujournals.com/eng/10.3785/j.issn.1008-973X.2026.09.003     OR     https://www.zjujournals.com/eng/Y2026/V60/I9/1862


融合几何约束特征提取与多奖励协同优化的仿生步态控制

针对四足机器人在复杂地形中运动稳定性与环境适应性不足的问题,提出融合几何约束特征提取与多奖励协同优化的仿生步态控制方法. 构建基于动态时间规整(DTW)与卷积处理的仿生步态编码器模块,通过逆运动学解算、关节速度微分计算来提取时空关联的步态特征. 结合微分运动学驱动的几何约束机制,将任务空间轨迹与关节力矩的耦合映射嵌入近端策略优化(PPO)网络来提升关节运动协调性与采样效率.设计融合能效因子与李雅普诺夫稳定性的分层奖励函数,通过指数型能效奖励和稳定性奖励实现仿生步态的多目标协同控制. 实验结果表明,所提方法在仿真环境中的碰撞平均值、适应性损失等优于对比算法,且有效改善了关节触地现象. 在真机复杂地形部署测试中,该方法有效提升了四足机器人的运动稳定性与环境适应性.


关键词: 步态特征提取,  几何约束,  分层奖励机制,  深度强化学习,  近端策略优化 
Fig.1 Framework of bio-inspired gait control method
Fig.2 Schematic diagram of geometric constraint relationship
Fig.3 Application flow of PPO algorithm
Fig.4 Gait learning process of quadruped robot
Fig.5 Variation graph of multi-dimensional performance metrics of quadruped robot
Fig.6 Variation graph of key kinematic performance reward metrics of quadruped robot
Fig.7 Simulation comparison of locomotion performance of quadruped robot using different algorithms
Fig.8 Comparison of mean collision, adaptive loss, and mean total reward in simulation training of quadruped robot using different algorithms
Fig.9 Comparison graph of key kinematic performance reward metrics of quadruped robot using different algorithms
Fig.10 Comparison of locomotion performance of quadruped robot using different algorithms in sandy terrain scenario
Fig.11 Comparison of locomotion performance of quadruped robot using different algorithms in uncultivated terrain scenarios
算法名称$ {t_{\text{e}}} $/s$ {c_{\text{e}}} $
walk-these-ways-go26.117.51
SAC6.059.19
所提方法6.833.76
Tab.1 Comparison of computational efficiency and collision performance of different algorithms in quadruped robot simulation training
[1]   谭民, 王硕 机器人技术研究进展[J]. 自动化学报, 2013, 39 (7): 963- 972
TAN Min, WANG Shuo Research progress on robotics[J]. Acta Automatica Sinica, 2013, 39 (7): 963- 972
[2]   贾庆轩, 袁博楠, 陈钢, 等 关节锁定空间机械臂负载操作能力评估与轨迹规划[J]. 控制与决策, 2020, 35 (1): 243- 249
JIA Qingxuan, YUAN Bonan, CHEN Gang, et al Load carrying capacity evaluation and task trajectory planning of space manipulator with the locked joint[J]. Control and Decision, 2020, 35 (1): 243- 249
[3]   尤波, 曲伟健, 李佳钰 面向双操作者的六足机器人共享遥操作[J]. 控制与决策, 2022, 37 (11): 2769- 2778
YOU Bo, QU Weijian, LI Jiayu Shared teleoperation of hexapod robot for dual operators[J]. Control and Decision, 2022, 37 (11): 2769- 2778
[4]   郭非, 汪首坤, 王军政 轮足复合移动机器人运动规划发展现状及关键技术分析[J]. 控制与决策, 2022, 37 (6): 1433- 1444
GUO Fei, WANG Shoukun, WANG Junzheng Development status and key technology analysis for motion planning of wheel-legged hybrid mobile robot[J]. Control and Decision, 2022, 37 (6): 1433- 1444
[5]   GUO H, ZHU J, CHEN Y E-LOAM: LiDAR odometry and mapping with expanded local structural information[J]. IEEE Transactions on Intelligent Vehicles, 2023, 8 (2): 1911- 1921
doi: 10.1109/TIV.2022.3151665
[6]   ASADI F, KHORRAM M, MOOSAVIAN S A A. CPG-based gait transition of a quadruped robot [C]// Proceedings of the 3rd RSI International Conference on Robotics and Mechatronics. Tehran: IEEE, 2015: 210–215.
[7]   DING Y, PANDALA A, LI C, et al Representation-free model predictive control for dynamic motions in quadrupeds[J]. IEEE Transactions on Robotics, 2021, 37 (4): 1154- 1171
doi: 10.1109/TRO.2020.3046415
[8]   汪首坤, 刘大江, 郭非, 等 基于动力学模型预测控制的Stewart型轮足腿控制方法[J]. 北京理工大学学报, 2021, 41 (4): 418- 424
WANG Shoukun, LIU Dajiang, GUO Fei, et al Stewart type wheel-foot leg control method based on dynamics and model predictive control[J]. Transactions of Beijing Institute of Technology, 2021, 41 (4): 418- 424
[9]   朱晓庆, 陈江涛, 张思远, 等 基于深度仲裁策略的四足机器人步态学习[J]. 北京理工大学学报, 2023, 43 (11): 1197- 1204
ZHU Xiaoqing, CHEN Jiangtao, ZHANG Siyuan, et al Gait learning of quadruped robot based on deep arbitration strategy[J]. Transactions of Beijing Institute of Technology, 2023, 43 (11): 1197- 1204
[10]   张思远, 朱晓庆, 阮晓钢, 等 基于环境反馈机制的四足机器人运动技能学习[J]. 控制与决策, 2024, 39 (5): 1461- 1468
ZHANG Siyuan, ZHU Xiaoqing, RUAN Xiaogang, et al Motor skill learning of quadruped robot based on environmental feedback mechanism[J]. Control and Decision, 2024, 39 (5): 1461- 1468
[11]   BUTYREV L, EDELHÄUβER T, MUTSCHLER C. Deep reinforcement learning for motion planning of mobile robots [EB/OL]. (2019-12-19) [2025-04-04]. https://arxiv.org/abs/1912.09260.
[12]   XIONG J, WANG Q, YANG Z, et al. Parametrized deep Q-networks learning: reinforcement learning with discrete-continuous hybrid action space [EB/OL]. (2018-10-10) [2025-04-04]. https://arxiv.org/abs/1810.06394.
[13]   VAN HASSELT H, GUEZ A, SILVER D. Deep reinforcement learning with double Q-Learning [C]// Thirtieth AAAI Conference on Artificial Intelligence. Phoenix: AAAI Press, 2016: 2094–2100.
[14]   WANG Z, SCHAUL T, HESSEL M, et al. Dueling network architectures for deep reinforcement learning [C]// Proceedings of the 33rd International Conference on Machine Learning. New York: JMLR, 2016: 1995–2003.
[15]   MNIH V, BADIA A P, MIRZA M, et al. Asynchronous methods for deep reinforcement learning [C]// Proceedings of the 33rd International Conference on Machine Learning. New York: JMLR, 2016: 1928–1937.
[16]   LILLICRAP T P, HUNT J J, PRITZEL A, et al. Continuous control with deep reinforcement learning [EB/OL]. (2019-07-05) [2025-04-04]. https://arxiv.org/abs/1509.02971.
[17]   GANGAPURWALA S, GEISERT M, ORSOLINO R, et al RLOC: terrain-aware legged locomotion using reinforcement learning and optimal control[J]. IEEE Transactions on Robotics, 2022, 38 (5): 2908- 2927
doi: 10.1109/TRO.2022.3172469
[18]   董豪, 杨静, 李少波, 等 基于深度强化学习的机器人运动控制研究进展[J]. 控制与决策, 2022, 37 (2): 278- 292
DONG Hao, YANG Jing, LI Shaobo, et al Research progress of robot motion control based on deep reinforcement learning[J]. Control and Decision, 2022, 37 (2): 278- 292
[19]   HAARNOJA T, HA S, ZHOU A, et al. Learning to walk via deep reinforcement learning [C]// Robotics: Science and Systems XV. Freiburg im Breisgau: RSS Foundation, 2019.
[20]   FUCHIOKA Y. Imitating optimized trajectories for dynamic quadruped behaviors [D]. Vancouver: University of British Columbia, 2023.
[21]   ISCEN A, CALUWAERTS K, TAN J, et al. Policies modulating trajectory generators [C]// Proceedings of the 2nd Conference on Robot Learning. Zurich: PMLR, 2018: 916–926.
[22]   SCHULMAN J, WOLSKI F, DHARIWAL P, et al. Proximal policy optimization algorithms [EB/OL]. (2017-08-28) [2025-04-04]. https://arxiv.org/abs/1707.06347.
[23]   RUDIN N, HOELLER D, REIST P, et al. Learning to walk in minutes using massively parallel deep reinforcement learning [C]// Proceedings of the 5th Conference on Robot Learning. London: PMLR, 2021: 91–100.
[24]   MARGOLIS G B, AGRAWAL P. Walk these ways: tuning robot control for generalization with multiplicity of behavior [C]// Proceedings of the 6th Conference on Robot Learning. Auckland: PMLR, 2022: 22–31.
[25]   朱晓庆, 刘鑫源, 阮晓钢, 等 融合元学习和PPO算法的四足机器人运动技能学习方法[J]. 控制理论与应用, 2024, 41 (1): 155- 162
ZHU Xiaoqing, LIU Xinyuan, RUAN Xiaogang, et al A quadruped robot kinematic skill learning method integrating meta-learning and PPO algorithms[J]. Control Theory & Applications, 2024, 41 (1): 155- 162
[26]   罗彪, 胡天萌, 周育豪, 等 多智能体强化学习控制与决策研究综述[J]. 自动化学报, 2025, 51 (3): 510- 539
LUO Biao, HU Tianmeng, ZHOU Yuhao, et al Survey on multi-agent reinforcement learning for control and decision-making[J]. Acta Automatica Sinica, 2025, 51 (3): 510- 539
[27]   LIN S, QIAO G, TAI Y, et al. HWC-Loco: a hierarchical whole-body control approach to robust humanoid locomotion [EB/OL]. (2025-05-18) [2025-06-08]. https://arxiv.org/abs/2503.00923.
[28]   WAN C, WANG L, PHOHA V V A survey on gait recognition[J]. ACM Computing Surveys, 2018, 51 (5): 89
[29]   史晓国, 云静, 张钰莹, 等 步态识别研究综述[J]. 计算机工程与应用, 2025, 61 (14): 65- 87
SHI Xiaoguo, YUN Jing, ZHANG Yuying, et al Review of gait recognition research[J]. Computer Engineering and Applications, 2025, 61 (14): 65- 87
[30]   ALOTAIBI M, MAHMOOD A Improved gait recognition based on specialized deep convolutional neural network[J]. Computer Vision and Image Understanding, 2017, 164: 103- 110
doi: 10.1016/j.cviu.2017.10.004
[31]   LAL A, NITHYAKANI P. Gait speed based individual recognition model using deep 2-D convolutional neural network [C]// Proceedings of the International Conference on Computer Communication and Informatics. Coimbatore: IEEE, 2023: 1–6.
[32]   WANG Y, DU B, SHEN Y, et al. EV-gait: event-based robust gait recognition using dynamic vision sensors [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Long Beach: IEEE, 2019: 6351–6360.
[33]   YE D, FAN C, MA J, et al. BigGait: learning gait representation you want by large vision models [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle: IEEE, 2024: 200–210.
[34]   FAN C, PENG Y, CAO C, et al. GaitPart: temporal part-based model for gait recognition [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle: IEEE, 2020: 14213–14221.
[35]   GAO S, YUN J, ZHAO Y, et al Gait-D: skeleton-based gait feature decomposition for gait recognition[J]. IET Computer Vision, 2022, 16 (2): 111- 125
doi: 10.1049/cvi2.12070
[36]   MATHIS A, MAMIDANNA P, CURY K M, et al DeepLabCut: markerless pose estimation of user-defined body parts with deep learning[J]. Nature Neuroscience, 2018, 21 (9): 1281- 1289
doi: 10.1038/s41593-018-0209-y
[37]   NATH T, MATHIS A, CHEN A C, et al Using DeepLabCut for 3D markerless pose estimation across species and behaviors[J]. Nature Protocols, 2019, 14 (7): 2152- 2176
doi: 10.1038/s41596-019-0176-0
[38]   LAUER J, ZHOU M, YE S, et al Multi-animal pose estimation, identification and tracking with DeepLabCut[J]. Nature Methods, 2022, 19 (4): 496- 504
doi: 10.1038/s41592-022-01443-0
[39]   YE S, FILIPPOVA A, LAUER J, et al SuperAnimal pretrained pose estimation models for behavioral analysis[J]. Nature Communications, 2024, 15: 5165
doi: 10.1038/s41467-024-48792-2
[40]   MATHIS A, BIASI T, SCHNEIDER S, et al. Pretraining boosts out-of-domain robustness for pose estimation [C]// Proceedings of the IEEE Winter Conference on Applications of Computer Vision. Waikoloa: IEEE, 2021: 1858–1867.
[41]   薛均晓, 陈金浦, 董博威, 等. 基于深度确定性策略梯度算法的航空母舰舰载机导引路径规划与仿真[J/OL]. 计算机辅助设计与图形学学报, 2025: 1–14. (2025-02-07) [2025-04-04]. https://kns.cnki.net/kcms/detail/detail.aspx?dbcode=CJFQ&dbname=CJFD&filename=JSJF2025020600B.
XUE Junxiao, CHEN Jinpu, DONG Bowei, et al. Guidance path planning and simulation of aircraft carrier based on depth deterministic strategy gradient algorithm [J/OL]. Journal of Computer-Aided Design & Computer Graphics, 2025: 1–14. (2025-02-07) [2025-04-04]. https://kns.cnki.net/kcms/detail/detail.aspx?dbcode=CJFQ&dbname=CJFD&filename=JSJF2025020600B.
[1] Lang CHEN,Zengli LIU,Xuanzhi ZHAO. Decision-making for multi-USV adversarial encirclement based on adaptive curriculum reinforcement learning[J]. Journal of ZheJiang University (Engineering Science), 2026, 60(7): 1369-1380.
[2] Yong CHEN,Xizhi DU,Yiwei JIANG,Wenchao YI,Zhi PEI,Zuzhen JI. Deep reinforcement learning models and algorithms for single-machine scheduling considering deteriorated maintenance[J]. Journal of ZheJiang University (Engineering Science), 2026, 60(7): 1528-1538.
[3] Qiao LUO,Jun CHEN,Jie GENG,Chaoyang TANG,Chunyun FU. Review of energy management strategies for hybrid electric vehicles based on driving information[J]. Journal of ZheJiang University (Engineering Science), 2026, 60(7): 1539-1556.
[4] Yiwei ZHANG,Xin CUI,Qinghui ZHAO,Yan CHEN. Collaborative content caching optimization in UAV-assisted internet of vehicle based on NOMA[J]. Journal of ZheJiang University (Engineering Science), 2026, 60(6): 1289-1298.
[5] Qingqing YANG,Runpeng TANG,Yi PENG. Joint waveform and phase shift design in integrated sensing and communication systems[J]. Journal of ZheJiang University (Engineering Science), 2026, 60(4): 906-914.
[6] Jiale LIU,Yali XUE,Shan CUI,Jun HONG. TD3 mapless navigation algorithm guided by dynamic window approach[J]. Journal of ZheJiang University (Engineering Science), 2025, 59(8): 1671-1679.
[7] Kun HAO,Xuan MENG,Xiaofang ZHAO,Zhisheng LI. 3D underwater AUV path planning method integrating adaptive potential field method and deep reinforcement learning[J]. Journal of ZheJiang University (Engineering Science), 2025, 59(7): 1451-1461.
[8] Wei ZHAO,Wanzhi ZHANG,Jialin HOU,Rui HOU,Yuhua LI,Lejun ZHAO,Jin Cheng. Path planning of agricultural robots based on improved deep reinforcement learning algorithm[J]. Journal of ZheJiang University (Engineering Science), 2025, 59(7): 1492-1503.
[9] Mingfang ZHANG,Jian MA,Nale ZHAO,Li WANG,Ying LIU. Intelligent connected vehicle motion planning at unsignalized intersections based on deep reinforcement learning[J]. Journal of ZheJiang University (Engineering Science), 2024, 58(9): 1923-1934.
[10] Baolin YE,Ruitao SUN,Weimin WU,Bin CHEN,Qing YAO. Traffic signal control method based on asynchronous advantage actor-critic[J]. Journal of ZheJiang University (Engineering Science), 2024, 58(8): 1671-1680.
[11] Meng ZHANG,Dian-hai WANG,Sheng JIN. Deep reinforcement learning approach to signal control combined with domain experience[J]. Journal of ZheJiang University (Engineering Science), 2023, 57(12): 2524-2532.
[12] Yu-feng JIANG,Dong-sheng CHEN. Assembly strategy for large-diameter peg-in-hole based on deep reinforcement learning[J]. Journal of ZheJiang University (Engineering Science), 2023, 57(11): 2210-2216.
[13] Xia HUA,Xin-qing WANG,Ting RUI,Fa-ming SHAO,Dong WANG. Vision-driven end-to-end maneuvering object tracking of UAV[J]. Journal of ZheJiang University (Engineering Science), 2022, 56(7): 1464-1472.
[14] Zhi-min LIU,Bao-Lin YE,Yao-dong ZHU,Qing YAO,Wei-min WU. Traffic signal control method based on deep reinforcement learning[J]. Journal of ZheJiang University (Engineering Science), 2022, 56(6): 1249-1256.
[15] Qi-lin DENG,Juan LU,Yong-hui CHEN,Jian FENG,Xiao-ping LIAO,Jun-yan MA. Optimization method of CNC milling parameters based on deep reinforcement learning[J]. Journal of ZheJiang University (Engineering Science), 2022, 56(11): 2145-2155.