Please wait a minute...
浙江大学学报(工学版)  2026, Vol. 60 Issue (9): 1862-1871    DOI: 10.3785/j.issn.1008-973X.2026.09.003
机械工程     
融合几何约束特征提取与多奖励协同优化的仿生步态控制
连远锋1,2(),范树玉1,张本哲1,王森1
1. 中国石油大学(北京) 人工智能学院,北京 102249
2. 中国石油大学(北京) 石油数据挖掘北京市重点实验室,北京 102249
Bio-inspired gait control integrating geometric constraint feature extraction and multi-reward cooperative optimization
Yuanfeng LIAN1,2(),Shuyu FAN1,Benzhe ZHANG1,Sen WANG1
1. College of Artificial Intelligence, China University of Petroleum, Beijing 102249, China
2. Beijing Key Laboratory of Petroleum Data Mining, China University of Petroleum, Beijing 102249, China
 全文: PDF(4099 KB)   HTML
摘要:

针对四足机器人在复杂地形中运动稳定性与环境适应性不足的问题,提出融合几何约束特征提取与多奖励协同优化的仿生步态控制方法. 构建基于动态时间规整(DTW)与卷积处理的仿生步态编码器模块,通过逆运动学解算、关节速度微分计算来提取时空关联的步态特征. 结合微分运动学驱动的几何约束机制,将任务空间轨迹与关节力矩的耦合映射嵌入近端策略优化(PPO)网络来提升关节运动协调性与采样效率.设计融合能效因子与李雅普诺夫稳定性的分层奖励函数,通过指数型能效奖励和稳定性奖励实现仿生步态的多目标协同控制. 实验结果表明,所提方法在仿真环境中的碰撞平均值、适应性损失等优于对比算法,且有效改善了关节触地现象. 在真机复杂地形部署测试中,该方法有效提升了四足机器人的运动稳定性与环境适应性.

关键词: 步态特征提取几何约束分层奖励机制深度强化学习近端策略优化    
Abstract:

A new bio-inspired gait control method was proposed to address the challenges of insufficient locomotion stability and environmental adaptability of quadruped robots in complex terrains. A bio-inspired gait encoder module based on dynamic time warping (DTW) and convolutional processing was constructed to extract spatiotemporally correlated gait features through inverse kinematics solution and joint velocity differentiation calculation. A geometric constraint mechanism driven by differential kinematics was incorporated, and the coupled mapping between task-space trajectories and joint torques was embedded into the proximal policy optimization (PPO) network to enhance the joint motion coordination and the sampling efficiency. A hierarchical reward function combining energy efficiency factors and the Lyapunov stability was designed, the and multi-objective collaborative control of bio-inspired gaits was achieved via exponential energy efficiency rewards and stability rewards. The experimental results demonstrated that the proposed method outperformed the comparative algorithms in terms of indicators such as average collision value and adaptive loss function in the simulation environment, and effectively mitigated the joint ground contact phenomenon. This method significantly enhanced the motion stability and the environmental adaptability of the quadruped robot in the deployment test of the real robot in complex terrains.

Key words: gait feature extraction    geometric constraint    hierarchical reward mechanism    deep reinforcement learning    proximal policy optimization
收稿日期: 2025-08-14 出版日期: 2026-07-20
CLC:  TP 393  
作者简介: 连远锋(1977—),男,教授,从事图像处理、虚拟现实、机器视觉、数字孪生、深度学习、具身智能和油气多模态大模型研究. orcid.org/0000-0002-1801-2507. E-mail:lianyuanfeng@cup.edu.cn
服务  
把本文推荐给朋友
加入引用管理器
E-mail Alert
作者相关文章  
连远锋
范树玉
张本哲
王森

引用本文:

连远锋,范树玉,张本哲,王森. 融合几何约束特征提取与多奖励协同优化的仿生步态控制[J]. 浙江大学学报(工学版), 2026, 60(9): 1862-1871.

Yuanfeng LIAN,Shuyu FAN,Benzhe ZHANG,Sen WANG. Bio-inspired gait control integrating geometric constraint feature extraction and multi-reward cooperative optimization. Journal of ZheJiang University (Engineering Science), 2026, 60(9): 1862-1871.

链接本文:

https://www.zjujournals.com/eng/CN/10.3785/j.issn.1008-973X.2026.09.003        https://www.zjujournals.com/eng/CN/Y2026/V60/I9/1862

图 1  仿生步态控制方法框架
图 2  几何约束关系示意图
图 3  PPO算法应用流程
图 4  四足机器人步态学习过程图
图 5  四足机器人的多维度性能指标变化图
图 6  四足机器人关键运动性能奖励指标变化图
图 7  采用不同算法的四足机器人运动性能仿真对比
图 8  采用不同算法的四足机器人在仿真训练中的碰撞均值、适应性损失及总奖励均值对比
图 9  采用不同算法的四足机器人的关键运动性能奖励指标对比图
图 10  沙地场景下采用不同算法的四足机器人的运动性能对比
图 11  荒地场景下采用不同算法的四足机器人的运动性能对比
算法名称$ {t_{\text{e}}} $/s$ {c_{\text{e}}} $
walk-these-ways-go26.117.51
SAC6.059.19
所提方法6.833.76
表 1  四足机器人仿真训练中不同算法的计算效率与碰撞性能对比
1 谭民, 王硕 机器人技术研究进展[J]. 自动化学报, 2013, 39 (7): 963- 972
TAN Min, WANG Shuo Research progress on robotics[J]. Acta Automatica Sinica, 2013, 39 (7): 963- 972
2 贾庆轩, 袁博楠, 陈钢, 等 关节锁定空间机械臂负载操作能力评估与轨迹规划[J]. 控制与决策, 2020, 35 (1): 243- 249
JIA Qingxuan, YUAN Bonan, CHEN Gang, et al Load carrying capacity evaluation and task trajectory planning of space manipulator with the locked joint[J]. Control and Decision, 2020, 35 (1): 243- 249
3 尤波, 曲伟健, 李佳钰 面向双操作者的六足机器人共享遥操作[J]. 控制与决策, 2022, 37 (11): 2769- 2778
YOU Bo, QU Weijian, LI Jiayu Shared teleoperation of hexapod robot for dual operators[J]. Control and Decision, 2022, 37 (11): 2769- 2778
4 郭非, 汪首坤, 王军政 轮足复合移动机器人运动规划发展现状及关键技术分析[J]. 控制与决策, 2022, 37 (6): 1433- 1444
GUO Fei, WANG Shoukun, WANG Junzheng Development status and key technology analysis for motion planning of wheel-legged hybrid mobile robot[J]. Control and Decision, 2022, 37 (6): 1433- 1444
5 GUO H, ZHU J, CHEN Y E-LOAM: LiDAR odometry and mapping with expanded local structural information[J]. IEEE Transactions on Intelligent Vehicles, 2023, 8 (2): 1911- 1921
doi: 10.1109/TIV.2022.3151665
6 ASADI F, KHORRAM M, MOOSAVIAN S A A. CPG-based gait transition of a quadruped robot [C]// Proceedings of the 3rd RSI International Conference on Robotics and Mechatronics. Tehran: IEEE, 2015: 210–215.
7 DING Y, PANDALA A, LI C, et al Representation-free model predictive control for dynamic motions in quadrupeds[J]. IEEE Transactions on Robotics, 2021, 37 (4): 1154- 1171
doi: 10.1109/TRO.2020.3046415
8 汪首坤, 刘大江, 郭非, 等 基于动力学模型预测控制的Stewart型轮足腿控制方法[J]. 北京理工大学学报, 2021, 41 (4): 418- 424
WANG Shoukun, LIU Dajiang, GUO Fei, et al Stewart type wheel-foot leg control method based on dynamics and model predictive control[J]. Transactions of Beijing Institute of Technology, 2021, 41 (4): 418- 424
9 朱晓庆, 陈江涛, 张思远, 等 基于深度仲裁策略的四足机器人步态学习[J]. 北京理工大学学报, 2023, 43 (11): 1197- 1204
ZHU Xiaoqing, CHEN Jiangtao, ZHANG Siyuan, et al Gait learning of quadruped robot based on deep arbitration strategy[J]. Transactions of Beijing Institute of Technology, 2023, 43 (11): 1197- 1204
10 张思远, 朱晓庆, 阮晓钢, 等 基于环境反馈机制的四足机器人运动技能学习[J]. 控制与决策, 2024, 39 (5): 1461- 1468
ZHANG Siyuan, ZHU Xiaoqing, RUAN Xiaogang, et al Motor skill learning of quadruped robot based on environmental feedback mechanism[J]. Control and Decision, 2024, 39 (5): 1461- 1468
11 BUTYREV L, EDELHÄUβER T, MUTSCHLER C. Deep reinforcement learning for motion planning of mobile robots [EB/OL]. (2019-12-19) [2025-04-04]. https://arxiv.org/abs/1912.09260.
12 XIONG J, WANG Q, YANG Z, et al. Parametrized deep Q-networks learning: reinforcement learning with discrete-continuous hybrid action space [EB/OL]. (2018-10-10) [2025-04-04]. https://arxiv.org/abs/1810.06394.
13 VAN HASSELT H, GUEZ A, SILVER D. Deep reinforcement learning with double Q-Learning [C]// Thirtieth AAAI Conference on Artificial Intelligence. Phoenix: AAAI Press, 2016: 2094–2100.
14 WANG Z, SCHAUL T, HESSEL M, et al. Dueling network architectures for deep reinforcement learning [C]// Proceedings of the 33rd International Conference on Machine Learning. New York: JMLR, 2016: 1995–2003.
15 MNIH V, BADIA A P, MIRZA M, et al. Asynchronous methods for deep reinforcement learning [C]// Proceedings of the 33rd International Conference on Machine Learning. New York: JMLR, 2016: 1928–1937.
16 LILLICRAP T P, HUNT J J, PRITZEL A, et al. Continuous control with deep reinforcement learning [EB/OL]. (2019-07-05) [2025-04-04]. https://arxiv.org/abs/1509.02971.
17 GANGAPURWALA S, GEISERT M, ORSOLINO R, et al RLOC: terrain-aware legged locomotion using reinforcement learning and optimal control[J]. IEEE Transactions on Robotics, 2022, 38 (5): 2908- 2927
doi: 10.1109/TRO.2022.3172469
18 董豪, 杨静, 李少波, 等 基于深度强化学习的机器人运动控制研究进展[J]. 控制与决策, 2022, 37 (2): 278- 292
DONG Hao, YANG Jing, LI Shaobo, et al Research progress of robot motion control based on deep reinforcement learning[J]. Control and Decision, 2022, 37 (2): 278- 292
19 HAARNOJA T, HA S, ZHOU A, et al. Learning to walk via deep reinforcement learning [C]// Robotics: Science and Systems XV. Freiburg im Breisgau: RSS Foundation, 2019.
20 FUCHIOKA Y. Imitating optimized trajectories for dynamic quadruped behaviors [D]. Vancouver: University of British Columbia, 2023.
21 ISCEN A, CALUWAERTS K, TAN J, et al. Policies modulating trajectory generators [C]// Proceedings of the 2nd Conference on Robot Learning. Zurich: PMLR, 2018: 916–926.
22 SCHULMAN J, WOLSKI F, DHARIWAL P, et al. Proximal policy optimization algorithms [EB/OL]. (2017-08-28) [2025-04-04]. https://arxiv.org/abs/1707.06347.
23 RUDIN N, HOELLER D, REIST P, et al. Learning to walk in minutes using massively parallel deep reinforcement learning [C]// Proceedings of the 5th Conference on Robot Learning. London: PMLR, 2021: 91–100.
24 MARGOLIS G B, AGRAWAL P. Walk these ways: tuning robot control for generalization with multiplicity of behavior [C]// Proceedings of the 6th Conference on Robot Learning. Auckland: PMLR, 2022: 22–31.
25 朱晓庆, 刘鑫源, 阮晓钢, 等 融合元学习和PPO算法的四足机器人运动技能学习方法[J]. 控制理论与应用, 2024, 41 (1): 155- 162
ZHU Xiaoqing, LIU Xinyuan, RUAN Xiaogang, et al A quadruped robot kinematic skill learning method integrating meta-learning and PPO algorithms[J]. Control Theory & Applications, 2024, 41 (1): 155- 162
26 罗彪, 胡天萌, 周育豪, 等 多智能体强化学习控制与决策研究综述[J]. 自动化学报, 2025, 51 (3): 510- 539
LUO Biao, HU Tianmeng, ZHOU Yuhao, et al Survey on multi-agent reinforcement learning for control and decision-making[J]. Acta Automatica Sinica, 2025, 51 (3): 510- 539
27 LIN S, QIAO G, TAI Y, et al. HWC-Loco: a hierarchical whole-body control approach to robust humanoid locomotion [EB/OL]. (2025-05-18) [2025-06-08]. https://arxiv.org/abs/2503.00923.
28 WAN C, WANG L, PHOHA V V A survey on gait recognition[J]. ACM Computing Surveys, 2018, 51 (5): 89
29 史晓国, 云静, 张钰莹, 等 步态识别研究综述[J]. 计算机工程与应用, 2025, 61 (14): 65- 87
SHI Xiaoguo, YUN Jing, ZHANG Yuying, et al Review of gait recognition research[J]. Computer Engineering and Applications, 2025, 61 (14): 65- 87
30 ALOTAIBI M, MAHMOOD A Improved gait recognition based on specialized deep convolutional neural network[J]. Computer Vision and Image Understanding, 2017, 164: 103- 110
doi: 10.1016/j.cviu.2017.10.004
31 LAL A, NITHYAKANI P. Gait speed based individual recognition model using deep 2-D convolutional neural network [C]// Proceedings of the International Conference on Computer Communication and Informatics. Coimbatore: IEEE, 2023: 1–6.
32 WANG Y, DU B, SHEN Y, et al. EV-gait: event-based robust gait recognition using dynamic vision sensors [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Long Beach: IEEE, 2019: 6351–6360.
33 YE D, FAN C, MA J, et al. BigGait: learning gait representation you want by large vision models [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle: IEEE, 2024: 200–210.
34 FAN C, PENG Y, CAO C, et al. GaitPart: temporal part-based model for gait recognition [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle: IEEE, 2020: 14213–14221.
35 GAO S, YUN J, ZHAO Y, et al Gait-D: skeleton-based gait feature decomposition for gait recognition[J]. IET Computer Vision, 2022, 16 (2): 111- 125
doi: 10.1049/cvi2.12070
36 MATHIS A, MAMIDANNA P, CURY K M, et al DeepLabCut: markerless pose estimation of user-defined body parts with deep learning[J]. Nature Neuroscience, 2018, 21 (9): 1281- 1289
doi: 10.1038/s41593-018-0209-y
37 NATH T, MATHIS A, CHEN A C, et al Using DeepLabCut for 3D markerless pose estimation across species and behaviors[J]. Nature Protocols, 2019, 14 (7): 2152- 2176
doi: 10.1038/s41596-019-0176-0
38 LAUER J, ZHOU M, YE S, et al Multi-animal pose estimation, identification and tracking with DeepLabCut[J]. Nature Methods, 2022, 19 (4): 496- 504
doi: 10.1038/s41592-022-01443-0
39 YE S, FILIPPOVA A, LAUER J, et al SuperAnimal pretrained pose estimation models for behavioral analysis[J]. Nature Communications, 2024, 15: 5165
doi: 10.1038/s41467-024-48792-2
40 MATHIS A, BIASI T, SCHNEIDER S, et al. Pretraining boosts out-of-domain robustness for pose estimation [C]// Proceedings of the IEEE Winter Conference on Applications of Computer Vision. Waikoloa: IEEE, 2021: 1858–1867.
41 薛均晓, 陈金浦, 董博威, 等. 基于深度确定性策略梯度算法的航空母舰舰载机导引路径规划与仿真[J/OL]. 计算机辅助设计与图形学学报, 2025: 1–14. (2025-02-07) [2025-04-04]. https://kns.cnki.net/kcms/detail/detail.aspx?dbcode=CJFQ&dbname=CJFD&filename=JSJF2025020600B.
XUE Junxiao, CHEN Jinpu, DONG Bowei, et al. Guidance path planning and simulation of aircraft carrier based on depth deterministic strategy gradient algorithm [J/OL]. Journal of Computer-Aided Design & Computer Graphics, 2025: 1–14. (2025-02-07) [2025-04-04]. https://kns.cnki.net/kcms/detail/detail.aspx?dbcode=CJFQ&dbname=CJFD&filename=JSJF2025020600B.
[1] 陈浪,刘增力,赵宣植. 基于自适应课程强化学习的多无人艇对抗围捕决策[J]. 浙江大学学报(工学版), 2026, 60(7): 1369-1380.
[2] 陈勇,杜习之,姜一炜,易文超,裴植,纪祖臻. 考虑劣化维护的单机调度深度强化学习模型和算法[J]. 浙江大学学报(工学版), 2026, 60(7): 1528-1538.
[3] 罗桥,陈俊,耿杰,唐朝阳,傅春耘. 基于行车信息的混合动力汽车能量管理策略综述[J]. 浙江大学学报(工学版), 2026, 60(7): 1539-1556.
[4] 张艺炜,崔鑫,赵庆慧,陈燕. 无人机辅助车联网NOMA协同缓存优化[J]. 浙江大学学报(工学版), 2026, 60(6): 1289-1298.
[5] 杨青青,唐润朋,彭艺. 通信感知一体化系统中的联合波形与相移设计[J]. 浙江大学学报(工学版), 2026, 60(4): 906-914.
[6] 柳佳乐,薛雅丽,崔闪,洪君. 动态窗口法引导的TD3无地图导航算法[J]. 浙江大学学报(工学版), 2025, 59(8): 1671-1679.
[7] 郝琨,孟璇,赵晓芳,李志圣. 融合自适应势场法和深度强化学习的三维水下AUV路径规划方法[J]. 浙江大学学报(工学版), 2025, 59(7): 1451-1461.
[8] 赵威,张万枝,侯加林,侯瑞,李玉华,赵乐俊,程进. 基于改进深度强化学习算法的农业机器人路径规划[J]. 浙江大学学报(工学版), 2025, 59(7): 1492-1503.
[9] 张名芳,马健,赵娜乐,王力,刘颖. 无信号交叉口处基于深度强化学习的智能网联车辆运动规划[J]. 浙江大学学报(工学版), 2024, 58(9): 1923-1934.
[10] 叶宝林,孙瑞涛,吴维敏,陈滨,姚青. 基于异步优势演员-评论家的交通信号控制方法[J]. 浙江大学学报(工学版), 2024, 58(8): 1671-1680.
[11] 张萌,王殿海,金盛. 结合领域经验的深度强化学习信号控制方法[J]. 浙江大学学报(工学版), 2023, 57(12): 2524-2532.
[12] 姜玉峰,陈东生. 基于深度强化学习的大口径轴孔装配策略[J]. 浙江大学学报(工学版), 2023, 57(11): 2210-2216.
[13] 华夏,王新晴,芮挺,邵发明,王东. 视觉感知的无人机端到端目标跟踪控制技术[J]. 浙江大学学报(工学版), 2022, 56(7): 1464-1472.
[14] 刘智敏,叶宝林,朱耀东,姚青,吴维敏. 基于深度强化学习的交通信号控制方法[J]. 浙江大学学报(工学版), 2022, 56(6): 1249-1256.
[15] 邓齐林,鲁娟,陈勇辉,冯健,廖小平,马俊燕. 基于深度强化学习的数控铣削加工参数优化方法[J]. 浙江大学学报(工学版), 2022, 56(11): 2145-2155.