Please wait a minute...
浙江大学学报(工学版)  2026, Vol. 60 Issue (9): 1980-1990    DOI: 10.3785/j.issn.1008-973X.2026.09.015
计算机技术、自动控制技术     
基于混合空洞注意力机制的图卷积强化学习
宋莉1,2(),宛袁玉3,*(),宋明黎2
1. 浙大城市学院 计算机与计算科学学院,浙江 杭州 310015
2. 浙江大学 计算机科学与技术学院,浙江 杭州 310027
3. 浙江大学 软件学院,浙江 宁波 315100
Graph convolutional reinforcement learning via hybrid dilated attention mechanism
Li SONG1,2(),Yuanyu WAN3,*(),Mingli SONG2
1. School of Computer and Computing Science, Hangzhou City University, Hangzhou 310015, China
2. College of Computer Science and Technology, Zhejiang University, Hangzhou 310027, China
3. School of Software Technology, Zhejiang University, Ningbo 315100, China
 全文: PDF(1974 KB)   HTML
摘要:

针对复杂高动态多智能体强化学习中多样化动态信息捕捉困难、计算复杂的难题,提出基于混合空洞注意力的多智能体图卷积强化学习算法,以平衡计算复杂度和感受野大小,提高学习效率. 获取能够捕捉多智能体动态交互的图动态矩阵. 混合空洞注意力机制利用Bagging融合了多头注意力和多尺度空洞注意力机制的优点,使模型有效地学习不同类型特征. 在损失函数中添加正则化项以促进模型稳定性并解决过拟合问题. 为了平衡动作的探索与利用,采用改进的暂时拓展贪心方法选择多智能体动作. 仿真结果表明,在复杂动态环境中,所提方法的策略优化准确性和稳定性优于现有方法.

关键词: 多智能体图强化学习混合空洞注意力机制暂时拓展贪心方法策略优化    
Abstract:

A multi-agent graph convolutional reinforcement learning based on hybrid dilated attention was proposed to address the challenges of capturing diverse dynamic information and computational complexity in complex high-dynamic multi-agent reinforcement learning. The multi-agent graph convolutional reinforcement learning with ensemble-based hybrid dilated attention was proposed to balance computational complexity and receptive field size in order to enhance learning efficiency. The dynamics matrix of the underlying graph was obtained to capture the dynamic interactions of multiple agents. The hybrid dilated attention mechanism used the Bagging method to combine the advantages of the multi-head attention and the multi-scale dilated attention mechanisms in order to enable the model to effectively learn different types of features. A regularization was added into the loss function to enhance the stability of the model and address the overfitting problem. The improved temporally extended greedy method was utilized to choose the multi-agent’s actions in order to achieve the exploration and exploitation trade-off of actions. The simulation results demonstrate that the proposed method achieves higher accuracy and greater stability in policy optimization than the existing approaches in complex dynamic environments.

Key words: multi-agent    graph reinforcement learning    hybrid dilated attention mechanism    temporally-extended greedy method    policy optimization
收稿日期: 2025-08-17 出版日期: 2026-07-20
CLC:  TP 181  
基金资助: 浙江省自然科学基金联合基金资助项目(LHZSD24F020001);国家自然科学基金资助项目(62306275, U20B2066);浙江大学上海高等研究院繁星科学基金资助项目(SN-ZJU-SIAS-001);中央高校基本科研业务费专项资金资助项目(226-2023-00048).
通讯作者: 宛袁玉     E-mail: slili516@zjsru.edu.cn;wanyy@zju.edu.cn
作者简介: 宋莉(1991—),女,博士,从事图多智能体强化学习算法研究. orcid.org/0000-0003-2616-5156. E-mail:slili516@zjsru.edu.cn
服务  
把本文推荐给朋友
加入引用管理器
E-mail Alert
作者相关文章  
宋莉
宛袁玉
宋明黎

引用本文:

宋莉,宛袁玉,宋明黎. 基于混合空洞注意力机制的图卷积强化学习[J]. 浙江大学学报(工学版), 2026, 60(9): 1980-1990.

Li SONG,Yuanyu WAN,Mingli SONG. Graph convolutional reinforcement learning via hybrid dilated attention mechanism. Journal of ZheJiang University (Engineering Science), 2026, 60(9): 1980-1990.

链接本文:

https://www.zjujournals.com/eng/CN/10.3785/j.issn.1008-973X.2026.09.015        https://www.zjujournals.com/eng/CN/Y2026/V60/I9/1980

图 1  基于混合空洞注意力机制的多智能体强化学习框架
图 2  Surviving环境
图 3  所提算法与基准算法的分数对比
图 4  不同图正则化系数下所提出算法的分数
图 5  Surviving环境中所提算法和基准算法的结果
算法损失值KL-lossattenq成功率
传统GCRL6.15×10163.99×1016–3.93×10158.800.68
GCRL-MHA7.54×10135.02×1016–4.87×10159.210.71
GCRL-MSDA6.82×10104.58×1016–4.89×101511.590.89
GCRL-EDA6.88×1074.54×1016–4.93×101512.020.92
表 1  所提算法与基准算法收敛时的实验结果
图 6  不同策略和图正则化的消融实验结果
图 7  学习过程中变化图的动态节点
轮次scores
GCRL-MHAGCRL-MSDAGCRL-EDA
169.56118.05130.85
1076.1079.9454.12
2020.4525.3621.26
3024.1322.8722.41
表 2  GCRL-MHA、GCRL-MSDA、GCRL-EDA的分数
图 8  高速路环境下所提算法与基准算法的实验结果
图 9  交汇路口环境中所提算法与基准算法的q值和学习率
图 10  所提出算法在快速高速路环境下的学习率、q值、GPU温度及利用率
图 11  十字路口城市车辆间的部分交互图网络
1 张萌, 王殿海, 金盛 结合领域经验的深度强化学习信号控制方法[J]. 浙江大学学报: 工学版, 2023, 57 (12): 2524- 2532
ZHANG Meng, WANG Dianhai, JIN Sheng Deep reinforcement learning approach to signal control combined with domain experience[J]. Journal of Zhejiang University: Engineering Science, 2023, 57 (12): 2524- 2532
2 CHU K F, LAM A Y, LI V O Traffic signal control using end-to-end off-policy deep reinforcement learning[J]. IEEE Transactions on Intelligent Transportation Systems, 2022, 23 (7): 7184- 7195
doi: 10.1109/TITS.2021.3067057
3 郭方洪, 伍泽芃, 杨淏, 等 基于个性化联邦强化学习的异构多微网能量调度[J]. 自动化学报, 2025, 51 (9): 2072- 2084
GUO Fanghong, WU Zepeng, YANG Hao, et al Energy scheduling of heterogeneous multi-microgrid based on personalized federated reinforcement learning[J]. Acta Automatica Sinica, 2025, 51 (9): 2072- 2084
doi: 10.16383/j.aas.c250130
4 GAUTIER P, LAURENT J, DIGUET J P Deep Q-learning-based dynamic management of a robotic cluster[J]. IEEE Transactions on Automation Science and Engineering, 2023, 20 (4): 2503- 2515
doi: 10.1109/TASE.2022.3205651
5 SHI D, LI L, OHTSUKI T, et al Make smart decisions faster: deciding D2D resource allocation via stackelberg game guided multi-agent deep reinforcement learning[J]. IEEE Transactions on Mobile Computing, 2022, 21 (12): 4426- 4438
doi: 10.1109/TMC.2021.3085206
6 MUNIKOTI S, AGARWAL D, DAS L, et al Challenges and opportunities in deep reinforcement learning with graph neural networks: a comprehensive review of algorithms and applications[J]. IEEE Transactions on Neural Networks and Learning Systems, 2024, 35 (11): 15051- 15071
doi: 10.1109/tnnls.2023.3283523
7 荣垂田, 田浩辉, 杜方 多目标深度强化学习驱动的数据库系统参数优化技术[J]. 软件学报, 2025, 36 (12): 5512- 5536
RONG Chuitian, TIAN Haohui, DU Fang Technique for database system parameter optimization using multi-objective deep reinforcement learning[J]. Journal of Software, 2025, 36 (12): 5512- 5536
doi: 10.13328/j.cnki.jos.007405
8 ZHAO X Y, WU C. Large-scale machine learning cluster scheduling via multi-agent graph reinforcement learning [C]// Thirty-Second AAAI Conference on Artificial Intelligence. Menlo Park: AAAI, 2022: 25082515.
9 PU Z Q, WANG H M, LIU Z, et al Attention enhanced reinforcement learning for multi agent cooperation[J]. IEEE Transactions on Neural Networks and Learning Systems, 2023, 34 (11): 8235- 8249
doi: 10.1109/TNNLS.2022.3146858
10 李岳珩, 谢广明 集中训练分布执行下的多智能体强化学习综述[J]. 控制理论与应用, 2025, 42 (11): 2114- 2124
LI Yueheng, XIE Guangming Review of multi-agent reinforcement learning under centralized training with decentralized execution[J]. Control Theory and Applications, 2025, 42 (11): 2114- 2124
doi: 10.7641/CTA.2025.50009
11 XING Q, XU Y, CHEN Z, et al A graph reinforcement learning-based decision-making platform for real-time charging navigation of urban electric vehicles[J]. IEEE Transactions on Industrial Informatics, 2023, 19 (3): 3284- 3295
doi: 10.1109/TII.2022.3210264
12 HOUIDI O, BAKRI S, ZEGHLACHE D. Multi-agent graph convolutional reinforcement learning for intelligent load balancing [C]// NOMS 2022-2022 IEEE/IFIP Network Operations and Management Symposium. Piscataway: IEEE, 2022: 1-6.
13 LIN W Y, SONG Y Z, RUAN B K, et al Temporal difference-aware graph convolutional reinforcement learning for multi-intersection traffic signal control[J]. IEEE Transactions on Intelligent Transportation Systems, 2023, 25 (1): 327- 337
14 LUO J, LI C F, FAN Q Q, et al A graph convolutional encoder and multi-head attention decoder network for tsp via reinforcement learning[J]. Engineering Applications of Artificial Intelligence, 2022, 112: 104848
doi: 10.1016/j.engappai.2022.104848
15 DU X Q, CHEN H C, XING Y H, et al A contrastive enhanced ensemble framework for efficient multi-agent reinforcement learning[J]. Expert Systems with Applications, 2024, 245: 123158
doi: 10.1016/j.eswa.2024.123158
16 HU Y F, FU J J, WEN G H Graph soft actor-critic reinforcement learning for large-scale distributed multirobot coordination[J]. IEEE Transactions on Neural Networks and Learning Systems, 2025, 36 (1): 665- 676
doi: 10.1109/TNNLS.2023.3329530
17 XIE S R, ZHANG H, YU H, et al ET-HF: a novel information sharing model to improve multi-agent cooperation[J]. Knowledge-based Systems, 2022, 257: 109916
doi: 10.1016/j.knosys.2022.109916
18 罗佳, 李朝锋 基于残差图卷积网络与深度强化学习的需求可拆分车辆路径优化算法[J]. 控制理论与应用, 2024, 41 (6): 1123- 1136
LUO Jia, LI Chaofeng The split delivery vehicle routing optimization with the residual graph convolutional network and deep reinforcement learning[J]. Control Theory and Applications, 2024, 41 (6): 1123- 1136
doi: 10.7641/CTA.2023.21040
19 GOECKNER A, SUI Y Y, MARTINET N, et al. Graph neural network-based reinforcement learning for multi-agent systems [C]// Thirty-Sixth Conference on Neural Information Processing Systems. Cambridge: MIT Press, 2022: 1-8.
20 MA X, XIE Y, CHIGAN C Graph convolutional network based multi-objective meta-deep Q-learning for eco-routing[J]. IEEE Transactions on Intelligent Transportation Systems, 2024, 25 (7): 7323- 7338
doi: 10.1109/TITS.2023.3348034
21 ELLIS J D, IQBAL R, YOSHIMATSU K Deep Q-learning-based molecular graph generation for chemical structure prediction from infrared spectra[J]. IEEE Transactions on Artificial Intelligence, 2024, 5 (2): 634- 646
doi: 10.1109/TAI.2023.3287947
22 ROKHFOROZ P, MONTAZERI M, FINK O Multi-agent reinforcement learning with graph convolutional neural networks for optimal bidding strategies of generation units in electricity markets[J]. Expert Systems with Applications, 2023, 225: 120010
doi: 10.1016/j.eswa.2023.120010
23 XIA L Q, LIANG Y S, LENG J W, et al Maintenance planning recommendation of complex industrial equipment based on knowledge graph and graph neural network[J]. Reliability Engineering and System Safety, 2023, 232: 109068
doi: 10.1016/j.ress.2022.109068
24 YE Y, JI S H Sparse graph attention networks[J]. IEEE Transactions on Knowledge and Data Engineering, 2023, 35 (1): 905- 916
doi: 10.1109/TKDE.2021.3072345
25 FAWAZ H, LESCA J, QUANG P T, et al Graph convolutional reinforcement learning for collaborative queuing agents[J]. IEEE Transactions on Network and Service Management, 2023, 20 (2): 13631377
doi: 10.1109/tnsm.2022.3226605
26 JIANG J C, DUN C, HUANG T J, et al. Graph convolutional reinforcement learning [C]// International Conference on Learning Representations. Washington DC: [s.n.], 2020: 1-13.
[1] 连远锋,范树玉,张本哲,王森. 融合几何约束特征提取与多奖励协同优化的仿生步态控制[J]. 浙江大学学报(工学版), 2026, 60(9): 1862-1871.
[2] 曹宁博,万启超,赵利英,李梓萌,黄宝林. 集成多智能体强化学习和最大压强控制的混行交叉口协同控制[J]. 浙江大学学报(工学版), 2026, 60(8): 1819-1831.
[3] 陈浪,刘增力,赵宣植. 基于自适应课程强化学习的多无人艇对抗围捕决策[J]. 浙江大学学报(工学版), 2026, 60(7): 1369-1380.
[4] 汪洋,刘红超,田池,吴兵,张笛. 航线交换机制下多船避碰的策略学习与博弈决策[J]. 浙江大学学报(工学版), 2026, 60(5): 964-976.
[5] 张静,王祎,陈子龙,李云松. 多目标多智能体路径规划方法[J]. 浙江大学学报(工学版), 2025, 59(8): 1689-1697.
[6] 黎文皓,季彦婕,吴浩,贾叶雯,张水潮. 基于多智体的自动驾驶汽车停车空载收费策略[J]. 浙江大学学报(工学版), 2024, 58(8): 1659-1670.
[7] 王卓,李永强,冯宇,冯远静. 两方零和马尔科夫博弈策略梯度算法及收敛性分析[J]. 浙江大学学报(工学版), 2024, 58(3): 480-491.
[8] 薛雅丽,叶金泽,李寒雁. 基于改进强化学习的多智能体追逃对抗[J]. 浙江大学学报(工学版), 2023, 57(8): 1479-1486.
[9] 李勇,柳富强,孙柏青,张秋豪,杨俊友. 日常养老情境的异构多机器人动态多任务分配[J]. 浙江大学学报(工学版), 2022, 56(9): 1806-1814.
[10] 赵永胜,李瑞祥,牛娜娜,赵志勇. 数字孪生驱动的机身形状控制方法[J]. 浙江大学学报(工学版), 2022, 56(7): 1457-1463.
[11] 徐小高,夏莹杰,朱思雨,邝砾. 基于强化学习的多路口可变车道协同控制方法[J]. 浙江大学学报(工学版), 2022, 56(5): 987-994, 1005.
[12] 张盼,丁华,张颖而,李冰凝,皇甫江涛,金仲和. 基于信息共享的多智能体自主电子干扰系统[J]. 浙江大学学报(工学版), 2022, 56(1): 75-83.
[13] 邵杭蕾,张冬梅. 基于静态输出反馈协议的多智能体系统同步[J]. 浙江大学学报(工学版), 2020, 54(7): 1308-1315.
[14] 董如良, 杨强, 颜文俊. 多智能体协同寻优的主动配网动态拓扑重构[J]. 浙江大学学报(工学版), 2015, 49(10): 1982-1989.
[15] 娄柯, 齐斌, 穆文英, 崔宝同. 基于反馈控制策略的多智能体蜂拥控制[J]. J4, 2013, 47(10): 1758-1763.