Please wait a minute...
Journal of ZheJiang University (Engineering Science)  2026, Vol. 60 Issue (9): 1980-1990    DOI: 10.3785/j.issn.1008-973X.2026.09.015
    
Graph convolutional reinforcement learning via hybrid dilated attention mechanism
Li SONG1,2(),Yuanyu WAN3,*(),Mingli SONG2
1. School of Computer and Computing Science, Hangzhou City University, Hangzhou 310015, China
2. College of Computer Science and Technology, Zhejiang University, Hangzhou 310027, China
3. School of Software Technology, Zhejiang University, Ningbo 315100, China
Download: HTML     PDF(1974KB) HTML
Export: BibTeX | EndNote (RIS)      

Abstract  

A multi-agent graph convolutional reinforcement learning based on hybrid dilated attention was proposed to address the challenges of capturing diverse dynamic information and computational complexity in complex high-dynamic multi-agent reinforcement learning. The multi-agent graph convolutional reinforcement learning with ensemble-based hybrid dilated attention was proposed to balance computational complexity and receptive field size in order to enhance learning efficiency. The dynamics matrix of the underlying graph was obtained to capture the dynamic interactions of multiple agents. The hybrid dilated attention mechanism used the Bagging method to combine the advantages of the multi-head attention and the multi-scale dilated attention mechanisms in order to enable the model to effectively learn different types of features. A regularization was added into the loss function to enhance the stability of the model and address the overfitting problem. The improved temporally extended greedy method was utilized to choose the multi-agent’s actions in order to achieve the exploration and exploitation trade-off of actions. The simulation results demonstrate that the proposed method achieves higher accuracy and greater stability in policy optimization than the existing approaches in complex dynamic environments.



Key wordsmulti-agent      graph reinforcement learning      hybrid dilated attention mechanism      temporally-extended greedy method      policy optimization     
Received: 17 August 2025      Published: 20 July 2026
CLC:  TP 181  
Fund:  浙江省自然科学基金联合基金资助项目(LHZSD24F020001);国家自然科学基金资助项目(62306275, U20B2066);浙江大学上海高等研究院繁星科学基金资助项目(SN-ZJU-SIAS-001);中央高校基本科研业务费专项资金资助项目(226-2023-00048).
Corresponding Authors: Yuanyu WAN     E-mail: slili516@zjsru.edu.cn;wanyy@zju.edu.cn
Cite this article:

Li SONG,Yuanyu WAN,Mingli SONG. Graph convolutional reinforcement learning via hybrid dilated attention mechanism. Journal of ZheJiang University (Engineering Science), 2026, 60(9): 1980-1990.

URL:

https://www.zjujournals.com/eng/10.3785/j.issn.1008-973X.2026.09.015     OR     https://www.zjujournals.com/eng/Y2026/V60/I9/1980


基于混合空洞注意力机制的图卷积强化学习

针对复杂高动态多智能体强化学习中多样化动态信息捕捉困难、计算复杂的难题,提出基于混合空洞注意力的多智能体图卷积强化学习算法,以平衡计算复杂度和感受野大小,提高学习效率. 获取能够捕捉多智能体动态交互的图动态矩阵. 混合空洞注意力机制利用Bagging融合了多头注意力和多尺度空洞注意力机制的优点,使模型有效地学习不同类型特征. 在损失函数中添加正则化项以促进模型稳定性并解决过拟合问题. 为了平衡动作的探索与利用,采用改进的暂时拓展贪心方法选择多智能体动作. 仿真结果表明,在复杂动态环境中,所提方法的策略优化准确性和稳定性优于现有方法.


关键词: 多智能体,  图强化学习,  混合空洞注意力机制,  暂时拓展贪心方法,  策略优化 
Fig.1 Framework of multi-agent reinforcement learning with hybrid dilated attention
Fig.2 Surviving environment
Fig.3 Comparison of scores between proposed algorithms and benchmark algorithms
Fig.4 Scores of proposed algorithm under different graph regularization coefficients
Fig.5 Results of proposed algorithm and benchmark algorithm in surviving environment
算法损失值KL-lossattenq成功率
传统GCRL6.15×10163.99×1016–3.93×10158.800.68
GCRL-MHA7.54×10135.02×1016–4.87×10159.210.71
GCRL-MSDA6.82×10104.58×1016–4.89×101511.590.89
GCRL-EDA6.88×1074.54×1016–4.93×101512.020.92
Tab.1 Experimental results when proposed algorithm and benchmark algorithms converge
Fig.6 Results of ablation experiments with different policies and graph regularization
Fig.7 Variation graphs of dynamic nodes in learning processes
轮次scores
GCRL-MHAGCRL-MSDAGCRL-EDA
169.56118.05130.85
1076.1079.9454.12
2020.4525.3621.26
3024.1322.8722.41
Tab.2 Scores of GCRL-MHA, GCRL-MSDA, GCRL-EDA
Fig.8 Results of proposed algorithm and benchmark algorithms in highway environments
Fig.9 q value and learning rate of proposed algorithm and benchmark algorithm in Intersection-v0
Fig.10 Learning rate, q value, GPU temperature and utilization of proposed algorithm in highway-fast-v0 environment
Fig.11 Partial interaction graph network between urban vehicles at intersection
[1]   张萌, 王殿海, 金盛 结合领域经验的深度强化学习信号控制方法[J]. 浙江大学学报: 工学版, 2023, 57 (12): 2524- 2532
ZHANG Meng, WANG Dianhai, JIN Sheng Deep reinforcement learning approach to signal control combined with domain experience[J]. Journal of Zhejiang University: Engineering Science, 2023, 57 (12): 2524- 2532
[2]   CHU K F, LAM A Y, LI V O Traffic signal control using end-to-end off-policy deep reinforcement learning[J]. IEEE Transactions on Intelligent Transportation Systems, 2022, 23 (7): 7184- 7195
doi: 10.1109/TITS.2021.3067057
[3]   郭方洪, 伍泽芃, 杨淏, 等 基于个性化联邦强化学习的异构多微网能量调度[J]. 自动化学报, 2025, 51 (9): 2072- 2084
GUO Fanghong, WU Zepeng, YANG Hao, et al Energy scheduling of heterogeneous multi-microgrid based on personalized federated reinforcement learning[J]. Acta Automatica Sinica, 2025, 51 (9): 2072- 2084
doi: 10.16383/j.aas.c250130
[4]   GAUTIER P, LAURENT J, DIGUET J P Deep Q-learning-based dynamic management of a robotic cluster[J]. IEEE Transactions on Automation Science and Engineering, 2023, 20 (4): 2503- 2515
doi: 10.1109/TASE.2022.3205651
[5]   SHI D, LI L, OHTSUKI T, et al Make smart decisions faster: deciding D2D resource allocation via stackelberg game guided multi-agent deep reinforcement learning[J]. IEEE Transactions on Mobile Computing, 2022, 21 (12): 4426- 4438
doi: 10.1109/TMC.2021.3085206
[6]   MUNIKOTI S, AGARWAL D, DAS L, et al Challenges and opportunities in deep reinforcement learning with graph neural networks: a comprehensive review of algorithms and applications[J]. IEEE Transactions on Neural Networks and Learning Systems, 2024, 35 (11): 15051- 15071
doi: 10.1109/tnnls.2023.3283523
[7]   荣垂田, 田浩辉, 杜方 多目标深度强化学习驱动的数据库系统参数优化技术[J]. 软件学报, 2025, 36 (12): 5512- 5536
RONG Chuitian, TIAN Haohui, DU Fang Technique for database system parameter optimization using multi-objective deep reinforcement learning[J]. Journal of Software, 2025, 36 (12): 5512- 5536
doi: 10.13328/j.cnki.jos.007405
[8]   ZHAO X Y, WU C. Large-scale machine learning cluster scheduling via multi-agent graph reinforcement learning [C]// Thirty-Second AAAI Conference on Artificial Intelligence. Menlo Park: AAAI, 2022: 25082515.
[9]   PU Z Q, WANG H M, LIU Z, et al Attention enhanced reinforcement learning for multi agent cooperation[J]. IEEE Transactions on Neural Networks and Learning Systems, 2023, 34 (11): 8235- 8249
doi: 10.1109/TNNLS.2022.3146858
[10]   李岳珩, 谢广明 集中训练分布执行下的多智能体强化学习综述[J]. 控制理论与应用, 2025, 42 (11): 2114- 2124
LI Yueheng, XIE Guangming Review of multi-agent reinforcement learning under centralized training with decentralized execution[J]. Control Theory and Applications, 2025, 42 (11): 2114- 2124
doi: 10.7641/CTA.2025.50009
[11]   XING Q, XU Y, CHEN Z, et al A graph reinforcement learning-based decision-making platform for real-time charging navigation of urban electric vehicles[J]. IEEE Transactions on Industrial Informatics, 2023, 19 (3): 3284- 3295
doi: 10.1109/TII.2022.3210264
[12]   HOUIDI O, BAKRI S, ZEGHLACHE D. Multi-agent graph convolutional reinforcement learning for intelligent load balancing [C]// NOMS 2022-2022 IEEE/IFIP Network Operations and Management Symposium. Piscataway: IEEE, 2022: 1-6.
[13]   LIN W Y, SONG Y Z, RUAN B K, et al Temporal difference-aware graph convolutional reinforcement learning for multi-intersection traffic signal control[J]. IEEE Transactions on Intelligent Transportation Systems, 2023, 25 (1): 327- 337
[14]   LUO J, LI C F, FAN Q Q, et al A graph convolutional encoder and multi-head attention decoder network for tsp via reinforcement learning[J]. Engineering Applications of Artificial Intelligence, 2022, 112: 104848
doi: 10.1016/j.engappai.2022.104848
[15]   DU X Q, CHEN H C, XING Y H, et al A contrastive enhanced ensemble framework for efficient multi-agent reinforcement learning[J]. Expert Systems with Applications, 2024, 245: 123158
doi: 10.1016/j.eswa.2024.123158
[16]   HU Y F, FU J J, WEN G H Graph soft actor-critic reinforcement learning for large-scale distributed multirobot coordination[J]. IEEE Transactions on Neural Networks and Learning Systems, 2025, 36 (1): 665- 676
doi: 10.1109/TNNLS.2023.3329530
[17]   XIE S R, ZHANG H, YU H, et al ET-HF: a novel information sharing model to improve multi-agent cooperation[J]. Knowledge-based Systems, 2022, 257: 109916
doi: 10.1016/j.knosys.2022.109916
[18]   罗佳, 李朝锋 基于残差图卷积网络与深度强化学习的需求可拆分车辆路径优化算法[J]. 控制理论与应用, 2024, 41 (6): 1123- 1136
LUO Jia, LI Chaofeng The split delivery vehicle routing optimization with the residual graph convolutional network and deep reinforcement learning[J]. Control Theory and Applications, 2024, 41 (6): 1123- 1136
doi: 10.7641/CTA.2023.21040
[19]   GOECKNER A, SUI Y Y, MARTINET N, et al. Graph neural network-based reinforcement learning for multi-agent systems [C]// Thirty-Sixth Conference on Neural Information Processing Systems. Cambridge: MIT Press, 2022: 1-8.
[20]   MA X, XIE Y, CHIGAN C Graph convolutional network based multi-objective meta-deep Q-learning for eco-routing[J]. IEEE Transactions on Intelligent Transportation Systems, 2024, 25 (7): 7323- 7338
doi: 10.1109/TITS.2023.3348034
[21]   ELLIS J D, IQBAL R, YOSHIMATSU K Deep Q-learning-based molecular graph generation for chemical structure prediction from infrared spectra[J]. IEEE Transactions on Artificial Intelligence, 2024, 5 (2): 634- 646
doi: 10.1109/TAI.2023.3287947
[22]   ROKHFOROZ P, MONTAZERI M, FINK O Multi-agent reinforcement learning with graph convolutional neural networks for optimal bidding strategies of generation units in electricity markets[J]. Expert Systems with Applications, 2023, 225: 120010
doi: 10.1016/j.eswa.2023.120010
[23]   XIA L Q, LIANG Y S, LENG J W, et al Maintenance planning recommendation of complex industrial equipment based on knowledge graph and graph neural network[J]. Reliability Engineering and System Safety, 2023, 232: 109068
doi: 10.1016/j.ress.2022.109068
[24]   YE Y, JI S H Sparse graph attention networks[J]. IEEE Transactions on Knowledge and Data Engineering, 2023, 35 (1): 905- 916
doi: 10.1109/TKDE.2021.3072345
[25]   FAWAZ H, LESCA J, QUANG P T, et al Graph convolutional reinforcement learning for collaborative queuing agents[J]. IEEE Transactions on Network and Service Management, 2023, 20 (2): 13631377
doi: 10.1109/tnsm.2022.3226605
[26]   JIANG J C, DUN C, HUANG T J, et al. Graph convolutional reinforcement learning [C]// International Conference on Learning Representations. Washington DC: [s.n.], 2020: 1-13.
[1] Yuanfeng LIAN,Shuyu FAN,Benzhe ZHANG,Sen WANG. Bio-inspired gait control integrating geometric constraint feature extraction and multi-reward cooperative optimization[J]. Journal of ZheJiang University (Engineering Science), 2026, 60(9): 1862-1871.
[2] Ningbo CAO,Qichao WAN,Liying ZHAO,Zimeng LI,Baolin HUANG. Collaborative control of mixed traffic intersections integrating multi-agent reinforcement learning and maximum pressure control[J]. Journal of ZheJiang University (Engineering Science), 2026, 60(8): 1819-1831.
[3] Lang CHEN,Zengli LIU,Xuanzhi ZHAO. Decision-making for multi-USV adversarial encirclement based on adaptive curriculum reinforcement learning[J]. Journal of ZheJiang University (Engineering Science), 2026, 60(7): 1369-1380.
[4] Yang WANG,Hongchao LIU,Chi TIAN,Bing WU,Di ZHANG. Multi-ship collision avoidance via route exchange mechanism: strategy learning and game-theoretic decision making[J]. Journal of ZheJiang University (Engineering Science), 2026, 60(5): 964-976.
[5] Jing ZHANG,Yi WANG,Zilong CHEN,Yunsong LI. Multi-goal multi-agent path finding algorithm[J]. Journal of ZheJiang University (Engineering Science), 2025, 59(8): 1689-1697.
[6] Wenhao LI,Yanjie JI,Hao WU,Yewen JIA,Shuichao ZHANG. Empty-load charging strategy for autonomous vehicle parking based on multi-agent system[J]. Journal of ZheJiang University (Engineering Science), 2024, 58(8): 1659-1670.
[7] Zhuo WANG,Yongqiang LI,Yu FENG,Yuanjing FENG. Policy gradient algorithm and its convergence analysis for two-player zero-sum Markov games[J]. Journal of ZheJiang University (Engineering Science), 2024, 58(3): 480-491.
[8] Yong LI,Fu-qiang LIU,Bai-qing SUN,Qiu-hao ZHANG,Jun-you YANG. Dynamic multi-task allocation of heterogeneous multi-robots for daily elderly care scenarios[J]. Journal of ZheJiang University (Engineering Science), 2022, 56(9): 1806-1814.
[9] Xiao-gao XU,Ying-jie XIA,Si-yu ZHU,Li KUANG. Cooperative control algorithm of multi-intersection variable-direction lanes based on reinforcement learning[J]. Journal of ZheJiang University (Engineering Science), 2022, 56(5): 987-994, 1005.
[10] Pan ZHANG,Hua DING,Ying-er ZHANG,Bing-ning LI,Jiang-tao HUANG-FU,Zhong-he JIN. Multi-agent autonomous electronic jamming system based on information sharing[J]. Journal of ZheJiang University (Engineering Science), 2022, 56(1): 75-83.
[11] Hang-lei SHAO,Dong-mei ZHANG. Synchronization of multi-agent systems based on static output feedback protocol[J]. Journal of ZheJiang University (Engineering Science), 2020, 54(7): 1308-1315.
[12] Xia-sheng SHI,Rong-hao ZHENG,Zhi-yun LIN,Gang-feng YAN. Saddle dynamic based distributed algorithm for economic dispatch problem[J]. Journal of ZheJiang University (Engineering Science), 2020, 54(4): 678-683.
[13] AI Jie-qing, GAO Ji, PENG Yan-bin, ZHENG Zhi-jun. Negotiation decision model based on transductive
support vector machine
[J]. Journal of ZheJiang University (Engineering Science), 2012, 46(6): 967-973.
[14] YANG Hong-Yong, LI Xiao. Formation control of multiagent with diverse delays[J]. Journal of ZheJiang University (Engineering Science), 2010, 44(7): 1355-1360.
[15] OU Li-Yong, DU Shu-Xin. Multi-agent based coordination of public detection resources[J]. Journal of ZheJiang University (Engineering Science), 2010, 44(1): 81-86+123.