Please wait a minute...
浙江大学学报(工学版)  2026, Vol. 60 Issue (8): 1819-1831    DOI: 10.3785/j.issn.1008-973X.2026.08.021
交通工程     
集成多智能体强化学习和最大压强控制的混行交叉口协同控制
曹宁博1(),万启超1,赵利英2,*(),李梓萌1,黄宝林1
1. 长安大学 运输工程学院,陕西 西安 710061
2. 西安理工大学 经济与管理学院,陕西 西安 710048
Collaborative control of mixed traffic intersections integrating multi-agent reinforcement learning and maximum pressure control
Ningbo CAO1(),Qichao WAN1,Liying ZHAO2,*(),Zimeng LI1,Baolin HUANG1
1. College of Transportation Engineering, Chang’an University, Xi’an 710061, China
2. School of Economics and Management, Xi’an University of Technology, Xi’an 710048, China
 全文: PDF(1505 KB)   HTML
摘要:

针对混行交叉口包含网联自动驾驶车(CAVs)、人驾车辆(HDVs)及行人的交通控制问题,提出融合多智能体近端策略优化(MAPPO)与最大压强控制(MPC)的协同控制方法. 通过构建分布式部分可观测马尔可夫决策过程(Dec-POMDP),设计分层状态空间与动作空间,并引入能够同时考虑安全性、通行效率、相位切换频率及规则依从性的奖励函数. 采用集中式训练与分布式执行框架,基于SUMO仿真平台在低、中、高交通流量场景下对模型性能进行验证. 结果表明,所提出的MAPPO-MPC方法能够显著提升吞吐量(最高提升44.3%),降低排队长度(最高降低49.3%)、车辆平均延误(最高降低43.2%)和行人等待时间(最高降低43.53%),且在CAV渗透率超过50%时表现出更显著的性能优势,优于传统Webster方法以及基线模型.

关键词: 交通运输工程混行交叉口多智能体强化学习最大压强控制    
Abstract:

A collaborative control approach integrating multi-agent proximal policy optimization (MAPPO) and maximum pressure control (MPC) was developed to address traffic control challenges at mixed traffic intersections involving connected autonomous vehicles (CAVs), human-driven vehicles (HDVs), and pedestrians. A hierarchical state space and an action space were designed through the formulation of a decentralized partially observable Markov decision process (Dec-POMDP). Reward functions considering the balance among safety, traffic efficiency, phase switching frequency, and regulatory compliance were introduced. The proposed model was validated under low, medium, and high traffic flow conditions using a centralized training with decentralized execution framework on the SUMO simulation platform. Experimental results demonstrated that the proposed MAPPO-MPC method significantly improved the throughput (by up to 44.3%) while reducing the queue length (by up to 49.3%), average vehicle delay (by up to 43.2%), and pedestrian waiting time (by up to 43.53%). Moreover, the model exhibited more substantial performance advantages when the CAV penetration rate exceeded 50%, outperforming the traditional Webster-based methods and the baseline models.

Key words: transportation engineering    mixed traffic intersection    multi-agent reinforcement learning    maximum pressure control
收稿日期: 2025-07-14 出版日期: 2026-07-16
CLC:  U 492.3  
基金资助: 中央高校基本科研业务资助项目(300102345602);陕西省自然科学基础研究计划资助项目(2024JC-YBMS-376);陕西省社会科学基金资助项目(2021R025);陕西省教育厅科学研究计划资助项目(23JK0557);陕西省社会科学基金资助项目(2022R028).
通讯作者: 赵利英     E-mail: 819868226@qq.com;lyzhao@xaut.edu.cn
作者简介: 曹宁博(1987—),男,讲师,从事自动驾驶汽车和行人安全研究. orcid.org/0000-0002-6630-0466. E-mail:819868226@qq.com
服务  
把本文推荐给朋友
加入引用管理器
E-mail Alert
作者相关文章  
曹宁博
万启超
赵利英
李梓萌
黄宝林

引用本文:

曹宁博,万启超,赵利英,李梓萌,黄宝林. 集成多智能体强化学习和最大压强控制的混行交叉口协同控制[J]. 浙江大学学报(工学版), 2026, 60(8): 1819-1831.

Ningbo CAO,Qichao WAN,Liying ZHAO,Zimeng LI,Baolin HUANG. Collaborative control of mixed traffic intersections integrating multi-agent reinforcement learning and maximum pressure control. Journal of ZheJiang University (Engineering Science), 2026, 60(8): 1819-1831.

链接本文:

https://www.zjujournals.com/eng/CN/10.3785/j.issn.1008-973X.2026.08.021        https://www.zjujournals.com/eng/CN/Y2026/V60/I8/1819

图 1  混行交叉口示意图
图 2  多智能体强化学习过程
图 3  MAPPO算法框架
图 4  MAPPO算法中第$ i' $个智能体的训练回合
图 5  混行交叉口管理方法框架图
智能体动作控制参数描述
CAVs0~20目标速度[0,2,···,40] m/s离散速度档位动态调整加速度
TL0相位保持维持或调整当前绿灯相位时长
1相位顺序切换按预设顺序切换至下一相位
2相位优先切换激活最大压强车道的绿灯相位
HDVs0~4IDM计算的加速度车速保持、加速、减速等动态调整
表 1  动作空间参数表
图 6  混行信号控制交叉口仿真场景
场景周期/st/s
东西直行东西左转南北直行南北左转
低流量7222 +424 +438 +440 +4
中流量10536 +439 +444 +447 +4
高流量13042 +445 +450 +453 +4
表 2  交叉口基础信号配时
流量密度EE
(veh·h?1)
SE
(veh·h?1)
WE
(veh·h?1)
NE
(veh·h?1)
NP
(per·h?1)
360650380575200
81011408901120400
995153010801580600
表 3  交叉口交通流量
参数数值参数数值参数数值参数数值
$ {w_{\mathrm{s}}} $0.5$ {k_{{\text{acc}}}} $/(s2·m?1)0.1$ {k_{{\text{flow}}}} $0.1$ {C_{{\text{re}}{{\text{d}}^\prime }}} $40
$ {C_{{\text{col}}}} $?1 000$ {a_{{\text{th}}}} $/(m·s?2)2$ {w_{{\mathrm{s}}'}} $0.4$ {k_{{\text{gree}}{{\text{n}}^\prime }}} $/(s·m?1)0.04
$ {k_{{\text{dist}}}} $/m?210$ {k_{{\text{jerk}}}} $/(s4·m?1)0.2$ {C_{{\text{co}}{{\text{l}}^\prime }}} $?1 000$ {k_{{\text{sudde}}{{\text{n}}^\prime }}} $/(s2·m?1)0.05
$ {d_{{\text{min}}}} $/m1$ {k_\delta } $0.05$ {k_{{\text{dis}}{{\text{t}}^\prime }}} $/m?25$ {w_{\mathrm{h}}} $0.15
$ \epsilon $/m0.1$ {\delta _{\max }} $/radπ/4$ {d_{{{\min }^\prime }}} $/m1$ \lambda $0.3
$ {k_{{\text{ped}}}} $/m?320$ {w_{\mathrm{p}}} $0.05$ {d_{{\text{saf}}{{\text{e}}^\prime }}} $/m1.5$ {k_{{\text{run - yellow}}}} $1
$ {d_{{\text{ped\_safe}}}} $/m3$ {C_{{\text{ped}}}} $?500$ {w_{{\mathrm{e}}'}} $0.3$ {k_{{\text{stop - green}}}} $0.5
$ {d_{{\text{ped\_min}}}} $/m0.5$ {k_{{\text{yield}}}} $5$ {k_{{\mathrm{v}}'}} $/(s·m?1)0.12$ {k_{{\text{honk - ped}}}} $0.3
$ {t_{{\text{react}}}} $/s1.5$ {k_{{\text{honk}}}} $1$ {v_{{\text{pref}}}} $/(m·s?1)限速×1.15$ {k_{{\text{distract}}}} $0.1
$ {d_0} $/m2$ {k_{{\text{pred}}}} $/m?20.5$ {k_{{\text{pro}}{{\text{g}}^\prime }}} $/m?10.01$ {w_{\Pr }} $1
$ {w_{\mathrm{e}}} $0.3$ {d_{{\text{cross\_safe}}}} $/m10$ {k_{{\text{delay}}}} $/s?10.02$ {w_{{\text{phase - change}}}} $0.5
$ {k_{\mathrm{v}}} $/(s·m?1)0.1$ {w_{\mathrm{l}}} $0.05$ {k_{{\text{detour}}}} $/m?10.001$ {w_{{\text{throughput}}}} $1.2
$ {v_{{\text{target}}}} $/(m·s?1)限速$ {C_{{\text{red}}}} $50$ {w_{{\mathrm{c}}'}} $0.05$ {w_{{\text{emergency}}}} $0.3
$ {k_{{\text{prog}}}} $/m?10.01$ {v_{\max }} $/(m·s?1)40$ k_{{\text{acc}}}^{'} $/(s2·m?1)0.05$ {w_{{\text{queue}}}} $0.8
$ {k_{{\text{idle}}}} $/s?10.05$ {k_{{\text{green}}}} $/(s·m?1)0.05$ {a_{{\text{t}}{{\text{h}}^\prime }}} $/(m·s?2)3$ {w_{{\text{wait}}}} $0.3
$ {v_{\min }} $/(m·s?1)2$ {v_{{\text{lim}}}} $/(m·s?1)限速$ k_{{\text{jerk}}}^{'} $/(s4·m?1)0.1$ w_{{\text{phase - change}}}^{'} $0.4
$ {k_{{\text{route}}}} $/ m?10.001$ {k_{{\text{yellow}}}} $/(s2·m?1)0.1$ {k_{\delta '}} $0.03$ \gamma $0.99
$ {w_{\mathrm{r}}} $0.1$ {k_{{\text{antic}}}} $0.2$ {w_{{\mathrm{p}}'}} $0.05学习率0.000 01
$ {k_{{\text{lane}}}} $/ m?10.1$ {w_{{\mathrm{co}}}} $0.05$ {C_{{\text{pe}}{{\text{d}}^\prime }}} $?500总训练回合数1 500
$ {k_{{\text{signal}}}} $0.5$ {k_{{\text{v2x}}}} $2$ {k_{{\text{hon}}{{\text{k}}^\prime }}} $0.5$ \epsilon $0.2
$ {k_{{\text{illegal}}}} $1$ {k_{{\text{cross}}}} $1$ {k_{{\text{ped}}\_{\text{avoid}}}} $2SGD迭代次数5
$ {w_{\mathrm{c}}} $0.05$ {k_{{\text{block}}}} $1$ {w_{{\mathrm{l}}'}} $0.05$ B $100
表 4  训练参数表
图 7  不同算法训练回合平均奖励值
图 8  模型成功率、安全率、效率和平均奖励值对比图
图 9  交叉口车辆排队、延误和吞吐量对比图
渗透率流量t/s
优化前MAPPO-
MPC
MADDPG-
MPC
MAPPOMADDPG
10%45.039.841.342.843.5
55.053.553.854.254.5
65.063.864.064.364.5
30%41.832.535.237.838.5
54.752.252.853.553.9
64.562.162.863.563.8
50%38.424.330.333.435.0
54.551.552.253.053.5
63.559.460.862.062.5
70%35.120.726.629.731.5
53.848.249.851.252.0
60.953.855.557.858.8
90%31.717.923.626.328.0
52.543.845.848.249.5
56.748.350.553.054.2
表 5  行人等待时间表
1 WU J, HUANG Z, LV C Uncertainty-aware model-based reinforcement learning: methodology and application in autonomous driving[J]. IEEE Transactions on Intelligent Vehicles, 2022, 8 (1): 194- 203
2 KIRAN B R, SOBH I, TALPAERT V, et al Deep reinforcement learning for autonomous driving: a survey[J]. IEEE Transactions on Intelligent Transportation Systems, 2021, 23 (6): 1- 18
3 ZHAO Z, XUN J, WEN X, et al. Safe reinforcement learning for single train trajectory optimization via shield SARSA [J]. IEEE Transactions on Intelligent Transportation Systems, 24(1), 412–428.
4 罗彪, 胡天萌, 周育豪, 等 多智能体强化学习控制与决策研究综述[J]. 自动化学报, 2025, 51 (3): 510- 539
LUO Biao, HU Tianmeng, ZHOU Yuhao, et al Survey on multi-agent reinforcement learning for control and decision-making[J]. Acta Automat Sin, 2025, 51 (3): 510- 539
doi: 10.16383/j.aas.c240392
5 PENG B, KESKIN M F, KULCSÁR B, et al Connected autonomous vehicles for improving mixed traffic efficiency in unsignalized intersections with deep reinforcement learning[J]. Communications in Transportation Research, 2021, 1 (1): 5- 17
6 许曼晨, 于镝, 赵理, 等 基于MAPPO的无信号灯交叉口自动驾驶决策[J]. 吉林大学学报: 信息科学版, 2024, 42 (5): 790- 798
XU Manchen, YU Di, ZHAO Li, et al Autonomous driving decision-making at signal-free intersections based on MAPPO[J]. Journal of Jilin University: Information Science Edition, 2024, 42 (5): 790- 798
doi: 10.3969/j.issn.1671-5896.2024.05.003
7 WEI H, CHEN C, ZHENG G, et al. Presslight: learning max pressure control to coordinate traffic signals in arterial network [C]// 2019 ACM 25th SIGKDD International Conference on Knowledge Discovery and Data Mining. Anchorage: ACM, 2019: 1290–1298.
8 BDEIR A, BOEDER S, DERNEDDE T, et al. RP-DQN: An application of Q-learning to vehicle routing problems [C]// German Conference on Artificial Intelligence. Cham: Springer, 2021: 3–16.
9 SHARIF A, MARIJAN D. Evaluating the robustness of deep reinforcement learning for autonomous policies in a multi-agent urban driving environment [C]// 2022 IEEE 22th International Conference on Software Quality, Reliability and Security. Guangzhou: IEEE, 2022: 785–796.
10 YU C, WANG X, XU X, et al Distributed multiagent coordinated learning for autonomous driving in highways based on dynamic coordination graphs[J]. IEEE Transactions on Intelligent Transportation Systems, 2020, 21 (2): 735- 748
doi: 10.1109/TITS.2019.2893683
11 孙彧, 曹雷, 陈希亮, 等 多智能体深度强化学习研究综述[J]. 计算机工程与应用, 2020, 56 (5): 13- 24
SUN Yu, CAO Lei, CHEN Xiliang, et al Overview of multi-agent deep reinforcement learning[J]. Computer Engineering and Applications, 2020, 56 (5): 13- 24
12 ZHANG X, LI S, WANG B, et al. Multi-vehicle collaborative lane changing based on multi-agent reinforcement learning [C]// 2024 IEEE Intelligent Vehicles Symposium. Jeju Island: IEEE, 2024: 1214–1221.
13 ZHOU W, CHEN D, YAN J, et al Multi-agent reinforcement learning for cooperative lane changing of connected and autonomous vehicles in mixed traffic[J]. Autonomous Intelligent Systems, 2022, 2 (1): 1- 11
doi: 10.1007/s43684-021-00019-7
14 MATE K, BECSI T Multi-agent reinforcement learning for highway platooning[J]. Electronics, 2023, 12 (24): 1–13
15 SCHESTER L, ORTIZ L E. A systematic study of multi-agent deep reinforcement learning for safe and robust autonomous highway ramp entry [EB/OL]. [2025–01–17]. https://arxiv.org/pdf/2411.14593.
16 SCHULMAN J , WOLSKI F, Dhariwal P, et al. Proximal policy optimization algorithms [EB/OL]. [2025–07–10]. https://arxiv.org/pdf/1707.06347.
17 YU C, VELU A, VINITSKY E, et al The surprising effectiveness of PPO in cooperative multi-agent games[J]. Advances in Neural Information Processing Systems, 2021, 35: 24611- 24624
18 KRAEMER L, BANERJEE B Multi-agent reinforcement learning as a rehearsal for decentralized planning[J]. Neurocomputing, 2016, 190: 82- 94
doi: 10.1016/j.neucom.2016.01.031
[1] 陈浪,刘增力,赵宣植. 基于自适应课程强化学习的多无人艇对抗围捕决策[J]. 浙江大学学报(工学版), 2026, 60(7): 1369-1380.
[2] 汪洋,刘红超,田池,吴兵,张笛. 航线交换机制下多船避碰的策略学习与博弈决策[J]. 浙江大学学报(工学版), 2026, 60(5): 964-976.