Please wait a minute...
浙江大学学报(工学版)  2026, Vol. 60 Issue (10): 2141-2152    DOI: 10.3785/j.issn.1008-973X.2026.10.007
计算机技术与控制工程     
基于动态核感知的无人机视角路面病害检测方法
张建刚1(),李肖1,冯丹丹2
1. 兰州交通大学 数理学院,甘肃 兰州 730070
2. 兰州交通大学 电子与信息工程学院,甘肃 兰州 730070
Dynamic kernel perception for pavement distress detection in UAV inspection
Jiangang ZHANG1(),Xiao LI1,Dandan FENG2
1. School of Mathematics and Physics, Lanzhou Jiaotong University, Lanzhou 730070, China
2. School of Electronic and Information Engineering, Lanzhou Jiaotong University, Lanzhou 730070, China
 全文: PDF(3699 KB)   HTML
摘要:

针对无人机视角路面病害检测中上下文信息丢失、多尺度病害共存以及细节还原不足等问题,构建基于动态核感知的路面病害检测模型,DKP-YOLO. 为了缓解复杂场景裂缝形态多变与上下文信息易丢失的问题,提出轻量化裂缝上下文模块,该模块通过异构并行深度可分离卷积与动态特征融合机制增强特征区分度. 针对复杂背景与多尺度病害的挑战,设计路面病害感知模块,利用并行多尺度卷积与双重注意力机制提升多尺度病害感知能力. 为了克服传统卷积难以建模长程依赖和对微小目标响应不足的局限,提出大核感知特征融合模块,通过整合超大感受野卷积与通道-空间注意力捕捉全局上下文关系. 为了减少上采样过程中细节丢失与语义模糊,在颈部网络引入轻量级动态上采样算子. 在UAV-PDD2023数据集上的实验表明,本研究模型在显著提升检测精度的同时,实现了模型复杂度的有效降低. 与基线模型相比,本研究模型在轻量化和检测性能上取得更好的平衡,更适合无人机视角路面病害目标检测.

关键词: 路面病害检测无人机航拍图像动态核感知多尺度特征融合复杂背景抑制    
Abstract:

A detection model named dynamic kernel perception YOLO (DKP-YOLO) was constructed to address the problems of contextual information loss, multi-scale distress coexistence, and insufficient detail restoration in pavement distress detection from a UAV perspective. A lightweight crack context module was proposed to mitigate the variability of crack morphology and the loss of contextual information in complex scenes. This module enhanced feature discriminability through heterogeneous parallel depthwise separable convolutions and a dynamic feature fusion mechanism. In order to tackle the challenges of complex backgrounds and multi-scale distress, a pavement distress perception module was designed, which improved the multi-scale distress perception capability using parallel multi-scale convolutions and a dual-attention mechanism. A large-kernel perception feature fusion module was introduced to overcome the limitations of traditional convolutions in modeling long-range dependencies and responding to small targets. This module captured global contextual relationships by integrating extra-large receptive field convolution with channel-spatial attention. Finally, a lightweight dynamic upsampling operator was incorporated into the neck network to reduce detail loss and semantic ambiguity during upsampling. Experiments on the UAV-PDD2023 dataset showed that the proposed model significantly improved detection accuracy while effectively reducing model complexity. Compared with the baseline model, DKP-YOLO achieved a better balance between being lightweight and maintaining high detection performance, making it more suitable for UAV-based pavement distress detection.

Key words: pavement distress detection    UAV aerial imagery    dynamic kernel perception    multi-scale feature fusion    complex background suppression
收稿日期: 2025-12-09 出版日期: 2026-07-28
CLC:  U 418.6  
基金资助: 甘肃省重点研发计划资助项目(25YFGA047);甘肃省基础研究创新群体资助项目(25JRRA805);鄂尔多斯市重点研发计划资助项目(YF20250245).
作者简介: 张建刚(1978—),男,教授,博士,从事非线性动力系统分析及控制、随机动力系统稳定性分析及控制等研究. orcid.org/0000-0002-8523-6623. E-mail:zhangjg7715776@126.com
服务  
把本文推荐给朋友
加入引用管理器
E-mail Alert
作者相关文章  
张建刚
李肖
冯丹丹

引用本文:

张建刚,李肖,冯丹丹. 基于动态核感知的无人机视角路面病害检测方法[J]. 浙江大学学报(工学版), 2026, 60(10): 2141-2152.

Jiangang ZHANG,Xiao LI,Dandan FENG. Dynamic kernel perception for pavement distress detection in UAV inspection. Journal of ZheJiang University (Engineering Science), 2026, 60(10): 2141-2152.

链接本文:

https://www.zjujournals.com/eng/CN/10.3785/j.issn.1008-973X.2026.10.007        https://www.zjujournals.com/eng/CN/Y2026/V60/I10/2141

图 1  DKP-YOLO结构图
图 2  通道注意力示意图
图 3  空间注意力、小目标特异性注意力示意图
图 4  动态上采样算子示意图
参数设置值
优化器AdamW
训练轮次300
初始学习率1×10?3
输入图像尺寸640×640
动量因子0.937
早停轮次50
表 1  实验参数配置
图 5  改进前、后模型混淆矩阵
病害类别P1/%P2/%R1/%R2/%mAP@0.51/%mAP@0.52/%mAP@0.5:0.951/%mAP@0.5:0.952/%
注:X1表示YOLO12-n相关指标,X2表示DKP-YOLO相关指标
所有类别78.090.4 (+12.4)71.780.4 (+8.7)76.084.9 (+8.9)48.361.9 (+13.6)
网状裂缝82.793.5 (+10.8)71.491.6 (+20.2)79.693.7 (+14.1)51.568.4 (+16.9)
纵向裂缝81.795.0 (+13.3)74.581.4 (+6.9)79.489.9 (+10.5)48.262.9 (+14.7)
斜向裂缝72.592.9 (+20.4)62.974.3 (+11.4)66.482.6 (+16.2)39.656.3 (+16.7)
坑洞74.385.4 (+11.1)60.058.5 (?1.5)57.560.1 (+2.6)33.046.0 (+13.0)
修补72.680.3 (+7.7)82.192.9 (+10.8)88.392.8 (+4.5)63.572.3 (+8.8)
横向裂缝84.695.4 (+10.8)79.583.8 (+4.3)84.990.7 (+5.8)54.165.6 (+11.5)
表 2  改进前、后模型各指标数据对比
图 6  改进前、后模型训练损失与评估指标曲线对比
模型P/%R/%mAP@0.5/%mAP@0.5:
0.95/%
F1/%FPS/(帧·s?1)
1)注:加粗字体表示每列中的最优值
YOLO12n
(基线模型)
78.071.776.048.374.7212.7
+A81.769.775.948.575.2217.5
+B91.274.782.155.482.1197.4
+C89.880.184.660.484.6236.1
+A+B89.476.582.355.182.4201.3
+A+C90.378.484.058.583.9239.7
+B+C${{\boldsymbol{92.3}}}^{1)} $78.385.060.784.7189.8
+A+B+C91.980.084.559.685.5197.2
+A+B+C+D90.480.484.961.985.1225.8
表 3  各改进模块的消融实验结果
模型P/%R/%mAP@0.5:0.95/%F1/%FPS/(帧·s?1)Params/106GFLOPs/109Weights/MB
1)注:加粗字体表示每列中的最优值
SSD69.254.442.560.940.032.3128.949.1
Faster-RCNN73.959.246.365.714.0136.8370.2108.2
RT-DETR70.356.434.162.677.04.2130.583.0
UAV-YOLO84.369.248.676.078.010.525.319.4
LFDS-YOLO90.475.753.082.495.04.420.28.7
YOLOv5-n82.171.450.076.3257.92.57.25.3
YOLOv6-n72.256.132.663.1253.24.211.98.4
YOLOv8-n90.374.154.381.4263.83.08.26.3
YOLOv9-s90.61)75.157.682.1196.07.327.415.3
YOLOv9-t82.266.044.473.2248.82.17.94.7
YOLOv10-n86.274.155.079.7261.12.78.45.6
YOLO11-n90.677.956.583.7256.32.66.45.3
YOLO12-n78.071.748.374.7212.72.66.25.3
YOLO13-n82.664.744.972.6191.52.56.45.4
DKP-YOLO90.480.461.985.1225.82.45.55.2
表 4  各模型检测性能评价指标
模型P/%R/%mAP@0.5/%mAP@0.5:0.95/%F1/%Params/MGFLOPs/109FPS/(帧·s?1)
YOLOv10-n59.957.359.738.058.62.78.4206.1
YOLO11-n64.160.463.238.962.22.66.4197.5
YOLO12-n66.156.762.036.261.02.66.2174.3
DKP-YOLO67.470.773.141.369.02.45.5191.8
表 5  泛化试验结果
图 7  UAV-PDD2023数据集上不同模型检测效果图
图 8  UAPD数据集上不同模型检测效果图
图 9  不同检测模型热力图对比
1 中华人民共和国交通运输部. 2024年交通运输行业发展统计公报[EB/OL]. (2025–06–12) [2025–07–09]. https://xxgk.mot.gov.cn/2020/jigou/zhghs/202506/t20250610_4170228.html.
2 《中国公路学报》编辑部 中国路面工程学术研究综述·2024[J]. 中国公路学报, 2024, 37 (3): 1- 49
Editorial Department of China Journal of Highway and Transport Review on China’s pavement engineering research: 2024[J]. China Journal of Highway and Transport, 2024, 37 (3): 1- 49
doi: 10.19815/j.jace.2025.04061
3 ALZAHRANI B, OUBBATI O S, BARNAWI A, et al UAV assistance paradigm: state-of-the-art in applications and challenges[J]. Journal of Network and Computer Applications, 2020, 166: 102706
doi: 10.1016/j.jnca.2020.102706
4 ZHANG Y, ZUO Z, XU X, et al Road damage detection using UAV images based on multi-level attention mechanism[J]. Automation in Construction, 2022, 144: 104613
doi: 10.1016/j.autcon.2022.104613
5 WANG F, ZOU Y, CHEN X, et al Rapid in-flight image quality check for UAV-enabled bridge inspection[J]. ISPRS Journal of Photogrammetry and Remote Sensing, 2024, 212: 230- 250
doi: 10.1016/j.isprsjprs.2024.05.008
6 MA X, LI Y, YANG Z, et al Lightweight network for millimeter-level concrete crack detection with dense feature connection and dual attention[J]. Journal of Building Engineering, 2024, 94: 109821
doi: 10.1016/j.jobe.2024.109821
7 杨豪, 刘李彦, 张军辉, 等 环境适应性优化的轻量化多尺度道路裂缝检测[J]. 中国公路学报, 2025, 38 (7): 118- 134
YANG Hao, LIU Liyan, ZHANG Junhui, et al Environmental adaptability optimization lightweight multi-scale road crack detection[J]. China Journal of Highway and Transport, 2025, 38 (7): 118- 134
8 MA X, DAI X, BAI Y, et al. Rewrite the stars [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle: IEEE, 2024: 5694–5703.
9 YANG F, ZHANG L, YU S, et al Feature pyramid and hierarchical boosting network for pavement crack detection[J]. IEEE Transactions on Intelligent Transportation Systems, 2020, 21 (4): 1525- 1535
doi: 10.1109/TITS.2019.2910595
10 LIU Z, LIN Y, CAO Y, et al. Swin transformer: hierarchical vision transformer using shifted windows [C]// Proceedings of the IEEE/CVF International Conference on Computer Vision. Montreal: IEEE, 2022: 9992–10002.
11 WANG L, YOON K J Knowledge distillation and student-teacher learning for visual intelligence: a review and new outlooks[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022, 44 (6): 3048- 3068
doi: 10.1109/TPAMI.2021.3055564
12 LIU H, MIAO X, MERTZ C, et al. CrackFormer: transformer network for fine-grained crack detection [C]// Proceedings of the IEEE/CVF International Conference on Computer Vision. Montreal: IEEE, 2022: 3763–3772.
13 DUNG C V, ANH L D Autonomous concrete crack detection using deep fully convolutional neural network[J]. Automation in Construction, 2019, 99: 52- 58
doi: 10.1016/j.autcon.2018.11.028
14 QU Z, CHEN W, WANG S Y, et al A crack detection algorithm for concrete pavement based on attention mechanism and multi-features fusion[J]. IEEE Transactions on Intelligent Transportation Systems, 2022, 23 (8): 11710- 11719
doi: 10.1109/TITS.2021.3106647
15 PAN Y, ZHANG G, ZHANG L A spatial-channel hierarchical deep learning network for pixel-level automated crack detection[J]. Automation in Construction, 2020, 119: 103357
doi: 10.1016/j.autcon.2020.103357
16 TAN M, PANG R, LE Q V. EfficientDet: scalable and efficient object detection [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle: IEEE, 2020: 10778–10787.
17 HOU Q, ZHOU D, FENG J. Coordinate attention for efficient mobile network design [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Nashville: IEEE, 2021: 13708–13717.
18 SHAN J, JIANG W, FENG X Bridging cross-domain and cross-resolution gaps for UAV-based pavement crack segmentation[J]. Automation in Construction, 2025, 174: 106141
doi: 10.1016/j.autcon.2025.106141
19 TIAN Y, YE Q, DOERMANN D. YOLOv12: attention-centric real-time object detectors [EB/OL]. [2025–02–18]. https://arxiv.org/abs/2502.12524.
20 ZHAO H, KONG X, HE J, et al. Efficient image super-resolution using pixel attention [C]// Computer Vision – ECCV 2020 Workshops. Cham: Springer, 2020: 56–72.
21 LAU K W, PO L M, REHMAN Y A U Large separable kernel attention: rethinking the large kernel attention design in CNN[J]. Expert Systems with Applications, 2024, 236: 121352
doi: 10.1016/j.eswa.2023.121352
22 LIU W, LU H, FU H, et al. Learning to upsample by learning to sample [C]// Proceedings of the IEEE/CVF International Conference on Computer Vision. Paris: IEEE, 2024: 6004–6014.
23 YAN H, ZHANG J UAV-PDD2023: a benchmark dataset for pavement distress detection based on UAV images[J]. Data in Brief, 2023, 51: 109692
doi: 10.1016/j.dib.2023.109692
24 ZHAO Y, LV W, XU S, et al. DETRs beat YOLOs on real-time object detection [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle: IEEE, 2024: 16965–16974.
25 WANG C Y, YEH I H, MARK LIAO H Y. YOLOv9: learning what you want toLearn using programmable gradient information [C]//Computer Vision – ECCV 2024. Cham: Springer, 2025: 1–21.
26 WANG A, CHEN H, LIU L H, et al. YOLOv10: real-time end-to-end object detection [C]// 38th Conference on Neural Information Processing Systems. Vancouver: Curran Associates, 2024: 107984–108011.
27 ZHU J, ZHONG J, MA T, et al Pavement distress detection using convolutional neural networks with images captured via UAV[J]. Automation in Construction, 2022, 133: 103991
doi: 10.1016/j.autcon.2021.103991
[1] 李国燕,于威,梅玉鹏,张明辉,王新强. 全局局部特征融合的遥感图像建筑物提取[J]. 浙江大学学报(工学版), 2026, 60(5): 1100-1108.
[2] 李国燕,李鹏辉,刘榕,梅玉鹏,张明辉. 融合多尺度分辨率和带状特征的遥感道路提取[J]. 浙江大学学报(工学版), 2026, 60(3): 585-593.
[3] 刘登峰,郭文静,陈世海. 基于内容引导注意力的车道线检测网络[J]. 浙江大学学报(工学版), 2025, 59(3): 451-459.
[4] 朵琳,殷瑜,段威,张芸,任勇. 基于改进YOLOv8的船舶目标检测算法[J]. 浙江大学学报(工学版), 2025, 59(11): 2379-2388.
[5] 刘欢,李云红,张蕾涛,郭越,苏雪平,朱耀麟,侯乐乐. 基于MA-ConvNext网络和分步关系知识蒸馏的苹果叶片病害识别[J]. 浙江大学学报(工学版), 2024, 58(9): 1757-1767.
[6] 刘毅,陈一丹,高琳,洪姣. 基于多尺度特征融合的轻量化道路提取模型[J]. 浙江大学学报(工学版), 2024, 58(5): 951-959.
[7] 曹寅,秦俊平,高彤,马千里,任家琪. 基于生成对抗网络的文本两阶段生成高质量图像方法[J]. 浙江大学学报(工学版), 2024, 58(4): 674-683.
[8] 于楠晶,范晓飚,邓天民,冒国韬. 基于多头自注意力的复杂背景船舶检测算法[J]. 浙江大学学报(工学版), 2022, 56(12): 2392-2402.
[9] 杨栋杰,高贤君,冉树浩,张广斌,王萍,杨元维. 基于多重多尺度融合注意力网络的建筑物提取[J]. 浙江大学学报(工学版), 2022, 56(10): 1924-1934.
[10] 陈智超,焦海宁,杨杰,曾华福. 基于改进MobileNet v2的垃圾图像分类算法[J]. 浙江大学学报(工学版), 2021, 55(8): 1490-1499.