Please wait a minute...
浙江大学学报(工学版)  2026, Vol. 60 Issue (9): 1890-1900    DOI: 10.3785/j.issn.1008-973X.2026.09.006
计算机技术、自动控制技术     
基于多尺度特征聚合的航拍图像检测算法
李珺(),丁彬彬,史维娟,杨琳
兰州交通大学 电子与信息工程学院,甘肃 兰州 730070
Aerial image detection algorithm based on multiscale feature aggregation
Jun LI(),Binbin DING,Weijuan SHI,Lin YANG
School of Electronic and Information Engineering, Lanzhou Jiaotong University, Lanzhou 730070, China
 全文: PDF(3870 KB)   HTML
摘要:

针对无人机(UAV)航拍图像检测任务中存在的密集小目标特征混淆导致漏检率高的问题,在YOLOv11的基础上提出基于多尺度特征聚合的航拍图像检测算法FDL-YOLO. 设计基于频率动态卷积的特征提取模块C3k2-FDCN,通过动态调整卷积核频率响应特性,增强对小目标高频细节特征的敏感度. 使用动态自适应尺度融合模块,重构特征金字塔结构并新增160×160高分辨率特征图,实现跨尺度特征的自适应深度融合. 在网络输入端引入LoG-Stem层,增强输入图像边缘特征和强化目标轮廓信息. 在检测头前端构建浅层细节融合模块,通过通道空间双重注意力机制缓解密集目标场景下的特征混淆. 采用基于依赖图的结构剪枝策略,在保留关键特征通道的前提下优化模型参数,实现检测精度与轻量化部署的平衡. 在VisDrone2019数据集上的实验结果表明,利用改进后的算法,准确率提升5.9%,召回率提升8.2%,mAP50提升8.6%~47.2%,参数量下降了70%. 该算法能够有效应对无人机航拍图像目标检测任务中的挑战.

关键词: YOLOv11航拍图像小目标检测多尺度特征结构剪枝轻量化    
Abstract:

An aerial image detection algorithm named FDL-YOLO based on multi-scale feature aggregation was proposed based on YOLOv11 aiming at the problem of high miss detection rate caused by feature confusion of dense small objects in unmanned aerial vehicle (UAV) aerial image detection task. A feature extraction module C3k2-FDCN based on frequency dynamic convolution was designed. The sensitivity to high-frequency detail features of small objects was enhanced by dynamically adjusting the frequency response characteristic of convolution kernel. A dynamic adaptive scale integration module was adopted to reconstruct the feature pyramid structure and add a new 160×160 high-resolution feature map. Then adaptive deep fusion of cross-scale features was realized. A LoG-Stem layer was introduced at the network input to enhance edge features of input image and strengthen target contour information. A shallow detail fusion module was constructed at the front end of the detection head to alleviate feature confusion in dense object scenario through dual channel-spatial attention mechanism. A dependency graph-based structured pruning strategy was employed to optimize model parameters while retaining key feature channel. Then a balance between detection accuracy and lightweight deployment was achieved. The experimental results on the VisDrone2019 dataset showed that the improved algorithm increased accuracy by 5.9%, recall by 8.2%, and mAP50 by 8.6% to 47.2%, with 70% reduction in parameters. The algorithm can effectively cope with the challenges in UAV aerial image object detection task.

Key words: YOLOv11    aerial image    small target detection    multi-scale feature    structural pruning    lightweighting
收稿日期: 2025-07-03 出版日期: 2026-07-20
CLC:  TP 391  
基金资助: 国家自然科学基金资助项目(62241204).
作者简介: 李珺(1974—),女,副教授,博士,从事图像处理和智能计算研究. orcid.org/0009-0006-9519-2069. E-mail:lijane@mail.lzjtu.cn
服务  
把本文推荐给朋友
加入引用管理器
E-mail Alert
作者相关文章  
李珺
丁彬彬
史维娟
杨琳

引用本文:

李珺,丁彬彬,史维娟,杨琳. 基于多尺度特征聚合的航拍图像检测算法[J]. 浙江大学学报(工学版), 2026, 60(9): 1890-1900.

Jun LI,Binbin DING,Weijuan SHI,Lin YANG. Aerial image detection algorithm based on multiscale feature aggregation. Journal of ZheJiang University (Engineering Science), 2026, 60(9): 1890-1900.

链接本文:

https://www.zjujournals.com/eng/CN/10.3785/j.issn.1008-973X.2026.09.006        https://www.zjujournals.com/eng/CN/Y2026/V60/I9/1890

图 1  FDL-YOLO的网络结构
图 2  频率动态卷积的结构图
图 3  C3k2-FDCN的结构图
图 4  改进的网络结构
图 5  维度感知选择性集成模块
图 6  拉普拉斯高斯输入预处理模块
图 7  浅层细节融合模块
图 8  依赖剪枝执行框图
图 9  依赖剪枝过程
模型P/%R/%mAP50/%mAP50-95/%Np/106FLOPs/109v/(帧·s?1)
n42.732.732.310.72.66.3153
s49.737.538.623.19.421.3148
m54.842.043.826.820.067.790
l54.542.644.027.125.386.680
x56.744.846.128.856.8194.556
表 1  不同YOLOv11模型的表现结果
方法C3k2-FDCNNet ImproveLog-StemSDFMPrune
YOLOv11s
方法①
方法②
方法③
方法④
方法⑤
方法⑥
方法⑦
方法⑧
表 2  消融实验的配置
方法P/%R/%mAP50/%mAP75/%mAP50-95/%Np/106F1FLOPs/109Sm/MBv/(帧·s?1)
YOLOv11s49.737.538.623.623.19.442.221.318.3148
方法①50.239.139.524.223.89.543.021.218.4146
方法②53.341.643.127.026.37.746.231.915.5149
方法③50.139.039.824.824.19.443.630.218.5144
方法④52.439.540.725.124.512.144.530.823.6159
方法⑤54.841.244.127.827.07.946.534.515.0141
方法⑥54.543.745.528.928.07.948.037.915.5123
方法⑦56.144.947.029.828.98.349.844.916.3106
方法⑧55.645.747.230.029.12.949.828.912.0122
表 3  消融实验结果
图 10  各个模块的精度提升
剪枝比例/%P/%R/%mAP50/%mAP50-95/%Np/106FLOPs/109Sm/MBv/(帧·s?1)
056.144.947.028.98.344.916.3106
3055.046.447.429.33.231.213.0120
3555.645.747.229.12.928.912.0122
4055.544.546.728.72.726.711.1123
4555.244.046.128.32.524.510.2130
5054.843.845.828.02.322.39.4133
5555.342.545.327.82.221.19.1137
表 4  模型剪枝比例的对比
模型mAP50/%
PedestrianPeopleBicycleMotorTricycleAwning-triTruckCarBusVan总体值
SSD18.79.05.019.111.715.533.163.247.230.025.3
Faster R-CNN20.914.87.321.214.08.819.551.030.529.721.8
CenterNet22.620.614.623.720.117.421.359.737.924.026.2
YOLOv5s39.231.410.638.418.29.826.272.639.933.732.0
YOLOv6s37.229.88.939.623.614.832.578.151.242.635.8
YOLOv8s42.231.612.043.326.414.636.079.256.843.738.6
YOLOv9s38.432.810.342.226.915.034.378.051.543.237.3
YOLOv10s43.134.014.145.027.415.335.279.955.144.739.4
YOLOv11s42.431.811.843.326.614.735.379.456.344.938.7
YOLOv12s41.030.611.342.525.614.233.478.755.643.237.6
YOLOv13s40.631.011.741.126.713.933.578.650.243.537.1
Hyper-YOLOs43.032.813.243.929.215.136.779.561.044.139.8
RT-DETR-r1851.144.217.555.529.217.133.584.756.649.043.8
DMF-YOLO[15]54.543.718.353.533.419.340.885.761.351.546.2
HSF-YOLO[16]53.542.017.753.233.117.736.784.361.249.044.8
YOLO-GAIS[17]48.555.423.852.538.724.435.477.949.439.243.2
FDL-YOLOs55.345.520.355.034.420.741.485.561.952.347.2
表 5  各个类别的精度对比
模型P/%R/%mAP50/%mAP50-
95/%
Np/
106
FLOPs/
109
YOLOv8n42.932.932.819.03.08.1
YOLOv8s49.638.138.623.111.128.5
YOLOv8m53.341.642.726.025.878.7
YOLOv10n43.833.333.219.02.36.5
YOLOv10s49.939.039.423.37.221.4
YOLOv10m53.641.642.925.915.358.9
YOLOv11n42.732.732.310.72.66.3
YOLOv11s49.737.538.623.19.421.3
YOLOv11m54.842.043.826.820.067.7
YOLOv12n41.832.431.618.125.15.8
YOLOv12s49.536.537.622.59.119.3
YOLOv12m53.240.742.225.719.659.5
YOLOv13n42.931.531.418.12.46.1
YOLOv13s47.637.237.122.19.020.1
YOLOv13l53.641.942.525.826.984.4
Hyper-YOLOn44.835.135.020.63.69.5
Hyper-YOLOs51.238.939.824.013.533.8
Hyper-YOLOm53.741.542.526.030.791.8
RT-DETR-r1858.541.543.826.619.857.0
RT-DETR-r3461.545.547.229.031.188.8
AAPW-YOLOn[18]49.937.338.622.42.111.7
PC-YOLOn[19]46.835.436.121.51.9
YOLO-S3DTn[20]47.437.937.921.12.911.0
Eagle-YOLOs[21]53.643.342.925.016.6
PARE-YOLOs[22]60.645.446.328.4
RPS-YOLO[2]55.344.346.328.113.337.7
FDL-YOLOn49.239.339.924.11.110.2
FDL-YOLOs55.645.747.229.12.928.9
表 6  不同算法的检测性能对比
模型P/%R/%mAP50/%mAP50-95/%
YOLOv8s78.680.982.466.7
YOLOv10s75.680.582.766.3
YOLOv11s76.181.082.866.9
YOLOv12s75.281.782.666.3
FDL-YOLOs79.381.683.767.3
表 7  不同模型在SIMD数据集上的性能指标对比
图 11  SIMD数据集的精度对比
模型P/%R/%mAP50/%mAP50-95/%
YOLOv8s41.828.725.98.2
YOLOv10s39.327.924.27.7
YOLOv11s44.327.926.08.3
YOLOv12s43.427.425.48.1
FDL-YOLOs46.534.229.89.5
表 8  TinyPerson数据集结果的对比
图 12  TinyPerson数据集的精度对比
图 13  热力图结果的对比
场景Nd
YOLOv11sFDL-YOLOs
目标密集241300
光线暗淡4861
大小目标4469
背景复杂46102
表 9  改进算法在不同场景下的检测数量对比
图 14  改进算法的检测结果对比
1 BAE M H, PARK S W, PARK J, et al YOLO-RACE: reassembly and convolutional block attention for enhanced dense object detection[J]. Pattern Analysis and Applications, 2025, 28 (2): 90
doi: 10.1007/s10044-025-01471-4
2 LEI P, WANG C, LIU P RPS-YOLO: a recursive pyramid structure-based YOLO network for small object detection in unmanned aerial vehicle scenarios[J]. Applied Sciences, 2025, 15 (4): 2039
doi: 10.3390/app15042039
3 张志豪, 厉小润, 陈淑涵 基于改进YOLO11的无人机航拍图像小目标检测算法[J]. 液晶与显示, 2025, 40 (6): 915- 930
ZHANG Zhihao, LI Xiaorun, CHEN Shuhan Small object detection algorithm in UAV aerial images based on improved YOLO11[J]. Chinese Journal of Liquid Crystals and Displays, 2025, 40 (6): 915- 930
4 BI J, LI K, ZHENG X, et al SPDC-YOLO: an efficient small target detection network based on improved YOLOv8 for drone aerial image[J]. Remote Sensing, 2025, 17 (4): 685
doi: 10.3390/rs17040685
5 WANG Q, YE G, CHEN Q, et al A UAV perspective based lightweight target detection and tracking algorithm for intelligent transportation[J]. Complex and Intelligent Systems, 2025, 11 (1): 73
doi: 10.1007/s40747-024-01687-7
6 童庆洪, 谌海云, 袁杰敏 低光照增强和自适应特征融合的小目标检测算法[J]. 激光与光电子学进展, 2025, 62 (20): 2028008
TONG Qinghong, CHEN Haiyun, YUAN Jiemin Small object detection algorithm based on low-light enhancement and adaptive feature fusion[J]. Laser and Optoelectronics Progress, 2025, 62 (20): 2028008
doi: 10.3788/LOP250914
7 MA F, ZHANG R, ZHU B, et al A lightweight UAV target detection algorithm based on improved YOLOv8s model[J]. Scientific Reports, 2025, 15 (1): 15352
doi: 10.1038/s41598-025-00341-7
8 HAO X, LI T Lightweight small target detection algorithm based on YOLOv8 network improvement[J]. IEEE Access, 2025, 13: 14051- 14062
doi: 10.1109/ACCESS.2025.3529835
9 MA P, FEI H, JIA D, et al YOLOFLY: a consumer-centric framework for efficient object detection in UAV imagery[J]. Electronics, 2025, 14 (3): 498
doi: 10.3390/electronics14030498
10 CHEN L, GU L, LI L, et al. Frequency dynamic convolution for dense image prediction [C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Nashville: IEEE, 2025: 30178–30188.
11 XU S, ZHENG S, XU W, et al. HCF-net: hierarchical context fusion network for infrared small object detection [C]//Proceedings of the IEEE International Conference on Multimedia and Expo. Niagara Falls: IEEE, 2024: 1–6.
12 LU W, CHEN S B, LI H D, et al. LEGNet: a lightweight edge-Gaussian network for low-quality remote sensing image object detection [EB/OL]. [2026-04-13]. https://arxiv.org/abs/2503.14012.
13 TANG L, ZHANG H, XU H, et al Rethinking the necessity of image fusion in high-level vision tasks: a practical infrared and visible image fusion network based on progressive semantic injection and scene fidelity[J]. Information Fusion, 2023, 99: 101870
doi: 10.1016/j.inffus.2023.101870
14 FANG G, MA X, SONG M, et al. DepGraph: towards any structural pruning [C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Vancouver: IEEE, 2023: 16091–16101.
15 贺智轩, 陈里里, 王翔, 等 DMF-YOLOv11: 基于改进YOLOv11n的无人机航拍图像目标检测算法[J]. 计算机工程与应用, 2025, 61 (14): 88- 100
HE Zhixuan, CHEN Lili, WANG Xiang, et al DMF-YOLOv11: target detection algorithm for UAV images based on improved YOLOv11n[J]. Computer Engineering and Applications, 2025, 61 (14): 88- 100
doi: 10.3778/j.issn.1002-8331.2502-0223
16 张轩宇, 周思航, 黄健, 等 基于高阶空间特征提取的无人机航拍小目标检测算法[J]. 计算机工程与应用, 2025, 61 (12): 210- 221
ZHANG Xuanyu, ZHOU Sihang, HUANG Jian, et al High-order spatial feature extraction based small target detection for UAV aerial photographs[J]. Computer Engineering and Applications, 2025, 61 (12): 210- 221
doi: 10.3778/j.issn.1002-8331.2407-0230
17 李凯璇, 刘晓锋, 陈强, 等 YOLOv8-GAIS: 一种改进的无人机航拍目标检测算法[J]. 光电工程, 2025, 52 (4): 75- 88
LI Kaixuan, LIU Xiaofeng, CHEN Qiang, et al YOLOv8-GAIS: improved object detection algorithm for UAV aerial photography[J]. Opto-Electronic Engineering, 2025, 52 (4): 75- 88
18 WU Y, MU X, SHI H, et al An object detection model AAPW-YOLO for UAV remote sensing images based on adaptive convolution and reconstructed feature fusion[J]. Scientific Reports, 2025, 15: 16214
doi: 10.1038/s41598-025-00239-4
19 WANG Z, SU Y, KANG F, et al PC-YOLO11s: a lightweight and effective feature extraction method for small target image detection[J]. Sensors, 2025, 25 (2): 348
doi: 10.3390/s25020348
20 GAO P, LI Z YOLO-S3DT: a small target detection model for UAV images based on YOLOv8[J]. Computers, Materials and Continua, 2025, 82 (3): 4555- 4572
21 WANG D, GAO Z, FANG J, et al Eagle-YOLOv8: UAV object detection inspired by the eagle-eye vision system[J]. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2025, 18: 9432- 9447
doi: 10.1109/JSTARS.2025.3554821
[1] 火久元,姜灏,常琛. 基于改进EAST的高速公路交通标识文本检测[J]. 浙江大学学报(工学版), 2026, 60(9): 1881-1889.
[2] 肖剑,杨小苑,何昕泽,陈林,胡欣. 基于全局信息感知的轻量级螺纹钢表面缺陷检测算法[J]. 浙江大学学报(工学版), 2026, 60(7): 1438-1451.
[3] 于鑫淼,夏楠,江佳鸿,郝子莹,把云胜. 基于多尺度特征相似性匹配的低照度目标检测[J]. 浙江大学学报(工学版), 2026, 60(7): 1464-1474.
[4] 李国燕,于威,梅玉鹏,张明辉,王新强. 全局局部特征融合的遥感图像建筑物提取[J]. 浙江大学学报(工学版), 2026, 60(5): 1100-1108.
[5] 宋耀莲,彭驰,唐菁敏,赵宣植,虞贵财. 基于融合注意力机制的光学遥感图像小目标检测算法[J]. 浙江大学学报(工学版), 2026, 60(4): 763-771.
[6] 武晓春,郭宁. 基于HMARU-net的隧道渗漏水轻量化检测方法[J]. 浙江大学学报(工学版), 2026, 60(3): 468-477.
[7] 李彬彬,张超,覃涛,陈昌盛,刘兴艳,杨靖. 面向光伏电站建设的移动端人体跌倒检测方法[J]. 浙江大学学报(工学版), 2026, 60(3): 546-555.
[8] 李国燕,李鹏辉,刘榕,梅玉鹏,张明辉. 融合多尺度分辨率和带状特征的遥感道路提取[J]. 浙江大学学报(工学版), 2026, 60(3): 585-593.
[9] 雒伟群,陆敬蔚,吴佳缔,梁钰迎,申传鹏,朱睿. 藏区高原典型环境地形目标的轻量化检测模型[J]. 浙江大学学报(工学版), 2026, 60(3): 594-603.
[10] 刘慧,王防修,王意,黄淄博,苏晨. 轻量级改进RT-DETR的葡萄叶片病害检测算法[J]. 浙江大学学报(工学版), 2026, 60(3): 604-613.
[11] 方芳,严军,郭红想,王勇. 基于时空注意力机制的轻量级脑纹识别算法[J]. 浙江大学学报(工学版), 2026, 60(3): 633-642.
[12] 孟昱煜,孔垂乐,火久元,武泽宇. 重构YOLOv11的无人机小目标检测算法[J]. 浙江大学学报(工学版), 2026, 60(2): 303-312.
[13] 谢章郁,杨杰,欧阳嗣源,曾阳剑. 动态场景下融合YOLOv11n目标检测的优化ORB-SLAM3算法[J]. 浙江大学学报(工学版), 2026, 60(2): 313-321.
[14] 闫光辉,黄霄,常文文. 基于脑电多尺度特征和图神经网络的紧急制动行为识别[J]. 浙江大学学报(工学版), 2026, 60(2): 404-414.
[15] 董超群,汪战,廖平,谢帅,荣玉杰,周靖淞. 轻量化YOLOv5s-OCG的轨枕裂纹检测算法[J]. 浙江大学学报(工学版), 2025, 59(9): 1838-1845.