Please wait a minute...
Journal of ZheJiang University (Engineering Science)  2026, Vol. 60 Issue (9): 1890-1900    DOI: 10.3785/j.issn.1008-973X.2026.09.006
    
Aerial image detection algorithm based on multiscale feature aggregation
Jun LI(),Binbin DING,Weijuan SHI,Lin YANG
School of Electronic and Information Engineering, Lanzhou Jiaotong University, Lanzhou 730070, China
Download: HTML     PDF(3870KB) HTML
Export: BibTeX | EndNote (RIS)      

Abstract  

An aerial image detection algorithm named FDL-YOLO based on multi-scale feature aggregation was proposed based on YOLOv11 aiming at the problem of high miss detection rate caused by feature confusion of dense small objects in unmanned aerial vehicle (UAV) aerial image detection task. A feature extraction module C3k2-FDCN based on frequency dynamic convolution was designed. The sensitivity to high-frequency detail features of small objects was enhanced by dynamically adjusting the frequency response characteristic of convolution kernel. A dynamic adaptive scale integration module was adopted to reconstruct the feature pyramid structure and add a new 160×160 high-resolution feature map. Then adaptive deep fusion of cross-scale features was realized. A LoG-Stem layer was introduced at the network input to enhance edge features of input image and strengthen target contour information. A shallow detail fusion module was constructed at the front end of the detection head to alleviate feature confusion in dense object scenario through dual channel-spatial attention mechanism. A dependency graph-based structured pruning strategy was employed to optimize model parameters while retaining key feature channel. Then a balance between detection accuracy and lightweight deployment was achieved. The experimental results on the VisDrone2019 dataset showed that the improved algorithm increased accuracy by 5.9%, recall by 8.2%, and mAP50 by 8.6% to 47.2%, with 70% reduction in parameters. The algorithm can effectively cope with the challenges in UAV aerial image object detection task.



Key wordsYOLOv11      aerial image      small target detection      multi-scale feature      structural pruning      lightweighting     
Received: 03 July 2025      Published: 20 July 2026
CLC:  TP 391  
Fund:  国家自然科学基金资助项目(62241204).
Cite this article:

Jun LI,Binbin DING,Weijuan SHI,Lin YANG. Aerial image detection algorithm based on multiscale feature aggregation. Journal of ZheJiang University (Engineering Science), 2026, 60(9): 1890-1900.

URL:

https://www.zjujournals.com/eng/10.3785/j.issn.1008-973X.2026.09.006     OR     https://www.zjujournals.com/eng/Y2026/V60/I9/1890


基于多尺度特征聚合的航拍图像检测算法

针对无人机(UAV)航拍图像检测任务中存在的密集小目标特征混淆导致漏检率高的问题,在YOLOv11的基础上提出基于多尺度特征聚合的航拍图像检测算法FDL-YOLO. 设计基于频率动态卷积的特征提取模块C3k2-FDCN,通过动态调整卷积核频率响应特性,增强对小目标高频细节特征的敏感度. 使用动态自适应尺度融合模块,重构特征金字塔结构并新增160×160高分辨率特征图,实现跨尺度特征的自适应深度融合. 在网络输入端引入LoG-Stem层,增强输入图像边缘特征和强化目标轮廓信息. 在检测头前端构建浅层细节融合模块,通过通道空间双重注意力机制缓解密集目标场景下的特征混淆. 采用基于依赖图的结构剪枝策略,在保留关键特征通道的前提下优化模型参数,实现检测精度与轻量化部署的平衡. 在VisDrone2019数据集上的实验结果表明,利用改进后的算法,准确率提升5.9%,召回率提升8.2%,mAP50提升8.6%~47.2%,参数量下降了70%. 该算法能够有效应对无人机航拍图像目标检测任务中的挑战.


关键词: YOLOv11,  航拍图像,  小目标检测,  多尺度特征,  结构剪枝,  轻量化 
Fig.1 Network structure of FDL-YOLO
Fig.2 Structure diagram of frequency dynamic convolution
Fig.3 Structure diagram of C3k2-FDCN
Fig.4 Improved network structure
Fig.5 Dimension aware selective integration module
Fig.6 Laplacian of Gaussian input preprocessing module
Fig.7 Superficial detail fusion module
Fig.8 DepGraph pruning execution diagram
Fig.9 DepGraph pruning process
模型P/%R/%mAP50/%mAP50-95/%Np/106FLOPs/109v/(帧·s?1)
n42.732.732.310.72.66.3153
s49.737.538.623.19.421.3148
m54.842.043.826.820.067.790
l54.542.644.027.125.386.680
x56.744.846.128.856.8194.556
Tab.1 Performance result of different YOLOv11 models
方法C3k2-FDCNNet ImproveLog-StemSDFMPrune
YOLOv11s
方法①
方法②
方法③
方法④
方法⑤
方法⑥
方法⑦
方法⑧
Tab.2 Configuration of ablation experiment
方法P/%R/%mAP50/%mAP75/%mAP50-95/%Np/106F1FLOPs/109Sm/MBv/(帧·s?1)
YOLOv11s49.737.538.623.623.19.442.221.318.3148
方法①50.239.139.524.223.89.543.021.218.4146
方法②53.341.643.127.026.37.746.231.915.5149
方法③50.139.039.824.824.19.443.630.218.5144
方法④52.439.540.725.124.512.144.530.823.6159
方法⑤54.841.244.127.827.07.946.534.515.0141
方法⑥54.543.745.528.928.07.948.037.915.5123
方法⑦56.144.947.029.828.98.349.844.916.3106
方法⑧55.645.747.230.029.12.949.828.912.0122
Tab.3 Result of ablation experiment
Fig.10 Improvement in accuracy of each module
剪枝比例/%P/%R/%mAP50/%mAP50-95/%Np/106FLOPs/109Sm/MBv/(帧·s?1)
056.144.947.028.98.344.916.3106
3055.046.447.429.33.231.213.0120
3555.645.747.229.12.928.912.0122
4055.544.546.728.72.726.711.1123
4555.244.046.128.32.524.510.2130
5054.843.845.828.02.322.39.4133
5555.342.545.327.82.221.19.1137
Tab.4 Comparison of model pruning ratio
模型mAP50/%
PedestrianPeopleBicycleMotorTricycleAwning-triTruckCarBusVan总体值
SSD18.79.05.019.111.715.533.163.247.230.025.3
Faster R-CNN20.914.87.321.214.08.819.551.030.529.721.8
CenterNet22.620.614.623.720.117.421.359.737.924.026.2
YOLOv5s39.231.410.638.418.29.826.272.639.933.732.0
YOLOv6s37.229.88.939.623.614.832.578.151.242.635.8
YOLOv8s42.231.612.043.326.414.636.079.256.843.738.6
YOLOv9s38.432.810.342.226.915.034.378.051.543.237.3
YOLOv10s43.134.014.145.027.415.335.279.955.144.739.4
YOLOv11s42.431.811.843.326.614.735.379.456.344.938.7
YOLOv12s41.030.611.342.525.614.233.478.755.643.237.6
YOLOv13s40.631.011.741.126.713.933.578.650.243.537.1
Hyper-YOLOs43.032.813.243.929.215.136.779.561.044.139.8
RT-DETR-r1851.144.217.555.529.217.133.584.756.649.043.8
DMF-YOLO[15]54.543.718.353.533.419.340.885.761.351.546.2
HSF-YOLO[16]53.542.017.753.233.117.736.784.361.249.044.8
YOLO-GAIS[17]48.555.423.852.538.724.435.477.949.439.243.2
FDL-YOLOs55.345.520.355.034.420.741.485.561.952.347.2
Tab.5 Comparison of accuracy among various categories
模型P/%R/%mAP50/%mAP50-
95/%
Np/
106
FLOPs/
109
YOLOv8n42.932.932.819.03.08.1
YOLOv8s49.638.138.623.111.128.5
YOLOv8m53.341.642.726.025.878.7
YOLOv10n43.833.333.219.02.36.5
YOLOv10s49.939.039.423.37.221.4
YOLOv10m53.641.642.925.915.358.9
YOLOv11n42.732.732.310.72.66.3
YOLOv11s49.737.538.623.19.421.3
YOLOv11m54.842.043.826.820.067.7
YOLOv12n41.832.431.618.125.15.8
YOLOv12s49.536.537.622.59.119.3
YOLOv12m53.240.742.225.719.659.5
YOLOv13n42.931.531.418.12.46.1
YOLOv13s47.637.237.122.19.020.1
YOLOv13l53.641.942.525.826.984.4
Hyper-YOLOn44.835.135.020.63.69.5
Hyper-YOLOs51.238.939.824.013.533.8
Hyper-YOLOm53.741.542.526.030.791.8
RT-DETR-r1858.541.543.826.619.857.0
RT-DETR-r3461.545.547.229.031.188.8
AAPW-YOLOn[18]49.937.338.622.42.111.7
PC-YOLOn[19]46.835.436.121.51.9
YOLO-S3DTn[20]47.437.937.921.12.911.0
Eagle-YOLOs[21]53.643.342.925.016.6
PARE-YOLOs[22]60.645.446.328.4
RPS-YOLO[2]55.344.346.328.113.337.7
FDL-YOLOn49.239.339.924.11.110.2
FDL-YOLOs55.645.747.229.12.928.9
Tab.6 Comparison of detection performance of different algorithms
模型P/%R/%mAP50/%mAP50-95/%
YOLOv8s78.680.982.466.7
YOLOv10s75.680.582.766.3
YOLOv11s76.181.082.866.9
YOLOv12s75.281.782.666.3
FDL-YOLOs79.381.683.767.3
Tab.7 Comparison of performance metrics of different models on SIMD dataset
Fig.11 Comparison of accuracy of SIMD dataset
模型P/%R/%mAP50/%mAP50-95/%
YOLOv8s41.828.725.98.2
YOLOv10s39.327.924.27.7
YOLOv11s44.327.926.08.3
YOLOv12s43.427.425.48.1
FDL-YOLOs46.534.229.89.5
Tab.8 Comparison of TinyPerson dataset result
Fig.12 Comparison of accuracy of TinyPerson dataset
Fig.13 Comparison of heat map result
场景Nd
YOLOv11sFDL-YOLOs
目标密集241300
光线暗淡4861
大小目标4469
背景复杂46102
Tab.9 Comparison of detection quantity of improved algorithm under different scenario
Fig.14 Comparison of detection result of improved algorithm
[1]   BAE M H, PARK S W, PARK J, et al YOLO-RACE: reassembly and convolutional block attention for enhanced dense object detection[J]. Pattern Analysis and Applications, 2025, 28 (2): 90
doi: 10.1007/s10044-025-01471-4
[2]   LEI P, WANG C, LIU P RPS-YOLO: a recursive pyramid structure-based YOLO network for small object detection in unmanned aerial vehicle scenarios[J]. Applied Sciences, 2025, 15 (4): 2039
doi: 10.3390/app15042039
[3]   张志豪, 厉小润, 陈淑涵 基于改进YOLO11的无人机航拍图像小目标检测算法[J]. 液晶与显示, 2025, 40 (6): 915- 930
ZHANG Zhihao, LI Xiaorun, CHEN Shuhan Small object detection algorithm in UAV aerial images based on improved YOLO11[J]. Chinese Journal of Liquid Crystals and Displays, 2025, 40 (6): 915- 930
[4]   BI J, LI K, ZHENG X, et al SPDC-YOLO: an efficient small target detection network based on improved YOLOv8 for drone aerial image[J]. Remote Sensing, 2025, 17 (4): 685
doi: 10.3390/rs17040685
[5]   WANG Q, YE G, CHEN Q, et al A UAV perspective based lightweight target detection and tracking algorithm for intelligent transportation[J]. Complex and Intelligent Systems, 2025, 11 (1): 73
doi: 10.1007/s40747-024-01687-7
[6]   童庆洪, 谌海云, 袁杰敏 低光照增强和自适应特征融合的小目标检测算法[J]. 激光与光电子学进展, 2025, 62 (20): 2028008
TONG Qinghong, CHEN Haiyun, YUAN Jiemin Small object detection algorithm based on low-light enhancement and adaptive feature fusion[J]. Laser and Optoelectronics Progress, 2025, 62 (20): 2028008
doi: 10.3788/LOP250914
[7]   MA F, ZHANG R, ZHU B, et al A lightweight UAV target detection algorithm based on improved YOLOv8s model[J]. Scientific Reports, 2025, 15 (1): 15352
doi: 10.1038/s41598-025-00341-7
[8]   HAO X, LI T Lightweight small target detection algorithm based on YOLOv8 network improvement[J]. IEEE Access, 2025, 13: 14051- 14062
doi: 10.1109/ACCESS.2025.3529835
[9]   MA P, FEI H, JIA D, et al YOLOFLY: a consumer-centric framework for efficient object detection in UAV imagery[J]. Electronics, 2025, 14 (3): 498
doi: 10.3390/electronics14030498
[10]   CHEN L, GU L, LI L, et al. Frequency dynamic convolution for dense image prediction [C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Nashville: IEEE, 2025: 30178–30188.
[11]   XU S, ZHENG S, XU W, et al. HCF-net: hierarchical context fusion network for infrared small object detection [C]//Proceedings of the IEEE International Conference on Multimedia and Expo. Niagara Falls: IEEE, 2024: 1–6.
[12]   LU W, CHEN S B, LI H D, et al. LEGNet: a lightweight edge-Gaussian network for low-quality remote sensing image object detection [EB/OL]. [2026-04-13]. https://arxiv.org/abs/2503.14012.
[13]   TANG L, ZHANG H, XU H, et al Rethinking the necessity of image fusion in high-level vision tasks: a practical infrared and visible image fusion network based on progressive semantic injection and scene fidelity[J]. Information Fusion, 2023, 99: 101870
doi: 10.1016/j.inffus.2023.101870
[14]   FANG G, MA X, SONG M, et al. DepGraph: towards any structural pruning [C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Vancouver: IEEE, 2023: 16091–16101.
[15]   贺智轩, 陈里里, 王翔, 等 DMF-YOLOv11: 基于改进YOLOv11n的无人机航拍图像目标检测算法[J]. 计算机工程与应用, 2025, 61 (14): 88- 100
HE Zhixuan, CHEN Lili, WANG Xiang, et al DMF-YOLOv11: target detection algorithm for UAV images based on improved YOLOv11n[J]. Computer Engineering and Applications, 2025, 61 (14): 88- 100
doi: 10.3778/j.issn.1002-8331.2502-0223
[16]   张轩宇, 周思航, 黄健, 等 基于高阶空间特征提取的无人机航拍小目标检测算法[J]. 计算机工程与应用, 2025, 61 (12): 210- 221
ZHANG Xuanyu, ZHOU Sihang, HUANG Jian, et al High-order spatial feature extraction based small target detection for UAV aerial photographs[J]. Computer Engineering and Applications, 2025, 61 (12): 210- 221
doi: 10.3778/j.issn.1002-8331.2407-0230
[17]   李凯璇, 刘晓锋, 陈强, 等 YOLOv8-GAIS: 一种改进的无人机航拍目标检测算法[J]. 光电工程, 2025, 52 (4): 75- 88
LI Kaixuan, LIU Xiaofeng, CHEN Qiang, et al YOLOv8-GAIS: improved object detection algorithm for UAV aerial photography[J]. Opto-Electronic Engineering, 2025, 52 (4): 75- 88
[18]   WU Y, MU X, SHI H, et al An object detection model AAPW-YOLO for UAV remote sensing images based on adaptive convolution and reconstructed feature fusion[J]. Scientific Reports, 2025, 15: 16214
doi: 10.1038/s41598-025-00239-4
[19]   WANG Z, SU Y, KANG F, et al PC-YOLO11s: a lightweight and effective feature extraction method for small target image detection[J]. Sensors, 2025, 25 (2): 348
doi: 10.3390/s25020348
[20]   GAO P, LI Z YOLO-S3DT: a small target detection model for UAV images based on YOLOv8[J]. Computers, Materials and Continua, 2025, 82 (3): 4555- 4572
[21]   WANG D, GAO Z, FANG J, et al Eagle-YOLOv8: UAV object detection inspired by the eagle-eye vision system[J]. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2025, 18: 9432- 9447
doi: 10.1109/JSTARS.2025.3554821
[1] Jiuyuan HUO,Hao JIANG,Chen CHANG. Text detection of highway traffic sign based on improved EAST[J]. Journal of ZheJiang University (Engineering Science), 2026, 60(9): 1881-1889.
[2] Xinmiao YU,Nan XIA,Jiahong JIANG,Ziying HAO,Yunsheng BA. Low-light target detection based on multi-scale feature similarity matching[J]. Journal of ZheJiang University (Engineering Science), 2026, 60(7): 1464-1474.
[3] Guoyan LI,Wei YU,Yupeng MEI,Minghui ZHANG,Xinqiang WANG. Building extraction from remote sensing images with global-local feature fusion[J]. Journal of ZheJiang University (Engineering Science), 2026, 60(5): 1100-1108.
[4] Yaolian SONG,Chi PENG,Jingmin TANG,Xuanzhi ZHAO,Guicai YU. Small object detection algorithm for optical remote sensing images based on fusion attention mechanism[J]. Journal of ZheJiang University (Engineering Science), 2026, 60(4): 763-771.
[5] Guoyan LI,Penghui LI,Rong LIU,Yupeng MEI,Minghui ZHANG. Remote sensing road extraction by fusing multi-scale resolution and strip feature[J]. Journal of ZheJiang University (Engineering Science), 2026, 60(3): 585-593.
[6] Yuyu MENG,Chuile KONG,Jiuyuan HUO,Zeyu WU. UAV small target detection algorithm based on reconstruction of YOLOv11[J]. Journal of ZheJiang University (Engineering Science), 2026, 60(2): 303-312.
[7] Zhangyu XIE,Jie YANG,Siyuan OUYANG,Yangjian ZENG. Optimized ORB-SLAM3 algorithm incorporating YOLOv11n object detection for dynamic scenes[J]. Journal of ZheJiang University (Engineering Science), 2026, 60(2): 313-321.
[8] Guanghui YAN,Xiao HUANG,Wenwen CHANG. Emergency braking behavior recognition based on EEG multi-scale features and graph neural networks[J]. Journal of ZheJiang University (Engineering Science), 2026, 60(2): 404-414.
[9] Yahong ZHAI,Yaling CHEN,Longyan XU,Yu GONG. Improved YOLOv8s lightweight small target detection algorithm of UAV aerial image[J]. Journal of ZheJiang University (Engineering Science), 2025, 59(8): 1708-1717.
[10] Xinyu WEI,Lei RAO,Guangyu FAN,Niansheng CHEN,Songlin CHENG,Dingyu YANG. High-precision real-time semantic segmentation network for UAV remote sensing images[J]. Journal of ZheJiang University (Engineering Science), 2025, 59(7): 1411-1420.
[11] Dengfeng LIU,Wenjing GUO,Shihai CHEN. Content-guided attention-based lane detection network[J]. Journal of ZheJiang University (Engineering Science), 2025, 59(3): 451-459.
[12] Junyin WANG,Bin WEN,Yanjun SHEN,Jun ZHANG,Zihao WANG. Surface defect detection method for aluminum profiles based on improved YOLOv7-tiny[J]. Journal of ZheJiang University (Engineering Science), 2025, 59(3): 523-534.
[13] Jianghao CHEN,Jun YANG. Object detection for multi-source remote sensing fused images based on depthwise separable convolution[J]. Journal of ZheJiang University (Engineering Science), 2025, 59(12): 2545-2555.
[14] Bing YANG,Chuyang XU,Jinliang YAO,Xueqin XIANG. 3D hand pose estimation method based on monocular RGB images[J]. Journal of ZheJiang University (Engineering Science), 2025, 59(1): 18-26.
[15] Kang HAN,Hongfei ZHAN,Junhe YU,Rui WANG. Rolling bearing fault diagnosis based on dilated convolution and enhanced multi-scale feature adaptive fusion[J]. Journal of ZheJiang University (Engineering Science), 2024, 58(6): 1285-1295.