|
|
|
| Data-efficient lightweight model for real-time semantic segmentation |
Hongwei LI( ),Jianyang DING,Zhilei CHAI*( ) |
| School of Artificial Intelligence and Computer Science, Jiangnan University, Wuxi 214122, China |
|
|
|
Abstract An efficient lightweight model, U-SAM, was proposed aiming at the requirement for lightweight, high-precision and high-speed performance in real-time semantic segmentation task on edge device. A multi-scale dilated convolution fusion module was designed in the encoder stage based on an improved U-Net architecture in order to enhance feature representation capability and significantly reduce computational complexity. Channel attention and pixel attention mechanism were introduced in the decoder stage in order to achieve effective modeling of key feature information and cross-scale fusion. Self-attention mechanism and convolution operation were integrated in the bottleneck section in order to efficiently capture long-range dependency and improve the generalization capability of the model. Experiments on the small-scale CamVid dataset demonstrated that U-SAM required only about 1.3×106 parameters in order to achieve segmentation accuracy of 70.2% mIoU, with inference speed reaching 69 frame/s. The overall performance of U-SAM is significantly superior to several mainstream lightweight real-time semantic segmentation models, providing a high-performance solution for real-time semantic segmentation application in resource-constrained scenarios such as autonomous driving, intelligent surveillance and remote sensing satellite.
|
|
Received: 29 November 2025
Published: 20 July 2026
|
|
|
| Fund: 国家自然科学基金资助项目(61972180);中央高校基本科研业务费资助项目(JUSRP202501072);无锡市科技发展基金资助项目(K20241027). |
|
Corresponding Authors:
Zhilei CHAI
E-mail: 6233114016@stu.jiangnan.edu.cn;zlchai@jiangnan.edu.cn
|
用于实时语义分割的数据高效轻量模型
针对边缘设备实时语义分割任务对轻量化、高精度与高速度的要求,提出高效的轻量级模型U-SAM. 该模型基于改进的U-Net架构,在编码阶段设计多尺度空洞卷积融合模块,以增强特征表达能力并显著降低计算复杂度. 在解码阶段引入通道注意力与像素注意力机制,实现关键特征信息的有效建模与跨尺度融合. 在瓶颈部分融合自注意力机制与卷积操作,高效捕捉长距离依赖关系并提升模型泛化能力. 在小规模数据集CamVid上的实验表明,U-SAM仅需约1.3×106的参数量,即可达到70.2%的mIoU分割精度,推理速度达到69帧/s. U-SAM的综合性能显著优于多种主流轻量级实时语义分割模型,为自动驾驶、智能监控、遥感卫星等资源受限场景下的实时语义分割应用提供了高性能解决方案.
关键词:
语义分割,
实时性,
轻量模型,
小规模数据集,
空洞卷积,
注意力机制
|
|
| [1] |
何淼楹, 崔宇超 面向自动驾驶的交通场景语义分割[J]. 计算机应用, 2021, 41 (Suppl.1): 25- 30 HE Miaoying, CUI Yuchao Semantic segmentation of traffic scenes for autonomous driving[J]. Journal of Computer Applications, 2021, 41 (Suppl.1): 25- 30
doi: 10.11772/j.issn.1001-9081.2020071114
|
|
|
| [2] |
熊彬, 张双德 基于改进PSPNet的卫星遥感图像建筑物语义分割算法[J]. 遥感信息, 2023, 38 (4): 73- 79 XIONG Bin, ZHANG Shuangde Semantic segmentation algorithm for buildings in satellite remote sensing images based on improved PSPNet[J]. Remote Sensing Information, 2023, 38 (4): 73- 79
|
|
|
| [3] |
HE K, ZHANG X, REN S, et al. Deep residual learning for image recognition [C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Las Vegas: IEEE, 2016: 770–778.
|
|
|
| [4] |
LONG J, SHELHAMER E, DARRELL T. Fully convolutional networks for semantic segmentation [C]// Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. New York: IEEE, 2015: 3431–3440.
|
|
|
| [5] |
YANG M, YU K, ZHANG C, et al. DenseASPP for semantic segmentation in street scenes [C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Salt Lake City: IEEE, 2018: 3684–3692.
|
|
|
| [6] |
WANG W, XIE E, LI X, et al PVT v2: improved baselines with Pyramid Vision Transformer[J]. Computational Visual Media, 2022, 8 (3): 415- 424
doi: 10.1007/s41095-022-0274-8
|
|
|
| [7] |
GAO R. Rethinking dilated convolution for real-time semantic segmentation [EB/OL]. (2021-11-18) [2025-03-28]. https://arxiv.org/abs/2111.09957.
|
|
|
| [8] |
GAO G, XU G, LI J, et al FBSNet: a fast bilateral symmetrical network for real-time semantic segmentation[J]. IEEE Transactions on Multimedia, 2022, 25: 3273- 3283
doi: 10.1109/tmm.2022.3157995
|
|
|
| [9] |
ZHOU Q, WANG L, GAO G, et al Boundary-guided lightweight semantic segmentation with multi-scale semantic context[J]. IEEE Transactions on Multimedia, 2024, 26: 7887- 7900
doi: 10.1109/TMM.2024.3372835
|
|
|
| [10] |
VASWANI A, SHAZEER N, PARMAR N, et al Attention is all you need[J]. Advances in Neural Information Processing Systems, 2017, 30: 6000- 6010
doi: 10.1007/978-3-031-84300-6_13
|
|
|
| [11] |
DOSOVITSKIY A, BEYER L, KOLESNIKOV A, et al. An image is worth 16x16 words: transformers for image recognition at scale [EB/OL]. (2020-10-22) [2025-03-23]. https://arxiv.org/abs/2010.11929.
|
|
|
| [12] |
HUANG Z, WANG X, HUANG L, et al. CCNet: criss-cross attention for semantic segmentation [C]//Proceedings of the IEEE/CVF International Conference on Computer Vision. Seoul: IEEE, 2020: 603–612.
|
|
|
| [13] |
CHEN L, FU Y, GU L, et al Frequency-aware feature fusion for dense image prediction[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024, 46 (12): 10763- 10780
doi: 10.1109/TPAMI.2024.3449959
|
|
|
| [14] |
DINH B D, NGUYEN T T, TRAN T T, et al. 1M parameters are enough? A lightweight CNN-based model for medical image segmentation [C]//Proceedings of the Asia Pacific Signal and Information Processing Association Annual Summit and Conference. Taipei: IEEE, 2023: 1279–1284.
|
|
|
| [15] |
ZHOU Z, RAHMAN SIDDIQUEE M M, TAJBAKHSH N, et al. UNet++: a nested U-Net architecture for medical image segmentation [C]//Deep Learning in Medical Image Analysis and Multimodal Learning for Clinical Decision Support. Cham: Springer, 2018: 3–11.
|
|
|
| [16] |
RONNEBERGER O, FISCHER P, BROX T. U-Net: convolutional networks for biomedical image segmentation [C]//Medical Image Computing and Computer-Assisted Intervention. Cham: Springer, 2015: 234–241.
|
|
|
| [17] |
OKTAY O, SCHLEMPER J, LE FOLGOC L, et al. Attention U-Net: learning where to look for the pancreas [EB/OL]. (2018-04-11) [2025-03-25]. https://arxiv.org/abs/1804.03999.
|
|
|
| [18] |
HAN Z, JIAN M, WANG G G ConvUNeXt: an efficient convolution neural network for medical image segmentation[J]. Knowledge-Based Systems, 2022, 253: 109512
doi: 10.1016/j.knosys.2022.109512
|
|
|
| [19] |
BADRINARAYANAN V, KENDALL A, CIPOLLA R SegNet: a deep convolutional encoder-decoder architecture for image segmentation[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2017, 39 (12): 2481- 2495
doi: 10.1109/TPAMI.2016.2644615
|
|
|
| [20] |
LI H, XIONG P, FAN H, et al. Dfanet: deep feature aggregation for real-time semantic segmentation [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. New York: IEEE, 2019: 9522–9531.
|
|
|
| [21] |
ZHAO H, QI X, SHEN X, et al. Icnet for real-time semantic segmentation on high-resolution images [C]// Proceedings of the European Conference on Computer Vision. Cham: Springer, 2018: 405–420.
|
|
|
| [22] |
FAN J, WANG F, CHU H, et al MLFNet: multi-level fusion network for real-time semantic segmentation of autonomous driving[J]. IEEE Transactions on Intelligent Vehicles, 2022, 8 (1): 756- 767
doi: 10.1109/wcnc61545.2025.10978834
|
|
|
| [23] |
HE J, LIANG S, WU X, et al MGSeg: multiple granularity-based real-time semantic segmentation network[J]. IEEE Transactions on Image Processing, 2021, 30: 7200- 7214
doi: 10.1109/TIP.2021.3102509
|
|
|
| [24] |
PASZKE A, CHAURASIA A, KIM S, et al. ENet: a deep neural network architecture for real-time semantic segmentation [EB/OL]. (2016-06-07) [2025-03-06]. https://arxiv.org/abs/1606.02147.
|
|
|
| [25] |
LI S, YAN Q, ZHOU X, et al NDNet: spacewise multiscale representation learning via neighbor decoupling for real-time driving scene parsing[J]. IEEE Transactions on Neural Networks and Learning Systems, 2024, 35 (6): 7884- 7898
doi: 10.1109/TNNLS.2022.3221745
|
|
|
| [26] |
LI G, YUN I, KIM J, et al. DABNet: depth-wise asymmetric bottleneck for real-time semantic segmentation [EB/OL]. (2019-07-25) [2025-02-15]. https://arxiv.org/abs/1907.11357.
|
|
|
| [27] |
LIU J, ZHOU Q, QIANG Y, et al. FDDWNet: a lightweight convolutional neural network for real-time semantic segmentation [C]//2020 IEEE International Conference on Acoustics, Speech and Signal Processing. Barcelona: IEEE, 2020: 2373–2377.
|
|
|
| [28] |
HU X, XU S, JING L Lightweight attention-guided redundancy-reuse network for real-time semantic segmentation[J]. IET Image Processing, 2023, 17 (9): 2649- 2658
doi: 10.1049/ipr2.12816
|
|
|
| [29] |
吴马靖, 张永爱, 林珊玲, 等 基于BiLevelNet的实时语义分割算法[J]. 光电工程, 2024, 51 (5): 240030 WU Majing, ZHANG Yongai, LIN Shanling, et al Real-time semantic segmentation algorithm based on BiLevelNet[J]. Opto-Electronic Engineering, 2024, 51 (5): 240030
doi: 10.12086/oee.2024.240030
|
|
|
| [30] |
ZHANG T, ZHUANG Y, WANG G, et al Multiscale semantic fusion-guided fractal convolutional object detection network for optical remote sensing imagery[J]. IEEE Transactions on Geoscience and Remote Sensing, 2022, 60: 5608720
doi: 10.1109/tgrs.2021.3108476
|
|
|
|
Viewed |
|
|
|
Full text
|
|
|
|
|
Abstract
|
|
|
|
|
Cited |
|
|
|
|
| |
Shared |
|
|
|
|
| |
Discussed |
|
|
|
|