Please wait a minute...
浙江大学学报(工学版)  2026, Vol. 60 Issue (9): 1872-1880    DOI: 10.3785/j.issn.1008-973X.2026.09.004
机械工程     
基于注意力增强的跨模态多层融合网络的物体位姿估计
杨恒(),王韶涵,董青,赵科渊,杨明亮
太原科技大学 机械工程学院,山西 太原 030024
Object pose estimation based on attention enhancement cross-modal multilayer fusion network
Heng YANG(),Shaohan WANG,Qing DONG,Keyuan ZHAO,Mingliang YANG
School of Mechanical Engineering, Taiyuan University of Science and Technology, Taiyuan 030024, China
 全文: PDF(1443 KB)   HTML
摘要:

提出鲁棒的端到端6D位姿估计网络. 利用层次化特征融合机制,使得RGB-D数据有效融合. 网络通过注意力增强模块提取像素级特征,通过多尺度网络对融合后的像素级多模态特征进行处理,得到多尺度多模态特征. 将多模态特征分别送至后续模块中,得到对应的全局特征. 将多模态特征与全局特征进行拼接,形成包含全局信息的多尺度密集判别特征. 根据多尺度密集判别特征得到最终位姿,实现精确的6D位姿估计. 在LineMOD和YCB-Video数据集上开展多项实验,所提方法的表现出色,特别是在处理YCB-Video数据集中具有挑战性的光滑和无纹理物体时表现突出. 在LineMOD数据集上,所提方法的精度优于现有的代表性方法DenseFusion,处理速度显著提升.

关键词: 6D位姿深度学习神经网络跨模态融合注意力机制    
Abstract:

A robust end-to-end 6D pose estimation network was proposed. A hierarchical feature fusion mechanism was used to effectively integrate RGB-D data. The pixel-level feature was extracted via an attention enhancement module. Then a multi-scale network was employed to process the fused pixel-level multimodal feature, yielding multi-scale multimodal features. The multimodal features were subsequently fed into subsequent modules to obtain corresponding global features. The global features were concatenated with the multimodal features to form multi-scale dense discriminative features that encode global information. The final pose was derived from multi-scale dense discriminative features, enabling accurate 6D pose estimation. Extensive experiments on the LineMOD and YCB-Video datasets demonstrate that the proposed method achieves superior performance, particularly in handling challenging smooth and textureless objects in the YCB-Video dataset. The method outperforms the representative approach DenseFusion in terms of accuracy on the LineMOD dataset, while achieving significantly higher processing speed.

Key words: 6D pose    deep learning    neural network    cross-modal fusion    attention mechanism
收稿日期: 2025-07-17 出版日期: 2026-07-20
CLC:  TP 183  
基金资助: 国家自然科学基金资助项目(52105269);省部级科技计划资助项目(202402150101006).
作者简介: 杨恒(1982—),男,副教授,从事智能机械装备关键技术的研究. orcid.org/0009-0004-1920-8677. E-mail:93328173@qq.com
服务  
把本文推荐给朋友
加入引用管理器
E-mail Alert
作者相关文章  
杨恒
王韶涵
董青
赵科渊
杨明亮

引用本文:

杨恒,王韶涵,董青,赵科渊,杨明亮. 基于注意力增强的跨模态多层融合网络的物体位姿估计[J]. 浙江大学学报(工学版), 2026, 60(9): 1872-1880.

Heng YANG,Shaohan WANG,Qing DONG,Keyuan ZHAO,Mingliang YANG. Object pose estimation based on attention enhancement cross-modal multilayer fusion network. Journal of ZheJiang University (Engineering Science), 2026, 60(9): 1872-1880.

链接本文:

https://www.zjujournals.com/eng/CN/10.3785/j.issn.1008-973X.2026.09.004        https://www.zjujournals.com/eng/CN/Y2026/V60/I9/1872

图 1  从RGB-D图像中提取像素级特征
图 2  DeepFusion6D网络的整体结构
图 3  AFE的整体结构
图 4  提取的多尺度密集判别特征中每个模块的输出大小
方法输入数据ADD(-S)分数/%
ApeBenCameraCanCatDrillerDuckEggboxGlueHoleIronLampPhone平均值
BB8[19]RGB40.291.755.863.762.973.744.358.141.067.284.776.954.062.3
PVNet[21]43.699.986.995.579.396.452.699.295.781.998.999.392.486.3
DeepIM[28]77.297.893.796.882.095.077.897.199.352.898.398.087.989.0
TexPose[36]80.999.094.899.792.697.483.494.993.479.399.898.378.991.7
SSD-6D+ICP[20]RGB-D65.080.078.086.070.073.066.010010049.078.073.079.076.7
CFFM+CWTM[23]88.291.593.493.894.491.689.497.399.189.395.794.293.793.2
Per-pixel Densefusion[13]79.884.176.786.888.977.876.210099.079.092.192.088.086.1
Iterative-Densefusion[13]92.193.294.393.196.587.092.010010092.197.095.192.894.3
本文方法(仅使用多尺度模块)93.889.690.590.291.190.389.510010084.793.191.589.089.4
本文方法(仅使用注意力增强模块)89.390.189.789.290.289.888.710098.489.390.289.790.089.0
本文方法95.493.195.993.295.094.290.310010092.296.595.194.895.2
表 1  linenmod数据集上ADD(-S)分数的定量评估
方法ADD(-S)分数/%
master-
chef-
can
cracker-
box
sugar-
box
tomato-
soup-
can
mustard-
bottle
tuna-
fish-
can
pudding-
box
gelatin-
box
potted-
meat-
can
Bananapitcher-
base
bleach-
cleaner
BowlMugpower-
drill
wood-
block
Scissorlarge-
marker
large-
clamp
extra-
large-
clamp
foam-
brick
平均值
PointFusion[12]99.862.295.496.984.099.896.710088.570.579.865.024.199.822.818.235.980.450.020.110074.1
PVN3D[31]10091.610096.910010010010093.699.710099.454.999.899.680.295.699.774.948.810093.2
DenseFusion[13]10099.510096.910010010010091.310010010098.810098.794.610010079.276.310096.8
CFFM+CWTM[23]10095.310095.710099.093.910090.893.410097.853.298.297.481.394.796.768.362.599.693.5
FoundationPose[14]99.198.799.398.999.599.098.899.498.599.298.698.398.098.798.497.597.198.296.996.599.199.0
本文方法
(仅使用多
尺度模块)
98.593.198.089.194.397.393.999.992.495.197.493.692.194.194.693.291.096.288.183.299.494.3
本文方法
(仅使用注意
力增强模块)
99.194.998.790.695.797.894.710093.796.398.195.193.595.395.894.792.597.590.385.899.795.2
本文方法10097.599.093.597.598.596.510096.097.598.097.099.297.096.598.096.598.085.082.010097.1
方法AUC/%
master-
chef-
can
cracker-
box
sugar-
box
tomato-
soup-
can
mustard-
bottle
tuna-
fish-
can
pudding-
box
gelatin-
box
potted-
meat-
can
Bananapitcher-
base
bleach-
cleaner
BowlMugpower-
drill
wood-
block
Scissorlarge-
marker
large-
clamp
extra-
large-
clamp
foam-
brick
平均值
PointFusion[12]90.980.590.491.988.593.887.595.086.484.785.581.075.794.271.568.176.787.965.960.491.883.9
PVN3D[31]95.892.798.294.598.697.197.998.892.797.197.896.981.095.098.287.691.797.275.264.497.293.0
DenseFusion[13]96.495.597.594.697.296.696.598.191.396.697.195.888.297.196.089.795.297.572.969.892.593.1
CFFM+CWTM[23]94.392.298.793.997.194.893.399.191.691.195.286.583.890.594.884.392.888.673.262.393.390.5
FoundationPose[14]96.594.296.895.097.496.194.597.093.896.394.092.591.894.393.189.588.292.087.585.495.993.2
本文方法
(仅使用多
尺度模块)
93.984.393.780.990.891.389.297.885.987.993.689.479.589.187.285.478.789.373.568.794.088.5
本文方法
(仅使用注意
力增强模块)
94.786.794.482.792.092.790.398.388.289.794.891.082.790.688.987.782.391.075.271.595.589.7
本文方法96.998.698.695.096.594.593.199.090.391.895.593.488.392.590.510085.598.482.580.597.093.7
表 2  YCB-Video数据集上ADD(-S)分数和ADD(-S) AUC的定量评估
模型配置Np/106FLOPs/109tin/msAUC/%ADD(-S)
分数/%
Baseline78.035.916.483.590.1
+Multi-scale91.046.221.585.892.5
+AFE(Attention)81.037.517.688.995.3
完整模型(本文方法)94.047.922.393.797.1
表 3  YCB-Video数据集上的效率与精度消融实验
图 5  PoseCNN和地面真实分割的‘large_clamp’ 和‘extra_large_clamp’的ADD(-S)分数-阈值曲线
图 6  不同方法在YCB-Video数据集上随遮挡水平增加的性能变化
图 7  DeepFusion6D在LINEMOD数据集上估计物体6D位姿的可视化结果
1 JIN D, ZHANG R, LI Y, et al. Mitigating latency effects on subjective experience in robot teleoperation using a VR-enabled virtual spring [C]//Proceedings of the IEEE International Symposium on Mixed and Augmented Reality. Bellevue: IEEE, 2024: 1276–1282.
2 ZHAO Y, WANG B, YU H, et al. Research on the self-powered flexible non-contact distance/contact pressure dual-mode sensor for manipulator based on triboelectric nanogenerator [C]//Proceedings of the China Automation Congress. Qingdao: IEEE, 2024: 437–441.
3 PAWAR U S, KULKARNI R A. Navigation with augmented reality [C]//Proceedings of the 1st International Conference on AIML: Applications for Engineering and Technology. Pune: IEEE, 2025: 1–6.
4 DIAS P, MARQUES B, FELIX I, et al. Creating asynchronous augmented reality instructions through a virtual reality framework [C]//Proceedings of the IEEE Conference on Virtual Reality and 3D User Interfaces Abstracts and Workshops. Saint Malo: IEEE, 2025: 1065–1071.
5 KIM J, JEON J, PARK J, et al. Continuous performance improvement of infrastructure guidance service for autonomous cooperative driving: focusing on data-centric AI [C]//Proceedings of the International Conference on Electronics, Information, and Communication. Taipei: IEEE, 2024: 1–4.
6 FENG Y, SHEN H, SHAN Z, et al Semantic communication for edge intelligence enabled autonomous driving system[J]. IEEE Network, 2025, 39 (2): 149- 157
doi: 10.1109/MNET.2024.3468328
7 CHEN J, LEI B, SONG Q, et al. A hierarchical graph network for 3D object detection on point clouds [C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle: IEEE, 2020: 389–398.
8 LU P, SHI W, QIAO X Multi-view image-based 3D reconstruction in indoor scenes: a survey[J]. ZTE Communications, 2024, 22 (3): 91- 98
9 BESL P J, MCKAY N D Method for registration of 3-D shapes[J]. Sensor Fusion IV: Control Paradigms and Data Structures, 1992, 1611: 586- 606
doi: 10.1109/cisp.2012.6469977
10 ZHAO R, ZHANG Y, WEN W, et al E-TPE: EfficientThumbnail: preserving encryption for privacy protection in visual sensor networks[J]. ACM Transactions on Sensor Networks, 2024, 20 (4): 1- 26
11 HANH N T, CUONG V D, TUNG D D, et al. H-LSHADE: an efficient hybrid approach for solving heterogeneous target coverage in visual sensor networks [C]//International Symposium on Information and Communication Technology. Singapore: Springer, 2024: 297–307.
12 XU D, ANGUELOV D, JAIN A. PointFusion: deep sensor fusion for 3D bounding box estimation [C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Salt Lake City: IEEE, 2018: 244–253.
13 WANG C, XU D, ZHU Y, et al. DenseFusion: 6D object pose estimation by iterative dense fusion [C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Long Beach: IEEE, 2020: 3338–3347.
14 WEN B, YANG W, KAUTZ J, et al. FoundationPose: unified 6D pose estimation and tracking of novel objects [C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle: IEEE, 2024: 17868–17879.
15 GOPAL G Y, AMER M A. Separable self and mixed attention transformers for efficient object tracking [C]//Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. Waikoloa: IEEE, 2024: 6708–6717.
16 LEI X, LU W, YONG J, et al RFF-PoseNet: a 6D object pose estimation network based on robust feature fusion in complex scenes[J]. Electronics, 2024, 13 (17): 3518
doi: 10.3390/electronics13173518
17 LV Z, GUO Y, YANG S, et al Multiscale feature fusion with self-attention for efficient 6D pose estimation[J]. Algorithms, 2025, 18 (4): 207
doi: 10.3390/a18040207
18 XIANG Y, SCHMIDT T, NARAYANAN V, et al. PoseCNN: a convolutional neural network for 6D object pose estimation in cluttered scenes [C]//Proceedings of the Robotics: Science and Systems XIV. Robotics: Science and Systems Foundation, 2018.
19 RAD M, LEPETIT V. BB8: a scalable, accurate, robust to partial occlusion method for predicting the 3D poses of challenging objects without using depth [C]//Proceedings of the IEEE International Conference on Computer Vision. Venice: IEEE, 2017: 3848–3856.
20 KEHL W, MANHARDT F, TOMBARI F, et al. SSD-6D: making RGB-based 3D detection and 6D pose estimation great again [C]//Proceedings of the IEEE International Conference on Computer Vision. Venice: IEEE, 2017: 1530–1538.
21 PENG S, LIU Y, HUANG Q, et al. PVNet: pixel-wise voting network for 6DoF pose estimation [C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Long Beach: IEEE, 2020: 4556–4565.
22 LEPETIT V, MORENO-NOGUER F, FUA P EPnP: an accurate O(n) solution to the PnP problem[J]. International Journal of Computer Vision, 2009, 81 (2): 155- 166
doi: 10.1007/s11263-008-0152-6
23 JIANG M, ZHANG L, WANG X, et al 6D object pose estimation based on cross-modality feature fusion[J]. Sensors, 2023, 23 (19): 8088
doi: 10.3390/s23198088
24 ZENG C, KWONG S. Dual swin-transformer based mutual interactive network for RGB-D salient object detection [EB/OL]. [2025-08-20]. https://arxiv.org/abs/2206.03105
25 WU P, LI W, YAN M 3D scene reconstruction based on improved ICP algorithm[J]. Microprocessors and Microsystems, 2020, 75: 103064
doi: 10.1016/j.micpro.2020.103064
26 WOHLHART P, LEPETIT V. Learning descriptors for object recognition and 3D pose estimation [C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Boston: IEEE, 2015: 3109–3118.
27 KENDALL A, GRIMES M, CIPOLLA R. PoseNet: a convolutional network for real-time 6-DOF camera relocalization [C]//Proceedings of the IEEE International Conference on Computer Vision. Santiago: IEEE, 2016: 2938–2946.
28 LI Y, WANG G, JI X, et al DeepIM: deep iterative matching for 6D pose estimation[J]. International Journal of Computer Vision, 2020, 128 (3): 657- 678
doi: 10.1007/s11263-019-01250-9
29 PARK K, PATTEN T, VINCZE M. Pix2Pose: pixel-wise coordinate regression of objects for 6D pose estimation [C]//Proceedings of the IEEE/CVF International Conference on Computer Vision. Seoul: IEEE, 2019: 7667–7676.
30 HOSSEINI JAFARI O, MUSTIKOVELA S K, PERTSCH K, et al. iPose: instance-aware 6D pose estimation of partly occluded objects [C]// 2018 Asian Conference on Computer Vision. Cham: Springer, 2019: 477–492.
31 HE Y, SUN W, HUANG H, et al. PVN3D: a deep point-wise 3D keypoints voting network for 6DoF pose estimation [C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle: IEEE, 2020: 11629–11638.
32 QI C R, YI L, SU H, et al. PointNet++: deep hierarchical feature learning on point sets in a metric space [C]//Advances in Neural Information Processing Systems. Long Beach: NIPS, 2017: 5100–5109.
33 HU J, SHEN L, SUN G. Squeeze-and-excitation networks [C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Salt Lake City: IEEE, 2018: 7132–7141.
34 WOO S, PARK J, LEE J Y, et al. CBAM: convolutional block attention module [C]//European Conference on Computer Vision. Munich: Springer, 2018: 3–19.
35 HINTERSTOISSER S, CAGNIAT C, IIII V L, et al. Multimodal templates for real-time detection of texture-less objects in heavily cluttered scenes [C]//International Conference on Computer Vision. Barcelona: IEEE, 2011: 858–865.
[1] 赵明华,吕雨萱,吕佳豪,陈逸飞,石程,胡静. 基于多尺度外观运动融合的视频异常检测[J]. 浙江大学学报(工学版), 2026, 60(8): 1730-1738.
[2] 范业博,李逸成,刘勇,张薇. 面向图神经网络谣言检测器的高效攻击模型[J]. 浙江大学学报(工学版), 2026, 60(8): 1739-1748.
[3] 刘慧雪,刘新倩,王帅,赵川,丁建国. 融合日志语义与图对抗学习的APT检测框架[J]. 浙江大学学报(工学版), 2026, 60(8): 1749-1759.
[4] 刘慧霞,王荣景,高宇轩,曹猛,盛苏英,马友鹏. 基于自适应神经网络的无信号交叉口车辆无死锁调度算法[J]. 浙江大学学报(工学版), 2026, 60(8): 1832-1840.
[5] 徐恺蔚,KHIZER BIN TALIBHafiz,曹衍龙,许源平,许志杰,宋景春. 基于光流和卷积视觉Transformer的轻量级微表情识别[J]. 浙江大学学报(工学版), 2026, 60(7): 1381-1391.
[6] 杨艳春,李佳龙. 基于多路协同与空谱先验的高光谱与多光谱图像融合[J]. 浙江大学学报(工学版), 2026, 60(7): 1427-1437.
[7] 苏挺,许月萍,WANGQuanjun,钟华,蒋建群. 基于物理信息的城市洪涝多阶段替代模型[J]. 浙江大学学报(工学版), 2026, 60(7): 1557-1566.
[8] 张乃洲,赵云超,曹薇,张啸剑. 基于多视图跨模态特征融合的图像描述生成[J]. 浙江大学学报(工学版), 2026, 60(6): 1205-1212.
[9] 李云红,张琪琪,陈锦妮,陈伟重,苏雪平,梁成名. 基于生成对抗网络和坐标注意力机制的文本生成图像算法[J]. 浙江大学学报(工学版), 2026, 60(6): 1213-1220.
[10] 郑文军,黎志昆,韩守飞. 知识增强图Transformer的方面级情感分析[J]. 浙江大学学报(工学版), 2026, 60(6): 1269-1276.
[11] 李国燕,于威,梅玉鹏,张明辉,王新强. 全局局部特征融合的遥感图像建筑物提取[J]. 浙江大学学报(工学版), 2026, 60(5): 1100-1108.
[12] 林洪彬,吕思进,王晨阳,蔡天放,骆鹏伟. 基于残差/梯度高斯自适应采样的径向基网络[J]. 浙江大学学报(工学版), 2026, 60(5): 1119-1127.
[13] 李子祥,陆克成,蔡海兵,解伟帅,张广东. 基于触觉仿生技术的黑暗环境混凝土裂缝检测[J]. 浙江大学学报(工学版), 2026, 60(5): 915-925.
[14] 吴杰,韩贝林,张舣航,邹超,辛莉峰,黄仕平. 微调稳定扩散模型的裂缝图像数据集扩充方法[J]. 浙江大学学报(工学版), 2026, 60(5): 926-934.
[15] 宋耀莲,彭驰,唐菁敏,赵宣植,虞贵财. 基于融合注意力机制的光学遥感图像小目标检测算法[J]. 浙江大学学报(工学版), 2026, 60(4): 763-771.