Please wait a minute...
Journal of ZheJiang University (Engineering Science)  2026, Vol. 60 Issue (9): 1872-1880    DOI: 10.3785/j.issn.1008-973X.2026.09.004
    
Object pose estimation based on attention enhancement cross-modal multilayer fusion network
Heng YANG(),Shaohan WANG,Qing DONG,Keyuan ZHAO,Mingliang YANG
School of Mechanical Engineering, Taiyuan University of Science and Technology, Taiyuan 030024, China
Download: HTML     PDF(1443KB) HTML
Export: BibTeX | EndNote (RIS)      

Abstract  

A robust end-to-end 6D pose estimation network was proposed. A hierarchical feature fusion mechanism was used to effectively integrate RGB-D data. The pixel-level feature was extracted via an attention enhancement module. Then a multi-scale network was employed to process the fused pixel-level multimodal feature, yielding multi-scale multimodal features. The multimodal features were subsequently fed into subsequent modules to obtain corresponding global features. The global features were concatenated with the multimodal features to form multi-scale dense discriminative features that encode global information. The final pose was derived from multi-scale dense discriminative features, enabling accurate 6D pose estimation. Extensive experiments on the LineMOD and YCB-Video datasets demonstrate that the proposed method achieves superior performance, particularly in handling challenging smooth and textureless objects in the YCB-Video dataset. The method outperforms the representative approach DenseFusion in terms of accuracy on the LineMOD dataset, while achieving significantly higher processing speed.



Key words6D pose      deep learning      neural network      cross-modal fusion      attention mechanism     
Received: 17 July 2025      Published: 20 July 2026
CLC:  TP 183  
Fund:  国家自然科学基金资助项目(52105269);省部级科技计划资助项目(202402150101006).
Cite this article:

Heng YANG,Shaohan WANG,Qing DONG,Keyuan ZHAO,Mingliang YANG. Object pose estimation based on attention enhancement cross-modal multilayer fusion network. Journal of ZheJiang University (Engineering Science), 2026, 60(9): 1872-1880.

URL:

https://www.zjujournals.com/eng/10.3785/j.issn.1008-973X.2026.09.004     OR     https://www.zjujournals.com/eng/Y2026/V60/I9/1872


基于注意力增强的跨模态多层融合网络的物体位姿估计

提出鲁棒的端到端6D位姿估计网络. 利用层次化特征融合机制,使得RGB-D数据有效融合. 网络通过注意力增强模块提取像素级特征,通过多尺度网络对融合后的像素级多模态特征进行处理,得到多尺度多模态特征. 将多模态特征分别送至后续模块中,得到对应的全局特征. 将多模态特征与全局特征进行拼接,形成包含全局信息的多尺度密集判别特征. 根据多尺度密集判别特征得到最终位姿,实现精确的6D位姿估计. 在LineMOD和YCB-Video数据集上开展多项实验,所提方法的表现出色,特别是在处理YCB-Video数据集中具有挑战性的光滑和无纹理物体时表现突出. 在LineMOD数据集上,所提方法的精度优于现有的代表性方法DenseFusion,处理速度显著提升.


关键词: 6D位姿,  深度学习,  神经网络,  跨模态融合,  注意力机制 
Fig.1 Extracting pixel-level feature from RGB-D image
Fig.2 Overall structure of DeepFusion6D network
Fig.3 Overall structure of AFE
Fig.4 Output size of each module in extracted multi-scale dense discriminative features
方法输入数据ADD(-S)分数/%
ApeBenCameraCanCatDrillerDuckEggboxGlueHoleIronLampPhone平均值
BB8[19]RGB40.291.755.863.762.973.744.358.141.067.284.776.954.062.3
PVNet[21]43.699.986.995.579.396.452.699.295.781.998.999.392.486.3
DeepIM[28]77.297.893.796.882.095.077.897.199.352.898.398.087.989.0
TexPose[36]80.999.094.899.792.697.483.494.993.479.399.898.378.991.7
SSD-6D+ICP[20]RGB-D65.080.078.086.070.073.066.010010049.078.073.079.076.7
CFFM+CWTM[23]88.291.593.493.894.491.689.497.399.189.395.794.293.793.2
Per-pixel Densefusion[13]79.884.176.786.888.977.876.210099.079.092.192.088.086.1
Iterative-Densefusion[13]92.193.294.393.196.587.092.010010092.197.095.192.894.3
本文方法(仅使用多尺度模块)93.889.690.590.291.190.389.510010084.793.191.589.089.4
本文方法(仅使用注意力增强模块)89.390.189.789.290.289.888.710098.489.390.289.790.089.0
本文方法95.493.195.993.295.094.290.310010092.296.595.194.895.2
Tab.1 Quantitative assessment of ADD(-S) score on linenmod dataset
方法ADD(-S)分数/%
master-
chef-
can
cracker-
box
sugar-
box
tomato-
soup-
can
mustard-
bottle
tuna-
fish-
can
pudding-
box
gelatin-
box
potted-
meat-
can
Bananapitcher-
base
bleach-
cleaner
BowlMugpower-
drill
wood-
block
Scissorlarge-
marker
large-
clamp
extra-
large-
clamp
foam-
brick
平均值
PointFusion[12]99.862.295.496.984.099.896.710088.570.579.865.024.199.822.818.235.980.450.020.110074.1
PVN3D[31]10091.610096.910010010010093.699.710099.454.999.899.680.295.699.774.948.810093.2
DenseFusion[13]10099.510096.910010010010091.310010010098.810098.794.610010079.276.310096.8
CFFM+CWTM[23]10095.310095.710099.093.910090.893.410097.853.298.297.481.394.796.768.362.599.693.5
FoundationPose[14]99.198.799.398.999.599.098.899.498.599.298.698.398.098.798.497.597.198.296.996.599.199.0
本文方法
(仅使用多
尺度模块)
98.593.198.089.194.397.393.999.992.495.197.493.692.194.194.693.291.096.288.183.299.494.3
本文方法
(仅使用注意
力增强模块)
99.194.998.790.695.797.894.710093.796.398.195.193.595.395.894.792.597.590.385.899.795.2
本文方法10097.599.093.597.598.596.510096.097.598.097.099.297.096.598.096.598.085.082.010097.1
方法AUC/%
master-
chef-
can
cracker-
box
sugar-
box
tomato-
soup-
can
mustard-
bottle
tuna-
fish-
can
pudding-
box
gelatin-
box
potted-
meat-
can
Bananapitcher-
base
bleach-
cleaner
BowlMugpower-
drill
wood-
block
Scissorlarge-
marker
large-
clamp
extra-
large-
clamp
foam-
brick
平均值
PointFusion[12]90.980.590.491.988.593.887.595.086.484.785.581.075.794.271.568.176.787.965.960.491.883.9
PVN3D[31]95.892.798.294.598.697.197.998.892.797.197.896.981.095.098.287.691.797.275.264.497.293.0
DenseFusion[13]96.495.597.594.697.296.696.598.191.396.697.195.888.297.196.089.795.297.572.969.892.593.1
CFFM+CWTM[23]94.392.298.793.997.194.893.399.191.691.195.286.583.890.594.884.392.888.673.262.393.390.5
FoundationPose[14]96.594.296.895.097.496.194.597.093.896.394.092.591.894.393.189.588.292.087.585.495.993.2
本文方法
(仅使用多
尺度模块)
93.984.393.780.990.891.389.297.885.987.993.689.479.589.187.285.478.789.373.568.794.088.5
本文方法
(仅使用注意
力增强模块)
94.786.794.482.792.092.790.398.388.289.794.891.082.790.688.987.782.391.075.271.595.589.7
本文方法96.998.698.695.096.594.593.199.090.391.895.593.488.392.590.510085.598.482.580.597.093.7
Tab.2 Quantitative evaluation of ADD(-S) score and ADD(-S) AUC on YCB-Video dataset
模型配置Np/106FLOPs/109tin/msAUC/%ADD(-S)
分数/%
Baseline78.035.916.483.590.1
+Multi-scale91.046.221.585.892.5
+AFE(Attention)81.037.517.688.995.3
完整模型(本文方法)94.047.922.393.797.1
Tab.3 Ablation study on efficiency and accuracy on YCB-Video dataset
Fig.5 ADD(-S) score-threshold curve for 'large_clamp' and 'extra_large_clamp' split by PoseCNN and Ground Truth
Fig.6 Performance change of different methods on YCB-Video dataset with increasing level of occlusion
Fig.7 Visualization of DeepFusion6D’s estimated 6D object pose on LINEMOD dataset
[1]   JIN D, ZHANG R, LI Y, et al. Mitigating latency effects on subjective experience in robot teleoperation using a VR-enabled virtual spring [C]//Proceedings of the IEEE International Symposium on Mixed and Augmented Reality. Bellevue: IEEE, 2024: 1276–1282.
[2]   ZHAO Y, WANG B, YU H, et al. Research on the self-powered flexible non-contact distance/contact pressure dual-mode sensor for manipulator based on triboelectric nanogenerator [C]//Proceedings of the China Automation Congress. Qingdao: IEEE, 2024: 437–441.
[3]   PAWAR U S, KULKARNI R A. Navigation with augmented reality [C]//Proceedings of the 1st International Conference on AIML: Applications for Engineering and Technology. Pune: IEEE, 2025: 1–6.
[4]   DIAS P, MARQUES B, FELIX I, et al. Creating asynchronous augmented reality instructions through a virtual reality framework [C]//Proceedings of the IEEE Conference on Virtual Reality and 3D User Interfaces Abstracts and Workshops. Saint Malo: IEEE, 2025: 1065–1071.
[5]   KIM J, JEON J, PARK J, et al. Continuous performance improvement of infrastructure guidance service for autonomous cooperative driving: focusing on data-centric AI [C]//Proceedings of the International Conference on Electronics, Information, and Communication. Taipei: IEEE, 2024: 1–4.
[6]   FENG Y, SHEN H, SHAN Z, et al Semantic communication for edge intelligence enabled autonomous driving system[J]. IEEE Network, 2025, 39 (2): 149- 157
doi: 10.1109/MNET.2024.3468328
[7]   CHEN J, LEI B, SONG Q, et al. A hierarchical graph network for 3D object detection on point clouds [C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle: IEEE, 2020: 389–398.
[8]   LU P, SHI W, QIAO X Multi-view image-based 3D reconstruction in indoor scenes: a survey[J]. ZTE Communications, 2024, 22 (3): 91- 98
[9]   BESL P J, MCKAY N D Method for registration of 3-D shapes[J]. Sensor Fusion IV: Control Paradigms and Data Structures, 1992, 1611: 586- 606
doi: 10.1109/cisp.2012.6469977
[10]   ZHAO R, ZHANG Y, WEN W, et al E-TPE: EfficientThumbnail: preserving encryption for privacy protection in visual sensor networks[J]. ACM Transactions on Sensor Networks, 2024, 20 (4): 1- 26
[11]   HANH N T, CUONG V D, TUNG D D, et al. H-LSHADE: an efficient hybrid approach for solving heterogeneous target coverage in visual sensor networks [C]//International Symposium on Information and Communication Technology. Singapore: Springer, 2024: 297–307.
[12]   XU D, ANGUELOV D, JAIN A. PointFusion: deep sensor fusion for 3D bounding box estimation [C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Salt Lake City: IEEE, 2018: 244–253.
[13]   WANG C, XU D, ZHU Y, et al. DenseFusion: 6D object pose estimation by iterative dense fusion [C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Long Beach: IEEE, 2020: 3338–3347.
[14]   WEN B, YANG W, KAUTZ J, et al. FoundationPose: unified 6D pose estimation and tracking of novel objects [C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle: IEEE, 2024: 17868–17879.
[15]   GOPAL G Y, AMER M A. Separable self and mixed attention transformers for efficient object tracking [C]//Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. Waikoloa: IEEE, 2024: 6708–6717.
[16]   LEI X, LU W, YONG J, et al RFF-PoseNet: a 6D object pose estimation network based on robust feature fusion in complex scenes[J]. Electronics, 2024, 13 (17): 3518
doi: 10.3390/electronics13173518
[17]   LV Z, GUO Y, YANG S, et al Multiscale feature fusion with self-attention for efficient 6D pose estimation[J]. Algorithms, 2025, 18 (4): 207
doi: 10.3390/a18040207
[18]   XIANG Y, SCHMIDT T, NARAYANAN V, et al. PoseCNN: a convolutional neural network for 6D object pose estimation in cluttered scenes [C]//Proceedings of the Robotics: Science and Systems XIV. Robotics: Science and Systems Foundation, 2018.
[19]   RAD M, LEPETIT V. BB8: a scalable, accurate, robust to partial occlusion method for predicting the 3D poses of challenging objects without using depth [C]//Proceedings of the IEEE International Conference on Computer Vision. Venice: IEEE, 2017: 3848–3856.
[20]   KEHL W, MANHARDT F, TOMBARI F, et al. SSD-6D: making RGB-based 3D detection and 6D pose estimation great again [C]//Proceedings of the IEEE International Conference on Computer Vision. Venice: IEEE, 2017: 1530–1538.
[21]   PENG S, LIU Y, HUANG Q, et al. PVNet: pixel-wise voting network for 6DoF pose estimation [C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Long Beach: IEEE, 2020: 4556–4565.
[22]   LEPETIT V, MORENO-NOGUER F, FUA P EPnP: an accurate O(n) solution to the PnP problem[J]. International Journal of Computer Vision, 2009, 81 (2): 155- 166
doi: 10.1007/s11263-008-0152-6
[23]   JIANG M, ZHANG L, WANG X, et al 6D object pose estimation based on cross-modality feature fusion[J]. Sensors, 2023, 23 (19): 8088
doi: 10.3390/s23198088
[24]   ZENG C, KWONG S. Dual swin-transformer based mutual interactive network for RGB-D salient object detection [EB/OL]. [2025-08-20]. https://arxiv.org/abs/2206.03105
[25]   WU P, LI W, YAN M 3D scene reconstruction based on improved ICP algorithm[J]. Microprocessors and Microsystems, 2020, 75: 103064
doi: 10.1016/j.micpro.2020.103064
[26]   WOHLHART P, LEPETIT V. Learning descriptors for object recognition and 3D pose estimation [C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Boston: IEEE, 2015: 3109–3118.
[27]   KENDALL A, GRIMES M, CIPOLLA R. PoseNet: a convolutional network for real-time 6-DOF camera relocalization [C]//Proceedings of the IEEE International Conference on Computer Vision. Santiago: IEEE, 2016: 2938–2946.
[28]   LI Y, WANG G, JI X, et al DeepIM: deep iterative matching for 6D pose estimation[J]. International Journal of Computer Vision, 2020, 128 (3): 657- 678
doi: 10.1007/s11263-019-01250-9
[29]   PARK K, PATTEN T, VINCZE M. Pix2Pose: pixel-wise coordinate regression of objects for 6D pose estimation [C]//Proceedings of the IEEE/CVF International Conference on Computer Vision. Seoul: IEEE, 2019: 7667–7676.
[30]   HOSSEINI JAFARI O, MUSTIKOVELA S K, PERTSCH K, et al. iPose: instance-aware 6D pose estimation of partly occluded objects [C]// 2018 Asian Conference on Computer Vision. Cham: Springer, 2019: 477–492.
[31]   HE Y, SUN W, HUANG H, et al. PVN3D: a deep point-wise 3D keypoints voting network for 6DoF pose estimation [C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle: IEEE, 2020: 11629–11638.
[32]   QI C R, YI L, SU H, et al. PointNet++: deep hierarchical feature learning on point sets in a metric space [C]//Advances in Neural Information Processing Systems. Long Beach: NIPS, 2017: 5100–5109.
[33]   HU J, SHEN L, SUN G. Squeeze-and-excitation networks [C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Salt Lake City: IEEE, 2018: 7132–7141.
[34]   WOO S, PARK J, LEE J Y, et al. CBAM: convolutional block attention module [C]//European Conference on Computer Vision. Munich: Springer, 2018: 3–19.
[35]   HINTERSTOISSER S, CAGNIAT C, IIII V L, et al. Multimodal templates for real-time detection of texture-less objects in heavily cluttered scenes [C]//International Conference on Computer Vision. Barcelona: IEEE, 2011: 858–865.
[1] Minghua ZHAO,Yuxuan LYU,Jiahao LYU,Yifei CHEN,Cheng SHI,Jing HU. Video anomaly detection based on multi-scale appearance and motion fusion[J]. Journal of ZheJiang University (Engineering Science), 2026, 60(8): 1730-1738.
[2] Yebo FAN,Yicheng LI,Yong LIU,Wei ZHANG. Efficient attack model targeting GNN-based rumor detectors[J]. Journal of ZheJiang University (Engineering Science), 2026, 60(8): 1739-1748.
[3] Huixue LIU,Xinqian LIU,Shuai WANG,Chuan ZHAO,Jianguo DING. Graph-adversarial learning framework integrating log semantics for APT detection[J]. Journal of ZheJiang University (Engineering Science), 2026, 60(8): 1749-1759.
[4] Huixia LIU,Rongjing WANG,Yuxuan GAO,Meng CAO,Suying SHENG,Youpeng MA. Adaptive neural network-based deadlock-free vehicle scheduling algorithm for unsignalized intersections[J]. Journal of ZheJiang University (Engineering Science), 2026, 60(8): 1832-1840.
[5] Kaiwei XU,Hafiz KHIZER BIN TALIB,Yanlong CAO,Yuanping XU,Zhijie XU,Jingchun SONG. Lightweight micro-expression recognition based on optical flow and convolutional vision Transformer[J]. Journal of ZheJiang University (Engineering Science), 2026, 60(7): 1381-1391.
[6] Yanchun YANG,Jialong LI. Multi-path collaboration-based and spatial-spectral prior-based hyperspectral and multispectral image fusion[J]. Journal of ZheJiang University (Engineering Science), 2026, 60(7): 1427-1437.
[7] Ting SU,Yueping XU,Quanjun WANG,Hua ZHONG,Jianqun JIANG. Physics-informed multi-stage surrogate model for urban flooding[J]. Journal of ZheJiang University (Engineering Science), 2026, 60(7): 1557-1566.
[8] Naizhou ZHANG,Yunchao ZHAO,Wei CAO,Xiaojian ZHANG. Image captioning generation based on multiple-view cross-modal feature fusion[J]. Journal of ZheJiang University (Engineering Science), 2026, 60(6): 1205-1212.
[9] Yunhong LI,Qiqi ZHANG,Jinni CHEN,Weichong CHEN,Xueping SU,Chengming LIANG. Text-to-image generation algorithm based on generative adversarial network and coordinate attention mechanism[J]. Journal of ZheJiang University (Engineering Science), 2026, 60(6): 1213-1220.
[10] Wenjun ZHENG,Zhikun LI,Shoufei HAN. Aspect-based sentiment analysis via knowledge-enhanced graph Transformer[J]. Journal of ZheJiang University (Engineering Science), 2026, 60(6): 1269-1276.
[11] Guoyan LI,Wei YU,Yupeng MEI,Minghui ZHANG,Xinqiang WANG. Building extraction from remote sensing images with global-local feature fusion[J]. Journal of ZheJiang University (Engineering Science), 2026, 60(5): 1100-1108.
[12] Hongbin LIN,Sijin LV,Chenyang WANG,Tianfang CAI,Pengwei LUO. Radial basis network based on residual/gradient Gaussian adaptive sampling[J]. Journal of ZheJiang University (Engineering Science), 2026, 60(5): 1119-1127.
[13] Zixiang LI,Kecheng LU,Haibing CAI,Weishuai XIE,Guangdong ZHANG. Concrete crack detection in dark environments based on biomimetic tactile technology[J]. Journal of ZheJiang University (Engineering Science), 2026, 60(5): 915-925.
[14] Jie WU,Beilin HAN,Yihang ZHANG,Chao ZOU,Lifeng XIN,Shiping HUANG. Method for augmenting crack image datasets via fine-tuning of stable diffusion models[J]. Journal of ZheJiang University (Engineering Science), 2026, 60(5): 926-934.
[15] Yaolian SONG,Chi PENG,Jingmin TANG,Xuanzhi ZHAO,Guicai YU. Small object detection algorithm for optical remote sensing images based on fusion attention mechanism[J]. Journal of ZheJiang University (Engineering Science), 2026, 60(4): 763-771.