|
|
|
| Object pose estimation based on attention enhancement cross-modal multilayer fusion network |
Heng YANG( ),Shaohan WANG,Qing DONG,Keyuan ZHAO,Mingliang YANG |
| School of Mechanical Engineering, Taiyuan University of Science and Technology, Taiyuan 030024, China |
|
|
|
Abstract A robust end-to-end 6D pose estimation network was proposed. A hierarchical feature fusion mechanism was used to effectively integrate RGB-D data. The pixel-level feature was extracted via an attention enhancement module. Then a multi-scale network was employed to process the fused pixel-level multimodal feature, yielding multi-scale multimodal features. The multimodal features were subsequently fed into subsequent modules to obtain corresponding global features. The global features were concatenated with the multimodal features to form multi-scale dense discriminative features that encode global information. The final pose was derived from multi-scale dense discriminative features, enabling accurate 6D pose estimation. Extensive experiments on the LineMOD and YCB-Video datasets demonstrate that the proposed method achieves superior performance, particularly in handling challenging smooth and textureless objects in the YCB-Video dataset. The method outperforms the representative approach DenseFusion in terms of accuracy on the LineMOD dataset, while achieving significantly higher processing speed.
|
|
Received: 17 July 2025
Published: 20 July 2026
|
|
|
| Fund: 国家自然科学基金资助项目(52105269);省部级科技计划资助项目(202402150101006). |
基于注意力增强的跨模态多层融合网络的物体位姿估计
提出鲁棒的端到端6D位姿估计网络. 利用层次化特征融合机制,使得RGB-D数据有效融合. 网络通过注意力增强模块提取像素级特征,通过多尺度网络对融合后的像素级多模态特征进行处理,得到多尺度多模态特征. 将多模态特征分别送至后续模块中,得到对应的全局特征. 将多模态特征与全局特征进行拼接,形成包含全局信息的多尺度密集判别特征. 根据多尺度密集判别特征得到最终位姿,实现精确的6D位姿估计. 在LineMOD和YCB-Video数据集上开展多项实验,所提方法的表现出色,特别是在处理YCB-Video数据集中具有挑战性的光滑和无纹理物体时表现突出. 在LineMOD数据集上,所提方法的精度优于现有的代表性方法DenseFusion,处理速度显著提升.
关键词:
6D位姿,
深度学习,
神经网络,
跨模态融合,
注意力机制
|
|
| [1] |
JIN D, ZHANG R, LI Y, et al. Mitigating latency effects on subjective experience in robot teleoperation using a VR-enabled virtual spring [C]//Proceedings of the IEEE International Symposium on Mixed and Augmented Reality. Bellevue: IEEE, 2024: 1276–1282.
|
|
|
| [2] |
ZHAO Y, WANG B, YU H, et al. Research on the self-powered flexible non-contact distance/contact pressure dual-mode sensor for manipulator based on triboelectric nanogenerator [C]//Proceedings of the China Automation Congress. Qingdao: IEEE, 2024: 437–441.
|
|
|
| [3] |
PAWAR U S, KULKARNI R A. Navigation with augmented reality [C]//Proceedings of the 1st International Conference on AIML: Applications for Engineering and Technology. Pune: IEEE, 2025: 1–6.
|
|
|
| [4] |
DIAS P, MARQUES B, FELIX I, et al. Creating asynchronous augmented reality instructions through a virtual reality framework [C]//Proceedings of the IEEE Conference on Virtual Reality and 3D User Interfaces Abstracts and Workshops. Saint Malo: IEEE, 2025: 1065–1071.
|
|
|
| [5] |
KIM J, JEON J, PARK J, et al. Continuous performance improvement of infrastructure guidance service for autonomous cooperative driving: focusing on data-centric AI [C]//Proceedings of the International Conference on Electronics, Information, and Communication. Taipei: IEEE, 2024: 1–4.
|
|
|
| [6] |
FENG Y, SHEN H, SHAN Z, et al Semantic communication for edge intelligence enabled autonomous driving system[J]. IEEE Network, 2025, 39 (2): 149- 157
doi: 10.1109/MNET.2024.3468328
|
|
|
| [7] |
CHEN J, LEI B, SONG Q, et al. A hierarchical graph network for 3D object detection on point clouds [C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle: IEEE, 2020: 389–398.
|
|
|
| [8] |
LU P, SHI W, QIAO X Multi-view image-based 3D reconstruction in indoor scenes: a survey[J]. ZTE Communications, 2024, 22 (3): 91- 98
|
|
|
| [9] |
BESL P J, MCKAY N D Method for registration of 3-D shapes[J]. Sensor Fusion IV: Control Paradigms and Data Structures, 1992, 1611: 586- 606
doi: 10.1109/cisp.2012.6469977
|
|
|
| [10] |
ZHAO R, ZHANG Y, WEN W, et al E-TPE: EfficientThumbnail: preserving encryption for privacy protection in visual sensor networks[J]. ACM Transactions on Sensor Networks, 2024, 20 (4): 1- 26
|
|
|
| [11] |
HANH N T, CUONG V D, TUNG D D, et al. H-LSHADE: an efficient hybrid approach for solving heterogeneous target coverage in visual sensor networks [C]//International Symposium on Information and Communication Technology. Singapore: Springer, 2024: 297–307.
|
|
|
| [12] |
XU D, ANGUELOV D, JAIN A. PointFusion: deep sensor fusion for 3D bounding box estimation [C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Salt Lake City: IEEE, 2018: 244–253.
|
|
|
| [13] |
WANG C, XU D, ZHU Y, et al. DenseFusion: 6D object pose estimation by iterative dense fusion [C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Long Beach: IEEE, 2020: 3338–3347.
|
|
|
| [14] |
WEN B, YANG W, KAUTZ J, et al. FoundationPose: unified 6D pose estimation and tracking of novel objects [C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle: IEEE, 2024: 17868–17879.
|
|
|
| [15] |
GOPAL G Y, AMER M A. Separable self and mixed attention transformers for efficient object tracking [C]//Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. Waikoloa: IEEE, 2024: 6708–6717.
|
|
|
| [16] |
LEI X, LU W, YONG J, et al RFF-PoseNet: a 6D object pose estimation network based on robust feature fusion in complex scenes[J]. Electronics, 2024, 13 (17): 3518
doi: 10.3390/electronics13173518
|
|
|
| [17] |
LV Z, GUO Y, YANG S, et al Multiscale feature fusion with self-attention for efficient 6D pose estimation[J]. Algorithms, 2025, 18 (4): 207
doi: 10.3390/a18040207
|
|
|
| [18] |
XIANG Y, SCHMIDT T, NARAYANAN V, et al. PoseCNN: a convolutional neural network for 6D object pose estimation in cluttered scenes [C]//Proceedings of the Robotics: Science and Systems XIV. Robotics: Science and Systems Foundation, 2018.
|
|
|
| [19] |
RAD M, LEPETIT V. BB8: a scalable, accurate, robust to partial occlusion method for predicting the 3D poses of challenging objects without using depth [C]//Proceedings of the IEEE International Conference on Computer Vision. Venice: IEEE, 2017: 3848–3856.
|
|
|
| [20] |
KEHL W, MANHARDT F, TOMBARI F, et al. SSD-6D: making RGB-based 3D detection and 6D pose estimation great again [C]//Proceedings of the IEEE International Conference on Computer Vision. Venice: IEEE, 2017: 1530–1538.
|
|
|
| [21] |
PENG S, LIU Y, HUANG Q, et al. PVNet: pixel-wise voting network for 6DoF pose estimation [C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Long Beach: IEEE, 2020: 4556–4565.
|
|
|
| [22] |
LEPETIT V, MORENO-NOGUER F, FUA P EPnP: an accurate O(n) solution to the PnP problem[J]. International Journal of Computer Vision, 2009, 81 (2): 155- 166
doi: 10.1007/s11263-008-0152-6
|
|
|
| [23] |
JIANG M, ZHANG L, WANG X, et al 6D object pose estimation based on cross-modality feature fusion[J]. Sensors, 2023, 23 (19): 8088
doi: 10.3390/s23198088
|
|
|
| [24] |
ZENG C, KWONG S. Dual swin-transformer based mutual interactive network for RGB-D salient object detection [EB/OL]. [2025-08-20]. https://arxiv.org/abs/2206.03105
|
|
|
| [25] |
WU P, LI W, YAN M 3D scene reconstruction based on improved ICP algorithm[J]. Microprocessors and Microsystems, 2020, 75: 103064
doi: 10.1016/j.micpro.2020.103064
|
|
|
| [26] |
WOHLHART P, LEPETIT V. Learning descriptors for object recognition and 3D pose estimation [C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Boston: IEEE, 2015: 3109–3118.
|
|
|
| [27] |
KENDALL A, GRIMES M, CIPOLLA R. PoseNet: a convolutional network for real-time 6-DOF camera relocalization [C]//Proceedings of the IEEE International Conference on Computer Vision. Santiago: IEEE, 2016: 2938–2946.
|
|
|
| [28] |
LI Y, WANG G, JI X, et al DeepIM: deep iterative matching for 6D pose estimation[J]. International Journal of Computer Vision, 2020, 128 (3): 657- 678
doi: 10.1007/s11263-019-01250-9
|
|
|
| [29] |
PARK K, PATTEN T, VINCZE M. Pix2Pose: pixel-wise coordinate regression of objects for 6D pose estimation [C]//Proceedings of the IEEE/CVF International Conference on Computer Vision. Seoul: IEEE, 2019: 7667–7676.
|
|
|
| [30] |
HOSSEINI JAFARI O, MUSTIKOVELA S K, PERTSCH K, et al. iPose: instance-aware 6D pose estimation of partly occluded objects [C]// 2018 Asian Conference on Computer Vision. Cham: Springer, 2019: 477–492.
|
|
|
| [31] |
HE Y, SUN W, HUANG H, et al. PVN3D: a deep point-wise 3D keypoints voting network for 6DoF pose estimation [C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle: IEEE, 2020: 11629–11638.
|
|
|
| [32] |
QI C R, YI L, SU H, et al. PointNet++: deep hierarchical feature learning on point sets in a metric space [C]//Advances in Neural Information Processing Systems. Long Beach: NIPS, 2017: 5100–5109.
|
|
|
| [33] |
HU J, SHEN L, SUN G. Squeeze-and-excitation networks [C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Salt Lake City: IEEE, 2018: 7132–7141.
|
|
|
| [34] |
WOO S, PARK J, LEE J Y, et al. CBAM: convolutional block attention module [C]//European Conference on Computer Vision. Munich: Springer, 2018: 3–19.
|
|
|
| [35] |
HINTERSTOISSER S, CAGNIAT C, IIII V L, et al. Multimodal templates for real-time detection of texture-less objects in heavily cluttered scenes [C]//International Conference on Computer Vision. Barcelona: IEEE, 2011: 858–865.
|
|
|
|
Viewed |
|
|
|
Full text
|
|
|
|
|
Abstract
|
|
|
|
|
Cited |
|
|
|
|
| |
Shared |
|
|
|
|
| |
Discussed |
|
|
|
|