|
|
|
| Large model-assisted noise-robust unsupervised skeleton representation learning |
Rui CAI1,2( ),Qinhan GUO1,Zihao ZHANG1,Jianfeng DONG1,2,*( ),Rong ZHANG1,2,Baolong LIU1,2,Xun WANG1,2 |
1. School of Computer Science and Technology, Zhejiang Gongshang University, Hangzhou 310035, China 2. Zhejiang Key Laboratory of Big Data and Future E-Commerce Technology, Hangzhou 310035, China |
|
|
|
Abstract A noise-robust unsupervised human skeleton representation learning method based on multimodal large models was proposed to address the challenge that existing methods struggled to capture subtle differences in actions (e.g., paper cutting and looking at a phone). The action description texts generated by multimodal large models were incorporated to provide key action information for training. A joint noise estimation approach based on outlier sample detection and information entropy was developed to provide high-quality auxiliary textual signals for skeleton representation learning and improve the noise robustness of multimodal large model. The experimental results on three public datasets showed that text-assisted learning enabled the model to significantly outperform the existing methods in downstream tasks of action recognition and retrieval, while demonstrating strong knowledge transfer capability of the model, particularly in low-resource scenarios. The method has good generalization ability and is suitable for most existing unsupervised skeleton representation learning frameworks, while its joint noise estimation method is also effective in cross-modal retrieval tasks.
|
|
Received: 16 July 2025
Published: 20 July 2026
|
|
|
| Fund: 国家自然科学基金资助项目(62306278,62472385,62302449);浙江省自然科学基金重点资助项目(LZ23F020004);中国科协青年人才托举工程资助项目(2022QNRC001);浙江省自然科学基金资助项目(LQ23F020008,LQ23F020009);浙江省“尖兵”“领雁”研发攻关计划资助项目(2024C01110);浙江省属高校基本科研业务费专项资金资助项目(FR2402ZD). |
|
Corresponding Authors:
Jianfeng DONG
E-mail: cairuics@gmail.com;dongjf24@gmail.com
|
大模型辅助的噪声鲁棒无监督骨架表示学习
针对现有方法难以捕捉细微动作差异(如剪纸、看手机)的问题,提出基于多模态大模型的噪声鲁棒无监督人体骨架表示学习方法,引入由多模态大模型生成的动作描述文本,为训练提供关键的动作信息. 为了提升多模态大模型的噪声鲁棒性,提出基于异常样本检测以及信息熵的联合噪声估计方法,为骨架表示学习提供高质量的辅助文本信号. 在3个公开数据集上的实验结果表明,文本辅助学习使模型在动作识别与检索下游任务中的性能显著优于HiCo等现有方法,并展现出较强的知识迁移能力,尤其在低资源场景下. 该方法具有良好的泛化性,适用于多数现有无监督骨架表示学习框架,其中的联合噪声估计方法在跨模态检索任务中也同样有效.
关键词:
人体骨架动作识别,
大模型,
无监督学习,
多模态数据,
噪声评估
|
|
| [1] |
FEI H, WU S, ZHANG M, et al Enhancing video-language representations with structural spatio-temporal alignment[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024, 46 (12): 7701- 7719
doi: 10.1109/TPAMI.2024.3393452
|
|
|
| [2] |
DONG J, SUN S, LIU Z, et al. Hierarchical contrast for unsupervised skeleton-based action representation learning [C]// Proceedings of the AAAI Conference on Artificial Intelligence. Washington, D. C. : AAAI Press, 2023: 525–533.
|
|
|
| [3] |
WANG X, MU Y Localized linear temporal dynamics for self-supervised skeleton action recognition[J]. IEEE Transactions on Multimedia, 2024, 26: 10189- 10199
doi: 10.1109/TMM.2024.3405712
|
|
|
| [4] |
HE Z, LV J, FANG S Representation modeling learning with multi-domain decoupling for unsupervised skeleton-based action recognition[J]. Neurocomputing, 2024, 582: 127495
doi: 10.1016/j.neucom.2024.127495
|
|
|
| [5] |
WU W, HUA Y, ZHENG C, et al. Skeletonmae: spatial-temporal masked autoencoders for self-supervised skeleton action recognition [C]// Proceedings of the IEEE International Conference on Multimedia and Expo Workshops. Brisbane: IEEE, 2023: 224–229.
|
|
|
| [6] |
ABDELFATTAH M, ALAHI A. S-JEPA: a joint embedding predictive architecture for skeletal action recognition [C]// Proceedings of the European Conference on Computer Vision. Milan: Springer, 2024: 367–384.
|
|
|
| [7] |
RAO H, XU S, HU X, et al Augmented skeleton based contrastive action learning with momentum LSTM for unsupervised action recognition[J]. Information Sciences, 2021, 569: 90- 109
doi: 10.1016/j.ins.2021.04.023
|
|
|
| [8] |
QU H, CAI Y, LIU J. LLMs are good action recognizers [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle: IEEE, 2024: 18395–18406.
|
|
|
| [9] |
XIANG W, LI C, ZHOU Y, et al. Generative action description prompts for skeleton-based action recognition [C]// Proceedings of the IEEE/CVF International Conference on Computer Vision. Paris: IEEE, 2023: 10242–10251.
|
|
|
| [10] |
HU E J, SHEN Y, WALLIS P, et al. LoRA: low-rank adaptation of large language models [EB/OL]. (2021-10-16) [2025-07-10]. https://arxiv.org/abs/2106.09685.
|
|
|
| [11] |
ZHU A, KE Q, GONG M, et al. Part-aware unified representation of language and skeleton for zero-shot action recognition [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle: IEEE, 2024: 18761–18770.
|
|
|
| [12] |
LIU S, ZENG Z, REN T, et al. Grounding DINO: marrying DINO with grounded pre-training for open-set object detection [C]// Proceedings of the European Conference on Computer Vision. Milan: Springer, 2024: 38–55.
|
|
|
| [13] |
LIU H, LI C, WU Q, et al. Visual instruction tuning [C]// Proceedings of the 37th International Conference on Neural Information Processing Systems. New Orleans: Curran Associates Inc, 2023: 34892–34916.
|
|
|
| [14] |
CHEN Y, HE T, FU J, et al Vision-language meets the skeleton: progressively distillation with cross-modal knowledge for 3D action representation learning[J]. IEEE Transactions on Multimedia, 2025, 27: 2293- 2303
doi: 10.1109/TMM.2024.3521718
|
|
|
| [15] |
CHEN T, KORNBLITH S, NOROUZI M, et al. A simple framework for contrastive learning of visual representations [EB/OL]. (2020-07-01) [2025-07-10]. https://arxiv.org/abs/2002.05709.
|
|
|
| [16] |
ANDONIAN A, CHEN S, HAMID R. Robust cross-modal representation learning with progressive self-distillation [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. New Orleans: IEEE, 2022: 16409–16420.
|
|
|
| [17] |
CAO W, ZHANG A, HE Z, et al Hierarchical spatial-temporal masked contrast for skeleton action recognition[J]. IEEE Transactions on Artificial Intelligence, 2024, 5 (11): 5801- 5814
doi: 10.1109/TAI.2024.3430260
|
|
|
| [18] |
SHAHROUDY A, LIU J, NG T T, et al. NTU RGB+D: a large scale dataset for 3D human activity analysis [C]// Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Las Vegas: IEEE, 2016: 1010–1019.
|
|
|
| [19] |
LIU J, SHAHROUDY A, PEREZ M, et al NTU RGB+ D 120: a large-scale benchmark for 3D human activity understanding[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2020, 42 (10): 2684- 2701
doi: 10.1109/TPAMI.2019.2916873
|
|
|
| [20] |
LIU C, HU Y, LI Y, et al. PKU-MMD: a large scale benchmark for continuous multi-modal human action understanding [EB/OL]. (2017-03-28) [2025-07-10]. https://arxiv.org/abs/1703.07475.
|
|
|
| [21] |
GUO T, LIU M, LIU H, et al Improving self-supervised action recognition from extremely augmented skeleton sequences[J]. Pattern Recognition, 2024, 150: 110333
doi: 10.1016/j.patcog.2024.110333
|
|
|
| [22] |
YANG D, WANG Y, DANTCHEVA A, et al View-invariant skeleton action representation learning via motion retargeting[J]. International Journal of Computer Vision, 2024, 132 (7): 2351- 2366
doi: 10.1007/s11263-023-01967-8
|
|
|
| [23] |
YANG S, LIU J, LU S, et al Self-supervised 3D action representation learning with skeleton cloud colorization[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024, 46 (1): 509- 524
doi: 10.1109/TPAMI.2023.3325463
|
|
|
| [24] |
SHAH A, ROY A, SHAH K, et al. HaLP: hallucinating latent positives for skeleton-based self-supervised learning of actions [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Vancouver: IEEE, 2023: 18846–18856.
|
|
|
| [25] |
SUN S, LIU D, DONG J, et al. Unified multi-modal unsupervised representation learning for skeleton-based action understanding [C]// Proceedings of the 31st ACM International Conference on Multimedia. Ottawa: ACM, 2023: 2973–2984.
|
|
|
| [26] |
HU J, HOU Y, GUO Z, et al Global and local contrastive learning for self-supervised skeleton-based action recognition[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2024, 34 (11): 10578- 10589
doi: 10.1109/TCSVT.2024.3410301
|
|
|
| [27] |
ZHANG J, LIN L, LIU J. Prompted contrast with masked motion modeling: towards versatile 3D action representation learning [C]// Proceedings of the 31st ACM International Conference on Multimedia. Ottawa: ACM, 2023: 7175–7183.
|
|
|
| [28] |
LIN L, ZHANG J, LIU J Mutual information driven equivariant contrastive learning for 3D action representation learning[J]. IEEE Transactions on Image Processing, 2024, 33: 1883- 1897
doi: 10.1109/TIP.2024.3372451
|
|
|
|
Viewed |
|
|
|
Full text
|
|
|
|
|
Abstract
|
|
|
|
|
Cited |
|
|
|
|
| |
Shared |
|
|
|
|
| |
Discussed |
|
|
|
|