Please wait a minute...
Journal of ZheJiang University (Engineering Science)  2026, Vol. 60 Issue (9): 1972-1979    DOI: 10.3785/j.issn.1008-973X.2026.09.014
    
Large model-assisted noise-robust unsupervised skeleton representation learning
Rui CAI1,2(),Qinhan GUO1,Zihao ZHANG1,Jianfeng DONG1,2,*(),Rong ZHANG1,2,Baolong LIU1,2,Xun WANG1,2
1. School of Computer Science and Technology, Zhejiang Gongshang University, Hangzhou 310035, China
2. Zhejiang Key Laboratory of Big Data and Future E-Commerce Technology, Hangzhou 310035, China
Download: HTML     PDF(2113KB) HTML
Export: BibTeX | EndNote (RIS)      

Abstract  

A noise-robust unsupervised human skeleton representation learning method based on multimodal large models was proposed to address the challenge that existing methods struggled to capture subtle differences in actions (e.g., paper cutting and looking at a phone). The action description texts generated by multimodal large models were incorporated to provide key action information for training. A joint noise estimation approach based on outlier sample detection and information entropy was developed to provide high-quality auxiliary textual signals for skeleton representation learning and improve the noise robustness of multimodal large model. The experimental results on three public datasets showed that text-assisted learning enabled the model to significantly outperform the existing methods in downstream tasks of action recognition and retrieval, while demonstrating strong knowledge transfer capability of the model, particularly in low-resource scenarios. The method has good generalization ability and is suitable for most existing unsupervised skeleton representation learning frameworks, while its joint noise estimation method is also effective in cross-modal retrieval tasks.



Key wordshuman skeleton-based action recognition      large model      unsupervised learning      multimodal data      noise estimation     
Received: 16 July 2025      Published: 20 July 2026
CLC:  TP 391  
Fund:  国家自然科学基金资助项目(62306278,62472385,62302449);浙江省自然科学基金重点资助项目(LZ23F020004);中国科协青年人才托举工程资助项目(2022QNRC001);浙江省自然科学基金资助项目(LQ23F020008,LQ23F020009);浙江省“尖兵”“领雁”研发攻关计划资助项目(2024C01110);浙江省属高校基本科研业务费专项资金资助项目(FR2402ZD).
Corresponding Authors: Jianfeng DONG     E-mail: cairuics@gmail.com;dongjf24@gmail.com
Cite this article:

Rui CAI,Qinhan GUO,Zihao ZHANG,Jianfeng DONG,Rong ZHANG,Baolong LIU,Xun WANG. Large model-assisted noise-robust unsupervised skeleton representation learning. Journal of ZheJiang University (Engineering Science), 2026, 60(9): 1972-1979.

URL:

https://www.zjujournals.com/eng/10.3785/j.issn.1008-973X.2026.09.014     OR     https://www.zjujournals.com/eng/Y2026/V60/I9/1972


大模型辅助的噪声鲁棒无监督骨架表示学习

针对现有方法难以捕捉细微动作差异(如剪纸、看手机)的问题,提出基于多模态大模型的噪声鲁棒无监督人体骨架表示学习方法,引入由多模态大模型生成的动作描述文本,为训练提供关键的动作信息. 为了提升多模态大模型的噪声鲁棒性,提出基于异常样本检测以及信息熵的联合噪声估计方法,为骨架表示学习提供高质量的辅助文本信号. 在3个公开数据集上的实验结果表明,文本辅助学习使模型在动作识别与检索下游任务中的性能显著优于HiCo等现有方法,并展现出较强的知识迁移能力,尤其在低资源场景下. 该方法具有良好的泛化性,适用于多数现有无监督骨架表示学习框架,其中的联合噪声估计方法在跨模态检索任务中也同样有效.


关键词: 人体骨架动作识别,  大模型,  无监督学习,  多模态数据,  噪声评估 
Fig.1 Framework diagram of noise-robust unsupervised skeleton representation learning based on multi-modal large model
Fig.2 Illustration of noise estimation method based on outlier sample detection
Fig.3 Illustration of noise estimation method based on information entropy
方法A/%
N60
x-sub
N60
x-view
N120
x-sub
N120
x-setup
P-II
x-sub
AimCLR++[21]77.281.565.567.841.2
ViA[22]78.185.869.266.9
3s-C-F[23]79.187.269.270.849.8
HaLP[24]79.786.871.172.243.5
HiCo[2]81.188.672.874.149.4
HSTM-Transformer[17]81.890.473.574.646.9
UmURL[25]82.389.873.574.352.1
Skeleton-logoCLR[26]82.487.272.873.554.7
KTCL[3]82.489.474.474.555.5
RMMD[4]83.090.575.275.850.6
PCM3[27]83.990.476.577.551.5
Lin等[28]83.990.375.777.2
C2VL[14]84.489.876.078.752.6
本研究方法85.892.578.179.654.2
Tab.1 Comparison of skeleton-based action recognition performance of proposed method and other methods on NTU RGB+D 60, NTU RGB+D 120 and PKU-MMD II datasets
方法A/%
N60
x-sub
N60
x-view
N120
x-sub
N120
x-setup
KTCL[3]67.184.758.959.8
HiCo[2]68.384.856.659.1
HSTM-Transformer[17]70.088.357.759.8
RMMD[4]70.987.458.461.8
HaLP[24]65.886.355.859.0
UmURL[25]71.388.358.560.9
本研究方法73.590.261.865.5
Tab.2 Comparison of skeleton-based action retrieval performance of proposed method and other methods on NTU RGB+D 60 and NTU RGB+D 120 datasets
方法A/%
rs=1%rs=10%rv=1%rv=10%
3s-C-F[23]52.376.553.181.3
PCM3[27]53.877.753.182.8
HiCo[2]54.473.054.878.3
Lin等[28]55.278.954.982.9
UmURL[25]58.176.558.381.5
CMD[29]50.675.453.080.2
RMMD[4]58.377.156.081.0
本研究方法62.179.963.384.5
Tab.3 Comparison of semi-supervised learning performance
方法A/%
N60N120
HaLP[24]54.855.4
RMMD[4]55.0
HSTM-Transformer[17]55.655.5
HiCo[2]56.355.4
Lin等[28]56.557.8
3s-C-F[23]58.1
CMD[29]56.057.0
UmURL[25]58.257.6
KTCL[3]58.4
本研究方法59.559.0
Tab.4 Performance comparison when transferred to PKU-MMD II dataset
方法A/%
rt=0%rt=30%rt=60%rt=90%rt=100%
基线模型78.778.778.778.778.7
本研究方法79.679.379.078.878.7
Tab.5 Stability analysis of proposed method and baseline model against textual noise on NTU RGB+D 120 dataset
骨架模态文本模态A/%
x-subx-setup
77.378.7
71.874.3
78.179.6
Tab.6 Ablation experiment on data of two modalities
文本t方法1方法2A/%
x-subx-setup
77.378.7
76.678.2
77.278.9
77.679.3
78.179.6
Tab.7 Ablation experiment on noise estimation methods for NTU RGB+D 120 dataset
[1]   FEI H, WU S, ZHANG M, et al Enhancing video-language representations with structural spatio-temporal alignment[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024, 46 (12): 7701- 7719
doi: 10.1109/TPAMI.2024.3393452
[2]   DONG J, SUN S, LIU Z, et al. Hierarchical contrast for unsupervised skeleton-based action representation learning [C]// Proceedings of the AAAI Conference on Artificial Intelligence. Washington, D. C. : AAAI Press, 2023: 525–533.
[3]   WANG X, MU Y Localized linear temporal dynamics for self-supervised skeleton action recognition[J]. IEEE Transactions on Multimedia, 2024, 26: 10189- 10199
doi: 10.1109/TMM.2024.3405712
[4]   HE Z, LV J, FANG S Representation modeling learning with multi-domain decoupling for unsupervised skeleton-based action recognition[J]. Neurocomputing, 2024, 582: 127495
doi: 10.1016/j.neucom.2024.127495
[5]   WU W, HUA Y, ZHENG C, et al. Skeletonmae: spatial-temporal masked autoencoders for self-supervised skeleton action recognition [C]// Proceedings of the IEEE International Conference on Multimedia and Expo Workshops. Brisbane: IEEE, 2023: 224–229.
[6]   ABDELFATTAH M, ALAHI A. S-JEPA: a joint embedding predictive architecture for skeletal action recognition [C]// Proceedings of the European Conference on Computer Vision. Milan: Springer, 2024: 367–384.
[7]   RAO H, XU S, HU X, et al Augmented skeleton based contrastive action learning with momentum LSTM for unsupervised action recognition[J]. Information Sciences, 2021, 569: 90- 109
doi: 10.1016/j.ins.2021.04.023
[8]   QU H, CAI Y, LIU J. LLMs are good action recognizers [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle: IEEE, 2024: 18395–18406.
[9]   XIANG W, LI C, ZHOU Y, et al. Generative action description prompts for skeleton-based action recognition [C]// Proceedings of the IEEE/CVF International Conference on Computer Vision. Paris: IEEE, 2023: 10242–10251.
[10]   HU E J, SHEN Y, WALLIS P, et al. LoRA: low-rank adaptation of large language models [EB/OL]. (2021-10-16) [2025-07-10]. https://arxiv.org/abs/2106.09685.
[11]   ZHU A, KE Q, GONG M, et al. Part-aware unified representation of language and skeleton for zero-shot action recognition [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle: IEEE, 2024: 18761–18770.
[12]   LIU S, ZENG Z, REN T, et al. Grounding DINO: marrying DINO with grounded pre-training for open-set object detection [C]// Proceedings of the European Conference on Computer Vision. Milan: Springer, 2024: 38–55.
[13]   LIU H, LI C, WU Q, et al. Visual instruction tuning [C]// Proceedings of the 37th International Conference on Neural Information Processing Systems. New Orleans: Curran Associates Inc, 2023: 34892–34916.
[14]   CHEN Y, HE T, FU J, et al Vision-language meets the skeleton: progressively distillation with cross-modal knowledge for 3D action representation learning[J]. IEEE Transactions on Multimedia, 2025, 27: 2293- 2303
doi: 10.1109/TMM.2024.3521718
[15]   CHEN T, KORNBLITH S, NOROUZI M, et al. A simple framework for contrastive learning of visual representations [EB/OL]. (2020-07-01) [2025-07-10]. https://arxiv.org/abs/2002.05709.
[16]   ANDONIAN A, CHEN S, HAMID R. Robust cross-modal representation learning with progressive self-distillation [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. New Orleans: IEEE, 2022: 16409–16420.
[17]   CAO W, ZHANG A, HE Z, et al Hierarchical spatial-temporal masked contrast for skeleton action recognition[J]. IEEE Transactions on Artificial Intelligence, 2024, 5 (11): 5801- 5814
doi: 10.1109/TAI.2024.3430260
[18]   SHAHROUDY A, LIU J, NG T T, et al. NTU RGB+D: a large scale dataset for 3D human activity analysis [C]// Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Las Vegas: IEEE, 2016: 1010–1019.
[19]   LIU J, SHAHROUDY A, PEREZ M, et al NTU RGB+ D 120: a large-scale benchmark for 3D human activity understanding[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2020, 42 (10): 2684- 2701
doi: 10.1109/TPAMI.2019.2916873
[20]   LIU C, HU Y, LI Y, et al. PKU-MMD: a large scale benchmark for continuous multi-modal human action understanding [EB/OL]. (2017-03-28) [2025-07-10]. https://arxiv.org/abs/1703.07475.
[21]   GUO T, LIU M, LIU H, et al Improving self-supervised action recognition from extremely augmented skeleton sequences[J]. Pattern Recognition, 2024, 150: 110333
doi: 10.1016/j.patcog.2024.110333
[22]   YANG D, WANG Y, DANTCHEVA A, et al View-invariant skeleton action representation learning via motion retargeting[J]. International Journal of Computer Vision, 2024, 132 (7): 2351- 2366
doi: 10.1007/s11263-023-01967-8
[23]   YANG S, LIU J, LU S, et al Self-supervised 3D action representation learning with skeleton cloud colorization[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024, 46 (1): 509- 524
doi: 10.1109/TPAMI.2023.3325463
[24]   SHAH A, ROY A, SHAH K, et al. HaLP: hallucinating latent positives for skeleton-based self-supervised learning of actions [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Vancouver: IEEE, 2023: 18846–18856.
[25]   SUN S, LIU D, DONG J, et al. Unified multi-modal unsupervised representation learning for skeleton-based action understanding [C]// Proceedings of the 31st ACM International Conference on Multimedia. Ottawa: ACM, 2023: 2973–2984.
[26]   HU J, HOU Y, GUO Z, et al Global and local contrastive learning for self-supervised skeleton-based action recognition[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2024, 34 (11): 10578- 10589
doi: 10.1109/TCSVT.2024.3410301
[27]   ZHANG J, LIN L, LIU J. Prompted contrast with masked motion modeling: towards versatile 3D action representation learning [C]// Proceedings of the 31st ACM International Conference on Multimedia. Ottawa: ACM, 2023: 7175–7183.
[28]   LIN L, ZHANG J, LIU J Mutual information driven equivariant contrastive learning for 3D action representation learning[J]. IEEE Transactions on Image Processing, 2024, 33: 1883- 1897
doi: 10.1109/TIP.2024.3372451
[1] Minghua ZHAO,Yuxuan LYU,Jiahao LYU,Yifei CHEN,Cheng SHI,Jing HU. Video anomaly detection based on multi-scale appearance and motion fusion[J]. Journal of ZheJiang University (Engineering Science), 2026, 60(8): 1730-1738.
[2] Xuemei ZHANG,Ying SUN,Xueying ZHANG. Speech emotion recognition with unsupervised graph contrastive learning[J]. Journal of ZheJiang University (Engineering Science), 2026, 60(4): 782-790.
[3] Linghao ZHANG,Haibo TAN,He ZHAO,Zhong CHEN,Haotian CHENG,Zhiyu MA. CompuDEX: blockchain-based large model fine-tuning compute-power sharing platform[J]. Journal of ZheJiang University (Engineering Science), 2026, 60(1): 1-18.
[4] Guangming WANG,Zhengyao BAI,Shuai SONG,Yue’e XU. Lightweight multimodal data fusion network for auxiliary diagnosis of Alzheimer’s disease[J]. Journal of ZheJiang University (Engineering Science), 2025, 59(1): 39-48.
[5] Shancheng TANG,Jianhui LU,Ying ZHANG,Zicheng JIN,Anxin ZHAO. Unsupervised surface defect detection of magnetic tile for repair of suspected area defects[J]. Journal of ZheJiang University (Engineering Science), 2024, 58(4): 718-728.
[6] Jun YANG,Jin-tai LI,Zhi-ming GAO. Unsupervised co-calculation on correspondence of three-dimensional shape collections[J]. Journal of ZheJiang University (Engineering Science), 2022, 56(10): 1935-1947.
[7] Peng ZHANG,Zi-du TIAN,Hao WANG. Flight parameter data anomaly detection method based on improved generative adversarial network[J]. Journal of ZheJiang University (Engineering Science), 2022, 56(10): 1967-1976.
[8] WU Ping, CHEN Liang, ZHOU Wei, GUO Ling-ling. Online subspace identification based on principal component analysis and noise estimation[J]. Journal of ZheJiang University (Engineering Science), 2018, 52(9): 1694-1701.
[9] CHEN Xiao-hong,WANG Wei-dong. A HDTV video de-noising algorithm based on spatial-temporal filtering[J]. Journal of ZheJiang University (Engineering Science), 2013, 47(5): 853-859.