Please wait a minute...
浙江大学学报(工学版)  2026, Vol. 60 Issue (9): 1972-1979    DOI: 10.3785/j.issn.1008-973X.2026.09.014
计算机技术、自动控制技术     
大模型辅助的噪声鲁棒无监督骨架表示学习
蔡蕊1,2(),郭钦涵1,张子豪1,董建锋1,2,*(),张荣1,2,刘宝龙1,2,王勋1,2
1. 浙江工商大学 计算机科学与技术学院,浙江 杭州 310035
2. 浙江省大数据与未来电子商务技术重点实验室,浙江 杭州 310035
Large model-assisted noise-robust unsupervised skeleton representation learning
Rui CAI1,2(),Qinhan GUO1,Zihao ZHANG1,Jianfeng DONG1,2,*(),Rong ZHANG1,2,Baolong LIU1,2,Xun WANG1,2
1. School of Computer Science and Technology, Zhejiang Gongshang University, Hangzhou 310035, China
2. Zhejiang Key Laboratory of Big Data and Future E-Commerce Technology, Hangzhou 310035, China
 全文: PDF(2113 KB)   HTML
摘要:

针对现有方法难以捕捉细微动作差异(如剪纸、看手机)的问题,提出基于多模态大模型的噪声鲁棒无监督人体骨架表示学习方法,引入由多模态大模型生成的动作描述文本,为训练提供关键的动作信息. 为了提升多模态大模型的噪声鲁棒性,提出基于异常样本检测以及信息熵的联合噪声估计方法,为骨架表示学习提供高质量的辅助文本信号. 在3个公开数据集上的实验结果表明,文本辅助学习使模型在动作识别与检索下游任务中的性能显著优于HiCo等现有方法,并展现出较强的知识迁移能力,尤其在低资源场景下. 该方法具有良好的泛化性,适用于多数现有无监督骨架表示学习框架,其中的联合噪声估计方法在跨模态检索任务中也同样有效.

关键词: 人体骨架动作识别大模型无监督学习多模态数据噪声评估    
Abstract:

A noise-robust unsupervised human skeleton representation learning method based on multimodal large models was proposed to address the challenge that existing methods struggled to capture subtle differences in actions (e.g., paper cutting and looking at a phone). The action description texts generated by multimodal large models were incorporated to provide key action information for training. A joint noise estimation approach based on outlier sample detection and information entropy was developed to provide high-quality auxiliary textual signals for skeleton representation learning and improve the noise robustness of multimodal large model. The experimental results on three public datasets showed that text-assisted learning enabled the model to significantly outperform the existing methods in downstream tasks of action recognition and retrieval, while demonstrating strong knowledge transfer capability of the model, particularly in low-resource scenarios. The method has good generalization ability and is suitable for most existing unsupervised skeleton representation learning frameworks, while its joint noise estimation method is also effective in cross-modal retrieval tasks.

Key words: human skeleton-based action recognition    large model    unsupervised learning    multimodal data    noise estimation
收稿日期: 2025-07-16 出版日期: 2026-07-20
CLC:  TP 391  
基金资助: 国家自然科学基金资助项目(62306278,62472385,62302449);浙江省自然科学基金重点资助项目(LZ23F020004);中国科协青年人才托举工程资助项目(2022QNRC001);浙江省自然科学基金资助项目(LQ23F020008,LQ23F020009);浙江省“尖兵”“领雁”研发攻关计划资助项目(2024C01110);浙江省属高校基本科研业务费专项资金资助项目(FR2402ZD).
通讯作者: 董建锋     E-mail: cairuics@gmail.com;dongjf24@gmail.com
作者简介: 蔡蕊(1992—),男,副研究员,从事自然语言处理相关研究. orcid.org/0009-0006-2823-3051. E-mail:cairuics@gmail.com
服务  
把本文推荐给朋友
加入引用管理器
E-mail Alert
作者相关文章  
蔡蕊
郭钦涵
张子豪
董建锋
张荣
刘宝龙
王勋

引用本文:

蔡蕊,郭钦涵,张子豪,董建锋,张荣,刘宝龙,王勋. 大模型辅助的噪声鲁棒无监督骨架表示学习[J]. 浙江大学学报(工学版), 2026, 60(9): 1972-1979.

Rui CAI,Qinhan GUO,Zihao ZHANG,Jianfeng DONG,Rong ZHANG,Baolong LIU,Xun WANG. Large model-assisted noise-robust unsupervised skeleton representation learning. Journal of ZheJiang University (Engineering Science), 2026, 60(9): 1972-1979.

链接本文:

https://www.zjujournals.com/eng/CN/10.3785/j.issn.1008-973X.2026.09.014        https://www.zjujournals.com/eng/CN/Y2026/V60/I9/1972

图 1  基于多模态大模型的噪声鲁棒无监督骨架表示学习框架图
图 2  基于异常样本检测的噪声估计方法示意图
图 3  基于信息熵的噪声估计方法示意图
方法A/%
N60
x-sub
N60
x-view
N120
x-sub
N120
x-setup
P-II
x-sub
AimCLR++[21]77.281.565.567.841.2
ViA[22]78.185.869.266.9
3s-C-F[23]79.187.269.270.849.8
HaLP[24]79.786.871.172.243.5
HiCo[2]81.188.672.874.149.4
HSTM-Transformer[17]81.890.473.574.646.9
UmURL[25]82.389.873.574.352.1
Skeleton-logoCLR[26]82.487.272.873.554.7
KTCL[3]82.489.474.474.555.5
RMMD[4]83.090.575.275.850.6
PCM3[27]83.990.476.577.551.5
Lin等[28]83.990.375.777.2
C2VL[14]84.489.876.078.752.6
本研究方法85.892.578.179.654.2
表 1  NTU RGB+D 60、NTU RGB+D 120以及PKU-MMD II数据集上所提方法与其他方法的骨架动作识别性能比较
方法A/%
N60
x-sub
N60
x-view
N120
x-sub
N120
x-setup
KTCL[3]67.184.758.959.8
HiCo[2]68.384.856.659.1
HSTM-Transformer[17]70.088.357.759.8
RMMD[4]70.987.458.461.8
HaLP[24]65.886.355.859.0
UmURL[25]71.388.358.560.9
本研究方法73.590.261.865.5
表 2  NTU RGB+D 60和NTU RGB+D 120数据集上所提方法与其他方法的骨架动作检索性能对比
方法A/%
rs=1%rs=10%rv=1%rv=10%
3s-C-F[23]52.376.553.181.3
PCM3[27]53.877.753.182.8
HiCo[2]54.473.054.878.3
Lin等[28]55.278.954.982.9
UmURL[25]58.176.558.381.5
CMD[29]50.675.453.080.2
RMMD[4]58.377.156.081.0
本研究方法62.179.963.384.5
表 3  半监督学习性能比较
方法A/%
N60N120
HaLP[24]54.855.4
RMMD[4]55.0
HSTM-Transformer[17]55.655.5
HiCo[2]56.355.4
Lin等[28]56.557.8
3s-C-F[23]58.1
CMD[29]56.057.0
UmURL[25]58.257.6
KTCL[3]58.4
本研究方法59.559.0
表 4  迁移到PKU-MMD II数据集上的性能对比
方法A/%
rt=0%rt=30%rt=60%rt=90%rt=100%
基线模型78.778.778.778.778.7
本研究方法79.679.379.078.878.7
表 5  NTU RGB+D 120数据集上所提方法与基线模型对文本噪声的稳定性分析
骨架模态文本模态A/%
x-subx-setup
77.378.7
71.874.3
78.179.6
表 6  2类模态数据的消融实验
文本t方法1方法2A/%
x-subx-setup
77.378.7
76.678.2
77.278.9
77.679.3
78.179.6
表 7  NTU RGB+D 120数据集上的噪声估计方法消融实验
1 FEI H, WU S, ZHANG M, et al Enhancing video-language representations with structural spatio-temporal alignment[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024, 46 (12): 7701- 7719
doi: 10.1109/TPAMI.2024.3393452
2 DONG J, SUN S, LIU Z, et al. Hierarchical contrast for unsupervised skeleton-based action representation learning [C]// Proceedings of the AAAI Conference on Artificial Intelligence. Washington, D. C. : AAAI Press, 2023: 525–533.
3 WANG X, MU Y Localized linear temporal dynamics for self-supervised skeleton action recognition[J]. IEEE Transactions on Multimedia, 2024, 26: 10189- 10199
doi: 10.1109/TMM.2024.3405712
4 HE Z, LV J, FANG S Representation modeling learning with multi-domain decoupling for unsupervised skeleton-based action recognition[J]. Neurocomputing, 2024, 582: 127495
doi: 10.1016/j.neucom.2024.127495
5 WU W, HUA Y, ZHENG C, et al. Skeletonmae: spatial-temporal masked autoencoders for self-supervised skeleton action recognition [C]// Proceedings of the IEEE International Conference on Multimedia and Expo Workshops. Brisbane: IEEE, 2023: 224–229.
6 ABDELFATTAH M, ALAHI A. S-JEPA: a joint embedding predictive architecture for skeletal action recognition [C]// Proceedings of the European Conference on Computer Vision. Milan: Springer, 2024: 367–384.
7 RAO H, XU S, HU X, et al Augmented skeleton based contrastive action learning with momentum LSTM for unsupervised action recognition[J]. Information Sciences, 2021, 569: 90- 109
doi: 10.1016/j.ins.2021.04.023
8 QU H, CAI Y, LIU J. LLMs are good action recognizers [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle: IEEE, 2024: 18395–18406.
9 XIANG W, LI C, ZHOU Y, et al. Generative action description prompts for skeleton-based action recognition [C]// Proceedings of the IEEE/CVF International Conference on Computer Vision. Paris: IEEE, 2023: 10242–10251.
10 HU E J, SHEN Y, WALLIS P, et al. LoRA: low-rank adaptation of large language models [EB/OL]. (2021-10-16) [2025-07-10]. https://arxiv.org/abs/2106.09685.
11 ZHU A, KE Q, GONG M, et al. Part-aware unified representation of language and skeleton for zero-shot action recognition [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle: IEEE, 2024: 18761–18770.
12 LIU S, ZENG Z, REN T, et al. Grounding DINO: marrying DINO with grounded pre-training for open-set object detection [C]// Proceedings of the European Conference on Computer Vision. Milan: Springer, 2024: 38–55.
13 LIU H, LI C, WU Q, et al. Visual instruction tuning [C]// Proceedings of the 37th International Conference on Neural Information Processing Systems. New Orleans: Curran Associates Inc, 2023: 34892–34916.
14 CHEN Y, HE T, FU J, et al Vision-language meets the skeleton: progressively distillation with cross-modal knowledge for 3D action representation learning[J]. IEEE Transactions on Multimedia, 2025, 27: 2293- 2303
doi: 10.1109/TMM.2024.3521718
15 CHEN T, KORNBLITH S, NOROUZI M, et al. A simple framework for contrastive learning of visual representations [EB/OL]. (2020-07-01) [2025-07-10]. https://arxiv.org/abs/2002.05709.
16 ANDONIAN A, CHEN S, HAMID R. Robust cross-modal representation learning with progressive self-distillation [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. New Orleans: IEEE, 2022: 16409–16420.
17 CAO W, ZHANG A, HE Z, et al Hierarchical spatial-temporal masked contrast for skeleton action recognition[J]. IEEE Transactions on Artificial Intelligence, 2024, 5 (11): 5801- 5814
doi: 10.1109/TAI.2024.3430260
18 SHAHROUDY A, LIU J, NG T T, et al. NTU RGB+D: a large scale dataset for 3D human activity analysis [C]// Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Las Vegas: IEEE, 2016: 1010–1019.
19 LIU J, SHAHROUDY A, PEREZ M, et al NTU RGB+ D 120: a large-scale benchmark for 3D human activity understanding[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2020, 42 (10): 2684- 2701
doi: 10.1109/TPAMI.2019.2916873
20 LIU C, HU Y, LI Y, et al. PKU-MMD: a large scale benchmark for continuous multi-modal human action understanding [EB/OL]. (2017-03-28) [2025-07-10]. https://arxiv.org/abs/1703.07475.
21 GUO T, LIU M, LIU H, et al Improving self-supervised action recognition from extremely augmented skeleton sequences[J]. Pattern Recognition, 2024, 150: 110333
doi: 10.1016/j.patcog.2024.110333
22 YANG D, WANG Y, DANTCHEVA A, et al View-invariant skeleton action representation learning via motion retargeting[J]. International Journal of Computer Vision, 2024, 132 (7): 2351- 2366
doi: 10.1007/s11263-023-01967-8
23 YANG S, LIU J, LU S, et al Self-supervised 3D action representation learning with skeleton cloud colorization[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024, 46 (1): 509- 524
doi: 10.1109/TPAMI.2023.3325463
24 SHAH A, ROY A, SHAH K, et al. HaLP: hallucinating latent positives for skeleton-based self-supervised learning of actions [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Vancouver: IEEE, 2023: 18846–18856.
25 SUN S, LIU D, DONG J, et al. Unified multi-modal unsupervised representation learning for skeleton-based action understanding [C]// Proceedings of the 31st ACM International Conference on Multimedia. Ottawa: ACM, 2023: 2973–2984.
26 HU J, HOU Y, GUO Z, et al Global and local contrastive learning for self-supervised skeleton-based action recognition[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2024, 34 (11): 10578- 10589
doi: 10.1109/TCSVT.2024.3410301
27 ZHANG J, LIN L, LIU J. Prompted contrast with masked motion modeling: towards versatile 3D action representation learning [C]// Proceedings of the 31st ACM International Conference on Multimedia. Ottawa: ACM, 2023: 7175–7183.
28 LIN L, ZHANG J, LIU J Mutual information driven equivariant contrastive learning for 3D action representation learning[J]. IEEE Transactions on Image Processing, 2024, 33: 1883- 1897
doi: 10.1109/TIP.2024.3372451
[1] 赵明华,吕雨萱,吕佳豪,陈逸飞,石程,胡静. 基于多尺度外观运动融合的视频异常检测[J]. 浙江大学学报(工学版), 2026, 60(8): 1730-1738.
[2] 张雪梅,孙颖,张雪英. 基于无监督图对比学习的语音情感识别[J]. 浙江大学学报(工学版), 2026, 60(4): 782-790.
[3] 张凌浩,谭海波,赵赫,陈中,程昊天,马志宇. CompuDEX:基于区块链的大模型微调算力共享平台[J]. 浙江大学学报(工学版), 2026, 60(1): 1-18.
[4] 王光明,柏正尧,宋帅,徐月娥. 阿尔茨海默病辅助诊断的多模态数据融合轻量级网络[J]. 浙江大学学报(工学版), 2025, 59(1): 39-48.
[5] 唐善成,逯建辉,张莹,金子成,赵安新. 修复缺陷嫌疑区域的无监督磁瓦表面缺陷检测[J]. 浙江大学学报(工学版), 2024, 58(4): 718-728.
[6] 杨军,李金泰,高志明. 无监督的三维模型簇对应关系协同计算[J]. 浙江大学学报(工学版), 2022, 56(10): 1935-1947.
[7] 张鹏,田子都,王浩. 基于改进生成对抗网络的飞参数据异常检测方法[J]. 浙江大学学报(工学版), 2022, 56(10): 1967-1976.