| 计算机技术、自动控制技术 |
|
|
|
|
| 大模型辅助的噪声鲁棒无监督骨架表示学习 |
蔡蕊1,2( ),郭钦涵1,张子豪1,董建锋1,2,*( ),张荣1,2,刘宝龙1,2,王勋1,2 |
1. 浙江工商大学 计算机科学与技术学院,浙江 杭州 310035 2. 浙江省大数据与未来电子商务技术重点实验室,浙江 杭州 310035 |
|
| Large model-assisted noise-robust unsupervised skeleton representation learning |
Rui CAI1,2( ),Qinhan GUO1,Zihao ZHANG1,Jianfeng DONG1,2,*( ),Rong ZHANG1,2,Baolong LIU1,2,Xun WANG1,2 |
1. School of Computer Science and Technology, Zhejiang Gongshang University, Hangzhou 310035, China 2. Zhejiang Key Laboratory of Big Data and Future E-Commerce Technology, Hangzhou 310035, China |
引用本文:
蔡蕊,郭钦涵,张子豪,董建锋,张荣,刘宝龙,王勋. 大模型辅助的噪声鲁棒无监督骨架表示学习[J]. 浙江大学学报(工学版), 2026, 60(9): 1972-1979.
Rui CAI,Qinhan GUO,Zihao ZHANG,Jianfeng DONG,Rong ZHANG,Baolong LIU,Xun WANG. Large model-assisted noise-robust unsupervised skeleton representation learning. Journal of ZheJiang University (Engineering Science), 2026, 60(9): 1972-1979.
链接本文:
https://www.zjujournals.com/eng/CN/10.3785/j.issn.1008-973X.2026.09.014
或
https://www.zjujournals.com/eng/CN/Y2026/V60/I9/1972
|
| 1 |
FEI H, WU S, ZHANG M, et al Enhancing video-language representations with structural spatio-temporal alignment[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024, 46 (12): 7701- 7719
doi: 10.1109/TPAMI.2024.3393452
|
| 2 |
DONG J, SUN S, LIU Z, et al. Hierarchical contrast for unsupervised skeleton-based action representation learning [C]// Proceedings of the AAAI Conference on Artificial Intelligence. Washington, D. C. : AAAI Press, 2023: 525–533.
|
| 3 |
WANG X, MU Y Localized linear temporal dynamics for self-supervised skeleton action recognition[J]. IEEE Transactions on Multimedia, 2024, 26: 10189- 10199
doi: 10.1109/TMM.2024.3405712
|
| 4 |
HE Z, LV J, FANG S Representation modeling learning with multi-domain decoupling for unsupervised skeleton-based action recognition[J]. Neurocomputing, 2024, 582: 127495
doi: 10.1016/j.neucom.2024.127495
|
| 5 |
WU W, HUA Y, ZHENG C, et al. Skeletonmae: spatial-temporal masked autoencoders for self-supervised skeleton action recognition [C]// Proceedings of the IEEE International Conference on Multimedia and Expo Workshops. Brisbane: IEEE, 2023: 224–229.
|
| 6 |
ABDELFATTAH M, ALAHI A. S-JEPA: a joint embedding predictive architecture for skeletal action recognition [C]// Proceedings of the European Conference on Computer Vision. Milan: Springer, 2024: 367–384.
|
| 7 |
RAO H, XU S, HU X, et al Augmented skeleton based contrastive action learning with momentum LSTM for unsupervised action recognition[J]. Information Sciences, 2021, 569: 90- 109
doi: 10.1016/j.ins.2021.04.023
|
| 8 |
QU H, CAI Y, LIU J. LLMs are good action recognizers [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle: IEEE, 2024: 18395–18406.
|
| 9 |
XIANG W, LI C, ZHOU Y, et al. Generative action description prompts for skeleton-based action recognition [C]// Proceedings of the IEEE/CVF International Conference on Computer Vision. Paris: IEEE, 2023: 10242–10251.
|
| 10 |
HU E J, SHEN Y, WALLIS P, et al. LoRA: low-rank adaptation of large language models [EB/OL]. (2021-10-16) [2025-07-10]. https://arxiv.org/abs/2106.09685.
|
| 11 |
ZHU A, KE Q, GONG M, et al. Part-aware unified representation of language and skeleton for zero-shot action recognition [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle: IEEE, 2024: 18761–18770.
|
| 12 |
LIU S, ZENG Z, REN T, et al. Grounding DINO: marrying DINO with grounded pre-training for open-set object detection [C]// Proceedings of the European Conference on Computer Vision. Milan: Springer, 2024: 38–55.
|
| 13 |
LIU H, LI C, WU Q, et al. Visual instruction tuning [C]// Proceedings of the 37th International Conference on Neural Information Processing Systems. New Orleans: Curran Associates Inc, 2023: 34892–34916.
|
| 14 |
CHEN Y, HE T, FU J, et al Vision-language meets the skeleton: progressively distillation with cross-modal knowledge for 3D action representation learning[J]. IEEE Transactions on Multimedia, 2025, 27: 2293- 2303
doi: 10.1109/TMM.2024.3521718
|
| 15 |
CHEN T, KORNBLITH S, NOROUZI M, et al. A simple framework for contrastive learning of visual representations [EB/OL]. (2020-07-01) [2025-07-10]. https://arxiv.org/abs/2002.05709.
|
| 16 |
ANDONIAN A, CHEN S, HAMID R. Robust cross-modal representation learning with progressive self-distillation [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. New Orleans: IEEE, 2022: 16409–16420.
|
| 17 |
CAO W, ZHANG A, HE Z, et al Hierarchical spatial-temporal masked contrast for skeleton action recognition[J]. IEEE Transactions on Artificial Intelligence, 2024, 5 (11): 5801- 5814
doi: 10.1109/TAI.2024.3430260
|
| 18 |
SHAHROUDY A, LIU J, NG T T, et al. NTU RGB+D: a large scale dataset for 3D human activity analysis [C]// Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Las Vegas: IEEE, 2016: 1010–1019.
|
| 19 |
LIU J, SHAHROUDY A, PEREZ M, et al NTU RGB+ D 120: a large-scale benchmark for 3D human activity understanding[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2020, 42 (10): 2684- 2701
doi: 10.1109/TPAMI.2019.2916873
|
| 20 |
LIU C, HU Y, LI Y, et al. PKU-MMD: a large scale benchmark for continuous multi-modal human action understanding [EB/OL]. (2017-03-28) [2025-07-10]. https://arxiv.org/abs/1703.07475.
|
| 21 |
GUO T, LIU M, LIU H, et al Improving self-supervised action recognition from extremely augmented skeleton sequences[J]. Pattern Recognition, 2024, 150: 110333
doi: 10.1016/j.patcog.2024.110333
|
| 22 |
YANG D, WANG Y, DANTCHEVA A, et al View-invariant skeleton action representation learning via motion retargeting[J]. International Journal of Computer Vision, 2024, 132 (7): 2351- 2366
doi: 10.1007/s11263-023-01967-8
|
| 23 |
YANG S, LIU J, LU S, et al Self-supervised 3D action representation learning with skeleton cloud colorization[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024, 46 (1): 509- 524
doi: 10.1109/TPAMI.2023.3325463
|
| 24 |
SHAH A, ROY A, SHAH K, et al. HaLP: hallucinating latent positives for skeleton-based self-supervised learning of actions [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Vancouver: IEEE, 2023: 18846–18856.
|
| 25 |
SUN S, LIU D, DONG J, et al. Unified multi-modal unsupervised representation learning for skeleton-based action understanding [C]// Proceedings of the 31st ACM International Conference on Multimedia. Ottawa: ACM, 2023: 2973–2984.
|
| 26 |
HU J, HOU Y, GUO Z, et al Global and local contrastive learning for self-supervised skeleton-based action recognition[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2024, 34 (11): 10578- 10589
doi: 10.1109/TCSVT.2024.3410301
|
| 27 |
ZHANG J, LIN L, LIU J. Prompted contrast with masked motion modeling: towards versatile 3D action representation learning [C]// Proceedings of the 31st ACM International Conference on Multimedia. Ottawa: ACM, 2023: 7175–7183.
|
| 28 |
LIN L, ZHANG J, LIU J Mutual information driven equivariant contrastive learning for 3D action representation learning[J]. IEEE Transactions on Image Processing, 2024, 33: 1883- 1897
doi: 10.1109/TIP.2024.3372451
|
|
Viewed |
|
|
|
Full text
|
|
|
|
|
Abstract
|
|
|
|
|
Cited |
|
|
|
|
| |
Shared |
|
|
|
|
| |
Discussed |
|
|
|
|