|
|
|
| Multimodal gait recognition based on pose induction |
Wei GAO1( ),Zhidong YU1,Yuanyuan CHANG2,Wang MIAO1,Yongliang YIN2,Hongxia DENG1,*( ) |
1. College of Computer Science and Technology (College of Data Science), Taiyuan University of Technology, Taiyuan 030024, China 2. School of Physical Education and Health Engineering, Taiyuan University of Technology, Taiyuan 030024, China |
|
|
|
Abstract A pose-induced multimodal gait recognition method was proposed and a new multimodal dataset named CASIA-V containing real 3D human poses was constructed, in order to achieve multi-view fine-grained gait recognition in complex environments. The dataset was collected using multiple synchronized devices, capturing 93 subjects walking under three conditions (normal, carrying a bag, and wearing a coat) from 11 different viewpoints. It contained 11 253 RGB sequences along with corresponding motion-capture poses and depth maps. At the methodological level, a spatiotemporal alignment strategy was introduced to match motion-capture poses with video-based poses. A separately trainable 3D Pose-Induced Module (3D-SIM) was designed to extract real pose semantics, and a dual-branch model named SIGait was constructed to fuse pose and silhouette features. The experiment showed that, under the cross-view evaluation protocol of the CASIA-V dataset (excluding samples from the same view), SIGait achieved a Rank-1 accuracy of 97.7% in the normal walking scenario. Compared with the baseline method MSAFF, SIGait improved the Rank-1 accuracy by 5.1 and 7.1 percentage points in the carrying-bag and wearing-coat scenarios, respectively. Notably, the inference process did not require motion-capture data, validating that real 3D pose effectively enhanced gait recognition in complex scenarios. The results further showed that real 3D pose semantics enhanced multimodal feature representation and maintained stable recognition performance under view variations and appearance interference.
|
|
Received: 16 December 2025
Published: 28 July 2026
|
|
|
| Fund: 山西省中央引导地方科技发展资金项目(YDZJSX2022A016);山西省重点研发计划资助项目(2022ZDYF128);山西省科技战略项目(202404030401080). |
|
Corresponding Authors:
Hongxia DENG
E-mail: gw190428@163.com;denghongxia@tyut.edu.cn
|
基于姿态诱导的多模态步态识别
为了实现复杂环境下步态的多视角细粒度识别,提出基于姿态诱导的多模态步态识别方法,并构建包含真实三维姿态的多模态数据集(CASIA-V). 该数据集通过多设备同步采集93名受试者在正常、背包、穿大衣行走下的11视角数据,包含11253个RGB序列及对应动捕姿态与深度图. 在方法层面,提出时空对齐策略实现动捕与视频姿态匹配,设计可独立训练的三维姿态诱导模块(3D-SIM)提取真实姿态语义,并构建双分支模型SIGait融合姿态与轮廓特征. 实验表明,在CASIA-V数据集的跨视角评估协议下(排除相同视角样本对),SIGait在正常行走场景下的Rank-1准确率达97.7%;相较基线方法MSAFF,在背包与穿大衣场景下Rank-1准确率分别提升5.1和7.1个百分点,且推理过程无需动捕数据,验证了真实三维姿态对复杂场景步态识别的有效促进. 结果同时显示,真实三维姿态语义能增强多模态特征表达,使识别系统在视角变化与外观干扰条件下保持稳定识别性能.
关键词:
步态识别,
多模态数据集,
三维姿态,
时空对齐,
姿态诱导,
特征融合,
惯性动捕
|
|
| [1] |
MAHMOUD M, KASEM M S, KANG H S A comprehensive survey of masked faces: recognition, detection, and unmasking[J]. Applied Sciences, 2024, 14 (19): 8781
doi: 10.3390/app14198781
|
|
|
| [2] |
JIA Z, HUANG C, WANG Z, et al Finger recovery transformer: toward better incomplete fingerprint identification[J]. IEEE Transactions on Information Forensics and Security, 2024, 19: 8860- 8874
doi: 10.1109/TIFS.2024.3419690
|
|
|
| [3] |
KUEHLKAMP A, BOYD A, CZAJKA A, et al. Interpretable deep learning-based forensic iris segmentation and recognition [C]// Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision Workshops. Waikoloa: IEEE, 2022: 359–368.
|
|
|
| [4] |
赵晓东, 刘作军, 陈玲玲, 等 下肢假肢穿戴者跑动步态识别方法[J]. 浙江大学学报: 工学版, 2018, 52 (10): 1980- 1988 ZHAO Xiaodong, LIU Zuojun, CHEN Lingling, et al Approach of running gait recognition for lower limb amputees[J]. Journal of Zhejiang University: Engineering Science, 2018, 52 (10): 1980- 1988
doi: 10.3785/j.issn.1008-973X.2018.10.018
|
|
|
| [5] |
CHAO H, WANG K, HE Y, et al GaitSet: cross-view gait recognition through utilizing gait as a deep set[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022, 44 (7): 3467- 3478
doi: 10.1109/tpami.2021.3057879
|
|
|
| [6] |
FAN C, PENG Y, CAO C, et al. GaitPart: temporal part-based model for gait recognition [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle: IEEE, 2020: 14213–14221.
|
|
|
| [7] |
HUANG Z, XUE D, SHEN X, et al. 3D local convolutional neural networks for gait recognition [C]// Proceedings of the IEEE/CVF International Conference on Computer Vision. Montreal: IEEE, 2022: 14900–14909.
|
|
|
| [8] |
LIN B, ZHANG S, YU X. Gait recognition via effective global-local feature representation and local temporal aggregation [C]// Proceedings of the IEEE/CVF International Conference on Computer Vision. Montreal: IEEE, 2022: 14628–14636.
|
|
|
| [9] |
WU Z, HUANG Y, WANG L, et al A comprehensive study on cross-view gait based human identification with deep CNNs[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2017, 39 (2): 209- 226
doi: 10.1109/TPAMI.2016.2545669
|
|
|
| [10] |
HUANG X, ZHU D, WANG H, et al. Context-sensitive temporal feature learning for gait recognition [C]// Proceedings of the IEEE/CVF International Conference on Computer Vision. Montreal: IEEE, 2022: 12889–12898.
|
|
|
| [11] |
WANG M, GUO X, LIN B, et al. DyGait: exploiting dynamic representations for high-performance gait recognition [C]// Proceedings of the IEEE/CVF International Conference on Computer Vision. Paris: IEEE, 2024: 13378–13387.
|
|
|
| [12] |
LIAO R, YU S, AN W, et al A model-based gait recognition method with body pose and human prior knowledge[J]. Pattern Recognition, 2020, 98: 107069
doi: 10.1016/j.patcog.2019.107069
|
|
|
| [13] |
TEEPE T, KHAN A, GILG J, et al. Gaitgraph: graph convolutional network for skeleton-based gait recognition [C]// Proceedings of the IEEE International Conference on Image Processing. Anchorage: IEEE, 2021: 2314–2318.
|
|
|
| [14] |
FU Y, MENG S, HOU S, et al. GPGait: generalized pose-based gait recognition [C]// Proceedings of the IEEE/CVF International Conference on Computer Vision. Paris: IEEE, 2024: 19538–19547.
|
|
|
| [15] |
ZHANG C, CHEN X P, HAN G Q, et al Spatial transformer network on skeleton-based gait recognition[J]. Expert Systems, 2023, 40 (6): e13244
doi: 10.1111/exsy.13244
|
|
|
| [16] |
GUO H, JI Q. Physics-augmented autoencoder for 3D skeleton-based gait recognition [C]// Proceedings of the IEEE/CVF International Conference on Computer Vision. Paris: IEEE, 2024: 19570–19581.
|
|
|
| [17] |
吴越, 梁铮, 高巍, 等 基于SMPL模态分解与嵌入融合的多模态步态识别[J]. 浙江大学学报: 工学版, 2026, 60 (1): 52- 60 WU Yue, LIANG Zheng, GAO Wei, et al Multi-modal gait recognition based on SMPL model decomposition and embedding fusion[J]. Journal of Zhejiang University: Engineering Science, 2026, 60 (1): 52- 60
|
|
|
| [18] |
ZOU S, XIONG J, FAN C, et al. A multi-stage adaptive feature fusion neural network for multimodal gait recognition [C]// Proceedings of the IEEE International Joint Conference on Biometrics. Ljubljana: IEEE, 2024: 1–10.
|
|
|
| [19] |
LI G, GUO L, ZHANG R, et al TransGait: multimodal-based gait recognition with set transformer[J]. Applied Intelligence, 2023, 53 (2): 1535- 1547
doi: 10.1007/s10489-022-03543-y
|
|
|
| [20] |
ZHENG W, ZHU H, ZHENG Z, et al GaitSTR: gait recognition with sequential two-stream refinement[J]. IEEE Transactions on Biometrics, Behavior, and Identity Science, 2024, 6 (4): 528- 538
doi: 10.1109/TBIOM.2024.3390626
|
|
|
| [21] |
FAN C, MA J, JIN D, et al. SkeletonGait: gait recognition using skeleton maps [C]// Proceedings of the AAAI Conference on Artificial Intelligence. Vancouver: AAAI Press, 2024: 1662–1669.
|
|
|
| [22] |
YU S, TAN D, TAN T. A framework for evaluating the effect of view angle, clothing and carrying condition on gait recognition [C]// Proceedings of the 18th International Conference on Pattern Recognition. Hong Kong: IEEE, 2006: 441–444.
|
|
|
| [23] |
TAKEMURA N, MAKIHARA Y, MURAMATSU D, et al Multi-view large population gait dataset and its performance evaluation for cross-view gait recognition[J]. IPSJ Transactions on Computer Vision and Applications, 2018, 10 (1): 4
doi: 10.1186/s41074-018-0039-6
|
|
|
| [24] |
ZHENG J, LIU X, LIU W, et al. Gait recognition in the wild with dense 3D representations and a benchmark [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. New Orleans: IEEE, 2022: 20196–20205.
|
|
|
| [25] |
ZHU Z, GUO X, YANG T, et al. Gait recognition in the wild: a benchmark [C]// Proceedings of the IEEE/CVF International Conference on Computer Vision. Montreal: IEEE, 2021: 14769–14779.
|
|
|
| [26] |
KIPF T N, WELLING M. Semi-supervised classification with graph convolutional networks [EB/OL]. [2025–11–23]. https://arxiv.org/abs/1609.02907.
|
|
|
| [27] |
LI J, ZHANG Y, SHAN H, et al. Gaitcotr: improved spatial-temporal representation for gait recognition with a hybrid convolution-transformer framework [C]// 2023 IEEE International Conference on Acoustics, Speech and Signal Processing. Rhodes Island: IEEE, 2023: 1–5.
|
|
|
| [28] |
SUN K, XIAO B, LIU D, et al. Deep high-resolution representation learning for human pose estimation [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Long Beach: IEEE, 2020: 5686–5696.
|
|
|
| [29] |
PAVLLO D, FEICHTENHOFER C, GRANGIER D, et al. 3D human pose estimation in video with temporal convolutions and semi-supervised training [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Long Beach: IEEE, 2020: 7745–7754.
|
|
|
| [30] |
LIN S, RYABTSEV A, SENGUPTA S, et al. Real-time high-resolution background matting [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Nashville: IEEE, 2021: 8758–8767.
|
|
|
| [31] |
TAKEMURA N, MAKIHARA Y, MURAMATSU D, et al Multi-view large population gait dataset and its performance evaluation for cross-view gait recognition[J]. IPSJ Transactions on Computer Vision and Applications, 2018, 10 (1): 4
doi: 10.1186/s41074-018-0039-6
|
|
|
|
Viewed |
|
|
|
Full text
|
|
|
|
|
Abstract
|
|
|
|
|
Cited |
|
|
|
|
| |
Shared |
|
|
|
|
| |
Discussed |
|
|
|
|