Please wait a minute...
浙江大学学报(工学版)  2026, Vol. 60 Issue (8): 1670-1677    DOI: 10.3785/j.issn.1008-973X.2026.08.006
计算机技术     
全局学习扩展的可见光-红外行人重识别
郭子强1(),肖璇1,陶浩然1,王少荣1,2,*()
1. 北京林业大学 信息学院,北京 100083
2. 国家林业草原林业智能信息处理工程技术研究中心,北京 100083
Global learning-expanded visible-infrared person re-identification
Ziqiang GUO1(),Xuan XIAO1,Haoran TAO1,Shaorong WANG1,2,*()
1. School of Information Science and Technology, Beijing Forest University, Beijing 100083, China
2. Engineering Research Center for Forestry-oriented Intelligent Information Processing of National Forestry and Grassland Administration, Beijing 100083, China
 全文: PDF(2304 KB)   HTML
摘要:

在可见光-红外行人重识别任务中,由于训练样本有限,且可见光与红外图像之间存在较大差异,如何挖掘多样化的跨模态信息并实现特征对齐是该任务的关键挑战. 为此,提出全局学习扩展的嵌入增强网络,有效融合视觉变换器与卷积神经网络的优势,协同捕获图像的局部细节与全局结构信息. 通过生成多样化的嵌入特征,引入特征图结构对齐机制,进一步增强模型对全局结构的建模能力,减小可见光与红外图像之间的模态差异,从而学习到更加丰富且判别性强的特征表示. 实验结果表明,相较于基线模型DEEN,所提方法在SYSU-MM01数据集的全搜索模式和LLCM数据集的可见光检索红外模式下的Rank-1准确率分别提高了4.2和9.1个百分点. 该方法通过聚合全局空间信息,生成了更具判别性的特征嵌入,显著增强了网络性能.

关键词: 可见光-红外行人重识别跨模态检索全局特征特征对齐嵌入增强    
Abstract:

In visible-infrared person re-identification, the limited number of training samples and the significant discrepancy between visible and infrared images pose key challenges in mining diverse cross-modal information and achieving feature alignment. An embedding enhancement network based on global learning expansion was proposed to address the challenges, which effectively integrated the strengths of both vision Transformer and CNN to jointly capture local details and global structural information from images. By generating diversified embedding features and introducing the feature map structural alignment mechanism, the model’s ability to model global structures was further enhanced and the modality gap between visible and infrared images was reduced, thereby enabling more comprehensive and discriminative feature representations to be learned. Experimental results demonstrated that, compared to the baseline model DEEN, the proposed method achieved improvements of 4.2 and 9.1 percentage points in Rank-1 accuracy under the full-search mode of the SYSU-MM01 dataset and the visible-light-to-infrared retrieval mode of the LLCM dataset, respectively. By aggregating global spatial information, the proposed method generated more discriminative feature embeddings, significantly boosting the network performance.

Key words: visible-infrared person re-identification    cross-modal retrieval    global feature    feature alignment    embedding enhancement
收稿日期: 2025-07-15 出版日期: 2026-07-16
CLC:  TP 391.41  
通讯作者: 王少荣     E-mail: ZQGuo@bjfu.edu.cn;shaorongwang@hotmail.com
作者简介: 郭子强(1998—),男,硕士生,从事计算机视觉研究. orcid.org/0009-0005-8324-1559. E-mail:ZQGuo@bjfu.edu.cn
服务  
把本文推荐给朋友
加入引用管理器
E-mail Alert
作者相关文章  
郭子强
肖璇
陶浩然
王少荣

引用本文:

郭子强,肖璇,陶浩然,王少荣. 全局学习扩展的可见光-红外行人重识别[J]. 浙江大学学报(工学版), 2026, 60(8): 1670-1677.

Ziqiang GUO,Xuan XIAO,Haoran TAO,Shaorong WANG. Global learning-expanded visible-infrared person re-identification. Journal of ZheJiang University (Engineering Science), 2026, 60(8): 1670-1677.

链接本文:

https://www.zjujournals.com/eng/CN/10.3785/j.issn.1008-973X.2026.08.006        https://www.zjujournals.com/eng/CN/Y2026/V60/I8/1670

图 1  全局学习扩展网络的整体架构
图 2  水平方向上的循环卷积
图 3  全局特征学习损失示意图
方法全搜索室内搜索
R-1/%R-10/%R-20/%mAP/%R-1/%R-10/%R-20/%mAP/%
DART[30]68.796.499.066.372.597.899.578.2
CAJ[31]69.995.798.566.976.397.999.580.4
MPANet[32]70.696.298.868.276.798.299.681.0
MMN[33]70.696.299.066.976.297.299.379.6
DCLNet[34]70.865.373.576.8
MAUM[4]71.768.877.081.9
DEEN[16]74.797.699.271.880.399.099.883.3
HOS-Net[35]75.674.284.286.7
SAAI[36]75.977.083.288.0
MUN[37]76.297.873.879.498.182.1
MSCLNet[38]77.097.699.271.678.599.399.981.2
PartMix[17]77.874.681.584.4
IDKL[18]81.497.498.979.987.198.399.389.4
GLE78.998.499.676.186.099.399.788.1
表 1  SYSU-MMO1数据集上GLE与先进方法的性能比较
方法可见光检索红外红外检索可见光
R-1/%R-10/%R-20/%mAP/%R-1/%R-10/%R-20/%mAP/%
DART[30]83.675.782.073.8
CAJ[31]85.095.597.579.184.895.397.577.8
MPANet[32]82.880.783.780.9
MMN[33]91.697.798.984.187.596.098.180.5
DCLNet[34]81.274.378.070.6
MAUM[4]87.985.187.084.3
DEEN[16]91.197.898.985.189.596.898.483.4
HOS-Net[35]94.790.493.389.2
SAAI[36]91.191.592.192.0
CMT[21]95.298.887.392.097.999.184.5
MUN[37]95.298.987.291.998.085.0
IDKL[18]94.790.294.290.4
GLE94.998.899.690.590.298.099.189.2
表 2  RegDB数据集上GLE与先进方法的性能比较
方法可见光检索红外红外检索可见光
R-1/%R-10/%R-20/%mAP/%R-1/%R-10/%R-20/%mAP/%
DDAG[15]48.079.286.152.340.371.479.648.4
AGW[7]51.581.587.955.343.674.682.451.8
LbA[39]50.884.391.155.643.878.286.653.1
CAJ[31]56.585.390.959.848.879.585.356.6
DART[30]60.487.191.963.252.280.787.059.8
MMN[33]59.988.593.662.752.581.688.458.9
DEEN[16]62.590.394.765.854.984.990.962.9
HOS-Net[35]64.967.956.463.2
IDKL[18]72.266.470.765.2
GLE71.692.596.256.957.786.392.164.4
表 3  LLCM数据集上GLE与先进方法的性能比较
FDEGFLGFAR-1/%mAP/%
×××74.771.8
××75.672.6
×76.973.5
××76.272.5
×77.774.7
78.976.1
表 4  GLE中各模块的消融实验结果
图 4  损失函数中各个超参数对模型性能的影响
图 5  类内、类间距离和特征分布可视化
图 6  LLCM数据集上GLE与基线模型DEEN的检索结果对比
1 ZHENG Y, TANG S, TENG G, et al. Online pseudo label generation by hierarchical cluster dynamics for adaptive person re-identification [C]// Proceedings of the IEEE/CVF International Conference on Computer Vision. Montreal: IEEE, 2021: 8351–8361.
2 PU N, CHEN W, LIU Y, et al. Lifelong person re-identification via adaptive knowledge accumulation [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Nashville: IEEE, 2021: 7897–7906.
3 TAN L, DAI P, JI R, et al. Dynamic prototype mask for occluded person re-identification [C]// Proceedings of the 30th ACM International Conference on Multimedia. Lisboa: ACM, 2022: 531–540.
4 LIU J, SUN Y, ZHU F, et al. Learning memory-augmented unidirectional metrics for cross-modality person re-identification [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. New Orleans: IEEE, 2022: 19344–19353.
5 YANG B, YE M, CHEN J, et al. Augmented dual-contrastive aggregation learning for unsupervised visible-infrared person re-identification [C]// Proceedings of the 30th ACM International Conference on Multimedia. Lisboa: ACM, 2022: 2843–2851.
6 WEI X, LI D, HONG X, et al. Co-attentive lifting for infrared-visible person re-identification [C]// Proceedings of the 28th ACM International Conference on Multimedia. Seattle: ACM, 2020: 1028–1037.
7 YE M, SHEN J, LIN G, et al Deep learning for person re-identification: a survey and outlook[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022, 44 (6): 2872- 2893
doi: 10.1109/TPAMI.2021.3054775
8 REN M, HE L, LIAO X, et al. Learning instance-level spatial-temporal patterns for person re-identification [C]// Proceedings of the IEEE/CVF International Conference on Computer Vision. Montreal: IEEE, 2021: 14910–14919.
9 SUN Y, ZHENG L, YANG Y, et al. Beyond part models: person retrieval with refined part pooling (and a strong convolutional baseline) [C]// Proceedings of the European Conference on Computer Vision. Munich: Springer, 2018: 501–518.
10 HE S, LUO H, WANG P, et al. TransReID: Transformer-based object re-identification [C]// Proceedings of the IEEE/CVF International Conference on Computer Vision. Montreal: IEEE, 2021: 14993–15002.
11 ZHANG L, LIU Z, ZHANG W, et al Style uncertainty based self-paced meta learning for generalizable person re-identification[J]. IEEE Transactions on Image Processing, 2023, 32: 2107- 2119
doi: 10.1109/TIP.2023.3263112
12 FU X, HUANG F, ZHOU Y, et al Cross-modal cross-domain dual alignment network for RGB-infrared person re-identification[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2022, 32 (10): 6874- 6887
doi: 10.1109/TCSVT.2022.3173263
13 LIANG W, WANG G, LAI J, et al Homogeneous-to-heterogeneous: unsupervised learning for RGB-infrared person re-identification[J]. IEEE Transactions on Image Processing, 2021, 30: 6392- 6407
doi: 10.1109/TIP.2021.3092578
14 WU A, ZHENG W S, YU H X, et al. RGB-infrared cross-modality person re-identification [C]// Proceedings of the IEEE International Conference on Computer Vision. Venice: IEEE, 2017: 5390–5399.
15 YE M, SHEN J, CRANDALL D J, et al. Dynamic dual-attentive aggregation learning for visible-infrared person re-identification [C]// Proceedings of the European Conference on Computer Vision. Glasgow: Springer, 2020: 229–247.
16 ZHANG Y, WANG H. Diverse embedding expansion network and low-light cross-modality benchmark for visible-infrared person re-identification [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Vancouver: IEEE, 2023: 2153–2162.
17 KIM M, KIM S, PARK J, et al. PartMix: regularization strategy to learn part discovery for visible-infrared person re-identification [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Vancouver: IEEE, 2023: 18621–18632.
18 REN K, ZHANG L. Implicit discriminative knowledge learning for visible-infrared person re-identification [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle: IEEE, 2024: 393–402.
19 YANG B, CHEN J, YE M. Top-K visual tokens Transformer: selecting tokens for visible-infrared person re-identification [C]// Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing. Rhodes Island: IEEE, 2023: 1–5.
20 CHEN C, YE M, QI M, et al Structure-aware positional Transformer for visible-infrared person re-identification[J]. IEEE Transactions on Image Processing, 2022, 31: 2352- 2364
doi: 10.1109/TIP.2022.3141868
21 JIANG K, ZHANG T, LIU X, et al. Cross-modality Transformer for visible-infrared person re-identification [C]// Proceedings of the European Conference on Computer Vision. Tel Aviv: Springer, 2022: 480–496.
22 ZHAO J, WANG H, ZHOU Y, et al Spatial-channel enhanced Transformer for visible-infrared person re-identification[J]. IEEE Transactions on Multimedia, 2023, 25: 3668- 3680
doi: 10.1109/TMM.2022.3163847
23 ZHANG H, HU W, WANG X. ParC-Net: position aware circular convolution with merits from ConvNets and Transformer [C]// Proceedings of the European Conference on Computer Vision. Tel Aviv: Springer, 2022: 613–630.
24 CHEN Y, DAI X, CHEN D, et al. Mobile-former: bridging MobileNet and Transformer [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. New Orleans: IEEE, 2022: 5260–5269.
25 HE K, ZHANG X, REN S, et al. Deep residual learning for image recognition [C]// Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Las Vegas: IEEE, 2016: 770–778.
26 LI X, LU Y, LIU B, et al. Counterfactual intervention feature transfer for visible-infrared person re-identification [C]// Proceedings of the European Conference on Computer Vision. Tel Aviv: Springer, 2022: 381–398.
27 LUO H, GU Y, LIAO X, et al. Bag of tricks and a strong baseline for deep person re-identification [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops. Long Beach: IEEE, 2019: 1487–1495.
28 HERMANS A, BEYER L, LEIBE B. In defense of the triplet loss for person re-identification [EB/OL]. (2017–11–21) [2024–10–25]. https://arxiv.org/abs/1703.07737.
29 ZHONG Z, ZHENG L, KANG G, et al Random erasing data augmentation[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2020, 34 (7): 13001- 13008
doi: 10.1609/aaai.v34i07.7000
30 YANG M, HUANG Z, HU P, et al. Learning with twin noisy labels for visible-infrared person re-identification [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. New Orleans: IEEE, 2022: 14288–14297.
31 YE M, RUAN W, DU B, et al. Channel augmented joint learning for visible-infrared recognition [C]// Proceedings of the IEEE/CVF International Conference on Computer Vision. Montreal: IEEE, 2021: 13547–13556.
32 WU Q, DAI P, CHEN J, et al. Discover cross-modality nuances for visible-infrared person re-identification [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Nashville: IEEE, 2021: 4328–4337.
33 ZHANG Y, YAN Y, LU Y, et al. Towards a unified middle modality learning for visible-infrared person re-identification [C]// Proceedings of the 29th ACM International Conference on Multimedia. [S.l.]: ACM, 2021: 788–796.
34 SUN H, LIU J, ZHANG Z, et al. Not all pixels are matched: dense contrastive learning for cross-modality person re-identification [C]// Proceedings of the 30th ACM International Conference on Multimedia. Lisboa: ACM, 2022: 5333–5341.
35 QIU L, CHEN S, YAN Y, et al High-order structure based middle-feature learning for visible-infrared person re-identification[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2024, 38 (5): 4596- 4604
doi: 10.1609/aaai.v38i5.28259
36 FANG X, YANG Y, FU Y. Visible-infrared person re-identification via semantic alignment and affinity inference [C]// Proceedings of the IEEE/CVF International Conference on Computer Vision. Paris: IEEE, 2023: 11236–11245.
37 YU H, CHENG X, PENG W, et al. Modality unifying network for visible-infrared person re-identification [C]// Proceedings of the IEEE/CVF International Conference on Computer Vision. Paris: IEEE, 2023: 11151–11161.
38 ZHANG Y, ZHAO S, KANG Y, et al. Modality synergy complement learning with cascaded aggregation for visible-infrared person re-identification [C]// Proceedings of the European Conference on Computer Vision. Tel Aviv: Springer, 2022: 462–479.
39 PARK H, LEE S, LEE J, et al. Learning by aligning: visible-infrared person re-identification using cross-modal correspondences [C]// Proceedings of the IEEE/CVF International Conference on Computer Vision. Montreal: IEEE, 2021: 12026–12035.
[1] 梁礼明,王成斌,钟奕,陈林俊,吴健. 基于轻量高频Transformer与特征互补融合的视网膜血管分割[J]. 浙江大学学报(工学版), 2026, 60(7): 1392-1403.
[2] 吴越,梁铮,高巍,杨茂达,赵培森,邓红霞,常媛媛. 基于SMPL模态分解与嵌入融合的多模态步态识别[J]. 浙江大学学报(工学版), 2026, 60(1): 52-60.
[3] 陈捷丰,姚金良. 基于自相似嵌入和全局特征重排序的图像检索方法[J]. 浙江大学学报(工学版), 2025, 59(6): 1130-1139.
[4] 艾青林,杨佳豪,崔景瑞. 基于自适应增殖数据增强与全局特征融合的小目标行人检测[J]. 浙江大学学报(工学版), 2023, 57(10): 1933-1944.
[5] 徐泽鑫,段立娟,王文健,恩擎. 基于上下文特征融合的代码漏洞检测方法[J]. 浙江大学学报(工学版), 2022, 56(11): 2260-2270.