Please wait a minute...
浙江大学学报(工学版)  2026, Vol. 60 Issue (9): 1953-1961    DOI: 10.3785/j.issn.1008-973X.2026.09.012
计算机技术、自动控制技术     
基于空间自适应Transformer的人脸素描图像生成
万伟国1(),钟凯晨1,张迎梅1,姚丽2,包书鸿1,龚大伟3
1. 江西财经大学 软件与物联网工程学院,江西 南昌 330032
2. 江西财经大学 教务处,江西 南昌 330013
3. 江西科技师范大学 智能制造学院,江西 南昌 330038
Face photo-to-sketch synthesis based on spatially adaptive Transformer
Weiguo WAN1(),Kaichen ZHONG1,Yingmei ZHANG1,Li YAO2,Shuhong BAO1,Dawei GONG3
1. School of Software and Internet of Things Engineering, Jiangxi University of Finance and Economics, Nanchang 330032, China
2. Office of Educational Administration, Jiangxi University of Finance and Economics, Nanchang 330013, China
3. School of Intelligent Manufacturing, Jiangxi Science and Technology Normal University, Nanchang 330038, China
 全文: PDF(2412 KB)   HTML
摘要:

针对现有人脸素描图像生成方法存在的生成图像结构失真、纹理模糊及伪影残留等问题,提出基于空间自适应Transformer的生成网络(SAT-StyTR). 在编码器中构建双模态特征解耦模块(DFDM),通过并行的人脸内容与素描风格特征提取分支,实现跨模态特征的解耦表征. 设计行-列向量动态融合模块(VDFM),引入双向门控机制,以实现行、列向量的自适应特征融合,显著增强局部细节保留能力. 引入多尺度感知判别器(MSPD),利用层级式梯度监督机制在多语义粒度上优化素描图像纹理的真实感. 在CUHK、AR、XM2VTS和CUFSF数据集上的实验结果表明,所提方法在SSIM、PSNR、MSE、LPIPS和FID等指标上均优于Mamba-ST等现有方法,其FID值在4个数据集上比次优方法分别降低了19.3%、1.0%、15.5%和10.0%. 视觉对比分析进一步验证了所提方法的生成结果在五官轮廓等关键区域上具有更优的结构一致性与细节还原度.

关键词: 人脸素描图像生成空间自适应Transformer双模态特征解耦行-列向量动态融合多尺度感知判别器(MSPD)    
Abstract:

A spatially adaptive Transformer-based generation network (SAT-StyTR) was proposed to address the issues of existing face photo-to-sketch synthesis methods, such as structural distortion, blurred textures, and artifact residue. A dual-modal feature disentanglement module (DFDM) was constructed in the encoder. The parallel branches of face content and sketch style feature extraction were employed in DFDM to achieve disentangled representation of cross-modal features. A row-column vector dynamic fusion module (VDFM) was designed, and a bidirectional gating mechanism was introduced in the VDFM to enable adaptive feature fusion of row-column vectors. Then the capability of local detail retention was significantly improved. A multi-scale perceptive discriminator (MSPD) was incorporated to optimize the realism of generated sketch textures at multiple semantic levels by leveraging the hierarchical gradient supervision mechanism. The experimental results on the CUHK, AR, XM2VTS and CUFSF datasets demonstrated that the proposed method outperformed existing methods such as Mamba-ST on SSIM, PSNR, MSE, LPIPS, and FID metrics, with its FID values improved by 19.3%, 1.0%, 15.5%, and 10.0% over those of the second-best method on the four datasets, respectively. Visual comparative analysis further validated that superior structural consistency and detail restoration were achieved in critical regions, such as facial contours.

Key words: face photo-to-sketch synthesis    spatially adaptive Transformer    dual-modal feature disentanglement    row-column vector dynamic fusion    multi-scale perceptive discriminator (MSPD)
收稿日期: 2025-07-16 出版日期: 2026-07-20
CLC:  TP 393  
基金资助: 国家自然科学基金资助项目(62261025,62362032);江西省自然科学基金资助项目(20242BAB25076,20232BAB212015).
作者简介: 万伟国(1991—),男,副教授,博士,从事计算机视觉、图像处理研究. orcid.org/0000-0002-3537-979X. E-mail:wanweiguo@jxufe.edu.cn
服务  
把本文推荐给朋友
加入引用管理器
E-mail Alert
作者相关文章  
万伟国
钟凯晨
张迎梅
姚丽
包书鸿
龚大伟

引用本文:

万伟国,钟凯晨,张迎梅,姚丽,包书鸿,龚大伟. 基于空间自适应Transformer的人脸素描图像生成[J]. 浙江大学学报(工学版), 2026, 60(9): 1953-1961.

Weiguo WAN,Kaichen ZHONG,Yingmei ZHANG,Li YAO,Shuhong BAO,Dawei GONG. Face photo-to-sketch synthesis based on spatially adaptive Transformer. Journal of ZheJiang University (Engineering Science), 2026, 60(9): 1953-1961.

链接本文:

https://www.zjujournals.com/eng/CN/10.3785/j.issn.1008-973X.2026.09.012        https://www.zjujournals.com/eng/CN/Y2026/V60/I9/1953

图 1  基于空间自适应Transformer的人脸素描图像生成网络(SAT-StyTR)示意图
图 2  双模态特征解耦模块结构
图 3  多尺度感知判别器结构
图 4  不同方法在CUHK、AR、XM2VTS和CUFSF数据集上的定性实验结果对比
模型CUHK数据集AR数据集
SSIMPSNR/dBMSELPIPSFIDSSIMPSNR/dBMSELPIPSFID
pix2pix0.658730.7056.300.173189.080.594629.6871.830.215099.92
U-GAT-IT0.555129.7569.280.3206131.850.473429.2877.790.4122170.04
DCLGAN0.647930.4858.560.214972.460.612829.7870.460.189397.96
文献[15]模型0.650630.1264.260.189371.270.604729.8069.310.218082.45
CAP-VSTNet0.638930.4359.340.222882.410.577829.6571.060.2710117.35
Mamba-ST0.536529.6972.030.346075.400.604129.0681.170.237984.68
SAT-StyTR0.667630.8953.930.174557.510.621329.8967.880.204381.66
模型XM2VTS数据集CUFSF数据集
SSIMPSNR/dBMSELPIPSFIDSSIMPSNR/dBMSELPIPSFID
pix2pix0.452729.2278.110.251462.130.361528.3694.940.277054.89
U-GAT-IT0.435929.3076.700.4471137.140.398128.2796.950.273142.63
DCLGAN0.498329.2278.280.284791.560.375028.3595.120.258444.22
文献[15]模型0.497829.4773.950.273263.650.382328.2597.260.272532.56
CAP-VSTNet0.496829.2378.150.271079.810.365528.10100.650.290961.21
Mamba-ST0.453628.3794.810.344974.160.401528.2298.010.281849.03
SAT-StyTR0.511629.5273.190.243652.480.423628.5191.740.235629.30
表 1  不同方法在CUHK、AR、XM2VTS和CUFSF数据集上的定量指标对比
方法SSIMPSNR/dBMSELPIPS
DFDM0.657030.4459.630.1830
DFDM+MSPD0.667730.7256.730.1788
SAT-StyTR0.669230.8953.930.1745
表 2  所提模块的消融实验结果
方法SSIMPSNR/dBMSELPIPS
w/o branch0.461427.96104.160.4442
w/ branch0.511629.5273.190.2436
表 3  DFDM中双模态分支的有效性
图 5  双模态分支结构的消融实验可视化结果
卷积核SSIMPSNR/dBMSELPIPS
5×50.508529.4873.660.2517
7×70.510729.2178.430.2485
13×13 & 3×30.511629.5273.190.2436
表 4  大、小核卷积交替结构的有效性
图 6  所提模型在复杂场景下生成的素描图像
模型Np/106tinf/s
pix2pix57.180.039
U-GAT-IT664.830.088
DCLGAN29.270.159
文献[15]模型58.900.035
CAP-VSTNet4.090.210
Mamba-ST40.020.076
SAT-StyTR83.740.041
表 5  不同方法的参数量和平均推理时间对比
1 GAO L, LIU F L, CHEN S Y, et al SketchFaceNeRF: sketch-based facial generation and editing in neural radiance fields[J]. ACM Transactions on Graphics, 2023, 42 (4): 159
2 翁丽芬, 李晨阳, 许华荣 基于GAN的分步合成人脸素描生成算法[J]. 计算机辅助设计与图形学学报, 2023, 35 (9): 1363- 1373
WENG Lifen, LI Chenyang, XU Huarong Stepwise synthetic face sketch generation algorithm based on GAN[J]. Journal of Computer-Aided Design & Computer Graphics, 2023, 35 (9): 1363- 1373
doi: 10.3724/SP.J.1089.2023.19582
3 ABDEL-AZIZ H G M, EBEID H M, MOSTAFA M G M. An unsupervised method for face photo-sketch synthesis and recognition [C]// Proceedings of the 7th International Conference on Information and Communication Systems. Irbid: IEEE, 2016: 221–226.
4 GALEA C, FARRUGIA R A. Face photo-sketch recognition using local and global texture descriptors [C]// Proceedings of the 24th European Signal Processing Conference. Budapest: IEEE, 2016: 2240–2244.
5 LI J, YU X, PENG C, et al Adaptive representation-based face sketch-photo synthesis[J]. Neurocomputing, 2017, 269: 152- 159
doi: 10.1016/j.neucom.2016.10.095
6 LIU Z S, SIU W C, CHAN H A. Learn to sketch: a fast approach for universal photo sketch [C]// Proceedings of the Asia-Pacific Signal and Information Processing Association Annual Summit and Conference. Tokyo: IEEE, 2021: 1450–1457.
7 ZHANG L, LIN L, WU X, et al. End-to-end photo-sketch generation via fully convolutional representation learning [C]// Proceedings of the 5th ACM on International Conference on Multimedia Retrieval. Shanghai: ACM, 2015: 627–634.
8 GALEA C, FARRUGIA R A Forensic face photo-sketch recognition using a deep learning-based architecture[J]. IEEE Signal Processing Letters, 2017, 24 (11): 1586- 1590
doi: 10.1109/LSP.2017.2749266
9 ZHU M, LI J, WANG N, et al A deep collaborative framework for face photo-sketch synthesis[J]. IEEE Transactions on Neural Networks and Learning Systems, 2019, 30 (10): 3096- 3108
doi: 10.1109/TNNLS.2018.2890018
10 曹林, 王震, 杜康宁, 等 基于层次对比生成对抗网络的非配对素描人脸合成[J]. 中国科技论文, 2024, 19 (6): 715- 723
CAO Lin, WANG Zhen, DU Kangning, et al Unpaired sketch face synthesis based on hierarchical contrast generative adversarial network[J]. China Sciencepaper, 2024, 19 (6): 715- 723
doi: 10.3969/j.issn.2095-2783.2024.06.011
11 ZHU M, LIANG C, WANG N, et al. A sketch-Transformer network for face photo-sketch synthesis [C]// Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence. Montreal: IJCAI, 2021: 1352–1358.
12 CHEN C, YE M, QI M, et al. Sketch Transformer: asymmetrical disentanglement learning from dynamic synthesis [C]// Proceedings of the 30th ACM International Conference on Multimedia. Lisboa: ACM, 2022: 4012–4020.
13 YU W, ZHU M, WANG N, et al An efficient Transformer based on global and local self-attention for face photo-sketch synthesis[J]. IEEE Transactions on Image Processing, 2023, 32: 483- 495
doi: 10.1109/TIP.2022.3229614
14 陈长武, 曹林, 郭亚男, 等 基于域自适应均值网络的素描人脸识别方法[J]. 计算机应用与软件, 2023, 40 (4): 107- 115
CHEN Changwu, CAO Lin, GUO Yanan, et al Sketch face recognition based on domain adaptation mean network[J]. Computer Applications and Software, 2023, 40 (4): 107- 115
15 SHI Z, WAN W Transformer-based adversarial network for semi-supervised face sketch synthesis[J]. Journal of Visual Communication and Image Representation, 2024, 102: 104204
doi: 10.1016/j.jvcir.2024.104204
16 DENG Y, TANG F, DONG W, et al. StyTr2: image style transfer with Transformers [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. New Orleans: IEEE, 2022: 11316–11326.
17 SUN J, LIU X, BÄCK T, et al Learning adaptive differential evolution algorithm from optimization experiences by policy gradient[J]. IEEE Transactions on Evolutionary Computation, 2021, 25 (4): 666- 680
doi: 10.1109/TEVC.2021.3060811
18 LI Y, CHEN X, YANG B, et al. DeepFacePencil: creating face images from freehand sketches [C]// Proceedings of the 28th ACM International Conference on Multimedia. Seattle: ACM, 2020: 991–999.
19 KARNEWAR A, WANG O. MSG-GAN: multi-scale gradients for generative adversarial networks [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle: IEEE, 2020: 7796–7805.
20 WANG X, TANG X Face photo-sketch synthesis and recognition[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2009, 31 (11): 1955- 1967
doi: 10.1109/TPAMI.2008.222
21 ZHANG W, WANG X, TANG X. Coupled information-theoretic encoding for face photo-sketch recognition [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Colorado Springs: IEEE, 2011: 513–520.
22 KINGMA D P, BA J. Adam: a method for stochastic optimization [EB/OL]. (2017-01-30) [2025-07-15]. https://arxiv.org/abs/1412.6980.
23 XIONG R, YANG Y, HE D, et al. On layer normalization in the Transformer architecture [C]// Proceedings of the 37th International Conference on Machine Learning. [S.l.]: PMLR, 2020: 10524–10533.
24 GOODFELLOW I, BENGIO Y, COURVILLE A. Optimization for training deep models [M]// Deep learning. Cambridge: MIT Press, 2016: 291–293.
25 ISOLA P, ZHU J Y, ZHOU T, et al. Image-to-image translation with conditional adversarial networks [C]// Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Honolulu: IEEE, 2017: 5967–5976.
26 KIM J, KIM M, KANG H, et al. U-GAT-IT: unsupervised generative attentional networks with adaptive layer-instance normalization for image-to-image translation [EB/OL]. (2020-04-08) [2025-07-15]. https://arxiv.org/abs/1907.10830.
27 HAN J, SHOEIBY M, PETERSSON L, et al. Dual contrastive learning for unsupervised image-to-image translation [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops. Nashville: IEEE, 2021: 746–755.
28 WEN L, GAO C, ZOU C. CAP-VSTNet: content affinity preserved versatile style transfer [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Vancouver: IEEE, 2023: 18300–18309.
29 BOTTI F, ERGASTI A, ROSSI L, et al. Mamba-ST: state space model for efficient style transfer [EB/OL]. (2024-09-16) [2025-07-15]. https://arxiv.org/abs/2409.10385.
30 WANG Z, BOVIK A C, SHEIKH H R, et al Image quality assessment: from error visibility to structural similarity[J]. IEEE Transactions on Image Processing, 2004, 13 (4): 600- 612
doi: 10.1109/TIP.2003.819861
31 KORHONEN J, YOU J. Peak signal-to-noise ratio revisited: is simple beautiful? [C]// Proceedings of the Fourth International Workshop on Quality of Multimedia Experience. Melbourne: IEEE, 2012: 37–38.
32 MARMOLIN H Subjective MSE measures[J]. IEEE Transactions on Systems, Man, and Cybernetics, 1986, 16 (3): 486- 489
doi: 10.1109/TSMC.1986.4308985
33 KETTUNEN M, HÄRKÖNEN E, LEHTINEN J. E-LPIPS: robust perceptual image similarity via random transformation ensembles [EB/OL]. (2019-06-11) [2025-07-15]. https://arxiv.org/abs/1906.03973.
[1] 连远锋,范树玉,张本哲,王森. 融合几何约束特征提取与多奖励协同优化的仿生步态控制[J]. 浙江大学学报(工学版), 2026, 60(9): 1862-1871.
[2] 韩翔宇,韩伟,王英丞,张可臻,王峰年,姚明宇. 基于氢氧化物的热化学储能体系研究进展[J]. 浙江大学学报(工学版), 2026, 60(8): 1611-1626.
[3] 段孟滨,白国星,孟宇,顾青,汪振,伊力夏提·伊力哈木江,刘绍冲. 基于A*与多参考点MPC的差动机器人路径规划与跟踪控制[J]. 浙江大学学报(工学版), 2026, 60(8): 1627-1637.
[4] 常天根,田国富,唐媛媛,曹明学. 考虑主客观因素的自动驾驶模型预测控制参数优化[J]. 浙江大学学报(工学版), 2026, 60(8): 1638-1649.
[5] 王美佳,张帆,王桦,张明丽,张彩明. 面向长期时间序列预测的多尺度双流架构[J]. 浙江大学学报(工学版), 2026, 60(8): 1770-1781.
[6] 彭商濂,娄颖,冯丽. 基于核心实体与局部子图的知识图谱增量更新方法[J]. 浙江大学学报(工学版), 2026, 60(8): 1809-1818.
[7] 李佳磊,郝伟,王士发,梅童,马文来. 基于自适应增益的三旋翼无人机超螺旋滑模抗扰容错控制[J]. 浙江大学学报(工学版), 2026, 60(7): 1475-1481.
[8] 刘柯良,陈坚,邱智宣,张雨,康波,安萌,申秀敏. 建成环境对违法停车需求的非线性影响模型[J]. 浙江大学学报(工学版), 2026, 60(6): 1196-1204.
[9] 娄世猛,邵玉斌,杜庆治,唐菁敏,张赜涛. 基于KAN与CKAN优化的医学图像分割模型[J]. 浙江大学学报(工学版), 2026, 60(6): 1277-1288.
[10] 廖文碧,郑梦莲,俞自涛. 考虑客流时空分布的航站楼空调系统热湿-新风双层运行优化[J]. 浙江大学学报(工学版), 2026, 60(5): 1006-1015.
[11] 王骏骋,章世伟. 磁流变减振器力学模型的模糊综合评价方法[J]. 浙江大学学报(工学版), 2026, 60(5): 1047-1058.
[12] 罗杰,杨鉴. 基于特征映射模型的情感语音合成方法[J]. 浙江大学学报(工学版), 2026, 60(5): 1092-1099.
[13] 林洪彬,吕思进,王晨阳,蔡天放,骆鹏伟. 基于残差/梯度高斯自适应采样的径向基网络[J]. 浙江大学学报(工学版), 2026, 60(5): 1119-1127.
[14] 陈思如,舒元超. 多模态大模型边缘部署与推理加速技术综述[J]. 浙江大学学报(工学版), 2026, 60(4): 723-737.
[15] 蔡智,周正东,袁晓曦,杨泽毅,袁梦瑶. 基于KAN和U-Net网络的颌面结构全景分割方法[J]. 浙江大学学报(工学版), 2026, 60(4): 772-781.