|
|
|
| Face photo-to-sketch synthesis based on spatially adaptive Transformer |
Weiguo WAN1( ),Kaichen ZHONG1,Yingmei ZHANG1,Li YAO2,Shuhong BAO1,Dawei GONG3 |
1. School of Software and Internet of Things Engineering, Jiangxi University of Finance and Economics, Nanchang 330032, China 2. Office of Educational Administration, Jiangxi University of Finance and Economics, Nanchang 330013, China 3. School of Intelligent Manufacturing, Jiangxi Science and Technology Normal University, Nanchang 330038, China |
|
|
|
Abstract A spatially adaptive Transformer-based generation network (SAT-StyTR) was proposed to address the issues of existing face photo-to-sketch synthesis methods, such as structural distortion, blurred textures, and artifact residue. A dual-modal feature disentanglement module (DFDM) was constructed in the encoder. The parallel branches of face content and sketch style feature extraction were employed in DFDM to achieve disentangled representation of cross-modal features. A row-column vector dynamic fusion module (VDFM) was designed, and a bidirectional gating mechanism was introduced in the VDFM to enable adaptive feature fusion of row-column vectors. Then the capability of local detail retention was significantly improved. A multi-scale perceptive discriminator (MSPD) was incorporated to optimize the realism of generated sketch textures at multiple semantic levels by leveraging the hierarchical gradient supervision mechanism. The experimental results on the CUHK, AR, XM2VTS and CUFSF datasets demonstrated that the proposed method outperformed existing methods such as Mamba-ST on SSIM, PSNR, MSE, LPIPS, and FID metrics, with its FID values improved by 19.3%, 1.0%, 15.5%, and 10.0% over those of the second-best method on the four datasets, respectively. Visual comparative analysis further validated that superior structural consistency and detail restoration were achieved in critical regions, such as facial contours.
|
|
Received: 16 July 2025
Published: 20 July 2026
|
|
|
| Fund: 国家自然科学基金资助项目(62261025,62362032);江西省自然科学基金资助项目(20242BAB25076,20232BAB212015). |
基于空间自适应Transformer的人脸素描图像生成
针对现有人脸素描图像生成方法存在的生成图像结构失真、纹理模糊及伪影残留等问题,提出基于空间自适应Transformer的生成网络(SAT-StyTR). 在编码器中构建双模态特征解耦模块(DFDM),通过并行的人脸内容与素描风格特征提取分支,实现跨模态特征的解耦表征. 设计行-列向量动态融合模块(VDFM),引入双向门控机制,以实现行、列向量的自适应特征融合,显著增强局部细节保留能力. 引入多尺度感知判别器(MSPD),利用层级式梯度监督机制在多语义粒度上优化素描图像纹理的真实感. 在CUHK、AR、XM2VTS和CUFSF数据集上的实验结果表明,所提方法在SSIM、PSNR、MSE、LPIPS和FID等指标上均优于Mamba-ST等现有方法,其FID值在4个数据集上比次优方法分别降低了19.3%、1.0%、15.5%和10.0%. 视觉对比分析进一步验证了所提方法的生成结果在五官轮廓等关键区域上具有更优的结构一致性与细节还原度.
关键词:
人脸素描图像生成,
空间自适应Transformer,
双模态特征解耦,
行-列向量动态融合,
多尺度感知判别器(MSPD)
|
|
| [1] |
GAO L, LIU F L, CHEN S Y, et al SketchFaceNeRF: sketch-based facial generation and editing in neural radiance fields[J]. ACM Transactions on Graphics, 2023, 42 (4): 159
|
|
|
| [2] |
翁丽芬, 李晨阳, 许华荣 基于GAN的分步合成人脸素描生成算法[J]. 计算机辅助设计与图形学学报, 2023, 35 (9): 1363- 1373 WENG Lifen, LI Chenyang, XU Huarong Stepwise synthetic face sketch generation algorithm based on GAN[J]. Journal of Computer-Aided Design & Computer Graphics, 2023, 35 (9): 1363- 1373
doi: 10.3724/SP.J.1089.2023.19582
|
|
|
| [3] |
ABDEL-AZIZ H G M, EBEID H M, MOSTAFA M G M. An unsupervised method for face photo-sketch synthesis and recognition [C]// Proceedings of the 7th International Conference on Information and Communication Systems. Irbid: IEEE, 2016: 221–226.
|
|
|
| [4] |
GALEA C, FARRUGIA R A. Face photo-sketch recognition using local and global texture descriptors [C]// Proceedings of the 24th European Signal Processing Conference. Budapest: IEEE, 2016: 2240–2244.
|
|
|
| [5] |
LI J, YU X, PENG C, et al Adaptive representation-based face sketch-photo synthesis[J]. Neurocomputing, 2017, 269: 152- 159
doi: 10.1016/j.neucom.2016.10.095
|
|
|
| [6] |
LIU Z S, SIU W C, CHAN H A. Learn to sketch: a fast approach for universal photo sketch [C]// Proceedings of the Asia-Pacific Signal and Information Processing Association Annual Summit and Conference. Tokyo: IEEE, 2021: 1450–1457.
|
|
|
| [7] |
ZHANG L, LIN L, WU X, et al. End-to-end photo-sketch generation via fully convolutional representation learning [C]// Proceedings of the 5th ACM on International Conference on Multimedia Retrieval. Shanghai: ACM, 2015: 627–634.
|
|
|
| [8] |
GALEA C, FARRUGIA R A Forensic face photo-sketch recognition using a deep learning-based architecture[J]. IEEE Signal Processing Letters, 2017, 24 (11): 1586- 1590
doi: 10.1109/LSP.2017.2749266
|
|
|
| [9] |
ZHU M, LI J, WANG N, et al A deep collaborative framework for face photo-sketch synthesis[J]. IEEE Transactions on Neural Networks and Learning Systems, 2019, 30 (10): 3096- 3108
doi: 10.1109/TNNLS.2018.2890018
|
|
|
| [10] |
曹林, 王震, 杜康宁, 等 基于层次对比生成对抗网络的非配对素描人脸合成[J]. 中国科技论文, 2024, 19 (6): 715- 723 CAO Lin, WANG Zhen, DU Kangning, et al Unpaired sketch face synthesis based on hierarchical contrast generative adversarial network[J]. China Sciencepaper, 2024, 19 (6): 715- 723
doi: 10.3969/j.issn.2095-2783.2024.06.011
|
|
|
| [11] |
ZHU M, LIANG C, WANG N, et al. A sketch-Transformer network for face photo-sketch synthesis [C]// Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence. Montreal: IJCAI, 2021: 1352–1358.
|
|
|
| [12] |
CHEN C, YE M, QI M, et al. Sketch Transformer: asymmetrical disentanglement learning from dynamic synthesis [C]// Proceedings of the 30th ACM International Conference on Multimedia. Lisboa: ACM, 2022: 4012–4020.
|
|
|
| [13] |
YU W, ZHU M, WANG N, et al An efficient Transformer based on global and local self-attention for face photo-sketch synthesis[J]. IEEE Transactions on Image Processing, 2023, 32: 483- 495
doi: 10.1109/TIP.2022.3229614
|
|
|
| [14] |
陈长武, 曹林, 郭亚男, 等 基于域自适应均值网络的素描人脸识别方法[J]. 计算机应用与软件, 2023, 40 (4): 107- 115 CHEN Changwu, CAO Lin, GUO Yanan, et al Sketch face recognition based on domain adaptation mean network[J]. Computer Applications and Software, 2023, 40 (4): 107- 115
|
|
|
| [15] |
SHI Z, WAN W Transformer-based adversarial network for semi-supervised face sketch synthesis[J]. Journal of Visual Communication and Image Representation, 2024, 102: 104204
doi: 10.1016/j.jvcir.2024.104204
|
|
|
| [16] |
DENG Y, TANG F, DONG W, et al. StyTr2: image style transfer with Transformers [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. New Orleans: IEEE, 2022: 11316–11326.
|
|
|
| [17] |
SUN J, LIU X, BÄCK T, et al Learning adaptive differential evolution algorithm from optimization experiences by policy gradient[J]. IEEE Transactions on Evolutionary Computation, 2021, 25 (4): 666- 680
doi: 10.1109/TEVC.2021.3060811
|
|
|
| [18] |
LI Y, CHEN X, YANG B, et al. DeepFacePencil: creating face images from freehand sketches [C]// Proceedings of the 28th ACM International Conference on Multimedia. Seattle: ACM, 2020: 991–999.
|
|
|
| [19] |
KARNEWAR A, WANG O. MSG-GAN: multi-scale gradients for generative adversarial networks [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle: IEEE, 2020: 7796–7805.
|
|
|
| [20] |
WANG X, TANG X Face photo-sketch synthesis and recognition[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2009, 31 (11): 1955- 1967
doi: 10.1109/TPAMI.2008.222
|
|
|
| [21] |
ZHANG W, WANG X, TANG X. Coupled information-theoretic encoding for face photo-sketch recognition [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Colorado Springs: IEEE, 2011: 513–520.
|
|
|
| [22] |
KINGMA D P, BA J. Adam: a method for stochastic optimization [EB/OL]. (2017-01-30) [2025-07-15]. https://arxiv.org/abs/1412.6980.
|
|
|
| [23] |
XIONG R, YANG Y, HE D, et al. On layer normalization in the Transformer architecture [C]// Proceedings of the 37th International Conference on Machine Learning. [S.l.]: PMLR, 2020: 10524–10533.
|
|
|
| [24] |
GOODFELLOW I, BENGIO Y, COURVILLE A. Optimization for training deep models [M]// Deep learning. Cambridge: MIT Press, 2016: 291–293.
|
|
|
| [25] |
ISOLA P, ZHU J Y, ZHOU T, et al. Image-to-image translation with conditional adversarial networks [C]// Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Honolulu: IEEE, 2017: 5967–5976.
|
|
|
| [26] |
KIM J, KIM M, KANG H, et al. U-GAT-IT: unsupervised generative attentional networks with adaptive layer-instance normalization for image-to-image translation [EB/OL]. (2020-04-08) [2025-07-15]. https://arxiv.org/abs/1907.10830.
|
|
|
| [27] |
HAN J, SHOEIBY M, PETERSSON L, et al. Dual contrastive learning for unsupervised image-to-image translation [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops. Nashville: IEEE, 2021: 746–755.
|
|
|
| [28] |
WEN L, GAO C, ZOU C. CAP-VSTNet: content affinity preserved versatile style transfer [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Vancouver: IEEE, 2023: 18300–18309.
|
|
|
| [29] |
BOTTI F, ERGASTI A, ROSSI L, et al. Mamba-ST: state space model for efficient style transfer [EB/OL]. (2024-09-16) [2025-07-15]. https://arxiv.org/abs/2409.10385.
|
|
|
| [30] |
WANG Z, BOVIK A C, SHEIKH H R, et al Image quality assessment: from error visibility to structural similarity[J]. IEEE Transactions on Image Processing, 2004, 13 (4): 600- 612
doi: 10.1109/TIP.2003.819861
|
|
|
| [31] |
KORHONEN J, YOU J. Peak signal-to-noise ratio revisited: is simple beautiful? [C]// Proceedings of the Fourth International Workshop on Quality of Multimedia Experience. Melbourne: IEEE, 2012: 37–38.
|
|
|
| [32] |
MARMOLIN H Subjective MSE measures[J]. IEEE Transactions on Systems, Man, and Cybernetics, 1986, 16 (3): 486- 489
doi: 10.1109/TSMC.1986.4308985
|
|
|
| [33] |
KETTUNEN M, HÄRKÖNEN E, LEHTINEN J. E-LPIPS: robust perceptual image similarity via random transformation ensembles [EB/OL]. (2019-06-11) [2025-07-15]. https://arxiv.org/abs/1906.03973.
|
|
|
|
Viewed |
|
|
|
Full text
|
|
|
|
|
Abstract
|
|
|
|
|
Cited |
|
|
|
|
| |
Shared |
|
|
|
|
| |
Discussed |
|
|
|
|