Please wait a minute...
浙江大学学报(工学版)  2026, Vol. 60 Issue (8): 1730-1738    DOI: 10.3785/j.issn.1008-973X.2026.08.012
计算机技术     
基于多尺度外观运动融合的视频异常检测
赵明华1,2(),吕雨萱1,吕佳豪1(),陈逸飞1,石程1,胡静1
1. 西安理工大学 计算机科学与工程学院,陕西 西安 710048
2. 陕西省网络计算与安全技术重点实验室,陕西 西安 710048
Video anomaly detection based on multi-scale appearance and motion fusion
Minghua ZHAO1,2(),Yuxuan LYU1,Jiahao LYU1(),Yifei CHEN1,Cheng SHI1,Jing HU1
1. School of Computer Science and Engineering, Xi’an University of Technology, Xi’an 710048, China
2. Shaanxi Key Laboratory of Network Computing and Security Technology, Xi’an 710048, China
 全文: PDF(2340 KB)   HTML
摘要:

现有的基于帧预测的视频异常检测方法仅关注外观特征而忽略运动信息,难以同时捕捉局部变化与全局异常,为此,提出基于多尺度外观运动融合的视频异常检测方法. 设计基于Inception模块的双流编码器,用于提取多尺度的外观特征与光流特征;将不同尺度外观和运动特征的融合特征输入动态原型单元模块,进行正常特征表征学习;在外观编码器和解码器之间的跳跃连接过程中添加基于CNN-Transformer的混合注意力机制,实现全局上下文信息与局部细节信息的融合,以提升检测性能. 在3个基准数据集UCSD Ped2、CUHK Avenue和ShanghaiTech上进行广泛实验和消融研究,AUC值分别达到了99.1%、88.5%和74.5%,充分验证了所提方法的有效性和优越性.

关键词: 视频异常检测卷积神经网络时序运动特征混合注意力机制无监督学习    
Abstract:

A new video anomaly detection method was proposed to address the issue that existing frame prediction-based video anomaly detection methods focus only on appearance features while ignoring motion information and are difficult to simultaneously capture local changes and global anomalies. A dual-stream encoder based on the Inception module was designed to extract multi-scale appearance and optical flow features. The fused features of appearance and motion features at different scales were input into a dynamic prototype unit module to learn the representations of normal features. A hybrid attention mechanism based on CNN-Transformer was added to the skip connection process between the appearance encoder and decoder to realize the integration of global contextual information and local detailed information, therefore enhancing the detection performance. Extensive experiments and ablation studies were conducted on three benchmark datasets: UCSD Ped2, CUHK Avenue, and ShanghaiTech. The AUC values of 99.1%, 88.5%, and 74.5% were achieved, respectively, demonstrating the effectiveness and superiority of the proposed method.

Key words: video anomaly detection    convolutional neural network    temporal motion feature    hybrid attention mechanism    unsupervised learning
收稿日期: 2025-07-11 出版日期: 2026-07-16
CLC:  TP 331.41  
基金资助: 陕西省自然科学基金资助项目(2024JC-ZDXM-35, 2024JC-YBMS-573, 2024JC-YBMS-458);陕西省科协青年人才托举计划资助项目(20240146);西安理工大学博士创新基金资助项目(BC202621).
作者简介: 赵明华(1979—),女,教授,从事计算机视觉与模式识别研究. orcid.org/0000-0001-8062-2982. E-mail:zhaominghua@xaut.edu.cn
服务  
把本文推荐给朋友
加入引用管理器
E-mail Alert
作者相关文章  
赵明华
吕雨萱
吕佳豪
陈逸飞
石程
胡静

引用本文:

赵明华,吕雨萱,吕佳豪,陈逸飞,石程,胡静. 基于多尺度外观运动融合的视频异常检测[J]. 浙江大学学报(工学版), 2026, 60(8): 1730-1738.

Minghua ZHAO,Yuxuan LYU,Jiahao LYU,Yifei CHEN,Cheng SHI,Jing HU. Video anomaly detection based on multi-scale appearance and motion fusion. Journal of ZheJiang University (Engineering Science), 2026, 60(8): 1730-1738.

链接本文:

https://www.zjujournals.com/eng/CN/10.3785/j.issn.1008-973X.2026.08.012        https://www.zjujournals.com/eng/CN/Y2026/V60/I8/1730

图 1  基于多尺度外观运动融合的视频异常检测方法框架图
图 2  混合注意力机制(GLmix)结构图
图 3  正常和异常样本示例
数据集Nf_tr/Nf_teNv_tr/Nv_teNs
Ped22 550 / 2 01016 / 121
Avenue15 328 / 15 32416 / 211
SHT274 515 / 42 883330 / 10713
表 1  3个数据集的帧数、视频数与场景数统计
类型方法AUC/%
Ped2AvenueSHT
重建MemAE[1]94.183.371.2
Astrid等[36]94.884.972.5
GMFC-VAE[37]92.283.4
CR-AE[38]95.673.1
MESDnet[35]95.686.373.2
预测Park等[9]97.088.570.5
Liu等[8]95.485.172.8
MPN[19]96.989.573.8
Conv-VRNN[12]96.185.8
FSCN[13]92.885.5
AMMC-Net[11]96.686.673.7
Li等[16]96.587.173.6
PDM-Net[39]97.788.174.2
GroupGAN[40]96.688.174.2
AMAE[41]97.488.273.6
本研究99.188.574.5
表 2  所提方法与先进方法在AUC指标上的对比
图 4  3个数据集上的视频异常检测误差可视化图
图 5  异常分数变化曲线
${E_{\mathrm{m}}}$DPUMASFFGLmixAUC/%
Ped2SHT
98.173.1
97.573.9
97.374.1
96.974.0
99.174.5
表 3  不同模块组合在不同数据集上的消融实验结果
模块AUC/%
Ped2SHT
${E_{\mathrm{m}}}$98.474.1
DPU97.473.9
MASFF97.073.8
GLmix97.973.6
表 4  单一模块在不同数据集上的消融实验结果
方法Np/MFPS/(帧·s?1)
MemAE[1]6.531
Li等[16]30
Liu等[8]349.025
本研究17.837
表 5  所提方法和部分先进方法的参数量和处理速率对比
1 GONG D, LIU L, LE V, et al. Memorizing normality to detect anomaly: memory-augmented deep autoencoder for unsupervised anomaly detection [C]// Proceedings of the IEEE/CVF International Conference on Computer Vision. Seoul: IEEE, 2019: 1705–1714.
2 RAMACHANDRA B, JONES M J, VATSAVAI R R A survey of single-scene video anomaly detection[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022, 44 (5): 2293- 2312
3 YE M, SHEN J, LIN G, et al Deep learning for person re-identification: a survey and outlook[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022, 44 (6): 2872- 2893
4 SABOKROU M, FATHY M, HOSEINI M, et al. Real-time anomaly detection and localization in crowded scenes [C]// Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops. Boston: IEEE, 2015: 56–62.
5 HASAN M, CHOI J, NEUMANN J, et al. Learning temporal regularity in video sequences [C]// Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Las Vegas: IEEE, 2016: 733–742.
6 FENG J, LIANG Y, LI L. Anomaly detection in videos using two-stream autoencoder with post hoc interpretability [J]. Computational Intelligence and Neuroscience, 2021: 7367870.
7 CHEN D, YUE L, CHANG X, et al NM-GAN: noise-modulated generative adversarial network for video anomaly detection[J]. Pattern Recognition, 2021, 116: 107969
doi: 10.1016/j.patcog.2021.107969
8 LIU W, LUO W, LIAN D, et al. Future frame prediction for anomaly detection: a new baseline [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Salt Lake City: IEEE, 2018: 6536–6545.
9 PARK H, NOH J, HAM B. Learning memory-guided normality for anomaly detection [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle: IEEE, 2020: 14360–14369.
10 梁家菲, 李婷, 杨佳琪, 等 融合自注意力和自编码器的视频异常检测[J]. 中国图象图形学报, 2023, 28 (4): 1029- 1040
LIANG Jiafei, LI Ting, YANG Jiaqi, et al Video anomaly detection by fusing self-attention and autoencoder[J]. Journal of Image and Graphics, 2023, 28 (4): 1029- 1040
doi: 10.11834/jig.211147
11 CAI R, ZHANG H, LIU W, et al Appearance-motion memory consistency network for video anomaly detection[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2021, 35 (2): 938- 946
doi: 10.1609/aaai.v35i2.16177
12 LU Y, KUMAR K M, NABAVI S S, et al. Future frame prediction using convolutional VRNN for anomaly detection [C]// Proceedings of the 16th IEEE International Conference on Advanced Video and Signal Based Surveillance. Taipei: IEEE, 2019: 1–8.
13 WU P, LIU J, LI M, et al Fast sparse coding networks for anomaly detection in videos[J]. Pattern Recognition, 2020, 107: 107515
doi: 10.1016/j.patcog.2020.107515
14 LU Y, YU F, REDDY M K K, et al. Few-shot scene-adaptive anomaly detection [C]// European Conference on Computer Vision. [S.l.]: Springer, 2020: 125–141.
15 LYU J, ZHAO M, HU J, et al Bidirectional skip-frame prediction for video anomaly detection with intra-domain disparity-driven attention[J]. Pattern Recognition, 2026, 170: 112010
doi: 10.1016/j.patcog.2025.112010
16 LI D, NIE X, GONG R, et al Multi-branch GAN-based abnormal events detection via context learning in surveillance videos[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2023, 34 (5): 3439- 3450
17 LUO W, LIU W, GAO S. A revisit of sparse coding based anomaly detection in stacked RNN framework [C]// Proceedings of the IEEE International Conference on Computer Vision. Venice: IEEE, 2017: 341–349.
18 SZEGEDY C, IOFFE S, VANHOUCKE V, et al. Inception-v4, inception-ResNet and the impact of residual connections on learning [C]// Proceedings of the 31st AAAI Conference on Artificial Intelligence. San Francisco: AAAI Press, 2017: 4278–4284.
19 LV H, CHEN C, CUI Z, et al. Learning normal dynamics in videos with meta prototype network [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Nashville: IEEE, 2021: 15420–15429.
20 LI Z, LI N, JIANG K, et al. Superpixel masking and inpainting for self-supervised anomaly detection [C]// Proceedings of the British Machine Vision Conference. [S.l.]: BMVA Press, 2020.
21 GOODFELLOW I, POUGET-ABADIE J, MIRZA M, et al Generative adversarial networks[J]. Communications of the ACM, 2020, 63 (11): 139- 144
doi: 10.1145/3422622
22 ZAHEER M Z, LEE J H, MAHMOOD A, et al Stabilizing adversarially learned one-class novelty detection using pseudo anomalies[J]. IEEE Transactions on Image Processing, 2022, 31: 5963- 5975
doi: 10.1109/TIP.2022.3204217
23 CHEN Y, LIU J, PENG L, et al Auto-encoding variational Bayes[J]. Cambridge Explorations in Arts and Sciences, 2024, 2 (1): 1
24 GRAVES A, GRAVES A. Long short-term memory [M]// Supervised sequence labelling with recurrent neural networks. Berlin, Heidelberg: Springer, 2012: 37–45.
25 CHEN D, WANG P, YUE L, et al Anomaly detection in surveillance video based on bidirectional prediction[J]. Image and Vision Computing, 2020, 98: 103915
doi: 10.1016/j.imavis.2020.103915
26 SUN S, GONG X. Hierarchical semantic contrast for scene-aware video anomaly detection [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Vancouver: IEEE, 2023: 22846–22856.
27 LIU Z, NIE Y, LONG C, et al. A hybrid video anomaly detection framework via memory-augmented flow reconstruction and flow-guided frame prediction [C]// Proceedings of the IEEE/CVF International Conference on Computer Vision. Montreal: IEEE, 2021: 13568–13577.
28 DOSOVITSKIY A, FISCHER P, ILG E, et al. FlowNet: learning optical flow with convolutional networks [C]// Proceedings of the IEEE International Conference on Computer Vision. Santiago: IEEE, 2015: 2758–2766.
29 ZHONG Y, CHEN X, HU Y, et al Bidirectional spatio-temporal feature learning with multiscale evaluation for video anomaly detection[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2022, 32 (12): 8285- 8296
doi: 10.1109/TCSVT.2022.3190539
30 MAHADEVAN V, LI W, BHALODIA V, et al. Anomaly detection in crowded scenes [C]// Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition. San Francisco: IEEE, 2010: 1975–1981.
31 LU C, SHI J, JIA J. Abnormal event detection at 150 FPS in MATLAB [C]// Proceedings of the IEEE International Conference on Computer Vision. Sydney: IEEE, 2013: 2720–2727.
32 LOSHCHILOV I, HUTTER F. SGDR: stochastic gradient descent with warm restarts [EB/OL]. (2017–05–03) [2025–04–06]. https://arxiv.org/abs/1608.03983.
33 PASZKE A, GROSS S, CHINTALA S, et al. Automatic differentiation in PyTorch [EB/OL]. (2017–10–29) [2025–04–06]. https://openreview.net/forum?id=BJJsrmfCZ.
34 ZHANG X, FANG J, YANG B, et al Hybrid attention and motion constraint for anomaly detection in crowded scenes[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2023, 33 (5): 2259- 2274
doi: 10.1109/TCSVT.2022.3221622
35 FANG Z, ZHOU J T, XIAO Y, et al Multi-encoder towards effective anomaly detection in videos[J]. IEEE Transactions on Multimedia, 2020, 23: 4106- 4116
36 ASTRID M, ZAHEER M Z, LEE J Y, et al. Learning not to reconstruct anomalies [EB/OL]. (2021–10–24) [2025–04–06]. https://arxiv.org/abs/2110.09742.
37 FAN Y, WEN G, LI D, et al Video anomaly detection and localization via Gaussian mixture fully convolutional variational autoencoder[J]. Computer Vision and Image Understanding, 2020, 195: 102920
doi: 10.1016/j.cviu.2020.102920
38 WANG B, YANG C Video anomaly detection based on convolutional recurrent AutoEncoder[J]. Sensors, 2022, 22 (12): 4647
doi: 10.3390/s22124647
39 HUANG C, WEN J, LIU C, et al. Long short-term dynamic prototype alignment learning for video anomaly detection [C]// Proceedings of the 33rd International Joint Conference on Artificial Intelligence. Jeju: Curran Associates Inc, 2024: 866–874.
40 SUN Z, WANG P, ZHENG W, et al Dual GroupGAN: an unsupervised four-competitor (2V2) approach for video anomaly detection[J]. Pattern Recognition, 2024, 153: 110500
doi: 10.1016/j.patcog.2024.110500
[1] 苏挺,许月萍,WANGQuanjun,钟华,蒋建群. 基于物理信息的城市洪涝多阶段替代模型[J]. 浙江大学学报(工学版), 2026, 60(7): 1557-1566.
[2] 张雪梅,孙颖,张雪英. 基于无监督图对比学习的语音情感识别[J]. 浙江大学学报(工学版), 2026, 60(4): 782-790.
[3] 方芳,严军,郭红想,王勇. 基于时空注意力机制的轻量级脑纹识别算法[J]. 浙江大学学报(工学版), 2026, 60(3): 633-642.
[4] 闫光辉,黄霄,常文文. 基于脑电多尺度特征和图神经网络的紧急制动行为识别[J]. 浙江大学学报(工学版), 2026, 60(2): 404-414.
[5] 魏新雨,饶蕾,范光宇,陈年生,程松林,杨定裕. 用于无人机遥感图像的高精度实时语义分割网络[J]. 浙江大学学报(工学版), 2025, 59(7): 1411-1420.
[6] 王立红,刘新倩,李静,冯志全. 基于联邦学习和时空特征融合的网络入侵检测方法[J]. 浙江大学学报(工学版), 2025, 59(6): 1201-1210.
[7] 张梦瑶,周杰,李文婷,赵勇. 结合全局信息和局部信息的三维网格分割框架[J]. 浙江大学学报(工学版), 2025, 59(5): 912-919.
[8] 钱新宇,谢清林,陶功权,温泽峰. 基于多结构数据驱动的车轮扁疤定量识别方法[J]. 浙江大学学报(工学版), 2025, 59(4): 688-697.
[9] 安春丽,张碧玲,赵国安,王博,刘岩. 基于FFT-CNN-GCN的电网故障诊断[J]. 浙江大学学报(工学版), 2025, 59(10): 2205-2212.
[10] 付承彪,庄清源,田安红. 基于多通道卷积方式的土壤重金属镍含量定量预测[J]. 浙江大学学报(工学版), 2025, 59(10): 2221-2228.
[11] 赖凌轩,柳景青,周一粟,李秀娟. 基于时频卷积神经网络的供水管道漏损识别[J]. 浙江大学学报(工学版), 2025, 59(1): 196-204.
[12] 王海军,王涛,俞慈君. 基于递归量化分析的CFRP超声检测缺陷识别方法[J]. 浙江大学学报(工学版), 2024, 58(8): 1604-1617.
[13] 李劲业,李永强. 融合知识图谱的时空多图卷积交通流量预测[J]. 浙江大学学报(工学版), 2024, 58(7): 1366-1376.
[14] 邢志伟,朱书杰,李彪. 基于改进图卷积神经网络的航空行李特征感知[J]. 浙江大学学报(工学版), 2024, 58(5): 941-950.
[15] 唐善成,逯建辉,张莹,金子成,赵安新. 修复缺陷嫌疑区域的无监督磁瓦表面缺陷检测[J]. 浙江大学学报(工学版), 2024, 58(4): 718-728.