Submitted:
18 August 2025
Posted:
19 August 2025
You are already at the latest version
Abstract
Keywords:
1. Introduction
2. Related Work
3. Materials and Methods
3.1. Overall Framework
3.2. Input Preprocessing

- frames: a 4D tensor of shape ,
- label: a scalar indicating the normal (0) or abnormal (1) category.
3.3. Handling and Evaluation of UCSD Ped2 Dataset
3.4. Training Environments
3.5. Generator Architecture
3.6. Discriminator Architecture
3.7. Loss Function
Generator Loss: DASLoss
Pixel Reconstruction Loss () :
Feature Matching Loss () :
Temporal Smoothness Loss () :
Perceptual Loss () :
Discriminator Loss
3.8. Anomaly Scoring and Thresholding
- and are the ground-truth and reconstructed frames at time t.
- denotes the discriminator’s final intermediate feature layer.
- and are weighting coefficients for pixel and feature components.
4. Results
4.1. Evaluation Metrics and Setup
- F1-Score. The F1-score is the harmonic mean of precision and recall, and is defined as:
- Average Precision (AP). AP measures the area under the precision-recall curve and is defined as:
- Area Under the Curve (AUC). AUC corresponds to the area under the Receiver Operating Characteristic (ROC) curve and is defined as:
| Parameter | Value / Description |
|---|---|
| Input frame size | |
| Input sequence length | 6 frames (flow) + 6 frames (RGB) |
| Batch size | 4 |
| Optimizer | Adam |
| Learning rate | |
| Training epochs | 400 (with early stopping) |
| 1.0 | |
| 0.1 | |
| 0.1 | |
| 0.01 |
4.2. Quantitative Results on Benchmark Datasets
5. Discussion
5.1. Performance Analysis Across Datasets
5.2. Architectural Contributions
6. Conclusions and Future Work
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Gong, D.; Liu, L.; Le, V.; Saha, B.; Mansour, M.R.; Venkatesh, S.; van den Hengel, A. Memorizing Normality to Detect Anomaly: Memory-Augmented Deep Autoencoder for Unsupervised Anomaly Detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Korea, 27 October–2 November 2019; pp. 1705–1714.
- Sabokrou, M.; Khalooei, M.; Fayyaz, M.; Adeli, E. Adversarially Learned One-Class Classifier for Novelty Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–22 June 2018; pp. 3379–3388.
- Liu, W.; Luo, W.; Lian, D.; Gao, S. Future Frame Prediction for Anomaly Detection – A New Baseline. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–22 June 2018; pp. 6536–6545.
- Nguyen, H.; Liu, V.; Prasad, S.; Tran, D. Anomaly Detection in Video Sequence with Appearance-Motion Correspondence. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), Seoul, Korea, 27 October–2 November 2019; pp. 1272–1281.
- Sultani, W.; Chen, C.; Shah, M. Real-World Anomaly Detection in Surveillance Videos. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–22 June 2018; pp. 6479–6488.
- Tran, D.; Wang, H.; Torresani, L.; Ray, J.; LeCun, Y.; Paluri, M. A Closer Look at Spatiotemporal Convolutions for Action Recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–22 June 2018; pp. 6450–6459.
- Zolfaghari, M.; Singh, K.; Brox, T. ECO: Efficient Convolutional Network for Online Video Understanding. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; pp. 695–712.
- Luo, W.; Liu, W.; Gao, S. A Revisit of Sparse Coding Based Anomaly Detection in Stacked RNN Framework. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy, 22–29 October 2017; pp. 341–349.
- Dosovitskiy, A.; Fischer, P.; Ilg, E.; Hausser, P.; Hazirbas, C.; Golkov, V.; van der Smagt, P.; Cremers, D.; Brox, T. FlowNet: Learning Optical Flow with Convolutional Networks. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), Santiago, Chile, 7–13 December 2015; pp. 2758–2766.
- Sun, D.; Yang, X.; Liu, M.Y.; Kautz, J. PWC-Net: CNNs for Optical Flow Using Pyramid, Warping, and Cost Volume. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–22 June 2018; pp. 8934–8943.
- Lu, C.; Shi, J.; Jia, J. Abnormal Event Detection at 150 FPS in MATLAB. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), Sydney, NSW, Australia, 1–8 December 2013; pp. 2720–2727.
- Carreira, J.; Zisserman, A. Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; pp. 6299–6308.
- Park, H.; Kim, T.; Oh, J.; Kim, C.; Kim, G. MGAFlow: Motion-Guided Attention Flow Network for Video Anomaly Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada, 18–22 June 2023; pp. 12861–12870.
- Georgescu, M.I.; Ionescu, R.T.; Popescu, M.; Khan, F.S.; Shao, L. Anomaly Detection in Video via Self-Supervised and Multi-Task Learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Virtual, 19–25 June 2021; pp. 14366–14375.
- Cho, M.; Kim, M.; Shin, Y. Weakly-Supervised Video Anomaly Detection via Robust Temporal Feature Alignment. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada, 18–22 June 2023; pp. 14213–14222.
- Xu, H.; Zhang, J.; Cai, J.; Rezatofighi, H.; Yu, F.; Tao, D.; Geiger, A. Unifying Flow, Stereo and Depth Estimation. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2023. doi:10.1109/TPAMI.2023.3242203.
- Nievas, E.B.; Suarez, O.D.; Garcia, G.B.; Sukthankar, R. Violence detection in video using computer vision techniques. In Computer Analysis of Images and Patterns (CAIP), Seville, Spain, 29–31 August 2011; pp. 332–339.
- Mahadevan, V.; Li, W.; Bhalodia, V.; Vasconcelos, N. Anomaly detection in crowded scenes. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), San Francisco, CA, USA, 13–18 June 2010; pp. 1975–1981.
- Wu, J.; Zhang, W.; Gao, C.; He, Z.; Jiao, J. Not only look, but also listen: Learning multimodal violence detection under weak supervision. In Proceedings of the European Conference on Computer Vision (ECCV), Glasgow, UK, 23–28 August 2020; pp. 322–339.
- Yang, Z.; Wu, Y.; Huang, Q.; Loy, C.C. DiffusionAD: Masked Spatiotemporal Autoencoding for Video Anomaly Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada, 18–22 June 2023; pp. 3055–3064.
- Akçay, S.; Atapour-Abarghouei, A.; Breckon, T.P. Ganomaly: Semi-supervised Anomaly Detection via Adversarial Training. In Proceedings of the Asian Conference on Computer Vision (ACCV), Perth, Australia, 2–6 December 2018; pp. 622–637.
- Tian, Y.; Chen, X.; Wu, J.; Zha, Z.J.; Zhang, Z. Cross-feature Attention with Hierarchical Local-Global Modeling for Video Anomaly Detection. In Proceedings of the 29th ACM International Conference on Multimedia (MM), Virtual Event, 20–24 October 2021; pp. 1126–1135.
- Zhou, W.; Wu, W.; Lin, W. Spatio-temporal Predictive Memory Network for Video Anomaly Detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 11–17 October 2021; pp. 13929–13938.
- Yu, P.; Lu, H.; Li, L. STADNet: Spatial-Temporal-Appearance Distribution-Aware Network for Video Anomaly Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 19–24 June 2022; pp. 20833–20842.
- Tang, Y.; Wang, L.; Liu, L.; Li, B.; Wu, Y.; Gao, J. Integrating Multimodal and Temporal Perception for Dynamic Event Understanding. In Proceedings of the European Conference on Computer Vision (ECCV), Glasgow, UK, 23–28 August 2020; pp. 312–328.
- Shi, X.; Chen, Z.; Wang, H.; Yeung, D.Y.; Wong, W.K.; Woo, W.C. Convolutional LSTM Network: A Machine Learning Approach for Precipitation Nowcasting. In Advances in Neural Information Processing Systems (NeurIPS), Montreal, QC, Canada, 7–12 December 2015; pp. 802–810.
- Tang, Y.; Wu, J.; Song, Y.; Liu, Y.; Liu, W. RTFM: A General Framework for Video Anomaly Detection via Robust Temporal Feature Magnitude Learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 11–17 October 2021; pp. 15804–15813.
- Kim, J.; Grauman, K. Observe locally, infer globally: A space-time MRF for detecting abnormal activities with incremental updates. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Miami, FL, USA, 20–25 June 2009; pp. 2921–2928.
- Hasan, M.; Choi, J.; Neumann, J.; Roy-Chowdhury, A.K.; Davis, L.S. Learning temporal regularity in video sequences. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 26 June–1 July 2016; pp. 733–742.
- Ji, S.; Xu, W.; Yang, M.; Yu, K. 3D convolutional neural networks for human action recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2013, 35(1), 221–231.
- Ionescu, R.T.; Smeureanu, S.; Alexe, B.; Popescu, M. Object-centric auto-encoders and dummy anomalies for abnormal event detection in video. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 15–20 June 2019; pp. 7842–7851.
- Ravanbakhsh, M.; Nabi, M.; Sangineto, E.; Sebe, N. Plug-and-play CNN for crowd motion analysis: An application in abnormal event detection. In Proceedings of the IEEE Winter Conference on Applications of Computer Vision (WACV), Lake Tahoe, NV, USA, 12–15 March 2018; pp. 1689–1698.
- Salimans, T.; Goodfellow, I.; Zaremba, W.; Cheung, V.; Radford, A.; Chen, X. Improved techniques for training GANs. In Advances in Neural Information Processing Systems (NeurIPS), Barcelona, Spain, 5–10 December 2016; pp. 2234–2242.
- Mathieu, M.; Couprie, C.; LeCun, Y. Deep multi-scale video prediction beyond mean square error. In Proceedings of the International Conference on Learning Representations (ICLR), San Juan, Puerto Rico, 2–4 May 2016.
- Villegas, R.; Yang, J.; Hong, S.; Lin, X.; Lee, H. Decomposing motion and content for natural video sequence prediction. In Proceedings of the International Conference on Learning Representations (ICLR), Toulon, France, 24–26 April 2017.
- Johnson, J.; Alahi, A.; Fei-Fei, L. Perceptual losses for real-time style transfer and super-resolution. In Proceedings of the European Conference on Computer Vision (ECCV), Amsterdam, The Netherlands, 11–14 October 2016; pp. 694–711.
- Mao, X.; Li, Q.; Xie, H.; Lau, R.Y.; Wang, Z.; Smolley, S.P. Least squares generative adversarial networks. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy, 22–29 October 2017; pp. 2794–2802.
- Youden, W.J. Index for rating diagnostic tests. Cancer, 1950, 3(1), 32–35.





| Model | Method | AP (%) |
|---|---|---|
| MGAFlow [13] | Unsupervised | 75.3 |
| DiffusionAD [20] | Unsupervised | 68.0 |
| STPM [23] | Unsupervised | 61.0 |
| MemAE [1] | Unsupervised | 63.0 |
| Flashback [14] | Zero-shot | 75.1 |
| Our model | Unsupervised | 80.5 |
| Model | Method | AUC | F1-score |
|---|---|---|---|
| AnoGAN [2] | Unsupervised | 0.63 | 0.57 |
| GANomaly [21] | Semi-supervised | 0.71 | 0.59 |
| MemAE [1] | Unsupervised | 0.73 | 0.65 |
| CFA-HLGAtt [22] | Unsupervised | 0.85 | 0.82 |
| RTFM [27] | Weakly-supervised | 0.85 | 0.73 |
| ConvLSTM-AE [8] | Unsupervised | 0.89 | 0.86 |
| Our model | Unsupervised | 0.92 | 0.85 |
| Model | Method | AUC |
|---|---|---|
| MPPCA [28] | Unsupervised | 0.69 |
| MDT [18] | Unsupervised | 0.83 |
| Conv2D-AE [29] | Unsupervised | 0.85 |
| ConvLSTM-AE [8] | Unsupervised | 0.88 |
| Conv3D-AE [30] | Unsupervised | 0.91 |
| STADNet [24] | Unsupervised | 0.95 |
| CR-AE [31] | Unsupervised | 0.96 |
| Optical Flow + STC + GAN [32] | Unsupervised | 0.97 |
| Multi-level Memory-augmented AMC [25] | Unsupervised | 0.99 |
| Our model | Unsupervised | 0.96 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).