Submitted:
01 November 2023
Posted:
02 November 2023
You are already at the latest version
Abstract

Keywords:
1. Introduction
1.1. Problem
1.2. Motivation
1.3. Objetives
- Generate a dataset from real surveillance cameras, which integrates characteristics of real human violence.
- Develop an effective model in terms of accuracy based on attention and temporal fusion mechanisms.
- Develop an efficient model in terms of the number of parameters and FLOPS based on temporal changes and 2D CNN.
- Develop a compact model for recognizing violent human actions in video surveillance, with minimal latency times close to real-time.
1.4. Contributions
- A model for recognizing violent human actions, which enables its use in a real scenario and in real-time.
- An efficient model in recognizing violent human actions, in terms of the number of parameters and FLOPS. In turn, the model is effective in recognition, in terms of accuracy, whose results contribute to the state-of-the-art.
- Likewise, our proposal contributes to the scientific community by generating and publishing a dataset oriented to the domain of video surveillance.
1.5. Work organization
2. Related work
2.1. Related work to the proposal
2.1.1. Extraction of regions of interest
2.1.2. Short-duration spatiotemporal feature extraction
2.1.3. Global spatiotemporal feature extraction
2.2. Reference backbones
2.3. Benchmark
2.3.1. Benchmark on classic datasets
2.3.2. Benchmark on RWF-2000 dataset
3. Proposal
3.1. Proposal architecture
3.2. Spatial Motion Extractor module(SME)
3.3. Short Temporal Extractor module (STE)
3.4. Global Temporal Extractor (GTE)
4. Results
4.1. Datasets
4.1.1. Hockey Fights dataset
4.1.2. Movies dataset
4.1.3. RWF-2000 dataset
4.1.4. VioPeru dataset
- The violent scene involves two people, several people, or crowds.
- Cameras have different resolutions.
- There is a different proportion of the violent scene concerning the size of the frames; that is, the violent action is large or so small that it can go unnoticed by the human eye.
- Violent human actions occur primarily at night when lighting can negatively influence detection.
- Violence in video surveillance is not only made up of fights; there is also looting, vandalism, violent protests, attacks on property, and confrontations between groups of people.
- Occlusion is typical in video surveillance; that is, the same people, trees, and vehicles, among others, cover violent scenes.
4.2. Model configuration
- Learning rate: , for all datasets.
- Batchsize: 2.
- Number of epochs: 100.
- Optimizer Adam, with textitEpsilon: , weight decay: , and to calculate the loss function: Cross Entropy.
- One-Cycle Learning Rate Scheduler, with min-lr: , patience: 2, and factor: 0.5.
4.3. Results evaluation
4.3.1. Results evaluation on classical datasets
4.3.2. Results evaluation on RWF-2000 datasets

4.3.3. Results evaluation on VioPeru dataset
4.3.4. Results evaluation in Real-Time
5. Ablation study
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- LeCun, Y.; Bengio, Y.; Hinton, G. Deep learning. nature 2015, 521, 436–444. [Google Scholar] [CrossRef]
- Tran, D.; Wang, H.; Torresani, L.; Ray, J.; LeCun, Y.; Paluri, M. A closer look at spatiotemporal convolutions for action recognition. In Proceedings of the Proceedings of the IEEE conference on Computer Vision and Pattern Recognition; 2018; pp. 6450–6459. [Google Scholar]
- Shou, Z.; Wang, D.; Chang, S.F. Temporal action localization in untrimmed videos via multi-stage cnns. In Proceedings of the Proceedings of the IEEE conference on computer vision and pattern recognition; 2016; pp. 1049–1058. [Google Scholar]
- Xu, D.; Yan, Y.; Ricci, E.; Sebe, N. Detecting anomalous events in videos by learning deep representations of appearance and motion. Computer Vision and Image Understanding 2017, 156, 117–127. [Google Scholar] [CrossRef]
- Qiu, Z.; Yao, T.; Mei, T. Learning spatio-temporal representation with pseudo-3d residual networks. In Proceedings of the proceedings of the IEEE International Conference on Computer Vision; 2017; pp. 5533–5541. [Google Scholar]
- Gao, Y.; Liu, H.; Sun, X.; Wang, C.; Liu, Y. Violence detection using oriented violent flows. Image and vision computing 2016, 48, 37–41. [Google Scholar] [CrossRef]
- Deniz, O.; Serrano, I.; Bueno, G.; Kim, T.K. Fast violence detection in video. In Proceedings of the 2014 international conference on computer vision theory and applications (VISAPP). IEEE; 2014; Vol. 2, pp. 478–485. [Google Scholar]
- Bilinski, P.; Bremond, F. Human violence recognition and detection in surveillance videos. In Proceedings of the 2016 13th IEEE International Conference on Advanced Video and Signal Based Surveillance (AVSS). IEEE; 2016; pp. 30–36. [Google Scholar]
- Zhang, T.; Jia, W.; He, X.; Yang, J. Discriminative dictionary learning with motion weber local descriptor for violence detection. IEEE transactions on circuits and systems for video technology 2016, 27, 696–709. [Google Scholar] [CrossRef]
- Deb, T.; Arman, A.; Firoze, A. Machine cognition of violence in videos using novel outlier-resistant vlad. In Proceedings of the 2018 17th IEEE International Conference on Machine Learning and Applications (ICMLA). IEEE; 2018; pp. 989–994. [Google Scholar]
- Ji, S.; Xu, W.; Yang, M.; Yu, K. 3D convolutional neural networks for human action recognition. IEEE transactions on pattern analysis and machine intelligence 2012, 35, 221–231. [Google Scholar] [CrossRef] [PubMed]
- Simonyan, K.; Zisserman, A. Two-stream convolutional networks for action recognition in videos. Advances in neural information processing systems 2014, 27. [Google Scholar]
- Tran, D.; Bourdev, L.; Fergus, R.; Torresani, L.; Paluri, M. Learning spatiotemporal features with 3d convolutional networks. In Proceedings of the Proceedings of the IEEE international conference on computer vision; 2015; pp. 4489–4497. [Google Scholar]
- Wang, X.; Girshick, R.; Gupta, A.; He, K. Non-local neural networks. In Proceedings of the Proceedings of the IEEE conference on computer vision and pattern recognition; 2018; pp. 7794–7803. [Google Scholar]
- Dong, Z.; Qin, J.; Wang, Y. Multi-stream deep networks for person to person violence detection in videos. In Proceedings of the Pattern Recognition: 7th Chinese Conference, CCPR 2016, Chengdu, China, 5-7 November 2016; Springer: Proceedings, Part I 7; pp. 517–531. [Google Scholar]
- Zhou, P.; Ding, Q.; Luo, H.; Hou, X. Violent interaction detection in video based on deep learning. In Proceedings of the Journal of physics: conference series. IOP Publishing, Vol. 844; 2017; p. 012044. [Google Scholar]
- Serrano, I.; Deniz, O.; Espinosa-Aranda, J.L.; Bueno, G. Fight recognition in video using hough forests and 2D convolutional neural network. IEEE Transactions on Image Processing 2018, 27, 4787–4797. [Google Scholar] [CrossRef] [PubMed]
- Sudhakaran, S.; Lanz, O. Learning to detect violent videos using convolutional long short-term memory. In Proceedings of the 2017 14th IEEE international conference on advanced video and signal based surveillance (AVSS). IEEE; 2017; pp. 1–6. [Google Scholar]
- Hanson, A.; Pnvr, K.; Krishnagopal, S.; Davis, L. Bidirectional convolutional lstm for the detection of violence in videos. In Proceedings of the Proceedings of the European conference on computer vision (ECCV) workshops; 2018; pp. 0–0. [Google Scholar]
- Ullah, F.U.M.; Obaidat, M.S.; Ullah, A.; Muhammad, K.; Hijji, M.; Baik, S.W. A comprehensive review on vision-based violence detection in surveillance videos. ACM Computing Surveys 2023, 55, 1–44. [Google Scholar] [CrossRef]
- Bermejo Nievas, E.; Deniz Suarez, O.; Bueno García, G.; Sukthankar, R. Violence detection in video using computer vision techniques. In Proceedings of the Computer Analysis of Images and Patterns: 14th International Conference, CAIP 2011, Seville, Spain, 29-31 August 2011, Proceedings, Part II 14. Springer, 2011; pp. 332–339.
- Cheng, M.; Cai, K.; Li, M. RWF-2000: an open large scale video database for violence detection. In Proceedings of the 2020 25th International Conference on Pattern Recognition (ICPR). IEEE; 2021; pp. 4183–4190. [Google Scholar]
- Ulutan, O.; Rallapalli, S.; Torres, C.; Srivatsa, M.; Manjunath, B. Actor Conditioned Attention Maps for Video Action Detection. In the IEEE Winter Conference on Applications of Computer Vision (WACV), 2020.
- Zhang, C.; Zou, Y.; Chen, G.; Gan, L. Pan: Towards fast action recognition via learning persistence of appearance. arXiv preprint arXiv:2008.03462, arXiv:2008.03462 2020.
- Carreira, J.; Zisserman, A. Quo vadis, action recognition? a new model and the kinetics dataset. In Proceedings of the proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; 2017; pp. 6299–6308. [Google Scholar]
- Lee, M.; Lee, S.; Son, S.; Park, G.; Kwak, N. Motion feature network: Fixed motion filter for action recognition. In Proceedings of the Proceedings of the European Conference on Computer Vision (ECCV), 2018; pp. 387–403.
- Xie, S.; Sun, C.; Huang, J.; Tu, Z.; Murphy, K. Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification. In Proceedings of the Proceedings of the European conference on computer vision (ECCV), 2018; pp. 305–321.
- Lin, J.; Gan, C.; Han, S. Tsm: Temporal shift module for efficient video understanding. In Proceedings of the Proceedings of the IEEE/CVF international conference on computer vision, 2019; pp. 7083–7093.
- Huillcen Baca, H.A.; de Luz Palomino Valdivia, F.; Solis, I.S.; Cruz, M.A.; Caceres, J.C.G. Human Violence Recognition in Video Surveillance in Real-Time. In Proceedings of the Future of Information and Communication Conference. Springer; 2023; pp. 783–795. [Google Scholar]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Deep residual learning for image recognition. In Proceedings of the Proceedings of the IEEE conference on computer vision and pattern recognition, 2016; pp. 770–778.
- Szegedy, C.; Vanhoucke, V.; Ioffe, S.; Shlens, J.; Wojna, Z. Rethinking the inception architecture for computer vision. In Proceedings of the Proceedings of the IEEE conference on computer vision and pattern recognition, 2016; pp. 2818–2826.
- Huang, G.; Liu, Z.; Van Der Maaten, L.; Weinberger, K.Q. Densely connected convolutional networks. In Proceedings of the Proceedings of the IEEE conference on computer vision and pattern recognition, 2017; pp. 4700–4708.
- Iandola, F.N.; Han, S.; Moskewicz, M.W.; Ashraf, K.; Dally, W.J.; Keutzer, K. SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and< 0.5 MB model size. arXiv preprint arXiv:1602.07360, 2016. [Google Scholar]
- Sandler, M.; Howard, A.; Zhu, M.; Zhmoginov, A.; Chen, L.C. Mobilenetv2: Inverted residuals and linear bottlenecks. In Proceedings of the Proceedings of the IEEE conference on computer vision and pattern recognition, 2018; pp. 4510–4520.
- Howard, A.; Sandler, M.; Chu, G.; Chen, L.C.; Chen, B.; Tan, M.; Wang, W.; Zhu, Y.; Pang, R.; Vasudevan, V.; et al. Searching for mobilenetv3. In Proceedings of the Proceedings of the IEEE/CVF international conference on computer vision, 2019; pp. 1314–1324.
- Tan, M.; Le, Q. Efficientnet: Rethinking model scaling for convolutional neural networks. In Proceedings of the International conference on machine learning. PMLR; 2019; pp. 6105–6114. [Google Scholar]
- Tan, M.; Chen, B.; Pang, R.; Vasudevan, V.; Sandler, M.; Howard, A.; Le, Q.V. Mnasnet: Platform-aware neural architecture search for mobile. In Proceedings of the Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019; pp. 2820–2828.
- Tang, Y.; Han, K.; Guo, J.; Xu, C.; Xu, C.; Wang, Y. GhostNetv2: enhance cheap operation with long-range attention. Advances in Neural Information Processing Systems 2022, 35, 9969–9982. [Google Scholar]
- Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020. [Google Scholar]
- Russakovsky, O.; Deng, J.; Su, H.; Krause, J.; Satheesh, S.; Ma, S.; Huang, Z.; Karpathy, A.; Khosla, A.; Bernstein, M.; et al. Imagenet large scale visual recognition challenge. International journal of computer vision 2015, 115, 211–252. [Google Scholar] [CrossRef]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention is all you need. Advances in neural information processing systems 2017, 30. [Google Scholar]
- Singh, S.; Dewangan, S.; Krishna, G.S.; Tyagi, V.; Reddy, S.; Medi, P.R. Video vision transformers for violence detection. arXiv preprint arXiv:2209.03561, 2022. [Google Scholar]
- Li, J.; Jiang, X.; Sun, T.; Xu, K. Efficient violence detection using 3d convolutional neural networks. In Proceedings of the 2019 16th IEEE International Conference on Advanced Video and Signal Based Surveillance (AVSS). IEEE; 2019; pp. 1–8. [Google Scholar]
- Huillcen Baca, H.A.; Gutierrez Caceres, J.C.; de Luz Palomino Valdivia, F. Efficiency in human actions recognition in video surveillance using 3D CNN and DenseNet. In Proceedings of the Future of Information and Communication Conference. Springer; 2022; pp. 342–355. [Google Scholar]
- Hassner, T.; Itcher, Y.; Kliper-Gross, O. Violent flows: Real-time detection of violent crowd behavior. In Proceedings of the 2012 IEEE computer society conference on computer vision and pattern recognition workshops. IEEE; 2012; pp. 1–6. [Google Scholar]
- Mumtaz, N.; Ejaz, N.; Habib, S.; Mohsin, S.M.; Tiwari, P.; Band, S.S.; Kumar, N. An overview of violence detection techniques: current challenges and future directions. Artificial intelligence review 2023, 56, 4641–4666. [Google Scholar] [CrossRef]
- Islam, Z.; Rukonuzzaman, M.; Ahmed, R.; Kabir, M.H.; Farazi, M. Efficient two-stream network for violence detection using separable convolutional lstm. In Proceedings of the 2021 International Joint Conference on Neural Networks (IJCNN). IEEE; 2021; pp. 1–8. [Google Scholar]
- Su, Y.; Lin, G.; Zhu, J.; Wu, Q. Human interaction learning on 3d skeleton point clouds for video violence recognition. In Proceedings of the Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, 23–28 August 2020, Proceedings, Part IV 16. Springer, 2020; pp. 74–90.












| Model | Accuracy (%) | Number of parameters (M) | FLOPS (G) |
|---|---|---|---|
| Resnet50 | 76 | 25,6 | 3,8 |
| InceptionV3 | 78,8 | 23,2 | 5,0 |
| DenseNet121 | 74 | 8,0 | 2,8 |
| Squeezenet | 57,5 | 1,25 | 0,83 |
| GhostNetV2 | 75,3 | 12,3 | 0,39 |
| EffcientNetB0 | 78 | 5,3 | 1,8 |
| MobileNetV2 | 72,6 | 3,4 | 0,3 |
| MobileNetV3 L | 76,6 | 7,5 | 0,36 |
| MasNet | 75,2 | 3,9 | 0,315 |
| Vision Transformers (VIT - Huge) | 88,55 | 632 | - |
| Method | Hockey Fight dataset |
Movies dataset |
Violent Flow dataset |
|---|---|---|---|
| ViF + OViF [6] | 87.5 ± 1.7% | - | 88 ± 2.45% |
| Random Transform [7] | 90.1 ± 0% | 98.9 ± 0.22% | - |
| STIFV [8] | 93.4% | 99% | 96.4% |
| MoIWLD [9] | 96.8 ± 1.04% | - | 93.19 ± 0.12% |
| OR-VLAD [10] | 98.2 ± 0.76% | 100 ± 0% | 93.09 ± 1.14% |
| Three streams + LSTM [15] | 93.9% | - | - |
| FightNet [16] | 97.0% | 100% | - |
| Hough Forests + CNN [7] | 94.6±0.6% | 99±0.5% | - |
| ConvLSTM [18] | 97.1±0.55% | 100±0% | 94.57±2.34% |
| Bi-ConvLSTM [19] | 98.1±0.58% | 100±0% | 93.87±2.58% |
| 3D CNN end to end [43] | 98.3±0.81% | 100±0% | 97.17±0.95% |
| 3D-DenseNet(2,6,12,8) [44] | 97.0% | 100% | 90% |
| SA+TA [29] | 97.2% | 100% | - |
| Model | Accuracy(%) | #Params (M) | FLOPs(G) |
|---|---|---|---|
| C3D (Tran et al.)[13] | 82,75 | 94,8 | 40,04 |
| I3D + RGB (Carreira et al.)[25] | 85,57 | 12,3 | 55,7 |
| I3D + Two Stream (Carreira et al.)[25] | 81,75 | 24,6 | - |
| I3D + Optical Flow (Carreira et al.)[25] | 75,5 | 12,3 | - |
| ConvLSTM (Sudhakaran et al.)[18] | 77 | 94,8 | 14,4 |
| Flow Gated Network (Cheng et al.)[22] | 87,25 | 0,27 | - |
| SA+TA (Huillcen et al.)[29] | 87,75 | 5,29 | 4,17 |
| SepConvLSTM (Islam et al.) [47] | 89,75 | 0,33 | 1.93 |
| Method | Hockey Fight dataset |
Movies dataset |
Violent Flow dataset |
|---|---|---|---|
| ViF + OViF [6] | 87.5 ± 1.7% | - | 88 ± 2.45% |
| Random Transform [7] | 90.1 ± 0% | 98.9 ± 0.22% | - |
| STIFV [8] | 93.4% | 99% | 96.4% |
| MoIWLD [9] | 96.8 ± 1.04% | - | 93.19 ± 0.12% |
| OR-VLAD [10] | 98.2 ± 0.76% | 100 ± 0% | 93.09 ± 1.14% |
| Three streams + LSTM [15] | 93.9% | - | - |
| FightNet [16] | 97.0% | 100% | - |
| Hough Forests + CNN [7] | 94.6±0.6% | 99±0.5% | - |
| ConvLSTM [18] | 97.1±0.55% | 100±0% | 94.57±2.34% |
| Bi-ConvLSTM [19] | 98.1±0.58% | 100±0% | 93.87±2.58% |
| 3D CNN end to end [43] | 98.3±0.81% | 100±0% | 97.17±0.95% |
| 3D-DenseNet(2,6,12,8) [44] | 97.0% | 100% | 90% |
| SA+TA [29] | 97.2% | 100% | - |
| Proposal | 98.2% | 100% | - |
| Model | Accuracy(%) | #Params (M) | FLOPs(G) |
|---|---|---|---|
| C3D (Tran et al.)[13] | 82,75 | 94,8 | 40,04 |
| I3D + RGB (Carreira et al.)[25] | 85,57 | 12,3 | 55,7 |
| I3D + Two Stream (Carreira et al.)[25] | 81,75 | 24,6 | - |
| I3D + Optical Flow (Carreira et al.)[25] | 75,5 | 12,3 | - |
| ConvLSTM (Sudhakaran et al.)[18] | 77 | 94,8 | 14,4 |
| Flow Gated Network (Cheng et al.)[22] | 87,25 | 0,27 | - |
| SA+TA (Huillcen et al.)[29] | 87,75 | 5,29 | 4,17 |
| SepConvLSTM (Islam et al.) [47] | 89,75 | 0,33 | 1.93 |
| Proposal | 88,5 | 3,51 | 3,15 |
| Model | Accuracy(%) |
|---|---|
| SepConvLSTM (Islam et al.) [47] | 73,21 |
| Proposal | 89,29 |
| Proposal Variations | RWF-2000 Accuracy(%) | VioPeru Accuracy(%) | Parameteres (M) | FLOPS (G) |
|---|---|---|---|---|
| Proposal with EfficientNet B0 backbone | 88.25 | 87.5 | 5.29 | 4.17 |
| Proposal with MobileNet V2 backbone | 88.5 | 89.29 | 3.51 | 3.15 |
| Proposal with MobileNet V3 backbone | 88.25 | 89 | 7.62 | 4.1 |
| Proposal with MNasNet backbone | 75.25 | 62.5 | 2.22 | 1.13 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2023 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).