Document Type : Original/Review Paper

Author

University of Mazandaran

Abstract

The rapid growth of intelligent surveillance systems has increased the demand for accurate and efficient criminal activity recognition methods capable of operating in real-world environments. Although conventional deep learning and object detection frameworks have demonstrated promising performance, they often struggle to capture long-range contextual dependencies and complex interactions present in surveillance scenes. To address these limitations, this study proposes a hybrid deep learning framework that combines the real-time detection capability of YOLOv10 with the global contextual modeling power of Vision Transformers (ViT). An attention-guided feature fusion mechanism is introduced to effectively integrate local spatial representations extracted by YOLOv10 with global semantic features generated by the transformer architecture. The proposed framework is evaluated on the UCF-Crime dataset, which consists of fourteen categories of normal and criminal activities, including burglary, robbery, assault, vandalism, shoplifting, and abuse. Surveillance videos are converted into image sequences and analyzed under two experimental scenarios: (I) a standalone YOLOv10 model and (II) the proposed Attention-Guided YOLOv10-ViT framework. Performance is assessed using accuracy, precision, recall, and F1-score metrics. Experimental results show that the standalone YOLOv10 model achieves an overall classification accuracy of 88.07%, outperforming the previously reported YOLOv8 baseline. More importantly, the proposed hybrid framework attains an accuracy of 93.45%, exceeding both YOLOv10 and earlier YOLOv8-ViT architectures. The improvement is particularly evident in challenging scenarios involving occlusion, illumination changes, cluttered backgrounds, and crowded environments. The results demonstrate that integrating YOLOv10, Transformers, and attention-guided feature fusion provides a scalable, robust, and real-time solution for intelligent surveillance and public monitoring applications.

Keywords

Main Subjects

[1] W. Sultani, C. Chen, and M. Shah, "Real-world anomaly detection in surveillance videos, " in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), Salt Lake City, UT, USA, pp. 6479–6488, 2018, doi: 10.1109/CVPR.2018.00678.
 
[2] H.-T. Duong, V.-T. Le, and V. T. Hoang, "Deep learning-based anomaly detection in video surveillance: A survey, " Sensors, vol. 23, no. 11, p. 5024, 2023, doi: 10.3390/s23115024.
 
[3] Y. Zhang, X. Li, and H. Wang, "MTFL: Multi-timescale feature learning for weakly-supervised anomaly detection in surveillance videos, " arXiv preprint arXiv:2410.05900, 2024, doi: 10.48550/arXiv.2410.05900.
 
[4] E. Dilek and M. Dener, "An overview of transformers for video anomaly detection, " Neural Computing and Applications, vol. 37, pp. 17825–17857, 2025, doi: 10.1007/s00521-025-11218-1.
 
[5] K. Boekhoudt, A. Matei, M. Aghaei, and E. Talavera, "HR-Crime: Human-related anomaly detection in surveillance videos, " in Proc. 19th Int. Conf. Comput. Anal. Images Patterns (CAIP), LNCS. Cham, Switzerland: Springer, pp. 164–174, 2021, doi: 10.1007/978-3-030-89131-2_15.
 
[6] U. V. Navalgund and P. K., "Crime intention detection system using deep learning, " in Proc. Int. Conf. Circuits Syst. Digit. Enterp. Technol. (ICCSDET), Kottayam, India, pp. 1–6, 2018, doi: 10.1109/ICCSDET.2018.8821168.
 
[7] V. Mandalapu, L. Elluri, P. Vyas, and N. Roy, "Crime prediction using machine learning and deep learning: A systematic review and future directions, " IEEE Access, vol. 11, pp. 60153–60170, 2023, doi: 10.1109/ACCESS.2023.3286344.
 
[8] K. Saifullah, M. M. Alam, P. M. Joy, J. I. Hasan, S. Islam, and N. Hossain, "Deep learning based crime detection and resource creation approach from Bengali voice calls, " Data Science, 2023.
 
[9] S. M. Divya, G. S. Priya, R. Abitha, K. Sirisha, A. Manikanta, and K. Jayanth, "Automated crime intention detection using deep learning, " Int. Res. J. Modernization Eng. Technol. Sci., vol. 4, no. 6, 2022.
 
[10] K. Shoeb and Y. R. Devi, "Real time crime detection using deep learning, " Int. Res. J. Eng. Technol. (IRJET), vol. 10, no. 12, 2023.
 
[11] P. Sivakumar, J. V., and R. R. K. S., "Real time crime detection using deep learning algorithm, " in Proc. Int. Conf. Syst. Comput. Autom. Netw. (ICSCAN), Puducherry, India, pp. 1–5, 2021, doi: 10.1109/ICSCAN53069.2021.9526393.
 
[12] K. H. W. and K. H. B., "Prediction of crime occurrence from multi-modal data using deep learning, " PLoS One, vol. 12, no. 4.
 
[13] M. Mukto, M. Hasan, M. Al Mahmud, I. Haque, A. Ahmed, T. Jabid, S. Md. Ali, M. R. A. Rashid, M. M. Islam, and M. Islam, "Design of a real-time crime monitoring system using deep learning techniques, " Intelligent Systems with Applications, vol. 21, p. 200311, 2024, doi: 10.1016/j.iswa.2023.200311.
 
[14] A. O. Hashi, A. A. Abdirahman, M. A. Elmi, and O. E. R. Rodriguez, "Deep learning models for crime intention detection using object detection, " Int. J. Adv. Comput. Sci. Appl. (IJACSA), vol. 14, no. 4, 2023, doi: 10.14569/IJACSA.2023.0140434.
 
[15] S. Jebur, K. Hussein, and H. Hoomod, "Abnormal behavior detection in video surveillance using Inception v3 transfer learning approaches, " Iraqi J. Comput. Commun. Control Syst. Eng., vol. 23, no. 2, pp. 210–221, 2023, doi: 10.33103/uot.ijccce.23.2.16.
 
[16] Z. Dorrani, "Anomaly detection in emerging crimes with deep autoencoder architecture, " Contrib. Sci. Technol. Eng., vol. 2, no. 3, pp. 45–56, 2025, doi: 10.22080/cste.2025.28900.1023.
 
[17] Y. Qian et al., "UCF-Crime-DVS: A novel event-based dataset for video anomaly detection with spiking neural networks, " in Proc. AAAI Conf. Artif. Intell., vol. 39, pp. 6577–6585, 2025.
 
[18] R. Mageed and H. Hatem, "An efficient deep learning based approaches for crime activities classification in surveillance videos, " J. Inf. Syst. Eng. Manag., vol. 10, no. 26s, 2025, doi: 10.52783/jisem.v10i26s.4246.
 
[19] C. Maradana, M. An, and A. Rasheed, "Human activity recognition and abnormality detection using deep learning, " in Proc. 6th Int. Conf. Pattern Recognit. Intell. Syst. (PRIS 2024), pp. 87–92, 2024, doi: 10.1145/3689218.3689231.
 
[20] A.-D. Matei, E. Talavera, and M. Aghaei, "Crime scene classification from skeletal trajectory analysis in surveillance settings, " Engineering Applications of Artificial Intelligence, vol. 141, p. 109800, 2025, doi: 10.1016/j.engappai.2024.109800.
 
[21] Kaggle, "UCF Crime Dataset, " [Online]. Available: https://www.kaggle.com/datasets/odins0n/ucf-crime-dataset. [Accessed: Apr. 29, 2026].
 
[22] N. Cauli and D. Reforgiato Recupero, "Survey on video data augmentation techniques for deep learning models, " Future Internet, vol. 14, no. 3, p. 93, 2022, doi: 10.3390/fi14030093.
 
[23] S. Yun, S. J. Oh, B. Heo, D. Han, J. Kim, and J. Kim, "VideoMix: Rethinking data augmentation for video classification, " arXiv preprint arXiv:2012.03457, 2020, doi: 10.48550/arXiv.2012.03457.
 
[24] S. Mavaddati and M. Razavi, "A CNN-LSTM-based approach for classification and quality detection of rice varieties, " Journal of AI and Data Mining, vol. 12, no. 4, pp. 473–485, 2024, doi: 10.22044/jadm.2024.15282.2631.
 
[25] S. Mavaddati, "A hybrid approach for brain tumor classification: Enhancing MRI-based diagnosis with CNN-transformer synergy,” Journal of AI and Data Mining, vol. 14, no. 1, pp. 37–49, 2026, doi: 10.22044/jadm.2025.16554.2779.
 
[26] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, "An image is worth 16×16 words: Transformers for image recognition at scale, " in Proc. Int. Conf. Learn. Represent. (ICLR), 2021, doi: 10.48550/arXiv.2010.11929.
 
[27] Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, "Swin Transformer: Hierarchical vision transformer using shifted windows, " in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), pp. 10012–10022, 2021, doi: 10.1109/ICCV48922.2021.00986.
 
[28] E. Moradi, "A Novel Fault Prediction Technique for Oil-Immersed Transformers Based on Advanced Gradient Boosting and Particle Swarm Optimization (PSO), " Journal of AI and Data Mining, vol. 14, no. 1, pp. 25–35, 2026, doi: 10.22044/jadm.2025.16214.2745.