Submitted:
29 October 2024
Posted:
30 October 2024
You are already at the latest version
Abstract
Keywords:
1. Introduction
1.1. Related Work
1.2. Motivation
1.3. Our Work
- Introduction of Lightweight Convolution (GhostConv): We first implement lightweight convolution within the YOLOv5 pedestrian detection algorithm, significantly reducing the model's computational complexity and the number of parameters. This approach maintains excellent feature extraction capabilities while enhancing the efficiency of real-time detection.
- Design of an Improved C2f Module: By constructing an enhanced C2f module, we optimize the feature fusion process to address the low accuracy in pedestrian detection resulting from variations in environmental scales. This improvement enhances the network's adaptability across diverse scenes, thereby significantly boosting pedestrian detection performance.
- Incorporation of Coordinate Attention: We introduce a coordinate attention mechanism to enhance the model's ability to capture key target locations. This innovation facilitates more accurate localization of pedestrians in complex backgrounds, reducing the incidence of missed and false detections while improving overall detection accuracy.
- Design of the WiseIoU Loss Function: The WiseIoU Loss function is employed in bounding box calculations, integrating the overlapping region, centroid distance, and aspect ratio. This design enhances the model's adaptability to various detection scenarios, ensuring robust performance in dynamic environments.
- Experimental Validation and Performance Enhancement: Rigorous experiments conducted on public datasets reveal that the detection accuracy of the CCW-YOLO algorithm reaches 95.6%, which is an improvement of 8.7% over the original YOLOv5s algorithm. This advancement significantly enhances the detection of small objects in images and effectively addresses the issues of misdetection and omission in complex scenes.
2. The Proposed Approach
2.1. Improved Backbone Network
2.1.1. GhostConv Structure
2.1.2. C2f Module
2.2. Coordinate Attention Mechanism
2.3. Improved WiseIoU Loss
3. Experimental Results
3.1. Dataset Description
3.2. Evaluation Metrics
- TP (True Positives): These are true cases, meaning positive instances that are accurately identified as such by the model.
- FP (False Positives): These refer to pseudo-positive cases, which are instances that the model incorrectly identifies as positive, although they are actually negative.
- FN (False Negatives): These are pseudo-negative cases, meaning instances that are actually positive but have been incorrectly identified as negative by the model.
- TN (True Negatives): These are true negative cases, referring to instances that are accurately identified as negative by the model.
3.3. Comparison Experiment
3.4. Visual Assessment
4. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Kesting, Arne, Martin Treiber, and Dirk Helbing. Enhanced intelligent driver model to access the impact of driving strategies on traffic capacity. Philosophical Transactions of the Royal Society A: Mathematical. Physical and Engineering Sciences 2010, 368, 4585–4605. [Google Scholar]
- Xia, Qin, et al. Test scenario design for intelligent driving system ensuring coverage and effectiveness. International Journal of Automotive Technology 2018, 19, 751–758. [Google Scholar] [CrossRef]
- Gui, Guan, et al. Machine learning aided air traffic flow analysis based on aviation big data. IEEE Transactions on Vehicular Technology 2020, 69, 4817–4826. [Google Scholar] [CrossRef]
- Cui, Xinyi, et al. 3d haar-like features for pedestrian detection. 2007 IEEE International Conference on Multimedia and Expo. IEEE, 2007.
- Wei, Yun, Qing Tian, and Teng Guo. An improved pedestrian detection algorithm integrating haar-like features and hog descriptors. Advances in Mechanical Engineering 2013, 5, 546206. [Google Scholar] [CrossRef]
- Zhou, Hongzhi, and Gan Yu. Research on pedestrian detection technology based on the SVM classifier trained by HOG and LTP features. Future Generation Computer Systems 2021, 125, 604–615. [Google Scholar] [CrossRef]
- Cai, Yingfeng, et al. Research on pedestrian detection technology based on improved DPM model. 2017 IEEE 7th Annual International Conference on CYBER Technology in Automation, Control, and Intelligent Systems (CYBER). IEEE, 2017.
- Khemmar, Redouane, et al. Real time pedestrian detection-based faster hog/dpm and deep learning approach. SITIS-International Conference on Signal Image Technology & Internet Based Systems. 2019.
- Szarvas, Mate, et al. Pedestrian detection with convolutional neural networks. IEEE Proceedings. Intelligent Vehicles Symposium, 2005.. IEEE, 2005.
- Masita, Katleho L., Ali N. Hasan, and Satyakama Paul. Pedestrian detection using R-CNN object detector. 2018 IEEE Latin American Conference on Computational Intelligence (LA-CCI). IEEE, 2018.
- Zhang, Shifeng, et al. Occlusion-aware R-CNN: Detecting pedestrians in a crowd. Proceedings of the European conference on computer vision (ECCV). 2018.
- Hung, Goon Li, et al. Faster R-CNN deep learning model for pedestrian detection from drone images. SN Computer Science 2020, 1, 1–9. [Google Scholar]
- Hsu, Wei-Yen, and Wen-Yen Lin. Ratio-and-scale-aware YOLO for pedestrian detection. IEEE transactions on image processing 2020, 30, 934–947. [Google Scholar]
- Fan, Di, et al. Improved ssd-based multi-scale pedestrian detection algorithm. Advances in 3D Image and Graphics Representation, Analysis, Computing and Information Technology: Algorithms and Applications, Proceedings of IC3DIT 2019, Volume 2. Springer Singapore, 2020.
- Liu, Congqiang, Haosen Wang, and Chunjian Liu. Double Mask R-CNN for Pedestrian Detection in a Crowd. Mobile Information Systems 2022, 2022, 4012252. [Google Scholar]
- Zhang, Shanshan, Christian Bauckhage, and Armin B. Cremers. Informed haar-like features improve pedestrian detection. Proceedings of the IEEE conference on computer vision and pattern recognition. 2014.
- Gawande, Ujwalla, Kamal Hajari, and Yogesh Golhar. Pedestrian detection and tracking in video surveillance system: issues, comprehensive review, and challenges. Recent Trends in Computational Intelligence 2020, 1–24.
- Mao, Xiao-Jiao, et al. Enhanced deformable part model for pedestrian detection via joint state inference. 2015 IEEE International Conference on Image Processing (ICIP). IEEE, 2015.
- Cao, Jiale, et al. From handcrafted to deep features for pedestrian detection: A survey. IEEE transactions on pattern analysis and machine intelligence 2021, 44, 4913–4934. [Google Scholar]
- Sha, Mingzhi, and Azzedine Boukerche. Performance evaluation of CNN-based pedestrian detectors for autonomous vehicles. Ad Hoc Networks 2022, 128, 102784. [Google Scholar] [CrossRef]
- Xie, Jin, et al. Count-and similarity-aware R-CNN for pedestrian detection. Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XVII 16. Springer International Publishing, 2020.
- Maity, Madhusri, Sriparna Banerjee, and Sheli Sinha Chaudhuri. Faster r-cnn and yolo based vehicle detection: A survey. 2021 5th international conference on computing methodologies and communication (ICCMC). IEEE, 2021.
- Murthy, Chintakindi Balaram, Mohammad Farukh Hashmi, and Avinash G. Keskar. Optimized MobileNet+ SSD: a real-time pedestrian detection on a low-end edge device. International Journal of Multimedia Information Retrieval 2021, 10, 171–184. [Google Scholar] [CrossRef]
- Huang, Lincai, Zhiwen Wang, and Xiaobiao Fu. Pedestrian detection using RetinaNet with multi-branch structure and double pooling attention mechanism. Multimedia Tools and Applications 2024, 83, 6051–6075. [Google Scholar] [CrossRef]
- Shao, Yifan, et al. Aero-YOLO: An Efficient Vehicle and Pedestrian Detection Algorithm Based on Unmanned Aerial Imagery. Electronics 2024, 13, 1190. [Google Scholar] [CrossRef]
- Gao, Fei, et al. Improved YOLOX for pedestrian detection in crowded scenes. Journal of Real-Time Image Processing 2023, 20, 24. [Google Scholar] [CrossRef]
- Zhao, Siqi, et al. Improved YOLOv5 Algorithm for Intensive Pedestrian Detection. International Conference on Computational & Experimental Engineering and Sciences. Cham: Springer Nature Switzerland, 2023.
- Cao, Jinshan, et al. GCL-YOLO: A GhostConv-based lightweight yolo network for UAV small object detection. Remote Sensing 2023, 15, 4932. [Google Scholar] [CrossRef]
- Pan, Jingmin, et al. C2F-YOLO: A Coarse-to-Fine Object Detection Framework Based on YOLO. Proceedings of the 2024 3rd Asia Conference on Algorithms, Computing and Machine Learning. 2024.
- Xie, Chao, Hongyu Zhu, and Yeqi Fei. Deep coordinate attention network for single image super-resolution. IET Image Processing 2022, 16, 273–284. [Google Scholar] [CrossRef]
- Tong, Zanjia, et al. Wise-IoU: bounding box regression loss with dynamic focusing mechanism. arXiv preprint arXiv:2301 (2023).
- Taiana, Matteo, Jacinto C. Nascimento, and Alexandre Bernardino. An improved labelling for the INRIA person data set for pedestrian detection. Pattern Recognition and Image Analysis: 6th Iberian Conference, IbPRIA 2013, Funchal, Madeira, Portugal, June 5-7, 2013. Proceedings 6. Springer Berlin Heidelberg, 2013.







| Network Models | Precision(%) | Recall(%) | mAP_0.5(%) | mAP_0.5:0.95(%) |
|---|---|---|---|---|
| YOLOv5s | 88.3 | 75.8 | 86.9 | 56.8 |
| YOLOv5s+C2f | 91.5 | 75.4 | 90.2 | 58.7 |
| Network Models | Precision(%) | Recall(%) | mAP_0.5(%) | mAP_0.5:0.95(%) |
|---|---|---|---|---|
| YOLOv5s | 88.3 | 75.8 | 86.2 | 56.8 |
| YOLOv5s+SE | 89.7 | 80.2 | 90.2 | 56.4 |
| YOLOv5s+CA | 92.3 | 81.4 | 92.4 | 62.3 |
| Network Models | Precision(%) | Recall(%) | mAP_0.5(%) | mAP_0.5:0.95(%) |
|---|---|---|---|---|
| YOLOv5s+CIoU | 88.3 | 75.8 | 86.9 | 56.8 |
| YOLOv5s+EIoU | 88.6 | 78.5 | 88.4 | 55.9 |
| YOLOv5s+WiseIoU | 89.8 | 80.3 | 89.6 | 58.5 |
| YOLOv5s | C2f | CA | WiseIoU | mAP_0.5(%) | mAP_0.5:0.95(%) |
|---|---|---|---|---|---|
| √ | 86.9 | 56.8 | |||
| √ | √ | 90.2 | 58.7 | ||
| √ | √ | 92.4 | 62.3 | ||
| √ | √ | 89.6 | 58.5 | ||
| √ | √ | √ | 93.1 | 63.6 | |
| √ | √ | √ | 90.1 | 58.9 | |
| √ | √ | √ | 92.6 | 62.7 | |
| √ | √ | √ | √ | 95.6 | 65.9 |
| Models | Precision(%) | Recall(%) | mAP_0.5(%) | mAP_0.5:0.95(%) |
|---|---|---|---|---|
| Faster R-CNN | 90.5 | 76.6 | 84.9 | 57.7 |
| YOLOv5s | 88.3 | 75.8 | 86.9 | 56.8 |
| YOLOv7 | 91.5 | 75.4 | 90.2 | 58.7 |
| YOLOv8 | 92.3 | 81.4 | 92.4 | 62.3 |
| CCW-YOLO | 95.6 | 83.5 | 94.2 | 65.9 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2024 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).