Submitted:
15 May 2024
Posted:
16 May 2024
You are already at the latest version
Abstract
Keywords:
1. Introduction
1.1. Contributions of the Study
- We have developed a series of YOLO models characterized by their exceptional precision scores in the detection of traffic officers. These models are poised for integration into AVs, thereby augmenting the capabilities of these vehicles to effectively navigate real-world scenarios.
- To evaluate the performance and behavior of these models, we have devised a comprehensive 6-phase methodology. This approach aids us in identifying the most suitable model for deployment within AVs, ensuring optimal functionality and safety in practical applications.
2. Related Work
3. Methodology
3.1. Model Selection
3.1.1. YOLO v3
3.1.2. YOLOv5
- Backbone: This component is equipped with pre-trained networks designed to extract essential features from input images. In the case of YOLOv5, the chosen backbone is the CSP-Darknet53. This configuration involves convolutional layers comprised of both residual and dense blocks, strategically engineered to enhance the flow of information within the network and alleviate the issue of vanishing gradients.
- Neck: The neck component plays a crucial role in feature extraction and pyramidal scaling to effectively handle objects of varying sizes and scales. YOLOv5 employs the Path Aggregation Network (PANet) within the neck, which optimizes information flow and aids in precise pixel localization, particularly when engaged in mask prediction tasks. Furthermore, the SPP component within the neck enhances feature aggregation, ensuring a consistent output length without sacrificing information throughput.
- Head/Prediction: A similar dynamic to that of YOLOv4 and YOLOv3 is maintained in the main component of YOLOv5. This entails the utilization of three prediction layers that play a pivotal role in determining bounding boxes and object identification. These prediction layers are instrumental in detecting and characterizing objects within the input data.

3.1.3. YOLOv8
3.2. Proposed Methodology
3.2.1. Phase 1: Dataset Collection
3.2.2. Phase 2: Data Preparation
3.2.3. Phase 3: Augmentation and Filter Application
3.2.4. Phase 4: Model Selection and Training
3.2.5. Quality Measurements
- Precision: Refers to the spread of values obtained from magnitude measurements. Precision is inversely proportional to dispersion, meaning that if precision is high, dispersion is minimal.
- Recall: Also known as the true positive rate. It represents the quantity of positives identified correctly.
- F1: This metric represents a summary of both precision and recall in a single metric.
3.2.6. Phase 5: Evaluation and Results





3.2.7. Phase 6: Behavior Analysis
4. Results and Discussion
5. Conclusion
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| AI | Artificial Intelligence |
| AV | Autonomous Vehicle |
| CCTV | Closed-Circuit Television |
| FPS | Frames Per Second |
| GCN | Graph-Based Networks |
| GPU | Graphics Processing Unit |
| LiDAR | Light Detection and Ranging |
| mAP | mean Average Precision |
| MOTA | Multiple Object Tracking Accuracy |
| OCD | Osteochondritis Dissecans |
| PANet | Path Aggregation Network |
| RNN | Recurrent Neural Network |
| SPP | Spatial Pyramid Pooling |
| SPPF | Spatial Pixel Pair Features |
| TCN | Temporal Convolutional Network |
| UAV | Unmanned Aerial Vehicle |
| YOLO | You Only Look Once |
References
- Yeong, D.J.; Velasco-Hernandez, G.; Barry, J.; Walsh, J. Sensor and Sensor Fusion Technology in Autonomous Vehicles: A Review. Sensors 2021, 21. [Google Scholar] [CrossRef] [PubMed]
- Vargas, J.; Alsweiss, S.; Toker, O.; Razdan, R.; Santos, J. An Overview of Autonomous Vehicles Sensors and Their Vulnerability to Weather Conditions. Sensors 2021, 21. [Google Scholar] [CrossRef] [PubMed]
- Peng, L.; Wang, H.; Li, J. Uncertainty Evaluation of Object Detection Algorithms for Autonomous Vehicles. Automotive Innovation 2021, 4, 241–252. [Google Scholar] [CrossRef]
- Vargas, J.; Alsweiss, S.; Toker, O.; Razdan, R.; Santos, J. An Overview of Autonomous Vehicles Sensors and Their Vulnerability to Weather Conditions. Sensors 2021, 21. [Google Scholar] [CrossRef] [PubMed]
- Parekh, D.; Poddar, N.; Rajpurkar, A.; Chahal, M.; Kumar, N.; Joshi, G.P.; Cho, W. A Review on Autonomous Vehicles: Progress, Methods and Challenges. Electronics 2022, 11. [Google Scholar] [CrossRef]
- Gupta, A.; Anpalagan, A.; Guan, L.; Khwaja, A.S. Deep learning for object detection and scene perception in self-driving cars: Survey, challenges, and open issues. Array 2021, 10, 100057. [Google Scholar] [CrossRef]
- Stockem Novo, A.; Hürten, C.; Baumann, R.; Sieberg, P. Self-evaluation of automated vehicles based on physics, state-of-the-art motion prediction and user experience. Scientific Reports 2023, 13, 12692. [Google Scholar] [CrossRef] [PubMed]
- Idrovo-Berrezueta, P.; Dutan-Sanchez, D.; Hurtado-Ortiz, R.; Robles-Bykbaev, V. Data Analysis Architecture using Techniques of Machine Learning for the Prediction of the Quality of Blood Fonations against the Hepatitis C Virus. 2022 IEEE International Autumn Meeting on Power, Electronics and Computing (ROPEC); IEEE: Ixtapa, Mexico, 2022; pp. 1–7. [Google Scholar] [CrossRef]
- Idrovo-Berrezueta, P.; Dutan-Sanchez, D.; Robles-Bykbaev, V. Comparison of Transfer Learning vs. Hyperparameter Tuning to Improve Neural Networks Precision in the Early Detection of Pneumonia in Chest X-Rays. In Information Technology and Systems; Rocha, A., Ferras, C., Ibarra, W., Eds.; Springer International Publishing: Cham, 2023; Vol. 691, pp. 263–272.Series Title: Lecture Notes in Networks and Systems. [Google Scholar] [CrossRef]
- He, J.; Zhang, C.; He, X.; Dong, R. Visual Recognition of traffic police gestures with convolutional pose machine and handcrafted features. Neurocomputing 2020, 390, 248–259. [Google Scholar] [CrossRef]
- Mishra, A.; Kim, J.; Cha, J.; Kim, D.; Kim, S. Authorized Traffic Controller Hand Gesture Recognition for Situation-Aware Autonomous Driving. Sensors 2021, 21. [Google Scholar] [CrossRef] [PubMed]
- Wiederer, J.; Bouazizi, A.; Kressel, U.; Belagiannis, V. Traffic Control Gesture Recognition for Autonomous Vehicles. 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2020, pp. 10676–10683. [CrossRef]
- Self-driving car stopped by San Francisco police, 2022.
- Public perceptions of autonomous vehicle safety: An international comparison, 2020. Publisher: Elsevier. [CrossRef]
- Sharma, N.; Baral, S.; Paing, M.P.; Chawuthai, R. Parking Time Violation Tracking Using YOLOv8 and Tracking Algorithms. Sensors 2023, 23, 5843. [Google Scholar] [CrossRef] [PubMed]
- Roboflow. Everything you need to build and deploy computer vision models. https://roboflow.com/. [Accessed 25-08-2023].
- Yasamorn, A.; Wongcharoen, A.; Joochim, C. Object Detection of Pedestrian Crossing Accident Using Deep Convolutional Neural Networks. 2022 Research, Invention, and Innovation Congress: Innovative Electricals and Electronics (RI2C); IEEE: Bangkok, Thailand, 2022; pp. 297–303. [Google Scholar] [CrossRef]
- Menon, A.; Omman, B.; S, A. Pedestrian Counting Using Yolo V3. 2021 International Conference on Innovative Trends in Information Technology (ICITIIT); IEEE: Kottayam, India, 2021; pp. 1–9. [Google Scholar] [CrossRef]
- Wei, C.; Tan, Z.; Qing, Q.; Zeng, R.; Wen, G. Fast Helmet and License Plate Detection Based on Lightweight YOLOv5. Sensors 2023, 23, 4335. [Google Scholar] [CrossRef] [PubMed]
- Avupati, S.L.; Harshitha, A.; Jeedigunta, S.P.; Sai Chikitha Chowdary, D.; Pushpa, B. Traffic Rules Violation Detection using YOLO and HAAR Cascade. 2023 9th International Conference on Advanced Computing and Communication Systems (ICACCS); IEEE: Coimbatore, India, 2023; pp. 1159–1163. [Google Scholar] [CrossRef]
- Wang, J.; Song, Q.; Hou, M.; Jin, G. Infrared Image Object Detection of Vehicle and Person Based on Improved YOLOv5. In Web and Big Data. APWeb-WAIM 2022 International Workshops; Yang, S.; Islam, S., Eds.; Springer Nature Singapore: Singapore, 2023. Vol. 1784, pp.175–187. Series Title: Communications in Computerand Information Science. [CrossRef]
- Nepal, U.; Eslamiat, H. Comparing YOLOv3, YOLOv4 and YOLOv5 for Autonomous Landing Spot Detection in Faulty UAVs. Sensors 2022, 22, 464. [Google Scholar] [CrossRef] [PubMed]
- Inui, A.; Mifune, Y.; Nishimoto, H.; Mukohara, S.; Fukuda, S.; Kato, T.; Furukawa, T.; Tanaka, S.; Kusunose, M.; Takigami, S.; Ehara, Y.; Kuroda, R. Detection of Elbow OCD in the Ultrasound Image by Artificial Intelligence Using YOLOv8. Applied Sciences 2023, 13, 7623. [Google Scholar] [CrossRef]
- GitHub - ultralytics/ultralytics: NEW - YOLOv8 in PyTorch > ONNX > OpenVINO > CoreML > TFLite — github.com. https://github.com/ultralytics/ultralytics. [Accessed 25-09-2023].
- Mota-Delfin, C.; López-Canteñs, G.D.J.; López-Cruz, I.L.; Romantchik-Kriuchkova, E.; Olguín-Rojas, J.C. Detection and Counting of Corn Plants in the Presence of Weeds with Convolutional Neural Networks. Remote Sensing 2022, 14, 4892. [Google Scholar] [CrossRef]




| YOLO version | Ref. | Graphics card NVIDIA | Dataset detection objective | Precision | Recall | mAP 0.50 |
|---|---|---|---|---|---|---|
| YOLO v3 | [17] | T4 (16GB) | Pedestrian Crossing | 0.979 | - | 0.559 |
| [18] | GetForce 920MX | Pedestrian | 0.954 | 0.93 | 0.933 | |
| YOLO v5 | [19] | GeForce RTX 3080 Ti | Electric bikes, helmets, and license plates | - | - | 0.870 |
| [20] | - | Traffic violations | - | - | 0.995 | |
| [21] | Xavier NX | Vehicles and pedestrians | - | - | 0.884 | |
| [22] | GeForce RTX 2070 SUPER | Landing spot | 0.707 | 0.611 | 0.633 | |
| YOLO v8 | [23] | GeForce RTX 3050 | Elbow osteochondritis dissecans | 0.991 | 0.9975 | 0.787 |
| [15] | Quadro P4000 | Parking time violations | - | - | 0.539 |
| Images | Train | Validate | Test |
|---|---|---|---|
| 1862 | 1734 | 81 | 47 |
| Version | Parameters |
|---|---|
| YOLO v3 | learning_rate = 0.01, momentum = 0.937, weight_decay = 0.0005, warmup_epochs = 3.0 |
| YOLO v5 | learning_rate = 0.01, momentum = 0.937, weight_decay = 0.0005, warmup_epochs = 3.0 |
| YOLO v8 | learning_rate = 0.01, momentum = 0.937, weight_decay = 0.0005, warmup_epochs = 3.0 |
| GPU | CPU | Memory |
|---|---|---|
| Nvidia A100 SXM4 120 GB | 16 | 32 GB |
| Model | Version | F1-Score | Confidence Score | Training Time (Hours) | Precision | Recall | mAP 0.50 | mAP 0.50-0.95 |
|---|---|---|---|---|---|---|---|---|
| YOLO v3 | Tiny | 0.85 | 0.320 | 0.380 | 0.907 | 0.866 | 0.906 | 0.464 |
| Small | 0.91 | 0.237 | 0.601 | 0.971 | 0.951 | 0.968 | 0.654 | |
| Medium | 0.93 | 0.349 | 0.605 | 0.964 | 0.950 | 0.961 | 0.629 | |
| YOLO v5 | Nano | 0.95 | 0.248 | 0.381 | 0.989 | 0.959 | 0.971 | 0.691 |
| Small | 0.95 | 0.103 | 0.383 | 0.991 | 0.982 | 0.985 | 0.734 | |
| Medium | 0.95 | 0.179 | 0.411 | 0.983 | 0.959 | 0.985 | 0.750 | |
| Large | 0.95 | 0.356 | 0.810 | 0.982 | 0.967 | 0.983 | 0.757 | |
| X-Large | 0.96 | 0.577 | 0.823 | 0.990 | 0.971 | 0.986 | 0.755 | |
| YOLO v8 | Nano | 0.93 | 0.332 | 0.265 | 1.0 | 1.0 | 0.979 | 0.758 |
| Small | 0.95 | 0.402 | 0.211 | 0.982 | 0.975 | 0.978 | 0.743 | |
| Medium | 0.96 | 0.402 | 0.345 | 0.991 | 0.974 | 0.986 | 0.780 | |
| Large | 0.94 | 0.442 | 0.510 | 0.974 | 0.965 | 0.977 | 0.782 | |
| X-Large | 0.95 | 0.631 | 0.533 | 0.981 | 0.967 | 0.977 | 0.782 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2024 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).