Submitted:
06 August 2024
Posted:
07 August 2024
You are already at the latest version
Abstract

Keywords:
1. Introduction
1.1. Pothole Detection Importance for Autonomous Vehicles
1.2. Human Response to Potholes
1.3. Pothole Detection Methods
1.4. Challenges in Potholes Detection
1.5. Proposed Method: Distant Pothole Detection with Vision and Perspective Transformation
2. Related Work
2.1. Vision Approaches
2.2. Perspective Transformation Technique
3. Methodology
3.1. Framework Overview
3.1.1. Perspective Transformation Motivation
3.1.2. YOLOv5: Key Features and Functionality
3.2. Automated Perspective Transformation Algorithm
| Algorithm 1:Automatic Perspective Transformation for Images and Bounding Boxes |
![]() |
- Initialize lists: Store coordinates and dimensions of bounding boxes for all images, including minimum and maximum x and y coordinates, width, and height for each bounding box.
- Read bounding boxes: Extract bounding box data from each image’s corresponding label file, calculate all ROI boundary points, and update the respective lists.
- Calculate offsets: Determine the ROI offsets using a specific value and the max width and height of all bounding boxes to define a slightly larger ROI.
- Determine ROI corners: Use the minimum and maximum coordinates from the lists, along with the calculated offsets, to determine the corners of the ROI. These corners are the source points () for the perspective transformation.
- Clip ROI corners: Ensure ROI corners stay within image boundaries.
- Define target points: Set target points () based on the image dimensions, representing the transformed image corners.
- Calculate transformation matrix: Compute the perspective transformation matrix M using the source and target points. This matrix is used to transform the coordinates of the ROI to the new perspective.
- Transform images and bounding boxes: Apply the transformation matrix M to each image and its bounding boxes. This involves transforming the image and adjusting the bounding box coordinates accordingly. The transformed images and bounding boxes are then saved.
4. Experiment Design
4.1. Evaluation Dataset
4.2. Evaluation Metrics
4.3. Evaluation Strategy
- Far:
- Medium:
- Near:
- Far:
- Medium:
- Near:
- Far:
- Medium:
- Near:
4.4. Implementation Settings
5. Results and Discussion
5.1. Experiment 1: Naive vs. Fixed Cropping vs. Automated Transformation Approach
5.2. Experiment 2: Effects of Network Complexity/Scale on Performance
5.3. Experiment 3: Ablation Study
- YOLOv5’s Augmentations Only, No Negative Images: This setup utilized only YOLOv5’s augmentation step without negative images, as illustrated in the first row of Table 3.
- YOLOv5’s Augmentations with Negative Images: This setup included negative images alongside YOLOv5’s augmentation step, shown in the second row of the table.
- Manual Preprocessing and YOLOv5’s Augmentation, Positive Images Only: This configuration combined manual preprocessing augmentations with YOLOv5’s augmentations, using only positive images. It achieved the best results among all setups.
- Manual Preprocessing and YOLOv5’s Augmentation with Negative Images: This setup used both manual and YOLOv5’s augmentations, incorporating negative images into the positive dataset. It resulted in the lowest performance metrics.
6. Conclusion
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Potholes could be ‘self-repairing’ in the next 30 years, say experts — independent.co.uk. Available online: https://www.independent.co.uk/travel/news-and-advice/potholes-fix-repair-self-repairing-roads-transport-infrastructure-maintenance-a8719716.html (accessed on 01 08 2024).
- Visual Expert. Reaction Time. 2024. Available online: https://www.visualexpert.com/Resources/reactiontime.html (accessed on 01 08 2024).
- Remodel or Move. Should You Hit the Brakes When Going Over a Pothole? 2023. Available online: https://www.remodelormove.com/should-you-hit-the-brakes-when-going-over-a-pothole/ (accessed on 01 08 2024).
- Balakuntala, S.; Venkatesh, S. An intelligent system to detect, avoid and maintain potholes: A graph theoretic approach. arXiv, arXiv:1305.5522 2013.
- University of Minnesota Twin Cities. Talking Potholes. 2023. Available online: https://twin-cities.umn.edu/news-events/talking-potholes-u-m (accessed on 01 08 2024).
- Kim, Y.M.; Kim, Y.G.; Son, S.Y.; Lim, S.Y.; Choi, B.Y.; Choi, D.H. Review of recent automated pothole-detection methods. Applied Sciences 2022, 12, 5320. [Google Scholar] [CrossRef]
- Eduzaurus. Pothole Detection Methods. 2023. Available online: https://eduzaurus.com/free-essay-samples/pothole-detection-methods/ (accessed on 01 08 2024).
- Geoawesomeness. Application of Mobile LiDAR on Pothole Detection. 2013. Available online: https://geoawesomeness.com/eo-hub/application-of-mobile-lidar-on-pothole-detection/ (accessed on 01 08 2024).
- Samczynski, P.; Giusti, E. Recent Advancements in Radar Imaging and Sensing Technology; MDPI, 2021.
- Outsight. How Does LiDAR Compare to Cameras and Radars? 2023. Available online: https://www.outsight.ai/insights/how-does-lidar-compares-to-cameras-and-radars (accessed on 01 08 2024).
- Zhang, J.; Zhang, J.; Chen, B.; Gao, J.; Ji, S.; Zhang, X.; Wang, Z. A perspective transformation method based on computer vision. In Proceedings of the 2020 IEEE International Conference on Artificial Intelligence and Computer Applications (ICAICA); 2020; pp. 765–768. [Google Scholar] [CrossRef]
- Nienaber, S.; Booysen, M.J.; Kroon, R. Detecting potholes using simple image processing techniques and real-world footage. In Proceedings of the 34th South Africa Transport Conference (SATC) Pretoria, South Africa, 6–9 July 2015. [Google Scholar]
- Pereira, V.; Tamura, S.; Hayamizu, S.; Fukai, H. A deep learning-based approach for road pothole detection in timor leste. In Proceedings of the 2018 IEEE International Conference on Service Operations and Logistics, and Informatics (SOLI). IEEE; 2018; pp. 279–284. [Google Scholar]
- Chen, H.; Yao, M.; Gu, Q. Pothole detection using location-aware convolutional neural networks. International Journal of Machine Learning and Cybernetics 2020, 11, 899–911. [Google Scholar] [CrossRef]
- Kumar, A.; Kalita, D.J.; Singh, V.P.; et al. A modern pothole detection technique using deep learning. In Proceedings of the 2nd International Conference on Data, Engineering and Applications (IDEA). IEEE; 2020; pp. 1–5. [Google Scholar]
- Dhiman, A.; Klette, R. Pothole detection using computer vision and learning. IEEE Transactions on Intelligent Transportation Systems 2019, 21, 3536–3550. [Google Scholar] [CrossRef]
- Dhiman, A.; Chien, H.J.; Klette, R. Road surface distress detection in disparity space. In Proceedings of the 2017 International Conference on Image and Vision Computing New Zealand (IVCNZ). IEEE; 2017; pp. 1–6. [Google Scholar]
- Maeda, H.; Kashiyama, T.; Sekimoto, Y.; Seto, T.; Omata, H. Generative adversarial network for road damage detection. Computer-Aided Civil and Infrastructure Engineering 2021, 36, 47–60. [Google Scholar] [CrossRef]
- Salaudeen, H.; Çelebi, E. Pothole Detection Using Image Enhancement GAN and Object Detection Network. Electronics 2022, 11, 1882. [Google Scholar] [CrossRef]
- Tan, M.; Pang, R.; Le, Q.V. Efficientdet: Scalable and efficient object detection. In Proceedings of the Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp.; pp. 10781–10790.
- Shaghouri, A.A.; Alkhatib, R.; Berjaoui, S. Real-time pothole detection using deep learning. arXiv, 2021; arXiv:2107.06356. [Google Scholar]
- Bučko, B.; Lieskovská, E.; Zábovská, K.; Zábovskỳ, M. Computer vision based pothole detection under challenging conditions. Sensors 2022, 22, 8878. [Google Scholar] [CrossRef] [PubMed]
- Rastogi, R.; Kumar, U.; Kashyap, A.; Jindal, S.; Pahwa, S. A comparative evaluation of the deep learning algorithms for pothole detection. In Proceedings of the 2020 IEEE 17th India Council International Conference (INDICON). IEEE; 2020; pp. 1–6. [Google Scholar]
- Kocur, V. Perspective transformation for accurate detection of 3d bounding boxes of vehicles in traffic surveillance. In Proceedings of the Proceedings of the 24th Computer Vision Winter Workshop; 2019; Vol.2, pp. 33–41. [Google Scholar]
- Lee, W.Y.; Jovanov, L.; Philips, W. Multi-View Target Transformation for Pedestrian Detection. In Proceedings of the Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) Workshops, January 2023, pp.; pp. 90–99.
- Wang, K.; Fang, B.; Qian, J.; Yang, S.; Zhou, X.; Zhou, J. Perspective Transformation Data Augmentation for Object Detection. IEEE Access 2020, 8, 4935–4943. [Google Scholar] [CrossRef]
- Hou, Y.; Zheng, L.; Gould, S. Multiview Detection with Feature Perspective Transformation. In Proceedings of the Computer Vision – ECCV 2020; Vedaldi, A.; Bischof, H.; Brox, T.; Frahm, J.M., Eds., Cham; 2020; pp. 1–18. [Google Scholar]
- Jocher, G. YOLOv5 by Ultralytics. 2020. Available online: https://github.com/ultralytics/yolov5 (accessed on 4 8 2024). [CrossRef]
- Liu, W.; Anguelov, D.; Erhan, D.; Szegedy, C.; Reed, S.; Fu, C.Y.; Berg, A.C. Ssd: Single shot multibox detector. In Proceedings of the Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, 11–14 October 2016; Proceedings, Part I 14. Springer, 2016; pp. 21–37. [Google Scholar]
- Redmon, J.; Farhadi, A. Yolov3: An incremental improvement. arXiv, 2018; arXiv:1804.02767. [Google Scholar]
- Liu, K.; Fu, Z.; Jin, S.; Chen, Z.; Zhou, F.; Jiang, R.; Chen, Y.; Ye, J. ESOD: Efficient Small Object Detection on High-Resolution Images. arXiv, 2024; arXiv:2407.16424. [Google Scholar]
- Saponara, S.; Elhanashi, A. Impact of image resizing on deep learning detectors for training time and model performance. In Proceedings of the International Conference on Applications in Electronics Pervading Industry, Environment and Society. Springer; 2021; pp. 10–17. [Google Scholar]
- Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; et al. Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems 2019, 32. [Google Scholar]
- Wang, C.Y.; Bochkovskiy, A.; Liao, H.Y.M. YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors. In Proceedings of the Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp.; pp. 7464–7475.
- Jocher, G.; Chaurasia, A.; Qiu, J. Ultralytics YOLO. 2023. Available online: https://github.com/ultralytics/ultralytics (accessed on 4 8 2024).
- Nozick, V. Multiple view image rectification. In Proceedings of the 2011 1st International Symposium on Access Spaces (ISAS). IEEE; 2011; pp. 277–282. [Google Scholar]
- El Shair, Z.; Rawashdeh, S. High-temporal-resolution event-based vehicle detection and tracking. Optical Engineering 2023, 62, 031209–031209. [Google Scholar] [CrossRef]
- Kocur, V.; Ftáčnik, M. Detection of 3D bounding boxes of vehicles using perspective transformation for accurate speed measurement. Machine Vision and Applications 2020, 31, 62. [Google Scholar] [CrossRef]
- Nienaber, S.; Kroon, R.; Booysen, M.J. A comparison of low-cost monocular vision techniques for pothole distance estimation. In Proceedings of the 2015 IEEE symposium series on computational Intelligence. IEEE; 2015; pp. 419–426. [Google Scholar]
- Quach, L.D.; Quoc, K.N.; Quynh, A.N.; Ngoc, H.T. Evaluating the effectiveness of YOLO models in different sized object detection and feature-based classification of small objects. Journal of Advances in Information Technology 2023, 14, 907–917. [Google Scholar] [CrossRef]





| Approach | Pothole Distance | Metric (%) | ||||
|---|---|---|---|---|---|---|
| AP50:95 | AP50 | AP75 | ARmax=1 | ARmax=10 | ||
| Image as Is | All | 17.3 | 41.8 | 10.5 | 15.4 | 22.6 |
| Near | 20.8 | 49.4 | 14.0 | 18.9 | 27.0 | |
| Medium | 16.8 | 42.4 | 9.3 | 18.8 | 23.5 | |
| Far | 4.5 | 11.9 | 1.5 | 5.8 | 6.6 | |
| Bottom Cropped | All | 17.9 | 42.5 | 11.4 | 15.9 | 23.3 |
| Near | 21.4 | 49.0 | 14.1 | 19.4 | 27.3 | |
| Medium | 18.7 | 46.2 | 11.5 | 20.6 | 25.9 | |
| Far | 3.9 | 12.4 | 1.5 | 5.5 | 6.6 | |
| Double Cropped | All | 16.7 | 43.7 | 9.0 | 15.0 | 22.0 |
| Near | 21.3 | 53.5 | 11.6 | 19.3 | 27.0 | |
| Medium | 17.6 | 46.9 | 10.2 | 20.9 | 23.9 | |
| Far | 8.8 | 25.4 | 4.3 | 10.4 | 13.4 | |
| Auto Transformation | All | 25.2 | 54.2 | 19.3 | 20.0 | 32.0 |
| Near | 27.1 | 57.1 | 22.6 | 22.8 | 33.6 | |
| Medium | 30.0 | 64.2 | 23.8 | 31.6 | 37.2 | |
| Far | 17.0 | 39.2 | 10.5 | 19.0 | 25.0 | |
| Approach | Object Detection Model | Parameters (M) | FLOPs (G) | Pothole Distance | Metric (%) | ||||
|---|---|---|---|---|---|---|---|---|---|
| AP50:95 | AP50 | AP75 | ARmax=1 | ARmax=10 | |||||
| Image As Is | YOLOv5-Small | 7.2 | 16.5 | All | 17.3 | 41.8 | 10.5 | 15.4 | 22.6 |
| Near | 20.8 | 49.4 | 14.0 | 18.9 | 27.0 | ||||
| Medium | 16.8 | 42.4 | 9.3 | 18.8 | 23.5 | ||||
| Far | 4.5 | 11.9 | 1.5 | 5.8 | 6.6 | ||||
| YOLOv5-Medium | 21.2 | 49.0 | All | 18.3 | 43.4 | 12.6 | 16.2 | 23.6 | |
| Near | 21.5 | 49.0 | 16.1 | 19.3 | 27.0 | ||||
| Medium | 18.9 | 45.6 | 12.0 | 20.6 | 26.1 | ||||
| Far | 5.1 | 16.2 | 1.8 | 7.6 | 8.5 | ||||
| YOLOv5-Large | 46.5 | 109.1 | All | 18.4 | 43.3 | 12.6 | 15.7 | 23.8 | |
| Near | 22.4 | 50.3 | 16.5 | 19.5 | 28.3 | ||||
| Medium | 17.7 | 45.2 | 10.6 | 20.1 | 25.0 | ||||
| Far | 4.6 | 13.6 | 1.5 | 6.4 | 7.6 | ||||
| Auto Transformation | YOLOv5-Small | 7.2 | 16.5 | All | 25.2 | 54.2 | 19.3 | 20.0 | 32.0 |
| Near | 27.1 | 57.1 | 22.6 | 22.8 | 33.6 | ||||
| Medium | 30.0 | 64.2 | 23.8 | 31.6 | 37.2 | ||||
| Far | 17.0 | 39.2 | 10.5 | 19.0 | 25.0 | ||||
| YOLOv5-Medium | 21.2 | 49.0 | All | 23.2 | 51.2 | 17.2 | 19.2 | 29.4 | |
| Near | 25.8 | 55.0 | 20.1 | 22.5 | 31.7 | ||||
| Medium | 27.4 | 57.0 | 21.5 | 28.8 | 33.6 | ||||
| Far | 15.9 | 39.8 | 10.7 | 16.7 | 22.4 | ||||
| YOLOv5-Large | 46.5 | 109.1 | All | 24.8 | 54.6 | 18.1 | 20.0 | 31.4 | |
| Near | 26.9 | 55.9 | 22.0 | 22.9 | 33.5 | ||||
| Medium | 27.9 | 60.8 | 19.3 | 31.0 | 34.9 | ||||
| Far | 18.8 | 45.2 | 11.9 | 18.7 | 25.2 | ||||
| Configuration | Metric (%) | |||||
|---|---|---|---|---|---|---|
| Preproc. Augs. | Neg. Images | AP50:95 | AP50 | AP75 | ARmax=1 | ARmax=10 |
| 24.5 | 53.6 | 17.9 | 18.9 | 31.2 | ||
| √ | 23.7 | 52.3 | 17.1 | 19.1 | 30.3 | |
| √ | 25.2 | 54.2 | 19.3 | 20.0 | 32.0 | |
| √ | √ | 23.4 | 51.5 | 17.3 | 19.0 | 29.7 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2024 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
