Submitted:
28 November 2025
Posted:
28 November 2025
Read the latest preprint version here
Abstract
Keywords:
1. Introduction
2. Related Works
3. Proposed Method
3.1. Datasets
3.2. Proposed Framework
3.3. Road Marking Detection (RMD)

3.4. Road Marking Segmentation (RMS)
- Skip connections and the shallow-kernel decoder preserve edges and narrow strokes that are easily lost in downsampling;
- VGG16’s pretrained filters accelerate convergence and help when the target appearance varies due to paint aging and surface texture;
- The model is compact, exportable, and decoupled from detection (RMD) proposes instances and IDs; RMS focuses solely on pixel-accurate shapes, keeping the pipeline modular and easy to maintain.
3.5. Damage Estimation
- Bayesian pixel labeling and decision threshold
3.6. Localization
- Reference model construction
- 2.
- Assigning 3D coordinates to new detections
4. Results
4.1. Road Marking Detection Results
4.2. Road Marking Segmentation Results
4.3. Damage Estimation Methods
- Otsu thresholding: Otsu selects the gray level maximizing the between-class variance computed from the image histogram [32]:where and denote the cumulative class probability and mean for class . In practice, Otsu can be sensitive to class imbalance and weak bimodality;
- K-means (): With two clusters, -means minimizes the within-cluster sum of squares; in 1-D the decision boundary is the midpoint of the two centroids. The classical objective and Lloyd-type algorithm are well documented (Lloyd, 1982), and improved seeding strategies such as -means++ enhance robustness [33].
- Direct GMM on raw pixels: A two-component Gaussian mixture is fitted by the EM algorithm, and labels are assigned by the MAP decision rule. Mixture modeling texts provide extensive discussion of convergence/local optimum issues in finite mixtures [34].
4.4. Localization
5. Conclusion and Discussion
- End-to-end marking assessment from drone imagery. We present a practical workflow that localizes, segments, and rates the condition of road markings, producing georeferenced outputs suitable for maintenance planning.
- Detector-guided, segmentation-refined delineation. Detections from YOLOv9 guide crops, and a standalone U-Net refines boundaries and thin structures to obtain pixel-level masks.
- Distribution-based damage quantification. We operationalize KDE/GMM within each segmented instance to partition intact versus distressed appearances and derive an interpretable damage ratio for each object.
- Generalization on unseen areas. We evaluate the pipeline on urban scenes not used in training and report standard object detection and segmentation metrics alongside stability analyses of the damage ratio, demonstrating readiness for operational deployment (contextualized by the role of marking conditions in ADAS and safety).
Funding
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Babić, D.; Fiolić, M.; Babić, D.; Gates, T. Road Markings and Their Impact on Driver Behaviour and Road Safety: A Systematic Review of Current Findings. J. Adv. Transp. 2020, 2020, 1–19. [Google Scholar] [CrossRef]
- Mahlberg, J.A.; Sakhare, R.S.; Li, H.; Mathew, J.K.; Bullock, D.M.; Surnilla, G.C. Prioritizing Roadway Pavement Marking Maintenance Using Lane Keep Assist Sensor Data. Sensors 2021, 21, 6014. [Google Scholar] [CrossRef] [PubMed]
- Ronneberger, O.; Fischer, P.; Brox, T. U-Net: Convolutional Networks for Biomedical Image Segmentation 2015.
- Rosenblatt, M. Remarks on Some Nonparametric Estimates of a Density Function. Ann. Math. Stat. 1956, 27, 832–837. [Google Scholar] [CrossRef]
- Parzen, E. On Estimation of a Probability Density Function and Mode. Ann. Math. Stat. 1962, 33, 1065–1076. [Google Scholar] [CrossRef]
- Dempster, A.P.; Laird, N.M.; Rubin, D.B. Maximum Likelihood from Incomplete Data Via the EM Algorithm. J. R. Stat. Soc. Ser. B Stat. Methodol. 1977, 39, 1–22. [Google Scholar] [CrossRef]
- Guan, H.; Lei, X.; Yu, Y.; Zhao, H.; Peng, D.; Marcato Junior, J.; Li, J. Road Marking Extraction in UAV Imagery Using Attentive Capsule Feature Pyramid Network. Int. J. Appl. Earth Obs. Geoinformation 2022, 107, 102677. [Google Scholar] [CrossRef]
- Bu, T.; Zhu, J.; Ma, T. A UAV Photography–Based Detection Method for Defective Road Marking. J. Perform. Constr. Facil. 2022, 36, 04022035. [Google Scholar] [CrossRef]
- Wu, J.; Liu, W.; Maruyama, Y. Street View Image-Based Road Marking Inspection System Using Computer Vision and Deep Learning Techniques. Sensors 2024, 24, 7724. [Google Scholar] [CrossRef] [PubMed]
- Yaseen, M. What Is YOLOv8: An In-Depth Exploration of the Internal Features of the Next-Generation Object Detector 2024.
- Song, L.; Zhao, F.; Han, J.; Li, S.; Hu, J. Road Marking Detection from UAV Perspective Based on Improved YOLOv3. In Proceedings of the International Conference on Image, Signal Processing, and Pattern Recognition (ISPP 2024); Bilas Pachori, R., Chen, L., Eds.; SPIE: Guangzhou, China, 13 June 2024; p. 30. [Google Scholar]
- Wu, J.; Liu, W.; Maruyama, Y. Automated Road-Marking Segmentation via a Multiscale Attention-Based Dilated Convolutional Neural Network Using the Road Marking Dataset. Remote Sens. 2022, 14, 4508. [Google Scholar] [CrossRef]
- Dong, Z.; Zhang, H.; Zhang, A.A.; Liu, Y.; Lin, Z.; He, A.; Ai, C. Intelligent Pixel-Level Pavement Marking Detection Using 2D Laser Pavement Images. Measurement 2023, 219, 113269. [Google Scholar] [CrossRef]
- Wei, C.; Li, S.; Wu, K.; Zhang, Z.; Wang, Y. Damage Inspection for Road Markings Based on Images with Hierarchical Semantic Segmentation Strategy and Dynamic Homography Estimation. Autom. Constr. 2021, 131, 103876. [Google Scholar] [CrossRef]
- Redmon, J.; Divvala, S.; Girshick, R.; Farhadi, A. You Only Look Once: Unified, Real-Time Object Detection 2015.
- Lin, T.-Y.; Dollar, P.; Girshick, R.; He, K.; Hariharan, B.; Belongie, S. Feature Pyramid Networks for Object Detection. In Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR); Honolulu, HI, July 2017; pp. 936–944. [Google Scholar]
- Liu, S.; Qi, L.; Qin, H.; Shi, J.; Jia, J. Path Aggregation Network for Instance Segmentation 2018.
- Tian, Z.; Shen, C.; Chen, H.; He, T. FCOS: Fully Convolutional One-Stage Object Detection. In Proceedings of the 2019 IEEE/CVF International Conference on Computer Vision (ICCV); Seoul, Korea (South), October 2019; pp. 9626–9635. [Google Scholar]
- Simonyan, K.; Zisserman, A. Very Deep Convolutional Networks for Large-Scale Image Recognition 2014.
- Silverman, B.W. Density Estimation for Statistics and Data Analysis, 1st ed.; Routledge, 2018; ISBN 978-1-315-14091-9. [Google Scholar]
- Schonberger, J.L.; Frahm, J.-M. Structure-from-Motion Revisited. In Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, June 2016; pp. 4104–4113. [Google Scholar]
- Ito, K.; Ito, T.; Aoki, T. PM-MVS: PatchMatch Multi-View Stereo. Mach. Vis. Appl. 2023, 34, 32. [Google Scholar] [CrossRef]
- Rusu, R.B.; Cousins, S. 3D Is Here: Point Cloud Library (PCL). In Proceedings of the 2011 IEEE International Conference on Robotics and Automation; Shanghai, China, May 2011; pp. 1–4. [Google Scholar]
- Lindeberg, T. Scale Invariant Feature Transform. Scholarpedia 2012, 7, 10491. [Google Scholar] [CrossRef]
- Lowe, D.G. Distinctive Image Features from Scale-Invariant Keypoints. Int. J. Comput. Vis. 2004, 60, 91–110. [Google Scholar] [CrossRef]
- Shepard, D. A Two-Dimensional Interpolation Function for Irregularly-Spaced Data. In Proceedings of the Proceedings of the 1968 23rd ACM national conference on -; ACM Press: Not Known. 1968; 517–524. [Google Scholar]
- Khozaimi, A.; Darti, I.; Anam, S.; Kusumawinahyu, W.M. Hybrid Dense-UNet201 Optimization for Pap Smear Image Segmentation Using Spider Monkey Optimization 2025.
- Zhang, L.; Li, X.; Arnab, A.; Yang, K.; Tong, Y.; Torr, P.H.S. Dual Graph Convolutional Network for Semantic Segmentation 2019.
- Xiao, T.; Liu, Y.; Zhou, B.; Jiang, Y.; Sun, J. Unified Perceptual Parsing for Scene Understanding 2018.
- Chen, L.-C.; Zhu, Y.; Papandreou, G.; Schroff, F.; Adam, H. Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation 2018.
- Zhao, H.; Shi, J.; Qi, X.; Wang, X.; Jia, J. Pyramid Scene Parsing Network 2016.
- Otsu, N. A Threshold Selection Method from Gray-Level Histograms. IEEE Trans. Syst. Man Cybern. 1979, 9, 62–66. [Google Scholar] [CrossRef]
- Choo, D.; Grunau, C.; Portmann, J.; Rozhoň, V. K-Means++: Few More Steps Yield Constant Approximation 2020.
- McLachlan, G.; Peel, D. Finite Mixture Models; Wiley Series in Probability and Statistics; 1st ed.; Wiley, 2000; ISBN 978-0-471-00626-8.
- Soille, P. Morphological Image Analysis: Principles and Applications; 2. ed., corr. 2. print.; Springer: Berlin Heidelberg, 2010; ISBN 978-3-642-07696-1. [Google Scholar]
- Kuhn, J.W.; Padgett, W.J.; Surles, J.G. Absolute Error Criteria for Bandwidth Selection in Density Estimation from Censored Data. J. Stat. Comput. Simul. 2001, 70, 215–230. [Google Scholar] [CrossRef]












| Overlap | Sidelap | Flight Altitude | Ground Sampling Distance (GSD) |
|---|---|---|---|
| 80% | 70% | 100 meters | 0.6 centimeters |
| Sensor Size | Resolution | Focal Length | Image Format |
|---|---|---|---|
| 13.2 mm x 8.8 mm | 5,472 x 3,648 pixels | 9 milimeters | JPEG |
| Model | Precision (%) | Recall (%) | F1 (%) | mAP50 (%) | mAP50-95 (%) | Speed (ms per image) |
GFLOPs |
|---|---|---|---|---|---|---|---|
| YOLOv8x | 95.506 | 90.592 | 92.984 | 94.764 | 64.256 | 0.6 | 257.5 |
| YOLOv9e | 95.363 | 91.982 | 93.642 | 95.497 | 65.553 | 0.8 | 189.2 |
| YOLOv10x | 95.451 | 90.669 | 92.999 | 95.803 | 64.757 | 0.1 | 160.0 |
| YOLOv11x | 95.556 | 91.66 | 93.567 | 95.693 | 64.478 | 0.5 | 172.6 |
| YOLOv12x | 95.402 | 91.583 | 93.454 | 95.433 | 64.749 | 1.3 | 198.6 |
| Model | mIoU | F1 | MAE | ||||||
|---|---|---|---|---|---|---|---|---|---|
| Train | Val | Unseen | Train | Val | Unseen | Train | Val | Unseen | |
| Densenet201-UNet [27] | 95.09 | 95.30 | 93.88 | 97.46 | 97.57 | 96.81 | 1.73 | 1.65 | 2.33 |
| Unet [3] | 94.11 | 94.58 | 93.84 | 96.97 | 96.93 | 96.72 | 1.95 | 1.86 | 2.25 |
| VGG16 - UNet | 95.53 | 95.60 | 94.21 | 97.69 | 97.73 | 97.00 | 1.49 | 1.42 | 2.12 |
| GCN [28] | 93.5 8 | 93.72 | 92.58 | 96.72 | 96.71 | 95.95 | 2.12 | 2.10 | 2.82 |
| Upernet [29] | 94.87 | 95.18 | 93.79 | 97.36 | 97.28 | 96.59 | 1.71 | 1.62 | 2.28 |
| Deeplab HDC DUC [30] | 90.6 | 90.76 | 90.44 | 95.08 | 94.91 | 94.84 | 3.12 | 3.09 | 3.53 |
| PSPnet [31] | 69.77 | 71.76 | 72.86 | 80.89 | 82.44 | 83.47 | 3.32 | 3.35 | 4.34 |
| Method | IAE median | ISE median |
|---|---|---|
| KDE/GMM | 0.021 | 0.002 |
| Otsu | 0.063 | 0.086 |
| k-means (k=2) | 0.063 | 0.087 |
| GMM-direct | 0.067 | 0.098 |
| Identfication (ID) | Location | Condition (Pecentage of Damage %) |
|
|---|---|---|---|
| Latitude | Longtitude | ||
| Crosswalk | 37° 28' 55.9" N | 127° 06' 16.38" N | 3.5 |
| Left Turn | 37° 29' 15.89" N | 127° 06' 04.04" N | 20.0 |
| Right Turn | 37° 29' 13.67" N | 127° 06' 08.32" N | 7.1 |
| Right Turn | 37° 29' 13.12" N | 127° 06' 08.32" N | 53.8 |
| Straight and Left Turn | 37° 28' 56.78" N | 127° 06' 16.38" N | 25.4 |
| Straight Arrow | 37° 28' 56.80" N | 127° 06' 16.52" N | 58.5 |
| Straight Arrow | 37° 28' 56.82" N | 127° 06' 16.63" N | 2.7 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).