Submitted:
12 September 2025
Posted:
15 September 2025
You are already at the latest version
Abstract
Keywords:
1. Introduction
- To address the critical challenge of capturing long-range dependencies and complex contextual information in expansive aerial imagery, a Transformer module is incorporated into the network. This enhancement enables the model to more accurately identify and delineate defects ranging from minute cracks to large-area potholes, which is a common scenario in UAV inspections.
- To overcome the limitations of handling extreme scale variations inherent in UAV perspectives, a multi-level feature pyramid network is employed to effectively fuse features across different scales. Coupled with an optimized detection head structure, this ensures robust performance on targets of various sizes while maintaining computational efficiency for detecting small targets.
- A Spatial-Channel Interaction Module (SCIM) is designed, building upon the Transformer and feature pyramid network. This module facilitates simultaneous capture of global and local features by jointly modeling spatial and channel information, significantly enhancing the feature representation power for complex aerial scenes.
2. Related Works
2.1. Traditional Road Inspection Methods
2.2. Deep Learning-Based Object Detection
2.3. UAV-Based Visual Inspection
3. Method
3.1. Multi-Scale Feature Representation and Optimization Strategies for Aerial Imagery
3.2. Spatial and Channel Interaction Optimization
4. Experiments and Results
4.1. UAV Image Dataset and Experimental Setup
4.2. Ablation Study on Component Effectiveness
4.3. Comparative Evaluation with State-of-the-Art Models
5. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Zhou, Y.; Yue, Y.; Yan, B.; et al. Collaborative Target Tracking Algorithm for Multi-Agent Based on MAPPO and BCTD. Drones, 2025, 9, 521. [Google Scholar] [CrossRef]
- Yang, B.; Tao, T.; Wu, W.; et al. MultiDistiller: Efficient Multimodal 3D Detection via Knowledge Distillation for Drones and Autonomous Vehicles. Drones, 2025, 9, 322. [Google Scholar] [CrossRef]
- Huang, Y.; Fan, J.Y.; Hu, J.Z.Y. TBi-YOLOv5: A surface defect detection model for crane wire with Bottleneck Transformer and small target detection layer. Proceedings of the Institution of Mechanical Engineers, Part C. Journal of mechanical engineering science, 2024, 238, 2425–2438. [Google Scholar] [CrossRef]
- Su, Y.; Deng, J.; Sun, R.; et al. A Unified Transformer Framework for Group-based Segmentation: Co-Segmentation, Co-Saliency Detection and Video Salient Object Detection. 2022, 26, 313–325. [Google Scholar] [CrossRef]
- Wang, A.; Ren, C.; Zhao, S.M.S. Attention guided multi-level feature aggregation network for camouflaged object detection. Image and vision computing, 2024, 144, 1. [Google Scholar] [CrossRef]
- Li, H.; Zhang, R.; Pan, Y.; et al. LR-FPN: Enhancing Remote Sensing Object Detection with Location Refined Feature Pyramid Network. IEEE, 2024, 1-8.
- Ha, T.T.; Chaisomphob, T. Automated Localization and Classification of Expressway Pole-Like Road Facilities from Mobile Laser Scanning Data. Advances in Civil Engineering 2020, 2020, 5016783.1–5016783.18. [Google Scholar]
- Ma, Y.; Lei, W.; Pang, Z.; et al. Rebar Clutter Suppression and Road Defects Localization in GPR B-Scan Images Based on SuppRebar-GAN and EC-Yolov7 Networks. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62, 1–14. [Google Scholar] [CrossRef]
- Lv, Y.; Wang, G.; Hu, X. MACHINE LEARNING BASED ROAD DETECTION FROM HIGH RESOLUTION IMAGERY. ISPRS - International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences, 2016, 41, 891–898. [Google Scholar]
- Zhang, H.; Shao, F.; Chu, W.; et al. Faster R-CNN based on frame difference and spatiotemporal context for vehicle detection. Signal, Image and Video Processing, 2024, 18, 7013–7027. [Google Scholar] [CrossRef]
- Mohd Yusof, N.; Sophian, A.; Mohd Zaki, H.F.; et al. Assessing the performance of YOLOv5, YOLOv6, and YOLOv7 in road defect detection and classification: a comparative study. Bulletin of Electrical Engineering & Informatics, 2024, 13, 350. [Google Scholar]
- Zhang, B.; Fang, S.; Li, Z. Research on Surface Defect Detection of Rare-Earth Magnetic Materials Based on Improved SSD. Complexity, 2021, 2021, 1–10. [Google Scholar] [CrossRef]
- Zhang, L.; Yan, S.F.; Hong, J.; et al. An improved defect recognition framework for casting based on DETR algorithm. Journal of Iron and Steel Research, International Edition, 2023, 30, 949–959. [Google Scholar] [CrossRef]
- Zhu, W.; Zhang, H.; Zhang, C.; Zhu, X.; Guan, Z.; Jia, J. Surface defect detection and classification of steel using an efficient Swin Transformer. 2023, 57, 1572. [Google Scholar] [CrossRef]
- Wu, Y.; Liao, K.Y.; Chen, J.; et al. D-former: a U-shaped Dilated Transformer for 3D medical image segmentation. Neural Computing and Applications, 2022, 35, 1931–1944. [Google Scholar] [CrossRef]
- Wang, X.; Gao, H.; Jia, Z.; et al. A road defect detection algorithm incorporating partially transformer and multiple aggregate trail attention mechanisms. IOP Publishing Ltd, 2024, 36, 026003. [Google Scholar] [CrossRef]
- Kim, G.I.; Yoo, H.; Cho, H.J.; et al. Defect Detection Model Using Time Series Data Augmentation and Transformation. Computers, Materials & Continua, 2024, 78, 1713. [Google Scholar]
- Jiang, T.Y.; Liu, Z.Y.; Zhang, G.Z. YOLOv5s-road: Road surface defect detection under engineering environments based on CNN-transformer and adaptively spatial feature fusion. Measurement, 2025, 242, 115990. [Google Scholar] [CrossRef]
- Wang, J.; Meng, R.; Huang, Y.; et al. Road defect detection based on improved YOLOv8s model. Scientific Reports, 2024, 14, 1. [Google Scholar] [CrossRef]
- Fang, Z.; Shi, Z.; Wang, X.; et al. Roadbed Defect Detection from Ground Penetrating Radar B-scan Data Using Faster RCNN. 2020, 131–137. [Google Scholar] [CrossRef]
- Sadhin, A.H.; Mohd Hashim, S.Z.; Samma, H.; et al. YOLO: A Competitive Analysis of Modern Object Detection Algorithms for Road Defects Detection Using Drone Images. Baghdad Science Journal, 2024, 21, 2167. [Google Scholar] [CrossRef]
- Arya, D.; Maeda, H.; Ghosh, S.K.; et al. Global Road Damage Detection: A Large-Scale Dataset and Benchmark. IEEE Transactions on Intelligent Transportation Systems, 2022, 23, 23294–23305. [Google Scholar]
- Kim, G.I.; Yoo, H.; Cho, H.J. Lightweight vision transformer for real-time road damage detection in UAV imagery. IEEE Geoscience and Remote Sensing Letters, 2023, 20, 6003205. [Google Scholar]
- Jinsheng Xiao, Haowen Guo, Jian Zhou, Tao Zhao, Qiuze Yu, Yunhua Chen, Zhongyuan Wang, Tiny object detection with context enhancement and feature purification, Expert Systems with Applications (ESWA), 2023, January, vol 211, 118665.




| Backbone | Transformer | SCIM | mAP@0.5 | Precision | Recall | Params (M) |
|---|---|---|---|---|---|---|
| CNN (Baseline) | ✗ | ✗ | 0.7357 | 0.7111 | 0.7172 | 7.2 |
| CNN | ✓ | ✗ | 0.8157 | 0.7904 | 0.7831 | 18.5 |
| CNN | ✓ | ✓ | 0.9128 | 0.904 | 0.8328 | 21.3 |
| Model | mAP@0.5 | Inference Time (ms) | Platform |
|---|---|---|---|
| RoadNet (Ours) | 0.9128 | 210.1 | CPU (Intel i9-13900K) |
| Faster R-CNN [20] | 0.4283 | 1221.3 | CPU (Intel i9-13900K) |
| YOLOv5s [18] | 0.7225 | 220.4 | CPU (Intel i9-13900K) |
| YOLOv5l [21] | 0.7234 | 993.8 | CPU (Intel i9-13900K) |
| YOLOv8s [19] | 0.7284 | 330.6 | CPU (Intel i9-13900K) |
| DETR [16] | 0.7012 | 1850.5 | CPU (Intel i9-13900K) |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).