Submitted:
16 July 2026
Posted:
17 July 2026
You are already at the latest version
Abstract
Keywords:
1. Introduction
- 1.
- A DA-C3k2 block is proposed to enhance the ability of lightweight backbone networks to represent directional continuity and weak texture cues in thin cracks.
- 2.
- A crack visual analysis framework with a shared backbone and task-decoupled supervision is constructed, allowing direction-aware features to serve both crack localization and mask prediction without forcing detection, segmentation, and measurement into a single optimization target.
- 3.
- Systematic experiments are conducted on detection, segmentation, external transferability, and inference efficiency to verify the effects of DA-C3k2 and Tversky loss on crack mask quality, boundary metrics, and generalization.
2. Related Work
2.1. UAV-Based Bridge Crack Inspection
2.2. Deep Learning-Based Crack Detection
2.3. Crack Segmentation and Lightweight Mask Prediction
2.4. Directional Feature Modeling and Imbalance-Aware Losses
3. Proposed DA-C3k2-Based Crack Detection and Segmentation Framework
3.1. Overall Framework

3.2. Shared Backbone and Task-Decoupled Supervision
3.3. Direction-Aware C3k2 Block

3.4. Detection Head and Segmentation Head
3.5. Tversky Loss for Sparse Crack Masks
3.6. Implementation Details
3.7. Application Example: Mask-Derived Geometric Estimation
4. Experiments
4.1. Datasets and Evaluation Protocols
| Experiment | Dataset/source | Classes | Split size | Purpose |
|---|---|---|---|---|
| Detection | UAV-PDD2023 crack subset | 4 crack types | 1680/240/480 | Train/val/test |
| Segmentation | Crack segmentation benchmark | Crack/background | 7704/963/963 | Original mask evaluation |
| External transfer | COCO-format crack export | Crack/background | 1239/200/112 | Validation split for transfer and speed |
4.2. Training Settings
4.3. Detection Results
4.4. Detection Ablation Study
4.5. Segmentation Ablation Results

4.6. Segmentation Comparison with Larger Semantic Segmentation Models

4.7. External Dataset Transferability Evaluation

4.8. CPU Inference Efficiency
4.9. Qualitative Analysis
5. Discussion
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| BF1 | Boundary-F1 score |
| B-IoU | Boundary intersection over union |
| CA | Coordinate attention |
| CE | Cross-entropy |
| DA-C3k2 | Direction-Aware C3k2 block |
| GFLOPs | Giga floating-point operations |
| mAP | Mean average precision |
| UAV | Unmanned aerial vehicle |
References
- Deng, J.; Singh, A.; Zhou, Y.; Lu, Y.; Lee, V.C.S. Review on computer vision-based crack detection and quantification methodologies for civil structures. Constr. Build. Mater. 2022, 356, 129238. [Google Scholar] [CrossRef]
- Ai, D.; Jiang, G.; Lam, S.K.; He, P.; Li, C. Computer vision framework for crack detection of civil infrastructure-A review. Eng. Appl. Artif. Intell. 2023, 117, 105478. [Google Scholar] [CrossRef]
- Luo, K.; Kong, X.; Zhang, J.; Hu, J.; Li, J.; Tang, H. Computer Vision-Based Bridge Inspection and Monitoring: A Review. Sensors 2023, 23, 7863. [Google Scholar] [CrossRef] [PubMed]
- Aliyari, M.; Droguett, E.L.; Ayele, Y.Z. UAV-Based Bridge Inspection via Transfer Learning. Sustainability 2021, 13, 11359. [Google Scholar] [CrossRef]
- Saeed, M.S. Unmanned Aerial Vehicle for Automatic Detection of Concrete Crack using Deep Learning. In Proceedings of the 2021 2nd International Conference on Robotics, Electrical and Signal Processing Techniques (ICREST), 2021; pp. 624–628. [Google Scholar] [CrossRef]
- Phan, T.N.; Nguyen, H.H.; Ha, T.T.H.; Thai, H.T.; Le, K.H. Deep Learning Models for UAV-Assisted Bridge Inspection: A YOLO Benchmark Analysis. In Proceedings of the 2024 International Conference on Advanced Technologies for Communications (ATC), 2024. [Google Scholar] [CrossRef]
- Cha, Y.J.; Choi, W.; Buyukozturk, O. Deep Learning-Based Crack Damage Detection Using Convolutional Neural Networks. In Computer-Aided Civil and Infrastructure Engineering; 2017. [Google Scholar] [CrossRef]
- Zhang, L.; Yang, F.; Zhang, Y.D.; Zhu, Y.J. Road crack detection using deep convolutional neural network. In Proceedings of the 2016 IEEE International Conference on Image Processing (ICIP), 2016. [Google Scholar] [CrossRef]
- Xu, H.; Su, X.; Wang, Y.; Cai, H.; Cui, K.; Chen, X. Automatic Bridge Crack Detection Using a Convolutional Neural Network. Appl. Sci. 2019, 9, 2867. [Google Scholar] [CrossRef]
- Liu, Y.; Zhou, T.; Xu, J.; Hong, Y.; Pu, Q.; Wen, X. Rotating Target Detection Method of Concrete Bridge Crack Based on YOLO v5. Appl. Sci. 2023, 13, 11118. [Google Scholar] [CrossRef]
- Zou, X.; Jiang, S.; Yang, J.; Huang, X. Concrete Bridge Crack Detection Based on YOLO v8s in Complex Background. In Proceedings of International Conference on Image, Vision and Intelligent Systems 2023 (ICIVIS 2023); Springer Nature Singapore, 2024; pp. 436–443. [Google Scholar] [CrossRef]
- Yu, A.; Gao, Y.; Xiong, Y.; Liu, W.; She, J. Mobile-YOLO: A Lightweight YOLO for Road Crack Detection on Mobile Devices. J. Adv. Comput. Intell. Intell. Inform. 2025, 29, 1443–1453. [Google Scholar] [CrossRef]
- Ronneberger, O.; Fischer, P.; Brox, T. U-Net: Convolutional Networks for Biomedical Image Segmentation. In Medical Image Computing and Computer-Assisted Intervention – MICCAI 2015; Springer International Publishing, 2015; pp. 234–241. [Google Scholar] [CrossRef]
- Zou, Q.; Zhang, Z.; Li, Q.; Qi, X.; Wang, Q.; Wang, S. DeepCrack: Learning Hierarchical Convolutional Features for Crack Detection. IEEE Trans. Image Process. 2019, 28, 1498–1512. [Google Scholar] [CrossRef] [PubMed]
- Liu, Y.; Yao, J.; Lu, X.; Xie, R.; Li, L. DeepCrack: A deep hierarchical feature learning architecture for crack segmentation. Neurocomputing 2019, 338, 139–153. [Google Scholar] [CrossRef]
- Sun, X.; Xie, Y.; Jiang, L.; Cao, Y.; Liu, B. DMA-Net: DeepLab With Multi-Scale Attention for Pavement Crack Segmentation. IEEE Trans. Intell. Transp. Syst. 2022, 23, 18392–18403. [Google Scholar] [CrossRef]
- Abraham, N.; Khan, N.M. A Novel Focal Tversky Loss Function With Improved Attention U-Net for Lesion Segmentation. In Proceedings of the 2019 IEEE 16th International Symposium on Biomedical Imaging (ISBI 2019); IEEE, Apr 2019; pp. 683–687. [Google Scholar] [CrossRef]
- Wang, M.; Liu, Y.; Hao, Y.; Gong, S.; Qiu, C.; Cai, R.; Deng, Y. Research on road crack detection algorithm based on YOLO-SW. PeerJ Comput. Sci. 2026, 12, e3783. [Google Scholar] [CrossRef]
- Chen, L.C.; Zhu, Y.; Papandreou, G.; Schroff, F.; Adam, H. Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation. In Computer Vision – ECCV 2018; Springer International Publishing, 2018; pp. 833–851. [Google Scholar] [CrossRef]
- Yang, L.; Bai, S.; Liu, Y.; Yu, H. Multi-scale triple-attention network for pixelwise crack segmentation. Autom. Constr. 2023, 150, 104853. [Google Scholar] [CrossRef]
- Yue, B.; Dang, J.; Sun, Q.; Wang, Y.; Min, Y.; Wang, F. TSPCS-net: Two-stage pavement crack segmentation network based on encoder-decoder architecture. Eng. Appl. Artif. Intell. 2025, 141, 109840. [Google Scholar] [CrossRef]
- Lin, T.Y.; Goyal, P.; Girshick, R.; He, K.; Dollar, P. Focal Loss for Dense Object Detection. In Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV); IEEE, Oct 2017; pp. 2999–3007. [Google Scholar] [CrossRef]
- Milletari, F.; Navab, N.; Ahmadi, S.A. V-Net: Fully Convolutional Neural Networks for Volumetric Medical Image Segmentation. In Proceedings of the 2016 Fourth International Conference on 3D Vision (3DV); IEEE, Oct 2016; pp. 565–571. [Google Scholar] [CrossRef]
- Cho, H.; Yoon, H.J.; Jung, J.Y. Image-Based Crack Detection Using Crack Width Transform (CWT) Algorithm. IEEE Access 2018, 6, 60100–60114. [Google Scholar] [CrossRef]
- Yu, L.; He, S.; Liu, X.; Jiang, S.; Xiang, S. Intelligent Crack Detection and Quantification in the Concrete Bridge: A Deep Learning-Assisted Image Processing Approach. Adv. Civ. Eng. 2022. [Google Scholar] [CrossRef]
- Li, C.; Qin, H.; Tang, Y.; Zhao, H.; Pan, S.; Liu, J.; Luo, W. An Image-Based Concrete-Crack-Width Measurement Method Using Skeleton Pruning and the Edge-OrthoBoundary Algorithm. Buildings 2025, 15, 2489. [Google Scholar] [CrossRef]


| Protocol | Item | Setting |
|---|---|---|
| Detection | Model and schedule | YOLO11n with DA-C3k2; input size 640×640; 300 epochs; batch size 256; seed 0. |
| Detection | Optimization and loss | optimizer=auto; lr0=0.01; lrf=0.001; momentum 0.937; weight decay 0.0005; loss gains: box=7.5, cls=1.0, dfl=1.5. |
| Detection | Augmentation | mosaic=1.0; close_mosaic=50; HSV h/s/v=0.015/0.7/0.4; translate=0.1; scale=0.5; degrees=0; shear=0; fliplr=0.5; flipud=0; mixup/cutmix=0/0. |
| Detection | Validation | warmup epochs=3.0; validation NMS IoU=0.7; AMP enabled. |
| Segmentation | Model and schedule | YOLO11-Seg with DA-C3k2; input size 640×640; 80 fine-tuning epochs; batch size 64; seed 42. |
| Segmentation | Optimization and loss | optimizer=auto; lr0=0.001; lrf=0.1; momentum 0.937; weight decay 0.0005; loss gains: box=7.5, cls=0.5, dfl=1.5; Tversky weight , , . |
| Segmentation | Augmentation | mosaic=1.0; close_mosaic=10; HSV h/s/v=0.015/0.7/0.4; translate=0.1; scale=0.5; degrees=10; shear=2; fliplr=0.5; flipud=0.5; mixup/cutmix=0/0. |
| Segmentation | Validation and application thresholds | warmup epochs=1.0; validation NMS IoU=0.7; application-example thresholds: confidence=0.25, NMS IoU=0.5, mask threshold=0.5; AMP enabled. |
| Tversky search | Model and schedule | YOLO11-Seg; input size 640×640; 80 epochs; batch size 256; seed 42. |
| Tversky search | Optimization and loss | optimizer=auto; lr0=0.001; lrf=0.1; Tversky weight . |
| Tversky search | Search space and augmentation | , searched from 0.1/0.9 to 0.5/0.5; the remaining augmentation settings follow the segmentation protocol; AMP enabled. |
| Method | Params/M | GFLOPs | Latency/ms | FPS | mAP50 | mAP50–95 |
|---|---|---|---|---|---|---|
| YOLOv8n | 3.20 | 8.7 | 109.54 | 9.13 | 0.626 | 0.316 |
| YOLOv10n | 2.30 | 6.7 | 114.03 | 8.77 | 0.487 | 0.216 |
| YOLO11n baseline | 2.59 | 6.4 | 213.71 | 4.68 | 0.670 | 0.334 |
| YOLO12n | 2.60 | 6.3 | 234.40 | 4.27 | 0.371 | 0.147 |
| DA-C3k2 | 2.96 | 6.8 | 220.64 | 4.53 | 0.749 | 0.382 |
| Configuration | Params/M | GFLOPs | mAP50 | mAP50–95 |
|---|---|---|---|---|
| Baseline C3k2 | 2.59 | 6.4 | 0.670 | 0.334 |
| Orientation decomposition | 2.86 | 6.7 | 0.687 | 0.353 |
| Orientation + calibration | 2.89 | 6.8 | 0.683 | 0.347 |
| Orientation + gated fusion | 2.94 | 6.8 | 0.693 | 0.351 |
| Three-stage DA-C3k2 | 2.97 | 6.9 | 0.711 | 0.360 |
| Baseline C3k2 + final protocol | 2.59 | 6.4 | 0.733 | 0.373 |
| DA-C3k2 + final protocol | 2.96 | 6.8 | 0.749 | 0.382 |
| Class | Baseline AP | DA-C3k2 AP | Absolute gain | Relative gain |
|---|---|---|---|---|
| Alligator crack | 0.729 | 0.823 | +0.094 | +12.9% |
| Longitudinal crack | 0.741 | 0.754 | +0.013 | +1.8% |
| Oblique crack | 0.591 | 0.638 | +0.047 | +8.0% |
| Transverse crack | 0.712 | 0.755 | +0.043 | +6.0% |
| Configuration | Loss | mAP50(M) | mAP50–95(M) | Prec(M) | Recall(M) |
|---|---|---|---|---|---|
| Baseline | CE | 0.5562 | 0.1257 | 0.6997 | 0.5467 |
| Baseline | Tversky | 0.5707 | 0.1295 | 0.7377 | 0.5269 |
| Baseline+DA-C3k2 | CE | 0.5861 | 0.1400 | 0.7301 | 0.5623 |
| Baseline+DA-C3k2 | Tversky | 0.5889 | 0.1406 | 0.7352 | 0.5483 |
| Method | Params/M | Dice | IoU |
|---|---|---|---|
| DA-C3k2 lightweight model | 2.96 | 0.6931 | 0.5603 |
| U-Net ResNet34 | 24.40 | 0.7560 | 0.6079 |
| DeepLabV3+ ResNet50 | 26.70 | 0.7522 | 0.6030 |
| Configuration | Dice | IoU | BF1 | B-IoU | Latency/ms | FPS |
|---|---|---|---|---|---|---|
| Baseline+CE | 0.4899 | 0.3543 | 0.4950 | 0.3234 | 154.97 | 6.45 |
| Baseline+Tversky | 0.4966 | 0.3615 | 0.4907 | 0.3160 | 201.49 | 4.96 |
| Baseline+DA-C3k2+CE | 0.5068 | 0.3701 | 0.5100 | 0.3288 | 133.58 | 7.49 |
| Baseline+DA-C3k2+Tversky | 0.5225 | 0.3811 | 0.5182 | 0.3439 | 138.94 | 7.20 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).