Submitted:
01 August 2026
Posted:
03 August 2026
You are already at the latest version
Abstract
Keywords:
1. Introduction
Proposed Hybrid Pipeline for Railway Sleeper Inspection
2. Materials and Methods
2.1. Field Investigations and Defect Characterization: LFI Arezzo Corridor
- Construction-related conditions: Prestressing tendon misalignment and incomplete concrete compaction, leading to internal tensile concentrations and void formation.
- Wheel–rail interaction in curved track geometries: Inducing lateral loading, torsional deformation, and uneven ballast reactions that translate into repetitive tensile cycling.
- Maintenance-induced overstressing: Particularly from ballast tamping, track lifting, and rail replacement operations imposing localized point loads.
- Moisture ingress below the rail plate (piastrina): Facilitating freeze-thaw expansion and potential corrosion of embedded prestressing reinforcement.
2.2. Field Case Study
2.3. Image-Based Classification of Railway Sleeper Conditions
2.4. Object Detection
3. Results
3.1. Classification Results and Model Interpretation
3.1.1. Explainability via Grad-CAM—Image Classification
3.2. Object Detection Metrics and Performance Comparison






4. Conclusions
Funding
Data Availability Statement
Conflicts of Interest
References
- Akyon, F.C.; Altinuc, S.O.; Temizel, A. Slicing aided hyper inference and fine-tuning for small object detection. In Proceedings of the 2022 IEEE International Conference on Image Processing (ICIP), Bordeaux, France, 16–19 October 2022. [Google Scholar] [CrossRef]
- Ali, L.; Alnajjar, F.; Jassmi, H.A.; Gocho, M.; Khan, W.; Serhani, M.A. Performance evaluation of deep CNN-based crack detection and localization techniques for concrete structures. Sensors 2021, 21, 1688. [Google Scholar] [CrossRef] [PubMed]
- Apostolopoulos, I.D.; Papathanasiou, N.D.; Papandrianos, N.; Papageorgiou, E.; Apostolopoulos, D.J. Innovative attention-based explainable feature-fusion VGG19 network for characterising myocardial perfusion imaging SPECT polar maps in patients with suspected coronary artery disease. Appl. Sci. 2023, 13, 8839. [Google Scholar] [CrossRef]
- Awan, M.R.; McClory, C. Deep learning and image data-based surface cracks recognition of laser nitrided titanium alloy. Results Eng. 2024, 22, 102003. [Google Scholar] [CrossRef]
- British Standards Institution. BS EN 13230-1:2016; 2016 Railway Applications—Track—Concrete Sleepers and Bearers—Part 1: General Requirements. BSI: London, UK.
- Cha, Y.J.; Choi, W.; Büyüköztürk, O. Deep learning-based crack damage detection using convolutional neural networks. Comput.-Aided Civ. Infrastruct. Eng. 2017, 32, 361–378. [Google Scholar] [CrossRef]
- Chandramouli, A.; Song, H.; Liu, M.; Damai, A.; Narman, H.S.; Alzarrad, A. Deep learning approaches for railroad infrastructure monitoring: Comparing YOLO and vision transformers for defect detection. In In Proceedings of the 2025 IEEE 16th Annual Ubiquitous Computing, Electronics & Mobile Communication Conference (UEMCON), 2025; pp. 205–211. [Google Scholar] [CrossRef]
- Chen, J.; Liu, Z.; Wang, H.; Núñez, A.; Han, Z. Automatic defect detection of fasteners on the catenary support device using deep convolutional neural network. IEEE Trans. Instrum. Meas. 2018, 67, 257–269. [Google Scholar] [CrossRef]
- Chintalapally, A.; Bejjam, A. AI-driven railway maintenance using object detection and segmentation. In Proceedings of the SHM 2025 Conference, 2025; Available online: https://dpi-proceedings.com/index.php/shm2025/article/view/37361 (accessed on 16 July 2026).
- Dang, M.; et al. Transformer-based defect detection under complex backgrounds. Autom. Constr. Complete author list and DOI must be verified before submission. 2023, 152, 104873. [Google Scholar]
- Dietterich, T.G. Approximate statistical tests for comparing supervised classification learning algorithms. Neural Comput. 1998, 10, 1895–1923. [Google Scholar] [CrossRef] [PubMed]
- Ding, F. Crack detection in infrastructure using transfer learning, spatial attention, and genetic algorithm optimization. arXiv 2024, arXiv:2411.17140. [Google Scholar] [CrossRef]
- Dong, X.; Liu, Y.; Dai, J. Concrete surface crack detection algorithm based on improved YOLOv8. Sensors 2024, 24, 5252. [Google Scholar] [CrossRef] [PubMed]
- Dorafshan, S.; Thomas, R.J.; Maguire, M. Comparison of deep convolutional neural networks and edge detectors for crack detection. Constr. Build. Mater. 2018, 186, 1031–1045. [Google Scholar] [CrossRef]
- Esveld; 2001.
- He, K.; Zhang, X.; Ren, S.; Sun, J. Deep residual learning for image recognition. In Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27–30 June 2016; pp. 770–778. [Google Scholar] [CrossRef]
- Kaewunruen; Remennikov. 2011.
- Kandhro, I.A.; Manickam, S.; Fatima, K.; Uddin, M.; Malik, U.; Naz, A.; Dandoush, A. Performance evaluation of E-VGG19 model: Enhancing real-time skin cancer detection and classification. Heliyon 2024, 10, e31488. [Google Scholar] [CrossRef] [PubMed]
- Kaveh, H.; Alhajj, R. Recent advances in crack detection technologies for structures: A survey of 2022–2023 literature. Front. Built Environ. 2024, 10, 1321634. [Google Scholar] [CrossRef]
- Kumar, A.; Harsha, S.P. A systematic literature review of defect detection in railways using machine vision-based inspection methods. Int. J. Transp. Sci. Technol. 2024, 18, 207–206. [Google Scholar] [CrossRef]
- La Ferroviaria Italiana S.p.A. Prescrizione di Esercizio LFI n. 01/2025: Procedura d’Interfaccia. Norme per la Circolazione dei Convogli da e verso il Raccordo Baraclit nella Stazione di Bibbiena; La Ferroviaria Italiana S.p.A.: Arezzo, Italy, 14 January 2025. [Google Scholar]
- Li, S.; Zhao, X.; Zhou, G. Automatic pixel-level multiple crack detection of concrete structures. Comput.-Aided Civ. Infrastruct. Eng. 2019, 34, 616–634. [Google Scholar] [CrossRef]
- Li, D.; You, R.; Kaewunruen, S. Crack propagation assessment of time-dependent concrete degradation of prestressed concrete sleepers. Sustainability 2022, 14, 3217. [Google Scholar] [CrossRef]
- Nasimov, R.; Cho, Y.I. Smart city infrastructure monitoring with a hybrid vision transformer for micro crack detection. Sensors 2025, 25, 5079. [Google Scholar] [CrossRef] [PubMed]
- Paramanandham, N.; Koppad, D.; Anbalagan, S. Vision-based crack detection in concrete structures using cutting-edge deep learning techniques. Trait. Signal 2022, 39, 385–395. [Google Scholar] [CrossRef]
- Philip, R.E.; Andrushia, A.D.; Nammalvar, A.; Gurupatham, B.G.A.; Roy, K. A comparative study on crack detection in concrete walls using transfer learning techniques. J. Compos. Sci. 2023, 7, 169. [Google Scholar] [CrossRef]
- Regione Toscana. Atto di imposizione obbligo di servizio per la gestione dell’infrastruttura ferroviaria regionale (LFI). Available online: https://www.regione.toscana.it/-/lfi.
- Selvaraju, R.R.; Cogswell, M.; Das, A.; Vedantam, R.; Parikh, D.; Batra, D. Grad-CAM: Visual explanations from deep networks via gradient-based localization. In Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy, 22–29 October 2017; pp. 618–626. [Google Scholar] [CrossRef]
- Simonyan, K.; Zisserman, A. Very deep convolutional networks for large-scale image recognition. In Proceedings of the 3rd International Conference on Learning Representations (ICLR 2015), San Diego, CA, USA, 7–9 May 2015. [Google Scholar] [CrossRef]
- Ultralytics. Performance Metrics Deep Dive: mAP, AP50/AP75, Precision, Recall, Confusion Matrix. Available online: https://docs.ultralytics.com/guides/yolo-performance-metrics/.
- Wan, Y.; Wang, H.; Lu, L.; Zhang, Y.; Chen, M. An improved real-time detection transformer model for traffic safety facility inspection. Sustainability 2024, 16, 10172. [Google Scholar] [CrossRef]
- Wei, X.; Yang, Z.; Liu, Y.; Wei, D.; Jia, L.; Li, Y. Railway track fastener defect detection based on image processing and deep learning techniques: A comparative study. Eng. Appl. Artif. Intell. 2019, 80, 66–81. [Google Scholar] [CrossRef]
- Yang, Q.; Shi, W.; Chen, J.; Lin, W. Deep convolution neural network-based transfer learning method for civil infrastructure crack detection. Autom. Constr. 2020, 116, 103199. [Google Scholar] [CrossRef]
- Zeiler, M.D.; Fergus, R. Visualizing and understanding convolutional networks. In Computer Vision—ECCV 2014; Fleet, D., Pajdla, T., Schiele, B., Tuytelaars, T., Eds.; Springer: Cham, Switzerland, 2014; Volume 8689, pp. 818–833. [Google Scholar] [CrossRef]
- Zhang, L.; Yang, F.; Zhang, Y.D.; Zhu, Y.J. Road crack detection using deep convolutional neural network. In Proceedings of the 2016 IEEE International Conference on Image Processing (ICIP), Phoenix, AZ, USA, 25–28 September 2016; pp. 3708–3712. [Google Scholar] [CrossRef]
- Zhao, Y.; Lv, W.; Xu, S.; Wei, J.; Wang, G.; Dang, Q.; Liu, Y.; Chen, J. DETRs beat YOLOs on real-time object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 17–21 June 2024; pp. 16965–16974. [Google Scholar]





| Location | Traffic | Axle Load |
| Baraclit siding | 10 trains/year | 20 t/axle |
| Location | Traffic | Axle Load |
| Parameter | Description |
| Total images | 289 |
| In-service sleepers | 232 |
| Dump-yard sleepers | 57 |
| Imaging conditions | Natural light; ballast intact; no cleaning |
| Resolution (input) | 224×224 (classification), 640×640 (detection) |
| Annotation format | Binary labels (classification); COCO (detection) |
| Parameter | Value |
| Optimizer | Adam |
| Learning rate | 1×10−4 |
| Loss function | Binary cross-entropy |
| Batch size | 16 |
| Epochs | 30–40 |
| Train-Test split | 80% / 20% |
| Evaluation metrics | Accuracy, Precision, Recall, F1-score, Confusion matrix |
| Model | Class | Precision | Recall | F1-Score | Accuracy (%) |
| VGG16 | Railway track | 0.93 | 1.00 | 0.96 | 95.0 |
| VGG16 | Sleeper dump | 1.00 | 0.81 | 0.90 | |
| ResNet50 | Railway track | 0.87 | 0.98 | 0.92 | 88.0 |
| ResNet50 | Sleeper dump | 0.91 | 0.62 | 0.74 | |
| VGG19 Baseline | Railway track | 0.98 | 1.00 | 0.99 | 98.2 |
| VGG19 Baseline | Sleeper dump | 1.00 | 0.94 | 0.97 | |
| VGG19 Block-5 | Railway track | 0.93 | 1.00 | 0.96 | 95.0 |
| VGG19 Block-5 | Sleeper dump | 1.00 | 0.81 | 0.90 | |
| VGG19 GAP+BN | Railway track | 0.81 | 0.95 | 0.88 | 81.0 |
| VGG19 GAP+BN | Sleeper dump | 0.78 | 0.44 | 0.56 |
| Reference | Application Domain | Best Reported Model |
Performance Reported |
Agreement with Present Study |
| Awan & McClory, 2024 [4] | Surface crack detection | VGG19 > ResNet-50 | Accuracy = 99% | Confirms VGG19 = 98.2% best |
| Paramanandham et al. 2022 [25] | Concrete cracks | VGG16/VGG19 > ResNet-50 | Accuracy > 99% | Supports VGG superiority |
| Chen et al. 2018 [8] | Railway fasteners | VGG16 stable, ResNet-50 overfits | ResNet mAP < 25% | Explains low ResNet recall |
| Wei et al. 2019 [32] | Railway fasteners | VGG16 effective | Accuracy > 97% | Aligns with VGG16 = 95% |
| Ali et al. 2021 [2] | Concrete cracks | CNN/VGG > ResNet | Accuracy > 95% | Matches VGG advantage |
| Kandhro et al. 2024 [18] | Skin cancer | Baseline VGG weaker than tuned Enhanced VGG (E-VGG) | Accuracy = 88% | Explains GAP+BN degradation |
| Apostolopoulos et al. 2023 [3] | Medical imaging | Deep ResNet underperforms small data | Accuracy = 0.66–0.70 | Confirms depth penalty |
| Present study | Railway sleeper imagery | VGG19 baseline best | 98.2% accuracy | Fully validated |
| Model | Precision | Recall | F1-Score | mAP@0.5 | mAP@0.5–0.95 | Inference Speed (FPS) |
| YOLOv11 | 0.70 | 0.68 | 0.69 | 0.65 | 0.45 | ~90 |
| RT-DETR | 0.58 | 0.55 | 0.56 | 0.56 | 0.42 | ~35 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).