Submitted:
28 July 2025
Posted:
11 August 2025
You are already at the latest version
Abstract
Keywords:
1. Introduction
2. Materials and Methods
2.1. Literature Review
2.2. Dataset
2.2.1. ISO 6346 Standard
- Owner code: Consists of 3 letters.
-
Category identifier: Designates the type of container, most common ones are:
- -
- J: Detachable freight container equipment
- -
- U: General purpose container
- -
- Z: Trailers and chassis
- Serial number: Consists of 6 digits, being the unique identifier of the container within that operator’s fleet.
- Check digit: A single digit, calculated mathematically in order to verify the integrity of the code.
- 1st character: Overall length.
- 2nd character: Overall height.
- 3rd–4th characters: Type and additional features (e.g.,General-G0 to G3; Tank-T0 to T9).
2.2.2. Dataset Description
2.3. Hybrid Pipeline Architecture
2.3.1. Detection Stage
2.3.2. Recognition Stage
2.3.3. Training Protocol
Hardware and Frameworks
- GPU: NVIDIA GeForce RTX 1080 Ti (11GB VRAM)
-
TrOCR Framework: Hugging Face Transformers v4.46.3
- -
- Base Model: trocr-base-stage1 from Hugging Face
- -
- WandB Integration: Full training metrics logging
-
YOLOv7 Framework: Custom PyTorch implementation
- -
- Input Resolution: 512x512 pixels
- -
- Augmentations: Mosaic (1.0), MixUp (0.15), FlipLR (0.5)
- -
- WandB Integration: Full training metrics logging
Hyperparameter Configuration
TrOCR Specifications
-
Architecture:
- -
- Vision Encoder-Decoder with 384M parameters
- -
- ViT encoder processes 384×384 images-16px patches
- -
- Transformer decoder: 1024 dim, 16 attention heads
- -
- Maximum sequence length: 64 tokens
- -
- Gradient clipping at 1.0 norm
- -
- Max Grad Norm: 1.0
-
Training Configuration:
- -
- Mixed FP16 precision (O1)
- -
- Gradient checkpointing enabled
- -
- Early stopping on validation CER
YOLOv7 Specifications
-
Architecture:
- -
- Darknet-based with 36.5M parameters
- -
-
Loss components:
- *
- Bounding box: 0.05 (CIoU)
- *
- Classification: 0.3
- *
- Objectness: 0.7
- -
- Optimal Transport Assignment (OTA) enabled
-
Training Configuration:
- -
- FP32 precision
-
Training Variants:
- -
- Frozen: First 50 layers fixed
- -
- Unfrozen: Entire network trainable
3. Results and Discussion
3.1. Evaluation Metrics
3.2. YOLOv7 Training
3.3. TrOCR Training
3.4. Comparative Results
4. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
Abbreviations
| AdamW | Adaptive Moment Estimation with Weight Decay |
| CER | Character Error Rate |
| CIoU | Complete Intersection over Union |
| EMA | Exponential Moving Average |
| FP16 | 16-bit Floating Point Precision |
| IoU | Intersection over Union |
| mAP | mean Average Precision |
| OCR | Optical Character Recognition |
| OTA | Optimal Transport Assignment |
| SGD | Stochastic Gradient Descent |
| Seq2Seq | Sequence-to-Sequence |
| TrOCR | Transformer-based Optical Character Recognition |
| ViT | Vision Transformer |
| WandB | Weights and Biases |
| WER | Word Error Rate |
| YOLOv7 | You Only Look Once version 7 |
References
- United Nations Conference on Trade and Development. Review of maritime transport 2024: Navigating maritime chokepoints (UNCTAD/RMT/2024 and Corr.1). United Nations, 2024. Available online: https://unctad.org/publication/review-maritime-transport-2024.
- de la Peña Zarzuelo, I.; Freire Soeane, M.J.; López Bermúdez, B. Industry 4.0 in the port and maritime industry: A literature review. Journal of Industrial Information Integration 2020, 20, 100173. [CrossRef]
- Yang, Y.; Gai, T.; Cao, M.; Zhang, Z.; Zhang, H.; Wu, J. Application of Group Decision Making in Shipping Industry 4.0: Bibliometric Analysis, Trends, and Future Directions. Systems 2023, 11, 69. https://www.mdpi.com/2079-8954/11/2/69/htm.
- Filom, S.; Amiri, A.M.; Razavi, S. Applications of machine learning methods in port operations—A systematic literature review. Transportation Research Part E: Logistics and Transportation Review 2022, 161, 102722. [CrossRef]
- Port Equipment Manufacturers Association. Information paper: OCR in ports and terminals (PEMA-IP4). Available online: https://www.pema.org/wp-content/uploads/2022/09/PEMA-IP4-OCR-in-Ports-and-Terminals.pdf (accessed on 2025).
- Smith, R. An Overview of the Tesseract OCR Engine. In Proceedings of the Ninth International Conference on Document Analysis and Recognition (ICDAR 2007), Curitiba, Brazil, 23–26 September 2007; Volume 2, pp. 629-633. [CrossRef]
- Vedhaviyassh, D.R.; Sudhan, R.; Saranya, G.; Safa, M.; Arun, D. Comparative Analysis of EasyOCR and TesseractOCR for Automatic License Plate Recognition using Deep Learning Algorithm. 6th International Conference on Electronics, Communication and Aerospace Technology, ICECA 2022 - Proceedings 2022, pp. 966-971. [CrossRef]
- Raj, R.; Kos, A. A Comprehensive Study of Optical Character Recognition. In Proceedings of the 2022 29th International Conference on Mixed Design of Integrated Circuits and System (MIXDES), 2022, pp. 151-154. [CrossRef]
- Baek, Y.; Lee, B.; Han, D.; Yun, S.; Lee, H. Character Region Awareness for Text Detection. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition 2019, 2019-June, 9357–9366. arXiv:1904.01941v1.
- Wang, C.-Y.; Bochkovskiy, A.; Liao, H.-Y.M. YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 7464-7475.
- Baek, J.; Kim, G.; Lee, J.; Park, S.; Han, D.; Yun, S.; Oh, S.J.; Lee, H. What Is Wrong With Scene Text Recognition Model Comparisons? Dataset and Model Analysis. Proceedings of the IEEE International Conference on Computer Vision 2019, 2019-October, 4714-4722. arXiv:1904.01906v4.
- Feng, X.; Wang, Z.; Liu, T. Port container number recognition system based on improved YOLO and CRNN Algorithm. Proceedings - International Conference on Artificial Intelligence and Electromechanical Automation, AIEA 2020 2020, 72-77. [CrossRef]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention Is All You Need. Advances in Neural Information Processing Systems 2017, 2017-December, 5999–6009. arXiv:1706.03762v7.
- Lin, T.; Wang, Y.; Liu, X.; Qiu, X. A Survey of Transformers. AI Open 2022, 3, 111-132. [CrossRef]
- Li, M.; Lv, T.; Chen, J.; Cui, L.; Lu, Y.; Florencio, D.; Zhang, C.; Li, Z.; Wei, F. TrOCR: Transformer-based Optical Character Recognition with Pre-trained Models. Proceedings of the 37th AAAI Conference on Artificial Intelligence, AAAI 2023 2021, 37, 13094–13102. arXiv:2109.10282v5.
- ISO. ISO 6346:2022 - Freight containers — Coding, identification and marking. Available online: https://www.iso.org/standard/83558.html (accessed on 2025).
- Bureau International des Containers. BIC Code - The Standard for Container Identification. Available online: https://www.bic-code.org/ (accessed on may 2025).
- Lin, B. ContainerNumber-OCR: Container Number Recognition Based on YOLOv7 and CRNN. GitHub repository, 2023. Available online: https://github.com/lbf4616/ContainerNumber-OCR (accessed on February 2025).
- Dutta, A.; Zisserman, A. VGG Image Annotator (VIA). Version 2.0.12. Available online: http://www.robots.ox.ac.uk/~vgg/software/via/ (accessed on 30 May 2025).
- Hugging Face. TrOCR Model Documentation. 2024. Available online: https://huggingface.co/docs/transformers/model_doc/trocr (accessed on March 2025).





| Parameter | TrOCR | YOLOv7 (Frozen) | YOLOv7 (Unfrozen) |
|---|---|---|---|
| Batch Size | 8 | 16 | 16 |
| Learning Rate | 5e-5 | 0.01 | 0.01 |
| LR Schedule | Linear Warmup | Cosine Annealing | Cosine Annealing |
| Warmup Steps | 500 steps | 3 epochs | 3 epochs |
| Epochs | 25 | 50 | 50 |
| Optimizer | AdamW | SGD | SGD |
| Momentum | - | 0.937 | 0.937 |
| Weight Decay | 0 | 0.0005 | 0.0005 |
| Model | Precision (P) | Recall (R) | mAP@0.5 | mAP@0.5:0.95 |
|---|---|---|---|---|
| YOLOv7 (Frozen Backbone) | 94.8 | 95.4 | 96.9 | 65.3 |
| YOLOv7 (Original) | 91.1 | 79.7 | 84.0 | 56.3 |
| Model | CER | WER | Character Accuracy | Exact Match Score |
|---|---|---|---|---|
| TrOCR | 0.0244 | 0.1066 | 97.56% | 89.34% |
| Detection Recall | Detection Precision | mAP@50 | mAP@95 | Vertical Recognition Accccuracy | Horizontal Recognition Accccuracy | |
|---|---|---|---|---|---|---|
| Yolov7+TrOCR | 96.77% | 99.40% | 96.69% | 70.84% | 99.11% | 96.86% |
| Yolov7+Tesseract | 87.66% | 99.34% | 87.57% | 64.41% | 2.58% | 73.25% |
| Yolov7+EasyOCR | 89.48% | 99.54% | 89.48% | 65.84% | 15.31% | 64.23% |
| Tesseract | 11.76% | 0.41% | 2.15% | 0.77% | 2.04% | 33.48% |
| EasyOCR | 88.73% | 16.47% | 27.09% | 10.02% | 13.85% | 71.17% |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).