Submitted:
05 August 2026
Posted:
05 August 2026
You are already at the latest version
Abstract
Keywords:
1. Introduction
2. Materials and Methods
2.1. Hardware Platform

2.2. Host Development Environment

2.3. ZCU104 FPGA Platform Setup
2.4. Project Workspace Preparation
2.5. Deployment Workflow for the Optimised RAAGR2-Net Model
- 1.
- Model preparation The trained RAAGR2-Net segmentation model was prepared on the host development environment installed on the virtual machine. The model weights obtained from the training stage were put together with the required preprocessing scripts and configuration files so that they could be processed by the Vitis AI 3.0 toolchain.
- 2.
- Calibration dataset preparation A representative subset of MRI images from the BraTS dataset was selected to serve as calibration data. This dataset is required during quantisation, the end result of quantization is converting floating-point operations to lower precision representations.
- 3.
- Model quantisation The initial trained model based on floating-point values, was processed using the Vitis AI quantisation tools to convert the network into an INT8 representation suitable for FPGA execution. This in effect reduces computational cost and memory requirements while maintaining acceptable segmentation accuracy. .
- 4.
- Model compilation After quantisation, the model was compiled using the Vitis AI compiler. The compiler transforms the quantised model into an executable representation compatible with the Deep Processing Unit (DPU) architecture available on the ZCU104 FPGA platform.
- 5.
- Deployment on FPGA The compiled model is transferred to the ZCU104 board where inference will be executed using the Vitis AI 3.0 runtime environment. This stage allows the evaluation of inference latency, throughput, and computational efficiency of the optimised RAAGR2-Net model on edge hardware.
3. Results
3.1. Segmentation Performance Before Deployment
3.1.1. FPGA Inference Results
3.2. Deployability of the Pruned Model
4. Discussion
5. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
Abbreviations
| MRI | Magnetic resonance imaging |
| WT | whole tumour |
| ET | Enhancing Tumour |
| TC | Tumour core |
| BRATS | Brain Tumour Segmentation |
| FPGA | Field-programmable gate arrays |
| DPU | Deep Processing Unit |
| FPS | Frames per second |
References
- Xilinx Inc. ZCU104 Evaluation Board User Guide. UG1267, v1.1, October 9, 2018. Available online: https://www.xilinx.com/support/documentation/boards_and_kits/zcu104/ug1267-zcu104-eval-bd.pdf (accessed on 12 March 2026).
- AMD. Vitis AI User Guide: Runtime (VART). UG1414, 2023. Available online: https://docs.amd.com/r/3.0-English/ug1414-vitis-ai/Vitis-AI-Runtime (accessed on 12 March 2026).
- AMD. Setting Up the ZCU102/ZCU104/KV260/VCK190 Evaluation Board. Vitis AI User Guide (UG1414), 2023. Available online: https://docs.amd.com/r/3.0-English/ug1414-vitis-ai/Setting-Up-the-ZCU102/ZCU104/KV260/VCK190-Evaluation-Board (accessed on 12 March 2026).
- Menze, B.H.; Jakab, A.; Bauer, S.; Kalpathy-Cramer, J.; Farahani, K.; Kirby, J.; Burren, Y.; Porz, N.; Slotboom, J.; Wiest, R.; et al. The Multimodal Brain Tumor Image Segmentation Benchmark (BRATS). IEEE Trans. Med. Imaging 2015, 34(10), 1993–2024. [Google Scholar] [CrossRef] [PubMed]
- Bakas, S.; Reyes, M.; Jakab, A.; Bauer, S.; Rempfler, M.; Crimi, A.; Shinohara, R.T.; Berger, C.; Ha, S.M.; Rozycki, M.; et al. Identifying the Best Machine Learning Algorithms for Brain Tumor Segmentation. Neuro-Oncology 2018, 20(10), 1399–1410. [Google Scholar]
- Ronneberger, O.; Fischer, P.; Brox, T. U-Net: Convolutional Networks for Biomedical Image Segmentation. In Proceedings of MICCAI, 2015, pp. 234–241.
- Rehman, A.; Khan, S.; et al. RAAGR2-Net: A Deep Learning Model for Brain Tumor Segmentation. IEEE Access 2021. [Google Scholar] [CrossRef] [PubMed]
- Venieris, S.I.; Bouganis, C.-S. Toolflows for Mapping Convolutional Neural Networks on FPGAs. ACM Comput. Surv. 2018, 51(3). [Google Scholar] [CrossRef]
- Han, S.; Mao, H.; Dally, W.J. Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding. In Proceedings of ICLR, 2016.
- Molchanov, P.; Tyree, S.; Karras, T.; Aila, T.; Kautz, J. Pruning Convolutional Neural Networks for Resource Efficient Inference. In Proceedings of ICLR, 2017.
- AMD Xilinx. Vitis AI User Guide. AMD Xilinx, 2023. Available online: https://xilinx.github.io/Vitis-AI/ (accessed on 12 March 2026).
- AMD Xilinx. Vitis AI Docker Development Environment Documentation. AMD Xilinx, 2022. Available online: https://xilinx.github.io/Vitis-AI/docs/ (accessed on March 2026).
- AMD. Vitis AI User Guide (UG1414): Deep Learning Processor Unit (DPU). AMD, 2023. Available online: https://docs.amd.com/r/3.0-English/ug1414-vitis-ai/Deep-Learning-Processor-Unit (accessed on 12 March 2026).

| Metric | Base Model | Pruned Model | Change |
|---|---|---|---|
| Loss | 0.0573 | 0.0637 | +0.0064 |
| Mean IoU | 0.7900 | 0.8218 | +0.0318 |
| Dice | 0.9858 | 0.9883 | +0.0025 |
| TC Dice | 0.8169 | 0.8578 | +0.0409 |
| ET Dice | 0.7896 | 0.8131 | +0.0235 |
| WT Dice | 0.8424 | 0.8635 | +0.0211 |
| Metric | Original | Pruned |
|---|---|---|
| Number of samples | 131 | 131 |
| Average latency | 63.08 ms | 66.01 ms |
| Std latency | 29.48 ms | 18.67 ms |
| Min latency | 50.40 ms | 47.49 ms |
| Max latency | 393.66 ms | 165.83 ms |
| Throughput | 15.85 FPS | 15.15 FPS |
| Total inference time | 8.26 s | 8.65 s |
| Metric | Original | Pruned |
|---|---|---|
| Number of samples | 131 | 131 |
| Average latency | 75.28 ms | 67.58 ms |
| Std latency | 84.79 ms | 9.24 ms |
| Min latency | 56.13 ms | 54.71 ms |
| Max latency | 1040.28 ms | 100.91 ms |
| Throughput | 13.28 FPS | 14.80 FPS |
| Total inference time | 9.86 s | 8.85 s |
| Metric | Value |
|---|---|
| Number of samples | 131 |
| Average latency | 69.16 ms |
| Std latency | 0.08 ms |
| Min latency | 68.96 ms |
| Max latency | 69.57 ms |
| Throughput | 14.46 FPS |
| Total inference time | 9.06 s |
| Platform | Model | Avg Latency (ms) | Std (ms) | Min (ms) | Max (ms) | FPS | Total Inference Time (s) |
|---|---|---|---|---|---|---|---|
| CPU | Original | 63.08 | 29.48 | 50.40 | 393.66 | 15.85 | 8.26 |
| CPU | Pruned | 66.01 | 18.67 | 47.49 | 165.83 | 15.15 | 8.65 |
| GPU | Original | 75.28 | 84.79 | 56.13 | 1040.28 | 13.28 | 9.86 |
| GPU | Pruned | 67.58 | 9.24 | 54.71 | 100.91 | 14.80 | 8.85 |
| ZCU104 FPGA | Pruned | 69.16 | 0.08 | 68.96 | 69.57 | 14.46 | 9.06 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).