Submitted:
17 June 2024
Posted:
18 June 2024
You are already at the latest version
Abstract
Keywords:
1. Introduction
2. Background
2.1. Floating-Point Format
2.2. Custom Formats
2.2.1. RTL Libraries
2.2.2. HLS Libraries
3. Implementation Methodology
3.1. Development Flow
3.2. Standard Floating-Point Operations
3.3. CuFP: Detailed Implementation

3.3.1. Primary Operations


3.4. Customized Vector Operations
3.4.1. Vector Summation

3.4.2. Dot-Product Operation

3.4.3. Matrix-Vector Multiplication

3.5. CuFP Automation
4. Experimental Results
4.1. Primary Operations
4.2. Comparative Analysis of CuFP
4.3. Evaluation of Dedicated Vector Operations
4.3.1. Vector Summation
4.3.2. Dot-Product
4.3.3. Matrix-Vector Multiplication
5. Discussion
5.1. Comparison with Existing Solutions
5.2. Challenges and Limitations
6. Conclusion
References
- Uguen, Y.; Dinechin, F.D.; Lezaud, V.; Derrien, S. Application-Specific Arithmetic in High-Level Synthesis Tools. ACM Trans. Archit. Code Optim. 2020, 17. [CrossRef]
- Xilinx. UG902: Vivado Design Suite User Guide, 2021.
- Xilinx. PG060: Floating-point operator v7.1-logicore ip product guide, 2020.
- Intel. Floating-Point IP Cores User Guide, 2023.
- Cherubin, S.; Cattaneo, D.; Chiari, M.; Bello, A.D.; Agosta, G. TAFFO: Tuning Assistant for Floating to Fixed Point Optimization. IEEE Embedded Systems Letters 2020, 12, 5–8. [CrossRef]
- Agosta, G. Precision tuning of mathematically intensive programs: A comparison study between fixed point and floating point representations. Master’s thesis, School of Industrial and Information Engineering, Politecnico di Milano, 2021.
- Cattaneo, D.; Chiari, M.; Fossati, N.; Cherubin, S.; Agosta, G. Architecture-aware Precision Tuning with Multiple Number Representation Systems. 2021 58th ACM/IEEE Design Automation Conference (DAC), 2021, pp. 673–678. [CrossRef]
- Thomas, D.B. Compile-Time Generation of Custom-Precision Floating-Point IP using HLS Tools. IEEE Symposium on Computer Arithmetic (ARITH), 2019, pp. 192–193. [CrossRef]
- Langhammer, M.; VanCourt, T. FPGA Floating Point Datapath Compiler. IEEE Symposium on Field Programmable Custom Computing Machines, 2009, pp. 259–262. [CrossRef]
- Ould-Bachir, T.; David, J.P. Self-Alignment Schemes for the Implementation of Addition-Related Floating-Point Operators. ACM Trans. Reconfigurable Technol. Syst. 2013, 6. [CrossRef]
- Montaño, F.; Ould-Bachir, T.; David, J.P. A Latency-Insensitive Design Approach to Programmable FPGA-Based Real-Time Simulators. Electronics 2020, 9. [CrossRef]
- Xilinx. UG579: UltraScale Architecture DSP Slice, 2021.
- Fang, X.; Leeser, M. Open-Source Variable-Precision Floating-Point Library for Major commercial FPGAs 2016. 9. [CrossRef]
- de Dinechin, F.; Pasca, B. Designing Custom Arithmetic Data Paths with FloPoCo. IEEE Design & Test of Computers 2011, 28, 18–27. [CrossRef]
- Thomas, D.B. Templatised Soft Floating-Point for High-Level Synthesis. IEEE Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM), 2019, pp. 227–235. [CrossRef]
- de Dinechin, Florent, K.M. Application-Specific Arithmetic; Springer International Publishing, 2023.
- Böttcher, A.; Kumm, M.; de Dinechin, F. Resource Optimal Truncated Multipliers for FPGAs. 2021 IEEE 28th Symposium on Computer Arithmetic (ARITH), 2021, pp. 102–109. [CrossRef]
- Böttcher, A.; Kumm, M. Towards Globally Optimal Design of Multipliers for FPGAs. IEEE Transactions on Computers 2023, 72, 1261–1273. [CrossRef]
- Fiorito, M.; Curzel, S.; Ferrandi, F. TrueFloat: A Templatized Arithmetic Library for HLS Floating-Point Operators. Embedded Computer Systems: Architectures, Modeling, and Simulation; Silvano, C.; Pilato, C.; Reichenbach, M., Eds., 2023, pp. 486–493.
- Perera, A.; Nilsen, R.; Haugan, T.; Ljokelsoy, K. A Design Method of an Embedded Real-Time Simulator for Electric Drives using Low-Cost System-on-Chip Platform. PCIM Europe digital days 2021; International Exhibition and Conference for Power Electronics, Intelligent Motion, Renewable Energy and Energy Management, 2021, pp. 1–8.
- Zamiri, E.; Sanchez, A.; Yushkova, M.; Martínez-García, M.S.; de Castro, A. Comparison of Different Design Alternatives for Hardware-in-the-Loop of Power Converters. Electronics 2021. [CrossRef]
- Hajizadeh, F.; Alavoine, L.; Ould-Bachir, T.; Sirois, F.; David, J.P. FPGA-Based FDNE Models for the Accurate Real-Time Simulation of Power Systems in Aircrafts. 2023 12th International Conference on Renewable Energy Research and Applications (ICRERA), 2023, pp. 344–348. [CrossRef]
- IEEE Standard for Floating-Point Arithmetic. IEEE Std 754-2008 2008, pp. 1–70.
- Sanchez, A.; Todorovich, E.; De Castro, A. Exploring the Limits of Floating-Point Resolution for Hardware-In-the-Loop Implemented with FPGAs. Electronics 2018, 7. [CrossRef]
- Martínez-García, M.S.; de Castro, A.; Sanchez, A.; Garrido, J. Analysis of Resolution in Feedback Signals for Hardware-in-the-Loop Models of Power Converters. Electronics 2019, 8. [CrossRef]
- Wang, X.; Leeser, M. VFloat: A Variable Precision Fixed- and Floating-Point Library for Reconfigurable Hardware. ACM Trans. Reconfigurable Technol. Syst. 2010, 3. [CrossRef]
- de Dinechin, F. Reflections on 10 years of FloPoCo. ARITH 2019 - 26th IEEE Symposium on Computer Arithmetic; , 2019; pp. 1–3. [CrossRef]
- Bansal, S.; Hsiao, H.; Czajkowski, T.; Anderson, J.H. High-level synthesis of software-customizable floating-point cores. 2018 Design, Automation & Test in Europe Conference & Exhibition (DATE), 2018, pp. 37–42. [CrossRef]
- Uguen, Y.; de Dinechin, F.; Derrien, S. Bridging high-level synthesis and application-specific arithmetic: The case study of floating-point summations. 2017 27th International Conference on Field Programmable Logic and Applications (FPL), 2017, pp. 1–8. [CrossRef]
- Ferrandi, F.; Castellana, V.G.; Curzel, S.; Fezzardi, P.; Fiorito, M.; Lattuada, M.; Minutoli, M.; Pilato, C.; Tumeo, A. Invited: Bambu: An Open-Source Research Framework for the High-Level Synthesis of Complex Applications. ACM/IEEE Design Automation Conference (DAC), 2021, pp. 1327–1330. [CrossRef]
- Parhami, B. Computer Arithmetic: Algorithms and Hardware Designs; Oxford series in electrical and computer engineering, Oxford University Press, 2010.
- Xilinx. UG1399: Vitis High-Level Synthesis User Guide, 2023.








| Format | Word Size | Exponent Width | Fraction Width | Bias |
|---|---|---|---|---|
| (Wfp) | (We) | (Wf) | ||
| Half-Precision | 16 | 5 | 10 | 15 |
| Single-Precision | 32 | 8 | 23 | 127 |
| Double-Precision | 64 | 11 | 52 | 1,023 |
| Quadruple-Precision | 128 | 15 | 112 | 16,383 |
| Operation | Variant | # of Stage | Max Frequency (MHz) |
Resource Utilization | ||
|---|---|---|---|---|---|---|
| DSP | LUT | FF | ||||
| CuFP (r) | 6 | 420 | 0 | 486 | 414 | |
| Sum | CuFP (t) | 6 | 436 | 0 | 410 | 314 |
| Vendor | 6 | 375 | 2 | 219 | 277 | |
| Flopoco | 6 | 332 | 0 | 278 | 331 | |
| CuFP (r) | 3 | 436 | 2 | 41 | 135 | |
| Mul | CuFP (t) | 3 | 468 | 2 | 22 | 118 |
| Vendor | 3 | 284 | 3 | 71 | 133 | |
| Flopoco | 3 | 338 | 2 | 71 | 98 | |
| Variant | Vector Size | # of Stages | Resource Utilization | ||
|---|---|---|---|---|---|
| DSP | LUT | FF | |||
| CuFP | 4 | 4 | 0 | 1,071 | 441 |
| 8 | 5 | 0 | 2,110 | 685 | |
| 16 | 5 | 0 | 3,675 | 1,277 | |
| 32 | 6 | 0 | 7,420 | 3,050 | |
| Vendor IP | 4 | 8 | 6 | 701 | 709 |
| 8 | 12 | 14 | 1,642 | 1,645 | |
| 16 | 16 | 30 | 3,522 | 3,513 | |
| 32 | 20 | 62 | 7,283 | 7,245 | |
| Variant | Vector Size | # of Stages | Resource Utilization | ||
|---|---|---|---|---|---|
| DSP | LUT | FF | |||
| CuFP | 4 | 5 | 8 | 1,427 | 562 |
| 8 | 5 | 16 | 2,629 | 1,085 | |
| 16 | 6 | 32 | 4,736 | 2,350 | |
| 32 | 7 | 64 | 9,117 | 4,831 | |
| Vendor IP | 4 | 10 | 18 | 1,057 | 1,095 |
| 8 | 14 | 38 | 2,354 | 2,415 | |
| 16 | 18 | 78 | 4,947 | 5,051 | |
| 32 | 22 | 158 | 10,138 | 10,319 | |
| Variant | Vector Size | # of Stages | Resource Utilization | ||
|---|---|---|---|---|---|
| DSP | LUT | FF | |||
| CuFP | 4 | 5 | 32 | 5,230 | 1,831 |
| 8 | 5 | 128 | 20,393 | 6,759 | |
| 16 | 6 | 512 | 80,974 | 29,320 | |
| 32 | 7 | 2,048 | 294,019 | 120,679 | |
| Vendor IP | 4 | 10 | 72 | 4,240 | 3,960 |
| 8 | 14 | 304 | 18,859 | 17,416 | |
| 16 | 18 | 1,248 | 79,188 | 72,836 | |
| 32 | 22 | 5,056 | 324,261 | 297,720 | |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2024 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).