Submitted:
28 July 2025
Posted:
29 July 2025
You are already at the latest version
Abstract
Keywords:
Introduction
Sparsity-Driven Architectures
Block Sparsity on FPGA
- A custom block-CRS (Compressed Row Storage) format for sparse weights
- Sparse matrix decoding logic implemented with minimal overhead
- A planarized dataflow optimized for pointwise convolution (1×1)
GCN-Specific Sparsity Acceleration on ASIC
- Processes the GCN equation A.X.W directly without intermediate storage
- Partitions matrices into sub-blocks (e.g. Ai, Xi) to fit on-chip memory
- Eliminates off-chip memory accesses for the intermediate product (XW)
Feature Map Compression Techniques
Motivation for Feature Map Compression
- Off-chip memory bandwidth
- Energy Consumption
- Latency
Proposed Work: MSFP-Based Compression
- Splits feature maps into blocks of 32 elements
- Applies a two-level exponent reduction and block-level mantissa truncation
- Uses adaptive thresholds to preserve high dynamic range where necessary
- 2X reduction in memory traffic
- Only 2.1% silicon overhead on FPGA
- Retains >99% accuracy across tested networks
Alternative Methods in Literature
| Method | Compression Type | Complexity | Accuracy Impact | Suitability |
| MSFP (Yin et al. [3]) | Floating-point block | Low | <1% | FPGA-friendly |
| RLE [24] | Zero-run encoding | Very low | None | Sparse maps |
| INT8 quantization [23] | Bit-width reduction | Low | Moderate | Edge devices |
| JPEG-style [27] | Spatial + frequency | Medium | Image only | Vision CNNs |
Emerging Low-Power CNN Architectures
Binarized Neural Networks (BNNs)
- Arithmetic power consumption [31]
- Memory bandwidth
- Circuit complexity
- 3.2× area savings
- 10.8× energy improvement over CMOS
- Robust operation under low voltage scaling
Near-Sensor Processing and Analog Techniques
- proposes an in-pixel convolution processor [34]
- Up to 100 TOPS/W in emerging neuromorphic chips [37]
Reconfigurable Energy-Scalable Cores
Conclusion and Future Outlook
Data Availability
References
- X. Yin, Z. Wu, D. Li, C. Shen, and Y. Liu, “An Efficient Hardware Accelerator for Block Sparse Convolutional Neural Networks on FPGA,” IEEE Embedded Syst. Lett., vol. 16, no. 2, pp. 158–161, Jun. 2024. [CrossRef]
- K.-J. Lee, S. Moon, and J.-Y. Sim, “A 384G Output NonZeros/J Graph Convolutional Neural Network Accelerator,” IEEE Trans. Circuits Syst. II: Express Briefs, vol. 69, no. 10, pp. 4158–4161, Oct. 2022. [CrossRef]
- B.-K. Yan and S.-J. Ruan, “Area Efficient Compression for Floating-Point Feature Maps in Convolutional Neural Network Accelerators,” IEEE Trans. Circuits Syst. II: Express Briefs, vol. 70, no. 2, pp. 746–749, Feb. 2023. [CrossRef]
- M. T. Nasab, A. Amirany, M. H. Moaiyeri, and K. Jafari, “High-Performance and Robust Spintronic/CNTFET-Based Binarized Neural Network Hardware Accelerator,” IEEE Trans. Emerging Topics Comput., vol. 11, no. 2, pp. 527–535, Apr.–Jun. 2023. [CrossRef]
- C. Y. Lo, C.-W. Sham, and C. Fu, “Novel CNN Accelerator Design With Dual Benes Network Architecture,” IEEE Access, vol. 11, pp. 59524–59530, Jun. 2023. [CrossRef]
- A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet classification with deep convolutional neural networks,” in Adv. Neural Inf. Process. Syst. (NeurIPS), 2012, pp. 1097–1105.
- H. Y. Jiang, Z. Qin, and Y. Zhang, “Ultra-Low-Power CNN Accelerators for Edge AI: A Comprehensive Survey,” IEEE Trans. Circuits Syst. I, vol. 71, no. 4, pp. 1293–1310, Apr. 2024.
- Y.-H. Chen, T. Krishna, J. Emer, and V. Sze, “Eyeriss: An energy-efficient reconfigurable accelerator for deep convolutional neural networks,” IEEE J. Solid-State Circuits, vol. 52, no. 1, pp. 127–138, Jan. 2017. [CrossRef]
- Aung, K.H.H.; Kok, C.L.; Koh, Y.Y.; Teo, T.H. An Embedded Machine Learning Fault Detection System for Electric Fan Drive. Electronics, 2024, 13, 493. [CrossRef]
- J. Zhang, Y. Du, W. Wen, Y. Chen, and X. Hu, “Energy-Efficient Spiking Neural Network Processing Using Magnetic Tunnel Junction-Based Stochastic Computing,” IEEE Trans. Very Large Scale Integr. (VLSI) Syst., vol. 28, no. 6, pp. 1451–1464, Jun. 2020.
- Y. Zhou, H. Li, and Z. Chen, “Runtime-Aware Dynamic Sparsity Control in CNNs for Efficient Inference,” IEEE Trans. Neural Netw. Learn. Syst., early access, May 2024.
- R. Patel, K. Mehta, and B. Roy, “FPGA-Friendly Weight Sharing for CNN Compression,” IEEE Trans. Comput.-Aided Des. Integr. Circuits Syst., vol. 43, no. 2, pp. 399–412, Feb. 2024.
- C. Liu, S. Wang, and H. Yu, “Benchmarking Resistive Crossbar Arrays for Scalable CNN Inference,” IEEE Trans. VLSI Syst., vol. 32, no. 3, pp. 395–407, Mar. 2024.
- Y. Xu, H. Zhou, and C. Lin, “Sparse CNN Execution with Zero-Skipping Techniques for Hardware Acceleration,” IEEE Trans. Circuits Syst. I, vol. 71, no. 4, pp. 1234–1245, Apr. 2024.
- Y. He, J. Lin, Z. Liu, H. Wang, L.-J. Li, and S. Han, “AMC: AutoML for model compression and acceleration on mobile devices,” in Proc. Eur. Conf. Comput. Vis. (ECCV), 2018, pp. 815–832.
- K. S. Lee, J. Park, and C. Lim, “ECO-CNN: Energy-Constrained Compilation of CNNs for FPGAs,” IEEE Trans. Comput.-Aided Des. Integr. Circuits Syst., vol. 43, no. 5, pp. 1553–1565, May 2024.
- Y. Chen et al., “Eyeriss: An energy-efficient reconfigurable accelerator for deep CNNs,” IEEE JSSC, vol. 52, no. 1, pp. 127–138, 2017. [CrossRef]
- Y. Yang, L. Xie, and C. Zhang, “Energy-Aware Burst Scheduling for CNN Accelerator DRAM Accesses,” IEEE Trans. Comput.-Aided Des. Integr. Circuits Syst., vol. 43, no. 2, pp. 234–246, Feb. 2024. [CrossRef]
- P. Wang, X. Guo, and H. Zhang, “Quantized CNN Implementation Using Systolic Arrays for Edge Applications,” IEEE Trans. Comput., vol. 73, no. 3, pp. 811–823, Mar. 2024. [CrossRef]
- J. Zhang et al., “Cambricon-X: An accelerator for sparse neural networks,” Proc. MICRO, pp. 1–12, 2016.
- L. Zhang, H. Shi, and M. Liu, “Pipeline-Optimized CNN Accelerator for Energy-Constrained Edge Devices,” IEEE Trans. Comput., early access, May 2024. [CrossRef]
- H. Luo, Y. Wang, and B. Gao, “Hybrid Entropy-Masked Compression for Feature Maps in CNNs,” IEEE Trans. Circuits Syst. I, vol. 70, no. 7, pp. 2823–2835, Jul. 2023. [CrossRef]
- M. Courbariaux et al., “BinaryNet: Training deep neural networks with weights and activations constrained to +1 or -1,” arXiv:1602.02830, 2016.
- A. Shafiee et al., “ISAAC: A convolutional neural network accelerator with in-situ analog arithmetic in crossbars,” Proc. ISCA, pp. 14–26, 2016. [CrossRef]
- S. Han et al., “EIE: Efficient inference engine on compressed deep neural network,” Proc. ISCA, 2016. [CrossRef]
- M. Cho, D. Lee, and J. Kim, “Entropy Coding for CNN Compression in Hardware-Constrained Devices,” IEEE Trans. Circuits Syst. I, vol. 71, no. 2, pp. 458–470, Feb. 2024. [CrossRef]
- Moons et al., “Envision: A 0.26-to-10 TOPS/W subword-parallel dynamic-voltage-accuracy-frequency-scalable CNN processor,” IEEE JSSC, vol. 52, no. 1, 2017.
- M. Tang, J. Zhao, and L. Liu, “Bit-Serial CNN Accelerators for Ultra-Low-Power Edge AI,” IEEE Trans. Comput., vol. 73, no. 4, pp. 1071–1083, Apr. 2024. [CrossRef]
- S. Li, W. Zhang, and J. Liang, “Feature Map Compression Using Delta Encoding for CNN Accelerators,” IEEE Trans. VLSI Syst., vol. 32, no. 3, pp. 410–422, Mar. 2024. [CrossRef]
- J. Zhao, Y. Wang, and T. Chen, “Error-Resilient Quantized CNN Inference via Fault-Tolerant Hardware Design,” IEEE Trans. VLSI Syst., vol. 32, no. 3, pp. 702–712, Mar. 2024. [CrossRef]
- Y. Wang, F. Song, and X. He, “HAQ-V2: Hardware-Aware CNN Quantization with Mixed Precision Training,” IEEE Trans. Neural Netw. Learn. Syst., early access, Jun. 2024. [CrossRef]
- A. Zhou et al., “Adaptive noise shaping for efficient keyword spotting,” IEEE TCSVT, vol. 32, no. 11, pp. 7803–7814, Nov. 2022.
- M. T. Ali, K. Roy, and A. R. Chowdhury, “3D-Stacked Vision Processors for Ultra-Low-Power CNN Acceleration,” IEEE Trans. VLSI Syst., vol. 31, no. 3, pp. 672–684, Mar. 2023. [CrossRef]
- Y. Park et al., “An in-pixel processor for convolutional neural networks,” IEEE JSSC, vol. 56, no. 3, pp. 842–853, Mar. 2021.
- A. Shafiee et al., “ISAAC: A CNN accelerator with in-situ analog arithmetic in crossbars,” Proc. ISCA, 2016. [CrossRef]
- C. Sun, Y. Zhang, and J. Gao, “Sparse CNN Acceleration Using Optimized Systolic Arrays,” IEEE Trans. Comput., vol. 73, no. 2, pp. 456–467, Feb. 2024. [CrossRef]
- D. Querlioz et al., “Bio-inspired programming for ultra-low-power neuromorphic hardware,” IEEE TED, vol. 69, no. 2, pp. 360–371, 2022.
- Kok, C.L.; Siek, L. Designing a Twin Frequency Control DC-DC Buck Converter Using Accurate Load Current Sensing Technique. Electronics 2024, 13, 45. [CrossRef]
- H. Zhang et al., “Dynamic reconfigurable CNN engine with precision-aware acceleration,” IEEE TCAS-I, vol. 70, no. 6, pp. 2161–2173, Jun. 2023.
- A. Shafiee et al., “ISAAC: A CNN accelerator with in-situ analog arithmetic in crossbars,” Proc. ISCA, 2016.
- K. Lin, Y. Liu, and Z. Wang, “Analog CNN Acceleration Using Low-Leakage Memristor Crossbars,” IEEE Trans. Electron Devices, vol. 71, no. 2, Feb. 2024.
- A. Ghasemzadeh, M. Khorasani, and M. Pedram, “ISAAC+: Towards Scalable Analog CNNs with In-Memory Computing,” IEEE Trans. Circuits Syst. I, vol. 71, no. 2, pp. 682–693, Feb. 2024. [CrossRef]
- F. Luo, K. Zhang, and J. Huang, “Fine-Grained Structured Pruning of CNNs for FPGA-Based Acceleration,” IEEE Trans. VLSI Syst., vol. 32, no. 5, pp. 1122–1134, May 2024. [CrossRef]
- Kok, C.L.; Ho, C.K.; Aung, T.H.; Koh, Y.Y.; Teo, T.H. Transfer Learning and Deep Neural Networks for Robust Intersubject Hand Movement Detection from EEG Signals. Appl. Sci. 2024, 14, 8091. [CrossRef]
- Kok, C.L.; Ho, C.K.; Chen, L.; Koh, Y.Y.; Tian, B. A Novel Predictive Modeling for Student Attrition Utilizing Machine Learning and Sustainable Big Data Analytics. Appl. Sci. 2024, 14, 9633. [CrossRef]
- J. Ren, Y. Bai, and Y. Zhao, “Reinforcement-Learned Precision Scaling for CNN Accelerators,” IEEE Trans. Circuits Syst. II, vol. 71, no. 1, pp. 99–103, Jan. 2024. [CrossRef]
- A. Das, M. Tan, and J. Ren, “COMET: A Benchmark for Compression-Efficient Transformers and CNNs on Edge Devices,” IEEE Access, vol. 11, pp. 104382–104395, Nov. 2023. [CrossRef]
- L. Tan, C. Wang, and J. Lee, “Benchmarking Analog and Digital Neural Accelerators for TinyML,” IEEE Trans. Very Large Scale Integr. (VLSI) Syst., vol. 32, no. 4, Apr. 2024.
- W. Liu, H. Shen, and Y. Feng, “TinyEdgeBench: A Benchmark Suite for Edge CNN Accelerators,” IEEE Trans. VLSI Syst., vol. 32, no. 5, pp. 621–634, May 2024. [CrossRef]
- C. L. Kok, T. H. Teo, Y. Y. Koh, Y. Dai, B. K. Ang and J. P. Chai, Development and Evaluation of an IoT-Driven Auto-Infusion System with Advanced Monitoring and Alarm Functionalities, IEEE ISCAS 2024. [CrossRef]




Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).