Submitted:
30 January 2025
Posted:
30 January 2025
You are already at the latest version
Abstract
Keywords:
1. Introduction
2. Methods
2.1. Overall Design Scheme
2.2. Accuracy Predictor Based on Multi-Head Self-Attention
2.3. Data Generator Based on Transformer Encoder and Decoder
2.4. The Process of Training and Using the Accuracy Predictor
3. Implementation and Experimental Results
3.1. Experimental Setting
3.2. Dataset
3.3. Main Results
3.4. Ablation Study
4. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Zhou X; Wu W. Unmanned system swarm intelligence and its research progresses. Microelectronics & Computer, vol. 38, no. 12, 2021, pp. 1-7. [CrossRef]
- Tang L; Ma Z; Li S; Wang Z. The present situation and developing trends of space-based intelligent computing technology. Microelectronics & Computer, vol. 39, no. 4, 2022, pp. 1-8. [CrossRef]
- Protsenko V; Kryzhanovskiy V; Filippov A. Quantization-Friendly Winograd Transformations for Convolutional Neural Networks. European Conference on Computer Vision. Springer, Cham, 2025. [CrossRef]
- Zhou S; Wu Y; Ni Z; Zhou X; Wen H; Zou Y. Dorefa-net: Training low bitwidth convolutional neural networks with low bitwidth gradients. arXiv preprint arXiv:1606.06160, 2016.
- Nagel M.; Fournarakis M.; Amjad R. A.; Bondarenko Y.; Baalen M. V.; Blankevoort T.. A White Paper on Neural Network Quantization. arXiv preprint arXiv:2106.08295, 2021.
- Gholami A.; Kim S.; Dong Z.; Yao Z.; Mahoney M. W.; Keutzer K.. A Survey of Quantization Methods for Efficient Neural Network Inference. arXiv preprint arXiv.2103.13630, 2021.
- Agustsson E; Mentzer F; Tschannen M; et al. Soft-to-hard vector quantization for end-to-end learning compressible representations. arXiv preprint arXiv:1704.00648, 2017.
- Agustsson E; Theis L. Universally Quantized Neural Compression. 2020. https://doi.org/10.48550/arXiv.2006.09952. [CrossRef]
- Kosunalp S. FPGA-QNN: Quantized Neural Network Hardware Acceleration on FPGAs. Applied Sciences, 2025, 15. [CrossRef]
- Zhang Y; Wang R; Zhang Y; et al. Mixed precision quantization of silicon optical neural network chip. Optics Communications, 2025, 574. [CrossRef]
- Pei S; Wang J; Chen Y M. DPQ: dynamic pseudo-mean mixed-precision quantization for pruned neural network. Machine learning, 2024, 113(7):4099-4112. [CrossRef]
- Diao H; Hao Y; Xu S; et al. Implementation of Lightweight Convolutional Neural Networks via Layer-Wise Differentiable Compression. Multidisciplinary Digital Publishing Institute, 2021(10). [CrossRef]
- Hubara I; Nahshan Y; Hanani Y; et al. Accurate post training quantization with small calibration sets. International Conference on Machine Learning, PMLR, 2021, pp. 4466–4475.
- Cai Y; Yao Z; Dong Z; et al. A novel zero shot quantization framework. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 13169–13178.
- Osamaa A; Gadallaha S.; et al. Chaotic neural network quantization and its robustness against adversarial attacks. Knowledge-BasedSystems, 286, 2024.
- Liang T; Glossner J; Wang L; et al. Pruning and quantization for deep neural network acceleration: A survey. arXiv preprint arXiv:2101.09671, 2021.
- Wang Y; Liu Q. AQA: An Adaptive Post-Training Quantization Method for Activations of CNNs. IEEE Transactions on Computers, 2024. https://doi.org/10.1109/TC.2024.3398503. [CrossRef]
- Yang D; He N; Hu X; Yuan Z; et al. Post-training Quantization for Re-parameterization via Coarse & Fine Weight Splitting. Journal of Systems Architecture, 147, 2024. [CrossRef]
- Dong P; Li L; Wei Z; et al. Emq: Evolving training-free proxies for automated mixed precision quantization. Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023: 17076-17086.
- Tang C; Ouyang K; Chai Z; et al. SEAM: Searching Transferable Mixed-Precision Quantization Policy through Large Margin Regularization. Proceedings of the 31st ACM International Conference on Multimedia, 2023, 7971-7980.
- Dong Z; Yao Z; Cai Y; Arfeen D; Gholami A; Mahoney M.W.; Keutzer K. Hawq-v2: Hessian aware trace-weighted quantization of neural networks. Advances in neural information processing systems, 2020.
- Tang C; Ouyang K; Wang Z; et al. Mixed-Precision Neural Network Quantization via Learned Layer-wise Importance. 2022. [CrossRef]
- Choukroun Y; Kravchik E; Kisilev P. Low-bit quantization of neural networks for efficient inference. IEEE/CVF International Conference on Computer Vision Workshop (ICCVW), IEEE, 2019, pp. 3009–3018.
- Wang P; Chen Q; He X; et al. Towards accurate post-training network quantization via bit-split and stitching, in: International Conference on Machine Learning, PMLR, 2020, pp. 9847–9856.
- Wang T.; Zhu J. Y.; Torralba A.; Efros, A. A.. Dataset distillation. arXiv preprint arXiv:1811.10959, 2018.
- Zhao B.; Mopuri K. R.; Bilen H. Dataset condensation with gradient matching. In International Conference on Learning Representations, 2021.
- Sajedi A.; Khaki S.; Amjadian E.; Liu L. Z.; Lawryshyn Y. A.; Plataniotis K. N.. Datadam: Efficient dataset distillation with attention matching. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 17097-17107.
- Guo Z.; Wang K.; Cazenavette G.; Li Hui, Zhang Kaipeng; You Yang. Towards lossless dataset distillation via difficulty-aligned trajectory matching. In The Twelfth International Conference on Learning Representations, 2023.
- Jacob B., Kligys S.; Chen B.; Zhu M.; Tang M.; Howard A; Adam H; Kalenichenko D. Quantization and training of neural networks for efficient integer-arithmetic-only inference. arXiv preprint arXiv:1712.05877, 2017.
- Wang T; Wang K; Cai H; et al. APQ: Joint Search for Network Architecture, Pruning and Quantization Policy. in Proc. IEEE Computer Society Conference on Computer Vision and Pattern Recognition, vol. 2006.08509, pp. 2075-2084, 2020.
- Wang P; Yang A; Men R; et al. Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework. in Proc. International Conference on Machine Learning, pp. 23318-23340, 2022.
- Vaswani A.; Shazeer N.; Parmar N.; Uszkoreit J.; Jones L.; Gomez A. N, Kaiser L. Attention Is All You Need. arXiv, 2017. [CrossRef]
- Dosovitskiy A.; Beyer L.; Kolesnikov A.; Weissenborn D.; Zhai Xiaohua. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In International Conference on Learning Representations, 2021.
- Krishna T; Murali E; Venkatram V; Arun K. Neural Architecture Search for Transformers: A Survey. IEEE Access, 2022, 10, pp. 108374-108412. [CrossRef]
- Liu, Y.; Zhang, Y.; Wang, Y.; Hou, F.; et al. A survey of transformers in computer vision. arXiv preprint arXiv:2111.06091, 2021.
- Carion Nicolas; et al. End-to-end object detection with transformers. In Proceedings of the European Conference on Computer Vision (ECCV), 2020.
- Yuan L.; Chen Y.; Wang T.; Yu W.; Shi Y.; Tay F. E; et al. Tokens-to-token vit: Training vision transformers from scratch on imagenet. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021.
- Liu Z.; Lin Y.; Cao Y.; Hu H.; Wei Y.; Zhang Z.; et al. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021.
- Wang W.; Xie E.; Li X.; Fan D. P.; Song K.; Liang D.; et al. PVT v2: Improved baselines with Pyramid Vision Transformer. Computational Visual Media, 2022. [CrossRef]
- Chang H; Zhang H; Jiang L; et al. MaskGIT: Masked Generative Image Transformer. 2022. [CrossRef]
- Yang Y; Newsam S. Bag-Of-Visual-Words and Spatial Extensions for Land-Use Classification. ACM SIGSPATIAL International Conference on Advances in Geographic Information Systems (ACM GIS), 2010.
- Krizhevsky A; Hinton G. Learning multiple layers of features from tiny images. Handbook of Systemic Autoimmune Diseases, 2009, 1(4).






| Model Index |
label Accuracy | Prediction Accuracy | Prediction Error |Prediction Accuracy- label Accuracy| |
|---|---|---|---|
| 1 | 84.31% | 81.49% | 2.82% |
| 2 | 82.39% | 79.38% | 3.01% |
| 3 | 85.55% | 82.27% | 3.28% |
| 4 | 82.42% | 80.11% | 2.31% |
| 5 | 78.40% | 75.99% | 2.41% |
| Average | - | - | 2.77% |
|
Model Index |
label Accuracy | Prediction Accuracy |
Prediction Error |Prediction Accuracy- label Accuracy| |
| 1 | 88.58% | 86.75% | 1.83% |
| 2 | 87.21% | 84.59% | 2.62% |
| 3 | 86.38% | 84.11% | 2.27% |
| 4 | 83.31% | 81.17% | 2.14% |
| 5 | 85.13% | 82.01% | 3.12% |
| Average | - | - | 2.40% |
| Model Index |
label Accuracy | Prediction Accuracy | Prediction Error |Prediction Accuracy- label Accuracy| |
|---|---|---|---|
| 1 | 87.55% | 85.32% | 2.23% |
| 2 | 87.55% | 86.04% | 1.51% |
| 3 | 85.45% | 83.12% | 2.33% |
| 4 | 87.16% | 86.03% | 1.13% |
| 5 | 85.45% | 83.21% | 2.24% |
| Average | - | - | 1.89% |
| Model | Average Prediction Error (Part 1) |
Average Prediction Accuracy (Part 1 + Part 2) |
|---|---|---|
| MobileNetV2 | 2.05% | 2.77% |
| ResNet50 | 1.91% | 2.40% |
| VGG16 | 1.41% | 1.89% |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).