Submitted:
21 July 2025
Posted:
22 July 2025
You are already at the latest version
Abstract
Keywords:
1. Introduction
2. Related Work
3. Datasets
- ImageNet-R [25]: Introduced in 2021, ImageNet-R is a robustness benchmark containing 30 000 validation images across 200 classes, curated from artistic renditions, cartoons, graffiti and other renditions of the original ImageNet classes. We resize all images to and apply the standard ImageNet normalization (, ). We report single image (batch size = 1) latency and batch-16 throughput for ResNet-50.
- Qasper [26]: Released at EMNLP 2021, Qasper is a question-answering dataset over NLP research papers, consisting of 3 049 annotated QA pairs on 1 049 documents, plus 36 000+ unannotated examples. We use the publicly provided BERT small tokenizer, limit inputs to 384 tokens, and measure end-to-end inference latency (batch size=1) and throughput (batch size=16).
4. Methodology
5. Experimental Evaluation
- CPU: Intel Xeon E5-2670 (8 cores @ 2.6 GHz, AVX2)
- GPU: NVIDIA Tesla V100 (5120 CUDA cores, 16 GB HBM2)
- NPU: Mobile Edge NPU (128 ALUs, 4 GB on-chip SRAM)
- Unfused MLFast: The same inference engine with no operator fusion.
- Vendor Library: NVIDIA TensorRT on GPU, Eigen-optimized C++ on CPU, and the NPU’s proprietary SDK.
6. Discussion
- Portability: The same fusion logic and UIR serve CPU, GPU, and NPU with no changes, reducing engineering effort.
- Cost Model Accuracy: Offline-profiled device parameters enable reliable performance predictions; minor deviations (< 5 %) from actual latencies validate the model’s fidelity.
- Scalability: Dynamic programming over the UIR scales linearly with graph size; end-to-end planning adds less than 5 ms even on large networks.
- Limitations and Future Work:
- Runtime Adaptation: Current profiling is static; integrating online performance feedback could further refine fusion decisions for varying workloads.
- Memory Constraints: Extremely large fusion groups may exceed on-chip memory, suggesting an opportunity for multi-stage tiling and split-fusion strategies.
- Extensibility: While UOF supports new devices via profiles and templates, automating profile collection and codegen template synthesis would streamline onboarding of novel accelerators.
7. Conclusions
References
- Wolf, M.; Lam, M.S. A Data Locality Optimizing Algorithm. In Proceedings of the Proceedings of the ACM SIGPLAN 1991 Conference on Programming Language Design and Implementation (PLDI). ACM, 1991; pp. 30–44. [Google Scholar]
- Kennedy, K.; Allen, J.R. Optimal Loop Fusion in Linear Time. In Proceedings of the Proceedings of the 1986 ACM/IEEE Conference on Supercomputing. IEEE Computer Society, 1986; pp. 302–311. [Google Scholar]
- TensorFlow XLA Team. XLA: Optimizing Compiler for Machine Learning. In Proceedings of the Google I/O, 2018. Available online: https://www.tensorflow.org/xla.
- Chen, T.; Moreau, T.; Jiang, Z.; Zheng, L.; Yan, E.; Shen, H.; Cowen, D.; Wang, Y.; Hu, L.; Ceze, L.; et al. TVM: An Automated End-to-End Optimizing Compiler for Deep Learning. In Proceedings of the 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI); 2018; pp. 578–594. [Google Scholar]
- ONNX Runtime Team. ONNX Runtime: Cross-platform, High-performance Scoring Engine for Open Neural Network Exchange (ONNX) Models. arXiv preprint 2020, arXiv:2006.14802 2020.
- NVIDIA Corporation. TensorRT: High-Performance Deep Learning Inference Optimizer and Runtime. In Proceedings of the NVIDIA Deep Learning Institute Workshop; 2016. [Google Scholar]
- Lattner, C.; et al. MLIR: A Compiler Infrastructure for the End of Moore’s Law. In Proceedings of the Proceedings of the ACM SIGPLAN International Workshop on Machine Learning and Programming Languages (MAPL); 2019; pp. 1–12. [Google Scholar]
- Roesch, D.; Chen, T.; Herrmann, M.; Shenker, S.; Onufryk, O.; et al. Glow: Graph Lowering Compiler Techniques for Neural Networks. In Proceedings of the Proceedings of the 2018 Conference on Systems and Machine Learning (SysML); 2018. [Google Scholar]
- Chen, T.; Moreau, T.; Xu, Z.; Zheng, L.; Yan, E.; Shen, H.; Cowen, D.; Wang, Y.; Hu, L.; Ceze, L.; et al. Relay: A High-Level IR for Machine Learning. In Proceedings of the Proceedings of the Workshop on MLIR for Tailored Software and Hardware (SYSML); 2019. [Google Scholar]
- Li, M.; Zhao, K.; Guo, Y.; Kan, X.; Zhang, L.; Li, K.; Wang, X.; et al. Heterogeneous Task Scheduling for Distributed Machine Learning. In Proceedings of the Proceedings of the Fourth Workshop on Hot Topics in Operating Systems (HotOS); 2019. [Google Scholar]
- Jia, X.; Liu, J.; Yuan, H.; He, Y.; Chen, K.; Taylor, G. FlexFlow: A Flexible Dataflow Programming Model for Distributed Deep Learning. In Proceedings of the Proceedings of the 2019 USENIX Annual Technical Conference (USENIX ATC); 2019. [Google Scholar]
- Wang, C.; Gong, J. Intelligent agricultural greenhouse control system based on internet of things and machine learning. arXiv preprint 2024, arXiv:2402.09488 2024. [Google Scholar]
- Wang, C.; Quach, H.T. Exploring the Effect of Sequence Smoothness on Machine Learning Accuracy. In Proceedings of the International Conference On Innovative Computing And Communication; Springer, 2024; pp. 475–494. [Google Scholar]
- Li, C.; Zheng, H.; Sun, Y.; Wang, C.; Yu, L.; Chang, C.; Tian, X.; Liu, B. Enhancing multi-hop knowledge graph reasoning through reward shaping techniques. In Proceedings of the 2024 4th International Conference on Machine Learning and Intelligent Systems Engineering (MLISE); IEEE, 2024; pp. 1–5. [Google Scholar]
- Liu, H.; Wang, C.; Zhan, X.; Zheng, H.; Che, C. Enhancing 3D Object Detection by Using Neural Network with Self-adaptive Thresholding. In Proceedings of the Proceedings of the 2nd International Conference on Software Engineering and Machine Learning; 2024; Vol. 67. [Google Scholar]
- Liu, M.; Sui, M.; Nian, Y.; Wang, C.; Zhou, Z. CA-BERT: Leveraging Context Awareness for Enhanced Multi-Turn Chat Interaction. In Proceedings of the 2024 5th International Conference on Big Data & Artificial Intelligence & Software Engineering (ICBASE), 2024, IEEE; pp. 388–392.
- Wang, C.; Sui, M.; Sun, D.; Zhang, Z.; Zhou, Y. Theoretical analysis of meta reinforcement learning: Generalization bounds and convergence guarantees. In Proceedings of the Proceedings of the International Conference on Modeling, Natural Language Processing and Machine Learning; 2024; pp. 153–159. [Google Scholar]
- Wang, C.; Yang, Y.; Li, R.; Sun, D.; Cai, R.; Zhang, Y.; Fu, C. Adapting llms for efficient context processing through soft prompt compression. In Proceedings of the Proceedings of the International Conference on Modeling, Natural Language Processing and Machine Learning; 2024; pp. 91–97. [Google Scholar]
- Wu, T.; Wang, Y.; Quach, N. Advancements in natural language processing: Exploring transformer-based architectures for text understanding. arXiv preprint 2025, arXiv:2503.20227 2025. [Google Scholar]
- Gao, Z. Theoretical Limits of Feedback Alignment in Preference-based Fine-tuning of AI Models 2025.
- Gao, Z. Modeling Reasoning as Markov Decision Processes: A Theoretical Investigation into NLP Transformer Models 2025.
- Gao, Z. Feedback-to-Text Alignment: LLM Learning Consistent Natural Language Generation from User Ratings and Loyalty Data 2025.
- Sang, Y. Robustness of Fine-Tuned LLMs under Noisy Retrieval Inputs. Preprints 2025. [Google Scholar] [CrossRef]
- Quach, N.; Wang, Q.; Gao, Z.; Sun, Q.; Guan, B.; Floyd, L. Reinforcement learning approach for integrating compressed contexts into knowledge graphs. In Proceedings of the 2024 5th International Conference on Computer Vision, Image and Deep Learning (CVIDL); IEEE, 2024; pp. 862–866. [Google Scholar]
- Hendrycks, D.; Mu, S.; Dietterich, T.G. The Many Faces of Robustness: Benchmarking Neural Network Robustness to Common Corruptions, Perturbations, and Subtextures. In Proceedings of the Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV); 2021. ImageNet-R. [Google Scholar]
- Lo, D.; Neeraj, K.; Joshi, M.; Thomas, L.; Biderman, S.; Lenhert, F. Qasper: A Dataset of Information-Seeking Questions and Answers on Research Papers. In Proceedings of the Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP); 2021; pp. 4963–4979. [Google Scholar]




| Model / Device | Latency (ms) | Throughput | ||
| Unfused | UOF | Unfused | UOF | |
| ResNet-50 / CPU | 120 | 75 | 8.3 | 13.3 |
| ResNet-50 / GPU | 25 | 12 | 40 | 83 |
| ResNet-50 / NPU | 60 | 20 | 16 | 50 |
| BERT-small / CPU | 200 | 120 | 5.0 | 8.3 |
| BERT-small / GPU | 45 | 22 | 22 | 45 |
| BERT-small / NPU | 80 | 28 | 12.5 | 35.7 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).