Submitted:
24 October 2023
Posted:
25 October 2023
You are already at the latest version
Abstract
Keywords:
0. Introduction
- Proposes a methodology for determining the limits of neural network performance on edge devices
- Analyzes the performance, both runtime and energy, of an edge TPU on both fully connected and Convolutional Neural Networks as compared to a mobile CPU
- Assesses the performance impact of modifications made to convolutional neural networks as part of transfer learning
1. Background and Related Work
1.1. Deep Neural Networks
1.2. Tensor Processing
2. Methodology
2.1. Convolutional Neural Networks
2.2. Exploring Adjustments to CNN Models
2.2.1. Extracting the Baseline Models from Tensorflow Lite
2.2.2. Generating Deeper Models
2.2.3. Generating Wider Models
3. Experimental Results
3.1. Convolutional Neural Networks
3.2. Transfer Learning
3.3. Input and Output Size
3.3.1. Depth Extensions
3.3.2. Wide Extensions
3.3.3. Off-Chip Memory
4. Discussion
4.1. Prefer Single-Core for Long Term Energy Efficiency but Multi-Core for Energy Efficiency Per Inference
4.2. Prefer a TPU for Convolutional Neural Networks
4.3. For Edge TPUs, prefer model depth when possible for convolutional networks
4.4. Avoid performance cliffs for transfer learning with fully-connected input layers
4.5. Scale Width at Input Layers to Exploit Parallelism on the Edge TPU
4.6. Limitations
5. Conclusions
Author Contributions
Funding
Acknowledgments
Conflicts of Interest
Abbreviations
| CPU | Central Processing Unit |
| TPU | Tensor Processing Unit |
| CNN | Convolutional Neural Network |
References
- "What makes TPUs fine-tuned for deep learning?" Google. [Online]. Available: https://cloud.google.com/blog/products/ai-machine-learning/ what-makes-tpus-fine-tuned-for-deep-learning/. [Accessed: 07-May-2022].
- "Frequently asked questions," Coral. [Online]. Available: https://coral.ai/docs/edgetpu/faq/what-is-the- edge-tpu/. [Accessed: 07-May-2022].
- "Cloud TPU Performance Guide," Google. [Online]. Available: https://cloud.google.com/tpu/docs/performance-guide. [Accessed: 01-Sept-2023].
- D. Xu, M. Zheng, L. Jiang, C. Gu, R. Tan, and P. Cheng. "Lightweight and Unobtrusive Data Obfuscation at IoT Edge for Remote Inference," IEEE Internet of Things Journal, vol 7, 2020, pp. 9540-9551. [CrossRef]
- A. Yazdanbakshsh, K. Seshadri, B. Akin, J. Laudon, R. Narayanaswami. "An Evaluation of Edge TPU Accelerators for Convolutional Neural Networks," 2020, https://arxiv.org/abs/2102.10423. [CrossRef]
- “Edge TPU performance benchmarks,” Coral. [Online]. Available: https://coral.ai/docs/edgetpu/benchmarks/. [Accessed: 08-May-2022].
- N. P. Jouppi et al., "In-datacenter performance analysis of a tensor processing unit," 2017 ACM/IEEE 44th Annual International Symposium on Computer Architecture (ISCA), 2017, pp. 1-12. [CrossRef]
- “Dev Board Datasheet,” Coral. [Online]. Available: https://coral.ai/docs/dev-board/datasheet/. [Accessed: 10-May-2022].
- “Arm Cortex-A53 MPCore Processor Technical Reference Manual,” ARM. [Online]. Available: https://developer.arm.com/documentation/ddi0500/latest/. [Accessed: 10-May-2022].
- Baller, S.P.; Jindal, A.; Chadha, M.; Gerndt, M. "DeepEdgeBench: Benchmarking Deep Neural Networks on Edge Devices." In Proceedings of the 2021 IEEE International Conference on Cloud Engineering (IC2E), San Francisco, CA, USA, 4–8 October 2021; pp. 20–30. [CrossRef]
- Dominguez-Morales JP, Duran-Lopez L, Gutierrez-Galan D, Rios-Navarro A, Linares-Barranco A, Jimenez-Fernandez A. Wildlife Monitoring on the Edge: A Performance Evaluation of Embedded Neural Networks on Microcontrollers for Animal Behavior Classification. Sensors. 2021; 21(9):2975. [CrossRef]
- B. Kim, S. Lee, A. R. Trivedi and W. J. Song, Energy-Efficient Acceleration of Deep Neural Networks on Realtime-Constrained Embedded Edge Devices, in IEEE Access, vol. 8, pp. 216259-216270, 2020. [CrossRef]
- “Energizer L91 Ultimate Lithium Product Datasheet,” Energizer. [Online]. Available: https://data.energizer.com/pdfs/l91.pdf. [Accessed: 10-May-2022].
- A. Yazdanbakhsh, K. Seshadri, B. Akin, J, Laudon, and R. Narayanaswami, “An evaluation of Edge TPU accelerators for convolutional neural networks,” arXiv preparing arXiv:2102.10423, 2021. [CrossRef]
- “Gen7i Transient Recorder and Data Acquisition System,” Durham Instruments. [Online]. Available: https://disensors.com/product/gen7i-transient-recorder-and-data-acquisition-system/. [Accessed: 20-June-2022].
- "Trained TensorFlow models for the Edge TPU," Coral. [Online]. Available: https://coral.ai/models/. [Accessed: 29-Jan-2022].
- F. Zhuang et al., "A Comprehensive Survey on Transfer Learning," in Proceedings of the IEEE, vol. 109, no. 1, pp. 43-76, Jan. 2021. [CrossRef]
- Samira Pouyanfar, Saad Sadiq, Yilin Yan, Haiman Tian, Yudong Tao, Maria Presa Reyes, Mei-Ling Shyu, Shu-Ching Chen, and S. S. Iyengar. 2018. A Survey on Deep Learning: Algorithms, Techniques, and Applications. ACM Comput. Surv. 51, 5, Article 92 (September 2019), 36 pages. [CrossRef]
- Alom, M.Z., Taha, T.M., Yakopcic, C., Westberg, S., Sidike, P., Nasrin, M.S., Hasan, M., Van Essen, B.C., Awwal, A.A.S., Asari, V.K. A State-of-the-Art Survey on Deep Learning Theory and Architectures. Electronics 2019, 8, 292. [CrossRef]
- G. Bebis and M. Georgiopoulos, "Feed-forward neural networks," in IEEE Potentials, vol. 13, no. 4, pp. 27-31, Oct.-Nov. 1994. [CrossRef]
- Z. Li, F. Liu, W. Yang, S. Peng and J. Zhou, "A Survey of Convolutional Neural Networks: Analysis, Applications, and Prospects," in IEEE Transactions on Neural Networks and Learning Systems, vol. 33, no. 12, pp. 6999-7019, Dec. 2022. [CrossRef]
- Jaehoon Lee, Yasaman Bahri, Roman Novak, Sam Schoenholz, Jeffrey Pennington, and Jascha Sohl-dickstein. Deep neural networks as gaussian processes. In International Conference on Learning Representations, 2018.
- Roman Novak, Lechao Xiao, Jaehoon Lee, Yasaman Bahri, Greg Yang, Jiri Hron, Daniel A. Abolafia, Jeffrey Pennington, and Jascha Sohl-Dickstein. Bayesian deep convolutional networks with many channels are gaussian processes. In International Conference on Learning Representations, 2019.
- Jaehoon Lee, Lechao Xiao, Samuel Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington. Wide Neural Networks of Any Depth Evolve as Linear Models Under Gradient Descent. In NeurIPS 2019. [CrossRef]
- Mavrovouniotis, M., Yang, S. Training neural networks with ant colony optimization algorithms for pattern classification. Soft Computing 19, 1511–1522 (2015). [CrossRef]
- Beheshti, Z., Shamsuddin, S.M.H., Beheshti, E. et al. Enhancement of artificial neural network learning using centripetal accelerated particle swarm optimization for medical diseases diagnosis. Soft Computing 18, 2253–2270 (2014). [CrossRef]
- Ashraf Mohamed Hemeida, Somaia Awad Hassan, Al-Attar Ali Mohamed, Salem Alkhalaf, Mountasser Mohamed Mahmoud, Tomonobu Senjyu, and Ayman Bahaa El-Din. Nature-inspired algorithms for feed-forward neural network classifiers: A survey of one decade of research. Ain Shams Engineering Journal, Volume 11, Issue 3, 2020, Pages 659-675, ISSN 2090-4479. [CrossRef]
- Christian Szegedy, Sergey Ioffe, Vincent Vanhoucke, and Alexander A. Alemi. 2017. Inception-v4, inception-ResNet and the impact of residual connections on learning. In Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence (AAAI’17). AAAI Press, 4278–4284. [CrossRef]
- Yann LeCun and Yoshua Bengio. 1998. Convolutional networks for images, speech, and time series. The handbook of brain theory and neural networks. MIT Press, Cambridge, MA, USA, 255–258.
- Eugene Charniak. 2019. Introduction to Deep Learning. The MIT Press.
- F. Wang et al., "Residual Attention Network for Image Classification," 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 2017, pp. 6450-6458. [CrossRef]
- C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens and Z. Wojna, "Rethinking the Inception Architecture for Computer Vision," 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 2016, pp. 2818-2826. [CrossRef]
- Nsight Compute Occupancy Calculator. [Online]. Available: https://docs.nvidia.com/nsight-compute/NsightCompute/index.html#occupancy-calculator. [Accessed: 9-Jun-2023].
- Seyedehfaezeh Hosseininoorbin, Siamak Layeghy, Brano Kusy, Raja Jurdak, Marius Portmann. Exploring Edge TPU for deep feed-forward neural networks. Internet of Things, Volume 22, 2023, 100749, ISSN 2542-6605. [CrossRef]
- Y. Ni, Y. Kim, T. Rosing and M. Imani, "Online Performance and Power Prediction for Edge TPU via Comprehensive Characterization," 2022 Design, Automation & Test in Europe Conference & Exhibition (DATE), Antwerp, Belgium, 2022, pp. 612-615. [CrossRef]
- Seyedehfaezeh Hosseininoorbin, Siamak Layeghy, Mohanad Sarhan, Raja Jurdak, Marius Portmann. Exploring edge TPU for network intrusion detection in IoT. Journal of Parallel and Distributed Computing, Volume 179, 2023, 104712, ISSN 0743-7315. [CrossRef]
- A. A. Asyraaf Jainuddin, Y. C. Hou, M. Z. Baharuddin and S. Yussof, "Performance Analysis of Deep Neural Networks for Object Classification with Edge TPU," 2020 8th International Conference on Information Technology and Multimedia (ICIMU), Selangor, Malaysia, 2020, pp. 323-328. [CrossRef]
- C. DeLozier, F. Rooney, J. Jung, J. A. Blanco, R. Rakvic and J. Shey, "A Performance Analysis of Deep Neural Network Models on an Edge Tensor Processing Unit," 2022 International Conference on Electrical, Computer and Energy Technologies (ICECET), Prague, Czech Republic, 2022, pp. 1-6. [CrossRef]
- M. Tan, R. Pang and Q. Le, "EfficientDet: Scalable and Efficient Object Detection," in 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 2020 pp. 10778-10787. [CrossRef]
- M. Tan, Q. Le, "EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks," in Proceedings of the 36th Internation Conference on Machine Learning, Long Beach, CA, PMLR 97, 2019.
- C. Szegedy, et al., "Going deeper with convolutions," in 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA, 2015 pp. 1-9. [CrossRef]
- Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. Mobilenetv2: Inverted residuals and linear bottlenecks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4510–4520, 2018. [CrossRef]
- Andrew Howard, Mark Sandler, Grace Chu, Liang-Chieh Chen, Bo Chen, Mingxing Tan, Weijun Wang, Yukun Zhu, Ruoming Pang, Vijay Vasudevan, et al. Searching for mobilenetv3. In Proceedings of the IEEE International Conference on Computer Vision, pages 1314–1324, 2019. [CrossRef]
- Howard, A. G., Zhu, M., Chen, B., Kalenichenko, D., Wang, W., Weyand, T., Adam, H. (2017). MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications. CoRR, abs/1704.04861. Retrieved from http://arxiv.org/abs/1704.04861. [CrossRef]
- MobileNet, MobileNetV2, and MobileNetV3. [Online]. Available: https://keras.io/api/applications/mobilenet/. [Accessed: 28-Sept-23].
- Tensorflow 2 Detection Model Zoo. [Online]. Available: https://github.com/tensorflow/models/blob/ master/research/object_detection/g3doc/tf2_detection_zoo.md. [Accessed: 28-Sept-23].
- Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation. Liang-Chieh Chen and Yukun Zhu and George Papandreou and Florian Schroff and Hartwig Adam. arXiv:1802.02611, 2018. [CrossRef]
| 1 |
nodes in the fully-connected layer |

















| Coral TPU [8] | Coral CPU [8,9] | |
| Processor | Google Edge TPU | Cortex-A53 Quad-core |
| Frequency | 480 MHz | 3.01 GHz |
| RAM | 4 GB DDR4 | 4 GB DDR4 |
| Operation Type | Fixed Point | Floating Point |
| Operations/s | 4 Trillion (8-bit) | 32 Billion (32-bit) |
| Model Name | Input | Output | GFLOP | Layers | % Parallel Layers |
| EfficientDet320 [39] | 320x320x3 | 90 | 2,323.1 | 266 | 36% |
| EfficientDet384 [39] | 384x384x3 | 90 | 4,272.3 | 321 | 31% |
| EfficientDet448 [39] | 448x448x3 | 90 | 6,806.7 | 356 | 29% |
| EfficientDet512 [39] | 512x512x3 | 90 | 13,117.4 | 423 | 29% |
| EfficientDet640 [39] | 640x640x3 | 90 | 23,671.8 | 423 | 29% |
| EfficientNetS [40] | 244x244x3 | 1000 | 2,991.1 | 66 | 0% |
| EfficientNetM [40] | 240x240x3 | 1000 | 4,598.3 | 86 | 0% |
| EfficientNetL [40] | 300x300x3 | 1000 | 11,752.0 | 97 | 0% |
| InceptionV1 [41] | 244x244x3 | 1000 | 2,167.2 | 83 | 53% |
| InceptionV2 [32] | 244x244x3 | 1000 | 2,708.2 | 98 | 46% |
| InceptionV3 [32] | 299x299x3 | 1000 | 7,347.8 | 132 | 47% |
| InceptionV4 [28] | 299x299x3 | 1000 | 15,666.1 | 205 | 32% |
| MobileDetSSDLite [42] | 320x320x3 | 90 | 2,437.3 | 136 | 32% |
| MobileDetV1 [44] | 300x300x3 | 90 | 1,929.1 | 75 | 44% |
| MobileDetV2Coco [42] | 300x300x3 | 90 | 1,494.4 | 110 | 30% |
| MobileDetV2Face [42] | 320x320x3 | 90 | 1,524.6 | 132 | 25% |
| TF2MobileDetV1 [46] | 640x640x3 | 90 | 67,482.4 | 104 | 56% |
| TF2MobileDetV2 [46] | 300x300x3 | 90 | 1,407.5 | 101 | 24% |
| DLV3DM05MobileNet [47] | 513x513x3 | 20 | 2,276.7 | 72 | 0% |
| DLV3MobileNet [47] | 513x513x3 | 20 | 5,343.5 | 72 | 0% |
| KerasMobileNet128 [45] | 128x128x3 | 37 | 1,350.8 | 76 | 10% |
| KerasMobileNet256 [45] | 256x256x3 | 37 | 5,390.1 | 76 | 10% |
| MobileNet0.25 [44] | 128x128x3 | 1000 | 37.8 | 31 | 0% |
| MobileNet0.5 [44] | 160x160x3 | 1000 | 155.1 | 31 | 0% |
| MobileNet0.75 [44] | 192x192x3 | 1000 | 412.1 | 31 | 0% |
| MobileNet1.0 [44] | 224x224x3 | 1000 | 912.6 | 31 | 0% |
| MobileNetV2Bird [42] | 224x224x3 | 900 | 652.3 | 65 | 0% |
| MobileNetV2Plant [42] | 224x224x3 | 2000 | 658.6 | 65 | 0% |
| MobileNetV2 [42] | 224x224x3 | 1000 | 652.3 | 66 | 0% |
| TF2MobileNetV1 [46] | 224x224x3 | 1000 | 840.3 | 33 | 0% |
| TF2MobileNetV2 [46] | 224x224x3 | 1000 | 614.8 | 68 | 0% |
| TF2MobileNetV3 [46] | 224x224x3 | 1000 | 1,280.9 | 79 | 0% |
| FNN240L810N [38] | 100x1 | 9 | 565.3 | 242 | 0% |
| Model | TPU Speedup (14 FC Nodes) | TPU Speedup (224 FC Nodes) |
| efficientnet | 33.0x | 0.99x |
| inception | 17.9x | 0.94x |
| mobilnet | 13.3x | 0.85x |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2023 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).