Submitted:
15 October 2025
Posted:
15 October 2025
You are already at the latest version
Abstract
Keywords:
1. Introduction
2. Historical and Theoretical Background
2.1. Early Developments in Neural Representation
2.2. Capsule Networks Development
2.3. Transformer and GPT Architecture
3. Representation Mechanisms: Vectorial Entity Encoding
3.1. Capsules: Activity Vectors as Instantiation Parameters
3.2. GPT Token Embeddings
3.3. Summary
4. Information Routing and Consensus Mechanisms
4.1. Dynamic Routing-by-Agreement in Capsule Networks
4.2. Self-Attention in GPT
4.3. Conceptual Parallels
5. Hierarchical Information Flow and Abstraction
5.1. Capsule Network Hierarchies
5.2. Multi-Layer Transformer Representations
5.3. Comparative Insight
6. Limitations and Challenges
6.1. Capsule Network Limitations
6.2. GPT Limitations
7. Discussion: Towards Hybrid Models
8. Conclusion
References
- Sabour, Sara, Nicholas Frosst, and Geoffrey E. Hinton. "Dynamic routing between capsules." Advances in neural information processing systems 30 (2017). [CrossRef]
- Vaswani, Ashish, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. "Attention is all you need." Advances in neural information processing systems 30 (2017).
- Radford, Alec, Rafal Jozefowicz, and Ilya Sutskever. "Learning to generate reviews and discovering sentiment." (2018). [CrossRef]
- Radford, Alec, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. "Language models are unsupervised multitask learners." OpenAI blog 1, no. 8 (2019): 9.
- Radford, Alec, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. "Improving language understanding by generative pre-training." (2018): 3.
- Radford, A., et al. (2018–2023). Various publications on GPT architectures. OpenAI.
- Hinton, Geoffrey, Yoshua Bengio, Demis Hassabis, Sam Altman, D. Amodei, D. Song, T. Lieu, B. Gates, Y. Q. Zhang, and I. Sutskever. "Statement on AI risk." Center for AI safety 31, no. 5 (2023): 2023.
- Bender, Emily M., Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. "On the dangers of stochastic parrots: Can language models be too big?." In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, pp. 610-623. 2021. [CrossRef]
- Khodadadzadeh, Massoud, Xuemei Ding, Priyanka Chaurasia, and Damien Coyle. "A hybrid capsule network for hyperspectral image classification." IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing 14 (2021): 11824-11839. [CrossRef]
- Huang, Zhongzhan, Mingfu Liang, Jinghui Qin, Shanshan Zhong, and Liang Lin. "Understanding self-attention mechanism via dynamical system perspective." In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 1412-1422. 2023. [CrossRef]
- Achtibat, Reduan, Sayed Mohammad Vakilzadeh Hatefi, Maximilian Dreyer, Aakriti Jain, Thomas Wiegand, Sebastian Lapuschkin, and Wojciech Samek. "Attnlrp: attention-aware layer-wise relevance propagation for transformers." arXiv preprint arXiv:2402.05602 (2024). arXiv:2402.05602. [CrossRef]
- Jafari, Nafiseh, Mohammad Reza Besharati, and Maryam Hourali. "SELM: Software Engineering of Machine Learning Models." In New Trends in Intelligent Software Methodologies, Tools and Techniques, pp. 48-54. IOS Press, 2021. [CrossRef]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).