Submitted:
21 April 2025
Posted:
22 April 2025
You are already at the latest version
Abstract
Keywords:
1. Introduction
2. Related Work
3. Methods
3.1. Humanoid Social Robot Pepper
3.2. Cognitive Model
3.3. Basic Settings of the LLM
4. Content Retrieval from the Declarative Memory of ACT-R
4.1. Example Application 1: Recollection Based on Text
4.1.1. Prompting the LLM
4.1.2. Results
4.2. Example Application 2: Recollection Based on Vision
4.2.1. Prompting the LLM
4.2.2. Results
5. Discussion
6. Conclusion
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
Abbreviations
| ACT-R | Adaptive Control of Thought-Rational |
| API | Application Programming Interface |
| AI | Artificial Intelligence |
| ACo | Artificial Cognition |
| CL | Continual Learning |
| CoALA | Cognitive Architectures for Language Agents |
| GPT | Generative Pretrained Transformer |
| HRI | Human-Robot Interaction |
| IBL | Instance-based Learning Theory |
| LLM | Large Language Model |
| LMM | Large Multi-modal Models |
| LTM | Long-Term Memory |
| RAG | Retrieval-Augmented Generation |
| TCP/IP | Transmission Control Protocol/Internet Protocol |
| VLM | Vision-Language Model |
References
- Bisaz, R.; Travaglia, A.; Alberini, C. The Neurobiological Bases of Memory Formation: From Physiological Conditions to Psychopathology. Psychopathology 2014, 47. [Google Scholar] [CrossRef] [PubMed]
- Zhang, J. How memories are stored in the brain: the declarative memory model, 2024, [arXiv:q-bio.NC/2403.17985]. [CrossRef]
- Myers, N.E.; Stokes, M.G.; Nobre, A.C. Prioritizing Information during Working Memory: Beyond Sustained Internal Attention. Trends in Cognitive Sciences 2017, 21, 449–461. [Google Scholar] [CrossRef] [PubMed]
- Braun, E.K.; Wimmer, G.E.; Shohamy, D. Retroactive and graded prioritization of memory by reward. Nature Communications 2018, 9, 2041–1723. [Google Scholar] [CrossRef] [PubMed]
- Sandini, G.; Sciutti, A.; Morasso, P. Artificial cognition vs. artificial intelligence for next-generation autonomous robotic agents. Frontiers in Computational Neuroscience 2024, 18. [Google Scholar] [CrossRef]
- Dubey, S.; Ghosh, R.; Dubey, M.; Chatterjee, S.; Das, S.; Benito-León, J. Redefining Cognitive Domains in the Era of ChatGPT: A Comprehensive Analysis of Artificial Intelligence’s Influence and Future Implications. Medical Research Archives 2024, 12. [Google Scholar] [CrossRef]
- Janik, R.A. Aspects of human memory and Large Language Models, 2024, [arXiv:cs.CL/2311.03839]. [CrossRef]
- Jeong, H.; Lee, H.; Kim, C.; Shin, S. A Survey of Robot Intelligence with Large Language Models. Applied Sciences 2024, 14. [Google Scholar] [CrossRef]
- Ghosh, A.; Acharya, A.; Saha, S.; Jain, V.; Chadha, A. Exploring the Frontier of Vision-Language Models: A Survey of Current Methodologies and Future Directions, 2024, [arXiv:cs.CV/2404.07214]. [CrossRef]
- Kawaharazuka, K.; Obinata, Y.; Kanazawa, N.; Okada, K.; Inaba, M. Robotic Applications of Pre-Trained Vision-Language Models to Various Recognition Behaviors. In Proceedings of the 2023 IEEE-RAS 22nd International Conference on Humanoid Robots (Humanoids). IEEE; 2023; pp. 1–8. [Google Scholar] [CrossRef]
- Tan, R.; Yang, H.; Jiao, S.; Shan, L.; Tao, C.; Jiao, R. Smart Robot Manipulation using GPT-4o Vision. In Proceedings of the 7th European Industrial Engineering and Operations Management Conference, Augsburg, Germany. IEOM Society International; 2024. [Google Scholar] [CrossRef]
- Kurup, U.; Lebiere, C. What can cognitive architectures do for robotics? Biologically Inspired Cognitive Architectures 2012, 2, 88–99. [Google Scholar] [CrossRef]
- Anderson, J.R.; Bothell, D.; Byrne, M.D.; Douglass, S.; Lebiere, C.; Qin, Y. An integrated theory of the mind 2004. 111, 1036–1060. [CrossRef]
- Thomson, R.; Lebiere, C.; Anderson, J.R.; Staszewski, J. A general instance-based learning framework for studying intuitive decision-making in a cognitive architecture. Journal of Applied Research in Memory and Cognition 2015, 4, 180–190, Modeling and Aiding Intuition in Organizational Decision Making. [Google Scholar] [CrossRef]
- Gonzalez, C.; Dutt, V.; Lebiere, C. Validating instance-based learning mechanisms outside of ACT-R. Journal of Computational Science 2013, 4, 262–268, PEDISWESA 2011 and Sc. computing for Cog. Sciences. [Google Scholar] [CrossRef]
- Gonzalez, C.; Lerch, J.F.; Lebiere, C. Instance-based learning in dynamic decision making. Cognitive Science 2003, 27, 591–635. [Google Scholar] [CrossRef]
- Werk, A.; Scholz, S.; Sievers, T.; Russwinkel, N. How to Provide a Dynamic Cognitive Person Model of a Human Collaboration Partner to a Pepper Robot. ICCM 2024. [Google Scholar]
- Kahneman, D. Thinking, fast and slow; Farrar, Straus and Giroux: New York, 2011. [Google Scholar]
- Whitehill, J. Understanding ACT-R - an Outsider’s Perspective, 2013.
- Quam, C.; Wang, A.; Maddox, W.T.; Golisch, K.; Lotto, A. Procedural-Memory, Working-Memory, and Declarative-Memory Skills Are Each Associated With Dimensional Integration in Sound-Category Learning. Frontiers in Psychology 2018, 9. [Google Scholar] [CrossRef] [PubMed]
- Eichenbaum, H. Hippocampus: Cognitive Processes and Neural Representations that Underlie Declarative Memory. Neuron 2004, 44, 109–120. [Google Scholar] [CrossRef] [PubMed]
- Hayne, H.; Boniface, J.; Barr, R. The Development of Declarative Memory in Human Infants: Age-Related Changes in Deferred Imitation. Behavioral neuroscience 2000, 114, 77–83. [Google Scholar] [CrossRef] [PubMed]
- OpenAI. The most powerful platform for building AI products. Technical report, 2025.
- Carnegie Mellon University. ACT-R Software. Technical report, 2024.
- Sievers, T.; Russwinkel, N. How to use a cognitive architecture for a dynamic person model with a social robot in human collaboration. CEUR Workshop Proceedings 2024. [Google Scholar]
- Lykov, A.; Litvinov, M.; Konenkov, M.; Prochii, R.; Burtsev, N.; Abdulkarim, A.A.; Bazhenov, A.; Berman, V.; Tsetserukou, D. CognitiveDog: Large Multimodal Model Based System to Translate Vision and Language into Action of Quadruped Robot, New York, NY, USA, 2024; HRI ’24, p. 712–716. [CrossRef]
- Wang, J.; Shi, E.; Hu, H.; Ma, C.; Liu, Y.; Wang, X.; Yao, Y.; Liu, X.; Ge, B.; Zhang, S. Large language models for robotics: Opportunities, challenges, and perspectives. Journal of Automation and Intelligence 2024. [Google Scholar] [CrossRef]
- Hu, Y.; Lin, F.; Zhang, T.; Yi, L.; Gao, Y. Look Before You Leap: Unveiling the Power of GPT-4V in Robotic Vision-Language Planning, 2023, [arXiv:cs.RO/2311.17842]. [CrossRef]
- Wake, N.; Kanehira, A.; Sasabuchi, K.; Takamatsu, J.; Ikeuchi, K. GPT-4V(ision) for Robotics: Multimodal Task Planning from Human Demonstration, 2024, [arXiv:cs.RO/2311.12015]. [CrossRef]
- Yoshida, T.; Baba, S.; Masumori, A.; Ikegami, T. Minimal Self in Humanoid Robot "Alter3" Driven by Large Language Model, 2024, [arXiv:cs.RO/2406.11420]. [CrossRef]
- Leidner, D. Toward Robotic Metacognition: Redefining Self-Awareness in an Era of Vision-Language Models. 40th Anniversary of the IEEE Conference on Robotics and Automation 2024. [Google Scholar]
- Niu, Q.; Liu, J.; Bi, Z.; Feng, P.; Peng, B.; Chen, K.; Li, M.; Yan, L.K.; Zhang, Y.; Yin, C.H.; et al. Large Language Models and Cognitive Science: A Comprehensive Review of Similarities, Differences, and Challenges, 2024.
- Wu, S.; Oltramari, A.; Francis, J.; Giles, C.L.; Ritter, F.E. Cognitive LLMs: Towards Integrating Cognitive Architectures and Large Language Models for Manufacturing Decision-making, 2024.
- González-Santamarta, M.Á.; Rodríguez-Lera, F.J.; Guerrero-Higueras, Á.M.; Matellán-Olivera, V. Integration of Large Language Models within Cognitive Architectures for Autonomous Robots, 2024, [arXiv:cs.RO/2309.14945]. [CrossRef]
- Sievers, T.; Russwinkel, N.; Möller, R. Grounding a Social Robot’s Understanding of Words with Associations in a Cognitive Architecture. In Proceedings of the Proceedings of the 17th International Conference on Agents and Artificial Intelligence - Volume 3: ICAART. INSTICC, SciTePress, 2025, pp. 406–410. [CrossRef]
- Knowles, K.; Witbrock, M.; Dobbie, G.; Yogarajan, V. A Proposal for a Language Model Based Cognitive Architecture. Proceedings of the AAAI Symposium Series 2024. [Google Scholar] [CrossRef]
- Leivada, E.; Marcus, G.; Günther, F.; Murphy, E. A Sentence is Worth a Thousand Pictures: Can Large Language Models Understand Hum4n L4ngu4ge and the W0rld behind W0rds?, 2024.
- Sun, R. Can A Cognitive Architecture Fundamentally Enhance LLMs? Or Vice Versa?, 2024.
- He, Z.; Lin, W.; Zheng, H.; Zhang, F.; Jones, M.W.; Aitchison, L.; Xu, X.; Liu, M.; Kristensson, P.O.; Shen, J. Human-inspired Perspectives: A Survey on AI Long-term Memory, 2025, [arXiv:cs.AI/2411.00489]. [CrossRef]
- Jiang, X.; Li, F.; Zhao, H.; Wang, J.; Shao, J.; Xu, S.; Zhang, S.; Chen, W.; Tang, X.; Chen, Y.; et al. Long Term Memory: The Foundation of AI Self-Evolution, 2024, [arXiv:cs.AI/2410.15665]. [CrossRef]
- Hou, Y.; Tamoto, H.; Miyashita, H. "My agent understands me better": Integrating Dynamic Human-like Memory Recall and Consolidation in LLM-Based Agents. In Proceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems, New York, NY, USA, 2024; CHI EA ’24. [Google Scholar] [CrossRef]
- umers, T.R.; Yao, S.; Narasimhan, K.; Griffiths, T.L. Cognitive Architectures for Language Agents, 2024, [arXiv:cs.AI/2309.02427]. [CrossRef]
- Gao, Y.; Xiong, Y.; Gao, X.; Jia, K.; Pan, J.; Bi, Y.; Dai, Y.; Sun, J.; Wang, M.; Wang, H. Retrieval-Augmented Generation for Large Language Models: A Survey, 2024, [arXiv:cs.CL/2312.10997]. [CrossRef]
- Sanmartin, D. KG-RAG: Bridging the Gap Between Knowledge and Creativity, 2024, [arXiv:cs.AI/2405.12035]. [CrossRef]
- Janik, R.A. Aspects of human memory and Large Language Models, 2024.
- Aldebaran, United Robotics Group and Softbank Robotics. Pepper. Technical report, 2025.
- Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. Language Models are Few-Shot Learners, 2020. [CrossRef]
- Aldebaran, United Robotics Group and Softbank Robotics. Pepper SDK for Android. Technical report, 2025.
- Sreedharan, S.; Kulkarni, A.; Kambhampati, S. Explainable Human-AI Interaction: A Planning Perspective, 2024, [arXiv:cs.AI/2405.15804]. [CrossRef]
- Ayub, A.; De Francesco, Z.; Mehta, J.; Yaakoub Agha, K.; Holthaus, P.; Nehaniv, C.L.; Dautenhahn, K. A Human-Centered View of Continual Learning: Understanding Interactions, Teaching Patterns, and Perceptions of Human Users Toward a Continual Learning Robot in Repeated Interactions. J. Hum.-Robot Interact. 2024, 13. [Google Scholar] [CrossRef]
- jakdot. pyactr. Technical report, 2025.
- navel robotics GmbH. navel. Technical report, 2025.







Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).