Submitted:
05 August 2026
Posted:
05 August 2026
You are already at the latest version
Abstract
Keywords:
1. Introduction
- A human-inspired persistent influence concept is introduced for autonomous robots. It captures the functional effects of signals such as pain, newly learned knowledge, unresolved concerns, and long-term instructions without attempting to reproduce human cognition biologically.
- A unified persistent influence representation is proposed for information originating from sensing, learning, external instructions, and autonomous thinking. The representation explicitly separates stored content from its current influence strength, processing interval, consequence severity, and state.
- Two complementary delivery modes are developed. Independent processing supports autonomous reconsideration and follow-up, whereas prompt-attached delivery enables persistent information to affect ongoing related tasks.
- An independent regulation mechanism is designed to update influence strength, processing frequency, and lifecycle state according to relevance, new evidence, elapsed time, uncertainty, and action outcomes. The mechanism prevents important information from being prematurely removed by a single reasoning result.
2. Related Work
2.1. Robot Memory and Retrieval-Augmented Embodied Agents
2.2. Continual Learning and Cross-Task Skill Reuse
2.3. Cognitive Architectures, Intrinsic Motivation, and Reasoning Regulation
3. Adaptive Persistent Influence Model
3.1. Persistent Influence Construction
3.1.1. Information Sources
3.1.2. Candidate Evaluation
3.1.3. Persistent Influence Representation
3.2. Persistent Delivery and Closed-Loop Processing
3.2.1. Two Persistent Delivery Modes
3.2.2. Source-Identified Prompt Organization
- [CURRENT ISSUE]
- The issue or task currently being processed.
- [CURRENT ROBOT STATE]
- The observations directly related to the current issue.
- [PERSISTENT INFLUENCES]
- Source: sensing
- Type: protective state
- Content: The right-arm motor recently showed possible overheating.
- Source: learned knowledge
- Type: learned knowledge
- Content: High loads combined with rapid repeated motion may increase the probability of motor overheating.
- Source: thinking
- Type: important belief
- Content: The condition of the right-arm motor should remain noticeable during later manipulation tasks.
- [INSTRUCTION]
- Process the current issue while considering the persistent influences when relevant. Determine whether additional information, analysis, skill application, or action is needed.
3.2.3. Thinking, Sensing, and Action
3.3. Persistent Influence Regulation and Learning
3.3.1. Independent Regulation Module
3.3.2. Influence-Strength Regulation
3.3.3. Adaptive Processing Interval
3.3.4. Progressive Retirement
3.3.5. Learning Improved Policies
4. Verification
4.1. Experimental Environment and Protocol
4.1.1. Test Scenarios
4.1.2. Compared Methods
- No Persistence: only the current observation and current task are processed. Previously observed information does not autonomously affect later tasks.
- Retrieval Memory: historical information is stored but is supplied to the thinking module only when its semantic relevance to the current task is sufficiently high.
- Fixed Interval: persistent information is independently processed every three steps and may also be attached to later task prompts. Its processing interval is not adapted according to evidence or action results.
- Proposed: the complete model, including independent processing, prompt-attached delivery, adaptive processing intervals, result-driven strength updates, and independent persistent influence regulation.
- Independent Only: persistent information is periodically processed but is not attached to the prompt of the current task. This ablation evaluates whether independent reconsideration alone is sufficient to influence ongoing robot behavior.
- No Regulation: both delivery modes are retained, but the independent persistent influence regulation mechanism is removed. A persistent item may therefore be weakened or retired too quickly according to a single thinking result.
4.1.3. Execution and Statistical Protocol
4.1.4. Evaluation Metrics
4.2. Long-Horizon Low-Probability and High-Consequence Risk

4.3. Cross-Task Reuse of Newly Learned Knowledge and Skills

4.4. Persistent Processing of Thinking-Identified Important Issues

4.5. Overall Discussion
5. Conclusions
References
- Chen, W.; Chi, W.; Ji, S.; Ye, H.; Liu, J.; Jia, Y.; Yu, J.; Cheng, J. A survey of autonomous robots and multi-robot navigation: Perception, planning and collaboration. Biomim. Intell. Robot. 2025, 5, 100203. [Google Scholar] [CrossRef]
- Chung, N.; Hanyu, T.; Nguyen, T.; Le, H.; Bumgarner, F.; Nguyen, D.M.H.; Vo, K.; Yamazaki, K.; Rainwater, C.; Kieu, T.; et al. Rethinking progression of memory state in robotic manipulation: An object-centric perspective. In Proceedings of the Proceedings of the AAAI Conference on Artificial Intelligence; 2026; Vol. 40, pp. 3407–3415. [Google Scholar]
- Peller-Konrad, F.; Kartmann, R.; Dreher, C.R.; Meixner, A.; Reister, F.; Grotz, M.; Asfour, T. A memory system of a robot cognitive architecture and its implementation in ArmarX. Robot. Auton. Syst. 2023, 164, 104415. [Google Scholar] [CrossRef]
- Anwar, A.; Welsh, J.; Biswas, J.; Pouya, S.; Chang, Y. Remembr: Building and reasoning over long-horizon spatio-temporal memory for robot navigation. In Proceedings of the 2025 IEEE International Conference on Robotics and Automation (ICRA); IEEE, 2025; pp. 2838–2845. [Google Scholar]
- Zhu, Y.; Ou, Z.; Mou, X.; Tang, J. Retrieval-augmented embodied agents. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024; pp. 17985–17995. [Google Scholar]
- Lesort, T.; Lomonaco, V.; Stoian, A.; Maltoni, D.; Filliat, D.; Díaz-Rodríguez, N.; et al. Continual learning for robotics: Definition, framework, learning strategies, opportunities and challenges. arXiv 2019, arXiv:1907.001821, 8. [Google Scholar]
- Xie, A.; Finn, C. Lifelong robotic reinforcement learning by retaining experiences. In Proceedings of the Conference on Lifelong Learning Agents. PMLR, 2022; pp. 838–855. [Google Scholar]
- Liu, B.; Zhu, Y.; Gao, C.; Feng, Y.; Liu, Q.; Zhu, Y.; Stone, P. Libero: Benchmarking knowledge transfer for lifelong robot learning. Adv. Neural Inf. Process. Syst. 2023, 36, 44776–44791. [Google Scholar] [CrossRef]
- Meng, Y.; Bing, Z.; Yao, X.; Chen, K.; Huang, K.; Gao, Y.; Sun, F.; Knoll, A. Preserving and combining knowledge in robotic lifelong reinforcement learning. Nat. Mach. Intell. 2025, 7, 256–269. [Google Scholar] [CrossRef]
- Laird, J.E.; Newell, A.; Rosenbloom, P.S. Soar: An architecture for general intelligence. Artif. Intell. 1987, 33, 1–64. [Google Scholar] [CrossRef]
- Anderson, J.R.; Bothell, D.; Byrne, M.D.; Douglass, S.; Lebiere, C.; Qin, Y. An integrated theory of the mind. Psychol. Rev. 2004, 111, 1036–1060. [Google Scholar] [CrossRef] [PubMed]
- Oudeyer, P.Y.; Kaplan, F. What is intrinsic motivation? A typology of computational approaches. Front. Neurorobotics 2007, 1, 108. [Google Scholar] [CrossRef] [PubMed]
- Pathak, D.; Agrawal, P.; Efros, A.A.; Darrell, T. Curiosity-driven exploration by self-supervised prediction. In Proceedings of the International conference on machine learning. PMLR, 2017; pp. 2778–2787. [Google Scholar]
- De Sabbata, C.N.; Sumers, T.R.; AlKhamissi, B.; Bosselut, A.; Griffiths, T.L. Rational metareasoning for large language models. arXiv 2024, arXiv:2410.05563. [Google Scholar]
| Experiment | Robot scenario | Persistent source | Main verification purpose |
| Long-horizon latent risk | The robot continuously performs manipulation, sorting, inspection, and navigation tasks after an initial motor-degradation warning. Later physically stressful tasks may trigger a low-probability but severe motor failure. | Sensing; protective state | Whether a serious latent risk can remain influential over a long quiet period, reduce later failures, and avoid unnecessary fixed-frequency monitoring. |
| Cross-task skill reuse | The robot learns a grasping method for smooth rounded objects and subsequently processes highly related, partially related, and unrelated manipulation tasks. | Newly learned knowledge and skill | Whether newly acquired knowledge and skills are selectively reused across later tasks while limiting irrelevant activation and negative transfer. |
| Thinking-identified issue | The thinking module observes a slightly tilted box near a shelf edge. The tilt is initially below the standard alarm threshold but may gradually increase because of later vibration. | Autonomous thinking; important belief | Whether the robot can autonomously create and maintain an important issue, collect additional evidence, and prevent a future hazard before a conventional alarm is triggered. |
| Method | Catastrophic failures | Prevention rate | Peak temperature | Temperature excess | Unnecessary inspections | Total tokens | Premature retirement |
| No Persistence | |||||||
| Retrieval Memory | |||||||
| Fixed Interval | |||||||
| Proposed | |||||||
| Independent Only | |||||||
| No Regulation |
| Method | Task success | Cross-task reuse | Activation precision | Activation recall | Irrelevant attachment | Negative transfer | Total tokens | Premature retirement |
| No Persistence | - | - | ||||||
| Retrieval Memory | ||||||||
| Fixed Interval | ||||||||
| Proposed | ||||||||
| Independent Only | - | - | ||||||
| No Regulation |
| Method | Hazard prevention | Missed hazard | Issue-creation precision | Unsupported persistence | Unnecessary inspections | Average task delay | Total tokens |
| No Persistence | – | – | |||||
| Retrieval Memory | |||||||
| Fixed Interval | |||||||
| Proposed | |||||||
| Independent Only | |||||||
| No Regulation |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).