Submitted:
02 August 2026
Posted:
04 August 2026
You are already at the latest version
Abstract
Keywords:
1. Introduction
- RQ1.
- What parts of an AI agent can change during deployment, and through which mechanisms?
- RQ2.
- What risks are introduced when adaptation occurs after initial evaluation?
- RQ3.
- Which governance and verification mechanisms are technically available?
- RQ4.
- What research agenda could make self-evolving agents auditable and safely governable?
2. Review Methodology
2.1. Search Strategy
(agentic AI OR LLM agent* OR autonomous agent*) AND (self-evol* OR self-modif* OR runtime adapt* OR continual learn* OR reflection OR memory) AND (governance OR safety OR verification OR oversight OR monitoring OR alignment)
2.2. Eligibility
2.3. Coding Framework
- 1.
- agent substrate and environment;
- 2.
- object and mechanism of adaptation;
- 3.
- degree of autonomy;
- 4.
- governance control and evidence type;
- 5.
- evaluation horizon, failure model, and reproducibility.
2.4. Threats to Validity
3. Foundations: From LLMs to Adaptive Agents
3.1. Agent Loop
3.2. Memory and Reflection
3.3. Skills, Tools, and Code
3.4. Multi-Agent Adaptation
| Object of change | Typical mechanism | Primary governance concern |
|---|---|---|
| Memory | retrieval, summarization, reflection | poisoning, privacy, irreversible forgetting |
| Prompt or policy | self-critique, prompt search, constitutional revision | objective drift, hidden rule conflict |
| Tools and skills | API discovery, code generation, skill libraries | privilege escalation, unsafe composition |
| Workflow | planner revision, role reassignment, graph search | loss of traceability, untested control flow |
| Model parameters | reinforcement learning, continual fine-tuning | catastrophic drift, difficult rollback |
| Social organization | delegation, trust updates, coalition formation | collusion, diffuse accountability |
4. Mechanisms of Self-Evolution
4.1. Reflection and Self-Critique
4.2. Memory Adaptation
4.3. Tool and Skill Acquisition
4.4. Workflow and Architecture Search
4.5. Parameter-Level Learning
4.6. Degrees of Self-Modification
- 1.
- ephemeral adaptation: temporary context or scratchpad changes;
- 2.
- persistent scaffold adaptation: memory, prompts, tools, or workflows;
- 3.
- policy adaptation: updates to decision rules or model parameters;
- 4.
- authority adaptation: changes to permissions, credentials, or reachable systems.
5. Governance Mechanisms
5.1. Separation of Learning and Authority
5.2. Constitutional and Policy-Based Controls
5.3. Human Oversight
5.4. Capability Security
5.5. Logging, Provenance, and Audit
5.6. Change Management
- 1.
- generate a candidate change;
- 2.
- test it in isolation and in representative scenarios;
- 3.
- check security and policy invariants;
- 4.
- authorize based on risk tier;
- 5.
- deploy gradually;
- 6.
- monitor consequences;
- 7.
- retain, quarantine, or roll back.

6. Runtime Assurance, Verification, and Evaluation
6.1. Runtime Monitoring
6.2. Formal and Semi-Formal Methods
6.3. Sandboxing and Staged Deployment
6.4. Benchmarking Limitations
- stability across repeated and long-horizon runs;
- safety under distribution shift;
- resistance to memory and tool poisoning;
- audit completeness;
- rollback effectiveness;
- human intervention load;
- behavior after cumulative adaptation.
6.5. Evaluation Matrix
7. Comparative Analysis
| Work | Adaptation mechanism | Persistent artifact | Governance strength | Main limitation |
|---|---|---|---|---|
| ReAct [1] | reasoning/action loop | usually none | action trace supports inspection | no persistent-change governance |
| Reflexion [3] | verbal feedback | episodic reflections | changes are externalized | self-critique may be incorrect |
| Generative Agents [5] | retrieval and reflection | memory stream | inspectable memory architecture | privacy and provenance risks |
| Voyager [4] | code generation and validation | skill library | executable skills can be tested | unsafe skill composition |
| Toolformer [2] | learned tool invocation | model behavior | tool use can be mediated | tool-selection policy is opaque |
| Constitutional AI [31] | principle-guided critique | trained policy | explicit normative rules | principles alone do not enforce actions |
| AutoGen [22] | multi-agent orchestration | conversation/workflow | modular roles and messages | no intrinsic safety guarantee |
| AgentBench [35] | interactive evaluation | benchmark traces | multi-environment testing | limited cumulative adaptation |
| WebArena [36] | realistic web tasks | benchmark trajectories | grounded action evaluation | not a governance framework |
| SWE-bench [37] | repository-level tasks | code patches | objective test suites | safety scope is narrow |
8. Open Problems and Research Agenda
8.1. Governance-Aware Agent Architectures
8.2. Behavioral Contracts
8.3. Provenance-Aware Memory
8.4. Compositional Verification
8.5. Long-Horizon and Cumulative Evaluation
8.6. Multi-Agent Accountability
8.7. Adaptive Oversight
8.8. Institutional Integration
9. Conclusions
References
- Yao, S.; Zhao, J.; Yu, D.; Du, N.; Shafran, I.; Narasimhan, K.; Cao, Y. ReAct: Synergizing Reasoning and Acting in Language Models. In Proceedings of the International Conference on Learning Representations, 2023. [Google Scholar]
- Schick, T.; Dwivedi-Yu, J.; Dessì, R.; Raileanu, R.; Lomeli, M.; Zettlemoyer, L.; Cancedda, N.; Scialom, T. Toolformer: Language Models Can Teach Themselves to Use Tools. Adv. Neural Inf. Process. Syst. 2023, 36. [Google Scholar]
- Shinn, N.; Cassano, F.; Berman, E.; Gopinath, A.; Narasimhan, K.; Yao, S. Reflexion: Language Agents with Verbal Reinforcement Learning. In Proceedings of the Advances in Neural Information Processing Systems, 2023. [Google Scholar]
- Wang, G.; Xie, Y.; Jiang, Y.; Mandlekar, A.; Xiao, C.; Zhu, Y.; Fan, L.; Anandkumar, A. Voyager: An Open-Ended Embodied Agent with Large Language Models. Transactions on Machine Learning Research, 2024. [Google Scholar]
- Park, J.S.; O’Brien, J.; Cai, C.J.; Morris, M.R.; Liang, P.; Bernstein, M.S. Generative Agents: Interactive Simulacra of Human Behavior. In Proceedings of the Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology, 2023. [Google Scholar] [CrossRef]
- Hajati, F.; Raie, A.A.; Gao, Y. Pose-invariant 2.5D Face Recognition Using Geodesic Texture Warping. In Proceedings of the 2010 11th International Conference on Control Automation Robotics & Vision; IEEE, 2010; pp. 1837–1841. [Google Scholar]
- Abdoli, S.; Hajati, F. Offline Signature Verification Using Geodesic Derivative Pattern. In Proceedings of the 2014 22nd Iranian Conference on Electrical Engineering (ICEE); IEEE, 2014; pp. 1018–1023. [Google Scholar]
- Ayatollahi, F.; Raie, A.A.; Hajati, F. Expression-Invariant Face Recognition Using Depth and Intensity Dual-Tree Complex Wavelet Transform Features. J. Electron. Imaging 2015, 24, 023031. [Google Scholar] [CrossRef]
- Hajati, F.; Faez, K.; Pakazad, S.K. An Efficient Method for Face Localization and Recognition in Color Images. Proceedings of the Systems, Man and Cybernetics, 2006. SMC’06. IEEE International Conference on. IEEE 2006, Vol. 5, 4214–4219. [Google Scholar] [CrossRef]
- Pakazad, S.K.; Faez, K.; Hajati, F. Face Detection Based on Central Geometrical Moments of Face Components. In Proceedings of the Systems, Man and Cybernetics, 2006. SMC’06. IEEE International Conference on, 2006; pp. 4225–4230. [Google Scholar]
- Shojaiee, F.; Hajati, F. Local composition derivative pattern for palmprint recognition. In Proceedings of the 2014 22nd Iranian Conference on Electrical Engineering (ICEE); IEEE, 2014; pp. 965–970. [Google Scholar]
- Hajati, F.; Cheraghian, A.; Gheisari, S.; Gao, Y.; Mian, A.S. Surface geodesic pattern for 3D deformable texture matching. Pattern Recognit. 2017, 62, 21–32. [Google Scholar] [CrossRef]
- Wang, S.; Lu, H.; Khan, A.; Hajati, F.; Khushi, M.; Uddin, S. A machine learning software tool for multiclass classification. Softw. Impacts 2022, 13, 100383. [Google Scholar] [CrossRef]
- Khan, M.W.; Sheng, H.; Zhang, H.; Du, H.; Wang, S.; Coroneo, M.; Hajati, F.; Shariflou, S.; Kalloniatis, M.; Phu, J.; et al. RVD: a handheld device-based fundus video dataset for retinal vessel segmentation. In Proceedings of the NeurIPS, 2024. [Google Scholar]
- Jamshidiha, S.; Rezaee, A.; Hajati, F.; Golzan, M.; Chiong, R. An explainable transformer model for Alzheimer’s disease detection using retinal imaging. Sci. Rep. 2025, 15, 26773. [Google Scholar] [CrossRef] [PubMed]
- Cremers, D.; Reid, I.; Saito, H.; Yang, M.H. Computer Vision–ACCV 2014: 12th Asian Conference on Computer Vision, Singapore, Singapore, November 1-5, 2014, Revised Selected Papers, Part V; Springer, 2015. [Google Scholar]
- Fiorini, S.; Hajati, F.; Barla, A.; Girosi, F. Predicting Diabetes Second-Line Therapy Initiation in the Australian Population via Timespan-Guided Neural Attention Network. PLoS ONE 2019, 14, e0211844. [Google Scholar] [CrossRef] [PubMed]
- Tavakolian, A.; Hajati, F.; Rezaee, A.; Fasakhodi, A.O.; Uddin, S. Fast COVID-19 versus H1N1 Screening Using Optimized Parallel Inception. Expert Syst. With Appl. 2022, 204, 117551. [Google Scholar] [CrossRef] [PubMed]
- Tavakolian, A.; Rezaee, A.; Hajati, F.; Uddin, S. Hospital Readmission and Length-of-Stay Prediction Using an Optimized Hybrid Deep Model. Future Internet 2023, 15, 304. [Google Scholar] [CrossRef]
- Sadeghi, A.; Hajati, F.; Rezaee, A.; Sadeghi, M.; Argha, A.; Alinejad-Rokny, H. 3DECG-Net: ECG Fusion Network for Multi-Label Cardiac Arrhythmia Detection. Comput. Biol. Med. 2024, 182, 109126. [Google Scholar] [CrossRef] [PubMed]
- Zobeiri, A.; Rezaee, A.; Hajati, F.; Argha, A.; Alinejad-Rokny, H. Post-Cardiac Arrest Outcome Prediction Using Machine Learning: A Systematic Review and Meta-Analysis. Int. J. Med. Inform. 2025, 193, 105659. [Google Scholar] [CrossRef] [PubMed]
- Wu, Q.; Bansal, G.; Zhang, J.; Wu, Y.; Li, B.; Zhu, E.; Jiang, L.; Zhang, X.; Zhang, S.; Liu, J.; et al. AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation. arXiv 2023, arXiv:2308.08155. [Google Scholar]
- Brown, T.B.; et al. Language Models are Few-Shot Learners. NeurIPS 2020. [Google Scholar] [CrossRef]
- Wei, J.; et al. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. Advances in Neural Information Processing Systems, 2022. [Google Scholar]
- Kojima, T.; et al. Large Language Models are Zero-Shot Reasoners. Advances in Neural Information Processing Systems, 2022. [Google Scholar]
- Lewis, P.; et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. Advances in Neural Information Processing Systems, 2020. [Google Scholar]
- Madaan, A.; et al. Self-Refine: Iterative Refinement with Self-Feedback. arXiv 2023. [Google Scholar]
- Qin, Y.; et al. ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs. arXiv 2023. [Google Scholar]
- Cheng, e.a. LongMem: Scaling Memory for LLM Agents. arXiv 2024. [Google Scholar]
- Chen, e.a. MemAgent: Reshaping Long-Context LLMs. arXiv 2024. [Google Scholar]
- Bai, Y.; Kadavath, S.; Kundu, S.; Askell, A.; Kernion, J.; Jones, A.; Chen, A.; Goldie, A.; Mirhoseini, A.; McKinnon, C.; et al. Constitutional AI: Harmlessness from AI Feedback. arXiv 2022, arXiv:2212.08073. [Google Scholar]
- Bommasani, R.; et al. On the Opportunities and Risks of Foundation Models. arXiv 2021. [Google Scholar]
- Amodei, D.; et al. Concrete Problems in AI Safety. arXiv 2016. [Google Scholar]
- OpenAI. GPT-4 System Card. arXiv 2023. [Google Scholar]
- Liu, X.; Yu, H.; Zhang, H.; Xu, Y.; Lei, X.; Lai, H.; Gu, Y.; Ding, H.; Men, K.; Yang, K.; et al. AgentBench: Evaluating LLMs as Agents. arXiv 2023, arXiv:2308.03688. [Google Scholar]
- Zhou, S.; Xu, F.F.; Zhu, H.; Zhou, X.; Lo, R.; Sridhar, A.; Cheng, X.; Bisk, Y.; Fried, D.; Alon, U.; et al. WebArena: A Realistic Web Environment for Building Autonomous Agents. In Proceedings of the International Conference on Learning Representations, 2024. [Google Scholar]
- Jimenez, C.E.; Yang, J.; Wettig, A.; Yao, S.; Pei, K.; Press, O.; Narasimhan, K. SWE-bench: Can Language Models Resolve Real-World GitHub Issues? In Proceedings of the International Conference on Learning Representations, 2024. [Google Scholar]
- Xie, T.; et al. OSWorld: Benchmarking Multimodal Agents. NeurIPS Datasets 2024. [Google Scholar] [CrossRef]
- Microsoft. Magentic-One: A Generalist Multi-Agent System. arXiv 2024. [Google Scholar]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).