Submitted:
27 August 2026
Posted:
28 August 2026
You are already at the latest version
Abstract
Generalist agents are evolving from task-oriented systems into autonomous, self-evolving systems that acquire reusable skills, maintain long-term memory, interact with external environments, and refine themselves through recursive workflows. This shift creates a lifecycle security challenge beyond conventional model-centric or component-wise perspectives: vulnerabilities may arise from acquired external capabilities, be amplified through the orchestration of skills, memory, and tools, materialize as consequential actions, and persist through feedback across future evolution. Security for generalist agents is therefore inherently a lifecycle problem rather than a collection of isolated component- level risks. Motivated by these observations, we present a lifecycle survey organized around three stages: provenance, orchestration, and execution. We trace how risks originate, propagate, accumulate over long-horizon interactions, and persist through feedback. Across this lifecycle, we map four coupled attack surfaces: skill supply chains, user inputs, long-term memory, and external environments, and organize evaluations and defenses by where safety evidence is collected and interventions occur. Evaluations span interactive execution and offline trajectory auditing, while defenses cover input and context filtering, decision and control integrity, and runtime monitoring and enforcement. Finally, motivated by the growing importance of recursive and self-evolving AI, we identify adversarial risk generation, fine-grained safety attribution, and adversarial-feedback-driven continual safety evolution as three connected directions that form a closed loop of risk discovery, diagnosis, and mitigation toward safe recursive self-improvement agents. The latest papers and repositories are maintained at Pandora-Agent.
Keywords:
generalist agent
; agent security
; self-evolving agent
; attack and defense
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.