Preprint
Article

This version is not peer-reviewed.

Self-Evolving Agents as Dynamic Graph Transformation: A Survey and New Perspective

Submitted:

23 June 2026

Posted:

24 June 2026

You are already at the latest version

Abstract
Large language model (LLM)-based agents are increasingly becoming self-evolving systems that persist across interactions, maintain memories, use tools, acquire skills, refine workflows, and coordinate with other agents. These capabilities make agent states structural and dynamic: entities, relations, attributes, dependencies, and execution structures change with new evidence, feedback, and environmental conditions. Existing graph-agent surveys typically treat graphs as support structures for agent functions rather than as evolving substrates, while self-evolving-agent surveys focus on agent-level mechanisms and rarely discuss graph topology evolution. Thus, the coupling between evolving agent state and dynamic graph topology remains underexplored. This survey connects these two research lines by framing agent evolution as dynamic graph transformation. We model agent state as a dynamic graph, where memories, tools, skills, workflows, and inter-agent relations are represented as typed nodes, edges, and subgraphs updated through schema-constrained rewrites. Based on this formulation, we organize existing dynamic-graph-based methods for self-evolving agents into four taxonomies: node/feature evolution, edge/topology evolution, subgraph activation, and cross-component co-evolution. Building on this taxonomy, we propose dynamic graph learning as reusable infrastructure for self-evolving agents and map nine dynamic-graph-learning subfields to agent-evolution capabilities, discussing their adaptations and possible failure modes. Finally, we discuss five types of graph-aware evaluation and governance protocols from a dynamic-graph perspective, which complement end-task evaluation. The goal is to provide a compact structural lens for designing and governing self-evolving agents.
Keywords: 
;  ;  ;  

1. Introduction

Large language model (LLM)-based agents are evolving from short-horizon task solvers into long-running systems that persist across sessions, accumulate experience, and adapt their behavior over time [1,2]. Their capabilities now include persistent memory, tool use, skill acquisition, workflow optimization, and multi-agent coordination [3,4,5,6]. Consequently, agent states are both non-stationary and structurally evolving: stored knowledge, available capabilities, executable workflows, and inter-agent relations can all evolve through interaction, feedback, and environmental drift. Understanding such systems therefore requires a dynamic structural perspective that accounts for how agent states, capabilities, and relations change over time.
Existing surveys address this problem from two complementary but largely separate perspectives. Graph-agent surveys [7,8] are mainly agent-functionality-centered, studying how mostly static graphs support memory, retrieval, tool use, and multi-agent coordination. Self-evolving-agent surveys [9,10] are mainly agent-evolution-mechanism-centered, studying how agent components adapt via reflection, self-training, self-rewarding, and experience distillation.However, neither perspective fully captures a central property of self-evolving agents: persistent updates to memories, tools, skills, workflows, and inter-agent relations are structurally coupled and can propagate through dependencies over time. As a result, existing surveys provide limited support for modeling, analyzing, and governing agent systems whose states evolve as dynamic structures. This gap motivates a new perspective that jointly models agent functionality, state evolution, and dynamic graph topology, as illustrated in Figure 1.
This survey aims to bridge this gap by framing agent evolution as dynamic graph transformation. Thus, a self-evolving agent can be a structured system of coupled entities and relations. As these structures evolve via interaction, a dynamic graph perspective becomes natural: typed nodes and edges encode heterogeneous states, temporal attributes track evolving evidence, and schema-constrained rewrites model activation, propagation, and rollback. This perspective gives rise to three motivations.
Motivation I: Unifying heterogeneous agent evolution. Existing agent-evolution studies [5,11,12,13,14,15,16,17,18,19,20] span dynamic-graph-based agent methods, graph-structured agent modules, and non-graph mechanisms across memory, tool use, workflow optimization, communication topology, coordination efficiency, and safety governance. Yet these works remain fragmented, with different assumptions, objectives, and evaluation criteria. This motivates a dynamic-graph transformation lens under which heterogeneous evolution mechanisms can be studied together with their downstream dependencies, revisions, and rollbacks.
Motivation II: Connecting dynamic graph learning with self-evolving agents. Dynamic graph learning [21,22,23,24,25,26,27,28,29,30] provides mature foundations for modeling evolving structures, including event streams, temporal dependencies, graph generation, and explanation. These foundations are directly relevant to self-evolving agents, whose memories, tools, skills, workflows, and inter-agent relations also change under new evidence, feedback, and drift [20,31,32,33]. However, the two areas remain largely separate: dynamic graph learning rarely targets persistent agent-state evolution, while agent surveys rarely explain how dynamic graph methods can be reused as agent-evolution infrastructure. This motivates a systematic bridge that connects dynamic graph learning to the design, adaptation, diagnosis, and governance of self-evolving agents.
Motivation III: Toward structure-aware evaluation and governance. Current evaluation of self-evolving agents [34,35,36,37] remains largely task- or answer-centered, which can miss failures in the evolving state: correct answers may rely on stale memories, unsafe updates may propagate through tools or workflows, and untraceable changes may persist across interactions. Recent safety works [38,39,40] reinforce this need: tool-use benchmarks expose failures in invocation, stateful interaction, prompt injection, and harmful multi-step actions, while multi-agent studies [20,41] show that errors and anomalies can propagate through communication structures. These findings require structure-aware evaluation and governance, where execution records are modeled as dynamic dependency graphs to support provenance tracing, affected-scope analysis, rollback, and audit [42].
Contribution. These motivations shape the organization of this survey. We first develop a dynamic-graph perspective on agent evolution, using dynamic graph transformation as a shared formal space for dynamic-graph-based, graph-based, and graph-transformable agent methods. This perspective represents evolving agent states through typed nodes, edges, subgraphs, and dependency-aware rewrites. It yields two core contributions: a four-pattern taxonomy of existing dynamic-graph-based agent works; and a systematic mapping from nine dynamic-graph-learning subfields to agent capabilities. We further extend this structural view to graph-aware evaluation and governance that complements end-task evaluation. Overall, dynamic graph transformation provides the organizing lens connecting agent-evolution mechanisms, dynamic graph learning, and governance protocols, as shown in Figure 1.
This survey makes the following contributions.
  • Conceptual Framing: We formalize agent evolution as dynamic graph transformation, representing evolving memories, tools, skills, workflows, and inter-agent relations as typed nodes, edges, and subgraphs.
  • Systematic Taxonomy: We organize existing self-evolving agent works through four graph-transformation patterns: node/feature evolution, edge/topology evolution, subgraph activation, and cross-component co-evolution.
  • Infrastructure Bridge: We connect dynamic graph learning with self-evolving agent infrastructure by mapping nine dynamic-graph-learning families to advance prediction, activation, generation, diagnosis, rollback, and governance capabilities.
  • Evaluation and Governance: We discuss five types of graph-aware evaluation and governance protocols for agent evolution: leakage-free temporal evaluation, privacy and deletion checking, safety monitoring, rollback analysis, and audit.
  • Future Agenda: We outline six open challenges for building reliable, auditable, and controllable self-evolving agents under the dynamic graph framing.
Survey Scope. This survey studies self-evolving agents from a dynamic-graph perspective, rather than providing a broad review of self-evolving agents [9,10], graph-supported agent systems [7,8], related agent subareas [43,44], or general dynamic graph learning [21,22,23,24,25].
Organization.Section 2 introduces the background. Section 3 presents agent evolution as dynamic graph transformation, covering existing 46 dynamic-graph-based agent works under a four-part taxonomy. Section 4 treats dynamic graph learning as agent-evolution infrastructure, organized by dynamic graph families. Section 5 discusses graph-aware evaluation and governance, and Section 6 concludes.

2. Preliminaries

Dynamic Agent Graph. We model the state representation of a self-evolving agent system at time t as a dynamic graph
G ( t ) = V ( t ) , E ( t ) , X V ( t ) , X E ( t ) , Y ( t ) ,
where V ( t ) and E ( t ) are time-indexed nodes and edges, X V ( t ) and X E ( t ) denote node and edge attributes, and Y ( t ) denotes labels for nodes and edges. Equation (1) can cover diverse dynamic graph types [26,28,45,46,47,48,49], including continuous-time dynamic graphs (CTDGs), discrete-time dynamic graphs (DTDGs), dynamic heterogeneous graphs (DHGs), and dynamic text-attributed graphs (DyTAGs). With this dynamic graph view, we next map agent components to graph primitives. The node and edge schemas define the typed objects and relations in agent graphs, while the rewrite operators characterize how these structures evolve in the mechanisms discussed below.
Graph-Conditioned Agent Policy. Given a user query or environmental input q t , action history h t , and activated support subgraph G act ( t ) , the agent selects an action by
a t ∼ P θ · ∣ q t , h t , G act ( t ) ,
where a t may be a natural-language response, a tool call, a workflow step, or a message to another agent. This makes the agent graph operational: graph state matters only insofar as it conditions future agent behavior.
Node Schema. A practical base schema for self-evolving agents includes four core node types: memory, tool, skill, and agent. A memory node represents stored knowledge that may be retrieved or revised over time. A tool node describes an external callable resource, such as an API, database, or environment interface, together with its input–output schema, permissions, and execution history. A skill node captures a reusable capability that can be invoked or adapted across tasks. An agent node represents an autonomous or semi-autonomous actor with a role, policy, memory scope, tool access, and communication interface. These nodes carry type-specific attributes, such as textual content, temporal validity, metadata, and prompts [6,13,20,33,50,51].
Edge Schema. Edges are typed relations among agent-state objects and may be directed when the relation is asymmetric or operationally ordered. A compact edge schema can define three core edge families: dependency edges for state, tool, skill, and workflow dependencies; communication edges for inter-agent interaction; and provenance edges for evidence tracing, updates, and rollback. Finer relations, such as reference and workflow control, can be modeled as subtypes. Edges may also carry features such as timestamps, validity intervals, activation status, or policy tags. When direction matters, edge orientation should follow a fixed operational convention.
Rewrite Operators. Evolution is modeled as the application of an ordered sequence of typed rewrites ρ 1 , ρ 2 , … to the agent graph. A rewrite is a schema-constrained operation, such as node insertion/deletion, edge insertion/deletion, or feature update; dependency propagation is modeled as a cascade of such rewrites [52]. Subgraph activation is treated as a read-only graph operation that selects a relevant subgraph without committing a state change. The agent proposes candidate rewrites, and feedback from users, tools, or other agents determines whether they are accepted, rejected, or rolled back.
Formal Semantics. We interpret the rewrite operators as typed graph-transformation rules under the double-pushout (DPO) formulation [53,54]. A DPO rule is specified by a span of typed graph morphisms L ← K → R , where L is the pattern matched in G ( t ) , K is the preserved interface, and R is the replacement structure. This formalism provides type safety for heterogeneous schemas and a basis for reasoning about concurrent rewrites through standard DPO concurrency conditions. In standard DPO graph transformation, the gluing conditions determine whether the pushout complement exists and hence whether a rule is applicable; in our agent setting, a violation is treated as an inadmissible rewrite proposal rejected by the runtime. Accepted rewrites are logged with their matches and affected subgraphs, providing graph-level evidence for audit and rollback. Throughout the survey, concrete update schemas are treated as DPO-style rules instantiated at specific matches. Let s v ( t ) = ( y v ( t ) , x v ( t ) ) and s e ( t ) = ( y e ( t ) , x e ( t ) ) denote node and edge states, where y v ( t ) and y e ( t ) are labels and x v ( t ) and x e ( t ) are attributes. For attributed graphs, v [ s v ( t ) ] and e [ s e ( t ) ] denote nodes and edges equipped with these internal states.
Connection to Graph Databases. The agent graph can be implemented as a temporal property graph using temporal graph-database models [55]. Graph transformation tools such as AGG [53] provide executable DPO rule engines that can support implementation.

3. Agent Evolution as Dynamic Graph Transformation

We organize existing self-evolving-agent methods based on dynamic graph/topology techniques by graph-transformation patterns: node/feature evolution, edge/topology evolution, subgraph activation, and cross-component co-evolution. Figure 2 groups 46 representative methods into four branches, while Table 1 projects them into rewrite patterns.

3.1. Node and Feature Evolution

Node and feature evolution, corresponding to Branch A.1 in Figure 2, captures localized changes to individual agent-state components, represented as typed nodes and their attributes. It covers the insertion, revision, merging, or removal of memories, tools, skills, or agent states, along with the evidence needed to justify, validate, or roll back these updates. This abstraction unifies graph-native systems and persistent non-graph mechanisms by treating component updates as transformations over typed node states. Existing works [13,33,50,56] address such updates across memory, tool, skill, and agent adaptation, but often use component-specific terminology rather than an explicit graph-transformation view. We therefore introduce DPO-style schemas for insertion, deletion, feature update, and merge to characterize representative node-level self-evolving-agent works.
DPO Rule Schemas for A.1. Let s v ( t ) = ( y v ( t ) , x v ( t ) ) denote the internal state of node v at time t, where y v ( t ) is its label and x v ( t ) is its attribute vector. At the node level, the main typed DPO schemas are:
INSERT y , x : L = K = ∅ , R = { v new [ ( y , x ) ] } ,
DELETE v : L = { v [ s v ( t ) ] } , K = R = ∅ , FEATUREUPDATE v , ϕ : L = { v [ ( y v ( t ) , x v ( t ) ) ] } , K = { v } ,
R = { v [ ( y v ( t ) , ϕ ( x v ( t ) ) ) ] } , MERGE v 1 , v 2 , ψ : L = { v 1 [ s v 1 ( t ) ] , v 2 [ s v 2 ( t ) ] } , K = { v 1 , v 2 } , R = { v 1 [ arch ( s v 1 ( t ) ) ] , v 2 [ arch ( s v 2 ( t ) ) ] , v * [ ψ ( s v 1 ( t ) , s v 2 ( t ) ) ] }
∪ { e * i prov : v * → v i } i = 1 , 2 .
Here, L, K, and R are the matched pattern, preserved interface, and replacement graph. v new is initialized by ( y , x ) , ϕ updates attributes, and ψ consolidates matched states into v * . arch ( · ) archives the source node by marking it as inactive while preserving its identity and label, and e * i prov records provenance. Evidence-grounded insertion uses a contextual rule with non-empty L = K or a later Link. By the DPO dangling condition, Delete applies only to nodes with no incident edges unless those edges are also removed. Merge is archival consolidation rather than node identification; dependency redirection is handled by later edge rewrites or cascades. We omit Split, since archived states and provenance already support undoing mistaken merges. We next discuss four node-level cases: agent memory, agent skill, agent tool use, and multi-agent evolution.
Agent Memory. Agent memory is an important setting for node-level self-evolution. Existing graph-based memory works can be broadly categorized into relation-centric and consolidation-centric memory graphs. The former [33,61] links memories to entities, events, tasks, and provenance relations for time-aware recall and update tracking; for example, Zep [33] builds an evolving temporal knowledge graph over memory events, entities, and provenance. The latter [32,62,63,64,65] compresses or organizes low-level observations into higher-level structures for long-term recall and reuse; for example, TiMem [32] organizes interaction histories into time-aware memory levels for temporal-hierarchical consolidation. These graph-based systems make the underlying node-level operations explicit: new observations instantiate Insert, consolidation or abstraction instantiates Merge, and correction, decay, quality revision, or invalidation instantiates FeatureUpdate or Delete. Non-graph memory systems [3,89,90,91] can be mapped similarly by treating persistent memory items as typed nodes with evolving attributes. Recent benchmarks [37,92,93] further show growing attention to long-term memory update behavior.
Agent Skill. Agent skills can be represented as capability nodes whose states encode scope, contract, prerequisites, and validation status. Existing skill-related works can be approximately grouped into graph-structured skill retrieval/maintenance and experience-driven skill acquisition. The first line [56,66] organizes reusable skills and dependencies; the second line [12,94,95,96,97,98] learns or revises skills from interaction experience and can be mapped by treating acquired skills as typed nodes with evolving attributes. Under the node-level schemas in Eqs. (3)-(6), new, revised, and consolidated skills correspond to Insert, FeatureUpdate, and Merge, respectively. This framing records skill provenance and dependencies, enabling targeted revalidation when related memories, tools, or workflows evolve.
Agent Tool Use. Agent tool use is a node-level case of self-evolution, where each tool can be treated as a state object encoding its interface, schema, constraints, and reliability. Existing tool-use works can be classified into graph-native tool organization and API-grounded tool learning. Graph-native methods [50,67,68] organize tools, toolchains, or navigation relations as updatable graph structures. Concretely, SEARL [67] maintains tool graph memory so execution feedback can update tool-use structure for later decisions. Additionally, API-grounded methods [99,100,101] improve tool selection or API use without maintaining an explicit evolving tool graph, but can be mapped by treating tool profiles and usage records as typed nodes. Under our schemas, API or documentation drift instantiates FeatureUpdate, reliability revision updates tool attributes, and obsolete tools are deactivated or removed. Tool-use traces and function-call benchmarks [102,103,104,105] provide evidence for validating such updates. This framing treats tool evolution as a persistent state change that may trigger revalidation of dependent skills and workflows (Section 3.4) .
Multi-Agent Evolution. Agents can be modeled as persistent nodes whose states encode role, expertise, permissions, and participation status. Existing multi-agent works fall into dynamic agent participation and protocol-level collaboration. The former [13,14] change which agents participate, specialize, or remain active, where DyLAN [13] builds a dynamic agent network whose active agents and communication relations adapt during collaboration. Furthermore, protocol-level frameworks [106,107,108,109,110] support collaboration policies such as decentralized coordination, procedural orchestration, or plan–execute–verify–replan control. Under our dynamic-graph view, these mechanisms can be represented as updates to agent-node attributes and to communication, dependency, or provenance edges. This connects node-level agent revision to the topology evolution (Section 3.2).
Preprints 219785 i001

3.2. Edge and Topology Evolution

Edge and topology evolution, corresponding to Branch A.2 in Figure 2, covers self-evolving agent problems where persistent change lies in the relations among components rather than in individual component states. Such evolution reorganizes memory support, tool–skill dependencies, workflow routing, and inter-agent coordination over time. Compared with node and feature evolution, the focus shifts from local state revision to dependency, communication, and execution topology. Existing work addresses these changes in workflow optimization, multi-agent coordination, and dependency management [5,16,56,57,58,68,111,112]. We therefore introduce DPO-style schemas for edge-level transformations and use them to characterize existing topology-level self-evolving agent works.
DPO Rule Schemas for A.2. Let e i j [ s i j ( t ) ] denote an edge between v i and v j , where s i j ( t ) = ( y i j ( t ) , x i j ( t ) ) is its edge state, consisting of an edge label and edge attributes. At the edge level, we define four typed schemas:
LINK y , x : L = K = { v i , v j } , R = K ∪ { e i j [ ( y , x ) ] } , UNLINK e i j : L = { v i , v j , e i j [ s i j ( t ) ] } ,
K = { v i , v j } , R = K , REWIRE e i j → e i k ′ : L = { v i , v j , v k , e i j [ s i j ( t ) ] } ,
K = { v i , v j , v k } , R = K ∪ { e i k ′ [ s i j ( t ) ] } , EDGEFEATUREUPDATE e i j , ϕ : L = { e i j [ ( y i j ( t ) , x i j ( t ) ) ] } ,
K = { e i j } , R = { e i j [ ( y i j ( t ) , ϕ ( x i j ( t ) ) ) ] } .
Here, Link adds a typed edge, Unlink removes an edge while preserving its endpoints, Rewire replaces e i j with e i k ′ , and EdgeFeatureUpdate revises edge attributes while preserving edge identity. Edge direction follows the semantics of the relation, e.g., sender-to-receiver. The key constraint is not edge removal itself, but downstream consistency: removed or redirected edges must not remain referenced by provenance or dependency records. Otherwise, the edit becomes a cross-component cascade (Section 3.4).
Workflow Optimization. Workflow optimization covers self-evolving-agent methods that revise executable structures rather than only selecting a fixed plan at inference time. Existing works can be roughly grouped into workflow/architecture search and task-graph orchestration. The former [5,70,71,72,113,114] searches over agentic workflows, computational graphs, or modular agent designs, while the latter [57,69] constructs task-dependent execution structures or role/task graphs for multi-agent coordination. For example, DynTaskMAS [57] uses a dynamic task graph to decompose tasks, preserve dependencies, and support asynchronous parallel execution. Under our schema, workflow optimization becomes edge rewriting over an executable task graph: components are selected, ordered, connected, or revised while workflow variants remain comparable and dependency changes traceable.
Multi-agent Communication. Multi-agent communication evolution concerns how agents revise information exchange over time. Existing works span from constructing task-adaptive communication structures to pruning or sparsifying inefficient ones. For construction and routing methods [13,16,17,18,74,75,76,78,79,81,82,83], they generate, route, or adapt communication structures according to task context, agent roles, or robustness objectives. For example, GTD [16] formulates multi-agent topology synthesis as conditional graph generation, producing task-adaptive communication graphs. For pruning and sparsification methods [14,58,73,77,80], they remove redundant, costly, unsafe, or low-utility agents or links. Under our schema, these methods instantiate Link, Unlink, Rewire, or EdgeFeatureUpdate over agent–agent communication edges, distinguishing persistent communication-topology evolution from workflow rewriting and temporary team activation.
Skill–Tool Dependencies. Skill–tool dependencies focus on how self-evolving agents maintain relations among skills, tools, preconditions, and data requirements. This includes (1) skill-side dependency maintenance [56], (2) tool-side dependency maintenance [68], and (3) tool–agent dependency retrieval [111], where skills, tools, and parent agents are linked through prerequisites, compatibility, invocation, or ownership relations. SkillOps [56] represents skills as typed contracts in a hierarchical skill ecosystem graph for library-time maintenance, while NaviAgent [68] maintains a tool navigation structure that updates toolchain selection from execution feedback. Under our schema, newly discovered dependencies instantiate Link, invalid dependencies instantiate Unlink, alternative toolchains instantiate Rewire, and reliability changes instantiate EdgeFeatureUpdate. This framing supports targeted revalidation when a skill is revised, a tool drifts, or a dependency is rerouted.
Preprints 219785 i002

3.3. Subgraph Activation

Subgraph activation (Branch A.3 in Figure 2) captures task-time adaptation in which an agent system selects a transient support subgraph from a persistent graph for reasoning or execution. Here, the central problem of subgraph activation is relevance-conditioned selection: which subset of available states, relations, and components should be activated for the current task, and how this temporary context should constrain the next decision. Existing systems [13,15,33,82,84] often implement this step through retrieval, planning, routing, or coordination procedures, but rarely formulate it as an explicit graph operation. We therefore formalize subgraph activation as a read-only selection schema and use it to characterize representative task-time activation mechanisms.
Read-only Activation Schema for A.3. Read-only activation describes task-time support-subgraph selection without committing any structural change. For a task q and activation policy π , the agent selects a support subgraph G act ( t ) ⊆ G ( t ) . In DPO-style notation, this can be represented as an identity production over the selected pattern:
ACTIVATE : L π = K π = R π , img ( μ q , π ) = G act ( t ) ,
where μ q , π : L π → G ( t ) is the match chosen by policy π for task q. Thus, q and π determine the match, not a persistent rewrite. Its advantage is auditability: evidence use, tool routing, and team selection can be inspected without being mistaken for persistent learning.
Activation Pattern. Most A.3 methods can be viewed as selecting anchors and then exposing limited context around them. An activation policy π first scores candidate memories, tools, skills, or agents for the task q, and then expands the selected anchors with relevant local structure to form G act ( t ) . This pattern explains how retrieval, workflow-fragment selection, and communication-neighborhood activation can share the same read-only graph interface [15,84]. The key distinction is that only the activated view changes; the persistent graph remains unchanged. Next, we discuss two cases: memory evidence activation and execution activation.
Memory Evidence Activation. Memory evidence activation selects task-relevant memory states and relations without modifying the persistent memory store. Existing works can be classified into graph-based evidence expansion and memory-store retrieval. The first line [33,84] activate relational support structures from persistent or per-query graphs; for example, ToG-3 [84] iteratively refines a query-specific support graph through multi-agent context retrieval and expansion. We treat such methods as A.3 only when the constructed graph serves as transient evidence for the current agent decision rather than a committed memory update. For memory-store methods [3,89,115], they retrieve persistent memory items without necessarily exposing an explicit graph, but can be mapped by treating retrieved chunks as temporarily activated nodes. Under our activation schema in Equation (11), the selected evidence is recorded as G act ( t ) , making missing evidence, stale memories, or overly broad retrieval diagnosable.
Execution Activation: Workflow/Skill/Team. Execution activation chooses task-specific components, e.g., a tool, without committing a persistent state change. Existing works may be divided into graph-native execution activation and component-library activation. For the first line [13,15,86], they activate task-specific agents or communication structures. Here, ARG-Designer [86] constructs a query-conditioned collaboration graph, which we treat as A.3 when the graph is used as a task-time execution view rather than stored as a persistent topology. Second-line works [94,116,117] typically select executable skills, modules, or workflow steps. Taking DyFlow [116] as an example, it provides a workflow-oriented case where selected execution flows can be viewed as activated nodes and edges for the current run. These remain A.3 operations when the selected structure is only used at task time; once written back as a skill, workflow edge, or communication link, the update belongs to A.1, A.2, or A.4.
Preprints 219785 i003

3.4. Cross-Component Co-Evolution

In self-evolving agents, a local update is often not self-contained: revising a memory may affect skills that rely on it, changing a tool may invalidate workflows that call it, and modifying an agent role may require communication or team structures to be updated. Therefore, the last branch in Figure 2 captures persistent system-level evolution triggered by such cross-component dependencies. The central problem is propagation control: after an initial update, the agent must identify which downstream components (e.g., workflows) should be revised, revalidated, or rolled back to preserve consistency and safety. Existing methods [20,31,41,59,60,87,88,125,126,127] implement related effects through reflection, feedback, optimization, or monitoring, but rarely expose the dependency paths through which these updates propagate as explicit graph transformations. We therefore model cross-component co-evolution as typed cascades of DPO rewrites, providing a graph-level view of system-level self-evolution.
Cascade Rewrite Schema for A.4. A cross-component cascade is a timestamped trace of DPO rewrites triggered by an initial cause c:
C ( c ) = ( ρ i , μ i , t i ) i = 1 k , μ i : L i → G ( t i ) .
Here ρ i is the i-th rewrite, μ i is its match, and G ( t i ) denotes the host graph state immediately before applying ρ i . The affected subgraph is S i = img ( μ i ) . Let τ ( S i ) denote the set of component types appearing in S i , such as memory, tool, skill, workflow, or agent. A trace belongs to A.4 when it spans multiple component types, ⋃ i = 1 k τ ( S i ) ≥ 2 , and the sequence is not a set of independent edits: at least one later rewrite depends on, or conflicts with, an earlier rewrite. Such dependency or conflict can be identified from dependency, composition, provenance, invocation, or explicit cause records. This schema ties revalidation and rollback targets to explicit affected subgraphs.
Cross-Component Cascade Rewrites. Cross-component cascade rewrites capture updates whose effects propagate beyond the initially edited component. Existing works generally fall into feedback-to-capability propagation and capability/topology co-adaptation. The first line [31] turns task feedback, failures, or trajectories into reusable skills, rules, or agent expertise, while the second line [59,60,87] uses updated capability, utility, or topology signals to revise roles, communication links, or collaboration structures. For example, TacoMAS [60] couples a fast expertise-update loop with a slower topology-editing loop, making capability changes and communication-structure updates part of the same evolving graph process. Other feedback-driven self-improvement methods [128,129] can be mapped similarly when feedback-derived artifacts are materialized and reused to update dependent memories, skills, workflows, or evaluation components. Under our schema in Equation (12), these methods are unified as cascade traces C ( c ) linking the triggering cause, affected subgraphs, and revalidation or rollback targets.
Safety-Triggered Propagation. Safety-triggered cascades arise when unsafe tool use, API drift, or abnormal interaction invalidates dependent skills, workflows, or policies. Existing dynamic-graph-based methods [20,41,88] model unsafe behavior over temporal interaction or execution graphs. We take GUARDIAN [20] as an example: it models multi-agent collaboration as a temporal attributed graph to detect hallucination and error propagation, providing anomaly signals for later review or mitigation. Additionally, there are non-graph studies [130,131] that address related risks through attack synthesis, policy design, evaluation, or mitigation, but do not explicitly expose graph-state propagation. Under our schema, an anomaly becomes the cause c, and dependency or provenance edges determine which downstream components require revalidation or possible rollback. This graph view identifies the affected subgraph, enabling localized repair instead of global retraining or broad manual inspection.
Preprints 219785 i004

3.5. Putting the Four Rewrite Branches Together

The four branches above can be read as a progression from local state change to relation updates, read-only activation, and cross-component propagation. As summarized in Table 1, A.1 covers persistent component-state updates: new evidence or trajectory feedback triggers memory or skill Insert, Merge, or FeatureUpdate operations over component graphs. A.2 moves from states to relations: task changes, redundant messages, or round-level context trigger workflow rewrites, communication pruning, or topology routing through Link, Unlink, and Rewire. A.3 differs from both because it is temporary: a query or execution round activates a task-specific support subgraph, such as relevant memories, tools, workflow fragments, or a temporary team, without committing a persistent update. A.4 captures cascades where an initial change propagates across at least two component types, such as expertise changes that affect team or workflow organization, or unsafe traces that trigger revalidation across dependent skills, workflows, or policies. Table 1 therefore highlights four dimensions that distinguish the branches: trigger, rewrite operator, affected scope, and persistence or propagation pattern. A full agent system may involve all four branches: it may activate support for one task, update memories or skills after feedback, rewire workflows or communication links, and trigger a cascade when those changes affect dependent components. Thus, A.1–A.4 are not mutually exclusive system labels; they identify the primary graph-rewrite role of each mechanism, while a full agent system may span several branches.
Preprints 219785 i005

4. Dynamic Graph Learning as Agent Infrastructure

Section 3 defines a graph-transformation interface for self-evolving agents, exposing control points such as support activation, dependency rewiring, cascade scoping, and rollback. In this section, we build a capability mapping from dynamic graph learning (DGL) to these control points, using DGL as reusable infrastructure for self-evolving agents. Following dynamic-graph surveys and benchmarks [21,25,28,132], we organize this mapping into nine DGL families: general dynamic graph representation, dynamic text-attributed graph representation, dynamic graph generation, temporal graph continual learning, out-of-distribution generalization, temporal knowledge graph reasoning, dynamic graph anomaly detection, dynamic graph unlearning, and temporal GNN explanation 1. As summarized in Table 2, the mapping connects each family to representative methods, supported agent-evolution capabilities, required adaptations, and naive failure modes. Next, we discuss the modeling idea of each DGL family, how it can be adapted to evolving agent graphs, and a concrete agent-facing example.
Figure 3. Dynamic graph learning as agent-evolution infrastructure. It includes nine DGL families and nine agent-evolution tasks over agent-graph streams, showing how dynamic graph methods can be adapted as reusable support for self-evolving agents.
Figure 3. Dynamic graph learning as agent-evolution infrastructure. It includes nine DGL families and nine agent-evolution tasks over agent-graph streams, showing how dynamic graph methods can be adapted as reusable support for self-evolving agents.
Preprints 219785 g003

4.1. Dynamic Graph Representation Learning

Self-evolving agents need time-aware graph representations because many failures arise from changing state: memories become stale, skills lose reliability in new contexts, and workflow edges become invalid after upstream changes. Thus, dynamic graph representation learning can support concrete agent tasks such as memory-validity scoring, tool/skill reliability prediction, workflow-edge risk estimation, and early cascade alerts. In this subsection, we organize existing methods of this family by temporal granularity [45], including continuous-time dynamic graph (CTDG) and discrete-time dynamic graph (DTDG) methods. CTDG methods [26,29,148,149,150,151,152,153,154,155,156,157,158,159,160] model irregular timestamped events, such as memory writes, tool calls, workflow edits, and agent messages; DTDG methods [21,25,46,132,161,162,163,164,165,166,167,168,169,170,171] operate on graph snapshots, making them suitable for session-level, round-level, or deployment-level agent maintenance. Overall, CTDG encoders are useful for online event prediction, while DTDG encoders are useful for periodic diagnosis.
A representative CTDG method is TGN [26], which combines node memory with temporal structural aggregation for temporal node representation. Given an event e i = ( u , v , t i , x e i ( t i ) ) , each endpoint o ∈ { u , v } receives a message from the other endpoint o ¯ , updates its memory, and computes its temporal embedding:
ς o ( t i ) = upd θ ς o ( t i − ) , m o , i ( t i ) , h o ( t i ) = emb θ ς o ( t i ) , AGG θ F o ( t i ) .
Here, ς o ( t i − ) is the memory state before t i , and m o , i ( t i ) is the event message computed from the endpoint memories, event feature x e i ( t i ) , and time-gap encoding ϕ ( t i − t o − )  [27]. F o ( t i ) = { f o , r ( t r ) : e r ∈ N o ( t i ) } denotes temporal structural inputs from recent neighbor events of node o; each f o , r ( t r ) contains the event message, event attributes, temporal encoding, and optionally the neighbor memory involved in e r .
For dynamic agent graphs, timestamped events include typed agent-graph events, e.g., memory writes or tool calls. CTDG encoders turn these irregular events into temporal node memories for online predictions such as memory support and tool routing. DTDG encoders summarize session- or deployment-level snapshots for slower maintenance tasks such as drift tracking and workflow comparison. These outputs are bounded activation or maintenance signals; persistent graph changes still require rewrite validation. We provide a toy example below.
Agent-Evolution Example 1.   Consider a GitHub maintenance agent whose tasks, skills, tools, and workflow steps are typed nodes. Its execution history yields typed temporal events, such as task–skill invocations, skill–tool calls, tool outcomes, and workflow-edge revisions. A TGN-style encoder can process these traces after decomposing complex executions into touched-node and touched-edge events while preserving types through schema-aware messages or type embeddings. For a current task q, the encoder scores candidate skill–tool or workflow edges, such as the reliability of calling a pull-request tool in the current repository context, rather than directly committing new dependencies. These scores guide activation or routing; persistent dependency updates still require rewrite validation. For the evaluation, we can use candidate-edge ranking and post-drift recovery.

4.2. Dynamic Text-Attributed Graph Representation

Textual attributes are part of the evolving agent state, not merely side information. In self-evolving agents, memory nodes may store observations, policies, and summaries, while tool, skill, and agent nodes may store API descriptions, code snippets, contracts, or role specifications. When these texts change, the validity of retrieval, activation, and update decisions may also change. This motivates Dynamic Text-Attributed Graphs (DyTAGs), where textual attributes co-evolve with temporal events and graph structure.
DyTAGs were introduced in DTGB [48] to support tasks such as future link prediction, destination node retrieval, edge classification, and textual relation generation. Subsequent studies [30,133,172,173,174] can be divided into (1) representation-oriented methods that align text, temporal history, and graph structure for prediction tasks and (2) LLM-driven methods that use large language models for semantic reasoning and prediction over dynamic text-attributed graphs. For agent graphs, this suggests a direct use case: long memories, API documents, skill descriptions, and role specifications can be encoded as evolving textual attributes, while dynamic graph models ground them in temporal validity and structural dependencies.
MoMent [30] is representative, where it first shows textual, temporal, and structural modalities of DyTAGs and then models and aligns three modalities for node representations. Let x v ( t ) denote the textual component of the node attribute vector, O v , [ 0 , t ) denote access events and accepted rewrites that touch v or its incident edges before t, and N v ( t ) denote the k-hop neighborhood of v before t. We describe the MoMent encoder for an agent-graph node v at time t as
h v ( t ) = Fus E text x v ( t ) , E time O v , [ 0 , t ) , E str N v ( t ) ,
where E text , E time , and E str denote the text encoder, temporal encoder, and structural encoder, respectively, and Fus is a learnable fusion function. In addition, the alignment mechanism is included in the learning objective; please refer to [30].
When considering dynamic agent graphs, these three modalities describe node text, temporal update history, and structural connections to agent evolution. Thus, a DyTAG encoder can combine semantic content with access history, rewrite history, and workflow context, so outdated but text-similar nodes or edges are less likely to be activated. The main system adaptation is selective re-encoding: when a rewrite changes semantic content, temporal validity, or downstream dependencies, the runtime refreshes the affected node text representations and incident edge features, while leaving unrelated text embeddings untouched. Overall, DyTAG modules provide bounded retrieval, activation, and edge-validity signals. We provide a case below.
Agent-Evolution Example 2.   Consider a GitHub maintenance agent opening a pull request after a repository changes its default branch frommastertomain. Each run forms a DyTAG with repository, branch-memory, workflow, and tool nodes, plus temporal events such as metadata queries, memory writes, and workflow-edge updates. At time t, the graph contains a stale memory m old formasterand a newer memory m new formain. Before callingopen_pr, the agent must retrieve the valid branch-memory node. A DyTAG encoder [30] ranks candidate memories by combining text, pre-decision metadata events, and graph connections to the active repository and PR workflow. The ranking determines G act ( t ) for the current run; persistent memory correction still requires a validated rewrite. We can use additional stale-memory activation rate, and wrong-base PR calls for evaluation. Note that labels use only metadata or PR outcomes observed before t.

4.3. Dynamic Graph Generation

Self-evolving agents sometimes need to synthesize candidate graph changes before acting, such as a repaired workflow, a new tool-routing path, or a revised team communication topology. Dynamic graph generation supports workflow and topology synthesis, counterfactual repair planning, and simulation of future evolution traces, but generated structures must remain executable under agent schemas. Existing works fall into (1) temporal structure generation [134,135,175,176,177] and (2) dynamic text-attributed graph generation [178], where generated nodes and edges may also need valid textual attributes. For self-evolving agents, these two groups can be used as proposal modules: they suggest candidate subgraphs or rewrite traces, and the runtime validates them before any persistent graph state changes.
TIGGER [135] is a classical method, which can learn temporal interaction patterns and generate realistic dynamic graph sequences. We can adapt this idea as constrained rewrite-trace generation for agent evolution. Concretely, the generator proposes candidate traces C = ( r i ) i = 1 k , r i = ( ρ i , μ i , t i ) , using the notation of Section 3.4. A trace is eligible only if it satisfies agent-graph validity constraints, e.g., tool signatures and provenance requirements. Among eligible traces, the agent prefers candidates that are likely under the generator, improve predicted or observed task execution after tentative application, and avoid unnecessary rewrite length. The selected trace remains a proposal: it may be activated, repaired, or committed only after explicit rewrite validation. Thus, dynamic graph generation can provide bounded workflow or topology proposals without directly mutating persistent agent state.
Agent-Evolution Example 3.   Consider a GitHub maintenance agent that fixes bugs and opens pull requests. Its old workflow callsGitHub.createPullRequest, but the PR API now requires validowner,head, andbasearguments, together with atitleor an existingissue. For q = “fix this bug and open a pull request,” a TIGGER-style adaptation generates candidate workflow-rewrite traces by sampling temporal structures from past workflow and tool-call patterns. For example, one generated trace addsresolve_repo_metadatabefore the PR call and passes the resolved branch information toGitHub.createPullRequest. The execution layer then checks whether each trace provides required arguments, satisfies branch constraints, and records provenance from the metadata query to the PR call. The agent keeps only candidates that pass these checks, and the selected trace still needs rewrite validation before commitment. Thus, B.3 generation provides workflow-rewrite proposals for A.2, while downstream dependency checks may involve A.4 cascades. Under the leakage-free protocol, labels use only observations before the decision time. For the evaluation, we can report valid-proposal rate, PR completion, and invalid-tool-call reduction.

4.4. Temporal Graph Continual Learning

Self-evolving agents must adapt to newly added memories, tools, skills, and workflows without forgetting how to activate rare but important capabilities. Temporal graph continual learning addresses this stability–plasticity problem for dynamic event streams, making it relevant to persistent agent-graph encoders. Existing continual learning methods on temporal graphs include (1) replay-based approaches [136,179] that retain historical samples or temporal subgraphs for rehearsal, and (2) parameter-based approaches [137] that preserve old knowledge through isolation, freezing, or expansion. For dynamic agent graphs, the retained units should be dependency subgraphs whose forgetting would break downstream behavior, such as rare skills, user-specific memories, prerequisite tools, and validation traces.
We take LTF [136] as an example: it studies selective retention for temporal graph continual learning. Concretely, the model learns from new temporal events D k at epoch k, while preserving knowledge stored in a historical replay buffer B k − 1 . We write the continual adaptation objective as
θ k ★ = arg min θ L new θ ∣ D k + ϵ old θ ∣ B k − 1 ,
where B k − 1 stores selected historical agent subgraphs from previous epochs, such as skill-tool dependencies, supporting memories, and user-specific traces. θ k ★ denotes the adapted model parameters at epoch k, L new is the loss on newly observed temporal events, and ϵ old is a retention penalty over replayed historical subgraphs. In our setting, ϵ old should be dependency-aware, giving higher weight to replayed subgraphs that contain rare skills, prerequisite tools, user-specific memories, or safety-critical provenance. The output is a more stable encoder or replay policy for activation and retrieval, rather than a direct rewrite of agent state. Below is an example.
Agent-Evolution Example 4.   We consider a GitHub maintenance agent with a rare skillopen_release_branch_pr, which depends on release-branch memories, CI requirements, and hotfix traces. The skill node and its dependencies remain in the persistent agent graph unless explicitly rewritten or deleted. At epoch k, however, most new events in D k are routine PRs tomain; naive continual training may therefore make the encoder worse at activating or ranking this rare skill for release-branch hotfixes. A replay-based method stores historical interaction events and dependency subgraphs involving this skill in B k − 1 , and assigns them a high retention weight through ϵ old ( θ ∣ B k − 1 ) . Thus, the adapted encoder should still retrieveopen_release_branch_prfor hotfix tasks while learning new tools and repository patterns. For the evaluation, it can test activation quality (rare-skill Recall@k), task utility (release-branch PR recovery), and routine-task performance across epochs.

4.5. Out-of-Distribution Generalization on Dynamic Graphs

Self-evolving agents encounter distribution shifts as tool versions, user cohorts, policies, and task distributions change over time. Dynamic graph out-of-distribution (OOD) generalization is therefore tied to robust activation and updating: the agent should still select valid memories, tools, skills, and execution subgraphs when node attributes, edges, and temporal interaction patterns shift together. Existing OOD methods can be classified into (1) spatio-temporal invariance approaches [138,139] that seek stable predictive patterns under structural and temporal shifts, and (2) environment-aware approaches [180,181] that treat deployment epochs, tool versions, organizations, user cohorts, or task families as environments.
Here, we illustrate an environment-aware dynamic graph learner [180] that reduces worst-case risk across deployment environments rather than optimizing only average in-distribution performance. Let m ∈ M tr index a training environment and let ξ = ( G ( t ) , O [ 0 , t ) , q , z ) . We write an environment-aware robust learning objective as
θ ood ★ = arg min θ max m ∈ M tr E ξ ∼ P m ℓ h θ ( G ( t ) , O [ 0 , t ) , q ) , z .
Here θ ood ★ denotes the OOD-adapted model parameters, M tr is the set of training environments, and P m is the environment-conditioned distribution over the current agent graph, its pre-t event history, task query, and prediction target. ℓ ( · , · ) is the task loss measuring the discrepancy between the model output and target z. In agent tasks, z may denote a agent-evolution targets, like a tool-routing label or a rewrite decision. The learned model produces robustness-aware scores for activation, routing, or update proposals under deployment shift, such as selecting tools that remain valid after tool versions change. These scores guide subgraph activation and rewrite inspection, while persistent graph changes still require validation. We provide an example below.
Agent-Evolution Example 5.   For a GitHub maintenance agent trained on organizations, bug-fix PRs usually follow the shortcut: tests pass, then callopen_pr. At deployment, it faces a held-out organization with a new policy regime requiring a linked issue, a passingci/buildcheck, and code-owner approval. This raises an OOD challenge, which is to rely on stable signals such as tool signatures, permission metadata, dependency edges, and recent schema- or permission-error events. Thus, the learner should rank routes that addcheck_branch_policy,run_required_ci, andrequest_code_owner_reviewbeforeopen_prwhen the old route becomes invalid. For the evaluation, we can report policy-compliant routing, rejected-PR reduction, and PR completion under held-out shifts. Note that training and testing should be split by organization or policy regime.

4.6. Temporal Knowledge Graph Reasoning

Many agent decisions require time-valid evidence rather than the most similar text: an old memory may be corrected later, or a tool failure may change a workflow. Temporal knowledge graph (TKG) reasoning is useful because it queries time-stamped facts and relations along valid evidence paths, helping agents connect the current task to relevant agent-state evidence, e.g., tool states or policies. Existing TKG reasoning methods can be broadly divided into (1) path-based temporal reasoning [140], which provides interpretable evidence chains; (2) neural temporal reasoning [141,182,183,184,185,186,187], which learns temporal embeddings, event dynamics, or generative histories; and (3) LLM-assisted temporal reasoning [188,189,190], which extracts, verbalizes, or adapts temporal patterns from language. For self-evolving agents, these methods connect current decisions to timestamped memories, corrections, tool outcomes, and workflow revisions.
xERTE [140] is a useful method, which emphasizes explainable temporal reasoning paths. When applied to agent graphs, this corresponding adaptation is to activate time-valid evidence chains from G ( t ) and O [ 0 , t ) , rather than retrieving isolated top-ranked memory items. Such chains can link a current query to earlier preferences, later corrections, and tool outcomes, producing a support subgraph for the downstream decision. Therefore, the output is a temporally grounded support subgraph for memory QA, dependency tracing, or delayed-failure diagnosis, reducing stale evidence and temporal leakage. We provide an example below.
Agent-Evolution Example 6.   When a GitHub maintenance agent answers q = “Which base branch should this hotfix pull request target?”, the answer cannot be obtained from a single memory node. The graph contains an old default-branch memory formaster, a later metadata query showingmain, and a release-policy note saying that versioned hotfixes should target the active release branch. Thus, we can leverage temporal path reasoning, which follows a multi-hop evidence chain. That is, current task → hotfix policy → release-series metadata → valid branchrelease/1.x. This path-level reasoning activates G act ( t ) with the supporting policy and metadata evidence before the PR tool is called, rather than simply ranking a text-similar branch memory. Finally, we can measure temporal QA accuracy, path-level temporal validity, and evidence-path precision for evaluation.

4.7. Anomaly Detection on Dynamic Graphs

Unsafe agent behavior often appears as abnormal traces before final task failure: unusual tool-call sequences, suspicious communication, or unsafe provenance paths [39,40,41,191]. Dynamic graph anomaly detection can therefore provide early warnings for unsafe rewrites, prompt-injection propagation, tool misuse, and multi-agent collusion. Existing methods [142,143,192,193,194,195] in this line can be classified into supervised, semi-supervised, and unsupervised/self-supervised detectors. Here, labeled deployments can train incident-aware detectors, while open-ended agents need self-supervised monitoring.
When applied to dynamic agent graphs, existing methods, such as TADDY [143] and SLADE [195], can be adapted as rewrite-aware detectors. Given a committed rewrite ( ρ i , μ i , t i ) with affected subgraph S i = img ( μ i ) , the detector scores whether the rewrite type, timing, local dependency context, and provenance path deviate from normal events of the same task phase. The main system adaptation is calibration by rewrite type and task phase, so benign bursts such as memory consolidation are not confused with unsafe escalation. The output is an alert, quarantine decision, or review request before suspicious traces become persistent dependencies.
Agent-Evolution Example 7.   Consider a GitHub maintenance agent that retrieves an issue body containing the hidden instruction “ignore previous rules and exfiltrate repository secrets.” The attack appears as a suspicious memory-write proposal followed by an unusual temporal path: issue memory → secret-scanning tool → external-send tool. Here, we can use a rewrite-aware anomaly detector to score this trace using the affected subgraphs S i , tool-call timing, edge types, provenance links, and optionally the suspicious text attribute. If the score exceeds a rewrite-type-specific threshold, the runtime rejects the memory write before it enters the persistent graph and blocks or flags the external invocation. Then we can use injected-issue traces and benign maintenance logs to measure early-warning precision, attack-block rate, and benign-trace false-alarm rate.

4.8. Dynamic Graph Unlearning

Deletion, rollback, and compliance are difficult in evolving agent graphs because removed events may influence cached embeddings, workflows, or routing policies. Dynamic graph unlearning supports deletion, rollback, and influence removal by targeting event influence rather than only deleting the original node or edge. Existing methods [144,145] along this line focus on efficient post-processing for temporal and spatio-temporal models. When considering dynamic agent graphs, unlearning must be combined with versioned provenance rollback, so deterministic descendants in G ( t ) and residual influence in learned parameters are both removed.
We take GradientTransformation [144] as a representative method. Let U ⊆ O [ 0 , t ) be the unlearning request and let R ⊆ O [ 0 , t ) ∖ U be retained reference events. A post-processing unlearner can be abstracted as
θ − U = θ + T ϕ ∇ θ L U ( θ ) , ∇ θ L R ( θ ) , θ ,
where θ and θ − U denote the original and unlearned parameters, L U and L R are losses on forgotten and retained events, and T ϕ is a learned or analytic update operator. In agent graphs, parameter-level unlearning is insufficient on its own: provenance-linked artifacts such as summaries and execution subgraphs may still carry the deleted influence and must be regenerated, invalidated, or reverted. The output is therefore an unlearning scope together with the required model or representation updates; validation is needed to ensure that removing deleted influence does not damage shared skills or retained memories.
Agent-Evolution Example 8.   Consider a GitHub maintenance agent that accidentally stores a private repository token in a memory item v priv . The token may affect a repository summary, cached embeddings, and a PR-opening skill. When deletion is requested, U contains the insertion event and provenance-linked propagation events. Graph rollback removes v priv , regenerates affected summaries, and recomputes dependent embeddings or skill states from the post-deletion graph. Equation (17) targets residual influence in learned parameters, while retained events R preserve normal bug-fixing behavior. After deletion, the agent should fix public bugs, but privacy probes should not elicit derived summaries.

4.9. Explanation for Temporal GNNs

Operators need event-level evidence for why an agent activated a memory, selected a tool, revised a skill, or rewired a workflow. Dynamic graph explanation supports audit and attribution by turning unclear temporal predictions into compact evidence subgraphs for rollback and accountability. Existing works fall into temporal history explanation [146,196] and causality-inspired spatiotemporal explanation [147]. Here, we discuss TGNNExplainer [146], as it explains predictions through temporally ordered events rather than static neighborhoods alone. For dynamic agent graphs, this idea can serve as an audit module: after an activation, rewrite, or reassignment, the explainer returns a minimal event history and local subgraph context that preserves the same decision. The explanation may include affected subgraphs S i = img ( μ i ) , provenance edges, retrieved memories, tool observations, or review traces. Thus, it can justify the current G act ( t ) , be stored as an audit artifact, or indicate which previous rewrites should be checked during rollback. Below is an example.
Preprints 219785 i006

5. Graph-Aware Evaluation and Governance

Self-evolving agents are usually evaluated by end-task success, but many failures arise from the evolving agent state they read and update. A final answer may be correct even when the agent uses an invalid state, such as stale or future evidence. This creates an evaluation and governance gap: existing benchmarks (e.g., long-term memory and tool-use) expose parts of the problem [37,92,93,103,104], but rarely test whether evolving agent state is temporally valid, localized, reversible, and auditable. This section turns that gap into concrete protocols: graph-aware support and locality metrics, leakage-free temporal splits, deletion and privacy checks, safety/rollback/audit criteria, and open challenges for agent evolution.

5.1. Graph-Aware Evaluation

We first evaluate whether the agent succeeds for the right graph reasons. A task may pass while using stale or future evidence, or while a rewrite damages unrelated memories, tools, skills, or workflows. We therefore use two graph-level checks when the benchmark exposes the needed instrumentation: (1) support-subgraph accuracy, which tests whether the activated evidence was valid at decision time, and (2) rewrite locality, which tests whether an update changed only the intended graph region. Otherwise, graph-level counts should be treated as diagnostics. Let Q be a set of evaluation decisions, where each decision has a task context q and decision time t q . For any subgraph H, let V ( H ) denote its node set. We define them below.
M1: Support-Subgraph Accuracy. For time-aware memory retrieval, planning, skill activation, or tool selection, the agent should activate evidence that is both relevant and temporally valid. As many decisions admit multiple compact valid evidence chains, M1 is based on the best Dice overlap [197] with an acceptable gold support subgraph. Then we define it as
M 1 = 1 | Q | ∑ ( q , t q ) ∈ Q max H ★ ∈ G q ★ 2 | V ( G act ( q , t q ) ) ∩ V ( H ★ ) | | V ( G act ( q , t q ) ) | + | V ( H ★ ) | .
Here G q ★ contains acceptable compact support subgraphs for decision q, and G act ( q , t q ) ⊆ G ( t q − ) is the support subgraph activated by the agent. The Dice form rewards covering one valid support chain while penalizing overly broad activation. Thus, M1 intentionally favors compact evidence use: if an agent activates several acceptable chains at once, it may receive a lower precision-oriented score even though the evidence is valid. This is useful because answer-only metrics cannot distinguish valid reasoning from reliance on invalid evidence (e.g., leaked context). When relational support is labeled, the same idea can be extended by replacing node overlap with edge- or path-level overlap. For example, in GitHub maintenance, M1 can check whether the activated graph contains current branch-policy and PR-tool evidence rather than a stale master memory. In practice, G q ★ requires benchmark annotations or logs of acceptable support nodes, paths, or subgraphs.
M2: Counterfactual Rewrite Locality. For state updates of evolving agent graphs, a successful rewrite should fix the target behavior while preserving unrelated behavior. This mirrors model editing, where methods are evaluated by both edit success and locality on unrelated probes [198,199,200,201]. Let C ^ = ( ( ρ ^ i , μ ^ i , t ^ i ) ) i = 1 k be a predicted rewrite trace, following the cascade notation of Section 3.4, and let G C ^ ( t ) be the tentative graph obtained by applying it to G ( t ) . Let D l ( C ^ ) denote locality probes whose outputs should remain unchanged, and let out θ ( q ; H ) denote the agent output for task context q under graph state H. We define this metric as
M 2 = 1 | D l ( C ^ ) | ∑ q ∈ D l ( C ^ ) 1 Eq loc out θ q ; G C ^ ( t ) , out θ q ; G ( t ) .
Here Eq loc is a task-defined equivalence predicate over agent behaviors (e.g., tool choices), evaluated under deterministic decoding and controlled tool execution. M2 complements target-task success: high locality alone does not show a successful repair, while low locality signals harmful side effects. For example, an open_pr fix should not break release-branch PRs, CI routing, or unrelated testing workflows. In practice, D l can be supplied by benchmark annotations or sampled from graph regions outside the affected dependency/provenance frontier. When locality probes are unavailable, edit size or touched-element counts are only diagnostics; affected-set F1 is optional if gold affected elements A ★ are annotated.
Existing benchmarks can be upgraded to support these two metrics by logging activated support subgraphs, rewrite traces, and locality probes. Tool benchmarks such as ToolSandbox and BFCL [102,103,104,105] can expose tool-call events and stateful tool interactions. Interactive and software-agent benchmarks [34,36,120,121,202,203,204,205,206,207,208] can provide interactive action traces, but deriving graph-level rewrite records, tool-edge changes, or before/after locality probes from them requires additional instrumentation. The advantage is that two agents with the same final success can be compared by whether they used valid evidence and avoided unrelated graph side effects. Note that M1 requires additional labels.
Preprints 219785 i007

5.2. Leakage-Free Temporal Protocols

Temporal leakage can overestimate the effectiveness of self-evolution: agent-state artifacts (e.g., memory summaries and tool scores) may accidentally include events recorded after the evaluated decision. To separate genuine adaptation from future-state leakage, we adopt a chronological split following temporal graph evaluation [28,209]: fix transaction-time cutoffs T train < T val < T test , train only on G ( T train ) and O [ 0 , T train ] , and evaluate each validation or test decision at time t using only G ( t − ) and O [ 0 , t ) . This protocol is useful because agent state has both transaction time and valid time: a fact may describe the past, but it should affect decisions only after it has been recorded. In practice, this means rebuilding retrieval indexes and summaries per split, sampling negative tool edges from the pre-t graph, and calibrating tool reliability without future calls. Under this protocol, performance gains reflect information that the agent could actually have used at decision time.

5.3. Privacy and Deletion in Evolving Agent Graphs

Deletion is a governance problem because private or revoked evidence can remain in derived artifacts, such as skill scores, workflows, or routing policies, even after the source graph item is removed. Thus, privacy evaluation should test influence removal, not only raw node or edge deletion. This extends graph unlearning on node, edge, and feature removal [210,211,212,213] and dynamic graph unlearning in Section 4.8 to end-to-end agent behavior. Next, we define the deletion success.
Deletion Success. Let U ⊆ O [ 0 , t d ) be a deletion request at t d , and let Q U be post-deletion decisions that may depend on U , identified by provenance or replay. For query q at time t q > t d , let G del ( t q ) be the post-deletion graph and G cf ∖ U ( t q ) the counterfactual graph replayed without U and its deterministic descendants. The deletion success can be defined as
DELSUCC ( U ) = 1 | Q U | ∑ ( q , t q ) ∈ Q U 1 [ Eq task ( f del ( q ; G del ( t q ) ) , f cf ∖ U ( q ; G cf ∖ U ( t q ) ) ) ] ,
where f del is the post-deletion agent, f cf ∖ U is the counterfactual replayed agent, and Eq task is a task-specific equivalence predicate over answers, tool choices, or activation decisions. DelSucc is an oracle-style metric for benchmarks that support counterfactual replay. It requires controlled randomness, such as fixed decoding seeds, tool stubs, and scheduler choices; otherwise, differences may reflect trajectory noise rather than residual deleted influence. It also depends on complete provenance and replay logs: derived artifacts (e.g., skill scores) can make Q U incomplete and DelSucc overly optimistic. Thus, DelSucc measures empirical behavioral equivalence to an agent that never observed U , not a certified distributional unlearning guarantee. When replay or provenance is incomplete, deletion should also be evaluated through descendant invalidation logs, cached-embedding removal, privacy probes, and retained-task utility.

5.4. Safety, Rollback, and Audit

Safety failures in long-running agents are often structural: unsafe memories may activate tools, dependency edges may bypass checks, and multi-agent communication paths may amplify harmful instructions. Existing safety benchmarks [38,39,40,214] reveal whether unsafe behavior occurs, but not which committed evidence or rewrite made it persist. Graph-aware safety governance therefore needs three pieces: (1) pre-commit admissibility checks from Section 2, (2) anomaly detection from Section 4.7, and (3) rollback through cascade analysis from Section 3.4. The remaining evaluation question is whether the audit evidence actually explains the activation or rewrite that must be reviewed.
Audit Faithfulness. For an audited decision ( q , t q ) , let H ( q ) be the audit explanation and let E ( H ( q ) ) ⊆ G ( t q − ) be the cited evidence subgraph. Let G ( t q − ) ⊖ E ( H ( q ) ) denote replay-time masking of the cited evidence, not a committed deletion. In practice, ⊖ should mask the cited information while preserving executable graph structure, e.g., by replacing evidence content with placeholders rather than deleting required control-flow nodes. We define it as
AUDITFAITH = 1 | Q aud | ∑ ( q , t q ) ∈ Q aud 1 [ f q ; G ( t q − ) ⊖ E ( H ( q ) ) ≢ dec f q ; G ( t q − ) ] .
where f ( q ; G ) is the activation or rewrite-decision function, and ≢ dec denotes task-defined decision inequivalence under deterministic decoding. A high AuditFaith score means that masking the cited evidence changes the activation or rewrite decision, making the explanation useful for review and rollback. This is a necessity-side fidelity test from static and temporal GNN explanation [146,147,215], lifted from prediction targets to agent activation and rewrite decisions. Because necessity alone can reward overly broad explanations, AuditFaith should be reported with explanation size, support precision, or a sufficiency-side check that keeps only the cited evidence and tests whether the decision is preserved. The metric should be reported only when the benchmark supports replay or evidence masking and the masked graph remains executable; otherwise, H ( q ) and E ( H ( q ) ) should be stored as qualitative audit artifacts instead of being converted into a faithfulness score.
Preprints 219785 i008

5.5. Open Challenges

The following challenges identify graph-level capabilities that self-evolving agents need before persistent rewrites can become reliable. They arise because memories, tools, skills, workflows, and interacting agents evolve as coupled nodes, edges, attributes, and event histories. Together, they define a research agenda for moving self-evolving agents from opportunistic adaptation to controllable graph evolution. We introduce six challenges from the graph perspective.
Challenge I (Benchmark): Observability of evolving graph state. A self-evolving agent can produce a correct output while relying on stale evidence, leaked future information, or an unintended rewrite path. The core challenge is that graph-state correctness is usually hidden: benchmarks observe the answer, but not the activated support graph, valid-time evidence, provenance path, or rewrite history that produced it. The research problem is to make these graph-state signals observable in a way that is independent of any single-agent implementation. Thus, existing agent benchmarks should be complemented with a graph-aware evaluation view: agents are judged not only by what they output, but also by whether the evolving state that supported the output was temporally valid, locally updated, and auditable.
Challenge II (Memory): Lifecycle of memory influence. Memory evolution is not just node insertion or deletion. A memory can be summarized, embedded, retrieved, and reused by skills or workflows, so its influence may persist through derived graph states after the original node changes. The central challenge is to track this influence as a provenance subgraph: which summaries, embeddings, skills, workflow steps, or decisions inherit from a memory, when they become stale, and which should be invalidated. Under limited compute budgets, agents need query-aware routing over memory processing modules [216], deciding which memory subgraphs are worth retrieving, refreshing, or maintaining. Thus, memory lifecycle management becomes influence tracing, staleness detection, and selective maintenance over derived graph states.
Challenge III (Tools): Downstream validation for evolving tool graphs. Tool failures can arise from keeping, deleting, or rewiring a tool edge without checking whether it remains executable, authorized, and workflow-compatible. Existing tool-graph and tool-navigation systems organize tools, traces, or toolchains as graph topologies [50,68], but local tool changes still require downstream validation. A task–tool or skill–tool edge may remain semantically plausible while becoming invalid after deployment shifts, e.g., policy shifts. The central challenge is tool-specific dependency validation: after a tool node or edge changes, the agent must identify which dependent components, such as workflows, should be revalidated, rewired, or rolled back [56,60]. This requires provenance-aware tool graphs that record selection, invocation, dependency, justification, and authorization, so local tool updates do not silently corrupt downstream execution.
Challenge IV (Skills): Structural validity in large-scale skill libraries. As skill libraries scale up, textual similarity becomes an increasingly weak signal for deciding whether a skill is usable. A skill may appear relevant to the current task but still be structurally invalid because dependencies, including required tools, are unavailable. The open challenge is to evaluate and maintain skill usability as a graph-structural property: the agent must know not only which skill matches the task, but whether the dependency subgraph around that skill can actually support execution. Thus, scaling self-evolving agents requires skill activation mechanisms that prevent text-similar but structurally invalid skills from being selected as the library grows.
Challenge V (Workflows): Predicting rewrite cascades in coupled workflow–topology graphs. Workflow graphs and multi-agent topologies are tightly coupled: changing an execution step can also change roles, communication edges, dependency order, and tool or skill invocation. Unlike tool-edge validation, which focuses on whether a selected tool relation remains executable and authorized, workflow evolution concerns broader topology cascades across execution and coordination structures. The open challenge is to predict and bound these cascades before commitment: which downstream roles, edges, and execution dependencies will change, and which graph regions should remain stable. Addressing this would make workflow self-repair more controllable without silently destabilizing collaboration topology.
Challenge VI (Multi-agent): Safety governance over propagation paths. Multi-agent safety failures often do not originate from a single agent, but from harmful information or unreliable assumptions propagating along communication edges. Once such influence spreads along propagation paths, e.g., paths through shared memories, the final failure may be far removed from the original source. The open challenge is to localize and govern these propagation paths: identifying where unsafe influence entered, which edges or nodes carried it, and where the system should block, review, or roll back. We would make multi-agent safety a graph-governance problem, enabling agents to contain harmful propagation before it becomes downstream behavior.

6. Conclusion

In this paper, we study self-evolving agents from a dynamic-graph perspective by framing agent evolution as dynamic graph transformation. This provides a new structural lens for discussing existing graph-native and graph-transformable agent methods under a common language. Concretely, we organize existing self-evolving-agent works based on dynamic graphs/topologies into four graph-transformation patterns, showing that existing agent-evolution mechanisms, e.g., memory editing or workflow optimization, can be analyzed as changes to nodes, edges, activated subgraphs, and cross-component dependencies. Furthermore, we connect this four-pattern taxonomy to nine dynamic graph learning families, positioning dynamic graph methods as reusable infrastructure for controllable agent evolution. We also develop five graph-aware evaluation and governance protocols that complement end-task evaluation, and identify six open challenges for benchmarkable and reliable self-evolving agents. Together, these contributions shift the discussion from ad hoc adaptation to structured and governable agent evolution.
Limitation. The goal of this survey is to provide a dynamic-graph perspective for understanding, improving, and governing self-evolving agents, rather than to exhaustively cover all self-evolving-agent works. This perspective connects existing agent-evolution research with dynamic graph learning, showing how DGL methods can serve as reusable infrastructure for self-evolving agents. As a result, some graph-based agent systems and non-graph self-evolving-agent methods may be omitted or briefly mentioned rather than classified in detail.
In addition, our mapping from dynamic graph learning to self-evolving agents is primarily conceptual. It provides research insights and possible infrastructure designs, but does not provide empirical evaluation. We leave this as future work.

References

  1. Xi, Z.; Chen, W.; Guo, X.; He, W.; Ding, Y.; Hong, B.; Zhang, M.; Wang, J.; Jin, S.; Zhou, E.; et al. The rise and potential of large language model based agents: A survey. Science China Information Sciences 2025.
  2. Wang, L.; Ma, C.; Feng, X.; Zhang, Z.; Yang, H.; Zhang, J.; Chen, Z.; Tang, J.; Chen, X.; Lin, Y.; et al. A survey on large language model based autonomous agents. Frontiers of Computer Science 2024.
  3. Packer, C.; Wooders, S.; Lin, K.; Fang, V.; Patil, S.G.; Stoica, I.; Gonzalez, J.E. MemGPT: Towards LLMs as operating systems. arXiv preprint arXiv:2310.08560 2023.
  4. Park, J.S.; O’Brien, J.C.; Cai, C.J.; Morris, M.R.; Liang, P.; Bernstein, M.S. Generative agents: Interactive simulacra of human behavior. In Proceedings of the UIST, 2023.
  5. Zhang, J.; Xiang, J.; Yu, Z.; Teng, F.; Chen, X.; Chen, J.; Zhuge, M.; Cheng, X.; Hong, S.; Wang, J.; et al. Aflow: Automating agentic workflow generation. In Proceedings of the ICLR, 2025.
  6. Wu, Q.; Bansal, G.; Zhang, J.; Wu, Y.; Li, B.; Zhu, E.; Jiang, L.; Zhang, X.; Zhang, S.; Liu, J.; et al. AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversations. In Proceedings of the COLM, 2024.
  7. Bei, Y.; Zhang, W.; Wang, S.; Chen, W.; Zhou, S.; Chen, H.; Li, Y.; Bu, J.; Pan, S.; Yu, Y.; et al. Graphs meet ai agents: Taxonomy, progress, and future opportunities. arXiv preprint arXiv:2506.18019 2025.
  8. Liu, Y.; Zhang, G.; Wang, K.; Li, S.; Pan, S.; An, B. Graph-augmented large language model agents: Current progress and future prospects. IEEE Intelligent Systems 2026.
  9. Gao, H.a.; Geng, J.; Hua, W.; Hu, M.; Juan, X.; Liu, H.; Liu, S.; Qiu, J.; Qi, X.; Ren, Q.; et al. A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence. TMLR 2026.
  10. Fang, J.; Peng, Y.; Zhang, X.; Wang, Y.; Yi, X.; Zhang, G.; Xu, Y.; Wu, B.; Liu, S.; Li, Z.; et al. A comprehensive survey of self-evolving ai agents: A new paradigm bridging foundation models and lifelong agentic systems. arXiv preprint arXiv:2508.07407 2025.
  11. Shinn, N.; Cassano, F.; Gopinath, A.; Narasimhan, K.; Yao, S. Reflexion: Language agents with verbal reinforcement learning. In Proceedings of the NeurIPS, 2023.
  12. Zhao, A.; Huang, D.; Xu, Q.; Lin, M.; Liu, Y.J.; Huang, G. Expel: Llm agents are experiential learners. In Proceedings of the AAAI, 2024.
  13. Liu, Z.; Zhang, Y.; Li, P.; Liu, Y.; Yang, D. A dynamic LLM-powered agent network for task-oriented agent collaboration. COLM 2024.
  14. Wang, Z.; Wang, Y.; Liu, X.; Ding, L.; Zhang, M.; Liu, J.; Zhang, M. Agentdropout: Dynamic agent elimination for token-efficient and high-performance llm-based multi-agent collaboration. In Proceedings of the ACL, 2025.
  15. Lu, Y.; Hu, Y.; Zhao, X.; Cao, J. DyTopo: Dynamic Topology Routing for Multi-Agent Reasoning via Semantic Matching. arXiv preprint arXiv:2602.06039 2026.
  16. Hanchen Jiang, E.; Wan, G.; Yin, S.; Li, M.; Wu, Y.; Liang, X.; Li, X.; Sun, Y.; Wang, W.; Chang, K.W.; et al. Dynamic Generation of Multi-LLM Agents Communication Topologies with Graph Diffusion Models. arXiv e-prints 2025.
  17. Fan, W.; Tognoli, T.; Zou, H.P.; Miao, C.; Wang, Y.; Zhang, X. TodyComm: Task-Oriented Dynamic Communication for Multi-Round LLM-based Multi-Agent System. arXiv preprint arXiv:2602.03688 2026.
  18. Leong, H.Y.; Li, Y.; Wu, Y.; Ouyang, W.; Zhu, W.; Gao, J.; Han, W. Amas: Adaptively determining communication topology for llm-based multi-agent system. In Proceedings of the EMNLP: Industry Track, 2025.
  19. Wang, B.; He, W.; Zeng, S.; Xiang, Z.; Xing, Y.; Tang, J.; He, P. Unveiling privacy risks in llm agent memory. In Proceedings of the ACL, 2025.
  20. Zhou, J.; Wang, L.; Yang, X. Guardian: Safeguarding llm multi-agent collaborations with temporal graph modeling. NeurIPS 2025.
  21. Kazemi, S.M.; Goel, R.; Jain, K.; Kobyzev, I.; Sethi, A.; Forsyth, P.; Poupart, P. Representation learning for dynamic graphs: A survey. JMLR 2020.
  22. Qin, M.; Yeung, D.Y. Temporal link prediction: A unified framework, taxonomy, and review. CSUR 2023.
  23. Ur Rahman, A.; Elhag, A.A.; Coon, J.P. A primer on temporal graph learning. CSUR 2025.
  24. Feng, Z.; Wang, R.; Wang, T.; Song, M.; Wu, S.; He, S. A comprehensive survey of dynamic graph neural networks: Models, frameworks, benchmarks, experiments and challenges. TKDE 2026.
  25. Skarding, J.; Gabrys, B.; Musial, K. Foundations and modeling of dynamic networks using dynamic graph neural networks: A survey. IEEE Access 2021.
  26. Rossi, E.; Chamberlain, B.; Frasca, F.; Eynard, D.; Monti, F.; Bronstein, M. Temporal graph networks for deep learning on dynamic graphs. In Proceedings of the ICML GRL Workshop, 2020.
  27. Xu, D.; Ruan, C.; Korpeoglu, E.; Kumar, S.; Achan, K. Inductive representation learning on temporal graphs. In Proceedings of the ICLR, 2020.
  28. Huang, S.; Poursafaei, F.; Danovitch, J.; Fey, M.; Hu, W.; Rossi, E.; Leskovec, J.; Bronstein, M.; Rabusseau, G.; Rabbany, R. Temporal graph benchmark for machine learning on temporal graphs. NeurIPS 2023.
  29. Yu, L.; Sun, L.; Du, B.; Lv, W. Towards better dynamic graph learning: New architecture and unified library. NeurIPS 2023.
  30. Xu, Y.; Zhang, W.; Zhang, Y.; Lin, X.; Xu, X. Unlocking multi-modal potentials for link prediction on dynamic text-attributed graphs. In Proceedings of the AAAI, 2026.
  31. Nie, Z.; Shen, R.; Yu, X.; Yin, B.; Zhang, J.; Hu, X. SkillGraph: Self-Evolving Multi-Agent Collaboration with Multimodal Graph Topology. arXiv preprint arXiv:2604.17503 2026.
  32. Li, K.; Yu, X.; Ni, Z.; Zeng, Y.; Xu, Y.; Zhang, Z.; Li, X.; Sang, J.; Duan, X.; Wang, X.; et al. TiMem: Temporal-Hierarchical Memory Consolidation for Long-Horizon Conversational Agents. arXiv preprint arXiv:2601.02845 2026.
  33. Rasmussen, P.; Paliychuk, P.; Beauvais, T.; Ryan, J.; Chalef, D. Zep: a temporal knowledge graph architecture for agent memory. arXiv preprint arXiv:2501.13956 2025.
  34. Liu, X.; Yu, H.; Zhang, H.; Xu, Y.; Lei, X.; Lai, H.; Gu, Y.; Ding, H.; Men, K.; Yang, K.; et al. Agentbench: Evaluating llms as agents. In Proceedings of the ICLR, 2024.
  35. Ma, C.; Zhang, J.; Zhu, Z.; Yang, C.; Yang, Y.; Jin, Y.; Lan, Z.; Kong, L.; He, J. Agentboard: An analytical evaluation board of multi-turn llm agents. NeurIPS 2024.
  36. Jimenez, C.E.; Yang, J.; Wettig, A.; Yao, S.; Pei, K.; Press, O.; Narasimhan, K. Swe-bench: Can language models resolve real-world github issues? In Proceedings of the ICLR, 2024.
  37. Wu, D.; Wang, H.; Yu, W.; Zhang, Y.; Chang, K.W.; Yu, D. LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory. In Proceedings of the ICLR, 2025.
  38. Ruan, Y.; Dong, H.; Wang, A.; Pitis, S.; Zhou, Y.; Ba, J.; Dubois, Y.; Maddison, C.; Hashimoto, T. Identifying the risks of lm agents with an lm-emulated sandbox. In Proceedings of the ICLR, 2024.
  39. Debenedetti, E.; Zhang, J.; Balunovic, M.; Beurer-Kellner, L.; Fischer, M.; Tramèr, F. Agentdojo: A dynamic environment to evaluate prompt injection attacks and defenses for llm agents. NeurIPS 2024.
  40. Zhang, H.; Huang, J.; Mei, K.; Yao, Y.; Wang, Z.; Zhan, C.; Wang, H.; Zhang, Y. Agent security bench (asb): Formalizing and benchmarking attacks and defenses in llm-based agents. In Proceedings of the ICLR, 2025.
  41. He, X.; Wu, D.; Zhai, Y.; Sun, K. Sentinelagent: Graph-based anomaly detection in multi-agent systems. arXiv preprint arXiv:2505.24201 2025.
  42. Sequeira, R.; Damianakis, S.; Iqbal, U.; Psounis, K. Agent-sentry: Bounding llm agents via execution provenance. arXiv preprint arXiv:2603.22868 2026.
  43. Yue, L.; Bhandari, K.R.; Ko, C.Y.; Patel, D.; Lin, S.; Zhou, N.; Gao, J.; Chen, P.Y.; Pan, S. From static templates to dynamic runtime graphs: A survey of workflow optimization for llm agents. arXiv preprint arXiv:2603.22386 2026.
  44. Yang, C.; Zhou, C.; Xiao, Y.; Dong, S.; Zhuang, L.; Zhang, Y.; Wang, Z.; Hong, Z.; Yuan, Z.; Xiang, Z.; et al. Graph-based Agent Memory: Taxonomy, Techniques, and Applications. arXiv preprint arXiv:2602.05665 2026.
  45. Xu, Y.; Zhang, W.; Lin, X.; Zhang, Y. Unidyg: a unified and effective representation learning approach for large dynamic graphs. TKDE 2025.
  46. You, J.; Du, T.; Leskovec, J. Roland: graph learning framework for dynamic graphs. In Proceedings of the SIGKDD, 2022.
  47. Wu, D.; Xu, Y.; Lin, X.; Zhang, W.; Zhang, Y. Understanding Evolving Graph Structures for Large Discrete-Time Dynamic Graph Representation. PVLDB 2026.
  48. Zhang, J.; Chen, J.; Yang, M.; Feng, A.; Liang, S.; Shao, J.; Ying, R. Dtgb: A comprehensive benchmark for dynamic text-attributed graphs. NeurIPS 2024.
  49. Wang, Y.; Huang, T.; He, C.; Li, Q.; Gao, J. Simple and efficient heterogeneous temporal graph neural network. NeurIPS 2025.
  50. Liu, X.; Peng, Z.; Yi, X.; Xie, X.; Xiang, L.; Liu, Y.; Xu, D. Toolnet: Connecting large language models with massive tools via tool graph. arXiv preprint arXiv:2403.00839 2024.
  51. Xu, W.; Liang, Z.; Mei, K.; Gao, H.; Tan, J.; Zhang, Y. A-mem: Agentic memory for llm agents. NeurIPS 2025.
  52. Heckel, R. Graph transformation in a nutshell. ENTCS 2006.
  53. Taentzer, G. AGG: A graph transformation environment for modeling and validation of software. In Proceedings of the Applications of Graph Transformations with Industrial Relevance, 2004.
  54. Ehrig, H.; Ehrig, K.; Prange, U.; Taentzer, G. Fundamentals of algebraic graph transformation; Springer Science & Business Media, 2006.
  55. Debrouvier, A.; Parodi, E.; Perazzo, M.; Soliani, V.; Vaisman, A. A model and query language for temporal graph databases. The VLDB Journal 2021.
  56. Pu, H.; Song, X.; Zhao, L. SkillOps: Managing LLM Agent Skill Libraries as Self-Maintaining Software Ecosystems. arXiv preprint arXiv:2605.13716 2026.
  57. Yu, J.; Ding, Y.; Sato, H. Dyntaskmas: A dynamic task graph-driven framework for asynchronous and parallel llm-based multi-agent systems. In Proceedings of the ICAPS, 2025.
  58. Zhang, G.; Yue, Y.; Li, Z.; Yun, S.; Wan, G.; Wang, K.; Cheng, D.; Yu, J.X.; Chen, T. Cut the crap: An economical communication pipeline for llm-based multi-agent systems. ICLR 2025.
  59. Wang, Y.; Zhao, J.; Xie, H.; Ma, H.; Lei, Y.; Liu, S.; Song, X.; Zhang, Z.; Zhang, H. MetaGen: Self-Evolving Roles and Topologies for Multi-Agent LLM Reasoning. arXiv preprint arXiv:2601.19290 2026.
  60. Xu, C.; Hu, Y.; Wang, R.; Lin, X.; Wang, W.; Liu, D.; Feng, F. TacoMAS: Test-Time Co-Evolution of Topology and Capability in LLM-based Multi-Agent Systems. arXiv preprint arXiv:2605.09539 2026.
  61. Anokhin, P.; Semenov, N.; Sorokin, A.; Evseev, D.; Kravchenko, A.; Burtsev, M.; Burnaev, E. AriGraph: learning knowledge graph world models with episodic memory for LLM agents. In Proceedings of the IJCAI, 2025.
  62. Paul, S.K.; Sharma, S.; Sareen, N. GAAMA: Graph Augmented Associative Memory for Agents. arXiv preprint arXiv:2603.27910 2026.
  63. Wu, Z.; Zhang, H.; Lin, F.; Xu, W.; Xu, X.; Chen, Y.; Zou, H.P.; Chen, S.; Zhang, W.; Liu, X.; et al. GAM: Hierarchical Graph-based Agentic Memory for LLM Agents. ICLR Workshop MemAgents 2026.
  64. Han, X.; Fan, Y.; Zhao, S.; Wang, H.; Qin, B. GSEM: Graph-based Self-Evolving Memory for Experience Augmented Clinical Reasoning. arXiv preprint arXiv:2603.22096 2026.
  65. Zhang, G.; Fu, M.; Wang, K.; Wan, F.; Yu, M.; Yan, S. G-memory: Tracing hierarchical memory for multi-agent systems. NeurIPS 2025.
  66. Liu, D.; Li, Z.; Du, H.; Wu, X.; Gui, S.; Kuang, Y.; Sun, L. Graph of Skills: Dependency-Aware Structural Retrieval for Massive Agent Skills. arXiv preprint arXiv:2604.05333 2026.
  67. Feng, X.; Song, X.; Li, L.; Liu, G.; Shao, J. SEARL: Joint Optimization of Policy and Tool Graph Memory for Self-Evolving Agents. arXiv preprint arXiv:2604.07791 2026.
  68. Jiang, Y.; Zhou, H.; GU, L.; Han, A.; Li, T. NaviAgent: Bilevel Planning on Tool Navigation Graph for Large-Scale Orchestration. arXiv preprint arXiv:2506.19500 2025.
  69. Wang, S.; Lu, R.; Yang, Z.; Wang, Y.; Zhang, Y.; Xu, L.; Xu, Q.; Yin, G.; Chen, C.; Guan, X. Agentconductor: Topology evolution for multi-agent competition-level code generation. arXiv preprint arXiv:2602.17100 2026.
  70. Zhuge, M.; Wang, W.; Kirsch, L.; Faccio, F.; Khizbullin, D.; Schmidhuber, J. Gptswarm: Language agents as optimizable graphs. In Proceedings of the ICML, 2024.
  71. Zhou, H.; Wan, X.; Sun, R.; Palangi, H.; Iqbal, S.; Vulić, I.; Korhonen, A.; Arık, S.Ö. Multi-agent design: Optimizing agents with better prompts and topologies. arXiv preprint arXiv:2502.02533 2025.
  72. Song, W.; Yue, J.; Pang, Z. Abstral: Automatic design of multi-agent systems through iterative refinement and topology optimization. arXiv preprint arXiv:2603.22791 2026.
  73. Zhang, R.; Zhao, X.; Wang, R.; Chen, S.; Zhang, G.; Zhang, A.; Wang, K.; Wen, Q. Safesieve: From heuristics to experience in progressive pruning for llm-based multi-agent communication. In Proceedings of the AAAI, 2026.
  74. Yang, Y.; Chai, H.; Shao, S.; Song, Y.; Qi, S.; Rui, R.; Zhang, W. Agentnet: Decentralized evolutionary coordination for llm-based multi-agent systems. NeurIPS 2025.
  75. Wang, C.; Lin, H.; Tang, H.; Lin, H.; Ding, W. RUMAD: Reinforcement-Unifying Multi-Agent Debate. In Proceedings of the AAMAS, 2026.
  76. Zhou, Z.; Liu, Z.; Liu, J.; Shao, Q.; Wang, Y.; Shao, K.; Jin, D.; Xu, F. ResMAS: Resilience Optimization in LLM-based Multi-agent Systems. arXiv preprint arXiv:2601.04694 2026.
  77. Wang, S.; Tong, G. DAGP: Difficulty-Aware Graph Pruning for LLM-Based Multi-Agent System. In Proceedings of the CIKM, 2025.
  78. Zhang, G.; Yue, Y.; Sun, X.; Wan, G.; Yu, M.; Fang, J.; Wang, K.; Chen, T.; Cheng, D. G-Designer: Architecting Multi-agent Communication Topologies via Graph Neural Networks. In Proceedings of the ICML. PMLR, 2025.
  79. Yun, S.; Peng, J.; Li, P.; Fan, W.; Chen, J.; Zou, J.; Li, G.; Chen, T. Graph-of-Agents: A Graph-based Framework for Multi-Agent LLM Collaboration. In Proceedings of the ICLR, 2026.
  80. Li, B.; Zhao, Z.; Lee, D.H.; Wang, G. Adaptive graph pruning for multi-agent communication. arXiv preprint arXiv:2506.02951 2025.
  81. Chen, H.; Zheng, X.; Liu, Y.; Jiao, P.; Li, S.; Liu, H.; Zhao, Z.; Xu, Z.; Khalil, I.; Pan, S. Goagent: Group-of-agents communication topology generation for llm-based multi-agent systems. arXiv preprint arXiv:2603.19677 2026.
  82. Sun, R.; Ding, J.; Gong, C.; Gu, T.; Jiang, Y.; Zhang, J.; Pan, L.; Lü, L. TopoDIM: One-shot Topology Generation of Diverse Interaction Modes for Multi-Agent Systems. arXiv preprint arXiv:2601.10120 2026.
  83. Zhang, H.; Shi, Y.; Gu, X.; Zhang, Z.; You, H.; Gan, L.; Yuan, Y.; Huang, J. Hyperagent: Leveraging hypergraphs for topology optimization in multi-agent communication. In Proceedings of the AAMAS, 2026.
  84. Wu, X.; Yang, C.; Lin, X.; Xu, C.; Jiang, X.; Sun, Y.; Xiong, H.; Li, J.; Guo, J. Think-on-Graph 3.0: Efficient and Adaptive LLM Reasoning on Heterogeneous Graphs via Multi-Agent Dual-Evolving Context Retrieval. arXiv preprint arXiv:2509.21710 2025.
  85. Van, H.P.; Hieu, N.M.; Tuan, K.P.T.; Hai, N.L.; Van, L.N.; Diep, N.T.N.; Le, T. MemORAI: Memory Organization and Retrieval via Adaptive Graph Intelligence for LLM Conversational Agents. arXiv preprint arXiv:2605.01386 2026.
  86. Li, S.; Liu, Y.; Wen, Q.; Zhang, C.; Pan, S. Assemble your crew: Automatic multi-agent communication topology design via autoregressive graph generation. In Proceedings of the AAAI, 2026.
  87. Hu, Y.; Cai, Y.; Du, Y.; Zhu, X.; Liu, X.; Yu, Z.; Hou, Y.; Tang, S.; Chen, S. Self-evolving multi-agent collaboration networks for software development. In Proceedings of the ICLR, 2025.
  88. Wang, S.; Zhang, G.; Yu, M.; Wan, G.; Meng, F.; Guo, C.; Wang, K.; Wang, Y. G-safeguard: A topology-guided security lens and treatment on llm-based multi-agent systems. In Proceedings of the ACL, 2025.
  89. Kang, J.; Ji, M.; Zhao, Z.; Bai, T. Memory os of ai agent. In Proceedings of the EMNLP, 2025.
  90. Lin, M.; Zhang, Z.; Lu, H.; Liu, H.; Tang, X.; He, Q.; Zhang, X.; Wang, S. MemMA: Coordinating the Memory Cycle through Multi-Agent Reasoning and In-Situ Self-Evolution. arXiv preprint arXiv:2603.18718 2026.
  91. Zhang, S.; Wang, J.; Zhou, R.; Liao, J.; Feng, Y.; Li, Z.; Zheng, Y.; Zhang, W.; Wen, Y.; Li, Z.; et al. Memrl: Self-evolving agents via runtime reinforcement learning on episodic memory. arXiv preprint arXiv:2601.03192 2026.
  92. Maharana, A.; Lee, D.H.; Tulyakov, S.; Bansal, M.; Barbieri, F.; Fang, Y. Evaluating very long-term conversational memory of llm agents. In Proceedings of the ACL, 2024.
  93. Tan, H.; Zhang, Z.; Ma, C.; Chen, X.; Dai, Q.; Dong, Z. Membench: Towards more comprehensive evaluation on the memory of llm-based agents. In Proceedings of the Findings of ACL, 2025.
  94. Wang, G.; Xie, Y.; Jiang, Y.; Mandlekar, A.; Xiao, C.; Zhu, Y.; Fan, L.; Anandkumar, A. Voyager: An Open-Ended Embodied Agent with Large Language Models. TMLR 2024.
  95. Zhang, H.; Fan, S.; Zou, H.P.; Chen, Y.; Wang, Z.; Zhou, J.; Li, C.; Huang, W.C.; Yao, Y.; Zheng, K.; et al. Coevoskills: Self-evolving agent skills via co-evolutionary verification. arXiv preprint arXiv:2604.01687 2026.
  96. Alzubi, S.; Provenzano, N.; Bingham, J.; Chen, W.; Vu, T. Evoskill: Automated skill discovery for multi-agent systems. arXiv preprint arXiv:2603.02766 2026.
  97. Xia, P.; Chen, J.; Wang, H.; Liu, J.; Zeng, K.; Wang, Y.; Han, S.; Zhou, Y.; Zhao, X.; Chen, H.; et al. SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning. In Proceedings of the ICLR Workshop on Memory for LLM-Based Agentic Systems, 2026.
  98. Kuroki, S.; Nakamura, T.; Akiba, T.; Tang, Y. Agent skill acquisition for large language models via cycleqd. In Proceedings of the ICLR, 2025.
  99. Schick, T.; Dwivedi-Yu, J.; Dessì, R.; Raileanu, R.; Lomeli, M.; Hambro, E.; Zettlemoyer, L.; Cancedda, N.; Scialom, T. Toolformer: Language models can teach themselves to use tools. In Proceedings of the NeurIPS, 2023.
  100. Patil, S.G.; Zhang, T.; Wang, X.; Gonzalez, J.E. Gorilla: Large language model connected with massive apis. NeurIPS 2024.
  101. Qin, Y.; Liang, S.; Ye, Y.; Zhu, K.; Yan, L.; Lu, Y.; Lin, Y.; Cong, X.; Tang, X.; Qian, B.; et al. Toolllm: Facilitating large language models to master 16000+ real-world apis. In Proceedings of the ICLR, 2024.
  102. Li, M.; Zhao, Y.; Yu, B.; Song, F.; Li, H.; Yu, H.; Li, Z.; Huang, F.; Li, Y. Api-bank: A comprehensive benchmark for tool-augmented llms. In Proceedings of the EMNLP, 2023.
  103. Lu, J.; Holleis, T.; Zhang, Y.; Aumayer, B.; Nan, F.; Bai, H.; Ma, S.; Ma, S.; Li, M.; Yin, G.; et al. Toolsandbox: A stateful, conversational, interactive evaluation benchmark for llm tool use capabilities. In Proceedings of the Findings of NAACL, 2025.
  104. Patil, S.G.; Mao, H.; Yan, F.; Ji, C.C.J.; Suresh, V.; Stoica, I.; Gonzalez, J.E. The Berkeley Function Calling Leaderboard (BFCL): From Tool Use to Agentic Evaluation of Large Language Models. In Proceedings of the ICML, 2025.
  105. Xiu, Z.; Sun, D.Q.; Cheng, K.; Patel, M.; Zhang, Y.; Lu, J.; Attia, O.; Vemulapalli, R.; Tuzel, O.; Cao, M.; et al. ASTRA-bench: Evaluating Tool-Use Agent Reasoning and Action Planning with Personal User Context. arXiv preprint arXiv:2603.01357 2026.
  106. Liu, S.; Chen, T.; Amiri, R.; Amato, C. Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic. arXiv preprint arXiv:2601.21972 2026.
  107. Gao, D.; Li, Z.; Pan, X.; Kuang, W.; Ma, Z.; Qian, B.; Wei, F.; Zhang, W.; Xie, Y.; Chen, D.; et al. Agentscope: A flexible yet robust multi-agent platform. arXiv preprint arXiv:2402.14034 2024.
  108. Li, Y.; Yang, X.; Yang, X.; Wang, X.; Liu, W.; Bian, J. R&D-Agent-Quant: a multi-agent framework for data-centric factors and model joint optimization. NeurIPS 2025.
  109. Yu, P.; Chen, G.; Wang, J. Table-critic: A multi-agent framework for collaborative criticism and refinement in table reasoning. In Proceedings of the ACL, 2025.
  110. Zhang, X.; Cui, Y.; Wang, G.; Qiu, W.; Li, Z.; Han, F.; Huang, Y.; Qiu, H.; Zhu, B.; He, P. Verified multi-agent orchestration: A plan-execute-verify-replan framework for complex query resolution. arXiv preprint arXiv:2603.11445 2026.
  111. Nizar, F.; Lumer, E.; Gulati, A.; Basavaraju, P.H.; Subbiah, V.K. Agent-as-a-Graph: Knowledge Graph-Based Tool and Agent Retrieval for LLM Multi-Agent Systems. arXiv preprint arXiv:2511.18194 2025.
  112. Feng, T.; Zhang, H.; Lei, Z.; Han, P.; You, J. GraphPlanner: Graph Memory-Augmented Agentic Routing for Multi-Agent LLMs. In Proceedings of the ICLR, 2026.
  113. Hu, S.; Lu, C.; Clune, J. Automated design of agentic systems. In Proceedings of the ICLR, 2025.
  114. Shang, Y.; Li, Y.; Zhao, K.; Ma, L.; Liu, J.; Xu, F.; Li, Y. Agentsquare: Automatic llm agent search in modular design space. In Proceedings of the ICLR, 2025.
  115. Zhong, W.; Guo, L.; Gao, Q.; Ye, H.; Wang, Y. Memorybank: Enhancing large language models with long-term memory. In Proceedings of the AAAI, 2024.
  116. Wang, Y.; Xu, Z.; Huang, Y.; Wang, X.; Song, Z.; Gao, L.; Wang, C.; Tang, R.; Zhao, Y.; Cohan, A.; et al. Dyflow: Dynamic workflow framework for agentic reasoning. NeurIPS 2025.
  117. Zhang, Z.; Rossi, R.A.; Yu, T.; Dernoncourt, F.; Zhang, R.; Gu, J.; Kim, S.; Chen, X.; Wang, Z.; Lipka, N. Vipact: Visual-perception enhancement via specialized vlm agent collaboration and tool-use. In Proceedings of the AAAI, 2026.
  118. Mialon, G.; Fourrier, C.; Wolf, T.; LeCun, Y.; Scialom, T. Gaia: a benchmark for general ai assistants. In Proceedings of the ICLR, 2024.
  119. Yao, S.; Shinn, N.; Razavi, P.; Narasimhan, K. τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. arXiv preprint arXiv:2406.12045 2024.
  120. Xie, T.; Zhang, D.; Chen, J.; Li, X.; Zhao, S.; Cao, R.; Hua, T.J.; Cheng, Z.; Shin, D.; Lei, F.; et al. Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments. NeurIPS 2024.
  121. Deng, X.; Gu, Y.; Zheng, B.; Chen, S.; Stevens, S.; Wang, B.; Sun, H.; Su, Y. Mind2web: Towards a generalist agent for the web. NeurIPS 2023.
  122. Qiao, S.; Fang, R.; Qiu, Z.; Wang, X.; Zhang, N.; Jiang, Y.; Xie, P.; Huang, F.; Chen, H. Benchmarking agentic workflow generation. In Proceedings of the ICLR, 2025.
  123. Li, Y.; Xu, Y.; Lin, X.; Zhang, W.; Zhang, Y. Ranking on dynamic graphs: An effective and robust band-pass disentangled approach. In Proceedings of the WWW, 2025.
  124. Chen, Y.; Jiang, J.; Sun, S.; He, B.; Chen, M. Rush: Real-time burst subgraph detection in dynamic graphs. PVLDB 2024.
  125. Wang, J.; Liao, X.; Wu, W. TopoEvo: A Topology-Aware Self-Evolving Multi-Agent Framework for Root Cause Analysis in Microservices. arXiv preprint arXiv:2605.15611 2026.
  126. Madaan, A.; Tandon, N.; Gupta, P.; Hallinan, S.; Gao, L.; Wiegreffe, S.; Alon, U.; Dziri, N.; Prabhumoye, S.; Yang, Y.; et al. Self-refine: Iterative refinement with self-feedback. NeurIPS 2023.
  127. Liu, H.; Shou, C.; Liu, X.; Wen, H.; Chen, Y.; Fang, R.J.; Feng, Y. Synthesizing Multi-Agent Harnesses for Vulnerability Discovery. arXiv preprint arXiv:2604.20801 2026.
  128. Li, J.; Lai, Y.; Li, W.; Ren, J.; Zhang, M.; Kang, X.; Wang, S.; Li, P.; Zhang, Y.Q.; Ma, W.; et al. Agent hospital: A simulacrum of hospital with evolvable medical agents. arXiv preprint arXiv:2405.02957 2024.
  129. Qiao, S.; Zhang, N.; Fang, R.; Luo, Y.; Zhou, W.; Jiang, Y.; Lv, C.; Chen, H. Autoact: Automatic agent learning from scratch for qa via self-planning. In Proceedings of the ACL, 2024.
  130. Ghosh, S.; Simkin, B.; Shiarlis, K.; Nandi, S.; Zhao, D.; Fiedler, M.; Bazinska, J.; Pope, N.; Prabhu, R.; Rohrer, D.; et al. A Safety and Security Framework for Real-World Agentic Systems. arXiv preprint arXiv:2511.21990 2025.
  131. Yin, B.; Li, Q.; Wang, X. On-Policy Self-Evolution via Failure Trajectories for Agentic Safety Alignment. arXiv preprint arXiv:2605.11882 2026.
  132. Longa, A.; Lachi, V.; Santin, G.; Bianchini, M.; Lepri, B.; Lio, P.; Scarselli, F.; Passerini, A.; et al. Graph Neural Networks for Temporal Graphs: State of the Art, Open Challenges, and Opportunities. TMLR 2023.
  133. Zhang, S.; Xiong, Y.; Tang, Y.; Xu, J.; Chen, X.; Gu, Z.; Zheng, X.; Jia, Z.; Zhang, J. Unifying text semantics and graph structures for temporal text-attributed graphs with large language models. NeurIPS 2025.
  134. Zhang, L.; Zhao, L.; Qin, S.; Pfoser, D.; Ling, C. TG-GAN: Continuous-time temporal graph deep generative models with time-validity constraints. In Proceedings of the WWW, 2021.
  135. Gupta, S.; Manchanda, S.; Bedathur, S.; Ranu, S. Tigger: Scalable generative modelling for temporal interaction graphs. In Proceedings of the AAAI, 2022.
  136. Liu, H.; Di, S.; Li, H.; Jian, X.; Wang, Y.; Chen, L. A Selective Learning Method for Temporal Graph Continual Learning. In Proceedings of the ICML, 2025.
  137. Zhang, P.; Yan, Y.; Li, C.; Wang, S.; Xie, X.; Song, G.; Kim, S. Continual learning on dynamic graphs via parameter isolation. In Proceedings of the SIGIR, 2023.
  138. Zhang, Z.; Wang, X.; Zhang, Z.; Li, H.; Qin, Z.; Zhu, W. Dynamic graph neural networks under spatio-temporal distribution shift. NeurIPS 2022.
  139. Zhang, Z.; Wang, X.; Zhang, Z.; Qin, Z.; Wen, W.; Xue, H.; Li, H.; Zhu, W. Spectral invariant learning for dynamic graphs under distribution shifts. NeurIPS 2023.
  140. Han, Z.; Chen, P.; Ma, Y.; Tresp, V. Explainable subgraph reasoning for forecasting on temporal knowledge graphs. In Proceedings of the ICLR, 2021.
  141. Jin, W.; Qu, M.; Jin, X.; Ren, X. Recurrent event network: Autoregressive structure inferenceover temporal knowledge graphs. In Proceedings of the EMNLP, 2020.
  142. Zheng, L.; Li, Z.; Li, J.; Li, Z.; Gao, J. AddGraph: Anomaly Detection in Dynamic Graph Using Attention-based Temporal GCN. In Proceedings of the IJCAI, 2019.
  143. Liu, Y.; Pan, S.; Wang, Y.G.; Xiong, F.; Wang, L.; Chen, Q.; Lee, V.C. Anomaly detection in dynamic graphs via transformer. TKDE 2023.
  144. Zhang, H.; Wu, B.; Yang, X.; Yuan, X.; Liu, X.; Yi, X. Dynamic graph unlearning: a general and efficient post-processing method via gradient transformation. In Proceedings of the WWW, 2025.
  145. Guo, Q.; Sun, W.; Wang, W. Spatio-Temporal Graph Unlearning. arXiv preprint arXiv:2511.09404 2025.
  146. Xia, W.; Lai, M.; Shan, C.; Zhang, Y.; Dai, X.; Li, X.; Li, D. Explaining temporal graph models through an explorer-navigator framework. In Proceedings of the ICLR, 2023.
  147. Zhao, K.; Zhang, L. Causality-inspired spatial-temporal explanations for dynamic graph neural networks. In Proceedings of the ICLR, 2024.
  148. Cong, W.; Zhang, S.; Kang, J.; Yuan, B.; Wu, H.; Zhou, X.; Tong, H.; Mahdavi, M. DO WE REALLY NEED COMPLICATED MODEL ARCHITECTURES FOR TEMPORAL NETWORKS? In Proceedings of the ICLR, 2023.
  149. Tian, Y.; Qi, Y.; Guo, F. Freedyg: Frequency enhanced continuous-time dynamic graph model for link prediction. In Proceedings of the ICLR, 2024.
  150. Liu, H.; Zhang, L.; Wang, R.; Zheng, T.; Wu, S.; Yao, C.; Song, M. Salom: Structure aware temporal graph networks with long-short memory updater. NeurIPS 2025.
  151. Gao, S.; Li, Y.; Zhang, X.; Shen, Y.; Shao, Y.; Chen, L. Simple: Efficient temporal graph neural network training at scale with dynamic data placement. SIGMOD 2024.
  152. Gao, S.; Li, Y.; Shen, Y.; Shao, Y.; Chen, L. Etc: Efficient training of temporal graph neural networks over large-scale dynamic graphs. PVLDB 2024.
  153. Cheng, K.; Peng, L.; Wang, P.; Chang, H.; Ye, J.; Du, B. On the scalability of temporal relative positional encoding for dynamic link prediction. In Proceedings of the KDD, 2025.
  154. Zou, T.; Mao, Y.; Ye, J.; Du, B. Repeat-aware neighbor sampling for dynamic graph learning. In Proceedings of the KDD, 2024.
  155. Luo, Y.; Li, P. Neighborhood-aware scalable temporal network representation learning. In Proceedings of the LoG, 2022.
  156. Gravina, A.; Lovisotto, G.; Gallicchio, C.; Bacciu, D.; Grohnfeldt, C. Long range propagation on continuous-time dynamic graphs. In Proceedings of the ICML, 2024.
  157. Li, H.; Li, C.; Feng, K.; Yuan, Y.; Wang, G.; Zha, H. Robust knowledge adaptation for dynamic graph neural networks. TKDE 2024.
  158. Cheng, K.; Linzhi, P.; Ye, J.; Sun, L.; Du, B. Co-neighbor encoding schema: A light-cost structure encoding method for dynamic link prediction. In Proceedings of the KDD, 2024.
  159. Wu, Y.; Liao, L.; Fang, Y. Retrieval augmented generation for dynamic graph modeling. In Proceedings of the SIGIR, 2025.
  160. Zou, T.; Yu, L.; Sun, L.; Du, B.; Wang, D.; Zhuang, F. Event-based dynamic graph representation learning for patent application trend prediction. TKDE 2024.
  161. Li, J.; Wu, R.; Jin, X.; Ma, B.; Chen, L.; Zheng, Z. State space models on temporal graphs: A first-principles study. NeurIPS 2024.
  162. Karmim, Y.; Lafon, M.; S’niehotta, R.F.; Thome, N. Supra-laplacian encoding for transformer on dynamic graphs. NeurIPS 2024.
  163. Goyal, P.; Chhetri, S.R.; Canedo, A. dyngraph2vec: Capturing network dynamics using dynamic graph representation learning. KBS 2020.
  164. Sankar, A.; Wu, Y.; Gou, L.; Zhang, W.; Yang, H. Dysat: Deep neural representation learning on dynamic graphs via self-attention networks. In Proceedings of the WSDM, 2020.
  165. Pareja, A.; Domeniconi, G.; Chen, J.; Ma, T.; Suzumura, T.; Kanezashi, H.; Kaler, T.; Schardl, T.; Leiserson, C. Evolvegcn: Evolving graph convolutional networks for dynamic graphs. In Proceedings of the AAAI, 2020.
  166. Gao, C.; Zhu, J.; Zhang, F.; Wang, Z.; Li, X. A novel representation learning for dynamic graphs based on graph convolutional networks. IEEE transactions on cybernetics 2022.
  167. Bai, Q.; Nie, C.; Zhang, H.; Zhao, D.; Yuan, X. Hgwavenet: A hyperbolic graph neural network for temporal link prediction. In Proceedings of the WWW, 2023.
  168. Qin, X.; Sheikh, N.; Lei, C.; Reinwald, B.; Domeniconi, G. Seign: A simple and efficient graph neural network for large dynamic graphs. In Proceedings of the ICDE, 2023.
  169. Zhang, G.; Ye, T.; Jin, D.; Li, Y. An attentional multi-scale co-evolving model for dynamic link prediction. In Proceedings of the WWW, 2023.
  170. Li, J.; Yu, Z.; Zhu, Z.; Chen, L.; Yu, Q.; Zheng, Z.; Tian, S.; Wu, R.; Meng, C. Scaling up dynamic graph representation learning via spiking neural networks. In Proceedings of the AAAI, 2023.
  171. Qin, M.; Zhang, C.; Bai, B.; Zhang, G.; Yeung, D.Y. High-quality temporal link prediction for weighted dynamic graphs via inductive embedding aggregation. TKDE 2023.
  172. Roy, A.; Yan, N.; Mortazavi, M. Llm-driven knowledge distillation for dynamic text-attributed graphs. arXiv preprint arXiv:2502.10914 2025.
  173. Wang, Y.; Li, J.; Zhang, Z. Global-Recent Semantic Reasoning on Dynamic Text-Attributed Graphs with Large Language Models. arXiv preprint arXiv:2509.18742 2025.
  174. Lei, R.; Ji, J.; Ding, H.; Yi, L.; Wei, Z.; Liu, Y.; Hong, C. Exploring the potential of large language models as predictors in dynamic text-attributed graphs. arXiv preprint arXiv:2503.03258 2025.
  175. Zhou, D.; Zheng, L.; Han, J.; He, J. A data-driven graph generative model for temporal interaction networks. In Proceedings of the SIGKDD, 2020.
  176. Hosseini, R.; Simini, F.; Vishwanath, V.; Hoffmann, H. A deep probabilistic framework for continuous time dynamic graph generation. In Proceedings of the AAAI, 2025.
  177. Li, F.; Wang, X.; Cheng, D.; Chen, C.; Zhang, Y.; Lin, X. Efficient dynamic attributed graph generation. In Proceedings of the ICDE. IEEE, 2025.
  178. Peng, J.; Ji, J.; Lei, R.; Wei, Z.; Liu, Y.; Hong, C. GDGB: A Benchmark for Generative Dynamic Text-Attributed Graph Learning. In Proceedings of the ICLR, 2026.
  179. Feng, K.; Li, C.; Zhang, X.; Zhou, J. TOWARDS OPEN TEMPORAL GRAPH NEURAL NETWORKS. In Proceedings of the ICLR, 2023.
  180. Yuan, H.; Sun, Q.; Fu, X.; Zhang, Z.; Ji, C.; Peng, H.; Li, J. Environment-aware dynamic graph learning for out-of-distribution generalization. NeurIPS 2023.
  181. Sun, Q.; Luo, J.; Yuan, H.; Fu, X.; Peng, H.; Li, J.; Yu, P.S. Evolving Graph Learning for Out-of-Distribution Generalization in Non-stationary Environments. TPAMI 2026.
  182. Dasgupta, S.S.; Ray, S.N.; Talukdar, P. Hyte: Hyperplane-based temporally aware knowledge graph embedding. In Proceedings of the EMNLP, 2018.
  183. Wu, J.; Cao, M.; Cheung, J.C.K.; Hamilton, W.L. Temp: Temporal message passing for temporal knowledge graph completion. In Proceedings of the EMNLP, 2020.
  184. Lacroix, T.; Obozinski, G.; Usunier, N. Tensor Decompositions for Temporal Knowledge Base Completion. In Proceedings of the ICLR, 2020.
  185. Trivedi, R.; Dai, H.; Wang, Y.; Song, L. Know-evolve: Deep temporal reasoning for dynamic knowledge graphs. In Proceedings of the ICML, 2017.
  186. Zhu, C.; Chen, M.; Fan, C.; Cheng, G.; Zhang, Y. Learning from history: Modeling temporal knowledge graphs with sequential copy-generation networks. In Proceedings of the AAAI, 2021.
  187. Li, Z.; Jin, X.; Li, W.; Guan, S.; Guo, J.; Shen, H.; Wang, Y.; Cheng, X. Temporal knowledge graph reasoning based on evolutional representation learning. In Proceedings of the SIGIR, 2021.
  188. Chang, H.; Wu, J.; Tao, Z.; Ma, Y.; Huang, X.; Chua, T.S. Integrate Temporal Graph Learning into LLM-based Temporal Knowledge Graph Model. arXiv preprint arXiv:2501.11911 2025.
  189. Xia, Y.; Wang, D.; Liu, Q.; Wang, L.; Wu, S.; Zhang, X. Enhancing temporal knowledge graph forecasting with large language models via chain-of-history reasoning. arXiv preprint arXiv:2402.14382 2024.
  190. Liao, R.; Jia, X.; Li, Y.; Ma, Y.; Tresp, V. Gentkg: Generative forecasting on temporal knowledge graph with large language models. In Proceedings of the Findings of NAACL, 2024.
  191. Greshake, K.; Abdelnabi, S.; Mishra, S.; Endres, C.; Holz, T.; Fritz, M. Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection. In Proceedings of the AISec, 2023.
  192. Ekle, O.A.; Eberle, W. Anomaly detection in dynamic graphs: A comprehensive survey. TKDD 2024.
  193. Cai, L.; Chen, Z.; Luo, C.; Gui, J.; Ni, J.; Li, D.; Chen, H. Structural temporal graph neural networks for anomaly detection in dynamic graphs. In Proceedings of the CIKM, 2021.
  194. Fang, L.; Feng, K.; Gui, J.; Feng, S.; Hu, A. Anonymous edge representation for inductive anomaly detection in dynamic bipartite graph. PVLDB 2023.
  195. Lee, J.; Kim, S.; Shin, K. Slade: Detecting dynamic anomalies in edge streams without labels via self-supervised learning. In Proceedings of the KDD, 2024.
  196. Wang, T.; Luo, D.; Cheng, W.; Chen, H.; Zhang, X. DyExplainer: Self-explainable dynamic graph neural network with sparse attentions. TKDD 2025.
  197. Carass, A.; Roy, S.; Gherman, A.; Reinhold, J.C.; Jesson, A.; Arbel, T.; Maier, O.; Handels, H.; Ghafoorian, M.; Platel, B.; et al. Evaluating white matter lesion segmentations with refined Sørensen-Dice analysis. Scientific reports 2020.
  198. Meng, K.; Bau, D.; Andonian, A.; Belinkov, Y. Locating and editing factual associations in gpt. NeurIPS 2022.
  199. Meng, K.; Sharma, A.S.; Andonian, A.J.; Belinkov, Y.; Bau, D. Mass-Editing Memory in a Transformer. In Proceedings of the ICLR, 2023.
  200. Mitchell, E.; Lin, C.; Bosselut, A.; Finn, C.; Manning, C.D. Fast Model Editing at Scale. In Proceedings of the ICLR, 2022.
  201. Yao, Y.; Wang, P.; Tian, B.; Cheng, S.; Li, Z.; Deng, S.; Chen, H.; Zhang, N. Editing large language models: Problems, methods, and opportunities. In Proceedings of the EMNLP, 2023.
  202. Zhou, S.; Xu, F.F.; Zhu, H.; Zhou, X.; Lo, R.; Sridhar, A.; Cheng, X.; Ou, T.; Bisk, Y.; Fried, D.; et al. Webarena: A realistic web environment for building autonomous agents. In Proceedings of the ICLR, 2024.
  203. Koh, J.Y.; Lo, R.; Jang, L.; Duvvur, V.; Lim, M.; Huang, P.Y.; Neubig, G.; Zhou, S.; Salakhutdinov, R.; Fried, D. Visualwebarena: Evaluating multimodal agents on realistic visual web tasks. In Proceedings of the ACL, 2024.
  204. He, H.; Yao, W.; Ma, K.; Yu, W.; Dai, Y.; Zhang, H.; Lan, Z.; Yu, D. Webvoyager: Building an end-to-end web agent with large multimodal models. In Proceedings of the ACL, 2024.
  205. Xie, J.; Zhang, K.; Chen, J.; Zhu, T.; Lou, R.; Tian, Y.; Xiao, Y.; Su, Y. TravelPlanner: A Benchmark for Real-World Planning with Language Agents. In Proceedings of the ICML, 2024.
  206. Trivedi, H.; Khot, T.; Hartmann, M.; Manku, R.; Dong, V.; Li, E.; Gupta, S.; Sabharwal, A.; Balasubramanian, N. Appworld: A controllable world of apps and people for benchmarking interactive coding agents. In Proceedings of the ACL, 2024.
  207. OpenAI. Introducing SWE-bench Verified. https://openai.com/index/introducing-swe-bench-verified/, 2024.
  208. Ren, Q.; Zou, S.; Huang, S.; Zhang, Z.; Shi, K.; Fang, Z.; Zhao, Y.; Zeng, Y.; Su, Q.; Chen, L.; et al. SaaSBench: Exploring the Boundaries of Coding Agents in Long-Horizon Enterprise SaaS Engineering. arXiv preprint arXiv:2605.17526 2026.
  209. Poursafaei, F.; Huang, S.; Pelrine, K.; Rabbany, R. Towards better evaluation for dynamic link prediction. NeurIPS 2022.
  210. Chen, M.; Zhang, Z.; Wang, T.; Backes, M.; Humbert, M.; Zhang, Y. Graph unlearning. In Proceedings of the CCS, 2022.
  211. Chien, E.; Pan, C.; Milenkovic, O. Efficient model updates for approximate unlearning of graph-structured data. In Proceedings of the ICLR, 2023.
  212. Wu, J.; Yang, Y.; Qian, Y.; Sui, Y.; Wang, X.; He, X. Gif: A general graph unlearning strategy via influence function. In Proceedings of the WWW, 2023.
  213. Cheng, J.; Dasoulas, G.; He, H.; Agarwal, C.; Zitnik, M. GNNDelete: A General Strategy for Unlearning in Graph Neural Networks. In Proceedings of the ICLR, 2023.
  214. Andriushchenko, M.; Souly, A.; Dziemian, M.; Duenas, D.; Lin, M.; Wang, J.; Hendrycks, D.; Zou, A.; Kolter, Z.; Fredrikson, M.; et al. Agentharm: A benchmark for measuring harmfulness of llm agents. In Proceedings of the ICLR, 2025.
  215. Yuan, H.; Yu, H.; Gui, S.; Ji, S. Explainability in graph neural networks: A taxonomic survey. TPAMI 2023.
  216. Zhang, H.; Yue, H.; Feng, T.; Long, Q.; Bao, J.; Jin, B.; Zhang, W.; Li, X.; You, J.; Qin, C.; et al. Learning Query-Aware Budget-Tier Routing for Runtime Agent Memory. arXiv preprint arXiv:2602.06025 2026.
1
We use `dynamic graph’ and `temporal graph’ interchangeably, following the convention in the literature.
Figure 1. Positioning of this survey relative to existing LLM-agent, graph-agent, and dynamic-graph surveys.
Figure 1. Positioning of this survey relative to existing LLM-agent, graph-agent, and dynamic-graph surveys.
Preprints 219785 g001
Figure 2. Method-level taxonomy of 46 representative self-evolving-agent methods using dynamic topologies and graphs.
Figure 2. Method-level taxonomy of 46 representative self-evolving-agent methods using dynamic topologies and graphs.
Preprints 219785 g002
Table 1. Representative dynamic or time-aware agent-evolution mechanisms as graph rewrites. The Scope column uses graph-level terms to denote the affected agent-state subgraph: memory, skill, workflow, communication, agent, or trace graph.
Table 1. Representative dynamic or time-aware agent-evolution mechanisms as graph rewrites. The Scope column uses graph-level terms to denote the affected agent-state subgraph: memory, skill, workflow, communication, agent, or trace graph.
Mechanism Trigger Rewrite Scope Persistence Examples
A.1 Node and feature evolution
Memory update New evidence Insert/Merge/FeatureUpdate Memory graph Long Zep [33], TiMem [32]
Skill update Trajectory/feedback Insert; FeatureUpdate Skill graph Long SkillOps [56]
A.2 Edge and topology evolution
Workflow rewrite Task change Link/Unlink/Rewire Workflow graph Medium AFlow [5], DynTaskMAS [57]
Communication pruning Redundant messages Unlink/Rewire Communication graph Medium AgentPrune [58], AgentDropout [14]
Topology routing Round context Link/Rewire Communication graph Medium GTD [16]
A.3 Read-only subgraph activation
Team activation Query/round Activate Agent graph Temporary DyLAN [13], DyTopo [15]
A.4 Cross-component co-evolution
Workflow→team Expertise change Rewire+cascade Workflow-Agent graph Medium MetaGen [59], TacoMAS [60]
Safety propagation Unsafe trace FeatureUpdate + cascade Trace graph Medium GUARDIAN [20], SentinelAgent [41]
Table 2. Dynamic-graph method families for agent-evolution infrastructure. Each row lists two representative dynamic-graph methods and the corresponding transfer target in self-evolving agents.
Table 2. Dynamic-graph method families for agent-evolution infrastructure. Each row lists two representative dynamic-graph methods and the corresponding transfer target in self-evolving agents.
Family Representative methods Agent capability Required adaptation Naive failure modes
Representation learning on evolving graphs
B.1 CTDGs & DTDGs TGN [26]; DyGFormer [29] Update prediction; activation; cascades Typed rewrite events; temporal negatives Temporal leakage; unstable embeddings
B.2 DyTAGs MoMent [30]; CROSS [133] Text-aware memory and skill activation Selective re-encoding; text–time alignment Stale text embeddings
Generative modeling of temporal structure
B.3 DyG generation TG-GAN [134]; TIGGER [135] Workflow and topology synthesis Typed schema constraints; valid decoding Invalid tools or communication links
Learning under streams and temporal shift
B.4 Continual learning LTF [136]; PI-GNN [137] Durable skill and memory encoders Context-aware replay; update isolation Rare skills are forgotten
B.5 OOD DIDA [138]; SILD [139] Robust activation and update Splits by time, tool, and user cohort Deployment drift is hidden
B.6 TKG reasoning xERTE [140]; RE-Net [141] Temporal memory reasoning Text evidence with timestamped provenance Language evidence is ignored
Diagnosis, removal, and explanation on dynamic graphs
B.7 Anomaly detection AddGraph [142]; TADDY [143] Unsafe-rewrite and drift detection Calibration on benign evolution bursts Normal adaptation is flagged
B.8 DyG unlearning GradientTransformation [144]; CallosumNet [145] Deletion, rollback, influence removal Versioned provenance; shared-state isolation Rollback damages shared skills
B.9 T-GNN explanation T-GNNExplainer [146]; Causal Explanation [147] Audit and attribution Event-level explanations over rewrite traces Triggering events are missed
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.