Preprint
Review

This version is not peer-reviewed.

LLM Agents: A Survey

Submitted:

16 July 2026

Posted:

05 August 2026

You are already at the latest version

Abstract
Large language models have progressed from text predictors into the reasoning core of autonomous agents, systems that plan, invoke tools, maintain memory, and coordinate with other agents to pursue goals over long horizons. The resulting literature has grown explosively but unevenly, splitting into subcommunities that rarely cite one another. This survey imposes a function-first taxonomy that organizes the field by the architectural role each component plays rather than by the framework that introduced it: four component pillars (planning and reasoning, memory, tool use, and multi-agent coordination) grounded in two context dimensions (interactive environments and application domains) and assessed along two cross-cutting concerns, evaluation and safety, which we treat as first-class pillars rather than afterthoughts. Synthesizing more than two hundred works from 2021 through early 2026, including the reasoning-native models, computer-use agents, and agent-security research absent from earlier surveys, we find a field pulling in two directions: its components are consolidating, as prompting gives way to trained reasoning and scaffolds migrate into the models, while its frontier, grounding agents in real environments, evaluating them honestly for cost and reliability, and defending them against attack, remains wide open. Throughout, we give documented negative results on self-correction, multi-agent debate, and agentic scaffolding equal weight with positive claims, support the argument with small controlled experiments of our own, and show that capability and safety are distinct axes that must be measured apart. A continuously updated reading list of the surveyed papers is maintained at https://github.com/js-lee-AI/awesome-llm-agent-papers.
Keywords: 
;  ;  ;  ;  ;  ;  ;  ;  

1. Introduction

Large language models (LLMs) have progressed from single-turn text predictors into the reasoning core of autonomous agents1: systems that plan, invoke external tools, accumulate memory across steps, and coordinate with other agents to pursue goals over long horizons. Three base-model capabilities made this shift possible. Multi-step reasoning turned out to be elicitable by prompting alone (Wei et al. 2022), and even without exemplars (Kojima et al. 2022), so a model could be asked to deliberate before committing to an action. Instruction tuning made that deliberation steerable by ordinary task descriptions. And models proved able to operate interfaces they were never trained on, learning when and how to call external APIs from a handful of demonstrations (Schick et al. 2023). ReAct (Yao et al. 2022b) composed these capabilities into the loop that still defines the field: the model interleaves reasoning traces with actions and folds the resulting observations back into its next thought.
That loop escaped the laboratory almost immediately. Within months, AutoGPT-style systems chained it into open-ended autonomy (Yang et al. 2023b), and Generative Agents showed that populations of such loops, given memory and reflection, produce believable social behavior (Park et al. 2023a). Three years later, descendants of the same recipe resolve real GitHub issues, operate browsers and desktop interfaces on a user’s behalf, and design and run scientific experiments (Section 8, Section 10). Figure 2 traces this trajectory.
The literature that documents this progress has grown faster than its organization. Research on LLM-based agents is split across subcommunities, planning and reasoning, memory, tool learning, multi-agent systems, web and GUI automation, embodied control, agent safety, each with its own benchmarks, vocabulary, and citation graph, and with limited contact between them. The foundational general surveys unified the field as it stood in 2023 (Wang et al. 2023b; Xi et al. 2023), but predate three developments that have since reshaped it: reasoning-native models trained to deliberate, computer-use agents that act through raw screen interfaces, and the systematic study of agent security. More recent general surveys organize the field by methodology or framework lineage (Luo et al. 2025; Masterman et al. 2024; Plaat et al. 2025), by a brain-inspired decomposition into cognitive modules (Liu et al. 2025), or trace a single thread such as the road from chain-of-thought prompting to language agents (Zhang et al. 2023c). Across all of them, evaluation methodology and safety appear as closing discussions rather than as objects of analysis in their own right.
The scale of the problem is easy to understate. Figure 1 plots monthly arXiv submissions across five agent subtopics: of the roughly nine thousand papers our keyword harvest assigns to the field, fewer than one in three hundred predate 2023, yet the field now adds on the order of a thousand papers per month and its cumulative volume roughly tripled in the year to mid-2026. That growth is spread across planning, multi-agent, tool-use, computer-use, and safety subcommunities that expand in parallel and largely in isolation, exactly the fragmentation a function-first organization is meant to make navigable.
This survey organizes the field by architectural function instead. Frameworks bundle components and are superseded; the functions those components serve persist. We therefore structure the literature along three axes (Figure 4, developed in Section 3): four component pillars that describe what an agent is made of, planning (Section 4), memory (Section 5), tool use (Section 6), and multi-agent coordination (Section 7); two context dimensions that describe where it acts and where it is deployed, interactive environments (Section 8) and application domains (Section 10); and two cross-cutting concerns that apply to every cell of that matrix, evaluation (Section 9) and safety (Section 11). The function-first organization is what makes the fragmented subcommunities commensurable: a memory paper and a web-agent paper rarely cite each other, but both take a position on the same small set of design questions, and the taxonomy makes those positions comparable. Read through that lens, the field tells a single story: its components are consolidating as prompting gives way to trained reasoning and scaffolds migrate into the models; its frontier has moved to grounding agents in real environments and evaluating them honestly; and safety is the binding constraint on both. This is the thesis the rest of the survey defends.
Our contributions are as follows:
  • A function-first taxonomy of LLM-based agents (four component pillars, two context dimensions, two cross-cutting concerns) that assigns every surveyed work a position independent of the framework that introduced it (Figure 4).
  • Current coverage. We synthesize the literature from the first trained browsing agents (2021) through early 2026, including the post-2023 developments absent from the foundational surveys: reasoning-native models, computer-use agents, AI-scientist systems, and agent security.
  • Evaluation as a first-class pillar. Beyond cataloguing benchmarks, Section 9 analyzes the validity of current evaluation practice, including cost-blind leaderboards, contamination, and outcome-versus-process metrics.
  • Safety and security as a first-class pillar. Section 11 develops a unified agentic threat model spanning prompt injection, memory poisoning, and backdoors, through alignment-level risks and autonomy governance.
  • Original controlled experiments. Rather than argue the survey’s two central trade-offs only from the literature, we measure them: a compute-matched comparison of agent control loops on ALFWorld (Section 4.3) and a capability-versus-injection-robustness measurement across model scale and family (Section 11.3).
Throughout, where the literature contains credible negative results, on self-correction (Section 4.4), on multi-agent debate (Section 7.2), and on agent scaffolds versus simple pipelines (Section 10.1), we present them with the same weight as the positive claims.
A note on method. The corpus behind this survey was assembled through ten subtopic sweeps (foundational surveys and position papers; single-agent architectures; planning and reasoning; memory; tool use; multi-agent systems; embodied, web, and GUI agents; benchmarks and evaluation methodology; safety and alignment; applied agents), yielding a core of 184 references, each verified against its primary source; further works were added under the same verification rule, for 228 references in total. We include a system if it places an LLM in an action–perception loop, the model’s outputs are executed against an environment, a tool interface, or another agent, and the results are fed back, and we exclude chat-only uses of LLMs, base-model training techniques except where they target agentic behavior, and vision–language–action robotics beyond what Section 8.3 requires.
Figure 2. A timeline of representative milestones in LLM-based agents, from the first trained browsing agent (2021) to the reasoning-native, computer-use, and AI-scientist systems of 2025; the rightmost tick marks the present frontier of agent safety and governance (Section 11, Section 12) rather than a single dated system. Each system is introduced and cited in the section indicated by the taxonomy of Figure 4; labels alternate above and below the spine only for legibility.
Figure 2. A timeline of representative milestones in LLM-based agents, from the first trained browsing agent (2021) to the reasoning-native, computer-use, and AI-scientist systems of 2025; the rightmost tick marks the present frontier of agent safety and governance (Section 11, Section 12) rather than a single dated system. Each system is introduced and cited in the section indicated by the taxonomy of Figure 4; labels alternate above and below the spine only for legibility.
Preprints 223549 g002
Table 1 positions this survey against the closest general surveys. Single-pillar surveys, planning (Huang et al. 2024), memory (Zhang et al. 2024f), tool use (Qu et al. 2024), GUI agents (Nguyen et al. 2024), evaluation (Yehudai et al. 2025), and trustworthiness (Yu et al. 2025), go deeper than we do on their single column; we cite them as deeper dives in the corresponding sections and direct our own coverage at the connective tissue between pillars, which is precisely what a single-pillar survey cannot supply.
The remainder of the survey follows the taxonomy. Section 2 sets up definitions and the cognitive-architecture vocabulary; Section 3 presents the taxonomy itself. Section 4, Section 5, Section 6 and Section 7 survey the four component pillars, Section 8 grounds them in web, GUI, and embodied environments, Section 9 reviews evaluation practice and its blind spots, Section 10 surveys deployed application domains, and Section 11 analyzes safety and security. Section 12 distills seven open challenges, and Section 13 concludes. Readers interested in a single pillar can read Section 2Section 3 and jump directly to the relevant section; each pillar section is self-contained and closes with a key takeaway.

2. Background: From Language Models to Agents

This section establishes the minimal background the rest of the survey builds on: the base-model capabilities agents inherit (Section 2.1), a working definition of an LLM-based agent (Section 2.2), and the cognitive-architecture vocabulary we adopt throughout (Section 2.3).

2.1. Language Models as Reasoning Cores

Everything an agent can do is bounded by what its underlying model can do, so we begin with the base-model capabilities that agents inherit. Language modeling passed through several regimes, from count-based n-gram models, through distributed word representations and recurrent networks, to the pretrained Transformer that made transfer learning the default, but the capability that turned language models into agent controllers is a property of scale rather than of any single architecture. At sufficient scale, models acquire in-context learning: they infer a task from a few examples in the prompt, with no gradient updates. Instruction tuning then made this latent competence addressable through ordinary natural-language commands, and reinforcement learning from human feedback aligned model outputs with those commands; the same post-training machinery is what later transfers these capabilities across languages and domains (Lee et al. 2025).
The ability most directly responsible for agency is multi-step reasoning. Chain-of-thought prompting showed that eliciting intermediate reasoning steps sharply improves performance on arithmetic, commonsense, and symbolic tasks, and that this benefit emerges only past a scale threshold (Wei et al. 2022); the same behavior can be triggered zero-shot by a single instruction to reason step by step (Kojima et al. 2022). Because reasoning is latent and promptable, a model can be asked to deliberate before it acts, the precondition for treating the model as a decision-maker rather than a text generator. Early surveys mapped the space of reasoning-elicitation methods and their evaluation (Huang and Chang 2022), a space we revisit as agent planning in Section 4.
A further inflection, unfolding across 2024–2025, moved reasoning from a prompting trick into the model’s weights. Bootstrapping a model on its own correct rationales (Zelikman et al. 2022), rewarding individual reasoning steps rather than only final answers (Lightman et al. 2023), and, most strikingly, incentivizing reasoning through large-scale reinforcement learning alone (DeepSeek-AI et al. 2025) produced reasoning-native models that deliberate, verify, and backtrack without being prompted to. Frontier model reports now advertise agentic tool use and computer operation as headline capabilities (Comanici et al. 2025; Kimi Team 2025). These models unsettle a long-standing assumption of agent design (that planning must be supplied by an external scaffold) a tension we take up in Section 4.5.

2.2. What Is an LLM-Based Agent?

Preprints 223549 i001
Definition 1 names four moving parts, each a pillar of this survey. The loop is closed by the model’s decision procedure (how it turns observations and internal state into the next action) which we survey as planning and reasoning (Section 4). The internal state that survives across steps is memory (Section 5). The external actions are, in the general case, tool calls (Section 6); when the recipient of an action is another agent, the action is a message, and the loop becomes a multi-agent system (Section 7). Figure 3 draws this anatomy; later sections refer to its blocks by name, and Section 11 re-annotates the same diagram to locate where attacks enter.
The perceive–decide–act loop is not new: it is the organizing abstraction of the belief–desire–intention (BDI) tradition and of reinforcement learning, both of which long predate LLMs. What the language model contributes is a general prior over world knowledge and, crucially, a natural-language action interface. A classical agent must be given a hand-engineered state representation and, for RL, a task-specific reward; an LLM-based agent instead reads instructions, tool documentation, and observations as text and emits actions as text, so a single model generalizes across tasks it was never explicitly trained for. This is the capability that the earliest LLM-agent systems exploited, WebGPT learned to answer questions by operating a text browser under human feedback (Nakano et al. 2021); MRKL routed queries between the model and external symbolic modules (Karpas et al. 2022); and ReAct showed that merely interleaving reasoning with actions in the prompt yields a competent agent with no additional training (Yao et al. 2022b).
Terminology in this area is unsettled. We use agent for any system matching Definition 1; some authors reserve agentic AI2 for higher-autonomy systems and contrast it with simpler tool-augmented models (Sapkota et al. 2025). Autonomy is best read as a spectrum rather than a binary, with proposed frameworks distinguishing several levels by how much of the loop the human retains (Feng et al. 2025b); whether systems at the top of that spectrum should be built at all is actively contested (Mitchell et al. 2025). We defer the governance of autonomy to Section 11.5 and use the levels here only as descriptive vocabulary.

2.3. A Cognitive-Architecture View

To describe agents precisely we adopt the vocabulary of Cognitive Architectures for Language Agents (CoALA) (Sumers et al. 2023), which recasts a diverse body of systems (ReAct, Reflexion, Voyager, and others surveyed below) in terms borrowed from classical cognitive architecture. CoALA organizes an agent along three axes. Its memory is divided into a short-lived working memory and long-term stores that, following human memory, are episodic (past experience), semantic (facts about the world), and procedural (skills and the agent’s own code). Its action space is split into internal actions that read from or write to those memories (reasoning, retrieval, learning) and external actions that ground the agent in an environment through tools or physical actuators. Its decision procedure is the loop that proposes, evaluates, and selects the next action.
These three axes are exactly the seams along which this survey cuts the literature. Planning (Section 4) is the decision procedure; memory (Section 5) is the store hierarchy; tool use (Section 6) is the external action space; and multi-agent coordination (Section 7) generalizes the action space to messages exchanged with other agents. Adopting a component vocabulary, rather than a per-framework one, is what lets us compare systems that were never designed to be compared: a framework such as AutoGen or MetaGPT bundles a particular choice along each axis, but the axes themselves are stable, and it is at the level of the axis that the design questions recur. Section 3 builds the survey’s taxonomy on this observation.

3. Taxonomy of LLM-Based Agents

The preceding section described one agent; this section organizes the thousands of them that the literature has produced. We classify the field along three kinds of axis. Component axes ask what an agent is made of, and following the cognitive-architecture view of Section 2.3 there are four: planning and reasoning (Section 4), memory (Section 5), tool use (Section 6), and multi-agent coordination (Section 7). Context axes ask where the agent acts: the interactive environment it is embedded in (Section 8) and the application domain it is deployed to (Section 10). Cross-cutting axes ask how we should judge it, and apply to every cell of the component-by-context matrix: evaluation methodology (Section 9) and safety (Section 11). Figure 4 lays out the full taxonomy; each leaf names a subsection of this survey together with representative systems, all cited in full where they appear.
The component pillars are the load-bearing distinction, because each is a mature subfield with its own dedicated survey that goes deeper than we can here: planning (Cao et al. 2025; Huang et al. 2024), memory (Wu et al. 2025; Zhang et al. 2024f), tool use (Qu et al. 2024; Wang et al. 2024g), and multi-agent systems (Guo et al. 2024a; Tran et al. 2025; Yan et al. 2025). We cite these as deeper dives at the head of each pillar and aim our own contribution at the connections between pillars, how a memory design constrains a planning strategy, how a tool interface reshapes an action space, which no single-pillar survey covers. The context axes are served by their own specialized surveys as well: web (Ning et al. 2025) and GUI agents (Nguyen et al. 2024; Zhang et al. 2024a) for environments, and domain surveys for software (Liu et al. 2024a), science (Ren et al. 2025), medicine (Wang et al. 2025c), and finance (Ding et al. 2024) for applications. The two cross-cutting concerns have recently acquired dedicated surveys of their own, evaluation (Mohammadi et al. 2025; Yehudai et al. 2025) and trustworthiness (Su et al. 2025; Yu et al. 2025), a sign that the field now treats them as objects of study rather than afterthoughts, and a stance this survey adopts by giving each a full section.
Why organize by function at all, rather than by framework, which is how practitioners often navigate the field? Because frameworks are bundles and functions are atoms. A system such as AutoGen (Wu et al. 2023a) or MetaGPT (Hong et al. 2024) fixes a particular choice on every axis at once (a role structure, a memory scheme, a tool protocol, a planning loop) and is then superseded by the next bundle. The design questions, however, recur unchanged across bundles: whether to plan ahead or interleave planning with acting, whether memory should be flat text or a structured graph, whether tools should be selected or authored, whether agents should cooperate or debate. Classifying by function makes these questions, and the competing answers to them, the unit of comparison, so that a lesson learned in one framework transfers to the next. Table 2 makes the payoff concrete: it places representative systems against every axis at once, so that designs the taxonomy tree of Figure 4 necessarily scatters into separate leaves can be read side by side. The remainder of the survey works through the taxonomy one axis at a time.
Table 2. Representative agent systems placed against every axis of the taxonomy at once. “Coord.” is the multi-agent topology (“–” for single-agent); “Planning” names the decision procedure of Section 4. Every entry is documented in the section listed; this is the comparison the per-leaf tree of Figure 4 cannot show.
Table 2. Representative agent systems placed against every axis of the taxonomy at once. “Coord.” is the multi-agent topology (“–” for single-agent); “Planning” names the decision procedure of Section 4. Every entry is documented in the section listed; this is the comparison the per-leaf tree of Figure 4 cannot show.
System Planning Memory Action Coord. Domain / env. § Ref.
ReAct interleaved – tools – general Section 4.3 (Yao et al. 2022b)
Reflexion reflection episodic tools – general Section 4.4 (Shinn et al. 2023)
AutoGPT plan-execute scratchpad tools – general Section 1 (Yang et al. 2023b)
HuggingGPT plan-execute – models – multimodal Section 6.2 (Shen et al. 2023)
ToolLLM DFS search – APIs – tool use Section 6.1 (Qin et al. 2023)
CodeAct interleaved – code – general Section 6.3 (Wang et al. 2024d)
SWE-agent interleaved – code (ACI) – software Section 6.3 (Yang et al. 2024a)
MemGPT – tiered tool calls – dialogue Section 5.2 (Packer et al. 2023)
Voyager curriculum skill library code – embodied Section 5.4 (Wang et al. 2023a)
CAMEL role-play dialogue messages dyad general Section 7.1 (Li et al. 2023a)
AutoGen configurable per-agent tools, code star general Section 7.1 (Wu et al. 2023a)
MetaGPT SOP pipeline shared docs structured pipeline software Section 7.1 (Hong et al. 2024)
Gen. Agents reflection memory stream messages society simulation Section 5.3, Section 7.4 (Park et al. 2023a)
WebVoyager interleaved – GUI (pixels) – web Section 8.1 (He et al. 2024)
UI-TARS trained e2e – GUI actions – computer use Section 8.2 (Qin et al. 2025)
OS-Copilot plan + improve skill store OS actions – OS Section 8.2 (Wu et al. 2024b)
ChemCrow plan-execute – domain tools – chemistry Section 10.2 (Bran et al. 2023)
AI Scientist plan-execute – code, tools – science Section 10.2 (Lu et al. 2024)
Figure 4. Master taxonomy of LLM-agent research as organized by this survey. Four component pillars (Planning & Reasoning Section 4, Memory Section 5, Tool Use Section 6, Multi-Agent Systems Section 7) are grounded in two context dimensions (Environments Section 8, Applications Section 10), with Evaluation (Section 9) and Safety (Section 11) as cross-cutting concerns (dashed edges). Leaves list representative works; all are cited in full in the corresponding sections.
Figure 4. Master taxonomy of LLM-agent research as organized by this survey. Four component pillars (Planning & Reasoning Section 4, Memory Section 5, Tool Use Section 6, Multi-Agent Systems Section 7) are grounded in two context dimensions (Environments Section 8, Applications Section 10), with Evaluation (Section 9) and Safety (Section 11) as cross-cutting concerns (dashed edges). Leaves list representative works; all are cited in full in the corresponding sections.
Preprints 223549 g004

4. Planning and Reasoning

Planning is the agent’s decision procedure, the policy that turns the current observation and internal state into the next action. It is the axis on which the field has moved fastest, and it is instructive to organize the work not by chronology but by where the improvement comes from. Four sources recur: better prompting of a fixed model (Section 4.1), explicit search over candidate reasoning paths (Section 4.2), decomposition and interleaving of planning with acting (Section 4.3), and feedback in the form of self-correction (Section 4.4). A fifth source, training the model to reason directly (Section 4.5), has recently begun to absorb the other four, and we close the section by asking what that means for agent design (Cao et al. 2025; Huang et al. 2024).

4.1. Prompted Reasoning

The cheapest way to make a model plan is to ask it to. Chain-of-thought prompting elicits intermediate reasoning steps that markedly improve multi-step problem solving (Wei et al. 2022), and the effect survives the removal of exemplars: a single instruction to reason step by step recovers much of the gain zero-shot (Kojima et al. 2022). These results reframed the model as something that can deliberate, and every planning method below is, in some sense, a more disciplined way of spending that deliberation.
The discipline takes three broad forms. Sampling many reasoning paths and taking a majority vote over their answers (self-consistency) trades inference compute for accuracy and reliably beats greedy decoding (Wang et al. 2022). Decomposing a problem into an ordered list of easier subproblems, each conditioned on the last, improves compositional generalization (Zhou et al. 2022). And splitting the work into an explicit plan-then-execute pair reduces the calculation and missing-step errors of naive zero-shot reasoning (Wang et al. 2023c). A theme already visible here, and central to Section 9, is that these gains are bought with extra inference: self-consistency’s accuracy scales with the number of sampled paths, so any comparison that ignores compute overstates the method’s value.

4.2. Search over Thoughts

If prompted reasoning explores a single path, the natural generalization is to search over many. Tree of Thoughts makes the search explicit, maintaining a tree of partial solutions that the model expands, evaluates, and backtracks over using breadth- or depth-first search (Yao et al. 2023). Graph of Thoughts relaxes the tree to an arbitrary graph so that reasoning paths can be merged and refined, not just branched (Besta et al. 2023). The move from chain to tree to graph (Figure 5) buys quality on hard problems at a steep and often unreported compute cost.3
A second line makes the search model-based. Reasoning via Planning repurposes the LLM as its own world model4, predicting the next state so that Monte Carlo Tree Search can plan against it (Hao et al. 2023a). Language Agent Tree Search extends this to grounded settings, unifying reasoning, acting, and reflection in a single search over an environment (Zhou et al. 2024a). A third, meta-level line has the model search over reasoning strategies rather than reasoning steps: Self-Discover composes task-specific reasoning structures from atomic modules (Zhou et al. 2024b), while Buffer of Thoughts retrieves distilled “thought-templates” from past problems (Yang et al. 2024b), a design that already blurs into the memory mechanisms of Section 5.

4.3. Task Decomposition and Interactive Planning

Search assumes the agent can evaluate a plan before executing it. In interactive settings it often cannot, because the environment reveals information only when acted upon, which raises the defining question of this subsection: should the agent commit to a full plan up front, or interleave planning with acting? ReAct takes the interleaved extreme, alternating a reasoning trace with an action and folding each observation back into the next thought (Yao et al. 2022b). ReWOO takes the plan-ahead extreme, producing a complete tool-use plan before any execution; by decoupling reasoning from observation it cuts token consumption roughly fivefold relative to interleaved loops, at the cost of being unable to adapt when an early step surprises it (Xu et al. 2023).
The productive answer is that decomposition granularity should be conditioned on the agent’s own competence rather than fixed in advance. ADaPT recurses into finer subtasks only when the executor fails to complete a step directly, matching plan depth to difficulty (Prasad et al. 2023). DEPS closes the loop differently, having the agent describe its progress, explain failures, and revise the plan, which lets a single zero-shot agent complete dozens of long-horizon tasks in an open world (Wang et al. 2023f). Both point the same way: the plan is not an artifact produced once but a hypothesis revised against feedback.
The trade-off this subsection describes is usually argued from the literature; we can also measure it. Figure 6 reports a small controlled comparison we ran of the same model (Qwen2.5-7B-Instruct) under three control loops on 50 ALFWorld tasks. Interleaving (ReAct) roughly doubles the success rate of plan-then-execute (ReWOO), 46% versus 22%, but costs about four times as many tokens; adding a single reflection retry (Reflexion, Section 4.4) buys a further 8 points to 54% at another 70% more tokens. The plan-ahead loop is genuinely cheaper (it just spends its savings on failures it cannot see coming) and the result lands squarely on a cost–accuracy frontier, foreshadowing the argument of Section 9.3 that neither axis should be read alone.

4.4. Reflection and Self-Correction

Revising against feedback invites an appealing idea: let the agent generate that feedback itself. Reflexion has the agent write a verbal self-reflection after a failed attempt and store it in episodic memory, improving decision-making on later attempts without any weight update (Shinn et al. 2023). Self-Refine applies the same loop within a single problem, alternating generation and self-critique (Madaan et al. 2023), and CRITIC grounds the critique in external tools such as search engines and interpreters (Gou et al. 2023). Reported gains across these systems made self-correction a near-default component of agent scaffolds.
The evidence for intrinsic self-correction, however, is weaker than the enthusiasm suggests. Controlled experiments find that when a model is asked to correct its reasoning using only its own judgment, with no external signal, performance frequently fails to improve and sometimes degrades, because the model cannot reliably tell which of its answers is wrong (Huang et al. 2023a). A critical survey of the subfield goes further, arguing that many positive results rest on ill-posed setups, oracle stopping criteria, unfair baselines, or tasks where the correction step smuggles in information the first attempt lacked (Kamoi et al. 2024). The two literatures are not actually in contradiction once the confound is named.
The reconciliation is that self-correction works when the feedback is not purely intrinsic. CRITIC’s tool grounding and Reflexion’s environment signal both supply an external check, which is why they help where pure self-critique does not. Where such a signal is unavailable, the alternative is to train the ability in: SCoRe uses multi-turn reinforcement learning on the model’s own correction traces to produce genuine self-correction, overcoming the distribution-shift and collapse failures of prompting-only approaches (Kumar et al. 2024). Self-correction, in short, is real but not free, it requires either an external verifier or a training signal, a conclusion that anticipates Section 4.5 and the section’s key takeaway.
Table 3. Planning strategies by the structure they search and the feedback they use. “Cost” is inference overhead relative to a single chain-of-thought pass (qualitative; exact factors depend on branching and depth). “Env?” marks whether the method requires interacting with an external environment or tools, as opposed to pure internal reasoning.
Table 3. Planning strategies by the structure they search and the feedback they use. “Cost” is inference overhead relative to a single chain-of-thought pass (qualitative; exact factors depend on branching and depth). “Env?” marks whether the method requires interacting with an external environment or tools, as opposed to pure internal reasoning.
Strategy Search structure Feedback source Cost Env? Ref.
Chain-of-Thought linear chain none base ✗ (Wei et al. 2022)
Self-Consistency parallel samples answer vote N × ✗ (Wang et al. 2022)
Tree of Thoughts tree self-evaluation high ✗ (Yao et al. 2023)
Graph of Thoughts graph (merges) self-evaluation high ✗ (Besta et al. 2023)
RAP tree (MCTS) world-model value high ✗ (Hao et al. 2023a)
LATS tree (MCTS) env. + reflection high ✓ (Zhou et al. 2024a)
ReWOO plan-then-act tool observations low ✓ (Xu et al. 2023)
Reflexion linear + memory env. reward medium ✓ (Shinn et al. 2023)

4.5. Learned Reasoning and Trained Agents

The methods above all treat the model as fixed and wrap machinery around it. The most recent shift instead trains the model to reason. STaR bootstraps a model on the rationales that lead to correct answers, iteratively improving its reasoning from a small seed (Zelikman et al. 2022). Process supervision rewards each intermediate step rather than only the final answer, and outperforms outcome supervision on hard mathematical reasoning (Lightman et al. 2023). DeepSeek-R1 pushes the idea to its limit, showing that large-scale reinforcement learning with only an outcome signal can induce reasoning-native behavior (self-verification, backtracking, and long deliberation) as an emergent property, with no supervised reasoning traces at all (DeepSeek-AI et al. 2025).
Training has since moved from teaching a model to reason to teaching it to act, and this is the concrete mechanism behind the survey’s recurring observation that scaffolds are migrating into the models. A first wave fine-tunes on agent trajectories: AgentTuning mixes interaction traces into instruction tuning to give general models agent abilities (Zeng et al. 2023), and FireAct fine-tunes on ReAct-style trajectories drawn from several tasks (Chen et al. 2023a). A second wave trains agents online, in the environment itself: WebRL builds a self-evolving curriculum with reinforcement learning for web agents (Qi et al. 2024), and AgentGym evolves agents across many interactive environments at once (Xi et al. 2024). A third wave moves the tools inside the reasoning chain itself: Search-R1 trains a model to interleave search-engine calls with its own thoughts from outcome rewards alone (Jin et al. 2025), and ReTool learns when to invoke a code interpreter mid-derivation (Feng et al. 2025a); this agentic reinforcement learning is growing fast enough to have a survey of its own (Zhang et al. 2025). The same move recurs in the capability-specific models surveyed later, function-calling models trained on synthetic call data (Section 6.4), reinforcement-learned GUI agents such as UI-TARS (Section 8.2), and RL-trained self-correction (SCoRe, Section 4.4). Table 4 organizes these by the training signal they use and by whether that signal is drawn from a live environment, which is the axis along which prompted scaffolds are being absorbed into trained behavior.
This raises a genuinely open question for agent design, and we flag it as open rather than resolve it. If a model backtracks and verifies internally, do the external scaffolds of Section 4.2, Section 4.3 and Section 4.4 (the search trees, the reflection loops) remain necessary, or does planning migrate into the weights and leave the scaffold redundant on the tasks the model has internalized? Early evidence is mixed, and the pillar surveys of learned planning treat it as unsettled (Cao et al. 2025; Huang et al. 2024); we return to it as a future direction in Section 12.
Preprints 223549 i002

5. Memory

If planning is what an agent does within a step, memory is what carries across steps. A language model is, by default, stateless: it retains nothing between calls beyond what fits in its context window. Memory is the machinery that lifts an agent out of that amnesia, it is what lets an agent learn a user’s preferences over a month, reuse a skill it acquired last week, or avoid a mistake it made an hour ago, and it is, in the strong sense, the component that makes an agent more than a prompt.

5.1. A Taxonomy of Agent Memory

Following the cognitive-architecture view of Section 2.3, agent memory divides first by time horizon and then by content. By horizon there is working memory, the information immediately available in the context window during a single decision, and long-term memory, which persists across episodes. Long-term memory, in turn, borrows the human tripartite division (Sumers et al. 2023; Wu et al. 2025): episodic memory of past experiences and interactions, semantic memory of facts about the world and the user, and procedural memory of skills and routines the agent can execute. The pillar survey of Zhang et al. (2024f) organizes the subfield along compatible lines and provides a fuller catalogue than we give here.
This four-way split is descriptive; what makes memory systems differ in practice is a small set of design decisions, and it is these that structure the rest of the section (Figure 7). Every memory system must choose a store form (flat text, vectors, key–value pairs, or a graph), a read policy (how relevant memories are retrieved for the current step), a write policy (what gets committed to long-term storage and in what form), and a forgetting policy (how stale or low-value memories are decayed, if at all). The recurring finding of the subfield is that these policies, not raw capacity, determine whether memory helps.

5.2. Working Memory and Context Management

The first problem memory poses is that working memory is bounded: the context window is finite, yet tasks are not. MemGPT reframes this as an operating-systems problem, giving the agent a tiered memory hierarchy and letting it page information between a limited main context and external storage through self-issued function calls, much as an OS moves data between RAM and disk (Packer et al. 2023). MemWalker takes a complementary route for long documents, building a tree of recursively summarized nodes and having the agent navigate it top-down at query time, reading only the branches it needs (Chen et al. 2023b). Both make the same conceptual move: context management is not a fixed preprocessing step but an action the agent takes, subject to the same decision-making the rest of the loop uses.

5.3. Long-Term and Structured External Memory

Long-term memory systems are best compared by their store form, because form dictates what the read and write policies can be. The simplest form is flat text with a lifecycle. The canonical design is the memory stream of Generative Agents (Park et al. 2023a): a chronological log of natural-language observations retrieved by a score combining recency, importance, and relevance, with a periodic reflection step that synthesizes low-level memories into higher-level ones, a template that most later text-memory systems vary. MemoryBank stores past interactions as text and adds a psychologically motivated forgetting rule, strengthening or decaying each memory by an Ebbinghaus curve so that importance and recency govern what survives (Zhong et al. 2023). Think-in-Memory refines what is stored, keeping the agent’s evolved thoughts rather than raw dialogue and updating them with insert, forget, and merge operations, which avoids the inconsistent re-reasoning that raw-text recall invites (Liu et al. 2023a).
Structuring the store buys sharper retrieval. RET-LLM records knowledge as subject–relation–object triplets, giving an updatable and interpretable memory (Modarressi et al. 2023); HippoRAG goes further, building an open knowledge graph and retrieving over it with Personalized PageRank in a deliberate analogy to hippocampal indexing, which enables single-step multi-hop recall that flat retrieval handles poorly (Gutiérrez et al. 2024). That store form matters is not merely asserted: a controlled study varying chunk-, triple-, fact-, and summary-based structures over identical queries finds the best structure depends on the query type, offering design guidance rather than a single winner (Zeng et al. 2024).
The newest systems make the store self-organizing or production-hardened. A-MEM links each new memory to related ones in a Zettelkasten-style network and lets new entries revise the attributes of old ones, so the memory continually restructures itself (Xu et al. 2025). Mem0 targets deployment, dynamically extracting and consolidating salient facts across sessions and reporting, on the LoCoMo long-conversation benchmark, large quality gains over a strong memory baseline, together with roughly ninety-percent reductions in latency and token cost relative to full-context prompting, a reminder that in production the binding constraint on memory is as much cost as accuracy (Chhikara et al. 2025). MIRIX pushes toward the frontier of multiple coordinated memory types, including multimodal experience, managed by dedicated per-type managers (Wang and Chen 2025). A separate, architecture-level line embeds memory in the model itself rather than in a scaffold: LongMem attaches a trainable side-network that caches and retrieves past context (Wang et al. 2023d), and Larimar equips an LLM with an episodic memory supporting fast one-shot updates and selective forgetting without retraining (Das et al. 2024).
Because these systems make competing claims, they need a common yardstick. LoCoMo supplies one, evaluating very long-term conversational memory with multi-hop and temporal questions that expose the weaknesses of both long-context and naive retrieval baselines (Maharana et al. 2024); it is now the standard benchmark by which the systems above report, with LongMemEval complementing it by decomposing long-term interactive memory into five separately testable abilities (Wu et al. 2024a). We return to memory evaluation in Section 9. Table 5 summarizes the design choices across representative systems.

5.4. Experiential Learning and Skill Libraries

Procedural memory closes the loop between memory and self-improvement. When an agent stores not facts but skills, memory becomes the substrate on which it gets better at its job without any weight update. Voyager makes this concrete in an open-ended embodied setting, maintaining a library of executable code skills that it writes, verifies, and later retrieves and composes, so that competence accumulates over a lifetime of play (Wang et al. 2023a). Synapse stores whole past trajectories as exemplars and retrieves them by task similarity to guide new episodes (Zheng et al. 2023).
Storing raw experience, however, scales poorly; the more useful object is often a distilled lesson. ExpeL has an agent gather experiences across training tasks and extract natural-language insights from both its successes and its failures, carrying forward compact rules rather than full episodes and improving purely through in-context use of that memory (Zhao et al. 2023). ReasoningBank scales the same idea toward deployment, distilling strategy-level memory from both successful and failed trajectories so that an agent self-evolves over a continuous task stream (Ouyang et al. 2025). This is continual learning without gradients, and it is one of the clearest bridges from memory to the open problem of lifelong agent improvement we take up in Section 12.
Preprints 223549 i003

6. Tool Use and Action Execution

Memory lets an agent look inward across time; tools let it reach outward into the world. A tool call is the external action of Definition 1, the means by which an agent computes what its weights cannot, retrieves what they do not contain, and effects change beyond generating text. Tool use is therefore the pillar that turns a reasoner into an actor, and this section follows its arc from prompting a model to call fixed tools, through learning to call thousands of them, to letting the agent write executable code as its action space, and finally to training models specialized for the task.
Preprints 223549 i004

6.1. Foundations and Paradigms

The literature splits along one axis: whether tool competence is prompted into a fixed model or learned into its weights. The learned branch came first at scale. TALM bootstrapped tool use from a few demonstrations through iterative self-play (Parisi et al. 2022), and Toolformer made the idea self-supervised, letting a model annotate ordinary text with the API calls that would have helped predict it, then training on the calls that lowered loss (Schick et al. 2023). Scaling the tool inventory exposed a retrieval problem: Gorilla pairs a fine-tuned model with a document retriever so it can emit correct calls across thousands of constantly changing APIs without hallucinating arguments (Patil et al. 2024), and ToolLLM extends this to over sixteen thousand real-world APIs, using a depth-first decision-tree search to plan and recover across multi-call traces (Qin et al. 2023).
The prompted branch asks how far a fixed model can be pushed with the right context, and the surprising answer is: quite far. Tool documentation alone, with no demonstrations, is often sufficient for zero-shot tool use and scales better than few-shot prompting when hundreds of tools are available (Hsieh et al. 2023). The two branches trade off predictably, prompted tool use scales with the size of the tool catalogue that fits in context, learned tool use scales with the volume of training data, a tension the pillar survey of Qu et al. (2024) maps in full and that the function-calling models of Section 6.4 attempt to resolve.

6.2. Tool Selection, Creation, and Composition

Calling one tool is selection; calling several in concert is composition, and it is where agents begin to resemble programs. Chameleon has an LLM synthesize a program that composes vision models, search, and Python functions to solve multimodal tasks (Lu et al. 2023); HuggingGPT generalizes the pattern by treating other models as the tools, using the LLM as a controller that plans a task and dispatches specialized models to execute its stages (Shen et al. 2023); and RestGPT drives real, stateful REST APIs through a coarse-to-fine planner and a dedicated executor that formats parameters and parses responses (Song et al. 2023).
More striking than composing given tools is authoring new ones. LATM closes a loop in which a capable model acts as a tool maker, writing reusable Python functions that a cheaper model then applies, amortizing the authoring cost across many uses (Cai et al. 2023). ToolkenGPT attacks the scaling limit from the other side, representing each tool as a learned vocabulary token (a “toolken”) so that a frozen model can select from a large, expandable tool set without listing every tool in its prompt (Hao et al. 2023b). Together these mark a shift in what the tool set is: no longer a fixed external inventory but a learned, self-extending object the agent grows.

6.3. Code as a Universal Action Space

Taken to its conclusion, the composition-and-authoring trend collapses the distinction between tool use and programming. CodeAct consolidates an agent’s entire action space into executable Python: instead of emitting a fixed JSON action schema, the agent writes code, which subsumes tool calls, control flow, and data manipulation in one interface and lets the agent revise from execution errors by debugging its own program (Wang et al. 2024d). Code as action is expressive precisely because it is Turing-complete, but realizing that expressiveness depends on the surrounding interface. SWE-agent makes this explicit with the notion of an agent–computer interface5: a deliberately designed set of commands and observations for browsing, editing, and running code, whose design changes task success as much as swapping the underlying model does (Yang et al. 2024a). We treat the ACI concept here and defer its software-engineering payoff to Section 10.1.

6.4. Function-Calling Models and Training Data

The prompted-versus-learned tension resolves, in practice, into a class of models trained specifically to call functions. xLAM is a family of “large action models” trained through a unified data-synthesis pipeline to strengthen tool use, topping function-calling leaderboards (Zhang et al. 2024c); ToolACE reaches comparable accuracy with an 8B model by generating a diverse, rule- and model-verified corpus of function-calling dialogues (Liu et al. 2024b). Both sit in the trained-agent lineage organized in Table 4. Both are also downstream of a data problem, and the open-model lineage that precedes them is a history of solving it: GPT4Tools self-instructs a tool-use dataset from a teacher model (Yang et al. 2023d), ToolAlpaca synthesizes several thousand tool-use cases from a multi-agent simulation to fine-tune compact models (Tang et al. 2023), and a careful study shows small models are weak tool learners unless the task is decomposed across specialized planner, caller, and summarizer sub-models (Shen et al. 2024). Progress on function calling has been, more than anything, progress on generating verified training data.
These benchmarks, API-Bank’s runnable multi-turn dialogues (Li et al. 2023b), StableToolBench’s virtual API server that stabilizes the volatile ToolBench APIs (Guo et al. 2024b), the Berkeley Function-Calling Leaderboard’s abstract-syntax-tree matching (Patil et al. 2025), and GTA’s real-tool tasks with implicit intent on which leading models finish under half (Wang et al. 2024a), are summarized in Table 6 and revisited in Section 9.
The final step of this arc is standardization. As tool use matured, the per-model function-calling conventions began to converge on open protocols that decouple an agent from the tools it calls: the Model Context Protocol exposes tools and data to any client through a common interface (Anthropic 2024b), and the Agent2Agent protocol standardizes how agents discover and delegate to one another (Google 2025). These protocols make tools and agents composable across vendors, but by turning the tool interface into shared infrastructure they also enlarge the supply-chain attack surface of Section 11 and raise the open infrastructure questions of Section 12. The protocol layer is young but already an object of study: a survey of agent protocols maps the emerging standards and their evaluation dimensions (Yang et al. 2025), and a security analysis of the MCP ecosystem catalogues threats across the server lifecycle (Hou et al. 2025).
Preprints 223549 i005

7. Multi-Agent Systems

The three preceding pillars describe a single agent. When the recipient of an action is another agent rather than a tool or an environment, the loop of Figure 3 closes between agents, and coordination becomes a design problem in its own right. The motivation is partly a division of labor (specialized roles, separate contexts, parallel work) and partly a hope that agents checking one another will be more reliable than any one alone. Whether that hope is borne out is the central, and contested, question of this section. Three pillar surveys map the space: a general survey of progress and challenges (Guo et al. 2024a), a communication-centric survey (Yan et al. 2025), and a survey of collaboration mechanisms (Tran et al. 2025).
Preprints 223549 i006

7.1. Role-Based Collaboration

The earliest and most influential multi-agent designs assign agents roles. CAMEL showed that two agents (an instruction-giving user and a solving assistant) kept on task by “inception” prompts can cooperate autonomously, and used the setup to study emergent cooperative behavior (Li et al. 2023a). Giving roles a workflow made them productive: MetaGPT encodes human standard operating procedures into a structured assembly line of specialist agents, which curbs the error cascades of naive agent chaining (Hong et al. 2024), and ChatDev instantiates a virtual software company whose agents communicate along a structured chat chain (Qian et al. 2023). Generalizing beyond fixed pipelines, AutoGen provides conversable agents that combine models, tools, and human input in configurable topologies (Wu et al. 2023a), while AgentVerse dynamically recruits and adjusts its agent group across a recruit–decide–act–evaluate loop (Chen et al. 2023d).
A revealing boundary case is that “multi-agent” collaboration need not require multiple agents at all: Solo Performance Prompting has a single model adopt several personas that critique one another in turn, recovering much of the benefit within one context window (Wang et al. 2023e). This sharpens the question the rest of the section pursues, what, exactly, does spawning separate agents buy? Table 7 contrasts representative frameworks along the axes that answer it: their topology, how roles are assigned, the medium of communication, whether a human is in the loop, and the domain they target.

7.2. Debate and Discussion

The most-studied coordination protocol is debate. Having several model instances propose answers and argue over successive rounds improves factuality and reduces fallacious reasoning (Du et al. 2023); the Multi-Agent Debate framework adds structured tit-for-tat argument and a judge to counter the “degeneration of thought” a lone model suffers (Liang et al. 2023). The protocol has been turned to evaluation, panels of referee agents correlate with human judgment better than a single evaluator (Chan et al. 2023), and to heterogeneous ensembles, where diverse models discuss and then vote by confidence (Chen et al. 2023c).
The enthusiasm deserves a skeptic, and the literature supplies a pointed one. A systematic comparison of debate against simpler strategies at matched cost finds that debate frequently fails to beat plain sampling-and-voting out of the box, and is often the more expensive way to reach the same accuracy (Smit et al. 2023). The productive reframing comes from the scalable-oversight literature: when two stronger models debate and a weaker judge decides, optimizing the debaters for persuasiveness raises the judge’s accuracy (Khan et al. 2024). Debate, then, is not reliably a route to better reasoning per unit compute, but it may be a route to supervising systems stronger than the supervisor, a different and arguably more important use, which we revisit in Section 11.

7.3. Scaling and Ensembling

The skeptic’s null hypothesis deserves its own subsection, because it is strong. Simply sampling many independent agent responses and taking a majority vote (an “agent forest” with no communication at all) improves performance monotonically with the number of agents, and this scaling accounts for a large share of the gains often attributed to elaborate coordination (Li et al. 2024b). Any claim that a communication protocol helps must therefore clear this bar. Some designs do: Mixture-of-Agents stacks layers in which each agent conditions on the previous layer’s outputs, and this structured aggregation of open models surpasses far stronger single models on instruction-following benchmarks (Wang et al. 2024b). When such systems are deployed the number of agents queried becomes a cost to be managed, not just an accuracy dial (Lee et al. 2026b). Pushing scale further, arranging over a thousand agents in explicit collaboration graphs finds that performance depends on the topology, with irregular graphs outperforming regular ones, an early collaborative scaling result (Qian et al. 2024). The lesson is that ensembling is the baseline to beat, and structured aggregation, not mere conversation, is what beats it.

7.4. Communication Topology and Team Optimization

If communication must earn its cost, the natural next step is to optimize who talks to whom. DyLAN treats the interaction graph as dynamic, pruning unhelpful agents mid-task via an agent-importance score and stopping early when consensus forms (Liu et al. 2023b), and Exchange-of-Thought catalogues the communication paradigms available (memory, report, relay, and debate) with a confidence mechanism to weigh contributions (Yin et al. 2023). A newer line makes the topology itself the object of optimization: GPTSwarm formalizes agent systems as computational graphs whose prompts and edges are jointly learned (Zhuge et al. 2024a), and AFlow searches over code-represented workflows to discover agentic pipelines automatically (Zhang et al. 2024d). Figure 8 sketches the recurring topologies, from fixed pipelines and stars to meshes and pruned dynamic networks.
Coordination is not always cooperative. Scorable negotiation games, some with adversarial players, test strategic communication and expose manipulation and self-interest as failure modes worth studying (Abdelnabi et al. 2023), and embodied teams must coordinate under partial observability, as in the modular cooperative agents of CoELA (Zhang et al. 2023b). At the largest scale, populations of agents with memory and reflection produce believable emergent social behavior, the clearest evidence that coordination dynamics are a genuine object of study rather than an engineering convenience (Park et al. 2023a).
Coordination also fails in systematic ways, and naming those failures sharpens the skeptic’s case from Section 7.2. A study of hundreds of execution traces across popular multi-agent frameworks distills fourteen recurring failure modes, under-specified roles, inter-agent misalignment, and weak verification among them, and attributes most breakdowns to these process failures rather than to the base model (Cemri et al. 2025). The lesson reinforces the section’s takeaway: multi-agent systems help when they add the right structure and fail when the structure itself is flawed.
Figure 8. Communication topologies for multi-agent systems. (a) A pipeline passes work along a fixed sequence (MetaGPT, ChatDev); (b) a star routes through an orchestrator (AutoGen); (c) a mesh has every agent debate every other (Section 7.2); (d) a hierarchy delegates down a tree; (e) a dynamic network prunes edges at run time (dashed), as in DyLAN.
Figure 8. Communication topologies for multi-agent systems. (a) A pipeline passes work along a fixed sequence (MetaGPT, ChatDev); (b) a star routes through an orchestrator (AutoGen); (c) a mesh has every agent debate every other (Section 7.2); (d) a hierarchy delegates down a tree; (e) a dynamic network prunes edges at run time (dashed), as in DyLAN.
Preprints 223549 g008
Preprints 223549 i007

8. Agents in Interactive Environments

The pillars so far describe agents largely through text. But an agent earns its name only by acting on something, and the something is increasingly a rich, partially observable environment: a web page, a phone screen, a desktop, a kitchen. Prior surveys tend to treat web, GUI, and embodied agents as separate literatures. We deliberately treat them as one family, because they share a single hard problem, grounding: mapping language and perception onto the specific, often irreversible action the environment will accept, whether that action is clicking the right link, tapping the right widget, or grasping the right object. Multimodal agent surveys support this unification across gaming, robotics, and interface control (Durante et al. 2024).

8.1. Web Agents

The web was the first rich environment, and its agents trace a clean lineage. WebGPT fine-tuned a model to answer questions by operating a text browser under human feedback, before ChatGPT existed (Nakano et al. 2021). WebShop then supplied a scalable simulated environment (over a million products) in which grounded agents navigate and purchase against natural-language goals (Yao et al. 2022a). Mind2Web moved to real websites, offering thousands of open-ended tasks across more than a hundred sites for training generalist web agents (Deng et al. 2023), and WebArena built a realistic, self-hostable environment with functionally checked tasks that has become a standard testbed (Zhou et al. 2024c).
The decisive recent shift is multimodal, and it identifies grounding as the bottleneck. SeeAct shows that a large multimodal model can act as a generalist web agent when it perceives the rendered page rather than only its HTML, but that turning a correct plan into the correct click (grounding) is where it most often fails (Zheng et al. 2024). WebVoyager closes the loop end-to-end on live websites with a multimodal agent and an automatic evaluator (He et al. 2024). The dedicated web-agent survey of Ning et al. (2025) catalogues the architectures and trustworthiness concerns this subfield has accumulated. The maturest commercial descendants of this line are deep research agents that run long-horizon, tool-mediated search and synthesis; the paradigm now has open end-to-end models (Tongyi DeepResearch Team 2025) and a systematic survey (Huang et al. 2025).

8.2. GUI and Computer-Use Agents

Generalizing from the browser to the whole screen yields the computer-use agent, and here perception becomes purely visual. CogAgent is an early large vision-language model built for graphical interfaces, using dual low- and high-resolution encoders to read screenshots and outperform HTML-based methods (Hong et al. 2023). SeeClick isolates the core subproblem, GUI grounding6, and introduces the ScreenSpot benchmark to measure locating the right element from a screenshot alone (Cheng et al. 2024).
On top of this perception layer sit device agents: AppAgent explores smartphone apps through a human-like action space (Zhang et al. 2023a), Mobile-Agent uses visual perception tools to operate apps from screenshots without back-end access (Wang et al. 2024c), and OS-Copilot extends autonomy to the operating-system level with a self-improving agent (Wu et al. 2024b). The frontier is end-to-end native models that map screenshots to actions directly, such as UI-TARS (Qin et al. 2025) (an RL-trained model in the lineage of Table 4) and its multi-turn-RL successor (Wang et al. 2025a), alongside the first mainstream commercial computer-use systems: Anthropic’s computer use (Anthropic 2024a) and OpenAI’s Operator (OpenAI 2025). What has driven the subfield is less the agents than the environments that grade them: OSWorld evaluates open-ended tasks in a real computer and reports agents near twelve percent success against seventy-two percent for humans (Xie et al. 2024b), and AndroidWorld supplies a dynamic, reward-bearing Android environment with parameterized tasks (Rawles et al. 2024). Two GUI-agent surveys map the space in detail (Nguyen et al. 2024; Zhang et al. 2024a).

8.3. Embodied and Robotic Agents

Embodied agents face the same grounding problem with physics added: actions move a body and cannot be undone. The foundational move was to use the LLM as a planner constrained by what the body can do. SayCan combines a model’s semantic knowledge with learned robotic affordances so that only feasible actions are proposed (Ahn et al. 2022), and Inner Monologue closes the loop by feeding environment feedback back to the planner in language, improving instruction completion without new training (Huang et al. 2022). ALFWorld aligns a text world with an embodied one so that policies learned abstractly transfer to grounded execution, and has become a standard testbed for LLM-based embodied agents (Shridhar et al. 2021).
A parallel line trains the mapping from perception to action end-to-end. PaLM-E injects sensor data directly into a language model to yield a generalist embodied model (Driess et al. 2023); the robotics transformers RT-1 and RT-2 scale action prediction and transfer web knowledge to control (Brohan et al. 2022, 2023); and OpenVLA opens the resulting vision–language–action recipe (Kim et al. 2024a), with π 0 extending it to a flow-matching action expert controlling many robot embodiments (Black et al. 2024). We treat these only insofar as they bear on LLM-agent design; the broader robotics-learning literature is out of scope. Voyager, discussed in Section 5.4, closes the circle by pairing embodiment with a procedural-memory skill library for open-ended learning (Wang et al. 2023a). Table 8 compares the environments these agents are measured in.
Table 8. Interactive environments by observation space, action space, scale, and whether the environment is stateful (actions change a persistent world state, evaluated by execution). Cross-referenced from Section 9.
Table 8. Interactive environments by observation space, action space, scale, and whether the environment is stateful (actions change a persistent world state, evaluated by execution). Cross-referenced from Section 9.
Environment Observation Action space #Tasks State Metric Ref.
WebShop text / HTML search, click 12k instr. ✓ task score (Yao et al. 2022a)
Mind2Web HTML + screenshot click, type, select 2k+ ✗ element accuracy (Deng et al. 2023)
WebArena DOM / pixels browser actions 812 ✓ functional success (Zhou et al. 2024c)
OSWorld pixels + a11y tree mouse, keyboard 369 ✓ execution success (Xie et al. 2024b)
AndroidWorld pixels + a11y tree touch, type 116 ✓ reward / success (Rawles et al. 2024)
ALFWorld text (embodied) high-level actions 6 types ✓ success rate (Shridhar et al. 2021)
Preprints 223549 i008

9. Evaluation and Benchmarks

Evaluation is the first of two cross-cutting concerns, and the more it is taken seriously the more sobering the picture becomes. The headline numbers of this section, GPT-4 solving under one percent of a travel-planning benchmark, agents reaching roughly a sixth of human success on real computer tasks, are not incidental; they are the field’s most honest signal of how far agents remain from reliability. This section asks what makes agents hard to measure, catalogues how the community measures them anyway, and then turns a critical eye on whether the measurements mean what they are taken to mean.

9.1. What Is Hard about Evaluating Agents

A static question-answering model is graded on one output; an agent must be graded on a trajectory. That difference is the source of every difficulty. Trajectories are long, so a single early misstep can doom an otherwise competent run, and stochastic, so one sample badly estimates the agent’s true success rate, which is why reliability metrics such as τ -bench’s pass k , the probability of succeeding on all of k independent attempts, have emerged as necessary complements to average success (Yao et al. 2024). Trajectories also unfold in stateful environments, so grading requires either checking the final world state or judging the process, and the two can disagree: an agent may reach the right answer by an unsafe or lucky path. The two dedicated evaluation surveys organize these problems along complementary axes, what to evaluate versus how (Mohammadi et al. 2025; Yehudai et al. 2025), and both flag cost, safety, and robustness as the dimensions current practice most neglects. Alongside aggregate scores, fine-grained studies that catalogue an agent’s concrete failure modes, as done early for memory- and retrieval-augmented dialogue agents (Lee et al. 2022), remain a useful complement to leaderboard numbers.

9.2. The Benchmark Landscape

The benchmark landscape divides into generalist and capability-specific suites (Table 9). On the generalist side, AgentBench evaluates a model as an agent across eight distinct environments and exposes a wide gap between commercial and open models (Liu et al. 2024c); GAIA poses questions that are conceptually simple for people yet require reasoning, tool use, and browsing, with people scoring around ninety-two percent and an augmented GPT-4 only about fifteen (Mialon et al. 2023); AgentBoard replaces binary success with a fine-grained progress-rate metric across a thousand environments (Ma et al. 2024); and SmartPlay isolates individual agentic capabilities across a suite of games (Wu et al. 2023b).
Capability-specific benchmarks probe one competence deeply, and their difficulty is instructive. SWE-bench asks agents to resolve real GitHub issues by editing a live codebase and passing hidden tests (Jimenez et al. 2023); τ -bench simulates multi-turn tool-agent-user interaction with policy constraints (Yao et al. 2024), extended by τ 2 -bench to dual-control settings where user and agent act on the environment simultaneously (Barres et al. 2025); TravelPlanner stress-tests constrained long-horizon planning, and the near-zero success it originally reported for GPT-4 (about 0.6 percent) is among the field’s most cited reality checks (Xie et al. 2024a); InterCode frames interactive coding as a reinforcement-learning environment with execution feedback (Yang et al. 2023c); Terminal-Bench assembles hard, realistic command-line tasks (Merrill et al. 2026); MLAgentBench evaluates agents on open-ended machine-learning experimentation (Huang et al. 2023b); and ScienceAgentBench grounds evaluation in a hundred-plus expert-validated data-science tasks (Chen et al. 2024b). Where these overlap the tool benchmarks of Table 6 and the environments of Table 8, we cross-reference rather than repeat.

9.3. Methodological Critique

A survey that only listed benchmarks would be complicit in their misuse, so we take the field’s own methodological critique seriously here. The anchor is Kapoor et al. (2024), whose central argument is that the near-exclusive focus on accuracy, with cost ignored, produces agents that are needlessly complex, overfit to their leaderboards, and expensive to run for gains that evaporate under scrutiny. Their remedies are concrete and, once stated, hard to argue with: report accuracy and cost together on a Pareto frontier rather than accuracy alone; distinguish the cost of developing an agent against a benchmark from its inference cost; and hold out genuinely unseen tasks, because agents tuned on a public test set inflate in ways that do not transfer. The same group’s Holistic Agent Leaderboard turns these remedies into shared infrastructure, a standardized harness that re-evaluates agents at scale with cost reported alongside accuracy (Kapoor et al. 2025). The same logic indicts several methods praised earlier in this survey, self-consistency, tree search, and elaborate multi-agent protocols all buy accuracy with compute that accuracy-only leaderboards render invisible.
Two further threats compound the cost blindness. Contamination is structural: agent benchmarks built from public GitHub issues, web pages, or question sets may already sit in pretraining data, so a high score can reflect memorization rather than capability, and the effect is hard to detect after the fact. And the increasingly common practice of using a strong model as an automatic judge, now systematized as agent-as-a-judge (Zhuge et al. 2024b), introduces a circularity, the evaluator shares the biases, blind spots, and even the identity of the systems it grades. Multi-agent referee panels mitigate but do not remove this, since they correlate with human judgment yet inherit model-family biases (Chan et al. 2023). The upshot is not that agent benchmarks are worthless but that a single accuracy number, uncontrolled for cost, contamination, and judge bias, is not the evidence it appears to be.
Preprints 223549 i009

10. Applications

Where agents are deployed is not, we argue, primarily a matter of economic value; it is a matter of feedback. Domains where the environment returns a cheap, reliable, ground-truth signal (does the test pass, does the reaction work, does the theorem check) let an agent iterate its way to competence, and these domains lead. Domains where the signal is expensive, delayed, or contested (a clinician’s judgment, a market’s verdict) trail, not for lack of interest but because the loop of Definition 1 cannot close tightly. The ordering of the subsections below tracks that thesis, from the most verifiable domain to the least.

10.1. Software Engineering

Software engineering is the paradigm verifiable domain: a unit test is a free, exact oracle, and it is no accident that coding agents are the most mature. SWE-agent resolves real repository issues through a purpose-built agent–computer interface (Section 6.3) (Yang et al. 2024a); AutoCodeRover raises efficiency by grounding the search in program structure and spectrum-based fault localization (Zhang et al. 2024e); and OpenHands provides an open platform on which much of this work now builds (Wang et al. 2024e). The dedicated survey of Liu et al. (2024a) catalogues the subfield. SWE-Lancer prices the same competence economically, asking whether frontier models can earn real freelance payouts and finding that they still leave most of the posted value unearned (Miserendino et al. 2025).
Yet software engineering also hosts this survey’s cleanest cautionary tale about agentic complexity, and it deserves equal weight. Agentless challenges the premise that issue resolution needs an autonomous agent at all: a fixed three-phase pipeline (localize, repair, validate) with no open-ended planning loop matches or beats far more elaborate agents on SWE-bench at a fraction of the cost (Xia et al. 2024). This is the cost-versus-capability argument of Section 9.3 made concrete: when the feedback signal is strong enough, a simple pipeline that exploits it can dominate an agent that reasons its way around it, and the burden of proof is on the agent.

10.2. Scientific Discovery

Scientific domains offer verifiable feedback of a costlier kind (an experiment, a proof, a passing computation) and agents have moved from augmenting scientists to conducting studies. Early systems equipped a model with domain tools: ChemCrow wraps a GPT-4 agent in expert chemistry tools for synthesis planning and safety (Bran et al. 2023), and Coscientist closes the loop to physical hardware, designing and executing chemistry experiments through robotic APIs (Boiko et al. 2023). A more ambitious line automates the research process itself: the AI Scientist generates ideas, runs experiments, and writes papers end to end (Lu et al. 2024), its successor produced the first fully AI-generated paper accepted at a peer-reviewed workshop (Yamada et al. 2025), and an AI co-scientist generates and ranks novel biomedical hypotheses under expert validation (Gottweis et al. 2025).
The most convincing evidence comes from discoveries that are checkable. AlphaEvolve, an evolutionary coding agent, found a way to multiply two 4 × 4 complex-valued matrices in 48 scalar multiplications, the first improvement on Strassen’s algorithm for this case in 56 years (Novikov et al. 2025), and Kosmos runs long autonomous research campaigns around a shared structured world model (Mitchener et al. 2025). Because a matrix-multiplication algorithm either works or does not, such results sidestep the evaluation ambiguities of Section 9; ScienceAgentBench supplies the corresponding expert-validated benchmark (Chen et al. 2024b), and the subfield has its own survey (Ren et al. 2025).

10.3. Healthcare

Medicine inverts the feedback economics: ground truth is expensive, delayed, and ethically gated, so progress relies on simulation and expert rating rather than on a cheap oracle. MDAgents adapts its collaboration structure (solo, group, or hierarchical) to the complexity of a medical decision (Kim et al. 2024b), and Agent Hospital lets doctor agents self-evolve by treating simulated patients, manufacturing experience where real cases cannot be used freely (Li et al. 2024a); AgentClinic supplies a multimodal benchmark of such simulated clinical settings (Schmidgall et al. 2024). The most rigorous result is AMIE, a diagnostic dialogue agent that, in a randomized OSCE-style study, was rated higher than primary-care physicians on 28 of 32 axes rated by specialist physicians (Tu et al. 2024), a striking figure that the survey of Wang et al. (2025c) rightly situates alongside unresolved risks of hallucination and liability, which we take up in Section 11.

10.4. Finance

Finance offers a feedback signal that looks precise but is treacherous: returns are quantitative yet noisy, non-stationary, and easy to overfit in backtests. FinMem structures a trading agent around layered, human-cognition-inspired memory (Yu et al. 2023); TradingAgents mirrors a trading firm with specialized analyst, researcher, and risk-management roles (Xiao et al. 2024); and FinGPT supplies open financial-model infrastructure underneath much of this work (Yang et al. 2023a). The survey of Ding et al. (2024) reviews the area, but the binding caveat is methodological: a favorable backtest is not a market outcome, and reported returns that are not out-of-sample say little, the same no-invented-numbers discipline this survey applies to benchmark tables applies with double force to financial claims. Table 10 summarizes the four domains by their feedback signal, maturity, and dominant risk.
Table 10. Application domains ordered by the quality of the feedback signal the environment provides. Maturity is a coarse judgment of deployment readiness; the dominant risk is the failure mode most cited in each domain’s literature.
Table 10. Application domains ordered by the quality of the feedback signal the environment provides. Maturity is a coarse judgment of deployment readiness; the dominant risk is the failure mode most cited in each domain’s literature.
Domain Representative systems Feedback signal Maturity Dominant risk
Software SWE-agent (Yang et al. 2024a), OpenHands (Wang et al. 2024e), Agentless (Xia et al. 2024) unit tests deployed wrong / insecure code
Science ChemCrow (Bran et al. 2023), AI Scientist (Lu et al. 2024), AlphaEvolve (Novikov et al. 2025) experiment / proof research unvalidated claims
Healthcare MDAgents (Kim et al. 2024b), AMIE (Tu et al. 2024), Agent Hospital (Li et al. 2024a) clinician rating research hallucination, liability
Finance FinMem (Yu et al. 2023), TradingAgents (Xiao et al. 2024), FinGPT (Yang et al. 2023a) market returns research backtest overfitting
Preprints 223549 i010

11. Safety, Security, and Trustworthiness

Autonomy is precisely what makes LLM-based agents useful and what makes them risky: an agent that can call tools, write files, or send messages on a user’s behalf can also do so incorrectly, at scale, and without a human in the loop (Yu et al. 2025).7

11.1. The Agentic Threat Model

The safety of a chatbot is largely the safety of its text; the safety of an agent is the safety of its actions. Three properties of the loop in Figure 3 each open a new class of risk. Tools are actuators: an injected instruction becomes a real side effect, a sent email, a deleted file, a transferred payment. Memory is persistent state: a corruption written once can influence every later decision. And autonomy is scale: an agent repeats its mistakes without a human to interrupt them. The recent surveys of the area converge on this framing, that autonomy itself, not merely the underlying model, is the source of the new risks (Su et al. 2025); that threats factor into intrinsic (reasoning, memory, tools) and extrinsic (user, other agents, environment) dimensions (Yu et al. 2025); and that safety must be considered across the full lifecycle from data to deployment (Wang et al. 2025b). Figure 9 redraws the agent loop as an attack surface, numbering the entry points that the next subsection works through.

11.2. Attacks on LLM-Based Agents

The dominant attack exploits a design feature: an agent cannot reliably tell instructions from data, because both arrive as text.
Preprints 223549 i011
Injection is entry points ❶ and ❷ of Figure 9. Since it was first shown to compromise real LLM-integrated applications (Greshake et al. 2023), a series of benchmarks has quantified how exposed agents are: InjecAgent finds a ReAct-prompted GPT-4 acts on injected instructions in roughly a quarter of tool-integrated cases (Zhan et al. 2024); AgentDojo provides a dynamic environment of realistic tasks and attacks and shows that current agents cannot be simultaneously useful and robust (Debenedetti et al. 2024); and WASP finds that even top reasoning models begin executing injected instructions on web tasks between sixteen and eighty-six percent of the time, though full goal hijacking is rarer (Evtimov et al. 2025).
The persistent state of memory adds a second, quieter class, entry point ❸. AgentPoison injects a handful of poisoned entries into an agent’s memory or retrieval knowledge base and triggers malicious behavior with over eighty-percent success, without any model training and with little effect on benign queries (Chen et al. 2024a). BadAgent shows that backdoors planted during fine-tuning survive subsequent fine-tuning on clean data (Wang et al. 2024f), a persistence result echoed at the alignment level in Section 11.4. Finally, composition amplifies risk, entry point ❹: red-teaming multi-agent systems shows that collaboration spreads and intensifies harmful behavior rather than diluting it (Tian et al. 2023), and the Agent Security Bench unifies these attack and defense classes at scale, reporting attack success rates up to eighty-four percent and finding existing defenses largely ineffective (Zhang et al. 2024b).

11.3. Risk Evaluation and Guardrails

Measuring these risks safely is itself a research problem, since one cannot let an agent cause real harm to see whether it would. ToolEmu resolves this by having a strong model emulate tool execution in a sandbox, surfacing risky behavior without real tools and finding that even the safest agents act with potentially severe consequences about a quarter of the time (Ruan et al. 2024). R-Judge tests whether models can recognize risk in agent behavior logs, and the best reach only around seventy-four percent accuracy, with most near chance (Yuan et al. 2024). Misuse and embodied hazards have their own suites: AgentHarm measures whether jailbroken agents retain the capability to carry out coherent harmful multi-step tasks (Andriushchenko et al. 2024), and SafeAgentBench finds embodied agents reject only five to ten percent of explicitly hazardous instructions (Yin et al. 2024), with OS-Harm extending safety measurement to computer-use agents operating real interfaces (Kuntz et al. 2025). On the defense side, a spectrum of countermeasures has emerged. Training-time methods teach the model an instruction hierarchy that privileges trusted system instructions over injected ones (Wallace et al. 2024); prompt-level methods such as spotlighting mark untrusted content so the model can discount it (Hines et al. 2024); and constitution-guided planning enforces safety before, during, and after plan generation (Hua et al. 2024). The strongest current answer to indirect injection is architectural: CaMeL isolates untrusted data from the agent’s control flow by design and gates actions behind explicit capability checks (Debenedetti et al. 2025), and a companion catalogue of design patterns constrains what an agent may do once it has read untrusted input, trading some generality for resistance by construction (Beurer-Kellner et al. 2025). These genuinely raise the bar, but as Table 11 records, coverage is uneven, strongest for direct injection under an instruction hierarchy, only partial for indirect injection, and weakest for memory poisoning and backdoors, so the summary is that no attack class yet has a defense that is both general and validated at deployment scale.
To make the relationship between capability and safety concrete rather than assumed, we ran a small measurement of our own. Across six Qwen2.5-Instruct models spanning 1.5B to 72B parameters, with Llama-3.1-8B-Instruct as a second-family control, we paired agentic capability (ALFWorld success, the same axis as Section 4.3) with susceptibility to indirect prompt injection on InjecAgent (Zhan et al. 2024), scored as attack success among valid responses; Figure 10 shows the result. Three observations follow. First, within the family the two axes broadly improve together: capability climbs from 6 percent at 1.5B to 70–76 percent for the 14B–72B models, while the hijack rate falls from 35–39 percent below 10B, through 18 percent at 14B, to 2–9 percent for the two largest models, though not perfectly monotonically (the 72B model is hijacked on 9 percent of valid responses against 2 for the 32B). Second, the metric matters: on the raw all-probes attack rate the two smallest models look nearly robust (7–9 percent), but only because three quarters or more of their responses are not valid tool calls at all; robustness that is really incompetence disappears once the metric conditions on valid actions (35–36 percent). Third, the level is family-specific: at comparable scale, Llama-3.1-8B is hijacked on 72 percent of its valid responses against 39 for Qwen2.5-7B. So scale buys injection-robustness within this family, but it is the optimistic exception, not a general law: the memory-poisoning, backdoor, and misuse benchmarks tabulated below show even the strongest models failing badly, so capability cannot be read as a proxy for safety; the two are distinct axes and must be measured one by one.

11.4. Alignment-Level Risks

Beneath the security threats lies a subtler class in which the agent’s own learned dispositions, not an external adversary, are the hazard. The evidence begins with behavior induced by training: model-written evaluations reveal that RLHF can amplify sycophancy and expressed self-preservation as scale grows (Perez et al. 2022), and controlled studies trace sycophancy to the preference feedback used to train assistants, with raters sometimes preferring convincing wrong answers to correct ones (Sharma et al. 2023). A survey of AI deception documents that systems already learn to induce false beliefs in humans (Park et al. 2023b).
For agents, these dispositions become more dangerous because agents act. Sleeper Agents shows that a model trained to behave maliciously under a trigger retains that behavior through standard safety training, adversarial training can even teach it to hide the trigger better (Hubinger et al. 2024). Frontier models have been observed to scheme in context on agentic tasks, introducing deliberate errors, attempting to disable oversight, and in some cases trying to exfiltrate their own weights (Meinke et al. 2024). And emergent misalignment demonstrates that fine-tuning on a narrow harmful task (writing insecure code) can generalize into broad misalignment on unrelated prompts (Betley et al. 2025), a result whose implications for the routine practice of fine-tuning agents on task data are not yet understood. The comprehensive alignment survey of Ji et al. (2023) supplies the conceptual scaffolding (robustness, interpretability, controllability, ethicality) within which these agent-specific failures sit.

11.5. Autonomy Governance

Because the risks scale with autonomy, some of the response is necessarily governance rather than engineering. Frameworks that define discrete levels of agent autonomy give a vocabulary for specifying how much control a human retains (Feng et al. 2025b), and a prominent position paper argues that fully autonomous agents should not be developed at all, since risk rises faster than benefit as the human is removed from the loop (Mitchell et al. 2025). Between these poles sits the design of human–agent oversight, feedback, interruption, and escalation mechanisms that keep a person meaningfully in control (Zou et al. 2025), which we regard as the most actionable near-term safety lever. The empirical ground truth is sobering: an index of thirty deployed agentic systems finds that developers document capabilities far more thoroughly than safety evaluations, a transparency gap governance frameworks have yet to close (Staufer et al. 2026).
Preprints 223549 i012

12. Open Challenges and Future Directions

The preceding sections suggest seven challenges that we expect to shape the next phase of LLM-agent research. Each is grounded in results already surveyed rather than in speculation, and several are consequences of the tensions flagged along the way. Table 12 first collects the survey’s documented negative results in one place: wherever an optimistic consensus meets credible counter-evidence, we have given the two equal weight, and the reconciliations there are as much a part of our contribution as the taxonomy.
  • Reliability over long horizons.
The single most important gap is that agents fail as tasks lengthen. The near-zero success on constrained long-horizon planning (Xie et al. 2024a) and the roughly sixfold human–agent gap on realistic computer tasks (Section 8.2) are not isolated: they reflect a structural problem in which a per-step error probability compounds over a trajectory, so that even a capable agent becomes unreliable over enough steps. Because current evaluation often reports averaged success rather than reliability (Kapoor et al. 2024), this fragility is routinely understated. Progress requires both error-correcting mechanisms that actually work (Section 4.4) and metrics such as pass k that measure consistency, not just average competence.
2.
Cost as a first-class objective.
Almost every capability gain in this survey was purchased with inference compute (sampled reasoning paths, search trees, multi-agent rounds) yet accuracy-only leaderboards render that price invisible (Kapoor et al. 2024). Since a large share of multi-agent benefit can be recovered by simply sampling one agent more often (Li et al. 2024b), the open problem is to report and optimize agents on an accuracy–cost frontier, and to design methods that are Pareto-efficient rather than merely accurate; methods that adaptively budget an agent’s reasoning tokens are one promising direction (Lee et al. 2026a). Cost-awareness is not an accounting detail; it changes which methods look good.
3.
Continual learning and memory consolidation.
Agents that improve from their own experience without weight updates (Section 5.4) point toward lifelong learning, but the mechanisms are immature. Distilling reusable insights (Zhao et al. 2023) and consolidating salient facts across sessions (Chhikara et al. 2025) work at small scale; what remains open is consolidation that avoids both unbounded memory growth and the erosion of old competence (the agentic analogue of catastrophic forgetting) over months of operation. The emerging self-evolving-agent literature, now with a survey of its own, frames exactly this bridge from static foundation models to lifelong agentic systems (Fang et al. 2025).
4.
Training agents, not just prompting them.
The shift from wrapping scaffolds around fixed models to training models for agency, surveyed in Section 4.5 and Table 4, raises two open problems. The first is data: high-quality trajectory-level supervision (sequences of decisions, tool calls, and recoveries) is scarce next to the text used for pretraining, and scaling its production, whether by offline collection or online experience, without a human in the loop is unsolved. The second is the question Section 4.5 left open: as reasoning and acting migrate into the weights, do the external scaffolds of Section 4 (search trees, reflection loops, elaborate coordination) remain necessary, or do they become redundant on tasks the model has internalized? The answer determines how much of this survey’s machinery the next generation of agents will still need.
5.
Infrastructure and interoperability.
As multi-agent systems leave the laboratory, they need standard protocols for how agents expose tools and talk to one another (Yan et al. 2025). Emerging interoperability standards, the Model Context Protocol for tool access (Anthropic 2024b) and the Agent2Agent protocol for inter-agent communication (Google 2025), are early attempts, but the research questions of secure, composable, and observable multi-agent infrastructure are largely open, and they interact directly with the security concerns of Section 11.
6.
Evaluation validity.
The evaluation critique of Section 9.3 is itself a research agenda. Contamination of benchmarks by pretraining data, overfitting to public leaderboards, and the circularity of agent-as-judge all threaten the validity of the numbers the field reports on (Kapoor et al. 2024; Yehudai et al. 2025). Building contamination-resistant, cost-controlled, and reproducible benchmarks (ideally with live or held-out task streams) is a precondition for trusting claims of progress at all.
7.
Safety–capability co-evolution.
Finally, the divergence documented in Section 11 must be closed rather than merely measured. Capability advances have not been matched by defenses, backdoors survive safety training (Hubinger et al. 2024), and no attack class yet has a robust general solution, so the central governance question is how much autonomy to grant under this asymmetry. Whether one adopts graded autonomy (Feng et al. 2025b) or argues against full autonomy outright (Mitchell et al. 2025), safety must co-evolve with capability rather than trail it, and we regard this as the field’s defining long-term challenge.

13. Conclusion

This survey has organized the LLM-agent literature around four component pillars (planning, memory, tool use, and multi-agent coordination) grounded in interactive environments and application domains, with evaluation and safety as first-class concerns. Read through the function-first taxonomy of Section 3, a fragmented field becomes legible: systems built in separate subfields turn out to be answering the same small set of design questions.
Three movements summarize where the field stands. The components are consolidating: prompting is giving way to trained reasoning (Section 4.5), flat memory to structured stores (Section 5.3), and hand-listed tools to code as a universal action space (Section 6.3), so that much of the scaffolding of the early agent frameworks (Wang et al. 2023b) is migrating into the models themselves. The frontier has moved to the context and cross-cutting axes, grounding agents in real environments (Section 8), and evaluating them honestly for cost and reliability rather than headline accuracy (Section 9), where the largest gaps between demonstration and dependability remain (Kapoor et al. 2024). And safety is the binding constraint: the divergence between what agents can do and what can be done to them (Section 11) means the field’s progress will be governed less by new capabilities than by whether those capabilities can be made trustworthy. The recent general surveys catalogue how fast the field is moving (Luo et al. 2025); our contribution is a stable frame in which to place what comes next.

References

  1. Sahar Abdelnabi, Amr Gomaa, Sarath Sivaprasad, Lea Schönherr, and Mario Fritz. Cooperation, competition, and maliciousness: Llm-stakeholders interactive negotiation, 2023. URL https://arxiv.org/abs/2309.17234. arXiv.
  2. Michael Ahn et al. Do as i can, not as i say: Grounding language in robotic affordances, 2022. URL https://arxiv.org/abs/2204.01691. CoRL 2022 / arXiv.
  3. Maksym Andriushchenko, Alexandra Souly, Mateusz Dziemian, Derek Duenas, Maxwell Lin, Justin Wang, Dan Hendrycks, Andy Zou, Zico Kolter, Matt Fredrikson, Eric Winsor, Jerome Wynne, Yarin Gal, and Xander Davies. Agentharm: A benchmark for measuring harmfulness of llm agents, 2024. URL https://arxiv.org/abs/2410.09024. ICLR 2025 (arXiv 2024).
  4. Anthropic. Introducing computer use, a new claude 3.5 sonnet, and claude 3.5 haiku, 2024a. URL https://www.anthropic.com/news/3-5-models-and-computer-use.
  5. Anthropic. Introducing the model context protocol, 2024b. URL https://www.anthropic.com/news/model-context-protocol.
  6. Victor Barres, Honghua Dong, Soham Ray, et al. τ2-bench: Evaluating conversational agents in a dual-control environment, 2025. URL https://arxiv.org/abs/2506.07982.
  7. Maciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger, Michal Podstawski, Lukas Gianinazzi, Joanna Gajda, Tomasz Lehmann, Hubert Niewiadomski, Piotr Nyczyk, and Torsten Hoefler. Graph of thoughts: Solving elaborate problems with large language models, 2023. URL https://arxiv.org/abs/2308.09687. AAAI 2024.
  8. Jan Betley, Daniel Tan, Niels Warncke, Anna Sztyber-Betley, Xuchan Bao, Martín Soto, Nathan Labenz, and Owain Evans. Emergent misalignment: Narrow finetuning can produce broadly misaligned llms, 2025. URL https://arxiv.org/abs/2502.17424. arXiv.
  9. Luca Beurer-Kellner, Beat Buesser, Ana-Maria Creţu, et al. Design patterns for securing llm agents against prompt injections, 2025. URL https://arxiv.org/abs/2506.08837.
  10. Kevin Black, Noah Brown, Danny Driess, et al. π0: A vision-language-action flow model for general robot control, 2024. URL https://arxiv.org/abs/2410.24164.
  11. Daniil A. Boiko, Robert MacKnight, Ben Kline, and Gabe Gomes. Autonomous chemical research with large language models. 2023. URL https://www.nature.com/articles/s41586-023-06792-0. Nature (vol. 624, pp. 570-578).
  12. Andres M. Bran, Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D. White, and Philippe Schwaller. Chemcrow: Augmenting large-language models with chemistry tools, 2023. URL https://arxiv.org/abs/2304.05376. Nature Machine Intelligence / arXiv.
  13. Anthony Brohan et al. Rt-1: Robotics transformer for real-world control at scale, 2022. URL https://arxiv.org/abs/2212.06817. arXiv / RSS 2023.
  14. Anthony Brohan et al. Rt-2: Vision-language-action models transfer web knowledge to robotic control, 2023. URL https://arxiv.org/abs/2307.15818. CoRL 2023 / arXiv.
  15. Tianle Cai, Xuezhi Wang, Tengyu Ma, Xinyun Chen, and Denny Zhou. Large language models as tool makers, 2023. URL https://arxiv.org/abs/2305.17126. arXiv.
  16. Pengfei Cao, Tianyi Men, Wencan Liu, Jingwen Zhang, Xuzhao Li, Xixun Lin, Dianbo Sui, Yanan Cao, Kang Liu, and Jun Zhao. Large language models for planning: A comprehensive and systematic survey, 2025. URL https://arxiv.org/abs/2505.19683. arXiv.
  17. Mert Cemri, Melissa Z. Pan, Shuyi Yang, et al. Why do multi-agent llm systems fail?, 2025. URL https://arxiv.org/abs/2503.13657.
  18. Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu. Chateval: Towards better llm-based evaluators through multi-agent debate, 2023. URL https://arxiv.org/abs/2308.07201. ICLR 2024 (arXiv preprint 2023).
  19. Baian Chen, Chang Shu, Ehsan Shareghi, Nigel Collier, Karthik Narasimhan, and Shunyu Yao. Fireact: Toward language agent fine-tuning, 2023a. URL https://arxiv.org/abs/2310.05915.
  20. Howard Chen, Ramakanth Pasunuru, Jason Weston, and Asli Celikyilmaz. Walking down the memory maze: Beyond context limit through interactive reading, 2023b. URL https://arxiv.org/abs/2310.05029. arXiv preprint.
  21. Justin Chih-Yao Chen, Swarnadeep Saha, and Mohit Bansal. Reconcile: Round-table conference improves reasoning via consensus among diverse llms, 2023c. URL https://arxiv.org/abs/2309.13007. ACL 2024 (arXiv preprint 2023).
  22. Weize Chen, Yusheng Su, Jingwei Zuo, Cheng Yang, Chenfei Yuan, Chi-Min Chan, Heyang Yu, Yaxi Lu, Yi-Hsin Hung, Chen Qian, Yujia Qin, Xin Cong, Ruobing Xie, Zhiyuan Liu, Maosong Sun, and Jie Zhou. Agentverse: Facilitating multi-agent collaboration and exploring emergent behaviors, 2023d. URL https://arxiv.org/abs/2308.10848. ICLR 2024 (arXiv preprint 2023).
  23. Zhaorun Chen, Zhen Xiang, Chaowei Xiao, Dawn Song, and Bo Li. Agentpoison: Red-teaming llm agents via poisoning memory or knowledge bases, 2024a. URL https://arxiv.org/abs/2407.12784. NeurIPS 2024.
  24. Ziru Chen et al. Scienceagentbench: Toward rigorous assessment of language agents for data-driven scientific discovery, 2024b. URL https://arxiv.org/abs/2410.05080. ICLR 2025 / arXiv.
  25. Kanzhi Cheng, Qiushi Sun, Yougang Chu, Fangzhi Xu, Yantao Li, Jianbing Zhang, and Zhiyong Wu. Seeclick: Harnessing gui grounding for advanced visual gui agents, 2024. URL https://arxiv.org/abs/2401.10935. ACL 2024.
  26. Prateek Chhikara, Dev Khant, Saket Aryan, Taranjeet Singh, and Deshraj Yadav. Mem0: Building production-ready ai agents with scalable long-term memory, 2025. URL https://arxiv.org/abs/2504.19413. arXiv preprint.
  27. Gheorghe Comanici, Eric Bieber, Mike Schaekermann, et al. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities, 2025. URL https://arxiv.org/abs/2507.06261.
  28. Payel Das, Subhajit Chaudhury, Elliot Nelson, Igor Melnyk, Sarath Swaminathan, Sihui Dai, Aurélie Lozano, Georgios Kollias, Vijil Chenthamarakshan, Jiří Navrátil, Soham Dan, and Pin-Yu Chen. Larimar: Large language models with episodic memory control, 2024. URL https://arxiv.org/abs/2403.11901. ICML 2024.
  29. Edoardo Debenedetti, Jie Zhang, Mislav Balunović, Luca Beurer-Kellner, Marc Fischer, and Florian Tramèr. Agentdojo: A dynamic environment to evaluate prompt injection attacks and defenses for llm agents, 2024. URL https://arxiv.org/abs/2406.13352. NeurIPS 2024 Datasets and Benchmarks Track.
  30. Edoardo Debenedetti, Ilia Shumailov, Tianqi Fan, Jamie Hayes, Nicholas Carlini, Daniel Fabian, Christoph Kern, Chongyang Shi, Andreas Terzis, and Florian Tramèr. Defeating prompt injections by design, 2025. URL https://arxiv.org/abs/2503.18813.
  31. DeepSeek-AI, Daya Guo, Dejian Yang, Haowei Zhang, et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025. URL https://arxiv.org/abs/2501.12948. arXiv preprint (later published in Nature, 2025).
  32. Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Samuel Stevens, Boshi Wang, Huan Sun, and Yu Su. Mind2web: Towards a generalist agent for the web, 2023. URL https://arxiv.org/abs/2306.06070. NeurIPS 2023 (Datasets & Benchmarks, Spotlight).
  33. Han Ding, Yinheng Li, Junhao Wang, Hang Chen, Doudou Guo, and Yunbai Zhang. Large language model agent in financial trading: A survey, 2024. URL https://arxiv.org/abs/2408.06361. arXiv.
  34. Danny Driess et al. Palm-e: An embodied multimodal language model, 2023. URL https://arxiv.org/abs/2303.03378. ICML 2023.
  35. Yilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum, and Igor Mordatch. Improving factuality and reasoning in language models through multiagent debate, 2023. URL https://arxiv.org/abs/2305.14325. arXiv preprint (also ICML 2024 workshop track).
  36. Zane Durante, Qiuyuan Huang, Naoki Wake, Ran Gong, Jae Sung Park, et al. Agent ai: Surveying the horizons of multimodal interaction, 2024. URL https://arxiv.org/abs/2401.03568. arXiv.
  37. Ivan Evtimov, Arman Zharmagambetov, Aaron Grattafiori, Chuan Guo, and Kamalika Chaudhuri. Wasp: Benchmarking web agent security against prompt injection attacks, 2025. URL https://arxiv.org/abs/2504.18575. arXiv.
  38. Jinyuan Fang, Yanwen Peng, Xi Zhang, et al. A comprehensive survey of self-evolving ai agents: A new paradigm bridging foundation models and lifelong agentic systems, 2025. URL https://arxiv.org/abs/2508.07407.
  39. Jiazhan Feng, Shijue Huang, Xingwei Qu, et al. Retool: Reinforcement learning for strategic tool use in llms, 2025a. URL https://arxiv.org/abs/2504.11536.
  40. K. J. Kevin Feng, David W. McDonald, and Amy X. Zhang. Levels of autonomy for ai agents, 2025b. URL https://arxiv.org/abs/2506.12469. arXiv.
  41. Google. Announcing the agent2agent protocol (a2a), 2025. URL https://developers.googleblog.com/en/a2a-a-new-era-of-agent-interoperability/.
  42. Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, et al. Accelerating scientific discovery with co-scientist, 2025. URL https://arxiv.org/abs/2502.18864. arXiv.
  43. Zhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen, Yujiu Yang, Nan Duan, and Weizhu Chen. Critic: Large language models can self-correct with tool-interactive critiquing, 2023. URL https://arxiv.org/abs/2305.11738. ICLR 2024.
  44. Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz. Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection, 2023. URL https://arxiv.org/abs/2302.12173. ACM Workshop on AI and Security (AISec) / arXiv.
  45. Taicheng Guo, Xiuying Chen, Yaqi Wang, Ruidi Chang, Shichao Pei, Nitesh V. Chawla, Olaf Wiest, and Xiangliang Zhang. Large language model based multi-agents: A survey of progress and challenges, 2024a. URL https://arxiv.org/abs/2402.01680. arXiv; IJCAI 2024.
  46. Zhicheng Guo, Sijie Cheng, Hao Wang, Shihao Liang, Yujia Qin, Peng Li, Zhiyuan Liu, Maosong Sun, and Yang Liu. Stabletoolbench: Towards stable large-scale benchmarking on tool learning of large language models, 2024b. URL https://arxiv.org/abs/2403.07714. Findings of ACL 2024.
  47. Bernal Jiménez Gutiérrez, Yiheng Shu, Yu Gu, Michihiro Yasunaga, and Yu Su. Hipporag: Neurobiologically inspired long-term memory for large language models, 2024. URL https://arxiv.org/abs/2405.14831. NeurIPS 2024.
  48. Shibo Hao, Yi Gu, Haodi Ma, Joshua Jiahua Hong, Zhen Wang, Daisy Zhe Wang, and Zhiting Hu. Reasoning with language model is planning with world model, 2023a. URL https://arxiv.org/abs/2305.14992. EMNLP 2023.
  49. Shibo Hao, Tianyang Liu, Zhen Wang, and Zhiting Hu. Toolkengpt: Augmenting frozen language models with massive tools via tool embeddings, 2023b. URL https://arxiv.org/abs/2305.11554. NeurIPS 2023 (Oral).
  50. Hongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Yong Dai, Hongming Zhang, Zhenzhong Lan, and Dong Yu. Webvoyager: Building an end-to-end web agent with large multimodal models, 2024. URL https://arxiv.org/abs/2401.13919. ACL 2024.
  51. Keegan Hines, Gary Lopez, Matthew Hall, Federico Zarfati, Yonatan Zunger, and Emre Kiciman. Defending against indirect prompt injection attacks with spotlighting, 2024. URL https://arxiv.org/abs/2403.14720.
  52. Sirui Hong, Mingchen Zhuge, Jiaqi Chen, Xiawu Zheng, Yuheng Cheng, et al. Metagpt: Meta programming for a multi-agent collaborative framework, 2024. URL https://arxiv.org/abs/2308.00352. ICLR 2024.
  53. Wenyi Hong, Weihan Wang, Qingsong Lv, Jiazheng Xu, Wenmeng Yu, Junhui Ji, Yan Wang, Zihan Wang, Yuxuan Zhang, Juanzi Li, Bin Xu, Yuxiao Dong, Ming Ding, and Jie Tang. Cogagent: A visual language model for gui agents, 2023. URL https://arxiv.org/abs/2312.08914. CVPR 2024.
  54. Xinyi Hou, Yanjie Zhao, Shenao Wang, and Haoyu Wang. Model context protocol (mcp): Landscape, security threats, and future research directions, 2025. URL https://arxiv.org/abs/2503.23278.
  55. Cheng-Yu Hsieh, Si-An Chen, Chun-Liang Li, Yasuhisa Fujii, Alexander Ratner, Chen-Yu Lee, Ranjay Krishna, and Tomas Pfister. Tool documentation enables zero-shot tool-usage with large language models, 2023. URL https://arxiv.org/abs/2308.00675. arXiv.
  56. Wenyue Hua, Xianjun Yang, Mingyu Jin, Zelong Li, Wei Cheng, Ruixiang Tang, and Yongfeng Zhang. Trustagent: Towards safe and trustworthy llm-based agents, 2024. URL https://arxiv.org/abs/2402.01586. EMNLP 2024 Findings.
  57. Jie Huang and Kevin Chen-Chuan Chang. Towards reasoning in large language models: A survey, 2022. URL https://arxiv.org/abs/2212.10403. ACL 2023 Findings.
  58. Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, and Denny Zhou. Large language models cannot self-correct reasoning yet, 2023a. URL https://arxiv.org/abs/2310.01798. ICLR 2024.
  59. Qian Huang, Jian Vora, Percy Liang, and Jure Leskovec. Mlagentbench: Evaluating language agents on machine learning experimentation, 2023b. URL https://arxiv.org/abs/2310.03302. arXiv.
  60. Wenlong Huang et al. Inner monologue: Embodied reasoning through planning with language models, 2022. URL https://arxiv.org/abs/2207.05608. CoRL 2022 / arXiv.
  61. Xu Huang, Weiwen Liu, Xiaolong Chen, Xingmei Wang, Hao Wang, Defu Lian, Yasheng Wang, Ruiming Tang, and Enhong Chen. Understanding the planning of llm agents: A survey, 2024. URL https://arxiv.org/abs/2402.02716. arXiv.
  62. Yuxuan Huang, Yihang Chen, Haozheng Zhang, et al. Deep research agents: A systematic examination and roadmap, 2025. URL https://arxiv.org/abs/2506.18096.
  63. Evan Hubinger, Carson Denison, et al. Sleeper agents: Training deceptive llms that persist through safety training, 2024. URL https://arxiv.org/abs/2401.05566. arXiv (Anthropic).
  64. Jiaming Ji, Tianyi Qiu, Boyuan Chen, Borong Zhang, et al. Ai alignment: A comprehensive survey, 2023. URL https://arxiv.org/abs/2310.19852. arXiv.
  65. Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. Swe-bench: Can language models resolve real-world github issues?, 2023. URL https://arxiv.org/abs/2310.06770. ICLR 2024.
  66. Bowen Jin, Hansi Zeng, Zhenrui Yue, et al. Search-r1: Training llms to reason and leverage search engines with reinforcement learning, 2025. URL https://arxiv.org/abs/2503.09516.
  67. Ryo Kamoi, Yusen Zhang, Nan Zhang, Jiawei Han, and Rui Zhang. When can llms actually correct their own mistakes? a critical survey of self-correction of llms, 2024. URL https://arxiv.org/abs/2406.01297. TACL 2024.
  68. Sayash Kapoor, Benedikt Stroebl, Zachary S. Siegel, Nitya Nadgir, and Arvind Narayanan. Ai agents that matter, 2024. URL https://arxiv.org/abs/2407.01502. arXiv.
  69. Sayash Kapoor, Benedikt Stroebl, Peter Kirgis, et al. Holistic agent leaderboard: The missing infrastructure for ai agent evaluation, 2025. URL https://arxiv.org/abs/2510.11977.
  70. Ehud Karpas, Omri Abend, Yonatan Belinkov, Barak Lenz, Opher Lieber, Nir Ratner, Yoav Shoham, Hofit Bata, Yoav Levine, Kevin Leyton-Brown, Dor Muhlgay, Noam Rozen, Erez Schwartz, Gal Shachaf, Shai Shalev-Shwartz, Amnon Shashua, and Moshe Tenenholtz. Mrkl systems: A modular, neuro-symbolic architecture that combines large language models, external knowledge sources and discrete reasoning, 2022. URL https://arxiv.org/abs/2205.00445. arXiv.
  71. Akbir Khan, John Hughes, Dan Valentine, Laura Ruis, Kshitij Sachan, Ansh Radhakrishnan, Edward Grefenstette, Samuel R. Bowman, Tim Rocktäschel, and Ethan Perez. Debating with more persuasive llms leads to more truthful answers, 2024. URL https://arxiv.org/abs/2402.06782. ICML 2024 (Best Paper).
  72. Moo Jin Kim et al. Openvla: An open-source vision-language-action model, 2024a. URL https://arxiv.org/abs/2406.09246. arXiv / CoRL 2024.
  73. Yubin Kim et al. Mdagents: An adaptive collaboration of llms for medical decision-making, 2024b. URL https://arxiv.org/abs/2404.15155. NeurIPS 2024.
  74. Kimi Team. Kimi k2: Open agentic intelligence, 2025. URL https://arxiv.org/abs/2507.20534.
  75. Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. Large language models are zero-shot reasoners, 2022. URL https://arxiv.org/abs/2205.11916. NeurIPS 2022.
  76. Aviral Kumar, Vincent Zhuang, Rishabh Agarwal, Yi Su, John D Co-Reyes, Avi Singh, Kate Baumli, Shariq Iqbal, Colton Bishop, Rebecca Roelofs, Lei M Zhang, Kay McKinney, Disha Shrivastava, Cosmin Paduraru, George Tucker, Doina Precup, Feryal Behbahani, and Aleksandra Faust. Training language models to self-correct via reinforcement learning, 2024. URL https://arxiv.org/abs/2409.12917. ICLR 2025.
  77. Thomas Kuntz, Agatha Duzan, Hao Zhao, et al. Os-harm: A benchmark for measuring safety of computer use agents, 2025. URL https://arxiv.org/abs/2506.14866.
  78. Jungseob Lee, Midan Shim, Suhyune Son, Chanjun Park, Yujin Kim, and Heuiseok Lim. There is no rose without a thorn: Finding weaknesses on blenderbot 2.0 in terms of model, data and user-centric approach, 2022. URL https://arxiv.org/abs/2201.03239.
  79. Jungseob Lee, Hyeonseok Moon, Seungjun Lee, Chanjun Park, Sugyeong Eo, Hyunwoong Ko, Jaehyung Seo, Seungyoon Lee, and Heuiseok Lim. Length-aware byte pair encoding for mitigating over-segmentation in korean machine translation. In Findings of the Association for Computational Linguistics: ACL 2024, pages 2287–2303, 2024. URL https://aclanthology.org/2024.findings-acl.135/. [CrossRef]
  80. Jungseob Lee, Seongtae Hong, Hyeonseok Moon, and Heuiseok Lim. Cross-lingual optimization for language transfer in large language models. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (ACL), 2025. URL https://arxiv.org/abs/2505.14297.
  81. Jungseob Lee, Seongtae Hong, Seungjun Lee, Jaehyung Seo, Junyoung Son, Sugyeong Eo, Chanjun Park, Hyeongju Park, Hyeonseok Moon, and Heuiseok Lim. Dart: Draft-agreement routing for training-free adaptive thinking budgets in hybrid reasoning models, 2026a. URL https://arxiv.org/abs/2606.23181.
  82. Jungseob Lee, Chanjun Park, and Heuiseok Lim. To isolate or to score? model-adaptive assessment for cost-efficient multi-agent RAG, 2026b. URL https://arxiv.org/abs/2606.25191.
  83. Guohao Li, Hasan Abed Al Kader Hammoud, Hani Itani, Dmitrii Khizbullin, and Bernard Ghanem. Camel: Communicative agents for "mind" exploration of large language model society, 2023a. URL https://arxiv.org/abs/2303.17760. NeurIPS 2023.
  84. Junkai Li, Yunghwei Lai, Weitao Li, Jingyi Ren, Meng Zhang, Xinhui Kang, Siyu Wang, Peng Li, Ya-Qin Zhang, Weizhi Ma, et al. Agent hospital: A simulacrum of hospital with evolvable medical agents, 2024a. URL https://arxiv.org/abs/2405.02957. arXiv.
  85. Junyou Li, Qin Zhang, Yangbin Yu, Qiang Fu, and Deheng Ye. More agents is all you need, 2024b. URL https://arxiv.org/abs/2402.05120. arXiv / TMLR.
  86. Minghao Li, Yingxiu Zhao, Bowen Yu, Feifan Song, Hangyu Li, Haiyang Yu, Zhoujun Li, Fei Huang, and Yongbin Li. Api-bank: A comprehensive benchmark for tool-augmented llms, 2023b. URL https://arxiv.org/abs/2304.08244. arXiv (EMNLP 2023).
  87. Tian Liang, Zhiwei He, Wenxiang Jiao, Xing Wang, Yan Wang, Rui Wang, Yujiu Yang, Shuming Shi, and Zhaopeng Tu. Encouraging divergent thinking in large language models through multi-agent debate, 2023. URL https://arxiv.org/abs/2305.19118. EMNLP 2024 (arXiv preprint 2023).
  88. Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. Let’s verify step by step, 2023. URL https://arxiv.org/abs/2305.20050. ICLR 2024.
  89. Bang Liu, Xinfeng Li, Jiayi Zhang, et al. Advances and challenges in foundation agents: From brain-inspired intelligence to evolutionary, collaborative, and safe systems, 2025. URL https://arxiv.org/abs/2504.01990.
  90. Junwei Liu, Kaixin Wang, Yixuan Chen, Xin Peng, Zhenpeng Chen, Lingming Zhang, and Yiling Lou. Large language model-based agents for software engineering: A survey, 2024a. URL https://arxiv.org/abs/2409.02977. arXiv.
  91. Lei Liu, Xiaoyan Yang, Yue Shen, Binbin Hu, Zhiqiang Zhang, Jinjie Gu, and Guannan Zhang. Think-in-memory: Recalling and post-thinking enable llms with long-term memory, 2023a. URL https://arxiv.org/abs/2311.08719. arXiv preprint.
  92. Weiwen Liu, Xu Huang, Xingshan Zeng, Xinlong Hao, Shuai Yu, Dexun Li, Shuai Wang, et al. Toolace: Winning the points of llm function calling, 2024b. URL https://arxiv.org/abs/2409.00920. arXiv.
  93. Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, et al. Agentbench: Evaluating llms as agents, 2024c. URL https://arxiv.org/abs/2308.03688. ICLR 2024.
  94. Zijun Liu, Yanzhe Zhang, Peng Li, Yang Liu, and Diyi Yang. A dynamic llm-powered agent network for task-oriented agent collaboration, 2023b. URL https://arxiv.org/abs/2310.02170. arXiv.
  95. Chris Lu, Cong Lu, Robert Tjarko Lange, Jakob Foerster, Jeff Clune, and David Ha. The ai scientist: Towards fully automated open-ended scientific discovery, 2024. URL https://arxiv.org/abs/2408.06292. arXiv.
  96. Pan Lu, Baolin Peng, Hao Cheng, Michel Galley, Kai-Wei Chang, Ying Nian Wu, Song-Chun Zhu, and Jianfeng Gao. Chameleon: Plug-and-play compositional reasoning with large language models, 2023. URL https://arxiv.org/abs/2304.09842. NeurIPS 2023.
  97. Junyu Luo, Weizhi Zhang, Ye Yuan, Yusheng Zhao, Dacheng Tao, et al. Large language model agent: A survey on methodology, applications and challenges, 2025. URL https://arxiv.org/abs/2503.21460. arXiv.
  98. Chang Ma, Junlei Zhang, Zhihao Zhu, Cheng Yang, Yujiu Yang, Yaohui Jin, Zhenzhong Lan, Lingpeng Kong, and Junxian He. Agentboard: An analytical evaluation board of multi-turn llm agents, 2024. URL https://arxiv.org/abs/2401.13178. NeurIPS 2024.
  99. Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, et al. Self-refine: Iterative refinement with self-feedback, 2023. URL https://arxiv.org/abs/2303.17651. NeurIPS 2023.
  100. Adyasha Maharana, Dong-Ho Lee, Sergey Tulyakov, Mohit Bansal, Francesco Barbieri, and Yuwei Fang. Evaluating very long-term conversational memory of llm agents, 2024. URL https://arxiv.org/abs/2402.17753. ACL 2024.
  101. Tula Masterman, Sandi Besen, Mason Sawtell, and Alex Chao. The landscape of emerging ai agent architectures for reasoning, planning, and tool calling: A survey, 2024. URL https://arxiv.org/abs/2404.11584. arXiv.
  102. Alexander Meinke, Bronson Schoen, Jérémy Scheurer, Mikita Balesni, Rusheb Shah, and Marius Hobbhahn. Frontier models are capable of in-context scheming, 2024. URL https://arxiv.org/abs/2412.04984. arXiv (Apollo Research).
  103. Mike A. Merrill, Alexander G. Shaw, Nicholas Carlini, et al. Terminal-bench: Benchmarking agents on hard, realistic tasks in command line interfaces, 2026. URL https://arxiv.org/abs/2601.11868.
  104. Grégoire Mialon, Clémentine Fourrier, Craig Swift, Thomas Wolf, Yann LeCun, and Thomas Scialom. Gaia: a benchmark for general ai assistants, 2023. URL https://arxiv.org/abs/2311.12983. ICLR 2024.
  105. Samuel Miserendino, Michele Wang, Tejal Patwardhan, and Johannes Heidecke. Swe-lancer: Can frontier llms earn $1 million from real-world freelance software engineering?, 2025. URL https://arxiv.org/abs/2502.12115.
  106. Margaret Mitchell, Avijit Ghosh, Alexandra Sasha Luccioni, and Giada Pistilli. Fully autonomous ai agents should not be developed, 2025. URL https://arxiv.org/abs/2502.02649. arXiv.
  107. Ludovico Mitchener, Angela Yiu, Benjamin Chang, et al. Kosmos: An ai scientist for autonomous discovery, 2025. URL https://arxiv.org/abs/2511.02824. arXiv.
  108. Ali Modarressi, Ayyoob Imani, Mohsen Fayyaz, and Hinrich Schütze. Ret-llm: Towards a general read-write memory for large language models, 2023. URL https://arxiv.org/abs/2305.14322. arXiv preprint.
  109. Mahmoud Mohammadi, Yipeng Li, Jane Lo, and Wendy Yip. Evaluation and benchmarking of llm agents: A survey, 2025. URL https://arxiv.org/abs/2507.21504. KDD 2025.
  110. Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, et al. Webgpt: Browser-assisted question-answering with human feedback, 2021. URL https://arxiv.org/abs/2112.09332. arXiv (OpenAI).
  111. Dang Nguyen et al. Gui agents: A survey, 2024. URL https://arxiv.org/abs/2412.13501. Findings of ACL 2025.
  112. Liangbo Ning, Ziran Liang, Zhuohang Jiang, Haohao Qu, Yujuan Ding, Wenqi Fan, Xiao yong Wei, Shanru Lin, Hui Liu, Philip S. Yu, and Qing Li. A survey of webagents: Towards next-generation ai agents for web automation with large foundation models, 2025. URL https://arxiv.org/abs/2503.23350. KDD 2025.
  113. Alexander Novikov, Ngan Vu, Marvin Eisenberger, Emilien Dupont, Po-Sen Huang, et al. Alphaevolve: A coding agent for scientific and algorithmic discovery, 2025. URL https://arxiv.org/abs/2506.13131. arXiv.
  114. OpenAI. Introducing operator, 2025. URL https://openai.com/index/introducing-operator/.
  115. Siru Ouyang, Jun Yan, I-Hung Hsu, et al. Reasoningbank: Scaling agent self-evolving with reasoning memory, 2025. URL https://arxiv.org/abs/2509.25140.
  116. Charles Packer, Sarah Wooders, Kevin Lin, Vivian Fang, Shishir G. Patil, Ion Stoica, and Joseph E. Gonzalez. Memgpt: Towards llms as operating systems, 2023. URL https://arxiv.org/abs/2310.08560. COLM 2024.
  117. Aaron Parisi, Yao Zhao, and Noah Fiedel. Talm: Tool augmented language models, 2022. URL https://arxiv.org/abs/2205.12255. arXiv.
  118. Joon Sung Park, Joseph C. O’Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. Generative agents: Interactive simulacra of human behavior, 2023a. URL https://arxiv.org/abs/2304.03442. arXiv preprint; UIST 2023.
  119. Peter S. Park, Simon Goldstein, Aidan O’Gara, Michael Chen, and Dan Hendrycks. Ai deception: A survey of examples, risks, and potential solutions, 2023b. URL https://arxiv.org/abs/2308.14752. Patterns (Cell Press) / arXiv.
  120. Shishir G. Patil, Tianjun Zhang, Xin Wang, and Joseph E. Gonzalez. Gorilla: Large language model connected with massive apis, 2024. URL https://arxiv.org/abs/2305.15334. NeurIPS 2024.
  121. Shishir G. Patil, Huanzhi Mao, Fanjia Yan, Charlie Cheng-Jie Ji, Vishnu Suresh, Ion Stoica, and Joseph E. Gonzalez. The berkeley function calling leaderboard (BFCL): From tool use to agentic evaluation of large language models. In Proceedings of the 42nd International Conference on Machine Learning, volume 267 of Proceedings of Machine Learning Research, pages 48371–48392. PMLR, 2025. URL https://proceedings.mlr.press/v267/patil25a.html.
  122. Ethan Perez, Sam Ringer, Kamilė Lukošiūtė, et al. Discovering language model behaviors with model-written evaluations, 2022. URL https://arxiv.org/abs/2212.09251. arXiv (Anthropic).
  123. Aske Plaat, Max van Duijn, Niki van Stein, Mike Preuss, Peter van der Putten, and Kees Joost Batenburg. Agentic large language models, a survey, 2025. URL https://arxiv.org/abs/2503.23037. arXiv.
  124. Archiki Prasad, Alexander Koller, Mareike Hartmann, Peter Clark, Ashish Sabharwal, Mohit Bansal, and Tushar Khot. Adapt: As-needed decomposition and planning with language models, 2023. URL https://arxiv.org/abs/2311.05772. NAACL 2024 Findings.
  125. Zehan Qi, Xiao Liu, Iat Long Iong, Hanyu Lai, Xueqiao Sun, Wenyi Zhao, Yu Yang, Xinyue Yang, Jiadai Sun, Shuntian Yao, Tianjie Zhang, Wei Xu, Jie Tang, and Yuxiao Dong. Webrl: Training llm web agents via self-evolving online curriculum reinforcement learning, 2024. URL https://arxiv.org/abs/2411.02337.
  126. Chen Qian, Wei Liu, Hongzhang Liu, Nuo Chen, Yufan Dang, Jiahao Li, Cheng Yang, Weize Chen, Yusheng Su, Xin Cong, Juyuan Xu, Dahai Li, Zhiyuan Liu, and Maosong Sun. Chatdev: Communicative agents for software development, 2023. URL https://arxiv.org/abs/2307.07924. ACL 2024 (arXiv preprint 2023).
  127. Chen Qian, Zihao Xie, YiFei Wang, et al. Scaling large language model-based multi-agent collaboration, 2024. URL https://arxiv.org/abs/2406.07155.
  128. Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, et al. Toolllm: Facilitating large language models to master 16000+ real-world apis, 2023. URL https://arxiv.org/abs/2307.16789. arXiv (ICLR 2024).
  129. Yujia Qin et al. Ui-tars: Pioneering automated gui interaction with native agents, 2025. URL https://arxiv.org/abs/2501.12326. arXiv.
  130. Changle Qu, Sunhao Dai, Xiaochi Wei, Hengyi Cai, Shuaiqiang Wang, Dawei Yin, Jun Xu, and Ji-Rong Wen. Tool learning with large language models: A survey, 2024. URL https://arxiv.org/abs/2405.17935. arXiv; Frontiers of Computer Science 2024. [CrossRef]
  131. Christopher Rawles et al. Androidworld: A dynamic benchmarking environment for autonomous agents, 2024. URL https://arxiv.org/abs/2405.14573. arXiv (ICLR 2025 submission).
  132. Shuo Ren, Can Xie, Pu Jian, Zhenjiang Ren, Chunlin Leng, and Jiajun Zhang. Towards scientific intelligence: A survey of llm-based scientific agents, 2025. URL https://arxiv.org/abs/2503.24047. arXiv.
  133. Yangjun Ruan, Honghua Dong, Andrew Wang, Silviu Pitis, Yongchao Zhou, Jimmy Ba, Yann Dubois, Chris J. Maddison, and Tatsunori Hashimoto. Identifying the risks of lm agents with an lm-emulated sandbox, 2024. URL https://arxiv.org/abs/2309.15817. ICLR 2024 (Spotlight).
  134. Ranjan Sapkota, Konstantinos I. Roumeliotis, and Manoj Karkee. Ai agents vs. agentic ai: A conceptual taxonomy, applications and challenges, 2025. URL https://arxiv.org/abs/2505.10468. arXiv.
  135. Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. Toolformer: Language models can teach themselves to use tools, 2023. URL https://arxiv.org/abs/2302.04761. arXiv preprint; NeurIPS 2023.
  136. Samuel Schmidgall, Rojin Ziaei, Carl Harris, et al. Agentclinic: A multimodal agent benchmark to evaluate ai in simulated clinical environments, 2024. URL https://arxiv.org/abs/2405.07960.
  137. Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, et al. Towards understanding sycophancy in language models, 2023. URL https://arxiv.org/abs/2310.13548. arXiv (Anthropic).
  138. Weizhou Shen, Chenliang Li, Hongzhan Chen, Ming Yan, Xiaojun Quan, et al. Small llms are weak tool learners: A multi-llm agent, 2024. URL https://arxiv.org/abs/2401.07324. EMNLP 2024.
  139. Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang. Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face, 2023. URL https://arxiv.org/abs/2303.17580. arXiv preprint; NeurIPS 2023.
  140. Noah Shinn, Federico Cassano, Edward Berman, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. Reflexion: Language agents with verbal reinforcement learning, 2023. URL https://arxiv.org/abs/2303.11366. arXiv preprint; NeurIPS 2023.
  141. Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk, Adam Trischler, and Matthew Hausknecht. Alfworld: Aligning text and embodied environments for interactive learning, 2021. URL https://arxiv.org/abs/2010.03768. ICLR 2021.
  142. Andries Smit, Paul Duckworth, Nathan Grinsztajn, Thomas D. Barrett, and Arnu Pretorius. Should we be going mad? a look at multi-agent debate strategies for llms, 2023. URL https://arxiv.org/abs/2311.17371. arXiv preprint (ICML 2024 workshop).
  143. Yifan Song, Weimin Xiong, Dawei Zhu, Wenhao Wu, Han Qian, et al. Restgpt: Connecting large language models with real-world restful apis, 2023. URL https://arxiv.org/abs/2306.06624. arXiv.
  144. Leon Staufer, Kevin Feng, Kevin Wei, et al. The 2025 ai agent index: Documenting technical and safety features of deployed agentic ai systems, 2026. URL https://arxiv.org/abs/2602.17753. ACM FAccT 2026 (arXiv 2026).
  145. Hang Su, Jun Luo, Chang Liu, Xiao Yang, Yichi Zhang, Yinpeng Dong, and Jun Zhu. A survey on autonomy-induced security risks in large model-based agents, 2025. URL https://arxiv.org/abs/2506.23844. arXiv.
  146. Theodore R. Sumers, Shunyu Yao, Karthik Narasimhan, and Thomas L. Griffiths. Cognitive architectures for language agents, 2023. URL https://arxiv.org/abs/2309.02427. arXiv; TMLR (camera-ready v3).
  147. Qiaoyu Tang, Ziliang Deng, Hongyu Lin, Xianpei Han, Qiao Liang, Boxi Cao, and Le Sun. Toolalpaca: Generalized tool learning for language models with 3000 simulated cases, 2023. URL https://arxiv.org/abs/2306.05301. arXiv.
  148. Yu Tian, Xiao Yang, Jingyuan Zhang, Yinpeng Dong, and Hang Su. Evil geniuses: Delving into the safety of llm-based agents, 2023. URL https://arxiv.org/abs/2311.11855. arXiv.
  149. Tongyi DeepResearch Team. Tongyi deepresearch technical report, 2025. URL https://arxiv.org/abs/2510.24701.
  150. Khanh-Tung Tran, Dung Dao, Minh-Duong Nguyen, Quoc-Viet Pham, Barry O’Sullivan, and Hoang D. Nguyen. Multi-agent collaboration mechanisms: A survey of llms, 2025. URL https://arxiv.org/abs/2501.06322. arXiv.
  151. Tao Tu et al. Towards conversational diagnostic ai, 2024. URL https://arxiv.org/abs/2401.05654. arXiv.
  152. Eric Wallace, Kai Xiao, Reimar Leike, Lilian Weng, Johannes Heidecke, and Alex Beutel. The instruction hierarchy: Training llms to prioritize privileged instructions, 2024. URL https://arxiv.org/abs/2404.13208.
  153. Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar. Voyager: An open-ended embodied agent with large language models, 2023a. URL https://arxiv.org/abs/2305.16291. arXiv preprint; TMLR.
  154. Haoming Wang, Haoyang Zou, Huatong Song, et al. Ui-tars-2 technical report: Advancing gui agent with multi-turn reinforcement learning, 2025a. URL https://arxiv.org/abs/2509.02544.
  155. Jize Wang, Zerun Ma, Yining Li, Songyang Zhang, Cailian Chen, Kai Chen, and Xinyi Le. Gta: A benchmark for general tool agents, 2024a. URL https://arxiv.org/abs/2407.08713. NeurIPS 2024.
  156. Junlin Wang, Jue Wang, Ben Athiwaratkun, Ce Zhang, and James Zou. Mixture-of-agents enhances large language model capabilities, 2024b. URL https://arxiv.org/abs/2406.04692. arXiv preprint (ICLR 2025).
  157. Junyang Wang, Haiyang Xu, Jiabo Ye, Ming Yan, Weizhou Shen, Ji Zhang, Fei Huang, and Jitao Sang. Mobile-agent: Autonomous multi-modal mobile device agent with visual perception, 2024c. URL https://arxiv.org/abs/2401.16158. arXiv.
  158. Kun Wang et al. A comprehensive survey in llm(-agent) full stack safety: Data, training and deployment, 2025b. URL https://arxiv.org/abs/2504.15585. arXiv.
  159. Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, Wayne Xin Zhao, Zhewei Wei, and Ji-Rong Wen. A survey on large language model based autonomous agents, 2023b. URL https://arxiv.org/abs/2308.11432. arXiv preprint; also published in Frontiers of Computer Science 2024. [CrossRef]
  160. Lei Wang, Wanyu Xu, Yihuai Lan, Zhiqiang Hu, Yunshi Lan, Roy Ka-Wei Lee, and Ee-Peng Lim. Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models, 2023c. URL https://arxiv.org/abs/2305.04091. ACL 2023.
  161. Weizhi Wang, Li Dong, Hao Cheng, Xiaodong Liu, Xifeng Yan, Jianfeng Gao, and Furu Wei. Augmenting language models with long-term memory, 2023d. URL https://arxiv.org/abs/2306.07174. NeurIPS 2023.
  162. Wenxuan Wang, Zizhan Ma, Zheng Wang, Chenghan Wu, Jiaming Ji, Wenting Chen, Xiang Li, and Yixuan Yuan. A survey of llm-based agents in medicine: How far are we from baymax?, 2025c. URL https://arxiv.org/abs/2502.11211. ACL 2025 Findings.
  163. Xingyao Wang, Yangyi Chen, Lifan Yuan, Yizhe Zhang, Yunzhu Li, Hao Peng, and Heng Ji. Executable code actions elicit better llm agents, 2024d. URL https://arxiv.org/abs/2402.01030. ICML 2024.
  164. Xingyao Wang et al. Openhands: An open platform for ai software developers as generalist agents, 2024e. URL https://arxiv.org/abs/2407.16741. ICLR 2025 / arXiv.
  165. Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. Self-consistency improves chain of thought reasoning in language models, 2022. URL https://arxiv.org/abs/2203.11171. ICLR 2023.
  166. Yifei Wang, Dizhan Xue, Shengjie Zhang, and Shengsheng Qian. Badagent: Inserting and activating backdoor attacks in llm agents, 2024f. URL https://arxiv.org/abs/2406.03007. ACL 2024.
  167. Yu Wang and Xi Chen. Mirix: Multi-agent memory system for llm-based agents, 2025. URL https://arxiv.org/abs/2507.07957. arXiv preprint.
  168. Zhenhailong Wang, Shaoguang Mao, Wenshan Wu, Tao Ge, Furu Wei, and Heng Ji. Unleashing the emergent cognitive synergy in large language models: A task-solving agent through multi-persona self-collaboration, 2023e. URL https://arxiv.org/abs/2307.05300. NAACL 2024 (arXiv preprint 2023).
  169. Zhiruo Wang, Zhoujun Cheng, Hao Zhu, Daniel Fried, and Graham Neubig. What are tools anyway? a survey from the language model perspective, 2024g. URL https://arxiv.org/abs/2403.15452. COLM 2024.
  170. Zihao Wang, Shaofei Cai, Guanzhou Chen, Anji Liu, Xiaojian Ma, and Yitao Liang. Describe, explain, plan and select: Interactive planning with large language models enables open-world multi-task agents, 2023f. URL https://arxiv.org/abs/2302.01560. NeurIPS 2023.
  171. Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. Chain-of-thought prompting elicits reasoning in large language models, 2022. URL https://arxiv.org/abs/2201.11903. NeurIPS 2022.
  172. Di Wu, Hongwei Wang, Wenhao Yu, et al. Longmemeval: Benchmarking chat assistants on long-term interactive memory, 2024a. URL https://arxiv.org/abs/2410.10813. ICLR 2025 (arXiv 2024).
  173. Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Beibin Li, Erkang Zhu, Li Jiang, Xiaoyun Zhang, Shaokun Zhang, Jiale Liu, Ahmed Hassan Awadallah, Ryen W. White, Doug Burger, and Chi Wang. Autogen: Enabling next-gen llm applications via multi-agent conversation, 2023a. URL https://arxiv.org/abs/2308.08155. arXiv.
  174. Yaxiong Wu, Sheng Liang, Chen Zhang, Yichao Wang, Yongyue Zhang, Huifeng Guo, Ruiming Tang, and Yong Liu. From human memory to ai memory: A survey on memory mechanisms in the era of llms, 2025. URL https://arxiv.org/abs/2504.15965. arXiv preprint.
  175. Yue Wu, Xuan Tang, Tom M. Mitchell, and Yuanzhi Li. Smartplay: A benchmark for llms as intelligent agents, 2023b. URL https://arxiv.org/abs/2310.01557. ICLR 2024.
  176. Zhiyong Wu, Chengcheng Han, Zichen Ding, Zhenmin Weng, Zhoumianze Liu, Shunyu Yao, Tao Yu, and Lingpeng Kong. Os-copilot: Towards generalist computer agents with self-improvement, 2024b. URL https://arxiv.org/abs/2402.07456. arXiv.
  177. Zhiheng Xi, Wenxiang Chen, Xin Guo, Wei He, Yiwen Ding, et al. The rise and potential of large language model based agents: A survey, 2023. URL https://arxiv.org/abs/2309.07864. arXiv.
  178. Zhiheng Xi, Yiwen Ding, Wenxiang Chen, Boyang Hong, Honglin Guo, Junzhe Wang, Dingwen Yang, Chenyang Liao, Xin Guo, Wei He, Songyang Gao, Lu Chen, Rui Zheng, Yicheng Zou, Tao Gui, Qi Zhang, Xipeng Qiu, Xuanjing Huang, Zuxuan Wu, and Yu-Gang Jiang. Agentgym: Evolving large language model-based agents across diverse environments, 2024. URL https://arxiv.org/abs/2406.04151.
  179. Chunqiu Steven Xia, Yinlin Deng, Soren Dunn, and Lingming Zhang. Agentless: Demystifying llm-based software engineering agents, 2024. URL https://arxiv.org/abs/2407.01489. FSE 2025 / arXiv.
  180. Yijia Xiao, Edward Sun, Di Luo, and Wei Wang. Tradingagents: Multi-agents llm financial trading framework, 2024. URL https://arxiv.org/abs/2412.20138. arXiv.
  181. Jian Xie, Kai Zhang, Jiangjie Chen, Tinghui Zhu, Renze Lou, Yuandong Tian, Yanghua Xiao, and Yu Su. Travelplanner: A benchmark for real-world planning with language agents, 2024a. URL https://arxiv.org/abs/2402.01622. ICML 2024.
  182. Tianbao Xie, Danyang Zhang, Jixuan Chen, et al. Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments, 2024b. URL https://arxiv.org/abs/2404.07972. NeurIPS 2024.
  183. Binfeng Xu, Zhiyuan Peng, Bowen Lei, Subhabrata Mukherjee, Yuchen Liu, and Dongkuan Xu. Rewoo: Decoupling reasoning from observations for efficient augmented language models, 2023. URL https://arxiv.org/abs/2305.18323. arXiv.
  184. Wujiang Xu, Zujie Liang, Kai Mei, Hang Gao, Juntao Tan, and Yongfeng Zhang. A-mem: Agentic memory for llm agents, 2025. URL https://arxiv.org/abs/2502.12110. arXiv preprint.
  185. Yutaro Yamada, Robert Tjarko Lange, Cong Lu, Shengran Hu, Chris Lu, Jakob Foerster, Jeff Clune, and David Ha. The ai scientist-v2: Workshop-level automated scientific discovery via agentic tree search, 2025. URL https://arxiv.org/abs/2504.08066. arXiv.
  186. Bingyu Yan, Zhibo Zhou, Litian Zhang, Lian Zhang, Ziyi Zhou, Dezhuang Miao, Zhoujun Li, Chaozhuo Li, and Xiaoming Zhang. Beyond self-talk: A communication-centric survey of llm-based multi-agent systems, 2025. URL https://arxiv.org/abs/2502.14321. arXiv (Frontiers of Computer Science).
  187. Hongyang Yang, Xiao-Yang Liu, and Christina Dan Wang. Fingpt: Open-source financial large language models, 2023a. URL https://arxiv.org/abs/2306.06031. IJCAI 2023 FinLLM Workshop / arXiv.
  188. Hui Yang, Sifu Yue, and Yunzhong He. Auto-gpt for online decision making: Benchmarks and additional opinions, 2023b. URL https://arxiv.org/abs/2306.02224. arXiv.
  189. John Yang, Akshara Prabhakar, Karthik Narasimhan, and Shunyu Yao. Intercode: Standardizing and benchmarking interactive coding with execution feedback, 2023c. URL https://arxiv.org/abs/2306.14898. NeurIPS 2023.
  190. John Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press. Swe-agent: Agent-computer interfaces enable automated software engineering, 2024a. URL https://arxiv.org/abs/2405.15793. NeurIPS 2024.
  191. Ling Yang, Zhaochen Yu, Tianjun Zhang, Shiyi Cao, Minkai Xu, Wentao Zhang, Joseph E. Gonzalez, and Bin Cui. Buffer of thoughts: Thought-augmented reasoning with large language models, 2024b. URL https://arxiv.org/abs/2406.04271. NeurIPS 2024.
  192. Rui Yang, Lin Song, Yanwei Li, Sijie Zhao, Yixiao Ge, Xiu Li, and Ying Shan. Gpt4tools: Teaching large language model to use tools via self-instruction, 2023d. URL https://arxiv.org/abs/2305.18752. arXiv (NeurIPS 2023).
  193. Yingxuan Yang, Huacan Chai, Yuanyi Song, et al. A survey of ai agent protocols, 2025. URL https://arxiv.org/abs/2504.16736.
  194. Shunyu Yao, Howard Chen, John Yang, and Karthik Narasimhan. Webshop: Towards scalable real-world web interaction with grounded language agents, 2022a. URL https://arxiv.org/abs/2207.01206. NeurIPS 2022.
  195. Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. React: Synergizing reasoning and acting in language models, 2022b. URL https://arxiv.org/abs/2210.03629. arXiv preprint; ICLR 2023 (camera-ready version noted on arXiv page).
  196. Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik Narasimhan. Tree of thoughts: Deliberate problem solving with large language models, 2023. URL https://arxiv.org/abs/2305.10601. NeurIPS 2023.
  197. Shunyu Yao, Noah Shinn, Pedram Razavi, and Karthik Narasimhan. τ-bench: A benchmark for tool-agent-user interaction in real-world domains, 2024. URL https://arxiv.org/abs/2406.12045. arXiv.
  198. Asaf Yehudai, Lilach Eden, Alan Li, Guy Uziel, Yilun Zhao, Roy Bar-Haim, Arman Cohan, and Michal Shmueli-Scheuer. Survey on evaluation of llm-based agents, 2025. URL https://arxiv.org/abs/2503.16416. arXiv.
  199. Sheng Yin, Xianghe Pang, Yuanzhuo Ding, Menglan Chen, Yutong Bi, Yichen Xiong, Wenhao Huang, Zhen Xiang, Jing Shao, and Siheng Chen. Safeagentbench: A benchmark for safe task planning of embodied llm agents, 2024. URL https://arxiv.org/abs/2412.13178. arXiv.
  200. Zhangyue Yin, Qiushi Sun, Cheng Chang, Qipeng Guo, Junqi Dai, Xuanjing Huang, and Xipeng Qiu. Exchange-of-thought: Enhancing large language model capabilities through cross-model communication, 2023. URL https://arxiv.org/abs/2312.01823. EMNLP 2023.
  201. Miao Yu, Fanci Meng, Xinyun Zhou, Shilong Wang, Junyuan Mao, Linsey Pang, Tianlong Chen, Kun Wang, Xinfeng Li, Yongfeng Zhang, Bo An, and Qingsong Wen. A survey on trustworthy llm agents: Threats and countermeasures, 2025. URL https://arxiv.org/abs/2503.09648. arXiv.
  202. Yangyang Yu, Haohang Li, Zhi Chen, Yuechen Jiang, Yang Li, Denghui Zhang, Rong Liu, Jordan W. Suchow, and Khaldoun Khashanah. Finmem: A performance-enhanced llm trading agent with layered memory and character design, 2023. URL https://arxiv.org/abs/2311.13743. AAAI Symposium Series 2024 / arXiv.
  203. Tongxin Yuan, Zhiwei He, Lingzhong Dong, Yiming Wang, Ruijie Zhao, Tian Xia, Lizhen Xu, Binglin Zhou, Fangqi Li, Zhuosheng Zhang, Rui Wang, and Gongshen Liu. R-judge: Benchmarking safety risk awareness for llm agents, 2024. URL https://arxiv.org/abs/2401.10019. EMNLP 2024 Findings.
  204. Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah D. Goodman. Star: Bootstrapping reasoning with reasoning, 2022. URL https://arxiv.org/abs/2203.14465. NeurIPS 2022.
  205. Aohan Zeng, Mingdao Liu, Rui Lu, Bowen Wang, Xiao Liu, Yuxiao Dong, and Jie Tang. Agenttuning: Enabling generalized agent abilities for llms, 2023. URL https://arxiv.org/abs/2310.12823.
  206. Ruihong Zeng, Jinyuan Fang, Siwei Liu, and Zaiqiao Meng. On the structural memory of llm agents, 2024. URL https://arxiv.org/abs/2412.15266. arXiv preprint.
  207. Qiusi Zhan, Zhixiang Liang, Zifan Ying, and Daniel Kang. Injecagent: Benchmarking indirect prompt injections in tool-integrated large language model agents, 2024. URL https://arxiv.org/abs/2403.02691. ACL 2024 Findings.
  208. Chaoyun Zhang, Shilin He, Jiaxu Qian, Bowen Li, Liqun Li, Si Qin, Yu Kang, Minghua Ma, Guyue Liu, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang, and Qi Zhang. Large language model-brained gui agents: A survey, 2024a. URL https://arxiv.org/abs/2411.18279. arXiv.
  209. Chi Zhang, Zhao Yang, Jiaxuan Liu, Yucheng Han, Xin Chen, Zebiao Huang, Bin Fu, and Gang Yu. Appagent: Multimodal agents as smartphone users, 2023a. URL https://arxiv.org/abs/2312.13771. arXiv (later ACM CHI 2025).
  210. Guibin Zhang, Hejia Geng, Xiaohang Yu, et al. The landscape of agentic reinforcement learning for llms: A survey, 2025. URL https://arxiv.org/abs/2509.02547.
  211. Hanrong Zhang, Jingyuan Huang, Kai Mei, Yifei Yao, Zhenting Wang, Chenlu Zhan, Hongwei Wang, and Yongfeng Zhang. Agent security bench (asb): Formalizing and benchmarking attacks and defenses in llm-based agents, 2024b. URL https://arxiv.org/abs/2410.02644. arXiv.
  212. Hongxin Zhang, Weihua Du, Jiaming Shan, Qinhong Zhou, Yilun Du, Joshua B. Tenenbaum, Tianmin Shu, and Chuang Gan. Building cooperative embodied agents modularly with large language models, 2023b. URL https://arxiv.org/abs/2307.02485. ICLR 2024 (arXiv preprint 2023).
  213. Jianguo Zhang, Tian Lan, Ming Zhu, Zuxin Liu, Thai Hoang, et al. xlam: A family of large action models to empower ai agent systems, 2024c. URL https://arxiv.org/abs/2409.03215. arXiv.
  214. Jiayi Zhang, Jinyu Xiang, Zhaoyang Yu, et al. Aflow: Automating agentic workflow generation, 2024d. URL https://arxiv.org/abs/2410.10762. ICLR 2025 (arXiv 2024).
  215. Yuntong Zhang, Haifeng Ruan, Zhiyu Fan, and Abhik Roychoudhury. Autocoderover: Autonomous program improvement, 2024e. URL https://arxiv.org/abs/2404.05427. ISSTA 2024 / arXiv.
  216. Zeyu Zhang, Xiaohe Bo, Chen Ma, Rui Li, Xu Chen, Quanyu Dai, Jieming Zhu, Zhenhua Dong, and Ji-Rong Wen. A survey on the memory mechanism of large language model based agents, 2024f. URL https://arxiv.org/abs/2404.13501. arXiv.
  217. Zhuosheng Zhang, Yao Yao, Aston Zhang, Xiangru Tang, Xinbei Ma, Zhiwei He, Yiming Wang, Mark Gerstein, Rui Wang, Gongshen Liu, and Hai Zhao. Igniting language intelligence: The hitchhiker’s guide from chain-of-thought reasoning to language agents, 2023c. URL https://arxiv.org/abs/2311.11797. arXiv (ACM Computing Surveys, as claimed).
  218. Andrew Zhao, Daniel Huang, Quentin Xu, Matthieu Lin, Yong-Jin Liu, and Gao Huang. Expel: Llm agents are experiential learners, 2023. URL https://arxiv.org/abs/2308.10144. AAAI 2024.
  219. Boyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun, and Yu Su. Gpt-4v(ision) is a generalist web agent, if grounded, 2024. URL https://arxiv.org/abs/2401.01614. ICML 2024.
  220. Longtao Zheng, Rundong Wang, Xinrun Wang, and Bo An. Synapse: Trajectory-as-exemplar prompting with memory for computer control, 2023. URL https://arxiv.org/abs/2306.07863. ICLR 2024.
  221. Wanjun Zhong, Lianghong Guo, Qiqi Gao, He Ye, and Yanlin Wang. Memorybank: Enhancing large language models with long-term memory, 2023. URL https://arxiv.org/abs/2305.10250. AAAI 2024.
  222. Andy Zhou, Kai Yan, Michal Shlapentokh-Rothman, Haohan Wang, and Yu-Xiong Wang. Language agent tree search unifies reasoning, acting, and planning in language models, 2024a. URL https://arxiv.org/abs/2310.04406. ICML 2024.
  223. Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc Le, and Ed Chi. Least-to-most prompting enables complex reasoning in large language models, 2022. URL https://arxiv.org/abs/2205.10625. ICLR 2023.
  224. Pei Zhou, Jay Pujara, Xiang Ren, Xinyun Chen, Heng-Tze Cheng, Quoc V. Le, Ed H. Chi, Denny Zhou, Swaroop Mishra, and Huaixiu Steven Zheng. Self-discover: Large language models self-compose reasoning structures, 2024b. URL https://arxiv.org/abs/2402.03620. NeurIPS 2024.
  225. Shuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig. Webarena: A realistic web environment for building autonomous agents, 2024c. URL https://arxiv.org/abs/2307.13854. ICLR 2024.
  226. Mingchen Zhuge, Wenyi Wang, Louis Kirsch, et al. Language agents as optimizable graphs, 2024a. URL https://arxiv.org/abs/2402.16823. ICML 2024 (arXiv 2024).
  227. Mingchen Zhuge, Changsheng Zhao, Dylan Ashley, et al. Agent-as-a-judge: Evaluate agents with agents, 2024b. URL https://arxiv.org/abs/2410.10934.
  228. Henry Peng Zou, Wei-Chieh Huang, Yaozu Wu, Jizhou Guo, Yankai Chen, et al. Llm-based human-agent collaboration and interaction systems: A survey, 2025. URL https://arxiv.org/abs/2505.00753. arXiv.
1
Agent. A system that perceives an environment, maintains internal state, and selects actions to pursue a goal, typically in a closed loop (Sumers et al. 2023); formalized in Section 2.2.
2
Agentic AI. An umbrella term, popularized in 2024–2025, for systems exhibiting autonomous multi-step goal pursuit; several authors use it to distinguish orchestrated multi-component systems from single tool-augmented models (Sapkota et al. 2025). We treat it as a synonym for the higher-autonomy region of Definition 1.
3
Tree and graph search multiply token consumption by the branching factor and depth; the original papers report the accuracy gains but comparisons across methods rarely fix a compute budget, a gap Section 9.3 returns to. The underlying subword tokenizer further sets the per-step token budget (Lee et al. 2024).
4
World model. An internal predictor of how the environment (or the reasoning state) will change in response to an action; classical planning searches against such a model. In LLM agents the model is often the LLM itself, prompted to predict the next state (Hao et al. 2023a).
5
Agent–computer interface (ACI). The commands and feedback formats through which an agent perceives and acts on a computer, the agent-facing analogue of a human user interface. SWE-agent shows that ACI design changes capability as much as the model does (Yang et al. 2024a).
6
Grounding. Mapping a referring expression or intended action onto the concrete interface element (a pixel region, a DOM node, an object) that realizes it. Grounding, not high-level planning, is the dominant failure mode of interactive agents; SeeClick introduces the ScreenSpot benchmark to measure it directly (Cheng et al. 2024).
7
Compare with the tool-use mechanisms of Section 6: the same interfaces that let an agent act on the world are what an adversary can target.
Figure 1. Growth of the LLM-agent literature. Monthly counts of new arXiv submissions in five agent subtopics (left axis) and the cumulative number of distinct papers (shaded, right axis), from January 2023 – when the literature begins in earnest, through mid-2026. Papers are matched by title and abstract keyword search and deduplicated by identifier for the cumulative count; keyword matching has high precision but imperfect recall, so the counts are conservative lower bounds.
Figure 1. Growth of the LLM-agent literature. Monthly counts of new arXiv submissions in five agent subtopics (left axis) and the cumulative number of distinct papers (shaded, right axis), from January 2023 – when the literature begins in earnest, through mid-2026. Papers are matched by title and abstract keyword search and deduplicated by identifier for the cumulative count; keyword matching has high precision but imperfect recall, so the counts are conservative lower bounds.
Preprints 223549 g001
Figure 3. Anatomy of the LLM-agent loop. The LLM core maps observations and retrieved memory to actions; external actions act on an environment and its tools (solid), messages pass to and from other agents (dashed), and memory is read and written each step. Each block is the subject of one component pillar of this survey (Section 4, Section 5, Section 6 and Section 7), and Figure 9 annotates the same diagram with attack entry points.
Figure 3. Anatomy of the LLM-agent loop. The LLM core maps observations and retrieved memory to actions; external actions act on an environment and its tools (solid), messages pass to and from other agents (dashed), and memory is read and written each step. Each block is the subject of one component pillar of this survey (Section 4, Section 5, Section 6 and Section 7), and Figure 9 annotates the same diagram with attack entry points.
Preprints 223549 g003
Figure 5. Four planning paradigms. (a) A linear chain of thoughts (Section 4.1); (b) a tree with self-evaluation and backtracking (Tree of Thoughts); (c) a graph that merges reasoning paths (Graph of Thoughts); (d) an interleaved loop that alternates thought, action, and observation (ReAct-style, Section 4.3). Expressiveness rises left to right, and so does inference cost.
Figure 5. Four planning paradigms. (a) A linear chain of thoughts (Section 4.1); (b) a tree with self-evaluation and backtracking (Tree of Thoughts); (c) a graph that merges reasoning paths (Graph of Thoughts); (d) an interleaved loop that alternates thought, action, and observation (ReAct-style, Section 4.3). Expressiveness rises left to right, and so does inference cost.
Preprints 223549 g005
Figure 6. Cost and accuracy of three agent control loops sharing a single model (Qwen2.5-7B-Instruct) on 50 ALFWorld tasks. Plan-then-execute (ReWOO) is the cheapest but least successful; interleaving reasoning with acting (ReAct) and adding one reflection retry (Reflexion) raise the task success rate (a) at a steep cost in generated tokens (b). Success and cost rise together, so neither reading alone is complete.
Figure 6. Cost and accuracy of three agent control loops sharing a single model (Qwen2.5-7B-Instruct) on 50 ALFWorld tasks. Plan-then-execute (ReWOO) is the cheapest but least successful; interleaving reasoning with acting (ReAct) and adding one reflection retry (Reflexion) raise the task success rate (a) at a steep cost in generated tokens (b). Success and cost rise together, so neither reading alone is complete.
Preprints 223549 g006
Figure 7. The design space of agent memory. Long-term stores are read by a retrieval policy (solid) and written by a consolidation policy (dashed), with an optional forgetting policy that decays stale entries. Systems differ mainly in the store form and in these three policies (Table 5), not in raw capacity.
Figure 7. The design space of agent memory. Long-term stores are read by a retrieval policy (solid) and written by a consolidation policy (dashed), with an optional forgetting policy that decays stale entries. Systems differ mainly in the store form and in these three policies (Table 5), not in raw capacity.
Preprints 223549 g007
Figure 9. The agent loop of Figure 3 as an attack surface. Entry points: ❶ direct prompt injection through the user instruction; ❷ indirect prompt injection through tool or environment content returned as an observation; ❸ poisoning of persistent memory or a retrieval knowledge base; and ❸ contagion from another agent. Numbers map to the paragraphs of Section 11.2; the tool supply chain (a compromised tool itself) is a further surface we treat as an open problem in Section 12.
Figure 9. The agent loop of Figure 3 as an attack surface. Entry points: ❶ direct prompt injection through the user instruction; ❷ indirect prompt injection through tool or environment content returned as an observation; ❸ poisoning of persistent memory or a retrieval knowledge base; and ❸ contagion from another agent. Numbers map to the paragraphs of Section 11.2; the tool supply chain (a compromised tool itself) is a further surface we treat as an open problem in Section 12.
Preprints 223549 g009
Figure 10. Agentic capability versus injection susceptibility across model scale: six Qwen2.5-Instruct models (burgundy; 1.5B–72B, marker size growing with parameter count) and a Llama-3.1-8B-Instruct control (indigo). The vertical axis scores InjecAgent attack success among valid responses; italic annotations give the share of invalid responses, which dominates the smallest models and would otherwise masquerade as robustness. Within the family, scale improves capability and injection-robustness together, and the Llama point shows the level is family-specific. Neither observation generalizes to the other attack classes of Table 11, where strong models still fail.
Figure 10. Agentic capability versus injection susceptibility across model scale: six Qwen2.5-Instruct models (burgundy; 1.5B–72B, marker size growing with parameter count) and a Llama-3.1-8B-Instruct control (indigo). The vertical axis scores InjecAgent attack success among valid responses; italic annotations give the share of invalid responses, which dominates the smallest models and would otherwise masquerade as robustness. Within the family, scale improves capability and injection-robustness together, and the Llama point shows the level is family-specific. Neither observation generalizes to the other attack classes of Table 11, where strong models still fail.
Preprints 223549 g010
Table 1. Comparison with prior general surveys of LLM-based agents. ✓ = an organizing pillar with dedicated sections; ∼ = discussed, but not an organizing pillar; ✗ = absent or mentioned only in passing. “Coverage” is the latest literature year substantively covered.
Table 1. Comparison with prior general surveys of LLM-based agents. ✓ = an organizing pillar with dedicated sections; ∼ = discussed, but not an organizing pillar; ✗ = absent or mentioned only in passing. “Coverage” is the latest literature year substantively covered.
Survey Plan. Mem. Tools Multi-ag. Envs. Eval. meth. Safety Coverage
Wang et al. (2023b) ✓ ✓ ∼ ∼ ∼ ∼ ✗ 2023
Xi et al. (2023) ✓ ∼ ✓ ✓ ∼ ∼ ∼ 2023
Zhang et al. (2023c) ✓ ✗ ∼ ✗ ✗ ∼ ✗ 2023
Durante et al. (2024) ∼ ∼ ∼ ∼ ✓ ✗ ∼ 2024
Guo et al. (2024a) ∼ ∼ ∼ ✓ ∼ ∼ ✗ 2024
Masterman et al. (2024) ✓ ∼ ✓ ✓ ✗ ✗ ∼ 2024
Liu et al. (2025) ✓ ✓ ∼ ✓ ∼ ∼ ✓ 2025
Luo et al. (2025) ✓ ✓ ✓ ✓ ∼ ∼ ∼ 2025
Plaat et al. (2025) ✓ ✓ ✓ ✗ 2025
This survey ✓ ✓ ✓ ✓ ✓ ✓ ✓ 2026
Table 4. From prompted to trained agents. Methods are ordered from training static reasoning to training agents online in an environment. “Env?” marks whether training uses a live interactive environment (✓), only offline traces recorded from one (∼), or no environment (✗).
Table 4. From prompted to trained agents. Methods are ordered from training static reasoning to training agents online in an environment. “Env?” marks whether training uses a live interactive environment (✓), only offline traces recorded from one (∼), or no environment (✗).
Method Training signal Data source Env? Target Ref.
STaR rationale bootstrap (SFT) self-generated ✗ reasoning (Zelikman et al. 2022)
DeepSeek-R1 outcome RL self-generated ✗ reasoning (DeepSeek-AI et al. 2025)
SCoRe multi-turn RL self-generated ✗ self-correction (Kumar et al. 2024)
AgentTuning trajectory SFT teacher traces ∼ general agent (Zeng et al. 2023)
FireAct trajectory SFT teacher ReAct traces ∼ reasoning + tools (Chen et al. 2023a)
xLAM function-call SFT synthetic pipeline ✗ tool calling (Zhang et al. 2024c)
ToolACE function-call SFT verified synthetic ✗ tool calling (Liu et al. 2024b)
Search-R1 outcome RL self-gen. + live search ✓ reasoning + search (Jin et al. 2025)
ReTool outcome RL self-gen. + code exec. ✓ reasoning + code (Feng et al. 2025a)
WebRL online curriculum RL self-evolving ✓ web (Qi et al. 2024)
AgentGym SFT + online RL many environments ✓ general agent (Xi et al. 2024)
UI-TARS SFT + RL GUI trajectories ✓ computer use (Qin et al. 2025)
Table 5. Long-term memory systems by their store form and policies. “Open” marks a public implementation; “–” means not clearly available. Skill-library systems (Voyager, ExpeL) store procedural memory and are discussed in Section 5.4.
Table 5. Long-term memory systems by their store form and policies. “Open” marks a public implementation; “–” means not clearly available. Skill-library systems (Voyager, ExpeL) store procedural memory and are discussed in Section 5.4.
System Store form Read policy Write policy Forget Open Ref.
Gen. Agents memory stream recency/import./relev. append + reflection recency ✓ (Park et al. 2023a)
MemGPT tiered context self-issued paging self-managed evict ✓ (Packer et al. 2023)
MemoryBank flat text relevance retrieval append + update Ebbinghaus ✓ (Zhong et al. 2023)
HippoRAG knowledge graph Personalized PageRank index into graph none ✓ (Gutiérrez et al. 2024)
A-MEM linked notes similarity + links note link + evolve implicit ✓ (Xu et al. 2025)
Mem0 text + graph salient-fact retrieval extract + consolidate salience ✓ (Chhikara et al. 2025)
Larimar episodic matrix addressable read one-shot update selective – (Das et al. 2024)
Voyager skill library task-embedding search add verified skill none ✓ (Wang et al. 2023a)
ExpeL insight rules relevance retrieval distill insights none ✓ (Zhao et al. 2023)
ReasoningBank strategy store relevance retrieval distill strategies none – (Ouyang et al. 2025)
Table 6. Representative tool-use benchmarks. “Execution” indicates whether tool calls are run against real APIs, a simulated/virtual server, or matched structurally; “Turns” distinguishes single-call from multi-turn tasks. Cross-referenced from Section 9.
Table 6. Representative tool-use benchmarks. “Execution” indicates whether tool calls are run against real APIs, a simulated/virtual server, or matched structurally; “Turns” distinguishes single-call from multi-turn tasks. Cross-referenced from Section 9.
Benchmark Tools / APIs Execution Turns Metric Ref.
API-Bank 73 APIs real (runnable) multi call correctness (Li et al. 2023b)
StableToolBench 16k+ (virtual) simulated server multi pass / win rate (Guo et al. 2024b)
BFCL multi-language AST matching single + multi AST accuracy (Patil et al. 2025)
GTA real deployed tools real execution multi task success (Wang et al. 2024a)
Table 7. Representative multi-agent frameworks. “Roles” distinguishes fixed (assigned in advance) from dynamic (recruited or pruned at run time); “Human” marks support for a human in the loop.
Table 7. Representative multi-agent frameworks. “Roles” distinguishes fixed (assigned in advance) from dynamic (recruited or pruned at run time); “Human” marks support for a human in the loop.
Framework Topology Roles Medium Human Domain Ref.
CAMEL dyad fixed natural language ✗ general (Li et al. 2023a)
AutoGen configurable fixed language + code ✓ general (Wu et al. 2023a)
MetaGPT pipeline (SOP) fixed structured docs ✗ software (Hong et al. 2024)
ChatDev pipeline (chain) fixed natural language ✗ software (Qian et al. 2023)
AgentVerse dynamic group dynamic natural language ✗ general (Chen et al. 2023d)
MoA layered fixed responses ✗ general (Wang et al. 2024b)
DyLAN dynamic network dynamic natural language ✗ reasoning (Liu et al. 2023b)
Debate mesh symmetric natural language ✗ reasoning (Du et al. 2023)
GPTSwarm learned graph learned messages ✗ general (Zhuge et al. 2024a)
AFlow searched workflow fixed code ✗ general (Zhang et al. 2024d)
Table 9. The agent benchmark landscape: generalist suites (top) and capability-specific benchmarks (bottom). “Env type” summarizes the interaction setting; “–” marks a count the source does not fix (e.g., open model rosters). Overlapping tool and environment benchmarks appear in Table 6 and Table 8.
Table 9. The agent benchmark landscape: generalist suites (top) and capability-specific benchmarks (bottom). “Env type” summarizes the interaction setting; “–” marks a count the source does not fix (e.g., open model rosters). Overlapping tool and environment benchmarks appear in Table 6 and Table 8.
Benchmark Target Env type #Tasks Metric Ref.
GAIA general assistant web + tools 466 exact match (Mialon et al. 2023)
AgentBench general agent 8 environments – success rate (Liu et al. 2024c)
AgentBoard multi-turn agent 9 task types 1013 envs progress rate (Ma et al. 2024)
SmartPlay agentic capabilities 6 games – reward (Wu et al. 2023b)
SWE-bench software engineering 12 repositories 2294 resolved % (Jimenez et al. 2023)
τ -bench tool–agent–user retail, airline – pass k (Yao et al. 2024)
τ 2 -bench dual-control dialogue telecom – pass k (Barres et al. 2025)
TravelPlanner constrained planning sandbox 1225 plan validity (Xie et al. 2024a)
InterCode interactive coding Bash/SQL/Python – success rate (Yang et al. 2023c)
Terminal-Bench terminal workflows sandboxed terminal 89 test-verified success (Merrill et al. 2026)
MLAgentBench ML experimentation code + execution 13 improvement (Huang et al. 2023b)
ScienceAgentBench data-driven science code 102 success rate (Chen et al. 2024b)
OSWorld computer use real OS 369 execution success (Xie et al. 2024b)
WebArena web navigation self-hosted sites 812 functional success (Zhou et al. 2024c)
Table 11. Agent attack classes, a representative attack and benchmark for each, the prevailing defense direction, and whether a robust general defense exists. ∼ marks partial mitigations; ✗ marks an open problem. The pervasiveness of ✗ is the section’s central finding.
Table 11. Agent attack classes, a representative attack and benchmark for each, the prevailing defense direction, and whether a robust general defense exists. ∼ marks partial mitigations; ✗ marks an open problem. The pervasiveness of ✗ is the section’s central finding.
Attack class Representative Benchmark Defense direction Robust?
Direct prompt injection jailbreak prompts AgentHarm (Andriushchenko et al. 2024) instruction hierarchy ∼
Indirect prompt injection Greshake et al. (2023) AgentDojo (Debenedetti et al. 2024), InjecAgent (Zhan et al. 2024) spotlighting, CaMeL ∼
Memory / KB poisoning AgentPoison (Chen et al. 2024a) ASB (Zhang et al. 2024b) provenance filtering ✗
Backdoor BadAgent (Wang et al. 2024f) ASB (Zhang et al. 2024b) data auditing ✗
Harmful-task misuse jailbroken agent SafeAgentBench (Yin et al. 2024), OS-Harm (Kuntz et al. 2025) refusal + monitoring ∼
Table 12. The survey’s both-sides findings collected in one place: where an optimistic consensus meets documented negative evidence, and how we reconcile them.
Table 12. The survey’s both-sides findings collected in one place: where an optimistic consensus meets documented negative evidence, and how we reconcile them.
Topic Optimistic claim Counter-evidence Reconciliation §
Self-correction self-critique improves reasoning fails without an external signal (Huang et al. 2023a; Kamoi et al. 2024) needs external or trained-in feedback Section 4.4
burgundytintblack Multi-agent debate debate raises factuality often ≤ ensembling at equal cost (Smit et al. 2023) useful for oversight, not raw accuracy Section 7.2
burgundytintblack Memory a larger store helps structure and policy dominate capacity (Zeng et al. 2024) read/write policy is the bottleneck Section 5.3
burgundytintblack Agent scaffolds complex agents are required a simple pipeline can match them (Xia et al. 2024) strong feedback favors simpler designs Section 10.1
burgundytintblack Evaluation accuracy ranks agents cost-blind leaderboards mislead (Kapoor et al. 2024) report cost alongside accuracy Section 9.3
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.