Submitted:
16 July 2026
Posted:
17 July 2026
You are already at the latest version
Abstract
Keywords:
1. Introduction
Why long-horizon task execution is hard for agents.
- Sparse, delayed rewards and irreversibility. Many long-horizon tasks return a reward only at the end, leaving the agent with an extremely sparse learning signal over its intermediate decisions [18]; moreover, the longer an agent acts, the more likely it is to take a risky, irreversible action, so monitoring its execution safety and recoverability becomes essential.
Why long-horizon agents need a new taxonomy.
- Limited and unsystematic coverage. Existing papers typically examine a single technical component of long-horizon agents in isolation, such as memory [19], context engineering [20,21], agent harnesses [14,22], self-evolution [23,24], or agentic reinforcement learning [18]. Each is a genuine piece of the long-horizon agent, but none offers a comprehensive account of what such an agent is, nor of how to build an efficient and trustworthy long-horizon agent that stays reliable across a full trajectory.
- Unclear definitions and overlapping terminology. As interest in capable agents grows, the concept itself has become broad and blurred. Related notions such as long-context, long-running, autonomous, and self-evolving agents are routinely used interchangeably, even though each emphasizes a different facet, and reported gains entangle changes in the model, the harness, the data pipeline, and the evaluation protocol. Most works cannot say precisely what counts as long-horizon agency or how these techniques relate, which leaves improvements hard to attribute and hard to transfer.
- Foundation. How should long-horizon agency be formalized, and how does it differ from long-running execution, autonomy, and self-evolution?
- Evolution. How did the control surface widen from prompts, to context, to runtime agent systems?
- Harness. How does externalized harness engineering keep long-horizon agent execution specifiable, steerable, verifiable, and recoverable at runtime?
- Optimization. How can internalized model optimization build long-horizon competence across the agent’s architectural substrate, data and environment synthesis, and full training lifecycle?
- Application. What forms do long-horizon agents take across domains, and how should their long-horizon capability be systematically evaluated?
- Frontier. What open problems define the next stage as horizons grow?
Contributions and Roadmap.
- 1.
- Foundation (sec:prelim). We formalize long-horizon agency as a coupled model–harness decision process and organize its difficulty into three nested task levels (H1–H3) with matching capabilities (C1–C3): intra-context reasoning, cross-context memory, and cross-task experience. Building on METR’s data, we note that the frontier time horizon is scaling with an accelerating doubling time. We further separate long-horizon agency from long-running execution, autonomy, and self-evolution, giving the field a shared vocabulary for attributing where long-horizon competence comes from.
- 2.
- Evolution (sec:evolution). We trace how the control surface migrated from prompts, to context, to runtime systems, framing today’s long-horizon agents as the current endpoint of a capability-driven trajectory and clarifying why control first moved outward before it could be internalized.
- 3.
- Harness (sec:externalization). We organize externalized harness engineering by component: loops and workflows, context and memory, tools, MCP, and skills, orchestration and multi-agent control, hooks and middleware, and verification. Together these components show how a runtime layer keeps long execution specifiable, steerable, verifiable, and recoverable.
- 4.
- Optimization (sec:internalization). We organize internalized model optimization along the pipeline, from the architectural substrate and data and environment synthesis to agentic pre/mid-training, fine-tuning, reinforcement learning with long-horizon credit assignment, on-policy distillation, and self-evolution, tracing how capabilities first scaffolded through externalized harness engineering are progressively internalized into the policy.
- 5.
- Application (sec:patterns). Organized by the agent–environment interface, we map long-horizon pressures across software engineering, information seeking, computer use, multimodal agents, and general-purpose agents, revealing a common long-horizon core beneath surface-level domain differences, and consolidate horizon-aware evaluation resources.
- 6.
- Frontier (sec:frontiers). We organize open problems into four axes, namely evolution, effectiveness, efficiency, and trustworthiness, spanning nine concrete directions that range from self-evolving and transferable harnesses, to real-world environments and error robustness, to cost- and modality-aware efficiency, and to safety, governance, and embodiment, marking what must be solved as horizons keep growing.
Scope and Boundaries.
2. Foundation: Formalizing Long-Horizon Agents
- What is a long-horizon agent? (ssec:formal-agent) We define it by the coupled logical dependencies of a task rather than by elapsed time, and cast the agent as a base policy coupled to a surrounding harness, so long-horizon agency becomes a property of the model–harness system rather than of the foundation model alone.
- How does its difficulty decompose? (ssec:horizon-dims) We use how far execution stretches beyond one working context as an observable proxy for the underlying dependency burden, giving three nested levels (H1–H3) and their capabilities (C1–C3).
- How fast is the frontier moving? (ssec:scaling-law) We quantify the rising time horizon of frontier agents as an empirical scaling trend and compare its full-period and recent-window doubling times.
- How does it differ from neighboring notions? (ssec:vs-related) We separate long-horizon agency from long-running execution, autonomy, and self-evolution, showing that it integrates these rather than reducing to any one.
2.1. Definition
Long-horizon Tasks.
Long-horizon Agents.
The Long-horizon Agent–Environment Loop.
2.2. Three Levels of Long-Horizon Tasks and Capabilities
- H1 — Intra-context, within one window (∼minutes). A single task can be completed within one working context window, yet it is far from a one-shot query: it requires composing many interdependent decisions into one coherent trajectory, with repeated interaction with the environment and, where relevant, the user to test and revise the reasoning direction until the goal is reached [10,44].
- H2 — Cross-context, across windows/sessions (∼hours to days). A single task far exceeds one context window or session and must be carried across multiple context windows or by several agents working in concert. The agent(s) must therefore actively compress and externalize historical context, track interaction state in real time, and resume progress without assuming that the full history remains visible at every step [9,12,13].
- H3 — Cross-task, open-ended task stream (∼lifelong). Tasks arrive as an open-ended stream rather than a closed episode. A single overarching goal may therefore span many tasks, each of which may itself require multiple context windows to complete. In deployment, such settings often demand persistent, and sometimes lifelong, operation across heterogeneous environments, tools, and sessions [18,23,139,142].
- C1 — Intra-context interactive reasoning (for H1). Within a single context, sustain continuous reasoning through interaction with the environment and the user: take actions, absorb feedback, and revise the trajectory so that many interdependent decisions compose into one coherent path toward the goal [10,44].
- C3 — Cross-task experience accumulation (for H3). Building on C2, turn experience into lasting competence across an open-ended task stream: switch among complex tasks, each of which may itself span multiple context windows, accumulate reusable skills [143], generalize from past experience [144], and continually improve as goals, environments, and data distributions shift [23,139,142].
2.3. Scaling the Time Horizon of Frontier-Agent Execution
2.4. Long-Horizon Agent vs. Neighboring Concepts
2.4.1. Long-Horizon Agent vs. Long-Running Agent
Overlap.
Distinctions.
2.4.2. Long-Horizon Agent vs. Autonomous Agent
Overlap.
Distinctions.
2.4.3. Long-Horizon Agent vs. Self-Evolving Agent
Overlap.
Distinctions.
3. Evolution: From Prompting to Runtime
3.1. Stage I: Prompt Engineering (2020–2023)
Chain-of-Thought Reasoning.
Decomposition, Planning, and Search.
Prompt Optimization and Internalization.
3.2. Stage II: Context Engineering (2023–2025)
Retrieval-Augmented Generation.
Tool Use and Function Calling.
Long Context, Memory, and Context Engineering.
3.3. Stage III: Runtime Harnesses (2025–Present)
Agent Loops.
Harness Engineering.
Internalizing the Loop.
4. Harness: Externalizing Long-Horizon Capability
- 1.
- Loops and Workflows (ssec:ext-loop): specifies how execution moves across states, including when the model is called, how actions are taken, and how control returns after observations.
- 2.
- Context and Memory (ssec:ext-context): manages the state available to the model, from the active context of the current run to persistent memory across windows, tasks, and sessions.
- 3.
- Tools, MCP, and Skills (ssec:ext-tools): defines how the agent accesses external capabilities, represents actions, and reuses acquired procedures.
- 4.
- Orchestration (ssec:ext-orch): organizes work beyond a single agent loop, including task decomposition, role assignment, routing, and inter-agent coordination.
- 5.
- Hooks and Middleware (ssec:hooks): provides predefined interception points in the runtime where execution can be inspected, modified, paused, or handed off.
- 6.
- Verification (ssec:verification): checks whether states, actions, outputs, or side effects satisfy task, safety, and quality constraints before execution continues or results are accepted.
4.1. Loops and Workflows
4.1.1. Linear Workflows
4.1.2. Plan-Execute Workflows
4.1.3. Branching Workflows
4.2. Context and Memory
4.2.1. Working Context
Context Discard.
Context Compression.
Context Selection.
4.2.2. Persistent Memory
Memory Contents.
Memory Operations.
4.3. Tools, MCP, and Skills
4.3.1. Tool Interfaces and Protocols
Function Calling.
Tool Protocols and MCP.
Multi-turn Tool Use.
4.3.2. Active Tool Discovery
Tool Retrieval.
Programmatic Tool Use.
4.3.3. Skill Libraries
Procedural Skills.
Skill Acquisition.
Skill Library Evolution.
4.4. Orchestration
4.4.1. Decomposition and Roles
4.4.2. Coordination Topologies
4.4.3. Orchestration Optimization
4.4.4. Agent Protocols
4.5. Hooks and Middleware
4.5.1. Pre-Defined Rule-Based Hooks
- Input Boundary. Operating before data ingestion by the core language model, this mechanism intercepts incoming raw data streams, including user instructions, prompt templates, and retrieved external context. Its primary objective is to filter out adversarial prompt injections, jailbreaks, and malformed system directives designed to compromise the model’s behavioral alignment. Systems implementing this boundary [374] deploy independent, token-level or semantic static filters to reject non-compliant inputs before they can influence the model’s internal attention state.
- Tool-call Boundary. Acting as a pre-execution firewall, this mechanism inspects tool arguments before execution to detect dangerous options (e.g., rm -rf) and common attack patterns such as code injection or unauthorized file access. AEGIS [249] instantiates this design by interposing a policy engine on every tool call, validating arguments against a curated set of known vulnerability patterns, including path traversal, with negligible latency overhead.
- Protocol Boundary. Enforced at the inter-component communication layer, this boundary requires every request to carry explicit credentials and receive per-action authorization, ensuring no message is implicitly trusted regardless of its origin. Representative systems include Authenticated Workflows [376], which attaches cryptographic tokens to protocol messages for authenticity verification, and MagenticUI [377], which mandates explicit user approval before dispatching agent actions.
4.5.2. Custom User-Defined Hooks
4.5.3. Runtime-Adaptive Hooks
4.6. Verification
4.6.1. Assessment Targets
4.6.2. Verification Levels
4.6.3. Verifier Strategies
5. Optimization: Internalizing Long-Horizon Capability
- 1.
- Architectural Substrate (ssec:int-arch): makes long trajectories efficient to train on and decode from.
- 2.
- Data and Environment Synthesis (ssec:int-data): constructs verifiable long-horizon experience, including tasks, environments, and trajectories, that later stages can use for training and feedback.
- 3.
- Pre-/Mid-training (ssec:int-pretrain): builds action–observation and reasoning priors, so agentic behavior becomes native rather than prompted.
- 4.
- Fine-tuning (ssec:int-sft): instills behavioral and tool-use discipline from curated agentic trajectories.
- 5.
- Reinforcement Learning (ssec:int-rl): optimizes whole trajectories and confronts long-horizon credit assignment under sparse, delayed reward.
- 6.
- On-policy Distillation (ssec:int-opd): consolidates the policy on its own state distribution, reducing distribution shift and stabilizing earlier learned behaviors.
- 7.
- Self-evolution (ssec:int-evolving): turns self-generated experience, feedback, and task-environment updates into continuing policy improvement.
5.1. Architectural Substrate
5.1.1. Explicit-Context Architectures
5.1.2. Compressed-State Architectures
5.1.3. Hybrid Architectures
5.1.4. High-Throughput Mechanisms
5.2. Data and Environment Synthesis
5.2.1. Task Synthesis
Dependency-chain Construction.
Graph-structured Information Construction.
State-transition Construction.
Curriculum-driven Construction.
5.2.2. Environment Synthesis
Symbolic Environments.
Neural Environments and World Models.
Neural–Symbolic Hybrid Environments.
5.2.3. Trajectory Synthesis
Environment-grounded Execution Records.
Policy-aligned Rollouts.
Step-level Credit Assignment and Failure Reuse.
History Compression and Experience Reuse.
5.3. Pre-Training and Mid-Training
5.3.1. Reasoning Priors
5.3.2. Long-Context State
5.3.3. Multimodal Perception
5.3.4. Data Mixture and Training Strategies
5.4. Fine-Tuning
5.4.1. Instruction Selection and Mixing
Trajectory Quality Over Quantity.
Compact Reasoning Demonstrations.
Executable Tool-use Data and Mixture.
5.4.2. Curriculum Learning
Staging by Horizon.
Staging by Difficulty.
5.4.3. Distillation
Strong-to-weak Distillation.
Self-distillation and Reflective Self-improvement.
5.5. Agentic Reinforcement Learning
5.5.1. Credit Assignment
Outcome Credit Assignment.
Process Credit Assignment.
Rubric Credit Assignment.
5.5.2. Policy Optimization
Clipping-based Policy Optimization.
Turn-level Policy Optimization.
Entropy-based Policy Optimization.
5.5.3. Sampling Strategy
Trajectory-level Rollout.
Tree-structured Sampling.
Budget-aware and Pruning Strategies.
5.5.4. Interaction Patterns
Hierarchical Interaction Frameworks.
Multi-agent Frameworks.
Memory-state Learning.
5.6. On-Policy Distillation
5.6.1. Teacher-Guided OPD
Supervision Scheduling.
Teacher Signals.
5.6.2. Self-Improving OPD
Privileged Information.
Environment Feedback.
5.7. Self-Evolution
5.7.1. Offline Self-Evolution
5.7.2. Online Self-Evolution
5.7.3. Agent–Environment Co-Evolution
6. Application: Long-Horizon Agents in Practice
- 1.
- Software Engineering Agents (ssec:apps-code): act over code repositories and execution environments, where edits can often be checked through tests, compilers, and runtime outputs.
- 2.
- Information-seeking Agents (ssec:apps-research): interact with web documents, search engines, and knowledge sources, where the challenge is to gather and synthesize evidence without a direct execution oracle.
- 3.
- Computer-use Agents (ssec:apps-cua): operate graphical interfaces through screenshots, accessibility trees, and low-level actions, inferring and manipulating partially observable interface state.
- 4.
- Multimodal Agents (ssec:apps-multimodal): reason over long visual, auditory, and textual streams, where relevant information is distributed across time and modalities.
- 5.
- General-purpose Agents (ssec:apps-general): combine several such interfaces within a single workflow, drawing on heterogeneous tools, permissions, and external services.
- What information and actions does the interface expose?
- What kinds of state and feedback must persist across a long trajectory?
- What capabilities must the agent or harness add to compensate for what the environment does not provide directly?
6.1. Software Engineering
6.1.1. Repository Grounding
6.1.2. Workflow-Level Planning
6.1.3. Feedback-Driven Repair
6.2. Information Seeking
6.2.1. Deep Search
6.2.2. Wide Search
6.2.3. Multimodal Grounding
6.2.4. Research Synthesis
6.3. Computer Use
6.3.1. Browser Agents
6.3.2. Desktop GUI Agents
6.3.3. Mobile Agents
6.4. Multimodal Agents
6.4.1. Multimodal Understanding
6.4.2. Multimodal Generation
6.4.3. Omnimodal Agency
6.5. General-Purpose Agents
6.5.1. Personal Assistants
6.5.2. Embodied Agents and World Models
6.5.3. Productive Agents
Professional Services.
Scientific Discovery and Reasoning.
6.6. Benchmarks and Resources
7. Frontier: Open Challenges and Outlooks
7.1. Evolution: Generalization and Learning Over Time
7.1.1. Self-Evolving Harness and Agents
Status and Challenges.
Outlook.
7.1.2. Harness Generalization and Transferability
Status and Challenges.
Outlook.
7.1.3. Continual and Lifelong Learning
Status and Challenges.
Outlook.
7.2. Effectiveness: Acting in Realistic Environments
7.2.1. Real-World Environment Interaction
Status and Challenges.
Outlook.
7.2.2. From Digital to Embodied Agents
Status and Challenges.
Outlook.
7.3. Efficiency: Spending Compute, Context, and Modality Budgets
7.3.1. Cost- and Budget-Aware Agency
Status and Challenges.
Outlook.
7.3.2. Multimodal and Omni Harnesses
Status and Challenges.
Outlook.
7.4. Trustworthiness: Robust and Governed Autonomy
7.4.1. Reflection and Error Robustness
Status and Challenges.
Outlook.
7.4.2. Safety and Governance
Status and Challenges.
Outlook.
8. Conclusions
References
- Jimenez, C.E.; Yang, J.; Wettig, A.; Yao, S.; Pei, K.; Press, O.; Narasimhan, K.R. SWE-bench: Can Language Models Resolve Real-world Github Issues? In Proceedings of the The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024; 2024. Available online: https://openreview.net/.
- Yang, J.; Jimenez, C.E.; Wettig, A.; Lieret, K.; Yao, S.; Narasimhan, K.; Press, O. SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering. In Proceedings of the Advances in Neural Information Processing Systems 37: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024; Vancouver, BC, Canada, Globersons, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J.M., Zhang, C., Eds.; 10 - 15 December 2024. [Google Scholar]
- Mialon, G.; Fourrier, C.; Wolf, T.; LeCun, Y.; Scialom, T. GAIA: a benchmark for General AI Assistants. In Proceedings of the The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024; 2024. Available online: https://openreview.net/.
- Wei, J.; Yang, Y.; Zhang, X.; Chen, Y.; Zhuang, X.; Gao, Z.; Zhou, D.; Wang, G.; Gao, Z.; Cao, J.; et al. From AI for Science to Agentic Science: A Survey on Autonomous Scientific Discovery. CoRR 2025, abs/2508.14111, [2508.14111. [Google Scholar] [CrossRef]
- Zhou, S.; Xu, F.F.; Zhu, H.; Zhou, X.; Lo, R.; Sridhar, A.; Cheng, X.; Ou, T.; Bisk, Y.; Fried, D.; et al. WebArena: A Realistic Web Environment for Building Autonomous Agents. In Proceedings of the The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024; 2024. Available online: https://openreview.net/.
- Xie, T.; Zhang, D.; Chen, J.; Li, X.; Zhao, S.; Cao, R.; Hua, T.J.; Cheng, Z.; Shin, D.; Lei, F.; et al. OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments. In Proceedings of the Advances in Neural Information Processing Systems 37: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024; Vancouver, BC, Canada, Globersons, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J.M., Zhang, C., et al., Eds.; 10 - 15 December 2024. [Google Scholar]
- Durante, Z.; Huang, Q.; Wake, N.; Gong, R.; Park, J.S.; Sarkar, B.; Taori, R.; Noda, Y.; Terzopoulos, D.; Choi, Y.; et al. Agent AI: Surveying the Horizons of Multimodal Interaction. CoRR 2024, abs/2401.03568, 2401.03568. [Google Scholar] [CrossRef]
- OpenAI. Harness engineering: leveraging Codex in an agent-first world. 2026. [Google Scholar]
- Anthropic. Effective Harnesses for Long-Running Agents; 2025. [Google Scholar]
- Kwa, T.; West, B.; Becker, J.; Deng, A.; Garcia, K.; Hasin, M.; Jawhar, S.; Kinniment, M.; Rush, N.; von Arx, S.; et al. Measuring AI Ability to Complete Long Software Tasks. In Proceedings of the Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025; San Diego, CA, USA, December 2-7, 2025 / Mexico City, Mexico, Belgrave, D., Zhang, C., Montoya, L.N., Lin, H., Pascanu, R., Koniusz, P., Ghassemi, M., Chen, N., Ruíz, I.V.M., Loaiza-Bonilla, A., Eds.; 30 November 2025. [Google Scholar]
- G.L.M. GLM-5: from Vibe Coding to Agentic Engineering. CoRR 2026, abs/2602.15763, [2602.15763. [Google Scholar] [CrossRef]
- Liu, N.F.; Lin, K.; Hewitt, J.; Paranjape, A.; Bevilacqua, M.; Petroni, F.; Liang, P. Lost in the Middle: How Language Models Use Long Contexts. Trans. Assoc. Comput. Linguist. 2024, 12, 157–173. [Google Scholar] [CrossRef]
- Research, C. Context Rot: When Long Contexts Hurt LLM Performance, 2025.
- Pan, L.; Zou, L.; Guo, S.; Ni, J.; Zheng, H. Natural-Language Agent Harnesses. CoRR 2026, abs/2603.25723, 2603.25723. [Google Scholar] [CrossRef]
- Ding, D.; Liu, S.; Yang, E.; Lin, J.; Chen, Z.; Dou, S.; Guo, H.; Cheng, W.; Zhao, P.; Xiao, C.; et al. OctoBench: Benchmarking Scaffold-Aware Instruction Following in Repository-Grounded Agentic Coding. In Proceedings of the Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2026; San Diego, California, United States, Liakata, M., Moreira, V.P., Zhang, J., Jurgens, D., Eds.; Association for Computational Linguistics, 2-7 July 2026; Volume 2026, pp. 5958–5978. [Google Scholar]
- Arike, R.; Donoway, E.; Bartsch, H.; Hobbhahn, M. Technical Report: Evaluating Goal Drift in Language Model Agents. CoRR 2025, abs/2505.02709, 2505.02709. [Google Scholar] [CrossRef]
- Sinha, A.; Arun, A.; Goel, S.; Staab, S.; Geiping, J. The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs. In Proceedings of the The Fourteenth International Conference on Learning Representations, 2026. [Google Scholar]
- Zhang, G.; Geng, H.; Yu, X.; Yin, Z.; Zhang, Z.; Tan, Z.; Zhou, H.; Li, Z.; Xue, X.; Li, Y.; et al. The Landscape of Agentic Reinforcement Learning for LLMs: A Survey. In Trans. Mach. Learn. Res.; 2026. [Google Scholar]
- Wu, Y.; Liang, S.; Zhang, C.; Wang, Y.; Zhang, Y.; Guo, H.; Tang, R.; Liu, Y. From Human Memory to AI Memory: A Survey on Memory Mechanisms in the Era of LLMs. CoRR 2025, abs/2504.15965, [2504.15965. [Google Scholar] [CrossRef]
- Mei, L.; Yao, J.; Ge, Y.; Wang, Y.; Bi, B.; Cai, Y.; Liu, J.; Li, M.; Li, Z.; Zhang, D.; et al. A Survey of Context Engineering for Large Language Models. CoRR 2025, abs/2507.13334, [2507.13334. [Google Scholar] [CrossRef]
- Zhang, Y.; Liu, Z.; Zhu, J.; Wang, S.; Chen, X.; Huang, H.; Kuang, J.; Chen, S.; Shen, A.; Wu, H.; et al. From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI. CoRR 2026, abs/2606.14502, [2606.14502. [Google Scholar] [CrossRef]
- Zhou, C.; Chai, H.; Chen, W.; Guo, Z.; Shan, R.; Song, Y.; Xu, T.; Yang, Y.; Yu, A.; Zhang, W.; et al. Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering. CoRR 2026, abs/2604.08224, 2604.08224. [Google Scholar] [CrossRef]
- Gao, H.; Geng, J.; Hua, W.; Hu, M.; Juan, X.; Liu, H.; Liu, S.; Qiu, J.; Qi, X.; Ren, Q.; et al. A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence. Trans. Mach. Learn. Res. 2026. [Google Scholar]
- Xiang, Z.; Yang, C.; Chen, Z.; Wei, Z.; Tang, Y.; Teng, Z.; Peng, Z.; Li, Z.; Huang, C.; He, Y.; et al. A Systematic Survey of Self-Evolving Agents: From Model-Centric to Environment-Driven Co-Evolution. Technical report, TechRxiv (no arXiv version available). 2026. [Google Scholar] [CrossRef]
- Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Ichter, B.; Xia, F.; Chi, E.H.; Le, Q.V.; Zhou, D. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. In Proceedings of the Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022; New Orleans, LA, USA, Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., Oh, A., Eds.; 28 November 2022. [Google Scholar]
- Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C.L.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al. Training language models to follow instructions with human feedback. In Proceedings of the Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022; New Orleans, LA, USA, Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., Oh, A., Eds.; 28 November 2022. [Google Scholar]
- Wang, L.; Xu, W.; Lan, Y.; Hu, Z.; Lan, Y.; Lee, R.K.; Lim, E. Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models. In Proceedings of the Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics; Toronto, Canada, Rogers, A., Boyd-Graber, J.L., Okazaki, N., Eds.; Association for Computational Linguistics, 9-14 July 2023; Volume 1, pp. 2609–2634. [Google Scholar] [CrossRef]
- Zhou, D.; Schärli, N.; Hou, L.; Wei, J.; Scales, N.; Wang, X.; Schuurmans, D.; Cui, C.; Bousquet, O.; Le, Q.V.; et al. Least-to-Most Prompting Enables Complex Reasoning in Large Language Models. In Proceedings of the The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023; 2023. Available online: https://openreview.net/.
- Madaan, A.; Tandon, N.; Gupta, P.; Hallinan, S.; Gao, L.; Wiegreffe, S.; Alon, U.; Dziri, N.; Prabhumoye, S.; Yang, Y.; et al. Self-Refine: Iterative Refinement with Self-Feedback. In Proceedings of the Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023; New Orleans, LA, USA, Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., Levine, S., Eds.; 10 - 16 December 2023. [Google Scholar]
- Lewis, P.; Perez, E.; Piktus, A.; Petroni, F.; Karpukhin, V.; Goyal, N.; Küttler, H.; Lewis, M.; Yih, W.; Rocktäschel, T.; et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In Proceedings of the Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020; virtual, Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., Lin, H., Eds.; 2020. [Google Scholar]
- Edge, D.; Trinh, H.; Cheng, N.; Bradley, J.; Chao, A.; Mody, A.; Truitt, S.; Larson, J. From Local to Global: A Graph RAG Approach to Query-Focused Summarization. CoRR 2024, abs/2404.16130, [2404.16130. [Google Scholar] [CrossRef]
- Sarthi, P.; Abdullah, S.; Tuli, A.; Khanna, S.; Goldie, A.; Manning, C.D. RAPTOR: Recursive Abstractive Processing for Tree-Organized Retrieval. In Proceedings of the The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024; 2024. Available online: https://openreview.net/.
- Gao, L.; Ma, X.; Lin, J.; Callan, J. Precise Zero-Shot Dense Retrieval without Relevance Labels. In Proceedings of the Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics; Toronto, Canada, Rogers, A., Boyd-Graber, J.L., Okazaki, N., Eds.; Association for Computational Linguistics, 9-14 July 2023; Volume 1, pp. 1762–1777. [Google Scholar] [CrossRef]
- Asai, A.; Wu, Z.; Wang, Y.; Sil, A.; Hajishirzi, H. Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection. In Proceedings of the The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024; 2024. Available online: https://openreview.net/.
- Wu, X.; Li, K.; Zhao, Y.; Zhang, L.; Ou, L.; Yin, H.; Zhang, Z.; Jiang, Y.; Xie, P.; Huang, F.; et al. ReSum: Unlocking Long-Horizon Search Intelligence via Context Summarization. CoRR 2025, abs/2509.13313, 2509.13313. [Google Scholar] [CrossRef]
- Yao, S.; Zhao, J.; Yu, D.; Du, N.; Shafran, I.; Narasimhan, K.R.; Cao, Y. ReAct: Synergizing Reasoning and Acting in Language Models. In Proceedings of the The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023; 2023. Available online: https://openreview.net/.
- Schick, T.; Dwivedi-Yu, J.; Dessì, R.; Raileanu, R.; Lomeli, M.; Hambro, E.; Zettlemoyer, L.; Cancedda, N.; Scialom, T. Toolformer: Language Models Can Teach Themselves to Use Tools. In Proceedings of the Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023; New Orleans, LA, USA, Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., Levine, S., Eds.; 10 - 16 December 2023. [Google Scholar]
- Wang, X.; Chen, Y.; Yuan, L.; Zhang, Y.; Li, Y.; Peng, H.; Ji, H. Executable Code Actions Elicit Better LLM Agents. In Proceedings of the Forty-first International Conference on Machine Learning, ICML 2024 PMLR / OpenReview.net; Vienna, Austria, Salakhutdinov, R., Kolter, Z., Heller, K.A., Weller, A., Oliver, N., Scarlett, J., Berkenkamp, F., Eds.; Proceedings of Machine Learning Research; 21-27 July 2024; Vol. 235, pp. 50208–50232. [Google Scholar]
- Gravitas, S. AutoGPT: Build, Deploy, and Run AI Agents, 2023.
- Wang, X.; Li, B.; Song, Y.; Xu, F.F.; Tang, X.; Zhuge, M.; Pan, J.; Song, Y.; Li, B.; Singh, J.; et al. OpenHands: An Open Platform for AI Software Developers as Generalist Agents. In Proceedings of the The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025; 2025. Available online: https://openreview.net/.
- Xu, B.; Peng, Z.; Lei, B.; Mukherjee, S.; Liu, Y.; Xu, D. ReWOO: Decoupling Reasoning from Observations for Efficient Augmented Language Models. CoRR 2023, abs/2305.18323, [2305.18323. [Google Scholar] [CrossRef]
- Yao, S.; Yu, D.; Zhao, J.; Shafran, I.; Griffiths, T.; Cao, Y.; Narasimhan, K. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. In Proceedings of the Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023; New Orleans, LA, USA, Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., Levine, S., Eds.; 10 - 16 December 2023. [Google Scholar]
- Zhou, A.; Yan, K.; Shlapentokh-Rothman, M.; Wang, H.; Wang, Y. Language Agent Tree Search Unifies Reasoning, Acting, and Planning in Language Models. In Proceedings of the Forty-first International Conference on Machine Learning, ICML 2024 PMLR / OpenReview.net; Vienna, Austria, Salakhutdinov, R., Kolter, Z., Heller, K.A., Weller, A., Oliver, N., Scarlett, J., Berkenkamp, F., Eds.; Proceedings of Machine Learning Research; 21-27 July 2024; Vol. 235, pp. 62138–62160. [Google Scholar]
- Shinn, N.; Cassano, F.; Gopinath, A.; Narasimhan, K.; Yao, S. Reflexion: language agents with verbal reinforcement learning. In Proceedings of the Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023; New Orleans, LA, USA, Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., Levine, S., Eds.; 10 - 16 December 2023. [Google Scholar]
- Packer, C.; Fang, V.; Patil, S.G.; Lin, K.; Wooders, S.; Gonzalez, J.E. MemGPT: Towards LLMs as Operating Systems. CoRR 2023, abs/2310.08560, [2310.08560. [Google Scholar] [CrossRef]
- Zhong, W.; Guo, L.; Gao, Q.; Ye, H.; Wang, Y. MemoryBank: Enhancing Large Language Models with Long-Term Memory. In Proceedings of the Thirty-Eighth AAAI Conference on Artificial Intelligence, AAAI 2024, Thirty-Sixth Conference on Innovative Applications of Artificial Intelligence, IAAI 2024, Fourteenth Symposium on Educational Advances in Artificial Intelligence, EAAI 2024, February 20-27, 2024; Vancouver, Canada, Wooldridge, M.J., Dy, J.G., Natarajan, S., Eds.; AAAI Press, 2024; pp. 19724–19731. [Google Scholar] [CrossRef]
- Kang, J.; Ji, M.; Zhao, Z.; Bai, T. Memory OS of AI Agent. In Proceedings of the Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025; Suzhou, China, Christodoulopoulos, C., Chakraborty, T., Rose, C., Peng, V., Eds.; Association for Computational Linguistics, 4-9 November 2025; Volume 2025, pp. 25961–25970. [Google Scholar] [CrossRef]
- Yu, H.; Chen, T.; Feng, J.; Chen, J.; Dai, W.; Yu, Q.; Zhang, Y.; Ma, W.; Liu, J.; Wang, M.; et al. MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent. CoRR 2025, abs/2507.02259, 2507.02259. [Google Scholar] [CrossRef]
- Patil, S.G.; Zhang, T.; Wang, X.; Gonzalez, J.E. Gorilla: Large Language Model Connected with Massive APIs. In Proceedings of the Advances in Neural Information Processing Systems 37: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024; Vancouver, BC, Canada, Globersons, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J.M., Zhang, C., Eds.; 10 - 15 December 2024. [Google Scholar]
- Patil, S.G.; Mao, H.; Yan, F.; Ji, C.C.; Suresh, V.; Stoica, I.; Gonzalez, J.E. The Berkeley Function Calling Leaderboard (BFCL): From Tool Use to Agentic Evaluation of Large Language Models. In Proceedings of the Forty-second International Conference on Machine Learning, ICML 2025; Vancouver, BC, Canada, Singh, A., Fazel, M., Hsu, D., Lacoste-Julien, S., Berkenkamp, F., Maharaj, T., Wagstaff, K., Zhu, J., Eds.; PMLR / OpenReview.net; Proceedings of Machine Learning Research; 13-19 July 2025; Vol. 267. [Google Scholar]
- Qin, Y.; Liang, S.; Ye, Y.; Zhu, K.; Yan, L.; Lu, Y.; Lin, Y.; Cong, X.; Tang, X.; Qian, B.; et al. ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs. In Proceedings of the The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024; 2024. Available online: https://openreview.net/.
- Contributors, M.C.P. Specification - Model Context Protocol. 2025. [Google Scholar]
- Hong, S.; Zhuge, M.; Chen, J.; Zheng, X.; Cheng, Y.; Wang, J.; Zhang, C.; Wang, Z.; Yau, S.K.S.; Lin, Z.; et al. MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework. In Proceedings of the The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024; 2024. Available online: https://openreview.net/.
- Qian, C.; Liu, W.; Liu, H.; Chen, N.; Dang, Y.; Li, J.; Yang, C.; Chen, W.; Su, Y.; Cong, X.; et al. ChatDev: Communicative Agents for Software Development. In Proceedings of the Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics; Bangkok, Thailand, Ku, L., Martins, A., Srikumar, V., Eds.; Association for Computational Linguistics, 2024; Volume 1, pp. 15174–15186. [Google Scholar] [CrossRef]
- Li, G.; Hammoud, H.; Itani, H.; Khizbullin, D.; Ghanem, B. CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society. In Proceedings of the Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023; New Orleans, LA, USA, Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., Levine, S., Eds.; 10 - 16 December 2023. [Google Scholar]
- Wu, Q.; Bansal, G.; Zhang, J.; Wu, Y.; Li, B.; Zhu, E.; Jiang, L.; Zhang, X.; Zhang, S.; Liu, J.; et al. AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation. arXiv 2023, arXiv:cs. [Google Scholar]
- LangChain. langgraph: Build resilient agents. 2024. [Google Scholar]
- Contributors, A. Agent2Agent (A2A) Protocol Specification. 2025. [Google Scholar]
- Anthropic. Hooks reference - Claude Code Docs. 2025.
- OpenAI. Hooks | ChatGPT Learn. 2025. [Google Scholar]
- LangChain. How Middleware Lets You Customize Your Agent Harness; 2026. [Google Scholar]
- Chen, Z.; Kang, M.; Li, B. ShieldAgent: Shielding Agents via Verifiable Safety Policy Reasoning. In Proceedings of the Forty-second International Conference on Machine Learning, ICML 2025 PMLR / OpenReview.net; Vancouver, BC, Canada, Singh, A., Fazel, M., Hsu, D., Lacoste-Julien, S., Berkenkamp, F., Maharaj, T., Wagstaff, K., Zhu, J., Eds.; Proceedings of Machine Learning Research; 13-19 July 2025; Vol. 267. [Google Scholar]
- Gou, Z.; Shao, Z.; Gong, Y.; Shen, Y.; Yang, Y.; Duan, N.; Chen, W. CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing. In Proceedings of the The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024; 2024. Available online: https://openreview.net/.
- Andriushchenko, M.; Souly, A.; Dziemian, M.; Duenas, D.; Lin, M.; Wang, J.; Hendrycks, D.; Zou, A.; Kolter, J.Z.; Fredrikson, M.; et al. AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents. In Proceedings of the The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025; 2025. Available online: https://openreview.net/.
- Levy, I.; Wiesel, B.; Marreed, S.; Oved, A.; Yaeli, A.; Shlomov, S. ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents. CoRR 2024, abs/2410.06703, 2410.06703. [Google Scholar] [CrossRef]
- OWASP GenAI Security Project. OWASP Top 10 for Agentic Applications for 2026, 2025.
- DeepSeek-AI. DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model. CoRR 2024, abs/2405.04434, [2405.04434. [Google Scholar] [CrossRef]
- DeepSeek-AI. DeepSeek-V3 Technical Report. CoRR 2024, abs/2412.19437, [2412.19437. [Google Scholar] [CrossRef]
- Yuan, J.; Gao, H.; Dai, D.; Luo, J.; Zhao, L.; Zhang, Z.; Xie, Z.; Wei, Y.; Wang, L.; Xiao, Z.; et al. Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention. In Proceedings of the Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics; Vienna, Austria, Che, W., Nabende, J., Shutova, E., Pilehvar, M.T., Eds.; 27 July 2025; Volume 1, pp. 23078–23097. [Google Scholar] [CrossRef]
- Dao, T.; Fu, D.Y.; Ermon, S.; Rudra, A.; Ré, C. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness. In Proceedings of the Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022; New Orleans, LA, USA, Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., Oh, A., Eds.; 28 November 2022. [Google Scholar]
- Liu, J.; Su, J.; Yao, X.; Jiang, Z.; Lai, G.; Du, Y.; Qin, Y.; Xu, W.; Lu, E.; Yan, J.; et al. Muon is Scalable for LLM Training. CoRR 2025, abs/2502.16982, [2502.16982. [Google Scholar] [CrossRef]
- Gu, A.; Dao, T. Mamba: Linear-Time Sequence Modeling with Selective State Spaces. CoRR 2023, abs/2312.00752, [2312.00752. [Google Scholar] [CrossRef]
- Tao, Z.; Wu, J.; Yin, W.; Zhang, J.; Li, B.; Shen, H.; Li, K.; Zhang, L.; Wang, X.; Jiang, Y.; et al. WebShaper: Agentically Data Synthesizing via Information-Seeking Formalization. CoRR 2025, abs/2507.15061, 2507.15061. [Google Scholar] [CrossRef]
- Pan, J.; Wang, X.; Neubig, G.; Jaitly, N.; Ji, H.; Suhr, A.; Zhang, Y. Training Software Engineering Agents and Verifiers with SWE-Gym. In Proceedings of the Forty-second International Conference on Machine Learning, ICML 2025; Vancouver, BC, Canada, Singh, A., Fazel, M., Hsu, D., Lacoste-Julien, S., Berkenkamp, F., Maharaj, T., Wagstaff, K., Zhu, J., Eds.; PMLR / OpenReview.net; Proceedings of Machine Learning Research; 13-19 July 2025; Vol. 267. [Google Scholar]
- Yang, J.; Lieret, K.; Jimenez, C.E.; Wettig, A.; Khandpur, K.; Zhang, Y.; Hui, B.; Press, O.; Schmidt, L.; Yang, D. SWE-smith: Scaling Data for Software Engineering Agents. In Proceedings of the Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025; San Diego, CA, USA, December 2-7, 2025 / Mexico City, Mexico, Belgrave, D., Zhang, C., Montoya, L.N., Lin, H., Pascanu, R., Koniusz, P., Ghassemi, M., Chen, N., Ruíz, I.V.M., Loaiza-Bonilla, A., Eds.; 30 November 2025. [Google Scholar]
- Xu, Z.; Soria, A.M.; Tan, S.; Roy, A.; Agrawal, A.S.; Poovendran, R.; Panda, R. TOUCAN: Synthesizing 1.5M Tool-Agentic Data from Real-World MCP Environments. CoRR 2025, abs/2510.01179, 2510.01179. [Google Scholar] [CrossRef]
- Song, X.; Chang, H.; Dong, G.; Zhu, Y.; Wen, J.; Dou, Z. EnvScaler: Scaling Tool-Interactive Environments for LLM Agent via Programmatic Synthesis. In Proceedings of the Findings of the Association for Computational Linguistics, ACL 2026; San Diego, California, United States, Liakata, M., Moreira, V.P., Zhang, J., Jurgens, D., Eds.; Association for Computational Linguistics, 2-7 July 2026; Volume 2026, pp. 8326–8357. [Google Scholar]
- Team, Q. Qwen3 Technical Report. CoRR 2025, abs/2505.09388, 2505.09388. [Google Scholar] [CrossRef]
- Cao, R.; Chen, M.; Chen, J.; Cui, Z.; Feng, Y.; Hui, B.; Jing, Y.; Li, K.; Li, M.; Lin, J.; et al. Qwen3-Coder-Next Technical Report. CoRR 2026, abs/2603.00729, [2603.00729. [Google Scholar] [CrossRef]
- GLM. GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models. CoRR 2025, abs/2508.06471, 2508.06471. [Google Scholar] [CrossRef]
- Team, K. Kimi K2: Open Agentic Intelligence. CoRR 2025, abs/2507.20534, 2507.20534. [Google Scholar] [CrossRef]
- Wang, S.; Ouyang, X.; Xu, T.; Hu, Y.; Liu, J.; Chen, G.; Zhang, T.; Zheng, J.; Yang, K.; Ren, X.; et al. OPUS: Towards Efficient and Principled Data Selection in Large Language Model Pre-training in Every Iteration. CoRR 2026, abs/2602.05400, 2602.05400. [Google Scholar] [CrossRef]
- Zeng, A.; Liu, M.; Lu, R.; Wang, B.; Liu, X.; Dong, Y.; Tang, J. AgentTuning: Enabling Generalized Agent Abilities for LLMs. In Proceedings of the Findings of the Association for Computational Linguistics, ACL 2024, Bangkok, Thailand and virtual meeting Association for Computational Linguistics; Ku, L., Martins, A., Srikumar, V., Eds.; 11-16 August 2024; Vol. ACL 2024, Findings of ACL, pp. 3053–3077. [Google Scholar] [CrossRef]
- Liu, W.; Huang, X.; Zeng, X.; Hao, X.; Yu, S.; Li, D.; Wang, S.; Gan, W.; Liu, Z.; Yu, Y.; et al. ToolACE: Winning the Points of LLM Function Calling. In Proceedings of the The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025; 2025. Available online: https://openreview.net/.
- Wu, J.; Li, B.; Fang, R.; Yin, W.; Zhang, L.; Wang, Z.; Tao, Z.; Zhang, D.; Xi, Z.; Tang, R.; et al. WebDancer: Towards Autonomous Information Seeking Agency. In Proceedings of the Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025; San Diego, CA, USA, December 2-7, 2025 / Mexico City, Mexico, Belgrave, D., Zhang, C., Montoya, L.N., Lin, H., Pascanu, R., Koniusz, P., Ghassemi, M., Chen, N., Ruíz, I.V.M., Loaiza-Bonilla, A., Eds.; 30 November 2025. [Google Scholar]
- Shao, Z.; Wang, P.; Zhu, Q.; Xu, R.; Song, J.; Zhang, M.; Li, Y.K.; Wu, Y.; Guo, D. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models. CoRR 2024, abs/2402.03300, 2402.03300. [Google Scholar] [CrossRef]
- Qian, C.; Acikgoz, E.C.; He, Q.; Wang, H.; Chen, X.; Hakkani-Tur, D.; Tur, G.; Ji, H. ToolRL: Reward is All Tool Learning Needs. In Proceedings of the Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025; San Diego, CA, USA, December 2-7, 2025 / Mexico City, Mexico, Belgrave, D., Zhang, C., Montoya, L.N., Lin, H., Pascanu, R., Koniusz, P., Ghassemi, M., Chen, N., Ruíz, I.V.M., Loaiza-Bonilla, A., Eds.; 30 November 2025. [Google Scholar]
- Wang, Z.; Wang, K.; Wang, Q.; Zhang, P.; Li, L.; Yang, Z.; Jin, X.; Yu, K.; Nguyen, M.N.; Liu, L.; et al. RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning. CoRR 2025, abs/2504.20073, 2504.20073. [Google Scholar] [CrossRef]
- Jin, B.; Zeng, H.; Yue, Z.; Wang, D.; Zamani, H.; Han, J. Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning. CoRR 2025, abs/2503.09516, 2503.09516. [Google Scholar] [CrossRef]
- Zhang, Y.; Zeng, Y.; Li, Q.; Hu, Z.; Han, K.; Zuo, W. Tool-R1: Sample-Efficient Reinforcement Learning for Agentic Tool Use. CoRR 2025, abs/2509.12867, 2509.12867. [Google Scholar] [CrossRef]
- Dong, G.; Mao, H.; Ma, K.; Bao, L.; Chen, Y.; Wang, Z.; Chen, Z.; Du, J.; Wang, H.; Zhang, F.; et al. Agentic Reinforced Policy Optimization. CoRR 2025, abs/2507.19849, [2507.19849. [Google Scholar] [CrossRef]
- Agarwal, R.; Vieillard, N.; Zhou, Y.; Stanczyk, P.; Garea, S.R.; Geist, M.; Bachem, O. On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes. In Proceedings of the The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024; 2024. Available online: https://openreview.net/.
- Tunstall, L.; Beeching, E.; Lambert, N.; Rajani, N.; Rasul, K.; Belkada, Y.; Huang, S.; von Werra, L.; Fourrier, C.; Habib, N.; et al. Zephyr: Direct Distillation of LM Alignment. CoRR 2023, abs/2310.16944, 2310.16944. [Google Scholar] [CrossRef]
- Lyu, Y.; Wang, C.; Huang, J.; Xu, T. From Correction to Mastery: Reinforced Distillation of Large Language Model Agents. CoRR 2025, abs/2509.14257, 2509.14257. [Google Scholar] [CrossRef]
- Gu, Y.; Dong, L.; Wei, F.; Huang, M. MiniLLM: On-Policy Distillation of Large Language Models. arXiv 2026, arXiv:cs. [Google Scholar]
- Wang, J.; Liu, Y.; Chen, J.; Hu, X.; Zhang, Q.; Cao, Y.; Wang, J.; Yang, H.; Xie, Y.; Chen, Q. MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate. CoRR 2026, abs/2605.01347, [2605.01347. [Google Scholar] [CrossRef]
- Yuan, Q.; Lou, J.; Yu, X.; Lin, H.; Sun, L.; Han, X.; Lu, Y. Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation. CoRR 2026, abs/2605.18740, 2605.18740. [Google Scholar] [CrossRef]
- Zhao, A.; Huang, D.; Xu, Q.; Lin, M.; Liu, Y.; Huang, G. ExpeL: LLM Agents Are Experiential Learners. In Proceedings of the Thirty-Eighth AAAI Conference on Artificial Intelligence, AAAI 2024, Thirty-Sixth Conference on Innovative Applications of Artificial Intelligence, IAAI 2024, Fourteenth Symposium on Educational Advances in Artificial Intelligence, EAAI 2024, February 20-27, 2024; Vancouver, Canada, Wooldridge, M.J., Dy, J.G., Natarajan, S., Eds.; AAAI Press, 2024; pp. 19632–19642. [Google Scholar] [CrossRef]
- Zhai, Y.; Tao, S.; Chen, C.; Zou, A.; Chen, Z.; Fu, Q.; Mai, S.; Yu, L.; Deng, J.; Cao, Z.; et al. AgentEvolver: Towards Efficient Self-Evolving Agent System. CoRR 2025, abs/2511.10395, 2511.10395. [Google Scholar] [CrossRef]
- Zhao, A.; Wu, Y.; Wu, T.; Xu, Q.; Yue, Y.; Lin, M.; Wang, S.; Wu, Q.; Zheng, Z.; Huang, G. Absolute Zero: Reinforced Self-play Reasoning with Zero Data. In Proceedings of the Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025; San Diego, CA, USA, December 2-7, 2025 / Mexico City, Mexico, Belgrave, D., Zhang, C., Montoya, L.N., Lin, H., Pascanu, R., Koniusz, P., Ghassemi, M., Chen, N., Ruíz, I.V.M., Loaiza-Bonilla, A., Eds.; 30 November 2025. [Google Scholar]
- Hu, S.; Lu, C.; Clune, J. Automated Design of Agentic Systems. In Proceedings of the The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025; 2025. Available online: https://openreview.net/.
- Zhang, J.; Hu, S.; Lu, C.; Lange, R.T.; Clune, J. Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents. CoRR 2025, abs/2505.22954, [2505.22954. [Google Scholar] [CrossRef]
- Zhang, Y.; Ruan, H.; Fan, Z.; Roychoudhury, A. AutoCodeRover: Autonomous Program Improvement. In Proceedings of the Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis, ISSTA 2024; Vienna, Austria, Christakis, M., Pradel, M., Eds.; ACM, 16-20 September 2024; Volume 2024, pp. 1592–1604. [Google Scholar] [CrossRef]
- Xia, C.S.; Deng, Y.; Dunn, S.; Zhang, L. Agentless: Demystifying LLM-based Software Engineering Agents. CoRR 2024, abs/2407.01489, 2407.01489. [Google Scholar] [CrossRef]
- Li, X.; Dong, G.; Jin, J.; Zhang, Y.; Zhou, Y.; Zhu, Y.; Zhang, P.; Dou, Z. Search-o1: Agentic Search-Enhanced Large Reasoning Models. In Proceedings of the Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025; Suzhou, China, Christodoulopoulos, C., Chakraborty, T., Rose, C., Peng, V., Eds.; Association for Computational Linguistics, 4-9 November 2025; Volume 2025, pp. 5420–5438. [Google Scholar] [CrossRef]
- He, G.; Yang, Z.; Liu, J.; Xu, B.; Hou, L.; Li, J. WebSeer: Training Deeper Search Agents through Reinforcement Learning with Self-Reflection. CoRR 2025, abs/2510.18798, [2510.18798. [Google Scholar] [CrossRef]
- Li, S.; Tang, Y.; Wang, Y.; Li, P.; Chen, X. ReSeek: A Self-Correcting Framework for Search Agents with Instructive Rewards. CoRR 2025, abs/2510.00568, 2510.00568. [Google Scholar] [CrossRef]
- Hu, Y.; Ma, R.; Fan, Y.; Shi, J.; Cao, Z.; Zhou, Y.; Yuan, J.; Zhang, S.; Feng, S.; Yan, X.; et al. FlowSearch: Advancing Deep Research with Dynamic Structured Knowledge Flow. In Proceedings of the Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2026; San Diego, California, United States, Liakata, M., Moreira, V.P., Zhang, J., Jurgens, D., Eds.; Association for Computational Linguistics, 2-7 July 2026; Volume 2026, pp. 21212–21245. [Google Scholar]
- Li, X.; Jin, J.; Dong, G.; Qian, H.; Wu, Y.; Wen, J.; Zhu, Y.; Dou, Z. WebThinker: Empowering Large Reasoning Models with Deep Research Capability. In Proceedings of the Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025; San Diego, CA, USA, December 2-7, 2025 / Mexico City, Mexico, Belgrave, D., Zhang, C., Montoya, L.N., Lin, H., Pascanu, R., Koniusz, P., Ghassemi, M., Chen, N., Ruíz, I.V.M., Loaiza-Bonilla, A., Eds.; 30 November 2025. [Google Scholar]
- Deng, X.; Gu, Y.; Zheng, B.; Chen, S.; Stevens, S.; Wang, B.; Sun, H.; Su, Y. Mind2Web: Towards a Generalist Agent for the Web. In Proceedings of the Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023; New Orleans, LA, USA, Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., Levine, S., Eds.; 10 - 16 December 2023. [Google Scholar]
- Qin, Y.; Ye, Y.; Fang, J.; Wang, H.; Liang, S.; Tian, S.; Zhang, J.; Li, J.; Li, Y.; Huang, S.; et al. UI-TARS: Pioneering Automated GUI Interaction with Native Agents. CoRR 2025, abs/2501.12326, 2501.12326. [Google Scholar] [CrossRef]
- Rawles, C.; Clinckemaillie, S.; Chang, Y.; Waltz, J.; Lau, G.; Fair, M.; Li, A.; Bishop, W.E.; Li, W.; Campbell-Ajala, F.; et al. AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents. In Proceedings of the The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025; 2025. Available online: https://openreview.net/.
- Wang, X.; Zhang, Y.; Zohar, O.; Yeung-Levy, S. VideoAgent: Long-Form Video Understanding with Large Language Model as Agent. In Proceedings of the Computer Vision - ECCV 2024 - 18th European Conference Proceedings, Part LXXX; Milan, Italy, Leonardis, A., Ricci, E., Roth, S., Russakovsky, O., Sattler, T., Varol, G., Eds.; Springer; Lecture Notes in Computer Science; 29 September 2024; Vol. 15138, pp. 58–76. [Google Scholar] [CrossRef]
- Zhang, X.; Jia, Z.; Guo, Z.; Li, J.; Li, B.; Li, H.; Lu, Y. Deep Video Discovery: Agentic Search with Tool Use for Long-form Video Understanding. In Proceedings of the Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025; San Diego, CA, USA, December 2-7, 2025 / Mexico City, Mexico, Belgrave, D., Zhang, C., Montoya, L.N., Lin, H., Pascanu, R., Koniusz, P., Ghassemi, M., Chen, N., Ruíz, I.V.M., Loaiza-Bonilla, A., Eds.; 30 November 2025. [Google Scholar]
- Lin, J.; Wu, J.; Liu, J.; Sun, X.; Wang, Z.; Yu, X.; Luo, J.; Liu, Z.; Barsoum, E. VideoSeek: Long-Horizon Video Agent with Tool-Guided Seeking. CoRR 2026, abs/2603.20185, [2603.20185. [Google Scholar] [CrossRef]
- Lin, H.; Shi, Y.; Geng, T.; Zhao, W.; Wang, W.; Singh, R.P. Agent-Omni: Test-Time Multimodal Reasoning via Model Coordination for Understanding Anything. CoRR 2025, abs/2511.02834, 2511.02834. [Google Scholar] [CrossRef]
- Li, X.; Jiao, W.; Jin, J.; Wang, S.; Dong, G.; Jin, J.; Wang, H.; Wang, Y.; Wen, J.; Lu, Y.; et al. OmniGAIA: Towards Native Omni-Modal AI Agents. CoRR 2026, abs/2602.22897, [2602.22897. [Google Scholar] [CrossRef]
- Fourney, A.; Bansal, G.; Mozannar, H.; Tan, C.; Salinas, E.; Zhu, E.; Niedtner, F.; Proebsting, G.; Bassman, G.; Gerrits, J.; et al. Magentic-One: A Generalist Multi-Agent System for Solving Complex Tasks. CoRR 2024, abs/2411.04468, 2411.04468. [Google Scholar] [CrossRef]
- Liu, X.; Yu, H.; Zhang, H.; Xu, Y.; Lei, X.; Lai, H.; Gu, Y.; Ding, H.; Men, K.; Yang, K.; et al. AgentBench: Evaluating LLMs as Agents. In Proceedings of the The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024; 2024. Available online: https://openreview.net/.
- Yao, S.; Shinn, N.; Razavi, P.; Narasimhan, K. τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. CoRR 2024, abs/2406.12045, [2406.12045. [Google Scholar] [CrossRef]
- Chan, J.S.; Chowdhury, N.; Jaffe, O.; Aung, J.; Sherburn, D.; Mays, E.; Starace, G.; Liu, K.; Maksin, L.; Patwardhan, T.; et al. MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering. In Proceedings of the The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025; 2025. Available online: https://openreview.net/.
- Starace, G.; Jaffe, O.; Sherburn, D.; Aung, J.; Chan, J.S.; Maksin, L.; Dias, R.; Mays, E.; Kinsella, B.; Thompson, W.; et al. PaperBench: Evaluating AI’s Ability to Replicate AI Research. In Proceedings of the Forty-second International Conference on Machine Learning, ICML 2025 PMLR / OpenReview.net; Vancouver, BC, Canada, Singh, A., Fazel, M., Hsu, D., Lacoste-Julien, S., Berkenkamp, F., Maharaj, T., Wagstaff, K., Zhu, J., Eds.; Proceedings of Machine Learning Research; 13-19 July 2025; Vol. 267. [Google Scholar]
- Merrill, M.A.; Shaw, A.G.; Carlini, N.; Li, B.; Raj, H.; Bercovich, I.; Shi, L.; Shin, J.Y.; Walshe, T.; Buchanan, E.K.; et al. Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces. CoRR 2026, abs/2601.11868, 2601.11868. [Google Scholar] [CrossRef]
- Costa, I. AgentSpawn: Adaptive Multi-Agent Collaboration Through Dynamic Spawning for Long-Horizon Code Generation. CoRR 2026, abs/2602.07072, 2602.07072. [Google Scholar] [CrossRef]
- Zhang, G.; Yu, H.; Yang, K.; Wu, B.; Huang, F.; Li, Y.; Yan, S. EvoRoute: Experience-Driven Self-Routing LLM Agent Systems. In Proceedings of the Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics; San Diego, California, United States, Liakata, M., Moreira, V.P., Zhang, J., Jurgens, D., Eds.; Association for Computational Linguistics, 2-7 July 2026; Volume 1, pp. 38213–38225. [Google Scholar]
- Novikov, A.; Vu, N.; Eisenberger, M.; Dupont, E.; Huang, P.; Wagner, A.Z.; Shirobokov, S.; Kozlovskii, B.; Ruiz, F.J.R.; Mehrabian, A.; et al. AlphaEvolve: A coding agent for scientific and algorithmic discovery. CoRR 2025, abs/2506.13131, 2506.13131. [Google Scholar] [CrossRef]
- Huang, C.; Yu, W.; Wang, X.; Zhang, H.; Li, Z.; Li, R.; Huang, J.; Mi, H.; Yu, D. R-Zero: Self-Evolving Reasoning LLM from Zero Data. CoRR 2025, abs/2508.05004, 2508.05004. [Google Scholar] [CrossRef]
- Xu, F.F.; Song, Y.; Li, B.; Tang, Y.; Jain, K.; Bao, M.; Wang, Z.Z.; Zhou, X.; Guo, Z.; Cao, M.; et al. TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks. CoRR 2024, abs/2412.14161, 2412.14161. [Google Scholar] [CrossRef]
- Kim, M.J.; Pertsch, K.; Karamcheti, S.; Xiao, T.; Balakrishna, A.; Nair, S.; Rafailov, R.; Foster, E.P.; Sanketi, P.R.; Vuong, Q.; et al. OpenVLA: An Open-Source Vision-Language-Action Model. In Proceedings of the Conference on Robot Learning PMLR; Munich, Germany, Agrawal, P., Kroemer, O., Burgard, W., Eds.; Proceedings of Machine Learning Research; 6-9 November 2024; Vol. 270, pp. 2679–2713. [Google Scholar]
- Chen, L.; Zaharia, M.; Zou, J. FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance. In Trans. Mach. Learn. Res.; 2024. [Google Scholar]
- Ong, I.; Almahairi, A.; Wu, V.; Chiang, W.; Wu, T.; Gonzalez, J.E.; Kadous, M.W.; Stoica, I. RouteLLM: Learning to Route LLMs with Preference Data. CoRR 2024, abs/2406.18665, [2406.18665. [Google Scholar] [CrossRef]
- Muennighoff, N.; Yang, Z.; Shi, W.; Li, X.L.; Fei-Fei, L.; Hajishirzi, H.; Zettlemoyer, L.; Liang, P.; Candès, E.J.; Hashimoto, T. s1: Simple test-time scaling. In Proceedings of the Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025 Association for Computational Linguistics; Suzhou, China, Christodoulopoulos, C., Chakraborty, T., Rose, C., Peng, V., Eds.; 4-9 November 2025; Volume 2025, pp. 20275–20321. [Google Scholar] [CrossRef]
- Lin, Y.; Wang, Z.; Liu, M.; Shan, Y.; Bai, L.; Zhang, J.; Jin, X.; Chen, B.; Su, J.; Wang, X.; et al. BAGEN: Are LLM Agents Budget-Aware? CoRR 2026. abs/2606.00198 2606.00198. [CrossRef]
- Gao, P.; Tian, Z.; Meng, X.; Wang, X.; Hu, R.; Xiao, Y.; Liu, Y.; Zhang, Z.; Chen, J.; Gao, C.; et al. Trae Agent: An LLM-based Agent for Software Engineering with Test-time Scaling. CoRR 2025, abs/2507.23370, 2507.23370. [Google Scholar] [CrossRef]
- Ichter, B.; Brohan, A.; Chebotar, Y.; Finn, C.; Hausman, K.; Herzog, A.; Ho, D.; Ibarz, J.; Irpan, A.; Jang, E.; et al. Do As I Can, Not As I Say: Grounding Language in Robotic Affordances. In Proceedings of the Conference on Robot Learning, CoRL 2022 PMLR; Auckland, New Zealand, Liu, K., Kulic, D., Ichnowski, J., Eds.; Proceedings of Machine Learning Research; 14-18 December 2022; Vol. 205, pp. 287–318. [Google Scholar]
- Mou, Y.; Xue, Z.; Li, L.; Liu, P.; Zhang, S.; Ye, W.; Shao, J. ToolSafe: Enhancing Tool Invocation Safety of LLM-based agents via Proactive Step-level Guardrail and Feedback. In Proceedings of the Findings of the Association for Computational Linguistics, ACL 2026; San Diego, California, United States, Liakata, M., Moreira, V.P., Zhang, J., Jurgens, D., Eds.; Association for Computational Linguistics, 2-7 July 2026; pp. 37125–37153. [Google Scholar]
- Zhang, H.; Liu, X.; Lv, B.; Sun, X.; Jing, B.; Iong, I.L.; Hou, Z.; Qi, Z.; Lai, H.; Xu, Y.; et al. AgentRL: Scaling Agentic Reinforcement Learning with a Multi-Turn, Multi-Task Framework. CoRR 2025, abs/2510.04206, 2510.04206. [Google Scholar] [CrossRef]
- Wang, L.; Ma, C.; Feng, X.; Zhang, Z.; Yang, H.; Zhang, J.; Chen, Z.; Tang, J.; Chen, X.; Lin, Y.; et al. A survey on large language model based autonomous agents. Front. Comput. Sci. 2024, 18, 186345. [Google Scholar] [CrossRef]
- Hu, Y.; Liu, S.; Yue, Y.; Zhang, G.; Liu, B.; Zhu, F.; Lin, J.; Guo, H.; Dou, S.; Xi, Z.; et al. Memory in the Age of AI Agents. CoRR 2025, abs/2512.13564, 2512.13564. [Google Scholar] [CrossRef]
- Kaelbling, L.P.; Littman, M.L.; Cassandra, A.R. Planning and Acting in Partially Observable Stochastic Domains. Artif. Intell. 1998, 101, 99–134. [Google Scholar] [CrossRef]
- Sutton, R.S.; Barto, A.G. Reinforcement Learning: An Introduction, 2nd ed.; MIT Press: Cambridge, MA, 2018. [Google Scholar]
- Sutton, R.S.; Silver, D. Welcome to the Era of Experience, 2025.
- Wang, G.; Xie, Y.; Jiang, Y.; Mandlekar, A.; Xiao, C.; Zhu, Y.; Fan, L.; Anandkumar, A. Voyager: An Open-Ended Embodied Agent with Large Language Models. In Trans. Mach. Learn. Res.; 2024. [Google Scholar]
- Ouyang, S.; Yan, J.; Hsu, I.; Chen, Y.; Jiang, K.; Wang, Z.; Han, R.; Le, L.T.; Daruki, S.; Tang, X.; et al. ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory. CoRR 2025, abs/2509.25140, [2509.25140. [Google Scholar] [CrossRef]
- METR. Time Horizon 1.1. 2026.
- Whitfill, P.; Snodin, B.; Becker, J. Forecasting AI Time Horizon Under Compute Slowdowns. CoRR 2025, abs/2511.19492, 2511.19492. [Google Scholar] [CrossRef]
- Hoffmann, J.; Borgeaud, S.; Mensch, A.; Buchatskaya, E.; Cai, T.; Rutherford, E.; de Las Casas, D.; Hendricks, L.A.; Welbl, J.; Clark, A.; et al. Training Compute-Optimal Large Language Models. CoRR 2022, abs/2203.15556, [2203.15556. [Google Scholar] [CrossRef]
- Feng, K.J.K.; McDonald, D.W.; Zhang, A.X. Levels of Autonomy for AI Agents. CoRR 2025, abs/2506.12469, 2506.12469. [Google Scholar] [CrossRef]
- OpenAI. Introducing Operator. 2025. [Google Scholar]
- Brown, T.B.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. Language Models are Few-Shot Learners. In Proceedings of the Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020; virtual, Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., Lin, H., Eds.; 6-12 December 2020. [Google Scholar]
- Liu, P.; Yuan, W.; Fu, J.; Jiang, Z.; Hayashi, H.; Neubig, G. Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing. ACM Comput. Surv. 2023, 55, 195:1–195:35. [Google Scholar] [CrossRef]
- Dong, Q.; Li, L.; Dai, D.; Zheng, C.; Ma, J.; Li, R.; Xia, H.; Xu, J.; Wu, Z.; Chang, B.; et al. A Survey on In-context Learning. In Proceedings of the Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024; Miami, FL, USA, Al-Onaizan, Y., Bansal, M., Chen, Y., Eds.; Association for Computational Linguistics, 12-16 November 2024; pp. 1107–1128. [Google Scholar] [CrossRef]
- Nye, M.I.; Andreassen, A.J.; Gur-Ari, G.; Michalewski, H.; Austin, J.; Bieber, D.; Dohan, D.; Lewkowycz, A.; Bosma, M.; Luan, D.; et al. Show Your Work: Scratchpads for Intermediate Computation with Language Models. CoRR 2021, abs/2112.00114. [Google Scholar]
- Kojima, T.; Gu, S.S.; Reid, M.; Matsuo, Y.; Iwasawa, Y. Large Language Models are Zero-Shot Reasoners. In Proceedings of the Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022; New Orleans, LA, USA, Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., Oh, A., Eds.; 28 November 2022. [Google Scholar]
- Zhang, Z.; Zhang, A.; Li, M.; Smola, A. Automatic Chain of Thought Prompting in Large Language Models. In Proceedings of the The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023; 2023. Available online: https://openreview.net/.
- Fu, Y.; Peng, H.; Sabharwal, A.; Clark, P.; Khot, T. Complexity-Based Prompting for Multi-step Reasoning. In Proceedings of the The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023; 2023. Available online: https://openreview.net/.
- Khot, T.; Trivedi, H.; Finlayson, M.; Fu, Y.; Richardson, K.; Clark, P.; Sabharwal, A. Decomposed Prompting: A Modular Approach for Solving Complex Tasks. In Proceedings of the The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023; 2023. Available online: https://openreview.net/.
- Wang, X.; Wei, J.; Schuurmans, D.; Le, Q.V.; Chi, E.H.; Narang, S.; Chowdhery, A.; Zhou, D. Self-Consistency Improves Chain of Thought Reasoning in Language Models. In Proceedings of the The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023; 2023. Available online: https://openreview.net/.
- Gao, L.; Madaan, A.; Zhou, S.; Alon, U.; Liu, P.; Yang, Y.; Callan, J.; Neubig, G. PAL: Program-aided Language Models. In Proceedings of the International Conference on Machine Learning, ICML 2023 PMLR; Honolulu, Hawaii, USA, Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., Scarlett, J., Eds.; Proceedings of Machine Learning Research; 23-29 July 2023; Vol. 202, pp. 10764–10799. [Google Scholar]
- Chen, W.; Ma, X.; Wang, X.; Cohen, W.W. Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks. In Trans. Mach. Learn. Res.; 2023. [Google Scholar]
- Huang, W.; Xia, F.; Xiao, T.; Chan, H.; Liang, J.; Florence, P.; Zeng, A.; Tompson, J.; Mordatch, I.; Chebotar, Y.; et al. Inner Monologue: Embodied Reasoning through Planning with Language Models. In Proceedings of the Conference on Robot Learning, CoRL 2022 PMLR; Auckland, New Zealand, Liu, K., Kulic, D., Ichnowski, J., Eds.; Proceedings of Machine Learning Research; 14-18 December 2022; Vol. 205, pp. 1769–1782. [Google Scholar]
- Shridhar, M.; Yuan, X.; Côté, M.; Bisk, Y.; Trischler, A.; Hausknecht, M.J. ALFWorld: Aligning Text and Embodied Environments for Interactive Learning. In Proceedings of the 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021; 2021. Available online: https://openreview.net/.
- Sun, H.; Zhuang, Y.; Kong, L.; Dai, B.; Zhang, C. AdaPlanner: Adaptive Planning from Feedback with Language Models. In Proceedings of the Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023; New Orleans, LA, USA, Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., Levine, S., Eds.; 10 - 16 December 2023. [Google Scholar]
- Prasad, A.; Koller, A.; Hartmann, M.; Clark, P.; Sabharwal, A.; Bansal, M.; Khot, T. ADaPT: As-Needed Decomposition and Planning with Language Models. In Proceedings of the Findings of the Association for Computational Linguistics: NAACL 2024 Association for Computational Linguistics; Mexico City, Mexico, Duh, K., Gómez-Adorno, H., Bethard, S., Eds.; Findings of ACL; 16-21 June 2024; Vol. NAACL 2024, pp. 4226–4252. [Google Scholar] [CrossRef]
- Wang, Z.; Cai, S.; Liu, A.; Ma, X.; Liang, Y. Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents. CoRR 2023, abs/2302.01560, 2302.01560. [Google Scholar] [CrossRef]
- Lu, Y.; Bartolo, M.; Moore, A.; Riedel, S.; Stenetorp, P. Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity. In Proceedings of the Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics; Dublin, Ireland, Muresan, S., Nakov, P., Villavicencio, A., Eds.; Association for Computational Linguistics, 22-27 May 2022; Volume 1, pp. 8086–8098. [Google Scholar] [CrossRef]
- Zhao, Z.; Wallace, E.; Feng, S.; Klein, D.; Singh, S. Calibrate Before Use: Improving Few-shot Performance of Language Models. In Proceedings of the Proceedings of the 38th International Conference on Machine Learning, ICML 2021 PMLR; Virtual Event, Meila, M., Zhang, T., Eds.; Proceedings of Machine Learning Research; 18-24 July 2021; Vol. 139, pp. 12697–12706. [Google Scholar]
- Min, S.; Lyu, X.; Holtzman, A.; Artetxe, M.; Lewis, M.; Hajishirzi, H.; Zettlemoyer, L. Rethinking the Role of Demonstrations: What Makes In-Context Learning Work? In Proceedings of the Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, EMNLP 2022; Abu Dhabi, United Arab Emirates, Goldberg, Y., Kozareva, Z., Zhang, Y., Eds.; Association for Computational Linguistics, 7-11 December 2022; pp. 11048–11064. [Google Scholar] [CrossRef]
- Zhou, Y.; Muresanu, A.I.; Han, Z.; Paster, K.; Pitis, S.; Chan, H.; Ba, J. Large Language Models are Human-Level Prompt Engineers. In Proceedings of the The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023; 2023. Available online: https://openreview.net/.
- Pryzant, R.; Iter, D.; Li, J.; Lee, Y.T.; Zhu, C.; Zeng, M. Automatic Prompt Optimization with "Gradient Descent" and Beam Search. In Proceedings of the Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing; EMNLP 2023, Singapore, Bouamor, H., Pino, J., Bali, K., Eds.; Association for Computational Linguistics, 6-10 December 2023; pp. 7957–7968. [Google Scholar] [CrossRef]
- Reynolds, L.; McDonell, K. Prompt Programming for Large Language Models: Beyond the Few-Shot Paradigm. In Proceedings of the CHI ’21: CHI Conference on Human Factors in Computing Systems Virtual Event / Yokohama Japan; May 8-13, 2021, Extended Abstracts, Kitamura, Y., Quigley, A., Isbister, K., Igarashi, T., Eds.; ACM, 2021; pp. 314:1–314:7. [Google Scholar] [CrossRef]
- Petroni, F.; Rocktäschel, T.; Lewis, P.; Bakhtin, A.; Wu, Y.; Miller, A.H.; Riedel, S. Language Models as Knowledge Bases? arXiv 2019, arXiv:cs. [Google Scholar]
- Roberts, A.; Raffel, C.; Shazeer, N. How Much Knowledge Can You Pack Into the Parameters of a Language Model? In Proceedings of the Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing; EMNLP 2020, Online, Webber, B., Cohn, T., He, Y., Liu, Y., Eds.; Association for Computational Linguistics, 16-20 November 2020; pp. 5418–5426. [Google Scholar] [CrossRef]
- Kandpal, N.; Deng, H.; Roberts, A.; Wallace, E.; Raffel, C. Large Language Models Struggle to Learn Long-Tail Knowledge. In Proceedings of the International Conference on Machine Learning, ICML 2023 PMLR; Honolulu, Hawaii, USA, Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., Scarlett, J., Eds.; Proceedings of Machine Learning Research; 23-29 July 2023; Vol. 202, pp. 15696–15707. [Google Scholar]
- Khandelwal, U.; Levy, O.; Jurafsky, D.; Zettlemoyer, L.; Lewis, M. Generalization through Memorization: Nearest Neighbor Language Models. In Proceedings of the 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020; 2020. Available online: https://openreview.net/.
- Guu, K.; Lee, K.; Tung, Z.; Pasupat, P.; Chang, M. REALM: Retrieval-Augmented Language Model Pre-Training. CoRR 2020, abs/2002.08909. [Google Scholar]
- Karpukhin, V.; Oguz, B.; Min, S.; Lewis, P.; Wu, L.; Edunov, S.; Chen, D.; Yih, W. Dense Passage Retrieval for Open-Domain Question Answering. In Proceedings of the Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing; EMNLP 2020, Online, Webber, B., Cohn, T., He, Y., Liu, Y., Eds.; Association for Computational Linguistics, 16-20 November 2020; pp. 6769–6781. [Google Scholar] [CrossRef]
- Izacard, G.; Grave, E. Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering. In Proceedings of the Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume; EACL 2021, Online, Merlo, P., Tiedemann, J., Tsarfaty, R., Eds.; Association for Computational Linguistics, 19 - 23 April 2021; pp. 874–880. [Google Scholar] [CrossRef]
- Borgeaud, S.; Mensch, A.; Hoffmann, J.; Cai, T.; Rutherford, E.; Millican, K.; van den Driessche, G.; Lespiau, J.; Damoc, B.; Clark, A.; et al. Improving Language Models by Retrieving from Trillions of Tokens. In Proceedings of the International Conference on Machine Learning, ICML 2022 PMLR; Baltimore, Maryland, USA, Chaudhuri, K., Jegelka, S., Song, L., Szepesvári, C., Niu, G., Sabato, S., Eds.; Proceedings of Machine Learning Research; 17-23 July 2022; Vol. 162, pp. 2206–2240. [Google Scholar]
- Izacard, G.; Lewis, P.; Lomeli, M.; Hosseini, L.; Petroni, F.; Schick, T.; Dwivedi-Yu, J.; Joulin, A.; Riedel, S.; Grave, E. Atlas: Few-shot Learning with Retrieval Augmented Language Models. J. Mach. Learn. Res. 2023, 24, 251:1–251:43. [Google Scholar]
- Trivedi, H.; Balasubramanian, N.; Khot, T.; Sabharwal, A. MuSiQue: Multihop Questions via Single-hop Question Composition, 2022. arXiv arXiv:cs.
- Jiang, Z.; Xu, F.F.; Gao, L.; Sun, Z.; Liu, Q.; Dwivedi-Yu, J.; Yang, Y.; Callan, J.; Neubig, G. Active Retrieval Augmented Generation. In Proceedings of the Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing; EMNLP 2023, Singapore, Bouamor, H., Pino, J., Bali, K., Eds.; Association for Computational Linguistics, 6-10 December 2023; pp. 7969–7992. [Google Scholar] [CrossRef]
- Tang, Q.; Deng, Z.; Lin, H.; Han, X.; Liang, Q.; Sun, L. ToolAlpaca: Generalized Tool Learning for Language Models with 3000 Simulated Cases. CoRR 2023, abs/2306.05301, 2306.05301. [Google Scholar] [CrossRef]
- Guo, Z.; Cheng, S.; Wang, H.; Liang, S.; Qin, Y.; Li, P.; Liu, Z.; Sun, M.; Liu, Y. StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models. In Association for Computational Linguistics; Proceedings of the Findings of the Association for Computational Linguistics, ACL 2024, Bangkok, Thailand and virtual meeting, Ku, L., Martins, A., Srikumar, V., Eds.; Findings of ACL; 2024; Vol. ACL 2024, pp. 11143–11156. [Google Scholar] [CrossRef]
- Nakano, R.; Hilton, J.; Balaji, S.; Wu, J.; Ouyang, L.; Kim, C.; Hesse, C.; Jain, S.; Kosaraju, V.; Saunders, W.; et al. WebGPT: Browser-assisted question-answering with human feedback. CoRR 2021, abs/2112.09332, 2112.09332. [Google Scholar]
- Qin, Y.; Hu, S.; Lin, Y.; Chen, W.; Ding, N.; Cui, G.; Zeng, Z.; Zhou, X.; Huang, Y.; Xiao, C.; et al. Tool Learning with Foundation Models. ACM Comput. Surv. 2025, 57, 101:1–101:40. [Google Scholar] [CrossRef]
- Qu, C.; Dai, S.; Wei, X.; Cai, H.; Wang, S.; Yin, D.; Xu, J.; Wen, J. Tool learning with large language models: a survey. Front. Comput. Sci. 2025, 19, 198343. [Google Scholar] [CrossRef]
- Shen, Y.; Song, K.; Tan, X.; Li, D.; Lu, W.; Zhuang, Y. HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face. In Proceedings of the Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023; New Orleans, LA, USA, Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., Levine, S., Eds.; 10 - 16 December 2023. [Google Scholar]
- Beltagy, I.; Peters, M.E.; Cohan, A. Longformer: The Long-Document Transformer. CoRR 2020 2004, abs/2004.05150. [Google Scholar]
- Zaheer, M.; Guruganesh, G.; Dubey, K.A.; Ainslie, J.; Alberti, C.; Ontañón, S.; Pham, P.; Ravula, A.; Wang, Q.; Yang, L.; et al. Big Bird: Transformers for Longer Sequences. In Proceedings of the Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020; virtual, Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., Lin, H., Eds.; 6-12 December 2020. [Google Scholar]
- Team, G.; Georgiev, P.; Lei, V.I.; Burnell, R.; Bai, L.; Gulati, A.; Tanzer, G.; Vincent, D.; Pan, Z.; Wang, S.; et al. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv 2024, arXiv:2403.05530. [Google Scholar]
- Bai, Y.; Lv, X.; Zhang, J.; Lyu, H.; Tang, J.; Huang, Z.; Du, Z.; Liu, X.; Zeng, A.; Hou, L.; et al. LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding. In Proceedings of the Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics; ACL 2024, Ku, L., Martins, A., Srikumar, V., Eds.; Association for Computational Linguistics, 11-16 August 2024; Volume 1, pp. 3119–3137. [Google Scholar] [CrossRef]
- Hsieh, C.; Sun, S.; Kriman, S.; Acharya, S.; Rekesh, D.; Jia, F.; Zhang, Y.; Ginsburg, B. RULER: What’s the Real Context Size of Your Long-Context Language Models? CoRR 2024. abs/2404.06654 2404.06654. [CrossRef]
- An, C.; Gong, S.; Zhong, M.; Zhao, X.; Li, M.; Zhang, J.; Kong, L.; Qiu, X. L-Eval: Instituting Standardized Evaluation for Long Context Language Models. In Proceedings of the Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics; Bangkok, Thailand, Ku, L., Martins, A., Srikumar, V., Eds.; Association for Computational Linguistics, 11-16 August 2024; Volume 1, pp. 14388–14411. [Google Scholar] [CrossRef]
- Zhang, X.; Chen, Y.; Hu, S.; Xu, Z.; Chen, J.; Hao, M.K.; Han, X.; Thai, Z.L.; Wang, S.; Liu, Z.; et al. ∞Bench: Extending Long Context Evaluation Beyond 100K Tokens. CoRR 2024, abs/2402.13718, [2402.13718. [Google Scholar] [CrossRef]
- Park, J.S.; O’Brien, J.C.; Cai, C.J.; Morris, M.R.; Liang, P.; Bernstein, M.S. Generative Agents: Interactive Simulacra of Human Behavior. In Proceedings of the Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology, UIST 2023; San Francisco, CA, USA, Follmer, S., Han, J., Steimle, J., Riche, N.H., Eds.; ACM, 29 October 2023- 1 November 2023; pp. 2:1–2:22. [Google Scholar] [CrossRef]
- Shi, F.; Chen, X.; Misra, K.; Scales, N.; Dohan, D.; Chi, E.H.; Schärli, N.; Zhou, D. Large Language Models Can Be Easily Distracted by Irrelevant Context. In Proceedings of the International Conference on Machine Learning, ICML 2023 PMLR; Honolulu, Hawaii, USA, Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., Scarlett, J., Eds.; Proceedings of Machine Learning Research; 23-29 July 2023; Vol. 202, pp. 31210–31227. [Google Scholar]
- Anthropic. Effective Context Engineering for AI Agents; 2025. [Google Scholar]
- Jiang, H.; Wu, Q.; Lin, C.; Yang, Y.; Qiu, L. LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models. In Proceedings of the Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing; EMNLP 2023, Singapore, Bouamor, H., Pino, J., Bali, K., Eds.; Association for Computational Linguistics, 6-10 December 2023; pp. 13358–13376. [Google Scholar] [CrossRef]
- Li, Y.; Dong, B.; Guerin, F.; Lin, C. Compressing Context to Enhance Inference Efficiency of Large Language Models. In Proceedings of the Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing; EMNLP 2023, Singapore, Bouamor, H., Pino, J., Bali, K., Eds.; Association for Computational Linguistics, 6-10 December 2023; pp. 6342–6353. [Google Scholar] [CrossRef]
- Yoon, C.; Lee, T.; Hwang, H.; Jeong, M.; Kang, J. CompAct: Compressing Retrieved Documents Actively for Question Answering. In Proceedings of the Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024; Miami, FL, USA, Al-Onaizan, Y., Bansal, M., Chen, Y., Eds.; Association for Computational Linguistics, 12-16 November 2024; pp. 21424–21439. [Google Scholar] [CrossRef]
- Hu, M.; Chen, T.; Chen, Q.; Mu, Y.; Shao, W.; Luo, P. HiAgent: Hierarchical Working Memory Management for Solving Long-Horizon Agent Tasks with Large Language Model. In Proceedings of the Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics; Vienna, Austria, Che, W., Nabende, J., Shutova, E., Pilehvar, M.T., Eds.; 2025; Volume 1, pp. 32779–32798. [Google Scholar] [CrossRef]
- Kang, M.; Chen, W.; Han, D.; Inan, H.A.; Wutschitz, L.; Chen, Y.; Sim, R.; Rajmohan, S. ACON: Optimizing Context Compression for Long-horizon LLM Agents. CoRR 2025, abs/2510.00615, 2510.00615. [Google Scholar] [CrossRef]
- Zhou, Z.; Qu, A.; Wu, Z.; Kim, S.; Prakash, A.; Rus, D.; Zhao, J.; Low, B.K.H.; Liang, P.P. MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents. CoRR 2025, abs/2506.15841, 2506.15841. [Google Scholar] [CrossRef]
- Zhang, Q.; Hu, C.; Upasani, S.; Ma, B.; Hong, F.; Kamanuru, V.; Rainton, J.; Wu, C.; Ji, M.; Li, H.; et al. Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models. CoRR 2025, abs/2510.04618, 2510.04618. [Google Scholar] [CrossRef]
- Yao, S.; Chen, H.; Yang, J.; Narasimhan, K. WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents. In Proceedings of the Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022; New Orleans, LA, USA, Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., Oh, A., Eds.; 28 November 2022. [Google Scholar]
- Nakajima, Y. BabyAGI, 2023.
- Kim, J.; Rhee, S.; Kim, M.; Kim, D.; Lee, S.; Sung, Y.; Jung, K. ReflAct: World-Grounded Decision Making in LLM Agents via Goal-State Reflection. In Proceedings of the Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025; Suzhou, China, Christodoulopoulos, C., Chakraborty, T., Rose, C., Peng, V., Eds.; Association for Computational Linguistics, 4-9 November 2025; pp. 33433–33465. [Google Scholar] [CrossRef]
- Yoran, O.; Amouyal, S.J.; Malaviya, C.; Bogin, B.; Press, O.; Berant, J. AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks? In Proceedings of the Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024; Miami, FL, USA, Al-Onaizan, Y., Bansal, M., Chen, Y., Eds.; Association for Computational Linguistics, 12-16 November 2024; pp. 8938–8968. [Google Scholar] [CrossRef]
- Zheng, B.; Gou, B.; Kil, J.; Sun, H.; Su, Y. GPT-4V(ision) is a Generalist Web Agent, if Grounded. In Proceedings of the Forty-first International Conference on Machine Learning, ICML 2024 PMLR / OpenReview.net; Vienna, Austria, Salakhutdinov, R., Kolter, Z., Heller, K.A., Weller, A., Oliver, N., Scarlett, J., Berkenkamp, F., Eds.; Proceedings of Machine Learning Research; 21-27 July 2024; Vol. 235, pp. 61349–61385. [Google Scholar]
- Trivedi, H.; Balasubramanian, N.; Khot, T.; Sabharwal, A. Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions. In Proceedings of the Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics; Toronto, Canada, Rogers, A., Boyd-Graber, J.L., Okazaki, N., Eds.; Association for Computational Linguistics, 9-14 July 2023; Volume 1, pp. 10014–10037. [Google Scholar] [CrossRef]
- Zheng, Y.; Fu, D.; Hu, X.; Cai, X.; Ye, L.; Lu, P.; Liu, P. DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments. In Proceedings of the Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025 Association for Computational Linguistics; Suzhou, China, Christodoulopoulos, C., Chakraborty, T., Rose, C., Peng, V., Eds.; 4-9 November 2025; pp. 414–431. [Google Scholar] [CrossRef]
- Wang, Z.Z.; Mao, J.; Fried, D.; Neubig, G. Agent Workflow Memory. arXiv 2024, arXiv:cs. [Google Scholar]
- Li, Z.; Xu, S.; Mei, K.; Hua, W.; Rama, B.; Raheja, O.; Wang, H.; Zhu, H.; Zhang, Y. AutoFlow: Automated Workflow Generation for Large Language Model Agents. CoRR 2024, abs/2407.12821, [2407.12821. [Google Scholar] [CrossRef]
- LangChain. Improving Deep Agents with harness engineering; 2026. [Google Scholar]
- Qin, T.; Chen, Q.; Wang, S.; Xing, H.; Zhu, K.; Zhu, H.; Shi, D.; Liu, X.; Zhang, G.; Liu, J.; et al. Flash-Searcher: Fast and Effective Web Agents via DAG-Based Parallel Execution. CoRR 2025, abs/2509.25301, 2509.25301. [Google Scholar] [CrossRef]
- Contributors, M.C.P. Model Context Protocol (MCP). 2025. [Google Scholar]
- Contributors, A. AGENTS.md Community Specification. 2025. [Google Scholar]
- Anthropic. Equipping agents for the real world with Agent Skills. 2025. [Google Scholar]
- Anthropic. Harness design for long-running application development. 2026. [Google Scholar]
- Uchibeke, U. Before the Tool Call: Deterministic Pre-Action Authorization for Autonomous AI Agents. CoRR 2026, abs/2603.20953, [2603.20953. [Google Scholar] [CrossRef]
- Anthropic. Scaling Managed Agents: Decoupling the brain from the hands; 2026. [Google Scholar]
- Bühler, C.; Biagiola, M.; Grazia, L.D.; Salvaneschi, G. Securing AI Agent Execution. CoRR 2025, abs/2510.21236, 2510.21236. [Google Scholar] [CrossRef]
- Anthropic. Overview - Claude Code Docs. 2025. [Google Scholar]
- OpenClaw. openclaw: Your own personal AI assistant. 2025.
- Research, N. hermes-agent: The agent that grows with you. 2025. [Google Scholar]
- Anthropic. Computer Use Tool — Claude API Docs. 2025. [Google Scholar]
- Zhang, J.; Xiang, J.; Yu, Z.; Teng, F.; Chen, X.; Chen, J.; Zhuge, M.; Cheng, X.; Hong, S.; Wang, J.; et al. AFlow: Automating Agentic Workflow Generation. In Proceedings of the The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025; 2025. Available online: https://openreview.net/.
- Wang, Y.; Liu, S.; Fang, J.; Meng, Z. EvoAgentX: An Automated Framework for Evolving Agentic Workflows. In Proceedings of the Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025 - System Demonstrations; Suzhou, China, Habernal, I., Schulam, P., Tiedemann, J., Eds.; Association for Computational Linguistics, 4-9 November 2025; Volume 2025, pp. 643–655. [Google Scholar] [CrossRef]
- Liu, S.; Fang, J.; Zhou, H.; Wang, Y.; Meng, Z. SEW: Self-Evolving Agentic Workflows for Automated Code Generation. CoRR 2025, abs/2505.18646, 2505.18646. [Google Scholar] [CrossRef]
- Zhuge, M.; Wang, W.; Kirsch, L.; Faccio, F.; Khizbullin, D.; Schmidhuber, J. GPTSwarm: Language Agents as Optimizable Graphs. In Proceedings of the Forty-first International Conference on Machine Learning, ICML 2024 PMLR / OpenReview.net; Vienna, Austria, Salakhutdinov, R., Kolter, Z., Heller, K.A., Weller, A., Oliver, N., Scarlett, J., Berkenkamp, F., Eds.; Proceedings of Machine Learning Research; 21-27 July 2024; Vol. 235, pp. 62743–62767. [Google Scholar]
- Wei, Y.; Duchenne, O.; Copet, J.; Carbonneaux, Q.; Zhang, L.; Fried, D.; Synnaeve, G.; Singh, R.; Wang, S.I. SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution. In Proceedings of the Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025; San Diego, CA, USA, December 2-7, 2025 / Mexico City, Mexico, Belgrave, D., Zhang, C., Montoya, L.N., Lin, H., Pascanu, R., Koniusz, P., Ghassemi, M., Chen, N., Ruíz, I.V.M., Loaiza-Bonilla, A., Eds.; 30 November 2025. [Google Scholar]
- Sun, S.; Song, H.; Huang, L.; Jiang, J.; Le, R.; Lv, Z.; Chen, Z.; Hu, Y.; Luo, W.; Zhao, W.X.; et al. SWE-World: Building Software Engineering Agents in Docker-Free Environments. CoRR 2026, abs/2602.03419, 2602.03419. [Google Scholar] [CrossRef]
- Wang, Y.; Xie, T.; Shen, K.; Wang, M.; Yang, L. RLAnything: Forge Environment, Policy, and Reward Model in Completely Dynamic RL System. CoRR 2026, abs/2602.02488, 2602.02488. [Google Scholar] [CrossRef]
- Wang, H.; Zou, H.; Song, H.; Feng, J.; Fang, J.; Lu, J.; Liu, L.; Luo, Q.; Liang, S.; Huang, S.; et al. UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning. arXiv 2025, arXiv:cs. [Google Scholar]
- Wang, D.; Cheng, M.; Li, Q.; Yu, S.; Ouyang, J.; Liu, Q. Claw-R1: A Step-Level Data Middleware System for Agentic Reinforcement Learning. CoRR 2026, abs/2606.09138, [2606.09138. [Google Scholar] [CrossRef]
- Dong, G.; Lu, J.; Huang, J.; Zhong, W.; Liu, L.; Huang, S.; Li, Z.; Zhao, Y.; Song, X.; Li, X.; et al. Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence. CoRR 2026, abs/2604.18292, [2604.18292. [Google Scholar] [CrossRef]
- Sun, Y.; Han, X.; Zhang, W.; Pang, Y.; Wang, T.; Cao, Y.; Huang, Y.; Duroiu, C.; Zhang, H.; Lin, J.; et al. Agents’ Last Exam. CoRR 2026, abs/2606.05405, 2606.05405. [Google Scholar] [CrossRef]
- Li, J.; Zhao, W.; Zhao, J.; Zeng, W.; Wu, H.; Wang, X.; Ge, R.; Cao, Y.; Huang, Y.; Liu, W.; et al. The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution. CoRR 2025, abs/2510.25726, [2510.25726. [Google Scholar] [CrossRef]
- Wu, Z.; Liu, X.; Zhang, X.; Chen, L.; Meng, F.; Du, L.; Zhao, Y.; Zhang, F.; Ye, Y.; Wang, J.; et al. MCPMark: A Benchmark for Stress-Testing Realistic and Comprehensive MCP Use. CoRR 2025, abs/2509.24002, [2509.24002. [Google Scholar] [CrossRef]
- Ye, B.; Li, R.; Yang, Q.; Liu, Y.; Yao, L.; Lv, H.; Xie, Z.; An, C.; Li, L.; Kong, L.; et al. Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents. arXiv 2026, arXiv:cs. [Google Scholar]
- Meng, F.; Du, L.; Wu, Z.; Chen, G.; Liu, X.; Liao, J.; Jiang, C.; Wan, Z.; Gu, J.; Zhou, P.; et al. ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents. CoRR 2026, abs/2604.23781, 2604.23781. [Google Scholar] [CrossRef]
- Opsahl-Ong, K.; Singhvi, A.; Collins, J.; Zhou, I.; Wang, C.; Baheti, A.; Oertell, O.; Portes, J.; Havens, S.; Elsen, E.; et al. OfficeQA Pro: An Enterprise Benchmark for End-to-End Grounded Reasoning. CoRR 2026, abs/2603.08655, 2603.08655. [Google Scholar] [CrossRef]
- Jin, J.; Hu, Y.; Qiu, K.; Dai, Q.; Luo, C.; Dong, G.; Li, X.; Zhao, T.; Ma, X.; Zhang, G.; et al. Toward Generalist Autonomous Research via Hypothesis-Tree Refinement. CoRR 2026, abs/2606.11926, [2606.11926. [Google Scholar] [CrossRef]
- Chhikara, P.; Khant, D.; Aryan, S.; Singh, T.; Yadav, D. Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory. In Proceedings of the ECAI 2025 - 28th European Conference on Artificial Intelligence; Bologna, Italy - Including 14th Conference on Prestigious Applications of Intelligent Systems (PAIS 2025), Lynce, I., Murano, N., Vallati, M., Villata, S., Chesani, F., Milano, M., Omicini, A., Dastani, M., Eds.; IOS Press; Frontiers in Artificial Intelligence and Applications; 25-30 October 2025; Vol. 413, pp. 2993–3000. [Google Scholar] [CrossRef]
- Du, Y.; Wei, F.; Zhang, H. AnyTool: Self-Reflective, Hierarchical Agents for Large-Scale API Calls. In Proceedings of the Forty-first International Conference on Machine Learning, ICML 2024 PMLR / OpenReview.net; Vienna, Austria, Salakhutdinov, R., Kolter, Z., Heller, K.A., Weller, A., Oliver, N., Scarlett, J., Berkenkamp, F., Eds.; Proceedings of Machine Learning Research; 21-27 July 2024; Vol. 235, pp. 11812–11829. [Google Scholar]
- Wang, R.; Han, X.; Ji, L.; Wang, S.; Baldwin, T.; Li, H. ToolGen: Unified Tool Retrieval and Calling via Generation. In Proceedings of the The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025; 2025. Available online: https://openreview.net/.
- Liang, Y.; Zhong, R.; Xu, H.; Jiang, C.; Zhong, Y.; Fang, R.; Gu, J.; Deng, S.; Yao, Y.; Wang, M.; et al. SkillNet: Create, Evaluate, and Connect AI Skills. CoRR 2026, abs/2603.04448, 2603.04448. [Google Scholar] [CrossRef]
- Yuan, A.; Su, Z.; Zhao, Y. AEGIS: No Tool Call Left Unchecked - A Pre-Execution Firewall and Audit Layer for AI Agents. CoRR 2026, abs/2603.12621, [2603.12621. [Google Scholar] [CrossRef]
- Wu, Y.; Roesner, F.; Kohno, T.; Zhang, N.; Iqbal, U. IsolateGPT: An Execution Isolation Architecture for LLM-Based Agentic Systems. In Proceedings of the 32nd Annual Network and Distributed System Security Symposium, NDSS 2025, San Diego, California, USA, February 24-28, 2025; The Internet Society, 2025. [Google Scholar]
- Wang, H.; Poskitt, C.M.; Sun, J. AgentSpec: Customizable Runtime Enforcement for Safe and Reliable LLM Agents. CoRR 2025, abs/2503.18666, 2503.18666. [Google Scholar] [CrossRef]
- Yang, R.; Fu, M.; Tantithamthavorn, C.; Arora, C.; Gulmammadova, G.; Chua, J.J. AdaptiveGuard: Towards Adaptive Runtime Safety for LLM-Powered Software. In Proceedings of the 40th IEEE/ACM International Conference on Automated Software Engineering, ASE 2025, Seoul, Korea, Republic of, November 16-20, 2025; IEEE; 2025, pp. 3381–3391. [Google Scholar] [CrossRef]
- Luo, W.; Dai, S.; Liu, X.; Banerjee, S.; Sun, H.; Chen, M.; Xiao, C. AGrail: A Lifelong Agent Guardrail with Effective and Adaptive Safety Detection. In Proceedings of the Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics; Vienna, Austria, Che, W., Nabende, J., Shutova, E., Pilehvar, M.T., Eds.; Association for Computational Linguistics, 2025; Volume 1, pp. 8104–8139. [Google Scholar] [CrossRef]
- Manakul, P.; Liusie, A.; Gales, M.J.F. SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models. In Proceedings of the Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing; EMNLP 2023, Singapore, Bouamor, H., Pino, J., Bali, K., Eds.; Association for Computational Linguistics, 6-10 December 2023; pp. 9004–9017. [Google Scholar] [CrossRef]
- Miculicich, L.; Parmar, M.; Palangi, H.; Dvijotham, K.D.; Montanari, M.; Pfister, T.; Le, L.T. VeriGuard: Enhancing LLM Agent Safety via Verified Code Generation. CoRR 2025, abs/2510.05156, 2510.05156. [Google Scholar] [CrossRef]
- Lightman, H.; Kosaraju, V.; Burda, Y.; Edwards, H.; Baker, B.; Lee, T.; Leike, J.; Schulman, J.; Sutskever, I.; Cobbe, K. Let’s Verify Step by Step. In Proceedings of the The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024; 2024. Available online: https://openreview.net/.
- Zheng, L.; Chiang, W.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E.P.; et al. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. In Proceedings of the Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023; New Orleans, LA, USA, Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., Levine, S., Eds.; 10 - 16 December 2023. [Google Scholar]
- Wang, P.; Li, L.; Shao, Z.; Xu, R.; Dai, D.; Li, Y.; Chen, D.; Wu, Y.; Sui, Z. Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations. In Proceedings of the Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics; Bangkok, Thailand, Ku, L., Martins, A., Srikumar, V., Eds.; Association for Computational Linguistics, 2024; Volume 1, pp. 9426–9439. [Google Scholar] [CrossRef]
- Sareen, K.; Moss, M.M.; Sordoni, A.; Agarwal, R.; Hosseini, A. Putting the Value Back in RL: Better Test-Time Scaling by Unifying LLM Reasoners With Verifiers. CoRR 2025, abs/2505.04842, 2505.04842. [Google Scholar] [CrossRef]
- Yu, C.; Cheng, Z.; Cui, H.; Gao, Y.; Luo, Z.; Wang, Y.; Zheng, H.; Zhao, Y. A Survey on Agent Workflow - Status and Future. CoRR 2025, abs/2508.01186, 2508.01186. [Google Scholar] [CrossRef]
- Sumers, T.R.; Yao, S.; Narasimhan, K.; Griffiths, T.L. Cognitive Architectures for Language Agents. In Trans. Mach. Learn. Res.; 2024. [Google Scholar]
- Zhang, X.; Cui, Y.; Wang, G.; Qiu, W.; Li, Z.; Han, F.; Huang, Y.; Qiu, H.; Zhu, B.; He, P. Verified Multi-Agent Orchestration: A Plan-Execute-Verify-Replan Framework for Complex Query Resolution. CoRR 2026, abs/2603.11445, [2603.11445. [Google Scholar] [CrossRef]
- Yao, Y.; Zhu, H.; Wang, P.; Ren, J.; Yang, X.; Chen, Q.; Li, X.; Shi, D.; Li, J.; Wang, Q.; et al. O-Researcher: An Open Ended Deep Research Model via Multi-Agent Distillation and Agentic RL. CoRR 2026, abs/2601.03743, 2601.03743. [Google Scholar] [CrossRef]
- Besta, M.; Blach, N.; Kubicek, A.; Gerstenberger, R.; Podstawski, M.; Gianinazzi, L.; Gajda, J.; Lehmann, T.; Niewiadomski, H.; Nyczyk, P.; et al. Graph of Thoughts: Solving Elaborate Problems with Large Language Models; Vancouver, Canada, Wooldridge, M.J., Dy, J.G., Natarajan, S., Eds.; Proceedings of the Thirty-Eighth AAAI Conference on Artificial Intelligence, AAAI 2024, Thirty-Sixth Conference on Innovative Applications of Artificial Intelligence, IAAI 2024, Fourteenth Symposium on Educational Advances in Artificial Intelligence, EAAI 2024, February 20-27, 2024; AAAI Press, 2024; pp. 17682–17690. [Google Scholar] [CrossRef]
- Choi, J.; Kim, H.; Ong, H.; Jang, M.; Kim, D.; Kim, J.; Yoon, Y. ReAcTree: Hierarchical LLM Agent Trees with Control Flow for Long-Horizon Task Planning. CoRR 2025, abs/2511.02424, 2511.02424. [Google Scholar] [CrossRef]
- Li, J.; Le, H.; Zhou, Y.; Xiong, C.; Savarese, S.; Sahoo, D. CodeTree: Agent-guided Tree Search for Code Generation with Large Language Models. In Proceedings of the Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL 2025 - Volume 1: Long Papers; Albuquerque, New Mexico, USA, Chiruzzo, L., Ritter, A., Wang, L., Eds.; Association for Computational Linguistics, 29 April 2025; Volume 2025, pp. 3711–3726. [Google Scholar] [CrossRef]
- Koh, J.Y.; McAleer, S.M.; Fried, D.; Salakhutdinov, R. Tree Search for Language Model Agents. Trans. Mach. Learn. Res. 2025, 2025. [Google Scholar]
- Ro, Y.; Qiu, H.; Goiri, Í.; Fonseca, R.; Bianchini, R.; Akella, A.; Wang, Z.; Erez, M.; Choukse, E. Sherlock: Reliable and Efficient Agentic Workflow Execution. CoRR 2025, abs/2511.00330, 2511.00330. [Google Scholar] [CrossRef]
- Li, X. Chain-in-Tree: Back to Sequential Reasoning in LLM Tree Search. In Proceedings of the Findings of the Association for Computational Linguistics, ACL 2026; San Diego, California, United States, Liakata, M., Moreira, V.P., Zhang, J., Jurgens, D., Eds.; Association for Computational Linguistics, 2-7 July 2026; Volume 2026, pp. 4365–4392. [Google Scholar]
- DeepSeek-AI. DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models. CoRR 2025, abs/2512.02556, 2512.02556. [Google Scholar] [CrossRef]
- Vercel. We Removed 80% of Our Agent’s Tools. 2025. [Google Scholar]
- Bai, S.; Bing, L.; Lei, L.; Li, R.; Li, X.; Lin, X.; Min, E.; Su, L.; Wang, B.; Wang, L.; et al. MiroThinker-1.7 & H1: Towards Heavy-Duty Research Agents via Verification. CoRR 2026, abs/2603.15726, [2603.15726. [Google Scholar] [CrossRef]
- Chen, G.; Qiao, Z.; Chen, X.; Yu, D.; Xu, H.; Zhao, W.X.; Song, R.; Yin, W.; Yin, H.; Zhang, L.; et al. IterResearch: Rethinking Long-Horizon Agents with Interaction Scaling. arXiv 2026, arXiv:cs. [Google Scholar]
- Wu, Y.; Zheng, Y.; Xu, T.; Zhang, Z.; Yu, Y.; Zhu, J.; Ma, C.; Lin, B.; Dong, B.; Zhu, H.; et al. ContextBudget: Budget-Aware Context Management for Long-Horizon Search Agents. CoRR 2026, abs/2604.01664, 2604.01664. [Google Scholar] [CrossRef]
- Sun, W.; Lu, M.; Ling, Z.; Liu, K.; Yao, X.; Yang, Y.; Chen, J. Scaling Long-Horizon LLM Agent via Context-Folding. CoRR 2025, abs/2510.11967, [2510.11967. [Google Scholar] [CrossRef]
- Ye, R.; Zhang, Z.; Li, K.; Yin, H.; Tao, Z.; Zhao, Y.; Su, L.; Zhang, L.; Qiao, Z.; Wang, X.; et al. AgentFold: Long-Horizon Web Agents with Proactive Context Management. CoRR 2025, abs/2510.24699, 2510.24699. [Google Scholar] [CrossRef]
- Han, D.; Couturier, C.; Díaz, D.M.; Zhang, X.; Rühle, V.; Rajmohan, S. LEGOMem: Modular Procedural Memory for Multi-agent LLM Systems for Workflow Automation. CoRR 2025, abs/2510.04851, 2510.04851. [Google Scholar] [CrossRef]
- Wang, Z.; Chen, H.; Wang, J.; Wei, W. Memex(RL): Scaling Long-Horizon LLM Agents via Indexed Experience Memory. CoRR 2026, abs/2603.04257, 2603.04257. [Google Scholar] [CrossRef]
- Dou, S.; Zhang, M.; Yin, Z.; Huang, C.; Shen, Y.; Wang, J.; Chen, J.; Ni, Y.; Ye, J.; Zhang, C.; et al. CL-bench: A Benchmark for Context Learning. CoRR 2026, abs/2602.03587, 2602.03587. [Google Scholar] [CrossRef]
- Hu, Y.; Qian, H.; Wang, S.; Liu, J.; Zhao, T.; Li, X.; Liu, Z.; Dou, Z. AgentFugue: Agent Scaling for Long-Horizon Tasks through Collective Reasoning. CoRR 2026, abs/2605.24486, [2605.24486. [Google Scholar] [CrossRef]
- Zhang, Y.; Shu, J.; Ma, Y.; Lin, X.; Wu, S.; Sang, J. Memory as Action: Autonomous Context Curation for Long-Horizon Agentic Tasks. In Proceedings of the Findings of the Association for Computational Linguistics, ACL 2026; San Diego, California, United States, Liakata, M., Moreira, V.P., Zhang, J., Jurgens, D., Eds.; Association for Computational Linguistics, 2-7 July 2026; Volume 2026, pp. 19149–19164. [Google Scholar]
- Hu, Y.; Qian, H.; Wang, S.; Liu, J.; Zhao, Z.; Tan, J.; Liu, Z.; Dou, Z. SAM: State-Adaptive Memory for Long-Horizon Reasoning Agent. CoRR 2026, abs/2605.24468, [2605.24468. [Google Scholar] [CrossRef]
- Cursor. Cursor Docs — Rules. 2025.
- Gutierrez, B.J.; Shu, Y.; Gu, Y.; Yasunaga, M.; Su, Y. HippoRAG: Neurobiologically Inspired Long-Term Memory for Large Language Models. In Proceedings of the Advances in Neural Information Processing Systems 37: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024; Vancouver, BC, Canada, Globersons, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J.M., Zhang, C., Eds.; 10 - 15 December 2024. [Google Scholar]
- Xu, W.; Liang, Z.; Mei, K.; Gao, H.; Tan, J.; Zhang, Y. A-MEM: Agentic Memory for LLM Agents. CoRR 2025, abs/2502.12110, 2502.12110. [Google Scholar] [CrossRef]
- Zheng, L.; Wang, R.; Wang, X.; An, B. Synapse: Trajectory-as-Exemplar Prompting with Memory for Computer Control. In Proceedings of the The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024; 2024. Available online: https://openreview.net/.
- Liang, X.; Tao, M.; Xia, Y.; Wang, J.; Li, K.; Wang, Y.; He, Y.; Yang, J.; Shi, T.; Wang, Y.; et al. SAGE: Self-evolving Agents with Reflective and Memory-augmented Abilities. Neurocomputing 2025, 647, 130470. [Google Scholar] [CrossRef]
- Yang, L.; Yu, Z.; Zhang, T.; Cao, S.; Xu, M.; Zhang, W.; Gonzalez, J.E.; Cui, B. Buffer of Thoughts: Thought-Augmented Reasoning with Large Language Models. In Proceedings of the Advances in Neural Information Processing Systems 37: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024; Vancouver, BC, Canada, Globersons, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J.M., Zhang, C., Eds.; 10 - 15 December 2024. [Google Scholar]
- Tang, X.; Qin, T.; Peng, T.; Zhou, Z.; Shao, D.; Du, T.; Wei, X.; Xia, P.; Wu, F.; Zhu, H.; et al. Agent KB: Leveraging Cross-Domain Experience for Agentic Problem Solving. CoRR 2025, abs/2507.06229, 2507.06229. [Google Scholar] [CrossRef]
- Zhang, Z.; Dai, Q.; Bo, X.; Ma, C.; Li, R.; Chen, X.; Zhu, J.; Dong, Z.; Wen, J. A Survey on the Memory Mechanism of Large Language Model-based Agents. ACM Trans. Inf. Syst. 2025, 43, 155:1–155:47. [Google Scholar] [CrossRef]
- Tan, J.; Dou, Z.; Zhang, L.; Hu, Y.; Cheng, Y.; Wen, J. MemSifter: Offloading LLM Memory Retrieval via Outcome-Driven Proxy Reasoning. CoRR 2026, abs/2603.03379, 2603.03379. [Google Scholar] [CrossRef]
- Hu, Y.; Liu, J.; Tan, J.; Zhu, Y.; Dou, Z. Memory Matters More: Event-Centric Memory as a Logic Map for Agent Searching and Reasoning. In Proceedings of the Findings of the Association for Computational Linguistics, ACL 2026; San Diego, California, United States, Liakata, M., Moreira, V.P., Zhang, J., Jurgens, D., Eds.; Association for Computational Linguistics, 2-7 July 2026; Volume 2026, pp. 22389–22407. [Google Scholar]
- Rasmussen, P.; Paliychuk, P.; Beauvais, T.; Ryan, J.; Chalef, D. Zep: A Temporal Knowledge Graph Architecture for Agent Memory. CoRR 2025, abs/2501.13956, 2501.13956. [Google Scholar] [CrossRef]
- Cai, Y.; Guo, X.; Huang, X.; Du, J.; Huang, C.; Huang, W.; Ma, W.; Hu, Y.; Zeng, A.; Tang, J.; et al. From Player to Master: Enhancing Test-Time Learning of LLM Agents via Reinforcement Learning over Memory. CoRR 2026, abs/2606.08656, [2606.08656. [Google Scholar] [CrossRef]
- Zhang, G.; Fu, M.; Wan, G.; Yu, M.; Wang, K.; Yan, S. G-Memory: Tracing Hierarchical Memory for Multi-Agent Systems. CoRR 2025, abs/2506.07398, [2506.07398. [Google Scholar] [CrossRef]
- Ni, J.; Liu, Y.; Liu, X.; Sun, Y.; Zhou, M.; Cheng, P.; Wang, D.; Zhao, E.; Jiang, X.; Jiang, G. Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills. CoRR 2026, abs/2603.25158, 2603.25158. [Google Scholar] [CrossRef]
- Luo, J.; Tian, Y.; Cao, C.; Luo, Z.; Lin, H.; Li, K.; Kong, C.; Yang, R.; Ma, J. From Storage to Experience: A Survey on the Evolution of LLM Agent Memory Mechanisms. In Proceedings of the Findings of the Association for Computational Linguistics, ACL 2026; San Diego, California, United States, Liakata, M., Moreira, V.P., Zhang, J., Jurgens, D., Eds.; Association for Computational Linguistics, 2-7 July 2026; pp. 41622–41652. [Google Scholar]
- Hu, C.; Gao, X.; Zhou, Z.; Xu, D.; Bai, Y.; Li, X.; Zhang, H.; Li, T.; Zhang, C.; Bing, L.; et al. EverMemOS: A Self-Organizing Memory Operating System for Structured Long-Horizon Reasoning. In Proceedings of the Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics; Liakata, M., Moreira, V.P., Zhang, J., Jurgens, D., Eds.; San Diego, California, United States, 2026; Volume 1, pp. 45836–45853. [Google Scholar] [CrossRef]
- Wang, Y.; Chen, X. MIRIX: Multi-Agent Memory System for LLM-Based Agents. arXiv 2025, arXiv:cs. [Google Scholar]
- Tran, H.; Yao, Z.; Tran, N.L.; Yang, Z.; Ouyang, F.; Han, S.; Rahimi, R.; Yu, H. PRIME: Planning and Retrieval-Integrated Memory for Enhanced Reasoning. In Proceedings of the Fortieth AAAI Conference on Artificial Intelligence, Thirty-Eighth Conference on Innovative Applications of Artificial Intelligence, Sixteenth Symposium on Educational Advances in Artificial Intelligence; AAAI 2026, Singapore, Koenig, S., Jenkins, C., Taylor, M.E., Eds.; AAAI Press, 20-27 January 2026; pp. 33268–33276. [Google Scholar] [CrossRef]
- Sun, H.; Zeng, S.; Zhang, B. H-MEM: Hierarchical Memory for High-Efficiency Long-Term Reasoning in LLM Agents. In Proceedings of the Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics; Rabat, Morocco, Demberg, V., Inui, K., Marquez, L., Eds.; 2026; Volume 1, pp. 341–350. [Google Scholar] [CrossRef]
- Qian, H.; Liu, Z.; Zhang, P.; Mao, K.; Lian, D.; Dou, Z.; Huang, T. MemoRAG: Boosting Long Context Processing with Global Memory-Enhanced Retrieval Augmentation. In Proceedings of the Proceedings of the ACM on Web Conference 2025, New York, NY, USA, 2025; WWW ’25, pp. 2366–2377. [Google Scholar] [CrossRef]
- Wu, D.; Wang, H.; Yu, W.; Zhang, Y.; Chang, K.; Yu, D. LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory. In Proceedings of the The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025; 2025. Available online: https://openreview.net/.
- Jiang, D.; Li, Y.; Wei, S.; Yang, J.; Kishore, A.; Zhao, A.; Kang, D.; Hu, X.; Chen, F.; Li, Q.; et al. Anatomy of Agentic Memory: Taxonomy and Empirical Analysis of Evaluation and System Limitations. CoRR 2026, abs/2602.19320, 2602.19320. [Google Scholar] [CrossRef]
- Huang, W.; Zhang, W.; Liang, Y.; Bei, Y.; Chen, Y.; Feng, T.; Pan, X.; Tan, Z.; Wang, Y.; Wei, T.; et al. Rethinking Memory Mechanisms of Foundation Agents in the Second Half: A Survey. CoRR 2026, abs/2602.06052, [2602.06052. [Google Scholar] [CrossRef]
- Xu, H.; Li, C.; Ma, X.; Ou, X.; Zhang, Z.; He, T.; Liu, X.; Wang, Z.; Liang, J.; Chu, Z.; et al. The Evolution of Tool Use in LLM Agents: From Single-Tool Call to Multi-Tool Orchestration. CoRR 2026, abs/2603.22862, [2603.22862. [Google Scholar] [CrossRef]
- Liu, Z.; Hoang, T.; Zhang, J.; Zhu, M.; Lan, T.; Kokane, S.; Tan, J.; Yao, W.; Liu, Z.; Feng, Y.; et al. APIGen: Automated PIpeline for Generating Verifiable and Diverse Function-Calling Datasets. In Proceedings of the Advances in Neural Information Processing Systems 37: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024; Vancouver, BC, Canada, Globersons, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J.M., Zhang, C., et al., Eds.; 10 - 15 December 2024. [Google Scholar]
- Yang, Y.; Chai, H.; Song, Y.; Qi, S.; Wen, M.; Li, N.; Liao, J.; Hu, H.; Lin, J.; Chang, G.; et al. A Survey of AI Agent Protocols. CoRR 2025, abs/2504.16736, [2504.16736. [Google Scholar] [CrossRef]
- Bandi, C.; Hertzberg, B.; Boo, G.; Polakam, T.; Da, J.; Hassaan, S.; Sharma, M.; Park, A.; Hernandez, E.; Rambado, D.; et al. MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers. CoRR 2026, abs/2602.00933, 2602.00933. [Google Scholar] [CrossRef]
- Li, Y.; Yang, X.; Wang, L.; Luo, W.; Chen, H. ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox. CoRR 2026, abs/2605.10787, [2605.10787. [Google Scholar] [CrossRef]
- Luo, Z.; Shen, Z.; Yang, W.; Zhao, Z.; Jwalapuram, P.; Saha, A.; Sahoo, D.; Savarese, S.; Xiong, C.; Li, J. MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers. CoRR 2025, abs/2508.14704, [2508.14704. [Google Scholar] [CrossRef]
- Sigdel, A.; Baral, R. Schema First Tool APIs for LLM Agents: A Controlled Study of Tool Misuse, Recovery, and Budgeted Performance. CoRR 2026, abs/2603.13404, [2603.13404. [Google Scholar] [CrossRef]
- Lu, J.; Holleis, T.; Zhang, Y.; Aumayer, B.; Nan, F.; Bai, H.; Ma, S.; Ma, S.; Li, M.; Yin, G.; et al. ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities. In Proceedings of the Findings of the Association for Computational Linguistics: NAACL 2025 Association for Computational Linguistics; Albuquerque, New Mexico, USA, Chiruzzo, L., Ritter, A., Wang, L., Eds.; Findings of ACL; 2025; Vol. NAACL 2025, pp. 1160–1183. [Google Scholar] [CrossRef]
- Shepard, D.; Salimans, R. AutomationBench. CoRR 2026, abs/2604.18934, 2604.18934. [Google Scholar] [CrossRef]
- He, W.; Sun, Y.; Hao, H.; Hao, X.; Xia, Z.; Gu, Q.; Han, C.; Zhao, D.; Su, H.; Zhang, K.; et al. VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications. CoRR 2025, abs/2509.26490, [2509.26490. [Google Scholar] [CrossRef]
- Yu, P.; Liu, W.; Yang, Y.; Li, J.; Zhang, Z.; Feng, X.; Zhang, F. Benchmarking LLM Tool-Use in the Wild. CoRR 2026, abs/2604.06185, 2604.06185. [Google Scholar] [CrossRef]
- Li, K.; Shi, J.; Xiao, Y.; Jiang, M.; Sun, J.; Wu, Y.; Fu, D.; Xia, S.; Cai, X.; Xu, T.; et al. AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts. In Proceedings of the Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2026; San Diego, California, United States, Liakata, M., Moreira, V.P., Zhang, J., Jurgens, D., Eds.; Association for Computational Linguistics, 2-7 July 2026; Volume 2026, pp. 7422–7440. [Google Scholar]
- Luo, H.; Zhang, H.; Zhang, X.; Wang, H.; Qin, Z.; Lu, W.; Ma, G.; He, H.; Xie, Y.; Zhou, Q.; et al. UltraHorizon: Benchmarking Agent Capabilities in Ultra Long-Horizon Scenarios. CoRR 2025, abs/2509.21766, [2509.21766. [Google Scholar] [CrossRef]
- Gan, T.; Sun, Q. RAG-MCP: Mitigating Prompt Bloat in LLM Tool Selection via Retrieval-Augmented Generation. CoRR 2025, abs/2505.03275, 2505.03275. [Google Scholar] [CrossRef]
- Fei, X.; Zheng, X.; Feng, H. MCP-Zero: Active Tool Discovery for Autonomous LLM Agents. CoRR 2025, abs/2506.01056, 2506.01056. [Google Scholar] [CrossRef]
- Mastouri, M.; Ksontini, E.; Kessentini, W. Making REST APIs Agent-Ready: From OpenAPI to Model Context Protocol Servers for Tool-Augmented LLMs. CoRR 2025, abs/2507.16044, 2507.16044. [Google Scholar] [CrossRef]
- Anthropic. Introducing advanced tool use on the Claude Developer Platform. 2025. [Google Scholar]
- Li, X.; Jiao, W.; Jin, J.; Dong, G.; Jin, J.; Wang, Y.; Wang, H.; Zhu, Y.; Wen, J.; Lu, Y.; et al. DeepAgent: A General Reasoning Agent with Scalable Toolsets. In Proceedings of the Proceedings of the ACM Web Conference 2026, WWW 2026 originally scheduled for April 13-17, 2026, rescheduled for June 29 - July 3, 2026; Dubai, United Arab Emirates, Hacid, H., Maarek, Y., Bonchi, F., Guy, I., Yilmaz, E., Eds.; ACM, 2026; pp. 2219–2230. [Google Scholar] [CrossRef]
- Anthropic. Code execution with MCP: building more efficient AI agents. 2025. [Google Scholar]
- Felendler, Y.; Gandhi, P.A.; Habler, I.; Elovici, Y.; Shabtai, A. From Tool Orchestration to Code Execution: A Study of MCP Design Choices. CoRR 2026, abs/2602.15945, [2602.15945. [Google Scholar] [CrossRef]
- Xu, R.; Yan, Y. Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward. CoRR 2026, abs/2602.12430, [2602.12430. [Google Scholar] [CrossRef]
- Li, Z.; Zhang, H.; Wei, C.; Lu, P.; Nie, P.; Lu, Y.; Bai, Y.; Feng, S.; Zhu, H.; Zhong, M.; et al. Beyond Semantic Similarity: Rethinking Retrieval for Agentic Search via Direct Corpus Interaction. CoRR 2026, abs/2605.05242, 2605.05242. [Google Scholar] [CrossRef]
- Subramanian, S.; Akinfaderin, A.; Zhang, Y.; Singh, I.; Khanuja, M.; Singh, S.; Tanke, M.L. Keyword search is all you need: Achieving RAG-Level Performance without vector databases using agentic tool use. CoRR 2026, abs/2602.23368, [2602.23368. [Google Scholar] [CrossRef]
- Sutawika, L.; Soni, A.B.; R, B.S.R.; Gandhi, A.; Yassine, T.; Vijayvargiya, S.; Li, Y.; Zhou, X.; Zhang, Y.; Maben, L.M.; et al. CodeScout: An Effective Recipe for Reinforcement Learning of Code Search Agents. CoRR 2026, abs/2603.17829, 2603.17829. [Google Scholar] [CrossRef]
- Zhou, Y.; Shu, W.; Su, Y.; Du, W.; Fang, Y.; Lin, X. A Comprehensive Survey on Agent Skills: Taxonomy, Techniques, and Applications. CoRR 2026, abs/2605.07358, [2605.07358. [Google Scholar] [CrossRef]
- Lu, J.; Kong, Z.; Wang, Y.; Fu, R.; Wan, H.; Yang, C.; Lou, W.; Sun, H.; Wang, L.; Jiang, Y.; et al. Beyond Static Tools: Test-Time Tool Evolution for Scientific Reasoning. CoRR 2026, abs/2601.07641, 2601.07641. [Google Scholar] [CrossRef]
- Huang, X.; Chen, J.; Fei, Y.; Li, Z.; Schwaller, P.; Ceder, G. CASCADE: Cumulative Agentic Skill Creation through Autonomous Development and Evolution. CoRR 2025, abs/2512.23880, 2512.23880. [Google Scholar] [CrossRef]
- Sun, Z.; Liu, Z.; Zang, Y.; Cao, Y.; Dong, X.; Wu, T.; Lin, D.; Wang, J. SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience. CoRR 2025, abs/2508.04700, 2508.04700. [Google Scholar] [CrossRef]
- Agashe, S.; Wong, K.; Tu, V.; Yang, J.; Li, A.; Wang, X.E. Agent S2: A Compositional Generalist-Specialist Framework for Computer Use Agents. CoRR 2025, abs/2504.00906, 2504.00906. [Google Scholar] [CrossRef]
- Chen, S.; Gai, J.; Zhou, R.; Zhang, J.; Zhu, T.; Li, J.; Wang, K.; Wang, Z.; Chen, Z.; Kaleb, K.; et al. SkillCraft: Can LLM Agents Learn to Use Tools Skillfully? CoRR 2026, abs/2603.00718, 2603.00718. [Google Scholar] [CrossRef]
- Mi, Q.; Ma, Z.; Yang, M.; Li, H.; Wang, Y.; Zhang, H.; Wang, J. ProcMEM: Learning Reusable Procedural Memory from Experience via Non-Parametric PPO for LLM Agents. CoRR 2026, abs/2602.01869, [2602.01869. [Google Scholar] [CrossRef]
- Li, X. When Single-Agent with Skills Replace Multi-Agent Systems and When They Fail. CoRR 2026, abs/2601.04748, 2601.04748. [Google Scholar] [CrossRef]
- Zhou, H.; Chen, Y.; Guo, S.; Yan, X.; Lee, K.H.; Wang, Z.; Lee, K.Y.; Zhang, G.; Shao, K.; Yang, L.; et al. Memento: Fine-tuning LLM Agents without Fine-tuning LLMs. CoRR 2025, abs/2508.16153, [2508.16153. [Google Scholar] [CrossRef]
- Qiu, J.; Qi, X.; Zhang, T.; Juan, X.; Guo, J.; Lu, Y.; Wang, Y.; Yao, Z.; Ren, Q.; Jiang, X.; et al. Alita: Generalist Agent Enabling Scalable Agentic Reasoning with Minimal Predefinition and Maximal Self-Evolution. CoRR 2025, abs/2505.20286, [2505.20286. [Google Scholar] [CrossRef]
- Yang, Y.; Gong, Z.; Huang, W.; Yang, Q.; Zhou, Z.; Huang, Z.; Li, Y.; Gao, X.; Dai, Q.; Liu, B.; et al. SkillOpt: Executive Strategy for Self-Evolving Agent Skills. CoRR 2026, abs/2605.23904, 2605.23904. [Google Scholar] [CrossRef]
- Chen, M.F.; Roberts, N.; Bhatia, K.; Wang, J.; Zhang, C.; Sala, F.; Ré, C. Skill-it! A data-driven skills framework for understanding and training language models. In Proceedings of the Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023; New Orleans, LA, USA, Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., Levine, S., Eds.; 10 - 16 December 2023. [Google Scholar]
- Li, X.; Li, M.; Bao, K.; Ma, Y.; Wang, W.; Liu, D.; Feng, F. SkillGraph: Skill-Augmented Reinforcement Learning for Agents via Evolving Skill Graphs. CoRR 2026, abs/2605.12039, [2605.12039. [Google Scholar] [CrossRef]
- Xia, T.; Hu, L.; Sun, Y.; Xu, M.; Xu, L.; Wang, S.; Xu, W.; Jiang, J. GraSP: Graph-Structured Skill Compositions for LLM Agents. CoRR 2026, abs/2604.17870, 2604.17870. [Google Scholar] [CrossRef]
- Lu, Z.; Yao, Z.; Wu, J.; Han, C.; Gu, Q.; Cai, X.; Lu, W.; Xiao, J.; Zhuang, Y.; Shen, Y. SKILL0: In-Context Agentic Reinforcement Learning for Skill Internalization. CoRR 2026, abs/2604.02268, 2604.02268. [Google Scholar] [CrossRef]
- Tran, K.; Dao, D.; Nguyen, M.; Pham, Q.; O’Sullivan, B.; Nguyen, H.D. Multi-Agent Collaboration Mechanisms: A Survey of LLMs. CoRR 2025, abs/2501.06322, 2501.06322. [Google Scholar] [CrossRef]
- Guo, T.; Chen, X.; Wang, Y.; Chang, R.; Pei, S.; Chawla, N.V.; Wiest, O.; Zhang, X. Large Language Model Based Multi-agents: A Survey of Progress and Challenges. In Proceedings of the Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence, IJCAI 2024, Jeju, South Korea, August 3-9, 2024; 2024; pp. 8048–8057. Available online: https://www.ijcai.org/.
- Cemri, M.; Pan, M.Z.; Yang, S.; Agrawal, L.A.; Chopra, B.; Tiwari, R.; Keutzer, K.; Parameswaran, A.G.; Klein, D.; Ramchandran, K.; et al. Why Do Multi-Agent LLM Systems Fail? In Proceedings of the Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025; San Diego, CA, USA, December 2-7, 2025 / Mexico City, Mexico, Belgrave, D., Zhang, C., Montoya, L.N., Lin, H., Pascanu, R., Koniusz, P., Ghassemi, M., Chen, N., Ruíz, I.V.M., Loaiza-Bonilla, A., Eds.; 30 November 2025. [Google Scholar]
- Wang, Y.; Wu, Z.; Yao, J.; Su, J. TDAG: A multi-agent framework based on dynamic Task Decomposition and Agent Generation. Neural Netw. 2025, 185, 107200. [Google Scholar] [CrossRef]
- Li, A.; Xie, Y.; Li, S.; Tsung, F.; Ding, B.; Li, Y. Agent-Oriented Planning in Multi-Agent Systems. In Proceedings of the The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025; 2025. Available online: https://openreview.net/.
- Wang, J.; Wang, J.; Athiwaratkun, B.; Zhang, C.; Zou, J. Mixture-of-Agents Enhances Large Language Model Capabilities. In Proceedings of the The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025; 2025. Available online: https://openreview.net/.
- Chen, W.; Su, Y.; Zuo, J.; Yang, C.; Yuan, C.; Chan, C.; Yu, H.; Lu, Y.; Hung, Y.; Qian, C.; et al. AgentVerse: Facilitating Multi-Agent Collaboration and Exploring Emergent Behaviors. In Proceedings of the The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024; 2024. Available online: https://openreview.net/.
- Liu, Z.; Zhang, Y.; Li, P.; Liu, Y.; Yang, D. Dynamic LLM-Agent Network: An LLM-agent Collaboration Framework with Agent Team Optimization. CoRR 2023, abs/2310.02170, 2310.02170. [Google Scholar] [CrossRef]
- Qian, C.; Xie, Z.; Wang, Y.; Liu, W.; Zhu, K.; Xia, H.; Dang, Y.; Du, Z.; Chen, W.; Yang, C.; et al. Scaling Large Language Model-based Multi-Agent Collaboration. In Proceedings of the The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025; 2025. Available online: https://openreview.net/.
- Chivukula, A.; Somasundaram, J.; Somasundaram, V. Agint: Agentic Graph Compilation for Software Engineering Agents. CoRR 2025, abs/2511.19635, 2511.19635. [Google Scholar] [CrossRef]
- Cheng, P.; Hu, T.; Xu, H.; Zhang, Z.; Dai, Y.; Han, L.; Du, N.; Li, X. Self-playing Adversarial Language Game Enhances LLM Reasoning. In Proceedings of the Advances in Neural Information Processing Systems 37: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024; Vancouver, BC, Canada, Globersons, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J.M., Zhang, C., Eds.; 10 - 15 December 2024. [Google Scholar]
- Ni, K.; Hua, W.; Shi, X.; Guo, J.; Chang, S.; Chen, T. Chimera: Latency- and Performance-Aware Multi-agent Serving for Heterogeneous LLMs. CoRR 2026, abs/2603.22206, [2603.22206. [Google Scholar] [CrossRef]
- Zhou, H.; Chan, H.Y. ORCH: many analyses, one merge - a deterministic multi-agent orchestrator for discrete-choice reasoning with EMA-guided routing. Front. Artif. Intell. 2026, 9. [Google Scholar] [CrossRef]
- Yue, Y.; Zhang, G.; Liu, B.; Wan, G.; Wang, K.; Cheng, D.; Qi, Y. MasRouter: Learning to Route LLMs for Multi-Agent Systems. In Proceedings of the Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025; Vienna, Austria, Che, W., Nabende, J., Shutova, E., Pilehvar, M.T., Eds.; Association for Computational Linguistics, 27 July 2025; Volume 2025, pp. 15549–15572. [Google Scholar] [CrossRef]
- Ruan, J.; Xu, Z.; Peng, Y.; Ren, F.; Yu, Z.; Liang, X.; Xiang, J.; Chen, Y.; Liu, B.; Wu, C.; et al. AOrchestra: Automating Sub-Agent Creation for Agentic Orchestration. CoRR 2026, abs/2602.03786, 2602.03786. [Google Scholar] [CrossRef]
- Zhang, W.; Cui, C.; Zhao, Y.; Hu, R.; Liu, Y.; Zhou, Y.; An, B. AgentOrchestra: A Hierarchical Multi-Agent Framework for General-Purpose Task Solving. CoRR 2025, abs/2506.12508, [2506.12508. [Google Scholar] [CrossRef]
- Dang, Y.; Qian, C.; Luo, X.; Fan, J.; Xie, Z.; Shi, R.; Chen, W.; Yang, C.; Che, X.; Tian, Y.; et al. Multi-Agent Collaboration via Evolving Orchestration. In Proceedings of the Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025; San Diego, CA, USA, December 2-7, 2025 / Mexico City, Mexico, Belgrave, D., Zhang, C., Montoya, L.N., Lin, H., Pascanu, R., Koniusz, P., Ghassemi, M., Chen, N., Ruíz, I.V.M., Loaiza-Bonilla, A., Eds.; 30 November 2025. [Google Scholar]
- Zhang, Y.; Lin, C.; Tang, S.; Chen, H.; Zhou, S.; Ma, Y.; Tresp, V. SwarmAgentic: Towards Fully Automated Agentic System Generation via Swarm Intelligence. In Proceedings of the Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025; Suzhou, China, Christodoulopoulos, C., Chakraborty, T., Rose, C., Peng, V., Eds.; Association for Computational Linguistics, 4-9 November 2025; Volume 2025, pp. 1778–1818. [Google Scholar] [CrossRef]
- Qu, A.; Zheng, H.; Zhou, Z.; Yan, Y.; Tang, Y.; Ong, S.Y.; Hong, F.; Zhou, K.; Jiang, C.; Kong, M.; et al. CORAL: Towards Autonomous Multi-Agent Evolution for Open-Ended Discovery. CoRR 2026, abs/2604.01658, 2604.01658. [Google Scholar] [CrossRef]
- Dochkina, V. Drop the Hierarchy and Roles: How Self-Organizing LLM Agents Outperform Designed Structures. CoRR 2026, abs/2603.28990, [2603.28990. [Google Scholar] [CrossRef]
- Research, I. Agent Communication Protocol (ACP). 2025. [Google Scholar]
- Ehtesham, A.; Singh, A.; Gupta, G.K.; Kumar, S. A survey of agent interoperability protocols: Model Context Protocol (MCP), Agent Communication Protocol (ACP), Agent-to-Agent Protocol (A2A), and Agent Network Protocol (ANP). CoRR 2025, abs/2505.02279, 2505.02279. [Google Scholar] [CrossRef]
- Ranjan, R.; Gupta, S.; Singh, S.N. LOKA Protocol: A Decentralized Framework for Trustworthy and Ethical AI Agent Ecosystems. CoRR 2025, abs/2504.10915, 2504.10915. [Google Scholar] [CrossRef]
- Marro, S.; Malfa, E.L.; Wright, J.; Li, G.; Shadbolt, N.; Wooldridge, M.J.; Torr, P. A Scalable Communication Protocol for Networks of Large Language Models. CoRR 2024, abs/2410.11905, [2410.11905. [Google Scholar] [CrossRef]
- Rahmani, L.; Minarsch, D.; Ward, J. Peer-to-peer Autonomous Agent Communication Network. In Proceedings of the AAMAS ’21: 20th International Conference on Autonomous Agents and Multiagent Systems; Virtual Event, United Kingdom, Dignum, F., Lomuscio, A., Endriss, U., Nowé, A., Eds.; ACM, 3-7 May 2021; pp. 1037–1045. [Google Scholar] [CrossRef]
- Li, J.; Zhang, Q.; Yu, Y.; Fu, Q.; Ye, D. More Agents Is All You Need. In Trans. Mach. Learn. Res.; 2024. [Google Scholar]
- Lee, Y.; Yen, H.; Ye, X.; Chen, D. Agentic Aggregation for Parallel Scaling of Long-Horizon Agentic Tasks. CoRR 2026, abs/2604.11753, [2604.11753. [Google Scholar] [CrossRef]
- Lee, N.; Erdogan, L.E.; John, C.J.; Krishnapillai, S.; Mahoney, M.W.; Keutzer, K.; Gholami, A. Agentic Test-Time Scaling for WebAgents. CoRR 2026, abs/2602.12276, [2602.12276. [Google Scholar] [CrossRef]
- Abdelnabi, S.; Greshake, K.; Mishra, S.; Endres, C.; Holz, T.; Fritz, M. Not What You’ve Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. In Proceedings of the Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security, AISec 2023; Copenhagen, Denmark, Pintor, M., Chen, X., Tramèr, F., Eds.; ACM, 30 November 2023; pp. 79–90. [Google Scholar] [CrossRef]
- Shamsujjoha, M.; Lu, Q.; Zhao, D.; Zhu, L. Swiss Cheese Model for AI Safety: A Taxonomy and Reference Architecture for Multi-Layered Guardrails of Foundation Model Based Agents. In Proceedings of the 22nd IEEE International Conference on Software Architecture, ICSA 2025, Odense, Denmark, March 31 - April 4, 2025; IEEE; 2025, pp. 37–48. [Google Scholar] [CrossRef]
- Marchand, R.; Catháin, A.Ó.; Wynne, J.; Giavridis, P.M.; Deverett, S.; Wilkinson, J.; Gwartz, J.; Coppock, H. Quantifying Frontier LLM Capabilities for Container Sandbox Escape. CoRR 2026, abs/2603.02277, 2603.02277. [Google Scholar] [CrossRef]
- Rajagopalan, M.; Rao, V. Authenticated Workflows: A Systems Approach to Protecting Agentic AI. CoRR 2026, abs/2602.10465, [2602.10465. [Google Scholar] [CrossRef]
- Mozannar, H.; Bansal, G.; Tan, C.; Fourney, A.; Dibia, V.; Chen, J.; Gerrits, J.; Payne, T.; Maldaner, M.K.; Grunde-McLaughlin, M.; et al. Magentic-UI: Towards Human-in-the-loop Agentic Systems. CoRR 2025, abs/2507.22358, [2507.22358. [Google Scholar] [CrossRef]
- Rebedea, T.; Dinu, R.; Sreedhar, M.N.; Parisien, C.; Cohen, J. NeMo Guardrails: A Toolkit for Controllable and Safe LLM Applications with Programmable Rails. In Proceedings of the Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023 - System Demonstrations; Singapore, Feng, Y., Lefever, E., Eds.; Association for Computational Linguistics, 6-10 December 2023; pp. 431–445. [Google Scholar] [CrossRef]
- Kamath, A.; Zhang, S.; Xu, C.; Ugare, S.; Singh, G.; Misailovic, S. Enforcing Temporal Constraints for LLM Agents. CoRR 2025, abs/2512.23738, 2512.23738. [Google Scholar] [CrossRef]
- Shi, T.; He, J.; Wang, Z.; Li, H.; Wu, L.; Guo, W.; Song, D. Progent: Securing AI Agents with Privilege Control. arXiv 2026, arXiv:cs. [Google Scholar]
- Palumbo, N.; Choudhary, S.; Choi, J.; Amir, G.; Chalasani, P.; Jha, S. Formal Policy Enforcement for Real-World Agentic Systems. arXiv 2026, arXiv:cs. [Google Scholar]
- Debenedetti, E.; Shumailov, I.; Fan, T.; Hayes, J.; Carlini, N.; Fabian, D.; Kern, C.; Shi, C.; Terzis, A.; Tramèr, F. Defeating Prompt Injections by Design. CoRR 2025, abs/2503.18813, 2503.18813. [Google Scholar] [CrossRef]
- Xiang, Z.; Zheng, L.; Li, Y.; Hong, J.; Li, Q.; Xie, H.; Zhang, J.; Xiong, Z.; Xie, C.; Yang, C.; et al. GuardAgent: Safeguard LLM Agents via Knowledge-Enabled Reasoning. In Proceedings of the Forty-second International Conference on Machine Learning, ICML 2025 PMLR / OpenReview.net; Vancouver, BC, Canada, Singh, A., Fazel, M., Hsu, D., Lacoste-Julien, S., Berkenkamp, F., Maharaj, T., Wagstaff, K., Zhu, J., Eds.; Proceedings of Machine Learning Research; 13-19 July 2025; Vol. 267. [Google Scholar]
- Kumar, T.; Tripathy, A.; Saranathan, G.; Foltin, M.; Bhattacharya, S.; Hinchley, S.; Bahls, D.M.; Brookshire, D.; Kaplan, L.; Wisniewski, R.W. InfrastructureSentinel: Policy Enforced Guardrails for Secure MCP-driven Infrastructure Agents. In Proceedings of the Fortieth AAAI Conference on Artificial Intelligence, Thirty-Eighth Conference on Innovative Applications of Artificial Intelligence, Sixteenth Symposium on Educational Advances in Artificial Intelligence, AAAI 2026, Singapore; Koenig, S., Jenkins, C., Taylor, M.E., Eds.; AAAI Press, 20-27 January 2026; pp. 40295–40301. [Google Scholar] [CrossRef]
- Inan, H.; Upasani, K.; Chi, J.; Rungta, R.; Iyer, K.; Mao, Y.; Tontchev, M.; Hu, Q.; Fuller, B.; Testuggine, D.; et al. Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations. CoRR 2023, abs/2312.06674, 2312.06674. [Google Scholar] [CrossRef]
- Kholkar, G.; Ahuja, R. Policy-as-Prompt: Turning AI Governance Rules into Guardrails for AI Agents. arXiv 2025, arXiv:cs. [Google Scholar]
- Chennabasappa, S.; Nikolaidis, C.; Song, D.; Molnar, D.; Ding, S.; Wan, S.; Whitman, S.; Deason, L.; Doucette, N.; Montilla, A.; et al. LlamaFirewall: An open source guardrail system for building secure AI agents. CoRR 2025, abs/2505.03574, 2505.03574. [Google Scholar] [CrossRef]
- Farquhar, S.; Kossen, J.; Kuhn, L.; Gal, Y. Detecting hallucinations in large language models using semantic entropy. Nat. 2024, 630, 625–630. [Google Scholar] [CrossRef]
- Zhang, J.; Choubey, P.K.; Huang, K.; Xiong, C.; Wu, C. Agentic Uncertainty Quantification. CoRR 2026, abs/2601.15703, 2601.15703. [Google Scholar] [CrossRef]
- Zhao, Q.; Zhao, X.; Liu, Y.; Cheng, W.; Sun, Y.; Oishi, M.; Osaki, T.; Matsuda, K.; Yao, H.; Chen, H. SAUP: Situation Awareness Uncertainty Propagation on LLM Agent. CoRR 2024, abs/2412.01033, 2412.01033. [Google Scholar] [CrossRef]
- Wang, H.; Poskitt, C.M.; Wei, J.; Sun, J. ProbGuard: Probabilistic Runtime Monitoring for LLM Agent Safety. arXiv 2026, arXiv:cs. [Google Scholar]
- Liu, Y.; Zhang, C.; Han, Z.; Liu, H.; Wang, Y.; Yu, Y.; Wang, X.; Yin, Y. TrajAD: Trajectory Anomaly Detection for Trustworthy LLM Agents. CoRR 2026, abs/2602.06443, 2602.06443. [Google Scholar] [CrossRef]
- He, X.; Wu, D.; Zhai, Y.; Sun, K. SentinelAgent: Graph-based Anomaly Detection in Multi-Agent Systems. CoRR 2025, abs/2505.24201, 2505.24201. [Google Scholar] [CrossRef]
- Deng, Y.; Yang, Y.; Zhang, J.; Wang, W.; Li, B. Enhancing LLM Safety Through a Theoretical Minimax Game Lens. arXiv 2026, arXiv:cs. [Google Scholar]
- Zhang, Z.; Cui, S.; Lu, Y.; Zhou, J.; Yang, J.; Wang, H.; Huang, M. Agent-SafetyBench: Evaluating the Safety of LLM Agents. CoRR 2024, abs/2412.14470, 2412.14470. [Google Scholar] [CrossRef]
- Ruan, Y.; Dong, H.; Wang, A.; Pitis, S.; Zhou, Y.; Ba, J.; Dubois, Y.; Maddison, C.J.; Hashimoto, T. Identifying the Risks of LM Agents with an LM-Emulated Sandbox. In Proceedings of the The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024; 2024. Available online: https://openreview.net/.
- Su, Y.; Xu, K.; Gao, Y.; Yang, F.; Liu, C.; Yang, M.; Xu, T. Neuro-Symbolic Verification on Instruction Following of LLMs. CoRR 2026, abs/2601.17789, 2601.17789. [Google Scholar] [CrossRef]
- Du, Y.; Li, S.; Torralba, A.; Tenenbaum, J.B.; Mordatch, I. Improving Factuality and Reasoning in Language Models through Multiagent Debate. In Proceedings of the Forty-first International Conference on Machine Learning, ICML 2024 PMLR / OpenReview.net; Vienna, Austria, Salakhutdinov, R., Kolter, Z., Heller, K.A., Weller, A., Oliver, N., Scarlett, J., Berkenkamp, F., Eds.; Proceedings of Machine Learning Research; 21-27 July 2024; Vol. 235, pp. 11733–11763. [Google Scholar]
- Lifshitz, S.; McIlraith, S.A.; Du, Y. Multi-Agent Verification: Scaling Test-Time Compute with Multiple Verifiers. CoRR 2025, abs/2502.20379, 2502.20379. [Google Scholar] [CrossRef]
- Xi, Z.; Liao, C.; Li, G.; Zhang, Z.; Chen, W.; Wang, B.; Jin, S.; Zhou, Y.; Guan, J.; Wu, W.; et al. AgentPRM: Process Reward Models for LLM Agents via Step-Wise Promise and Progress. In Proceedings of the Proceedings of the ACM Web Conference 2026, WWW 2026 originally scheduled for April 13-17, 2026, rescheduled for June 29 - July 3, 2026; Dubai, United Arab Emirates, Hacid, H., Maarek, Y., Bonchi, F., Guy, I., Yilmaz, E., Eds.; ACM, 2026; pp. 4184–4195. [Google Scholar] [CrossRef]
- Zhuge, M.; Zhao, C.; Ashley, D.R.; Wang, W.; Khizbullin, D.; Xiong, Y.; Liu, Z.; Chang, E.; Krishnamoorthi, R.; Tian, Y.; et al. Agent-as-a-Judge: Evaluate Agents with Agents. In Proceedings of the Forty-second International Conference on Machine Learning, ICML 2025 PMLR / OpenReview.net; Vancouver, BC, Canada, Singh, A., Fazel, M., Hsu, D., Lacoste-Julien, S., Berkenkamp, F., Maharaj, T., Wagstaff, K., Zhu, J., et al., Eds.; Proceedings of Machine Learning Research; 13-19 July 2025; Vol. 267. [Google Scholar]
- Yuan, L.; Li, W.; Chen, H.; Cui, G.; Ding, N.; Zhang, K.; Zhou, B.; Liu, Z.; Peng, H. Free Process Rewards without Process Labels. In Proceedings of the Forty-second International Conference on Machine Learning, ICML 2025; Vancouver, BC, Canada, Singh, A., Fazel, M., Hsu, D., Lacoste-Julien, S., Berkenkamp, F., Maharaj, T., Wagstaff, K., Zhu, J., Eds.; PMLR / OpenReview.net; Proceedings of Machine Learning Research; 13-19 July 2025; Vol. 267. [Google Scholar]
- Zhang, D.; Zhoubian, S.; Hu, Z.; Yue, Y.; Dong, Y.; Tang, J. ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search. In Proceedings of the Advances in Neural Information Processing Systems 37: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024; Vancouver, BC, Canada, Globersons, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J.M., Zhang, C., Eds.; 10 - 15 December 2024. [Google Scholar]
- Snell, C.V.; Lee, J.; Xu, K.; Kumar, A. Scaling LLM Test-Time Compute Optimally Can be More Effective than Scaling Parameters for Reasoning. In Proceedings of the The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025; 2025. Available online: https://openreview.net/.
- Laboratory, S.A.I. AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security. CoRR 2026, abs/2601.18491, 2601.18491. [Google Scholar] [CrossRef]
- Laboratory, S.A.I. AgentDoG 1.5: A Lightweight and Scalable Alignment Framework for AI Agent Safety and Security. CoRR 2026, abs/2605.29801, 2605.29801. [Google Scholar] [CrossRef]
- Yuan, T.; He, Z.; Dong, L.; Wang, Y.; Zhao, R.; Xia, T.; Xu, L.; Zhou, B.; Li, F.; Zhang, Z.; et al. R-Judge: Benchmarking Safety Risk Awareness for LLM Agents. In Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2024 Association for Computational Linguistics; Miami, Florida, USA, Al-Onaizan, Y., Bansal, M., Chen, Y., Eds.; Findings of ACL; 2024; Vol. EMNLP 2024, pp. 1467–1490. [Google Scholar] [CrossRef]
- Guo, C.; Liu, X.; Xie, C.; Zhou, A.; Zeng, Y.; Lin, Z.; Song, D.; Li, B. RedCode: Risky Code Execution and Generation Benchmark for Code Agents. In Proceedings of the Advances in Neural Information Processing Systems 37: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024; Vancouver, BC, Canada, Globersons, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J.M., Zhang, C., Eds.; 10 - 15 December 2024. [Google Scholar]
- Sun, Q.; Li, M.; Liu, Z.; Xie, Z.; Xu, F.; Yin, Z.; Cheng, K.; Li, Z.; Ding, Z.; Liu, Q.; et al. OS-Sentinel: Towards Safety-Enhanced Mobile GUI Agents via Hybrid Validation in Realistic Workflows. In Proceedings of the Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics; San Diego, California, United States, Liakata, M., Moreira, V.P., Zhang, J., Jurgens, D., Eds.; Association for Computational Linguistics, 2-7 July 2026; Volume 1, pp. 9529–9553. [Google Scholar]
- Deng, Y.; Zhang, X.; Zhang, W.; Yuan, Y.; Ng, S.; Chua, T. On the Multi-turn Instruction Following for Conversational Web Agents. In Proceedings of the Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics; Bangkok, Thailand, Ku, L., Martins, A., Srikumar, V., Eds.; Association for Computational Linguistics, 11-16 August 2024; Volume 1, pp. 8795–8812. [Google Scholar] [CrossRef]
- Menon, A.; Saebo, M.; Crosse, T.; Gibson, S.; Jang, E.; Cruz, D. Inherited Goal Drift: Contextual Pressure Can Undermine Agentic Goals. CoRR 2026, abs/2603.03258, [2603.03258. [Google Scholar] [CrossRef]
- Cobbe, K.; Kosaraju, V.; Bavarian, M.; Chen, M.; Jun, H.; Kaiser, L.; Plappert, M.; Tworek, J.; Hilton, J.; Nakano, R.; et al. Training Verifiers to Solve Math Word Problems. CoRR 2021, abs/2110.14168, [2110.14168. [Google Scholar]
- Uesato, J.; Kushman, N.; Kumar, R.; Song, H.F.; Siegel, N.Y.; Wang, L.; Creswell, A.; Irving, G.; Higgins, I. Solving math word problems with process- and outcome-based feedback. CoRR 2022, abs/2211.14275, [2211.14275. [Google Scholar] [CrossRef]
- Zhang, L.; Hosseini, A.; Bansal, H.; Kazemi, M.; Kumar, A.; Agarwal, R. Generative Verifiers: Reward Modeling as Next-Token Prediction. In Proceedings of the The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025; 2025. Available online: https://openreview.net/.
- Zhao, J.; Liu, R.; Zhang, K.; Zhou, Z.; Gao, J.; Li, D.; Lyu, J.; Qian, Z.; Qi, B.; Li, X.; et al. GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning. In Proceedings of the Fortieth AAAI Conference on Artificial Intelligence, Thirty-Eighth Conference on Innovative Applications of Artificial Intelligence, Sixteenth Symposium on Educational Advances in Artificial Intelligence; AAAI 2026, Singapore, Koenig, S., Jenkins, C., Taylor, M.E., Eds.; AAAI Press, 20-27 January 2026; pp. 34932–34940. [Google Scholar] [CrossRef]
- Khalifa, M.; Agarwal, R.; Logeswaran, L.; Kim, J.; Peng, H.; Lee, M.; Lee, H.; Wang, L. Process Reward Models That Think. In Trans. Mach. Learn. Res.; 2026. [Google Scholar]
- Chae, H.; Kim, S.; Cho, J.; Kim, S.; Moon, S.; Hwangbo, G.; Lim, D.; Kim, M.; Hwang, Y.; Gwak, M.; et al. Web-Shepherd: Advancing PRMs for Reinforcing Web Agents. CoRR 2025, abs/2505.15277, 2505.15277. [Google Scholar] [CrossRef]
- Gandhi, S.; Tsay, J.; Ganhotra, J.; Kate, K.; Rizk, Y. When Agents go Astray: Course-Correcting SWE Agents with PRMs. CoRR 2025, abs/2509.02360, 2509.02360. [Google Scholar] [CrossRef]
- Han, H.; Xie, J.; Ma, X.; Zhu, W.; Zhang, Z.; Long, Z.; Chen, H.; Ye, Q. SWE-TRACE: Optimizing Long-Horizon SWE Agents Through Rubric Process Reward Models and Heuristic Test-Time Scaling. CoRR 2026, abs/2604.14820, [2604.14820. [Google Scholar] [CrossRef]
- Li, M.; Zeng, Q.; Fang, T.; Liang, Z.; Song, L.; Liu, Q.; Mi, H.; Yu, D. Verified Critical Step Optimization for LLM Agents. In Proceedings of the Findings of the Association for Computational Linguistics, ACL 2026; San Diego, California, United States, Liakata, M., Moreira, V.P., Zhang, J., Jurgens, D., Eds.; Association for Computational Linguistics, 2-7 July 2026; Volume 2026, pp. 39627–39639. [Google Scholar]
- Li, D.; Yao, Y.; Tan, Z.; Liu, H.; Guo, R. ToolPRMBench: Evaluating and Advancing Process Reward Models for Tool-using Agents. In Proceedings of the Findings of the Association for Computational Linguistics, ACL 2026; San Diego, California, United States, Liakata, M., Moreira, V.P., Zhang, J., Jurgens, D., Eds.; Association for Computational Linguistics, 2-7 July 2026; Volume 2026, pp. 12378–12391. [Google Scholar]
- Zhang, J.; Fu, Z.; Xi, Z.; Jing, W.; Chai, M.; He, W.; Zhang, G.; Fan, C.; An, C.; Chen, W.; et al. AgentV-RL: Scaling Reward Modeling with Agentic Verifier. In Proceedings of the Findings of the Association for Computational Linguistics, ACL 2026; San Diego, California, United States, Liakata, M., Moreira, V.P., Zhang, J., Jurgens, D., Eds.; Association for Computational Linguistics, 2-7 July 2026; pp. 23078–23100. [Google Scholar]
- You, R.; Cai, H.; Zhang, C.; Xu, Q.; Liu, M.; Yu, T.; Li, Y.; Li, W. Agent-as-a-Judge. CoRR 2026, abs/2601.05111, 2601.05111. [Google Scholar] [CrossRef]
- Qian, Y.; Zhang, S.; Zhou, Y.; Ding, H.; Socolinsky, D.; Zhang, Y. CollabEval: Enhancing LLM-as-a-Judge via Multi-Agent Collaboration. CoRR 2026, abs/2603.00993, 2603.00993. [Google Scholar] [CrossRef]
- Tie, G.; Yuan, Z.; Zhao, Z.; Hu, C.; Gu, T.; Zhang, R.; Zhang, S.; Wu, J.; Tu, X.; Jin, M.; et al. Can LLMs Correct Themselves? A Benchmark of Self-. Correction in LLMs. In Proceedings of the Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025, San Diego, CA, USA, December 2-7, 2025 / Mexico City, Mexico, November 30 - December 5, 2025; Belgrave, D.; Zhang, C.; Montoya, L.N.; Lin, H.; Pascanu, R.; Koniusz, P.; Ghassemi, M.; Chen, N.; Ruíz, I.V.M.; Loaiza-Bonilla, A., Eds., 2025.
- Shum, K.; Hui, B.; Chen, J.; Zhang, L.; W., X.; Yang, J.; Huang, Y.; Lin, J.; He, J. SWE-RM: Execution-free Feedback For Software Engineering Agents. CoRR 2025, abs/2512.21919, 2512.21919. [Google Scholar] [CrossRef]
- Dhuliawala, S.; Komeili, M.; Xu, J.; Raileanu, R.; Li, X.; Celikyilmaz, A.; Weston, J. Chain-of-Verification Reduces Hallucination in Large Language Models. In Association for Computational Linguistics; Proceedings of the Findings of the Association for Computational Linguistics, ACL 2024, Bangkok, Thailand and virtual meeting, Ku, L., Martins, A., Srikumar, V., Eds.; Findings of ACL; 2024; Vol. ACL 2024, pp. 3563–3578. [Google Scholar] [CrossRef]
- Cui, C.; Huang, J.; Wang, S.; Zheng, L.; Kong, Q.; Zeng, Z. Agentic Reward Modeling: Verifying GUI Agent via Online Proactive Interaction. CoRR 2026, abs/2602.00575, 2602.00575. [Google Scholar] [CrossRef]
- Guan, X.; Zhang, L.L.; Liu, Y.; Shang, N.; Sun, Y.; Zhu, Y.; Yang, F.; Yang, M. rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking. In Proceedings of the Forty-second International Conference on Machine Learning, ICML 2025; Vancouver, BC, Canada, Singh, A., Fazel, M., Hsu, D., Lacoste-Julien, S., Berkenkamp, F., Maharaj, T., Wagstaff, K., Zhu, J., Eds.; PMLR / OpenReview.net; Proceedings of Machine Learning Research; 13-19 July 2025; Vol. 267. [Google Scholar]
- Misaki, K.; Inoue, Y.; Imajuku, Y.; Kuroki, S.; Nakamura, T.; Akiba, T. Wider or Deeper? Scaling LLM Inference-Time Compute with Adaptive Branching Tree Search. CoRR 2025, abs/2503.04412, 2503.04412. [Google Scholar] [CrossRef]
- DeepSeek-AI. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. CoRR 2025, abs/2501.12948, 2501.12948. [Google Scholar] [CrossRef]
- Wang, Y.; Ji, P.; Yang, C.; Li, K.; Hu, M.; Li, J.; Sartoretti, G. MCTS-Judge: Test-Time Scaling in LLM-as-a-Judge for Code Correctness Evaluation. CoRR 2025, abs/2502.12468, 2502.12468. [Google Scholar] [CrossRef]
- Zhang, H.; Wang, R.; Ji, Y.; Kwak, M.G.; Wu, X.; Li, C.; Zhang, L.; Shi, W.; Peng, Y.; Wang, Y. Scaling Medical Reasoning Verification via Tool-Integrated Reinforcement Learning. CoRR 2026, abs/2601.20221, 2601.20221. [Google Scholar] [CrossRef]
- Tay, Y.; Dehghani, M.; Bahri, D.; Metzler, D. Efficient Transformers: A Survey. ACM Comput. Surv. 2023, 55, 109:1–109:28. [Google Scholar] [CrossRef]
- Lepikhin, D.; Lee, H.; Xu, Y.; Chen, D.; Firat, O.; Huang, Y.; Krikun, M.; Shazeer, N.; Chen, Z. GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding. In Proceedings of the 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021; 2021. Available online: https://openreview.net/.
- Fedus, W.; Zoph, B.; Shazeer, N. Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity. J. Mach. Learn. Res. 2022, 23, 120:1–120:39. [Google Scholar]
- Lenz, B.; Lieber, O.; Arazi, A.; Bergman, A.; Manevich, A.; Peleg, B.; Aviram, B.; Almagor, C.; Fridman, C.; Padnos, D.; et al. Jamba: Hybrid Transformer-Mamba Language Models. In Proceedings of the The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025; 2025. Available online: https://openreview.net/.
- Peng, B.; Alcaide, E.; Anthony, Q.; Albalak, A.; Arcadinho, S.; Biderman, S.; Cao, H.; Cheng, X.; Chung, M.; Derczynski, L.; et al. RWKV: Reinventing RNNs for the Transformer Era. In Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore Association for Computational Linguistics; Bouamor, H., Pino, J., Bali, K., Eds.; Findings of ACL; 6-10 December 2023; Vol. EMNLP 2023, pp. 14048–14077. [Google Scholar] [CrossRef]
- Zhang, Y.; Lin, Z.; Yao, X.; Hu, J.; Meng, F.; Liu, C.; Men, X.; Yang, S.; Li, Z.; Li, W.; et al. Kimi Linear: An Expressive, Efficient Attention Architecture. CoRR 2025, abs/2510.26692, 2510.26692. [Google Scholar] [CrossRef]
- Li, Y.; Wei, F.; Zhang, C.; Zhang, H. EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty. In Proceedings of the Forty-first International Conference on Machine Learning, ICML 2024 PMLR / OpenReview.net; Vienna, Austria, Salakhutdinov, R., Kolter, Z., Heller, K.A., Weller, A., Oliver, N., Scarlett, J., Berkenkamp, F., Eds.; Proceedings of Machine Learning Research; 21-27 July 2024; Vol. 235, pp. 28935–28948. [Google Scholar]
- Ainslie, J.; Lee-Thorp, J.; de Jong, M.; Zemlyanskiy, Y.; Lebrón, F.; Sanghai, S. GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints. In Proceedings of the Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing; EMNLP 2023, Singapore, Bouamor, H., Pino, J., Bali, K., Eds.; Association for Computational Linguistics, 6-10 December 2023; pp. 4895–4901. [Google Scholar] [CrossRef]
- Shi, D.; Cao, J.; Chen, Q.; Sun, W.; Li, W.; Lu, H.; Dong, F.; Qin, T.; Zhu, K.; Liu, M.; et al. TaskCraft: Automated Generation of Agentic Tasks. CoRR 2025, abs/2506.10055, [2506.10055. [Google Scholar] [CrossRef]
- Xi, Z.; Huang, J.; Liao, C.; Huang, B.; Guo, H.; Liu, J.; Zheng, R.; Ye, J.; Zhang, J.; Chen, W.; et al. AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning. CoRR 2025, abs/2509.08755, 2509.08755. [Google Scholar] [CrossRef]
- Peng, B.; Quesnelle, J.; Fan, H.; Shippole, E. YaRN: Efficient Context Window Extension of Large Language Models. In Proceedings of the The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024; 2024. Available online: https://openreview.net/.
- Chen, Y.; Qian, S.; Tang, H.; Lai, X.; Liu, Z.; Han, S.; Jia, J. LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models. In Proceedings of the The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024; 2024. Available online: https://openreview.net/.
- Wang, P.; Bai, S.; Tan, S.; Wang, S.; Fan, Z.; Bai, J.; Chen, K.; Liu, X.; Wang, J.; Ge, W.; et al. Qwen2-VL: Enhancing Vision-Language Model’s Perception of the World at Any Resolution. CoRR 2024, abs/2409.12191, [2409.12191. [Google Scholar] [CrossRef]
- Zhu, J.; Wang, W.; Chen, Z.; Liu, Z.; Ye, S.; Gu, L.; Tian, H.; Duan, Y.; Su, W.; Shao, J.; et al. InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models. CoRR 2025, abs/2504.10479, [2504.10479. [Google Scholar] [CrossRef]
- Xie, S.M.; Pham, H.; Dong, X.; Du, N.; Liu, H.; Lu, Y.; Liang, P.; Le, Q.V.; Ma, T.; Yu, A.W. DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining. In Proceedings of the Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023; New Orleans, LA, USA, Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., Levine, S., Eds.; 10 - 16 December 2023. [Google Scholar]
- Xiao, Y.; Jiang, M.; Sun, J.; Li, K.; Lin, J.; Zhuang, Y.; Zeng, J.; Xia, S.; Hua, Q.; Li, X.; et al. LIMI: Less is More for Agency. CoRR 2025, abs/2509.17567, [2509.17567. [Google Scholar] [CrossRef]
- Parashar, S.; Gui, S.; Li, X.; Ling, H.; Vemuri, S.; Olson, B.; Li, E.; Zhang, Y.; Caverlee, J.; Kalathil, D.; et al. Curriculum Reinforcement Learning from Easy to Hard Tasks Improves LLM Reasoning. CoRR 2025, abs/2506.06632, [2506.06632. [Google Scholar] [CrossRef]
- Kang, M.; Jeong, J.; Lee, S.; Cho, J.; Hwang, S.J. Distilling LLM Agent into Small Models with Retrieval and Code Tools. In Proceedings of the Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025; San Diego, CA, USA, December 2-7, 2025 / Mexico City, Mexico, Belgrave, D., Zhang, C., Montoya, L.N., Lin, H., Pascanu, R., Koniusz, P., Ghassemi, M., Chen, N., Ruíz, I.V.M., Loaiza-Bonilla, A., Eds.; 30 November 2025. [Google Scholar]
- Yuan, S.; Chen, Z.; Xi, Z.; Ye, J.; Du, Z.; Chen, J. Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training. CoRR 2025, abs/2501.11425, 2501.11425. [Google Scholar] [CrossRef]
- Zhou, Y.; Li, S.; Liu, S.; Fang, W.; Zhao, J.; Yang, J.; Lv, J.; Zhang, K.; Zhou, Y.; Lu, H.; et al. Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General LLM Reasoning. CoRR 2025, abs/2508.16949, [2508.16949. [Google Scholar] [CrossRef]
- Yu, Q.; Zhang, Z.; Zhu, R.; Yuan, Y.; Zuo, X.; Yue, Y.; Dai, W.; Fan, T.; Liu, G.; Liu, J.; et al. DAPO: An Open-Source LLM Reinforcement Learning System at Scale. In Proceedings of the Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025; San Diego, CA, USA, December 2-7, 2025 / Mexico City, Mexico, Belgrave, D., Zhang, C., Montoya, L.N., Lin, H., Pascanu, R., Koniusz, P., Ghassemi, M., Chen, N., Ruíz, I.V.M., Loaiza-Bonilla, A., Eds.; 30 November 2025. [Google Scholar]
- Feng, L.; Xue, Z.; Liu, T.; An, B. Group-in-Group Policy Optimization for LLM Agent Training. CoRR 2025, abs/2505.10978, [2505.10978. [Google Scholar] [CrossRef]
- Ji, Y.; Ma, Z.; Wang, Y.; Chen, G.; Chu, X.; Wu, L. Tree Search for LLM Agent Reinforcement Learning. CoRR 2025, abs/2509.21240, 2509.21240. [Google Scholar] [CrossRef]
- Wang, J.; Zhang, W.; Shi, W.; Li, Y.; Cheng, J. TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents. CoRR 2026, abs/2604.24005, [2604.24005. [Google Scholar] [CrossRef]
- Hu, Z.; Liu, W.; Qu, X.; Yue, X.; Chen, C.; Wang, Z.; Cheng, Y. Divide and Conquer: Grounding LLMs as Efficient Decision-Making Agents via Offline Hierarchical Reinforcement Learning. In Proceedings of the Forty-second International Conference on Machine Learning, ICML 2025; Vancouver, BC, Canada, Singh, A., Fazel, M., Hsu, D., Lacoste-Julien, S., Berkenkamp, F., Maharaj, T., Wagstaff, K., Zhu, J., Eds.; PMLR / OpenReview.net; Proceedings of Machine Learning Research; 13-19 July 2025; Vol. 267. [Google Scholar]
- Yan, S.; Yang, X.; Huang, Z.; Nie, E.; Ding, Z.; Li, Z.; Ma, X.; Bi, J.; Kersting, K.; Pan, J.Z.; et al. Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning. In Proceedings of the Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2026; San Diego, California, United States, Liakata, M., Moreira, V.P., Zhang, J., Jurgens, D., Eds.; Association for Computational Linguistics, 2-7 July 2026; Volume 2026, pp. 12805–12825. [Google Scholar]
- Li, C.; Qiang, R.; Huang, J.; Gao, C.; Zhang, C.; He, N.; Dai, B. Revisiting DAgger in the Era of LLM-Agents. CoRR 2026, abs/2605.12913, [2605.12913. [Google Scholar] [CrossRef]
- Zhang, Y.; Zhu, Y.; Chong, W.; Tu, S.; Zhang, Q.; Chai, J.; Wang, X.; Lin, W.; Yin, G.; Zhao, D. π-Play: Multi-Agent Self-Play via Privileged Self-Distillation without External Data. CoRR 2026, abs/2604.14054, 2604.14054. [Google Scholar] [CrossRef]
- He, Y.; Kaur, S.; Bhaskar, A.; Yang, Y.; Liu, J.; Ri, N.; Fowl, L.; Panigrahi, A.; Chen, D.; Arora, S. Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision. CoRR 2026, abs/2604.12002, [2604.12002. [Google Scholar] [CrossRef]
- Zelikman, E.; Wu, Y.; Mu, J.; Goodman, N.D. STaR: Bootstrapping Reasoning With Reasoning. In Proceedings of the Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022; New Orleans, LA, USA, Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., Oh, A., Eds.; 28 November 2022. [Google Scholar]
- Xia, P.; Zeng, K.; Liu, J.; Qin, C.; Wu, F.; Zhou, Y.; Xiong, C.; Yao, H. Agent0: Unleashing Self-Evolving Agents from Zero Data via Tool-Integrated Reasoning. CoRR 2025, abs/2511.16043, 2511.16043. [Google Scholar] [CrossRef]
- Lu, S.; Wang, Z.; Zhang, H.; Wu, Q.; Gan, L.; Zhuang, C.; Gu, J.; Lin, T. Don’t Just Fine-tune the Agent, Tune the Environment. CoRR 2025, abs/2510.10197, 2510.10197. [Google Scholar] [CrossRef]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, L.; Polosukhin, I. Attention is All you Need. In Proceedings of the Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017; Long Beach, CA, USA, Guyon, I., von Luxburg, U., Bengio, S., Wallach, H.M., Fergus, R., Vishwanathan, S.V.N., Garnett, R., Eds.; 4-9 December 2017; pp. 5998–6008. [Google Scholar]
- Jiang, A.Q.; Sablayrolles, A.; Mensch, A.; Bamford, C.; Chaplot, D.S.; de Las Casas, D.; Bressand, F.; Lengyel, G.; Lample, G.; Saulnier, L.; et al. Mistral 7B. CoRR 2023. abs/2310.06825 2310.06825. [CrossRef]
- Team, G. Gemma 2: Improving Open Language Models at a Practical Size. CoRR 2024, abs/2408.00118, 2408.00118. [Google Scholar] [CrossRef]
- Team, G. Gemma 3 Technical Report. CoRR 2025, abs/2503.19786, 2503.19786. [Google Scholar] [CrossRef]
- Lu, E.; Jiang, Z.; Liu, J.; Du, Y.; Jiang, T.; Hong, C.; Liu, S.; He, W.; Yuan, E.; Wang, Y.; et al. MoBA: Mixture of Block Attention for Long-Context LLMs. CoRR 2025, abs/2502.13189, [2502.13189. [Google Scholar] [CrossRef]
- Yang, A.; Yang, B.; Zhang, B.; Hui, B.; Zheng, B.; Yu, B.; Li, C.; Liu, D.; Huang, F.; Wei, H.; et al. Qwen2.5 Technical Report. CoRR 2024, abs/2412.15115, [2412.15115. [Google Scholar] [CrossRef]
- Zeng, A.; Xu, B.; Wang, B.; Zhang, C.; Yin, D.; Rojas, D.; Feng, G.; Zhao, H.; Lai, H.; Yu, H.; et al. ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools. CoRR 2024, abs/2406.12793, [2406.12793. [Google Scholar] [CrossRef]
- Wang, S.; Li, B.Z.; Khabsa, M.; Fang, H.; Ma, H. Linformer: Self-Attention with Linear Complexity. CoRR 2020, abs/2006.04768. [Google Scholar]
- Choromanski, K.M.; Likhosherstov, V.; Dohan, D.; Song, X.; Gane, A.; Sarlós, T.; Hawkins, P.; Davis, J.Q.; Mohiuddin, A.; Kaiser, L.; et al. Rethinking Attention with Performers. In Proceedings of the 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021; 2021. Available online: https://openreview.net/.
- Sun, Y.; Dong, L.; Huang, S.; Ma, S.; Xia, Y.; Xue, J.; Wang, J.; Wei, F. Retentive Network: A Successor to Transformer for Large Language Models. CoRR 2023, abs/2307.08621, [2307.08621. [Google Scholar] [CrossRef]
- Gu, A.; Goel, K.; Ré, C. Efficiently Modeling Long Sequences with Structured State Spaces. In Proceedings of the The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022; 2022. Available online: https://openreview.net/.
- Patro, B.N.; Agneeswaran, V.S. Mamba-360: Survey of state space models as transformer alternative for long sequence modelling: Methods, Applications, and Challenges. Eng. Appl. Artif. Intell. 2025, 159, 111279. [Google Scholar] [CrossRef]
- Tiezzi, M.; Casoni, M.; Betti, A.; Gori, M.; Melacci, S. State-space modeling in long sequence processing: a survey on recurrence in the transformer era. Neural Netw. 2026, 193, 108039. [Google Scholar] [CrossRef]
- Leviathan, Y.; Kalman, M.; Matias, Y. Fast Inference from Transformers via Speculative Decoding. In Proceedings of the International Conference on Machine Learning, ICML 2023 PMLR; Honolulu, Hawaii, USA, Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., Scarlett, J., Eds.; Proceedings of Machine Learning Research; 23-29 July 2023; Vol. 202, pp. 19274–19286. [Google Scholar]
- Gloeckle, F.; Idrissi, B.Y.; Rozière, B.; Lopez-Paz, D.; Synnaeve, G. Better & Faster Large Language Models via Multi-token Prediction. In Proceedings of the Forty-first International Conference on Machine Learning, ICML 2024 PMLR / OpenReview.net; Vienna, Austria, Salakhutdinov, R., Kolter, Z., Heller, K.A., Weller, A., Oliver, N., Scarlett, J., Berkenkamp, F., Eds.; Proceedings of Machine Learning Research; 21-27 July 2024; Vol. 235, pp. 15706–15734. [Google Scholar]
- Shazeer, N. Fast Transformer Decoding: One Write-Head is All You Need. CoRR 2019, abs/1911.02150. [Google Scholar]
- Song, Y.; Zhang, Z.; Luo, C.; Gao, P.; Xia, F.; Luo, H.; Li, Z.; Yang, Y.; Yu, H.; Qu, X.; et al. Seed Diffusion: A Large-Scale Diffusion Language Model with High-Speed Inference. CoRR 2025, abs/2508.02193, 2508.02193. [Google Scholar] [CrossRef]
- Cai, T.; Li, Y.; Geng, Z.; Peng, H.; Lee, J.D.; Chen, D.; Dao, T. Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads. In Proceedings of the Forty-first International Conference on Machine Learning, ICML 2024 PMLR / OpenReview.net; Vienna, Austria, Salakhutdinov, R., Kolter, Z., Heller, K.A., Weller, A., Oliver, N., Scarlett, J., Berkenkamp, F., Eds.; Proceedings of Machine Learning Research; 21-27 July 2024; Vol. 235, pp. 5209–5235. [Google Scholar]
- Huang, Y.; Li, S.; Liu, M.; Liu, W.; Huang, S.; Fan, Z.; Chan, H.P.; Fung, Y.R. Environment Scaling for Interactive Agentic Experience Collection: A Survey. CoRR 2025, abs/2511.09586, 2511.09586. [Google Scholar] [CrossRef]
- Wang, X.J.; Bai, H.; Sun, Y.; Wang, H.; Zhang, S.; Hu, W.; Schroder, M.; Mutlu, B.; Song, D.; Nowak, R.D. The Long-Horizon Task Mirage? Diagnosing Where and Why Agentic Systems Break. CoRR 2026, abs/2604.11978, [2604.11978. [Google Scholar] [CrossRef]
- Kim, S.; Cho, J.; Kwak, B.; Kwon, T.; Wang, L.; Yang, N.; Zhang, X.; Wei, F.; Yeo, J. On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length. CoRR 2026, abs/2605.02572, 2605.02572. [Google Scholar] [CrossRef]
- Xie, J.; Xu, D.; Zhao, X.; Song, D. AgentSynth: Scalable Task Generation for Generalist Computer-Use Agents. CoRR 2025, abs/2506.14205, 2506.14205. [Google Scholar] [CrossRef]
- Wu, J.; Yin, W.; Jiang, Y.; Wang, Z.; Xi, Z.; Fang, R.; Zhang, L.; He, Y.; Zhou, D.; Xie, P.; et al. WebWalker: Benchmarking LLMs in Web Traversal. In Proceedings of the Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics; Vienna, Austria, Che, W., Nabende, J., Shutova, E., Pilehvar, M.T., Eds.; Association for Computational Linguistics, 2025; Volume 1, pp. 10290–10305. [Google Scholar] [CrossRef]
- Li, K.; Zhang, Z.; Yin, H.; Zhang, L.; Ou, L.; Wu, J.; Yin, W.; Li, B.; Tao, Z.; Wang, X.; et al. WebSailor: Navigating Super-human Reasoning for Web Agent. CoRR 2025, abs/2507.02592, 2507.02592. [Google Scholar] [CrossRef]
- Li, K.; Zhang, Z.; Yin, H.; Ye, R.; Zhao, Y.; Zhang, L.; Ou, L.; Zhang, D.; Wu, X.; Wu, J.; et al. WebSailor-V2: Bridging the Chasm to Proprietary Agents via Synthetic Data and Scalable Reinforcement Learning. CoRR 2025, abs/2509.13305, [2509.13305. [Google Scholar] [CrossRef]
- Li, Z.; Guan, X.; Zhang, B.; Huang, S.; Zhou, H.; Lai, S.; Yan, M.; Jiang, Y.; Xie, P.; Huang, F.; et al. WebWeaver: Structuring Web-Scale Evidence with Dynamic Outlines for Open-Ended Deep Research. CoRR 2025, abs/2509.13312, 2509.13312. [Google Scholar] [CrossRef]
- Jain, N.; Shetty, M.; Zhang, T.; Han, K.; Sen, K.; Stoica, I. R2E: Turning any Github Repository into a Programming Agent Environment. In Proceedings of the Forty-first International Conference on Machine Learning, ICML 2024 PMLR / OpenReview.net; Vienna, Austria, Salakhutdinov, R., Kolter, Z., Heller, K.A., Weller, A., Oliver, N., Scarlett, J., Berkenkamp, F., Eds.; Proceedings of Machine Learning Research; 21-27 July 2024; Vol. 235, pp. 21196–21224. [Google Scholar]
- Chen, A.; Zhang, C.; Liu, J.; Chen, J.; Du, C.; Li, Y.; Zhong, M.; Wang, Q.; Zhu, Z.; Song, J.; et al. DIVE: Scaling Diversity in Agentic Task Synthesis for Generalizable Tool Use. CoRR 2026, abs/2603.11076, [2603.11076. [Google Scholar] [CrossRef]
- Keren, T.; Calderon, N.; Yehudai, A.; Perlitz, Y.; Shmueli-Scheuer, M.; Reichert, R. A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks. CoRR 2026, abs/2605.28556, [2605.28556. [Google Scholar] [CrossRef]
- Lin, Y.; Wang, H.; Wu, S.; Fan, L.; Pan, F.; Zhao, S.; Tu, D. CLI-Gym: Scalable CLI Task Generation via Agentic Environment Inversion. CoRR 2026, abs/2602.10999, [2602.10999. [Google Scholar] [CrossRef]
- Qi, Z.; Liu, X.; Iong, I.L.; Lai, H.; Sun, X.; Sun, J.; Yang, X.; Yang, Y.; Yao, S.; Xu, W.; et al. WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning. In Proceedings of the The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025; 2025. Available online: https://openreview.net/.
- Su, H.; Sun, R.; Yoon, J.; Yin, P.; Yu, T.; Arik, S.Ö. Learn-by-interact: A Data-Centric Framework For Self-Adaptive Agents in Realistic Environments. In Proceedings of the The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025; 2025. Available online: https://openreview.net/.
- Chen, X.; Qiao, Z.; Chen, G.; Su, L.; Zhang, Z.; Wang, X.; Xie, P.; Huang, F.; Zhou, J.; Jiang, Y. AgentFrontier: Expanding the Capability Frontier of LLM Agents with ZPD-Guided Data Synthesis. CoRR 2025, abs/2510.24695, 2510.24695. [Google Scholar] [CrossRef]
- Mai, S.; Zhai, Y.; Chen, Z.; Chen, C.; Zou, A.; Tao, S.; Liu, Z.; Ding, B. CuES: A Curiosity-driven and Environment-grounded Synthesis Framework for Agentic RL. CoRR 2025, abs/2512.01311, 2512.01311. [Google Scholar] [CrossRef]
- Fan, L.; Wang, G.; Jiang, Y.; Mandlekar, A.; Yang, Y.; Zhu, H.; Tang, A.; Huang, D.; Zhu, Y.; Anandkumar, A. MineDojo: Building Open-Ended Embodied Agents with Internet-Scale Knowledge. In Proceedings of the Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022; New Orleans, LA, USA, Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., Oh, A., Eds.; 28 November 2022. [Google Scholar]
- Zhao, J.; Chen, G.; Meng, F.; Li, M.; Chen, J.; Xu, H.; Sun, Y.; Zhao, X.; Song, R.; Zhang, Y.; et al. Immersion in the GitHub Universe: Scaling Coding Agents to Mastery. CoRR 2026, abs/2602.09892, 2602.09892. [Google Scholar] [CrossRef]
- Zeng, Y.; Li, S.; Dong, D.; Xu, R.; Chen, Z.; Zheng, L.; Li, Y.; Zhou, Z.; Zhao, H.; Tian, L.; et al. SWE-Hub: A Unified Production System for Scalable, Executable Software Engineering Tasks. CoRR 2026, abs/2603.00575, 2603.00575. [Google Scholar] [CrossRef]
- Cheng, Z.; Wang, H.; Liu, Z.; Wang, X.; Zhu, X.; Guo, Y.; Lin, W.; Pan, J.Z.; Wang, Y. Terminal-World: Scaling Terminal-Agent Environments via Agent Skills. CoRR 2026, abs/2605.20876, [2605.20876. [Google Scholar] [CrossRef]
- Cheng, D.; Huang, S.; Gu, Y.; Song, H.; Chen, G.; Dong, L.; Zhao, W.X.; Wen, J.R.; Wei, F. Computer Environments Elicit General Agentic Intelligence in LLMs. arXiv 2026, arXiv:cs. [Google Scholar]
- Ouyang, M.; Hu, S.; Lin, K.Q.; Ng, H.T.; Shou, M.Z. GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents. CoRR 2026, abs/2604.07429, [2604.07429. [Google Scholar] [CrossRef]
- Jia, H.; Liao, J.; Zhang, X.; Xu, H.; Xie, T.; Jiang, C.; Yan, M.; Liu, S.; Ye, W.; Huang, F. OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use Agents. CoRR 2025, abs/2510.24563, 2510.24563. [Google Scholar] [CrossRef]
- Wu, Y.; Peng, Y.; Chen, Y.; Ruan, J.; Zhuang, Z.; Yang, C.; Zhang, J.; Chen, M.; Tseng, Y.; Yu, Z.; et al. AutoWebWorld: Synthesizing Infinite Verifiable Web Environments via Finite State Machines. CoRR 2026, abs/2602.14296, [2602.14296. [Google Scholar] [CrossRef]
- Zhang, Z.; Wang, Z.; Zhang, X.; Guo, Z.; Li, J.; Li, B.; Lu, Y. InfiniteWeb: Scalable Web Environment Synthesis for GUI Agent Training. In Proceedings of the Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2026; San Diego, California, United States, Liakata, M., Moreira, V.P., Zhang, J., Jurgens, D., Eds.; Association for Computational Linguistics, 2-7 July 2026; Volume 2026, pp. 28465–28492. [Google Scholar]
- Fang, R.; Cai, S.; Li, B.; Wu, J.; Li, G.; Yin, W.; Wang, X.; Wang, X.; Su, L.; Zhang, Z.; et al. Towards General Agentic Intelligence via Environment Scaling. In Proceedings of the Findings of the Association for Computational Linguistics, ACL 2026; San Diego, California, United States, Liakata, M., Moreira, V.P., Zhang, J., Jurgens, D., Eds.; Association for Computational Linguistics, 2-7 July 2026; Volume 2026, pp. 17610–17621. [Google Scholar]
- Xu, R.; Zhuang, Y.; Zhong, Y.; Yu, Y.; Wang, Z.; Tang, X.; Wu, H.; Wang, M.D.; Ruan, P.; Yang, D.; et al. MedAgentGym: A Scalable Agentic Training Environment for Code-Centric Reasoning in Biomedical Data Science. arXiv 2025, arXiv:cs. [Google Scholar]
- Shen, Y.; Yang, Y.; Xi, Z.; Hu, B.; Sha, H.; Zhang, J.; Peng, Q.; Shang, J.; Huang, J.; Fan, Y.; et al. SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents. CoRR 2026, abs/2602.12984, 2602.12984. [Google Scholar] [CrossRef]
- Wang, D.; Cheng, M.; Liu, Q.; Yu, S.; Liu, Z.; Guo, Z. PaperArena: An Evaluation Benchmark for Tool-Augmented Agentic Reasoning on Scientific Literature. CoRR 2025, abs/2510.10909, [2510.10909. [Google Scholar] [CrossRef]
- Alonso, E.; Jelley, A.; Micheli, V.; Kanervisto, A.; Storkey, A.J.; Pearce, T.; Fleuret, F. Diffusion for World Modeling: Visual Details Matter in Atari. In Proceedings of the Advances in Neural Information Processing Systems 37: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024; Vancouver, BC, Canada, Globersons, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J.M., Zhang, C., Eds.; 10 - 15 December 2024. [Google Scholar]
- Valevski, D.; Leviathan, Y.; Arar, M.; Fruchter, S. Diffusion Models Are Real-Time Game Engines. In Proceedings of the The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025; 2025. Available online: https://openreview.net/.
- Guo, J.; Ye, Y.; He, T.; Wu, H.; Jiang, Y.; Pearce, T.; Bian, J. MineWorld: a Real-Time and Open-Source Interactive World Model on Minecraft. CoRR 2025, abs/2504.08388, 2504.08388. [Google Scholar] [CrossRef]
- Zhang, Y.; Peng, C.; Wang, B.; Wang, P.; Zhu, Q.; Kang, F.; Jiang, B.; Gao, Z.; Li, E.; Liu, Y.; et al. Matrix-Game: Interactive World Foundation Model. CoRR 2025, abs/2506.18701, [2506.18701. [Google Scholar] [CrossRef]
- Chi, X.; Fan, C.; Zhang, H.; Qi, X.; Zhang, R.; Chen, A.; Chan, C.; Xue, W.; Liu, Q.; Zhang, S.; et al. Empowering World Models with Reflection for Embodied Video Prediction. In Proceedings of the Forty-second International Conference on Machine Learning, ICML 2025; Vancouver, BC, Canada, Singh, A., Fazel, M., Hsu, D., Lacoste-Julien, S., Berkenkamp, F., Maharaj, T., Wagstaff, K., Zhu, J., Eds.; PMLR / OpenReview.net; Proceedings of Machine Learning Research; 13-19 July 2025; Vol. 267. [Google Scholar]
- Jang, J.; Ye, S.; Lin, Z.; Xiang, J.; Bjorck, J.; Fang, Y.; Hu, F.; Huang, S.; Kundalia, K.; Lin, Y.C.; et al. DreamGen: Unlocking Generalization in Robot Learning through Video World Models. arXiv 2025, arXiv:cs. [Google Scholar]
- Ye, S.; Ge, Y.; Zheng, K.; Gao, S.; Yu, S.; Kurian, G.; Indupuru, S.; Tan, Y.L.; Zhu, C.; Xiang, J.; et al. World Action Models are Zero-shot Policies. CoRR 2026, abs/2602.15922, 2602.15922. [Google Scholar] [CrossRef]
- Agarwal, N.; Ali, A.; Bala, M.; Balaji, Y.; Barker, E.; Cai, T.; Chattopadhyay, P.; Chen, Y.; Cui, Y.; Ding, Y.; et al. Cosmos World Foundation Model Platform for Physical AI. CoRR 2025, abs/2501.03575, 2501.03575. [Google Scholar] [CrossRef]
- Gao, S.; Zhou, S.; Du, Y.; Zhang, J.; Gan, C. AdaWorld: Learning Adaptable World Models with Latent Actions. In Proceedings of the Forty-second International Conference on Machine Learning, ICML 2025; Vancouver, BC, Canada, Singh, A., Fazel, M., Hsu, D., Lacoste-Julien, S., Berkenkamp, F., Maharaj, T., Wagstaff, K., Zhu, J., Eds.; PMLR / OpenReview.net; Proceedings of Machine Learning Research; 13-19 July 2025; Vol. 267. [Google Scholar]
- Gu, Y.; Zhang, K.; Ning, Y.; Zheng, B.; Gou, B.; Xue, T.; Chang, C.; Srivastava, S.; Xie, Y.; Qi, P.; et al. Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents. Trans. Mach. Learn. Res. 2025, 2025. [Google Scholar]
- Wang, Y.; Yin, D.; Cui, Y.; Zheng, R.; Li, Z.; Lin, Z.; Wu, D.; Wu, X.; Ye, C.; Zhou, Y.; et al. LLMs as Scalable, General-Purpose Simulators For Evolving Digital Agent Training. CoRR 2025, abs/2510.14969, 2510.14969. [Google Scholar] [CrossRef]
- Fang, T.; Zhang, H.; Zhang, Z.; Ma, K.; Yu, W.; Mi, H.; Yu, D. WebEvolver: Enhancing Web Agent Self-Improvement with Co-evolving World Model. In Proceedings of the Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025; Suzhou, China, Christodoulopoulos, C., Chakraborty, T., Rose, C., Peng, V., Eds.; Association for Computational Linguistics, 4-9 November 2025; pp. 8959–8975. [Google Scholar] [CrossRef]
- Xiao, Z.; Tu, J.; Zou, C.; Zuo, Y.; Li, Z.; Wang, P.; Yu, B.; Huang, F.; Lin, J.; Liu, Z. WebWorld: A Large-Scale World Model for Web Agent Training. CoRR 2026, abs/2602.14721, 2602.14721. [Google Scholar] [CrossRef]
- Zheng, Y.; Zhong, L.; Wang, Y.; Dai, R.; Liu, K.; Chu, X.; Lv, L.; Torr, P.; Lin, K.Q. Code2World: A GUI World Model via Renderable Code Generation. CoRR 2026, abs/2602.09856, 2602.09856. [Google Scholar] [CrossRef]
- Cao, Y.; Zhong, Y.; Zeng, Z.; Zheng, L.; Huang, J.; Qiu, H.; Shi, P.; Mao, W.; Guanglu, W. MobileDreamer: Generative Sketch World Model for GUI Agent. CoRR 2026, abs/2601.04035, 2601.04035. [Google Scholar] [CrossRef]
- Zuo, Y.; Xiao, Z.; Sheng, L.; Huang, F.; Tu, J.; Liu, Y.; Tang, T.; Hu, X.; Su, Y.; Lan, Q.; et al. Qwen-AgentWorld: Language World Models for General Agents. CoRR 2026, abs/2606.24597, [2606.24597. [Google Scholar] [CrossRef]
- Guo, J.; Yang, L.; Chen, P.; Xiao, Q.; Wang, Y.; Juan, X.; Qiu, J.; Shen, K.; Wang, M. GenEnv: Difficulty-Aligned Co-Evolution Between LLM Agents and Environment Simulators. CoRR 2025, abs/2512.19682, 2512.19682. [Google Scholar] [CrossRef]
- Assran, M.; Duval, Q.; Misra, I.; Bojanowski, P.; Vincent, P.; Rabbat, M.G.; LeCun, Y.; Ballas, N. Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, 2023; IEEE; 2023, pp. 15619–15629. [Google Scholar] [CrossRef]
- Ghaemi, H.; Muller, E.B.; Bakhtiari, S. seq-JEPA: Autoregressive Predictive Learning of Invariant-Equivariant World Models. CoRR 2025, abs/2505.03176, 2505.03176. [Google Scholar] [CrossRef]
- Baldassarre, F.; Szafraniec, M.; Terver, B.; Khalidov, V.; Massa, F.; LeCun, Y.; Labatut, P.; Seitzer, M.; Bojanowski, P. Back to the Features: DINO as a Foundation for Video World Models. CoRR 2025, abs/2507.19468, [2507.19468. [Google Scholar] [CrossRef]
- Assran, M.; Bardes, A.; Fan, D.; Garrido, Q.; Howes, R.; Komeili, M.; Muckley, M.J.; Rizvi, A.; Roberts, C.; Sinha, K.; et al. V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning. CoRR 2025, abs/2506.09985, [2506.09985. [Google Scholar] [CrossRef]
- Trivedi, H.; Khot, T.; Hartmann, M.; Manku, R.; Dong, V.; Li, E.; Gupta, S.; Sabharwal, A.; Balasubramanian, N. AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents. In Proceedings of the Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics; Bangkok, Thailand, Ku, L., Martins, A., Srikumar, V., Eds.; Association for Computational Linguistics, 2024; Volume 1, pp. 16022–16076. [Google Scholar] [CrossRef]
- Lei, F.; Yang, Y.; Sun, W.; Lin, D. MCPVerse: An Expansive, Real-World Benchmark for Agentic Tool Use. CoRR 2025, abs/2508.16260, 2508.16260. [Google Scholar] [CrossRef]
- Wang, Z.; Chang, Q.; Patel, H.; Biju, S.; Wu, C.; Liu, Q.; Ding, A.; Rezazadeh, A.; Shah, A.; Bao, Y.; et al. MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers. CoRR 2025, abs/2508.20453, [2508.20453. [Google Scholar] [CrossRef]
- Chae, H.; Park, J.; Ritter, A. Safe and Scalable Web Agent Learning via Recreated Websites. CoRR 2026, abs/2603.10505, [2603.10505. [Google Scholar] [CrossRef]
- Zhu, Y.; Huang, Y.; Qin, S.; Huang, Z.; Zhang, S.; Zhang, X. MedMCP-Calc: Benchmarking LLMs for Realistic Medical Calculator Scenarios via MCP Integration. In Proceedings of the Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics; San Diego, California, United States, Liakata, M., Moreira, V.P., Zhang, J., Jurgens, D., Eds.; Association for Computational Linguistics, 2-7 July 2026; Volume 1, pp. 4845–4873. [Google Scholar]
- Zhang, C.; Gan, Z.; Zhu, L.; Pang, Y.; Zhang, Q.; Zhang, R. FinMTM: A Multi-Turn Multimodal Benchmark for Financial Reasoning and Agent Evaluation. CoRR 2026, abs/2602.03130, 2602.03130. [Google Scholar] [CrossRef]
- Lù, X.H.; Kazemnejad, A.; Meade, N.; Patel, A.; Shin, D.; Zambrano, A.; Stanczak, K.; Shaw, P.; Pal, C.J.; Reddy, S. AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories. CoRR 2025, abs/2504.08942, 2504.08942. [Google Scholar] [CrossRef]
- Li, Z.; Jiang, D.; Ma, X.; Zhang, H.; Nie, P.; Zhang, Y.; Zou, K.; Xie, J.; Zhang, Y.; Chen, W. OpenResearcher: A Fully Open Pipeline for Long-Horizon Deep Research Trajectory Synthesis. CoRR 2026 2027, abs/2603.20278. [Google Scholar] [CrossRef]
- Tian, X.; Wang, H.; Chen, S.; Zhou, H.; Yu, K.; Zhang, Y.; Ouyang, J.; Yin, J.; Chen, J.; Guo, B.; et al. ASTRA: Automated Synthesis of agentic Trajectories and Reinforcement Arenas. CoRR 2026, abs/2601.21558, [2601.21558. [Google Scholar] [CrossRef]
- He, Y.; Chawla, P.; Souri, Y.; Som, S.; Song, X. WebSTAR: Scalable Data Synthesis for Computer Use Agents with Step-Level Filtering. In Proceedings of the Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics; San Diego, California, United States, Liakata, M., Moreira, V.P., Zhang, J., Jurgens, D., Eds.; Association for Computational Linguistics, 2-7 July 2026; Volume 1, pp. 516–533. [Google Scholar]
- Wei, J.; Zhao, Y.; Ni, K.; Cohan, A. Anchor: Branch-Point Data Generation for GUI Agents. In Proceedings of the Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2026; San Diego, California, United States, Liakata, M., Moreira, V.P., Zhang, J., Jurgens, D., Eds.; Association for Computational Linguistics, 2-7 July 2026; Volume 2026, pp. 17031–17047. [Google Scholar]
- Shao, R.; Gao, R.; Xie, B.; Li, Y.; Zhou, K.; Wang, S.; Guan, W.; Chen, G. HATS: Hardness-Aware Trajectory Synthesis for GUI Agents. CoRR 2026, abs/2603.12138, [2603.12138. [Google Scholar] [CrossRef]
- Gao, J.; Fu, W.; Xie, M.; Xu, S.; He, C.; Mei, Z.; Zhu, B.; Wu, Y. Beyond Ten Turns: Unlocking Long-Horizon Agentic Search with Large-Scale Asynchronous RL. CoRR 2025, abs/2508.07976, [2508.07976. [Google Scholar] [CrossRef]
- Zhang, J.; Huang, C.; Guo, F.; Li, Z.; Shi, K.; Jiang, M.; Yu, J.; Shang, S.; Gao, S. DPEPO: Diverse Parallel Exploration Policy Optimization for LLM-based Agents. In Proceedings of the Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics; San Diego, California, United States, Liakata, M., Moreira, V.P., Zhang, J., Jurgens, D., Eds.; Association for Computational Linguistics, 2-7 July 2026; Volume 1, pp. 46367–46389. [Google Scholar]
- Luo, X.; Zhang, Y.; He, Z.; Wang, Z.; Zhao, S.; Li, D.; Qiu, L.K.; Yang, Y. Agent Lightning: Train ANY AI Agents with Reinforcement Learning. CoRR 2025, abs/2508.03680, 2508.03680. [Google Scholar] [CrossRef]
- Yuan, H.; Xu, Z.; Wang, H.; Yi, X.; Gao, J.; Zhang, X.; Wang, Y.; Yu, C.; Wu, Y. Verifiable Process Rewards for Agentic Reasoning. CoRR 2026, abs/2605.10325, 2605.10325. [Google Scholar] [CrossRef]
- Li, S.; Huang, Y.; Liu, Z.; Li, Y.; Fu, J.; Zhao, L.; Bian, J.; Zhang, L.; Zhang, J.; Wang, R. GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation. CoRR 2026, abs/2605.11853, [2605.11853. [Google Scholar] [CrossRef]
- Ding, L. AgentHER: Hindsight Experience Replay for LLM Agent Trajectory Relabeling. CoRR 2026, abs/2603.21357, [2603.21357. [Google Scholar] [CrossRef]
- Yang, P.; Chen, W.; Zheng, A.Y.; Li, X.; Li, X.; Tu, H.; Xiao, J.; Pang, Y.; Zhang, D.; Li, F.; et al. AOI: Turning Failed Trajectories into Training Signals for Autonomous Cloud Diagnosis. CoRR 2026, abs/2603.03378, 2603.03378. [Google Scholar] [CrossRef]
- Zhang, T.; Popa, A.; Xu, Y.; Song, R.; Dimitriadis, D. PIVOT: Bridging Planning and Execution in LLM Agents via Trajectory Refinement. CoRR 2026, abs/2605.11225, [2605.11225. [Google Scholar] [CrossRef]
- Chen, M.; Wang, J.; Liu, Z.; Wang, Y.; Wang, Q. From Failed Trajectories to Reliable LLM Agents: Diagnosing and Repairing Harness Flaws. CoRR 2026, abs/2606.06324, [2606.06324. [Google Scholar] [CrossRef]
- Yan, S.; Bahloul, A.; Nie, E.; Schwarzmann, S.; Trivisonno, R.; Tresp, V.; Ma, Y. Memory-R2: Fair Credit Assignment for Long-Horizon Memory-Augmented LLM Agents. CoRR 2026, abs/2605.21768, [2605.21768. [Google Scholar] [CrossRef]
- Chen, S.; Wong, S.; Chen, L.; Tian, Y. Extending Context Window of Large Language Models via Positional Interpolation. CoRR 2023, abs/2306.15595, 2306.15595. [Google Scholar] [CrossRef]
- Gao, T.; Wettig, A.; Yen, H.; Chen, D. How to Train Long-Context Language Models (Effectively). In Proceedings of the Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025; Vienna, Austria, July 27 - August 1, Che, W., Nabende, J., Shutova, E., Pilehvar, M.T., Eds.; Association for Computational Linguistics, 2025; Volume 2025, pp. 7376–7399. [Google Scholar] [CrossRef]
- Xu, C.; Ping, W.; Xu, P.; Liu, Z.; Wang, B.; Shoeybi, M.; Li, B.; Catanzaro, B. From 128K to 4M: Efficient Training of Ultra-Long Context Large Language Models. In Proceedings of the Findings of the Association for Computational Linguistics, ACL 2026; San Diego, California, United States, Liakata, M., Moreira, V.P., Zhang, J., Jurgens, D., Eds.; Association for Computational Linguistics, 2-7 July 2026; Volume 2026, pp. 13122–13133. [Google Scholar]
- Bai, Y.; Zhang, J.; Lv, X.; Zheng, L.; Zhu, S.; Hou, L.; Dong, Y.; Tang, J.; Li, J. LongWriter: Unleashing 10, 000+ Word Generation from Long Context LLMs. In Proceedings of the The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025; 2025. Available online: https://openreview.net/.
- Vodrahalli, K.; Ontanon, S.; Tripuraneni, N.; Xu, K.; Jain, S.; Shivanna, R.; Hui, J.; Dikkala, N.; Kazemi, M.; Fatemi, B.; et al. Michelangelo: Long Context Evaluations Beyond Haystacks via Latent Structure Queries. CoRR 2024, abs/2409.12640, [2409.12640. [Google Scholar] [CrossRef]
- Chen, L.; Xie, W.; Liang, Y.; He, H.; Zhao, H.; Yang, Z.; Huang, Z.; Wu, H.; Lu, H.; Charles, Y.; et al. BabyVision: Visual Reasoning Beyond Language. CoRR 2026, abs/2601.06521, 2601.06521. [Google Scholar] [CrossRef]
- Vo, A.; Nguyen, K.; Taesiri, M.R.; Dang, V.T.; Nguyen, A.T.; Kim, D. Vision Language Models are Biased. CoRR 2025, abs/2505.23941, [2505.23941. [Google Scholar] [CrossRef]
- Rahmanzadehgervi, P.; Bolton, L.; Taesiri, M.R.; Nguyen, A.T. Vision language models are blind: Failing to translate detailed visual features into words. arXiv 2025, arXiv:cs. [Google Scholar]
- Tong, S.; Liu, Z.; Zhai, Y.; Ma, Y.; LeCun, Y.; Xie, S. Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024, Seattle, WA, USA, June 16-22, 2024; IEEE; 2024, pp. 9568–9578. [Google Scholar] [CrossRef]
- Tong, P.; Brown, E.; Wu, P.; Woo, S.; Iyer, A.; Akula, S.C.; Yang, S.; Yang, J.; Middepogu, M.; Wang, Z.; et al. Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs. In Proceedings of the Advances in Neural Information Processing Systems 37: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024; Vancouver, BC, Canada, Globersons, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J.M., Zhang, C., Eds.; 10 - 15 December 2024. [Google Scholar]
- Wang, W.; Gao, Z.; Gu, L.; Pu, H.; Cui, L.; Wei, X.; Liu, Z.; Jing, L.; Ye, S.; Shao, J.; et al. InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency. CoRR 2025, abs/2508.18265, [2508.18265. [Google Scholar] [CrossRef]
- Guo, D.; Wu, F.; Zhu, F.; Leng, F.; Shi, G.; Chen, H.; Fan, H.; Wang, J.; Jiang, J.; Wang, J.; et al. Seed1.5-VL Technical Report. arXiv 2025, arXiv:cs. [Google Scholar]
- ByteDance Seed Team. Seed1.6 Tech Introduction. 2025. [Google Scholar]
- LLM-Core Team, Xiaomi. MiMo: Unlocking the Reasoning Potential of Language Model – From Pretraining to Posttraining. CoRR 2025, abs/2505.07608, 2505.07608. [Google Scholar] [CrossRef]
- LLM-Core Team, Xiaomi. MiMo-VL Technical Report. CoRR 2025, abs/2506.03569, [2506.03569. [Google Scholar] [CrossRef]
- Qwen Team. Qwen3.5-Omni Technical Report. CoRR 2026, abs/2604.15804, [2604.15804. [Google Scholar] [CrossRef]
- Team, K. Kimi K2.5: Visual Agentic Intelligence. CoRR 2026, abs/2602.02276, 2602.02276. [Google Scholar] [CrossRef]
- Li, J.; Fang, A.; Smyrnis, G.; Ivgi, M.; Jordan, M.; Gadre, S.Y.; Bansal, H.; Guha, E.K.; Keh, S.S.; Arora, K.; et al. DataComp-LM: In search of the next generation of training sets for language models. In Proceedings of the Advances in Neural Information Processing Systems 37: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024; Vancouver, BC, Canada, Globersons, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J.M., Zhang, C., Eds.; 10 - 15 December 2024. [Google Scholar]
- Penedo, G.; Kydlícek, H.; Allal, L.B.; Lozhkov, A.; Mitchell, M.; Raffel, C.A.; von Werra, L.; Wolf, T. The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale. In Proceedings of the Advances in Neural Information Processing Systems 37: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024; Vancouver, BC, Canada, Globersons, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J.M., Zhang, C., Eds.; 10 - 15 December 2024. [Google Scholar]
- Soldaini, L.; Kinney, R.; Bhagia, A.; Schwenk, D.; Atkinson, D.; Authur, R.; Bogin, B.; Chandu, K.R.; Dumas, J.; Elazar, Y.; et al. Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research. In Proceedings of the Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics; Bangkok, Thailand, Ku, L., Martins, A., Srikumar, V., Eds.; Association for Computational Linguistics, 2024; Volume 1, pp. 15725–15788. [Google Scholar] [CrossRef]
- Groeneveld, D.; Beltagy, I.; Walsh, E.P.; Bhagia, A.; Kinney, R.; Tafjord, O.; Jha, A.H.; Ivison, H.; Magnusson, I.; Wang, Y.; et al. OLMo: Accelerating the Science of Language Models. In Proceedings of the Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers); ACL 2024, Ku, L., Martins, A., Srikumar, V., Eds.; Association for Computational Linguistics, 11-16 August 2024; Volume 2024, pp. 15789–15809. [Google Scholar] [CrossRef]
- Ye, J.; Liu, P.; Sun, T.; Zhan, J.; Zhou, Y.; Qiu, X. Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance. In Proceedings of the The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025; 2025. Available online: https://openreview.net/.
- Chen, Z.; Liu, K.; Wang, Q.; Zhang, W.; Liu, J.; Lin, D.; Chen, K.; Zhao, F. Agent-FLAN: Designing Data and Methods of Effective Agent Tuning for Large Language Models. In Association for Computational Linguistics; Proceedings of the Findings of the Association for Computational Linguistics, ACL 2024, Bangkok, Thailand and virtual meeting, Ku, L., Martins, A., Srikumar, V., Eds.; Findings of ACL; 2024; Vol. ACL 2024, pp. 9354–9366. [Google Scholar] [CrossRef]
- Song, Y.; Xiong, W.; Zhao, X.; Zhu, D.; Wu, W.; Wang, K.; Li, C.; Peng, W.; Li, S. AgentBank: Towards Generalized LLM Agents via Fine-Tuning on 50000+ Interaction Trajectories. In Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2024 Association for Computational Linguistics; Miami, Florida, USA, Al-Onaizan, Y., Bansal, M., Chen, Y., Eds.; Findings of ACL; 12-16 November 2024; Vol. EMNLP 2024, pp. 2124–2141. [Google Scholar] [CrossRef]
- Chen, B.; Shu, C.; Shareghi, E.; Collier, N.; Narasimhan, K.; Yao, S. FireAct: Toward Language Agent Fine-tuning. CoRR 2023, abs/2310.05915, 2310.05915. [Google Scholar] [CrossRef]
- Chen, Z.; Li, M.; Huang, Y.; Du, Y.; Fang, M.; Zhou, T. ATLAS: Agent Tuning via Learning Critical Steps. In Proceedings of the Findings of the Association for Computational Linguistics, ACL 2025 Association for Computational Linguistics; Vienna, Austria, Che, W., Nabende, J., Shutova, E., Pilehvar, M.T., Eds.; 27 July 2025; Vol. ACL 2025, Findings of ACL, pp. 25334–25349. [Google Scholar] [CrossRef]
- Ye, Y.; Huang, Z.; Xiao, Y.; Chern, E.; Xia, S.; Liu, P. LIMO: Less is More for Reasoning. CoRR 2025, abs/2502.03387, 2502.03387. [Google Scholar] [CrossRef]
- Li, X.; Zou, H.; Liu, P. LIMR: Less is More for RL Scaling. CoRR 2025, abs/2502.11886, 2502.11886. [Google Scholar] [CrossRef]
- Li, X.; Ma, Y.; Li, C.; Zhu, F.; Yu, Y.; Bao, K.; Wang, W.; Feng, F.; Liu, D. Unified Data Selection for LLM Reasoning. CoRR 2026, abs/2605.22389, [2605.22389. [Google Scholar] [CrossRef]
- Ding, B.; Chen, Y.; Lyu, J.; Yuan, J.; Zhu, Q.; Tian, S.; Zhu, D.; Wang, F.; Deng, H.; Mi, F.; et al. Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning. In Proceedings of the Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2026; San Diego, California, United States, Liakata, M., Moreira, V.P., Zhang, J., Jurgens, D., Eds.; Association for Computational Linguistics, 2-7 July 2026; Volume 2026, pp. 33081–33106. [Google Scholar]
- Zhang, Y.; Xiong, G.; Li, H.; Zhao, W. EDGE: Efficient Data Selection for LLM Agents via Guideline Effectiveness. In Proceedings of the Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence, IJCAI 2025, Montreal, Canada; 2025. ijcai.org, August 16-22; 2025, pp. 8384–8392. [CrossRef]
- Song, Y.; Ramaneti, K.; Sheikh, Z.; Chen, Z.; Gou, B.; Xie, T.; Xu, Y.; Zhang, D.; Gandhi, A.; Yang, F.; et al. Agent Data Protocol: Unifying Datasets for Diverse, Effective Fine-tuning of LLM Agents. CoRR 2025, abs/2510.24702, [2510.24702. [Google Scholar] [CrossRef]
- Liu, J.; Kong, Z.; Dong, P.; Yang, C.; Li, T.; Tang, H.; Yuan, G.; Niu, W.; Zhang, W.; Zhao, P.; et al. Structured Agent Distillation for Large Language Model. CoRR 2025, abs/2505.13820, [2505.13820. [Google Scholar] [CrossRef]
- Luo, Y.; Jin, Y.; Yu, W.; Zhang, M.; Kumar, S.; Li, X.; Xu, W.; Chen, X.; Wang, J. AgentArk: Distilling Multi-Agent Intelligence into a Single LLM Agent. CoRR 2026, abs/2602.03955, 2602.03955. [Google Scholar] [CrossRef]
- Xiong, W.; Song, Y.; Zhao, X.; Wu, W.; Wang, X.; Wang, K.; Li, C.; Peng, W.; Li, S. Watch Every Step! LLM Agent Learning via Iterative Step-level Process Refinement. In Proceedings of the Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024; Miami, FL, USA, Al-Onaizan, Y., Bansal, M., Chen, Y., Eds.; Association for Computational Linguistics, 12-16 November 2024; pp. 1556–1572. [Google Scholar] [CrossRef]
- Zhang, Y.; Li, S.; Yu, C.; Lu, Q.; Jin, S.; Dong, C.; Liu, H.; Hong, I.; Li, X.; Shi, Z.; et al. Learning with Rare Success but Rich Feedback via Reflection-Enhanced Self-Distillation. CoRR 2026, abs/2605.12741, 2605.12741. [Google Scholar] [CrossRef]
- Liu, H.; Zhang, Y.; Li, X.; Lyu, B.; Shang, J. HERO: Hindsight-Enhanced Reflection from Environment Observations for Agentic Self-Distillation. CoRR 2026, abs/2606.11559, 2606.11559. [Google Scholar] [CrossRef]
- Jiang, P.; Lin, J.; Cao, L.; Tian, R.; Kang, S.; Wang, Z.; Sun, J.; Han, J. DeepRetrieval: Hacking Real Search Engines and Retrievers with Large Language Models via Reinforcement Learning. CoRR 2025, abs/2503.00223, 2503.00223. [Google Scholar] [CrossRef]
- Dong, G.; Chen, Y.; Li, X.; Jin, J.; Qian, H.; Zhu, Y.; Mao, H.; Zhou, G.; Dou, Z.; Wen, J. Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning. CoRR 2025, abs/2505.16410, [2505.16410. [Google Scholar] [CrossRef]
- Zhang, Y.; Huang, H.; Song, Z.; Zhao, Z.; Zhang, Q.; Zhu, Y.; Zhao, D. CriticSearch: Fine-Grained Credit Assignment for Search Agents via a Retrospective Critic. In Proceedings of the Findings of the Association for Computational Linguistics, ACL 2026; San Diego, California, United States, Liakata, M., Moreira, V.P., Zhang, J., Jurgens, D., Eds.; Association for Computational Linguistics, 2-7 July 2026; pp. 12272–12290. [Google Scholar]
- Gunjal, A.; Wang, A.; Lau, E.; Nath, V.; Liu, B.; Hendryx, S. Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains. CoRR 2025, abs/2507.17746, [2507.17746. [Google Scholar] [CrossRef]
- Liu, T.; Xu, R.; Yu, T.; Hong, I.; Yang, C.; Zhao, T.; Wang, H. OpenRubrics: Towards Scalable Synthetic Rubric Generation for Reward Modeling and LLM Alignment. In Proceedings of the Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics; San Diego, California, United States, Liakata, M., Moreira, V.P., Zhang, J., Jurgens, D., Eds.; Association for Computational Linguistics, 2-7 July 2026; Volume 1, pp. 17417–17437. [Google Scholar]
- Shao, R.; Asai, A.; Shen, S.Z.; Ivison, H.; Kishore, V.; Zhuo, J.; Zhao, X.; Park, M.; Finlayson, S.G.; Sontag, D.A.; et al. DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research. CoRR 2025, abs/2511.19399, [2511.19399. [Google Scholar] [CrossRef]
- Hu, J. REINFORCE++: A Simple and Efficient Approach for Aligning Large Language Models. CoRR 2025, abs/2501.03262, 2501.03262. [Google Scholar] [CrossRef]
- Liu, Z.; Chen, C.; Li, W.; Qi, P.; Pang, T.; Du, C.; Lee, W.S.; Lin, M. Understanding R1-Zero-Like Training: A Critical Perspective. CoRR 2025, abs/2503.20783, [2503.20783. [Google Scholar] [CrossRef]
- Zheng, C.; Liu, S.; Li, M.; Chen, X.; Yu, B.; Gao, C.; Dang, K.; Liu, Y.; Men, R.; Yang, A.; et al. Group Sequence Policy Optimization. CoRR 2025, abs/2507.18071, 2507.18071. [Google Scholar] [CrossRef]
- Li, J.; Zhou, P.; Meng, R.; Vadera, M.P.; Li, L.; Li, Y. Turn-PPO: Turn-Level Advantage Estimation with PPO for Improved Multi-Turn RL in Agentic LLMs. In Proceedings of the Findings of the Association for Computational Linguistics: EACL 2026 Association for Computational Linguistics; Rabat, Morocco, Demberg, V., Inui, K., Marquez, L., Eds.; Findings of ACL, 24-29 March 2026; pp. 6227–6243. [Google Scholar] [CrossRef]
- Wang, D.; Li, Q.; Cheng, M.; Ouyang, J.; Yu, S.; Liu, Q.; Chen, E. StepPO: Step-Aligned Policy Optimization for Agentic Reinforcement Learning. CoRR 2026, abs/2604.18401, [2604.18401. [Google Scholar] [CrossRef]
- Cui, G.; Zhang, Y.; Chen, J.; Yuan, L.; Wang, Z.; Zuo, Y.; Li, H.; Fan, Y.; Chen, H.; Chen, W.; et al. The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models. CoRR 2025, abs/2505.22617, [2505.22617. [Google Scholar] [CrossRef]
- Li, Q.; Xue, R.; Wang, J.; Zhou, M.; Li, Z.; Ji, X.; Wang, Y.; Liu, M.; Yang, Z.; Qiu, M.; et al. CURE: Critical-Token-Guided Re-Concatenation for Entropy-Collapse Prevention. CoRR 2025, abs/2508.11016, [2508.11016. [Google Scholar] [CrossRef]
- Xu, W.; Zhao, W.; Wang, Z.; Li, Y.; Jin, C.; Jin, M.; Mei, K.; Wang, K.; Metaxas, D.N. EPO: Entropy-regularized Policy Optimization for LLM Agents Reinforcement Learning. CoRR 2025, abs/2509.22576, 2509.22576. [Google Scholar] [CrossRef]
- Song, H.; Jiang, J.; Min, Y.; Chen, J.; Chen, Z.; Zhao, W.X.; Fang, L.; Wen, J. R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning. CoRR 2025, abs/2503.05592, 2503.05592. [Google Scholar] [CrossRef]
- Hao, S.; Gu, Y.; Ma, H.; Hong, J.J.; Wang, Z.; Wang, D.Z.; Hu, Z. Reasoning with Language Model is Planning with World Model. In Proceedings of the Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing; EMNLP 2023, Singapore, Bouamor, H., Pino, J., Bali, K., Eds.; Association for Computational Linguistics, 6-10 December 2023; pp. 8154–8173. [Google Scholar] [CrossRef]
- Li, Y.; Gu, Q.; Wen, Z.; Li, Z.; Xing, T.; Guo, S.; Zheng, T.; Zhou, X.; Qu, X.; Zhou, W.; et al. TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling. CoRR 2025, abs/2508.17445, [2508.17445. [Google Scholar] [CrossRef]
- Li, W.; Qu, B.; Pan, B.; Zhang, J.; Liu, Z.; Zhang, P.; Chen, W.; Zhang, B. LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent. CoRR 2026, abs/2604.17931, 2604.17931. [Google Scholar] [CrossRef]
- Xia, P.; Chen, J.; Wang, H.; Liu, J.; Zeng, K.; Wang, Y.; Han, S.; Zhou, Y.; Zhao, X.; Chen, H.; et al. SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning. CoRR 2026, abs/2602.08234, 2602.08234. [Google Scholar] [CrossRef]
- Chang, Q.; Zhang, Z.; Hu, P.; Ma, J.; Pan, Y.; Zhang, J.; Du, J.; Liu, Q.; Gao, J. THOR: Tool-Integrated Hierarchical Optimization via RL for Mathematical Reasoning. CoRR 2025, abs/2509.13761, 2509.13761. [Google Scholar] [CrossRef]
- Zhao, Y.; Hu, L.; Wang, Y.; Hou, M.; Zhang, H.; Ding, K.; Zhao, J. Stronger Together: On-Policy Reinforcement Learning for Collaborative LLMs. CoRR 2025, abs/2510.11062, [2510.11062. [Google Scholar] [CrossRef]
- Mo, Z.; Li, X.; Chen, Y.; Bing, L. Multi-Agent Tool-Integrated Policy Optimization. CoRR 2025, abs/2510.04678, 2510.04678. [Google Scholar] [CrossRef]
- Liu, B.; Guertler, L.; Yu, S.; Liu, Z.; Qi, P.; Balcells, D.; Liu, M.; Tan, C.; Shi, W.; Lin, M.; et al. SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning. CoRR 2025, abs/2506.24119, 2506.24119. [Google Scholar] [CrossRef]
- Li, X.; Zou, H.; Liu, P. ToRL: Scaling Tool-Integrated RL. CoRR 2025, abs/2503.23383, 2503.23383. [Google Scholar] [CrossRef]
- Wei, Z.; Yao, W.; Liu, Y.; Zhang, W.; Lu, Q.; Qiu, L.; Yu, C.; Xu, P.; Zhang, C.; Yin, B.; et al. WebAgent-R1: Training Web Agents via End-to-End Multi-Turn Reinforcement Learning. In Proceedings of the Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025; Suzhou, China, Christodoulopoulos, C., Chakraborty, T., Rose, C., Peng, V., Eds.; Association for Computational Linguistics, 4-9 November 2025; pp. 7909–7928. [Google Scholar] [CrossRef]
- Chen, K.; Cusumano-Towner, M.F.; Huval, B.; Petrenko, A.; Hamburger, J.; Koltun, V.; Krähenbühl, P. Reinforcement Learning for Long-Horizon Interactive LLM Agents. CoRR 2025, abs/2502.01600, 2502.01600. [Google Scholar] [CrossRef]
- Fang, Z.; Sun, R. AdaTIR: Adaptive Tool-Integrated Reasoning via Difficulty-Aware Policy Optimization. CoRR 2026, abs/2601.14696, 2601.14696. [Google Scholar] [CrossRef]
- Song, S.; Ma, C.; Cheng, Z.; Lei, S.; Li, M.; Zeng, Y.; Tou, H.; Jia, K. EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance. CoRR 2025, abs/2509.23730, [2509.23730. [Google Scholar] [CrossRef]
- Chen, Y.; Dong, G.; Dou, Z. ET-Agent: Incentivizing Effective Tool-Integrated Reasoning Agent via Behavior Calibration. In Proceedings of the Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2026; San Diego, California, United States, Liakata, M., Moreira, V.P., Zhang, J., Jurgens, D., Eds.; Association for Computational Linguistics, 2-7 July 2026; Volume 2026, pp. 7337–7359. [Google Scholar]
- Wen, T.; Dong, G.; Dou, Z. SmartSearch: Process Reward-Guided Query Refinement for Search Agents. CoRR 2026, abs/2601.04888, 2601.04888. [Google Scholar] [CrossRef]
- Li, J.; Wang, Y.; Yan, Q.; Tian, Y.; Xu, Z.; Song, H.; Xu, P.; Cheong, L.L. SALT: Step-level Advantage Assignment for Long-horizon Agents via Trajectory Graph. In Proceedings of the Findings of the Association for Computational Linguistics: EACL 2026 Association for Computational Linguistics; Rabat, Morocco, Demberg, V., Inui, K., Marquez, L., Eds.; Findings of ACL, 24-29 March 2026; pp. 4709–4725. [Google Scholar] [CrossRef]
- Liu, W.; Jin, J.; Huang, Z.; Wen, T.; Dong, G.; Zhao, Z.; Zhu, Y.; Dou, Z.; Wen, J.R. The Rules of the Game: A Survey of Rubrics for Large Language Models. 2026. [Google Scholar]
- Sharma, M.; Zhang, C.B.C.; Bandi, C.; Wang, C.; Aich, A.; Nghiem, H.; Rabbani, T.; Htet, Y.; Jang, B.; Basu, S.; et al. ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research Agents. CoRR 2025, abs/2511.07685, 2511.07685. [Google Scholar] [CrossRef]
- Huang, Z.; Zhuang, Y.; Lu, G.; Qin, Z.; Xu, H.; Zhao, T.; Peng, R.; Hu, J.; Shen, Z.; Hu, X.; et al. Reinforcement Learning with Rubric Anchors. CoRR 2025, abs/2508.12790, [2508.12790. [Google Scholar] [CrossRef]
- Zhang, J.; Wang, Z.; Gui, L.; Sathyendra, S.M.; Jeong, J.; Veitch, V.; Wang, W.; He, Y.; Liu, B.; Jin, L. Chasing the Tail: Effective Rubric-based Reward Modeling for Large Language Model Post-Training. CoRR 2025, abs/2509.21500, [2509.21500. [Google Scholar] [CrossRef]
- Raghavendra, M.; Gunjal, A.; Liu, B.; He, Y. Agentic Rubrics as Contextual Verifiers for SWE Agents. In Proceedings of the Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2026; San Diego, California, United States, Liakata, M., Moreira, V.P., Zhang, J., Jurgens, D., Eds.; Association for Computational Linguistics, 2-7 July 2026; Volume 2026, pp. 15265–15290. [Google Scholar]
- Gangwani, I.; Bansal, A. Rubric as Reward: Decomposing Verification Signals for Logical Reasoning in GRPO. In Proceedings of the ICLR 2026 Workshop on Logical Reasoning of Large Language Models, 2026. [Google Scholar]
- Kim, S.; Shin, J.; Choi, Y.; Jang, J.; Longpre, S.; Lee, H.; Yun, S.; Shin, S.; Kim, S.; Thorne, J.; et al. Prometheus: Inducing Fine-Grained Evaluation Capability in Language Models. In Proceedings of the The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024; 2024. Available online: https://openreview.net/.
- Anugraha, D.; Tang, Z.; Miranda, L.J.V.; Zhao, H.; Farhansyah, M.R.; Kuwanto, G.; Wijaya, D.; Winata, G.I. R3: Robust Rubric-Agnostic Reward Models. CoRR 2025, abs/2505.13388, 2505.13388. [Google Scholar] [CrossRef]
- Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; Klimov, O. Proximal Policy Optimization Algorithms. CoRR 2017, abs/1707.06347, [1707.06347. [Google Scholar]
- Schulman, J.; Moritz, P.; Levine, S.; Jordan, M.I.; Abbeel, P. High-Dimensional Continuous Control Using Generalized Advantage Estimation. In Proceedings of the 4th International Conference on Learning Representations, ICLR 2016 Conference Track Proceedings; San Juan, Puerto Rico, Bengio, Y., LeCun, Y., Eds.; 2-4 May 2016. [Google Scholar]
- MiniMax. MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention. CoRR 2025, abs/2506.13585, 2506.13585. [Google Scholar] [CrossRef]
- Su, Z.; Pan, L.; Lv, M.; Li, Y.; Hu, W.; Zhang, F.; Gai, K.; Zhou, G. CE-GPPO: Coordinating Entropy via Gradient-Preserving Clipping Policy Optimization in Reinforcement Learning. In Proceedings of the Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics; San Diego, California, United States, Liakata, M., Moreira, V.P., Zhang, J., Jurgens, D., Eds.; Association for Computational Linguistics, 2-7 July 2026; Volume 1, pp. 37747–37760. [Google Scholar]
- Xiao, C.; Zhang, M.; Cao, Y. BNPO: Beta Normalization Policy Optimization. CoRR 2025, abs/2506.02864, 2506.02864. [Google Scholar] [CrossRef]
- Chu, X.; Huang, H.; Zhang, X.; Wei, F.; Wang, Y. GPG: A Simple and Strong Reinforcement Learning Baseline for Model Reasoning. CoRR 2025, abs/2504.02546, 2504.02546. [Google Scholar] [CrossRef]
- Hao, Y.; Dong, L.; Wu, X.; Huang, S.; Chi, Z.; Wei, F. On-Policy RL with Optimal Reward Baseline. CoRR 2025, abs/2505.23585, 2505.23585. [Google Scholar] [CrossRef]
- Liu, Z.; Meng, F.; Du, L.; Zhou, Z.; Yu, C.; Shao, W.; Zhang, Q. CPGD: Toward Stable Rule-based Reinforcement Learning for Language Models. CoRR 2025, abs/2505.12504, [2505.12504. [Google Scholar] [CrossRef]
- Zou, T.; Zhang, X.; Yu, H.; Wang, M.; Huang, F.; Li, Y. EIFBENCH: Extremely Complex Instruction Following Benchmark for Large Language Models. In Proceedings of the Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025; Suzhou, China, Christodoulopoulos, C., Chakraborty, T., Rose, C., Peng, V., Eds.; Association for Computational Linguistics, 4-9 November 2025; Volume 2025, pp. 20930–20953. [Google Scholar] [CrossRef]
- He, S.; Feng, L.; Wei, Q.; Cheng, X.; Feng, L.; An, B. Hierarchy-of-Groups Policy Optimization for Long-Horizon Agentic Tasks. CoRR 2026, abs/2602.22817, 2602.22817. [Google Scholar] [CrossRef]
- Li, Y.; Zhou, Q.; Duan, W.; Chen, L. When Denser Credit Is Not Enough: Evidence-Calibrated Policy Optimization for Long-Horizon LLM Agent Training. CoRR 2026, abs/2606.05885, [2606.05885. [Google Scholar] [CrossRef]
- Wu, F.; Zhu, W.; Zhang, Y.; Chatterjee, S.; Zhu, J.; Mo, F.; Luo, R.; Gao, J. PORTool: Tool-Use LLM Training with Rewarded Tree. CoRR 2025, abs/2510.26020, 2510.26020. [Google Scholar] [CrossRef]
- Peng, J.; Liu, Y.; Zhou, R.; Fleming, C.; Wang, Z.; García, A.; Hong, M. HiPER: Hierarchical Reinforcement Learning with Explicit Credit Assignment for Large Language Model Agents. CoRR 2026, abs/2602.16165, [2602.16165. [Google Scholar] [CrossRef]
- Tan, H.; Yang, X.; Chen, H.; Shao, J.; Wen, Y.; Shen, Y.; Luo, W.; Du, X.; Guo, L.; Li, Y. Hindsight Credit Assignment for Long-Horizon LLM Agents. CoRR 2026, abs/2603.08754, 2603.08754. [Google Scholar] [CrossRef]
- Xu, H.; Zhao, S.; Wu, X.; Luu, A.T. Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization. In Proceedings of the Findings of the Association for Computational Linguistics, ACL 2026; San Diego, California, United States, Liakata, M., Moreira, V.P., Zhang, J., Jurgens, D., Eds.; Association for Computational Linguistics, 2-7 July 2026; Volume 2026, pp. 17759–17771. [Google Scholar]
- Shen, H. On Entropy Control in LLM-RL Algorithms. CoRR 2025, abs/2509.03493, 2509.03493. [Google Scholar] [CrossRef]
- Zhang, X.; Yuan, X.; Huang, D.; You, W.; Hu, C.; Ruan, J.; Chen, K.; Hu, X. Revisiting Entropy Regularization: Adaptive Coefficient Unlocks Its Potential for LLM Reinforcement Learning. In Proceedings of the Findings of the Association for Computational Linguistics, ACL 2026; San Diego, California, United States, Liakata, M., Moreira, V.P., Zhang, J., Jurgens, D., Eds.; Association for Computational Linguistics, 2-7 July 2026; pp. 18005–18020. [Google Scholar]
- Dong, G.; Bao, L.; Wang, Z.; Zhao, K.; Li, X.; Jin, J.; Yang, J.; Mao, H.; Zhang, F.; Gai, K.; et al. Agentic Entropy-Balanced Policy Optimization. CoRR 2025, abs/2510.14545, [2510.14545. [Google Scholar] [CrossRef]
- Zhang, K.; Zuo, Y.; He, B.; Sun, Y.; Liu, R.; Jiang, C.; Fan, Y.; Tian, K.; Jia, G.; Li, P.; et al. A Survey of Reinforcement Learning for Large Reasoning Models. CoRR 2025, abs/2509.08827, 2509.08827. [Google Scholar] [CrossRef]
- Wang, W.; Xiong, S.; Chen, G.; Gao, W.; Guo, S.; He, Y.; Huang, J.; Liu, J.; Li, Z.; Li, X.; et al. Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library. CoRR 2025, abs/2506.06122, 2506.06122. [Google Scholar] [CrossRef]
- Qiao, Z.; Chen, G.; Chen, X.; Yu, D.; Yin, W.; Wang, X.; Zhang, Z.; Li, B.; Yin, H.; Li, K.; et al. WebResearcher: Unleashing unbounded reasoning capability in Long-Horizon Agents. CoRR 2025, abs/2509.13309, 2509.13309. [Google Scholar] [CrossRef]
- Zong, Z.; Chen, D.; Li, Y.; Yi, Q.; Zhou, B.; Li, C.; Qian, B.; Chen, P.; Jiang, J. AT2PO: Agentic Turn-based Policy Optimization via Tree Search. arXiv 2026, arXiv:cs. [Google Scholar]
- Zhao, Y.; Huang, W.; Wang, S.; Zhao, R.; Chen, C.; Shu, Y.; Qin, C. Training Multi-Turn Search Agent via Contrastive Dynamic Branch Sampling. CoRR 2026, abs/2602.03719, 2602.03719. [Google Scholar] [CrossRef]
- Chen, Y.; Dong, G.; Dou, Z. Toward Effective Tool-Integrated Reasoning via Self-Evolved Preference Learning. CoRR 2025, abs/2509.23285, 2509.23285. [Google Scholar] [CrossRef]
- Zhen, S.; Yu, Y.; Guo, R.; Cheng, N.; Deng, Y. Hierarchical Reinforcement Learning with Augmented Step-Level Transitions for LLM Agents. In Proceedings of the Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2026; San Diego, California, United States, Liakata, M., Moreira, V.P., Zhang, J., Jurgens, D., Eds.; Association for Computational Linguistics, 2-7 July 2026; Volume 2026, pp. 7035–7053. [Google Scholar]
- Liu, S.; Liang, Z.; Lyu, X.; Amato, C. LLM Collaboration with Multi-Agent Reinforcement Learning. In Proceedings of the Fortieth AAAI Conference on Artificial Intelligence, Thirty-Eighth Conference on Innovative Applications of Artificial Intelligence, Sixteenth Symposium on Educational Advances in Artificial Intelligence; AAAI 2026, Singapore, Koenig, S., Jenkins, C., Taylor, M.E., Eds.; AAAI Press, 20-27 January 2026; pp. 32150–32158. [Google Scholar] [CrossRef]
- Hong, H.; Yin, J.; Wang, Y.; Liu, J.; Chen, Z.; Yu, A.; Li, J.; Ye, Z.; Xiao, H.; Chen, Y.; et al. Multi-Agent Deep Research: Training Multi-Agent Systems with M-GRPO. CoRR 2025, abs/2511.13288, [2511.13288. [Google Scholar] [CrossRef]
- Zhang, Z.; Li, X.; Lin, Y.; Liu, H.; Chandradevan, R.; Wu, L.; Lin, M.; Wang, F.; Tang, X.; He, Q.; et al. Unlocking the Power of Multi-Agent LLM for Reasoning: From Lazy Agents to Deliberation. CoRR 2025, abs/2511.02303, 2511.02303. [Google Scholar] [CrossRef]
- Lowe, R.; Wu, Y.; Tamar, A.; Harb, J.; Abbeel, P.; Mordatch, I. Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments. In Proceedings of the Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017; Long Beach, CA, USA, Guyon, I., von Luxburg, U., Bengio, S., Wallach, H.M., Fergus, R., Vishwanathan, S.V.N., Garnett, R., Eds.; 4-9 December 2017; pp. 6379–6390. [Google Scholar]
- Rashid, T.; Samvelyan, M.; de Witt, C.S.; Farquhar, G.; Foerster, J.N.; Whiteson, S. QMIX: Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement Learning. In Proceedings of the Proceedings of the 35th International Conference on Machine Learning, ICML 2018 PMLR; Stockholmsmässan, Stockholm, Sweden, Dy, J.G., Krause, A., Eds.; Proceedings of Machine Learning Research; 10-15 July 2018; Vol. 80, pp. 4292–4301. [Google Scholar]
- Yu, C.; Velu, A.; Vinitsky, E.; Gao, J.; Wang, Y.; Bayen, A.M.; Wu, Y. The Surprising Effectiveness of PPO in Cooperative Multi-Agent Games. In Proceedings of the Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022; New Orleans, LA, USA, Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., Oh, A., Eds.; 28 November 2022. [Google Scholar]
- Chen, G.; Qiao, Z.; Wang, W.; Yu, D.; Chen, X.; Sun, H.; Liao, M.; Fan, K.; Jiang, Y.; Xie, P.; et al. MARS: Optimizing Dual-System Deep Research via Multi-Agent Reinforcement Learning. CoRR 2025, abs/2510.04935, 2510.04935. [Google Scholar] [CrossRef]
- Han, A.; Hu, J.; Wei, P.; Zhang, Z.; Guo, Y.; Lu, J.; Chen, Z.; Zhang, Z. JoyAgents-R1: Accelerating Multi-Agent Evolution Dynamics with Variance-Reduction Group Relative Policy Optimization. 2026. [Google Scholar]
- Liang, T.; He, Z.; Jiao, W.; Wang, X.; Wang, Y.; Wang, R.; Yang, Y.; Shi, S.; Tu, Z. Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate. In Proceedings of the Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024; Miami, FL, USA, Al-Onaizan, Y., Bansal, M., Chen, Y., Eds.; Association for Computational Linguistics, 12-16 November 2024; pp. 17889–17904. [Google Scholar] [CrossRef]
- Zhang, M.; Luo, H.; Shen, T.; Lin, Q.; Tang, X.; Mao, R.; Cambria, E. FlowSteer: Interactive Agentic Workflow Orchestration via End-to-End Reinforcement Learning. CoRR 2026, abs/2602.01664, 2602.01664. [Google Scholar] [CrossRef]
- Liu, J.; Li, Y.; Zhang, C.; Li, J.; Chen, A.; Ji, K.; Cheng, W.; Wu, Z.; Du, C.; Xu, Q.; et al. WebExplorer: Explore and Evolve for Training Long-Horizon Web Agents. CoRR 2025, abs/2509.06501, 2509.06501. [Google Scholar] [CrossRef]
- Torabi, F.; Warnell, G.; Stone, P. Behavioral Cloning from Observation. In Proceedings of the Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence, IJCAI 2018, July 13-19, 2018; Stockholm, Sweden, Lang, J., Ed.; Ed. ijcai.org, 2018; pp. 4950–4957. [Google Scholar] [CrossRef]
- Christiano, P.F.; Leike, J.; Brown, T.B.; Martic, M.; Legg, S.; Amodei, D. Deep Reinforcement Learning from Human Preferences. In Proceedings of the Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017; Long Beach, CA, USA, Guyon, I., von Luxburg, U., Bengio, S., Wallach, H.M., Fergus, R., Vishwanathan, S.V.N., Garnett, R., Eds.; 4-9 December 2017; pp. 4299–4307. [Google Scholar]
- Wu, Y.; Cai, Z.; Ning, L.; Wang, H.; Chen, Z.; Tang, Y.; Chen, H. LiteGUI: Distilling Compact GUI Agents with Reinforcement Learning. CoRR 2026, abs/2605.07505, [2605.07505. [Google Scholar] [CrossRef]
- Li, F.; Zhang, H.; Huang, H.; Wang, J.; Hao, J.; Yuan, K.; Li, M.; Zhang, M.; Xu, P.; Zhuang, W.; et al. KAT-Coder-V2 Technical Report. CoRR 2026, abs/2603.27703, 2603.27703. [Google Scholar] [CrossRef]
- Zhong, Q.; Zheng, M.; Song, M.; Lin, X.; Sun, J.; Jiang, H.; Wang, X.; Fang, J. SOD: Step-wise On-policy Distillation for Small Language Model Agents. CoRR 2026, abs/2605.07725, [2605.07725. [Google Scholar] [CrossRef]
- Li, G.; Zheng, M.; Song, M.; Liu, R.; Yang, T.; Sun, J.; Zhong, Q.; Guo, H.; Fang, J.; Zhang, D.; et al. On-Policy Distillation with Curriculum Turn-level Guidance for Multi-turn Agents. CoRR 2026, abs/2606.15912, [2606.15912. [Google Scholar] [CrossRef]
- Zhang, Y.; Lin, X.; Wu, C. StepOPSD: Step-Aware Online Preference Distillation for Agent Reinforcement Learning. CoRR 2026, abs/2605.27140, [2605.27140. [Google Scholar] [CrossRef]
- Zhang, G.; Xu, X.; Yue, Y.; Su, Z.; Zhou, W.; You, X.; Yan, S. OPD-Evolver: Cultivating Holistic Agent Evolver via On-Policy Distillation. CoRR 2026, abs/2606.17628, [2606.17628. [Google Scholar] [CrossRef]
- Hübotter, J.; Lübeck, F.; Behric, L.; Baumann, A.; Bagatella, M.; Marta, D.; Hakimi, I.; Shenfeld, I.; Buening, T.K.; Guestrin, C.; et al. Reinforcement Learning via Self-Distillation. CoRR 2026, abs/2601.20802, 2601.20802. [Google Scholar] [CrossRef]
- Wang, H.; Wang, G.; Xiao, H.; Zhou, Y.; Pan, Y.; Wang, J.; Xu, K.; Wen, Y.; Ruan, X.; Chen, X.; et al. Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents. CoRR 2026, abs/2604.10674, [2604.10674. [Google Scholar] [CrossRef]
- Chen, X.; Zhang, S.; Wu, J. f-OPD: Stabilizing Long-Horizon On-Policy Distillation with Freshness-Aware Control. CoRR 2026, abs/2605.17862, 2605.17862. [Google Scholar] [CrossRef]
- Xu, Y.; Sang, H.; Zhou, Z.; He, R.; Wang, Z.; Geramifard, A. TIP: Token Importance in On-Policy Distillation. CoRR 2026, abs/2604.14084, 2604.14084. [Google Scholar] [CrossRef]
- Yu, X.; Yang, C.; Yu, C.; Huang, L.; An, Z.; Xu, Y. Online Policy Distillation with Decision-Attention. In Proceedings of the International Joint Conference on Neural Networks, IJCNN 2024, Yokohama, Japan, June 30 - July 5, 2024; IEEE; 2024, pp. 1–8. [Google Scholar] [CrossRef]
- Fu, Y.; Huang, H.; Jiang, K.; Zhu, Y.; Zhao, D. Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes. CoRR 2026, abs/2603.25562, [2603.25562. [Google Scholar] [CrossRef]
- Fang, J.; Peng, Y.; Zhang, X.; Wang, Y.; Yi, X.; Zhang, G.; Xu, Y.; Wu, B.; Liu, S.; Li, Z.; et al. A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems. CoRR 2025, abs/2508.07407, [2508.07407. [Google Scholar] [CrossRef]
- Zweiger, A.; Pari, J.; Guo, H.; Akyürek, E.; Kim, Y.; Agrawal, P. Self-Adapting Language Models. CoRR 2025, abs/2506.10943, 2506.10943. [Google Scholar] [CrossRef]
- Li, C.; Xue, M.; Zhang, Z.; Yang, J.; Zhang, B.; Yu, B.; Hui, B.; Lin, J.; Wang, X.; Liu, D. START: Self-taught Reasoner with Tools. In Proceedings of the Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025; Suzhou, China, Christodoulopoulos, C., Chakraborty, T., Rose, C., Peng, V., Eds.; Association for Computational Linguistics, 4-9 November 2025; pp. 13512–13553. [Google Scholar] [CrossRef]
- Zhao, W.; Yüksekgönül, M.; Wu, S.; Zou, J.Y. SiriuS: Self-improving Multi-agent Systems via Bootstrapped Reasoning. In Proceedings of the Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025; San Diego, CA, USA, December 2-7, 2025 / Mexico City, Mexico, Belgrave, D., Zhang, C., Montoya, L.N., Lin, H., Pascanu, R., Koniusz, P., Ghassemi, M., Chen, N., Ruíz, I.V.M., Loaiza-Bonilla, A., Eds.; 30 November 2025. [Google Scholar]
- Robeyns, M.; Szummer, M.; Aitchison, L. A Self-Improving Coding Agent. CoRR 2025, abs/2504.15228, 2504.15228. [Google Scholar] [CrossRef]
- Pourcel, J.; Colas, C.; Oudeyer, P. Self-Improving Language Models for Evolutionary Program Synthesis: A Case Study on ARC-AGI. In Proceedings of the Forty-second International Conference on Machine Learning, ICML 2025; Vancouver, BC, Canada, Singh, A., Fazel, M., Hsu, D., Lacoste-Julien, S., Berkenkamp, F., Maharaj, T., Wagstaff, K., Zhu, J., Eds.; PMLR / OpenReview.net; Proceedings of Machine Learning Research; 13-19 July 2025; Vol. 267. [Google Scholar]
- Ge, Y.; Romeo, S.; Cai, J.; Sunkara, M.; Zhang, Y. SAMULE: Self-Learning Agents Enhanced by Multi-level Reflection. In Proceedings of the Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025; Suzhou, China, Christodoulopoulos, C., Chakraborty, T., Rose, C., Peng, V., Eds.; Association for Computational Linguistics, 4-9 November 2025; Volume 2025, pp. 16591–16610. [Google Scholar] [CrossRef]
- Zeng, Z.; Ivison, H.; Wang, Y.; Yuan, L.; Li, S.S.; Ye, Z.; Li, S.; He, J.; Zhou, R.; Chen, T.; et al. RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments. CoRR 2025, abs/2511.07317, [2511.07317. [Google Scholar] [CrossRef]
- Xue, X.; Zhou, Y.; Zhang, G.; Zhang, Z.; Li, Y.; Zhang, C.; Yin, Z.; Torr, P.; Ouyang, W.; Bai, L. CoMAS: Co-Evolving Multi-Agent Systems via Interaction Rewards. CoRR 2025, abs/2510.08529, 2510.08529. [Google Scholar] [CrossRef]
- Wang, S.; Jiao, Z.; Zhang, Z.; Peng, Y.; Ze, X.; Yang, B.; Wang, W.; Wei, H.; Zhang, L. Socratic-Zero: Bootstrapping Reasoning via Data-Free Agent Co-evolution. CoRR 2025, abs/2509.24726, [2509.24726. [Google Scholar] [CrossRef]
- Wang, Z.; Shi, T.; He, J.; Cai, M.; Zhang, J.; Song, D. CyberGym: Evaluating AI Agents’ Cybersecurity Capabilities with Real-World Vulnerabilities at Scale. CoRR 2025, abs/2506.02548, 2506.02548. [Google Scholar] [CrossRef]
- Zhang, C.; Li, Y.; Xu, C.; Liu, J.; Liu, A.; Hu, S.; Wu, D.; Huang, G.; Li, K.; Yi, Q.; et al. ArtifactsBench: Bridging the Visual-Interactive Gap in LLM Code Generation Evaluation. CoRR 2025, abs/2507.04952, 2507.04952. [Google Scholar] [CrossRef]
- Raghavendra, M.; Dan, S.; Calvo, M.R.; He, Y.Y.; Mols, J.B.; Anand, G.; McCollum, C.; Arakelyan, E.; Bharadwaj, V.; Park, A.; et al. SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution. CoRR 2026, abs/2605.08366, [2605.08366. [Google Scholar] [CrossRef]
- Yang, J.; Lieret, K.; Ma, J.; Thakkar, P.; Pedchenko, D.; Sootla, S.; McMilin, E.; Yin, P.; Hou, R.; Synnaeve, G.; et al. ProgramBench: Can Language Models Rebuild Programs From Scratch? CoRR 2026. abs/2605.03546 2605.03546. [CrossRef]
- Ding, J.; Long, S.; Pu, C.; Zhou, H.; Gao, H.; Gao, X.; He, C.; Hou, Y.; Hu, F.; Li, Z.; et al. NL2Repo-Bench: Towards Long-Horizon Repository Generation Evaluation of Coding Agents. CoRR 2025, abs/2512.12730, 2512.12730. [Google Scholar] [CrossRef]
- Deng, X.; Da, J.; Pan, E.; He, Y.Y.; Ide, C.; Garg, K.; Lauffer, N.; Park, A.; Pasari, N.; Rane, C.; et al. SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? CoRR 2025, abs/2509.16941, 2509.16941. [Google Scholar] [CrossRef]
- Li, J.; Li, G.; Zhao, Y.; Li, Y.; Liu, H.; Zhu, H.; Wang, L.; Liu, K.; Fang, Z.; Wang, L.; et al. DevEval: A Manually-Annotated Code Generation Benchmark Aligned with Real-World Code Repositories. In Proceedings of the Findings of the Association for Computational Linguistics, ACL 2024, Bangkok, Thailand and virtual meeting, August 11-16, 2024 Association for Computational Linguistics; Ku, L., Martins, A., Srikumar, V., Eds.; Findings of ACL; 2024; Vol. ACL 2024, pp. 3603–3614. [Google Scholar] [CrossRef]
- Duston, T.; Xin, S.; Sun, Y.; Zan, D.; Li, A.; Xin, S.; Shen, K.; Chen, Y.; Sun, Q.; Zhang, G.; et al. AInsteinBench: Benchmarking Coding Agents on Scientific Repositories. CoRR 2025, abs/2512.21373, [2512.21373. [Google Scholar] [CrossRef]
- Miserendino, S.; Wang, M.; Patwardhan, T.; Heidecke, J. SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering? In Proceedings of the Forty-second International Conference on Machine Learning, ICML 2025; Vancouver, BC, Canada, Singh, A., Fazel, M., Hsu, D., Lacoste-Julien, S., Berkenkamp, F., Maharaj, T., Wagstaff, K., Zhu, J., Eds.; PMLR / OpenReview.net; Proceedings of Machine Learning Research; 13-19 July 2025; Vol. 267. [Google Scholar]
- Liu, S.; Jiang, B.; Yang, J.; Li, Y.; Guo, J.; Liu, X.; Dai, B. Context as a Tool: Context Management for Long-Horizon SWE-Agents. In Proceedings of the Findings of the Association for Computational Linguistics, ACL 2026; San Diego, California, United States, Liakata, M., Moreira, V.P., Zhang, J., Jurgens, D., Eds.; Association for Computational Linguistics, 2-7 July 2026; Volume 2026, pp. 20604–20617. [Google Scholar]
- Stripe. Minions: Stripe’s One-Shot, End-to-End Coding Agents; 2026. [Google Scholar]
- Krishna, S.; Krishna, K.; Mohananey, A.; Schwarcz, S.; Stambler, A.; Upadhyay, S.; Faruqui, M. Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented Generation. In Proceedings of the Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL 2025 -; Albuquerque, New Mexico, USA, Chiruzzo, L., Ritter, A., Wang, L., Eds.; Association for Computational Linguistics, 2025; Volume 1, pp. 4745–4759. [Google Scholar] [CrossRef]
- Huang, Y.; Chen, Y.; Zhang, H.; Li, K.; Fang, M.; Yang, L.; Li, X.; Shang, L.; Xu, S.; Hao, J.; et al. Deep Research Agents: A Systematic Examination And Roadmap. CoRR 2025, abs/2506.18096, [2506.18096. [Google Scholar] [CrossRef]
- Zhou, P.; Leon, B.; Ying, X.; Zhang, C.; Shao, Y.; Ye, Q.; Chong, D.; Jin, Z.; Xie, C.; Cao, M.; et al. BrowseComp-ZH: Benchmarking Web Browsing Ability of Large Language Models in Chinese. CoRR 2025, abs/2504.19314, 2504.19314. [Google Scholar] [CrossRef]
- Hu, L.; Jiao, J.; Liu, J.; Ren, Y.; Wen, Z.; Zhang, K.; Zhang, X.; Gao, X.; He, T.; Hu, F.; et al. FinSearchComp: Towards a Realistic, Expert-Level Evaluation of Financial Search and Reasoning. CoRR 2025, abs/2509.13160, [2509.13160. [Google Scholar] [CrossRef]
- Lin, M.; Wu, Z.; Xu, Z.; Liu, H.; Tang, X.; He, Q.; Aggarwal, C.C.; Liu, H.; Zhang, X.; Wang, S. A Comprehensive Survey on Reinforcement Learning-based Agentic Search: Foundations, Roles, Optimizations, Evaluations, and Applications. CoRR 2025, abs/2510.16724, [2510.16724. [Google Scholar] [CrossRef]
- Tang, Q.; Xiang, H.; Yu, L.; Yu, B.; Lu, Y.; Han, X.; Sun, L.; Zhang, W.; Wang, P.; Liu, S.; et al. Beyond Turn Limits: Training Deep Search Agents with Dynamic Context Window. CoRR 2025, abs/2510.08276, [2510.08276. [Google Scholar] [CrossRef]
- Deng, Y.; Wang, G.; Ying, Z.; Wu, X.; Lin, J.; Xiong, W.; Dai, Y.; Yang, S.; Zhang, Z.; Wang, Q.; et al. Atom-Searcher: Enhancing Agentic Deep Research via Fine-Grained Atomic Thought Reward. CoRR 2025, abs/2508.12800, [2508.12800. [Google Scholar] [CrossRef]
- Shao, J.; Miao, Y.; Zhang, W.; Luo, B. FoldAct: Efficient and Stable Context Folding for Long-Horizon Search Agents. CoRR 2025, abs/2512.22733, [2512.22733. [Google Scholar] [CrossRef]
- Dai, Y.; Wang, G.; Wang, Y.; Dou, K.; Zhou, K.; Zhang, Z.; Yang, S.; Tang, F.; Yin, J.; Zeng, P.; et al. EviNote-RAG: Enhancing RAG Models via Answer-Supportive Evidence Notes. CoRR 2025, abs/2509.00877, 2509.00877. [Google Scholar] [CrossRef]
- Ying, S.; Wang, Z.; Peng, Y.; Chen, J.; Wu, Y.; Lin, H.; He, D.; Liu, S.; Yu, G.; Piao, Y.; et al. Retrieval-Infused Reasoning Sandbox: A Benchmark for Decoupling Retrieval and Reasoning Capabilities. CoRR 2026, abs/2601.21937, [2601.21937. [Google Scholar] [CrossRef]
- Pham, T.; Nguyen, N.; Zunjare, P.; Chen, W.; Tseng, Y.; Vu, T. SealQA: Raising the Bar for Reasoning in Search-Augmented Language Models. CoRR 2025, abs/2506.01062, 2506.01062. [Google Scholar] [CrossRef]
- Yang, C.; Wu, X.; Lin, X.; Xu, C.; Jiang, X.; Sun, Y.; Li, J.; Xiong, H.; Guo, J. GraphSearch: An Agentic Deep Searching Workflow for Graph Retrieval-Augmented Generation. CoRR 2025, abs/2509.22009, 2509.22009. [Google Scholar] [CrossRef]
- LangChain. Open Deep Research. 2025. [Google Scholar]
- Wang, S.; Xia, Q.; Wang, V.; Herberttli; Bobsimons; Dou, Z. Laser: Governing Long-Horizon Agentic Search via Structured Protocol and Context Register. CoRR 2025, abs/2512.20458, 2512.20458. [Google Scholar] [CrossRef]
- Xu, Z.; Xu, Z.; Zhang, R.; Zhu, C.; Yu, S.; Liu, W.; Zhang, Q.; Ding, W.; Yu, C.; Wang, Y. WideSeek-R1: Exploring Width Scaling for Broad Information Seeking via Multi-Agent Reinforcement Learning. CoRR 2026, abs/2602.04634, 2602.04634. [Google Scholar] [CrossRef]
- Huang, Z.; Ren, H.; Yuan, X.; Wang, J.; Jiang, Z.; Xu, K.; He, S.; Zhao, J.; Liu, K. WideSeek: Advancing Wide Research via Multi-Agent Scaling. CoRR 2026, abs/2602.02636, 2602.02636. [Google Scholar] [CrossRef]
- Lee, K.Y.; Huang, Y.; He, Z.; Zhou, H.; Luo, W.; Shao, K.; Fang, M.; Wang, J. InfoSeeker: A Scalable Hierarchical Parallel Agent Framework for Web Information Seeking. CoRR 2026, abs/2604.02971, 2604.02971. [Google Scholar] [CrossRef]
- Song, X.; Zhang, L.; Zhao, K.; Zhu, Y.; Wang, Z.; Dong, G.; Yang, J.; Li, H.; Gai, K.; Wen, J.R.; et al. WebSwarm: Recursive Multi-Agent Orchestration for Deep-and-Wide Web Search. arXiv 2026, arXiv:cs. [Google Scholar]
- Lan, T.; Henry, F.; Zhu, B.; Jia, Q.; Ren, J.; Pu, Q.; Li, H.; Wang, L.; Xu, Z.; Luo, W. Table-as-Search: Agentic Information Seeking is Table Completion. In Proceedings of the Findings of the Association for Computational Linguistics, ACL 2026; San Diego, California, United States, Liakata, M., Moreira, V.P., Zhang, J., Jurgens, D., Eds.; Association for Computational Linguistics, 2-7 July 2026; pp. 15088–15107. [Google Scholar]
- Xie, J.; Chen, Z.; Zhang, R.; Wan, X.; Li, G. Large Multimodal Agents: A Survey. arXiv 2024, arXiv:cs. [Google Scholar]
- Ma, Y.; Zang, Y.; Chen, L.; Chen, M.; Jiao, Y.; Li, X.; Lu, X.; Liu, Z.; Ma, Y.; Dong, X.; et al. MMLONGBENCH-DOC: Benchmarking Long-context Document Understanding with Visualizations. In Proceedings of the Advances in Neural Information Processing Systems 37: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024; Vancouver, BC, Canada, Globersons, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J.M., Zhang, C., Eds.; 10 - 15 December 2024. [Google Scholar]
- Deng, C.; Yuan, J.; Bu, P.; Wang, P.; Li, Z.; Xu, J.; Li, X.; Gao, Y.; Song, J.; Zheng, B.; et al. LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and Locating. In Proceedings of the Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics; Vienna, Austria, Che, W., Nabende, J., Shutova, E., Pilehvar, M.T., Eds.; 27 July 2025; Volume 1, pp. 1135–1159. [Google Scholar] [CrossRef]
- Tang, J.; Liu, Q.; Ye, Y.; Lu, J.; Wei, S.; Wang, A.; Lin, C.; Feng, H.; Zhao, Z.; Wang, Y.; et al. MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering. In Proceedings of the Findings of the Association for Computational Linguistics, ACL 2025 Association for Computational Linguistics; Vienna, Austria, Che, W., Nabende, J., Shutova, E., Pilehvar, M.T., Eds.; Findings of ACL; 27 July 2025; Vol. ACL 2025, pp. 7748–7763. [Google Scholar] [CrossRef]
- Ouyang, L.; Qu, Y.; Zhou, H.; Zhu, J.; Zhang, R.; Lin, Q.; Wang, B.; Zhao, Z.; Jiang, M.; Zhao, X.; et al. OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2025, Nashville, TN, USA, June 11-15, 2025; Computer Vision Foundation / IEEE; 2025, pp. 24838–24848. [Google Scholar] [CrossRef]
- Landeghem, J.V.; Powalski, R.; Tito, R.; Jurkiewicz, D.; Blaschko, M.B.; Borchmann, L.; Coustaty, M.; Moens, S.; Pietruszka, M.; Anckaert, B.; et al. Document Understanding Dataset and Evaluation (DUDE). In Proceedings of the IEEE/CVF International Conference on Computer Vision, ICCV 2023, Paris, France, October 1-6, 2023; IEEE; 2023, pp. 19471–19483. [Google Scholar] [CrossRef]
- Fu, L.; Kuang, Z.; Song, J.; Huang, M.; Yang, B.; Li, Y.; Zhu, L.; Luo, Q.; Wang, X.; Lu, H.; et al. OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning. In Proceedings of the Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025; San Diego, CA, USA, December 2-7, 2025 / Mexico City, Mexico, Belgrave, D., Zhang, C., Montoya, L.N., Lin, H., Pascanu, R., Koniusz, P., Ghassemi, M., Chen, N., Ruíz, I.V.M., Loaiza-Bonilla, A., Eds.; 30 November 2025. [Google Scholar]
- Geng, X.; Xia, P.; Zhang, Z.; Wang, X.; Wang, Q.; Ding, R.; Wang, C.; Wu, J.; Zhao, Y.; Li, K.; et al. WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent. CoRR 2025, abs/2508.05748, [2508.05748. [Google Scholar] [CrossRef]
- Chen, S.; Feng, K.; Chen, H.; Huang, W.; Dai, D.; Shou, Q.; Lin, Y.; Yue, X.; Gao, S.; Pang, T. OpenSearch-VL: An Open Recipe for Frontier Multimodal Search Agents. CoRR 2026, abs/2605.05185, 2605.05185. [Google Scholar] [CrossRef]
- Du, Y.; Liu, Z.; Peng, J.; Wu, J.; Li, J.; Li, J.; Zhao, W.X.; Wen, J. Towards Long-horizon Agentic Multimodal Search. CoRR 2026, abs/2604.12890, [2604.12890. [Google Scholar] [CrossRef]
- Zhang, R.; Sun, Q.; Song, C.; Qi, Y.; Zheng, Z. VSearcher: Long-Horizon Multimodal Search Agent via Reinforcement Learning. CoRR 2026, abs/2603.02795, 2603.02795. [Google Scholar] [CrossRef]
- Li, Y.; Li, Y.; Wang, X.; Jiang, Y.; Zhang, Z.; Zheng, X.; Wang, H.; Zheng, H.; Huang, F.; Zhou, J.; et al. Benchmarking Multimodal Retrieval Augmented Generation with Dynamic VQA Dataset and Self-adaptive Planning Agent. In Proceedings of the The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025; 2025. Available online: https://openreview.net/.
- Zhang, C.; Dong, G.; Yang, X.; Dou, Z. Towards Mixed-Modal Retrieval for Universal Retrieval-Augmented Generation. CoRR 2025, abs/2510.17354, 2510.17354. [Google Scholar] [CrossRef]
- Liu, C.; Yu, X.; Chang, Z.; Huang, Z.; Zhang, S.; Lian, H.; Wang, K.; Xu, R.; Hu, S.; Hou, J.; et al. Watching, Reasoning, and Searching: A Video Deep Research Benchmark on Open Web for Agentic Video Reasoning. CoRR 2026, abs/2601.06943, [2601.06943. [Google Scholar] [CrossRef]
- Liang, Z.; Shu, Y.; Liu, X.; Qin, M.; Liang, K.; Sebe, N.; Liu, Z.; Liao, L. Video-Browser: Towards Agentic Open-web Video Browsing, 2026. arXiv arXiv:cs.
- Fan, S.; Guo, M.; Yang, S. Agentic Keyframe Search for Video Question Answering. CoRR 2025, abs/2503.16032, 2503.16032. [Google Scholar] [CrossRef]
- Du, M.; Xu, B.; Zhu, C.; Wang, X.; Mao, Z. DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents. CoRR 2025, abs/2506.11763, 2506.11763. [Google Scholar] [CrossRef]
- Zhang, W.; Li, X.; Zhang, Y.; Jia, P.; Wang, Y.; Guo, H.; Liu, Y.; Zhao, X. Deep Research: A Survey of Autonomous Research Agents. CoRR 2025, abs/2508.12752, [2508.12752. [Google Scholar] [CrossRef]
- Zhu, C.; Xu, B.; Du, M.; Wang, S.; Wang, X.; Mao, Z.; Zhang, Y. FS-Researcher: Test-Time Scaling for Long-Horizon Research Tasks with File-System-Based Agents. In Proceedings of the Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics; San Diego, California, United States, Liakata, M., Moreira, V.P., Zhang, J., Jurgens, D., Eds.; Association for Computational Linguistics, 2-7 July 2026; Volume 1, pp. 6353–6373. [Google Scholar]
- Li, Y.; Chen, W.; Yan, Y.; Li, M.; Mei, S.; Wang, X.; Liu, K.; Cong, X.; Wang, S.; Zhang, Z.; et al. AgentCPM-Report: Interleaving Drafting and Deepening for Open-Ended Deep Research. CoRR 2026, abs/2602.06540, [2602.06540. [Google Scholar] [CrossRef]
- Venkit, P.N.; Laban, P.; Zhou, Y.; Huang, K.; Mao, Y.; Wu, C. DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence. CoRR 2025, abs/2509.04499, 2509.04499. [Google Scholar] [CrossRef]
- Wei, J.; Yang, C.; Song, X.; Lu, Y.; Hu, N.; Huang, J.; Tran, D.; Peng, D.; Liu, R.; Huang, D.; et al. Long-form factuality in large language models. In Proceedings of the Advances in Neural Information Processing Systems 37: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024; Vancouver, BC, Canada, Globersons, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J.M., Zhang, C., Eds.; 10 - 15 December 2024. [Google Scholar]
- Min, S.; Krishna, K.; Lyu, X.; Lewis, M.; Yih, W.; Koh, P.W.; Iyyer, M.; Zettlemoyer, L.; Hajishirzi, H. FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation. In Proceedings of the Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing; EMNLP 2023, Singapore, Bouamor, H., Pino, J., Bali, K., Eds.; Association for Computational Linguistics, 6-10 December 2023; pp. 12076–12100. [Google Scholar] [CrossRef]
- Tang, Z.; Zhou, X.; Liu, Y.; Li, L.; Wu, Y.; Wang, W.; Huang, H.; Zhou, W.; Zhou, J.; Song, J.; et al. Workspace-Bench 1.0: Benchmarking AI Agents on Workspace Tasks with Large-Scale File Dependencies. CoRR 2026, abs/2605.03596, 2605.03596. [Google Scholar] [CrossRef]
- Ding, S.; Dai, X.; Xing, L.; Ding, S.; Liu, Z.; JingYi, Y.; Yang, P.; Zhang, Z.; Wei, X.; Fang, X.; et al. WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation. CoRR 2026, abs/2605.10912, [2605.10912. [Google Scholar] [CrossRef]
- Yang, B.; Jin, K.; Wu, Z.; Liu, Z.; Sun, Q.; Li, Z.; Xie, J.; Liu, Z.; Xu, F.; Cheng, K.; et al. OS-Symphony: A Holistic Framework for Robust and Generalist Computer-Using Agents. In Proceedings of the Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics; San Diego, California, United States, Liakata, M., Moreira, V.P., Zhang, J., Jurgens, D., Eds.; Association for Computational Linguistics, 2-7 July 2026; Volume 1, pp. 22300–22330. [Google Scholar]
- Zhang, Y.; Wang, Y.; Zhu, Y.; Du, P.; Miao, J.; Lu, X.; Xu, W.; Hao, Y.; Cai, S.; Wang, X.; et al. ClawBench: Can AI Agents Complete Everyday Online Tasks? CoRR 2026. abs/2604.08523 2604.08523. [CrossRef]
- Yu, T.; Zhang, Z.; Lyu, Z.; Gong, J.; Yi, H.; Wang, X.; Zhou, Y.; Yang, J.; Nie, P.; Huang, Y.; et al. BrowserAgent: Building Web Agents with Human-Inspired Web Browsing Actions. In Trans. Mach. Learn. Res.; 2026. [Google Scholar]
- Hu, X.; Xiong, T.; Yi, B.; Wei, Z.; Xiao, R.; Chen, Y.; Ye, J.; Tao, M.; Zhou, X.; Zhao, Z.; et al. OS Agents: A Survey on MLLM-based Agents for Computer, Phone and Browser Use. In Proceedings of the Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics; Vienna, Austria, Che, W., Nabende, J., Shutova, E., Pilehvar, M.T., Eds.; Association for Computational Linguistics, 27 July 2025; Volume 1, pp. 7436–7465. [Google Scholar] [CrossRef]
- Murty, S.; Zhu, H.; Bahdanau, D.; Manning, C.D. NNetNav: Unsupervised Learning of Browser Agents Through Environment Interaction in the Wild. arXiv 2025, arXiv:cs. [Google Scholar]
- He, K.; Wang, Z.; Zhuang, C.; Gu, J. Recon-Act: A Self-Evolving Multi-Agent Browser-Use System via Web Reconnaissance, Tool Generation, and Task Execution. CoRR 2025, abs/2509.21072, 2509.21072. [Google Scholar] [CrossRef]
- Lu, Y.; Wang, Z.; Huang, J.; Liu, H.; Gesi, J.; Han, Y.; Fu, S.; Zheng, T.; Tang, X.; Luo, C.; et al. WEBSERV: A Full-Stack and RL-Ready Web Environment for Training Web Agents at Scale. arXiv 2026, arXiv:cs. [Google Scholar]
- Wang, J.; Zhou, J.; Zhang, W.; Wang, T.; Liu, W.; Zhang, Z.; Lou, X.; Zhang, W.; Deng, H.; Wang, J. ColorBrowserAgent: Complex Long-Horizon Browser Agent with Adaptive Knowledge Evolution. arXiv 2026, arXiv:cs. [Google Scholar]
- Software, T. Computer Use Benchmark (CUB). 2025. [Google Scholar]
- Seed, B. Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields. CoRR 2026, abs/2606.11042, [2606.11042. [Google Scholar] [CrossRef]
- Li, K.; Meng, Z.; Lin, H.; Luo, Z.; Tian, Y.; Ma, J.; Huang, Z.; Chua, T. ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use. In Proceedings of the Proceedings of the 33rd ACM International Conference on Multimedia, MM 2025; Dublin, Ireland, Gurrin, C., Schoeffmann, K., Zhang, M., Rossetto, L., Rudinac, S., Dang-Nguyen, D., Cheng, W., Chen, P., Benois-Pineau, J., Eds.; ACM, 27-31 October 2025; pp. 8778–8786. [Google Scholar] [CrossRef]
- Gou, B.; Wang, R.; Zheng, B.; Xie, Y.; Chang, C.; Shu, Y.; Sun, H.; Su, Y. Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents. In Proceedings of the The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025; 2025. Available online: https://openreview.net/.
- Wu, Q.; Cheng, K.; Yang, R.; Zhang, C.; Yang, J.; Jiang, H.; Mu, J.; Peng, B.; Qiao, B.; Tan, R.; et al. GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents. CoRR 2025, abs/2506.03143, 2506.03143. [Google Scholar] [CrossRef]
- Rong, B.; Shen, Z.; Wang, Q.; Kang, P.; Xu, Y.; Wei, Y.; Wu, H.; Zhao, Z.; Pei, L.; Jiang, L. AliyunConsoleAgent: Training Web Agents in Real-World Cloud Environments via Distillation and Reinforcement Learning. CoRR 2026, abs/2606.09447, [2606.09447. [Google Scholar] [CrossRef]
- Jiang, Z.; An, L.; Liu, Y.; Ji, J.; Wu, Q.; Andreas, J.; Zhang, Y.; Chang, S. VISUALSKILL: Multimodal Skills for Computer-Use Agents. CoRR 2026, abs/2606.18448, [2606.18448. [Google Scholar] [CrossRef]
- OpenAI. From model to agent: Equipping the Responses API with a computer environment. 2026. [Google Scholar]
- Sun, J.; Hua, Z.; Xia, Y. AutoEval: A Practical Framework for Autonomous Evaluation of Mobile Agents. CoRR 2025, abs/2503.02403, 2503.02403. [Google Scholar] [CrossRef]
- Liu, G.; Zhao, P.; Liu, L.; Chen, Z.; Chai, Y.; Liang, Y.; Wang, W.; Chen, S.; Lu, Z.; Ren, S.; et al. LearnAct: Few-Shot Mobile GUI Agent with a Unified Demonstration Benchmark. In Proceedings of the Findings of the Association for Computational Linguistics, ACL 2026; San Diego, California, United States, Liakata, M., Moreira, V.P., Zhang, J., Jurgens, D., Eds.; Association for Computational Linguistics, 2-7 July 2026; pp. 29820–29843. [Google Scholar]
- Kong, Q.; Zhang, X.; Yang, Z.; Gao, N.; Liu, C.; Tong, P.; Cai, C.; Zhou, H.; Zhang, J.; Chen, L.; et al. MobileWorld: Benchmarking Autonomous Mobile Agents in Agent-User Interactive and MCP-Augmented Environments. In Proceedings of the Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2026; San Diego, California, United States, Liakata, M., Moreira, V.P., Zhang, J., Jurgens, D., Eds.; Association for Computational Linguistics, 2-7 July 2026; Volume 2026, pp. 6142–6167. [Google Scholar]
- Wang, Z.; Xu, H.; Wang, J.; Zhang, X.; Yan, M.; Zhang, J.; Huang, F.; Ji, H. Mobile-Agent-E: Self-Evolving Mobile Assistant for Complex Tasks. CoRR 2025, abs/2501.11733, 2501.11733. [Google Scholar] [CrossRef]
- Ye, J.; Zhang, X.; Xu, H.; Liu, H.; Wang, J.; Zhu, Z.; Zheng, Z.; Gao, F.; Cao, J.; Lu, Z.; et al. Mobile-Agent-v3: Fundamental Agents for GUI Automation. CoRR 2025, abs/2508.15144, 2508.15144. [Google Scholar] [CrossRef]
- Xu, H.; Zhang, X.; Liu, H.; Wang, J.; Zhu, Z.; Zhou, S.; Hu, X.; Gao, F.; Cao, J.; Wang, Z.; et al. Mobile-Agent-v3.5: Multi-platform Fundamental GUI Agents. CoRR 2026, abs/2602.16855, 2602.16855. [Google Scholar] [CrossRef]
- Shi, Y.; Yu, W.; Li, Z.; Wang, Y.; Zhang, H.; Liu, N.; Mi, H.; Yu, D. MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment. CoRR 2025, abs/2507.05720, 2507.05720. [Google Scholar] [CrossRef]
- Dai, G.; Jiang, S.; Cao, T.; Li, Y.; Yang, Y.; Tan, R.; Li, M.; Qiu, L. Advancing Mobile GUI Agents: A Verifier-Driven Approach to Practical Deployment. CoRR 2025, abs/2503.15937, [2503.15937. [Google Scholar] [CrossRef]
- Wang, Z.; Yu, W.; Ren, X.; Zhang, J.; Zhao, Y.; Saxena, R.; Cheng, L.; Wong, G.Y.; See, S.; Minervini, P.; et al. MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly. In Proceedings of the Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025; San Diego, CA, USA, December 2-7, 2025 / Mexico City, Mexico, Belgrave, D., Zhang, C., Montoya, L.N., Lin, H., Pascanu, R., Koniusz, P., Ghassemi, M., Chen, N., Ruíz, I.V.M., Loaiza-Bonilla, A., Eds.; 30 November 2025. [Google Scholar]
- Chen, G.; Liu, Y.; Huang, Y.; Pei, B.; Xu, J.; He, Y.; Lu, T.; Wang, Y.; Wang, L. CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding. In Proceedings of the The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025; 2025. Available online: https://openreview.net/.
- Li, J.; Wang, J.; Tan, M.; Wang, H.; Yan, C.; Shi, L.; Cai, J.; Jiang, X.; Hu, Y. CrossVid: A Comprehensive Benchmark for Evaluating Cross-Video Reasoning in Multimodal Large Language Models. In Proceedings of the Fortieth AAAI Conference on Artificial Intelligence, Thirty-Eighth Conference on Innovative Applications of Artificial Intelligence, Sixteenth Symposium on Educational Advances in Artificial Intelligence; AAAI 2026, Singapore, Koenig, S., Jenkins, C., Taylor, M.E., Eds.; AAAI Press, 20-27 January 2026; pp. 6244–6252. [Google Scholar] [CrossRef]
- Ma, W.; Ren, W.; Jia, Y.; Li, Z.; Nie, P.; Zhang, G.; Chen, W. VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation. CoRR 2025, abs/2505.14640, 2505.14640. [Google Scholar] [CrossRef]
- Su, Z.; Gao, J.; Guo, H.; Liu, Z.; Zhang, L.; Geng, X.; Huang, S.; Xia, P.; Jiang, G.; Wang, C.; et al. AgentVista: Evaluating Multimodal Agents in Ultra-Challenging Realistic Visual Scenarios. CoRR 2026, abs/2602.23166, [2602.23166. [Google Scholar] [CrossRef]
- Tao, K.; Zheng, Y.; Jia, X.; Du, W.; Shao, K.; Wang, H.; Chen, X.; Jin, X.; Zhu, J.; Yu, B.; et al. LVOmniBench: Pioneering Long Audio-Video Understanding Evaluation for Omnimodal LLMs. CoRR 2026, abs/2603.19217, [2603.19217. [Google Scholar] [CrossRef]
- Padlewski, P.; Bain, M.; Henderson, M.; Zhu, Z.; Relan, N.; Pham, H.; Ong, D.; Aleksiev, K.; Ormazabal, A.; Phua, S.; et al. Vibe-Eval: A hard evaluation suite for measuring progress of multimodal language models. CoRR 2024, abs/2405.02287, 2405.02287. [Google Scholar] [CrossRef]
- Cheng, X.; Zhang, W.; Zhang, S.; Yang, J.; Guan, X.; Wu, X.; Li, X.; Zhang, G.; Liu, J.; Mai, Y.; et al. SimpleVQA: Multimodal Factuality Evaluation for Multimodal Large Language Models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, ICCV 2025, Honolulu, HI, USA, October 19-25, 2025; IEEE; 2025, pp. 4637–4646. [Google Scholar] [CrossRef]
- Guan, T.; Liu, F.; Wu, X.; Xian, R.; Li, Z.; Liu, X.; Wang, X.; Chen, L.; Huang, F.; Yacoob, Y.; et al. Hallusionbench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024, Seattle, WA, USA, June 16-22, 2024; IEEE; 2024, pp. 14375–14385. [Google Scholar] [CrossRef]
- Cao, M.; Hu, P.; Wang, Y.; Gu, J.; Tang, H.; Zhao, H.; Wang, C.; Dong, J.; Yu, W.; Zhang, G.; et al. Video SimpleQA: Towards Factuality Evaluation in Large Video Language Models. In Proceedings of the Fortieth AAAI Conference on Artificial Intelligence, Thirty-Eighth Conference on Innovative Applications of Artificial Intelligence, Sixteenth Symposium on Educational Advances in Artificial Intelligence; AAAI 2026, Singapore, January 20-27, 2026, Koenig, S., Jenkins, C., Taylor, M.E., Eds.; AAAI Press, 2026; pp. 2616–2624. [Google Scholar] [CrossRef]
- Zou, Y.; Jin, S.; Deng, A.; Zhao, Y.; Wang, J.; Chen, C. A.I.R.: Enabling Adaptive, Iterative, and Reasoning-based Frame Selection For Video Question Answering. CoRR 2025, abs/2510.04428, 2510.04428. [Google Scholar] [CrossRef]
- Ren, X.; Xu, L.; Xia, L.; Wang, S.; Yin, D.; Huang, C. VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos. In Proceedings of the Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.1, KDD 2026; Jeju Island, Korea, Parthasarathy, S., Gleich, D.F., Zhang, X., Tok, W.H., Farooq, F., He, Q., Singh, A.K., Wang, H., Liu, Y., Eds.; ACM, 9-13 August 2026; pp. 2390–2401. [Google Scholar] [CrossRef]
- Xue, Z.; Zhang, J.; Xie, X.; Cai, Y.; Liu, Y.; Li, X.; Tao, D. AdaVideoRAG: Omni-Contextual Adaptive Retrieval-Augmented Efficient Long Video Understanding. CoRR 2025, abs/2506.13589, [2506.13589. [Google Scholar] [CrossRef]
- Zeng, X.; Qiu, K.; Zhang, Q.; Li, X.; Wang, J.; Li, J.; Yan, Z.; Tian, K.; Tian, M.; Zhao, X.; et al. StreamForest: Efficient Online Video Understanding with Persistent Event Memory. In Proceedings of the Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025; San Diego, CA, USA, December 2-7, 2025 / Mexico City, Mexico, Belgrave, D., Zhang, C., Montoya, L.N., Lin, H., Pascanu, R., Koniusz, P., Ghassemi, M., Chen, N., Ruíz, I.V.M., Loaiza-Bonilla, A., Eds.; 30 November 2025. [Google Scholar]
- Zhao, B.; Wu, C.; Li, D.; Meng, H.; Li, J.; Zhang, J.; Zhou, J.; Lin, J.; Gao, K.; Cao, K.; et al. Qwen-Image-2.0 Technical Report. arXiv 2026, arXiv:cs. [Google Scholar]
- Team, Q. Qwen VLo: From “Understanding” the World to “Depicting” It; 2025. [Google Scholar]
- Gao, Y.; Gong, L.; Guo, Q.; Hou, X.; Lai, Z.; Li, F.; Li, L.; Lian, X.; Liao, C.; Liu, L.; et al. Seedream 3.0 Technical Report. CoRR 2025, abs/2504.11346, [2504.11346. [Google Scholar] [CrossRef]
- Seedance Team. Seedance 2.0: Advancing Video Generation for World Complexity. arXiv 2026, arXiv:2604.14148. [Google Scholar]
- Chen, X.; Wu, Z.; Liu, X.; Pan, Z.; Liu, W.; Xie, Z.; Yu, X.; Ruan, C. Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling. CoRR 2025, abs/2501.17811, 2501.17811. [Google Scholar] [CrossRef]
- Deng, C.; Zhu, D.; Li, K.; Gou, C.; Li, F.; Wang, Z.; Zhong, S.; Yu, W.; Nie, X.; Song, Z.; et al. Emerging Properties in Unified Multimodal Pretraining. CoRR 2025, abs/2505.14683, [2505.14683. [Google Scholar] [CrossRef]
- Wang, P.; Shi, Y.; Lian, X.; Zhai, Z.; Xia, X.; Xiao, X.; Huang, W.; Yang, J. SeedEdit 3.0: Fast and High-Quality Generative Image Editing. CoRR 2025, abs/2506.05083, 2506.05083. [Google Scholar] [CrossRef]
- He, J.; Ye, J.; Huang, Z.; Jiang, D.; Zhang, C.; Zhu, L.; Zhang, R.; Zhang, X.; Li, W. Mind-Brush: Integrating Agentic Cognitive Search and Reasoning into Image Generation. CoRR 2026, abs/2602.01756, [2602.01756. [Google Scholar] [CrossRef]
- Jiang, K.; Wang, Y.; Zhou, J.; Li, P.; Liu, Z.; Xie, C.; Chen, Z.; Zheng, Y.; Zhang, W. GenAgent: Scaling Text-to-Image Generation via Agentic Multimodal Reasoning. CoRR 2026, abs/2601.18543, 2601.18543. [Google Scholar] [CrossRef]
- Chen, S.; Shou, Q.; Chen, H.; Zhou, Y.; Feng, K.; Hu, W.; Zhang, Y.; Lin, Y.; Huang, W.; Song, M.; et al. Unify-Agent: A Unified Multimodal Agent for World-Grounded Image Synthesis. CoRR 2026, abs/2603.29620, 2603.29620. [Google Scholar] [CrossRef]
- Sun, W.; Wang, Z.; Hu, Z.; Wang, C.; Li, H.; Chen, W. MUSE: A Multi-agent Framework for Unconstrained Story Envisioning via Closed-Loop Cognitive Orchestration. CoRR 2026, abs/2602.03028, 2602.03028. [Google Scholar] [CrossRef]
- Yan, L.; Zhang, Y.; Pan, B.; Zheng, X.; Qian, J.; Wu, A.; Li, W.; Lyu, C. Crayotter: Traceable Multi-Agent Workflows for Long-Form Video Editing. CoRR 2026, abs/2606.07636, [2606.07636. [Google Scholar] [CrossRef]
- Zhang, X.; Zhang, X.; Wu, Y.; Cao, Y.; Zhang, R.; Chu, R.; Yang, L.; Yang, Y. Generative Universal Verifier as Multimodal Meta-Reasoner. CoRR 2025, abs/2510.13804, [2510.13804. [Google Scholar] [CrossRef]
- Chen, W.; Yu, K.; Tian, B.; Song, J.; Liang, S.; Jia, H.; Cheng, K.; Li, H.; Yuan, K.; Wang, L.; et al. MemoGen: Can Past Experience Improve Future Text-to-Image Generation? CoRR 2026. abs/2606.03243 2606.03243. [CrossRef]
- Xu, X.; Mei, J.; Li, C.; Wu, Y.; Yan, M.; Lai, S.; Zhang, J.; Wu, M. MM-StoryAgent: Immersive Narrated Storybook Video Generation with a Multi-Agent Paradigm across Text, Image and Audio. CoRR 2025, abs/2503.05242, [2503.05242. [Google Scholar] [CrossRef]
- Zhang, C.; Dong, G.; Liu, Y.; Zhao, T.; Dou, Z. Towards Verifiable Multimodal Deep Research: A Multi-Agent Harness for Interleaved Report Generation. CoRR 2026, abs/2605.29861, 2605.29861. [Google Scholar] [CrossRef]
- Xu, J.; Guo, Z.; He, J.; Hu, H.; He, T.; Bai, S.; Chen, K.; Wang, J.; Fan, Y.; Dang, K.; et al. Qwen2.5-Omni Technical Report. CoRR 2025, abs/2503.20215, [2503.20215. [Google Scholar] [CrossRef]
- Xu, J.; Guo, Z.; Hu, H.; Chu, Y.; Wang, X.; He, J.; Wang, Y.; Shi, X.; He, T.; Zhu, X.; et al. Qwen3-Omni Technical Report. arXiv 2025, arXiv:cs. [Google Scholar]
- Zhang, F.; Zhang, V.; Qian, S.; Li, H.; Wu, H.; Wu, J.; Zhou, D.; Zhu, Z.; Lian, Z.; Wang, X.; et al. Orchestra-o1: Omnimodal Agent Orchestration. CoRR 2026, abs/2606.13707, [2606.13707. [Google Scholar] [CrossRef]
- Du, P. OmniNova:A General Multimodal Agent Framework. CoRR 2025, abs/2503.20028, 2503.20028. [Google Scholar] [CrossRef]
- Xing, Z.; Xu, R.; Wang, Y.; He, J.; Ma, Z.; Yang, Q.; Chu, Y.; Xu, J.; Lin, J.; Fu, C.; et al. Native Active Perception as Reasoning for Omni-Modal Understanding. CoRR 2026, abs/2606.19341, 2606.19341. [Google Scholar] [CrossRef]
- Xu, K.; Wang, Y.; Cheng, Z.; Liu, H.; Wang, Y.; Wang, Y. Agentic Active Omni-Modal Perception for Multi-Hop Audio-Visual Reasoning. CoRR 2026, abs/2605.28192, 2605.28192. [Google Scholar] [CrossRef]
- Li, X.; Ming, R.; Setlur, P.; Paladugu, A.; Tang, A.; Kang, H.; Shao, S.; Jin, R.; Xiong, C. Benchmark Test-Time Scaling of General LLM Agents. CoRR 2026, abs/2602.18998, [2602.18998. [Google Scholar] [CrossRef]
- Zhang, Y.; Jiang, S.; Li, R.; Tu, J.; Su, Y.; Deng, L.; Guo, X.; Lv, C.; Lin, J. DeepPlanning: Benchmarking Long-Horizon Agentic Planning with Verifiable Constraints. In Proceedings of the Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics; San Diego, California, United States, Liakata, M., Moreira, V.P., Zhang, J., Jurgens, D., Eds.; Association for Computational Linguistics, 2-7 July 2026; Volume 1, pp. 7377–7407. [Google Scholar]
- Vidgen, B.; Mann, A.; Fennelly, A.; Stanly, J.W.; Rothman, L.; Burstein, M.; Benchek, J.; Ostrofsky, D.; Ravichandran, A.; Sur, D.; et al. APEX-Agents. CoRR 2026, abs/2601.14242, 2601.14242. [Google Scholar] [CrossRef]
- Chen, X.; Zhu, J.; Li, P.; Wang, H.; Yang, S.; Guo, M. PresentBench: A Fine-Grained Rubric-Based Benchmark for Slide Generation. CoRR 2026, abs/2603.07244, [2603.07244. [Google Scholar] [CrossRef]
- Ma, Z.; Zhang, B.; Zhang, J.; Yu, J.; Zhang, X.; Zhang, X.; Luo, S.; Wang, X.; Tang, J. SpreadsheetBench: Towards Challenging Real World Spreadsheet Manipulation. In Proceedings of the Advances in Neural Information Processing Systems 37: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024; Vancouver, BC, Canada, Globersons, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J.M., Zhang, C., Eds.; 10 - 15 December 2024. [Google Scholar]
- OpenAI. Introducing Codex. 2025. [Google Scholar]
- Zhu, N.; Wang, H.; Zhou, J.; Chen, F.; Zhang, S.; Chen, G.; Liu, C.; Wu, J.; Chen, W.; Mou, X.; et al. SemaClaw: A Step Towards General-Purpose Personal AI Agents through Harness Engineering. CoRR 2026, abs/2604.11548, [2604.11548. [Google Scholar] [CrossRef]
- Manus Team. Manus: Hands On AI. 2025. [Google Scholar]
- Yeh, C.; Wang, C.; Tong, S.; Cheng, T.Y.; Wang, R.; Chu, T.; Zhai, Y.; Chen, Y.; Gao, S.; Ma, Y. Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs. In Proceedings of the Fortieth AAAI Conference on Artificial Intelligence, Thirty-Eighth Conference on Innovative Applications of Artificial Intelligence, Sixteenth Symposium on Educational Advances in Artificial Intelligence; AAAI 2026, Singapore, Koenig, S., Jenkins, C., Taylor, M.E., Eds.; AAAI Press, 20-27 January 2026; pp. 12000–12008. [Google Scholar] [CrossRef]
- Yang, S.; Xu, R.; Xie, Y.; Yang, S.; Li, M.; Lin, J.; Zhu, C.; Chen, X.; Duan, H.; Yue, X.; et al. MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence. CoRR 2025, abs/2505.23764, [2505.23764. [Google Scholar] [CrossRef]
- Du, M.; Wu, B.; Li, Z.; Huang, X.; Wei, Z. EmbSpatial-Bench: Benchmarking Spatial Understanding for Embodied Tasks with Large Vision-Language Models. In Proceedings of the Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics; Bangkok, Thailand, Ku, L., Martins, A., Srikumar, V., Eds.; Association for Computational Linguistics, 11-16 August 2024; Volume 2, pp. 346–355. [Google Scholar] [CrossRef]
- Wong, L.H.K.; Kang, X.; Bai, K.; Zhang, J. A Survey of Robotic Navigation and Manipulation with Physics Simulators in the Era of Embodied AI. CoRR 2025, abs/2505.01458, 2505.01458. [Google Scholar] [CrossRef]
- Zhou, E.; An, J.; Chi, C.; Han, Y.; Rong, S.; Zhang, C.; Wang, P.; Wang, Z.; Huang, T.; Sheng, L.; et al. RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics. CoRR 2025, abs/2506.04308, 2506.04308. [Google Scholar] [CrossRef]
- Zheng, Z.; Yan, X.; Chen, Z.; Wang, J.; Lim, Q.Z.E.; Tenenbaum, J.B.; Gan, C. ContPhy: Continuum Physical Concept Learning and Reasoning from Videos. In Proceedings of the Forty-first International Conference on Machine Learning, ICML 2024 PMLR / OpenReview.net; Vienna, Austria, Salakhutdinov, R., Kolter, Z., Heller, K.A., Weller, A., Oliver, N., Scarlett, J., Berkenkamp, F., Eds.; Proceedings of Machine Learning Research; 21-27 July 2024; Vol. 235, pp. 61526–61558. [Google Scholar]
- Chow, W.; Mao, J.; Li, B.; Seita, D.; Guizilini, V.C.; Wang, Y. PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding. In Proceedings of the The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025; 2025. Available online: https://openreview.net/.
- Team, G.R. Gemini Robotics: Bringing AI into the Physical World. CoRR 2025, abs/2503.20020, 2503.20020. [Google Scholar] [CrossRef]
- Bjorck, J.; Castañeda, F.; Cherniadev, N.; Da, X.; Ding, R.; Fan, L.; Fang, Y.; Fox, D.; Hu, F.; Huang, S.; et al. GR00T N1: An Open Foundation Model for Generalist Humanoid Robots. CoRR 2025, abs/2503.14734, 2503.14734. [Google Scholar] [CrossRef]
- Intelligence, P.; Black, K.; Brown, N.; Darpinian, J.; Dhabalia, K.; Driess, D.; Esmail, A.; Equi, M.; Finn, C.; Fusai, N.; et al. π0.5: a Vision-Language-Action Model with Open-World Generalization. CoRR 2025, abs/2504.16054, [2504.16054. [Google Scholar] [CrossRef]
- Wang, Q.; Li, M.; Guan, J.; Ye, J.; Xie, S.; Liu, Y.; Chen, J.; Liang, Z.; Zhang, J.; Hu, X.; et al. Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments. CoRR 2026, abs/2605.30280, 2605.30280. [Google Scholar] [CrossRef]
- Yuan, H.; Liang, Z.; Chen, A.; Wang, Y.; Li, H.; Lin, P.; Huang, Y.; Lei, Z.; Zhang, T.; Zhang, J.; et al. Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models. CoRR 2026, abs/2606.17846, [2606.17846. [Google Scholar] [CrossRef]
- Zhang, J.; Zhou, G.; Yin, H.; Huang, Y.; Lei, Z.; Peng, Q.; Yuan, H.; Zhang, J.; Guo, X.; Chen, X.; et al. Qwen-RobotNav Technical Report: A Scalable Navigation Model Designed for an Agentic Navigation System. CoRR 2026, abs/2606.18112, 2606.18112. [Google Scholar] [CrossRef]
- Zhou, G.; Pan, H.; LeCun, Y.; Pinto, L. DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning. In Proceedings of the Forty-second International Conference on Machine Learning, ICML 2025; Vancouver, BC, Canada, Singh, A., Fazel, M., Hsu, D., Lacoste-Julien, S., Berkenkamp, F., Maharaj, T., Wagstaff, K., Zhu, J., Eds.; PMLR / OpenReview.net; Proceedings of Machine Learning Research; 13-19 July 2025; Vol. 267. [Google Scholar]
- NVIDIA. World Simulation with Video Foundation Models for Physical AI. CoRR 2025, abs/2511.00062, 2511.00062. [Google Scholar] [CrossRef]
- Wu, Z.; Gao, J. OSCAR: Omni-Embodiment Action-Conditioned World Model for Robotics. CoRR 2026, abs/2606.04463, 2606.04463. [Google Scholar] [CrossRef]
- Zhou, P.; Chen, S.; Chen, D.; Wang, J.; Jin, R.; Zhu, B.; Pan, Y.; Gu, S.; Wang, K.; Nan, S.; et al. τ0-WM: A Unified Video-Action World Model for Robotic Manipulation. CoRR 2026, abs/2606.01027, [2606.01027. [Google Scholar] [CrossRef]
- Zhou, T.; Wang, P.; Wu, Y.; Yang, H. FinRobot: AI Agent for Equity Research and Valuation with Large Language Models. CoRR 2024, abs/2411.08804, [2411.08804. [Google Scholar] [CrossRef]
- Bigeard, A.; Nashold, L.; Krishnan, R.; Wu, S. Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks. CoRR 2025, abs/2508.00828, 2508.00828. [Google Scholar] [CrossRef]
- Jin, J.; Zhang, Y.; Xu, Y.; Qian, H.; Zhu, Y.; Dou, Z. FinSight: Towards Real-World Financial Deep Research. In Proceedings of the Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2026; San Diego, California, United States, Liakata, M., Moreira, V.P., Zhang, J., Jurgens, D., Eds.; Association for Computational Linguistics, 2-7 July 2026; Volume 2026, pp. 5868–5894. [Google Scholar]
- Li, H.; Chen, J.; Yang, J.; Ai, Q.; Jia, W.; Liu, Y.; Lin, K.; Wu, Y.; Yuan, G.; Hu, Y.; et al. LegalAgentBench: Evaluating LLM Agents in Legal Domain. In Proceedings of the Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics; Vienna, Austria, Che, W., Nabende, J., Shutova, E., Pilehvar, M.T., Eds.; Association for Computational Linguistics, 2025; Volume 1, pp. 2322–2344. [Google Scholar] [CrossRef]
- Grupen, N.; Pereyra, G.; Pereyra, J. Introducing Harvey’s Legal Agent Benchmark. In Harvey Blog; 2026. [Google Scholar]
- Geng, H.; Liu, L. Parthenon Law: A Self-Evolving Legal-Agent Framework. CoRR 2026, abs/2606.04602, 2606.04602. [Google Scholar] [CrossRef]
- Yang, X.; Deng, C.; Wen, T.; Xie, B.; Dou, Z. LawThinker: A Deep Research Legal Agent in Dynamic Environments. CoRR 2026, abs/2602.12056, 2602.12056. [Google Scholar] [CrossRef]
- Zuo, Y.; Qu, S.; Li, Y.; Chen, Z.; Zhu, X.; Hua, E.; Zhang, K.; Ding, N.; Zhou, B. MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding. In Proceedings of the Forty-second International Conference on Machine Learning, ICML 2025; Vancouver, BC, Canada, Singh, A., Fazel, M., Hsu, D., Lacoste-Julien, S., Berkenkamp, F., Maharaj, T., Wagstaff, K., Zhu, J., Eds.; PMLR / OpenReview.net; Proceedings of Machine Learning Research; 13-19 July 2025; Vol. 267. [Google Scholar]
- Arora, R.K.; Wei, J.; Hicks, R.S.; Bowman, P.; Candela, J.Q.; Tsimpourlas, F.; Sharman, M.; Shah, M.; Vallone, A.; Beutel, A.; et al. HealthBench: Evaluating Large Language Models Towards Improved Human Health. CoRR 2025, abs/2505.08775, 2505.08775. [Google Scholar] [CrossRef]
- Jiang, Y.; Black, K.C.; Geng, G.; Park, D.; Zou, J.; Ng, A.Y.; Chen, J.H. MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents. arXiv 2025, arXiv:cs. [Google Scholar]
- Zhu, Y.; He, Z.; Hu, H.; Zheng, X.; Zhang, X.; Wang, J.; Gao, J.; Ma, L.; Yu, L. MedAgentBoard: Benchmarking Multi-Agent Collaboration with Conventional Methods for Diverse Medical Tasks. In Proceedings of the Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025; San Diego, CA, USA, December 2-7, 2025 / Mexico City, Mexico, Belgrave, D., Zhang, C., Montoya, L.N., Lin, H., Pascanu, R., Koniusz, P., Ghassemi, M., Chen, N., Ruíz, I.V.M., Loaiza-Bonilla, A., Eds.; 30 November 2025. [Google Scholar]
- Liu, R.; Mohiuddin, I.Q.; Schoeffler, A.J.; Renduchintala, K.; Nayak, A.; Vemu, P.L.; Vedak, S.; Black, K.C.; Havlik, J.L.; Ogunmola, I.; et al. PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments. CoRR 2026, abs/2605.02240, 2605.02240. [Google Scholar] [CrossRef]
- Ashraf, T.; Jeong, H.; Thoker, F.M.; Ghanem, B. MedCTA: A Benchmark for Clinical Tool Agents. CoRR 2026, abs/2606.11702, [2606.11702. [Google Scholar] [CrossRef]
- Yang, Q.; Liu, Y.; Li, J.; Bai, J.; Chen, H.; Chen, K.; Duan, T.; Dong, J.; Hu, X.; Jia, Z.; et al. $OneMillion-Bench: How Far are Language Agents from Human Experts? CoRR 2026. abs/2603.07980 2603.07980. [CrossRef]
- Patwardhan, T.; Dias, R.; Proehl, E.; Kim, G.; Wang, M.; Watkins, O.; Fishman, S.P.; Aljubeh, M.; Thacker, P.; Fauconnet, L.; et al. GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks. CoRR 2025, abs/2510.04374, 2510.04374. [Google Scholar] [CrossRef]
- Wang, M.; Lin, R.; Hu, K.; Jiao, J.; Chowdhury, N.; Chang, E.; Patwardhan, T. FrontierScience: Evaluating AI’s Ability to Perform Expert-Level Scientific Tasks. CoRR 2026, abs/2601.21165, [2601.21165. [Google Scholar] [CrossRef]
- Zhou, J.; Chen, J.; Hao, L.; Cao, D.; Wang, Z.; Chen, Q.; Fu, C.; Chen, J.; Wu, Y.; Zhang, G.; et al. BABE: Biology Arena BEnchmark. CoRR 2026, abs/2602.05857, 2602.05857. [Google Scholar] [CrossRef]
- Zheng, T.; Deng, Z.; Tsang, H.T.; Wang, W.; Bai, J.; Wang, Z.; Song, Y. From Automation to Autonomy: A Survey on Large Language Models in Scientific Discovery. In Proceedings of the Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025; Suzhou, China, Christodoulopoulos, C., Chakraborty, T., Rose, C., Peng, V., Eds.; Association for Computational Linguistics, 4-9 November 2025; Volume 2025, pp. 17733–17750. [Google Scholar] [CrossRef]
- Rank, B.; Bhatnagar, H.; Prabhu, A.; Eisenberg, S.; Nguyen, K.; Bethge, M.; Andriushchenko, M. PostTrainBench: Can LLM Agents Automate LLM Post-Training? CoRR 2026. abs/2603.08640 2603.08640. [CrossRef]
- Lu, C.; Lu, C.; Lange, R.T.; Foerster, J.N.; Clune, J.; Ha, D. The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery. CoRR 2024, abs/2408.06292, [2408.06292. [Google Scholar] [CrossRef]
- Yamada, Y.; Lange, R.T.; Lu, C.; Hu, S.; Lu, C.; Foerster, J.N.; Clune, J.; Ha, D. The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search. CoRR 2025, abs/2504.08066, [2504.08066. [Google Scholar] [CrossRef]
- Mitchener, L.; Yiu, A.; Chang, B.; Bourdenx, M.; Nadolski, T.; Sulovari, A.; Landsness, E.C.; Barabasi, D.L.; Narayanan, S.; Evans, N.; et al. Kosmos: An AI Scientist for Autonomous Discovery. CoRR 2025, abs/2511.02824, 2511.02824. [Google Scholar] [CrossRef]
- Zhang, B.; Feng, S.; Yan, X.; Yuan, J.; Yu, Z.; He, X.; Huang, S.; Hou, S.; Nie, Z.; Wang, Z.; et al. NovelSeek: When Agent Becomes the Scientist - Building Closed-Loop System from Hypothesis to Verification. CoRR 2025, abs/2505.16938, [2505.16938. [Google Scholar] [CrossRef]
- Boiko, D.A.; MacKnight, R.; Kline, B.; Gomes, G. Autonomous chemical research with large language models. Nat. 2023, 624, 570–578. [Google Scholar] [CrossRef]
- Ghareeb, A.E.; Chang, B.; Mitchener, L.; Yiu, A.; Szostkiewicz, C.J.; Laurent, J.M.; Razzak, M.T.; White, A.D.; Hinks, M.M.; Rodriques, S.G. Robin: A multi-agent system for automating scientific discovery. CoRR 2025, abs/2505.13400, 2505.13400. [Google Scholar] [CrossRef]
- Swanson, K.; Wu, W.; Bulaong, N.L.; Pak, J.E.; Zou, J. The Virtual Lab of AI agents designs new SARS-CoV-2 nanobodies. Nat. 2025, 646, 716–723. [Google Scholar] [CrossRef]
- Zheng, K.; Han, J.M.; Polu, S. miniF2F: a cross-system benchmark for formal Olympiad-level mathematics. In Proceedings of the The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022; 2022. Available online: https://openreview.net/.
- Tsoukalas, G.; Lee, J.; Jennings, J.; Xin, J.; Ding, M.; Jennings, M.; Thakur, A.; Chaudhuri, S. PutnamBench: Evaluating Neural Theorem-Provers on the Putnam Mathematical Competition. In Proceedings of the Advances in Neural Information Processing Systems 37: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024; Vancouver, BC, Canada, Globersons, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J.M., Zhang, C., Eds.; 10 - 15 December 2024. [Google Scholar]
- Chen, J.; Chen, W.; Du, J.; Hu, J.; Jiang, Z.; Jie, A.; Jin, X.; Jin, X.; Li, C.; Shi, W.; et al. Seed-Prover 1.5: Mastering Undergraduate-Level Theorem Proving via Learning from Experience. CoRR 2025, abs/2512.17260, [2512.17260. [Google Scholar] [CrossRef]
- Baba, K.; Liu, C.; Kurita, S.; Sannai, A. Prover Agent: An Agent-based Framework for Formal Mathematical Proofs. CoRR 2025, abs/2506.19923, 2506.19923. [Google Scholar] [CrossRef]
- Xin, H.; Li, L.; Jin, X.; Fleuriot, J.; Li, W. APE-Bench: Evaluating Automated Proof Engineering for Formal Math Libraries. arXiv 2026, arXiv:cs. [Google Scholar]
- Wang, E.Y.; Motwani, S.R.; Roggeveen, J.V.; Hodges, E.; Jayalath, D.; London, C.; Ramakrishnan, K.; Cipcigan, F.; Torr, P.; Abate, A. HorizonMath: Measuring AI Progress Toward Mathematical Discovery with Automatic Verification. CoRR 2026, abs/2603.15617, [2603.15617. [Google Scholar] [CrossRef]
- Tian, R.; Ye, Y.; Qin, Y.; Cong, X.; Lin, Y.; Pan, Y.; Wu, Y.; Hui, H.; Liu, W.; Liu, Z.; et al. DebugBench: Evaluating Debugging Capability of Large Language Models. Proceedings of the Findings of the Association for Computational Linguistics, ACL 2024, Bangkok, Thailand and virtual meeting 2024, 2024, 4173–4198. [Google Scholar] [CrossRef]
- Wei, J.; Sun, Z.; Papay, S.; McKinney, S.; Han, J.; Fulford, I.; Chung, H.W.; Passos, A.T.; Fedus, W.; Glaese, A. BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents. CoRR 2025, abs/2504.12516, 2504.12516. [Google Scholar] [CrossRef]
- Li, B.; Zhang, B.; Zhang, D.; Huang, F.; Li, G.; Chen, G.; Yin, H.; Wu, J.; Zhou, J.; Li, K.; et al. Tongyi DeepResearch Technical Report. CoRR 2025, abs/2510.24701, [2510.24701. [Google Scholar] [CrossRef]
- Team, M.; Bai, S.; Bing, L.; Chen, C.; Chen, G.; Chen, Y.; Chen, Z.; Chen, Z.; Dai, J.; Dong, X.; et al. MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling. CoRR 2025, abs/2511.11793, [2511.11793. [Google Scholar] [CrossRef]
- Wong, R.; Wang, J.; Zhao, J.; Chen, L.; Gao, Y.; Zhang, L.; Zhou, X.; Wang, Z.; Xiang, K.; Zhang, G.; et al. WideSearch: Benchmarking Agentic Broad Info-Seeking. CoRR 2025, abs/2508.07999, 2508.07999. [Google Scholar] [CrossRef]
- Zhu, Y.; Zhang, X.; Zhang, M.; Jin, J.; Zhang, L.; Song, X.; Zhao, K.; Zeng, W.; Tang, R.; Li, H.; et al. GISA: A Benchmark for General Information-Seeking Assistant. CoRR 2026, abs/2602.08543, 2602.08543. [Google Scholar] [CrossRef]
- Tian, S.; Zhang, Z.; Chen, L.; Liu, Z. MMInA: Benchmarking Multihop Multimodal Internet Agents. In Proceedings of the Findings of the Association for Computational Linguistics, ACL 2025, Vienna, Austria, July 27 - August 1; 2025, pp. 13682–13697. [CrossRef]
- Li, S.; Bu, X.; Wang, W.; Liu, J.; Dong, J.; He, H.; Lu, H.; Zhang, H.; Jing, C.; Li, Z.; et al. MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents. CoRR 2025, abs/2508.13186, [2508.13186. [Google Scholar] [CrossRef]
- Team, D. DeepConsult: A Deep Research Benchmark for Consulting and Business Queries. GitHub repository, 2025. Available online: https://github.com/youdotcom-oss/ydc-deep-research-evals.
- He, H.; Yao, W.; Ma, K.; Yu, W.; Dai, Y.; Zhang, H.; Lan, Z.; Yu, D. WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models. In Proceedings of the Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics; Bangkok, Thailand, Ku, L., Martins, A., Srikumar, V., Eds.; Association for Computational Linguistics, 11-16 August 2024; Volume 1, pp. 6864–6890. [Google Scholar] [CrossRef]
- Cheng, L.; Duan, J.; Wang, Y.R.; Fang, H.; Li, B.; Huang, Y.; Wang, E.; Eftekhar, A.; Lee, J.; Yuan, W.; et al. PointArena: Probing Multimodal Grounding Through Language-Guided Pointing. CoRR 2025, abs/2505.09990, 2505.09990. [Google Scholar] [CrossRef]
- Seed, B. Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity. arXiv 2026, arXiv:cs. [Google Scholar]
- Xu, Y.; Liu, X.; Sun, X.; Cheng, S.; Yu, H.; Lai, H.; Zhang, S.; Zhang, D.; Tang, J.; Dong, Y. AndroidLab: Training and Systematic Benchmarking of Android Autonomous Agents. In Proceedings of the Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1; 2025, pp. 2144–2166. [CrossRef]
- Mangalam, K.; Akshulakov, R.; Malik, J. EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding. In Proceedings of the Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023, 2023. [Google Scholar]
- Li, K.; Wang, Y.; He, Y.; Li, Y.; Wang, Y.; Liu, Y.; Wang, Z.; Xu, J.; Chen, G.; Lou, P.; et al. MVBench: A Comprehensive Multi-modal Video Understanding Benchmark. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024, Seattle, WA, USA, June 16-22, 2024; IEEE; 2024, pp. 22195–22206. [Google Scholar] [CrossRef]
- Fu, C.; Dai, Y.; Luo, Y.; Li, L.; Ren, S.; Zhang, R.; Wang, Z.; Zhou, C.; Shen, Y.; Zhang, M.; et al. Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2025, Nashville, TN, USA, June 11-15; 2025, pp. 24108–24118. [CrossRef]
- Wu, H.; Li, D.; Chen, B.; Li, J. LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding. In Proceedings of the Advances in Neural Information Processing Systems 37: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, 2024. [Google Scholar]
- Lin, J.; Fang, Z.; Chen, C.; Wan, Z.; Luo, F.; Li, P.; Liu, Y.; Sun, M. StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding. CoRR 2024, abs/2411.03628, [2411.03628. [Google Scholar] [CrossRef]
- Li, B.; Lin, Z.; Pathak, D.; Li, J.; Fei, Y.; Wu, K.; Ling, T.; Xia, X.; Zhang, P.; Neubig, G.; et al. GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation. CoRR 2024, abs/2406.13743, [2406.13743. [Google Scholar] [CrossRef]
- Li, C.; Chen, Y.; Ji, Y.; Xu, J.; Cui, Z.; Li, S.; Zhang, Y.; Tang, J.; Song, Z.; Zhang, D.; et al. OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs. CoRR 2025, abs/2510.10689, [2510.10689. [Google Scholar] [CrossRef]
- Krantz, J.; Wijmans, E.; Majumdar, A.; Batra, D.; Lee, S. Beyond the Nav-Graph: Vision-and-Language Navigation in Continuous Environments. In Proceedings of the Computer Vision - ECCV 2020 - 16th European Conference, Glasgow, UK, August 23-28, 2020; Proceedings, Part XXVIII, 2020, pp. 104–120. [Google Scholar] [CrossRef]
- Yang, R.; Chen, H.; Zhang, J.; Zhao, M.; Qian, C.; Wang, K.; Wang, Q.; Koripella, T.V.; Movahedi, M.; Li, M.; et al. EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents. CoRR 2025, abs/2502.09560, 2502.09560. [Google Scholar] [CrossRef]
- Bran, A.M.; Cox, S.; Schilter, O.; Baldassari, C.; White, A.D.; Schwaller, P. Augmenting large language models with chemistry tools. Nat. Mac. Intell. 2024, 6, 525–535. [Google Scholar] [CrossRef]
- He, M.; Jain, A.; Kumar, A.; Tu, V.; Bakshi, S.; Patro, S.; Rajani, N. YC-Bench: Benchmarking AI Agents for Long-Term Planning and Consistent Execution. CoRR 2026, abs/2604.01212, 2604.01212. [Google Scholar] [CrossRef]
- Xia, P.; Chen, J.; Yang, X.; Tu, H.; Liu, J.; Xiong, K.; Han, S.; Qiu, S.; Ji, H.; Zhou, Y.; et al. MetaClaw: Just Talk - An Agent That Meta-Learns and Evolves in the Wild. CoRR 2026, abs/2603.17187, [2603.17187. [Google Scholar] [CrossRef]
- Hao, Z.; Wang, H.; Luo, J.; Zhang, J.; Zhou, Y.; Lin, Q.; Wang, C.; Dong, H.; Chen, J. ReCreate: Reasoning and Creating Domain Agents Driven by Experience. In Proceedings of the Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics; San Diego, California, United States, Liakata, M., Moreira, V.P., Zhang, J., Jurgens, D., Eds.; Association for Computational Linguistics, 2-7 July 2026; Volume 1, pp. 31018–31046. [Google Scholar]
- Zhang, H.; Long, Q.; Bao, J.; Feng, T.; Zhang, W.; Yue, H.; Wang, W. MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents. CoRR 2026, abs/2602.02474, 2602.02474. [Google Scholar] [CrossRef]
- Wang, S.; Tan, Z.; Chen, Z.; Zhou, S.; Chen, T.; Li, J. AnyMAC: Cascading Flexible Multi-Agent Collaboration via Next-Agent Prediction. In Proceedings of the Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025; Suzhou, China, Christodoulopoulos, C., Chakraborty, T., Rose, C., Peng, V., Eds.; Association for Computational Linguistics, 4-9 November 2025; Volume 2025, pp. 11555–11567. [Google Scholar] [CrossRef]
- Wang, Z.; Wang, Y.; Liu, X.; Ding, L.; Zhang, M.; Liu, J.; Zhang, M. AgentDropout: Dynamic Agent Elimination for Token-Efficient and High-Performance LLM-Based Multi-Agent Collaboration. In Proceedings of the Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics; Vienna, Austria, Che, W., Nabende, J., Shutova, E., Pilehvar, M.T., Eds.; 27 July 2025; Volume 1, pp. 24013–24035. [Google Scholar] [CrossRef]
- Lee, Y.; Nair, R.; Zhang, Q.; Lee, K.; Khattab, O.; Finn, C. Meta-Harness: End-to-End Optimization of Model Harnesses. CoRR 2026, abs/2603.28052, [2603.28052. [Google Scholar] [CrossRef]
- Lou, X.; Lázaro-Gredilla, M.; Dedieu, A.; Wendelken, C.; Lehrach, W.; Murphy, K.P. AutoHarness: improving LLM agents by automatically synthesizing a code harness. CoRR 2026, abs/2603.03329, 2603.03329. [Google Scholar] [CrossRef]
- Mei, K.; Li, Z.; Xu, S.; Ye, R.; Ge, Y.; Zhang, Y. AIOS: LLM Agent Operating System. CoRR 2024, abs/2403.16971, 2403.16971. [Google Scholar] [CrossRef]
- Ye, H.; He, X.; Arak, V.; Dong, H.; Song, G. Meta Context Engineering via Agentic Skill Evolution. CoRR 2026, abs/2601.21557, [2601.21557. [Google Scholar] [CrossRef]
- Wang, L.; Zhang, X.; Su, H.; Zhu, J. A Comprehensive Survey of Continual Learning: Theory, Method and Application. IEEE Trans. Pattern Anal. Mach. Intell. 2024, 46, 5362–5383. [Google Scholar] [CrossRef]
- Chen, H.; Sun, Z.; Ye, H.; Li, K.; Lin, X. Continual Learning in Large Language Models: Methods, Challenges, and Opportunities. CoRR 2026, abs/2603.12658, [2603.12658. [Google Scholar] [CrossRef]
- Liang, J.; Huang, W.; Xia, F.; Xu, P.; Hausman, K.; Ichter, B.; Florence, P.; Zeng, A. Code as Policies: Language Model Programs for Embodied Control. In Proceedings of the IEEE International Conference on Robotics and Automation, ICRA 2023, London, UK, May 29 - June 2, 2023; IEEE; 2023, pp. 9493–9500. [Google Scholar] [CrossRef]
- Driess, D.; Xia, F.; Sajjadi, M.S.M.; Lynch, C.; Chowdhery, A.; Ichter, B.; Wahid, A.; Tompson, J.; Vuong, Q.; Yu, T.; et al. PaLM-E: An Embodied Multimodal Language Model. In Proceedings of the International Conference on Machine Learning, ICML 2023 PMLR; Honolulu, Hawaii, USA, Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., Scarlett, J., Eds.; Proceedings of Machine Learning Research; 23-29 July 2023; Vol. 202, pp. 8469–8488. [Google Scholar]
- Zitkovich, B.; Yu, T.; Xu, S.; Xu, P.; Xiao, T.; Xia, F.; Wu, J.; Wohlhart, P.; Welker, S.; Wahid, A.; et al. RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control. In Proceedings of the Conference on Robot Learning, CoRL 2023 PMLR; Atlanta, GA, USA, Tan, J., Toussaint, M., Darvish, K., Eds.; Proceedings of Machine Learning Research; 6-9 November 2023; Vol. 229, pp. 2165–2183. [Google Scholar]
- Black, K.; Brown, N.; Driess, D.; Esmail, A.; Equi, M.; Finn, C.; Fusai, N.; Groom, L.; Hausman, K.; Ichter, B.; et al. π0: A Vision-Language-Action Flow Model for General Robot Control. CoRR 2024, abs/2410.24164, [2410.24164. [Google Scholar] [CrossRef]
- Huang, J.; Chen, X.; Mishra, S.; Zheng, H.S.; Yu, A.W.; Song, X.; Zhou, D. Large Language Models Cannot Self-Correct Reasoning Yet. In Proceedings of the The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024; 2024. Available online: https://openreview.net/.
- Zou, W.; Dong, M.; Calvo, M.R.; Chang, S.; Guo, J.; Lee, D.; Niu, X.; Ma, X.; Qi, Y.; Jiang, J. Poison Once, Exploit Forever: Environment-Injected Memory Poisoning Attacks on Web Agents. CoRR 2026, abs/2604.02623, 2604.02623. [Google Scholar] [CrossRef]
- Lin, X.; Liu, Y.; Chen, Y.; Wu, Y.; Ning, Y.; Liu, Y.; Sun, N.; Zhang, S.; Chong, B.; Zhou, C.; et al. SafeHarness: Lifecycle-Integrated Security Architecture for LLM-based Agent Deployment. CoRR 2026, abs/2604.13630, 2604.13630. [Google Scholar] [CrossRef]
- Sigdel, A.; Baral, R. Guardrails as Infrastructure: Policy-First Control for Tool-Orchestrated Workflows. CoRR 2026, abs/2603.18059, [2603.18059. [Google Scholar] [CrossRef]
| 1 | METR live dashboard, https://metr.org/time-horizons/, accessed May 8, 2026. |
| 2 | Devin (Cognition): https://devin.ai. |
| 3 | |
| 4 | Cursor Rules: https://cursor.com/docs/rules.md. |
| 5 |
AGENTS.md: https://agents.md/. |
| 6 | Codex, Agent Approvals & Security: https://developers.openai.com/codex/agent-approvals-security. Anthropic, Claude Code documentation: https://code.claude.com/docs/en/overview. |
| 7 | Anthropic, Claude Code Hooks reference: https://code.claude.com/docs/en/hooks. |
| 8 | |
| 9 | browser-use: https://github.com/browser-use/browser-use. |
| 10 | Open Interpreter: https://github.com/OpenInterpreter/open-interpreter. |
| 11 | OpenManus: https://github.com/FoundationAgents/OpenManus. |
| 12 | |
| 13 | |
| 14 | |
| 15 | |
| 16 |














| Stage (era) | Control unit & surface | Key technical routes | Long-horizon profile | ||
|---|---|---|---|---|---|
| C1 | C2 | C3 | |||
|
I. Prompt engineering (2020–2023) |
Single querylanguage of the prompt | • Chain-of-thought reasoning• Decomposition, planning, and search• Prompt optimization and internalization | ◗ | ||
|
II. Context engineering (2023–2025) |
Conditioned callinformation in the context | • Retrieval-augmented generation• Tool use and function calling• Long context, memory, and context engineering | ◗ | ◗ | |
|
III. Runtime harnesses (2025–Present) |
Whole trajectoryruntime around the model | • Agent loops (reason–act–observe)• Harness engineering• Internalizing the loop | ● | ● | ◗ |
| Component (§) | Harness responsibility | Representative mechanisms | Representative systems |
|---|---|---|---|
| Loop & Workflow (§Section 4.1) | Schedules model calls, action execution, and observation feedback across steps | • Linear workflows • Plan–execute • Branching |
ReAct [36], Self-Refine [29] ReWOO [41], Arbor [244] Tree-of-Thoughts [42], LATS [43] |
| Context & Memory (§Section 4.2) | Bounds in-run state and persists what must survive across windows and sessions | • Working context • Factual memory |
ReSum [35], MemAgent [48] Mem0 [245], MemoryBank [46] ExpeL [98], ReasoningBank [144] |
| Tools, MCP & Skills (§Section 4.3) | Interfaces model decisions with external capabilities and reusable procedures | • Tool interfaces & MCP • Active discovery • Skill libraries |
Toolformer [37], MCP [52] AnyTool [246], ToolGen [247] Voyager [143], SkillNet [248] |
| Orchestration (§Section 4.4) | Decomposes goals, assigns roles, and coordinates work beyond a single loop | • Decomposition • Topologies • Routing & protocols |
MetaGPT [53], CAMEL [55] Magentic-One [118], LangGraph [57] GPTSwarm [231], A2A [58] |
| Hooks & Middleware (§Section 4.5) | Enforces allow, block, or hand-off at predefined points on the action path | • Rule-based hooks • User-defined policies • Adaptive guards |
AEGIS [249], IsolateGPT [250] AgentSpec [251], ShieldAgent [62] AdaptiveGuard [252], AGrail [253] |
| Verification (§Section 4.6) | Scores states and outputs for correctness, safety, and task fit before continuing or accepting results | • Assessment targets • Verification levels • Verifier strategies |
SelfCheckGPT [254], VeriGuard [255] PRMs [256], LLM-as-a-Judge [257] Math-Shepherd [258], RLV [259] |
| Method | Type | State form | Train-free | Trigger | Key mechanism |
|---|---|---|---|---|---|
| Working Context | |||||
| Effective Harnesses [9] | Discard | Turns | ✔ | Boundary | Reset context, drop parsed tool output |
| Deep Agents [215] | Discard | Files | ✔ | Boundary | Offload state to files |
| Tool pruning [271] | Discard | Schemas | ✔ | Start-of-run | Prune tools/schemas before prompt |
| LLMLingua [199] | Compress | Summary | ✘ | Per-payload | Token-level prompt compression |
| CompAct [201] | Compress | Summary | ✘ | Per-payload | Actively compress retrieved docs |
| ReSum [35] | Compress | Summary | ✔ | Boundary | Learned trajectory summarization |
| MEM1 [204] | Compress | Summary | ✘ | Per-step | RL fuses memory and reasoning |
| MemAgent [48] | Compress | Summary | ✘ | Per-step | RL over fixed-size memory |
| ACON [203] | Compress | Summary | ✔ | Boundary | Optimizes context-compression policy |
| ContextBudget [274] | Compress | Summary | ✘ | Per-step | Budget-aware context allocation |
| HiAgent [202] | Compress | Plan | ✔ | Boundary | Subgoal-scoped working memory |
| Context-Folding [275] | Compress | Summary | ✘ | Per-step | Folds history under a budget |
| AgentFold [276] | Compress | Summary | ✘ | Per-step | Proactive history folding |
| ACE [205] | Select | NL-note | ✔ | On-demand | Retrieves, evolves context entries |
| LEGOMem [277] | Select | Summary | ✔ | On-demand | Modular per-subtask procedural memory |
| Memex(RL) [278] | Select | Vector | ✘ | On-demand | RL over indexed experience memory |
| Memory-as-Action [281] | Mixed | Summary | ✘ | Per-step | Memory read/write as actions |
| Persistent Memory | |||||
| CLAUDE.md [224] | Factual | NL-note | ✔ | Start-of-run | Project-rule file loaded per run |
| Cursor Rules [283] | Factual | NL-note | ✔ | Start-of-run | Persistent project rules per run |
| AGENTS.md [218] | Factual | NL-note | ✔ | Start-of-run | Portable project-knowledge spec |
| MemoryBank [46] | Factual | Vector | ✔ | On-demand | Store with forgetting-curve updates |
| Mem0 [245] | Factual | Graph | ✔ | On-demand | Linked long-term fact store |
| MemoryOS [47] | Factual | Mixed | ✔ | On-demand | OS-style tiered memory |
| HippoRAG [284] | Factual | Graph | ✔ | On-demand | Graph memory for associative recall |
| A-MEM [285] | Factual | Graph | ✔ | On-demand | Self-organizing memory notes |
| Reflexion [44] | Experiential | NL-note | ✔ | After-failure | Verbal lessons from failures |
| ExpeL [98] | Experiential | NL-note | ✔ | On-demand | Reuses insights from trajectories |
| Synapse [286] | Experiential | Summary | ✔ | On-demand | Trajectory-as-exemplar prompting |
| SAGE [287] | Experiential | NL-note | ✔ | On-demand | Reflective memory-augmented evolution |
| Buffer of Thoughts [288] | Experiential | NL-note | ✔ | On-demand | Reusable thought templates |
| ReasoningBank [144] | Experiential | NL-note | ✔ | On-demand | Distills reusable reasoning strategies |
| AWM [213] | Experiential | Code-skill | ✔ | Boundary | Induces workflows from traces |
| Agent KB [289] | Experiential | NL-note | ✔ | On-demand | Cross-domain experience knowledge base |
| Voyager [143] | Experiential | Code-skill | ✔ | On-demand | Reusable executable skill library |
| MemGPT [45] | Mixed | Mixed | ✔ | On-demand | OS-style context/memory paging |
| Generative Agents [196] | Mixed | Summary | ✔ | On-demand | Memory stream with reflection |
| Method | Train-free | Key mechanism |
|---|---|---|
| Assessment Targets | ||
| SelfCheckGPT [254] | ✔ | Sampling consistency for factuality verification |
| VeriGuard [255] | ✔ | Verified code generation for action safety |
| ToolSafe [136] | ✘ | RL-trained proactive tool-invocation safety guardrail |
| NSVIF [397] | ✔ | Neuro-symbolic constraint checking for instruction faithfulness |
| Multiagent Debate [398] | ✔ | Inter-model deliberation for factual agreement |
| Multi-Agent Verification [399] | ✔ | Independent aspect verifiers for cross-verifier agreement |
| Verification Levels | ||
| Let’s Verify Step by Step [256] | ✘ | Human-supervised process rewards for intermediate steps |
| AgentPRM [400] | ✘ | Per-decision process rewards for agent trajectories |
| CRITIC [63] | ✔ | Tool-interactive stepwise verify-then-correct loop |
| LLM-as-a-Judge [257] | ✔ | Model-based evaluation of final outputs |
| Agent-as-a-Judge [401] | ✔ | Tool-augmented assessment after task completion |
| Reflexion [44] | ✔ | End-of-trial verbal reflection with episodic memory |
| Verifier Strategies | ||
| Math-Shepherd [258] | ✘ | Monte Carlo process labels for verifier training |
| Implicit PRM [402] | ✘ | Process rewards from policy logits without extra training |
| Self-Consistency [158] | ✔ | Majority voting over diverse reasoning paths |
| LATS [43] | ✔ | Tree search with LLM-generated value estimates |
| ReST-MCTS* [403] | ✘ | PRM-guided self-training via iterative tree search |
| Compute-optimal scaling [404] | ✘ | Difficulty-adaptive compute allocation with process verifiers |
| Stage (§) | Internalized capability | Representative mechanisms | Representative works |
|---|---|---|---|
| Architectural Substrate (§Section 5.1) | Affordably representing and serving long histories | • Explicit-context • Compressed-state • Hybrid memory • High-throughput |
Longformer [189], Big Bird [190] Mamba [72], RWKV [438] Jamba [437], Kimi Linear [439] EAGLE [440], GQA [441] |
| Data & Environment Synthesis (§Section 5.2) | Executable, verifiable long-horizon experience | • Task synthesis • Environment synthesis • Trajectory synthesis |
TaskCraft [442], WebShaper [73] SWE-Gym [74], EnvScaler [77] TOUCAN [76], AgentGym-RL [443] |
| Pre-/Mid-training (§Section 5.3) | Reasoning and perception priors | • Reasoning priors • Long-context priors • Multimodal priors • Data-mixture design |
DeepSeek-V3 [68], Qwen3-Coder [79] YaRN [444], LongLoRA [445] Qwen2-VL [446], InternVL3 [447] DoReMi [448], OPUS [82] |
| Fine-tuning (§Section 5.4) | Behavioral and tool-use discipline | • Selection & mixing • Curriculum • Distillation |
AgentTuning [83], LIMI [449] ScalingInter-RL [443], E2H Reasoner [450] Agent Distillation [451], Agent-R [452] |
| Reinforcement Learning (§Section 5.5) | Feedback-driven long-horizon decisions | • Credit assignment • Policy optimization • Sampling strategy • Interaction patterns |
ToolRL [87], RuscaRL [453] DAPO [454], GiGPO [455] Tree-GRPO [456], TCOD [457] GLIDER [458], Memory-R1 [459] |
| On-policy Distillation (§Section 5.6) | Consolidating behavior on the policy’s own states | • Teacher-guided • Self-improving |
GKD [92], DAgger-LLM [460] -Play [461], Self-Distillation Zero [462] |
| Self-evolution (§Section 5.7) | Cross-task experience accumulation | • Offline bootstrap • Online interaction • Co-evolution |
STaR [463], rStar-Math [429] R-Zero [127], Agent0 [464] Env Tuning [465], Agent-World [237] |
| Method | Type | Key Mechanism | Resource |
|---|---|---|---|
| Credit Assignment | |||
| DeepSeekMath [86] | Outcome Reward | Terminal reward for math reasoning |
GitHub
|
| Search-R1 [89] | Outcome Reward | Answer correctness as terminal reward |
GitHub
|
| DeepRetrieval [593] | Outcome Reward | Retrieval-metric terminal reward |
GitHub
|
| ToolRL [87] | Process Reward | Decomposes tool calls into scored aspects |
GitHub
|
| Tool-Star [594] | Process Reward | Monitors answer and tool usage |
GitHub
|
| CriticSearch [595] | Process Reward | Retrospective critic for fine-grained credit | − |
| RaR [596] | Rubric Reward | Expert checklists used directly as rewards | − |
| OpenRubrics [597] | Rubric Reward | Rubrics induced from preference contrasts | − |
| DR Tulu [598] | Rubric Reward | Co-evolving rubrics in deep research |
GitHub
|
| RuscaRL [453] | Rubric Reward | Rubrics as decaying exploration scaffolds |
GitHub
|
| Policy Optimization | |||
| REINFORCE++ [599] | Ratio Clipping | Global advantage normalization |
GitHub
|
| DAPO [454] | Ratio Clipping | Decoupled clip-higher/lower |
GitHub
|
| Dr.GRPO [600] | Ratio Clipping | Removes length/std normalization bias |
GitHub
|
| GSPO [601] | Ratio Clipping | Sequence-level ratio and clipping | − |
| Turn-PPO [602] | Credit Granularity | Turn-level advantage estimation | − |
| StepPO [603] | Credit Granularity | Step-aligned advantage |
GitHub
|
| GiGPO [455] | Credit Granularity | Episode- and step-level advantage |
GitHub
|
| Clip-Cov [604] | Entropy Regularization | Covariance-gated updates control entropy |
GitHub
|
| CURE [605] | Entropy Regularization | Critical-token-guided re-concatenation |
GitHub
|
| EPO [606] | Entropy Regularization | Trajectory entropy regularization |
GitHub
|
| Sampling Strategy | |||
| R1-searcher [607] | Trajectory Rollout | On-policy search-agent trajectory rollout |
GitHub
|
| WebDancer [85] | Trajectory Rollout | Web-search trajectory pipeline |
GitHub
|
| WebRL [496] | Trajectory Rollout | Self-evolving curriculum web trajectories |
GitHub
|
| Tree of Thoughts [42] | Tree Sampling | Organizes thoughts as a search tree |
GitHub
|
| RAP [608] | Tree Sampling | Planning as tree search with world model |
GitHub
|
| LATS [43] | Tree Sampling | Language-agent tree search over actions |
GitHub
|
| TreePO [609] | Tree Sampling | Reuses inference compute across tree paths |
GitHub
|
| LiteResearcher [610] | Budget-Aware Pruning | Difficulty/pass-rate-based pruning |
GitHub
|
| TCOD [457] | Budget-Aware Pruning | Trajectory-depth curriculum |
GitHub
|
| Interaction Patterns | |||
| GLIDER [458] | Hierarchical Framework | Grounds LLMs as decision-making agents |
GitHub
|
| SkillRL [611] | Hierarchical Framework | Evolves a skill library from failures |
GitHub
|
| THOR [612] | Hierarchical Framework | Joint episode- and step-level optimization |
GitHub
|
| AT-GRPO [613] | Multi-agent Framework | Per-agent, per-turn advantage grouping |
GitHub
|
| MATPO [614] | Multi-agent Framework | Multiple roles in one LLM over joint rollouts |
GitHub
|
| SPIRAL [615] | Multi-agent Framework | Role-conditioned advantages |
GitHub
|
| Memory-R1 [459] | Memory-state Learning | Separates memory manager and answer agent |
GitHub
|
| Memory-as-Action [281] | Memory-state Learning | Working-memory editing as policy action |
GitHub
|
| Agent Workflow Memory [213] | Memory-state Learning | Induces reusable workflows |
GitHub
|
| Agent Lightning [548] | Memory-state Learning | Traces workflows into trainable samples |
GitHub
|
| Work | Capability | Type | Description | Date | GitHub Repo |
|---|---|---|---|---|---|
| Software Engineering | |||||
| AutoCodeRover [103] | Repository Grounding | System | Structure-aware fault localization and patching | 2024.04 | https://github.com/AutoCodeRoverSG/auto-code-rover |
| SWE-agent [2] | Repository Grounding | System | Agent–computer interface for repositories | 2024.05 | https://github.com/SWE-agent/SWE-agent |
| OpenHands [40] | Repository Grounding | System | Persistent coding workspace | 2024.07 | https://github.com/All-Hands-AI/OpenHands |
| NL2Repo-Bench [696] | Repository Grounding | Benchmark | From-scratch repository generation | 2025.12 | https://github.com/multimodal-art-projection/NL2RepoBench |
| OctoBench [15] | Repository Grounding | Benchmark | Scaffold-aware instruction following | 2026.01 | https://github.com/MiniMax-AI/mini-vela |
| SWE-bench [1] | Workflow-level Planning | Benchmark | Repository issue resolution | 2023.10 | https://github.com/SWE-bench/SWE-bench |
| Claude Code [224] | Workflow-level Planning | System | Autonomous code editing | 2025.02 | https://github.com/anthropics/claude-code |
| Trae Agent [134] | Workflow-level Planning | System | Ensemble agent for repository issue resolution | 2025.07 | https://github.com/bytedance/trae-agent |
| SWE-bench Pro [697] | Workflow-level Planning | Benchmark | Enterprise repository issue resolution | 2025.09 | https://github.com/scaleapi/SWE-bench_Pro-os |
| Terminal-Bench 2.0 [123] | Workflow-level Planning | Benchmark | Hard terminal-based tasks | 2025.11 | https://github.com/laude-institute/terminal-bench |
| Aider7 | Feedback-driven Repair | System | Iterative pair-programming loop | 2023.06 | https://github.com/Aider-AI/aider |
| DebugBench [864] | Feedback-driven Repair | Benchmark | Multi-language bug diagnosis | 2024.01 | https://github.com/thunlp/DebugBench |
| Agentless [104] | Feedback-driven Repair | System | Localize-then-repair pipeline | 2024.07 | https://github.com/OpenAutoCoder/Agentless |
| SWE-Gym [74] | Feedback-driven Repair | System | Executable training environments | 2024.10 | https://github.com/SWE-Gym/SWE-Gym |
| SWE-smith [75] | Feedback-driven Repair | System | Repair-trajectory synthesis | 2025.04 | https://github.com/SWE-bench/SWE-smith |
| Information Seeking | |||||
| Search-o1 [105] | Deep Search | System | Retrieval and document reading | 2025.01 | https://github.com/sunnynexus/Search-o1 |
| BrowseComp [865] | Deep Search | Benchmark | Persistent web browsing | 2025.04 | https://github.com/openai/simple-evals |
| WebDancer [85] | Deep Search | System | Information-seeking policy | 2025.05 | https://github.com/Alibaba-NLP/DeepResearch |
| Tongyi DeepResearch [866] | Deep Search | System | Agentic deep-research model | 2025.10 | https://github.com/Alibaba-NLP/DeepResearch |
| MiroThinker [867] | Deep Search | System | Interactive-scaling research agent | 2025.11 | https://github.com/MiroMindAI/MiroThinker |
| Open Deep Research [715] | Wide Search | System | Evidence aggregation and synthesis | 2025.06 | https://github.com/langchain-ai/open_deep_research |
| WideSearch [868] | Wide Search | Benchmark | Structured information collection | 2025.08 | https://github.com/ByteDance-Seed/WideSearch |
| FlowSearch [108] | Wide Search | System | Knowledge-flow orchestration | 2025.10 | https://github.com/InternScience/InternAgent |
| GISA [869] | Wide Search | Benchmark | Deep reasoning and broad information aggregation | 2026.02 | https://github.com/RUC-NLPIR/GISA |
| MMInA [870] | Multimodal Grounding | Benchmark | Multi-hop multimodal browsing | 2024.04 | https://github.com/shulin16/MMInA |
| MM-BrowseComp [871] | Multimodal Grounding | Benchmark | Multimodal evidence seeking | 2025.08 | https://github.com/MMBrowseComp/MM-BrowseComp |
| WebWatcher [729] | Multimodal Grounding | System | Vision-language deep research | 2025.08 | https://github.com/Alibaba-NLP/DeepResearch |
| WebThinker [109] | Research Synthesis | System | Autonomous research workflow | 2025.04 | https://github.com/RUC-NLPIR/WebThinker |
| DeepResearch Bench [738] | Research Synthesis | Benchmark | Research-report quality evaluation | 2025.06 | https://github.com/Ayanami0730/deep_research_bench |
| WebWeaver [491] | Research Synthesis | System | Evidence organization and synthesis | 2025.09 | https://github.com/Alibaba-NLP/DeepResearch |
| DeepConsult [872] | Research Synthesis | Benchmark | Consulting and business deep research | ~2025 | https://github.com/youdotcom-oss/ydc-deep-research-evals |
| Computer Use | |||||
| Mind2Web [110] | Browser Agents | Benchmark | Cross-domain web demonstrations | 2023.06 | https://github.com/OSU-NLP-Group/Mind2Web |
| WebArena [5] | Browser Agents | Benchmark | Realistic interactive websites | 2023.07 | https://github.com/web-arena-x/webarena |
| WebVoyager [873] | Browser Agents | Benchmark | End-to-end live-web agent evaluation | 2024.01 | https://github.com/MinorJerry/WebVoyager |
| browser-use8 | Browser Agents | System | Browser control through agent loop | 2024.10 | https://github.com/browser-use/browser-use |
| BrowserAgent [749] | Browser Agents | System | Human-inspired native browsing | 2025.10 | https://github.com/TIGER-AI-Lab/BrowserAgent |
| OSWorld [6] | Desktop GUI Agents | Benchmark | Desktop task environment | 2024.04 | https://github.com/xlang-ai/OSWorld |
| UI-TARS [111] | Desktop GUI Agents | System | Native visual GUI perception | 2025.01 | https://github.com/bytedance/UI-TARS |
| ScreenSpot-Pro [757] | Desktop GUI Agents | Benchmark | High-resolution professional GUI grounding | 2025.01 | https://github.com/likaixin2000/ScreenSpot-Pro-GUI-Grounding |
| PointArena [874] | Desktop GUI Agents | Benchmark | Language-guided visual pointing | 2025.05 | https://github.com/PointArena/PointArena |
| Seed 2.0 [875] | Desktop GUI Agents | System | Frontier multimodal foundation model | 2026.01 | https://github.com/ByteDance-Seed/Seed2.0 |
| AndroidWorld [112] | Mobile Agents | Benchmark | Dynamic Android environment | 2024.05 | https://github.com/google-research/android_world |
| AndroidLab [876] | Mobile Agents | Benchmark | Android agent training framework | 2024.10 | https://github.com/THUDM/Android-Lab |
| Mobile-Agent-v3 [767] | Mobile Agents | System | Fundamental GUI-automation agents | 2025.08 | https://github.com/X-PLUG/MobileAgent |
| Mobile-Agent-v3.5 [768] | Mobile Agents | System | Multi-platform fundamental GUI agents | 2026.01 | https://github.com/X-PLUG/MobileAgent |
| Multimodal Agents | |||||
| EgoSchema [877] | Multimodal Understanding | Benchmark | Egocentric video question answering | 2023.08 | https://github.com/egoschema/EgoSchema |
| MVBench [878] | Multimodal Understanding | Benchmark | Comprehensive video understanding | 2023.11 | https://github.com/OpenGVLab/Ask-Anything |
| VideoAgent [113] | Multimodal Understanding | System | Agentic evidence selection over long video | 2024.03 | https://github.com/wxh1996/VideoAgent |
| Video-MME [879] | Multimodal Understanding | Benchmark | Long-video understanding | 2024.05 | https://github.com/BradyFU/Video-MME |
| LongVideoBench [880] | Multimodal Understanding | Benchmark | Long-video and text understanding | 2024.07 | https://github.com/longvideobench/LongVideoBench |
| StreamingBench [881] | Multimodal Understanding | Benchmark | Streaming-video understanding | 2024.11 | https://github.com/THUNLP-MT/StreamingBench |
| DVD [114] | Multimodal Understanding | System | Tool-using video deep research | 2025.05 | https://github.com/microsoft/DeepVideoDiscovery |
| VideoSeek [115] | Multimodal Understanding | System | Tool-augmented active perception in long video | 2026.03 | https://github.com/jylins/videoseek |
| GenAI-Bench [882] | Multimodal Generation | Benchmark | Text-to-visual generation evaluation | 2024.06 | https://github.com/TIGER-AI-Lab/GenAI-Bench |
| MM-StoryAgent [799] | Multimodal Generation | System | Text–image–audio story generation | 2025.03 | https://github.com/X-PLUG/MM_StoryAgent |
| MUSE [795] | Multimodal Generation | System | Plan–execute–verify story-video generation | 2026.02 | https://github.com/sunwenzhang1996/MUSE |
| Ptah [800] | Multimodal Generation | System | Verifier-guided multimodal report generation | 2026.05 | https://github.com/SnowNation101/Ptah |
| CrayOtter [796] | Multimodal Generation | System | Verification-driven multimodal generation | 2026.06 | https://github.com/idwts/Crayotter |
| OmniVideoBench [883] | Omnimodal Agency | Benchmark | Audio-visual evaluation | 2025.10 | https://github.com/NJU-LINK/OmniVideoBench |
| Agent-Omni [116] | Omnimodal Agency | System | Omnimodal reasoning agent | 2025.11 | https://github.com/huawei-lin/Agent-Omni |
| OmniGAIA [117] | Omnimodal Agency | Benchmark | Cross-modal tool-use reasoning | 2026.02 | https://github.com/RUC-NLPIR/OmniGAIA |
| AgentVista [775] | Omnimodal Agency | Benchmark | Multimodal agent evaluation | 2026.02 | https://github.com/hkust-nlp/AgentVista |
| Orchestra-o1 [803] | Omnimodal Agency | System | Cross-modal orchestration agent | 2026.06 | https://github.com/zfkarl/Orchestra-o1 |
| General-Purpose Agents | |||||
| ToolBench [51] | Personal Assistants | Benchmark | Large-scale tool use without env state | 2023.07 | https://github.com/OpenBMB/ToolBench |
| Open Interpreter9 | Personal Assistants | System | Local computer and code execution | 2023.07 | https://github.com/OpenInterpreter/open-interpreter |
| AgentBench [119] | Personal Assistants | Benchmark | Multi-environment agent evaluation | 2023.08 | https://github.com/THUDM/AgentBench |
| GAIA [3] | Personal Assistants | Benchmark | General assistant tasks | 2023.11 | https://github.com/UKGovernmentBEIS/inspect_evals/tree/main/src/inspect_evals/gaia |
| -bench [120] | Personal Assistants | Benchmark | Stateful tool–agent–user interaction | 2024.06 | https://github.com/sierra-research/tau-bench |
| AppWorld [534] | Personal Assistants | Benchmark | Stateful app-based personal tasks | 2024.07 | https://github.com/StonyBrookNLP/appworld |
| BFCL [50] | Personal Assistants | Benchmark | Function-calling leaderboard | 2024.08 | https://github.com/ShishirPatil/gorilla |
| OpenManus10 | Personal Assistants | System | General-purpose agent workflows | 2025.03 | https://github.com/FoundationAgents/OpenManus |
| MCP-Mark [240] | Personal Assistants | Benchmark | Realistic MCP tool-use stress test | 2025.09 | https://github.com/eval-sys/mcpmark |
| Toolathlon [239] | Personal Assistants | Benchmark | Diverse realistic tool environments | 2025.10 | https://github.com/hkust-nlp/Toolathlon |
| MCP-Atlas [309] | Personal Assistants | Benchmark | Large-scale tool-use competency | 2026.02 | https://github.com/scaleapi/mcp-atlas |
| Claw-Eval [241] | Personal Assistants | Benchmark | Trustworthy autonomous-agent evaluation | 2026.04 | https://github.com/claw-eval/claw-eval |
| ClawMark [242] | Personal Assistants | Benchmark | Multi-day multimodal living-world eval | 2026.04 | https://github.com/evolvent-ai/clawmark |
| Agents’ Last Exam [238] | Personal Assistants | Benchmark | Frontier professional-workflow evaluation | 2026.06 | https://github.com/rdi-berkeley/agents-last-exam |
| VLN-CE [884] | Embodied & World Models | Benchmark | Continuous vision-language navigation | 2020.04 | https://github.com/jacobkrantz/VLN-CE |
| OpenVLA [129] | Embodied & World Models | System | Open vision-language-action model | 2024.06 | https://github.com/openvla/openvla |
| DINO-WM [828] | Embodied & World Models | System | Feature-space world model | 2024.11 | https://github.com/gaoyuezhou/dino_wm |
| EmbodiedBench [885] | Embodied & World Models | Benchmark | Embodied task evaluation suite | 2025.02 | https://github.com/EmbodiedBench/EmbodiedBench |
| GR00T N1 [823] | Embodied & World Models | System | Humanoid vision-language-action model | 2025.03 | https://github.com/NVIDIA/Isaac-GR00T |
| [824] | Embodied & World Models | System | Open-world VLA model | 2025.04 | https://github.com/Physical-Intelligence/openpi |
| V-JEPA 2 [533] | Embodied & World Models | System | Latent video-prediction world model | 2025.06 | https://github.com/facebookresearch/vjepa2 |
| Cosmos-Predict2.5 [520] | Embodied & World Models | System | Video world simulation | 2025.10 | https://github.com/nvidia-cosmos/cosmos-predict2.5 |
| ChemCrow [886] | Productive Agents | System | Tool-augmented chemistry reasoning | 2023.04 | https://github.com/ur-whitelab/chemcrow-public |
| FinRobot [832] | Productive Agents | System | Financial research and analysis | 2024.05 | https://github.com/AI4Finance-Foundation/FinRobot |
| MLE-bench [121] | Productive Agents | Benchmark | Machine-learning engineering tasks | 2024.10 | https://github.com/openai/mle-bench |
| LegalAgentBench [835] | Productive Agents | Benchmark | Multi-step legal tasks | 2024.12 | https://github.com/CSHaitao/LegalAgentBench |
| MedAgentBench [841] | Productive Agents | Benchmark | Clinical EHR workflows | 2025.01 | https://github.com/stanfordmlgroup/MedAgentBench |
| METR [10,145] | Productive Agents | Benchmark | Frontier time-horizon measurement | 2025.03 | https://github.com/METR/eval-analysis-public |
| AI Scientist-v2 [852] | Productive Agents | System | Autonomous scientific discovery | 2025.04 | https://github.com/SakanaAI/AI-Scientist-v2 |
| PaperBench [122] | Productive Agents | Benchmark | End-to-end research-paper replication | 2025.04 | https://github.com/openai/preparedness |
| AlphaEvolve [126] | Productive Agents | System | Algorithmic discovery agent | 2025.05 | https://github.com/google-deepmind/alphaevolve_results |
| Finance Agent Benchmark [833] | Productive Agents | Benchmark | Real-world financial analysis tasks | 2025.08 | https://github.com/vals-ai/finance-agent |
| GDPval [846] | Productive Agents | Benchmark | Real-world economic task value | 2025.10 | https://github.com/UKGovernmentBEIS/inspect_evals/tree/main/src/inspect_evals/gdpval |
| Kosmos [853] | Productive Agents | System | Cross-domain discovery agent | 2025.11 | https://github.com/EdisonScientific/kosmos-figures |
| APEX-Agents [809] | Productive Agents | Benchmark | Office multi-step task execution | 2026.01 | https://github.com/Mercor-Intelligence/archipelago |
| OfficeQA Pro [243] | Productive Agents | Benchmark | Enterprise grounded reasoning | 2026.03 | https://github.com/databricks/officeqa |
| PresentBench [810] | Productive Agents | Benchmark | Rubric-based slide generation | 2026.03 | https://github.com/PresentBench/PresentBench |
| OneMillion Bench [845] | Productive Agents | Benchmark | Human-level economic tasks | 2026.03 | https://github.com/humanlaya/OneMillion-Bench |
| Workspace-Bench [745] | Productive Agents | Benchmark | Workspace file-dependency tasks | 2026.05 | https://github.com/OpenDataBox/Workspace-Bench |
| YC-Bench [887] | Productive Agents | Benchmark | Long-term planning consistency | 2026.06 | https://github.com/collinear-ai/yc-bench |
| Frontier | Open challenges | Outlook |
|---|---|---|
| I. Evolution — generalize across runtimes and keep learning over time | ||
| Self-evolving Harness & Agents | • Hand-set optimization metric• Limited gains on out-of-distribution tasks• Overfitting and drift over long runs | • Adaptively choose optimization objective and scope• Autonomous evolution guided by feedback• Lifelong, non-overfitting evolution |
| Harness Generalization and Transferability | • Models tied to specific harnesses• Performance varies across model–harness pairings• Portable standards still ad hoc | • Multi-harness (hybrid-harness) training• Portable skills / experience at inference time• A standard harness protocol |
| Continual & Lifelong Learning | • External experience is noisy and unverified• Catastrophic forgetting during weight updates | • Learning only from validated experience• Safe online learning from billion-user logs |
| II. Effectiveness — act reliably in realistic, verifiable environments | ||
| Real-world Environment Interaction | • Slow, costly, and irreversible interactions with real-world systems• Synthesized Envs & world-models may be unfaithful | • Measure, audit, and optimize sim-to-real faithfulness• Fusion of symbolic and neural simulation environments |
| From Digital to Embodied Agents | • Deep planning vs. real-time control• Limited modeling of physical dynamics & laws• Coarse vs. fine-grained feedback | • Async, hierarchical harness (deliberation vs. reflex)• Predict physical outcomes with world models• Integrate feedback across timescales |
| III. Efficiency — spend compute, context, and modality budgets | ||
| Cost- & Budget-aware Agency | • Model selection is poorly matched to task difficulty• No calibrated cost / infeasibility sense• No runtime enforcement of ceilings• No law linking budget to success | • Perceive cost, difficulty, infeasibility• Act mid-run: shorten, re-route, or stop• Harness-enforced ceilings & cost-aware routing• A budget-vs-success scaling law |
| Multimodal & Omni Harness | • Multimodality bolted on; no vision-native harness• Visual-token budgeting is heuristic• Cross-modal verification is unreliable | • A vision-native, unified omni harness substrate• Per-modality granularity control• Cross-modal verification & omni-modal routing |
| IV. Trustworthiness — stay robust and governable as the horizon grows | ||
| Reflection & Error Robustness | • Failures detected late under noisy, delayed feedback• Intrinsic self-correction is unreliable• Errors compound into goal drift | • Early failure detection & path-switching• External-anchored, calibrated self-knowledge• Align to human course-correction |
| Safety & Governance | • Injected error & hazardous experience reuse• No unified safety-verification standard• Self-evolution erodes invariants; biased self-checks | • Untrusted external input; gate experience reuse• Score action safety beside effectiveness• Independent, invariant-preserving verification |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
