Submitted:
18 September 2026
Posted:
20 September 2026
You are already at the latest version
Abstract
Large language model (LLM) agents can learn through interaction with external environments, obtain training signals and reusable experience from task execution. Scaling this form of learning requires a sufficient supply of interactive environments and tasks grounded in those environments, with explicit objectives and completion criteria. We refer to these complementary resources as interactive learning resources. This survey reviews their automated construction at scale and continuous adaptation through three interconnected themes: environment synthesis, task synthesis, and learning-driven evolution of both. For environment synthesis, we examine functional components, construction paradigms, and intrinsic quality assurance. For task synthesis, we review intent formation, grounding, completion verification, and task quality assurance. We then examine how a target agent's performance and learning progress inform assessments of resource suitability and guide adaptation to its evolving capabilities and learning needs across successive learning rounds. Finally, we discuss open challenges concerning the order of environment and task synthesis, the trade-off between intrinsic quality and learning utility, long-term resource evolution, and quality assurance at scale.
Keywords:
LLM agents
; environment synthesis
; task synthesis
1. Introduction
Large language model (LLM) agents extend language models from passive text generation to systems that can perceive, reason, and act through interactions with external environments. As agents are deployed for complex interactive tasks, enabling them to improve through experience has become a central research problem. Interaction allows an agent to attempt a task, observe the consequences of its actions, and obtain evidence about its execution. Such evidence provides two complementary sources of improvement: (i) execution outcomes can supply training signals, such as rewards for reinforcement learning; and (ii) trajectories preserve successful strategies, failures, and intermediate decisions that can be distilled into reusable experience to support self-improvement(Chen et al. 2026b; Dong et al. 2026; Guo et al. 2026b).
Interactive Learning Resources Learning through interaction requires both environments and tasks. Environments provide the states, dynamics, interfaces, and runtimes in which agents can act; tasks specify what agents should accomplish, under which concrete conditions, and how completion is assessed. As shown in Figure 1, we use interactive learning resources as an umbrella term for these two complementary resource types. An environment can support many tasks, and task synthesis can reuse an existing environment or accompany the construction of a new one. Their shared role in supporting goal-directed practice motivates studying their construction together.
Scalable Resource Construction Constructing interactive learning resources at scale through manual effort remains challenging. Environments are costly to build, deploy, reset, and maintain, while manually authored tasks offer limited coverage of objectives, states, and workflows. These constraints motivate large-scale automated synthesis across tool use, web navigation, GUI interaction, terminal interaction, and software engineering (Hu et al. 2025; Song et al. 2026; Sullivan et al. 2025; Tu et al. 2026; Xi et al. 2026). We organize environment synthesis around three primary paradigms for constructing reusable interaction substrates: model-based simulation, reconstruction and adaptation of existing systems, and creation from scratch. We structure task synthesis around three core components: task intent synthesis, environment grounding and task instantiation, and task completion verification. Together, these construction approaches expand the availability and diversity of interactive learning resources, providing broader opportunities for agents to develop capabilities across varied settings and workflows.
However, expanding resource supply is useful only if quality can be maintained as construction scales. Unlike small collections that can be individually curated and inspected, large-scale automated synthesis cannot depend on exhaustive manual checking. It requires scalable mechanisms to ensure that environments remain executable, semantically coherent, and robust, that tasks are grounded and solvable, and that the resulting collections provide sufficient diversity, structural complexity, and fidelity. To this end, execution-based testing, consistency checking, filtering, and repair therefore become integral parts of synthesis pipelines rather than optional checks after generation (Shi et al. 2026a; Tu et al. 2026; Zhang et al. 2026a). Accordingly, Section 3 and Section 4 examine both the technical routes for automated construction and the quality requirements that synthesized environments and tasks must satisfy.
Learning-Driven Resource Evolution Construction does not end with an initial pool of valid environments and tasks. As a target agent learns through interaction, its changing capabilities create new requirements for what should be synthesized next. This introduces a learner-dependent dimension of resource assessment: beyond whether environments and tasks satisfy their intrinsic quality requirements, the question becomes how well they match the agent’s current capabilities and learning needs. Success rates, failure patterns, and learning progress can help identify tasks that are already mastered or remain too difficult, capabilities that receive insufficient practice, and limitations in what the available environments allow the agent to learn (Dong et al. 2026; Huang et al. 2026a; Kang et al. 2026; Meng et al. 2026). The agent’s interaction evidence thus provides a basis for assessing which resources should be retained, revised, or expanded.
These assessments can guide subsequent synthesis by generating new objectives, adjusting task distributions, modifying environment states, tools, or dynamics, and constructing additional environments. Repeating interaction, assessment, adaptation, and learning allows resource construction to respond to the learner’s evolving needs rather than only expanding a fixed collection. The technical challenge is to translate agent feedback into appropriate construction decisions while preserving resource quality and broad coverage across successive rounds. Section 5 examines this learning-driven evolution of tasks and environments, in which agent-relative assessment provides signals for deciding what to construct or revise next. Such signals guide adaptation, while their value ultimately depends on whether the resulting resources support further agent improvement.
Survey Overview Building on this perspective, this survey provides a structured account of how interactive learning resources are constructed, assessed, and continually adapted. Section 3 reviews environment synthesis through its functional components, major construction paradigms, and intrinsic quality requirements. Section 4 examines task synthesis from intent formation and environment grounding to task-completion verification and task quality assurance. Section 5 then extends the analysis beyond initial construction to learning-driven resource evolution, studying how evidence from a target agent supports capability estimation, agent-relative resource assessment, and subsequent task and environment adaptation. In Section 6, we finally discuss open challenges concerning synthesis order, quality assessment, learning utility, long-term evolution, and quality assurance at scale.
Relation to Existing Surveys Li et al. (2026b) study agentic environment engineering through environment modeling, synthesis, evaluation, and applications, including task synthesis and evolution within that broader lifecycle. Zeng et al. (2026) study agentic data generation, organizing environments, task signals, trajectories, and verifiers through the Accuracy–Complexity–divErsity (ACE) lens. Complementary to these perspectives, we focus on the technical routes and requirements for the large-scale automated construction of interactive learning resources. We organize environment synthesis and task synthesis as two complementary subjects, examine the quality assurance needed to make their outputs usable at scale, and extend the analysis to continuing resource evolution informed by agent learning. Detailed agent-learning and self-improvement algorithms are covered by complementary surveys (Gao et al. 2026; Zhang et al. 2026b).
2. Preliminaries
This section defines interactive learning resources, distinguishes their construction from their execution, and clarifies the assessment roles and scope of the survey.
Interactive Learning Resources We use interactive learning resources to refer to the environments and tasks that support LLM agents in learning through goal-directed interaction. Environments supply the conditions and capabilities for interaction, while tasks provide objectives, execution conditions, and completion criteria. They can be constructed separately or jointly and reused across multiple executions. A single environment can support many tasks; a task intent can also be instantiated in different environments, provided its objects, parameters, and requirements are grounded in each one.
Environment An environment is the interactive substrate within which an agent acts and observes the consequences of its actions. It maintains the state relevant to interaction, determines how that state changes, exposes actions and observations through agent-facing interfaces, and provides the runtime in which interaction is executed. It determines what can happen. For a compact formal view, we write
where S is the state space, A the available actions, O the observations exposed to the agent, P the transition dynamics, and R the runtime substrate.
Task A task specifies what an agent should accomplish within a concrete environment. We write
where g is the objective, the initial condition, c the grounded parameters and constraints, and v the intended completion semantics. For the feedback-driven learning considered here, task construction must also supply or identify a procedure V that operationalizes these semantics. This procedure may be generated with the task, instantiated from a template, or inherited from an existing environment.
Resource Execution and Verification Executing a task T in environment E with policy produces an interaction trajectory x, which a verifier assesses:
Here, x records the observations, actions, and available execution evidence, and y is a completion judgment such as a success label, score, or step-wise assessment. The task semantics v specify what completion means, whereas V implements its assessment. The resulting trajectories and feedback are products of using the resources, rather than the resources themselves. They may supply training rewards, experience-selection feedback, or evidence for subsequent resource adaptation.
Scope This survey focuses on the large-scale automated synthesis of environments and tasks for LLM agent learning and evolution. We include methods that create interactive environments, reconstruct or adapt existing systems, simulate environment behavior, or generate task intents, grounding, execution conditions, and completion checks. We also examine how target-agent evidence guides the assessment and adaptation of these resources over learning rounds, with environment-centric evolution schedules providing a complementary point of comparison.
Our primary objects are the reusable environments and tasks made available for interaction. Trajectory synthesis and rollout collection are included when they contribute mechanisms for constructing these resources. Hand-crafted benchmarks and static instruction collections provide adjacent context unless they introduce scalable construction mechanisms. Literature collection is summarized in Appendix A, and Appendix B provides the literature mappings.
3. Environment Synthesis
Environment synthesis concerns the automatic and scalable construction of reusable environments that can sustain continued agent interaction. This section presents their functional components and synthesis targets, organizes construction paradigms by how the interaction substrate is obtained, and examines intrinsic environment quality.
3.1. Environment Components
From the perspective of agent interaction, an environment needs to specify what the world currently contains, how it can evolve, what the agent can observe and manipulate, and how such interaction is executed. Accordingly, we characterize an environment through four components:
- Environment State represents the persistent configuration of the world at a given point in interaction, including the entities, resources, relations, and application states relevant to future actions. It should preserve the consequences of previous interactions so that later decisions are conditioned on the evolving world state.
- Environment Dynamics defines how the world state can evolve, including the rules, constraints, and side effects that govern state transitions. These transitions may be triggered by agent actions as well as by internal processes, external events, or other actors.
- Interaction Tools specifies how the agent observes and operates on the environment through interfaces such as tools, GUIs, browsers, or shells. It determines the agent-facing observation and action space.
- Runtime Infrastructure provides the execution substrate that makes the environment usable in practice. It supports reliable initialization, execution, state persistence, reset, isolation, and repeated interaction at scale.
3.2. Environment Synthesis Paradigms
According to how the underlying interaction substrate is obtained, we categorize existing approaches into three paradigms: (i) Model-based Environment Simulation, which replaces a fully implemented backend with a generative model that produces environment responses and transitions during interaction; (ii) Existing-System Reconstruction and Adaptation, which converts existing systems, artifacts, or interaction traces into reusable environments; and (iii) Environment Creation from Scratch, which constructs a new executable environment from seeds or specifications.
Paradigm 1: Model-based Environment Simulation This paradigm obtains an interactive environment by configuring a generative model as its response and transition mechanism:
where M is the generative simulator, z denotes optional environment specifications such as tool descriptions or API schemas, and configures the model with the interfaces and runtime needed for repeated interaction. During execution, M produces responses or predicts state transitions from the interaction history and the current action. The simulated state may remain implicit rather than being exposed as an accessible state object. Two common, non-exclusive implementation choices are specification grounding, which conditions simulation on tool descriptions, API schemas, or task context (Lee et al. 2026; Li et al. 2025), and learning from interaction data, which distills response behavior or transition dynamics (Chen et al. 2026b; Ding et al. 2026; Xu et al. 2026b). A simulator can combine both: MirrorAPI uses API descriptions together with a specialized model trained on request–response examples (Guo et al. 2025). Hybrid systems can additionally combine live execution with synthesized world engines (Xi et al. 2026). This paradigm scales interaction without implementing every state transition, but its black-box transitions are difficult to inspect and can accumulate semantic or state inconsistencies over long horizons. These limitations make deterministic execution, debugging, and reliable reward attribution harder than in explicitly executable environments.
Paradigm2: Existing-System Reconstruction and Adaptation This paradigm obtains an interactive environment by reconstructing or adapting semantics already present in software systems, applications, repositories, or prior interactions:
where B denotes an existing system or its artifacts, H denotes optional execution histories, and reconstructs or adapts these sources into a reusable environment. The resulting environment inherits or recovers interaction semantics from an existing source. One family focuses on runtime adaptation, resolving dependencies, containerizing software, or exposing authentic services through stable interfaces (Guo et al. 2026a; Xu et al. 2026a). TerminalTraj also constructs Dockerized environments from repositories before generating environment-aligned tasks and trajectories (Wu et al. 2026b). A second family performs system reconstruction, reproducing existing applications or websites as lightweight, resettable, and verifiable replicas (Cao et al. 2026; Chae et al. 2026). More recent work reconstructs executable workspaces from environment histories or trajectories: CLI-Gym inverts healthy environment histories into failure states, while Terminal-Universe recovers reusable terminal environments from accumulated tool-execution traces (Lin et al. 2026; Wu et al. 2026a). These routes inherit useful semantics and fidelity from their sources, but their coverage remains constrained by the systems or traces available for reconstruction and by the cost of adapting heterogeneous dependencies.
Paradigm3: Environment Creation from Scratch This paradigm obtains an interactive environment by constructing a new implementation from a seed or specification:
where z denotes a scenario description, capability specification, schema, or other design input, and materializes the environment’s state, dynamics, interfaces, and runtime. Recent systems span planning worlds, tool/API ecosystems, Web applications, terminal containers, software repositories, and computer-use environments (Cai et al. 2025; Fang et al. 2026; Hu et al. 2025; Sullivan et al. 2025; Tu et al. 2026; Wang et al. 2026a; Wu et al. 2026d; Zhu et al. 2025). A complete synthesis pipeline typically contains five stages.
Seed and environment specification. The intended structure and functionality of the environment can be anchored by semantic seeds such as domains or business scenarios, capability-grounded seeds such as schemas, protocols, tool ecosystems, or skill structures, and reference-grounded seeds such as source applications, code artifacts, or interface designs (Cai et al. 2025; Hu et al. 2025; Jeong and Yoon 2026; Shi et al. 2026b; Song et al. 2026; Zhang et al. 2026c). The choice of seed controls both the freedom of synthesis and how strongly later states, dynamics, and interfaces inherit externally defined semantics.
State construction and persistence. The specification must then be instantiated as explicit state that can be read, modified, persisted, and reset. Structured databases, object stores, application backends, or generated project files provide an external source of truth rather than leaving state implicit in model context (Fang et al. 2026; Song et al. 2026; Tu et al. 2026; Wang et al. 2026c; Xi et al. 2026; Zhu et al. 2025). State lifecycle mechanisms further support inspection, isolation, and repeated episodes, as in task-specific state injection or session-scoped state management (Wang et al. 2026a).
Dynamics construction. Executable environments realize dynamics by translating preconditions, update rules, side effects, and constraints into programmatic operations over state. Local transitions may be implemented as functions or database transactions, while interconnected environments require shared-state constraints and cross-service invariants (Jeong and Yoon 2026; Song et al. 2026; Tu et al. 2026; Wang et al. 2026c; Wu et al. 2026d). Programmatic state machines and controllers provide an alternative way to make transition structure explicit and verifiable (Wu et al. 2026d; Xi et al. 2026).
Tool and interface synthesis. Internal capabilities must be exposed through actions that the agent can invoke. Methods range from synthesizing tool schemas and implementations, to binding APIs to persistent backends, to inheriting protocol constraints from existing tool ecosystems (Castellani et al. 2025; Fang et al. 2026; Shi et al. 2026b; Sullivan et al. 2025; Wang et al. 2026c). The common objective is to align declared interfaces with meaningful, state-changing environment behavior.
Runtime packaging. Finally, state, dynamics, and interfaces are assembled into runtimes that support initialization, isolation, reset, and repeated execution. Terminal and software environments typically use containers or sandboxes, whereas Web environments package self-contained frontends and backends (Gandhi et al. 2026; Peng et al. 2026; Shen et al. 2026; Wu et al. 2026d; Zhang et al. 2026c; Zhu et al. 2026; Zhu et al. 2025).
3.3. Environment Quality Assurance
Quality assurance assesses the environment itself and the environment collection as reusable learning resources. A large collection can remain unreliable or limited if instances fail to execute, implement inconsistent transitions, remain highly redundant, support only trivial interactions, behave unreliably, or deviate substantially from the intended setting. We summarize six intrinsic dimensions: executability, semantic correctness, diversity, complexity, robustness, and fidelity.
Executability Executability asks whether a generated environment can be successfully built, launched, interacted with, and reset. Existing pipelines combine static checks, runtime testing, executable probing, and repair loops before environments are admitted for training (Cai et al. 2025; Guo et al. 2026a; Shen et al. 2026; Shi et al. 2026a; Tu et al. 2026; Wu et al. 2026d; Zhu et al. 2026).
Semantic Correctness Semantic correctness requires environment behavior to conform to intended interface, transition, and system-level semantics. Assurance mechanisms include specification–implementation auditing, transition probing, explicit invariants, finite-state constraints, and backend consistency checks (Castellani et al. 2025; Jeong and Yoon 2026; Song et al. 2026; Tu et al. 2026; Wu et al. 2026d; Zhang et al. 2026a).
Diversity Diversity measures whether synthesis broadens the available environments and interaction capabilities rather than producing superficial variants. Methods diversify semantic domains, tool/API structures, state schemas, workflow compositions, and whole environment ecosystems, often with explicit coverage objectives or open-ended discovery (Castellani et al. 2025; Dong et al. 2026; Fang et al. 2026; Song et al. 2026; Sullivan et al. 2025; Tu et al. 2026).
Complexity Complexity characterizes the structural dependencies an environment can support, including action-level composition, persistent state coupling, cross-service dependencies, and navigational or programmatic structure (Jeong and Yoon 2026; Shi et al. 2026b; Sullivan et al. 2025; Xi et al. 2026; Zhang et al. 2026c).
Robustness Robustness concerns whether an environment remains stable and recoverable across repeated and concurrent interaction. Isolation, deterministic reset, snapshots, session-scoped state, monitoring, and incremental repair reduce contamination and infrastructure-induced variation across rollouts (Cao et al. 2026; Guo et al. 2026a; Shen et al. 2026; Wang et al. 2026a).
Fidelity Fidelity concerns how faithfully a synthetic environment preserves the properties of the systems or scenarios it is intended to represent. Grounding may occur at protocol/schema level, through direct system or website reconstruction, or through authentic resources and reference designs (Cao et al. 2026; Chae et al. 2026; Shi et al. 2026b; Xu et al. 2026a; Zhang et al. 2026c). Recent evidence further shows that structural and semantic defects in synthetic worlds can materially reduce downstream learning utility, motivating explicit verification rather than relying on surface realism alone (Zhang et al. 2026a).
4. Task Synthesis
Task synthesis constructs the objectives and execution specifications through which agents use environments for goal-directed learning. We first discuss Task Intent Synthesis, which forms objectives, and then Environment Grounding and Task Instantiation, which binds them to concrete states, resources, parameters, and runtime conditions. Task Verification supplies procedures that assess whether an execution meets the objective and constraints, making completion checks part of task construction. Task Quality Assurance instead evaluates the resulting tasks and task collection as resources.
4.1. Task Intent Synthesis
Task intent synthesis forms the objective that the agent is expected to accomplish. We separate intent formation from full task instantiation for analysis. We organize existing approaches according to the primary source from which the task intent is derived, and discuss seed-driven, tool-driven, state-driven, and trajectory-driven methods in turn.
Seed-driven Seed-driven methods start from externally provided descriptions of what kinds of tasks should be generated, such as seed instructions, capability requirements, domain descriptions, personas, or a small number of exemplars. Rather than only paraphrasing seeds, recent methods structure them into task profiles, hierarchical compositions, or difficulty-controlled design spaces (Hua et al. 2026; Pan et al. 2026; Pi et al. 2026; Shi et al. 2025; Xie et al. 2026; Zhao et al. 2026). This route offers direct control over the desired task distribution, but feasibility must still be established when the intent is grounded into a concrete environment.
Tool-driven Tool-driven methods derive task intents from the affordance structure exposed by an environment. The underlying action space may be organized as a flat tool collection, a dependency graph, a workflow/program structure, or a skill graph, and task synthesis composes objectives over these structures. Representative mechanisms include tool-sequence exploration, protocol or dependency graphs, multi-skill composition, skill-graph path sampling, and topology-aware tool sampling (Cheng et al. 2026; Dong et al. 2026; Fan et al. 2026a; Keren et al. 2026; Shi et al. 2026b; Tan et al. 2026; Tu et al. 2026; Xu et al. 2026a). The defining signal is therefore not merely which tools exist, but how their capabilities can be composed into meaningful workflows.
State-driven State-driven methods formulate task intents from a concrete environment state. Once an environment has been instantiated with specific entities, resources, records, or configurations, the generator uses this state to determine objectives that are meaningful and feasible under current conditions. Database- and backend-grounded systems generate scenarios from instantiated state and available operations, thereby avoiding nonexistent entities or violated preconditions (Song et al. 2026; Tu et al. 2026; Wang et al. 2026c; Xi et al. 2026).
Trajectory-driven Trajectory-driven methods obtain task intents from interaction traces that have already been executed in the environment. They first explore or execute the environment, observe a path of actions and state transitions, and then abstract or reverse-engineer a task from this evidence (Chen et al. 2026a; Murty et al. 2024; Ramrakhya et al. 2026; Sun et al. 2025; Wang et al. 2026d). Terminal-Universe extends this idea further by reconstructing a reusable environment from a trajectory before exploring the recovered workspace for new tasks (Wu et al. 2026a). Because the proposed intent is supported by execution evidence, trajectory-driven synthesis offers strong grounding, although its coverage is ultimately constrained by what the exploration or trajectory source exposes.
4.2. Environment Grounding and Task Instantiation
Once a task intent has been formed, it must be instantiated with concrete objects, parameters, and initial conditions that actually exist in the environment. This grounding process is strongly shaped by the underlying interaction substrate. We therefore organize existing methods into two broad settings: (i) Web/Service Environments, where tasks are grounded in websites, online services, and backend states; and (ii) Computer/Software Environments, where tasks are instantiated over applications, filesystems, repositories, and executable workspaces. Within each setting, different interfaces—such as browser actions, APIs, GUIs, or command lines—further determine how the task is concretely instantiated.
Web/Service Environments For tools and services, grounding mainly requires selecting executable capabilities, resolving valid entities and arguments, and satisfying dependencies among successive calls. These constraints are often derived from schemas, tool graphs, backend state, or explicitly synthesized transition structures (Chen et al. 2026a; Fang et al. 2026; Shi et al. 2026b; Song et al. 2026; Tu et al. 2026; Wang et al. 2026c; Xu et al. 2026a). Browser-based Web environments additionally require tasks to align with reachable pages, sessions, navigation structure, and persistent backend state (Chae et al. 2026; Huang et al. 2026b; Murty et al. 2024; Wang et al. 2026b; Wu et al. 2026d). Across both cases, grounding binds abstract objectives to resources and interaction paths that the environment can actually support.
Computer/Software Environments Computer and software tasks require a concrete executable workspace around the intended objective. GUI settings instantiate applications, files, user data, and task-specific application states, while terminal and software-engineering settings materialize repositories, dependencies, services, containers, and tests (Hua et al. 2026; Lin et al. 2026; Lv et al. 2026; Ramrakhya et al. 2026; Shi et al. 2026a; Wang et al. 2026a; Wu et al. 2026b; Zhao et al. 2026). More recent pipelines generate or reconstruct complete workspaces and their verifiers from scratch, from source skills, or from prior trajectories (Du et al. 2026; Gandhi et al. 2026; Li et al. 2026c; Pan et al. 2026; Shen et al. 2026; Wu et al. 2026a; Zhu et al. 2026; Zhu et al. 2025). Grounding ultimately requires aligning the task with a realizable software state and a reproducible execution context.
4.3. Task Verification
Task completion verification assesses a particular execution against the intended objective and constraints. We characterize verification along three complementary design dimensions: when verification is performed, what evidence is inspected, and how the final judgment is made.
Formally, we characterize a verifier as
where t specifies when verification is performed, e specifies what evidence is inspected, and j specifies how that evidence is converted into a judgment y such as a success label, score, or reward.
Final Outcome vs. Step-wise Verification Final verification determines success from the outcome observed after an interaction terminates and is appropriate when task semantics are fully captured by the final state or output (Hua et al. 2026; Lei et al. 2026; Tu et al. 2026; Wang et al. 2026a). Step-wise verification instead evaluates selected intermediate conditions when correctness depends on mandatory checkpoints, ordering constraints, forbidden operations, or side effects (Shi et al. 2026a; Song et al. 2026; Wu et al. 2026d). It provides denser supervision but requires more explicit specification of which intermediate behaviors matter.
Evidence Sources and Checking MechanismsState-based verification checks structured environment states through predicates, state differences, invariants, or target-state matching (Chae et al. 2026; Lei et al. 2026; Tu et al. 2026; Wang et al. 2026a). Path-based verification examines action sequences and intermediate states when correctness depends on how an outcome is reached (Chen et al. 2026a; Wang et al. 2026b). Test-based verification instead describes an executable checking mechanism: unit tests, integration tests, compilers, or shell scripts may inspect states, outputs, paths, or several of these together. Such checking is particularly common in terminal and software environments (Gandhi et al. 2026; Hua et al. 2026; Shen et al. 2026; Shi et al. 2026a; Wu et al. 2026b; Zhu et al. 2026; Zhu et al. 2025).
Deterministic and Rubric-based JudgmentDeterministic verification uses rules, assertions, exact matching, state predicates, or executable tests to produce reproducible judgments, whereas rubric-based verification uses LLM or multimodal judges for objectives that are difficult to formalize programmatically (Cao et al. 2026; Hua et al. 2026; Lei et al. 2026; Lv et al. 2026; Tan et al. 2026; Xi et al. 2026; Xue et al. 2026). The former offers stronger auditability; the latter extends coverage to open-ended semantics but introduces calibration, consistency, and reward-exploitation risks.
Verifier Reliability and Downstream Uses A verifier should faithfully reflect the intended task semantics, remain robust to alternative valid solutions, and avoid exploitable mismatches between the objective and checking procedure. Reliability failures arise when rules or tests capture only part of the intended behavior, when path constraints reject valid alternatives, or when learned judges are inconsistent (Bercovich 2026; Chen et al. 2026a; Shi et al. 2026a; Wang et al. 2026b). Large-scale synthesis therefore increasingly couples verifier generation with executable validation, adversarial checking, or consistency auditing(Meng et al. 2026; Shen et al. 2026; Tu et al. 2026; Zhang et al. 2026a).
A verifier should provide machine-usable assessments of task execution, such as success/failure labels, scalar scores, or step-wise judgments, optionally accompanied by diagnostic feedback on unmet criteria. These outputs can be used directly or transformed into (i) reward signals for reinforcement learning; (ii) feedback and selection signals for agent self-evolution, supporting trajectory selection, experience extraction, and the evaluation of candidate memory, skill, or workflow updates; and (iii) capability- and resource-assessment signals, when aggregated with interaction evidence to assess difficulty, capability gaps, and learning suitability for the target agent. The third use guides task and environment adaptation in Section 5; an execution score is not itself a score of the task’s intrinsic quality or learning utility. For detailed agent-side learning and update mechanisms, we refer readers to surveys on agentic reinforcement learning (Zhang et al. 2026b) and self-evolving agents (Gao et al. 2026).
4.4. Task Quality Assurance
Task quality assurance evaluates the generated task and task collection, rather than judging a particular execution. It asks whether a valid solution exists under the specified conditions, whether the collection covers diverse objectives and workflows, how structurally complex its tasks are, and how faithfully they reflect the intended uses. We organize these intrinsic properties as Solvability, Diversity, Complexity, and Fidelity.
Solvability Solvability requires that at least one valid solution exists under the given environment and initial conditions. Evidence can come from reachability analysis, reference solutions, successful rollouts, golden-state construction, and solver-based calibration (Lei et al. 2026; Li et al. 2026c; Meng et al. 2026; Pan et al. 2026; Shen et al. 2026; Shi et al. 2026a; Tu et al. 2026). An executable test suite alone does not establish solvability: it must be paired with a valid witness or other feasibility evidence, and a solver failure is not a proof that no solution exists.
Diversity Diversity concerns whether tasks cover new capabilities, workflows, tool combinations, state contexts, or interaction patterns rather than merely paraphrasing instructions. Methods use semantic deduplication, skill/tool coverage, structured combination sampling, graph/path sampling, persona variation, and coverage-guided generation (Chen et al. 2026a; Fan et al. 2026a; Ivison et al. 2026; Keren et al. 2026; Shi et al. 2025; Tan et al. 2026; Xie et al. 2026).
Complexity Complexity describes structural dependencies within a task, such as subgoal coupling, cross-tool or cross-application dependencies, persistent-state span, dependency depth, branching, delayed consequences, and recovery requirements. Existing methods increase complexity through compositional task expansion, graph-based workflow construction, multi-hop grounding, or recursive task/environment rewriting (Fan et al. 2026a; Huang et al. 2026b; Keren et al. 2026; Li et al. 2026c; Shi et al. 2025; Xie et al. 2026).
Fidelity Task fidelity concerns whether synthetic tasks reflect real user needs, realistic workflows, and the target-domain task distribution. Assurance mechanisms include real-demand discovery, grounding in logs or websites, preserving source intent, authentic-resource exploration, and synthetic-to-real transfer evaluation (Chae et al. 2026; Huang et al. 2026b; Shi et al. 2026a; Xu et al. 2026a; Zhao et al. 2026). Fidelity is alignment between generated objectives and the workflows agents are expected to encounter.
Figure 3.
Development of scalable environment synthesis, task synthesis, and learning-driven resource evolution for LLM agents.
Figure 3.
Development of scalable environment synthesis, task synthesis, and learning-driven resource evolution for LLM agents.

5. Learning-Driven Interactive Learning Resource Evolution
Section 3 and Section 4 examine the construction and intrinsic quality of interactive learning resources. Once a target agent uses these resources, its behavior provides a further basis for evaluation: how well the available environments and tasks match its current capabilities and learning needs. The same interactions can inform both capability estimation, which characterizes the learner, and agent-relative resource assessment, which characterizes the suitability of its practice. The central agent-conditioned loop uses these assessments to adapt environments or tasks, then reassesses them after further interaction and learning. This section focuses on that feedback-driven process.
For agent-conditioned evolution, a compact view is
followed by learning on interactions from the updated resources,
Here, and denote the environment and task collections, and contains the target policy’s interaction evidence and verification outcomes. The estimate characterizes current capability, while summarizes resource assessments such as learner-relative difficulty, failure-prone regions, and coverage of learning needs. The assessments guide adaptation U, while L denotes the subsequent learning update.
Capability Estimation Agent-conditioned evolution estimates the learner’s capabilities from task success rates, learning progress, and contrasts between successful and failed trajectories or heterogeneous solvers. These signals can reveal missing capabilities, identify failure-prone regions, or locate a solver-relative learnable frontier (Dong et al. 2026; Huang et al. 2026a; Kang et al. 2026; Meng et al. 2026). Read from the resource side, the same evidence indicates which tasks are already mastered, which expose current weaknesses, and which environment capabilities remain underused or insufficiently represented in training. Task difficulty is therefore agent-relative, unlike structural complexity or the existence of a valid solution. Agent-relative resource assessment turns these observations into judgments about what the learner should encounter next. Task-level signals can guide the selection or revision of individual objectives, whereas evidence aggregated over tasks and workflows can identify which environments offer suitable practice or need expansion. The purpose is to estimate learning relevance: low success can reflect a capability gap, an invalid environment or task, or an unreliable verifier.
Task and Environment Adaptation Given capability estimates and resource assessments, the synthesis system decides which learning resources should change. Task-side adaptation can generate new objectives, compose additional subgoals, vary conditions, or shift sampling toward the current frontier (Lv et al. 2026; Meng et al. 2026; Wu et al. 2026c; Xue et al. 2026). Environment-side adaptation changes states, tools, dynamics, or whole executable worlds when the current substrate does not expose the capabilities required for further practice (Huang et al. 2026a; Kang et al. 2026; Liu et al. 2026; Shen et al. 2026; Shi et al. 2026c). Related benchmark-construction work such as ProEvolve makes environment changes programmable through graph transformations (Li et al. 2026a); it supplies an adaptation mechanism without itself demonstrating a continual policy-training loop.
Continuous Evolution Repeated interaction and learning change both capability estimates and resource suitability. A task that previously exposed a useful weakness may become routine, while newly available environment capabilities may require additional tasks. Agent-conditioned evolution therefore uses current-policy rollouts, verifier rewards, or capability gaps to regenerate tasks or reshape environments near the moving learning frontier (Dong et al. 2026; Guo et al. 2026b; Huang et al. 2026a; Liu et al. 2026; Shen et al. 2026; Wu et al. 2026c). Resource assessment is repeated within this process. RLVE provides a related precursor: its manually engineered verifiable environments procedurally generate problems and adapt their difficulty distributions to the policy’s evolving capabilities (Zeng et al. 2025).
6. Discussion and Future Directions
The preceding sections show that scaling interactive learning resources involves more than generating larger numbers of environments and tasks. Several broader questions remain unresolved regarding how these artifacts should be constructed, evaluated, evolved, and maintained.
Synthesis Order A fundamental question is how environment, task, and verification should be ordered during construction. Environment-first pipelines provide strong grounding because tasks are generated from capabilities and states that already exist, and one environment can support many tasks (Hu et al. 2025; Lei et al. 2026; Ramrakhya et al. 2026; Shi et al. 2026a; Wang et al. 2026b). Requirement-first approaches instead begin from desired capabilities or task demands and construct the necessary environment afterward (Zhao et al. 2026), while shared-specification approaches derive multiple artifacts from common primitives to improve cross-artifact consistency (Cheng et al. 2026; Ivanov and Rana 2026; Wang et al. 2026a).
Intrinsic Quality vs. Learning UtilitySection 3.3 and Section 4.4 summarize intrinsic quality dimensions for synthesized environments and tasks, while Section 5 examines learner-relative signals used to assess and adapt them. Two questions remain open: (i) how can intrinsic properties and learner-relative suitability be measured accurately and consistently at scale; and (ii) to what extent do improvements in these assessments predict actual learning gains? Existing studies provide initial evidence that environment diversity, executable correctness, and defect repair can materially affect downstream generalization (Sullivan et al. 2025; Tu et al. 2026; Zhang et al. 2026a), but controlled studies that isolate the causal contribution of individual quality dimensions remain limited.
Long-Term Task–Environment EvolutionSection 5 considers how agent feedback can drive repeated task and environment adaptation, but long-term evolution introduces additional control problems. A synthesis system must decide which mastered interactions should be retired or retained, which rare failures remain valuable, when to expand environment capabilities rather than only generate new tasks, and how to combine agent-conditioned evolution with environment-centric exploration without narrowing coverage over time.
Quality Assurance at Scale Quality assurance becomes harder as collections of interactive learning resources grow and change. Updates to states, interfaces, dynamics, or tests can invalidate previously accepted tasks or verifiers, while repeated synthesis can introduce duplication, stale checks, or reward loopholes. Execution-driven repair, auditing, dependency-aware updates, and adversarial verifier checks provide useful building blocks (Bercovich 2026; Castellani et al. 2025; Guo et al. 2026a; Li et al. 2026a; Shi et al. 2026a), but scalable synthesis ultimately requires persistent maintenance.
7. Conclusion
This survey reviews the large-scale automated construction and continuing adaptation of interactive learning resources for LLM agents. Environments supply the conditions for action, while tasks specify objectives, execution conditions, and completion criteria. We organize environment synthesis around components, construction paradigms, and intrinsic quality assurance, and task synthesis around intent formation, grounding, completion verification, and task quality. When a target agent uses the resources, its performance and learning progress enable a further, agent-relative assessment of their suitability, which guides task and environment evolution. This resource perspective connects the two synthesis subjects through their common purpose: providing reliable and continually relevant opportunities for agent learning through interaction.
Appendix A. Literature Collection
The literature scope covers work publicly available up to September 16, 2026. Candidate papers were collected from public scholarly sources, including arXiv, ACL Anthology, OpenReview, and proceedings of major NLP, machine learning, and data-mining venues. Search terms combined agent environment synthesis, generation, or evolution with task synthesis, executable or verifiable environments, and terminal, Web, GUI, or tool agents. The pool was expanded through reference-list snowballing and targeted searches around recurring method families and newly emerging systems.
The core inclusion criterion is an automatic, scalable mechanism for constructing or expanding environments or grounded tasks. This includes simulation, reconstruction/adaptation, creation from scratch, grounded task generation, and repeated task/environment adaptation. Downstream use does not determine eligibility: evaluation-oriented methods are included when they contribute such construction mechanisms, while trajectory-generation systems are included for their task/environment construction components. Static instruction synthesis, fixed-task trajectory collection, and manually authored resources are treated as adjacent context unless they provide relevant scalable construction mechanisms. This collection supports a structured technical synthesis rather than a claim of exhaustive coverage. Appendix B maps representative systems to the survey taxonomy; inclusion in a table does not imply that every component of a system falls within the core scope.
Appendix B. Literature Mapping
The three tables map representative literature to Section 3, Section 4, and Section 5. A work may appear in multiple tables, and combined labels denote multiple sources or routes. The marker † denotes a boundary or precursor case.
Appendix B.1. Environment Synthesis Literature
“Model-based simulation,” “Recon./adapt.,” and “From scratch” abbreviate the three paradigms in Section 3.2. Combined routes are listed where needed; hybrid operation, recursive extension, adaptive evolution, and repair are described as mechanisms rather than additional paradigms. Here, reconstruction/adaptation can also reuse previously generated artifacts. Table A1 records how substrates are obtained.
Table A1.
Representative literature on automated environment synthesis. Routes denote substrate provenance; adaptation, evolution, and repair mechanisms are described separately.
Table A1.
Representative literature on automated environment synthesis. Routes denote substrate provenance; adaptation, evolution, and repair mechanisms are described separately.
| Work | Setting | Route | Primary synthesis target / mechanism | Quality or verification signal |
|---|---|---|---|---|
| AgentGen Hu et al. 2025 | Planning | From scratch | PDDL environment and planning-task generation from inspiration corpora | Executable planning checks |
| gg-bench Verma et al. 2025 | Games | From scratch | Generated game descriptions materialized as Gym environments for evaluation | Game rules / executable evaluation |
| RandomWorld Sullivan et al. 2025 | Tool | From scratch | Procedural generation of interactive tools and compositional worlds | Programmatic checks |
| StableToolBench-MirrorAPI Guo et al. 2025 | API | Model-based simulation | API-conditioned response simulation with a model trained on request–response data | Simulated API consistency |
| Simia Li et al. 2025 | Tool / API | Model-based simulation | Reasoning-model environment simulator conditioned on tools and history | Simulated feedback |
| AgentScaler Fang et al. 2026 | Tool / API | From scratch | Programmatic state, database, and tool materialization from API structures | State / rule-based checks |
| AutoForge Cai et al. 2025 | Tool / API | From scratch | Automated construction of simulated executable environments | Executable / verifiable tasks |
| CodeGym Du et al. 2026 | Code / tool | Recon./adapt. | Conversion of static coding problems into interactive Gym-style environments | Programmatic rewards |
| SWE-Playground Zhu et al. 2025 | SWE | From scratch | Generated software projects, repositories, and task-specific workspaces | Unit tests |
| DreamGym Chen et al. 2026b | Reasoning / tool | Model-based simulation | Learned experience model for synthetic transitions and adaptive challenges | Real/synthetic calibration |
| Endless Terminals Gandhi et al. 2026 | Terminal | From scratch | Containerized terminal environments with generated services and dependencies | Executable tests |
| DynaWeb Ding et al. 2026 | Web | Model-based simulation | Learned Web transition/world model | Model-based transition prediction |
| MEnvAgent Guo et al. 2026a | SWE | Recon./adapt. | Automated dependency resolution and Dockerized runtime construction | Build/run/repair checks |
| Environment-free API Simulation Lee et al. 2026 | API | Model-based simulation | Stateful API response simulation from specifications | Simulated response consistency |
| ScaleEnv Tu et al. 2026 | Tool | From scratch | Interactive environments expanded through tool-dependency structures | Procedural/executable tests |
| Agent World Model Wang et al. 2026c | Tool / API | From scratch | SQL-backed state and executable tool interfaces | State-based checks |
| C-World Xi et al. 2026 | Computer use | Model-based simulation + recon./adapt. | Hybrid live-API execution and World-Engine simulation for training/evaluation | Mixed rule/model checks |
| AutoWebWorld Wu et al. 2026d | Web | From scratch | FSM-grounded generation of executable websites | FSM/programmatic verification |
| InfiniteWeb Zhang et al. 2026c | Web | From scratch | Multi-page Web application synthesis from references/specifications | Executable app/reset checks |
| GUI-GENESIS Cao et al. 2026 | GUI | Recon./adapt. | Lightweight reconstruction of real GUI applications | Code-native assertions |
| CLI-Gym Lin et al. 2026 | Terminal | Recon./adapt. | Failure-state inversion and environment-intensive task packaging | Executable tests |
| TermiGen Zhu et al. 2026 | Terminal | From scratch | Generated terminal containers and executable workspaces | Executable verifier |
| EnvScaler Song et al. 2026 | Tool / API | From scratch | Programmatic state, dynamics, and interface synthesis | Probing / state checks |
| VeriEnv Chae et al. 2026 | Web | Recon./adapt. | Executable, resettable reconstruction of real websites | Deterministic state checks |
| Agent-World Dong et al. 2026 | Tool / API | From scratch | Discovery and synthesis of heterogeneous executable environments | Verifiable tasks / RL feedback |
| Terminal-World Cheng et al. 2026 | Terminal | From scratch | Skill-based joint construction of tasks, environments, and trajectories | Executable bundles |
| EnvFactory Xu et al. 2026a | Tool / API | Recon./adapt. | Discovery and validation of stateful executable tool environments | Executability checks |
| LiteCoder-Terminal Peng et al. 2026 | Terminal | From scratch | Long-horizon terminal environment generation and packaging | Runtime checks |
| SynthTools Castellani et al. 2025 | Tool | From scratch | Tool generation, simulation, and audit | Tool audit |
| CUA-Gym Wang et al. 2026a | Computer use | From scratch + recon./adapt. | Synthetic mock applications plus task-specific states in computer-use runtimes | State / judge |
| Tmax Ivison et al. 2026 | Terminal | From scratch | Large-scale diversified terminal environments and verifiers | Diversified verifiers |
| SETA Shen et al. 2026 | Terminal | From scratch | Automated terminal environment construction with later adaptive evolution | Unified executable verifier |
| GAIS Shi et al. 2026b | Tool / MCP | From scratch | Protocol/tool-dependency-grounded interaction synthesis | Grounded interaction checks |
| Meta-Task Pan et al. 2026 | Terminal | From scratch | Automated executable task/workspace generation | Sandbox execution |
| RST Li et al. 2026c | Terminal | Recon./adapt. | Recursive extension of verified seed-task artifacts and workspaces | Executable tests |
| AppDeltaWorld Xu et al. 2026b | Mobile GUI | Model-based simulation | Delta-code world model for GUI transitions | Transition prediction |
| FACET Shi et al. 2026a | Terminal | From scratch | Source-skill-grounded container construction, execution, and targeted repair | Container + tests |
| SPADE Liu et al. 2026 | General / tool | From scratch | Adaptive self-play generation of executable reset/step environments | Executable interaction / regret |
| EnvHarness Huang et al. 2026a | Multi-domain | Recon./adapt. | Programmable wrappers that reshape frozen environments | Original verifier + fresh rollouts |
| EvoEnv Shi et al. 2026c | Reasoning | From scratch | Adaptive generation and filtering of executable reasoning environments | Oracle + staged checks |
| Terminal-Universe Wu et al. 2026a | Terminal | Recon./adapt. | Environment recovery from trajectories and execution artifacts | Executable workspace |
| Environment Evolution Fan et al. 2026b | Terminal | Recon./adapt. | Environment-centric transformations of seed environments along difficulty directions | Executable environment checks |
| Trustworthy Worlds Zhang et al. 2026a | Web | From scratch | Synthetic Web environment construction with pre-training verification and repair | Structural/state predicates |
| AgentMercury Jeong and Yoon 2026 | Business / tool | From scratch | Persistent multi-service environments with shared state | Cross-service invariants |
| TerminalTraj Wu et al. 2026b | Terminal | Recon./adapt. | Repository-derived Docker environments and environment-aligned task instances | Executable validation code |
Appendix B.2. Task Synthesis Literature
The intent-source column in Table A2 uses Seed, Tool, State, and Trajectory as defined in Section 4.1, followed by the specific source or control mechanism. Externally supplied skill descriptions are seeds; exposed skill/tool composition structures are tool-based sources. Verification entries summarize checking mechanisms and may mix evidence sources with executable or model-based checks; they are not an exhaustive inventory of each system’s verification modes.
Table A2.
Representative literature on automated task synthesis. Intent sources follow the four-source taxonomy; qualifiers explain composition, grounding, or feedback-guided control.
Table A2.
Representative literature on automated task synthesis. Intent sources follow the four-source taxonomy; qualifiers explain composition, grounding, or feedback-guided control.
| Work | Setting | Intent source / qualifier | Grounding / instantiation | Verification |
|---|---|---|---|---|
| AgentGen Hu et al. 2025 | Planning | Seed; generated-world grounding | PDDL states and actions | Executable plan |
| NNetNav Murty et al. 2024 | Web | Trajectory | Browser interaction trace | Execution outcome |
| OS-Genesis Sun et al. 2025 | GUI | Trajectory | Executed GUI interaction | Execution / state |
| TaskCraft Shi et al. 2025 | Tool | Seed; compositional expansion | Multi-tool executable setting | Executable trajectory |
| AgentSynth Xie et al. 2026 | Computer use | Seed; subtask composition | Existing computer-use runtime | Successful execution |
| RandomWorld Sullivan et al. 2025 | Tool | Tool; composition | Procedurally generated tool world | Programmatic |
| AgentScaler Fang et al. 2026 | Tool / API | Tool / state; domain structure | Generated database state and executable APIs | State / rule based |
| AutoForge Cai et al. 2025 | Tool / API | Seed; task/domain requirements | Synthesized executable environment | Executable / verifiable |
| SWE-Playground Zhu et al. 2025 | SWE | Seed; project proposal | Generated repository/workspace | Unit tests |
| ScaleEnv Tu et al. 2026 | Tool | Tool; dependency graph | Generated interactive tool environment | Executable action checks |
| Agent World Model Wang et al. 2026c | Tool / API | State / tool; backend structure | SQL-backed environment state | State based |
| C-World Xi et al. 2026 | Computer use | Seed; workflow/constraints | Live or synthesized computer world | Mixed rule + judge |
| AutoWebWorld Wu et al. 2026d | Web | State; goal/FSM | Generated website state graph | FSM / programmatic |
| InfiniteWeb Zhang et al. 2026c | Web | Seed; reference design | Synthesized Web application | Executable app/reset |
| CLI-Gym Lin et al. 2026 | Terminal | State; injected failure | Reconstructed terminal workspace | Executable tests |
| TermiGen Zhu et al. 2026 | Terminal | Seed; task/container specification | Synthesized executable workspace | Executable verifier |
| EnvScaler Song et al. 2026 | Tool / API | State / tool; rules | Generated database state | State / probing |
| VeriEnv Chae et al. 2026 | Web | State; reconstructed site | Executable cloned website | Deterministic state checks |
| DIVE Chen et al. 2026a | Tool | Trajectory | Real tool execution | Execution-grounded |
| Agent-World Dong et al. 2026 | Tool / API | Tool / state; capability-guided | Discovered/generated environments | Verifiable tasks |
| SkillSynth Fan et al. 2026a | Terminal | Tool; skill graph/path | Executable terminal workspace | Executable task |
| Terminal-World Cheng et al. 2026 | Terminal | Tool; skill primitives | Skill-conditioned environment/task bundle | Executable bundle |
| GTA Huang et al. 2026b | Web | Seed / state; site graph | Reachable pages and Web paths | Grounded execution |
| Trajectory2Task Wang et al. 2026d | Tool | Trajectory | Executed multi-turn tool trace | Executable / verifiable |
| CUA-Gym Wang et al. 2026a | Computer use | Seed / state; task templates | Task-specific application state | State / judge |
| TASTE Keren et al. 2026 | Tool | Tool; sequences / dependencies | Tool-sequence instantiation into executable benchmark tasks | Executability / scoring |
| Tmax Ivison et al. 2026 | Terminal | Seed; taxonomy/persona | Generated terminal environment | Diversified verifiers |
| SETA Shen et al. 2026 | Terminal | Seed; diverse sources | Synthesized terminal environment | Unified executable verifier |
| GAIS Shi et al. 2026b | Tool / MCP | Tool; dependency graph | Protocol-grounded tool ecosystem | Grounded interaction |
| ScaleCUA Lv et al. 2026 | Computer use | State / trajectory; exploration | Docker-interaction task generation; later frontier sampling | Executable judges |
| NexForge Zhao et al. 2026 | Terminal | Seed; capability requirements | Task-specific executable workspace | Executable tests |
| Meta-Task Pan et al. 2026 | Terminal | Seed; task profile | Generated sandbox/workspace | Sandbox execution |
| ACuRL Xue et al. 2026 | Computer use | Trajectory; feedback-guided | Existing GUI environment | Model/rubric based |
| SKT Tan et al. 2026 | Tool / skill | Tool; skill composition | Existing tool/skill configurations | Verified data |
| State2State Lei et al. 2026 | General | State; reachable targets | Existing environment state | Target-state match |
| RST Li et al. 2026c | Terminal | Seed; reference solution | Recursively expanded executable workspace | Executable tests |
| CalibForge Meng et al. 2026 | Terminal | Seed; solver-calibrated | Existing/synthesized workspace | Solver + executable tests |
| Envs-FORGE Wu et al. 2026c | Terminal | Seed; reward-guided rewriting | Jointly rewritten task and environment artifacts | Gold-verified bundle |
| FACET Shi et al. 2026a | Terminal | Seed; source skill/intent | Repaired executable container | Container + tests |
| Terminal-Universe Wu et al. 2026a | Terminal | Trajectory / state; recovered workspace | Reconstructed reusable workspace | Executable workspace |
| AutoPlay Ramrakhya et al. 2026 | GUI | Trajectory / state; exploration | Discovered application state | Execution outcome |
| SynWeaver Wang et al. 2026b | Web | Trajectory / state; site prior | Existing website structure | Path / state based |
| AgentMercury Jeong and Yoon 2026 | Business / tool | Seed / state; scenario | Persistent multi-service state | Cross-service invariants |
| TerminalTraj Wu et al. 2026b | Terminal | Seed / tool; repository docs/scripts | Repository-derived Docker environments | State-based pytest |
| CLI-Universe Hua et al. 2026 | Terminal | Seed; capability taxonomy | Evidence-refined blueprints realized in Docker environments | Rubric-gated tests; reference solution; fail-to-pass checks |
| Terminal-Task-Gen Pi et al. 2026 | Terminal | Seed; problems/skill taxonomy | Generated inputs and tasks in shared domain-specific Docker images | Weighted pytest checks; optional seed-solution expectations |
Appendix B.3. Learning-Driven Evolution Literature
Table A3 summarizes methods in which evidence from agent interaction, solver behavior, or learning progress is used to assess and adapt tasks, environments, or their distributions over time. These signals support the capability estimation and agent-relative resource assessment discussed in Section 5, and are then translated into changes to the available interactive learning resources. We focus on agent-conditioned evolution, where resource adaptation is informed by the behavior or performance of a target learner or solver. Related mechanisms that provide programmable adaptation but do not themselves close a learner-driven training loop are marked with †. Iterative environment generation that follows predefined transformation directions without using current-agent evidence is treated as a boundary case closer to environment construction rather than as a core form of learning-driven evolution. Pre-training defect repair, such as Trustworthy Worlds, is therefore recorded in Table A1 rather than counted as learning-driven evolution.
Table A3.
Representative methods for learning-driven task and environment evolution. The table separates the agent or solver evidence used for assessment, the resource being adapted, and the mechanism through which adaptation is performed. Works marked with † are precursor or boundary mechanisms that do not themselves instantiate a complete learner-driven evolution loop.
Table A3.
Representative methods for learning-driven task and environment evolution. The table separates the agent or solver evidence used for assessment, the resource being adapted, and the mechanism through which adaptation is performed. Works marked with † are precursor or boundary mechanisms that do not themselves instantiate a complete learner-driven evolution loop.
| Work | Setting | Agent / solver evidence | Adaptation target | Adaptation mechanism |
|---|---|---|---|---|
| DreamGym Chen et al. 2026b | Reasoning / tool | Real interaction experience / model calibration | Environment challenges | Continual update of the experience model and generated challenges |
| RLVE† Zeng et al. 2025 | Reasoning | Agent success rate | Procedural task difficulty | Capability-adaptive curriculum over manually engineered environments |
| GenEnv Guo et al. 2026b | General | Policy performance / curriculum reward | Environment difficulty | Difficulty-aligned agent–environment co-evolution |
| ProEvolve† Li et al. 2026a | General | Benchmark state / transformation specification | Environment and benchmark artifacts | Programmable graph transformations for benchmark evolution |
| Agent-World Dong et al. 2026 | Tool / API | Capability gaps / interaction performance | Tasks and environments | Capability-guided discovery of new environments and verifiable tasks |
| TRACE Kang et al. 2026 | General | Successful vs. failed trajectories | Tasks and environments | Capability-targeted mini-environment synthesis |
| ScaleCUA Lv et al. 2026 | Computer use | Task success / capability frontier | Task sampling distribution | Frontier-based task sampling during online RL |
| SETA Shen et al. 2026 | Terminal | Agent performance / training progress | Terminal tasks and environments | Environment and task mutation during continual training |
| ACuRL Xue et al. 2026 | Computer use | Previous interaction feedback | Curriculum tasks | Continual regeneration of curriculum tasks |
| CalibForge Meng et al. 2026 | Terminal | Solver behavior / pass rate | Task difficulty | Solver-relative calibration toward the learnable frontier |
| Envs-FORGE Wu et al. 2026c | Terminal | Verifier reward / seed pass rate | Tasks, fixtures, oracle, tests, and runtime | Frontier-optimized joint rewriting of task and environment artifacts |
| SPADE Liu et al. 2026 | General / tool | Regret / agent performance | Executable environments | Self-play generation near the learner’s capability frontier |
| EnvHarness Huang et al. 2026a | Multi-domain | Policy weaknesses extracted from trajectories | Environment wrappers, rules, and observations | Harness synthesis targeted at observed policy weaknesses |
| EvoEnv Shi et al. 2026c | Reasoning | Agent-relative difficulty / novelty | Executable reasoning environments | Generate and filter environment programs around the current learning frontier |
References
- Ivan Bercovich. 2026. What makes a good terminal-agent benchmark task: A guideline for adversarial, difficult, and legible evaluation design. Preprint, arXiv:2604.28093.
- Shihao Cai, Runnan Fang, Jialong Wu, Baixuan Li, Xinyu Wang, Yong Jiang, Liangcai Su, Liwen Zhang, Wenbiao Yin, Zhen Zhang, Fuli Feng, Pengjun Xie, and Xiaobin Wang. 2025. AutoForge: Automated environment synthesis for agentic reinforcement learning. Preprint, arXiv:2512.22857.
- Yuan Cao, Dezhi Ran, Mengzhou Wu, Yuzhe Guo, Xin Chen, Ang Li, Gang Cao, Gong Zhi, Hao Yu, Linyi Li, Wei Yang, and Tao Xie. 2026. GUI-GENESIS: Automated synthesis of efficient environments with verifiable rewards for GUI agent post-training. Preprint, arXiv:2602.14093.
- Tommaso Castellani, Naimeng Ye, Daksh Mittal, Thomson Yen, and Hongseok Namkoong. 2025. SynthTools: A framework for scaling synthetic tools for agent development. Preprint, arXiv:2511.09572.
- Hyungjoo Chae, Jungsoo Park, and Alan Ritter. 2026. Safe and scalable web agent learning via recreated websites. Preprint, arXiv:2603.10505.
- Aili Chen, Chi Zhang, Junteng Liu, Jiangjie Chen, Chengyu Du, Yunji Li, Ming Zhong, Qin Wang, Zhengmao Zhu, Jiayuan Song, Ke Ji, Junxian He, Pengyu Zhao, and Yanghua Xiao. 2026a. DIVE: Scaling diversity in agentic task synthesis for generalizable tool use. Preprint, arXiv:2603.11076.
- Zhaorun Chen, Zhuokai Zhao, Kai Zhang, Bo Liu, Qi Qi, Yifan Wu, Tarun Kalluri, Sara Cao, Yuanhao Xiong, Haibo Tong, Huaxiu Yao, Hengduo Li, Jiacheng Zhu, Xian Li, Dawn Song, Bo Li, Jason Weston, and Dat Huynh. 2026b. Scaling agent learning via experience synthesis. In International Conference on Learning Representations.
- Zihao Cheng, Hongru Wang, Zeming Liu, Xinyi Wang, Xiangrong Zhu, Yuhang Guo, Wei Lin, Jeff Z. Pan, and Yunhong Wang. 2026. Terminal-World: Scaling terminal-agent environments via agent skills. Preprint, arXiv:2605.20876.
- Hang Ding, Peidong Liu, Junqiao Wang, Ziwei Ji, Meng Cao, Rongzhao Zhang, Lynn Ai, Eric Yang, Tianyu Shi, and Lei Yu. 2026. DynaWeb: Model-based reinforcement learning of web agents. Preprint, arXiv:2601.22149.
- Guanting Dong, Junting Lu, Junjie Huang, Wanjun Zhong, Longxiang Liu, Shijue Huang, Zhenyu Li, Yang Zhao, Xiaoshuai Song, Xiaoxi Li, Jiajie Jin, Yutao Zhu, Hanbin Wang, Fangyu Lei, Qinyu Luo, Mingyang Chen, Zehui Chen, Jiazhan Feng, Ji-Rong Wen, and Zhicheng Dou. 2026. Agent-World: Scaling real-world environment synthesis for evolving general agent intelligence. Preprint, arXiv:2604.18292.
- Weihua Du, Hailei Gong, Zhan Ling, Kang Liu, Lingfeng Shen, Xuesong Yao, Yufei Xu, Dingyuan Shi, Yiming Yang, and Jiecao Chen. 2026. Generalizable end-to-end tool-use RL with synthetic CodeGym. In International Conference on Learning Representations.
- Zhiyuan Fan, Tinghao Yu, Yuanjun Cai, Jiangtao Guan, Yun Yang, Dingxin Hu, Jiang Zhou, Xing Wu, Zhuo Han, Feng Zhang, and Lilin Wang. 2026a. Toward scalable terminal task synthesis via skill graphs. Preprint, arXiv:2604.25727.
- Zhiyuan Fan, Tinghao Yu, Yuanjun Cai, Jiang Zhou, Jiangtao Guan, Jincheng Liu, Yun Yang, Dingxin Hu, Zhuo Han, Xing Wu, Feng Zhang, and Lilin Wang. 2026b. Environment evolution for terminal agents. Preprint, arXiv:2609.04128.
- Runnan Fang, Shihao Cai, Baixuan Li, Jialong Wu, Guangyu Li, Wenbiao Yin, Xinyu Wang, Xiaobin Wang, Liangcai Su, Zhen Zhang, Shibin Wu, Zhengwei Tao, Yong Jiang, Pengjun Xie, Ningyu Zhang, Fei Huang, Wentao Zhang, and Jingren Zhou. 2026. Towards general agentic intelligence via environment scaling. In Findings of the Association for Computational Linguistics: ACL 2026, pages 17610–17621. Association for Computational Linguistics.
- Kanishk Gandhi, Shivam Garg, Noah D. Goodman, and Dimitris Papailiopoulos. 2026. Endless terminals: Scaling RL environments for terminal agents. Preprint, arXiv:2601.16443.
- Huan-ang Gao, Jiayi Geng, Wenyue Hua, Mengkang Hu, Xinzhe Juan, Hongzhang Liu, Shilong Liu, Jiahao Qiu, Xuan Qi, Yiran Wu, Hongru Wang, Han Xiao, Yuhang Zhou, Shaokun Zhang, Jiayi Zhang, Jinyu Xiang, Yixiong Fang, Qiwen Zhao, Dongrui Liu, and 8 others. 2026. A survey of self-evolving agents: What, when, how, and where to evolve on the path to artificial super intelligence. Transactions on Machine Learning Research.
- Chuanzhe Guo, Jingjing Wu, Sijun He, Yang Chen, Zhaoqi Kuang, Shilong Fan, Bingjin Chen, Siqi Bao, Jing Liu, Hua Wu, Qingfu Zhu, Wanxiang Che, and Haifeng Wang. 2026a. MEnvAgent: Scalable polyglot environment construction for verifiable software engineering. Preprint, arXiv:2601.22859.
- Jiacheng Guo, Ling Yang, Peter Chen, Qixin Xiao, Yinjie Wang, Xinzhe Juan, Jiahao Qiu, Ke Shen, and Mengdi Wang. 2026b. GenEnv: Difficulty-aligned co-evolution between LLM agents and environment simulators. In International Conference on Learning Representations.
- Zhicheng Guo, Sijie Cheng, Yuchen Niu, Hao Wang, Sicheng Zhou, Wenbing Huang, and Yang Liu. 2025. StableToolBench-MirrorAPI: Modeling tool environments as mirrors of 7,000+ real-world APIs. In Findings of the Association for Computational Linguistics: ACL 2025, pages 5247–5270. Association for Computational Linguistics.
- Mengkang Hu, Pu Zhao, Can Xu, Qingfeng Sun, Jian-Guang Lou, Qingwei Lin, Ping Luo, and Saravan Rajmohan. 2025. AgentGen: Enhancing planning abilities for large language model based agent via environment and task generation. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pages 496–507.
- Zhanbo Hua, Yifan Yao, Weihao Xie, Yongchi Zhao, Minghao Liu, Ruizhi Qiu, Zhewei Huang, Zun Wang, Yiyan Ji, Yunhai Ye, Letian Zhu, Xinping Lei, Han Li, Zhiyuan Ma, Zili Wang, Zhaoxiang Zhang, and Jiaheng Liu. 2026. CLI-Universe: Towards verifiable task synthesis engine for terminal agents. Preprint, arXiv:2606.22883.
- Chengsong Huang, Zifeng Wang, Rujun Han, Jun Yan, Yanfei Chen, Zoey CuiZhu, Ke Jiang, Peng Xia, Han Yu, Yufan Zhuang, Yifei Ming, Jiaqi Pan, Bhavana Dalvi Mishra, Jiaxin Huang, Burak Gokturk, Tomas Pfister, and Chen-Yu Lee. 2026a. EnvHarness: Awakening static worlds for agent learning. Preprint, arXiv:2608.19880.
- Tenghao Huang, Kung-Hsiang Huang, Prafulla Kumar Choubey, Yilun Zhou, Muhao Chen, Jonathan May, and Chien-Sheng Wu. 2026b. GTA: Generating long-horizon tasks for web agents at scale. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 18805–18820. Association for Computational Linguistics.
- Maksim Ivanov and Abhijay Rana. 2026. Anchor: Mitigating artifact drift in agent benchmark generation. Preprint, arXiv:2605.26321. Presented at the RLEval Workshop, ACM CAIS 2026 (non-archival).
- Hamish Ivison, Junjie Oscar Yin, Rulin Shao, Teng Xiao, Nathan Lambert, and Hannaneh Hajishirzi. 2026. Tmax: A simple recipe for terminal agents. Preprint, arXiv:2606.23321.
- Minbyul Jeong and Chanwoong Yoon. 2026. AgentMercury: Your agent can synthesize verifiable environments for business scenarios at scale. Preprint, arXiv:2608.20634.
- Hangoo Kang, Tarun Suresh, Jon Saad-Falcon, and Azalia Mirhoseini. 2026. TRACE: Capability-targeted agentic training. Preprint, arXiv:2604.05336.
- Tomer Keren, Nitay Calderon, Asaf Yehudai, Yotam Perlitz, Michal Shmueli-Scheuer, and Roi Reichart. 2026. A matter of TASTE: Improving coverage and difficulty of agent benchmarks. Preprint, arXiv:2605.28556.
- Seanie Lee, Sanjoy Chowdhury, Chao Jiang, Cheng-Yu Hsieh, Ting-Yao Hu, Alexander T. Toshev, Oncel Tuzel, and Raviteja Vemulapalli. 2026. Environment-free synthetic data generation for API-calling agents. Preprint, arXiv:2607.16900.
- Xuanyu Lei, Yiqi Zhu, Chenliang Li, Kaiming Liu, Peng Li, Ming Yan, Jieping Ye, Ya-Qin Zhang, and Yang Liu. 2026. State2State: Environment-derived mid-training for LLM agents. Preprint, arXiv:2608.04934.
- Guangrui Li, Yaochen Xie, Yi Liu, Ziwei Dong, Xingyuan Pan, Tianqi Zheng, Jason Choi, Michael J. Morais, Binit Jha, Shaunak Mishra, Bingrou Zhou, Chen Luo, Monica Xiao Cheng, and Dawn Song. 2026a. The world won’t stay still: Programmable evolution for agent benchmarks. Preprint, arXiv:2603.05910.
- Jiachun Li, Zhuoran Jin, Tianyi Men, Yupu Hao, Kejian Zhu, Lingshuai Wang, Dongqi Huang, Longxiang Wang, Shengjia Hua, Lu Wang, Jinshan Gao, Hongbang Yuan, Ruilin Xu, Kang Liu, and Jun Zhao. 2026b. Agentic environment engineering for large language models: A survey of environment modeling, synthesis, evaluation, and application. Preprint, arXiv:2606.12191.
- Yuetai Li, Huseyin A. Inan, Xiang Yue, Wei-Ning Chen, Lukas Wutschitz, Janardhan Kulkarni, Radha Poovendran, Robert Sim, and Saravan Rajmohan. 2025. Simulating environments with reasoning models for agent training. Preprint, arXiv:2511.01824.
- Zhongzhi Li, Yucheng Shi, Zongxia Li, Ruhan Wang, Anhao Li, Zixun Huang, Junyao Yang, Lei Ke, Ninghao Liu, Haitao Mi, and Leowei Liang. 2026c. Recursive synthesis for long-horizon terminal tasks. Preprint, arXiv:2608.05466.
- Yusong Lin, Haiyang Wang, Shuzhe Wu, Lue Fan, Feiyang Pan, Sanyuan Zhao, and Dandan Tu. 2026. CLI-Gym: Scalable CLI task generation via agentic environment inversion. Preprint, arXiv:2602.10999.
- Bo Liu, Simon Yu, Yiding Jiang, Ao Qu, Andrew Zhao, Zichen Liu, Junsu Kim, Zijian Zhou, Seungone Kim, Tongzheng Ren, Mickel Liu, Hanfei Yu, Zhaorun Chen, Weiyan Shi, Paul Pu Liang, Luke Zettlemoyer, Yejin Choi, and Natasha Jaques. 2026. SPADE: Self-play in adaptive synthetic executable environments. Preprint, arXiv:2608.19197.
- Bowen Lv, Xiao Liu, Yanyu Ren, Hanyu Lai, Bohao Jing, Hanchen Zhang, Yanxiao Zhao, Shuntian Yao, Jie Tang, and Yuxiao Dong. 2026. SCALECUA: Scaling computer use agents with verifiable task synthesis and efficient online RL. Preprint, arXiv:2607.11185.
- Fanzhe Meng, Guoxin Chen, Jiale Zhao, Shuang Sun, Zhiyu Lin, Wayne Xin Zhao, Ruihua Song, Ji-Rong Wen, and Kai Jia. 2026. CalibForge: Adversarial solver calibration for scaling learnable terminal tasks. Preprint, arXiv:2608.06352.
- Shikhar Murty, Hao Zhu, Dzmitry Bahdanau, and Christopher D. Manning. 2024. NNetNav: Unsupervised learning of browser agents through environment interaction in the wild. Preprint, arXiv:2410.02907.
- Zhihong Pan, Jiyuan He, Kai Zhang, Yupeng Han, Ze Liu, Yuze Zhao, Yongcong Ye, and Zhaohua Yang. 2026. Meta-Task: Turning terminal task synthesis into a terminal task for scalable agent training. Preprint, arXiv:2607.27929.
- Xiaoxuan Peng, Kaiqi Zhang, Xinyu Lu, Boxi Cao, Yaojie Lu, Hongyu Lin, Xianpei Han, and Le Sun. 2026. LiteCoder-Terminal: Scaling long-horizon terminal environments for learning language agents. Preprint, arXiv:2605.29559.
- Renjie Pi, Grace Lam, Mohammad Shoeybi, Pooya Jannaty, Bryan Catanzaro, and Wei Ping. 2026. On data engineering for scaling LLM terminal capabilities. Preprint, arXiv:2602.21193.
- Ram Ramrakhya, Andrew Szot, Omar Attia, Yuhao Yang, Anh Nguyen, Bogdan Mazoure, Zhe Gan, Harsh Agrawal, and Alexander Toshev. 2026. Scaling synthetic task generation for agents via exploration. In International Conference on Learning Representations.
- Qijia Shen, Zhiqi Huang, Vamsidhar Kamanuru, Aznaur Aliev, Jay Rainton, Ahmed Awelkair, Zhichen Zeng, Jiajun Li, Shi Dong, Yueming Yuan, Boyuan Ma, Qizheng Zhang, Jiwei Fu, Yuzhen Mao, Wendong Fan, Ping Nie, Philip Torr, Bernard Ghanem, Changran Hu, and 3 others. 2026. SETA: Scaling environments for terminal agents. Preprint, arXiv:2607.10891.
- Dingfeng Shi, Jingyi Cao, Qianben Chen, Weichen Sun, Weizhen Li, Hongxuan Lu, Fangchen Dong, Tianrui Qin, King Zhu, Minghao Yang, Jian Yang, Ge Zhang, Jiaheng Liu, Changwang Zhang, Jun Wang, Yuchen Eleanor Jiang, and Wangchunshu Zhou. 2025. TaskCraft: Automated generation of agentic tasks. Preprint, arXiv:2506.10055.
- Kou Shi, Zun Wang, Qisheng Su, Shiting Huang, Ziao Zhang, Zhen Fang, Qingnan Ren, Jin Liu, Yu Zeng, Yiming Zhao, Lin Chen, Zehui Chen, and Feng Zhao. 2026a. FACET: Preserving source intent and executable state in terminal task synthesis. Preprint, arXiv:2608.18580.
- Wenhang Shi, Jinhao Dong, Yiren Chen, Zhe Zhao, Shuqing Bian, Wei Lu, and Xiaoyong Du. 2026b. Scaling agentic capabilities via grounded interaction synthesis. In Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pages 4252–4263.
- Yucheng Shi, Zhenwen Liang, Kishan Panaganti, Dian Yu, Wenhao Yu, and Haitao Mi. 2026c. Learning to build the environment: Self-evolving reasoning RL via verifiable environment synthesis. Preprint, arXiv:2605.14392.
- Xiaoshuai Song, Haofei Chang, Guanting Dong, Yutao Zhu, Ji-Rong Wen, and Zhicheng Dou. 2026. EnvScaler: Scaling tool-interactive environments for LLM agent via programmatic synthesis. In Findings of the Association for Computational Linguistics: ACL 2026, pages 8326–8357. Association for Computational Linguistics.
- Michael Sullivan, Mareike Hartmann, and Alexander Koller. 2025. Procedural environment generation for tool-use agents. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 18544–18562. Association for Computational Linguistics.
- Qiushi Sun, Kanzhi Cheng, Zichen Ding, Chuanyang Jin, Yian Wang, Fangzhi Xu, Zhenyu Wu, Chengyou Jia, Liheng Chen, Zhoumianze Liu, Ben Kao, Guohao Li, Junxian He, Yu Qiao, and Zhiyong Wu. 2025. OS-genesis: Automating GUI agent trajectory construction via reverse task synthesis. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5555–5579. Association for Computational Linguistics.
- Zelin Tan, Yiqun Zhang, Hao Li, Zhiyao Cui, Hejia Geng, Shao Zhang, Hangfan Zhang, Yang Chen, Xiaosong Wang, Lilong Wang, Zhenfei Yin, Shuyue Hu, Chen Zhang, and Lei Bai. 2026. SKT: Skill-use training at scale via verified synthetic data generation. Preprint, arXiv:2608.02287.
- Dunwei Tu, Hongyan Hao, Hansi Yang, Yihao Chen, Yi-Kai Zhang, Zhikang Xia, Yu Yang, Yueqing Sun, Xingchen Liu, Furao Shen, Qi Gu, Hui Su, and Xunliang Cai. 2026. ScaleEnv: Scaling environment synthesis from scratch for generalist interactive tool-use agent training. Preprint, arXiv:2602.06820.
- Vivek Verma, David Huang, William Chen, Dan Klein, and Nicholas Tomlin. 2025. Measuring general intelligence with generated games. Preprint, arXiv:2505.07215.
- Bowen Wang, Dunjie Lu, Junli Wang, Tianyi Bai, Shixuan Liu, Zhipeng Zhang, Haiquan Wang, Hao Hu, Tianbao Xie, Shuai Bai, Dayiheng Liu, Que Shen, Junyang Lin, and Tao Yu. 2026a. CUA-Gym: Scaling verifiable training environments and tasks for computer-use agents. Preprint, arXiv:2605.25624.
- Ruitao Wang, Yuwen Hao, and Menglin Yang. 2026b. SynWeaver: Website-prior task and trajectory co-synthesis for web agents. Preprint, arXiv:2608.12429.
- Zhaoyang Wang, Canwen Xu, Boyi Liu, Yite Wang, Siwei Han, Zhewei Yao, Huaxiu Yao, and Yuxiong He. 2026c. Agent world model: Infinity synthetic environments for agentic reinforcement learning. Preprint, arXiv:2602.10090.
- Ziyi Wang, Yuxuan Lu, Yimeng Zhang, Pei Chen, Ziwei Dong, Jing Huang, Jiri Gesi, Xianfeng Tang, Chen Luo, Qun Liu, Yisi Sang, Hanqing Lu, Manling Li, Jin Lai, and Dakuo Wang. 2026d. Trajectory2Task: Training robust tool-calling agents with synthesized yet verifiable data for complex user intents. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 44021–44044. Association for Computational Linguistics.
- Jie Wu, Zhenru Zhang, Beichen Zhang, Xuwu Wang, Yuhui Su, Mouxiang Chen, Peng Wang, Zhihai Wang, Que Shen, Hao Zhou, An Yang, Fei Huang, Yujiu Yang, and Dayiheng Liu. 2026a. Terminal-Universe: Turning agent trajectories into scalable terminal environments. Preprint, arXiv:2609.04148.
- Siwei Wu, Yizhi Li, Yuyang Song, Wei Zhang, Yang Wang, Riza Batista-Navarro, Xian Yang, Mingjie Tang, Bryan Dai, Jian Yang, and Chenghua Lin. 2026b. Large-scale terminal agentic trajectory generation from dockerized environments. In Forty-third International Conference on Machine Learning.
- Xiaojun Wu, Cehao Yang, Honghao Liu, Xueyuan Lin, Zhichao Shi, Hao Zhou, Xuhui Jiang, Chengjin Xu, Jia Li, and Jian Guo. 2026c. Envs-FORGE: Frontier-optimized reward-grounded environment synthesis for agent RL. Preprint, arXiv:2608.14312.
- Yifan Wu, Yiran Peng, Yiyu Chen, Jianhao Ruan, Zijie Zhuang, Cheng Yang, Jiayi Zhang, Man Chen, Yenchi Tseng, Zhaoyang Yu, Liang Chen, Yuyao Zhai, Bang Liu, Chenglin Wu, and Yuyu Luo. 2026d. AutoWebWorld: Synthesizing infinite verifiable web environments via finite state machines. Preprint, arXiv:2602.14296.
- Ziqiao Xi, Shuang Liang, Qi Liu, Jiaqing Zhang, Letian Peng, Fang Nan, Meshal Nayim, Tianhui Zhang, Rishika Mundada, Lianhui Qin, Biwei Huang, and Kun Zhou. 2026. C-world: A computer use agent environment creator. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 43202–43216. Association for Computational Linguistics.
- Jingxu Xie, Dylan Xu, Xuandong Zhao, and Dawn Song. 2026. AgentSynth: Scalable task generation for generalist computer-use agents. In International Conference on Learning Representations.
- Minrui Xu, Zilin Wang, Mengyi Deng, Zhiwei Li, Zhicheng Yang, Xiao Zhu, Yinhong Liu, Boyu Zhu, Baiyu Huang, Chao Chen, Heyuan Deng, Fei Mi, Lifeng Shang, Xingshan Zeng, and Zhijiang Guo. 2026a. EnvFactory: Scaling tool-use agents via executable environments synthesis and robust RL. Preprint, arXiv:2605.18703.
- Weikai Xu, Yunren Feng, Haoxiang Lei, Kun Huang, Yuxuan Liu, Kang Zhao, Xiaolin Hu, Shuo Shang, and Bo An. 2026b. AppDeltaWorld: Transition-grounded delta code world model for mobile GUI agents. Preprint, arXiv:2608.05891.
- Tianci Xue, Zeyi Liao, Tianneng Shi, Zilu Wang, Kai Zhang, Dawn Song, Yu Su, and Huan Sun. 2026. Autonomous continual learning for environment adaptation of computer-use agents. Preprint, arXiv:2602.10356.
- Xingshan Zeng, Zishan Xu, Boju Zhang, Yuzhou Wu, Lingzhi Wang, Jianghao Lin, Liangyou Li, Yasheng Wang, Lifeng Shang, Xin Jiang, Weinan Zhang, Yong Yu, Qun Liu, and Weiwen Liu. 2026. What makes good agentic data? an ACE lens on data generation for LLM agents. Preprint, arXiv:2608.27260.
- Zhiyuan Zeng, Hamish Ivison, Yiping Wang, Lifan Yuan, Shuyue Stella Li, Zhuorui Ye, Siting Li, Jacqueline He, Runlong Zhou, Tong Chen, Chenyang Zhao, Yulia Tsvetkov, Simon Shaolei Du, Natasha Jaques, Hao Peng, Pang Wei Koh, and Hannaneh Hajishirzi. 2025. RLVE: Scaling up reinforcement learning for language models with adaptive verifiable environments. Preprint, arXiv:2511.07317.
- Chenghao Zhang, Canran Xiao, SaiSai Hu, and Dan Roth. 2026a. Training needs trustworthy worlds: Verified synthetic web environments for agent learning. Preprint, arXiv:2608.21898.
- Guibin Zhang, Hejia Geng, Xiaohang Yu, Zhenfei Yin, Zaibin Zhang, Zelin Tan, Heng Zhou, Zhongzhi Li, Xiangyuan Xue, Yijiang Li, Yifan Zhou, Yang Chen, Chen Zhang, Yutao Fan, Zihu Wang, Songtao Huang, Francisco Piedrahita-Velez, Yue Liao, Hongru Wang, and 6 others. 2026b. The landscape of agentic reinforcement learning for LLMs: A survey. Transactions on Machine Learning Research.
- Ziyun Zhang, Zezhou Wang, Xiaoyi Zhang, Zongyu Guo, Jiahao Li, Bin Li, and Yan Lu. 2026c. InfiniteWeb: Scalable web environment synthesis for GUI agent training. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 28465–28492. Association for Computational Linguistics.
- Jiarong Zhao, Zhikai Lei, Zhiheng Xi, Rui Zheng, Hang Yan, Jie Zhou, Qin Chen, and Liang He. 2026. NexForge: Scaling executable agent tasks via requirement-first synthesis. Preprint, arXiv:2607.14186.
- Kaijie Zhu, Yuzhou Nie, Yijiang Li, Yiming Huang, Jialian Wu, Jiang Liu, Ximeng Sun, Zhenfei Yin, Lun Wang, Zicheng Liu, Emad Barsoum, William Yang Wang, and Wenbo Guo. 2026. TermiGen: High-fidelity environment and robust trajectory synthesis for terminal agents. Preprint, arXiv:2602.07274.
- Yiqi Zhu, Apurva Gandhi, and Graham Neubig. 2025. Training versatile coding agents in synthetic environments. Preprint, arXiv:2512.12216.
Figure 1.
Environments and tasks form interactive learning resources. Synthesizing them at scale supports agent interaction that yields training rewards and reusable experience for self-improvement.
Figure 1.
Environments and tasks form interactive learning resources. Synthesizing them at scale supports agent interaction that yields training rewards and reusable experience for self-improvement.

Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.