Preprint
Review

This version is not peer-reviewed.

Scaling Interactive Learning Resources for LLM Agents: A Survey of Environment and Task Synthesis

  † Equal contribution.

Submitted:

18 September 2026

Posted:

20 September 2026

You are already at the latest version

Abstract
Large language model (LLM) agents can learn through interaction with external environments, obtain training signals and reusable experience from task execution. Scaling this form of learning requires a sufficient supply of interactive environments and tasks grounded in those environments, with explicit objectives and completion criteria. We refer to these complementary resources as interactive learning resources. This survey reviews their automated construction at scale and continuous adaptation through three interconnected themes: environment synthesis, task synthesis, and learning-driven evolution of both. For environment synthesis, we examine functional components, construction paradigms, and intrinsic quality assurance. For task synthesis, we review intent formation, grounding, completion verification, and task quality assurance. We then examine how a target agent's performance and learning progress inform assessments of resource suitability and guide adaptation to its evolving capabilities and learning needs across successive learning rounds. Finally, we discuss open challenges concerning the order of environment and task synthesis, the trade-off between intrinsic quality and learning utility, long-term resource evolution, and quality assurance at scale.
Keywords: 
;  ;  

1. Introduction

Large language model (LLM) agents extend language models from passive text generation to systems that can perceive, reason, and act through interactions with external environments. As agents are deployed for complex interactive tasks, enabling them to improve through experience has become a central research problem. Interaction allows an agent to attempt a task, observe the consequences of its actions, and obtain evidence about its execution. Such evidence provides two complementary sources of improvement: (i) execution outcomes can supply training signals, such as rewards for reinforcement learning; and (ii) trajectories preserve successful strategies, failures, and intermediate decisions that can be distilled into reusable experience to support self-improvement(Chen et al. 2026b; Dong et al. 2026; Guo et al. 2026b).
Interactive Learning Resources Learning through interaction requires both environments and tasks. Environments provide the states, dynamics, interfaces, and runtimes in which agents can act; tasks specify what agents should accomplish, under which concrete conditions, and how completion is assessed. As shown in Figure 1, we use interactive learning resources as an umbrella term for these two complementary resource types. An environment can support many tasks, and task synthesis can reuse an existing environment or accompany the construction of a new one. Their shared role in supporting goal-directed practice motivates studying their construction together.
Scalable Resource Construction Constructing interactive learning resources at scale through manual effort remains challenging. Environments are costly to build, deploy, reset, and maintain, while manually authored tasks offer limited coverage of objectives, states, and workflows. These constraints motivate large-scale automated synthesis across tool use, web navigation, GUI interaction, terminal interaction, and software engineering (Hu et al. 2025; Song et al. 2026; Sullivan et al. 2025; Tu et al. 2026; Xi et al. 2026). We organize environment synthesis around three primary paradigms for constructing reusable interaction substrates: model-based simulation, reconstruction and adaptation of existing systems, and creation from scratch. We structure task synthesis around three core components: task intent synthesis, environment grounding and task instantiation, and task completion verification. Together, these construction approaches expand the availability and diversity of interactive learning resources, providing broader opportunities for agents to develop capabilities across varied settings and workflows.
However, expanding resource supply is useful only if quality can be maintained as construction scales. Unlike small collections that can be individually curated and inspected, large-scale automated synthesis cannot depend on exhaustive manual checking. It requires scalable mechanisms to ensure that environments remain executable, semantically coherent, and robust, that tasks are grounded and solvable, and that the resulting collections provide sufficient diversity, structural complexity, and fidelity. To this end, execution-based testing, consistency checking, filtering, and repair therefore become integral parts of synthesis pipelines rather than optional checks after generation (Shi et al. 2026a; Tu et al. 2026; Zhang et al. 2026a). Accordingly, Section 3 and Section 4 examine both the technical routes for automated construction and the quality requirements that synthesized environments and tasks must satisfy.
Learning-Driven Resource Evolution Construction does not end with an initial pool of valid environments and tasks. As a target agent learns through interaction, its changing capabilities create new requirements for what should be synthesized next. This introduces a learner-dependent dimension of resource assessment: beyond whether environments and tasks satisfy their intrinsic quality requirements, the question becomes how well they match the agent’s current capabilities and learning needs. Success rates, failure patterns, and learning progress can help identify tasks that are already mastered or remain too difficult, capabilities that receive insufficient practice, and limitations in what the available environments allow the agent to learn (Dong et al. 2026; Huang et al. 2026a; Kang et al. 2026; Meng et al. 2026). The agent’s interaction evidence thus provides a basis for assessing which resources should be retained, revised, or expanded.
These assessments can guide subsequent synthesis by generating new objectives, adjusting task distributions, modifying environment states, tools, or dynamics, and constructing additional environments. Repeating interaction, assessment, adaptation, and learning allows resource construction to respond to the learner’s evolving needs rather than only expanding a fixed collection. The technical challenge is to translate agent feedback into appropriate construction decisions while preserving resource quality and broad coverage across successive rounds. Section 5 examines this learning-driven evolution of tasks and environments, in which agent-relative assessment provides signals for deciding what to construct or revise next. Such signals guide adaptation, while their value ultimately depends on whether the resulting resources support further agent improvement.
Survey Overview Building on this perspective, this survey provides a structured account of how interactive learning resources are constructed, assessed, and continually adapted. Section 3 reviews environment synthesis through its functional components, major construction paradigms, and intrinsic quality requirements. Section 4 examines task synthesis from intent formation and environment grounding to task-completion verification and task quality assurance. Section 5 then extends the analysis beyond initial construction to learning-driven resource evolution, studying how evidence from a target agent supports capability estimation, agent-relative resource assessment, and subsequent task and environment adaptation. In Section 6, we finally discuss open challenges concerning synthesis order, quality assessment, learning utility, long-term evolution, and quality assurance at scale.
Relation to Existing Surveys Li et al. (2026b) study agentic environment engineering through environment modeling, synthesis, evaluation, and applications, including task synthesis and evolution within that broader lifecycle. Zeng et al. (2026) study agentic data generation, organizing environments, task signals, trajectories, and verifiers through the Accuracy–Complexity–divErsity (ACE) lens. Complementary to these perspectives, we focus on the technical routes and requirements for the large-scale automated construction of interactive learning resources. We organize environment synthesis and task synthesis as two complementary subjects, examine the quality assurance needed to make their outputs usable at scale, and extend the analysis to continuing resource evolution informed by agent learning. Detailed agent-learning and self-improvement algorithms are covered by complementary surveys (Gao et al. 2026; Zhang et al. 2026b).

2. Preliminaries

This section defines interactive learning resources, distinguishes their construction from their execution, and clarifies the assessment roles and scope of the survey.
Interactive Learning Resources We use interactive learning resources to refer to the environments and tasks that support LLM agents in learning through goal-directed interaction. Environments supply the conditions and capabilities for interaction, while tasks provide objectives, execution conditions, and completion criteria. They can be constructed separately or jointly and reused across multiple executions. A single environment can support many tasks; a task intent can also be instantiated in different environments, provided its objects, parameters, and requirements are grounded in each one.
Environment An environment is the interactive substrate within which an agent acts and observes the consequences of its actions. It maintains the state relevant to interaction, determines how that state changes, exposes actions and observations through agent-facing interfaces, and provides the runtime in which interaction is executed. It determines what can happen. For a compact formal view, we write
E = ( S , A , O , P , R ) ,
where S is the state space, A the available actions, O the observations exposed to the agent, P the transition dynamics, and R the runtime substrate.
Task A task specifies what an agent should accomplish within a concrete environment. We write
T = ( g , s 0 , c , v ) ,
where g is the objective, s 0 the initial condition, c the grounded parameters and constraints, and v the intended completion semantics. For the feedback-driven learning considered here, task construction must also supply or identify a procedure V that operationalizes these semantics. This procedure may be generated with the task, instantiated from a template, or inherited from an existing environment.
Resource Execution and Verification Executing a task T in environment E with policy π produces an interaction trajectory x, which a verifier assesses:
x p π ( · E , T ) , y = V ( x ; E , T ) .
Here, x records the observations, actions, and available execution evidence, and y is a completion judgment such as a success label, score, or step-wise assessment. The task semantics v specify what completion means, whereas V implements its assessment. The resulting trajectories and feedback are products of using the resources, rather than the resources themselves. They may supply training rewards, experience-selection feedback, or evidence for subsequent resource adaptation.
Scope This survey focuses on the large-scale automated synthesis of environments and tasks for LLM agent learning and evolution. We include methods that create interactive environments, reconstruct or adapt existing systems, simulate environment behavior, or generate task intents, grounding, execution conditions, and completion checks. We also examine how target-agent evidence guides the assessment and adaptation of these resources over learning rounds, with environment-centric evolution schedules providing a complementary point of comparison.
Our primary objects are the reusable environments and tasks made available for interaction. Trajectory synthesis and rollout collection are included when they contribute mechanisms for constructing these resources. Hand-crafted benchmarks and static instruction collections provide adjacent context unless they introduce scalable construction mechanisms. Literature collection is summarized in Appendix A, and Appendix B provides the literature mappings.
Figure 2. Organization of the survey around interactive learning resources: environment synthesis (Section 3), task synthesis (Section 4), and learning-driven task and environment evolution (Section 5).
Figure 2. Organization of the survey around interactive learning resources: environment synthesis (Section 3), task synthesis (Section 4), and learning-driven task and environment evolution (Section 5).
Preprints 233934 g002

3. Environment Synthesis

Environment synthesis concerns the automatic and scalable construction of reusable environments that can sustain continued agent interaction. This section presents their functional components and synthesis targets, organizes construction paradigms by how the interaction substrate is obtained, and examines intrinsic environment quality.

3.1. Environment Components

From the perspective of agent interaction, an environment needs to specify what the world currently contains, how it can evolve, what the agent can observe and manipulate, and how such interaction is executed. Accordingly, we characterize an environment through four components:
  • Environment State represents the persistent configuration of the world at a given point in interaction, including the entities, resources, relations, and application states relevant to future actions. It should preserve the consequences of previous interactions so that later decisions are conditioned on the evolving world state.
  • Environment Dynamics defines how the world state can evolve, including the rules, constraints, and side effects that govern state transitions. These transitions may be triggered by agent actions as well as by internal processes, external events, or other actors.
  • Interaction Tools specifies how the agent observes and operates on the environment through interfaces such as tools, GUIs, browsers, or shells. It determines the agent-facing observation and action space.
  • Runtime Infrastructure provides the execution substrate that makes the environment usable in practice. It supports reliable initialization, execution, state persistence, reset, isolation, and repeated interaction at scale.

3.2. Environment Synthesis Paradigms

According to how the underlying interaction substrate is obtained, we categorize existing approaches into three paradigms: (i) Model-based Environment Simulation, which replaces a fully implemented backend with a generative model that produces environment responses and transitions during interaction; (ii) Existing-System Reconstruction and Adaptation, which converts existing systems, artifacts, or interaction traces into reusable environments; and (iii) Environment Creation from Scratch, which constructs a new executable environment from seeds or specifications.
Paradigm 1: Model-based Environment Simulation This paradigm obtains an interactive environment by configuring a generative model as its response and transition mechanism:
E = F sim ( M , z ) ,
where M is the generative simulator, z denotes optional environment specifications such as tool descriptions or API schemas, and F sim configures the model with the interfaces and runtime needed for repeated interaction. During execution, M produces responses or predicts state transitions from the interaction history and the current action. The simulated state may remain implicit rather than being exposed as an accessible state object. Two common, non-exclusive implementation choices are specification grounding, which conditions simulation on tool descriptions, API schemas, or task context (Lee et al. 2026; Li et al. 2025), and learning from interaction data, which distills response behavior or transition dynamics (Chen et al. 2026b; Ding et al. 2026; Xu et al. 2026b). A simulator can combine both: MirrorAPI uses API descriptions together with a specialized model trained on request–response examples (Guo et al. 2025). Hybrid systems can additionally combine live execution with synthesized world engines (Xi et al. 2026). This paradigm scales interaction without implementing every state transition, but its black-box transitions are difficult to inspect and can accumulate semantic or state inconsistencies over long horizons. These limitations make deterministic execution, debugging, and reliable reward attribution harder than in explicitly executable environments.
Paradigm2: Existing-System Reconstruction and Adaptation This paradigm obtains an interactive environment by reconstructing or adapting semantics already present in software systems, applications, repositories, or prior interactions:
E = F adapt ( B , H ) ,
where B denotes an existing system or its artifacts, H denotes optional execution histories, and F adapt reconstructs or adapts these sources into a reusable environment. The resulting environment inherits or recovers interaction semantics from an existing source. One family focuses on runtime adaptation, resolving dependencies, containerizing software, or exposing authentic services through stable interfaces (Guo et al. 2026a; Xu et al. 2026a). TerminalTraj also constructs Dockerized environments from repositories before generating environment-aligned tasks and trajectories (Wu et al. 2026b). A second family performs system reconstruction, reproducing existing applications or websites as lightweight, resettable, and verifiable replicas (Cao et al. 2026; Chae et al. 2026). More recent work reconstructs executable workspaces from environment histories or trajectories: CLI-Gym inverts healthy environment histories into failure states, while Terminal-Universe recovers reusable terminal environments from accumulated tool-execution traces (Lin et al. 2026; Wu et al. 2026a). These routes inherit useful semantics and fidelity from their sources, but their coverage remains constrained by the systems or traces available for reconstruction and by the cost of adapting heterogeneous dependencies.
Paradigm3: Environment Creation from Scratch This paradigm obtains an interactive environment by constructing a new implementation from a seed or specification:
E = F create ( z ) ,
where z denotes a scenario description, capability specification, schema, or other design input, and F create materializes the environment’s state, dynamics, interfaces, and runtime. Recent systems span planning worlds, tool/API ecosystems, Web applications, terminal containers, software repositories, and computer-use environments (Cai et al. 2025; Fang et al. 2026; Hu et al. 2025; Sullivan et al. 2025; Tu et al. 2026; Wang et al. 2026a; Wu et al. 2026d; Zhu et al. 2025). A complete synthesis pipeline typically contains five stages.
Seed and environment specification. The intended structure and functionality of the environment can be anchored by semantic seeds such as domains or business scenarios, capability-grounded seeds such as schemas, protocols, tool ecosystems, or skill structures, and reference-grounded seeds such as source applications, code artifacts, or interface designs (Cai et al. 2025; Hu et al. 2025; Jeong and Yoon 2026; Shi et al. 2026b; Song et al. 2026; Zhang et al. 2026c). The choice of seed controls both the freedom of synthesis and how strongly later states, dynamics, and interfaces inherit externally defined semantics.
State construction and persistence. The specification must then be instantiated as explicit state that can be read, modified, persisted, and reset. Structured databases, object stores, application backends, or generated project files provide an external source of truth rather than leaving state implicit in model context (Fang et al. 2026; Song et al. 2026; Tu et al. 2026; Wang et al. 2026c; Xi et al. 2026; Zhu et al. 2025). State lifecycle mechanisms further support inspection, isolation, and repeated episodes, as in task-specific state injection or session-scoped state management (Wang et al. 2026a).
Dynamics construction. Executable environments realize dynamics by translating preconditions, update rules, side effects, and constraints into programmatic operations over state. Local transitions may be implemented as functions or database transactions, while interconnected environments require shared-state constraints and cross-service invariants (Jeong and Yoon 2026; Song et al. 2026; Tu et al. 2026; Wang et al. 2026c; Wu et al. 2026d). Programmatic state machines and controllers provide an alternative way to make transition structure explicit and verifiable (Wu et al. 2026d; Xi et al. 2026).
Tool and interface synthesis. Internal capabilities must be exposed through actions that the agent can invoke. Methods range from synthesizing tool schemas and implementations, to binding APIs to persistent backends, to inheriting protocol constraints from existing tool ecosystems (Castellani et al. 2025; Fang et al. 2026; Shi et al. 2026b; Sullivan et al. 2025; Wang et al. 2026c). The common objective is to align declared interfaces with meaningful, state-changing environment behavior.
Runtime packaging. Finally, state, dynamics, and interfaces are assembled into runtimes that support initialization, isolation, reset, and repeated execution. Terminal and software environments typically use containers or sandboxes, whereas Web environments package self-contained frontends and backends (Gandhi et al. 2026; Peng et al. 2026; Shen et al. 2026; Wu et al. 2026d; Zhang et al. 2026c; Zhu et al. 2026; Zhu et al. 2025).

3.3. Environment Quality Assurance

Quality assurance assesses the environment itself and the environment collection as reusable learning resources. A large collection can remain unreliable or limited if instances fail to execute, implement inconsistent transitions, remain highly redundant, support only trivial interactions, behave unreliably, or deviate substantially from the intended setting. We summarize six intrinsic dimensions: executability, semantic correctness, diversity, complexity, robustness, and fidelity.
Executability Executability asks whether a generated environment can be successfully built, launched, interacted with, and reset. Existing pipelines combine static checks, runtime testing, executable probing, and repair loops before environments are admitted for training (Cai et al. 2025; Guo et al. 2026a; Shen et al. 2026; Shi et al. 2026a; Tu et al. 2026; Wu et al. 2026d; Zhu et al. 2026).
Semantic Correctness Semantic correctness requires environment behavior to conform to intended interface, transition, and system-level semantics. Assurance mechanisms include specification–implementation auditing, transition probing, explicit invariants, finite-state constraints, and backend consistency checks (Castellani et al. 2025; Jeong and Yoon 2026; Song et al. 2026; Tu et al. 2026; Wu et al. 2026d; Zhang et al. 2026a).
Diversity Diversity measures whether synthesis broadens the available environments and interaction capabilities rather than producing superficial variants. Methods diversify semantic domains, tool/API structures, state schemas, workflow compositions, and whole environment ecosystems, often with explicit coverage objectives or open-ended discovery (Castellani et al. 2025; Dong et al. 2026; Fang et al. 2026; Song et al. 2026; Sullivan et al. 2025; Tu et al. 2026).
Complexity Complexity characterizes the structural dependencies an environment can support, including action-level composition, persistent state coupling, cross-service dependencies, and navigational or programmatic structure (Jeong and Yoon 2026; Shi et al. 2026b; Sullivan et al. 2025; Xi et al. 2026; Zhang et al. 2026c).
Robustness Robustness concerns whether an environment remains stable and recoverable across repeated and concurrent interaction. Isolation, deterministic reset, snapshots, session-scoped state, monitoring, and incremental repair reduce contamination and infrastructure-induced variation across rollouts (Cao et al. 2026; Guo et al. 2026a; Shen et al. 2026; Wang et al. 2026a).
Fidelity Fidelity concerns how faithfully a synthetic environment preserves the properties of the systems or scenarios it is intended to represent. Grounding may occur at protocol/schema level, through direct system or website reconstruction, or through authentic resources and reference designs (Cao et al. 2026; Chae et al. 2026; Shi et al. 2026b; Xu et al. 2026a; Zhang et al. 2026c). Recent evidence further shows that structural and semantic defects in synthetic worlds can materially reduce downstream learning utility, motivating explicit verification rather than relying on surface realism alone (Zhang et al. 2026a).

4. Task Synthesis

Task synthesis constructs the objectives and execution specifications through which agents use environments for goal-directed learning. We first discuss Task Intent Synthesis, which forms objectives, and then Environment Grounding and Task Instantiation, which binds them to concrete states, resources, parameters, and runtime conditions. Task Verification supplies procedures that assess whether an execution meets the objective and constraints, making completion checks part of task construction. Task Quality Assurance instead evaluates the resulting tasks and task collection as resources.

4.1. Task Intent Synthesis

Task intent synthesis forms the objective that the agent is expected to accomplish. We separate intent formation from full task instantiation for analysis. We organize existing approaches according to the primary source from which the task intent is derived, and discuss seed-driven, tool-driven, state-driven, and trajectory-driven methods in turn.
Seed-driven Seed-driven methods start from externally provided descriptions of what kinds of tasks should be generated, such as seed instructions, capability requirements, domain descriptions, personas, or a small number of exemplars. Rather than only paraphrasing seeds, recent methods structure them into task profiles, hierarchical compositions, or difficulty-controlled design spaces (Hua et al. 2026; Pan et al. 2026; Pi et al. 2026; Shi et al. 2025; Xie et al. 2026; Zhao et al. 2026). This route offers direct control over the desired task distribution, but feasibility must still be established when the intent is grounded into a concrete environment.
Tool-driven Tool-driven methods derive task intents from the affordance structure exposed by an environment. The underlying action space may be organized as a flat tool collection, a dependency graph, a workflow/program structure, or a skill graph, and task synthesis composes objectives over these structures. Representative mechanisms include tool-sequence exploration, protocol or dependency graphs, multi-skill composition, skill-graph path sampling, and topology-aware tool sampling (Cheng et al. 2026; Dong et al. 2026; Fan et al. 2026a; Keren et al. 2026; Shi et al. 2026b; Tan et al. 2026; Tu et al. 2026; Xu et al. 2026a). The defining signal is therefore not merely which tools exist, but how their capabilities can be composed into meaningful workflows.
State-driven State-driven methods formulate task intents from a concrete environment state. Once an environment has been instantiated with specific entities, resources, records, or configurations, the generator uses this state to determine objectives that are meaningful and feasible under current conditions. Database- and backend-grounded systems generate scenarios from instantiated state and available operations, thereby avoiding nonexistent entities or violated preconditions (Song et al. 2026; Tu et al. 2026; Wang et al. 2026c; Xi et al. 2026).
Trajectory-driven Trajectory-driven methods obtain task intents from interaction traces that have already been executed in the environment. They first explore or execute the environment, observe a path of actions and state transitions, and then abstract or reverse-engineer a task from this evidence (Chen et al. 2026a; Murty et al. 2024; Ramrakhya et al. 2026; Sun et al. 2025; Wang et al. 2026d). Terminal-Universe extends this idea further by reconstructing a reusable environment from a trajectory before exploring the recovered workspace for new tasks (Wu et al. 2026a). Because the proposed intent is supported by execution evidence, trajectory-driven synthesis offers strong grounding, although its coverage is ultimately constrained by what the exploration or trajectory source exposes.

4.2. Environment Grounding and Task Instantiation

Once a task intent has been formed, it must be instantiated with concrete objects, parameters, and initial conditions that actually exist in the environment. This grounding process is strongly shaped by the underlying interaction substrate. We therefore organize existing methods into two broad settings: (i) Web/Service Environments, where tasks are grounded in websites, online services, and backend states; and (ii) Computer/Software Environments, where tasks are instantiated over applications, filesystems, repositories, and executable workspaces. Within each setting, different interfaces—such as browser actions, APIs, GUIs, or command lines—further determine how the task is concretely instantiated.
Web/Service Environments For tools and services, grounding mainly requires selecting executable capabilities, resolving valid entities and arguments, and satisfying dependencies among successive calls. These constraints are often derived from schemas, tool graphs, backend state, or explicitly synthesized transition structures (Chen et al. 2026a; Fang et al. 2026; Shi et al. 2026b; Song et al. 2026; Tu et al. 2026; Wang et al. 2026c; Xu et al. 2026a). Browser-based Web environments additionally require tasks to align with reachable pages, sessions, navigation structure, and persistent backend state (Chae et al. 2026; Huang et al. 2026b; Murty et al. 2024; Wang et al. 2026b; Wu et al. 2026d). Across both cases, grounding binds abstract objectives to resources and interaction paths that the environment can actually support.
Computer/Software Environments Computer and software tasks require a concrete executable workspace around the intended objective. GUI settings instantiate applications, files, user data, and task-specific application states, while terminal and software-engineering settings materialize repositories, dependencies, services, containers, and tests (Hua et al. 2026; Lin et al. 2026; Lv et al. 2026; Ramrakhya et al. 2026; Shi et al. 2026a; Wang et al. 2026a; Wu et al. 2026b; Zhao et al. 2026). More recent pipelines generate or reconstruct complete workspaces and their verifiers from scratch, from source skills, or from prior trajectories (Du et al. 2026; Gandhi et al. 2026; Li et al. 2026c; Pan et al. 2026; Shen et al. 2026; Wu et al. 2026a; Zhu et al. 2026; Zhu et al. 2025). Grounding ultimately requires aligning the task with a realizable software state and a reproducible execution context.

4.3. Task Verification

Task completion verification assesses a particular execution against the intended objective and constraints. We characterize verification along three complementary design dimensions: when verification is performed, what evidence is inspected, and how the final judgment is made.
Formally, we characterize a verifier as
V = ( t , e , j ) , y = V ( x ; E , T ) ,
where t specifies when verification is performed, e specifies what evidence is inspected, and j specifies how that evidence is converted into a judgment y such as a success label, score, or reward.
Final Outcome vs. Step-wise Verification Final verification determines success from the outcome observed after an interaction terminates and is appropriate when task semantics are fully captured by the final state or output (Hua et al. 2026; Lei et al. 2026; Tu et al. 2026; Wang et al. 2026a). Step-wise verification instead evaluates selected intermediate conditions when correctness depends on mandatory checkpoints, ordering constraints, forbidden operations, or side effects (Shi et al. 2026a; Song et al. 2026; Wu et al. 2026d). It provides denser supervision but requires more explicit specification of which intermediate behaviors matter.
Evidence Sources and Checking MechanismsState-based verification checks structured environment states through predicates, state differences, invariants, or target-state matching (Chae et al. 2026; Lei et al. 2026; Tu et al. 2026; Wang et al. 2026a). Path-based verification examines action sequences and intermediate states when correctness depends on how an outcome is reached (Chen et al. 2026a; Wang et al. 2026b). Test-based verification instead describes an executable checking mechanism: unit tests, integration tests, compilers, or shell scripts may inspect states, outputs, paths, or several of these together. Such checking is particularly common in terminal and software environments (Gandhi et al. 2026; Hua et al. 2026; Shen et al. 2026; Shi et al. 2026a; Wu et al. 2026b; Zhu et al. 2026; Zhu et al. 2025).
Deterministic and Rubric-based JudgmentDeterministic verification uses rules, assertions, exact matching, state predicates, or executable tests to produce reproducible judgments, whereas rubric-based verification uses LLM or multimodal judges for objectives that are difficult to formalize programmatically (Cao et al. 2026; Hua et al. 2026; Lei et al. 2026; Lv et al. 2026; Tan et al. 2026; Xi et al. 2026; Xue et al. 2026). The former offers stronger auditability; the latter extends coverage to open-ended semantics but introduces calibration, consistency, and reward-exploitation risks.
Verifier Reliability and Downstream Uses A verifier should faithfully reflect the intended task semantics, remain robust to alternative valid solutions, and avoid exploitable mismatches between the objective and checking procedure. Reliability failures arise when rules or tests capture only part of the intended behavior, when path constraints reject valid alternatives, or when learned judges are inconsistent (Bercovich 2026; Chen et al. 2026a; Shi et al. 2026a; Wang et al. 2026b). Large-scale synthesis therefore increasingly couples verifier generation with executable validation, adversarial checking, or consistency auditing(Meng et al. 2026; Shen et al. 2026; Tu et al. 2026; Zhang et al. 2026a).
A verifier should provide machine-usable assessments of task execution, such as success/failure labels, scalar scores, or step-wise judgments, optionally accompanied by diagnostic feedback on unmet criteria. These outputs can be used directly or transformed into (i) reward signals for reinforcement learning; (ii) feedback and selection signals for agent self-evolution, supporting trajectory selection, experience extraction, and the evaluation of candidate memory, skill, or workflow updates; and (iii) capability- and resource-assessment signals, when aggregated with interaction evidence to assess difficulty, capability gaps, and learning suitability for the target agent. The third use guides task and environment adaptation in Section 5; an execution score is not itself a score of the task’s intrinsic quality or learning utility. For detailed agent-side learning and update mechanisms, we refer readers to surveys on agentic reinforcement learning (Zhang et al. 2026b) and self-evolving agents (Gao et al. 2026).

4.4. Task Quality Assurance

Task quality assurance evaluates the generated task and task collection, rather than judging a particular execution. It asks whether a valid solution exists under the specified conditions, whether the collection covers diverse objectives and workflows, how structurally complex its tasks are, and how faithfully they reflect the intended uses. We organize these intrinsic properties as Solvability, Diversity, Complexity, and Fidelity.
Solvability Solvability requires that at least one valid solution exists under the given environment and initial conditions. Evidence can come from reachability analysis, reference solutions, successful rollouts, golden-state construction, and solver-based calibration (Lei et al. 2026; Li et al. 2026c; Meng et al. 2026; Pan et al. 2026; Shen et al. 2026; Shi et al. 2026a; Tu et al. 2026). An executable test suite alone does not establish solvability: it must be paired with a valid witness or other feasibility evidence, and a solver failure is not a proof that no solution exists.
Diversity Diversity concerns whether tasks cover new capabilities, workflows, tool combinations, state contexts, or interaction patterns rather than merely paraphrasing instructions. Methods use semantic deduplication, skill/tool coverage, structured combination sampling, graph/path sampling, persona variation, and coverage-guided generation (Chen et al. 2026a; Fan et al. 2026a; Ivison et al. 2026; Keren et al. 2026; Shi et al. 2025; Tan et al. 2026; Xie et al. 2026).
Complexity Complexity describes structural dependencies within a task, such as subgoal coupling, cross-tool or cross-application dependencies, persistent-state span, dependency depth, branching, delayed consequences, and recovery requirements. Existing methods increase complexity through compositional task expansion, graph-based workflow construction, multi-hop grounding, or recursive task/environment rewriting (Fan et al. 2026a; Huang et al. 2026b; Keren et al. 2026; Li et al. 2026c; Shi et al. 2025; Xie et al. 2026).
Fidelity Task fidelity concerns whether synthetic tasks reflect real user needs, realistic workflows, and the target-domain task distribution. Assurance mechanisms include real-demand discovery, grounding in logs or websites, preserving source intent, authentic-resource exploration, and synthetic-to-real transfer evaluation (Chae et al. 2026; Huang et al. 2026b; Shi et al. 2026a; Xu et al. 2026a; Zhao et al. 2026). Fidelity is alignment between generated objectives and the workflows agents are expected to encounter.
Figure 3. Development of scalable environment synthesis, task synthesis, and learning-driven resource evolution for LLM agents.
Figure 3. Development of scalable environment synthesis, task synthesis, and learning-driven resource evolution for LLM agents.
Preprints 233934 g003

5. Learning-Driven Interactive Learning Resource Evolution

Section 3 and Section 4 examine the construction and intrinsic quality of interactive learning resources. Once a target agent uses these resources, its behavior provides a further basis for evaluation: how well the available environments and tasks match its current capabilities and learning needs. The same interactions can inform both capability estimation, which characterizes the learner, and agent-relative resource assessment, which characterizes the suitability of its practice. The central agent-conditioned loop uses these assessments to adapt environments or tasks, then reassesses them after further interaction and learning. This section focuses on that feedback-driven process.
For agent-conditioned evolution, a compact view is
c t = F ( H t ) , q t = Q ( E t , T t , H t , c t ) ,
( E t + 1 , T t + 1 ) = U ( E t , T t , c t , q t ) ,
followed by learning on interactions from the updated resources,
π t + 1 = L ( π t , E t + 1 , T t + 1 ) .
Here, E t and T t denote the environment and task collections, and H t contains the target policy’s interaction evidence and verification outcomes. The estimate c t characterizes current capability, while q t summarizes resource assessments such as learner-relative difficulty, failure-prone regions, and coverage of learning needs. The assessments guide adaptation U, while L denotes the subsequent learning update.
Capability Estimation Agent-conditioned evolution estimates the learner’s capabilities from task success rates, learning progress, and contrasts between successful and failed trajectories or heterogeneous solvers. These signals can reveal missing capabilities, identify failure-prone regions, or locate a solver-relative learnable frontier (Dong et al. 2026; Huang et al. 2026a; Kang et al. 2026; Meng et al. 2026). Read from the resource side, the same evidence indicates which tasks are already mastered, which expose current weaknesses, and which environment capabilities remain underused or insufficiently represented in training. Task difficulty is therefore agent-relative, unlike structural complexity or the existence of a valid solution. Agent-relative resource assessment turns these observations into judgments about what the learner should encounter next. Task-level signals can guide the selection or revision of individual objectives, whereas evidence aggregated over tasks and workflows can identify which environments offer suitable practice or need expansion. The purpose is to estimate learning relevance: low success can reflect a capability gap, an invalid environment or task, or an unreliable verifier.
Task and Environment Adaptation Given capability estimates and resource assessments, the synthesis system decides which learning resources should change. Task-side adaptation can generate new objectives, compose additional subgoals, vary conditions, or shift sampling toward the current frontier (Lv et al. 2026; Meng et al. 2026; Wu et al. 2026c; Xue et al. 2026). Environment-side adaptation changes states, tools, dynamics, or whole executable worlds when the current substrate does not expose the capabilities required for further practice (Huang et al. 2026a; Kang et al. 2026; Liu et al. 2026; Shen et al. 2026; Shi et al. 2026c). Related benchmark-construction work such as ProEvolve makes environment changes programmable through graph transformations (Li et al. 2026a); it supplies an adaptation mechanism without itself demonstrating a continual policy-training loop.
Continuous Evolution Repeated interaction and learning change both capability estimates and resource suitability. A task that previously exposed a useful weakness may become routine, while newly available environment capabilities may require additional tasks. Agent-conditioned evolution therefore uses current-policy rollouts, verifier rewards, or capability gaps to regenerate tasks or reshape environments near the moving learning frontier (Dong et al. 2026; Guo et al. 2026b; Huang et al. 2026a; Liu et al. 2026; Shen et al. 2026; Wu et al. 2026c). Resource assessment is repeated within this process. RLVE provides a related precursor: its manually engineered verifiable environments procedurally generate problems and adapt their difficulty distributions to the policy’s evolving capabilities (Zeng et al. 2025).

6. Discussion and Future Directions

The preceding sections show that scaling interactive learning resources involves more than generating larger numbers of environments and tasks. Several broader questions remain unresolved regarding how these artifacts should be constructed, evaluated, evolved, and maintained.
Synthesis Order A fundamental question is how environment, task, and verification should be ordered during construction. Environment-first pipelines provide strong grounding because tasks are generated from capabilities and states that already exist, and one environment can support many tasks (Hu et al. 2025; Lei et al. 2026; Ramrakhya et al. 2026; Shi et al. 2026a; Wang et al. 2026b). Requirement-first approaches instead begin from desired capabilities or task demands and construct the necessary environment afterward (Zhao et al. 2026), while shared-specification approaches derive multiple artifacts from common primitives to improve cross-artifact consistency (Cheng et al. 2026; Ivanov and Rana 2026; Wang et al. 2026a).
Intrinsic Quality vs. Learning UtilitySection 3.3 and Section 4.4 summarize intrinsic quality dimensions for synthesized environments and tasks, while Section 5 examines learner-relative signals used to assess and adapt them. Two questions remain open: (i) how can intrinsic properties and learner-relative suitability be measured accurately and consistently at scale; and (ii) to what extent do improvements in these assessments predict actual learning gains? Existing studies provide initial evidence that environment diversity, executable correctness, and defect repair can materially affect downstream generalization (Sullivan et al. 2025; Tu et al. 2026; Zhang et al. 2026a), but controlled studies that isolate the causal contribution of individual quality dimensions remain limited.
Long-Term Task–Environment EvolutionSection 5 considers how agent feedback can drive repeated task and environment adaptation, but long-term evolution introduces additional control problems. A synthesis system must decide which mastered interactions should be retired or retained, which rare failures remain valuable, when to expand environment capabilities rather than only generate new tasks, and how to combine agent-conditioned evolution with environment-centric exploration without narrowing coverage over time.
Quality Assurance at Scale Quality assurance becomes harder as collections of interactive learning resources grow and change. Updates to states, interfaces, dynamics, or tests can invalidate previously accepted tasks or verifiers, while repeated synthesis can introduce duplication, stale checks, or reward loopholes. Execution-driven repair, auditing, dependency-aware updates, and adversarial verifier checks provide useful building blocks (Bercovich 2026; Castellani et al. 2025; Guo et al. 2026a; Li et al. 2026a; Shi et al. 2026a), but scalable synthesis ultimately requires persistent maintenance.

7. Conclusion

This survey reviews the large-scale automated construction and continuing adaptation of interactive learning resources for LLM agents. Environments supply the conditions for action, while tasks specify objectives, execution conditions, and completion criteria. We organize environment synthesis around components, construction paradigms, and intrinsic quality assurance, and task synthesis around intent formation, grounding, completion verification, and task quality. When a target agent uses the resources, its performance and learning progress enable a further, agent-relative assessment of their suitability, which guides task and environment evolution. This resource perspective connects the two synthesis subjects through their common purpose: providing reliable and continually relevant opportunities for agent learning through interaction.

Appendix A. Literature Collection

The literature scope covers work publicly available up to September 16, 2026. Candidate papers were collected from public scholarly sources, including arXiv, ACL Anthology, OpenReview, and proceedings of major NLP, machine learning, and data-mining venues. Search terms combined agent environment synthesis, generation, or evolution with task synthesis, executable or verifiable environments, and terminal, Web, GUI, or tool agents. The pool was expanded through reference-list snowballing and targeted searches around recurring method families and newly emerging systems.
The core inclusion criterion is an automatic, scalable mechanism for constructing or expanding environments or grounded tasks. This includes simulation, reconstruction/adaptation, creation from scratch, grounded task generation, and repeated task/environment adaptation. Downstream use does not determine eligibility: evaluation-oriented methods are included when they contribute such construction mechanisms, while trajectory-generation systems are included for their task/environment construction components. Static instruction synthesis, fixed-task trajectory collection, and manually authored resources are treated as adjacent context unless they provide relevant scalable construction mechanisms. This collection supports a structured technical synthesis rather than a claim of exhaustive coverage. Appendix B maps representative systems to the survey taxonomy; inclusion in a table does not imply that every component of a system falls within the core scope.

Appendix B. Literature Mapping

The three tables map representative literature to Section 3, Section 4, and Section 5. A work may appear in multiple tables, and combined labels denote multiple sources or routes. The marker † denotes a boundary or precursor case.

Appendix B.1. Environment Synthesis Literature

“Model-based simulation,” “Recon./adapt.,” and “From scratch” abbreviate the three paradigms in Section 3.2. Combined routes are listed where needed; hybrid operation, recursive extension, adaptive evolution, and repair are described as mechanisms rather than additional paradigms. Here, reconstruction/adaptation can also reuse previously generated artifacts. Table A1 records how substrates are obtained.
Table A1. Representative literature on automated environment synthesis. Routes denote substrate provenance; adaptation, evolution, and repair mechanisms are described separately.
Table A1. Representative literature on automated environment synthesis. Routes denote substrate provenance; adaptation, evolution, and repair mechanisms are described separately.
Work Setting Route Primary synthesis target / mechanism Quality or verification signal
AgentGen Hu et al. 2025 Planning From scratch PDDL environment and planning-task generation from inspiration corpora Executable planning checks
gg-bench Verma et al. 2025 Games From scratch Generated game descriptions materialized as Gym environments for evaluation Game rules / executable evaluation
RandomWorld Sullivan et al. 2025 Tool From scratch Procedural generation of interactive tools and compositional worlds Programmatic checks
StableToolBench-MirrorAPI Guo et al. 2025 API Model-based simulation API-conditioned response simulation with a model trained on request–response data Simulated API consistency
Simia Li et al. 2025 Tool / API Model-based simulation Reasoning-model environment simulator conditioned on tools and history Simulated feedback
AgentScaler Fang et al. 2026 Tool / API From scratch Programmatic state, database, and tool materialization from API structures State / rule-based checks
AutoForge Cai et al. 2025 Tool / API From scratch Automated construction of simulated executable environments Executable / verifiable tasks
CodeGym Du et al. 2026 Code / tool Recon./adapt. Conversion of static coding problems into interactive Gym-style environments Programmatic rewards
SWE-Playground Zhu et al. 2025 SWE From scratch Generated software projects, repositories, and task-specific workspaces Unit tests
DreamGym Chen et al. 2026b Reasoning / tool Model-based simulation Learned experience model for synthetic transitions and adaptive challenges Real/synthetic calibration
Endless Terminals Gandhi et al. 2026 Terminal From scratch Containerized terminal environments with generated services and dependencies Executable tests
DynaWeb Ding et al. 2026 Web Model-based simulation Learned Web transition/world model Model-based transition prediction
MEnvAgent Guo et al. 2026a SWE Recon./adapt. Automated dependency resolution and Dockerized runtime construction Build/run/repair checks
Environment-free API Simulation Lee et al. 2026 API Model-based simulation Stateful API response simulation from specifications Simulated response consistency
ScaleEnv Tu et al. 2026 Tool From scratch Interactive environments expanded through tool-dependency structures Procedural/executable tests
Agent World Model Wang et al. 2026c Tool / API From scratch SQL-backed state and executable tool interfaces State-based checks
C-World Xi et al. 2026 Computer use Model-based simulation + recon./adapt. Hybrid live-API execution and World-Engine simulation for training/evaluation Mixed rule/model checks
AutoWebWorld Wu et al. 2026d Web From scratch FSM-grounded generation of executable websites FSM/programmatic verification
InfiniteWeb Zhang et al. 2026c Web From scratch Multi-page Web application synthesis from references/specifications Executable app/reset checks
GUI-GENESIS Cao et al. 2026 GUI Recon./adapt. Lightweight reconstruction of real GUI applications Code-native assertions
CLI-Gym Lin et al. 2026 Terminal Recon./adapt. Failure-state inversion and environment-intensive task packaging Executable tests
TermiGen Zhu et al. 2026 Terminal From scratch Generated terminal containers and executable workspaces Executable verifier
EnvScaler Song et al. 2026 Tool / API From scratch Programmatic state, dynamics, and interface synthesis Probing / state checks
VeriEnv Chae et al. 2026 Web Recon./adapt. Executable, resettable reconstruction of real websites Deterministic state checks
Agent-World Dong et al. 2026 Tool / API From scratch Discovery and synthesis of heterogeneous executable environments Verifiable tasks / RL feedback
Terminal-World Cheng et al. 2026 Terminal From scratch Skill-based joint construction of tasks, environments, and trajectories Executable bundles
EnvFactory Xu et al. 2026a Tool / API Recon./adapt. Discovery and validation of stateful executable tool environments Executability checks
LiteCoder-Terminal Peng et al. 2026 Terminal From scratch Long-horizon terminal environment generation and packaging Runtime checks
SynthTools Castellani et al. 2025 Tool From scratch Tool generation, simulation, and audit Tool audit
CUA-Gym Wang et al. 2026a Computer use From scratch + recon./adapt. Synthetic mock applications plus task-specific states in computer-use runtimes State / judge
Tmax Ivison et al. 2026 Terminal From scratch Large-scale diversified terminal environments and verifiers Diversified verifiers
SETA Shen et al. 2026 Terminal From scratch Automated terminal environment construction with later adaptive evolution Unified executable verifier
GAIS Shi et al. 2026b Tool / MCP From scratch Protocol/tool-dependency-grounded interaction synthesis Grounded interaction checks
Meta-Task Pan et al. 2026 Terminal From scratch Automated executable task/workspace generation Sandbox execution
RST Li et al. 2026c Terminal Recon./adapt. Recursive extension of verified seed-task artifacts and workspaces Executable tests
AppDeltaWorld Xu et al. 2026b Mobile GUI Model-based simulation Delta-code world model for GUI transitions Transition prediction
FACET Shi et al. 2026a Terminal From scratch Source-skill-grounded container construction, execution, and targeted repair Container + tests
SPADE Liu et al. 2026 General / tool From scratch Adaptive self-play generation of executable reset/step environments Executable interaction / regret
EnvHarness Huang et al. 2026a Multi-domain Recon./adapt. Programmable wrappers that reshape frozen environments Original verifier + fresh rollouts
EvoEnv Shi et al. 2026c Reasoning From scratch Adaptive generation and filtering of executable reasoning environments Oracle + staged checks
Terminal-Universe Wu et al. 2026a Terminal Recon./adapt. Environment recovery from trajectories and execution artifacts Executable workspace
Environment Evolution Fan et al. 2026b Terminal Recon./adapt. Environment-centric transformations of seed environments along difficulty directions Executable environment checks
Trustworthy Worlds Zhang et al. 2026a Web From scratch Synthetic Web environment construction with pre-training verification and repair Structural/state predicates
AgentMercury Jeong and Yoon 2026 Business / tool From scratch Persistent multi-service environments with shared state Cross-service invariants
TerminalTraj Wu et al. 2026b Terminal Recon./adapt. Repository-derived Docker environments and environment-aligned task instances Executable validation code

Appendix B.2. Task Synthesis Literature

The intent-source column in Table A2 uses Seed, Tool, State, and Trajectory as defined in Section 4.1, followed by the specific source or control mechanism. Externally supplied skill descriptions are seeds; exposed skill/tool composition structures are tool-based sources. Verification entries summarize checking mechanisms and may mix evidence sources with executable or model-based checks; they are not an exhaustive inventory of each system’s verification modes.
Table A2. Representative literature on automated task synthesis. Intent sources follow the four-source taxonomy; qualifiers explain composition, grounding, or feedback-guided control.
Table A2. Representative literature on automated task synthesis. Intent sources follow the four-source taxonomy; qualifiers explain composition, grounding, or feedback-guided control.
Work Setting Intent source / qualifier Grounding / instantiation Verification
AgentGen Hu et al. 2025 Planning Seed; generated-world grounding PDDL states and actions Executable plan
NNetNav Murty et al. 2024 Web Trajectory Browser interaction trace Execution outcome
OS-Genesis Sun et al. 2025 GUI Trajectory Executed GUI interaction Execution / state
TaskCraft Shi et al. 2025 Tool Seed; compositional expansion Multi-tool executable setting Executable trajectory
AgentSynth Xie et al. 2026 Computer use Seed; subtask composition Existing computer-use runtime Successful execution
RandomWorld Sullivan et al. 2025 Tool Tool; composition Procedurally generated tool world Programmatic
AgentScaler Fang et al. 2026 Tool / API Tool / state; domain structure Generated database state and executable APIs State / rule based
AutoForge Cai et al. 2025 Tool / API Seed; task/domain requirements Synthesized executable environment Executable / verifiable
SWE-Playground Zhu et al. 2025 SWE Seed; project proposal Generated repository/workspace Unit tests
ScaleEnv Tu et al. 2026 Tool Tool; dependency graph Generated interactive tool environment Executable action checks
Agent World Model Wang et al. 2026c Tool / API State / tool; backend structure SQL-backed environment state State based
C-World Xi et al. 2026 Computer use Seed; workflow/constraints Live or synthesized computer world Mixed rule + judge
AutoWebWorld Wu et al. 2026d Web State; goal/FSM Generated website state graph FSM / programmatic
InfiniteWeb Zhang et al. 2026c Web Seed; reference design Synthesized Web application Executable app/reset
CLI-Gym Lin et al. 2026 Terminal State; injected failure Reconstructed terminal workspace Executable tests
TermiGen Zhu et al. 2026 Terminal Seed; task/container specification Synthesized executable workspace Executable verifier
EnvScaler Song et al. 2026 Tool / API State / tool; rules Generated database state State / probing
VeriEnv Chae et al. 2026 Web State; reconstructed site Executable cloned website Deterministic state checks
DIVE Chen et al. 2026a Tool Trajectory Real tool execution Execution-grounded
Agent-World Dong et al. 2026 Tool / API Tool / state; capability-guided Discovered/generated environments Verifiable tasks
SkillSynth Fan et al. 2026a Terminal Tool; skill graph/path Executable terminal workspace Executable task
Terminal-World Cheng et al. 2026 Terminal Tool; skill primitives Skill-conditioned environment/task bundle Executable bundle
GTA Huang et al. 2026b Web Seed / state; site graph Reachable pages and Web paths Grounded execution
Trajectory2Task Wang et al. 2026d Tool Trajectory Executed multi-turn tool trace Executable / verifiable
CUA-Gym Wang et al. 2026a Computer use Seed / state; task templates Task-specific application state State / judge
TASTE Keren et al. 2026 Tool Tool; sequences / dependencies Tool-sequence instantiation into executable benchmark tasks Executability / scoring
Tmax Ivison et al. 2026 Terminal Seed; taxonomy/persona Generated terminal environment Diversified verifiers
SETA Shen et al. 2026 Terminal Seed; diverse sources Synthesized terminal environment Unified executable verifier
GAIS Shi et al. 2026b Tool / MCP Tool; dependency graph Protocol-grounded tool ecosystem Grounded interaction
ScaleCUA Lv et al. 2026 Computer use State / trajectory; exploration Docker-interaction task generation; later frontier sampling Executable judges
NexForge Zhao et al. 2026 Terminal Seed; capability requirements Task-specific executable workspace Executable tests
Meta-Task Pan et al. 2026 Terminal Seed; task profile Generated sandbox/workspace Sandbox execution
ACuRL Xue et al. 2026 Computer use Trajectory; feedback-guided Existing GUI environment Model/rubric based
SKT Tan et al. 2026 Tool / skill Tool; skill composition Existing tool/skill configurations Verified data
State2State Lei et al. 2026 General State; reachable targets Existing environment state Target-state match
RST Li et al. 2026c Terminal Seed; reference solution Recursively expanded executable workspace Executable tests
CalibForge Meng et al. 2026 Terminal Seed; solver-calibrated Existing/synthesized workspace Solver + executable tests
Envs-FORGE Wu et al. 2026c Terminal Seed; reward-guided rewriting Jointly rewritten task and environment artifacts Gold-verified bundle
FACET Shi et al. 2026a Terminal Seed; source skill/intent Repaired executable container Container + tests
Terminal-Universe Wu et al. 2026a Terminal Trajectory / state; recovered workspace Reconstructed reusable workspace Executable workspace
AutoPlay Ramrakhya et al. 2026 GUI Trajectory / state; exploration Discovered application state Execution outcome
SynWeaver Wang et al. 2026b Web Trajectory / state; site prior Existing website structure Path / state based
AgentMercury Jeong and Yoon 2026 Business / tool Seed / state; scenario Persistent multi-service state Cross-service invariants
TerminalTraj Wu et al. 2026b Terminal Seed / tool; repository docs/scripts Repository-derived Docker environments State-based pytest
CLI-Universe Hua et al. 2026 Terminal Seed; capability taxonomy Evidence-refined blueprints realized in Docker environments Rubric-gated tests; reference solution; fail-to-pass checks
Terminal-Task-Gen Pi et al. 2026 Terminal Seed; problems/skill taxonomy Generated inputs and tasks in shared domain-specific Docker images Weighted pytest checks; optional seed-solution expectations

Appendix B.3. Learning-Driven Evolution Literature

Table A3 summarizes methods in which evidence from agent interaction, solver behavior, or learning progress is used to assess and adapt tasks, environments, or their distributions over time. These signals support the capability estimation and agent-relative resource assessment discussed in Section 5, and are then translated into changes to the available interactive learning resources. We focus on agent-conditioned evolution, where resource adaptation is informed by the behavior or performance of a target learner or solver. Related mechanisms that provide programmable adaptation but do not themselves close a learner-driven training loop are marked with †. Iterative environment generation that follows predefined transformation directions without using current-agent evidence is treated as a boundary case closer to environment construction rather than as a core form of learning-driven evolution. Pre-training defect repair, such as Trustworthy Worlds, is therefore recorded in Table A1 rather than counted as learning-driven evolution.
Table A3. Representative methods for learning-driven task and environment evolution. The table separates the agent or solver evidence used for assessment, the resource being adapted, and the mechanism through which adaptation is performed. Works marked with † are precursor or boundary mechanisms that do not themselves instantiate a complete learner-driven evolution loop.
Table A3. Representative methods for learning-driven task and environment evolution. The table separates the agent or solver evidence used for assessment, the resource being adapted, and the mechanism through which adaptation is performed. Works marked with † are precursor or boundary mechanisms that do not themselves instantiate a complete learner-driven evolution loop.
Work Setting Agent / solver evidence Adaptation target Adaptation mechanism
DreamGym Chen et al. 2026b Reasoning / tool Real interaction experience / model calibration Environment challenges Continual update of the experience model and generated challenges
RLVE Zeng et al. 2025 Reasoning Agent success rate Procedural task difficulty Capability-adaptive curriculum over manually engineered environments
GenEnv Guo et al. 2026b General Policy performance / curriculum reward Environment difficulty Difficulty-aligned agent–environment co-evolution
ProEvolve Li et al. 2026a General Benchmark state / transformation specification Environment and benchmark artifacts Programmable graph transformations for benchmark evolution
Agent-World Dong et al. 2026 Tool / API Capability gaps / interaction performance Tasks and environments Capability-guided discovery of new environments and verifiable tasks
TRACE Kang et al. 2026 General Successful vs. failed trajectories Tasks and environments Capability-targeted mini-environment synthesis
ScaleCUA Lv et al. 2026 Computer use Task success / capability frontier Task sampling distribution Frontier-based task sampling during online RL
SETA Shen et al. 2026 Terminal Agent performance / training progress Terminal tasks and environments Environment and task mutation during continual training
ACuRL Xue et al. 2026 Computer use Previous interaction feedback Curriculum tasks Continual regeneration of curriculum tasks
CalibForge Meng et al. 2026 Terminal Solver behavior / pass rate Task difficulty Solver-relative calibration toward the learnable frontier
Envs-FORGE Wu et al. 2026c Terminal Verifier reward / seed pass rate Tasks, fixtures, oracle, tests, and runtime Frontier-optimized joint rewriting of task and environment artifacts
SPADE Liu et al. 2026 General / tool Regret / agent performance Executable environments Self-play generation near the learner’s capability frontier
EnvHarness Huang et al. 2026a Multi-domain Policy weaknesses extracted from trajectories Environment wrappers, rules, and observations Harness synthesis targeted at observed policy weaknesses
EvoEnv Shi et al. 2026c Reasoning Agent-relative difficulty / novelty Executable reasoning environments Generate and filter environment programs around the current learning frontier

References

  1. Ivan Bercovich. 2026. What makes a good terminal-agent benchmark task: A guideline for adversarial, difficult, and legible evaluation design. Preprint, arXiv:2604.28093.
  2. Shihao Cai, Runnan Fang, Jialong Wu, Baixuan Li, Xinyu Wang, Yong Jiang, Liangcai Su, Liwen Zhang, Wenbiao Yin, Zhen Zhang, Fuli Feng, Pengjun Xie, and Xiaobin Wang. 2025. AutoForge: Automated environment synthesis for agentic reinforcement learning. Preprint, arXiv:2512.22857.
  3. Yuan Cao, Dezhi Ran, Mengzhou Wu, Yuzhe Guo, Xin Chen, Ang Li, Gang Cao, Gong Zhi, Hao Yu, Linyi Li, Wei Yang, and Tao Xie. 2026. GUI-GENESIS: Automated synthesis of efficient environments with verifiable rewards for GUI agent post-training. Preprint, arXiv:2602.14093.
  4. Tommaso Castellani, Naimeng Ye, Daksh Mittal, Thomson Yen, and Hongseok Namkoong. 2025. SynthTools: A framework for scaling synthetic tools for agent development. Preprint, arXiv:2511.09572.
  5. Hyungjoo Chae, Jungsoo Park, and Alan Ritter. 2026. Safe and scalable web agent learning via recreated websites. Preprint, arXiv:2603.10505.
  6. Aili Chen, Chi Zhang, Junteng Liu, Jiangjie Chen, Chengyu Du, Yunji Li, Ming Zhong, Qin Wang, Zhengmao Zhu, Jiayuan Song, Ke Ji, Junxian He, Pengyu Zhao, and Yanghua Xiao. 2026a. DIVE: Scaling diversity in agentic task synthesis for generalizable tool use. Preprint, arXiv:2603.11076.
  7. Zhaorun Chen, Zhuokai Zhao, Kai Zhang, Bo Liu, Qi Qi, Yifan Wu, Tarun Kalluri, Sara Cao, Yuanhao Xiong, Haibo Tong, Huaxiu Yao, Hengduo Li, Jiacheng Zhu, Xian Li, Dawn Song, Bo Li, Jason Weston, and Dat Huynh. 2026b. Scaling agent learning via experience synthesis. In International Conference on Learning Representations.
  8. Zihao Cheng, Hongru Wang, Zeming Liu, Xinyi Wang, Xiangrong Zhu, Yuhang Guo, Wei Lin, Jeff Z. Pan, and Yunhong Wang. 2026. Terminal-World: Scaling terminal-agent environments via agent skills. Preprint, arXiv:2605.20876.
  9. Hang Ding, Peidong Liu, Junqiao Wang, Ziwei Ji, Meng Cao, Rongzhao Zhang, Lynn Ai, Eric Yang, Tianyu Shi, and Lei Yu. 2026. DynaWeb: Model-based reinforcement learning of web agents. Preprint, arXiv:2601.22149.
  10. Guanting Dong, Junting Lu, Junjie Huang, Wanjun Zhong, Longxiang Liu, Shijue Huang, Zhenyu Li, Yang Zhao, Xiaoshuai Song, Xiaoxi Li, Jiajie Jin, Yutao Zhu, Hanbin Wang, Fangyu Lei, Qinyu Luo, Mingyang Chen, Zehui Chen, Jiazhan Feng, Ji-Rong Wen, and Zhicheng Dou. 2026. Agent-World: Scaling real-world environment synthesis for evolving general agent intelligence. Preprint, arXiv:2604.18292.
  11. Weihua Du, Hailei Gong, Zhan Ling, Kang Liu, Lingfeng Shen, Xuesong Yao, Yufei Xu, Dingyuan Shi, Yiming Yang, and Jiecao Chen. 2026. Generalizable end-to-end tool-use RL with synthetic CodeGym. In International Conference on Learning Representations.
  12. Zhiyuan Fan, Tinghao Yu, Yuanjun Cai, Jiangtao Guan, Yun Yang, Dingxin Hu, Jiang Zhou, Xing Wu, Zhuo Han, Feng Zhang, and Lilin Wang. 2026a. Toward scalable terminal task synthesis via skill graphs. Preprint, arXiv:2604.25727.
  13. Zhiyuan Fan, Tinghao Yu, Yuanjun Cai, Jiang Zhou, Jiangtao Guan, Jincheng Liu, Yun Yang, Dingxin Hu, Zhuo Han, Xing Wu, Feng Zhang, and Lilin Wang. 2026b. Environment evolution for terminal agents. Preprint, arXiv:2609.04128.
  14. Runnan Fang, Shihao Cai, Baixuan Li, Jialong Wu, Guangyu Li, Wenbiao Yin, Xinyu Wang, Xiaobin Wang, Liangcai Su, Zhen Zhang, Shibin Wu, Zhengwei Tao, Yong Jiang, Pengjun Xie, Ningyu Zhang, Fei Huang, Wentao Zhang, and Jingren Zhou. 2026. Towards general agentic intelligence via environment scaling. In Findings of the Association for Computational Linguistics: ACL 2026, pages 17610–17621. Association for Computational Linguistics.
  15. Kanishk Gandhi, Shivam Garg, Noah D. Goodman, and Dimitris Papailiopoulos. 2026. Endless terminals: Scaling RL environments for terminal agents. Preprint, arXiv:2601.16443.
  16. Huan-ang Gao, Jiayi Geng, Wenyue Hua, Mengkang Hu, Xinzhe Juan, Hongzhang Liu, Shilong Liu, Jiahao Qiu, Xuan Qi, Yiran Wu, Hongru Wang, Han Xiao, Yuhang Zhou, Shaokun Zhang, Jiayi Zhang, Jinyu Xiang, Yixiong Fang, Qiwen Zhao, Dongrui Liu, and 8 others. 2026. A survey of self-evolving agents: What, when, how, and where to evolve on the path to artificial super intelligence. Transactions on Machine Learning Research.
  17. Chuanzhe Guo, Jingjing Wu, Sijun He, Yang Chen, Zhaoqi Kuang, Shilong Fan, Bingjin Chen, Siqi Bao, Jing Liu, Hua Wu, Qingfu Zhu, Wanxiang Che, and Haifeng Wang. 2026a. MEnvAgent: Scalable polyglot environment construction for verifiable software engineering. Preprint, arXiv:2601.22859.
  18. Jiacheng Guo, Ling Yang, Peter Chen, Qixin Xiao, Yinjie Wang, Xinzhe Juan, Jiahao Qiu, Ke Shen, and Mengdi Wang. 2026b. GenEnv: Difficulty-aligned co-evolution between LLM agents and environment simulators. In International Conference on Learning Representations.
  19. Zhicheng Guo, Sijie Cheng, Yuchen Niu, Hao Wang, Sicheng Zhou, Wenbing Huang, and Yang Liu. 2025. StableToolBench-MirrorAPI: Modeling tool environments as mirrors of 7,000+ real-world APIs. In Findings of the Association for Computational Linguistics: ACL 2025, pages 5247–5270. Association for Computational Linguistics.
  20. Mengkang Hu, Pu Zhao, Can Xu, Qingfeng Sun, Jian-Guang Lou, Qingwei Lin, Ping Luo, and Saravan Rajmohan. 2025. AgentGen: Enhancing planning abilities for large language model based agent via environment and task generation. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pages 496–507.
  21. Zhanbo Hua, Yifan Yao, Weihao Xie, Yongchi Zhao, Minghao Liu, Ruizhi Qiu, Zhewei Huang, Zun Wang, Yiyan Ji, Yunhai Ye, Letian Zhu, Xinping Lei, Han Li, Zhiyuan Ma, Zili Wang, Zhaoxiang Zhang, and Jiaheng Liu. 2026. CLI-Universe: Towards verifiable task synthesis engine for terminal agents. Preprint, arXiv:2606.22883.
  22. Chengsong Huang, Zifeng Wang, Rujun Han, Jun Yan, Yanfei Chen, Zoey CuiZhu, Ke Jiang, Peng Xia, Han Yu, Yufan Zhuang, Yifei Ming, Jiaqi Pan, Bhavana Dalvi Mishra, Jiaxin Huang, Burak Gokturk, Tomas Pfister, and Chen-Yu Lee. 2026a. EnvHarness: Awakening static worlds for agent learning. Preprint, arXiv:2608.19880.
  23. Tenghao Huang, Kung-Hsiang Huang, Prafulla Kumar Choubey, Yilun Zhou, Muhao Chen, Jonathan May, and Chien-Sheng Wu. 2026b. GTA: Generating long-horizon tasks for web agents at scale. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 18805–18820. Association for Computational Linguistics.
  24. Maksim Ivanov and Abhijay Rana. 2026. Anchor: Mitigating artifact drift in agent benchmark generation. Preprint, arXiv:2605.26321. Presented at the RLEval Workshop, ACM CAIS 2026 (non-archival).
  25. Hamish Ivison, Junjie Oscar Yin, Rulin Shao, Teng Xiao, Nathan Lambert, and Hannaneh Hajishirzi. 2026. Tmax: A simple recipe for terminal agents. Preprint, arXiv:2606.23321.
  26. Minbyul Jeong and Chanwoong Yoon. 2026. AgentMercury: Your agent can synthesize verifiable environments for business scenarios at scale. Preprint, arXiv:2608.20634.
  27. Hangoo Kang, Tarun Suresh, Jon Saad-Falcon, and Azalia Mirhoseini. 2026. TRACE: Capability-targeted agentic training. Preprint, arXiv:2604.05336.
  28. Tomer Keren, Nitay Calderon, Asaf Yehudai, Yotam Perlitz, Michal Shmueli-Scheuer, and Roi Reichart. 2026. A matter of TASTE: Improving coverage and difficulty of agent benchmarks. Preprint, arXiv:2605.28556.
  29. Seanie Lee, Sanjoy Chowdhury, Chao Jiang, Cheng-Yu Hsieh, Ting-Yao Hu, Alexander T. Toshev, Oncel Tuzel, and Raviteja Vemulapalli. 2026. Environment-free synthetic data generation for API-calling agents. Preprint, arXiv:2607.16900.
  30. Xuanyu Lei, Yiqi Zhu, Chenliang Li, Kaiming Liu, Peng Li, Ming Yan, Jieping Ye, Ya-Qin Zhang, and Yang Liu. 2026. State2State: Environment-derived mid-training for LLM agents. Preprint, arXiv:2608.04934.
  31. Guangrui Li, Yaochen Xie, Yi Liu, Ziwei Dong, Xingyuan Pan, Tianqi Zheng, Jason Choi, Michael J. Morais, Binit Jha, Shaunak Mishra, Bingrou Zhou, Chen Luo, Monica Xiao Cheng, and Dawn Song. 2026a. The world won’t stay still: Programmable evolution for agent benchmarks. Preprint, arXiv:2603.05910.
  32. Jiachun Li, Zhuoran Jin, Tianyi Men, Yupu Hao, Kejian Zhu, Lingshuai Wang, Dongqi Huang, Longxiang Wang, Shengjia Hua, Lu Wang, Jinshan Gao, Hongbang Yuan, Ruilin Xu, Kang Liu, and Jun Zhao. 2026b. Agentic environment engineering for large language models: A survey of environment modeling, synthesis, evaluation, and application. Preprint, arXiv:2606.12191.
  33. Yuetai Li, Huseyin A. Inan, Xiang Yue, Wei-Ning Chen, Lukas Wutschitz, Janardhan Kulkarni, Radha Poovendran, Robert Sim, and Saravan Rajmohan. 2025. Simulating environments with reasoning models for agent training. Preprint, arXiv:2511.01824.
  34. Zhongzhi Li, Yucheng Shi, Zongxia Li, Ruhan Wang, Anhao Li, Zixun Huang, Junyao Yang, Lei Ke, Ninghao Liu, Haitao Mi, and Leowei Liang. 2026c. Recursive synthesis for long-horizon terminal tasks. Preprint, arXiv:2608.05466.
  35. Yusong Lin, Haiyang Wang, Shuzhe Wu, Lue Fan, Feiyang Pan, Sanyuan Zhao, and Dandan Tu. 2026. CLI-Gym: Scalable CLI task generation via agentic environment inversion. Preprint, arXiv:2602.10999.
  36. Bo Liu, Simon Yu, Yiding Jiang, Ao Qu, Andrew Zhao, Zichen Liu, Junsu Kim, Zijian Zhou, Seungone Kim, Tongzheng Ren, Mickel Liu, Hanfei Yu, Zhaorun Chen, Weiyan Shi, Paul Pu Liang, Luke Zettlemoyer, Yejin Choi, and Natasha Jaques. 2026. SPADE: Self-play in adaptive synthetic executable environments. Preprint, arXiv:2608.19197.
  37. Bowen Lv, Xiao Liu, Yanyu Ren, Hanyu Lai, Bohao Jing, Hanchen Zhang, Yanxiao Zhao, Shuntian Yao, Jie Tang, and Yuxiao Dong. 2026. SCALECUA: Scaling computer use agents with verifiable task synthesis and efficient online RL. Preprint, arXiv:2607.11185.
  38. Fanzhe Meng, Guoxin Chen, Jiale Zhao, Shuang Sun, Zhiyu Lin, Wayne Xin Zhao, Ruihua Song, Ji-Rong Wen, and Kai Jia. 2026. CalibForge: Adversarial solver calibration for scaling learnable terminal tasks. Preprint, arXiv:2608.06352.
  39. Shikhar Murty, Hao Zhu, Dzmitry Bahdanau, and Christopher D. Manning. 2024. NNetNav: Unsupervised learning of browser agents through environment interaction in the wild. Preprint, arXiv:2410.02907.
  40. Zhihong Pan, Jiyuan He, Kai Zhang, Yupeng Han, Ze Liu, Yuze Zhao, Yongcong Ye, and Zhaohua Yang. 2026. Meta-Task: Turning terminal task synthesis into a terminal task for scalable agent training. Preprint, arXiv:2607.27929.
  41. Xiaoxuan Peng, Kaiqi Zhang, Xinyu Lu, Boxi Cao, Yaojie Lu, Hongyu Lin, Xianpei Han, and Le Sun. 2026. LiteCoder-Terminal: Scaling long-horizon terminal environments for learning language agents. Preprint, arXiv:2605.29559.
  42. Renjie Pi, Grace Lam, Mohammad Shoeybi, Pooya Jannaty, Bryan Catanzaro, and Wei Ping. 2026. On data engineering for scaling LLM terminal capabilities. Preprint, arXiv:2602.21193.
  43. Ram Ramrakhya, Andrew Szot, Omar Attia, Yuhao Yang, Anh Nguyen, Bogdan Mazoure, Zhe Gan, Harsh Agrawal, and Alexander Toshev. 2026. Scaling synthetic task generation for agents via exploration. In International Conference on Learning Representations.
  44. Qijia Shen, Zhiqi Huang, Vamsidhar Kamanuru, Aznaur Aliev, Jay Rainton, Ahmed Awelkair, Zhichen Zeng, Jiajun Li, Shi Dong, Yueming Yuan, Boyuan Ma, Qizheng Zhang, Jiwei Fu, Yuzhen Mao, Wendong Fan, Ping Nie, Philip Torr, Bernard Ghanem, Changran Hu, and 3 others. 2026. SETA: Scaling environments for terminal agents. Preprint, arXiv:2607.10891.
  45. Dingfeng Shi, Jingyi Cao, Qianben Chen, Weichen Sun, Weizhen Li, Hongxuan Lu, Fangchen Dong, Tianrui Qin, King Zhu, Minghao Yang, Jian Yang, Ge Zhang, Jiaheng Liu, Changwang Zhang, Jun Wang, Yuchen Eleanor Jiang, and Wangchunshu Zhou. 2025. TaskCraft: Automated generation of agentic tasks. Preprint, arXiv:2506.10055.
  46. Kou Shi, Zun Wang, Qisheng Su, Shiting Huang, Ziao Zhang, Zhen Fang, Qingnan Ren, Jin Liu, Yu Zeng, Yiming Zhao, Lin Chen, Zehui Chen, and Feng Zhao. 2026a. FACET: Preserving source intent and executable state in terminal task synthesis. Preprint, arXiv:2608.18580.
  47. Wenhang Shi, Jinhao Dong, Yiren Chen, Zhe Zhao, Shuqing Bian, Wei Lu, and Xiaoyong Du. 2026b. Scaling agentic capabilities via grounded interaction synthesis. In Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pages 4252–4263.
  48. Yucheng Shi, Zhenwen Liang, Kishan Panaganti, Dian Yu, Wenhao Yu, and Haitao Mi. 2026c. Learning to build the environment: Self-evolving reasoning RL via verifiable environment synthesis. Preprint, arXiv:2605.14392.
  49. Xiaoshuai Song, Haofei Chang, Guanting Dong, Yutao Zhu, Ji-Rong Wen, and Zhicheng Dou. 2026. EnvScaler: Scaling tool-interactive environments for LLM agent via programmatic synthesis. In Findings of the Association for Computational Linguistics: ACL 2026, pages 8326–8357. Association for Computational Linguistics.
  50. Michael Sullivan, Mareike Hartmann, and Alexander Koller. 2025. Procedural environment generation for tool-use agents. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 18544–18562. Association for Computational Linguistics.
  51. Qiushi Sun, Kanzhi Cheng, Zichen Ding, Chuanyang Jin, Yian Wang, Fangzhi Xu, Zhenyu Wu, Chengyou Jia, Liheng Chen, Zhoumianze Liu, Ben Kao, Guohao Li, Junxian He, Yu Qiao, and Zhiyong Wu. 2025. OS-genesis: Automating GUI agent trajectory construction via reverse task synthesis. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5555–5579. Association for Computational Linguistics.
  52. Zelin Tan, Yiqun Zhang, Hao Li, Zhiyao Cui, Hejia Geng, Shao Zhang, Hangfan Zhang, Yang Chen, Xiaosong Wang, Lilong Wang, Zhenfei Yin, Shuyue Hu, Chen Zhang, and Lei Bai. 2026. SKT: Skill-use training at scale via verified synthetic data generation. Preprint, arXiv:2608.02287.
  53. Dunwei Tu, Hongyan Hao, Hansi Yang, Yihao Chen, Yi-Kai Zhang, Zhikang Xia, Yu Yang, Yueqing Sun, Xingchen Liu, Furao Shen, Qi Gu, Hui Su, and Xunliang Cai. 2026. ScaleEnv: Scaling environment synthesis from scratch for generalist interactive tool-use agent training. Preprint, arXiv:2602.06820.
  54. Vivek Verma, David Huang, William Chen, Dan Klein, and Nicholas Tomlin. 2025. Measuring general intelligence with generated games. Preprint, arXiv:2505.07215.
  55. Bowen Wang, Dunjie Lu, Junli Wang, Tianyi Bai, Shixuan Liu, Zhipeng Zhang, Haiquan Wang, Hao Hu, Tianbao Xie, Shuai Bai, Dayiheng Liu, Que Shen, Junyang Lin, and Tao Yu. 2026a. CUA-Gym: Scaling verifiable training environments and tasks for computer-use agents. Preprint, arXiv:2605.25624.
  56. Ruitao Wang, Yuwen Hao, and Menglin Yang. 2026b. SynWeaver: Website-prior task and trajectory co-synthesis for web agents. Preprint, arXiv:2608.12429.
  57. Zhaoyang Wang, Canwen Xu, Boyi Liu, Yite Wang, Siwei Han, Zhewei Yao, Huaxiu Yao, and Yuxiong He. 2026c. Agent world model: Infinity synthetic environments for agentic reinforcement learning. Preprint, arXiv:2602.10090.
  58. Ziyi Wang, Yuxuan Lu, Yimeng Zhang, Pei Chen, Ziwei Dong, Jing Huang, Jiri Gesi, Xianfeng Tang, Chen Luo, Qun Liu, Yisi Sang, Hanqing Lu, Manling Li, Jin Lai, and Dakuo Wang. 2026d. Trajectory2Task: Training robust tool-calling agents with synthesized yet verifiable data for complex user intents. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 44021–44044. Association for Computational Linguistics.
  59. Jie Wu, Zhenru Zhang, Beichen Zhang, Xuwu Wang, Yuhui Su, Mouxiang Chen, Peng Wang, Zhihai Wang, Que Shen, Hao Zhou, An Yang, Fei Huang, Yujiu Yang, and Dayiheng Liu. 2026a. Terminal-Universe: Turning agent trajectories into scalable terminal environments. Preprint, arXiv:2609.04148.
  60. Siwei Wu, Yizhi Li, Yuyang Song, Wei Zhang, Yang Wang, Riza Batista-Navarro, Xian Yang, Mingjie Tang, Bryan Dai, Jian Yang, and Chenghua Lin. 2026b. Large-scale terminal agentic trajectory generation from dockerized environments. In Forty-third International Conference on Machine Learning.
  61. Xiaojun Wu, Cehao Yang, Honghao Liu, Xueyuan Lin, Zhichao Shi, Hao Zhou, Xuhui Jiang, Chengjin Xu, Jia Li, and Jian Guo. 2026c. Envs-FORGE: Frontier-optimized reward-grounded environment synthesis for agent RL. Preprint, arXiv:2608.14312.
  62. Yifan Wu, Yiran Peng, Yiyu Chen, Jianhao Ruan, Zijie Zhuang, Cheng Yang, Jiayi Zhang, Man Chen, Yenchi Tseng, Zhaoyang Yu, Liang Chen, Yuyao Zhai, Bang Liu, Chenglin Wu, and Yuyu Luo. 2026d. AutoWebWorld: Synthesizing infinite verifiable web environments via finite state machines. Preprint, arXiv:2602.14296.
  63. Ziqiao Xi, Shuang Liang, Qi Liu, Jiaqing Zhang, Letian Peng, Fang Nan, Meshal Nayim, Tianhui Zhang, Rishika Mundada, Lianhui Qin, Biwei Huang, and Kun Zhou. 2026. C-world: A computer use agent environment creator. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 43202–43216. Association for Computational Linguistics.
  64. Jingxu Xie, Dylan Xu, Xuandong Zhao, and Dawn Song. 2026. AgentSynth: Scalable task generation for generalist computer-use agents. In International Conference on Learning Representations.
  65. Minrui Xu, Zilin Wang, Mengyi Deng, Zhiwei Li, Zhicheng Yang, Xiao Zhu, Yinhong Liu, Boyu Zhu, Baiyu Huang, Chao Chen, Heyuan Deng, Fei Mi, Lifeng Shang, Xingshan Zeng, and Zhijiang Guo. 2026a. EnvFactory: Scaling tool-use agents via executable environments synthesis and robust RL. Preprint, arXiv:2605.18703.
  66. Weikai Xu, Yunren Feng, Haoxiang Lei, Kun Huang, Yuxuan Liu, Kang Zhao, Xiaolin Hu, Shuo Shang, and Bo An. 2026b. AppDeltaWorld: Transition-grounded delta code world model for mobile GUI agents. Preprint, arXiv:2608.05891.
  67. Tianci Xue, Zeyi Liao, Tianneng Shi, Zilu Wang, Kai Zhang, Dawn Song, Yu Su, and Huan Sun. 2026. Autonomous continual learning for environment adaptation of computer-use agents. Preprint, arXiv:2602.10356.
  68. Xingshan Zeng, Zishan Xu, Boju Zhang, Yuzhou Wu, Lingzhi Wang, Jianghao Lin, Liangyou Li, Yasheng Wang, Lifeng Shang, Xin Jiang, Weinan Zhang, Yong Yu, Qun Liu, and Weiwen Liu. 2026. What makes good agentic data? an ACE lens on data generation for LLM agents. Preprint, arXiv:2608.27260.
  69. Zhiyuan Zeng, Hamish Ivison, Yiping Wang, Lifan Yuan, Shuyue Stella Li, Zhuorui Ye, Siting Li, Jacqueline He, Runlong Zhou, Tong Chen, Chenyang Zhao, Yulia Tsvetkov, Simon Shaolei Du, Natasha Jaques, Hao Peng, Pang Wei Koh, and Hannaneh Hajishirzi. 2025. RLVE: Scaling up reinforcement learning for language models with adaptive verifiable environments. Preprint, arXiv:2511.07317.
  70. Chenghao Zhang, Canran Xiao, SaiSai Hu, and Dan Roth. 2026a. Training needs trustworthy worlds: Verified synthetic web environments for agent learning. Preprint, arXiv:2608.21898.
  71. Guibin Zhang, Hejia Geng, Xiaohang Yu, Zhenfei Yin, Zaibin Zhang, Zelin Tan, Heng Zhou, Zhongzhi Li, Xiangyuan Xue, Yijiang Li, Yifan Zhou, Yang Chen, Chen Zhang, Yutao Fan, Zihu Wang, Songtao Huang, Francisco Piedrahita-Velez, Yue Liao, Hongru Wang, and 6 others. 2026b. The landscape of agentic reinforcement learning for LLMs: A survey. Transactions on Machine Learning Research.
  72. Ziyun Zhang, Zezhou Wang, Xiaoyi Zhang, Zongyu Guo, Jiahao Li, Bin Li, and Yan Lu. 2026c. InfiniteWeb: Scalable web environment synthesis for GUI agent training. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 28465–28492. Association for Computational Linguistics.
  73. Jiarong Zhao, Zhikai Lei, Zhiheng Xi, Rui Zheng, Hang Yan, Jie Zhou, Qin Chen, and Liang He. 2026. NexForge: Scaling executable agent tasks via requirement-first synthesis. Preprint, arXiv:2607.14186.
  74. Kaijie Zhu, Yuzhou Nie, Yijiang Li, Yiming Huang, Jialian Wu, Jiang Liu, Ximeng Sun, Zhenfei Yin, Lun Wang, Zicheng Liu, Emad Barsoum, William Yang Wang, and Wenbo Guo. 2026. TermiGen: High-fidelity environment and robust trajectory synthesis for terminal agents. Preprint, arXiv:2602.07274.
  75. Yiqi Zhu, Apurva Gandhi, and Graham Neubig. 2025. Training versatile coding agents in synthetic environments. Preprint, arXiv:2512.12216.
Figure 1. Environments and tasks form interactive learning resources. Synthesizing them at scale supports agent interaction that yields training rewards and reusable experience for self-improvement.
Figure 1. Environments and tasks form interactive learning resources. Synthesizing them at scale supports agent interaction that yields training rewards and reusable experience for self-improvement.
Preprints 233934 g001
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.