Preprint
Article

This version is not peer-reviewed.

ReAgent: Rethinking Agent Training Through Scenario Co-Design

Submitted:

21 August 2026

Posted:

21 August 2026

You are already at the latest version

Abstract
Agentic reinforcement learning scales training by expanding programmatic environments and synthesizing diverse tasks. We show that this scaling can paradoxically narrow training coverage through how data are constructed. Existing pipelines often construct tasks from executable paths or successful workflows, inducing a construction-induced selection bias toward feasible requests. In real-world applications, however, users often state goals without observing the environment state or environment-specific constraints, leaving such interactions underrepresented in training. This bias shows that scaling training around executable tasks alone cannot adequately represent the user--environment relations encountered in real-world applications. For such interactions, the existence of a conflict, its valid resolutions, and the corresponding supervision must be jointly determined with respect to the same environment instance. We therefore rethink agent training by taking the training scenario, rather than the executable task alone, as the unit of construction and verification. We introduce ReAgent to instantiate this view through scenario co-design, using environment-supported conflicts and valid resolutions as the common basis for jointly defining the task, user behavior, and supervision. ReAgent constructs feasible and conflict scenarios over execution-tested programmatic environments. Across two model scales and three agent benchmarks, ReAgent yields consistent gains spanning executable, blocked, and partially resolvable requests.
Keywords: 
;  
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.