Preprint
Review

This version is not peer-reviewed.

Agentic Visual Generation: A Survey

Submitted:

30 August 2026

Posted:

02 September 2026

You are already at the latest version

Abstract
Visual generation is becoming a general component of artificial intelligence infrastructure, supporting the production and editing of images, videos, three-dimensional assets, data-grounded graphics, structured documents, interactive interfaces, and other visual artifacts. Generative and multimodal models have substantially improved single-pass rendering, yet as applications extend to multi-step visual creation, persistent weaknesses in spatial precision, identity persistence, temporal coherence, editable structure, factual fidelity, and physical plausibility, among other cross-step requirements, accumulate across steps. Meeting these requirements is therefore a problem of creation control: model scaling, prompt refinement, additional conditioning, repeated sampling, and other output-level refinements can improve individual outputs, but the process remains open-loop whenever runtime observations do not change subsequent actions, allowing early errors to persist or propagate across later artifacts and actions. This control problem has motivated agentic visual generation (AVG): goal-driven visual creation organized as a closed loop over a multi-step trajectory, in which, at the decision points exposed to the controller (which operation or tool to invoke, which repair target to address, whether to continue or stop, and related decisions), the selection among available alternatives depends on runtime observations of the evolving visual artifact, its task environment, and the interaction history. Its defining behavioral criterion is a demonstrable observation--action dependency: different runtime observations lead to different selected actions. This agentic direction has become an increasingly prominent trend in visual creation, yet few existing surveys systematically organize this rapidly growing area. To fill this gap, this survey presents a unified account of agentic visual generation. Specifically, it first defines and formalizes AVG around the observation--action dependency, distinguishes it from one-shot generators, fixed workflows, and agent-augmented workflows, and grades autonomy on five cumulative levels (L1--L5). A six-component analytical framework (goal understanding, specification, and planning; memory; tool; perception; action; and cross-task self-improvement) then organizes a cross-domain comparison spanning image, video, 3D/CAD, scientific-visualization, document, UI/Web, and other visual-creation domains. Building on this comparison, the survey analyzes training from trajectory supervision and policy initialization through reinforcement learning from multimodal and executable feedback to experience reuse with skill abstraction and transfer, and organizes evaluation into four complementary levels (artifact quality, goal and constraint satisfaction, trajectory and decision quality, and system and human-centered evaluation). Finally, research directions follow three paths: expanding the scope of visual agency, increasing the intelligence of visual agents, and establishing reliable evaluation.
Keywords: 
;  ;  ;  ;  ;  ;  ;  
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.