Preprint
Review

This version is not peer-reviewed.

From Trajectories to Experience: A Survey of Experience-Driven LLM Agents

Submitted:

09 October 2026

Posted:

10 October 2026

You are already at the latest version

Abstract
Large language model (LLM) agents record in their trajectories which strategies failed and how they recovered, yet an agent that resolves a problem in one episode may repeat the same diagnostic detour in the next. Work on reflection, agent memory, workflow learning, and skill acquisition addresses this gap under different names, making methods and their gains hard to compare. We define agent experience as action-guiding information grounded in task-execution trajectories and retained to guide later episodes where it applies; it may be a trajectory kept as a precedent, a semantic lesson, or a reusable procedure. We review methods through four lifecycle functions: formation selects or derives candidate experience from trajectories, organization makes it accessible, utilization applies it to current decisions or internalizes it into model parameters, and maintenance admits, consolidates, revises, and retires it as evidence accumulates. A recurring source of failure is guidance that loses its link to its applicability conditions and supporting evidence: it is over-generalized during formation, misapplied during utilization, and left unrevised during maintenance. We organize evaluation around intrinsic experience quality, the experience utilization process, and downstream effects, and separate three improvement claims: fixed experience reuse, feedback-driven sustained improvement, and improvement of the learning mechanism. The reviewed studies most clearly support fixed experience reuse on tasks with tests or reference answers; evidence for the other two claims remains partial. Across seven application settings, we derive recommendations on whether reuse justifies its cost, what to retain, how strongly experience should influence a task, and when to revise it. We close with eight open problems, from learning with human expert records to correcting propagated errors and testing whether gains accumulate.
Keywords: 
;  
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.