Preprint
Article

This version is not peer-reviewed.

A Survey on AI for AI: When the Improver Becomes the Improvee

Deheng Ye  †
,
Zheng Zhang  †
,
Hao Wang  *
,
Chunyan Miao  *

  † Equal contribution.

Submitted:

13 September 2026

Posted:

17 September 2026

You are already at the latest version

Abstract
AI for AI concerns processes that use AI to improve AI systems and their development methods. This Review examines how AI contributes to these improvement processes, including model learning and harness engineering. We focus on cases where the improver itself changes. We compare updates to experience generators, evaluation mechanisms and programs that produce improvements. A system may combine these paths. Their effects depend on the feedback used, the experience kept for later rounds and the external checks that remain fixed. We examine their benefits, failure modes and costs. We distinguish changes to an improver from gains in its ability to improve AI in later rounds. We then ask when improvement becomes recursive, how gains scale with resources, how long progress can continue and whether improvers generalize across AI systems.
Keywords: 
;  ;  ;  ;  ;  

1. Introduction

Building an AI system involves choosing examples, judging behavior and testing changes. Foundation models increasingly help with this work. Self-Instruct [1] uses a model to generate instruction data. InstructGPT [2] trains a responding model, or policy, using scores from a learned reward model. Automated Design of Agentic Systems (ADAS) [3] uses a language-model agent to propose and test other agents. These systems place AI on both sides of an improvement process. One component supplies data, feedback or proposed changes to help improve another.
The same component can act as both improver and improvee. In STaR [4], a model learns from selected reasoning traces and then produces new traces. These traces record the steps used to reach an answer. A self-rewarding model [5] can supply both answers and the judgments used to train its next version. Promptbreeder [6] changes prompts that guide later revisions of other prompts, and the Darwin Gödel Machine [7] changes agent code that can produce further modifications. The resulting model or program may be better at a task. A further question is whether it has become better at improving AI. That question concerns the process that produces a result, as well as the result itself. The process can change through model learning, harness engineering, or both.
There are practical reasons to study this distinction. Better experience generators may reduce the need for new labels. Better evaluators may guide learning on tasks with costly human judgments. Better design agents may make AI development accessible with less manual trial and error. Yet the process can also repeat its own mistakes, select easier tasks or spend more resources to find a good candidate. A higher final score alone cannot show which of these changes helped. We must identify what changed, how it was used later and what the comparison actually tests.
Existing surveys provide important context. Reviews of language-model self-evolution [8] describe how experience is collected, refined and used for updates. Agent surveys [9,10] organize changes to models, memory, tools and workflows. Recent reviews [11,12,13] also examine recursive improvement, changing evaluators and the reliability of agents over long tasks. We build on these views through a focused comparison of three functional roles. Our aim is to explain how changing an experience generator, an evaluation mechanism or an improvement program affects the next round of work.
We first describe the improvement process and its three AI roles, then examine updates to the improver. We connect gains to limits in task coverage, evaluation, saved experience and cost, and compare results from representative systems. We close with questions about recursion, how gains scale with resources, how long progress can continue and transfer across improvees. Our central question is when AI can make its own improvement process more effective over time.

2. Anatomy of AI for AI

2.1. Scope and Objects

In this Review, AI for AI (AI4AI) refers to processes that use AI to improve AI systems or the methods used to develop and update them [13]. We examine how foundation models and agents built on them carry out these processes. They generate learning experience, provide evaluation feedback, or propose and carry out system changes. “Improvement” describes the intended goal. Tests are needed to show whether a change achieves that goal.
AI systems can take two roles within this process. One system may take both. The first is the improver. It may be a large language model, a vision-language model, or an agent built around such a model. The second is the improvee, the target AI system that receives the proposed improvement. It may be another foundation model, a smaller model or an agent system. For example, Genesys [14] uses language-model agents to design language-model architectures. CLOVA [15] uses a language model to coordinate visual tools and their updates. The improvee need not have the same form as the improver. Multi-agent frameworks such as AutoGen [16] and MetaGPT [17] show why combining agents alone does not mean that the system learns.
Data, prompts, memory, tools and code carry changes into a system. We do not treat each of these objects as a separate AI. When a procedure changes, we identify which AI system uses it to make later improvements. A generated dataset is relevant because it is used to train a model. A revised prompt is relevant because it can change later model or agent behavior. Agent interfaces can also change task performance without changing model weights, as SWE-agent [18] demonstrates.
Harness engineering concerns the software around a model, including prompts, memory, tools and rules for running and checking actions. It falls within AI4AI when AI helps propose, evaluate or apply changes to that software. Self-evolution concerns a system using experience to change itself, so the two can overlap. In Self-Harness [19], a fixed model proposes changes to the harness through which it operates. Accepted changes are used in later rounds. Human choices about goals and approval can still guide an AI4AI process.
Figure 1 groups methods by the work their improvers do. We then examine how those improvers can change.
Figure 2 connects these roles to the structure of this Review. We first describe how AI supplies improvement. We then examine changes to the improver, their possible benefits and failures, and how to assess their value.

2.2. Roles, Feedback and Updates

Five questions help us compare systems. Who guides the change? What changes? What feedback guides it? What is kept for later use? What stays fixed? In Self-Instruct [1], for example, a fixed generator supplies selected examples that are kept for training another model. These questions separate a system’s goal from the signal available during an update. A model score, a test result and a human judgment may all guide selection, but they support different claims about quality.
Saving a change makes later use possible, but does not show that it is used again. STaR [4] keeps training examples and model checkpoints, which save model weights, for later rounds. Reflexion [60] stores written feedback in memory. REvolve [45] keeps reward programs that guide policy training. A component can therefore change even when all language-model weights stay fixed. Changing model weights, in turn, does not show that the method used to update them has improved.
The fixed parts should be named. They may include a proof checker, the task objective, a selection rule, the starting data or a higher-level optimizer. Keeping these parts fixed helps compare results across rounds and can limit unwanted changes. It can also limit which improvements a system can find. These fixed parts should be reported when describing how much work the system can do on its own.

2.3. Boundaries and Timescales

We must state which components belong to the system and when we check whether a change is used again. Self-Refine [65] revises an answer through feedback from a fixed model. Self-debugging [66] and code-repair methods [67] similarly use execution or explanation to improve a current output. These related methods are useful. If the changes are discarded when the task ends, they do not establish a reusable improvement in the model or agent. If the system saves a correction in memory and uses it on later tasks, it may have changed beyond the current answer [68,69,70].
We must also state the time period used for analysis. One trial, one training step, one generated candidate and one independent run are different units. An optimizer may change during a training stage and remain fixed during deployment. metaTextGrad [62] illustrates this distinction. It learns one optimizer through a second optimization process that stays fixed. For self-modifying programs, we need to track which version produces each later change [63]. These boundaries prevent a growing conversation or a long search from being mistaken for repeated improvement of the improver.
We use recursive self-improvement when a system changes how it produces improvements and uses the changed mechanism to help update itself again. An optimizer trained by a separate, fixed process has a different update structure. A recursive structure does not establish better improvement ability or gains that grow faster over time.

3. Foundation Models and Agents as Improvers

The same three functions organize both fixed and changing improvers. We first describe what each function contributes. An AI4AI method can be useful while the component that supplies improvement remains fixed during the improvee’s update. Table 1 groups representative methods by the work they supply.

3.1. Training Experience

A foundation model can turn a small set of examples into a larger training resource. Self-Instruct [1] generates instructions, inputs and responses, filters them and uses the resulting collection for instruction tuning. Later work [22,71] extends model-generated instruction data to dialogue and information retrieval tasks. Models also supply formal statements, code and interface examples whose quality can be checked by formal rules or tests [34,72,73]. Here the model contributes examples that would otherwise require human effort. The learner may be the generator’s next version or a different model [74]. Toolformer [75] generates candidate tool calls for training, while Distilling Step-by-Step [76] uses model-produced explanations to train smaller models.
Training experience can include correct answers, explanations, failed attempts, corrections and interaction histories [4,77,78]. Selecting examples can make the training data better fit the learner’s needs. AP2O-Coder [28] uses preferences related to errors and regular checks to adjust replay, or the reuse of earlier training examples. Active instruction tuning [26] selects useful tasks based on how the model’s responses change, while curricula that account for the learner’s ability [27] change the training tasks as its abilities change. A curriculum is a plan for which tasks a learner sees during training. Selecting useful experience and generating new experience are distinct operations. Both influence what is learned, and both need evaluation beyond the size of the resulting dataset.
For web agents and agents that act in physical or simulated environments, exploration supplies action sequences and feedback [30,31,79,80]. Experience can later train another model, support a memory lookup or become a reusable skill. These uses differ in cost, in what is kept and in how it is used later. The main test is whether this experience helps the improvee learn or act better than experience from a suitable alternative source.

3.2. Evaluation Feedback

Evaluation provides feedback that can guide selection and learning. InstructGPT [2] uses a reward model trained on human comparisons to guide policy training. The reward model represents human feedback, but during this stage it acts as the AI improver. Direct preference optimization [81] trains on preferred and less preferred responses without fitting a separate reward model for that step. If a model supplies those preferences, it still contributes evaluation to the broader process. The training algorithm alone does not identify the source of judgment. RLAIF [39] studies using AI preferences in place of human feedback.
Feedback can take several forms. Scores rank final responses. Process verifiers check intermediate steps, while critic models identify problems and suggest repairs [42,43,82]. Reward code converts an environment state into a training signal, while written scoring criteria guide evaluation of open-ended work [45,47]. Visual and web tasks require checks that connect language to the relevant images or interactions [83,84]. The form of feedback affects what errors a learner can detect. A final score may identify failure without explaining which action caused it.
Agreement with another judge is not enough to make an evaluator useful. Its feedback must help the improvee choose or learn better behavior. Work on critics trained through correction outcomes [41] directly studies this link. Evaluation without reference answers and evaluation based on the effects of an output [85,86] offer further ways to test quality when complete answers are difficult to obtain. Each still needs to explain what makes an outcome useful or a reference reliable.

3.3. AI System Design

A model can propose changes to prompts [87], tools, architecture or the way several agents work together. TextGrad [52] uses language feedback to improve components of a larger system. ProTeGi [88] and GEPA [89] also use written feedback to improve prompts, while AFlow [53] searches over agent workflows. ADAS [3] searches over agent programs, while Dolphin [90] and TusoAI [91] automate parts of AI research and development. The proposed design can be run, tested, saved and revised. AI-assisted changes to prompts, tools and agent workflows are forms of harness engineering.
Architecture design gives a clear example in which the two AI roles differ. Genesys [14] and NADER [56] use language-model agents to propose neural-network architectures, and RZ-NAS [57] uses a model to review designs and assess architectures at low cost. Models can also write procedures for selecting tensor-network structures or acquisition functions, which choose what to test next in Bayesian optimization [92,93]. We include these works because a foundation model helps construct an AI system or an AI-development procedure. We do not separately review traditional architecture or hyperparameter search.
Search history can improve proposals while the designer’s weights and code stay fixed. Compare proposals made with and without that history under the same search budget. A better final design does not by itself show a better design process.

4. When the Improver Becomes the Improvee

The components that perform the three roles in Section 3 can themselves be updated. A generator can learn from its own data. An evaluator can change its judgments. An optimizer can revise the instructions or code it uses to make improvements. The key change occurs when the updated component is used again to improve AI. These paths describe what is updated and how it is used later. They are not levels of intelligence.

4.1. Experience Generators

A model can use successful attempts to build a training set and then generate the next set with its updated weights. STaR [4] provides a clear example. It produces reasoning traces, keeps traces that lead to known correct answers and trains on them. When direct generation fails, the correct answer helps the model produce an explanation for training. The new checkpoint then generates the next round of traces. Training may restart from the original weights. Selected examples still carry information from one round to the next. SPIN [94] uses responses from earlier model versions in contrast with a fixed human demonstration set. Quiet-STaR [95] trains the model to generate reasoning that helps predict later text. These feedback sources impose different limits on what self-generated experience can teach.
The benefit depends on what later generators add. More correct traces may reduce wasted training, but repeating easy examples adds few new kinds of problems. Spend Wisely [96] studies how to divide the computing budget for generating outputs during repeated self-training. Its findings show why the generation budget matters for learning. Related approaches [25,97,98] train models to revise outputs over several attempts, use search histories to improve program generation, or add formal proof data when progress slows. Their shared resource is experience that the current model can produce and a later model can use. Work on evolutionary algorithm discovery [99] similarly updates the language model used for later search. These methods use different feedback and update schedules, so one self-training round can have very different costs.
Task generation changes which problems the learner sees, not only how it solves a fixed set. STP [23] alternates between proposing mathematical claims and trying to prove them. A formal proof checker tests proposed solutions, while a stored collection keeps proofs from the three most recent rounds. The updated models thus affect both future tasks and future solutions. LeanAgent [100] and work on learning formal mathematics through intrinsic motivation [101] also link the supply of problems or proofs to a changing learner. Formal checking gives a strong correctness test within the proof system. It does not show that the chosen problems cover the mathematics the user cares about.
Absolute Zero [24] uses code execution to support task proposal and problem solving. One model takes both roles, and code execution provides rewards without a supplied set of reasoning examples. This reduces dependence on a fixed human task collection. It still depends on the initial model, the allowed task format, the code execution system and the training objective. Shared weights make it harder to identify the cause of a gain. Removing the proposer loss does not keep the proposer fixed, because solver training still changes the shared model. A separate, fixed generator would more directly test the value of updating the task generator.
For interactive agents, experience includes states, actions and consequences. OpenWebVoyager [30] repeats exploration, filtering and training over three rounds. The updated web agent produces later action sequences, while a fixed GPT-4o filter checks experience and the initial imitation data remain available. AdvEvo-MARL [32] instead updates attacking and defending agents, so each helps shape the other’s future experience. These designs make experience responsive to current abilities. They also leave different parts fixed. A fixed judge does not prevent the explorer from changing, and a changing opponent does not by itself improve the rule that trains either agent. A comparison with a fixed collection of examples tests a different question from a comparison with a fixed generator that can produce new examples.

4.2. Evaluation Mechanisms

During policy training with a reward model, the reward model can remain fixed while the policy changes. Self-Rewarding Language Models [5] remove this separation by using the current model to generate answers and judge candidate responses. The resulting preferences train a later version through direct preference optimization. That version then supplies both responses and judgments. The main analysis compares an initial supervised model with the next two trained versions, and also tests a fourth version. Agreement with human preferences on response pairs held out for testing rises from 78.7% to 80.4% and then 81.7%. Other measures of judgment do not all improve, so the gains depend on what is measured and in which setting.
This arrangement can let evaluation keep pace with stronger responses. It can also make the generator and judge repeat the same errors. Calibrated Self-Rewarding Vision Language Models [38] use external visual signals to check and adjust the model’s own preferences. The shared model changes, but the visual reference provides a check beyond the model’s confidence. Training a separate critic model offers another approach. CTRL [41] trains code critics using the success of corrections made by a fixed generator, and tests their use with other generators. This links critic quality to the usefulness of its feedback. Agreement with a reference score and the ability to help a learner are related but distinct evaluation goals.
An evaluator can also change without updating model weights. REvolve [45] uses a language model to revise reward programs from feedback, then uses those programs to train policies. The reward code is an active part of the evaluation mechanism. ROSKA [46] combines reward design with the reuse of policy knowledge across iterations. Keeping a trained policy can reduce the cost of further training. However, the apparent quality of a reward program then depends on what the policy has already learned. A reward program that works well with a previously trained policy may work poorly with a new one. Comparisons should report both the reward program and the policy used before further training.
For tasks without a simple execution test, evaluation can use a rubric, meaning a written set of scoring criteria. DR Tulu [47] updates rubrics and the stored collection of rubrics during reinforcement learning for deep research. The models that generate and apply these criteria remain fixed. What changes is the written guidance used to judge evidence and responses during later training. RuCL [48] adjusts the contribution of different rubric levels as a multimodal learner develops. By contrast, using a fixed rubric to revise an answer at inference time can improve that answer without changing the evaluator [102]. What the method updates matters more than whether its name includes terms such as self-evolution.
A higher internal score is hard to interpret unless changes in the scoring process are also recorded. The response may have improved, the scoring scale may have changed, or the evaluator may have become easier to satisfy. Keeping reference judgments fixed and testing external outcomes can help separate these explanations. So can testing whether the updated evaluator helps train a new learner. These checks allow evaluation to change while measuring whether the changes help.

4.3. Rules and Programs

The third path changes how proposed improvements are produced. At the prompt level, a task prompt tells a model how to solve a problem, whereas a mutation prompt tells it how to revise a task prompt. Promptbreeder [6] updates both within a collection of candidate prompts. A saved mutation prompt guides later proposals, so the search changes part of the method used to propose improvements. A control that removes mutation-prompt updates uses the same number of evaluations. This supports the value of updating those prompts in the tested setting. Equal evaluation counts do not imply equal token use, and the base model and higher-level search framework remain fixed.
metaTextGrad [62] uses one optimizer to improve other language-model optimizers. It changes their prompts and how they are combined. The optimizer that directs these changes stays fixed. The learned optimizer can then improve other programs, including with different models or datasets. A related distinction appears in memory systems. MemoPilot [103] trains a component that updates memory for a fixed task-solving agent. Learning how to update memory differs from adding one more memory entry. Both can improve later behavior, but they change different parts of the process.
STOP [63] uses an optimizer program that can revise its own code. Given a program and a measure of its quality, the optimizer asks a language model to produce better code. The input can also be the optimizer’s own code. A revised version then performs later optimization. The study compares resulting optimizers with the starting optimizer and tests a selected improved version on five new tasks. This is evidence about a reusable improvement procedure, beyond a better answer on the original task. The reported experiments also show failures with weaker models and declines in some iterations. Using an optimizer to revise itself may help, but a new version can also be worse.
The Darwin Gödel Machine, or DGM [7], applies program change to a coding agent. Agents revise their own tools and code. The resulting agents are saved in an archive and can be selected for later changes. The agent’s harness changes while its model weights remain fixed. In the reported SWE-bench setup, the best discovered agent raises task success from 20.0% to 50.0%. This measures the final agent’s task performance. It does not measure how well that agent improves later agents. A control keeps the initial agent fixed as the system that proposes changes. The reported search has 80 rounds of candidate generation across different agents. This does not mean that one agent improves 80 times in a row. Costs also differ across the main search and this control. The results support useful self-modification within the tested setting, but do not establish the size of its advantage under the same budget.
Huxley-Gödel Machine, or HGM [64], extends this search by choosing which saved agents to modify and where to spend evaluations. Its comparison under 800 evaluations helps examine how a limited evaluation budget is used. The allocation rules remain fixed, even as agent code changes. ADAS [3] shows why the agent being designed and its designer must be considered separately. A fixed meta-agent designs agents and consults an archive of past designs. The archive can supply useful context for later design attempts, but finding a better agent does not show that the meta-agent’s code or model has improved. DGM, HGM and ADAS should therefore be compared component by component, rather than grouped by the quality of their final agents alone.
A system may combine these paths. An updated agent may generate better experience, write a new scoring program and change the prompt used for its next revision. Combining these functions can reduce human work, but makes it harder to identify which update caused a gain. We need to test which saved change helps later improvement, and under what feedback and budget conditions. A loop alone does not answer this question.
Updating one component and changing how several components work together can have different effects. Updating a generator lets it respond to the learner’s current limits. Updating an evaluator may let it identify errors in the generator’s new outputs. Updating the modification program can change how both are used. If all three change at once, however, a later gain has several possible causes. A more accurate judge may reject examples that an earlier generator could already produce. A better generator may instead make an unchanged judge appear more reliable because the candidates become easier to distinguish.
These alternatives suggest different update schedules. A generator can change frequently while a reference evaluator stays fixed for measurement. A new evaluator can first be tested on stored outputs from several generator versions. A revised optimizer can then be compared using the same data and evaluator before the full loop resumes. Testing these updates separately may help identify useful changes and gains that depend on several changes together. This is a proposed experiment, not a requirement for how useful systems must be built.

5. Gains and Failure Modes

An updated improver may better meet the learner’s current needs, but it may also carry errors into later rounds. We examine these gains and risks together because the same design choices often affect both. Figure 3 shows three ways errors and bias can carry into later rounds.

5.1. Experience Quality and Coverage

A stronger generator can supply correct examples that were previously too rare to obtain. However, filtering for success also favors tasks the model already finds easy. Work on imbalance between frequent and rare cases in vision-language self-improvement [104] observes this effect and tests changes to sampling. Generating new tasks, adjusting the curriculum and maintaining variety can help cover cases beyond those already solved [23,36,105]. The useful question is how much new learning value a round adds. The fraction of correct samples alone cannot answer it. The performance of LIMA [106] with a small, carefully selected training set also shows why data volume alone does not measure learning value. Controlled self-training studies [107] report gains on harder tasks and longer sequences. Evaluation should track which kinds of problems the learner can now solve.
Research on model collapse [108,109] explains how repeated training on model-generated data can reduce quality or diversity. Some kinds of data may become poorly represented or disappear from later outputs. Its conclusions depend on assumptions about sampling, model fitting and the mixture of real and synthetic data. They do not imply that every verified self-training loop must fail. Keeping earlier data while adding new data can avoid collapse in the settings studied by Gerstgrasser et al. [110]. Selecting data based on preferences can also support improvement under stated assumptions about quality signals and the separation between peaks in the data distribution [111]. These results show why the choice of data matters. Keeping earlier data helps preserve the range of examples available for training. Filtering and sampling determine which examples receive the most training attention.
Correctness and diversity should therefore be recorded separately. Execution can establish that a program passes the available tests. It cannot show that all useful behaviors appear in the training set. For agents, a successful action sequence may even contain weak reasoning that rewards based only on final outcomes do not discourage. Guided Thought Reinforcement [112] studies this problem in visual environments. Process checks and broader task coverage can address different failures, so one should not be treated as a substitute for the other. Verification bounds for small transformers [113] and reasoning analysis in SPARKLE [114] show why evaluation should test error detection and repair as well as final accuracy.

5.2. Evaluation Adaptation and Drift

A fixed evaluator can become less useful when a learner produces responses unlike those seen during evaluator training. Updating the evaluator or its criteria may help distinguish stronger candidates. Yet model-generated comments and scores can also be wrong. Studies of text and visual judges [115,116,117,118] report sensitivity to presentation, misleading content and visual manipulation. The results apply to the tested judges and settings. They establish reasons to test a scoring process, not a permanent limit on all model-based evaluation. ReaLMistake [119] tests error detection, while text and figure-caption studies [120,121] examine agreement with human ratings. MMStar [122] checks whether correct answers depend on visual input.
The risk is greater when a model judges its own work. The model may make an error while answering and then fail to notice it while judging. Reward overoptimization studies [123] show a problem in controlled experiments. Optimizing the reward used for training, called a proxy reward, can eventually lower a separate reference reward. Other work [124] finds that reward models with similar accuracy on a fixed dataset can lead to policies of different quality after training. Improving agreement on a fixed test set is therefore not enough to establish a better training signal. Inference-time alignment analysis [125] links coverage to reward-model error under explicit assumptions. MONA [126] and capability-seeking RL experiments [127] show why tests should check whether policies gain reward in ways that miss the intended goal.
External checks help when their own limits are stated. Feedback based on visual observations can help correct errors [83], and calibrated self-rewarding [38] uses an external visual reference. Formal proofs and executable tests give stronger checks for claims they can express. They remain incomplete measures of open-ended usefulness. A small, fixed reference set can reveal evaluation drift, meaning changes in scoring standards, while criteria for individual tasks evolve. Periodic independent review can then test whether new criteria preserve the intended objective. This supports adaptation without allowing each version to define its own success.

5.3. Memory, Archives and Error Accumulation

Memory and archives make earlier work available after an update. ExpeL [61] illustrates how agents can extract reusable lessons from interaction. Reflexion [60] keeps written feedback, whereas DGM [7] keeps different agent programs that can be run again. Other agents store mistake summaries, contextual experience or reusable task knowledge [128,129,130,131,132,133]. Reusing useful experience can reduce repeated failures. It also makes the choice of what to store, retrieve and remove part of the learning process. A wrong memory may be reused with more confidence simply because it has appeared before.
Storage capacity alone does not solve this problem. Research on software agents [134] finds that storing an entire task attempt as one memory can hide useful differences when similar tasks need different steps. MemoryAgentBench [135] tests memory retrieval, learning during use, understanding information spread across a long history and forgetting selected information. These functions can conflict. Keeping more history preserves experience but can add outdated or misleading context. Archives keep alternative candidates that would be lost if only the latest version were saved. They also add storage, evaluation and selection costs. Their benefit depends on later access to useful candidates, not only archive size. Replay experiments [136] examine how much earlier learning is preserved when storage space is ample. A paper proposing separate memory components [137] suggests separating fast updates to stored context from slower model-weight changes. Tests should measure later adaptation as well as the ability to recall information.
A saved change may also affect behavior on tasks that were not tested. Your Agent May Misevolve [138] studies safety loss under changes to model parameters, memory and workflows. Related work [139] reports more tool hallucinations after reasoning improvements, and agent security evaluations [140] include attacks on memory and tools. Such results call for checks across several objectives [141]. They do not show that self-improvement always reduces safety. They show why task gains should not be used as a complete measure of a changed system.
Saving earlier versions supports direct comparisons. Researchers can rerun an earlier generator, compare old and new rubrics on the same responses, or test a modification rule on agents it did not recently produce. Records should also include the tasks, feedback and resources used to produce each version.

5.4. Transfer and Cost

Transfer has two meanings here. A discovered agent may solve a new task, or the updated designer may produce a better agent for that task. ADAS [3] reports transfer of discovered designs. STOP [63] tests an improved optimizer on new optimization tasks, while metaTextGrad [62] studies reuse of learned optimizers. These experiments test different things. A model or program that works on new tasks is useful. It does not show that the method used to produce it can also work on new tasks.
A loop uses resources to generate candidates, obtain feedback, train, test and save system states. Work on meta-agent search [142] and synthetic-data bootstrapping [96] shows why the choice of context and the computing budget affect the outcome. A method can use fewer evaluations while each evaluation becomes more expensive. Reusing policy weights can reduce training costs, while larger memories can increase inference costs. Comparisons should report both useful performance at a fixed budget and the budget needed to reach a fixed level of performance. The cost of learning an improver should then be spread over a stated number of future uses. Its practical value therefore depends on how often it will be used, not only on its best search result. Strategy auctions [143] study how to assign work to models of different sizes, while test-time discovery [144] makes the cost of adapting to one problem explicit. These results show why reports should separate the cost of reusable learning from the cost of work on one improvee. Code-refinement studies [145] and noisy evolution-strategy analysis [146] show a tradeoff between proposing more candidates and evaluating each more carefully. The useful balance depends on noise and cost.

6. Evaluation and Outlook

6.1. Evaluating Updated Improvers

Evaluating an improvee and evaluating its improver answer different questions. First, we need to identify what changes in the improver, what is kept and how it is used later. A separate comparison tests whether the updated improver produces better improvees under comparable conditions. Table 2 applies this distinction to representative systems. It records positive results as well as the limits of the comparisons.
A direct experiment uses two copies of the same improvee. One receives changes from the updated improver, the other from an earlier version of that improver kept fixed. Both improvers have comparable budgets, access to feedback and task samples. We then compare the quality of the resulting improvees. Independent tests should not guide training or the choice of updates. Figure 4 illustrates this comparison. Removing one loss term does not keep a component fixed if it shares weights with another component that is still being trained. Giving one improver a larger archive also makes it harder to separate better modification from more context.
The existing literature supplies parts of this evidence. Self-Rewarding [5] tests judgments against human preferences kept for evaluation. Promptbreeder [6] tests what happens when mutation-prompt updates are removed. STOP [63] tests an improved optimizer on new tasks. DGM [7] includes a control that uses the fixed initial agent to make changes, and HGM [64] compares search within an evaluation limit. These results support specific kinds of improved improvers. They cannot be ranked directly because they update different components, use different budgets and measure different outcomes.
Benchmarks can support this comparison without replacing it. MLAgentBench [147] and MLE-bench [148] test work on machine-learning tasks. VeRO [149] records versions, observations and rewards when agents optimize other agents. Software benchmarks [150,151,152] test software fixes, debugging or code restructuring. Their scores measure what an agent can do in a defined environment. An improvement study must additionally identify how the evaluated agent was obtained, including unsuccessful proposals and the cost of search. Repeated trials are needed because one successful agent does not show how reliably the process produces good agents.
Repeatedly using test results to select updates can favor those tests. Fresh tests help check progress beyond them. Work on dynamic benchmarks [153] studies this problem under explicit assumptions. LiveCodeBench [154] and checks for evaluation data appearing in training [155] address related concerns. Changes to benchmarks and hosted models should be recorded separately from changes to the improver. Work on API-based evaluation [156] also recommends recording model versions and checking for changes in behavior. These records help distinguish changes in the improvement process from changes in the conditions used to test it.

6.2. Future Directions

The comparisons above support four broader questions about improvement processes.
When does improvement become recursive? A process becomes recursive when changes to its improvement mechanism help produce its own later updates. Promptbreeder [6] and STOP [63] provide examples, but such a loop does not show that improvement ability grows. Which conditions allow a changed mechanism to produce better later updates? Studies should separate changes to generators, evaluators and update procedures, while tracking which goals and checks stay fixed. This can reveal when progress depends on a changing improver and when it depends mainly on a fixed external process.
Is there an improvement scaling law? Do gains follow a predictable pattern as compute, feedback and improver ability increase? This is an open empirical question. Budgets must include generation, evaluation, training and failed proposals, as well as the cost of creating the improver. Budget-aware coding benchmarks [157] illustrate why the costs of calls and tests matter. At each budget, comparing fixed and updated improvers can test whether changing the improver adds value beyond additional search. Any proposed relationship should be tested on tasks and improvees beyond those used to fit it.
What limits the improvement horizon? Here, the improvement horizon is the number of update rounds over which a process can sustain useful progress under stated resource limits and independent tests. It differs from the length of a single agent task. Gains need not occur in every round. Progress may stop because new experience becomes scarce, evaluation becomes unreliable, saved errors accumulate or costs exceed benefits. Studies should identify which limits appear first and whether changes to data, evaluation or memory extend progress. The horizon therefore depends on the full process, not only on the model.
Can an improver generalize across improvees? A broadly useful improver could reduce the need to design a separate update process for each AI system. STOP [63] and metaTextGrad [62] offer early tests, but the question is what transfers. A learned update rule, a stored archive and a finished agent are different results. Studies should apply the same improver to unseen improvees, including different model families and agent harnesses. Any further adjustment and its cost should be reported separately. This would help distinguish general improvement methods from solutions tied to one system.

7. Conclusions

AI4AI concerns the processes through which AI improves AI systems and their development methods. Foundation models and agents contribute through experience, feedback and system changes. Model learning and harness engineering both fit this view. When the improver becomes the improvee, these changes can affect how later improvements are produced.
The reviewed systems show useful changes within specific settings. They do not yet establish that recursive updates lead to sustained gains across AI systems. Evaluation must test both the resulting improvees and the process that produced them. This provides a basis for studying when improvement becomes recursive, how gains depend on resources, what limits continued progress and which methods transfer. The broader aim is to understand and build improvement processes that remain useful as models, tasks and conditions change.

Data Availability Statement

The companion website for this Review is available at https://3dagentworld.github.io/AI4AI-survey.

References

  1. Wang, Y.; Kordi, Y.; Mishra, S.; Liu, A.; Smith, N.A.; Khashabi, D.; Hajishirzi, H. Self-Instruct: Aligning Language Models with Self-Generated Instructions. In Proceedings of the Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Toronto, Canada, 2023; pp. 13484–13508. [Google Scholar] [CrossRef]
  2. Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al. Training language models to follow instructions with human feedback. In Proceedings of the Advances in Neural Information Processing Systems; Curran Associates, Inc., 2022; Vol. 35, pp. 27730–27744. [Google Scholar] [CrossRef]
  3. Hu, S.; Lu, C.; Clune, J. Automated Design of Agentic Systems. In Proceedings of the The Thirteenth International Conference on Learning Representations, 2025. [Google Scholar]
  4. Zelikman, E.; Wu, Y.; Mu, J.; Goodman, N. STaR: Bootstrapping Reasoning With Reasoning. In Proceedings of the Advances in Neural Information Processing Systems, 2022. [Google Scholar]
  5. Yuan, W.; Pang, R.Y.; Cho, K.; Li, X.; Sukhbaatar, S.; Xu, J.; Weston, J.E. Self-Rewarding Language Models. In Proceedings of the Proceedings of the 41st International Conference on Machine Learning. PMLR, 21–27 Jul 2024, Vol. 235, Proceedings of Machine Learning Research, pp. 57905–57923.
  6. Fernando, C.; Banarse, D.S.; Michalewski, H.; Osindero, S.; Rocktäschel, T. Promptbreeder: Self-Referential Self-Improvement via Prompt Evolution. In Proceedings of the Proceedings of the 41st International Conference on Machine Learning. PMLR, 21–27 Jul 2024, Vol. 235, Proceedings of Machine Learning Research, pp. 13481–13544.
  7. Zhang, J.; Hu, S.; Lu, C.; Lange, R.T.; Clune, J. Darwin Gödel Machine: Open-Ended Evolution of Self-Improving Agents. In Proceedings of the The Fourteenth International Conference on Learning Representations, 2026. [Google Scholar]
  8. Tao, Z.; Lin, T.E.; Chen, X.; Li, H.; Wu, Y.; Li, Y.; Jin, Z.; Huang, F.; Tao, D.; Zhou, J. A Survey on Self-Evolution of Large Language Models. arXiv 2024, arXiv:2404.14387. [Google Scholar] [CrossRef]
  9. ang Gao, H.; Geng, J.; Hua, W.; Hu, M.; Juan, X.; Liu, H.; Liu, S.; Qiu, J.; Qi, X.; Ren, Q.; et al. A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence. Trans. Mach. Learn. Res. Survey Certification. 2026. [Google Scholar]
  10. Fang, J.; Peng, Y.; Zhang, X.; Wang, Y.; Yi, X.; Zhang, G.; Xu, Y.; Wu, B.; Liu, S.; Li, Z.; et al. A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems. arXiv 2025, arXiv:2508.07407. [Google Scholar] [CrossRef]
  11. Chen, M.; Wang, L.; Qu, B. Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops. arXiv 2026, arXiv:2607.07663. [Google Scholar] [CrossRef]
  12. Wang, K.; Hou, W.; Yan, Y.; Jia, H.; Zhong, Z.; Ren, B.; Chen, Y.; Zhang, S.; Shen, Y.; Xiao, J.; et al. Diving into Reliable Self-Evolving Agents: A Survey. OpenReview Archive preprint 2026. [Google Scholar]
  13. Wu, K.; Lyu, H.; Luo, Z.; Wang, C.; Ye, S.; Lin, J.; Ji, X.; Jiang, B.; Wang, S.; Wang, Z.; et al. AI4AI Survey: From Long-Horizon Agents to Recursive Self-Improvement—Definitions, Reliable Horizons, and Open Problems. Preprints.org 2026. Preprint, version 1, posted 28 August 2026.
  14. Cheng, J.; Clark, P.; Richardson, K. Language Modeling by Language Models. In Proceedings of the The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025. [Google Scholar]
  15. Gao, Z.; Du, Y.; Zhang, X.; Ma, X.; Han, W.; Zhu, S.C.; Li, Q. CLOVA: A Closed-LOop Visual Assistant with Tool Usage and Update. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2024; pp. 13258–13268. [Google Scholar]
  16. Wu, Q.; Bansal, G.; Zhang, J.; Wu, Y.; Li, B.; Zhu, E.; Jiang, L.; Zhang, X.; Zhang, S.; Liu, J.; et al. AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversations. In Proceedings of the First Conference on Language Modeling, 2024. [Google Scholar]
  17. Hong, S.; Zhuge, M.; Chen, J.; Zheng, X.; Cheng, Y.; Wang, J.; Zhang, C.; Wang, Z.; Yau, S.K.S.; Lin, Z.; et al. MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework. In Proceedings of the The Twelfth International Conference on Learning Representations, 2024. [Google Scholar]
  18. Yang, J.; Jimenez, C.E.; Wettig, A.; Lieret, K.; Yao, S.; Narasimhan, K.; Press, O. SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering. In Proceedings of the NeurIPS, 2024. [Google Scholar]
  19. Zhang, H.; Zhang, S.; Li, K.; Zhang, C.; Chen, Y.; Zhang, Y.; Bai, L.; Hu, S. Self-Harness: Harnesses That Improve Themselves. arXiv 2026, arXiv:cs.CL/2606.09498. [Google Scholar]
  20. Xu, C.; Sun, Q.; Zheng, K.; Geng, X.; Zhao, P.; Feng, J.; Tao, C.; Lin, Q.; Jiang, D. WizardLM: Empowering Large Pre-Trained Language Models to Follow Complex Instructions. In Proceedings of the The Twelfth International Conference on Learning Representations, 2024. [Google Scholar]
  21. Sun, Z.; Shen, Y.; Zhou, Q.; Zhang, H.; Chen, Z.; Cox, D.D.; Yang, Y.; Gan, C. Principle-Driven Self-Alignment of Language Models from Scratch with Minimal Human Supervision. In Proceedings of the NeurIPS, 2023. [Google Scholar]
  22. Askari, A.; Petcu, R.; Meng, C.; Aliannejadi, M.; Abolghasemi, A.; Kanoulas, E.; Verberne, S. SOLID: Self-seeding and Multi-intent Self-instructing LLMs for Generating Intent-aware Information-Seeking Dialogs. In Proceedings of the Findings of the Association for Computational Linguistics: NAACL 2025, Albuquerque, New Mexico, 2025; pp. 6390–6410. [Google Scholar] [CrossRef]
  23. Dong, K.; Ma, T. STP: Self-play LLM Theorem Provers with Iterative Conjecturing and Proving. In Proceedings of the Forty-second International Conference on Machine Learning, 2025. [Google Scholar]
  24. Zhao, A.; Wu, Y.; Yue, Y.; Wu, T.; Xu, Q.; Yue, Y.; Lin, M.; Wang, S.; Wu, Q.; Zheng, Z.; et al. Absolute Zero: Reinforced Self-play Reasoning with Zero Data. In Proceedings of the The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025. [Google Scholar]
  25. Lin, Y.; Tang, S.; Lyu, B.; Yang, Z.; Chung, J.H.; Zhao, H.; Jiang, L.; Geng, Y.; Ge, J.; Sun, J.; et al. Goedel-Prover-V2: Scaling Formal Theorem Proving with Scaffolded Data Synthesis and Self-Correction. In Proceedings of the The Fourteenth International Conference on Learning Representations, 2026. [Google Scholar]
  26. Kung, P.N.; Yin, F.; Wu, D.; Chang, K.W.; Peng, N. Active Instruction Tuning: Improving Cross-Task Generalization by Training on Prompt Sensitive Tasks. In Proceedings of the Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Singapore, 2023; pp. 1813–1829. [Google Scholar] [CrossRef]
  27. Li, Y.; Lu, T.; Li, Y.; Chen, Y.; Huang, W.C.; Jiang, W.; Wang, H.; Zheng, H.T.; Yu, P.S. Teaching According to Talents! Instruction Tuning LLMs with Competence-Aware Curriculum Learning. In Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2025, Suzhou, China, 2025; pp. 11724–11741. [Google Scholar] [CrossRef]
  28. Zhang, J.; Xia, W.; Dong, H.; Lin, Q.; Cao, J. AP2O-Coder: Adaptively Progressive Preference Optimization for Reducing Compilation and Runtime Errors in LLM-Generated Code. In Proceedings of the Proceedings of the AAAI Conference on Artificial Intelligence, 2026; pp. 34701–34709. [Google Scholar] [CrossRef]
  29. Zhan, R.; Li, Y.; Wang, Z.; Qu, X.; Liu, D.; Shao, J.; Wong, D.F.; Cheng, Y. ExGRPO: Learning to Reason from Experience. In Proceedings of the The Fourteenth International Conference on Learning Representations, 2026. [Google Scholar]
  30. He, H.; Yao, W.; Ma, K.; Yu, W.; Zhang, H.; Fang, T.; Lan, Z.; Yu, D. OpenWebVoyager: Building Multimodal Web Agents via Iterative Real-World Exploration, Feedback and Optimization. In Proceedings of the Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Vienna, Austria, 2025; pp. 27545–27564. [Google Scholar] [CrossRef]
  31. Yang, Q.; Wang, X.; Perszyk, D.; Wang, Y.X. Self-Guided Hierarchical Exploration for Generalist Foundation Model Web Agents. In Proceedings of the The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025. [Google Scholar]
  32. Pan, Z.; Zhang, Y.; Liu, Z.; Tang, Y.Y.; Zhang, Z.; Luo, H.; Xu, C.; Han, Y.; Zhang, J.; Wu, D.; et al. AdvEvo-MARL: Shaping Internalized Safety through Adversarial Co-Evolution in Multi-Agent Reinforcement Learning. In Proceedings of the Forty-third International Conference on Machine Learning, 2026. [Google Scholar]
  33. Liu, B.; Yu, S.; Liu, Z.; Guertler, L.; Qi, P.; Balcells, D.; Liu, M.; Tan, C.; Shi, W.; Lin, M.; et al. SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning. In Proceedings of the The Fourteenth International Conference on Learning Representations, 2026. [Google Scholar]
  34. Wu, J.; Schoop, E.; Leung, A.; Barik, T.; Bigham, J.; Nichols, J. UICoder: Finetuning Large Language Models to Generate User Interface Code through Automated Feedback. In Proceedings of the Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), Mexico City, Mexico, 2024; pp. 7511–7525. [Google Scholar] [CrossRef]
  35. Qiu, H.; Gao, M.; Qian, L.; Pan, K.; Yu, Q.; Li, J.; Wang, W.; Tang, S.; Zhuang, Y.; Chua, T.S. STEP: Enhancing Video-LLMs’ Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2025; pp. 3284–3294. [Google Scholar]
  36. Zhao, H.; Shen, J.; Zhang, Y.; Gao, S.; Liu, K.; Ma, T.; Zheng, F.; Lin, D.; Zhang, W.; Chen, K. Achieving Olympia-Level Geometry Large Language Model Agent via Complexity Boosting Reinforcement Learning. In Proceedings of the The Fourteenth International Conference on Learning Representations, 2026. [Google Scholar]
  37. Wang, X.; Dai, W.; Qiao, K.; Wang, K.; Chen, P.; Cao, G.; Kangqin; Wang, Z.; Zhang, X.; Liu, Y.; et al. GUI0: Self-Evolving Foundational GUI Agents in Super App Ecosystems. In Proceedings of the Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), San Diego, California, United States, 2026; pp. 44348–44361. [Google Scholar] [CrossRef]
  38. Zhou, Y.; Fan, Z.; Cheng, D.; Yang, S.; Chen, Z.; Cui, C.; Wang, X.; Li, Y.; Zhang, L.; Yao, H. Calibrated Self-Rewarding Vision Language Models. In Proceedings of the The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024. [Google Scholar]
  39. Lee, H.; Phatale, S.; Mansoor, H.; Mesnard, T.; Ferret, J.; Lu, K.R.; Bishop, C.; Hall, E.; Carbune, V.; Rastogi, A.; et al. RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback. In Proceedings of the Forty-first International Conference on Machine Learning, 2024. [Google Scholar]
  40. Asai, A.; Wu, Z.; Wang, Y.; Sil, A.; Hajishirzi, H. Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection. In Proceedings of the The Twelfth International Conference on Learning Representations, 2024. [Google Scholar]
  41. Xie, Z.; chen, J.; Chen, L.; Mao, W.; Xu, J.; Kong, L. Teaching Language Models to Critique via Reinforcement Learning. In Proceedings of the Forty-second International Conference on Machine Learning, 2025. [Google Scholar]
  42. Setlur, A.; Nagpal, C.; Fisch, A.; Geng, X.; Eisenstein, J.; Agarwal, R.; Agarwal, A.; Berant, J.; Kumar, A. Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning. In Proceedings of the The Thirteenth International Conference on Learning Representations, 2025. [Google Scholar]
  43. Yang, R.; Ye, F.; Li, J.; Yuan, S.; Zhang, Y.; Tu, Z.; Li, X.; Yang, D. The Lighthouse of Language: Enhancing LLM Agents via Critique-Guided Improvement. In Proceedings of the The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025. [Google Scholar]
  44. Ma, Y.J.; Liang, W.; Wang, G.; Huang, D.A.; Bastani, O.; Jayaraman, D.; Zhu, Y.; Fan, L.; Anandkumar, A. Eureka: Human-Level Reward Design via Coding Large Language Models. In Proceedings of the The Twelfth International Conference on Learning Representations, 2024. [Google Scholar]
  45. HAZRA, R.; Sygkounas, A.; Persson, A.; Loutfi, A.; Martires, P.Z.D. REvolve: Reward Evolution with Large Language Models using Human Feedback. In Proceedings of the The Thirteenth International Conference on Learning Representations, 2025. [Google Scholar]
  46. Huang, C.; Chang, Y.; Lin, J.; Liang, J.; Zeng, R.; Li, J. Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution. In Proceedings of the Proceedings of the AAAI Conference on Artificial Intelligence; 2025; pp. 14576–14584. [Google Scholar] [CrossRef]
  47. Shao, R.; Asai, A.; Shen, S.Z.; Ivison, H.; Kishore, V.; Zhuo, J.; Zhao, X.; Park, M.; Finlayson, S.G.; Sontag, D.; et al. DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research. In Proceedings of the Forty-third International Conference on Machine Learning, 2026. [Google Scholar]
  48. Chen, Y.; Li, J.; Chen, L.; Gong, Z.; Li, J.; Qin, Z.; Chang, H.; Zhang, L.; Xu, A.; Yang, Z.; et al. RuCL: Stratified Rubric-Based Curriculum Learning for Multimodal Large Language Model Reasoning. In Proceedings of the Forty-third International Conference on Machine Learning, 2026. [Google Scholar]
  49. Yang, C.; Wang, X.; Lu, Y.; Liu, H.; Le, Q.V.; Zhou, D.; Chen, X. Large Language Models as Optimizers. In Proceedings of the The Twelfth International Conference on Learning Representations, 2024. [Google Scholar]
  50. Guo, Q.; Wang, R.; Guo, J.; Li, B.; Song, K.; Tan, X.; Liu, G.; Bian, J.; Yang, Y. Connecting Large Language Models with Evolutionary Algorithms Yields Powerful Prompt Optimizers. In Proceedings of the The Twelfth International Conference on Learning Representations, 2024. [Google Scholar]
  51. Khattab, O.; Singhvi, A.; Maheshwari, P.; Zhang, Z.; Santhanam, K.; A, S.V.; Haq, S.; Sharma, A.; Joshi, T.T.; Moazam, H.; et al. DSPy: Compiling Declarative Language Model Calls into State-of-the-Art Pipelines. In Proceedings of the The Twelfth International Conference on Learning Representations, 2024. [Google Scholar]
  52. Yuksekgonul, M.; Bianchi, F.; Boen, J.; Liu, S.; Lu, P.; Huang, Z.; Guestrin, C.; Zou, J. Optimizing generative AI by backpropagating language model feedback. Nature 2025, 639, 609–616. [Google Scholar] [CrossRef] [PubMed]
  53. Zhang, J.; Xiang, J.; Yu, Z.; Teng, F.; Chen, X.H.; Chen, J.; Zhuge, M.; Cheng, X.; Hong, S.; Wang, J.; et al. AFlow: Automating Agentic Workflow Generation. In Proceedings of the The Thirteenth International Conference on Learning Representations, 2025. [Google Scholar]
  54. Yang, J.; Wan, G.; Zhang, M.; Ye, M. RAAS: LLM Agentic System Architecture Search with GRPO. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026; pp. 34470–34479. [Google Scholar]
  55. Wei, Y.; Huang, Z.; Xu, R.; Wang, H.; XING, W.W. EvoMAS: Heuristics in the Loop—Evolving Smarter Agentic Workflows. In Proceedings of the Forty-third International Conference on Machine Learning, 2026. [Google Scholar]
  56. Yang, Z.; Zeng, W.; Jin, S.; Qian, C.; Luo, P.; Liu, W. NADER: Neural Architecture Design via Multi-Agent Collaboration. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2025; pp. 4452–4461. [Google Scholar]
  57. Ji, Z.; Zhu, G.; Yuan, C.; Huang, Y. RZ-NAS: Enhancing LLM-guided Neural Architecture Search via Reflective Zero-Cost Strategy. In Proceedings of the Forty-second International Conference on Machine Learning, 2025. [Google Scholar]
  58. Ding, K.;Wang, Y.; Xiang, S. EvoVLMA: Evolutionary Vision-Language Model Adaptation. In Proceedings of the Proceedings of the 33rdACMInternational Conference on Multimedia. ACM, 2025,MM’25, p. 4619–4628. [CrossRef]
  59. Wang, G.; Xie, Y.; Jiang, Y.; Mandlekar, A.; Xiao, C.; Zhu, Y.; Fan, L.; Anandkumar, A. Voyager: An Open-Ended Embodied Agent with Large Language Models. Transactions on Machine Learning Research, 2024. [Google Scholar]
  60. Shinn, N.; Cassano, F.; Gopinath, A.; Narasimhan, K.R.; Yao, S. Reflexion: Language agents with verbal reinforcement learning. In Proceedings of the Thirty-seventh Conference on Neural Information Processing Systems, 2023. [Google Scholar]
  61. Zhao, A.; Huang, D.; Xu, Q.; Lin, M.; Liu, Y.J.; Huang, G. ExpeL: LLM Agents Are Experiential Learners. In Proceedings of the AAAI, 2024; pp. 19632–19642. [Google Scholar]
  62. Xu, G.; Yuksekgonul, M.; Guestrin, C.; Zou, J. metaTextGrad: Automatically optimizing language model optimizers. In Proceedings of the The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025. [Google Scholar]
  63. Zelikman, E.; Lorch, E.; Mackey, L.; Kalai, A.T. Self-Taught Optimizer (STOP): Recursively Self-Improving Code Generation. In Proceedings of the First Conference on Language Modeling, 2024. [Google Scholar]
  64. Wang, W.; Piękos, P.; Nanbo, L.; Laakom, F.; Chen, Y.; Ostaszewski, M.; Zhuge, M.; Schmidhuber, J. Huxley-G\”odel Machine: Human-Level Coding Agent Development by an Approximation of the Optimal Self-Improving Machine. In Proceedings of the The Fourteenth International Conference on Learning Representations, 2026. [Google Scholar]
  65. Madaan, A.; Tandon, N.; Gupta, P.; Hallinan, S.; Gao, L.; Wiegreffe, S.; Alon, U.; Dziri, N.; Prabhumoye, S.; Yang, Y.; et al. Self-Refine: Iterative Refinement with Self-Feedback. In Proceedings of the Thirty-seventh Conference on Neural Information Processing Systems, 2023. [Google Scholar]
  66. Chen, X.; Lin, M.; Schärli, N.; Zhou, D. Teaching Large Language Models to Self-Debug. In Proceedings of the The Twelfth International Conference on Learning Representations, 2024. [Google Scholar]
  67. Olausson, T.X.; Inala, J.P.; Wang, C.; Gao, J.; Solar-Lezama, A. Is Self-Repair a Silver Bullet for Code Generation? In Proceedings of the The Twelfth International Conference on Learning Representations, 2024. [Google Scholar]
  68. Tandon, N.; Madaan, A.; Clark, P.; Yang, Y. Learning to repair: Repairing model output errors after deployment using a dynamic memory of feedback. In Proceedings of the Findings of the Association for Computational Linguistics: NAACL 2022, Seattle, United States, 2022; pp. 339–352. [Google Scholar] [CrossRef]
  69. Gao, J.; Ding, X.; Cui, Y.; Zhao, J.; Wang, H.; Liu, T.; Qin, B. Self-Evolving GPT: A Lifelong Autonomous Experiential Learner. In Proceedings of the Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Bangkok, Thailand, 2024; pp. 6385–6432. [Google Scholar] [CrossRef]
  70. Lan, Y.; Hu, Z.; Wang, L.; Wang, Y.; Ye, D.; Zhao, P.; Lim, E.P.; Xiong, H.; Wang, H. LLM-Based Agent Society Investigation: Collaboration and Confrontation in Avalon Gameplay. In Proceedings of the Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Miami, Florida, USA, 2024; pp. 128–145. [Google Scholar] [CrossRef]
  71. Mao, K.; Liu, Z.; Qian, H.; Mo, F.; Deng, C.; Dou, Z. RAG-Studio: Towards In-Domain Adaptation of Retrieval Augmented Generation Through Self-Alignment. In Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2024, Miami, Florida, USA, 2024; pp. 725–735. [Google Scholar] [CrossRef]
  72. Wu, Y.; Jiang, A.Q.; Li, W.; Rabe, M.N.; Staats, C.E.; Jamnik, M.; Szegedy, C. Autoformalization with Large Language Models. In Proceedings of the Advances in Neural Information Processing Systems, 2022. [Google Scholar]
  73. Haluptzok, P.; Bowers, M.; Kalai, A.T. Language Models Can Teach Themselves to Program Better. In Proceedings of the The Eleventh International Conference on Learning Representations, 2023. [Google Scholar]
  74. Zhang, Z.; Xiao, N.; Chai, Q.; Ye, D.;Wang, H. MultiMind: EnhancingWerewolf Agents with Multimodal Reasoning and Theory of Mind. In Proceedings of the Proceedings of the 33rd ACM International Conference on Multimedia. ACM, 2025, MM ’25, pp. 5824–5833. [CrossRef]
  75. Schick, T.; Dwivedi-Yu, J.; Dessì, R.; Raileanu, R.; Lomeli, M.; Hambro, E.; Zettlemoyer, L.; Cancedda, N.; Scialom, T. Toolformer: Language Models Can Teach Themselves to Use Tools. In Proceedings of the NeurIPS, 2023. [Google Scholar]
  76. Hsieh, C.Y.; Li, C.L.; Yeh, C.k.; Nakhost, H.; Fujii, Y.; Ratner, A.; Krishna, R.; Lee, C.Y.; Pfister, T. Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes. In Proceedings of the Findings of the Association for Computational Linguistics: ACL 2023, Toronto, Canada, 2023; pp. 8003–8017. [Google Scholar] [CrossRef]
  77. Tong, Y.; Li, D.; Wang, S.; Wang, Y.; Teng, F.; Shang, J. Can LLMs Learn from Previous Mistakes? Investigating LLMs’ Errors to Boost for Reasoning. In Proceedings of the Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Bangkok, Thailand, 2024; pp. 3065–3080. [Google Scholar] [CrossRef]
  78. Cho, J.; Kang, D.; Kim, H.; Lee, G. Self-Correcting Code Generation Using Small Language Models. In Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2025, Suzhou, China, 2025; pp. 2345–2368. [Google Scholar] [CrossRef]
  79. Yang, Y.; Zhou, T.; Li, K.; Tao, D.; Li, L.; Shen, L.; He, X.; Jiang, J.; Shi, Y. Embodied Multi-Modal Agent trained by an LLM from a Parallel TextWorld. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2024; pp. 26275–26285. [Google Scholar]
  80. Zhang, Z.; Lan, Y.; Chen, Y.; Wang, L.; Wang, X.; Wang, H. DVM: Towards Controllable LLM Agents in Social Deduction Games. In Proceedings of the ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP); IEEE, 2025; pp. 1–5. [Google Scholar] [CrossRef]
  81. Rafailov, R.; Sharma, A.; Mitchell, E.; Manning, C.D.; Ermon, S.; Finn, C. Direct Preference Optimization: Your Language Model is Secretly a Reward Model. In Proceedings of the Advances in Neural Information Processing Systems; Curran Associates, Inc., 2023; Vol. 36, pp. 53728–53741. [Google Scholar] [CrossRef]
  82. Rajaee, S.; Pratik, K.; Cesa, G.; Behboodi, A. Local Look-Ahead Guidance via Verifier-in-the-Loop for Automated Theorem Proving. In Proceedings of the Findings of the Association for Computational Linguistics: ACL 2025, Vienna, Austria, 2025; pp. 16023–16040. [Google Scholar] [CrossRef]
  83. Liao, Y.H.; Mahmood, R.; Fidler, S.; Acuna, D. Can Large Vision-Language Models Correct Semantic Grounding Errors By Themselves? In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2025; pp. 14667–14678. [Google Scholar]
  84. Li, C.; Zheng, Y.; Huang, X.; Fang, T.; Xu, J.; Chen, L.; Song, Y.; Hu, H. WebDevJudge: Evaluating (M)LLMs as Critiques for Web Development Quality. In Proceedings of the The Fourteenth International Conference on Learning Representations, 2026. [Google Scholar]
  85. Deutsch, D.; Dror, R.; Roth, D. On the Limitations of Reference-Free Evaluations of Generated Text. In Proceedings of the Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Abu Dhabi, United Arab Emirates, 2022; pp. 10960–10977. [Google Scholar] [CrossRef]
  86. Son, G.; Yang, D.; Patel, H.L.; Ko, H.; Agarwal, A.; Ahn, S.; Lee, K.H.; Yu, Y. Judging What We Cannot Solve: A Consequence-Based Approach for Oracle-Free Evaluation of Research-Level Math. In Proceedings of the Forty-third International Conference on Machine Learning, 2026. [Google Scholar]
  87. Zhang, Z.; Zhao, P.; Ye, D.; Wang, H. Enhancing Jailbreak Attacks on LLMs via Persona Prompts, 2026. arXiv arXiv:cs.
  88. Pryzant, R.; Iter, D.; Li, J.; Lee, Y.; Zhu, C.; Zeng, M. Automatic Prompt Optimization with “Gradient Descent” and Beam Search. In Proceedings of the Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Singapore, 2023; pp. 7957–7968. [Google Scholar] [CrossRef]
  89. Agrawal, L.A.; Tan, S.; Soylu, D.; Ziems, N.; Khare, R.; Opsahl-Ong, K.; Singhvi, A.; Shandilya, H.; Ryan, M.J.; Jiang, M.; et al. GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning. In Proceedings of the The Fourteenth International Conference on Learning Representations, 2026. [Google Scholar]
  90. Yuan, J.; Yan, X.; Zhang, B.; Chen, T.; Shi, B.; Ouyang, W.; Qiao, Y.; Bai, L.; Zhou, B. Dolphin: Moving Towards Closed-loop Auto-research through Thinking, Practice, and Feedback. In Proceedings of the Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Vienna, Austria, 2025; pp. 21768–21789. [Google Scholar] [CrossRef]
  91. Turcan, A.; Huang, K.; Li, L.; Zhang, M.J. TusoAI: Agentic Optimization for Scientific Methods. In Proceedings of the The Fourteenth International Conference on Learning Representations, 2026. [Google Scholar]
  92. Zeng, J.; Li, C.; Sun, Z.; Zhao, Q.; Zhou, G. tnGPS: Discovering Unknown Tensor Network Structure Search Algorithms via Large Language Models (LLMs). In Proceedings of the Proceedings of the 41st International Conference on Machine Learning. PMLR, 21–27 Jul 2024, Vol. 235, Proceedings of Machine Learning Research, pp. 58329–58347.
  93. Aglietti, V.; Ktena, I.; Schrouff, J.; Sgouritsa, E.; Ruiz, F.; Malek, A.; Bellot, A.; Chiappa, S. FunBO: Discovering Acquisition Functions for Bayesian Optimization with FunSearch. In Proceedings of the Forty-second International Conference on Machine Learning, 2025. [Google Scholar]
  94. Chen, Z.; Deng, Y.; Yuan, H.; Ji, K.; Gu, Q. Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models. In Proceedings of the Proceedings of the 41st International Conference on Machine Learning. PMLR, 21–27 Jul 2024, Vol. 235, Proceedings of Machine Learning Research, pp. 6621–6642.
  95. Zelikman, E.; Harik, G.R.; Shao, Y.; Jayasiri, V.; Haber, N.; Goodman, N. Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking. In Proceedings of the First Conference on Language Modeling, 2024. [Google Scholar]
  96. Yang, P.; Feng, Y.; Chen, Z.; Wu, Y.; Li, Z. Spend Wisely: Maximizing Post-Training Gains in Iterative Synthetic Data Bootstrapping. In Proceedings of the The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025. [Google Scholar]
  97. Qu, Y.; Zhang, T.; Garg, N.; Kumar, A. Recursive Introspection: Teaching Language Model Agents How to Self-Improve. In Proceedings of the The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024. [Google Scholar]
  98. Pourcel, J.; Colas, C.; Oudeyer, P.Y. Self-Improving Language Models for Evolutionary Program Synthesis: A Case Study on ARC-AGI. In Proceedings of the Forty-second International Conference on Machine Learning, 2025. [Google Scholar]
  99. Šurina, A.; Mansouri, A.; Quaedvlieg, L.C.; Seddas, A.; Viazovska, M.; Abbe, E.; Gulcehre, C. Algorithm Discovery With LLMs: Evolutionary Search Meets Reinforcement Learning. In Proceedings of the Second Conference on Language Modeling, 2025. [Google Scholar]
  100. Kumarappan, A.; Tiwari, M.; Song, P.; George, R.J.; Xiao, C.; Anandkumar, A. LeanAgent: Lifelong Learning for Formal Theorem Proving. In Proceedings of the The Thirteenth International Conference on Learning Representations, 2025. [Google Scholar]
  101. Poesia, G.; Broman, D.; Haber, N.; Goodman, N. Learning Formal Mathematics From Intrinsic Motivation. In Proceedings of the The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024. [Google Scholar]
  102. Wan, Y.; Fang, T.; LI, Z.; Huo, Y.; Wang, W.; Mi, H.; Yu, D.; Lyu, M.R. Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification. In Proceedings of the Findings of the Association for Computational Linguistics: ACL 2026, San Diego, California, United States, 2026; pp. 24822–24835. [Google Scholar] [CrossRef]
  103. Cai, Y.; Guo, X.; Huang, X.; Du, J.; can huang; Huang, W.; Ma, W.; Hu, Y.; Zeng, A.; Tang, J.; et al. From Player to Master: Enhancing Test-Time Learning of LLM Agents via Reinforcement Learning over Memory. In Proceedings of the Forty-third International Conference on Machine Learning, 2026. [Google Scholar]
  104. Guo, X.; Xi, Z.; Ding, Y.; Zhai, Y.; Shi, X.; Cai, X.; Gui, T.; Zhang, Q.; Huang, X. Counteracting the Matthew Effect in Self-Improvement of LVLMs through Head-Tail Re-balancing. In Proceedings of the Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), San Diego, California, United States, 2026; pp. 22104–22121. [Google Scholar] [CrossRef]
  105. Tsoukalas, G.; Saha, R.; Thakur, A.; Reguyal, S.; Chaudhuri, S. Learning Interestingness in Automated Mathematical Theory Formation. In Proceedings of the The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025. [Google Scholar]
  106. Zhou, C.; Liu, P.; Xu, P.; Iyer, S.; Sun, J.; Mao, Y.; Ma, X.; Efrat, A.; Yu, P.; YU, L.; et al. LIMA: Less Is More for Alignment. In Proceedings of the Thirty-seventh Conference on Neural Information Processing Systems, 2023. [Google Scholar]
  107. Lee, N.; Cai, Z.; Schwarzschild, A.; Lee, K.; Papailiopoulos, D. Self-Improving Transformers Overcome Easy-to-Hard and Length Generalization Challenges. In Proceedings of the Forty-second International Conference on Machine Learning, 2025. [Google Scholar]
  108. Shumailov, I.; Shumaylov, Z.; Zhao, Y.; Papernot, N.; Anderson, R.; Gal, Y. AI models collapse when trained on recursively generated data. Nature 2024, 631, 755–759. [Google Scholar] [CrossRef] [PubMed]
  109. Fu, S.; Zhang, S.; Wang, Y.; Tian, X.; Tao, D. Towards Theoretical Understandings of Self-Consuming Generative Models. In Proceedings of the Proceedings of the 41st International Conference on Machine Learning. PMLR, 21–27 Jul 2024, Vol. 235, Proceedings of Machine Learning Research, pp. 14228–14255.
  110. Gerstgrasser, M.; Schaeffer, R.; Dey, A.; Rafailov, R.; Pai, D.B.; Sleight, H.; Hughes, J.; Korbak, T.; Agrawal, R.; Gromov, A.; et al. Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data. In Proceedings of the ECCV 2024 Workshop The Dark Side of Generative AIs and Beyond, 2024. [Google Scholar]
  111. Falahati, A.; Amiri, M.M.; Larson, K.; Golab, L. Curated Synthetic Data Doesn’t Have to Collapse: A Theoretical Study of Generative Retraining with Pluralistic Preferences. In Proceedings of the Forty-third International Conference on Machine Learning, 2026. [Google Scholar]
  112. Wei, T.; Yang, Y.; Xing, J.; Shi, Y.; Lu, Z.; Ye, D. GTR: Guided Thought Reinforcement Prevents Thought Collapse in RL-based VLM Agent Training. In Proceedings of the Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), October 2025; pp. 18855–18865. [Google Scholar]
  113. Yu, Z.; Xia, W.; Yan, X.; XU, B.; Zhang, H.; Du, Y.; Wang, J. Self-Verifying Reflection Helps Transformers with CoT Reasoning. In Proceedings of the The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025. [Google Scholar]
  114. Wang, J.; Ming, Y.; Ke, Z.; Xiong, C.; Joty, S.; Albarghouthi, A.; Sala, F. Beyond Accuracy: Dissecting Mathematical Reasoning for LLMs Under Reinforcement Learning. In Proceedings of the The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025. [Google Scholar]
  115. Zheng, L.; Chiang, W.L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E.P.; et al. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. In Proceedings of the NeurIPS, 2023. [Google Scholar]
  116. Shen, C.; Cheng, L.; Nguyen, X.P.; You, Y.; Bing, L. Large Language Models are Not Yet Human-Level Evaluators for Abstractive Summarization. In Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, 2023; pp. 4215–4233. [Google Scholar] [CrossRef]
  117. Chen, G.H.; Chen, S.; Liu, Z.; Jiang, F.; Wang, B. Humans or LLMs as the Judge? A Study on Judgement Bias. In Proceedings of the Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Miami, Florida, USA, 2024; pp. 8301–8327. [Google Scholar] [CrossRef]
  118. Hwang, Y.; Lee, D.; Min, K.; Kang, T.; Kim, Y.; Jung, K. Fooling the LVLM Judges: Visual Biases in LVLM-Based Evaluation. In Proceedings of the Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, Suzhou, China, 2025; pp. 23186–23205. [Google Scholar] [CrossRef]
  119. Kamoi, R.; Das, S.S.S.; Lou, R.; Ahn, J.J.; Zhao, Y.; Lu, X.; Zhang, N.; Zhang, Y.; Zhang, H.R.; Vummanthala, S.R.; et al. Evaluating LLMs at Detecting Errors in LLM Responses. In Proceedings of the First Conference on Language Modeling, 2024. [Google Scholar]
  120. Chiang, C.H.; Lee, H.y. A Closer Look into Using Large Language Models for Automatic Evaluation. In Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, 2023; pp. 8928–8942. [Google Scholar] [CrossRef]
  121. Hsu, T.Y.; Huang, C.Y.; Rossi, R.; Kim, S.; Giles, C.L.; Huang, T.H.K. GPT-4 as an Effective Zero-Shot Evaluator for Scientific Figure Captions. In Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, 2023; pp. 5464–5474. [Google Scholar] [CrossRef]
  122. Chen, L.; Li, J.; Dong, X.; Zhang, P.; Zang, Y.; Chen, Z.; Duan, H.; Wang, J.; Qiao, Y.; Lin, D.; et al. Are We on the Right Way for Evaluating Large Vision-Language Models? In Proceedings of the The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024. [Google Scholar]
  123. Gao, L.; Schulman, J.; Hilton, J. Scaling Laws for Reward Model Overoptimization. In Proceedings of the Proceedings of the 40th International Conference on Machine Learning. PMLR, 23–29 Jul 2023, Vol. 202, Proceedings of Machine Learning Research, pp. 10835–10866.
  124. Wen, X.; Lou, J.; Lu, Y.; Lin, H.; XingYu; Lu, X.; He, B.; Han, X.; Zhang, D.; Sun, L. Rethinking Reward Model Evaluation: Are We Barking up the Wrong Tree? In Proceedings of the The Thirteenth International Conference on Learning Representations, 2025. [Google Scholar]
  125. Huang, A.; Block, A.; Liu, Q.; Jiang, N.; Krishnamurthy, A.; Foster, D.J. Is Best-of-N the Best of Them? Coverage, Scaling, and Optimality in Inference-Time Alignment. In Proceedings of the Forty-second International Conference on Machine Learning, 2025. [Google Scholar]
  126. Farquhar, S.; Varma, V.; Lindner, D.; Elson, D.; Biddulph, C.; Goodfellow, I.; Shah, R. MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking. In Proceedings of the Forty-second International Conference on Machine Learning, 2025. [Google Scholar]
  127. Zhou, Y.; Huang, Y.; Bao, H.; Guo, K.; Liang, Z.; Chen, P.Y.; Gao, T.; Geyer, W.; Moniz, N.; Chawla, N.V.; et al. Alignment Risks from Capability-Seeking RL Training. In Proceedings of the Forty-third International Conference on Machine Learning, 2026. [Google Scholar]
  128. Su, X.; Zhang, Y.; Luo, H.; Liu, X.; Huang, L. Mistake Notebook Learning: Batch-Clustered Failures for Training-Free Agent Adaptation. In Proceedings of the Findings of the Association for Computational Linguistics: ACL 2026, San Diego, California, United States, 2026; pp. 14629–14645. [Google Scholar] [CrossRef]
  129. Liu, Y.; Si, C.; Narasimhan, K.R.; Yao, S. Contextual Experience Replay for Self-Improvement of Language Agents. In Proceedings of the Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Vienna, Austria, 2025; pp. 14179–14198. [Google Scholar] [CrossRef]
  130. Ge, Y.; Romeo, S.; Cai, J.; Sunkara, M.; Zhang, Y. SAMULE: Self-Learning Agents Enhanced by Multi-level Reflection. In Proceedings of the Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, Suzhou, China, 2025; pp. 16591–16610. [Google Scholar] [CrossRef]
  131. Zhang, Z.; Bu, W.; Pan, K.; Miao, B.; Zhang, W.; Wang, G.; Ji, W.; Tang, R.; Li, J.; Tang, S. Evolving Generalist Virtual Agents with Generative and Associative Memory. In Proceedings of the Proceedings of the AAAI Conference on Artificial Intelligence, 2026; pp. 13006–13014. [Google Scholar] [CrossRef]
  132. Chai, Q.; Shen, W.; Yao, N.; Xia, Y.; Zhao, K.; Ma, J.; Lin, G.; Wang, H. EvolveNav: Proactive Preflection and Self-Evolving Memory for Zero-Shot Object Goal Navigation. arXiv 2026, arXiv:cs.AI/2606.18235. [Google Scholar]
  133. Zhang, Z.; He, J.; Cai, Y.; Ye, D.; Zhao, P.; Feng, R.; Wang, H. Genesis: Evolving Attack Strategies for LLM Web Agent Red-Teaming. arXiv 2026, arXiv:cs.AI/2510.18314. [Google Scholar]
  134. Shen, K.; Zhang, J.; Sun, C.; Zeng, W.; Yue, Y. Structurally Aligned Subtask-Level Memory for Software Engineering Agents. In Proceedings of the Forty-third International Conference on Machine Learning, 2026. [Google Scholar]
  135. Hu, Y.; Wang, Y.; McAuley, J. Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions. In Proceedings of the The Fourteenth International Conference on Learning Representations, 2026. [Google Scholar]
  136. Cho, D.; Moon, T.; Chunara, R.; Cho, K.; Cha, S. Forget Forgetting: Continual Learning in a World of Abundant Memory. In Proceedings of the The Fourteenth International Conference on Learning Representations, 2026. [Google Scholar]
  137. Dorovatas, V.; Schwerin, M.; Bagdanov, A.D.; Caccia, L.; Carta, A.; Charlin, L.; Hammer, B.; Hayes, T.L.; Hess, T.; Kanan, C.; et al. Position: Modular Memory is the Key to Continual Learning Agents. In Proceedings of the Forty-third International Conference on Machine Learning Position Paper Track, 2026. [Google Scholar]
  138. Shao, S.; Ren, Q.; Liu, D.; Qian, C.; Wei, B.; Guo, D.; JingYi, Y.; Song, X.; Zhang, L.; Zhang, W.; et al. Your Agent May Misevolve: Emergent Risks in Self-evolving LLM Agents. In Proceedings of the The Fourteenth International Conference on Learning Representations, 2026. [Google Scholar]
  139. Yin, C.; Sha, Z.; Cui, S.; Meng, C.; Li, Z. The Reasoning Trap: How Enhancing LLM Reasoning Amplifies Tool Hallucination. In Proceedings of the Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), San Diego, California, United States, 2026; pp. 8310–8328. [Google Scholar] [CrossRef]
  140. Zhang, H.; Huang, J.; Mei, K.; Yao, Y.; Wang, Z.; Zhan, C.; Wang, H.; Zhang, Y. Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents. In Proceedings of the The Thirteenth International Conference on Learning Representations, 2025. [Google Scholar]
  141. Agnihotri, A.; Jain, R.; Ramachandran, D.; Wen, Z. Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models. In Proceedings of the Forty-third International Conference on Machine Learning, 2026. [Google Scholar]
  142. El, B.; Yuksekgonul, M.; Zou, J. Inefficiencies of Meta Agents for Agent Design. In Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2025, Suzhou, China, 2025; pp. 20815–20824. [Google Scholar] [CrossRef]
  143. Alazraki, L.; Shen, W.F.; Bachrach, Y.; Mathur, A. Scaling Small Agents Through Strategy Auctions. In Proceedings of the Forty-third International Conference on Machine Learning, 2026. [Google Scholar]
  144. Yuksekgonul, M.; Koceja, D.; Li, X.; Bianchi, F.; McCaleb, J.; Wang, X.; Kautz, J.; Choi, Y.; Zou, J.; Guestrin, C.; et al. Learning to Discover at Test Time. In Proceedings of the Forty-third International Conference on Machine Learning, 2026. [Google Scholar]
  145. Tang, H.; Hu, K.; Zhou, J.P.; Zhong, S.C.; Zheng, W.L.; Si, X.; Ellis, K. Code Repair with LLMs gives an Exploration-Exploitation Tradeoff. In Proceedings of the The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024. [Google Scholar]
  146. Wang, S.; Lu, Z. Depth over Fidelity in Fixed-Budget Noisy Evolution Strategies. In Proceedings of the Forty-third International Conference on Machine Learning, 2026. [Google Scholar]
  147. Huang, Q.; Vora, J.; Liang, P.; Leskovec, J. MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation. In Proceedings of the Proceedings of the 41st International Conference on Machine Learning. PMLR, 21–27 Jul 2024, Vol. 235, Proceedings of Machine Learning Research, pp. 20271–20309.
  148. Chan, J.S.; Chowdhury, N.; Jaffe, O.; Aung, J.; Sherburn, D.; Mays, E.; Starace, G.; Liu, K.; Maksin, L.; Patwardhan, T.; et al. MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering. In Proceedings of the The Thirteenth International Conference on Learning Representations, 2025. [Google Scholar]
  149. Ursekar, V.; Shanker, A.; Chatrath, V.; Xue, Y.; Denton, S.M. VeRO: A Harness for Agents to Optimize Agents. In Proceedings of the Forty-third International Conference on Machine Learning, 2026. [Google Scholar]
  150. Jimenez, C.E.; Yang, J.; Wettig, A.; Yao, S.; Pei, K.; Press, O.; Narasimhan, K. SWE-bench: Can Language Models Resolve Real-world Github Issues? In Proceedings of the International Conference on Learning Representations, 2024, Vol. 2024, pp. 54107–54157.
  151. Tian, R.; Ye, Y.; Qin, Y.; Cong, X.; Lin, Y.; Pan, Y.; Wu, Y.; Haotian, H.; Weichuan, L.; Liu, Z.; et al. DebugBench: Evaluating Debugging Capability of Large Language Models. In Proceedings of the Findings of the Association for Computational Linguistics: ACL 2024, Bangkok, Thailand, 2024; pp. 4173–4198. [Google Scholar] [CrossRef]
  152. Gautam, D.; Garg, S.; Jang, J.; Sundaresan, N.; Moghaddam, R.Z. RefactorBench: Evaluating Stateful Reasoning in Language Agents Through Code. In Proceedings of the The Thirteenth International Conference on Learning Representations, 2025. [Google Scholar]
  153. Shirali, A.; Abebe, R.; Hardt, M. A Theory of Dynamic Benchmarks. In Proceedings of the The Eleventh International Conference on Learning Representations, 2023. [Google Scholar]
  154. Jain, N.; Han, K.; Gu, A.; Li, W.D.; Yan, F.; Zhang, T.; Wang, S.; Solar-Lezama, A.; Sen, K.; Stoica, I. LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code. In Proceedings of the The Thirteenth International Conference on Learning Representations, 2025. [Google Scholar]
  155. Safarzadeh, M.; Patel, H.L.; Oroojlooy, A.; Horwood, G.; Roth, D. SPENCE: A Syntactic Probe for Detecting Contamination in NL2SQL Benchmarks. In Proceedings of the Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), San Diego, California, United States, 2026; pp. 20222–20239. [Google Scholar] [CrossRef]
  156. Gupta, U. Position: ICML Should Treat Hosted LLM APIs as Versioned Dependencies and Require Drift-Audit Artifacts. In Proceedings of the Forty-third International Conference on Machine Learning Position Paper Track, 2026. [Google Scholar]
  157. Zhou, L.; Shi, J.; Gao, J.; Wang, D. Credit-Budgeted ICPC-Style Coding: When Agents Must Pay for Every Decision. In Proceedings of the The Fourteenth International Conference on Learning Representations, 2026. [Google Scholar]
Figure 1. Taxonomy of AI4AI by the role that supplies improvement. The tree links three roles to functional groups and representative methods. Each method is linked to its reference. A method may serve several roles. Its position identifies a function. It does not show that the improver changes or becomes better at improving AI.
Figure 1. Taxonomy of AI4AI by the role that supplies improvement. The tree links three roles to functional groups and representative methods. Each method is linked to its reference. A method may serve several roles. Its position identifies a function. It does not show that the improver changes or becomes better at improving AI.
Preprints 233114 g001
Figure 2. Overview of AI for AI and the scope of this Review. An AI4AI process links an improver to an improvee through experience, feedback or system changes. The Review follows these processes from fixed to changing improvers, then considers gains, failures and evaluation. A changed improver does not necessarily produce better improvees.
Figure 2. Overview of AI for AI and the scope of this Review. An AI4AI process links an improver to an improvee through experience, feedback or system changes. The Review follows these processes from fixed to changing improvers, then considers gains, failures and evaluation. A changed improver does not necessarily produce better improvees.
Preprints 233114 g002
Figure 3. How errors and bias can carry into later rounds. (a) Success-based selection can narrow training coverage. (b) A mistaken judgment admits a wrong answer into training. (c) A saved rule conflicts with a new task. These examples illustrate mechanisms rather than measured outcomes.
Figure 3. How errors and bias can carry into later rounds. (a) Success-based selection can narrow training coverage. (b) A mistaken judgment admits a wrong answer into training. (c) A saved rule conflicts with a new task. These examples illustrate mechanisms rather than measured outcomes.
Preprints 233114 g003
Figure 4. Two questions for evaluating an updated improver. (a) Trace changes kept in the improver and its later use. (b) Compare fixed and updated improvers using two copies of the same improvee under comparable conditions. Independent tests assess the resulting improvees. Repeated comparisons with other improvees examine quality, cost, transfer and failures.
Figure 4. Two questions for evaluating an updated improver. (a) Trace changes kept in the improver and its later use. (b) Compare fixed and updated improvers using two copies of the same improvee under comparable conditions. Independent tests assess the resulting improvees. Repeated comparisons with other improvees examine quality, cost, transfer and failures.
Preprints 233114 g004
Table 1. Three roles through which AI supplies improvement. Each row shows a function, selected examples and their contribution. A system can serve more than one role.
Table 1. Three roles through which AI supplies improvement. Each row shows a function, selected examples and their contribution. A system can serve more than one role.
Function Examples Contribution
Learning experience
Training data Self-Instruct [1], STaR [4] Instructions and reasoning traces
Interactive practice OpenWebVoyager [30], AdvEvo-MARL [32] Action sequences and opponent tasks
Evaluation feedback
Scores and comments Self-Rewarding [5], CTRL [41] Preferences and repair suggestions
Rewards and criteria REvolve [45], DR Tulu [47] Reward code and written criteria
System changes
AI system design TextGrad [52], ADAS [3], Genesys [14] Prompts, workflows and architectures
Tools and memory Voyager [59], Reflexion [60], CLOVA [15] Reusable code and experience
Improvement procedures Promptbreeder [6], STOP [63], DGM [7] Prompts and code that guide updates
Table 2. Results from representative studies of updated improvers. Each row lists the updated component, the reported evaluation and a limit on what the result shows.
Table 2. Results from representative studies of updated improvers. Each row lists the updated component, the reported evaluation and a limit on what the result shows.
System Updated component Evaluation Key limitation
STaR [4] Reasoning generator Gains over training rounds Needs fixed generator at equal cost
Absolute Zero [24] Shared proposer and solver Training without proposer loss Shared weights link both updates
Self-Rewarding [5] Shared responder and judge Checks against human preferences Judge change not tested separately
REvolve [45] Reward code Comparison of reward search Designer model stays fixed
Promptbreeder [6] Mutation prompts Control with equal evaluations Equal token cost not shown
metaTextGrad [62] Optimizer prompts and structure Tests across models and datasets The optimizer updating it stays fixed
STOP [63] Optimizer code Tests on five new tasks Some updates reduce performance
DGM [7] Agent tools and code Tests of agent updates and archive Control uses a different budget
ADAS [3] Agent designs and archive Tests of designs on new tasks Designer model and code stay fixed
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.