Submitted:
12 June 2026
Posted:
12 June 2026
You are already at the latest version
Abstract
Keywords:
1. Introduction
- We present the first survey of hint-based RL, consolidating works that introduce external textual signals into rollout-group construction to overcome zero-advantage failure.
- We propose a taxonomy of hints across two granularity levels, covering sample-level (trajectory-based and scaffold-based) and task-level (static and evolving experience bases).
- We further review domain-specific instantiations and discuss the boundaries, open problems, and future directions of hint-based RL.
2. Preliminary
3. Sample-Level Hints
3.1. Trajectory-Based Hints
3.1.1. Trajectory Injection
Reference Prefix
Self-Replay Prefix
Hybrid Prefix
3.1.2. Trajectory Continuation
Scheduled Prefix Decay
Adaptive Prefix Control
Intra-group Prefix Mixing
3.1.3. Trajectory Replacement
Unconditional Augmentation
Triggered Substitution
Reconstructive Repair
3.1.4. Trajectory Optimization
3.2. Scaffold-Based Hints
3.2.1. Scaffold Injection
Answer-Level Scaffold
Solution-Blueprint Scaffold
Knowledge-Level Scaffold
Format-Level Scaffold
3.2.2. Scaffold Replacement
Pedagogical Guidance
Critique-Driven Refinement
Action-Level Guidance
3.2.3. Scaffold Optimization
Auxiliary Supervision
Distribution Alignment
4. Task-Level Hints
4.1. Static Experience Base
Demonstration-Based Experience
Skill-based Experience
4.2. Evolving Experience Base
4.2.1. Accumulative Experience Base
Reflection Experience
Strategy Experience
4.2.2. Curated Experience Base
Skill Evolution
Experience Revision
4.2.3. Optimized Experience Base
Experience-Derived Utility
Trainable Experience Module
5. Domain-Specific Hints
6. Discussion
6.1. Scope and Boundaries
| Offline | Online | ||||||
|---|---|---|---|---|---|---|---|
| Human | Base | Teacher | Old | Current | Teacher | Total | |
| Inj. | 3 | 1 | 16 | 6 | 10 | 2 | 38 |
| Cont. | 1 | 1 | 8 | 1 | 1 | 0 | 12 |
| Repl. | 1 | 1 | 9 | 1 | 9 | 2 | 23 |
| Opt. | 3 | 2 | 3 | 3 | 3 | 6 | 20 |
| Total | 8 | 5 | 36 | 11 | 23 | 10 | 93 |
6.2. Cross-Level Analysis of Hints
6.3. Challenges and Future Directions
7. Conclusions
Limitations
Appendix A
Appendix A.1. Taxonomy and Representative Methods

Appendix A.2. Comparison Tables and Field Definitions
- Hint Content. This field names the explicit textual signal used by a method, such as a reference prefix, repaired trajectory, critique, workflow, skill, memory entry, executable feedback, or preference signal. It corresponds to the hint object that directly affects rollout generation, rollout repair, or the RL update.
- Source. This field records where the hint comes from. On-policy indicates that the hint is generated by the trained model itself. Off-policy indicates that the hint is produced with an external model, teacher model, external annotated trajectory, or external reference solution. Hybrid corresponds to Off&On-policy, where hint construction explicitly depends on both the trained model and an external model or teacher. External verifiers, rule-based scorers, and Math-Verify style correctness checkers are treated as reward or validation mechanisms. They do not change the hint source by themselves.
- Trigger. This field records when the hint intervenes in RL training. We use Always, Stage-wise Curriculum, and Specific Rules to distinguish persistent intervention, curriculum-dependent intervention, and rule-triggered intervention.
- Inference. This field records whether hints are still used after training. It supports the deployment discussion in Appendix A.4, where training-time hints may be removed, retained, or partially retained at inference.
- Objective. This field records the concrete optimization objective or training recipe reported by each method. It may include GRPO-family objectives, DAPO, REINFORCE, or method-specific variants, as well as auxiliary terms such as SFT, Kullback-Leibler divergence (KL), negative log-likelihood (NLL), Jensen-Shannon divergence (JSD), or self-distillation losses (SDL). Note that objectives labeled as GRPO in the table may also refer to GRPO variants adapted to specific methods.
- Retrieval / Use. This field appears in the task-level table and records how stored experience is selected or applied, such as random sampling, similarity retrieval, category retrieval, top-K retrieval, or population sampling.
- Experience Base Operation. This field in the task-level table records whether the experience repository is fixed or updated through operations including addition, update, deletion, merging, pruning, or reconstruction.
| Method | Hint Content | Source | Trigger | Inference | Objective |
|---|---|---|---|---|---|
| LLM | |||||
| StepHint [38] | step-level reference prefixes | Off-policy | Always | No | GRPO |
| ANCHOR [7] | verified reference trajectory anchor | Off-policy | Always | No | GRPO |
| LUFFY [15] | off-policy expert reasoning trace | Off-policy | Always | No | GRPO |
| RAVR [45] | ground-truth answer | Off-policy | Always | No | GRPO+KL |
| CoVRL [46] | ground-truth answer | Off-policy | Always | No | GRPO+KL+NLL |
| MeRF [55] | natural-language reward specification | Off-policy | Always | No | GRPO |
| RAVR [45] | answer-guided reference distribution | Off-policy | Always | No | GRPO+KL |
| CoVRL [46] | answer-guided mixed trajectory distribution | Off-policy | Always | No | GRPO+KL+NLL |
| POPE [8] | short useful reference prefix | Off-policy | Rules | No | GRPO |
| SEELE [23] | reference prefix with learned length control | Off-policy | Rules | No | GRPO+SFT |
| BREAD [25] | episode-level expert prefix | Off-policy | Rules | No | GRPO |
| G2RPO-A [33] | trajectory prefix adjusted by reward trend | Off-policy | Rules | No | GRPO |
| TRAPO [37] | expert prefix selected within the rollout group | Off-policy | Rules | No | GRPO+SFT |
| AMPO [39] | selected teacher reference trajectories | Off-policy | Rules | No | GRPO |
| HAPO [41] | reference trajectory replacing a low-reward rollout | Off-policy | Rules | No | GRPO+SFT |
| HDPO [47] | ground-truth answer | Off-policy | Rules | No | GRPO+JSD |
| Guide-GRPO [48] | teacher-generated high-level guidance | Off-policy | Rules | No | GRPO |
| KnowRL [18] | minimal sufficient knowledge points | Off-policy | Rules | Opt. | GRPO |
| HINT [57] | teacher-generated heuristic hints | Off-policy | Rules | No | GRPO |
| RGR-GRPO [61] | unmet rubric criteria | Off-policy | Rules | No | GRPO |
| HDPO [47] | answer-conditioned teacher distribution | Off-policy | Rules | No | GRPO+JSD |
| UFT [31] | reference prefix with scheduled decay | Off-policy | Curriculum | No | GRPO+KL |
| PieceHint [50] | high-value reasoning pieces | Off-policy | Curriculum | No | GRPO |
| RuscaRL [56] | checklist-style rubric criteria | Off-policy | Curriculum | No | GRPO |
| ThinkTuning [67] | teacher reflection tokens | Off-policy | Curriculum | No | GRPO |
| QuestA [21] | partial reference solution prefix | Off-policy | Rules+Curriculum | No | GRPO |
| LLM | |||||
| CCL [22] | threshold-calibrated reference prefix | Off-policy | Rules+Curriculum | No | GRPO |
| GHPO [24] | stagewise reference prefix | Off-policy | Rules+Curriculum | No | GRPO |
| Prefix-RFT [16] | offline demonstration prefix | Off-policy | Rules+Curriculum | No | Dr.GRPO |
| Scaf-GRPO [17] | hierarchical knowledge, planning, and solution hints | Off-policy | Rules+Curriculum | No | GRPO |
| RPO [28] | truncated historical response prefix | On-policy | Always | No | GRPO/DAPO |
| HiPO [26] | successful self-prefix from current batch | On-policy | Rules | No | GRPO |
| PROS [27] | uncertainty-truncated historical correct prefix | On-policy | Rules | No | GRPO |
| HiPO [26] | successful self-hinted trajectory group | On-policy | Rules | No | GRPO |
| LTE [60] | wrong-answer negative hints | On-policy | Rules | No | GRPO |
| Failure-Prefix [29] | rare failure prefix from saturated samples | On-policy | Rules+Curriculum | No | GRPO |
| ExPO [43] | answer-conditioned self-explanation trajectory | Hybrid | Always | No | GRPO+SFT |
| RLAD [51] | reasoning abstraction | Hybrid | Always | Yes | DAPO+RFT+SFT |
| ECHO [64] | diagnostic feedback from a trainable critic | Hybrid | Always | No | GRPO |
| RLTF [68] | natural-language critique feedback | Hybrid | Always | Opt. | GRPO+CE+AWR |
| CORE [30] | answer-stripped prefix from a successful peer policy | Hybrid | Rules | No | GRPO |
| SCOPE [42] | teacher-corrected suffix after a valid prefix | Hybrid | Rules | No | GRPO+NLL |
| A2D [52] | decomposer-generated sub-questions | Hybrid | Rules | No | GRPO+NLL |
| SAGEscaf [53] | self-generated plan or decomposition | Hybrid | Rules | No | GRPO |
| HiLL [59] | hinter-generated pedagogical hints | Hybrid | Rules | No | GRPO |
| GOLF [62] | group-level feedback from critiques | Hybrid | Rules | Opt. | GRPO |
| MEL [66] | verified meta-experience target | Hybrid | Rules | No | GRPO+NLL |
| EvoCoT [32] | CoT prefix constructed from final answer | Hybrid | Curriculum | No | GRPO |
| BHA [36] | distribution-aligned teacher prefix | Hybrid | Rules+Curriculum | No | DAPO |
| MENTOR [44] | mixed-policy rollout with expert token intervention | Hybrid | Rules+Curriculum | No | GRPO |
| NuRL [54] | self-generated abstract knowledge cue | Hybrid | Rules+Curriculum | No | GRPO |
| VLLM | |||||
| ADHint [34] | reference prefix selected by query difficulty | Off-policy | Rules | No | GRPO |
| Hint-GRPO [35] | stepwise reference prefix | Off-policy | Rules | No | GRPO |
| S-GRPO [40] | single verified reference trajectory | Off-policy | Rules | No | GRPO |
| AVATAR [49] | precomputed strategy hints | Off-policy | Rules | No | GRPO+SFT |
| KEPO [58] | answer-conditioned teacher reasoning hints | Off-policy | Rules | No | GRPO+KL+SFT |
| Agent | |||||
| InfoFlow [65] | pathfinding hint used as a search query | Off-policy | Rules | No | SFT+GRPO+KL |
| R3L [63] | reflection diagnosis with corrective guidance | Hybrid | Rules | No | GRPO |
| Method | Hint Content | Source | Retrieval / Use | Experience Base Operation | Trigger | Inference | Objective |
|---|---|---|---|---|---|---|---|
| LLM | |||||||
| CBRL [69] | Cross-sample solved demonstrations | Off-policy | Label sampling | Fixed | Curriculum | No | GRPO |
| TemplateRL [71] | Problem-solving templates | Hybrid | PCC matching | Fixed | Always | Yes | Dr.GRPO |
| ICPO [70] | Expert problem-solution demonstrations | Hybrid | Random sampling | Fixed | Rules | No | GRPO |
| DGO [85] | Correct and wrong reasoning experience | Hybrid | Experience retrieval | Add/Update/Delete | Curriculum | Yes | GRPO+SFT |
| Agent | |||||||
| SkillRL [20] | Task-level skills | Off-policy | Category retrieval | Add/Update | Always | Yes | GRPO+SFT |
| CRMWeaver [77] | Workflow guidelines | Off-policy | Top-1 similarity | Add/Update | Rules | Yes | DAPO+SFT |
| SKILL0 [19] | Procedural skills | Off-policy | Skill selection | Fixed | Rules+Curriculum | No | GRPO |
| INSPO [87] | Evolved instructions | Off-policy | Population sampling | Add/Delete/Reconstruct | Rules+Curriculum | Yes | GRPO |
| COS-PLAY [81] | Long-horizon game skills | On-policy | Skill retrieval | Add/Update/Merge/Delete | Always | Yes | GRPO+SFT |
| SAGEexp [82] | Executable function skills | On-policy | Skill retrieval | Add/Update | Always | Yes | GRPO+SFT |
| UMEM [93] | Editable memory-style experience entries | On-policy | Top-K retrieval | Add/Update | Rules | Yes | GRPO |
| EvolveR [72] | Guiding and cautionary reflections | On-policy | Explicit search | Add/Update/Merge/Prune | Rules | Yes | GRPO+SFT |
| RetroAgent [73] | Retrospective lessons | On-policy | SimUtil-UCB | Add/Update | Rules | Yes | GRPO |
| ERL [74] | Failure-correction reflections | On-policy | Threshold retry | Add | Rules | No | GRPO+SFT |
| SGE [76] | Successful and failed strategies | On-policy | FIFO sampling | Add/Delete | Rules | No | GRPO |
| SLEA-RL [91] | Step-level strategies and warnings | On-policy | Cluster retrieval | Add/Delete | Rules | Yes | GRPO |
| MAGE [75] | Interaction history and reflections | On-policy | Meta-context | Temporary aggregation | Curriculum | Yes | GiGPO |
| PEARL [89] | Time-management strategies | On-policy | StrategyHub | Add/Update/Delete | Rules+Curriculum | Yes | GRPO |
| Comp.RL [86] | Reusable trajectory-derived experience | Hybrid | Search-and-ask | Add/Update/Merge/Delete | Always | Yes | Split-GRPO |
| D2Skill [90] | Task and step skills | Hybrid | Utility comparison | Add/Update/Prune | Always | Yes | GRPO |
| Skill-SD [92] | Teacher-side task skills | Hybrid | Teacher scoring | Add/Update | Always | No | GRPO+SDL |
| IntPro [78] | User intent patterns with explanations | Hybrid | Tool retrieval | Add/Update | Rules | Yes | GRPO+SFT |
| MetaClaw [79] | Failure-driven skill instructions | Hybrid | Embedding search | Add | Rules | Yes | GRPO |
| BEPA [84] | Per-task successful trajectories | Hybrid | Cache lookup | Add/Update | Rules | No | GRPO |
| AgentEvolver [88] | Natural-language experience entries | Hybrid | Vector retrieval | Add/Update | Rules | Yes | GRPO |
| Trainable G.Mem. [94] | Graph-structured meta-cognition | Hybrid | Graph retrieval | Add/Merge/Update | Rules | Yes | REINFORCE |
| K2-Agent [83] | Mobile-control rules and demonstrations | Hybrid | Type retrieval | Add/Update/Merge | Curriculum | Yes | C-GRPO |
| ARISE [80] | Seed and generated skills | Hybrid | Manager selection | Add/Update/Delete | Rules+Curriculum | Yes | GRPO |
| Method | Hint Content | Source | Trigger | Inference | Objective |
|---|---|---|---|---|---|
| LLM | |||||
| TaoSR-AGRL [100] | Dimension-level business labels | Off-policy | Rules | No | GRPO |
| Kevin [95] | Execution feedback and prior kernel attempts | On-policy | Always | Yes | GRPO |
| SGS [96] | Synthetic theorem or subproblem | On-policy | Rules | No | REINFORCE |
| VeriRole [101] | Role evidence snippets | Hybrid | Always | Yes | GRPO |
| VLLM | |||||
| C2F-Thinker [97] | Ground-truth polarity label | Off-policy | Rules | No | GRPO+SFT |
| Agent | |||||
| WebGen-Agent [98] | Execution, visual, and GUI feedback | Off-policy | Rules | Yes | Step-GRPO+SFT |
| UI-S1 [99] | Expert GUI action patch | Off-policy | Rules | No | Semi-online GRPO+SFT |
| CRMWeaver [77] | Workflow guideline | Off-policy | Rules | Yes | DAPO+SFT |
| COS-PLAY [81] | Skill protocol | On-policy | Always | Yes | GRPO+SFT |
| PEARL [89] | Preference strategy | On-policy | Rules+Curriculum | Yes | GRPO |
Appendix A.3. Category-Level Comparison of Hints
trajectory-based,
scaffold-based. Task-level methods:
static experience base,
evolving experience base. Methods spanning two construction sources appear in both columns. “–” denotes unexplored combinations. † Method spans two construction sources and appears in both corresponding columns.
trajectory-based,
scaffold-based. Task-level methods:
static experience base,
evolving experience base. Methods spanning two construction sources appear in both columns. “–” denotes unexplored combinations. † Method spans two construction sources and appears in both corresponding columns.| Offline Construction | Online Construction | |||||
|---|---|---|---|---|---|---|
| Utilization | Human | Base Policy | Teacher | Old Policy | Current Policy | Teacher |
| Hint Injection |
GHPO MeRF CBRL† |
NuRL |
QuestA, POPE CCL, SEELE BREAD Guide-GRPO PieceHint, AVATAR KnowRL, RuscaRL RLAD† SKILL0, CBRL† CRMWeaver, IntPro SkillRL |
PROS, RPO Failure-Prefix† EvolveR RetroAgent, SGE |
CORE Failure-Prefix† SAGEscaf RLAD† MAGE, ARISE COS-PLAY SAGEexp Comp.RL, PEARL |
MetaClaw INSPO |
| Hint Continuation | K2-Agent† | EvoCoT† |
UFT, Prefix-RFT G2RPO-A, ADHint Hint-GRPO, BHA TRAPO, StepHint |
EvoCoT† | K2-Agent† | – |
| Hint Replacement | ICPO | BEPA† |
ANCHOR, LUFFY AMPO, S-GRPO HAPO HINT, Scaf-GRPO InfoFlow RGR-GRPO† |
BEPA† |
HiPO, SCOPE† LTE, R3L ECHO, HiLL GOLF† RGR-GRPO† ERL |
SCOPE† GOLF† |
| Hint Optimization |
RAVR, CoVRL HDPO |
ExPO A2D† |
MENTOR DGPO A2D† |
DGO† SLEA-RL Trainable G.Mem. |
MEL DGO† UMEM |
KEPO ThinkTuning RLTF AgentEvolver D2Skill, Skill-SD |
Appendix A.4. Hint Availability at Inference
References
- Wen, X.; Liu, Z.; Zheng, S.; Ye, S.; Wu, Z.; Wang, Y.; Xu, Z.; Liang, X.; Li, J.; Miao, Z.; et al. Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs. In Proceedings of the The Fourteenth International Conference on Learning Representations, 2026. [Google Scholar]
- Xu, F.; Hao, Q.; Zong, Z.; Wang, J.; Zhang, Y.; Wang, J.; Lan, X.; Gong, J.; Ouyang, T.; Meng, F.; et al. Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models. arXiv 2025, arXiv:cs. [Google Scholar]
- Shao, Z.; Wang, P.; Zhu, Q.; Xu, R.; Song, J.; Bi, X.; Zhang, H.; Zhang, M.; Li, Y.K.; Wu, Y.; et al. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models. arXiv 2024, arXiv:cs. [Google Scholar]
- Zhang, K.; Zuo, Y.; He, B.; Sun, Y.; Liu, R.; Jiang, C.; Fan, Y.; Tian, K.; Jia, G.; Li, P.; et al. A Survey of Reinforcement Learning for Large Reasoning Models. arXiv 2025, arXiv:cs. [Google Scholar]
- Wang, J.; Zhang, Z.; He, Y.; Zhang, Z.; Song, X.; Song, Y.; Shi, T.; Li, Y.; Xu, H.; Wu, K.; et al. Enhancing Code LLMs with Reinforcement Learning in Code Generation: A Survey. arXiv 2025, arXiv:cs. [Google Scholar]
- Zhang, G.; Geng, H.; Yu, X.; Yin, Z.; Zhang, Z.; Tan, Z.; Zhou, H.; Li, Z.Z.; Xue, X.; Li, Y.; et al. The Landscape of Agentic Reinforcement Learning for LLMs: A Survey. Trans. Mach. Learn. Res. Survey Certification. 2026. [Google Scholar] [CrossRef]
- Liu, J.; Dhole, K.; Wang, Y.; Wen, H.; Zhang, S.; Mao, H.; Li, G.; Varshney, N.; Liu, J.; Pan, X. Toward Honest Language Models for Deductive Reasoning. arXiv 2025, arXiv:cs. [Google Scholar]
- Qu, Y.; Setlur, A.; Smith, V.; Salakhutdinov, R.; Kumar, A. POPE: Learning to Reason on Hard Problems via Privileged On-Policy Exploration. arXiv 2026, arXiv:cs. [Google Scholar]
- Nie, S.; Ding, S.; Zhang, W.; Yu, L.; Yang, T.; Chen, Y.; Yin, W.; Sun, Y.; Wu, H.; Liu, T. ATTNPO: Attention-Guided Process Supervision for Efficient Reasoning. arXiv 2026, arXiv:cs. [Google Scholar]
- Ai, Z.; Shan, Z.; Ai, X.; Tang, J.; Hu, H.; Lu, P. SHAPE: Stage-aware Hierarchical Advantage via Potential Estimation for LLM Reasoning. arXiv 2026, arXiv:cs. [Google Scholar]
- Zheng, C.; Zhu, J.; Ou, Z.; Chen, Y.; Zhang, K.; Shan, R.; Zheng, Z.; Yang, M.; Lin, J.; Yu, Y.; et al. A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models. arXiv 2026, arXiv:cs. [Google Scholar]
- Yu, Q.; Zhang, Z.; Zhu, R.; Yuan, Y.; Zuo, X.; Yue, Y.; Dai, W.; Fan, T.; Liu, G.; Liu, L.; et al. DAPO: An Open-Source LLM Reinforcement Learning System at Scale. arXiv 2025, arXiv:cs. [Google Scholar]
- Li, Y.; Gu, Q.; Wen, Z.; Li, Z.; Xing, T.; Guo, S.; Zheng, T.; Zhou, X.; Qu, X.; Zhou, W.; et al. TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling. arXiv 2025, arXiv:cs. [Google Scholar]
- Yue, Y.; Chen, Z.; Lu, R.; Zhao, A.; Wang, Z.; Yue, Y.; Song, S.; Huang, G. Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model? arXiv 2025, arXiv:cs. [Google Scholar]
- Yan, J.; Li, Y.; Hu, Z.; Wang, Z.; Cui, G.; Qu, X.; Cheng, Y.; Zhang, Y. Learning to Reason under Off-Policy Guidance. In Proceedings of the The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2026. [Google Scholar]
- Huang, Z.; Cheng, T.; Qiu, Z.; Wang, Z.; Xu, Y.; Ponti, E.M.; Titov, I. Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling. arXiv 2025, arXiv:cs. [Google Scholar]
- Zhang, X.; Wu, S.; Zhu, Y.; Tan, H.; Yu, S.; He, Z.; Jia, J. Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM Reasoning. In Proceedings of the The Fourteenth International Conference on Learning Representations, 2026. [Google Scholar]
- Yu, L.; Yang, T.; Ding, S.; Jin, R.; Gu, N.; Hao, X.; Nie, S.; Xiong, D.; Yin, W.; Sun, Y.; et al. KnowRL: Boosting LLM Reasoning via Reinforcement Learning with Minimal-Sufficient Knowledge Guidance. arXiv 2026, arXiv:cs. [Google Scholar]
- Lu, Z.; Yao, Z.; Wu, J.; Han, C.; Gu, Q.; Cai, X.; Lu, W.; Xiao, J.; Zhuang, Y.; Shen, Y. SKILL0: In-Context Agentic Reinforcement Learning for Skill Internalization. arXiv 2026, arXiv:cs. [Google Scholar]
- Xia, P.; Chen, J.; Wang, H.; Liu, J.; Zeng, K.; Wang, Y.; Han, S.; Zhou, Y.; Zhao, X.; Chen, H.; et al. SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning. arXiv 2026, arXiv:cs. [Google Scholar]
- Li, J.; Lin, H.; Lu, H.; Wen, K.; Yang, Z.; Gao, J.; Wu, Y.; Zhang, J. QuestA: Expanding Reasoning Capacity in LLMs via Question Augmentation. In Proceedings of the The Fourteenth International Conference on Learning Representations, 2026. [Google Scholar]
- Wu, M.; Qian, Q.; Liu, W.; Wang, X.; Huang, Z.; Liang, D.; Miao, L.; Dou, S.; Lv, C.; Wang, Z.; et al. Progressive Mastery: Customized Curriculum Learning with Guided Prompting for Mathematical Reasoning. arXiv 2025, arXiv:cs. [Google Scholar]
- Li, Z.; Sun, Z.; Zhao, J.; Min, E.; Zeng, Y.; Wu, H.; Cai, H.; Wang, S.; Yin, D.; Chen, X.; et al. Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding. arXiv 2025, arXiv:cs. [Google Scholar]
- Liu, Z.; Gong, C.; Fu, X.; Liu, Y.; Chen, R.; Hu, S.; Zhang, S.; Liu, R.; Zhang, Q.; Tu, D. GHPO: Adaptive Guidance for Stable and Efficient LLM Reinforcement Learning. arXiv 2025, arXiv:cs. [Google Scholar]
- Zhang, X.; Huang, Z.; Li, Y.; Ni, C.; Chen, J.; Oymak, S. BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning. In Proceedings of the The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2026. [Google Scholar]
- Qiyuan, D.; Chen, K.; Zhang, M.; Xu, Z. HiPO: Self-Hint Policy Optimization for RLVR. In Proceedings of the The Fourteenth International Conference on Learning Representations, 2026. [Google Scholar]
- Huang, B.; Wan, X. PROS: Towards Compute-Efficient RLVR via Rollout Prefix Reuse. In Proceedings of the The Fourteenth International Conference on Learning Representations, 2026. [Google Scholar]
- Yi, H.; Wang, X.; zhang, Z.; Zong, T.; Wang, Y.; Xie, J.; Yu, T.; Jin, H.; Xu, K.; Chen, F.; et al. RPO:Reinforcement Fine-Tuning with Partial Reasoning Optimization. arXiv 2026, arXiv:cs. [Google Scholar]
- Kim, M.; Shrestha, S.; Ross, K. Training Reasoning Models on Saturated Problems via Failure-Prefix Conditioning. arXiv 2026, arXiv:cs. [Google Scholar]
- Mishra, K.; Aubakirov, M.; Takac, M.; Lukas, N.; Lahlou, S. CORE: Collaborative Reasoning via Cross Teaching. arXiv 2026, arXiv:cs. [Google Scholar]
- Liu, M.; Farina, G.; Ozdaglar, A.E. UFT: Unifying Supervised and Reinforcement Fine-Tuning. In Proceedings of the The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2026. [Google Scholar]
- Liu, H.; Li, J.; Dong, Y.; Yu, C.; Chen, T.; Wang, L.; Tao, Y.; Gu, B.; Li, G. EvoCoT: Overcoming the Exploration Bottleneck in Reinforcement Learning. arXiv 2026, arXiv:cs. [Google Scholar]
- Guo, Y.; Deng, W.; Cheng, Z.; Tang, X. G2RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance. arXiv 2025, arXiv:cs. [Google Scholar]
- Zhang, F.; Tan, Z.; Ma, X.; Dong, Z.; Leng, X.; Zhao, J.; Sun, X.; Yang, Y. ADHint: Adaptive Hints with Difficulty Priors for Reinforcement Learning. arXiv 2026, arXiv:cs. [Google Scholar]
- Huang, Q.; Dai, W.; Liu, J.; He, W.; Jiang, H.; Song, M.; Chen, J.; Yao, C.; Song, J. Boosting MLLM Reasoning with Text-Debiased Hint-GRPO. In Proceedings of the Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), October 2025; pp. 4848–4857. [Google Scholar]
- Xie, P.X.; Lin, C.Y.; Yang, C.L. Mitigating Distribution Sharpening in Math RLVR via Distribution-Aligned Hint Synthesis and Backward Hint Annealing. arXiv 2026, arXiv:cs. [Google Scholar]
- Su, M.; Guan, J.; Gu, Y.; Huang, M.; Wang, H. Trust-Region Adaptive Policy Optimization. In Proceedings of the The Fourteenth International Conference on Learning Representations, 2026. [Google Scholar]
- Zhang, K.; Lv, A.; Li, J.; Wang, Y.; Wang, F.; Hu, H.; Yan, R. StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason. arXiv 2025, arXiv:cs. [Google Scholar]
- Yuan, X.; Ding, Y.; Bin, Y.; Shao, W.; Cai, J.; Song, J.; Yang, Y.; Shen, H.T. More Than One Teacher: Adaptive Multi-Guidance Policy Optimization for Diverse Exploration. arXiv 2025, arXiv:cs. [Google Scholar]
- Yan, Y.; Tang, K.; Chen, S.; Xu, K.; Hu, D.; Yu, Q.; Hu, P. S-GRPO: Unified Post-Training for Large Vision-Language Models. arXiv 2026, arXiv:cs. [Google Scholar]
- Wu, Y.; Wang, K.; Chen, D.; Wei, K. Hindsight-Anchored Policy Optimization: Turning Failure into Feedback in Sparse Reward Settings. arXiv 2026, arXiv:cs. [Google Scholar]
- Ren, Y.; Zhang, H.; Xiao, L.; Zhang, X.; Huang, J.; Qiu, J.; Yu, B.; Chen, Q.; Liu, L. Recycling Failures: Salvaging Exploration in RLVR via Fine-Grained Off-Policy Guidance. arXiv 2026, arXiv:cs. [Google Scholar]
- Zhou, R.; Li, S.; Zhang, A.; Leqi, L. ExPO: Unlocking Hard Reasoning with Self-Explanation-Guided Reinforcement Learning. In Proceedings of the The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2026. [Google Scholar]
- Jiang, Z.; Han, J.; li, tingyun; Wang, X.; Jiang, S.; Dai, Z.; Shuguang, M.; Yu, F.; Liang, J.; Xiao, Y. Selective Expert Guidance for Effective and Diverse Exploration in Reinforcement Learning of LLMs. In Proceedings of the The Fourteenth International Conference on Learning Representations, 2026. [Google Scholar]
- Lin, T.; Zhao, X.; Zhang, X.; Long, R.; Xu, Y.; Jiang, Z.; Su, W.; Zheng, B. RAVR: Reference-Answer-guided Variational Reasoning for Large Language Models. arXiv 2025, arXiv:cs. [Google Scholar]
- Wen, X.; Lou, J.; Liu, Y.; Lin, H.; He, B.; Han, X.; Sun, L.; Lu, Y.; Zhang, D. Coupled Variational Reinforcement Learning for Language Model General Reasoning. arXiv 2026, arXiv:cs. [Google Scholar]
- Ding, K. HDPO: Hybrid Distillation Policy Optimization via Privileged Self-Distillation. arXiv 2026, arXiv:cs. [Google Scholar]
- Nath, V.; Lau, E.; Gunjal, A.; Sharma, M.; Baharte, N.; Hendryx, S. Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models. arXiv 2025, arXiv:cs. [Google Scholar]
- Kulkarni, Y.; Fazli, P. AVATAR: Reinforcement Learning to See, Hear, and Reason Over Video. arXiv 2026, arXiv:cs. [Google Scholar]
- Fang, Y.; Lin, J.; Fu, X.; Qin, C.; Shi, H. Placing Puzzle Pieces Where They Matter: A Question Augmentation Framework for Reinforcement Learning, 2026. arXiv arXiv:cs.
- Qu, Y.; Singh, A.; Lee, Y.; Setlur, A.; Salakhutdinov, R.; Finn, C.; Kumar, A. RLAD: Training LLMs to Discover Abstractions for Solving Reasoning Problems. arXiv 2025, arXiv:cs. [Google Scholar]
- Chen, Z.; Qin, X.; Zhao, W.X.; Wu, Y.; Wen, J.R. Adaptive Ability Decomposing for Unlocking Large Reasoning Model Effective Reinforcement Learning. arXiv 2026, arXiv:cs. [Google Scholar]
- Liao, B.; Dong, H.; Xu, X.; Monz, C.; Bian, J. Self-Hinting Language Models Enhance Reinforcement Learning. arXiv 2026, arXiv:cs. [Google Scholar]
- Chen, J.; PENG, X.; Choubey, P.K.; Huang, K.H.; Zhang, J.; Bansal, M.; Wu, C.S. Nudging the Boundaries of LLM Reasoning. In Proceedings of the The Fourteenth International Conference on Learning Representations, 2026. [Google Scholar]
- Zhang, J.; Ma, G.; Liu, S.; Wang, H.; Huang, J.; Lin, T.E.; Huang, F.; Li, Y.; Tao, D. A Simple "Motivation" Can Enhance Reinforcement Finetuning of Large Reasoning Models. arXiv 2026, arXiv:cs. [Google Scholar]
- Zhou, Y.; Li, S.; Liu, S.; Fang, W.; Zhang, K.; Zhao, J.; Yang, J.; Zhou, Y.; Lv, J.; Zheng, T.; et al. Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General LLM Reasoning. arXiv 2026, arXiv:cs. [Google Scholar]
- Wang, X.; Han, J.; Jiang, Z.; Li, T.; Liang, J.; Jiang, S.; Dai, Z.; Ma, S.; Yu, F.; Xiao, Y. HINT: Helping Ineffective Rollouts Navigate Towards Effectiveness. arXiv 2025, arXiv:cs. [Google Scholar]
- Yang, F.; Meng, R.; Qi, T.D.; Ezzati, A.; Wen, Y. KEPO: Knowledge-Enhanced Preference Optimization for Reinforcement Learning with Reasoning. arXiv 2026, arXiv:cs. [Google Scholar]
- Xia, Y.; Xu, C.; Yao, Z.; McAuley, J.; He, Y. Learning to Hint for Reinforcement Learning. arXiv 2026, arXiv:cs. [Google Scholar]
- Tang, C.; Huang, H.Y.; Liu, W.; Bai, C.; Yang, S.; Wu, Y. Do Not Step Into the Same River Twice: Learning to Reason from Trial and Error. arXiv 2026, arXiv:cs. [Google Scholar]
- Bi, B.; Liu, S.; Wang, Y.; Tong, S.; Mei, L.; Ge, Y.; Xu, Y.; Guo, J.; Cheng, X. Reward and Guidance through Rubrics: Promoting Exploration to Improve Multi-Domain Reasoning. arXiv 2025, arXiv:cs. [Google Scholar]
- Huang, L.; Cheng, X.; Zhao, C.; Shen, G.; Yang, J.; Feng, X.; Gu, Y.; Yu, X.; Qin, B. Bootstrapping Exploration with Group-Level Natural Language Feedback in Reinforcement Learning. arXiv 2026, arXiv:cs. [Google Scholar]
- Shi, W.; Chen, Y.; Li, Z.; Pan, X.; Sun, Y.; Xu, J.; Zhou, X.; Li, Y. R3L: Reflect-then-Retry Reinforcement Learning with Language-Guided Exploration, Pivotal Credit, and Positive Amplification. arXiv 2026, arXiv:cs. [Google Scholar]
- Li, Z.; Jiang, L.; Hu, Y.; Zeng, X.; Li, Y.; Zhang, X.; Chen, G.; Pan, Z.; Li, X.; Liu, Y. No More Stale Feedback: Co-Evolving Critics for Open-World Agent Learning. arXiv 2026, arXiv:cs. [Google Scholar]
- Luo, K.; Qian, H.; Liu, Z.; Xia, Z.; Xiao, S.; Bao, S.; Zhao, J.; Liu, K. InfoFlow: Reinforcing Search Agent Via Reward Density Optimization. arXiv 2025, arXiv:cs. [Google Scholar]
- Huang, S.; Li, Z.; Zeng, Y.; Ren, Q.; Fang, Z.; Su, Q.; Shi, K.; Chen, L.; Chen, Z.; Zhao, F. Internalizing Meta-Experience into Memory for Guided Reinforcement Learning in Large Language Models. arXiv 2026, arXiv:cs. [Google Scholar]
- Rrv, A.; Dineen, J.; Handa, D.; Uddin, M.N.; Parmar, M.; Baral, C.; Zhou, B. ThinkTuning: Instilling Cognitive Reflections without Distillation. In Proceedings of the Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing; Suzhou, China, Christodoulopoulos, C., Chakraborty, T., Rose, C., Peng, V., Eds.; 2025; pp. 31248–31262. [Google Scholar] [CrossRef]
- Song, Y.; Chen, L.; Tajwar, F.; Munos, R.; Pathak, D.; Bagnell, J.A.; Singh, A.; Zanette, A. Expanding the Capabilities of Reinforcement Learning via Text Feedback. arXiv 2026, arXiv:cs. [Google Scholar]
- Agashe, S.; Srinivasa, J.; Liu, G.; Kompella, R.; Wang, X.E. Context Bootstrapped Reinforcement Learning. arXiv 2026, arXiv:cs. [Google Scholar]
- Huang, H.Y.; Tang, C.; Liu, W.; Bai, C.; Yang, S.; Wu, Y. Think Outside the Policy: In-Context Steered Policy Optimization. arXiv 2026, arXiv:cs. [Google Scholar]
- Wu, J.; Liao, C.; Feng, M.; Zhang, S.; Wen, Z.; Luo, H.; Yang, L.; Xu, H.; Tao, J. TemplateRL: Structured Template-Guided Reinforcement Learning for LLM Reasoning. arXiv 2025, arXiv:cs. [Google Scholar]
- Wu, R.; Wang, X.; Mei, J.; Cai, P.; Fu, D.; Yang, C.; Wen, L.; Yang, X.; Shen, Y.; Wang, Y.; et al. EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle. arXiv 2025, arXiv:cs. [Google Scholar]
- Zhang, X.; Liu, Z.; Zhang, Y.; Hu, X.; Shao, W. RetroAgent: From Solving to Evolving via Retrospective Dual Intrinsic Feedback. arXiv 2026, arXiv:cs. [Google Scholar]
- Shi, T.; Chen, S.; Jiang, B.; Song, L.; Yang, L.; Zhao, J. Experiential Reinforcement Learning. arXiv 2026, arXiv:cs. [Google Scholar]
- Yang, lu; Xu, Z.; Xie, M.; Gao, J.; shok, zhao; Wang, Y.; Wu, Y. MAGE: Meta-Reinforcement Learning for Language Agents toward Strategic Exploration and Exploitation. In Proceedings of the ICLR 2026 Workshop on Lifelong Agents: Learning, Aligning, Evolving, 2026. [Google Scholar]
- Szot, A.; Kirchhof, M.; Attia, O.; Toshev, A. Expanding LLM Agent Boundaries with Strategy-Guided Exploration, 2026. arXiv arXiv:cs.
- Lai, Y.; Yang, Y.; Wu, J.; Mo, F.; Wang, Z.; Liang, T.; Lin, J.; Yang, K. CRMWeaver: Building Powerful Business Agent via Agentic RL and Shared Memories. arXiv 2025, arXiv:cs. [Google Scholar]
- Liu, G.; Wu, M.; Zhang, P.; Zhang, Y.; Shu, Y.; Huang, X.; Tu, K.; Gu, N.; Zhang, L.; Wang, Q.; et al. IntPro: A Proxy Agent for Context-Aware Intent Understanding via Retrieval-conditioned Inference. arXiv 2026, arXiv:cs. [Google Scholar]
- Xia, P.; Chen, J.; Yang, X.; Tu, H.; Liu, J.; Xiong, K.; Han, S.; Qiu, S.; Ji, H.; Zhou, Y.; et al. MetaClaw: Just Talk – An Agent That Meta-Learns and Evolves in the Wild, 2026. arXiv arXiv:cs.
- Li, Y.; Miao, R.; Qi, Z.; Lan, T. ARISE: Agent Reasoning with Intrinsic Skill Evolution in Hierarchical Reinforcement Learning. arXiv 2026, arXiv:cs. [Google Scholar]
- Wu, X.; Li, Z.; Shi, G.; Duffy, A.; Marques, T.; Olson, M.L.; Zhou, T.; Manocha, D. Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks. arXiv 2026, arXiv:cs. [Google Scholar]
- Wang, J.; Yan, Q.; Wang, Y.; Tian, Y.; Mishra, S.S.; Xu, Z.; Gandhi, M.; Xu, P.; Cheong, L.L. Reinforcement Learning for Self-Improving Agent with Skill Library. arXiv 2026, arXiv:cs. [Google Scholar]
- Wu, Z.; Mo, D.; Lu, H.; Xing, J.; Liu, J.; Jing, Y.; Li, K.; Shao, K.; HAO, J.; Shi, Y. K²-Agent: Co-Evolving Know-What and Know-How for Hierarchical Mobile Device Control. In Proceedings of the The Fourteenth International Conference on Learning Representations, 2026. [Google Scholar]
- Wang, Z.; Zhang, Z.; Zhang, X.; Qian, Z.; Lu, Y. From Off-Policy to On-Policy: Enhancing GUI Agents via Bi-level Expert-to-Policy Assimilation. arXiv 2026, arXiv:cs. [Google Scholar]
- Bai, F.; Chen, Z.; Hao, C.; Yang, M.; Tao, R.; Dai, B.; Zhao, W.X.; Yang, J.; Xu, H. Towards Effective Experiential Learning: Dual Guidance for Utilization and Internalization. arXiv 2026, arXiv:cs. [Google Scholar]
- Muhtar, D.; Liu, J.; Gao, W.; Wang, W.; Xiong, S.; Huang, J.; Yang, S.; Su, W.; Wang, J.; Pan, L.; et al. Complementary Reinforcement Learning. arXiv 2026, arXiv:cs. [Google Scholar]
- Zhou, H.; Wan, X.; Vulić, I.; Korhonen, A. Agentic Policy Optimization via Instruction-Policy Co-Evolution. arXiv 2026, arXiv:cs. [Google Scholar]
- Zhai, Y.; Tao, S.; Chen, C.; Zou, A.; Chen, Z.; Fu, Q.; Mai, S.; Yu, L.; Deng, J.; Cao, Z.; et al. AgentEvolver: Towards Efficient Self-Evolving Agent System. arXiv 2025, arXiv:cs. [Google Scholar]
- Li, B.; Kim, J.; Qian, C.; Chen, X.; Anzenberg, E.; Kundapur, N.; Ji, H. PEARL: Self-Evolving Assistant for Time Management with Reinforcement Learning. arXiv 2026, arXiv:cs. [Google Scholar]
- Tu, S.; Xu, C.; Zhang, Q.; Zhang, Y.; Lan, X.; Li, L.; Zhao, D. Dynamic Dual-Granularity Skill Bank for Agentic RL. arXiv 2026, arXiv:cs. [Google Scholar]
- Wang, P.Z.; Jiang, S. SLEA-RL: Step-Level Experience Augmented Reinforcement Learning for Multi-Turn Agentic Training. arXiv 2026, arXiv:cs. [Google Scholar]
- Wang, H.; Wang, G.; Xiao, H.; Zhou, Y.; Pan, Y.; Wang, J.; Xu, K.; Wen, Y.; Ruan, X.; Chen, X.; et al. Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents. arXiv 2026, arXiv:cs. [Google Scholar]
- Ye, Y.; Jiang, H.; Jiang, F.; Lan, T.; Du, Y.; Fu, B.; Shi, X.; Jia, Q.; Wang, L.; Luo, W. UMEM: Unified Memory Extraction and Management Framework for Generalizable Memory. arXiv 2026, arXiv:cs. [Google Scholar]
- Xia, S.; Xu, Z.; Chai, J.; Fan, W.; Song, Y.; Wang, X.; Yin, G.; Lin, W.; Zhang, H.; Wang, J. From Experience to Strategy: Empowering LLM Agents with Trainable Graph Memory. arXiv 2025, arXiv:cs. [Google Scholar]
- Baronio, C.; Marsella, P.; Pan, B.; Guo, S.; Alberti, S. Kevin: Multi-Turn RL for Generating CUDA Kernels. arXiv 2025, arXiv:cs. [Google Scholar]
- Bailey, L.; Wen, K.; Dong, K.; Hashimoto, T.; Ma, T. Scaling Self-Play with Self-Guidance. arXiv 2026, arXiv:cs. [Google Scholar]
- Luo, M.; Yang, Z.; Long, J.; Sun, J.; Liu, Y.; Mai, S. C2F-Thinker: Coarse-to-Fine Reasoning with Hint-Guided Reinforcement Learning for Multimodal Sentiment Analysis. arXiv 2026, arXiv:cs. [Google Scholar]
- Lu, Z.; Ren, H.; Yang, Y.; Wang, K.; Zong, Z.; Pan, J.; Zhan, M.; Li, H. WebGen-Agent: Enhancing Interactive Website Generation with Multi-Level Feedback and Step-Level Reinforcement Learning. arXiv 2025, arXiv:cs. [Google Scholar]
- Lu, Z.; Ye, J.; Tang, F.; Shen, Y.; Xu, H.; Zheng, Z.; Lu, W.; Yan, M.; Huang, F.; Xiao, J.; et al. UI-S1: Advancing GUI Automation via Semi-online Reinforcement Learning. arXiv 2025, arXiv:cs. [Google Scholar]
- Yang, J.; Jin, Y.; Jiao, P.; Dong, C.; Huang, Z.; Yao, S.; Zhou, X.; Ou, D.; Tang, H. TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance. In Proceedings of the Proceedings of the ACM Web Conference 2026, New York, NY, USA, 2026; WWW ’26, pp. 7955–7966. [Google Scholar] [CrossRef]
- Wang, Z.; Sun, K.; Wu, B.; qun yu; Li, Y.; Chen, X.; Wang, B. VeriRole: Verifiable Role-Awareness through Hint-Guided Reinforcement Learning. In Proceedings of the The Fourteenth International Conference on Learning Representations, 2026. [Google Scholar]
- Ji, Y.; Ma, Z.; Wang, Y.; Chen, G.; Chu, X.; Wu, L. Tree Search for LLM Agent Reinforcement Learning. In Proceedings of the The Fourteenth International Conference on Learning Representations, 2026. [Google Scholar]
- Dong, G.; Mao, H.; Ma, K.; Bao, L.; Chen, Y.; Wang, Z.; Chen, Z.; Du, J.; Wang, H.; Zhang, F.; et al. Agentic Reinforced Policy Optimization. In Proceedings of the The Fourteenth International Conference on Learning Representations, 2026. [Google Scholar]
- Guo, Y.; Hu, T.; Sun, Z.; Lin, Y. Less Noise, More Voice: Reinforcement Learning for Reasoning via Instruction Purification, 2026. arXiv arXiv:cs.
- Zhang, K.; Yao, Q.; Liu, S.; Zhang, W.; Cen, M.; Zhou, Y.; Fang, W.; Zhao, Y.; Lai, B.; Song, M. Replay Failures as Successes: Sample-Efficient Reinforcement Learning for Instruction Following. arXiv 2025, arXiv:cs. [Google Scholar]
- Chen, J.C.Y.; Prasad, A.; Khan, Z.; Singh, J.; Tian, R.; Stengel-Eskin, E.; Bansal, M. Cog-DRIFT: Exploration on Adaptively Reformulated Instances Enables Learning from Hard Reasoning Problems. arXiv 2026, arXiv:cs. [Google Scholar]
- Nourzad, N.; Joe-Wong, C. MIRA: Memory-Integrated Reinforcement Learning Agent with Limited LLM Guidance. In Proceedings of the The Fourteenth International Conference on Learning Representations, 2026. [Google Scholar]
- Xu, Y.; Potje, G.; Shandilya, S.; Yuan, T.; de Oliveira Nunes, L.; Agarwal, R.; Asgari, S.; Atkinson, A.; Kıcıman, E.; Lu, S.; et al. SibylSense: Adaptive Rubric Learning via Memory Tuning and Adversarial Probing. arXiv 2026, arXiv:cs. [Google Scholar]
- Sheng, L.; Ma, W.; Hong, R.; Wang, X.; Zhang, A.; Chua, T.S. Reinforcing Chain-of-Thought Reasoning with Self-Evolving Rubrics. arXiv 2026, arXiv:cs. [Google Scholar]
- Gu, N.; Yang, C.; Si, Q.; Qin, C.; Yao, D.; Fu, P.; Lin, Z.; Wang, W.; Duan, N.; Wang, J. Co-Evolving Policy Distillation. arXiv 2026, arXiv:cs. [Google Scholar]
- Yang, C.; Qin, C.; Si, Q.; Chen, M.; Gu, N.; Yao, D.; Lin, Z.; Wang, W.; Wang, J.; Duan, N. Self-Distilled RLVR, 2026. arXiv arXiv:cs.
- Song, M.; Zheng, M. A Survey of On-Policy Distillation for Large Language Models. arXiv 2026, arXiv:cs. [Google Scholar]
| 1 | We use the subscripts scaf and exp to distinguish SAGE variants that share the same name. |


Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).