Submitted:
05 October 2026
Posted:
08 October 2026
You are already at the latest version
Abstract
Language-model systems now rewrite their own memory, code, and weights, and small language models now run on phones and embedded boards. This paper asks whether the two developments have met: can a system improve itself within the memory, energy, thermal, latency, and connectivity budgets of a device? We define edge self-improvement as a verified update loop under a timevarying device budget, and we classify 63 works along three axes: the improved component, where computation ran in the reported evaluation, and the verification tier. For the 14 works that carry the argument we read the full text. Three findings emerge. First, none of the works we examined that update prompts, code, or weights from self-generated supervision was evaluated on edge hardware, including those motivated by edge deployment; on-device evidence exists only for memory-level loops and for weight training with external supervision. Second, the only loop we found that ran on edge hardware with energy instrumentation relies on the weakest verification tier, and its decisive gains are in cost rather than accuracy. Third, benchmarks that score improvement do not score device cost, and benchmarks that measure devices do not score improvement. We trace seven limitations to their causes and propose energy-per-accepted-update metrics and a reporting checklist.
Keywords:
recursive self-improvement
; edge AI
; on-device learning
; small language models
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.