Submitted:
11 July 2026
Posted:
14 July 2026
You are already at the latest version
Abstract
Keywords:
1. Introduction
1.1. Why Token Consumption Matters for Agents
- 1.
- Cost. Because the full transcript is re-sent each turn, cumulative input spend can grow quadratically in the number of turns absent caching (a sum of monotonically growing per-turn contexts; Figure 1b); prompt caching softens the constant but not the trend. Unaudited community telemetry for OpenClaw attributes an estimated 40–50% of typical token usage to context accumulation alone, and a single user reported that session context occupied 56–58% of a 400K-token window—roughly 230,000 tokens re-sent with every message LaoZhang AI (2026).
- 2.
- 3.
- 4.
- Accuracy degradation (“context rot”). Performance falls as inputs grow. Liu et al. (2024a) show a U-shaped accuracy curve in which performance “degrades significantly” when relevant information sits in the middle of long inputs (multi-document QA fell from ∼75% when the gold document was first to ∼55% when it was tenth of twenty, across GPT-3.5-Turbo, GPT-4, Claude-1.3, MPT-30B, and LongChat-13B). Chroma’s 2025 Context Rot report extends this to 18 frontier models (including GPT-4.1, Claude 4, Gemini 2.5, and Qwen3), finding that every model degrades as input length grows, even at 50K tokens within a 1M-token window Hong et al. (2025). Controlled probes sharpen the diagnosis: on NoLiMa, which removes literal needle–question lexical overlap, 11 of 13 models advertising ≥128K context fall below half their short-context accuracy at just 32K tokens (GPT-4o drops from 99.3% to 69.7%) Modarressi et al. (2025); RULER finds only about half of 17 evaluated long-context LMs sustain acceptable accuracy at 32K despite claiming 32K+ windows Hsieh et al. (2024); and BABILong reports that models effectively use only 10–20% of their context on facts scattered across long inputs Kuratov et al. (2024). The effective window is thus far smaller than the advertised one—keeping context small often preserves more usable signal than filling it, though the degradation is highly task- and model-dependent (some models and tasks stay near-flat), a heterogeneity we revisit when discussing evaluation in Section 7.
1.2. Scope and Named Systems
- 1.
- A unified agent-context lifecycle model—admission, placement, compaction, recovery, reuse, and governance—that frames context compression as full-session management rather than one-shot prompt shortening (Section 3); governance is included as a normative stage—motivated by compression–security interactions but not yet instantiated in any surveyed production system.
- 2.
- 3.
- A production-agent comparison of how deployed general-purpose agents (OpenClaw, Hermes, Manus, Claude Code, OpenHands) instantiate these choices in practice (Section 5).
- 4.
- A synthesis of the 2025–2026 shift from compression-as-scaffold to compression-as-policy, where the compaction step becomes a reinforcement-learned action (Section 6).
- 5.
- An evaluation agenda that jointly measures task success, peak and total tokens, latency, cache hit rate, and recovery/provenance, rather than compression ratio alone (Section 7).
| Survey family | Primary focus |
|---|---|
| Prompt compression Li et al. (2025b) | Prompt representation and hard/soft compression mechanisms for a single call. |
| KV-cache / serving Li et al. (2025a) | Serving-time memory, throughput, eviction, and quantization of the attention cache. |
| Agent memory Zhang et al. (2024d); Du (2026) | Persistent memory mechanisms as a write–manage–read loop. |
| Context engineering Mei et al. (2025) | The broad acquisition–processing–management pipeline, with compression as one component. |
| This survey | Full-session token economics and compression decisions in tool-using agents. |
2. What Makes Agent Compression Different
3. A Lifecycle Taxonomy for Agent Context Compression
- Training-free vs. model-modifying. Token-space methods leave the model untouched and run on any closed API; embedding- and KV-space methods require model access or self-hosting. This axis directly shapes the recommendations of Section 8.
- External scaffold vs. learned policy. Classically compression is a wrapper around an unmodified agent; an emerging line instead trains the agent to compress its own context (Section 6). Learned-policy compression cuts across families and is the field’s most active frontier.
4. Method Families as Implementation Choices
4.1. Natural-Language Compaction
4.1.1. Summarization-Based Compression
4.1.2. Token Pruning/Selective Context Dropping
- LLMLingua Jiang et al. (2023a): a coarse-to-fine pipeline (budget controller + iterative token-level perplexity pruning + distribution alignment) using a small LM (e.g. GPT-2/LLaMA-7B). It reports “up to 20x compression with only a 1.5 point performance drop” across GSM8K, BBH, ShareGPT, and Arxiv-March23 (LLaMA-7B compressor, GPT-3.5-Turbo target).
- LongLLMLingua Jiang et al. (2024b): adds query-aware coarse-to-fine compression and document reordering for long-context RAG, boosting performance by up to 21.4% with ∼4× fewer tokens on NaturalQuestions, achieving a 94.0% cost reduction on LooGLE, and accelerating end-to-end latency 1.4–2.6× at 2–6× compression of ∼10K-token prompts.
- LLMLingua-2 Pan et al. (2024): reframes compression as token classification with a BERT-sized bidirectional encoder distilled from GPT-4, giving task-agnostic compression that is 3–6× faster than LLMLingua and reduces end-to-end latency by up to 2.9× at 2–5× compression.
4.2. Soft and Latent Compression
4.3. KV-Cache Compression and Serving-Time Controls
4.4. Retrieval and External-Memory Compression
4.4.1. Context Offloading (Filesystem as External Memory)
4.4.2. Hierarchical/Virtual Memory (MemGPT, Letta)
4.4.3. Retrieval-Augmented Generation (RAG) for Context Reduction
4.5. Sub-Agent Isolation
5. Production Patterns in General-Purpose Agents
5.1. Trajectory Compaction in Deployed Agents
- 1.
- Agent ContextCompressor (primary), fires at 50% of context using accurate API-reported token counts.
- 2.
- Gateway session hygiene (safety net), fires at 85% using a rough character estimate, to catch sessions that ballooned between turns (e.g. overnight on Telegram/Discord).

5.2. Budget Guards and Mundane Controls
5.3. Recoverable Context by Design
5.4. Prompt-Cache-Aware Layouts
5.5. Production Agents at a Glance
6. Learned, Agent-Native Compression: The Frontier
6.1. Compression as Action
6.2. Memory as Action
6.3. Credit Assignment for Compression Decisions
6.4. Budget-Conditioned Managers
6.5. Why the Frontier is Agent-Specific
6.6. Comparison of Methods
7. Evaluation: From Compression Ratio to Agent Utility
8. Design Guidelines and Recommendations
- 1.
- First, cheap structural fixes (highest ROI). Route low-value turns (heartbeats, status checks) to cheap models; reset sessions on task boundaries; truncate large tool outputs; keep always-injected bootstrap/system files short. Community write-ups report (unaudited) 80–95% cost reductions from these alone LaoZhang AI (2026); Hayes (2026); we present this ordering as a practitioner heuristic pending controlled measurement, not an established result. Threshold to revisit: if context still exceeds ∼50% of the window mid-task.
- 2.
- Then, enable trajectory compaction with caching. Adopt an OpenHands-/Hermes-style condenser that fires at 50–80% utilization, protects the system prompt and a recent tail, summarizes the middle with a structured template, and updates (not re-creates) prior summaries. Align pruning TTL with prompt-cache TTL to avoid cache-miss spikes Nous Research (2026a); All Hands AI (2025a); Wu (2026). Benchmark to track: task success rate and peak tokens jointly; back off compression aggressiveness if success drops >2–3 points.
- 3.
- Add recoverable offloading for large artifacts. Drop verbose content (web pages, file dumps) from context while keeping URLs/paths so it can be re-fetched; let the agent maintain a todo.md/scratchpad Ji (2025); Bhavsar (2025). Prefer this over destructive truncation whenever the artifact may be needed later.
- 4.
- Isolate exploration in sub-agents for research/codebase-exploration tasks, returning only 1–2K-token summaries to the orchestrator—accepting the (community-estimated) ∼3.5× coordination overhead only when the task genuinely parallelizes Hadfield et al. (2025); Hayes (2026).
- 5.
- 6.
- Reserve soft/embedding and KV-cache compression for self-hosted open-weight deployments. Gisting/ICAE/500x and H2O/SnapKV give the largest ratios but need model access; they are not actionable for closed-API agents, which should instead rely on provider prompt caching and the above Mu et al. (2023); Ge et al. (2024); Zhang et al. (2023).
9. Open Challenges
Reproducibility and Provenance Statement
Conflicts of Interest
References
- All Hands AI. OpenHands context condensation for more efficient AI agents. All Hands AI (OpenHands) Blog, April 2025a. URL https://www.openhands.dev/blog/openhands-context-condensensation-for-more-efficient-ai-agents. Published 9 April 2025; accessed June 2026.
- All Hands AI. Context condenser (LLMSummarizingCondenser / RollingCondenser). OpenHands Documentation, 2025b. URL https://docs.openhands.dev/sdk/guides/context-condenser. Accessed 2026-06-24.
- All Hands AI. Sub-agent delegation. OpenHands SDK Documentation, 2025c. URL https://docs.openhands.dev/sdk/guides/agent-delegation. Accessed 2026-06-24. Sub-agents as tools with isolated context via DelegateTool / TaskToolSet.
- Anthropic. Prompt caching. Claude API Documentation, 2024. URL https://platform.claude.com/docs/en/build-with-claude/prompt-caching. Cache writes 1.25 × (5-min) / 2.0 × (1-hour) base input; cache reads 0.1 × (90% discount). Accessed 2026-06-24.
- Anthropic. Introducing advanced tool use on the Claude developer platform. Anthropic Engineering Blog, November 2025a. URL https://www.anthropic.com/engineering/advanced-tool-use. Published 24 November 2025. Authored by Bin Wu and the Claude Developer Platform team. On-page title rendered in sentence case: “Introducing advanced tool use on the Claude Developer Platform”; accessed June 2026.
- Anthropic. Code execution with MCP: Building more efficient agents. Anthropic Engineering Blog, November 2025b. URL https://www.anthropic.com/engineering/code-execution-with-mcp. Published 4 November 2025. Authored by Adam Jones and Conor Kelly. On-page title rendered in sentence case: “Code execution with MCP: Building more efficient agents”; accessed June 2026.
- Anthropic. Effective context engineering for AI agents. Anthropic Engineering Blog, September 2025c. URL https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents. Published 29 September 2025; accessed June 2026.
- Anthropic. Scaling managed agents: Decoupling the brain from the hands. Anthropic Engineering Blog, April 2026a. URL https://www.anthropic.com/engineering/managed-agents. Published 8 April 2026; accessed June 2026.
- Anthropic. Claude opus 4.6. Anthropic News, February 2026b. URL https://www.anthropic.com/news/claude-opus-4-6. Published 5 February 2026. Context compaction triggered at a configurable threshold (50K-token example in benchmark methodology); accessed June 2026.
- AtlasPA. OpenClaw context optimizer. GitHub repository, 2026. URL https://github.com/AtlasPA/openclaw-context-optimizer. Cited for an advertised 40–60% token reduction via intelligent context compression; accessed June 2026.
- Yushi Bai, Shangqing Tu, Jiajie Zhang, Hao Peng, Xiaozhi Wang, Xin Lv, Shulin Cao, Jiazheng Xu, Lei Hou, Yuxiao Dong, Jie Tang, and Juanzi Li. LongBench v2: Towards deeper understanding and reasoning on realistic long-context multitasks, 2024. URL https://arxiv.org/abs/2412.15204.
- Nick Baumann. How to think about context engineering in Cline. Cline Blog, August 2025. URL https://cline.bot/blog/how-to-think-about-context-engineering-in-cline. Published 19 August 2025. On-page title: “How to Think about Context Engineering in Cline”; accessed June 2026.
- Gabriele Berton, Jayakrishnan Unnikrishnan, Son Tran, and Mubarak Shah. Compllm: Compression for long context q&a, 2025. URL https://arxiv.org/abs/2509.19228.
- Pratik Bhavsar. Deep dive into context engineering for agents. Galileo AI Blog, https://galileo.ai/blog/context-engineering-for-agents, sep 2025. Accessed 2026-06-24.
- Zefan Cai, Yichi Zhang, Bofei Gao, Yuliang Liu, Yucheng Li, Tianyu Liu, Keming Lu, Wayne Xiong, Yue Dong, Junjie Hu, and Wen Xiao. PyramidKV: Dynamic KV cache compression based on pyramidal information funneling. 2024. URL https://arxiv.org/abs/2406.02069.
- centminmod. explain-openclaw: Multi-AI documentation for OpenClaw (architecture, token/context optimization, security). GitHub repository, 2026. URL https://github.com/centminmod/explain-openclaw. Cited for OpenClaw’s tokenizer-free ∼4-chars-per-token estimation, tool-result truncation, contextPruning, and bootstrap-injection size caps (see 06-optimizations/cost-token-optimization.md); accessed June 2026.
- Alex Chen, Renato Geh, Aditya Grover, Guy Van den Broeck, and Daniel Israel. The pitfalls of KV cache compression, 2025. URL https://arxiv.org/abs/2510.00231.
- Qianben Chen, Tianrui Qin, King Zhu, Qiexiang Wang, Chengjun Yu, Shu Xu, Jiaqi Wu, Jiayu Zhang, Xinpeng Liu, Xin Gui, Jingyi Cao, Piaohong Wang, Dingfeng Shi, He Zhu, Tiannan Wang, Yuqing Wang, Maojia Song, Tianyu Zheng, Ge Zhang, Jian Yang, Jiaheng Liu, Minghao Liu, Yuchen Eleanor Jiang, and Wangchunshu Zhou. Search more, think less: Rethinking long-horizon agentic search for efficiency and generalization, 2026. URL https://arxiv.org/abs/2602.22675.
- Xin Cheng, Xun Wang, Xingxing Zhang, Tao Ge, Si-Qing Chen, Furu Wei, Huishuai Zhang, and Dongyan Zhao. xrag: Extreme context compression for retrieval-augmented generation with one token. In Advances in Neural Information Processing Systems (NeurIPS), 2024. URL https://arxiv.org/abs/2405.13792.
- Alexis Chevalier, Alexander Wettig, Anirudh Ajith, and Danqi Chen. Adapting language models to compress contexts. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 3829–3846. Association for Computational Linguistics, 2023. URL https://aclanthology.org/2023.emnlp-main.232/.
- Prateek Chhikara, Dev Khant, Saket Aryan, Taranjeet Singh, and Deshraj Yadav. Mem0: Building production-ready AI agents with scalable long-term memory. arXiv preprint arXiv:2504.19413, 2025. URL https://arxiv.org/abs/2504.19413.
- Nadezhda Chirkova, Thibault Formal, Vassilina Nikoulina, and Stéphane Clinchant. Provence: Efficient and robust context pruning for retrieval-augmented generation. In International Conference on Learning Representations (ICLR), 2025. URL https://arxiv.org/abs/2501.16214.
- DeepSeek-AI. DeepSeek-V2: A strong, economical, and efficient mixture-of-experts language model, 2024. URL https://arxiv.org/abs/2405.04434. Introduces Multi-head Latent Attention (MLA): low-rank latent compression of the KV cache.
- Hongchao Du, Shangyu Wu, Qiao Li, Riwei Pan, Jinheng Li, Youcheng Sun, and Chun Jason Xue. ClawMobile: Rethinking smartphone-native agentic systems, 2026. URL https://arxiv.org/abs/2602.22942. Cited for ClawMobile’s memory mechanism, built on OpenClaw.
- Pengfei Du. Memory for autonomous llm agents: Mechanisms, evaluation, and emerging frontiers, 2026. URL https://arxiv.org/abs/2603.07670.
- Yuan Feng, Junlin Lv, Yukun Cao, Xike Xie, and S. Kevin Zhou. Ada-KV: Optimizing KV cache eviction by adaptive budget allocation for efficient LLM inference. In Advances in Neural Information Processing Systems (NeurIPS), 2025. URL https://arxiv.org/abs/2407.11550.
- Tao Ge, Jing Hu, Lei Wang, Xun Wang, Si-Qing Chen, and Furu Wei. In-context autoencoder for context compression in a large language model. In International Conference on Learning Representations (ICLR), 2024. URL https://arxiv.org/abs/2307.06945.
- Jeremy Hadfield, Barry Zhang, Kenneth Lien, Florian Scholz, Jeremy Fox, and Daniel Ford. How we built our multi-agent research system. Anthropic Engineering Blog, June 2025. URL https://www.anthropic.com/engineering/multi-agent-research-system. Published 13 June 2025; accessed June 2026.
- Ellie Grace Hayes. How to reduce your OpenClaw API costs by 90% or more. LumaDock (blog), 2026. URL https://lumadock.com/tutorials/openclaw-cost-optimization-budgeting. Community write-up cited for ∼3.5 × multi-agent coordination token overhead and heartbeat cost figures; accessed June 2026.
- Kelly Hong, Anton Troynikov, and Jeff Huber. Context rot: How increasing input tokens impacts LLM performance. Technical report, Chroma, July 2025. URL https://research.trychroma.com/context-rot. Accessed June 2026.
- Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, Yang Zhang, and Boris Ginsburg. RULER: What’s the real context size of your long-context language models? In Conference on Language Modeling (COLM), 2024. URL https://arxiv.org/abs/2404.06654.
- Taeho Hwang, Sukmin Cho, Soyeong Jeong, Hoyun Song, SeungYoon Han, and Jong C. Park. EXIT: Context-aware extractive compression for enhancing retrieval-augmented generation. In Findings of the Association for Computational Linguistics: ACL 2025. Association for Computational Linguistics, 2025. URL https://arxiv.org/abs/2412.12559.
- Haonian Ji, Kaiwen Xiong, Siwei Han, Peng Xia, Shi Qiu, Yiyang Zhou, Jiaqi Liu, Jinlong Li, Bingzhou Li, Zeyu Zheng, Cihang Xie, and Huaxiu Yao. ClawArena: Benchmarking AI agents in evolving information environments, 2026. URL https://arxiv.org/abs/2604.04202. Cited for MetaClaw’s distilled procedural-skill injection that adapts behavior without weight updates.
- Yichao Ji. Context engineering for AI agents: Lessons from building Manus. Manus Blog, https://manus.im/blog/Context-Engineering-for-AI-Agents-Lessons-from-Building-Manus, jul 2025. Accessed 2026-06-24.
- Huiqiang Jiang, Qianhui Wu, Chin-Yew Lin, Yuqing Yang, and Lili Qiu. LLMLingua: Compressing prompts for accelerated inference of large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 13358–13376, Singapore, December 2023a. Association for Computational Linguistics. URL https://aclanthology.org/2023.emnlp-main.825/. [CrossRef]
- Huiqiang Jiang, Qianhui Wu, Chin-Yew Lin, Yuqing Yang, and Lili Qiu. LLMLingua: Innovating LLM efficiency with prompt compression. Microsoft Research Blog, December 2023b. URL https://www.microsoft.com/en-us/research/blog/llmlingua-innovating-llm-efficiency-with-prompt-compression/. Accessed: 2026-06-24.
- Huiqiang Jiang, Yucheng Li, Chengruidong Zhang, Qianhui Wu, Xufang Luo, Surin Ahn, Zhenhua Han, Amir H. Abdi, Dongsheng Li, Chin-Yew Lin, Yuqing Yang, and Lili Qiu. MInference 1.0: Accelerating pre-filling for long-context LLMs via dynamic sparse attention. In Advances in Neural Information Processing Systems (NeurIPS), 2024a. URL https://arxiv.org/abs/2407.02490.
- Huiqiang Jiang, Qianhui Wu, Xufang Luo, Dongsheng Li, Chin-Yew Lin, Yuqing Yang, and Lili Qiu. LongLLMLingua: Accelerating and enhancing LLMs in long context scenarios via prompt compression. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1658–1677, Bangkok, Thailand, August 2024b. Association for Computational Linguistics. URL https://aclanthology.org/2024.acl-long.91/. [CrossRef]
- Greg Kamradt. Needle in a haystack – pressure testing LLMs. https://github.com/gkamradt/LLMTest_NeedleInAHaystack, 2023. Open-source long-context recall evaluation; accessed June 2026.
- Jiazheng Kang, Mingming Ji, Zhe Zhao, and Ting Bai. Memory OS of AI agent. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2025a. URL https://arxiv.org/abs/2506.06326. Oral. arXiv:2506.06326.
- Minki Kang, Wei-Ning Chen, Dongge Han, Huseyin A. Inan, Lukas Wutschitz, Yanzhi Chen, Robert Sim, and Saravan Rajmohan. ACON: Optimizing context compression for long-horizon LLM agents. arXiv preprint arXiv:2510.00615, 2025b. URL https://arxiv.org/abs/2510.00615. Accepted to ICML 2026.
- Jang-Hyun Kim, Jinuk Kim, Sangwoo Kwon, Jae W. Lee, Sangdoo Yun, and Hyun Oh Song. Kvzip: Query-agnostic kv cache compression with context reconstruction. In Advances in Neural Information Processing Systems (NeurIPS), 2025. URL https://arxiv.org/abs/2505.23416.
- Yuri Kuratov, Aydar Bulatov, Petr Anokhin, Ivan Rodkin, Dmitry Sorokin, Artyom Sorokin, and Mikhail Burtsev. BABILong: Testing the limits of LLMs with long context reasoning-in-a-haystack. In Advances in Neural Information Processing Systems (NeurIPS), Datasets and Benchmarks Track, 2024. URL https://arxiv.org/abs/2406.10149.
- LaoZhang AI. OpenClaw cost optimization: Complete token-management guide. Blog, 2026. URL https://blog.laozhang.ai/en/posts/openclaw-cost-optimization-token-management. Community write-up cited for context-accumulation telemetry: 40–50% of token usage and a session occupying 56–58% of a 400K window (∼230K tokens re-sent per message). Vendor blog; figures are self-reported; accessed June 2026.
- Letta AI. Letta: The platform for building stateful agents (formerly MemGPT). https://www.letta.com/, 2024. Successor framework to MemGPT; OS-inspired tiered memory (core/recall/archival) with background sleeptime memory compaction. Accessed 2026-06-24.
- Ang Li, Sean McLeish, Haozhe Chen, Nimit Kalra, Zaiqian Chen, Artem Gazizov, Venkata Anoop Suhas Kumar Morisetty, Bhavya Kailkhura, Harshitha Menon, Zhuang Liu, Brian R. Bartoldson, Tom Goldstein, Sanae Lotfi, Micah Goldblum, and Pavel Izmailov. End-to-end context compression at scale, 2026a. URL https://arxiv.org/abs/2606.09659.
- Haoyang Li, Yiming Li, Anxin Tian, Tianhao Tang, Zhanchao Xu, Xuejia Chen, Nicole Hu, Wei Dong, Qing Li, and Lei Chen. A survey on large language model acceleration based on KV cache management. Transactions on Machine Learning Research (TMLR), 2025a. URL https://arxiv.org/abs/2412.19442.
- Ruoran Li, Xinghua Zhang, Haiyang Yu, Shitong Duan, Xiang Li, Wenxin Xiang, Chonghua Liao, Xudong Guo, Yongbin Li, and Jinli Suo. MemPO: Self-memory policy optimization for long-horizon agents, 2026b. URL https://arxiv.org/abs/2603.00680.
- Xiaozhe Li, Tianyi Lyu, Yizhao Yang, Liang Shan, Siyi Yang, Ligao Zhang, Zhuoyi Huang, Qingwen Liu, and Yang Li. Escaping the context bottleneck: Active context curation for llm agents via reinforcement learning, 2026c. URL https://arxiv.org/abs/2604.11462.
- Yucheng Li, Bo Dong, Chenghua Lin, and Frank Guerin. Compressing context to enhance inference efficiency of large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6342–6353, Singapore, December 2023. Association for Computational Linguistics. URL https://aclanthology.org/2023.emnlp-main.391/. [CrossRef]
- Yuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh, Acyr Locatelli, Hanchen Ye, Tianle Cai, Patrick Lewis, and Deming Chen. SnapKV: LLM knows what you are looking for before generation. In Advances in Neural Information Processing Systems (NeurIPS), 2024. URL https://arxiv.org/abs/2404.14469.
- Zongqian Li, Yinhong Liu, Yixuan Su, and Nigel Collier. Prompt compression for large language models: A survey. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, 2025b. URL https://aclanthology.org/2025.naacl-long.368/.
- Zongqian Li, Yixuan Su, and Nigel Collier. 500xcompressor: Generalized prompt compression for large language models. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (ACL). Association for Computational Linguistics, 2025c. URL https://arxiv.org/abs/2408.03094.
- Tobias Lindenbauer et al. The complexity trap: Simple observation masking is as efficient as LLM summarization for agent context management. In Proceedings of the 4th Deep Learning for Code (DL4Code) Workshop at NeurIPS, 2025. URL https://arxiv.org/abs/2508.21433. JetBrains Research.
- Jim Liu. OpenClaw context management guide: Prevent memory loss and token waste. OpenAI Tools Hub (third-party blog; unaffiliated with OpenAI), 2026. URL https://www.openaitoolshub.org/en/blog/openclaw-context-management-guide. Community write-up cited for the caveat that lowering historyLimit can increase tokens via memory_search/grep/read amplification; accessed June 2026.
- Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. Lost in the middle: How language models use long contexts. Transactions of the Association for Computational Linguistics, 12:157–173, 2024a. URL https://aclanthology.org/2024.tacl-1.9/. [CrossRef]
- Xin Liu, Runsong Zhao, Pengcheng Huang, Xinyu Liu, Junyi Xiao, Chunyang Xiao, Tong Xiao, Shengxiang Gao, Zhengtao Yu, and Jingbo Zhu. Autoencoding-free context compression for llms via contextual semantic anchors, 2025. URL https://arxiv.org/abs/2510.08907.
- Zirui Liu, Jiayi Yuan, Hongye Jin, Shaochen Zhong, Zhaozhuo Xu, Vladimir Braverman, Beidi Chen, and Xia Hu. KIVI: A tuning-free asymmetric 2bit quantization for KV cache. In Proceedings of the 41st International Conference on Machine Learning (ICML), 2024b. URL https://arxiv.org/abs/2402.02750.
- Miao Lu, Weiwei Sun, Weihua Du, Zhan Ling, Xuesong Yao, Kang Liu, and Jiecao Chen. Scaling LLM multi-turn RL with end-to-end summarization-based context management. arXiv preprint arXiv:2510.06727, 2025. URL https://arxiv.org/abs/2510.06727.
- Lvzhou Luo, Yixuan Cao, and Ping Luo. AttnComp: Attention-guided adaptive context compression for retrieval-augmented generation. In Findings of the Association for Computational Linguistics: EMNLP 2025, pages 8456–8472, 2025. URL https://arxiv.org/abs/2509.17486. arXiv:2509.17486.
- Wei Luo, Yi Huang, Songchen Ma, Huanyu Qu, Jiang Cai, and Mingkun Xu. Meta-soft: Leveraging composable meta-tokens for context-preserving kv cache compression, 2026. URL https://arxiv.org/abs/2605.22337.
- Lingrui Mei, Jiayu Yao, Yuyao Ge, Yiwei Wang, Baolong Bi, Yujun Cai, Jiazhi Liu, Mingyu Li, Zhong-Zhi Li, Duzhen Zhang, Chenlin Zhou, Jiayi Mao, Tianze Xia, Jiafeng Guo, and Shenghua Liu. A survey of context engineering for large language models. arXiv preprint arXiv:2507.13334, 2025. URL https://arxiv.org/abs/2507.13334.
- Microsoft Research. LLMLingua series: Effectively deliver information to LLMs via prompt compression. Microsoft Research Project Page, 2024. URL https://www.microsoft.com/en-us/research/project/llmlingua/. Project site mirror at https://llmlingua.com/. Accessed: 2026-06-24.
- Ali Modarressi, Hanieh Deilamsalehy, Franck Dernoncourt, Trung Bui, Ryan A. Rossi, Seunghyun Yoon, and Hinrich Schütze. NoLiMa: Long-context evaluation beyond literal matching. In International Conference on Machine Learning (ICML), 2025. URL https://arxiv.org/abs/2502.05167.
- Jesse Mu, Xiang Lisa Li, and Noah Goodman. Learning to compress prompts with gist tokens. In Advances in Neural Information Processing Systems (NeurIPS), volume 36, 2023. URL https://arxiv.org/abs/2304.08467.
- Nous Research. Hermes Agent: Context compression and caching. Developer documentation, 2026a. URL https://hermes-agent.nousresearch.com/docs/developer-guide/context-compression-and-caching. Cited for the dual ContextEngine/ContextCompressor design (50% agent / 85% gateway triggers), the multi-phase compaction, and the “system_and_3” prompt-cache breakpoint strategy. Source: https://github.com/NousResearch/hermes-agent; accessed June 2026.
- Nous Research. Hermes Agent. Project website, 2026b. URL https://hermes-agent.nousresearch.com/. Cited for isolated sub-agents and Python-RPC turns that collapse multi-step pipelines into “zero-context-cost” turns. Repository: https://github.com/NousResearch/hermes-agent; accessed June 2026.
- OpenAI. Prompt caching in the API. OpenAI, October 2024. URL https://openai.com/index/api-prompt-caching/. Announced 1 October 2024. Automatic 50% discount on cached input tokens for prompts ≥1,024 tokens; accessed June 2026.
- OpenClaw. The OpenClaw ecosystem. Official project website, 2026a. URL https://openclaw.ai/ecosystem/. Cited for the family of Claw-ecosystem variants (the official page describes a federation of ∼70 projects). Associated GitHub star-count growth figures are community-reported and not independently corroborated; accessed June 2026.
- OpenClaw. OpenClaw documentation: Agent workspace and bootstrap files. Official documentation, 2026b. URL https://docs.openclaw.ai/concepts/agent-workspace. Cited for SOUL.md/AGENTS.md/MEMORY.md/HEARTBEAT.md bootstrap injection; heartbeat-run mechanics at https://docs.openclaw.ai/gateway/heartbeat; accessed June 2026.
- OpenClaw. OpenClaw documentation: Token use and costs. Official documentation, 2026c. URL https://docs.openclaw.ai/reference/token-use. OpenClaw token-accounting docs: cache reads are “significantly cheaper than input tokens,” deferring the exact rate to provider (Anthropic) pricing, where a cache read is ∼10% of the input-token price; accessed June 2026.
- Charles Packer, Sarah Wooders, Kevin Lin, Vivian Fang, Shishir G. Patil, Ion Stoica, and Joseph E. Gonzalez. MemGPT: Towards LLMs as operating systems. arXiv preprint arXiv:2310.08560, 2023. URL https://arxiv.org/abs/2310.08560.
- Zhuoshi Pan, Qianhui Wu, Huiqiang Jiang, Menglin Xia, Xufang Luo, Jue Zhang, Qingwei Lin, Victor Rühle, Yuqing Yang, Chin-Yew Lin, H. Vicky Zhao, Lili Qiu, and Dongmei Zhang. LLMLingua-2: Data distillation for efficient and faithful task-agnostic prompt compression. In Findings of the Association for Computational Linguistics: ACL 2024, pages 963–981, Bangkok, Thailand, August 2024. Association for Computational Linguistics. URL https://aclanthology.org/2024.findings-acl.57/. [CrossRef]
- Joon Sung Park, Joseph C. O’Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (UIST ’23), San Francisco, CA, USA, 2023. Association for Computing Machinery. URL https://dl.acm.org/doi/10.1145/3586183.3606763. [CrossRef]
- plgonzalezrx8. OpenClaw context manager skill. GitHub repository (OpenClaw-Skills), 2026. URL https://github.com/plgonzalezrx8/OpenClaw-Skills/blob/master/context-manager/SKILL.md. Cited for AI-summary-and-replace at 70–80% context usage with a JSONL backup written to memory/compressed/ before session reset; accessed June 2026.
- Jielin Qiu, Zuxin Liu, Zhiwei Liu, Rithesh Murthy, Jianguo Zhang, Haolin Chen, Shiyu Wang, Ming Zhu, Liangwei Yang, Juntao Tan, Roshan Ram, Akshara Prabhakar, Tulika Awalgaonkar, Zixiang Chen, Zhepeng Cen, Cheng Qian, Shelby Heinecke, Weiran Yao, Silvio Savarese, Caiming Xiong, and Huan Wang. Locobench-agent: An interactive benchmark for LLM agents in long-context software engineering, 2025. URL https://arxiv.org/abs/2511.13998.
- Preston Rasmussen, Pavlo Paliychuk, Travis Beauvais, Jack Ryan, and Daniel Chalef. Zep: A temporal knowledge graph architecture for agent memory. arXiv preprint arXiv:2501.13956, 2025. URL https://arxiv.org/abs/2501.13956.
- David Rau, Shuai Wang, Hervé Déjean, and Stéphane Clinchant. Context embeddings for efficient answer generation in rag. In Proceedings of the 18th ACM International Conference on Web Search and Data Mining (WSDM). ACM, 2025. URL https://arxiv.org/abs/2407.09252. [CrossRef]
- Benjamin Rombaut. Inside the scaffold: A source-code taxonomy of coding agent architectures. arXiv preprint arXiv:2604.03515, 2026. URL https://arxiv.org/abs/2604.03515.
- Cosmo Santoni. Contextual memory virtualisation: DAG-based state management and structurally lossless trimming for LLM agents, 2026. URL https://arxiv.org/abs/2602.22402. Cited for the observed 132K→2.3K-token (98%) autocompaction reduction and DAG-based context virtualization.
- Significant-Gravitas (AutoGPT). Vector memory revamp: Removing external vector databases from AutoGPT. GitHub Pull Request #4208, https://github.com/Significant-Gravitas/AutoGPT/pull/4208, 2023. Replaced external vector-DB backends with a JSON store and brute-force NumPy similarity search; merged 2023-05-25. Accessed 2026-06-24.
- Calvin Smith. Improve performance of LLM summarizing condenser. OpenHands GitHub Pull Request #6597, 2025. URL https://github.com/OpenHands/OpenHands/pull/6597. Latency ∼8s with condenser vs 16s baseline at iteration 100; 200 vs 203 SWE-bench Verified instances; accessed June 2026.
- Theodore R. Sumers, Shunyu Yao, Karthik Narasimhan, and Thomas L. Griffiths. Cognitive architectures for language agents. Transactions on Machine Learning Research (TMLR), 2024. URL https://arxiv.org/abs/2309.02427. arXiv:2309.02427.
- Weiwei Sun, Miao Lu, Zhan Ling, Kang Liu, Xuesong Yao, Yiming Yang, and Jiecao Chen. Scaling long-horizon llm agent via context-folding. arXiv preprint arXiv:2510.11967, 2025. URL https://arxiv.org/abs/2510.11967.
- Sijun Tan, Xiuyu Li, Shishir G. Patil, Ziyang Wu, Tianjun Zhang, Kurt Keutzer, Joseph E. Gonzalez, and Raluca Ada Popa. Lloco: Learning long contexts offline. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP). Association for Computational Linguistics, 2024. URL https://arxiv.org/abs/2404.07979.
- Hanlin Tang, Yang Lin, Jing Lin, Qingsen Han, Shikuan Hong, Yiwu Yao, and Gongyi Wang. RazorAttention: Efficient KV cache compression through retrieval heads. In International Conference on Learning Representations (ICLR), 2025. URL https://arxiv.org/abs/2407.15891.
- Jiaming Tang, Yilong Zhao, Kan Zhu, Guangxuan Xiao, Baris Kasikci, and Song Han. Quest: Query-aware sparsity for efficient long-context LLM inference. In Proceedings of the 41st International Conference on Machine Learning (ICML), 2024. URL https://arxiv.org/abs/2406.10774.
- Nikhil Verma. Active context compression: Autonomous memory management in LLM agents. arXiv preprint arXiv:2601.07190, 2026. URL https://arxiv.org/abs/2601.07190. The agent/system is referred to as the “Focus Agent”.
- Qingyue Wang, Yanhe Fu, Yanan Cao, Shuai Wang, Zhiliang Tian, and Liang Ding. Recursively summarizing enables long-term dialogue memory in large language models. arXiv preprint arXiv:2308.15022, 2023. URL https://arxiv.org/abs/2308.15022. Journal version in Neurocomputing (2025). [CrossRef]
- Xingyao Wang, Simon Rosenberg, Juan Michelini, Calvin Smith, Hoang Tran, Engel Nyst, Rohit Malhotra, Xuhui Zhou, Valerie Chen, Robert Brennan, and Graham Neubig. The OpenHands software agent SDK: A composable and extensible foundation for production agents. arXiv preprint arXiv:2511.03690, 2025a. URL https://arxiv.org/abs/2511.03690.
- Yu Wang, Ryuichi Takanobu, Zhiqi Liang, Yuzhen Mao, Yuanzhe Hu, Julian McAuley, and Xiaojian Wu. Mem-α: Learning memory construction via reinforcement learning, 2025b. URL https://arxiv.org/abs/2509.25911.
- Yuhang Wang, Yuling Shi, Mo Yang, Rongrui Zhang, Shilin He, Heng Lian, Yuting Chen, Siyu Ye, Kai Cai, and Xiaodong Gu. SWE-Pruner: Self-adaptive context pruning for coding agents. arXiv preprint arXiv:2601.16746, 2026. URL https://arxiv.org/abs/2601.16746.
- Keying Wu. How to cut 90% of your OpenClaw token usage. AI Agents Hub (blog), 2026. URL https://www.aiagentshub.net/blog/openclaw-token-cost-optimization-guide. Community write-up cited for heartbeat token/cost estimates (a 5-min heartbeat carrying 50K tokens ≈600K input tokens/hour) and cache-TTL misalignment cost spikes; accessed June 2026.
- Xixi Wu, Kuan Li, Yida Zhao, Liwen Zhang, Litu Ou, Huifeng Yin, Zhongwang Zhang, Xinmiao Yu, Dingchu Zhang, Yong Jiang, Pengjun Xie, Fei Huang, Minhao Cheng, Shuai Wang, Hong Cheng, and Jingren Zhou. ReSum: Unlocking long-horizon search intelligence via context summarization. arXiv preprint arXiv:2509.13313, 2025. URL https://arxiv.org/abs/2509.13313.
- Yong Wu, YanZhao Zheng, TianZe Xu, ZhenTao Zhang, YuanQiang Yu, JiHuai Zhu, Chao Ma, BinBin Lin, BaoHua Dong, HangCheng Zhu, RuoHui Huang, and Gang Yu. ContextBudget: Budget-aware context management for long-horizon search agents, 2026. URL https://arxiv.org/abs/2604.01664.
- Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis. Efficient streaming language models with attention sinks. In International Conference on Learning Representations (ICLR), 2024. URL https://arxiv.org/abs/2309.17453.
- Guangxuan Xiao, Jiaming Tang, Jingwei Zuo, Junxian Guo, Shang Yang, Haotian Tang, Yao Fu, and Song Han. DuoAttention: Efficient long-context LLM inference with retrieval and streaming heads. In International Conference on Learning Representations (ICLR), 2025. URL https://arxiv.org/abs/2410.10819.
- Fangyuan Xu, Weijia Shi, and Eunsol Choi. RECOMP: Improving retrieval-augmented LMs with compression and selective augmentation. In International Conference on Learning Representations (ICLR), 2024. URL https://arxiv.org/abs/2310.04408. arXiv:2310.04408.
- Wujiang Xu, Zujie Liang, Kai Mei, Hang Gao, Juntao Tan, and Yongfeng Zhang. A-MEM: Agentic memory for LLM agents. In Advances in Neural Information Processing Systems (NeurIPS), 2025. URL https://arxiv.org/abs/2502.12110. arXiv:2502.12110.
- Yilong Xu, Zhi Zheng, Xiang Long, Yujun Cai, and Yiwei Wang. Self-manager: Parallel agent loop for long-form deep research, 2026. URL https://arxiv.org/abs/2601.17879.
- Jiangze Yan, Yi Shen, Wenjing Zhang, Jieyun Huang, Zhaoxiang Liu, Ning Wang, Kai Wang, and Shiguo Lian. HiMPO: Hindsight-informed memory policy optimization for less-entangled credit in long-horizon agents, 2026. URL https://arxiv.org/abs/2606.16285.
- Sikuan Yan, Xiufeng Yang, Zuchao Huang, Ercong Nie, Zifeng Ding, Zonggen Li, Xiaowen Ma, Jinhe Bi, Kristian Kersting, Jeff Z. Pan, Hinrich Schütze, Volker Tresp, and Yunpu Ma. Memory-R1: Enhancing large language model agents to manage and utilize memories via reinforcement learning, 2025. URL https://arxiv.org/abs/2508.19828.
- Walden Yan. Don’t build multi-agents. Cognition Blog, June 2025. URL https://cognition.com/blog/dont-build-multi-agents. Published 12 June 2025; accessed June 2026.
- Rui Ye, Zhongwang Zhang, Kuan Li, Huifeng Yin, Zhengwei Tao, Yida Zhao, Liangcai Su, Liwen Zhang, Zile Qiao, Xinyu Wang, Pengjun Xie, Fei Huang, Siheng Chen, Jingren Zhou, and Yong Jiang. Agentfold: Long-horizon web agents with proactive context management. arXiv preprint arXiv:2510.24699, 2025. URL https://arxiv.org/abs/2510.24699.
- Howard Yen, Tianyu Gao, Minmin Hou, Ke Ding, Daniel Fleischer, Peter Izsak, Moshe Wasserblat, and Danqi Chen. HELMET: How to evaluate long-context language models effectively and thoroughly. In International Conference on Learning Representations (ICLR), 2025. URL https://arxiv.org/abs/2410.02694.
- Lu Yi, Runlin Lei, Liuyi Yao, Yuexiang Xie, Yuyang Li, Wenhao Zhang, Zhewei Wei, Yaliang Li, and Jian-Yun Nie. Learning agent-compatible context management for long-horizon tasks, 2026. URL https://arxiv.org/abs/2605.30785.
- Hongli Yu, Tinghong Chen, Jiangtao Feng, Jiangjie Chen, Weinan Dai, Qiying Yu, Ya-Qin Zhang, Wei-Ying Ma, Jingjing Liu, Mingxuan Wang, and Hao Zhou. MemAgent: Reshaping long-context LLM with multi-conv RL-based memory agent, 2025. URL https://arxiv.org/abs/2507.02259.
- Yi Yu, Liuyi Yao, Yuexiang Xie, Qingquan Tan, Jiaqi Feng, Yaliang Li, and Libing Wu. Agentic memory: Learning unified long-term and short-term memory management for large language model agents, 2026. URL https://arxiv.org/abs/2601.01885.
- Qianhao Yuan, Jie Lou, Zichao Li, Jiawei Chen, Yaojie Lu, Hongyu Lin, Le Sun, Debing Zhang, and Xianpei Han. MemSearcher: Training LLMs to reason, search and manage memory via end-to-end reinforcement learning, 2025. URL https://arxiv.org/abs/2511.02805.
- Peitian Zhang, Zheng Liu, Shitao Xiao, Ninglu Shao, Qiwei Ye, and Zhicheng Dou. Long context compression with activation beacon, 2024a. URL https://arxiv.org/abs/2401.03462.
- Qianchi Zhang, Hainan Zhang, Liang Pang, Hongwei Zheng, and Zhiming Zheng. AdaComp: Extractive context compression with adaptive predictor for retrieval-augmented large language models. arXiv preprint arXiv:2409.01579, 2024b. URL https://arxiv.org/abs/2409.01579.
- Qizheng Zhang, Changran Hu, Shubhangi Upasani, Boyuan Ma, Fenglu Hong, Vamsidhar Kamanuru, Jay Rainton, Chen Wu, Mengmeng Ji, Hanchen Li, Urmish Thakker, James Zou, and Kunle Olukotun. Agentic context engineering: Evolving contexts for self-improving language models. arXiv preprint arXiv:2510.04618, 2025a. URL https://arxiv.org/abs/2510.04618.
- Xinrong Zhang, Yingfa Chen, Shengding Hu, Zihang Xu, Junhao Chen, Moo Khai Hao, Xu Han, Zhen Leng Thai, Shuo Wang, Zhiyuan Liu, and Maosong Sun. ∞bench: Extending long context evaluation beyond 100K tokens. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL), pages 15262–15277, 2024c. URL https://aclanthology.org/2024.acl-long.814/.
- Yihao Zhang, Zeming Wei, Xiaokun Luan, Chengcan Wu, Zhixin Zhang, Jiangrong Wu, Haolin Wu, Huanran Chen, Jun Sun, and Meng Sun. ClawWorm: Self-propagating attacks across LLM agent ecosystems. 2026. URL https://arxiv.org/abs/2603.15727. Cited for the controlled-testbed demonstration of persistent, self-propagating compromise in OpenClaw and its analysis of flat context trust and unconditional configuration loading as architectural root causes.
- Yuxiang Zhang, Jiangming Shu, Ye Ma, Xueyuan Lin, Shangxi Wu, and Jitao Sang. Memory as action: Autonomous context curation for long-horizon agentic tasks. arXiv preprint arXiv:2510.12635, 2025b. URL https://arxiv.org/abs/2510.12635.
- Zeyu Zhang, Xiaohe Bo, Chen Ma, Rui Li, Xu Chen, Quanyu Dai, Jieming Zhu, Zhenhua Dong, and Ji-Rong Wen. A survey on the memory mechanism of large language model based agents. arXiv preprint arXiv:2404.13501, 2024d. URL https://arxiv.org/abs/2404.13501.
- Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianmin Zheng, Ruisi Cai, Zhao Song, Yuandong Tian, Christopher Ré, Clark Barrett, Zhangyang Wang, and Beidi Chen. H2O: Heavy-hitter oracle for efficient generative inference of large language models. In Advances in Neural Information Processing Systems (NeurIPS), 2023. URL https://arxiv.org/abs/2306.14048.
- Yiqing Zhou, Yu Lei, Shuzheng Si, Qingyan Sun, Wei Wang, Yifei Wu, Hao Wen, Gang Chen, Fanchao Qi, and Maosong Sun. From context to edus: Faithful and structured context compression via elementary discourse unit decomposition, 2025a. URL https://arxiv.org/abs/2512.14244.
- Zijian Zhou, Ao Qu, Zhaoxuan Wu, Sunghwan Kim, Alok Prakash, Daniela Rus, Jinhua Zhao, Bryan Kian Hsiang Low, and Paul Pu Liang. MEM1: Learning to synergize memory and reasoning for efficient long-horizon agents, 2025b. URL https://arxiv.org/abs/2506.15841.
- Jialiang Zhu, Gongrui Zhang, Xiaolong Ma, Lin Xu, Miaosen Zhang, Ruiqi Yang, Song Wang, Kai Qiu, Zhirong Wu, Qi Dai, Ruichun Ma, Bei Liu, Yifan Yang, Chong Luo, Zhengyuan Yang, Linjie Li, Lijuan Wang, Weizhu Chen, Xin Geng, and Baining Guo. Re-trac: Recursive trajectory compression for deep search agents, 2026. URL https://arxiv.org/abs/2602.02486.
- Zylos Research. AI agent context compression: Strategies for long-running sessions. Commentary / blog, 2026. URL https://zylos.ai/research/2026-02-28-ai-agent-context-compression-strategies/. Commentary cited for the 2026 context-window plateau and the shift to “hybrid compression+caching, and memory-augmented architectures.”; accessed June 2026.







| Stage | Question | Representative mechanisms |
|---|---|---|
| Admission | What is allowed into context? | tool-output caps, schema filtering, retrieval top-k, observation masking |
| Placement | Where should information live? | in-window text, soft tokens, KV cache, filesystem, vector memory, sub-agent context |
| Compaction | How is old context reduced? | summarization, token pruning, learned memory writes, KV eviction/quantization |
| Recovery | Can deleted detail be reconstructed? | raw transcript replay, file pointers, retrieval stores, KVzip-style reconstruction |
| Reuse | Can the prefix be cached or shared? | prompt caching, stable head/tail layout, sub-agent summaries |
| Governance | Can provenance and security be audited? | structured summaries, source spans, privilege-aware compression |
| Agent | Trigger | Protected span | Compaction target | Recovery source | Cache interaction |
|---|---|---|---|---|---|
| Hermes Nous Research (2026a) | 50% (token) / 85% (char) | first 3 msgs + token-budgeted tail (≥20 msgs) | structured middle summary | — | Y (system_and_3 breakpoints; update-not-rewrite) |
| OpenClaw centminmod (2026); plgonzalezrx8 (2026) | /compact; 70–80% (skill) | — | /compact summary | JSONL backup in memory/compressed/ | Y (contextPruning aligned to cache-ttl) |
| Claude Code Anthropic (2026a) | ∼80% (server-side) | — | server-side summary | memory tool | — |
| OpenHands All Hands AI (2025a) | condenser-dependent | — | view-level markers | raw event stream (replayable) | trade-off noted: condensation lowers cache utilization |
| Manus Ji (2025) | — | — | filesystem offload | filesystem (re-read on demand) | — |
| Cline Baumann (2025) | LLM-decided (condense tool) | Focus-Chain checklist | condense summary | — | — |
| Devin/Cognition Yan (2025) | — | — | single-thread compression LLM | — | — |
| Cluster | What is learned | Representative methods | Space / level | Lifecycle stages |
|---|---|---|---|---|
| Compression as action (§6.1) | when to fold a sub-trajectory and what summary to reset the window to | Context-Folding Sun et al. (2025), AgentFold Ye et al. (2025), SUPO Lu et al. (2025), ReSum Wu et al. (2025) | token space; trajectory | compaction (fold→summary); placement (window reset) |
| Memory as action (§6.2) | when to write / update / delete an explicit memory state | MEM1 Zhou et al. (2025b), Memory-as-Action Zhang et al. (2025b), MemPO Li et al. (2026b), AgeMem Yu et al. (2026) | token space; memory | compaction (memory rewrite); placement (<mem>-only prompt) |
| Credit assignment (§6.3) | how to attribute reward to a single memory write | HiMPO Yan et al. (2026) | token space; memory | compaction (credit for memory writes) |
| Budget-conditioned / external managers (§6.4) | compression level given a remaining-token budget; or an external manager over a frozen agent | ContextBudget Wu et al. (2026), AdaCoM Yi et al. (2026), ContextCurator Li et al. (2026c), RE-TRAC Zhu et al. (2026) | token space; trajectory / memory | compaction (budget-conditioned); placement (manager edits context) |
| Method (family) | Space | Reported ratio / savings | Trade-offs |
|---|---|---|---|
| LLMLingua Jiang et al. (2023a) | Token (pruning) | up to 20× prompt tokens; pt perf. | Perplexity-based; coarse, may drop key tokens |
| LongLLMLingua Jiang et al. (2024b) | Token (pruning) | ∼4× fewer tokens, +21.4% perf. (NQ); 94% cost cut (LooGLE); 1.4–2.6× faster | Query-aware; needs small LM for scoring |
| LLMLingua-2 Pan et al. (2024) | Token (classif.) | 2–5× prompt tokens; up to 2.9× lower latency | Task-agnostic, faithful; needs trained encoder |
| RECOMP Xu et al. (2024) | Token (RAG) | extractive & abstractive summaries | Over-compression hurts multi-hop QA |
| AttnComp Luo et al. (2025) | Token (RAG) | 17× avg; +1.9 pt acc.; 49% latency | Needs attention access |
| xRAG Cheng et al. (2024) | Embedding (RAG) | 1 token/doc; 3.53× fewer FLOPs | Retriever–LLM bridge training |
| Gisting Mu et al. (2023) | Embedding | up to 26× prompt tokens; 40% FLOPs; 4.2% faster | Short prompts only; requires fine-tuning |
| AutoCompressor Chevalier et al. (2023) | Embedding | up to 30,720-token contexts | Slow training; tuned tokens tied to model |
| ICAE Ge et al. (2024) | Embedding | 4× prompt tokens; >2× (up to >7× cached) speedup; ∼20GB mem. saved | Needs LoRA encoder; data-leakage concerns |
| 500xCompressor Li et al. (2025c) | Embedding (KV) | 6–480× prompt tokens; retains 62–73% | Large capability loss at extreme ratios |
| Activation Beacon Zhang et al. (2024a) | Embedding/KV | 4K→400K ctx; 8× KV mem; ∼2× faster | Needs training; sliding-window stream |
| StreamingLLM Xiao et al. (2024) | KV cache (evict) | 4M+ tokens; attention sinks | Content-agnostic; drops middle tokens |
| H2O Zhang et al. (2023) | KV cache (evict) | up to 29× tput; 1.9× latency; 5× mem | Eviction permanent; weaker on reasoning |
| SnapKV Li et al. (2024) | KV cache (evict) | 3.6× faster gen; 8.2× mem (16K) | Per-head selection cost |
| KIVI Liu et al. (2024b) | KV cache (quant) | 2-bit; ∼2.6× mem; 2.35–3.47× tput | Quantization error at low bits |
| KVzip Kim et al. (2025) | KV cache (recov) | 3–4× size; ∼2× latency; query-agnostic | Reconstruction-pass overhead |
| Method (family) | Space | Reported ratio / savings | Trade-offs |
|---|---|---|---|
| ACON Kang et al. (2025b) | Trajectory (abs.) | 26–54% peak-token cut | Compression LLM call overhead |
| Focus Verma (2026) | Trajectory (agent) | 22.7% (18–57%) tokens, equal acc. () | Needs aggressive prompting; task-dependent |
| OpenHands condenser All Hands AI (2025a) | Trajectory (abs.) | ∼50% cost; latency 16→8 s | Lower cache utilization |
| Hermes ContextCompressor Nous Research (2026a) | Trajectory (abs.) | 95K→45K tokens ex.; +75% via caching | Lossy; summary-model context-size dependency |
| Manus filesystem Ji (2025); Bhavsar (2025) | Offload (recoverable) | Restorable offload; ∼100:1 input:output token ratio (motivation), not a compression ratio | Retrieval latency; summarization recall risk |
| MemGPT Packer et al. (2023) | Memory hierarchy | “unbounded” effective context | Paging logic complexity; retrieval precision |
| Mem0 Chhikara et al. (2025) | Memory hierarchy | 90%+ token/cost cut; 91% lower p95 (LoCoMo) | Extraction/retrieval precision |
| Context-Folding Sun et al. (2025) | Trajectory (learned) | ∼10× smaller active ctx; parity acc. | Requires RL training (FoldGRPO) |
| MEM1 Zhou et al. (2025b) | Memory (learned) | constant-size state; perf at less mem | Fixed-size state; requires RL |
| SUPO Lu et al. (2025) | Trajectory (learned) | effective horizon (4K→32K window) | Requires RL (GRPO); summary quality |
| MemPO Li et al. (2026b) | Memory (learned) | F1; 67–73% fewer tokens | Requires RL; per-step <mem> write |
| ContextCurator Li et al. (2026c) | Trajectory (learned) | ∼8× ctx (46.7K→6.6K), +acc. | External curator; requires RL |
| Metric | What it captures | Why it is needed |
|---|---|---|
| Task success / answer quality | Whether the agent still solves the task after compression | Compression that degrades the outcome is not a saving |
| Peak context tokens | Largest prefill the agent ever pays for | Bounds worst-case cost and context-window pressure |
| Total session input tokens | Cumulative repeated-prefill bill across all turns | The quantity actually billed over a long-horizon session |
| Output tokens & tool-call count | Extra generation and tool work induced by compression | Summarization and re-retrieval can offset prefill savings |
| Latency per turn & total wall time | End-to-end responsiveness | Compaction and recovery add wall-clock cost users feel |
| Cache write/read ratio & hit rate | How compression interacts with prompt caching | Compaction invalidates prefixes and can trigger cache-miss spikes |
| Recovery rate | Whether omitted detail is retrievable when later needed | Lossy compaction must not cause delayed-relevance failures |
| Drift / faithfulness | Whether iterated summaries stay faithful to the source | Compounding summarization errors silently corrupt state |
| Security / provenance | Whether privileged or untrusted boundaries survive compaction | Merged provenance enables prompt-injection and privilege leakage |
| Benchmark type | What it measures | Why insufficient alone | Needed agent addition |
|---|---|---|---|
| Needle / retrieval probes | Positional recall | Weak predictor of tool agents | Add delayed dependency after tool use |
| LongBench-style QA | Long-context answer quality | Single-shot, not looped | Add total-session token accounting |
| SWE / web agents | Task completion | Scaffold-dependent | Standardize cache and tool-output logging |
| Memory benchmarks | Multi-session recall | May ignore cost | Add budgeted memory pressure |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).