Preprint
Review

This version is not peer-reviewed.

Memory in Vision-Language Models: Taxonomy, Mechanisms, and Applications

  † These authors contributed equally to this work.

Submitted:

15 September 2026

Posted:

21 September 2026

You are already at the latest version

Abstract
Vision-language models (VLMs) have achieved remarkable progress in multimodal understanding and generation by integrating visual and linguistic representations. However, most current VLMs lack explicit mechanisms for persistent memory, limiting their ability to maintain contextual coherence, accumulate knowledge over time, and support long-term reasoning across extended interactions. To address these limitations, a diverse set of memory mechanisms has emerged, including latent memory, key-value caches, external memory stores, retrieval-augmented memory, and hybrid memory systems. Despite rapid advances, the design space of memory in VLMs remains fragmented, and a unified understanding of its design and application remains lacking. In this survey, we provide a comprehensive review of memory mechanisms in VLMs from a system-oriented perspective. We introduce a novel four-dimensional (4D) taxonomy that organizes existing approaches along four orthogonal aspects: when memory is maintained (temporal scope), where memory is stored (storage location), what memory encodes (information stored), and how memory is accessed, updated, and utilized (memory operations). Using this taxonomy as a unifying framework, we systematically analyze representative memory-enhanced VLM architectures, review evaluation protocols for memory capabilities, and summarize key application areas, including long-video understanding, multimodal dialogue, embodied agents, and robotic reasoning. We further discuss key challenges such as scalability, memory efficiency, continual updating and forgetting, and multimodal grounding, and outline promising directions toward adaptive and unified memory systems. Overall, this survey provides a comprehensive foundation for understanding memory in VLMs and serves as a roadmap for developing next-generation multimodal systems with persistent memory and long-term reasoning capabilities. A repository associated with this survey is available at https://github.com/Xzcv-hub/Awesome-Memory-in-VLM.
Keywords: 
;  ;  ;  ;  ;  ;  ;  
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.