Preprint
Article

This version is not peer-reviewed.

Robust Vision RAG: Mitigating Retrieval Poisoning with Semantic Coherence Refinement

Submitted:

07 October 2026

Posted:

10 October 2026

You are already at the latest version

Abstract
Retrieval-Augmented Generation (RAG) enhances text-to-image diffusion models by grounding generation in retrieved visual exemplars, but recent studies reveal that multimodal retrieval pipelines are highly vulnerable to poisoning attacks. When adversaries corrupt the retrieval database, semantically mismatched exemplars such as retrieving images of cats for prompts requesting dogs can mislead diffusion models into generating incorrect or misleading outputs. We identify this failure mode as a breakdown of semantic coherence between the text prompt and retrieved visual context. To address this issue, we propose a score-based semantic coherence refinement module that explicitly evaluates prompt-image consistency, refines misaligned prompt components, and re-retrieves corrected exemplars prior to diffusion. Acting as a multimodal feedback loop, the proposed method prevents poisoned retrieval from propagating semantic errors into the generative process. Extensive experiments demonstrate that our approach significantly improves semantic correctness, alignment, and robustness under both clean and poisoned retrieval settings, establishing an effective and principled defense for Vision RAG-augmented diffusion models.
Keywords: 
;  ;  ;  
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.