Submitted:
26 September 2026
Posted:
28 September 2026
You are already at the latest version
Abstract
Knowledge distillation-based unsupervised anomaly detection has achieved strong performance in industrial inspection. However, existing Reverse Distillation (RD) methods suffer from two fundamental limitations: the bottleneck compresses semantically entangled CNN features into a one-class representation that inadequately characterizes normal structural regularity, and the pairwise distillation objective leaves the co-learned teacher-student geometry semantically underconstrained. To address these issues, we propose Semantically Grounded Reverse Distillation (SGRD), which integrates a frozen vision foundation model (DINOv2 in this work) as a fixed semantic reference into the RD framework. Specifically, a foundation-aware semantic bottleneck conditions the bottleneck encoding pathway with foundation-model patch semantics to produce a structurally coherent one-class code. Furthermore, a foundation-semantic representation anchoring module jointly constrains both teacher and student representations toward the foundation-model semantic subspace, preventing semantically arbitrary representation geometry. Experiments on MVTec AD achieve 98.5% image-level AUROC and 98.3% pixel-level AUROC, while experiments on VisA achieve 99.8% image-level AUROC and 94.2% pixel-level AUPRO, with the proposed method attaining the strongest overall average performance among the compared methods on both benchmarks.
Keywords:
anomaly detection
; knowledge distillation
; vision foundation model
; reverse distillation
; semantic grounding
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.