Submitted:
21 August 2026
Posted:
28 August 2026
You are already at the latest version
Abstract
Industrial anomaly detection supports automated quality inspection. Traditionally, each product has required its own detector, trained on images without defects and evaluated on benchmarks with fixed classes. Recent studies use vision encoders pretrained on web data, multimodal large language models, expert mixtures, and generative models to relax this setup. They transfer visual features, add reasoning through language, share models across categories and modalities, or generate abnormal training data. Existing taxonomies centered on reconstruction, embedding, distillation, and memory do not describe these developments well. We therefore organize the literature published over the past three years according to five sources of information that replace training on normal samples for each category, namely visual, reasoning, geometric and multimodal, universal, and synthesis priors. Reported results show progress on curated benchmarks, but evidence remains limited under domain shift, protocols for open world settings, and operating points with low false positive rates. We also analyze recent benchmarks and recurring deployment bottlenecks. Industrial video grounded in physical dynamics provides an emerging direction beyond inspection of single frames. The evidence points to a change in research focus, although current methods do not yet provide a general replacement for separate training on normal data for each product in deployed inspection.
Keywords:
industrial anomaly detection
; foundation models
; vision-language models
; zero-shot anomaly detection
; anomaly synthesis
; benchmark evaluation
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.