Submitted:
09 September 2026
Posted:
14 September 2026
You are already at the latest version
Abstract
World models are increasingly used to predict and simulate how environments evolve, yet their evaluation remains fragmented across video generation, autonomous driving, and robotics. World-model benchmarks differ in four respects: what capability is evaluated, how the model is evaluated, which metrics are used, and where the data come from. Thus, similar scores may reflect different evidence and do not necessarily support the same capability claims. We present an evaluation-centric survey of 102 representative benchmarks released between 2018 and 2026. We characterize them through four coupled dimensions: Evaluation Target, Evaluation Protocol, Evaluation Metrics, and Evaluation Data. Evaluation Target encompasses visual and temporal quality, spatial and state consistency, long-horizon memory and state persistence, physical plausibility, causal and counterfactual reasoning, control fidelity and interactive dynamics, and functional utility. Our analysis reveals that evaluation targets are expanding toward Control Fidelity and Interactive Dynamics and Functional Utility. However, benchmarks in the former category predominantly use open-loop protocols, leaving reliability under feedback-driven interaction insufficiently tested. These findings highlight the need to align evaluation protocols, metrics, and data with the capability claims they are intended to support. We identify priorities for evaluating functional utility across downstream roles, developing intervention-based and closed-loop protocols, combining prediction-level and downstream outcome metrics, constructing multimodal action-grounded data, and building standardized and reproducible evaluation toolkits. By clarifying what each benchmark measures and which claims its evidence can support, this survey provides a roadmap toward credible and comparable world-model evaluation. The survey website is available at https://world-model-benchmarks.github.io/.
Keywords:
world models
; evaluation benchmarks
; evaluation taxonomy
; embodied intelligence
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.