Preprint
Review

This version is not peer-reviewed.

Diffusion-Based World Models: A Survey

Submitted:

12 September 2026

Posted:

14 September 2026

You are already at the latest version

Abstract
World models have become a key foundation for artificial general intelligence and autonomous agents, enabling perception, understanding, and reasoning via internal representations. Diffusion models, empowered by high-fidelity generation and flexible conditional modeling, have become an important avenue for building representation-based world models. Despite rapid progress, the community still lacks a dedicated and systematic survey of diffusion-based world models that systematically analyzes their distinctive strengths and inherent limitations and clearly formulates the key scientific problems that must be addressed going forward. To bridge this gap, we present a comprehensive and structured review of diffusion-based world models. We first summarize the principles that connect diffusion processes to world modeling, and organize mainstream methodological frameworks and representative systems across diverse application scenarios. We then synthesize research progress, intrinsic characteristics, and commonly used datasets, providing a unified taxonomy to position existing works. Importantly, we offer an in-depth assessment of diffusion models for world imitation, highlighting their advantages (e.g., multimodal hypothesis coverage, controllable conditional generation, and high-fidelity generation) and limitations (e.g., computational inefficiency, long-horizon consistency challenges, error accumulation, and evaluation bottlenecks). Building on this analysis, we identify critical open problems and future research directions. Overall, this survey provides a consolidated reference framework and a clear research roadmap for diffusion-based world models, facilitating rapid understanding and guiding sustained progress in this emerging area. https://github.com/energy588/Diffusion-based-World-Models.
Keywords: 
;  ;  ;  ;  

1. Introduction

1.1. Background

World models aim to learn an interactive internal simulator that captures environment dynamics—both natural evolution and action-induced transitions—in a compact latent space, thereby enabling predictive and controllable reasoning for intelligent agents. By closing the loop between past observations and behavioral feedback, a world model can generate or forecast future observations, infer latent states, maintain temporal consistency, and support queryable rollouts and counterfactual reasoning for planning and decision-making. Existing research has followed several paradigms. Early approaches adopted an “encoder–dynamics–decoder” latent generative framework, combining CNN feature extraction with VAE and RNN dynamics (e.g., World Models [1]), which is simple and efficient but often suffers from blurry reconstructions and error accumulation over long horizons. To better model uncertainty, RSSM integrates deterministic recurrence with stochastic latent variables (the Dreamer series [2,3,4]), enabling efficient imagination in latent space, yet still facing latent bias and distribution-shifted rollouts. GAN-based methods improve visual fidelity (e.g., GameGAN [5]) but are notoriously unstable and prone to mode collapse, with limited physical consistency and controllability. With the rise of large-scale pretraining, discrete tokenization and autoregressive Transformers recast world modeling as sequence prediction (VideoGPT [6], Genie-1 [7]), strengthening long-range dependencies and scalability, while being constrained by tokenization artifacts, exposure bias, and inference cost. Complementarily, representation-prediction paradigms such as JEPA [8] perform prediction and alignment in latent space to avoid pixel-level generative redundancy, offering more stable training but lacking an explicit generative rollout interface and direct controllability. They use adaptive multimodal control, these models can generate world simulations and find use in various world-to-world transfer use cases [9,10]. Owing to their demonstrated capability for high-fidelity, temporally consistent scene generation in representative large-scale video generation systems [11], as shown in Figure 2, diffusion models have rapidly become the dominant paradigm for building world models.
Figure 1. Survey at A Glance. (1) Preliminaries. We provide a theoretical analysis of existing classes of diffusion models, establishing a principled foundation for their role in world modeling. (2) Task & Data. We present a taxonomy of world model applications across three major domains and summarize the key publicly available datasets. (3) Characteristics. We systematically characterize DWMs through five key properties and analyze the strengths and limitations of each. (4) Challenges and future trends. We review the outstanding challenges in DWMs and offer insights into potential future research directions.
Figure 1. Survey at A Glance. (1) Preliminaries. We provide a theoretical analysis of existing classes of diffusion models, establishing a principled foundation for their role in world modeling. (2) Task & Data. We present a taxonomy of world model applications across three major domains and summarize the key publicly available datasets. (3) Characteristics. We systematically characterize DWMs through five key properties and analyze the strengths and limitations of each. (4) Challenges and future trends. We review the outstanding challenges in DWMs and offer insights into potential future research directions.
Preprints 233002 g001
Figure 2. The number of research papers on DWMs published between 2018 and 2026.
Figure 2. The number of research papers on DWMs published between 2018 and 2026.
Preprints 233002 g002

1.2. Motivation

Despite the rapid surge of interest in world models, existing surveys [12,13,14,15,16,17,18] largely remain at the level of high-level concepts and application landscapes, or focus on the practical value in specific vertical domains. For instance, Ding et al. [12] provide a comprehensive overview and frame world models from two primary perspectives–understanding the world and predicting the future. Tu et al. [13]summarize the role of world models in reshaping the autonomous driving paradigm. And Long et al. [19] highlight that integrating physical simulators with world models may underpin the next generation of embodied intelligence. However, these reviews rarely offer a dedicated and structured discussion of the fundamental role diffusion has played in the evolution of world modeling. In fact, diffusion models have become a prevailing backbone by virtue of their strong capability in high-fidelity generation [20,21], multimodal conditional fusion [22,23], and long-horizon temporal coherence [24,25], substantially improving the representational and generative quality of world models for complex scenes. Meanwhile, the diffusion paradigm also introduces key bottlenecks that hinder progress toward truly interactive world models, including high inference latency due to iterative sampling [26,27], limited logical compositionality and controllable generalization [28,29], and insufficient understanding of real-world physical laws and causal structure[30]–leading to deficits in predictability and verifiability. Motivated by these gaps, our survey focuses on diffusion-based world models (DWMs), systematically characterizes their core advantages and inherent limitations, distills the key scientific problems that must be addressed to build efficient, controllable, physically consistent, and interactive world models, and outlines a structured roadmap for future research.

1.3. Task Classification for DWMs

In this review, we comprehensively map the research landscape of DWMs, aiming to foster more targeted advancements in this field. As illustrated in the classification diagram, we have categorized the research domain into three core dimensions, summarizing the key roles of diffusion mechanisms across various scenarios. In autonomous driving, addressing the issue of ambiguous pedestrian and obstacle representations caused by "mean reversion" in early world models, approaches such as Vista[31] and GAIA-1[32] enhance generative fidelity and decision safety under extreme conditions by leveraging diffusion’s high fidelity advantages.Within embodied intelligence, diffusion models overcome traditional architectures’ loss of fine-grained detail. By leveraging universal simulation of UniSim[33] and visual priors of Vidar[34], researchers have established high-fidelity, physically consistent visual feedback loops.This mechanism of "guiding decisions through visual predictions" enables agents to effectively acquire complex manipulation skills through diffusion-generated virtual rehearsals. In general domains, addressing error accumulation and long-term sequence breakdown in autoregressive predictions, diffusion models—represented by general video generation research such as Sora[11] and Streetscapes[35]—achieve logically consistent generation of multi-frame high-definition videos, laying the technical foundation for general-purpose physical simulators.Therefore, comprehensively analysing the current state and future potential of DWMs across these domains not only aligns with prevailing trends in generative AI but also holds critical significance for realising AGI, as shown in Figure 3.

1.4. Characteristics of DWMs

In this review, we further define the core characteristics and underlying mechanisms that DWMs should possess. As illustrated in the characteristic classification diagram, we categorize these into five key properties:
Long-Term Spatial-Temporal Span: Refers to the model’s capacity to maintain physical evolution logic over extended durations. This enables agents to conduct deep spatio-temporal reasoning (e.g., long-range traffic condition prediction) and forms the foundation for forward-looking strategic planning.
Multimodal Fusion: Supports integration of multi-source data including text, motion, audio, and radar. Through cross-modal semantic compensation in latent spaces, the model maintains robust and accurate environmental representation even when partial signals are ambiguous.
Interactivity: Emphasises a dynamic feedback loop of "action-perception". The model updates its state in real-time based on external interventions, enabling agents to learn implicit physical knowledge and align with user preferences within a risk-free virtual environment.
Spatio-Temporal Consistency: Encompasses spatial geometric stability and temporal sequence logical continuity. High consistency is pivotal in suppressing generated “hallucinations,” ensuring reliable predicted trajectories and providing a deterministic baseline for downstream decision-making.
Environmental Diversification: Endows the model with generalization capabilities across geography, climate, and culture. By synthesising rare edge conditions (e.g., extreme weather), the model effectively addresses the long-tail problem in real-world data collection, reducing training costs.

1.5. Key Contributions

The main contributions of this survey are:
(1) We propose a perspective: the influence and role of diffusion on world models, specifically how diffusion models are integrated into world modeling.
(2) We Systematically outlines the three primary application domains and five defining characteristics of DWMs.
(3) We articulate the challenges facing DWMs and explore diffusion’s future impact on world modeling, particularly in the context of physics-compliant generative AI.
Concurrently, we examine how diffusion’s inherent limitations can be optimized and mitigated within world models.
The remainder of this paper is organized as follows: Section 2 introduces the implementation principles and definitions of diffusion probability models, alongside the diffusion models primarily employed by DWMs. Section 3 details recent research advancements and relevant open-source datasets across three key research domains for DWMs, alongside diffusion’s inherent strengths and weaknesses within these domains and their implications for World Models. Section 4 provides a detailed description of the five key characteristics of DWMs. Section 5 outlines the unresolved issues and challenges facing DWMs, while proposing future research directions, as shown in Figure 1.

2. Preliminaries

As generative probabilistic model based on non-equilibrium thermodynamics[36], diffusion models learn complex data distributions by simulating the injection and elimination of noise. Within the architecture of World Models, diffusion models are progressively replacing traditional Variational Autoencoders (VAEs) and Generative Adversarial Networks (GANs) to form the core foundation of environmental dynamics kernels.
Figure 4. A Principles of Diffusion Probabilistic Models.
Figure 4. A Principles of Diffusion Probabilistic Models.
Preprints 233002 g004
A Principles of Diffusion Probabilistic Models. The principles of diffusion probability models trace back to DDPM [20], the theoretical cornerstone of the diffusion paradigm. Its core concept involves learning data generation through a bidirectional stochastic process: a forward diffusion step transforming data into noise, and a reverse denoising step reconstructing data from noise.
In the forward diffusion process, given a sample x 0 ∼ q ( x ) from the true distribution, the forward process constitutes a Markov chain. At each time step t ∈ [ 1 , T ] ,the model adds a small Gaussian noise ϵ to the data according to a predefined variance adjustment β t . Its state transition probability is defined as:
q x t ∣ x t − 1 = N x t ; 1 − β t x t − 1 , β t I
Through reparameterization, the state x t at any time t can be directly derived from x 0 without iterative computation:
x t = α ¯ t x 0 + 1 − α ¯ t ϵ , ϵ ∼ N ( 0 , I )
where α t = 1 − β t , and α ¯ t = ∏ i = 1 t α i . When T → ∞ , x T ultimately converges to a standard normal distribution.
The inverse process aims to learn how to reconstruct the original data from the pure noise map p x T = N ( 0 , I ) . Since the true inverse transfer distribution q x t − 1 ∣ x t is mathematically non-analyzable, a neural network p θ x t − 1 ∣ x t is typically employed to approximate it:
p θ x t − 1 ∣ x t = N x t − 1 ; μ θ x t , t , ∑ θ x t , t
where θ represents the learnable parameters of the neural network μ θ denotes the predicted mean, and Σ θ x t , t denotes the predicted variance.
This paper posits that the integration of diffusion models into the world modeling research framework signifies a milestone paradigm shift from deterministic prediction toward probabilistic generation. Fundamentally, this evolution re-contextualizes the task of ’future prediction’ from primitive conditional regression to sophisticated conditional probability distribution estimation.
Early iterations of world models typically operated under the assumption that state transitions followed simple Gaussian distributions. Consequently, their objective functions were conventionally formulated based on MSE.Nevertheless, real-world physical environments exhibit pronounced stochasticity and inherent multimodality (e.g., the probabilistic nature of navigational maneuvers at intersections). Conventional MSE-based loss functions constrain the model to output the expectation of potential outcomes, thereby precipitating severe ’mean-blurring’ artifacts. In contrast, diffusion models leverage iterative denoising processes—grounded in score matching or variational inference—to characterize highly intricate non-Gaussian distributions and precisely encapsulate divergent latent evolutionary trajectories.
Furthermore, the intrinsic robustness manifested by diffusion models substantially fortifies the spatiotemporal consistency of world models during long-horizon forecasting. Moreover, they empower world models with the capacity to operate directly upon high-dimensional, pixel-level manifolds. Albeit challenging in terms of inference efficiency and computational overhead, their superior performance in complex distribution modeling has established them as a foundational component in the architecture of contemporary world models.
Diffusion Probability Model: The Foundation for Realizing World Models. As a pivotal milestone in the domain of diffusion probabilistic models, DDPM[20] established the foundational paradigm for generative modeling. Subsequently, various advancements have emerged upon this basis, encompassing conditional diffusion models characterized by robust conditioning mechanisms, Latent Diffusion Models (LDMs) that significantly enhance computational tractability, and the Diffusion Transformer (DiT), which integrates Transformer architectures to augment long-range modeling capacities. These evolutionary developments have not only elevated the fidelity and controllability of synthesized imagery but also—by virtue of their profound representational capacity for high-dimensional distributions—provided the quintessential generative substrate and analytical framework for constructing world models endowed with complex environmental simulation and multimodal perceptual capabilities.
Conditional diffusion models. Future state forecasting in world models is intrinsically a constrained probabilistic inference process, necessitating the holistic integration of conditional inputs such as current environmental features and agent-induced actions. As Eq.4, conditional diffusion models incorporate contextual variables y (e.g., actions, linguistic descriptors, or trajectories) to approximate the conditional distribution p ( x ∣ y ) , thereby effectively characterizing intricate multimodal conditional distributions. This mechanism not only bolsters the semantic alignment of synthesized content but also endows world models with controllability, enabling the precise modulation of predicted future trajectories via the strategic adjustment of input conditioning.
Consequently, the corresponding reverse denoising process is reformulated to be contingent upon the probability distribution conditioned on these auxiliary variables.
p θ x t − 1 ∣ x t , y = N x t − 1 ; μ θ x t , t , y , Σ θ x t , t , y
where y functions as a contextual condition that imposes a constraint on the diffusion process, thereby facilitating the mapping of state-action pairs into the reverse transition kernel of the diffusion framework. This mechanism not only realizes physics-informed deterministic evolution but also, during the denoising iterations, transforms ’future state prognostication’ into conditional probability distribution sampling, thus safeguarding spatiotemporal coherence and task-oriented consistency.
To augment the logical alignment between the synthesized outcomes and the conditioning variable y, world models ubiquitously adopt Classifier-Free Guidance (CFG)[37], as formulated in the following expression:
ϵ ˜ θ x t , t , y = ( 1 + ω ) ϵ θ x t , t , y − ω ϵ θ x t , t , ∅
where ω denotes the guidance weight, ϵ θ x t , t , y represents the conditional prediction, and ϵ θ x t , t , y signifies the unconditional prediction. During the inference phase, this technique aggregates conditional and unconditional predictive estimates via linear interpolation, thereby markedly enhancing the governing influence of action directives over the synthesized environmental feedback.
Latent Diffusion Model. To augment inference efficiency, researchers introduced the LDM[22], which relocates generative tasks from high-dimensional pixel space to a compact, low-dimensional latent manifold by incorporating a perceptual compression paradigm. Specifically, the LDM characterizes the underlying distributional regularities by minimizing a latent-space denoising objective function, as formulated below:
L L D M : = E E ( x ) , ϵ ∼ N ( 0 , 1 ) , t [ | | ϵ − ϵ θ ( z t , t ) | | 2 2 ]
where z t denotes the perturbed latent representation and t represents the diffusion timesteps. By modeling dynamical transitions within the latent manifold, LDMs efficiently capture intricate causal rationales and long-term spatiotemporal dependencies, providing a rigorous architectural scaffolding for constructing large-scale, multitask generative interactive environments. Compared to early world models where modeling directly in high-dimensional pixel space resulted in prohibitive computational overhead and perceptual blurring artifacts, this paradigm enables the precise simulation of physical laws within a semantically-dense space, significantly enhancing inference efficiency and the structural veridicality of predicted trajectories.
Diffusion Transformer. To fortify the capacity of world models to capture long-range spatiotemporal dependencies, researchers have further incorporated the DiT[25] architecture. The advent of DiT signifies a paradigmatic shift in generative dynamics modeling. it eschews traditional convolutional residual blocks, opting instead for a monolithic Transformer architecture possessing a global receptive field as the denoising backbone. This transition not only establishes a novel precursor for generative models but also leverages its superior parameter scalability to markedly enhance modeling precision and spatiotemporal coherence within intricate dynamic environments.
The pivotal innovation resides in partitioning the latent representation z into a sequence of continuous spatiotemporal patches, utilizing self-attention mechanisms to distill global contextual information. The underlying computational formulation is expressed as follows:
Attention ( Q , K , V ) = softmax Q K T d k V
where Q, K, and V denote the query, key, and value matrices, respectively. Compared to antecedent architectures, the conspicuous advantage of DiT lies in its extraordinary scaling properties, whereby generative fidelity and computational resource allocation exhibit improvements that strictly adhere to scaling laws as the depth or width of the Transformer layers increases. Furthermore, this query-key-based global interaction mechanism, synergized with its high-capacity representation learning, profoundly augments the capability of world models to simulate large-scale, complex real-world scenarios and maintain long-term temporal coherence, this enables the model to implicitly distill and adhere to the macroscopic evolutionary regularities of the physical world, thereby providing a rigorous foundation for the development of comprehensive world models.

3. Task Classification of DWMs

This chapter commences with a systematic dissection of the core datasets foundational to the training of world models. Subsequently, predicated on the divergence of task-specific constraints and interaction characteristics, we establish a taxonomical framework that categorizes diffusion-based world models into three primary application domains: autonomous driving, embodied intelligence, and general-purpose scenarios, as delineated in Table 2. This classification not only encapsulates contemporary mainstream research trajectories but also elucidates the evolutionary pathways of world models under heterogeneous operational constraints.
In the realm of autonomous driving, diffusion models function as high-fidelity spatiotemporal simulators, dedicated to synthesizing consistent long-tail scenarios while adhering to stringent traffic regulations and multi-sensor alignment constraints. Within the context of embodied intelligence, these models are repurposed as policy representors. by accurately modeling the inherent multimodality of action distributions, they facilitate dexterous, closed-loop control for agents engaged in complex physical contact tasks. Finally, in general-purpose scenarios, diffusion models serve as a probabilistic foundation for universal physical laws, through representation learning on massive, heterogeneous video corpora, they implicitly distill cross-domain physical priors and underlying causal logic.

3.1. Datasets

This section delineates the datasets instrumental to the development of diffusion-based world models across three pivotal domains: autonomous driving, embodied intelligence, and general-purpose scenarios, as summarized in Table 1. These benchmarks encompass a heterogeneous range of environments—including urban transit, pedestrian-level street views, simulated gaming sequences, and motion trajectories—thereby establishing a robust empirical foundation for training world models within these specialized fields.

3.1.1. Dataset Classification

Autonomous Driving. Large-scale perception datasets, such as nuScenes[41] and the Waymo Open Dataset[43,44], have laid the groundwork for driving generation by integrating multimodal sensor suites with HD Maps. Building upon these, nuScenes-Occupancy[41] introduces fine-grained 3D occupancy grid annotations, facilitating a deeper structural understanding of spatiotemporal evolution in 3D space. For closed-loop evaluation, frameworks like nuPlan[42] and NAVSIM[46] synergize real-world trajectories with kinematic signals to support policy-driven planning simulations. Furthermore, OpenDV-YouTube[80] leverages web-scale video data to provide unprecedented diversity and scale, significantly enhancing the generalizability of pre-trained models.
Embodied Intelligence. In the realm of robotic manipulation, cross-platform datasets such as Open-X Embodiment[57] and RT-1[56] aggregate millions of trajectories to define universal action spaces. Conversely, CALVIN[81] and Language Table[58] focus on long-horizon logical reasoning by conditioning actions on linguistic instructions. To capture human-centric interaction priors, Ego4D[55] and EgoDex[65] utilize first-person video streams to imbue agents with common-sense physical knowledge. Additionally, Bridge Data[59] provides dense, domain-specific trajectories that facilitate efficient fine-tuning and sim-to-real transfer on physical robotic systems.
General-Purpose Scenarios. Interactive virtual environments, including Atari[82], Minecraft, and CSGO, serve as the bedrock for research into interactive generation due to their high dynamics and complex underlying logic in open-world sequences. MiniGrid[78] emphasizes the assessment of agent generalization through procedurally generated layouts. Regarding visual dynamics, WebVid-10M[83] and SSv2[70] reinforce the modeling of physical causality through dense text-video alignment. Concurrently, datasets such as UCF-101[67], MSR-VTT[69], and the animated FlintstonesHD[77] combine short-term action recognition with semantic captions, supporting language-conditioned video modeling and synthetic content generation.

3.1.2. Limitation

While the aforementioned datasets collectively delineate a hierarchical cognitive roadmap—spanning from virtual simulations to real-world dynamic interactions—several fundamental bottlenecks persist that impede the realization of robust, generalized world models.
Missing Action Tags. Despite providing vast visual streams, benchmarks such as UCF101[67], MSR-VTT[69], and SSv2[70] are inherently "observational" rather than "interventional." The lack of explicit action labels constrains models to merely fitting pixel-level statistical transition regularities. Devoid of an action-effect causal grounding, models are susceptible to statistical hallucinations; they may generate visually coherent sequences that lack an understanding of the underlying forces driving dynamical evolution.
Discrete and Continuous. Existing datasets exhibit starkly divergent physical priors. For instance, Minecraft operates on discrete, voxel-based grid dynamics where causality is largely symbolic and rule-oriented. Conversely, autonomous driving datasets like Waymo Open Motion[43] and nuScenes[41] are predicated on continuous manifolds and rigid-body dynamics. This profound ontological gap between "discrete block-logic" and "continuous spatiotemporal fields" renders physical common sense—such as collision detection and inertial evolution—difficult to reconcile or transfer across domains via current feature-alignment methodologies.
Fragmentation of Domain-Specific Physical Knowledge. Current robotic datasets (e.g., RT-1[56], BridgeData V2[62], Open X-Embodiment[57]) focus primarily on tabletop manipulation, whereas navigation datasets like TartanDrive[60] or SCAND[54] emphasize terrain interaction and large-scale spatial attributes. In the absence of a universal physical representation protocol, physical laws remain sequestered within specific domains. This fragmentation prevents the effective distillation of logical reasoning (learned in environments like Procgen[74]) into robust spatiotemporal predictive capabilities required for complex traffic scenarios in Argoverse2[39].
The Risks of Overusing Synthetic Data. To overcome the challenge of scarce real-world samples in long-tail scenarios, current research increasingly favors generative probabilistic modeling techniques such as diffusion models to synthesize virtual data. While these models can accurately fit visual statistical distributions and provide excellent sample diversity, their limitations remain significant when constructing rigorous "world models." Existing generative algorithms often struggle to capture the details of entropy increase in irreversible physical processes (such as the evolutionary mechanism of material fragmentation at the moment of collision) and tend to smooth out high-frequency noise and physical anomalies in the real world. However, these anomalies are precisely the key to improving the perceptual robustness of world models.[84] Furthermore, since synthetic data inherently lacks deep modeling of underlying physical consistency, if the iteration of the world model relies excessively on such virtual samples, the model may fall into a self-evolving logical loop, leading to "model collapse."

3.2. Diffusion-Based World Model for Autonomous Driving

Diffusion-based world models are catalyzing a fundamental paradigm shift in autonomous driving, transitioning from passive perception toward interactive generative simulation. Their core utility lies in their superior capacity for high-dimensional distribution modeling, which effectively relieves the representation bottlenecks encountered by traditional predictive models when navigating multimodal uncertainties and sparse, long-tail edge cases. By formulating complex traffic dynamics as a conditional diffusion process, these frameworks not only achieve breakthroughs in multiview spatial consistency but also establish a temporal evolution logic grounded in physical intuition[85]. For example, cutting-edge research such as Drive-WM[86], DriveDreamer-2[87], and Vista[31] integrate map topology, agent decision-making, and sensor data to place autonomous driving systems in an interactive and predictable "generative simulator," providing crucial prior support for achieving end-to-end closed-loop planning.
To systematically abstract how the diffusion model reconstructs the future state prediction and high-fidelity scene generation capabilities of the autonomous driving world model, this section defines the state evolution of the world model as Eq.8:
z t + 1 = W ϕ z t , a t , D θ ( ϵ ∣ z t , a t , { C p h y , C g e o , C b e h } )
where z t + 1 represents the predicted latent world features for the subsequent timestep. The operator W ϕ denotes the world model backbone, responsible for capturing deterministic linear evolution and inertial regularities within the traffic flow. In contrast, D θ serves as the diffusion operator; it performs denoising and reconstruction of the noise vector ϵ under the guidance of multidimensional constraint variables C , thereby enabling probabilistic modeling of the multimodal uncertainties inherent in future scenes. The action vector a t encapsulates the ego-vehicle’s driving maneuvers, enhancing the system’s controllability and facilitating counterfactual reasoning. Finally, the constraint set C is decomposed into three pivotal components: C p h y (physical constraints) to mitigate visual artifacts and violations of physical laws; C g e o (geometric constraints) to ensure topological consistency across heterogeneous views; and C b e h (behavioral/causal constraints) to enable responsive environmental feedback conditioned on diverse driving instructions.
Figure 5. Overview of DWMs for autonomous driving. By sensing the surrounding environment, it performs dynamic path planning.
Figure 5. Overview of DWMs for autonomous driving. By sensing the surrounding environment, it performs dynamic path planning.
Preprints 233002 g005
While the proposed mathematical formulation formalizes the structural integration of diffusion models within the world modeling framework, the efficacy of high-fidelity environmental synthesis remains fundamentally contingent upon the architectural design of constraint operators. Therefore, this section divides existing research into three core evolutionary directions based on the constraint mechanisms imposed by the model during the generation process. Physical and spatiotemporal constraints aim to ensure the dynamic consistency and physical rationality of the generated sequences over long time spans; geometric and multidimensional perception alignment focuses on establishing a fine mapping relationship between high-dimensional generation distributions and heterogeneous sensor data and map topology; while behavioral conditional and causal reasoning generation strives to construct world models with closed-loop interaction capabilities, endowing the system with counterfactual reasoning capabilities for long-tail scenarios by causally decoupling vehicle decisions from environmental evolution.
Physical and Spatiotemporal Constraints. In the domain of autonomous driving, characterized by its stringent emphasis on determinism and safety-criticality, the fundamental challenge of such world models resides in ensuring that synthesized sequences do not diverge from physical laws across extended temporal horizons. Formally, research in this trajectory focuses on conditioning the diffusion operator D θ through explicit geometric priors and temporal memory. To mitigate the spatial misalignment issues prevalent in multiview generation, researchers have integrated explicit 3D structural constraints within the diffusion latent space. For instance, DiST-4D[88] ensures spatial consistency and temporal extrapolation stability in multiview video generation by decoupling spatiotemporal diffusion processes and introducing metric depth as a core geometric representation without requiring per-scene optimization. WoVoGen[89] advances this further by employing 4D World Volumes as conditioning factors; by utilizing voxelized occupancy networks as inputs, it integrates 3D geometric priors into the denoising process to synthesize highly coherent multi-camera videos, effectively resolving object vanishing or deformation across camera perspectives. Addressing the predictive collapse induced by error accumulation in long-term forecasting, Vista[31] proposes a latent replacement strategy and introduces a motion instance loss to penalize structural distortion during the diffusion process. Orbis[90] optimizes continuous flow matching and long-term prediction modules, significantly enhancing forecasting accuracy in complex urban scenarios while maintaining inference efficiency. Epona[91] constructs an autoregressive diffusion framework with a decoupled trajectory-visual dual-stream DiT architecture, allowing each denoising step to align precisely with preceding states, thereby extending generative stability to the minute level. Furthermore, DriveDreamer-2[87] leverages LLMs to augment the semantic perception of diffusion models, combining structured layouts with LLM-derived driving semantics to provide controlled guidance for the denoising process. This approach not only improves the simulation fidelity of dynamic agent behaviors but also ensures logical coherence in extreme weather or rare long-tail scenarios.
Despite these advancements in visual coherence, such models continue to face three critical challenges. First, the majority of current frameworks remain "visual approximations" rather than true "dynamical simulations." Their deterministic transition backbones are typically fitted via black-box neural networks lacking explicit classical mechanical constraints—such as mass or friction—rendering the generated physical evolution unconvincing under extreme maneuvering conditions. Second, notwithstanding the extended horizons achieved by Epona[91] and Orbis[90], the causal chain between ego-actions and environmental response tends to rupture in multiagent game-theoretic scenarios as the prediction horizon expands; consequently, models often produce "plausible pixel motion" while losing the causal determinism of action-consequence relationships. Finally, the iterative nature of diffusion models entails prohibitive computational overhead. Although Orbis[90] has optimized efficiency, achieving millisecond-level generative latency while preserving physical consistency remains the primary bottleneck for deploying these models in real-time onboard systems for end-to-end closed-loop testing.
Geometric and Multidimensional Perception Alignment. In the architectural formulation of diffusion-based world models, ensuring that synthesized pixel distributions strictly adhere to rigorous 3D spatial logic is paramount for achieving high-fidelity simulation. From a mathematical perspective, the crux of this technical trajectory lies in the profound deconstruction and reformulation of conditioning variables within the diffusion refinement operator. By formalizing C g e o as a multidimensional representation set—encompassing 3D bounding boxes, High-Definition (HD) maps, and camera intrinsic/extrinsic parameters—the diffusion process is guided toward constrained sampling within the latent space. For instance, MagicDrive[92] systematically investigated the imposition of 3D geometric controls during the diffusion denoising process for the first time. Through specifically engineered encoding strategies, it transforms camera parameters and 3D bounding boxes into conditional embeddings for the diffusion operator and introduces cross-view attention mechanisms, enabling the model to share geometric priors across disparate viewpoints and ensuring the strict spatial alignment of surround-view imagery. Building upon this, MagicDrive-V2[93] further augmented the generative resolution and long-term coherence of the world model. By adopting an MVDiT architecture, it replaces conventional convolutional denoising backbones with Transformers possessing superior global modeling capacities, and coupled with adaptive control mechanisms, the diffusion process responds more precisely to dynamic geometric constraints, significantly enhancing geometric realism in complex intersection scenarios. Panacea[94] focuses on the consistency of panoramic synthesis by embedding a panoramic consistency module within the diffusion framework, utilizing Bird’s-Eye-View (BEV) sequences as the primary driver to ensure that pixel distributions across overlapping camera regions remain statistically continuous and coherent during the generation of surround-view video.
To elevate diffusion-based world models from mere "image synthesis" to comprehensive "world-state representations," researchers have begun exploring occupancy grids or 3D voxels as intermediate representations to achieve physical alignment across multimodal data. For example, UniScene[95] proposed a hierarchical modeling paradigm centered on occupancy networks. During prediction, it first utilizes a diffusion model to generate consistent semantic occupancy grids, which then serve as core physical constraints to synchronously guide the diffusion generation of multiview video and LiDAR point clouds. This methodology ensures a profound unification of heterogeneous sensor data at the underlying physical representation layer, demonstrating the remarkable potential of diffusion models in joint multimodal modeling. LiDARCrafter[96] specifically targets the 4D modeling of LiDAR sequences by designing a triple-branch diffusion network. By independently processing object structures, motion trajectories, and spatial geometry during the denoising process, it ensures that synthesized LiDAR sequences maintain both temporal coherence and precise alignment with the environment’s topological structure, thereby providing high-value non-visual modal priors for autonomous driving systems.
Despite the substantial progress achieved in spatial alignment, several pressing challenges remain. First, as the dimensionality of conditioning variables escalates, the conditional branches of the diffusion operator exhibit prohibitive complexity. Maintaining model lightweighting without compromising the sensitivity to intricate geometric constraints remains a critical bottleneck. Second, rigid constraint mechanisms often precipitate a diminution in the diversity of synthesized samples. Retaining the distributional manifold of the diffusion model in representing long-tail scenarios—such as extreme precipitation, fog, or abrupt illumination transitions—under stringent geometric alignment remains a formidable challenge for robust world modeling. Finally, while existing multimodal world models perform adequately in static alignment, they struggle with the cross-modal synchronization of highly dynamic entities (e.g., the instantaneous correspondence of high-speed vehicles between imagery and LiDAR). Due to the inherent stochasticity of separate diffusion branches, minute spatiotemporal misalignments persist, highlighting a current deficiency in joint cross-modal noise distribution constraints.
Behavioral Conditional and Causal Reasoning This type of constraint mechanism-driven diffusion world model, by establishing a deep causal coupling between "actions" and "states," propels autonomous driving simulation from simple video synthesis to a paradigm shift toward "digital twin evolution" with closed-loop interactive capabilities. The fundamental value of this paradigm lies in its endowment of the model with the capacity for counterfactual reasoning within complex, high-dimensional spaces, enabling the system to accurately characterize the perturbation logic of ego-vehicle decisions on environmental dynamical evolution. Pioneering research focused on enabling diffusion models to distill driving regularities from massive datasets, with the GAIA[32,97] series representing foundational contributions in this trajectory. Specifically, GAIA-2[97] abstracts world modeling as a high-fidelity sequential modeling task in the latent space. Its pivotal innovation involves utilizing a DiT architecture to jointly encode action sequences, textual descriptors, and visual features. By leveraging self-supervised pre-training on large-scale real-world driving datasets, GAIA-2[97] achieves precise responsiveness of the diffusion refinement operator to complex driving instructions, demonstrating that diffusion-based world models can manifest emergent deep causal cognition regarding traffic regulations and scene evolution. In contrast, Drive-WM[86] proposes a multi-scenario "imagination" pathway based on world models, synthesizing multiple parallel future trajectories by injecting distinct candidate action sequences during the diffusion denoising process. This "prediction-evaluation-decision" mechanism essentially utilizes the diffusion model to simulate counterfactual outcomes under varying interventions, thereby providing critical reward signals for end-to-end closed-loop planning.
Simultaneously, to address action alignment challenges in interactive simulation, DriVerse[98] and Imagine-2-Drive[99] introduce trajectory-prompt augmentation and diffusion policy models, respectively, effectively mitigating predictive drift and policy degradation in long-horizon forecasting. UniFuture[100] and CVD-STORM[101] approach the problem from a unified modeling perspective. Specifically, UniFuture[100] establishes a synchronized evolution logic of "action-geometry-representation" by sharing latent spaces for imagery and depth during the diffusion process; meanwhile, CVD-STORM[101] utilizes a spatiotemporal reconstruction VAE to enhance the 4D consistency of the diffusion operator, ensuring that synthesized environmental topologies strictly adhere to causal laws during complex interactive maneuvers. Nevertheless, maintaining rigorous causal consistency across extended temporal horizons and achieving real-time closed-loop feedback remains a critical challenge and a representation bottleneck for advancing toward high-level autonomous driving. Future research should incorporate causal intervention theory to explore the use of structural causal models (SCMs) for explicitly constraining the generative process. this would enable the model to precisely differentiate between inherent environmental evolution and the chain reactions induced by ego-vehicle actions, ultimately facilitating the construction of world models endowed with genuine physical common sense and logical reasoning capabilities.
Table 2. Feature Classification. AD: Autonomous Driving, EI: Embodied Intelligence, GD: General Domains
Table 2. Feature Classification. AD: Autonomous Driving, EI: Embodied Intelligence, GD: General Domains
Areas Method Publication Code Characteristics
long-horizons Multimodal Interactive Consistency Diverse environments
AD ADriver-I [102] arXiv’23 ✓ ✓ ✓ ✓ ✓
AD MotionDiffuser[103] CVPR’23 ✓ ✓ ✓
AD GAIA-1[32] arXiv’23 ✓ ✓ ✓ ✓ ✓
AD DriveDreamer[104] ECCV’24 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
AD Vista[31] NeurIPS’24 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
AD Drive-WM[86] CVPR’24 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
AD MagicDrive [92] ICLR’24 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
AD Panacea [94] CVPR’24 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
AD SceneDiffuser [105] NeurIPS’24 ✓ ✓ ✓ ✓ ✓
AD WoVoGen [89] ECCV’24 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
AD UrbanWorld [106] arXiv’24 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
AD GAIA-2[97] arXiv’25 ✓ ✓ ✓ ✓ ✓
AD Epona[91] ICCV’25 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
AD DriveDreamer4D[107] CVPR’25 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
AD MagicDrive-V2[93] ICCV’25 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
AD AVD2[108] ICRA’25 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
AD DiST-4D[88] ICCV’25 Preprints 233002 i001 ✓ ✓ ✓ ✓
AD DriVerse [98] ACM MM’25 Preprints 233002 i001 ✓ ✓ ✓ ✓
AD InfiniCube [109] ICCV’25 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
AD MaskGWM [110] CVPR’25 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
AD SceneDiffuser++ [111] CVPR’25 ✓ ✓ ✓ ✓ ✓
AD UniScene [95] CVPR’25 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
AD Imagine-2-Drive [99] IROS’25 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
AD ImagiDrive[112] arXiv’25 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
AD UniFuture [100] ICRA’26 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
AD DriveLaw [113] CVPR’26 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
AD SimScale [114] CVPR’26 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
EI 3D-VLA[115] ICML’24 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
EI AVID[116] arXiv’24 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
EI UniSim[33] ICLR’24 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
EI PhysDreamer[117] ECCV’24 Preprints 233002 i001 ✓ ✓ ✓ ✓
EI RoboDreamer[118] ICML’24 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
EI Cosmos-Transfer1[9] arXiv’25 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
EI DiWA[119] CoRL’25 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
EI DreamGen[120] arXiv’25 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
EI LaDi-WM[121] CoRL’25 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
EI NWM[122] CVPR’25 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
EI Vid2World[123] arXiv’25 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
EI CoT-VLA[124] CVPR’25 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
EI EchoWorld[125] CVPR’25 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
EI GWM[126] ICCV’25 Preprints 233002 i001 ✓ ✓ ✓ ✓
EI MoWM[127] arXiv’25 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
EI PlayerOne[128] arXiv’25 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
EI VISTAv2[129] arXiv’25 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
EI PEVA[130] arXiv’25 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
EI WMPO[131] arXiv’25 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
EI World4RL[132] arXiv’25 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
EI GenEx[133] arXiv’25 ✓ ✓ ✓ ✓ ✓
EI Motion Prompting[134] CVPR’25 Preprints 233002 i001 ✓ ✓ ✓ ✓
EI ABot-PhysWorld[135] CVPR’26 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
EI VideoWorld2[136] CVPR’26 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
GD NUWA-XL[77] arXiv’23 Preprints 233002 i001 ✓ ✓ ✓ ✓
GD DiffDreamer[137] ICCV’23 Preprints 233002 i001 ✓ ✓ ✓ ✓
GD ConsistI2V[138] TMLR’24 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
GD DIAMOND[139] NeurIPS’24 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
GD PEEKABOO[140] CVPR’24 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
GD WorldGPT[141] arXiv’24 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
GD Pandora[142] arXiv’24 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
GD StoryWeaver[143] AAAI’24 Preprints 233002 i001 ✓ ✓ ✓ ✓
GD Sora[11] Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
GD Genie-2[144] Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
GD Genie-3 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
GD GameNGen[145] ICLR’25 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
GD GEM[146] CVPR’25 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
GD LongVie 2[147] arXiv’25 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
GD MorphoSim[148] arXiv’25 Preprints 233002 i001 ✓ ✓ ✓ ✓
GD OmniWorld[149] arXiv’25 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
GD Voyager[150] arXiv’25 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
GD YUME[151] arXiv’25 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
GD ASTRA[152] arXiv’26 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
GD FantasyWorld[153] ICLR’26 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
GD CoECT[152] CVPR’26 Preprints 233002 i001 ✓ ✓ ✓ ✓
GD FreeLOC[154] CVPR’26 Preprints 233002 i001 ✓ ✓ ✓
GD NeoVerse[155] CVPR’26 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
GD ProPhy[156] CVPR’26 Preprints 233002 i001 ✓ ✓ ✓ ✓
GD VerseCrafter[157] CVPR’26 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓
GD WorldForge[158] CVPR’26 Preprints 233002 i001 ✓ ✓ ✓ ✓ ✓

3.3. Embodied Intelligence

Within the field of embodied intelligence, diffusion-based world models are propelling agents from conventional ’policy imitation’ toward a general cognitive stage characterized by the capacity for internal world evolution. Diverging from traditional methodologies reliant on deterministic mappings, diffusion models—by virtue of their superior representational capacity for high-dimensional, nonlinear spatiotemporal distributions and complex multimodal distributions—abstract environmental evolution in embodied interactions as a conditional generative process within the latent space. This paradigm shift endows robotic systems with anthropomorphic ’mental simulation’ capabilities, enabling the internal rehearsal of multiple potential future trajectories prior to decision execution. For instance, pioneering research exemplified by the π series[159][160], Vidar[34], and LaDi-WM[121] leverages the diffusion process to model intricate spatiotemporal dynamics, establishing high-fidelity interactive simulation substrates and formalizing the ’generation-is-planning’ core logic. This not only fundamentally alleviates technical bottlenecks such as unsmooth action sequences and multimodal distribution collapse that are common in embodied tasks, but also provides the theoretical and technical scaffolding for constructing general-purpose embodied foundation models predicated on world models as a dynamical engine.
Delving into its logical core, this paradigm evolution can be characterized as a controlled state reasoning process in the latent space as Eq.9.
z t + 1 = f ϕ ( z t , a t , D θ ( ϵ ∣ c t , P v i d e o ) )
Where z t + 1 denotes the predicted world state for the subsequent timestep, and f ϕ represents the backbone dynamics function of the world model, responsible for integrating the current state, action, and synthesized stochastic evolution terms to ensure spatiotemporal consistency. z t is the current perceptual state, a t denotes the current embodied action input, and c t signifies the conditional embedding, encompassing global objectives, environmental constraints, or cross-modal priors, while P v i d e o represents the video prior, and D θ denotes the diffusion generative operator, tasked with capturing the nondeterministic and multimodal distributional components of environmental evolution. Serving as a stochastic injection source, this operator enables the world model to synthesize high-fidelity future prognostications that transcend mere linear extrapolations.
Figure 6. Overview of DWMs for embodied intelligence. Embodied agents achieve a closed–loop architecture of "decision–planning–control" through world models, unifying dynamic modeling of three–dimensional environments, prediction of interactive properties, and execution of actions.
Figure 6. Overview of DWMs for embodied intelligence. Embodied agents achieve a closed–loop architecture of "decision–planning–control" through world models, unifying dynamic modeling of three–dimensional environments, prediction of interactive properties, and execution of actions.
Preprints 233002 g006
Within the conceptual framework of this formulation, recent advancements have systematically propelled the technological evolution of diffusion-based world models in the domain of embodied intelligence through the mathematical refinement of the diffusion operator, the structural reconfiguration of the state space, and the augmentation of conditional embeddings. Regarding the enhancement of generative efficiency and behavioral generalization, research exemplified by the π series has established foundational milestones for embodied foundation models. Specifically, π 0 [159] addresses the computational bottleneck of prolonged inference latency inherent in conventional diffusion models by innovatively integrating flow-matching techniques into the VLA architecture, thereby reconfiguring the diffusion process via a deterministic vector field to facilitate cross-morphological control across massive heterogeneous datasets. This modification essentially optimizes the transport trajectory of the diffusion operator, validating the superior generalization capabilities of the diffusion paradigm in synthesizing continuous behavioral trajectories. Subsequently, π 0.5 [160] leverages synergistic co-training on vast heterogeneous data to facilitate the resolution of long-horizon dexterous manipulation tasks—such as kitchen organization and garment folding—significantly elevating the task success rate of agents in unstructured, open-world environments. Furthermore, DiWA[119] utilizes a pre-trained world model dynamics backbone to fine-tune diffusion-based robotic primitives on offline data via reinforcement learning, leveraging internal world model evolution to ameliorate the suboptimal sample efficiency and safety constraints typically encountered by diffusion policies during physical deployment.
Regarding the construction of world models as simulation substrates, the research nexus centers on leveraging large-scale video datasets to augment the generative fidelity of diffusion factors. For instance, UniSim[33] constructs a universal interactive simulator capable of responding to high-level semantic directives and low-level motor primitives by curating internet-scale video and multimodal robotic corpora. Vidar[34] advances this paradigm further by integrating large-scale video diffusion priors with Masked Inverse Dynamics Models (MIDM). specifically, it utilizes diffusion-synthesized future visual streams to deduce explicit physical constraints of the environment, thereby facilitating precise action forecasting in novel environments with minimal demonstrations and bypassing the traditional over-reliance on embodiment-specific data. To mitigate the challenge of physical hallucinations, the research focus has shifted toward the alignment of diffusion processes within the latent manifold. LaDi-WM[121] innovatively eschews pixel-space synthesis in favor of performing diffusion prediction within an aligned latent space that integrates geometric descriptors from DINO and semantic features from CLIP. This architectural reconfiguration of the state space enables the diffusion operator to circumvent pixel-level noise and concentrate on modeling the evolutionary regularities of high-dimensional features, significantly enhancing the physical veridicality of long-horizon predictions. Concurrently, NWM[122] transforms this predictive capacity directly into a decision operator, demonstrating the robust potential for closed-loop navigation and the ’generation-as-planning’ paradigm in unstructured environments by simulating multiple trajectories within the latent space and evaluating their goal-directed efficacy.
Advantages and Limitations. In contrast to conventional deterministic dynamical models, the primary merit of the diffusion paradigm resides in its superior representational capacity for multimodal distributions and high-dimensional continuous action spaces. By decomposing the synthesis of high-dimensional continuous actions into a sequential denoising trajectory, the diffusion process effectively ameliorates prevalent bottlenecks in embodied tasks, such as "suboptimal action smoothness" and ’mode collapse across multitask distributions,’ thereby facilitating the acquisition of robust feature representations from massive heterogeneous datasets. Furthermore, the potential of diffusion models to serve as universal interactive simulation substrates is increasingly being elucidated. By synergizing diffusion priors with inverse dynamics or latent space alignment methodologies, these models can not only synthesize future visual trajectories consistent with physical principles but also function as a "world surrogate" providing zero-shot closed-loop supervision for downstream policies, as exemplified by Vidar[34], LaDi-WM[121], and NWM[122]. This "video-as-planning" trajectory surmounts the traditional overreliance of embodied intelligence on labor-intensive manual annotations.
Nevertheless, establishing real-time closed-loop control via diffusion-based world models within the embodied domain remains a formidable challenge. The foremost obstacle is the inherent tension between inference latency and real-time operational requirements. While embodied agents typically necessitate millisecond-level decision feedback, the fundamental mechanism of diffusion models involves denoising via a T-step Markov chain, rendering their generative throughput significantly inferior to the environment’s dynamic rate of change.The secondary concern involves the absence of physical consistency. Although diffusion models yield visually compelling video streams, they remain essentially stochastic approximations of pixel or feature distributions, frequently manifesting content that violates first-order physical principles due to the lack of explicit dynamic grounding. Such "hallucinations," when utilized as a world model to guide robotic planning, may precipitate erroneous or hazardous maneuvers, which proves catastrophic in embodied tasks requiring high-precision manipulation.

3.4. General Domains

In general-purpose domains, diffusion-based world models demonstrate a conspicuous trajectory of transitioning from "high-dimensional pixel synthesis" toward "generative physical simulation." By leveraging the representational capacity of diffusion processes for complex stochastic distributions, these models—absent explicit physical formalizations—manifest emergent preliminary simulation capabilities regarding spatiotemporal coherence, object constancy, and fundamental causal regularities through unsupervised learning on large-scale corpora. This paradigm not only exhibits extraordinary efficacy in large-scale visual simulations such as Sora[11] and Genie-2[144], but has also attained substantive advancements in closed-loop interaction and localized controllable synthesis via frameworks like DIAMOND[139] and PEEKABOO[140]. Concurrently, NUWA-XL[77], DiffDreamer[137], and StateSpaceDiffuser[161], respectively from the dimensions of temporal span, scene scale, geometric constraints, and computational efficiency, collectively expand the technical boundaries of diffusion models, establishing them as a critical nexus bridging digital twin environments and the perceptual grounding of the physical world.
To systematically deconstruct the underlying technical evolutionary logic of diffusion-based world models within general-purpose domains, this section formalizes the simulation mechanics of such models regarding the physical world into the following macroscopic spatiotemporal evolution equation:
z t + 1 = W ϕ ( Z < t + 1 , τ , D θ ( ϵ ∣ C ) )
where z t + 1 denotes the predicted future state representation; W ϕ represents the backbone inference framework of the world model, primarily responsible for the integration of spatiotemporal context, interactive directives; the diffusion generative operator, Z < t + 1 signifies the long-range spatiotemporal representation, encompassing all historical information from the initial onset to the current timestamp; τ denotes the interaction operator, comprising textual guidance, visual conditioning, and action instructions; D θ serves as the diffusion refinement operator, representing a diffusion generative architecture parameterized by θ , whose functional role is to render "logical predictions" into "physically veridical landscapes"; ϵ represents the stochastic noise; C encapsulates the physical and geometric constraints.
Figure 7. Overview of DWMs for general domains. DWMs evolve from pixel–level synthesis to generative physical simulation, indicating a potential pathway toward AGI.
Figure 7. Overview of DWMs for general domains. DWMs evolve from pixel–level synthesis to generative physical simulation, indicating a potential pathway toward AGI.
Preprints 233002 g007
This equation deconstructs the simulation process of intricate world models into a synergistic operation between backbone inference and diffusion refinement; it not only aligns with contemporary mainstream architectural designs but also delineates the complete trajectory from "comprehension" to "imagination" for general-purpose world models. Predicated on this mathematical framework, this section establishes a taxonomical framework across three complementary dimensions: representational capacity, constraint mechanisms, and interaction paradigms. First, it explores "Long-range Spatiotemporal Representation and Architecture Expansion," the core of which resides in leveraging diffusion processes to model high-dimensional manifolds to ensure the consistency of high-dimensional pixel manifolds across ultra-long sequences. Second, it investigates "Physics rule and geometric constraints," dissecting the methodologies for injecting the structural invariance of the objective world and kinematic priors into generative distributions to facilitate embodied physical simulation. Finally, it analyzes "Interaction Dynamics and Closed-loop Evolution," examining how diffusion models respond conditionally to discrete or continuous action directives, thereby realizing causal inference and interactive feedback within complex environments.
Long-Range Spatiotemporal Representation and Architecture Expansion. In the context of general dynamical simulation, constructing high-fidelity latent manifolds that span extended temporal horizons remains a fundamental challenge for diffusion-based world models. The efficacy of long-range simulation is contingent upon the long-term memory capacity of the backbone architecture and the spatiotemporal generative stability of the diffusion refinement operator.
Research exemplified by Sora[11] demonstrates that applying DiT architectures to spatiotemporal latent spaces yields significant scaling law effects, imbuing models with preliminary capabilities for complex physical interaction modeling. However, the quadratic complexity of traditional self-attention mechanisms poses a substantial computational bottleneck as sequence lengths escalate. To address this, Po et al.[162] introduced State Space Models (SSMs) with linear complexity. By employing a block-wise SSM scanning mechanism, this work significantly expands context perception boundaries while maintaining the representational power of diffusion models, enabling consistent predictions for complex long-range tasks with reduced inference overhead.
Addressing the challenge of semantic drift in video generation, NUWA-XL[77] proposes a hierarchical parallel architecture termed "Diffusion over Diffusion." This framework decomposes the diffusion process into global keyframe diffusion and local interpolation diffusion, where a global model first synthesizes a macroscopic logical skeleton, followed by recursive content filling by local models. This hierarchical methodology ameliorates the serial bottlenecks associated with traditional autoregressive generation, facilitating ultra-long video synthesis. Concurrently, LongVie2[147] adopts a degradation-aware autoregressive training strategy; by integrating dense and sparse control signals, it effectively mitigates iterative error accumulation during continuous generation—spanning three to five minutes—ensuring that dynamical evolution remains anchored in physical reality.
Maintaining cross-frame entity constancy is imperative for complex narratives and multidimensional spatial modeling. StoryWeaver[143] reinforces identity invariance in multi-character interaction scenarios through a global semantic alignment mechanism. For three-dimensional consistency, ConsistI2V[138] introduces spatiotemporal alignment and layout guidance within the image-to-video diffusion process, enhancing the anchoring effect of the initial frame on subsequent state evolution. Furthermore, OmniWorld[149] leverages geometric supervision from multimodal and multi-domain datasets to inject authentic camera motion parameters and depth constraints, enabling diffusion world models to evolve from pure pixel synthesis into 4D space-aware dynamical simulators.
While these architectural extensions have made significant strides in enhancing visual coherence, future research must facilitate a paradigm shift from "visual continuity" to "physical veridicality." First, although current structural expansions maintain pixel-level smoothness, they still encounter inevitable semantic entropy increase during ultra-long-term simulations. Future trajectories should move beyond mere feature stacking and explore the explicit embedding of anchors governed by causal conservation laws within the diffusion latent space to ensure the stability of physical attributes following complex interactions. Second, regarding the inherent non-uniformity of real-world physical evolution—characterized by low-frequency stability in backgrounds versus high-frequency nonlinearity at interaction centers—the development of adaptive spatiotemporal sampling frameworks will be pivotal. By introducing hierarchical perception mechanisms, world models can achieve high-precision simulation of local microscopic collisions while maintaining global macroscopic consistency, thereby overcoming precision bottlenecks in general-domain simulation under constrained computational resources.
Physics Rule and Geometric Constraints. In the research lineage of diffusion-based world models, the induction of physical laws and geometric constraints demarcates a paradigmatic evolution from "high-dimensional probabilistic distribution fitting" toward "structured dynamical simulation." The pivotal challenge resides in ameliorating physical hallucinations and spatial topological collapse resulting from the deficiency of physical priors in the long-horizon predictions of diffusion models. This section aims to investigate the construction of sophisticated constraint mechanisms to ensure that the diffusion refinement operator, while rendering deterministic logical skeletons into physical reality, adheres to the geometric invariance and conservation laws of the objective world.
At the level of geometric consistency, addressing the content drift prevalent in purely data-driven models such as Sora[11] during viewpoint transformations, contemporary research ensures object constancy by injecting 3D structural priors into the diffusion process. For instance, DiffDreamer[137] proposes a conditional diffusion architecture for monocular scene extrapolation tasks, embedding 3D conditioning within the diffusion process to enable the model to maintain stringent spatiotemporal stability during scene completion and extrapolation. Its essence involves projecting input 2D imagery into 3D space and utilizing the diffusion operator for iterative refinement upon geometrically aligned representations, effectively mitigating the content drift associated with conventional video generation under camera ego-motion. The Martian World Model[163] demonstrates that in extremely sparse data regimes, explicit geometric reconstruction and depth map constraints serve as the bedrock for maintaining the fidelity of diffusion-based synthesis. By utilizing geometric parameters derived from 3D reconstruction as robust constraint signals within the diffusion process, it ensures that synthesized dynamic videos exhibit physical measurability.
Beyond static geometric constraints, world models must comprehend material properties, deformations, and complex nonrigid physical interactions. PhyWorld[164] introduces a physics-inspired synthesis framework that leverages structured data generated by physics-based simulators to guide the diffusion model. Within this architecture, the diffusion model transcends mere learning of interpixel correlations. Instead, by training on simulator-generated trajectories, it learns to simulate the stress and strain responses of deformable bodies under conditional guidance. Furthermore, MorphoSim[148] achieves precise manipulation of 4D scene attributes during the diffusion process via feature field distillation techniques. By integrating linguistic instructions and trajectory guidance into the constraint terms, it endows the diffusion world model with the capability to edit object position, color, and morphology during interaction, significantly enhancing the physical interpretability and interaction precision of the generative space.
Although breakthroughs have been achieved in geometric alignment and short-range simulation, realizing a true generative physics engine necessitates addressing fundamental contradictions. First, contemporary models frequently oscillate between "diminished diversity induced by rigid geometric constraints" and "physical collapse resulting from weak constraints." A prospective breakthrough involves the development of adaptive constraint-weight frameworks, enabling models to dynamically adjust the injection intensity of geometric priors based on the physical entropy of the scene—maintaining rigid geometric conservation in static backgrounds while permitting plausible topological deformations during complex interactions. Second, existing constraints are predominantly predicated on visually observable geometric features, whereas real-world evolution is governed by latent physical quantities such as mass, friction, and damping. Future research should focus on latent physical quantity mining—specifically, inverting and characterizing the intrinsic dynamical parameters of objects solely from visual diffusion processes without reliance on explicit simulators. This paradigmatic shift from "simulating phenomena" toward "comprehending first principles" will serve as the decisive hallmark of whether diffusion world models can supersede traditional physics engines in general-purpose domains.
Interaction Dynamics and Closed-Loop Evolution. Interaction dynamics characterize the core capacity of diffusion-based world models to facilitate closed-loop evolution, marking a fundamental qualitative transformation from pure "video synthesis" toward "interactive simulation." In macroscopic dynamic equations, interaction directives—such as peripheral inputs, camera trajectories, or semantic instructions—act directly upon the backbone architecture of the world model, while the diffusion refinement operator, governed by the synergy of stochastic noise and physical constraints, ensures the visual veridicality of the interactive feedback.
The capacity of diffusion models to represent complex stochastic distributions enables the world model to achieve fine-grained dynamic evolution within continuous latent manifolds through conditional responses to interaction directives. For instance, DIAMOND[139] demonstrates that within Atari environments, the high-frequency visual fidelity captured by the diffusion world model directly dictates the cognitive upper bound of the agent regarding action consequences, effectively ameliorating the bottlenecks of detail attrition inherent in traditional discrete latent space models. Even more groundbreaking, GameNGen[145] achieves real-time simulation of the complex 3D interactive environment DOOM at 20 FPS on a single commodity chip for the first time by introducing conditional augmentation of historical observations and actions within the diffusion process, thereby validating the feasibility of diffusion models as neural game engines.
To enhance interaction precision, researchers have explored the introduction of granular manipulation mechanisms within the diffusion pipeline. PEEKABOO[140] leverages the internal attention mechanisms of diffusion models to achieve zero-shot interactive control via spatiotemporal masking. Absent additional training, it permits users to inject positional and trajectory information into physical constraints through control masks, inducing the diffusion model to synthesize semantically coherent interactive behaviors within localized regions. Building upon this, Yume[151] develops the Masked Video Diffusion Transformer (MVDT), which captures long-horizon background consistency through memory modules and supports the real-time triggering of diffusion manifold evolution via peripheral directives, constructing an infinitely extendable interactive world space.
Robust interaction must be predicated on stable spatial representations to prevent causal drift during autonomous exploration. Voyager[150] achieves long-range consistent scene generation by coupling camera motion parameters within the diffusion architecture; furthermore, by jointly predicting RGB and depth information, it augments the model’s 3D geometric perception, ensuring that interaction outcomes maintain physical constancy. FantasyWorld[153] introduces an explicit geometric branch to synergistically model video latent variables and implicit 3D fields in a single forward pass, utilizing geometric constraints to guide the diffusion process and providing measurable interactive feedback for the navigation and task reasoning of embodied agents. Additionally, WorldGPT[141] utilizes LLMs to enhance the coherence of action instructions in long-range tasks, optimizing the causal inference capabilities of diffusion world models via semantically augmented prompting mechanisms.
Despite recent breakthroughs in closed-loop simulation within specific domains, realizing truly universal interactive simulators necessitates addressing two primary bottlenecks. First, contemporary models suffer from a "visual inertia" phenomenon, wherein the model exhibits a bias toward continuing historical visual streams while marginalizing abrupt action interventions. Future research must incorporate causal intervention analysis to resolve causal drift in long-range simulation by decoupling inherent environmental evolution from agent-induced perturbations. Second, existing diffusion world models are largely constrained by synchronous sampling at fixed frame rates, rendering it difficult to balance the computational requirements of low-frequency semantic decision-making and high-frequency physical collisions. Future research should prioritize the development of asynchronous diffusion evolution architectures that permit the model to dynamically modulate sampling frequencies according to interaction intensity—maintaining low-power semantic progression during quiescence and triggering localized high-frequency diffusion reconstruction during impulsive collisions or high-precision manipulations. This "dual-process" simulation paradigm, characterized by fast and slow thinking, represents the requisite trajectory for diffusion world models toward large-scale, real-time physical simulation systems.

4. Diffusion-Based World Model Characteristics

Subsequent to a systematic exposition of the taxonomic framework and open-source dataset for diffusion-based world models, this chapter conducts an in-depth exploration of the core technical attributes manifested by this paradigm in modeling the physical world. Contrasted with traditional architectures predicated on VAEs or purely Autoregressive frameworks, diffusion-based world models not only facilitate a qualitative leap in visual fidelity but also exhibit pronounced advantages in characterizing complex environmental regularities. This chapter dissects the mechanisms through which diffusion-based world models construct high-fidelity, physically consistent universal simulation substrates across five dimensions: Long-Term Evolution, Multimodal Fusion, Interactive Dynamics, Spatiotemporal Coherence, and Multi-environment.

4.1. Long-Term Evolution

Long-term evolution capability serves as a pivotal metric for evaluating the proficiency of a world model in simulating the dynamical logic of the physical world. A robust world model should transcend restricted short-range visual extrapolation, necessitating a capacity for long-term forecasting that adheres to environmental dynamical regularities, thereby ensuring that synthesized sequences maintain high alignment with real-world temporal progression in terms of object kinematics, illumination transitions, and logical inference. In the context of world modeling, long-term temporal characteristics represent not merely quantitative parameters for extending generative duration, but also the fundamental requisite for surmounting recursive drift and establishing causal closure. Within the closed-loop simulations of embodied intelligence or autonomous driving, tasks typically span tens of seconds to several minutes, where any infinitesimal single-step predictive bias may accumulate exponentially over time, precipitating the collapse of environmental logic.
Figure 8. Two examples illustrating the long–horizon prediction behaviors of DWMs in embodied scenarios. The red–marked case exhibits two failure modes: (i) hallucinated object generation, where the number of bananas increases from one at 2 s to two at 4 s; and (ii) temporal inconsistency, where the banana moves before the robotic arm executes the corresponding action. In contrast, the green–marked case demonstrates correct and temporally coherent behavior.
Figure 8. Two examples illustrating the long–horizon prediction behaviors of DWMs in embodied scenarios. The red–marked case exhibits two failure modes: (i) hallucinated object generation, where the number of bananas increases from one at 2 s to two at 4 s; and (ii) temporal inconsistency, where the banana moves before the robotic arm executes the corresponding action. In contrast, the green–marked case demonstrates correct and temporally coherent behavior.
Preprints 233002 g008
Given the susceptibility of traditional autoregressive models to single-step error accumulation, NUWA-XL[77] employs nonlinear processing of historical frames. Essentially, this evolves the diffusion process into a hierarchical "coarse-to-fine" parallel diffusion paradigm. While this strategy effectively ameliorates semantic rupture in long-range generation, thus facilitating long-term temporal synthesis, it remains deficient in theoretical verification regarding the maintenance of cross-hierarchical consistency. Conversely, Vista[31] prioritizes the enhancement of geometric constraints within the operator. It injects historical frames as robust priors via a latent variable replacement mechanism, supplemented by a triangular classifier-free guidance scheme[165] to dynamically modulate the denoising trajectory. This is analogous to imposing manifold calibration on predictive outcomes through geometric constraints at each denoising step, thereby suppressing perceptual degradation in closed-loop simulations exceeding 15 seconds. Epona[91] and LongScape[166] balance efficiency and stability through the decoupled architectural design of their respective diffusion models. Epona[91] utilizes a chain-forward training paradigm coupling autoregression with diffusion—a design that mitigates the issue of attention dispersion encountered by monolithic diffusion architectures when processing ultra-long contexts. In contrast, LongScape[166] introduces a context-aware Mixture-of-Experts (MoE) model, adaptively activating expert modules tailored to distinct action segments. This design essentially modulates the expressive capacity of the diffusion operator based on the semantic complexity of action directives, thereby suppressing context drift.
Diffusion models, through iterative denoising processes, are capable of capturing highly intricate environmental dynamical features and exhibit superior robustness when handling multimodal uncertainties across long temporal horizons. For instance, GAIA-1[32] and Drive-WM[86] demonstrate that in autonomous driving scenarios, diffusion models can maintain high spatial consistency across camera views and action-conditioned consistency spanning several seconds, with synthesized image fidelity significantly outperforming Prediction models based on discrete tokens. UniSim[33], trained on large-scale video datasets, exhibits the capacity to maintain physical law consistency during long-term evolution, while Sora[11], leveraging DiT, further proves that large-scale parameterization can achieve ultra-long video generation with 3D spatial consistency, effectively ameliorating prevalent issues such as object flickering and distortion.
Despite the exceptional representational power demonstrated by diffusion models in single-frame synthesis, they still confront a tripartite challenge in long-range autoregressive evolution: spatiotemporal consistency decay, a deficiency in geometric constraints, and prohibitive inference latency. First, compound errors in recursive prediction lead to severe semantic drift[32,86]. Even when single-step predictive deviations are marginal, the recursive accumulation of such biases—in the absence of explicit external calibration—induces the sampling trajectory to rapidly deviate from the authentic physical manifold. Although NUWA-XL[77] attempts to integrate hierarchical diffusion architectures to fortify global spatiotemporal constraints, the maintenance of cross-hierarchical semantic homeostasis during ultra-long sequence synthesis remains devoid of rigorous theoretical guarantees. Second, purely diffusion-based architectures struggle to maintain geometric constancy in the absence of explicit state-space constraints. As noted in LongScape[166], since models typically perform statistical fitting within pixel or latent spaces rather than explicit modeling of 3D physical structures, long-range reconstruction is prone to losing fine-grained geometric information from the initial state, resulting in object flickering or topological deformation. Finally, the inherent iterative sampling mechanism of diffusion models results in severe computational complexity and real-time bottlenecks. The escalation of sequence length not only imposes a linear GPU memory load but also augments inference latency due to the multistep denoising iterations. While DIAMOND[139] attempts to embed diffusion world models within reinforcement learning training loops, its sampling rate remains insufficient to support large-scale, real-time closed-loop simulations. This constraint necessitates a reliance on sophisticated distillation techniques or computational resource stacking for diffusion models in interactive environments requiring high-frequency feedback (such as the real-time gaming environments discussed in GameNGen[145]). This imbalance between efficiency and precision circumscribes the transition of diffusion world models from offline generation toward real-time embodied decision-making.

4.2. Multimodal Fusion

Multimodal input constitutes the quintessential attribute that facilitates the transition of world models from rudimentary video synthesizers into "general cognitive foundations." This attribute enables the seamless integration of heterogeneous information from disparate sources—encompassing visual observations (imagery/video), semantic directives (textual corpora), structural constraints (HD maps/3D bounding boxes), and control signals (actions). The utility of this multimodal capacity resides in providing multidimensional conditional guidance for the predictive processes of the model, thereby empowering it to transcend pixel-level statistical correlations and distill the latent semantic logic and governing physical principles underlying visual sequences. For world models, multimodal characteristics represent an indispensable prerequisite for attaining high degrees of controllability and situational awareness. In the context of intricate physical interactions, unimodal visual input is frequently insufficient to mitigate predictive stochasticity, only through the synergistic application of macroscopic linguistic guidance and microscopic dynamical constraints can the model accurately "collapse" the expansive state space into deterministic, physically plausible evolutionary outcomes.
Regarding technical implementation and empirical utility, multimodal integration imbues world models with substantial robustness and heightened control precision. Within the 3D-VLA[115] framework, the cross-modal alignment function is instantiated as an "interaction tokenization" process, which systematically aligns 3D point clouds with textual directives within the latent manifold of LLMs, thereby imbuing the diffusion process with spatial reasoning capabilities within 3D physical domains. This integration not only expands the generative dimensionality but, more pivotally, establishes a closed-loop modeling paradigm—extending from perception to actuation—mediated by action directives. In the realm of autonomous driving, the decoupling capacity of the GAIA-1[32] diffusion framework when processing high-dimensional actions ensures that synthesized driving scenarios exhibit rigorous causal grounding. Furthermore, DriveDreamer[104] enhances the structural determinism of cross-modal alignment within complex traffic flows by incorporating structured constraints such as HDMaps, ensuring that synthesized sequences adhere to user intent while remaining strictly anchored by physical perspective and traffic regulations. Additionally, MotionDiffuser[103] elucidates the intrinsic advantages of diffusion models in managing predictive uncertainty. It synergizes historical trajectories with cartographic information and steers the denoising process via the introduction of differentiable cost functions during the inference phase. Consequently, the cross-modal alignment function in this context transcends static mapping, functioning as dynamic probabilistic guidance that enables the model to derive multiple logically consistent interaction strategies from a unified initial state.
Figure 9. Multimodal fusion characteristics. DWMs possess the capability for multimodal feature fusion, enabling the alignment of heterogeneous data modalities in real–world environments and serving as an indispensable prerequisite for achieving high levels of controllability and situational awareness.
Figure 9. Multimodal fusion characteristics. DWMs possess the capability for multimodal feature fusion, enabling the alignment of heterogeneous data modalities in real–world environments and serving as an indispensable prerequisite for achieving high levels of controllability and situational awareness.
Preprints 233002 g009
Nevertheless, the practical deployment of multimodal fusion remains beset by formidable challenges, specifically the inherent heterogeneity of modal representations and the escalating computational complexity induced by informational redundancy. Firstly, multimodal data exhibit fundamental disparities in mathematical structure and information density, achieving cross-modal representation alignment within a unified feature space, such as for discrete text tags, continuous high-bit image pixels, and sparse LiDAR point clouds[96], remains a challenge for fields like autonomous driving. Owing to the semantic gap between low-level sensory observations and high-level cognitive decision-making, the fusion process is susceptible to the attrition of fine-grained information and may precipitate modality imbalance, a phenomenon wherein the model exhibits a disproportionate reliance on dominant discriminative modalities at the expense of subordinate ones. Secondly, informational redundancy and stochastic noise are pervasive in multimodal corpora, where naive all-to-all attention mechanisms incur a computational complexity that scales exponentially with the cardinality of modalities. Consequently, there is a critical exigency for efficacious feature extraction and pre-compression mechanisms to ensure that the model selectively propagates only the essential, mutually complementary features across modalities.

4.3. Interactive Dynamics

Dynamic interactivity constitutes the defining attribute that distinguishes world models from traditional video generation frameworks, marking a fundamental transition from a passive perceptual observer to an active environmental emulator. This capability allows the model to dynamically predict and synthesize the successive spatiotemporal evolution of environmental states based on externally input action sequences or interactive directives, thereby constructing a closed-loop interactive framework. For world models, interactivity serves not only as the substrate for embodied policy rehearsal but also as the core embodiment of the model’s counterfactual reasoning capacity. By simulating multiple future trajectories under divergent action sequences, the model facilitates agent-environment exploration and path planning within risk-averse virtual manifolds. This "action-conditioned denoising process" reconfigures the diffusion paradigm from stochastic pixel synthesis into a dynamic response governed by physical causal laws, thereby providing a high-fidelity "imagination-driven" training environment for reinforcement learning.
Figure 10. Example of interactive dynamics. The interactive dynamics of DWMs in autonomous driving scenarios, where such interactions directly influence the accuracy and success rate of trajectory prediction among surrounding agents.
Figure 10. Example of interactive dynamics. The interactive dynamics of DWMs in autonomous driving scenarios, where such interactions directly influence the accuracy and success rate of trajectory prediction among surrounding agents.
Preprints 233002 g010
In terms of methodological trajectories and empirical utility, various innovative mechanisms have been explored to address heterogeneous levels of interactive requirements. UniSim[33] prioritizes the mapping bottleneck between action interaction and diffusion, demonstrating that the diffusion paradigm can serve as a unified substrate for both components. By jointly training on massive internet-scale video and annotated robotic datasets, it encodes kinematic displacement vectors or high-level directives into robust conditional signals, enabling the diffusion process to achieve fine-grained pixel-level reconfiguration according to divergent action signals rather than mere stochastic denoising. It successfully transmutes highly abstract action vectors into robust directional steerage within the diffusion process, achieving superlative interactive fidelity. Simultaneously, GAIA-1[32] and Drive-WM[86] bolster the control of interactive perturbations over spatiotemporal evolution within the autonomous driving domain. Notably, Drive-WM[86] leverages the multiview generative affordances of diffusion models to ensure that perceptual sensor streams synthesized following specific driving maneuvers are not only temporally coherent but also strictly adhere to driving semantics across disparate camera perspectives.
The distinct characteristic of PhysDreamer[117] resides in its explicit distillation of physical priors, demonstrating that interactivity is contingent not only upon exogenous action inputs but also upon the model’s comprehension of physical common sense. Eschewing the explicit modeling of partial differential equations, it distills physical priors from large-scale video diffusion models. When an infinitesimal impulse is applied, the model can automatically infer material stiffness and elastic moduli via these distilled priors to synthesize physically veridical deformation responses. Diverging from large-scale retraining methodologies, PEEKABOO[140] introduces a parsimonious trajectory that bypasses the need for retraining the diffusion term by executing interactive interventions through fine-grained architectural manipulations. Its underlying mechanism involves injecting spatial masks into the attention layers of the diffusion process, which is mathematically equivalent to imposing a rigid interactive constraint within the generative formulation. By forcing a reconfiguration of feature weights during noise prediction, the model adaptively responds to user interaction directives in real-time. Furthermore, RoboDreamer[118] addresses the generalization bottlenecks of interactive intervention when confronted with complex tasks. By decomposing intricate interactive instructions into foundational behavioral primitives, the diffusion operator can synthesize plausible future evolutions during the denoising process based on primitive combinations when handling out-of-distribution task compositions. This not only enhances the model’s robustness against ambiguous instructions but also provides multimodal alternative trajectories for the agent’s simulation-driven training.
However, the realization of a fully interactive world model remains beset by critical challenges, such as interactive signal discrimination and high-order semantic reasoning. First, within complex dynamic environments, pixel-level disturbances induced by subtle interactive actions are easily obfuscated by inherent environmental noise, rendering it difficult for the model to disentangle spontaneous environmental evolution from action-induced perturbations. Second, although diffusion models excel at modeling multimodal distributions, ensuring a strong deterministic logical association between action directives and generative outcomes in real-time continuous interaction remains a formidable challenge. Finally, contemporary research is largely confined to passive responses guided by pixels or low-order trajectories, lacking a profound comprehension of the environment’s underlying logic and functional affordances. This deficiency in semantic-aware interactivity implies that while the model can execute explicit directives such as "turn left," it struggles to comprehend and respond to highly abstract, reasoning-dependent task objectives such as "finding shelter from the rain." Consequently, an ideal interactive world model necessitates not only robust generative fidelity but also the construction of a deep understanding regarding the ontological and functional logic of the physical world.

4.4. Spatiotemporal Coherence

Spatiotemporal consistency serves as the fundamental cornerstone for constructing stable virtual realities within world models, ensuring that synthesized visual sequences exhibit logical rigor in terms of spatial structure, entity identity, and temporal evolution. Within the context of world modeling, consistency represents not merely a heuristic for enhancing visual fidelity but a quintessential requisite for the model’s functionality as a high-fidelity simulator. Should the model fail to maintain object constancy during kinetic processes, or should the background undergo distortion and collapse during viewpoint transitions, agents will be precluded from distilling stable physical laws or causal relationships, potentially yielding erroneous decision logic predicated on spurious feedback from hallucinations. Consequently, world models must reinforce consistency to instantiate a predictive and reliable digital environment, transforming discrete frame synthesis into a spatiotemporally continuous dynamic simulation.
Regarding methodological implementation, the maintenance of consistency yields significant simulation gains and has precipitated a variety of innovative computational paradigms. The crux of ConsistI2V[138] lies in its profound insight into the propensity of diffusion models to exhibit "state amnesia" regarding the initial frame during long-horizon sequences. This work modifies the attention mechanisms within the diffusion process, mandating pixel-level feature interaction between the diffusion operator and the inaugural frame during the generation of each subsequent frame. Such a design effectively curtails visual drift in long-range generation and ameliorates the prevalent "identity flipping" issue in diffusion models. However, it may induce kinetic rigidity during large-scale object displacements due to excessive reliance on the initial frame. Conversely, GAIA-1[32] utilizes a large-scale Transformer to perform physical logic inference within a discrete token space, subsequently employing these logical tokens as conditions to guide the video diffusion model in pixel rendering. The architectural advantage of this approach lies in relieving the diffusion model of the computational burden associated with complex physical inference, allowing it to specialize exclusively in visual representation. This divide-and-conquer strategy enables GAIA-1[32] to synthesize logically consistent autonomous driving scenarios spanning several minutes, achieving dual consistency in both macro-layout and micro-detail.
The innovation of Vid2World[123] resides in its attempt to integrate local dynamics with diffusion perception within the established formulation. Whereas traditional diffusion models focus on acausal snapshot generation, Vid2World[123] performs a causal reformulation of the pre-trained model, endowing it with the capacity for autoregressive prediction conditioned on action sequences. Consequently, the diffusion term transcends its role as a mere renderer, having internalized the governing laws of physical evolution. Upon the input of divergent driving maneuvers, the diffusion process synthesizes diverse trajectories with strong causal grounding, providing a reliable closed-loop simulation environment for the agent. WorldGPT[141], in contrast, intervenes via the constraint variables of the diffusion perceptual term by introducing an LLM as a high-level controller, which generates granular task scripts to fortify conditional constraints for diffusion-based synthesis. This methodology effectively introduces semantic navigation to spatiotemporal evolution, ensuring that during long-horizon operations, the progression of each frame is not only visually coherent but also strictly aligned with human-interpretable task logic.
Figure 11. Two types of spatiotemporal coherence in world models. The first is temporal consistency: if preserved, the model can maintain accurate predictions over extended horizons; otherwise, errors accumulate over time, and this leads to a change in the object. The second is physical consistency, where violations of physical laws, such as implausible soccer ball dynamics, may occur.
Figure 11. Two types of spatiotemporal coherence in world models. The first is temporal consistency: if preserved, the model can maintain accurate predictions over extended horizons; otherwise, errors accumulate over time, and this leads to a change in the object. The second is physical consistency, where violations of physical laws, such as implausible soccer ball dynamics, may occur.
Preprints 233002 g011
Although the aforementioned mechanisms have achieved significant breakthroughs in generative fidelity and visual coherence, diffusion-based world models confront fundamental bottlenecks in ensuring rigorous physical law consistency. Current limitations are primarily manifested in the deficiency of explicit geometric awareness, perspective distortion during dynamic interaction, and the failure to model "object permanence." Specifically, most existing models operate within a 2D latent space or rely exclusively on temporal-dimension convolutions and attention mechanisms to simulate temporal evolution. Such a paradigm essentially performs conditional fitting of pixel distributions, lacking an explicit understanding of the underlying 3D Euclidean spatial structure. Consequently, in interactive tasks involving substantial viewpoint translation or rotation, models struggle to adhere to geometric constraints, frequently resulting in scale drift or perspective errors within the synthesized scenes[86]. For instance, in autonomous driving simulations, the variation in the geometric dimensions of distal objects as the vehicle approaches often deviates from real-world physical regularities. Furthermore, when handling complex interactive scenarios involving the intersection or overlapping of multiple entities, models frequently fail to effectively differentiate between physical occlusion and stochastic fusion. This lack of semantic parsing capability leads to frequent identity flipping or attribute drift once entities are de-occluded[94]. This reflects a critical modeling deficiency within the current diffusion generative paradigm regarding the fundamental cognitive concept of "object permanence"—specifically, the model’s failure to truly internalize the spatiotemporal stability of an object as an independent entity.

4.5. Multi-Environment

The proactive synthesis of diversified environments represents a pivotal stride for world models toward AGI, as it endows models with the capacity to autonomously construct simulation spaces characterized by heterogeneous visual styles, geographic features, and meteorological conditions based on high-level directives or structured priors. The utility of this attribute for world models lies in its ability to transcend the constraints of conventional simulators tethered to preset assets and elevate "data playback" to "scene synthesis," thereby positioning the model as an inexhaustible content source that provides massive training corpora—including long-tail scenarios—for systems such as autonomous driving and embodied robotics. The exigent requirement for this characteristic in diffusion-based world models stems from the fact that the complexity of the physical world significantly outstrips the coverage of static datasets. To cultivate agents endowed with robust generalization capabilities, models must proactively evolve extreme, rare, or cross-regional simulation environments; by iteratively refining policies within these diversified virtual spaces in a safe and efficient manner, the model fundamentally addresses the challenges of high acquisition costs and non-uniform distribution inherent in real-world data.
Figure 12. Multi-environment characteristics. DWMs demonstrate notable strengths in capturing environmental stochasticity and enabling cross–domain generalization.
Figure 12. Multi-environment characteristics. DWMs demonstrate notable strengths in capturing environmental stochasticity and enabling cross–domain generalization.
Preprints 233002 g012
Regarding methodological trajectories and empirical utility, recently emerged state-of-the-art models demonstrate a diverse range of technical solutions spanning from semantic guidance to spatial geometric constraints. The primary contribution of DriveDreamer-2[87] resides in mitigating the uncertainty of diffusion models when processing ambiguous semantics. By incorporating LLMs as semantic converters, it translates vague user instructions into highly structured logical chains, thereby enhancing generative precision under adverse weather conditions (e.g., rain or snow) and complex interaction cases, which effectively ameliorates the instability of traditional diffusion models under rare semantic compositions. Conversely, Streetscapes[35] demonstrates the elicitation of the generative potential of diffusion models through stabilized geometric constraints. It formalizes the geometric fidelity of space by leveraging maps and heightmaps; under these constraints, global style transfer is achieved by varying regional priors within the diffusion process upon a fixed topological structure. This validates that diffusion-based world models can synthesize infinitely diversified environmental textures while preserving physical spatial constraints, thereby adapting to the simulation requirements of disparate geographic regions.
While diffusion-based world models exhibit pronounced advantages in capturing environmental stochasticity and cross-domain generalization, achieving full-scene coverage and modeling extreme-value distributions remains a formidable challenge. Firstly, contemporary synthesis of diversified environments is frequently confined to the stochastic combination of visual appearances—such as weather and illumination—rather than a substantive expansion of the underlying environmental state space. The diffusion process predominantly manifests as nonlinear interpolation on high-dimensional manifolds, which renders generative diversity heavily contingent upon the training dataset. When confronted with requirements for diversifying underlying topological structures or environmental rules, models often struggle to synthesize structurally novel environments, leading to significant homogeneity in the underlying logic of the generative outcomes. Secondly, the generalization capacity of world models remains bounded when encountering unseen environmental features. Although pre-training imbues models with a degree of universality, significant distribution shifts between source and target domains during cross-domain transfer often lead to a precipitous decline in generative quality due to a lack of robust disentangled representations. This latency in adaptation to novel scenarios impedes the seamless deployment of world models in expansive, dynamic, and diversified environments.

6. Conclusion

This paper provides a systematic review of the current state of diffusion-based world model research, thoroughly elucidating the pivotal role of diffusion models in advancing world models toward high fidelity and multimodality. It meticulously traces the application progress of this paradigm across three key domains–autonomous driving, embodied intelligence, and general simulation–while summarizing and evaluating representative research outcomes. Furthermore, this paper distills five core characteristics inherent to DWMs and analyzes how these traits drive model evolution. Finally, addressing common deficiencies inherited from diffusion models into world models and the resulting challenges, it prospectively proposes future research directions, aiming to provide foundational reference for sustained innovation in this field.

References

  1. Ha, D.; Schmidhuber, J. World models. arXiv 2018, arXiv:1803.101222, 440. [Google Scholar]
  2. Hafner, D.; Lillicrap, T.; Ba, J.; Norouzi, M. Dream to control: Learning behaviors by latent imagination. arXiv 2019, arXiv:1912.01603. [Google Scholar]
  3. Hafner, D.; Lillicrap, T.; Norouzi, M.; Ba, J. Mastering atari with discrete world models. arXiv 2020, arXiv:2010.02193. [Google Scholar]
  4. Hafner, D.; Pasukonis, J.; Ba, J.; Lillicrap, T. Mastering diverse domains through world models. arXiv 2023, arXiv:2301.04104. [Google Scholar]
  5. Kim, S.W.; Zhou, Y.; Philion, J.; Torralba, A.; Fidler, S. Learning to simulate dynamic environments with gamegan. In Proceedings of the Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020; pp. 1231–1240. [Google Scholar]
  6. Yan, W.; Zhang, Y.; Abbeel, P.; Srinivas, A. Videogpt: Video generation using vq-vae and transformers. arXiv 2021, arXiv:2104.10157. [Google Scholar]
  7. Bruce, J.; Dennis, M.D.; Edwards, A.; Parker-Holder, J.; Shi, Y.; Hughes, E.; Lai, M.; Mavalankar, A.; Steigerwald, R.; Apps, C.; et al. Genie: Generative interactive environments. In Proceedings of the Forty-first International Conference on Machine Learning, 2024. [Google Scholar]
  8. LeCun, Y.; et al. A path towards autonomous machine intelligence version 0.9. 2, 2022-06-27. Open Rev. 2022, 62, 1–62. [Google Scholar]
  9. Alhaija, H.A.; Alvarez, J.; Bala, M.; Cai, T.; Cao, T.; Cha, L.; Chen, J.; Chen, M.; Ferroni, F.; Fidler, S.; et al. Cosmos-transfer1: Conditional world generation with adaptive multimodal control. arXiv 2025, arXiv:2503.14492. [Google Scholar]
  10. Ren, X.; Lu, Y.; Cao, T.; Gao, R.; Huang, S.; Sabour, A.; Shen, T.; Pfaff, T.; Wu, J.Z.; Chen, R.; et al. Cosmos-drive-dreams: Scalable synthetic driving data generation with world foundation models. arXiv 2025, arXiv:2506.09042. [Google Scholar]
  11. Karaarslan, E.; Aydın, Ö. OpenAI Sora: generate impressive videos with text instructions. SSRN Electron. J. 2024. [Google Scholar] [CrossRef]
  12. Ding, J.; Zhang, Y.; Shang, Y.; Zhang, Y.; Zong, Z.; Feng, J.; Yuan, Y.; Su, H.; Li, N.; Sukiennik, N.; et al. Understanding world or predicting future? a comprehensive survey of world models. ACM Comput. Surv. 2025, 58, 1–38. [Google Scholar] [CrossRef]
  13. Tu, S.; Zhou, X.; Liang, D.; Jiang, X.; Zhang, Y.; Li, X.; Bai, X. The role of world models in shaping autonomous driving: A comprehensive survey. arXiv 2025, arXiv:2502.10498. [Google Scholar]
  14. Kong, L.; Yang, W.; Mei, J.; Liu, Y.; Liang, A.; Zhu, D.; Lu, D.; Yin, W.; Hu, X.; Jia, M.; et al. 3D and 4D world modeling: A survey. arXiv 2025, arXiv:2509.07996. [Google Scholar]
  15. Li, X.; He, X.; Zhang, L.; Wu, M.; Li, X.; Liu, Y. A comprehensive survey on world models for embodied ai. arXiv 2025, arXiv:2510.16732. [Google Scholar]
  16. Zhang, P.F.; Cheng, Y.; Sun, X.; Wang, S.; Li, F.; Zhu, L.; Shen, H.T. A step toward world models: A survey on robotic manipulation. arXiv 2025, arXiv:2511.02097. [Google Scholar]
  17. Liu, Y.; Chen, W.; Bai, Y.; Liang, X.; Li, G.; Gao, W.; Lin, L. Aligning cyber space with physical world: A comprehensive survey on embodied ai. IEEE/ASME Transactions on Mechatronics 2025. [Google Scholar] [CrossRef]
  18. Feng, T.; Wang, W.; Yang, Y. A survey of world models for autonomous driving. arXiv 2025, arXiv:2501.11260. [Google Scholar]
  19. Long, X.; Zhao, Q.; Zhang, K.; Zhang, Z.; Wang, D.; Liu, Y.; Shu, Z.; Lu, Y.; Wang, S.; Wei, X.; et al. A survey: Learning embodied intelligence from physical simulators and world models. arXiv 2025, arXiv:2507.00917. [Google Scholar]
  20. Ho, J.; Jain, A.; Abbeel, P. Denoising diffusion probabilistic models. Adv. Neural Inf. Process. Syst. 2020, 33, 6840–6851. [Google Scholar]
  21. Dhariwal, P.; Nichol, A. Diffusion models beat gans on image synthesis. Adv. Neural Inf. Process. Syst. 2021, 34, 8780–8794. [Google Scholar]
  22. Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; Ommer, B. High-resolution image synthesis with latent diffusion models. In Proceedings of the Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022; pp. 10684–10695. [Google Scholar]
  23. Brooks, T.; Holynski, A.; Efros, A.A. Instructpix2pix: Learning to follow image editing instructions. In Proceedings of the Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023; pp. 18392–18402. [Google Scholar]
  24. Ho, J.; Salimans, T.; Gritsenko, A.; Chan, W.; Norouzi, M.; Fleet, D.J. Video diffusion models. Adv. Neural Inf. Process. Syst. 2022, 35, 8633–8646. [Google Scholar] [CrossRef]
  25. Peebles, W.; Xie, S. Scalable diffusion models with transformers. In Proceedings of the Proceedings of the IEEE/CVF international conference on computer vision, 2023; pp. 4195–4205. [Google Scholar]
  26. Salimans, T.; Ho, J. Progressive distillation for fast sampling of diffusion models. arXiv 2022, arXiv:2202.00512. [Google Scholar]
  27. Xiao, Z.; Kreis, K.; Vahdat, A. Tackling the generative learning trilemma with denoising diffusion GANs. arXiv 2021, arXiv:2112.07804. [Google Scholar]
  28. Liu, N.; Li, S.; Du, Y.; Torralba, A.; Tenenbaum, J.B. Compositional visual generation with composable diffusion models. In Proceedings of the European conference on computer vision, 2022; Springer; pp. 423–439. [Google Scholar]
  29. Feng, W.; He, X.; Fu, T.J.; Jampani, V.; Akula, A.; Narayana, P.; Basu, S.; Wang, X.E.; Wang, W.Y. Training-free structured diffusion guidance for compositional text-to-image synthesis. arXiv 2022, arXiv:2212.05032. [Google Scholar]
  30. Brooks, T.; Peebles, B.; Holmes, C.; DePue, W.; Guo, Y.; Jing, L.; Schnurr, D.; Taylor, J.; Luhman, T.; Luhman, E.; et al. Video generation models as world simulators. 2024, 3, p. 3. https://openai.com/research/video-generation-models-as-world-simulators.
  31. Gao, S.; Yang, J.; Chen, L.; Chitta, K.; Qiu, Y.; Geiger, A.; Zhang, J.; Li, H. Vista: A generalizable driving world model with high fidelity and versatile controllability. Adv. Neural Inf. Process. Syst. 2024, 37, 91560–91596. [Google Scholar] [CrossRef]
  32. Hu, A.; Russell, L.; Yeo, H.; Murez, Z.; Fedoseev, G.; Kendall, A.; Shotton, J.; Corrado, G. Gaia-1: A generative world model for autonomous driving. arXiv 2023, arXiv:2309.17080. [Google Scholar]
  33. Yang, S.; Du, Y.; Ghasemipour, K.; Tompson, J.; Kaelbling, L.; Schuurmans, D.; Abbeel, P. Learning interactive real-world simulators. arXiv 2023, arXiv:2310.06114. [Google Scholar]
  34. Feng, Y.; Tan, H.; Mao, X.; Xiang, C.; Liu, G.; Huang, S.; Su, H.; Zhu, J. Vidar: Embodied video diffusion model for generalist manipulation. arXiv 2025, arXiv:2507.12898. [Google Scholar]
  35. Deng, B.; Tucker, R.; Li, Z.; Guibas, L.; Snavely, N.; Wetzstein, G. Streetscapes: Large-scale consistent street view generation using autoregressive video diffusion. In Proceedings of the ACM SIGGRAPH 2024 Conference Papers, 2024; pp. 1–11. [Google Scholar]
  36. Sohl-Dickstein, J.; Weiss, E.; Maheswaranathan, N.; Ganguli, S. Deep unsupervised learning using nonequilibrium thermodynamics. In Proceedings of the International conference on machine learning. pmlr, 2015; pp. 2256–2265. [Google Scholar]
  37. Ho, J.; Salimans, T. Classifier-free diffusion guidance. arXiv 2022, arXiv:2207.12598. [Google Scholar]
  38. Sun, P.; Kretzschmar, H.; Dotiwalla, X.; Chouard, A.; Patnaik, V.; Tsui, P.; Guo, J.; Zhou, Y.; Chai, Y.; Caine, B.; et al. Scalability in perception for autonomous driving: Waymo open dataset. In Proceedings of the Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020; pp. 2446–2454. [Google Scholar]
  39. Wilson, B.; Qi, W.; Agarwal, T.; Lambert, J.; Singh, J.; Khandelwal, S.; Pan, B.; Kumar, R.; Hartnett, A.; Pontes, J.K.; et al. Argoverse 2: Next generation datasets for self-driving perception and forecasting. arXiv 2023, arXiv:2301.00493. [Google Scholar]
  40. Fu, J.; Kumar, A.; Nachum, O.; Tucker, G.; Levine, S. D4rl: Datasets for deep data-driven reinforcement learning. arXiv 2020, arXiv:2004.07219. [Google Scholar]
  41. Caesar, H.; Bankiti, V.; Lang, A.H.; Vora, S.; Liong, V.E.; Xu, Q.; Krishnan, A.; Pan, Y.; Baldan, G.; Beijbom, O. nuscenes: A multimodal dataset for autonomous driving. In Proceedings of the Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020; pp. 11621–11631. [Google Scholar]
  42. Caesar, H.; Kabzan, J.; Tan, K.S.; Fong, W.K.; Wolff, E.; Lang, A.; Fletcher, L.; Beijbom, O.; Omari, S. nuplan: A closed-loop ml-based planning benchmark for autonomous vehicles. arXiv 2021, arXiv:2106.11810. [Google Scholar]
  43. Ettinger, S.; Cheng, S.; Caine, B.; Liu, C.; Zhao, H.; Pradhan, S.; Chai, Y.; Sapp, B.; Qi, C.R.; Zhou, Y.; et al. Large scale interactive motion forecasting for autonomous driving: The waymo open motion dataset. In Proceedings of the Proceedings of the IEEE/CVF international conference on computer vision, 2021; pp. 9710–9719. [Google Scholar]
  44. Chen, K.; Ge, R.; Qiu, H.; Ai-Rfou, R.; Qi, C.; Zhou, X.; Yang, Z.; Ettinger, S.; Sun, P.; Leng, Z.; et al. Womd-lidar: Raw sensor dataset benchmark for motion forecasting. In Proceedings of the 2024 IEEE International Conference on Robotics and Automation (ICRA); IEEE, 2024; pp. 4766–4773. [Google Scholar]
  45. Li, K.; Chen, K.; Wang, H.; Hong, L.; Ye, C.; Han, J.; Chen, Y.; Zhang, W.; Xu, C.; Yeung, D.Y.; et al. Coda: A real-world road corner case dataset for object detection in autonomous driving. In Proceedings of the European Conference on Computer Vision, 2022; Springer; pp. 406–423. [Google Scholar]
  46. Dauner, D.; Hallgarten, M.; Li, T.; Weng, X.; Huang, Z.; Yang, Z.; Li, H.; Gilitschenski, I.; Ivanovic, B.; Pavone, M.; et al. Navsim: Data-driven non-reactive autonomous vehicle simulation and benchmarking. Adv. Neural Inf. Process. Syst. 2024, 37, 28706–28719. [Google Scholar] [CrossRef]
  47. Yang, J.; Gao, S.; Qiu, Y.; Chen, L.; Li, T.; Dai, B.; Chitta, K.; Wu, P.; Zeng, J.; Luo, P.; et al. Generalized predictive model for autonomous driving. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024; pp. 14662–14672. [Google Scholar]
  48. Hirose, N.; Sadeghian, A.; Vázquez, M.; Goebel, P.; Savarese, S. Gonet: A semi-supervised deep learning approach for traversability estimation. In Proceedings of the 2018 IEEE/RSJ international conference on intelligent robots and systems (IROS); IEEE, 2018; pp. 3044–3051. [Google Scholar]
  49. Hirose, N.; Sadeghian, A.; Xia, F.; Martín-Martín, R.; Savarese, S. Vunet: Dynamic scene view synthesis for traversability estimation using an rgb camera. IEEE Robot. Autom. Lett. 2019, 4, 2062–2069. [Google Scholar] [CrossRef]
  50. Damen, D.; Doughty, H.; Farinella, G.M.; Fidler, S.; Furnari, A.; Kazakos, E.; Moltisanti, D.; Munro, J.; Perrett, T.; Price, W.; et al. Scaling egocentric vision: The epic-kitchens dataset. In Proceedings of the Proceedings of the European conference on computer vision (ECCV), 2018; pp. 720–736. [Google Scholar]
  51. Ebert, F.; Dasari, S.; Lee, A.X.; Levine, S.; Finn, C. Robustness via retrying: Closed-loop robotic manipulation with self-supervised learning. In Proceedings of the Conference on robot learning. PMLR, 2018; pp. 983–993. [Google Scholar]
  52. Miech, A.; Zhukov, D.; Alayrac, J.B.; Tapaswi, M.; Laptev, I.; Sivic, J. Howto100m: Learning a text-video embedding by watching hundred million narrated video clips. In Proceedings of the Proceedings of the IEEE/CVF international conference on computer vision, 2019; pp. 2630–2640. [Google Scholar]
  53. Shah, D.; Eysenbach, B.; Kahn, G.; Rhinehart, N.; Levine, S. Rapid exploration for open-world navigation with latent goal models. arXiv 2021, arXiv:2104.05859. [Google Scholar]
  54. Karnan, H.; Nair, A.; Xiao, X.; Warnell, G.; Pirk, S.; Toshev, A.; Hart, J.; Biswas, J.; Stone, P. Socially compliant navigation dataset (scand): A large-scale dataset of demonstrations for social navigation. IEEE Robot. Autom. Lett. 2022, 7, 11807–11814. [Google Scholar] [CrossRef]
  55. Grauman, K.; Westbury, A.; Byrne, E.; Chavis, Z.; Furnari, A.; Girdhar, R.; Hamburger, J.; Jiang, H.; Liu, M.; Liu, X.; et al. Ego4d: Around the world in 3,000 hours of egocentric video. In Proceedings of the Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022; pp. 18995–19012. [Google Scholar]
  56. Brohan, A.; Brown, N.; Carbajal, J.; Chebotar, Y.; Dabis, J.; Finn, C.; Gopalakrishnan, K.; Hausman, K.; Herzog, A.; Hsu, J.; et al. Rt-1: Robotics transformer for real-world control at scale. arXiv 2022, arXiv:2212.06817. [Google Scholar]
  57. O’Neill, A.; Rehman, A.; Maddukuri, A.; Gupta, A.; Padalkar, A.; Lee, A.; Pooley, A.; Gupta, A.; Mandlekar, A.; Jain, A.; et al. Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0. In Proceedings of the 2024 IEEE International Conference on Robotics and Automation (ICRA); IEEE, 2024; pp. 6892–6903. [Google Scholar]
  58. Lynch, C.; Wahid, A.; Tompson, J.; Ding, T.; Betker, J.; Baruch, R.; Armstrong, T.; Florence, P. Interactive language: Talking to robots in real time. IEEE Robotics and Automation Letters, 2023. [Google Scholar]
  59. Ebert, F.; Yang, Y.; Schmeckpeper, K.; Bucher, B.; Georgakis, G.; Daniilidis, K.; Finn, C.; Levine, S. Bridge data: Boosting generalization of robotic skills with cross-domain datasets. arXiv 2021, arXiv:2109.13396. [Google Scholar]
  60. Triest, S.; Sivaprakasam, M.; Wang, S.J.; Wang, W.; Johnson, A.M.; Scherer, S. Tartandrive: A large-scale dataset for learning off-road dynamics models. In Proceedings of the 2022 International Conference on Robotics and Automation (ICRA); IEEE, 2022; pp. 2546–2552. [Google Scholar]
  61. Liu, B.; Zhu, Y.; Gao, C.; Feng, Y.; Liu, Q.; Zhu, Y.; Stone, P. Libero: Benchmarking knowledge transfer for lifelong robot learning. Adv. Neural Inf. Process. Syst. 2023, 36, 44776–44791. [Google Scholar] [CrossRef]
  62. Walke, H.R.; Black, K.; Zhao, T.Z.; Vuong, Q.; Zheng, C.; Hansen-Estruch, P.; He, A.W.; Myers, V.; Kim, M.J.; Du, M.; et al. Bridgedata v2: A dataset for robot learning at scale. In Proceedings of the Conference on Robot Learning. PMLR, 2023; pp. 1723–1736. [Google Scholar]
  63. Li, X.; Li, P.; Liu, M.; Wang, D.; Liu, J.; Kang, B.; Ma, X.; Kong, T.; Zhang, H.; Liu, H. Towards Generalist Robot Policies: What Matters in Building Vision-Language-Action Models. arXiv 2024, arXiv:2412.14058. [Google Scholar]
  64. Wu, K.; Hou, C.; Liu, J.; Che, Z.; Ju, X.; Yang, Z.; Li, M.; Zhao, Y.; Xu, Z.; Yang, G.; et al. Robomind: Benchmark on multi-embodiment intelligence normative data for robot manipulation. arXiv 2024, arXiv:2412.13877. [Google Scholar]
  65. Hoque, R.; Huang, P.; Yoon, D.J.; Sivapurapu, M.; Zhang, J. Egodex: Learning dexterous manipulation from large-scale egocentric video. arXiv 2025, arXiv:2505.11709. [Google Scholar]
  66. Chen, T.; Chen, Z.; Chen, B.; Cai, Z.; Liu, Y.; Li, Z.; Liang, Q.; Lin, X.; Ge, Y.; Gu, Z.; et al. Robotwin 2.0: A scalable data generator and benchmark with strong domain randomization for robust bimanual robotic manipulation. arXiv 2025, arXiv:2506.18088. [Google Scholar]
  67. Soomro, K.; Zamir, A.R.; Shah, M. Ucf101: A dataset of 101 human actions classes from videos in the wild. arXiv 2012, arXiv:1212.0402. [Google Scholar]
  68. Perazzi, F.; Pont-Tuset, J.; McWilliams, B.; Van Gool, L.; Gross, M.; Sorkine-Hornung, A. A benchmark dataset and evaluation methodology for video object segmentation. In Proceedings of the Proceedings of the IEEE conference on computer vision and pattern recognition, 2016; pp. 724–732. [Google Scholar]
  69. Xu, J.; Mei, T.; Yao, T.; Rui, Y. Msr-vtt: A large video description dataset for bridging video and language. In Proceedings of the Proceedings of the IEEE conference on computer vision and pattern recognition, 2016; pp. 5288–5296. [Google Scholar]
  70. Goyal, R.; Ebrahimi Kahou, S.; Michalski, V.; Materzynska, J.; Westphal, S.; Kim, H.; Haenel, V.; Fruend, I.; Yianilos, P.; Mueller-Freitag, M.; et al. The" something something" video database for learning and evaluating visual common sense. In Proceedings of the Proceedings of the IEEE international conference on computer vision, 2017; pp. 5842–5850. [Google Scholar]
  71. Fan, H.; Lin, L.; Yang, F.; Chu, P.; Deng, G.; Yu, S.; Bai, H.; Xu, Y.; Liao, C.; Ling, H. Lasot: A high-quality benchmark for large-scale single object tracking. In Proceedings of the Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019; pp. 5374–5383. [Google Scholar]
  72. Fan, H.; Bai, H.; Lin, L.; Yang, F.; Chu, P.; Deng, G.; Yu, S.; Harshit; Huang, M.; Liu, J.; et al. Lasot: A high-quality large-scale single object tracking benchmark. Int. J. Comput. Vis. 2021, 129, 439–461. [Google Scholar] [CrossRef]
  73. Mildenhall, B.; Srinivasan, P.P.; Ortiz-Cayon, R.; Kalantari, N.K.; Ramamoorthi, R.; Ng, R.; Kar, A. Local light field fusion: Practical view synthesis with prescriptive sampling guidelines. ACM Trans. Graph. (ToG) 2019, 38, 1–14. [Google Scholar]
  74. Cobbe, K.; Hesse, C.; Hilton, J.; Schulman, J. Leveraging procedural generation to benchmark reinforcement learning. In Proceedings of the International conference on machine learning. PMLR, 2020; pp. 2048–2056. [Google Scholar]
  75. Barron, J.T.; Mildenhall, B.; Verbin, D.; Srinivasan, P.P.; Hedman, P. Mip-nerf 360: Unbounded anti-aliased neural radiance fields. In Proceedings of the Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022; pp. 5470–5479. [Google Scholar]
  76. Menapace, W.; Lathuiliere, S.; Siarohin, A.; Theobalt, C.; Tulyakov, S.; Golyanik, V.; Ricci, E. Playable environments: Video manipulation in space and time. In Proceedings of the Proceedings of the ieee/cvf conference on computer vision and pattern recognition, 2022; pp. 3584–3593. [Google Scholar]
  77. Yin, S.; Wu, C.; Yang, H.; Wang, J.; Wang, X.; Ni, M.; Yang, Z.; Li, L.; Liu, S.; Yang, F.; et al. Nuwa-xl: Diffusion over diffusion for extremely long video generation. In Proceedings of the Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers); 2023; pp. 1309–1320. [Google Scholar]
  78. Chevalier-Boisvert, M.; Dai, B.; Towers, M.; Perez-Vicente, R.; Willems, L.; Lahlou, S.; Pal, S.; Castro, P.S.; Terry, J.K. Minigrid & miniworld: Modular & customizable reinforcement learning environments for goal-oriented tasks. Adv. Neural Inf. Process. Syst. 2023, 36, 73383–73394. [Google Scholar] [CrossRef]
  79. Jay, N.; Nguyen, H.M.; Hoang, T.D.; Haimes, J. Evaluating precise geolocation inference capabilities of vision language models. arXiv 2025, arXiv:2502.14412. [Google Scholar]
  80. Yang, J.; Gao, S.; Qiu, Y.; Chen, L.; Li, T.; Dai, B.; Chitta, K.; Wu, P.; Zeng, J.; Luo, P.; et al. GenAD: Generalized Predictive Model for Autonomous Driving. arXiv 2024, arXiv:2403.09630. [Google Scholar]
  81. Mees, O.; Hermann, L.; Rosete-Beas, E.; Burgard, W. Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks. IEEE Robot. Autom. Lett. 2022, 7, 7327–7334. [Google Scholar] [CrossRef]
  82. Bellemare, M.G.; Naddaf, Y.; Veness, J.; Bowling, M. The arcade learning environment: An evaluation platform for general agents. J. Artif. Intell. Res. 2013, 47, 253–279. [Google Scholar] [CrossRef]
  83. Bain, M.; Nagrani, A.; Varol, G.; Zisserman, A. Frozen in time: A joint video and image encoder for end-to-end retrieval. In Proceedings of the Proceedings of the IEEE/CVF international conference on computer vision, 2021; pp. 1728–1738. [Google Scholar]
  84. Wang, L.; Yang, G.; Yang, L.; Zhang, X.; Song, Z.; Chen, Y.; Liu, L.; Gao, J.; Li, Z.; Yang, Q.; et al. S2R-bench: A sim-to-real evaluation benchmark for autonomous driving. Sci. Data 2025. [Google Scholar] [CrossRef] [PubMed]
  85. Wang, X.; Zhu, Z.; Huang, G.; Wang, B.; Chen, X.; Lu, J. Worlddreamer: Towards general world models for video generation via predicting masked tokens. arXiv 2024, arXiv:2401.09985. [Google Scholar]
  86. Wang, Y.; He, J.; Fan, L.; Li, H.; Chen, Y.; Zhang, Z. Driving into the future: Multiview visual forecasting and planning with world model for autonomous driving. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024; pp. 14749–14759. [Google Scholar]
  87. Zhao, G.; Wang, X.; Zhu, Z.; Chen, X.; Huang, G.; Bao, X.; Wang, X. Drivedreamer-2: Llm-enhanced world models for diverse driving video generation. In Proceedings of the AAAI Conference on Artificial Intelligence; 2025; Vol. 39, pp. 10412–10420. [Google Scholar]
  88. Guo, J.; Ding, Y.; Chen, X.; Chen, S.; Li, B.; Zou, Y.; Lyu, X.; Tan, F.; Qi, X.; Li, Z.; et al. Dist-4d: Disentangled spatiotemporal diffusion with metric depth for 4d driving scene generation. In Proceedings of the Proceedings of the IEEE/CVF International Conference on Computer Vision, 2025; pp. 27231–27241. [Google Scholar]
  89. Lu, J.; Huang, Z.; Yang, Z.; Zhang, J.; Zhang, L. Wovogen: World volume-aware diffusion for controllable multi-camera driving scene generation. In Proceedings of the European conference on computer vision, 2024; Springer; pp. 329–345. [Google Scholar]
  90. Mousakhan, A.; Mittal, S.; Galesso, S.; Farid, K.; Brox, T. Orbis: Overcoming challenges of long-horizon prediction in driving world models. arXiv 2025, arXiv:2507.13162. [Google Scholar]
  91. Zhang, K.; Tang, Z.; Hu, X.; Pan, X.; Guo, X.; Liu, Y.; Huang, J.; Yuan, L.; Zhang, Q.; Long, X.X.; et al. Epona: Autoregressive diffusion world model for autonomous driving. In Proceedings of the Proceedings of the IEEE/CVF International Conference on Computer Vision, 2025; pp. 27220–27230. [Google Scholar]
  92. Gao, R.; Chen, K.; Xie, E.; Hong, L.; Li, Z.; Yeung, D.Y.; Xu, Q. Magicdrive: Street view generation with diverse 3d geometry control. In Proceedings of the International Conference on Learning Representations; 2024; Vol. 2024, pp. 22841–22860. [Google Scholar]
  93. Gao, R.; Chen, K.; Xiao, B.; Hong, L.; Li, Z.; Xu, Q. MagicDrive-V2: High-resolution long video generation for autonomous driving with adaptive control. In Proceedings of the Proceedings of the IEEE/CVF International Conference on Computer Vision, 2025; pp. 28135–28144. [Google Scholar]
  94. Wen, Y.; Zhao, Y.; Liu, Y.; Jia, F.; Wang, Y.; Luo, C.; Zhang, C.; Wang, T.; Sun, X.; Zhang, X. Panacea: Panoramic and controllable video generation for autonomous driving. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024; pp. 6902–6912. [Google Scholar]
  95. Li, B.; Guo, J.; Liu, H.; Zou, Y.; Ding, Y.; Chen, X.; Zhu, H.; Tan, F.; Zhang, C.; Wang, T.; et al. Uniscene: Unified occupancy-centric driving scene generation. In Proceedings of the Proceedings of the computer vision and pattern recognition conference, 2025; pp. 11971–11981. [Google Scholar]
  96. Liang, A.; Liu, Y.; Yang, Y.; Lu, D.; Li, L.; Kong, L.; Zhao, H.; Ooi, W.T. LiDARCrafter: Dynamic 4D world modeling from LiDAR sequences. In Proceedings of the AAAI Conference on Artificial Intelligence; 2026; Vol. 40, pp. 18406–18414. [Google Scholar]
  97. Russell, L.; Hu, A.; Bertoni, L.; Fedoseev, G.; Shotton, J.; Arani, E.; Corrado, G. Gaia-2: A controllable multi-view generative world model for autonomous driving. arXiv 2025, arXiv:2503.20523. [Google Scholar]
  98. Li, X.; Wu, C.; Yang, Z.; Xu, Z.; Zhang, Y.; Liang, D.; Wan, J.; Wang, J. DriVerse: Navigation world model for driving simulation via multimodal trajectory prompting and motion alignment. In Proceedings of the Proceedings of the 33rd ACM International Conference on Multimedia, 2025; pp. 9753–9762. [Google Scholar]
  99. Garg, A.; Krishna, K.M. Imagine-2-Drive: Leveraging high-fidelity world models via multi-modal diffusion policies. In Proceedings of the 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS); IEEE, 2025; pp. 4188–4195. [Google Scholar]
  100. Liang, D.; Zhang, D.; Zhou, X.; Tu, S.; Feng, T.; Li, X.; Zhang, Y.; Du, M.; Tan, X.; Bai, X. Seeing the future, perceiving the future: A unified driving world model for future generation and perception. arXiv 2025, arXiv:2503.13587. [Google Scholar]
  101. Zhang, T.; Liu, Y.; Guo, Z.; Guo, Y.; Ni, J.; Ding, C.; Xu, D.; Lu, L.; Wu, Z. CVD-STORM: Cross-View Video Diffusion with Spatial-Temporal Reconstruction Model for Autonomous Driving. arXiv 2025, arXiv:2510.07944. [Google Scholar]
  102. Jia, F.; Mao, W.; Liu, Y.; Zhao, Y.; Wen, Y.; Zhang, C.; Zhang, X.; Wang, T. Adriver-i: A general world model for autonomous driving. arXiv 2023, arXiv:2311.13549. [Google Scholar]
  103. Jiang, C.; Cornman, A.; Park, C.; Sapp, B.; Zhou, Y.; Anguelov, D.; et al. Motiondiffuser: Controllable multi-agent motion prediction using diffusion. In Proceedings of the Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023; pp. 9644–9653. [Google Scholar]
  104. Wang, X.; Zhu, Z.; Huang, G.; Chen, X.; Zhu, J.; Lu, J. Drivedreamer: Towards real-world-drive world models for autonomous driving. In Proceedings of the European conference on computer vision, 2024; Springer; pp. 55–72. [Google Scholar]
  105. Jiang, C.M.; Bai, Y.; Cornman, A.; Davis, C.; Huang, X.; Jeon, H.; Kulshrestha, S.; Lambert, J.; Li, S.; Zhou, X.; et al. Scenediffuser: Efficient and controllable driving simulation initialization and rollout. Adv. Neural Inf. Process. Syst. 2024, 37, 55729–55760. [Google Scholar] [CrossRef]
  106. Shang, Y.; Lin, Y.; Zheng, Y.; Fan, H.; Ding, J.; Feng, J.; Chen, J.; Tian, L.; Li, Y. Urbanworld: An urban world model for 3d city generation. arXiv 2024, arXiv:2407.11965. [Google Scholar]
  107. Zhao, G.; Ni, C.; Wang, X.; Zhu, Z.; Zhang, X.; Wang, Y.; Huang, G.; Chen, X.; Wang, B.; Zhang, Y.; et al. Drivedreamer4d: World models are effective data machines for 4d driving scene representation. In Proceedings of the Proceedings of the computer vision and pattern recognition conference, 2025; pp. 12015–12026. [Google Scholar]
  108. Li, C.; Zhou, K.; Liu, T.; Wang, Y.; Zhuang, M.; Gao, H.a.; Jin, B.; Zhao, H. Avd2: Accident video diffusion for accident video description. In Proceedings of the 2025 IEEE International Conference on Robotics and Automation (ICRA); IEEE, 2025; pp. 13289–13296. [Google Scholar]
  109. Lu, Y.; Ren, X.; Yang, J.; Shen, T.; Wu, Z.; Gao, J.; Wang, Y.; Chen, S.; Chen, M.; Fidler, S.; et al. Infinicube: Unbounded and controllable dynamic 3d driving scene generation with world-guided video models. In Proceedings of the Proceedings of the IEEE/CVF International Conference on Computer Vision, 2025; pp. 27272–27283. [Google Scholar]
  110. Ni, J.; Guo, Y.; Liu, Y.; Chen, R.; Lu, L.; Wu, Z. Maskgwm: A generalizable driving world model with video mask reconstruction. In Proceedings of the Proceedings of the Computer Vision and Pattern Recognition Conference, 2025; pp. 22381–22391. [Google Scholar]
  111. Tan, S.; Lambert, J.; Jeon, H.; Kulshrestha, S.; Bai, Y.; Luo, J.; Anguelov, D.; Tan, M.; Jiang, C.M. SceneDiffuser++: City-scale traffic simulation via a generative world model. In Proceedings of the Proceedings of the Computer Vision and Pattern Recognition Conference, 2025; pp. 1570–1580. [Google Scholar]
  112. Li, J.; Zhang, B.; Jin, X.; Deng, J.; Zhu, X.; Zhang, L. ImagiDrive: A Unified Imagination-and-Planning Framework for Autonomous Driving. arXiv 2025, arXiv:2508.11428. [Google Scholar]
  113. Xia, T.; Li, Y.; Zhou, L.; Yao, J.; Xiong, K.; Sun, H.; Wang, B.; Ma, K.; Chen, G.; Ye, H.; et al. Drivelaw: Unifying planning and video generation in a latent driving world. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026; pp. 39701–39712. [Google Scholar]
  114. Tian, H.; Li, T.; Liu, H.; Yang, J.; Qiu, Y.; Li, G.; Wang, J.; Gao, Y.; Zhang, Z.; Wang, L.; et al. Simscale: Learning to drive via real-world simulation at scale. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026; pp. 36365–36374. [Google Scholar]
  115. Zhen, H.; Qiu, X.; Chen, P.; Yang, J.; Yan, X.; Du, Y.; Hong, Y.; Gan, C. 3d-vla: A 3d vision-language-action generative world model. arXiv 2024, arXiv:2403.09631. [Google Scholar]
  116. Rigter, M.; Gupta, T.; Hilmkil, A.; Ma, C. Avid: Adapting video diffusion models to world models. arXiv 2024, arXiv:2410.12822. [Google Scholar]
  117. Zhang, T.; Yu, H.X.; Wu, R.; Feng, B.Y.; Zheng, C.; Snavely, N.; Wu, J.; Freeman, W.T. Physdreamer: Physics-based interaction with 3d objects via video generation. In Proceedings of the European Conference on Computer Vision, 2024; Springer; pp. 388–406. [Google Scholar]
  118. Zhou, S.; Du, Y.; Chen, J.; Li, Y.; Yeung, D.Y.; Gan, C. Robodreamer: Learning compositional world models for robot imagination. arXiv 2024, arXiv:2404.12377. [Google Scholar]
  119. Chandra, A.L.; Nematollahi, I.; Huang, C.; Welschehold, T.; Burgard, W.; Valada, A. Diwa: Diffusion policy adaptation with world models. arXiv 2025, arXiv:2508.03645. [Google Scholar]
  120. Jang, J.; Ye, S.; Lin, Z.; Xiang, J.; Bjorck, J.; Fang, Y.; Hu, F.; Huang, S.; Kundalia, K.; Lin, Y.C.; et al. Dreamgen: Unlocking generalization in robot learning through video world models. arXiv 2025, arXiv:2505.12705. [Google Scholar]
  121. Huang, Y.; Zhang, J.; Zou, S.; Liu, X.; Hu, R.; Xu, K. LaDi-WM: A Latent Diffusion-based World Model for Predictive Manipulation. arXiv 2025, arXiv:2505.11528. [Google Scholar]
  122. Bar, A.; Zhou, G.; Tran, D.; Darrell, T.; LeCun, Y. Navigation world models. In Proceedings of the Proceedings of the Computer Vision and Pattern Recognition Conference, 2025; pp. 15791–15801. [Google Scholar]
  123. Huang, S.; Wu, J.; Zhou, Q.; Miao, S.; Long, M. Vid2world: Crafting video diffusion models to interactive world models. arXiv 2025, arXiv:2505.14357. [Google Scholar]
  124. Zhao, Q.; Lu, Y.; Kim, M.J.; Fu, Z.; Zhang, Z.; Wu, Y.; Li, Z.; Ma, Q.; Han, S.; Finn, C.; et al. Cot-vla: Visual chain-of-thought reasoning for vision-language-action models. In Proceedings of the Proceedings of the Computer Vision and Pattern Recognition Conference, 2025; pp. 1702–1713. [Google Scholar]
  125. Yue, Y.; Wang, Y.; Jiang, H.; Liu, P.; Song, S.; Huang, G. Echoworld: Learning motion-aware world models for echocardiography probe guidance. In Proceedings of the Proceedings of the Computer Vision and Pattern Recognition Conference, 2025; pp. 25993–26003. [Google Scholar]
  126. Lu, G.; Jia, B.; Li, P.; Chen, Y.; Wang, Z.; Tang, Y.; Huang, S. Gwm: Towards scalable gaussian world models for robotic manipulation. In Proceedings of the Proceedings of the IEEE/CVF International Conference on Computer Vision, 2025; pp. 9263–9274. [Google Scholar]
  127. Yu, Y.; Jin, X.; Shang, Y.; Zhang, X.; Su, H.; Wu, W.; Li, Y. Mowm: Mixture-of-world-models for embodied planning via latent-to-pixel feature modulation. arXiv 2025, arXiv:2509.21797. [Google Scholar]
  128. Tu, Y.; Luo, H.; Chen, X.; Bai, X.; Wang, F.; Zhao, H. Playerone: Egocentric world simulator. Adv. Neural Inf. Process. Syst. 2026, 38, 145235–145261. [Google Scholar]
  129. Huang, Y.; Jiang, X.; Gao, X.; Wu, M.; Tu, Z. VISTAv2: World Imagination for Indoor Vision-and-Language Navigation. arXiv 2025, arXiv:2512.00041. [Google Scholar]
  130. Bai, Y.; Tran, D.; Bar, A.; LeCun, Y.; Darrell, T.; Malik, J. Whole-body conditioned egocentric video prediction. Adv. Neural Inf. Process. Syst. 2026, 38, 164375–164418. [Google Scholar]
  131. Zhu, F.; Yan, Z.; Hong, Z.; Shou, Q.; Ma, X.; Guo, S. Wmpo: World model-based policy optimization for vision-language-action models. arXiv 2025, arXiv:2511.09515. [Google Scholar]
  132. Jiang, Z.; Liu, K.; Qin, Y.; Tian, S.; Zheng, Y.; Zhou, M.; Yu, C.; Li, H.; Zhao, D. World4rl: Diffusion world models for policy refinement with reinforcement learning for robotic manipulation. arXiv 2025, arXiv:2509.19080. [Google Scholar]
  133. Lu, T.; Shu, T.; Yuille, A.; Khashabi, D.; Chen, J. Genex: Generating an explorable world. In Proceedings of the International Conference on Learning Representations; 2025; Vol. 2025, pp. 52310–52335. [Google Scholar]
  134. Geng, D.; Herrmann, C.; Hur, J.; Cole, F.; Zhang, S.; Pfaff, T.; Lopez-Guevara, T.; Aytar, Y.; Rubinstein, M.; Sun, C.; et al. Motion prompting: Controlling video generation with motion trajectories. In Proceedings of the Proceedings of the Computer Vision and Pattern Recognition Conference, 2025; pp. 1–12. [Google Scholar]
  135. Chen, Y.; Chen, R.; Huo, D.; Yang, Y.; Qi, D.; Liu, H.; Lin, T.; Zeng, S.; Xiao, J.; Chang, X.; et al. Abot-physworld: Interactive world foundation model for robotic manipulation with physics alignment. arXiv 2026, arXiv:2603.23376. [Google Scholar]
  136. Ren, Z.; Wei, Y.; Yu, X.; Luo, G.; Zhao, Y.; Kang, B.; Feng, J.; Jin, X. Videoworld 2: Learning transferable knowledge from real-world videos. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026; pp. 40569–40580. [Google Scholar]
  137. Cai, S.; Chan, E.R.; Peng, S.; Shahbazi, M.; Obukhov, A.; Van Gool, L.; Wetzstein, G. Diffdreamer: Towards consistent unsupervised single-view scene extrapolation with conditional diffusion models. In Proceedings of the Proceedings of the IEEE/CVF international conference on computer vision, 2023; pp. 2139–2150. [Google Scholar]
  138. Ren, W.; Yang, H.; Zhang, G.; Wei, C.; Du, X.; Huang, W.; Chen, W. Consisti2v: Enhancing visual consistency for image-to-video generation. arXiv 2024, arXiv:2402.04324. [Google Scholar]
  139. Alonso, E.; Jelley, A.; Micheli, V.; Kanervisto, A.; Storkey, A.; Pearce, T.; Fleuret, F. Diffusion for world modeling: Visual details matter in atari. Adv. Neural Inf. Process. Syst. 2024, 37, 58757–58791. [Google Scholar] [CrossRef]
  140. Jain, Y.; Nasery, A.; Vineet, V.; Behl, H. Peekaboo: Interactive video generation via masked-diffusion. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024; pp. 8079–8088. [Google Scholar]
  141. Ge, Z.; Huang, H.; Zhou, M.; Li, J.; Wang, G.; Tang, S.; Zhuang, Y. Worldgpt: Empowering llm as multimodal world model. In Proceedings of the Proceedings of the 32nd ACM International Conference on Multimedia, 2024; pp. 7346–7355. [Google Scholar]
  142. Xiang, J.; Liu, G.; Gu, Y.; Gao, Q.; Ning, Y.; Zha, Y.; Feng, Z.; Tao, T.; Hao, S.; Shi, Y.; et al. Pandora: Towards general world model with natural language actions and video states. arXiv 2024, arXiv:2406.09455. [Google Scholar]
  143. Zhang, J.; Tang, J.; Zhang, R.; Lv, T.; Sun, X. Storyweaver: A unified world model for knowledge-enhanced story character customization. In Proceedings of the Proceedings of the AAAI Conference on Artificial Intelligence; 2025; Vol. 39, pp. 9951–9959. [Google Scholar]
  144. Parker-Holder, J.; Ball, P.; Bruce, J.; Dasagi, V.; Holsheimer, K.; Kaplanis, C.; Moufarek, A.; Scully, G.; Shar, J.; Shi, J.; et al. Genie 2: A large-scale foundation world model. 2024, 2. https://deepmind.google/discover/blog/genie-2-a-large-scale-foundation-world-model.
  145. Valevski, D.; Leviathan, Y.; Arar, M.; Fruchter, S. Diffusion models are real-time game engines. In Proceedings of the International Conference on Learning Representations; 2025; Vol. 2025, pp. 73754–73776. [Google Scholar]
  146. Hassan, M.; Stapf, S.; Rahimi, A.; Rezende, P.; Haghighi, Y.; Brüggemann, D.; Katircioglu, I.; Zhang, L.; Chen, X.; Saha, S.; et al. Gem: A generalizable ego-vision multimodal world model for fine-grained ego-motion, object dynamics, and scene composition control. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025; pp. 22404–22415. [Google Scholar]
  147. Gao, J.; Chen, Z.; Liu, X.; Zhuang, J.; Xu, C.; Feng, J.; Qiao, Y.; Fu, Y.; Si, C.; Liu, Z. LongVie 2: Multimodal Controllable Ultra-Long Video World Model. arXiv 2025, arXiv:2512.13604. [Google Scholar]
  148. He, X.; Zhou, S.; Venkateswaran, T.; Zheng, K.; Wan, Z.; Kadambi, A.; Wang, X.E. MorphoSim: An Interactive, Controllable, and Editable Language-guided 4D World Simulator. arXiv 2025, arXiv:2510.04390. [Google Scholar]
  149. Zhou, Y.; Wang, Y.; Zhou, J.; Chang, W.; Guo, H.; Li, Z.; Ma, K.; Li, X.; Wang, Y.; Zhu, H.; et al. Omniworld: A multi-domain and multi-modal dataset for 4d world modeling. arXiv 2025, arXiv:2509.12201. [Google Scholar]
  150. Huang, T.; Zheng, W.; Wang, T.; Liu, Y.; Wang, Z.; Wu, J.; Jiang, J.; Li, H.; Lau, R.; Zuo, W.; et al. Voyager: Long-range and world-consistent video diffusion for explorable 3d scene generation. ACM Trans. Graph. (TOG) 2025, 44, 1–15. [Google Scholar] [CrossRef]
  151. Mao, X.; Lin, S.; Li, Z.; Li, C.; Peng, W.; He, T.; Pang, J.; Chi, M.; Qiao, Y.; Zhang, K. Yume: An interactive world generation model. arXiv 2025, arXiv:2507.17744. [Google Scholar]
  152. Zhu, Y.; Feng, J.; Zheng, W.; Gao, Y.; Tao, X.; Wan, P.; Zhou, J.; Lu, J. Astra: General Interactive World Model with Autoregressive Denoising. arXiv 2025, arXiv:2512.08931. [Google Scholar]
  153. Dai, Y.; Jiang, F.; Wang, C.; Xu, M.; Qi, Y. Fantasyworld: Geometry-consistent world modeling via unified video and 3d prediction. arXiv 2025, arXiv:2509.21657. [Google Scholar]
  154. Tian, J.; Song, C.; Cheng, W.; Zhang, C. Free-lunch long video generation via layer-adaptive ood correction. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026; pp. 1973–1982. [Google Scholar]
  155. Yang, Y.; Fan, L.; Shi, Z.; Peng, J.; Wang, F.; Zhang, Z. NeoVerse: Enhancing 4D World Model with in-the-wild Monocular Videos. arXiv 2026, arXiv:2601.00393. [Google Scholar]
  156. Wang, Z.; Hu, P.; Wang, J.; Zhang, T.J.; Cheng, Y.; Chen, L.; Yan, Y.; Jiang, Z.; Li, H.; Liang, X. Prophy: Progressive physical alignment for dynamic world simulation. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026; pp. 14492–14501. [Google Scholar]
  157. Zheng, S.; Yin, M.; Hu, W.; Li, X.; Shan, Y.; Fu, Y. Versecrafter: Dynamic realistic video world model with 4d geometric control. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026; pp. 40277–40290. [Google Scholar]
  158. Song, C.; Yang, Y.; Zhao, T.; Li, R.; Zhang, C. Taming Video Models for 3D and 4D Generation via Zero-Shot Camera Control. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026; pp. 40352–40363. [Google Scholar]
  159. Black, K.; Brown, N.; Driess, D.; Esmail, A.; Equi, M.; Finn, C.; Fusai, N.; Groom, L.; Hausman, K.; Ichter, B.; et al. π0: A Vision-Language-Action Flow Model for General Robot Control. arXiv 2024, arXiv:2410.24164. [Google Scholar]
  160. Intelligence, P.; Black, K.; Brown, N.; Darpinian, J.; Dhabalia, K.; Driess, D.; Esmail, A.; Equi, M.; Finn, C.; Fusai, N.; et al. π0.5: A Vision-Language-Action Model with Open-World Generalization. arXiv 2025, arXiv:2504.16054. [Google Scholar]
  161. Savov, N.; Kazemi, N.; Zhang, D.; Paudel, D.P.; Wang, X.; Gool, L.V. Statespacediffuser: Bringing long context to diffusion world models. Adv. Neural Inf. Process. Syst. 2026, 38, 68865–68898. [Google Scholar]
  162. Po, R.; Nitzan, Y.; Zhang, R.; Chen, B.; Dao, T.; Shechtman, E.; Wetzstein, G.; Huang, X. Long-context state-space video world models. In Proceedings of the Proceedings of the IEEE/CVF International Conference on Computer Vision, 2025; pp. 8733–8744. [Google Scholar]
  163. Li, L.; Fan, Z.; Cong, W.; Liu, X.; Yin, Y.; Foutter, M.; Pan, P.; You, C.; Wang, Y.; Wang, Z.; et al. Martian world model: Controllable video synthesis with physically accurate 3d reconstructions. Adv. Neural Inf. Process. Syst. 2026, 38. [Google Scholar]
  164. Kang, B.; Yue, Y.; Lu, R.; Lin, Z.; Zhao, Y.; Wang, K.; Huang, G.; Feng, J. How far is video generation from world model: A physical law perspective. arXiv 2024, arXiv:2411.02385. [Google Scholar]
  165. Voleti, V.; Yao, C.H.; Boss, M.; Letts, A.; Pankratz, D.; Tochilkin, D.; Laforte, C.; Rombach, R.; Jampani, V. Sv3d: Novel multi-view synthesis and 3d generation from a single image using latent video diffusion. In Proceedings of the European Conference on Computer Vision, 2024; Springer; pp. 439–457. [Google Scholar]
  166. Shang, Y.; Jin, L.; Ma, Y.; Zhang, X.; Gao, C.; Wu, W.; Li, Y. LongScape: Advancing Long-Horizon Embodied World Models with Context-Aware MoE. arXiv 2025, arXiv:2509.21790. [Google Scholar]
  167. Zhang, Y.; Gong, S.; Xiong, K.; Ye, X.; Li, X.; Tan, X.; Wang, F.; Huang, J.; Wu, H.; Wang, H. BEVWorld: A multimodal world simulator for autonomous driving via scene-level BEV latents. arXiv 2024, arXiv:2407.05679. [Google Scholar]
  168. Guo, X.; Ding, C.; Dou, H.; Zhang, X.; Tang, W.; Wu, W. Infinitydrive: Breaking time limits in driving world models. arXiv 2024, arXiv:2412.01522. [Google Scholar]
  169. Ding, Z.; Zhang, A.; Tian, Y.; Zheng, Q. Diffusion world model: Future modeling beyond step-by-step rollout for offline reinforcement learning. arXiv 2024, arXiv:2402.03570. [Google Scholar]
  170. Ha, D.; Schmidhuber, J. Recurrent world models facilitate policy evolution. Adv. Neural Inf. Process. Syst. 2018, 31. [Google Scholar]
  171. Chu, Z.; Zhang, L.; Sun, Y.; Xue, S.; Wang, Z.; Qin, Z.; Ren, K. Sora detector: A unified hallucination detection for large text-to-video models. arXiv 2024, arXiv:2405.04180. [Google Scholar]
  172. Chen, X.; Song, D.; Gui, H.; Wang, C.; Zhang, N.; Jiang, Y.; Huang, F.; Lv, C.; Zhang, D.; Chen, H. Factchd: Benchmarking fact-conflicting hallucination detection. arXiv 2023, arXiv:2310.12086. [Google Scholar]
  173. Sanchez, P.; Tsaftaris, S.A. Diffusion causal models for counterfactual estimation. arXiv 2022, arXiv:2202.10166. [Google Scholar]
  174. Janner, M.; Du, Y.; Tenenbaum, J.B.; Levine, S. Planning with diffusion for flexible behavior synthesis. arXiv 2022, arXiv:2205.09991. [Google Scholar]
  175. Song, J.; Meng, C.; Ermon, S. Denoising diffusion implicit models. arXiv 2020, arXiv:2010.02502. [Google Scholar]
  176. Chi, C.; Xu, Z.; Feng, S.; Cousineau, E.; Du, Y.; Burchfiel, B.; Tedrake, R.; Song, S. Diffusion policy: Visuomotor policy learning via action diffusion. Int. J. Robot. Res. 2025, 44, 1684–1704. [Google Scholar] [CrossRef]
  177. Lu, C.; Zhou, Y.; Bao, F.; Chen, J.; Li, C.; Zhu, J. Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps. Adv. Neural Inf. Process. Syst. 2022, 35, 5775–5787. [Google Scholar] [CrossRef]
  178. Lu, C.; Zhou, Y.; Bao, F.; Chen, J.; Li, C.; Zhu, J. Dpm-solver++: Fast solver for guided sampling of diffusion probabilistic models. Mach. Intell. Res. 2025, 22, 730–751. [Google Scholar] [CrossRef]
  179. Zheng, K.; Lu, C.; Chen, J.; Zhu, J. Dpm-solver-v3: Improved diffusion ode solver with empirical model statistics. Adv. Neural Inf. Process. Syst. 2023, 36, 55502–55542. [Google Scholar] [CrossRef]
  180. Wang, F.Y.; Zhou, H.; Yuan, L.; Woo, S.; Gong, B.; Han, B.; Yang, M.H.; Zhang, H.; Zhu, Y.; Liu, T.; et al. Image Diffusion Preview with Consistency Solver. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026; pp. 43271–43280. [Google Scholar]
  181. Song, Y.; Dhariwal, P.; Chen, M.; Sutskever, I. Consistency Models. In Proceedings of the Proceedings of the 40th International Conference on Machine Learning; PMLR; Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., Scarlett, J., Eds.; Proceedings of Machine Learning Research, 23–29 Jul 2023; Vol. 202, pp. 32211–32252. [Google Scholar]
  182. Chai, W.; Zheng, D.; Cao, J.; Chen, Z.; Wang, C.; Ma, C. SpeedUpNet: A Plug-and-Play Adapter Network for Accelerating Text-to-Image Diffusion Models. In Proceedings of the European Conference on Computer Vision, 2024; Springer; pp. 181–196. [Google Scholar]
  183. Blattmann, A.; Dockhorn, T.; Kulal, S.; Mendelevitch, D.; Kilian, M.; Lorenz, D.; Levi, Y.; English, Z.; Voleti, V.; Letts, A.; et al. Stable video diffusion: Scaling latent video diffusion models to large datasets. arXiv 2023, arXiv:2311.15127. [Google Scholar]
  184. Wang, R.; He, K. Diffuse and disperse: Image generation with representation regularization. arXiv 2025, arXiv:2506.09027. [Google Scholar]
  185. Rigter, M.; Yamada, J.; Posner, I. World models via policy-guided trajectory diffusion. arXiv 2023, arXiv:2312.08533. [Google Scholar]
  186. Berdica, U.; Li, K.; Beukman, M.; Goldie, A.D.; Fellows, M.; Maiolino, P.; Foerster, J.N. Investigating Online RL in World Models.
  187. Yu, X.; Peng, B.; Xu, R.; Shen, Y.; He, P.; Nath, S.; Singh, N.; Gao, J.; Yu, Z. Reinforcement World Model Learning for LLM-based Agents. arXiv 2026, arXiv:2602.05842. [Google Scholar]
  188. Wu, J.; Yin, S.; Feng, N.; Long, M. Rlvr-world: Training world models with reinforcement learning. Adv. Neural Inf. Process. Syst. 2026, 38, 125312–125350. [Google Scholar]
  189. Hansen, N.; Su, H.; Wang, X. Td-mpc2: Scalable, robust world models for continuous control. In Proceedings of the International Conference on Learning Representations; 2024; Vol. 2024, pp. 47376–47405. [Google Scholar]
  190. Reed, S.; Zolna, K.; Parisotto, E.; Colmenarejo, S.G.; Novikov, A.; Barth-Maron, G.; Gimenez, M.; Sulsky, Y.; Kay, J.; Springenberg, J.T.; et al. A generalist agent. arXiv 2022, arXiv:2205.06175. [Google Scholar]
  191. Li, C.; Krause, A.; Hutter, M. Uncertainty-Aware Robotic World Model Makes Offline Model-Based Reinforcement Learning Work on Real Robots. arXiv 2025, arXiv:2504.16680. [Google Scholar]
  192. Brohan, A.; Brown, N.; Carbajal, J.; Chebotar, Y.; Chen, X.; Choromanski, K.; Ding, T.; Driess, D.; Dubey, A.; Finn, C.; et al. Rt-2: Vision-language-action models transfer web knowledge to robotic control, 2023. arXiv 2024, arXiv:2307.158181, 2. [Google Scholar]
  193. Janner, M.; Fu, J.; Zhang, M.; Levine, S. When to trust your model: Model-based policy optimization. Adv. Neural Inf. Process. Syst. 2019, 32. [Google Scholar]
  194. Wang, Y.; Syed, R.; Wu, F.; Zhang, M.; Onol, A.; Barreiros, J.; Nayyeri, H.; Dear, T.; Zhang, H.; Li, Y. Interactive world simulator for robot policy training and evaluation. arXiv 2026, arXiv:2603.08546. [Google Scholar]
  195. He, H.; Zhang, Y.; Lin, L.; Xu, Z.; Pan, L. Pre-trained video generative models as world simulators. In Proceedings of the AAAI Conference on Artificial Intelligence; 2026; Vol. 40, pp. 4645–4653. [Google Scholar]
  196. Li, C.; Krause, A.; Hutter, M. Robotic world model: A neural network simulator for robust policy optimization in robotics. arXiv 2025, arXiv:2501.10100. [Google Scholar]
  197. of Robotics, I.F. Top Global Robotics Trends 2026, 2026.
  198. Mildenhall, B.; Srinivasan, P.P.; Tancik, M.; Barron, J.T.; Ramamoorthi, R.; Ng, R. Nerf: Representing scenes as neural radiance fields for view synthesis. Commun. ACM 2021, 65, 99–106. [Google Scholar]
  199. Kirillov, A.; Mintun, E.; Ravi, N.; Mao, H.; Rolland, C.; Gustafson, L.; Xiao, T.; Whitehead, S.; Berg, A.C.; Lo, W.Y.; et al. Segment anything. In Proceedings of the Proceedings of the IEEE/CVF international conference on computer vision, 2023; pp. 4015–4026. [Google Scholar]
  200. Tancik, M.; Casser, V.; Yan, X.; Pradhan, S.; Mildenhall, B.; Srinivasan, P.P.; Barron, J.T.; Kretzschmar, H. Block-nerf: Scalable large scene neural view synthesis. In Proceedings of the Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022; pp. 8248–8258. [Google Scholar]
  201. Driess, D.; Xia, F.; Sajjadi, M.S.; Lynch, C.; Chowdhery, A.; Ichter, B.; Wahid, A.; Tompson, J.; Vuong, Q.; Yu, T.; et al. Palm-e: An embodied multimodal language model. arXiv 2023, arXiv:2303.03378. [Google Scholar]
  202. Kerbl, B.; Kopanas, G.; Leimkühler, T.; Drettakis, G.; et al. 3d gaussian splatting for real-time radiance field rendering. ACM Trans. Graph. 2023, 42, 139–1. [Google Scholar] [CrossRef]
  203. Turki, H.; Zhang, J.Y.; Ferroni, F.; Ramanan, D. Suds: Scalable urban dynamic scenes. In Proceedings of the Proceedings of the ieee/cvf conference on computer vision and pattern recognition, 2023; pp. 12375–12385. [Google Scholar]
  204. Zha, D.; Bhat, Z.P.; Lai, K.H.; Yang, F.; Jiang, Z.; Zhong, S.; Hu, X. Data-centric artificial intelligence: A survey. ACM Comput. Surv. 2025, 57, 1–42. [Google Scholar] [CrossRef]
  205. Jakubik, J.; Vössing, M.; Kühl, N.; Walk, J.; Satzger, G. Data-Centric Artificial Intelligence: J. Jakubik et al. Bus. Inf. Syst. Eng. 2024, 66, 507–515. [Google Scholar]
  206. Kaiser, L.; Babaeizadeh, M.; Milos, P.; Osinski, B.; Campbell, R.H.; Czechowski, K.; Erhan, D.; Finn, C.; Kozakowski, P.; Levine, S.; et al. Model-based reinforcement learning for atari. arXiv 2019, arXiv:1903.00374. [Google Scholar]
  207. Zha, D.; Bhat, Z.P.; Lai, K.H.; Yang, F.; Hu, X. Data-centric ai: Perspectives and challenges. In Proceedings of the Proceedings of the 2023 SIAM international conference on data mining (SDM), 2023; SIAM; pp. 945–948. [Google Scholar]
  208. Team, G.; Ye, A.; Wang, B.; Ni, C.; Huang, G.; Zhao, G.; Li, H.; Zhu, J.; Li, K.; Xu, M.; et al. Gigaworld-0: World models as data engine to empower embodied ai. arXiv 2025, arXiv:2511.19861. [Google Scholar]
  209. Zhang, J.; Jiang, M.; Dai, N.; Lu, T.; Uzunoglu, A.; Zhang, S.; Wei, Y.; Wang, J.; Patel, V.M.; Liang, P.P.; et al. World-in-world: World models in a closed-loop world. arXiv 2025, arXiv:2510.18135. [Google Scholar]
  210. Liu, G.; Deng, Y.; Liu, Z.; Jia, K. GS-World: An Efficient, Engine-driven Learning Paradigm for Pursuing Embodied Intelligence using World Models of Generative Simulation. Authorea Preprints 2025. [Google Scholar]

Short Biography of Authors

Preprints 233002 i002 Gang Wang received the Ph. D degree in computer science from Beihang University, Beijing, China, in 2018. He was a visiting scholar with Imperial College London, London, UK, from September 2022 to September 2023. He is currently an Associate Professor with the School of Computing and Data Engineering, NingboTech University, Ningbo, China. His current research interests include image/video processing, multimedia communication, machine learning and hardware optimization.
Preprints 233002 i003 Zhen Liu received the BS degree in electronic information engineering with Hangzhou Dianzi University, Hangzhou, China in 2025. Now, he is currently working toward the MS degree in computer technology, Zhejiang Sci-Tech University. His research interests include machine learning and computer vision.
Preprints 233002 i004 Mingliang Zhou received a Ph.D. degree in computer science from Beihang University, Beijing, China, in 2017. He was a Postdoctoral researcher with the Department of Computer Science, City University of Hong Kong, Hong Kong, China, from September 2017 to September 2019. He was a Postdoctoral Fellow with the State Key Lab of Internet of Things for Smart City, University of Macau, Macau, China, from October 2019 to October 2021. He is currently an Associate Professor with the School of Computer Science, Chongqing University, Chongqing, China. His research interests include image and video coding, perceptual image processing, multimedia signal processing, rate control, multimedia communication, machine learning, and optimization.
Preprints 233002 i005 Ziying Song received the B.S. degree from Hebei Normal University of Science and Technology, China, in 2019, the M.S. degree from Hebei University of Science and Technology, China, in 2022, and the Ph.D. degree in computer science and technology from Beijing Jiaotong University, China, in March 2026. He is currently an Assistant Professor with the School of Artificial Intelligence, Yanshan University, China. His research interests include autonomous driving, end-to-end autonomous driving, world models, embodied intelligence, and VLA models.
Preprints 233002 i006 Yugui Zhang is currently an assistant researcher at the Institute of Semiconductors, Chinese Academy of Sciences. He received his master’s and doctoral degrees from Beihang University and his bachelor’s degree from Xinjiang University. His research interests include image enhancement, object detection, person reidentification, human pose estimation, and model compression and acceleration.
Preprints 233002 i007 Lei Yang (Member, IEEE) received the M.S. degree from the Robotics Institute, Beihang University, China, in 2018, and the Ph.D. degree from the School of Vehicle and Mobility, Tsinghua University, China, in 2024. From 2018 to 2020, he joined the Autonomous Driving R&D Department of JD.COM as an algorithm researcher. Currently, he is a research fellow with the School of Mechanical and Aerospace Engineering, Nanyang Technological University, Singapore. His current research interests include autonomous driving, 3D scene understanding and world model.
Preprints 233002 i008 Yuanyan Tang (Life Fellow, IEEE) is currently the Chair Professor of the Faculty of Science and Technology, University of Macau, Macau, China, and a Professor/an Adjunct Professor/an Honorary Professor at several institutes, including Chongqing University, Chongqing, China, Concordia University, Montreal, QC, Canada, and Hong Kong Baptist University, Hong Kong, China. His current interests include wavelets, pattern recognition, and image processing. He has published more than 400 academic articles and is the author/co-author of more than 25 monographs/books/book chapters. Dr. Tang is an IAPR Fellow. He is the Founder and the Editor-in-Chief of International Journal of Wavelets, Multiresolution and Information Processing (IJWMIP), and an associate editor of several international journals. He is the Founder and the Chair of the Pattern Recognition Committee in IEEE SMC. He served as the general chair, the program chair, and a committee member for many international conferences. He is the Founder and the General Chair of the Series International Conferences on Wavelets Analysis and Pattern Recognition (ICWAPRs). He is the Founder and the Chair of Macau Branch of the International Association of Pattern Recognition (IAPR).
Preprints 233002 i009 Zheng Zhu is currently the Co-founder and Chief Scientist at GigaAI. During 2019-2021, he was a post-doc fellow at Tsinghua University. Before that, he received Ph.D. degree from Institute of Automation, Chinese Academy of Sciences in 2019. During 2016-2019, he was research interns at SenseTime, Horizon Robotics, and DeepGlint. He has co-authored more than 70 top journal and conference papers mainly on computer vision and robotics problems, such as world model, AIGC, autonomous driving, face recognition, and visual tracking. He has more than 20,000 Google Scholar citations to his work, including SiamRPN (3,600+ citations), DaSiamRPN (1,800+ citations) and BEVDet (1,200+ citations). His work DaSiamRPN is included in famous OpenCV Library. He was recognised as one of the Top 2% Scientists Worldwide by Stanford University (2022, 2023, 2024, 2025). He won the Best Student Paper Award and CCF Outstanding Paper Award in PRCV 2025 (CCF-C Conference). He ranked the 1st on NIST-FRVT Masked Face Recognition, won the COCO Keypoint Detection Challenge in ECCV 2020 and Visual Object Tracking (VOT) Real-Time Challenge in ECCV 2018. He was the Area Chair of ICLR, NeurIPS, CoRL and AAAI.
Preprints 233002 i010 Lin Gu is a research scientist with RIKEN AIP, Japan with specific interest and expertise in the application of Artificial Intelligence in Medical Imaging and Computational Photography. Before moving to Japan, he was a postdoctoral research fellow with A*STAR, Singapore working on machine learning on biomedical imaging. His research interests include simulating human’s neocortex to enhance AI and especially its application on medical fields.
Preprints 233002 i011 Guang Yang (Senior Member, IEEE) received an M.Sc. degree in vision imaging and virtual environments from the Department of Computer Science, in 2006, and a Ph.D. degree in medical image analysis jointly from the CMIC, Department of Computer Science and Medical Physics, in 2012, both from University College London. He is currently an Advanced Research Fellow with NHLI, Imperial College London, and a Honourary Senior Lecturer with the School of Biomedical Engineering and Imaging Sciences, King’s College London. He is the Head of the Smart Imaging Laboratory funded by UKRI, BHF and ERC. His research interests include pattern recognition, machine learning, and medical image processing and analysis.
Figure 3. Statistical Milestones in Diffusion World Models
Figure 3. Statistical Milestones in Diffusion World Models
Preprints 233002 g003
Table 1. Dataset for three major fields.Env:ENVIRONMENTS,Traj:TRAJECTORIES
Table 1. Dataset for three major fields.Env:ENVIRONMENTS,Traj:TRAJECTORIES
Dataset Year Domain Data Type Env Data Size Task/Skill/Content Notes URL
Waymo Open Perception dataset[38] 2019 AD Lidar, Image, etc 1.15K Scenes 230K Frames Object Detection, etc 2D, 3D, etc Preprints 233002 i001
Argoverse 2[39] 2019 AD Lidar, Traj, Map, etc 6 Cities 1K Sequences, 6M Lidar Frames, etc 3D Object Detection, etc 3D, etc Preprints 233002 i001
D4RL[40] 2020 AD Traj, Image, etc 7 Env, e.g., Mazes, Kitchens, etc. 1M steps, 25 Traj – – Preprints 233002 i001
nuScenes[41] 2020 AD Lidar, radar, Image, Traj, etc 1K Scenes 1.4M Images, 40K Keyframes, etc 3D Object Detection/Tracking, etc Scene Descriptions, 3D, etc Preprints 233002 i001
Nuplan[42] 2021 AD Lidar, Image, Traj, etc 4 Cities 1.5K H Open–loop, Closed–loop, Metrics Object Tracks, etc Preprints 233002 i001
Waymo Open Motion dataset[43][44] 2021 AD Lidar, Traj, Map, etc 6 Cities, 100K Scenes 7.64M Traj, 574H Motion Forecasting Speed, 3D, etc Preprints 233002 i001Preprints 233002 i001
CODA[45] 2022 AD Image, Lidar 3 countries, 1.5K Scenes 1.5K Images Object Detection, etc Bounding Boxes Preprints 233002 i001
NAVSIM[46] 2024 AD Lidar, Image, Map, etc 396 independent scenes 115K Samples, 120H End–to–End Planning 3D Preprints 233002 i001
OpenDV–2K[47] 2024 AD Video, Text, Traj 40 Cities 65.1M Frames, 2059H Planning, Prediction, etc Language, Traj, etc Preprints 233002 i001
GS [48][49] 2018 EI Video, Image, etc 42 Buildings 255K Images, 25.2H Navigation – Preprints 233002 i001
EPIC–KITCHENS[50] 2018 EI Video, Action Narrations Kitchen 432 Sequences/55 H 323 Tasks, e.g., Cook, Clean Egocentric Preprints 233002 i001
Robotic Pushing[51] 2018 EI Video, Gripper Pose Laboratory 59k Push – Preprints 233002 i001
Howto100M[52] 2019 EI Video, Annotation 12 Env., e.g., Home, Garden 136M 23k Tasks, e.g., Cook, Mark – Preprints 233002 i001
RECON[53] 2021 EI Traj, Video, etc 9 Env., e.g., Parking Lot, Cafeteria 5K+ Traj, 100H+ Navigation – Preprints 233002 i001
SCAND[54] 2022 EI Lidar, Traj, Text, Video, etc Outdoor, Indoor 138 Traj, 8.7H Videos, etc – 12 Types, e.g., Driving with traffic Preprints 233002 i001
Ego4D[55] 2022 EI Video, Annotation, Audio Dailylife, Outdoor, Indoor 3,670H Diverse Tasks, e.g., flip, lift 4D, Egocentric Preprints 233002 i001
RT–1[56] 2022 EI Video Office, Kitchen 130K 744 Tasks, e.g., Pick, Move – Preprints 233002 i001
Open X–Embodiment[57] 2022 EI Video, Depth, 3D Information 311 Scenes, e.g., Household 1M+ Traj 527 Skills, 160266 Tasks, e.g., Pick 3D Preprints 233002 i001
Language Table[58] 2022 EI Video Simulated & Real–world Tabletop 600k Traj Diverse tasks/skills, e.g., Push – Preprints 233002 i001
Bridge Data[59] 2022 EI Video Toy & Real Kitchens, Toy Sinks 7.2K Traj 71 Tasks, e.g., Flip, Put – Preprints 233002 i001
TartanDrive[60] 2022 EI Image, Traj, Map, etc Outdoor 184K Samples, 630 Traj, 5H Dynamics Prediction – Preprints 233002 i001
LIBERO[61] 2023 EI Language instructions, Traj, etc Kitchen, Tabletop, etc. 6.5K Traj 130 Tasks – Preprints 233002 i001
BridgeData V2[62] 2023 EI video 24 Env., e.g., Kitchens, Tabletops 60.1K Traj 13 Tasks, e.g., Sweep – Preprints 233002 i001
CALVIN[63] 2024 EI Traj, Language Instructions, etc Tabletop 24K Demos 34 Tasks, e.g., Opening – Preprints 233002 i001
RoboMind[64] 2025 EI Text, Traj, etc 5 Env, e.g., Home, Kitchen 107K Traj,305.5H 479 Tasks, e.g., Grab, Open Multi–view, Language Preprints 233002 i001
EgoDex[65] 2025 EI Video,Text,Traj, etc Tabletop 90M Frames, 829H 194 Tasks, e.g.,Washing a cup Language, etc Preprints 233002 i001
RoboTwin2.0[66] 2025 EI Video,Text,Traj 5 Env 100K+ Traj 50 Tasks, e.g., Stack Bowls Semantic, etc Preprints 233002 i001
UCF101[67] 2012 GD Video,Audio 101 Different Actions 13,320 clips, 27H Action Recognition – Preprints 233002 i001
DAVIS[68] 2016 GD Video 4 Env 50 Video Sequences Object Segmentation Challenge Attribute Preprints 233002 i001
MSR–VTT[69] 2016 GD Video,Text,Audio 20 Different Scenarios 10K clips, 41.2H Video to Text Natural Language Preprints 233002 i001
SSv2[70] 2017 GD Video,Text,Audio Daily Env 108,499 short videos 174 categories, e.g.,Pick up Structured Captions, etc Preprints 233002 i001
LaSOT[71][72] 2019 GD Video,Text 14 Attributes, e.g., Out–of–View 1.55K Videos Single Object Tracking Bounding Box, etc Preprints 233002 i001Preprints 233002 i001
LLFF[73] 2019 GD Image 8 real–world scenarios – view synthesis – Preprints 233002 i001
Procgen[74] 2020 GD Image 16 unique environments – 16 different games – Preprints 233002 i001
Mip–NeRF 360[75] 2022 GD Video, Image, etc 9 Scenarios – View Synthesis – Preprints 233002 i001
Minecraft&Tennis [76] 2022 GD Image, Video, etc Tennis fields, Game 13H Playable 3D Video Generation, etc 2D, Camera Poses, etc Preprints 233002 i001
FlintstonesHD[77] 2023 GD Video,Text – 166 Episodes of Cartoon Series Long Video Generation – Preprints 233002 i001
MiniGrid&Miniworld [78] 2023 GD Image, Text, etc 2D, 3D – Pick up,GoToObject, etc – Preprints 233002 i001
Google Street View[79] 2025 GD Image 88 countries 1602 Images Geolocation Inference, etc Geographic Coordinates, etc Preprints 233002 i001
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.