Preprint
Article

This version is not peer-reviewed.

ThunderWorld: Efficient World Model Generation with Selective FP4 Refinement via Native GEMMs

Submitted:

05 October 2026

Posted:

09 October 2026

You are already at the latest version

Abstract
We study whether world models can recover high-quality generation through dense four-bit floating-point (FP4) matrix computation without sparse matrix instructions or token/channel reordering. ThunderWorld is a training-free framework that corrects four-bit weight-and-activation (W4A4) error using additional dense FP4 products formed from quantization residuals. Block-aligned supports jointly govern residual storage, online preparation, and dense execution under a compute budget, preserving every base attention block. Controlled attention diagnostics support structure-informed selection over random and window policies at matched repair computation. On Cosmos3 Edge and Nano across RBench and PAI-Bench-G, a configuration capped at twice base W4A4 matrix work keeps mean task-score gaps within 0.08 RBench points and 0.37 PAI-Bench-G percentage points of BF16. It reduces LPIPS, a perceptual distance to same-input, same-seed BF16 generations, by 3.3-7.8% versus ARCQuant (2.0), with 1.17-1.88x its throughput on RTX PRO 6000 and Jetson AGX Thor. These results provide a numerical and systems basis for world-model inference on devices with native FP4 Tensor Cores.
Keywords: 
;  ;  ;  ;  ;  ;  
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.