Preprint
Article

This version is not peer-reviewed.

Kinematics-Gated Post-Training Quantization: A Biomechanical Feedback Loop for Mixed-Precision Compression of a Video-Diffusion Transformer

Submitted:

12 August 2026

Posted:

18 August 2026

You are already at the latest version

Abstract
Post-training quantization (PTQ) compresses large video-diffusion transformers without fine-tuning, but it is validated almost exclusively against frame-averaged reconstruction quality, which is not designed to express whether the generated human motion remains physically plausible: a video can be close to its full-precision reference frame by frame and still depict implausible motion. We study this gap for DynamiCtrl, a 5.77-billionparameter pose-conditioned video-diffusion transformer, and close a feedback loop around an automatic biomechanical evaluation pipeline. Kinematics enters not as a training loss but as a deterministic, gradient-free kinematic gate: a seven-check pass/fail acceptance test over acceleration variance, limb-length stability, normalised jitter, constraint violations, mean and 95th-percentile pose error, and motion suppression. Through the loop, the gate’s per-check verdicts and the depth of the offending layers rank layer sensitivity and construct a mixed-precision map, a role played elsewhere by gradient-based curvature heuristics. We run the loop with two W8A8 algorithms—ViDiT-Q with SmoothQuant folding and a calibration-matched naive baseline—across 66 quantized videos in a fullmodel grid, a layer-group sensitivity sweep, and a gate-guided hybrid. Early blocks dominate kinematic degradation, with relative acceleration deltas several times larger than mid or late blocks; protecting them raises the ViDiT-Q pass rate from 2 of 12 clips to 5 of 12 without fine-tuning. The gate also exposes distinct failure modes—smoothness for ViDiT-Q, constraint violations for the naive baseline—motivating a temporal quantization hypothesis. Kinematic plausibility, fed back into compression, is an effective training-free signal for mixed-precision mapping of human-video diffusion transformers.
Keywords: 
;  ;  ;  ;  ;  ;  ;  
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.