Submitted:
28 August 2026
Posted:
30 August 2026
You are already at the latest version
Abstract
Muon orthogonalizes the momentum matrix, granting every singular direction of the update equal trust. EchoMuon prices that trust by each direction's echo: its support in a second, slower momentum buffer. This paper reports what that buys, and what measuring such a gate honestly costs. The gate is a contraction, not a reallocation. Its multiplier never exceeds one in 5,760 logged layer-steps, and the arm carrying it sits higher on the learning-rate axis: over four cells it is penalized 3.7x less than ungated Muon for a step twice too large, and 1.6x more for one half too small, without exception in either direction. A shared learning-rate grid therefore does not compare the two fairly: when the grid sits above the joint optimum, the contracting arm is rescued and the other is not, by an amount comparable to the margins being measured. Our initial hypotheses about the gate were wrong on this point, so we re-ran the full suite under corrected selection. All eight CIFAR arm-cells had picked the bottom edge of a single shared grid; widening it and selecting on a held-out split cuts those margins by 42 to 80% and leaves none individually significant. Tiny ImageNet, whose optimum the original grid did contain, is unchanged at +1.65 pp (t = +7.4) pooled over its two cells, and on the clean cell a norm-matched scalar control reproduces only 30% of the gain, so what remains is directional rather than a step-size effect. Three further boundaries are measured. The FineWeb advantage falls from -0.031 nats at 0.3 tokens per parameter to -0.003 at 0.9, verified not to be a learning-rate artifact. Across ten vision cells the margin scales with the baseline's error rate and crosses zero near 90% accuracy. On byte-level text EchoMuon wins on clean enwik8 (t=-10.1) and loses only under input corruption, which replaces the tokenization boundary we first proposed. We release 1,815 runs, the re-run suite behind these numbers, and the reporting practice we recommend: publish the margin by which each learning rate beat its runner-up. Here a pick that won by 0.0026 in validation loss reversed under an equally valid split, with a 1.45 pp consequence on test accuracy.
Keywords:
optimization
; Muon
; orthogonalized momentum
; learning-rate selection
; experimental methodology
; neural network training
; gated optimizers
; reproducibility
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.