Simulated LLM-pretraining loss-spike probability for 448 configurations across 7 optimizers (AdamW, naive/scaled/8-bit Muon, naive/stabilized SOAP, Gluon), 4 batch-size regimes, and 4 model scales, grounded almost entirely in 2025-2026 optimizer research (Muon, SOAP, Gluon). Forward-looking task predicts a loss spike by the end of training from only the first 10% of training steps.
Simulates loss-spike risk grounded in the 2025 wave of matrix-aware optimizer research. 448 configs across 7 optimizers (naive vs. fixed Muon/SOAP variants, plus Gluon and 8-bit Muon), 4 batch regimes, 4 model scales (0.5B-70B). Ten sources, nine from 2025 and one a July 2026 technical report, in optimizer_stability_theory.py — including "Muon is Scalable for LLM Training" (Moonlight model paper), a formal Muon convergence analysis, and a unified SOAP/Muon/AdamW stability study identifying and fixing SOAP's large-batch loss spikes.
Lets A Team Monitoring A Long, Expensive Pretraining Job Estimate From Just The First 10% Of Steps Whether Their Optimizer Choice Is Headed For A Loss Spike.
Attribution 4.0 International (CC BY- 4.0)
To preview this file, you need to be a registered user. Please complete the registration process to gain access and continue viewing the content.
© 2026 - Copyright AIKosh. All rights reserved.