Simulated router-training-dynamics dataset, 216 runs by 200 steps, for six open and research Mixture-of-Experts architecture archetypes modeled on Switch, GShard, GLaM, Mixtral, DeepSeekMoE, and OLMoE (2021-2024), grounded in the load-balancing-loss and router-z-loss literature (2017-2024). Forward-looking task predicts end-of-training routing collapse from only the first 25 percent of the training curve.
This dataset simulates the self-reinforcing load-imbalance dynamics of Mixture-of-Experts token routers during training. The simulation sweeps six architecture archetypes approximating published open-weight and research MoE configs, Switch Transformer (2022), GShard (2021), GLaM (2022), Mixtral 8x7B (2024), DeepSeekMoE (2024), and OLMoE (2024), across load-balancing-loss weight, expert-capacity factor, and random seed, tracking routing entropy, max-expert token share, and capacity-factor token-drop rate at every one of 200 simulated training steps per run. The underlying router-dynamics equations are grounded in ten papers spanning 2017 to 2024: the original sparsely-gated MoE paper (2017), GShard and V-MoE (2021), Switch Transformers, GLaM, and ST-MoE (2022), Mixtral, DeepSeekMoE, Soft MoE, and OLMoE (2024). All ten sources and every equation are documented in moe_routing_theory.py and the README Research basis. The simulated model architectures reflect 2024-2025 open-weight releases; the mechanisms grounding the equations date back to 2017 since the 2024 models build directly on that earlier work.
Lets Ml Training Engineers Estimate, From Only The Early Portion Of A Training Run, Whether A Given Moe Architecture And Load-balancing Hyperparameter Combination Is Headed For Routing Collapse, Before Committing The Compute For A Full Run.
Attribution 4.0 International (CC BY- 4.0)
To preview this file, you need to be a registered user. Please complete the registration process to gain access and continue viewing the content.
© 2026 - Copyright AIKosh. All rights reserved.