Simulated Q-value overestimation trajectories and final policy divergence outcomes for 2475 offline reinforcement learning configurations across eleven methods (BCQ, BEAR, BRAC, CQL, IQL, TD3+BC, EDAC, MOPO, RvS, COMBO, and a naive baseline), five D4RL-style dataset quality categories, and dataset size, grounded in the offline RL extrapolation-error literature. Forward-looking task predicts end-of-training divergence from only an early-training checkpoint.
This dataset simulates the central failure mode of offline reinforcement learning: a Q-function queried on out-of-distribution actions during the Bellman backup compounds unchecked error across bootstrapped updates until the policy diverges. The simulation spans 2475 configurations across five task archetypes, five D4RL-style dataset quality categories, eleven methods with different OOD-correction mechanisms, and three dataset sizes. Twelve grounding sources are documented in offline_rl_theory.py and the README Research basis, including BCQ, CQL, IQL, EDAC, MOPO, and COMBO.
Lets Ml Engineers Estimate, From An Early-training Checkpoint, Whether A Given Offline Rl Method And Dataset-quality Combination Is Headed For Q-value Divergence, Before The Failure Becomes Visually Obvious On A Training Curve.
Attribution 4.0 International (CC BY- 4.0)
To preview this file, you need to be a registered user. Please complete the registration process to gain access and continue viewing the content.
© 2026 - Copyright AIKosh. All rights reserved.