Simulated true-vs-proxy reward score trajectories for 300 reward-optimization configurations across five tasks, five mitigation methods, and reward-model quality, grounded in the reward-model overoptimization scaling-law literature (Gao et al 2023), InstructGPT, DPO, and Constitutional AI. Forward-looking task predicts eventual reward hacking from only early-optimization observations, before the classic Goodhart's Law decline becomes visible.
This dataset simulates reward hacking: a policy's true-objective score rises with optimization pressure only up to a point, then plateaus and declines, while its proxy reward keeps climbing, a clean instance of Goodhart's Law. The simulation spans 300 configurations across five tasks with varied specification-gap severity and five mitigation methods including a KL-divergence penalty, reward-model ensembling, DPO, and Constitutional AI-style feedback. Eleven grounding sources are documented in reward_hacking_theory.py and the README Research basis, eleven of twelve peer-reviewed.
Lets Ml Engineers Monitoring An Rlhf Or Applied-rl Training Run Estimate, From Early-optimization Observations Alone, Whether A Given Task And Mitigation-method Combination Is Headed For Reward Hacking, Before The Still-climbing Proxy Reward Curve Gives Any Visual Hint Of Trouble.
Attribution 4.0 International (CC BY- 4.0)
To preview this file, you need to be a registered user. Please complete the registration process to gain access and continue viewing the content.
© 2026 - Copyright AIKosh. All rights reserved.