Simulated true-accuracy-vs-search-width curves for 180 configurations across 5 reasoning tasks and 6 test-time-compute methods (weak/strong-verifier best-of-N, beam search, self-consistency, PRM-guided search, verification-based finetuning), grounded in 2023-2026 research on search-time reward overoptimization. Forward-looking task predicts overoptimization at a large search budget from a cheap small-budget check.
Simulates a search-time instance of Goodhart's Law: the verifier/proxy score always rises with more search, but true accuracy rises then declines past a critical width for most methods. 180 configs, 6 methods. Ten sources (2021-2026) in test_time_compute_theory.py, including Gao et al.'s reward-overoptimization scaling laws, a formal best-of-N analysis, an ICML 2025 spotlight paper, and a 2026 paper extending the finding to beam search via extreme value theory.
Lets A Team Decide How Much Inference-time Search Budget Is Worth Spending With A Given Verifier, Before Discovering More Compute Made Things Worse.
Attribution 4.0 International (CC BY- 4.0)
To preview this file, you need to be a registered user. Please complete the registration process to gain access and continue viewing the content.
© 2026 - Copyright AIKosh. All rights reserved.