Indian Flag
Government Of India
A-
A
A+
ORGANISATION
Test-Time-Compute-Scaling-Risk-Dataset

Test-Time-Compute-Scaling-Risk-Dataset

Simulated true-accuracy-vs-search-width curves for 180 configurations across 5 reasoning tasks and 6 test-time-compute methods (weak/strong-verifier best-of-N, beam search, self-consistency, PRM-guided search, verification-based finetuning), grounded in 2023-2026 research on search-time reward overoptimization. Forward-looking task predicts overoptimization at a large search budget from a cheap small-budget check.

About Dataset

Simulates a search-time instance of Goodhart's Law: the verifier/proxy score always rises with more search, but true accuracy rises then declines past a critical width for most methods. 180 configs, 6 methods. Ten sources (2021-2026) in test_time_compute_theory.py, including Gao et al.'s reward-overoptimization scaling laws, a formal best-of-N analysis, an ICML 2025 spotlight paper, and a 2026 paper extending the finding to beam search via extreme value theory.

Purpose of Dataset

Lets A Team Decide How Much Inference-time Search Budget Is Worth Spending With A Given Verifier, Before Discovering More Compute Made Things Worse.

Activity Overview Activity Overview

  • Downloads0
  • Downloads 0
  • File Size 337.10 KB
  • Views 4

Tags Tags

  • Large Language Model

License Control License Control

Attribution 4.0 International (CC BY- 4.0)

test-time-compute-scaling-risk-dataset.json ( 337.10 KB )


To preview this file, you need to be a registered user. Please complete the registration process to gain access and continue viewing the content.

Data Quality Score BetaData Quality Score Beta

Version Control Version Control

FolderVersion 1(337.10 KB)
  • JAI MALI·2 day(s) ago
    • application/json
      test-time-compute-scaling-risk-dataset.json