Rung 04 Evaluation & MeasurementSwitch rungClose

In scope: benchmark design and critique, matched-budget controls, contamination and overfitting to the eval set, agent benchmarks, the eval-to-deployment gap, measurement methodology. Out: the findings themselves, which belong to whichever rung they move.
| Type | Source | Published |
|---|---|---|
| ANALYSIS | AI evals are becoming the new compute bottleneck Avijit Ghosh, Yifan Mai, Georgia Channing, Leshem Choshen · EvalEval Coalition (Hugging Face community article) Prices measuring itself: ~$40,000 for one agent leaderboard (HAL, 21,730 rollouts, 9 models x 9 benchmarks, single seed), $2,829 for one GAIA run, a 33x cost spread on identical tasks from scaffold choice (Exgentic, $22,000 sweep), and ~$320,000 once every cell is re-run 8 times for reliability. Static benchmarks compress 100-200x without changing rankings; agent benchmarks only 2-3.5x. Compiled figures from papers and the live HAL leaderboard, GPU time converted at $2.50/H100-hr and $1.50/A10-hr. | 2026-04-29 |