Skip to content
Rung 04 Evaluation & MeasurementSwitch rung
Rung 04 / How you know it works

Evaluation & Measurement

How you know it works — and whether the measurement can be trusted.

7
Sources
3
Concepts
2
Entities
A brushed steel dial gauge with a copper needle, on paper.

In scope: benchmark design and critique, matched-budget controls, contamination and overfitting to the eval set, agent benchmarks, the eval-to-deployment gap, measurement methodology. Out: the findings themselves, which belong to whichever rung they move.

Three stacked sieves of tightening aperture above the small sample that survives.
BenchmarksMatched budgetContamination
Analysis only1Show all
Sources compiled for this topic
TypeSourcePublished
ANALYSISAI evals are becoming the new compute bottleneck
Avijit Ghosh, Yifan Mai, Georgia Channing, Leshem Choshen · EvalEval Coalition (Hugging Face community article)

Prices measuring itself: ~$40,000 for one agent leaderboard (HAL, 21,730 rollouts, 9 models x 9 benchmarks, single seed), $2,829 for one GAIA run, a 33x cost spread on identical tasks from scaffold choice (Exgentic, $22,000 sweep), and ~$320,000 once every cell is re-run 8 times for reliability. Static benchmarks compress 100-200x without changing rankings; agent benchmarks only 2-3.5x. Compiled figures from papers and the live HAL leaderboard, GPU time converted at $2.50/H100-hr and $1.50/A10-hr.

2026-04-29