
In scope: benchmark design and critique, matched-budget controls, contamination and overfitting to the eval set, agent benchmarks, the eval-to-deployment gap, measurement methodology. Out: the findings themselves, which belong to whichever rung they move.
Evaluation & Measurement
This rung is about how you know a thing works, and whether the measurement can be trusted. Its scope is benchmark design and the critique of it — matched-budget controls, contamination and overfitting to the eval set, agent benchmarks, and the gap between a benchmark score and what happens in deployment. The findings themselves belong to whichever rung they move; what belongs here is the question of whether those findings were measured properly.
What this rung prices
A buyer pays for a finished task, not for tokens, so every claim on the path from a token price to a task price rests on a benchmark number holding up. This rung is where that number is checked, and the check has one hard result: capability is not a scalar. Across up to 12 frontier models and 7 benchmarks — software engineering, mathematics, medicine, cybersecurity — run under controlled inference-scaling interventions, evaluations that report a single restrictive budget understate what a model can do, and understate newer models more, because a newer model's higher ceiling only opens at a generous budget (How Inference Compute Shapes Frontier LLM Evaluation; that close read reached abstract level only, so this rests on the abstract and the per-benchmark numbers were never captured). If a score moves with the budget, then score and spend are the same axis, and a leaderboard that publishes the first without the second cannot be converted into a cost per finished task at all. That is what this rung prices: it decides whether any other rung's number is a measurement or a marketing figure.
The load-bearing findings
- A benchmark score without a stated compute budget is not a measurement. The recommendation the inference-compute paper actually makes is to report capability as a function of inference-time compute rather than as a fixed score, and it names the three interventions it swept: expanded token budgets, context compaction, and repeated submission attempts. Repeated submissions helped broadly; the others were domain-dependent (How Inference Compute Shapes Frontier LLM Evaluation, read at abstract level).
- Capability is best read as a doubling rate, not a score, and the rate is itself moving. In the only multi-year independent series here — 30+ models tested 2022–2025 by a body that sells no model — apprentice-level cyber task success rose from about 9–10% in late 2023 to about 50%, and the length of cyber task a model can complete unassisted was doubling roughly every 8 months, with later internal data (February 2026) putting that at about every 4.7 months since late 2024 (Frontier AI Trends Report (AISI)).
- The same series shows how far a controlled setting can carry a claim, and where it stops. Self-replication success across five frontier models rose from under 5% in early 2023 to over 60% by mid-2025, but models were strongest at the early stages (getting compute and money) and weakest at the later ones (replicating onto compute, holding persistent access) — so the assessment is that real-world autonomous replication is unlikely today and the gains live in simplified environments (Frontier AI Trends Report (AISI)). A number that only exists inside the harness that produced it is the exact failure mode this rung exists to catch.
- Safeguard measurements move in both directions at once. The same report records a 40× increase in expert effort needed to find biological-misuse jailbreaks between two models released six months apart, and universal jailbreaks found in every system tested (Frontier AI Trends Report (AISI)). One metric improving by a large factor while the binary "can it be broken at all" stays at yes is a reminder that picking the metric is most of the argument.
- Benchmark design work outside the stack still demonstrates the method. Two video-generation benchmarks sit on this rung: one evaluating 12,600 generated videos for physical realism and finding systematic failures on rigid-body collisions, fluid dynamics and gravity (PhyWorldBench), and one scoring Sora-2 at about 64% and Veo-3 at about 58.7% on its Phenomenon Congruency metric (VideoScience-Bench). They are honest benchmark construction, but they measure text-to-video physics, not the inference stack this Atlas prices — they were ingested before the rung was rescoped, and they are counted here for accuracy, not for coverage.
The standing that all of this converts into is set out on agent evaluation benchmarks: a result is only interpretable if it discloses the harness, the search and inference budget, and held-out generalisation.
What we do not know yet
- The rung is thin, and thinner than its file count. Four sources, two of which are about video generation. On the rung's own through-line — measuring the inference stack — the working count is two. Three things named in scope have no coverage at all: contamination and overfitting to the eval set, the eval-to-deployment gap, and agent benchmarks treated as a design problem rather than a list.
- The rung owns a shape but not a number. The inference-compute study establishes the direction — scores rise with budget, and the gap widens for newer models — but its close read explicitly records that the per-benchmark numeric grid was not captured. So how far a score actually moves across a budget sweep is not held here, and capturing that grid from the source already ingested is the cheapest available improvement (How Inference Compute Shapes Frontier LLM Evaluation).
- Whether matched-budget control changes published rankings. The argument that it should is strong and general, but nobody here has re-run a model leaderboard under a single fixed budget, and no source on this rung reports that having been done.
- Whether benchmark versions are comparable to each other. Nothing on this rung publishes a version-to-version calibration, so a score quoted without its benchmark version cannot be placed against a score quoted with a different one.
- Nobody here has run any of it. Every number above is read, not measured. That matters more on this rung than on any other, because the rung's whole claim is that other people's numbers should not be taken at face value — and it is currently making that claim entirely on other people's numbers.
Read next
- how-inference-compute-shapes-frontier-llm-evaluation — the rung's central claim: a score without a budget is not a measurement.
- aisi-frontier-ai-trends-report-2025 — the only multi-year independent capability series here, and the rung's most credible instrument.
- phyworldbench-physics-benchmark — benchmark construction done well, on a subject off this rung's through-line; useful as method, not as coverage.