Evaluation & Measurement — Frontier
Evaluation & Measurement — Frontier
The open questions on the rung that decides whether any other rung's number can be believed. Compiled 2026-08-28 from four close-reads. This is a thin rung and the compile does not disguise it — see the coverage note at the end.
The binding constraint, named
A benchmark score without a stated compute budget is not a measurement.
The evidence is direct: across up to 12 frontier models and 7 benchmarks (software engineering, maths, medicine, cybersecurity) under controlled inference-scaling interventions — larger token budgets, context compaction, repeated attempts — fixed single-budget evaluations increasingly understate newer models, because newer models have a higher ceiling at generous budgets (How Inference Compute Shapes Frontier LLM Evaluation). The paper's recommendation is the constraint stated positively: report capability as a function of inference-time compute, not as a scalar.
That finding is why this rung exists on an inference-stack Atlas at all. If capability is a function of budget, then score and cost are the same axis — and a leaderboard that publishes one without the other cannot be converted into a task price. Every rung downstream that wants to turn a token price into a task price depends on this one refusing scalar scores.
The number this rung owns
It does not have one yet, and that is the honest answer.
The closest candidate is the 12-model × 7-benchmark inference-scaling grid (as-of 2026-06-16), but the close-read explicitly records that the per-benchmark numeric grid was not captured — only the qualitative direction. So this rung currently owns a shape (capability rises with budget, and the gap widens for newer models) without owning a number.
Closing that is the single most valuable ingest available here.
Active questions
| Question | State of evidence | What would settle it |
|---|---|---|
| How much does a score move across the budget sweep? | Direction established, magnitude not held here. The one source with the grid did not have it extracted. Externally, the picture is worse than "uncertain" — the same model has been reported moving 73% → 89% while cost moved $0.91 → $2.34, with no model change at all (Rails Foundation study, held on the research sheet, not yet ingested to this rung). | Capturing the per-benchmark grid from the existing source, then one owned replication. |
| Are benchmark versions comparable at all? | Actively not. Terminal-Bench 3.0 compresses scores roughly 2.5x against 2.1 on the same task family, and both versions circulate in vendor material without the version prominently attached. (Research sheet, 2026-08-20 — not yet an ingested source here.) | A version-to-version calibration table. Nobody publishes one. |
| Does a frontier capability trend survive independent testing? | Partly. AISI's is the only multi-year independent series here — 30+ models tested 2022–2025 across cyber, biology, chemistry, autonomy/self-replication, software engineering and safeguards. It is the rung's most credible instrument precisely because the tester owns no model. | Continuation of the series, and any second national institute publishing a comparable one. |
| Is contamination or eval-set overfitting measurable in the open? | Unknown — the rung holds nothing on it, despite it being named in scope in LENS.md. | Any source at all. |
| Does matched-budget control change published rankings? | Open and high-stakes. The harnesses rung already records that automatic harness evolution does not beat simple test-time scaling under matched budget — the same control, applied to model leaderboards, has not been done publicly. | A leaderboard re-run under a single fixed budget. |
Coverage note — read this before citing the rung
Four sources, and two of them are off the rung's through-line. PhyWorldBench (12,600 videos, physical realism in text-to-video) and VideoScience-Bench (Sora-2, Veo-3 and peers reaching only ~58–64% on Phenomenon Congruency) are genuine benchmark-design work, but they measure video generation physics, not the inference stack this Atlas exists to price. They were ingested 2026-04-22, before the rung was rescoped.
So the working count for this rung's actual purpose is two sources, not four. Read sourceCount: 4 in _meta.json with that in mind — it is accurate as a file count and misleading as a coverage signal.
Three things in scope have zero coverage: contamination and overfitting to the eval set, the eval-to-deployment gap, and agent benchmarks as a design problem (the one concept page notwithstanding).
Standing note
Connor's seat here is read it. No evaluation in this rung has been run by him.
This is the rung where that matters most, because its whole claim is that other people's numbers should not be taken at face value — and a rung making that argument on four sources, two off-topic, and no owned number is making it on borrowed authority. Deep-dive menu #2 (the frozen-harness benchmark on his own corpus) would convert this rung and the serving rung in the same run.