PAPER2026-06-16·Not stated verbatim on abstract page (roster overlaps UK AI Security Institute / frontier-eval groups)·arXiv 2606.17930

How Inference Compute Shapes Frontier LLM Evaluation

Jessica McFadyen, Ole Jorgensen, Harry Coppock, Kevin Wei, Cozmin Ududec
COMPILED NOTES

Across up to 12 frontier models x 7 benchmarks (SWE/math/medicine/cyber) under controlled inference-scaling interventions (token budget, context compaction, repeated attempts), fixed single-budget evals increasingly UNDERSTATE newer models, which have a higher ceiling at generous budgets; recommends reporting capability as a function of inference-time compute. Bridges test-time-compute, the matched-budget eval floor (§12), and inference-economics (§4).

How Inference Compute Shapes Frontier LLM Evaluation

Abstract

Verbatim (opening): "AI evaluations are shifting toward harder tasks that benefit from longer trajectories involving tool use and iterative problem solving. As a result, performance is increasingly sensitive to the amount and allocation of compute available at test time ('inference compute')."

The paper argues evaluations that report a single restrictive budget can understate model capability, and evaluates up to 12 frontier models across seven benchmarks (software engineering, mathematics, medicine, cybersecurity) under controlled inference-scaling interventions to show how much the reported number depends on the budget.

Key Contributions

  • Frames inference budget as a first-class evaluation axis, not a hidden constant — capability should be reported as a function of inference-time compute.
  • Three controlled interventions studied: (1) expanded token budgets; (2) context compaction; (3) repeated submission attempts (model-guided or feedback-directed).
  • Finding — fixed-budget evals increasingly underestimate frontier models, and the underestimation grows for newer models, which have a higher ceiling that only opens up at generous budgets.
  • Intervention efficacy is domain-dependent — repeated submissions help broadly; other techniques are domain-specific.

Methodology

Up to 12 frontier models evaluated across seven benchmarks spanning SWE, math, medicine, and cybersecurity, each run under controlled inference-scaling interventions (token-budget expansion, context compaction, repeated attempts). By sweeping the budget rather than fixing it, the study traces each model's capability-vs-compute curve instead of a point estimate.

Results

  • Larger token budgets substantially improve performance across multiple domains.
  • Newer models show a higher performance ceiling at generous budgets, unlocking harder tasks that fixed-budget scoring misses.
  • Different benchmarks benefit from different scaling methods; repeated submissions broadly help while others are domain-specific.
  • Core recommendation: report capability as a function of inference-time compute rather than a single fixed score. (Per-benchmark numeric grid not captured in this ingest.)

Limitations

Abstract-level numeric grounding only — the per-model, per-benchmark curves are not captured here. As with any inference-scaling study, results depend on the specific interventions and budget ranges chosen. Preprint, not peer-reviewed.

Relation to the KB

Bridges three frontier sections. It is the evaluation-methodology counterpart to §12's matched-budget floor (from Rethinking Harness Evolution): if agentic results must control for inference budget, this paper shows how much the budget moves the number and in which direction (up, for newer models). It also feeds §4 (Inference Economics) — if capability is a function of spend, "how good is model X" and "how much does it cost" collapse into one curve, strengthening the per-task (vs per-token) pricing argument. And it complements the test-time-compute strand (LoopFormer / latent reasoning) by measuring test-time compute's effect at the evaluation layer rather than the architecture layer.


Source: arXiv:2606.17930 — How Inference Compute Shapes Frontier LLM Evaluation, McFadyen, Jorgensen, Coppock, Wei, Ududec, 16 June 2026. Abstract page retrieved 2026-07-23.

RELATED · IN THE BASE
How Inference Compute Shapes Frontier LLM Evaluation | Knowledge Base | MenFem