Rung 11 / How you know it works

Evaluation & Measurement

How you know it works — and whether the measurement can be trusted.

6
Sources
1
Concepts
1
Entities
Paper5Report1AnalysisNews
A brushed steel dial gauge with a copper needle, on paper.

In scope: benchmark design and critique, matched-budget controls, contamination and overfitting to the eval set, agent benchmarks, the eval-to-deployment gap, measurement methodology. Out: the findings themselves, which belong to whichever rung they move.

Three stacked sieves of tightening aperture above the small sample that survives.
BenchmarksMatched budgetContamination
Sources compiled for this topic
TypeSourcePublished
PAPERPhyWorldBench: A Comprehensive Evaluation of Physical Realism in Text-to-Video Models
Multiple authors · Multiple

12,600-video empirical benchmark — quantifies systematic physics failures in Sora and peer generative video models

2025-07-17
PAPERVideoScience-Bench: Benchmarking Scientific Understanding and Reasoning for Video Generation
Multiple authors · Multiple

Sora-2 ~64% / Veo-3 ~58.7% on Phenomenon Congruency — quantifies how far frontier video models are from ground-truth physical realism

2025-12-03
PAPERHow Inference Compute Shapes Frontier LLM Evaluation
Jessica McFadyen, Ole Jorgensen, Harry Coppock, Kevin Wei, Cozmin Ududec · Not stated verbatim on abstract page (roster overlaps UK AI Security Institute / frontier-eval groups)

Across up to 12 frontier models x 7 benchmarks (SWE/math/medicine/cyber) under controlled inference-scaling interventions (token budget, context compaction, repeated attempts), fixed single-budget evals increasingly UNDERSTATE newer models, which have a higher ceiling at generous budgets; recommends reporting capability as a function of inference-time compute. Bridges test-time-compute, the matched-budget eval floor (§12), and inference-economics (§4).

2026-06-16
PAPERBenchmarking the Benchmarks: A Validity Audit of Tool-Calling Evaluation
Vishvesh Bhat, Jay Vaghasiya, Muhammad Ahmed Mohsin, Asad Aali · Not stated on the arXiv abstract page

Across 496 expert-reviewed tasks on BFCL v4, tau2-Bench, LiveMCPBench and MCP-Atlas: 92 evaluator-human disagreements (18.5% misalignment). In LiveMCPBench, 23 repeats of the SAME setup score 57.9%-76.8% — an 18.9-point spread, wide enough to reorder a leaderboard. Taxonomy of deterministic vs LLM-judge failure families.

2026-06-30
PAPERQuantifying construct validity in large language model evaluations
Ryan Othniel Kearns · Not stated on the arXiv abstract page (thesis)

Structured capabilities model (latent factors + scaling laws jointly) fitted to 4,395 models x 19 BBH subtasks from OpenLLM Leaderboard v2. Logistic form: RMSEA 0.0623 / CFI 0.9836 / SRMR 0.0110 vs pure EFA 0.0876 / 0.9737 / 0.0134. OOD prediction beats observational scaling laws but the magnitude did not render on fetch — unquantified here.

2026-02-17
REPORTFrontier AI Trends Report (AISI)
UK AI Security Institute (AISI) · UK AI Security Institute (AISI)

First public AISI assessment of frontier-model capability trends (30+ models, 2022–2025): cyber apprentice tasks 9%→50%, cyber time-horizon doubling ~8mo (accelerating to ~4.7mo), RepliBench <5%→>60%, 1hr SWE tasks >40%, models surpass biology-PhD baseline; universal jailbreaks in every system.

2025-12-18
Evaluation & Measurement | Knowledge Base | MenFem