Rung 11 / How you know it works

Evaluation & Measurement

How you know it works — and whether the measurement can be trusted.

6
Sources
1
Concepts
1
Entities
Paper5Report1AnalysisNews
A brushed steel dial gauge with a copper needle, on paper.

In scope: benchmark design and critique, matched-budget controls, contamination and overfitting to the eval set, agent benchmarks, the eval-to-deployment gap, measurement methodology. Out: the findings themselves, which belong to whichever rung they move.

Three stacked sieves of tightening aperture above the small sample that survives.
BenchmarksMatched budgetContamination
Paper only5Show all →
Sources compiled for this topic
TypeSourcePublished
PAPERPhyWorldBench: A Comprehensive Evaluation of Physical Realism in Text-to-Video Models
Multiple authors · Multiple

12,600-video empirical benchmark — quantifies systematic physics failures in Sora and peer generative video models

2025-07-17
PAPERVideoScience-Bench: Benchmarking Scientific Understanding and Reasoning for Video Generation
Multiple authors · Multiple

Sora-2 ~64% / Veo-3 ~58.7% on Phenomenon Congruency — quantifies how far frontier video models are from ground-truth physical realism

2025-12-03
PAPERHow Inference Compute Shapes Frontier LLM Evaluation
Jessica McFadyen, Ole Jorgensen, Harry Coppock, Kevin Wei, Cozmin Ududec · Not stated verbatim on abstract page (roster overlaps UK AI Security Institute / frontier-eval groups)

Across up to 12 frontier models x 7 benchmarks (SWE/math/medicine/cyber) under controlled inference-scaling interventions (token budget, context compaction, repeated attempts), fixed single-budget evals increasingly UNDERSTATE newer models, which have a higher ceiling at generous budgets; recommends reporting capability as a function of inference-time compute. Bridges test-time-compute, the matched-budget eval floor (§12), and inference-economics (§4).

2026-06-16
PAPERBenchmarking the Benchmarks: A Validity Audit of Tool-Calling Evaluation
Vishvesh Bhat, Jay Vaghasiya, Muhammad Ahmed Mohsin, Asad Aali · Not stated on the arXiv abstract page

Across 496 expert-reviewed tasks on BFCL v4, tau2-Bench, LiveMCPBench and MCP-Atlas: 92 evaluator-human disagreements (18.5% misalignment). In LiveMCPBench, 23 repeats of the SAME setup score 57.9%-76.8% — an 18.9-point spread, wide enough to reorder a leaderboard. Taxonomy of deterministic vs LLM-judge failure families.

2026-06-30
PAPERQuantifying construct validity in large language model evaluations
Ryan Othniel Kearns · Not stated on the arXiv abstract page (thesis)

Structured capabilities model (latent factors + scaling laws jointly) fitted to 4,395 models x 19 BBH subtasks from OpenLLM Leaderboard v2. Logistic form: RMSEA 0.0623 / CFI 0.9836 / SRMR 0.0110 vs pure EFA 0.0876 / 0.9737 / 0.0134. OOD prediction beats observational scaling laws but the magnitude did not render on fetch — unquantified here.

2026-02-17
Evaluation & Measurement | Knowledge Base | MenFem