
In scope: benchmark design and critique, matched-budget controls, contamination and overfitting to the eval set, agent benchmarks, the eval-to-deployment gap, measurement methodology. Out: the findings themselves, which belong to whichever rung they move.
| Type | Source | Published |
|---|---|---|
| PAPER | PhyWorldBench: A Comprehensive Evaluation of Physical Realism in Text-to-Video Models Multiple authors · Multiple 12,600-video empirical benchmark — quantifies systematic physics failures in Sora and peer generative video models | 2025-07-17 |
| PAPER | VideoScience-Bench: Benchmarking Scientific Understanding and Reasoning for Video Generation Multiple authors · Multiple Sora-2 ~64% / Veo-3 ~58.7% on Phenomenon Congruency — quantifies how far frontier video models are from ground-truth physical realism | 2025-12-03 |
| PAPER | How Inference Compute Shapes Frontier LLM Evaluation Jessica McFadyen, Ole Jorgensen, Harry Coppock, Kevin Wei, Cozmin Ududec · Not stated verbatim on abstract page (roster overlaps UK AI Security Institute / frontier-eval groups) Across up to 12 frontier models x 7 benchmarks (SWE/math/medicine/cyber) under controlled inference-scaling interventions (token budget, context compaction, repeated attempts), fixed single-budget evals increasingly UNDERSTATE newer models, which have a higher ceiling at generous budgets; recommends reporting capability as a function of inference-time compute. Bridges test-time-compute, the matched-budget eval floor (§12), and inference-economics (§4). | 2026-06-16 |
| PAPER | Benchmarking the Benchmarks: A Validity Audit of Tool-Calling Evaluation Vishvesh Bhat, Jay Vaghasiya, Muhammad Ahmed Mohsin, Asad Aali · Not stated on the arXiv abstract page Across 496 expert-reviewed tasks on BFCL v4, tau2-Bench, LiveMCPBench and MCP-Atlas: 92 evaluator-human disagreements (18.5% misalignment). In LiveMCPBench, 23 repeats of the SAME setup score 57.9%-76.8% — an 18.9-point spread, wide enough to reorder a leaderboard. Taxonomy of deterministic vs LLM-judge failure families. | 2026-06-30 |
| PAPER | Quantifying construct validity in large language model evaluations Ryan Othniel Kearns · Not stated on the arXiv abstract page (thesis) Structured capabilities model (latent factors + scaling laws jointly) fitted to 4,395 models x 19 BBH subtasks from OpenLLM Leaderboard v2. Logistic form: RMSEA 0.0623 / CFI 0.9836 / SRMR 0.0110 vs pure EFA 0.0876 / 0.9737 / 0.0134. OOD prediction beats observational scaling laws but the magnitude did not render on fetch — unquantified here. | 2026-02-17 |
| REPORT | Frontier AI Trends Report (AISI) UK AI Security Institute (AISI) · UK AI Security Institute (AISI) First public AISI assessment of frontier-model capability trends (30+ models, 2022–2025): cyber apprentice tasks 9%→50%, cyber time-horizon doubling ~8mo (accelerating to ~4.7mo), RepliBench <5%→>60%, 1hr SWE tasks >40%, models surpass biology-PhD baseline; universal jailbreaks in every system. | 2025-12-18 |