11HOW YOU KNOW IT WORKS· RISING

Evaluation & Measurement

How you know it works — and whether the measurement can be trusted. In scope: benchmark design and critique, matched-budget controls, contamination and overfitting to the eval set, agent benchmarks, eval-to-deployment gap, measurement methodology. Out: the findings themselves, which belong to whichever rung they move.

4SOURCES
1CONCEPTS
0ENTITIES
SOURCE MIX
3 P1 R0 A0 N
ACTIVITY · 20W
BenchmarksMatched budgetContamination
PAPER
2025-07-17

PhyWorldBench: A Comprehensive Evaluation of Physical Realism in Text-to-Video Models

Multiple authors · Multiple

12,600-video empirical benchmark — quantifies systematic physics failures in Sora and peer generative video models

PAPER
2025-12-03

VideoScience-Bench: Benchmarking Scientific Understanding and Reasoning for Video Generation

Multiple authors · Multiple

Sora-2 ~64% / Veo-3 ~58.7% on Phenomenon Congruency — quantifies how far frontier video models are from ground-truth physical realism

PAPER
2026-06-16

How Inference Compute Shapes Frontier LLM Evaluation

Jessica McFadyen, Ole Jorgensen, Harry Coppock, Kevin Wei, Cozmin Ududec · Not stated verbatim on abstract page (roster overlaps UK AI Security Institute / frontier-eval groups)

Across up to 12 frontier models x 7 benchmarks (SWE/math/medicine/cyber) under controlled inference-scaling interventions (token budget, context compaction, repeated attempts), fixed single-budget evals increasingly UNDERSTATE newer models, which have a higher ceiling at generous budgets; recommends reporting capability as a function of inference-time compute. Bridges test-time-compute, the matched-budget eval floor (§12), and inference-economics (§4).

REPORT
2025-12-18

Frontier AI Trends Report (AISI)

UK AI Security Institute (AISI) · UK AI Security Institute (AISI)

First public AISI assessment of frontier-model capability trends (30+ models, 2022–2025): cyber apprentice tasks 9%→50%, cyber time-horizon doubling ~8mo (accelerating to ~4.7mo), RepliBench <5%→>60%, 1hr SWE tasks >40%, models surpass biology-PhD baseline; universal jailbreaks in every system.

Evaluation & Measurement | Knowledge Base | MenFem