Evaluation & Measurement
How you know it works — and whether the measurement can be trusted. In scope: benchmark design and critique, matched-budget controls, contamination and overfitting to the eval set, agent benchmarks, eval-to-deployment gap, measurement methodology. Out: the findings themselves, which belong to whichever rung they move.
PhyWorldBench: A Comprehensive Evaluation of Physical Realism in Text-to-Video Models
12,600-video empirical benchmark — quantifies systematic physics failures in Sora and peer generative video models
VideoScience-Bench: Benchmarking Scientific Understanding and Reasoning for Video Generation
Sora-2 ~64% / Veo-3 ~58.7% on Phenomenon Congruency — quantifies how far frontier video models are from ground-truth physical realism
How Inference Compute Shapes Frontier LLM Evaluation
Across up to 12 frontier models x 7 benchmarks (SWE/math/medicine/cyber) under controlled inference-scaling interventions (token budget, context compaction, repeated attempts), fixed single-budget evals increasingly UNDERSTATE newer models, which have a higher ceiling at generous budgets; recommends reporting capability as a function of inference-time compute. Bridges test-time-compute, the matched-budget eval floor (§12), and inference-economics (§4).
Frontier AI Trends Report (AISI)
First public AISI assessment of frontier-model capability trends (30+ models, 2022–2025): cyber apprentice tasks 9%→50%, cyber time-horizon doubling ~8mo (accelerating to ~4.7mo), RepliBench <5%→>60%, 1hr SWE tasks >40%, models surpass biology-PhD baseline; universal jailbreaks in every system.