Rung 04 Evaluation & MeasurementSwitch rungClose
Evaluation & Measurement — Timeline
Dates are publication dates of the underlying source, not ingest dates.
2026
September
- Sep 24 — [Compile] Timeline first written. The eval-cost article is ingested and compiled, and the two validity sources ingested on Sep 11 are compiled. (Index)
June
- Jun 30 — [Preprint] Benchmarking the Benchmarks: tool-calling graders disagree with expert reviewers on 18.5% of 496 tasks; 23 repeats of one LiveMCPBench set-up span 57.9%–76.8%. (Benchmarking the Benchmarks; Benchmark Validity)
- Jun 16 — [Paper] How Inference Compute Shapes Frontier LLM Evaluation: fixed single-budget scores understate newer models (up to 12 models × 7 benchmarks). (Inference compute)
April
- Apr 29 — [Analysis] EvalEval Coalition prices measuring: ~$40,000 for one agent leaderboard run once, $2,829 for one GAIA run, a 33× cost spread from agent set-up, ~$320,000 with 8 repeats. (EvalEval; The Cost of Evaluation)
February
- Feb 17 — [Preprint] Quantifying construct validity: a combined latent-factor and scaling-law model fitted to 4,395 Open LLM Leaderboard models fits better than either alone. (Construct validity)
2025
December
- Dec 18 — [Report] UK AI Security Institute's Frontier AI Trends Report: 30+ models tested 2022–2025; cyber task length doubling roughly every 8 months. (AISI; AISI entity)
- Dec 03 — [Paper] VideoScience-Bench: Sora-2 ~64%, Veo-3 ~58.7% on Phenomenon Congruency. (VideoScience-Bench)
July
- Jul 17 — [Paper] PhyWorldBench: 12,600 generated videos checked for physical realism. (PhyWorldBench)