Skip to content
Rung 04 Evaluation & MeasurementSwitch rung

Evaluation & Measurement — Timeline

Dates are publication dates of the underlying source, not ingest dates.

2026

September

  • Sep 24 — [Compile] Timeline first written. The eval-cost article is ingested and compiled, and the two validity sources ingested on Sep 11 are compiled. (Index)

June

  • Jun 30 — [Preprint] Benchmarking the Benchmarks: tool-calling graders disagree with expert reviewers on 18.5% of 496 tasks; 23 repeats of one LiveMCPBench set-up span 57.9%–76.8%. (Benchmarking the Benchmarks; Benchmark Validity)
  • Jun 16 — [Paper] How Inference Compute Shapes Frontier LLM Evaluation: fixed single-budget scores understate newer models (up to 12 models × 7 benchmarks). (Inference compute)

April

  • Apr 29 — [Analysis] EvalEval Coalition prices measuring: ~$40,000 for one agent leaderboard run once, $2,829 for one GAIA run, a 33× cost spread from agent set-up, ~$320,000 with 8 repeats. (EvalEval; The Cost of Evaluation)

February

  • Feb 17 — [Preprint] Quantifying construct validity: a combined latent-factor and scaling-law model fitted to 4,395 Open LLM Leaderboard models fits better than either alone. (Construct validity)

2025

December

  • Dec 18 — [Report] UK AI Security Institute's Frontier AI Trends Report: 30+ models tested 2022–2025; cyber task length doubling roughly every 8 months. (AISI; AISI entity)
  • Dec 03 — [Paper] VideoScience-Bench: Sora-2 ~64%, Veo-3 ~58.7% on Phenomenon Congruency. (VideoScience-Bench)

July

  • Jul 17 — [Paper] PhyWorldBench: 12,600 generated videos checked for physical realism. (PhyWorldBench)