Rung 11 / How you know it works

Evaluation & Measurement

How you know it works — and whether the measurement can be trusted.

6
Sources
1
Concepts
1
Entities
Paper5Report1AnalysisNews
A brushed steel dial gauge with a copper needle, on paper.

In scope: benchmark design and critique, matched-budget controls, contamination and overfitting to the eval set, agent benchmarks, the eval-to-deployment gap, measurement methodology. Out: the findings themselves, which belong to whichever rung they move.

Three stacked sieves of tightening aperture above the small sample that survives.
BenchmarksMatched budgetContamination
Report only1Show all →
Sources compiled for this topic
TypeSourcePublished
REPORTFrontier AI Trends Report (AISI)
UK AI Security Institute (AISI) · UK AI Security Institute (AISI)

First public AISI assessment of frontier-model capability trends (30+ models, 2022–2025): cyber apprentice tasks 9%→50%, cyber time-horizon doubling ~8mo (accelerating to ~4.7mo), RepliBench <5%→>60%, 1hr SWE tasks >40%, models surpass biology-PhD baseline; universal jailbreaks in every system.

2025-12-18
Evaluation & Measurement | Knowledge Base | MenFem