
In scope: benchmark design and critique, matched-budget controls, contamination and overfitting to the eval set, agent benchmarks, the eval-to-deployment gap, measurement methodology. Out: the findings themselves, which belong to whichever rung they move.
Report only1Show all →
| Type | Source | Published |
|---|---|---|
| REPORT | Frontier AI Trends Report (AISI) UK AI Security Institute (AISI) · UK AI Security Institute (AISI) First public AISI assessment of frontier-model capability trends (30+ models, 2022–2025): cyber apprentice tasks 9%→50%, cyber time-horizon doubling ~8mo (accelerating to ~4.7mo), RepliBench <5%→>60%, 1hr SWE tasks >40%, models surpass biology-PhD baseline; universal jailbreaks in every system. | 2025-12-18 |