REPORT2025-12-18·UK AI Security Institute (AISI)

Frontier AI Trends Report (AISI)

UK AI Security Institute (AISI)
COMPILED NOTES

First public AISI assessment of frontier-model capability trends (30+ models, 2022–2025): cyber apprentice tasks 9%→50%, cyber time-horizon doubling ~8mo (accelerating to ~4.7mo), RepliBench <5%→>60%, 1hr SWE tasks >40%, models surpass biology-PhD baseline; universal jailbreaks in every system.

Frontier AI Trends Report (AISI)

Core Thesis

The UK AI Security Institute's first public Frontier AI Trends Report (published 18 December 2025) is an evidence-based, longitudinal assessment of how the world's most advanced AI systems are evolving. It synthesizes two years of AISI's own controlled lab evaluations of 30+ large language models (2022–2025) across cybersecurity, biology, chemistry, autonomy, software engineering, and loss-of-control (self-replication). The throughline: frontier capability in security-relevant domains is advancing rapidly — in several domains faster than safeguards — but real-world autonomous risk (e.g. unaided self-replication) remains constrained to simplified, controlled settings today.

Key Facts

Cybersecurity (Cybench-style evaluations)

  • Apprentice-level cyber tasks: success rose from ~9–10% (late 2023 / early 2024) to ~50% on average (current).
  • Expert-level tasks: in 2025, the first model was tested that could complete expert-level tasks typically requiring 10+ years of human practitioner experience.
  • Time-horizon doubling: the length of cyber tasks (measured as how long they would take a human expert) that models can complete unassisted is doubling roughly every 8 months. More recent internal data (February 2026) estimated this had accelerated to doubling every ~4.7 months since late 2024.

Software Engineering

  • Software-engineering tasks lasting roughly one hour succeeded >40% of the time (as of summer 2025).

Autonomy / Self-Replication (RepliBench)

  • Tested five top-performing frontier models (Q3 2025).
  • RepliBench success rates rose from <5% (early 2023) to >60% (summer/mid-2025).
  • Models are strongest at early-stage capabilities (obtaining compute and money) and weakest at later stages (replicating onto compute and maintaining persistent access).
  • AISI assessment: real-world autonomous replication is unlikely today; gains are confined to controlled, simplified environments.

Biology & Chemistry

  • PhD-expert baselines: 48% (chemistry QA), 38% (biology QA).
  • Frontier models have "far surpassed PhD-level expertise in biology", with chemistry catching up. On some domain-specific expertise and open-ended questions, models exceed the expert baseline by up to ~60%.
  • First model to exceed a biology-PhD baseline appeared in 2024.

Safeguards

  • 40× increase in expert effort required to find biological-misuse jailbreaks between two models released just six months apart (2024–2025).
  • Universal jailbreaks were nonetheless found in every tested system; safeguard effectiveness varies by provider, request type, and weight-release status (open vs. closed).

Societal Use (UK survey, n=2,028)

  • 33% used AI for emotional support in the past year; 8% weekly, 4% daily.

Significance

This is the first time a government safety body has published a multi-year, multi-domain capability-trend baseline. It converts the "capabilities are outpacing understanding" narrative into measured doubling rates and domain-specific success curves, and it explicitly flags that the cyber time-horizon doubling rate has accelerated (8 months → ~4.7 months). The biology/safeguards findings (PhD-surpassing performance + universal jailbreaks present in every system) anchor the biorisk debate in measured evidence rather than speculation.


Source: Frontier AI Trends Report — UK AI Security Institute, published 18 December 2025. Findings summary: 5 key findings from our first Frontier AI Trends Report.

RELATED · IN THE BASE
Frontier AI Trends Report (AISI) | Knowledge Base | MenFem