official government evaluation institute (UK AI Security Institute) · catastrophic / loss-of-control risk document

UK’s AI Security Institute: self-replication test scores jumped from 5% to 60%

Origin Lane: aiUnited Kingdom (AISI / DSIT) · global frontier labs

The UK AI Security Institute’s first public Frontier AI Trends Report — built from two years of government evaluations across more than 30 frontier systems — says capabilities that would be needed to evade human control are improving, with self-replication evaluation success rising from about 5% to 60%, while cyber and biology skills race past expert baselines.

UK AI Security Institute evaluator locus (London / UK) — not a model-lab GPS. Source map

Key facts

Who
UK AI Security Institute (AISI) under DSIT — evaluations since Nov 2023 across >30 frontier systems
Product
Frontier AI Trends Report — first public evidence-based synthesis (cyber, chem/bio, autonomy, loss-of-control precursors, safeguards)
Self-replication (RepliBench)
Early-2023 highest success <5% (11 of 20 evals) → summer 2025 two frontier models >60% (Figure 16); stronger on early obtain-compute/money than later replicate/persist; no spontaneous sandbagging/self-replication evidence yet
AGI framing (AISI)
Institute states it is plausible the observed trend may lead to capabilities widely acknowledged as AGI or otherwise transformative AI
Cyber
Apprentice-level ~10% early 2024 → ~50% average now; 2025 first model completing expert-level (10+ years human); task-length doubling ~every eight months
Chemistry & biology
Exceed PhD expert baselines on open-ended QA (up to ~+60% relative); protocol generation accurate from late 2024 and wet-lab feasible; troubleshooting up to ~90% better than experts
Safeguards
Universal jailbreaks for every system tested; some bio-misuse defenses needed ~40× more expert effort between two models six months apart
Companion Mar 2026
Multi-step cyber ranges — best run 22/32 steps; performance scales with test-time compute (10M→100M tokens, gains up to 59%); ICS range still limited (avg 1.2–1.4/7)
Live
Desk cites AISI measured results and AISI’s own AGI/loss-of-control language — no invented doom dial

Self-replication evaluations that barely cleared 5% success in early 2023 cleared 60% for top models by summer 2025 — and the same institute says AGI-class capabilities are a plausible near-path outcome of the trend they measured.

Desk reading of UK AISI Frontier AI Trends Report

Note

Britain’s AI Security Institute spent two years breaking and measuring frontier models, then published the scoreboard. Self-replication evaluations that barely cleared 5% success in early 2023 cleared 60% for top models by summer 2025. Cyber tasks that once needed a human apprentice now get beaten half the time; expert-level items started falling in 2025; biology helpers already outscore PhD baselines and draft lab protocols that work wet. Safeguards improved in places — and still yielded universal jailbreaks everywhere AISI looked. The institute’s own pages say the trajectory could reach what people call AGI. Desk is not inventing a takeover clock. It is reading the government’s ruler.

Attribution: UK AI Security Institute — Frontier AI Trends Report (aisi.gov.uk; PDF object last-modified 16 Dec 2025). Companion same-institute evaluation: Measuring AI Agents’ Progress on Multi-Step Cyber Attack Scenarios, 16 Mar 2026.

Why it matters

This is not a product launch blog. It is a national security evaluator telling the public that the precursor skills for losing control — self-replication pieces, long-horizon cyber agency, expert-beating wet-lab help — are moving fast while jailbreaks still exist for every tested stack.

Sources

Official data. Live values go to the HUD / source product.

Daily board