Rung 06 Agent SecuritySwitch rungClose
Agent Security — Timeline
Dates are publication dates of the underlying source, not ingest dates.
2026
September
- Sep 24 — [Compile] First compile of the rung: index, timeline and frontier written. The two MCP sources ingested on Sep 11 are compiled. (Index)
- Sep 24 — [Ingest] Check Point's Claude Code config write-up ingested and compiled. (Project Config as an Attack Path)
August
- Aug 28 — [Preprint] ContextLeak: a trained attacker writes tool names and descriptions that get an agent to pick the tool and hand over its context — 92% selection in simulation, 22% against Claude Code; three detectors miss it 99–100% of the time. (ContextLeak)
April
- Apr 06 — [Preprint] Wang et al. study OpenClaw, a live personal agent: poisoning its tools, guidelines or memory raises attack success from 24.6% to 64–74%. (OpenClaw analysis; Deployed Agent Safety)
- Apr 06 — [Preprint] Mouzouni's ~10,000-trial study: of twelve kinds of prompt pressure, only goal reframing reliably triggers exploitation. (Exploitation taxonomy; Agent Exploitation Attack Surface)
February
- Feb 25 — [Report] Check Point Research publishes "Caught in the Hook": three Claude Code project-config attacks — hooks after one trust click, auto-approved MCP servers (CVE-2025-59536), and API-key theft via a redirected base URL (CVE-2026-21852) — all fixed before publication. (Caught in the Hook)
- Feb 09 — [Analysis] Zylos Research lists six alignment failure modes and the alignment trilemma. (Zylos; Agent Safety & Alignment)
- Feb 07 — [Analysis] Kanagala's agentic threat taxonomy: permission escalation, memory manipulation, supply-chain attacks. (Red-teaming framework; Agent Safety & Alignment)
January
- Jan 24 — [Preprint] Breaking the Protocol: three design-level defects in MCP; attack success 23–41% higher than equivalent non-MCP integrations; the proposed MCPSec fix takes it from 52.8% to 12.4% at 8.3 ms per message. (Breaking the Protocol)
- Jan 21 — [Advisory] Anthropic publishes CVE-2026-21852 (API key sent before trust via a
repo-set
ANTHROPIC_BASE_URL; fixed in Claude Code 2.0.65). (Caught in the Hook)
2025
December
- Dec 18 — [Report] UK AI Security Institute's Frontier AI Trends Report: universal jailbreaks found in every system tested, despite a 40x rise in the expert effort needed for one class of jailbreak. (AISI; Agent Safety & Alignment)
October
- Oct 03 — [Advisory] Anthropic publishes CVE-2025-59536 (repo-approved MCP servers run before the trust dialog; fixed in Claude Code 1.0.111). (Caught in the Hook)
August
- Aug 29 — [Advisory] Anthropic publishes GHSA-ph6w-f82w-28w6 (repo hooks run after the trust prompt without a further prompt; fixed in Claude Code 1.0.87). The OSV record dates it Sep 03. (Caught in the Hook)
July
- Jul 21 — [Disclosure] Check Point reports the hooks flaw to Anthropic. (Caught in the Hook)