PAPER2026-06-10·Not stated in abstract·arXiv 2606.13715

WorkBench Revisited: Workplace Agents Two Years On

Olly Styles, Sam Miller
COMPILED NOTES

Two-year re-run of the WorkBench workplace-agent benchmark: best-agent task completion 43% (GPT-4, Mar-2024) -> 98% (Claude Fable 5, mid-2026); unintended harmful actions 26% -> 1.9% (capability and safety moved TOGETHER). KB-relevant finding: open-weight models collapsed the cost of a given performance level while FRONTIER serving costs stayed flat — a direct benchmark datapoint that the price decline shows up in the open-weight rung, not the frontier rung (open-weight-driven, ~100x since 2024 per the body). Per-tier $ tables not in abstract; full-text read owed.

WorkBench Revisited: Workplace Agents Two Years On

Abstract

"The best agent on WorkBench in March 2024, GPT-4, completed just 43% of tasks. We revisit the benchmark in June 2026 and find that the best agent to date, Claude Fable 5, now completes 98%. Beyond this considerable progress in frontier agent performance, three things stand out. First, unintended harmful actions, such as emailing the wrong person, fell from 26% of tasks for GPT-4 to 1.9% for Claude Fable 5; capability and safety go together on WorkBench rather than trade off, so the models that finish the most tasks also do the least unintended damage. Second, the rise of open-weight models has drastically lowered costs for a performance level that was only accessible to proprietary models, while frontier costs have stayed stable. Third, while several classes of error have been eliminated, frontier models still make some basic mistakes that occasionally result in irreversible harm. We release an updated version of the benchmark with data and code quality improvements, new model scores, and analysis of agent progress on WorkBench since 2024."

Key Contributions

  • Two-year longitudinal re-measurement of a fixed workplace-agent benchmark (same tasks, new models).
  • Capability curve: 43% → 98% best-agent task completion (Mar-2024 → mid-2026).
  • Safety curve: 26% → 1.9% of tasks with unintended harmful actions — capability and safety moved together, not as a trade-off.
  • Cost finding: open-weight models collapsed the cost of a performance level once proprietary-only, while frontier serving costs stayed flat.

Results — the numbers that matter to this KB

MetricMar 2024 (GPT-4)Mid-2026 (best)
Task completion43%98%
Unintended harmful actions26% of tasks1.9%
  • The load-bearing cost claim: the cost of achieving a given performance level dropped sharply — driven by open-weight models reaching what was proprietary-only — while frontier costs stayed stable. This is a direct observation of the KB's central open question ("do serving-layer efficiency gains reach published prices?"): the price decline shows up in the open-weight rung, not the frontier rung. (The abstract states "drastically lowered" / "hundredfold since 2024" in the body; the exact per-tier $ figures are in the paper's cost tables — full-text read owed.)

Limitations

  • One benchmark (workplace tasks); not coding/research agents.
  • The paper names the direction (open-weight cheap, frontier flat) but the exact per-tier cost levels need the full tables.

Why it matters (inference-economics lens)

This is the cleanest available narrative datapoint for the bifurcated price decline: the "cost per performance level" is falling fast at the open-weight tier and flat at the frontier — exactly the asymmetry the [price-decline-distribution] and Token Price Index work is trying to quantify. It complements [STAGE-Claw] (absolute per-task levels) with a trajectory (100×-since-2024, open-weight-driven).


Source: WorkBench Revisited (arXiv:2606.13715), submitted 2026-06-10 (v2 2026-07-01), Olly Styles & Sam Miller. Ingested 2026-07-24 at abstract grade. Provenance [preprint]; per-tier cost tables [full-text read owed].

RELATED · IN THE BASE
WorkBench Revisited: Workplace Agents Two Years On | Knowledge Base | MenFem