Skip to content
Rung 01 Harnesses & Agent SystemsSwitch rung

Harnesses & Agent Systems — Timeline

Dates are publication dates of the underlying source, not ingest dates. First written 2026-09-24; it covers the sources that carry this rung's cost and measurement findings, not every source in the register (raw/_sources.json holds all 23).

2026

September

  • Sep 24 — [Compile] The Scaffold Effect paper ingested and compiled; the Harness Effect and Less Context sources ingested 2026-09-11 are compiled; the rung's first cost-per-task page is written. (Cost per Completed Task)

August

  • Aug 05 — [Report] Prime Intellect's Prime Agent: a self-improving harness claiming wins on 9 of 11 rows, with no matched-budget control. (Prime Agent)

July

  • Jul 14 — [Paper] Rethinking the evaluation of harness evolution: evolved harnesses do not consistently beat plain test-time scaling at matched budget. (Rethinking)
  • Jul 10 — [Report] Prime Intellect's verifiers release splits taskset, harness and runtime. (verifiers)
  • Jul 08 — [Preprint] The Harness Effect (Writer): an orchestration swap cuts cost per task 41%, $0.21 → $0.12, across six models; vendor-run. (The Harness Effect)

June

  • Jun 08 — [Preprint] The Scaffold Effect in Coding Agents: ~40× tokens per solved task from harness choice, 0–8 points of pass rate. arXiv lists it as submitted Jun 08 under a July ID. (Scaffold Effect)
  • Jun 08 — [Preprint] Less Context, Better Agents (Microsoft): trimmed history plus a summary reaches 91.6% completion on 553,374 tokens against 71.0% on 1,480,996 for full history. (Less Context)

May

  • May 27 — [Paper] Harness-Bench: 106 tasks, 5,194 trajectories, harness effects isolated. (Harness-Bench)
  • May 07 — [Paper] Stop Comparing LLM Agents Without Disclosing the Harness: harness variance 7.80× model variance on a 3×3 coding factorial. (Stop Comparing)