Rung 01 Harnesses & Agent SystemsSwitch rungClose
Harnesses & Agent Systems — Timeline
Dates are publication dates of the underlying source, not ingest dates. First written
2026-09-24; it covers the sources that carry this rung's cost and measurement findings, not every
source in the register (raw/_sources.json holds all 23).
2026
September
- Sep 24 — [Compile] The Scaffold Effect paper ingested and compiled; the Harness Effect and Less Context sources ingested 2026-09-11 are compiled; the rung's first cost-per-task page is written. (Cost per Completed Task)
August
- Aug 05 — [Report] Prime Intellect's Prime Agent: a self-improving harness claiming wins on 9 of 11 rows, with no matched-budget control. (Prime Agent)
July
- Jul 14 — [Paper] Rethinking the evaluation of harness evolution: evolved harnesses do not consistently beat plain test-time scaling at matched budget. (Rethinking)
- Jul 10 — [Report] Prime Intellect's verifiers release splits taskset, harness and runtime. (verifiers)
- Jul 08 — [Preprint] The Harness Effect (Writer): an orchestration swap cuts cost per task 41%, $0.21 → $0.12, across six models; vendor-run. (The Harness Effect)
June
- Jun 08 — [Preprint] The Scaffold Effect in Coding Agents: ~40× tokens per solved task from harness choice, 0–8 points of pass rate. arXiv lists it as submitted Jun 08 under a July ID. (Scaffold Effect)
- Jun 08 — [Preprint] Less Context, Better Agents (Microsoft): trimmed history plus a summary reaches 91.6% completion on 553,374 tokens against 71.0% on 1,480,996 for full history. (Less Context)
May
- May 27 — [Paper] Harness-Bench: 106 tasks, 5,194 trajectories, harness effects isolated. (Harness-Bench)
- May 07 — [Paper] Stop Comparing LLM Agents Without Disclosing the Harness: harness variance 7.80× model variance on a 3×3 coding factorial. (Stop Comparing)