Measuring Harness Effects

Active Frontier
Sign in to track mastery·Sign in
harnessagent-evaluationmethodology

Measuring Harness Effects

If the harness dominates the result, isolating it becomes the central experimental problem. The reference design is Harness-Bench: 106 sandboxed offline tasks and 5,194 execution trajectories, with task, budget and evaluation held fixed while only the harness varies across model backends.

Fixing the budget is the load-bearing control, and it is the one most often omitted. Without it, a "better harness" can simply be one that spends more inference — the improvement is real but it is bought, not designed, and it disappears the moment cost is held constant.

That control produces the field's most useful negative result. Automatic harness evolution — letting the system improve its own scaffolding — does not consistently outperform simple test-time-scaling baselines under matched feedback and inference budget, and what gains appear generalize poorly to held-out tasks (measured on Terminal-Bench 2.1 with GPT-5.4 and Claude Opus 4.6). Self-improving scaffolding is not yet a free lunch; much of its apparent advantage is unmatched spend.

Benchmarks themselves remain unsettled: a unified taxonomy covering roughly 60 agent benchmarks exists, which is itself evidence that the field has not converged on what to measure.

Key Claims

  • Harness-Bench: 106 tasks, 5,194 trajectories, isolating configuration-level harness effects by fixing task/budget/eval. Evidence: strong (paper) (Harness-Bench)
  • Harness evolution does not beat test-time-scaling baselines under matched budget, and generalizes poorly off-distribution. Evidence: strong (paper) (Rethinking Harness Evolution)
  • ~60 agent benchmarks catalogued in a unified taxonomy — the field has not converged on a measurement standard. Evidence: moderate (paper) (From LLM Reasoning to Autonomous Agents)
Measuring Harness Effects | KB | MenFem