
In scope: scaffolding and orchestration, tool use, context and memory management, retries and recovery, agent evaluation and benchmarks, failure modes at step three, agent security and the exploitation surface, and cost per completed task. The load-bearing finding this rung already holds: for long-horizon agentic work the harness is often a stronger performance determinant than the model.
Analysis only1Show all →
| Type | Source | Published |
|---|---|---|
| ANALYSIS | Harness Engineering for Self-Improvement Lilian Weng · Lil'Log (self-published — page states no affiliation) Survey of ~39 refs. Lin et al. (2026): harness-UPDATING capability FLAT across Qwen2-32B->Opus 4.6 while harness BENEFIT is NON-MONOTONIC, middle-tier models gaining most. DGM: SWE-bench Verified 20%->50%, Polyglot 14.2%->30.7%. RE-Bench: agents 4x humans at a 2h budget but humans win at 8h and 32h. | 2026-07-04 |