Harnesses — Frontier
Last updated Invalid Date
Harnesses & Agent Systems — Frontier
The open questions on the rung that turns a token price into a task price. Compiled from the 15 sources moved here when the rung was split out of models on 2026-08-05.
Active questions
| Question | State of evidence | What would settle it |
|---|---|---|
| Does the harness advantage survive at matched budget? | Partly answered, and the answer is uncomfortable. Harness-induced variance is 7.80× model-induced variance — but automatic harness evolution does not beat simple test-time-scaling under matched feedback/inference budget. So harness design matters enormously while harness self-improvement may be mostly unmatched spend. | Matched-budget replications of self-improving scaffolds on held-out task families. |
| Is compositional reasoning a model limit or a harness limit? | Open, and it is the field's central bottleneck. Best agents reach 37.2% on 1,400 navigation tasks with navigation errors at 27–52% of failures — per-step tool use is fine, step selection is not. Explicit problem-model construction helps against CoT and ReAct, which points at the harness. | Whether frontier model releases move the navigation number at all, or only the per-step numbers. |
| How model-dependent is the exploitation surface? | Wide open, and consequential. Goal reframing was the sole reliable trigger of 12 dimensions — yet GPT-4.1 was completely immune across 1,850 trials. If immunity is reproducible, "agents are exploitable" is the wrong granularity of claim. | Cross-model replication of the 12-dimension battery on current frontier models. |
| Does memory curation pay for itself? | Under-measured. The write-manage-read frame makes manage the decisive stage, and A-MEM shows a design, but the KB holds no cost-per-task figure for curation against a naive append-and-retrieve store. | A matched-budget comparison of curated vs append-only memory on a long-horizon benchmark. |
| What is the actual cost-per-completed-task spread across harnesses? | The rung's biggest hole. Everything here is measured in accuracy and variance ratios; nothing in this topic reports dollars or tokens per finished job. That is the number the brand's whole argument needs. | Instrumented runs of the same task across harnesses, reporting tokens and dollars to completion including retries and failures. |
What this rung does not yet hold
- No cost-per-task measurement. See the last row above. This is the gap most worth closing, and the one Connor's own
ran itseat can close where a literature review cannot. - No production-scale failure data beyond the OpenClaw deployment study — the evidence base is dominated by offline benchmarks.
- Nothing on human-in-the-loop harnesses — approval gates, escalation policy, supervision cost. Given how MenFem's own desks operate, that absence is conspicuous.
Standing note
This is the one rung where Connor's seat is ran it — he operates desks, subagents and workflows daily. Per the teaching doctrine that is the standing a teaching-level claim requires, and the cost-per-task gap above is precisely where running it beats reading about it. The converse binds too: a claim here he has not run is coverage, not teaching.