Harnesses — Frontier

Last updated Invalid Date

Harnesses & Agent Systems — Frontier

The open questions on the rung that turns a token price into a task price. Compiled from the 15 sources moved here when the rung was split out of models on 2026-08-05.

Active questions

QuestionState of evidenceWhat would settle it
Does the harness advantage survive at matched budget?Partly answered, and the answer is uncomfortable. Harness-induced variance is 7.80× model-induced variance — but automatic harness evolution does not beat simple test-time-scaling under matched feedback/inference budget. So harness design matters enormously while harness self-improvement may be mostly unmatched spend.Matched-budget replications of self-improving scaffolds on held-out task families.
Is compositional reasoning a model limit or a harness limit?Open, and it is the field's central bottleneck. Best agents reach 37.2% on 1,400 navigation tasks with navigation errors at 27–52% of failures — per-step tool use is fine, step selection is not. Explicit problem-model construction helps against CoT and ReAct, which points at the harness.Whether frontier model releases move the navigation number at all, or only the per-step numbers.
How model-dependent is the exploitation surface?Wide open, and consequential. Goal reframing was the sole reliable trigger of 12 dimensions — yet GPT-4.1 was completely immune across 1,850 trials. If immunity is reproducible, "agents are exploitable" is the wrong granularity of claim.Cross-model replication of the 12-dimension battery on current frontier models.
Does memory curation pay for itself?Under-measured. The write-manage-read frame makes manage the decisive stage, and A-MEM shows a design, but the KB holds no cost-per-task figure for curation against a naive append-and-retrieve store.A matched-budget comparison of curated vs append-only memory on a long-horizon benchmark.
What is the actual cost-per-completed-task spread across harnesses?The rung's biggest hole. Everything here is measured in accuracy and variance ratios; nothing in this topic reports dollars or tokens per finished job. That is the number the brand's whole argument needs.Instrumented runs of the same task across harnesses, reporting tokens and dollars to completion including retries and failures.

What this rung does not yet hold

  • No cost-per-task measurement. See the last row above. This is the gap most worth closing, and the one Connor's own ran it seat can close where a literature review cannot.
  • No production-scale failure data beyond the OpenClaw deployment study — the evidence base is dominated by offline benchmarks.
  • Nothing on human-in-the-loop harnesses — approval gates, escalation policy, supervision cost. Given how MenFem's own desks operate, that absence is conspicuous.

Standing note

This is the one rung where Connor's seat is ran it — he operates desks, subagents and workflows daily. Per the teaching doctrine that is the standing a teaching-level claim requires, and the cost-per-task gap above is precisely where running it beats reading about it. The converse binds too: a claim here he has not run is coverage, not teaching.

Frontier — Harnesses & Agent Systems | KB | MenFem