# Cost Per Task (Cost-of-Pass)

Canonical URL: https://menfem.com/kb/inference-economics/concepts/cost-per-task
Knowledge base topic: [Inference & Token-Pricing Economics](https://menfem.com/kb/inference-economics)
Frontier status: active
Tags: inference-economics, cost-per-task, cost-of-pass, accuracy-adjusted-cost, tokenomics, value-vs-cost

---

This topic has said, repeatedly, that the commercially meaningful unit is **cost per task**, not price per token — and until this compile it held no framework for it, only Gartner's second-hand 5–30x agentic multiplier. Two 2026 economics preprints supplied the missing layer: one gives the **measurable definition** (Cost-of-Pass), the other the **theory of why the token is the wrong unit** (AI Tokenomics). Both are single-author-or-small-team preprints, **not peer-reviewed**, and **read from their abstract pages only** — the measured trajectories are real, the frameworks are definitional.

**As of the 2026-07-24 compile the topic's long-missing *absolute dated level* is finally in.** [STAGE-Claw](../../raw/stage-claw-state-based-agent-benchmark-cost-per-task.md) (arXiv 2606.10394) measures **~$0.17–$6.55 per task across 11 frontier models** on a 40-task state-based agent benchmark — the first concrete 2026 $/task range this KB holds — and [WorkBench Revisited](../../raw/workbench-revisited-workplace-agents-two-years-on.md) (arXiv 2606.13715) adds the **two-year trajectory** (task completion 43%→98%, cost of a given performance level collapsing at the *open-weight* tier while frontier costs stayed flat). Both are ingested at **abstract grade**; STAGE-Claw's specific cost table is **search-surfaced, full-text confirmation owed**. What the level does *not* yet cover: a frontier/proprietary per-task cost *series*, and any workload beyond personal-computing/workplace agents.

## The definition — price a correct answer, not a token

Cost-of-Pass (Erol, El, Suzgun, Yuksekgonul & Zou; arXiv 2504.13359; submitted 2025-04-17, revised 2026-02-26) formalizes the unit:

- **Cost-of-pass** = "the expected monetary cost of generating a *correct* solution." Operationally, per-attempt inference cost divided by the probability the attempt is correct (accuracy). A cheap-per-token model that is often wrong can carry a *higher* cost-of-pass than an expensive model that is usually right.
- **Frontier cost-of-pass** = the minimum cost-of-pass across all available models — *or* across models **and human expert(s)**. This makes "is the AI economically worth it vs a human, on this task?" a computable question per task type.

The distinction is not cosmetic. Per-token price and cost-per-correct-answer can move in *opposite* directions: a model can get cheaper per token and more expensive per solved task (if accuracy drops), or dearer per token and cheaper per task (if accuracy rises enough). That decoupling is exactly what the reasoning-model tier does — see below.

## The measured findings (Cost-of-Pass, across real models over time)

- **Task-type specialization is economic, not just capability.** Lightweight models win on basic quantitative tasks; large models win on knowledge-intensive tasks; **reasoning models win on complex quantitative problems despite higher per-token cost** — because their accuracy lowers cost *per correct answer*. *Evidence: measured.*
- **Frontier cost-of-pass for complex quantitative tasks "roughly halved every few months."** *Evidence: measured trajectory; the exact model set and window need the full PDF to date precisely.*
- **Model-level innovation — not inference-time techniques (sampling, budget-aware decoding) — drives the primary cost-efficiency gains.** Budget-aware methods (e.g. TALE-EP) help at the margin, not as the main lever. *Evidence: measured.*

## The measured absolute levels (added 2026-07-24)

Cost-of-Pass and AI Tokenomics gave the topic the *unit* but no *number*. Two agentic-benchmark preprints now attach one.

**STAGE-Claw (Liang et al., arXiv 2606.10394, 2026-06-09)** — a state-based personal-agent benchmark of **40 tasks run across 11 frontier models**, scored on the correctness of the final *system state* rather than the text of the response. Its cost axis is why it matters here:

- **Measured API cost per task ≈ $0.17 (cheapest, MiniMax-tier) to ≈ $6.55 (Claude-Opus-4.7-tier)** — a ~38x spread across the 11 models on the same 40 tasks. This is the first **dated, absolute, per-task dollar level** for a real 2026 agentic workload anywhere in this KB. *Evidence: measured, but search-surfaced from the paper's cost table (the fetched abstract references "costs" without tabulating them) — full-text confirmation owed; API list-price, personal-computing scenarios only.*
- Time cost per task ≈ **3–15 minutes** per model (same provenance caveat).

The ~38x model spread is the cost-of-pass thesis in dollars: the expensive-per-task model at the top of the range is the frontier reasoning tier that Cost-of-Pass says you buy *because* accuracy lowers cost-per-correct-answer — STAGE-Claw prices both ends of that trade for the first time.

**WorkBench Revisited (Styles & Miller, arXiv 2606.13715, 2026-06-10)** re-runs a fixed workplace-agent benchmark two years on and supplies the *trajectory* under the level:

- Best-agent task completion **43% (GPT-4, Mar-2024) → 98% (Claude Fable 5, mid-2026)**; unintended harmful actions **26% → 1.9%** — capability and safety moved *together*, not as a trade-off. *Evidence: measured (abstract-grade).*
- The load-bearing cost finding: **open-weight models collapsed the cost of a performance level that was once proprietary-only, while frontier serving costs stayed flat** (~100x since 2024, open-weight-driven, per the body; exact per-tier $ tables need the full text). This is a *bifurcated* decline — the same asymmetry recorded on [the price-decline distribution](./price-decline-distribution.md), where WorkBench's main treatment lives.

Together: STAGE-Claw gives the **absolute level** ($0.17–$6.55/task, one date, one benchmark family), WorkBench gives the **trajectory and its shape** (falling fast at the open-weight tier, flat at the frontier). Neither closes the frontier-proprietary cost *series* that [Frontier 1](../frontier.md) still tracks — the per-task cost is now anchored for agentic benchmarks, not for GPT-5.x / Claude / Gemini serving as a time series.

## The theory — token expenditure ≠ economic value

AI Tokenomics (Quanyan Zhu; arXiv 2606.24616; 2026-06-10) proposes "AI tokenomics" as a formal field — "how tokens are generated, consumed, priced, allocated, and optimized" — and lands the load-bearing distinction that reframes the whole margin debate:

- **Token expenditure and economic value are distinct quantities.** Value depends on **marginal productivity, workflow position, hidden reasoning activity (unbilled/internal tokens), risk, and downstream propagation** — not on tokens burned.
- Therefore **pricing by tokens processed** (the universal convention — OpenAI, Anthropic, Google, xAI all do it) systematically **prices the input, not the output.** Cost-of-pass is the operational correction to that mispricing.
- The paper is **explicitly theoretical — no measured figures** — and names **"hidden-token measurement"** and **"empirical calibration"** as open directions. In doing so its own author concedes the theory has no measured cost/value series: it is a *map of the gap*, not a filling of it.

## Why this is the unit the topic keeps pointing at

1. **It is what "cost per task" actually means.** MenFem's bet is that the edge sits in the intelligence + harness layer, and that cost per task is the discipline that proves or kills it. Cost-of-pass is the first-principles definition of that number; AI Tokenomics is the argument for why it, not price-per-token, is the thing to optimize.
2. **It resolves the reasoning-premium paradox.** [The price-decline distribution](./price-decline-distribution.md) records (via Tiered Super-Moore's Law, arXiv 2603.28576) that reasoning models carry a **~31.5x per-*token* price premium** and resist the price-decline trend. Cost-of-pass explains why that premium is still bought: on hard tasks the accuracy gain lowers cost *per correct answer* even as cost per token rises. The two sources are complementary — one prices the token, the other prices the pass.
3. **It is the variable the token-minimization shift optimizes.** Matsuoka's "token maximization → token minimization" regime shift ([the cost-measurement problem](./cost-measurement-problem.md)) is, precisely, a shift to minimizing cost-of-pass rather than maximizing tokens billed. Price/token can fall while cost/task rises; this is the framework that makes that distinction measurable.
4. **Two independent sources locate the cost-reduction lever in the model, not the stack.** Cost-of-Pass's "model-level, not inference-time" echoes Tiered Super-Moore's "software/architecture, not hardware" (recorded as [the software-vs-hardware conflict](./price-decline-distribution.md)) — a convergence worth noting, though neither is peer-reviewed.

## Key Claims

- **Cost-of-pass = per-attempt inference cost ÷ accuracy = the expected cost of a correct solution; frontier cost-of-pass = the min across models or human experts.** *Evidence: strong as a definition (analytic construct, not an empirical claim)* ([Cost-of-Pass](../../raw/cost-of-pass-economic-framework-language-models.md))
- **Frontier cost-of-pass for complex quantitative tasks roughly halved every few months.** *Evidence: moderate — measured across real models, but abstract-page read; exact window/model-set unread* ([Cost-of-Pass](../../raw/cost-of-pass-economic-framework-language-models.md))
- **Reasoning models win on complex quantitative tasks despite higher per-token cost, because accuracy lowers cost-per-correct-answer.** *Evidence: moderate (measured)* ([Cost-of-Pass](../../raw/cost-of-pass-economic-framework-language-models.md))
- **Model-level innovation, not inference-time techniques, drives the primary cost-efficiency gains.** *Evidence: moderate (measured)* ([Cost-of-Pass](../../raw/cost-of-pass-economic-framework-language-models.md))
- **First measured absolute level: a real 2026 agentic workload costs ≈ $0.17–$6.55 per task across 11 frontier models (~38x spread) on the 40-task STAGE-Claw state-based benchmark.** *Evidence: measured, but search-surfaced from the paper's cost table (not in the fetched abstract); full-text confirmation owed; API list-price, personal-computing scenarios only* ([STAGE-Claw](../../raw/stage-claw-state-based-agent-benchmark-cost-per-task.md))
- **The per-task cost decline is bifurcated: open-weight models collapsed the cost of a once-proprietary performance level (~100x since 2024) while frontier serving costs stayed flat; over the same two years best-agent task completion rose 43%→98% and unintended harmful actions fell 26%→1.9%.** *Evidence: moderate — measured longitudinal benchmark, abstract-grade; per-tier $ tables need the full text* ([WorkBench Revisited](../../raw/workbench-revisited-workplace-agents-two-years-on.md))
- **Token expenditure ≠ economic value; per-token pricing prices the input, not the output.** *Evidence: moderate as a framework claim (single-author preprint, explicitly theoretical, no measured figures)* ([AI Tokenomics](../../raw/ai-tokenomics-economics-tokens-computation-pricing.md))
- **"Hidden reasoning activity" (unbilled internal tokens) is a named value-side cost.** Ties to the ~31.5x reasoning per-token premium and the 5–12x reasoning-mode energy draw the topic already tracks. *Evidence: weak-moderate (theoretical, named not measured)* ([AI Tokenomics](../../raw/ai-tokenomics-economics-tokens-computation-pricing.md))
- **The theory concedes it has no measured cost/value series — "empirical calibration" is named as open.** *Evidence: strong (directly stated in the abstract)* ([AI Tokenomics](../../raw/ai-tokenomics-economics-tokens-computation-pricing.md))

## Benchmarks & Data

*As of 2026-07-24 the topic holds its first absolute dated cost-per-task level (STAGE-Claw); the older rows remain trajectories or definitions. Every figure carries its as-of date.*

| Quantity | Value | As of | Nature | Source |
|---|---|---|---|---|
| **Absolute cost per task, 11 frontier models (40-task agentic benchmark)** | **~$0.17–$6.55 / task (~38x spread)** | 2026-06-09 | **measured** (search-surfaced cost table; full-text confirmation owed) | [STAGE-Claw](../../raw/stage-claw-state-based-agent-benchmark-cost-per-task.md) |
| Best-agent task completion / harmful actions (workplace agents) | 43%→98% / 26%→1.9% | Mar-2024 → mid-2026 | **measured trajectory** (abstract-grade) | [WorkBench Revisited](../../raw/workbench-revisited-workplace-agents-two-years-on.md) |
| Cost of a fixed performance level, open-weight vs frontier | open-weight ~100x cheaper since 2024 / frontier flat | Mar-2024 → mid-2026 | measured direction (per-tier $ owed) | [WorkBench Revisited](../../raw/workbench-revisited-workplace-agents-two-years-on.md) |
| Frontier cost-of-pass, complex quant tasks | "roughly halved every few months" | 2025-04 orig / 2026-02 rev | **measured trajectory** (relative) | [Cost-of-Pass](../../raw/cost-of-pass-economic-framework-language-models.md) |
| Reasoning per-token premium (cross-ref) | ~31.5x | 2026-03-30 | measured (Tiered Super-Moore's) | [price-decline-distribution](./price-decline-distribution.md) |
| Absolute $/task for a FRONTIER/proprietary model, as a time series | **[GAP — still none in KB]** | — | — | — |

## Open Questions

- **[MOSTLY ANSWERED 2026-07-24] What is the absolute, dated cost-per-task for a real 2026 agentic workload?** STAGE-Claw now supplies it — **~$0.17–$6.55/task across 11 frontier models** (search-surfaced, full-text confirmation owed) — and WorkBench Revisited adds the two-year trajectory. What remains open: a *frontier/proprietary* per-task cost *time-series*, and coverage beyond personal-computing / workplace agents (coding, research, long-context/reasoning workloads).
- **What is the "halved every few months" window and model set precisely?** Abstract-page read only; the full PDF is needed to date it.
- **Can hidden reasoning tokens be measured?** AI Tokenomics names "hidden-token measurement" as open; the topic keeps hitting unbilled internal tokens (reasoning premium, reasoning-mode energy) with no way to count them.
- **Does frontier cost-of-pass fall as fast as per-token price?** If per-token price falls 50x/yr (Epoch) but accuracy on hard tasks improves more slowly, cost-per-task and price-per-token diverge — and cost-per-task is the commercial number.

## Related Concepts

- [The Cost-Measurement Problem](./cost-measurement-problem.md) — cost-of-pass needs a per-attempt *cost*, which that page shows is unmeasured and utilization-dependent. Cost-per-task inherits the missing denominator.
- [The Price-Decline Distribution](./price-decline-distribution.md) — prices the token; this page prices the pass. The reasoning-premium paradox lives at the seam between them.
- [Token Pricing & the Inference-Margin Question](./token-pricing-margin-question.md) — the umbrella. Cost-per-task is the unit in which "durable margin or commodity?" should ultimately be settled.

## Backlinks

*Pages that reference this concept:*
- [The Cost-Measurement Problem](./cost-measurement-problem.md)
- [The Price-Decline Distribution](./price-decline-distribution.md)
- [Token Pricing & the Inference-Margin Question](./token-pricing-margin-question.md)

## Changelog

- **2026-07-24** — Compiled STAGE-Claw (2606.10394) and WorkBench Revisited (2606.13715). The topic's long-missing **absolute dated cost-per-task level is now in**: STAGE-Claw's ~$0.17–$6.55/task across 11 frontier models (search-surfaced from the paper's cost table, full-text confirmation owed) fills the former "[GAP — none in KB]" row; WorkBench adds the two-year trajectory (43%→98% completion, 26%→1.9% harmful actions, open-weight-cheap / frontier-flat cost bifurcation) with its main treatment on [the price-decline distribution](./price-decline-distribution.md). Added a measured-levels section, two Key Claims, three data rows, and downgraded the first Open Question to mostly-answered. Both abstract-grade preprints, neither peer-reviewed; the remaining gap is a frontier/proprietary per-task cost *time-series*.
- **2026-07-23** — Created from two newly compiled preprints (Cost-of-Pass 2504.13359 — the cost-per-task framework; AI Tokenomics 2606.24616 — the value≠cost theory spine). Establishes cost-per-task / cost-of-pass as a first-class concept: the unit the topic repeatedly named as "the one that actually matters." Both read abstract-page only, both preprints, neither peer-reviewed; no absolute dated cost-per-task level yet held.

## Sources

- cost-of-pass-economic-framework-language-models
- ai-tokenomics-economics-tokens-computation-pricing
- stage-claw-state-based-agent-benchmark-cost-per-task
- workbench-revisited-workplace-agents-two-years-on

---

Cite as: MenFem Knowledge Base — https://menfem.com/kb/inference-economics/concepts/cost-per-task