PAPER2025-04-17·Stanford University (inferred from author roster; not stated on abstract page)·arXiv 2504.13359

Cost-of-Pass: An Economic Framework for Evaluating Language Models

Mehmet Hamza Erol, Batu El, Mirac Suzgun, Mert Yuksekgonul, James Zou
COMPILED NOTES

Formalizes cost-per-TASK: 'cost-of-pass' = expected monetary cost of a CORRECT solution (per-token cost / accuracy); 'frontier cost-of-pass' = min across models or human experts. Frontier cost-of-pass for complex quantitative tasks 'roughly halved every few months'; model-level (not inference-time) innovation drives gains. The canonical framework for the unit this topic said actually matters. Original 2025-04-17, revised 2026-02-26.

Cost-of-Pass: An Economic Framework for Evaluating Language Models

What this is

The canonical academic framework for cost per task — the unit this topic's own frontier named as "the unit that actually matters" and had zero data on. Rather than pricing a token, it prices a correct answer. Submitted April 2025, revised February 2026 (so it is being actively maintained into the current window). It is the methodological backbone the KB was missing: a principled way to combine accuracy and inference cost into one economic number.

The framework

  • Cost-of-pass = "the expected monetary cost of generating a correct solution." Operationally, per-attempt inference cost divided by the probability the attempt is correct (accuracy). A cheap-per-token model that is often wrong can have a higher cost-of-pass than an expensive model that is usually right.
  • Frontier cost-of-pass = the minimum cost-of-pass achievable across all available models — or across models and human expert(s). This makes "is the AI economically worth it vs a human?" a computable question per task type.

Key findings (measured across real models over time)

  • Task-type specialization is economic, not just capability: lightweight models win on basic quantitative tasks; large models win on knowledge-intensive tasks; reasoning models win on complex quantitative problems despite higher per-token cost — because their higher accuracy lowers cost per correct answer.
  • Frontier cost-of-pass for complex quantitative tasks "roughly halved every few months."
  • Model-level innovations — not inference-time techniques (sampling, budget-aware decoding) — drive the primary cost-efficiency gains. Budget-aware methods (e.g. TALE-EP) help at the margin but are not the main lever.

Why it matters here

  • Closes the cost-per-task gap with a real framework. The KB held only Gartner's second-hand 5–30x agentic multiplier. Cost-of-pass gives a first-principles, measured way to state cost/task and compare models and humans on it.
  • Resolves the reasoning-premium paradox. Tiered Super-Moore's Law (2603.28576) shows reasoning models carry a ~31.5x per-token price premium and resist the price-decline trend. Cost-of-pass explains why that premium is still bought: on hard tasks, the accuracy gain lowers cost per correct answer even as cost per token rises. The two sources are complementary — one prices the token, the other prices the pass.
  • Directly engages token-minimization. Under the token-maximization → token-minimization regime shift (Matsuoka), the optimized variable is exactly cost-of-pass, not price-per-token — price/token can fall while cost/task rises. This is the framework that makes that distinction measurable.
  • "Model-level not inference-time" echoes Tiered Super-Moore's "software/architecture, not hardware" decomposition — two independent sources locating the cost-reduction lever in the model, not the serving stack or the silicon.

Limitations

Abstract-page read only. Earliest-dated source in the topic (2025-04 original), though revised 2026-02; the specific model set and the "halved every few months" window need the full PDF to date precisely. Cost figures are relative-trajectory, not absolute dated $/task levels.


Source: Cost-of-Pass: An Economic Framework for Evaluating Language Models by Erol, El, Suzgun, Yuksekgonul & Zou, arXiv 2504.13359, submitted 2025-04-17 / revised 2026-02-26

RELATED · IN THE BASE
Cost-of-Pass: An Economic Framework for Evaluating Language Models | Knowledge Base | MenFem