# Token Pricing & the Inference-Margin Question

Canonical URL: https://menfem.com/kb/inference-economics/concepts/token-pricing-margin-question
Knowledge base topic: [Inference & Token-Pricing Economics](https://menfem.com/kb/inference-economics)
Frontier status: active
Tags: token-pricing, inference-economics, gross-margin, commodity-infrastructure

---

The central open question in AI-infrastructure economics: does inference settle into a durable-margin business, or does it become low-margin commodity infrastructure with value captured elsewhere in the stack (as happened to mobile carriers)? Two independent, well-evidenced-but-incomplete answers are in tension this week, plus one attributed data point that cuts the other way.

**Empirically, token prices are falling fast but unevenly.** Epoch AI's cross-benchmark study finds a **median 50x/year** decline in the price required to hit a fixed performance bar, with the rate ranging from **9x to 900x/year** depending on which benchmark and threshold is used. The often-quoted "GPT-4-level" milestone (GPQA Diamond, PhD-level science) specifically falls at **40x/year — below the all-benchmark median**, not above it; restricting to models released after January 2024 pushes the median up to **200x/year**, suggesting the pace is accelerating, though that faster regime has too short a track record to call durable. *Evidence: strong (methodology-documented benchmark study).*

**Structurally, Ben Evans (2026-07-09) argues this instability is the point, not a bug.** Inference currently runs at **40–50% gross margins**, but that figure excludes training cost, "which is currently far larger than revenue" — meaning headline inference profitability is not lab profitability. Evans frames the outcome as resting on four unresolved questions (frontier adoption breadth, whether frontier capability progress can outrun efficiency/pricing pressure, competitive concentration vs. fragmentation, and where value is captured — model layer vs. wrapping software) and reaches for the **mobile-data-network analogy**: cellular traffic grew "several orders of magnitude" into a ~$1T/year industry on ~$200B capex, and "the stocks have gone nowhere" — the carriers built the pipes, other layers captured the value. He also notes TSMC's near-monopoly on frontier logic manufacturing still nets it "$53bn [net income], less than half of Apple alone" as a second monopoly-without-majority-value-capture example. *Evidence: moderate (single analyst essay, no primary financial data cited).*

**Patel's attributed claim points the other way for at least one lab.** On Podcast Alpha (2026-07-10), Dylan Patel of SemiAnalysis claims **Anthropic is FCF-positive, past $50B ARR, at 70%+ gross margin** — well above Evans' generic 40–50% figure. This is **Patel claims/estimates, unverified**, and carries a disclosed conflict: SemiAnalysis is reportedly an Anthropic enterprise customer (~$7M/year on Claude Code), giving Patel motivated visibility but not independence. *Evidence: weak (single attributed source, disclosed conflict, no primary financials, paywalled beyond preview).*

**The tension is the story, not a resolved answer.** Evans' 40-50% figure is presented as an industry-wide, training-cost-excluded estimate; Patel's 70%+ figure is lab-specific (Anthropic only) and unverified. They are not strictly contradictory — a leading lab could plausibly outperform an industry-wide average — but the 30-point gap is exactly the kind of number that needs a primary-source correction (an actual Anthropic financial disclosure) before it can be treated as fact either way.

**As of 2026-07-22, both margin figures are undercut at the same point: neither has a cost denominator.** Patil's concurrency-aware measurement (arXiv 2606.11690, 2026-06-10) shows that on *identical H100 hardware*, effective cost spans **$0.21 to $15.25 per million output tokens** — a **2.5-24x** penalty across ordinary enterprise loads (1-10 rps) and up to **36.3x near idle** — driven entirely by offered request rate. Neither Evans' 40-50% nor Patel's 70%+ states a utilization assumption, a cost unit, or an as-of date. This does not make either figure *wrong*; it makes both **unfalsifiable as stated**. The gap between them (30 points) is smaller than the range Patil measures on one SKU. Before this question can be answered, the denominator has to be defined — which is now its own concept, [the cost-measurement problem](./cost-measurement-problem.md). *(Ratio findings are the robust part; the absolute $ level rests on a $/GPU-hour input not disclosed in the abstract.)*

**As of 2026-07-23, a measured cost floor exists — for one open model, not for any frontier lab.** DigitalOcean measured a 70B open model (Llama-3.3-70B FP8) on a single H200 at **$20.32/M at batch=1 → $0.45/M at batch=128** (~44x from traffic shape alone; 2026-07-08, throughput measured, $/M derived at $3.44/GPU-hr list). This is the topic's first measured cost-to-serve series and it independently reproduces Patil's utilization thesis in dollars. **But it changes the margin question only for open models.** The Evans-vs-Patel debate is about serving GPT-5.x / Claude / Gemini — for which **no measured cost floor exists anywhere in this KB.** So the sharpest remaining gap is precise: *there is now a measured cost floor for one open 70B model and none for any frontier/proprietary model, so the Evans-40% vs Patel-70% lab-margin debate stays unfalsifiable.* One corroborating cross-check does land: the cheapest independent-provider list price (Together $0.65/M batch, 70B) sits **on** DigitalOcean's measured ~$0.45–$0.64/M floor — thin-to-zero margin at the batch tier for open models ([independent providers](../entities/independent-inference-providers.md)).

**The unit itself is probably wrong — and there is now a framework for the right one.** AI Tokenomics (Zhu, arXiv 2606.24616, 2026-06-10, theoretical) argues **token expenditure ≠ economic value**: per-token pricing prices the input, not the output. Cost-of-Pass (Erol et al., arXiv 2504.13359, rev 2026-02) gives the operational correction — **cost per *task*** = per-token cost ÷ accuracy — and finds frontier cost-of-pass on hard quantitative tasks "roughly halved every few months," driven by model-level (not inference-time) innovation. This is the unit in which "durable margin or commodity?" should ultimately be settled, and it is now [its own concept](./cost-per-task.md). The catch the KB keeps hitting: cost-of-pass needs a per-attempt *cost*, which is exactly the missing, utilization-dependent denominator above. *(Both abstract-page reads; both preprints; no absolute dated cost-per-task level held.)*

**The May-2024 competition break weakly favours Evans' pole.** Tiered Super-Moore's Law (Du, arXiv 2603.28576) dates a structural shift (Chow F=5.74, p=0.005) from technology-driven to *competition-driven* price decline, with concentration falling (HHI 4,558 → 2,086). That supports Evans' "compresses toward marginal cost as competition bites" reading over the "rotating-landlord oligopoly" pole — and puts a date on it. It also carries the topic's new first-order conflict: Du attributes ~103.7% of the decline to software/TFP and –0.9% to GPU hardware, cutting against the memory-as-destiny framing (see [the depreciation conveyor](./depreciation-conveyor.md) and [frontier Conflict 7](../frontier.md#conflicts-on-the-record)). *(Estimated econometrics, abstract-page, single-author preprint.)*

**The first quantitative model of the question prices it as a coin flip.** Matsuoka's preprint (arXiv 2607.07207, 2026-07-08) is the topic's first paper-grade engagement with "commodity infrastructure or durable margin?", and its five-scenario probability set puts the two poles **tied at 25% each** — Rotating Landlord Oligopoly 25%, Commoditization Crash 25%, then Jevons Absorption 20%, System-Layer Re-differentiation 18%, Geopolitical Bifurcation 12%. These are author-assigned priors from a single-author preprint with no stated elicitation method, so they are weak evidence about the world — but they are a clear statement that the question is **not currently decidable**, from the one source that tried hardest to decide it. The mechanism it contributes — the **depreciation conveyor**, under which the entrant-incumbent cost gap never closes (3.2x in 2026, 1.9x in 2027, re-widening to 3-4x by 2029-30, all modelled) — argues for durable *relative* advantage among incumbents even inside a commoditizing market, and is tracked at [the depreciation conveyor](./depreciation-conveyor.md). Note that the two questions are separable: inference can become a low-margin commodity in aggregate while incumbents remain permanently cheaper than entrants.

**One primary data point now cuts against the pure-commodity reading — the sell-side rung.** When OpenAI shipped GPT-5.6 (2026-07-16), two of its three variants landed on the *exact* prior-generation price rung: Sol (frontier) at $5/$30 = GPT-5.5's rung, Terra (strong) at $2.50/$15 = GPT-5.4's rung (Luna, the new mid variant at $1/$6, is a *new* rung, the honest exception). Capability rose; the published price for the tier did not fall. A provider holding its rung while the model improves is capturing the capability gain rather than passing it through as a lower price — the opposite of what commodity infrastructure does at the moment of a release. This does not resolve the margin question (held list prices say nothing about the cost side or about realized, discount-adjusted revenue), but it is a primary, own-dataset observation, and it sits directly alongside Epoch's buy-side decline data: the rung holding at the top is the same mechanism that makes the price of *fixed* capability collapse below it. These two halves are now first-class concepts — [the price-decline distribution](./price-decline-distribution.md) (buy-side) and [price-rung persistence](./price-rung-persistence.md) (sell-side).

## Key Claims

- **Median inference price decline: 50x/year, range 9x-900x/year** across 6 benchmarks. *Evidence: strong* ([Epoch AI](../../raw/epoch-ai-llm-inference-price-trends.md))
- **GPT-4-level PhD-science milestone: 40x/year — below the all-benchmark median.** *Evidence: strong* ([Epoch AI](../../raw/epoch-ai-llm-inference-price-trends.md))
- **Post-Jan-2024-only median: 200x/year** (accelerating, short track record). *Evidence: moderate* ([Epoch AI](../../raw/epoch-ai-llm-inference-price-trends.md))
- **Inference gross margin ~40-50%, excludes training cost which exceeds revenue.** *Evidence: moderate* ([Ben Evans](../../raw/ben-evans-ways-to-think-about-token-pricing.md))
- **Mobile-data analogy: ~$1T/yr traffic on ~$200B capex, carrier stocks flat — value captured elsewhere.** *Evidence: moderate (historical analogy, not AI-specific data)* ([Ben Evans](../../raw/ben-evans-ways-to-think-about-token-pricing.md))
- **TSMC near-monopoly nets $53B net income, less than half of Apple alone.** *Evidence: moderate* ([Ben Evans](../../raw/ben-evans-ways-to-think-about-token-pricing.md))
- **Anthropic FCF-positive, $50B+ ARR, 70%+ gross margin.** *Evidence: weak — Patel claims/estimates, unverified, disclosed conflict of interest* ([Dylan Patel / Podcast Alpha](../../raw/dylan-patel-podcast-alpha-2026-07-10.md))
- **GPT-5.6 held two of three prior-generation price rungs to the cent** (Sol $5/$30, Terra $2.50/$15; Luna a new mid rung) — capability up, list price flat, cutting against the pure-commodity reading. *Evidence: strong (own dataset, primary vs OpenAI's pricing page 2026-07-16)* ([TPI](../../raw/tpi-gpt-5-6-rung-hold-2026-07-16.md))
- **Neither margin figure is falsifiable as stated — no utilization assumption, no cost unit, no as-of date on either.** *Evidence: strong (directly observable absence in both sources)* ([Evans](../../raw/ben-evans-ways-to-think-about-token-pricing.md), [Patel](../../raw/dylan-patel-podcast-alpha-2026-07-10.md))
- **Effective cost on identical H100 hardware spans $0.21-$15.25 / M output tokens (as of 2026-06) — a spread wider than the entire Evans-vs-Patel margin disagreement.** *Evidence: moderate (measured ratios; absolute level rests on an undisclosed $/GPU-hour input)* ([Patil](../../raw/beyond-per-token-pricing-concurrency-aware-cost-methodology.md))
- **A measured cost floor now exists for one open 70B model (~$0.45/M saturated, H200, 2026-07-08) — but for NO frontier/proprietary model.** The lab-margin debate is about GPT-5.x / Claude / Gemini serving; that cost is unmeasured, so Evans-40% vs Patel-70% stays unfalsifiable. *Evidence: strong (measured floor exists; frontier absence directly observable)* ([DigitalOcean](../../raw/measured-h200-token-serving-cost-digitalocean.md))
- **The commercially right unit is cost per TASK (cost-of-pass = per-token cost ÷ accuracy), not price per token — token expenditure ≠ economic value.** No absolute dated cost-per-task level is held. *Evidence: moderate as frameworks (two abstract-page preprints, neither peer-reviewed)* ([Cost-of-Pass](../../raw/cost-of-pass-economic-framework-language-models.md), [AI Tokenomics](../../raw/ai-tokenomics-economics-tokens-computation-pricing.md))
- **A May-2024 break dates the decline shifting from technology- to competition-driven (HHI 4,558→2,086), weakly favouring Evans' compression pole; the same source attributes ~103.7% of the decline to software/TFP and –0.9% to hardware (CONFLICT with memory-as-destiny, unresolved).** *Evidence: estimated econometrics, abstract-page, single-author preprint* ([Du](../../raw/tiered-super-moores-law-inference-price-evolution.md))
- **The first quantitative scenario model of this question ties the two poles at 25% each** (Commoditization Crash 25% / Rotating Landlord Oligopoly 25%; Jevons Absorption 20%, System-Layer Re-differentiation 18%, Geopolitical Bifurcation 12%). *Evidence: weak (author-assigned priors, single-author preprint, no elicitation method)* ([Matsuoka](../../raw/memory-scarcity-open-models-ai-industry-restructuring-2026-2030.md))
- **Buildout solvency modelled as a two-condition corridor: ~2x annual token-demand growth for four years AND sticky premium pricing.** *Evidence: weak-moderate (modelled)* ([Matsuoka](../../raw/memory-scarcity-open-models-ai-industry-restructuring-2026-2030.md))
- **Inference cost may pass through to economy-wide inflation at near-unit elasticity (US slope 0.087, HAC s.e. 0.021, data to 2026-04).** *Evidence: weak — unreplicated single-author preprint; undisclosed cost series; implausibly clean scaling fit* ([ICPC preprint](../../raw/economics-of-ai-inference-phillips-curve.md))

## Open Questions

- **What does a FRONTIER model cost to serve?** A measured cost floor now exists for one open 70B (DigitalOcean, ~$0.45/M) but for no frontier/proprietary model — which is what the Evans-vs-Patel debate is actually about. Until that exists, both lab-margin figures stay unfalsifiable *even with an open-model floor in hand*. Now the topic's sharpest single gap — see [the cost-measurement problem](./cost-measurement-problem.md).
- **Is the right unit even the token?** Cost-per-task (cost-of-pass) and the value≠cost theory say no — see [cost per task](./cost-per-task.md). No absolute dated cost-per-task level is held for a real 2026 workload.
- Is Anthropic's reported 70%+ margin (Patel, attributed) real, or an artifact of Patel's customer-side vantage point? No *audited* primary financial disclosure exists in this KB to check it; a confidential draft S-1 gives a date-certain resolution path (see [frontier.md](../frontier.md)).
- Does the post-2024 200x/year price-decline pace persist, or was 2024-2025 a one-off efficiency unlock (e.g. MoE, quantization, speculative decoding) that won't repeat?
- Which of Evans' four variables (frontier adoption breadth, frontier-progress-vs-efficiency race, competitive concentration, value-capture layer) is moving fastest right now?
- **Are "does inference commoditize?" and "does the incumbent cost advantage persist?" the same question?** Matsuoka's conveyor implies not — a market can commoditize on price while incumbents stay structurally cheaper than entrants.
- **Is the right unit the token at all?** Matsuoka argues $/PB of bandwidth delivered is model-agnostic for bandwidth-bound decode; every other figure in this topic is per-token.

## Related Concepts

- [The Price-Decline Distribution](./price-decline-distribution.md) — the buy-side empirical: how fast the price of *fixed* capability falls (Epoch's median 50x/yr, 9x-900x range). Split out from this umbrella concept.
- [Price-Rung Persistence](./price-rung-persistence.md) — the sell-side empirical: providers hold nominal per-token rungs while capability rises (GPT-5.6). The mirror image of the decline distribution.
- [The Cost-Measurement Problem](./cost-measurement-problem.md) — the missing denominator under every margin claim on this page. Logically upstream of this question.
- [The Depreciation Conveyor & Vintage Economics](./depreciation-conveyor.md) — the first quantitative structural answer offered: incumbents stay cheaper than entrants regardless of where price settles. Now also carries the software-vs-hardware conflict.
- [Cost Per Task (Cost-of-Pass)](./cost-per-task.md) — the unit in which this question should ultimately be settled: price a correct answer, not a token.
- [Inference Cost as a Macro Input](./inference-cost-macro-channel.md) — the "does this matter outside the lab P&L?" branch. Speculative, one preprint.
- [HBM4 Memory Architecture](../../../hardware/wiki/concepts/hbm4-memory-architecture.md) — memory cost/availability is a direct input to inference cost structure; see Patel's attributed structural-shortage claim there.
- [Co-Packaged Optics](../../../optical-computing/wiki/concepts/co-packaged-optics.md) — interconnect cost/timing is a second direct input to inference cost structure; see Patel's attributed 2029-vs-2027 ramp claim there.

## Backlinks

*Pages that reference this concept:*
- [Anthropic](../../../ai/wiki/entities/anthropic.md) — Patel's attributed margin claim lives on the Anthropic entity page

## Changelog

- **2026-07-23** — Compiled the 6-source discovery sweep against the umbrella. Added: a measured open-model cost floor exists (DigitalOcean H200, ~$0.45/M) but no frontier-model floor — so the lab-margin debate stays unfalsifiable (now the sharpest gap); cost-per-task / cost-of-pass named as the commercially right unit (new concept, value≠cost from AI Tokenomics); the cheapest independent-provider list price sitting on the measured floor (thin batch margin); and the May-2024 competition break (weakly favouring Evans) plus its software-vs-hardware conflict. Four new source links, new Key Claims and Open Questions.
- **2026-07-22** — Compiled 3 arXiv preprints (Matsuoka 2607.07207, Patil 2606.11690, ICPC 2605.20281). Both margin figures are now recorded as unfalsifiable-as-stated (no cost denominator); added Matsuoka's 25%/25% scenario tie and the two-condition solvency corridor; spun out three new concepts — cost-measurement-problem, depreciation-conveyor, inference-cost-macro-channel. This page stays the umbrella; the cost denominator now sits logically above it.
- **2026-07-16** — Split the Epoch price-decline data into its own concept ([price-decline-distribution](./price-decline-distribution.md)) and added its sell-side mirror ([price-rung-persistence](./price-rung-persistence.md)) from the GPT-5.6 rung-hold. Added the GPT-5.6 rung as a primary data point cutting against the pure-commodity reading; this page is now the umbrella question over the two empirical concepts.
- **2026-07-14** — Initial compilation from 3 sources (Ben Evans essay, Epoch AI data insight, Patel/Podcast Alpha attributed claims). First entry in this new topic.

## Sources

- ben-evans-ways-to-think-about-token-pricing
- epoch-ai-llm-inference-price-trends
- dylan-patel-podcast-alpha-2026-07-10
- memory-scarcity-open-models-ai-industry-restructuring-2026-2030
- beyond-per-token-pricing-concurrency-aware-cost-methodology
- economics-of-ai-inference-phillips-curve
- measured-h200-token-serving-cost-digitalocean
- epoch-inference-compute-crunch-serving-capacity
- tiered-super-moores-law-inference-price-evolution
- cost-of-pass-economic-framework-language-models
- ai-tokenomics-economics-tokens-computation-pricing
- independent-inference-provider-pricing-matrix-q2-2026

---

Cite as: MenFem Knowledge Base — https://menfem.com/kb/inference-economics/concepts/token-pricing-margin-question