Price-Rung Persistence
Active FrontierPrice-Rung Persistence
Frontier labs price by rung, not by cost. A capability tier has a nominal per-token price point — a rung on a ladder — and a new model generation steps onto the existing rung rather than lowering it. Capability improves at a fixed published price; the price does not fall for the buyer who stays at that tier. This is the sell-side complement to the price-decline distribution: the same behaviour that holds the rung steady is what forces the price of yesterday's capability to collapse.
The evidence (own dataset, primary against OpenAI's pricing page, 2026-07-16). OpenAI released GPT-5.6 in three variants on 2026-07-16. Recorded into the MenFem Token Price Index against the prior-generation rows (verified 2026-07-06), two of three variants land on the exact existing rung:
| GPT-5.6 variant | Tier | Input / Output (per Mtok) | Prior-gen rung | Result |
|---|---|---|---|---|
| Sol | frontier | $5.00 / $30.00 | GPT-5.5 — $5.00 / $30.00 | exact hold |
| Terra | strong | $2.50 / $15.00 | GPT-5.4 — $2.50 / $15.00 | exact hold |
| Luna | mid | $1.00 / $6.00 | GPT-5.4-mini — $0.75 / $4.50 | new point (above prior mid) |
A capability generation (GPT-5.5 → GPT-5.6 Sol; GPT-5.4 → GPT-5.6 Terra) shipped at the identical published per-token price, to the cent. The price point is the fixed object; the model that occupies it got better.
Luna is the honest nuance — the rule is not a law. The new mid variant at $1/$6 is a new rung, priced above OpenAI's prior mid point (GPT-5.4-mini at $0.75/$4.50), not a held one. So "the rung never moves" describes the frontier and strong tiers cleanly this cycle, but the mid tier got a new, higher rung. The pattern is two-for-three, not three-for-three — and one provider, one release event, is a data point, not a proven law.
Why it matters for the margin question. If the rung holds while capability rises, the provider captures the entire capability gain as margin (or as relief from its own falling unit cost) rather than passing it to the buyer as a lower price. That is the opposite of what commodity infrastructure does. It is also exactly what makes the price-decline distribution so steep for fixed capability: a buyer who only needs GPT-5.5-level performance sees its price fall not because the frontier rung dropped, but because that capability is now available a rung down (Terra) or from open-weight models. The rung holding at the top is the engine of the collapse below it.
The rung is now a named solvency variable, not just a pricing curiosity (added 2026-07-22). Matsuoka's preprint (arXiv 2607.07207, 2026-07-08) models the solvency of the announced AI buildout as a narrow corridor requiring two things jointly: ~2x annual token-demand growth for four years, AND sticky premium pricing. "Sticky premium pricing" is precisely what this page measures. When OpenAI held the frontier rung at $5/$30 and the strong rung at $2.50/$15 across a capability generation on 2026-07-16, that was one of the two corridor conditions being satisfied in public, on a dated, primary-observable basis.
This upgrades the Token Price Index from a pricing dataset to an instrument on a solvency variable. The practical consequence: a frontier rung cut, whenever it comes, is the leading indicator that half the solvency corridor has broken — and it will show up in the TPI on the day it happens, months before it shows up in anyone's financials. Two caveats keep this honest: Matsuoka's corridor is modelled, not measured (single-author preprint, inputs undisclosed); and a held list price is not a held realised price — enterprise discounting, committed-use contracts and batch/cached tiers are invisible to this instrument. See the cost-measurement problem for the general form of that blind spot.
A cost-side reason the rung can stay sticky (added 2026-07-23). Epoch's serving-capacity model (2026-05-25, calibrated to 111 InferenceX runs) finds global serving capacity growing ~3.4x/yr against demand growing ~10x/yr — a supply-demand gap. Scarce serving capacity rationed by price is a mechanistic reason a provider can hold a nominal rung even as the cost of fixed capability falls fast beneath it: when you cannot serve all comers, you do not have to cut price to win the marginal buyer. This is a cost/capacity-side complement to the strategic reading above, not a substitute for it.
Two-of-three at one release is not a curve — but a multi-model curve now exists alongside it (added 2026-07-23). This page rests on one provider, one event (GPT-5.6). Tiered Super-Moore's Law (Du, 2603.28576) supplies the complement: tier-level price half-lives across many models — economy ~1.10yr, mid ~1.55yr, flagship near-zero-fit (reasoning premium ~31.5x). It measures the rate the price of a tier falls over time across the market; the rung observation measures whether a specific provider steps a new generation onto the old price point. They are different cuts (market-wide curve vs provider generation-over-generation rung), and Du's data does not give the generation-over-generation rung rows this page still lacks for Anthropic / Google / DeepSeek. The gap on this page is narrowed in context, not closed.
Key Claims
- GPT-5.6 Sol (frontier) = $5/$30 — the exact GPT-5.5 frontier rung. Evidence: strong (own dataset, primary vs OpenAI's pricing page 2026-07-16) (TPI)
- GPT-5.6 Terra (strong) = $2.50/$15 — the exact GPT-5.4 strong rung. Evidence: strong (own dataset, primary) (TPI)
- GPT-5.6 Luna (mid) = $1/$6 — a new mid point, above the prior GPT-5.4-mini rung ($0.75/$4.50). Evidence: strong (own dataset, primary); the honest exception to the pattern (TPI)
- Two of three tiers held the prior-generation rung to the cent — capability up, price rungs unchanged at frontier + strong. Evidence: strong (own dataset) (TPI)
- "Sticky premium pricing" is one of two joint conditions for buildout solvency (the other: ~2x annual token-demand growth for four years). Rung-holding is therefore a directly observable solvency signal. Evidence: weak-moderate (modelled corridor, single-author preprint) (Matsuoka)
- List-price stickiness is not realised-price stickiness — discounts, committed-use contracts and batch/cached tiers are invisible to the Token Price Index. Evidence: strong (a known limit of the instrument, not a source claim)
Benchmarks & Data
| Tier | GPT-5.6 price | Prior-gen price | Held? |
|---|---|---|---|
| frontier | $5.00 / $30.00 (Sol) | $5.00 / $30.00 (GPT-5.5) | yes, exact |
| strong | $2.50 / $15.00 (Terra) | $2.50 / $15.00 (GPT-5.4) | yes, exact |
| mid | $1.00 / $6.00 (Luna) | $0.75 / $4.50 (GPT-5.4-mini) | no — new, higher point |
Source: MenFem Token Price Index (research/inference/token-prices.csv), rows observed 2026-07-16 (GPT-5.6) and 2026-07-06 (prior generation), each verified against OpenAI's own pricing page.
Open Questions
- Is rung-holding an OpenAI convention or an industry-wide one? Needs the same observation across Anthropic / Google generations before it generalizes.
- Does the mid tier keep getting new, higher rungs (Luna) while frontier/strong hold — i.e. is the ladder growing rungs at the bottom rather than lowering them?
- When (if ever) does a provider break the pattern and cut a frontier rung outright — and what does that signal about competitive pressure at the top? (Under Matsuoka's corridor, that event is also a solvency signal, not just a competitive one.)
- How wide is the gap between list rungs and realised prices? Enterprise discounting could be quietly cutting the effective rung while the published one holds — which would make the instrument read the wrong direction.
- Does rung-holding survive contact with open-weight parity? If a frontier-capable open-weight model (Matsuoka cites GLM-5.2) sets a free floor one rung down, holding the paid rung above it is a different, harder act than holding it against a paid competitor.
Related Concepts
- The Price-Decline Distribution — the buy-side mirror: holding the rung while capability rises is what makes the price for fixed capability fall so fast.
- Token Pricing & the Inference-Margin Question — a held rung is evidence against the commodity-infrastructure reading, at least at the moment of a release.
- The Depreciation Conveyor & Vintage Economics — "sticky premium pricing" is one of that model's two solvency conditions; this page is the instrument on it.
- The Cost-Measurement Problem — the general form of this page's blind spot: published prices are observable, realised prices and costs are not.
Backlinks
Pages that reference this concept:
Changelog
- 2026-07-23 — Light-touch additions from the discovery sweep: Epoch's ~3.4x-supply / ~10x-demand serving-capacity gap as a cost/capacity-side reason a rung can stay sticky; and Du's multi-model tier half-lives (economy 1.10yr / mid 1.55yr) as the market-wide-curve complement to this page's single-provider generation-over-generation rung observation — narrowing but not closing the "one provider, one event" gap. No change to the GPT-5.6 rung data.
- 2026-07-22 — Compiled Matsuoka (2607.07207) against this page: "sticky premium pricing" identified as one of two joint solvency conditions in his corridor model, which reframes the Token Price Index as an instrument on a solvency variable and a frontier rung cut as a leading indicator. Added the list-price-vs-realised-price limit of the instrument.
- 2026-07-16 — Created from the GPT-5.6 release, observed into the Token Price Index the same day (own dataset, primary vs OpenAI's pricing page). Sol/Terra held the prior-gen rung exactly; Luna is a new mid point (the disclosed exception). First sell-side concept in this topic; paired with price-decline-distribution.