OpenAI cuts GPT-5.6 Luna 80% and Terra 20%, and says why (announcement + API changelog + pricing page)
Luna -80% ($1.00/$6.00 -> $0.20/$1.20 per Mtok in/out) and Terra -20% ($2.50/$15 -> $2.00/$12) on 2026-07-30, with a vendor-named software cause: 20% lower serving cost from GPU kernel work and 15%+ better token generation from speculative decoding. Same changelog: launch 2026-07-09; Sol cut to $4/$20 on 2026-08-21 (promotional, at least through 2026-11-21); GPT-6 Sol $2/$10 on 2026-09-22.
What this is, and how deeply it was read
On 30 July 2026, three weeks after the GPT-5.6 family launched, OpenAI cut the price of its two cheaper GPT-5.6 models and said what made it possible. This close-read covers that one event from the OpenAI surfaces that could be read on 2026-09-24:
| Surface | Read? | What it gave |
|---|---|---|
| Community forum announcement (moderator VeitB, 2026-07-30 18:03 UTC) | Full (thread JSON) | The stated cause, the two cuts, Fast mode |
| OpenAI API changelog | Full for Jul 9 – Sep 22 | Launch date, cut date and size, the later Sol cut, GPT-6 prices |
| OpenAI API pricing page (live, 2026-09-24) | Full for the GPT-5.6 and GPT-6 rows | Today's prices, used to confirm the new rungs |
| @OpenAI on X, 2026-07-29 21:21 UTC (via public mirror) | Full (short post) | The stated cause, in OpenAI's own account |
| openai.com blog, "Advancing the price-performance frontier with GPT-5.6" | Not read — HTTP 403; archive copy unreachable | — |
The 2026-09-05 discovery pass took these prices from secondary reporting and marked them [unverified]. They are now checked against OpenAI's own pricing page and changelog, and they hold.
The cut
OpenAI's changelog, 30 July: "Starting July 30, GPT-5.6 Luna costs 80% less, while GPT-5.6 Terra costs 20% less."
| Model | Before (launch price, per million tokens) | After (read on OpenAI's pricing page 2026-09-24) | Change |
|---|---|---|---|
| GPT-5.6 Luna | $1.00 in · $0.10 cached · $6.00 out | $0.20 in · $0.02 cached · $1.20 out | −80% on every line |
| GPT-5.6 Terra | $2.50 in · $0.25 cached · $15.00 out | $2.00 in · $0.20 cached · $12.00 out | −20% on every line |
| GPT-5.6 Sol | $5.00 in · $0.50 cached · $30.00 out | unchanged on 2026-07-30 (but see 21 August below) | — |
The "before" prices are the launch rows the rung already held (Token Price Index, 2026-07-16); a forum reply tabulating old against new gives the same numbers. The "after" prices were read directly from OpenAI's pricing page and match the changelog's 80% and 20%.
Why OpenAI says it could cut
From the forum announcement: "OpenAI tasked GPT-5.6 with optimizing its own runtime efficiency. The results: 20% lower serving costs through production GPU kernel improvements. More than 15% better token-generation efficiency through improved speculative decoding. These improvements are being passed on to everyone using the API, Codex, and ChatGPT."
OpenAI's own post on X the evening before says the same, naming the model that did the work: "we applied GPT-5.6 Sol to advance the frontier of efficiency by making itself more efficient to run. ... 20% lower serving costs from production GPU kernel improvements. 15%+ better token-generation efficiency from improved speculative decoding."
In plain words: two software changes — faster GPU code, and a better way of drafting several words at once and checking them together (speculative decoding) — made each token cheaper to produce on the same chips. Nothing in the disclosure mentions new hardware.
What the numbers do not add up to. OpenAI names a ~20% cost fall and cut the price of one model by 20% and another by 80%. A 20% serving saving explains Terra's 20% cut. It does not explain Luna's 80%: the extra 60 points are a pricing decision, not a cost pass-through. The disclosure gives no dollar cost for either model, so the gap cannot be measured, only noted.
Fast mode (same announcement)
From the changelog: Fast mode "replaces our Priority Processing offering. For GPT-5.6 Sol, Fast mode now delivers up to 2.5× faster speeds than standard processing at twice the price." On the pricing page the Fast (priority) rows for GPT-5.6 are exactly double the standard rows (Sol $8/$40 against the current $4/$20; Terra $4/$24; Luna $0.40/$2.40). Speed is now its own priced line: the same model, the same tokens, twice the price for up to 2.5× the speed.
What came next — from the same changelog (dated observations, not the 30 July event)
These were read from the same OpenAI changelog and pricing page. They are recorded here because they change what the 30 July cut means for the rung's claims; each carries its own date.
- 9 July 2026 — launch. "Released the GPT-5.6 model family, including GPT-5.6 Sol ... GPT-5.6 Terra ... and GPT-5.6 Luna." This settles the launch date the rung had as "unsettled": 9 July, not 16 July. (16 July is when the Token Price Index recorded the rows.)
- 21 August 2026 — the frontier model is cut too. "GPT-5.6 Sol now costs $4 per million input tokens and $20 per million output tokens, representing 20% lower input pricing and 33% lower output pricing. GPT-5.6 Sol's promotional pricing is available at least through November 21, 2026." The word promotional matters: this is a dated discount, not (yet) a permanent new price. Confirmed on the pricing page 2026-09-24: Sol $4.00 in · $0.40 cached · $20.00 out.
- 22 September 2026 — the next generation arrives below the old frontier price. "Released GPT-6 Sol (gpt-6-sol) and GPT-6 Luna (gpt-6-luna) ... GPT-6 Sol: $2 input, $0.20 cached input, and $10 output. GPT-6 Luna: $0.10 input, $0.01 cached input, and $0.50 output." The pricing page also lists GPT-6 Astra at $10 in · $1 cached · $50 out — a new, higher top price, the same $10/$50 point Anthropic's Fable sits at.
Through-line: how this moves a number on the path from token price to task price
- Token price, cheap tier: the price of a Luna-class token fell 80% in one day. A job that used 1 million input and 200,000 output tokens on Luna cost $1.00 + $1.20 = $2.20 before 30 July and $0.20 + $0.24 = $0.44 after (arithmetic on OpenAI's list prices; no caching or batch discount).
- Cause: the named cause is software on existing chips, which bears on the rung's recorded software-versus-hardware conflict: it is a vendor-stated case of the cost fall coming from kernels and decoding, not silicon.
- Rungs: it is the primary source under the rung's record that OpenAI's strong and mid rungs broke on 30 July, and the same changelog shows the frontier rung moving on 21 August.
Limitations
- Vendor source. OpenAI is reporting its own costs, as percentages with no dollar base. The 20% and 15% cannot be checked from outside.
- Two announcement surfaces, one missing. The formal openai.com post was not read. The forum post is by a community moderator flagged as staff on the forum; the same claims appear in OpenAI's own X post and changelog, which is why the cause is attributed to OpenAI rather than to the moderator.
- List prices only. Batch (half price), Flex, caching and negotiated rates are not what a given customer paid.
- The Sol cut is promotional, with no stated end beyond "at least through November 21, 2026". Cite it with that caveat.
- Prices go stale fast. Every figure here is as of the date shown; never cite one as today's price without re-reading the page.
Sources: OpenAI community forum announcement (2026-07-30) · OpenAI API changelog · OpenAI API pricing · @OpenAI on X (2026-07-29) · OpenAI blog post (not read, 403). All read 2026-09-24.