The Energy Cost of Tokens (Photons = Tokens)
The Energy Cost of Tokens (Photons = Tokens)
Every other cost lever in this topic — GPU-hours, HBM bandwidth, memory vintage — sits downstream of one physical input the KB had, until this compile, no source for: electricity. The topic's frontier tracked a named hole — "no energy / power-cost pass-through leg" — because it held web-literature figures (~5×10⁻⁴ Wh/token, reasoning-mode 5–12×, ~$254/mo electricity per H100) but nothing that connected energy to tokens inside a source. Photons = Tokens (Litowitz, Polson & Sokolov, arXiv 2603.06630, 2026-02-23) fills the leg — partly — by treating the token as a physical quantity with a measurable thermodynamic cost and building a supply/demand balance sheet for global token production.
It is filed early stage, and the read-depth is honest: abstract-grade, single preprint, not peer-reviewed, and the figures are explicitly order-of-magnitude policy-framing estimates, not a measured per-SKU cost series. What it changes is that the topic now has an energy→token bridge inside a citable source, and a clear statement of where the physical floor beneath the price sits.
The method — arithmetic, not rhetoric
The paper's stated move is to apply MacKay's (2009) Sustainable Energy — Without the Hot Air method (reframe an energy debate as arithmetic) to AI computation. It defines the token — the elementary unit of LLM input/output — as a physical quantity bounded from two directions:
- A thermodynamic floor via Landauer's principle — irreversible bit operations have a minimum energy cost set by temperature (kT ln 2 per erased bit).
- An information ceiling via Shannon's channel capacity — the maximum information a channel can carry.
Between those bounds, plus current infrastructure data, it constructs a supply-and-demand balance sheet for global token production, and from it derives a "finite question budget": the number of meaningful queries humanity can direct at AI systems under physical, information-theoretic, and economic constraints.
The numbers — the energy → token bridge
All three figures are the paper's own, at current efficiency, order-of-magnitude:
- 326 TWh — the projected 2028 US AI energy allocation — could support
- ≈ 6.5×10¹⁷ tokens per year, or
- ≈ 225,000 tokens per person per day — the per-capita token budget that energy allocation implies —
- > 3 orders of magnitude above estimated mid-2024 utilization.
The moral the paper draws (the contrarian teaching point): at projected 2028 US energy, token volume is abundant, not scarce. Energy is therefore not the near-term binding constraint on how many tokens can be produced. The paper relocates the true scarcity from capacity to direction — "the decisive variable is not how many questions can be answered but which questions are worth asking," a problem of agency that computation alone cannot solve. That is a direct answer to anyone who models the AI economy as energy-gated in the near term: on these numbers, it is demand-worthiness-gated, not supply-gated.
DERIVED (desk, flagged as NOT an independent confirmation). The implied efficiency behind the bridge is 326 TWh ÷ 6.5×10¹⁷ ≈ 5.0×10⁻⁴ Wh/token — which lands exactly on the bottom-up ~5×10⁻⁴ Wh/token web-literature figure the frontier gap already logged. Read this as a consistency check on the units, not as two independent measurements agreeing: the top-down figure was almost certainly built from that per-token efficiency, so the agreement is definitional, not corroborative. The honest statement is that the top-down (TWh → token-count) and bottom-up (Wh/token) framings are the same coin.
What it closes, and what it does not
Closes: the KB now holds an energy source and an energy→token volume bridge — the "no energy leg ingested" half of the frontier gap.
Does not close: this is a top-down TWh → token-count basis, not a bottom-up $/kWh → $/token per-SKU cost pass-through. The paper prices energy against token volume, not token dollars. So it bounds the pass-through leg only loosely — if energy is abundant relative to token demand, the electricity term is unlikely to be the binding variable in near-term $/token — but it does not give a measured $/kWh → $/token conversion, and it does not touch the reasoning-mode 5–12× energy multiplier that the same web literature flags (reasoning workloads are the growing share, and "Photons = Tokens" is a single-rate abstraction that averages over them). That measured $/token energy conversion, and a peer-reviewed energy source (the Joule/Cell paper is still 403'd — see frontier.md), remain owed.
Where it sits against the rest of the topic
- It is the energy analog of Matsuoka's $/PB unit (the cost-measurement problem). Matsuoka reframes decode cost in bandwidth delivered; Photons = Tokens reframes it in power delivered. Both are attempts to price the token in a model-agnostic physical unit rather than a billing convention. The value chain the paper maps — photon → atom → chip → power → token → question — is the same stack the depreciation/memory story runs on, read from the energy end.
- It puts a physical floor under the price series. The price-decline distribution measures how fast the price of a token falls; this page says the thermodynamic cost has a Landauer floor and the volume has an energy ceiling — the price can keep falling toward, but not below, the physical floor.
- It is a volume/abundance claim, which is the mirror image of the depreciation conveyor's cost-structure claim: one says tokens are physically cheap and abundant, the other says incumbents stay structurally cheaper than entrants. Both can hold at once — abundant tokens with a durable incumbent cost edge.
- It applies Coase's theory of the firm and the durable-goods-monopoly problem to locate where economic rent concentrates in that chain, and connects measurement limits to a Goodhart's-law / Heisenberg parallel and to Arrow's impossibility result for efficient information pricing — theory-driven, not empirical, and recorded as such.
Key Claims
- The token is a physical quantity with a measurable thermodynamic cost — a Landauer floor (min energy per irreversible bit op) and a Shannon information ceiling. Evidence: strong as a framework (first-principles physics); its application to real serving is the modelled part (Photons = Tokens)
- The projected 2028 US AI energy allocation of 326 TWh could support ≈ 6.5×10¹⁷ tokens/year ≈ 225,000 tokens/person/day at current efficiency — > 3 orders of magnitude above mid-2024 utilization. Evidence: moderate — the paper's own order-of-magnitude estimate, abstract-grade, single preprint; "at current efficiency" is load-bearing and unstated in the abstract beyond the ~5×10⁻⁴ Wh/token implied rate (Photons = Tokens)
- On these numbers, energy is NOT the near-term binding constraint on token volume; the binding constraint is which questions are worth asking (direction, not capacity). Evidence: moderate as a framing conclusion (follows from the balance sheet; a policy argument, not a measurement) (Photons = Tokens)
- The paper supplies a top-down TWh → token-count basis, NOT a bottom-up $/kWh → $/token per-SKU cost pass-through. Evidence: strong (directly observable in what was read — the estimates are token-volume, not token-dollars) (Photons = Tokens)
- "Photons = tokens" is a single-rate modelling abstraction; real serving efficiency varies by model/hardware/workload, with reasoning-mode drawing 5–12× more. Evidence: the abstraction is stated by the paper; the 5–12× multiplier is web-literature the frontier gap logs, not a Photons = Tokens measurement (Photons = Tokens)
- DERIVED (desk, not independent): 326 TWh ÷ 6.5×10¹⁷ ≈ 5.0×10⁻⁴ Wh/token, matching the bottom-up web-literature per-token energy figure — a units consistency check, not two measurements agreeing. Evidence: derived arithmetic over the paper's own figures; the agreement is definitional, likely because the per-token rate is the paper's input
Benchmarks & Data
All figures are the paper's own order-of-magnitude estimates at "current efficiency," as of the 2026-02-23 preprint. None is a measured per-SKU cost.
| Quantity | Value | As of | Nature | Source |
|---|---|---|---|---|
| Projected 2028 US AI energy allocation | 326 TWh | 2028 (projected) | projection (paper input) | Photons = Tokens |
| Tokens/year that supports (current efficiency) | ≈ 6.5×10¹⁷ | 2026-02-23 | modelled (order-of-magnitude) | Photons = Tokens |
| Per-capita token budget | ≈ 225,000 tokens/person/day | 2026-02-23 | modelled (order-of-magnitude) | Photons = Tokens |
| Headroom vs mid-2024 utilization | > 3 orders of magnitude | 2026-02-23 | modelled comparison | Photons = Tokens |
| Implied per-token energy (DERIVED) | ≈ 5.0×10⁻⁴ Wh/token | 2026-02-23 | derived (= paper input; not independent) | desk, over Photons = Tokens |
| $/kWh → $/token pass-through | [GAP — still none in KB] | — | — | — |
Open Questions
- What is the bottom-up $/kWh → $/token conversion for a real 2026 SKU? This page gives energy→token volume; the topic still has no measured dollar energy cost per token. The web literature has the pieces (~5×10⁻⁴ Wh/token × a $/kWh datacenter power price) but no ingested source does the multiply.
- How does the reasoning-mode 5–12× energy multiplier change the budget? The 225k/person/day figure is a single-rate average; if the token mix shifts to reasoning-heavy workloads (the growing share), the effective per-capita budget shrinks materially, and the "energy is abundant" conclusion weakens.
- Is the 326 TWh → 6.5×10¹⁷ chain robust to its efficiency assumption? The whole bridge rests on ~5×10⁻⁴ Wh/token; a 2–3× shift in serving efficiency (either direction) moves the token budget proportionally. The abstract does not disclose the efficiency-improvement path assumed to 2028.
- Does the "direction, not capacity" conclusion survive a peer-reviewed energy source? The one directly on-thesis peer-reviewed paper (Joule/Cell) is still blocked; until it is read, the abundance claim rests on a single preprint.
Related Concepts
- The Cost-Measurement Problem — this page is the energy physical unit (power → token); that page holds Matsuoka's bandwidth physical unit ($/PB, bandwidth → token). Both reject the per-token billing convention as the true cost unit.
- The Depreciation Conveyor & Vintage Economics — the value chain "photon → atom → chip → power → token" is the same stack the memory/vintage cost story runs on, read from the energy end; abundant tokens and a durable incumbent cost edge can coexist.
- The Price-Decline Distribution — the price of a token falls fast; this page puts the thermodynamic floor under how far it can fall.
- Inference Cost as a Macro Input — the ICPC prices inference cost into the macro price level via a composite that already includes an electricity term; this page is the physical grounding for that electricity term.
Backlinks
Pages that reference this concept:
Changelog
- 2026-07-24 — Created from the newly compiled Photons = Tokens preprint (arXiv 2603.06630, abstract-grade). Fills the frontier's "no energy pass-through leg" gap partly: supplies a top-down energy→token volume bridge (326 TWh → ≈6.5×10¹⁷ tokens/yr → ≈225,000 tokens/person/day, >3 OOM above mid-2024), and the contrarian moral that energy is not the near-term binding constraint on token volume — direction is. Records what it does NOT close: a bottom-up $/kWh → $/token per-SKU cost, and the reasoning-mode 5–12× multiplier. Single preprint, not peer-reviewed, order-of-magnitude policy-framing figures throughout.