The Economics of AI Inference: Inflation Dynamics, Welfare Costs, and Optimal Monetary Policy under the Inference-Cost Phillips Curve
First attempt to model AI inference cost as a first-order marginal-cost input to the macroeconomy: an augmented New Keynesian 'Inference-Cost Phillips Curve' (ICPC), estimated via GMM against 2022-2026 US + G7 panel data (slope coefficient 0.087, near-unit-elasticity pass-through). Speculative but a genuinely novel lens for where token-pricing economics could matter beyond the AI-lab P&L question this topic otherwise tracks.
The Economics of AI Inference: Inflation Dynamics, Welfare Costs, and Optimal Monetary Policy under the Inference-Cost Phillips Curve
Read discipline — debt cleared. The prior version of this file was built from the arXiv abstract only and flagged
lambda-bar(the economy-wide AI-inference share of marginal cost — the size term, without which the slope has no magnitude) as unreported. It IS reported in the full text:lambda-bar = 0.18(Table I, calibrated on US monthly data 2022:M01–2026:M04). So the economic magnitude of the channel can now be judged — and 0.18 is a suspiciously large figure (see critique). Full 6-page PDF read directly 2026-07-23.
Abstract
"The operating cost of large language model (LLM) inference, dominated by GPU compute and electricity, has become a first-order input price across a rapidly expanding set of consumer-facing services. We develop a New Keynesian framework augmented with an AI inference cost wedge and derive the Inference-Cost Phillips Curve (ICPC), a closed-form modification of the Phillips curve in which the slope on the output gap, kappa, and the inference pass-through coefficient, kappa_inf, are explicit functions of the cross-sectional AI intensity distribution and the Calvo stickiness parameter theta. We prove existence and generic uniqueness of the ICPC; show that algorithmic dynamic pricing phirho attenuates the demand slope by a factor (1 - phirho) and amplifies the inference pass-through by a factor (1 + phirho); and establish a welfare decomposition, a mean-field inflation limit, an impossibility of information-constrained implementation, a sqrt(T)-consistency result for the two-step GMM estimator, and a Lucas-style closed-form welfare cost of inference-induced inflation. We calibrate the model to U.S. monthly data on cloud GPU compute prices, electricity prices, and core CPI between 2022:M01 and 2026:M04 and estimate kappa-hat_inf = 0.087 (HAC s.e. 0.021), implying that AI inference cost shocks account for between 0.18 and 0.41 percentage points of headline inflation over the sample. A near-linear scaling regression log10 kappa-hat_inf = a + b log10 lambda-bar yields b-hat = 0.987 with R^2 = 0.998. A reduced-form G7 monthly panel for 2022:M01-2026:M04 delivers a within-group estimate b-hat^G7 = 0.094 (Driscoll-Kraay HAC s.e. 0.026) with R^2_within = 0.927, indistinguishable from the U.S. baseline. The framework rationalizes a compute-price indexing component in the Taylor rule with response volatility Delta C = (1/2) gamma (kappa^ALG_inf)^2 sigma^2_inf/(1-betatheta)^2 and an inference-adjusted optimal inflation target pit = -lambda-bar kappa E_t[c^inf{t+1}]/(1 - beta*theta)."
THE LOAD-BEARING DETAIL — lambda-bar (FOUND)
lambda-bar = 0.18 — Table I ("Calibrated parameters for the ICPC, U.S. monthly, 2022:M01–2026:M04, accessed 2026-05-18"), labelled "λ̄ (AI int.)", used in Eq. (3).
lambda-bar is the economy-wide average AI-inference share of firm-level marginal cost: from Eq. (1), firm i's real marginal cost is mc_{i,t} = w_t − a_t + λ_i c^inf_t, with λ_i ∈ [0,1] the firm's AI intensity; Assumption 1 gives lambda-bar = ∫₀¹ λ dF_λ(λ). It is the multiplier that turns the standard Calvo-Yun slope κ into the inference pass-through: kappa*_inf = lambda-bar · kappa (Theorem 1, Eq. 3).
Why this matters and why 0.18 is a red flag. The whole macro channel scales with lambda-bar. At 0.18 the paper is asserting that, on average across the economy over 2022–2026, ~18% of firm-level marginal cost is AI inference. That is implausibly large as a literal economy-wide figure for that window (AI inference was a rounding error in most firms' cost of goods through 2024–25). The paper offers no calibration source or derivation for the 0.18 — it appears only as a Table I parameter. Table II's data sources cover the inference-cost series and CPI/output-gap inputs, not the share. So the size term that makes the entire result economically meaningful is a stipulated number, not an estimated one, and it is set at a value that flatters the channel. Treat 0.18 as an assumption doing very heavy lifting.
Full Table I (calibrated parameters)
| Parameter | Value | Use |
|---|---|---|
| θ (Calvo stickiness) | 0.75 | Eq. 3 |
| β (discount) | 0.996 | Eq. 3 |
| λ̄ (AI intensity) | 0.18 | Eq. 3 |
| φ (algorithmic penetration) | 0.32 | Def. 1 |
| ρ (collusion / near-collusive responsiveness) | 0.20 | Def. 1 |
| ω (welfare loss weight) | 0.50 | Eq. 6 |
Note φ·ρ = 0.064 as calibrated (the "algorithmic pricing intensity" that attenuates the demand slope by (1−φρ) and amplifies inference pass-through by (1+φρ), Theorem 2).
Key Contributions (the paper lists NINE)
- Inference cost as a macro input. Firm marginal cost augmented with an AI-inference share
lambda-bar, yielding structural slopekappa*_inf = lambda-bar · kappa. - Algorithmic attenuation/amplification (Theorem 2):
kappa^ALG = (1−φρ)κ,kappa^ALG_inf = (1+φρ)κ*_inf— algorithmic dynamic pricing dampens the demand-side slope and sharpens the cost pass-through. - A policy rule: optimal response coefficient under commitment
psi*_inf = (1 + φρ) · lambda-bar · κ(Corollary 2), plus a generalized Taylor principle and an inference-adjusted optimal inflation targetpi*_t = −lambda-bar κ E_t[c^inf_{t+1}]/(1−βθ)(Corollary 3). - A welfare decomposition (Hicks-Kaldor / Theorem 3), a mean-field inflation limit via a Fokker-Planck representation (Theorem 4), and an impossibility result (Theorem 5): no incentive-compatible mechanism using only firm-level observations can implement the planner-optimal response when AI intensity is private.
- A Lucas-style closed-form welfare cost
Delta C* = (1/2) γ (kappa^ALG_inf)² σ²_inf/(1−βθ)²(Theorem 7). - Empirical: two-step GMM (
sqrt(T)-consistency, Theorem 6) on US data + a G7 reduced-form panel.
Results (ESTIMATED — econometric, data window 2022:M01 – 2026:M04, accessed 2026-05-18)
Table III — two-step GMM estimates (US)
| Coefficient | Estimate | HAC s.e. |
|---|---|---|
| κ | 0.041 | 0.012 |
| κ_inf | 0.087 | 0.021 |
| φρ | 0.064 | 0.018 |
| η_inf (inflation-variance share attributable to inference shocks) | 0.27 | 0.06 |
- η_inf = 0.27 (Proposition 1 upper bound): ~27% of unconditional inflation variance attributable to inference-cost shocks. Combined with κ_inf = 0.087, the paper reads this as 0.18–0.41 pp of headline inflation over the sample.
- Internal-consistency note: κ_inf = 0.087 is NOT simply λ̄·κ with the reported κ = 0.041 (that product is ~0.0074). The reported κ = 0.041 is the algorithmic-attenuated slope
κ^ALG, and the closed-form κ*_inf = λ̄κ refers to the pre-attenuation structural κ; the mapping between the calibrated λ̄=0.18 and the estimated κ_inf=0.087 runs through Theorem 2's (1+φρ) amplification and the underlying (pre-attenuation) κ, not the attenuated 0.041. The paper does not print the pre-attenuation κ explicitly.
Scaling and panel
| Quantity | Estimate | Note |
|---|---|---|
Scaling regression b-hat (Table IV) | 0.987 (R² = 0.998, a = −2.41) | log10 κ̂_inf on log10 λ̄ over 50 resampled subwindows; near-unit elasticity |
G7 panel b-hat^G7 (Table V) | 0.094 (Driscoll-Kraay HAC s.e. 0.026) | N=7, T×N = 52×7, R²_within = 0.927; ξ^G7 = 0.038 (0.014) |
| Cross-country homogeneity | b̂^G7 ≈ κ̂_inf, "statistically indistinguishable" | G7 = Canada, France, Germany, Italy, Japan, UK, US |
(The Wald p = 0.78 for cross-country homogeneity quoted in the previous abstract-derived version is consistent with this; the PDF states the two estimates are statistically indistinguishable.)
How the AI-inference cost series was constructed (previously an open question — NOW ANSWERED)
Table II data sources (all accessed 2026-05-18): Headline CPI, Core CPI (BEA [26] / FRED [27]), Output gap (CBO [28], quarterly → monthly), a GPU price index, electricity (IEA [29]), and expected inflation π^e_{t+1}. The compute-price inputs are constructed from the Stanford AI Index 2025 [30] and the Epoch AI compute database [31]. For the G7 panel the country inference-cost index c^inf_{j,t} is "the simple average of the standardized GPU compute price and electricity price for country j in month t."
Critical read (strengthens, does not remove, the concern). The inference-cost series is a GPU-list-price + electricity composite — an input-cost proxy, NOT a utilization-adjusted delivered inference cost. It therefore inherits exactly the defect Patil (2606.11690) documents: list-price-based cost is utilization-naive, and true effective cost moves up to 36× on identical hardware with load. A macro pass-through estimated on a list-price composite measures the pass-through of hardware list prices + power, not of what inference actually costs to deliver. This is the load-bearing measurement, and it is a proxy.
Limitations / critique
- λ̄ = 0.18 is stipulated, not estimated, and is implausibly large as an economy-wide 2022–2026 average AI-inference cost share. The economic magnitude of the entire channel rests on it.
- R² = 0.998 over 50 rolling windows is too clean for a macro relationship. Overlapping rolling windows are not independent observations; a near-perfect scaling fit is the signature of a mechanically-constructed regressor (κ̂_inf regressed on λ̄, which by Theorem 1 are proportional by construction — so the scaling regression is close to tautological) as much as of a real elasticity.
- The cost series is a list-price/electricity proxy (see above) — utilization-naive by construction.
- 2022–2026 is dominated by post-pandemic and energy inflation; identification of an AI-specific channel across that window is not addressed.
- Single-author preprint, no peer review, no located replication. Affiliation IS stated (Stockholm University) — the prior "no affiliation" note was an abstract-page artifact and is corrected.
Source: The Economics of AI Inference: Inflation Dynamics, Welfare Costs, and Optimal Monetary Policy under the Inference-Cost Phillips Curve by Gustav Olaf Yunus Laitinen-Fredriksson Lundström-Imanov (Department of Economics, Stockholm University), arXiv:2605.20281 [econ.GN], submitted 2026-05-19. Full 6-page PDF read directly 2026-07-23.