Inference Cost as a Macro Input (the Inference-Cost Phillips Curve)

Sign in to track mastery·Sign in
inference-economicsmacroeconomicsphillips-curvemonetary-policypass-throughspeculative

Inference Cost as a Macro Input (the Inference-Cost Phillips Curve)

Everything else in this topic treats inference cost as a line in an AI lab's P&L. This concept holds the one source that treats it as an input to the price level of the whole economy — and it is filed as early stage / speculative, because the lens is novel and the evidence is one unreplicated preprint.

The idea. Laitinen-Fredriksson Lundström-Imanov (arXiv 2605.20281, 2026-05-19) augments a New Keynesian Phillips curve so that firms' marginal cost of producing differentiated goods includes an AI-inference component lambda-bar, giving a structural slope kappa*_inf = lambda-bar * kappa (κ being the standard Calvo-Yun slope). The honest feature of this construction is that it makes its own stakes explicit: the macro channel is only as large as lambda-bar, the share of marginal cost that is inference. If inference is a rounding error in economy-wide marginal cost, the ICPC says so; it does not assume AI matters.

The empirical pass. Two-step GMM on US monthly data 2022:M01-2026:M04 with Newey-West HAC errors and a Hansen J-test recovers kappa-hat_inf = 0.087 (HAC s.e. 0.021), within one standard error of the structural prediction. A G7 reduced-form panel with Driscoll-Kraay errors gives 0.094 (s.e. 0.026); a Wald test fails to reject cross-country homogeneity (p = 0.78). A scaling regression across 50 rolling windows returns b-hat = 0.987 (R² = 0.998), read by the author as near-unit-elasticity pass-through.

Why this KB does not believe it yet. Three specific reasons — all sharpened, not softened, by the full-text read (2026-07-23):

  1. R² = 0.998 across 50 rolling windows is too clean for a macro relationship. Overlapping rolling windows are not independent observations, and the scaling regression regresses κ̂_inf on λ̄ — quantities that are proportional by construction under Theorem 1 (κ*_inf = λ̄·κ), so a near-unit slope at R²≈1 is close to tautological, not corroborating.
  2. The inference-cost series is a list-price proxy — now disclosed, and it is the wrong basis. The full text (Table II) builds c^inf_t from a GPU price index (Stanford AI Index 2025 + Epoch AI compute database) averaged with IEA electricity prices — an input-cost proxy, not a utilization-adjusted delivered cost. Per the cost-measurement problem, that is exactly the utilization-naive basis Patil shows mis-states true cost by up to 36×. So the macro pass-through is estimated on hardware list prices + power, not on what inference costs to deliver.
  3. lambda-bar IS reported — and it is implausibly large. The full text calibrates λ̄ = 0.18 (Table I): ~18% of firm-level marginal cost, economy-wide average, is AI inference over 2022–2026. That is not credible as a literal figure for that window, and the paper gives no calibration source for it — it is a stipulated Table I parameter. The size term that makes the whole channel economically meaningful is an assumption set at a value that flatters the result. So the magnitude can now be judged, and judging it lowers confidence rather than raising it.

Additionally: 2022-2026 is a window dominated by post-pandemic and energy inflation, and the paper does not address how an AI-specific channel is identified against that.

Why it is worth keeping anyway. If inference cost genuinely passes through to consumer prices at near-unit elasticity, then the token price index this desk maintains stops being an industry metric and becomes a macro one — and the direction of the trade changes with it. That is a large enough consequence to be worth tracking at low confidence rather than discarding. It is also the only source in this topic that asks where inference economics matters outside the AI-lab P&L.

Key Claims

  • AI inference cost can be modelled as a first-order marginal-cost input to the macroeconomy via an augmented New Keynesian Phillips curve, slope kappa*_inf = lambda-bar * kappa. Evidence: weak (single-author preprint, no peer review, no replication located) (ICPC preprint)
  • US empirical slope kappa-hat_inf = 0.087 (HAC s.e. 0.021), data window 2022:M01-2026:M04. Evidence: weak — estimated, but the underlying inference-cost series construction is undisclosed (ICPC preprint)
  • G7 panel coefficient 0.094 (s.e. 0.026); cross-country homogeneity not rejected (Wald p = 0.78). Evidence: weak (ICPC preprint)
  • Near-unit-elasticity pass-through claimed from a scaling regression (b-hat 0.987, R² 0.998 over 50 rolling windows). Evidence: weak — the fit quality is a red flag, not a strength; overlapping windows are not independent (ICPC preprint)
  • An optimal-policy coefficient under commitment psi*_inf = (1 + phi*rho) * lambda-bar * kappa and a generalized Taylor principle for the inference-augmented economy. Evidence: weak (theoretical derivation, unverified) (ICPC preprint)
  • lambda-bar is calibrated at 0.18 (Table I) — i.e. ~18% of economy-wide firm marginal cost is assumed to be AI inference over 2022–2026. Evidence: weak — calibrated with no stated source, and implausibly large for the window; the entire channel's economic magnitude rests on it (ICPC preprint)

Benchmarks & Data

QuantityValueAs ofNature
US slope kappa-hat_inf0.087 (HAC s.e. 0.021)data to 2026-04; preprint 2026-05-19estimated
G7 panel b-hat^G70.094 (s.e. 0.026)preprint 2026-05-19estimated
Scaling regression b-hat0.987 (R² 0.998, 50 windows)preprint 2026-05-19estimated
Cross-country homogeneityWald p = 0.78preprint 2026-05-19test statistic
lambda-bar (inference share of marginal cost)0.18Table I (calibrated)CALIBRATED, no source given — implausibly large
η_inf (inflation-variance share from inference shocks)0.27 (HAC s.e. 0.06)Table IIIestimated → 0.18-0.41pp of headline inflation

Open Questions

  • What is lambda-bar? RESOLVED 2026-07-23: λ̄ = 0.18 (Table I, calibrated). The channel now has a size — but 0.18 is implausibly large and unsourced, which is itself the finding. New question: on what basis was 0.18 chosen? The paper does not say.
  • How was the AI-inference cost series constructed? RESOLVED: a GPU-list-price index (Stanford AI Index 2025 + Epoch AI) averaged with IEA electricity — an input-cost proxy, utilization-naive by construction (Table II). This does not remove the concern; it confirms it.
  • Does anyone replicate this? No replication, citation, or critical response has been located. A single preprint proposing a new macro curve is a hypothesis.
  • Is the 2022-2026 window identifiable for an AI-specific channel, given post-pandemic and energy inflation dominate it?

Related Concepts

Backlinks

Pages that reference this concept:

Changelog

  • 2026-07-23 — Full-text touch-up (complete 6-page PDF read). Load-bearing detail FOUND: λ̄ = 0.18 (Table I, calibrated, no source, implausibly large) — the channel's size term, previously "unreported." Also resolved the cost-series question: it is a GPU-list-price + IEA-electricity composite (utilization-naive). Both findings lower confidence. Affiliation corrected (Stockholm University). No change to the early-stage/speculative status.
  • 2026-07-22 — Created from the ICPC preprint (arXiv 2605.20281), ingested this compile. Filed as early-stage/speculative with three named reasons for low confidence (implausible R², undisclosed cost series, unreported lambda-bar). First macro-lens concept in this topic.
Inference Cost as a Macro Input (the Inference-Cost Phillips Curve) | KB | MenFem