ANALYSIS2026-04-24·Digital Applied

AI Inference Providers Compared: Q2 2026 Pricing Matrix (independent-provider observed pricing)

Digital Applied Team
COMPILED NOTES

First independent-provider pricing in the KB: on ONE model (Llama 4 70B output), observed Q2-2026 LIST prices spread 6x — Together $0.65 (batch) to Cerebras $4.20 (wafer-scale); throughput spread 5-7x (Groq 750 tps, Cerebras 620 vs H100 ~115). Cheapest list ($0.65/M batch) sits at DigitalOcean's measured ~$0.45-0.64/M cost floor => thin-to-zero margin at batch tier. Opens the layer where marginal-cost pricing shows up first.

AI Inference Providers Compared: Q2 2026 Pricing Matrix

What this is

The KB's first look at the independent inference-provider layer — the Together / Fireworks / Groq / Cerebras cohort that serves the same open-weight models on thinner margins than frontier labs. This layer is where the commoditization thesis is most testable: if inference pricing is compressing toward marginal cost, it shows up here first, because these providers have no proprietary frontier model to subsidize price. Observed Q2-2026 list prices (published 2026-04-24), from provider pages + Artificial Analysis + the author's own April-2026 testing.

Observed pricing — same model (Llama 4 70B, output), as-of Q2 2026

ProviderTier$/M output tokensNotes
Together AIBatch$0.6560-min latency, lowest priority
Together AIReserved$0.95Committed capacity, 1-mo minimum
Fireworks AIServerless$1.20On-demand, real-time SLA
OctoAIServerless$1.50On-demand, broad model coverage
Anyscale EndpointsEnterprise$2.10SOC 2 / HIPAA included
ReplicateOn-demand$2.55Container-based
GroqLPU$3.20Specialty hardware
CerebrasWafer-scale$4.20Highest tier listed

Price spread: ~6x on the same model ($0.65 → $4.20). The article's framing: "Same model, 6× pricing spread — pick by total cost-of-answer" — pricing reflects latency tolerance, SLA, and hardware specialization, not pure per-token cost.

Observed throughput (output decode, tok/s)

  • Groq LPU: ~750 tps (5–7x a typical H100 endpoint)
  • Cerebras: ~620 tps
  • Together / Fireworks (H100): ~115 tps

Market context (from the surrounding independent-provider literature, logged with this source)

  • The serverless-inference market consolidated around ~7 providers by Q2 2026 (Together, Fireworks, Anyscale, Groq, Cerebras, Replicate, OctoAI).
  • Reported ARRs (secondary, not in this source's table): Together ~$1B, Fireworks ~$800M (May 2026, up from ~$305M end-2025), Baseten ~$600M (reportedly raising at ~$11B); Cerebras IPO'd 2026-05-14 at a ~$66B day-one cap. Evidence: weak-moderate, secondary — logged for context, not verified here.
  • The margin mechanic these providers live or die on: "an H100 at $3.50/hr sitting unused for four hours generates $14 of pure cost with zero revenue" — idle time is the existential margin problem, the same utilization lever DigitalOcean and Patil measure.

Why it matters here

  • Closes the independent-provider gap the topic flagged as the place "marginal-cost pricing would show up first." The 6x same-model spread is dominated by latency/SLA/hardware tier, not raw token cost — evidence that even at this thin-margin layer, price is still a product-differentiation lever, not yet pinned to marginal cost.
  • Anchors the batch/reserved discount structure the KB flagged as "list vs realised price unmeasured": Together's own batch ($0.65) vs reserved ($0.95) vs Fireworks serverless ($1.20) shows the within-market discount ladder for latency-tolerant work — a partial, observed look at the list-vs-realised gap.
  • Cross-checks the measured cost floor. Together's $0.65/M batch for a 70B model sits right at DigitalOcean's measured ~$0.45–$0.64/M cost-to-serve for a 70B model at high utilization — i.e. the cheapest observed list price is close to the measured cost floor, consistent with thin-to-zero margin at the batch tier and real margin only at the SLA/specialty tiers.

Limitations

Analysis-grade blog, not research. List prices, not realised/contract prices. Single model (Llama 4 70B). Baseten named in the KB gap but absent here. ARR/valuation context is secondary and unverified. Reasoning-model and long-context pricing not covered.


Source: AI Inference Providers Compared: Q2 2026 Pricing Matrix, Digital Applied, 2026-04-24

RELATED · IN THE BASE
AI Inference Providers Compared: Q2 2026 Pricing Matrix (independent-provider observed pricing) | Knowledge Base | MenFem