AI Inference Providers Compared: Q2 2026 Pricing Matrix (independent-provider observed pricing)
First independent-provider pricing in the KB: on ONE model (Llama 4 70B output), observed Q2-2026 LIST prices spread 6x — Together $0.65 (batch) to Cerebras $4.20 (wafer-scale); throughput spread 5-7x (Groq 750 tps, Cerebras 620 vs H100 ~115). Cheapest list ($0.65/M batch) sits at DigitalOcean's measured ~$0.45-0.64/M cost floor => thin-to-zero margin at batch tier. Opens the layer where marginal-cost pricing shows up first.
AI Inference Providers Compared: Q2 2026 Pricing Matrix
What this is
The KB's first look at the independent inference-provider layer — the Together / Fireworks / Groq / Cerebras cohort that serves the same open-weight models on thinner margins than frontier labs. This layer is where the commoditization thesis is most testable: if inference pricing is compressing toward marginal cost, it shows up here first, because these providers have no proprietary frontier model to subsidize price. Observed Q2-2026 list prices (published 2026-04-24), from provider pages + Artificial Analysis + the author's own April-2026 testing.
Observed pricing — same model (Llama 4 70B, output), as-of Q2 2026
| Provider | Tier | $/M output tokens | Notes |
|---|---|---|---|
| Together AI | Batch | $0.65 | 60-min latency, lowest priority |
| Together AI | Reserved | $0.95 | Committed capacity, 1-mo minimum |
| Fireworks AI | Serverless | $1.20 | On-demand, real-time SLA |
| OctoAI | Serverless | $1.50 | On-demand, broad model coverage |
| Anyscale Endpoints | Enterprise | $2.10 | SOC 2 / HIPAA included |
| Replicate | On-demand | $2.55 | Container-based |
| Groq | LPU | $3.20 | Specialty hardware |
| Cerebras | Wafer-scale | $4.20 | Highest tier listed |
Price spread: ~6x on the same model ($0.65 → $4.20). The article's framing: "Same model, 6× pricing spread — pick by total cost-of-answer" — pricing reflects latency tolerance, SLA, and hardware specialization, not pure per-token cost.
Observed throughput (output decode, tok/s)
- Groq LPU: ~750 tps (5–7x a typical H100 endpoint)
- Cerebras: ~620 tps
- Together / Fireworks (H100): ~115 tps
Market context (from the surrounding independent-provider literature, logged with this source)
- The serverless-inference market consolidated around ~7 providers by Q2 2026 (Together, Fireworks, Anyscale, Groq, Cerebras, Replicate, OctoAI).
- Reported ARRs (secondary, not in this source's table): Together ~$1B, Fireworks ~$800M (May 2026, up from ~$305M end-2025), Baseten ~$600M (reportedly raising at ~$11B); Cerebras IPO'd 2026-05-14 at a ~$66B day-one cap. Evidence: weak-moderate, secondary — logged for context, not verified here.
- The margin mechanic these providers live or die on: "an H100 at $3.50/hr sitting unused for four hours generates $14 of pure cost with zero revenue" — idle time is the existential margin problem, the same utilization lever DigitalOcean and Patil measure.
Why it matters here
- Closes the independent-provider gap the topic flagged as the place "marginal-cost pricing would show up first." The 6x same-model spread is dominated by latency/SLA/hardware tier, not raw token cost — evidence that even at this thin-margin layer, price is still a product-differentiation lever, not yet pinned to marginal cost.
- Anchors the batch/reserved discount structure the KB flagged as "list vs realised price unmeasured": Together's own batch ($0.65) vs reserved ($0.95) vs Fireworks serverless ($1.20) shows the within-market discount ladder for latency-tolerant work — a partial, observed look at the list-vs-realised gap.
- Cross-checks the measured cost floor. Together's $0.65/M batch for a 70B model sits right at DigitalOcean's measured ~$0.45–$0.64/M cost-to-serve for a 70B model at high utilization — i.e. the cheapest observed list price is close to the measured cost floor, consistent with thin-to-zero margin at the batch tier and real margin only at the SLA/specialty tiers.
Limitations
Analysis-grade blog, not research. List prices, not realised/contract prices. Single model (Llama 4 70B). Baseten named in the KB gap but absent here. ARR/valuation context is secondary and unverified. Reasoning-model and long-context pricing not covered.
Source: AI Inference Providers Compared: Q2 2026 Pricing Matrix, Digital Applied, 2026-04-24