Independent Inference Providers
organizationIndependent Inference Providers
Type: Market segment (a cohort of companies, not one entity)
The Together / Fireworks / Anyscale / Groq / Cerebras / Replicate / OctoAI cohort — providers that serve the same open-weight models as the frontier labs, on thinner margins, with no proprietary frontier model to subsidize price. This is the layer where the commoditization thesis is most testable: if inference pricing is compressing toward marginal cost, it shows up here first, because these firms cannot cross-subsidize a token from a frontier product. By Q2 2026 the serverless-inference market had consolidated around ~7 such providers.
This page is the inference-economics lens on the cohort — pricing behaviour and unit-economics only. It is a market-structure stub, not a set of company dossiers.
Observed pricing — one model (Llama 4 70B, output), as-of Q2 2026
Digital Applied's Q2-2026 matrix (published 2026-04-24; provider pages + Artificial Analysis + the author's own April-2026 testing) — observed LIST prices, not realised/contract prices, single model:
| Provider | Tier | $/M output tokens | Notes |
|---|---|---|---|
| Together AI | Batch | $0.65 | 60-min latency, lowest priority |
| Together AI | Reserved | $0.95 | Committed capacity, 1-mo minimum |
| Fireworks AI | Serverless | $1.20 | On-demand, real-time SLA |
| OctoAI | Serverless | $1.50 | On-demand, broad model coverage |
| Anyscale Endpoints | Enterprise | $2.10 | SOC 2 / HIPAA included |
| Replicate | On-demand | $2.55 | Container-based |
| Groq | LPU | $3.20 | Specialty hardware |
| Cerebras | Wafer-scale | $4.20 | Highest tier listed |
~6x list-price spread on the same model ($0.65 → $4.20), dominated by latency tolerance, SLA, and hardware specialization — not raw per-token cost. Even at this thin-margin layer, price is still a product-differentiation lever, not yet pinned to marginal cost.
Observed throughput (output decode, tok/s)
- Groq LPU: ~750 tps (5–7x a typical H100 endpoint)
- Cerebras: ~620 tps
- Together / Fireworks (H100): ~115 tps
The load-bearing cross-check — cheapest list sits on the measured cost floor
Together's $0.65/M batch for a 70B model sits right at DigitalOcean's measured ~$0.45–$0.64/M cost-to-serve for a 70B open model at high utilization (the cost-measurement problem). So the cheapest observed list price is close to the measured cost floor — consistent with thin-to-zero margin at the batch tier, and real margin appearing only at the SLA / specialty-hardware tiers. This is the first place in the KB where an observed price and a measured cost floor can be laid side by side, and they nearly touch.
The margin mechanic these providers live or die on: "an H100 at $3.50/hr sitting unused for four hours generates $14 of pure cost with zero revenue" — idle time is the existential margin problem, the same utilization lever DigitalOcean and Patil measure.
Market context (secondary / unverified — logged, not established)
- Reported ARRs: Together ~$1B, Fireworks ~$800M (May 2026, up from ~$305M end-2025), Baseten ~$600M (reportedly raising at ~$11B). Cerebras IPO'd 2026-05-14 at a ~$66B day-one cap. Evidence: weak-moderate, secondary — not in the pricing source's table, not verified here.
- Baseten is named in this topic's gaps but absent from the Q2-2026 matrix — its pricing is not in the KB.
Key Contributions (to this topic)
- First independent-provider pricing data in the KB: a ~6x same-model list spread and the observed batch/reserved/serverless discount ladder (matrix).
- The cross-check that the cheapest list price sits on the measured open-70B cost floor → thin batch-tier margin (DigitalOcean measured cost).
Mentioned In
- The Cost-Measurement Problem — the measured cost floor this cohort's cheapest list price sits on.
- Token Pricing & the Inference-Margin Question — the layer where marginal-cost pricing would show up first.
- The Price-Decline Distribution — observed prices at the thin-margin edge of the buy-side decline.
Related Entities
- OpenAI (pricing) — the frontier-lab pricing behaviour this cohort undercuts on open models; rung-holding is a frontier act, thin-margin serving is the independent-provider act.
- vllm-cost-meter — the instrument for measuring the cost floor these providers price against.
Limitations
Single source is an analysis-grade blog (not research); list prices, not realised/contract prices; single model (Llama 4 70B); no reasoning-model or long-context tiers; Baseten absent; ARR/valuation context secondary and unverified. As-of Q2 2026 (published 2026-04-24) — prices decay fast.
Changelog
- 2026-07-23 — Created from the Digital Applied Q2-2026 pricing matrix (independent-provider observed pricing) cross-checked against DigitalOcean's measured H200 cost floor. First representation of the Together/Fireworks/Groq/Cerebras independent-provider layer in this topic — the layer where marginal-cost pricing would appear first, and where the cheapest observed list price is shown sitting on the measured cost floor.