Independent Inference Providers

organization
independent-providersinference-economicsobserved-pricingprice-spreadmarginal-cost

Independent Inference Providers

Type: Market segment (a cohort of companies, not one entity)

The Together / Fireworks / Anyscale / Groq / Cerebras / Replicate / OctoAI cohort — providers that serve the same open-weight models as the frontier labs, on thinner margins, with no proprietary frontier model to subsidize price. This is the layer where the commoditization thesis is most testable: if inference pricing is compressing toward marginal cost, it shows up here first, because these firms cannot cross-subsidize a token from a frontier product. By Q2 2026 the serverless-inference market had consolidated around ~7 such providers.

This page is the inference-economics lens on the cohort — pricing behaviour and unit-economics only. It is a market-structure stub, not a set of company dossiers.

Observed pricing — one model (Llama 4 70B, output), as-of Q2 2026

Digital Applied's Q2-2026 matrix (published 2026-04-24; provider pages + Artificial Analysis + the author's own April-2026 testing) — observed LIST prices, not realised/contract prices, single model:

ProviderTier$/M output tokensNotes
Together AIBatch$0.6560-min latency, lowest priority
Together AIReserved$0.95Committed capacity, 1-mo minimum
Fireworks AIServerless$1.20On-demand, real-time SLA
OctoAIServerless$1.50On-demand, broad model coverage
Anyscale EndpointsEnterprise$2.10SOC 2 / HIPAA included
ReplicateOn-demand$2.55Container-based
GroqLPU$3.20Specialty hardware
CerebrasWafer-scale$4.20Highest tier listed

~6x list-price spread on the same model ($0.65 → $4.20), dominated by latency tolerance, SLA, and hardware specialization — not raw per-token cost. Even at this thin-margin layer, price is still a product-differentiation lever, not yet pinned to marginal cost.

Observed throughput (output decode, tok/s)

  • Groq LPU: ~750 tps (5–7x a typical H100 endpoint)
  • Cerebras: ~620 tps
  • Together / Fireworks (H100): ~115 tps

The load-bearing cross-check — cheapest list sits on the measured cost floor

Together's $0.65/M batch for a 70B model sits right at DigitalOcean's measured ~$0.45–$0.64/M cost-to-serve for a 70B open model at high utilization (the cost-measurement problem). So the cheapest observed list price is close to the measured cost floor — consistent with thin-to-zero margin at the batch tier, and real margin appearing only at the SLA / specialty-hardware tiers. This is the first place in the KB where an observed price and a measured cost floor can be laid side by side, and they nearly touch.

The margin mechanic these providers live or die on: "an H100 at $3.50/hr sitting unused for four hours generates $14 of pure cost with zero revenue" — idle time is the existential margin problem, the same utilization lever DigitalOcean and Patil measure.

Market context (secondary / unverified — logged, not established)

  • Reported ARRs: Together ~$1B, Fireworks ~$800M (May 2026, up from ~$305M end-2025), Baseten ~$600M (reportedly raising at ~$11B). Cerebras IPO'd 2026-05-14 at a ~$66B day-one cap. Evidence: weak-moderate, secondary — not in the pricing source's table, not verified here.
  • Baseten is named in this topic's gaps but absent from the Q2-2026 matrix — its pricing is not in the KB.

Key Contributions (to this topic)

  • First independent-provider pricing data in the KB: a ~6x same-model list spread and the observed batch/reserved/serverless discount ladder (matrix).
  • The cross-check that the cheapest list price sits on the measured open-70B cost floor → thin batch-tier margin (DigitalOcean measured cost).

Mentioned In

Related Entities

  • OpenAI (pricing) — the frontier-lab pricing behaviour this cohort undercuts on open models; rung-holding is a frontier act, thin-margin serving is the independent-provider act.
  • vllm-cost-meter — the instrument for measuring the cost floor these providers price against.

Limitations

Single source is an analysis-grade blog (not research); list prices, not realised/contract prices; single model (Llama 4 70B); no reasoning-model or long-context tiers; Baseten absent; ARR/valuation context secondary and unverified. As-of Q2 2026 (published 2026-04-24) — prices decay fast.

Changelog

  • 2026-07-23 — Created from the Digital Applied Q2-2026 pricing matrix (independent-provider observed pricing) cross-checked against DigitalOcean's measured H200 cost floor. First representation of the Together/Fireworks/Groq/Cerebras independent-provider layer in this topic — the layer where marginal-cost pricing would appear first, and where the cheapest observed list price is shown sitting on the measured cost floor.