
In scope: batching and scheduling, KV-cache management and eviction, quantization at serving time, speculative decoding, prefill/decode disaggregation, paged attention, routing, throughput and latency engineering. Out: what the model can do (models), what a finished task costs (inference-economics), the silicon it runs on (hardware).
Analysis only1Show all →
| Type | Source | Published |
|---|---|---|
| ANALYSIS | Token Economics Across Traffic Profiles on Dedicated GPUs (measured H200 serving-cost benchmark) Vinayak Baranwal · DigitalOcean First MEASURED cost-per-token series in this KB: single H200 ($3.44/GPU-hr, Llama-3.3-70B FP8, vLLM 0.24.0), swept batch 1->256, yields $20.32 -> $0.45 per M output tokens (~44x spread on ONE SKU from traffic shape alone), independently reproducing Patil's utilization thesis with dollar-anchored levels. As-of 2026-07-08; MEASURED throughput, cost derived at list rate. | 2026-07-08 |