Rung 02 Serving & RuntimeSwitch rungClose
Serving & Runtime — Timeline
Dates are publication dates of the underlying source, not ingest dates. First written
2026-09-24; raw/_sources.json holds all 18 sources.
2026
September
- Sep 24 — [Compile] The batch speculative-decoding correctness paper ingested and compiled; the latency model and the vLLM disaggregation post (ingested Sep 11) compiled. (Speculative Decoding)
August
- Aug 31 — [Preprint, v4] Correctness Forensics for Batch Speculative Decoding (first posted 2025-10-26 as "Batch Speculative Decoding Done Right"), accepted to Findings of EMNLP 2026: published batch code at 0.0–3.5% exact match; corrected versions 90.8–97.3%. (Correctness Forensics)
July
- Jul 20 — [Preprint] C²KV: compressed, reusable KV cache blocks; up to 17× long-context speed-up. (C²KV)
- Jul 09 — [Preprint] StreamDQ: dequantization on the HBM base die (SK hynix). (StreamDQ)
- Jul 08 — [Analysis] DigitalOcean's H200 sweep: $20.32 to $0.45 per million output tokens from traffic shape alone. (H200 cost)
- Jul 06 — [Preprint] DSpark: 60–85% faster per-user generation in DeepSeek's own fleet. (DSpark)
May
- May 25 — [Report] Epoch AI: serving capacity ~3.4× a year against demand ~10×. (Epoch)
- May 14 — [Preprint] An interpretable latency model for speculative decoding: speed-ups fade with load unless acceptance is ≥ ~90%. (Latency Model)
April
- Apr 07 — [Report] vLLM / AMD MORI-IO: splitting prefill and decode takes on-time requests from 26–30 to 70–73 of 100 on one 8× MI300X node. (Next-Level Inference)
2025
October
- Oct 26 — [Preprint, v1] "Batch Speculative Decoding Done Right" first posted. (Correctness Forensics)