Skip to content
Rung 02 Serving & RuntimeSwitch rung

Serving & Runtime — Timeline

Dates are publication dates of the underlying source, not ingest dates. First written 2026-09-24; raw/_sources.json holds all 18 sources.

2026

September

  • Sep 24 — [Compile] The batch speculative-decoding correctness paper ingested and compiled; the latency model and the vLLM disaggregation post (ingested Sep 11) compiled. (Speculative Decoding)

August

  • Aug 31 — [Preprint, v4] Correctness Forensics for Batch Speculative Decoding (first posted 2025-10-26 as "Batch Speculative Decoding Done Right"), accepted to Findings of EMNLP 2026: published batch code at 0.0–3.5% exact match; corrected versions 90.8–97.3%. (Correctness Forensics)

July

  • Jul 20 — [Preprint] C²KV: compressed, reusable KV cache blocks; up to 17× long-context speed-up. (C²KV)
  • Jul 09 — [Preprint] StreamDQ: dequantization on the HBM base die (SK hynix). (StreamDQ)
  • Jul 08 — [Analysis] DigitalOcean's H200 sweep: $20.32 to $0.45 per million output tokens from traffic shape alone. (H200 cost)
  • Jul 06 — [Preprint] DSpark: 60–85% faster per-user generation in DeepSeek's own fleet. (DSpark)

May

  • May 25 — [Report] Epoch AI: serving capacity ~3.4× a year against demand ~10×. (Epoch)
  • May 14 — [Preprint] An interpretable latency model for speculative decoding: speed-ups fade with load unless acceptance is ≥ ~90%. (Latency Model)

April

  • Apr 07 — [Report] vLLM / AMD MORI-IO: splitting prefill and decode takes on-time requests from 26–30 to 70–73 of 100 on one 8× MI300X node. (Next-Level Inference)

2025

October

  • Oct 26 — [Preprint, v1] "Batch Speculative Decoding Done Right" first posted. (Correctness Forensics)