Skip to content
Rung 07 Compilers & KernelsSwitch rung

Compilers & Kernels — Timeline

Dates are publication dates of the underlying source unless marked as an ingest or compile. Kernel baselines age in weeks, so every performance figure carries its as-of date.

2026

September

  • Sep 24 — [Ingest + Compile] The batch-1 decode study and Meta's Triton autoWS design note ingested and compiled: 5 sources, 5 concepts, 4 entities. TritonForge skipped. (Index)
  • Sep 24 — [Compile] First compile of the rung: 3 sources, 4 concepts, 3 entities. The FlashMLA source, ingested in August but missing from the registry, was registered and compiled. (Index)
  • Sep 11 — [Ingest] FlashAttention-4 and ReQAT ingested — the rung's first registered sources.

August

  • Aug 21 — [Ingest] DeepSeek's FlashMLA FP8 deep dive close-read, with the repository read at HEAD 15f13e50 (2026-07-27). (Source)

June

  • Jun 14 — [Conference] ReQAT (ICML 2026): FP4 reasoning failure concentrates on low-entropy tokens; W4A4KV4 NVFP4 matches full-precision fine-tuning on R1-Qwen-14B (65.94% vs 65.46%, AIME-120) but trails on R1-Llama-8B; 3.9× throughput on DGX Spark, 3.1× on B200. (Source · Concept)

May

  • May 28 — [Preprint] "Memory-Bound but Not Bandwidth-Limited": at batch 1, H100 reaches ~27% of its memory-bandwidth limit against ~81% on L4; CUDA Graphs 1.259× on H100 (N=10, 95% CI 1.253–1.267), 1.028× on L4; on L4, int4 via ExLlamaV2 17.36 ms vs AutoAWQ+Marlin 45.24 ms per token. (Source · Concept)

March

  • Mar 5 — [Preprint] FlashAttention-4: on Blackwell, non-matmul work exceeds matmul work by 25–60% in attention; 1,613 TFLOPs/s at 71% utilisation on B200 BF16, up to 1.3× cuDNN 9.13 and 2.7× Triton; 20–30× faster compiles in CuTe-DSL. (Source · Concept)

January

  • Jan 8 — [Technical report] Meta's Triton team publishes the design and roadmap for automatic warp specialization (autoWS): 1.5–2× over stock Triton on B200 flash-attention forward, cuDNN still 10–20% ahead (vendor-reported). (Source · Concept)

2025

October

  • Oct 1 — [Technical report] FlashMLA README figures set: sparse FP8 decode 410 TFLOPS on H800 SXM5, up to 350 TFLOPS on B200 ("not really optimized yet"). Not restated since. (Source · Concept)

September

  • Sep 29 — [Technical report] DeepSeek publishes the FlashMLA FP8 sparse decoding deep dive for DeepSeek-V3.2 (context doubled to 128K): dequantization, not the tensor cores, bounds decode on Hopper (≥50 vs 34 cycles per KV token); "crossover" lifts decode 250 → 410 TFLOPS (1.64×). (Source · Concept)

April

  • Apr 22 — [Technical report] DeepSeek's companion FlashMLA dense-decode deep dive supplies the roofline the September piece builds on: decode is compute-bound on H800 once heads × query length reaches ~128; up to 80% tensor-core utilisation and 3 TB/s. Read as part of the September source. (Source)