Rung 07 Compilers & KernelsSwitch rungClose
Compilers & Kernels — Timeline
Dates are publication dates of the underlying source unless marked as an ingest or compile. Kernel baselines age in weeks, so every performance figure carries its as-of date.
2026
September
- Sep 24 — [Ingest + Compile] The batch-1 decode study and Meta's Triton autoWS design note ingested and compiled: 5 sources, 5 concepts, 4 entities. TritonForge skipped. (Index)
- Sep 24 — [Compile] First compile of the rung: 3 sources, 4 concepts, 3 entities. The FlashMLA source, ingested in August but missing from the registry, was registered and compiled. (Index)
- Sep 11 — [Ingest] FlashAttention-4 and ReQAT ingested — the rung's first registered sources.
August
- Aug 21 — [Ingest] DeepSeek's FlashMLA FP8 deep dive close-read, with the repository read at
HEAD
15f13e50(2026-07-27). (Source)
June
- Jun 14 — [Conference] ReQAT (ICML 2026): FP4 reasoning failure concentrates on low-entropy tokens; W4A4KV4 NVFP4 matches full-precision fine-tuning on R1-Qwen-14B (65.94% vs 65.46%, AIME-120) but trails on R1-Llama-8B; 3.9× throughput on DGX Spark, 3.1× on B200. (Source · Concept)
May
- May 28 — [Preprint] "Memory-Bound but Not Bandwidth-Limited": at batch 1, H100 reaches ~27% of its memory-bandwidth limit against ~81% on L4; CUDA Graphs 1.259× on H100 (N=10, 95% CI 1.253–1.267), 1.028× on L4; on L4, int4 via ExLlamaV2 17.36 ms vs AutoAWQ+Marlin 45.24 ms per token. (Source · Concept)
March
- Mar 5 — [Preprint] FlashAttention-4: on Blackwell, non-matmul work exceeds matmul work by 25–60% in attention; 1,613 TFLOPs/s at 71% utilisation on B200 BF16, up to 1.3× cuDNN 9.13 and 2.7× Triton; 20–30× faster compiles in CuTe-DSL. (Source · Concept)
January
- Jan 8 — [Technical report] Meta's Triton team publishes the design and roadmap for automatic warp specialization (autoWS): 1.5–2× over stock Triton on B200 flash-attention forward, cuDNN still 10–20% ahead (vendor-reported). (Source · Concept)
2025
October
- Oct 1 — [Technical report] FlashMLA README figures set: sparse FP8 decode 410 TFLOPS on H800 SXM5, up to 350 TFLOPS on B200 ("not really optimized yet"). Not restated since. (Source · Concept)
September
- Sep 29 — [Technical report] DeepSeek publishes the FlashMLA FP8 sparse decoding deep dive for DeepSeek-V3.2 (context doubled to 128K): dequantization, not the tensor cores, bounds decode on Hopper (≥50 vs 34 cycles per KV token); "crossover" lifts decode 250 → 410 TFLOPS (1.64×). (Source · Concept)
April
- Apr 22 — [Technical report] DeepSeek's companion FlashMLA dense-decode deep dive supplies the roofline the September piece builds on: decode is compute-bound on H800 once heads × query length reaches ~128; up to 80% tensor-core utilisation and 3 TB/s. Read as part of the September source. (Source)