Skip to content
Rung 07 Compilers & KernelsSwitch rung
Rung 07 / Model graph to silicon

Compilers & Kernels

The layer between a model graph and the silicon.

5
Sources
5
Concepts
4
Entities
Paper2Report3AnalysisNews
A steel press die with a copper blank beneath it, on paper.

In scope: CUDA and Triton, kernel fusion and autotuning, compiler stacks (XLA, TVM, Inductor), custom kernels for attention and quantization, portability across accelerators. Out: the serving system above it (serving), the chip below it (hardware).

Five loose operators entering a taper and leaving as one fused copper solid.
CUDATritonKernel fusionPortability
Sources compiled for this topic
TypeSourcePublished
PAPERFlashAttention-4: Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling
Ted Zadouri, Markus Hoehnerbach, Jay Shah, Timmy Liu, Vijay Thakkar, Tri Dao · Not stated on the arXiv abstract page (FlashAttention lineage)

Names which bound moved: on Blackwell, shared-memory traffic and exponential operations exceed MMA compute by 25-60% for typical attention workloads — the tensor cores are no longer the constraint. Reaches 1,613 TFLOPs/s at 71% utilisation on B200 BF16 (1.3x cuDNN 9.13, 2.7x Triton); CuTe-DSL compiles 20-30x faster than C++ templates (2.5s vs 55s fwd). Deterministic backward costs up to 25% of speed.

2026-03-05
PAPERMemory-Bound but Not Bandwidth-Limited: The Physical AI Inference Gap in Batch-1 LLM Decode
Josef Chen · KAIKAKU

At batch 1, three 7-8B models across four NVIDIA GPUs (44 cells): H100 reaches ~27% of its memory-bandwidth floor vs ~81% on L4. Launch overhead is the gap — CUDA Graphs 1.259x on H100 (95% CI 1.253-1.267, N=10, pre-registered thresholds) vs 1.028x on L4. On L4 the int4 KERNEL is the lever: ExLlamaV2 17.36 ms vs AutoAWQ+Marlin 45.24 ms vs bf16 62.32 ms per token. H100+Graphs 11.78 ms vs L4+ExLlamaV2 17.36 ms at $3.50 vs $0.30/hr (May 2026): paper says ~6x cheaper per token on L4; its own inputs give ~7.9x (~$11.45 vs ~$1.45 per Mtok, KB arithmetic). Preprint, single author, no Blackwell/AMD/Apple.

2026-05-28
REPORTA Deep Dive Into The Flash MLA FP8 Decoding Kernel on Hopper
Shengyu Liu (git author of both commits to this document); FlashMLA authors of record per the repository citation block: Jiashi Li, Shengyu Liu · DeepSeek (deepseek-ai)

An FP8-KV-cache attention decoding kernel on Hopper is bound not by tensor-core throughput but by dequantization on the CUDA cores: 50 cycles of format conversion per KV token against 34 cycles of MMA. "Crossover" — two CTAs in a cluster of 2 each dequantize half the KV block and share it over Distributed Shared Memory — lifts measured decode throughput 250 -> 410 TFLOPS on H800 SXM5 (1.64x, as of 2025-09-29) with silicon, algorithm and numerics unchanged. Vendor self-measurement; mechanism verified in shipped source. Same kernel scores 350 TFLOPS on B200, below the older H800. Engineering blog in the repository, filed as technical-report (company research blog). No accuracy evaluation of the FP8 cache format.

2025-09-29
REPORTWarp Specialization in Triton: Design and Roadmap
Manman Ren, Nick Riasanovsky, Neil Dhar, Hongtao Yu, Jie Liu, Partha Kanuparthy, Shane Nay · Meta (Triton autoWS team), published on the PyTorch blog

VENDOR-AUTHORED (the team shipping the feature). Design of automatic warp specialization in the Triton compiler — data partitioning, loop scheduler, partition scheduler, buffer creation, memory planner, channel/code partitioner. On B200 flash-attention forward: 1.5-2x over stock Triton, near Gluon/cuDNN, cuDNN still ahead by 10-20% (one sentence, no configurations or absolute TFLOPS). Experimental; attention forward only; Hopper and Blackwell only; AMD on the roadmap.

2026-01-08
REPORTReQAT: Achieving Full-Precision Reasoning Accuracy with 4-bit Floating-Point Quantization-Aware Training
Janghwan Lee, Sihwa Lee, Jinseok Kim, Yongjik Kim, Jieun Lim, Jinwook Oh, Jungwook Choi · Not stated on the arXiv abstract page

FP4 reasoning failure concentrates on LOW-ENTROPY tokens (digits, operators) — shown by a logit-noise experiment where perturbing only low-entropy predictions collapses accuracy and perturbing high-entropy ones barely moves it. W4A4KV4 NVFP4 + ReQAT hits 65.94% on AIME-120 vs BF16 FT 65.46% on R1-Qwen-14B — but TRAILS BF16 FT on R1-Llama-8B (41.85% vs 48.75%). 3.9x throughput on DGX Spark, 3.1x on B200. ICML 2026.

2026-06-14