In scope: CUDA and Triton, kernel fusion and autotuning, compiler stacks (XLA, TVM, Inductor), custom kernels for attention and quantization, portability across accelerators. Out: the serving system above it (serving), the chip below it (hardware).
No sources yet. This rung is open and scoped, but nothing has been read into it — no papers registered, no concepts compiled. It will fill from the top of the reading queue rather than from a summary of other people’s summaries.
← Back to the atlas