06MODEL GRAPH TO SILICON· STEADY

Compilers & Kernels

The layer between a model graph and the silicon. In scope: CUDA and Triton, kernel fusion and autotuning, compiler stacks (XLA, TVM, Inductor), custom kernels for attention and quantization, portability across accelerators. Out: the serving system above it (serving), the chip below it (hardware).

0SOURCES
0CONCEPTS
0ENTITIES
SOURCE MIX
0 P0 R0 A0 N
ACTIVITY · 20W
CUDATritonKernel fusionPortability