← Study
Rung 3 · Gated

Inference Engineering

Philip Kiely · Baseten · 2026freeSource ↗

Inference as a SYSTEM — serving, batching, KV cache, the production cost surface. It sits above the memory rung because KV-cache and batching arithmetic IS memory-hierarchy arithmetic: taken first, cost-per-token is a number you look up; taken here, it is a number you can derive and therefore contest.

0 / 8 chapters
Progress
Allocated
Quota
Read it
Seat
Gated
State

Second technical slot, once the memory rung closes.

Gate

Opens when the memory rung (L2a, L9, L12–L15) closes — the hardware has to be under it.

Computer Architecture

Segments

The named subset — the scope above is the count
RefTitleDoneArtifactsNote
ch1Inference·NBLABGMEBTL0/0
ch2Prerequisites·NBLABGMEBTL0/0
ch3Models·NBLABGMEBTL0/0Arithmetic intensity — pairs with Comp Arch L9.
ch4Hardware·NBLABGMEBTL0/0
ch5Software·NBLABGMEBTL0/0
ch6Techniques·NBLABGMEBTL0/0
ch7Modalities·NBLABGMEBTL0/0
ch8Production·NBLABGMEBTL0/2

Artifacts

Consume → do → output
NBNotebookLABLabGMEGameBTLBottle0/3
  • The token-price → task-price bridgeNot yet — nothing has landed for this slot.
  • The memory roofline, verifiedEarmarked for inference-benchbuilding on the workshop; not counted until it ships there.
  • ops:byte and KV-cache formulas, spacedNot yet — nothing has landed for this slot.

Feeds

What closing this sharpens

The reading

Sections of the catalogue this unit draws on
  • §3Inference Efficiency and KV Cache8 papersServing · Inference economics
  • §4Sparse Attention and Long Context9 papersServing · Models

The papers

The rows behind the sections above